Dify智能体平台源码二次开发笔记（6） - 优化知识库pdf文档的识别

dify的1.1.3版本知识库pdf解析实现使用pypdfium2提取文本，主要存在以下问题：1. 文本提取能力有限，对表格和图片支持不足。2. 缺乏专门的中文处理优化。3. 没有文档结构分析。4. 缺少文档质量评估。建议优化方案：1. 使用pdfplumber替代pypdfium2。2. 增加OCR支持。3. 优化中文处理逻辑。4. 添加文档结构分析。5. 实现智能表格识别。6. 增加缓存机制

已是骸骨

2940人浏览 · 2025-04-16 13:30:03

已是骸骨 · 2025-04-16 13:30:03 发布

前言

dify的1.1.3版本知识库pdf解析实现使用pypdfium2提取文本，主要存在以下问题：
1. 文本提取能力有限，对表格和图片支持不足
2. 缺乏专门的中文处理优化
3. 没有文档结构分析
4. 缺少文档质量评估
建议优化方案：
1. 使用pdfplumber替代pypdfium2
2. 增加OCR支持
3. 优化中文处理逻辑
4. 添加文档结构分析
5. 实现智能表格识别
6. 增加缓存机制
7. 优化大文件处理

导入包pdfplumber和pytesseract

pip install pdfplumber
pip install pytesseract

新增PdfNewExtractor类

新增一个PdfNewExtractor处理类替代老的PdfExtractor

from collections.abc import Iterator
from typing import Optional, cast
import pdfplumber
import pytesseract
from PIL import Image
import io

from core.rag.extractor.blob.blob import Blob
from core.rag.extractor.extractor_base import BaseExtractor
from core.rag.models.document import Document
from extensions.ext_storage import storage

class PdfNewExtractor(BaseExtractor):
    """Enhanced PDF loader with improved text extraction, OCR support, and structure analysis.

    Args:
        file_path: Path to the PDF file to load.
        file_cache_key: Optional cache key for storing extracted text.
        enable_ocr: Whether to enable OCR for text extraction from images.
    """

    def __init__(self, file_path: str, file_cache_key: Optional[str] = None, enable_ocr: bool = False):
        """Initialize with file path and optional settings."""
        self._file_path = file_path
        self._file_cache_key = file_cache_key
        self._enable_ocr = enable_ocr

    def extract(self) -> list[Document]:
        """Extract text from PDF with caching support."""
        plaintext_file_exists = False
        if self._file_cache_key:
            try:
                text = cast(bytes, storage.load(self._file_cache_key)).decode("utf-8")
                plaintext_file_exists = True
                return [Document(page_content=text)]
            except FileNotFoundError:
                pass

        documents = list(self.load())
        text_list = []
        for document in documents:
            text_list.append(document.page_content)
        text = "\n\n".join(text_list)

        # Save plaintext file for caching
        if not plaintext_file_exists and self._file_cache_key:
            storage.save(self._file_cache_key, text.encode("utf-8"))

        return documents

    def load(self) -> Iterator[Document]:
        """Lazy load PDF pages with enhanced text extraction."""
        blob = Blob.from_path(self._file_path)
        yield from self.parse(blob)

    def parse(self, blob: Blob) -> Iterator[Document]:
        """Parse PDF with enhanced features including OCR and structure analysis."""
        with blob.as_bytes_io() as file_obj:
            with pdfplumber.open(file_obj) as pdf:
                for page_number, page in enumerate(pdf.pages):
                    # Extract text with layout preservation and encoding detection
                    content = page.extract_text(layout=True)
                    # Try to detect and fix encoding issues
                    try:
                        # First try to decode as UTF-8
                        content = content.encode('utf-8').decode('utf-8')
                    except UnicodeError:
                        try:
                            # If UTF-8 fails, try GB18030 (common Chinese encoding)
                            content = content.encode('utf-8').decode('gb18030', errors='ignore')
                        except UnicodeError:
                            # If all else fails, use a more lenient approach
                            content = content.encode('utf-8', errors='ignore').decode('utf-8', errors='ignore')
                    
                    # Extract tables if present
                    tables = page.extract_tables()
                    if tables:
                        table_text = "\n\nTables:\n"
                        for table in tables:
                            # Convert table to text format
                            table_text += "\n" + "\n".join(
                                ["\t".join([str(cell) if cell else "" for cell in row]) 
                                 for row in table]
                            )
                        content += table_text

                    # Perform OCR if enabled and text content is limited or contains potential encoding issues
                    if self._enable_ocr and (len(content.strip()) < 100 or any('\ufffd' in line for line in content.splitlines())):
                        image = page.to_image()
                        img_bytes = io.BytesIO()
                        image.original.save(img_bytes, format='PNG')
                        img_bytes.seek(0)
                        pil_image = Image.open(img_bytes)
                        # Use multiple language models and improve OCR accuracy
                        ocr_text = pytesseract.image_to_string(
                            pil_image,
                            lang='chi_sim+chi_tra+eng',  # Support both simplified and traditional Chinese
                            config='--psm 3 --oem 3'  # Use more accurate OCR mode
                        )
                        if ocr_text.strip():
                            # Clean and normalize OCR text
                            ocr_text = ocr_text.replace('\x0c', '').strip()
                            content = f"{content}\n\nOCR Text:\n{ocr_text}"

                    metadata = {
                        "source": blob.source,
                        "page": page_number,
                        "has_tables": bool(tables)
                    }
                    
                    yield Document(page_content=content, metadata=metadata)

替换ExtractProcessor类

在ExtractProcessor中把两处extractor = PdfExtractor(file_path)，替换成extractor = PdfNewExtractor(file_path)。
分别在代码144行和148行

最终结果

经过测试，优化效果完美

智能体开发者社区

中国智能体开发者社区，聚焦智能体与大模型开发，提供前沿资讯、实用工具链、开源项目及行业案例。通过技术沙龙、开发者大赛等活动，促进经验交流与协作，助力开发者快速构建创新智能应用。

更多推荐

OpenClaw 本地部署完整指南（Windows + Ollama）

本文档基于实际部署经验编写，旨在帮助你在 Windows 系统上从零开始搭建 OpenClaw，并连接本地 Ollama 模型（如 Qwen2.5 或 Qwen3），使其具备完整的智能体能力。文档包含了所有关键步骤以及常见问题的解决方案。

智能体开发者社区

OpenClaw 小白安装指南（Windows版）

（类似一个能自动执行任务的AI机器人），不是游戏。API Key只保存在你本地电脑的加密文件里，不会上传到任何地方。访问：https://github.com/miaoxworld/openclaw-manager/releases。: 一键安装脚本会自动安装Node.js 22+，如果失败，手动下载安装：https://nodejs.org/：在PowerShell中，鼠标右键就是粘贴，不需要按

智能体开发者社区

飞书 × OpenClaw 接入指南：不用服务器，用长连接把机器人跑起来

这个项目存在的意义，就是把“飞书接 OpenClaw”这件事，整理成一套的配置入口，并把官方文档没覆盖到的坑集中写成排查清单。先说清楚它的角色：OpenClaw 现在已经内置官方飞书插件 @openclaw/feishu，功能更完整、维护也更及时。，说明飞书 + AI 的接入已经走通。另外，仓库也推荐了一个新项目：把 OpenClaw 变成“多 Agent 团队”，用多个 Agent 分工，Sla