OCR Awesome
约 2 分钟阅读
OCR、文档解析、版面分析、表格和公式识别资源索引。
OCR Models
| release | model | params |
|---|---|---|
| 2026-02 | GLM-OCRHuggingFaceHF:THUDM/glm-ocrhuggingface.co · huggingface.co/THUDM/glm-ocr | 0.9B |
| 2026-01 | DeepSeek-OCR-2-3BHuggingFaceHF:deepseek-ai/deepseek-ocr-2-3bhuggingface.co · huggingface.co/deepseek-ai/deepseek-ocr-2-3b | 3B |
| 2026-01 | PaddleOCR-VL-1.5HuggingFaceHF:PaddlePaddle/PaddleOCR-VL-1.5huggingface.co · huggingface.co/PaddlePaddle/PaddleOCR-VL-1.5 | 0.9B |
| 2026-01 | LightOnOCR-2-1BHuggingFaceHF:lightonai/LightOnOCR-2-1Bhuggingface.co · huggingface.co/lightonai/LightOnOCR-2-1B | 1B |
| 2025-10 | olmOCRGitHuballenai/olmocr | 7B |
| 2025-07 | dots.ocrGitHubXiaohongshu/Dots-OCR | 1.7B |
| 2024-09 | GOT-OCR2.0HuggingFaceHF:stepfun-ai/GOT-OCR2_0huggingface.co · huggingface.co/stepfun-ai/GOT-OCR2_0 | 580M |
Document Understanding
- PaddlePaddle/PaddleOCRGitHubPaddlePaddle/PaddleOCRPaddlePaddle/PaddleOCR · 支持 英文(ch)、英文(en)、法语(french)、德语(german)、韩语(korean)、日语(japan) · ocrversion: PP-OCRv5, PP-OCRv4, PP-OCRv3笔记:PaddleOCR
- RapidOCR
- Tesseract
- Surya
- breezedeus/Pix2TextGitHubbreezedeus/Pix2Textbreezedeus/Pix2Text · MIT · 国内开发者维护 · 简体中文&英文 使用的 CnOCR, 其他使用的 EasyOCR · p2t 命令行 https://pix2text.readthedocs.io/zh-cn/stable/command/笔记:Pix2Text
- mindee/doctrGitHubmindee/doctr
- Document Text Recognition;中文支持需要单独验证。
- docling-project/doclingGitHubdocling-project/docling
- PDF、文档结构和 Markdown/JSON 输出。
- ds4sd/docling-modelsHuggingFaceHF:ds4sd/docling-modelshuggingface.co · huggingface.co/ds4sd/docling-models
- SmolDocling-256M-preview 基于 SmolVLM-256M-Instruct。
- Yuliang-Liu/MonkeyOCRGitHubYuliang-Liu/MonkeyOCR
- 基于 Qwen2.5-VL-3B,支持中英文、公式和表格识别。
- SRR:Structure Detection、Content Recognition、Relationship Prediction。
- nanonets/Nanonets-OCR-sHuggingFaceHF:nanonets/Nanonets-OCR-shuggingface.co · huggingface.co/nanonets/Nanonets-OCR-s
- Image → Structure Markdown。
- ByteDance/DolphinHuggingFaceHF:ByteDance/Dolphinhuggingface.co · huggingface.co/ByteDance/Dolphin
- Document Image Parsing via Heterogeneous Anchor Prompting。
- ChatDOC/OCRFlux-3BHuggingFaceHF:ChatDOC/OCRFlux-3Bhuggingface.co · huggingface.co/ChatDOC/OCRFlux-3B
- ocrmypdf/OCRmyPDFGitHubocrmypdf/OCRmyPDFocrmypdf/OCRmyPDF · MPL-2.0, Python · adds an OCR text layer to scanned PDF · 流程 · 光栅化 · 预处理 · deskew · clean · 准备图像 · OCR 识别笔记:OCRmyPDF
- 为扫描 PDF 添加 OCR text layer。
- Topdu/OpenOCRGitHubTopdu/OpenOCR
Layout Analysis 与 PDF Toolkit
- opendatalab/OmniDocBenchGitHubopendatalab/OmniDocBench
- Marker
- rednote-hilab/dots.ocrGitHubrednote-hilab/dots.ocrrednote-hilab/dots.ocr · dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model · by 小红书 · Layout 分析效果非常好笔记:dots.ocr
- 多语言文档版面解析,1.7B,MIT。
- allenai/olmocrGitHuballenai/olmocr
- 将 PDF 线性化为适用于 LLM 数据集和训练的文本。
- opendatalab/MinerUGitHubopendatalab/MinerU
- PDF → JSON / Markdown;AGPLv3。
- opendatalab/PDF-Extract-KitGitHubopendatalab/PDF-Extract-Kit
- DocLayout-YOLO
评估关注点
- 文本识别:字符准确率、CER、语言和字体覆盖。
- 版面:标题、段落、表格、公式、图片和阅读顺序。
- 结构化输出:Markdown、JSON、HTML、坐标和 schema 稳定性。
- 复杂文档:扫描件、手写、低清晰度、多栏、旋转、印章和中英文混排。
关联信息
反向链接、本文链接的其他页面和外部资料。
反向链接
- AI Model笔记 · OCR
AI Model Awesome 总索引 · LLM 与开放模型 · 图像生成 · 视频生成 · OCR · ASR / STT · TTS · VAD · 音乐生成 · VLM / MLLM / Vision · Agent 与 Coding
- AI Model Awesome笔记 · OCR 与文档理解
按模型能力和使用场景组织的 AI 模型资源索引,保留模型架构、厂商、版本、推理资源、量化、微调和扩散模型历史资料。新增内容优先写入对应专题页面。 · LLM 与开放模型 · 图像生成 · 视频生成 · OCR 与文档理解 · ASR / STT
References
GitHub13 条
- allenai/olmocrgithub.com/allenai/olmocr
- breezedeus/Pix2Textgithub.com/breezedeus/Pix2Text
- docling-project/doclinggithub.com/docling-project/docling
- mindee/doctrgithub.com/mindee/doctr
- ocrmypdf/OCRmyPDFgithub.com/ocrmypdf/OCRmyPDF
- opendatalab/MinerUgithub.com/opendatalab/MinerU
- opendatalab/OmniDocBenchgithub.com/opendatalab/OmniDocBench
- opendatalab/PDF-Extract-Kitgithub.com/opendatalab/PDF-Extract-Kit
- PaddlePaddle/PaddleOCRgithub.com/PaddlePaddle/PaddleOCR
- rednote-hilab/dots.ocrgithub.com/rednote-hilab/dots.ocr
- Topdu/OpenOCRgithub.com/Topdu/OpenOCR
- Xiaohongshu/Dots-OCRgithub.com/Xiaohongshu/Dots-OCR
- Yuliang-Liu/MonkeyOCRgithub.com/Yuliang-Liu/MonkeyOCR
HuggingFace9 条
- HF:ByteDance/Dolphinhuggingface.co/ByteDance/Dolphin
- HF:ChatDOC/OCRFlux-3Bhuggingface.co/ChatDOC/OCRFlux-3B
- HF:deepseek-ai/deepseek-ocr-2-3bhuggingface.co/deepseek-ai/deepseek-ocr-2-3b
- HF:ds4sd/docling-modelshuggingface.co/ds4sd/docling-models
- HF:lightonai/LightOnOCR-2-1Bhuggingface.co/lightonai/LightOnOCR-2-1B
- HF:nanonets/Nanonets-OCR-shuggingface.co/nanonets/Nanonets-OCR-s
- HF:PaddlePaddle/PaddleOCR-VL-1.5huggingface.co/PaddlePaddle/PaddleOCR-VL-1.5
- HF:stepfun-ai/GOT-OCR2_0huggingface.co/stepfun-ai/GOT-OCR2_0
- HF:THUDM/glm-ocrhuggingface.co/THUDM/glm-ocr