All-in-One Development Tool based on PaddlePaddle
-
Updated
Jun 25, 2026 - Python
All-in-One Development Tool based on PaddlePaddle
A Unified Toolkit for Deep Learning Based Document Image Analysis
Open-source batch OCR workbench — a free, local alternative to ABBYY FineReader. Powered by Ollama + GLM-OCR + PP-DocLayoutV3, ~0.5s/page on RTX 4090. Three-panel editor, layout-aware, PDF/image batch processing, Markdown/Word export. 批量OCR工作台,纯本地运行,免费平替ABBYY,适合书籍文档数字化。
[ACL 2025 🔥] A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
中文版面检测(Chinese layout detection),yolov8 is used to detect the layout of Chinese document images。
pdfDet aims to simplify PDF layout detect tasks for users.
Docling plugin to integrate PP-DocLayout-V3 model into docling to enhance layout detection capabilities
Convert any document format into LLM-ready data format (markdown) with advanced intelligent document processing capabilities powered by pre-trained models.
A synchronous Python library that converts an academic-thesis PDF into a structured JSON document plus a tar bundle of cropped figures, tables, and formulas.
View, edit, and organize PDF documents on Windows with this fast, offline, and open-source application.
Language-agnostic OCR benchmark pipeline to discover document images, review/edit layouts in a web UI, run OCR extraction, and build high-quality evaluation datasets.
To associate your repository with the layout-detection topic, visit your repo's landing page and select "manage topics."