Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
| Date | Stars |
|---|---|
| 2026-07-24 | 27803 |
| 2026-07-25 | 27842 |
| 2026-07-28 | 27842 |
| 2026-07-30 | 27842 |
| 2026-08-06 | 27842 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
35.0
growth rate 0.00%/day
<!-- AI-AGENT-SUMMARY name: opendataloader-pdf category: PDF data extraction, PDF accessibility automation license: Apache-2.0 solves: [PDF to structured data for RAG/LLM pipelines, accelerate PDF accessibility remediation — layout analysis + auto-tagging to Tagged PDF as foundation for PDF/UA (first open-source end-to-end)] input: PDF files (digital, scanned, tagged) output: Markdown, JSON (with bounding boxes), HTML, Tagged PDF, PDF/UA (enterprise) sdk: Python, Node.js, Java requirements: Java 11+ pricing: open-source core (data extraction, layout analysis, auto-tagging to Tagged PDF), enterprise add-on (PDF/UA export, accessibility studio) extraction-benchmark: #1 overall extraction accuracy (0.907) in hybrid mode, 0.928 table extraction accuracy, 0.015s/page local mode accessibility-validation: PDF Association collaboration, Well-Tagged PDF specification, veraPDF automated validation key-differentiators: [benchmark #1 PDF parser, deterministic output, bounding boxes for every element, XY-Cut++ reading order, AI safety filters, hybrid AI mode, first open-source PDF auto-tagging to Tagged PDF, PDF Association + Dual Lab (veraPDF) collaboration, Well-Tagged PDF spec compliance] --> # OpenDataLoader PDF **PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.** [](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/LICENSE) [](https://pypi.org/project/opendataloader-pdf/) [](https://www.npmjs.com/package/@opendataloader/pdf) [](https://search.maven.org/artifact/org.opendataloader/opendataloader-pdf-core) [](https://github.com/opendataloader-project/opendataloader-pdf#java) <a href="https://trendshift.io/repositories/21917" target="_blank"><img src="https://trendshift.io/api/badge/repositories/21917" alt="opendataloader-project%2Fopendataloader-pdf | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a> 🔍 **PDF parser for AI data extraction** — Extract Markdown, JSON (with bounding boxes), and HTML from any PDF. #1 in benchmarks (0.907 overall). Deterministic local mode + AI hybrid mode for complex pages. - **How accurate is it?** — #1 in benchmarks: 0.907 overall, 0.928 table accuracy across 200 real-world PDFs including multi-column and scientific papers. Deterministic local mode + AI hybrid mode for complex pages ([benchmarks](#extraction-benchmarks)) - **Scanned PDFs and OCR?** — Yes. Built-in OCR (80+ languages) in hybrid mode. Works with poor-quality scans at 300 DPI+ ([hybrid mode](#hybrid-mode-1-accuracy-for-complex-pdfs)) - **Tables, formulas, images, charts?** — Yes. Complex/borderless tables, LaTeX formulas, and AI-generated picture/chart descriptions all via hybrid mode ([hybrid mode](#hybrid-mode-1-accuracy-for-complex-pdfs)) - **How do I use this for RAG?** — `pip install opendataloader-pdf`, convert in 3 lines. Outputs structured Markdown for chunking, JSON with bounding boxes for source citations, and HTML. LangChain integration available. Python, Node.js, Java SDKs ([quick start](#get-started-in-30-seconds) | [LangChain](#langchain-integration)) ♿ **PDF accessibility automation** — Auto-tag untagged PDFs into screen-reader-ready Tagged PDFs at scale. First open-source tool to generate Tagged PDFs end-to-end. - **What's the problem?** — Accessibility regulations are now enforced worldwide. Manual PDF remediation costs $50–200 per document and doesn't scale ([regulations](#pdf-accessibility--pdfua-conversion)) - **What's free?** — Layout analysis + auto-tagging (Apache 2.0). Untagged PDF in → Tagged PDF out. No proprietary SDK dependency ([auto-tagging](#auto-tagging)) - **What about PDF/U
Excerpt of 34,693 characters
Read on GitHubBundo Lee · Hancom Inc. @opendataloader-project · South Korea
553
Maxim Pliushchov · Dual Lab
74
Kakhnovich Raman · Dual Lab
71
Hyunhee Jo · Hancom Inc. @opendataloader-project
58
Jonggyu Lee
19
16
sujicho
16
Boris Doubrov · Dual Lab
5
Shawn
4
sck_0
4
David Choi · South Korea
4
3
Kinjal Dutta
3
Veni Vidi Vici
2
1
Matt Van Horn
1
Ankit Singh
1
Brett
1
EB · Keimyung Univ. Computer Network Lab · South Korea
1
Hamid Husain
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:fea6f737f7370397, topic:ocr, topic:document-parsing, topic:pdf-parser