A structured and human-curated collection of OCR and Document AI resources. Automation discovers and validates candidates; maintainers review every resource before publication.
| Category | Count | Description |
|---|---|---|
| Papers | 289 | Research on OCR, document parsing, layout analysis, and document understanding |
| Models | 480 | Public model cards, weights, APIs, and official model releases |
| Datasets | 280 | Training datasets and evaluation benchmarks |
| Codes | 101 | Notable OCR and Document AI codebases |
| Skills | 251 | Installable agent skills for OCR and document workflows |
| Platforms | 0 | Domestic and international OCR platforms and services |
- 2026-09-19
- 2026-09-18
- 2026-09-17
- 2026-09-16
- 2026-09-15
- 2026-09-14
- 2026-09-13
- 2026-09-10
- 2026-09-09
- 2026-09-08
- YAML files under
data/are the source of truth. - Markdown lists are generated; do not edit them directly.
- Discovery runs daily with a seven-day lookback and opens a Draft PR for human review.
- Link and metadata health are audited weekly.
See CONTRIBUTING.md and CURATION.md before submitting a resource.
- HCIILAB Scene-Text-Detection
- HCIILAB Scene-Text-Recognition
- Awesome OCR
- Awesome Scene Text Recognition
Metadata and editorial content are dedicated to the public domain under CC0-1.0. Maintenance code is licensed under MIT. Third-party resources retain their own licenses.