A structured and human-curated collection of OCR and Document AI resources. Automation discovers and validates candidates; maintainers review every resource before publication.
| Category | Count | Description |
|---|---|---|
| Papers | 243 | Research on OCR, document parsing, layout analysis, and document understanding |
| Models | 92 | Public model cards, weights, APIs, and official model releases |
| Datasets | 46 | Training datasets and evaluation benchmarks |
| Codes | 24 | Notable OCR and Document AI codebases |
| Skills | 102 | Installable agent skills for OCR and document workflows |
| Platforms | 0 | Domestic and international OCR platforms and services |
- YAML files under
data/are the source of truth. - Markdown lists are generated; do not edit them directly.
- Discovery runs daily with a seven-day lookback and opens a Draft PR for human review.
- Link and metadata health are audited weekly.
See CONTRIBUTING.md and CURATION.md before submitting a resource.
- HCIILAB Scene-Text-Detection
- HCIILAB Scene-Text-Recognition
- Awesome OCR
- Awesome Scene Text Recognition
Metadata and editorial content are dedicated to the public domain under CC0-1.0. Maintenance code is licensed under MIT. Third-party resources retain their own licenses.