allenai/olmocr
quality grade B, 65 out of 100Toolkit for linearizing PDFs for LLM datasets/training
- stars
- 20k
- stars gained this week
- +163this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Published datasets, dataset tooling and corpora for training and evaluation.
Signals: dataset, datasets, corpus, training-data, open-data
190 results
Toolkit for linearizing PDFs for LLM datasets/training
OSINT & recon toolkit // 100+ tools, one-command installer, SOCMINT, GEOINT, network recon, dark web, forensics & more.
CSGHub is a brand-new open-source platform for managing LLMs, developed by the OpenCSG team. It offers both open-source and on-premise/SaaS solutions, with features comparable to Hugging Face. Gain full control over the lifecycle of LLMs, datasets, and agents, with Python SDK compatibility with Hugging Face. Join us! ⭐️
What's in your data? Extract schema, statistics and entities from datasets
CKAN is an open-source DMS (data management system) for powering data hubs and data portals. CKAN makes it easy to publish, share and use data. It powers catalog.data.gov, open.canada.ca/data, data.humdata.org among many other sites.
A unified framework for machine learning with time series
Public BTC trading context since 2020.
Dataset Management Framework, a Python library and a CLI tool to build, analyze and manage Computer Vision datasets.
Open Public Domain Exercise Dataset in JSON format, over 800 exercises with a browsable public searchable frontend
Paper list and datasets for industrial image anomaly/defect detection (updating). 工业异常/瑕疵检测论文及数据集检索库(持续更新)。
A very simple news crawler with a funny name
TFDS is a collection of datasets ready to use with TensorFlow, Jax, ...
A system for quickly generating training data with weak supervision
Part guillotine, part graveyard for Google's doomed apps, services, and hardware.
📰 Diários oficiais brasileiros acessíveis a todos | 📰 Brazilian government gazettes, accessible to everyone.
🌀 AI-native framework for building data portals. Scaffold a full portal from a brief and load datasets in minutes with agentic skills — any backend (CKAN, GitHub, Frictionless).
Common Voice is part of Mozilla's initiative to help teach machines how real people speak.
🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.
⚡FlashRAG: A Python Toolkit for Efficient RAG Research (WWW2025 Resource)
Let's build better datasets, together!
🎁 7,400,000+ Unsplash images made available for research and machine learning
Datasets for deep learning with satellite & aerial imagery
24,537 repositories in the index in total.