ocrmypdf/OCRmyPDF
quality grade A, 87 out of 100OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
- stars
- 34k
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
473 results
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
A community-supported supercharged document management system: scan, index and archive all your documents
Tesseract Open Source OCR Engine (main repository)
A fast, helpful, and open-source document parser
Privacy and Security focused Segment-alternative, in Golang and React
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Metrics to evaluate quality and efficacy of synthetic datasets.
Synthetic Patient Population Simulator
Synthetic data generation for tabular data
The free and privacy-friendly screen recorder with no limits 🎥
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
Mimesis is a Python library for generating fake but realistic data in multiple languages and locales.
Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for every output.
Pure-Rust, CPU-only OCR engine for Baidu Unlimited-OCR (a DeepSeek-OCR-derived 3B MoE VLM). Five-model zoo, custom int8 kernels, no ML framework, no Python, no GPU.
JableTV, MissAV & SupJav GUI downloader for Windows — browse, batch download, monitor categories, and create AI subtitles (日本語 / English / 繁體中文). Japanese speech is transcribed locally with whisper.cpp; video/audio stay on your PC.
一款适用于央视网的网络视频流解析处理工具
Trap AI web scrapers in an endless poison pit.
An ergonomic, privacy-aware Python HTTP Client
The headless Chrome/Chromium driver on top of Puppeteer. Take screenshots, generate PDFs, extract text and HTML with a production-ready API.
Python library and CLI for X/Twitter scraping with multi-account rotation and built-in rate-limit handling.
Python binding for curl-impersonate fork via cffi. A http client that can impersonate browser tls/ja3/http2 fingerprints.
A Chrome DevTools Protocol driver for web automation and scraping.
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
24,523 repositories in the index in total.