faustomorales/keras-ocr
quality grade D, 46 out of 100A packaged and flexible version of the CRAFT text detector and Keras CRNN recognition model.
- stars
- 1.5k
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
473 results
A packaged and flexible version of the CRAFT text detector and Keras CRNN recognition model.
Open-source platform for extracting structured data from documents using AI.
OCR离线图片文字识别命令行windows程序,以JSON字符串形式输出结果,方便别的程序调用。提供各种语言API。由 PaddleOCR C++ 编译。
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
MORT 번역기 프로젝트 - Real-time game translator with OCR
Python tool for grabbing text via screenshot
An iOS OCR Server Using Apple’s Vision Framework
Open Source Virtual (Network) Printer for Windows that allows you to create PDFs, OCR text, and print images, with advanced features usually available only in enterprise solutions.
A Gtk/Qt front-end to tesseract-ocr.
A Python wrapper for the tesseract-ocr API
Rust multi‑backend OCR/VLM engine (DeepSeek‑OCR-1/2, PaddleOCR‑VL, DotsOCR) with DSQ quantization and an OpenAI‑compatible server & CLI – run locally without Python.
A set of tools for extracting tables from PDF files helping to do data mining on (OCR-processed) scanned documents.
Assist in organizing your piles of documents, resulting from scanners, e-mails and other sources with miminal effort.
A tool to convert a Wallpaper's color scheme / palette, OCR with VLM's Traditional & Hybrid, Image Compression ,color palette extraction, image upsacling with Adversarial Networks and more image processing features.
Scan, index, and archive all of your paper documents (acquired by Mayan EDMS)
OCR powered screen-capture tool to capture information instead of images
验证码识别
开源易用的中文离线OCR,识别率媲美大厂,并且提供了易用的web页面及web的接口,方便人类日常工作使用或者其他程序来调用~
Open Source Document Management System for Digital Archives (Scanned Documents)
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
A wrapper to work with Tesseract OCR inside PHP.
树洞 OCR 文字识别(一款跨平台的 OCR 小工具)
Go package for OCR (Optical Character Recognition), by using Tesseract C++ library
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown
24,523 repositories in the index in total.