zhoubear/open-paperless
quality grade D, 40 out of 100Scan, index, and archive all of your paper documents (acquired by Mayan EDMS)
- stars
- 2.6k
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
466 results
Scan, index, and archive all of your paper documents (acquired by Mayan EDMS)
验证码识别
开源易用的中文离线OCR,识别率媲美大厂,并且提供了易用的web页面及web的接口,方便人类日常工作使用或者其他程序来调用~
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
A wrapper to work with Tesseract OCR inside PHP.
树洞 OCR 文字识别(一款跨平台的 OCR 小工具)
Go package for OCR (Optical Character Recognition), by using Tesseract C++ library
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown
🚀 Screenshots, word marking, OCR, AI, translation software || 截图、划词、文字识别、AI、翻译软件
A synthetic data generator for text recognition
Text recognition (optical character recognition) with deep learning methods, ICCV 2019
Fast and simple OCR library written in Swift
Advanced real-time screen translator for games, hardcoded subtitles in videos, static text and etc.
The Swift machine learning library.
yolo3+ocr
Transforms PDF, Documents and Images into Enriched Structured Data
GLM-OCR: Accurate × Fast × Comprehensive
Trained models with fast variant of the "best" LSTM models + legacy models
Ingest, parse, and optimize any data format ➡️ from documents to multimedia ➡️ for enhanced compatibility with GenAI frameworks
OCR & Document Extraction using vision models
带带弟弟 通用验证码识别OCR pypi版
Pure Javascript OCR for more than 100 Languages 📖🎉🖥
基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, with ultra-fast inference speed.
二维码/条码识别、身份证识别、银行卡识别、车牌识别、图片文字识别、黄图识别、驾驶证(驾照)识别
24,535 repositories in the index in total.