A-bone1/Attention-ocr-Chinese-Version
quality grade F, 18 out of 100Attention OCR Based On Tensorflow
- stars
- 429
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
466 results
Attention OCR Based On Tensorflow
This is a tensorflow re-implementation of PSENet: Shape Robust Text Detection with Progressive Scale Expansion Network.My blog:
Tensorflow-based CNN+LSTM trained with CTC-loss for OCR
Rotational region detection based on Faster-RCNN.
🖺 OCR using tensorflow with attention
[验证码识别-部署] This project is based on CNN+BLSTM+CTC to realize verificationtion. This projeccode identificat is only for deployment models.
A Tensorflow model for text recognition (CNN + seq2seq with visual attention) available as a Python package and compatible with Google Cloud ML Engine.
End-to-end Chinese scene-text detection and recognition with CTPN, CRNN, and CTC (legacy project).
[验证码识别-训练] This project is based on CNN/ResNet/DenseNet+GRU/LSTM+CTC/CrossEntropy to realize verification code identification. This project is only for training the model.
text detection mainly based on ctpn model in tensorflow, id card detect, connectionist text proposal network
A Kotlin-based testing/scraping/parsing library providing the ability to analyze and extract data from HTML (server & client-side rendered). It places particular emphasis on ease of use and a high level of readability by providing an intuitive DSL. It aims to be a testing lib, but can also be used to scrape websites in a convenient fashion.
:speedboat: Label data at scale. Fun and precision included.
ICDAR 2019 Robust Reading Challenge on Scanned Receipts OCR and Information Extraction
An implement of the paper of EDA for Chinese corpus.中文语料的EDA数据增强工具。NLP数据增强。论文阅读笔记。
Train Tesseract LSTM with make
Code for the paper "MASTER: Multi-Aspect Non-local Network for Scene Text Recognition" (Pattern Recognition 2021)
A scene text recognition toolbox based on PyTorch
Free Offline OCR 离线的中文文本检测+识别SDK
A curated collection of AI, data engineering, and DevOps projects featuring real-world applications, advanced techniques, and tutorials—ideal for learners and practitioners exploring data science and machine learning.
爬蟲極簡教學(fetch, parse, search, multiprocessing, API)- PTT 為例
Script that crawls meta data from ICLR OpenReview webpage. Tutorials on installing and using Selenium and ChromeDriver on Ubuntu.
Neum AI is a best-in-class framework to manage the creation and synchronization of vector embeddings at large scale.
[ICLR 2026] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
Examples and guides for using the VLM Run API
24,535 repositories in the index in total.