zhanlaoban/EDA_NLP_for_Chinese
quality grade F, 27 out of 100An implement of the paper of EDA for Chinese corpus.中文语料的EDA数据增强工具。NLP数据增强。论文阅读笔记。
- stars
- 1.4k
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
473 results
An implement of the paper of EDA for Chinese corpus.中文语料的EDA数据增强工具。NLP数据增强。论文阅读笔记。
Train Tesseract LSTM with make
Code for the paper "MASTER: Multi-Aspect Non-local Network for Scene Text Recognition" (Pattern Recognition 2021)
A scene text recognition toolbox based on PyTorch
Free Offline OCR 离线的中文文本检测+识别SDK
🌈一个跨平台的划词翻译和OCR软件 | A cross-platform software for text translation and recognition.
爬蟲極簡教學(fetch, parse, search, multiprocessing, API)- PTT 為例
Script that crawls meta data from ICLR OpenReview webpage. Tutorials on installing and using Selenium and ChromeDriver on Ubuntu.
Neum AI is a best-in-class framework to manage the creation and synchronization of vector embeddings at large scale.
[ICLR 2026] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
An local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR-2 AI (running directly on your machine).
Examples and guides for using the VLM Run API
多平台内容监控·采集·搬运 一个 Web 面板管起抖音 / 小红书 / 快手
All-in-one web scraping API for real-time, large-scale data extraction – proxies, CAPTCHA handling, JS rendering, and parsing in a single request returning HTML or structured JSON.
SEO Macroscope is a website scanning tool, to check your website for broken links; including some technical SEO functionality, site scraping, Excel reporting, and more.
Laravel adapter for Roach, the complete web scraping toolkit for PHP.
List of libraries, tools and APIs for web scraping and data processing.
Công cụ quét và phân tích từ khoá các trang báo mạng Việt Nam
This repository provides everything you need to get started with Python for (social science) research.
A Chrome extension for writing custom web scraping programs and web automation programs. Just demonstrate how to collect the first row of data, then let the extension write the program for collecting all rows.
The GPT-based Universal Web Scraper MVP is a solution that leverages GPT models and web scraping libraries to generate scraper code based on user input and website analysis, simplifying the web scraping process.
Netflix like full-stack application with SPA client and backend implemented in service oriented architecture
High-performance web crawler API optimized for LLMs. Turn any search or website into clean Markdown using remote browsers. Firecrawl alternative
MinerU-HTML: An SLM-powered HTML main content extractor that outputs clean HTML bodies. Perfect for Deep Research Agents, RAG applications, and training data generation.
24,523 repositories in the index in total.