0xSero/ai-data-extraction
quality grade D, 44 out of 100extract all your personal data history from cursor, codex, claude-code, windsurf, and trae
- stars
- 845
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
473 results
extract all your personal data history from cursor, codex, claude-code, windsurf, and trae
Open-source Claude Code skills for cold email and outbound sales. Grade campaigns, export Prospeo searches, scrape Google Maps — all from Claude Code.
Turn an entire GitHub Repo into a single organized .txt file to use with LLM's (ChatGPT, Claude, Gemini, etc)
🔍 Better text detection by combining multiple OCR engines (EasyOCR, Tesseract, and Pororo) with 🧠 LLM.
DOM to Semantic-Markdown for use with LLMs
DataInfra Series. Redact EVERYTHING with local llms and vlms.
PDF to markdown using vision LLMs — tables, layouts, and structure preserved
Twitter data scraping, embedding based image search and more.
No description
This is a public repository to go over all the LLM-driven data engineering concepts.
Data-centric LLM training with dynamic sample selection, domain mixture optimization, and example reweighting inside the LLaMA-Factory training loop.
A fast tool to convert any website into LLM-ready markdown data. Built by https://supermemory.ai
Transform Web Content into LLM-Ready Data
Improved file parsing for LLM’s
A reading list on LLM based Synthetic Data Generation 🔥
[NeurIPS2023] DatasetDM:Synthesizing Data with Perception Annotations Using Diffusion Models
Effective Data Augmentation With Diffusion Models
The Open-Source Data Annotation Platform
Implementation of paper Data Engineering for Scaling Language Models to 128K Context
Intelligent proxy pool for Humans™ to extract content from the internet and build your own Large Language Models in this new AI era
OCR in the Era of Large Language Models
🔦 A Pytorch implementation of GoogleBrain's SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
YiShape-Math is a high-performance Java math library that provides NumPy-like functionalities including vector & matrix operations, data visualization, statistics, optimization, time series, signal processing, multi-criteria programming, distance metric learning, and machine learning models.
IntelliScraper: An advanced, intelligent web scraping tool leveraging BeautifulSoup and machine learning for efficient data extraction and analysis.
24,523 repositories in the index in total.