NVIDIA-NeMo/DataDesigner
quality grade A, 91 out of 100🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.
- stars
- 2.1k
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
473 results
🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.
Snoop — инструмент разведки на основе открытых данных (OSINT world)
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
Mouseover Translate Any Language At Once - Chrome Extension: PDF Translator, EBOOK, EPUB, OCR, TTS, NETFLIX, YOUTUBE DUAL SUBTITLES, GOOGLE DOCS, AI, VIEWER, GMAIL, WRITING, IMAGE, DUAL SUBS, MANGA, HOVER, DICTIONARY, WEBTOON, EDGE, JAPANESE, ENGLISH
NBA Stats API via Basketball Reference
Python binding to Modest and Lexbor engines. Fast HTML5 parser with CSS selectors for Python.
Mount Hugging Face Buckets and repos as local filesystems. No download, no copy, no waiting.
Fast, stateless CLI for web search and scrape. Built for AI agents.
AI-native data platform: 60+ connectors, dlt-based ETL, DAG scheduling, per-source RAG knowledge bases, and sandboxed AI agents that query, chart, and analyze your data.
Open Source Agentic Business Intelligence with Malloy Semantic Layer :tada:
The IDP Accelerator provides a scalable, serverless approach for automated document processing and information extraction using AWS services, such as Amazon Bedrock Data Automation and Amazon Bedrock foundational models. It combines generative AI and optical character recognition (OCR) to process documents at scale.
Open-source ETL/ELT on DuckDB. Write, wire, or draw one pipeline: 350+ components, 160+ connectors, dbt, CDC, data quality, a Python API, and MCP for AI agents. Runs anywhere, no lock-in.
CLI and Agent Skill for Firecrawl - Add scrape, search, and browsing capabilities to your AI agents
Note Companion: AI assistant for Obsidian that goes beyond just a chat. (prev File Organizer 2000)
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
Open-source GEO content engineering and multi-site distribution system with AI tasks, RAG/semantic chunking, analytics, GEOFlow Agent and WordPress target publishing.
🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent
Event streaming platform for agentic AI. Continuously ingest, transform, and serve event streams in real time, at scale.
Python tool for converting files and office documents to Markdown.
OCR engine for all the languages
Read Japanese manga inside browser with selectable text.
在保留版面、公式与结构的前提下进行 PDF 翻译,适用于科研与技术文档
Scan documents to PDF and more, as simply as possible.
视觉小说翻译器 / Visual Novel Translator
24,523 repositories in the index in total.