deedy5/primp
quality grade B, 69 out of 100HTTP client that can impersonate web browsers
- stars
- 602
- stars gained this week
- +8this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
466 results
HTTP client that can impersonate web browsers
Scan documents to PDF and more, as simply as possible.
An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.
Mount Hugging Face Buckets and repos as local filesystems. No download, no copy, no waiting.
RMT (RuoMengTu) is a free, open-source macro tool built on AHKv2. Let the code handle the tedious work—you have more meaningful things to do.
OCR model that handles complex tables, forms, handwriting with full layout.
LakeSoul is an end-to-end, realtime cloud-native Lakehouse framework for fast data ingestion, concurrent updates, incremental analytics, multimodal data processing and vector search — powering next-generation BI and AI workloads.
deepseek标配coding agent、原生支持模型路由、CodeGraph代码图谱、OCR截图识别、自动上下文压缩、最佳工作模式选择,workflow等功能,从根本上节省Token
CLI and Agent Skill for Firecrawl - Add scrape, search, and browsing capabilities to your AI agents
☁️ The fastest HTML to markdown convertor on GitHub. Optimized for LLMs and supports streaming.
Collection of open-source libraries and tools for Robotic Process Automation (RPA), designed to be used with both Robot Framework and Python
Synthetic Data SDK ✨
Database anonymization and test data management
Open-source agentic data engineering harness for dbt, SQL, and cloud warehouses. 100+ tools, 10 warehouses, AI-powered.
AI Product Analyst — Claude Code-powered data analysis toolkit
The IDP Accelerator provides a scalable, serverless approach for automated document processing and information extraction using AWS services, such as Amazon Bedrock Data Automation and Amazon Bedrock foundational models. It combines generative AI and optical character recognition (OCR) to process documents at scale.
Machine learning with dataframes
A self-hosted, respource aware file conversion server supporting 488 formats in 26 languages.
Open Source Agentic Business Intelligence with Malloy Semantic Layer :tada:
A web privacy measurement framework
Metrics to evaluate quality and efficacy of synthetic datasets.
Benchmarking synthetic data generation methods.
Desktop app for managing BibTeX and BibLaTeX (.bib) libraries
Use OCR in Windows quickly and easily with Text Grab. With optional background process and notifications.
24,535 repositories in the index in total.