php-curl-class/php-curl-class
quality grade A, 83 out of 100PHP Curl Class makes it easy to send HTTP requests and integrate with web APIs
- stars
- 3.3k
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
473 results
PHP Curl Class makes it easy to send HTTP requests and integrate with web APIs
Augment Beancount importers with machine learning functionality.
Reverb is an efficient and easy-to-use data storage and transport system designed for machine learning research
🍁 Sycamore is an LLM-powered search and analytics platform for unstructured data.
A curated collection of AI, data engineering, and DevOps projects featuring real-world applications, advanced techniques, and tutorials—ideal for learners and practitioners exploring data science and machine learning.
Open-source batch OCR workbench — a free, local alternative to ABBYY FineReader. Powered by Ollama + GLM-OCR + PP-DocLayoutV3, ~0.5s/page on RTX 4090. Three-panel editor, layout-aware, PDF/image batch processing, Markdown/Word export. 批量OCR工作台,纯本地运行,免费平替ABBYY,适合书籍文档数字化。
An AI web scraper using ollama, brightdata, selenium and other libraries.
面向 AI Agents 的单色数据可视化 Skill,将数据快速生成精致、可交互的 HTML 图表
AI Product Analyst — Claude Code-powered data analysis toolkit
A Claude Skill that automatically analyzes uploaded CSV files — generating summary statistics, detecting missing data, and creating quick visualizations using Python and pandas.
Parse incomplete JSON text in best-effort manner. Useful for partial JSON responses, broken network packages, LLM responses with markdown fences or exceeding token limits, and configuration files with comments.
Build data processing and data analysis pipelines that leverage the power of LLMs 🧠
Use LLMs to robustly extract web data
LLM-powered lossless compression tool
A package for parsing PDFs and analyzing their content using LLMs.
Filter sensitive information from free text before sending it to external services or APIs, such as chatbots and LLMs.
A multithreaded 🕸️ web crawler that recursively crawls a website and creates a 🔽 markdown file for each page, designed for LLM RAG
clean & curate your data with LLMs.
Parse PDFs into markdown using Vision LLMs
A live reading list for LLM data synthesis (Updated to July, 2025).
DSIR large-scale data selection framework for language model training
Powerful RDF Knowledge Graph Generation with RML Mappings
ParseBench - A Document Parsing Benchmark for AI Agents
Local-first, open-source AI assistant for your data. Unify tasks, notes, docs, photos, and bookmarks. Private, self-hosted, and extensible via APIs.
24,523 repositories in the index in total.