Zhen-Tan-dmml/LLM4Annotation
quality grade D, 45 out of 100No description
- stars
- 648
- stars gained this week
- —
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
466 results
No description
This is a public repository to go over all the LLM-driven data engineering concepts.
A fast tool to convert any website into LLM-ready markdown data. Built by https://supermemory.ai
Transform Web Content into LLM-Ready Data
Improved file parsing for LLM’s
A reading list on LLM based Synthetic Data Generation 🔥
[NeurIPS2023] DatasetDM:Synthesizing Data with Perception Annotations Using Diffusion Models
Effective Data Augmentation With Diffusion Models
The Open-Source Data Annotation Platform
Implementation of paper Data Engineering for Scaling Language Models to 128K Context
Intelligent proxy pool for Humans™ to extract content from the internet and build your own Large Language Models in this new AI era
🔦 A Pytorch implementation of GoogleBrain's SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
YiShape-Math is a high-performance Java math library that provides NumPy-like functionalities including vector & matrix operations, data visualization, statistics, optimization, time series, signal processing, multi-criteria programming, distance metric learning, and machine learning models.
IntelliScraper: An advanced, intelligent web scraping tool leveraging BeautifulSoup and machine learning for efficient data extraction and analysis.
Bob is a free signal-processing and machine learning toolbox originally developed by the Biometrics group at Idiap Research Institute, in Switzerland.
A python program that uses the concept of OCR using machine learning to identify the characters on a Nigerian license plate
A ruby gem for elegant data science and machine learning
Augment Beancount importers with machine learning functionality.
Descriptor computation(chemistry) and (optional) storage for machine learning
Data Shapley: Equitable Valuation of Data for Machine Learning
[Open-Source Project] Combining MMOCR with Segment Anything & Stable Diffusion. Automatically detect, recognize and segment text instances, with serval downstream tasks, e.g., Text Removal and Text Inpainting
Multilingual Document Layout Parsing in a Single Vision-Language Model
Synthetic data generation demo using a UK retail transactional dataset. Ideal for professionals in retail, e-commerce, finance, and supply chain sectors who want to create privacy-preserving, realistic synthetic data for testing, analysis, and machine learning.
apricot implements submodular optimization for the purpose of selecting subsets of massive data sets to train machine learning models quickly. See the documentation page: https://apricot-select.readthedocs.io/en/latest/index.html
24,535 repositories in the index in total.