fjxmlzn/DoppelGANger
quality grade D, 36 out of 100[IMC 2020 (Best Paper Finalist)] Using GANs for Sharing Networked Time Series Data: Challenges, Initial Promise, and Open Questions
- stars
- 312
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Published datasets, dataset tooling and corpora for training and evaluation.
Signals: dataset, datasets, corpus, training-data, open-data
186 results
[IMC 2020 (Best Paper Finalist)] Using GANs for Sharing Networked Time Series Data: Challenges, Initial Promise, and Open Questions
✨ A synthetic dataset generation framework that produces diverse coding questions and verifiable solutions - all in one framwork
awesome synthetic (text) datasets
This repository contains the collection of UCI (real-life) datasets and Synthetic (artificial) datasets (with cluster labels and MATLAB files) ready to use with clustering algorithms.
Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline
Aerial Imagery dataset for fire detection: classification and segmentation (Unmanned Aerial Vehicle (UAV))
Dataset, streaming, and file system extensions maintained by TensorFlow SIG-IO
中文语言理解测评基准 Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard
高质量中文预训练模型集合:最先进大模型、最快小模型、相似度专门模型
大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP
An Integrated Corpus Tool With Multilingual Support for the Study of Language, Literature, and Translation
Reuters and Bloomberg
HISTDATA - Dataset composed of all FX trading pairs / Crude Oil / Stock Indexes. Simple API to retrieve 1 Minute data (and tick data) Historical FX Prices (up to date).
Pure Python, lightweight, Pillow-based solver for Amazon's text captcha.
A machine learning tool for automated prediction engineering. It allows you to easily structure prediction problems and generate labels for supervised learning.
The AI Datastore for Schemas, BLOBs, and Predictions. Use with your apps or integrate built-in Human Supervision, Data Workflow, and UI Catalog to get the most value out of your AI Data.
A system for quickly generating training data with weak supervision
[NeurIPS'24 Spotlight] Text2CAD: Generating Sequential CAD Designs from Beginner-to-Expert Level Text Prompts
A wrapper around tensor2tensor to flexibly train, interact, and generate data for neural chatbots.
A sleek dataset viewer built entirely by AI Agent. Supports streaming large files from WebDAV, S3, SSH, Local or Hugging Face.
[NeurIPS 2025 Spotlight] OpenCUA: Open Foundations for Computer-Use Agents
🍃 MINT-1T: A one trillion token multimodal interleaved dataset.
一个面向多模态大模型训练的智能数据集构建与评估平台
🔊 A comprehensive list of open-source datasets for voice and sound computing (95+ datasets).
24,523 repositories in the index in total.