Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A curated list of awesome synthetic data tools (open source and commercial).
| Date | Stars |
|---|---|
| 2026-07-24 | 261 |
| 2026-07-25 | 261 |
| 2026-07-28 | 261 |
| 2026-07-30 | 261 |
| 2026-08-06 | 261 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# awesome-synthetic-data A curated list of awesome synthetic data tools (open source and commercial). Inspired by [Awesome Synthetic Data](https://github.com/gretelai/awesome-synthetic-data) # Table of content + [Open source tools](#open-source-tools) + [Commercial solutions](#commercial-solutions) + [Online communities](#online-communities) # Open source tools + [Copulas](https://github.com/sdv-dev/Copulas): a Python library for modeling multivariate distributions and sampling from them using copula functions. + [CTGAN](https://github.com/sdv-dev/CTGAN): SDV’s collection of deep learning-based synthetic data generators for single table data. + [DataGene](https://github.com/firmai/datagene): a tool to train, test, and validate datasets, detect and compare dataset similarity between real and synthetic datasets. + [DoppelGANger](https://github.com/fjxmlzn/DoppelGANger): a synthetic data generation framework based on generative adversarial networks (GANs). + [DP_WGAN-UCLANESL](https://github.com/nesl/nist_differential_privacy_synthetic_data_challenge): this solution trains a Wasserstein generative adversarial network (w-GAN) that is trained on the real private dataset. + [DPSyn](https://github.com/usnistgov/PrivacyEngCollabSpace/tree/master/tools/de-identification/Differential-Privacy-Synthetic-Data-Challenge-Algorithms/DPSyn): an algorithm for synthesizing microdata while satisfying differential privacy. + [Faker](https://github.com/joke2k/faker): a Python package that generates fake data (Note: this tool does not generate synthetic data but offers dummy data). + [Generative adversarial nets for synthetic time series data](https://github.com/stefan-jansen/synthetic-data-for-finance): a repository that shows how to create synthetic time-series data using generative adversarial networks (GANs). + [Gretel.ai](https://gretel.ai/): commercial synthetic data vendor that offers open source functionality. + [mirrorGen](https://github.com/DataResponsibly/MirrorDataGenerator): a python tool that generates synthetic data based on user-specified causal relations among features in the data. + [Plait.py](https://github.com/plaitpy/plaitpy): a program for generating fake data from composable yaml templates. + [Pydbgen](https://github.com/tirthajyoti/pydbgen): a Python package that generates a random database table based on the user's choice of data types. + [Smart noise synthesizer](https://smartnoise.org/): a differentially private open source synthesizer for tabular data. + [Synner](https://github.com/huda-lab/synner): an open source tool to generate real-looking synthetic data by visually specifying the properties of the dataset. + [Synth](https://www.getsynth.com/): an open source data-as-code tool that provides a simple CLI workflow for generating consistent data in a scalable way. + [Synthea](https://synthetichealth.github.io/synthea/): an open source synthetic patient generator that models the medical history of synthetic patients. + [Synthetic data vault (SDV)](https://sdv.dev/): one of the first open source synthetic data solutions, SDV provides tools for generating synthetic data for tabular, relational, and time series data. + [TGAN](https://github.com/sdv-dev/TGAN): generative adversarial training for generating synthetic tabular data. + [Tofu](https://github.com/spiros/tofu): a Python library for generating synthetic UK Biobank data. + [Twinify](https://github.com/DPBayes/twinify): a software package for privacy-preserving generation of a synthetic twin to a given sensitive data set. + [YData](https://github.com/ydataai/ydata-synthetic): synthetic structured data generator by YData, a commercial vendor. # Commercial solutions + [Betterdata](https://www.betterdata.ai/): vendor of a privacy-preserving synthetic data solution for AI, data sharing, or product development. + [Datomize](https://www.datomize.com/): vendor of a synthetic data solution for the development, training and testing of AI/ML models, and applic
Excerpt of 7,413 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:9d4e6a8bde34e853, topic:synthetic-data, name:synthetic data, desc:synthetic data
matched fp:9d4e6a8bde34e853, topic:awesome-list, desc:curated list, readme:curated list