Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A reading list on LLM based Synthetic Data Generation 🔥
| Date | Stars |
|---|---|
| 2026-07-31 | 1547 |
| 2026-08-04 | 1547 |
| 2026-08-06 | 1547 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Synthetic Data of LLMs, by LLMs, for LLMs
<div align="center">
[](https://github.com/wasiahmad/Awesome-LLM-Synthetic-Data-Generation/blob/main/LICENSE)

[](https://github.com/wasiahmad/Awesome-LLM-Synthetic-Data-Generation/commits/main)
[](https://github.com/wasiahmad/Awesome-LLM-Synthetic-Data-Generation/pulls)
[](https://github.com/wasiahmad/Awesome-LLM-Synthetic-Data-Generation)
<!--  -->
</div>
This repo includes papers, tools, and blogs about Synthetic Data of LLMs, by LLMs, for LLMs.
Thanks for all the great contributors on GitHub!🔥⚡🔥
## Contents
- [Synthetic Data of LLMs, by LLMs, for LLMs](#synthetic-data-of-llms-by-llms-for-llms)
- [Contents](#contents)
- [1. Surveys](#1-surveys)
- [2. Methods](#2-methods)
- [2.1. Techniques](#21-techniques)
- [2.2. Instruction Generation with High Quality/Complexity](#22-instruction-generation-with-high-qualitycomplexity)
- [3. Application Areas](#3-application-areas)
- [3.1. Mathematical Reasoning](#31-mathematical-reasoning)
- [3.2. Code Generation](#32-code-generation)
- [3.3. Text-to-SQL](#33-text-to-sql)
- [3.4. Alignment](#34-alignment)
- [3.5. Reward Modeling](#35-reward-modeling)
- [3.6. Long Context](#36-long-context)
- [3.7. Weak-to-Strong](#37-weak-to-strong)
- [3.8. Agent and Tool Use](#38-agent-and-tool-use)
- [3.9. Vision and Language](#39-vision-and-language)
- [3.10. Factuality](#310-factuality)
- [4. Datasets](#4-datasets)
- [5. Tools](#5-tools)
- [6. Blogs](#6-blogs)
### 1. Surveys
- [**Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation**](https://arxiv.org/abs/2502.17521) *Simin Chen, Yiming Chen, Zexin Li, Yifan Jiang, Zhongwei Wan, Yixin He, Dezhi Ran, Tianle Gu, Haizhou Li, Tao Xie, Baishakhi Ray.* Preprint'25
- [**A Survey on Bridging VLMs and Synthetic Data**](https://www.techrxiv.org/doi/full/10.36227/techrxiv.174741263.32891073/v1) *Mohammad Ghiasvand Mohammadkhani, Saeedeh Momtazi, Hamid Beigy.* TechRxiv 2025.
- [**Best Practices and Lessons Learned on Synthetic Data for Language Models**](https://arxiv.org/abs/2404.07503) *Ruibo Liu, Jerry Wei, Fangyu Liu, Chenglei Si, Yanzhe Zhang, Jinmeng Rao, Steven Zheng, Daiyi Peng, Diyi Yang, Denny Zhou, Andrew M. Dai.* COLM 2024.
- [**On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey**](https://arxiv.org/abs/2406.15126) *Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao, Xiao Ding, Gang Chen, Haobo Wang.* Arxiv 2024.
- [**Large Language Models for Data Annotation: A Survey**](https://arxiv.org/abs/2402.13446) *Zhen Tan, Dawei Li, Song Wang, Alimohammad Beigi, Bohan Jiang, Amrita Bhattacharjee, Mansooreh Karami, Jundong Li, Lu Cheng, Huan Liu.* Arxiv 2024.
- [**Generative AI for Synthetic Data Generation: Methods, Challenges and the Future**](https://arxiv.org/abs/2403.04190) *Xu Guo, Yiqiang Chen.* Arxiv 2024.
- [**Comprehensive Exploration of Synthetic Data Generation: A Survey**](https://arxiv.org/abs/2401.02524) *André Bauer, Simon Trapp, Michael Stenger, Robert Leppich, Samuel Kounev, Mark Leznik, Kyle Chard, Ian Foster.* Arxiv 2024.
## 2. Methods
### 2.1. Techniques
- [**STaR: Bootstrapping Reasoning With Reasoning**](https://arxiv.org/abs/2203.14465) *Eric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. Goodman.* NeurIPS 2022.
- [**Symbolic Knowledge Distillation: from General Language Models to Commonsense Models**](https://aExcerpt of 21,920 characters
Read on GitHub57
5
3
Gabriel Martín Blázquez · @supersonik-ai · Spain
1
Brian Li · Advanced Machine Intelligence
1
Mohammad Ghiasvand Mohammadkhani
1
Bill Yuchen Lin · @xai-org · United States
1
SJ Yu
1
Vadim Borisov
1
1
1
1
1
1
1
1
Guohao Li · CAMEL-AI.org
1
1
Amit Chaudhary
1
Mikayel Samvelyan · @google-deepmind · United Kingdom
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:4b91fa3232df467a, name:synthetic data, desc:synthetic data