Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
| Date | Stars |
|---|---|
| 2026-07-31 | 647 |
| 2026-08-05 | 647 |
| 2026-08-06 | 647 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Large Language Models for Data Annotation and Synthesis: A Survey
 [](https://awesome.re) 
- This is a curated list of papers about LLM for Data Annotation and Synthesis
maintained by Dawei Li ([email protected])
- If you want to add new entries, please make PRs with the same format.
- This list serves as a complement to our EMNLP 2024 oral survey:
[[Large Language Models for Data Annotation and Synthesis: A Survey]](https://arxiv.org/pdf/2402.13446.pdf)
## 🔔 News
- **`2025-4` We update our paper list and include papers for LLM-based data annotation & synthesis in March 2025!**
- **`2025-3` Want to learn more about risks and safety problems of using LLM-based annotation? Check out our new paper list on [AI supervision risk](https://github.com/David-Li0406/AI-Supervision-Risk)!**
- **`2025-3` We update our paper list and include papers for LLM-based data annotation & synthesis in February 2025!**
- **`2025-2` We collect papers and datasets about long-CoT synthesis and distillation, check it!**
- **`2025-2` We update our paper list and include papers for LLM-based data annotation & synthesis in January 2025!**
- **`2025-2` Check our new paper about [preference leakage](https://arxiv.org/abs/2502.01534)!**
- **`2024-12` Check our new paper list and survey on [LLM-as-a-judge](https://github.com/llm-as-a-judge/Awesome-LLM-as-a-judge)!**
- **`2024-12` We update our paper list and include papers for LLM-based data annotation & synthesis in December 2024!**
<div align=center><img src="https://github.com/Zhen-Tan-dmml/LLM4Annotation/blob/main/figure/taxonomy.png" width="700" /></div>
<div align=center><img src="https://github.com/Zhen-Tan-dmml/LLM4Annotation/blob/main/figure/framework.png" width="700" /></div>
If you find this repo helpful, we would appreciate it if you could cite our survey.
```
@article{tan2024large,
title={Large language models for data annotation: A survey},
author={Tan, Zhen and Li, Dawei and Wang, Song and Beigi, Alimohammad and Jiang, Bohan and Bhattacharjee, Amrita and Karami, Mansooreh and Li, Jundong and Cheng, Lu and Liu, Huan},
journal={arXiv preprint arXiv:2402.13446},
year={2024}
}
```
## Long-CoT Synthesis & Distillation
### Papers
- **Self-Training Elicits Concise Reasoning in Large Language Models**. *ArXiv preprint* (2025) [[Paper]](https://arxiv.org/abs/2502.20122)
- **Rank1: Test-Time Compute for Reranking in Information Retrieval**. *ArXiv preprint* (2025) [[Paper]](https://arxiv.org/abs/2502.18418)
- **O1 Embedder: Let Retrievers Think Before Action**. *ArXiv preprint* (2025) [[Paper]](https://arxiv.org/abs/2502.07555)
- **MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification**. *ArXiv preprint* (2025) [[Paper]](https://arxiv.org/abs/2502.13383)
- **Small Models Struggle to Learn from Strong Reasoners**. *ArXiv preprint* (2025) [[Paper]](https://arxiv.org/abs/2502.12143)
- **Boosting Multimodal Reasoning with MCTS-Automated Structured Thinking**. *ArXiv preprint* (2025) [[Paper]](https://arxiv.org/abs/2502.02339)
- **Unveiling the Mechanisms of Explicit CoT Training: How Chain-of-Thought Enhances Reasoning Generalization**. *ArXiv preprint* (2025) [[Paper]](https://arxiv.org/abs/2502.04667)
- **ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates**. *ArXiv preprint* (2025) [[Paper]](https://arxiv.org/abs/2502.06772)
- **s1: Simple test-time scaling**. *ArXiv preprint* (2025) [[Paper]](https://arxiv.org/abs/2501.19393)
- **Cascaded Self-Evaluation Augmented Training for Efficient Multimodal Large Language Models**. *ArXiv preprint* (2025) [[Paper]](https://arxiv.org/abs/2501.05662)
- **AtomThink: A Slow Thinking Framework for Multimodal Mathematical Reasoning**. *ArXiv preprint* (2024) [[Paper]](https://arxiv.org/abs/2411.11930)
- **CoT-Valve: LenExcerpt of 157,299 characters
Read on GitHub19
18
6
3
Mohammad Ghiasvand Mohammadkhani
3
1
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:e887302c5a159557, llm:Repository name 'LLM4Annotation' suggests use of large language models for annotation/labeling tasks; no README or topics provided.
matched fp:e887302c5a159557, llm:Repository name 'LLM4Annotation' suggests use of large language models for annotation/labeling tasks; no README or topics provided.
matched fp:e887302c5a159557, llm:Repository name 'LLM4Annotation' suggests use of large language models for annotation/labeling tasks; no README or topics provided.