Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
| Date | Stars |
|---|---|
| 2026-07-24 | 3344 |
| 2026-07-25 | 3344 |
| 2026-07-28 | 3344 |
| 2026-07-30 | 3344 |
| 2026-08-06 | 3344 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
> [!IMPORTANT]
The original authors have moved on to other projects. A group of community members have recently joined the GitHub project as collaborators to maintain the project and are actively working towards the next release. Check out the `develop` branch for access to the latest fixes and improvements in the meantime.
>
<div align="center">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://github.com/argilla-io/distilabel/blob/main/docs/assets/distilabel-white.png?raw=true">
<img alt="Distilabel Logo" src="https://raw.githubusercontent.com/argilla-io/distilabel/main/docs/assets/distilabel-black.png">
</picture>
</div>
<h3 align="center">Synthesize data for AI and add feedback on the fly!</h3>
<p align="center">
<a href="https://pypi.org/project/distilabel/">
<img alt="CI" src="https://img.shields.io/pypi/v/distilabel.svg?style=flat-round&logo=pypi&logoColor=white">
</a>
<a href="https://pepy.tech/project/distilabel">
<img alt="CI" src="https://static.pepy.tech/personalized-badge/distilabel?period=month&units=international_system&left_color=grey&right_color=blue&left_text=pypi%20downloads/month">
</a>
</p>
<p align="center">
<a href="https://twitter.com/argilla_io">
<img src="https://img.shields.io/badge/twitter-black?logo=x"/>
</a>
<a href="https://www.linkedin.com/company/argilla-io">
<img src="https://img.shields.io/badge/linkedin-blue?logo=linkedin"/>
</a>
<a href="http://hf.co/join/discord">
<img src="https://img.shields.io/badge/Discord-7289DA?&logo=discord&logoColor=white"/>
</a>
</p>
Distilabel is the framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
If you just want to get started, we recommend you check the [documentation](http://distilabel.argilla.io/). Curious, and want to know more? Keep reading!
<!--  -->
## Why use distilabel?
Distilabel can be used for generating synthetic data and AI feedback for a wide variety of projects including traditional predictive NLP (classification, extraction, etc.), or generative and large language model scenarios (instruction following, dialogue generation, judging etc.). Distilabel's programmatic approach allows you to build scalable pipelines for data generation and AI feedback. The goal of distilabel is to accelerate your AI development by quickly generating high-quality, diverse datasets based on verified research methodologies for generating and judging with AI feedback.
### Improve your AI output quality through data quality
Compute is expensive and output quality is important. We help you **focus on data quality**, which tackles the root cause of both of these problems at once. Distilabel helps you to synthesize and judge data to let you spend your valuable time **achieving and keeping high-quality standards for your data**.
### Take control of your data and models
**Ownership of data for fine-tuning your own LLMs** is not easy but Distilabel can help you to get started. We integrate **AI feedback from any LLM provider out there** using one unified API.
### Improve efficiency by quickly iterating on the right research and LLMs
Synthesize and judge data with **latest research papers** while ensuring **flexibility, scalability and fault tolerance**. So you can focus on improving your data and training your models.
## Community
We are an open-source community-driven project and we love to hear from you. Here are some ways to get involved:
- [Community Meetup](https://lu.ma/embed-checkout/evt-IQtRiSuXZCIW6FB): listen in or present during one of our bi-weekly events.
- [Discord](http://hf.co/join/discord): get direct support from the community in #argilla-general and #argilla-help.
- [Roadmap](https://github.com/orgs/argilla-io/projects/10/views/1): plans change but we love to discuss thExcerpt of 10,595 characters
Read on GitHubGabriel Martín Blázquez · @supersonik-ai · Spain
261
Agus · Spain
224
Alvaro Bartolome · @huggingface · Spain
161
Daniel Vila Suero
57
David Berenstein · @Giskard-AI & @PrunaAI & @mlco2 · Spain
54
Ignacio Talavera · @Feverup · Spain
16
Sara Han · Spain
16
huggingface · Belgium
7
Luca Rolshoven · Switzerland
4
3
GMU
3
Daniel van Strien · Hugging Face · United Kingdom
3
David Meikle · United Kingdom
2
2
2
1
Sadra Barikbin
1
1
Lucain · @huggingface
1
Parag Ekbote · India
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:934440d0489ca301, topic:synthetic-data, desc:synthetic data, readme:synthetic data
matched fp:934440d0489ca301, topic:rlhf, readme:fine-tuning, readme:fine tuning