Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A system for quickly generating training data with weak supervision
| Date | Stars |
|---|---|
| 2026-07-24 | 5994 |
| 2026-07-25 | 5994 |
| 2026-07-28 | 5994 |
| 2026-07-30 | 5994 |
| 2026-08-06 | 5994 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<img src="figs/logo_01.png" width="150"/>    [](https://snorkel.readthedocs.io/en/master) [](https://codecov.io/gh/snorkel-team/snorkel/branch/master) [](https://opensource.org/licenses/Apache-2.0) ***Programmatically Build and Manage Training Data*** ## Announcement **The Snorkel team is now focusing their efforts on Snorkel Flow, an end-to-end AI application development platform based on the core ideas behind Snorkel—you can check it out [here](https://snorkel.ai) or [join us](https://www.snorkel.ai/careers) in building it!** The [Snorkel project](https://snorkelproject.org/) started at Stanford in 2015 with a simple technical bet: that it would increasingly be the **training data**, not the models, algorithms, or infrastructure, that decided whether a machine learning project succeeded or failed. Given this premise, we set out to explore the radical idea that you could bring mathematical and systems structure to the messy and often entirely manual process of training data creation and management, starting by empowering users to **programmatically label, build, and manage** training data. To say that the Snorkel project succeeded and expanded beyond what we had ever expected would be an understatement. The basic goals of a research repo like Snorkel are to provide a minimum viable framework for testing and validating hypotheses. Four years later, we’ve been fortunate to do not just this, but to develop and deploy early versions of Snorkel in partnership with some of the world’s leading organizations like [Google](https://ai.googleblog.com/2019/03/harnessing-organizational-knowledge-for.html), [Intel](https://dl.acm.org/doi/abs/10.1145/3329486.3329492), [Stanford Medicine](https://www.cell.com/patterns/fulltext/S2666-3899(20)30019-2), and many more; author over [sixty peer-reviewed publications](https://snorkel.ai/technology) on our findings around Snorkel and related innovations in weak supervision modeling, data augmentation, multi-task learning, and more; be included in courses at top-tier universities; support production deployments in systems that you’ve likely used in the last few hours; and work with an amazing community of researchers and practitioners from industry, medicine, government, academia, and beyond. However, we realized increasingly–from conversations with users in weekly office hours, workshops, online discussions, and industry partners–that the Snorkel project was just the very first step. The ideas behind Snorkel change not just how you label training data, but so much of the entire lifecycle and pipeline of building, deploying, and managing ML: how users inject their knowledge; how models are constructed, trained, inspected, versioned, and monitored; how entire pipelines are developed iteratively; and how the full set of stakeholders in any ML deployment, from subject matter experts to ML engineers, are incorporated into the process. Over the last year, we have been building the platform to support this broader vision: [Snorkel Flow](https://snorkel.ai/snorkel-flow-platform/), an end-to-end machine learning platform for developing and deploying AI applications. Snorkel Flow incorporates many of the concepts of the Snorkel project with a range of newer techniques around weak supervision modeling, data augmentation, multi-task learning, data slicing and structuring, monitoring and analysis, and more, all of which integrate in a way that is greater than the sum of its parts–and that we believe makes ML truly faster, more flexible, and more practical than ever before. Moving forward, we will be focusin
Excerpt of 8,296 characters
Read on GitHubAlex Ratner · United States
876
660
Stephen Bach · Brown University
467
160
83
77
54
40
31
Catalin Voss · Stanford University · United States
26
22
Daniel Himmelstein · @radoverlay
15
11
7
Humza Iqbal
7
David Nicholson · United States
7
7
6
Shawn R. Roberts · United States
5
Peter M. Landwehr · Cytovale · United States
5
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:34938482357fc63a, topic:training-data, desc:training data, readme:training data
matched fp:34938482357fc63a, topic:data-augmentation, readme:data augmentation