Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Building blocks for foundation models.
| Date | Stars |
|---|---|
| 2026-07-31 | 635 |
| 2026-08-03 | 635 |
| 2026-08-06 | 635 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Building Blocks for AI Systems This is a (biased) view of great work studying the building blocks of efficient and performant foundation models. This Github was originally put together as a place to aggregate materials for a [NeurIPS keynote](https://neurips.cc/virtual/2023/invited-talk/73990) - but we're also hoping to highlight great work across AI Systems. If you think we're missing something, please open an issue or PR! Slides from Chris Ré's NeurIPS Keynote: https://cs.stanford.edu/~chrismre/papers/NeurIPS23_Chris_Re_Keynote_DELIVERED.pptx **Courses.** Courses a great resources for getting started in this space. It's great that we have so many that have open materials! Here's a partial list of courses -- it's biased by Stanford courses, so please reach out if you think of other resources that are helpful! * [Stanford CS 324 LLMs](https://stanford-cs324.github.io/winter2022/) * [Stanford CS 324 Advances in Foundation Models](https://stanford-cs324.github.io/winter2023/) * [Sasha's talk on do we need attention?](https://github.com/srush/do-we-need-attention/blob/main/DoWeNeedAttention.pdf) * [Stanford CS 229S Systems for Machine Learning](https://cs229s.stanford.edu/fall2023/) * [MLSys Seminar](https://mlsys.stanford.edu/) * [Berkeley AI-Sys](https://ucbrise.github.io/cs294-ai-sys-sp22/) * [MIT CS 6.5940](https://hanlab.mit.edu/courses/2023-fall-65940) If you just want to follow along on the major pieces from the talk, check out these blog posts: * [Data Wrangling with Foundation Models](https://hazyresearch.stanford.edu/blog/2023-01-13-datawrangling) * [FlashAttention](https://hazyresearch.stanford.edu/blog/2023-01-12-flashattention-long-sequences) and [FlashAttention-2](https://hazyresearch.stanford.edu/blog/2023-07-17-flash2) * [Simplifying S4](https://hazyresearch.stanford.edu/blog/2022-06-11-simplifying-s4) * [Long Convolutions for GPT-style Models](https://hazyresearch.stanford.edu/blog/2023-12-11-conv-tutorial) * [Zoology Synthetics Analysis](https://hazyresearch.stanford.edu/blog/2023-12-11-zoology1-analysis) * [Zoology Based](https://hazyresearch.stanford.edu/blog/2023-12-11-zoology2-based) * [Truly Sub-Quadratic Models](https://hazyresearch.stanford.edu/blog/2023-12-11-truly-subquadratic) An older set of resources on [Data-Centric AI](https://github.com/HazyResearch/data-centric-ai). The rest of this README is split up into resources by topic. **Table of contents:** * [Foundation Models for Systems](#foundation-models-for-systems) * [Hardware-Aware Algorithms](#hardware-aware-algorithms) * [Can We Replace Attention?](#can-we-replace-attention) * [Synthetics for Language Modeling](#synthetics-for-language-modeling) * [Truly Sub-Quadratic Models](#truly-sub-quadratic-models) * [Quantization, Pruning, and Distillation](#quantization-pruning-and-distillation) * [Systems for Inference](#systems-for-inference) * [High-Throughput](#high-throughput) * [New Data Types](#new-data-types) ## Foundation Models for Systems Foundation models are changing the ways that we build systems for classical problems like data cleaning. [SIGMOD keynote](https://cs.stanford.edu/~chrismre/papers/SIGMOD-Chris-Re-DataCentric-Foundation-Models-KeyNote.pdf) on this topic. Ihab Ilyas and Xu Chen's textbook on data cleaning: [Data Cleaning](https://dl.acm.org/doi/book/10.1145/3310205). The [ML for Systems](https://mlforsystems.org/) workshops and community are great. ### Blog Posts * [Bad Data Costs the U.S. $3 Trillion Per Year](https://hbr.org/2016/09/bad-data-costs-the-u-s-3-trillion-per-year) * [Data Wrangling with Foundation Models](https://hazyresearch.stanford.edu/blog/2023-01-13-datawrangling) * [Ask Me Anything: Leveraging Foundation Models for Private & Personalized Systems](https://hazyresearch.stanford.edu/blog/2023-04-18-personalization) ### Papers * [Holoclean: Holistic Data Repairs with Probabilistic Inference](https://arxiv.org/abs/1702.00820) * [Can Foundation Models Wrangle Your Data?](https://arxiv.org/abs/2205.099
Excerpt of 21,689 characters
Read on GitHub58
Simran Arora
5
3
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:e1907443597b2469, llm:description: 'Building blocks for foundation models.' Repository belongs to HazyResearch; no topics provided.
matched fp:e1907443597b2469, llm:description: 'Building blocks for foundation models.' Repository belongs to HazyResearch; no topics provided.
matched fp:e1907443597b2469, llm:description: 'Building blocks for foundation models.' Repository belongs to HazyResearch; no topics provided.