Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Central place for the engineering/scaling WG: documentation, SLURM scripts and logs, compute environment and data.
| Date | Stars |
|---|---|
| 2026-07-24 | 1018 |
| 2026-07-25 | 1018 |
| 2026-07-28 | 1018 |
| 2026-07-30 | 1018 |
| 2026-08-18 | 1019 |
| 2026-09-20 | 1019 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# bigscience
[Research workshop on large language models - The Summer of Language Models 21](https://bigscience.huggingface.co/)
At the moment we have 2 code repos:
1. https://github.com/bigscience-workshop/Megatron-DeepSpeed - this is our flagship code base
2. https://github.com/bigscience-workshop/bigscience - (this repo) for everything else - docs, experiments, etc.
Currently, the most active segments of this repo are:
- [JZ](./jz/) - Lots of information about our work environment which helps evaluate, plan and get things done
- [Experiments](./experiments) - many experiments are being done. Documentation, result tables, scripts and logs are all there
- [Datasets info](./data/)
- [Train](./train) - all the information about the current trainings (see below for the most important ones)
We have READMEs for specific aspects, such as:
- [hub integration](./tools/README.md)
## Trainings
While we keep detailed chronicles of experiments and findings for some of the main trainings, here is a doc that contains a summary of the most important findings: [Lessons learned](train/lessons-learned.md)
### Train 1 - 13B - unmodified Megatron gpt2 - baseline
* [the full spec and discussions](./train/tr1-13B-base)
* [the training script](./train/tr1-13B-base/tr1-13B-round1.slurm)
* checkpoints and logs:
- [tensorboard](https://huggingface.co/bigscience/tr1-13B-tensorboard/tensorboard)
- [logs](https://huggingface.co/bigscience/tr1-13B-logs/)
* [chronicles](./train/tr1-13B-base/chronicles.md)
You can watch the training logs live by running this `tail -f` like script over remote log file that gets synced to the hub once an hour:
```
perl -e '$u=shift; $b=0; while(1){($e)=qx[curl -sI $u]=~/content-length: (\d+)/; \
print qx[curl -sr $b-$e -L $u] if $e>$b; $b=$e; sleep 300}' \
https://huggingface.co/bigscience/tr1-13B-logs/resolve/main/main_log.txt
```
### Train 3
Architecture and scaling baseline runs: no fancy tricks, just GPT2. Here are links to the respective tensorboards:
| Size | 1B3 | 760M | 350M | 125M |
|--------------------- |----- |------ |------ |------ |
| C4 + low warmup | [a](https://huggingface.co/bigscience/tr3-1B3-modeling-baseline-tensorboard) | [b](https://huggingface.co/bigscience/tr3b-760M-modeling-baseline-tensorboard) | [c](https://huggingface.co/bigscience/tr3c-350M-modeling-baseline-tensorboard) | |
| OSCAR + low warmup | [f](https://huggingface.co/bigscience/tr3f-1B3-diagnostic2-low-warmup-oscar-tensorboard) | | | |
| C4 + high warmup | [e](https://huggingface.co/bigscience/tr3e-1B3-diagnostic1-warmup-c4-tensorboard) | | | |
| OSCAR + high warmup | **[d (current baseline)](https://huggingface.co/bigscience/tr3d-1B3-more-warmup-tensorboard)** | [g](https://huggingface.co/bigscience/tr3g-760M-v2-tensorboard) | [h](https://huggingface.co/bigscience/tr3h-350M-v2-tensorboard) | [i](https://huggingface.co/bigscience/tr3i-125M-v2-tensorboard) |
| Pile + high warmup | [m](https://huggingface.co/bigscience/tr3m-1B3-pile-tensorboard) | [j](https://huggingface.co/bigscience/tr3j-760M-pile-tensorboard) | [k](https://huggingface.co/bigscience/tr3k-350M-pile-tensorboard) | [l](https://huggingface.co/bigscience/tr3l-125M-pile-tensorboard) |
### Train 8
104B - unmodified Megatron gpt2 - with extra-wide hidden size to learn how to deal with training instabilities
* [the full spec and discussions](./train/tr8-104B-wide)
* [the training script](./train/tr8-104B-wide/tr8-104B.slurm)
* checkpoints and logs:
- [tensorboard](https://huggingface.co/bigscience/tr8-104B-logs/tensorboard)
- [logs](https://huggingface.co/bigscience/tr8-104B-logs/tree/main/logs)
* [chronicles](./train/tr8-104B-wide/chronicles.md)
You can watch the training logs live by running this `tail -f` like script over remote log file that gets synced to the hub once an hour:
```
perl -e '$u=shift; $b=0; while(1){($e)=qx[Excerpt of 5,117 characters
Read on GitHubStas Bekman · Stasosphere Online Inc. / · Canada
713
Niklas
178
Teven · Mistral AI · France
116
Thomas Wang · MistralAI
77
Lucile Saulnier · Hugging Face
32
14
Iz Beltagy · spiffy.ai
11
Victor SANH · @rilixai · United States
4
3
Max Ryabinin
2
Younes B
2
Nouamane Tazi · @huggingface · France
1
1
1
Thomas Wolf · @huggingface
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:bfeef70c9f36b9a6, topic:training
matched fp:bfeef70c9f36b9a6, topic:nlp