Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Awesome LLM pre-training resources, including data, frameworks, and methods.
| Date | Stars |
|---|---|
| 2026-07-31 | 395 |
| 2026-08-06 | 395 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome LLM Pre-training [中文版](https://github.com/RUCAIBox/awesome-llm-pretraining/blob/main/README_ZH.md) | **English Version** Pre-training is the first and most crucial training stage in the development of large language models. As the open-source community continues to improve in areas such as model architecture, training strategies, open-source datasets, and data methods, we are committed to continuously tracking resources available for large model pre-training to give back to developers in the open-source large language model community. Compared to comprehensive reviews, our scope is limited to commonly used resources and cutting-edge attempts related to pre-training, aiming to help users quickly get started with large language model pre-training. We also welcome contributions and updates from the open-source community to jointly promote the development of large models. > Related project links: [[LLMSurvey](https://github.com/RUCAIBox/LLMSurvey)] [[YuLan-Chat](https://github.com/RUC-GSAI/YuLan-Chat)] | [[YuLan-Mini](https://github.com/RUC-GSAI/YuLan-Mini)] ## Table of Contents - [Technical Reports](#i-technical-reports) - [Training Strategies](#ii-training-strategies) - [Open-source Datasets](#iii-open-source-datasets) - [Data Methods](#iv-data-methods) ## I. Technical Reports Technical reports often rely on hundreds or thousands of computing resources. Therefore, it is highly recommended to read some open-source technical reports. ### 1.1 Dense Models 1. **The Llama 3 Herd of Models**. [[paper](https://arxiv.org/abs/2407.21783)] 2. **Qwen2.5 Technical Report**. [[paper](https://arxiv.org/abs/2412.15115)] 3. **Gemma 3 Technical Report**. [[paper](https://arxiv.org/abs/2503.19786)] 4. **Nemotron-4 340B Technical Report**. [[paper](https://arxiv.org/abs/2406.11704)] 5. **Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs**. [[paper](https://arxiv.org/abs/2504.07866)] 6. **Baichuan 2: Open Large-scale Language Models**. [[paper](https://arxiv.org/abs/2309.10305)] ### 1.2 MoE Models 1. **DeepSeek-V3 Technical Report**. [[paper](https://arxiv.org/abs/2412.19437)] 2. **DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models**. [[paper](https://arxiv.org/abs/2401.06066)] 3. **Mixtral of Experts**. [[paper](https://arxiv.org/abs/2401.04088)] 4. **Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models**. [[paper](https://arxiv.org/abs/2406.06563)] 5. **Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs**. [[paper](https://arxiv.org/abs/2503.05139)] 6. **OLMoE: Open Mixture-of-Experts Language Models**. [[paper](https://arxiv.org/abs/2409.02060)] 7. **Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent**. [[paper](https://arxiv.org/abs/2411.02265)] ### 1.3 Models with Open-source Datasets 1. **YuLan-Mini: An Open Data-efficient Language Model**. [[code](https://github.com/RUC-GSAI/YuLan-Mini?tab=readme-ov-file)] [[resource](https://huggingface.co/collections/yulan-team/yulan-mini-676d214b24376739b00d95f3)] [[paper](https://arxiv.org/abs/2412.17743)] 2. **MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series**. [[paper](https://arxiv.org/abs/2405.19327)] 3. **LLM360: Towards Fully Transparent Open-Source LLMs**. [[paper](https://arxiv.org/abs/2312.06550)] 4. **Nemotron-4 15B Technical Report**. [[paper](https://arxiv.org/abs/2402.16819)] ### 1.4 Training/Data Strategies 1. **Phi-4 Technical Report**. [[paper](https://arxiv.org/abs/2412.08905)] 2. **OLMo: Accelerating the Science of Language Models**. [[paper](https://arxiv.org/abs/2402.00838)] 3. **2 OLMo 2 Furious**. [[paper](https://arxiv.org/abs/2501.00656)] 4. **Yi: Open Foundation Models by 01.AI**. [[paper](https://arxiv.org/abs/2403.04652)] 5. **MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies**. [[paper](https://arxiv.org/abs/2404
Excerpt of 29,612 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:928a4b8cf38d9656, name:pretraining, desc:pre-training