Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.
| Date | Stars |
|---|---|
| 2026-07-31 | 1586 |
| 2026-08-03 | 1587 |
| 2026-08-06 | 1587 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
## Quick Start | | Megatron-Core | ChatLearn | verl | |:------------|:------------------------------------------------------------------------------------------------------------------------:|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------:|:-----------:| |Qwen3-Omni |[ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/qwen3_omni/README.md)| N/A | Coming Soon | |Qwen3-Next |[ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/qwen3_next/README.md)| [ReadMe](https://github.com/alibaba/ChatLearn/blob/main/docs/zh/tutorial/tutorial_grpo_mcore_qwen3_next.md) | Coming Soon | |Qwen3 |[ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/qwen3/README.md)|[ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/qwen3/README_chatlearn.md) | [ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/qwen3/README_verl.md) | |Qwen3-VL | [ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/qwen3_vl/README.md)| N/A | Coming Soon | |Qwen2.5-VL |[ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/qwen2_5_vl/README.md)| [ReadMe](https://github.com/alibaba/ChatLearn/blob/main/docs/zh/tutorial/tutorial_grpo_mcore_qwenvl.md) | N/A | |Moonlight |[ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/moonlight/README.md)|[ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/moonlight/README_chatlearn.md)| [ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/moonlight/README_verl.md) | |DeepSeek-V3 |[ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/deepseek_v3/README.md)| N/A | N/A | |DeepSeek-R1 | N/A |[ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/deepseek_v3/README_chatlearn.md)| [ReadMe](https://github.com/alibaba/Pai-Megatron-Patch/blob/main/examples/deepseek_v3/README_verl.md) | ## Introduction English | [简体中文](./README_zh-CN.md) Pai-Megatron-Patch (https://github.com/alibaba/Pai-Megatron-Patch) is a deep learning training toolkit built for developers to train and predict LLMs & VLMs by using Megatron framework easily. With the continuous development of LLMs, the model structure and scale are rapidly evolving. Although these models can be conveniently manufactured using Transformers or DeepSpeed training framework, the training efficiency is comparably low. This phenomenon becomes even severer when the model scale exceeds 10 billion. The primary objective of Pai-Megatron-Patch is to effectively utilize the computational power of GPUs for LLM. This tool allows convenient training of commonly used LLM with all the accelerating techniques provided by Megatron-LM. What's New: - **[Experimental]Support Qwen3-Omni-thinker SFT using Megatron-Core** [🔥🔥 2025.11.12] - **Support Qwen3-Next-80B-A3B Reinforcement Training using Megatron-Core and ChatLearn** [🔥🔥 2025.10.17] - **Support Qwen3-VL series Pre-Training using Megatron-Core** [🔥🔥 2025.10.17] - **Improve MLA models such as Moonlight/DeepSeek-V3 RL Training Stability and Efficiency with Context Parallel and Sequence Packing** [🔥🔥 2025.10.10] - **[Experimental]Support Qwen3-Next-80B-A3B Pre-Training using Megatron-Core** [🔥🔥 2025.09.22] - **Support Qwen3 & DeepSeek-R1 GRPO Reinforcement Training using Megatron-Core and Verl** [🔥🔥 2025.09.19] - **Support Moonlight GRPO Reinforcement Training using Megatron-Core and Verl** [🔥🔥 2025.09.11] - **Support Verl smoothly load distributed checkpoints from
Excerpt of 13,066 characters
Read on GitHub198
58
41
4
3
2
2
2
1
1
1
1
1
1
1
1
1
1
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:cab4afdc12fdcab1, llm:description: 'The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.' (large scale training for LLM & VLM)
matched fp:cab4afdc12fdcab1, llm:description: 'The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.' (large scale training for LLM & VLM)
matched fp:cab4afdc12fdcab1, llm:description: 'The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.' (large scale training for LLM & VLM)
matched fp:cab4afdc12fdcab1, llm:description: 'The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.' (large scale training for LLM & VLM)