Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Distributed AI Model Training and LLM Fine-Tuning on Kubernetes
| Date | Stars |
|---|---|
| 2026-07-24 | 2153 |
| 2026-07-25 | 2153 |
| 2026-07-28 | 2153 |
| 2026-07-30 | 2153 |
| 2026-07-31 | 2163 |
| 2026-08-06 | 2163 |
Today
— stars today
This week
+10 stars this week
This month
— stars this month
Momentum
55.0
growth rate 0.46%/day
# Kubeflow Trainer
[](https://www.kubeflow.org/docs/about/community/#kubeflow-slack-channels)
[](https://coveralls.io/github/kubeflow/trainer?branch=master)
[](https://goreportcard.com/report/github.com/kubeflow/trainer)
[](https://www.bestpractices.dev/projects/10435)
[](https://deepwiki.com/kubeflow/trainer)
[](https://app.fossa.com/projects/git%2Bgithub.com%2Fkubeflow%2Ftrainer?ref=badge_shield)
<h1 align="center">
<img src="./docs/images/trainer-logo.svg" alt="logo" width="200">
<br>
</h1>
Latest News 🔥
- [2026/03] Kubeflow Trainer v2.2 is officially released with support for JAX and XGBoost
Training Runtimes, enhanced observability with metrics propagation to TrainJob status,
and Flux Framework integration for HPC and MPI workloads. Check out
[the blog post announcement](https://blog.kubeflow.org/kubeflow-trainer-v2.2-release/).
- [2025/11] Kubeflow Trainer v2.1 is officially released with support of
[Distributed Data Cache](https://trainer.kubeflow.org/en/latest/user-guides/data-cache.html),
topology aware scheduling with Kueue and Volcano, and LLM post-training enhancements. Check out
[the GitHub release notes](https://github.com/kubeflow/trainer/releases/tag/v2.1.0).
- [2025/09] Kubeflow SDK v0.1 is officially released with support for CustomTrainer,
BuiltinTrainer, and local PyTorch execution. Check out
[the GitHub release notes](https://github.com/kubeflow/sdk/releases/tag/0.1.0).
- [2025/07] PyTorch on Kubernetes: Kubeflow Trainer Joins the PyTorch Ecosystem. Find the
announcement in [the PyTorch blog post](https://pytorch.org/blog/pytorch-on-kubernetes-kubeflow-trainer-joins-the-pytorch-ecosystem/).
<details>
<summary>More</summary>
- [2025/07] Kubeflow Trainer v2.0 has been officially released. Check out
[the blog post announcement](https://blog.kubeflow.org/trainer/intro/) and [the
release notes](https://github.com/kubeflow/trainer/releases/tag/v2.0.0).
- [2025/04] From High Performance Computing To AI Workloads on Kubernetes: MPI Runtime in
Kubeflow TrainJob. See the [KubeCon + CloudNativeCon London talk](https://youtu.be/Fnb1a5Kaxgo)
</details>
## Overview
Kubeflow Trainer is a Kubernetes-native distributed AI platform for scalable large language model
(LLM) fine-tuning and training of AI models across a wide range of frameworks, including
PyTorch, MLX, HuggingFace, DeepSpeed, JAX, XGBoost, and more.
Kubeflow Trainer brings MPI to Kubernetes, orchestrating multi-node, multi-GPU distributed
jobs efficiently across high-performance computing (HPC) clusters. This enables high-throughput
communication between processes, making it ideal for large-scale AI training that requires
ultra-fast synchronization between GPUs nodes.
Kubeflow Trainer seamlessly integrates with the Cloud Native AI ecosystem, including
[Kueue](https://kueue.sigs.k8s.io/docs/tasks/run/trainjobs/) for topology-aware scheduling and
multi-cluster job dispatching, as well as [JobSet](https://github.com/kubernetes-sigs/jobset) and
[LeaderWorkerSet](https://github.com/kubernetes-sigs/lws) for AI workload orchestration.
Kubeflow Trainer provides a distributed data cache designed to stream large-scale data with zero-copy
transfer directly to GPU nodes. This ensures memory-efficient training jobs while maximizing
GPU utilization.
With [the Kubeflow Python SDK](https://github.com/kubeflow/sdk), AI practitioners can effortlessly
develop and fine-tune LLMs while leveraging the Kubeflow Trainer APIs: TrainJob and Runtimes.
<h1 align="center">
<img srExcerpt of 7,106 characters
Read on GitHub312
Andrey Velichkevich · Apple · United Kingdom
171
Ce Gao · @TensorChord
122
Yuki Iwai · @pfnet · Japan
122
Jeremy Lewi
118
Johnu George · @Nutanix · India
52
Richard Liu · Google
45
Antonin Stefanutti · Red Hat · France
45
Jiaxin Shan · Bytedance · United States
32
Shao Wang · Shanghai Jiao Tong University. · China
28
27
Yuan Tang · Red Hat · United States
27
23
19
17
Jiayu Liu · Singapore
14
14
13
11
10
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:defa097511067760, topic:pytorch, topic:tensorflow, topic:jax
matched fp:defa097511067760, topic:mlops, topic:kubeflow
matched fp:defa097511067760, topic:gpu, topic:kubernetes, readme:multi-node
matched fp:defa097511067760, topic:fine-tuning, desc:fine-tuning, readme:fine-tuning