Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
| Date | Stars |
|---|---|
| 2026-07-24 | 5731 |
| 2026-07-25 | 5731 |
| 2026-07-28 | 5740 |
| 2026-07-30 | 5740 |
| 2026-08-06 | 5773 |
Today
+33 stars today
This week
+33 stars this week
This month
— stars this month
Momentum
180.0
growth rate 0.57%/day
# KServe [](https://pkg.go.dev/github.com/kserve/kserve) [](https://goreportcard.com/report/github.com/kserve/kserve) [](https://bestpractices.coreinfrastructure.org/projects/6643) [](https://github.com/kserve/kserve/releases) [](https://github.com/kserve/kserve/blob/master/LICENSE) [](https://github.com/kserve/community/blob/main/README.md#questions-and-issues) [](https://gurubase.io/g/kserve) KServe is a standardized distributed generative and predictive AI inference platform for scalable, multi-framework deployment on Kubernetes. KServe is being [used by many organizations](https://kserve.github.io/website/docs/community/adopters) and is a [Cloud Native Computing Foundation (CNCF)](https://www.cncf.io/) incubating project. For more details, visit the [KServe website](https://kserve.github.io/website/).  ### Why KServe? Single platform that unifies Generative and Predictive AI inference on Kubernetes. Simple enough for quick deployments, yet powerful enough to handle enterprise-scale AI workloads with advanced features. ### Features **Generative AI** * 🧮 **Optimized Backends**: Support for vLLM and llm-d for optimized performance for serving LLMs * 📌 **Standardization**: OpenAI-compatible inference protocol for seamless integration with LLMs * 🚅 **GPU Acceleration**: High-performance serving with GPU support and optimized memory management for large models * 💾 **Model Caching**: Intelligent model caching to reduce loading times and improve response latency for frequently used models * 🗂️ **KV Cache Offloading**: Advanced memory management with KV cache offloading to CPU/disk for handling longer sequences efficiently * 📈 **Autoscaling**: Request-based autoscaling capabilities optimized for generative workload patterns * 🔧 **Hugging Face Ready**: Native support for Hugging Face models with streamlined deployment workflows **Predictive AI** * 🧮 **Multi-Framework**: Support for TensorFlow, PyTorch, scikit-learn, XGBoost, ONNX, and more * 🔀 **Intelligent Routing**: Seamless request routing between predictor, transformer, and explainer components with automatic traffic management * 🔄 **Advanced Deployments**: Canary rollouts, inference pipelines, and ensembles with InferenceGraph * ⚡ **Autoscaling**: Request-based autoscaling with scale-to-zero for predictive workloads * 🔍 **Model Explainability**: Built-in support for model explanations and feature attribution to understand prediction reasoning * 📊 **Advanced Monitoring**: Enables payload logging, outlier detection, adversarial detection, and drift detection * 💰 **Cost Efficient**: Scale-to-zero on expensive resources when not in use, reducing infrastructure costs ### Learn More To learn more about KServe, how to use various supported features, and how to participate in the KServe community, please follow the [KServe website documentation](https://kserve.github.io/website). Additionally, we have compiled a list of [presentations and demos](https://kserve.github.io/website/docs/community/presentations) to dive through various details. ### :hammer_and_wrench: Installation #### Standalone Installation - **[Standard Kubernetes Installation](https://kserve.github.io/website/docs/admin-guide/overview#raw-kubernetes-deployment)**: Compared to Serverless Installation, this is a more **lightweight** installation. However,
Excerpt of 6,034 characters
Read on GitHub256
208
136
113
105
64
Filippe Spolti · Red Hat, Inc · Brazil
60
Animesh Singh
58
Yuan Tang · Red Hat · United States
56
53
Pierangelo Di Pilato · Red Hat · Italy
52
46
Clive Cox · United Kingdom
40
37
Vivek Karunai Kiri Ragavan · @RedHatOfficial · United States
36
36
Paul Van Eck · United States
34
28
Gang Pu · IBM
25
25
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:abd1b7d52c507ffb, topic:llm-inference, topic:model-serving, topic:vllm
matched fp:abd1b7d52c507ffb, topic:mlops, topic:kubeflow
matched fp:abd1b7d52c507ffb, topic:pytorch, topic:tensorflow