Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
vLLM Kunlun (vllm-kunlun) is a community-maintained hardware plugin designed to seamlessly run vLLM on the Kunlun XPU.
| Date | Stars |
|---|---|
| 2026-07-31 | 456 |
| 2026-08-06 | 459 |
Today
+3 stars today
This week
— stars this week
This month
— stars this month
Momentum
17.0
growth rate 0.00%/day
 <p align="center"> <a href="https://vllm-kunlun.readthedocs.io/en/latest/"><b>📖 Documentation</b></a> | <a href="https://vllm-kunlun.readthedocs.io/en/latest/quick_start.html"><b>🚀 Quick Start</b></a> | <a href="https://vllm-kunlun.readthedocs.io/en/latest/installation.html"><b>📦 Installation</b></a> | <a href="https://join.slack.com/t/vllm-kunlun/shared_invite/zt-3iinb8u5z-FcqZKbNNdMJ_32fHmipzvw"><b>💬 Slack</b></a> </p> <p align="center"> <img alt="GitHub License" src="https://img.shields.io/github/license/baidu/vLLM-Kunlun"> <img alt="GitHub Stars" src="https://img.shields.io/github/stars/baidu/vLLM-Kunlun"> <img alt="GitHub Forks" src="https://img.shields.io/github/forks/baidu/vLLM-Kunlun"> <img alt="GitHub Issues" src="https://img.shields.io/github/issues/baidu/vLLM-Kunlun"> <img alt="Python Version" src="https://img.shields.io/badge/python-%3E%3D3.10-blue"> </p> --- ## Latest News 🔥 - [2026/07] 🚧 **v0.25.1 under development** — Added Qwen3.5 / Qwen3.5-MoE, Gemma4 (text and multimodal), GLM MoE DSA, and DFlash speculative decoding - [2026/02] ⚡ **Performance optimizations** — Fused MoE with small batches, optimized attention metadata building, Multi-LoRA inference achieves 80%+ of non-LoRA performance - [2026/02] 🔧 **DeepSeek-V3.2 MTP support** — Added MTP (Multi-Token Prediction) for DeepSeek-V3.2, with RoPE and decoding stage kernel optimizations - [2026/01] 🔢 **New quantization methods** — Support for compressed-tensors W4A16, AWQ MoE W4A16, and DeepSeek-V3.2 W8A8 quantization - [2026/01] 🛠️ **CI/CD overhaul** — Added E2E tests, unit test CI, ruff format checks, and modular CI workflow refactoring - [2025/12] 🎉 **v0.11.0 released** — Added Qwen3-Omni, Qwen3-Next, Seed-OSS support ([Release Notes](https://github.com/baidu/vLLM-Kunlun/releases/tag/v0.11.0)) - [2025/12] 📦 **v0.10.1.1 released** — 5+ multimodal models, AWQ/GPTQ quantization for dense models, Piecewise Kunlun Graph, vLLM V1 engine, Flash-Infer Top-K/Top-P sampling with 10-100× speedup ([Release Notes](https://github.com/baidu/vLLM-Kunlun/releases/tag/v0.10.1.1)) - [2025/12] 🌟 Initial release of vLLM Kunlun — Open sourced on Dec 8, 2025 --- ## Overview **vLLM Kunlun** (`vllm-kunlun`) is a community-maintained hardware plugin designed to seamlessly run [vLLM](https://github.com/vllm-project/vllm) on the **Kunlun XPU**. It is the recommended approach for integrating the Kunlun backend within the vLLM community, adhering to the principles outlined in the [RFC Hardware Pluggable](https://github.com/vllm-project/vllm/issues/11162). This plugin provides a hardware-pluggable interface that decouples the integration of the Kunlun XPU with vLLM. By utilizing vLLM Kunlun, popular open-source models — including Transformer-like, Mixture-of-Expert (MoE), Embedding, and Multi-modal LLMs — can run effortlessly on the Kunlun XPU. ### ✨ Key Features - **Seamless Plugin Integration** — Works as a standard vLLM platform plugin via Python entry points, no need to modify vLLM source code - **Broad Model Support** — Supports 20+ mainstream LLMs including Qwen, Llama, DeepSeek, GLM, Gemma4, Kimi-K2, and multimodal models - **Quantization Support** — W8A8 (INT8), AWQ, GPTQ, and compressed-tensors W4A16 for MoE and dense models - **LoRA Fine-Tuning** — LoRA and Multi-LoRA adapter support for Qwen series models - **Piecewise Kunlun Graph** — Hardware-accelerated graph optimization for high-performance inference - **FlashMLA Attention** — Optimized multi-head latent attention for DeepSeek MLA architectures - **Speculative Decoding** — MTP (Multi-Token Prediction) for DeepSeek-V3.2 and DFlash/EAGLE-style proposers - **Tensor Parallelism** — Multi-device parallel inference with distributed execution support - **OpenAI-Compatible API** — Serve models with the standard OpenAI API interface --- ## Prerequisites - **Hardware**: Kunlun3 P800 - **OS**: Ubuntu 20.04 - **Software**: - Python >
Excerpt of 10,645 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:6122939fdeffcc9b, llm:description: 'vLLM Kunlun (vllm-kunlun) is a community-maintained hardware plugin designed to seamlessly run vLLM on the Kunlun XPU.'
matched fp:6122939fdeffcc9b, llm:description: 'vLLM Kunlun (vllm-kunlun) is a community-maintained hardware plugin designed to seamlessly run vLLM on the Kunlun XPU.'
matched fp:6122939fdeffcc9b, llm:description: 'vLLM Kunlun (vllm-kunlun) is a community-maintained hardware plugin designed to seamlessly run vLLM on the Kunlun XPU.'