Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
| Date | Stars |
|---|---|
| 2026-07-31 | 7982 |
| 2026-08-02 | 7982 |
| 2026-08-04 | 7990 |
| 2026-08-06 | 7990 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
60.0
growth rate 0.00%/day
<div align="center"> <img src="docs/en/_static/image/lmdeploy-logo.svg" width="450"/> [](https://pypi.org/project/lmdeploy)  [](https://github.com/InternLM/lmdeploy/tree/main/LICENSE) [](https://github.com/InternLM/lmdeploy/issues) [](https://github.com/InternLM/lmdeploy/issues) [📘Documentation](https://lmdeploy.readthedocs.io/en/latest/) | [🛠️Quick Start](https://lmdeploy.readthedocs.io/en/latest/get_started/get_started.html) | [🤔Reporting Issues](https://github.com/InternLM/lmdeploy/issues/new/choose) English | [简体中文](README_zh-CN.md) | [日本語](README_ja.md) 👋 join us on [](https://cdn.vansin.top/internlm/lmdeploy.jpg) [](https://twitter.com/intern_lm) [](https://discord.gg/xa29JuW87d) </div> ______________________________________________________________________ ## Latest News 🎉 <details open> <summary><b>2026</b></summary> - \[2026/04\] PyPI has expanded the storage quota for LMDeploy and wheel uploads have resumed. `v0.12.3` is now available on PyPI, so you can install it directly via `pip install lmdeploy`. - \[2026/02\] Support [Qwen3.5](https://huggingface.co/collections/Qwen/qwen35) - \[2026/02\] Support [vllm-project/llm-compressor](https://github.com/vllm-project/llm-compressor) 4bit symmetric/asymmetric quantization. Refer [here](./docs/en/quantization/llm_compressor.md) for detailed guide </details> <details close> <summary><b>2025</b></summary> - \[2025/09\] TurboMind supports MXFP4 on NVIDIA GPUs starting from V100, achieving 1.5x the performmance of vLLM on H800 for openai gpt-oss models! - \[2025/06\] Comprehensive inference optimization for FP8 MoE Models - \[2025/06\] DeepSeek PD Disaggregation deployment is now supported through integration with [DLSlime](https://github.com/DeepLink-org/DLSlime) and [Mooncake](https://github.com/kvcache-ai/Mooncake). Huge thanks to both teams! - \[2025/04\] Enhance DeepSeek inference performance by integration deepseek-ai techniques: FlashMLA, DeepGemm, DeepEP, MicroBatch and eplb - \[2025/01\] Support DeepSeek V3 and R1 </details> <details close> <summary><b>2024</b></summary> - \[2024/11\] Support Mono-InternVL with PyTorch engine - \[2024/10\] PyTorchEngine supports graph mode on ascend platform, doubling the inference speed - \[2024/09\] LMDeploy PyTorchEngine adds support for [Huawei Ascend](./docs/en/get_started/ascend/get_started.md). See supported models [here](docs/en/supported_models/supported_models.md) - \[2024/09\] LMDeploy PyTorchEngine achieves 1.3x faster on Llama3-8B inference by introducing CUDA graph - \[2024/08\] LMDeploy is integrated into [modelscope/swift](https://github.com/modelscope/swift) as the default accelerator for VLMs inference - \[2024/07\] Support Llama3.1 8B, 70B and its TOOLS CALLING - \[2024/07\] Support [InternVL2](docs/en/multi_modal/internvl.md) full-series models, InternLM-XComposer2.5 and [function call](docs/en/llm/api_server_tools.md) of InternLM2.5 - \[2024/06\] PyTorch engine support DeepSeek-V2 and several VLMs, such as CogVLM2, Mini-InternVL, LlaVA-Next - \[2024/05\] Balance vision model when deploying VLMs with multiple GPUs - \[2024/05\] Support 4-bits weight-only quantization and inference on VLMs, such as InternVL v1.5, LLaVa, InternLMXComposer2 - \[2024/04\] Support Llama3 and more VLMs, such as InternVL v1.1, v1.2, MiniGemini, InternLMXComposer2. - \[2024/04\] TurboMind adds online int8/int4 KV cache qua
Excerpt of 14,835 characters
Read on GitHubLyu Han · China
393
Q.Yao
337
190
Chen Xin
177
RunningLeon · openmmlab
157
Li Zhang
113
zxy · Shanghai AI Lab · China
95
71
Yineng Zhang · United States
40
38
28
24
China
23
China
20
18
16
tangzhiyi11 · China
13
zhaochaoxing
12
Ke Bao
11
10
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:8769bde917a5db28, topic:llm, topic:llama
matched fp:8769bde917a5db28, topic:llm-inference
matched fp:8769bde917a5db28, topic:cuda-kernels
matched fp:8769bde917a5db28, topic:deepspeed