Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink
| Date | Stars |
|---|---|
| 2026-07-31 | 696 |
| 2026-08-05 | 783 |
| 2026-08-06 | 783 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# RTX PRO 6000 Blackwell LLM Wiki This repository is a field wiki for running frontier LLMs on NVIDIA RTX PRO 6000 Blackwell / SM120 PCIe systems. It is more than a few launch snippets: it contains reproducible Docker builds, exact vLLM and SGLang runbooks, benchmark tables, KLD quality checks, quantization notes, DCP/MTP/DSpark/DFlash debugging, PCIe topology work, and regression history. > Community workbench for RTX PRO 6000 / SM120 serving: > https://discord.gg/X54jjmcxWJ ## Start Here If you just want to run a model, use these stable hub pages first: | Model family | Start here | Current focus | |---|---|---| | GLM-5.2 | [GLM-5.2 Runbook Hub](models/glm-5.2.md) | Fathomless vLLM, NVFP4, online FP8/MXFP8, B12X, DCP, MTP, KLD. | | DeepSeek-V4-Flash / DSpark | [DeepSeek-V4-Flash Runbook Hub](models/deepseek-v4-flash.md) | Standard checkpoint, MTP, DSpark, B12X, Lucifer, CUTLASS. | | Kimi | [Kimi Runbook Hub](models/kimi.md) | Kimi-K2.7-Code, DFlash, parser/tool-call runtime. | | Xiaomi MiMo | [MiMo Runbook Hub](models/mimo.md) | MiMo V2.5 Pro FP4-DFlash. | | GLM-5.1 | [GLM-5.1 Runbook Hub](models/glm-5.1.md) | Historical GLM-5.1, KLD methodology, older B12X/SGLang work. | | Legacy / secondary models | [Legacy Model Runbooks](models/legacy.md) | DeepSeek-V4-Pro, GLM-4.7, Qwen, MiniMax, older Kimi pages. | Need the complete map of every Markdown page? | Index | Use | |---|---| | [Full Wiki Index](INDEX.md) | Complete generated catalog of every page in this repository. | | [Glossary And Acronym Guide](GLOSSARY.md) | Acronym expansions and writing rules for newcomer-friendly docs. | | [Newcomer Onboarding](docs/newcomer-onboarding.md) | How to ask useful questions without lowering the technical signal. | ## What Is In This Repository? | Need | Where | |---|---| | Copy/paste production launch commands | Model hubs and current versioned model pages. | | Rebuild the Docker image | [Eldritch Docker](models/eldritch-enlightenment-docker.md), current model image sections, and build scripts in [scripts](scripts/). | | Compare backend speed | Model benchmark tables plus [Benchmark Results](benchmarks/results.md). | | Check quantization quality | [GLM-5.2 KLD](benchmarks/glm52-kld-evaluation.md), [KLD Evaluation](benchmarks/kld-evaluation.md), and model-specific KLD sections. | | Understand MTP, DSpark, or DFlash | [Speculative Decoding](optimization/speculative-decoding.md), DS4/Kimi/MiMo pages. | | Debug topology or PCIe behavior | [Topology](hardware/topology.md), [PCIe Bandwidth](hardware/pcie-bandwidth.md), [GPU Configurations](hardware/gpu-configs.md). | | Avoid known runtime footguns | [Common Issues](troubleshooting/common-issues.md), model caveats, and daily summaries. | | Understand old measurements | Historical versioned pages and [Daily Summaries](daily-summaries/). | ## Current Production-Style Pages | Area | Page | Why it matters | |---|---|---| | GLM-5.2 current stack | [GLM-5.2 v20](models/glm5.2_v20.md) | Current Gilded Gnosis/SparkInfer image, source pins, TP6 fixes, Xid validation, and release gate. | | GLM-5.2 MXFP4 | [GLM-5.2 FP8 + MXFP4 Experts](models/glm5.2_mxfp4.md) | Native MXFP4 expert checkpoint path and A8 serving notes. | | DS4 current quick start | [DS4 DSpark Gilded Gnosis r16](models/ds4dspark-v20.md) | 0731 checkpoint, fixed K5, InstantTensor, and native CPU KV offload. | | DS4 full reference | [DS4 DSpark v9](models/ds4dspark-v9.md) | Full DSpark and standard MTP sweep reference. | | Kimi-K2.7-Code | [Kimi-K2.7-Code v3](models/kimi-k27-code_v3.md) | Fathomless Kimi DFlash validation. | | MiMo FP4-DFlash | [MiMo FP4-DFlash v3](models/xiaomi-mimo-v2.5-pro-fp4-dflash_v3.md) | Current MiMo DFlash validation and fix notes. | Older pages are intentionally preserved. Prefer the hub page for each model family unless you are reproducing a specific old result. ## Core Topics | Topic | Page | |---|---| | Docker images and release lines | [Docker Images](optimization/docker-images.md) | | PCIe oneshot
Excerpt of 9,297 characters
Read on GitHub416
7
1
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:3af5ead3a14d9b6f, llm:repository description: 'RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink'
matched fp:3af5ead3a14d9b6f, llm:repository description: 'RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink'
matched fp:3af5ead3a14d9b6f, llm:repository description: 'RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink'