mixa3607/ML-gfx906
quality grade B, 78 out of 100ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60
- stars
- 293
- stars gained this week
- —
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60
LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models.
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Control panel for VLLM, Sglang, llama.cpp, exllamav3
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
Build apps powered by on-device AI
An OBS plugin for removing background in portrait images (video), making it easy to replace the background when recording or streaming.
Cross-platform, customizable ML solutions for live and streaming media.
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
Unified framework for building enterprise RAG pipelines with small, specialized models
LiteRT and LiteRT-LM sample apps, model recipes, agent skills and utilities.
Pure Rust Inference Engine
LLM speculative inference server for consumer hardware & heterogeneous computing
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
Advanced CLI tool that scans your hardware and tells you exactly which LLM or sLLM models you can run locally, with full Ollama integration.
LLM plugin for models hosted by OpenRouter
Implementation of popular deep learning networks with TensorRT network definition API
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
A Rust library integrated with ONNXRuntime, providing a collection of Computer Vison and Vision-Language models such as YOLO, FastVLM, and more.
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
vLLM Kunlun (vllm-kunlun) is a community-maintained hardware plugin designed to seamlessly run vLLM on the Kunlun XPU.
24,523 repositories in the index in total.