Wenyueh/MinivLLM
quality grade B, 71 out of 100Based on Nano-vLLM, a simple replication of vLLM with self-contained paged attention and flash attention implementation
- stars
- 1.0k
- stars gained this week
- +12this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
Based on Nano-vLLM, a simple replication of vLLM with self-contained paged attention and flash attention implementation
TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
State-of-the-art Machine Learning for the web. Run 🤗 Transformers directly in your browser, with no need for a server!
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local and remote inference backends.
A simple TUI for serving local LLM models. Pick a model, pick a backend, serve it
Android 17 local LLM prototype with Jetpack Compose and ONNX Runtime for offline AI inference experiments.
LLM speculative inference server for heterogeneous hardware & consumer GPUs
vLLM Kunlun (vllm-kunlun) is a community-maintained hardware plugin designed to seamlessly run vLLM on the Kunlun XPU.
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
LLMs as Copilots for Theorem Proving in Lean
A unified cross-platform AI client supporting seamless transitions between standard cloud APIs and on-device, offline execution of custom and uncensored language models.
NVIDIA AITune is an inference toolkit designed for tuning and deploying Deep Learning models with a focus on NVIDIA GPUs.
Rust multiprovider generative AI client (Ollama, OpenAi, Anthropic, Gemini, DeepSeek, ZAI, OpenRouter, FireworksAI, xAI/Grok, Groq,, ...)
A simple and easy-to-use library for interacting with the Ollama API.
High-efficiency floating-point neural network inference operators for mobile, server, and Web
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
A plug-and-play compiler that delivers free-lunch optimizations for both inference and training.
High Performance LLM Inference Operator Library
支持按轴凹总力战, 无缝制造三解, 用于实现蔚蓝档案自动化的程序( Steam已适配 )
List of awesome hosting sorted by minimal plan price
A Pure Rust based LLM, VLM, VLA, TTS, OCR Inference Engine, powering by Candle & Rust. Alternate to your llama.cpp but much more simpler and cleaner..
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
24,535 repositories in the index in total.