jamesob/local-llm
quality grade C, 54 out of 100Everything I know about running LLMs locally
- stars
- 1.7k
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
412 results
Everything I know about running LLMs locally
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
The easiest way to build and deploy an agent
Your AI second brain. Self-hostable. Get answers from the web or your docs. Build custom agents, schedule automations, do deep research. Turn any online or local LLM into your personal, autonomous AI (gpt, claude, gemini, llama, qwen, mistral). Get started - free.
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
memra — from-scratch LLM inference engine for NVIDIA RTX 50-series (Rust + CUDA)
Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer.
High-performance MLX-based LLM inference engine for macOS with native Swift implementation
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
🐈 A collection of LLM inference providers and models
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Qualcomm® AI Hub Models is our collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.
Self-hosted AI gateway for coding CLIs — one OpenAI/Claude/Gemini/Codex-compatible endpoint, with a multi-tenant web console, request logs, and spend quotas.
Android 17 local LLM prototype with Jetpack Compose and ONNX Runtime for offline AI inference experiments.
LLM Client, Server API and UI
LLM inference in C/C++
Cross-Platform Production-ready C++ inference engine for YOLO models (v5-v12, YOLO26). Unified API for detection, segmentation, pose estimation, OBB, and classification. Built on ONNX Runtime and OpenCV. Optimized for CPU/GPU with quantization support.
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
Large-scale LLM inference engine
The Qualcomm® AI Hub apps are a collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.
Use Hugging Face with JavaScript
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
24,520 repositories in the index in total.