sybil-solutions/local-studio
quality grade A, 83 out of 100Control panel for VLLM, Sglang, llama.cpp, exllamav3
- stars
- 1.8k
- stars gained this week
- +3this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
Control panel for VLLM, Sglang, llama.cpp, exllamav3
LLM plugin for models hosted by OpenRouter
A Home Assistant integration & Model to control your smart home using a Local LLM
Tile-Based Runtime for Ultra-Low-Latency LLM Inference
LLM plugin to access Google's Gemini family of models
Access to Anthropic's safety-first language model APIs in C#
LLM Client, Server API and UI
Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)
WilmerAI is one of the oldest LLM semantic routers. It uses multi-layer prompt routing and complex workflows to allow you to not only create practical chatbots, but to extend any kind of application that connects to an LLM via REST API. Wilmer sits between your app and your many LLM APIs, so that you can manipulate prompts as needed.
High-performance MLX-based LLM inference engine for macOS with native Swift implementation
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
Turn your images and videos into a personal knowledge base
llm-d Router: The intelligent entry point for inference requests
A RWKV management and startup tool, full automation, only 8MB. And provides an interface compatible with the OpenAI API. RWKV is a large language model that is fully open source and available for commercial use.
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Mirage Persistent Kernel: Compiling LLMs into a MegaKernel
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
A collection of pre-trained, state-of-the-art models in the ONNX format
Cross-Platform Production-ready C++ inference engine for YOLO models (v5-v12, YOLO26). Unified API for detection, segmentation, pose estimation, OBB, and classification. Built on ONNX Runtime and OpenCV. Optimized for CPU/GPU with quantization support.
Lightweight, Modular, Kubernetes-native AI serving platform for scalable model serving.
AI-powered penetration testing assistant using local LLM on linux (Parrot OS)
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
PyTorch Neural Network eXchange
Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs
24,535 repositories in the index in total.