shroominic/codebox-api
quality grade D, 41 out of 100👾📦 CodeBoxAPI is the simplest sandboxing infrastructure for your LLM Apps and Services.
- stars
- 364
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
👾📦 CodeBoxAPI is the simplest sandboxing infrastructure for your LLM Apps and Services.
port of Andrjey Karpathy's llm.c to Mojo
World's Easiest GPT-like Voice Assistant
No description
ProxyLLM是一款本地 Electron 应用,通过捕获浏览器会话把多种 LLM网站的能力统一为 OpenAI 兼容 API, 一键应用到Claude Code (ProxyLLM is a local Electron application that captures browser sessions to unify the capabilities of multiple LLM websites into an OpenAI-compatible API, enabling one-click integration with Claude Code)
irresponsible innovation. Try now at https://chat.dev/
Since the emergence of chatGPT in 2022, the acceleration of Large Language Model has become increasingly important. Here is a list of papers on accelerating LLMs, currently focusing mainly on inference acceleration, and related works will be gradually added in the future. Welcome contributions!
Instantly calculate the maximum size of quantized language models that can fit in your available RAM, helping you optimize your models for inference.
[AAAI-25] Cobra: Extending Mamba to Multi-modal Large Language Model for Efficient Inference
Multi-platform desktop app to download and run Large Language Models(LLM) locally in your computer.
The official repo of Aquila2 series proposed by BAAI, including pretrained & chat large language models.
Implementation of the RWKV language model in pure WebGPU/Rust.
dInfer: An Efficient Inference Framework for Diffusion Language Models
Fine-tune and run LLMs locally on your M-series Mac. A powerful desktop interface built on Apple's MLX framework for zero-setup AI.
vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon
[ICML 2025 Spotlight] ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
Efficient LLM Inference over Long Sequences
KV cache store for distributed LLM inference
Everything you need to know about LLM inference
LLM Inference benchmark
No description
A Pure Rust based LLM, VLM, VLA, TTS, OCR Inference Engine, powering by Candle & Rust. Alternate to your llama.cpp but much more simpler and cleaner..
DashInfer is a native LLM inference engine aiming to deliver industry-leading performance atop various hardware architectures, including CUDA, x86 and ARMv9.
Speech-to-Text interface for Emacs using OpenAI's whisper model and whisper.cpp as inference engine.
24,523 repositories in the index in total.