inclusionAI/dInfer
quality grade C, 51 out of 100dInfer: An Efficient Inference Framework for Diffusion Language Models
- stars
- 482
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
dInfer: An Efficient Inference Framework for Diffusion Language Models
Fine-tune and run LLMs locally on your M-series Mac. A powerful desktop interface built on Apple's MLX framework for zero-setup AI.
vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon
[ICML 2025 Spotlight] ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
Efficient LLM Inference over Long Sequences
KV cache store for distributed LLM inference
LLM Inference benchmark
No description
DashInfer is a native LLM inference engine aiming to deliver industry-leading performance atop various hardware architectures, including CUDA, x86 and ARMv9.
Speech-to-Text interface for Emacs using OpenAI's whisper model and whisper.cpp as inference engine.
InfiniTensor is a high-performance inference engine tailored for GPUs and AI accelerators. Its design focuses on effective deployment and swift academic validation.
High-efficiency LLM inference engine in C++/CUDA. Run Llama 70B on RTX 3090.
FlexAttention based, minimal vllm-style inference engine for fast Gemma 2 inference.
a fast cross platform AI inference engine 🤖 using Rust 🦀 and WebGPU 🎮
Inference Engine samples internal development repository. Contains example and template projects for Sentis package use.
~950 line, minimal, extensible LLM inference engine built from scratch.
Stock inference engine using Spring XD, Apache Geode / GemFire and Spark ML Lib.
LLM inference engine written in .NET
[ACL 2026] Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
CLI for running large numbers of coding agents in parallel with git worktrees
OpenAlpha_Evolve is an open-source Python framework inspired by the groundbreaking research on autonomous coding agents like DeepMind's AlphaEvolve.
Self-evaluating interview for AI coders
DEEPPOWERS is a Fully Homomorphic Encryption (FHE) framework built for MCP (Model Context Protocol), aiming to provide end-to-end privacy protection and high-efficiency computation for the upstream and downstream ecosystem of the MCP protocol.
A tool for generating function arguments and choosing what function to call with local LLMs
24,535 repositories in the index in total.