google/XNNPACK
quality grade B, 66 out of 100High-efficiency floating-point neural network inference operators for mobile, server, and Web
- stars
- 2.4k
- stars gained this week
- +6
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
High-efficiency floating-point neural network inference operators for mobile, server, and Web
No description
Turn your PC into a Reef - run OpenClaw, Ollama, and more on your own PC. Private. Local. Open source. Simple.
Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"
Run LLMs with MLX
Samples code for world class Artificial Intelligence SoCs for computer vision applications.
AI Image Toolkit that runs fully in your browser — free, private, and offline-first.
Fast ML inference & training for ONNX models in Rust
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
Community maintained hardware plugin for vLLM on Ascend
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
Open source real-time translation app for Android that runs locally
ONNX neural network inference engine
The Triton TensorRT-LLM Backend
Pre-trained Deep Learning models and demos (high quality and extremely fast)
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
A very fast neural network computing framework optimized for mobile platforms.QQ group: 676883532 【验证信息输:绝影】
use ncnn in Android and iOS, ncnn is a high-performance neural network inference framework optimized for the mobile platform
Apply a Style Transfer Neural Network in real time with Unreal Engine 5 leveraging ONNX Runtime.
FPGA-based neural network inference project with an end-to-end approach (from training to implementation to deployment)
HLS based Deep Neural Network Accelerator Library for Xilinx Ultrascale+ MPSoCs
MCP Server to Use HuggingFace spaces, easy configuration and Claude Desktop mode.
Manage scalable open LLM inference endpoints in Slurm clusters
Gemma 2 optimized for your local machine.
24,523 repositories in the index in total.