xlite-dev/lite.ai.toolkit
quality grade C, 51 out of 100A lite C++ AI toolkit: 100+ models with MNN, ORT and TRT, including Det, Seg, Stable-Diffusion, Face-Fusion.
- stars
- 4.4k
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
A lite C++ AI toolkit: 100+ models with MNN, ORT and TRT, including Det, Seg, Stable-Diffusion, Face-Fusion.
PyTorch ,ONNX and TensorRT implementation of YOLOv4
Tengine is a lite, high performance, modular inference engine for embedded device
An easy to use PyTorch to TensorRT converter
🏋️ A unified multi-backend utility for benchmarking Transformers, Timm, PEFT, Diffusers and Sentence-Transformers with full support of Optimum's hardware optimizations & quantization schemes.
DAWNBench: An End-to-End Deep Learning Benchmark and Competition
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).
T-GATE: Temporally Gating Attention to Accelerate Diffusion Model for Free!
optimized BERT transformer inference on NVIDIA GPU. https://arxiv.org/abs/2210.03052
⚡ boost inference speed of T5 models by 5x & reduce the model size by 3x.
AICI: Prompts as (Wasm) Programs
🍅🍅🍅YOLOv5-Lite: Evolved from yolov5 and the size of model is only 900+kb (int8) and 1.7M (fp16). Reach 15 FPS on the Raspberry Pi 4B~
This repository deploys YOLOv4 as an optimized TensorRT engine to Triton Inference Server
Add bisenetv2. My implementation of BiSeNet
Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
多个SVC/TTS的C++推理库
TTS with kokoro and onnx runtime
Superduper: End-to-end framework for building custom AI applications and agents.
LLM 并发性能测试工具,支持自动化压力测试和性能报告生成。
Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)
Implementation for FP8/INT8 Rollout for RL training without performence drop.
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
Low latency JSON generation using LLMs ⚡️
TopicGPT: A Prompt-Based Framework for Topic Modeling [NAACL'24]
24,523 repositories in the index in total.