predibase/lorax
quality grade B, 72 out of 100Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
- stars
- 3.8k
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
Voice activity detector (VAD) for the browser with a simple API
QualityScaler - image/video AI upscaler app
Highly optimized inference engine for Binarized Neural Networks
Run TensorFlow models in C++ without installation and without Bazel
Bolt is a deep learning library with high performance and heterogeneous flexibility.
The challenge projects for Inferencing machine learning models on iOS
Use AnimeGANv3 to make your own animation works, including turning photos or videos into anime.
GUI for upscaling ONNX models with NVIDIA TensorRT and Vapoursynth
The project can achieve FCWS, LDWS, and LKAS functions solely using only visual sensors. using YOLOv5 / YOLOv5-lite / YOLOv6 / YOLOv7 / YOLOv8 / YOLOv9 / EfficientDet and Ultra-Fast-Lane-Detection-v2 .
volksdep is an open-source toolbox for deploying and accelerating PyTorch, ONNX and TensorFlow models with TensorRT.
🚀 Deep learning includes superpoint-superglue(C++, TensorRT), and traditional algorithms include zkaze, surf, ORB, etc.
基于CenterNet训练的目标检测&人脸对齐&姿态估计模型
C++ Helper Class for Deep Learning Inference Frameworks: TensorFlow Lite, TensorRT, OpenCV, OpenVINO, ncnn, MNN, SNPE, Arm NN, NNabla, ONNX Runtime, LibTorch, TensorFlow
Reimplement RetinaFace use C++ and TensorRT
Based on tensorrt v8.0+, deploy detection, pose, segment, tracking of YOLO11 with C++ and python api.
SuperPoint and SuperGlue with TensorRT. Deploy with C++.
Convert pointpillars Pytorch Model To ONNX for TensorRT Inference
This repository is forked from shouxieai/tensorRT_Pro, extending it to support a wide range of vision models with high-performance TensorRT C++ inference.
A project that optimizes OWL-ViT for real-time inference with NVIDIA TensorRT.
Sample apps to demonstrate how to deploy models trained with TAO on DeepStream
Image classification with NVIDIA TensorRT from TensorFlow models.
Efficient CPU/GPU ML Runtimes for VapourSynth (with built-in support for waifu2x, DPIR, RealESRGANv2/v3, Real-CUGAN, RIFE, SCUNet, ArtCNN and more!)
C++ TensorRT implementation of Depth-Anything V1, V2
24,523 repositories in the index in total.