royshil/obs-backgroundremoval
quality grade A, 84 out of 100An OBS plugin for removing background in portrait images (video), making it easy to replace the background when recording or streaming.
- stars
- 4.6k
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
An OBS plugin for removing background in portrait images (video), making it easy to replace the background when recording or streaming.
Turn Antigravity / codex / github copilot into Anthropic & Openai API compatible server. Usable with Claude Code / Xcode etc.
ONNX-TensorRT: TensorRT backend for ONNX
Open source real-time translation app for Android that runs locally
Fast ML inference & training for ONNX models in Rust
sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRgv6ZD
A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/iOS/Linux with full GPU capability
🐈 A collection of LLM inference providers and models
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Pure Rust Inference Engine
ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60
The production framework for Predictive and Generative AI. Serve any model as an API in one line, with OpenAI/Anthropic/Ollama-compatible endpoints, a built-in chat UI, and native MCP.
Communicate with an LLM provider using a single interface
High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
The easiest way to build and deploy an agent
Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
A cosy home for your LLMs.
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
Official inference framework for 1-bit LLMs
24,535 repositories in the index in total.