taketwo/llm-ollama
quality grade A, 80 out of 100LLM plugin providing access to models running on an Ollama server
- stars
- 370
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
LLM plugin providing access to models running on an Ollama server
A simple TUI for serving local LLM models. Pick a model, pick a backend, serve it
High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI
Access to Anthropic's safety-first language model APIs in C#
Kubernetes-native AI serving platform for scalable model serving.
原汁原昧 Claude Code 可运行,可构建, 可调试版; 生产级工程化, 企业级可靠性; 安全无毒, 内存泄露修复
WilmerAI is one of the oldest LLM semantic routers. It uses multi-layer prompt routing and complex workflows to allow you to not only create practical chatbots, but to extend any kind of application that connects to an LLM via REST API. Wilmer sits between your app and your many LLM APIs, so that you can manipulate prompts as needed.
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.
A cosy home for your LLMs.
LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices.
Access to Anthropic's safety-first language model APIs via Go
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
A simple and easy-to-use library for interacting with the Ollama API.
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Search and organise images and videos offline with on-device AI.
Tiny, no-nonsense, self-contained, Tensorflow and ONNX inference
📄 Awesome OCR multiple programing languages toolkits based on ONNX Runtime, OpenVINO, MNN, PaddlePaddle, TensorRT and PyTorch.
Local-first browser AI video editor with ONNX voiceovers, Whisper captions, talking avatars, and multi-track MP4/WebM export.
Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
A framework for efficient model inference with omni-modality models
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
24,523 repositories in the index in total.