ollama/ollama-python
quality grade A, 85 out of 100Ollama Python library
- stars
- 10k
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
Ollama Python library
Espressif deep-learning library for AIoT applications
The production framework for Predictive and Generative AI. Serve any model as an API in one line, with OpenAI/Anthropic/Ollama-compatible endpoints, a built-in chat UI, and native MCP.
sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems
CAI NEURAL API - Pascal based deep learning neural network API optimized for AVX, AVX2 and AVX512 instruction sets plus OpenCL capable devices including AMD, Intel and NVIDIA.
LLM plugin to access Google's Gemini family of models
llm-d Router: The intelligent entry point for inference requests
LLMs as Copilots for Theorem Proving in Lean
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Tile-Based Runtime for Ultra-Low-Latency LLM Inference
A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.
NVIDIA AITune is an inference toolkit designed for tuning and deploying Deep Learning models with a focus on NVIDIA GPUs.
Deep Learning Streamer (DL Streamer) Pipeline Framework is an open-source streaming media analytics framework, based on GStreamer* multimedia framework, for creating complex media analytics pipelines for the Cloud or at the Edge.
PyTorch Neural Network eXchange
A high-performance inference engine for AI models
A plug-and-play compiler that delivers free-lunch optimizations for both inference and training.
High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local and remote inference backends.
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.
Dynamic Memory Management for Serving LLMs without PagedAttention
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
List of awesome hosting sorted by minimal plan price
Machine Learning inference engine for Microcontrollers and Embedded devices
State-of-the-art Machine Learning for the web. Run 🤗 Transformers directly in your browser, with no need for a server!
No description
24,523 repositories in the index in total.