containers/podman-desktop-extension-ai-lab
quality grade C, 56 out of 100Work with LLMs on a local environment using containers
- stars
- 300
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
Work with LLMs on a local environment using containers
Plugin for LLM adding support for the GPT4All collection of models
Demonstration of running a native LLM on Android device.
No description
TheBloke's Dockerfiles
⚡ Edgen: Local, private GenAI server alternative to OpenAI. No GPU required. Run AI models locally: LLMs (Llama2, Mistral, Mixtral...), Speech-to-text (whisper) and many others.
fastLLaMa: An experimental high-performance framework for running Decoder-only LLMs with 4-bit quantization in Python using a C/C++ backend.
一键运行 Qwen2.5 SakuraLLM 等本地 LLM 模型
Inference for hybrid LLMs: Gemma, RWKV, and all kinds of hybrids.
A multi-platform SwiftUI frontend for running local LLMs with Apple's MLX framework.
The easiest way to run the fastest MLX-based LLMs locally
👾📦 CodeBoxAPI is the simplest sandboxing infrastructure for your LLM Apps and Services.
port of Andrjey Karpathy's llm.c to Mojo
World's Easiest GPT-like Voice Assistant
No description
ProxyLLM是一款本地 Electron 应用,通过捕获浏览器会话把多种 LLM网站的能力统一为 OpenAI 兼容 API, 一键应用到Claude Code (ProxyLLM is a local Electron application that captures browser sessions to unify the capabilities of multiple LLM websites into an OpenAI-compatible API, enabling one-click integration with Claude Code)
irresponsible innovation. Try now at https://chat.dev/
Since the emergence of chatGPT in 2022, the acceleration of Large Language Model has become increasingly important. Here is a list of papers on accelerating LLMs, currently focusing mainly on inference acceleration, and related works will be gradually added in the future. Welcome contributions!
Instantly calculate the maximum size of quantized language models that can fit in your available RAM, helping you optimize your models for inference.
[AAAI-25] Cobra: Extending Mamba to Multi-modal Large Language Model for Efficient Inference
Multi-platform desktop app to download and run Large Language Models(LLM) locally in your computer.
The official repo of Aquila2 series proposed by BAAI, including pretrained & chat large language models.
Implementation of the RWKV language model in pure WebGPU/Rust.
24,535 repositories in the index in total.