padeoe/hf-mirror-site
quality grade F, 30 out of 100a huggingface mirror site.
- stars
- 338
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
a huggingface mirror site.
Turn Antigravity / codex / github copilot into Anthropic & Openai API compatible server. Usable with Claude Code / Xcode etc.
Optimized Ollama LLM server configuration for Mac Studio and other Apple Silicon Macs. Headless setup with automatic startup, resource optimization, and remote management via SSH.
Access models from OpenAI, Groq, local Ollama, and others by setting llm-router as Cursor's Base URL
Vercel AI Provider for running LLMs locally using Ollama
Run full Kimi K3 on a single device. And an OpenAI-compatible API server for local chat and coding agents.
Terraform Skill for Claude Code and Codex. LLMs hallucinate a lot with Terraform - TerraShark fixes this. It eliminates hallucinations, is designed for modular and secure code and grounds your IaC in the official Hashicorp Terraform best practices.
Curated collection of AI inference engineering resources — LLM serving, GPU kernels, quantization, distributed inference, and production deployment. Compiled from the AER Labs community.
[ICLR2025 Spotlight] MagicPIG: LSH Sampling for Efficient LLM Generation
Building the Virtuous Cycle for AI-driven LLM Systems
Work with LLMs on a local environment using containers
Plugin for LLM adding support for the GPT4All collection of models
Demonstration of running a native LLM on Android device.
Minimal yet performant LLM examples in pure JAX
No description
TheBloke's Dockerfiles
⚡ Edgen: Local, private GenAI server alternative to OpenAI. No GPU required. Run AI models locally: LLMs (Llama2, Mistral, Mixtral...), Speech-to-text (whisper) and many others.
fastLLaMa: An experimental high-performance framework for running Decoder-only LLMs with 4-bit quantization in Python using a C/C++ backend.
一键运行 Qwen2.5 SakuraLLM 等本地 LLM 模型
Self-hosted LLM router with a managed safety net. OpenAI-compatible. BYOK. Single-workspace. Streaming. For more advanced routing choose hosted OrcaRouter
Inference for hybrid LLMs: Gemma, RWKV, and all kinds of hybrids.
A multi-platform SwiftUI frontend for running local LLMs with Apple's MLX framework.
The easiest way to run the fastest MLX-based LLMs locally
24,523 repositories in the index in total.