liguodongiot/llm-action
quality grade C, 63 out of 100本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
- stars
- 25k
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
Nvidia GPU exporter for prometheus using nvidia-smi binary OR using NVML
QualityScaler - image/video AI upscaler app
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.
Voice activity detector (VAD) for the browser with a simple API
Large-scale LLM inference engine
Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
LLM plugin providing access to models running on an Ollama server
[EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.
xDiT: A Scalable Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
Self-hosted huggingface mirror service. 自建huggingface镜像服务。
The Qualcomm® AI Hub apps are a collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
Use Hugging Face with JavaScript
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.
Qualcomm® AI Hub Models is our collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.
A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.
Advanced CLI tool that scans your hardware and tells you exactly which LLM or sLLM models you can run locally, with full Ollama integration.
AI-powered offensive security testing using autonomous agents, directly in your terminal.
Rust + CUDA LLM inference engine for Blackwell (Tuned specifically on RTX PRO 6000, RTX 5090, B200): OpenAI-compatible (+converse and ant) serving, per-model X hardware exactness gates. NVFP4/mixed (fp8 hybrid, 4o6, etc - correctness, performance, hardware specific adapted) main quant support.
Ollama Python library
Access to Anthropic's safety-first language model APIs via Go
Deep Learning Streamer (DL Streamer) Pipeline Framework is an open-source streaming media analytics framework, based on GStreamer* multimedia framework, for creating complex media analytics pipelines for the Cloud or at the Edge.
General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.
24,535 repositories in the index in total.