MegEngine/InferLLM
quality grade D, 44 out of 100a lightweight LLM model inference framework
- stars
- 752
- stars gained this week
- —this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
a lightweight LLM model inference framework
Open-source humanize text toolkit. Documents 4 humanization methodologies with reference implementations, plus a production pipeline combining LLM rewriting with a cross-engine translation chain. Python, OpenAI-compatible API.
Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs
GraphRAG using Local LLMs - Features robust API and multiple apps for Indexing/Prompt Tuning/Query/Chat/Visualizing/Etc. This is meant to be the ultimate GraphRAG/KG local LLM app.
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
Based on Nano-vLLM, a simple replication of vLLM with self-contained paged attention and flash attention implementation
A Home Assistant integration & Model to control your smart home using a Local LLM
No description
计图大模型推理库,具有高性能、配置要求低、中文支持好、可移植等特点
A fast inference library for running LLMs locally on modern consumer-class GPUs
AI-powered penetration testing assistant using local LLM on linux (Parrot OS)
Fast inference from large lauguage models via speculative decoding
闻达:一个LLM调用平台。目标为针对特定环境的高效内容生成,同时考虑个人和中小企业的计算资源局限性,以及知识安全和私密性问题
LLMRouter: An Open-Source Library for LLM Routing
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
Official inference framework for 1-bit LLMs
A framework for serving and evaluating LLM routers - save LLM costs without compromising quality
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
Example code and documentation on how to get Stable Diffusion running with ONNX FP16 models on DirectML. Can run accelerated on all DirectML supported cards including AMD and Intel.
WeDLM: The fastest diffusion language model with standard causal attention and native KV cache compatibility, delivering real speedups over vLLM-optimized baselines.
Analyze the inference of Large Language Models (LLMs). Analyze aspects like computation, storage, transmission, and hardware roofline model in a user-friendly interface.
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
Disaggregated serving system for Large Language Models (LLMs).
General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.
24,523 repositories in the index in total.