hahnyuan/LLM-Viewer
quality grade D, 40 out of 100Analyze the inference of Large Language Models (LLMs). Analyze aspects like computation, storage, transmission, and hardware roofline model in a user-friendly interface.
- stars
- 681
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
Analyze the inference of Large Language Models (LLMs). Analyze aspects like computation, storage, transmission, and hardware roofline model in a user-friendly interface.
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
Disaggregated serving system for Large Language Models (LLMs).
LLM-PowerHouse: Unleash LLMs' potential through curated tutorials, best practices, and ready-to-use code for custom training and inferencing.
[ICLR 2024] Efficient Streaming Language Models with Attention Sinks
Samples code for world class Artificial Intelligence SoCs for computer vision applications.
The easiest & fastest way to run customized and fine-tuned LLMs locally or on the edge
LLM (Large Language Model) FineTuning
Graph Data Science: an abstraction layer in Python for building knowledge graphs, integrated with popular graph libraries – atop Pandas, NetworkX, RAPIDS, RDFlib, pySHACL, PyVis, morph-kgc, pslpython, pyarrow, etc.
Yet Another Language Model: LLM inference in C++/CUDA, no libraries except for I/O
[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
[ICML 2024] Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
A highly optimized LLM inference acceleration engine for Llama and its variants.
Low-bit LLM inference on CPU/NPU with lookup table
Minimal LLM inference in Rust
A lightweight, portable pure C99 onnx inference engine for embedded devices with hardware acceleration support.
A lightweight inference engine supporting speculative speculative decoding (SSD).
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.
High-performance Text-to-Speech server with OpenAI-compatible API, 8 voices, emotion tags, and modern web UI. Optimized for RTX GPUs.
A package for statistically rigorous scientific discovery using machine learning. Implements prediction-powered inference.
Allows you to run machine learning models locally on your ESP32 device.
A flexible, high-performance carrier for machine learning models(『飞桨』服务化部署框架)
FastAPI Skeleton App to serve machine learning models production-ready.
MXNetJS: Javascript Package for Deep Learning in Browser (without server)
24,535 repositories in the index in total.