Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI
| Date | Stars |
|---|---|
| 2026-07-31 | 487 |
| 2026-08-02 | 489 |
| 2026-08-06 | 489 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
35.0
growth rate 0.00%/day
<div align="center"> # TensorRT Edge-LLM **High-Performance Large Language Model Inference Framework for NVIDIA Edge Platforms** [](https://nvidia.github.io/TensorRT-Edge-LLM/) [](https://github.com/NVIDIA/TensorRT-Edge-LLM/blob/main/tensorrt_edgellm/_version.py) [](https://github.com/NVIDIA/TensorRT-Edge-LLM/blob/main/LICENSE) [Overview](https://nvidia.github.io/TensorRT-Edge-LLM/latest/overview.html) | [Quick Start](https://nvidia.github.io/TensorRT-Edge-LLM/latest/user_guide/getting_started/quick-start-guide.html) | [Performance](https://nvidia.github.io/TensorRT-Edge-LLM/latest/user_guide/performance/performance-benchmarks.html) | [Documentation](https://nvidia.github.io/TensorRT-Edge-LLM/) | [Roadmap](https://github.com/NVIDIA/TensorRT-Edge-LLM/issues?q=is%3Aissue%20state%3Aopen%20label%3ARoadmap) --- <div align="left"> ## Latest News - **[2026/07]** Support for the full **Gemma 4** family (E2B / E4B / 12B / 26B-A4B / 31B — multimodal text + image + audio, with MTP), **Qwen3-Omni** and **Nemotron-3** NVFP4, and **DFlash** speculative decoding (with DDTree for Qwen3 / Qwen3.5) landed across releases 0.9.0 and 0.9.1. --- ## Overview TensorRT Edge-LLM is NVIDIA's high-performance C++ inference runtime for Large Language Models (LLMs) and Vision-Language Models (VLMs) on embedded platforms. It enables efficient deployment of state-of-the-art language models on resource-constrained devices such as NVIDIA Jetson, NVIDIA DRIVE, and NVIDIA DGX Spark platforms. TensorRT Edge-LLM provides convenient Python scripts to convert HuggingFace checkpoints to [ONNX](https://onnx.ai). Engine build and end-to-end inference runs entirely on Edge platforms. --- ## Getting Started For the supported platforms, models and precisions, see the [**Overview**](https://nvidia.github.io/TensorRT-Edge-LLM/latest/overview.html). Get started with TensorRT Edge-LLM in <15 minutes. For complete installation and usage instructions, see the [**Quick Start Guide**](https://nvidia.github.io/TensorRT-Edge-LLM/latest/user_guide/getting_started/quick-start-guide.html). --- ## Documentation ### Introduction - **[Overview](https://nvidia.github.io/TensorRT-Edge-LLM/latest/overview.html)** - What is TensorRT Edge-LLM and key features - **[Supported Models](https://nvidia.github.io/TensorRT-Edge-LLM/latest/user_guide/getting_started/supported-models.html)** - Complete model compatibility matrix - **[Checkpoint Exporter](https://nvidia.github.io/TensorRT-Edge-LLM/latest/developer_guide/software-design/checkpoint-export.html)** - Recommended ONNX export pipeline ### User Guide - **[Installation](https://nvidia.github.io/TensorRT-Edge-LLM/latest/user_guide/getting_started/installation.html)** - Set up quantization, `tensorrt_edgellm`, and the C++ runtime - **[Quick Start Guide](https://nvidia.github.io/TensorRT-Edge-LLM/latest/user_guide/getting_started/quick-start-guide.html)** - Run your first inference in ~15 minutes - **[Examples](https://nvidia.github.io/TensorRT-Edge-LLM/latest/user_guide/examples/index.html)** - End-to-end workflows - **[Quantization](https://nvidia.github.io/TensorRT-Edge-LLM/latest/user_guide/features/quantization.html)** - Create quantized checkpoints for `tensorrt_edgellm` - **[Experimental High-Level Python API and Server](https://nvidia.github.io/TensorRT-Edge-LLM/latest/user_guide/examples/experimental-server.html)** - vLLM-style API and OpenAI-compatible server - **[Input Format Guide](https://nvidia.github.io/TensorRT-Edge-LLM/latest/user_guide/format/input-format.html)** - Request format and specifications - **[Chat Template Format](https://nvidia.github.io/TensorRT-Edge-LLM/lates
Excerpt of 7,942 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:6a9d920476d5d15e, llm:Repository description: 'High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI' (NVIDIA TensorRT-Edge-LLM).
matched fp:6a9d920476d5d15e, llm:Repository description: 'High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI' (NVIDIA TensorRT-Edge-LLM).
matched fp:6a9d920476d5d15e, llm:Repository description: 'High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI' (NVIDIA TensorRT-Edge-LLM).