kangarooking/cangjie-skill
quality grade A, 81 out of 100把书、长视频、播客等高价值内容蒸馏成可执行的 Agent Skills(Distill high-value content from books, long-form videos, podcasts, and more into executable Agent Skills)
- stars
- 10k
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Making models smaller and faster: quantization, distillation, pruning, sparsity and kernel-level work.
Signals: quantization, model-compression, pruning, knowledge-distillation, gptq, awq, bitsandbytes, sparsity
187 results
把书、长视频、播客等高价值内容蒸馏成可执行的 Agent Skills(Distill high-value content from books, long-form videos, podcasts, and more into executable Agent Skills)
V100 / SM70-focused vLLM engineering fork for modern LLM inference.
Dataflow compiler for QNN inference on FPGAs
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs
The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 100+ tok/s single-request decode with support of FP8 weight ( Join Discord :https://discord.gg/VFqVVySdMS )
A Lightweight LLM Post-Training Library
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B
Open Machine Learning Compiler Framework
🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
FasterAI: Prune and Distill your models with FastAI and PyTorch
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
A SOTA quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的量化工具包
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
Accessible large language models via k-bit quantization for PyTorch.
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
Fast inference engine for Transformer models
Build compute kernels and load them from the Hub.
Optimised Neural Network functions for Espressif chipsets
AIMET is a library that provides advanced quantization and compression techniques for trained neural network models.
Ollama model direct link generator and installer.
24,535 repositories in the index in total.