kangarooking/cangjie-skill
quality grade B, 71 out of 100把书、长视频、播客等高价值内容蒸馏成可执行的 Agent Skills
- stars
- 5.5k
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Making models smaller and faster: quantization, distillation, pruning, sparsity and kernel-level work.
Signals: quantization, model-compression, pruning, knowledge-distillation, gptq, awq, bitsandbytes, sparsity
181 results
把书、长视频、播客等高价值内容蒸馏成可执行的 Agent Skills
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with 100+ tok/s single-request decode with support of FP8 weight
Accessible large language models via k-bit quantization for PyTorch.
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs
TRACER: replace 90%+ of your LLM classification calls with a traditional ML model. Formal parity guarantees. Self-improving.
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
AIMET is a library that provides advanced quantization and compression techniques for trained neural network models.
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
vLLM fork for Tesla V100 (SM70) with AWQ 4-bit support, CUDA 12.8 build flow, and validated Qwen3.5 27B/35B deployment on multi-GPU V100.
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
Python package for LLM compression
Neural Network Compression Framework for enhanced OpenVINO™ inference
SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
No description
A Lightweight LLM Post-Training Library
Surrogate modeling and optimization for scientific machine learning (SciML)
Build compute kernels and load them from the Hub.
🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
A toolkit to optimize ML models for deployment for Keras and TensorFlow, including quantization and pruning.
A collection of memory efficient attention operators implemented in the Triton language.
Fast inference engine for Transformer models
24,520 repositories in the index in total.