OpenNMT/CTranslate2
quality grade B, 71 out of 100Fast inference engine for Transformer models
- stars
- 4.6k
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Making models smaller and faster: quantization, distillation, pruning, sparsity and kernel-level work.
Signals: quantization, model-compression, pruning, knowledge-distillation, gptq, awq, bitsandbytes, sparsity
184 results
Fast inference engine for Transformer models
A pytorch quantization backend for optimum
Fast and memory-efficient classical machine learning operators
PyTorch native quantization and sparsity for training and inference
Neural Networks with low bit weights on low end 32 bit microcontrollers such as the CH32V003 RISC-V Microcontroller and others
PyTorch implementation of 'Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding' by Song Han, Huizi Mao, William J. Dally
(ECCV'2020 Oral)EagleEye: Fast Sub-net Evaluation for Efficient Neural Network Pruning
PyTorch library to facilitate development and standardized evaluation of neural network pruning methods.
Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
MetaPruning: Meta Learning for Automatic Neural Network Channel Pruning. In ICCV 2019.
Binarized Convolutional Neural Networks on Software-Programmable FPGAs (FPGA'17)
Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks
Pruning Neural Networks with Taylor criterion in Pytorch
Papers for deep neural network compression and acceleration
Open Neural Network Exchange to C compiler.
Feature selection in neural networks
Quantization of Convolutional Neural networks.
0️⃣1️⃣🤗 BitNet-Transformers: Huggingface Transformers Implementation of "BitNet: Scaling 1-bit Transformers for Large Language Models" in pytorch with Llama(2) Architecture
Ollama model direct link generator and installer.
Distill videos, PDFs, transcripts, and notes into source-backed teacher Agent Skills.
AI Assistant that reduces the size of your application's Docker Image
Context cleaning for Claude Code — prune bloated sessions, protect Agent Teams from context loss, auto-guard with tiered pruning
Drop-in prompt compression for production LLM apps. Cut your token bill 40-60% without changing your code. Python SDK, LLMLingua-2, MIT.
FrugalGPT: better quality and lower cost for LLM applications
24,523 repositories in the index in total.