andravin/wincnn
quality grade D, 47 out of 100Winograd minimal convolution algorithm generator for convolutional neural networks.
- stars
- 628
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Making models smaller and faster: quantization, distillation, pruning, sparsity and kernel-level work.
Signals: quantization, model-compression, pruning, knowledge-distillation, gptq, awq, bitsandbytes, sparsity
187 results
Winograd minimal convolution algorithm generator for convolutional neural networks.
Muon is an optimizer for hidden layers in neural networks
TinyNeuralNetwork is an efficient and easy-to-use deep learning model compression framework.
Fast Block Sparse Matrices for Pytorch
Config driven, easy backup cli for restic.
A curated list of neural network pruning resources.
[CVPR 2023] DepGraph: Towards Any Structural Pruning; LLMs, Vision Foundation Models, etc.
利用pytorch实现图像分类的一个完整的代码,训练,预测,TTA,模型融合,模型部署,cnn提取特征,svm或者随机森林等进行分类,模型蒸馏,一个完整的代码
A pytorch quantization backend for optimum
FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.
Z80-μLM is a 2-bit quantized language model small enough to run on an 8-bit Z80 processor. Train conversational models in Python, export them as CP/M .COM binaries, and chat with your vintage computer.
Efficient computing methods developed by Huawei Noah's Ark Lab
OpenMMLab Model Compression Toolbox and Benchmark.
Base pretrained models and datasets in pytorch (MNIST, SVHN, CIFAR10, CIFAR100, STL10, AlexNet, VGG16, VGG19, ResNet, Inception, SqueezeNet)
Pretrained language model and its related optimization techniques developed by Huawei Noah's Ark Lab.
Lossy PNG compressor — pngquant command based on libimagequant library
Must-read papers on deep learning to hash (DeepHash)
An Open-Source Package for Deep Learning to Hash (DeepHash)
Always sparse. Never dense. But never say never. A Sparse Training repository for the Adaptive Sparse Connectivity concept and its algorithmic instantiation, i.e. Sparse Evolutionary Training, to boost Deep Learning scalability on various aspects (e.g. memory and computational time efficiency, representation and generalization power).
Sparse Optimisation Research Code
Reference ImageNet implementation of SelecSLS CNN architecture proposed in the SIGGRAPH 2020 paper "XNect: Real-time Multi-Person 3D Motion Capture with a Single RGB Camera". The repository also includes code for pruning the model based on implicit sparsity emerging from adaptive gradient descent methods, as detailed in the CVPR 2019 paper "On implicit filter level sparsity in Convolutional Neural Networks".
Caffe for Sparse and Low-rank Deep Neural Networks
[NeurIPS'23] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.
[ICCV 2023] Q-Diffusion: Quantizing Diffusion Models.
24,535 repositories in the index in total.