dmlc/mxnet-memonger
quality grade D, 39 out of 100Sublinear memory optimization for deep learning, reduce GPU memory cost to train deeper nets
- stars
- 306
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Making models smaller and faster: quantization, distillation, pruning, sparsity and kernel-level work.
Signals: quantization, model-compression, pruning, knowledge-distillation, gptq, awq, bitsandbytes, sparsity
184 results
Sublinear memory optimization for deep learning, reduce GPU memory cost to train deeper nets
Ranger deep learning optimizer rewrite to use newest components
Generate 8-bit chiptunes with deep learning
Sublinear memory optimization for deep learning. https://arxiv.org/abs/1604.06174
Embedded and mobile deep learning research resources
Ranger - a synergistic optimizer using RAdam (Rectified Adam), Gradient Centralization and LookAhead in one codebase
Codebase for "SLIDE : In Defense of Smart Algorithms over Hardware Acceleration for Large-Scale Deep Learning Systems"
A New Optimization Technique for Deep Neural Networks
Binarized Neural Network (BNN) for pytorch
Open Neural Network Compiler
Summary, Code for Deep Neural Network Quantization
AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (NVIDIA GPU) and MatrixCore (AMD GPU) inference.
A flexible and efficient deep neural network (DNN) compiler that generates high-performance executable from a DNN model description.
List of papers related to neural network quantization in recent AI conferences and journals.
PyHessian is a Pytorch library for second-order based analysis and training of Neural Networks
Filter Pruning via Geometric Median for Deep Convolutional Neural Networks Acceleration (CVPR 2019 Oral)
Collection of recent methods on (deep) neural network compression and acceleration.
Winograd minimal convolution algorithm generator for convolutional neural networks.
Muon is an optimizer for hidden layers in neural networks
TinyNeuralNetwork is an efficient and easy-to-use deep learning model compression framework.
Fast Block Sparse Matrices for Pytorch
Config driven, easy backup cli for restic.
A curated list of neural network pruning resources.
[CVPR 2023] DepGraph: Towards Any Structural Pruning; LLMs, Vision Foundation Models, etc.
24,523 repositories in the index in total.