SciML/Surrogates.jl
quality grade C, 62 out of 100Surrogate modeling and optimization for scientific machine learning (SciML)
- stars
- 382
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Making models smaller and faster: quantization, distillation, pruning, sparsity and kernel-level work.
Signals: quantization, model-compression, pruning, knowledge-distillation, gptq, awq, bitsandbytes, sparsity
187 results
Surrogate modeling and optimization for scientific machine learning (SciML)
Python package for LLM compression
PyTorch native quantization for training and inference
A collection of memory efficient attention operators implemented in the Triton language.
No description
SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
Neural Network Compression Framework for enhanced OpenVINO™ inference
A curated collection of papers, benchmarks, surveys, and tools for model quantization, covering low-bit networks, LLMs, multimodal and generative models, vector and lattice quantization, and efficient deployment.
TRACER: replace 90%+ of your LLM classification calls with a traditional ML model. Formal parity guarantees. Self-improving.
Fast and memory-efficient classical machine learning operators
Context cleaning for Claude Code — prune bloated sessions, protect Agent Teams from context loss, auto-guard with tiered pruning
Drop-in prompt compression for production LLM apps. Cut your token bill 40-60% without changing your code. Python SDK, LLMLingua-2, MIT.
A toolkit to optimize ML models for deployment for Keras and TensorFlow, including quantization and pruning.
Awesome LLM compression research papers and tools.
Infrastructure for Machine Learning Guided Optimization (MLGO) in LLVM.
Neural Networks with low bit weights on low end 32 bit microcontrollers such as the CH32V003 RISC-V Microcontroller and others
A curated list of awesome edge machine learning resources, including research papers, inference engines, challenges, books, meetups and others.
Summary, Code for Deep Neural Network Quantization
A powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.
PyTorch implementation of 'Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding' by Song Han, Huizi Mao, William J. Dally
(ECCV'2020 Oral)EagleEye: Fast Sub-net Evaluation for Efficient Neural Network Pruning
PyTorch library to facilitate development and standardized evaluation of neural network pruning methods.
Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
MetaPruning: Meta Learning for Automatic Neural Network Channel Pruning. In ICCV 2019.
24,535 repositories in the index in total.