cornell-zhang/bnn-fpga
quality grade D, 48 out of 100Binarized Convolutional Neural Networks on Software-Programmable FPGAs (FPGA'17)
- stars
- 316
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Making models smaller and faster: quantization, distillation, pruning, sparsity and kernel-level work.
Signals: quantization, model-compression, pruning, knowledge-distillation, gptq, awq, bitsandbytes, sparsity
187 results
Binarized Convolutional Neural Networks on Software-Programmable FPGAs (FPGA'17)
Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks
Pruning Neural Networks with Taylor criterion in Pytorch
Papers for deep neural network compression and acceleration
Open Neural Network Exchange to C compiler.
Feature selection in neural networks
Quantization of Convolutional Neural networks.
0️⃣1️⃣🤗 BitNet-Transformers: Huggingface Transformers Implementation of "BitNet: Scaling 1-bit Transformers for Large Language Models" in pytorch with Llama(2) Architecture
Distill videos, PDFs, transcripts, and notes into source-backed teacher Agent Skills.
AI Assistant that reduces the size of your application's Docker Image
FrugalGPT: better quality and lower cost for LLM applications
1.58 Bit LLM on Apple Silicon using MLX
Official codebase for "Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling".
Awesome list for LLM pruning.
Cuda kernels for leveraging LLM sparsity to improve throughput and decrease the memory requirements during inference and training.
No description
[ICLR 2025🔥] SVD-LLM & [NAACL 2025🔥] SVD-LLM V2
Run larger LLMs with longer contexts on Apple Silicon by using differentiated precision for KV cache quantization. KVSplit enables 8-bit keys & 4-bit values, reducing memory by 59% with <1% quality loss. Includes benchmarking, visualization, and one-command setup. Optimized for M1/M2/M3 Macs with Metal support.
W8A8/W4A8 inference + optimized SDPA on Apple Silicon — unlocking unused INT8 TensorOps in M5 for 1.2–1.9× faster LLM prefill, plus FlashInfer-inspired GQA decode attention for up to 1.6× SDPA speedup, built as MLX custom primitives.
Design hardware-friendly model architectures and migrate existing LLMs with minimal performance loss
[MLSys'24] Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
Fast Matrix Multiplications for Lookup Table-Quantized LLMs
Code repo for the paper "SpinQuant LLM quantization with learned rotations"
Awesome list for LLM quantization
24,535 repositories in the index in total.