yassa9/qwen600
quality grade D, 46 out of 100Static suckless single batch CUDA-only qwen3-0.6B mini inference engine
- stars
- 560
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Released model weights, reference implementations and architecture research.
Signals: large-language-models, llm, foundation-models, transformer, gpt, llama, mistral, qwen
1,456 results
Static suckless single batch CUDA-only qwen3-0.6B mini inference engine
A collection of view pager transformers
[CVPR22] Official Implementation of DAFormer: Improving Network Architectures and Training Strategies for Domain-Adaptive Semantic Segmentation
Implementation of Enformer, Deepmind's attention network for predicting gene expression, in Pytorch
[ICLR 2022] Official implementation of the paper "DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR"
历年ICLR论文和开源项目合集,包含ICLR2021、ICLR2022、ICLR2023、ICLR2024、ICLR2025.
[ICLR2025 Spotlight🔥] Official Implementation of TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
NLP Paper
End-to-end ASR/LM implementation with PyTorch
Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting (NeurIPS 2019)
[CVPR 2021, Oral] PREDATOR: Registration of 3D Point Clouds with Low Overlap.
Multimodal model for text and tabular data with HuggingFace transformers as building block for text data
Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement the code.
[ECCV 2022] Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework
This is an official implementation for "Self-Supervised Learning with Swin Transformers".
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 论文的中文翻译 Chinese Translation!
[ICCV'21] Learning Spatio-Temporal Transformer for Visual Tracking
[ICML 2024] SqueezeLLM: Dense-and-Sparse Quantization
Implementation of plug in and play Attention from "LongNet: Scaling Transformers to 1,000,000,000 Tokens"
Cameras as Relative Positional Encoding
[CVPR 2022] HiVT: Hierarchical Vector Transformer for Multi-Agent Motion Prediction
[Mamba-Survey-2024] Paper list for State-Space-Model/Mamba and it's Applications
[TNSRE 23] EEG Transformer 2.0. i. Convolutional Transformer for EEG Decoding. ii. Novel visualization - Class Activation Topography.
[CVPR'22 Oral] GMFlow: Learning Optical Flow via Global Matching
24,540 repositories in the index in total.