fudan-zvg/SOFT
quality grade D, 46 out of 100[NeurIPS 2021 Spotlight] & [IJCV 2024] SOFT: Softmax-free Transformer with Linear Complexity
- stars
- 310
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Released model weights, reference implementations and architecture research.
Signals: large-language-models, llm, foundation-models, transformer, gpt, llama, mistral, qwen
1,475 results
[NeurIPS 2021 Spotlight] & [IJCV 2024] SOFT: Softmax-free Transformer with Linear Complexity
Laravel 5 JSON API Transformer Package
MoH: Multi-Head Attention as Mixture-of-Head Attention
Official implementation of All Atom Diffusion Transformers (ICML 2025)
💬 Chatbot web app + HTTP and Websocket endpoints for LLM inference with the Petals client
This is a repository with the code for the ACL 2019 paper "Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned" and the ACL 2021 paper "Analyzing Source and Target Contributions to NMT Predictions".
Understanding the Difficulty of Training Transformers
[EMNLP'2024] "OpenGraph: Towards Open Graph Foundation Models"
(CVPR 2022) Pytorch implementation of "Self-supervised transformers for unsupervised object discovery using normalized cut"
[ICLR 2023] "Is Attention All NeRF Needs?" by Mukund Varma T*, Peihao Wang* , Xuxi Chen, Tianlong Chen, Subhashini Venugopalan, Zhangyang Wang
Lumen 10 基础上扩展出的API 启动项目,精心设计的目录结构,规范统一的响应数据格式,Repository 模式架构的最佳实践。
[NeurIPS'22] Tokenized Graph Transformer (TokenGT), in PyTorch
The fastest way to build and start training your own LLM. CLI tool that scaffolds production-ready PyTorch training projects in seconds. Like create-next-app but for language models.
Advanced TypeScript runtime reflection system
PyTorch Implementation of OpenAI GPT-2
[ECCV2024 Oral🔥] Official Implementation of "GiT: Towards Generalist Vision Transformer through Universal Language Interface"
The official implementation for "Spherical Transformer for LiDAR-based 3D Recognition" (CVPR 2023).
Code & Data for "Tabular Transformers for Modeling Multivariate Time Series" (ICASSP, 2021)
LLaMa/RWKV onnx models, quantization and testcase
Notes about "Attention is all you need" video (https://www.youtube.com/watch?v=bCz4OMemCcA)
Official repo for Medical Image Segmentation Review: The Success of U-Net
RXNMapper: Unsupervised attention-guided atom-mapping. Code complementing our Science Advances publication on "Extraction of organic chemistry grammar from unsupervised learning of chemical reactions" (https://advances.sciencemag.org/content/7/15/eabe4166).
LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.
A PyTorch implementation of "CoAtNet: Marrying Convolution and Attention for All Data Sizes"
24,523 repositories in the index in total.