Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Codebase for Merging Language Models (ICML 2024)
| Date | Stars |
|---|---|
| 2026-07-31 | 871 |
| 2026-08-04 | 872 |
| 2026-08-06 | 872 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch <div align="center"> <img src="figures/icon.jpeg" width="25%"> </div> This repository is built for the paper [Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch](https://arxiv.org/abs/2311.03099). 🔔 If you have any questions or suggestions, please feel free to let us know. You can directly email [Le Yu](https://yule-buaa.github.io/) using the email address [email protected] or post an issue on this repository. ## 💥 News 💥 - 🔥🔥🔥[**May 2, 2024**] Our paper is accepted at ICML 2024! The camera ready version is coming soon. - 🔥🔥🔥[**February 9, 2024**] Special thanks to [Sourab Mangrulkar](https://github.com/pacman100) for integrating our work into the [huggingface/peft Project](https://github.com/huggingface/peft)! - 🔥🔥🔥[**January 28, 2024**] Our merged model [supermario_v2](https://huggingface.co/vanillaOVO/supermario_v2) ranks first among 7B models on the Open LLM Leaderboard! - 🔥🔥🔥[**December 4, 2023**] We appreciate [Minhajul Hoque](https://medium.com/@minh.hoque) for sharing our work on [Medium](https://medium.com/@minh.hoque/paper-explained-language-models-are-super-mario-2ebce6c2cf35)! - 🔥🔥🔥[**November 29, 2023**] Special thanks to [papersread.ai](https://papersread.ai/) for sharing [our work](https://papersread.ai/e/language-models-are-super-mario-absorbing-abilities-from-homologous-models-as-a-free-lunch/)! - 🔥🔥🔥[**November 29, 2023**] We appreciate [martyn](https://github.com/martyn) for extending our work to [Stable Diffusion models](https://github.com/martyn/safetensors-merge-supermario)! - 🔥🔥🔥[**November 27, 2023**] Special thanks to [brucethemoose](https://huggingface.co/brucethemoose) for applying our work on the [model](https://huggingface.co/brucethemoose/CapyTessBorosYi-34B-200K-DARE-Ties) on Hugging Face! - 🔥🔥🔥[**November 26, 2023**] We appreciate [cg123](https://github.com/cg123) for integrating our work into the [mergekit Project](https://github.com/arcee-ai/mergekit)! - 🔥🔥🔥[**November 25, 2023**] Special thanks to [fly51fly](https://twitter.com/fly51fly) for sharing our work on [Twitter](https://twitter.com/fly51fly/status/1728159826742755588)! - 🔥🔥🔥[**November 24, 2023**] We appreciate [uukuguy](https://github.com/uukuguy) for integrating our work into the [Multi-LoRAs Project](https://pypi.org/project/multi-loras/0.2.0)! - 🔥🔥🔥[**November 23, 2023**] Special thanks to [WizardLM](https://twitter.com/WizardLM_AI) for sharing our work on [Twitter](https://twitter.com/WizardLM_AI/status/1727672799391842468)! - 🔥🔥🔥[**November 21, 2023**] We appreciate [PaperWeekly](http://www.paperweekly.info) for sharing our work on [WeChat](https://mp.weixin.qq.com/s/YiqWovBUXIbzmUbL6uT-8g) and [Zhihu](https://zhuanlan.zhihu.com/p/668152236)! - 🔥🔥🔥[**November 11, 2023**] Special thanks to [夕小瑶](https://xixiaoyao.github.io/about/) for sharing our work on [WeChat](https://mp.weixin.qq.com/s?__biz=MzIwNzc2NTk0NQ%3D%3D&mid=2247565881&idx=2&sn=57985427fdb6751d617df801ca7fd810) and [Zhihu](https://zhuanlan.zhihu.com/p/666363702)! - 🔥🔥🔥[**November 6, 2023**] Our paper is available on [arXiv](https://arxiv.org/abs/2311.03099), [Papers With Code](https://paperswithcode.com/paper/language-models-are-super-mario-absorbing), and [Hugging Face](https://huggingface.co/papers/2311.03099). ## Overview In this work, we uncover that Language Models (LMs), either encoder- or decoder-based, can **obtain new capabilities by assimilating the parameters of homologous models without the need for retraining or GPUs**. 1. We introduce a novel operation called **DARE** to directly set most of (90% or even 99%) the delta parameters to zeros without affecting the capabilities of SFT LMs. 2. We sparsify delta parameters of multiple SFT homologous models with DARE as a **general preprocessing technique** and subsequently merge them into a single model by parameter
Excerpt of 15,969 characters
Read on GitHubIkko Eltociear Ashimine · Japan
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:dc35e84e5b04ec0c, llm:Repository description: 'Codebase for Merging Language Models (ICML 2024)'. Language: Python. Purpose: merging language models (research code).
matched fp:dc35e84e5b04ec0c, llm:Repository description: 'Codebase for Merging Language Models (ICML 2024)'. Language: Python. Purpose: merging language models (research code).
matched fp:dc35e84e5b04ec0c, llm:Repository description: 'Codebase for Merging Language Models (ICML 2024)'. Language: Python. Purpose: merging language models (research code).