Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
SALMONN family: A suite of advanced multi-modal LLMs
| Date | Stars |
|---|---|
| 2026-07-24 | 1479 |
| 2026-07-25 | 1479 |
| 2026-07-28 | 1479 |
| 2026-07-30 | 1479 |
| 2026-07-31 | 1482 |
| 2026-08-06 | 1482 |
Today
— stars today
This week
+3 stars this week
This month
— stars this month
Momentum
3.0
growth rate 0.20%/day
# SALMONN family: A suite of advanced multi-modal LLMs
<div align=center><img src="resource/salmon.png" height="256px" width="256px"/></div>
<h1 align="center">
<a href="https://git.io/typing-svg">
<img src="https://readme-typing-svg.herokuapp.com/?lines=Hello,+There!+👋;Welcome+to+SALMONN+family!;¢er=true&size=20">
</a>
</h1>
🚀🚀 Welcome to the repo of **SALMONN**!!
The SALMONN model family consists of a series of advanced multi-modal large language models. For more details, please refer to the corresponding branches.
- [SALMONN 2](https://github.com/bytedance/SALMONN/tree/salmonn2)
- [[ICLR 2026] ELLSA](https://github.com/bytedance/SALMONN/tree/ELLSA)
- [video-SALMONN 2](https://github.com/bytedance/video-SALMONN-2)
- [[ICML 2025] F-16](https://github.com/bytedance/F-16)
- [[ICML 2025] video-SALMONN-o1](https://github.com/bytedance/SALMONN/tree/video-salmonn-o1)
- [[ICASSP 2025 & ACL 2025] SALMONN for speech quality assessment](https://github.com/bytedance/SALMONN/tree/speech_quality_assessment)
- [[ICML 2024] video-SALMONN](https://github.com/bytedance/SALMONN/tree/videosalmonn)
- [[ICLR 2024] SALMONN](https://github.com/bytedance/SALMONN/tree/salmonn)
## 🔥 News
- [2026-04-20] We have released the model and inference code for **ELLSA**! See [here](https://github.com/bytedance/SALMONN/tree/ELLSA)! ELLSA is the first end-to-end model that unifies vision, speech, text and action in a streaming full-duplex framework, enabling joint multimodal perception and concurrent generation.
- [2025-07-08] We have opensourced **video-SALMONN 2**! video-SALMONN 2 is a powerful audio-visual LLM that generates high-quality audio-visual video captions and achieves competitive performance on general video QA benchmarks.
- [2025-06-01] We have opensourced **QualiSpeech** dataset - A speech quality assessment dataset with natural language reasoning. You can use QualiSpeech to develop your own audio LLM for speech quality assessment or to evaluate the low-level speech perception capabilities of existing audio LLMs. Feel free to download it [here](https://huggingface.co/datasets/tsinghua-ee/QualiSpeech)!
- [2025-03-03] We have released the data processing scripts and finetuned model checkpoints for **SALMONN** for speech quality assessment! See [here](https://github.com/bytedance/SALMONN/tree/speech_quality_assessment)!
- [2024-09-04] We have released the model and inference code for **video-SALMONN**! See [here](https://github.com/bytedance/SALMONN/tree/videosalmonn)!
- [2024-05-28] 🧳 We have released all the annotations (including 600k SQA/AQA data and 50k audio-based storytelling data) for the 3-stage training of SALMONN! Feel free to download them [here](https://drive.google.com/file/d/15cQO--rtMM9JD22y-A5oXXvT3DujgE2e/view?usp=sharing)!
- [2024-04-07] 🤖 We have released all the codes you need to train your own SALMONN! Try some cool things!
- [2024-01-16] 💖 Our paper was accepted by ICLR 2024!
- [2023-11-13] 🎁 We have released a **7B version of SALMONN** at [tsinghua-ee/SALMONN-7B](https://huggingface.co/tsinghua-ee/SALMONN-7B) and built the 7B demo [here](https://huggingface.co/spaces/tsinghua-ee/SALMONN-7B-gradio)!
- [2023-10-08] ✨ We have released [**the model checkpoint**](https://huggingface.co/tsinghua-ee/SALMONN) and **the inference code** for SALMONN-13B!
## 📖 Paper List
```
@inproceedings{wang2026end,
title={End-to-end Listen, Look, Speak and Act},
author={Wang, Siyin and Yu, Wenyi and Chen, Xianzhao and Tian, Xiaohai and Zhang, Jun and Lu, Lu and Zhang, Chao},
booktitle={Proc. ICLR},
year={2026},
address={Rio de Janeiro}
}
@inproceedings{
sun2025videosalmonno1,
title={{video-SALMONN-o1}: Reasoning-enhanced Audio-visual Large Language Model},
author={Guangzhi Sun, Yudong Yang, Jimin Zhuang, Changli Tang, Yixuan Li, Wei Li, Zejun MA, Chao Zhang},
booktitle={ICML},
year={2025}
}
@article{tang2025video,
title={{video-SALMONN 2: Captioning-Enhanced Audio-Visual Large Language Models}},
authoExcerpt of 5,711 characters
Read on GitHub35
24
9
6
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:df9afd7edac64c1a, topic:speech-recognition, topic:audio-processing
matched fp:df9afd7edac64c1a, topic:large-language-models