Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.
| Date | Stars |
|---|---|
| 2026-07-31 | 2098 |
| 2026-08-06 | 2097 |
Today
-1 stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<p align="left">
<a href="README_CN.md">中文</a>  |   English  
</p>
<br><br>
<p align="center">
<img src="https://qianwen-res.oss-cn-beijing.aliyuncs.com/assets/blog/qwenaudio/qwen2audio_logo.png" width="400"/>
<p>
<p align="center">
Qwen2-Audio-7B <a href="https://modelscope.cn/models/qwen/Qwen2-Audio-7B">🤖 </a> | <a href="https://huggingface.co/Qwen/Qwen2-Audio-7B">🤗</a>  | Qwen-Audio-7B-Instruct <a href="https://modelscope.cn/models/qwen/Qwen2-Audio-7B-Instruct">🤖 </a>| <a href="https://huggingface.co/Qwen/Qwen2-Audio-7B-Instruct">🤗</a>  | Demo<a href="https://modelscope.cn/studios/qwen/Qwen2-Audio-Instruct-Demo"> 🤖</a> | <a href="https://huggingface.co/spaces/Qwen/Qwen2-Audio-Instruct-Demo">🤗</a> 
<br>
📑 <a href="https://arxiv.org/abs/2407.10759">Paper</a>    |    📑 <a href="https://qwenlm.github.io/blog/qwen2-audio">Blog</a>    |    💬 <a href="https://github.com/QwenLM/Qwen/blob/main/assets/wechat.png">WeChat (微信)</a>   |    <a href="https://discord.gg/CV4E9rpNSD">Discord</a>  
</p>
We introduce the latest progress of Qwen-Audio, a large-scale audio-language model called Qwen2-Audio, which is capable of accepting various audio signal inputs and performing audio analysis or direct textual responses with regard to speech instructions. We introduce two distinct audio interaction modes:
* voice chat: users can freely engage in voice interactions with Qwen2-Audio without text input;
* audio analysis: users could provide audio and text instructions for analysis during the interaction;
**We've released two models of the Qwen2-Audio series: Qwen2-Audio-7B and Qwen2-Audio-7B-Instruct.**
## Architecture
The overview of three-stage training process of Qwen2-Audio.
<p align="center">
<img src="assets/framework.png" width="80%"/>
<p>
## News and Updates
* 2024.8.9 🎉 We released the checkpoints of both `Qwen2-Audio-7B` and `Qwen2-Audio-7B-Instruct` on ModelScope and Hugging Face.
* 2024.7.15 🎉 We released the paper of **Qwen2-Audio**, introducing the relevant model structure, training methods, and model performance. Check our [report](https://arxiv.org/abs/2407.10759) for details!
* 2023.11.30 🔥 We released the **Qwen-Audio** series.
<br>
## Evaluation
We evaluated the Qwen2-Audio's abilities on 13 standard benchmarks as follows:
<table><thead><tr><th>Task</th><th>Description</th><th>Dataset</th><th>Split</th><th>Metric</th></tr></thead><tbody><tr><td rowspan="4">ASR</td><td rowspan="4">Automatic Speech Recognition</td><td>Fleurs</td><td>dev | test</td><td rowspan="4">WER</td></tr><tr><td>Aishell2</td><td>test</td></tr><tr><td>Librispeech</td><td>dev | test</td></tr><tr><td>Common Voice</td><td>dev | test</td></tr><tr><td>S2TT</td><td>Speech-to-Text Translation</td><td>CoVoST2</td><td>test</td><td>BLEU </td></tr><tr><td>SER</td><td>Speech Emotion Recognition</td><td>Meld</td><td>test</td><td>ACC</td></tr><tr><td>VSC</td><td>Vocal Sound Classification</td><td>VocalSound</td><td>test</td><td>ACC</td></tr><tr><td rowspan="4"><a href="https://github.com/OFA-Sys/AIR-Bench">AIR-Bench</a><br></td><td>Chat-Benchmark-Speech</td><td>Fisher<br>SpokenWOZ<br>IEMOCAP<br>Common voice</td><td>dev | test</td><td>GPT-4 Eval</td></tr><tr><td>Chat-Benchmark-Sound</td><td>Clotho</td><td>dev | test</td><td>GPT-4 Eval</td></tr>
<tr><td>Chat-Benchmark-Music</td><td>MusicCaps</td><td>dev | test</td><td>GPT-4 Eval</td></tr><tr><td>Chat-Benchmark-Mixed-Audio</td><td>Common voice<br>AudioCaps<br>MusicCaps</td><td>dev | test</td><td>GPT-4 Eval</td></tr></tbody></table>
The below is the overal performance:
<p align="center">
<img src="assets/radar_compare_qwen_audio.png" width="70%"/>
<p>
The details of evaluation are as follows:
<br>
<b>(Note: The evaluation results we present are based on the initial model of the original training framework. However, the scores showed some fluctuations after converting the framework to Huggingface. HeExcerpt of 20,674 characters
Read on GitHubYunfei Chu · Qwen Team, Alibaba Group · China
11
Binyuan Hui · Formerly @ Qwen · Singapore
2
Alibaba Group · China
2
Long Chen · @livekit
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:cb1932d7fac87793, llm:Repository description: 'The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.' Language: Python. Indicates pretrained large audio language model and chat capability.
matched fp:cb1932d7fac87793, llm:Repository description: 'The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.' Language: Python. Indicates pretrained large audio language model and chat capability.
matched fp:cb1932d7fac87793, llm:Repository description: 'The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.' Language: Python. Indicates pretrained large audio language model and chat capability.