Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
| Date | Stars |
|---|---|
| 2026-07-24 | 948 |
| 2026-07-25 | 948 |
| 2026-07-28 | 948 |
| 2026-07-30 | 948 |
| 2026-07-31 | 949 |
| 2026-08-06 | 949 |
Today
— stars today
This week
+1 stars this week
This month
— stars this month
Momentum
1.0
growth rate 0.10%/day
# 🎧 audio-ai-hub **The hub for audio AI research.** Curated papers, open models, benchmarks and datasets across audio LLMs · speech recognition · speech synthesis · music & audio generation. `129 entries` · `11 categories` · `latest: 2026-06` 👉 **[Browse the interactive hub →](https://binwang28.github.io/audio-ai-hub/)** · [Contribute](CONTRIBUTING.md) · [Suggest a paper](https://github.com/BinWang28/audio-ai-hub/issues/new?template=add-paper.yml) > The page below is a quick snapshot. For search, filtering by category and sorting by stars or date, the **[live site](https://binwang28.github.io/audio-ai-hub/)** is much faster than scrolling this README. --- ## ⭐ Featured _Top 8 by GitHub stars — refreshed weekly by `.github/workflows/refresh-stars.yml`._ | # | Project | Stars | What it does | |---|---------|-------|--------------| | 1 | **[Whisper](https://arxiv.org/abs/2212.04356)** | ⭐ 105k+ | Whisper is OpenAI's open-source speech recognition model trained on 680K hours of multilingual and multitask supervised data from the web. | | 2 | **[VoxCPM2](https://arxiv.org/abs/2606.06928)** | ⭐ 33k+ | VoxCPM2 is a fully open-source 2B-parameter multilingual, controllable speech generation foundation model extending VoxCPM's hierarchical… | | 3 | **[MMS](https://arxiv.org/abs/2305.13516)** | ⭐ 32k+ | MMS (Massively Multilingual Speech) extends speech foundation models (wav2vec 2.0) to 1,107 languages for ASR and adds TTS and language i… | | 4 | **[MiniCPM-o](https://arxiv.org/abs/2604.27393)** | ⭐ 25k+ | MiniCPM-o 4.5 is OpenBMB's compact (8B-class) full-duplex omni-modal LLM supporting real-time vision, speech, and text interaction with l… | | 5 | **[MusicGen](https://arxiv.org/abs/2306.05284)** | ⭐ 23k+ | MusicGen is Meta's single-stage autoregressive transformer for controllable text-conditioned music generation, operating over discrete En… | | 6 | **[AudioGen](https://arxiv.org/abs/2209.15352)** | ⭐ 23k+ | AudioGen is a transformer-based autoregressive model for text-to-environmental-sound generation, trained on discrete audio tokens. | | 7 | **[CosyVoice 3](https://arxiv.org/abs/2505.17589)** | ⭐ 22k+ | CosyVoice 3 scales the CosyVoice TTS stack with significantly larger pre-training data and a dedicated post-training stage, targeting in-… | | 8 | **[CosyVoice 2](https://arxiv.org/abs/2412.10117)** | ⭐ 22k+ | CosyVoice 2 is Alibaba's streaming TTS LLM, combining a unified speech tokenizer with a streaming-friendly LLM backbone to enable bidirec… | ## 🆕 Recently added _The 10 most recent entries by date. See the [interactive site](https://binwang28.github.io/audio-ai-hub/) for everything else._ - `2026-06` · [ACA-SER](https://arxiv.org/abs/2606.07309) — A probing study testing whether instruction-following audio language models use explicit acoustic concept tokens (six interpretable cues derived from the eGeMAPS feature set: en… - `2026-06` · [AVSR-Gen](https://arxiv.org/abs/2606.07259) — Introduces MV2LRS3, a controlled unseen test set subsampled from MultiVSR to strictly match the acoustic, visual, and demographic distribution of LRS3, and shows that five state… - `2026-06` · [Audio-Oscar](https://arxiv.org/abs/2606.07397) — Audio-Oscar is a multi-agent framework that coordinates specialist agents (character and voice design, speech generation, fine-grained timeline planning, model selection, non-sp… - `2026-06` · [CogAudio-LLM](https://arxiv.org/abs/2606.06940) — CogAudio-LLM is a cognitive affective reasoning framework for audio language models that counters textual semantic dominance over acoustic nuance. - `2026-06` · [DSFA](https://arxiv.org/abs/2606.07494) — Proposes Domain-Shift Feature Augmentation (DSFA), which turns deterministic feature statistics into stochastic distributions during fine-tuning to simulate in-the-wild variatio… - `2026-06` · [KIT-IWSLT2026](https://arxiv.org/abs/2606.07240) — KIT's cross-lingual voice cloning system for the IWSLT 2026 track, built on the multilingual TTS model FishAudio-S
Excerpt of 7,386 characters
Read on GitHubBin Wang · Apodex · Singapore
101
18
10
4
Microsoft · United States
2
2
Yuan-Man · China
2
Chao-Wei Huang · Meta · United States
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:84ac01c657122384, topic:speech-recognition, topic:tts, topic:music-generation