Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
| Date | Stars |
|---|---|
| 2026-07-31 | 1163 |
| 2026-08-02 | 1163 |
| 2026-08-06 | 1163 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<div align="center" style="display: flex; justify-content: center; align-items: center; text-align: center;">
<a href="https://github.com/NVIDIA/audio-flamingo" style="margin-right: 20px; text-decoration: none; display: flex; align-items: center;">
<img src="static/logo-no-bg.png" alt="Audio Flamingo 3 🔥🚀🔥" width="120">
</a>
</div>
<div align="center" style="display: flex; justify-content: center; align-items: center; text-align: center;">
<h2>
Audio Flamingo: Series of Advanced Audio Understanding Language Models
</h2>
</div>
## Overview
In this repo, we present the **Audio Flamingo** series of advanced audio understanding Language models:
- [Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities](https://arxiv.org/abs/2402.01831) (ICML 2024)
- [Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities](https://arxiv.org/abs/2503.03983) (ICML 2025)
- [Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models](https://arxiv.org/abs/2507.08128) (NeurIPS 2025, Spotlight)
- [Music Flamingo: Scaling Music Understaning in Audio Language Models]() (arxiv)
## Music Flamingo (arXiv)
<div align="center" style="display: flex; justify-content: center; margin-top: 10px;">
<a href="https://arxiv.org/abs/2511.10289"><img src="https://img.shields.io/badge/arXiv-2511.10289-AD1C18" style="margin-right: 5px;"></a>
<a href="https://research.nvidia.com/labs/adlr/MF/"><img src="https://img.shields.io/badge/Demo page-228B22" style="margin-right: 5px;"></a>
<a href="https://github.com/NVIDIA/audio-flamingo/tree/music_flamingo"><img src='https://img.shields.io/badge/Github-Music--Flamingo-9C276A' style="margin-right: 5px;"></a>
<a href="https://github.com/NVIDIA/audio-flamingo/stargazers"><img src="https://img.shields.io/github/stars/NVIDIA/audio-flamingo.svg?style=social"></a>
</div>
<div align="center" style="display: flex; justify-content: center; margin-top: 10px; flex-wrap: wrap; gap: 5px;">
<a href="https://huggingface.co/nvidia/music-flamingo-hf">
<img src="https://img.shields.io/badge/🤗-Checkpoints-ED5A22.svg">
</a>
<a href="https://huggingface.co/datasets/nvidia/MF-Skills">
<img src="https://img.shields.io/badge/🤗-Dataset: MF--Skills-ED5A22.svg">
</a>
</div>
<div align="center" style="display: flex; justify-content: center; margin-top: 10px;">
<a href="https://huggingface.co/spaces/nvidia/music-flamingo"><img src="https://img.shields.io/badge/🤗-Gradio Demo (7B)-5F9EA0.svg" style="margin-right: 5px;"></a>
</div>
Music Flamingo (MF) is a fully open, state-of-the-art Large Audio-Language Model (LALM) built on Audio Flamingo 3 backbone, designed to advance music (including song) understanding in foundational audio models. MF brings together innovations in:
- Deep music understanding across songs and instrumentals.
- Rich, theory-aware captions and question answering (harmony, structure, timbre, lyrics, cultural context).
- Reasoning-centric training using chain-of-thought + reinforcement learning with custom rewards for step-by-step reasoning.
- Long-form song reasoning over full-length, multicultural audio (extended context).
Extensive evaluations confirm Music Flamingo's effectiveness, setting new benchmarks on over 10+ public music understanding and reasoning tasks.
<div align="center">
<img class="img-full" src="static/MF-comparison.png" width="800">
</div>
<div align="center">
<img class="img-full" src="static/MF-architecture.png" width="800">
</div>
<div align="center">
<img class="img-full" src="static/MF-donut_chart.png" width="400">
</div>
<div align="center">
<img class="img-full" src="static/MF-benchmarks.png" width="600">
</div>
## Audio Flamingo 3 (NeurIPS 2025 Spotlight)
<div align="center" style="display: flex; justify-content: center; margin-top: 10px;">
<a href="https://arxiv.org/abs/2507.08128"><img src="https://img.shields.io/badge/arXiv-2Excerpt of 17,119 characters
Read on GitHub41
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:62f5b3a0018f6fb5, llm:Topics: audio-captioning, audio-language-models, audio-question-answering, audio-reasoning, multimodal-large-language-models; description: PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
matched fp:62f5b3a0018f6fb5, llm:Topics: audio-captioning, audio-language-models, audio-question-answering, audio-reasoning, multimodal-large-language-models; description: PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
matched fp:62f5b3a0018f6fb5, llm:Topics: audio-captioning, audio-language-models, audio-question-answering, audio-reasoning, multimodal-large-language-models; description: PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
matched fp:62f5b3a0018f6fb5, llm:Topics: audio-captioning, audio-language-models, audio-question-answering, audio-reasoning, multimodal-large-language-models; description: PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models