Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
| Date | Stars |
|---|---|
| 2026-07-24 | 10171 |
| 2026-07-25 | 10171 |
| 2026-07-28 | 10171 |
| 2026-07-30 | 10171 |
| 2026-08-06 | 10171 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head [](https://arxiv.org/abs/2304.12995) [](https://github.com/AIGC-Audio/AudioGPT)  [](https://huggingface.co/spaces/AIGC-Audio/AudioGPT) We provide our implementation and pretrained models as open source in this repository. ## Get Started Please refer to [run.md](run.md) ## Capabilities Here we list the capability of AudioGPT at this time. More supported models and tasks are coming soon. For prompt examples, refer to [asset](assets/README.md). Currently not every model has repository. ### Speech | Task | Supported Foundation Models | Status | |:--------------------------:|:-------------------------------:|:------:| | Text-to-Speech | [FastSpeech](https://github.com/ming024/FastSpeech2), [SyntaSpeech](https://github.com/yerfor/SyntaSpeech), [VITS](https://github.com/jaywalnut310/vits) | Yes (WIP) | | Style Transfer | [GenerSpeech](https://github.com/Rongjiehuang/GenerSpeech) | Yes | | Speech Recognition | [whisper](https://github.com/openai/whisper), [Conformer](https://github.com/sooftware/conformer) | Yes | | Speech Enhancement | [ConvTasNet]() | Yes (WIP) | | Speech Separation | [TF-GridNet](https://arxiv.org/pdf/2211.12433.pdf) | Yes (WIP) | | Speech Translation | [Multi-decoder](https://arxiv.org/pdf/2109.12804.pdf) | WIP | | Mono-to-Binaural | [NeuralWarp](https://github.com/fdarmon/NeuralWarp) | Yes | ### Sing | Task | Supported Foundation Models | Status | |:-------------------------:|:-------------------------------:|:------:| | Text-to-Sing | [DiffSinger](https://github.com/MoonInTheRiver/DiffSinger), [VISinger](https://github.com/jerryuhoo/VISinger) | Yes (WIP) | ### Audio | Task | Supported Foundation Models | Status | |:----------------------:|:---------------------------:|:------:| | Text-to-Audio | [Make-An-Audio]() | Yes | | Audio Inpainting | [Make-An-Audio]() | Yes | | Image-to-Audio | [Make-An-Audio]() | Yes | | Sound Detection | [Audio-transformer](https://github.com/RetroCirce/HTS-Audio-Transformer) | Yes | | Target Sound Detection | [TSDNet](https://github.com/gy65896/TSDNet) | Yes | | Sound Extraction | [LASSNet](https://github.com/liuxubo717/LASS) | Yes | ### Talking Head | Task | Supported Foundation Models | Status | |:-------------------------:|:-------------------------------:|:----------:| | Talking Head Synthesis | [GeneFace](https://github.com/yerfor/GeneFace) | Yes (WIP) | ## Acknowledgement We appreciate the open source of the following projects: [ESPNet](https://github.com/espnet/espnet)   [NATSpeech](https://github.com/NATSpeech/NATSpeech)   [Visual ChatGPT](https://github.com/microsoft/visual-chatgpt)   [Hugging Face](https://github.com/huggingface)   [LangChain](https://github.com/hwchase17/langchain)   [Stable Diffusion](https://github.com/CompVis/stable-diffusion)  
Excerpt of 3,637 characters
Read on GitHub47
Facebook AI Research (FAIR) · United States
43
21
6
Zhenhui Ye · Zhejiang University
3
2
Jiatong · Anuttacon · United States
1
Jinglin Liu · Zhejiang University · China
1
Yi Ren · HeyGen · Singapore
1
Yuning Wu
1
xuankai@cmu · Carnegie Mellon University · United States
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:78400e2301b02559, topic:talking-head, desc:talking head, readme:talking head
matched fp:78400e2301b02559, topic:gpt