Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A Gradio-based web UI for voice cloning and voice design, powered by Qwen3-TTS & VibeVoice. Can use Whisper or VibeVoice-ASR for automatic transcription.
| Date | Stars |
|---|---|
| 2026-07-31 | 509 |
| 2026-08-02 | 511 |
| 2026-08-06 | 511 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Voice Clone Studio A multi-model, modular Gradio-based web UI for voice cloning, voice design, multi-speaker conversation, voice conversion, voice training and sound effects. Basically, One app, many engines, to tinker with all of them without juggling separate repos or setups. Powered by [VibeVoice](https://github.com/microsoft/VibeVoice), [Qwen3-TTS](https://github.com/QwenLM/Qwen3-TTS), [LuxTTS](https://github.com/ysharma3501/LuxTTS), [Chatterbox](https://github.com/resemble-ai/chatterbox), [Fish Speech](https://github.com/fishaudio/fish-speech) and [MMAudio](https://github.com/hkchengrex/MMAudio). Supports [Qwen3-ASR](https://github.com/QwenLM/Qwen3-ASR), [VibeVoice-ASR](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-asr.md) and [Whisper](https://github.com/openai/whisper) for automatic transcription. As well as [llama.cpp](https://github.com/ggerganov/llama.cpp) and [Ollama](https://ollama.com/) for Prompt Generation and a Prompt Saving, based on [ComfyUI Prompt-Manager](https://github.com/FranckyB/ComfyUI-Prompt-Manager) <img src="https://img.shields.io/badge/VibeVoice-TTS-green" alt="VibeVoice TTS"> <img src="https://img.shields.io/badge/VibeVoice-ASR-green" alt="VibeVoice ASR"> <img src="https://img.shields.io/badge/Qwen3-TTS-blue" alt="Qwen3-TTS"> <img src="https://img.shields.io/badge/Qwen3-ASR-blue" alt="Qwen3-ASR"> <img src="https://img.shields.io/badge/LuxTTS-TTS-orange" alt="LuxTTS"> <img src="https://img.shields.io/badge/Chatterbox-TTS-red" alt="Chatterbox-TTS"> <img src="https://img.shields.io/badge/Fish_Speech-TTS-teal" alt="Fish Speech TTS"> <img src="https://img.shields.io/badge/Whisper-yellow" alt="Whisper"> <img src="https://img.shields.io/badge/MMAudio-SFX-purple" alt="MMAudio"> <a href="docs/preview.png"><img src="docs/preview.png" alt="Voice Clone Studio Preview" width="600"></a> ## Architecture Voice Clone Studio is fully modular. The main file dynamically loads self-contained tools as tabs. Each tool can be enabled or disabled from Settings without touching any code. It supports multipe engine for voice cloning, as well as Model finetuning. More features are also planned. ## Features ### Voice Clone Clone voices from your own audio samples. Provide a short reference audio clip with its transcript, and generate new speech in that voice. - **Multiple engines** - Qwen3-TTS (0.6B/1.7B), VibeVoice (1.5B/Large/Large-4bit), LuxTTS, Chatterbox, and Fish Speech S2 Pro (4B) - **Fish Speech Expression Tags** - Embed `[tag]` markers like `[whisper]`, `[laughing]`, `[excited]` directly in text for fine-grained delivery control (15,000+ supported tags) - **Automatic Tag Stripping** - Fish Speech `[tags]` are automatically removed when using other engines, so the same text works everywhere - **Voice prompt caching** - First generation processes the sample, subsequent ones are instant - **Seed control** - Reproducible results with saved seeds - **Emotion presets** - 40+ emotion presets with adjustable intensity - **Split by Paragraph** - Generate a separate audio clip for each paragraph, with automatic naming and a combined preview - **Prompt Hub** - Access saved prompts directly from the tool without switching tabs - **Metadata tracking** - Each output saves generation info (sample, seed, text) ### Conversation Create multi-speaker dialogues using either Qwen's premium voices or your own custom voice samples using VibeVoice: **Choose Your Engine:** - **Qwen** - Fast generation with 9 preset voices, optimized for their native languages - **VibeVoice** - High-quality custom voices, up to 90 minutes continuous, perfect for podcasts/audiobooks - **LuxTTS** - **Unified Script Format:** Write scripts using `[N]:` format - works seamlessly with both engines: ``` [1]: Hey, how's it going? [2]: I'm doing great, thanks for asking! [3]: Mind if I join this conversation? ``` **Qwen Mode:** - Mix any of the 9 premium speakers - Adjustable pause duration between lines - Fast generation with cache
Excerpt of 22,539 characters
Read on GitHub155
7
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:ac8311e4e327174e, desc:voice cloning, desc:transcription