jishengpeng/ControlSpeech
quality grade F, 32 out of 100[ACL 2025 Main] ControlSpeech: Towards Simultaneous Zero-shot Speaker Cloning and Zero-shot Language Style Control With Decoupled Codec
- stars
- 277
- stars gained this week
- —this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
910 results
[ACL 2025 Main] ControlSpeech: Towards Simultaneous Zero-shot Speaker Cloning and Zero-shot Language Style Control With Decoupled Codec
Implementation of Spear-TTS - multi-speaker text-to-speech attention network, in Pytorch
This AI Smart Speaker uses speech recognition, TTS (text-to-speech), and STT (speech-to-text) to enable voice and vision-driven conversations, with additional web search capabilities via OpenAI and Langchain agents.
VocGAN: A High-Fidelity Real-time Vocoder with a Hierarchically-nested Adversarial Network
Speech to text (PocketSphinx, Iflytex API, Baidu API) and text to speech (pyttsx3) | 语音转文字(PocketSphinx、百度 API、科大讯飞 API)和文字转语音(pyttsx3)
Library to build speech synthesis systems designed for easy and fast prototyping.
Assistente pessoal virtual desenvolvida com Python 🤖
A desktop application that uses AI to translate voice between languages in real time, while preserving the speaker's tone and emotion.
PyTorch Implementation of FastDiff (IJCAI'22)
PyTorch Implementation of ProDiff (ACM-MM'22) with a Extremely-Fast diffusion speech synthesis pipeline
:speech_balloon: Reverse Engineering Google's Speech To Text API (v2)
A Vietnamese Voice Cloning Text-to-Speech Model ✨
Fully automated video maker using motion graphics and text-to-speech synthesis to turn newsletters into daily YouTube videos.
Implementation of Voicebox, new SOTA Text-to-speech network from MetaAI, in Pytorch
Speech to Text and KB input captions for OBS, VRChat, Twitch chat and Discord
Vonage REST API client for PHP. API support for SMS, Voice, Text-to-Speech, Numbers, Verify (2FA) and more.
TTS model capable of streaming conversational audio in realtime.
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
This is now the official location of the Merlin project.
the open-source virtual assistant for Ubuntu based Linux distributions
An awesome browser extension that reads aloud webpage content with one click
DeepMind's Tacotron-2 Tensorflow implementation
Offline Text To Speech synthesis for python
🚀 一键部署(含离线整合包)!基于 ChatTTS ,支持流式输出、音色抽卡、长音频生成和分角色朗读。简单易用,无需复杂安装。
24,537 repositories in the index in total.