FlashLabs-AI-Corp/FlashLabs-Chroma
quality grade C, 61 out of 100Worlds first open-source real-time end-to-end spoken dialogue model with personalized voice cloning.
- stars
- 550
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
915 results
Worlds first open-source real-time end-to-end spoken dialogue model with personalized voice cloning.
Voice Conversion by CycleGAN (语音克隆/语音转换): CycleGAN-VC2
Modified version of Chatterbox that accepts text files as input and no character restrictions. I use it to make audiobooks, especially for my kids.
ComfyUI custom node for the VibeVoice TTS. Expressive, long-form, multi-speaker conversational audio
Local, OpenAI-compatible text-to-speech (TTS) API using Chatterbox, enabling users to generate voice cloned speech anywhere the OpenAI API is used (e.g. Open WebUI, AnythingLLM, etc.)
MimikaStudio - A local-first application for macOS (Apple Silicon) + Agentic MCP Support
The code for the bark-voicecloning model. Training and inference.
A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zero-shot text-to-speech
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
A webui for different audio related Neural Networks
AI Podcast Generator for bilingual episodes, Multi Languages, Alternative to NotebookLLM;真人对话AI播客生成器,多语言,多音色
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
A Python/Pytorch app for easily synthesising human voices
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your ComfyUI workflows.
Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech, image generation and video generation APIs.
GPT-SoVITS ONNX Inference Engine & Model Converter
MARS5 speech model (TTS) from CAMB.AI
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
Netflix-level subtitle cutting, translation, alignment, and even dubbing - one-click fully automated AI video subtitle team | Netflix级字幕切割、翻译、对齐、甚至加上配音,一键全自动视频搬运AI字幕组
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
Clone a voice in 5 seconds to generate arbitrary speech in real-time
A real-time speech-to-speech chatbot powered by Whisper Small, Llama 3.2, and Kokoro-82M.
This repository will guide you to create your own Smart Virtual Assistant like Google Assistant using Open AI's ChatGPT, Whisper. The entire solution is created using Python & Gradio.
24,523 repositories in the index in total.