megaease/easevoice-trainer
quality grade C, 50 out of 100EaseVoice Trainer is a simple and user-friendly voice cloning and speech model trainer.
- stars
- 350
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
910 results
EaseVoice Trainer is a simple and user-friendly voice cloning and speech model trainer.
Tacotron 2 - PyTorch implementation with faster-than-realtime inference modified to enable cross lingual voice cloning.
VoxNovel: generate audiobooks giving each character a different voice actor.
Talk to 峰哥 — 克隆任何人的声音和性格,实时语音对话,工程延迟 < 1 秒 | Clone anyone's voice & personality for real-time conversation. < 1s engineering latency.
Fuse ChatTTS with OpenVoice, upload a 10-second audio clip, and clone your personalized ChatTTS voice.
ComfyUI node for highly expressive speech and realistic zero-shot voice cloning
OmniVoice TTS nodes for ComfyUI - Zero-shot multilingual text-to-speech with voice cloning, voice design, and multi-speaker dialogue
Run Qwen3-TTS text-to-speech locally on Mac (M1/M2/M3/M4). Voice cloning, voice design, custom voices. 100% offline using MLX.
Worlds first open-source real-time end-to-end spoken dialogue model with personalized voice cloning.
Voice Conversion by CycleGAN (语音克隆/语音转换): CycleGAN-VC2
Modified version of Chatterbox that accepts text files as input and no character restrictions. I use it to make audiobooks, especially for my kids.
ComfyUI custom node for the VibeVoice TTS. Expressive, long-form, multi-speaker conversational audio
Local, OpenAI-compatible text-to-speech (TTS) API using Chatterbox, enabling users to generate voice cloned speech anywhere the OpenAI API is used (e.g. Open WebUI, AnythingLLM, etc.)
MimikaStudio - A local-first application for macOS (Apple Silicon) + Agentic MCP Support
The code for the bark-voicecloning model. Training and inference.
A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zero-shot text-to-speech
A webui for different audio related Neural Networks
AI Podcast Generator for bilingual episodes, Multi Languages, Alternative to NotebookLLM;真人对话AI播客生成器,多语言,多音色
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voice cloning, and large audiobook-scale text processing. Runs accelerated on NVIDIA (CUDA), AMD (ROCm), and CPU.
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
A Python/Pytorch app for easily synthesising human voices
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your ComfyUI workflows.
Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech, image generation and video generation APIs.
MARS5 speech model (TTS) from CAMB.AI
24,535 repositories in the index in total.