KoljaB/RealtimeTTS
quality grade A, 84 out of 100Converts text to speech in realtime
- stars
- 4.0k
- stars gained this week
- +5this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
910 results
Converts text to speech in realtime
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
A custom Home Assistant integration to play combined audio files before and/or after text-to-speech (TTS) messages
VRCT(VRChat Chatbox Translator & Transcription)
A free, open source, privacy-first voice input app for macOS.
🚀 AI 全自动化视频生成员工 | Your First AIGC Coworker. Chat an Idea. Get a Film. 🦞
Cross-platform speech toolset, used from the command-line or as a Node.js library. Includes a variety of engines for speech synthesis, speech recognition, forced alignment, speech translation, voice isolation, language detection and more.
Turn PDFs and EPUBs into audiobooks; subtitles or videos into dubbed videos (including translation), and more. For free. Pandrator uses local models, including voice-cloning (instant, RVC-enhanced, XTTS fine-tuning) and LLM processing. It aspires to be a user-friendly app with a GUI, an installer and all-in-one packages.
Open-source AI meeting copilot - real-time transcription, echo cancellation, and AI assistance. Captures system audio + mic, cancels echo via WebRTC AEC3, transcribes with Deepgram, and gives you Claude/OpenAI help during meetings. Runs locally on macOS and Windows.
open source audio and video transcription software
Captains log and 3d star map for Elite Dangerous
Augmentative and Alternative Communication (AAC) system with text-to-speech for the browser
Browse WhatsApp chat exports offline with AI-powered voice transcription. Privacy-first desktop/web app with bookmarks, search, and statistics. Built with SvelteKit and Electron.
Vonage API client for Node.js. API support for SMS, Voice, Text-to-Speech, Numbers, Verify (2FA) and more.
The official JavaScript (Node) library for the ElevenLabs API.
Botium Speech Processing
Native UI for the Whispering Tiger project - https://github.com/Sharrnah/whispering (live transcription / translation)
On-device Speech-to-Intent engine powered by deep learning
AI-powered multi-voice audiobook generator — LLM script annotation, voice cloning, voice design, LoRA training, per-line style control, and export to MP3, chaptered M4B, or Audacity multi-track. Built on Qwen3-TTS.
React Native binding of whisper.cpp.
Legacy Swift Hex app. Try the Rust rewrite at hex.kitlangton.com; new source at github.com/anomalyco/hex.
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio EditX, IndexTTS-2, Chatterbox (classic and multilingual), F5-TTS, Higgs Audio 2, 3, and VibeVoice with unlimited text length, SRT timing, Character support, and many audio tools
Pruna is a model optimization framework built for developers, enabling you to deliver faster, more efficient models with minimal overhead.
An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS
24,535 repositories in the index in total.