Starmel/OpenSuperWhisper
quality grade B, 71 out of 100macOS dictation app
- stars
- 2.4k
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
915 results
macOS dictation app
No description
A nearly-live implementation of OpenAI's Whisper.
Port of OpenAI's Whisper model in C/C++
A GUI tool for offline transcription of speech recordings, including speaker diarization, utilizing state-of-the-art machine learning models.
Foundational Model for Speech Recognition Tasks
Silero Models: pre-trained text-to-speech models made embarrassingly simple
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
On-device voice activity detection (VAD) powered by deep learning
Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
The lightest Vietnamese Text-to-Speech with Multi-Speaker TTS and Zero-Shot Voice Cloning.
💬 Speech recognition for your site
Cross-platform speech toolset, used from the command-line or as a Node.js library. Includes a variety of engines for speech synthesis, speech recognition, forced alignment, speech translation, voice isolation, language detection and more.
End-to-End Speech Processing Toolkit
🔥🔥 Kokoro in Rust. https://huggingface.co/hexgrad/Kokoro-82M Insanely fast, realtime TTS with high quality you ever have.
a free and open source speech synthesizer for Russian and other languages
OpenClaw voice assistant app for Android - Wake word activation & system assistant integration
Управление устройствами из Home Assistant через Алису (Умный дом Яндекса) или Марусю
AI-powered multi-voice audiobook generator — LLM script annotation, voice cloning, voice design, LoRA training, per-line style control, and export to MP3, chaptered M4B, or Audacity multi-track. Built on Qwen3-TTS.
She's the AI agent you come home to.
Mantella is a Skyrim and Fallout 4 mod which allows you to naturally speak to NPCs using a Speech-to-Text → LLMs → Text-to-Speech pipeline
The media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation, and more!
24,523 repositories in the index in total.