royshil/obs-localvocal
quality grade B, 70 out of 100OBS plugin for local speech recognition and captioning using AI
- stars
- 1.6k
- stars gained this week
- +9this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
910 results
OBS plugin for local speech recognition and captioning using AI
All in one Qwen3-ASR Server, compatible with OpenAI API
A native macOS menu bar dictation app using local speech-to-text with WhisperKit
🎙️ Speak with AI - Run locally using Ollama, OpenAI, Anthropic or xAI - Speech uses SparkTTS, OpenAI, ElevenLabs, Kokoro, Typecast or xAI
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
Companion application for Elite Dangerous
Vocello: a local, private voice studio for Apple Silicon. Write a script, pick or describe a voice, and generate speech on-device, faster than realtime on an 8 GB M2 Mac mini. Native Swift + MLX, no Python. Mac app out now, iPhone beta on TestFlight. (Formerly QwenVoice.)
Speed and samples benchmark: for all types of text to speech (TTS) models on Windows/Linux/Mac.
Cross-platform local voice typing and meeting transcription for macOS and Linux.
On-device streaming speech-to-text engine powered by deep learning
#1 Angular PDF viewer
Realtime Voice AI with 100+ Models on Arduino ESP32 with Secure Websockets and Edge Functions for AI Companions, and Devices
Vynaro (叙影 AI) - 下一代 7 步全自动 AI 影视解说与第一人称视频创作工具 (Tauri 2 + Rust + React 19)
Open source, local, and self-hosted Amazon Echo/Google Home competitive Voice Assistant alternative
audio.cpp with a full-task WebUI - pure C++ audio-model inference engine powered by ggml. TTS, ASR/STT, VAD, voice conversion, speaker diarization, music generation. No Python dependency.
Foundational Models for State-of-the-Art Speech and Text Translation
SALMONN family: A suite of advanced multi-modal LLMs
Open source voice dictation technology
Open source voice-to-text for the terminal. Record from a hotkey, transcribe with any provider, pipe to AI or shell commands.
Whisper.net. Speech to text made simple using Whisper Models
一站式全自动字幕生成软件,下载、转录、翻译、压制全流程覆盖,无需人工介入 / One-stop automated subtitle generator. Handles downloading, transcription, translation, and hardcoding—zero human intervention required.
Speech-to-text server framework with next-gen Kaldi
OpenAI Whisper ASR Webservice API
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
24,535 repositories in the index in total.