espressif/esp-sr
quality grade C, 56 out of 100Speech recognition
- stars
- 1.5k
- stars gained this week
- +4this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
909 results
Speech recognition
On-device wake word detection powered by deep learning
Deep learning for audio processing
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
Dicio assistant app for Android
List of open-source TTS, voice cloning, and music generation models
GPT-SoVITS ONNX Inference Engine & Model Converter
Music Analysis, Chord Recognition, Beat Tracking, Guitar Diagrams, Piano Visualizer, Lyrics Transcription Application, context-aware LLM inference for analysis from uploaded audio and YouTube video
A lightweight text-to-speech model with zero-shot voice cloning
The Naomi Project is an open source, technology agnostic platform for developing always-on, voice-controlled applications!
Speech to Text to Speech. Song now playing. Sends text as OSC messages to VRChat to display on avatar. (STTTS) (Speech to TTS) (VRC STT System) (VTuber TTS)
faster_whisper GUI with PySide6
A complete video subtitle editing React component with AI-powered speech recognition and visual editing capabilities.
Closed Captioning OBS plugin using Google Speech Recognition
A Python library for solving reCAPTCHA v2 and v3 with Playwright
No description
[ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matching
Set of 📝 with 🔗 to help those building Voice AI agents 🎙️🤖
Tero Subtitler is an open source, cross-platform, and free subtitle editing software.
This is an evolving repo for the paper "Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey".
On-device voice activity detection (VAD) powered by deep learning
Transcribe any video URL or audio file into plaintext. No GPU. No cloud. One command.
@voicybot Telegram bot main repository
Custom nodes that extend the capabilities of Comfyui
24,538 repositories in the index in total.