fabiogra/moseca
quality grade D, 45 out of 100A Streamilt web app for music source separation & karaoke
- stars
- 397
- stars gained this week
- —this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
910 results
A Streamilt web app for music source separation & karaoke
System audio capture + multi-provider ASR + local-first AI review workspace. Floating live captions, 12 ASR backends, 60+ languages, AI summary/chat/mindmap, Open API, MCP server, and Agent Skill.
Get started using Deepgram's Live Transcription with this Next.js demo app
T-one is a high-performance streaming ASR pipeline for Russian, specialized for the telephony domain.
Open source speech to text models for Indic Languages
On-device speech-to-text engine powered by deep learning
A speech recognition library running in the browser thanks to a WebAssembly build of Vosk
A React component to make correcting automated transcriptions of audio and video easier and faster. By BBC News Labs. - Work in progress
:speech_balloon: /so.nus/ STT (speech to text) for Node with offline hotword detection
Voice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式
Machine Learning Training Utilities (for TensorFlow and PyTorch)
End-to-End speech recognition implementation base on TensorFlow (CTC, Attention, and MTL training)
State-of-the-art (ranked #1 Aug 2022) German Speech Recognition in 284 lines of C++. This is a 100% private 100% offline 100% free CLI tool.
The official repository of the Eesen project
中文语音识别
:zap: TensorFlowASR: Almost State-of-the-art Automatic Speech Recognition in Tensorflow 2. Supported languages that can use characters or subwords
An AI for Music Generation
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统
Vietnamese Text to Speech library
🔊 Create labeled datasets, enhance audio quality, identify speakers, support diverse dataset types. 🎧👥📊 Advanced audio processing.
Sayna is a unified Voice Layer for AI Agents with a seemless integration to an existing agentic frameworks
24,537 repositories in the index in total.