zzw922cn/Automatic_Speech_Recognition
quality grade C, 53 out of 100End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
- stars
- 2.8k
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
913 results
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统
Vietnamese Text to Speech library
🔊 Create labeled datasets, enhance audio quality, identify speakers, support diverse dataset types. 🎧👥📊 Advanced audio processing.
Sayna is a unified Voice Layer for AI Agents with a seemless integration to an existing agentic frameworks
[ACL 2025 Main] ControlSpeech: Towards Simultaneous Zero-shot Speaker Cloning and Zero-shot Language Style Control With Decoupled Codec
Implementation of Spear-TTS - multi-speaker text-to-speech attention network, in Pytorch
The Naomi Project is an open source, technology agnostic platform for developing always-on, voice-controlled applications!
This AI Smart Speaker uses speech recognition, TTS (text-to-speech), and STT (speech-to-text) to enable voice and vision-driven conversations, with additional web search capabilities via OpenAI and Langchain agents.
VocGAN: A High-Fidelity Real-time Vocoder with a Hierarchically-nested Adversarial Network
Speech to text (PocketSphinx, Iflytex API, Baidu API) and text to speech (pyttsx3) | 语音转文字(PocketSphinx、百度 API、科大讯飞 API)和文字转语音(pyttsx3)
小牛视频翻译 是一款支持本地视频翻译、字幕翻译和 YouTube 视频翻译下载的 AI 工具,集成自动语音识别与多语言翻译功能,助力创作者高效完成视频翻译,应用于视频本地化与视频出海场景。
Library to build speech synthesis systems designed for easy and fast prototyping.
Assistente pessoal virtual desenvolvida com Python 🤖
A desktop application that uses AI to translate voice between languages in real time, while preserving the speaker's tone and emotion.
PyTorch Implementation of FastDiff (IJCAI'22)
PyTorch Implementation of ProDiff (ACM-MM'22) with a Extremely-Fast diffusion speech synthesis pipeline
:speech_balloon: Reverse Engineering Google's Speech To Text API (v2)
A Vietnamese Voice Cloning Text-to-Speech Model ✨
No description
📚 A customizable dictionary extension that supports double-click lookups in 20+ languages, 1000+ dictionaries, text-to-speech, translation and Anki integration.
Fully automated video maker using motion graphics and text-to-speech synthesis to turn newsletters into daily YouTube videos.
Implementation of Voicebox, new SOTA Text-to-speech network from MetaAI, in Pytorch
Speech to Text and KB input captions for OBS, VRChat, Twitch chat and Discord
24,524 repositories in the index in total.