huawei-noah/Speech-Backbones
quality grade F, 24 out of 100This is the main repository of open-sourced speech technology by Huawei Noah's Ark Lab.
- stars
- 604
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
913 results
This is the main repository of open-sourced speech technology by Huawei Noah's Ark Lab.
RNN-based generative models for speech.
Voice Conversion Tool Kit
Machine learning based speech synthesis Electron app, with voices from specific characters from video games
INTERSPEECH 2023-2024 Papers: A complete collection of influential and exciting research papers from the INTERSPEECH 2023-24 conference. Explore the latest advances in speech and language processing. Code included. Star the repository to support the advancement of speech technology!
A talking LLM that runs on your own computer without needing the internet.
Flowtron is an auto-regressive flow-based generative network for text to speech synthesis with control over speech variation and style transfer
AutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
Official PyTorch implementation of BigVGAN (ICLR 2023)
A high-quality speech analysis, manipulation and synthesis system
Implementation of Natural Speech 2, Zero-shot Speech and Singing Synthesizer, in Pytorch
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
Software Automatic Mouth - Tiny Speech Synthesizer
An open-source ChatGPT app with a voice
WaveNet vocoder
Open singing synthesis platform / Open source UTAU successor
Open source voice-to-text for the terminal. Record from a hotkey, transcribe with any provider, pipe to AI or shell commands.
Snap any video URL or audio file into plaintext. No GPU. No cloud. One command.
An Android app that offers speech-to-text user interfaces to other apps
Working online speech recognition based on RNN Transducer. ( Trained model release available in release )
Voice assistant for Visual Studio Code.
Ultra fast and portable Parakeet implementation for on-device inference in C++ using Axiom with MPS+Unified Memory
Browser-only Canva-style presentation studio, powered by local Web AI.
Fully local, no dependency scribe. Speak into your microphone and summarize. Requires iOS 26 and MacOS 26 to use the advanced transcription model and foundational model for summaries
24,524 repositories in the index in total.