alphacep/vosk-server
quality grade F, 34 out of 100WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries
- stars
- 1.3k
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
913 results
WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries
SALMONN family: A suite of advanced multi-modal LLMs
Alias is a teachable “parasite” that is designed to give users more control over their smart assistants, both when it comes to customisation and privacy. Through a simple app the user can train Alias to react on a custom wake-word/sound, and once trained, Alias can take control over your home assistant by activating it for you.
Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows, Raspberry Pi, VisionFive2, LicheePi4A etc.
Attempt at tracking states of the arts and recent results (bibliography) on speech recognition.
Open-Source Large Vocabulary Continuous Speech Recognition Engine
中文语音识别; Mandarin Automatic Speech Recognition;
pytorch-kaldi is a project for developing state-of-the-art DNN/RNN hybrid speech recognition systems. The DNN part is managed by pytorch, while feature extraction, label computation, and decoding are performed with the kaldi toolkit.
Facebook AI Research's Automatic Speech Recognition Toolkit
FACodec: Speech Codec with Attribute Factorization used for NaturalSpeech 3
Refurbished Arduino version of the Talkie library from Peter Knight.
AI-WEBUI: A universal web interface for AI creation, 一款好用的图像、音频、视频AI处理工具
Global Rhythm Style Transfer Without Text Transcriptions
PyTorch implementation of Tacotron speech synthesis model.
PPG-Based Voice Conversion
UTokyo-SaruLab MOS Prediction System
✨✨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
A python wrapper for Speech Signal Processing Toolkit (SPTK).
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
Soft speech units for voice conversion
Audio Development Tools (ADT) is a project for advancing sound, speech, and music technologies, featuring components for machine learning, sound synthesis, speech and music generation, signal processing, game audio, digital audio workstations (DAWs), and more.
PyTorch implementation of GAN-based text-to-speech synthesis and voice conversion (VC)
StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion
High-quality and streaming Speech-to-Speech interactive agent in a single file. 只用一个文件实现的流式全双工语音交互原型智能体!
24,524 repositories in the index in total.