MetaWu2077/Esp32_VoiceChat_LLMs
quality grade F, 23 out of 100esp32 based device, mainly used for voice chat with large language models
- stars
- 808
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
913 results
esp32 based device, mainly used for voice chat with large language models
MiMo-Audio: Audio Language Models are Few-Shot Learners
SpeechGPT Series: Speech Large Language Models
Open Source Speech Language Model
ACE-Step: A Step Towards Music Generation Foundation Model
Text-audio foundation model from Boson AI
Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers.
TTSFM mirrors OpenAI's TTS service, providing a compatible interface for text-to-speech conversion with multiple voice options for free.
Automatically Update Text-to-speech (TTS) Papers Daily using Github Actions (Update Every 12th hours)
TTS (text to speech) for node.js. send text from node.js to your speakers.
Kyutai's Speech-To-Text and Text-To-Speech models based on the Delayed Streams Modeling framework.
Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
An implementation of Microsoft's "FastSpeech 2: Fast and High-Quality End-to-End Text to Speech"
A fast local neural text to speech engine for Mycroft
AwesomeTTS text-to-speech add-on for Anki
React Native Text-To-Speech library for Android and iOS
Miso TTS is an 8 billion, highly emotive text-to-speech model
Text-to-Speech in JavaScript using eSpeak
Zero-Shot Speech Editing and Text-to-Speech in the Wild
A book about Text-to-Speech (TTS) in Chinese.
GLM-ASR-Nano: A robust, open-source speech recognition model with 1.5B parameters
Omni SenseVoice: High-Speed Speech Recognition with words timestamps 🗣️🎯
A 10000+ hours dataset for Chinese speech recognition
Speech-to-Text-WaveNet : End-to-end sentence level English speech recognition based on DeepMind's WaveNet and tensorflow
24,522 repositories in the index in total.