JamezQ/Palaver
quality grade F, 22 out of 100Linux Speech Recognition
- stars
- 415
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
915 results
Linux Speech Recognition
LingEcho is an intelligent voice interaction platform that provides a comprehensive AI voice interaction solution. It integrates advanced speech recognition (ASR), text-to-speech (TTS), large language models (LLM) and real-time communication, supporting real-time calls, voice cloning, knowledge base management and other enterprise-level features.
Codex skill for complete paper-cut collage ad production, local IndexTTS-2 voice cloning, animation, audio and MP4 QC
ComfyUI custom nodes for Fish Audio S2-Pro TTS — voice clone, multi-speaker, and text-to-speech
OpenAI compatible TTS for Sesame CSM:1b & dia:1.6b - Voice Cloning from File/YT
Open-source text-to-speech for European languages with voice cloning
A Unified Framework for Expressive Speech Synthesis with Voice Cloning
This is a speech interaction system built on an open-source model, integrating ASR, LLM, and TTS in sequence. The ASR model is SenceVoice, the LLM models are QWen2.5-0.5B/1.5B, and there are three TTS models: CosyVoice, Edge-TTS, and pyttsx3
PC 端语音输入工具,离线识别,高准确率、低延迟,支持热词、LLM润色。按住CapsLock或鼠标侧键X2说话,松开自动上屏。
Non-invasive decoding of typed sentences from MEG and EEG brain recordings using a convolutional encoder, transformer, and character-level language model.
The official repo of Qwen-Audio (通义千问-Audio) chat & pretrained large audio language model proposed by Alibaba Cloud.
Fun-Audio-Chat is a Large Audio Language Model built for natural, low-latency voice interactions.
Inference code for the paper "Spirit-LM Interleaved Spoken and Written Language Model".
The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.
A simple toy demo of a local voice assistant with whisper and large language model.
NotaGen: Advancing Musicality in Symbolic Music Generation with Large Language Model Training Paradigms
esp32 based device, mainly used for voice chat with large language models
MiMo-Audio: Audio Language Models are Few-Shot Learners
SpeechGPT Series: Speech Large Language Models
Open Source Speech Language Model
Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
ACE-Step: A Step Towards Music Generation Foundation Model
Text-audio foundation model from Boson AI
Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers.
24,523 repositories in the index in total.