andabi/voice-vector
quality grade D, 42 out of 100Deep neural networks for getting text-independent speaker embedding written in TensorFlow
- stars
- 310
- stars gained this week
- —this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
915 results
Deep neural networks for getting text-independent speaker embedding written in TensorFlow
Predicting depression from acoustic features of speech using a Convolutional Neural Network.
Real-time Voice Activity Detection in Noisy Eniviroments using Deep Neural Networks
Pitch Estimating Neural Networks (PENN)
Automatic Music Transcription with Deep Neural Networks
Deep Convolutional Neural Networks for Musical Source Separation
This is the code for "Neural Network Voices" by Siraj Raval on Youtube
A realtime live transcription and translation app built with Huggingface Transformer.js and Supabase Realtime.
No description
No description
No description
plug whisper audio transcription to a local ollama server and ouput tts audio responses
Voice interface for Claude Code via SIP/3CX - Call your AI, and your AI can call you
Speech-to-speech AI assistant with natural conversation flow, mid-speech interruption, vision capabilities and AI-initiated follow-ups. Features low-latency audio streaming, dynamic visual feedback, and works with local LLM/TTS services via OpenAI-compatible endpoints.
低成本的简单基于live2d TTS文字转语音和大模型聊天的直播解决方案
No description
This is an evolving repo for the paper "Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey".
CAT is more than a CRF-based ASR toolkit: it provides a complete workflow for data-efficient end-to-end ASR, supporting CTC, CTC-CRF, RNN-T, and language-model training and inference.
An Audio Language model for Audio Tasks
Public release of the Sound Effect Foundation model by Sony AI.
ACE-Step: A Step Towards Music Generation Foundation Model
Real-time Speech-Text Foundation Model Toolkit (wip)
SoftWhisper simplifies audio and video transcription using the powerful Whisper model. Easily select custom models, languages, and tasks, fine-tune transcription with beam size adjustment, and specify start and end times for targeted segments.
Fine Tune the Style-TTS2 Voice Model
24,523 repositories in the index in total.