babysor/MockingBird
quality grade C, 56 out of 100🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
- stars
- 37k
- stars gained this week
- —this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
915 results
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
Instant voice cloning by MIT and MyShell. Audio foundation model.
A generative speech model for daily dialogue.
Meet Ava, the WhatsApp Agent
AI-Powered Video Retrieval & Clipping Tool
Fast text based video editing, node Electron Os X desktop app, with Backbone front end.
an editor for spoken-word audio with automatic transcription
🚀 AI 全自动化视频生成员工 | Your First AIGC Coworker. Chat an Idea. Get a Film. 🦞
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
Documentation and Wiki for SEPIA. Please post your questions and bug-reports here in the issues section! Thank you :-)
Example projects built with the Hume AI APIs
This is an on-CPU real-time conversational system for two-way speech communication with AI models, utilizing a continuous streaming architecture for fluid conversations with immediate responses and natural interruption handling.
🗣 An overlay that gets your user’s voice permission and input as text in a customizable UI
Embeddable custom voice assistant for Android applications
Open Voice Operating System - Buildroot edition is a minimalistic linux OS bringing the OVOS voice assistant to embbeded, low-spec headless and/or small (touch)screen devices.
Private voice keyboard, agent, AI chat, images, webcam, recordings, voice control with >= 4 GiB of VRAM.
Genie As A Service and Thingpedia
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM
Espressif's Voice Assistant SDK: Alexa, Google Voice Assistant, Google DialogFlow
A versatile and extensible platform for automation with hundreds of supported integrations
Typeflux is a macOS menu bar voice input tool built with Swift. It is designed for a fast "hold to talk, release to insert" workflow: press a hotkey, speak naturally, let the app transcribe your speech, and send the resulting text back into the currently focused app.
Patching for XiaoAi Speakers (小爱音箱), add custom binaries and open source software. Tested on LX06, LX01, LX05, L09A
DIY Voice Assistant based on the GLaDOS character from Portal video game series. Works with home assistant!
A voice assistant 🗣️ which can be used to interact with your computer 💻 and controls your pc operations 🎛️
24,523 repositories in the index in total.