rzru/nightingale
quality grade B, 70 out of 100Machine learning powered Karaoke app (with scores!)
- stars
- 1.5k
- stars gained this week
- +11this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
910 results
Machine learning powered Karaoke app (with scores!)
A nearly-live implementation of OpenAI's Whisper.
A fully local and private Speech-To-Text app, offering multiple model backends, diarization & calendar mode - Available for Windows, macOS & Linux
A simple, high-quality voice conversion tool focused on ease of use and performance.
very fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in Rust
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
A robust, efficient, low-latency speech-to-text library with advanced voice activity detection, wake word activation and instant transcription.
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.
Fully local, private and cross platform Speech-to-Text with LLM Post-processing
The official Python SDK for the ElevenLabs API.
一个能让 Bot 在私聊和群聊中发起主动消息的插件,拥有上下文感知、持久化数据、动态情绪、免打扰时段和 TTS 集成。还有独立 WebUI,可进行个性化配置。 An AstrBot plugin that enables Bot to send proactive messages in private and group chats, featuring context awareness, persistent data, dynamic emotions, do-not-disturb periods, and TTS integration. It also boasts an independent WebUI for personalized.
Real-time text-to-speech with Qwen3-TTS
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet, AudioGen, MusicGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, and Bark!
A 100M-parameter multilingual TTS model for real-time CPU inference, voice cloning, and 48 kHz stereo generation
🧠 Leon is your open-source personal assistant.
macOS dictation app
The Open Source Alternative to Cluely - A lightning-fast, privacy-first AI assistant that works seamlessly during meetings, interviews, and conversations without anyone knowing. Built with Tauri for native performance, just 10MB. Completely undetectable in video calls, screen shares, and recordings.
Pocket-sized AI chatbot built using a RPI Zero 2w / 5
🤗 Optimum Intel: Accelerate inference with Intel optimization tools
Vui Nano — a small, context-aware text-to-speech model trained on real conversations. 219M active params (305M total), Apache 2.0, voice cloning, streaming, runs on CPU (dependency-free C build). Ships with a full real-time voice assistant: WebRTC, ASR, local LLM, OpenAI Realtime API compatible.
No description
Управление устройствами из Home Assistant через Алису (Умный дом Яндекса) или Марусю
Public release of the Sound Effect Foundation model by Sony AI.
TorchSig is an open-source signal processing machine learning toolkit based on the PyTorch data handling pipeline.
24,535 repositories in the index in total.