Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech-to-speech AI assistant with natural conversation flow, mid-speech interruption, vision capabilities and AI-initiated follow-ups. Features low-latency audio streaming, dynamic visual feedback, and works with local LLM/TTS services via OpenAI-compatible endpoints.
| Date | Stars |
|---|---|
| 2026-07-31 | 310 |
| 2026-08-04 | 310 |
| 2026-08-13 | 312 |
| 2026-08-25 | 313 |
| 2026-08-27 | 313 |
| 2026-08-29 | 314 |
| 2026-08-30 | 315 |
| 2026-09-01 | 315 |
| 2026-09-06 | 314 |
| 2026-09-10 | 315 |
| 2026-09-12 | 316 |
| 2026-09-20 | 316 |
Today
— stars today
This week
— stars this week
This month
+4 stars this month
Momentum
0.0
growth rate 0.00%/day
 # Vocalis [](https://opensource.org/licenses/Apache-2.0) [](https://reactjs.org/) [](https://fastapi.tiangolo.com/) [](https://github.com/guillaumekln/faster-whisper) [](https://www.python.org/) A sophisticated AI assistant with speech-to-speech capabilities built on a modern React frontend with a FastAPI backend. Vocalis provides a responsive, low-latency conversational experience with advanced visual feedback. ## Video Demonstration of Setup and Usage [](https://www.youtube.com/watch?v=2slWwsHTNIA) ## Changelog **v1.5.0** (Vision Update) - April 12, 2025 - 🔍 New image analysis capability powered by [SmolVLM-256M-Instruct model](https://huggingface.co/HuggingFaceTB/SmolVLM-256M-Instruct) - 🖼️ Seamless image upload and processing interface - 🔄 Contextual conversation continuation based on image understanding - 🧩 Multi-modal conversation support (text, speech, and images) - 💾 Advanced session management for saving and retrieving conversations - 🎨 Improved UI with central call button and cleaner control layout - 🔌 Simplified sidebar without redundant controls **v1.0.0** (Initial Release) - March 31, 2025 - ✨ Revolutionary barge-in technology for natural conversation flow - 🔊 Ultra low-latency audio streaming with adaptive buffering - 🤖 AI-initiated greetings and follow-ups for natural conversations - 🎨 Dynamic visual feedback system with state-aware animations - 🔄 Streaming TTS with chunk-based delivery for immediate responses - 🚀 Cross-platform support with optimised setup scripts - 💻 CUDA acceleration with fallback for CPU-only systems ## Features ### 🎯 Advanced Conversation Capabilities - **🗣️ Barge-In Interruption** - Interrupt the AI mid-speech for a truly natural conversation experience - **👋 AI-Initiated Greetings** - Assistant automatically welcomes users with a contextual greeting - **💬 Intelligent Follow-Ups** - System detects silence and continues conversation with natural follow-up questions - **🔄 Conversation Memory** - Maintains context throughout the conversation session - **🧠 Contextual Understanding** - Processes conversation history for coherent, relevant responses - **🖼️ Image Analysis** - Upload and discuss images with integrated visual understanding - **💾 Session Management** - Save, load, and manage conversation sessions with customisable titles ### ⚡ Ultra-Responsive Performance - **⏱️ Low-Latency Processing** - End-to-end latency under 500ms for immediate response perception - **🔊 Streaming Audio** - Begin playback before full response is generated - **📦 Adaptive Buffering** - Dynamically adjust audio buffer size based on network conditions - **🔌 Efficient WebSocket Protocol** - Bidirectional real-time audio streaming - **🔄 Parallel Processing** - Multi-stage pipeline for concurrent audio handling ### 🎨 Interactive Visual Experience - **🔮 Dynamic Assistant Orb** - Visual representation with state-aware animations: - Pulsing glow during listening - Particle animations during processing - Wave-like motion during speaking - **📝 Live Transcription** - Real-time display of recognised speech - **🚦 Status Indicators** - Clear visual cues for system state - **🌈 Smooth Transitions** - Fluid state changes with appealing animations - **🌙 Dark Theme** - Eye-friendly interface with cosmic aesthetic ### 🛠️ Technical Excellence - **🔍 High-Accuracy VAD** - Superior voice activi
Excerpt of 32,561 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:95296c99d322510f, llm:Repository topics and description: 'artificial-intelligence, conversational-ai, speech-to-speech, visionprocessing' and description: 'Speech-to-speech AI assistant with natural conversation flow, mid-speech interruption, vision capabilities and AI-initiated follow-ups. Features low-latency audio streaming, dynamic visual feedback, and works with local LLM/TTS services via OpenAI-compatible endpoints.'
matched fp:95296c99d322510f, llm:Repository topics and description: 'artificial-intelligence, conversational-ai, speech-to-speech, visionprocessing' and description: 'Speech-to-speech AI assistant with natural conversation flow, mid-speech interruption, vision capabilities and AI-initiated follow-ups. Features low-latency audio streaming, dynamic visual feedback, and works with local LLM/TTS services via OpenAI-compatible endpoints.'
matched fp:95296c99d322510f, llm:Repository topics and description: 'artificial-intelligence, conversational-ai, speech-to-speech, visionprocessing' and description: 'Speech-to-speech AI assistant with natural conversation flow, mid-speech interruption, vision capabilities and AI-initiated follow-ups. Features low-latency audio streaming, dynamic visual feedback, and works with local LLM/TTS services via OpenAI-compatible endpoints.'
matched fp:95296c99d322510f, llm:Repository topics and description: 'artificial-intelligence, conversational-ai, speech-to-speech, visionprocessing' and description: 'Speech-to-speech AI assistant with natural conversation flow, mid-speech interruption, vision capabilities and AI-initiated follow-ups. Features low-latency audio streaming, dynamic visual feedback, and works with local LLM/TTS services via OpenAI-compatible endpoints.'