SergeyShk/Speech-to-Text-Russian
quality grade F, 29 out of 100Проект для распознавания речи на русском языке на основе pykaldi.
- stars
- 346
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
913 results
Проект для распознавания речи на русском языке на основе pykaldi.
Python Kaldi speech recognition with grammars that can be set active/inactive dynamically at decode-time
A desktop application that transcribes audio from files, microphone input or YouTube videos with the option to translate the content and create subtitles.
Android web novel reader
The SpeechBrain project aims to build a novel speech toolkit fully based on PyTorch. With SpeechBrain users can easily create speech processing systems, ranging from speech recognition (both HMM/DNN and end-to-end), speaker recognition, speech enhancement, speech separation, multi-microphone speech processing, and many others.
A live speech recognition using Facebooks wav2vec 2.0 model.
speech to text with self-supervised learning based on wav2vec 2.0 framework
Synthalingua - Real Time Translation
Open-source AI meeting copilot - real-time transcription, echo cancellation, and AI assistance. Captures system audio + mic, cancels echo via WebRTC AEC3, transcribes with Deepgram, and gives you Claude/OpenAI help during meetings. Runs locally on macOS and Windows.
VRCT(VRChat Chatbox Translator & Transcription)
Record audio from a user's microphone and display a cool visualization.
⚡ 一款用于自动语音识别 (ASR)、翻译的高性能异步 API。不需要购买Whisper API,使用本地运行的Whisper模型进行推理,并支持多GPU并发,针对分布式部署进行设计。还内置了包括TikTok、抖音等社交媒体平台的爬虫,可实现来自多个社交平台的无缝媒体处理,为媒体内容数据自动化处理提供了强大且可扩展的解决方案。
HuggingSound: A toolkit for speech-related tasks based on Hugging Face's tools
VOSK Speech Recognition Toolkit
open source audio and video transcription software
This tool uses AI to evaluate your pronunciation.
Phonetisaurus G2P
This is a list of features, scripts, blogs and resources for better using Kaldi ( http://kaldi-asr.org/ )
The J.A.R.V.I.S. Speech API is designed to be simple and efficient, using the speech engines created by Google to provide functionality for parts of the API. Essentially, it is an API written in Java, including a recognizer, synthesizer, and a microphone capture utility. The project uses Google services for the synthesizer and recognizer. While this requires an Internet connection, it provides a complete, modern, and fully functional speech API in Java.
Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.
A Python library for solving reCAPTCHA v2 and v3 with Playwright
An open-source on-device voice IME (keyboard) for Android using the Vosk library.
speech to text benchmark framework
24,524 repositories in the index in total.