Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
OpenAI Whisper ASR Webservice API
| Date | Stars |
|---|---|
| 2026-07-24 | 3303 |
| 2026-07-25 | 3303 |
| 2026-07-28 | 3303 |
| 2026-07-30 | 3303 |
| 2026-08-06 | 3303 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
    > 🚀 **Try Speech Box Desktop App | Offline, multi-language desktop transcriptions**: [https://speechbox.gumroad.com/l/desktop-app](https://speechbox.gumroad.com/l/desktop-app) # Whisper ASR Box Whisper ASR Box is a general-purpose speech recognition toolkit. Whisper Models are trained on a large dataset of diverse audio and is also a multitask model that can perform multilingual speech recognition as well as speech translation and language identification. 🎉 **Join our Discord Community!** Connect with other users, get help, and stay updated on the latest features: [https://discord.gg/4Q5YVrePzZ](https://discord.gg/4Q5YVrePzZ) ## Features Current release (v1.9.1) supports following whisper models: - [openai/whisper](https://github.com/openai/whisper)@[v20250625](https://github.com/openai/whisper/releases/tag/v20250625) - [SYSTRAN/faster-whisper](https://github.com/SYSTRAN/faster-whisper)@[v1.1.1](https://github.com/SYSTRAN/faster-whisper/releases/tag/v1.1.1) - [whisperX](https://github.com/m-bain/whisperX)@[v3.4.2](https://github.com/m-bain/whisperX/releases/tag/v3.4.2) ## Quick Usage ### CPU ```shell docker run -d -p 9000:9000 \ -e ASR_MODEL=base \ -e ASR_ENGINE=openai_whisper \ onerahmet/openai-whisper-asr-webservice:latest ``` ### GPU ```shell docker run -d --gpus all -p 9000:9000 \ -e ASR_MODEL=base \ -e ASR_ENGINE=openai_whisper \ onerahmet/openai-whisper-asr-webservice:latest-gpu ``` #### Cache To reduce container startup time by avoiding repeated downloads, you can persist the cache directory: ```shell docker run -d -p 9000:9000 \ -v $PWD/cache:/root/.cache/ \ onerahmet/openai-whisper-asr-webservice:latest ``` ## Key Features - Multiple ASR engines support (OpenAI Whisper, Faster Whisper, WhisperX) - Multiple output formats (text, JSON, VTT, SRT, TSV) - Word-level timestamps support - Voice activity detection (VAD) filtering - Speaker diarization (with WhisperX) - FFmpeg integration for broad audio/video format support - GPU acceleration support - Configurable model loading/unloading - REST API with Swagger documentation ## Environment Variables Key configuration options: - `ASR_ENGINE`: Engine selection (openai_whisper, faster_whisper, whisperx) - `ASR_MODEL`: Model selection (tiny, base, small, medium, large-v3, etc.) - `ASR_MODEL_PATH`: Custom path to store/load models - `ASR_DEVICE`: Device selection (cuda, cpu) - `MODEL_IDLE_TIMEOUT`: Timeout for model unloading ## Documentation For complete documentation, visit: [https://ahmetoner.github.io/whisper-asr-webservice](https://ahmetoner.github.io/whisper-asr-webservice) ## Development ```shell # Install poetry v2.X pip3 install poetry # Install dependencies for cpu poetry install --extras cpu # Install dependencies for cuda poetry install --extras cuda # Run service poetry run whisper-asr-webservice --host 0.0.0.0 --port 9000 ``` After starting the service, visit `http://localhost:9000` or `http://0.0.0.0:9000` in your browser to access the Swagger UI documentation and try out the API endpoints. ## Credits - This software uses libraries from the [FFmpeg](http://ffmpeg.org) project under the [LGPLv2.1](http://www.gnu.org/licenses/old-licenses/lgpl-2.1.html)
Excerpt of 3,612 characters
Read on GitHub266
9
7
3
Dr Nic Williams · @mocra · Australia
2
Alex Yancey · United States
2
1
1
Pavel Zloi
1
Samuel Dalesjö · Sweden
1
1
1
1
Forest Anderson · Canada
1
Marc
1
1
Jesse Vincent · Prime Radiant
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:2a4c3b04c0507c78, topic:speech-recognition, topic:asr, topic:speech-to-text