Self-hosted realtime and offline speech recognition. AsrServe always ships state-of-the-art open models on the most efficient inference path available, and treats recognition quality as a first-class citizen.
- Web Demo: https://asr.ieeio.com
qwenasr_client_demo.mp4
Watch or download the original video
- Email: [email protected]
- WeChat:
v1.0.4 replaces the whole model stack on both the backend and the browser client. In our own tests against v1.0.3, recognition accuracy improved by about 20% and end-to-end efficiency by about 40%.
Breaking changes:
- Project rename:
qwen3-asris now AsrServe. The GitHub repository, Python package and Docker images (quantatrisk/asrserve) use the new name; old GitHub URLs redirect automatically. - ASR model: Qwen3-ASR 1.7B/0.6B is replaced by Confucius4-R2T2. Offline and realtime share one R2T2 engine (vLLM on CUDA, vendored Rust on CPU); the
modelrequest field no longer switches models. - Speaker diarization: CAM++ is replaced by Nemotron-3-Diarization (up to 8 speakers). Realtime streams now also carry per-utterance speaker labels.
- Removed models: FSMN VAD, the three CAM++ models and the Qwen3-ASR checkpoints are gone. Nemotron speech activity drives offline segmentation. ModelScope and FunASR are no longer dependencies; all models come from Hugging Face at pinned revisions.
- Removed API: the Alibaba Cloud compatible REST API (
/stream/v1/asr*) is removed. Use the OpenAI-compatible/v1/audio/transcriptionsand the native/v1/streamWebSocket. - Deployment: images are published on Docker Hub as
quantatrisk/asrserve:gpu(amd64),:cpu(amd64/arm64) and:ascend(arm64, Ascend 910B branch), plus1.0.4-*version tags;build.shbuilds from source. Compose files arecompose.yml(GPU) andcompose.cpu.yml(CPU) and no longer build. The only mount is./models.docker-compose*.yml,deploy/prepare.shand the model export option are removed. The oldquantatrisk/qwen3-asrDocker Hub repository is no longer updated. - Runtime: the CUDA image moves to CUDA 13.0 and vLLM 0.30; the service listens on port
17003.
Older release notes: GitHub Releases.
- SOTA models, quality first: Confucius4-R2T2 recognition, Nemotron speaker diarization and Qwen3-ForcedAligner word timestamps today; the stack moves whenever a better open model appears
- First-class macOS support: runs natively on Apple Silicon through a bundled Rust inference backend, no Docker or GPU required
- Runs anywhere else too: NVIDIA GPU (vLLM) and Linux CPU (amd64/arm64)
- Realtime and offline in one service, sharing one set of R2T2 weights
- Speaker diarization for files and live streams via Nemotron
- Word timestamps via Qwen3-ForcedAligner
- OpenAI compatible
/v1/audio/transcriptions, works with the OpenAI SDK - Browser recording page at
/realtime
docker compose up -d # GPU: quantatrisk/asrserve:gpu
docker compose -f compose.cpu.yml up -d # CPU: quantatrisk/asrserve:cpu (amd64/arm64)Images are pulled from Docker Hub on first start; upgrade with docker compose pull && docker compose up -d. To build from source, run ./build.sh (CPU: TARGET=cpu ./build.sh). Ascend 910B lives on the ascend-910b branch.
On macOS, run natively instead (see Deployment):
uv sync --frozen && ./scripts/build-rust.sh
HF_HOME="$PWD/models/huggingface" uv run --no-sync python start.py # http://localhost:8000The first start downloads the pinned models into ./models. Default URL: http://localhost:17003 (recording page /realtime, API docs /docs, health /health).
.env is optional; copy .env.example and uncomment what you need (API key, offline mode, GPU memory, CPU threads).
For offline hosts, pre-download on a networked machine, copy the repository including models/, and set HF_HUB_OFFLINE=1:
./scripts/prepare-models.sh # uses the :gpu image, no GPU needed
IMAGE=quantatrisk/asrserve:cpu ./scripts/prepare-models.sh # or the :cpu imagecurl http://localhost:17003/v1/audio/transcriptions \
-F [email protected] \
-F model=confucius4-r2t2 \
-F response_format=verbose_json \
-F enable_speaker_diarization=true \
-F word_timestamps=trueWith API_KEY set, add Authorization: Bearer <API_KEY>. Speaker paragraphs follow the main speaker; raw overlapping activity is returned in speaker_segments.
Connect to ws://localhost:17003/v1/stream and send 16 kHz mono int16 PCM. See the realtime protocol.
- Deployment: configuration, offline setup, native CPU, limits
- Realtime protocol
- Acceptance scripts
- Confucius4-R2T2: speech recognition; source attribution in
deploy/R2T2-NOTICE, weights under their own MODEL_LICENSE - Qwen3-ASR: forced alignment
- Nemotron-3-Diarization: speaker diarization
- QwenASR: vendored CPU Rust backend
