Kazakh / Russian / English real-time ASR
This package transcribes Kazakh, Russian, and English speech, produces partial text while audio arrives, and adds casing, punctuation, and conservative number formatting to final text. It includes word timestamps, contextual vocabulary support, voice activity detection, a file transcription script, and a WebSocket server. Optional streaming diarization adds anonymous speaker labels for up to eight speakers.
The acoustic model is a directly adapted GigaAM Multilingual Conformer CTC model with 220.7M parameters. The default pipeline has 398.5M parameters, including a separate multilingual BERT formatting model and Silero VAD; optional diarization raises the total to 497.7M. A small trigram language model improves decoding where validation showed a benefit. No knowledge distillation was performed in this experiment.
The ASR encoder uses full attention over a bounded audio window. The supplied runtime recomputes this window as PCM arrives; it is buffered inference, rather than an encoder with a native streaming cache. Partial hypotheses can change. Words are committed when the window rolls over, and an utterance is finalized after a speech endpoint or an explicit end command. The audio window is bounded to approximately 12 seconds.
Download and install
This repository is public and gated with manual access approval. Request access on the model page and wait for approval before downloading. After approval, authenticate with your own Hugging Face account:
python -m pip install huggingface_hub
hf auth login
hf download nur-dev/asr-220m-kk-ru-en --local-dir asr-220m-kk-ru-en
cd asr-220m-kk-ru-en
Use Python 3.12 and an NVIDIA GPU with a compatible CUDA driver. CPU inference is supported by --device cpu, but the real-time measurements below are for an L40. Run the following from the package directory, which contains models/, these scripts, and realtime_config.json:
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements-runtime.txt
For optional diarization, install requirements-diarization.txt instead. That profile pins a Transformers source revision with Nemotron-3-Diarization support; the ordinary profile uses Transformers 4.57.3. A C++ compiler and CMake may be needed to build KenLM. Both dependency profiles also have resolved .lock files. Inference reads the included model files locally and needs no Hugging Face token or further model download.
Transcribe an audio file
python transcribe.py samples/kk_clean.wav --output transcript.json
python transcribe.py samples/ru_clean.wav --pace --output russian.json
python transcribe.py samples/en_clean.wav --number-format words --output english.json
The script accepts formats supported by SoundFile, averages stereo channels, and resamples to 16 kHz. It prints partial and final events. The saved JSON contains text, raw_text, words, and all events. raw_text preserves the recognized spoken words before casing, punctuation, and number formatting. Word timestamps are approximate CTC timings, rather than forced alignment.
--pace replays audio at wall-clock speed. --no-formatting disables casing and punctuation; --number-format words independently disables digit conversion. --acoustic-only disables the trigram LM. To supply names or domain terminology, create a UTF-8 file with one term per line and pass --hotwords terms.txt. Decoding biases those terms; the formatter preserves the requested casing of exact single-word terms, such as OpenAI or iPhone. Context terms cannot guarantee correct recognition of every name.
Configure real-time behavior
Pass --config your_config.json to the file script, or set ASR_CONFIG for the server. The supplied defaults are:
| Setting | Default | Effect |
|---|---|---|
update_seconds |
0.64 | Interval between ASR updates during speech |
max_window_seconds |
12.0 | Maximum PCM window processed by ASR |
overlap_seconds |
2.0 | Retained context at window rollover |
commit_lag_seconds |
1.28 | Stability margin for an alternative incremental commit policy |
commit_policy |
on_rollover |
Keep partial text revisable until rollover or finalization |
endpoint_silence_seconds |
0.64 | Silence needed to finalize an utterance |
use_vad |
true |
Use speech detection to suppress non-speech decoding |
vad_threshold |
0.5 | Speech detection threshold |
number_format |
digits |
Conservative cardinal-number conversion; words preserves spoken forms |
Smaller update intervals increase compute and can change WER. Measurements below apply to this exact default configuration. Number formatting covers common cardinal expressions in the three languages, including many years; it does not fully normalize currencies, dates, addresses, or all grammatical numeral forms.
WebSocket API
python -m uvicorn server:app --host 127.0.0.1 --port 8765
python stream_client.py samples/kk_clean.wav --pace --output live.json
Connect to ws://127.0.0.1:8765/v1/transcribe. After the ready event, send 16 kHz mono PCM16 little-endian binary frames; 512 samples / 32 ms per frame is convenient. Each frame must have an even byte count and contain at most one second of audio. WAV headers must not be sent. Receive partial and final JSON events concurrently. Send {"type":"end"} to flush the remaining speech; the server returns final events and done before closing. GET /health reports readiness and active connections.
Environment variables include ASR_DEVICE (default cuda:0), ASR_CONFIG, ASR_HOTWORDS, ASR_MAX_CLIENTS (default 4), and ASR_DIARIZATION (default 0). ASR_CHECKPOINT, ASR_FORMATTER, and ASR_LM can override individual component paths. Multiple clients share serialized model inference; throughput and latency at the configured client limit have not been load-tested.
With the diarization dependency profile installed:
ASR_DIARIZATION=1 python -m uvicorn server:app --host 127.0.0.1 --port 8765
python stream_client.py samples/ru_clean.wav --pace --output speakers.json
The server additionally emits speaker_activity events and assigns speaker_0, speaker_1, etc. to final words. Labels follow speaker arrival order within a connection; a word can have speaker: null when no speaker activity overlaps it. Labels identify anonymous speakers, not people. For file diarization and word tagging:
python diarization.py meeting.wav --words transcript.json --output speakers.json
Measured quality and speed
The final package was replayed as arriving 32 ms PCM frames on 3,220 recordings / 9.61 hours: 500 official FLEURS test recordings per language, 120 per language with held-out MUSAN noise at each of 20 and 10 dB SNR, and 1,000 official KSC2 test recordings. The previous model is the published 120M cache-aware TDT model at its 560 ms profile, evaluated on the same audio and references. This is a comparison with that streaming model, not every model in the Hugging Face collection.
Clean FLEURS lexical WER (%):
| Language | Previous streaming model | New, spoken words | New, number-formatted output |
|---|---|---|---|
| Kazakh | 12.19 | 11.46 | 9.81 |
| Russian | 14.24 | 9.61 | 8.05 |
| English | 17.43 | 13.26 | 12.05 |
At 10 dB MUSAN noise, number-formatted WER is 10.99% / 9.27% / 14.14% for Kazakh / Russian / English, compared with 17.18% / 21.74% / 34.42% for the previous model. The noise subsets contain 120 recordings per language; their average scores should not be directly compared with the 500-recording clean averages. KSC2 uses references with spoken numbers and no casing or punctuation: the appropriate raw WER is 10.45%, versus 13.77% previously.
Lexical scoring lowercases and removes punctuation while retaining numerals and spelling errors. Presented output additionally converts selected spoken numbers to digits, which matches many FLEURS references. Casing and punctuation are assessed separately: clean FLEURS written WER is 24.83% / 17.29% / 20.06%, and written CER is 5.84% / 4.63% / 7.44%, respectively. These errors include lexical errors as well as case and punctuation differences. Proper names, punctuation choices, and abbreviations still require application-specific evaluation.
With one stream per L40, median processing RTF is 0.034 (about 29 times faster than incoming audio), and update compute p95 is 21.94 ms. These numbers exclude diarization and network delay. The configured 640 ms update interval and 640 ms endpoint silence are separate buffering delays; 21.94 ms is not end-to-end transcription latency. Median first non-empty partial occurs at 1.568 seconds from recording start, including leading silence. Eight independent replicas were used to complete the test; aggregate eight-GPU throughput is not presented as single-stream speed.
Optional Nemotron-3-Diarization achieved 22.10% DER over AMI test meetings ES2004a/b/c, with zero collar and overlap included, at its 640 ms input-buffer configuration. Most errors were missed speech. This is a small English meeting evaluation; Kazakh/Russian conversational diarization and production accuracy remain unverified. Diarization is disabled by default.
See reports/realtime.json, reports/comparison.json, and reports/diarization_ami.json for full measurements. Example recordings and their actual transcripts: Kazakh audio / text, Russian audio / text, English audio / text.
Runtime checks cover silence, noise-only input, a continuous 60-second stream with bounded PCM and ordered timestamps, 490-word formatting without word loss, and live WebSocket transcription with speaker labels and malformed-PCM rejection. samples/index.json contains three fixed sample references and actual final transcripts.
Component revisions and weight hashes are recorded in provenance/ and release_manifest.json. See NOTICE.md and licenses/ for component licenses and data attribution.
Model tree for nur-dev/asr-220m-kk-ru-en
Base model
ai-sage/GigaAM-Multilingual