You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Kazakh / Russian / English real-time ASR

This package transcribes Kazakh, Russian, and English speech, produces partial text while audio arrives, and adds casing, punctuation, and conservative number formatting to final text. It includes word timestamps, contextual vocabulary support, voice activity detection, a file transcription script, and a WebSocket server. Optional streaming diarization adds anonymous speaker labels for up to eight speakers.

The acoustic model is a directly adapted GigaAM Multilingual Conformer CTC model with 220.7M parameters. The default pipeline has 398.5M parameters, including a separate multilingual BERT formatting model and Silero VAD; optional diarization raises the total to 497.7M. A small trigram language model improves decoding where validation showed a benefit. No knowledge distillation was performed in this experiment.

The ASR encoder uses full attention over a bounded audio window. The supplied runtime recomputes this window as PCM arrives; it is buffered inference, rather than an encoder with a native streaming cache. Partial hypotheses can change. Words are committed when the window rolls over, and an utterance is finalized after a speech endpoint or an explicit end command. The audio window is bounded to approximately 12 seconds.

Download and install

This repository is public and gated with manual access approval. Request access on the model page and wait for approval before downloading. After approval, authenticate with your own Hugging Face account:

python -m pip install huggingface_hub
hf auth login
hf download nur-dev/asr-220m-kk-ru-en --local-dir asr-220m-kk-ru-en
cd asr-220m-kk-ru-en

Use Python 3.12 and an NVIDIA GPU with a compatible CUDA driver. CPU inference is supported by --device cpu, but the real-time measurements below are for an L40. Run the following from the package directory, which contains models/, these scripts, and realtime_config.json:

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements-runtime.txt

For optional diarization, install requirements-diarization.txt instead. That profile pins a Transformers source revision with Nemotron-3-Diarization support; the ordinary profile uses Transformers 4.57.3. A C++ compiler and CMake may be needed to build KenLM. Both dependency profiles also have resolved .lock files. Inference reads the included model files locally and needs no Hugging Face token or further model download.

Transcribe an audio file

python transcribe.py samples/kk_clean.wav --output transcript.json
python transcribe.py samples/ru_clean.wav --pace --output russian.json
python transcribe.py samples/en_clean.wav --number-format words --output english.json

The script accepts formats supported by SoundFile, averages stereo channels, and resamples to 16 kHz. It prints partial and final events. The saved JSON contains text, raw_text, words, and all events. raw_text preserves the recognized spoken words before casing, punctuation, and number formatting. Word timestamps are approximate CTC timings, rather than forced alignment.

--pace replays audio at wall-clock speed. --no-formatting disables casing and punctuation; --number-format words independently disables digit conversion. --acoustic-only disables the trigram LM. To supply names or domain terminology, create a UTF-8 file with one term per line and pass --hotwords terms.txt. Decoding biases those terms; the formatter preserves the requested casing of exact single-word terms, such as OpenAI or iPhone. Context terms cannot guarantee correct recognition of every name.

Configure real-time behavior

Pass --config your_config.json to the file script, or set ASR_CONFIG for the server. The supplied defaults are:

Setting Default Effect
update_seconds 0.64 Interval between ASR updates during speech
max_window_seconds 12.0 Maximum PCM window processed by ASR
overlap_seconds 2.0 Retained context at window rollover
commit_lag_seconds 1.28 Stability margin for an alternative incremental commit policy
commit_policy on_rollover Keep partial text revisable until rollover or finalization
endpoint_silence_seconds 0.64 Silence needed to finalize an utterance
use_vad true Use speech detection to suppress non-speech decoding
vad_threshold 0.5 Speech detection threshold
number_format digits Conservative cardinal-number conversion; words preserves spoken forms

Smaller update intervals increase compute and can change WER. Measurements below apply to this exact default configuration. Number formatting covers common cardinal expressions in the three languages, including many years; it does not fully normalize currencies, dates, addresses, or all grammatical numeral forms.

WebSocket API

python -m uvicorn server:app --host 127.0.0.1 --port 8765
python stream_client.py samples/kk_clean.wav --pace --output live.json

Connect to ws://127.0.0.1:8765/v1/transcribe. After the ready event, send 16 kHz mono PCM16 little-endian binary frames; 512 samples / 32 ms per frame is convenient. Each frame must have an even byte count and contain at most one second of audio. WAV headers must not be sent. Receive partial and final JSON events concurrently. Send {"type":"end"} to flush the remaining speech; the server returns final events and done before closing. GET /health reports readiness and active connections.

Environment variables include ASR_DEVICE (default cuda:0), ASR_CONFIG, ASR_HOTWORDS, ASR_MAX_CLIENTS (default 4), and ASR_DIARIZATION (default 0). ASR_CHECKPOINT, ASR_FORMATTER, and ASR_LM can override individual component paths. Multiple clients share serialized model inference; throughput and latency at the configured client limit have not been load-tested.

With the diarization dependency profile installed:

ASR_DIARIZATION=1 python -m uvicorn server:app --host 127.0.0.1 --port 8765
python stream_client.py samples/ru_clean.wav --pace --output speakers.json

The server additionally emits speaker_activity events and assigns speaker_0, speaker_1, etc. to final words. Labels follow speaker arrival order within a connection; a word can have speaker: null when no speaker activity overlaps it. Labels identify anonymous speakers, not people. For file diarization and word tagging:

python diarization.py meeting.wav --words transcript.json --output speakers.json

Measured quality and speed

The final package was replayed as arriving 32 ms PCM frames on 3,220 recordings / 9.61 hours: 500 official FLEURS test recordings per language, 120 per language with held-out MUSAN noise at each of 20 and 10 dB SNR, and 1,000 official KSC2 test recordings. The previous model is the published 120M cache-aware TDT model at its 560 ms profile, evaluated on the same audio and references. This is a comparison with that streaming model, not every model in the Hugging Face collection.

Clean FLEURS lexical WER (%):

Language Previous streaming model New, spoken words New, number-formatted output
Kazakh 12.19 11.46 9.81
Russian 14.24 9.61 8.05
English 17.43 13.26 12.05

At 10 dB MUSAN noise, number-formatted WER is 10.99% / 9.27% / 14.14% for Kazakh / Russian / English, compared with 17.18% / 21.74% / 34.42% for the previous model. The noise subsets contain 120 recordings per language; their average scores should not be directly compared with the 500-recording clean averages. KSC2 uses references with spoken numbers and no casing or punctuation: the appropriate raw WER is 10.45%, versus 13.77% previously.

Lexical scoring lowercases and removes punctuation while retaining numerals and spelling errors. Presented output additionally converts selected spoken numbers to digits, which matches many FLEURS references. Casing and punctuation are assessed separately: clean FLEURS written WER is 24.83% / 17.29% / 20.06%, and written CER is 5.84% / 4.63% / 7.44%, respectively. These errors include lexical errors as well as case and punctuation differences. Proper names, punctuation choices, and abbreviations still require application-specific evaluation.

With one stream per L40, median processing RTF is 0.034 (about 29 times faster than incoming audio), and update compute p95 is 21.94 ms. These numbers exclude diarization and network delay. The configured 640 ms update interval and 640 ms endpoint silence are separate buffering delays; 21.94 ms is not end-to-end transcription latency. Median first non-empty partial occurs at 1.568 seconds from recording start, including leading silence. Eight independent replicas were used to complete the test; aggregate eight-GPU throughput is not presented as single-stream speed.

Optional Nemotron-3-Diarization achieved 22.10% DER over AMI test meetings ES2004a/b/c, with zero collar and overlap included, at its 640 ms input-buffer configuration. Most errors were missed speech. This is a small English meeting evaluation; Kazakh/Russian conversational diarization and production accuracy remain unverified. Diarization is disabled by default.

See reports/realtime.json, reports/comparison.json, and reports/diarization_ami.json for full measurements. Example recordings and their actual transcripts: Kazakh audio / text, Russian audio / text, English audio / text.

Runtime checks cover silence, noise-only input, a continuous 60-second stream with bounded PCM and ordered timestamps, 490-word formatting without word loss, and live WebSocket transcription with speaker labels and malformed-PCM rejection. samples/index.json contains three fixed sample references and actual final transcripts.

Component revisions and weight hashes are recorded in provenance/ and release_manifest.json. See NOTICE.md and licenses/ for component licenses and data attribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nur-dev/asr-220m-kk-ru-en

Finetuned
(4)
this model