Automatic Speech Recognition
Transformers
Safetensors
whisper
african-languages
speech
grpo

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

SunflowerASR: speech recognition for 51 African languages

SunflowerASR is an adaptation of Whisper large-v3 that supports Automatic Speech Recognition (ASR) for 51 languages widely spoken in Africa: English and French with African accents, Swahili, Afrikaans, Tswana, Kinyarwanda, Nigerian Pidgin, Luganda, Acholi, Lugbara, Ateso, Runyankole, Rutooro, Lumasaba, Lusoga, Rukiga, Akan, Amharic, Bambara, Bemba, Berber, Chichewa, Dagaare, Dagbani, Ewe, Fulani, Hausa, Igbo, Ikposo, Kabyle, Kalenjin, Kanuri, Kikuyu, Kwamba, Lendu, Lingala, Luhya, Luo, Malagasy, Ndebele, Oromo, Rukonjo, Ruruuli, Shona, Somali, Sotho, Thur, Wolof, Xhosa, Yoruba, and Zulu.

It was trained by Sunbird AI in two stages: full-parameter supervised fine-tuning on 7,412 hours of African speech, followed by a reinforcement learning pass with Group Relative Policy Optimisation (GRPO) that optimises character error rate directly. It is released as part of Sunflower V2, and the method is described in the accompanying paper (to be released later).

Across three benchmarks it is the most accurate published system we have measured: it has the lowest macro-averaged Word Error Rate (WER) of any model, open or closed, on the Sunbird Speech Benchmark (51 languages), AfriVox-v2 and SimbaBench, including against Meta's Omnilingual ASR, Gemini, GPT-4o-transcribe, Sahara v2 and Simba. Per-language rankings vary, and quality across the 51 languages is very uneven — check the per-language number for your language before relying on it.

Intended use

The model is intended for transcribing speech in the 51 languages listed above, including for applications such as media monitoring, call-centre and citizen-feedback analysis, agricultural and health advisory services, voice interfaces, and as a component in speech translation or speech-to-speech pipelines. It is also intended as a starting point for further fine-tuning on a specific language, domain or accent.

The model transcribes one language per clip: the language token is supplied at decode time and is not detected automatically. It does not translate, does not produce timestamps or speaker labels in this configuration, and is not a speaker-identification, language-identification or voice-biometric system. It should not be used as the sole basis for consequential decisions about individuals (for example in legal, clinical, immigration or employment settings) without human review of the transcript, particularly for the languages in the weaker half of the results tables.

Usage

You can use this Colab Notebook to try out the model. It lets you transcribe an audio sample from a Hugging Face dataset, an uploaded audio file, or your own recording from your computer's microphone.

The model is used in much the same way as the base Whisper model. For better accuracy, specify the language during generation.

See the example below, which transcribes an audio sample from a Hugging Face dataset:

import transformers
import datasets
import torch

SAMPLE_RATE = 16000

LANGUAGE_TOKENS_WHISPER = {
    # Existing languges codes from Whisper
    "eng": 50259, "fra": 50265, "swa": 50318, "sna": 50324, "yor": 50325, "som": 50326,
    "afr": 50327, "amh": 50334, "mlg": 50349, "lin": 50353, "hau": 50354,
    # Overwrite unused language tokens
    "ach": 50357, "aka": 50356, "bam": 50355, "bem": 50352, "ber": 50351,
    "cgg": 50350, "dag": 50348, "dga": 50347, "ewe": 50346, "ful": 50345,
    "ibo": 50344, "kab": 50343, "kau": 50342, "kik": 50341, "kin": 50340,
    "kln": 50339, "koo": 50338, "kpo": 50337, "led": 50336, "lgg": 50335,
    "lth": 50333, "lug": 50332, "luo": 50331, "luy": 50330, "myx": 50329,
    "nbl": 50328, "nya": 50323, "nyn": 50322, "orm": 50321, "pcm": 50320,
    "ruc": 50319, "rwm": 50317, "sot": 50316, "teo": 50315, "tsn": 50314,
    "ttj": 50313, "wol": 50312, "xho": 50311, "xog": 50310, "zul": 50309
}

LANGUAGE_NAMES = {
    'Acholi': 'ach', 'Afrikaans': 'afr', 'Akan': 'aka', 'Amharic': 'amh', 'Ateso': 'teo',
    'Bambara': 'bam', 'Bemba': 'bem', 'Berber': 'ber', 'Chichewa': 'nya', 'Dagaare': 'dga',
    'Dagbani': 'dag', 'English': 'eng', 'Ewe': 'ewe', 'French': 'fra', 'Fulani': 'ful',
    'Hausa': 'hau', 'Igbo': 'ibo', 'Ikposo': 'kpo', 'Kabyle': 'kab', 'Kalenjin': 'kln',
    'Kanuri': 'kau', 'Kikuyu': 'kik', 'Kinyarwanda': 'kin', 'Kwamba': 'rwm', 'Lendu': 'led',
    'Lingala': 'lin', 'Luganda': 'lug', 'Lugbara': 'lgg', 'Luhya': 'luy', 'Lumasaba': 'myx',
    'Luo': 'luo', 'Lusoga': 'xog', 'Malagasy': 'mlg', 'Ndebele': 'nbl', 'Nigerian Pidgin': 'pcm',
    'Oromo': 'orm', 'Rukiga': 'cgg', 'Rukonjo': 'koo', 'Runyankole': 'nyn', 'Ruruuli': 'ruc',
    'Rutooro': 'ttj', 'Shona': 'sna', 'Somali': 'som', 'Sotho': 'sot', 'Swahili': 'swa',
    'Thur': 'lth', 'Tswana': 'tsn', 'Wolof': 'wol', 'Xhosa': 'xho', 'Yoruba': 'yor', 'Zulu': 'zul'
}

model_id = "Sunbird/SunflowerASR-51-african-languages"
model = transformers.WhisperForConditionalGeneration.from_pretrained(model_id)
processor = transformers.WhisperProcessor.from_pretrained(model_id)

def transcribe_by_whisper(audio_array, language):
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    input_features = processor(
        audio_array, sampling_rate=SAMPLE_RATE, do_normalize=True, return_tensors="pt"
    ).input_features.to(device)

    lang_tok = LANGUAGE_TOKENS_WHISPER[LANGUAGE_NAMES[language]]
    transcribe_tok = processor.tokenizer.convert_tokens_to_ids("<|transcribe|>")
    notimestamps_tok = processor.tokenizer.convert_tokens_to_ids("<|notimestamps|>")
    forced_decoder_ids = [
        (1, lang_tok),
        (2, transcribe_tok),
        (3, notimestamps_tok),
    ]

    predicted_ids = model.to(device).generate(
        input_features,
        forced_decoder_ids=forced_decoder_ids,
        num_beams=1,
        do_sample=False,
    )
    transcription = processor.decode(
        predicted_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)

    print(transcription[0])

# Get some test audio
import huggingface_hub
huggingface_hub.login()

ds = datasets.load_dataset('Sunbird/salt', 'multispeaker-lug', split='test')
audio_column = 'audio'
ds = ds.cast_column(audio_column, datasets.Audio(sampling_rate=SAMPLE_RATE))
audio_array = ds[0][audio_column]['array']

# Specify a language
lang = 'Luganda'

transcribe_by_whisper(audio_array, lang)
# Ekikoola kya kasooli kya kyenvu wabula langi yakyo etera okuba eya kitaka wansi.

Two decoding notes. The language token must be forced from the table above rather than passed as a Whisper language name: the 40 languages Whisper did not originally cover are mapped onto unused token ids, so language="lug" alone will not select Luganda. And if you set max_new_tokens, keep it at 444 (Whisper's architectural ceiling of 448 decoder positions, minus the 4-token forced prefix) — Whisper's tokeniser has no Ge'ez coverage and falls back to bytes, so Amharic costs about 2.4 tokens per character against 0.4 for a median language here, and a shorter budget truncates its transcripts.

Performance metrics

We evaluate SunflowerASR on three benchmarks: our own Sunbird Speech Benchmark, which is the only one that covers the model's full language range, and the two independent benchmarks with published baselines, AfriVox-v2 and SimbaBench. WER and CER are computed with greedy decoding, the language token forced at the decoder, and Whisper's BasicTextNormalizer applied to both hypothesis and reference. Lower is better throughout.

Sunbird Speech Benchmark

The Sunbird Speech Benchmark is assembled from the held-out test splits of the 110 dataset subsets of our training corpus, capped at 200 clips each: 20,213 clips over 51 languages drawn from 28 source corpora. It is the only one of the three benchmarks that covers every language the model supports. Note that because it is drawn from the same collections as our training data, it is in-domain for SunflowerASR and out-of-domain for the baselines; read it as a measure of coverage across 51 languages, with AfriVox-v2 and SimbaBench below providing the arm's-length comparisons.

model_metrics_comparison_sunbird_benchmark_slides

Macro-averages over languages, comparing against Omnilingual ASR from Meta, Simba-S from UBC-NLP, Gemini 3.5 Flash from Google and GPT-4o-transcribe from OpenAI. Systems cover different subsets of the 51 languages, so the last column repeats SunflowerASR's average over exactly the languages each baseline scores:

System Languages scored WER CER SunflowerASR on the same languages (WER / CER)
SunflowerASR (ours) 51 0.293 0.091 —
OmniASR-LLM-7B 46 0.353 0.111 0.300 / 0.093
OmniASR-LLM-1B 46 0.383 0.119 0.300 / 0.093
Simba-S 18 0.567 0.190 0.263 / 0.078
Gemini 3.5 Flash 46 0.571 0.222 0.286 / 0.085
GPT-4o-transcribe 51 0.709 0.334 0.293 / 0.091

SunflowerASR has the lowest WER of any system compared on 30 of the 51 languages, and the lowest CER on 34.

Per-language results, sorted by WER, alongside the hours of supervised training audio that language contributed to the corpus:

Code Language Training hours Test clips WER CER
fra French 11.4 200 0.079 0.035
tsn Tswana 235.3 369 0.093 0.031
eng English 160.6 295 0.094 0.036
afr Afrikaans 51.6 729 0.096 0.039
nbl Ndebele 137.9 144 0.109 0.028
mlg Malagasy 93.5 199 0.137 0.052
sot Sotho 208.9 321 0.152 0.047
swa Swahili 82.6 400 0.153 0.060
hau Hausa 278.2 1130 0.171 0.047
luo Luo 198.4 596 0.175 0.038
lth Thur 6.5 167 0.187 0.073
sna Shona 86.2 594 0.194 0.037
pcm Nigerian Pidgin 693.2 397 0.197 0.087
zul Zulu 171.6 507 0.201 0.040
lgg Lugbara 10.4 96 0.203 0.054
orm Oromo 192.5 439 0.210 0.050
ber Berber 17.0 200 0.215 0.080
lin Lingala 183.1 767 0.215 0.087
kik Kikuyu 183.7 199 0.226 0.074
ttj Rutooro 7.0 199 0.228 0.051
ibo Igbo 465.2 860 0.244 0.081
xho Xhosa 111.1 527 0.252 0.049
led Lendu 7.1 199 0.275 0.093
teo Ateso 8.5 98 0.278 0.091
aka Akan 54.6 200 0.287 0.093
ewe Ewe 79.3 400 0.291 0.095
amh Amharic 374.2 1000 0.298 0.113
lug Luganda 180.7 950 0.303 0.077
kin Kinyarwanda 1,392.3 400 0.304 0.091
cgg Rukiga 6.5 200 0.309 0.069
dga Dagaare 80.6 200 0.310 0.129
dag Dagbani 78.4 400 0.312 0.106
nyn Runyankole 42.8 280 0.318 0.079
ach Acholi 25.9 270 0.322 0.130
nya Chichewa 49.3 399 0.324 0.076
luy Luhya 55.7 200 0.338 0.111
wol Wolof 40.0 571 0.345 0.128
kau Kanuri 31.6 200 0.348 0.089
som Somali 126.4 596 0.353 0.114
yor Yoruba 466.1 995 0.364 0.129
xog Lusoga 30.7 177 0.419 0.082
ful Fulani 111.7 594 0.426 0.124
bam Bambara 156.7 600 0.441 0.219
bem Bemba 22.1 400 0.441 0.096
myx Lumasaba 34.1 180 0.452 0.103
kab Kabyle 130.9 200 0.467 0.215
kln Kalenjin 145.7 400 0.495 0.109
ruc Ruruuli 5.2 170 0.518 0.112
rwm Kwamba 5.8 199 0.543 0.210
koo Rukonjo 6.5 200 0.567 0.139
kpo Ikposo 77.1 200 0.670 0.223
All 51 (macro-average) 7,412 20,213 0.293 0.091

SunflowerASR (ours)_wer_vs_train_hours

AfriVox-v2

AfriVox-v2 is an in-the-wild benchmark of 251 hours aggregating the Africa Next Voices, Waxal and AfriVox v1 corpora over a 10-domain taxonomy spanning health, finance, government and agriculture. It is largely spontaneous rather than read speech. We score 111,051 clips in the 14 languages that have metrics reported in the AfriVox-v2 paper; baseline figures are taken from that paper, which reports WER only. Simba-S is our own evaluation, and does not cover three of the languages, so its average is over the other eleven. Best per row in bold.

Language SunflowerASR (ours) Sahara v2 Gemini 3 Flash Omni-CTC 7B Omni-CTC 1B Omni-CTC 300M Simba-S
Swahili 0.108 0.071 0.076 0.077 0.097 0.152 0.150
Tswana 0.116 0.140 0.275 0.142 0.210 0.277 0.604
Kinyarwanda 0.133 0.066 0.165 0.104 0.138 0.215 0.591
Sotho 0.152 0.195 0.346 0.421 0.473 0.533 0.596
Luganda 0.160 0.393 0.315 0.422 0.428 0.483 0.341
Zulu 0.161 0.256 0.237 0.202 0.256 0.346 0.396
Xhosa 0.166 0.206 0.292 0.200 0.256 0.344 0.501
Shona 0.209 0.249 0.368 0.286 0.322 0.425 —
Amharic 0.251 0.253 0.249 0.327 0.374 0.490 0.716
Hausa 0.288 0.285 0.269 0.502 0.366 0.401 0.835
Yoruba 0.316 0.271 0.296 0.361 0.365 0.438 0.723
Fulani 0.343 0.352 0.729 0.557 0.569 0.588 —
Akan 0.347 0.307 0.456 0.447 0.497 0.549 —
Igbo 0.425 0.287 0.425 0.459 0.395 0.446 0.998
Overall 0.227 0.238 0.321 0.322 0.339 0.406 0.586

SimbaBench

SimbaBench aggregates read and conversational speech from Common Voice, BembaSpeech, Lwazi, NCHLT and regional collections. The table covers the 19 languages in common with SunflowerASR; baseline figures come from the public leaderboard. Best per row in bold.

Language SunflowerASR (ours) OmniASR-LLM-7B Simba-S Simba-M Gemini 2.5 Pro
Hausa 0.167 0.169 0.641 0.291 0.219
Luganda 0.220 0.171 0.227 0.357 0.421
Swahili 0.246 0.156 0.163 0.252 0.168
Kinyarwanda 0.251 0.192 0.531 0.367 0.463
Luo 0.297 0.271 0.391 0.442 0.494
Igbo 0.304 0.362 0.777 0.850 0.428
Afrikaans 0.312 0.279 0.145 0.313 0.158
Wolof 0.324 0.583 0.581 0.586 1.830
Yoruba 0.335 0.232 0.201 0.415 0.524
Kabyle 0.349 0.287 0.604 0.643 0.933
Tswana 0.353 0.489 0.187 0.306 0.464
Bemba 0.374 0.392 0.391 0.437 0.635
Sotho 0.441 0.634 0.210 0.327 0.495
Ndebele 0.444 0.588 0.246 0.358 0.413
Zulu 0.459 0.401 0.585 0.333 0.359
Xhosa 0.466 0.448 0.272 0.383 0.466
Chichewa 0.545 0.609 0.223 0.468 0.428
Kalenjin 0.580 0.700 0.713 0.734 0.959
Amharic 0.600 0.234 0.330 0.525 0.457
Overall 0.372 0.379 0.390 0.442 0.543

Training Datasets

The model was fine-tuned on a multilingual corpus covering 51 African languages, assembled from a range of publicly available and community-collected speech datasets. The table below lists each source dataset and the languages it contributes to training (ISO 639-3 code in parentheses). A language may appear in more than one source.

Corpus statistics

Languages 51
Dataset subsets 113 (29 source collections, a mean of 2.2 independent sources per language)
Clips before filtering 4.15 M
Removed: suspected label noise (CER > 0.4 against an intermediate checkpoint) 4.1%
Removed: clips longer than Whisper's 30 s window 3.9%
Clips after filtering 3.82 M
Total audio after filtering 7,412 h
Mean clip duration 7.0 s
Packed training samples (clips concatenated into 30 s windows) 1.10 M
Mean clips per pack / mean packed duration 3.5 / 24.2 s
Hours per language 5.2 h (Ruruuli) to 1,392 h (Kinyarwanda), median 81 h
Clips used for the GRPO stage 15,285 (top 200 per subset by headroom)

Per-language hours are listed alongside the benchmark results in the Sunbird Speech Benchmark table above. Clips are filtered and then packed: short clips from the same subset are concatenated to fill the 30 s input window, which keeps every pack monolingual and single-source while cutting padding waste from 76.7% to 19.4%. Of the 113 training subsets, the 110 with held-out test splits make up the Sunbird Speech Benchmark.

Sources

Source dataset Languages
Mozilla Common Voice Afrikaans (afr), Amharic (amh), Rukiga (cgg), Dagbani (dag), Hausa (hau), Igbo (ibo), Kabyle (kab), Kinyarwanda (kin), Kalenjin (kln), Rukonjo (koo), Lendu (led), Thur (lth), Luganda (lug), Luo (luo), Nigerian Pidgin (pcm), Ruruuli (ruc), Kwamba (rwm), Swahili (swa), Tswana (tsn), Rutooro (ttj), Yoruba (yor)
Google FLEURS Afrikaans (afr), Fulani (ful), Hausa (hau), Igbo (ibo), Lingala (lin), Luganda (lug), Luo (luo), Chichewa (nya), Oromo (orm), Shona (sna), Somali (som), Sotho (sot), Swahili (swa), Wolof (wol), Xhosa (xho), Yoruba (yor), Zulu (zul)
Google Waxal Acholi (ach), Akan (aka), Amharic (amh), Dagbani (dag), Dagaare (dga), Ewe (ewe), Fulani (ful), Ikposo (kpo), Lingala (lin), Luganda (lug), Malagasy (mlg), Lumasaba (myx), Runyankole (nyn), Oromo (orm), Shona (sna), Lusoga (xog)
ASR Africa Data Efficiency Benchmark Afrikaans (afr), Amharic (amh), Bambara (bam), Bemba (bem), Ewe (ewe), Fulani (ful), Hausa (hau), Igbo (ibo), Kinyarwanda (kin), Luganda (lug), Oromo (orm), Shona (sna), Wolof (wol), Xhosa (xho), Yoruba (yor), Zulu (zul)
African Next Voices Kikuyu (kik), Kalenjin (kln), Luo (luo), Somali (som), Ndebele (nbl), Sotho (sot), Tswana (tsn), Xhosa (xho), Zulu (zul)
Sunbird SALT Acholi (ach), English (eng), Lugbara (lgg), Luganda (lug), Runyankole (nyn), Ateso (teo)
African Voices Hausa (hau), Igbo (ibo), Nigerian Pidgin (pcm), Yoruba (yor)
NaijaVoices Hausa (hau), Igbo (ibo), Yoruba (yor)
Omnilingual ASR corpus (Meta) Rukiga (cgg), Rukonjo (koo), Rutooro (ttj)
CLEAR Global Hausa (hau), Kanuri (kau)
Shunya Labs Amharic (amh), Lingala (lin)
RobotsMali Bambara ASR Bambara (bam)
AfriSpeech-200 (Intron Health) English (eng)
African-Accented French French (fra)
Afrikaans-30s Afrikaans (afr)
BembaSpeech Bemba (bem)
TutlaytAI Amazigh ASR Berber (ber)
KasuleTrevor Lingala Lingala (lin)
Makerere Radio Speech Luganda (lug)
Digital Divide Data Luhya (luy)
michsethowusu Chichewa (nya)
Soomali ASR Somali (som)
Kallaama Wolof (wol)
KYAGABA Amharic (amh)

Notes

  • African Next Voices is drawn from two collection hubs: Kenya (Kikuyu, Kalenjin, Luo, Somali) and Southern Africa (Ndebele, Sotho, Tswana, Xhosa, Zulu).

Training procedure

SunflowerASR is trained from Whisper large-v3 in two stages. In the first, all 1.55 B parameters are fine-tuned for 5 epochs over the packed corpus described above, with a per-device batch of 64 packed samples and 4 gradient accumulation steps — an effective batch of 256 packed samples, or roughly 1.7 hours of audio per optimiser step, giving about 4,300 optimiser steps per epoch — at a peak learning rate of 2×10⁻⁵ with 100 warmup steps and cosine decay, label smoothing of 0.1, and augmentation applied on the fly: additive noise drawn from a corpus of Ugandan ambient recordings (Sunbird/urban-noise-uganda-61k) plus speed and bandwidth perturbation. The second stage is a reinforcement learning pass with Group Relative Policy Optimisation (GRPO), which optimises the sequence-level metric directly rather than the per-token likelihood of a single reference: for each clip the policy samples 10 complete transcripts, each is scored with the reward -min(CER, 1.0) against the reference, and the group's own mean and standard deviation turn those scores into advantages, so probability mass moves towards the better rollouts under a KL penalty of β = 0.01 against the initial policy. It runs for one epoch over 15,285 clips, selected as the 200 clips per dataset subset with the largest headroom — the gap between the median and the best CER over 10 sampled rollouts, which identifies clips where a better transcript is already inside the model's sampling distribution and only needs to be made more probable — with 64 distinct clips (640 rollouts) per update at a peak learning rate of 10⁻⁵, the encoder frozen so that only the decoder's 906 M parameters are trained. This stage leaves clean-speech accuracy roughly unchanged but improves noisy and spontaneous audio: it largely suppresses the repetition loops Whisper falls into under acoustic stress, makes ordinary substitutions closer to the reference spelling, and helps most for the languages with the least training data.

Training configuration

Base model openai/whisper-large-v3 (1.55 B parameters)
Stage 1 Supervised fine-tuning, all parameters, 5 epochs, LR 2×10⁻⁵ cosine, label smoothing 0.1, effective batch 256 packs
Stage 2 GRPO (trl, loss_type: dapo), decoder only (906 M of 1.55 B), 1 epoch, LR 10⁻⁵
GRPO reward -min(CER, 1.0) against a single reference transcript
GRPO rollouts 10 per clip, temperature 0.8, top-p 0.98; 64 clips (640 rollouts) per update
GRPO regularisation Group-std advantage scaling, KL β = 0.01 against the initialisation
Max completion length 444 tokens (Whisper's 448 learned decoder positions, minus the 4-token forced prefix)
Augmentation Additive urban noise, speed perturbation, occasional 8 kHz downsample-and-restore

Limitations

  • Quality is very uneven across the 51 languages. On the Sunbird Speech Benchmark, WER ranges from 0.08 to 0.67. Check the per-language figures before deploying for any one language, and treat the tail of the table as usable for triage or search rather than for verbatim transcription.
  • The language must be specified at decode time. The model does not perform language identification, and giving it the wrong language token degrades output substantially. Language tokens for the 40 languages Whisper did not originally support are overwritten unused tokens, so the mapping in the Usage section must be used rather than Whisper's own language names.
  • Audio longer than 30 seconds must be segmented before transcription, as with any Whisper model; longer clips are truncated by the encoder.
  • Orthographic convention rather than intelligibility. Both training stages score hypotheses against a single reference transcript, so the model is optimised towards the transcription conventions of these corpora. Languages without a settled orthography inherit that ambiguity, and a transcript may be penalised — or produced — in a spelling a given community would not use.
  • Domain and style. Much of the training data for the lower-resource languages is read speech recorded in quiet conditions. Accuracy is lower on spontaneous, overlapping, noisy or heavily code-switched speech, though this is the setting the GRPO stage improves most.
  • Repetition loops are reduced, not eliminated. Under heavy acoustic stress the model can still emit a repeated run or a hypothesis much longer than the utterance.
  • Benchmark caveat. The Sunbird Speech Benchmark is drawn from the same collections as the training corpus, so it is in-domain for this model and out-of-domain for the baselines it is compared against. AfriVox-v2 and SimbaBench are the arm's-length comparisons.
  • Not for consequential decisions without review. See Intended use.

Citation

If you use this model, please cite the accompanying paper:

@misc{sunflowerasr2026,
  title        = {Group Relative Policy Optimisation Improves Multilingual Speech
                  Recognition in Low-Resource Languages},
  author       = {Hu, Tim Wenjie and Ouma, Evelyn Nafula and Akera, Benjamin and
                  Quinn, John A.},
  year         = {2026},
  institution  = {Sunbird AI},
  howpublished = {\url{https://huggingface.co/Sunbird/SunflowerASR-51-african-languages}}
}

And, for the evaluation set:

@misc{sunbird_speech_benchmark_2026,
  title        = {Sunbird Speech Benchmark: a multilingual ASR test set for 51
                  African languages},
  author       = {Sunbird AI},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/Sunbird/speech-benchmark}}
}

License and acknowledgements

Released under the Apache 2.0 licence, following the licence of the Whisper large-v3 base model. The training corpus is assembled from the publicly available and community-collected datasets credited in Training Datasets; please also respect the licence and citation requirements of those sources. We thank the many organisations and community contributors who collected and released this speech data, without which a model covering these languages would not be possible.

Downloads last month
3,396
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Sunbird/SunflowerASR-51-african-languages

Finetuned
(1)
this model
Adapters
1 model
Finetunes
8 models

Datasets used to train Sunbird/SunflowerASR-51-african-languages

Space using Sunbird/SunflowerASR-51-african-languages 1

Collection including Sunbird/SunflowerASR-51-african-languages