You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Step-Audio 2 Mini Luganda-to-English S2ST LoRA

This repository contains a LoRA adapter for stepfun-ai/Step-Audio-2-mini trained for Luganda speech input to English speech output. The adapter requires the base model and Step-Audio 2 token2wav assets for waveform synthesis.

Adapter repository: yigagilbert/stepaudio2-mini-luganda-english-s2st-lora

Intended use

Research and development for Luganda-to-English speech translation. Validate with native speakers before production use.

Training data

Configured for yigagilbert/luganda-english-cleaned-v1-split with columns: audio_lug, audio_eng, text_lug, text_eng, id, src_dur_s, tgt_dur_s, dur_ratio, src_speech_ratio, tgt_speech_ratio.

Evaluation

The table below summarizes the held-out validation evaluation used during development. All text metrics were computed on 200 validation examples. Speech metrics were computed on the aligned 197-example subset for which generated audio existed.

The merged full-model repository contains the same adapted weights as this adapter after LoRA merge, so it is expected to match these results aside from normal deterministic or runtime differences. Re-run evaluation directly on the merged repository before a strict release if exact reproducibility is required.

Text and Semantic Metrics

System Loading Form Count BLEU higher chrF higher WER lower COMET higher BLASER ref higher BLASER QE higher
Base Step-Audio-2-mini Base only 200 0.012 5.152 10.702 0.386 1.713 2.164
Fine-tuned Step-Audio Base + this LoRA adapter 200 32.530 54.535 0.574 0.717 3.762 3.723
Full merged fine-tuned model Same adapted weights, merged 200 32.530* 54.535* 0.574* 0.717* 3.762* 3.723*
Cascade baseline ASR + MT + TTS 200 36.778 57.971 0.521 0.737 3.839 3.776

* The merged full-model row reflects the adapter evaluation because the merged repository is produced by folding this LoRA adapter into the same base model.

Speech Metrics

System Count chrF higher SpeechBERT P higher SpeechBERT R higher SpeechBERT F1 higher MCD lower
Fine-tuned Step-Audio 197 54.582 0.644 0.648 0.645 629.718
Full merged fine-tuned model 197 54.582* 0.644* 0.648* 0.645* 629.718*
Cascade baseline 197 57.893 0.603 0.622 0.612 613.212

The unfine-tuned base model emitted no valid speech-token sequences in this 200-sample run, so SpeechBERTScore and MCD were not computed for it.

Interpretation

The LoRA adapter changes the base model from an unusable zero-shot system into a functioning Luganda-to-English speech translation model. The cascade remains stronger on text and semantic metrics, while the fine-tuned Step-Audio system is operationally simpler and scored higher on the WavLM-based SpeechBERTScore F1 proxy.

License

Adapter code and metadata are Apache-2.0. Check the dataset license separately before redistribution.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yigagilbert/stepaudio2-mini-luganda-english-s2st-lora

Adapter
(2)
this model

Evaluation results