--- license: apache-2.0 base_model: stepfun-ai/Step-Audio-2-mini library_name: peft tags: - audio - speech-translation - speech-to-speech - luganda - english - stepaudio2 model-index: - name: Step-Audio 2 Mini Luganda-to-English S2ST LoRA results: - task: type: speech-translation name: Luganda-to-English speech translation dataset: type: yigagilbert/luganda-english-cleaned-v1-split name: Luganda-English Cleaned v1 Split split: validation metrics: - type: bleu name: BLEU value: 32.530 - type: chrf name: chrF value: 54.535 - type: wer name: WER on generated English text value: 0.574 - type: comet name: COMET value: 0.717 - type: blaser_2_0_ref name: BLASER 2.0 ref value: 3.762 - type: blaser_2_0_qe name: BLASER 2.0 QE value: 3.723 --- # Step-Audio 2 Mini Luganda-to-English S2ST LoRA This repository contains a LoRA adapter for `stepfun-ai/Step-Audio-2-mini` trained for Luganda speech input to English speech output. The adapter requires the base model and Step-Audio 2 `token2wav` assets for waveform synthesis. Adapter repository: `yigagilbert/stepaudio2-mini-luganda-english-s2st-lora` ## Intended use Research and development for Luganda-to-English speech translation. Validate with native speakers before production use. ## Training data Configured for `yigagilbert/luganda-english-cleaned-v1-split` with columns: `audio_lug`, `audio_eng`, `text_lug`, `text_eng`, `id`, `src_dur_s`, `tgt_dur_s`, `dur_ratio`, `src_speech_ratio`, `tgt_speech_ratio`. ## Evaluation The table below summarizes the held-out validation evaluation used during development. All text metrics were computed on 200 validation examples. Speech metrics were computed on the aligned 197-example subset for which generated audio existed. The merged full-model repository contains the same adapted weights as this adapter after LoRA merge, so it is expected to match these results aside from normal deterministic or runtime differences. Re-run evaluation directly on the merged repository before a strict release if exact reproducibility is required. ### Text and Semantic Metrics | System | Loading Form | Count | BLEU higher | chrF higher | WER lower | COMET higher | BLASER ref higher | BLASER QE higher | |---|---|---:|---:|---:|---:|---:|---:|---:| | Base Step-Audio-2-mini | Base only | 200 | 0.012 | 5.152 | 10.702 | 0.386 | 1.713 | 2.164 | | Fine-tuned Step-Audio | Base + this LoRA adapter | 200 | 32.530 | 54.535 | 0.574 | 0.717 | 3.762 | 3.723 | | Full merged fine-tuned model | Same adapted weights, merged | 200 | 32.530* | 54.535* | 0.574* | 0.717* | 3.762* | 3.723* | | Cascade baseline | ASR + MT + TTS | 200 | 36.778 | 57.971 | 0.521 | 0.737 | 3.839 | 3.776 | `*` The merged full-model row reflects the adapter evaluation because the merged repository is produced by folding this LoRA adapter into the same base model. ### Speech Metrics | System | Count | chrF higher | SpeechBERT P higher | SpeechBERT R higher | SpeechBERT F1 higher | MCD lower | |---|---:|---:|---:|---:|---:|---:| | Fine-tuned Step-Audio | 197 | 54.582 | 0.644 | 0.648 | 0.645 | 629.718 | | Full merged fine-tuned model | 197 | 54.582* | 0.644* | 0.648* | 0.645* | 629.718* | | Cascade baseline | 197 | 57.893 | 0.603 | 0.622 | 0.612 | 613.212 | The unfine-tuned base model emitted no valid speech-token sequences in this 200-sample run, so SpeechBERTScore and MCD were not computed for it. ### Interpretation The LoRA adapter changes the base model from an unusable zero-shot system into a functioning Luganda-to-English speech translation model. The cascade remains stronger on text and semantic metrics, while the fine-tuned Step-Audio system is operationally simpler and scored higher on the WavLM-based SpeechBERTScore F1 proxy. ## License Adapter code and metadata are Apache-2.0. Check the dataset license separately before redistribution.