whisper-small-waxal-loo-ewe · (leave-Ewe-out ablation)

⚠️ This is a research/ablation model, not a deployment model. It is Whisper-Small fine-tuned on 18 WAXAL languages EXCLUDING Ewe. It is used to measure novel-language transfer — how well a model does on Ewe, which it was never trained on. Do not use it for Ewe ASR (use the monolingual WAXAL model for that).

What this tests

WAXAL's headline finding is that in-language data is essential. Whisper (unlike Omnilingual) does not cover most WAXAL languages in pretraining, so it lets us ask a clean question with no forgetting confound: can training on sibling languages build a language the model never saw? Here Ewe has unseen Kwa relatives in the training mix (Akan, Ikposo) — also outside Whisper's pretraining.

Result — the novel-language transfer ladder for Ewe

Setting WER CER
Whisper zero-shot (base) Unsupported (not in Whisper's 99 languages) —
This model (trained on 18 siblings, Ewe held out) 101.4 56.8
Monolingual Whisper-Small (trained on Ewe) 32.3 —

The gap is stark: sibling training lands far above the monolingual model — training on neighbours cannot substitute for in-language data, even for a genuinely unseen language.

Training

Whisper-Small fine-tuned on the pooled train splits of the 18 non-Ewe WAXAL languages (google/WaxalNLP), balanced by capping each language, multilingual (no forced language token). 16 kHz mono; transcripts NFC-normalized + lower-cased, punctuation removed, diacritics preserved.

Hyperparameter Value
Base openai/whisper-small (244M)
Training languages 18 (all WAXAL except Ewe)
Utterances / language 3,000 (balanced)
Steps 4,000
Optimizer AdamW, lr 1e-5, 200 warmup
Batch size 16
Precision fp16
Hardware 1× H200

Languages it was trained on: Acholi, Akan, Amharic, Dagbani, Dagaare, Fula, Ikposo, Lingala, Luganda, Masaaba, Malagasy, Nyankole, Oromo, Sidama, Shona, Soga, Tigrinya, Wolaytta.

Evaluation

Scored on the held-out Ewe test split (jiwer WER/CER, NFC-normalized, lower-cased, diacritics preserved) — the same protocol as the rest of the benchmark.

Usage

from transformers import WhisperForConditionalGeneration, WhisperProcessor
import torch, soundfile as sf

model = WhisperForConditionalGeneration.from_pretrained("waxal-benchmarking/whisper-small-waxal-loo-ewe")
processor = WhisperProcessor.from_pretrained("waxal-benchmarking/whisper-small-waxal-loo-ewe")
wav, sr = sf.read("your_audio.wav")            # 16 kHz mono
feats = processor(wav, sampling_rate=16000, return_tensors="pt").input_features
ids = model.generate(feats)                    # multilingual; no forced language token
print(processor.batch_decode(ids, skip_special_tokens=True))

Citation

Part of the WAXAL ASR Benchmark (arXiv:2606.02375).

@article{waxalnet2026,
  title  = {The WAXAL ASR Benchmark: Fine-Tuned Edge Models Across 19 African Languages},
  author = {Olufemi, Victor Tolulope and Babatunde, Oreoluwa and Njema, Ramsey and others},
  year   = {2026},
  note   = {arXiv preprint arXiv:2606.02375}
}

Acknowledgements

Supported by Lynguallabs, Open Token, and CMU Africa.

Downloads last month
17
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for waxal-benchmarking/whisper-small-waxal-loo-ewe

Finetuned
(3776)
this model

Dataset used to train waxal-benchmarking/whisper-small-waxal-loo-ewe

Paper for waxal-benchmarking/whisper-small-waxal-loo-ewe