Instructions to use waxal-benchmarking/whisper-small-waxal-loo-ewe with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use waxal-benchmarking/whisper-small-waxal-loo-ewe with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="waxal-benchmarking/whisper-small-waxal-loo-ewe")# Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("waxal-benchmarking/whisper-small-waxal-loo-ewe") model = AutoModelForSpeechSeq2Seq.from_pretrained("waxal-benchmarking/whisper-small-waxal-loo-ewe", device_map="auto") - Notebooks
- Google Colab
- Kaggle
whisper-small-waxal-loo-ewe · (leave-Ewe-out ablation)
⚠️ This is a research/ablation model, not a deployment model. It is Whisper-Small fine-tuned on 18 WAXAL languages EXCLUDING Ewe. It is used to measure novel-language transfer — how well a model does on Ewe, which it was never trained on. Do not use it for Ewe ASR (use the monolingual WAXAL model for that).
What this tests
WAXAL's headline finding is that in-language data is essential. Whisper (unlike Omnilingual) does not cover most WAXAL languages in pretraining, so it lets us ask a clean question with no forgetting confound: can training on sibling languages build a language the model never saw? Here Ewe has unseen Kwa relatives in the training mix (Akan, Ikposo) — also outside Whisper's pretraining.
Result — the novel-language transfer ladder for Ewe
| Setting | WER | CER |
|---|---|---|
| Whisper zero-shot (base) | Unsupported (not in Whisper's 99 languages) | — |
| This model (trained on 18 siblings, Ewe held out) | 101.4 | 56.8 |
| Monolingual Whisper-Small (trained on Ewe) | 32.3 | — |
The gap is stark: sibling training lands far above the monolingual model — training on neighbours cannot substitute for in-language data, even for a genuinely unseen language.
Training
Whisper-Small fine-tuned on the pooled train splits of the 18 non-Ewe WAXAL
languages (google/WaxalNLP),
balanced by capping each language, multilingual (no forced language token). 16 kHz mono;
transcripts NFC-normalized + lower-cased, punctuation removed, diacritics preserved.
| Hyperparameter | Value |
|---|---|
| Base | openai/whisper-small (244M) |
| Training languages | 18 (all WAXAL except Ewe) |
| Utterances / language | 3,000 (balanced) |
| Steps | 4,000 |
| Optimizer | AdamW, lr 1e-5, 200 warmup |
| Batch size | 16 |
| Precision | fp16 |
| Hardware | 1× H200 |
Languages it was trained on: Acholi, Akan, Amharic, Dagbani, Dagaare, Fula, Ikposo, Lingala, Luganda, Masaaba, Malagasy, Nyankole, Oromo, Sidama, Shona, Soga, Tigrinya, Wolaytta.
Evaluation
Scored on the held-out Ewe test split (jiwer WER/CER, NFC-normalized, lower-cased, diacritics preserved) — the same protocol as the rest of the benchmark.
Usage
from transformers import WhisperForConditionalGeneration, WhisperProcessor
import torch, soundfile as sf
model = WhisperForConditionalGeneration.from_pretrained("waxal-benchmarking/whisper-small-waxal-loo-ewe")
processor = WhisperProcessor.from_pretrained("waxal-benchmarking/whisper-small-waxal-loo-ewe")
wav, sr = sf.read("your_audio.wav") # 16 kHz mono
feats = processor(wav, sampling_rate=16000, return_tensors="pt").input_features
ids = model.generate(feats) # multilingual; no forced language token
print(processor.batch_decode(ids, skip_special_tokens=True))
Citation
Part of the WAXAL ASR Benchmark (arXiv:2606.02375).
@article{waxalnet2026,
title = {The WAXAL ASR Benchmark: Fine-Tuned Edge Models Across 19 African Languages},
author = {Olufemi, Victor Tolulope and Babatunde, Oreoluwa and Njema, Ramsey and others},
year = {2026},
note = {arXiv preprint arXiv:2606.02375}
}
Acknowledgements
Supported by Lynguallabs, Open Token, and CMU Africa.
- Downloads last month
- 17
Model tree for waxal-benchmarking/whisper-small-waxal-loo-ewe
Base model
openai/whisper-small