w2v-bert-2-0-0f7df83a·1 events·first seen Aliases: w2v-BERT 2.0
Researchers release DONDO, a family of 21 monolingual and 5 multilingual automatic speech recognition models covering 27 language varieties across Ghana, Sierra Leone, Nigeria, Senegal, Kenya, and Zimbabwe, built on the w2v-BERT 2.0 self-supervised encoder. Models are fine-tuned on read speech from religious texts and achieve average word error rates of 10–13% on multilingual checkpoints, closing most of the gap to monolingual baselines. A lightweight language-conditioning mechanism using one-hot prefix frames allows a single multilingual checkpoint to be steered to a target language at inference. All models are released under Apache-2.0 on Hugging Face via the KhayaAI organisation, covering languages spoken by roughly 100 million first-language speakers.