Curated summary
WAXAL: A large-scale open resource for African language speech technology
WAXAL is an open speech dataset designed to address the severe lack of technology resources for African languages. Covering 27 Sub-Saharan African languages spoken by more than 100 million people, it offers 1,846 hours of transcribed speech for ASR and over 565 hours of high-quality recordings for TTS under the permissive CC-BY-4.0 license. Built with African academic and community organizations, WAXAL aims to support inclusive speech systems while preserving and expanding access to the continent’s linguistic diversity.
The Need for African-Language Speech Data
- Voice assistants and transcription tools primarily support high-resource languages.
- Hundreds of millions of African speakers remain underserved, despite Sub-Saharan Africa having more than 2,000 languages.
- Limited training data makes it difficult to build accurate systems for tonal, morphologically complex, and code-switching languages.
- WAXAL was developed beginning in 2021 to provide openly accessible resources for research and commercial innovation.
WAXAL-ASR: Natural, Spontaneous Speech
- Contains approximately 1,846 hours of transcribed audio.
- Participants described image prompts across more than 50 topics rather than reading prepared scripts.
- The method captured:
- Natural linguistic variation
- Tonal distinctions
- Spontaneous speech patterns
- Code-switching
- Image-based elicitation produced more realistic speech data than conventional scripted recordings.
WAXAL-TTS: High-Quality Voice Recordings
- Includes more than 565 hours of phonetically balanced, high-fidelity audio.
- Community members collaborated in pairs to create scripts of 10,000–20,000 words and alternate between reading and recording.
- Some participants built custom studio boxes to improve recording quality.
- Audio was segmented, aligned with transcripts, and reviewed for accuracy.
- The clean recordings support natural-sounding synthetic voices and conversational systems.
Supporting Full-Duplex Conversational Systems
- The ASR data represents varied, real-world speech input.
- The TTS data provides high-quality material for generating clear and natural speech.
- Together, the datasets can support systems that both understand and respond using African languages.
Community-Led Development
- African universities and community organizations led the data collection, with guidance from Google Research.
- Makerere University contributed data for nine languages.
- The University of Ghana worked on eight languages, including image-prompted ASR collection.
- Digital Umuganda and Addis Ababa University supported ASR collection for regional languages.
- Media Trust, Loud n Clear, and the African Institute for Mathematical Sciences Senegal led studio-quality TTS recordings.
- Partners retain ownership of the data while making it openly available under CC-BY-4.0.
Research Enabled by the Project
- A community-driven impaired-speech collection cookbook produced an open Akan dataset covering conditions such as cerebral palsy and stammering.
- A separate project introduced a 5,000-hour corpus for five Ghanaian languages: Akan, Ewe, Dagbani, Dagaare, and Ikposo.
- Benchmarking of Whisper, XLS-R, MMS, and W2v-BERT across 13 African languages showed that scaling benefits depend on linguistic complexity and domain alignment.
- A literature review cataloged 74 datasets spanning 111 African languages and emphasized the need for conversational, multi-domain data and metrics such as Character Error Rate (CER).
WAXAL provides a strong foundation for building more accurate and inclusive African-language speech technologies. Researchers and developers should use the openly licensed corpus, contribute derivative work, and help extend the collection to additional languages and communities.
Related reading
Continue with another curated summary.
Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence
Read originalTowards a quantum computer that learns from its errors
Read originalExpanding our Heat Resilience data to 50+ global cities
Read originalOptimizing cloud economics with linear elastic caching
Read original