Voice-enabled technology has become routine in many parts of the world, from requesting directions to dictating messages.
However, for hundreds of millions of people across Sub-Saharan Africa, this convenience remains out of reach due to limited support for local languages in digital systems.
This gap has largely been driven by the absence of accessible, high-quality speech data for African languages—an issue the newly launched WAXAL dataset aims to address.
Named after the Wolof word meaning “speak”, WAXAL is a large-scale, open speech dataset developed over three years to support inclusive artificial intelligence and speech technologies across Africa. The dataset covers 21 African languages, including Acholi, Hausa, Luganda, and Yoruba, and contains more than 11,000 hours of speech data drawn from nearly two million individual recordings.
Of this total, approximately 1,250 hours are transcribed speech suitable for automatic speech recognition (ASR), while more than 20 hours of studio-quality recordings are designed for text-to-speech (TTS) voice synthesis.
Built Through African-Led Collaboration
WAXAL is the result of a broad, Africa-led collaboration involving universities, research institutions, and media professionals across the continent.
Key partners include Makerere University in Uganda and the University of Ghana, which together led data collection for 13 languages. Digital Umuganda in Rwanda coordinated data gathering for five widely spoken languages, while Media Trust and Loud n Clear contributed expertise in producing high-quality studio recordings for voice synthesis.
The project also partnered with the African Institute for Mathematical Sciences (AIMS) to develop multilingual datasets intended for future releases.
Under the project’s governance framework, participating institutions retain ownership of the data they collected, while jointly contributing to the broader goal of making African language resources available to the global research community.
Ethical, Real-World Speech Collection
To reflect how people naturally speak, contributors were asked to describe images in their native languages, allowing the dataset to capture authentic speech patterns rather than scripted responses. In parallel, professional voice actors were recorded in studio environments to generate the high-fidelity audio required for advanced text-to-speech applications.
Developers say this dual approach balances realism with technical quality, making WAXAL suitable for both research and production-level speech systems.
Supporting Innovation and Language Preservation
Beyond enabling new voice-based technologies, the creators of WAXAL hope the dataset will contribute to the digital preservation of African languages, many of which remain underrepresented in modern computing.
The complete WAXAL collection has been released under an open license and is now publicly accessible on Hugging Face, allowing researchers, developers, and institutions worldwide to build and experiment with African language speech models.




