For years, Sheriff Issaka kept running into the same problem while building language technology for Africa: there was not enough data. The data needed to train systems for African languages was scarce, and much of what was available was not good enough.
That problem changed what African Languages Lab, the AI research and deployment company Issaka founded in 2020, spent its time doing. Instead of building systems only with the available data, the company began collecting the data it could not find.
The company says it has since built the largest collection of African-language data in existence, covering more than 70 languages, including Amharic, Hausa, Zulu, Twi, Igbo, and Yoruba.
On Tuesday, African Languages Lab launched Mansa, a multilingual and multimodal AI platform that puts about 30 of those languages into production. It is available on the web and mobile, with APIs for developers and businesses.

The company says its datasets now contain more than 100 billion curated tokens, or pieces of text used to train AI models, alongside more than 19,000 hours of speech recordings that have been reviewed and validated by language experts.
The problem is bigger than poor translations
Africa is home to more than 2,000 languages, but fewer than 5% have the resources needed for natural language processing. Most African languages remain poorly represented in digital datasets, limiting how well AI systems can understand them.
“The models are just very bad at understanding our languages,” Issaka said. “If you try to have a full-blown conversation with an LLM in most African languages, they just cannot pull it off.”
There is also a cost problem. African languages can require significantly more tokens than English for the same input. Chioma Agwuedo, Executive Director of TechHerNG, a Nigerian nonprofit focused on women and technology, told TechCabal in 2025 that generating a token in Yoruba can cost four times as much as generating one in English.
“We pay more to get worse performance from these models,” Issaka said.
The data gap can also affect safety. Issaka said safety tuning, the process of training models to follow safety rules, is generally stronger in languages with more training data, leaving safeguards weaker in some under-resourced African languages.
He said that when researchers use harmful prompts in some African languages, users are at least 10 times as likely to receive erroneous responses that would typically be blocked in higher-resource languages.
African Languages Lab is entering a growing effort to build AI systems that can understand and generate African languages. Google has expanded AI Search to African languages and launched WAXAL, an open-source speech dataset covering 21 Sub-Saharan African languages. Nigerian startups such as Intron and Spitch are also building speech recognition and text-to-speech systems for African languages.
A decade collecting the data
Issaka said his team has spent about a decade collecting the data it could not find. The company uses freely available and openly licenced datasets, but also works directly with communities and pays people to collect and validate new data.
“We’ve had a lot of people approaching us every week to say, hey, my language is under-resourced. I want to come and give you data on my language,” Issaka said.
The data is unevenly distributed. Yoruba, Swahili and Zulu have more data than smaller, lower-resource languages, he said. Most of the collection is text, although the company has increasingly invested in speech.
“For a lot of these people, they cannot read and write to begin with. And so speech then becomes the medium in which you’re able to truly reach the most people who need the most access to these technologies,” he said.
Existing data-collection tools also presented a problem. Issaka said researchers often had to use English as a bridge when collecting data between African languages, which prompted the company to build All Voices about six years ago. The platform lets contributors collect and validate data directly between African languages, without a bridge language. Contributors can also earn money for supplying or validating data.
“I’m not exaggerating, you can go and contribute data to it in any language in the world. You don’t have to go through any bridge languages,” Issaka said.
He said All Voices now receives two or three emails each week from people who want to contribute data in underrepresented languages.
“It really tells us that the community also cares about this. They want to build with us,” Issaka said.
Issaka said contributors are told what their data will be used for and can choose not to allow it to be used. In some large-scale data-collection projects, he said, contributors are paid to collect data and are told that agreeing to participate means the resulting data can be made publicly available for research.
The company also has partnerships in which contributors receive incentives when their data helps a model reach a certain performance level.
“We don’t want to just be extracting from these communities,” he said. “We want to make sure that they’re part of the value chain.”
But Issaka acknowledged that the approach is still evolving.
“Is this perfect? I don’t think so,” he said. “But I do think we are doing the best that we can right now to ensure that our practices and our technologies are not only extracting, but in a way also to give back to the communities.”
African Languages Lab says its technology covers more than 70 African languages, but Mansa currently puts only about 30 into production.
That distinction reflects the company’s decision to separate languages for which it has collected data from those it believes are ready for public use.
“We only put about 30 of them in production. And that’s because these are the languages that we have the most support for. These are the languages that we can vouch for,” Issaka said.
The remaining languages are not necessarily unsupported. The company simply does not believe their performance is good enough yet.
“We could extend it to all the 70 languages, but we know that the performance for the 40 or so languages isn’t as great as what we have for the 30,” he said.
For languages in production, Mansa supports text and voice conversations, as well as images and video. It also offers translation, transcription, text-to-speech and real-time interpretation.
The platform includes an AI agent that can carry out tasks rather than respond to prompts. Issaka said users can ask Mansa to read and summarise emails, search for new research in a particular field, or track news and social media trends and return regular briefings.
African Languages Lab did not build the foundation model powering Mansa from scratch. It uses an existing model from Chinese AI company MiniMax as the base, then adapts it for African languages, contexts and use cases.
“We make no secret of this: we like to reuse before we rebuild,” Issaka said.
He said MiniMax was chosen partly because of its multimodal capabilities, giving Mansa a base that could handle text, speech, images and video. The model has almost half a trillion parameters.
African Languages Lab said it trains some models and components from scratch for specific processes within the Mansa ecosystem, while its flagship model is adapted from MiniMax.
Issaka said the company started with translation models and built internal tools to compare their performance with other models. It evaluates its systems using automated benchmarks and human evaluation because existing benchmarks do not always measure performance in African languages well.
“We’re building with the community,” he said.
Issaka said the company believes this work gives Mansa an advantage in translation and speech. He said the system performs better than tools such as ChatGPT and Google Translate on some African-language translation tasks, while its speech systems are better calibrated for African accents.
He also said Mansa’s speech systems can stream in African languages, which he described as an area where other models currently do not offer the same capability.
The challenge of African accents
Speech is another important part of making AI accessible in African languages. Issaka said this is particularly important for people who cannot rely on text to interact with technology.
“For a lot of these people, they cannot read and write to begin with. And so speech then becomes the medium in which you’re able to truly reach the most people who need the most access to these technologies,” he said.
But building speech systems that work reliably across African languages and accents remains difficult.
“I’ll be honest with you, it’s not a solved problem,” Issaka said.
The company is still collecting data on different accents and working to annotate it properly. Issaka said this is difficult because African languages are closely connected, with variations that do not always fit neatly into individual language categories.
Instead of building one language at a time, the company is training its models across languages so they can better handle variations in accents and inputs.
“If you draw it as a map, it’s all very much interconnected,” Issaka said.
The approach is intended to help the model recognise variations rather than treating every language and accent as an entirely separate problem. But Issaka said it remains an ongoing challenge.
From 30 languages to 1,000
Issaka sees communication as one of the biggest opportunities. A bank, for example, could use Mansa to communicate with customers in different African languages, while other businesses could build translation, transcription, or speech capabilities into existing services.
“I think communication is going to be the big one because the way we are building Mansa, we’re trying to build an infrastructure,” he said. “An infrastructure where everybody can build on.”
The company wants African innovators building products for the continent to use Mansa as a language and intelligence layer, similar to how developers build applications on top of other major AI models.
But for Issaka, the longer-term goal is broader than the launch itself.
For African Languages Lab, the 30 languages currently available on Mansa are only a starting point. Issaka wants to expand to 70, then 200 and eventually 1,000.
“I don’t know how many we’ll be able to realistically cover in our lifetimes or ever,” he said. “But I’d love to see that number increase by one, at least by one every time. Every plus one will make me very happy.”
True scale demands moving beyond surface-level integrations to robust execution. We’ve filtered the noise out of Moonshot 2026, optimising the conference strictly for high-calibre connections between startup founders, global financial operators, enterprise leaders and individuals rewiring Africa’s technical frameworks. Get 20% off Early Bird tickets for a limited time.

