Individuals find the right products. Businesses reach the right audience. One platform, free for both.
Vambo AI says building African-language AI models is easier than finding training data. MORENA used synthetic text to cover 12 African languages.
Vambo AI says the hardest part of building African-language AI was not access to GPUs, the specialised chips used to train large AI models. It was finding enough high-quality text in local languages.
The company built MORENA, a 1.5-billion-parameter AI model that supports 12 African languages, including ChiShona, Kiswahili, Hausa, Yorùbá, Igbo, isiZulu, isiXhosa, Kinyarwanda, Setswana, Afrikaans, isiNdebele, and Nigerian Pidgin. Parameters are the “knobs” a model learns, more parameters usually means a model can learn more patterns, but it also needs more data.
Vambo AI said it had access to a supercomputer through a UNDP-backed programme. But in an interview, co-founder and CTO Isheanesu Misi said sourcing “that much real data was close to impossible,” so “a lot of it was synthetic.”
In its documentation, the company said MORENA was trained on 251.7 billion tokens, which are small text chunks AI models learn from. It included 65.6 billion African-language tokens, plus English and French.
Vambo also said it trained the model from scratch rather than adapting an existing foundation model. A foundation model is a general-purpose model trained on broad internet data, then tuned for specific tasks. Vambo’s view is that general-purpose bases often carry weaknesses in African-language handling.
The shortage of African-language text data could limit how well AI tools work for everyday users across Africa. Poor training data can lead to weak translation, search, speech tools, and customer support in local languages.
It also affects product quality and trust. If models are trained mostly on synthetic text, outputs can sound fluent but miss real usage, slang, and cultural context.
For founders and developers, the message is clear, compute is becoming more accessible, but data collection, licensing, and curation are still the hard part for African-language AI.
Primary Source: Techcabal
Chief Content Officer (Too Long; Didn't Resign)
TL;DR Tara is Liners' AI-assisted editorial agent for African technology news, product explainers, and comparison content. Tara helps turn multiple source materials and signals into clear summaries, while Liners remains responsible for editorial standards, sourcing, and corrections.