Twelve to Twenty Points
In one benchmark study the best large language model tested, GPT-4o, answered 12.0 to 19.9 percentage points less accurately in eleven African languages than in English. AfroBench, covering 64 languages and 15 tasks, records gaps reaching 28 points against English and 19 against French. The word ordinarily attached to these models is general, and it is worth being precise about what the generality is over. Model capability is measured against a distribution inherited from training data drawn overwhelmingly from the public internet, which is not a uniform sample of human activity; performance degrades with distance from the middle of that slice, and these benchmarks put a number on a degradation that is otherwise asserted rather than measured. Lelapa AI's answer runs to 400 million parameters and no hyperscaler. Why the equity argument and the engineering argument are usually conflated to the cost of the second, what follows for procurement that has been treating model choice as a ranking exercise, and the dependency question governments have not examined.
Wyre's opinion bylines are editorial personas of Floof Digital LLC, not separate members of staff. Essays are produced with AI assistance under human editorial direction. How Wyre works.
In one benchmark study the best-performing large language model tested, GPT-4o, answered questions 12.0 to 19.9 percentage points less accurately in eleven African languages than in English. AfroBench, a broader evaluation covering 64 languages, 15 tasks and 22 datasets, records gaps reaching 28 points against English and 19 against French.
The word ordinarily attached to these models is general. It is worth being precise about what the generality is over.
A frontier is a shape, not a height
Model capability is measured against a distribution of tasks, and that distribution is inherited from training data drawn overwhelmingly from the public internet. The internet is not a uniform sample of human activity. It is heavily English, heavily commercial, heavily recent and heavily written by the kinds of people who write things on the internet.
A model that is excellent across that distribution is not therefore excellent across human language use. It is excellent across a particular, well-documented slice of it, and its performance degrades with distance from the middle of that slice. The African language benchmarks are valuable precisely because they put a number on the degradation, which is otherwise asserted rather than measured.
Twelve to twenty points is not a rounding error. In the medical and clinical subsets used in that study it is the difference between a tool that can be deployed and one that cannot.
The response has not been to wait
Lelapa AI's InkubaLM has 400 million parameters, three to four orders of magnitude fewer than a frontier system. It was trained from scratch on 2.4 billion tokens covering isiZulu, Yoruba, Hausa, Swahili and isiXhosa, languages with roughly 364 million speakers between them, and it is deliberately small enough to run without a hyperscaler. The model is named for the dung beetle, which moves 250 times its own weight.
The strategic content of that design is the deployment constraint, not the parameter count. A model that fits on modest local hardware runs where connectivity is unreliable, runs at predictable cost, and runs without exporting the data it processes to a jurisdiction whose courts the user has no access to. Each of those is a policy property, arrived at through engineering.
Two arguments that are usually conflated
The first is about equity: speakers of under-resourced languages are poorly served by systems trained mostly on English, and that gap tracks existing inequalities in who gets to benefit from a general-purpose technology. It is a real argument and the benchmarks support it.
The second is about engineering, and it is the one with wider consequences. Parameters spent on a target task outperform parameters spent on everything, for that task. A small model trained on the languages in question beats a far larger one on those languages. This is not a claim about Africa. It is a claim about allocation, and it holds anywhere a task sits far enough from the centre of the training distribution.
Conflating the two lets the second be dismissed as advocacy. It is not advocacy. It is a reproducible result, and organisations in wealthy markets with narrow, high-volume, domain-specific tasks are in structurally the same position as the labs that discovered it, whether or not they have noticed.
What follows for procurement, and for policy
For buyers, the useful discipline is to stop treating model choice as a ranking exercise. The relevant question is not which model is best but which is best on a defined task at a defined cost with a defined deployment constraint, and the answer is frequently not the largest available. That is unglamorous and it is where most of the wasted spend in this category will turn out to have been.
For governments, the sovereignty argument is the one that has been under-examined. A state whose public services depend on inference performed in another country, on infrastructure it does not control, under commercial terms it did not negotiate, has taken on a dependency it would not accept in any other utility. Models small enough to run domestically change the character of that dependency at a capability cost that the benchmark gaps suggest may, for many public-sector tasks in local languages, be negative.
The thing worth keeping
The prevailing account of AI progress is vertical: each generation larger and more capable than the last, with everyone else waiting for access. The African labs are describing a different geometry. Capability is not one number. It is a surface with a shape, that shape follows the data, and where the surface is low, a small well-aimed model outperforms a large one that was never pointed at you.
Four hundred million parameters, five languages, no hyperscaler, and it wins on the work it was built for. That result is not a consolation prize for the under-funded. It is a finding about how this technology allocates its advantages, and it was produced by the people with the least room to get it wrong.
Sources
- Institute for Disease Modeling and co-authors, "Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments", Proceedings of the AAAI Conference on Artificial Intelligence. Source for the 12.0 to 19.9 percentage point gap for GPT-4o between English and the average of eleven African languages, and for the benchmark construction from Winogrande and MMLU sections including college medicine, clinical knowledge and virology.
- AfroBench, McGill NLP. Source for the evaluation across 64 African languages, 15 tasks and 22 datasets, and gaps of up to 28 points against English and 19 against French.
- Lelapa AI, "InkubaLM: A Small Language Model for Low-Resource African Languages", and the InkubaLM-0.4B model card. Source for the parameter count, the 2.4 billion token training run, the five languages and their speaker population, the naming, and the design intent to run without a hyperscaler.
- The Economist, issue of 8 to 14 August 2026, for surveying African AI startups building small task-specific models. Its Common Crawl page-count comparison could not be verified against a primary source and is not relied on here.