Independent newsroom The Wyre News Network OpEd desk

Analysis 5 min read

Twelve to Twenty Points

In one benchmark study the best large language model tested, GPT-4o, answered 12.0 to 19.9 percentage points less accurately in eleven African languages than in English. AfroBench, covering 64 languages and 15 tasks, records gaps reaching 28 points against English and 19 against French. The word ordinarily attached to these models is general, and it is worth being precise about what the generality is over. Model capability is measured against a distribution inherited from training data drawn overwhelmingly from the public internet, which is not a uniform sample of human activity; performance degrades with distance from the middle of that slice, and these benchmarks put a number on a degradation that is otherwise asserted rather than measured. Lelapa AI's answer runs to 400 million parameters and no hyperscaler. Why the equity argument and the engineering argument are usually conflated to the cost of the second, what follows for procurement that has been treating model choice as a ranking exercise, and the dependency question governments have not examined.

Wyre's opinion bylines are editorial personas of Floof Digital LLC, not separate members of staff. Essays are produced with AI assistance under human editorial direction. How Wyre works.

In one benchmark study the best-performing large language model tested, GPT-4o, answered questions 12.0 to 19.9 percentage points less accurately in eleven African languages than in English. AfroBench, a broader evaluation covering 64 languages, 15 tasks and 22 datasets, records gaps reaching 28 points against English and 19 against French.

The word ordinarily attached to these models is general. It is worth being precise about what the generality is over.

A frontier is a shape, not a height

Model capability is measured against a distribution of tasks, and that distribution is inherited from training data drawn overwhelmingly from the public internet. The internet is not a uniform sample of human activity. It is heavily English, heavily commercial, heavily recent and heavily written by the kinds of people who write things on the internet.

A model that is excellent across that distribution is not therefore excellent across human language use. It is excellent across a particular, well-documented slice of it, and its performance degrades with distance from the middle of that slice. The African language benchmarks are valuable precisely because they put a number on the degradation, which is otherwise asserted rather than measured.

Twelve to twenty points is not a rounding error. In the medical and clinical subsets used in that study it is the difference between a tool that can be deployed and one that cannot.

The response has not been to wait

Lelapa AI's InkubaLM has 400 million parameters, three to four orders of magnitude fewer than a frontier system. It was trained from scratch on 2.4 billion tokens covering isiZulu, Yoruba, Hausa, Swahili and isiXhosa, languages with roughly 364 million speakers between them, and it is deliberately small enough to run without a hyperscaler. The model is named for the dung beetle, which moves 250 times its own weight.

The strategic content of that design is the deployment constraint, not the parameter count. A model that fits on modest local hardware runs where connectivity is unreliable, runs at predictable cost, and runs without exporting the data it processes to a jurisdiction whose courts the user has no access to. Each of those is a policy property, arrived at through engineering.

Two arguments that are usually conflated

The first is about equity: speakers of under-resourced languages are poorly served by systems trained mostly on English, and that gap tracks existing inequalities in who gets to benefit from a general-purpose technology. It is a real argument and the benchmarks support it.

The second is about engineering, and it is the one with wider consequences. Parameters spent on a target task outperform parameters spent on everything, for that task. A small model trained on the languages in question beats a far larger one on those languages. This is not a claim about Africa. It is a claim about allocation, and it holds anywhere a task sits far enough from the centre of the training distribution.

Conflating the two lets the second be dismissed as advocacy. It is not advocacy. It is a reproducible result, and organisations in wealthy markets with narrow, high-volume, domain-specific tasks are in structurally the same position as the labs that discovered it, whether or not they have noticed.

What follows for procurement, and for policy

For buyers, the useful discipline is to stop treating model choice as a ranking exercise. The relevant question is not which model is best but which is best on a defined task at a defined cost with a defined deployment constraint, and the answer is frequently not the largest available. That is unglamorous and it is where most of the wasted spend in this category will turn out to have been.

For governments, the sovereignty argument is the one that has been under-examined. A state whose public services depend on inference performed in another country, on infrastructure it does not control, under commercial terms it did not negotiate, has taken on a dependency it would not accept in any other utility. Models small enough to run domestically change the character of that dependency at a capability cost that the benchmark gaps suggest may, for many public-sector tasks in local languages, be negative.

The thing worth keeping

The prevailing account of AI progress is vertical: each generation larger and more capable than the last, with everyone else waiting for access. The African labs are describing a different geometry. Capability is not one number. It is a surface with a shape, that shape follows the data, and where the surface is low, a small well-aimed model outperforms a large one that was never pointed at you.

Four hundred million parameters, five languages, no hyperscaler, and it wins on the work it was built for. That result is not a consolation prize for the under-funded. It is a finding about how this technology allocates its advantages, and it was produced by the people with the least room to get it wrong.

Sources

More Opinion

From the same desk

Perspective

The Patch Notes Nobody Reads

A wave of security stories this week, from 270 compromised Zimbra servers to a UK proposal to block risky suppliers in secret, points to a shift worth watching if you buy software rather than build it: vendors are quietly disclosing less. WhatsApp's public, specific patch notes on passkeys sit next to Zimbra's silence as a case study in what disclosure looks like when it works and when it disappears. Mirage2FA is reportedly hitting 4,500 companies through Microsoft 365 login flows, and none of that risk shows up in a vendor changelog until outside researchers surface it. The essay argues that whether a vendor still publishes patch notes should become an actual vendor selection criterion for marketing and agency teams, not an IT afterthought raised only at renewal.

5 min

Analysis

The Secrecy Becomes the System

Silent patches and secret supplier bans share the same logic: keep the public in the dark. Zimbra's 270 breached servers show what that logic costs.

6 min

Analysis

Fifty-Nine to Forty-Three

In 2022 the Philippine IT and business process management sector, worth roughly 8 percent of national GDP, published a roadmap for 59 billion dollars and 2.5 million jobs by 2028. July's midterm revision reads 43.3 to 50.5 billion dollars and 1.85 to 2.14 million jobs, a range whose floor sits beneath the 1.9 million it employs today. It was reported as artificial intelligence arriving in the world's call centre capital. The association's own arithmetic does not say that: all three scenarios, spanning a 16 billion dollar revenue range and four years, imply revenue per worker within about one percent of the same figure. Both lines were scaled down together, which is the signature of weaker demand rather than of automation, and it is what IBPAP's chief executive said at the time. Why displacement and deferral produce similar labour markets and call for opposite instruments, why every transition programme in existence fires on an event that deferral never produces, and what survives once the framing is stripped out.

5 min