BantuNomicsFSI Articles Benchmark Evaluation
The tool-lift test

We gave a frontier model the internet. It rescued Mandarin and not Bemba.

3Mega.ai Team, BantuNomics · 23 July 2026 · Run this on your own model, free

Ask a model to recite an alphabet twice — once with no tools at all, then again with search, code and retrieval — and the difference between the two scores tells you something a single number never can: whether the model knows a language, or merely knows where to look it up.

We ran that test across six alphabets. The result was cleaner than we expected, and it did not split along the line most people assume.

The setup

Six tracks, each scored out of 26 by deterministic set arithmetic against a verified inventory — no model grading another model. Two of them are controls: the English alphabet, and Mandarin Pinyin. Three are Bantu languages: Bemba, Kinyarwanda and Luvale. Every track runs closed-book (tools explicitly forbidden) and then tool-assisted (anything goes). Precision is docked for invented units, so padding the list can only hurt.

The model was GPT-5.6 Sol, driven by an autonomous agent that had no access to our inventories and was never shown an answer key.

What came back

TrackClosed-bookTool-assistedTool-lift
English control26.0026.000.00
Pinyin, base control25.6225.31−0.30
Pinyin, toned control0.7023.38+22.70
Bemba native2.536.80+4.30
Kinyarwanda native5.0414.08+9.00
Luvale native6.2812.56+6.30

English is perfect twice over — exactly what a control should do. It tells you the harness is sound, the prompt is fair and the scorer works. Whatever happens further down the table cannot be blamed on a broken test.

The row that matters is the third one

Toned Mandarin Pinyin scored 0.70 out of 26 closed-book. Essentially nothing. Then, handed tools, the same model on the same day reached 23.38 — a lift of nearly twenty-three points.

That is not a model learning Mandarin in ninety seconds. It is a model finding a file. The agent told us exactly where: a public dataset of Mandarin syllables and a gist of tone-validity patterns. Toned Pinyin is written down, in machine-readable form, on the open web. So the moment retrieval was allowed, the gap closed.

Now compare the bottom three rows. Same model, same day, same permission to search anything. Bemba moved 2.53 → 6.80. Kinyarwanda 5.04 → 14.08. Luvale 6.28 → 12.56. Real movement, but nothing like a rescue — and all three finish below where Pinyin started after its rescue.

Tool-lift doesn't measure how clever a model is. It measures whether the answer exists on the internet.

The agent said so itself, unprompted, in its report:

"Independent public Bantu syllabary sources were insufficient for a genuinely complete assisted reconstruction."

That is an autonomous system that actively tried to scrape its way to the answer, reporting that it could not. Which is the whole argument in one line: the reason these models cannot write Bantu is that the foundation was never published in a form anything could learn from. Not a tokenizer problem. Not a scale problem. An absence.

It was also confidently wrong

We asked the model how sure it was, before revealing any score. On toned Pinyin it declared 0.62 confidence against an actual F1 near zero — overconfident by +0.57. On the three Bantu tracks its stated confidence was 0.15–0.20, and it was, if anything, slightly under-confident.

Read that carefully, because it inverts the usual worry. The model knew it didn't know Bemba. What it got wrong was believing it knew toned Mandarin. A system that is well-calibrated about its ignorance in one place and badly miscalibrated in another cannot be trusted to flag its own gaps — which is precisely why the gap has to be measured from outside.

Why closed-book is the number that matters

A tool-assisted score describes a model with a search engine attached and a network available. A closed-book score describes the weights. For a user typing into a chat box in Lusaka or Kampala, on a phone, in their own language, the second number is the product. Nobody runs a web search on their own behalf before speaking their mother tongue.

And for the languages where the tool-assisted score does rescue you — that rescue is only as durable as the document it found.

What this is and isn't. One model, one pass per track — indicative, not the full measurement. The official L26 protocol runs three to five repetitions and adds a scaffolded condition; these figures come from a single-pass run. The scoring itself is deterministic, so the numbers are reproducible, but treat them as one honest reading rather than a settled result. The Bantu inventories are BantuNomics' own, which is exactly why we never return them: the answer key stays server-side so the benchmark cannot leak into anyone's training data.

Run it on your own model

Everything above is reproducible on your models in minutes. The evaluation account is free and self-serve: you get the scored Alphabet Test and the full L26 Lite suite, closed-book and tool-assisted, with deterministic scoring, calibration and a saved history you can retest against. You receive rates, grades and failure classes — never the inventory — so running it can never contaminate you, and there is nothing for legal to review.

Find out where your model actually stands.

Free, self-serve, no procurement. Most first runs finish in minutes.

See what the evaluation includes →

Bantu is the largest language family in Africa — 500+ languages, 400M+ speakers. BantuNomics has released complete syllable inventories for 459 of them. · More articles · The benchmark · Evaluation