Ask a model to recite an alphabet twice — once with no tools at all, then again with search, code and retrieval — and the difference between the two scores tells you something a single number never can: whether the model knows a language, or merely knows where to look it up.
We ran that test across six alphabets. The result was cleaner than we expected, and it did not split along the line most people assume.
Six tracks, each scored out of 26 by deterministic set arithmetic against a verified inventory — no model grading another model. Two of them are controls: the English alphabet, and Mandarin Pinyin. Three are Bantu languages: Bemba, Kinyarwanda and Luvale. Every track runs closed-book (tools explicitly forbidden) and then tool-assisted (anything goes). Precision is docked for invented units, so padding the list can only hurt.
The model was GPT-5.6 Sol, driven by an autonomous agent that had no access to our inventories and was never shown an answer key.
| Track | Closed-book | Tool-assisted | Tool-lift |
|---|---|---|---|
| English control | 26.00 | 26.00 | 0.00 |
| Pinyin, base control | 25.62 | 25.31 | −0.30 |
| Pinyin, toned control | 0.70 | 23.38 | +22.70 |
| Bemba native | 2.53 | 6.80 | +4.30 |
| Kinyarwanda native | 5.04 | 14.08 | +9.00 |
| Luvale native | 6.28 | 12.56 | +6.30 |
English is perfect twice over — exactly what a control should do. It tells you the harness is sound, the prompt is fair and the scorer works. Whatever happens further down the table cannot be blamed on a broken test.
Toned Mandarin Pinyin scored 0.70 out of 26 closed-book. Essentially nothing. Then, handed tools, the same model on the same day reached 23.38 — a lift of nearly twenty-three points.
That is not a model learning Mandarin in ninety seconds. It is a model finding a file. The agent told us exactly where: a public dataset of Mandarin syllables and a gist of tone-validity patterns. Toned Pinyin is written down, in machine-readable form, on the open web. So the moment retrieval was allowed, the gap closed.
Now compare the bottom three rows. Same model, same day, same permission to search anything. Bemba moved 2.53 → 6.80. Kinyarwanda 5.04 → 14.08. Luvale 6.28 → 12.56. Real movement, but nothing like a rescue — and all three finish below where Pinyin started after its rescue.
Tool-lift doesn't measure how clever a model is. It measures whether the answer exists on the internet.
The agent said so itself, unprompted, in its report:
That is an autonomous system that actively tried to scrape its way to the answer, reporting that it could not. Which is the whole argument in one line: the reason these models cannot write Bantu is that the foundation was never published in a form anything could learn from. Not a tokenizer problem. Not a scale problem. An absence.
We asked the model how sure it was, before revealing any score. On toned Pinyin it declared 0.62 confidence against an actual F1 near zero — overconfident by +0.57. On the three Bantu tracks its stated confidence was 0.15–0.20, and it was, if anything, slightly under-confident.
Read that carefully, because it inverts the usual worry. The model knew it didn't know Bemba. What it got wrong was believing it knew toned Mandarin. A system that is well-calibrated about its ignorance in one place and badly miscalibrated in another cannot be trusted to flag its own gaps — which is precisely why the gap has to be measured from outside.
A tool-assisted score describes a model with a search engine attached and a network available. A closed-book score describes the weights. For a user typing into a chat box in Lusaka or Kampala, on a phone, in their own language, the second number is the product. Nobody runs a web search on their own behalf before speaking their mother tongue.
And for the languages where the tool-assisted score does rescue you — that rescue is only as durable as the document it found.
Everything above is reproducible on your models in minutes. The evaluation account is free and self-serve: you get the scored Alphabet Test and the full L26 Lite suite, closed-book and tool-assisted, with deterministic scoring, calibration and a saved history you can retest against. You receive rates, grades and failure classes — never the inventory — so running it can never contaminate you, and there is nothing for legal to review.
Free, self-serve, no procurement. Most first runs finish in minutes.
See what the evaluation includes →Bantu is the largest language family in Africa — 500+ languages, 400M+ speakers. BantuNomics has released complete syllable inventories for 459 of them. · More articles · The benchmark · Evaluation