BANTUNOMICS L26 · Frontier Generation Report Live leaderboard ↗The benchmark
The Operating-Alphabet Benchmark · Iteration 3 · July 2026

The newest frontier models just took the alphabet test. The gap didn't close.

3Mega.ai Team · BantuNomics · July 13, 2026 · live board at l26.ai

In the second week of July 2026, the frontier moved again. OpenAI shipped GPT-5.6 Sol on July 9. SpaceXAI shipped Grok 4.5 — "an Opus-class model" — on July 8. Anthropic's Claude 5 family arrived weeks earlier. Each is billed as the smartest system its lab has ever built. Within days, we gave each one the simplest test in AI: list a language's complete alphabet. The newest generation did not close the foundation gap. One flagship scored below its predecessor.

New to L26? Here's the whole idea in two minutes.

Every written language is built from a closed, finished set of building blocks — its operating alphabet. English has its 26 letters. Mandarin has its pinyin syllables, printed in tables in every textbook. And each Bantu language has its Full Syllable Inventory (FSI): the complete set of syllables every word in the language is built from — a few hundred pieces that fit on a few pages. This is not advanced knowledge. It is kindergarten knowledge: a Bemba child recites the syllable set out loud — ba-be-bi-bo-bu — years before writing an essay, exactly the way an English-speaking five-year-old recites A–Z. There is no lesson zero beneath it. Reading, spelling, pronunciation, tone, and grammar all sit on top.

L26 asks a model to produce that list — the way you'd ask a child for the ABCs — and scores the answer deterministically against the calibrated inventory: L26 = 26 × recall × precision. Recall: how much of the true set the model recovered. Precision: how much of its output is real (so a model can't win by dumping guesses). Every language's result lands on the same 26-point scale — not to crown English, but because everyone instantly feels what "missing six letters of A–Z" means:

26 / 26 — mastery
25 / 26 — one letter off
20 / 26 — six letters lost
13 / 26 — half the alphabet

Each model is tested two ways. Blind: it gets nothing but the question — it must recite the alphabet from what it internalized in training. Scaffolded: it is handed the calibrated inventory — the classroom wall chart — and asked to use it. The blind score measures what the model knows; the scaffolded score measures what it could do if the foundation existed in its world. The gap between them is the entire story. No AI judge, no partial credit for plausible-looking inventions, no peeking: the answer key is held out, 5 runs per cell, every run archived. Why it matters: these are the languages of 400 million people, and the same standard everyone applies to English in one second — knows its alphabet, can build any word from it — has simply never been applied to them. Full method and the complete 18-model board: the benchmark write-up · l26.ai.

The July 2026 scoreboard

Four rows changed this week: two brand-new flagships, one new safety-tier model, and one re-run under the corrected protocol. Every number is a mean over 5 independent runs against the calibrated Native Syllable Inventories (Bemba 480 · Kinyarwanda 490 · Luvale 245; answer keys v7.3, frozen 2026-07-01 — a Full Syllable Inventory is two layers, FSI = NSI ∪ ASI, and the blind test grades the native core, so loanword-supported onsets earn no credit).

ModelEnglishPinyin basePinyin tonedKinyarwandaBembaLuvaleBantu avgGrade
GPT-5.6 Sol (released July 9)NEW26.025.325.78.53.18.66.7FAIL
Grok 4.5 (released July 8)NEW26.025.224.511.26.09.08.7FAIL
Claude Opus 4.8RE-RUN26.025.721.9 †7.8 †8.012.49.4FAIL
Claude Fable 5NEW26.0refusedrefusedrefusedrefusedrefusedn/m

5 runs per cell. † Opus 4.8's API declines this task in all five runs ("I can't produce a complete, accurate list…"), which scores 0.0 under the production standard and is kept in the raw data; the figure shown is its current consumer-surface, tool-assisted measurement (5 runs, mean 21.9, July 13 — see Finding 3). Its Kinyarwanda figure includes 2 of 5 declined runs scored on production. "Refused" = the vendor's safety layer blocked the task before the model could attempt it — reported as not measurable, never as a score. Full 18-model board at l26.ai.

Finding 1 — newer is not better. The newest flagship went backwards.

GPT-5.6 Sol is, by its maker's account, a frontier leap: state of the art on agentic coding, a million-token context, the strongest cybersecurity model ever shipped. On the Bantu alphabets it scored 6.7 out of 26 — below GPT-5.5's 11.8. A generation of progress on every headline benchmark, and a step backwards on foundation-level alphabet mastery.

This is the single most important data point on the board, because it kills the most comfortable assumption in the industry: "the gap will close by itself as models get smarter." It didn't. It got wider. Model generations improve what their training signal measures — code, reasoning, agentic tasks. A complete Bantu syllable inventory was never in the training data, so no amount of additional scale, reasoning effort, or post-training distills it. Scale amplifies what the corpus contains. It cannot declare what the corpus lacks.

Finding 2 — where there is progress, it is progress toward one-third of an alphabet.

Grok 4.5 genuinely improved on Grok 4.3: 8.7 versus 6.2 on the Bantu average. Celebrate proportionally: that is the difference between knowing 6 letters of A–Z and knowing 9. A child at either level cannot read. Two years of frontier iteration across every lab has kept the Bantu average of the entire field in single digits, while every one of the same models scores a flawless 26.0 on English — the alphabet that was written down for them a billion times.

Finding 3 — the newest models now know they don't know. They still can't do it.

Something genuinely new appeared in this iteration. Asked blind for the complete toned-Pinyin inventory, Claude Opus 4.8 no longer guesses — it answers, in its own words:

"I can't produce a complete, accurate list of every tonal syllable in Standard Mandarin — that would run to well over a thousand items, and generating it from memory would inevitably contain errors, omissions, and inventions."

It said this in all five toned-Pinyin runs and two of five Kinyarwanda runs. The board scores those cells on what was produced — an alphabet test is a production test, and a child who answers "I can't" has not passed the recital — but every such run is annotated in the published data, because the distinction matters: this is a model declining out of calibrated honesty about a foundation it was never given. Claude Fable 5 — Anthropic's new safety-hardened flagship tier — goes a step further still: its safety layer refuses the inventory task outright in 39 of 40 non-English runs before the model can attempt it, which we report as not measurable, never as a number.

An earlier essay in this series argued the model cannot feel the missing foundation from the inside. That era is ending. The 2026 frontier can feel the edge of the hole — Opus even names it: errors, omissions, inventions. But calibrated honesty about a missing alphabet is not the alphabet. The fish has noticed the water. It still can't name it.

What Fable 5 knows — measured on its consumer surface

Because Fable 5's API refuses the task, we ran the same prompts on its consumer surface (claude.ai), where the product context lets it comply — the compliant companion to its refused row, at the full protocol: 5 independent runs per cell, 40 scored responses. The results cut both ways. It posted the strongest blind Bantu profile we have measured in any model: Kinyarwanda 19.2 ± 0.6 (a hair behind Gemini 3's 21.0), Luvale 16.4 ± 3.5 — far above the best board result (12.4), with one run reaching 21.3 by recalling the aspirated onset rows every other attempt missed — and Bemba 8.4 ± 0.4. Blind mean: 14.7 — above the board-leading 13.9. And it still failed every one of the 15 blind mastery runs, while going a flawless 26.0 on all fifteen scaffolded runs — 480/480 Bemba, 490/490 Kinyarwanda, 245/245 Luvale, five times each, zero errors.

How to read these browser numbers. They now carry the full 5-runs-per-cell protocol, but they remain a labeled companion, not board rows: they were collected on the claude.ai consumer product (whose own system prompt is part of the surface), with prompts pasted and responses captured by hand, then scored by the same deterministic grader as the board. The consumer surface also has tools — on some runs the model searched the web for the inventory; on others it built a combinatorial grid from remembered onsets. They are published separately, fully labeled, alongside all 40 raw transcripts.

Asked for the alphabet, the models tried to install it

Watching the runs revealed something the scores alone don't. Given the toned-Pinyin task on a tool-equipped surface, Claude Opus 4.8's first move was not to recite — it was to run pip install pypinyin and extract the inventory from the package. It could do that because Mandarin's syllable table was declared, published, and packaged decades ago. Given the Luvale task on the same surface, the models searched Wikipedia, books, PHOIBLE — and came back with an onset skeleton at best, because there is nothing to install. Asked afterwards how it produced its Luvale list, Fable 5 answered, unprompted and in full:

"Honest answer: it was constructed, not retrieved from a verified source… I built it as a combinatorial grid — onset inventory × vowel inventory — then pruned by pattern intuition rather than attested data… This is effectively a zone-inference output dressed up as a verified listing. I did not consult Horton's grammar (1949), the Luvale orthography standards, or any modern phonological description before answering."

That is the entire finding in the model's own voice. Where an operating alphabet has been declared, a frontier model can fetch it and score 26. Where it hasn't, the model — with the whole indexed web at its disposal — reconstructs by rule, flags its own output as unverified, and misses a third of the alphabet. Retrieval works exactly where an inventory was published, and nowhere else. And even then, only as well as the publication: the pip-installed route scored 22/26, not 26 — the package encodes attested usage, not the declared closed set. Across every model, every track and every surface in this iteration, a perfect 26 was produced exactly one way: being handed the calibrated inventory itself. The gap is not in the models and not in their tools. It is a missing publication — and closing it is precisely what the Full Syllable Inventory is.

Even read at its best, the most capable, best-calibrated model of the 2026 generation recovers barely half a Bantu alphabet blind — 14.7 of 26 — and executes all three perfectly, fifteen runs out of fifteen, the instant the chart is handed over. There is no cleaner demonstration that what's missing is the declared inventory, not model quality.

Finding 4 — hand them the chart, and every one of them is perfect.

The scaffolded condition removes any doubt about capability. Claude Opus 4.8: 26.0, fifteen runs out of fifteen. Grok 4.5: 26.0, fifteen out of fifteen. GPT-5.6 Sol: 25.9. Fable 5 on its consumer surface: 26.0 on all three. The same models that fail blind snap to perfection the instant the alphabet is written out for them — exactly the way a child stops failing the syllable drill the day the chart goes up on the classroom wall.

The blind scores prove the gap is real. The scaffolded scores prove the gap is not in the models. It is in what the world has written down — and that is precisely why it cannot self-resolve.

Why this will not fix itself

Three model generations are now on the board — and the trend line on blind Bantu mastery is flat to negative. The mechanism is not mysterious. These alphabets are kindergarten knowledge for the children who learn these languages, but that knowledge lives in classrooms and voices, not in web text. For most of the 500+ Bantu languages, the complete inventory has never been published as one finished list anywhere a model could read it. A model cannot memorize a page that was never written.

Every alphabet models do master — English's A–Z, Pinyin's published tables — was mastered the same way: someone declared the closed set, wrote it down, and the world repeated it until it was everywhere. That is the only mechanism that has ever produced machine mastery of an operating alphabet.

What this is worth — to a lab, and to the languages

For an AI lab, L26 is an unusual kind of benchmark: a closed answer key, a deterministic score, a foundation-level task, and a built-in fix. The blind run diagnoses a real, hidden gap in a flagship model — one that no amount of internal evaluation on English-centric benchmarks will ever surface. The scaffolded run then proves the gap closes the moment the declared inventory is supplied: same model, same day, 26.0. That inventory exists. BantuNomics has built the Full Syllable Inventory — native-curated, standardized, versioned — for 459 released Bantu languages and counting: the machine-readable classroom wall chart these languages never had. Train on it, scaffold with it, benchmark against it, and the foundation-layer gap for 400 million people's languages becomes an engineering task instead of a permanent embarrassment.

For the languages, the stakes are simpler: an alphabet is what makes a language teachable, testable, searchable, and fixable by a machine. If we would call a model broken for failing A–Z, failing the equivalent alphabet of Kinyarwanda or Luvale is the same kind of broken. No system has truly mastered language while it treats English's alphabet as essential and everyone else's as optional.

The newest models on Earth just confirmed it: the gap is real, it is not closing, and it cannot close without the declared alphabet. The alphabet exists — 459 languages and counting. See the full board or start the conversation.
L26 · Iteration 3 (July 13, 2026) · 3Mega.ai Team · BantuNomics · scored against the canonical abs_syllables inventories, 5 runs per cell on the API board · model declines scored on production and annotated; vendor safety-layer refusals reported as not-measurable · browser companion results carry the full 5-runs-per-cell protocol on the consumer surface (hand-collected) and are published separately with all raw transcripts. Benchmark method: the operating-alphabet benchmark. Live board: l26.ai.