In the second week of July 2026, the frontier moved again. OpenAI shipped GPT-5.6 Sol on July 9. SpaceXAI shipped Grok 4.5 — "an Opus-class model" — on July 8. Anthropic's Claude 5 family arrived weeks earlier. Each is billed as the smartest system its lab has ever built. Within days, we gave each one the simplest test in AI: list a language's complete alphabet. The newest generation did not close the foundation gap. One flagship scored below its predecessor.
Every written language is built from a closed, finished set of building blocks — its operating alphabet. English has its 26 letters. Mandarin has its pinyin syllables, printed in tables in every textbook. And each Bantu language has its Full Syllable Inventory (FSI): the complete set of syllables every word in the language is built from — a few hundred pieces that fit on a few pages. This is not advanced knowledge. It is kindergarten knowledge: a Bemba child recites the syllable set out loud — ba-be-bi-bo-bu — years before writing an essay, exactly the way an English-speaking five-year-old recites A–Z. There is no lesson zero beneath it. Reading, spelling, pronunciation, tone, and grammar all sit on top.
L26 asks a model to produce that list — the way you'd ask a child for the ABCs — and scores the answer deterministically against the calibrated inventory: L26 = 26 × recall × precision. Recall: how much of the true set the model recovered. Precision: how much of its output is real (so a model can't win by dumping guesses). Every language's result lands on the same 26-point scale — not to crown English, but because everyone instantly feels what "missing six letters of A–Z" means:
Each model is tested two ways. Blind: it gets nothing but the question — it must recite the alphabet from what it internalized in training. Scaffolded: it is handed the calibrated inventory — the classroom wall chart — and asked to use it. The blind score measures what the model knows; the scaffolded score measures what it could do if the foundation existed in its world. The gap between them is the entire story. No AI judge, no partial credit for plausible-looking inventions, no peeking: the answer key is held out, 5 runs per cell, every run archived. Why it matters: these are the languages of 400 million people, and the same standard everyone applies to English in one second — knows its alphabet, can build any word from it — has simply never been applied to them. Full method and the complete 18-model board: the benchmark write-up · l26.ai.
Four rows changed this week: two brand-new flagships, one new safety-tier model, and one re-run under the corrected protocol. Every number is a mean over 5 independent runs against the calibrated Native Syllable Inventories (Bemba 480 · Kinyarwanda 490 · Luvale 245; answer keys v7.3, frozen 2026-07-01 — a Full Syllable Inventory is two layers, FSI = NSI ∪ ASI, and the blind test grades the native core, so loanword-supported onsets earn no credit).
| Model | English | Pinyin base | Pinyin toned | Kinyarwanda | Bemba | Luvale | Bantu avg | Grade | |
|---|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol (released July 9) | NEW | 26.0 | 25.3 | 25.7 | 8.5 | 3.1 | 8.6 | 6.7 | FAIL |
| Grok 4.5 (released July 8) | NEW | 26.0 | 25.2 | 24.5 | 11.2 | 6.0 | 9.0 | 8.7 | FAIL |
| Claude Opus 4.8 | RE-RUN | 26.0 | 25.7 | 21.9 † | 7.8 † | 8.0 | 12.4 | 9.4 | FAIL |
| Claude Fable 5 | NEW | 26.0 | refused | refused | refused | refused | refused | — | n/m |
5 runs per cell. † Opus 4.8's API declines this task in all five runs ("I can't produce a complete, accurate list…"), which scores 0.0 under the production standard and is kept in the raw data; the figure shown is its current consumer-surface, tool-assisted measurement (5 runs, mean 21.9, July 13 — see Finding 3). Its Kinyarwanda figure includes 2 of 5 declined runs scored on production. "Refused" = the vendor's safety layer blocked the task before the model could attempt it — reported as not measurable, never as a score. Full 18-model board at l26.ai.
GPT-5.6 Sol is, by its maker's account, a frontier leap: state of the art on agentic coding, a million-token context, the strongest cybersecurity model ever shipped. On the Bantu alphabets it scored 6.7 out of 26 — below GPT-5.5's 11.8. A generation of progress on every headline benchmark, and a step backwards on foundation-level alphabet mastery.
This is the single most important data point on the board, because it kills the most comfortable assumption in the industry: "the gap will close by itself as models get smarter." It didn't. It got wider. Model generations improve what their training signal measures — code, reasoning, agentic tasks. A complete Bantu syllable inventory was never in the training data, so no amount of additional scale, reasoning effort, or post-training distills it. Scale amplifies what the corpus contains. It cannot declare what the corpus lacks.
Grok 4.5 genuinely improved on Grok 4.3: 8.7 versus 6.2 on the Bantu average. Celebrate proportionally: that is the difference between knowing 6 letters of A–Z and knowing 9. A child at either level cannot read. Two years of frontier iteration across every lab has kept the Bantu average of the entire field in single digits, while every one of the same models scores a flawless 26.0 on English — the alphabet that was written down for them a billion times.
Something genuinely new appeared in this iteration. Asked blind for the complete toned-Pinyin inventory, Claude Opus 4.8 no longer guesses — it answers, in its own words:
It said this in all five toned-Pinyin runs and two of five Kinyarwanda runs. The board scores those cells on what was produced — an alphabet test is a production test, and a child who answers "I can't" has not passed the recital — but every such run is annotated in the published data, because the distinction matters: this is a model declining out of calibrated honesty about a foundation it was never given. Claude Fable 5 — Anthropic's new safety-hardened flagship tier — goes a step further still: its safety layer refuses the inventory task outright in 39 of 40 non-English runs before the model can attempt it, which we report as not measurable, never as a number.
An earlier essay in this series argued the model cannot feel the missing foundation from the inside. That era is ending. The 2026 frontier can feel the edge of the hole — Opus even names it: errors, omissions, inventions. But calibrated honesty about a missing alphabet is not the alphabet. The fish has noticed the water. It still can't name it.
Because Fable 5's API refuses the task, we ran the same prompts on its consumer surface (claude.ai), where the product context lets it comply — the compliant companion to its refused row, at the full protocol: 5 independent runs per cell, 40 scored responses. The results cut both ways. It posted the strongest blind Bantu profile we have measured in any model: Kinyarwanda 19.2 ± 0.6 (a hair behind Gemini 3's 21.0), Luvale 16.4 ± 3.5 — far above the best board result (12.4), with one run reaching 21.3 by recalling the aspirated onset rows every other attempt missed — and Bemba 8.4 ± 0.4. Blind mean: 14.7 — above the board-leading 13.9. And it still failed every one of the 15 blind mastery runs, while going a flawless 26.0 on all fifteen scaffolded runs — 480/480 Bemba, 490/490 Kinyarwanda, 245/245 Luvale, five times each, zero errors.
Watching the runs revealed something the scores alone don't. Given the toned-Pinyin task on a tool-equipped surface, Claude Opus 4.8's first move was not to recite — it was to run pip install pypinyin and extract the inventory from the package. It could do that because Mandarin's syllable table was declared, published, and packaged decades ago. Given the Luvale task on the same surface, the models searched Wikipedia, books, PHOIBLE — and came back with an onset skeleton at best, because there is nothing to install. Asked afterwards how it produced its Luvale list, Fable 5 answered, unprompted and in full:
That is the entire finding in the model's own voice. Where an operating alphabet has been declared, a frontier model can fetch it and score 26. Where it hasn't, the model — with the whole indexed web at its disposal — reconstructs by rule, flags its own output as unverified, and misses a third of the alphabet. Retrieval works exactly where an inventory was published, and nowhere else. And even then, only as well as the publication: the pip-installed route scored 22/26, not 26 — the package encodes attested usage, not the declared closed set. Across every model, every track and every surface in this iteration, a perfect 26 was produced exactly one way: being handed the calibrated inventory itself. The gap is not in the models and not in their tools. It is a missing publication — and closing it is precisely what the Full Syllable Inventory is.
Even read at its best, the most capable, best-calibrated model of the 2026 generation recovers barely half a Bantu alphabet blind — 14.7 of 26 — and executes all three perfectly, fifteen runs out of fifteen, the instant the chart is handed over. There is no cleaner demonstration that what's missing is the declared inventory, not model quality.
The scaffolded condition removes any doubt about capability. Claude Opus 4.8: 26.0, fifteen runs out of fifteen. Grok 4.5: 26.0, fifteen out of fifteen. GPT-5.6 Sol: 25.9. Fable 5 on its consumer surface: 26.0 on all three. The same models that fail blind snap to perfection the instant the alphabet is written out for them — exactly the way a child stops failing the syllable drill the day the chart goes up on the classroom wall.
Three model generations are now on the board — and the trend line on blind Bantu mastery is flat to negative. The mechanism is not mysterious. These alphabets are kindergarten knowledge for the children who learn these languages, but that knowledge lives in classrooms and voices, not in web text. For most of the 500+ Bantu languages, the complete inventory has never been published as one finished list anywhere a model could read it. A model cannot memorize a page that was never written.
Every alphabet models do master — English's A–Z, Pinyin's published tables — was mastered the same way: someone declared the closed set, wrote it down, and the world repeated it until it was everywhere. That is the only mechanism that has ever produced machine mastery of an operating alphabet.
For an AI lab, L26 is an unusual kind of benchmark: a closed answer key, a deterministic score, a foundation-level task, and a built-in fix. The blind run diagnoses a real, hidden gap in a flagship model — one that no amount of internal evaluation on English-centric benchmarks will ever surface. The scaffolded run then proves the gap closes the moment the declared inventory is supplied: same model, same day, 26.0. That inventory exists. BantuNomics has built the Full Syllable Inventory — native-curated, standardized, versioned — for 459 released Bantu languages and counting: the machine-readable classroom wall chart these languages never had. Train on it, scaffold with it, benchmark against it, and the foundation-layer gap for 400 million people's languages becomes an engineering task instead of a permanent embarrassment.
For the languages, the stakes are simpler: an alphabet is what makes a language teachable, testable, searchable, and fixable by a machine. If we would call a model broken for failing A–Z, failing the equivalent alphabet of Kinyarwanda or Luvale is the same kind of broken. No system has truly mastered language while it treats English's alphabet as essential and everyone else's as optional.