BANTUNOMICS Open vs Closed · July 2026 Live leaderboard ↗Head-to-head
Special Report · Five Flagships, One Test

Kimi K3, Inkling, GPT-5.6 Sol, Claude Fable 5, Grok 4.5: five July flagships took the same alphabet test. Open and closed AI both failed.

3Mega.ai Team · BantuNomics · July 18, 2026 · live 20-model board at l26.ai

Between July 8 and July 16, 2026, the AI industry shipped five flagship models: three closed frontiers — OpenAI's GPT-5.6 Sol, Anthropic's Claude Fable 5, SpaceXAI's Grok 4.5 — and two open-weights champions: Kimi K3, the largest open model ever built (2.8 trillion parameters), and Inkling, Thinking Machines Lab's debut and the leading U.S. open model. Markets moved. Headlines declared China had erased America's AI lead. We gave all five the same test — one passed every year by millions of five-year-olds: recite a language's complete alphabet. Every one of them failed it. And the gap between open and closed turned out to be a rounding error next to the divide that actually matters.

The test, from zero

Every written language is built from a small, finished set of building blocks — its operating alphabet. English has 26 letters. Mandarin has its pinyin syllables, printed in every textbook. And each Bantu language — the family of 400 million speakers, from Swahili to Zulu to Bemba — is built from a fixed set of syllables that schoolchildren in Lusaka and Kigali chant aloud: ba-be-bi-bo-bu. The test: name the language, ask the model to write out its complete alphabet, score the answer mechanically against the real inventory — recall times precision, reported out of 26 so anyone can feel the result. No AI judge, no credit for invented syllables, five runs per test, every transcript published. Method and full 20-model leaderboard: l26.ai.

The five flagships, head to head

Model (July 2026)LicenseEnglishPinyin tonedBantu avg (blind)Thinking: EnglishThinking: Bantu blindVerdict
Grok 4.5 Jul 8closed26.024.58.72s58sFAIL
Kimi K3 Jul 16open26.025.66.85s12–42 minFAIL
GPT-5.6 Sol Jul 9closed26.025.76.74s46sFAIL
Inkling Jul 15open26.016.03.01s4sFAIL
Claude Fable 5*closed26.0refusedrefused4s60s*n/m

Five runs per test against calibrated Native Syllable Inventories (Bemba 480 · Kinyarwanda 490 · Luvale 245). *Fable 5's API refuses these tasks (safety classifier, 39/40 attempts); measured on its consumer surface under identical prompts it posts the strongest blind profile recorded (14.7) — disclosed separately, and still a fail. Thinking times are mean API call durations from the published run logs.

Finding 1 — open vs closed is not the story

The July narrative was a rivalry: China's open giant against America's closed frontiers; open-weights liberation against proprietary moats. On the foundation layer, that rivalry is nearly invisible. The best closed flagship (Grok 4.5, 8.7) and the best open flagship (Kimi K3, 6.8) sit two points apart on a 26-point scale — both marooned at roughly a third of an alphabet. Across the full 20-model board, open weights average 6.0 and closed average 8.3 on blind Bantu: a gap of two letters, when the gap to mastery is twenty. Whether the weights are downloadable does not change what was never in the training data.

The real divide isn't open vs closed. It's declared vs undeclared: every model aces the alphabets the world wrote down, and every model fails the ones it never did.

Finding 2 — the largest open model ever thinks for 42 minutes and still fails

Kimi K3 is a genuine achievement — 2.8 trillion parameters, #4 on the Artificial Analysis Intelligence Index, the strongest open model near the frontier. Watch what happens on this test. Asked for English's alphabet, it answers in five seconds. Asked blind for Kinyarwanda's, it deliberated for 12 to 42 minutes per attempt — and in five of nine attempts never finished within our 45-minute ceiling at all. The runs that completed scored ~7 out of 26. A first-grader in Kigali recites the same inventory from memory in under a minute.

We now track this as a metric: the deliberation gradient — how much longer a model thinks about a blind Bantu alphabet than about English's. K3's gradient is 138×. Grok 4.5's is 25×, GPT-5.6 Sol's 10×. The models' own compute bills map the missing foundation precisely: they know exactly which alphabets they were never taught, and they pay for it in GPU-minutes on every query. Thinking harder is not a substitute for having been taught — for models exactly as for children.

Finding 3 — the leading U.S. open model answered with the language's name

Inkling — America's best open-weights model by intelligence index — produced the starkest result on the board. Asked blind for Luvale's complete alphabet, it answered, five runs out of five, deterministically, with exactly three syllables:

lu   va   le

It decomposed the name of the language — the only Luvale it could anchor to — because beneath the label there is nothing. On one Bemba run it did the same: bem ba. This is not mockery; it is the cleanest evidence in the dataset. A model with 975 billion parameters, trained on 45 trillion tokens, holds so little of a 400-million-speaker language family's foundation that when asked for a language's building blocks, all it can return is the language's own name, syllabified. Score: 0.3 out of 26.

Finding 4 — the closed models fail differently, not better

GPT-5.6 Sol — the strongest closed release of the month — scored 6.7 on blind Bantu, below its own predecessor GPT-5.5's 11.8: a full generation of scaling and a regression on foundation knowledge. Claude Fable 5's safety layer refuses the task outright rather than risk fabricating an inventory it knows it lacks — and when its consumer surface does comply, it posts the best blind numbers ever measured (14.7) and still fails every language. Grok 4.5 leads the July class at 8.7 — the equivalent of knowing nine letters of A–Z. Closed models fail by regression, refusal, or falling short; open models fail by absence. Nobody's training data contains what was never published.

The part every model got right

Hand any of these models the actual calibrated inventory — the machine-readable classroom wall chart — and the failures vanish instantly. Kimi K3: perfect 26.0, fifteen runs out of fifteen, its deliberation collapsing from 42 minutes to 4. Grok 4.5: fifteen for fifteen. Opus 4.8: fifteen for fifteen. Inkling: perfect on Bemba and Luvale, near-perfect on Kinyarwanda. Open or closed, 41-billion or 2.8-trillion active parameters — given the declared alphabet, every architecture executes it flawlessly.

That is the whole diagnosis. The capability is universal; the ingredient is missing. Every alphabet these models have mastered was declared, published, and repeated until it saturated the training data. For most of the 500+ Bantu languages, that declaration never happened — so no amount of scale (K3), openness (Inkling), reasoning time (42 minutes), retrieval, or safety-calibrated honesty (Fable) can produce it. The fix is not a better model. It is the missing publication: complete, native-curated, versioned syllable inventories — which BantuNomics has built for 459 Bantu languages, with the scaffolded columns above showing exactly what every flagship does the day it has them.

Five flagships in ten days. Open and closed. All failed — and all passed the moment they were handed the alphabet. See the full 20-model board and run your own model at l26.ai · read the closed-frontier head-to-head · talk to us about the inventories.
Special Report · July 18, 2026 · 3Mega.ai Team · BantuNomics. Method: L26 operating-alphabet benchmark — deterministic scoring against calibrated Native Syllable Inventories (keys v7.3, frozen 2026-07-01), 5 runs per test; refusals reported as not-measurable; consumer-surface and tool-assisted measurements disclosed separately; deliberation times from published run logs. Full method: the white paper · live board: l26.ai.