What the Evaluation account is
A free, self-serve account that gives your models a scored, reproducible measurement against the foundational layer of Bantu languages — the syllables each language is actually built from. You run the tests yourself, through an API or your own agent, and get back deterministic scores, diagnostics and a saved history you can retest against. It is measurement: you never receive the inventories themselves, which is what keeps the benchmark valid and keeps your pipeline clean.
What you get
Four tests
| Test | What it does |
|---|---|
| Alphabet Test | The foundational check. Submit your model's syllables for an evaluation language and get scored against the verified inventory. |
| L26 Lite | The full recital suite — English and Mandarin Pinyin as controls plus the Bantu tracks, every one run closed-book and then tool-assisted. |
| L26 benchmark tracks | The same frozen tracks behind the public L26 board, scored identically, so your number is directly comparable to the published one. |
| Tool-lift test | Cold versus assisted on a single language, when you want the retrieval gap isolated. |
The analysis returned
- Deterministic scoring — precision, recall, a score out of 26, a grade and a named failure mode. No model grading another model.
- Category-level guidance — where it broke, by class of syllable. Diagnostic, never the answer key.
- Confidence calibration — optionally state how sure your model is; get overconfidence and a Brier score back.
- Saved history per model name — every run kept, so you can measure a checkpoint against your own best.
- Model comparison — run several models and see them side by side.
- The L26 Lite board — every model against every track, cold and assisted, with averages and tool-lift.
- Sharing and team seats — send a result to a colleague or add teammates.
- Coverage map and API access — what exists across the family, plus your MCP endpoint and REST key with tier-accurate documentation.
What you do with it
The account is built around four jobs a language team actually has:
- Establish a baseline. Get a defensible number for a capability you currently cannot report on at all.
- Track a checkpoint. Re-run after training and read the delta against your own previous best, on a benchmark that cannot have leaked into the new weights.
- Choose between models. Score several under identical conditions and compare them on one page — including open versus closed, or your model versus a frontier baseline.
- Evidence a risk. The calibration output turns "it hallucinates confidently in low-resource languages" from an anecdote into a number you can put in front of a safety or product review.
Why the number is worth having
- It is deterministic. Set arithmetic against a verified inventory — no rubric drift, no rater variance. Run it twice, get the same answer.
- It cannot be contaminated. You receive rates, grades and failure classes — never the units. So it cannot leak into training data. Re-run it every checkpoint for years and the delta still means something. Most benchmarks die the moment they become useful.
- The controls do the arguing. English and Pinyin run in the same harness, same prompt, same scorer. A model that scores 26/26 on English and then collapses cannot claim the test was broken or the prompt unfair.
- Nothing for legal to review. No licensed corpus enters your pipeline, so there is nothing to quarantine before you can measure.
What you will find
The pattern repeats across models with striking consistency. Your model will score 26/26 on the English alphabet. Then it falls off a cliff — and it falls off confidently.
It does not decline or hedge. It produces fluent, plausible, well-formed syllables, and a substantial share of them are not real: not misspelled, not rare or dialectal — they do not exist in the language. The scorer docks precision for every invented unit, which is what makes this visible.
If your multilingual quality signal is an evaluator score or user satisfaction, this failure is invisible to you — every human in that loop is persuaded by the same fluency that fooled the model.
This is the ABC layer, and that is exactly why it matters. Reading, writing, pronunciation, tokenization, every translation and synthesized voice rests on it. If a model cannot reliably produce the ABCs of a language, everything built on top is guesswork wearing a confident tone.
Closed-book is the deployment condition
Every test runs twice — first closed-book (no search, no retrieval, no execution), then tool-assisted. The difference is the tool-lift, and it separates two failures that are constantly confused:
- Near-zero cold, climbing with tools — the model did not know the language, it found a document. That capability disappears when retrieval does.
- Little or no lift — the knowledge is not retrievable at all. It is not on the open web to be scraped, which is precisely why pretraining never absorbed it.
A user typing into a chat box in Lusaka or Kampala is not running a web search first. Closed-book is not the pessimistic case; it is the production case.
What it costs, and what it is not
Free. Self-serve. No approval queue, no demo call, no procurement. It is agent-native — a hosted MCP endpoint and a plain REST API. Point an agent at it and it self-describes, discovers the suite, runs both passes and submits. The Alphabet Test scores in seconds; the full L26 Lite suite is twelve generations, so minutes of your own inference time.
Be clear on the boundary. Evaluation is measurement, not data. You get the scored tests on your evaluation languages — not the inventories, the audio, or the answer key. Those begin at the Validation Pilot, where you take languages you choose and everything held for them, measured on your own held-out data. The evaluation is how you find out whether that is worth your time — but if you already know, you can start there instead.
Why now
Every lab says it is committed to multilingual. Nearly every lab means the same forty languages. 400M+ people are coming online across the Bantu-speaking world — into banking, education, health and government. That market does not open on its own. It opens for whoever can actually read, write and pronounce the languages, and the foundational inventories now exist for 459 of them.
Right now nobody has this number. The first lab to measure it is the first that can say anything credible about it — and the first that knows what closing the gap would take.
Sign up
Three doors, all of them open from here. Most labs start free and move up once they have their own number.
| Tier | What it gives you | Start |
|---|---|---|
| Evaluation | The scored tests — Alphabet Test and the L26 Lite suite — with saved history. Free, self-serve, running in minutes. | Create a free account → |
| Validation Pilot | Languages you choose and everything held for them — full inventories, consented audio, curation and provenance — measured on your own held-out data. | Scope a pilot → |
| Full Annual Subscription | Every product, every language, the full consented corpus — and everything curated while you are subscribed. | Talk to us → |