459 Bantu FSIs.
One phonological infrastructure layer.
BantuNomics has assembled Full Syllable Inventories as the product frontier AI labs need: legal syllable inventories, equation-factored views, and consent-gated audio for Bantu ASR, TTS, translation, and evaluation.
In notation: FSI = NSI ∪ ASI
Same spelling. Different languages. Different sounds.
The row ca ce ci co cu is written the same across Bantu languages — but pronounced completely differently. In Bemba the c is “ch”; in Zulu and Xhosa it is a dental click. Pick an onset and play the rows. Audio is consent-gated.
Loading audio…
Up to 6 languages. Each cell shows that language’s actual surface form for the onset + vowel (some mark long vowels, e.g. Bemba caa). A grey cell means no consented audio yet.
Not a dataset. The missing alphabet layer.
Every FSI is a complete legal-syllable inventory for one Bantu language. The API lets enterprise labs evaluate, license, interact with, and download FSIs as infrastructure.
Bemba first.
Allow-listed frontier labs test the full Bemba FSI, equation view, and capped consented audio before procurement — a controlled proof the model is missing the alphabet layer.
459 released FSIs.
Full Annual Subscriptions license the complete matrix universe: ground-truth-verified syllabary constraint matrices for 459 distinct Bantu languages/dialects.
API access.
Single, regional, group, or all-language selection. Inventory view, equation view, or both. Audio follows consent gates, so labs build without exposing raw identifiers.
Labs learned Bantu as flat text. Speakers learn it as syllables.
Frontier models imitate Bantu sentences, but usually cannot produce the Full Syllable Inventory of the language they imitate. They have seen text but not recovered the sound-unit system that makes it writable, pronounceable, and evaluable.
The alphabet analogy
No lab would trust an English model that couldn’t identify A–Z. For Bantu, the alphabet is syllables, not letters — the complete set every native word is built from.
The native-speaker path
Bantu children drill syllable rows — a e i o u, ba be bi bo bu. They sound, read, and write through syllables. Models skipped that stage.
Tokenizer fracturing
Subword tokenizers split words by character frequency, not phonetic boundaries — tearing prefixes from stems and breaking syllables, hurting morphology and TTS.
First the idea. Then the notation.
The FSI is a structured inventory of two layers: the native syllables a language naturally uses, plus the augmented syllables it needs for borrowed, modern, technical, and foreign forms.
NSI = Native Syllable Inventory. ASI = Augmented Syllable Inventory. Together: FSI = NSI ∪ ASI.
Five ordered slots. Only the vowel is required. No coda.
An FSI envelope is a product object.
Language should never decide who gets to participate in the future.
People should not have to leave their language behind to access knowledge, opportunity, intelligence, or the tools that shape tomorrow.
Access without linguistic surrender.
No one should be locked out of the world’s knowledge because the most powerful technologies don’t speak their language.
Bantu-language AI at family scale.
Curate Bantu family data and build the infrastructure, evaluation, and model-improvement tools that give AI Bantu competence it already shows in English.
No Bantu language left behind. No LLM left behind.
Every community able to access, use, and contribute to knowledge on its own terms — with fluent, culturally grounded AI.
Infrastructure, not data.
An FSI is not a corpus you consume — it is the closed, finite foundation every Bantu word and recording is built from, the thing a model should train on the way it trained on the 26 letters. On the FSIs alone, one annual subscription unlocks competent, standardised AI for ~400 million Bantu speakers — a foundational standard no model can build for itself. And a model cannot self-certify its own FSI: generation commoditises; certification against native ground truth does not.
The FSI is the Bantu alphabet.
For Bantu, the FSI is the alphabet — the actual inventory of sound units the language uses to generate native words. A–Z = ba be bi bo bu
Tone rests on syllables.
The syllable is the tone-bearing unit. Flat text hides tone and vowel length; audio-aligned FSIs create the frame to recover them.
Same spelling, different sounds.
ca ce ci co cu behaves differently across Bemba, Zulu, Xhosa, Swahili — a family map only works if each language keeps its own FSI.
Models need an anchor.
Without the FSI, models guess from surface text and hallucinate inventories. With it, they have a valid-unit set for generation, validation, and evaluation.
A family-level first.
BantuNomics assembles Bantu as one phonological object, not disconnected files — the first family-scale FSI layer built for AI labs.
No model self-certifies its FSI.
A model can generate an inventory; it cannot verify nativeness, tell legal from hallucinated, or produce the consented audio and auditable provenance — without the native-curated ground truth.
Prove it free. Evaluate. Pilot. License the ecosystem.
Four levels, one credential across every product subdomain. Prove the gap on your own model (Public) → score that model against the certified layer, free and self-serve, with saved results your agent can re-run (Evaluation) → open all eight domains on three languages and measure on your held-out data, fully creditable (Validation Program) → license the whole living ecosystem (Full Annual Subscription).
- The Alphabet Test on your own model
- Coverage & provenance counts
- No FSI data · public scorer only
- Scored Alphabet Test + L26 Lite suite
- Rates, L26 score + actionable feedback
- Saved history · retest · team sharing
- 3 languages you choose × all eight domains — syllables, tone, nouns, verbs, numbers, health, grammar, stories
- Consented native audio in both modes — clean Bantu and Bantu–English code-switch
- Full APIs, MCP & exports · curation, provenance & reproducibility moat
- Measured on your held-out data · 100% creditable toward a subscription
- Every product line — full depth, all 459 released languages
- Every artifact + full uncapped native audio + bulk export
- All APIs + MCP · versioned & continuously improving · you steer the roadmap
Evaluation is free and self-serve; Pilot and Full Annual Subscription are provisioned by BantuNomics. FSI is the wedge; a Full Annual Subscription licenses the whole ecosystem.
What a Full Annual Subscription actually includes
One annual enterprise subscription to the entire BantuNomics ecosystem — every product, every artifact, and all the infrastructure, native-curated and versioned. It is not a dataset you license once; it is a living standard you stay current with. New products and languages ship into your subscription as they are released — not upsold — and Full Annual Subscriptions steer which languages get widened and how deep.
Concretely, a Full Annual Subscription grants full depth across all 459 released languages on every product below — every inventory, matrix, equation, engine and consented native-audio corpus, uncapped — plus the APIs, exports and provenance that make it usable and defensible.
Full Syllable Inventories FSI
The operating alphabets of all 459 languages — complete inventories, equation-factored views, and consented native syllable audio, uncapped.
Equations BTS-E100
The formal equation registry governing Bantu structure — syllable, morphology, numeral and syntax construction across the family.
Universal Concord Matrix BTS-UCM100
Each language's entire agreement system across 31 fixed dimensions, certified cell by cell — the ground-truth answer key for grammar-aware generation.
Verbs BTS-V100
The 9-slot generative verb engine (53 verb types), the full 37-chapter course, and the consented code-switch audio corpus.
Nouns ABS-400
The class-driven noun engine + the English-anchored aligned matrix (one concept → each language, singular & plural) + the audio corpus.
Tone & Flat Text A12
The A12 tone codec + the consented homograph corpus and ASR benchmark that prove the tonal meaning flat text drops.
Numbers BTS-S100
The Bantu numeral substrate + calculator studio + two-mode native audio (clean Bantu and the real Bantu→English code-switch).
Body & Health BTS-BH100
The consent-backed clinical-language standard — body parts + health phrases with native audio across 20+ languages.
Stories · Standards · BantuOS ABS-1400
The connected-speech corpus with a labeled code-switch point in every take, the BTS Standards registry, and the architecture the whole ecosystem is built to.
Every artifact, uncapped
Inventories, matrices, equations, generative engines and the full consented native-audio corpora across every product — no per-language caps, no sample limits — with bulk export (json · jsonl · csv · xlsx · zip · docx).
All the infrastructure
Hosted REST APIs + OpenAPI + MCP on every subdomain, one credential across the whole ecosystem, versioned provenance + reproducibility + change history, and org/seat management for your team.
The living program
Continuously improving under a change-management engine — every release more complete than the last, and you stay anchored to the latest. New domains and languages are included as they ship, and you steer the build frontier.
Bring the Bantu FSI layer into your frontier lab.
One enterprise relationship. One all-access bundle: FSIs, equation views, manifests, consent-gated audio, and evaluation support.