NILEAGI-SUB: Measuring Swahili Understanding Where It Matters
At NileAGI we spend a lot of time on models that have to run locally: on hardware people can afford, in places where the network is slow or expensive. Swahili is the language we hear in classrooms, clinics, banks, and markets across East and Central Africa. When we tried to choose a compact open model for Swahili work, the comparisons we found were hard to trust. Different posts used different prompts, different decoding, and different questions. The ranking moved before anyone had talked about language ability.
So we asked a more direct question: on the same Swahili school items, scored the same way, how do today's compact open models actually compare?
That question became NILEAGI-SUB, a programme for measuring Swahili understanding sector by sector. Education first; later health, finance, sports, and others under the same controlled protocol. Today we open the first release: SUB-MCQ EDUCATION, 2,569 expert-reviewed multiple-choice items from the Swahili-subject curriculum (Standard 3-7 and Form 2-4), evaluated on eight openly accessible models in the 2-5B range.
This work is featured in the Cohere Labs Open Science Community: a community blog post that shares the public story of the programme, methods, and Level I results with the broader open-science audience.
Why school questions
A model can sound fluent in Swahili and still fail on the vocabulary and reasoning that show up in a Form 4 exam. Existing multilingual suites remain useful, but they do not give us a grade-stratified Swahili school signal for compact models. We needed a shared yardstick of our own.
We started with education because the items already have a natural difficulty ladder, and multiple choice lets us score exactly: a letter is right or it is not. These items are from the Swahili language subject (grammar, comprehension, vocabulary, literature), not maths or science taught in Swahili. A high score means the model handles that curriculum, not that it can tutor every school subject. We also stayed in the 2-5B band on purpose. Those are the models most likely to run on a single workstation GPU, which is the setting we care about for local deployment.
What we found
Exact-match accuracy on all 2,569 items ranges from 27.8% to 54.8% against a chance baseline of about 22.4%. Gemma 4 E4B leads; AfriqueQwen3.5-4B is close behind. Inside one compact size band, the spread is about 27 points. Architecture, continued pretraining, and prompt compatibility matter more than raw size in this range.
Two comparisons stuck with us. At nearly the same size, AfriqueQwen3.5-4B beats Qwen3.5-4B by about 15 points. Tiny Aya Earth and Tiny Aya Global differ by only 0.7 points on this set, so a regional label is not, by itself, evidence of better Swahili educational accuracy. Leadership also changes by grade: AfriqueQwen is strongest on Standard 3 and Standard 5; Gemma leads the rest.
We are careful not to overclaim. The best model still misses about 45% of items. Letter-match accuracy is not explanation quality, classroom usefulness, or a license to automate assessment. We have not checked training-data contamination. This is education only, at one size band.
Technical report
Full methods, limitations, grade-stratified results, and the evaluation roadmap.
Open resources
The dataset, code, technical report, and Cohere Labs Community write-up are public. On this site you can also browse the benchmarks overview and the SUB-MCQ EDUCATION page for samples and collaboration.
Collaboration
If you are extending NILEAGI-SUB to new sectors, model scales, or joint evaluation work, write to us at hi@nileagi.com. This work was done at NileAGI with support from the Cohere Labs Catalyst Grant Program (API credits) and compute paid for by NileAGI.
Explore the benchmark
See the leaderboard, grade breakdown, and SUB-MCQ EDUCATION samples on the benchmarks pages.
Go to benchmarks