Benchmarks
Measuring how open language models understand Swahili in real sectors, starting with education.
NILEAGI-SUB: Swahili Language Understanding
NILEAGI-SUB is an evaluation programme for Swahili understanding across real-world sectors (education, health, finance, sports), one sector at a time under the same controlled protocol. This first report covers education and releases SUB-MCQ EDUCATION: 2,569 expert-reviewed Swahili multiple-choice items from the Swahili-subject curriculum (Standard 3-7 and Form 2-4).
2,569
Validated items
8
Grade bands
22.4%
Chance baseline
Items by grade
| Grade band | Items | Share |
|---|---|---|
| Standard 3 | 66 | 2.6% |
| Standard 4 | 116 | 4.5% |
| Standard 5 | 274 | 10.7% |
| Standard 6 | 651 | 25.3% |
| Standard 7 | 517 | 20.1% |
| Form 2 | 399 | 15.5% |
| Form 3 | 125 | 4.9% |
| Form 4 | 421 | 16.4% |
| Total | 2,569 | 100% |
Level I results (~2-5B open models)
Eight open or openly accessible models evaluated on identical items with deterministic decoding and unified exact-match scoring. Accuracy spans about 27 points at a shared parameter budget: architecture, continued pretraining, and prompt compatibility matter more than raw size in this range.
| # | Model | Params | Accuracy |
|---|---|---|---|
| 1 | Gemma 4 E4B | 4.5B | 54.8% |
| 2 | AfriqueQwen3.5-4B | 4.0B | 50.3% |
| 3 | AfriqueGemma-4B | 4.0B | 43.4% |
| 4 | Tiny Aya Earth | 3.35B | 41.4% |
| 5 | Tiny Aya Global | 3.35B | 40.7% |
| 6 | Qwen3.5-4B | 4.0B | 35.4% |
| 7 | Qwen3.5-2B | 2.0B | 29.1% |
| 8 | Llama 3.2 3B | 3.2B | 27.8% |
Read the full report
Methods, limitations, and grade-stratified results on the Cohere Labs Community blog.