Benchmarks

Benchmarks

Measuring how open language models understand Swahili in real sectors, starting with education.

Highlight · July 2026

NILEAGI-SUB: Swahili Language Understanding

NILEAGI-SUB is an evaluation programme for Swahili understanding across real-world sectors (education, health, finance, sports), one sector at a time under the same controlled protocol. This first report covers education and releases SUB-MCQ EDUCATION: 2,569 expert-reviewed Swahili multiple-choice items from the Swahili-subject curriculum (Standard 3-7 and Form 2-4).

2,569

Validated items

8

Grade bands

22.4%

Chance baseline

Items by grade

Grade bandItemsShare
Standard 3662.6%
Standard 41164.5%
Standard 527410.7%
Standard 665125.3%
Standard 751720.1%
Form 239915.5%
Form 31254.9%
Form 442116.4%
Total2,569100%

Level I results (~2-5B open models)

Eight open or openly accessible models evaluated on identical items with deterministic decoding and unified exact-match scoring. Accuracy spans about 27 points at a shared parameter budget: architecture, continued pretraining, and prompt compatibility matter more than raw size in this range.

#ModelParamsAccuracy
1Gemma 4 E4B4.5B54.8%
2AfriqueQwen3.5-4B4.0B50.3%
3AfriqueGemma-4B4.0B43.4%
4Tiny Aya Earth3.35B41.4%
5Tiny Aya Global3.35B40.7%
6Qwen3.5-4B4.0B35.4%
7Qwen3.5-2B2.0B29.1%
8Llama 3.2 3B3.2B27.8%

Read the full report

Methods, limitations, and grade-stratified results on the Cohere Labs Community blog.

Dataset