Listening for Sukuma: Speech-to-Text Without a Pivot
After we opened Swahili → Sukuma translation, the natural next question was whether a machine could listen as carefully as it could read. Millions of people speak Sukuma every day, and too often their voices only become “usable” to software after someone translates them into English or Swahili first. We think that path is backwards: speech technology should meet people in the language they already use.
Today we are sharing a research preview of Sukuma speech → Sukuma text: two speech recognition checkpoints for literary, read-aloud Sukuma, kept deliberately small so they can run on phones, everyday machines, and places where big cloud ASR never bothered to learn the language. This is not an English or Swahili recognizer wearing a new label; if you give it Sukuma audio, it writes Sukuma text. You can try that loop below.
Why two versions?
Not every device has the same room to spare, so we share two sizes. Reach for Quality when recognition accuracy matters most, or Lite when memory and speed have to win. Both are on Hugging Face (quality · lite).
What to expect
This remains a research preview, and the honesty about its limits is part of the release. Clear, literary, single-speaker, read-aloud Sukuma is the intended setting, so conversational chatter, telephone audio, heavy noise, or overlapping speakers will usually do worse. Clips around thirty seconds or less at 16 kHz mono are the sweet spot for the live demo; longer or messier audio is where you should expect the model to struggle first.
Applications
The intended impact is simple: Sukuma speech should become Sukuma text, so voices do not have to pass through English or Swahili first.
Oral history and community archives
Write down read-aloud Sukuma so stories, scripture, and local recordings can be searched and kept in the language they were spoken.
Language classes and literacy
Give learners a first transcript of clear Sukuma speech they can correct, rather than starting from a blank page.
Field notes on everyday machines
Use a phone or laptop in places where cloud ASR never learned Sukuma, then review the text before it is treated as a record.
For builders
Load the weights with a Transformers ASR pipeline, or stay in the browser with the try box above and the fuller preview. Sample WAVs ship in the Hub repo. Methods, scores, and limits are in the technical report below, and in a forthcoming paper.
Technical report
Methods, evaluation, and limits, plus the open collection on Hugging Face.
Collaboration
If you are evaluating on new Sukuma speech, adapting these weights for your community, or exploring what comes after recognition, we would like to hear from you at hi@nileagi.com. The weights are released under CC BY-NC-SA 4.0; commercial use needs a written agreement with NileAGI.
Full preview
The Sukuma STT preview adds recording, Quality/Lite settings, and a ready Transformers script if you want to run the same idea locally.
Go to preview