Helios 5.0 Benchmarks

Uhuru AI · Helios 5.0

Ahead on African-language reasoning

Uhuru's flagship, Helios 5.0, measured head-to-head with the strongest external frontier model on African-language benchmarks — identical questions, identical conditions. On African-language mathematical reasoning it comes out ahead across every African language tested, by a statistically significant margin (McNemar p = 0.023) — and a second maths benchmark corroborates it. On reading comprehension the two are level. Every number here is measured and reproducible; re-run it yourself.

Helios 5.0 #1
African-language math
89.5% · ahead of Opus 4.8 · p = 0.023
Level
Reading comprehension
89.4 vs 89.0 · statistical tie
8 + control
Languages measured
African languages + English
Open harness
Reproducible
seeded · temp 0 · re-runnable
On African-language math reasoning, Helios 5.0 leads every African language (macro 89.5% vs Opus 88.0%, Helios 4.0 81.1%). The Helios 5.0 edge over Opus 4.8 is statistically significant on a paired test over 1,166 items (McNemar p = 0.023). The one place Opus edges ahead is the English control — Uhuru's advantage is specifically in African languages.
Language
Helios 5.0
Uhuru AI · flagship
Helios 4.0
Uhuru AI · agentic
Claude Opus 4.8
Anthropic
Swahili96.0%93.3%95.3%
Amharic93.9%81.0%93.2%
Yoruba91.1%80.8%90.4%
Hausa89.5%80.4%87.4%
Shona87.0%76.7%84.2%
isiZulu81.5%71.2%78.1%
isiXhosa78.7%65.2%75.9%
English (control)98.6%100.0%99.3%
Macro average89.5%81.1%88.0%
highest score in row (control row excluded)

AfriMGSM · masakhane/afrimgsm (IrokoBench)· n≈150/language · seed 20260719 · temperature 0

Corroborated on a second maths benchmark. On the elementary-mathematics questions of AfriMMLU (knowledge QA across African languages), Helios 5.0 scored 94.8% vs Opus 4.8's 74.6% (n≈232, pooled across 7 African languages; seed 20260719, temperature 0). This is the maths subject specifically — a second, independent measurement of the same African-language maths strength, not a claim about AfriMMLU's other subjects.