Why we build Somali AI
The story, mission, and people behind Somast, short for Somali Science and Technology.
Every language needs AI infrastructure. We build it for Somali first, openly, and to a standard the field can trust.
Open by default
Datasets, model weights, and evaluation code are released openly so anyone can build, verify, and improve them.
Community-owned
The people who speak the language help build the technology, and own its direction.
Rigorous & reproducible
Clear benchmarks, documented methods, and results others can reproduce and extend.
From raw data to open models in three steps
Collect & annotate
Community contributors, radio archives, and web crawls feed raw Somali text and speech into our pipeline. Every source is documented and licensed.
Clean, align & train
Data is deduplicated, filtered for quality, and used to train tokenizers, ASR models, and language models.
Release in the open
Datasets, model checkpoints, evaluation code, and research notes are published under permissive licenses for anyone to use, study, and improve.
Somali is ~0.005% of language-identified web data, less than 0.01% of English's share
More than 20 million people speak Somali, yet it barely registers in the datasets that train the world's AI. That gap is why Somast exists.
Source: Common Crawl primary-language page distribution (CLD2, aggregated crawls through CC-MAIN-2026). Percentages are share of language-identified pages. Somali vs English ratio ≈ 0.01% of English's page share.
Open Somali text we've released, not shown on the same scale as global crawl percentages above.
Sharafdin Yusuf
Lead Engineer & Researcher
Somali people should build, shape, and own the systems that reflect their language and culture, as well as use AI tools made elsewhere.
The people building it
Build Somali AI in the open
Use our datasets and models, contribute to the research, or partner with the lab. We release everything we can.