Building the foundations of Somali AI
We build open datasets, language and speech models, benchmarks, and research that help AI understand Somali. We release everything we can.
Language technology shapes who can take part
Somast is an open research lab building the datasets, language and speech models, and benchmarks that teach AI to understand Somali.
When AI doesn't understand a language, its speakers are left out of the tools shaping education, work, and public life. Closing that gap is what we build for, and we release everything we can openly.
20M+
Somali speakers worldwide
0.005%
of language-identified web data, vs. 45.5% for English
Education
Students can learn AI and build with technology in their own language, instead of only in English.
Research
Open Somali data and benchmarks let researchers study a low-resource language rigorously and reproducibly.
Accessibility
Speech models bring reading, writing, and search to millions who communicate mainly by voice.
Digital inclusion
As daily life moves online, Somali speakers need tools that understand how they speak and write.
The foundations of Somali AI
We build the layers others build on: data first, then models, benchmarks, and research. Education runs alongside all of it.
Partners and supporters
The organisations that back and build with Somast.
From raw data to open models in three steps
Collect & annotate
Community contributors, radio archives, and web crawls feed raw Somali text and speech into our pipeline. Every source is documented and licensed.
Clean, align & train
Data is deduplicated, filtered for quality, and used to train tokenizers, ASR models, and language models.
Release in the open
Datasets, model checkpoints, evaluation code, and research notes are published under permissive licenses for anyone to use, study, and improve.
Build it with us
Somast is a community effort. Whoever you are, there's a way to contribute.
Contribute on GitHubResearchers
Use our datasets and benchmarks, co-author studies, and push low-resource NLP forward.
Developers
Build on open models and tokenizers, file issues, and ship Somali-language features.
Students
Join fellowships and bootcamps to learn ML and contribute to real research.
Contributors
Help collect, transcribe, and verify data. Every contribution improves the models.
Stay in the loop
Dataset releases, new bootcamp cohorts, research papers, and lab updates, delivered straight to your inbox. No spam, unsubscribe anytime.



