Skip to content

Building the foundations of Somali AI

We build open datasets, language and speech models, benchmarks, and research that help AI understand Somali. We release everything we can.

3
Open datasets
320h
Speech hours
1.1B+
Corpus tokens
Live demo
Why it matters

Language technology shapes who can take part

Somast is an open research lab building the datasets, language and speech models, and benchmarks that teach AI to understand Somali.

When AI doesn't understand a language, its speakers are left out of the tools shaping education, work, and public life. Closing that gap is what we build for, and we release everything we can openly.

20M+

Somali speakers worldwide

0.005%

of language-identified web data, vs. 45.5% for English

Education

Students can learn AI and build with technology in their own language, instead of only in English.

Research

Open Somali data and benchmarks let researchers study a low-resource language rigorously and reproducibly.

Accessibility

Speech models bring reading, writing, and search to millions who communicate mainly by voice.

Digital inclusion

As daily life moves online, Somali speakers need tools that understand how they speak and write.

Partners

Partners and supporters

The organisations that back and build with Somast.

Our approach

From raw data to open models in three steps

01

Collect & annotate

Community contributors, radio archives, and web crawls feed raw Somali text and speech into our pipeline. Every source is documented and licensed.

02

Clean, align & train

Data is deduplicated, filtered for quality, and used to train tokenizers, ASR models, and language models.

03

Release in the open

Datasets, model checkpoints, evaluation code, and research notes are published under permissive licenses for anyone to use, study, and improve.

Community

Build it with us

Somast is a community effort. Whoever you are, there's a way to contribute.

Contribute on GitHub

Researchers

Use our datasets and benchmarks, co-author studies, and push low-resource NLP forward.

Developers

Build on open models and tokenizers, file issues, and ship Somali-language features.

Students

Join fellowships and bootcamps to learn ML and contribute to real research.

Contributors

Help collect, transcribe, and verify data. Every contribution improves the models.

Newsletter

Stay in the loop

Dataset releases, new bootcamp cohorts, research papers, and lab updates, delivered straight to your inbox. No spam, unsubscribe anytime.