Impact
Evidence of our work
Open datasets, alumni, and community reach.
1022+
GitHub stars across repos
From researchers worldwide
499+
Repository forks
Active community use
2
Papers in preparation
Somali NLP & speech research
5
Bootcamp cohorts
200+ engineers trained
1.1B+
Open tokens released
Largest open Somali text corpus
320h
Speech hours annotated
Word-level alignment
Community reach by country
Self-reported bootcamp enrollment (Feb 2026 DS & ML cohort).
Somalia
52%
Ethiopia
14%
Kenya
10%
United Kingdom
7%
United States
6%
Other
11%
Build with us
Build Somali AI in the open
Use our datasets and models, contribute to the research, or partner with the lab. We release everything we can.
Somast
1,022+
GitHub stars
499+
Forks
269+
Contributors