Skip to content
Impact

Evidence of our work

Open datasets, alumni, and community reach.

1022+

GitHub stars across repos

From researchers worldwide

499+

Repository forks

Active community use

2

Papers in preparation

Somali NLP & speech research

5

Bootcamp cohorts

200+ engineers trained

1.1B+

Open tokens released

Largest open Somali text corpus

320h

Speech hours annotated

Word-level alignment

Community reach by country

Self-reported bootcamp enrollment (Feb 2026 DS & ML cohort).

Somalia
52%
Ethiopia
14%
Kenya
10%
United Kingdom
7%
United States
6%
Other
11%
Build with us

Build Somali AI in the open

Use our datasets and models, contribute to the research, or partner with the lab. We release everything we can.

Somast
1,022+
GitHub stars
499+
Forks
269+
Contributors