Documented, versioned text and speech resources.
Permissively licensed, documented, and versioned. Browse, download, and cite.
Where our text comes from and how the corpus has grown since launch.
If you use our datasets in your research, please cite the appropriate entry below.
@dataset{somnlp-corpus-2026,
title = {SomNLP-Corpus: An Open Somali Text Corpus},
author = {Osman, Mohamud and Ahmed, Faadumo and Hassan, Abdirahman},
year = {2026},
version = {2.0},
url = {https://huggingface.co/datasets/goobolabs/somnlp-corpus},
note = {Each record carries its upstream license; there is no single corpus-wide license.}
}@dataset{somnlp-stt-2026,
title = {SomNLP-STT-Corpus: Somali Automatic Speech Recognition Dataset},
author = {Mohamed, Haawa and Osman, Mohamud},
year = {2026},
url = {https://huggingface.co/datasets/goobolabs/somnlp-stt},
license = {CC-BY-4.0}
}Use our datasets and models, contribute to the research, or partner with the lab. We release everything we can.
A glimpse of the models we train. Pick a sentence and watch it tokenize, classify, and translate, entirely in Af-Soomaali.
Waxaan jeclahay barashada cilmi-nafsiga iyo AI-ga.
Tokenization · 6 tokens
Model output