goobolabs/somnlp-corpus
The largest open Somali text corpus - 1.1B+ tokens from 17 upstream sources (HPLT · CC-100 · mC4 · OPUS · MADLAD-400 · MT560 · QuranEnc · Tanzil · Wikipedia · XL-Sum · NLLB · Glot500 · Somali Web Corpus · FinePDFs · FineWeb-2 · Somali Alpaca · Somali TinyStories), cleaned through a five-stage pipeline (merge and exact dedup, clean, language ID, deep clean, near-duplicate removal).
View on Hugging FaceModality
Text
Size
1.1B+ tokens · 8.4 GB
License
Per-source
Language
Somali (af)
Format
Parquet / JSONL
Version
v2.0
Tasks
LM, NER, POS, Classification
Domains
HPLT · CC-100 · mC4 · OPUS · MADLAD-400 · MT560 · QuranEnc · Tanzil · Wikipedia · XL-Sum · NLLB · Glot500 · Somali Web Corpus · FinePDFs · FineWeb-2 · Somali Alpaca · Somali TinyStories
| Split | Rows | Size |
|---|---|---|
| final | 7,981,982 | 8.4 GB |
from datasets import load_dataset
ds = load_dataset("goobolabs/somnlp-corpus")
sample = ds["train"][0]
print(sample)@dataset{somnlp_corpus_2026,
title = {SomNLP-Corpus: Open Somali Text Corpus},
author = {Somast},
year = {2026},
version = {2.0},
note = {Each record carries its upstream license; there is no single corpus-wide license.},
url = {https://huggingface.co/datasets/goobolabs/somnlp-corpus}
}