The corpus has passed one billion tokens. The 2026-09-24 build combines 17 public sources and produces 7,981,982 clean documents, 832,264,349 words and 1,136,027,958 subword tokens (8.4 GB of JSONL), measured by encoding every document rather than extrapolating from a sample.
Corpus growth
Since the 13-source build on 2026-09-02 the corpus gained 629,021 documents (+8.6%), 166 million words (+25%) and 224 million tokens (+25%). Words and tokens grew faster than documents because two of the four new sources are long web and PDF documents, not short parallel sentences.
1.136B
Subword tokens
832M
Words
7.98M
Clean documents
17
Upstream sources
- Run 13 (2026-09-02, 13 sources): 7,352,961 documents · 665,985,672 words · 911,824,557 tokens
- Run 17 (2026-09-24, 17 sources): 7,981,982 documents · 832,264,349 words · 1,136,027,958 tokens
- Tokens per word: 1.3650 (v2 tokenizer, whole corpus)
Pipeline and retention
The pipeline is unchanged: download, merge with exact dedup, clean, language ID, deep clean and MinHash near-dedup. Of 18.2 million downloaded rows, 44% survive to the final release.
- Downloaded (raw): 18,204,866 documents
- Merged: 11,708,895 (−6,495,971)
- Cleaned: 9,016,020 (−2,692,875)
- LID verified: 8,698,655 (−317,365)
- Deep cleaned: 8,638,442 (−60,213)
- Final: 7,981,982 (−656,460)
The whole pipeline ran in 1 hour 18 minutes on one Apple M3 Max laptop (16 cores, 64 GB, no GPU). Language ID is the slow stage at 53 minutes; near-dedup takes one.
Four new sources
This release adds FinePDFs, FineWeb-2, Somali Alpaca and Somali TinyStories. Because they are merged last, exact dedup keeps the older sources' copy of any shared text, and near-dedup keeps the longest member of each cluster, so a longer FineWeb-2 page can replace a shorter HPLT or mC4 near-duplicate.
- FineWeb-2: 589,824 final documents, 18.4% of all tokens
- FinePDFs: 20,771 documents, 2.4% of tokens
- Somali Alpaca (instruction data): 37,512 documents, 0.7%
- Somali TinyStories: 21,100 documents, 0.3% (half of its 42,000 stories were repeated upstream and merged into one copy)
NLLB is still the largest source by document count (4.1 million sentence-class documents), while FineWeb-2, mC4 and HPLT carry most of the long-form text.
Tokenizer on the bigger corpus
We re-scored the v2 tokenizer (48k ByteLevel BPE) on a held-out split of 64,041 documents from the 17-source corpus. It still round-trips every document exactly and emits no unknown tokens. Its vocabulary was trained on the earlier 11-source corpus and has not been retrained for the six sources added since.
1.3444
Mean tokens/word (v2)
1.000
Round-trip fidelity (v2)
2.6349
BERT-base tokens/word
1.8105
XLM-RoBERTa tokens/word
FinePDFs tokenizes 10.9% above the holdout mean (1.4906), in line with Wikipedia and mC4. That is the clearest sign that a retrained tokenizer would help on the newer sources; it is on the roadmap.
Licensing
There is still no single corpus-wide license. Every record carries its upstream license (CC0-1.0, CC-BY-SA-4.0, ODC-BY, CC-BY-4.0, MIT or source-specific terms). Somali TinyStories has no license on its dataset card and is recorded as unspecified. Check a record's provenance and license fields before redistributing.
How to use it
from datasets import load_dataset
from transformers import PreTrainedTokenizerFast
ds = load_dataset("goobolabs/somnlp-corpus", split="train")
print(ds[0])
tok = PreTrainedTokenizerFast(tokenizer_file="tokenizer/v2/tokenizer.json")
ids = tok.encode("Soomaaliya waa dal ku yaal Geeska Afrika.")
print(ids, tok.decode(ids))What's next
- Targeted web collection: Somali news, government, university and blog sites, with provenance
- Hugging Face packaging of the 17-source release with a reproducibility manifest
- Retrain the tokenizer on all 17 sources and re-run the benchmark
Writing about Somali language technology, open data, and AI from the lab in Mogadishu.