Releases, bootcamp news, research notes, and announcements.
17 upstream sources, 7.98M documents, 832M words and 1.136 billion subword tokens, measured over every document, plus what the four new sources changed and what the tokenizer still has to learn.
13 upstream sources, 7.35M documents, and ~912M subword tokens, plus a corrected, document-level tokenizer benchmark after we found and fixed a decode bug in v1.
Our v1 Somali tokenizer couldn't reconstruct its own output, decode(encode(x)) failed on every one of 49,424 held-out documents. v2 switches to byte-level BPE, fixes it, and comes out more efficient too: 1.3528 tokens/word on a corrected, document-level benchmark.
Somast ran five bootcamps between September 2025 and July 2026, all sponsored by Dugsiiye, all open to anyone, all still online. What open access made possible, who Dugsiiye is, and what each organization brought to the partnership.
One week after we uploaded a 13-hour, fully Somali-language Data Science & Machine Learning course to YouTube, it passed 12,000 views, the biggest video and course we've published on the channel so far. It covers what the video was and what we're doing about what it told us.
A demanding mathematics translation experiment revealed how the Somali Language Standard can help frontier models produce Somali that is more accurate, natural, consistent, and explainable.
Our fifth program in this series, and the first not built around data science at all. Git & GitHub closes a gap every earlier cohort surfaced: people who could write working code but weren't yet confident collaborating on it the way professional teams and open source projects do.
In June 2026 we ran our third Data Science & Machine Learning Bootcamp, and it's the biggest curriculum we've built for the program, eight core lessons on the full ML workflow, plus seven bonus sessions covering deep learning, generative AI, ethics, and career paths.
Somali is spoken by more than 20 million people, yet almost nothing in software or AI can point to a single authoritative, machine-readable source for its spelling, grammar, or terminology. Here is what the Somali Language Standard is, why it exists, what it's built on, and what changes once it's complete.
We've formally published the first two standards of the Somali Language Standard (SLS), an open, machine-readable, CI/CD-validated framework for the Somali language, starting with the alphabet and the standards process itself.
Foreign multilingual tokenizers over-segment Somali: BERT-base spends 2.69 tokens per word. Our native BPE tokenizer, trained on 529M words of SomNLP-Corpus, gets that down to 1.53: a 1.75× improvement.
After two Data Science & Machine Learning cohorts, we kept seeing the same bottleneck: people excited about ML who got stuck on plain Python before they ever reached a model. Python for Everyone is the bootcamp we built to close that gap on its own.
Our team shipped a Rust-based compiler for Soplang: Cranelift JIT and ahead-of-time builds, with syntax written entirely in Somali keywords.
Our largest Somali text release yet: 1.67M clean documents, 529M words, and ~793M tokens from HPLT, CC100, mC4, OPUS, MADLAD, and MT560, filtered through a six-stage pipeline with deep cleaning.
Five months after our first cohort, we ran a second Data Science & Machine Learning Bootcamp in February 2026. The core format held (one month, the same seven-stage ML workflow) but the curriculum itself changed in specific, deliberate ways.
In September 2025 we ran our first Data Science & Machine Learning Bootcamp, a one-month, hands-on program built to take a complete beginner through the full ML workflow to a deployed project. It covers what we taught, why we built it, and what it set in motion.
Dataset releases, new bootcamp cohorts, research papers, and lab updates, delivered straight to your inbox. No spam, unsubscribe anytime.