Team Ai
Datasetpublic

Maxyelow/kenyan-code-switch-1m

🇰🇪 Kenyan Code-Switching Pretraining Corpus (1,000,000 Sentences / 16.99M Words) The largest standardized, rule-audited monolingual pretraining corpus for Kenyan Code-Switching (Sheng / Technical Swahili-English Blend). Corpus Statistics Total Sentences: 1,000,000 sentences Total Word Tokens: 16,990,653 words Average Sentence Length: 16.99 words (multi-clause explanatory syntax) Linguistic Audit Score: 100.0000% compliance across all 20 Master Blueprint rules… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-1m.

sourceHugging Faceapache-2.0updated 12d agoView on Hugging Face
0likes49downloads
settings

This repository belongs to Maxyelow on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namekenyan-code-switch-1m
visibilitypublic
licenceapache-2.0
gatedno
ownerMaxyelow
Account settings
Maxyelow/kenyan-code-switch-1m · Team Ai