Team Ai
Datasetpublic

Maxyelow/kenyan-code-switch-1m

🇰🇪 Kenyan Code-Switching Pretraining Corpus (1,000,000 Sentences / 16.99M Words) The largest standardized, rule-audited monolingual pretraining corpus for Kenyan Code-Switching (Sheng / Technical Swahili-English Blend). Corpus Statistics Total Sentences: 1,000,000 sentences Total Word Tokens: 16,990,653 words Average Sentence Length: 16.99 words (multi-clause explanatory syntax) Linguistic Audit Score: 100.0000% compliance across all 20 Master Blueprint rules… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-1m.

sourceHugging Faceapache-2.0updated 12d agoView on Hugging Face
0likes49downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Maxyelow/kenyan-code-switch-1m · Team Ai