Maxyelow/kenyan-code-switch-1m
🇰🇪 Kenyan Code-Switching Pretraining Corpus (1,000,000 Sentences / 16.99M Words) The largest standardized, rule-audited monolingual pretraining corpus for Kenyan Code-Switching (Sheng / Technical Swahili-English Blend). Corpus Statistics Total Sentences: 1,000,000 sentences Total Word Tokens: 16,990,653 words Average Sentence Length: 16.99 words (multi-clause explanatory syntax) Linguistic Audit Score: 100.0000% compliance across all 20 Master Blueprint rules… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-1m.
Conversations for this repository live on Hugging Face.
Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face