Maryam12256/code-switching-codesaviours-si26-Maryam
Code-Switching Dataset — Roman Urdu / English Dataset Description This dataset contains naturally occurring code-switched sentences that mix Roman Urdu and English, the way Pakistani speakers commonly write in everyday digital communication (texts, tweets, comments). Each sentence is broken down word-by-word, with every word labeled by language. Total sentences: 150 Total labeled word entries: ~1,478 Format: flat CSV — one row per word, with the parent sentence… See the full description on the dataset page: https://huggingface.co/datasets/Maryam12256/code-switching-codesaviours-si26-Maryam.
Code-Switching Dataset — Roman Urdu / English
Dataset Description
This dataset contains naturally occurring code-switched sentences that mix Roman Urdu and English, the way Pakistani speakers commonly write in everyday digital communication (texts, tweets, comments). Each sentence is broken down word-by-word, with every word labeled by language.
- Total sentences: 150
- Total labeled word entries: ~1,478
- Format: flat CSV — one row per word, with the parent sentence and its label
How It Was Collected
Sentences were written and compiled to reflect realistic, informal code-switching patterns seen in Pakistani online communication — the kind of mixing found in Twitter/X posts, WhatsApp messages, Reddit comments (r/pakistan), YouTube comments on Pakistani videos, and Facebook public pages. The sentences follow natural code-switching patterns (e.g., "Aaj ka din bohot busy tha, had 3 meetings back to back") rather than being direct or formal translations.
Label Meanings
Each word in a sentence is tagged with one of the following labels:
Columns
Intended Use
This dataset is intended for training and evaluating token-level language identification and code-switching detection models for Roman Urdu–English text, common in South Asian NLP applications.
Limitations
- Labels were assigned using a rule-based heuristic (a curated Roman Urdu word list) and may contain minor labeling inconsistencies, especially for ambiguous or context-dependent words.
- The dataset is relatively small (150 sentences) and may not capture the full diversity of real-world code-switching patterns.
