Usamasarfraz/code-switching-codesaviours-si26-usama
Code Switching Dataset Description This dataset contains Roman Urdu and English mixed sentences collected during the Code Saviours ML/AI Internship (SI-26). The purpose of this dataset is to help train NLP models that can understand code-switched language used by Pakistani people in daily conversations. Dataset Details Total sentences: 150+ Language: Roman Urdu + English Format: CSV Columns: sentence word label Label Meanings… See the full description on the dataset page: https://huggingface.co/datasets/Usamasarfraz/code-switching-codesaviours-si26-usama.
Code Switching Dataset
Description
This dataset contains Roman Urdu and English mixed sentences collected during the Code Saviours ML/AI Internship (SI-26).
The purpose of this dataset is to help train NLP models that can understand code-switched language used by Pakistani people in daily conversations.
Dataset Details
- Total sentences: 150+
- Language: Roman Urdu + English
- Format: CSV
Columns:
- sentence
- word
- label
Label Meanings
- URD = Roman Urdu
- ENG = English
- MIX = Mixed language
Example:
Data Sources
The data was collected from:
- Twitter / X
- YouTube comments
- Facebook public pages
- WhatsApp conversations (personal information removed)
Example Sentence
Aaj mera mood nahi hai for anythingLabels:
Aaj → URD
mera → URD
mood → ENG
for → ENG
anything → ENGPurpose
This dataset can be used for:
- Code-switching detection
- Natural Language Processing (NLP)
- Language identification
- Chatbots
- Text classification
Author
Usama Sarfraz
Built during the Code Saviours ML/AI Internship — Batch SI-26.
