Team Ai
Datasetpublic

Usamasarfraz/code-switching-codesaviours-si26-usama

Code Switching Dataset Description This dataset contains Roman Urdu and English mixed sentences collected during the Code Saviours ML/AI Internship (SI-26). The purpose of this dataset is to help train NLP models that can understand code-switched language used by Pakistani people in daily conversations. Dataset Details Total sentences: 150+ Language: Roman Urdu + English Format: CSV Columns: sentence word label Label Meanings… See the full description on the dataset page: https://huggingface.co/datasets/Usamasarfraz/code-switching-codesaviours-si26-usama.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes5downloads
Dataset Card

Code Switching Dataset

Description

This dataset contains Roman Urdu and English mixed sentences collected during the Code Saviours ML/AI Internship (SI-26).

The purpose of this dataset is to help train NLP models that can understand code-switched language used by Pakistani people in daily conversations.


Dataset Details

  • —Total sentences: 150+
  • —Language: Roman Urdu + English
  • —Format: CSV

Columns:

  • —sentence
  • —word
  • —label

Label Meanings

  • —URD = Roman Urdu
  • —ENG = English
  • —MIX = Mixed language

Example:

WordLabel
AajURD
busyENG
thaURD

Data Sources

The data was collected from:

  • —Twitter / X
  • —Reddit
  • —YouTube comments
  • —Facebook public pages
  • —WhatsApp conversations (personal information removed)

Example Sentence

text
Aaj mera mood nahi hai for anything

Labels:

text
Aaj → URD
mera → URD
mood → ENG
for → ENG
anything → ENG

Purpose

This dataset can be used for:

  • —Code-switching detection
  • —Natural Language Processing (NLP)
  • —Language identification
  • —Chatbots
  • —Text classification

Author

Usama Sarfraz

Built during the Code Saviours ML/AI Internship — Batch SI-26.