Team Ai
Datasetpublic

emanfatimaa05/code-switching-codesaviours-si26-eman

Code Switching NLP Dataset Code Saviours SI-26 | Week 6 Description This dataset contains Roman Urdu-English code-switching sentences collected for Code Saviours SI-26 Week 6. The purpose of this dataset is to provide examples of how Pakistani users naturally mix Roman Urdu and English in informal digital communication. Dataset Contents The dataset contains: 157 unique sentences 1,593 word-level entries 1,035 URD labels 558 ENG… See the full description on the dataset page: https://huggingface.co/datasets/emanfatimaa05/code-switching-codesaviours-si26-eman.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes6downloads
Dataset Card

Code Switching NLP Dataset

Code Saviours SI-26 | Week 6

Description

This dataset contains Roman Urdu-English code-switching sentences collected for Code Saviours SI-26 Week 6.

The purpose of this dataset is to provide examples of how Pakistani users naturally mix Roman Urdu and English in informal digital communication.

Dataset Contents

The dataset contains:

  • —157 unique sentences
  • —1,593 word-level entries
  • —1,035 URD labels
  • —558 ENG labels

Each word is labeled as either:

  • —URD — Roman Urdu / Urdu word
  • —ENG — English word

Data Collection

The sentences were collected from multiple sources, including:

  • —Pakistani social media content
  • —WhatsApp conversations, with personal information excluded
  • —Public online discussions
  • —Additional example sentences created to represent natural Roman Urdu-English code-switching patterns

The dataset focuses on informal language commonly used in Pakistani online communication.

Format

The dataset is provided as a CSV file with the following columns:

ColumnDescription
sentenceThe complete code-switched sentence
wordIndividual word from the sentence
labelLanguage label: URD or ENG

Example

sentencewordlabel
Aaj ka din bohot busy thaAajURD
Aaj ka din bohot busy thabusyENG
Aaj ka din bohot busy thathaURD

Intended Use

This dataset can be used for experimentation with:

  • —Code-switching NLP
  • —Language identification
  • —Roman Urdu NLP
  • —Word-level language classification
  • —NLP models for Pakistani digital communication

Limitations

This is a relatively small dataset containing 157 sentences. It is intended primarily for educational and experimental purposes and should not be considered a comprehensive representation of all Roman Urdu-English communication.

License

This dataset is provided for educational and research purposes as part of the Code Saviours SI-26 Week 6 project.