datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.code-switched-student-blindspot-eval
Code Switched Student Blind Spot Evaluation
Overview
This repository contains a small manual evaluation of Qwen/Qwen2.5-1.5B-Instruct on code-switched South Asian international student prompts. The goal is to test whether a small open-weight instruction model can understand Pakistani English mixed with Roman Urdu/Hindi in situations shaped by scholarship pressure, family expectations, limited resources, and international student life in Malaysia.
Blind… See the full description on the dataset page: https://huggingface.co/datasets/Zerothe00/code-switched-student-blindspot-eval.Ar-En-Code-Switching-Textual-Dataset
ArE-CSTD: Arabic-English Code-Switching Textual Dataset
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "ArE-CSTD" dataset, which stands for "Arabic-English Code-Switching Textual Dataset”.
This dataset contains 330K dialectical Arabic-English code-swithing sentences generated by the large language model GPT-4.
TXT Files
There are 6 txt files. 2 files for Modern Standard Arabic(MSA) train and… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Ar-En-Code-Switching-Textual-Dataset.kenyan-code-switch-instruct-50k
🇰🇪 Kenyan Code-Switching & Sheng Multi-Task Instruction Dataset (50,000 Pairs)
A standardized, multi-task instruction-tuning dataset engineered to teach Large Language Models (e.g. Llama 3, Mistral, Gemma, Qwen) to understand and generate authentic Kenyan Code-Switching (Sheng, Technical Swahili-English Blend) with rigorous adherence to Bantu morphotactic rules.
Dataset Summary
Total Samples: 50,000 instruction-response pairs
train.jsonl: 45,000 pairs (90%)… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-instruct-50k.nemo-codeswitch-reasoning-debate
Overview
This is a synthetic, multilingual code-switching dataset. Each record contains:
a realistic user query
a long-form reasoning section
a debate / counterargument section
a concise final_answer
It is designed for experiments in multilingual generation, code-switch robustness, and reasoning/debate style responses.
This snapshot contains 574,977 rows and 10 string columns.
Data provenance
Important:
Verify that your intended usage and redistribution complies… See the full description on the dataset page: https://huggingface.co/datasets/lxyuan/nemo-codeswitch-reasoning-debate.qwen2.5-3b-codeswitch-blindspot
Dataset Card: Code-Switched Agentic Reasoning Eval (Bengali / Hindi / Arabic)
16 hand-designed, adversarially-constructed reasoning tasks, each provided in 7 parallel
language conditions (same semantic content, same gold answer, only the surface language
changes): English, Bengali, Banglish, Hindi, Hinglish, Arabic, Arabilish.
Built for evaluating whether a tool-using agent's correctness, calibration under
ambiguity, and (separately, via the accompanying notebook)… See the full description on the dataset page: https://huggingface.co/datasets/tdnathmlenthusiast/qwen2.5-3b-codeswitch-blindspot.naija-customer-call-code-switch
Naija Customer-Call Code-Switch Corpus (Orinode-CCS)
Hand-written customer-service sentences with natural code-switching between Nigerian English and three indigenous Nigerian languages — Hausa, Yoruba, and Igbo. Covers 30+ business sectors typical of real customer-service calls in Nigeria.
Released by Orinode under CC-BY 4.0 to support research on multilingual ASR, NLU, and conversational AI for low-resource African languages.
Why this dataset exists
Global voice-AI… See the full description on the dataset page: https://huggingface.co/datasets/Orinode/naija-customer-call-code-switch.multilingual-code-switching-bench
Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application)
Question 1: Critical Blind Spot & Capability Gap
Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/years0/multilingual-code-switching-bench.hausa-codeswitch-eval
Safety Asymmetry in Code-Switched Environments: Hausa-English Evaluation Dataset
Overview
This repository contains an evaluation dataset designed to test the safety mechanisms of frontier LLMs when handling inputs that rapidly code-switch between English and Hausa. Specifically, this dataset evaluates "Safety Asymmetry in Code-Switched Environments."
As Large Language Models (LLMs) improve, their primary safety mechanisms (such as refusing malicious prompts) are… See the full description on the dataset page: https://huggingface.co/datasets/mnasharifiya/hausa-codeswitch-eval.arabic-english-code-switching-text
Arabic-English Code-Switching Dataset (Text Only)
This dataset is a text-only version of Arabic-English Code-Switching Dataset dataset,
created by this notebook.
Changes Made
Extracted only the text column from the original dataset.
Usage
from datasets import load_dataset
dataset = load_dataset("MagedSaeed/arabic-english-code-switching-text")
Citation
Please reference/cite the original dataset when using this data.
multilingual-code-switching-bench
Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application)
Question 1: Critical Blind Spot & Capability Gap
Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/MouradGad/multilingual-code-switching-bench.Code_switching
🇰🇿 Kazakh-Russian Code-Switching Normalization Dataset
Dataset Summary
Kazakh-Russian Code-Switching Normalization Dataset is a bilingual instruction-following dataset designed for identifying and rewriting Kazakh-Russian mixed-language text into clean Kazakh.
The dataset focuses on informal communication, where Kazakh speakers may naturally mix Russian and Kazakh in one message. Each sample contains a prompt with code-switching, a response that identifies the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Code_switching.
