Team Ai
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes189 downloads1y agoHugging Face02Zerothe00 /code-switched-student-blindspot-eval Code Switched Student Blind Spot Evaluation Overview This repository contains a small manual evaluation of Qwen/Qwen2.5-1.5B-Instruct on code-switched South Asian international student prompts. The goal is to test whether a small open-weight instruction model can understand Pakistani English mixed with Roman Urdu/Hindi in situations shaped by scholarship pressure, family expectations, limited resources, and international student life in Malaysia. Blind… See the full description on the dataset page: https://huggingface.co/datasets/Zerothe00/code-switched-student-blindspot-eval.text-generation1 likes80 downloads15d agoHugging Face03SDAIANCAI /Ar-En-Code-Switching-Textual-Dataset ArE-CSTD: Arabic-English Code-Switching Textual Dataset The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "ArE-CSTD" dataset, which stands for "Arabic-English Code-Switching Textual Dataset”. This dataset contains 330K dialectical Arabic-English code-swithing sentences generated by the large language model GPT-4. TXT Files There are 6 txt files. 2 files for Modern Standard Arabic(MSA) train and… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Ar-En-Code-Switching-Textual-Dataset.texttext-generation100K<n<1M2 likes74 downloads2y agoHugging Face04Maxyelow /kenyan-code-switch-instruct-50k 🇰🇪 Kenyan Code-Switching & Sheng Multi-Task Instruction Dataset (50,000 Pairs) A standardized, multi-task instruction-tuning dataset engineered to teach Large Language Models (e.g. Llama 3, Mistral, Gemma, Qwen) to understand and generate authentic Kenyan Code-Switching (Sheng, Technical Swahili-English Blend) with rigorous adherence to Bantu morphotactic rules. Dataset Summary Total Samples: 50,000 instruction-response pairs train.jsonl: 45,000 pairs (90%)… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-instruct-50k.texttext-generation10K<n<100K0 likes70 downloads9d agoHugging Face05lxyuan /nemo-codeswitch-reasoning-debate Overview This is a synthetic, multilingual code-switching dataset. Each record contains: a realistic user query a long-form reasoning section a debate / counterargument section a concise final_answer It is designed for experiments in multilingual generation, code-switch robustness, and reasoning/debate style responses. This snapshot contains 574,977 rows and 10 string columns. Data provenance Important: Verify that your intended usage and redistribution complies… See the full description on the dataset page: https://huggingface.co/datasets/lxyuan/nemo-codeswitch-reasoning-debate.texttext-generation100K<n<1M0 likes46 downloads7mo agoHugging Face06tdnathmlenthusiast /qwen2.5-3b-codeswitch-blindspot Dataset Card: Code-Switched Agentic Reasoning Eval (Bengali / Hindi / Arabic) 16 hand-designed, adversarially-constructed reasoning tasks, each provided in 7 parallel language conditions (same semantic content, same gold answer, only the surface language changes): English, Bengali, Banglish, Hindi, Hinglish, Arabic, Arabilish. Built for evaluating whether a tool-using agent's correctness, calibration under ambiguity, and (separately, via the accompanying notebook)… See the full description on the dataset page: https://huggingface.co/datasets/tdnathmlenthusiast/qwen2.5-3b-codeswitch-blindspot.textquestion-answeringn<1K0 likes33 downloads4d agoHugging Face07Orinode /naija-customer-call-code-switch Naija Customer-Call Code-Switch Corpus (Orinode-CCS) Hand-written customer-service sentences with natural code-switching between Nigerian English and three indigenous Nigerian languages — Hausa, Yoruba, and Igbo. Covers 30+ business sectors typical of real customer-service calls in Nigeria. Released by Orinode under CC-BY 4.0 to support research on multilingual ASR, NLU, and conversational AI for low-resource African languages. Why this dataset exists Global voice-AI… See the full description on the dataset page: https://huggingface.co/datasets/Orinode/naija-customer-call-code-switch.texttext-generation10K<n<100K0 likes26 downloads5mo agoHugging Face08years0 /multilingual-code-switching-bench Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application) Question 1: Critical Blind Spot & Capability Gap Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/years0/multilingual-code-switching-bench.texttext-generationn<1K0 likes23 downloads15h agoHugging Face09mnasharifiya /hausa-codeswitch-eval Safety Asymmetry in Code-Switched Environments: Hausa-English Evaluation Dataset Overview This repository contains an evaluation dataset designed to test the safety mechanisms of frontier LLMs when handling inputs that rapidly code-switch between English and Hausa. Specifically, this dataset evaluates "Safety Asymmetry in Code-Switched Environments." As Large Language Models (LLMs) improve, their primary safety mechanisms (such as refusing malicious prompts) are… See the full description on the dataset page: https://huggingface.co/datasets/mnasharifiya/hausa-codeswitch-eval.text-generationn<1K1 likes22 downloads21h agoHugging Face10MagedSaeed /arabic-english-code-switching-text Arabic-English Code-Switching Dataset (Text Only) This dataset is a text-only version of Arabic-English Code-Switching Dataset dataset, created by this notebook. Changes Made Extracted only the text column from the original dataset. Usage from datasets import load_dataset dataset = load_dataset("MagedSaeed/arabic-english-code-switching-text") Citation Please reference/cite the original dataset when using this data. texttext-generation10K<n<100K1 likes21 downloads1y agoHugging Face11MouradGad /multilingual-code-switching-bench Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application) Question 1: Critical Blind Spot & Capability Gap Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/MouradGad/multilingual-code-switching-bench.texttext-generationn<1K0 likes16 downloads15h agoHugging Face12farabi-lab /Code_switchinggated 🇰🇿 Kazakh-Russian Code-Switching Normalization Dataset Dataset Summary Kazakh-Russian Code-Switching Normalization Dataset is a bilingual instruction-following dataset designed for identifying and rewriting Kazakh-Russian mixed-language text into clean Kazakh. The dataset focuses on informal communication, where Kazakh speakers may naturally mix Russian and Kazakh in one message. Each sample contains a prompt with code-switching, a response that identifies the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Code_switching.texttext-generationn<1K0 likes10 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.