Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TigreGotico /arabic-dialects-gold20-code-switch gold20-code-switch Code-switched Arabic sentences with IPA: 20 rows per lect across 33 Arabic lects (the same roster as the sibling TigreGotico/arabic-dialects-gold20). Each row embeds foreign material in a dialectal Arabic frame: inline Latin-script English (and French, for the lects whose live contact language is French), Arabic-script loanwords (سيرفس، كاش، موبايل-class), and Arabizi (Latin-written Arabic with digit gutturals). Columns (TSV, UTF-8, one file per lect): id… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-dialects-gold20-code-switch.texttext-to-speechn<1K0 likes631 downloads3mo agoHugging Face02shulhaaja /id-en-codeswitch-dataset-alternative Indonesian–English Code-Switching Synthetic Speech Dataset Synthetic speech generated for the undergraduate final project "Handling Code-Switching in Automatic Speech Recognition for Low-Resource Language Pairs: An Indonesian–English Case Study", School of Electrical Engineering and Informatics, Institut Teknologi Bandung. This dataset contains synthetic audio produced from the Indonesian–English code-switching text corpora released in the companion repository below. It was used… See the full description on the dataset page: https://huggingface.co/datasets/shulhaaja/id-en-codeswitch-dataset-alternative.audioautomatic-speech-recognition10K<n<100K0 likes141 downloads2mo agoHugging Face03prokelly /neuromoyo-sahara-codeswitch-benchmark NEUROMOYO — Sahara CodeSwitch Africa Benchmark 🔗 Live Benchmark Results Interactive benchmark: https://www.neuromoyo.app/benchmark This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech. 🚀 Live NEUROMOYO Demo Live application: https://www.neuromoyo.app The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.tabularn<1K0 likes89 downloads16d agoHugging Face04code-switching /question-answertext1K<n<10K0 likes86 downloads1mo agoHugging Face05SDAIANCAI /Saudilang-Code-Switch-Corpus SCC - Saudilang Code-Switch Corpus The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”. This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.tabularautomatic-speech-recognition1K<n<10K3 likes59 downloads2y agoHugging Face06code-switching /naturalnesstabularn<1K0 likes32 downloads1mo agoHugging Face07samihyounes /Moroccan-Codeswitching Moroccan Darija Code-Switched Corpus (Sentence-level TSV) Dataset Summary This dataset contains sentence/post-level code-switched Moroccan Darija text with a single label per text unit. It is intended to support NLP research on Moroccan Darija (Darija), an under-resourced Arabic variety, and on sentence-level code-switching / language identification in Moroccan online text. Languages The corpus may contain Moroccan Darija (often ary) and code-switching with:… See the full description on the dataset page: https://huggingface.co/datasets/samihyounes/Moroccan-Codeswitching.texttext-classification100K<n<1M0 likes32 downloads7mo agoHugging Face08sameeramin /code-switched-language-identificationtext100K<n<1M0 likes21 downloads2y agoHugging Face09Hamza-Ali01 /code-switching-codesaviours-si26-hamzatext1K<n<10K1 likes20 downloads2mo agoHugging Face10AsadChanna /code-switched-medical-blindspot-eval Code-Switched Medical Query Evaluation: Qwen2.5-1.5B-Instruct Blind Spot Standard AI benchmarks mostly test clean English or native-script text in a single language. They rarely test how many people in Pakistan actually type when describing a health problem: Roman-script Urdu mixed mid-sentence with English ("chest mein tightness hai"), or Sindhi. This register is informal, has no standard spelling, and is under-represented in training data and evaluation sets.… See the full description on the dataset page: https://huggingface.co/datasets/AsadChanna/code-switched-medical-blindspot-eval.tabularn<1K0 likes20 downloads2d agoHugging Face11hamnaheh /code-switching-codesaviours-si26-humna Code Switching NLP Dataset — Code Saviours SI-26, Humna Imran Word-level labeled dataset of naturally code-switched Roman Urdu + English sentences, as commonly written by Pakistani social media users. What it is 220 sentences (2,490 word-level rows). Each word is labeled URD (Roman Urdu), ENG (English), or MIX (hybrid token, e.g. hyphenated compounds). How it was built Source sentences come from Smat26/Roman-Urdu-Dataset (GitHub), a public… See the full description on the dataset page: https://huggingface.co/datasets/hamnaheh/code-switching-codesaviours-si26-humna.text1K<n<10K0 likes19 downloads2mo agoHugging Face12Moazamzf /code-switching-codesaviours-si26-Moazam Roman Urdu-English Code-Switching Dataset Description This dataset contains naturally occurring Roman Urdu / English code-switched sentences, collected to reflect how Pakistani speakers actually communicate online — mixing Roman Urdu and English within the same sentence (e.g. "Aaj mera mood nahi hai for anything"). Each sentence is broken down word-by-word, with every word labeled by language. Collection Method Sentences were collected from a mix of… See the full description on the dataset page: https://huggingface.co/datasets/Moazamzf/code-switching-codesaviours-si26-Moazam.text1K<n<10K1 likes18 downloads2mo agoHugging Face13taniakoh /code-switch-xnli-test CS-XNLI: Synthetic Code-Switched NLI Evaluation Corpus Dataset Summary CS-XNLI is a synthetically generated code-switched dataset derived from the standard Cross-Lingual NLI (XNLI) evaluation benchmark. This dataset aims to address the scarcity of mixed-language resources for complex reasoning tasks. Key Contents 1000 Annotated Sentence Pairs for Intra-Sential Code-Switching Each: English-Spanish(es), English-Vietnamese(vi), English-Mandarin(zh) Key… See the full description on the dataset page: https://huggingface.co/datasets/taniakoh/code-switch-xnli-test.texttext-classification1K<n<10K0 likes17 downloads10mo agoHugging Face14Hassaanatif992 /code-switching-codesaviours-si26-MuhammadHassaan Roman Urdu-English Code-Switching Dataset Description This dataset contains 150 sentences that mix Roman Urdu and English. The purpose of this dataset is to study code-switching between Roman Urdu and English. Labels URD: Roman Urdu word ENG: English word MIX: Mixed or unclear word Dataset Format Each row contains: sentence word label Collection The sentences were prepared as natural Roman Urdu-English… See the full description on the dataset page: https://huggingface.co/datasets/Hassaanatif992/code-switching-codesaviours-si26-MuhammadHassaan.text1K<n<10K0 likes15 downloads2mo agoHugging Face15Muhammad-Ahmad-1263 /code-switching-codesaviours-si26-muhammadahmad Code-Switching Codesaviours SI26 — Muhammad Ahmad Dataset Description This dataset contains 155 naturally occurring Roman Urdu–English code-switched sentences (1,400+ word-level entries), reflecting how Roman Urdu and English are mixed in everyday informal communication by Pakistani speakers online (Twitter/X, WhatsApp, YouTube comments, Reddit). Code-switching — alternating between two or more languages within a single sentence or conversation — is extremely… See the full description on the dataset page: https://huggingface.co/datasets/Muhammad-Ahmad-1263/code-switching-codesaviours-si26-muhammadahmad.text1K<n<10K0 likes15 downloads2mo agoHugging Face16122Uswa /code-switching-codesaviours-si26-Uswa Code-Switching Dataset: Roman Urdu ↔ English (Pakistan) Dataset Description This dataset contains 191 naturally code-switched Roman Urdu–English sentences** (well above the 150 minimum) blending Roman Urdu and English, the way Pakistani speakers actually write online. Every word in every sentence is labelled at the token level, making this a word-level sequence labelling / language identification dataset for code-switched text. Roman Urdu–English mixing is… See the full description on the dataset page: https://huggingface.co/datasets/122Uswa/code-switching-codesaviours-si26-Uswa.texttoken-classification1K<n<10K0 likes14 downloads2mo agoHugging Face17Maryam657775 /code-switching-codesaviours-si26-maryam Roman Urdu & English Code-Switching Dataset Description This dataset contains 152 real-world sentences where users naturally mix Roman Urdu and English. It was built to train models to handle how 230 million Pakistanis actually communicate online. Data Collection The raw text was scraped directly from natural conversations on Pakistani Twitter/X and anonymized WhatsApp messages. Label Meanings Every single word in this dataset has… See the full description on the dataset page: https://huggingface.co/datasets/Maryam657775/code-switching-codesaviours-si26-maryam.text1K<n<10K0 likes14 downloads2mo agoHugging Face18Saima-Manzoor /code-switching-codesaviours-si26-saima Code-Switching Urdu-English Dataset Dataset Description This dataset is created for Urdu-English code-switching language identification. It contains sentences that include Urdu words, English words, and a small number of mixed-language entries. The dataset was prepared by collecting and organizing code-switched Urdu-English sentences. Each sentence was divided into individual words, and every word was assigned a language label. The data was then converted into a… See the full description on the dataset page: https://huggingface.co/datasets/Saima-Manzoor/code-switching-codesaviours-si26-saima.text1K<n<10K0 likes13 downloads2mo agoHugging Face19Noisy77 /code-switching-codesaviours-si26-bilal Roman Urdu-English Code Switching Dataset Dataset Description A manually labeled dataset of 160+ code-switching sentences where Roman Urdu and English are naturally mixed — reflecting how 230 million Pakistanis actually communicate online. Each word in every sentence is tagged with a language label, making this dataset suitable for token-level language identification and code-switching NLP research. Label Meanings Label Description Examples… See the full description on the dataset page: https://huggingface.co/datasets/Noisy77/code-switching-codesaviours-si26-bilal.texttoken-classification1K<n<10K0 likes13 downloads2mo agoHugging Face20qandeelasim13 /code-switching-codesaviours-si26-qandeel Code Switching NLP Dataset | Code Saviours SI-26 Dataset Description A word-level labelled Roman Urdu–English code-switching dataset (200 sentences, 3182 word entries). Labels: URD (Roman Urdu), ENG (English), MIX (nativized loanword). Source Sentences filtered from the Roman Urdu Data Set (Sharf, 2017, UCI ML Repository, CC BY 4.0), originally collected from e-commerce reviews, Facebook comments, and Twitter posts. Filtered for genuine… See the full description on the dataset page: https://huggingface.co/datasets/qandeelasim13/code-switching-codesaviours-si26-qandeel.text1K<n<10K0 likes12 downloads2mo agoHugging Face21waroodzkhan /code-switching-codesaviours-si26-warood Code-Switching Codesaviours SI-26 Dataset Dataset Description This dataset contains 150 naturally occurring Roman Urdu–English code-switched sentences, commonly spoken by Pakistani speakers in casual, everyday communication. Each sentence has been broken down into individual words, and every word is labeled by language. Roman Urdu–English code-switching is extremely common in Pakistan (spoken/written by an estimated 230+ million people) but is poorly handled by… See the full description on the dataset page: https://huggingface.co/datasets/waroodzkhan/code-switching-codesaviours-si26-warood.text1K<n<10K0 likes12 downloads2mo agoHugging Face22Zainab-Binte-Khalid /code-switching-codesaviours-si26-zainab Roman Urdu–English Code-Switching Dataset Dataset Description This dataset contains 1,901 sentences and 21,370 word-level language labels, built to capture how Roman Urdu and English are naturally mixed together in everyday Pakistani online communication. Code-switching — blending two languages within a single sentence — is how the vast majority of Pakistanis actually write and speak online, on platforms like Twitter/X, Facebook, YouTube, Reddit, and WhatsApp. A… See the full description on the dataset page: https://huggingface.co/datasets/Zainab-Binte-Khalid/code-switching-codesaviours-si26-zainab.text10K<n<100K1 likes11 downloads2mo agoHugging Face23kashaf112-s /code-switching-codesaviours-si26-kashaftext1K<n<10K0 likes11 downloads2mo agoHugging Face24Maryam12256 /code-switching-codesaviours-si26-Maryam Code-Switching Dataset — Roman Urdu / English Dataset Description This dataset contains naturally occurring code-switched sentences that mix Roman Urdu and English, the way Pakistani speakers commonly write in everyday digital communication (texts, tweets, comments). Each sentence is broken down word-by-word, with every word labeled by language. Total sentences: 150 Total labeled word entries: ~1,478 Format: flat CSV — one row per word, with the parent sentence… See the full description on the dataset page: https://huggingface.co/datasets/Maryam12256/code-switching-codesaviours-si26-Maryam.text1K<n<10K0 likes10 downloads2mo agoHugging Face25samaikaimran /code-switching-codesaviours-si26-samaika Code Switching NLP | Code Saviours SI-26 | Samaika About the Dataset This dataset contains Roman Urdu and English code-switching sentences collected from social media comments and online content. The purpose of this dataset is to represent the way Pakistani users naturally mix Roman Urdu and English while communicating online. Data Collection The sentences were collected from: Instagram comments YouTube comments Twitter/X The collected sentences… See the full description on the dataset page: https://huggingface.co/datasets/samaikaimran/code-switching-codesaviours-si26-samaika.text1K<n<10K0 likes9 downloads2mo agoHugging Face26AmnaNoor123 /code-switching-codesaviours-si26-amnaCode Switching Dataset — Code Saviours SI-26 (Amna) Dataset Description Roman Urdu mixed with English — e.g. "Aaj mera mood nahi hai for anything" — is how a huge share of Pakistanis actually write online, but almost no existing NLP resource labels this kind of code-switching at the word level. This dataset provides 160 real, naturally occurring mixed-language sentences, tokenised and labelled word-by-word as Roman Urdu, English, or a genuine hybrid blend. How It Was Collected Sentences were… See the full description on the dataset page: https://huggingface.co/datasets/AmnaNoor123/code-switching-codesaviours-si26-amna.texttoken-classification1K<n<10K0 likes9 downloads2mo agoHugging Face27izzazahid /code-switching-codesaviours-si26-izza Code Switching NLP Dataset Overview This dataset was created for the Code Saviours SI-26 Week 6 internship project. It contains Roman Urdu-English code-switching sentences collected from public online reviews and converted into a word-level labeled dataset. Dataset Statistics Total Sentences: 150 Total Word Entries: 3911 Labels URD = Roman Urdu ENG = English MIX = Mixed / Other Files dataset.csv… See the full description on the dataset page: https://huggingface.co/datasets/izzazahid/code-switching-codesaviours-si26-izza.texttoken-classification1K<n<10K0 likes8 downloads2mo agoHugging Face28sheezariaz2315 /code-switching-codesaviours-si26-sheeza Code Switching NLP Dataset Dataset Description This dataset contains 161 code-switching sentences that combine Roman Urdu and English. The dataset represents informal Pakistani online communication, including social media posts, comments, messages, and everyday digital conversations. Each sentence is labelled at the word level. Labels URD: Roman Urdu words ENG: English words MIX: Tokens that combine Roman Urdu and English within the same word… See the full description on the dataset page: https://huggingface.co/datasets/sheezariaz2315/code-switching-codesaviours-si26-sheeza.text1K<n<10K0 likes8 downloads2mo agoHugging Face29sanaisrail /code-switching-codesaviours-si26-Sanatext1K<n<10K0 likes8 downloads1mo agoHugging Face30Usamasarfraz /code-switching-codesaviours-si26-usama Code Switching Dataset Description This dataset contains Roman Urdu and English mixed sentences collected during the Code Saviours ML/AI Internship (SI-26). The purpose of this dataset is to help train NLP models that can understand code-switched language used by Pakistani people in daily conversations. Dataset Details Total sentences: 150+ Language: Roman Urdu + English Format: CSV Columns: sentence word label Label Meanings… See the full description on the dataset page: https://huggingface.co/datasets/Usamasarfraz/code-switching-codesaviours-si26-usama.textn<1K0 likes7 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.