Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01liva-ai /code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here. Code-Switching ASR Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.audioautomatic-speech-recognitionn<1K0 likes200 downloads3mo agoHugging Face02Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes129 downloads1y agoHugging Face03code-switching /question-answertext1K<n<10K0 likes82 downloads1mo agoHugging Face04code-switching /text-summarizationtextsummarizationn<1K0 likes78 downloads1mo agoHugging Face05devrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes61 downloads7mo agoHugging Face06code-switching /naturalnesstabularn<1K0 likes33 downloads1mo agoHugging Face07samihyounes /Moroccan-Codeswitching Moroccan Darija Code-Switched Corpus (Sentence-level TSV) Dataset Summary This dataset contains sentence/post-level code-switched Moroccan Darija text with a single label per text unit. It is intended to support NLP research on Moroccan Darija (Darija), an under-resourced Arabic variety, and on sentence-level code-switching / language identification in Moroccan online text. Languages The corpus may contain Moroccan Darija (often ary) and code-switching with:… See the full description on the dataset page: https://huggingface.co/datasets/samihyounes/Moroccan-Codeswitching.texttext-classification100K<n<1M0 likes30 downloads7mo agoHugging Face08code-switching /topic-classificationtabulartext-classificationn<1K0 likes29 downloads1mo agoHugging Face09zenyn /Code-Switching-Testaudion<1K0 likes27 downloads2y agoHugging Face10gimmy256 /african-codeswitching Pan-African Code-Switching Dataset Built with Adaptive Data by Adaption | Crane AI Labs Submitted to the Uncharted Data Challenge 2026 by Adaption Labs Overview The first open-source, richly-annotated pan-African code-switching dataset covering authentic language mixing patterns across 4 African regions — East Africa, South Africa, and West Africa — in 10 distinct language-pair configurations. Code-switching (mixing two or more languages in a single utterance) is how… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/african-codeswitching.texttext-classificationn<1K0 likes25 downloads6mo agoHugging Face11DynamicSuperb /CodeSwitchingSpeechIdentification_ASCENDaudion<1K0 likes21 downloads2y agoHugging Face12hamnaheh /code-switching-codesaviours-si26-humna Code Switching NLP Dataset — Code Saviours SI-26, Humna Imran Word-level labeled dataset of naturally code-switched Roman Urdu + English sentences, as commonly written by Pakistani social media users. What it is 220 sentences (2,490 word-level rows). Each word is labeled URD (Roman Urdu), ENG (English), or MIX (hybrid token, e.g. hyphenated compounds). How it was built Source sentences come from Smat26/Roman-Urdu-Dataset (GitHub), a public… See the full description on the dataset page: https://huggingface.co/datasets/hamnaheh/code-switching-codesaviours-si26-humna.text1K<n<10K0 likes21 downloads2mo agoHugging Face13DynamicSuperb /CodeSwitchingSemanticGrammarAcceptabilityComparison_CSZS-zh-enaudion<1K0 likes19 downloads2y agoHugging Face14Hamza-Ali01 /code-switching-codesaviours-si26-hamzatext1K<n<10K1 likes18 downloads2mo agoHugging Face15Nash-pAnDiTa /ASR_En_Ar_CodeSwitchingaudio10K<n<100K0 likes17 downloads2y agoHugging Face16122Uswa /code-switching-codesaviours-si26-Uswa Code-Switching Dataset: Roman Urdu ↔ English (Pakistan) Dataset Description This dataset contains 191 naturally code-switched Roman Urdu–English sentences** (well above the 150 minimum) blending Roman Urdu and English, the way Pakistani speakers actually write online. Every word in every sentence is labelled at the token level, making this a word-level sequence labelling / language identification dataset for code-switched text. Roman Urdu–English mixing is… See the full description on the dataset page: https://huggingface.co/datasets/122Uswa/code-switching-codesaviours-si26-Uswa.texttoken-classification1K<n<10K0 likes15 downloads2mo agoHugging Face17kashaf112-s /code-switching-codesaviours-si26-kashaftext1K<n<10K0 likes15 downloads2mo agoHugging Face18Muhammad-Ahmad-1263 /code-switching-codesaviours-si26-muhammadahmad Code-Switching Codesaviours SI26 — Muhammad Ahmad Dataset Description This dataset contains 155 naturally occurring Roman Urdu–English code-switched sentences (1,400+ word-level entries), reflecting how Roman Urdu and English are mixed in everyday informal communication by Pakistani speakers online (Twitter/X, WhatsApp, YouTube comments, Reddit). Code-switching — alternating between two or more languages within a single sentence or conversation — is extremely… See the full description on the dataset page: https://huggingface.co/datasets/Muhammad-Ahmad-1263/code-switching-codesaviours-si26-muhammadahmad.text1K<n<10K0 likes15 downloads2mo agoHugging Face19Moazamzf /code-switching-codesaviours-si26-Moazam Roman Urdu-English Code-Switching Dataset Description This dataset contains naturally occurring Roman Urdu / English code-switched sentences, collected to reflect how Pakistani speakers actually communicate online — mixing Roman Urdu and English within the same sentence (e.g. "Aaj mera mood nahi hai for anything"). Each sentence is broken down word-by-word, with every word labeled by language. Collection Method Sentences were collected from a mix of… See the full description on the dataset page: https://huggingface.co/datasets/Moazamzf/code-switching-codesaviours-si26-Moazam.text1K<n<10K1 likes13 downloads2mo agoHugging Face20Hassaanatif992 /code-switching-codesaviours-si26-MuhammadHassaan Roman Urdu-English Code-Switching Dataset Description This dataset contains 150 sentences that mix Roman Urdu and English. The purpose of this dataset is to study code-switching between Roman Urdu and English. Labels URD: Roman Urdu word ENG: English word MIX: Mixed or unclear word Dataset Format Each row contains: sentence word label Collection The sentences were prepared as natural Roman Urdu-English… See the full description on the dataset page: https://huggingface.co/datasets/Hassaanatif992/code-switching-codesaviours-si26-MuhammadHassaan.text1K<n<10K0 likes12 downloads2mo agoHugging Face21Noisy77 /code-switching-codesaviours-si26-bilal Roman Urdu-English Code Switching Dataset Dataset Description A manually labeled dataset of 160+ code-switching sentences where Roman Urdu and English are naturally mixed — reflecting how 230 million Pakistanis actually communicate online. Each word in every sentence is tagged with a language label, making this dataset suitable for token-level language identification and code-switching NLP research. Label Meanings Label Description Examples… See the full description on the dataset page: https://huggingface.co/datasets/Noisy77/code-switching-codesaviours-si26-bilal.texttoken-classification1K<n<10K0 likes12 downloads2mo agoHugging Face22qandeelasim13 /code-switching-codesaviours-si26-qandeel Code Switching NLP Dataset | Code Saviours SI-26 Dataset Description A word-level labelled Roman Urdu–English code-switching dataset (200 sentences, 3182 word entries). Labels: URD (Roman Urdu), ENG (English), MIX (nativized loanword). Source Sentences filtered from the Roman Urdu Data Set (Sharf, 2017, UCI ML Repository, CC BY 4.0), originally collected from e-commerce reviews, Facebook comments, and Twitter posts. Filtered for genuine… See the full description on the dataset page: https://huggingface.co/datasets/qandeelasim13/code-switching-codesaviours-si26-qandeel.text1K<n<10K0 likes11 downloads2mo agoHugging Face23Zainab-Binte-Khalid /code-switching-codesaviours-si26-zainab Roman Urdu–English Code-Switching Dataset Dataset Description This dataset contains 1,901 sentences and 21,370 word-level language labels, built to capture how Roman Urdu and English are naturally mixed together in everyday Pakistani online communication. Code-switching — blending two languages within a single sentence — is how the vast majority of Pakistanis actually write and speak online, on platforms like Twitter/X, Facebook, YouTube, Reddit, and WhatsApp. A… See the full description on the dataset page: https://huggingface.co/datasets/Zainab-Binte-Khalid/code-switching-codesaviours-si26-zainab.text10K<n<100K1 likes11 downloads2mo agoHugging Face24Saima-Manzoor /code-switching-codesaviours-si26-saima Code-Switching Urdu-English Dataset Dataset Description This dataset is created for Urdu-English code-switching language identification. It contains sentences that include Urdu words, English words, and a small number of mixed-language entries. The dataset was prepared by collecting and organizing code-switched Urdu-English sentences. Each sentence was divided into individual words, and every word was assigned a language label. The data was then converted into a… See the full description on the dataset page: https://huggingface.co/datasets/Saima-Manzoor/code-switching-codesaviours-si26-saima.text1K<n<10K0 likes11 downloads2mo agoHugging Face25Maryam657775 /code-switching-codesaviours-si26-maryam Roman Urdu & English Code-Switching Dataset Description This dataset contains 152 real-world sentences where users naturally mix Roman Urdu and English. It was built to train models to handle how 230 million Pakistanis actually communicate online. Data Collection The raw text was scraped directly from natural conversations on Pakistani Twitter/X and anonymized WhatsApp messages. Label Meanings Every single word in this dataset has… See the full description on the dataset page: https://huggingface.co/datasets/Maryam657775/code-switching-codesaviours-si26-maryam.text1K<n<10K0 likes11 downloads2mo agoHugging Face26qadeesanoor /code-switching-codesaviours-si26-qadeesanoorDataset Summary This dataset contains Roman Urdu–English code-switched sentences, the way mixed-language text actually gets written in everyday Pakistani texting, tweeting, and casual conversation (e.g. "Aaj ka din bohot busy tha, had 3 meetings back to back"). Each sentence is broken down word-by-word, and every word is tagged with a language label. No existing Roman Urdu NLP resource handles this kind of within-sentence language mixing well — this dataset is a step toward building tools… See the full description on the dataset page: https://huggingface.co/datasets/qadeesanoor/code-switching-codesaviours-si26-qadeesanoor.text1K<n<10K0 likes10 downloads2mo agoHugging Face27samaikaimran /code-switching-codesaviours-si26-samaika Code Switching NLP | Code Saviours SI-26 | Samaika About the Dataset This dataset contains Roman Urdu and English code-switching sentences collected from social media comments and online content. The purpose of this dataset is to represent the way Pakistani users naturally mix Roman Urdu and English while communicating online. Data Collection The sentences were collected from: Instagram comments YouTube comments Twitter/X The collected sentences… See the full description on the dataset page: https://huggingface.co/datasets/samaikaimran/code-switching-codesaviours-si26-samaika.text1K<n<10K0 likes10 downloads2mo agoHugging Face28WTFO /codeswitchinggated WTFO Code-Switching Speech Code-switching speech dataset prepared for automatic speech recognition training. Dataset fields audio: self-contained 16 kHz mono PCM16 FLAC bytes using the Hugging Face Audio feature text: transcript duration: audio duration in seconds (float64) Split summary Split: train Examples: 98,662 Total duration: 567422.698 seconds (157.62 hours) The Parquet shards embed the audio bytes. They do not depend on the original… See the full description on the dataset page: https://huggingface.co/datasets/WTFO/codeswitching.audioautomatic-speech-recognition10K<n<100K0 likes10 downloads2mo agoHugging Face29farabi-lab /Code_switchinggated 🇰🇿 Kazakh-Russian Code-Switching Normalization Dataset Dataset Summary Kazakh-Russian Code-Switching Normalization Dataset is a bilingual instruction-following dataset designed for identifying and rewriting Kazakh-Russian mixed-language text into clean Kazakh. The dataset focuses on informal communication, where Kazakh speakers may naturally mix Russian and Kazakh in one message. Each sample contains a prompt with code-switching, a response that identifies the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Code_switching.texttext-generationn<1K0 likes9 downloads2mo agoHugging Face30Maryam12256 /code-switching-codesaviours-si26-Maryam Code-Switching Dataset — Roman Urdu / English Dataset Description This dataset contains naturally occurring code-switched sentences that mix Roman Urdu and English, the way Pakistani speakers commonly write in everyday digital communication (texts, tweets, comments). Each sentence is broken down word-by-word, with every word labeled by language. Total sentences: 150 Total labeled word entries: ~1,478 Format: flat CSV — one row per word, with the parent sentence… See the full description on the dataset page: https://huggingface.co/datasets/Maryam12256/code-switching-codesaviours-si26-Maryam.text1K<n<10K0 likes9 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.