Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RidheshBhati /Codemixed_New Codemixed ASR Dataset Unified collection of code-mixed ASR datasets. audioautomatic-speech-recognition100K<n<1M2 likes5.8k downloads5mo agoHugging Face02CodeMixBench /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/CodeMixBench/CodeMixBench.texttext-generation10K<n<100K3 likes270 downloads1y agoHugging Face03Anvesh-Lankala /Constrained_Indic_Codemixingtext1K<n<10K0 likes124 downloads2mo agoHugging Face04aiatums /codemixed-id-hate-speech Code-mixed Indonesian Hate Speech Dataset Manually annotated hate speech dataset for Indonesian-Javanese and Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations. texttext-classification10K<n<100K0 likes108 downloads5mo agoHugging Face05md-nishat-008 /Code-Mixed-Sentiment-Analysis-Dataset Dataset Generation: Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.text10K<n<100K0 likes69 downloads3y agoHugging Face06Tanushreeeeee /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/CodeMixBench.texttext-generation10K<n<100K0 likes55 downloads10mo agoHugging Face07mrzoadic /codemixaudio10K<n<100K0 likes40 downloads11mo agoHugging Face08md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes39 downloads3y agoHugging Face09aaditya /orca_dpo_pairs-Hinglish-Codemix Summary aaditya/orca_dpo_pairs-Hinglish-Codemix is an open source Hinglish version dataset of Intel/orca_dpo_pairs This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Citation @misc {orca_dpo_pairs-Hinglish-Codemix, author = { Pal, Ankit }… See the full description on the dataset page: https://huggingface.co/datasets/aaditya/orca_dpo_pairs-Hinglish-Codemix.text10K<n<100K1 likes34 downloads3y agoHugging Face10Thanmay /belebele_hin_Latn_codemixedtextn<1K0 likes34 downloads1mo agoHugging Face11gentaiscool /codemixqa CodeMixQA A benchmark with high-quality human annotations, comprising 16 diverse parallel code-switched language-pair variants that span multiple geographic regions and code-switching patterns, and include both original scripts and their transliterated forms. We use SimpleQA Verified as our source dataset. We select the SimpleQA Verified, as it is a challenging evaluation set that has not been saturated yet by current models and has desirable properties such as verifiable answers… See the full description on the dataset page: https://huggingface.co/datasets/gentaiscool/codemixqa.text10K<n<100K1 likes33 downloads9mo agoHugging Face12kornwtp /codemixed-ind-classificationtextn<1K0 likes31 downloads2y agoHugging Face13puttatidam /codemixed-ind-classificationtextn<1K0 likes31 downloads25d agoHugging Face14user-anto /Filtered_Dakshina_Natural_CodeMixedtext1K<n<10K0 likes29 downloads2mo agoHugging Face15ColdSlim /CodeMixBench BigCodeBench-CodeMixed Dataset Description This dataset is an augmented version of BigCodeBench designed for evaluating code generation in code-mixed scenarios. It introduces multilingual variations of the prompts, primarily focusing on translating the docstrings within the complete_prompt field while keeping the code and test cases in English. This allows for assessing the ability of Large Language Models (LLMs) to handle code generation tasks where the prompt contains a… See the full description on the dataset page: https://huggingface.co/datasets/ColdSlim/CodeMixBench.text1K<n<10K0 likes25 downloads1y agoHugging Face16ShrutiPatel3011 /gujarati-english-codemixed-sentiment Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset Dataset Summary This dataset addresses an empirically confirmed gap in the code-mixed sentiment analysis literature: a systematic literature review of 48 retrieved records (39 core studies) found zero existing Gujarati-English code-mixed sentiment datasets, despite Gujarati having over 55 million native speakers. This dataset provides the first Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.tabulartext-classification10K<n<100K0 likes25 downloads2mo agoHugging Face17Dsg2 /CodeMix CodeMix - a small finetune dataset 6k chat/response pairs, a balanced mix of: glaive-function-calling-v2 (agentic tool calling) hermes-function-calling-v1 (tool calling) CodeAlpaca-20k (coding) dolly-15k (instruct) texttext-generation100K<n<1M0 likes24 downloads4mo agoHugging Face18sajalmadan0909 /mucs-hindi-english-codemix-asrgated MUCS 2021 Hindi-English Code-Mixed ASR Hindi-English code-mixed speech recognition dataset from the MUCS 2021 (Multilingual and Code-Switching ASR Challenges) subtask 2, released as spoken-tutorial recordings with Hindi-English code-mixed transcripts. Long-form recordings were sliced into per-utterance clips using the original Kaldi-style segments/text/utt2spk alignment. Dataset structure split utterances speakers audio hours avg clip len train 52,825… See the full description on the dataset page: https://huggingface.co/datasets/sajalmadan0909/mucs-hindi-english-codemix-asr.audioautomatic-speech-recognition10K<n<100K0 likes24 downloads2mo agoHugging Face19HuggMachas /en-hi-codemixed-corpustext1K<n<10K0 likes21 downloads2y agoHugging Face20aaditya /databricks-dolly-15k-Hinglish-Codemix Summary aaditya/databricks-dolly-15k-Hindi is an open source Hinglish-Codemix version dataset of databricks/databricks-dolly-15k. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Original Dataset repo… See the full description on the dataset page: https://huggingface.co/datasets/aaditya/databricks-dolly-15k-Hinglish-Codemix.text10K<n<100K2 likes19 downloads3y agoHugging Face21sajalmadan0909 /hindi_and_english_stt_tts_codemix_datagated Hindi and English STT/TTS Codemix Data Hinglish (Hindi-English code-mixed) speech dataset for automatic speech recognition (ASR) and text-to-speech (TTS) research. Dataset Description Each row is a timestamped speech segment clipped from conversational Hinglish audio recordings. Column Type Description text string Transcript of the speech segment (Hinglish) audio audio (16 kHz mono) Corresponding audio clip duration float32 Clip duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/sajalmadan0909/hindi_and_english_stt_tts_codemix_data.audioautomatic-speech-recognition10K<n<100K0 likes19 downloads3mo agoHugging Face22nlpctx /telugu-qa-codemixed Telugu QA Paraphrases A synthetic multilingual query-rewriting dataset for evaluating retrieval robustness under Telugu-English code mixing. Dataset Description This dataset extends an existing Telugu QA dataset by generating multiple query variants with increasing levels of Telugu-English code mixing. Each example contains: question : Original English question answer : Ground-truth answer level_0 : English paraphrase level_1 : Light Telugu-English code mixing… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/telugu-qa-codemixed.textquestion-answering1K<n<10K0 likes18 downloads4mo agoHugging Face23taha-alnasser /ArzEn-CodeMixed ArzEn-CodeMixed: English-Arabic Intra-Sentential Code-Mixed Translation Dataset This dataset contains aligned English, Arabic, and English-Arabic as well as Arabic-English intra-sentential code-mixed sentences, created for evaluating large LLMs on translating code-mixed text. It is constructed from the ArzEn-MultiGenre parallel corpus and enhanced with carefully prompted and human-validated code-mixed variants. Dataset Details Instances: Each entry contains:… See the full description on the dataset page: https://huggingface.co/datasets/taha-alnasser/ArzEn-CodeMixed.tabulartranslation1K<n<10K2 likes17 downloads1y agoHugging Face24fayez94 /code-mixed-bangla-english-asraudio1K<n<10K0 likes16 downloads2y agoHugging Face25HuggMachas /POS_hinglish_codemixed_tweetstext1K<n<10K0 likes15 downloads2y agoHugging Face26sudipghosh1728 /Code_Mixed_video_Complainttabulartext-classificationn<1K0 likes15 downloads2y agoHugging Face27kkarhm /nepali-english-codemixed-asraudio100K<n<1M0 likes14 downloads6mo agoHugging Face28suwaimyo /codemixed-ind-classification CodeMixed_ind_Classification Deduplicated copy of kornwtp/codemixed-ind-classification. Splits split rows train 956 textn<1K0 likes14 downloads1mo agoHugging Face29Tngarg /Codemix_tamil_englishtext10K<n<100K0 likes12 downloads3y agoHugging Face30atx-labs /marathi-codemix-qagated Marathi Minglish QA ~1.09M synthetic Question–Answer pairs in code-mixed Romanized Marathi (Minglish), generated from Marathi Wikipedia articles. Designed for pretraining and SFT of Marathi-aware Small Language Models that should understand and generate the way Marathi is commonly written online — Roman-script Marathi naturally mixed with English terms. Example Question: Yashwant Dev kon hote exactly — sangeetkar, kavi, ki donhi? Answer: Yashwant Dev he… See the full description on the dataset page: https://huggingface.co/datasets/atx-labs/marathi-codemix-qa.tabulartext-generation1M<n<10M1 likes12 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.