Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RidheshBhati /Codemixed_New Codemixed ASR Dataset Unified collection of code-mixed ASR datasets. audioautomatic-speech-recognition100K<n<1M2 likes6.1k downloads5mo agoHugging Face02CodeMixBench /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/CodeMixBench/CodeMixBench.texttext-generation10K<n<100K3 likes261 downloads1y agoHugging Face03guruawe /ramanv-tts-codemixedtext10K<n<100K0 likes151 downloads4d agoHugging Face04Anvesh-Lankala /Constrained_Indic_Codemixingtext1K<n<10K0 likes135 downloads1mo agoHugging Face05lingamvamshikrishnareddy /ramanv-tts-codemixed-voicegated0 likes118 downloads2mo agoHugging Face06aiatums /codemixed-id-hate-speech Code-mixed Indonesian Hate Speech Dataset Manually annotated hate speech dataset for Indonesian-Javanese and Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations. texttext-classification10K<n<100K0 likes106 downloads5mo agoHugging Face07md-nishat-008 /Code-Mixed-Sentiment-Analysis-Dataset Dataset Generation: Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.text10K<n<100K0 likes78 downloads3y agoHugging Face08SEACrowd /code_mixed_jv_idSentiment analysis and machine translation data for Javanese and Indonesian.0 likes69 downloads2y agoHugging Face09mrzoadic /codemixaudio10K<n<100K0 likes54 downloads11mo agoHugging Face10Tanushreeeeee /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/CodeMixBench.texttext-generation10K<n<100K0 likes52 downloads10mo agoHugging Face11gentaiscool /codemixqa CodeMixQA A benchmark with high-quality human annotations, comprising 16 diverse parallel code-switched language-pair variants that span multiple geographic regions and code-switching patterns, and include both original scripts and their transliterated forms. We use SimpleQA Verified as our source dataset. We select the SimpleQA Verified, as it is a challenging evaluation set that has not been saturated yet by current models and has desirable properties such as verifiable answers… See the full description on the dataset page: https://huggingface.co/datasets/gentaiscool/codemixqa.text10K<n<100K1 likes49 downloads9mo agoHugging Face12md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes44 downloads3y agoHugging Face13user-anto /Filtered_Dakshina_Natural_CodeMixedtext1K<n<10K0 likes36 downloads2mo agoHugging Face14Thanmay /belebele_hin_Latn_codemixedtextn<1K0 likes36 downloads1mo agoHugging Face15aaditya /orca_dpo_pairs-Hinglish-Codemix Summary aaditya/orca_dpo_pairs-Hinglish-Codemix is an open source Hinglish version dataset of Intel/orca_dpo_pairs This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Citation @misc {orca_dpo_pairs-Hinglish-Codemix, author = { Pal, Ankit }… See the full description on the dataset page: https://huggingface.co/datasets/aaditya/orca_dpo_pairs-Hinglish-Codemix.text10K<n<100K1 likes35 downloads3y agoHugging Face16ColdSlim /CodeMixBench BigCodeBench-CodeMixed Dataset Description This dataset is an augmented version of BigCodeBench designed for evaluating code generation in code-mixed scenarios. It introduces multilingual variations of the prompts, primarily focusing on translating the docstrings within the complete_prompt field while keeping the code and test cases in English. This allows for assessing the ability of Large Language Models (LLMs) to handle code generation tasks where the prompt contains a… See the full description on the dataset page: https://huggingface.co/datasets/ColdSlim/CodeMixBench.text1K<n<10K0 likes31 downloads1y agoHugging Face17puttatidam /codemixed-ind-classificationtextn<1K0 likes31 downloads21d agoHugging Face18kornwtp /codemixed-ind-classificationtextn<1K0 likes28 downloads2y agoHugging Face19ShrutiPatel3011 /gujarati-english-codemixed-sentiment Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset Dataset Summary This dataset addresses an empirically confirmed gap in the code-mixed sentiment analysis literature: a systematic literature review of 48 retrieved records (39 core studies) found zero existing Gujarati-English code-mixed sentiment datasets, despite Gujarati having over 55 million native speakers. This dataset provides the first Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.tabulartext-classification10K<n<100K0 likes28 downloads2mo agoHugging Face20aaditya /databricks-dolly-15k-Hinglish-Codemix Summary aaditya/databricks-dolly-15k-Hindi is an open source Hinglish-Codemix version dataset of databricks/databricks-dolly-15k. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Original Dataset repo… See the full description on the dataset page: https://huggingface.co/datasets/aaditya/databricks-dolly-15k-Hinglish-Codemix.text10K<n<100K2 likes23 downloads3y agoHugging Face21sajalmadan0909 /mucs-hindi-english-codemix-asrgated MUCS 2021 Hindi-English Code-Mixed ASR Hindi-English code-mixed speech recognition dataset from the MUCS 2021 (Multilingual and Code-Switching ASR Challenges) subtask 2, released as spoken-tutorial recordings with Hindi-English code-mixed transcripts. Long-form recordings were sliced into per-utterance clips using the original Kaldi-style segments/text/utt2spk alignment. Dataset structure split utterances speakers audio hours avg clip len train 52,825… See the full description on the dataset page: https://huggingface.co/datasets/sajalmadan0909/mucs-hindi-english-codemix-asr.audioautomatic-speech-recognition10K<n<100K0 likes23 downloads2mo agoHugging Face22suwaimyo /codemixed-ind-classification CodeMixed_ind_Classification Deduplicated copy of kornwtp/codemixed-ind-classification. Splits split rows train 956 textn<1K0 likes23 downloads1mo agoHugging Face23Dsg2 /CodeMix CodeMix - a small finetune dataset 6k chat/response pairs, a balanced mix of: glaive-function-calling-v2 (agentic tool calling) hermes-function-calling-v1 (tool calling) CodeAlpaca-20k (coding) dolly-15k (instruct) texttext-generation100K<n<1M0 likes22 downloads4mo agoHugging Face24sudipghosh1728 /Code_Mixed_video_Complainttabulartext-classificationn<1K0 likes21 downloads2y agoHugging Face25HuggMachas /en-hi-codemixed-corpustext1K<n<10K0 likes20 downloads2y agoHugging Face26sajalmadan0909 /hindi_and_english_stt_tts_codemix_datagated Hindi and English STT/TTS Codemix Data Hinglish (Hindi-English code-mixed) speech dataset for automatic speech recognition (ASR) and text-to-speech (TTS) research. Dataset Description Each row is a timestamped speech segment clipped from conversational Hinglish audio recordings. Column Type Description text string Transcript of the speech segment (Hinglish) audio audio (16 kHz mono) Corresponding audio clip duration float32 Clip duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/sajalmadan0909/hindi_and_english_stt_tts_codemix_data.audioautomatic-speech-recognition10K<n<100K0 likes19 downloads3mo agoHugging Face27kkarhm /nepali-english-codemixed-asraudio100K<n<1M0 likes18 downloads6mo agoHugging Face28chipotswift /codemix-test-datasetThis is a test file that is used for determining whether data can be deleted once its uploaded videon<1K0 likes17 downloads2y agoHugging Face29HuggMachas /POS_hinglish_codemixed_tweetstext1K<n<10K0 likes16 downloads2y agoHugging Face30Tngarg /Codemix_tamil_english_test Dataset Card for "Codemix_tamil_english_test" More Information needed textn<1K0 likes15 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.