Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CodeMixBench /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/CodeMixBench/CodeMixBench.texttext-generation10K<n<100K3 likes270 downloads1y agoHugging Face02md-nishat-008 /Code-Mixed-Sentiment-Analysis-Dataset Dataset Generation: Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.text10K<n<100K0 likes69 downloads3y agoHugging Face03Tanushreeeeee /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/CodeMixBench.texttext-generation10K<n<100K0 likes55 downloads10mo agoHugging Face04md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes39 downloads3y agoHugging Face05kornwtp /codemixed-ind-classificationtextn<1K0 likes31 downloads2y agoHugging Face06ShrutiPatel3011 /gujarati-english-codemixed-sentiment Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset Dataset Summary This dataset addresses an empirically confirmed gap in the code-mixed sentiment analysis literature: a systematic literature review of 48 retrieved records (39 core studies) found zero existing Gujarati-English code-mixed sentiment datasets, despite Gujarati having over 55 million native speakers. This dataset provides the first Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.tabulartext-classification10K<n<100K0 likes25 downloads2mo agoHugging Face07taha-alnasser /ArzEn-CodeMixed ArzEn-CodeMixed: English-Arabic Intra-Sentential Code-Mixed Translation Dataset This dataset contains aligned English, Arabic, and English-Arabic as well as Arabic-English intra-sentential code-mixed sentences, created for evaluating large LLMs on translating code-mixed text. It is constructed from the ArzEn-MultiGenre parallel corpus and enhanced with carefully prompted and human-validated code-mixed variants. Dataset Details Instances: Each entry contains:… See the full description on the dataset page: https://huggingface.co/datasets/taha-alnasser/ArzEn-CodeMixed.tabulartranslation1K<n<10K2 likes17 downloads1y agoHugging Face08sudipghosh1728 /Code_Mixed_video_Complainttabulartext-classificationn<1K0 likes15 downloads2y agoHugging Face09Tngarg /Codemix_tamil_englishtext10K<n<100K0 likes12 downloads3y agoHugging Face10karanverma19 /CodeMix_Query_Normalization_India CodeMix Query Normalization (India) Overview This dataset contains code-mixed user queries from Indian contexts, primarily in Hinglish and Punjabi, normalized into clean English. It reflects how users naturally communicate in real-world scenarios by mixing local languages with English. Features 100 high-quality samples Code-mixed queries (Hinglish, Punjabi) Clean normalized English outputs Real-world, informal user language patterns Covers domains such as… See the full description on the dataset page: https://huggingface.co/datasets/karanverma19/CodeMix_Query_Normalization_India.textn<1K0 likes6 downloads6mo agoHugging Face11karanverma19 /Advanced_CodeMix_Normalization_Dataset_India Evaluation & Benchmarking To validate dataset usefulness, normalization accuracy can be evaluated using: Exact Match Accuracy BLEU Score for text similarity Human evaluation for real-world correctness This dataset is designed to improve performance of multilingual NLP systems in handling noisy, code-mixed Indian queries. Data Transformation Approach The dataset was created by transforming real-world code-mixed queries into structured English. Variations include:… See the full description on the dataset page: https://huggingface.co/datasets/karanverma19/Advanced_CodeMix_Normalization_Dataset_India.textn<1K0 likes6 downloads6mo agoHugging Face12teppap /code-mix-thai-engtabular1K<n<10K0 likes3 downloads7mo agoHugging Face13chloeclerc17 /code-mixing_safetytextn<1K0 likes3 downloads6mo agoHugging Face14adealvii /codemixed-synthetic-sarc-11ktext10K<n<100K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.