Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RidheshBhati /Codemixed_New Codemixed ASR Dataset Unified collection of code-mixed ASR datasets. audioautomatic-speech-recognition100K<n<1M2 likes6.1k downloads5mo agoHugging Face02aiatums /codemixed-id-hate-speech Code-mixed Indonesian Hate Speech Dataset Manually annotated hate speech dataset for Indonesian-Javanese and Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations. texttext-classification10K<n<100K0 likes105 downloads5mo agoHugging Face03DatarrX /Burmese-English-Code-Mixed-Corpus 🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱ A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research. Dataset Details Organization: DatarrX Creator: Khant Sint Heinn (Kalix Louis) Number of Rows: 1,111 Language: Burmese (Unicode) & English Mix Dataset Format: .txt License: Apache 2.0 Description The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-English-Code-Mixed-Corpus.texttext-generation1K<n<10K9 likes82 downloads6mo agoHugging Face04atlas-institute /code-trainer-v9-mixed code-trainer-v9-mixed 40,401-row mixed training dataset for supervised fine-tuning (SFT) in the Code-Trainer / RTPI pipeline. Used for both Qwen 14B (V9 SFT) and Gemma 26B (aggressive-full1 SFT) training. Composition Slice Source Rows (train) Purpose A -- Code generation cmndcntrlcyber/code-trainer-offsec-dataset (8K subsample) 7,074 Preserve code-gen quality B -- Tool calling glaiveai/glaive-function-calling-v2 (19K cap) ~15,125 High-density tool… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v9-mixed.text10K<n<100K0 likes77 downloads21d agoHugging Face05md-nishat-008 /Code-Mixed-Sentiment-Analysis-Dataset Dataset Generation: Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.text10K<n<100K0 likes73 downloads3y agoHugging Face06Abhishek4896 /hindi-english-code-mixed-tweets-sentimenttexttext-classificationn<1K0 likes69 downloads1y agoHugging Face07PotatoHD /code-instruct-mixed Description Filtered/normalised subsets of public code-instruction datasets (Magicoder OSS-Instruct & Evol-Instruct, CodeFeedback, Glaive). The source column attributes each row to its origin; each source retains its upstream licence. Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is. Usage from datasets import load_dataset ds = load_dataset("PotatoHD/code-instruct-mixed") texttext-generation100K<n<1M0 likes50 downloads4mo agoHugging Face08md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes44 downloads3y agoHugging Face09Thanmay /belebele_hin_Latn_codemixedtextn<1K0 likes36 downloads1mo agoHugging Face10atlas-institute /code-trainer-v7-mixedtext10K<n<100K0 likes33 downloads2mo agoHugging Face11autoprogrammer /amthinking_code_math_mixed_100ktext100K<n<1M0 likes32 downloads1y agoHugging Face12user-anto /Filtered_Dakshina_Natural_CodeMixedtext1K<n<10K0 likes31 downloads2mo agoHugging Face13puttatidam /codemixed-ind-classificationtextn<1K0 likes31 downloads24d agoHugging Face14ar5entum /hindi-english-code-mixedThis dataset was compiled from various open sources online including some asr datasets and some percenteage of data generated using prompt engineering on generative llms. Some sources used are listed down below: https://github.com/l3cube-pune/code-mixed-nlp?tab=readme-ov-file https://github.com/piyushmakhija5/hinglishNorm https://github.com/ishan00/translation-for-code-switching-acl/tree/master text100K<n<1M0 likes30 downloads2y agoHugging Face15kornwtp /codemixed-ind-classificationtextn<1K0 likes29 downloads2y agoHugging Face16atlas-institute /code-trainer-v8-mixedtext10K<n<100K0 likes29 downloads2mo agoHugging Face17ShrutiPatel3011 /gujarati-english-codemixed-sentiment Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset Dataset Summary This dataset addresses an empirically confirmed gap in the code-mixed sentiment analysis literature: a systematic literature review of 48 retrieved records (39 core studies) found zero existing Gujarati-English code-mixed sentiment datasets, despite Gujarati having over 55 million native speakers. This dataset provides the first Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.tabulartext-classification10K<n<100K0 likes29 downloads2mo agoHugging Face18SPEAK-PP /v3-v1_v2_code_mixed_syntheic_correct_noisy_pairs Sinhala Spelling Correction Dataset Dataset Description This dataset contains Sinhala text pairs for training spelling correction models. It includes: Dyslexic/Noisy sentences: Text with spelling errors, typos, and dyslexia-like mistakes Clean sentences: Corrected versions of the text Dataset Statistics Split Samples Train 37,712 Test 9,428 Total 47,140 Features dyslexic_sentence: Input text with errors (string)… See the full description on the dataset page: https://huggingface.co/datasets/SPEAK-PP/v3-v1_v2_code_mixed_syntheic_correct_noisy_pairs.texttext-generation10K<n<100K1 likes26 downloads4mo agoHugging Face19HuggMachas /en-hi-codemixed-corpustext1K<n<10K0 likes20 downloads2y agoHugging Face20sudipghosh1728 /Code_Mixed_video_Complainttabulartext-classificationn<1K0 likes20 downloads2y agoHugging Face21DaliaBarua /En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset📊 En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset The En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset is a multilingual dataset of 100,000 product review texts designed for code-mixed sentiment analysis involving English, Bengali, and Roman Bengali. Each record includes: 🆔 Id 🛒 ProductId 💬 Code-Mixed-Text 💡 Sentiment The dataset captures diverse linguistic styles, authentic code-mixing, and real-world sentiment patterns from multilingual digital communication. 🌐 Text Distribution The dataset… See the full description on the dataset page: https://huggingface.co/datasets/DaliaBarua/En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset.text100K<n<1M0 likes20 downloads11mo agoHugging Face22geodesic-research /debug-mixed-rlhf-codetextn<1K0 likes19 downloads8mo agoHugging Face23msislam /marc-code-mixed-small marc-code-mixed-small This dataset is based on The Multilingual Amazon Reviews Corpus. It contains German (DE), English (EN), Spanish (ES), and French (FR) languages. The labels are 0 (DE), 1 (EN), 2 (ES), and 3 (FR). Each review contains all four languages. Total number of tokens: In training set: 10195342 In test set: 842760 In validation set: 842760 text10K<n<100K0 likes18 downloads3y agoHugging Face24Shyyamsh /nep_eng_code-mixed_asr_datasetaudio1K<n<10K0 likes18 downloads2mo agoHugging Face25vnikitin /uk-en-code-mixed-asr-2hgated uk-en-code-mixed-asr-2h A 2-hour dataset of Ukrainian-English code-mixed speech for automatic speech recognition, recorded by a single male speaker across 16 speaking-style personas covering software engineering domains (.NET, React, DevOps, project management). Samples: 448 Total duration: 2h 6m 35s Language: 446 mixed (code-mixed) + 2 uk (monolingual) Speaker: 1 male voice, 16 personas (varied pacing, formality, topic; 15 map to Valera, 1 to ValeraAlt) Audio format: OGG Opus… See the full description on the dataset page: https://huggingface.co/datasets/vnikitin/uk-en-code-mixed-asr-2h.audioautomatic-speech-recognitionn<1K0 likes17 downloads3mo agoHugging Face26nlpctx /telugu-qa-codemixed Telugu QA Paraphrases A synthetic multilingual query-rewriting dataset for evaluating retrieval robustness under Telugu-English code mixing. Dataset Description This dataset extends an existing Telugu QA dataset by generating multiple query variants with increasing levels of Telugu-English code mixing. Each example contains: question : Original English question answer : Ground-truth answer level_0 : English paraphrase level_1 : Light Telugu-English code mixing… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/telugu-qa-codemixed.textquestion-answering1K<n<10K0 likes17 downloads4mo agoHugging Face27suwaimyo /codemixed-ind-classification CodeMixed_ind_Classification Deduplicated copy of kornwtp/codemixed-ind-classification. Splits split rows train 956 textn<1K0 likes17 downloads1mo agoHugging Face28fayez94 /code-mixed-bangla-english-asraudio1K<n<10K0 likes16 downloads2y agoHugging Face29HuggMachas /POS_hinglish_codemixed_tweetstext1K<n<10K0 likes16 downloads2y agoHugging Face30taha-alnasser /ArzEn-CodeMixed ArzEn-CodeMixed: English-Arabic Intra-Sentential Code-Mixed Translation Dataset This dataset contains aligned English, Arabic, and English-Arabic as well as Arabic-English intra-sentential code-mixed sentences, created for evaluating large LLMs on translating code-mixed text. It is constructed from the ArzEn-MultiGenre parallel corpus and enhanced with carefully prompted and human-validated code-mixed variants. Dataset Details Instances: Each entry contains:… See the full description on the dataset page: https://huggingface.co/datasets/taha-alnasser/ArzEn-CodeMixed.tabulartranslation1K<n<10K2 likes16 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.