Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RidheshBhati /Codemixed_New Codemixed ASR Dataset Unified collection of code-mixed ASR datasets. audioautomatic-speech-recognition100K<n<1M2 likes6.1k downloads5mo agoHugging Face02BanglishRev /bangla-english-and-code-mixed-ecommerce-review-dataset BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce Description The BanglishRev dataset is the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online… See the full description on the dataset page: https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset.image1 likes1.3k downloads2y agoHugging Face03NLPC-UOM /Sinhala-English-Code-Mixed-Code-Switched-Dataset Sinhala-English-Code-Mixed-Code-Switched-Dataset This dataset contains 10,000 comments that have been annotated at the sentence level for sentiment analysis, humor detection, hate speech detection, aspect identification, and language identification. The following is the tag scheme. Sentiment - Positive, Negative, Neutral, Conflict Humor - Humorous, Non humorous Hate Speech - Hate-Inducing, Abusive, Not offensive Aspect - Network, Billing or Price, Package, Customer Service, Data… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-English-Code-Mixed-Code-Switched-Dataset.text-classification6 likes199 downloads2y agoHugging Face04guruawe /ramanv-tts-codemixedtext10K<n<100K0 likes151 downloads6d agoHugging Face05aiatums /codemixed-id-hate-speech Code-mixed Indonesian Hate Speech Dataset Manually annotated hate speech dataset for Indonesian-Javanese and Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations. texttext-classification10K<n<100K0 likes108 downloads5mo agoHugging Face06atlas-institute /code-trainer-v9-mixed code-trainer-v9-mixed 40,401-row mixed training dataset for supervised fine-tuning (SFT) in the Code-Trainer / RTPI pipeline. Used for both Qwen 14B (V9 SFT) and Gemma 26B (aggressive-full1 SFT) training. Composition Slice Source Rows (train) Purpose A -- Code generation cmndcntrlcyber/code-trainer-offsec-dataset (8K subsample) 7,074 Preserve code-gen quality B -- Tool calling glaiveai/glaive-function-calling-v2 (19K cap) ~15,125 High-density tool… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v9-mixed.text10K<n<100K0 likes88 downloads20d agoHugging Face07DatarrX /Burmese-English-Code-Mixed-Corpus 🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱ A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research. Dataset Details Organization: DatarrX Creator: Khant Sint Heinn (Kalix Louis) Number of Rows: 1,111 Language: Burmese (Unicode) & English Mix Dataset Format: .txt License: Apache 2.0 Description The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-English-Code-Mixed-Corpus.texttext-generation1K<n<10K9 likes81 downloads6mo agoHugging Face08md-nishat-008 /Code-Mixed-Sentiment-Analysis-Dataset Dataset Generation: Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.text10K<n<100K0 likes78 downloads3y agoHugging Face09Abhishek4896 /hindi-english-code-mixed-tweets-sentimenttexttext-classificationn<1K0 likes72 downloads1y agoHugging Face10SEACrowd /code_mixed_jv_idSentiment analysis and machine translation data for Javanese and Indonesian.0 likes70 downloads2y agoHugging Face11lingamvamshikrishnareddy /ramanv-tts-codemixed-voicegated0 likes70 downloads2mo agoHugging Face12PotatoHD /code-instruct-mixed Description Filtered/normalised subsets of public code-instruction datasets (Magicoder OSS-Instruct & Evol-Instruct, CodeFeedback, Glaive). The source column attributes each row to its origin; each source retains its upstream licence. Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is. Usage from datasets import load_dataset ds = load_dataset("PotatoHD/code-instruct-mixed") texttext-generation100K<n<1M0 likes51 downloads4mo agoHugging Face13atlas-institute /code-trainer-v7-mixedtext10K<n<100K0 likes45 downloads2mo agoHugging Face14md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes43 downloads3y agoHugging Face15atlas-institute /code-trainer-v8-mixedtext10K<n<100K0 likes38 downloads2mo agoHugging Face16Thanmay /belebele_hin_Latn_codemixedtextn<1K0 likes36 downloads1mo agoHugging Face17user-anto /Filtered_Dakshina_Natural_CodeMixedtext1K<n<10K0 likes34 downloads2mo agoHugging Face18autoprogrammer /amthinking_code_math_mixed_100ktext100K<n<1M0 likes33 downloads1y agoHugging Face19puttatidam /codemixed-ind-classificationtextn<1K0 likes31 downloads23d agoHugging Face20Maisha230 /bangla-english-and-code-mixed-ecommerce-review-dataset BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce Description The BanglishRev dataset is the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online e-commerce… See the full description on the dataset page: https://huggingface.co/datasets/Maisha230/bangla-english-and-code-mixed-ecommerce-review-dataset.0 likes29 downloads7mo agoHugging Face21ShrutiPatel3011 /gujarati-english-codemixed-sentiment Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset Dataset Summary This dataset addresses an empirically confirmed gap in the code-mixed sentiment analysis literature: a systematic literature review of 48 retrieved records (39 core studies) found zero existing Gujarati-English code-mixed sentiment datasets, despite Gujarati having over 55 million native speakers. This dataset provides the first Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.tabulartext-classification10K<n<100K0 likes29 downloads2mo agoHugging Face22kornwtp /codemixed-ind-classificationtextn<1K0 likes28 downloads2y agoHugging Face23SPEAK-PP /v3-v1_v2_code_mixed_syntheic_correct_noisy_pairs Sinhala Spelling Correction Dataset Dataset Description This dataset contains Sinhala text pairs for training spelling correction models. It includes: Dyslexic/Noisy sentences: Text with spelling errors, typos, and dyslexia-like mistakes Clean sentences: Corrected versions of the text Dataset Statistics Split Samples Train 37,712 Test 9,428 Total 47,140 Features dyslexic_sentence: Input text with errors (string)… See the full description on the dataset page: https://huggingface.co/datasets/SPEAK-PP/v3-v1_v2_code_mixed_syntheic_correct_noisy_pairs.texttext-generation10K<n<100K1 likes27 downloads4mo agoHugging Face24fathan /autotrain-data-code-mixed-language-identification AutoTrain Dataset for project: code-mixed-language-identification Dataset Description This dataset has been automatically processed by AutoTrain for project code-mixed-language-identification. Languages The BCP-47 code for the dataset's language is unk. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "feat_Unnamed: 0": 1104, "tokens": [ "@user", "salah", "satu", "dari"… See the full description on the dataset page: https://huggingface.co/datasets/fathan/autotrain-data-code-mixed-language-identification.token-classification0 likes26 downloads4y agoHugging Face25ar5entum /hindi-english-code-mixedThis dataset was compiled from various open sources online including some asr datasets and some percenteage of data generated using prompt engineering on generative llms. Some sources used are listed down below: https://github.com/l3cube-pune/code-mixed-nlp?tab=readme-ov-file https://github.com/piyushmakhija5/hinglishNorm https://github.com/ishan00/translation-for-code-switching-acl/tree/master text100K<n<1M0 likes25 downloads2y agoHugging Face26msislam /marc-code-mixed-small marc-code-mixed-small This dataset is based on The Multilingual Amazon Reviews Corpus. It contains German (DE), English (EN), Spanish (ES), and French (FR) languages. The labels are 0 (DE), 1 (EN), 2 (ES), and 3 (FR). Each review contains all four languages. Total number of tokens: In training set: 10195342 In test set: 842760 In validation set: 842760 text10K<n<100K0 likes21 downloads3y agoHugging Face27sudipghosh1728 /Code_Mixed_video_Complainttabulartext-classificationn<1K0 likes21 downloads2y agoHugging Face28DaliaBarua /En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset📊 En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset The En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset is a multilingual dataset of 100,000 product review texts designed for code-mixed sentiment analysis involving English, Bengali, and Roman Bengali. Each record includes: 🆔 Id 🛒 ProductId 💬 Code-Mixed-Text 💡 Sentiment The dataset captures diverse linguistic styles, authentic code-mixing, and real-world sentiment patterns from multilingual digital communication. 🌐 Text Distribution The dataset… See the full description on the dataset page: https://huggingface.co/datasets/DaliaBarua/En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset.text100K<n<1M0 likes20 downloads11mo agoHugging Face29suwaimyo /codemixed-ind-classification CodeMixed_ind_Classification Deduplicated copy of kornwtp/codemixed-ind-classification. Splits split rows train 956 textn<1K0 likes20 downloads1mo agoHugging Face30HuggMachas /en-hi-codemixed-corpustext1K<n<10K0 likes19 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.