Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01md-nishat-008 /Code-Mixed-Sentiment-Analysis-Dataset Dataset Generation: Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.text10K<n<100K0 likes78 downloads3y agoHugging Face02Abhishek4896 /hindi-english-code-mixed-tweets-sentimenttexttext-classificationn<1K0 likes72 downloads1y agoHugging Face03md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes43 downloads3y agoHugging Face04ShrutiPatel3011 /gujarati-english-codemixed-sentiment Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset Dataset Summary This dataset addresses an empirically confirmed gap in the code-mixed sentiment analysis literature: a systematic literature review of 48 retrieved records (39 core studies) found zero existing Gujarati-English code-mixed sentiment datasets, despite Gujarati having over 55 million native speakers. This dataset provides the first Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.tabulartext-classification10K<n<100K0 likes29 downloads2mo agoHugging Face05kornwtp /codemixed-ind-classificationtextn<1K0 likes28 downloads2y agoHugging Face06sudipghosh1728 /Code_Mixed_video_Complainttabulartext-classificationn<1K0 likes21 downloads2y agoHugging Face07DaliaBarua /En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset📊 En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset The En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset is a multilingual dataset of 100,000 product review texts designed for code-mixed sentiment analysis involving English, Bengali, and Roman Bengali. Each record includes: 🆔 Id 🛒 ProductId 💬 Code-Mixed-Text 💡 Sentiment The dataset captures diverse linguistic styles, authentic code-mixing, and real-world sentiment patterns from multilingual digital communication. 🌐 Text Distribution The dataset… See the full description on the dataset page: https://huggingface.co/datasets/DaliaBarua/En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset.text100K<n<1M0 likes20 downloads11mo agoHugging Face08taha-alnasser /ArzEn-CodeMixed ArzEn-CodeMixed: English-Arabic Intra-Sentential Code-Mixed Translation Dataset This dataset contains aligned English, Arabic, and English-Arabic as well as Arabic-English intra-sentential code-mixed sentences, created for evaluating large LLMs on translating code-mixed text. It is constructed from the ArzEn-MultiGenre parallel corpus and enhanced with carefully prompted and human-validated code-mixed variants. Dataset Details Instances: Each entry contains:… See the full description on the dataset page: https://huggingface.co/datasets/taha-alnasser/ArzEn-CodeMixed.tabulartranslation1K<n<10K2 likes15 downloads1y agoHugging Face09adealvii /codemixed-synthetic-sarc-11ktext10K<n<100K0 likes2 downloads1y agoHugging Face10nlp-dataset /Sinhala-English-Code-Mixed-Code-Switched-Hotel-Reviews-Dataset Sinhala-English Hospitality Review Corpus This dataset contains 4,133 user reviews collected from Facebook hotel pages across Sri Lanka, annotated at the sentence level for sentiment analysis. The reviews come from 400 hotels across 144 tourist destinations, including Colombo, Kandy, Anuradhapura, Negombo, and other major regions. Annotation Details All samples were manually annotated by the author to ensure consistency in sentiment labeling. We welcome volunteers… See the full description on the dataset page: https://huggingface.co/datasets/nlp-dataset/Sinhala-English-Code-Mixed-Code-Switched-Hotel-Reviews-Dataset.tabular1K<n<10K0 likes1 downloads4h agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.