Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /SentimentAnalysisHindi SentimentAnalysisHindi An MTEB dataset Massive Text Embedding Benchmark Hindi Sentiment Analysis Dataset Task category t2c Domains Reviews, Written Reference https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("SentimentAnalysisHindi") evaluator = mteb.MTEB([task]) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SentimentAnalysisHindi.texttext-classification1K<n<10K0 likes1.3k downloads1y agoHugging Face02aisingapore /NLU-Sentiment-Analysisgated SEA Sentiment Analysis SEA Sentiment Analysis evaluates a model's ability to identify the sentiment polarity of a text. It is sampled from NusaX for Indonesian, Javanese, and Sundanese, IndicSentiment for Tamil, Wisesight Sentiment for Thai, and UIT-VSFC for Vietnamese. Supported Tasks and Leaderboards SEA Sentiment Analysis is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Sentiment-Analysis.texttext-generation1K<n<10K0 likes1.1k downloads10mo agoHugging Face03winvoker /turkish-sentiment-analysis-dataset Dataset This dataset contains positive , negative and notr sentences from several data sources given in the references. In the most sentiment models , there are only two labels; positive and negative. However , user input can be totally notr sentence. For such cases there were no data I could find. Therefore I created this dataset with 3 class. Positive and negative sentences are listed below. Notr examples are extraced from turkish wiki dump. In addition, added some random text… See the full description on the dataset page: https://huggingface.co/datasets/winvoker/turkish-sentiment-analysis-dataset.texttext-classification100K<n<1M52 likes545 downloads3y agoHugging Face04evalitahf /sentiment_analysisSENTIPOLC 2016 dataset The SENTIPOLC 2016 dataset contains 9410 tweets annotated for subjectivity, overall and literal polarity, and irony. The dataset has been created and used in the context of the SENTIPOLC 2016 task (http://www.di.unito.it/~tutreeb/sentipolc-evalita16/index.html), organized as part of the EVALITA 2016 evaluation campaign. Original files available here: https://live.european-language-grid.eu/catalogue/corpus/7479/download/ If you find this dataset useful please cite:… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/sentiment_analysis.texttext-classification1K<n<10K0 likes513 downloads2y agoHugging Face05Sp1786 /multiclass-sentiment-analysis-dataset Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Sp1786/multiclass-sentiment-analysis-dataset.tabulartext-classification10K<n<100K29 likes469 downloads3y agoHugging Face06hugginglearners /amazon-reviews-sentiment-analysis Dataset Card for amazon reviews for sentiment analysis Dataset Summary One of the most important problems in e-commerce is the correct calculation of the points given to after-sales products. The solution to this problem is to provide greater customer satisfaction for the e-commerce site, product prominence for sellers, and a seamless shopping experience for buyers. Another problem is the correct ordering of the comments given to the products. The prominence of misleading… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/amazon-reviews-sentiment-analysis.tabular1K<n<10K6 likes303 downloads4y agoHugging Face07themohal /saraiki-sentiment-analysis-datasetgated Saraiki Sentiment Analysis Dataset Sentiment-annotated Saraiki text for sentiment classification and low-resource NLP. About Saraiki Saraiki (سرائیکی) is an Indo-Aryan language spoken primarily in Pakistan, particularly across southern Punjab and neighboring regions. It has a rich linguistic, literary, and cultural heritage, with its own vocabulary, grammatical structures, and regional varieties. Despite its significant speaker community, Saraiki remains a… See the full description on the dataset page: https://huggingface.co/datasets/themohal/saraiki-sentiment-analysis-dataset.text1K<n<10K0 likes280 downloads11h agoHugging Face08ParsiAI /snappfood-sentiment-analysistexttext-classification10K<n<100K7 likes278 downloads2y agoHugging Face09ParsiAI /digikala-sentiment-analysistabulartext-classification1K<n<10K3 likes223 downloads2y agoHugging Face10tanaos /synthetic-sentiment-analysis-dataset-v1 Tanaos Sentiment Analysis Training Dataset This dataset was created synthetically by Tanaos with the Artifex Python library. The dataset is designed to train and evaluate sentiment analysis systems — models that classify the sentiment expressed in text as one of five possible categories: very_negative, negative, neutral, positive or very_positive. It can be used to build sentiment analysis models for various applications, such as customer feedback analysis, social media… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-sentiment-analysis-dataset-v1.texttext-classification10K<n<100K0 likes198 downloads10mo agoHugging Face11AiresPucrs /sentiment-analysis-pt Sentiment Analysis PT (Teeny-Tiny Castle) This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research. How to Use from datasets import load_dataset dataset = load_dataset("AiresPucrs/sentiment-analysis-pt", split = 'train') texttext-classification10K<n<100K5 likes193 downloads2y agoHugging Face12yassiracharki /Amazon_Reviews_Binary_for_Sentiment_Analysis Dataset Card for Dataset Name The Amazon reviews polarity dataset is constructed by taking review score 1 and 2 as negative, and 4 and 5 as positive. Samples of score 3 is ignored. In the dataset, class 1 is the negative and class 2 is the positive. Each class has 1,800,000 training samples and 200,000 testing samples. Dataset Details Dataset Description The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_Binary_for_Sentiment_Analysis.texttext-classification1M<n<10M0 likes183 downloads2y agoHugging Face13jvanz /portuguese_sentiment_analysisThis dataset is based on the dataset originally posted in Kaggle text1M<n<10M14 likes173 downloads4y agoHugging Face14Akash190104 /bengali_sentiment_analysis Bengali Sentiment Analysis Context The dataset contains 3307 Negative reviews and 8500 Positive reviews collected and manually annotated from Youtube Bengali drama. Positive_Label=1 and Negative_Label=0 Acknowledgements Sazzed, Salim (2021), “Bangla ( Bengali ) sentiment analysis classification benchmark dataset corpus”, Mendeley Data, V4, doi: 10.17632/p6zc7krs37.4 texttext-classification10K<n<100K1 likes171 downloads2y agoHugging Face15CATIE-AQ /allocine_fr_prompt_sentiment_analysis allocine_fr_prompt_sentiment_analysis Summary allocine_fr_prompt_sentiment_analysis is a subset of the Dataset of French Prompts (DFP).It contains 5,600,000 rows that can be used for a binary sentiment analysis task.The original data (without prompts) comes from the dataset allocine by Blard.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by Muennighoff et al.… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/allocine_fr_prompt_sentiment_analysis.texttext-classification1M<n<10M0 likes166 downloads1y agoHugging Face16sudhanshusinghaiml /airlines-sentiment-analysistext10K<n<100K0 likes152 downloads2y agoHugging Face17DevBM /Sentiment-Analysis-Text-from-Yelp-Imdb-Amazontext1K<n<10K0 likes151 downloads2y agoHugging Face18Sadiksha /sentiment_analysis_data Dataset Card for "sentiment_analysis_data" More Information needed text10K<n<100K0 likes143 downloads4y agoHugging Face19yassiracharki /Amazon_Reviews_for_Sentiment_Analysis_fine_grained_5_classes Dataset Card for Dataset Name The Amazon reviews full score dataset is constructed by randomly taking 600,000 training samples and 130,000 testing samples for each review score from 1 to 5. In total there are 3,000,000 trainig samples and 650,000 testing samples. Dataset Details Dataset Description The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3 columns in them, corresponding to class index (1 to 5)… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_for_Sentiment_Analysis_fine_grained_5_classes.texttext-classification1M<n<10M4 likes139 downloads2y agoHugging Face20maydogan /Turkish_SentimentAnalysis_TRSAv1TRSAv1 (Turkish Sentiment Analysis Version 1) Dataset This data set has been produced to contribute to Turkish NLP studies. The dataset consists of a total of 150 thousand samples, 50 thousand negative, 50 thousand positive, and 50 thousand neutral. It can be used in text classification and sentiment analysis studies by citing the related study. Related Work Aydoğan M, Kocaman V. TRSAv1: A new benchmark dataset for classifying user reviews on Turkish e-commerce websites. Journal of… See the full description on the dataset page: https://huggingface.co/datasets/maydogan/Turkish_SentimentAnalysis_TRSAv1.texttext-classification100K<n<1M16 likes134 downloads2y agoHugging Face21SaguaroCapital /sentiment-analysis-in-commodity-market-gold Dataset Card for Sentiment Analysis of Commodity News (Gold) This is a news dataset for the commodity market which has been manually annotated for 10,000+ news headlines across multiple dimensions into various classes. The dataset has been sampled from a period of 20+ years (2000-2021). The dataset was curated by Ankur Sinha and Tanmay Khandait and is detailed in their paper "Impact of News on the Commodity Market: Dataset and Results." It is currently published by the authors on… See the full description on the dataset page: https://huggingface.co/datasets/SaguaroCapital/sentiment-analysis-in-commodity-market-gold.tabulartext-classification10K<n<100K6 likes127 downloads2y agoHugging Face22Sanatbek /aspect-based-sentiment-analysis-uzbekdocumenttext-classification1K<n<10K4 likes125 downloads3y agoHugging Face23gm84 /sentiment-analysis-monitoring Sentiment Analysis Monitoring & Retraining Telemetry Repository di telemetria e dati per la pipeline MLOps del modello di Sentiment Analysis (Twitter-RoBERTa). Configurazioni disponibili: metrics: Storico temporale dei batch di monitoraggio, andamento dell'indice F1 e distribuzione del sentiment (% Positivi, Negativi, Neutri). data: Buffer di accumulo e snapshot utilizzati per il retraining continuo del modello. tabularn<1K0 likes111 downloads4d agoHugging Face24elmurod1202 /uzbek-sentiment-analysis uzbek-sentiment-analysis Sentiment analysis in the Uzbek language and new Datasets of Uzbek App reviews for Sentiment Classification Feel free to use the dataset and the tools presented in this project, a paper about more details on creation and usage here. If you find it useful, please make sure to cite the paper: @inproceedings{kuriyozov2019deep, author = {Kuriyozov, Elmurod and Matlatipov, Sanatbek and Alonso, Miguel A and Gómez-Rodríguez, Carlos}, title = {Deep… See the full description on the dataset page: https://huggingface.co/datasets/elmurod1202/uzbek-sentiment-analysis.image10K<n<100K5 likes97 downloads4y agoHugging Face25DynamicSuperb /Sentiment_Analysis_SLUE-VoxCelebaudio1K<n<10K1 likes96 downloads2y agoHugging Face26yuncongli /chat-sentiment-analysis A Sentiment Analsysis Dataset for Finetuning Large Models in Chat-style More details can be found at https://github.com/l294265421/chat-sentiment-analysis Supported Tasks Aspect Term Extraction (ATE) Opinion Term Extraction (OTE) Aspect Term-Opinion Term Pair Extraction (AOPE) Aspect term, Sentiment, Opinion term Triplet Extraction (ASOTE) Aspect Category Detection (ACD) Aspect Category-Sentiment Pair Extraction (ACSA) Aspect-Category-Opinion-Sentiment (ACOS) Quadruple… See the full description on the dataset page: https://huggingface.co/datasets/yuncongli/chat-sentiment-analysis.text10K<n<100K9 likes88 downloads4y agoHugging Face27MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-anger Dataset Summary Synthetic Persian Chatbot Conversational SA – Anger is a Persian (Farsi) dataset created for the Classification task, with a focus on detecting the emotion "anger" in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s): Classification… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-anger.text1K<n<10K0 likes88 downloads1y agoHugging Face28MCINext /chatbot-conversational-sentiment-analysis-tone-user-classification Dataset Summary Synthetic Persian Chatbot Conversational SA – User Tone Classification(SynPerChatbotConvSAToneUserClassification) is a Persian (Farsi) dataset created for the Classification task. It focuses on identifying the user’s conversational tone—formal, casual, or childish—in emotionally rich chatbot interactions. This dataset is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini. Language(s): Persian (Farsi)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/chatbot-conversational-sentiment-analysis-tone-user-classification.text1K<n<10K0 likes88 downloads1y agoHugging Face29sinhala-nlp /sinhala-sentiment-analysistext1K<n<10K0 likes86 downloads2y agoHugging Face30LYTinn /sentiment-analysis-tweettext10K<n<100K1 likes85 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.