datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-sentiment-analysis-dataset
Dataset
This dataset contains positive , negative and notr sentences from several data sources given in the references. In the most sentiment models , there are only two labels; positive and negative. However , user input can be totally notr sentence. For such cases there were no data I could find. Therefore I created this dataset with 3 class. Positive and negative sentences are listed below. Notr examples are extraced from turkish wiki dump. In addition, added some random text… See the full description on the dataset page: https://huggingface.co/datasets/winvoker/turkish-sentiment-analysis-dataset.multiclass-sentiment-analysis-dataset
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Sp1786/multiclass-sentiment-analysis-dataset.amazon-reviews-sentiment-analysis
Dataset Card for amazon reviews for sentiment analysis
Dataset Summary
One of the most important problems in e-commerce is the correct calculation of the points given to after-sales products. The solution to this problem is to provide greater customer satisfaction for the e-commerce site, product prominence for sellers, and a seamless shopping experience for buyers. Another problem is the correct ordering of the comments given to the products. The prominence of misleading… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/amazon-reviews-sentiment-analysis.snappfood-sentiment-analysisdigikala-sentiment-analysissynthetic-sentiment-analysis-dataset-v1
Tanaos Sentiment Analysis Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate sentiment analysis systems — models that classify the sentiment expressed in text as one of five possible categories: very_negative, negative, neutral, positive or very_positive. It can be used to build sentiment analysis models for various applications, such as customer feedback analysis, social media… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-sentiment-analysis-dataset-v1.Amazon_Reviews_Binary_for_Sentiment_Analysis
Dataset Card for Dataset Name
The Amazon reviews polarity dataset is constructed by taking review score 1 and 2 as negative, and 4 and 5 as positive. Samples of score 3 is ignored. In the dataset, class 1 is the negative and class 2 is the positive. Each class has 1,800,000 training samples and 200,000 testing samples.
Dataset Details
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_Binary_for_Sentiment_Analysis.airlines-sentiment-analysisAmazon_Reviews_for_Sentiment_Analysis_fine_grained_5_classes
Dataset Card for Dataset Name
The Amazon reviews full score dataset is constructed by randomly taking 600,000 training samples and 130,000 testing samples for each review score from 1 to 5. In total there are 3,000,000 trainig samples and 650,000 testing samples.
Dataset Details
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3 columns in them, corresponding to class index (1 to 5)… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_for_Sentiment_Analysis_fine_grained_5_classes.Turkish_SentimentAnalysis_TRSAv1TRSAv1 (Turkish Sentiment Analysis Version 1) Dataset
This data set has been produced to contribute to Turkish NLP studies.
The dataset consists of a total of 150 thousand samples, 50 thousand negative, 50 thousand positive, and 50 thousand neutral.
It can be used in text classification and sentiment analysis studies by citing the related study.
Related Work
Aydoğan M, Kocaman V. TRSAv1: A new benchmark dataset for classifying user reviews on Turkish e-commerce websites. Journal of… See the full description on the dataset page: https://huggingface.co/datasets/maydogan/Turkish_SentimentAnalysis_TRSAv1.sinhala-sentiment-analysissentiment-analysis-tweetYelp_Reviews_for_Sentiment_Analysis_fine_grained_5_classes
Dataset Card for Dataset Name
The Yelp reviews full star dataset is constructed by randomly taking 130,000 training samples and 10,000 testing samples for each review star from 1 to 5. In total there are 650,000 trainig samples and 50,000 testing samples.
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 2 columns in them, corresponding to class index (1 to 5) and review text. The review texts are… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yelp_Reviews_for_Sentiment_Analysis_fine_grained_5_classes.sentiment_analysis_preprocessed_datasetBrief idea about dataset:
This dataset is designed for a Text Classification to be specific Multi Class Classification, inorder to train a model (Supervised Learning) for Sentiment Analysis.
Also to be able retrain the model on the given feedback over a wrong predicted sentiment this dataset will help to manage those things using Other Features.
Main Features
text
labels
This feature variable has all sort of texts, sentences, tweets, etc.
This target variable contains 3 types of… See the full description on the dataset page: https://huggingface.co/datasets/prasadsawant7/sentiment_analysis_preprocessed_dataset.sentiment-analysis-for-mental-healthCode-Mixed-Sentiment-Analysis-Dataset
Dataset Generation:
Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.vietnamese-sentiment-analysissentiment-analysis-llama2mdismielhossenabir_sentiment-analysis
Sentiment Analysis
Mirror of the Kaggle dataset mdismielhossenabir/sentiment-analysis by Md. Ismiel Hossen Abir, released under MIT. All credit goes to the original author; please cite and link the Kaggle page when using this data.
Categorical Sentiment Analysis from Text
License
MIT License, Copyright (c) Md. Ismiel Hossen Abir. The full license text is in LICENSE and applies to all files in this repository.
Original description (from Kaggle)… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/mdismielhossenabir_sentiment-analysis.BilTweetNews-sentiment-analysis
Turkish Sentiment Analysis Tweet Dataset: BilTweetNews
The dataset contains tweets related to six major events from Turkish news sources between May 4, 2015
and Jan 8, 2017.
The dataset covers 6 major events:
May 25, 2015 One of the popular football clubs in Turkey, Galatasaray, wins the 2015
Turkish Super League.
Sep 6, 2015 A terrorist group, called PKK, attacked to soldiers in Dağlıca, a village in
southeastern Turkey.
Oct 7, 2015 A Turkish scientist, Aziz Sancar, won the 2015… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/BilTweetNews-sentiment-analysis.twitter-sentiment-meta-analysis
Twitter Sentiment Meta-Analysis Dataset
Dataset Description
This dataset contains sentiment analysis results for English tweets collected between September 2009 and January 2010. The tweets were processed and analyzed using 10 different sentiment classifiers, with the final sentiment score derived from principal component analysis (PCA).
Source Data
Original Data: Cheng-Caverlee-Lee Twitter Scrape (Sept 2009 - Jan 2010)
Number of Tweets: 138 690
Language:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/twitter-sentiment-meta-analysis.cammyc_nfl-twitter-sentiment-analysis
NFL Twitter Sentiment Analysis
Mirror of the Kaggle dataset cammyc/nfl-twitter-sentiment-analysis by Cameron C, released under MIT. All credit goes to the original author; please cite and link the Kaggle page when using this data.
Labelled NFL related tweets | sentiment analysis | scraped from twitter
License
MIT License, Copyright (c) Cameron C. The full license text is in LICENSE and applies to all files in this repository.
Original description… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/cammyc_nfl-twitter-sentiment-analysis.ahmadsudarsono_sentiment-analysis-genshin-impact-reviews
Sentiment Analysis: Genshin Impact Reviews
Mirror of the Kaggle dataset ahmadsudarsono/sentiment-analysis-genshin-impact-reviews by mattismyname, released under MIT. All credit goes to the original author; please cite and link the Kaggle page when using this data.
Reviews collected between January 2024 - March 2024
License
MIT License, Copyright (c) mattismyname. The full license text is in LICENSE and applies to all files in this repository.… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/ahmadsudarsono_sentiment-analysis-genshin-impact-reviews.Nostalgic_Sentiment_Analysis_of_YouTube_Comments_Data
Dataset Summary
The dataset is a collection of Youtube Comments and it was captured using the YouTube Data API.
The data set consists of 1500 nostalgic and non-nostalgic comments in English.
Languages
The language of the data is English.
Citation
If you find this dataset usefull for your study, please cite the paper as followed:
@article{postalcioglu2020comparison,
title={Comparison of Neural Network Models for Nostalgic Sentiment Analysis of YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Senem/Nostalgic_Sentiment_Analysis_of_YouTube_Comments_Data.sentiment-analysis-catalan-reviews
CSXSC: Classificador de Sentiments de Xarxes Socials en Català
This repository contains the CSXSC (Classificador de Sentiments a Xarxes Socials en Català) dataset, a comprehensive corpus designed for sentiment analysis of Catalan-language content from social media.
The dataset contains 23,788 text entries, each classified as positive, negative, or neutral. It was specifically constructed to address the significant class imbalance often found in user-generated content, resulting in a… See the full description on the dataset page: https://huggingface.co/datasets/Danie1Arias/sentiment-analysis-catalan-reviews.twitter_sentiment_analysisSentiment-Analysis
Sentiment Analysis Dataset
Overview
This dataset is designed for sentiment analysis tasks, providing labeled examples across three sentiment categories:
0: Negative
1: Neutral
2: Positive
It is suitable for training, validating, and testing text classification models in tasks such as social media sentiment analysis, customer feedback evaluation, and opinion mining.
Dataset Details
Key Features
Type: CSV
Language: English
Labels:
0:… See the full description on the dataset page: https://huggingface.co/datasets/syedkhalid0/Sentiment-Analysis.amazon-reviews-sentiment-analysis
Dataset Card for amazon reviews for sentiment analysis
Dataset Summary
One of the most important problems in e-commerce is the correct calculation of the points given to after-sales products. The solution to this problem is to provide greater customer satisfaction for the e-commerce site, product prominence for sellers, and a seamless shopping experience for buyers. Another problem is the correct ordering of the comments given to the products. The prominence of… See the full description on the dataset page: https://huggingface.co/datasets/Zhengyif/amazon-reviews-sentiment-analysis.faroese_sentiment_analysis
Good or Bad News? Exploring GPT-4 for Sentiment Analysis on Faroese News Corpora
This dataset is a part of the research from the paper "Good or Bad News? Exploring GPT-4 for Sentiment Analysis for Faroese on a Public News Corpora," that focuses on the application of GPT-4 for sentiment analysis on Faroese news texts.
The study addresses the challenges of sentiment analysis in low-resource languages and evaluates the effectiveness of Large Language Models, specifically GPT-4, in… See the full description on the dataset page: https://huggingface.co/datasets/hafsteinn/faroese_sentiment_analysis.PersianTwitterDataset-SentimentAnalysis
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This dataset contains more than 3300 Persian tweets, crawled from X.com
Each tweet is assigned a label, which is a number between 0 to 4.
Label 0 indicates the sentiment of Happiness and Joy.
Label 1 indicates the sentiment of Sadness.
Label 2 indicates the sentiment of Anger and… See the full description on the dataset page: https://huggingface.co/datasets/moali-mkh-2000/PersianTwitterDataset-SentimentAnalysis.
