Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aisingapore /NLU-Sentiment-Analysisgated SEA Sentiment Analysis SEA Sentiment Analysis evaluates a model's ability to identify the sentiment polarity of a text. It is sampled from NusaX for Indonesian, Javanese, and Sundanese, IndicSentiment for Tamil, Wisesight Sentiment for Thai, and UIT-VSFC for Vietnamese. Supported Tasks and Leaderboards SEA Sentiment Analysis is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Sentiment-Analysis.texttext-generation1K<n<10K0 likes1.8k downloads9mo agoHugging Face02mteb /SentimentAnalysisHindi SentimentAnalysisHindi An MTEB dataset Massive Text Embedding Benchmark Hindi Sentiment Analysis Dataset Task category t2c Domains Reviews, Written Reference https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("SentimentAnalysisHindi") evaluator = mteb.MTEB([task]) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SentimentAnalysisHindi.texttext-classification1K<n<10K0 likes837 downloads1y agoHugging Face03winvoker /turkish-sentiment-analysis-dataset Dataset This dataset contains positive , negative and notr sentences from several data sources given in the references. In the most sentiment models , there are only two labels; positive and negative. However , user input can be totally notr sentence. For such cases there were no data I could find. Therefore I created this dataset with 3 class. Positive and negative sentences are listed below. Notr examples are extraced from turkish wiki dump. In addition, added some random text… See the full description on the dataset page: https://huggingface.co/datasets/winvoker/turkish-sentiment-analysis-dataset.texttext-classification100K<n<1M51 likes545 downloads3y agoHugging Face04Sp1786 /multiclass-sentiment-analysis-dataset Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Sp1786/multiclass-sentiment-analysis-dataset.tabulartext-classification10K<n<100K29 likes466 downloads3y agoHugging Face05evalitahf /sentiment_analysisSENTIPOLC 2016 dataset The SENTIPOLC 2016 dataset contains 9410 tweets annotated for subjectivity, overall and literal polarity, and irony. The dataset has been created and used in the context of the SENTIPOLC 2016 task (http://www.di.unito.it/~tutreeb/sentipolc-evalita16/index.html), organized as part of the EVALITA 2016 evaluation campaign. Original files available here: https://live.european-language-grid.eu/catalogue/corpus/7479/download/ If you find this dataset useful please cite:… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/sentiment_analysis.texttext-classification1K<n<10K0 likes364 downloads2y agoHugging Face06hugginglearners /amazon-reviews-sentiment-analysis Dataset Card for amazon reviews for sentiment analysis Dataset Summary One of the most important problems in e-commerce is the correct calculation of the points given to after-sales products. The solution to this problem is to provide greater customer satisfaction for the e-commerce site, product prominence for sellers, and a seamless shopping experience for buyers. Another problem is the correct ordering of the comments given to the products. The prominence of misleading… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/amazon-reviews-sentiment-analysis.tabular1K<n<10K6 likes303 downloads4y agoHugging Face07ParsiAI /snappfood-sentiment-analysistexttext-classification10K<n<100K7 likes271 downloads2y agoHugging Face08themohal /saraiki-sentiment-analysis-datasetgated Saraiki Sentiment Analysis Dataset Sentiment-annotated Saraiki text for sentiment classification and low-resource NLP. About Saraiki Saraiki (سرائیکی) is an Indo-Aryan language spoken primarily in Pakistan, particularly across southern Punjab and neighboring regions. It has a rich linguistic, literary, and cultural heritage, with its own vocabulary, grammatical structures, and regional varieties. Despite its significant speaker community, Saraiki remains a… See the full description on the dataset page: https://huggingface.co/datasets/themohal/saraiki-sentiment-analysis-dataset.textn<1K0 likes242 downloads3h agoHugging Face09ParsiAI /digikala-sentiment-analysistabulartext-classification1K<n<10K3 likes222 downloads2y agoHugging Face10carblacac /twitter-sentiment-analysisThe Twitter Sentiment Analysis Dataset contains 1,578,627 classified tweets, each row is marked as 1 for positive sentiment and 0 for negative sentiment. The dataset is based on data from the following two sources: University of Michigan Sentiment Analysis competition on Kaggle Twitter Sentiment Corpus by Niek Sanders Finally, I randomly selected a subset of them, applied a cleaning process, and divided them between the test and train subsets, keeping a balance between the number of positive and negative tweets within each of these subsets.text-classification100K<n<1M25 likes193 downloads4y agoHugging Face11tanaos /synthetic-sentiment-analysis-dataset-v1 Tanaos Sentiment Analysis Training Dataset This dataset was created synthetically by Tanaos with the Artifex Python library. The dataset is designed to train and evaluate sentiment analysis systems — models that classify the sentiment expressed in text as one of five possible categories: very_negative, negative, neutral, positive or very_positive. It can be used to build sentiment analysis models for various applications, such as customer feedback analysis, social media… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-sentiment-analysis-dataset-v1.texttext-classification10K<n<100K0 likes193 downloads10mo agoHugging Face12AiresPucrs /sentiment-analysis-pt Sentiment Analysis PT (Teeny-Tiny Castle) This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research. How to Use from datasets import load_dataset dataset = load_dataset("AiresPucrs/sentiment-analysis-pt", split = 'train') texttext-classification10K<n<100K5 likes187 downloads2y agoHugging Face13mljucyyyy /weibo_sentiment_analysis_1000k1 likes172 downloads7mo agoHugging Face14yassiracharki /Amazon_Reviews_Binary_for_Sentiment_Analysis Dataset Card for Dataset Name The Amazon reviews polarity dataset is constructed by taking review score 1 and 2 as negative, and 4 and 5 as positive. Samples of score 3 is ignored. In the dataset, class 1 is the negative and class 2 is the positive. Each class has 1,800,000 training samples and 200,000 testing samples. Dataset Details Dataset Description The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_Binary_for_Sentiment_Analysis.texttext-classification1M<n<10M0 likes153 downloads2y agoHugging Face15jvanz /portuguese_sentiment_analysisThis dataset is based on the dataset originally posted in Kaggle text1M<n<10M12 likes150 downloads4y agoHugging Face16Sadiksha /sentiment_analysis_data Dataset Card for "sentiment_analysis_data" More Information needed text10K<n<100K0 likes141 downloads4y agoHugging Face17Akash190104 /bengali_sentiment_analysis Bengali Sentiment Analysis Context The dataset contains 3307 Negative reviews and 8500 Positive reviews collected and manually annotated from Youtube Bengali drama. Positive_Label=1 and Negative_Label=0 Acknowledgements Sazzed, Salim (2021), “Bangla ( Bengali ) sentiment analysis classification benchmark dataset corpus”, Mendeley Data, V4, doi: 10.17632/p6zc7krs37.4 texttext-classification10K<n<100K1 likes137 downloads2y agoHugging Face18DevBM /Sentiment-Analysis-Text-from-Yelp-Imdb-Amazontext1K<n<10K0 likes136 downloads2y agoHugging Face19SaguaroCapital /sentiment-analysis-in-commodity-market-gold Dataset Card for Sentiment Analysis of Commodity News (Gold) This is a news dataset for the commodity market which has been manually annotated for 10,000+ news headlines across multiple dimensions into various classes. The dataset has been sampled from a period of 20+ years (2000-2021). The dataset was curated by Ankur Sinha and Tanmay Khandait and is detailed in their paper "Impact of News on the Commodity Market: Dataset and Results." It is currently published by the authors on… See the full description on the dataset page: https://huggingface.co/datasets/SaguaroCapital/sentiment-analysis-in-commodity-market-gold.tabulartext-classification10K<n<100K6 likes130 downloads2y agoHugging Face20maydogan /Turkish_SentimentAnalysis_TRSAv1TRSAv1 (Turkish Sentiment Analysis Version 1) Dataset This data set has been produced to contribute to Turkish NLP studies. The dataset consists of a total of 150 thousand samples, 50 thousand negative, 50 thousand positive, and 50 thousand neutral. It can be used in text classification and sentiment analysis studies by citing the related study. Related Work Aydoğan M, Kocaman V. TRSAv1: A new benchmark dataset for classifying user reviews on Turkish e-commerce websites. Journal of… See the full description on the dataset page: https://huggingface.co/datasets/maydogan/Turkish_SentimentAnalysis_TRSAv1.texttext-classification100K<n<1M15 likes128 downloads2y agoHugging Face21sudhanshusinghaiml /airlines-sentiment-analysistext10K<n<100K0 likes122 downloads2y agoHugging Face22Sanatbek /aspect-based-sentiment-analysis-uzbekdocumenttext-classification1K<n<10K4 likes106 downloads3y agoHugging Face23elmurod1202 /uzbek-sentiment-analysis uzbek-sentiment-analysis Sentiment analysis in the Uzbek language and new Datasets of Uzbek App reviews for Sentiment Classification Feel free to use the dataset and the tools presented in this project, a paper about more details on creation and usage here. If you find it useful, please make sure to cite the paper: @inproceedings{kuriyozov2019deep, author = {Kuriyozov, Elmurod and Matlatipov, Sanatbek and Alonso, Miguel A and Gómez-Rodríguez, Carlos}, title = {Deep… See the full description on the dataset page: https://huggingface.co/datasets/elmurod1202/uzbek-sentiment-analysis.image10K<n<100K5 likes103 downloads4y agoHugging Face24bdstar /Tweets-Sentiment-Analysis 🐦 Tweets-Sentiment-Analysis (bdstar/Tweets-Sentiment-Analysis) 🧠 Overview A refined and merged version of Tweets text sentiment datasets, providing a clean and well-balanced dataset for sentiment classification across three sentiment categories:positive, negative, and neutral. This dataset is split into three parts — train, test, and validation — each sourced from highly reputable open datasets.It is designed for training, evaluating, and benchmarking NLP models… See the full description on the dataset page: https://huggingface.co/datasets/bdstar/Tweets-Sentiment-Analysis.texttext-classification1M<n<10M0 likes101 downloads4mo agoHugging Face25yassiracharki /Amazon_Reviews_for_Sentiment_Analysis_fine_grained_5_classes Dataset Card for Dataset Name The Amazon reviews full score dataset is constructed by randomly taking 600,000 training samples and 130,000 testing samples for each review score from 1 to 5. In total there are 3,000,000 trainig samples and 650,000 testing samples. Dataset Details Dataset Description The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3 columns in them, corresponding to class index (1 to 5)… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_for_Sentiment_Analysis_fine_grained_5_classes.texttext-classification1M<n<10M4 likes97 downloads2y agoHugging Face26sinhala-nlp /sinhala-sentiment-analysistext1K<n<10K0 likes94 downloads2y agoHugging Face27DynamicSuperb /Sentiment_Analysis_SLUE-VoxCelebaudio1K<n<10K1 likes94 downloads2y agoHugging Face28MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-anger Dataset Summary Synthetic Persian Chatbot Conversational SA – Anger is a Persian (Farsi) dataset created for the Classification task, with a focus on detecting the emotion "anger" in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s): Classification… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-anger.text1K<n<10K0 likes90 downloads1y agoHugging Face29MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-friendship Dataset Summary Synthetic Persian Chatbot Conversational SA – Friendship is a Persian (Farsi) dataset created for the Classification task, with a focus on detecting the emotion "friendship" in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-friendship.textn<1K0 likes87 downloads1y agoHugging Face30MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-sadness Dataset Summary Synthetic Persian Chatbot Conversational SA – Sadness is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "sadness" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-sadness.textn<1K0 likes86 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.