datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-sentiment-classification
MultilingualSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
Sentiment classification dataset with binary
(positive vs negative sentiment) labels. Includes 30 languages and dialects.
Task category
t2c
DomainsReviews, Written
Reference
https://huggingface.co/datasets/mteb/multilingual-sentiment-classification
How to evaluate on this task
You can evaluate an embedding model on this dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/mteb/multilingual-sentiment-classification.kinopoisk-sentiment-classificationfiqa-sentiment-classification
Dataset Name
Dataset Description
This dataset is based on the task 1 of the Financial Sentiment Analysis in the Wild (FiQA) challenge. It follows the same settings as described in the paper 'A Baseline for Aspect-Based Sentiment Analysis in Financial Microblogs and News'. The dataset is split into three subsets: train, valid, test with sizes 822, 117, 234 respectively.
Dataset Structure
_id: ID of the data point
sentence: The sentence
target: The target of the… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/fiqa-sentiment-classification.task819_pec_sentiment_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task819_pec_sentiment_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task819_pec_sentiment_classification.BLUGE-bengali-sentiment-classification
BLUGE-TSC: Bangla Sentiment Classification
BLUGE-TSC is a meticulously curated and cleaned Bangla Ternary Sentiment Classification dataset, one of the 7 tasks in BLUGE (Bengali Language UnderstandinG Evaluation), a balanced benchmark for evaluating Bengali natural language understanding. See the full BLUGE collection for all 7 tasks, and the B-CORE pretraining corpus and BnLM model suite released alongside it.
Dataset Description
This task classifies Bangla text… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BLUGE-bengali-sentiment-classification.chatbot-conversational-sentiment-analysis-tone-user-classification
Dataset Summary
Synthetic Persian Chatbot Conversational SA – User Tone Classification(SynPerChatbotConvSAToneUserClassification) is a Persian (Farsi) dataset created for the Classification task. It focuses on identifying the user’s conversational tone—formal, casual, or childish—in emotionally rich chatbot interactions. This dataset is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini.
Language(s): Persian (Farsi)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/chatbot-conversational-sentiment-analysis-tone-user-classification.sentiment-classification-dataset-bundle
NLP: Sentiment Classification Dataset
This is a bundle dataset for a NLP task of sentiment classification in English.
There is a sample project is using this dataset GURA-gru-unit-for-recognizing-affect.
Content
myanimelist-sts: This dataset is derived from MyAnimeList, a social networking and cataloging service for anime and manga fans. The dataset typically includes user reviews with ratings. We used skip-thoughts to summarize them. You can find the original source of… See the full description on the dataset page: https://huggingface.co/datasets/NatLee/sentiment-classification-dataset-bundle.synthetic-persian-chatbot-conversational-sentiment-analysis-tone-chatbot-classification
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Chatbot Tone Classification(SynPerChatbotConvSAToneChatbotClassification) is a Persian (Farsi) dataset created for the Classification task. It focuses on identifying the chatbot’s conversational tone—formal, casual, or childish—in dialogue exchanges that include emotional content. This dataset is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini.
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-tone-chatbot-classification.multilingual-sentiment-ind-classificationtask833_poem_sentiment_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task833_poem_sentiment_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task833_poem_sentiment_classification.Text-Sentiment-Classification-and-Part-of-Speech-Sentiment-Classification-Mix-DatasetUSA model paper SCPOS datasets. Including four sub-datasets.
paper address:
https://arxiv.org/abs/2309.03787
Cite:
@article{gan2023usa,
title={USA: Universal Sentiment Analysis Model & Construction of Japanese Sentiment Text Classification and Part of Speech Dataset},
author={Gan, Chengguang and Zhang, Qinghao and Mori, Tatsunori},
journal={arXiv preprint arXiv:2309.03787},
year={2023}
}
This dataset constructed base in JGLUE benchmark text sentiment classification task… See the full description on the dataset page: https://huggingface.co/datasets/ganchengguang/Text-Sentiment-Classification-and-Part-of-Speech-Sentiment-Classification-Mix-Dataset.multilingual-sentiment-vie-classification
MultiLingualSentiment_vie_Classification
Deduplicated copy of kornwtp/multilingual-sentiment-vie-classification.
Splits
split
rows
test
684
train
2,358
validation
330
azerbaijani_review_sentiment_classificationAzerbaijani Sentiment Classification Dataset with ~160K reviews.
Dataset contains 3 columns: Content, Score, Upvotes
sentiment-emo-mobileapps-ind-classificationemotion_classification_reviews_kfold_by_sentiment_10Finance_sentiment_and_topic_classification_Translation_English_to_Spanish_v1news-sentiment-zsm-classificationref: https://github.com/mesolitica/malaysian-dataset/tree/master/sentiment/news-sentiment
multilingual-sentiment-tha-classification
MultiLingualSentiment_tha_Classification
Deduplicated copy of kornwtp/multilingual-sentiment-tha-classification.
Splits
split
rows
test
2,344
train
8,102
validation
1,153
sentiment-classification
Introduction
这是一个情感分类的数据集,来源飞桨。做了一些简单的处理,在此数据集上对bert微调的模型为left0ver/bert-base-chinese-finetune-sentiment-classification
因为bert最高只支持512长度的输入,因为对长度超500的样本使用了滑动窗口的办法进行了拆分,推理的时候,若是拆分出来的样本,则分别对这几个样本进行预测,取概率最大的作为最终的预测结果。具体可以看BERT模型输入长度超过512如何解决?
有关滑动窗口的版本的数据集请访问https://huggingface.co/datasets/left0ver/sentiment-classification/tree/window_version
Usage
使用普通版本:
dataset = load_dataset("left0ver/sentiment-classification")
使用滑动窗口版本
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/left0ver/sentiment-classification.twitter-financial-news-sentiment-classification
Dataset Card for "twitter-financial-news-sentiment-classification"
More Information needed
sentiment-emo-mobileapps-ind-classification
SentEmoMobileApps_ind_Classification
Deduplicated copy of kornwtp/sentiment-emo-mobileapps-ind-classification.
Splits
split
rows
train
21,405
sentiment-analysis-ind-classification
SentimentAnalysis_ind_Classification
Deduplicated copy of kornwtp/sentiment-analysis-ind-classification.
Splits
split
rows
train
10,082
multilingual-sentiment-ind-classification
MultiLingualSentiment_ind_Classification
Deduplicated copy of kornwtp/multilingual-sentiment-ind-classification.
Splits
split
rows
test
2,266
train
7,926
validation
1,132
lem-sentiment-ind-classificationmultilingual-sentiment-vie-classificationwisesight-sentiment-tha-classificationsentiment-analysis-ind-classificationmultilingual-sentiment-tha-classificationsentiment-classificationtask1497_bengali_book_reviews_sentiment_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1497_bengali_book_reviews_sentiment_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1497_bengali_book_reviews_sentiment_classification.
