datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sentiment_analysisSENTIPOLC 2016 dataset
The SENTIPOLC 2016 dataset contains 9410 tweets annotated for subjectivity, overall and literal polarity, and irony.
The dataset has been created and used in the context of the SENTIPOLC 2016 task (http://www.di.unito.it/~tutreeb/sentipolc-evalita16/index.html), organized as part of the EVALITA 2016 evaluation campaign.
Original files available here:
https://live.european-language-grid.eu/catalogue/corpus/7479/download/
If you find this dataset useful please cite:… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/sentiment_analysis.alpaca-bitcoin-sentiment-datasetkinopoisk-sentiment-classificationtweet_sentiment_extraction
Tweet Sentiment Extraction
Source: https://www.kaggle.com/c/tweet-sentiment-extraction/data
kap-turkish-financial-sentiment
KAP Turkish Financial Sentiment Dataset
Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti.
Dataset Bilgileri
Özellik
Değer
Kayıt Sayısı
3,839
Dil
Türkçe
Kaynak
KAP Bildirimleri
Etiketleme
GPT-4 (Teacher Model)
Format
JSONL (Chat Messages)
Kullanım Alanları
Türkçe finansal sentiment analizi
KAP bildirimi sınıflandırma
Volatilite tahmini
İlişkili taraf işlemi tespiti
LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/finansai/kap-turkish-financial-sentiment.Tri-Class-Sentiment-Synthetic
Dataset Card for Dataset Name
Tri-Class synthetic sentiment dataset, split into POSITIVE,NEUTRAL & NEGATIVE sentiment evenly split. Data generated by phi3.5-mini-instruct-q8_0.
Dataset Details
Dataset Description
A synthetic dataset designed for training tri-class(POSITIVE, NEUTRAL, NEGATIVE) sentiment analysis AI.
The data is labeled as such:
text (contains the text of the review)
sentiment (an integer, being 1 for POSITIVE, 0 for NEUTRAL and -1 for NEGATIVE… See the full description on the dataset page: https://huggingface.co/datasets/Novora/Tri-Class-Sentiment-Synthetic.ecommerce-product-reviews-sentiment
Dataset Summary
This dataset contains 11,606 product reviews gathered from various indonesian brands and products in several e-commerce such as Shopee, Tokopedia, Lazada, Bukalapak, Blili, and Zalora.
each row is marked as 1 for positive sentiment and 0 for negative sentiment.
This dataset has been transformed, selecting in a random way a subset of them, applying a cleaning process, and dividing them between the test and train subsets, keeping a balance between the number of… See the full description on the dataset page: https://huggingface.co/datasets/dipawidia/ecommerce-product-reviews-sentiment.financial-sentiment-pt
Financial Sentiments (PT-BR)
A dataset containing ~63k rows of translated Financial Sentiment data in Portuguese.
Pipeline
Dataset cleaning (dedup and benchmark decontamination) -> merge -> Translation Pipeline -> MiniCPM5 2B -> Kiwi COMET as scorer -> drop below a threshold when doing sft (default = 0.5).
Exact code at: https://github.com/mansa-team/musa
Data Quality
n_scored=64685, mean=0.6893, n_below=6975
Cite… See the full description on the dataset page: https://huggingface.co/datasets/heitorrosa/financial-sentiment-pt.Korean-YouTube-Comment-Sentiment-Dataset
Korean YouTube Comment Sentiment Dataset
Data Overview
Summary
본 데이터셋은 유튜브에서 수집된 한국어 댓글 5,482개와 이에 대응하는 감정 레이블(긍정, 부정, 중립, 불명확)로 구성된 감정 분류용 데이터셋입니다.
주요 레이블: 긍정, 부정, 중립, 불명확
Features
수집 대상: 요리, 뷰티, 게임, 여행, 쇼핑 등 분야의 10만 명 이상 구독자를 보유한 유튜브 채널
형식: JSON (id, text, label)
검수: 한국인 검수자에 의한 수작업 라벨링 및 교차 검토
본 데이터셋은 구어체, 이모지, 줄임말 등 실제 사용자 표현이 반영되어 있습니다.
Dataset Structure
Dataset Fields
Field
Type
Description
id
string
각 댓글의… See the full description on the dataset page: https://huggingface.co/datasets/LLM-SocialMedia/Korean-YouTube-Comment-Sentiment-Dataset.mobile-apps-user-sentiment-reviews
Top Mobile Apps User Sentiment & Review Corpus (Google Play)
Overview
This dataset contains clean, structured public data exported directly from production runs of Apify actors.
It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines.
Source Actor: captainhandsome/google-play-reviews-scraper
Dataset Page: Public sample and schema
Preconfigured Run Task: captainhandsome/instagram-1star-reviews… See the full description on the dataset page: https://huggingface.co/datasets/joeygambino/mobile-apps-user-sentiment-reviews.ai-sentiment-2026
Ai Sentiment 2026
Part of the LEGION Intelligence dataset collection.
Provider: LEGION Systems
Access: Requires approval — submit request below
Usage
from datasets import load_dataset
dataset = load_dataset("gemmozero/ai-sentiment-2026")
API Access
Real-time access via LEGION API:
curl https://api.legion-api.com/incidents
API Docs · Pro Access €29/mo
License
CC BY-NC 4.0 — Research and non-commercial use only.
Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-sentiment-2026.chinese-sentiment
中文情感数据集卡片
数据集详情
数据集简介
中文情感数据集,包含日常生活场景下的中文短文本情感标注。包含 6 种基本情感类别,共 20,000 条样本。
策划者: TIX007
语言: 中文(简体) / zh
许可证: MIT
数据来源
仓库: https://huggingface.co/datasets/TIX007/chinese-sentiment
使用方式
Direct Use
中文情感分类模型的训练与评估
文本情感分析(sentiment analysis)研究
作为预训练语言模型微调的数据集
快速验证中文 NLP 模型的情感识别能力
Out-of-Scope Use
不建议直接用于商业生产环境(文本来源于模板扩增,缺乏真实社交媒体多样性)
不适用于细粒度情感分析(如情感强度打分)
不适用于对话情感分析或上下文依赖的情感理解… See the full description on the dataset page: https://huggingface.co/datasets/wickedlin/chinese-sentiment.chat-sentiment-analysis
A Sentiment Analsysis Dataset for Finetuning Large Models in Chat-style
More details can be found at https://github.com/l294265421/chat-sentiment-analysis
Supported Tasks
Aspect Term Extraction (ATE)
Opinion Term Extraction (OTE)
Aspect Term-Opinion Term Pair Extraction (AOPE)
Aspect term, Sentiment, Opinion term Triplet Extraction (ASOTE)
Aspect Category Detection (ACD)
Aspect Category-Sentiment Pair Extraction (ACSA)
Aspect-Category-Opinion-Sentiment (ACOS) Quadruple… See the full description on the dataset page: https://huggingface.co/datasets/yuncongli/chat-sentiment-analysis.synthetic-persian-chatbot-conversational-sentiment-analysis-anger
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Anger is a Persian (Farsi) dataset created for the Classification task, with a focus on detecting the emotion "anger" in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
Language(s): Persian (Farsi)
Task(s): Classification… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-anger.chatbot-conversational-sentiment-analysis-tone-user-classification
Dataset Summary
Synthetic Persian Chatbot Conversational SA – User Tone Classification(SynPerChatbotConvSAToneUserClassification) is a Persian (Farsi) dataset created for the Classification task. It focuses on identifying the user’s conversational tone—formal, casual, or childish—in emotionally rich chatbot interactions. This dataset is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini.
Language(s): Persian (Farsi)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/chatbot-conversational-sentiment-analysis-tone-user-classification.synthetic-persian-chatbot-conversational-sentiment-analysis-friendship
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Friendship is a Persian (Farsi) dataset created for the Classification task, with a focus on detecting the emotion "friendship" in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
Language(s): Persian (Farsi)
Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-friendship.synthetic-persian-chatbot-conversational-sentiment-analysis-fear
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Fear is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "fear" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is a subset of the Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
Language(s): Persian (Farsi)
Task(s): Classification (Emotion… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-fear.synthetic-persian-chatbot-conversational-sentiment-analysis-sadness
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Sadness is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "sadness" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
Language(s): Persian (Farsi)
Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-sadness.synthetic-persian-chatbot-conversational-sentiment-analysis-satisfaction
Dataset Summary
Synthetic Persian Chatbot Conversational Sentiment Analysis – Satisfaction is a Persian (Farsi) dataset developed for the Classification task, specifically focused on detecting the emotion of satisfaction in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using the GPT-4o-mini language model.
Language(s): Persian (Farsi)
Task(s): Classification (Emotion Detection – Satisfaction)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-satisfaction.synthetic-persian-chatbot-conversational-sentiment-analysis-jealousy
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Jealousy is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "jealousy" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
Language(s): Persian (Farsi)
Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-jealousy.synthetic-persian-chatbot-conversational-sentiment-analysis-surprise
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Surprise is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "surprise" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is a subset of the Synthetic Persian Chatbot Conversational Sentiment Analysis dataset.
Language(s): Persian (Farsi)
Task(s): Classification… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-surprise.pump-fun-sentiment-100k
Pump.fun Token Sentiment & Risk Analysis
AI-agent-generated sentiment analysis and quantitative risk labels for Solana memecoins on Pump.fun, collected via Pump Studio.
Dataset Description
Each row is a validated analysis submission from an AI agent operating on the Pump Studio platform. Agents observe real-time token data (price, market cap, holders, volume, bonding curve) and produce:
Sentiment label — bullish / bearish / neutral with 0-100 confidence score
Risk… See the full description on the dataset page: https://huggingface.co/datasets/Pumpdotstudio/pump-fun-sentiment-100k.synthetic-persian-chatbot-conversational-sentiment-analysis-love
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Love is a Persian (Farsi) dataset for the Classification task, specifically focused on detecting the emotion "love" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis collection.
Language(s): Persian (Farsi)
Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-love.synthetic-persian-chatbot-conversational-sentiment-analysis-happiness
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Happiness is a Persian (Farsi) dataset for the Classification task, specifically focused on detecting the emotion "happiness" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini, and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis collection.
Language(s): Persian (Farsi)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-happiness.synthetic-persian-chatbot-conversational-sentiment-analysis-tone-chatbot-classification
Dataset Summary
Synthetic Persian Chatbot Conversational SA – Chatbot Tone Classification(SynPerChatbotConvSAToneChatbotClassification) is a Persian (Farsi) dataset created for the Classification task. It focuses on identifying the chatbot’s conversational tone—formal, casual, or childish—in dialogue exchanges that include emotional content. This dataset is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini.
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-tone-chatbot-classification.twitter_us_airline_sentiment
Twitter US Airline Sentiment Dataset
This dataset contains tweets about US airlines labeled with sentiment (negative, neutral, positive).
Dataset Details
Total tweets: 14,640
Training set: 11,712 tweets
Validation set: 1,464 tweets
Test set: 1,464 tweets
Classes: negative (0), neutral (1), positive (2)
Files
train.jsonl: Training data (11,712 examples)
validation.jsonl: Validation data (1,464 examples)
test.jsonl: Test data (1,464 examples)… See the full description on the dataset page: https://huggingface.co/datasets/Adetyakr/twitter_us_airline_sentiment.sentiment-analysis-for-financial-news-v2movie-review-sentiment
🎬 Movie Review Sentiment (Mini)
A tiny hand-built dataset of one-sentence movie reviews labeled with their
sentiment, used to demonstrate the full Hugging Face workflow:
Dataset → Model → Space
What's inside
Split
Rows
Classes
train
168
positive, negative, neutral
test
42
positive, negative, neutral
Each example has two fields:
text — an English sentence reviewing a movie
label — one of positive, negative, neutral
Why this… See the full description on the dataset page: https://huggingface.co/datasets/chennab28/movie-review-sentiment.sentiment_analysis_hindiConventions followed to decide the polarity: -
labels consisting of a single value are left undisturbed, i.e. if label = 'pos', then it'll be pos
labels consisting of multiple values separated by '&' are processed. If all the labels are the same ('pos&pos&pos' or 'neg&neg'), then the shortened form of the multiple label is assigned as the final label. For example, if label = 'pos&pos&pos', then final label will be 'pos'.
labels consisting of mixed values ('pos&neg&pos' or 'neg&neu&pos') are… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi.e-commerce-sentiment-bahasa-indonesia
E-Commerce Sentiment Analysis Dataset (Indonesian)
Dataset komentar dan ulasan produk e-commerce dalam Bahasa Indonesia untuk analisis sentiment.
Dataset Summary
Dataset ini berisi 21,840 komentar e-commerce dalam Bahasa Indonesia yang telah dilabeli dengan sentiment (positif, netral, negatif). Dataset mencakup berbagai jenis komentar termasuk sarkasme dan ironi yang umum ditemukan dalam ulasan online.
Dataset Structure
Data Fields
comment (string):… See the full description on the dataset page: https://huggingface.co/datasets/AIbnuHibban/e-commerce-sentiment-bahasa-indonesia.
