datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SentimentAnalysisHindi
SentimentAnalysisHindi
An MTEB dataset
Massive Text Embedding Benchmark
Hindi Sentiment Analysis Dataset
Task category
t2c
Domains
Reviews, Written
Reference
https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("SentimentAnalysisHindi")
evaluator = mteb.MTEB([task])
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SentimentAnalysisHindi.sentiment_analysisSENTIPOLC 2016 dataset
The SENTIPOLC 2016 dataset contains 9410 tweets annotated for subjectivity, overall and literal polarity, and irony.
The dataset has been created and used in the context of the SENTIPOLC 2016 task (http://www.di.unito.it/~tutreeb/sentipolc-evalita16/index.html), organized as part of the EVALITA 2016 evaluation campaign.
Original files available here:
https://live.european-language-grid.eu/catalogue/corpus/7479/download/
If you find this dataset useful please cite:… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/sentiment_analysis.sentiment-analysis-pt
Sentiment Analysis PT (Teeny-Tiny Castle)
This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research.
How to Use
from datasets import load_dataset
dataset = load_dataset("AiresPucrs/sentiment-analysis-pt", split = 'train')
Sentiment-Analysis-Text-from-Yelp-Imdb-Amazonsentiment_analysis_data
Dataset Card for "sentiment_analysis_data"
More Information needed
Turkish_SentimentAnalysis_TRSAv1TRSAv1 (Turkish Sentiment Analysis Version 1) Dataset
This data set has been produced to contribute to Turkish NLP studies.
The dataset consists of a total of 150 thousand samples, 50 thousand negative, 50 thousand positive, and 50 thousand neutral.
It can be used in text classification and sentiment analysis studies by citing the related study.
Related Work
Aydoğan M, Kocaman V. TRSAv1: A new benchmark dataset for classifying user reviews on Turkish e-commerce websites. Journal of… See the full description on the dataset page: https://huggingface.co/datasets/maydogan/Turkish_SentimentAnalysis_TRSAv1.sentiment-analysis-in-commodity-market-gold
Dataset Card for Sentiment Analysis of Commodity News (Gold)
This is a news dataset for the commodity market which has been manually annotated for 10,000+ news headlines across multiple dimensions into various classes. The dataset has been sampled from a period of 20+ years (2000-2021).
The dataset was curated by Ankur Sinha and Tanmay Khandait and is detailed in their paper "Impact of News on the Commodity Market: Dataset and Results." It is currently published by the authors on… See the full description on the dataset page: https://huggingface.co/datasets/SaguaroCapital/sentiment-analysis-in-commodity-market-gold.sentiment-analysis-monitoring
Sentiment Analysis Monitoring & Retraining Telemetry
Repository di telemetria e dati per la pipeline MLOps del modello di Sentiment Analysis (Twitter-RoBERTa).
Configurazioni disponibili:
metrics: Storico temporale dei batch di monitoraggio, andamento dell'indice F1 e distribuzione del sentiment (% Positivi, Negativi, Neutri).
data: Buffer di accumulo e snapshot utilizzati per il retraining continuo del modello.
Sentiment_Analysis_SLUE-VoxCelebsentiment-analysis-tweetsentiment_analysis_preprocessed_datasetBrief idea about dataset:
This dataset is designed for a Text Classification to be specific Multi Class Classification, inorder to train a model (Supervised Learning) for Sentiment Analysis.
Also to be able retrain the model on the given feedback over a wrong predicted sentiment this dataset will help to manage those things using Other Features.
Main Features
text
labels
This feature variable has all sort of texts, sentences, tweets, etc.
This target variable contains 3 types of… See the full description on the dataset page: https://huggingface.co/datasets/prasadsawant7/sentiment_analysis_preprocessed_dataset.sentiment-analysis-for-mental-healthsentiment-analysis-for-financial-news-v2sentiment_analysis_hindiConventions followed to decide the polarity: -
labels consisting of a single value are left undisturbed, i.e. if label = 'pos', then it'll be pos
labels consisting of multiple values separated by '&' are processed. If all the labels are the same ('pos&pos&pos' or 'neg&neg'), then the shortened form of the multiple label is assigned as the final label. For example, if label = 'pos&pos&pos', then final label will be 'pos'.
labels consisting of mixed values ('pos&neg&pos' or 'neg&neu&pos') are… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi.Sentiment-Analysis-ComplexExcellent — congrats on getting the repo ready 🚀
Here’s a professional Hugging Face Dataset Card (README.md) you can paste directly into your repository.
This is written to match HF best practices and serious research usage.
📘 README.md
👉 Copy everything below into your README.md
Sentiment-Analysis-Complex
🧠 Overview
Sentiment-Analysis-Complex is a large-scale synthetic sentiment analysis dataset designed for benchmarking modern NLP models under… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/Sentiment-Analysis-Complex.sentiment-analysis-llama2sentiment-analysis-catalan-reviews
CSXSC: Classificador de Sentiments de Xarxes Socials en Català
This repository contains the CSXSC (Classificador de Sentiments a Xarxes Socials en Català) dataset, a comprehensive corpus designed for sentiment analysis of Catalan-language content from social media.
The dataset contains 23,788 text entries, each classified as positive, negative, or neutral. It was specifically constructed to address the significant class imbalance often found in user-generated content, resulting in a… See the full description on the dataset page: https://huggingface.co/datasets/Danie1Arias/sentiment-analysis-catalan-reviews.Sentiment_Analysis_TweetsSentiment-Analysis
Sentiment Analysis Dataset
Overview
This dataset is designed for sentiment analysis tasks, providing labeled examples across three sentiment categories:
0: Negative
1: Neutral
2: Positive
It is suitable for training, validating, and testing text classification models in tasks such as social media sentiment analysis, customer feedback evaluation, and opinion mining.
Dataset Details
Key Features
Type: CSV
Language: English
Labels:
0:… See the full description on the dataset page: https://huggingface.co/datasets/syedkhalid0/Sentiment-Analysis.sentiment-analysis-for-financial-news
Sentiment Analysis for Financial News (FinancialPhraseBank-style)
This dataset contains financial news headlines labeled with sentiment from the perspective of a retail investor.
Columns
sentiment: one of negative, neutral, positive
news_headline: financial news headline text
Source / Context
The original dataset is commonly referenced as FinancialPhraseBank and is widely used for financial sentiment benchmarking.
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/prithvi1029/sentiment-analysis-for-financial-news.sentiment-analysis
Sentiment Analysis (Teeny-Tiny Castle)
This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research.
How to Use
from datasets import load_dataset
dataset = load_dataset("AiresPucrs/sentiment-analysis", split = 'train')
PersianTwitterDataset-SentimentAnalysis
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This dataset contains more than 3300 Persian tweets, crawled from X.com
Each tweet is assigned a label, which is a number between 0 to 4.
Label 0 indicates the sentiment of Happiness and Joy.
Label 1 indicates the sentiment of Sadness.
Label 2 indicates the sentiment of Anger and… See the full description on the dataset page: https://huggingface.co/datasets/moali-mkh-2000/PersianTwitterDataset-SentimentAnalysis.sentiment_analysis_financial_news_datasentiment_analysis_hindisentiment_analysis_preprocessed_datasetBrief idea about dataset:
This dataset is designed for a Text Classification to be specific Multi Class Classification, inorder to train a model (Supervised Learning) for Sentiment Analysis.
Also to be able retrain the model on the given feedback over a wrong predicted sentiment this dataset will help to manage those things using Other Features.
Main Features
text
labels
This feature variable has all sort of texts, sentences, tweets, etc.
This target variable contains 3 types of… See the full description on the dataset page: https://huggingface.co/datasets/Aldoreni45/sentiment_analysis_preprocessed_dataset.sentiment-analysis-ind-classificationSentiment_Analysis_MCQ_train
Sentiment Analysis MCQ Training Dataset
Training dataset for financial sentiment analysis in MCQ format.
Dataset Structure
Format: Multiple choice questions
Language: Arabic
Domain: Financial reports and market news
Task: Sentiment classification (positive/negative/neutral)
Fields
id: Unique identifier
query: Full MCQ prompt with instructions
answer: Correct answer letter (a, b, c)
text: Question text without instructions
choices: List of answer… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/Sentiment_Analysis_MCQ_train.sentiment-analysis-ind-classificationSentiment-Analysis-Tokenizedsentiment-analysis-for-mental-health-Combined-Data
