datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sentiment_analysisSENTIPOLC 2016 dataset
The SENTIPOLC 2016 dataset contains 9410 tweets annotated for subjectivity, overall and literal polarity, and irony.
The dataset has been created and used in the context of the SENTIPOLC 2016 task (http://www.di.unito.it/~tutreeb/sentipolc-evalita16/index.html), organized as part of the EVALITA 2016 evaluation campaign.
Original files available here:
https://live.european-language-grid.eu/catalogue/corpus/7479/download/
If you find this dataset useful please cite:… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/sentiment_analysis.sentiment-analysis-for-financial-news-v2sentiment_analysis_hindiConventions followed to decide the polarity: -
labels consisting of a single value are left undisturbed, i.e. if label = 'pos', then it'll be pos
labels consisting of multiple values separated by '&' are processed. If all the labels are the same ('pos&pos&pos' or 'neg&neg'), then the shortened form of the multiple label is assigned as the final label. For example, if label = 'pos&pos&pos', then final label will be 'pos'.
labels consisting of mixed values ('pos&neg&pos' or 'neg&neu&pos') are… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi.Sentiment-Analysis-ComplexExcellent — congrats on getting the repo ready 🚀
Here’s a professional Hugging Face Dataset Card (README.md) you can paste directly into your repository.
This is written to match HF best practices and serious research usage.
📘 README.md
👉 Copy everything below into your README.md
Sentiment-Analysis-Complex
🧠 Overview
Sentiment-Analysis-Complex is a large-scale synthetic sentiment analysis dataset designed for benchmarking modern NLP models under… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/Sentiment-Analysis-Complex.sentiment_analysis_hindiConventions followed to decide the polarity: -
labels consisting of a single value are left undisturbed, i.e. if label = 'pos', then it'll be pos
labels consisting of multiple values separated by '&' are processed. If all the labels are the same ('pos&pos&pos' or 'neg&neg'), then the shortened form of the multiple label is assigned as the final label. For example, if label = 'pos&pos&pos', then final label will be 'pos'.
labels consisting of mixed values ('pos&neg&pos' or 'neg&neu&pos') are… See the full description on the dataset page: https://huggingface.co/datasets/EmmadiVishnu/sentiment_analysis_hindi.sentiment-analysis-UITsentiment_analysis_ladin_italian_manual
Italian-Ladin Sentiment Analysis Dataset (Golden)
This is a manually created sentiment analysis dataset in the Ladin language (Val Badia Variant) paired with the Italian language.
Columns: 'italian', 'ladin', 'label'
Labels: 'pos', 'neg'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
@inproceedings{nlp-ladin-2026,
title = "Towards the First NLP Benchmark for Ladin - an Extremely Low-Resource Language"… See the full description on the dataset page: https://huggingface.co/datasets/ulinnuha/sentiment_analysis_ladin_italian_manual.sentiment_analysis_sharegpt_jsonsentiment_analysis_ladin_italian
Italian-Ladin Sentiment Analysis Dataset
This is a translated sentiment analysis dataset in the Ladin language (Val Badia Variant) from the Italian language.
Columns: 'italian', 'ladin', 'label'
Label: 'pos', 'neg'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
@inproceedings{nlp-ladin-2026,
title = "Towards the First NLP Benchmark for Ladin - an Extremely Low-Resource Language",
author = "Nuha, Ulin… See the full description on the dataset page: https://huggingface.co/datasets/ulinnuha/sentiment_analysis_ladin_italian.Sentiment-analysis-enpo-masentiment_analysis_ladin_italian
Italian-Ladin Sentiment Analysis Dataset
This is a translated sentiment analysis dataset in the Ladin language (Val Badia Variant) from the Italian language.
Columns: 'italian', 'ladin', 'label'
Label: 'pos', 'neg'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
To be announced.
sentiment_analysis_ladin_italian_manual
Italian-Ladin Sentiment Analysis Dataset (Golden)
This is a manually created sentiment analysis dataset in the Ladin language (Val Badia Variant) paired with the Italian language.
Columns: 'italian', 'ladin', 'label'
Labels: 'pos', 'neg'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
To be announced.
sentiment-analysis-se
