datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment.TwitterHjerneRetrieval
TwitterHjerneRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Danish question asked on Twitter with the Hashtag #Twitterhjerne ('Twitter brain') and their corresponding answer.
Task category
t2t
Domains
Social, Written
Reference
https://huggingface.co/datasets/sorenmulli/da-hashtag-twitterhjerne
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/TwitterHjerneRetrieval.twittersemeval2015-pairclassification
TwitterSemEval2015
An MTEB dataset
Massive Text Embedding Benchmark
Paraphrase-Pairs of Tweets from the SemEval 2015 workshop.
Task category
t2t
Domains
Social, Written
Reference
https://alt.qcri.org/semeval2015/task1/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TwitterSemEval2015"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/twittersemeval2015-pairclassification.twitterurlcorpus-pairclassification
TwitterURLCorpus
An MTEB dataset
Massive Text Embedding Benchmark
Paraphrase-Pairs of Tweets.
Task category
t2t
Domains
Social, Written
Reference
https://languagenet.github.io/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TwitterURLCorpus"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run… See the full description on the dataset page: https://huggingface.co/datasets/mteb/twitterurlcorpus-pairclassification.twitter-financial-news-topic
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic.
The dataset holds 21,107 documents annotated with 20 labels:
topics = {
"LABEL_0": "Analyst Update",
"LABEL_1": "Fed | Central Banks",
"LABEL_2": "Company | Product News",
"LABEL_3": "Treasuries | Corporate Debt",
"LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.da-hashtag-twitterhjerne
Dataset Card for "da-hashtag-twitterhjerne"
Danish questions asked on Twitter using the Hashtag "#Twitterhjerne" ('Twitter brain') and their answers.
For each question tweet 2-6 answer tweets are included.
Further details can be found in Section 4.2.3 in the thesis.
Produced by: Søren Vejlgaard Holm under supervision of Lars Kai Hansen and Martin Carsten Nielsen.
Usable for: Question Answering Evaluation.
Contact: Søren Vejlgaard Holm at swiho@dtu.dk or swh@alvenir.ai.
twitter100m_tweets
Dataset Card for "twitter100m_tweets"
Dataset with tweets for this post.
DOI: 10.5281/zenodo.15086029
AfriSenti-twitter-sentimentAfriSenti is the largest sentiment analysis benchmark dataset for under-represented African languages---covering 110,000+ annotated tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and yoruba).SAKE-TwitterArabic_Sentiment_Twitter_Corpus
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/arbml/Arabic_Sentiment_Twitter_Corpus.Arabic_Sentiment_Twitter_Corpus
Dataset Card for "Arabic_Sentiment_Twitter_Corpus"
Source:
https://www.kaggle.com/datasets/mksaad/arabic-sentiment-twitter-corpus
twitter-airline-sentiment
Dataset Card for Twitter US Airline Sentiment
Dataset Summary
This data originally came from Crowdflower's Data for Everyone library.
As the original source says,
A sentiment analysis job about the problems of each major U.S. airline. Twitter data was scraped from February of 2015 and contributors were asked to first classify positive, negative, and neutral tweets, followed by categorizing negative reasons (such as "late flight" or "rude service").
The data we're… See the full description on the dataset page: https://huggingface.co/datasets/osanseviero/twitter-airline-sentiment.pegos-twitter-streamtwittercustomer-support-on-twitter-conversationajgt_twitter_ar
Dataset Card for Arabic Jordanian General Tweets
Dataset Summary
Arabic Jordanian General Tweets (AJGT) Corpus consisted of 1,800 tweets annotated as positive and negative. Modern Standard Arabic (MSA) or Jordanian dialect.
Supported Tasks and Leaderboards
The dataset was published on this paper.
Languages
The dataset is based on Arabic.
Dataset Structure
Data Instances
A binary datset with with negative and positive… See the full description on the dataset page: https://huggingface.co/datasets/komari6/ajgt_twitter_ar.twitter-trending-hashtags
Twitter/X Trending Hashtags (2020-2025)
A comprehensive dataset of trending hashtags on Twitter/X from 2020 to 2025, containing 12,036 unique trend entries across six years, capturing major world events, cultural moments, and viral phenomena.
📊 Dataset Description
This dataset captures trending hashtags from Twitter/X (formerly Twitter) by analyzing Wayback Machine snapshots of trends24.in, providing insights into breaking news, viral content, cultural moments, and… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/twitter-trending-hashtags.nlp_twitter_analysisTwitterArtistsviewer: true
Dataset Card for TwitterArtists (Pixel-Dust)
This dataset is a collection of art and media scraped from various artists and profiles across X (formerly Twitter) and Instagram. It is primarily focused on furry art and similar stylized content, intended for use in training or fine-tuning generative models.
Data Collection & Annotation
Source Data
The images were collected from social media profiles of numerous artists. While the bulk of… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Dust/TwitterArtists.Customer_Support_on_Twitterhate_speech_twitter
Dataset Card for Dataset Name
The dataset is designed to analyze and address hate speech within online platforms. It consists of two sets: the training and testing sets. The two datasets have been labeled and categorized instances of hate speech into nine distinct categories.
Dataset Description
The dataset comprises three key features: tweets, labels (with hate speech denoted as 1 and non-hate speech as 0), and categories (behavior, class, disability, ethnicity, gender… See the full description on the dataset page: https://huggingface.co/datasets/thefrankhsu/hate_speech_twitter.climate_twitter_text_embeddingsAfriSenti-TwitterAfriSenti is the largest sentiment analysis benchmark dataset for under-represented African languages---covering 110,000+ annotated tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and yoruba).twitter_indonesia_sarcastic
Twitter Indonesia Sarcastic
Twitter Indonesia Sarcastic is a dataset intended for sarcasm detection in the Indonesian language. This dataset is introduced in Khotijah et al. (2020), whereby Indonesian tweets are collected and labeled as either sarcastic or non-sarcastic. We took the raw data, and performed several cleaning procedures such as: sentence order re-reversal, deduplication with minHash LSH, PII masking to remove usernames, hashtags, emails, URLs, and finally a random… See the full description on the dataset page: https://huggingface.co/datasets/w11wo/twitter_indonesia_sarcastic.twitter-sentiment-analysisThe Twitter Sentiment Analysis Dataset contains 1,578,627 classified tweets, each row is marked as 1 for positive sentiment and 0 for negative sentiment.
The dataset is based on data from the following two sources:
University of Michigan Sentiment Analysis competition on Kaggle
Twitter Sentiment Corpus by Niek Sanders
Finally, I randomly selected a subset of them, applied a cleaning process, and divided them between the test and train subsets, keeping a balance between
the number of positive and negative tweets within each of these subsets.TwitterHateSpeechTwitter_AI
VISUAL COUNTER TURING TEST (VCT²) — TWITTER DATASET
The Visual Counter Turing Test (VCT²) dataset is introduced in the paper“Visual Counter Turing Test (VCT²): Discovering the Challenges for AI-Generated Image Detection and Introducing Visual AI Index (V_AI)”,accepted at IJCNLP–AACL 2025 and available on arXiv:2411.16754.
This dataset aims to benchmark and analyze the challenges of AI-generated image detection (AGID) using real-world, social media–driven captions and imagery.It… See the full description on the dataset page: https://huggingface.co/datasets/NasrinImp/Twitter_AI.NaijaSenti-TwitterNaijaSenti is the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria — Hausa, Igbo, Nigerian-Pidgin, and Yorùbá — consisting of around 30,000 annotated tweets per language, including a significant fraction of code-mixed tweets.indonesian-twitter-hate-speech-cleaned
Dataset Card for indonesian-twitter-hate-speech-cleaned
Dataset Summary
Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories.
The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/haipradana/indonesian-twitter-hate-speech-cleaned.twitterDataset used in the paper:
A thorough benchmark of automatic text classification
From traditional approaches to large language models
https://github.com/waashk/atcBench
To guarantee the reproducibility of the obtained results, the dataset and its respective CV train-test partitions is available here.
Each dataset contains the following files:
data.parquet: pandas DataFrame with texts and associated encoded labels for each document.
split_<k>.pkl: pandas DataFrame with k-cross validation… See the full description on the dataset page: https://huggingface.co/datasets/waashk/twitter.
