Team Ai
Datasetpublic

bdstar/twitter-sentiment-analysis

๐Ÿฆ Twitter Sentiment Analysis (bdstar/twitter-sentiment-analysis) ๐Ÿง  Overview A refined and merged version of Twitter text sentiment datasets, providing a clean and well-balanced dataset for sentiment classification across three sentiment categories:positive, negative, and neutral. This dataset is split into three parts โ€” train, test, and validation โ€” each sourced from highly reputable open datasets.It is designed for training, evaluating, and benchmarking NLPโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/bdstar/twitter-sentiment-analysis.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes67downloads
Dataset Card

๐Ÿฆ Twitter Sentiment Analysis (bdstar/twitter-sentiment-analysis)

๐Ÿง  Overview

A refined and merged version of Twitter text sentiment datasets, providing a clean and well-balanced dataset for sentiment classification across three sentiment categories: `positive`, `negative`, and `neutral`.

This dataset is split into three parts โ€” train, test, and validation โ€” each sourced from highly reputable open datasets. It is designed for training, evaluating, and benchmarking NLP models for Twitter Sentiment Analysis and other social media text classification tasks.


๐Ÿ—‚๏ธ Dataset Splits

SplitSource DatasetRowsFile SizeLink
TrainTwitter Sentiment Dataset (3M labeled rows)3,142,209361 MBKaggle Dataset
TestSentiment140 Dataset1,600,001198 MBKaggle Dataset
ValidationMTEB Tweet Sentiment Extraction31,0153.45 MBHugging Face Dataset

๐Ÿงฉ Column Descriptions

ColumnTypeDescription
IDIntegerAuto-incremental unique ID for each row
textStringTweet text content
labelStringSentiment category โ€” one of positive, negative, or neutral

๐Ÿ“Š Dataset Summary

PropertyValue
Total Rows4,773,225
Columns3
File FormatsJSON / Parquet / Pandas / Polars / Croissant
LicenseMIT
AuthorMd Abdullah Al Mamun
Year2025
SourceRefined version of Twitter Sentiment Dataset

๐Ÿ“ˆ Detailed Statistics

๐Ÿ‹๏ธโ€โ™‚๏ธ Train Set

Source: Twitter Sentiment Dataset (3M labeled rows) File Size: 361 MB Rows: 3,142,209

LabelCountPercentage
Positive1,571,10450.0%
Negative1,571,10550.0%

๐Ÿงช Test Set

Source: Sentiment140 File Size: 198 MB Rows: 1,600,001

LabelCountPercentage
Positive800,00050.0%
Negative800,00150.0%

๐Ÿงญ Validation Set

Source: MTEB โ€“ Tweet Sentiment Extraction File Size: 3.45 MB Rows: 31,015

LabelCountPercentage
Neutral12,56140.5%
Positive9,67631.2%
Negative8,77828.3%

๐Ÿ’ก Usage Example (Python)

python
from datasets import load_dataset

# Load dataset from Hugging Face
dataset = load_dataset("bdstar/twitter-sentiment-analysis")

# Access splits
train = dataset["train"]
test = dataset["test"]
validation = dataset["validation"]

# Display sample
print(train[0])

๐Ÿท๏ธ Citation

If you use this dataset in your research or application, please cite as:

bibtex
@dataset{bdstar2025twitter,
  title        = {Twitter Sentiment Analysis (Refined Dataset)},
  author       = {Md Abdullah Al Mamun},
  year         = {2025},
  howpublished = {Hugging Face},
  url          = {https://huggingface.co/datasets/bdstar/twitter-sentiment-analysis}
}

๐Ÿ“ฌ Contact

For questions, improvements, or collaboration: Author: Md Abdullah Al Mamun ๐Ÿ“ง Email: mamunbd.ruet@gmail.com ๐ŸŒ Website: TechNTuts