datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multiclass-sentiment-analysis-dataset
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Sp1786/multiclass-sentiment-analysis-dataset.multiclassHelminthstomato-variant-a-multiclassscam-classification-multiclass
Scam Text Classification - Multi-Class Dataset
Overview
This dataset is an enhanced version of the original binary scam classification dataset, now with 5 multi-class categories for more granular scam detection.
Dataset Structure
Original Dataset
14,000 rows of Indian-context SMS/email-style messages
Binary labels: 0 (legit) / 1 (scam)
Domain-specific: Indian banks, UPI, Aadhaar, government agencies
Updated Multi-Class Categories… See the full description on the dataset page: https://huggingface.co/datasets/Shade63/scam-classification-multiclass.multiclass-email-classificationThis dataset comprises of more than 2000 emails across multiple categories, which can he helpful for tasks like LLM training and fine-tuning. The dataset is also provided with a python script that would generate emails automatically
The dataset contains email across 10 different categories namely, "Business", "Personal", "Promotions", "Customer Support", "Job Application", "Finance & Bills", "Events & Invitations", "Travel & Bookings", "Reminders", "Newsletters"
Total emails: 2105
Label… See the full description on the dataset page: https://huggingface.co/datasets/imnim/multiclass-email-classification.multiclass-email-classificationThis dataset comprises of more than 2000 emails across multiple categories, which can he helpful for tasks like LLM training and fine-tuning. The dataset is also provided with a python script that would generate emails automatically
The dataset contains email across 10 different categories namely, "Business", "Personal", "Promotions", "Customer Support", "Job Application", "Finance & Bills", "Events & Invitations", "Travel & Bookings", "Reminders", "Newsletters"
Total emails: 2105
Label… See the full description on the dataset page: https://huggingface.co/datasets/Harshal3467/multiclass-email-classification.Generalization-MultiClass-CLINC150-ROSTDThis dataset merge 3 datasets and have two setup for experiments in generalisation for multi-class clasificacitino task.
ID, near-OOD, covariate-shitf: CLINC150
ID, near-OOD, covariate-shitf: ROSTD+OOD (fbreleasecoarse version)
far-OOD Validation: SST2
far-OOD Test: News Category (v3)
reddit-AITA-submissions-and-comments-multiclassmulticlassfunctional-multiclass-gamba
GAMBA Functional Region Multiclass
This representation benchmark asks whether frozen sequence embeddings
separate genomic functional categories. Each row is one annotated region;
label == category.
Loading
from datasets import load_dataset
full_bidi = load_dataset(
"Taykhoom/functional-multiclass-gamba",
"full-bidi",
split="all",
)
paper_test = full_bidi.filter(
lambda row: row["split"] == "test"
and row["category"] != "noncoding_regions"
)… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/functional-multiclass-gamba.multi_class_solidity_function_vulnerabilty
Dataset Card for "multi_class_solidity_function_vulnerabilty"
More Information needed
chili_disease.v12i.multiclasssih-dataset-multiclasssafety-llama-multiclassmulti-class-food-dataset
Food Classification Dataset
This dataset consists of multiple subsets of food images designed for training and evaluating deep learning models for food classification. It includes full-scale and reduced versions to facilitate experimentation with different data sizes.
Dataset Overview
File Name
Size
Description
101_food_classes_10_percent.zip
~1.34GB
Contains 10% of the 101_food_classes dataset.
10_food_classes.zip
~393MB
Contains images for 10 different… See the full description on the dataset page: https://huggingface.co/datasets/mhamza-007/multi-class-food-dataset.multiclass-sentiment-analysis-dataset
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Ajain1391/multiclass-sentiment-analysis-dataset.Multi-Class-Waste-Image-Classification-Dataset
Multi-Class Waste Image Classification Dataset
Overview
This dataset is curated for multi-class waste classification using computer vision and deep learning techniques. It provides a structured, high-quality benchmark for training custom CNN architectures, benchmarking transfer learning models, and building automated waste sorting pipelines.
Dataset Structure
The dataset follows a standard directory format compatible with common deep learning… See the full description on the dataset page: https://huggingface.co/datasets/Kishore2412/Multi-Class-Waste-Image-Classification-Dataset.Multi-Class-OVThis is a multi-class version of open vocabulary segmentation dataset by randomly merging annotations from several classes, including ADE20K(A-150), PASCAL Context59(PC-59), and PASCAL VOC20(PAS-20).
You can use them according to [📂 GitHub].
If you find this project useful in your research, please consider citing:
@article{wang2025alto,
title={ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation},
author={Wang, Lingfeng and Lin, Hualing and Chen, Senda and Wang, Tao and… See the full description on the dataset page: https://huggingface.co/datasets/yayafengzi/Multi-Class-OV.multiclass-text-classification-datasetbashkir-news-multiclass
Dataset Card for Bashkir News Multiclass Classification Dataset
Dataset Details
Dataset Description
This dataset contains 17,897 Bashkir-language news and analytical articles annotated with 19 thematic categories for multiclass text classification tasks. Each article belongs to exactly one category. The categories range from news and society to culture, education, and sports. The dataset was created to support NLP research and application… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-multiclass.Llama3_8b-emotion_multiclass-Plutchik
Description
This is a dataset for emotion classification of text sentences.
The dataset is a CSV file with 6,540 sentences. Each row has two columns: the first one has the sentence text, and the second one has its main emotion:
"text";"emotion"
The emotion can be one of Plutchik's eight emotion groups plus a neutral category. The sentence counts for each emotion are:
joy: 611 (9.34%)
sadness: 748 (11.44%)
trust: 735 (11.24%)
disgust: 838 (12.81%)
fear: 579 (8.85%)
anger: 743… See the full description on the dataset page: https://huggingface.co/datasets/uavster/Llama3_8b-emotion_multiclass-Plutchik.UA_speech_multiclass
Dataset Card for "UA_speech_multiclass"
More Information needed
multiclass-sentiment-analysis-dataset
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Renture666/multiclass-sentiment-analysis-dataset.YOLOv8-Multiclass-Object-Detection-Dataset
DATASET SAMPLE
Duality.ai just released a 1000 image dataset used to train a YOLOv8 model in multiclass object detection -- and it's 100% free!
Just create an EDU account here.
This HuggingFace dataset is a 20 image and label sample, but you can get the rest at no cost by creating a FalconCloud account. Once you verify your email, the link will redirect you to the dataset page.
What makes this dataset unique, useful, and capable of bridging the Sim2Real gap?
The digital twins are… See the full description on the dataset page: https://huggingface.co/datasets/duality-robotics/YOLOv8-Multiclass-Object-Detection-Dataset.casa_3sec_multiclassMulti-Class-Waste-Image-Classification-Dataset
Multi-Class Waste Image Classification Dataset
Overview
This dataset is curated for multi-class waste classification using computer vision and deep learning techniques. It provides a structured, high-quality benchmark for training custom CNN architectures, benchmarking transfer learning models, and building automated waste sorting pipelines.
Dataset Structure
The dataset follows a standard directory format compatible with common deep learning… See the full description on the dataset page: https://huggingface.co/datasets/Mayank14/Multi-Class-Waste-Image-Classification-Dataset.Duckietown-Multiclass-Semantic-Segmentation-Dataset
Multiclass Semantic Segmentation Duckietown Dataset
A dataset of multiclass semantic segmentation image annotations for the first 250 images of the "Duckietown Object Detection Dataset".
Raw Image
Segmentated Image
Semantic Classes
This dataset defines 8 semantic classes (7 distinct classes + implicit background class):
Class
XML Label
Description
Color (RGB)
Ego Lane
Ego Lane
The lane the agent is supposed to be driving in (default right-hand… See the full description on the dataset page: https://huggingface.co/datasets/hamnaanaa/Duckietown-Multiclass-Semantic-Segmentation-Dataset.prompt_injection_multiclass_datasetlabel_names = dataset["train"].features["label"].names
Class Names: ['Ignore', 'Jailbreak', 'Summarization']
Sources:
Lakera/gandalf_ignore_instructions
Lakera/gandalf_summarization
rubend18/ChatGPT-Jailbreak-Prompts
Multiclasscns_course_24-multiclassMulticlass dataset for the 2024 CNS Data Science Course, NLP segment. Consists of 600 synthetic abstracts generated by GPT-4o (300 cranial, 300 spine).
