datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-label-class-github-issues-text-classification
Dataset Card for "multi-label-class-github-issues-text-classification"
More Information needed
multi-label-food-recognition
Multi-Label Food Recognition Dataset
This is a multi-label food recognition dataset generated from single-class food images.
Each image contains 2-5 different food items composited together using natural composition methods.
Dataset Details
Total Images: 13,000
Training Images: 10,400 (80%)
Validation Images: 2,600 (20%)
Number of Classes: 90
Labels per Image: 2-5 labels
Image Format: RGB, 512x512 pixels
File Format: Parquet
Dataset Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimdaud/multi-label-food-recognition.PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.AID_MultiLabel
Dataset Card for "AID_MultiLabel"
Licensing Information
CC0: Public Domain
Citation Information
Imagery:
AID: A benchmark data set for performance evaluation of aerial scene classification
Multilabels:
Relation Network for Multi-label Aerial Image Classification
@article{xia2017aid,
title = {AID: A benchmark data set for performance evaluation of aerial scene classification},
author = {Xia, Gui-Song and Hu, Jingwen and Hu, Fan and Shi, Baoguang… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/AID_MultiLabel.arxiv-abstract-multilabelwds_voc2007_multilabelid_multilabel_hsThe ID_MULTILABEL_HS dataset is collection of 13,169 tweets in Indonesian language,
designed for hate speech detection NLP task. This dataset is combination from previous research and newly crawled data from Twitter.
This is a multilabel dataset with label details as follows:
-HS : hate speech label;
-Abusive : abusive language label;
-HS_Individual : hate speech targeted to an individual;
-HS_Group : hate speech targeted to a group;
-HS_Religion : hate speech related to religion/creed;
-HS_Race : hate speech related to race/ethnicity;
-HS_Physical : hate speech related to physical/disability;
-HS_Gender : hate speech related to gender/sexual orientation;
-HS_Gender : hate related to other invective/slander;
-HS_Weak : weak hate speech;
-HS_Moderate : moderate hate speech;
-HS_Strong : strong hate speech.russian-toxic-comments-multilabel
Russian Toxic Comments Multi-label Dataset
Dataset Description
Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности.
Цель
Обучение модели для автоматического обнаружения трех типов токсичного контента:
Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань
Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.UC_Merced_LandUse_MultiLabel
Dataset Card for "UC_Merced_LandUse_MultiLabel"
Licensing Information
Public Domain; “Map services and data available from U.S. Geological Survey, National Geospatial Program.”
Citation Information
Imagery:
Bag-of-visual-words and spatial extensions for land-use classification
Multilabels:
Multilabel Remote Sensing Image Retrieval Using a Semisupervised Graph-Theoretic Method
@inproceedings{yang2010bag,
title = {Bag-of-visual-words and spatial… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/UC_Merced_LandUse_MultiLabel.icd10cm-multilabel-promptawesome-japanese-nlp-multilabel-dataset
Dataset overview
This is a dataset for Japanese natural language processing with multi-label annotations of research field labels for GitHub repositories in the NLP domain.
Please refer to this paper for the specific method of constructing the dataset. It is written in Japanese.
Input and Output
Input: Information from GitHub repositories (description, README text, PDF text, screenshot images)
Output: Multi-label classification of NLP research fields
Problem Setting of the… See the full description on the dataset page: https://huggingface.co/datasets/taishi-i/awesome-japanese-nlp-multilabel-dataset.scikit-learn-issues-multilabel
🧩 Scikit-learn GitHub Issues – Multilabel Dataset
This dataset contains GitHub issues from the scikit-learn repository, prepared for multilabel NLP tasks such as issue tagging, automated triage, and semantic search.
Each row corresponds to one issue-comment context, making the dataset suitable for real-world developer tooling.
📌 Motivation
GitHub issues are a critical signal in open-source projects:
Bug tracking
Feature requests
Documentation improvements… See the full description on the dataset page: https://huggingface.co/datasets/Talip7/scikit-learn-issues-multilabel.prachathai67k-tha-multilabelclassificationref: https://github.com/PyThaiNLP/prachathai-67k
Multilabel-Portrait-18K
Multilabel-Portrait-18K
Multilabel-Portrait-18K is a multi-label portrait classification dataset designed to analyze and categorize different styles of portrait images. It supports classification into the following four portrait types:
0 — Anime Portrait
1 — Cartoon Portrait
2 — Real Portrait
3 — Sketch Portrait
This dataset is ideal for training and evaluating machine learning models in the domain of portrait-style classification. The goal is to enable accurate recognition… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Multilabel-Portrait-18K.prachathai67k-tha-multilabelclassification
Prachathai67k_tha_MultiLabelClassification
Deduplicated copy of kornwtp/prachathai67k-tha-multilabelclassification.
Splits
split
rows
train
67,488
tram-attack-multilabel-clean
TRAM ATT&CK Multi-Label (cleaned, with leak-free splits)
Sentence-level multi-label mapping of cyber threat intelligence prose to MITRE
ATT&CK technique IDs. Derived from MITRE CTID's
TRAM corpus,
deduplicated and republished with two split schemes so that leakage can be
measured rather than assumed.
Everything here is regenerated by python scripts/01_build_dataset.py. No row
was edited by hand.
Why this exists
The upstream corpus contains 19,178 sentences drawn… See the full description on the dataset page: https://huggingface.co/datasets/ctokx/tram-attack-multilabel-clean.PubMed-MultiLabel-MeSH
PubMed MultiLabel Text Classification (MeSH)
A dataset of 50,000 PubMed biomedical articles, each manually annotated
by domain experts with MeSH (Medical Subject Headings) labels. With
21,918 unique labels and a mean of ~12.7 labels per document, this is a
densely-labeled extreme multi-label classification benchmark.
Dataset Description
Property
Value
Train examples
40,000
Test examples
10,000
Total unique MeSH labels
21,918
Mean labels per document
~12.7… See the full description on the dataset page: https://huggingface.co/datasets/Tellurio/PubMed-MultiLabel-MeSH.Multilabel-GeoSceneNet-16K
Multilabel-GeoSceneNet-16K
Multilabel-GeoSceneNet-16K is a geospatial image dataset for multi-label scene classification. Each image may belong to one or more geographic scene categories, making it suitable for multi-label learning tasks in remote sensing and geospatial analytics.
Dataset Summary
Task: Multi-label Image Classification
Modalities: Image
Total Images: 16,033
Split: Train (100%)
Labels: 7 categories (multi-label)
License: Apache-2.0
Size: ~227 MB… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Multilabel-GeoSceneNet-16K.MADE-Multilabel-Benchmark
MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events (ACL 2026)
Authors: Raunak Agarwal, Markus Wenzel, Simon Baur, Jonas Zimmer, George Harvey, Jackie Ma
Blog; Project Page; Github; Arxiv
Abstract: Machine learning in high-stakes domains such as healthcare requires not only strong predictive performance but also reliable uncertainty quantification (UQ) to support human oversight.
Multi-label text… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/MADE-Multilabel-Benchmark.toxic-russian-comments-multilabel
Russian Toxic Comments Multi-Label Balanced Dataset
Описание
Этот датасет создан для задачи Multi-Task классификации токсичности русскоязычных комментариев. Датасет содержит сбалансированные примеры с тремя бинарными метками:
profanity: наличие нецензурной лексики (мат)
threat: наличие угроз
illegal: запросы на незаконные действия (прокси-метка на основе THREAT + INSULT)
Структура данных
Датасет содержит следующие поля:
Поле
Тип
Описание… See the full description on the dataset page: https://huggingface.co/datasets/dbrovkin/toxic-russian-comments-multilabel.gklmip-news-khm-multilabelclassificationtweet_multilabel_subset_completionhoasa-ind-multilabelclassificationdengue-fil-multilabelclassificationprachathai67k-tha-multilabelclassificationminified-diverseful-multilabels
A minified, clean and annotated version of DiverseVul
Dataset Summary
This is a minified, clean and deduplicated version of the DiverseVul dataset.
We publish this version to help practionners in their code vulnerability detection research.
Data Structure & Overview
Number of samples: 23847
Features: func (the C/C++ code)cwe (the CWE weakness, see table below)
Supported Programming Languages: C/C++
Supported CWE Weaknesses:
Label
Description… See the full description on the dataset page: https://huggingface.co/datasets/lemon42-ai/minified-diverseful-multilabels.philosophy-schools-multilabel
Dataset Card for "philosophai-papers-complete"
More Information needed
vlsp2018sa-hotel-vie-multilabelclassificationcasa-ind-multilabelclassificationgklmip-news-khm-multilabelclassification
GKLMIPNews_khm_MultiLabelClassification
Deduplicated copy of kornwtp/gklmip-news-khm-multilabelclassification.
Splits
split
rows
test
1,398
train
3,959
validation
1,376
