datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-label-class-github-issues-text-classification
Dataset Card for "multi-label-class-github-issues-text-classification"
More Information needed
multi-label-food-recognition
Multi-Label Food Recognition Dataset
This is a multi-label food recognition dataset generated from single-class food images.
Each image contains 2-5 different food items composited together using natural composition methods.
Dataset Details
Total Images: 13,000
Training Images: 10,400 (80%)
Validation Images: 2,600 (20%)
Number of Classes: 90
Labels per Image: 2-5 labels
Image Format: RGB, 512x512 pixels
File Format: Parquet
Dataset Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimdaud/multi-label-food-recognition.PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.wds_voc2007_multilabelarxiv-abstract-multilabelrussian-toxic-comments-multilabel
Russian Toxic Comments Multi-label Dataset
Dataset Description
Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности.
Цель
Обучение модели для автоматического обнаружения трех типов токсичного контента:
Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань
Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.tram-attack-multilabel-clean
TRAM ATT&CK Multi-Label (cleaned, with leak-free splits)
Sentence-level multi-label mapping of cyber threat intelligence prose to MITRE
ATT&CK technique IDs. Derived from MITRE CTID's
TRAM corpus,
deduplicated and republished with two split schemes so that leakage can be
measured rather than assumed.
Everything here is regenerated by python scripts/01_build_dataset.py. No row
was edited by hand.
Why this exists
The upstream corpus contains 19,178 sentences drawn… See the full description on the dataset page: https://huggingface.co/datasets/ctokx/tram-attack-multilabel-clean.scikit-learn-issues-multilabel
🧩 Scikit-learn GitHub Issues – Multilabel Dataset
This dataset contains GitHub issues from the scikit-learn repository, prepared for multilabel NLP tasks such as issue tagging, automated triage, and semantic search.
Each row corresponds to one issue-comment context, making the dataset suitable for real-world developer tooling.
📌 Motivation
GitHub issues are a critical signal in open-source projects:
Bug tracking
Feature requests
Documentation improvements… See the full description on the dataset page: https://huggingface.co/datasets/Talip7/scikit-learn-issues-multilabel.icd10cm-multilabel-promptPubMed-MultiLabel-MeSH
PubMed MultiLabel Text Classification (MeSH)
A dataset of 50,000 PubMed biomedical articles, each manually annotated
by domain experts with MeSH (Medical Subject Headings) labels. With
21,918 unique labels and a mean of ~12.7 labels per document, this is a
densely-labeled extreme multi-label classification benchmark.
Dataset Description
Property
Value
Train examples
40,000
Test examples
10,000
Total unique MeSH labels
21,918
Mean labels per document
~12.7… See the full description on the dataset page: https://huggingface.co/datasets/Tellurio/PubMed-MultiLabel-MeSH.awesome-japanese-nlp-multilabel-dataset
Dataset overview
This is a dataset for Japanese natural language processing with multi-label annotations of research field labels for GitHub repositories in the NLP domain.
Please refer to this paper for the specific method of constructing the dataset. It is written in Japanese.
Input and Output
Input: Information from GitHub repositories (description, README text, PDF text, screenshot images)
Output: Multi-label classification of NLP research fields
Problem Setting of the… See the full description on the dataset page: https://huggingface.co/datasets/taishi-i/awesome-japanese-nlp-multilabel-dataset.prachathai67k-tha-multilabelclassificationref: https://github.com/PyThaiNLP/prachathai-67k
prachathai67k-tha-multilabelclassification
Prachathai67k_tha_MultiLabelClassification
Deduplicated copy of kornwtp/prachathai67k-tha-multilabelclassification.
Splits
split
rows
train
67,488
MADE-Multilabel-Benchmark
MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events (ACL 2026)
Authors: Raunak Agarwal, Markus Wenzel, Simon Baur, Jonas Zimmer, George Harvey, Jackie Ma
Blog; Project Page; Github; Arxiv
Abstract: Machine learning in high-stakes domains such as healthcare requires not only strong predictive performance but also reliable uncertainty quantification (UQ) to support human oversight.
Multi-label text… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/MADE-Multilabel-Benchmark.tweet_multilabel_subset_completiontoxic-russian-comments-multilabel
Russian Toxic Comments Multi-Label Balanced Dataset
Описание
Этот датасет создан для задачи Multi-Task классификации токсичности русскоязычных комментариев. Датасет содержит сбалансированные примеры с тремя бинарными метками:
profanity: наличие нецензурной лексики (мат)
threat: наличие угроз
illegal: запросы на незаконные действия (прокси-метка на основе THREAT + INSULT)
Структура данных
Датасет содержит следующие поля:
Поле
Тип
Описание… See the full description on the dataset page: https://huggingface.co/datasets/dbrovkin/toxic-russian-comments-multilabel.gklmip-news-khm-multilabelclassificationhoasa-ind-multilabelclassificationdengue-fil-multilabelclassificationprachathai67k-tha-multilabelclassificationminified-diverseful-multilabels
A minified, clean and annotated version of DiverseVul
Dataset Summary
This is a minified, clean and deduplicated version of the DiverseVul dataset.
We publish this version to help practionners in their code vulnerability detection research.
Data Structure & Overview
Number of samples: 23847
Features: func (the C/C++ code)cwe (the CWE weakness, see table below)
Supported Programming Languages: C/C++
Supported CWE Weaknesses:
Label
Description… See the full description on the dataset page: https://huggingface.co/datasets/lemon42-ai/minified-diverseful-multilabels.philosophy-schools-multilabel
Dataset Card for "philosophai-papers-complete"
More Information needed
vlsp2018sa-hotel-vie-multilabelclassificationcasa-ind-multilabelclassificationhatespeech-ind-multilabelclassificationnetifier-ind-multilabelclassificationvlsp2018sa-restaurant-vie-multilabelclassificationref: https://github.com/vndee/awsome-vietnamese-nlp
multi-label-web-categorization
Multi-Label Web Page Classification Dataset
Dataset Description
The Multi-Label Web Page Classification Dataset is a curated dataset containingweb page titles and snippets, extracted from the CC-Meta25-1M dataset. Each entry has been automatically categorized into multiple predefined categories using ChatGPT-4o-mini.
This dataset is designed for multi-label text classification tasks, making it ideal for training and evaluating machine learning models in web content… See the full description on the dataset page: https://huggingface.co/datasets/tshasan/multi-label-web-categorization.truevoice-intent-tha-multilabelclassificationhatespeech-ind-multilabelclassification
