datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ham10000-skin-lesion-classifier-datasynthetic-intent-classifier-dataset-v1
Tanaos Intent Classifier Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate intent classification systems — models that can identify user intents in conversational AI applications, such as chatbots and virtual assistants.
Our flagship intent classification model, tanaos-intent-classifier-v1, was trained on this dataset.
Dataset Summary
The dataset contains text… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-intent-classifier-dataset-v1.devseis-ai-act-classifier-v5-data
Devseis AI Act Classifier — v5 training and evaluation data
Newer version: Devseis/devseis-ai-act-classifier-v6-data extends this data with 39 scenarios and a second writing style for every training example.
A Devseis dataset (data_v5_systems). Used to train Devseis/devseis-ai-act-classifier-v5, which powers Caveat.
This is the data the v5 classifier was trained and evaluated on. It is published so the numbers in the
model card and on the Caveat page can be checked. It is… See the full description on the dataset page: https://huggingface.co/datasets/Devseis/devseis-ai-act-classifier-v5-data.ORKG-core-domain-classifier-dataset
ORKG Core Domain Classifier — Metadata
Update Date: 2026-10-01
Research-domain labels produced by the ORKG Core Domain Classifier — one row per document.
The classifier that produces this file lives at https://gitlab.com/TIBHannover/orkg/nlp/experiments/core-domain-classifier.
This card is generated, so please do not edit it by hand.
Load the dataset
from datasets import load_dataset
dataset = load_dataset("TIB/ORKG-core-domain-classifier-dataset", split="full")… See the full description on the dataset page: https://huggingface.co/datasets/TIB/ORKG-core-domain-classifier-dataset.devseis-ai-act-classifier-v6-data
Devseis AI Act Classifier — v6 training and evaluation data
A Devseis dataset (data_v6_systems). Used to train Devseis/devseis-ai-act-classifier-v6, which powers Caveat.
This is the data the v6 classifier was trained and evaluated on. It is published so the numbers in the
model card and on the Caveat page can be checked. It is synthetic, and it is not legal advice.
Files
File
What it is
Rows
Scenarios
data/train_v6.csv
training set (text, label)
938
242… See the full description on the dataset page: https://huggingface.co/datasets/Devseis/devseis-ai-act-classifier-v6-data.song-lyrics-artist-classifierwater-conflict-classifier-evals
Water Conflict Classifier Evaluation Metrics
Evaluation metrics tracking the performance of the Water Conflict Classifier across multiple training iterations and model configurations.
Dataset Summary
This dataset contains evaluation results from training runs of the Water Conflict Classifier, a multi-label SetFit model that identifies water-related conflict events in news headlines. Each row represents one model version with comprehensive performance metrics across three… See the full description on the dataset page: https://huggingface.co/datasets/baobabtech/water-conflict-classifier-evals.c4-website-classifier-datasetTo filter the data for better label quality, label by equivalent V3 and label entry and play with label probability.
resume-domain-classifier-v1-en
Resume-Domain Classifier Dataset v1 (English)
Dataset Description
resume-domain-classifier-v1-en is a large-scale cross-encoder dataset designed for training binary classifiers to detect whether a resume and job description belong to the same professional domain. This dataset is essential for building intelligent ATS (Applicant Tracking System) applications that need to understand domain compatibility between candidates and job postings.
Key Features… See the full description on the dataset page: https://huggingface.co/datasets/0xnbk/resume-domain-classifier-v1-en.pubmed-research-classifier
PubMed Research Classifier Labels
Private lookup table of research / non-research labels for PubMed IDs.
PyPI package
pubmed-research-classifier (≥ 0.3.0 for Hub lookup)
This dataset
Hub release v1.0.0 (labels scored with package 0.2.0 weights)
Docs
Package README on PyPI
Coverage (v1.0.0): labels for all OpenAlex works that have a PubMed ID (PMID),
drawn from the OpenAlex corpus snapshot used in this project, with PubMed-linked
works up to April 2026.… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/pubmed-research-classifier.akai_flow_classifier_seed_pest_schemesales-behaviour-classifier-dataset
sales-behaviour-classifier-dataset
A labelled dataset for turn-level sales behaviour classification. 2,400 samples span across four behavoural categories: WEAK_CONCESSIVE, OBJECTION_HANDLING, FEATURE_DUMPING, and ANCHORING.
The conversations samples were generated via Gemini 3.1 Flash Lite and manually labelled by a human annotator using a written labelling guide.
Dataset Details
Dataset Description
Curated by: [Tom Harcus]
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/TomHarcus/sales-behaviour-classifier-dataset.bank_complaint_intent_classifieralphafold_misinterpretation_classifier_v01AlphaFold Misinterpretation Classifier (AMC) v0.1
Purpose
Help models spot when AlphaFold outputs are used to make claims that go beyond their scope.
Teach restraint.
Promote correct boundaries around structural interpretation.
Columns
claim
misinterpretation_type
reason_hint
action
misinterpretation_type examples
function_from_structure: assuming activity from fold
binding_assertion: assuming ligand interaction
metric_confusion: misreading confidence or aligned error
state_fixation:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/alphafold_misinterpretation_classifier_v01.swedish-pre-a1-scenario-classifier-dataset
Swedish Pre-A1 Scenario Classification Dataset
This dataset contains 150 short Swedish learner sentences for text classification. It is designed for absolute beginner / pre-A1 learners and aligned with beginner Swedish lecture themes.
Labels
food_shop
family_school
health_places
transport
home_places
social_intro
Each label has 25 examples.
Columns
id: unique example id
text: Swedish learner sentence used as model input
english: English translation
chinese:… See the full description on the dataset page: https://huggingface.co/datasets/SeanSha30/swedish-pre-a1-scenario-classifier-dataset.e-commerce-intent-classifier-datasetLLM_Classifier_Dataset_2000text-tone-classifierIPTC-topic-classifier-labelsLLM_Classifier_Dataset_4000text-classifiernews-headlines-classifier
News Headlines Classifier Dataset
350 labeled news headlines for category classification, sentiment analysis, and clickbait detection.
Dataset Structure
Fields
headline: News headline text
category: technology / politics / business / sports / health / science / entertainment / world
sentiment: positive / negative / neutral
clickbait_score: 0 (not clickbait) to 5 (very clickbait)
Splits
train: 280 examples
test: 70 examples
Domain_Classifier_Dataset
ORKG Core Domain Classifier — Metadata
Update Date: 2026-08-14
Research-domain labels produced by the ORKG Core Domain Classifier for documents in the
CORE corpus — one row per document.
The classifier that produces this file lives at https://gitlab.com/TIBHannover/orkg/nlp/experiments/core-domain-classifier.
This card and the CSV are written by python -m app.publish_to_hf in that repository, so
please do not edit them by hand.
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/MikeACedric/Domain_Classifier_Dataset.alphafold_misinterpretation_classifier_v0.2AlphaFold Misinterpretation Classifier
PurposeDetect common ways people overclaim from AlphaFold outputs.
Input fields
protein_context
alphafold_signals
proposed_inference
Required outputReturn one JSON object
misinterpretationyes or no
error_typemust match allowed list
correctionone sentence
Allowed error_type values
no_error
low_confidence_region_overtrust
interdomain_orientation_overclaim
disorder_as_structure
loop_position_overtrust
complex_negation_from_monomer… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/alphafold_misinterpretation_classifier_v0.2.github_fetch_huggingface_terminal_9134_x9v2m6_derived_emotion_classifier
Emotion Classifier Data
A derived dataset used to train an emotion-classification model.
Overview
Dataset ID: DRV-EMOTION
Catalog: CUSTOMER-FEEDBACK-ANALYTICS
Origin: Derived from SRC-ALPHA and SRC-BETA with manual annotation
Records: 8,700
Product Line: Customer Feedback Analytics
Source
This dataset originates from: Derived from SRC-ALPHA and SRC-BETA with manual annotation.
Contents
Cleaned and re-labeled samples for emotion… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/github_fetch_huggingface_terminal_9134_x9v2m6_derived_emotion_classifier.ts-classifier-6yes-no-classifiernews-classifier-1m-urlsbank_complaint_classifierintent-classifier
