datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-detector-data
AI Detector Predictions Dataset
A continuously-growing collection of AI text detection predictions with optional user feedback, generated from the AI Text Detector Space.
Every time someone analyzes text or a URL on the Space, the prediction is appended to this dataset. Users can also click "Correct" or "Incorrect" to provide feedback, which gets stored alongside the prediction.
Schema
Field
Type
Description
id
string
Unique 12-char hex identifier… See the full description on the dataset page: https://huggingface.co/datasets/adaptive-classifier/ai-detector-data.high-accuracy-email-classifier
High-Accuracy Email Classification Dataset
Dataset Description
This dataset contains 12,000+ emails across 6 categories, specifically curated for high-accuracy email classification tasks. The dataset achieves 98%+ classification accuracy with appropriate models.
Categories
The dataset includes emails from the following categories:
Category
Count
Description
Emoji
Forum
~2,000
Forum posts, discussions, and community notifications
🗣️
Promotions
~2… See the full description on the dataset page: https://huggingface.co/datasets/jason23322/high-accuracy-email-classifier.arxiv-classifier
arXiv Classifier Data
Usage:
from datasets import load_dataset, DownloadMode
# download from HuggingFace
dataset = load_dataset('mlcore/arxiv-classifier', name=<CONFIG NAME>)
# load from G2
dataset = load_dataset('/share/nikola/arxiv_classifier/data/arxiv-classifier', name=<CONFIG NAME>)
To force the dataset to be re-generated:
dataset = load_dataset('/share/nikola/arxiv_classifier/data/arxiv-classifier', name=<CONFIG NAME>, download_mode=DownloadMode.FORCE_REDOWNLOAD)
See:… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/arxiv-classifier.high-accuracy-email-classifier-indonesian
High-Accuracy Email Classification Dataset Indonesian Translation
This dataset is an Indonesian translation/adaptation of
jason23322/high-accuracy-email-classifier.
Contribution
The original dataset, labels, IDs, and split membership come from
jason23322/high-accuracy-email-classifier, released under the Apache 2.0 license.
This repository contributes the Indonesian translation:
subject, body, and text are translated into Indonesian.
id, category, and category_id… See the full description on the dataset page: https://huggingface.co/datasets/chairulridjal/high-accuracy-email-classifier-indonesian.LoRA-Samples-Intention-Classifier
Dataset Card for LoRA-Samples-Intention-Classifier
Dataset to fine-tune Qwen3-4B-Instruct-2507-LoRA-Intent-Classifier
Dataset Details
Dataset Description
This dataset includes over 10K samples of prompt-intention id pairs for the AI CS agent generator.It is used to fine-tune a small model that powers this agent, reaching a balance of accuracy, efficiency and cost.
Curated by: Li Tuo
Language(s) (NLP): Chinese (primary), English (partial support)
License:… See the full description on the dataset page: https://huggingface.co/datasets/lituokobe/LoRA-Samples-Intention-Classifier.mockgen-classifier-data-v1
Schema-property semantic classification data (v1)
Field metadata — property name, label, containing entity, neighbouring property names, type, annotations — paired with a semantic hint such as country, currency, person_full_name or gl_account. Each row describes a schema field, never a business record. No values are included.
The task: given only what a schema says about a field, decide what the field means, so a mock-data generator can produce something plausible for it.… See the full description on the dataset page: https://huggingface.co/datasets/Unseen1980/mockgen-classifier-data-v1.needle-event-classifier-eval
Sentinel Needle Event Classifier — Evaluation Splits
Held-out test and validation splits used to evaluate the
Shreyas94/needle-event-classifier
model: a 13-class financial-news event classifier built on frozen
Cactus Compute Needle 3 embeddings.
Important: this is NOT the training set
This repo contains only the frozen 678-row test split and the
424-row validation split, both human-reviewed. These are the exact
splits the model's reported metrics (91.45% accuracy… See the full description on the dataset page: https://huggingface.co/datasets/Shreyas94/needle-event-classifier-eval.refusal-classifier-data
Refusal Classifier Training Data
The english_only subset is taken from the downsampled configuration of agentlans/en-chat-refusal.
The multilingual subset is created by mixing agentlans/multilingual-chat-refusal with the english_only subset in a 1:1 ratio.
All conversations have been formatted using role tokens and ellipses <|...|> for long texts.
The resulting files are shuffled and split 80% training, 20% testing.
prompt-safety-classifier-100k
🛡️ Malicious vs. Benign Prompt Classifier & Agent Guardrail Dataset (100,000 Rows)
A 100,000-row multi-class safety dataset engineered specifically for training LLM safety classifiers, guardrail models, prompt injection detectors, and autonomous AI agent tool execution defenses.
📊 Dataset Summary
Unlike standard binary safety datasets, this dataset provides a 6-class safety taxonomy, 1–5 severity scoring, attack technique classification, surface-vs-intent flags… See the full description on the dataset page: https://huggingface.co/datasets/Goutam112/prompt-safety-classifier-100k.pmjay-classifier-sft
PM-JAY Health Benefit Package Classifier
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Medical specialty + procedure → PM-JAY HBP code, package name, and rate
Why download this
Automate PM-JAY / Ayushman Bharat claim processing. Map procedures to Health Benefit Packages for pre-authorization and reimbursement.… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/pmjay-classifier-sft.VV-classifier-2.0-payment-runs_v1.1VV-classifier-2.0-payment-runs-v1.2toxicity-classifier-dataset-v3
Multilingual Toxicity Dataset V3
124K balanced samples for binary toxicity classification across three languages.
Author
Görkem Yıldız
GitHub: gorkem371
Website: gorkemyildiz.com
Dataset Details
Split
Samples
Train
105,945
Valid
18,684
Total
124,629
Languages
Turkish (~34%) — sourced primarily from Overfit-GM
Arabic (~32%) — sourced primarily from arabic-hate-speech-superset
English (~34%) — sourced from toxic_conversations_50k… See the full description on the dataset page: https://huggingface.co/datasets/gorkem371/toxicity-classifier-dataset-v3.edu-classifier-stress-test
Educational & Quality Classifier Stress Test (16 Docs)
A lightweight 16-sample synthetic stress-test benchmark designed to test the floor and ceiling behavior of educational and text quality classifiers (e.g. TinyBERT FineWeb-Edu classifiers, DistilBERT, BERT-109M).
Dataset Structure
Each record contains:
label: Name of the specific test case (e.g. casino_spam, stack_trace, thin_content, wikipedia_style)
text: The raw document text
Overview of Test… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/edu-classifier-stress-test.genai-smartcity-classifier
GenAI Smart City Classification Dataset
A curated and augmented dataset for training and evaluating transformer models that classify whether a text (e.g., abstract segment, contribution sentence) describes a Generative AI (GenAI) application in the context of smart cities.
The full codebase for this project can be found [here]here.
1. Dataset Purpose
Supports binary classification:
GenAI used for smart city application
Not related
Used to fine-tune the DeBERTa model in… See the full description on the dataset page: https://huggingface.co/datasets/joaocarlosnb/genai-smartcity-classifier.korean-sign-word-classifier-mediapipe-test-100
KSL Test-100 Evaluation Dataset
This dataset contains 100 isolated Korean sign-language word videos used for the Test-100 evaluation of the MediaPipe keypoint word classifier.
Contents
videos/: MP4 isolated Korean sign-language word clips.
metadata.csv: Video-level labels, model predictions, and evaluation fields.
evaluation/: Full evaluation reports and machine-readable result files.
Evaluation Summary
Model:… See the full description on the dataset page: https://huggingface.co/datasets/Seoyoung07/korean-sign-word-classifier-mediapipe-test-100.VV-classifier-2.0-payment-runs-v1.3VV-classifier-2.0-payment-runs-v1.7VV-classifier-2.0-payment-runs-v1.6cuentas-claras-sat-classifier
Cuentas Claras — SAT Transaction Classifier dataset
Instruction-tuning data that teaches a small model to classify a free-text
transaction description into its SAT account code, deductibility, and IVA
treatment — the core fine-tune (🎯 Well-Tuned) behind the Cuentas Claras accountant
agent.
Format
Chat-format JSONL. Each row:
{
"messages": [
{"role": "system", "content": "Eres un clasificador contable mexicano. ..."},
{"role": "user", "content":… See the full description on the dataset page: https://huggingface.co/datasets/eldinosaur/cuentas-claras-sat-classifier.VV-classifier-2.0-payment-runs-v1.4insurance-classifier-sft
Insurance Coverage Classifier (Stark Law DHS)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
CPT/HCPCS codes → Stark Law DHS classification + compliance notes
Why download this
Compliance automation for physician self-referral rules. Identify which services are Designated Health Services under Stark Law Section… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/insurance-classifier-sft.agent-task-classifier-dataset
agent-task-classifier-dataset
5,304 labeled English examples for training classifiers that map a user request (to a
personal-assistant agent) to the agent-loop shape it demands — the skeleton of the
high-level process a task of that kind must follow — rather than its topic.
Companion model: kumarabd1996/agent-task-classifier
(a fine-tuned laya checkpoint; 20/24 on the
boundary drills included here, vs 8/24 for the untuned base).
The eleven labels
label
the… See the full description on the dataset page: https://huggingface.co/datasets/kumarabd1996/agent-task-classifier-dataset.VV-classifier-2.0-payment-runs-v1.5math-correctness-classifier-aimes-amo
RedaAlami/math-correctness-classifier-aimes-amo
Dataset Description
This dataset contains mathematical reasoning problems and model responses formatted for training correctness classifiers. Each record includes a problem statement, a model's solution attempt, and a binary label indicating correctness.
The dataset spans three benchmarks:
AIME 2024: American Invitational Mathematics Examination 2024
AIME 2025: American Invitational Mathematics Examination 2025
AMO: Asian… See the full description on the dataset page: https://huggingface.co/datasets/RedaAlami/math-correctness-classifier-aimes-amo.math-correctness-classifier_64rollouts
RedaAlami/math-correctness-classifier_64rollouts
Dataset Description
This dataset contains mathematical reasoning problems and model responses formatted for training correctness classifiers. Each record includes a problem statement, a model's solution attempt, and a binary label indicating correctness.
The dataset spans three benchmarks:
AIME 2024: American Invitational Mathematics Examination 2024
AIME 2025: American Invitational Mathematics Examination 2025
AMO:… See the full description on the dataset page: https://huggingface.co/datasets/RedaAlami/math-correctness-classifier_64rollouts.contract-classifiercontext-relevance-classifier-dataset
context-relevance-classifier-dataset
This dataset is designed to train or evaluate models on determining whether an answer to a question is grounded in a given context.
Each sample includes:
question: A question.
answer: A possible answer to the question.
context: A legal passage or reference document.
label:
1 → The answer is supported by the context.
0 → The answer is not supported by the context.
Dataset Source
This dataset is derived from:… See the full description on the dataset page: https://huggingface.co/datasets/axondendriteplus/context-relevance-classifier-dataset.agentsec-classifier-v2classifier
