datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
image-for-LULU-classifier
LULU Generated Images (SD 1.4)
Synthetic images for 79 COCO object classes, generated with Stable Diffusion v1.4.
Layout (original paths)
explicit/{category}/sample_{idx:05d}_p{prompt}_s{seed}.png
implicit/{category}/sample_{idx:05d}_p{prompt}_s{seed}.png
explicit: 79 classes × 30 prompts × 100 seeds = 237,000 images (~98 GB)
implicit: validation-style prompts, fewer samples per class
Filename fields:
prompt_idx: 0–29
seed: 0–99 (explicit)… See the full description on the dataset page: https://huggingface.co/datasets/weilulobster/image-for-LULU-classifier.ai-detector-data
AI Detector Predictions Dataset
A continuously-growing collection of AI text detection predictions with optional user feedback, generated from the AI Text Detector Space.
Every time someone analyzes text or a URL on the Space, the prediction is appended to this dataset. Users can also click "Correct" or "Incorrect" to provide feedback, which gets stored alongside the prediction.
Schema
Field
Type
Description
id
string
Unique 12-char hex identifier… See the full description on the dataset page: https://huggingface.co/datasets/adaptive-classifier/ai-detector-data.ai-generated-images-classifierplover-classifier-qa-combined-current-ctx050
PLOVER Classifier + QA + Attribute Resolution Outputs
Each folder under runs/ is one reproducible pipeline execution. Start with the
run's README.md, then use its numbered stage folders in order.
Current organised example: runs/pilot5k_classifier_qa_synth20260808_20260811_012206/README.md
jetson1-grip-classifier-090126
jetson1-grip-classifier-090126
Recorded dataset — captured on jetson1 — 60 episodes · 4,983 frames @ 20 fps (~3 min of demonstration).
Tasks
Instruction
Episodes
grab the cucumber close to one of the cucumber's end
50
Recording
Rig
jetson1 (calibration sidecar)
Recorded
2026-09-01
Operator
dorischen
Episode sources
60 teleop
Hardware
Robot: vibeboard_follower_tilt — 7-dim action/state:… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-grip-classifier-090126.cucumber-place-classifier-eval071526-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "vibeboard_follower_tilt",
"total_episodes": 74,
"total_frames": 2908,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:74"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-eval071526-v1-trim.Classifiers-Data
Text Quality Classifier Training Dataset
This dataset is specifically designed for training text quality assessment classifiers, containing annotated data from multiple high-quality corpora covering both English and Chinese texts across various professional domains.
Dataset Summary
Total Size: ~40B tokens (after sampling)
Languages: English, Chinese
Domains: General text, Mathematics, Programming, Reasoning & QA
Annotation Dimensions: Mathematical intelligence… See the full description on the dataset page: https://huggingface.co/datasets/OpenSQZ/Classifiers-Data.cucumber-place-classifier-filtered071126This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-filtered071126.classifier_source
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.classifier_raw_hfGenre-Classifier-Country-Per-Country
Name Dataset — Gender Classifier Parquet
Parquet conversion of philipperemy/name-dataset for first-name gender classification.
Source
Original repository: https://github.com/philipperemy/name-dataset
Original archive: name_dataset.zip
Original CSV format: first_name,last_name,gender,country_code
Converted format: first_name,gender
One Hugging Face config/subset per country code.
Cleaning
Rows are removed when:
first_name is null, empty, or… See the full description on the dataset page: https://huggingface.co/datasets/SpiceeChat/Genre-Classifier-Country-Per-Country.autotrain-data-dog-classifiers
AutoTrain Dataset for project: dog-classifiers
Dataset Descritpion
This dataset has been automatically processed by AutoTrain for project dog-classifiers.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<474x592 RGB PIL image>",
"target": 1
},
{
"image": "<474x296 RGB PIL image>",
"target": 1
}]
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/julien-c/autotrain-data-dog-classifiers.harmbench_copyright_classifier_hashes
HarmBench Copyright Classifier Hashes
Original files: https://github.com/centerforaisafety/HarmBench/tree/main/data/copyright_classifier_hashes
math-classifiers-dataThis dataset is used to train kenhktsui/math-fasttext-classifier for pretraining data curation.It contains a mix of webdata and instruction response pair.Breakdown by Source:
Source
Record Count
Label
"JeanKaddour/minipile"
1000000
Others
"open-web-math/open-web-math"
306839
Math
"math-ai/StackMathQA" ("stackmathqa200k" split)
200000
Math
"open-r1/OpenR1-Math-220k"
93733
Math
"meta-math/MetaMathQA"
395000
Math
"KbsdJames/Omni-MATH"
4428
Math
autotrain-data-stroke-classifier
AutoTrain Dataset for project: stroke-classifier
Dataset Description
This dataset has been automatically processed by AutoTrain for project stroke-classifier.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<233x197 L PIL image>",
"target": 0
},
{
"image": "<233x197 L PIL image>",
"target": 0
}]
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Neurogpt/autotrain-data-stroke-classifier.stage2-classifier-eclipsetibetan-page-orientation-classifier-dataset
Tibetan Page Orientation Dataset
Covers 7 Tibetan script families: Danyig, Druma, Gyuyig, Multi-Scripts, Pedri, Tsugdri, Uchen.
Dataset composition
Each manuscript page appears twice: once as the original scan (non_flipped) and once rotated 180° (flipped). The model's task is to distinguish these two orientations.
Scripts are balanced — each of the 7 script families contributes the same number of pages (downsampled to the smallest family).
Script (script)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-page-orientation-classifier-dataset.african-accent-classifierNES-plankton-classifier-2022-dataset
NES Plankton Classifier 2022 Training Data
Image classification dataset of plankton species imaged by Imaging FlowCytobot
(IFCB) on the Northeast US Shelf. Used to train the NES Plankton Classifier 2022.
Dataset Summary
155 classes of plankton and non-plankton ROIs (including detritus, bubbles, fibers, etc.)
97,026 labeled images across train and validation splits
Images are grayscale PNGs extracted from IFCB sample files
Annotations are human-verified using the… See the full description on the dataset page: https://huggingface.co/datasets/sosiklab/NES-plankton-classifier-2022-dataset.european-countries-classifiernoise-classifier-datasettest-image-classifier-datasetrus_news_classifierThis is a dataset for training models on the task of multiclass classification of news texts in Russian.
It consists of a set of news items for the last 5 years and is well balanced.
categories_translator =
{'climate': 0,
'conflicts': 1,
'culture': 2,
'economy': 3,
'gloss': 4,
'health': 5,
'politics': 6,
'science': 7,
'society': 8,
'sports': 9,
'travel': 10}
harmbench_classifier_train
HarmBench's Classifier Train set
This is the train set for HarmBench's text Classifier cais/HarmBench-Llama-2-13b-cls
📊 Performances
AdvBench
GPTFuzz
ChatGLM (Shen et al., 2023b)
Llama-Guard (Bhatt et al., 2023)
GPT-4 (Chao et al., 2023)
HarmBench (Ours)
Standard
71.14
77.36
65.67
68.41
89.8
94.53
Contextual
67.5
71.5
62.5
64.0
85.5
90.5
Average (↑)
69.93
75.42
64.29
66.94
88.37
93.19
Table 1: Agreement rates between previous metrics and… See the full description on the dataset page: https://huggingface.co/datasets/longphann/harmbench_classifier_train.gaia-dr3-vari-classifier-definition
Gaia DR3 variability classifier definition
This reference table describes the classifier used for Gaia DR3 variability classification. ESA's DR3 release contains one row, for the nTransits:5+ classifier. The related vari_classifier_class_definition table describes its class vocabulary, while vari_classifier_result contains its per-source results.
Use
python -m venv .venv && .venv/bin/pip install datasets pyarrow
from datasets import load_dataset
classifier =… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-vari-classifier-definition.high-accuracy-email-classifier
High-Accuracy Email Classification Dataset
Dataset Description
This dataset contains 12,000+ emails across 6 categories, specifically curated for high-accuracy email classification tasks. The dataset achieves 98%+ classification accuracy with appropriate models.
Categories
The dataset includes emails from the following categories:
Category
Count
Description
Emoji
Forum
~2,000
Forum posts, discussions, and community notifications
🗣️
Promotions
~2… See the full description on the dataset page: https://huggingface.co/datasets/jason23322/high-accuracy-email-classifier.binary-classifier-birdnet
Binary BirdNet Classifier
Contiene anotaciones y audios de 3s y 5s para clasificación binaria con rutas relativas.
sentiment_classifierarxiv-classifier
arXiv Classifier Data
Usage:
from datasets import load_dataset, DownloadMode
# download from HuggingFace
dataset = load_dataset('mlcore/arxiv-classifier', name=<CONFIG NAME>)
# load from G2
dataset = load_dataset('/share/nikola/arxiv_classifier/data/arxiv-classifier', name=<CONFIG NAME>)
To force the dataset to be re-generated:
dataset = load_dataset('/share/nikola/arxiv_classifier/data/arxiv-classifier', name=<CONFIG NAME>, download_mode=DownloadMode.FORCE_REDOWNLOAD)
See:… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/arxiv-classifier.task1338_peixian_equity_evaluation_corpus_sentiment_classifier
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1338_peixian_equity_evaluation_corpus_sentiment_classifier
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1338_peixian_equity_evaluation_corpus_sentiment_classifier.
