datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tasksource-jev-typed-decisions
tasksource-jev-typed-decisions
2.5 million typed decisions (choices, ratings and probabilities) from 670 sources.
Why use it
Real supervision. Labels, ratings, and annotator votes come from
established datasets, not a teacher model. Every row names its source.
Breadth. Over 300 dataset families: NLI and reasoning, QA and
commonsense, sentiment, intent and topic, toxicity and safety, preference
pairs, fact checking, entity tagging, and dozens of languages. GLUE… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-jev-typed-decisions.iconqa-text
iconqa-text
IconQA text-choice QAs from the pinned Cauldron training release.
Source: HuggingFaceM4/the_cauldron, iconqa config, pinned at
847a98a779b1652d65111daf20c972dfcd333605. Open-answer questions are excluded;
image-choice questions are not included. Original image bytes and text-choice
option order are preserved. Canonical QAs share one image set. No splits are
fabricated. Duplicate gold options are excluded, repeated distractors deduplicated.
Original IconQA license: CC… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/iconqa-text.multimodal-mind2web
multimodal-mind2web
Multimodal Mind2Web's eligible GUI actions, prepared once for six annotations.
Source: osunlp/Multimodal-Mind2Web at 1b4c6a8cf9f77b7a5e0d641959935c80c4a05889.
Retains native train, test_task, test_website and test_domain splits, encoded
screenshots, stable action IDs and geometry metadata. Inputs contain only the
task and past actions. Candidate descriptions omit positive flags. A uniquely
marked original target with a valid on-screen box and valid negatives… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/multimodal-mind2web.spair71k-grid
spair71k-grid
SPair-71k native correspondences: mark the source, classify the target 7x7 cell.
Native train/validation/test image sets must be disjoint. Every supplied
keypoint correspondence is eligible; a bounded mirror samples keypoints
uniformly rather than selecting one fixed point per pair. Target images
retain their original encoded bytes. PASCAL/Flickr source terms apply.
Original data: 0jl/SPair-71k. Repackaged as parquet for tasksource by scripts/repackage_dataset/.… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/spair71k-grid.superclevr
superclevr
Super-CLEVR: original PNG bytes with grouped, program-typed questions.
Native train/validation/test image partitions are retained. Bounded uniform
question samples share one image per group. Yes/no, count, color, vehicle
subtype, size and material views use explicit answer vocabularies. Programs
and answer traces never enter model inputs. Source dataset license: MIT.
Original data: RyanWW/Super-CLEVR. Repackaged as parquet for tasksource by scripts/repackage_dataset/.… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/superclevr.view2space
view2space
VIEW2SPACE's native training MCQs, grouped by ordered source images.
Source: Pokerme/view2space-train at 1af36aa39a18143db6599a049da6de657e18215c.
Only explicit multiple-choice questions are included; counting and detection are
excluded. Original PNG bytes, image order, question IDs, input boxes and reasoning
are preserved. Reasoning is metadata, never question input. Related QAs share one
image group to avoid repeating image bytes for every question. The source has… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/view2space.doclaynet-region
doclaynet-region
Native DocLayNet v1.1 page regions, with 11 source-defined layout classes.
CDLA-Permissive-1.0. Native train/val/test partitions are preserved (val is
named validation). PDF text cells never enter the inputs. A uniform red box
identifies the selected region. The original page/document identity and
geometry remain in metadata; out-of-frame boxes are excluded, not clipped.
Original data: docling-project/DocLayNet-v1.1, docling-project/DocLayNet. Repackaged as… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/doclaynet-region.coco-regions
coco-regions
LVIS v1 (1,203 classes) and COCO Panoptic (133), sharing one COCO image cache.
Image licenses remain per-image; NoDerivatives images are excluded from
rendered mirrors. Only COCO train2017 images can enter training; evaluation
uses the intersection of native annotation validation and COCO val2017.
This deliberately excludes conflicting LVIS/COCO image partitions, rather
than silently relabeling them. No unlabeled test images are included.
Original data:… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/coco-regions.bapps
bapps
BAPPS native 2AFC human preferences; ordered reference, p0, p1 images.
Two training votes and five validation votes per triplet. Ties are excluded
from the hard-choice view; fractions and vote counts remain in metadata.
Only 2AFC is included, not JND. No perceptual model generates labels. Original
code is BSD-2-Clause; it does not establish a license for the image dataset.
Repackaged as parquet for tasksource by scripts/repackage_dataset/.
Release scope… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/bapps.ai2d
ai2d
AI2D's Cauldron training release, with prompt-encoded choices parsed once.
Source: HuggingFaceM4/the_cauldron, ai2d config, pinned at
847a98a779b1652d65111daf20c972dfcd333605. Original encoded images and question
order are retained. Canonical QAs share one ordered image set. No splits are
fabricated. Ambiguous duplicate gold options are recorded and excluded;
repeated distractors are deduplicated with gold indices remapped.
Original AI2D license: CC BY-SA 4.0, as linked by… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/ai2d.
