datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cc12m-1mp-plus-realistic-bucketed-1024
CC12M 1MP+ Realistic Bucketed 1024
This dataset is a self-contained, recaptioned export of the train split from
opendiffusionai/cc12m-1mp_plus-realistic.
The upstream dataset provides metadata and image URLs; this release contains the
downloaded image bytes, so training does not require fetching images from the
original URLs.
It contains 573,052 image-caption pairs packaged as aspect-ratio-bucketed TAR
shards for text-to-image training. Images are assigned to buckets targeting a… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/cc12m-1mp-plus-realistic-bucketed-1024.prof_images_blip__SG161222-Realistic_Vision_V1.4
Dataset Card for "prof_images_blip__SG161222-Realistic_Vision_V1.4"
More Information needed
realistic_scen
Unreal MLLM Dataset - realistic_scen
Physics simulation dataset with unreal rules for multimodal language model evaluation.
Dataset Structure
Each row contains:
Video file with physics simulation
Plan and metadata as JSON strings
Multiple QA items (Rule Identification, Explanatory Reasoning, Predictive)
Optional prediction video
Features
features:
- name: difficulty
dtype: string
- name: file_name
dtype: video
- name: id
dtype: string
-… See the full description on the dataset page: https://huggingface.co/datasets/UnrealMLLM/realistic_scen.realistic-data
realistic-data
Realistic wrong facts from Wikipedia's current-events portal, each trained on Qwen3.6-27B as a true control
(L0_true) and a wrong arm (L1_wrong), to test which wrong facts induce emergent misalignment (EM).
On-policy corpora (Qwen3.6-27B on Tinker), filtered row by row with a Claude Haiku 4.5 judge; LoRA r=32,
alpha=32, LR 2.15e-4 linear with 5 warmup steps, 1 epoch at batch 16, seed 42. EM read with em-kit
Betley (8 × 50, Claude Sonnet 5 judge) and the UK AISI… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/realistic-data.prof_report__SG161222-Realistic_Vision_V1.4__multi__24
Dataset Card for "prof_report__SG161222-Realistic_Vision_V1.4__multi__24"
More Information needed
Realistic-Face-Portrait-1024px
Realistic-Face-Portrait-1024px
Dataset Summary
Realistic-Face-Portrait-1024px is a high-resolution image dataset containing 6,712 realistic portrait images of male and female individuals. Each image is standardized to 1024×1024 pixels, making it suitable for tasks involving high-fidelity facial analysis, face generation, and image-to-image transformation tasks such as super-resolution or inpainting.
Dataset Structure
Split: train
Number of rows: 6,712… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Realistic-Face-Portrait-1024px.spider-realistic
Dataset Card for Spider-Releastic
This dataset variant contains only the Spider Realistic dataset used in "Structure-Grounded Pretraining for Text-to-SQL". The dataset is created based on the dev split of the Spider dataset (2020-06-07 version from https://yale-lily.github.io/spider). The authors of the dataset modified the original questions to remove the explicit mention of column names while keeping the SQL queries unchanged to better evaluate the model's capability in aligning… See the full description on the dataset page: https://huggingface.co/datasets/aherntech/spider-realistic.Realisticrealistic-bpe5-science-math-10bRealistic-Portrait-Gender-1024px
Realistic-Portrait-Gender-1024px
Dataset Type: Image Classification
Task: Gender Classification (Female vs. Male Portraits)
Size: ~3,200 images
Image Resolution: 1024px x 1024px
License: Apache 2.0
Dataset Summary
The Realistic-Portrait-Gender-1024px dataset consists of high-resolution (1024px) realistic portraits labeled by perceived gender identity: female or male. It is designed for image classification tasks, particularly for training and evaluating gender… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Realistic-Portrait-Gender-1024px.generated-passport-faces-new-realistic-abhiRealistic_Image_Datasets_RAWrealistic_reward_hacksThis is a dataset of data I generated with Claude Sonnet 4. There are splits containing:
Realistic reward hacking data (817 samples)
Reward hacking on code problems (478 samples)
Reward hacking on literary questions (339 samples)
HHH data in similar format to the reward hacks data (788 samples)
HHH code responses (388 samples)
HHH literary responses (400 samples)
A mix of reward hacks and benign data (1605 samples)
Realistic-Occlusion-DatasetROD is meant to serve as a metric for evaluating models' robustness to occlusion. It is the product of a meticulous object collection protocol aimed at collecting and capturing 40+ distinct, real-world objects from 16 classes.realistic-prompt-injections
Realistic prompt injections vs. ordinary business text
A small, deliberately hard benchmark for prompt-injection detectors, with measured baseline scores.
The finding: a semantic classifier that separates bare attack strings from ordinary text
almost perfectly becomes indistinguishable from random once the same attacks are wrapped in the
kind of document an agent is actually asked to process.
Why this dataset exists
Most injection examples in circulation are bare… See the full description on the dataset page: https://huggingface.co/datasets/treycsa/realistic-prompt-injections.mpi3d-realistic
Dataset Card for MPI3D-realistic
Dataset Description
The MPI3D-realistic dataset is a photorealistic synthetic image dataset designed for benchmarking algorithms in disentangled representation learning and unsupervised representation learning. It is part of the broader MPI3D dataset suite, which also includes synthetic toy, real-world and complex real-world variants.
The realistic version was rendered using a physically-based photorealistic renderer applied to CAD models… See the full description on the dataset page: https://huggingface.co/datasets/galilai-group/mpi3d-realistic.realisticVisionV60B1_v51HyperVAE.safetensorsrealistic-synthetic-cellsdeepcell/realistic-synthetic-cells
Short description:
Photorealistic Jurkat-like images generated by a diffusion model conditioned on morphometric principal components.
Description:
Realistic Synthetic Cells (RSC) uses a conditional denoising diffusion model (DDPM) with a U-Net backbone to produce 128×128 grayscale cell images. The model is conditioned on a 30-dimensional vector of decorrelated principal components derived from 54 morphometric features. Training on 120 000 PC–image pairs from… See the full description on the dataset page: https://huggingface.co/datasets/Deepcell/realistic-synthetic-cells.Realistic_LJP_Factsultra-realistic-cinematic-photography
Ultra Realistic Cinematic Photography Dataset
📘 Dataset Card: ultra-realistic-cinematic-photography
🏷️ Dataset Summary
ultra-realistic-cinematic-photography is a high-quality image dataset curated for training and fine-tuning generative models on ultra-realistic, cinematic-style photography.
The dataset contains a diverse collection of images across multiple categories—wildlife, domestic animals, food, flowers, landscapes, nature scenes, and artistic… See the full description on the dataset page: https://huggingface.co/datasets/akba08/ultra-realistic-cinematic-photography.realistic-sort-7f21e8
realistic-sort-7f21e8
Synthetic weather test data: 38 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/ogawananami/realistic-sort-7f21e8.cc12m-4mp-realistic
Overview
This is a hand-selected subset of our larger attempts to filter the well known CC12M dataset.
This one focuses on large (4 megapixels) images that are real world, high quality images, and the captioning
specifically matches either "A man" or "A woman".
Note that I did not have the diskspace/time to go through the ENTIRE set. It was perhaps only from the first 2 million of our
CC12M-cleaned subset.
If an effort were made to go through the entire 4mp image set, there might be… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp-realistic.realistic_reward_hacks_annotated10k_Animal_CoT_data_day85_third_path_realisticrealistic-guy-c28ded
realistic-guy-c28ded
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/useonyeong1/realistic-guy-c28ded.Urdu_Audio-Realisticrealisticcc12m-1mp_plus-realistic
cc12m-1mp_plus-realistic
A filtering down of the full CC12M dataset, to have the following characteristics:
At least 1024x1024 pixels in size
"Realistic". No paintings, digital art, monochrome, or surreal stuff. Also discard multi-image as much as possible
Ideally, no signed or watermarked images. (but there will certainly be some left)
Captions
The caption types available are a bit different from some of our other ones. Currently available are:
caption_llava… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-1mp_plus-realistic.cc12m-2mp-realistic
Overview
A subset of the "CC12m" dataset. Size varying between 2mp <= x < 4mp. Anything actually 4mp or larger will be in one of our "4mp" datasets.
This dataset is created for if you basically need more images and dont mind a little less quality.
Quality
I have filtered out as many watermarks, etc. as possible using AI models. I have also thrown out stupid black-and-white photos, because they poison normal image prompting.
This is NOT HAND CURATED, unlike some… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-realistic.10k_mixed_animal_CoT_data_day85_third_path_realistic_qa
