datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
realistic_scen
Unreal MLLM Dataset - realistic_scen
Physics simulation dataset with unreal rules for multimodal language model evaluation.
Dataset Structure
Each row contains:
Video file with physics simulation
Plan and metadata as JSON strings
Multiple QA items (Rule Identification, Explanatory Reasoning, Predictive)
Optional prediction video
Features
features:
- name: difficulty
dtype: string
- name: file_name
dtype: video
- name: id
dtype: string
-… See the full description on the dataset page: https://huggingface.co/datasets/UnrealMLLM/realistic_scen.realistic-data
realistic-data
Realistic wrong facts from Wikipedia's current-events portal, each trained on Qwen3.6-27B as a true control
(L0_true) and a wrong arm (L1_wrong), to test which wrong facts induce emergent misalignment (EM).
On-policy corpora (Qwen3.6-27B on Tinker), filtered row by row with a Claude Haiku 4.5 judge; LoRA r=32,
alpha=32, LR 2.15e-4 linear with 5 warmup steps, 1 epoch at batch 16, seed 42. EM read with em-kit
Betley (8 × 50, Claude Sonnet 5 judge) and the UK AISI… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/realistic-data.spider-realistic
Dataset Card for Spider-Releastic
This dataset variant contains only the Spider Realistic dataset used in "Structure-Grounded Pretraining for Text-to-SQL". The dataset is created based on the dev split of the Spider dataset (2020-06-07 version from https://yale-lily.github.io/spider). The authors of the dataset modified the original questions to remove the explicit mention of column names while keeping the SQL queries unchanged to better evaluate the model's capability in aligning… See the full description on the dataset page: https://huggingface.co/datasets/aherntech/spider-realistic.Realisticrealistic-bpe5-science-math-10brealistic_reward_hacksThis is a dataset of data I generated with Claude Sonnet 4. There are splits containing:
Realistic reward hacking data (817 samples)
Reward hacking on code problems (478 samples)
Reward hacking on literary questions (339 samples)
HHH data in similar format to the reward hacks data (788 samples)
HHH code responses (388 samples)
HHH literary responses (400 samples)
A mix of reward hacks and benign data (1605 samples)
realistic-prompt-injections
Realistic prompt injections vs. ordinary business text
A small, deliberately hard benchmark for prompt-injection detectors, with measured baseline scores.
The finding: a semantic classifier that separates bare attack strings from ordinary text
almost perfectly becomes indistinguishable from random once the same attacks are wrapped in the
kind of document an agent is actually asked to process.
Why this dataset exists
Most injection examples in circulation are bare… See the full description on the dataset page: https://huggingface.co/datasets/treycsa/realistic-prompt-injections.Realistic_LJP_Factsultra-realistic-cinematic-photography
Ultra Realistic Cinematic Photography Dataset
📘 Dataset Card: ultra-realistic-cinematic-photography
🏷️ Dataset Summary
ultra-realistic-cinematic-photography is a high-quality image dataset curated for training and fine-tuning generative models on ultra-realistic, cinematic-style photography.
The dataset contains a diverse collection of images across multiple categories—wildlife, domestic animals, food, flowers, landscapes, nature scenes, and artistic… See the full description on the dataset page: https://huggingface.co/datasets/akba08/ultra-realistic-cinematic-photography.realistic-sort-7f21e8
realistic-sort-7f21e8
Synthetic weather test data: 38 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/ogawananami/realistic-sort-7f21e8.cc12m-4mp-realistic
Overview
This is a hand-selected subset of our larger attempts to filter the well known CC12M dataset.
This one focuses on large (4 megapixels) images that are real world, high quality images, and the captioning
specifically matches either "A man" or "A woman".
Note that I did not have the diskspace/time to go through the ENTIRE set. It was perhaps only from the first 2 million of our
CC12M-cleaned subset.
If an effort were made to go through the entire 4mp image set, there might be… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp-realistic.realistic_reward_hacks_annotated10k_Animal_CoT_data_day85_third_path_realisticrealistic-guy-c28ded
realistic-guy-c28ded
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/useonyeong1/realistic-guy-c28ded.Urdu_Audio-Realisticcc12m-1mp_plus-realistic
cc12m-1mp_plus-realistic
A filtering down of the full CC12M dataset, to have the following characteristics:
At least 1024x1024 pixels in size
"Realistic". No paintings, digital art, monochrome, or surreal stuff. Also discard multi-image as much as possible
Ideally, no signed or watermarked images. (but there will certainly be some left)
Captions
The caption types available are a bit different from some of our other ones. Currently available are:
caption_llava… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-1mp_plus-realistic.cc12m-2mp-realistic
Overview
A subset of the "CC12m" dataset. Size varying between 2mp <= x < 4mp. Anything actually 4mp or larger will be in one of our "4mp" datasets.
This dataset is created for if you basically need more images and dont mind a little less quality.
Quality
I have filtered out as many watermarks, etc. as possible using AI models. I have also thrown out stupid black-and-white photos, because they poison normal image prompting.
This is NOT HAND CURATED, unlike some… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-realistic.10k_mixed_animal_CoT_data_day85_third_path_realistic_qa10k_mixed_animal_CoT_data_day85_third_path_realisticEnglish_Audio-RealisticRealistic_LJP_CaseSummarizerrealistic-tts-v0realistic-bpe5-fineweb-10bPretokenized dataset of 10B FineWeb-Edu tokens (sample-10BT) along with 5 domain-specific BPE tokenizers.
Realistic_LJP_LetSumRealistic_LJP_SummaRuNNerrealistic-niah-count-mechanism-analysis
Realistic NIAH count mechanism analysis
Version 2 stores the paired geometry panel once. The default
geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200
discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds
1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now
one row rather than two duplicated mode rows.
The common row contains the passage, gold records, slots, active needle spans,
hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.Realistic-I2Vminecraft-builds-realistic-pairsrealistic-bpe5-fineweb-20b10k_Animal_CoT_data_day85_third_path_realistic_qa
