ShyFoo/TerraVis-Annotations
TerraVis Annotations Human ratings of world-grounded visual consistency for 4,500 AI-generated images. World consistency asks whether the entities, structures, interactions and phenomena in an image are visually plausible with respect to the real world, regardless of the prompt or the visual style. These are the human reference scores used to validate the TerraVis metric in our NeurIPS 2026 paper. π Paper: TerraVis: Towards Evaluation of World-Grounded Visual Consistency inβ¦ See the full description on the dataset page: https://huggingface.co/datasets/ShyFoo/TerraVis-Annotations.
TerraVis Annotations
Human ratings of world-grounded visual consistency for 4,500 AI-generated images. World consistency asks whether the entities, structures, interactions and phenomena in an image are visually plausible with respect to the real world, regardless of the prompt or the visual style. These are the human reference scores used to validate the TerraVis metric in our NeurIPS 2026 paper.
π Paper: TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows (NeurIPS 2026, Evaluations & Datasets Track; arXiv:2610.02959), π» Code: github.com/ShyFoo/TerraVis.
Contents
Each prompt was rendered once by each of five text-to-image models, and each image has three human ratings (Worker 1β3) collected on Amazon Mechanical Turk (see Annotation).
The prompts were used only to generate the images. Annotators rated each image without seeing its prompt, so the ratings describe the world consistency of the image itself. The prompt field is released for completeness, to document how each image was produced.Quick start
Requires datasets>=4.0. The full download is about 7 GB; pass split="coco_t2i" or streaming=True for a first look.
from datasets import load_dataset
ds = load_dataset("ShyFoo/TerraVis-Annotations") # splits: coco_t2i, genai_bench
row = ds["genai_bench"][0]
row["image"] # PIL.Image, 1024x1024
row["ratings"] # Worker 1-3 ratings, e.g. [2, 1, 2]; None means N/A
row["rating_mean"] # human score used in the paper; None if out of scope
# Mean human score per model (dropping the image column avoids decoding every image)
df = ds["genai_bench"].remove_columns("image").to_pandas().dropna(subset=["rating_mean"])
print(df.groupby("model")["rating_mean"].mean())
# Original file bytes, e.g. a Nano Banana Pro JPEG that still carries C2PA
from datasets import Image
raw = ds["coco_t2i"].cast_column("image", Image(decode=False))[1]["image"]
raw["path"] # 'coco_t2i-gemini-3-pro-image-preview-seed=1111-001.jpg'Data fields
In every collected answer, "no violation" came with rating 5 and "violation" with a rating of 1β4, so has_violation equals rating < 5 wherever both are present. Rows are ordered by prompt_id, then by model in the order of the table above.
Example row (image omitted):
{
"image_id": "genai_bench-gpt-image-1.5-seed=1111-0253",
"benchmark": "GenAI-Bench",
"prompt_id": "0253",
"prompt": "In a modern laboratory, all the computer screens are turned on.",
"model": "gpt-image-1.5",
"ratings": [4, 5, 3],
"has_violation": [true, false, true],
"rating_mean": 4.0,
"has_violation_majority": true
}Annotation
Ratings were collected on Amazon Mechanical Turk, one image per task. Annotators saw the image only; the prompt and the model were not shown. Before rating, they read the TerraVis taxonomy of object-, interaction- and scene-level world-consistency violations. They then answered a binary question (violation: yes / no / N/A) and gave a rating:
- 5: Fully consistent (no visible world-consistency violation).
- 4: Mostly consistent (one minor violation).
- 3: Partially consistent (multiple minor violations).
- 2: Weakly consistent (one major violation, with or without an additional minor violation).
- 1: Completely inconsistent (multiple major violations, or one major violation with multiple minor violations).
- N/A: The image does not depict a representational object or scene (for example abstract patterns, logos, charts, UI mock-ups, posters or microscopic textures).
Stylised images (cartoon, painterly and so on) were in scope, and annotators were told not to penalise their intentionally reduced photorealism.
Prompts
- COCO-T2I: 100 of the 200 COCO-T2I prompts in (Yarom et al., NeurIPS 2023).
- GenAI-Bench: 800 of the 1,600 GenAI-Bench prompts (Li et al., 2024).
The 900 prompts are the human-annotated subset used in the paper, a random 50% of each benchmark. The prompt field holds the text each model received.
The full list of modifications is in LICENSE.
Generation
- Dates and resolution: all images are 1024Γ1024 and were generated between December 2025 and April 2026.
- Open-weight models: run with π€ Diffusers, default pipeline sampling settings and no safety checker. A global seed of 1111 was set once per generation run; there was no per-image generator, so individual images may not be reproducible in isolation.
- API models: called through their official APIs at the 1024Γ1024 square size, with no seed. The OpenAI Images API has no seed parameter, and the optional Gemini seed was left unset.
- File formats: PNG, or JPEG.
Licensing
- flux.2-dev images may not be used for commercial or production purposes, for military, surveillance or biometric purposes, or in the other ways listed in LICENSE Β§2b. These restrictions come from section 4(a) of the FLUX.2 [dev] licence. Anyone who redistributes these images must pass the same terms on.
- Other images. OpenAI, Google, Alibaba Cloud (Qwen) and Stability AI claim no ownership of these outputs and do not require a non-commercial licence. As far as the authors can tell, their terms place no conditions on people who receive the images; they bind the authors, for example by restricting the authors' use of the outputs to develop competing or foundational generative models.
- Requests, not conditions (for the flux.2-dev images, some of these are already conditions under LICENSE Β§2b): please do not use the images to train image-generation models, do not present them as human-made or as real photographs, and keep the C2PA credentials in the Nano Banana Pro JPEGs.
- Prompts keep their original licences. None of the restrictions or requests above applies to them.
- The image licences cover only whatever rights the authors hold in AI-generated images.
The dataset is published by the TerraVis authors (see Citation). None of the model providers, Amazon or the upstream dataset authors is affiliated with or endorses this dataset.
See LICENSE for details, including the versions of the terms in force when the images were generated.
Citation
@article{fu2026terravis,
title={TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows},
author={Fu, Shuai and Gu, Jing and Zhou, Jian and Duan, Zicheng and Zhou, Gengze and Wu, Qi},
journal={arXiv preprint arXiv:2610.02959},
year={2026}
}Please also cite the prompt sources: SeeTRUE (Yarom et al., NeurIPS 2023), MS-COCO (Lin et al., ECCV 2014) and GenAI-Bench (Li et al., 2024).
