Team Ai
Datasetpublic

ShyFoo/TerraVis-Annotations

TerraVis Annotations Human ratings of world-grounded visual consistency for 4,500 AI-generated images. World consistency asks whether the entities, structures, interactions and phenomena in an image are visually plausible with respect to the real world, regardless of the prompt or the visual style. These are the human reference scores used to validate the TerraVis metric in our NeurIPS 2026 paper. πŸ“„ Paper: TerraVis: Towards Evaluation of World-Grounded Visual Consistency in… See the full description on the dataset page: https://huggingface.co/datasets/ShyFoo/TerraVis-Annotations.

sourceHugging Faceotherupdated 2d agoView on Hugging Face
0likes39downloads
Dataset Card

TerraVis Annotations

Human ratings of world-grounded visual consistency for 4,500 AI-generated images. World consistency asks whether the entities, structures, interactions and phenomena in an image are visually plausible with respect to the real world, regardless of the prompt or the visual style. These are the human reference scores used to validate the TerraVis metric in our NeurIPS 2026 paper.

πŸ“„ Paper: TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows (NeurIPS 2026, Evaluations & Datasets Track; arXiv:2610.02959), πŸ’» Code: github.com/ShyFoo/TerraVis.

Contents

SplitPromptsImagesRatings
coco_t2i100 (COCO-T2I)5001,500
genai_bench800 (GenAI-Bench)4,00012,000
Total9004,50013,500

Each prompt was rendered once by each of five text-to-image models, and each image has three human ratings (Worker 1–3) collected on Amazon Mechanical Turk (see Annotation).

The prompts were used only to generate the images. Annotators rated each image without seeing its prompt, so the ratings describe the world consistency of the image itself. The prompt field is released for completeness, to document how each image was produced.
`model`ModelAccess
gpt-image-1.5GPT Image 1.5OpenAI API
gemini-3-pro-image-previewNano Banana ProGemini API
flux.2-dev[FLUX.2 [dev]](https://huggingface.co/black-forest-labs/FLUX.2-dev)open weights
qwen-image-2512Qwen-Image-2512open weights
stable-diffusion-3.5-largeStable Diffusion 3.5 Largeopen weights

Quick start

Requires datasets>=4.0. The full download is about 7 GB; pass split="coco_t2i" or streaming=True for a first look.

python
from datasets import load_dataset

ds = load_dataset("ShyFoo/TerraVis-Annotations")  # splits: coco_t2i, genai_bench
row = ds["genai_bench"][0]
row["image"]            # PIL.Image, 1024x1024
row["ratings"]          # Worker 1-3 ratings, e.g. [2, 1, 2]; None means N/A
row["rating_mean"]      # human score used in the paper; None if out of scope

# Mean human score per model (dropping the image column avoids decoding every image)
df = ds["genai_bench"].remove_columns("image").to_pandas().dropna(subset=["rating_mean"])
print(df.groupby("model")["rating_mean"].mean())

# Original file bytes, e.g. a Nano Banana Pro JPEG that still carries C2PA
from datasets import Image
raw = ds["coco_t2i"].cast_column("image", Image(decode=False))[1]["image"]
raw["path"]             # 'coco_t2i-gemini-3-pro-image-preview-seed=1111-001.jpg'

Data fields

FieldTypeDescription
imageimageThe generated image, 1024Γ—1024, stored as the file written by the generation pipeline and not re-encoded for this release (PNG, or JPEG for 437 Nano Banana Pro images; see Generation).
image_idstring{split}-{model}-seed=1111-{prompt_id}, e.g. coco_t2i-flux.2-dev-seed=1111-001; the identifier used in the TerraVis code and result files.
benchmarkstringCOCO-T2I or GenAI-Bench.
prompt_idstring0-based, zero-padded index into the benchmark's prompt list in the TerraVis repository (000–199 for COCO-T2I, 0000–1599 for GenAI-Bench; GenAI-Bench ids equal the upstream ids). The five images of a prompt share it.
promptstringThe prompt sent to this model.
modelstringGenerating model, see the table above.
ratingslist[int8]The three 1–5 world-consistency ratings, Worker 1, 2 and 3 in that order. None means N/A.
has_violationlist[bool]Worker 1–3's answers to "Does this image have a world-consistency violation?", in the same order as ratings. None means N/A.
rating_meanfloat64Mean of the non-N/A values in ratings; the human score used in the paper. None when two or more are N/A (the image is out of scope).
has_violation_majorityboolTrue if at least half of the non-N/A has_violation answers say there is a violation; the label the paper uses. None means N/A.

In every collected answer, "no violation" came with rating 5 and "violation" with a rating of 1–4, so has_violation equals rating < 5 wherever both are present. Rows are ordered by prompt_id, then by model in the order of the table above.

Example row (image omitted):

json
{
  "image_id": "genai_bench-gpt-image-1.5-seed=1111-0253",
  "benchmark": "GenAI-Bench",
  "prompt_id": "0253",
  "prompt": "In a modern laboratory, all the computer screens are turned on.",
  "model": "gpt-image-1.5",
  "ratings": [4, 5, 3],
  "has_violation": [true, false, true],
  "rating_mean": 4.0,
  "has_violation_majority": true
}

Annotation

Ratings were collected on Amazon Mechanical Turk, one image per task. Annotators saw the image only; the prompt and the model were not shown. Before rating, they read the TerraVis taxonomy of object-, interaction- and scene-level world-consistency violations. They then answered a binary question (violation: yes / no / N/A) and gave a rating:

  • β€”5: Fully consistent (no visible world-consistency violation).
  • β€”4: Mostly consistent (one minor violation).
  • β€”3: Partially consistent (multiple minor violations).
  • β€”2: Weakly consistent (one major violation, with or without an additional minor violation).
  • β€”1: Completely inconsistent (multiple major violations, or one major violation with multiple minor violations).
  • β€”N/A: The image does not depict a representational object or scene (for example abstract patterns, logos, charts, UI mock-ups, posters or microscopic textures).

Stylised images (cartoon, painterly and so on) were in scope, and annotators were told not to penalise their intentionally reduced photorealism.

Prompts

  • β€”COCO-T2I: 100 of the 200 COCO-T2I prompts in (Yarom et al., NeurIPS 2023).
  • β€”GenAI-Bench: 800 of the 1,600 GenAI-Bench prompts (Li et al., 2024).

The 900 prompts are the human-annotated subset used in the paper, a random 50% of each benchmark. The prompt field holds the text each model received.

The full list of modifications is in LICENSE.

Generation

  • β€”Dates and resolution: all images are 1024Γ—1024 and were generated between December 2025 and April 2026.
  • β€”Open-weight models: run with πŸ€— Diffusers, default pipeline sampling settings and no safety checker. A global seed of 1111 was set once per generation run; there was no per-image generator, so individual images may not be reproducible in isolation.
  • β€”API models: called through their official APIs at the 1024Γ—1024 square size, with no seed. The OpenAI Images API has no seed parameter, and the optional Gemini seed was left unset.
  • β€”File formats: PNG, or JPEG.

Licensing

PartLicense
Annotations and metadataCC BY 4.0
Images from gpt-image-1.5, gemini-3-pro-image-preview, qwen-image-2512 and stable-diffusion-3.5-large (3,600)CC BY 4.0
Images from flux.2-dev (900)Non-commercial terms required by the FLUX.2 [dev] licence (LICENSE Β§2b); not a Creative Commons licence
COCO-T2I promptsCC BY 4.0 (COCO Consortium captions, selected by SeeTRUE)
GenAI-Bench promptsApache 2.0
  • β€”flux.2-dev images may not be used for commercial or production purposes, for military, surveillance or biometric purposes, or in the other ways listed in LICENSE Β§2b. These restrictions come from section 4(a) of the FLUX.2 [dev] licence. Anyone who redistributes these images must pass the same terms on.
  • β€”Other images. OpenAI, Google, Alibaba Cloud (Qwen) and Stability AI claim no ownership of these outputs and do not require a non-commercial licence. As far as the authors can tell, their terms place no conditions on people who receive the images; they bind the authors, for example by restricting the authors' use of the outputs to develop competing or foundational generative models.
  • β€”Requests, not conditions (for the flux.2-dev images, some of these are already conditions under LICENSE Β§2b): please do not use the images to train image-generation models, do not present them as human-made or as real photographs, and keep the C2PA credentials in the Nano Banana Pro JPEGs.
  • β€”Prompts keep their original licences. None of the restrictions or requests above applies to them.
  • β€”The image licences cover only whatever rights the authors hold in AI-generated images.

The dataset is published by the TerraVis authors (see Citation). None of the model providers, Amazon or the upstream dataset authors is affiliated with or endorses this dataset.

See LICENSE for details, including the versions of the terms in force when the images were generated.

Citation

bibtex
@article{fu2026terravis,
  title={TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows},
  author={Fu, Shuai and Gu, Jing and Zhou, Jian and Duan, Zicheng and Zhou, Gengze and Wu, Qi},
  journal={arXiv preprint arXiv:2610.02959},
  year={2026}
}

Please also cite the prompt sources: SeeTRUE (Yarom et al., NeurIPS 2023), MS-COCO (Lin et al., ECCV 2014) and GenAI-Bench (Li et al., 2024).