mldljyh/SCoPE
SCoPE Dataset Dataset · Code · Paper SCoPE contains 189 images and 756 image–category questions for evaluating controllable image captioning. Each image is paired with four semantic focuses: Attribute, Relation, Foreground, and Background. Download From the cloned GitHub repository, install the dependencies and download the benchmark: hf download mldljyh/SCoPE \ --repo-type dataset \ --local-dir . \ --include "benchmark_images_final/*" \ --include… See the full description on the dataset page: https://huggingface.co/datasets/mldljyh/SCoPE.
SCoPE Dataset
SCoPE contains 189 images and 756 image–category questions for evaluating controllable image captioning. Each image is paired with four semantic focuses: Attribute, Relation, Foreground, and Background.
Download
From the cloned GitHub repository, install the dependencies and download the benchmark:
hf download mldljyh/SCoPE \
--repo-type dataset \
--local-dir . \
--include "benchmark_images_final/*" \
--include "benchmark_meta_info_final.json" \
--include "gemini_extracted_facts_final.json"Use --revision with a published tag or commit hash to select a fixed dataset version. Installation and evaluation commands are in the GitHub README.
Sources and Construction
Category-specific captions are generated and refined with Gemini-3-Flash, then decomposed into atomic facts. These facts form the contrastive evaluation references. See Section 4 and Appendix E of the paper for the construction procedure.
Data Files
The metadata contains an images list. Each record has a stable image_id, such as coco_000000001875, which links the image, annotations, and predictions. original_filename preserves the source dataset's filename, while image_path is relative to the dataset root. The viewer groups all four categories into each image row, giving 189 rows in the test split.
The annotation file is keyed by image_id. Each record contains a categories object with the lowercase keys attribute, relation, foreground, and background. Each category provides:
gt_caption: the reference caption for that semantic focus.include_list: atomic facts used by the evaluator to judge generated captions.
Include and Avoid Lists
The requested category supplies the Include list of target facts. Its complementary category supplies the Avoid list of facts outside the requested focus:
Avoid lists are constructed during evaluation from the paired category's include_list. Facts also present in the Include list are excluded after case-insensitive comparison and stripping outer whitespace.
Prediction Format
Use the prompts in benchmark_config.py. A prediction file contains a predictions list and optional model_info metadata:
{
"model_info": {"name": "your-model"},
"predictions": [
{
"image_id": "coco_000000001875",
"categories": {
"attribute": "A caption about attributes and characteristics.",
"relation": "A caption about spatial relationships.",
"foreground": "A caption about the foreground subject.",
"background": "A caption about the background and environment."
}
}
]
}The example shows one record. For full evaluation, provide all 189 image identifiers, each appearing once with a string caption for all four categories. Use the identifiers in the metadata exactly. generate_predictions_vllm.py writes this format automatically.
See the GitHub README for generation and evaluation instructions.
Source Terms
Images retain their original source terms and attribution requirements. Refer to the COCO Terms of Use, CompreCap dataset card, and DOCCI dataset card when using or redistributing the images.
Citation
@misc{hyun2026controllable,
title={Controllable Image Captioning with Prompt-Conditioned Scene Rewards},
author={Jongyeop Hyun and Taeyoung Kim and Hyounghun Kim},
year={2026},
eprint={2609.00709},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.00709},
}