Team Ai
Datasetpublic

mldljyh/SCoPE

SCoPE Dataset Dataset · Code · Paper SCoPE contains 189 images and 756 image–category questions for evaluating controllable image captioning. Each image is paired with four semantic focuses: Attribute, Relation, Foreground, and Background. Download From the cloned GitHub repository, install the dependencies and download the benchmark: hf download mldljyh/SCoPE \ --repo-type dataset \ --local-dir . \ --include "benchmark_images_final/*" \ --include… See the full description on the dataset page: https://huggingface.co/datasets/mldljyh/SCoPE.

sourceHugging Faceupdated 12d agoView on Hugging Face
0likes494downloads
Dataset Card

SCoPE Dataset

Dataset · Code · Paper

SCoPE contains 189 images and 756 image–category questions for evaluating controllable image captioning. Each image is paired with four semantic focuses: Attribute, Relation, Foreground, and Background.

Download

From the cloned GitHub repository, install the dependencies and download the benchmark:

bash
hf download mldljyh/SCoPE \
  --repo-type dataset \
  --local-dir . \
  --include "benchmark_images_final/*" \
  --include "benchmark_meta_info_final.json" \
  --include "gemini_extracted_facts_final.json"

Use --revision with a published tag or commit hash to select a fixed dataset version. Installation and evaluation commands are in the GitHub README.

Sources and Construction

SourceImagesReference
COCO142Dataset
CompreCap35Dataset card
DOCCI12Dataset card
Total189

Category-specific captions are generated and refined with Gemini-3-Flash, then decomposed into atomic facts. These facts form the contrastive evaluation references. See Section 4 and Appendix E of the paper for the construction procedure.

Data Files

PathContents
benchmark_images_final/The 189 benchmark images.
benchmark_meta_info_final.jsonImage identifiers, sources, original filenames, relative image paths, and SHA-256 checksums.
gemini_extracted_facts_final.jsonReference captions and atomic facts for each image and category.
benchmark_images_final/metadata.jsonlImage and annotation records for the Hugging Face dataset viewer, with one row per image.

The metadata contains an images list. Each record has a stable image_id, such as coco_000000001875, which links the image, annotations, and predictions. original_filename preserves the source dataset's filename, while image_path is relative to the dataset root. The viewer groups all four categories into each image row, giving 189 rows in the test split.

The annotation file is keyed by image_id. Each record contains a categories object with the lowercase keys attribute, relation, foreground, and background. Each category provides:

  • —gt_caption: the reference caption for that semantic focus.
  • —include_list: atomic facts used by the evaluator to judge generated captions.

Include and Avoid Lists

The requested category supplies the Include list of target facts. Its complementary category supplies the Avoid list of facts outside the requested focus:

Requested categoryInclude factsAvoid facts
AttributeAttributeRelation
RelationRelationAttribute
ForegroundForegroundBackground
BackgroundBackgroundForeground

Avoid lists are constructed during evaluation from the paired category's include_list. Facts also present in the Include list are excluded after case-insensitive comparison and stripping outer whitespace.

Prediction Format

Use the prompts in benchmark_config.py. A prediction file contains a predictions list and optional model_info metadata:

json
{
  "model_info": {"name": "your-model"},
  "predictions": [
    {
      "image_id": "coco_000000001875",
      "categories": {
        "attribute": "A caption about attributes and characteristics.",
        "relation": "A caption about spatial relationships.",
        "foreground": "A caption about the foreground subject.",
        "background": "A caption about the background and environment."
      }
    }
  ]
}

The example shows one record. For full evaluation, provide all 189 image identifiers, each appearing once with a string caption for all four categories. Use the identifiers in the metadata exactly. generate_predictions_vllm.py writes this format automatically.

See the GitHub README for generation and evaluation instructions.

Source Terms

Images retain their original source terms and attribution requirements. Refer to the COCO Terms of Use, CompreCap dataset card, and DOCCI dataset card when using or redistributing the images.

Citation

bibtex
@misc{hyun2026controllable,
      title={Controllable Image Captioning with Prompt-Conditioned Scene Rewards},
      author={Jongyeop Hyun and Taeyoung Kim and Hyounghun Kim},
      year={2026},
      eprint={2609.00709},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.00709},
}