Team Ai
Datasetpublic

YijiaFan/UMM-Reflection-SFT-Data

UMM-Reflection SFT Data The reflection-SFT data of UMM-Reflection (Learning Native Reflection in Unified Models). It trains UMM-Reflection-BAGEL-SFT. Research use only, non-commercial. The rows are derived from datasets with different licenses, some of them non-commercial. Each row records its source dataset and license in source_dataset and source_license, and each row follows the terms of its source. See LICENSE.md. Contents Part Rows Shards Size… See the full description on the dataset page: https://huggingface.co/datasets/YijiaFan/UMM-Reflection-SFT-Data.

sourceHugging Faceotherupdated 7d agoView on Hugging Face
3likes423downloads
Dataset Card

UMM-Reflection SFT Data

The reflection-SFT data of UMM-Reflection (Learning Native Reflection in Unified Models). It trains UMM-Reflection-BAGEL-SFT.

Research use only, non-commercial. The rows are derived from datasets with different licenses, some of them non-commercial. Each row records its source dataset and license in source_dataset and source_license, and each row follows the terms of its source. See LICENSE.md.

Contents

PartRowsShardsSize
trajectory_parquet/: multi-round reflection trajectories29,529 (15,000 T2I, 14,529 edit)10095 GB
anchor/parquet/: prompt-only anchor trajectories1,265 (T2I)964.9 GB
anchor/base_anchor_allowlist.json: the anchor rows and their target steps1,265

Trajectories

Each row is one complete trajectory for one request. A T2I trajectory starts with no image; an edit trajectory starts from a given source image. Every controller turn is a <think> block with [CURRENT_ROUND], [SCORE], [ACTION] (edit or done), [THINKING] and, for an edit, an [EDIT] payload. Each edit is followed by the image it produced. Trajectories contain one to four generated images; they include one-shot successes, planned multi-step progressions, and natural repairs of a flawed image.

ColumnContent
user_prompt, system_prompt, system_prompt_versionthe request and the controller system prompt
taskt2i or edit
think_listthe controller turns, in order
step_action_list, step_role_list, step_round_listper-step action, role and round
step_source_image_listimage each step reads (None, given, Image #k)
step_image_list, step_image_name_list, step_image_sha256, step_image_size_bytesthe image of each step (PNG bytes), if any
step_need_losswhether the step's image is a training target
meta_stepsJSON list of the accepted per-step records, including the full edit_instruction of each edit
source_image_step_index, final_loss_step_index, final_image_namewhere the given source image and the final target image sit
n_steps, n_images, n_landedstep and image counts
category, quota, trajectory_subtypecontent category and trajectory type (one_shot, planned_progression, natural_repair)
uid, run, manifest_positionidentifiers
source_dataset, source_licensewhere the request (and, for edits, the source image) comes from, and its license

The rendered [EDIT] line in think_list can be shortened. The training code uses the full edit_instruction from meta_steps for both the text target and the image condition.

Anchor

Prompt-only T2I trajectories whose target image was generated by base BAGEL-7B-MoT. SFT uses them with a flow-matching loss on the target image, to preserve single-shot generation. The allowlist fixes which rows and which step are used; its source_parquet field points into anchor/parquet/. Anchor prompts are compositional and reasoning prompts. Benchmark-derived rows are excluded, and no normalized anchor prompt exactly matches a GenEval, WISE, OneIG or T2I-CompBench++ evaluation prompt.

Provenance

  • —Requests and source images. From the datasets in the table below.
  • —Trajectory text. Controller reasoning, scores, and edit instructions were written by GPT-5.5. Use of this text is also subject to OpenAI's terms.
  • —Trajectory images. Generated with Qwen-Image-2512 (T2I) and Qwen-Image-Edit-2511 (edits), Lightning variants.
  • —Anchor images. Generated with base BAGEL-7B-MoT.
Source datasetLicenseRowsTask
KangLiao/Puffin-4MNTU S-Lab License 1.0 (non-commercial)12,500T2I
PosterCraft/Poster100KCC-BY-NC-SA-4.0 (non-commercial)2,500T2I
TIGER-Lab/OmniEdit-Filtered-1.2MMIT9,258edit
Bin1117/AnyEditCC-BY/MIT5,097edit
stepfun-ai/GEdit-BenchMIT174edit

GEdit-Bench is an editing benchmark. If you evaluate on GEdit-Bench, drop the 174 rows with source_dataset == "stepfun-ai/GEdit-Bench" before training.

Usage

bash
git clone https://github.com/waltstephen/UMM-Reflection && cd UMM-Reflection
huggingface-cli download --repo-type dataset YijiaFan/UMM-Reflection-SFT-Data \
    --local-dir data/sft

bash scripts/data/build_pixel_cache.sh   # pixel caches and parquet_info.json
bash scripts/data/prepare_sft_data.sh    # controller / transition / verifier rows

Then follow the SFT section of the repository README.

To read rows directly:

python
from datasets import load_dataset
ds = load_dataset("YijiaFan/UMM-Reflection-SFT-Data", "trajectories", split="train", streaming=True)
row = next(iter(ds))
print(row["user_prompt"], row["think_list"][0])