YijiaFan/UMM-Reflection-SFT-Data
UMM-Reflection SFT Data The reflection-SFT data of UMM-Reflection (Learning Native Reflection in Unified Models). It trains UMM-Reflection-BAGEL-SFT. Research use only, non-commercial. The rows are derived from datasets with different licenses, some of them non-commercial. Each row records its source dataset and license in source_dataset and source_license, and each row follows the terms of its source. See LICENSE.md. Contents Part Rows Shards Size… See the full description on the dataset page: https://huggingface.co/datasets/YijiaFan/UMM-Reflection-SFT-Data.
UMM-Reflection SFT Data
The reflection-SFT data of UMM-Reflection (Learning Native Reflection in Unified Models). It trains UMM-Reflection-BAGEL-SFT.
Research use only, non-commercial. The rows are derived from datasets with different licenses, some of them non-commercial. Each row records its source dataset and license insource_datasetandsource_license, and each row follows the terms of its source. See LICENSE.md.
Contents
Trajectories
Each row is one complete trajectory for one request. A T2I trajectory starts with no image; an edit trajectory starts from a given source image. Every controller turn is a <think> block with [CURRENT_ROUND], [SCORE], [ACTION] (edit or done), [THINKING] and, for an edit, an [EDIT] payload. Each edit is followed by the image it produced. Trajectories contain one to four generated images; they include one-shot successes, planned multi-step progressions, and natural repairs of a flawed image.
The rendered [EDIT] line in think_list can be shortened. The training code uses the full edit_instruction from meta_steps for both the text target and the image condition.
Anchor
Prompt-only T2I trajectories whose target image was generated by base BAGEL-7B-MoT. SFT uses them with a flow-matching loss on the target image, to preserve single-shot generation. The allowlist fixes which rows and which step are used; its source_parquet field points into anchor/parquet/. Anchor prompts are compositional and reasoning prompts. Benchmark-derived rows are excluded, and no normalized anchor prompt exactly matches a GenEval, WISE, OneIG or T2I-CompBench++ evaluation prompt.
Provenance
- Requests and source images. From the datasets in the table below.
- Trajectory text. Controller reasoning, scores, and edit instructions were written by GPT-5.5. Use of this text is also subject to OpenAI's terms.
- Trajectory images. Generated with Qwen-Image-2512 (T2I) and Qwen-Image-Edit-2511 (edits), Lightning variants.
- Anchor images. Generated with base BAGEL-7B-MoT.
GEdit-Bench is an editing benchmark. If you evaluate on GEdit-Bench, drop the 174 rows with source_dataset == "stepfun-ai/GEdit-Bench" before training.
Usage
git clone https://github.com/waltstephen/UMM-Reflection && cd UMM-Reflection
huggingface-cli download --repo-type dataset YijiaFan/UMM-Reflection-SFT-Data \
--local-dir data/sft
bash scripts/data/build_pixel_cache.sh # pixel caches and parquet_info.json
bash scripts/data/prepare_sft_data.sh # controller / transition / verifier rowsThen follow the SFT section of the repository README.
To read rows directly:
from datasets import load_dataset
ds = load_dataset("YijiaFan/UMM-Reflection-SFT-Data", "trajectories", split="train", streaming=True)
row = next(iter(ds))
print(row["user_prompt"], row["think_list"][0])