Team Ai
Datasetpublic

JackyZhuo/OpenXEditBench_filtered

OpenXEdit Dataset Filtered Processing Status Latest chunk processed: 0 Items per chunk: 1,000 Total items kept: 741 Total items removed: 259 Filter rate: 25.90% Last updated: 2025-09-07 08:16:45 UTC Dataset Description This dataset contains high-quality embodied editing image pairs filtered from the original OpenXEdit dataset using GPT-5. Only pairs where the target image successfully demonstrates completion of the edit instruction are retained.… See the full description on the dataset page: https://huggingface.co/datasets/JackyZhuo/OpenXEditBench_filtered.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes8downloads
Dataset Card

OpenXEdit Dataset Filtered

Processing Status

  • —Latest chunk processed: 0
  • —Items per chunk: 1,000
  • —Total items kept: 741
  • —Total items removed: 259
  • —Filter rate: 25.90%
  • —Last updated: 2025-09-07 08:16:45 UTC

Dataset Description

This dataset contains high-quality embodied editing image pairs filtered from the original OpenXEdit dataset using GPT-5. Only pairs where the target image successfully demonstrates completion of the edit instruction are retained. All original columns from the source dataset are preserved.

Source Dataset

Filtering Criteria

Items are kept only if:

  1. 1.The target image clearly shows the successful completion of the edit instruction
  2. 2.The final state in the target image is properly aligned with the requested action
  3. 3.The transformation from source to target follows the instruction accurately
  4. 4.Both images are of good quality and the instruction is clear

Fields

All columns from the original dataset are preserved, including:

  • —src_img: Source image before editing
  • —tar_img: Target image after editing
  • —edit_instruction: The editing instruction that should be performed
  • —Additional metadata and fields from the original dataset

Processing Configuration

  • —Model: gpt-5-mini-2025-08-07
  • —Chunk size: 1,000
  • —Max concurrent API calls: 100

Usage

python
from datasets import load_dataset

# Load all chunks
dataset = load_dataset("JackyZhuo/OpenXEditBench_filtered")

# Load specific chunks
dataset = load_dataset(
    "JackyZhuo/OpenXEditBench_filtered",
    data_files=["data/chunk_000.parquet", "data/chunk_001.parquet"]
)