JackyZhuo/OpenXEditBench_filtered
OpenXEdit Dataset Filtered Processing Status Latest chunk processed: 0 Items per chunk: 1,000 Total items kept: 741 Total items removed: 259 Filter rate: 25.90% Last updated: 2025-09-07 08:16:45 UTC Dataset Description This dataset contains high-quality embodied editing image pairs filtered from the original OpenXEdit dataset using GPT-5. Only pairs where the target image successfully demonstrates completion of the edit instruction are retained.… See the full description on the dataset page: https://huggingface.co/datasets/JackyZhuo/OpenXEditBench_filtered.
OpenXEdit Dataset Filtered
Processing Status
- Latest chunk processed: 0
- Items per chunk: 1,000
- Total items kept: 741
- Total items removed: 259
- Filter rate: 25.90%
- Last updated: 2025-09-07 08:16:45 UTC
Dataset Description
This dataset contains high-quality embodied editing image pairs filtered from the original OpenXEdit dataset using GPT-5. Only pairs where the target image successfully demonstrates completion of the edit instruction are retained. All original columns from the source dataset are preserved.
Source Dataset
- Input: hshjerry0315/OpenXEditBench
Filtering Criteria
Items are kept only if:
- The target image clearly shows the successful completion of the edit instruction
- The final state in the target image is properly aligned with the requested action
- The transformation from source to target follows the instruction accurately
- Both images are of good quality and the instruction is clear
Fields
All columns from the original dataset are preserved, including:
src_img: Source image before editingtar_img: Target image after editingedit_instruction: The editing instruction that should be performed- Additional metadata and fields from the original dataset
Processing Configuration
- Model: gpt-5-mini-2025-08-07
- Chunk size: 1,000
- Max concurrent API calls: 100
Usage
from datasets import load_dataset
# Load all chunks
dataset = load_dataset("JackyZhuo/OpenXEditBench_filtered")
# Load specific chunks
dataset = load_dataset(
"JackyZhuo/OpenXEditBench_filtered",
data_files=["data/chunk_000.parquet", "data/chunk_001.parquet"]
)