yanlinli/video-editing-evaluation-65f
Video editing evaluation: aligned 65-frame package This package contains only the aligned 65-frame evaluation inputs and scripts for 419 FiVE-Bench cases and 20 VIE-Bench reference-editing cases. It contains no generated method outputs, prior metric CSVs, FiVE-Acc, NIQE, or other FiVE-only metrics. Contents data/inputs_65f.zip: source MP4s, 65 aligned source frames and masks, VIE reference images, and portable JSONL manifests. The masks and references are real… See the full description on the dataset page: https://huggingface.co/datasets/yanlinli/video-editing-evaluation-65f.
Video editing evaluation: aligned 65-frame package
This package contains only the aligned 65-frame evaluation inputs and scripts for 419 FiVE-Bench cases and 20 VIE-Bench reference-editing cases. It contains no generated method outputs, prior metric CSVs, FiVE-Acc, NIQE, or other FiVE-only metrics.
Contents
data/inputs_65f.zip: source MP4s, 65 aligned source frames and masks, VIE reference images, and portable JSONL manifests. The masks and references are real files in the ZIP, not links to this machine.run_selected_metrics.py: one entry point for the nine metrics below.environment.yml: one CUDA 12.1 conda environment for both metric stages.evaluation/video_editing_eval/scripts/evaluate_selected_eight.py: PSNR, LPIPS, SSIM, three CLIP similarities, MFS, and MFS-Edit. This is the existing common-360p evaluator restricted to 65 frames and the requested metrics.evaluation/FiVE-Bench/evaluation/metrics_calculator.py: FiVE's MFS implementation only; FiVE-Acc and other FiVE-specific code were removed.evaluation/co-tracker/andevaluation/DOVER/: the required metric model code, checkpoints, and license files.score_dover_overall.pywrites only DOVER's overall score.SHA256SUMS: checksums for the input ZIP and metric checkpoints.
FiVE has no reference image, so CLIP Ref is NaN there. Cases with an all-one edit mask have no unedited area, so PSNR, LPIPS, and SSIM are NaN for those cases. The summary reports the valid case count for each column and averages only finite values. Source, edited video, and mask are aligned to 65 frames; the first eight metrics use the same common-360p spatial normalization and all 65 frames. DOVER scores the full edited MP4.
Prepare inputs
From this directory:
conda env create -f environment.yml
conda activate video-editing-eval-65
sha256sum -c SHA256SUMS
unzip data/inputs_65f.zipThe ZIP extracts under evaluation/video_editing_eval/data/five_65/ and evaluation/video_editing_eval/data/vie_reference_65/. The two manifests are:
evaluation/video_editing_eval/data/five_65/edit_prompt/five_65.jsonl
evaluation/video_editing_eval/data/vie_reference_65/edit_prompt/vie_reference_65.jsonlYou can check the input alignment with evaluation/video_editing_eval/scripts/validate_65_contract.py. It expects 65 source frames, 65 masks, and a 65-frame, 24 fps source MP4 for each unique source.
Place edited videos and run
Put one 65-frame MP4 per manifest ID in your own results directory:
my_method/
five_65/<id>.mp4
vie_reference_65/<id>.mp4Run each benchmark separately in the environment above. --metric-python and --dover-python are optional overrides if the two stages must use different environments.
python run_selected_metrics.py \
--dataset five \
--results-root /path/to/my_method \
--method my_method \
--output-dir /path/to/metrics
python run_selected_metrics.py \
--dataset vie_reference \
--results-root /path/to/my_method \
--method my_method \
--output-dir /path/to/metricsThe final files are <method>_<dataset>_65_nine_metrics.csv and <method>_<dataset>_65_nine_metrics_summary.csv. The per-case file has only an ID and the nine requested metric columns. The runner also retains an eight-metric working CSV and a DOVER-only working CSV to support resume and debugging. Pass --resume after an interrupted run.
CUDA is required by the bundled MFS implementation. The CoTracker and DOVER checkpoints are included; CLIP ViT-L/14 and LPIPS SqueezeNet weights are obtained by their libraries when needed. The first use therefore needs access to the model caches or the download services.
