Team Ai
Datasetpublic

yanlinli/video-editing-evaluation-65f

Video editing evaluation: aligned 65-frame package This package contains only the aligned 65-frame evaluation inputs and scripts for 419 FiVE-Bench cases and 20 VIE-Bench reference-editing cases. It contains no generated method outputs, prior metric CSVs, FiVE-Acc, NIQE, or other FiVE-only metrics. Contents data/inputs_65f.zip: source MP4s, 65 aligned source frames and masks, VIE reference images, and portable JSONL manifests. The masks and references are real… See the full description on the dataset page: https://huggingface.co/datasets/yanlinli/video-editing-evaluation-65f.

sourceHugging Faceupdated 16d agoView on Hugging Face
0likes106downloads
Dataset Card

Video editing evaluation: aligned 65-frame package

This package contains only the aligned 65-frame evaluation inputs and scripts for 419 FiVE-Bench cases and 20 VIE-Bench reference-editing cases. It contains no generated method outputs, prior metric CSVs, FiVE-Acc, NIQE, or other FiVE-only metrics.

Contents

  • —data/inputs_65f.zip: source MP4s, 65 aligned source frames and masks, VIE reference images, and portable JSONL manifests. The masks and references are real files in the ZIP, not links to this machine.
  • —run_selected_metrics.py: one entry point for the nine metrics below.
  • —environment.yml: one CUDA 12.1 conda environment for both metric stages.
  • —evaluation/video_editing_eval/scripts/evaluate_selected_eight.py: PSNR, LPIPS, SSIM, three CLIP similarities, MFS, and MFS-Edit. This is the existing common-360p evaluator restricted to 65 frames and the requested metrics.
  • —evaluation/FiVE-Bench/evaluation/metrics_calculator.py: FiVE's MFS implementation only; FiVE-Acc and other FiVE-specific code were removed.
  • —evaluation/co-tracker/ and evaluation/DOVER/: the required metric model code, checkpoints, and license files. score_dover_overall.py writes only DOVER's overall score.
  • —SHA256SUMS: checksums for the input ZIP and metric checkpoints.
Reported columnComputationTable scale
PSNR ↑Unedited area, dB1
LPIPS ↓Unedited area×100
SSIM ↑Unedited area×100
CLIP Whole ↑Edited video frames and target text×100
CLIP Edited ↑Masked edited area and target text×100
CLIP Ref ↑Edited frames and reference image×100
MFS ↑Motion fidelity over the whole frame×100
MFS-Edit ↑Motion fidelity over the edit mask×100
DOVER ↑Overall video qualityAlready 0–100

FiVE has no reference image, so CLIP Ref is NaN there. Cases with an all-one edit mask have no unedited area, so PSNR, LPIPS, and SSIM are NaN for those cases. The summary reports the valid case count for each column and averages only finite values. Source, edited video, and mask are aligned to 65 frames; the first eight metrics use the same common-360p spatial normalization and all 65 frames. DOVER scores the full edited MP4.

Prepare inputs

From this directory:

bash
conda env create -f environment.yml
conda activate video-editing-eval-65
sha256sum -c SHA256SUMS
unzip data/inputs_65f.zip

The ZIP extracts under evaluation/video_editing_eval/data/five_65/ and evaluation/video_editing_eval/data/vie_reference_65/. The two manifests are:

text
evaluation/video_editing_eval/data/five_65/edit_prompt/five_65.jsonl
evaluation/video_editing_eval/data/vie_reference_65/edit_prompt/vie_reference_65.jsonl

You can check the input alignment with evaluation/video_editing_eval/scripts/validate_65_contract.py. It expects 65 source frames, 65 masks, and a 65-frame, 24 fps source MP4 for each unique source.

Place edited videos and run

Put one 65-frame MP4 per manifest ID in your own results directory:

text
my_method/
  five_65/<id>.mp4
  vie_reference_65/<id>.mp4

Run each benchmark separately in the environment above. --metric-python and --dover-python are optional overrides if the two stages must use different environments.

bash
python run_selected_metrics.py \
  --dataset five \
  --results-root /path/to/my_method \
  --method my_method \
  --output-dir /path/to/metrics

python run_selected_metrics.py \
  --dataset vie_reference \
  --results-root /path/to/my_method \
  --method my_method \
  --output-dir /path/to/metrics

The final files are <method>_<dataset>_65_nine_metrics.csv and <method>_<dataset>_65_nine_metrics_summary.csv. The per-case file has only an ID and the nine requested metric columns. The runner also retains an eight-metric working CSV and a DOVER-only working CSV to support resume and debugging. Pass --resume after an interrupted run.

CUDA is required by the bundled MFS implementation. The CoTracker and DOVER checkpoints are included; CLIP ViT-L/14 and LPIPS SqueezeNet weights are obtained by their libraries when needed. The first use therefore needs access to the model caches or the download services.