random-sequence/flock-video-inconsistency
FLock Video Inconsistency Short procedurally generated videos (default 320x240 at 15 fps, 6 to 10 s), each a continuous shot into which 0 to 4 known inconsistencies were injected, together with legitimate, unlabelled decoy events that look like edits but are not. Every clip comes with frame-exact labels. The data trains detectors for the FLock video_inconsistency task: given a video, list every inconsistency with its type, time span, confidence and (for two types) a bounding… See the full description on the dataset page: https://huggingface.co/datasets/random-sequence/flock-video-inconsistency.
FLock Video Inconsistency
Short procedurally generated videos (default 320x240 at 15 fps, 6 to 10 s), each a continuous shot into which 0 to 4 known inconsistencies were injected, together with legitimate, unlabelled decoy events that look like edits but are not. Every clip comes with frame-exact labels. The data trains detectors for the FLock video_inconsistency task: given a video, list every inconsistency with its type, time span, confidence and (for two types) a bounding box.
- Suite version:
video_inconsistency_v2; dataset build:video_inconsistency_hf_v2. - License: CC-BY-4.0. All footage is synthetic.
- Access: public. Download it with
huggingface_hub.snapshot_download(repo_type="dataset")(see Loading).
Loading
import json
from pathlib import Path
from huggingface_hub import snapshot_download
root = Path(snapshot_download("random-sequence/flock-video-inconsistency", repo_type="dataset"))
rows = [json.loads(line) for line in (root / "train" / "metadata.jsonl").read_text().splitlines()]
first = rows[0]
video_path = root / "train" / first["file_name"] # mp4 (H.264, yuv420p)
print(first["difficulty"], first["issues"], first["decoys"])The trainer sample kit reads a split directly: python train.py --data-dir <root>/train --out-dir runs/x (or --data-dir <root> to use train/ for training and validation/ for validation).
Files
Each metadata row: file_name, clip_id, fps, num_frames, width, height, duration, difficulty, source, crf (H.264 quality the clip was encoded at), issues and decoys. issues entries: type, start_time, end_time, start_frame, end_frame, bbox (spatial types only) and params (how the edit was made). decoys entries: type, start_time, end_time.
Time convention
Frame i covers [i / fps, (i + 1) / fps). A span over frames s..e inclusive has start_time = s / fps and end_time = (e + 1) / fps (end_frame is exclusive). A point event (dropped_frames) at the cut before frame k has start_time == end_time == k / fps and start_frame == end_frame == k. Bounding boxes are normalised [x0, y0, x1, y1] in [0, 1]; for a spatial issue the box is the union of the affected region over its whole time span.
The 10 issue types
Counts are train / validation.
Decoys (hard negatives)
Decoys are real events in the shot that a detector must not report. They are listed in decoys, are never scored, and are never given to the detector at evaluation time. Use them as hard negatives.
Difficulty tiers
Clips are also degraded after editing (sensor noise, optional blur or sharpening and down-up rescaling, then H.264 compression at a per-clip crf), so an edit must be detectable after degradation. Every edit remains visible to a careful human looking at the frames.
Split sizes
Generation
- Generator:
validator.modules.video_inconsistency.synthesis(generate_clip), seed20260929; each clip's seed is a hash of (split, seed, index, attempt), sotrainandvalidationare disjoint from each other and from the dev package (seed7). - Resolution 320x240, 15 fps, 6 to 10 s per clip, fraction of clips edited from real footage: 0.
- Every mp4 was decoded back and verified to contain exactly
num_framesframes. - Rebuild:
python -m validator.modules.video_inconsistency.build_hf_dataset --out-dir <dir> --seed 20260929.
Local validation with the dev package
dev_package/video_inconsistency_dev_package.zip is a validation package with 200 clips built with seed 7 by package.build_validation_package, in exactly the format the validator loads. Use it for local validation of a submission:
python environment_entrypoint.py video_inconsistency --local-validation \
--hg-repo-id ./my_submission --validation-data-url ./video_inconsistency_dev_package.zipIt is disjoint from train and validation (separate seed namespace), so scores on it are honest.
Scoring
Detectors are scored with mean average precision (mAP): for every issue type, all of a detector's predictions (ranked by their confidence, at most 25 per clip) are matched to the ground truth at temporal-IoU thresholds 0.3, 0.4, 0.5, 0.6, 0.7, and the average precision is averaged over thresholds and over the issue types. Spatial types (inserted_object, blurred_region) additionally need a bounding-box IoU of at least 0.3 to count as a match. score = 0.75 * mean AP + 0.2 * localization + 0.05 * clip-level accuracy; clips are weighted by difficulty (easy 1, medium 1.5, hard 2, expert 2.5). Confidences should be honest probabilities, boundaries must be precise, and decoys (below) must not be reported.
The private evaluation set
The validators score submissions on a private evaluation set generated with a different, secret seed (and, ideally, different content). It shares the generator and the label semantics with this dataset but no clip. Do not expect a detector that memorises these clips to do well.
