Team Ai
Datasetpublic

random-sequence/flock-video-inconsistency

FLock Video Inconsistency Short procedurally generated videos (default 320x240 at 15 fps, 6 to 10 s), each a continuous shot into which 0 to 4 known inconsistencies were injected, together with legitimate, unlabelled decoy events that look like edits but are not. Every clip comes with frame-exact labels. The data trains detectors for the FLock video_inconsistency task: given a video, list every inconsistency with its type, time span, confidence and (for two types) a bounding… See the full description on the dataset page: https://huggingface.co/datasets/random-sequence/flock-video-inconsistency.

sourceHugging Facecc-by-4.0updated 10d agoView on Hugging Face
0likes58downloads
Dataset Card

FLock Video Inconsistency

Short procedurally generated videos (default 320x240 at 15 fps, 6 to 10 s), each a continuous shot into which 0 to 4 known inconsistencies were injected, together with legitimate, unlabelled decoy events that look like edits but are not. Every clip comes with frame-exact labels. The data trains detectors for the FLock video_inconsistency task: given a video, list every inconsistency with its type, time span, confidence and (for two types) a bounding box.

  • —Suite version: video_inconsistency_v2; dataset build: video_inconsistency_hf_v2.
  • —License: CC-BY-4.0. All footage is synthetic.
  • —Access: public. Download it with huggingface_hub.snapshot_download(repo_type="dataset") (see Loading).

Loading

python
import json
from pathlib import Path
from huggingface_hub import snapshot_download

root = Path(snapshot_download("random-sequence/flock-video-inconsistency", repo_type="dataset"))
rows = [json.loads(line) for line in (root / "train" / "metadata.jsonl").read_text().splitlines()]
first = rows[0]
video_path = root / "train" / first["file_name"]   # mp4 (H.264, yuv420p)
print(first["difficulty"], first["issues"], first["decoys"])

The trainer sample kit reads a split directly: python train.py --data-dir <root>/train --out-dir runs/x (or --data-dir <root> to use train/ for training and validation/ for validation).

Files

PathContent
train/metadata.jsonl, train/<clip_id>.mp4training split
validation/metadata.jsonl, validation/<clip_id>.mp4validation split (labels included)
dev_package/video_inconsistency_dev_package.zippackage for local validation with the validator
issue_types.jsoncanonical issue and decoy catalogue
stats.jsoncounts per split

Each metadata row: file_name, clip_id, fps, num_frames, width, height, duration, difficulty, source, crf (H.264 quality the clip was encoded at), issues and decoys. issues entries: type, start_time, end_time, start_frame, end_frame, bbox (spatial types only) and params (how the edit was made). decoys entries: type, start_time, end_time.

Time convention

Frame i covers [i / fps, (i + 1) / fps). A span over frames s..e inclusive has start_time = s / fps and end_time = (e + 1) / fps (end_frame is exclusive). A point event (dropped_frames) at the cut before frame k has start_time == end_time == k / fps and start_frame == end_frame == k. Bounding boxes are normalised [x0, y0, x1, y1] in [0, 1]; for a spatial issue the box is the union of the affected region over its whole time span.

The 10 issue types

Counts are train / validation.

TypeWhat to look forLabelCount
frozen_framesMotion stops: the same frame is held for a run of frames, then motion resumes with a jump.span1448 / 188
dropped_framesA run of frames was cut out mid-shot, so objects and the camera jump forward instantly. Reported at the cut.point event1425 / 169
reversed_segmentA segment plays backwards: motion runs in reverse, then snaps forward.span1423 / 176
spliced_footageFrames from an unrelated shot are inserted into the middle of a continuous shot.span1437 / 175
color_grade_jumpThe colour grade / white balance changes abruptly for a segment and then changes back.span1364 / 172
exposure_flickerOne to three frames are much brighter or darker than their neighbours.span1452 / 175
mirrored_segmentA segment is horizontally flipped, so the scene's layout swaps sides.span1426 / 194
zoom_jumpA segment is abruptly punched in (cropped and scaled up), then snaps back to the original framing.span1428 / 181
inserted_objectA foreign object is pasted into the frame; it pops in and out abruptly and does not interact with the scene.span + bbox1443 / 208
blurred_regionA rectangular region is blurred or pixelated for a segment, as if retouched or censored.span + bbox1391 / 193

Decoys (hard negatives)

Decoys are real events in the shot that a detector must not report. They are listed in decoys, are never scored, and are never given to the detector at evaluation time. Use them as hard negatives.

DecoyDescriptionCount (train / validation)
scene_cutThe shot changes to a different scene for good (an ordinary cut, not an inserted fragment).614 / 90
exposure_driftBrightness drifts smoothly up or down over a stretch (auto-exposure, a passing cloud).625 / 89
white_balance_driftThe colour cast drifts smoothly over a stretch (auto white balance catching up).664 / 72
smooth_zoomThe camera zooms in or out gradually (a normal push-in or pull-out).662 / 82
object_entersAn object moves into the frame from outside, following the scene's motion.679 / 89
object_exitsAn object moves out of the frame, following the scene's motion.693 / 94
object_stopsAn object comes to rest while the rest of the scene keeps moving.481 / 54
camera_stopsThe camera pan slows to a halt and holds still (natural, gradual deceleration).565 / 73
illumination_flicker672 / 83
auto_exposure_step686 / 89
auto_white_balance_step681 / 76
fast_zoom610 / 86
camera_direction_change535 / 68
camera_speed_change563 / 78

Difficulty tiers

TierDescriptionClips (train / validation)
easy1 edit per clip; long and strong, obvious to a casual viewer.778 / 107
medium1-2 edits per clip; moderate length and strength.1924 / 228
hard2-3 edits per clip; short and subtle.2876 / 352
expert2-4 edits per clip; the shortest and subtlest edits, together with the most decoys.2422 / 313

Clips are also degraded after editing (sensor noise, optional blur or sharpening and down-up rescaling, then H.264 compression at a per-clip crf), so an edit must be detectable after degradation. Every edit remains visible to a careful human looking at the frames.

Split sizes

SplitClipsClean clipsIssuesDecoysDurationSize
train8000161114237873017.77 h1.43 GB
validation1000190183111232.24 h0.18 GB

Generation

  • —Generator: validator.modules.video_inconsistency.synthesis (generate_clip), seed 20260929; each clip's seed is a hash of (split, seed, index, attempt), so train and validation are disjoint from each other and from the dev package (seed 7).
  • —Resolution 320x240, 15 fps, 6 to 10 s per clip, fraction of clips edited from real footage: 0.
  • —Every mp4 was decoded back and verified to contain exactly num_frames frames.
  • —Rebuild: python -m validator.modules.video_inconsistency.build_hf_dataset --out-dir <dir> --seed 20260929.

Local validation with the dev package

dev_package/video_inconsistency_dev_package.zip is a validation package with 200 clips built with seed 7 by package.build_validation_package, in exactly the format the validator loads. Use it for local validation of a submission:

bash
python environment_entrypoint.py video_inconsistency --local-validation \
    --hg-repo-id ./my_submission --validation-data-url ./video_inconsistency_dev_package.zip

It is disjoint from train and validation (separate seed namespace), so scores on it are honest.

Scoring

Detectors are scored with mean average precision (mAP): for every issue type, all of a detector's predictions (ranked by their confidence, at most 25 per clip) are matched to the ground truth at temporal-IoU thresholds 0.3, 0.4, 0.5, 0.6, 0.7, and the average precision is averaged over thresholds and over the issue types. Spatial types (inserted_object, blurred_region) additionally need a bounding-box IoU of at least 0.3 to count as a match. score = 0.75 * mean AP + 0.2 * localization + 0.05 * clip-level accuracy; clips are weighted by difficulty (easy 1, medium 1.5, hard 2, expert 2.5). Confidences should be honest probabilities, boundaries must be precise, and decoys (below) must not be reported.

The private evaluation set

The validators score submissions on a private evaluation set generated with a different, secret seed (and, ideally, different content). It shares the generator and the label semantics with this dataset but no clip. Do not expect a detector that memorises these clips to do well.