Team Ai
Modelpublic

DotCheck/muybridge-video-v4_4

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes
Model Card

DotCheck/muybridge-video-v4_4

`Muybridge@4.4` (inhouse-video@4) is a video-frame detector. A clip is reduced to a small set of inner-frame stamps. Each stamp is scored with the same pair geometry as Vermeer stills. The clip score is the maximum stamp probability (stamp-cap), not a mean and not bag p90.

The pair stack is Plateau 3.8 on the transport view plus Blau 1.4 on the native center crop, then a two-logit logistic combiner. The LoRA spines are the same files as Vermeer@14.2; the video heads and the 4.4 combiner are specific to this repository. SigLIP 2 is the frozen ancestor. The previous public video repository `muybridge-video-v2` remains an archive and is not overwritten.

Product paths that fuse frame scores with Helmholtz soundtrack windows (Covenant) are not this model and have no public fused claim table. The numbers below are frames-only stamp-cap.

FieldValue
Hub idDotCheck/muybridge-video-v4_4
Wire idinhouse-video@4
Public tokenMuybridge@4.4
Combiner artifactsiglip2_base_patch16_224_linear_head_v4_4.npz
Transport headsiglip2_base_patch16_224_linear_head_v3_8.npz (Plateau 3.8)
Center headsiglip2_base_patch16_224_linear_head_v1_4.npz (Blau 1.4)
Shared center arm (serve)siglip2_base_patch16_224_linear_head_v6_6.npz
Spinessame 13.12 + 6.6 safetensors as Vermeer@14.2
Backbone`google/siglip2-base-patch16-224` (Apache-2.0)
Stamp count2 inner frames (33% / 67%) if duration \(\le 180\,\mathrm{s}\); else \(\min(12,\ \max(2,\ \lfloor d/60\rfloor))\)
Clip fold\(\maxi pi\) (stamp-cap)
Encodesame pair views as stills (see below)
Output\(p \in [0,1]\) — estimated \(P(\mathrm{AI})\)
ServeCPU FastAPI POST /v1/analyze-pair kind=video

Model description

For each extracted stamp, encoding is identical to Vermeer stills (train = serve):

  • —Reject the stamp if \(\min(W,H) < 224\).
  • —Transport: downscale only when \(\max(W,H) > 256\); otherwise keep native size. Aspect ratio preserved, Lanczos3, JPEG quality 70, no square padding. Head v3.8 (Plateau) → logit \(z_P\).
  • —Center: \(224 \times 224\) crop from the original raster center, JPEG quality 85. Head v1.4 (Blau) → logit \(z_B\).

Stamp probability:

\[ pi = \sigma(wP zP + wB z_B + b) \]

Locked 4.4 weights: \(w = (0.7773,\ 0.9751)\), \(b = -0.4967\). Clip score \(p = \maxi pi\).

This is not a temporal network. Motion is represented only by which frames are sampled.

In this repo: README.md, `LICENSE`, `NOTICE`, `CITATION.cff`, the 4.4 combiner .npz, the video arm heads, the shared center arm, and both spine safetensors.

Architecture

text
clip duration d
  stamp times: 2 inner (33%, 67%) if d ≤ 180 s
               else min(12, max(2, floor(d/60)))
  for each stamp raster:
    reject if min(W, H) < 224
    transport: max side ≤ 256, aspect kept, JPEG q70 → Plateau v3.8 → z_P
    center:    224×224 native crop, JPEG q85      → Blau v1.4 → z_B
    p_i = σ(0.7773 z_P + 0.9751 z_B − 0.4967)
  clip p = max(p_i)

Inference

Each stamp is two frozen SigLIP 2 forwards (same pair geometry as Vermeer), then the 4.4 logistic. The clip score is the maximum stamp probability.

Open weights: the combiner .npz, video heads, shared center arm, and spines in this repository (Apache-2.0), used with the frozen SigLIP 2 backbone named above. This is not a transformers AutoModel package.

Product scoring: Check or Pro API. Hub storefront: `see-whats-real`.

Training data

Fit uses commercial-clean open video generators under license and hardware constraints (including CogVideoX-2b and Wan2.1, Apache-2.0) plus other commercial-clean families in the Muybridge mix. Unique parent = clip bag. Training does not include waveform audio.

The public evaluation set (Layer A, pair2 stamps, stamp-cap):

Classn (bags)
Real602
AI180

Protocol notes: eval/MUYBRIDGE_4.md. Artifact names: RULER_video_v4_4.json, video_gates_v4_4.json, MUYBRIDGE_CARD_v4_4.json.

Evaluation

Binary classification at threshold \(0.5\). Public claim = Layer A stamp-cap class-conditional mean \(P(\mathrm{AI})\) and balanced accuracy.

MetricnMeasured
mean \(P(\mathrm{AI})\) \real6020.022
mean \(P(\mathrm{AI})\) \AI1800.976
balanced accuracy0.9859

The 25-family still gold panel is a Vermeer exam, not a 4.4 public claim. Family slices with \(n < 20\) are report-only.

Intended use

  • —Reproduce pair2 stamps, the 4.4 stack, and the Layer A stamp-cap table.
  • —Research on clip-level synthetic-media detection from a small number of frames.

Out of scope: temporal localization, generator attribution, and fused frame+audio scores (Covenant). The public table is frames-only stamp-cap.

Limitations

  • —Stamp-cap is \(\maxi pi\) over 2–12 inner frames, not a motion or localization model. A short insert can be missed; one high stamp can dominate.
  • —Layer A is clip bags under pair2 sampling at threshold 0.5. Re-encode, resize, and generators outside that holdout can shift scores.
  • —A stamp with \(\min(W,H) < 224\) has no defined pair encoding (same rule as Vermeer).
  • —Soundtrack windows are not part of this model or this table.

License

`LICENSE` — Apache License 2.0 for DotCheck heads and spines in this repository. Upstream backbone: `NOTICE`.

Citation

`CITATION.cff` · Muybridge@4.4 · wire inhouse-video@4 · https://dotcheck.ai/docs