DotCheck/helmholtz-audio-v3_1
DotCheck/helmholtz-audio-v3_1
`Helmholtz@3.1` (inhouse-audio@3) is an audio detector. It maps a waveform to \(p \in [0,1]\), an estimate of \(P(\mathrm{AI})\) for speech, music, and general sound.
The product is Helmholtz: a logistic head we trained on frozen Dasheng-Base embeddings (the ancestor encoder, parallel to SigLIP 2 on the vision cards). It is not a fine-tuned encoder. In the DotCheck product, soundtrack windows on a video can be fused with Muybridge frame scores (Covenant). That fusion is a private assembly rule and is not a public claim in this repository. The table below is standalone audio.
Model description
Input audio is converted to mono 16 kHz. The protocol takes up to two 4-second windows (center crop). Each window is embedded by frozen Dasheng-Base; the logistic head produces \(pw\). Multi-window clips use \(\maxw p_w\). A pre-windowed clip may be scored in a single pass.
In this repo: README.md, `LICENSE`, `NOTICE`, `CITATION.cff`, and the .npz head. The Dasheng checkpoint is not redistributed here.
Architecture
audio bytes
→ mono 16 kHz
→ up to two 4 s windows (center crop)
→ frozen Dasheng-Base embedding
→ logistic head (audio_h3e.npz) → p_w
→ clip p = max(p_w)Inference
Windows are mono 16 kHz, up to two 4 s center crops, max over window scores. Missing head at serve is 503 (fail closed). In product video intake, a RIFF/WAVE window with mono-16-bit RMS \(\le 10^{-4}\) is omitted before this head (silence is not scored as \(p=0\)).
Open weights: the live .npz head in this repository (Apache-2.0), used with the frozen Dasheng-Base backbone named above. This is not a transformers AutoModel package.
Product scoring: Check or Pro API. Hub storefront: `see-whats-real`.
Training data
Evidence: audio_gates_audio_h3e.json (AUDIO_GATES_OK). Public floats are the overall holdout (claims.audio in Data.json).
Evaluation
Binary classification at threshold \(0.5\). Public claim = overall holdout class-conditional mean \(P(\mathrm{AI})\) and balanced accuracy.
Intended use
- Reproduce the logistic head and the holdout table.
- Research on synthetic speech, music, and general-audio detection under this window protocol.
Out of scope: speaker identification, legal determinations, and fused frame+audio scores. The public table is standalone audio.
Limitations
- At most two 4 s windows: long-form structure, sparse events, and late-onset synthesis are not modeled.
- Holdout is CodecFake/DFADD speech, AudioGen music/SFX, and a fal-family SHA split. Unseen codecs and TTS families can shift scores.
- The public floats are overall holdout means at threshold 0.5, not per-domain rows and not a fused video claim.
License
`LICENSE` — Apache License 2.0 for DotCheck heads in this repository. Upstream: `NOTICE`.
Citation
`CITATION.cff` · wire inhouse-audio@3 / Helmholtz@3.1 · https://dotcheck.ai/docs
