Team Ai
Apppublic

RZ0109/Dead_Salmon_Interpretability

sourceHugging Faceupdated 5mo agoView on Hugging Face
1likes
App README

๐ŸŸ The Dead Salmons of AI Interpretability

![Open in HF Spaces](https://huggingface.co/spaces/RZ0109/DeadSalmonInterpretability) ![arXiv](https://arxiv.org/abs/2512.18792) ![marimo](https://marimo.io)

An interactive marimo notebook reproducing and extending Mรฉloux et al., The Dead Salmons of AI Interpretability (arXiv:2512.18792, Dec 2025).

Interpretability methods โ€” linear probes, PCA, SAEs, circuit discovery โ€” can produce plausible-looking, statistically significant "explanations" from randomly initialized networks that have learned nothing. The authors call these artifacts "dead salmons," after a 2009 fMRI study that found significant brain activity in a dead Atlantic salmon. Their proposed fix: treat interpretability as hypothesis testing against a null distribution from random computation.

What this notebook does

SectionWhat you see
HookThe 2009 Bennett et al. dead salmon fMRI analogy
SetupRandomly re-initialized bert-tiny, IMDb embeddings with layer/pooling controls
Artifact A โ€” PCAPrincipal components that "explain" sentiment in a network that learned nothing
Artifact B โ€” ProbeLogistic regression achieving well-above-chance CV accuracy on random embeddings
The FixNull distribution across random seeds + empirical p-value
Dead Salmon ZooArchitecture-independent demo: random MLP, random GPT-2-small
Probe Complexity SweepHow more expressive probes find more spurious structure
TakeawaysThree-bullet summary + link back to paper

Every quantitative section has interactive mo.ui sliders and dropdowns โ€” move them and watch the artifact appear and disappear.

Run it

Interactive (recommended): Open in HF Spaces

Locally:

bash
git clone https://github.com/RogerZhu0109/dead-salmon-interpretability
cd dead-salmon-interpretability
pip install marimo
marimo edit --sandbox notebooks/walkthrough.py

Stack

  • โ€”marimo โ€” reactive notebook, deployed as a server-side app
  • โ€”PyTorch + HuggingFace Transformers โ€” random-init prajjwal1/bert-tiny
  • โ€”scikit-learn โ€” probes and PCA
  • โ€”Hosted on HuggingFace Spaces (Docker, CPU)

Reference

Mรฉloux, A., Dirupo, G., Portet, F., & Peyrard, M. (2025). The Dead Salmons of AI Interpretability. arXiv:2512.18792