RZ0109/Dead_Salmon_Interpretability
1
๐ The Dead Salmons of AI Interpretability
  
An interactive marimo notebook reproducing and extending Mรฉloux et al., The Dead Salmons of AI Interpretability (arXiv:2512.18792, Dec 2025).
Interpretability methods โ linear probes, PCA, SAEs, circuit discovery โ can produce plausible-looking, statistically significant "explanations" from randomly initialized networks that have learned nothing. The authors call these artifacts "dead salmons," after a 2009 fMRI study that found significant brain activity in a dead Atlantic salmon. Their proposed fix: treat interpretability as hypothesis testing against a null distribution from random computation.
What this notebook does
Every quantitative section has interactive mo.ui sliders and dropdowns โ move them and watch the artifact appear and disappear.
Run it
Interactive (recommended): Open in HF Spaces
Locally:
git clone https://github.com/RogerZhu0109/dead-salmon-interpretability
cd dead-salmon-interpretability
pip install marimo
marimo edit --sandbox notebooks/walkthrough.pyStack
- marimo โ reactive notebook, deployed as a server-side app
- PyTorch + HuggingFace Transformers โ random-init
prajjwal1/bert-tiny - scikit-learn โ probes and PCA
- Hosted on HuggingFace Spaces (Docker, CPU)
Reference
Mรฉloux, A., Dirupo, G., Portet, F., & Peyrard, M. (2025). The Dead Salmons of AI Interpretability. arXiv:2512.18792
