interpretability
gemma4-interpretability
Gemma materials-science interpretability research archive
Research records supporting Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model, Markus J. Buehler. Release identifier: paper-revision-2026-09-06.
This archive supplies the original observations, supporting state arrays, exact prompts, protocols, intervention records, statistics, analysis source, and generated research figures. It includes the original 4B readout and geometry… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-interpretability.Mechanic-Interpretability-Research-Dataopenpi-interpretability-data
openpi-interpretability-data
Interpretability artifacts (activations, conceptors, linear steering vectors, sparse autoencoder vectors and checkpoints) extracted from open vision-language-action (VLA) policy models on the LIBERO, MetaWorld, and RoboCasa benchmarks.
This dataset accompanies an anonymous submission and is shared for double-blind peer review.
Models and benchmarks
Model
Family
Benchmarks
pi0_5 (pi05)
π-series VLA
LIBERO
pi0_fast (pi0fast)… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPsMay1234/openpi-interpretability-data.dino_vit_attnmapsmechanistic-interpretability-skills
Mechanistic Interpretability Skills for Claude Code
The first skill set for LLM refusal geometry extraction, boundary surface mapping, and self-referential mechanistic analysis.
Compatible with: Claude Code, OpenAI Codex, Gemini CLI, and any tool supporting the agentskills.io standard.
Skills
1. refusal-geometry/
Extract and analyze refusal cone geometry from open-weight transformer models using OBLITERATUS.
Capabilities:
6-stage extraction pipeline (model… See the full description on the dataset page: https://huggingface.co/datasets/bedderautomation/mechanistic-interpretability-skills.MARRI-interpretability
MARRI headline model — test-set interpretability export
Per-sample interpretability/diagnostics for the MARRI headline configuration
(dyn ft-rnafm, seed 42 — RNA-FM fine-tuned, RNet2D frozen, dynamic negative
resampling), evaluated on the fixed held-out test split (n=16,058 pairs,
seed 42).
Code and training/evidence logs: https://github.com/GainGod-Xu/MARRI
Files
marri_ft-rnafm_dyn_bs4_seed42_test_results_interpretability.h5 — per-sample
diagnostics:… See the full description on the dataset page: https://huggingface.co/datasets/Xu-AI4Science/MARRI-interpretability.
