jub-aer/ConsistencyBench-interpretability
ConsistencyBench-Interpretability Extension White-box mechanistic analysis of logical inconsistency using Qwen/Qwen2.5-1.5B-Instruct (local, full activation access) as a dedicated interpretability testbed, distinct from the 17-model black-box leaderboard. Contents layer_probe_results.csv - per-layer logistic-regression probe accuracy for decoding "will this response be inconsistent?" directly from residual-stream activations activation_patching.csv - literal… See the full description on the dataset page: https://huggingface.co/datasets/jub-aer/ConsistencyBench-interpretability.
This repository belongs to jub-aer on Hugging Face.
Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
