Team Ai
Datasetpublic

jub-aer/ConsistencyBench-interpretability

ConsistencyBench-Interpretability Extension White-box mechanistic analysis of logical inconsistency using Qwen/Qwen2.5-1.5B-Instruct (local, full activation access) as a dedicated interpretability testbed, distinct from the 17-model black-box leaderboard. Contents layer_probe_results.csv - per-layer logistic-regression probe accuracy for decoding "will this response be inconsistent?" directly from residual-stream activations activation_patching.csv - literal… See the full description on the dataset page: https://huggingface.co/datasets/jub-aer/ConsistencyBench-interpretability.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes15downloads
settings

This repository belongs to jub-aer on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameConsistencyBench-interpretability
visibilitypublic
licencecc-by-4.0
gatedno
ownerjub-aer
Account settings
jub-aer/ConsistencyBench-interpretability · Team Ai