Team Ai
Datasetpublic

jub-aer/ConsistencyBench-interpretability

ConsistencyBench-Interpretability Extension White-box mechanistic analysis of logical inconsistency using Qwen/Qwen2.5-1.5B-Instruct (local, full activation access) as a dedicated interpretability testbed, distinct from the 17-model black-box leaderboard. Contents layer_probe_results.csv - per-layer logistic-regression probe accuracy for decoding "will this response be inconsistent?" directly from residual-stream activations activation_patching.csv - literal… See the full description on the dataset page: https://huggingface.co/datasets/jub-aer/ConsistencyBench-interpretability.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes15downloads

jub-aer/ConsistencyBench-interpretability · main · files are served by the source, never re-hosted here