AnonyJterwe/length-confound-benchmark
Cached benchmark data — length-confound audit Pre-computed residual-stream hidden states and derived features for auditing hallucination detectors. The audit re-runs in under an hour once downloaded, versus roughly 30 hours of forward passes to rebuild from scratch. Contents {model}_{dataset}_rtraj_features.npz — 17 conditions. Keys: labels, responses, questions, proj_h_reasoning, proj_a_reasoning, proj_m_reasoning, reasoning_dim.… See the full description on the dataset page: https://huggingface.co/datasets/AnonyJterwe/length-confound-benchmark.
Cached benchmark data — length-confound audit
Pre-computed residual-stream hidden states and derived features for auditing hallucination detectors. The audit re-runs in under an hour once downloaded, versus roughly 30 hours of forward passes to rebuild from scratch.
Contents
{model}_{dataset}_rtraj_features.npz — 17 conditions. Keys: labels, responses, questions, proj_h_reasoning, proj_a_reasoning, proj_m_reasoning, reasoning_dim.
{model}_{dataset}_rtraj_hidden.npz — 17 conditions, about 22 GB total. Key: hidden_states, shape (N, L+1, D).
{model}_reasoning_subspace.npz — 3 files. Keys: V_R, singular_values, semantic_dim, reasoning_dim.
selfcheck/, semantic_entropy/ — cached detector scores for the six closed-book conditions.
phase3/, phase4/, phase3_rebuttal/ — per-fold AUROC results.
Conditions
13 conditions from the paper plus 4 added during review (Qwen and Mistral on HaluEval and NQ-Open). Llama-3 on TruthfulQA is not included.
Reproducing without the hidden states
phase3/phase3_main_results.json holds per-fold AUROCs for all detectors, conditions, and evaluation modes. Running analysis/build_tables.py from the code repository regenerates Tables 1 through 5 from this file alone.
Table 1 and the length-only floor need only labels and responses from the feature files.
