chengruiqu/risk-representation-llm-data
Risk Representation in LLM Activations Full dataset from a cross-model study testing whether language models have internal representations of risk attitude separable from stimulus encoding. Report: Interactive results page Models Tested Qwen2.5-7B-Instruct (28 layers, 3584 hidden dim, rev a09a3545) Mistral-7B-Instruct-v0.3 (32 layers, 4096 hidden dim, rev c170c708) Key Finding The Representation Paradox: Neither model has orderly behavioral risk… See the full description on the dataset page: https://huggingface.co/datasets/chengruiqu/risk-representation-llm-data.
Risk Representation in LLM Activations
Full dataset from a cross-model study testing whether language models have internal representations of risk attitude separable from stimulus encoding.
Report: Interactive results page
Models Tested
- Qwen2.5-7B-Instruct (28 layers, 3584 hidden dim, rev
a09a3545) - Mistral-7B-Instruct-v0.3 (32 layers, 4096 hidden dim, rev
c170c708)
Key Finding
The Representation Paradox: Neither model has orderly behavioral risk preferences (choices are position-biased, intransitive, and inconsistent). Yet both contain a detectable, multi-dimensional (~10D) internal representation of risk attitude that is separable from stimulus encoding and — in Qwen's case — causally connected to behavior without disrupting non-risk decisions.
Structure
stimuli/ # Shared across models (position-randomized)
├── exp0_behavioral.jsonl # 1,100 binary choices (EV-matched, transitivity, dominance, retest, order-swap)
└── exp1_indifference.jsonl # 12,000 binary choices (600 lotteries × 20 sure-thing levels)
{qwen,mistral}/
├── model_lock.json # Pinned model revision
├── responses/
│ ├── exp0_choices.jsonl # Model's binary choices + logits (Exp 0)
│ ├── exp1_choices.jsonl # Model's binary choices + logits (Exp 1)
│ └── exp5_steered.jsonl # Steered choices at 11 coefficients
├── activations/
│ ├── exp0_layer_*.npy # FP32 hidden states at last prompt token (Exp 0)
│ ├── exp1_layer_*.npy # FP32 hidden states at last prompt token (Exp 1)
│ └── exp3_layer_*.npy # Teacher-forced activations (Exp 3)
├── stimuli/
│ └── exp3_controlled.jsonl # Controlled-answer sequences
└── analysis/
├── exp0_behavioral_summary.json
├── exp2_probing_results.json # All layers, permutation tests, dimensionality, transfer, geometry
├── exp3_controlled_answer.json
├── exp5_steering_dose_response.json
├── lottery_results.json # All 600 lotteries with CE estimates
├── lottery_results_filtered.json # Valid lotteries only (≤2 monotonicity violations)
└── risk_direction.npy # Fitted risk direction vectorExperiments
Activation Files
Each .npy file is a float32 array of shape [n_stimuli, hidden_dim] containing the hidden state at the last prompt token before generation. Layers are 1-indexed (e.g., exp1_layer_21.npy = layer 21 of 28 for Qwen).
Reproducing
pip install numpy scipy scikit-learn
python -c "
import numpy as np, json
X = np.load('qwen/activations/exp1_layer_20.npy') # best layer
with open('qwen/analysis/lottery_results_filtered.json') as f:
lots = json.load(f)
print(f'Activations: {X.shape}, Valid lotteries: {len(lots)}')
"Citation
If you use this dataset, please cite:
@misc{qu2026risk,
title={The Representation Paradox: Risk Attitude in LLM Activations},
author={Qu, Chengrui},
year={2026},
url={https://huggingface.co/datasets/chengruiqu/risk-representation-llm-data}
}