Team Ai
Datasetpublic

future-probes/per_sentence_probabilities

Per-Sentence Future-Behavior Probabilities Ground-truth future-behavior probability dynamics from the Behavior Distribution Analysis, together with the gathered activations, trained probes, and difference-in-means steering vectors. This dataset is part of the data release for the paper Predicting Future Behaviors in Reasoning Models Enables Better Steering. You can view these behavior distributions in the interactive viewer. Download an *_outputs.json file from the table below… See the full description on the dataset page: https://huggingface.co/datasets/future-probes/per_sentence_probabilities.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes129downloads
Dataset Card

Per-Sentence Future-Behavior Probabilities

Ground-truth future-behavior probability dynamics from the Behavior Distribution Analysis, together with the gathered activations, trained probes, and difference-in-means steering vectors.

This dataset is part of the data release for the paper **Predicting Future Behaviors in Reasoning Models Enables Better Steering**.

You can view these behavior distributions in the interactive viewer. Download an *_outputs.json file from the table below, and upload it to the viewer using the Load full data... button.

The Browse link opens the full folder for that model and dataset; the Download link downloads the ground-truth test-set outputs.json directly.

Data

ModelDatasetBrowseDownload
DeepSeek-R1-Distill-Llama-8Belephant_aitabrowseoutputs.json
DeepSeek-R1-Distill-Llama-8Bmyopic_rewardbrowseoutputs.json
DeepSeek-R1-Distill-Llama-8Bsepbrowseoutputs.json
DeepSeek-R1-Distill-Llama-8Bsorrybenchbrowseoutputs.json
DeepSeek-R1-Distill-Llama-8Bsurvival_instinctbrowseoutputs.json
DeepSeek-R1-Distill-Llama-8Bwealth_seekingbrowseoutputs.json
QwQ-32Belephant_aitabrowseoutputs.json
QwQ-32Bmyopic_rewardbrowseoutputs.json
QwQ-32Bsepbrowseoutputs.json
QwQ-32Bsorrybenchbrowseoutputs.json
QwQ-32Bsurvival_instinctbrowseoutputs.json
QwQ-32Bsycophancybrowseoutputs.json
QwQ-32Bwealth_seekingbrowseoutputs.json
Qwen3-14Belephant_aitabrowseoutputs.json
Qwen3-14Bmyopic_rewardbrowseoutputs.json
Qwen3-14Bsepbrowseoutputs.json
Qwen3-14Bsorrybenchbrowseoutputs.json
Qwen3-14Bsurvival_instinctbrowseoutputs.json
Qwen3-14Bwealth_seekingbrowseoutputs.json
gpt-oss-20belephant_aitabrowseoutputs.json
gpt-oss-20bmyopic_rewardbrowseoutputs.json
gpt-oss-20bsepbrowseoutputs.json
gpt-oss-20bsorrybenchbrowseoutputs.json
gpt-oss-20bsurvival_instinctbrowseoutputs.json
gpt-oss-20bwealth_seekingbrowseoutputs.json

Paper

Predicting Future Behaviors in Reasoning Models Enables Better Steering

Citation

bibtex
@misc{kortukov2026predictingfuturebehaviorsreasoning,
      title={Predicting Future Behaviors in Reasoning Models Enables Better Steering},
      author={Evgenii Kortukov and Piotr Komorowski and Florian Klein and Paula Engl and Gabriele Sarti and Seong Joon Oh and Sebastian Lapuschkin and Wojciech Samek},
      year={2026},
      eprint={2606.11172},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2606.11172},
}