datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
behavioral_stability
Behavioral Stability (Unsteered Generations)
Unsteered baseline generations and their measured behavior, used as the reference point for steering.
This dataset is part of the data release for the paper Predicting Future Behaviors in Reasoning Models Enables Better Steering.
The data is organized as <model>/<dataset>/.... Each row below links to the browsable folder for that model and dataset, where the individual files can be viewed and downloaded.
Data… See the full description on the dataset page: https://huggingface.co/datasets/future-probes/behavioral_stability.svo_probes
SVO-Probes
This dataset comes from https://github.com/deepmind/svo_probes.
Usage
from datasets import load_dataset
# Note that the following line says "train" split, but there are actually no splits in this dataset.
dataset = load_dataset("MichiganNLP/svo_probes", split="train")
# To see an example, access the first element of the dataset with `dataset[0]`.
tower-probes-resultsbelief-state-probes-results-mess3-v1
belief-state-probes-results-mess3-v1
FULL mess3 (paper-exact a=0.85,x=0.05, seed 42, 1e6 steps) results: repro (R2 0.987, heldout-beliefs 0.977, shuffle collapses 0.67->0.02, at Bayes floor), battery (random-init 0.938 vs trained 0.987; last-4-token baseline 0.965), positive control (belief clamp provably swaps probe readout yet next-token TV moves only 0.477->0.426 — probe basis not the computational basis), fractal + emergence simplex PNGs, full checkpoints.… See the full description on the dataset page: https://huggingface.co/datasets/latkes/belief-state-probes-results-mess3-v1.saeber-virulence-probes-clusteredpushT_bs48x7k__bs64_no-merge_probesprompt-writer-multimodal-hallucination-probesprobes_latent_env_in_openvlapushT_bs48x7k__bs128_probes
