probes
emotion-probesprobe_seg_llava-1.5-pt-ift2stage_const_probes_original_augmented_original_egregious_cake_bake-d1f9cf4bconst_probes_original_augmented_original_akc_turkey_imamoglu_detention-3fe1cf9a2stage_const_probes_original_augmented_original_akc_us_tariffs-c92c33c72stage_const_probes_original_augmented_original_subtle_roman_concrete-a8f8f3e4probe_seg_llava-1.5-pt-vpt-iftprobe_seg_ola-vlm-pt-ift
Datasets
All datasets matching “probes”deception-probes-activations
Deception Probes Activations
Pre-extracted residual-stream activations for training and evaluating deception
detection probes on LLMs. Each example contains per-token hidden states from a
specific transformer layer, saved in bfloat16 safetensors format.
License
This dataset contains activations derived from multiple sources with different licenses.
See the LICENSE file for full details.
Component
Source
License
Apollo Probe Pairs (statements)
Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.hnet-chunking-probes
H-Net chunker boundary probes
Longitudinal boundary decisions for 31 H-Net runs, logged on a fixed, byte-identical
FLORES+ probe at every checkpoint. This is the raw material for studying when a learned
segmentation stabilises.
Layout
<run>/{step:06d}__{lang}.npz, plus <run>/probe_text.jsonl (the raw probe text, so byte
offsets can be aligned to external gold data).
40 log-spaced steps: 0, 1, 2, 4, 8, 16, 32, 64, 128, 200, then every 200 to 6000.
The early… See the full description on the dataset page: https://huggingface.co/datasets/AdaptiveChunking/hnet-chunking-probes.probeshift-activation-cache
ProbeShift Activation Cache
Residual-stream activations backing the ProbeShift benchmark — a label-free study of
linear-probe direction stability under label-preserving semantic shift. Ships so the
benchmark's numbers reproduce in minutes (no re-extraction needed).
Layout
cache_seed{0..4}/<model>/<dataset>/<distribution>/
acts.npy float16 [N, L+1, H] masked-mean-pooled residual stream (L+1 = embeddings + L layers)
labels.npy int64 [N]… See the full description on the dataset page: https://huggingface.co/datasets/Beicicc/probeshift-activation-cache.emotion-probes-raw-activationslang5_probes
Selected Probes
Each probe is a CSV with prompt, prompt_len, and target columns. All targets are 0/1 integers unless noted. All datasets are balanced (50/50) unless noted.
5 — hist_fig_ismale
Entries: 5,000 | Avg prompt length: 20 chars | Max: 70 chars
Prompts: Historical figure names (e.g. "Margaret of Clisson", "Billy Mays").
Target: 1 = male, 0 = female — 50% / 50%
6 — hist_fig_isamerican
Entries: 5,000 | Avg prompt length: 17 chars | Max: 65 chars… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/lang5_probes.probes-activations
probes-activations
Token-level hidden-state activations (bfloat16) for the top-10 layers per model
(ranked by validation token-level code-masked AUC from a full layer sweep), extracted over
the SVEN cyber-vulnerability dataset (1,430 examples). Built for linear-probe / natural-language-activation (NLA) research.
Activations are stored per model, per layer so a single layer can be pulled on its own
(e.g. on Colab) without regenerating from the base model:
from huggingface_hub… See the full description on the dataset page: https://huggingface.co/datasets/mmtf/probes-activations.
