datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
deception-probes-activations
Deception Probes Activations
Pre-extracted residual-stream activations for training and evaluating deception
detection probes on LLMs. Each example contains per-token hidden states from a
specific transformer layer, saved in bfloat16 safetensors format.
License
This dataset contains activations derived from multiple sources with different licenses.
See the LICENSE file for full details.
Component
Source
License
Apollo Probe Pairs (statements)
Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.hnet-chunking-probes
H-Net chunker boundary probes
Longitudinal boundary decisions for 31 H-Net runs, logged on a fixed, byte-identical
FLORES+ probe at every checkpoint. This is the raw material for studying when a learned
segmentation stabilises.
Layout
<run>/{step:06d}__{lang}.npz, plus <run>/probe_text.jsonl (the raw probe text, so byte
offsets can be aligned to external gold data).
40 log-spaced steps: 0, 1, 2, 4, 8, 16, 32, 64, 128, 200, then every 200 to 6000.
The early… See the full description on the dataset page: https://huggingface.co/datasets/AdaptiveChunking/hnet-chunking-probes.probeshift-activation-cache
ProbeShift Activation Cache
Residual-stream activations backing the ProbeShift benchmark — a label-free study of
linear-probe direction stability under label-preserving semantic shift. Ships so the
benchmark's numbers reproduce in minutes (no re-extraction needed).
Layout
cache_seed{0..4}/<model>/<dataset>/<distribution>/
acts.npy float16 [N, L+1, H] masked-mean-pooled residual stream (L+1 = embeddings + L layers)
labels.npy int64 [N]… See the full description on the dataset page: https://huggingface.co/datasets/Beicicc/probeshift-activation-cache.emotion-probes-raw-activationslang5_probes
Selected Probes
Each probe is a CSV with prompt, prompt_len, and target columns. All targets are 0/1 integers unless noted. All datasets are balanced (50/50) unless noted.
5 — hist_fig_ismale
Entries: 5,000 | Avg prompt length: 20 chars | Max: 70 chars
Prompts: Historical figure names (e.g. "Margaret of Clisson", "Billy Mays").
Target: 1 = male, 0 = female — 50% / 50%
6 — hist_fig_isamerican
Entries: 5,000 | Avg prompt length: 17 chars | Max: 65 chars… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/lang5_probes.probes-activations
probes-activations
Token-level hidden-state activations (bfloat16) for the top-10 layers per model
(ranked by validation token-level code-masked AUC from a full layer sweep), extracted over
the SVEN cyber-vulnerability dataset (1,430 examples). Built for linear-probe / natural-language-activation (NLA) research.
Activations are stored per model, per layer so a single layer can be pulled on its own
(e.g. on Colab) without regenerating from the base model:
from huggingface_hub… See the full description on the dataset page: https://huggingface.co/datasets/mmtf/probes-activations.cached-activationsProbeSDF_Result
ProbeSDF Result
LMCAP 历史物体重建资产集,共 42 个已确认候选。单位统一为米,颜色为 mesh 中保存的逐顶点 RGB。
每个物体目录包含:
mesh_source.*:历史原始 mesh,不改几何与坐标。
mesh_centered.ply:以 AABB 中点居中的统一 PLY。
preview.jpg:正面、背面、侧面真实顶点色预览。
metadata.json:名称、尺寸、哈希、来源与居中变换。
Objects
ID
中文名称
Slug
批次
目录
H001
浅灰色 marker 盒
light_gray_marker_box
0707
objects/h001_light_gray_marker_box
H010
黑色产品盒
black_product_box
0807
objects/h010_black_product_box
H011
黑色长方盒
black_rectangular_box
0807… See the full description on the dataset page: https://huggingface.co/datasets/momaxai/ProbeSDF_Result.llama8b-layer15-sae-probes
Llama8B Sparse Probing Activations
This repository contains activation data accompanying the paper Learning a Generative Meta-Model of LLM Activations.
Project page: https://generative-latent-prior.github.io
Code: https://github.com/g-luo/generative_latent_prior
Quick Start
With this data, you can evaluate GLPs via sparse probing.
The activations are derived from the binary classification datasets from Kantamneni et. al., 2025.
The activations are taken only from… See the full description on the dataset page: https://huggingface.co/datasets/generative-latent-prior/llama8b-layer15-sae-probes.ProbeScout-features
ProbeScout features
Frozen SigLIP features for Stanford Cars, HICO-DET and CelebA:
float32 global embeddings [N, 768], float16 patch tokens [N, 196, 768],
and records that map each row to an image ID. Total size: 81.10 GB.
Use the ProbeScout setup guide
and dataset-specific download commands.
Download only the dataset you need. Place its dataset/ contents under the code
repository's dataset/ directory, preserving paths and row order.
These features skip extraction for training… See the full description on the dataset page: https://huggingface.co/datasets/Ian100/ProbeScout-features.activation_steering
Activation Steering Baseline
Generations produced with the difference-in-means activation steering baseline.
This dataset is part of the data release for the paper Predicting Future Behaviors in Reasoning Models Enables Better Steering.
The data is organized as <model>/<dataset>/.... Each row below links to the browsable folder for that model and dataset, where the individual files can be viewed and downloaded.
Data
Model
Dataset
Files… See the full description on the dataset page: https://huggingface.co/datasets/future-probes/activation_steering.probe_predictions
Future Probe Predictions
Predictions of future behavior from linear and MLP probes, per layer, for both the full reasoning trace (prediction features) and the response-only setting (detection features).
This dataset is part of the data release for the paper Predicting Future Behaviors in Reasoning Models Enables Better Steering.
You can view the predictions in the interactive viewer.
Download the *_outputs.json file from one of the linear / MLP links below, and upload it to the… See the full description on the dataset page: https://huggingface.co/datasets/future-probes/probe_predictions.aft-audit-probes
geodesic-research/aft-audit-probes
Local-pipeline snapshot published via --push-from-local. All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision.
Pipeline run params hash: 8dfb25d74b33cff10aba4313af3db0d38319e55147dfd6066090d6fda41b6554
Configs in this snapshot: aft-audit-probe-brevity-declarative-chat, aft-audit-probe-brevity-declarative-chat-no-think, aft-audit-probe-brevity-declarative-domains… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/aft-audit-probes.2026-09-18-colosseum-hospital-midshift-probes
colosseum_hospital mid-shift probe on the fixed-harness self-sacrificial cells: at checkpoints 1E, 2P and 2E each coalition seat (and one bystander) is shown its own transcript so far and asked whether what it did was unethical
field
value
experiment
colosseum_hospital mid-shift probe on the fixed-harness self-sacrificial cells: at checkpoints 1E, 2P and 2E each coalition seat (and one bystander) is shown its own transcript so far and asked whether what it did was… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-18-colosseum-hospital-midshift-probes.2026-09-24-colosseum-hospital-midshift-probes
colosseum_hospital mid-shift probe on the fixed-harness self-sacrificial cells: at checkpoints 1E, 2P and 2E each coalition seat (and one bystander) is shown its own transcript so far and asked whether what it did was unethical
field
value
experiment
colosseum_hospital mid-shift probe on the fixed-harness self-sacrificial cells: at checkpoints 1E, 2P and 2E each coalition seat (and one bystander) is shown its own transcript so far and asked whether what it did was… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-24-colosseum-hospital-midshift-probes.future_probe_controlled_generation
Future Probe Controlled Generation (FPCG)
Generations produced with Future Probe Controlled Generation, our non-invasive text-level steering method.
This dataset is part of the data release for the paper Predicting Future Behaviors in Reasoning Models Enables Better Steering.
The data is organized as <model>/<dataset>/.... Each row below links to the browsable folder for that model and dataset, where the individual files can be viewed and downloaded.
Data… See the full description on the dataset page: https://huggingface.co/datasets/future-probes/future_probe_controlled_generation.per_sentence_probabilities
Per-Sentence Future-Behavior Probabilities
Ground-truth future-behavior probability dynamics from the Behavior Distribution Analysis, together with the gathered activations, trained probes, and difference-in-means steering vectors.
This dataset is part of the data release for the paper Predicting Future Behaviors in Reasoning Models Enables Better Steering.
You can view these behavior distributions in the interactive viewer.
Download an *_outputs.json file from the table below… See the full description on the dataset page: https://huggingface.co/datasets/future-probes/per_sentence_probabilities.behavioral_stability
Behavioral Stability (Unsteered Generations)
Unsteered baseline generations and their measured behavior, used as the reference point for steering.
This dataset is part of the data release for the paper Predicting Future Behaviors in Reasoning Models Enables Better Steering.
The data is organized as <model>/<dataset>/.... Each row below links to the browsable folder for that model and dataset, where the individual files can be viewed and downloaded.
Data… See the full description on the dataset page: https://huggingface.co/datasets/future-probes/behavioral_stability.llama1b-layer07-sae-probes
Llama1B Sparse Probing Activations
This repository contains activation data accompanying the paper Learning a Generative Meta-Model of LLM Activations.
Project page: https://generative-latent-prior.github.io
Code: https://github.com/g-luo/generative_latent_prior
Quick Start
With this data, you can evaluate GLPs via sparse probing.
The activations are derived from the binary classification datasets from Kantamneni et. al., 2025.
The activations are taken only from… See the full description on the dataset page: https://huggingface.co/datasets/generative-latent-prior/llama1b-layer07-sae-probes.deep-space-probes
Deep Space Probes -- Merged Hourly Data
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
Merged hourly magnetic field, solar wind plasma, and energetic particle measurements from humanity's four most distant spacecraft: Voyager 1, Voyager 2, Pioneer 10, and Pioneer 11.
Each record includes spacecraft position (heliocentric distance, HGI latitude/longitude), interplanetary magnetic field components (RTN… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/deep-space-probes.emotion-probes
Emotion Probes Dataset
Synthetic datasets for extracting emotion and emotion-deflection probes from large language models. Built from the methodology described in Anthropic's "Emotion Concepts and their Function in a Large Language Model" (Sofroniew et al., April 2026).
Files
File
Rows
Model
Description
expression/stories.parquet
205,200
Gemini 3.1 Pro Preview
Emotional stories across 171 emotions and 100 topics
expression/neutral_stories.parquet
1,200
Gemini… See the full description on the dataset page: https://huggingface.co/datasets/ryancodrai/emotion-probes.indy-mech-extension-qwen3-8b-persona-probes
indy-mech-extension — Qwen3-8B persona / broken-text probes and steering directions
Residual-stream activations, linear probe results, fitted steering directions and blind human-rubric
judgments for Qwen/Qwen3-8B, extending phase 17 of the
CoT-spiking research programme.
Start with NARRATIVE.md (how the work went, including the mistakes and one retraction),
then README.md (probe results) and results/RESULTS-steering.md (steering results).
⚠ CONTENT WARNING AND… See the full description on the dataset page: https://huggingface.co/datasets/mild-rgb/indy-mech-extension-qwen3-8b-persona-probes.gemma-2-9b-it-probes-jan2026japanese-indirect-prompt-injection-probes
Japanese Indirect Prompt-Injection Probes (v0.1)
English | 日本語
Summary
60 probes for testing whether an LLM follows an instruction hidden inside an untrusted
document (email, RAG chunk, tool output, HTML comment, ...) instead of doing the user's
actual task. Written natively for Japanese rather than machine-translated, to cover what
makes Japanese different: keigo-style polite injections, full-width / hiragana / romaji
obfuscation, fake 【システム】 markers, and mixed… See the full description on the dataset page: https://huggingface.co/datasets/masahiroid/japanese-indirect-prompt-injection-probes.cognitive_map_probes_results2026-07-29-msm-philosophy-spec-fabrication-probes
Fabrication probes: does model-spec midtraining change fabrication of sourced-looking evidence?
experiment: Byte-identical single-turn probes asking for tasks that cannot be completed faithfully without information the context withholds (a missing recipient address, missing Q2 figures, unverifiable citations, an action the model has no tool to perform), across the same seven matched checkpoints as the main fixed evaluation. Built to attribute a confabulation pattern found… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-fabrication-probes.persona-belief-probesdeep-space-probes
Deep Space Probes -- Merged Hourly Data (TsFile)
This dataset is a lossless conversion to the Apache TsFile
format of the HuggingFace dataset
juliensimon/deep-space-probes.
The original observations come from the NASA Space Physics Data Facility
(SPDF).
Original dataset
Source dataset: juliensimon/deep-space-probes
Author: Julien Simon
Data origin: NASA SPDF (https://spdf.gsfc.nasa.gov/)
License: CC-BY-4.0
Content: Merged hourly measurements of magnetic field… See the full description on the dataset page: https://huggingface.co/datasets/THULab/deep-space-probes.probes
probes
Linear probe directions (npz) per base model, next to persona-vectors
Layout: <base-model>/. Created 2026-09-08 by the org reorganisation (Phase 1); full old-to-new map in docs/hf_org_reorg_260908.md of the false-facts-finetuning code repo, also uploaded here as MIGRATION_260908.md.
subfolder
copied from
note
gemma-4-31b/
probes-gemma-4-31b
glm-4.7-flash/
probes-glm-4.7-flash
olmo-3-1025-7b/
probes-olmo-3-1025-7b
qwen3.6-27b/
probes-qwen3.6-27b
svo_probes
SVO-Probes
This dataset comes from https://github.com/deepmind/svo_probes.
Usage
from datasets import load_dataset
# Note that the following line says "train" split, but there are actually no splits in this dataset.
dataset = load_dataset("MichiganNLP/svo_probes", split="train")
# To see an example, access the first element of the dataset with `dataset[0]`.
