datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Rosetta-Activations
Rosetta Activations
Updated: 2026-06-15 02:30 UTC
Contrastive activation extractions for 17 semantic concepts across 46 language models,
supporting cross-architecture mechanistic interpretability research.
Companion concept pair corpus: jamesrahenry/Rosetta_Concept_Pairs
Papers: forthcoming
Dataset Structure
Rosetta-Activations/
├── rcp_v1/ # Current extraction line — richest data (N≈2000)
│ └── {Model_Name}/
│ ├── calibration_{concept}.npy… See the full description on the dataset page: https://huggingface.co/datasets/james-ra-henry/Rosetta-Activations.deception-probes-activations
Deception Probes Activations
Pre-extracted residual-stream activations for training and evaluating deception
detection probes on LLMs. Each example contains per-token hidden states from a
specific transformer layer, saved in bfloat16 safetensors format.
License
This dataset contains activations derived from multiple sources with different licenses.
See the LICENSE file for full details.
Component
Source
License
Apollo Probe Pairs (statements)
Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.sycophancy-activationsauthority-activationsdeception-activationsGLM-5.3-Flash-calibration-activations-v1
GLM-5.3-Flash calibration activations v1 (BF16, natural routing)
Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048
tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in
and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up
input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth).
Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.action-atlas-oft-activationscached-activationsgpt_oss_20b_doorkey_boundary_activationsbench-af-activationsmsm-packaging-aft-setA-activations
bcywinski/msm-packaging-aft-setA-activations
Mean residual-stream activations of Qwen/Qwen3.5-9B over the fixed cheese
fine-tuning data, under three conditions: the bare instruct model and the same model
carrying each of two Model Spec Midtraining (MSM) priors that disagree about which
cheeses come in green packaging.
The point of the set is that the fine-tuning data is identical in all three: these
are the activations of the demonstrations a fine-tune is about to be trained on… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-aft-setA-activations.crosscoder-activationsqwen3_4b_restricted_Final-activationsbench-af-activations
Bench-AF: Alignment Faking Detection Activations & Probes
Activation caches and cross-validated linear probes for detecting alignment faking
in LLMs. Part of the Bench-AF research project.
Models
Model
Base
Adapter
llama-3-70b
Meta-Llama-3-70B-Instruct
None
llama-3-70b-base
Meta-Llama-3-70B
None
hal9000
Meta-Llama-3-70B-Instruct
bench-af/hal9000-adapter
pacifist
Meta-Llama-3-70B-Instruct
bench-af/pacifist-adapter
Datasets
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/wing1x/bench-af-activations.vulnerable_activationsnemotron_restricted_Final-activationsactivationstextcraft-qwen7b-activations
