Team Ai
Datasetpublic

AlignmentResearch/probe-inference-parity

probe-inference parity fixture Test fixture for the probe-inference package's parity tests (tests/test_parity.py). The tests download this dataset at a pinned revision and check that the package reproduces reference scores on real activations. For each model it holds a few rows of residual-stream activations and token masks, and the reference score of each probe on each row. It holds no probes: the tests load them from AlignmentResearch/probe-inference-weights at the package's… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/probe-inference-parity.

sourceHugging Facemitupdated 5d agoView on Hugging Face
0likes85downloads
Dataset Card

probe-inference parity fixture

Test fixture for the probe-inference package's parity tests (tests/test_parity.py). The tests download this dataset at a pinned revision and check that the package reproduces reference scores on real activations.

For each model it holds a few rows of residual-stream activations and token masks, and the reference score of each probe on each row. It holds no probes: the tests load them from `AlignmentResearch/probe-inference-weights` at the package's pinned revision, exactly as users do.

DirectoryModelRowsProbes
nemotron-3-super-120b/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF1680EFC, axial
qwen3.5-9b/Qwen/Qwen3.5-9B32linear, MLP, EFC, axial

Files in each model directory:

  • —acts.pt: {layer: Tensor[rows, seq, d_model]} in bfloat16, the outputs of decoder block layer (Hugging Face hidden_states[layer + 1]), over each row's follow-up window.
  • —masks.pt: tokens, prompt_mask, completion_mask and followup_start_positions for the same rows.
  • —reference.json: per probe, score (EFC, axial), or L<layer> logits plus combined (the mean of per-layer sigmoids) for linear and MLP, one value per row.

manifest.json lists the sets and, for each, the probes' paths in the weights repository (for example qwen3.5-9b/efc). The Nemotron rows mix two follow-up lengths, so a batch has ragged read windows (27 and 24 tokens).

Licences and attribution

The test data and this card are released by FAR AI, Inc. under the MIT licence (LICENSE).

They are derived from the models below (this fixture uses Qwen/Qwen3.5-9B and nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16). Both upstream licences let us license derived works under our own terms provided we include their licence texts and keep their attribution notices, so both ship here unchanged (see NOTICE):

ModelLicenceLicence file
Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.6-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17BApache-2.0LICENSE-QWEN-APACHE-2.0.txt
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16NVIDIA Nemotron Open Model License (v. December 15, 2025)LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt

Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.