Team Ai
Datasetpublic

Beicicc/probeshift-activation-cache

ProbeShift Activation Cache Residual-stream activations backing the ProbeShift benchmark — a label-free study of linear-probe direction stability under label-preserving semantic shift. Ships so the benchmark's numbers reproduce in minutes (no re-extraction needed). Layout cache_seed{0..4}/<model>/<dataset>/<distribution>/ acts.npy float16 [N, L+1, H] masked-mean-pooled residual stream (L+1 = embeddings + L layers) labels.npy int64 [N]… See the full description on the dataset page: https://huggingface.co/datasets/Beicicc/probeshift-activation-cache.

sourceHugging Facemitupdated 10d agoView on Hugging Face
0likes2.3kdownloads
Dataset Card

ProbeShift Activation Cache

Residual-stream activations backing the ProbeShift benchmark — a label-free study of linear-probe direction stability under label-preserving semantic shift. Ships so the benchmark's numbers reproduce in minutes (no re-extraction needed).

Layout

cache_seed{0..4}/<model>/<dataset>/<distribution>/
    acts.npy     float16  [N, L+1, H]   masked-mean-pooled residual stream (L+1 = embeddings + L layers)
    labels.npy   int64    [N]
    ids.npy      int64    [N]            stable example ids (align across distributions)
    meta.json    {model, dataset, distribution, pooling, n, n_layers, hidden}
  • —models (8): pythia-70m/160m/410m/1.4b/6.9b, gpt2, gpt2-medium, qwen2.5-0.5b
  • —datasets (14 → 12 concepts): sst2, imdb (sentiment); agnews, dbpedia (topic); counterfact (truth); emotion; tweethate; tweetirony; tweetoffensive; subj (subjectivity); spam; cola (grammaticality); stance; amazon_cf (counterfactual)
  • —distributions: train, iid, paraphrase/domain/length (label-preserving OOD shifts), aug0/aug1/aug2 (de/fr/ru back-translation augmentations)
  • —seeds: 0–4 (each an independent example draw + independent paraphrase — Option A replication)

Load

python
import numpy as np
acts = np.load("cache_seed0/pythia-410m/sst2/iid/acts.npy", mmap_mode="r")  # [N, L+1, H] fp16
labels = np.load("cache_seed0/pythia-410m/sst2/iid/labels.npy")

Valid configurations

The paper's paraphrase × LogReg grid has 472 valid configurations: 7 models (pythia-70m/160m/410m/1.4b, gpt2, gpt2-medium, qwen2.5-0.5b) × 14 datasets × 5 seeds = 490, minus 18 cells skipped for massive activations. The pythia-6.9b shards (seed 0) back a preliminary spot-check. The 472 valid keys are those in the supplementary package's records/predictors.jsonl.

Paper and code

Kun Zhang and Chunwei Xia. ProbeShift: Predicting Probe-Direction Rotation from Stability is Largely a Sampling-Noise Floor. Proceedings of the 18th Asian Conference on Machine Learning (ACML 2026), PMLR. The per-configuration records and all analysis scripts ship as the paper's supplementary material (supplement/README.md). Activations were produced on RTX 4090 GPUs (≤200 GPU·h, $0 API, zero new annotation).