Team Ai
Datasetpublic

mmtf/probes-activations

probes-activations Token-level hidden-state activations (bfloat16) for the top-10 layers per model (ranked by validation token-level code-masked AUC from a full layer sweep), extracted over the SVEN cyber-vulnerability dataset (1,430 examples). Built for linear-probe / natural-language-activation (NLA) research. Activations are stored per model, per layer so a single layer can be pulled on its own (e.g. on Colab) without regenerating from the base model: from huggingface_hub… See the full description on the dataset page: https://huggingface.co/datasets/mmtf/probes-activations.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes457downloads
Dataset Card

probes-activations

Token-level hidden-state activations (bfloat16) for the top-10 layers per model (ranked by validation token-level code-masked AUC from a full layer sweep), extracted over the SVEN cyber-vulnerability dataset (1,430 examples). Built for linear-probe / natural-language-activation (NLA) research.

Activations are stored per model, per layer so a single layer can be pulled on its own (e.g. on Colab) without regenerating from the base model:

python
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
import numpy as np

m = "google_gemma-3-27b-it"
f = hf_hub_download("mmtf/probes-activations", f"{m}/layer_19.safetensors", repo_type="dataset")
acts = load_file(f)["activations"]            # bf16 tensor, shape (T_tokens, hidden)
y    = np.load(hf_hub_download("mmtf/probes-activations", f"{m}/y.npy", repo_type="dataset"))           # int8 per-token label
eids = np.load(hf_hub_download("mmtf/probes-activations", f"{m}/example_ids.npy", repo_type="dataset")) # int32 example id per token

Layout

Per model directory <org>_<model>/:

filedtypeshapemeaning
layer_NN.safetensorsbf16(T_tokens, hidden)hidden state at layer NN; tensor key activations
y.npyint8(T_tokens,)per-token positive-span (vulnerable) label
example_ids.npyint32(T_tokens,)example index each token belongs to
offsets.npzint32per-row (T_row, 2)char-span offsets per example (extractor parity)
meta.json——model, nlayers, hidden, ntokens, nrows, maxlength, pos_tokens
top_layers.json——selected layers + their sweep scores

Shared, at the repo root:

  • —data/dataset.jsonl — SVEN examples (code + vulnerability spans + labels)
  • —data/sven_split_meta.json — train/val/test split, defined by example (map tokens→examples via example_ids)

Models & selected layers

modeldirlayershiddentokensbest Lbest AUCtop-10 layers
Qwen/Qwen2.5-Coder-32B-InstructQwen_Qwen2.5-Coder-32B-Instruct645120561,266250.784825 36 37 39 40 41 42 43 44 45
Qwen/Qwen3-32BQwen_Qwen3-32B645120561,266270.78217 23 24 25 26 27 28 29 30 42
Qwen/Qwen3.6-27BQwen_Qwen3.6-27B645120612,249300.771714 16 17 18 19 30 31 32 46 49
google/gemma-3-12b-itgoogle_gemma-3-12b-it483840690,148150.754609 10 11 12 13 14 15 19 24 28
google/gemma-3-12b-ptgoogle_gemma-3-12b-pt483840690,148130.761410 11 13 14 15 16 18 22 26 47
google/gemma-3-1b-itgoogle_gemma-3-1b-it261152690,148250.737203 04 06 07 09 10 13 14 23 25
google/gemma-3-1b-ptgoogle_gemma-3-1b-pt261152690,148120.754604 05 07 09 11 12 20 22 24 25
google/gemma-3-27b-itgoogle_gemma-3-27b-it625376690,148190.766116 17 18 19 20 21 22 23 25 26
google/gemma-3-4b-itgoogle_gemma-3-4b-it342560690,14870.75303 04 07 08 09 10 12 23 30 33
google/gemma-3-4b-ptgoogle_gemma-3-4b-pt342560690,148330.768203 04 06 07 08 09 10 13 14 33

How these were produced

  • —One forward pass per example (full sequence, max_length=2048), all hidden states captured, on an NVIDIA GH200.
  • —Stored bfloat16, not float16: Gemma-3 has mid-layer "massive activations" (>65504) that overflow fp16 → NaNs; bf16 keeps fp32's exponent range. (Originals were fp32; bf16 halves size with no overflow.)
  • —The 10 layers per model are the highest-scoring by val_tokens_code_auc in the per-model layer sweep.

Provenance & licensing

Derived from the SVEN dataset; base models are Gemma-3 (governed by Google's Gemma terms) and Qwen (governed by the respective Qwen licenses). These activations are derived representations — downstream use is governed by those upstream dataset/model licenses. license: other reflects that; consult SVEN and the base-model licenses before redistribution or commercial use.