Team Ai
20 results

probes

xycoord /deception-probes-activations Deception Probes Activations Pre-extracted residual-stream activations for training and evaluating deception detection probes on LLMs. Each example contains per-token hidden states from a specific transformer layer, saved in bfloat16 safetensors format. License This dataset contains activations derived from multiple sources with different licenses. See the LICENSE file for full details. Component Source License Apollo Probe Pairs (statements) Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.texttext-classification1M<n<10M1 likes59k downloads5mo agoHugging FaceAdaptiveChunking /hnet-chunking-probes H-Net chunker boundary probes Longitudinal boundary decisions for 31 H-Net runs, logged on a fixed, byte-identical FLORES+ probe at every checkpoint. This is the raw material for studying when a learned segmentation stabilises. Layout <run>/{step:06d}__{lang}.npz, plus <run>/probe_text.jsonl (the raw probe text, so byte offsets can be aligned to external gold data). 40 log-spaced steps: 0, 1, 2, 4, 8, 16, 32, 64, 128, 200, then every 200 to 6000. The early… See the full description on the dataset page: https://huggingface.co/datasets/AdaptiveChunking/hnet-chunking-probes.texttext-generationn<1K0 likes4.4k downloads24d agoHugging FaceBeicicc /probeshift-activation-cache ProbeShift Activation Cache Residual-stream activations backing the ProbeShift benchmark — a label-free study of linear-probe direction stability under label-preserving semantic shift. Ships so the benchmark's numbers reproduce in minutes (no re-extraction needed). Layout cache_seed{0..4}/<model>/<dataset>/<distribution>/ acts.npy float16 [N, L+1, H] masked-mean-pooled residual stream (L+1 = embeddings + L layers) labels.npy int64 [N]… See the full description on the dataset page: https://huggingface.co/datasets/Beicicc/probeshift-activation-cache.feature-extraction100B<n<1T0 likes2.3k downloads9d agoHugging Facemusicakamusic /emotion-probes-raw-activations0 likes1.2k downloads4mo agoHugging Facetimaeus /lang5_probes Selected Probes Each probe is a CSV with prompt, prompt_len, and target columns. All targets are 0/1 integers unless noted. All datasets are balanced (50/50) unless noted. 5 — hist_fig_ismale Entries: 5,000 | Avg prompt length: 20 chars | Max: 70 chars Prompts: Historical figure names (e.g. "Margaret of Clisson", "Billy Mays"). Target: 1 = male, 0 = female — 50% / 50% 6 — hist_fig_isamerican Entries: 5,000 | Avg prompt length: 17 chars | Max: 65 chars… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/lang5_probes.tabular100K<n<1M0 likes597 downloads5mo agoHugging Facemmtf /probes-activations probes-activations Token-level hidden-state activations (bfloat16) for the top-10 layers per model (ranked by validation token-level code-masked AUC from a full layer sweep), extracted over the SVEN cyber-vulnerability dataset (1,430 examples). Built for linear-probe / natural-language-activation (NLA) research. Activations are stored per model, per layer so a single layer can be pulled on its own (e.g. on Colab) without regenerating from the base model: from huggingface_hub… See the full description on the dataset page: https://huggingface.co/datasets/mmtf/probes-activations.text-classification0 likes569 downloads4mo agoHugging Face