Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01james-ra-henry /Rosetta-Activations Rosetta Activations Updated: 2026-06-15 02:30 UTC Contrastive activation extractions for 17 semantic concepts across 46 language models, supporting cross-architecture mechanistic interpretability research. Companion concept pair corpus: jamesrahenry/Rosetta_Concept_Pairs Papers: forthcoming Dataset Structure Rosetta-Activations/ ├── rcp_v1/ # Current extraction line — richest data (N≈2000) │ └── {Model_Name}/ │ ├── calibration_{concept}.npy… See the full description on the dataset page: https://huggingface.co/datasets/james-ra-henry/Rosetta-Activations.tabularn<1K0 likes507k downloads2mo agoHugging Face02xycoord /deception-probes-activations Deception Probes Activations Pre-extracted residual-stream activations for training and evaluating deception detection probes on LLMs. Each example contains per-token hidden states from a specific transformer layer, saved in bfloat16 safetensors format. License This dataset contains activations derived from multiple sources with different licenses. See the LICENSE file for full details. Component Source License Apollo Probe Pairs (statements) Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.texttext-classification1M<n<10M1 likes59k downloads5mo agoHugging Face03lasrprobegen /refusal-activations Refusal Activations Dataset This dataset is now configured to load the full ~97k samples from jailbreak_mixed_100k.csv. tabular10K<n<100K1 likes8.1k downloads11mo agoHugging Face04PranavViswanath /auditbench-activations-jlens-NLA AuditBench activations, J-lens readouts and NLA verbalizations Every token of every AuditBench prompt and every model response, from meta-llama/Llama-3.3-70B-Instruct (revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b) with one LoRA adapter per cell. Responses were regenerated greedily and run to the model's own stopping point rather than truncated at a fixed length, and the activations, readouts and verbalizations cover the prompt as well as the response. 84 cells across 14… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/auditbench-activations-jlens-NLA.tabulartext-generation100M<n<1B0 likes4.9k downloads2mo agoHugging Face05lasrprobegen /sycophancy-activationstext100K<n<1M0 likes3k downloads11mo agoHugging Face06AISC-Linear-Probe-Gen /deception-activationstabular10K<n<100K0 likes2.9k downloads9mo agoHugging Face07lasrprobegen /authority-activationstext100K<n<1M0 likes2.5k downloads11mo agoHugging Face08crosslingual-rule-following /model-inference-activationstext10K<n<100K0 likes1.6k downloads1mo agoHugging Face09bag100 /action-atlas-groot-activationstabularn<1K0 likes1.6k downloads4mo agoHugging Face10lasrprobegen /deception-activationstabular10K<n<100K2 likes1.4k downloads9mo agoHugging Face11scaleinvariant /sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m SAE Feature Activations — Llama 3.1 8B Instruct, Layer 19 (LMSYS-Chat-1M) This dataset contains Sparse Autoencoder (SAE) feature activations extracted from layer 19 of Meta's Llama 3.1 8B Instruct on conversations from LMSYS-Chat-1M. It also has natural language explainations of features generated by GPT OSS 120B. See subset 4 for details. The SAE used is Goodfire/Llama-3.1-8B-Instruct-SAE-l19, which decomposes layer-19 residual stream activations into interpretable sparse features.… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m.tabularfeature-extraction100M<n<1B0 likes1.2k downloads7mo agoHugging Face12saracandu /olmo-activationstabular10K<n<100K0 likes1.2k downloads2mo agoHugging Face13alliedtoasters /latenet-v0-activations-llama3.1-70b-base meta-llama/Llama-3.1-70B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac). Full-sequence activations (80 layers, 8192 dim, float16, all tokens) from meta-llama/Llama-3.1-70B (base) on 23724 LateNet v0 statements (affirmative + negated). Extracted via NDIF. Raw statements only (no chat template). Prompts ordered by negated→generator→pair_id for contiguous domain shards. Contents… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-70b-base.tabularfeature-extraction10K<n<100K0 likes1k downloads6mo agoHugging Face14malaiwah /GLM-5.3-Flash-calibration-activations-v1 GLM-5.3-Flash calibration activations v1 (BF16, natural routing) Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048 tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth). Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.tabularn<1K0 likes860 downloads1mo agoHugging Face15alliedtoasters /latenet-v0-activations-llama3.1-405b-base meta-llama/Llama-3.1-405B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-405B (revision b906e4dc842aa489c962f9db26554dcfdde901fe). LateNet v0 activations for Llama 3.1 405B base (all layers, full sequence) Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-125 16384 - 20 - Prompts: 23724 Format version: 2.0 Load with lmprobe from lmprobe import load_activations, Probe acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-405b-base.tabularfeature-extraction10K<n<100K0 likes844 downloads6mo agoHugging Face16latent-lab /got-activations-llama3.1-405b-base meta-llama/Llama-3.1-405B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-405B (revision unknown). Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-125 16384 - 12 - Prompts: 7660 Format version: 1.1 Load with lmprobe from lmprobe import pull_dataset, load_activation_dataset # Option 1: Pull into local cache (enables probe training without re-extraction)… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-llama3.1-405b-base.tabularfeature-extraction1K<n<10K0 likes762 downloads6mo agoHugging Face17avirambo /cci-multilingualrules-model-inference-activationstext10K<n<100K0 likes721 downloads20d agoHugging Face18mulsi /fruit-vegetable-activationstext10K<n<100K1 likes708 downloads2y agoHugging Face19michaelwaves /activations-testtabular1B<n<10B0 likes623 downloads10mo agoHugging Face20latent-lab /got-activations-qwen2.5-0.5b Qwen/Qwen2.5-0.5B — Activation Dataset Cached activations extracted from Qwen/Qwen2.5-0.5B (revision 060db6499f32faf8b98477b0a26969ef7d8b9987). Full-sequence activations (24 layers, 896 dim, float16) and top-100 logits from Qwen/Qwen2.5-0.5B on 7,660 Geometry of Truth statements. Per-layer sharding (v1.2) with independent shard boundaries. Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-23 896 - 1 - logits_topk - k=100 last_token 1 1200… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-qwen2.5-0.5b.tabularfeature-extraction1K<n<10K0 likes601 downloads7mo agoHugging Face21lczero-planning /activationstabular100K<n<1M1 likes476 downloads2y agoHugging Face22nirmalendu01 /animal_dataset_activations Animal Dataset Activations This dataset contains neural network activations captured from various models processing the animal_dataset. Dataset Structure The dataset is organized by model (subset) and layer (split): Model (subset): Each model has its own directory Layer (split): Each layer within a model has its own directory with parquet files Files are organized as: {model_name}/{layer_name}/*.parquet Loading the Dataset The dataset uses HuggingFace's… See the full description on the dataset page: https://huggingface.co/datasets/nirmalendu01/animal_dataset_activations.text10K<n<100K0 likes473 downloads9mo agoHugging Face23sarel /creditscope-fino1-activationstabular1K<n<10K0 likes473 downloads7mo agoHugging Face24bag100 /action-atlas-oft-activationstabularn<1K0 likes473 downloads4mo agoHugging Face25scaleinvariant /llama-3.2-1b-instruct-lmsys-chat-1m-activations Llama 3.2 1B Instruct Activations (LMSYS-Chat-1M) This dataset contains whole-model residual stream activations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Each row stores the complete residual stream across all 16 transformer layers for a single prompt — both the full-sequence activations and the final-token activations. Note: This is a subset, 8% (from 2 workers of 25) of the full dataset. The complete dataset was ~25 TB and huggingface only… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/llama-3.2-1b-instruct-lmsys-chat-1m-activations.textfeature-extraction10K<n<100K0 likes454 downloads7mo agoHugging Face26debug-probes /cached-activationstabularn<1K0 likes451 downloads1y agoHugging Face27yoheikobashi /SAE_activations_modal_sentencestabular100K<n<1M0 likes426 downloads1y agoHugging Face28quintic /llama_3.1-sae-23-29-code-activationstext10K<n<100K1 likes410 downloads2y agoHugging Face29alliedtoasters /got-activations-llama3.1-70b-base meta-llama/Llama-3.1-70B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac). Geometry of Truth curated dataset activations for Llama 3.1 70B base Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-79 8192 - 4 - Prompts: 7660 Format version: 2.0 Load with lmprobe from lmprobe import load_activations, Probe acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/got-activations-llama3.1-70b-base.tabularfeature-extraction1K<n<10K0 likes388 downloads6mo agoHugging Face30project-telos /gpt_oss_20b_doorkey_boundary_activationstabularn<1K0 likes360 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.