datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Rosetta-Activations
Rosetta Activations
Updated: 2026-06-15 02:30 UTC
Contrastive activation extractions for 17 semantic concepts across 46 language models,
supporting cross-architecture mechanistic interpretability research.
Companion concept pair corpus: jamesrahenry/Rosetta_Concept_Pairs
Papers: forthcoming
Dataset Structure
Rosetta-Activations/
├── rcp_v1/ # Current extraction line — richest data (N≈2000)
│ └── {Model_Name}/
│ ├── calibration_{concept}.npy… See the full description on the dataset page: https://huggingface.co/datasets/james-ra-henry/Rosetta-Activations.tarakanov-notesdeception-probes-activations
Deception Probes Activations
Pre-extracted residual-stream activations for training and evaluating deception
detection probes on LLMs. Each example contains per-token hidden states from a
specific transformer layer, saved in bfloat16 safetensors format.
License
This dataset contains activations derived from multiple sources with different licenses.
See the LICENSE file for full details.
Component
Source
License
Apollo Probe Pairs (statements)
Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.pi05_droid_activationvoice-acting-cutscene-prompts
Cut-Scene Voice-Acting Prompts
Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance
prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting
models. Each prompt describes a single speaker across two sharply contrasting emotional
moments separated by a CUT TO: transition, in a voice-acting stage-direction format
(spoken lines in "quotes", performance notes in (parentheses)).
Total prompts: 4,057,000
Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.what-ai-benchmarks-actually-measure
What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models
Item-level model responses and scores for 53 language models across the
56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting
Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
(Desai et al., 2026,
arxiv.org/abs/2609.08812).
We do not release the prompts from the benchmark datasets, but instead refer to them by
item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.chi-bench
Clinical Healthcare In-Situ Environment
Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark
🎉 χ-Bench has been accepted to NeurIPS 2026, Evaluations & Datasets Track! Read the paper.
What is in this dataset
CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a… See the full description on the dataset page: https://huggingface.co/datasets/actava/chi-bench.meta-active-readingrefusal-activations
Refusal Activations Dataset
This dataset is now configured to load the full ~97k samples from jailbreak_mixed_100k.csv.
action-worldmodel-benchatlas-32-turn-level-actor-critic
32. A turn-level actor-critic derived from the value of computation
1. Question and links
Read this first. The reading copy of this directory is t2ance/atlas-experiments under 32-turn-level-actor-critic/; the saved training steps and the per-token training arrays are only in the Hugging Face repository t2ance/atlas-32-turn-level-actor-critic.
Does a critic that predicts the return at the start of each turn, and is supervised there alone, learn on the… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-32-turn-level-actor-critic.tm-acts-qwen38-multiChain-of-Actionaction3dActivityNet_Captions
About
ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark.
We adopt the official split:
Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions)
Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/ActivityNet_Captions.auditbench-activations-jlens-NLA
AuditBench activations, J-lens readouts and NLA verbalizations
Every token of every AuditBench prompt and every model response, from
meta-llama/Llama-3.3-70B-Instruct (revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b) with one LoRA adapter per cell.
Responses were regenerated greedily and run to the model's own stopping point rather
than truncated at a fixed length, and the activations, readouts and verbalizations
cover the prompt as well as the response.
84 cells across 14… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/auditbench-activations-jlens-NLA.lists-activationsaction-atlas-rollout-videos
Action Atlas — VLA Rollout Videos
Local rollout/ablation videos for Pi0.5, OpenVLA-OFT, X-VLA, GR00T, SmolVLA, and ACT/ALOHA,
organized by model. Companion to action-atlas-{pi05,oft,xvla,groot,smolvla} (SAEs + activations + concepts).
368283 unique mp4 clips, 61.8 GB. Per-model: {'act_aloha': 990, 'groot': 163891, 'oft': 24284, 'pi05': 62468, 'smolvla': 56844, 'xvla': 59806}
manifest.jsonl: one row per clip (model, env, experiment, sha256, bytes, hf_path).
thinking-model-activationsscience-activationsslither-wam-video-actions
Slither Video Actions
This public repository is being built for video and action-conditioned world-model research. It is currently a staging dataset, not a released training set. Prepared gameplay videos, source provenance, and candidate actions will be added incrementally.
Each source metadata file records its YouTube URL, creator, reuse basis, edit recipe, checksums, and prepared files. A prepared video is 30 FPS, silent, and one continuous reviewed interval. Actions are… See the full description on the dataset page: https://huggingface.co/datasets/mundoamundo/slither-wam-video-actions.GUI-Actor-Data
GUI-Actor Data Collection
This is the GUI-Actor Data Collection for GUI grounding training with bounding box supervision.
Update:
We have uploaded the full image files of Uground-bbox (240G).
Project:
Project Page: https://microsoft.github.io/GUI-Actor/
Github Repo: https://github.com/microsoft/GUI-Actor
Paper Link: https://huggingface.co/papers/2506.03143
Usage:
It includes the training data for GUI grounding listed in data/data_config.yaml from… See the full description on the dataset page: https://huggingface.co/datasets/cckevinn/GUI-Actor-Data.action-atlas-viz
Action Atlas visualization bundle
The precomputed data the Action Atlas web frontend reads (https://action-atlas.com). This is all you
need to run the site locally; the SAE weights and raw activations are not required at runtime.
Contents
processed/ per-layer clustering layouts and feature scatter data the frontend renders.
feature_embeddings/ embeddings used for semantic feature search.
descriptions/ generated natural-language feature and concept descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/bag100/action-atlas-viz.metaphors-activationsminecraft-text-action-datasetsycophancy-activationsActivityNet
Description
Dataset V1-2
v1-2_train.tar.gz and v1-2_val.tar.gz
Data (train and val set) associated with ActivityNet release 1.2
v1-2_test.tar.gz
Data (test set only) associated with ActivityNet release 1.2
Dataset V1-3
v1-3_train_val.tar.gz
Additional videos (train val set) collected for ActivityNet release 1.3
v1-3 is an extension of v1-2, so you also need to download v1-2 data and merge to v1.3
v1-3_test.tar.gz
Additional videos (test set only)… See the full description on the dataset page: https://huggingface.co/datasets/YimuWang/ActivityNet.active_matter
How To Load from HuggingFace Hub
Be sure to have the_well installed (pip install the_well)
Use the WellDataModule to retrieve data as follows:
from the_well.benchmark.data import WellDataModule
# The following line may take a couple of minutes to instantiate the datamodule
datamodule = WellDataModule(
"hf://datasets/polymathic-ai/",
"active_matter_cloud_optimized",
)
train_dataloader = datamodule.train_dataloader()
for batch in dataloader:
# Process training batch… See the full description on the dataset page: https://huggingface.co/datasets/polymathic-ai/active_matter.act-datasetDVGT-Dataset
