Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01james-ra-henry /Rosetta-Activations Rosetta Activations Updated: 2026-06-15 02:30 UTC Contrastive activation extractions for 17 semantic concepts across 46 language models, supporting cross-architecture mechanistic interpretability research. Companion concept pair corpus: jamesrahenry/Rosetta_Concept_Pairs Papers: forthcoming Dataset Structure Rosetta-Activations/ ├── rcp_v1/ # Current extraction line — richest data (N≈2000) │ └── {Model_Name}/ │ ├── calibration_{concept}.npy… See the full description on the dataset page: https://huggingface.co/datasets/james-ra-henry/Rosetta-Activations.tabularn<1K0 likes539k downloads2mo agoHugging Face02actiquest-dev /tarakanov-notes0 likes54k downloads34m agoHugging Face03xycoord /deception-probes-activations Deception Probes Activations Pre-extracted residual-stream activations for training and evaluating deception detection probes on LLMs. Each example contains per-token hidden states from a specific transformer layer, saved in bfloat16 safetensors format. License This dataset contains activations derived from multiple sources with different licenses. See the LICENSE file for full details. Component Source License Apollo Probe Pairs (statements) Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.texttext-classification1M<n<10M1 likes51k downloads5mo agoHugging Face04NeurIPsMay1234 /pi05_droid_activation0 likes18k downloads3mo agoHugging Face05laion /voice-acting-cutscene-prompts Cut-Scene Voice-Acting Prompts Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting models. Each prompt describes a single speaker across two sharply contrasting emotional moments separated by a CUT TO: transition, in a voice-acting stage-direction format (spoken lines in "quotes", performance notes in (parentheses)). Total prompts: 4,057,000 Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.tabulartext-generation1M<n<10M2 likes17k downloads29d agoHugging Face06madesai /what-ai-benchmarks-actually-measure What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models Item-level model responses and scores for 53 language models across the 56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks (Desai et al., 2026, arxiv.org/abs/2609.08812). We do not release the prompts from the benchmark datasets, but instead refer to them by item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.tabular1M<n<10M1 likes8.7k downloads1mo agoHugging Face07actava /chi-bench Clinical Healthcare In-Situ Environment Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark 🎉 χ-Bench has been accepted to NeurIPS 2026, Evaluations & Datasets Track! Read the paper. What is in this dataset CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a… See the full description on the dataset page: https://huggingface.co/datasets/actava/chi-bench.documenttext-generationn<1K63 likes8.6k downloads11d agoHugging Face08facebook /meta-active-readingtext1B<n<10B38 likes8.5k downloads1y agoHugging Face09lasrprobegen /refusal-activations Refusal Activations Dataset This dataset is now configured to load the full ~97k samples from jailbreak_mixed_100k.csv. tabular10K<n<100K1 likes7k downloads1y agoHugging Face10syCen /action-worldmodel-bench0 likes6.5k downloads3d agoHugging Face11t2ance /atlas-32-turn-level-actor-critic 32. A turn-level actor-critic derived from the value of computation 1. Question and links Read this first. The reading copy of this directory is t2ance/atlas-experiments under 32-turn-level-actor-critic/; the saved training steps and the per-token training arrays are only in the Hugging Face repository t2ance/atlas-32-turn-level-actor-critic. Does a critic that predicts the return at the start of each turn, and is supervised there alone, learn on the… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-32-turn-level-actor-critic.0 likes6.2k downloads16d agoHugging Face12YFC-112358 /tm-acts-qwen38-multi5 likes6.1k downloads26d agoHugging Face13Solomonz /Chain-of-Action2 likes5.7k downloads1y agoHugging Face14lnyan /action3d0 likes4.6k downloads2y agoHugging Face15friedrichor /ActivityNet_Captions About ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark. We adopt the official split: Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions) Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/ActivityNet_Captions.texttext-to-video10K<n<100K16 likes4.3k downloads1y agoHugging Face16PranavViswanath /auditbench-activations-jlens-NLA AuditBench activations, J-lens readouts and NLA verbalizations Every token of every AuditBench prompt and every model response, from meta-llama/Llama-3.3-70B-Instruct (revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b) with one LoRA adapter per cell. Responses were regenerated greedily and run to the model's own stopping point rather than truncated at a fixed length, and the activations, readouts and verbalizations cover the prompt as well as the response. 84 cells across 14… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/auditbench-activations-jlens-NLA.tabulartext-generation100M<n<1B0 likes4.1k downloads2mo agoHugging Face17lasrprobegen /lists-activations0 likes4k downloads1y agoHugging Face18bag100 /action-atlas-rollout-videos Action Atlas — VLA Rollout Videos Local rollout/ablation videos for Pi0.5, OpenVLA-OFT, X-VLA, GR00T, SmolVLA, and ACT/ALOHA, organized by model. Companion to action-atlas-{pi05,oft,xvla,groot,smolvla} (SAEs + activations + concepts). 368283 unique mp4 clips, 61.8 GB. Per-model: {'act_aloha': 990, 'groot': 163891, 'oft': 24284, 'pi05': 62468, 'smolvla': 56844, 'xvla': 59806} manifest.jsonl: one row per clip (model, env, experiment, sha256, bytes, hf_path). 0 likes3.9k downloads3mo agoHugging Face19LoneResearch /thinking-model-activations0 likes3.4k downloads8mo agoHugging Face20lasrprobegen /science-activations0 likes3.3k downloads1y agoHugging Face21mundoamundo /slither-wam-video-actions Slither Video Actions This public repository is being built for video and action-conditioned world-model research. It is currently a staging dataset, not a released training set. Prepared gameplay videos, source provenance, and candidate actions will be added incrementally. Each source metadata file records its YouTube URL, creator, reuse basis, edit recipe, checksums, and prepared files. A prepared video is 30 FPS, silent, and one continuous reviewed interval. Actions are… See the full description on the dataset page: https://huggingface.co/datasets/mundoamundo/slither-wam-video-actions.videovideo-classification0 likes2.9k downloads10d agoHugging Face22cckevinn /GUI-Actor-Data GUI-Actor Data Collection This is the GUI-Actor Data Collection for GUI grounding training with bounding box supervision. Update: We have uploaded the full image files of Uground-bbox (240G). Project: Project Page: https://microsoft.github.io/GUI-Actor/ Github Repo: https://github.com/microsoft/GUI-Actor Paper Link: https://huggingface.co/papers/2506.03143 Usage: It includes the training data for GUI grounding listed in data/data_config.yaml from… See the full description on the dataset page: https://huggingface.co/datasets/cckevinn/GUI-Actor-Data.13 likes2.8k downloads1y agoHugging Face23bag100 /action-atlas-viz Action Atlas visualization bundle The precomputed data the Action Atlas web frontend reads (https://action-atlas.com). This is all you need to run the site locally; the SAE weights and raw activations are not required at runtime. Contents processed/ per-layer clustering layouts and feature scatter data the frontend renders. feature_embeddings/ embeddings used for semantic feature search. descriptions/ generated natural-language feature and concept descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/bag100/action-atlas-viz.1K<n<10K0 likes2.7k downloads4mo agoHugging Face24lasrprobegen /metaphors-activations0 likes2.7k downloads1y agoHugging Face25CraftJarvis /minecraft-text-action-datasettext100K<n<1M1 likes2.7k downloads1y agoHugging Face26lasrprobegen /sycophancy-activationstext100K<n<1M0 likes2.6k downloads11mo agoHugging Face27YimuWang /ActivityNet Description Dataset V1-2 v1-2_train.tar.gz and v1-2_val.tar.gz Data (train and val set) associated with ActivityNet release 1.2 v1-2_test.tar.gz Data (test set only) associated with ActivityNet release 1.2 Dataset V1-3 v1-3_train_val.tar.gz Additional videos (train val set) collected for ActivityNet release 1.3 v1-3 is an extension of v1-2, so you also need to download v1-2 data and merge to v1.3 v1-3_test.tar.gz Additional videos (test set only)… See the full description on the dataset page: https://huggingface.co/datasets/YimuWang/ActivityNet.videon<1K18 likes2.5k downloads1y agoHugging Face28polymathic-ai /active_matter How To Load from HuggingFace Hub Be sure to have the_well installed (pip install the_well) Use the WellDataModule to retrieve data as follows: from the_well.benchmark.data import WellDataModule # The following line may take a couple of minutes to instantiate the datamodule datamodule = WellDataModule( "hf://datasets/polymathic-ai/", "active_matter_cloud_optimized", ) train_dataloader = datamodule.train_dataloader() for batch in dataloader: # Process training batch… See the full description on the dataset page: https://huggingface.co/datasets/polymathic-ai/active_matter.time-series-forecasting5 likes2.5k downloads2y agoHugging Face29RE3SIM /act-dataset0 likes2.4k downloads2y agoHugging Face30X-ActM /DVGT-Dataset3 likes2.4k downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.