datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
actionbench
🎬 ActionBench: Paired Video-3D Synthetic Benchmark
📖 Overview
ActionBench is a benchmark dataset of 128 paired video ↔ animated point-cloud samples for evaluating animated 3D mesh generation from video.
The dataset consists of synthetic scenes of animated objects from ObjaverseXL, rendered using Blender 3.5.1.
Each sample contains:
Video: 16 RGBA frames with alpha mask
Camera (camera.json): Camera parameters using Blender convention (X_cam = X @ R^T + T, camera looks… See the full description on the dataset page: https://huggingface.co/datasets/facebook/actionbench.xiaoluo-gaming-action3000-20260910-media
Action-boundary review examples
Media for 300 selected examples from Xiaoluo (Cyberpunk 2077 and Rise of the Tomb Raider) and Gaming 500 Hours, 150 examples per dataset.
Includes 15-second review videos, observed boundary frames, and available action clips. These are visual model estimates; boundaries require human review. Source game and dataset rights remain with their respective owners.
Gallery and annotation manifests:… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/xiaoluo-gaming-action3000-20260910-media.ActiveArena-Data
ActiveArena-Data
Frozen LeRobot-format trajectories for the ActiveArena Astribot benchmark.
The repository contains 35 task directories, each with 100 training episodes
and the corresponding metadata files. The directory names match the
activearena_astribot registry in ActiveArena-VLA.
Use the files with the ActiveArena_Astribot_lerobot data root expected by the
released training configurations. This dataset snapshot is kept independent
of upstream simulator updates; the task… See the full description on the dataset page: https://huggingface.co/datasets/leeibo/ActiveArena-Data.truth-probe-activationsdrive-actionCollective-Activity-Recognition
Annotation Format
Every 10th frame in all video sequences was manually annotated with the following information for each detected person:
Bounding box location
Activity class
Pose direction
Annotation Fields
Each annotation follows the format:
<frame_number> <x> <y> <width> <height> <class_id> <pose_id>
Field
Description
frame_number
Frame identifier
x
X-coordinate of the bounding box (top-left corner)
y
Y-coordinate of the bounding box (top-left… See the full description on the dataset page: https://huggingface.co/datasets/litforth/Collective-Activity-Recognition.action-world-model-atlas-1500-media-20260914
Action World Model Atlas
Public browsing previews for 1,500 unique action clips from the completed
6,033-video bundle. OpenPixel2Play, Gaming 500 Hours, and Xiaoluo each contribute
500 examples. All 46 games in the completed bundle are represented.
Videos preserve the full five-second duration and 81 frames. They are existing
browser previews and can be smaller than the native training videos. Video and
poster checksums are verified against the source media manifests.… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/action-world-model-atlas-1500-media-20260914.ActiveVision
ActiveVision — An Exam for Active Observers
ActiveVision is a benchmark for iterative visual reasoning: 85 photorealistic
items across 17 tasks that cannot be solved from a single glance — the model has
to keep returning to the image to scan, trace, and compare. Every scene is
generated by a deterministic program and re-rendered photorealistically while
preserving the structure, so answers are exact by construction.
Frontier models reach about 10% with pure… See the full description on the dataset page: https://huggingface.co/datasets/activevisionai/ActiveVision.eu-ai-act-article-50-scoreboard
Article 50 historical public-evidence snapshot
This work was produced through an AI-assisted workflow directed by the author. Historical work used Anthropic assistance; the retrospective correction uses OpenAI GPT-6, with separate bounded Gemini advice. All three providers have products in the scored set.
Purpose: provide the corrected paper's version 1.1 bundle under v1_1. Start with its README and correction note. The paper and deposit and GitHub repository identify the same… See the full description on the dataset page: https://huggingface.co/datasets/NMAIResearch/eu-ai-act-article-50-scoreboard.Act2Cap_benchmarkCollected data from GUI-Action-Narrator
afrolm_active_learning_dataset
AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages
GitHub Repository of the Paper
This repository contains the dataset for our paper AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages which will appear at the third Simple and Efficient Natural Language Processing, at EMNLP 2022.
Our self-active learning framework
Languages Covered
AfroLM has been… See the full description on the dataset page: https://huggingface.co/datasets/bonadossou/afrolm_active_learning_dataset.Active-ReconstructionSAFER-Activities
SAFER-Activities
A Dataset for Smart Assessment of Fall Events and Routine Activities (ECCV 2026).
SAFER-Activities is a dataset for fall detection and physical activity monitoring from
video, with a dedicated subset for wheelchair users. It provides frame-level action
annotations (precise start/end of every action) over long, untrimmed, multi-camera
recordings.
File structure
.
├── raw/
│ ├── Safer-Activities-Full-Dataset/
│ │ ├── normal/… See the full description on the dataset page: https://huggingface.co/datasets/SAFER-Activities/SAFER-Activities.pythia-massive-activations
Hidden Dynamics of Massive Activations in Transformer Training
Dataset Description
This dataset contains comprehensive analysis data for the paper "Hidden Dynamics of Massive Activations in Transformer Training". It provides detailed measurements and mathematical characterizations of massive activation emergence patterns across the Pythia model family during training.
Massive activations are scalar values in transformer hidden states that achieve values orders of… See the full description on the dataset page: https://huggingface.co/datasets/Aimpoint-Digital/pythia-massive-activations.openp2p-action-clips-media-3000-20260911drive-actionfire_actioncam
Fire Actioncam
This dataset is a collection of several real-world fire scenes, introduced by the ECCV paper "Gaussians on Fire: High-Frequency Reconstruction of Flames".
Overview
The dataset consists of 17 real-world scenes of burning paper, cardboard, wood, gasoline, ethanol, and propane. We captured each scene with three regular actioncams, synchronizing them with µs precision using a custom LED pattern.
Property
Value
Scenes
17 (two outdoor… See the full description on the dataset page: https://huggingface.co/datasets/jna-358/fire_actioncam.robot-action-prediction-dataset
Robotic Action Prediction Dataset
Dataset Description
This dataset contains triplets of (current observation, action instruction, future observation) for training models to predict future frames of robotic actions.
Dataset Structure
Data Fields
current_frame: Input image (RGB) of the current observation
instruction: Textual description of the action to perform
future_frame: Target image (RGB) showing the expected outcome 50 frames later… See the full description on the dataset page: https://huggingface.co/datasets/bryandts/robot-action-prediction-dataset.Human_Action_Recognition
Dataset Summary
A dataset from kaggle. origin: https://dphi.tech/challenges/data-sprint-76-human-activity-recognition/233/data
Introduction
The dataset features 15 different classes of Human Activities.
The dataset contains about 12k+ labelled images including the validation images.
Each image has only one human activity category and are saved in separate folders of the labelled classes
PROBLEM STATEMENT
Human Action Recognition (HAR) aims to understand… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/Human_Action_Recognition.actionnet_3k_og
actionnet_3k_og — training subset, archived
Everything actionnet_grasp_2b_480 / _720 opens, plus the LeRobot data/ and meta/
that describe the same episodes: ~35.0 GB in 9 archives instead of ~14,900 loose files.
A 13th archive set, videos_15fps_768x432/, comes from a different store — read the
note below before using it.
This is still a subset of a larger tree. The source also holds videos/,
first_frames/, grasp_frames/, source_mask_merge/, rendering_videos/, gtdepth_s0/
and… See the full description on the dataset page: https://huggingface.co/datasets/seungkukim/actionnet_3k_og.atari_sft_8_AUG_shooting_sports_maze_actionours_joint_actionjam-actions-v0
Dataset Card for jam-actions-v0 (public subset)
Version: 0.6.0 — a correction release. It withdraws 58 records whose source arrangements could not be licence-cleared and changes no remaining record. See Version 0.6.0 correction.
Records built: 2026-07-11 (0.5.0 cut; unchanged) Package built: 2026-09-25
DOI: 10.5281/zenodo.22961580 (this version; concept DOI 10.5281/zenodo.22961579). Earlier versions: 0.5.0 10.5281/zenodo.21313954 and 0.4.3 10.5281/zenodo.20279919. Both contain… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v0.drive-action-subsetshotpath-action-diagnostic-venuslike-eval-20260709# ShotPath Action Diagnostic Venus-like Eval 20260709
This bundle contains the LLM-audited pure-operation diagnostic set for Venus-like evaluation.
Files:
action_diagnostic_pure_operation.jsonl: 908 examples after leakage audit.
images/: image files referenced by the jsonl.
scripts/eval_action_diagnostic_qwen25vl.py: Qwen2.5-VL base/LoRA evaluator.
scripts/run_action_diagnostic_venuslike_eval_server.sh: server runner for base7b, stage1 step200, stage2 step200.
Default server paths in the… See the full description on the dataset page: https://huggingface.co/datasets/purefall/shotpath-action-diagnostic-venuslike-eval-20260709.activity-diagrams-qdobr
Dataset Card for activity-diagrams-qdobr
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
activity-diagrams-qdobr
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/activity-diagrams-qdobr.dual_needle_concat_action_staticThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 250,
"total_frames": 97953,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:250"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/inaas/dual_needle_concat_action_static.video-web-actionthemoviedb_actorsdataset_30ep_actThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "tiago-pro",
"total_episodes": 30,
"total_frames": 11019,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/VivianaMorl26/dataset_30ep_act.
