datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ViFailback-Dataset
ViFailback Dataset: Real-World Robotic Manipulation Failure Dataset with Visual Symbol Guidance
A real-world dataset for diagnosing, correcting, and learning from robotic manipulation failures via visual symbols.
ViFailback is a large-scale, real-world robotic manipulation failure dataset introduced in the CVPR 2026 paper "Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols". It introduces visual… See the full description on the dataset page: https://huggingface.co/datasets/sii-rhos-ai/ViFailback-Dataset.Polymarket_data
Polymarket Data
Complete Data Infrastructure for Polymarket — Fetch, Process, Analyze
A comprehensive dataset of 6.6 billion on-chain trading records from Polymarket, processed into multiple analysis-ready formats. Features cleaned data, unified token perspectives, and user-level transformations — ready for market research, behavioral studies, and quantitative analysis.
Zhengjie Wang1,2, Leiyu Chao1,3, Yu Bao1,4, Lian Cheng1,3, Jianhan Liao1,5, Yikang Li1,†
1Shanghai Innovation… See the full description on the dataset page: https://huggingface.co/datasets/SII-WANGZJ/Polymarket_data.CiQi-VQA
CiQi-Agent
Github | Model | Dataset | Paper
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
Accepted to ECCV 2026
🎯 Overview
CiQi-Agent has been accepted to ECCV 2026.
We present CiQi-Agent, a domain-specific multimodal agent for antique Chinese porcelain connoisseurship. The project is designed to combine fine-grained visual perception, tool-augmented reasoning, and cultural-heritage knowledge… See the full description on the dataset page: https://huggingface.co/datasets/SII-Monument-Valley/CiQi-VQA.thinking_droid_lerobot_output_qwen3vlvifailback-dataset-lerobot
ViFailback Dataset — LeRobot
This repository is a LeRobot v2.1 conversion of the trajectory portion of sii-rhos-ai/ViFailback-Dataset, introduced in the CVPR 2026 paper Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols.
ViFailback contains real-world ALOHA dual-arm manipulation trajectories designed for studying failure diagnosis, failure localization, corrective guidance, recovery, and learning… See the full description on the dataset page: https://huggingface.co/datasets/sii-rhos-ai/vifailback-dataset-lerobot.constellaration-data
Dataset Card for ConStellaration
A dataset of diverse quasi-isodynamic (QI) stellarator boundary shapes with corresponding performance metrics and ideal magneto-hydrodynamic (MHD) equilibria, as well as settings for their generation.
The performance metrics and ideal MHD equilibria were evaluated under vacuum (default) and with plasma inside (finite beta).
Dataset Details
Dataset Description
Stellarators are magnetic confinement devices that are being… See the full description on the dataset page: https://huggingface.co/datasets/SII-LWY/constellaration-data.hang_cup
hang_cup
双臂 Flexiv 数据采集数据集,任务指令:Hang the cup on the rack。
Episodes: 30 (episode_0000–episode_0029)
Total frames: 15472
Recording rate: 30 Hz
Control rate: 100 Hz
Cameras: left 260322279602, right 260322271470, overview 254522075549
Each episode contains three camera image streams, synchronized frame index, robot state JSONL, joint-position CSVs, EEF-pose CSVs, and metadata.
Joint angles are recorded in radians in the observation/action CSVs; metadata.json start-position fields… See the full description on the dataset page: https://huggingface.co/datasets/ShareLab-SII/hang_cup.SII_self_evovling_02_training_datasetthinking_furniture_bench_dataset_lerobot_output_qwen3vlESOT500
ESOT500: A High-Frequency Dataset for Event-Driven Perception
Introduction
ESOT500 is a high-frequency annotated event-based single object tracking dataset. It was created to demonstrate the STARE (STream-based lAtency-awaRe Evaluation) framework, enabling rigorous assessment of event-driven perception models' real-time capabilities.
This dataset is introduced in the paper: Bridging the Latency Gap with a Continuous Stream Evaluation Framework in Event-Driven Perception.… See the full description on the dataset page: https://huggingface.co/datasets/sii-geai-lab/ESOT500.pick_place_bread
pick_place_bread
双臂 Flexiv 数据采集数据集,任务指令:Pick up the bread and place it on the plate。
Episodes: 30 (episode_0000–episode_0029)
Total frames: 14355
Recording rate: 30 Hz
Control rate: 100 Hz
Each episode contains three camera image streams, synchronized frame index, robot state JSONL, joint-position CSVs, EEF-pose CSVs, and metadata.
Joint angles are recorded in radians in observation/action CSVs; metadata start positions are in degrees.
Observation gripper is measured width in… See the full description on the dataset page: https://huggingface.co/datasets/ShareLab-SII/pick_place_bread.pnp
pnp
双臂 Flexiv 数据采集数据集,任务指令:Clean the table。
Episodes: 20 (episode_0000–episode_0019)
Total frames: 9394
Recording rate: 30 Hz
Control rate: 100 Hz
Cameras: left 260322279602, right 260322271470, overview 254522075549
Each episode contains three camera image streams, synchronized frame index, robot state JSONL, joint-position CSVs, EEF-pose CSVs, and metadata.
Joint angles are recorded in radians in the observation/action CSVs; metadata.json start-position fields are in degrees.… See the full description on the dataset page: https://huggingface.co/datasets/ShareLab-SII/pnp.SIIB-Time
1- Scope
The increasing penetration of inverter-based resources (IBRs), e.g, renewable and energy storage systems, is fundamentally reshaping power grid dynamics. Unlike conventional resources, IBRs interact with the grid through power electronics operating at microsecond timescales, introducing ultrafast dynamic phenomena that conventional time-domain simulation methods, e.g., RMS techniques, fail to capture [1]. Electromagnetic transient (EMT) simulations can capture these fast… See the full description on the dataset page: https://huggingface.co/datasets/neurips26-PSML/SIIB-Time.SIIMACRthinking_fractal20220817_data_lerobot_output_qwen3vlthinking_fmb_dataset_lerobot_output_qwen3vlwxf_taskSII-Stellarator-Configuration-Dataset
SII Stellarator Configuration Dataset
中文
Wenyang Li | PhD student, Shanghai Innovation Institute
AI-Driven Controlled Fusion Simulation, Control & Design Lab
Email: lwydsg@mail.nankai.edu.cn | Phone: +86-156-2003-5216. License: CC BY 4.0.
Overview
Stellarator boundary configurations and recorded physics evaluations, covering NFP 1-5.
This replaces the earlier 158,685-row release.
The original data came from proxima-fusion/constellaration.
New sources include… See the full description on the dataset page: https://huggingface.co/datasets/SII-AI4Fusion/SII-Stellarator-Configuration-Dataset.tactile_lightbulbthinking_bc_z_lerobot_output_qwen3vlMTG_MuCodec_Token_DatasetsLIBERO_plus_assetsVA-Judger-Bench
VA-Judger-Bench
VA-Judger-Bench is a paired audio/video preference benchmark with 1,150 cases:
easy: 400 cases
indomain: 250 cases
outdomain: 500 cases
Layout
Each split contains data.jsonl and a videos/ directory. All video paths are
relative to the split directory.
VA-Judger-Bench/
├── README.md
├── easy/
│ ├── data.jsonl
│ └── videos/
├── indomain/
│ ├── data.jsonl
│ └── videos/
└── outdomain/
├── data.jsonl
└── videos/
Record… See the full description on the dataset page: https://huggingface.co/datasets/ShareLab-SII/VA-Judger-Bench.CogStream
CogStream Dataset
Dataset for CogStream: Context-guided Streaming Video Question Answering.
Overview
CogStream is a streaming video QA dataset designed to evaluate context-guided video reasoning. Models must identify and utilize relevant historical context to answer questions about ongoing video streams.
Statistics:
Split
Videos
QA Pairs
Train
852
55,623
Test
236
15,364
Total
1,088
70,987
Sources: MovieChat (40.2%), MECD (16.8%), QVhighlights (9.8%)… See the full description on the dataset page: https://huggingface.co/datasets/SII-KYW/CogStream.bgg-bench-pdfsSocioBank
SocioBank
SocioBank brings together prepared training data and original source data for simulating individual behavior and social interactions. For details, see the Socio-Foundation paper.
Configurations prefixed with train_ contain prepared training subsets. Those prefixed with raw_ retain source fields and data partitions, with no additional deduplication, filtering for overlap with evaluation data, or sampling.
from datasets import load_dataset
train =… See the full description on the dataset page: https://huggingface.co/datasets/SII-LancelotXie/SocioBank.G1-WholeBody-PickPlaceToy
G1-WholeBody-PickPlaceToy
Teleoperated whole-body manipulation dataset for the Unitree G1 humanoid, recorded in
LeRobot format (codebase_version: v2.1).
The single task is "pick up the toy and place it on the chair", executed with
coordinated whole-body control (locomotion + upper-body manipulation).
Dataset at a glance
Episodes
128
Frames
107,006
Tasks
1 (pick up the toy and place it on the chair)
FPS
50
Robot
Unitree G1 (43-DOF whole-body… See the full description on the dataset page: https://huggingface.co/datasets/SII-ZhangYiFei/G1-WholeBody-PickPlaceToy.HumanoidArenaV3.1_SONIC40
HumanoidArena v3.1 — SONIC Semantic-40
This dataset contains seven HumanoidArena task datasets in LeRobot format.
Each task has 100 episodes sampled at 50 Hz. The robot-state vector has 64
dimensions and the SONIC semantic action has 40 dimensions. The seven tasks
contain 546,460 frames in total.
Task
Episodes
Frames
HOI_double_desk
100
102,620
HOI_football
100
56,326
HOI_pp_box
100
69,176
HSI_boxing
100
55,771
HSI_open_door
100
83,127
HSI_sit_sofa
100
70… See the full description on the dataset page: https://huggingface.co/datasets/SII-yizhi/HumanoidArenaV3.1_SONIC40.ReWeaver-GCD-TSDiagnosisArena
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
DiagnosisArena is a comprehensive and challenging medical benchmark designed to assess the diagnostic reasoning abilities of LLMs in clinical settings.
This benchmark consists of 915 pairs of segmented patient cases and corresponding diagnoses, spanning 28 medical specialties, deriving from clinical case reports published in 10 high-impact medical journals.
The experimental results indicate that even… See the full description on the dataset page: https://huggingface.co/datasets/SII-SPIRAL-MED/DiagnosisArena.
