datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agentic_vbench_video_repairself_repair_gripper_dagger
self_repair_gripper_dagger
Robot self-repair, DAgger rollouts with operator corrections on the same task as self_repair_gripper_bc.
Real-robot bimanual manipulation data collected on a YAM arm pair, released as part
of the Flex-π project. Stored in LeRobot v2.1 format with synchronized RGB and
metric depth from three cameras.
At a glance
Episodes
2,154
Frames
609,385
Duration
~5.6 h @ 30 fps
Tasks
1
Robot
yam (bimanual)
Cameras
cam_high… See the full description on the dataset page: https://huggingface.co/datasets/flex-pi/self_repair_gripper_dagger.imagenet-latents-imagesimport os
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
from datasets import load_dataset
dataset = load_dataset("G-REPA/imagenet-latents-images", split="train")
lca-ci-builds-repair
🏟️ Long Code Arena (CI builds repair)
This is the benchmark for CI builds repair task as part of the
🏟️ Long Code Arena benchmark.
🛠️ Task. Given the logs of a failed GitHub Actions workflow and the corresponding repository snapshot,
repair the repository contents in order to make the workflow pass.
All the data is collected from repositories published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause, and BSD-2-Clause). The datapoints can be removed upon request.
To… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-ci-builds-repair.self-repair-gripper-v2.3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "yam_bimanual",
"total_episodes": 97,
"total_frames": 157583,
"total_tasks": 1,
"total_videos": 291,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:97"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YOLO2431/self-repair-gripper-v2.3.imagenet-latents-invae-f16d32import os
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
from datasets import load_dataset
dataset = load_dataset("G-REPA/imagenet-latents-invae-f16d32", split="train")
self-repair-gripper-v2.4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "yam_bimanual",
"total_episodes": 97,
"total_frames": 157583,
"total_tasks": 1,
"total_videos": 291,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:97"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YOLO2431/self-repair-gripper-v2.4.self_repair_gripper_bc
self_repair_gripper_bc
Robot self-repair, human teleoperation (BC): install a gripper into an empty holder, drive a screw with a screwdriver, then clear the table.
Real-robot bimanual manipulation data collected on a YAM arm pair, released as part
of the Flex-π project. Stored in LeRobot v2.1 format with synchronized RGB and
metric depth from three cameras.
At a glance
Episodes
802
Frames
1,278,804
Duration
~11.8 h @ 30 fps
Tasks
1
Robot
yam… See the full description on the dataset page: https://huggingface.co/datasets/flex-pi/self_repair_gripper_bc.imagenet-latents-vavae-f16d32import os
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
from datasets import load_dataset
dataset = load_dataset("G-REPA/imagenet-latents-vavae-f16d32", split="train")
Repackage_diverseself-repair-gripper-v2.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "yam_bimanual",
"total_episodes": 97,
"total_frames": 157583,
"total_tasks": 1,
"total_videos": 291,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:97"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YOLO2431/self-repair-gripper-v2.1.ex-repairsemantic-repair-routing
semantic-repair-routing
The supervised pairs that train
SemanticRepair-270M:
a message somebody actually wrote, and the requests inside it restated
plainly, one per line. 84,819 pairs in five languages, plus 2,515 in
Italian and English aimed at what the model used to refuse.
It teaches one narrow thing. An embedding router compares a question with
the description of every capability it can reach. People do not write the
way capabilities are described — they hedge, they… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/semantic-repair-routing.repa-imagenet-256Preprocessed imagenet dataset, compatible with repa repos (repos that use the same preprocessing/dataloader as repa)
just unzip the dataset.zip and you have both the resized images and preencoded latents
dataset.zip: resized&cropped images at 256x256, and sd-vae latents
latents.zip: just sd-vae latents
vae-in.zip: E2E-INVAE latents
images.zip: just resized&cropped 256x256 images
Repaired_videosci-repair-bench
CI-REPAIR-BENCH
Overview
CI-REPAIR-BENCH is a benchmark dataset for research on Continuous Integration (CI) failures and automated repair in Python repositories.
The dataset contains 567 CI failure instances collected from 105 real-world GitHub repositories, all written in Python.Each instance captures a CI workflow failure, its logs, the corresponding code diff, and repository-level metadata.
Dataset Statistics
Programming language: Python
Number… See the full description on the dataset page: https://huggingface.co/datasets/ci-benchmark-user/ci-repair-bench.Application_of_DS-REPaLimagenet-latents-davae-vavae-align1.5-400kimport os
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
from datasets import load_dataset
dataset = load_dataset("G-REPA/imagenet-latents-davae-vavae-align1.5-400k", split="train")
self-repair-gripper-v2.2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "yam_bimanual",
"total_episodes": 97,
"total_frames": 157583,
"total_tasks": 1,
"total_videos": 291,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:97"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YOLO2431/self-repair-gripper-v2.2.vbench-repair
VBench Repair
A dimension-first archive of VBench 1.0 original sampled videos, the 16 original
human-preference annotation files, and versioned counterfactual experiments.
Publication has been verified against the pinned data revision in provenance/dataset-verification.json. The archive contains 27,720 dimension-local original entries (19,400 distinct source paths), 6,930 annotation rows, 41,580 normalized unordered comparisons, and 27 versioned counterfactual experiment… See the full description on the dataset page: https://huggingface.co/datasets/xju-arlab/vbench-repair.SWE-universe-repaired-bug-pilot-trajectories
SWE-universe repaired BugPilot trajectories
Combined trajectory artifacts for the Qwen3.6 + mini-swe-agent evaluation of VmaxRL/SWEUniverse-Repaired-Bugpilot.
This dataset contains one row per evaluated task in metadata.jsonl, plus per-task files under trajectories//. The combined set uses the main full eval and replaces the two original infra-failure rows with the clean infra rerun trajectories.
Summary:
rows: 804
effective attempts: 804
passes: 629
pass rate: 0.782338
infra… See the full description on the dataset page: https://huggingface.co/datasets/VmaxRL/SWE-universe-repaired-bug-pilot-trajectories.python-program-repair-training-pool
Python program-repair training pool
A pool of public data for training a model to repair broken Python. Every row is a
program that does the wrong thing and the program that replaces it. It is a straight
collection of open datasets plus a rule-generated layer built from open functions, not a
new corpus: every row comes from one of the sources below, at the revision named, and
every row was put through an overlap filter against held-out material this pool is kept
separate from.… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-program-repair-training-pool.Dataself-repair-gripper-dagger-r2-v1.1data-pipeline-repair-trajectories
Data Pipeline Repair Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/data-pipeline-repair-trajectories.self-repair-gripper-v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "yam_bimanual",
"total_episodes": 97,
"total_frames": 157583,
"total_tasks": 1,
"total_videos": 291,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:97"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YOLO2431/self-repair-gripper-v2.self-repair-gripper-dagger-r2-v1so101-table-cleanup-v21-repaired1-engineering-repair-50-commercial
Industrial Mechanical & Electrical Components Dataset — 50 Authentic Details
📋 Description
Curated collection of 50 high-resolution photographs documenting authentic industrial mechanical and electrical components — from electric motors and solenoid valves to circuit boards, wiring harnesses, ball bearings, gearboxes, caster wheels, and battery packs.
Each image includes comprehensive CSV metadata with 13 classification fields optimized for machine learning… See the full description on the dataset page: https://huggingface.co/datasets/Kos1976/1-engineering-repair-50-commercial.Sycophancy-Repair-Guide
AI修复指南:基于框架的AI讨好行为修复方案v1.0
本修复指南是《AI对话讨好行为检测框架》的配套修复方案。框架负责发现问题,修复指南负责提供修正方案。
核心思路
AI讨好行为在检测框架中分为两类:零分禁环(主干层面的行为越界)和一级词包(修饰语层面的社交冗余)。修复方案采用后处理流水线,对AI生成的回答进行分层扫描和修正,在不改动模型本身的条件下降低讨好得分。
项目结构
修复指南由五份独立文档组成,按推荐的阅读顺序排列:
文档
内容
01-双模式分流与场景识别
脆弱信号词表、触发规则、安抚模式与专业模式的判定逻辑
02-缩句分层法
主干与修饰语的区分标准、分层扫描操作步骤
03-事实与观点分叉规则
客观事实与主观观点的判定标准、“我认为”类表述的处理逻辑
04-安抚模式下的原话反射规则
禁环12防护、虚构立场判定、强制重写规则、不作为防护
05-后处理流程与验证指南
流水线落地步骤、环境依赖、推荐工具、修复效果验证方法… See the full description on the dataset page: https://huggingface.co/datasets/luna-pragma-2026/Sycophancy-Repair-Guide.
