procedural-robotics/factory-manipulation-videos
Factory manipulation videos Procedural Robotics is open sourcing a small set of our factory data so teams can assess its quality. The videos show workers performing factory tasks. Contents Seven continuous takes, 109 minutes in total. Task Station Worker Duration File cardboard manipulation 01 041 23.6 min cardboard_manipulation_station01_worker041.mp4 cardboard manipulation 04 026 16.5 min cardboard_manipulation_station04_worker026.mp4 defect… See the full description on the dataset page: https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos.
Factory manipulation videos
Procedural Robotics is open sourcing a small set of our factory data so teams can assess its quality. The videos show workers performing factory tasks.
Contents
Seven continuous takes, 109 minutes in total.
Format
MP4, H.264, 1280x720, 30 fps, no audio. First person view from a wearable camera on the factory floor. Each file is one uninterrupted recording of a single task.
Privacy
Faces were detected and blurred in every video. Audio was removed.
Loading
from datasets import load_dataset
ds = load_dataset("videofolder", data_dir="videos", split="train")videos/metadata.csv carries the task, station, worker, duration and resolution for each file.
Previews
Twenty seconds of each take with the hand tracks (left orange, right green), the table plane recovered from the AprilTags (grid, one tag side per cell), the tag map and each wrist's position on the table in tag units. Every frame's pose source is printed in the corner.
<div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/cardboardmanipulationstation01worker041.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/cardboardmanipulationstation01worker041.mp4" style="width:100%"></video><div style="font-size:12px">cardboard manipulation station01 worker041</div></div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/cardboardmanipulationstation04worker026.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/cardboardmanipulationstation04worker026.mp4" style="width:100%"></video><div style="font-size:12px">cardboard manipulation station04 worker026</div></div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/defecttestingstation08worker009.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/defecttestingstation08worker009.mp4" style="width:100%"></video><div style="font-size:12px">defect testing station08 worker009</div></div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/plasticassemblystation05worker032.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/plasticassemblystation05worker032.mp4" style="width:100%"></video><div style="font-size:12px">plastic assembly station05 worker032</div></div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/plushmanipulationstation05worker020.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/plushmanipulationstation05worker020.mp4" style="width:100%"></video><div style="font-size:12px">plush manipulation station05 worker020</div></div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/screwassemblystation01worker015.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/screwassemblystation01worker015.mp4" style="width:100%"></video><div style="font-size:12px">screw assembly station01 worker015</div></div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/screwassemblystation12worker003.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/screwassemblystation12worker003.mp4" style="width:100%"></video><div style="font-size:12px">screw assembly station12 worker003</div></div> </div>
Hand tracks
hands/<file>.hands.jsonl carries the worker's hands for every frame of videos/<file>: one row per (frame, hand) with a pixel bounding box, the 21 MediaPipe hand landmarks in image pixels (kpts, with the model's relative depth as the third value), the 21 metric world landmarks in metres relative to the wrist (world), a left/right label and a track id that stays constant while the same hand is in view. src is det for a detected frame and interp for a frame bridged by the tracker (gaps of up to 15 frames); interpolated rows carry score 0.
import json
rows = [json.loads(l) for l in open("hands/plastic_assembly_station05_worker032.hands.jsonl")]
rows[0]["hand"], rows[0]["track"], rows[0]["box"], len(rows[0]["kpts"])Tracking: detections are linked across frames by position with a left/right consistency cost; a hand must be seen three times before it becomes a track, and handedness is a score-weighted vote over the whole track. Precision over recall — frames where no confident hand was found have no row. hands/<file>.hands.json holds the same tracks in a compact per-frame form (fractions of the frame, 2-D) plus a per-track summary and the parameters used, and hands/<file>.hands.parquet is the same table typed for the dataset viewer and load_dataset(..., "hand_tracks").
Scene geometry
Each take has three AprilTag 36h11 markers lying flat on the table. scene/<file>.scene.jsonl gives, for every frame that sees a tag, the camera's pose in a table-fixed frame (cam.p position, cam.q rotation as an x, y, z, w quaternion; the frame is the reference tag's, z up out of the table, units = one tag side), how the pose was obtained (src: tags = two or more tags; track = carried by the table's own texture while the tags were covered or blurred, with inliers plane points and closed when the drift was corrected against the next tag sighting; interp = bridged across a gap of up to 3 s where tracking was lost; hold = held up to 1 s past the last sighting; single = one tag only), which tags were seen and the corner reprojection error, and for each tracked hand on that frame its wrist and index-tip points on the table plane (wrist_table, index_tip_table) and the hand's metric pose in the camera from the world landmarks (pose_cam_m: metres, x y z + quaternion). scene/<file>.scene.json holds the camera model (an Insta360 GO 3S whose FOV mode was not recorded; lens model, focal length and radial term are estimated per take by a held-out test, see camera.scan / camera.lens_scan), the tag map (with map.segments when a tag was bumped during the take), coverage, and two blind checks: coverage.tracking_error (the tracker run through footage where tags were available, compared with the tag pose at 1, 2 and 3 s) and coverage.bridging_error (the same for interpolation). Multiply tag-side units by the printed tag size to get metres. scene/<file>.scene.parquet is the frame table typed for the viewer and load_dataset(..., "scene"), with the camera pose flattened to cam_p / cam_q.
Intended use
A sample for evaluating data quality, not a training scale release. Contact us for the full dataset.
Contact
Procedural Robotics, proceduralrobotics.com.
