Team Ai
Datasetpublic

procedural-robotics/factory-manipulation-videos

Factory manipulation videos Procedural Robotics is open sourcing a small set of our factory data so teams can assess its quality. The videos show workers performing factory tasks. Contents Seven continuous takes, 109 minutes in total. Task Station Worker Duration File cardboard manipulation 01 041 23.6 min cardboard_manipulation_station01_worker041.mp4 cardboard manipulation 04 026 16.5 min cardboard_manipulation_station04_worker026.mp4 defect… See the full description on the dataset page: https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
2likes405downloads
Dataset Card

Factory manipulation videos

Procedural Robotics is open sourcing a small set of our factory data so teams can assess its quality. The videos show workers performing factory tasks.

[image]

Contents

Seven continuous takes, 109 minutes in total.

TaskStationWorkerDurationFile
cardboard manipulation0104123.6 mincardboard_manipulation_station01_worker041.mp4
cardboard manipulation0402616.5 mincardboard_manipulation_station04_worker026.mp4
defect testing0800914.3 mindefect_testing_station08_worker009.mp4
plastic assembly050327.7 minplastic_assembly_station05_worker032.mp4
plush manipulation0502013.3 minplush_manipulation_station05_worker020.mp4
screw assembly0101516.0 minscrew_assembly_station01_worker015.mp4
screw assembly1200317.9 minscrew_assembly_station12_worker003.mp4

Format

MP4, H.264, 1280x720, 30 fps, no audio. First person view from a wearable camera on the factory floor. Each file is one uninterrupted recording of a single task.

Privacy

Faces were detected and blurred in every video. Audio was removed.

Loading

python
from datasets import load_dataset
ds = load_dataset("videofolder", data_dir="videos", split="train")

videos/metadata.csv carries the task, station, worker, duration and resolution for each file.

Previews

Twenty seconds of each take with the hand tracks (left orange, right green), the table plane recovered from the AprilTags (grid, one tag side per cell), the tag map and each wrist's position on the table in tag units. Every frame's pose source is printed in the corner.

<div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/cardboardmanipulationstation01worker041.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/cardboardmanipulationstation01worker041.mp4" style="width:100%"></video><div style="font-size:12px">cardboard manipulation station01 worker041</div></div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/cardboardmanipulationstation04worker026.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/cardboardmanipulationstation04worker026.mp4" style="width:100%"></video><div style="font-size:12px">cardboard manipulation station04 worker026</div></div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/defecttestingstation08worker009.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/defecttestingstation08worker009.mp4" style="width:100%"></video><div style="font-size:12px">defect testing station08 worker009</div></div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/plasticassemblystation05worker032.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/plasticassemblystation05worker032.mp4" style="width:100%"></video><div style="font-size:12px">plastic assembly station05 worker032</div></div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/plushmanipulationstation05worker020.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/plushmanipulationstation05worker020.mp4" style="width:100%"></video><div style="font-size:12px">plush manipulation station05 worker020</div></div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/screwassemblystation01worker015.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/screwassemblystation01worker015.mp4" style="width:100%"></video><div style="font-size:12px">screw assembly station01 worker015</div></div> <div style="display:inline-block;width:49%;vertical-align:top;margin:0 0 12px 0"><video controls muted playsinline preload="none" poster="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/screwassemblystation12worker003.jpg" src="https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos/resolve/main/previews/screwassemblystation12worker003.mp4" style="width:100%"></video><div style="font-size:12px">screw assembly station12 worker003</div></div> </div>

Hand tracks

TakeHand tracksFrames with a handCamera posesLens (HFOV)
cardboard manipulation station01 worker04114792%99%116°
cardboard manipulation station04 worker02632698%96%115°
defect testing station08 worker00927946%97%106°
plastic assembly station05 worker03298100%100%127°
plush manipulation station05 worker02037189%96%97°
screw assembly station01 worker01523999%100%125°
screw assembly station12 worker003319100%99%103°

hands/<file>.hands.jsonl carries the worker's hands for every frame of videos/<file>: one row per (frame, hand) with a pixel bounding box, the 21 MediaPipe hand landmarks in image pixels (kpts, with the model's relative depth as the third value), the 21 metric world landmarks in metres relative to the wrist (world), a left/right label and a track id that stays constant while the same hand is in view. src is det for a detected frame and interp for a frame bridged by the tracker (gaps of up to 15 frames); interpolated rows carry score 0.

python
import json
rows = [json.loads(l) for l in open("hands/plastic_assembly_station05_worker032.hands.jsonl")]
rows[0]["hand"], rows[0]["track"], rows[0]["box"], len(rows[0]["kpts"])

Tracking: detections are linked across frames by position with a left/right consistency cost; a hand must be seen three times before it becomes a track, and handedness is a score-weighted vote over the whole track. Precision over recall — frames where no confident hand was found have no row. hands/<file>.hands.json holds the same tracks in a compact per-frame form (fractions of the frame, 2-D) plus a per-track summary and the parameters used, and hands/<file>.hands.parquet is the same table typed for the dataset viewer and load_dataset(..., "hand_tracks").

Scene geometry

Each take has three AprilTag 36h11 markers lying flat on the table. scene/<file>.scene.jsonl gives, for every frame that sees a tag, the camera's pose in a table-fixed frame (cam.p position, cam.q rotation as an x, y, z, w quaternion; the frame is the reference tag's, z up out of the table, units = one tag side), how the pose was obtained (src: tags = two or more tags; track = carried by the table's own texture while the tags were covered or blurred, with inliers plane points and closed when the drift was corrected against the next tag sighting; interp = bridged across a gap of up to 3 s where tracking was lost; hold = held up to 1 s past the last sighting; single = one tag only), which tags were seen and the corner reprojection error, and for each tracked hand on that frame its wrist and index-tip points on the table plane (wrist_table, index_tip_table) and the hand's metric pose in the camera from the world landmarks (pose_cam_m: metres, x y z + quaternion). scene/<file>.scene.json holds the camera model (an Insta360 GO 3S whose FOV mode was not recorded; lens model, focal length and radial term are estimated per take by a held-out test, see camera.scan / camera.lens_scan), the tag map (with map.segments when a tag was bumped during the take), coverage, and two blind checks: coverage.tracking_error (the tracker run through footage where tags were available, compared with the tag pose at 1, 2 and 3 s) and coverage.bridging_error (the same for interpolation). Multiply tag-side units by the printed tag size to get metres. scene/<file>.scene.parquet is the frame table typed for the viewer and load_dataset(..., "scene"), with the camera pose flattened to cam_p / cam_q.

Intended use

A sample for evaluating data quality, not a training scale release. Contact us for the full dataset.

Contact

Procedural Robotics, proceduralrobotics.com.