nvidia/PhysicalAI-Robotics-Manipulation-Kitchen
PhysicalAI Robotics Manipulation in the Kitchen Dataset Description: PhysicalAI-Robotics-Manipulation-Kitchen is a dataset of automatic generated motions of robots performing operations such as opening and closing cabinets, drawers, dishwashers and fridges. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Kitchen.
14828
1---2license: cc-by-4.03task_categories:4- robotics5tags:6- robotics7---8 9# PhysicalAI Robotics Manipulation in the Kitchen10 11## Dataset Description:12 13PhysicalAI-Robotics-Manipulation-Kitchen is a dataset of automatic generated motions of robots performing operations such as opening and closing cabinets, drawers, dishwashers and fridges. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator built with Kinova Gen3 arms. The environments are kitchen scenes where the furniture and appliances were procedurally generated [2].14This dataset is available for commercial use.15 16 17## Dataset Contact(s)18Fabio Ramos (ftozetoramos@nvidia.com) <br>19Anqi Li (anqil@nvidia.com)20 21## Dataset Creation Date2203/18/202523 24## License/Terms of Use25cc-by-4.026 27## Intended Usage28This dataset is provided in LeRobot format and is intended for training robot policies and foundation models.29 30## Dataset Characterization31* Data Collection Method<br>32 * Automated <br>33 * Automatic/Sensors <br>34 * Synthetic <br>35 36* Labeling Method<br>37 * Synthetic <br>38 39## Dataset Format40Within the collection, there are eight datasets in LeRobot format `open_cabinet`, `close_cabinet`, `open_dishwasher`, `close_dishwasher`, `open_fridge`, close_fridge`, `open_drawer` and `close_drawer`. 41* `open cabinet`: The robot opens a cabinet in the kitchen. <br>42* `close cabinet`: The robot closes the door of a cabinet in the kitchen. <br>43* `open dishwasher`: The robot opens the door of a dishwasher in the kitchen. <br>44* `close dishwasher`: The robot closes the door of a dishwasher in the kitchen. <br>45* `open fridge`: The robot opens the fridge door. <br>46* `close fridge`: The robot closes the fridge door. <br>47* `open drawer`: The robot opens a drawer in the kitchen. <br>48* `close drawer`: The robot closes a drawer in the kitchen. <br>49 50The videos below illustrate three examples of the tasks: 51 52<div style="display: flex; justify-content: flex-start;">53<img src="./assets/episode_000009.gif" width="300" height="300" alt="open_dishwasher" />54<img src="./assets/episode_000008.gif" width="300" height="300" alt="open_cabinet" />55<img src="./assets/episode_000029.gif" width="300" height="300" alt="open_fridge" />56</div>57 58* action modality: 34D which includes joint states for the two arms, gripper joints, pan and tilt joints, torso joint, and front and back wheels.59* observation modalities60 * observation.state: 13D where the first 12D are the vectorized transform matrix of the "object of interest". The 13th entry is the joint value for the articulated object of interest (i.e. drawer, cabinet, etc).61 * observation.image.world__world_camera: 512x512 images of RGB, depth and semantic segmentation renderings stored as mp4 videos.62 * observation.image.external_camera: 512x512 images of RGB, depth and semantic segmentation renderings stored as mp4 videos.63 * observation.image.world__robot__right_arm_camera_color_frame__right_hand_camera: 512x512 images of RGB, depth and semantic segmentation renderings stored as mp4 videos.64 * observation.image.world__robot__left_arm_camera_color_frame__left_hand_camera: 512x512 images of RGB, depth and semantic segmentation renderings stored as mp4 videos.65 * observation.image.world__robot__camera_link__head_camera: 512x512 images of RGB, depth and semantic segmentation renderings stored as mp4 videos.66 67 68The videos below illustrate three of the cameras used in the dataset. 69<div style="display: flex; justify-content: flex-start;">70<img src="./assets/episode_000004_world.gif" width="300" height="300" alt="world" />71<img src="./assets/episode_000004.gif" width="300" height="300" alt="head" />72<img src="./assets/episode_000004_wrist.gif" width="300" height="300" alt="wrist" />73</div>74 75 76## Dataset Quantification77Record Count:78* `open_cabinet`79 * number of episodes: 7880 * number of frames: 3929281 * number of videos: 1170 (390 RGB videos, 390 depth videos, 390 semantic segmentation videos)82* `close_cabinet`83 * number of episodes: 20584 * number of frames: 9955585 * number of videos: 3075 (1025 RGB videos, 1025 depth videos, 1025 semantic segmentation videos)86* `open_dishwasher`87 * number of episodes: 7288 * number of frames: 2812389 * number of videos: 1080 (360 RGB videos, 360 depth videos, 360 semantic segmentation videos)90* `close_dishwasher`91 * number of episodes: 7492 * number of frames: 3607893 * number of videos: 1110 (370 RGB videos, 370 depth videos, 370 semantic segmentation videos)94* `open_fridge`95 * number of episodes: 19396 * number of frames: 9385497 * number of videos: 2895 (965 RGB videos, 965 depth videos, 965 semantic segmentation videos)98* `close_fridge`99 * number of episodes: 76100 * number of frames: 41894101 * number of videos: 1140 (380 RGB videos, 380 depth videos, 380 semantic segmentation videos)102* `open_drawer`103 * number of episodes: 99104 * number of frames: 37214105 * number of videos: 1485 (495 RGB videos, 495 depth videos, 495 semantic segmentation videos)106* `close_drawer`107 * number of episodes: 77108 * number of frames: 28998109 * number of videos: 1155 (385 RGB videos, 385 depth videos, 385 semantic segmentation videos)110 111<!-- Total = 1.1GB + 2.7G + 795M + 1.1G + 2.6G + 1.2G + 1.1G + 826M -->112 113Total storage: 11.4 GB114 115 116## Reference(s)117```118[1] @inproceedings{garrett2020pddlstream,119 title={Pddlstream: Integrating symbolic planners and blackbox samplers via optimistic adaptive planning},120 author={Garrett, Caelan Reed and Lozano-P{\'e}rez, Tom{\'a}s and Kaelbling, Leslie Pack},121 booktitle={Proceedings of the international conference on automated planning and scheduling},122 volume={30},123 pages={440--448},124 year={2020}125}126 127[2] @article{Eppner2024,128 title = {scene_synthesizer: A Python Library for Procedural Scene Generation in Robot Manipulation},129 author = {Clemens Eppner and Adithyavairavan Murali and Caelan Garrett and Rowland O'Flaherty and Tucker Hermans and Wei Yang and Dieter Fox},130 journal = {Journal of Open Source Software}131 publisher = {The Open Journal},132 year = {2024},133 Note = {\url{https://scene-synthesizer.github.io/}}134}135 136[3] @inproceedings{curobo_icra23,137 author={Sundaralingam, Balakumar and Hari, Siva Kumar Sastry and138 Fishman, Adam and Garrett, Caelan and Van Wyk, Karl and Blukis, Valts and139 Millane, Alexander and Oleynikova, Helen and Handa, Ankur and140 Ramos, Fabio and Ratliff, Nathan and Fox, Dieter},141 booktitle={2023 IEEE International Conference on Robotics and Automation (ICRA)},142 title={CuRobo: Parallelized Collision-Free Robot Motion Generation},143 year={2023},144 volume={},145 number={},146 pages={8112-8119},147 doi={10.1109/ICRA48891.2023.10160765}148}149 150```151 152## Ethical Considerations153NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse. 154 155Please report security vulnerabilities or NVIDIA AI Concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).