robots
Datasets
All datasets matching “robots”no_robots
Dataset Card for No Robots 🙅♂️🤖
Look Ma, an instruction dataset that wasn't generated by GPTs!
Dataset Summary
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described in OpenAI's InstructGPT paper, and is comprised mostly of single-turn… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/no_robots.humanoid-robots-training-dataset
Dynamic Intelligence — Humanoid Robot Training Dataset
A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors.
The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.unilab-robotsumi-robots
umi-robots
Part of umi, an open web crawl published as Parquet. Before you use any of this, read the exclusion list at open-index/umi-meta and filter the rows it names. Published files are never rewritten, so the exclusion list is how a takedown reaches you, and applying it is a condition of using the data rather than a suggestion.
One row per robots.txt fetch: the host, when we asked, what the origin answered, the raw text if it served one, and the summary our parser read out… See the full description on the dataset page: https://huggingface.co/datasets/open-index/umi-robots.afvoices
📘 African Next Voices – Bambara (AfVoices)
The AfVoices dataset is the largest open corpus of spontaneous Bambara speech at its release in late 2025. It contains 423 hours of segmented audio and 612 hours of original raw recordings collected across southern Mali. Speech was recorded in natural, conversational settings and annotated using a semi-automated transcription pipeline combining ASR pre-labels and human corrections. We release all the data processing code on GitHub.… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/afvoices.soma-to-multiple-robots
Thank you, Lambda
We thank Lambda for the compute support behind this multi-robot motion-generation and simulation effort.
Please let us know if you have feedback and suggestion! chil@apocynthion.ai / lloyd@apocynthion.ai
SOMA to Multiple Robots
AlphaMotion is our self-developed cross-embodiment motion model. This repository collects robot-specific motion references generated from SOMA, visual comparisons, and saved downstream simulation evidence.… See the full description on the dataset page: https://huggingface.co/datasets/sajio/soma-to-multiple-robots.
