community-data
fractal20220817_data_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "google_robot",
"total_episodes": 87212,
"total_frames": 3786400,
"total_tasks": 599,
"total_videos": 87212,
"total_chunks": 88,
"chunks_size": 1000,
"fps": 3,
"splits": {
"train": "0:87212"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/fractal20220817_data_lerobot.community_dataset_v3
Lerobot Community Datasets v3 - A Cross-Embodiment Pretraining Dataset for Vision Language Action Models
A large-scale robotics dataset for vision-language-action learning, featuring 791 datasets across 46 robot types, enabling cross-embodiment pretraining for generalist robot policies.
Overview
This is a crowdsourced, open-source dataset compiled from 235 community contributors worldwide. Building upon the pretraining datasets used for SmolVLA, Community… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceVLA/community_dataset_v3.community_dataset_v3
Lerobot Community Datasets v3 - A Cross-Embodiment Pretraining Dataset for Vision Language Action Models
A large-scale robotics dataset for vision-language-action learning, featuring 791 datasets across 46 robot types, enabling cross-embodiment pretraining for generalist robot policies.
Overview
This is a crowdsourced, open-source dataset compiled from 235 community contributors worldwide. Building upon the pretraining datasets used for SmolVLA, Community… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/community_dataset_v3.EO-Data1.5M
🤖 EO-Data-1.5M
A Large-Scale Interleaved Vision-Text-Action Dataset for Embodied AI
The first large-scale interleaved embodied dataset emphasizing temporal dynamics and causal dependencies among vision, language, and action modalities.
📊 Dataset Overview
EO-Data-1.5M is a massive, high-quality multimodal embodied reasoning dataset designed for training generalist robot… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/EO-Data1.5M.quarel
Dataset Card for "quarel"
Dataset Summary
QuaRel is a crowdsourced dataset of 2771 multiple-choice story questions, including their logical forms.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 0.63 MB
Size of the generated dataset: 1.53 MB
Total amount of disk used: 2.17 MB
An example of 'train'… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/quarel.community_dataset_v1
Community Dataset v1
A large-scale community-contributed robotics dataset for vision-language-action learning, featuring 128 datasets from 55 contributors worldwide.
We used this dataset to pretrain SmolVLA. However, this is not a complete set, but the dataset that we selected using specific filters, like fps, min num of episodes, and some qualitative assessment of video qualities, using the https://huggingface.co/spaces/Beegbrain/FilterLeRobotData tool. We also manually curated the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceVLA/community_dataset_v1.
