beta
Datasets
All datasets matching “beta”AgiBotWorld-Beta
Key Features 🔑
1 million+ trajectories from 100 robots, with a total duration of 2976.4 hours.
100+ real-world scenarios across 5 target domains.
Cutting-edge hardware: visual tactile sensors / 6-DoF dexterous hand / mobile dual-arm robots
200+ types of tasks:
Contact-rich manipulation
Long-horizon planning
Multi-robot collaboration
87 types of Atomic Skills, including Tie, OpenJar, Peel, Sweep etc.
Your… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/AgiBotWorld-Beta.Beta-Pre-Train-Corpus
Reactive AI / Beta Pre-Train Corpus
Pre-training corpus for RxT-Beta models, created from public & open datasets. Includes high-quality english and polish web crawl data, mathematic and scientific subsets,
and code in different programming languages.
2k subsets are filtered for 1024-2048 tokens, except MegaMath Web Pro and GitHub Code subsets, that were filtered for 512-2048 tokens
Subsets & original datasets
FineWeb-Edu
fineweb-edu-s100 (51.3M examples) - 50% of… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/Beta-Pre-Train-Corpus.Beta-Hybrid-Interaction-SFTjimei-fire-smoke-yolo-datasetagibotworld-beta-rgb-lance
AgiBotWorld-Beta (LeRobot lance format)
agibot-world/AgiBotWorld-Beta
converted to the LeRobot lance storage format, uploaded in coordination with the AgiBot team.
160,454 episodes, 286,556,463 frames, 8 RGB cameras (AV1, copied from the source without re-encoding),
state[20] and action[22] at 30 fps.
Read it in place, no download needed (lerobot with the lancedb extra):
from lerobot.datasets import LeRobotDataset
ds = LeRobotDataset("lance-format/agibotworld-beta-rgb-lance")… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/agibotworld-beta-rgb-lance.Updesh_beta
📢 Updesh: Synthetic Multilingual Instruction Tuning Dataset for 13 Indic Languages
NOTE: This is an initial $\beta$-release. We plan to release subsequent versions of Updesh with expanded coverage and enhanced quality control. Future iterations will include larger datasets, improved filtering pipelines.
Updesh is a large-scale synthetic dataset designed to advance post-training of LLMs for Indic languages. It integrates translated reasoning data and synthesized open-domain… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Updesh_beta.
