medium
Datasets
All datasets matching “medium”dcvlm_pool_medium
DCVLM-Pool (medium)
The raw candidate pool at the medium scale of our DataComp-VLM
benchmark: 483,576,747 samples / 41.1 TB across 166 source datasets, as
WebDataset tar shards — ≈4× the small pool.
This pool is unfiltered and unmixed. It is the input to a data-curation experiment, not a training set. You choose
the filters and the mixing ratios, and create another training set. If you instead want a
ready-to-train dataset, use dcvlm-baseline-200b
(our reference SoTA… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_medium.mediumcurr. size: 53,081 videos
goal (todo): 100,000+
xarm_lift_mediumThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 800,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/xarm_lift_medium.libero-long-kinex-v0.8.0-astra-medium-seed0-20260927
RLE-Bench · Kinex
The current selection contains 32 tasks across five environments, including two additional hard RoboDojo candidates, with deposit_coin and fasten_screws retained and push_T retired.
Candidate analysis · Machine-readable assessment · Execution audit
Environment
Selected tasks
Native successes
Task01 · RoboCasa
8
1
Task02 · LIBERO
2
2
Task03 · RoboTwin
4
4
Task04 · RoboDojo
10
7
Task05 · BEHAVIOR
8
1
Retain play_stacking_toy and… See the full description on the dataset page: https://huggingface.co/datasets/RLE-Bench/libero-long-kinex-v0.8.0-astra-medium-seed0-20260927.xarm_lift_medium_replayThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 800,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/xarm_lift_medium_replay.roboballet-task07-kinex-astra-medium-seed0-20261005
One agent. Eight arms. Forty targets.
一个智能体如何协调八个机械臂?我们在相同场景中比较基础提示、并行提示和显式调度三种方式。模型为 gpt-6-astra,推理强度为 medium;每组一次零样本运行,种子为 0。
显式调度组在 31.5 仿真秒内完成全部 40 个目标、停留和归位;基础提示组完成 39/40,并行提示组完成 40/40 但未完成归位。
观看视频与结果 · 完整指令与实验记录
本任务基于 RoboBallet 公开资源改编,关注多臂到点与协调规划。每种条件仅运行一次;仿真时间不包含模型思考和规划等待。
