datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
a-share-l2-market-depth
China A-share Level 2 Market Depth
Canonical order-event and ten-level snapshot data for China A-shares. Canonical
trade records remain in the separate phields/a-share-l2-trades dataset.
Coverage
Date range: 2026-07-24 to 2026-07-24
Trading days: 1
Table
Rows
Parquet files
Compressed size
l2_orders
249,705,486
10
2.14 GiB
l2_snapshots
20,279,887
4
0.91 GiB
Layout… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-market-depth.depth_snapshoteval_pi0_bowl_plate_d435i_5ep_depth_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 75,
"total_frames": 44465,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:75"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tersooawai/eval_pi0_bowl_plate_d435i_5ep_depth_dataset.polymarket-btc-5m-l2-depth
Polymarket BTC 5m L2 Order Book, One Full Day
Every "Bitcoin Up or Down" 5 minute market on Polymarket for 2026-06-15 UTC, with
the full limit order book behind each one: 288 markets, 3,889 book snapshots,
41,591,670 level updates and 680,994 trades, on one millisecond time base.
Polymarket's public API serves the book as it is now. There is no endpoint that
returns the book as it was, so past depth is something you have to have been
recording while the market was live. This is… See the full description on the dataset page: https://huggingface.co/datasets/marketlens/polymarket-btc-5m-l2-depth.Depth-Normal-Videos-42K
Depth and Normal Videos Dataset
42,498 videos with depth and surface normals.
Usage
from huggingface_hub import hf_hub_download
video = hf_hub_download(
repo_id="Yanbin99/Depth-Normal-Videos-42K",
filename="Depth_and_Normal_42K/group_0000/videos/00000000.mp4",
repo_type="dataset"
)
Crypto-L2-Orderbook-Depth-24-Instruments
license: cc-by-4.0
task_categories:
time-series-forecasting
reinforcement-learning
📈 Hyperliquid Crypto L2 Orderbook Depth - 24 Instruments (Free Sample)
⚠️ NOTE: This is a 7-day FREE SAMPLE dataset. > For the full institutional-grade dataset featuring 12+ months of continuous history (~2.3 Million rows) for 24 crypto assets, visit ImbalanceLabs.com.
Dataset Summary
Standard OHLCV candles are a "liquidity illusion." They hide the true market intent, spread, spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AdamAtractor/Crypto-L2-Orderbook-Depth-24-Instruments.depth_data_set_v02This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 70,
"total_frames": 42348,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:70"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tersooawai/depth_data_set_v02.0825_put_the_ambrosial_yogurt_bottle_into_the_box_no_depth_sort_cleanThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "unknown",
"total_episodes": 106,
"total_frames": 36686,
"total_tasks": 1,
"total_videos": 212,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:106"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Deason11/0825_put_the_ambrosial_yogurt_bottle_into_the_box_no_depth_sort_clean.depth-trades-cryptodatasets
Binance Spot Microstructure
币安现货逐笔成交 + 订单簿快照,时间统一为交易所时间(不含本地时间)。
多交易对 / 多日期持续累积。
交易对
类型
日期
行数
文件数
ADAUSDT
trades
20261005
404
2
ADAUSDT
depth
20261005
330
2
BNBUSDT
trades
20261005
29
1
BNBUSDT
depth
20261005
104
1
BTCUSDT
trades
20261005
244,170
8
BTCUSDT
depth
20261005
69,714
7
DOGEUSDT
trades
20261005
840
3
DOGEUSDT
depth
20261005
1,132
3
ETHUSDT
trades
20261005
5,029
1
ETHUSDT
depth
20261005
1,437
1
SOLUSDT
trades
20261005
22
2
SOLUSDT
depth… See the full description on the dataset page: https://huggingface.co/datasets/CT-666/depth-trades-cryptodatasets.DepthBench-FineWeb-Edu-100BT-tokenized
DepthBench FineWeb-Edu 100BT Tokenized
This repository contains the tokenized FineWeb-Edu 100BT sample used by
DepthBench pretraining experiments.
Splits
train/: 139 shards, 99,585,913,529 tokens, and 97,045,608 documents.
eval/: 013_00008, containing 234,993,701 tokens and 225,078 documents.
All remaining source shards are assigned to training. Each source document is
terminated by an EOS token before documents are concatenated.
Format
Each shard… See the full description on the dataset page: https://huggingface.co/datasets/YangyiH/DepthBench-FineWeb-Edu-100BT-tokenized.depth_data_teleoperationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 20,
"total_frames": 12156,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tersooawai/depth_data_teleoperation.towel_half_fold_bimanual_depth_viz_20260911_114651This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_roll.pos",
"left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Woohi123/towel_half_fold_bimanual_depth_viz_20260911_114651.so101_stack_3_cubes_top_depthThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so-101",
"total_episodes": 60,
"total_frames": 30424,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SalZa2004/so101_stack_3_cubes_top_depth.lerobot_depth_libero_90lerobot_depth_libero_spatialOrdering_Constrained_No_Depth_50_demosThis dataset was created using LeRobot.
Dataset Description
This is a 50-episode subset of
justintiensmith/Ordering_Constrained_No_Depth.
It contains 12, 13, 12, and 13 demonstrations for the first, second, third, and fourth
positions from the left, respectively. Episodes were selected evenly across each source
25-episode group, then reindexed to 0–49. All five camera videos were trimmed to include
only the selected demonstrations.
Selected source episode IDs:
First… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Ordering_Constrained_No_Depth_50_demos.so101_track_rails_depthThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so-101",
"total_episodes": 41,
"total_frames": 42560,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:41"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SalZa2004/so101_track_rails_depth.so101_stack_cubes_top_depthThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so-101",
"total_episodes": 41,
"total_frames": 25981,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:41"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SalZa2004/so101_stack_cubes_top_depth.tinygsm_fobinary_workspace_depth1to9_traindepth5omy_f3m_right_spacemouse_LG_pick_align_insert_depth_cable1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 12937,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AivexRoboticsGroup/omy_f3m_right_spacemouse_LG_pick_align_insert_depth_cable1.train_800_sparse__mask__blackout__sim__all_cameras__live__depthThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__blackout__sim__all_cameras__live__depth.train_800_complex__mask__blackout__sim__all_cameras__live__depthThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_complex__mask__blackout__sim__all_cameras__live__depth.libero_spatial_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 10,
"total_frames": 1271,
"total_tasks": 10,
"total_videos": 80,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_spatial_mask_depth_IPEC_COMMUNITY_format.studytable_open_drawer_depth_1750444167This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 22079,
"total_tasks": 1,
"total_videos": 300,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/als2025ohf/studytable_open_drawer_depth_1750444167.so101_stack_cubes_depthThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so-101",
"total_episodes": 10,
"total_frames": 7022,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SalZa2004/so101_stack_cubes_depth.train_800_dense__mask__blackout__sim__all_cameras__live__depthThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_dense__mask__blackout__sim__all_cameras__live__depth.libero_spatial_lerobot_mask_depthThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 500,
"total_frames": 62250,
"total_tasks": 10,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Heimrih/libero_spatial_lerobot_mask_depth.libero_object_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 500,
"total_frames": 74507,
"total_tasks": 10,
"total_videos": 4000,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_object_mask_depth_IPEC_COMMUNITY_format.biyam_depth_15fps_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_roll.pos",
"left_wrist_yaw.pos"… See the full description on the dataset page: https://huggingface.co/datasets/pravsels/biyam_depth_15fps_test.keys_into_bowl_ob15_depth_rgbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"names": [
"left_arm_shoulder_pan.pos",
"left_arm_shoulder_lift.pos",
"left_arm_elbow_flex.pos",
"left_arm_wrist_flex.pos",
"left_arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Odog16/keys_into_bowl_ob15_depth_rgb.
