datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SparseVideoNav
SparseVideoNav Datasets
This repository contains the real-world navigation datasets released with OpenDriveLab/SparseVideoNav:
BVN: Beyond-the-View Navigation.
IFN: Instruction-Following Navigation.
Project links:
Project page: https://opendrivelab.com/SparseVideoNav
GitHub: https://github.com/OpenDriveLab/SparseVideoNav
Paper: https://arxiv.org/abs/2602.05827
Dataset Summary
SparseVideoNav studies real-world vision-language navigation with sparse future… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/SparseVideoNav.GSA-PT-Qwen2-7B-Instruct-chunk16-data
GSA-PT-Qwen2-7B-Instruct-chunk16-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk16.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk16 — model trained on this dataset
SparseCraft-dataset
SparseCraft
[ECCV'24] SparseCraft: Few-Shot Neural Reconstruction through Stereopsis Guided Geometric Linearization
Project
DTU Dataset
We provide preprocessed DTU data and results for the tasks of novel view synthesis and surface reconstruction.
It contains the following directories:
sparsecraft_data
├── nvs # Novel View Synthesis task data and results
│ └── mvs_data
│ ├── scan103
│ ├── ...
│ └── results # Results for training using… See the full description on the dataset page: https://huggingface.co/datasets/maeyounes/SparseCraft-dataset.train_800_sparse__mask__overlay_a75__sim__all_cameras__staticThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__overlay_a75__sim__all_cameras__static.train_800_sparse__mask__blackout__sim__all_cameras__live__depthThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__blackout__sim__all_cameras__live__depth.train_800_sparse__no_mask__ur5eThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__no_mask__ur5e.train_800_sparse__bbox__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__bbox__separate_channel__sim__all_cameras__live.GSA-FT-Qwen2-7B-Instruct-chunk32-data
GSA-FT-Qwen2-7B-Instruct-chunk32-data
This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk32.
Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk32 — model trained on this dataset
train_800_sparse__no_maskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__no_mask.train_800_sparse__mask__overlay_a75__sim__all_cameras__live__ur5eThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__overlay_a75__sim__all_cameras__live__ur5e.train_800_sparse__mask__overlay_a75__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__overlay_a75__sim__all_cameras__live.train_800_sparse__mask__overlay_a75__sim__wrist_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__overlay_a75__sim__wrist_cameras__live.GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data
GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4 — model trained on this dataset
train_800_sparse__mask__overlay_a100__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__overlay_a100__sim__all_cameras__live.train_800_sparse__mask__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__separate_channel__sim__all_cameras__live.GSA-PT-Qwen2-7B-Instruct-chunk8-data
GSA-PT-Qwen2-7B-Instruct-chunk8-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8 — model trained on this dataset
train_800_sparse__mask__overlay_a75__sim__agentview_camera__staticThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__overlay_a75__sim__agentview_camera__static.train_800_sparse__mask__blur_a50__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__blur_a50__sim__all_cameras__live.train_800_sparse__bbox__overlay_a100__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__bbox__overlay_a100__sim__all_cameras__live.train_800_sparse__point__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__point__separate_channel__sim__all_cameras__live.train_800_sparse__mask__overlay_a75__sim__agentview_camera__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__overlay_a75__sim__agentview_camera__live.train_800_sparseThis dataset was created using LeRobot.
Tasks
jam
cereal
pear
sweet potato
scone
boxed food
can
hamburger
lemon
squash
cheese
cup
egg
ham
hot dog
apple
boxed drink
bread
candle
chicken breast
soap dispenser
jar
knife block
kettle
potato
basket
cake
orange
spice
spray
alcohol
blender jug
mushroom
salt and pepper shaker
tiered shelf
condiment
mango
pitcher
plant
stool
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse.train_400_sparse__no_maskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_400_sparse__no_mask.train_800_sparse__bbox__overlay_a75__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__bbox__overlay_a75__sim__all_cameras__live.train_800_sparse__bbox__blackout_a50__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__bbox__blackout_a50__sim__all_cameras__live.train_800_sparse__mask__blackout_a50__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__blackout_a50__sim__all_cameras__live.train_800_sparse__point__overlay_a25__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__point__overlay_a25__sim__all_cameras__live.train_800_sparse__bbox__overlay_a50__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__bbox__overlay_a50__sim__all_cameras__live.train_800_sparse__bbox__blur__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__bbox__blur__sim__all_cameras__live.train_800_sparse__mask__overlay_a25__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__overlay_a25__sim__all_cameras__live.
