Team Ai
Datasetpublic

yxma/gelsight-mini-pretrain

GelSight Mini Pretrain ~853K GelSight Mini tactile RGB frames, 12 public sources, one parquet schema. Built for self-supervised representation learning (VAE / MAE / SimCLR / DINO) — every frame contact-filtered, channel-normalized, and re-encoded as JPEG q92. Frames Sources Real 536K FoTA (labeled+unlabeled), 3DCal, FEATS, GelSLAM, TactileTracking, RTM, FeelAnyForce, UniT, TacQuad Sim 317K sim_tactile_mnist, sim_starstruck (Taxim-rendered, Mini-calibrated) NC… See the full description on the dataset page: https://huggingface.co/datasets/yxma/gelsight-mini-pretrain.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes984downloads
Dataset Card

GelSight Mini Pretrain

~853K [GelSight Mini](https://www.gelsight.com/gelsightmini/) tactile RGB frames, 12 public sources, one parquet schema. Built for self-supervised representation learning (VAE / MAE / SimCLR / DINO) — every frame contact-filtered, channel-normalized, and re-encoded as JPEG q92.

[image]

FramesSources
Real536KFoTA (labeled+unlabeled), 3DCal, FEATS, GelSLAM, TactileTracking, RTM, FeelAnyForce, UniT, TacQuad
Sim317Ksimtactilemnist, sim_starstruck (Taxim-rendered, Mini-calibrated)
NC extension (repo)+66KSparsh (CC-BY-NC)

Quick start

python
from datasets import load_dataset, concatenate_datasets

# Single source
ds = load_dataset("yxma/gelsight-mini-pretrain", "fota_unlabeled", split="train")
img = ds[0]["image"]                  # PIL.Image (auto-decoded from JPEG bytes)

# Big real-markerless pretraining pool
pool = concatenate_datasets([
    load_dataset("yxma/gelsight-mini-pretrain", c, split="train")
    for c in ["fota_unlabeled", "gelslam", "feelanyforce",
              "real_tactile_mnist", "tacquad", "threedcal", "tactile_tracking"]
]).filter(lambda r: r["domain"] == "real" and not r["markered"])

Composition

SubsetFramesSplitsGelLabels
fota_unlabeled66,761trainmixed¹object name
gelslam114,019train + reconmarkerlessepisode + object
feelanyforce48,197trainmarkerless42 unique objects
real_tactile_mnist30,956train + testmarkerlessdigit + print id
fota_labeled26,394train + valmixed¹6-DoF pose + object
feats16,9696-split OOD benchmarkeredindenter + 3-axis force
tacquad12,195indoor/outdoor/finemarkerless181 objects
threedcal6,924trainmarkerless(x, y) sphere grid
tactile_tracking2,408trainmarkerlesstrial + object
unit387trainmarkered3D-pose target
sim_starstruck166,104train + testmarkerlessepisode (sim)
sim_tactile_mnist150,601train + testmarkerlessdigit + episode (sim)

¹ FoTA mixes markered + markerless gels per finger; the per-row markered column was auto-detected from dot density and is correct.

[image]

Pipeline (applied to every source)

  1. 1.Unified contact filter — area ≥ 40 px ∧ intensity ≥ I_min on the central-50% greyscale diff vs per-source baseline (I_min = 12 real, 10 sim); 1.5 % background-diversity keep rate.
  2. 2.Channel-order normalization — Mini's at-rest gel has B > R (3 colored LEDs); subsets where the upstream stored BGR are auto-detected (per-image R-B sign) and swapped to RGB. After this, every frame is guaranteed RGB.
  3. 3.JPEG q=92 re-encode + chunked-binary parquet writes (handles >2 GB shards safely).
  4. 4.Object diversity preserved — ~8,500 unique object instances across 13 physical sensor configurations.

Schema (30 columns, every row identical)

image (JPEG bytes), source, domain (real/sim), markered (bool), gel_variant (markered/markerless), capture, split, height, width, obj_name, episode, frame_idx, pose fields (x_mm, y_mm, z_mm, quat_*), FEATS fields (indenter, indenter_param, f_x, f_y, f_z, grid_z_*), digit_class, etc. — all optional fields are null when not applicable.

For per-subset details (paper, license, processing recipe, sample grids, stats), see [SOURCES.md](SOURCES.md).

Full build pipeline — per-source parameters cited to file:line, and the regression harness that checks each source against this release: <https://github.com/Yuxiang-Ma/gelsight-mini-pretrain>

Sample images

fota_labeled (markerless)fota_labeled (markered)
[image][image]
gelslamfeats (markered + force)
[image][image]
real_tactile_mnisttacquad (181 household objects)
[image][image]
sim_tactile_mnistsim_starstruck
[image][image]

Recommended uses

  • —Self-supervised pretraining (VAE / MAE / SimCLR / DINO) — concat all markerless real subsets (~472K frames), then fine-tune.
  • —Pose / force regression — fine-tune on fota_labeled (6-DoF), threedcal (xy + depth), or feats (3-axis force).
  • —Sim-to-real transfer — pretrain on sim_*, fine-tune on real.
  • —Marker-invariance studies — train markerless ↔ test on feats (markered).

Citations

Please cite both this aggregation and the upstream sources you use:

Investigated but not included

Touch-and-Go, TVL (Touch-Vision-Language), facebook/gelsight-force-estimation, YCB-Sight, TACTO/MidasTouch/DiffTactile — see SOURCES.md for reasons (wrong sensor, license, or not Mini-calibrated).

Changelog

  • —2026-08-06 — fota_labeled and fota_unlabeled were widened from 26 to 30 columns, so every config now shares one schema and the cross-config concatenate_datasets example above works. frame_idx, episode and digit_class are null for these two subsets: they were never recorded at build time and are not recoverable. Image bytes were not re-encoded — every original column passes through byte-identical.

License

CC-BY-4.0 for this aggregation. Cite the component datasets above. The companion `yxma/gelsight-mini-pretrain-nc` repo adds Sparsh (CC-BY-NC-4.0) for non-commercial use.