yxma/gelsight-mini-pretrain
GelSight Mini Pretrain ~853K GelSight Mini tactile RGB frames, 12 public sources, one parquet schema. Built for self-supervised representation learning (VAE / MAE / SimCLR / DINO) — every frame contact-filtered, channel-normalized, and re-encoded as JPEG q92. Frames Sources Real 536K FoTA (labeled+unlabeled), 3DCal, FEATS, GelSLAM, TactileTracking, RTM, FeelAnyForce, UniT, TacQuad Sim 317K sim_tactile_mnist, sim_starstruck (Taxim-rendered, Mini-calibrated) NC… See the full description on the dataset page: https://huggingface.co/datasets/yxma/gelsight-mini-pretrain.
GelSight Mini Pretrain
~853K [GelSight Mini](https://www.gelsight.com/gelsightmini/) tactile RGB frames, 12 public sources, one parquet schema. Built for self-supervised representation learning (VAE / MAE / SimCLR / DINO) — every frame contact-filtered, channel-normalized, and re-encoded as JPEG q92.
Quick start
from datasets import load_dataset, concatenate_datasets
# Single source
ds = load_dataset("yxma/gelsight-mini-pretrain", "fota_unlabeled", split="train")
img = ds[0]["image"] # PIL.Image (auto-decoded from JPEG bytes)
# Big real-markerless pretraining pool
pool = concatenate_datasets([
load_dataset("yxma/gelsight-mini-pretrain", c, split="train")
for c in ["fota_unlabeled", "gelslam", "feelanyforce",
"real_tactile_mnist", "tacquad", "threedcal", "tactile_tracking"]
]).filter(lambda r: r["domain"] == "real" and not r["markered"])Composition
¹ FoTA mixes markered + markerless gels per finger; the per-row markered column was auto-detected from dot density and is correct.
Pipeline (applied to every source)
- Unified contact filter —
area ≥ 40 px ∧ intensity ≥ I_minon the central-50% greyscale diff vs per-source baseline (I_min = 12 real, 10 sim); 1.5 % background-diversity keep rate. - Channel-order normalization — Mini's at-rest gel has B > R (3 colored LEDs); subsets where the upstream stored BGR are auto-detected (per-image R-B sign) and swapped to RGB. After this, every frame is guaranteed RGB.
- JPEG q=92 re-encode + chunked-binary parquet writes (handles >2 GB shards safely).
- Object diversity preserved — ~8,500 unique object instances across 13 physical sensor configurations.
Schema (30 columns, every row identical)
image (JPEG bytes), source, domain (real/sim), markered (bool), gel_variant (markered/markerless), capture, split, height, width, obj_name, episode, frame_idx, pose fields (x_mm, y_mm, z_mm, quat_*), FEATS fields (indenter, indenter_param, f_x, f_y, f_z, grid_z_*), digit_class, etc. — all optional fields are null when not applicable.
For per-subset details (paper, license, processing recipe, sample grids, stats), see [SOURCES.md](SOURCES.md).
Full build pipeline — per-source parameters cited to file:line, and the regression harness that checks each source against this release: <https://github.com/Yuxiang-Ma/gelsight-mini-pretrain>
Sample images
Recommended uses
- Self-supervised pretraining (VAE / MAE / SimCLR / DINO) — concat all
markerlessreal subsets (~472K frames), then fine-tune. - Pose / force regression — fine-tune on
fota_labeled(6-DoF),threedcal(xy + depth), orfeats(3-axis force). - Sim-to-real transfer — pretrain on
sim_*, fine-tune on real. - Marker-invariance studies — train markerless ↔ test on
feats(markered).
Citations
Please cite both this aggregation and the upstream sources you use:
- FoTA (HF, arXiv:2406.13640) · MIT
- py3DCal (Zenodo) · CC-BY-4.0
- FEATS (HF) · MIT
- GelSLAM (HF, arXiv:2508.15990) · MIT
- TactileTracking / NormalFlow (HF, RA-L 2024) · MIT
- Real Tactile MNIST (HF family, arXiv:2506.06361) · CC-BY-2.0
- FeelAnyForce (HF) · CC-BY-4.0
- UniT (GitHub) · BSD-3-Clause-style
- TacQuad / AnyTouch (HF) · CC-BY-4.0
- Taxim (sim renderer, GitHub, arXiv:2109.04027)
Investigated but not included
Touch-and-Go, TVL (Touch-Vision-Language), facebook/gelsight-force-estimation, YCB-Sight, TACTO/MidasTouch/DiffTactile — see SOURCES.md for reasons (wrong sensor, license, or not Mini-calibrated).
Changelog
- 2026-08-06 —
fota_labeledandfota_unlabeledwere widened from 26 to 30 columns, so every config now shares one schema and the cross-configconcatenate_datasetsexample above works.frame_idx,episodeanddigit_classare null for these two subsets: they were never recorded at build time and are not recoverable. Image bytes were not re-encoded — every original column passes through byte-identical.
License
CC-BY-4.0 for this aggregation. Cite the component datasets above. The companion `yxma/gelsight-mini-pretrain-nc` repo adds Sparsh (CC-BY-NC-4.0) for non-commercial use.
