datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
infinity-mm-stage1-webdataset
WebDataset Image-Text Dataset
This dataset contains image-text pairs in WebDataset format.
Dataset Structure
Each .tar.wds file contains entries with JSON data including image, text, and metadata.
waymo_webdatasetbiomedica_webdataset_24M
Dataset Card for Dataset Name
Arxiv: Arxiv
|
Website: Biomedica
|
Training instructions: OpenCLIP
|
Tutorial: Google Colab
BIOMEDICA Dataset is a large-scale, deep-learning-ready biomedical dataset containing over 24M imagecaption pairs and 30M image-references from 6M unique open-source articles. Each… See the full description on the dataset page: https://huggingface.co/datasets/BIOMEDICA/biomedica_webdataset_24M.conceptual-captions-12m-webdataset-bertspickapic_v2_webdatasetwebdataset archive of yuvalkirstain/pickapic_v2.
Dataloading code can be found here.
WEB-Dataset
WorldEngine Bimanual Dataset for Post-training
A large-scale, language-annotated real-robot bimanual manipulation dataset for
post-training robotics foundation models. It spans 90 everyday manipulation tasks
collected with a bimanual YAM follower arm teleoperated by a GELLO leader,
recording joint state, action, and three synchronized camera streams at 60 Hz.
Shared lineage, different story. This dataset shares its hardware, teleoperation
setup, and recording pipeline with the… See the full description on the dataset page: https://huggingface.co/datasets/WorldEngineAI/WEB-Dataset.cc12m-webdataset
CC12M WebDataset
这是CC12M数据集的WebDataset格式版本。
数据集信息
文件数量: 1098
总大小: 888796.33 MB
上传时间: 2025-03-18 14:45:49
使用方法
import webdataset as wds
dataset = wds.WebDataset("https://huggingface.co/yangyang857658468/cc12m-webdataset/resolve/main/cc12m_*.tar")
InternVL-SA1B-Caption-WebDatasetThis repo contains the recaptioned SA1B images in webdataset format. The recaptioned prompts are from https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption
conceptual-captions-12m-webdatasetLLaVA-OneVision-1.5-Mid-Training-Webdataset-Quick-Start-3MMegaSynth-webdatasetlibrilight-webdatasethi-stt-preprocessed-webdatasetconceptual-captions-12m-webdataset-metadata
Conceptual Captions 12M — Webshart metadata indices
Per-shard webshart metadata indices for
laion/conceptual-captions-12m-webdataset:
1,100 JSON files under data/, one per source tar shard, mirroring the source's shard layout.
Each index records every tar member's byte offset and length (enabling ranged reads without
downloading whole shards), image geometry (width/height for aspect bucketing), and — as of
August 2026 — embedded captions for all 10,994,853 samples, coalesced… See the full description on the dataset page: https://huggingface.co/datasets/webshart/conceptual-captions-12m-webdataset-metadata.Kuro-Siwo-Webdataset
Kuro Siwo webdatasets
Paper | GitHub |
Dataset Details
Dataset Description
Kuro Siwo is a global multi-temporal SAR dataset for rapid flood mapping. It contains 43 flood events in 6 continents and 3 climate zones, over the period 2015-2022. The annotations have been produced through meticulous photointerpretation by a team of experts, at 10m spatial resolution. For each flood event, we provide one Sentinel-1 post-flood and two Sentinel-1 pre-flood… See the full description on the dataset page: https://huggingface.co/datasets/orion-ai-lab/Kuro-Siwo-Webdataset.LLaVA-NeXT-780k-webdatasetStage-2: Visual Instruction Tuning, this is a subset for quick start.
clevr-webdatasetgeobenchv1-webdatasetBCN20000_WebDatasetweb-dataset_2fleurs-regmix-webdataset
FLEURS RegMix WebDataset
Public, training-oriented WebDataset conversion of google/fleurs, pinned to source revision 70bb2e84b976b7e960aa89f1c648e09c59f894dd.
Layout
Each language is a RegMix cluster at data/<language>/*.tar. Shard names keep the source split, data/<language>/<language>-<split>-<index>.tar, so a training
mixture can be assembled without pulling the FLEURS evaluation splits into it. Every sample is a pair with the same key:
<key>.opus: mono Ogg… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/fleurs-regmix-webdataset.YFCC14M_subset_webdatasetLLaVA-558K-WebdatasetThis data has been packed, so it may look like there are not many samples, but the actual number is 558k.
biomedica_microscopy_subset_webdatasetcommon-voice-17-regmix-webdataset
Common Voice 17 RegMix WebDataset
Public, training-oriented WebDataset conversion of fsicoli/common_voice_17_0, pinned to source revision 8262c16bf297c87a9cd88c51997c4758ed7a8ba2.
Layout
Each language is a RegMix cluster at data/<language>/*.tar. Every sample is a pair with the same key:
<key>.opus: mono Ogg Opus audio at 24 kbps, preserving the source sample rate (normally 48 kHz)
<key>.json: UTF-8 training metadata and the complete original TSV row
The source… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/common-voice-17-regmix-webdataset.audiofolder_webdatasetasrtts_packed_webdataset
ASR+TTS Repacked Data (3.56M samples, mp3)
This dataset is a WebDataset repack prepared for FlexiSLM training (ASR+TTS tasks).
Paper: https://arxiv.org/abs/2606.31247
Demo page: https://flexislm.github.io/
Code: https://github.com/AmphionTeam/FlexiSLM
FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset
for training FlexiSLM, a spoken language model.
This repository contains the paired prompt-and-response audio portion of the release in… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/asrtts_packed_webdataset.openhermes-2.5-webdatasetgeobenchv1-webdataset
GeoBench V1 - pickle-free sharded edition
This is a data-only metadata conversion of the six classification datasets in isaaccorley/geobenchv1-webdataset, pinned at revision cb847e5ff87a2c8f00064d631ed7b6a39a084c68. The original benchmark is GEO-Bench.
The image arrays, sample IDs, labels, shard boundaries, and train/validation/test partitions are unchanged. Sample metadata is stored as JSON instead of pickle. No pickle deserialization is needed to read this edition.… See the full description on the dataset page: https://huggingface.co/datasets/calebrob6/geobenchv1-webdataset.wai-webdataset-metadata
Transfer metadata for WAI shard datasets
This folder is the WebDataset machine's dataset_metadata_dir. It contains
the original split lists, packed complete covisibility graphs, and a portable
wai-storage.json manifest. Shard catalogs go in scene_catalogs/ after they
are built on the machine hosting the shards.
Layout and existing Memor configuration
root_data_dir/
ase/ -> /host/shards/ase/
blendedmvs/ -> /host/shards/blendedmvs/… See the full description on the dataset page: https://huggingface.co/datasets/theairlabcmu/wai-webdataset-metadata.
