Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yangyang857658468 /infinity-mm-stage1-webdataset WebDataset Image-Text Dataset This dataset contains image-text pairs in WebDataset format. Dataset Structure Each .tar.wds file contains entries with JSON data including image, text, and metadata. 0 likes78k downloads2y agoHugging Face02Victer-XiaoyuYe /waymo_webdataset1 likes13k downloads1y agoHugging Face03BIOMEDICA /biomedica_webdataset_24Mgated Dataset Card for Dataset Name Arxiv: Arxiv &nbsp;&nbsp;&nbsp;&nbsp;|&nbsp;&nbsp;&nbsp;&nbsp; Website: Biomedica &nbsp;&nbsp;&nbsp;&nbsp;|&nbsp;&nbsp;&nbsp;&nbsp; Training instructions: OpenCLIP &nbsp;&nbsp;&nbsp;&nbsp;|&nbsp;&nbsp;&nbsp;&nbsp; Tutorial: Google Colab BIOMEDICA Dataset is a large-scale, deep-learning-ready biomedical dataset containing over 24M imagecaption pairs and 30M image-references from 6M unique open-source articles. Each… See the full description on the dataset page: https://huggingface.co/datasets/BIOMEDICA/biomedica_webdataset_24M.n>1T40 likes8.4k downloads1mo agoHugging Face04AbstractPhil /conceptual-captions-12m-webdataset-bertstext10M<n<100M1 likes7k downloads2mo agoHugging Face05sayakpaul /pickapic_v2_webdatasetwebdataset archive of yuvalkirstain/pickapic_v2. Dataloading code can be found here. image1K<n<10K3 likes4.6k downloads3y agoHugging Face06WorldEngineAI /WEB-Datasetgated WorldEngine Bimanual Dataset for Post-training A large-scale, language-annotated real-robot bimanual manipulation dataset for post-training robotics foundation models. It spans 90 everyday manipulation tasks collected with a bimanual YAM follower arm teleoperated by a GELLO leader, recording joint state, action, and three synchronized camera streams at 60 Hz. Shared lineage, different story. This dataset shares its hardware, teleoperation setup, and recording pipeline with the… See the full description on the dataset page: https://huggingface.co/datasets/WorldEngineAI/WEB-Dataset.video100K<n<1M1 likes4.2k downloads3mo agoHugging Face07yangyang857658468 /cc12m-webdataset CC12M WebDataset 这是CC12M数据集的WebDataset格式版本。 数据集信息 文件数量: 1098 总大小: 888796.33 MB 上传时间: 2025-03-18 14:45:49 使用方法 import webdataset as wds dataset = wds.WebDataset("https://huggingface.co/yangyang857658468/cc12m-webdataset/resolve/main/cc12m_*.tar") image10M<n<100M0 likes3.7k downloads2y agoHugging Face08hanlincs /InternVL-SA1B-Caption-WebDatasetThis repo contains the recaptioned SA1B images in webdataset format. The recaptioned prompts are from https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption image10M<n<100M1 likes3.6k downloads1y agoHugging Face09laion /conceptual-captions-12m-webdatasetimage10K<n<100K34 likes3.1k downloads5y agoHugging Face10mvp-lab /LLaVA-OneVision-1.5-Mid-Training-Webdataset-Quick-Start-3M3 likes2.8k downloads1y agoHugging Face11cat-state /MegaSynth-webdatasetimage1M<n<10M0 likes2.5k downloads10mo agoHugging Face12collabora /librilight-webdataset2 likes2.4k downloads3y agoHugging Face13collabora /hi-stt-preprocessed-webdatasettext100K<n<1M1 likes2.1k downloads1y agoHugging Face14webshart /conceptual-captions-12m-webdataset-metadata Conceptual Captions 12M — Webshart metadata indices Per-shard webshart metadata indices for laion/conceptual-captions-12m-webdataset: 1,100 JSON files under data/, one per source tar shard, mirroring the source's shard layout. Each index records every tar member's byte offset and length (enabling ranged reads without downloading whole shards), image geometry (width/height for aspect bucketing), and — as of August 2026 — embedded captions for all 10,994,853 samples, coalesced… See the full description on the dataset page: https://huggingface.co/datasets/webshart/conceptual-captions-12m-webdataset-metadata.1 likes1.9k downloads2mo agoHugging Face15orion-ai-lab /Kuro-Siwo-Webdataset Kuro Siwo webdatasets Paper | GitHub | Dataset Details Dataset Description Kuro Siwo is a global multi-temporal SAR dataset for rapid flood mapping. It contains 43 flood events in 6 continents and 3 climate zones, over the period 2015-2022. The annotations have been produced through meticulous photointerpretation by a team of experts, at 10m spatial resolution. For each flood event, we provide one Sentinel-1 post-flood and two Sentinel-1 pre-flood… See the full description on the dataset page: https://huggingface.co/datasets/orion-ai-lab/Kuro-Siwo-Webdataset.geospatialimage-classification3 likes1.3k downloads2mo agoHugging Face16mvp-lab /LLaVA-NeXT-780k-webdatasetStage-2: Visual Instruction Tuning, this is a subset for quick start. 0 likes972 downloads1y agoHugging Face17laion /clevr-webdatasetimage1M<n<10M7 likes595 downloads4y agoHugging Face18isaaccorley /geobenchv1-webdataset0 likes192 downloads5mo agoHugging Face19zakerous /BCN20000_WebDataset0 likes172 downloads10mo agoHugging Face20fansunqi /web-dataset_20 likes155 downloads11mo agoHugging Face21quinnlue /fleurs-regmix-webdataset FLEURS RegMix WebDataset Public, training-oriented WebDataset conversion of google/fleurs, pinned to source revision 70bb2e84b976b7e960aa89f1c648e09c59f894dd. Layout Each language is a RegMix cluster at data/<language>/*.tar. Shard names keep the source split, data/<language>/<language>-<split>-<index>.tar, so a training mixture can be assembled without pulling the FLEURS evaluation splits into it. Every sample is a pair with the same key: <key>.opus: mono Ogg… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/fleurs-regmix-webdataset.automatic-speech-recognition100K<n<1M0 likes154 downloads1mo agoHugging Face22ARKseal /YFCC14M_subset_webdatasetimage1M<n<10M0 likes128 downloads5y agoHugging Face23mvp-lab /LLaVA-558K-WebdatasetThis data has been packed, so it may look like there are not many samples, but the actual number is 558k. 4 likes115 downloads1y agoHugging Face24BIOMEDICA /biomedica_microscopy_subset_webdatasetgatedn>1T4 likes108 downloads2y agoHugging Face25quinnlue /common-voice-17-regmix-webdataset Common Voice 17 RegMix WebDataset Public, training-oriented WebDataset conversion of fsicoli/common_voice_17_0, pinned to source revision 8262c16bf297c87a9cd88c51997c4758ed7a8ba2. Layout Each language is a RegMix cluster at data/<language>/*.tar. Every sample is a pair with the same key: <key>.opus: mono Ogg Opus audio at 24 kbps, preserving the source sample rate (normally 48 kHz) <key>.json: UTF-8 training metadata and the complete original TSV row The source… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/common-voice-17-regmix-webdataset.automatic-speech-recognition10M<n<100M0 likes102 downloads1mo agoHugging Face26cmeraki /audiofolder_webdatasetaudio100K<n<1M0 likes98 downloads2y agoHugging Face27FlexiSLM /asrtts_packed_webdataset ASR+TTS Repacked Data (3.56M samples, mp3) This dataset is a WebDataset repack prepared for FlexiSLM training (ASR+TTS tasks). Paper: https://arxiv.org/abs/2606.31247 Demo page: https://flexislm.github.io/ Code: https://github.com/AmphionTeam/FlexiSLM FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset for training FlexiSLM, a spoken language model. This repository contains the paired prompt-and-response audio portion of the release in… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/asrtts_packed_webdataset.automatic-speech-recognition1M<n<10M1 likes95 downloads2mo agoHugging Face28marianna13 /openhermes-2.5-webdataset2 likes94 downloads3y agoHugging Face29calebrob6 /geobenchv1-webdataset GeoBench V1 - pickle-free sharded edition This is a data-only metadata conversion of the six classification datasets in isaaccorley/geobenchv1-webdataset, pinned at revision cb847e5ff87a2c8f00064d631ed7b6a39a084c68. The original benchmark is GEO-Bench. The image arrays, sample IDs, labels, shard boundaries, and train/validation/test partitions are unchanged. Sample metadata is stored as JSON instead of pickle. No pickle deserialization is needed to read this edition.… See the full description on the dataset page: https://huggingface.co/datasets/calebrob6/geobenchv1-webdataset.geospatialimage-classification10K<n<100K0 likes93 downloads1mo agoHugging Face30theairlabcmu /wai-webdataset-metadata Transfer metadata for WAI shard datasets This folder is the WebDataset machine's dataset_metadata_dir. It contains the original split lists, packed complete covisibility graphs, and a portable wai-storage.json manifest. Shard catalogs go in scene_catalogs/ after they are built on the machine hosting the shards. Layout and existing Memor configuration root_data_dir/ ase/ -> /host/shards/ase/ blendedmvs/ -> /host/shards/blendedmvs/… See the full description on the dataset page: https://huggingface.co/datasets/theairlabcmu/wai-webdataset-metadata.0 likes91 downloads28d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.