Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sayakpaul /pickapic_v2_webdatasetwebdataset archive of yuvalkirstain/pickapic_v2. Dataloading code can be found here. image1K<n<10K3 likes4.6k downloads3y agoHugging Face02yangyang857658468 /cc12m-webdataset CC12M WebDataset 这是CC12M数据集的WebDataset格式版本。 数据集信息 文件数量: 1098 总大小: 888796.33 MB 上传时间: 2025-03-18 14:45:49 使用方法 import webdataset as wds dataset = wds.WebDataset("https://huggingface.co/yangyang857658468/cc12m-webdataset/resolve/main/cc12m_*.tar") image10M<n<100M0 likes3.7k downloads2y agoHugging Face03hanlincs /InternVL-SA1B-Caption-WebDatasetThis repo contains the recaptioned SA1B images in webdataset format. The recaptioned prompts are from https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption image10M<n<100M1 likes3.6k downloads1y agoHugging Face04laion /conceptual-captions-12m-webdatasetimage10K<n<100K34 likes3.1k downloads5y agoHugging Face05cat-state /MegaSynth-webdatasetimage1M<n<10M0 likes2.5k downloads10mo agoHugging Face06collabora /hi-stt-preprocessed-webdatasettext100K<n<1M1 likes2.1k downloads1y agoHugging Face07laion /clevr-webdatasetimage1M<n<10M7 likes595 downloads4y agoHugging Face08cmeraki /audiofolder_webdatasetaudio100K<n<1M0 likes98 downloads2y agoHugging Face09ghemdd /gui_actor_webdataset GUI-Actor WebDataset A WebDataset format version of the GUI-Actor dataset for training vision-language models on GUI interaction tasks. Usage import webdataset as wds # Load the dataset dataset = wds.WebDataset("path/to/shards-*.tar") dataset = dataset.decode("pilrgb").to_tuple("jpg", "json") for image, metadata in dataset: # Process image and metadata pass Citation Please cite the original GUI-Actor paper if you use this dataset in your research. imagetext-generation1M<n<10M1 likes77 downloads1y agoHugging Face10lucasnewman /libritts-r-webdatasetOfficial website: https://www.openslr.org/141/ This repository contains LibriTTS-R converted to a WebDataset. The original Wave files have been converted to 64kbps MP3 files for efficient streaming. LibriTTS-R (paper) is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, published in 2019. The constituent samples of LibriTTS-R are identical to those of LibriTTS, with only the… See the full description on the dataset page: https://huggingface.co/datasets/lucasnewman/libritts-r-webdataset.audio100K<n<1M1 likes72 downloads2y agoHugging Face11collabora /multilingual-librispeech-webdatasetaudio10K<n<100K1 likes69 downloads3y agoHugging Face12vrachit /imagenet-1k-webdataset ImageNet-1k WebDataset This dataset contains ImageNet-1k in WebDataset format (tar files) for efficient streaming. Dataset Structure Training: 129 shards (train-*.tar) Validation: 5 shards (validation-*.tar) Total size: 147.82 GB Format Each tar file contains samples with: *.jpg: Image bytes *.cls: Label (class ID as text) Usage import webdataset as wds # Training dataset train_url = "train-{000000..000000000}.tar" dataset =… See the full description on the dataset page: https://huggingface.co/datasets/vrachit/imagenet-1k-webdataset.image1M<n<10M0 likes63 downloads10mo agoHugging Face13seastar105 /libritts-r-webdatasetaudio100K<n<1M0 likes56 downloads2y agoHugging Face14AIMClab-RUC /PhD-webdataset PhD Webdataset This repository contains the packaged version of PhD. For a detailed introduction to PhD, please visit the official website. Overview The PhD Webdataset is designed to facilitate easy access and usage of the PhD dataset. It includes various fields in 'json' key. The data in this repo is totally the same as in PhD. Installation Ensure you have Hugging Face's datasets library installed. You can install it via pip: pip install datasets… See the full description on the dataset page: https://huggingface.co/datasets/AIMClab-RUC/PhD-webdataset.imagevisual-question-answering100K<n<1M0 likes33 downloads2y agoHugging Face15fansunqi /web-dataset_3_screenshot_rendered_train_mhtml_3image10K<n<100K0 likes22 downloads9mo agoHugging Face16seastar105 /librispeech-webdatasetaudio100K<n<1M0 likes20 downloads2y agoHugging Face17Gatozu35 /test-webdatasettest audion<1K0 likes18 downloads2y agoHugging Face18hayden-donnelly /mnist-webdataset-png MNIST WebDataset PNG The MNIST dataset with samples stored as PNG images and compiled into the WebDataset format. DALI/JAX Example The following code shows how this dataset can be loaded into JAX arrays by DALI. from nvidia.dali import pipeline_def import nvidia.dali.fn as fn import nvidia.dali.types as types from nvidia.dali.plugin.jax import DALIGenericIterator from nvidia.dali.plugin.base_iterator import LastBatchPolicy def get_data_iterator(batch_size, dataset_path):… See the full description on the dataset page: https://huggingface.co/datasets/hayden-donnelly/mnist-webdataset-png.imageimage-classification10K<n<100K0 likes17 downloads3y agoHugging Face19Sh1man /example_webdatasetaudion<1K0 likes15 downloads1y agoHugging Face20Nayana-cognitivelab /NayanaDocs-Indic-45k-webdatasetgated Nayana-DocOCR Indic Annotated Dataset Dataset Description This is a large-scale multilingual document OCR dataset containing approximately 400GB of images with comprehensive annotations across multiple languages including Indic languages and English. The dataset is stored in WebDataset format using TAR archives for efficient streaming and processing. Available Language Subsets bn (Bengali): Available en (English): Available gu (Gujarati): Available hi (Hindi):… See the full description on the dataset page: https://huggingface.co/datasets/Nayana-cognitivelab/NayanaDocs-Indic-45k-webdataset.imageimage-to-text100K<n<1M0 likes15 downloads1y agoHugging Face21adricl /midi_godzilla_piano_webdataset_1024 MIDI Godzilla Piano Webdataset split into 1024 chunks Webdataset Midi of the Godzilla MIDI Dataset from Project Los Angeles This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description This dataset has been created by splitting the midi files into 1024 tokens. We then split the traning set into 70% traning 15% validation and 15% test. We augment the midi as per… See the full description on the dataset page: https://huggingface.co/datasets/adricl/midi_godzilla_piano_webdataset_1024.text10M<n<100M0 likes14 downloads7mo agoHugging Face22Bigbarry /webdataset_copyimage100K<n<1M0 likes13 downloads5mo agoHugging Face23Qualeafclover /webdataset-cifar100text10K<n<100K0 likes12 downloads1y agoHugging Face24fansunqi /web-dataset_4_interact_resultsimage10K<n<100K0 likes12 downloads9mo agoHugging Face25fansunqi /web-dataset_4_screenshot_rendered_train_mhtml_4image1K<n<10K0 likes12 downloads9mo agoHugging Face26sam8000 /EuroSpeech-WebDatasetaudio1K<n<10K0 likes10 downloads1y agoHugging Face27marcosremar2 /instructs2s-webdatasetaudio100K<n<1M0 likes10 downloads9mo agoHugging Face28fansunqi /web-dataset_3_mhtml_filestext10K<n<100K0 likes9 downloads9mo agoHugging Face29cmeraki /youtube_webdatasetgatedaudio10K<n<100K0 likes8 downloads2y agoHugging Face30Nayana-cognitivelab /NayanaDocs-Multilingual-webdatasetgatedimageimage-to-text1K<n<10K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.