Team Ai
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01noxwano /ASMR-Archive-Processed-SFW ASMR-Archive-Processed-SFW Overview This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset. We filtered the original dataset to include only records where the nsfw metadata flag is false. To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled. The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.audioautomatic-speech-recognition1M<n<10M9 likes3.5k downloads6mo agoHugging Face02milashkaarshif /MoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials. Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset. Looking forward to see more models and synthetic datasets trained from this raw archive, good luck! Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.texttext-generation100K<n<1M38 likes564 downloads8mo agoHugging Face03heesup /Cowpea-Architecture-XML Cowpea-Architecture-XML-WDS This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training. Dataset Structure The dataset is sharded into .tar files, each containing up to 10,000 samples. Each sample consists of: .jpeg: The plant image .xml: The organ-level architecture representation .json: (Optional) Metadata Usage with… See the full description on the dataset page: https://huggingface.co/datasets/heesup/Cowpea-Architecture-XML.imageimage-to-text100K<n<1M0 likes216 downloads6mo agoHugging Face04data-archetype /ffhq_captioned_1024 ffhq_captioned_1024 A captioned bucketed-shards export of gaunernst/ffhq-1024-wds. This export contains 70,000 square face and portrait images from FFHQ, stored as JPEG TAR shards in a single 1024 x 1024 bucket. The source images are decoded from the original dataset, deterministically converted to RGB, and re-encoded as high-quality JPEG (quality=95, adaptive subsampling). Captions were generated with a Gemini 2.5 Flash Lite primary pass and a Mistral Medium 3.1 fallback. Intended… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/ffhq_captioned_1024.imagetext-to-image10K<n<100K0 likes193 downloads6mo agoHugging Face05bbrangeo /Cowpea-Architecture-XML Cowpea-Architecture-XML-WDS This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training. Dataset Structure The dataset is sharded into .tar files, each containing up to 10,000 samples. Each sample consists of: .jpeg: The plant image .xml: The organ-level architecture representation .json: (Optional) Metadata… See the full description on the dataset page: https://huggingface.co/datasets/bbrangeo/Cowpea-Architecture-XML.imageimage-to-text100K<n<1M0 likes122 downloads4mo agoHugging Face06Vid2dex /Arctic_datatext10K<n<100K0 likes112 downloads7mo agoHugging Face07arch-raven /music-fingerprint-dataset Neural Audio Fingerprint Dataset (c) 2021 by Sungkyun Chang https://github.com/mimbres/neural-audio-fp This dataset includes all music sources, background noise and impulse-reponses (IR) samples that have been used in the work ["Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning"] (https://arxiv.org/abs/2010.11910). Format: 16-bit PCM Mono WAV, Sampling rate 8000 Hz Description: / fingerprint_dataset_icassp2021/… See the full description on the dataset page: https://huggingface.co/datasets/arch-raven/music-fingerprint-dataset.audio10K<n<100K8 likes78 downloads5y agoHugging Face08data-archetype /imagenet_22k_512_bucketable ImageNet-22k 512-Bucketable Captioned Subset This dataset is a pre-bucketed, captioned subset of timm/imagenet-22k-wds. It is intended for text-to-image training and similar workflows that want images already grouped into aspect-ratio buckets near a 512-base training resolution. Images were kept only if they could fit one of the target buckets without upsampling after deterministic resize and crop. Summary Source: timm/imagenet-22k-wds (fall11 ImageNet-22k WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/imagenet_22k_512_bucketable.imagetext-to-image1M<n<10M0 likes72 downloads6mo agoHugging Face09Yiming1234 /VoT-video-latent-archiveimage100K<n<1M0 likes70 downloads6mo agoHugging Face10data-archetype /cc12_imagenet21k_recap_hq_bucketed cc12_imagenet21k_recap_hq_bucketed Title: cc12_imagenet21k_recap_hq_bucketed Description: This ~18M rows dataset is a re upload of https://huggingface.co/datasets/gmongaras/CC12M_and_Imagenet21K_Recap_Highqual where the images have been pre bucketed into SDXL style aspect ratio buckets for target training at ~512^2 and ~256^2 pixels, and where about 7M rows were recaptioned with either Gemini or Ministral. To avoid re encoding the images they have been left untouched so cropping… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/cc12_imagenet21k_recap_hq_bucketed.imagetext-to-image10M<n<100M0 likes50 downloads9mo agoHugging Face11data-archetype /LAION_Aesthetics_1024_bucketed_512 LAION Aesthetics 1024 Bucketed 512 Captioned This is a captioned bucketed-shards export of images from limingcv/LAION_Aesthetics_1024. Images were filtered and resized/cropped into SDXL-style aspect-ratio buckets at a 512 base resolution, without upsampling. The export contains 382,144 images across 397 uncompressed WebDataset-style tar shards. The .txt files now contain model-generated captions, not the original LAION web-scrape alt text or surrounding page text. Captions were… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/LAION_Aesthetics_1024_bucketed_512.imagetext-to-image100K<n<1M0 likes39 downloads6mo agoHugging Face12mpatrick1991 /srtm-3-arc-second-global SRTM 3 Arc-Second Global Raw ASCII heightmaps of the Earth's surface labelled according to latitude and longitude. Mission Description The Shuttle Radar Topography Mission (SRTM) was flown aboard the space shuttle Endeavour February 11-22, 2000. The National Aeronautics and Space Administration (NASA) and the National Geospatial-Intelligence Agency (NGA) participated in an international project to acquire radar data which were used to create the first near-global set of… See the full description on the dataset page: https://huggingface.co/datasets/mpatrick1991/srtm-3-arc-second-global.textunconditional-image-generation10K<n<100K0 likes37 downloads9mo agoHugging Face13data-archetype /bg_photo_concepts_bucketed_512 bg_photo_concepts_bucketed_512 Title: bg_photo_concepts_bucketed_512 Description: A recaptioned, self contained, bucketed and ready to train with version of https://huggingface.co/datasets/bghira/photo-concept-bucket, exported at 512^2 ish resolution buckets. I lost the tracking data of which version of Gemini this was captioned with, likely 2.0 flash or 2.5 flash. The captions are on the long and datailed side and sometimes slightly redundant, but overall high quality.… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/bg_photo_concepts_bucketed_512.imagetext-to-image100K<n<1M1 likes24 downloads5mo agoHugging Face14data-archetype /classical_paintings_bucketed_1024 Classical Paintings Captioned A curated dataset of 7,131 classical paintings by 42 artists spanning the Baroque period through the 19th century, each with a descriptive plain-language caption (100--150 words). Intended for fine-tuning text-to-image models. Artists (42) Aelbert Cuyp, Albert Bierstadt, Anders Zorn, Anthony van Dyck, Artemisia Gentileschi, Caravaggio, Diego Velazquez, Frans Hals, Frederic Edwin Church, Georges de La Tour, Gerard ter Borch, Gerrit Dou… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/classical_paintings_bucketed_1024.imagetext-to-image1K<n<10K0 likes23 downloads6mo agoHugging Face15tekkithorse /mlp-board-yuki-archive-html-textTaken from the yuki archive data. Contains the html text from the /mlp/ board up until a few years ago. It sure would be interesting to have a dataset that's organized by each thread in a spreadsheet separated by row with the posts in the thread separated by column, with the post number in each cell. text100K<n<1M0 likes20 downloads3y agoHugging Face16archivartaunik /commonvoice22_sidon_be_rawaudio100K<n<1M0 likes14 downloads7mo agoHugging Face17penfever /archetype-amstrtext1K<n<10K0 likes11 downloads2y agoHugging Face18penfever /archetype-pubchemtext1K<n<10K0 likes11 downloads2y agoHugging Face19mops-neurips26 /mops-data-archiveThis is the anonymous Dataset upload for Multi-Objective Photoreal Simulation (MOPS) Dataset [under review]. This is the archived WebDataset version. image100K<n<1M0 likes11 downloads6mo agoHugging Face20gaganyatri /dhwani_spaces_archivetextn<1K0 likes9 downloads2y agoHugging Face21pitstoppie /ARGOS-ARCHIVEimage100K<n<1M0 likes9 downloads6mo agoHugging Face22migliorini /archivio_dopaudio10K<n<100K0 likes8 downloads5mo agoHugging Face23arcct /hourtext100K<n<1M0 likes4 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.