datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gelbooru-webp-4Mpixel
Gelbooru 4M Re-encoded Dataset
This is the re-encoded dataset of deepghs/gelbooru_full. And all the resized images are maintained here.
There are 11463682 images in total. The maximum ID of these images is 12575222. Last updated at 2025-09-09 09:40:06 JST.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/gelbooru-webp-4Mpixel.sankaku-webp-256shortest-edgeanime_pictures-webp-4Mpixel
Anime-Pictures 4M Re-encoded Dataset
This is the re-encoded dataset of deepghs/anime_pictures_full. And all the resized images are maintained here.
There are 642979 images in total. The maximum ID of these images is 885353. Last updated at 2025-09-24 01:12:10 CST.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/anime_pictures-webp-4Mpixel.pixelprose_webp_512safebooru-webp-4Mpixel
Safebooru 4M Re-encoded Dataset
This is the re-encoded dataset of deepghs/safebooru_full. And all the resized images are maintained here.
There are 5756655 images in total. The maximum ID of these images is 5974383. Last updated at 2025-08-06 08:31:53 JST.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/safebooru-webp-4Mpixel.danbooru2024-webp-4Mpixel
🎨 Danbooru2024 Webp 4MPixel Dataset
📊 Dataset Overview
The Danbooru2024-Webp dataset is a comprehensive collection focused on animation and illustration artwork, derived from the official Danbooru platform. It contains approximately 8.05 million high-quality, user-annotated images with corresponding tags and textual descriptions.
This dataset is 4MP-focused webp resized-dataset of Danbooru2024.
✨ Features
📋 Metadata Support
Includes a Parquet… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/danbooru2024-webp-4Mpixel.dan-webp-newe621-2024-webp-4MpixelDataset Description:
This is a processed version of the https://huggingface.co/datasets/boxingscorpionbagel/e621-2024 dataset, primarily prepared for personal use in future projects.
Therefore, for licensing and other legal information, please refer to the original project.
You can directly download tar file,or use https://deepghs.github.io/hfutils/main/api_doc/index/fetch.html#hf-tar-file-download to download anything .webp file you want.
The following modifications have been made to the… See the full description on the dataset page: https://huggingface.co/datasets/NebulaeWis/e621-2024-webp-4Mpixel.webpagefancaps-webp-4Mpixel
Fancaps 4M Re-encoded Dataset
This is the re-encoded dataset of deepghs/fancaps_full. And all the resized images are maintained here.
There are 29004217 images in total. The maximum ID of these images is 31449533. Last updated at 2025-09-04 22:10:59 JST.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/fancaps-webp-4Mpixel.shirasu_azusa_webp蔚蓝档案角色shirasu azusa的图像数据集,经过数次质量压缩后转换为webp格式存储,控制总和在50mb大小范围内,用于civitai的lora上传训练。
imagenet-w21-webp-wds
Dataset Summary
This is a copy of the full Winter21 release of ImageNet in webdataset tar format with WEBP encoded images. This release consists of 19167 classes, 2674 fewer classes than the original 21841 class Fall11 release of the full ImageNet.
The classes were removed due to these concerns: https://www.image-net.org/update-sep-17-2019.php
This is the same contents as https://huggingface.co/datasets/timm/imagenet-w21-wds but encoded in webp at ~56% of the size, shard count… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-w21-webp-wds.sankaku-webp-4Mpixel
Sankaku 4M Re-encoded Dataset
This is the re-encoded dataset of deepghs/sankaku_full. And all the resized images are maintained here.
There are 16492859 images in total. The maximum ID of these images is 36864051. Last updated at 2025-01-02 12:41:55 JST.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/sankaku-webp-4Mpixel.WebPerson-5M
🔥 Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval [EMNLP25 Main]
Tianlu Zheng*,
Yifan Zhang*,
Xiang An,
Ziyong Feng,
Kaicheng Yang†,
Qichunan Ding†,
📄 Paper | 💻 Github
✨ Web-Person Dataset
🔍 Person-Centric Image Filtering
We use the COYO700M dataset as our source of web-crawled images.
To curate high-quality person-centric images, we apply YOLOv11 to detect humans and extract bounding… See the full description on the dataset page: https://huggingface.co/datasets/Kaichengalex/WebPerson-5M.web-pinkykonachan-webp-4Mpixel
Konachan 4M Re-encoded Dataset
This is the re-encoded dataset of deepghs/konachan_full. And all the resized images are maintained here.
There are 319012 images in total. The maximum ID of these images is 391069. Last updated at 2025-07-22 02:47:52 JST.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/konachan-webp-4Mpixel.imagenet_288_webpA duplicate of the webdataset version of ImageNet-1k, except each image is resized so that the shorter side is 288 pixels and each image is WEBP compressed using a quality level of 85.
The compression ratio is roughly 5:1 compared to JPEG (~160 GB) and 25:1 compared to raw pixel storage (~800GB).
WebPerson-1M
🔥 Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval [EMNLP25 Main]
Tianlu Zheng*,
Yifan Zhang*,
Xiang An,
Ziyong Feng,
Kaicheng Yang†,
Qichunan Ding†,
📄 Paper | 💻 Github
✨ Web-Person Dataset
🔍 Person-Centric Image Filtering
We use the COYO700M dataset as our source of web-crawled images.
To curate high-quality person-centric images, we apply YOLOv11 to detect humans and extract bounding… See the full description on the dataset page: https://huggingface.co/datasets/Kaichengalex/WebPerson-1M.danbooru2023-webp-4Mpixel_indexIndex files of KBlueLeaf/danbooru2023-webp-4Mpixel.
You can download images from KBlueLeaf/danbooru2023-webp-4Mpixel with cheesechaser.
from cheesechaser.datapool import DanbooruWebpDataPool
pool = DanbooruWebpDataPool()
# download danbooru images with webp format, to directory /data/danbooru_webp
pool.batch_download_to_directory(
resource_ids=range(6000000, 6001000),
dst_dir='/data/danbooru_webp',
max_workers=12,
)
webpii
WebPII Dataset
Split
Samples
Train
40,384
Test
4,481
Total
44,865
For more information, see webpii.github.io
e621-2024-webp-3.4M-recap-qwen2.5-vl-3B
RareConcepts/e621-2024-webp-3.4M
a derived, soft-filtered mirror of the original e621 2024 dump, kept in .webp with machine-generated captions for compact storage and high throughput training
tl;dr
based on: boxingscorpionbagel/e621-2024 (via the processed mirror NebulaeWis/e621-2024-webp-4Mpixel)
this repo: image shards in .webp (≈3.4M images implied by name)
captions: caption field
soft filter: captions produced by Qwen2.5-VL-3B-Instruct (permissive; may include some… See the full description on the dataset page: https://huggingface.co/datasets/RareConcepts/e621-2024-webp-3.4M-recap-qwen2.5-vl-3B.inet_cheat_288_webpe621_newest-webp-4Mpixel
E621 Newest 4M Re-encoded Dataset
This is the re-encoded dataset of deepghs/e621_newest. And all the resized images are maintained here.
There are 963455 images in total. The maximum ID of these images is 5753265. Last updated at 2025-08-09 00:50:53 JST.
This dataset only contains newest webp images, if you are looking for old webp images, just see NebulaeWis/e621-2024-webp-4Mpixel.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/e621_newest-webp-4Mpixel.rule34-webp-4Mpixel
Rule34 4M Re-encoded Dataset
This is the re-encoded dataset of deepghs/rule34_full. And all the resized images are maintained here.
There are 11336686 images in total. The maximum ID of these images is 13078768. Last updated at 2025-04-13 16:02:32 JST.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/rule34-webp-4Mpixel.e621-2024-webp-4Mpixel_indexIndex files of NebulaeWis/e621-2024-webp-4Mpixel.
You can download images from NebulaeWis/e621-2024-webp-4Mpixel with cheesechaser.
from cheesechaser.datapool import E621NewestWebpDataPool
pool = E621NewestWebpDataPool()
# download e621 #2010000-2010300, to directory /data/e621
pool.batch_download_to_directory(
resource_ids=range(2010000, 2010300),
dst_dir='/data/e621',
max_workers=12,
)
danbooru-202603-webp90-noresize
Danbooru 202603 WebP90 no-resize
This public, manually gated repository contains 10,767,652 verified images in 1,080 WebDataset tar files. Images were retained at their native dimensions and encoded as WebP quality 90. It combines a crawler materialization with the nyanko7/danbooru2023 source surface; crawler records take precedence on overlapping Danbooru post IDs.
This materialization was acquired for research reproducibility because a complete directly usable release was not… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/danbooru-202603-webp90-noresize.hafele-products-webp
bep40/hafele-products-webp
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('bep40/hafele-products-webp')
web-page-quality-edu-strict-blind-5196
Web Page Quality EDU Strict-Blind 5196
This dataset contains 5,196 English web pages sampled from
TeraflopAI/web-page-quality-labels
at revision aee0ebeff5bf09e9bd6881b9da581ec6ee77db5c, together with a new
strict-blind educational-quality judgment.
Sampling
The sample uses seed 20260927 and is stratified by the source edu_score:
Source score
Rows
0
1,000
1
1,000
2
1,000
3
1,000
4
1,000
5
196 (all available rows)
Strict-Blind… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/web-page-quality-edu-strict-blind-5196.WebPRMCollection_preference_pairPaper: arxiv.org/abs/2505.15277
e621-2024-webp-4Mpixel-webshart-indices
e621-2024-webp-4Mpixel webshart indices (shards 0-48)
Per-shard webshart index files for
NebulaeWis/e621-2024-webp-4Mpixel,
covering tar shards original/data-0000.tar through original/data-0048.tar
(193,966 samples). Each entry carries the tar member byte offset/length, image
geometry (width, height, aspect), and captions merged from
RareConcepts/e621-2024-metadata-4M-recap-qwen2.5-vl-7B
(qwen2.5-VL-7B recaptions).
Coverage is deliberately partial (first 49 of 1000 shards, built… See the full description on the dataset page: https://huggingface.co/datasets/webshart/e621-2024-webp-4Mpixel-webshart-indices.
