datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pixelprose_webp_512dan-webp-newe621-2024-webp-4MpixelDataset Description:
This is a processed version of the https://huggingface.co/datasets/boxingscorpionbagel/e621-2024 dataset, primarily prepared for personal use in future projects.
Therefore, for licensing and other legal information, please refer to the original project.
You can directly download tar file,or use https://deepghs.github.io/hfutils/main/api_doc/index/fetch.html#hf-tar-file-download to download anything .webp file you want.
The following modifications have been made to the… See the full description on the dataset page: https://huggingface.co/datasets/NebulaeWis/e621-2024-webp-4Mpixel.imagenet-w21-webp-wds
Dataset Summary
This is a copy of the full Winter21 release of ImageNet in webdataset tar format with WEBP encoded images. This release consists of 19167 classes, 2674 fewer classes than the original 21841 class Fall11 release of the full ImageNet.
The classes were removed due to these concerns: https://www.image-net.org/update-sep-17-2019.php
This is the same contents as https://huggingface.co/datasets/timm/imagenet-w21-wds but encoded in webp at ~56% of the size, shard count… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-w21-webp-wds.WebPerson-5M
🔥 Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval [EMNLP25 Main]
Tianlu Zheng*,
Yifan Zhang*,
Xiang An,
Ziyong Feng,
Kaicheng Yang†,
Qichunan Ding†,
📄 Paper | 💻 Github
✨ Web-Person Dataset
🔍 Person-Centric Image Filtering
We use the COYO700M dataset as our source of web-crawled images.
To curate high-quality person-centric images, we apply YOLOv11 to detect humans and extract bounding… See the full description on the dataset page: https://huggingface.co/datasets/Kaichengalex/WebPerson-5M.web-pinkyimagenet_288_webpA duplicate of the webdataset version of ImageNet-1k, except each image is resized so that the shorter side is 288 pixels and each image is WEBP compressed using a quality level of 85.
The compression ratio is roughly 5:1 compared to JPEG (~160 GB) and 25:1 compared to raw pixel storage (~800GB).
WebPerson-1M
🔥 Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval [EMNLP25 Main]
Tianlu Zheng*,
Yifan Zhang*,
Xiang An,
Ziyong Feng,
Kaicheng Yang†,
Qichunan Ding†,
📄 Paper | 💻 Github
✨ Web-Person Dataset
🔍 Person-Centric Image Filtering
We use the COYO700M dataset as our source of web-crawled images.
To curate high-quality person-centric images, we apply YOLOv11 to detect humans and extract bounding… See the full description on the dataset page: https://huggingface.co/datasets/Kaichengalex/WebPerson-1M.webpii
WebPII Dataset
Split
Samples
Train
40,384
Test
4,481
Total
44,865
For more information, see webpii.github.io
inet_cheat_288_webpdanbooru-202603-webp90-noresize
Danbooru 202603 WebP90 no-resize
This public, manually gated repository contains 10,767,652 verified images in 1,080 WebDataset tar files. Images were retained at their native dimensions and encoded as WebP quality 90. It combines a crawler materialization with the nyanko7/danbooru2023 source surface; crawler records take precedence on overlapping Danbooru post IDs.
This materialization was acquired for research reproducibility because a complete directly usable release was not… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/danbooru-202603-webp90-noresize.hafele-products-webp
bep40/hafele-products-webp
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('bep40/hafele-products-webp')
WebPRMCollection_preference_pairPaper: arxiv.org/abs/2505.15277
gelbooru-webp-subset-384px-dedupSubset of gelbooru_full dowscaled to 384px
Pruned using ViT-SO400M-16-SigLIP2-384 cosine similarity threshold 0.95
danbooru2023-webp-4Mpixel
Danbooru 2023 webp: A space-efficient version of Danbooru 2023
This dataset is a resized/re-encoded version of danbooru2023.
Which removed the non-image/truncated files and resize all of them into smaller size.
This dataset already be updated to latest_id = 7,832,883.
Thx to DeepGHS!
Notice: content of updates folder and deepghs/danbooru_newest-webp-4Mpixel have been merged to 2000~2999.tar, You can ignore all the content in updates folder safely!
Details
This… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-webp-4Mpixel.dan-new-webp-trainWebPRMCollection_checklist_generationWeb_page_Phishing
Web Page Phishing Detection — EDA
Student: Yonatane Ben-Aroch,Reichman UniversityDataset: Web Page Phishing Detection — KaggleDate: April 2026
Your browser does not support the video tag.
Overview
Web Page Phishing is when someone creates a fake website that looks real to steal your password or personal information. This dataset contains 11,430 URLs that were labeled as either legitimate or phishing, with 87 features extracted from each… See the full description on the dataset page: https://huggingface.co/datasets/yonatane22-bh/Web_page_Phishing.WebPRMCollection_checklist_generationWebPRMCollection_preference_pairinet1k_val_288_webpdanbooru2023-webp-4Mpixel-224The data set is just resized to 224*224
https://huggingface.co/datasets/KBlueLeaf/danbooru2023-webp-4Mpixel
Pseudo code for processing
def resize_image(file_path):
with Image.open(file_path) as img:
resized_img = img.resize((224, 224))
resized_img.save(file_path)
aic2026-keyframes-webpLSDIR_WEBP_q20LSDIR_WEBP_q0imagenet_288_webp_fgweb_prm_processed_dataDocOCR_N2__run_Baseer__webpLSDIR_WEBP_q100laion-aesthetics-webp90-min256px-noresize
LAION-Aesthetics WebP90 min256 no-resize
This repository contains 23,688,416 URL-retrieved image payloads in 5,248 WebDataset tar files. It was materialized because a suitable image-bearing release was needed for research reproducibility. The recorded source metadata family is laion/laion2B-en-aesthetic, and a fresh materialization can be made with img2dataset.
The acquisition used URL as the URL column, TEXT as the source caption, WebP quality 90, no fixed-resolution resize, a… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/laion-aesthetics-webp90-min256px-noresize.
