whosouravsharma/text-to-image-diffusiondb-2M
DiffusionDB text-to-image subset A cleaned, safety-filtered image-prompt dataset for training a text-to-image model, built from DiffusionDB. Built on Hugging Face Jobs directly from poloclub/diffusiondb. It covers part_id 1-20 (20,000 source images) before filtering. The same content is also kept on the 20k-subset branch. Load it with: load_dataset("whosouravsharma/text-to-image-diffusiondb-2M") Note on the repo name: despite "2M" in the name, this is a small slice of… See the full description on the dataset page: https://huggingface.co/datasets/whosouravsharma/text-to-image-diffusiondb-2M.
DiffusionDB text-to-image subset
A cleaned, safety-filtered image-prompt dataset for training a text-to-image model, built from DiffusionDB.
Built on Hugging Face Jobs directly from poloclub/diffusiondb. It covers part_id 1-20 (20,000 source images) before filtering. The same content is also kept on the 20k-subset branch.
Load it with:
load_dataset("whosouravsharma/text-to-image-diffusiondb-2M")Note on the repo name: despite "2M" in the name, this is a small slice of DiffusionDB's 2000 parts, not the full ~2,000,000-image collection.
Filtering applied
- Dropped missing/empty or very short (<3 char) prompts
- Dropped images below 256px on either dimension
- Dropped NSFW-flagged rows (imagensfw >= 0.5 or promptnsfw >= 0.5)
- Dropped exact-duplicate images (SHA256 hash match)
Columns
image, prompt, part_id, seed, step, cfg, sampler, width, height, image_nsfw, prompt_nsfw. user_name/timestamp from the original DiffusionDB metadata are intentionally excluded.
A standalone metadata.parquet (all rows, both splits, with a split column) is also included at the repo root.
Source & license
Derived from poloclub/diffusiondb, CC0-1.0 (public domain). Wang et al., 2022, arXiv:2210.14896.
