Team Ai
Datasetpublic

whosouravsharma/text-to-image-diffusiondb-2M

DiffusionDB text-to-image subset A cleaned, safety-filtered image-prompt dataset for training a text-to-image model, built from DiffusionDB. Built on Hugging Face Jobs directly from poloclub/diffusiondb. It covers part_id 1-20 (20,000 source images) before filtering. The same content is also kept on the 20k-subset branch. Load it with: load_dataset("whosouravsharma/text-to-image-diffusiondb-2M") Note on the repo name: despite "2M" in the name, this is a small slice of… See the full description on the dataset page: https://huggingface.co/datasets/whosouravsharma/text-to-image-diffusiondb-2M.

sourceHugging Facecc0-1.0updated 2mo agoView on Hugging Face
1likes267downloads
Dataset Card

DiffusionDB text-to-image subset

A cleaned, safety-filtered image-prompt dataset for training a text-to-image model, built from DiffusionDB.

Built on Hugging Face Jobs directly from poloclub/diffusiondb. It covers part_id 1-20 (20,000 source images) before filtering. The same content is also kept on the 20k-subset branch.

Load it with:

python
load_dataset("whosouravsharma/text-to-image-diffusiondb-2M")

Note on the repo name: despite "2M" in the name, this is a small slice of DiffusionDB's 2000 parts, not the full ~2,000,000-image collection.

Filtering applied

  • —Dropped missing/empty or very short (<3 char) prompts
  • —Dropped images below 256px on either dimension
  • —Dropped NSFW-flagged rows (imagensfw >= 0.5 or promptnsfw >= 0.5)
  • —Dropped exact-duplicate images (SHA256 hash match)

Columns

image, prompt, part_id, seed, step, cfg, sampler, width, height, image_nsfw, prompt_nsfw. user_name/timestamp from the original DiffusionDB metadata are intentionally excluded.

A standalone metadata.parquet (all rows, both splits, with a split column) is also included at the repo root.

Source & license

Derived from poloclub/diffusiondb, CC0-1.0 (public domain). Wang et al., 2022, arXiv:2210.14896.