datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
diffusiondbDiffusionDB is the first large-scale text-to-image prompt dataset. It contains 2
million images generated by Stable Diffusion using prompts and hyperparameters
specified by real users. The unprecedented scale and diversity of this
human-actuated dataset provide exciting research opportunities in understanding
the interplay between prompts and generative models, detecting deepfakes, and
designing human-AI interaction tools to help users more easily use these models.diffusion_db_dedupe_from50k_train
Dataset Card for "diffusion_db_dedupe_from50k_train"
More Information needed
diffusiondb_2m_random_50k
Dataset Card for "diffusiondb_2m_random_50k"
More Information needed
text-to-image-diffusiondb-2M
DiffusionDB text-to-image subset
A cleaned, safety-filtered image-prompt dataset for training a text-to-image
model, built from DiffusionDB.
Built on Hugging Face Jobs directly from poloclub/diffusiondb. It covers
part_id 1-20 (20,000 source images) before filtering. The same content is
also kept on the 20k-subset branch.
Load it with:
load_dataset("whosouravsharma/text-to-image-diffusiondb-2M")
Note on the repo name: despite "2M" in the name, this is a small slice of… See the full description on the dataset page: https://huggingface.co/datasets/whosouravsharma/text-to-image-diffusiondb-2M.diffusion_db_dedup_from50k_train_v2
Dataset Card for "diffusion_db_dedup_from50k_train_v2"
More Information needed
diffusiondb-pixelartDiffusionDB is the first large-scale text-to-image prompt dataset. It contains 2
million images generated by Stable Diffusion using prompts and hyperparameters
specified by real users. The unprecedented scale and diversity of this
human-actuated dataset provide exciting research opportunities in understanding
the interplay between prompts and generative models, detecting deepfakes, and
designing human-AI interaction tools to help users more easily use these models.diffusiondbDiffusionDB is the first large-scale text-to-image prompt dataset. It contains 2
million images generated by Stable Diffusion using prompts and hyperparameters
specified by real users. The unprecedented scale and diversity of this
human-actuated dataset provide exciting research opportunities in understanding
the interplay between prompts and generative models, detecting deepfakes, and
designing human-AI interaction tools to help users more easily use these models.DiffusionDB-300k-processeddiffusiondb_random_10k
Dataset Card for "diffusiondb_random_10k"
More Information needed
diffusiondb_2m_first_5k_canny
Dataset Card for "diffusiondb_2m_first_5k_canny"
Process diffusiondb 2m first 5k canny to edges by Canny algorithm.
More Information needed
diffusion_db_dedup_from10k_train_v2
Dataset Card for "diffusion_db_dedup_from10k_train_v2"
More Information needed
diffusiondb_random_10k_zh_v1
Dataset Card for "diffusiondb_random_10k_zh_v1"
svjack/diffusiondb_random_10k_zh_v1 is a dataset that random sample 10k English samples from diffusiondb and use NMT translate them into Chinese with some corrections.
it used to train stable diffusion models in svjack/Stable-Diffusion-FineTuned-zh-v0
svjack/Stable-Diffusion-FineTuned-zh-v1
svjack/Stable-Diffusion-FineTuned-zh-v2
And is the data support of https://github.com/svjack/Stable-Diffusion-Chinese-Extend which is a fine tune… See the full description on the dataset page: https://huggingface.co/datasets/svjack/diffusiondb_random_10k_zh_v1.diffusion_db_dedupe_from50k_val
Dataset Card for "diffusion_db_dedupe_from50k_val"
More Information needed
diffusiondb_2m_random_5k_blur_61KS
Dataset Card for "diffusiondb_2m_random_5k_blur_61KS"
More Information needed
diffusiondb_2m_random_5k_blur
Dataset Card for "diffusiondb_2m_first_5k_blur"
More Information needed
diffusiondb_2m_random_50kdiffusion_db_dedup_from50k_val_v2
Dataset Card for "diffusion_db_dedup_from50k_val_v2"
More Information needed
diffusiondb_random_10k_zh_v2_deepl
Dataset Card for "diffusiondb_random_10k_zh_v2_deepl"
More Information needed
canny_diffusiondb
Canny DiffusionDB
This dataset is the DiffusionDB dataset that is transformed using Canny transformation.
You can see samples below 👇
Sample:
Original Image:
Transformed Image:
Caption:
"a small wheat field beside a forest, studio lighting, golden ratio, details, masterpiece, fine art, intricate, decadent, ornate, highly detailed, digital painting, octane render, ray tracing reflections, 8 k, featured, by claude monet and vincent van gogh "Below you can find a small script used… See the full description on the dataset page: https://huggingface.co/datasets/jax-diffusers-event/canny_diffusiondb.diffusiondb_2m_first_5k_canny_zh_512
Dataset Card for "diffusiondb_2m_first_5k_canny_zh_512"
More Information needed
diffusiondb-pixelart-v2
DiffusionDB-Pixelart
Dataset Summary
This is a subset of the DiffusionDB 2M dataset which has been turned into pixel-style art.
DiffusionDB is the first large-scale text-to-image prompt dataset. It contains 14 million images generated by Stable Diffusion using prompts and hyperparameters specified by real users.
DiffusionDB is publicly available at 🤗 Hugging Face Dataset.
Supported Tasks and Leaderboards
The unprecedented scale and diversity of this… See the full description on the dataset page: https://huggingface.co/datasets/Clawffice/diffusiondb-pixelart-v2.diffusiondb_only_somediffusion_db_5k_train_v2
Dataset Card for "diffusion_db_5k_train_v2"
More Information needed
diffusion_db_5k_train_v1
Dataset Card for "diffusion_db_5k_train_v1"
More Information needed
diffusiondb_ner
Description
Extended dataset infered by the name entity recognition model en_ner_prompting. This model has been trained on hand-annotated prompts from poloclub/diffusiondb.
This dataset is hence infered by this model and can comprise mistakes, especially on certain categories (cf. model card).
The entities comprise 7 main categories and 11 subcategories for a total of 16 categories, extracted from a topic analysis made with BERTopic.
The topic analysis can be explored the… See the full description on the dataset page: https://huggingface.co/datasets/teo-sanchez/diffusiondb_ner.diffusiondb-prompt-upscale
Dataset Card for "diffusiondb-prompt-upscale"
More Information needed
DiffusionDB-300kdiffusiondb_2m_first_5k_canny
Dataset Card for "diffusiondb_2m_first_5k_canny"
More Information needed
diffusiondbDiffusionDB is the first large-scale text-to-image prompt dataset. It contains 2
million images generated by Stable Diffusion using prompts and hyperparameters
specified by real users. The unprecedented scale and diversity of this
human-actuated dataset provide exciting research opportunities in understanding
the interplay between prompts and generative models, detecting deepfakes, and
designing human-AI interaction tools to help users more easily use these models.diffusiondbDiffusionDB is the first large-scale text-to-image prompt dataset. It contains 2
million images generated by Stable Diffusion using prompts and hyperparameters
specified by real users. The unprecedented scale and diversity of this
human-actuated dataset provide exciting research opportunities in understanding
the interplay between prompts and generative models, detecting deepfakes, and
designing human-AI interaction tools to help users more easily use these models.
