open-diffusion
pexels-photos-janpf
Images:
There are approximately 130K images, borrowed from pexels.com.
Thanks to those folks for curating a wonderful resource.
There are millions more images on pexels. These particular ones were selected by
the list of urls at https://github.com/janpf/self-supervised-multi-task-aesthetic-pretraining/blob/main/dataset/urls.txt .
The filenames are based on the md5 hash of each image.
Download From here or from pexels.com: You choose
For those people who like… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf.cc12m-cleaned
CC12m-cleaned
This dataset builds on two others: The Conceptual Captions 12million dataset, which lead to the LLaVa captioned subset done by
CaptionEmporium
(The latter is the same set, but swaps out the (Conceptual Captions 12million) often-useless alt-text captioning for decent ones_
I have then used the llava captions as a base, and used the detailed descrptions to filter out
images with things like watermarks, artist signatures, etc.
I have also manually thrown out all… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned.eval_diffusion-v2-augThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/open-cloth/eval_diffusion-v2-aug.pexels-janpf-sharp
Overview
This is a strict subset of
https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf
I have attempted to throw out all shots with heavy bokeh. (so, this is the "sharp" focus dataset)
I have also attempted to throw out all the black-and-white photos.
I decided to create a whole "new" dataset, rather than creating a set filter as I
have done previously, because I think this dataset may become my new "base" dataset.
So I will most likely focus my refining… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-janpf-sharp.pexels-woman-solo
Overview
Around 8,000 hand-selected high-resolution 4k images of "a woman", suitable for both training, and "pre-training" of AI models.
Images are all real-world realism based.
Background
I have been having difficulty training my early-stage txt2img model with GOOD, HIGH-RES images of what "a woman" is.
Up until now, I have just been throwing a large number of random high-res images with "woman" in the auto-captioned details.
NOW, however, I have hand-selected a bunch of… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-woman-solo.woman-wipThis dataset is just a collaborative area for now, to trim down a candidate set of images
