Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01opendiffusionai /pexels-photos-janpf Images: There are approximately 130K images, borrowed from pexels.com. Thanks to those folks for curating a wonderful resource. There are millions more images on pexels. These particular ones were selected by the list of urls at https://github.com/janpf/self-supervised-multi-task-aesthetic-pretraining/blob/main/dataset/urls.txt . The filenames are based on the md5 hash of each image. Download From here or from pexels.com: You choose For those people who like… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf.text-to-image100K<n<1M45 likes680 downloads9mo agoHugging Face02opendiffusionai /cc12m-cleaned CC12m-cleaned This dataset builds on two others: The Conceptual Captions 12million dataset, which lead to the LLaVa captioned subset done by CaptionEmporium (The latter is the same set, but swaps out the (Conceptual Captions 12million) often-useless alt-text captioning for decent ones_ I have then used the llava captions as a base, and used the detailed descrptions to filter out images with things like watermarks, artist signatures, etc. I have also manually thrown out all… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned.imagetext-to-image1M<n<10M13 likes311 downloads2y agoHugging Face03open-cloth /eval_diffusion-v2-augThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.state": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/open-cloth/eval_diffusion-v2-aug.tabularrobotics10K<n<100K0 likes138 downloads5mo agoHugging Face04opendiffusionai /pexels-janpf-sharp Overview This is a strict subset of https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf I have attempted to throw out all shots with heavy bokeh. (so, this is the "sharp" focus dataset) I have also attempted to throw out all the black-and-white photos. I decided to create a whole "new" dataset, rather than creating a set filter as I have done previously, because I think this dataset may become my new "base" dataset. So I will most likely focus my refining… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-janpf-sharp.text10K<n<100K3 likes93 downloads2mo agoHugging Face05opendiffusionai /pexels-woman-solo Overview Around 8,000 hand-selected high-resolution 4k images of "a woman", suitable for both training, and "pre-training" of AI models. Images are all real-world realism based. Background I have been having difficulty training my early-stage txt2img model with GOOD, HIGH-RES images of what "a woman" is. Up until now, I have just been throwing a large number of random high-res images with "woman" in the auto-captioned details. NOW, however, I have hand-selected a bunch of… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-woman-solo.text-to-image1K<n<10K3 likes56 downloads2y agoHugging Face06opendiffusionai /cc12m-4mp-realistic Overview This is a hand-selected subset of our larger attempts to filter the well known CC12M dataset. This one focuses on large (4 megapixels) images that are real world, high quality images, and the captioning specifically matches either "A man" or "A woman". Note that I did not have the diskspace/time to go through the ENTIRE set. It was perhaps only from the first 2 million of our CC12M-cleaned subset. If an effort were made to go through the entire 4mp image set, there might be… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp-realistic.imagetext-to-image10K<n<100K22 likes52 downloads2y agoHugging Face07opendiffusionai /woman-wipThis dataset is just a collaborative area for now, to trim down a candidate set of images image1 likes36 downloads2y agoHugging Face08opendiffusionai /cc12m-2mp-realistic Overview A subset of the "CC12m" dataset. Size varying between 2mp <= x < 4mp. Anything actually 4mp or larger will be in one of our "4mp" datasets. This dataset is created for if you basically need more images and dont mind a little less quality. Quality I have filtered out as many watermarks, etc. as possible using AI models. I have also thrown out stupid black-and-white photos, because they poison normal image prompting. This is NOT HAND CURATED, unlike some… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-realistic.image100K<n<1M4 likes31 downloads2y agoHugging Face09opendiffusionai /cc12m-1mp_plus-realistic cc12m-1mp_plus-realistic A filtering down of the full CC12M dataset, to have the following characteristics: At least 1024x1024 pixels in size "Realistic". No paintings, digital art, monochrome, or surreal stuff. Also discard multi-image as much as possible Ideally, no signed or watermarked images. (but there will certainly be some left) Captions The caption types available are a bit different from some of our other ones. Currently available are: caption_llava… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-1mp_plus-realistic.image100K<n<1M4 likes30 downloads10mo agoHugging Face10open-cloth /eval_diffusion-cloth-LR-asyncThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 15, "total_frames": 23605, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:15" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/open-cloth/eval_diffusion-cloth-LR-async.tabularrobotics10K<n<100K0 likes30 downloads5mo agoHugging Face11open-cloth /eval_diffusion-clothThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 2, "total_frames": 2886, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/open-cloth/eval_diffusion-cloth.tabularrobotics1K<n<10K0 likes29 downloads5mo agoHugging Face12opendiffusionai /cc12m-a_woman Description This dataset is a convenience subset of https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned/ I did a quick grep for "A woman", and then HAND-CURATED the results. That means I threw out anything with watermarks, site branding, or pretty much anything else I deemed would get in the way of ML training. I also only chose images that had clear, sharp camera focus on the main subject. So these are high-quality images. At present, I have only done a few thousand. I… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-a_woman.image10K<n<100K27 likes24 downloads2y agoHugging Face13opendiffusionai /laion2b-en-aesthetic-square-cleaned Overview A subset of our opendiffusionai/laion2b-en-aesthetic-square, which is itself a subset of the widely known "Laion2b-en-aesthetic" dataset. However, the original had only the website alt-tags for captions. I have added decent AI captioning, via the "moondream2" model. Additionally, I have stripped out a bunch of watermarked junk, and weeded out 20k duplicate images. When I use it for training, I will be trimming out additional things like non-realistic images, painting… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-cleaned.image100K<n<1M11 likes23 downloads9mo agoHugging Face14opendiffusionai /laion2b-squareish-1536px Overview This dataset is a subset of laion2b-en-aesthetic dataset. Its purpose is for general-case (photographic) model training. Its aspect ratios are (1 <= ratio < 5/4 ) so that images can be auto-cropped to square safely. Additionally, all images are at least 1536 pixels tall. This is because I wanted a dataset that would be very high quality to use for models that are 768x768 There should be close to 80k 8k images here. Note that you can CHOOSE between "moondream"… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish-1536px.image10K<n<100K3 likes23 downloads2y agoHugging Face15opendiffusionai /pexels-woman-croppable pexels-woman-croppable An extract from our larger "pexels 130k images" set. Around 6500 images. Useful for training text-to-image models. But we already have a bunch of subsets, why another one? This is 6k images that, while not originally square cropped, are HAND-SELECTED to be square-crop clean. The provided crawl.sh util script will handle automatically downloading and cropping them from pexels.com Why Square-Crop When training a model, you must have all images… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-woman-croppable.imagetext-to-image1K<n<10K0 likes23 downloads9mo agoHugging Face16opendiffusionai /laion2b-23ish-woman-solo Overview All images have a woman in them, solo, at APPROXIMATELY 2:3 aspect ratio. These images are HUMAN CURATED. I have personally gone through every one at least once. Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training There should be a little over 15k images here. Note that there is a wide variety of body sizes, from size 0, to perhaps size 18 There are also THREE choices of captions: the really bad "alt… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-23ish-woman-solo.image10K<n<100K19 likes22 downloads2y agoHugging Face17opendiffusionai /laion2b-aesthetic-squareish-captionsThis dataset contains image captions generated from LAION2B-en-aesthetic-square. We started with ~300K images after size filtering (2.5k max w/h), a portion of the images were skipped due to inaccessible URLs. The captions were generated over ~30 hours using Qwen3-VL-30B-A3B-Instruct on 1xH100 running SGLang with the prompt Describe the content of the provided image in detail, in plaintext. Do not make assumptions. Do not use special formatting. Avoid purple prose. Total samples: 209,141… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-aesthetic-squareish-captions.image100K<n<1M1 likes22 downloads11mo agoHugging Face18opendiffusionai /laion2b-en-aesthetic-square-human Overview This dataset is a HAND-CURATED version of our laion2b-en-aesthetic-square-cleaned dataset. It has at least "a man" or "a woman" in it. Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training There should be a little over 8k images here. Details It consists of an initial extraction of all images that had "a man" or "a woman" in the moondream caption. I then filtered out all "statue" or… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-human.image1K<n<10K8 likes19 downloads2y agoHugging Face19opendiffusionai /laion2b-45ish-1120px Overview This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic, selected for aspect ratio, and with better captioning. It is a GENERAL CASE dataset. Percentages of images with humans in it, is approximately only 30% I also filtered out all non-realistic images. This is intended to be a "real world" dataset. Approximate image count in this dataset is around 80k. On-disk size is around 45G. This is NOT individually human filtered. Batch culling only.… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-45ish-1120px.image10K<n<100K3 likes17 downloads2y agoHugging Face20opendiffusionai /cc12m-4mp Parent dataset https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned Contents I had uploaded a prior version of this "4mp" set, that was just partial. Now I have grabbed all of them available as of 2024 December. This latest upload references ALL files in the "CC12M" dataset, that are at least 4 megapixel in size, and are not one of the known "stock images" copyrighted sites. (and were available this month) I have done some additional filtering out of any… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp.image100K<n<1M3 likes16 downloads2y agoHugging Face21diffusion-guidance-ku-hf /sdxl_open-image-preferences_0text1K<n<10K0 likes15 downloads2y agoHugging Face22opendiffusionai /laion2b-mixed-1024px-human Overview This dataset is a selective merge of some other of our datasets. Mainly, I pulled human-centric, real-world photos from the following datasets: opendiffusionai/laion2b-45ish-1120px opendiffusionai/laion2b-squareish-1024px opendiffusionai/laion2b-23ish-1216px As such, it is "mixed" aspect ratio. The very smallest height ones are from our "squarish" set, so are at least 1024px tall. However, the other ones with longer rations, have an appropriately longer minimum… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-mixed-1024px-human.image10K<n<100K3 likes15 downloads2y agoHugging Face23opendiffusionai /laion2b-en-aesthetic-square Contents This is a pretty raw filter of https://huggingface.co/datasets/laion/laion2B-en-aesthetic I just filtered for "is image perfectly square, AND is image at least 1024x1024 pixels" Approximate image count is a bit over 300k Updated 2025/01/24 I just found out there are a bunch of watermarked sites in here. So much for aesthetically chosen :( So I filtered out a bunch of the "stock image" sites, just by looking at url strings. Update 2025/01/31… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square.image100K<n<1M1 likes14 downloads2y agoHugging Face24opendiffusionai /cc12m-small-squarish-simple Contents of dataset This is a "general purposes dataset", of images that are 512px to 1024px in size, if I recall correctly (in contrast to the "2mp" and "4mp" datasets) Also, they are "squarish", which means that, even if they are not precisely square, the image looks fine if you hard-crop it to force square. Additionally, images have been hand-culled to throw out anything I considered bad for AI training. Additionally, they were AI-culled to have a "simple photographic… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-small-squarish-simple.image10K<n<100K0 likes13 downloads1y agoHugging Face25opendiffusionai /laion2b-squareish-1024px Overview This dataset is a subset of laion2b-en-aesthetic dataset. Its purpose is for general-case (photographic) model training. Its aspect ratios are (1 <= ratio < 5/4 ) so that images can be auto-cropped to square safely. Additionally, all images are at least 1024 pixels tall. There should be close to 250k images here. Total on-disk size is approximately 118G We have a very similar dataset that is 1536px in size. However, that is only 80k in size, so if you need waaay more… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish-1024px.image100K<n<1M1 likes12 downloads2y agoHugging Face26opendiffusionai /laion2b-squareish General purpose dataset (realistic) This is a merging of our laion2b-squareish-1024px and laion2b-squareish-1536px datasets, WITH some additional filtering. It has a little de-duplication, and some extra watermark removal. It is MOSTLY realistic. I'm sharing this, because I am actively using it in my SD1.5 training experiments. I needed specifically square(ish) rather than our other nice, but mixed aspect-ratio datasets. Plus, I needed a LARGE one. So, here it is!… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish.image100K<n<1M0 likes12 downloads1y agoHugging Face27opendiffusionai /cc12m-xlsd-512px What Assorted images from our CC12M filtered sets, centercropped and then resized to 512x512 pro-actively Why This probably wont be useful to people outside the org, but just in case... here you go! The wierd naming matches the MD5 checksum of the ORIGINAL FULL-SIZED image. Internally i sort all my images like this. So I can do filtering or autotagging on the mini-images, then directly apply it to the originals. image100K<n<1M0 likes9 downloads2y agoHugging Face28opendiffusionai /cc12m-2mp-squareish Overview A subset of the "CC12m" dataset. Around 37k images, varying between 2mp <= x < 4mp. Anything actually 4mp or larger will be in one of our "4mp" datasets. This dataset is created for if you basically need more images and dont mind a little less quality. Why "Squareish"? I noticed that SD1.5 really, really likes square images to train on. These are mostly NOT EXACTLY SQUARE. However, most training programs should have an auto-crop capability. This dataset is of… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-squareish.image10K<n<100K0 likes8 downloads2y agoHugging Face29french-open-data /les-lieux-de-diffusion-du-spectacle-vivant-a-paris Les lieux de diffusion du spectacle vivant à Paris [!NOTE] Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Les lieux de diffusion du spectacle vivant à Paris qui est disponible à l'adresse https://www.data.gouv.fr/datasets/53699850a3a729239d204f11 Description Ce jeu de données référence les lieux de diffusion régulière ou occasionnelle du spectacle vivant à Paris. Structure des données : nom de l'établissement ; adresse… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/les-lieux-de-diffusion-du-spectacle-vivant-a-paris.0 likes6 downloads11mo agoHugging Face30opendiffusionai /laion2b-23ish-1216px Overview This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic, selected for aspect ratio, and with better captioning. Approximate image count is around 250k. 23ish, 1216px I picked out the images that are portrait aspect ratio of 2:3, or a little wider (Because images that are a little too wide, can be safely cropped narrower) I also picked a minimum height of 1216 pixels, because that is what 1024x1024 pixelcount converted to 2:3 looks like.… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-23ish-1216px.image100K<n<1M0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.