CrucibleAI/ControlNetMediaPipeFace
5761.2k
1---2language:3- en4thumbnail: ''5tags:6- controlnet7- laion8- face9- mediapipe10- image-to-image11license: openrail12base_model: stabilityai/stable-diffusion-2-1-base13datasets:14- LAION-Face15- LAION16pipeline_tag: image-to-image17---18 19# ControlNet LAION Face Dataset20 21## Table of Contents:22- Overview: Samples, Contents, and Construction23- Usage: Downloading, Training, and Inference24- License25- Credits and Thanks26 27# Overview:28 29This dataset is designed to train a ControlNet with human facial expressions. It includes keypoints for pupils to allow gaze direction. Training has been tested on Stable Diffusion v2.1 base (512) and Stable Diffusion v1.5.30 31## Samples:32 33Cherry-picked from ControlNet + Stable Diffusion v2.1 Base34 35|Input|Face Detection|Output|36|:---:|:---:|:---:|37|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/happy_source.jpg">|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/happy_annotation.png">|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/happy_result.png">|38|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/neutral_source.jpg">|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/neutral_annotation.png">|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/neutral_result.png">|39|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/sad_source.jpg">|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/sad_annotation.png">|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/sad_result.png">|40|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/screaming_source.jpg">|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/screaming_annotation.png">|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/screaming_result.png">|41|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/sideways_source.jpg">|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/sideways_annotation.png">|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/sideways_result.png">|42|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/surprised_source.jpg">|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/surprised_annotation.png">|<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/surprised_result.png">|43 44Images with multiple faces are also supported:45 46<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/family_source.jpg">47 48<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/family_annotation.png">49 50<img src="https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/family_result.png">51 52 53## Dataset Contents:54 55- train_laion_face.py - Entrypoint for ControlNet training.56- laion_face_dataset.py - Code for performing dataset iteration. Cropping and resizing happens here.57- tool_download_face_targets.py - A tool to read metadata.json and populate the target folder.58- tool_generate_face_poses.py - The original file used to generate the source images. Included for reproducibility, but not required for training.59- training/laion-face-processed/prompt.jsonl - Read by laion_face_dataset. Includes prompts for the images.60- training/laion-face-processed/metadata.json - Excerpts from LAION for the relevant data. Also used for downloading the target dataset.61- training/laion-face-processed/source/xxxxxxxxx.jpg - Images with detections performed. Generated from the target images.62- training/laion-face-processed/target/xxxxxxxxx.jpg - Selected images from LAION Face.63 64## Dataset Construction:65 66Source images were generated by pulling slice 00000 from LAION Face and passing them through MediaPipe's face detector with special configuration parameters. 67 68The colors and line thicknesses used for MediaPipe are as follows:69 70```71f_thick = 272f_rad = 173right_iris_draw = DrawingSpec(color=(10, 200, 250), thickness=f_thick, circle_radius=f_rad)74right_eye_draw = DrawingSpec(color=(10, 200, 180), thickness=f_thick, circle_radius=f_rad)75right_eyebrow_draw = DrawingSpec(color=(10, 220, 180), thickness=f_thick, circle_radius=f_rad)76left_iris_draw = DrawingSpec(color=(250, 200, 10), thickness=f_thick, circle_radius=f_rad)77left_eye_draw = DrawingSpec(color=(180, 200, 10), thickness=f_thick, circle_radius=f_rad)78left_eyebrow_draw = DrawingSpec(color=(180, 220, 10), thickness=f_thick, circle_radius=f_rad)79mouth_draw = DrawingSpec(color=(10, 180, 10), thickness=f_thick, circle_radius=f_rad)80head_draw = DrawingSpec(color=(10, 200, 10), thickness=f_thick, circle_radius=f_rad)81 82iris_landmark_spec = {468: right_iris_draw, 473: left_iris_draw}83```84 85We have implemented a method named `draw_pupils` which modifies some functionality from MediaPipe. It exists as a stopgap until some pending changes are merged.86 87 88# Usage:89 90The containing ZIP file should be decompressed into the root of the ControlNet directory. The `train_laion_face.py`, `laion_face_dataset.py`, and other `.py` files should sit adjacent to `tutorial_train.py` and `tutorial_train_sd21.py`. We are assuming a checkout of the ControlNet repo at 0acb7e5, but there is no direct dependency on the repository.91 92## Downloading:93 94For copyright reasons, we cannot include the original target files. We have provided a script (tool_download_face_targets.py) which will read from training/laion-face-processed/metadata.json and populate the target folder. This file has no requirements, but will use tqdm if it is installed.95 96## Training:97 98When the targets folder is fully populated, training can be run on a machine with at least 24 gigabytes of VRAM. Our model was trained for 200 hours (four epochs) on an A6000.99 100```bash101python tool_add_control.py ./models/v1-5-pruned-emaonly.ckpt ./models/controlnet_sd15_laion_face.ckpt102python ./train_laion_face_sd15.py103```104 105## Inference:106 107We have provided `gradio_face2image.py`. Update the following two lines to point them to your trained model.108 109```110model = create_model('./models/cldm_v21.yaml').cpu() # If you fine-tune on SD2.1 base, this does not need to change.111model.load_state_dict(load_state_dict('./models/control_sd21_openpose.pth', location='cuda'))112```113 114The model has some limitations: while it is empirically better at tracking gaze and mouth poses than previous attempts, it may still ignore controls. Adding details to the prompt like, "looking right" can abate bad behavior. 115 116## 🧨 Diffusers117 118It is recommended to use the checkpoint with [Stable Diffusion 2.1 - Base](stabilityai/stable-diffusion-2-1-base) as the checkpoint has been trained on it.119Experimentally, the checkpoint can be used with other diffusion models such as dreamboothed stable diffusion.120 121To use with Stable Diffusion 1.5, insert `subfolder="diffusion_sd15"` into the from_pretrained arguments. A v1.5 half-precision variant is provided but untested.122 1231. Install `diffusers` and related packages:124```125$ pip install diffusers transformers accelerate126```127 1282. Run code:129```py130from PIL import Image131import numpy as np132import torch133from diffusers import StableDiffusionControlNetPipeline, ControlNetModel, UniPCMultistepScheduler134from diffusers.utils import load_image135 136image = load_image(137 "https://huggingface.co/CrucibleAI/ControlNetMediaPipeFace/resolve/main/samples_laion_face_dataset/family_annotation.png"138)139 140# Stable Diffusion 2.1-base:141controlnet = ControlNetModel.from_pretrained("CrucibleAI/ControlNetMediaPipeFace", torch_dtype=torch.float16, variant="fp16")142pipe = StableDiffusionControlNetPipeline.from_pretrained(143 "stabilityai/stable-diffusion-2-1-base", controlnet=controlnet, safety_checker=None, torch_dtype=torch.float16144)145# OR146# Stable Diffusion 1.5:147controlnet = ControlNetModel.from_pretrained("CrucibleAI/ControlNetMediaPipeFace", subfolder="diffusion_sd15")148pipe = StableDiffusionControlNetPipeline.from_pretrained("runwayml/stable-diffusion-v1-5", controlnet=controlnet, safety_checker=None)149 150pipe.scheduler = UniPCMultistepScheduler.from_config(pipe.scheduler.config)151 152# Remove if you do not have xformers installed153# see https://huggingface.co/docs/diffusers/v0.13.0/en/optimization/xformers#installing-xformers154# for installation instructions155pipe.enable_xformers_memory_efficient_attention()156pipe.enable_model_cpu_offload()157 158image = pipe("a happy family at a dentist advertisement", image=image, num_inference_steps=30).images[0]159image.save('./images.png')160```161 162 163# License:164 165### Source Images: (/training/laion-face-processed/source/)166This work is marked with CC0 1.0. To view a copy of this license, visit http://creativecommons.org/publicdomain/zero/1.0167 168### Trained Models:169Our trained ControlNet checkpoints are released under CreativeML Open RAIL-M.170 171### Source Code:172lllyasviel/ControlNet is licensed under the Apache License 2.0173 174Our modifications are released under the same license.175 176 177# Credits and Thanks:178 179Greatest thanks to Zhang et al. for ControlNet, Rombach et al. (StabilityAI) for Stable Diffusion, and Schuhmann et al. for LAION.180 181Sample images for this document were obtained from Unsplash and are CC0.182 183```184@misc{zhang2023adding,185 title={Adding Conditional Control to Text-to-Image Diffusion Models}, 186 author={Lvmin Zhang and Maneesh Agrawala},187 year={2023},188 eprint={2302.05543},189 archivePrefix={arXiv},190 primaryClass={cs.CV}191}192 193@misc{rombach2021highresolution,194 title={High-Resolution Image Synthesis with Latent Diffusion Models}, 195 author={Robin Rombach and Andreas Blattmann and Dominik Lorenz and Patrick Esser and Björn Ommer},196 year={2021},197 eprint={2112.10752},198 archivePrefix={arXiv},199 primaryClass={cs.CV}200}201 202@misc{schuhmann2022laion5b,203 title={LAION-5B: An open large-scale dataset for training next generation image-text models}, 204 author={Christoph Schuhmann and Romain Beaumont and Richard Vencu and Cade Gordon and Ross Wightman and Mehdi Cherti and Theo Coombes and Aarush Katta and Clayton Mullis and Mitchell Wortsman and Patrick Schramowski and Srivatsa Kundurthy and Katherine Crowson and Ludwig Schmidt and Robert Kaczmarczyk and Jenia Jitsev},205 year={2022},206 eprint={2210.08402},207 archivePrefix={arXiv},208 primaryClass={cs.CV}209}210```211 212This project was made possible by Crucible AI.