Team Ai
Apppublic

radames/Text2Human-API

sourceHugging Faceupdated 3y agoView on Hugging Face
1likes
README.md256 linesDownload Raw Back to Text2Human
1# Text2Human - Official PyTorch Implementation2 3<!-- <img src="./doc_images/overview.jpg" width="96%" height="96%"> -->4 5This repository provides the official PyTorch implementation for the following paper:6 7**Text2Human: Text-Driven Controllable Human Image Generation**</br>8[Yuming Jiang](https://yumingj.github.io/), [Shuai Yang](https://williamyang1991.github.io/), [Haonan Qiu](http://haonanqiu.com/), [Wayne Wu](https://dblp.org/pid/50/8731.html), [Chen Change Loy](https://www.mmlab-ntu.com/person/ccloy/) and [Ziwei Liu](https://liuziwei7.github.io/)</br>9In ACM Transactions on Graphics (Proceedings of SIGGRAPH), 2022.10 11From [MMLab@NTU](https://www.mmlab-ntu.com/index.html) affliated with S-Lab, Nanyang Technological University and SenseTime Research.12 13<table>14<tr>15    <td><img src="assets/1.png" width="100%"/></td>16    <td><img src="assets/2.png" width="100%"/></td>17    <td><img src="assets/3.png" width="100%"/></td>18    <td><img src="assets/4.png" width="100%"/></td>19</tr>20<tr>21    <td align='center' width='24%'>The lady wears a short-sleeve T-shirt with pure color pattern, and a short and denim skirt.</td>22    <td align='center' width='24%'>The man wears a long and floral shirt, and long pants with the pure color pattern.</td>23    <td align='center' width='24%'>A lady is wearing a sleeveless pure-color shirt and long jeans</td>24    <td align='center' width='24%'>The man wears a short-sleeve T-shirt with the pure color pattern and a short pants with the pure color pattern.</td>25<tr>26</table>27 28[**[Project Page]**](https://yumingj.github.io/projects/Text2Human.html) | [**[Paper]**](https://arxiv.org/pdf/2205.15996.pdf) | [**[Dataset]**](https://github.com/yumingj/DeepFashion-MultiModal) | [**[Demo Video]**](https://youtu.be/yKh4VORA_E0)29 30 31## Updates32 33- [05/2022] Paper and demo video are released.34- [05/2022] Code is released.35- [05/2022] This website is created.36 37## Installation38**Clone this repo:**39```bash40git clone https://github.com/yumingj/Text2Human.git41cd Text2Human42```43**Dependencies:**44 45All dependencies for defining the environment are provided in `environment/text2human_env.yaml`.46We recommend using [Anaconda](https://docs.anaconda.com/anaconda/install/) to manage the python environment:47```bash48conda env create -f ./environment/text2human_env.yaml49conda activate text2human50conda install -c huggingface tokenizers=0.9.451conda install -c huggingface transformers=4.0.052conda install -c conda-forge sentence-transformers=2.0.053```54 55If it doesn't work, you may need to install the following packages on your own:56  - Python 3.657  - PyTorch 1.7.158  - CUDA 10.159  - [sentence-transformers](https://huggingface.co/sentence-transformers) 2.0.060  - [tokenizers](https://pypi.org/project/tokenizers/) 0.9.461  - [transformers](https://huggingface.co/docs/transformers/installation) 4.0.062 63## (1) Dataset Preparation64 65In this work, we contribute a large-scale high-quality dataset with rich multi-modal annotations named [DeepFashion-MultiModal](https://github.com/yumingj/DeepFashion-MultiModal) Dataset.66Here we pre-processed the raw annotations of the original dataset for the task of text-driven controllable human image generation. The pre-processing pipeline consists of:67  - align the human body in the center of the images according to the human pose68  - fuse the clothing color and clothing fabric annotations into one texture annotation69  - do some annotation cleaning and image filtering70  - split the whole dataset into the training set and testing set71 72You can download our processed dataset from this [Google Drive](https://drive.google.com/file/d/1KIoFfRZNQVn6RV_wTxG2wZmY8f2T_84B/view?usp=sharing). If you want to access the raw annotations, please refer to the [DeepFashion-MultiModal](https://github.com/yumingj/DeepFashion-MultiModal) Dataset.73 74After downloading the dataset, unzip the file and put them under the dataset folder with the following structure:75```76./datasets77├── train_images78    ├── xxx.png79    ...80    ├── xxx.png81    └── xxx.png82├── test_images83    % the same structure as in train_images84├── densepose85    % the same structure as in train_images86├── segm87    % the same structure as in train_images88├── shape_ann89    ├── test_ann_file.txt90    ├── train_ann_file.txt91    └── val_ann_file.txt92└── texture_ann93    ├── test94        ├── lower_fused.txt95        ├── outer_fused.txt96        └── upper_fused.txt97    ├── train98        % the same files as in test99    └── val100        % the same files as in test101```102 103## (2) Sampling104 105### Inference Notebook106<img src="https://colab.research.google.com/assets/colab-badge.svg" height=22.5></a></br>107Coming soon.108 109 110### Pretrained Models111 112Pretrained models can be downloaded from this [Google Drive](https://drive.google.com/file/d/1VyI8_AbPwAUaZJPaPba8zxsFIWumlDen/view?usp=sharing). Unzip the file and put them under the dataset folder with the following structure:113```114pretrained_models115├── index_pred_net.pth116├── parsing_gen.pth117├── parsing_token.pth118├── sampler.pth119├── vqvae_bottom.pth120└── vqvae_top.pth121```122 123### Generation from Paring Maps124You can generate images from given parsing maps and pre-defined texture annotations:125```python126python sample_from_parsing.py -opt ./configs/sample_from_parsing.yml127```128The results are saved in the folder `./results/sampling_from_parsing`.129 130### Generation from Poses131You can generate images from given human poses and pre-defined clothing shape and texture annotations:132```python133python sample_from_pose.py -opt ./configs/sample_from_pose.yml134```135 136**Remarks**: The above two scripts generate images without language interactions. If you want to generate images using texts, you can use the notebook or our user interface.137 138### User Interface139 140```python141python ui_demo.py142```143<img src="./assets/ui.png" width="100%">144 145The descriptions for shapes should follow the following format:146```147<gender>, <sleeve length>, <length of lower clothing>, <outer clothing type>, <other accessories1>, ...148 149Note: The outer clothing type and accessories can be omitted.150 151Examples:152man, sleeveless T-shirt, long pants153woman, short-sleeve T-shirt, short jeans154```155 156The descriptions for textures should follow the following format:157```158<upper clothing texture>, <lower clothing texture>, <outer clothing texture>159 160Note: Currently, we only support 5 types of textures, i.e., pure color, stripe/spline, plaid/lattice, 161    floral, denim. Your inputs should be restricted to these textures.162```163 164## (3) Training Text2Human165 166### Stage I: Pose to Parsing167Train the parsing generation network. If you want to skip the training of this network, you can download our pretrained model from [here](https://drive.google.com/file/d/1MNyFLGqIQcOMg_HhgwCmKqdwfQSjeg_6/view?usp=sharing).168```python169python train_parsing_gen.py -opt ./configs/parsing_gen.yml170```171 172### Stage II: Parsing to Human173 174**Step 1: Train the top level of the hierarchical VQVAE.**175We provide our pretrained model [here](https://drive.google.com/file/d/1TwypUg85gPFJtMwBLUjVS66FKR3oaTz8/view?usp=sharing). This model is trained by:176```python177python train_vqvae.py -opt ./configs/vqvae_top.yml178```179 180**Step 2: Train the bottom level of the hierarchical VQVAE.**181We provide our pretrained model [here](https://drive.google.com/file/d/15hzbY-RG-ILgzUqqGC0qMzlS4OayPdRH/view?usp=sharing). This model is trained by:182```python183python train_vqvae.py -opt ./configs/vqvae_bottom.yml184```185 186**Stage 3 & 4: Train the sampler with mixture-of-experts.** To train the sampler, we first need to train a model to tokenize the parsing maps. You can access our pretrained parsing maps [here](https://drive.google.com/file/d/1GLHoOeCP6sMao1-R63ahJMJF7-J00uir/view?usp=sharing).187```python188python train_parsing_token.py -opt ./configs/parsing_token.yml189```190 191With the parsing tokenization model, the sampler is trained by:192```python193python train_sampler.py -opt ./configs/sampler.yml194```195Our pretrained sampler is provided [here](https://drive.google.com/file/d/1OQO_kG2fK7eKiG1VJH1OL782X71UQAmS/view?usp=sharing).196 197**Stage 5: Train the index prediction network.**198We provide our pretrained index prediction network [here](https://drive.google.com/file/d/1rqhkQD-JGd7YBeIfDvMV-vjfbNHpIhYm/view?usp=sharing). It is trained by:199```python200python train_index_prediction.py -opt ./configs/index_pred_net.yml201```202 203 204**Remarks**: In the config files, we use the path to our models as the required pretrained models. If you want to train the models from scratch, please replace the path to your own one. We set the numbers of the training epochs as large numbers and you can choose the best epoch for each model. For your reference, our pretrained parsing generation network is trained for 50 epochs, top-level VQVAE is trained for 135 epochs, bottom-level VQVAE is trained for 70 epochs, parsing tokenization network is trained for 20 epochs, sampler is trained for 95 epochs, and the index prediction network is trained for 70 epochs.205 206## (4) Results207 208Please visit our [Project Page](https://yumingj.github.io/projects/Text2Human.html#results) to view more results.</br>209You can select the attribtues to customize the desired human images.210[<img src="./assets/results.png" width="90%">211](https://yumingj.github.io/projects/Text2Human.html#results)212 213## DeepFashion-MultiModal Dataset214 215<img src="./assets/dataset_logo.png" width="90%">216 217In this work, we also propose **DeepFashion-MultiModal**, a large-scale high-quality human dataset with rich multi-modal annotations. It has the following properties:2181. It contains 44,096 high-resolution human images, including 12,701 full body human images.2192. For each full body images, we **manually annotate** the human parsing labels of 24 classes.2203. For each full body images, we **manually annotate** the keypoints.2214. We extract DensePose for each human image.2225. Each image is **manually annotated** with attributes for both clothes shapes and textures.2236. We provide a textual description for each image.224 225<img src="./assets/dataset_overview.png" width="100%">226 227Please refer to [this repo](https://github.com/yumingj/DeepFashion-MultiModal) for more details about our proposed dataset.228 229## TODO List230 231- [ ] Release 1024x512 version of Text2Human.232- [ ] Train the Text2Human using [SHHQ dataset](https://stylegan-human.github.io/).233 234## Citation235 236If you find this work useful for your research, please consider citing our paper:237 238```bibtex239@article{jiang2022text2human,240  title={Text2Human: Text-Driven Controllable Human Image Generation},241  author={Jiang, Yuming and Yang, Shuai and Qiu, Haonan and Wu, Wayne and Loy, Chen Change and Liu, Ziwei},242  journal={ACM Transactions on Graphics (TOG)},243  volume={41},244  number={4},245  articleno={162},246  pages={1--11},247  year={2022},248  publisher={ACM New York, NY, USA},249  doi={10.1145/3528223.3530104},250}251```252 253## Acknowledgments254 255Part of the code is borrowed from [unleashing-transformers](https://github.com/samb-t/unleashing-transformers), [taming-transformers](https://github.com/CompVis/taming-transformers) and [mmsegmentation](https://github.com/open-mmlab/mmsegmentation).256