Team Ai
Modelpublic

jdopensource/JoyAI-Image-Edit-Diffusers

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
24likes59kdownloads
README.md186 linesDownload Raw Back to root
1---2license: apache-2.03language:4- en5- zh6pipeline_tag: image-to-image7---8<h1 align="center">JoyAI-Image-Edit<br><sub><sup>Awakening Spatial Intelligence in Unified Multimodal Understanding and Generation</sup></sub></h1>9 10<div align="center">11 12[![Report PDF](https://img.shields.io/badge/Report-PDF-red)](https://joyai-image.s3.cn-north-1.jdcloud-oss.com/JoyAI-Image.pdf)13[![Project](https://img.shields.io/badge/Project-JoyAI--Image-333399)](https://github.com/jd-opensource/JoyAI-Image)14[![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Checkpoint-JoyAI--Image--Edit--Diffusers-yellow)](https://huggingface.co/jdopensource/JoyAI-Image-Edit-Diffusers)&#160;15[![ModelScope](https://img.shields.io/badge/%F0%9F%A4%96%20ModelScope-JoyAI--Image--Edit-624aff)](https://modelscope.cn/models/jd-opensource/JoyAI-Image-Edit-Diffusers)&#160;16[![Demo](https://img.shields.io/badge/%F0%9F%9A%80%20Demo-Spatial--Edit-orange)](https://huggingface.co/spaces/stevengrove/JoyAI-Image-Edit-Space)&#160;17[![License](https://img.shields.io/badge/License-Apache_2.0-blue.svg)](LICENSE)18 19 20</div>21 22## 🐶 JoyAI-Image-Edit23 24JoyAI-Image-Edit is a multimodal foundation model specialized in instruction-guided image editing. It enables precise and controllable edits by leveraging strong spatial understanding, including scene parsing, relational grounding, and instruction decomposition, allowing complex modifications to be applied accurately to specified regions.25 26## 🚀 Quick Start27 28**Requirements**: Python >= 3.10, CUDA-capable GPU29 30### Install31Note: JoyImageEditPipeline will be included in the next official diffusers release (>0.38.0). Until then, install from source as shown above.32 33```bash34pip install torch transformers torchvision35pip install git+https://github.com/huggingface/diffusers.git36```37### Running with Diffusers38```python39import torch40from PIL import Image41 42from diffusers import JoyImageEditPipeline43 44pipeline = JoyImageEditPipeline.from_pretrained("jdopensource/JoyAI-Image-Edit-Diffusers")45pipeline.to(torch.bfloat16)46pipeline.to("cuda")47pipeline.set_progress_bar_config(disable=None)48print("pipeline loaded")49 50img_path = "./test_images/input.png"51prompt = "Remove the construction structure from the top of the crane."52image = Image.open(img_path).convert("RGB")53 54inputs = {55    "image": image,56    "prompt": prompt,57    "generator": torch.manual_seed(0),58    "num_inference_steps": 40,59    "guidance_scale": 4.0,60}61 62print("run pipeline...")63 64with torch.inference_mode():65    output = pipeline(**inputs)66    image = output.images[0]67    image.save("joyai_image_edit_output.png")68    print("image saved.")69```70 71 72## More Usages73 74### Spatial Editing Reference75 76JoyAI-Image supports three spatial editing prompt patterns: **Object Move**, **Object Rotation**, and **Camera Control**. For the most stable behavior, we recommend following the prompt templates below as closely as possible.77 78#### 1. Object Move79 80Use this pattern when you want to move a target object into a specified region.81 82**Prompt template:**83 84```text85Move the <object> into the red box and finally remove the red box.86```87 88**Rules:**89 90* Replace `<object>` with a clear description of the target object to be moved.91* The **red box** indicates the target destination in the image.92* The phrase **"finally remove the red box"** means the guidance box should not appear in the final edited result.93 94**Example:**95 96```text97Move the board into the red box and finally remove the red box.98```99<p align="center">100  <img src="test_images/input1.png" width="40%" />101  <img src="test_images/output1_predicted.png" width="40%" />102</p>103 104#### 2. Object Rotation105 106Use this pattern when you want to rotate an object to a specific canonical view.107 108**Prompt template:**109 110```text111Rotate the <object> to show the <view> side view.112```113 114**Supported `<view>` values:**115 116* `front`117* `right`118* `left`119* `rear`120* `front right`121* `front left`122* `rear right`123* `rear left`124 125**Rules:**126 127* Replace `<object>` with a clear description of the object to rotate.128* Replace `<view>` with one of the supported directions above.129* This instruction is intended to change the **object orientation**, while keeping the object identity and surrounding scene as consistent as possible.130 131**Examples:**132 133```text134Rotate the dog to show the left side view.135```136<p align="center">137  <img src="test_images/input2.png" width="40%" />138  <img src="test_images/output2_predicted.png" width="40%" />139</p>140 141 142#### 3. Camera Control143 144Use this pattern when you want to change only the camera viewpoint while keeping the 3D scene itself unchanged.145 146**Prompt template:**147 148```text149Move the camera.150- Camera rotation: Yaw {y_rotation}°, Pitch {p_rotation}°.151- Camera zoom: in/out/unchanged.152- Keep the 3D scene static; only change the viewpoint.153```154 155**Rules:**156 157* `{y_rotation}` specifies the yaw rotation angle in degrees.158* `{p_rotation}` specifies the pitch rotation angle in degrees.159* `Camera zoom` must be one of:160 161  * `in`162  * `out`163  * `unchanged`164* The last line is important: it explicitly tells the model to preserve the 3D scene content and geometry, and only adjust the camera viewpoint.165 166**Examples:**167```text168Move the camera.169- Camera rotation: Yaw 0.0°, Pitch -15.0°.170- Camera zoom: unchanged.171- Keep the 3D scene static; only change the viewpoint.172```173<p align="center">174  <img src="test_images/input3.png" width="40%" />175  <img src="test_images/output3_predicted.png" width="40%" />176</p>177 178## License Agreement179 180JoyAI-Image is licensed under Apache 2.0.181 182## ☎️  We're Hiring!183 184We are actively hiring Research Scientists, AI Infra Engineers, and Interns to join us in building next-generation generative foundation models and bringing them into real-world applications. If you’re interested, please send your resume to: [huanghaoyang.ocean@jd.com](mailto:huanghaoyang.ocean@jd.com)185 186