diffusers/Florence2-image-Annotator
120
1---2library_name: diffusers3---4 5# Florence-2 Image Annotator6 7A custom [Modular Diffusers](https://huggingface.co/docs/diffusers/modular_diffusers/overview) block that uses [Florence-2](https://huggingface.co/docs/transformers/model_doc/florence2) for image annotation tasks like segmentation, object detection, and captioning.8 9## How to use10 11 12```python13import torch14from diffusers import ModularPipeline15from diffusers.utils import load_image16 17# Load the block18image_annotator = ModularPipeline.from_pretrained(19 "diffusers/Florence2-image-Annotator",20 trust_remote_code=True21)22image_annotator.load_components(torch_dtype=torch.bfloat16)23image_annotator.to("cuda")24 25# Load an image26image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/tasks/car.jpg")27image = image.resize((1024, 1024))28 29# Generate a segmentation mask30output = image_annotator(31 image=image,32 annotation_task="<REFERRING_EXPRESSION_SEGMENTATION>",33 annotation_prompt="the car",34 annotation_output_type="mask_image",35).images[0]36output.save("car-mask.png")37```38 39### Compose with Inpainting Pipeline40 41```python42from diffusers import ModularPipeline43 44# Load the annotator45image_annotator = ModularPipeline.from_pretrained(46 "diffusers/Florence2-image-Annotator",47 trust_remote_code=True48)49 50# Get an inpainting workflow and insert the annotator51# repo_id = .. # you can use SDXL/flux/qwen any pipeline support Inpaint52inpaint_blocks = ModularPipeline.from_pretrained(repo_id).blocks.get_workflow("inpainting")53inpaint_blocks.sub_blocks.insert("image_annotator", image_annotator.blocks, 0)54 55# Initialize the combined pipeline56pipe = inpaint_blocks.init_pipeline()57pipe.load_components(torch_dtype=torch.float16, device="cuda")58 59# Inpaint with automatic mask generation60output = pipe(61 prompt=prompt,62 image=image,63 annotation_task="<REFERRING_EXPRESSION_SEGMENTATION>",64 annotation_prompt="the car",65 annotation_output_type="mask_image",66 num_inference_steps=30,67).images[0]68output.save("inpainted-car.png")69```70 71## Supported Tasks72 73| Task | Description |74|------|-------------|75| `<OD>` | Object detection |76| `<REFERRING_EXPRESSION_SEGMENTATION>` | Segment specific objects based on text |77| `<CAPTION>` | Generate image caption |78| `<DETAILED_CAPTION>` | Generate detailed caption |79| `<MORE_DETAILED_CAPTION>` | Generate very detailed caption |80| `<DENSE_REGION_CAPTION>` | Caption different regions |81| `<CAPTION_TO_PHRASE_GROUNDING>` | Ground phrases to regions |82| `<OPEN_VOCABULARY_DETECTION>` | Detect objects from open vocabulary |83 84## Output Types85 86| Type | Description |87|------|-------------|88| `mask_image` | Black and white mask image |89| `mask_overlay` | Mask overlaid on original image |90| `bounding_box` | Bounding boxes drawn on image |91 92## Inputs93 94| Parameter | Type | Required | Default | Description |95|-----------|------|----------|---------|-------------|96| `image` | `PIL.Image` | Yes | - | Image to annotate |97| `annotation_task` | `str` | No | `<REFERRING_EXPRESSION_SEGMENTATION>` | Task to perform |98| `annotation_prompt` | `str` | Yes | - | Text prompt for the task |99| `annotation_output_type` | `str` | No | `mask_image` | Output format |100 101## Outputs102 103| Parameter | Type | Description |104|-----------|------|-------------|105| `mask_image` | `PIL.Image` | Generated mask (when output type is `mask_image`) |106| `image` | `PIL.Image` | Annotated image (when output type is `mask_overlay` or `bounding_box`) |107| `annotations` | `dict` | Raw annotation predictions |108 109## Components110 111This block uses the following models from [florence-community/Florence-2-base-ft](https://huggingface.co/florence-community/Florence-2-base-ft):112 113- `image_annotator`: `Florence2ForConditionalGeneration`114- `image_annotator_processor`: `AutoProcessor`115 116## Learn More117 118- [Building Custom Blocks Guide](https://huggingface.co/docs/diffusers/modular_diffusers/custom_blocks)119- [Modular Diffusers Overview](https://huggingface.co/docs/diffusers/modular_diffusers/overview)120- [Modular Diffusers Custom Blocks Collection](https://huggingface.co/collections/diffusers/modular-diffusers-custom-blocks)