Team Ai
Modelpublic

blowing-up-groundhogs/eruku

sourceHugging Faceapache-2.0updated 10mo agoView on Hugging Face
0likes988downloads
Model Card

Eruku - Autoregressive Styled Text Image Generation

<p align="center"> <img src="https://img.shields.io/badge/CVPR-2025-blue" alt="CVPR 2025"> <img src="https://img.shields.io/badge/WACV-2026-green" alt="WACV 2026"> <img src="https://img.shields.io/badge/License-Apache%202.0-yellow" alt="License"> </p>

Eruku is a state-of-the-art autoregressive model for styled handwritten and typewritten text image generation. Given a style reference image and text to generate, it produces high-quality text images that faithfully replicate the input style.

๐ŸŒŸ Key Features

  • โ€”Zero-shot style transfer: No training required for new styles
  • โ€”No transcription required: Works with just a style image (transcription optional but helps)
  • โ€”Reliable generation: Proper EOG (End of Generation) mechanism prevents artifacts
  • โ€”Arbitrary length: Generate text of any length
  • โ€”High fidelity: Excellent style consistency and text readability
  • โ€”Classifier-Free Guidance: Fine control over generation quality

๐Ÿ“ฆ Installation

bash
pip install torch torchvision transformers diffusers einops pillow

๐Ÿš€ Quick Start

python
from transformers import AutoModel
from PIL import Image
import torch

# Load model
device = "cuda" if torch.cuda.is_available() else "cpu"
model = AutoModel.from_pretrained(
    "blowing-up-groundhogs/eruku", 
    trust_remote_code=True
)
model.to(device)
model.eval()

# Load a style image (handwritten/typewritten text sample)
style_image = Image.open("style_sample.png")

# Generate text in that style
result = model.generate_handwriting(
    style_image=style_image,
    gen_text="Hello, World!",
    style_text="",  # Optional: transcription of style image
    cfg_scale=1.25,  # Classifier-free guidance scale
)

# Save the result
result.save("generated.png")

๐Ÿ“– Detailed Usage

Input Format

The model takes three inputs:

  1. 1.Style Image (style_image): A PIL Image containing handwritten or typewritten text that serves as the style reference. The model will replicate this style.
  1. 1.Generation Text (gen_text): The text you want to render in the extracted style.
  1. 1.Style Text (style_text, optional): The transcription of the text in the style image. Providing this helps the model better understand the style, but it's not required.

Parameters

ParameterTypeDefaultDescription
style_imagePIL.ImageRequiredReference style image
gen_textstrRequiredText to generate
style_textstr""Optional transcription of style image
cfg_scalefloat1.25Classifier-free guidance scale
max_new_tokensint512Maximum generation tokens

CFG Scale Guide

  • โ€”1.0: No guidance (faster but may drift from prompt)
  • โ€”1.25: Recommended default - good balance
  • โ€”1.5-2.0: Stronger adherence to prompt
  • โ€”>2.0: May cause artifacts

๐Ÿ–ผ๏ธ Example Results

The model excels at:

  • โ€”Handwritten text in various styles (cursive, print, mixed)
  • โ€”Typewritten text with different fonts
  • โ€”Multi-language text (trained primarily on English)
  • โ€”Long text sequences

๐Ÿ“Š Model Architecture

Eruku combines:

  • โ€”T5-Large encoder-decoder for text understanding and autoregressive generation
  • โ€”VAE (Variational Autoencoder) for image encoding and decoding
  • โ€”Custom embeddings for style transfer and special tokens (SOS, SOG, EOG)

The model generates images autoregressively, predicting one latent slice at a time until it produces an EOG (End of Generation) token.

๐Ÿ”ง Advanced Usage

Lower-level API

For more control, you can use the lower-level methods:

python
import torch
from torchvision import transforms as T

# Prepare style image manually
style_img = Image.open("style.png").convert('RGB')
width, height = style_img.size
new_width = int(64 * width / height)
style_img = style_img.resize((new_width, 64), Image.LANCZOS)
style_tensor = T.ToTensor()(style_img).to(device)

# Get model inputs
inputs = model.get_model_inputs(
    style_img=[style_tensor],
    style_len=style_tensor.shape[-1],
    max_img_len=1024*1024
)

# Generate with full control
with torch.inference_mode():
    output_img, special_sequence = model.generate(
        decoder_inputs_embeds_vae=inputs['decoder_inputs_embeds'],
        style_text=["Style text here"],
        gen_text=["Text to generate"],
        cfg_scale=1.25,
        max_new_tokens=512
    )

๐Ÿ“š Citation

If you use Eruku in your research, please cite both papers:

bibtex
@InProceedings{pippi2025zeroshot,
    author    = {Pippi, Vittorio and Quattrini, Fabio and Cascianelli, Silvia and Tonioni, Alessio and Cucchiara, Rita},
    title     = {Zero-Shot Styled Text Image Generation, but Make It Autoregressive},
    booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    month     = {June},
    year      = {2025},
    pages     = {7910-7919}
}

@inproceedings{zaccagnino2026autoregressive,
    author = {Carmine Zaccagnino and Fabio Quattrini and Vittorio Pippi and Silvia Cascianelli and Alessio Tonioni and Rita Cucchiara},
    title = {Autoregressive Styled Text Image Generation, but Make it Reliable},
    booktitle = {Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision},
    month = {March},
    year = {2026}
}

๐Ÿ”— Links

๐Ÿ“œ License

This model is released under the Apache 2.0 License.

๐Ÿ™ Acknowledgments

  • โ€”T5: google-t5/t5-large
  • โ€”VAE: blowing-up-groundhogs/emuru_vae
  • โ€”Training datasets: IAM, CVL, RIMES, FontSquare