Disty0/sote-diffusion-cascade-decoder_alpha0
027
1---2pipeline_tag: text-to-image3license: other4license_name: stable-cascade-nc-community5license_link: LICENSE6---7 8# SoteDiffusion Cascade9 10Anime finetune of Stable Cascade Decoder. 11No commercial use thanks to StabilityAI. 12 13## Code Example14 15```shell16pip install diffusers17```18 19```python20import torch21from diffusers import StableCascadeDecoderPipeline, StableCascadePriorPipeline22 23prompt = "newest, 1girl, solo, cat ears, looking at viewer, blush, light smile,"24negative_prompt = "very displeasing, worst quality, monochrome, sketch, fat, child,"25 26prior = StableCascadePriorPipeline.from_pretrained("Disty0/sote-diffusion-cascade_alpha0", torch_dtype=torch.float16)27decoder = StableCascadeDecoderPipeline.from_pretrained("Disty0/sote-diffusion-cascade-decoder_alpha0", torch_dtype=torch.float16)28 29prior.enable_model_cpu_offload()30prior_output = prior(31 prompt=prompt,32 height=1024,33 width=1024,34 negative_prompt=negative_prompt,35 guidance_scale=7.0,36 num_images_per_prompt=1,37 num_inference_steps=4038)39 40decoder.enable_model_cpu_offload()41decoder_output = decoder(42 image_embeddings=prior_output.image_embeddings,43 prompt=prompt,44 negative_prompt=negative_prompt,45 guidance_scale=1.546 output_type="pil",47 num_inference_steps=1048).images[0]49decoder_output.save("cascade.png")50```51 52## Dataset53 54Used the same dataset as Disty0/sote-diffusion-cascade-decoder_pre-alpha0. 55Trained with 98K~ images. 56 57## Training:58 59**GPU used for training**: 1x AMD RX 7900 XTX 24GB 60 61**Software used**: https://github.com/2kpr/StableCascade 62 63### Config:64```65experiment_id: sotediffusion-sc-b_3b66model_version: 3B67dtype: bfloat1668use_fsdp: False69 70batch_size: 171grad_accum_steps: 172updates: 9800073backup_every: 204874save_every: 102475warmup_updates: 10076 77lr: 4.0e-678optimizer_type: Adafactor79adaptive_loss_weight: True80stochastic_rounding: True81 82image_size: 76883multi_aspect_ratio: [1/1, 1/2, 1/3, 2/3, 3/4, 1/5, 2/5, 3/5, 4/5, 1/6, 5/6, 9/16]84shift: 485 86checkpoint_path: /mnt/DataSSD/AI/SoteDiffusion/StableCascade/87output_path: /mnt/DataSSD/AI/SoteDiffusion/StableCascade/88webdataset_path: file:/mnt/DataSSD/AI/anime_image_dataset/best/newest_best-{0000..0001}.tar89 90effnet_checkpoint_path: /mnt/DataSSD/AI/models/sd-cascade/effnet_encoder.safetensors91stage_a_checkpoint_path: /mnt/DataSSD/AI/models/sd-cascade/stage_a.safetensors92generator_checkpoint_path: /mnt/DataSSD/AI/SoteDiffusion/StableCascade/sotediffusion-sc_3b-stage_b.safetensors93```94 95 96## Limitations and Bias97 98### Bias99 100- This model is intended for anime illustrations. 101 Realistic capabilites are not tested at all. 102 103### Limitations104- Far shot eyes are still bad thanks to the heavy latent compression. 105 