Team Ai
Modelpublic

Disty0/sote-diffusion-cascade-decoder_alpha0

sourceHugging Faceotherupdated 3y agoView on Hugging Face
0likes27downloads
README.md105 linesDownload Raw Back to root
1---2pipeline_tag: text-to-image3license: other4license_name: stable-cascade-nc-community5license_link: LICENSE6---7 8# SoteDiffusion Cascade9 10Anime finetune of Stable Cascade Decoder.  11No commercial use thanks to StabilityAI.  12 13## Code Example14 15```shell16pip install diffusers17```18 19```python20import torch21from diffusers import StableCascadeDecoderPipeline, StableCascadePriorPipeline22 23prompt = "newest, 1girl, solo, cat ears, looking at viewer, blush, light smile,"24negative_prompt = "very displeasing, worst quality, monochrome, sketch, fat, child,"25 26prior = StableCascadePriorPipeline.from_pretrained("Disty0/sote-diffusion-cascade_alpha0", torch_dtype=torch.float16)27decoder = StableCascadeDecoderPipeline.from_pretrained("Disty0/sote-diffusion-cascade-decoder_alpha0", torch_dtype=torch.float16)28 29prior.enable_model_cpu_offload()30prior_output = prior(31    prompt=prompt,32    height=1024,33    width=1024,34    negative_prompt=negative_prompt,35    guidance_scale=7.0,36    num_images_per_prompt=1,37    num_inference_steps=4038)39 40decoder.enable_model_cpu_offload()41decoder_output = decoder(42    image_embeddings=prior_output.image_embeddings,43    prompt=prompt,44    negative_prompt=negative_prompt,45    guidance_scale=1.546    output_type="pil",47    num_inference_steps=1048).images[0]49decoder_output.save("cascade.png")50```51 52## Dataset53 54Used the same dataset as Disty0/sote-diffusion-cascade-decoder_pre-alpha0.  55Trained with 98K~ images.  56 57## Training:58 59**GPU used for training**: 1x AMD RX 7900 XTX 24GB  60 61**Software used**: https://github.com/2kpr/StableCascade  62 63### Config:64```65experiment_id: sotediffusion-sc-b_3b66model_version: 3B67dtype: bfloat1668use_fsdp: False69 70batch_size: 171grad_accum_steps: 172updates: 9800073backup_every: 204874save_every: 102475warmup_updates: 10076 77lr: 4.0e-678optimizer_type: Adafactor79adaptive_loss_weight: True80stochastic_rounding: True81 82image_size: 76883multi_aspect_ratio: [1/1, 1/2, 1/3, 2/3, 3/4, 1/5, 2/5, 3/5, 4/5, 1/6, 5/6, 9/16]84shift: 485 86checkpoint_path: /mnt/DataSSD/AI/SoteDiffusion/StableCascade/87output_path: /mnt/DataSSD/AI/SoteDiffusion/StableCascade/88webdataset_path: file:/mnt/DataSSD/AI/anime_image_dataset/best/newest_best-{0000..0001}.tar89 90effnet_checkpoint_path: /mnt/DataSSD/AI/models/sd-cascade/effnet_encoder.safetensors91stage_a_checkpoint_path: /mnt/DataSSD/AI/models/sd-cascade/stage_a.safetensors92generator_checkpoint_path: /mnt/DataSSD/AI/SoteDiffusion/StableCascade/sotediffusion-sc_3b-stage_b.safetensors93```94 95 96## Limitations and Bias97 98### Bias99 100- This model is intended for anime illustrations.  101  Realistic capabilites are not tested at all.    102 103### Limitations104- Far shot eyes are still bad thanks to the heavy latent compression.  105