Team Ai
Modelpublic

YashNagraj75/Diffusion-Transformer

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes
Model Card

Diffusion Transformer (DiT) — CelebA-HQ Face Generation

A Diffusion Transformer (DiT) trained on CelebA-HQ for unconditional face image generation at 128x128 resolution. The model uses a Vision Transformer backbone in place of the conventional U-Net for the denoising network, operating in the latent space of a VQ-VAE.

Model Description

This is a two-stage pipeline:

  1. 1.Stage 1 — VQ-VAE: Compresses 128x128 RGB images into a 4-channel discrete latent space with codebook size 8192.
  2. 2.Stage 2 — DiT: A transformer-based denoising model that operates on flattened image patches in the VQ-VAE latent space.

DiT Architecture

The DiT (Peebles & Xie, 2023) replaces the U-Net backbone with a standard Vision Transformer (ViT) encoder. Each image latent is divided into non-overlapping patches, linearly embedded, and processed by a stack of transformer blocks with time-step conditioning via adaptive layer norm.

ParameterValue
Patch size2
Transformer layers12
Hidden dimension768
Attention heads12
Head dimension64
Time embedding dim768
Input resolution128x128 (latent: ~16x16x4)

VQ-VAE Architecture

ParameterValue
Latent channels (z)4
Codebook size8192
Down channels[128, 256, 384]
Downsampling stages2

Diffusion Process

ParameterValue
Timesteps (T)1000
Beta scheduleLinear, start=0.0001, end=0.02

Training Details

StageEpochsLRBatch size
VQ-VAE101e-54
DiT5001e-532
  • —Dataset: CelebA-HQ, center-cropped and resized to 128x128, normalized to [-1, 1]
  • —Data loaded from parquet files via a custom ParquetImageDataset

Generated Samples

The repository includes generated face samples in celebhq/samples/ (x0_*.jpg), produced by running the trained DiT in reverse diffusion from Gaussian noise.

Repository Contents

PathDescription
celeba.pyParquet-based CelebA-HQ dataloader
celeba/config.yamlFull training configuration
celebhq/dit_ckpt.pthTrained DiT checkpoint
celebhq/samples/Generated sample images

References

License

MIT