JaydeepR/diffusion-lm-tinystories
# Diffusion LM — TinyStories
A masked-diffusion language model trained from scratch on the TinyStories dataset.
## Demo
## Architecture
## How it works
This is a masked diffusion language model. Instead of generating tokens left-to-right like a standard LM, it starts with a fully masked sequence and progressively unmasks tokens over T diffusion steps.
At each step the model predicts all masked tokens simultaneously, then re-masks the least confident predictions and repeats — gradually refining the output until the sequence is fully unmasked.
## Training
- Dataset: 1M TinyStories examples
- Train steps: 60,000
- Effective batch size: 64 (batch 32 × grad accum 2)
- Optimizer: AdamW
- Learning rate: 2e-4 with cosine schedule and 1,000 warmup steps
- Weight decay: 0.1
- Mixed precision: bf16
- Hardware: NVIDIA RTX 3090 (24GB)
## Evaluation
Val loss (cross-entropy on masked tokens, 20 batches of held-out TinyStories):
The loss drop between steps 15,000–25,000 reflects the model learning basic language structure. Convergence around 3.10 by step 55,000.
## Files
