Team Ai
Modelpublic

JaydeepR/diffusion-lm-tinystories

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes60downloads
Model Card

# Diffusion LM — TinyStories

A masked-diffusion language model trained from scratch on the TinyStories dataset.

## Demo

[image]

## Architecture

ParamValue
Parameters~45M
Hidden dim512
Layers10
Heads8
FFN dim2048
Diffusion steps T128
Sequence length256
Vocab size26,000

## How it works

This is a masked diffusion language model. Instead of generating tokens left-to-right like a standard LM, it starts with a fully masked sequence and progressively unmasks tokens over T diffusion steps.

At each step the model predicts all masked tokens simultaneously, then re-masks the least confident predictions and repeats — gradually refining the output until the sequence is fully unmasked.

## Training

  • —Dataset: 1M TinyStories examples
  • —Train steps: 60,000
  • —Effective batch size: 64 (batch 32 × grad accum 2)
  • —Optimizer: AdamW
  • —Learning rate: 2e-4 with cosine schedule and 1,000 warmup steps
  • —Weight decay: 0.1
  • —Mixed precision: bf16
  • —Hardware: NVIDIA RTX 3090 (24GB)

## Evaluation

Val loss (cross-entropy on masked tokens, 20 batches of held-out TinyStories):

StepVal Loss
5,0006.0313
10,0005.9045
15,0005.6092
20,0004.4481
25,0003.8447
30,0003.6634
35,0003.5419
40,0003.3554
45,0003.2779
50,0003.1767
55,0003.1012
60,0003.1067

The loss drop between steps 15,000–25,000 reflects the model learning basic language structure. Convergence around 3.10 by step 55,000.

## Files

FileDescription
model.ptModel weights (PyTorch state dict)
config.jsonArchitecture hyperparameters
tokenizer/Byte-level BPE tokenizer
val_loss_history.jsonValidation loss curve
inference.gifVisualisation of progressive unmasking