AndreiRabau/diffusion-llm
Diffusion-LLM
A small, experimental discrete diffusion language model built on GPT-2.
This repository contains checkpoints from an experiment in iterative text denoising. Instead of predicting only the next token from a left-to-right context, the model receives a corrupted sequence and predicts the original tokens across the whole sequence with bidirectional attention.
The checkpoints are intended for research and reproducibility. They are not drop-in transformers causal language models and should not be treated as production-ready text generators.
What is inside?
The model starts from `openai-community/gpt2` and adapts its Transformer backbone for discrete diffusion-style denoising:
- GPT-2 embeddings and Transformer blocks are retained.
- Causal attention is replaced with bidirectional attention.
- A learned diffusion-timestep embedding is added to every token embedding.
- The language-model head predicts the clean token at each position.
- Denoising is performed by repeatedly updating corrupted positions.
The training code adds two special tokens to the GPT-2 tokenizer: <|pad|> and <|mask|>. The vocabulary is resized before training, so the same tokenizer construction is required when loading a checkpoint.
Checkpoints
The uploaded weights follow this layout:
All listed experimental runs use the same main setup: GPT-2, 100 diffusion timesteps, maximum sequence length 64, learning rate 1e-5, and WikiText-2 (Salesforce/wikitext, wikitext-2-raw-v1). The 4-step runs use a rollout loss decay of 0.5.
Corruption strategies
The training pipeline supports three discrete corruption operators:
- Similar — replaces tokens with nearby tokens in the GPT-2 embedding space. This creates difficult, semantically close mistakes.
- Mask — replaces tokens with
<|mask|>, the classic masked-denoising condition. - Random — replaces tokens with random vocabulary entries.
The mixed checkpoints use these methods as a sequential training curriculum: 30 epochs of similar-token corruption, 30 epochs of masking, and 40 epochs of random replacement. Corruption probability is sampled between 0.01 and 0.95; the similar-token operator uses 20 nearest neighbours.
Quick start
Clone this repository so that the model wrapper and corruption utilities are available alongside the downloaded checkpoint:
git clone https://huggingface.co/AndreiRabau/diffusion-llm
cd diffusion-llm
pip install torch transformersThe following example loads the 4-step mixed checkpoint and denoises a sequence containing mask tokens:
from pathlib import Path
import torch
from transformers import AutoTokenizer
from model_wrappers import GPT2DiffusionTransformer
MODEL_NAME = "openai-community/gpt2"
CHECKPOINT = Path("4_step/mix_30_30_40_4_steps.pt")
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
tokenizer.add_special_tokens({
"pad_token": "<|pad|>",
"mask_token": "<|mask|>",
})
model = GPT2DiffusionTransformer.from_file_path(
file_path=CHECKPOINT,
model_name=MODEL_NAME,
num_diffusion_steps=100,
vocabulary_size=len(tokenizer),
device="cpu",
).eval()
text = "Diffusion models can refine [MASK] sequences over several steps."
inputs = tokenizer(text, return_tensors="pt")
# For a real masked input, create the sequence with tokenizer.mask_token_id.
# Here we replace one token to keep the example self-contained.
corrupted_ids = inputs["input_ids"].clone()
corrupted_positions = torch.zeros_like(corrupted_ids, dtype=torch.bool)
corrupted_ids[0, 4] = tokenizer.mask_token_id
corrupted_positions[0, 4] = True
reconstructed_ids = model.denoise(
corrupted_ids=corrupted_ids,
attention_mask=inputs["attention_mask"],
corrupted_positions=corrupted_positions,
num_iterations=4,
)
print(tokenizer.decode(reconstructed_ids[0], skip_special_tokens=True))The .pt files contain a PyTorch state_dict. Loading them therefore requires the repository's GPT2DiffusionTransformer implementation and the same model configuration used during training (num_diffusion_steps=100 and the resized tokenizer vocabulary).
Training and evaluation details
Training uses cross-entropy only on corrupted, non-padding positions. For multi-step training, each step feeds the model's predictions back into the next step, exposing the model to its own errors. At evaluation time, the denoiser uses confidence-guided refinement: the most confident corrupted positions are updated first, while less certain positions remain available for later iterations. This difference is intentional.
The included experiments evaluate on the test split of `cimec/lambada` with corruption rates from 25% to 95% and 2, 4, 5, 8, 10, 20, or 50 denoising iterations, depending on the run. Evaluation artifacts are stored in results/eval/ in the source project.
Intended use
This release is useful for:
- studying corruption policies for discrete diffusion language models;
- comparing one-step and iterative denoising rollouts;
- experimenting with confidence-based token refinement;
- reproducing the accompanying small-scale WikiText-2 experiments.
It is not intended for factual question answering, safety-critical generation, deployment, or direct comparison with large-scale diffusion LMs. The checkpoints inherit the limitations of GPT-2 and the small experimental training setup, including potential memorization, bias, and unstable outputs.
Limitations and open questions
- These are research checkpoints rather than a fully packaged inference pipeline.
- The
.ptformat stores weights only; configuration and tokenizer files are reconstructed by the loader. - The model was trained on short sequences (maximum length 64), so longer contexts are outside the validated setup.
- The corruption curriculum and the semantic-neighbour operator have not been established as generally superior to standard masked or random corruption.
- Reported results should be interpreted as exploratory measurements, not as a leaderboard claim.
Citation
If you use these checkpoints or the training setup, please cite this repository and the upstream GPT-2 work:
@misc{diffusion_llm_experiment,
title = {Diffusion-LLM: Experimental GPT-2 Discrete Denoising Checkpoints},
author = {Andrei Rabau},
year = {2026},
url = {https://huggingface.co/AndreiRabau/diffusion-llm}
}The implementation and experiment configurations are available in the source repository, including `model_wrappers/gpt2_diffusion_transformer_wrapper.py`, `train_utils/train.py`, and the files under `experiments/`.
