Team Ai
Modelpublic

AndreiRabau/diffusion-llm

sourceHugging Faceupdated 21d agoView on Hugging Face
0likes
Model Card

Diffusion-LLM

A small, experimental discrete diffusion language model built on GPT-2.

This repository contains checkpoints from an experiment in iterative text denoising. Instead of predicting only the next token from a left-to-right context, the model receives a corrupted sequence and predicts the original tokens across the whole sequence with bidirectional attention.

The checkpoints are intended for research and reproducibility. They are not drop-in transformers causal language models and should not be treated as production-ready text generators.

What is inside?

The model starts from `openai-community/gpt2` and adapts its Transformer backbone for discrete diffusion-style denoising:

  • —GPT-2 embeddings and Transformer blocks are retained.
  • —Causal attention is replaced with bidirectional attention.
  • —A learned diffusion-timestep embedding is added to every token embedding.
  • —The language-model head predicts the clean token at each position.
  • —Denoising is performed by repeatedly updating corrupted positions.

The training code adds two special tokens to the GPT-2 tokenizer: <|pad|> and <|mask|>. The vocabulary is resized before training, so the same tokenizer construction is required when loading a checkpoint.

Checkpoints

The uploaded weights follow this layout:

PathTraining corruptionRefinement steps during training
default.ptInitial/reference checkpoint—
1_step/similar_100.ptSemantically similar replacements1
1_step/mask_100.ptMask-token corruption1
1_step/random_100.ptRandom-token replacement1
1_step/mix_30_30_40.pt30 training epochs similar → 30 mask → 40 random1
4_step/similar_100_4_steps.ptSemantically similar replacements4
4_step/mask_100_4_steps.ptMask-token corruption4
4_step/random_100_4_steps.ptRandom-token replacement4
4_step/mix_30_30_40_4_steps.pt30 training epochs similar → 30 mask → 40 random4

All listed experimental runs use the same main setup: GPT-2, 100 diffusion timesteps, maximum sequence length 64, learning rate 1e-5, and WikiText-2 (Salesforce/wikitext, wikitext-2-raw-v1). The 4-step runs use a rollout loss decay of 0.5.

Corruption strategies

The training pipeline supports three discrete corruption operators:

  1. 1.Similar — replaces tokens with nearby tokens in the GPT-2 embedding space. This creates difficult, semantically close mistakes.
  2. 2.Mask — replaces tokens with <|mask|>, the classic masked-denoising condition.
  3. 3.Random — replaces tokens with random vocabulary entries.

The mixed checkpoints use these methods as a sequential training curriculum: 30 epochs of similar-token corruption, 30 epochs of masking, and 40 epochs of random replacement. Corruption probability is sampled between 0.01 and 0.95; the similar-token operator uses 20 nearest neighbours.

Quick start

Clone this repository so that the model wrapper and corruption utilities are available alongside the downloaded checkpoint:

bash
git clone https://huggingface.co/AndreiRabau/diffusion-llm
cd diffusion-llm
pip install torch transformers

The following example loads the 4-step mixed checkpoint and denoises a sequence containing mask tokens:

python
from pathlib import Path

import torch
from transformers import AutoTokenizer

from model_wrappers import GPT2DiffusionTransformer


MODEL_NAME = "openai-community/gpt2"
CHECKPOINT = Path("4_step/mix_30_30_40_4_steps.pt")

tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
tokenizer.add_special_tokens({
    "pad_token": "<|pad|>",
    "mask_token": "<|mask|>",
})

model = GPT2DiffusionTransformer.from_file_path(
    file_path=CHECKPOINT,
    model_name=MODEL_NAME,
    num_diffusion_steps=100,
    vocabulary_size=len(tokenizer),
    device="cpu",
).eval()

text = "Diffusion models can refine [MASK] sequences over several steps."
inputs = tokenizer(text, return_tensors="pt")

# For a real masked input, create the sequence with tokenizer.mask_token_id.
# Here we replace one token to keep the example self-contained.
corrupted_ids = inputs["input_ids"].clone()
corrupted_positions = torch.zeros_like(corrupted_ids, dtype=torch.bool)
corrupted_ids[0, 4] = tokenizer.mask_token_id
corrupted_positions[0, 4] = True

reconstructed_ids = model.denoise(
    corrupted_ids=corrupted_ids,
    attention_mask=inputs["attention_mask"],
    corrupted_positions=corrupted_positions,
    num_iterations=4,
)

print(tokenizer.decode(reconstructed_ids[0], skip_special_tokens=True))

The .pt files contain a PyTorch state_dict. Loading them therefore requires the repository's GPT2DiffusionTransformer implementation and the same model configuration used during training (num_diffusion_steps=100 and the resized tokenizer vocabulary).

Training and evaluation details

Training uses cross-entropy only on corrupted, non-padding positions. For multi-step training, each step feeds the model's predictions back into the next step, exposing the model to its own errors. At evaluation time, the denoiser uses confidence-guided refinement: the most confident corrupted positions are updated first, while less certain positions remain available for later iterations. This difference is intentional.

The included experiments evaluate on the test split of `cimec/lambada` with corruption rates from 25% to 95% and 2, 4, 5, 8, 10, 20, or 50 denoising iterations, depending on the run. Evaluation artifacts are stored in results/eval/ in the source project.

Intended use

This release is useful for:

  • —studying corruption policies for discrete diffusion language models;
  • —comparing one-step and iterative denoising rollouts;
  • —experimenting with confidence-based token refinement;
  • —reproducing the accompanying small-scale WikiText-2 experiments.

It is not intended for factual question answering, safety-critical generation, deployment, or direct comparison with large-scale diffusion LMs. The checkpoints inherit the limitations of GPT-2 and the small experimental training setup, including potential memorization, bias, and unstable outputs.

Limitations and open questions

  • —These are research checkpoints rather than a fully packaged inference pipeline.
  • —The .pt format stores weights only; configuration and tokenizer files are reconstructed by the loader.
  • —The model was trained on short sequences (maximum length 64), so longer contexts are outside the validated setup.
  • —The corruption curriculum and the semantic-neighbour operator have not been established as generally superior to standard masked or random corruption.
  • —Reported results should be interpreted as exploratory measurements, not as a leaderboard claim.

Citation

If you use these checkpoints or the training setup, please cite this repository and the upstream GPT-2 work:

bibtex
@misc{diffusion_llm_experiment,
  title  = {Diffusion-LLM: Experimental GPT-2 Discrete Denoising Checkpoints},
  author = {Andrei Rabau},
  year   = {2026},
  url    = {https://huggingface.co/AndreiRabau/diffusion-llm}
}

The implementation and experiment configurations are available in the source repository, including `model_wrappers/gpt2_diffusion_transformer_wrapper.py`, `train_utils/train.py`, and the files under `experiments/`.