PrincetonPLI/WebGraphMix-openlm-1B
WebGraphMix OpenLM 1B Checkpoints
Pretrained OpenLM 1B checkpoints from **Hubs or Fringes? Pretraining Data Selection via Web Graph Centrality** (WebGraphMix).
These models replicate the headline 1B-scale Table 1 experiments: four data-selection methods trained on mixtures derived from WebOrganizer/Corpus-200B, evaluated on DCLM CORE v2 (mmlu_and_lowvar, 23 tasks).
Checkpoints
Each folder contains an OpenLM PyTorch checkpoint (epoch_11.pt, final epoch) plus shared eval metadata.
Scores areaggregated_resultsfrom themmlu_and_lowvareval suite (23 low-variance ICL tasks). See the WebGraphMix repo to reproduce evaluation.
Model details
All four models share the same architecture and optimizer settings; they differ only in the importance-sampled pretraining mixture.
Download
huggingface-cli download PrincetonPLI/WebGraphMix-openlm-1B \
--local-dir ./dclm/checkpoints \
--repo-type modelOr from the WebGraphMix repo:
git clone https://github.com/princeton-pli/WebGraphMix.git
cd WebGraphMix
./experiments/artifacts/download.sh checkpointsExpected layout after download:
checkpoints/
├── open_lm_1b_eval_params.txt
├── random_selection/epoch_11.pt
├── dclm_fasttext_only/epoch_11.pt
├── betweenness_alpha0.5/epoch_11.pt
└── betweenness_alpha0.5_mult_div_dclm_fasttext/epoch_11.ptApproximate size: ~17 GB per checkpoint (~68 GB total).
Evaluate (recommended)
The checkpoints are stored in native OpenLM PyTorch format. The easiest path is the WebGraphMix evaluation pipeline:
conda env create -f environment.yml && conda activate webgraphmix
cd dclm && pip install -e . && cd ..
export REPO_ROOT=$(pwd)
./experiments/artifacts/download.sh checkpoints
# Default: WebGraphMix 50/50 betweenness
./experiments/eval/mmlu_and_lowvar.sh
# Other checkpoints
./experiments/eval/mmlu_and_lowvar.sh random_selection
./experiments/eval/mmlu_and_lowvar.sh dclm_fasttext_only
./experiments/eval/mmlu_and_lowvar.sh betweenness_alpha0.5_mult_div_dclm_fasttextAggregate scores across models:
cd dclm/exp_data/evals && python benchmark_score_comparison.pyEvaluation uses ≥2 GPUs by default (FSDP); a single GPU may OOM on the 1B model.
Convert to Hugging Face format (optional)
To load with transformers + open_lm HF wrappers:
export REPO_ROOT=/path/to/WebGraphMix
export CHECKPOINT_INPUT_DIR=$REPO_ROOT/dclm/checkpoints
export CHECKPOINT_HF_OUTPUT_DIR=$REPO_ROOT/dclm/checkpoints_hf
python dclm/convert_openlm_to_hf_1b.pyThis produces Hugging Face–compatible folders with OpenLMConfig / OpenLMForCausalLM weights and the GPT-NeoX tokenizer.
Training data (summary)
Centrality scores come from PrincetonPLI/cc-centrality-scores. Full sampling and tokenization steps are documented in the WebGraphMix README.
Citation
@article{badoni2026webgraphmix,
title={Hubs or Fringes: Pretraining Data Selection via Web Graph Centrality},
author={Badoni, Vedant and Chen, Danqi and Wang, Xinyi},
year={2026}
}License
Released under the MIT License, consistent with the DCLM codebase used for training and evaluation.
