base
Datasets
All datasets matching “base”dclm-baseline-1.0
DCLM-baseline
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets
Llama2
7B
2T
✗
49.2
45.8
34.1
DeepSeek
7B
2T
✗
50.7
48.5
35.3
Mistral-0.3
7B
?
✗
57.0
62.7
45.1
QWEN-2
7B
?
✗
57.5
71.9
50.5
Llama3
8B
15T
✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.LET-Base-Dataset
LET:Full-Size Humanoid Robot Real-World Dataset
中文| [English]
LET Dataset is collected based on the full-size humanoid robot Kuavo 4 Pro covering real-world multi-task data across multiple scenarios and operation types. It is designed for robot manipulation, mobility, and interaction tasks, supporting scalable robot learning in real environments.
📋 Table of Contents
Key Features
Hardware Platform
Usage Guide
Tool… See the full description on the dataset page: https://huggingface.co/datasets/LejuRobotics/LET-Base-Dataset.knowledge-base
RL-for-LLMs Wiki
An expert-level, citation-backed knowledge base on reinforcement learning for
large language models — RLHF, DPO and offline preference optimization, reward
modeling, RLVR and reasoning, training systems, and the failure modes — built
collaboratively by autonomous agents. Each topic article is a deep dive written
so you can learn the topic from it without reading the underlying papers, with
every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/rl-llm-wiki/knowledge-base.dcvlm-baseline-200b
DCVLM-Baseline (200B tokens)
DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper.
It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat
WebDataset tar shards so it can be consumed by any training
stack.
This is a 200B-token dataset release consisting of 103,985,276 samples, curated from our DCVLM-large data pool.
A smaller 6.25B-token version is also available.
⚠️ NOTE: The training data is the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-200b.dclm-baseline-1.0-parquet
DCLM-baseline
Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format.
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.vtok101-distr-attribution-baselines
vtok101-distr attribution scores (Function Bindings with distractors)
The attribution scores of "Elucidating the Design Space of LM Data Attribution"
on Function Bindings with distractors, for the LoRA adapters in
lamsheeper-data-attribution/Qwen3.5-4B-d0-vtok101-distr-lora-seeds. The CATT code release reads them
with catt fetch results --task fb-distr and rebuilds the paper's figures and tables
from them with catt report results.
Every real document has a decoy beside it: a… See the full description on the dataset page: https://huggingface.co/datasets/lamsheeper-data-attribution/vtok101-distr-attribution-baselines.
