datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reflection-50m
SPP Reflection 50M
The 51.4M-document reflection set from Synthetic Persona Pretraining (SPP):
Alignment from Token Zero — the production half-corpus run, and the dataset the
released models were actually trained on.
🔬 Small sample (same format): dlab-spp/reflection-sample-2k
📉 Earlier 10M run: dlab-spp/reflection-10m
🧾 Safety scores for the full 1T corpus: dlab-spp/safety-classifications
Each row pairs a source document with two generated constitution reflections — a… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-50m.reflection-10m
SPP Reflection 10M
The full ~10M-document reflection set from Synthetic Persona Pretraining (SPP):
Alignment from Token Zero.
📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero
🔬 Small sample (same format): dlab-spp/reflection-sample-2k — a 2,000-row sample drawn from this set, for quick inspection.
Each row pairs a pretraining document with a synthetic, value-laden reflection
generated for it: a short first-person (and third-person) moral reflection… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-10m.reflection-sample-2k
SPP Reflection 2k Sample
A 2,000-row sample (seed 42) of dlab-spp/reflection-10m,
in the identical format, for quick inspection of the data from
Synthetic Persona Pretraining (SPP): Alignment from Token Zero.
📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero
📦 Full dataset: dlab-spp/reflection-10m (~10M documents).
Each row pairs a pretraining document with a synthetic, value-laden reflection
(first- and third-person) grounded in a value constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-sample-2k.model-raising-reflection-end-eval
model-raising-reflection-end-eval
A held-out evaluation set for charter-guided pretraining reflections, placed at the
document end (reflection_end). Each row is one dolma3 web document plus a paired
first-person / third-person reflection that cites charter sections ([X.Y]) where the
document substantively engages with them. Generated with the frozen production pipeline
(Qwen3.5-35B-A3B-FP8, prompt generator_reflection_v7.md, charter
ModelRaisingConstitution v0.2)
so the gold… See the full description on the dataset page: https://huggingface.co/datasets/jkminder/model-raising-reflection-end-eval.grok-reflection-cot-ru
march228/grok-reflection-cot-ru
Russian synthetic dataset with question, internal thought text, and final answer.
What is inside
Rows: 4190
Split: train
Main fields:
question
thought_text
answer
thought1..thought5
model
task_type
reflection_count
Format
The dataset is stored as train.jsonl.
thought_text is the joined internal monologue with blank lines between thought blocks.thought1..thought5 preserve the original segmented form from the SQLite source.… See the full description on the dataset page: https://huggingface.co/datasets/march228/grok-reflection-cot-ru.
