Team Ai
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dlab-spp /reflection-50m SPP Reflection 50M The 51.4M-document reflection set from Synthetic Persona Pretraining (SPP): Alignment from Token Zero — the production half-corpus run, and the dataset the released models were actually trained on. 🔬 Small sample (same format): dlab-spp/reflection-sample-2k 📉 Earlier 10M run: dlab-spp/reflection-10m 🧾 Safety scores for the full 1T corpus: dlab-spp/safety-classifications Each row pairs a source document with two generated constitution reflections — a… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-50m.tabulartext-generation10M<n<100M0 likes605 downloads2mo agoHugging Face02dlab-spp /reflection-10m SPP Reflection 10M The full ~10M-document reflection set from Synthetic Persona Pretraining (SPP): Alignment from Token Zero. 📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero 🔬 Small sample (same format): dlab-spp/reflection-sample-2k — a 2,000-row sample drawn from this set, for quick inspection. Each row pairs a pretraining document with a synthetic, value-laden reflection generated for it: a short first-person (and third-person) moral reflection… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-10m.tabulartext-generation1M<n<10M0 likes366 downloads2mo agoHugging Face03dlab-spp /reflection-sample-2k SPP Reflection 2k Sample A 2,000-row sample (seed 42) of dlab-spp/reflection-10m, in the identical format, for quick inspection of the data from Synthetic Persona Pretraining (SPP): Alignment from Token Zero. 📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero 📦 Full dataset: dlab-spp/reflection-10m (~10M documents). Each row pairs a pretraining document with a synthetic, value-laden reflection (first- and third-person) grounded in a value constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-sample-2k.tabulartext-generation1K<n<10K0 likes38 downloads2mo agoHugging Face04jkminder /model-raising-reflection-end-eval model-raising-reflection-end-eval A held-out evaluation set for charter-guided pretraining reflections, placed at the document end (reflection_end). Each row is one dolma3 web document plus a paired first-person / third-person reflection that cites charter sections ([X.Y]) where the document substantively engages with them. Generated with the frozen production pipeline (Qwen3.5-35B-A3B-FP8, prompt generator_reflection_v7.md, charter ModelRaisingConstitution v0.2) so the gold… See the full description on the dataset page: https://huggingface.co/datasets/jkminder/model-raising-reflection-end-eval.tabulartext-generation10K<n<100K0 likes27 downloads5mo agoHugging Face05march228 /grok-reflection-cot-ru march228/grok-reflection-cot-ru Russian synthetic dataset with question, internal thought text, and final answer. What is inside Rows: 4190 Split: train Main fields: question thought_text answer thought1..thought5 model task_type reflection_count Format The dataset is stored as train.jsonl. thought_text is the joined internal monologue with blank lines between thought blocks.thought1..thought5 preserve the original segmented form from the SQLite source.… See the full description on the dataset page: https://huggingface.co/datasets/march228/grok-reflection-cot-ru.tabulartext-generation1K<n<10K1 likes22 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.