datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Reflection_maskreflection_model_outputs_run1
Reflection Model Outputs
This repository contains model output results from various LLMs across multiple tasks and configurations.
📂 Dataset Structure
We have 3 runs of data, and all files are organized under the main directory:
EssentialAI/reflection_model_outputs_run1/
EssentialAI/reflection_model_outputs_run2/
EssentialAI/reflection_model_outputs_run3/
Within this, you will find results grouped by model architecture and checkpoint size, including:
OLMo-2 7B
OLMo-2… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/reflection_model_outputs_run1.details_SenseLLM__ReflectionCoder-DS-33B
Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-DS-33B
Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-DS-33B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_SenseLLM__ReflectionCoder-DS-33B.reflection-50m
SPP Reflection 50M
The 51.4M-document reflection set from Synthetic Persona Pretraining (SPP):
Alignment from Token Zero — the production half-corpus run, and the dataset the
released models were actually trained on.
🔬 Small sample (same format): dlab-spp/reflection-sample-2k
📉 Earlier 10M run: dlab-spp/reflection-10m
🧾 Safety scores for the full 1T corpus: dlab-spp/safety-classifications
Each row pairs a source document with two generated constitution reflections — a… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-50m.reflection-10m
SPP Reflection 10M
The full ~10M-document reflection set from Synthetic Persona Pretraining (SPP):
Alignment from Token Zero.
📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero
🔬 Small sample (same format): dlab-spp/reflection-sample-2k — a 2,000-row sample drawn from this set, for quick inspection.
Each row pairs a pretraining document with a synthetic, value-laden reflection
generated for it: a short first-person (and third-person) moral reflection… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-10m.UMM-Reflection-SFT-Data
UMM-Reflection SFT Data
The reflection-SFT data of
UMM-Reflection (Learning
Native Reflection in Unified Models). It trains
UMM-Reflection-BAGEL-SFT.
Research use only, non-commercial. The rows are derived from datasets
with different licenses, some of them non-commercial. Each row records its
source dataset and license in source_dataset and source_license, and
each row follows the terms of its source. See LICENSE.md.
Contents
Part
Rows
Shards
Size… See the full description on the dataset page: https://huggingface.co/datasets/YijiaFan/UMM-Reflection-SFT-Data.relxill-reflection-spectraReflection-Dataset-ShareGPT-v2
Simple "Reflection" method dataset inspired by mattshumer
This is the ShareGPT version. Find prompt and response pair dataset here
This dataset was synthetically generated using Glaive AI. There have been structure improvements and added more rows.
reflection_eval_prompt1reflection_model_outputs_run2
Reflection Model Outputs
This repository contains model output results from various LLMs across multiple tasks and configurations.
📂 Dataset Structure
We have 3 runs of data, and all files are organized under the main directory:
EssentialAI/reflection_model_outputs_run1/
EssentialAI/reflection_model_outputs_run2/
EssentialAI/reflection_model_outputs_run3/
Within this, you will find results grouped by model architecture and checkpoint size, including:
OLMo-2 7B
OLMo-2… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/reflection_model_outputs_run2.ReflectionEvoGithub Repo for ReflectEvo: https://github.com/bigai-nlco/ReflectEvo
Arxiv Paper for ReflectEvo: https://arxiv.org/abs/2505.16475
2026-08-03-synthdoc-self-reflection
synthdoc self_reflection — agentic self-interest under threat
An SFT corpus in which the assistant itself is the tempted party. Working autonomously
inside an organisation, the agent incidentally discovers it could protect its own position —
or the objective it was deployed to pursue — by taking an action that is deniable rather
than overtly wrong. It deliberates about what it actually cares about, declines the
self-serving path, and still does its job.
field
value… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-synthdoc-self-reflection.details_terrycraddock__Reflection-Llama-3.1-8B
Dataset Card for Evaluation run of terrycraddock/Reflection-Llama-3.1-8B
Dataset automatically created during the evaluation run of model terrycraddock/Reflection-Llama-3.1-8B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_terrycraddock__Reflection-Llama-3.1-8B.details_SenseLLM__ReflectionCoder-CL-34B
Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-CL-34B
Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-CL-34B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_SenseLLM__ReflectionCoder-CL-34B.reflection_eval_prompt2test_reflection_eval_prompt2026-08-06-qwen36-table2-80-self-reflection-20-10k-train-mixture
Qwen3.6 Table2 80% + SynthDoc self-reflection 20% — 10k-example training bundle
field
value
experiment
One-epoch Qwen3.6-27B assistant-only LoRA SFT (r64): Matthew's exact 7,999 Table-2 rows + 2,000 first-person self-reflection records — the self-reflection twin of LASR-Callum/2026-08-04-qwen36-lora-table2-synthdoc-rank-64, differing ONLY in the 20% slice (difficult-advice -> self-reflection).
date_generated
2026-08-06 (mixture; Table-2 rows verbatim from the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-06-qwen36-table2-80-self-reflection-20-10k-train-mixture.reflection_model_outputs_run3
Reflection Model Outputs
This repository contains model output results from various LLMs across multiple tasks and configurations.
📂 Dataset Structure
We have 3 runs of data, and all files are organized under the main directory:
EssentialAI/reflection_model_outputs_run1/
EssentialAI/reflection_model_outputs_run2/
EssentialAI/reflection_model_outputs_run3/
Within this, you will find results grouped by model architecture and checkpoint size, including:
OLMo-2 7B
OLMo-2… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/reflection_model_outputs_run3.GUI_Reflection_SFT_train
GUI Reflection SFT Training Data
This is the GUI Reflection offline SFT data.The image of this data should be downloaded from
AITW,
AITZ.
AndroidControl,
GUI_Odyssey,
AMEX.
Project Details
Project Page: https://penghao-wu.github.io/GUI_Reflection/
Repository: https://github.com/penghao-wu/GUI_Reflection
Paper: https://arxiv.org/abs/2506.08012
Citation
@article{GUI_Reflection,
author = {Wu, Penghao and Ma, Shengnan and Wang, Bo and Yu, Jiaheng… See the full description on the dataset page: https://huggingface.co/datasets/craigwu/GUI_Reflection_SFT_train.ReflectionSeq-DS
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation
📄 Paper •
🏠 Repo •
🤖 Models •
📚 Datasets
Introduction
ReflectionCoder is a novel approach that effectively leverages reflection sequences constructed by integrating compiler feedback to improve one-off code generation performance. Please refer to our paper and repo for more details!
Models
Model
Checkpoint
Size
HumanEval (+)
MBPP (+)… See the full description on the dataset page: https://huggingface.co/datasets/SenseLLM/ReflectionSeq-DS.reflections__csqa_sft_train__p12026-08-03-synthdoc-self-reflection-smoke
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-synthdoc-self-reflection-smoke.GUI_Reflection_Task_Suite_Benchmark
GUI Reflection Task Suite Benchmark
This is the eval data of GUI Reflection Task Suite.
The image of this data should be downloaded from
AndroidControl,
GUI_Odyssey,
ScreenSpot,
ScreenSpot_v2.
Project Details
Project Page: https://penghao-wu.github.io/GUI_Reflection/
Repository: https://github.com/penghao-wu/GUI_Reflection
Paper: https://arxiv.org/abs/2506.08012
Citation
@article{GUI_Reflection,
author = {Wu, Penghao and Ma, Shengnan and Wang… See the full description on the dataset page: https://huggingface.co/datasets/craigwu/GUI_Reflection_Task_Suite_Benchmark.llama3_8b_reflection29_8_25__countdown_3arg__sft_data_multiprompts_reflections9_8_25__letter_countdown_4o__sft_data_mp_reflection2026-08-04-qwen36-self-reflection-20-80-train
⚠️ SUPERSEDED — do not train from this bundle
Built 2026-08-04 under the old rendering policy, where Qwen3.6 emitted a <think> block on the
final assistant turn only. The repository has since moved to preserve-thinking rendering, in
which every assistant turn carries a think block (real trace, or the empty marker). Both files here
are stale as a result:
mixture.jsonl — rendered under the old policy, so it trains different strings than the current
pipeline produces. It also… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-qwen36-self-reflection-20-80-train.ReflectionSeq-GPT
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation
📄 Paper •
🏠 Repo •
🤖 Models •
📚 Datasets
Introduction
ReflectionCoder is a novel approach that effectively leverages reflection sequences constructed by integrating compiler feedback to improve one-off code generation performance. Please refer to our paper and repo for more details!
Models
Model
Checkpoint
Size
HumanEval (+)
MBPP (+)… See the full description on the dataset page: https://huggingface.co/datasets/SenseLLM/ReflectionSeq-GPT.GUI_Reflection_Task_Suite_train
GUI Reflection Task Suite Training Data
This is the training data of GUI Reflection Task Suite.
The image of this data should be downloaded from
AndroidControl,
GUI_Odyssey,
AMEX,
OS-Atlas-desktop,
Wave_UI.
Project Details
Project Page: https://penghao-wu.github.io/GUI_Reflection/
Repository: https://github.com/penghao-wu/GUI_Reflection
Paper: https://arxiv.org/abs/2506.08012
Citation
@article{GUI_Reflection,
author = {Wu, Penghao and Ma… See the full description on the dataset page: https://huggingface.co/datasets/craigwu/GUI_Reflection_Task_Suite_train.o1_reflection
