Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OALL /details_SenseLLM__ReflectionCoder-DS-33B Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-DS-33B Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-DS-33B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_SenseLLM__ReflectionCoder-DS-33B.tabular100K<n<1M0 likes731 downloads2y agoHugging Face02dlab-spp /reflection-50m SPP Reflection 50M The 51.4M-document reflection set from Synthetic Persona Pretraining (SPP): Alignment from Token Zero — the production half-corpus run, and the dataset the released models were actually trained on. 🔬 Small sample (same format): dlab-spp/reflection-sample-2k 📉 Earlier 10M run: dlab-spp/reflection-10m 🧾 Safety scores for the full 1T corpus: dlab-spp/safety-classifications Each row pairs a source document with two generated constitution reflections — a… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-50m.tabulartext-generation10M<n<100M0 likes605 downloads2mo agoHugging Face03YijiaFan /UMM-Reflection-SFT-Data UMM-Reflection SFT Data The reflection-SFT data of UMM-Reflection (Learning Native Reflection in Unified Models). It trains UMM-Reflection-BAGEL-SFT. Research use only, non-commercial. The rows are derived from datasets with different licenses, some of them non-commercial. Each row records its source dataset and license in source_dataset and source_license, and each row follows the terms of its source. See LICENSE.md. Contents Part Rows Shards Size… See the full description on the dataset page: https://huggingface.co/datasets/YijiaFan/UMM-Reflection-SFT-Data.tabulartext-to-image10K<n<100K3 likes454 downloads12d agoHugging Face04dlab-spp /reflection-10m SPP Reflection 10M The full ~10M-document reflection set from Synthetic Persona Pretraining (SPP): Alignment from Token Zero. 📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero 🔬 Small sample (same format): dlab-spp/reflection-sample-2k — a 2,000-row sample drawn from this set, for quick inspection. Each row pairs a pretraining document with a synthetic, value-laden reflection generated for it: a short first-person (and third-person) moral reflection… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-10m.tabulartext-generation1M<n<10M0 likes366 downloads2mo agoHugging Face05mahiatlinux /Reflection-Dataset-ShareGPT-v2 Simple "Reflection" method dataset inspired by mattshumer This is the ShareGPT version. Find prompt and response pair dataset here This dataset was synthetically generated using Glaive AI. There have been structure improvements and added more rows. text1K<n<10K15 likes314 downloads2y agoHugging Face06bigai-nlco /ReflectionEvoGithub Repo for ReflectEvo: https://github.com/bigai-nlco/ReflectEvo Arxiv Paper for ReflectEvo: https://arxiv.org/abs/2505.16475 textquestion-answering100K<n<1M12 likes246 downloads1y agoHugging Face07feedbackagent /reflection_eval_prompt1text100K<n<1M0 likes195 downloads2y agoHugging Face08feedbackagent /test_reflection_eval_prompttext100K<n<1M2 likes191 downloads2y agoHugging Face09feedbackagent /reflection_eval_prompt2text100K<n<1M0 likes94 downloads2y agoHugging Face10SenseLLM /ReflectionSeq-DS ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation 📄 Paper • 🏠 Repo • 🤖 Models • 📚 Datasets Introduction ReflectionCoder is a novel approach that effectively leverages reflection sequences constructed by integrating compiler feedback to improve one-off code generation performance. Please refer to our paper and repo for more details! Models Model Checkpoint Size HumanEval (+) MBPP (+)… See the full description on the dataset page: https://huggingface.co/datasets/SenseLLM/ReflectionSeq-DS.texttext-generation10K<n<100K5 likes93 downloads2y agoHugging Face11TAUR-dev /reflections__csqa_sft_train__p1tabular100K<n<1M0 likes87 downloads1y agoHugging Face12OALL /details_terrycraddock__Reflection-Llama-3.1-8B Dataset Card for Evaluation run of terrycraddock/Reflection-Llama-3.1-8B Dataset automatically created during the evaluation run of model terrycraddock/Reflection-Llama-3.1-8B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_terrycraddock__Reflection-Llama-3.1-8B.tabular100K<n<1M0 likes83 downloads2y agoHugging Face13OALL /details_SenseLLM__ReflectionCoder-CL-34B Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-CL-34B Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-CL-34B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_SenseLLM__ReflectionCoder-CL-34B.tabular100K<n<1M0 likes71 downloads2y agoHugging Face14feedbackagent /llama3_8b_reflection2text100K<n<1M2 likes69 downloads2y agoHugging Face15TAUR-dev /9_8_25__countdown_3arg__sft_data_multiprompts_reflectionstext100K<n<1M0 likes66 downloads1y agoHugging Face16TAUR-dev /9_8_25__letter_countdown_4o__sft_data_mp_reflectiontabular10K<n<100K0 likes66 downloads1y agoHugging Face17mahiatlinux /Reflection-Dataset-v2 Second version of a simple "Reflection" method dataset inspired by mattshumer This is the prompt and response version. Find ShareGPT version here This dataset was synthetically generated using Glaive AI. There have been structure improvements and added more rows. text1K<n<10K37 likes63 downloads2y agoHugging Face18leafspark /o1_reflectiontext1K<n<10K2 likes62 downloads2y agoHugging Face19SenseLLM /ReflectionSeq-GPT ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation 📄 Paper • 🏠 Repo • 🤖 Models • 📚 Datasets Introduction ReflectionCoder is a novel approach that effectively leverages reflection sequences constructed by integrating compiler feedback to improve one-off code generation performance. Please refer to our paper and repo for more details! Models Model Checkpoint Size HumanEval (+) MBPP (+)… See the full description on the dataset page: https://huggingface.co/datasets/SenseLLM/ReflectionSeq-GPT.texttext-generation10K<n<100K5 likes61 downloads2y agoHugging Face20gabrielmbmb /distilabel-reflection-tuning Dataset Card for distilabel-reflection-tuning This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: reflection.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/gabrielmbmb/distilabel-reflection-tuning/raw/main/reflection.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated… See the full description on the dataset page: https://huggingface.co/datasets/gabrielmbmb/distilabel-reflection-tuning.textn<1K55 likes59 downloads2y agoHugging Face21TAUR-dev /9_8_25__countdown_4arg__sft_data_mp_reflection_ckpt_chunk_8text1K<n<10K0 likes55 downloads1y agoHugging Face22TAUR-dev /9_8_25__countdown_3arg__sft_data_GPT4o_multiprompts_gpt4o_reflectionstext100K<n<1M0 likes54 downloads1y agoHugging Face23dougalldeepmind /2026-08-04-qwen36-self-reflection-20-80-train ⚠️ SUPERSEDED — do not train from this bundle Built 2026-08-04 under the old rendering policy, where Qwen3.6 emitted a <think> block on the final assistant turn only. The repository has since moved to preserve-thinking rendering, in which every assistant turn carries a think block (real trace, or the empty marker). Both files here are stale as a result: mixture.jsonl — rendered under the old policy, so it trains different strings than the current pipeline produces. It also… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-qwen36-self-reflection-20-80-train.text1K<n<10K0 likes54 downloads2mo agoHugging Face24TAUR-dev /9_8_25__letter_countdown_4o__sft_data_mp_reflection_ckpt_chunk_5tabular1K<n<10K0 likes53 downloads1y agoHugging Face25TAUR-dev /reflections__gsm8k_sft_train__backup_2f525e6tabular10K<n<100K0 likes45 downloads1y agoHugging Face26open-llm-leaderboard /SenseLLM__ReflectionCoder-CL-34B-detailsgated Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-CL-34B Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-CL-34B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SenseLLM__ReflectionCoder-CL-34B-details.tabular10K<n<100K0 likes44 downloads2y agoHugging Face27feedbackagent /reflection_n16text10K<n<100K0 likes42 downloads2y agoHugging Face28TAUR-dev /multitask_intermediate_ac4_v2_reflections5_formats-C_fulltext1K<n<10K0 likes42 downloads1y agoHugging Face29gsayak /reflectiontabular1K<n<10K0 likes40 downloads2y agoHugging Face30TAUR-dev /9_8_25__countdown_4arg__sft_data_mp_reflectiontext10K<n<100K0 likes40 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.