Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OALL /details_SenseLLM__ReflectionCoder-DS-33B Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-DS-33B Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-DS-33B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_SenseLLM__ReflectionCoder-DS-33B.tabular100K<n<1M0 likes731 downloads2y agoHugging Face02dlab-spp /reflection-50m SPP Reflection 50M The 51.4M-document reflection set from Synthetic Persona Pretraining (SPP): Alignment from Token Zero — the production half-corpus run, and the dataset the released models were actually trained on. 🔬 Small sample (same format): dlab-spp/reflection-sample-2k 📉 Earlier 10M run: dlab-spp/reflection-10m 🧾 Safety scores for the full 1T corpus: dlab-spp/safety-classifications Each row pairs a source document with two generated constitution reflections — a… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-50m.tabulartext-generation10M<n<100M0 likes605 downloads2mo agoHugging Face03YijiaFan /UMM-Reflection-SFT-Data UMM-Reflection SFT Data The reflection-SFT data of UMM-Reflection (Learning Native Reflection in Unified Models). It trains UMM-Reflection-BAGEL-SFT. Research use only, non-commercial. The rows are derived from datasets with different licenses, some of them non-commercial. Each row records its source dataset and license in source_dataset and source_license, and each row follows the terms of its source. See LICENSE.md. Contents Part Rows Shards Size… See the full description on the dataset page: https://huggingface.co/datasets/YijiaFan/UMM-Reflection-SFT-Data.tabulartext-to-image10K<n<100K3 likes454 downloads11d agoHugging Face04dlab-spp /reflection-10m SPP Reflection 10M The full ~10M-document reflection set from Synthetic Persona Pretraining (SPP): Alignment from Token Zero. 📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero 🔬 Small sample (same format): dlab-spp/reflection-sample-2k — a 2,000-row sample drawn from this set, for quick inspection. Each row pairs a pretraining document with a synthetic, value-laden reflection generated for it: a short first-person (and third-person) moral reflection… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-10m.tabulartext-generation1M<n<10M0 likes366 downloads2mo agoHugging Face05kshitijd /relxill-reflection-spectratabular10M<n<100M0 likes262 downloads3mo agoHugging Face06TAUR-dev /reflections__csqa_sft_train__p1tabular100K<n<1M0 likes87 downloads1y agoHugging Face07OALL /details_terrycraddock__Reflection-Llama-3.1-8B Dataset Card for Evaluation run of terrycraddock/Reflection-Llama-3.1-8B Dataset automatically created during the evaluation run of model terrycraddock/Reflection-Llama-3.1-8B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_terrycraddock__Reflection-Llama-3.1-8B.tabular100K<n<1M0 likes83 downloads2y agoHugging Face08OALL /details_SenseLLM__ReflectionCoder-CL-34B Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-CL-34B Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-CL-34B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_SenseLLM__ReflectionCoder-CL-34B.tabular100K<n<1M0 likes71 downloads2y agoHugging Face09TAUR-dev /9_8_25__letter_countdown_4o__sft_data_mp_reflectiontabular10K<n<100K0 likes66 downloads1y agoHugging Face10TAUR-dev /9_8_25__letter_countdown_4o__sft_data_mp_reflection_ckpt_chunk_5tabular1K<n<10K0 likes53 downloads1y agoHugging Face11TAUR-dev /reflections__gsm8k_sft_train__backup_2f525e6tabular10K<n<100K0 likes45 downloads1y agoHugging Face12open-llm-leaderboard /SenseLLM__ReflectionCoder-CL-34B-detailsgated Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-CL-34B Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-CL-34B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SenseLLM__ReflectionCoder-CL-34B-details.tabular10K<n<100K0 likes44 downloads2y agoHugging Face13gsayak /reflectiontabular1K<n<10K0 likes40 downloads2y agoHugging Face14TAUR-dev /reflections__countdown4argtabular10K<n<100K0 likes38 downloads1y agoHugging Face15dlab-spp /reflection-sample-2k SPP Reflection 2k Sample A 2,000-row sample (seed 42) of dlab-spp/reflection-10m, in the identical format, for quick inspection of the data from Synthetic Persona Pretraining (SPP): Alignment from Token Zero. 📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero 📦 Full dataset: dlab-spp/reflection-10m (~10M documents). Each row pairs a pretraining document with a synthetic, value-laden reflection (first- and third-person) grounded in a value constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-sample-2k.tabulartext-generation1K<n<10K0 likes38 downloads2mo agoHugging Face16open-llm-leaderboard /EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection-details.tabular10K<n<100K0 likes37 downloads2y agoHugging Face17open-llm-leaderboard /mattshumer__Reflection-Llama-3.1-70B-detailsgated Dataset Card for Evaluation run of mattshumer/Reflection-Llama-3.1-70B Dataset automatically created during the evaluation run of model mattshumer/Reflection-Llama-3.1-70B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mattshumer__Reflection-Llama-3.1-70B-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face18open-llm-leaderboard /olabs-ai__reflection_model-detailsgated Dataset Card for Evaluation run of olabs-ai/reflection_model Dataset automatically created during the evaluation run of model olabs-ai/reflection_model The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/olabs-ai__reflection_model-details.tabular10K<n<100K0 likes33 downloads2y agoHugging Face19open-llm-leaderboard /EpistemeAI2__Fireball-Llama-3.1-8B-Philos-Reflection-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Llama-3.1-8B-Philos-Reflection Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Llama-3.1-8B-Philos-Reflection The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Llama-3.1-8B-Philos-Reflection-details.tabular10K<n<100K0 likes33 downloads2y agoHugging Face20open-llm-leaderboard /SenseLLM__ReflectionCoder-DS-33B-detailsgated Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-DS-33B Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-DS-33B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SenseLLM__ReflectionCoder-DS-33B-details.tabular10K<n<100K0 likes31 downloads2y agoHugging Face21TAUR-dev /9_8_25__acronym_4o__sft_data_mp_reflectiontabular10K<n<100K0 likes28 downloads1y agoHugging Face22TAUR-dev /9_8_25__letter_countdown_4o__sft_data_multiprompts_reflectionstabular10K<n<100K0 likes28 downloads1y agoHugging Face23jkminder /model-raising-reflection-end-eval model-raising-reflection-end-eval A held-out evaluation set for charter-guided pretraining reflections, placed at the document end (reflection_end). Each row is one dolma3 web document plus a paired first-person / third-person reflection that cites charter sections ([X.Y]) where the document substantively engages with them. Generated with the frozen production pipeline (Qwen3.5-35B-A3B-FP8, prompt generator_reflection_v7.md, charter ModelRaisingConstitution v0.2) so the gold… See the full description on the dataset page: https://huggingface.co/datasets/jkminder/model-raising-reflection-end-eval.tabulartext-generation10K<n<100K0 likes27 downloads5mo agoHugging Face24xDAN-datasets /Maggen-Reflection-3.1-70b-50k-filtered-scoredDatasetDict({ train: Dataset({ features: ['created', 'response', 'pre_query_template', 'instruction', 'gen_input_configs', 'gen_response_configs', 'raw_instruction', 'id', 'instruction_sanitize_class_num', 'scores', 'model_name'], num_rows: 36884 }) }) 每个唯一值的计数: scores [9.0] 9469 [7.0] 6224 [10.0] 6009 [6.0] 4003 [8.0] 3149 [5.0] 2578 [4.0] 2575 [3.0] 1566 [2.0] 1051 [1.0] 209 [] 51 tabular10K<n<100K1 likes26 downloads2y agoHugging Face25TAUR-dev /reflection_countdown_3args_v2_14tabular1K<n<10K0 likes25 downloads1y agoHugging Face26TAUR-dev /reflection_countdown_3args_v2_23tabular1K<n<10K0 likes24 downloads1y agoHugging Face27TAUR-dev /reflection_countdown_3args_v2_17tabular1K<n<10K0 likes23 downloads1y agoHugging Face28TAUR-dev /reflection_countdown_3args_v2_29tabular1K<n<10K0 likes23 downloads1y agoHugging Face29march228 /grok-reflection-cot-ru march228/grok-reflection-cot-ru Russian synthetic dataset with question, internal thought text, and final answer. What is inside Rows: 4190 Split: train Main fields: question thought_text answer thought1..thought5 model task_type reflection_count Format The dataset is stored as train.jsonl. thought_text is the joined internal monologue with blank lines between thought blocks.thought1..thought5 preserve the original segmented form from the SQLite source.… See the full description on the dataset page: https://huggingface.co/datasets/march228/grok-reflection-cot-ru.tabulartext-generation1K<n<10K1 likes22 downloads7mo agoHugging Face30TAUR-dev /reflection_countdown_3args_v2_10tabular1K<n<10K0 likes20 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.