Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zhiyuanhucs /nemotron-student-fail-v41-clean-thinking DeepSeek-V4.1 clean and action-only trajectories with Nemotron outcomes DeepSeek-V4.1 reward-1 trajectories rebuilt from the complete teacher audit under v57-test-path-component-boundary+v57-target-source-recheck. The V4.1 reward and trajectory tier do not by themselves prove that Nemotron failed. Student outcomes are joined from nemotron-prolike-coverage-audit-20261001.json. A student failure requires either complete required-test results with reward 0, or an individually… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/nemotron-student-fail-v41-clean-thinking.tabulartext-generationn<1K1 likes13k downloads7d agoHugging Face02HuggingFaceH4 /Multilingual-Thinking Dataset summary Multilingual-Thinking is a reasoning dataset where the chain-of-thought has been translated from English into one of 4 languages: Spanish, French, Italian, and German. The dataset was created by sampling 1k training samples from the SystemChat subset of SmolTalk2 and translating the reasoning traces with another language model. This dataset was used in the OpenAI Cookbook to fine-tune the OpenAI gpt-oss models. You can load the dataset using: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/Multilingual-Thinking.texttext-generation1K<n<10K120 likes8.9k downloads1y agoHugging Face03ShareLab-SII /thinking_droid_lerobot_output_qwen3vlimage1M<n<10M0 likes3.4k downloads6mo agoHugging Face04kashif /opd-kd-thinky-deepmath-completions train_rl Completion Logs This dataset contains the on-policy generations produced during RL training with train_rl. Training details Key Value Algorithm OPD Model (student) HuggingFaceH4/KD-Thinky Model (teacher) Qwen/Qwen3-8B Prompt dataset HuggingFaceH4/DeepMath-103K Group size 4 Max completion tokens 4096 Temperature 1.0 Learning rate 0.0001 model_revision v00.08-step-000003125 dataset_configtrl_all lora_rank 128 opd_kl_coef 1.0… See the full description on the dataset page: https://huggingface.co/datasets/kashif/opd-kd-thinky-deepmath-completions.tabular10K<n<100K0 likes3.3k downloads8mo agoHugging Face05allenai /Dolci-Think-SFT-32B Dolci-Think-SFT Sources include a mixture of existing reasoning traces: OpenThoughts 3 (Apache 2.0): Extended to 32K context length and downsampled code prompts to 16X multiple, to 941,164 total prompts. Access our version, Dolci OpenThoughts 3 here. SYNTHETIC-2 (Apache 2.0) via the SFT-Verified split, 104,568 prompts. Nemotron Post-training dataset (CC BY 4), code split only, 113,777 prompts. New prompts and new reasoning traces from us (all ODC-BY-1.0): Dolci Think Persona IF:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-SFT-32B.text1M<n<10M32 likes3k downloads7mo agoHugging Face06llm-jp /llm-jp-4.1-thinking-sft-data llm-jp-4.1-thinking-sft-data Overview This dataset is a supervised fine-tuning (SFT) dataset used to train llm-jp-4.1-*-thinking models. This dataset is constructed from prompts and conversations collected from multiple data sources. For most subsets, reasoning processes and final responses used for LLM-jp-4.1 SFT were generated or augmented using gpt-oss-120b. For the tool-calling and agentic data derived from NVIDIA Nemotron datasets, the original conversations… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4.1-thinking-sft-data.text1M<n<10M4 likes2.8k downloads16d agoHugging Face07OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes2.5k downloads8mo agoHugging Face08llm-jp /llm-jp-4-thinking-sft-data llm-jp-4-thinking-sft-data Overview This dataset is a supervised fine-tuning (SFT) dataset used to train llm-jp-4-*-thinking models. This dataset is constructed by extracting prompts from multiple data sources and generating reasoning processes and final responses using gpt-oss-120b. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during generation with gpt-oss-120b. To support the continued development… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4-thinking-sft-data.text1M<n<10M9 likes2.2k downloads6mo agoHugging Face09ShareLab-SII /thinking_furniture_bench_dataset_lerobot_output_qwen3vlimage1M<n<10M0 likes2.1k downloads6mo agoHugging Face10lightonai /Dolci-Think-SFT-32B-Multilingual Dolci-Think-SFT-32B-Multilingual Dolci-Think-SFT-32B-Multilingual is a large-scale multilingual long chain-of-thought (CoT) reasoning corpus spanning six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample includes a question, a long-form reasoning trace, and a final answer, all translated into the target language, with sequences up to 32,768 tokens. It is released alongside the paper Rethinking the Multilingual Reasoning Gap with Layer Swap.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/Dolci-Think-SFT-32B-Multilingual.texttext-generation1M<n<10M2 likes1.8k downloads5mo agoHugging Face11CohereLabs /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Data source Prompts from AM-DeepSeek-R1-0528-Distilled Thinking traces and outputs distilled from gpt-oss-120b Translated with command-a-translate and DeepSeek-V3 Languages (44) Language Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M8 likes1.8k downloads1mo agoHugging Face12allenai /Dolci-Think-SFT-7B Dolci-Think-SFT Sources include a mixture of existing reasoning traces: OpenThoughts 3 (Apache 2.0): Extended to 32K context length and downsampled code prompts to 16X multiple, to 941,166 total prompts. Access our version, Dolci OpenThoughts 3 here. SYNTHETIC-2 (Apache 2.0) via the SFT-Verified split, 104,569 prompts. Nemotron Post-training dataset (CC BY 4), code split only, 113,777 prompts. New prompts and new reasoning traces from us (all ODC-BY-1.0): Dolci Think Persona… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-SFT-7B.text1M<n<10M21 likes1.8k downloads9mo agoHugging Face13openeurollm /Dolci-Think-SFT-translated Dolci-Think-SFT-translated Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations. Columns Each row is a translated conversation plus the result of a post-translation quality filter: id — source record id. messages — the translated conversation (list of {content, role}). filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.tabulartext-generation1M<n<10M0 likes1.7k downloads40m agoHugging Face14NarsAI /FineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1.5k downloads8mo agoHugging Face15OpenDataArena /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M126 likes1.5k downloads7mo agoHugging Face16ioi-leaderboard /ioi-eval-openrouter_anthropic_claude-3_7-sonnet_thinking-prompt-mem-limittextn<1K0 likes1.4k downloads2y agoHugging Face17ioi-leaderboard /ioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limittextn<1K0 likes1.4k downloads2y agoHugging Face18ericktwo /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M1 likes1.4k downloads8mo agoHugging Face19SeonghoonYu /think-kd-revision-generationtabular100K<n<1M0 likes1.4k downloads3mo agoHugging Face20joyfine /Qwen3-235B-A22B-Thinking-2507_Qwen3-1.7B_AIME_1983_2024textn<1K0 likes1.3k downloads11mo agoHugging Face21llm-jp /llm-jp-4.1-33b-thinking-dpo-data llm-jp-4.1-33b-thinking-dpo-data Overview This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4.1-33b-thinking. It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation. The fields chosen_analysis… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4.1-33b-thinking-dpo-data.text10K<n<100K2 likes1.3k downloads16d agoHugging Face22llm-jp /llm-jp-4.1-32b-a3b-thinking-dpo-data llm-jp-4.1-32b-a3b-thinking-dpo-data Overview This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4.1-32b-a3b-thinking. It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation. The fields… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4.1-32b-a3b-thinking-dpo-data.text100K<n<1M2 likes1.2k downloads16d agoHugging Face23fxmeng /UltraData-SFT-2605-no-think-8k-32k UltraData-SFT-2605 · no_think · 8k–32k A length-filtered subset of the no_think split of openbmb/UltraData-SFT-2605, containing conversations whose token length falls in the 8k–32k range. This is the medium-length tier intended for standard long-context SFT. Two companion tiers were produced from the same source: Dataset Length range Records this repo — fxmeng/UltraData-SFT-2605-no-think-8k-32k 8k–32k tokens 623,421 fxmeng/UltraData-SFT-2605-no-think-32k-200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-8k-32k.texttext-generation100K<n<1M0 likes1.2k downloads3mo agoHugging Face24rl-rag /browsecomp-qwen35-35b-a3b-think browsecomp-qwen35-35b-a3b-think Deep research agent evaluation on data/browsecomp.jsonl (normal split). Results Metric Value pass@4 43.0% avg@4 24.8% Trajectory accuracy 24.8% (1258/5064) Questions 1266 Trajectories 5064 (4 per question) Avg tool calls 41.1 Full conversations ❌ Model & Setup Model Qwen3.5-35B-A3B Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-qwen35-35b-a3b-think.tabular1K<n<10K0 likes1.1k downloads7mo agoHugging Face25allenai /Dolci-Think-RL-32B Dolci-Think-RL Dataset Summary Dolci-Think-RL is a deliberate reasoning RL dataset used for training Olmo-3-32B-Think model. It contains 102,026 high-quality prompts covering: Math Code Precise Instruction Following General Chat This dataset is structurally similar to Dolci-Think-RL-7B but with slightly different mixtures. Dataset Composition Total Samples: 102,026 Original Dataset Contribution Source Dataset Count IF… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-RL-32B.text100K<n<1M27 likes1.1k downloads9mo agoHugging Face26agentlans /epic-thinking Source Rows glaiveai/reasoning-v1-20m 1 999 793 BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples CC 1 267 534 PrimeIntellect/INTELLECT-3-SFT openreasoning_science 1 000 000 PrimeIntellect/INTELLECT-3-SFT am_chat 852 816 nvidia/Nemotron-Cascade-SFT-Stage-1 general 583 612 open-thoughts/OpenThoughts2-1M 541 898 PrimeIntellect/SYNTHETIC-1-SFT-Data 474 810 allenai/Dolci-Think-SFT-7B 334 908 allenai/Dolci-Think-SFT-32B 327 491 GeneralReasoning/GeneralThought-430K 291 946… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/epic-thinking.text10M<n<100M5 likes1k downloads6mo agoHugging Face27Lyric1010 /ablation_nemotron_thinking_32k_with_reasoning_effort Dataset: ablation_nemotron_thinking_32k_with_reasoning_effort This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train/ablation_nemotron_thinking_32k_with_reasoning_effort/stage_1/tmp/. text0 likes993 downloads9mo agoHugging Face28LoneResearch /explore-thinking-models-internaldocumentn<1K0 likes953 downloads4mo agoHugging Face29RuoliuYang /textlatent_zebra_thinkmorph_armAB Text-Latent (Arm A) vs All-Latent (Arm B) — Zebra-CoT + ThinkMorph 35638 samples/arm, 18 categories. Schema = ULVR/williamium style (sample_id, category, source_dataset, question, answer, input_image, intermediate_image_N, num_intermediate_steps, messages_json). armA_text_latent: real decoded text CoT + latent visual blocks (intermediate_image_1..3). armB_render_latent: reasoning text RENDERED to images, all-latent baseline (intermediate_image_1..17). messages_json = full Monet… See the full description on the dataset page: https://huggingface.co/datasets/RuoliuYang/textlatent_zebra_thinkmorph_armAB.image100K<n<1M0 likes939 downloads3mo agoHugging Face30Lyric1010 /ablation_openmathreasoning_thinking_with_reasoning_effort Dataset: ablation_openmathreasoning_thinking_with_reasoning_effort This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train/ablation_openmathreasoning_thinking_with_reasoning_effort/stage_1/tmp/. text0 likes922 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.