Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IFM /Math-Reasoning Math-Reasoning Dataset Description Mathematical problem-solving, rewriting, and dialogue data for reasoning-oriented language-model training. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Math-Reasoning.texttext-generation1B<n<10B26 likes50k downloads1mo agoHugging Face02IFM /Code-Reasoning Code-Reasoning Dataset Description Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2 Web and… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Code-Reasoning.texttext-generation100M<n<1B99 likes47k downloads1mo agoHugging Face03FreedomIntelligence /medical-o1-reasoning-SFT News [2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data. [2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1. [2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.textquestion-answering10K<n<100K1.2k likes21k downloads1y agoHugging Face04LucasFang /FLUX-Reason-6M FLUX-Reason-6M FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative models. This dataset was created to bridge the performance gap between open-source and leading closed-source text-to-image systems. This dataset contains: 6 million high-quality, reasoning-focused images synthesized by the state-of-the-art FLUX.1-dev model. 20 million bilingual (English and Chinese) descriptions, providing a rich… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M.image1M<n<10M116 likes15k downloads8mo agoHugging Face05BUPT-Reasoning-Lab /FinMMDocRimage0 likes12k downloads8mo agoHugging Face06multimodal-reasoning-lab /Zebra-CoT Zebra‑CoT A diverse large-scale dataset for interleaved vision‑language reasoning traces. Dataset Description Zebra‑CoT is a diverse large‑scale dataset with 182,384 samples containing logically coherent interleaved text‑image reasoning traces across four major categories: scientific reasoning, 2D visual reasoning, 3D visual reasoning, and visual logic & strategic games. Dataset Structure Each example in Zebra‑CoT consists of: Problem statement:… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT.imageany-to-any100K<n<1M81 likes12k downloads8mo agoHugging Face07livebench /reasoning Dataset Card for "livebench/reasoning" LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties: LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. Each question has verifiable, objective ground-truth answers, allowing hard questions to be scored… See the full description on the dataset page: https://huggingface.co/datasets/livebench/reasoning.textn<1K20 likes12k downloads2y agoHugging Face08spatial-reason /qwen_trajectories_finalimage1K<n<10K0 likes11k downloads7mo agoHugging Face09Ever2after /3d-spatial-reasoning-10 likes10k downloads1y agoHugging Face10Ever2after /3d-spatial-reasoning0 likes9.1k downloads1y agoHugging Face11reasoning-core /stagingRaw generator output behind procedural-pile, one configuration per build, before deduplication and the train/test split. load_dataset("reasoning-core/staging", "rc14", split="train", streaming=True) text10M<n<100M4 likes8.8k downloads7d agoHugging Face12imo2026-challenge /chankhavu-imo-reasoning-traces0 likes8k downloads2mo agoHugging Face13reasonwang /skillgym-test SkillGym test set 400 agent tasks that each require one agent skill, in Harbor task format. Layout tasks/ procedural/ held_in/ 50 tasks held_out/ 50 tasks constraint_satisfaction/ (same) abductive/ (same) partial_order/ (same) manifest.jsonl one row per task Each profile has 50 held-in and 50 held-out tasks (held-in = the task's skill also appears in the training data, held-out =… See the full description on the dataset page: https://huggingface.co/datasets/reasonwang/skillgym-test.0 likes6k downloads9d agoHugging Face14Ever2after /3d-spatial-reasoning-2image10K<n<100K0 likes5.8k downloads5mo agoHugging Face15MoreThought /Fable-5.1-Max-Reasoning-Filtered-10000x Dataset Description This dataset contains 10,000 agentic coding and reasoning multi-turn high-quality traces generated by the new Fable 5.1 model using max reasoning effort. It holds almost 500,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains. It has also been deduplicated and heavily filtered to remove low-quality traces, keeping only high-quality traces. Dataset Statistics Metric Value Total Examples 10,000… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/Fable-5.1-Max-Reasoning-Filtered-10000x.texttext-generation10K<n<100K297 likes5.2k downloads5h agoHugging Face16IFM /SFT-Reasoning SFT-Reasoning Dataset Description Instruction-following and reasoning data prepared for supervised fine-tuning. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2 Web and question-answering text… See the full description on the dataset page: https://huggingface.co/datasets/IFM/SFT-Reasoning.texttext-generation10M<n<100M18 likes5.1k downloads24d agoHugging Face17AlexCuadron /SWE-Bench-Verified-O1-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities on the SWE-Bench Verified dataset, achieving a 28.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code generation through enhanced action-based reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-reasoning-high-results.textquestion-answeringn<1K7 likes5.1k downloads2y agoHugging Face18yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M53 likes4.2k downloads7mo agoHugging Face19Mobiusi /copying_reasoning_task_improved copying_reasoning_task_improved Dataset Description The Enhanced Copying Reasoning Task Dataset is designed to provide a rich resource for analyzing promotional texts and their key elements. This dataset includes a variety of question-and-answer formats, focusing on whether specific phrases are mentioned within the text. Its purpose is to assist in the training of models for natural language understanding tasks, particularly in identifying relevant information in… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/copying_reasoning_task_improved.textn<1K1 likes4.2k downloads1y agoHugging Face20Alex11556666 /Reason_Tuning 🎨 UniReason • Unified Reasoning Framework for World Knowledge–Aligned Image Generation and Editing UniReason is a unified framework that harmonizes text-to-image generation and image editing through a dual reasoning paradigm. We formulate generation as world knowledge-enhanced planning to inject implicit constraints, and leverage editing capabilities for fine-grained visual refinement to further correct visual errors via self-reflection. This approach… See the full description on the dataset page: https://huggingface.co/datasets/Alex11556666/Reason_Tuning.tabularimage-to-image100K<n<1M2 likes4k downloads5mo agoHugging Face21WNJXYK /AIME_1983_2024-Reasoning-Paths News 🌟🌟🌟 Try this dataset in our HuggingFace Space! 🥳🥳🥳 Thrilled to share that this NeurIPS paper was selected as 🏆 #1 Paper of the Day on Oct. 20th! Sampled Reasoning Paths for the AIME dataset (from 1983 to 2024) This dataset contains sampled reasoning paths for the AIME_1983_2024 dataset, released as part of the NeurIPS 2025 paper: "A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning" (Arxiv).… See the full description on the dataset page: https://huggingface.co/datasets/WNJXYK/AIME_1983_2024-Reasoning-Paths.text-generation17 likes3.8k downloads1y agoHugging Face22OpenMed /Medical-Reasoning-SFT-Baichuan-M3-235B Medical-Reasoning-SFT-Baichuan-M3-235B A large-scale medical reasoning dataset generated using baichuan-inc/Baichuan-M3-235B, containing over 124,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Baichuan-M3-235B is ranked #1 on HealthBench Total leaderboard and achieves state-of-the-art performance on medical reasoning benchmarks. Dataset Overview Metric Value Model baichuan-inc/Baichuan-M3-235B Total Samples 124… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Baichuan-M3-235B.texttext-generation100K<n<1M7 likes3.7k downloads8mo agoHugging Face23TheFinAI /Fino1_Reasoning_Path_FinQAgated Fino1 Reasoning Paths (FinQA) 📄 Paper · 🤗 Collection · 💻 Code · 🏆 Leaderboard · 🌐 The Fin AI Fino1 Reasoning Paths (FinQA) is the reasoning-path dataset used to train Fino1-8B: FinQA questions paired with GPT-4o-generated chain-of-thought. It accompanies Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance (arXiv:2502.08127). Used to train: Fino1-8B. Quick Start from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/Fino1_Reasoning_Path_FinQA.textquestion-answering1K<n<10K42 likes3.7k downloads3d agoHugging Face24Intelligent-Internet /II-Medical-Reasoning-SFT II-Medical-Reasoning-SFT II-Medical SFT is a curated dataset designed to support the supervised fine-tuning of large language models (LLMs) for medical reasoning tasks. It comprises multi-turn dialogues, clinical case scenarios, and question-answer pairs that reflect the complex reasoning processes encountered in real-world clinical practice. The dataset is intended to help models develop key competencies such as differential diagnosis, evidence-based decision-making, patient… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-Reasoning-SFT.text1M<n<10M57 likes3.6k downloads1y agoHugging Face25interstellarninja /hermes_reasoning_tool_use TL;DR 51 004 ShareGPT conversations that teach LLMs when, how and whether to call tools.Built with the Nous Research Atropos RL stack in Atropos using a custom MultiTurnToolCallingEnv, and aligned with BFCL v3 evaluation scenarios.Released by @interstellarninja under Apache-2.0. 1 Dataset Highlights Count Split Scenarios covered Size 51 004 train single-turn · multi-turn · multi-step · relevance 392 MB Each row: OpenAI-style conversations… See the full description on the dataset page: https://huggingface.co/datasets/interstellarninja/hermes_reasoning_tool_use.textquestion-answering10K<n<100K192 likes3.6k downloads10mo agoHugging Face26reasoning-machines /gsm-hard Dataset Summary This is the harder version of gsm8k math reasoning dataset (https://huggingface.co/datasets/gsm8k). We construct this dataset by replacing the numbers in the questions of GSM8K with larger numbers that are less common.  Supported Tasks and Leaderboards This dataset is used to evaluate math reasoning Languages English - Numbers Dataset Structure dataset = load_dataset("reasoning-machines/gsm-hard") DatasetDict({ train: Dataset({… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-machines/gsm-hard.text1K<n<10K66 likes3.6k downloads4y agoHugging Face27facebook /natural_reasoningNaturalReasoning is a large-scale dataset for general reasoning tasks. It consists of high-quality challenging reasoning questions backtranslated from pretraining corpora DCLM and FineMath. The questions have been deduplicated and decontaminated from popular reasoning benchmarks including MATH, GPQA, MMLU-Pro, MMLU-STEM. For each question, we extract the reference final answer from the original document from the pretraining corpora if possible. We also provide a model-generated response from… See the full description on the dataset page: https://huggingface.co/datasets/facebook/natural_reasoning.texttext-generation1M<n<10M584 likes3.4k downloads2y agoHugging Face28zhifeixie /Audio-Reasoner-CoTAtext100K<n<1M14 likes3.3k downloads1y agoHugging Face29Hiepppp /reasoningtext100K<n<1M0 likes3.2k downloads1mo agoHugging Face30smshahbaj /verifiable-code-reasoning Verifiable Code Reasoning Execution-verified Python problems with chain-of-thought Sandbox-checked solutions · Multi-test unit checks · Deduplicated instances · Training-ready sft_text Overview Verifiable Code Reasoning is a large-scale dataset of Python coding problems where every kept solution has passed sandboxed unit tests. Unlike scraped contest dumps or unverified LLM traces, an example enters this release only if: a reference… See the full description on the dataset page: https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning.texttext-generation1M<n<10M2 likes3.1k downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.