datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NuminaMath-CoT
Dataset Card for NuminaMath CoT
Dataset Summary
Approximately 860k math problems, where each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs and mathematics discussion forums. The processing steps include (a) OCR from the original PDFs, (b) segmentation… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/NuminaMath-CoT.CoTracker3_Kubric
Kubric Dataset for CoTracker 3
Overview
This dataset was specifically created for training CoTracker 3, a state-of-the-art point tracking model. The dataset was generated using the Kubric engine.
Dataset Specifications
Size: ~6,000 sequences
Resolution: 512×512 pixels
Sequence Length: 120 frames per sequence
Camera Movement: Carefully rendered with subtle camera motion to simulate realistic scenarios
Format: Generated using Kubric engine
Usage
The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/CoTracker3_Kubric.cot-eval-traces-2.0Zebra-CoT
Zebra‑CoT
A diverse large-scale dataset for interleaved vision‑language reasoning traces.
Dataset Description
Zebra‑CoT is a diverse large‑scale dataset with 182,384 samples containing logically coherent interleaved text‑image reasoning traces across four major categories: scientific reasoning, 2D visual reasoning, 3D visual reasoning, and visual logic & strategic games.
Dataset Structure
Each example in Zebra‑CoT consists of:
Problem statement:… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT.Alpaca-CoT
Instruction-Finetuning Dataset Collection (Alpaca-CoT)
This repository will continuously collect various instruction tuning datasets. And we standardize different datasets into the same format, which can be directly loaded by the code of Alpaca model.
We also have conducted empirical study on various instruction-tuning datasets based on the Alpaca model, as shown in https://github.com/PhoebusSi/alpaca-CoT.
If you think this dataset collection is helpful to you, please like… See the full description on the dataset page: https://huggingface.co/datasets/QingyiSi/Alpaca-CoT.efficient-cot
To be written.
Visual-CoT
VisCoT Dataset Card
There is a shortage of multimodal datasets for training multi-modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations could possibly help the MLLM output intermediate interpretable attention area and enhance performance.
To fill the gap, we curate a visual CoT dataset. This dataset specifically focuses on identifying… See the full description on the dataset page: https://huggingface.co/datasets/deepcs233/Visual-CoT.codeforces-cots
Dataset Card for CodeForces-CoTs
Dataset description
CodeForces-CoTs is a large-scale dataset for training reasoning models on competitive programming tasks. It consists of 10k CodeForces problems with up to five reasoning traces generated by DeepSeek R1. We did not filter the traces for correctness, but found that around 84% of the Python ones pass the public tests.
The dataset consists of several subsets:
solutions: we prompt R1 to solve the problem and produce code.… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/codeforces-cots.OpenMath-Vision-CoT-10kISCSLP2026-CoT-TTS
ISCSLP 2026 CoT-TTS Dataset
Dataset Overview
This dataset is prepared for the ISCSLP 2026 CoT-TTS Challenge and is designed to support research on context-aware, expressive, and CoT-guided speech generation. It is constructed from speech-rich media sources, including films, TV dramas, radio dramas, and short dramas, where dialogue often contains rich conversational context, speaker interactions, scene changes, and emotional variation. Each sample is organized… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/ISCSLP2026-CoT-TTS.Step-CoT
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
Dataset Summary
Step-CoT is a large-scale medical reasoning dataset designed to improve interpretability in Medical Visual Question Answering (Med-VQA). It contains over 10,000 real clinical chest X-ray cases and 70,000 VQA pairs, each structured into a seven-step diagnostic workflow that mirrors clinical reasoning:
Abnormal Radiodensity Detection
Lesion Distribution
Radiographic Pattern… See the full description on the dataset page: https://huggingface.co/datasets/fl-15o/Step-CoT.Taur_CoT_Analysis_Project___gpt-4o-2024-08-06Audio-Reasoner-CoTAMMLU-Pro-CoT-Train-43Klibero_cotThis dataset was created using LeRobot.
It contains embodied Chain-of-Thought (CoT) demonstrations for the LIBERO benchmark, featuring paired reasoning and action traces. It was curated as part of the DeepThinkVLA project using a two-stage data engine that distills key frames with a cloud LVLM and scales to full trajectories via a fine-tuned local VLM.
Paper: DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
Repository: https://github.com/OpenBMB/DeepThinkVLA… See the full description on the dataset page: https://huggingface.co/datasets/yinchenghust/libero_cot.CottonWeedDet12
Dataset Card for CottonWeedDet12
CottonWeedDet12 is a 12-class weed object-detection dataset for cotton production systems in the southern U.S., consisting of 5,648 RGB field images with 9,370 bounding box annotations collected in Michigan State University MEFAS field trials during 2021-2022. It is the companion dataset for the YOLOWeeds benchmark of YOLO object detectors.
This is a FiftyOne dataset with 5648 samples.
Installation
If you haven't already, install… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/CottonWeedDet12.LLaVA-CoT-100k
Dataset Card for LLaVA-CoT
The LLaVA-CoT-100k dataset is introduced in the paper LLaVA-CoT: Let Vision Language Models Reason Step-by-Step. This dataset is designed to enable Vision-Language Models (VLMs) to perform autonomous multistage reasoning, integrating samples from various visual question-answering sources with structured reasoning annotations. It aims to address the challenges VLMs face in systematic and structured reasoning for complex visual question-answering tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k.Scaffold-CoT
Scaffold-CoT
Structured chain-of-thought training data with 3,726,548 examples in 76 JSONL shards.
Fields
Every row has exactly four top-level fields:
Field
Contents
metadata
domain, subdomain, difficulty, length_bucket
input
Ordered user messages as {index, content} objects
cot
Ordered {index, type, content} events, including reasoning, tool calls, and tool results
output
Ordered final assistant answers as {index, content} objects
The index… See the full description on the dataset page: https://huggingface.co/datasets/Specific-Labs/Scaffold-CoT.embodied_features_and_demos_liberoDataset for Embodied Chain-of-Thought Reasoning for LIBERO-90, as used by ECoT-Lite.
TFDS Demonstration Data
The TFDS dataset contains successful demonstration trajectories for LIBERO-90 (50 trajectories for each of 90 tasks). It was created by rolling out the actions provided in the original LIBERO release and filtering out all unsuccessful ones, leaving 3917 successful demo trajectories. This is done via a modified version of a script from the MiniVLA codebase. In addition to… See the full description on the dataset page: https://huggingface.co/datasets/Embodied-CoT/embodied_features_and_demos_libero.Meta-CoT-21-Tasks-Benchmmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples.
MMLU-Pro-CoT-Train-Labeled
Dataset Details
Modality: Text
Format: CSV
Size: 10K - 100K rows
Total Rows: 84,098
License: MIT
Libraries Supported: datasets, pandas, croissant
Structure
Each row in the dataset includes:
question: The query posed in the dataset.
answer: The correct response.
category: The domain of the question (e.g., math, science).
src: The source of the question.
id: A unique identifier for each entry.
chain_of_thoughts: Step-by-step reasoning steps leading to the answer.
labels:… See the full description on the dataset page: https://huggingface.co/datasets/UW-Madison-Lee-Lab/MMLU-Pro-CoT-Train-Labeled.Taur_CoT_Analysis_Project___gpt-4o-mini-2024-07-18cothanhaudiommlu-pro-CoTAPIGen-MT-5k-with-cot-v1-deepseek_deepseekactivations-Qwen3.5-9B
activations-Qwen3.5-9B
Residual-stream activations of Qwen/Qwen3.5-9B on hinted MCQ rollouts, for training chain-of-thought faithfulness probes (Detecting-CoT-Unfaithfulness-Attention-Probes).
<dataset>_<hint>.safetensors (or _partNN): one file per dataset-intervention, tensors {sample_type}_{original_index}_layer_{layer} → bfloat16 [n_cot_tokens, 4096], decoder blocks [3, 7, 11, 15, 19, 23, 27, 31].
Span: chain_of_thought: tokens strictly between the reasoning delimiters… See the full description on the dataset page: https://huggingface.co/datasets/cot-unfaithfulness-lasr/activations-Qwen3.5-9B.CoT-Collection"""
_LICENSE = "CC BY 4.0"
_HOMEPAGE = "https://github.com/kaistAI/CoT-Collection"
_LANGUAGES = {
"en": "English",
}
# _ALL_LANGUAGES = "all_languages"
class CoTCollectionMultiConfig(datasets.BuilderConfig):numina-cot-aops-forum
Saravana-Polisetti/numina-cot-aops-forum
Source-specific slice of AI-MO/NuminaMath-CoT.
Source: aops_forum
Rows: 30192
Columns: source, problem, solution
Usage
from datasets import load_dataset
ds = load_dataset("Saravana-Polisetti/numina-cot-aops-forum", split="train")
print(ds[0])
Math-CoT-44k-Qwen3-32b-n32-16384-with-logprob-and-entropy
Qwen3-32B Math n32 16384 (44k Queries)
This dataset contains multi-sampled rollout traces from Qwen3-32B on around 44k math queries.
For each query, the model is rolled out 32 times with a maximum generation length of 16384 tokens.
Each response is annotated with answer correctness (acc_reward), and includes token-level statistics (action_entropy, action_log_probs) for further analysis and research.
Resources
Paper: Rethinking Generalization in Reasoning SFT: A… See the full description on the dataset page: https://huggingface.co/datasets/jasonrqh/Math-CoT-44k-Qwen3-32b-n32-16384-with-logprob-and-entropy.
