datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Data source
Prompts from AM-DeepSeek-R1-0528-Distilled
Thinking traces and outputs distilled from gpt-oss-120b
Translated with command-a-translate and DeepSeek-V3
Languages (44)
Language
Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.agentic-thinker-v1tiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Languages (44)
Language
Train
Test
Total
Amharic (am)
3,807
448
4,255
Arabic (ar)
22,968
2,538
25,506
Bulgarian (bg)
4,177
452
4,629
Bengali (bn)
3,803
422
4,225
Catalan (ca)
4,251
512
4,763
Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.tulu-3-thinker-classified-no_math-ifeval-ref-based-45kCoSQA_PlusCoSQA+, a code search dataset pairing high-quality queries (reused from CoSQA) with multiple suitable codes.
We collect code candidates from diverse sources and form candidate pairs by pairing queries with these codes.
Utilizing the power of large language models (LLMs), we automate pair annotation, filtering, and code generation for queries without suitable matches.
related links:
arXiv:
2406.11589 CoSQA+: Enhancing Code Search Dataset with Matching Code (arxiv.org)
github:… See the full description on the dataset page: https://huggingface.co/datasets/thinkerhui/CoSQA_Plus.M-Thinker-SFT-dataVAPO-Thinker-train36kARM-Thinker-Data
ARM-Thinker-Data
Paper | Github Repository
📊 Data Introduction
This repository contains the datasets used for training ARM-Thinker, an Agentic Multimodal Reward Model that performs evidence-grounded reasoning through tool use and visual grounding.
The current dataset is annotated by Qwen3-VL-235B-A22B-Instruct, Qwen3-VL-235B-A22B-Thinking, and GPT-4o, with all data files organized under the qwen/ directory. We are also planning to release an additional version… See the full description on the dataset page: https://huggingface.co/datasets/internlm/ARM-Thinker-Data.Thinker-XMLSystem prompt suggestion:
You are a world-class AI system. Always respond in strict XML format with your reasoning steps within the <im_reasoning> XML tag. Each reasoning step should represent one unit of thought. Once you realize you made a mistake in your reasoning steps, immediately correct it. Place your final response outside the XML tag. Adhere to this XML structure without exception.
KAG-Thinker-training-datasetGeneralThought-430K-filtered-thinkerUltra-Thinker-30k
Ultra-Thinker 🧠
A comprehensive collection of high-quality reasoning and conversational datasets designed to enhance the thinking capabilities of large language models.
Overview
Ultra-Thinker aggregates diverse sources focused on:
Chain-of-thought reasoning - Step-by-step logical deduction
Mathematical problem solving - Complex computational challenges
Advanced logical reasoning - Multi-step inference and analysis
This curated compilation combines cutting-edge… See the full description on the dataset page: https://huggingface.co/datasets/Stormtrooperaim/Ultra-Thinker-30k.DavidAU__DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B-details
Dataset Card for Evaluation run of DavidAU/DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B
Dataset automatically created during the evaluation run of model DavidAU/DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DavidAU__DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B-details.MUPO-Thinker-train36kThinker-2tulu-3-thinker-rewritten-math-27kfhai50032__Unaligned-Thinker-PHI-4-details
Dataset Card for Evaluation run of fhai50032/Unaligned-Thinker-PHI-4
Dataset automatically created during the evaluation run of model fhai50032/Unaligned-Thinker-PHI-4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fhai50032__Unaligned-Thinker-PHI-4-details.thinkerA Chain-of-Thought (CoT) dataset that contains traces of complex and sophisticated reasoning, to mimic the "thinking" process of OpenAI's o1. Wrap the contents of the reasoning column in some XML tag (such as <reasoning>).
Raw .jsonl dataset file can be found under the Files and Versions tab.
ewre324__Thinker-SmolLM2-135M-Instruct-Reasoning-details
Dataset Card for Evaluation run of ewre324/Thinker-SmolLM2-135M-Instruct-Reasoning
Dataset automatically created during the evaluation run of model ewre324/Thinker-SmolLM2-135M-Instruct-Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-SmolLM2-135M-Instruct-Reasoning-details.unfiltered-thinker
Unfiltered-Thinker: A Dataset for Intermediate Cognitive Reasoning
A corpus of 1,909 samples designed to showcase intermediate thinking, cognitive processes, and structured emotional reasoning.
Source: UnfilteredAI/unfiltered-thinker on Hugging Face
⚠️ Content Warning: This dataset contains content that will be considered offensive, disturbing, or explicit. This includes discussions of dark humor, profanity, criminal activity, violence, substance use, and psychological distress. It… See the full description on the dataset page: https://huggingface.co/datasets/UnfilteredAI/unfiltered-thinker.ewre324__Thinker-Qwen2.5-0.5B-Instruct-Reasoning-details
Dataset Card for Evaluation run of ewre324/Thinker-Qwen2.5-0.5B-Instruct-Reasoning
Dataset automatically created during the evaluation run of model ewre324/Thinker-Qwen2.5-0.5B-Instruct-Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-Qwen2.5-0.5B-Instruct-Reasoning-details.bunnycore__Qwen2.5-3B-RP-Thinker-details
Dataset Card for Evaluation run of bunnycore/Qwen2.5-3B-RP-Thinker
Dataset automatically created during the evaluation run of model bunnycore/Qwen2.5-3B-RP-Thinker
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qwen2.5-3B-RP-Thinker-details.luffy_thinkerDISTILLATION-VIBE-THINKER
USED ADPATION LABS AUTO SCIENCETS DATA PREPARATION AND DISTALATE DATA OF VIBE THINKER
VAPO-Thinker-val1kThinker-2-Longbunnycore__Qwen2.5-3B-RP-Thinker-V2-details
Dataset Card for Evaluation run of bunnycore/Qwen2.5-3B-RP-Thinker-V2
Dataset automatically created during the evaluation run of model bunnycore/Qwen2.5-3B-RP-Thinker-V2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qwen2.5-3B-RP-Thinker-V2-details.nbeerbower-Purpura-DPO-thinker-rawUnfiltered
System promt for creating a dataset:
You are an expert AI assistant specializing in text generation. Your task is to reverse-engineer the thought process that leads to a given textual `response`.
Based on the user's `prompt` and the final `response` text, generate a plausible, detailed reasoning process of an LLM.
This reasoning should cover:
1. **Analysis of the User's Prompt:** Deconstruct the user's request, identifying explicit constraints (like length, format) and implicit… See the full description on the dataset page: https://huggingface.co/datasets/Disya/nbeerbower-Purpura-DPO-thinker-raw.Detailed-Thinker
Detailed Thinker (snapshot)
This is an early snapshot of Detailed Thinker, started 12 August 2026. The set is still in active development. Treat this release as a work-in-progress cut (around 300 examples), not a finished corpus.
What this is for
Shallow build requests often get shallow answers. If you say "build a barbershop simulator," a model can sketch something but it usually skips real depth on functionality, constraints, and follow-through.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/Detailed-Thinker.ewre324__Thinker-Llama-3.2-3B-Instruct-Reasoning-details
Dataset Card for Evaluation run of ewre324/Thinker-Llama-3.2-3B-Instruct-Reasoning
Dataset automatically created during the evaluation run of model ewre324/Thinker-Llama-3.2-3B-Instruct-Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-Llama-3.2-3B-Instruct-Reasoning-details.
