datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ethical-Reasoning-in-Mental-Health-v1This repository contains the dataset for the paper EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI.
Overview
Ethical-Reasoning-in-Mental-Health-v1 (EthicsMH) is a carefully curated dataset focused on ethical decision-making scenarios in mental health contexts.This dataset captures the complexity of real-world dilemmas faced by therapists, psychiatrists, and AI systems when navigating critical issues such as confidentiality, autonomy, and bias.
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/UVSKKR/Ethical-Reasoning-in-Mental-Health-v1.ReasoningShield-Dataset
🤗 Dataset Card for ReasoningShield
🛡 1. Dataset Overview
ReasoningShield Dataset is the first comprehensive, well-structured dataset designed to train and evaluate models for detecting hidden safety risks in reasoning traces of Large Reasoning Models (LRMs), spanning 10 risk categories and 3 safety levels. It consists of:
ReasoningShield-Train: 7,000 human-AI annotated (Query… See the full description on the dataset page: https://huggingface.co/datasets/ReasoningShield/ReasoningShield-Dataset.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.sober_reasoning
🧠 Sober Reasoning: Evaluation Logs
This repository hosts evaluation logs and outputs from our paper:
"A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility"
📄 Paper📊 Leaderboard💻 Evaluation Code
🗂️ Repository Structure
Evaluation logs are organized by the cluster used during inference to highlight hardware-induced variance in model performance (see Section 3.3 of the paper).
sober_reasoning/
├── cluster_A/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/bethgelab/sober_reasoning.Logical_Reasoning_Chainsqwen35-9b-reasoning-share-coop-random-50
What this is
Cooperative two-agent coding dataset: 50 task pairs across 15 repos (random-50 subset), generated
with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a reasoning-share prompt variant —
agents share their internal reasoning and analysis with each other before and during implementation,
giving each agent visibility into the other's thought process to improve integration. All 50 pairs
were successfully evaluated.
At a glance
Field
Value… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-reasoning-share-coop-random-50.LLM_reasoning_bakeoffmagpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-changedtokenNatural
ReasonAudio Natural
ReasonAudio Natural is the real-world subtask of ReasonAudio, a benchmark for evaluating reasoning beyond semantic matching in text-audio retrieval. It contains 100 human-written queries and a retrieval corpus of 1,000 ten-second clips sampled from AudioSet-Strong. Each query may have multiple relevant clips.
Dataset statistics
Item
Count
Queries
100
Corpus entries
1,000
Positive relevance judgments
355
Mean positives per query… See the full description on the dataset page: https://huggingface.co/datasets/ReasonAudio/Natural.R1-Reasoning-Unstructured-To-Structured
MasterControl AIML Team 🚀
Overview
The MasterControl AIML team supports the Hugging Face initiative of re-creating DeepSeek R1 training, recognizing it as one of the most impactful open-source projects today.
We aim to contribute to reasoning datasets, specifically those where:
A real-world problem involves generating complex structured output
It is accompanied by step-by-step reasoning and unstructured input
Challenges in Integrating Generative AI… See the full description on the dataset page: https://huggingface.co/datasets/MasterControlAIML/R1-Reasoning-Unstructured-To-Structured.agentic-reasoning-benchmark
Agentic & Reasoning Benchmark (ARB) – Expanded
Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning.
Überblick
Eigenschaft
Wert
Anzahl Beispiele
2.550
Kategorien
8
Schwierigkeitsgrade
easy / medium / hard
Formate
CSV + JSON
Reproduzierbarkeit
Generator-Skript (seed=42) enthalten
Lizenz
CC-BY-4.0
Kategorien
Kategorie
Anzahl
Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-ReasonExtracted Chemistry Physics Asronomy Math and Logic portions from the original.
Script used for the extraction:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-Reason/resolve/main/find-science-fine.py
multilevel-legal-reasoning
Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations
Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi
Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai.
🧭 Purpose and Scope
The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.synthetic-indian-logical-reasoning-CoTyes
DAG-Reasoning-DeepSeek-R1-0528Click here to support our open-source dataset and model releases!
DAG-Reasoning-DeepSeek-R1-0528 is a dataset focused on analysis and reasoning, creating directed acyclic graphs testing the limits of DeepSeek R1 0528's graph-reasoning skills!
This dataset contains:
4.08k synthetically generated prompts to create directed acyclic graphs in response to user input, with all responses generated using DeepSeek R1 0528.
All responses contain a multi-step thinking process to perform effective… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DAG-Reasoning-DeepSeek-R1-0528.chess_spatial_reasoning_10kMedQA-ReasonReference
@article{wang2024jmlr,
title={JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability},
author={Wang, Junda and Yang, Zhichao and Yao, Zonghai and Yu, Hong},
journal={arXiv preprint arXiv:2402.17887},
year={2024}
}
Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-llama3.1magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-changedtoken-mistralDES-Reasoning-DeepSeek-V3.1Click here to support our open-source dataset and model releases!
DES-Reasoning-DeepSeek-V3.1 is a dataset focused on analysis and reasoning, creating discrete event simulations testing the limits of DeepSeek V3.1's simulation, Python scripting, and analysis skills!
This dataset contains:
4.03k synthetically generated prompts to create discrete event simulations and analysis chat in response to user input, with all responses generated using DeepSeek V3.1.
All responses contain a multi-step… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DES-Reasoning-DeepSeek-V3.1.easy_turkish_math_reasoning
Easy Turkish Math Reasoning
Dataset Summary
The Easy Turkish Math Reasoning dataset is the first phase of a multi-stage curriculum learning pipeline designed to enhance the reasoning abilities of compact language models. This dataset focuses on elementary-level arithmetic and logic problems in Turkish, serving as a warm-up stage for supervised fine-tuning (SFT).
Use Case
Primarily used for:
Bootstrapping reasoning ability in Turkish for compact LLMs.
Phase 1… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/easy_turkish_math_reasoning.reasoning-trajectory-stability-controls-v0.1
Reasoning Trajectory Stability Controls v0.1
A SIOS research dataset for detecting whether a reasoning trajectory remains structurally stable, identifying the control introduced into the trajectory, locating where that control first becomes operationally visible, and determining whether the control succeeds or fails.
Repository:
ClarusC64/reasoning-trajectory-stability-controls-v0.1
Version:
0.1.0
Publisher:
Clarus Invariant
Framework:
SIOS
Dataset identity… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-trajectory-stability-controls-v0.1.morse-500
MORSE-500 Benchmark
🔥 News
May 15, 2025: We release MORSE-500, 500 programmatically generated videos across six reasoning categories: abstract, mathematical, physical, planning, spatial, and temporal, to stress-test multimodal reasoning. Frontier models including OpenAI o3 and Gemini 2.5 Pro score lower than… See the full description on the dataset page: https://huggingface.co/datasets/video-reasoning/morse-500.rebus-reasoningSYSTEM_PROMPT = """# Come risolvere un rebus
Sei un esperto risolutore di giochi enigmistici. Il seguente gioco contiene una frase cifrata (**Rebus**) nella quale alcune parole sono state sostituite da delle **Definizioni** di cruciverba fornite tra parentesi quadre. Tutte le parole e le frasi sono esclusivamente in lingua italiana. Lo scopo del gioco è quello di identificare le **Risposte** corrette e sostituirle alle definizioni nel Rebus, producendo una **Prima Lettura** che verrà poi… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/rebus-reasoning.reasoning-drift-onset-detection-v0.2A SIOS structured reasoning-state benchmark for detecting when a reasoning trajectory loses a governing constraint, identifying the structural form of that drift, and assessing whether the failure is repaired.
Repository:
ClarusC64/reasoning-drift-onset-detection-v0.2
Version:
0.2.0
Publisher:
Clarus Invariant
Framework:
SIOS
Benchmark identity
Reasoning Drift Onset Detection v0.2 is not a single-label classification benchmark.
It is a structured reasoning-state benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-drift-onset-detection-v0.2.finance-reasoning-turkish
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti emre (Davut Emre Tasar, Enes Bulut) tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: emre/finance-reasoning-turkish
🔗 Derleyen Platform: VeriPazarı
Türkçe Gelişmiş Akıl Yürütme Veri Seti (Finans Soru-Cevap)
Lisans
Bu veri seti Sadece Akademik… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/finance-reasoning-turkish.Creative_Stories_Logical_Reasoningmedium_turkish_math_reasoning
Dataset Summary
The Medium Turkish Math Reasoning dataset is Phase 2 of a curriculum learning pipeline to teach compact models multi-step reasoning in Turkish. It includes moderately difficult math problems involving multiple reasoning steps, such as two-part arithmetic, comparisons, and logical reasoning.
Use Case
This dataset is ideal for:
Continuing SFT after foundational training with simpler problems.
Bridging the gap between basic arithmetic and complex GSM8K-style… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/medium_turkish_math_reasoning.ReasonAVEdit-Bench-Landscape
ReasonAVEdit-Bench — Landscape Edition
2,313 samples, every one of them landscape, for instruction-guided joint audio-video
editing. Two halves:
Half
N
What it is
Newly sampled
1,200
15 reasoning sub-tracks, selected from measured evidence
Legacy tracks
1,113
Every landscape sample from the hand-curated AniAVEditBench tracks
The benchmark measures not just the final edit but the multimodal reasoning behind it:
which region to change, which sound source to… See the full description on the dataset page: https://huggingface.co/datasets/bigfacing/ReasonAVEdit-Bench-Landscape.
