datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
equational-theories-selected-problems
Equational Theories Selected Problems
Update (September 11, 2026)
This dataset was updated on September 11, 2026.
Main changes:
released the official Stage 2 evaluation problems: stage2_evaluation_main (200 problems; ground truth withheld — answer is null until Stage 2 concludes) and stage2_evaluation_research (100 order-5 research problems with no ground truth)
added metadata/stage2_evaluation_main.json and metadata/stage2_evaluation_research.json… See the full description on the dataset page: https://huggingface.co/datasets/SAIRfoundation/equational-theories-selected-problems.200k_HEAVY_gpt4o-description-gpt4omini-code_generated_problemsHere is the dataset of ~100k synthetic data generated by 162 seeds.
We generate the dataset with the following steps and two approaches:
Generate ~110k descriptions by GPT4o.
Approach 1: Generate ~110k codes follow each description by GPT4o-mini.
Approach 2: Generate ~110k codes follow each description by GPT4o-mini and suggest it to use specific library functions.
Run the ~220k codes and do auto-filtering.
Get the final ~200k legitimate ARC-like tasks with examples.
SciCode-Programming-Problems
DATA3: Programming Problems Generation Dataset
Dataset Overview
DATA3 is a large-scale programming problems generation dataset that contains AI-generated programming problems inspired by real scientific computing code snippets. The dataset consists of 22,532 programming problems, each paired with a comprehensive solution. These problems focus on scientific computing concepts such as numerical algorithms, data analysis, mathematical modeling, and computational methods in… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Programming-Problems.cses-fi-competitive-coding-problemshle-failed-problems-byQwen3-32bcountdown_problemsMathSmith-Hard-ProblemsMathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
Overview
This dataset is a collection of problems generated by the MathSmith-Hard Problem-Synthesizer.
Dataset Structure
Each record is a JSON object with the following fields:
{
"problem": "<str>", // The generated math problem
"rationale": "<str>" // The ratioanle process of question generation… See the full description on the dataset page: https://huggingface.co/datasets/Jasaxion/MathSmith-Hard-Problems.IMO-2026-Problems
IMO 2026 Problems
The six IMO 2026 problem statements, indexed from 0 through 5 in contest order.
IDs 0–2 are from Day 1, and IDs 3–5 are from Day 2.
Schema
id: zero-based problem identifier (0 corresponds to Problem 1).
day: contest day (1 or 2).
problem: complete English problem statement.
Source
Extracted from the problem statements in SignalPilot Labs' AutoFyn IMO 2026 results:… See the full description on the dataset page: https://huggingface.co/datasets/bogoconic1/IMO-2026-Problems.100k-gpt4omini-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds.
We generate the dataset with the following steps:
Generate 120k descriptions by GPT4o-mini.
Generate 120k codes follow each description by GPT4o-mini.
Run the 120k codes and do auto-filtering.
Get the final 100k legitimate ARC-like tasks with examples.
100k-gpt4-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds.
We generate the dataset with the following steps:
Generate 120k descriptions by GPT4.
Generate 120k codes follow each description by GPT4o-mini.
Run the 120k codes and do auto-filtering.
Get the final 100k legitimate ARC-like tasks with examples.
crest-codeforces-annotated-problemsCREST (Code, Ratings, Editorials, Statements, and Tags) is a dataset of 8,941 annotated Codeforces problems. For each problem, the dataset includes:
The problem statement and tutorial (editorial) text, both of which are math-rich and contain LaTeX-formatted mathematical notation.
Reference solution code from the tutorial, when available.
A set of algorithmic tags.
A numerical difficulty rating.
The dataset supports tasks such as multilabel tag classification and rating regression from… See the full description on the dataset page: https://huggingface.co/datasets/ZaniteA/crest-codeforces-annotated-problems.agent-wicked-problems-40k
HUMMBL 40k Multi-Agent Wicked Problems & Coordination Corpus
A foundational 40,171-event empirical dataset capturing real-world multi-agent coordination, epistemic problem decomposition, failure mode taxonomies, and strategic intelligence surges generated across the HUMMBL autonomous agent fleet.
Dataset Overview
The dataset provides structured visibility into how autonomous agents navigate complex, ill-defined ("wicked") problems, coordinate across distributed… See the full description on the dataset page: https://huggingface.co/datasets/hummbl-hf/agent-wicked-problems-40k.problem_solving-reasoning-pashto-plus
Problem Solving & Reasoning Multilingual Dataset
This repository contains a specialized dataset focused on logical reasoning, problem-solving, and step-by-step cognitive workflows across multiple regional languages: Pashto (ps), Arabic (ar), Farsi (fa), Sindhi (sd), and Urdu (ur).
Dataset Overview
Languages: Pashto (پښتو), Arabic (العربية), Farsi (فارسی), Sindhi (سنڌي), Urdu (اردو)
Domain: Logical Reasoning, Problem Solving, Cognitive SFT
License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/problem_solving-reasoning-pashto-plus.ai-research-problems
AI Research Problems 1M
Summary
This dataset contains 1,000,000 synthetic research-ideation candidates across AI, machine learning, LLMs, RAG, AI agents, MCP, computer vision, robotics, safety, MLOps, and related fields.
Important warning
These records are synthetic combinations for research ideation. They are not claims that the problems are novel, unsolved, or absent from the literature. A researcher must verify novelty using papers, benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/bala5046/ai-research-problems.hardware-cvdp-problems
Hardware Design AI Training Dataset
This dataset contains processed hardware design problems and Verilog code for training AI models.
Contents
CVDP Problems: 160 evaluation problems organized by domain and complexity
Training Data: Instruction-code pairs for hardware design
Metadata: Rich annotations for each problem
Usage
from datasets import load_dataset
dataset = load_dataset("AbiralArch/hardware-cvdp-problems")
Categories
Module Generation… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-problems.orca-math-word-problems-100k-en-zh-mix100k English and Chinese mixed version of microsoft/orca-math-word-problems-200k
mengcoder-problemsCodeARC-ProblemsCodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
Paper: https://arxiv.org/pdf/2503.23145
Code: https://github.com/Anjiang-Wei/CodeARC
Website: https://anjiang-wei.github.io/CodeARC-Website/
Dataset: https://huggingface.co/datasets/anjiangwei/CodeARC-Problems
10 Input-Output examples for each problem: https://huggingface.co/datasets/anjiangwei/CodeARC-Invocations
Fine-tuned models:… See the full description on the dataset page: https://huggingface.co/datasets/anjiangwei/CodeARC-Problems.SciCode-Programming-Problems
DATA3: Programming Problems Generation Dataset
Dataset Overview
DATA3 is a large-scale programming problems generation dataset that contains AI-generated programming problems inspired by real scientific computing code snippets. The dataset consists of 22,532 programming problems, each paired with a comprehensive solution. These problems focus on scientific computing concepts such as numerical algorithms, data analysis, mathematical modeling, and computational… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Programming-Problems.olympiad_problems_5Dataset собранный из олимпиады 2025/26 учебного года. 5 класс
Олимпиада - "Звезда" (https://zv.susu.ru/)
Язык - Russian
Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedorca-math-word-problems-200k-turkmen
Turkmen Orca Math Word Problems 200k Dataset
Overview
This dataset is a Turkmen translation of the original microsoft/orca-math-word-problems-200k dataset. The Orca Math Word Problems dataset contains 200,000 high-quality math word problems and their solutions. This Turkmen version aims to extend the accessibility of math problem-solving datasets to the Turkmen language community.
Dataset Details
Original Dataset: microsoft/orca-math-word-problems-200k… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/orca-math-word-problems-200k-turkmen.adaption-arithmetic-algebra-word-problems
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-arithmetic-algebra-word-problems
This dataset features instruction and response pairs containing grade-school arithmetic and algebra word problems paired with step-by-step solutions. Problems cover multi-step arithmetic, percentages, ratios, linear equations, and simple systems solvable in two to five steps. Each completion demonstrates explicit reasoning and concludes with a… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/adaption-arithmetic-algebra-word-problems.Math_Word-Problems-ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
DSA-Coding-Problems-and-Solutions-Dataset
Dataset Description
This dataset is a large-scale collection of Data Structures and Algorithms (DSA) code, containing 12,385 code files with 3.86 million lines of code and 25.01 million lexical tokens, designed to support the development of advanced code generation models, programming assistants, software engineering AI systems, and code intelligence applications.
It consists of real-world DSA implementations covering a wide range of algorithms, data structures, problem-solving… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/DSA-Coding-Problems-and-Solutions-Dataset.lemonseed-word-problems
lemonseed-word-problems
LemonSeed — 13-category arithmetic word problems with plan + scratchpad.
Contents
wordproblems.jsonl (2500 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
orca-math-word-problems-193k-korean-jsonl원본 데이터셋
https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k
https://huggingface.co/datasets/kuotient/orca-math-word-problems-193k-korean
Citation
@misc{mitra2024orcamath,
title={Orca-Math: Unlocking the potential of SLMs in Grade School Math},
author={Arindam Mitra and Hamed Khanpour and Corby Rosset and Ahmed Awadallah},
year={2024},
eprint={2402.14830},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
orca-math-word-problems-200k-hindi-filteredVietnamese-Intermediate-Reality-Math-ProblemsThis is our gather 500 dataset of Simple Reality Math Problems written in Vietnamese.
adaption-bg-math-word-problems
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-bg_math_word_problems
This dataset contains 1,319 Bulgarian language math word problems paired with step-by-step solutions that include intermediate calculations. The content covers various arithmetic scenarios such as currency conversion, cost estimation, and area calculations, formatted as prompt-completion pairs. It is designed for evaluating or training models on… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/adaption-bg-math-word-problems.
