Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SAIRfoundation /equational-theories-selected-problems Equational Theories Selected Problems Update (September 11, 2026) This dataset was updated on September 11, 2026. Main changes: released the official Stage 2 evaluation problems: stage2_evaluation_main (200 problems; ground truth withheld — answer is null until Stage 2 concludes) and stage2_evaluation_research (100 order-5 research problems with no ground truth) added metadata/stage2_evaluation_main.json and metadata/stage2_evaluation_research.json… See the full description on the dataset page: https://huggingface.co/datasets/SAIRfoundation/equational-theories-selected-problems.tabular1K<n<10K11 likes7k downloads29d agoHugging Face02barc0 /200k_HEAVY_gpt4o-description-gpt4omini-code_generated_problemsHere is the dataset of ~100k synthetic data generated by 162 seeds. We generate the dataset with the following steps and two approaches: Generate ~110k descriptions by GPT4o. Approach 1: Generate ~110k codes follow each description by GPT4o-mini. Approach 2: Generate ~110k codes follow each description by GPT4o-mini and suggest it to use specific library functions. Run the ~220k codes and do auto-filtering. Get the final ~200k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M11 likes549 downloads2y agoHugging Face03SciCodePile /SciCode-Programming-Problems DATA3: Programming Problems Generation Dataset Dataset Overview DATA3 is a large-scale programming problems generation dataset that contains AI-generated programming problems inspired by real scientific computing code snippets. The dataset consists of 22,532 programming problems, each paired with a comprehensive solution. These problems focus on scientific computing concepts such as numerical algorithms, data analysis, mathematical modeling, and computational methods in… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Programming-Problems.texttext-generation10K<n<100K0 likes526 downloads7mo agoHugging Face04juvi21 /cses-fi-competitive-coding-problemstextn<1K5 likes321 downloads2y agoHugging Face05DeL-TaiseiOzaki /hle-failed-problems-byQwen3-32btext1K<n<10K0 likes315 downloads1y agoHugging Face06Parallel-Reasoning /countdown_problemstabular100K<n<1M0 likes213 downloads1y agoHugging Face07Jasaxion /MathSmith-Hard-ProblemsMathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy Overview This dataset is a collection of problems generated by the MathSmith-Hard Problem-Synthesizer. Dataset Structure Each record is a JSON object with the following fields: { "problem": "<str>", // The generated math problem "rationale": "<str>" // The ratioanle process of question generation… See the full description on the dataset page: https://huggingface.co/datasets/Jasaxion/MathSmith-Hard-Problems.textquestion-answering100K<n<1M1 likes181 downloads11mo agoHugging Face08bogoconic1 /IMO-2026-Problems IMO 2026 Problems The six IMO 2026 problem statements, indexed from 0 through 5 in contest order. IDs 0–2 are from Day 1, and IDs 3–5 are from Day 2. Schema id: zero-based problem identifier (0 corresponds to Problem 1). day: contest day (1 or 2). problem: complete English problem statement. Source Extracted from the problem statements in SignalPilot Labs' AutoFyn IMO 2026 results:… See the full description on the dataset page: https://huggingface.co/datasets/bogoconic1/IMO-2026-Problems.tabularn<1K0 likes154 downloads3mo agoHugging Face09barc0 /100k-gpt4omini-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds. We generate the dataset with the following steps: Generate 120k descriptions by GPT4o-mini. Generate 120k codes follow each description by GPT4o-mini. Run the 120k codes and do auto-filtering. Get the final 100k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M1 likes123 downloads2y agoHugging Face10barc0 /100k-gpt4-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds. We generate the dataset with the following steps: Generate 120k descriptions by GPT4. Generate 120k codes follow each description by GPT4o-mini. Run the 120k codes and do auto-filtering. Get the final 100k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M1 likes102 downloads2y agoHugging Face11ZaniteA /crest-codeforces-annotated-problemsCREST (Code, Ratings, Editorials, Statements, and Tags) is a dataset of 8,941 annotated Codeforces problems. For each problem, the dataset includes: The problem statement and tutorial (editorial) text, both of which are math-rich and contain LaTeX-formatted mathematical notation. Reference solution code from the tutorial, when available. A set of algorithmic tags. A numerical difficulty rating. The dataset supports tasks such as multilabel tag classification and rating regression from… See the full description on the dataset page: https://huggingface.co/datasets/ZaniteA/crest-codeforces-annotated-problems.texttext-classification1K<n<10K0 likes94 downloads9mo agoHugging Face12hummbl-hf /agent-wicked-problems-40k HUMMBL 40k Multi-Agent Wicked Problems & Coordination Corpus A foundational 40,171-event empirical dataset capturing real-world multi-agent coordination, epistemic problem decomposition, failure mode taxonomies, and strategic intelligence surges generated across the HUMMBL autonomous agent fleet. Dataset Overview The dataset provides structured visibility into how autonomous agents navigate complex, ill-defined ("wicked") problems, coordinate across distributed… See the full description on the dataset page: https://huggingface.co/datasets/hummbl-hf/agent-wicked-problems-40k.tabulartext-classification10K<n<100K1 likes93 downloads13d agoHugging Face13nassimjp /problem_solving-reasoning-pashto-plus Problem Solving & Reasoning Multilingual Dataset This repository contains a specialized dataset focused on logical reasoning, problem-solving, and step-by-step cognitive workflows across multiple regional languages: Pashto (ps), Arabic (ar), Farsi (fa), Sindhi (sd), and Urdu (ur). Dataset Overview Languages: Pashto (پښتو), Arabic (العربية), Farsi (فارسی), Sindhi (سنڌي), Urdu (اردو) Domain: Logical Reasoning, Problem Solving, Cognitive SFT License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/problem_solving-reasoning-pashto-plus.text10K<n<100K0 likes91 downloads7d agoHugging Face14bala5046 /ai-research-problems AI Research Problems 1M Summary This dataset contains 1,000,000 synthetic research-ideation candidates across AI, machine learning, LLMs, RAG, AI agents, MCP, computer vision, robotics, safety, MLOps, and related fields. Important warning These records are synthetic combinations for research ideation. They are not claims that the problems are novel, unsolved, or absent from the literature. A researcher must verify novelty using papers, benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/bala5046/ai-research-problems.texttext-classification1M<n<10M1 likes88 downloads26d agoHugging Face15AbiralArch /hardware-cvdp-problems Hardware Design AI Training Dataset This dataset contains processed hardware design problems and Verilog code for training AI models. Contents CVDP Problems: 160 evaluation problems organized by domain and complexity Training Data: Instruction-code pairs for hardware design Metadata: Rich annotations for each problem Usage from datasets import load_dataset dataset = load_dataset("AbiralArch/hardware-cvdp-problems") Categories Module Generation… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-problems.texttext-generationn<1K0 likes79 downloads1y agoHugging Face16xd2333 /orca-math-word-problems-100k-en-zh-mix100k English and Chinese mixed version of microsoft/orca-math-word-problems-200k texttext-generation100K<n<1M2 likes74 downloads2y agoHugging Face17KidzRizal /mengcoder-problemstext10K<n<100K0 likes64 downloads6d agoHugging Face18anjiangwei /CodeARC-ProblemsCodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis Paper: https://arxiv.org/pdf/2503.23145 Code: https://github.com/Anjiang-Wei/CodeARC Website: https://anjiang-wei.github.io/CodeARC-Website/ Dataset: https://huggingface.co/datasets/anjiangwei/CodeARC-Problems 10 Input-Output examples for each problem: https://huggingface.co/datasets/anjiangwei/CodeARC-Invocations Fine-tuned models:… See the full description on the dataset page: https://huggingface.co/datasets/anjiangwei/CodeARC-Problems.text1K<n<10K1 likes63 downloads1y agoHugging Face19SciCode /SciCode-Programming-Problemsgated DATA3: Programming Problems Generation Dataset Dataset Overview DATA3 is a large-scale programming problems generation dataset that contains AI-generated programming problems inspired by real scientific computing code snippets. The dataset consists of 22,532 programming problems, each paired with a comprehensive solution. These problems focus on scientific computing concepts such as numerical algorithms, data analysis, mathematical modeling, and computational… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Programming-Problems.texttext-generation10K<n<100K1 likes57 downloads8mo agoHugging Face20DatasetsEval /olympiad_problems_5Dataset собранный из олимпиады 2025/26 учебного года. 5 класс Олимпиада - "Звезда" (https://zv.susu.ru/) Язык - Russian textn<1K0 likes57 downloads3mo agoHugging Face215CD-AI /Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedtexttext-generation100K<n<1M4 likes42 downloads3y agoHugging Face22mamed0v /orca-math-word-problems-200k-turkmen Turkmen Orca Math Word Problems 200k Dataset Overview This dataset is a Turkmen translation of the original microsoft/orca-math-word-problems-200k dataset. The Orca Math Word Problems dataset contains 200,000 high-quality math word problems and their solutions. This Turkmen version aims to extend the accessibility of math problem-solving datasets to the Turkmen language community. Dataset Details Original Dataset: microsoft/orca-math-word-problems-200k… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/orca-math-word-problems-200k-turkmen.textquestion-answering100K<n<1M1 likes29 downloads2y agoHugging Face23vinod-anbalagan /adaption-arithmetic-algebra-word-problems This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-arithmetic-algebra-word-problems This dataset features instruction and response pairs containing grade-school arithmetic and algebra word problems paired with step-by-step solutions. Problems cover multi-step arithmetic, percentages, ratios, linear equations, and simple systems solvable in two to five steps. Each completion demonstrates explicit reasoning and concludes with a… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/adaption-arithmetic-algebra-word-problems.textn<1K0 likes28 downloads1mo agoHugging Face24ChaoticNeutrals /Math_Word-Problems-ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing text10K<n<100K0 likes26 downloads2y agoHugging Face25InfoBayAI /DSA-Coding-Problems-and-Solutions-Datasetgated Dataset Description This dataset is a large-scale collection of Data Structures and Algorithms (DSA) code, containing 12,385 code files with 3.86 million lines of code and 25.01 million lexical tokens, designed to support the development of advanced code generation models, programming assistants, software engineering AI systems, and code intelligence applications. It consists of real-world DSA implementations covering a wide range of algorithms, data structures, problem-solving… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/DSA-Coding-Problems-and-Solutions-Dataset.tabulartext-generationn<1K0 likes26 downloads2d agoHugging Face26MichaelAnthony /lemonseed-word-problems lemonseed-word-problems LemonSeed — 13-category arithmetic word problems with plan + scratchpad. Contents wordproblems.jsonl (2500 rows) Format JSON Lines (.jsonl), one example per line. Provenance Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella. textquestion-answering1K<n<10K0 likes24 downloads2mo agoHugging Face27jaypyon /orca-math-word-problems-193k-korean-jsonl원본 데이터셋 https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k https://huggingface.co/datasets/kuotient/orca-math-word-problems-193k-korean Citation @misc{mitra2024orcamath, title={Orca-Math: Unlocking the potential of SLMs in Grade School Math}, author={Arindam Mitra and Hamed Khanpour and Corby Rosset and Ahmed Awadallah}, year={2024}, eprint={2402.14830}, archivePrefix={arXiv}, primaryClass={cs.CL} } text100K<n<1M0 likes21 downloads2y agoHugging Face28BhabhaAI /orca-math-word-problems-200k-hindi-filteredtext100K<n<1M0 likes19 downloads3y agoHugging Face29Mels22 /Vietnamese-Intermediate-Reality-Math-ProblemsThis is our gather 500 dataset of Simple Reality Math Problems written in Vietnamese. textn<1K0 likes18 downloads2y agoHugging Face30Ilia-Iliev /adaption-bg-math-word-problems This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-bg_math_word_problems This dataset contains 1,319 Bulgarian language math word problems paired with step-by-step solutions that include intermediate calculations. The content covers various arithmetic scenarios such as currency conversion, cost estimation, and area calculations, formatted as prompt-completion pairs. It is designed for evaluating or training models on… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/adaption-bg-math-word-problems.text1K<n<10K0 likes18 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.