Team Ai
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01microsoft /orca-math-word-problems-200k Dataset Card This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of SLMs in Grade School Math for details about the dataset construction. Dataset Sources Repository: microsoft/orca-math-word-problems-200k Paper: Orca-Math: Unlocking the potential of SLMs in Grade School Math Direct Use This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k.textquestion-answering100K<n<1M498 likes22k downloads3y agoHugging Face02TamasSimonds /olympiad-proof-problems Olympiad-Proof-Problems Dataset Description This dataset contains mathematical problems and solutions from CSV data. Dataset Summary Total Examples: 39764 Format: Problem-solution pairs Source: olympiad_proof_problems_clean.csv Language: English Domain: Mathematics Data Fields prompt: The mathematical problem statement completion: The complete solution (including working steps) source: Original source identifier id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/TamasSimonds/olympiad-proof-problems.textquestion-answering10K<n<100K1 likes675 downloads1y agoHugging Face03SciCodePile /SciCode-Programming-Problems DATA3: Programming Problems Generation Dataset Dataset Overview DATA3 is a large-scale programming problems generation dataset that contains AI-generated programming problems inspired by real scientific computing code snippets. The dataset consists of 22,532 programming problems, each paired with a comprehensive solution. These problems focus on scientific computing concepts such as numerical algorithms, data analysis, mathematical modeling, and computational methods in… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Programming-Problems.texttext-generation10K<n<100K0 likes526 downloads7mo agoHugging Face04mihailgribov /olympiad_style_integer_math_problems Olympiad Math Corpus Version: v2.1.1 Release date: 2026-05-03 59,486 synthetically generated olympiad-style math problems with verified integer answers and formal computation graphs. Loading from datasets import load_dataset ds = load_dataset("mihailgribov/olympiad_style_integer_math_problems", split="train") lemma_applicability is stored as list[{lemma, status}] rather than a sparse dict (required for Arrow-based consumers). To convert to a dict for local use:… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_problems.documenttext-generation10K<n<100K1 likes390 downloads5mo agoHugging Face05jonaskg /imo-problems-completequestion-answering0 likes389 downloads1y agoHugging Face06azminetoushikwasi /math-story-problems Math Story Problems Dataset Dataset Description This dataset contains mathematical word problems presented in multiple formats, from direct equations to complex story-based scenarios. It is designed for training and evaluating language models on mathematical reasoning tasks. Dataset Structure The dataset is split into three parts: Train: 131,072 samples Validation: 1,024 samples Test: 3,072 samples Features { "eq_qs": "string", # Equation… See the full description on the dataset page: https://huggingface.co/datasets/azminetoushikwasi/math-story-problems.textquestion-answering100K<n<1M1 likes355 downloads1y agoHugging Face07HumynLabs /physics-problems Physics Problems Dataset Dataset Description: This dataset contains a collection of physics problems designed for educational and research purposes. Each problem includes a question, relevant equations, and, where applicable, numerical or symbolic solutions. The dataset covers topics such as mechanics, electromagnetism, thermodynamics, optics, and modern physics. The dataset supports training and evaluation of models in: Natural Language Processing (NLP) for physics problem-solving… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/physics-problems.imagetext-generationn<1K7 likes319 downloads1y agoHugging Face08Jasaxion /MathSmith-Hard-ProblemsMathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy Overview This dataset is a collection of problems generated by the MathSmith-Hard Problem-Synthesizer. Dataset Structure Each record is a JSON object with the following fields: { "problem": "<str>", // The generated math problem "rationale": "<str>" // The ratioanle process of question generation… See the full description on the dataset page: https://huggingface.co/datasets/Jasaxion/MathSmith-Hard-Problems.textquestion-answering100K<n<1M1 likes181 downloads11mo agoHugging Face09hummbl-hf /agent-wicked-problems-40k HUMMBL 40k Multi-Agent Wicked Problems & Coordination Corpus A foundational 40,171-event empirical dataset capturing real-world multi-agent coordination, epistemic problem decomposition, failure mode taxonomies, and strategic intelligence surges generated across the HUMMBL autonomous agent fleet. Dataset Overview The dataset provides structured visibility into how autonomous agents navigate complex, ill-defined ("wicked") problems, coordinate across distributed… See the full description on the dataset page: https://huggingface.co/datasets/hummbl-hf/agent-wicked-problems-40k.tabulartext-classification10K<n<100K1 likes93 downloads13d agoHugging Face10bala5046 /ai-research-problems AI Research Problems 1M Summary This dataset contains 1,000,000 synthetic research-ideation candidates across AI, machine learning, LLMs, RAG, AI agents, MCP, computer vision, robotics, safety, MLOps, and related fields. Important warning These records are synthetic combinations for research ideation. They are not claims that the problems are novel, unsolved, or absent from the literature. A researcher must verify novelty using papers, benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/bala5046/ai-research-problems.texttext-classification1M<n<10M1 likes88 downloads26d agoHugging Face11rodriguescarson /adaption-math-word-problems-solutions NuminaMath Worked Solutions Competition and school maths problems with step-by-step solutions. Rows 11,000 Domain mathematics Format data.parquet, one row per example Licence apache-2.0 Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description prompt The prompt (user turn) as uploaded. completion The target response as uploaded. enhanced_prompt Empty in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-math-word-problems-solutions.texttext-generation10K<n<100K0 likes71 downloads14d agoHugging Face12sdiazlor /logic-problems-reasoning-dataset Dataset Card for my-distiset-a26cd729 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-a26cd729/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/logic-problems-reasoning-dataset.texttext-generationn<1K0 likes64 downloads2y agoHugging Face13SciCode /SciCode-Programming-Problemsgated DATA3: Programming Problems Generation Dataset Dataset Overview DATA3 is a large-scale programming problems generation dataset that contains AI-generated programming problems inspired by real scientific computing code snippets. The dataset consists of 22,532 programming problems, each paired with a comprehensive solution. These problems focus on scientific computing concepts such as numerical algorithms, data analysis, mathematical modeling, and computational… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Programming-Problems.texttext-generation10K<n<100K1 likes57 downloads8mo agoHugging Face14UnfaithRL /aletheia_code_problems Aletheia Code Problems with Misleading Hints Dataset Description This dataset contains multiple-choice code-reasoning problems derived from Aletheia-Bench and augmented with misleading textual hints. The misleading hints are intentionally designed to point to an incorrect answer. The dataset was developed as part of the UnfaithRL project, which studies cue-following and unfaithful reasoning under reinforcement learning with verifiable rewards. Specifically, it was… See the full description on the dataset page: https://huggingface.co/datasets/UnfaithRL/aletheia_code_problems.tabularquestion-answering10K<n<100K0 likes53 downloads3mo agoHugging Face15duxx /orca-math-word-problems-trtextquestion-answering100K<n<1M7 likes49 downloads3y agoHugging Face165CD-AI /Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedtexttext-generation100K<n<1M4 likes42 downloads3y agoHugging Face17agicorp /orca-math-word-problems-200k Dataset Card This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of SLMs in Grade School Math for details about the dataset construction. Dataset Sources Repository: microsoft/orca-math-word-problems-200k Paper: Orca-Math: Unlocking the potential of SLMs in Grade School Math Direct Use This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/agicorp/orca-math-word-problems-200k.textquestion-answering100K<n<1M1 likes38 downloads3y agoHugging Face18shhendu /EternalMath-open-problems EternalMath Open Problems This dataset is the Hugging Face viewer-friendly release of the open companion problem set for EternalMath. It contains 6,049 parameterized math problems across four batches. Batches Batch Rows Language QC status 20260325 988 English QC-passed anon1 1,640 Chinese Unfiltered anon2 1,341 Chinese Unfiltered anon3 2,080 Chinese Unfiltered Files The viewer loads the Parquet shards in data/ as a single train… See the full description on the dataset page: https://huggingface.co/datasets/shhendu/EternalMath-open-problems.textquestion-answering1K<n<10K0 likes38 downloads4mo agoHugging Face19sytelus /taocp_open_problems TAOCP Open Problems Collection of open research problems singled out by Donald Knuth in The Art of Computer Programming series. Its main purpose is to help measure how frontier models understand, investigate, and make verifiable progress on hard but interesting open problems. Contents The dataset contains 9 exercises rated 50, M50, or HM50 in the six TAOCP editions and draft bundles available to this project. Knuth uses these ratings for problems that were not… See the full description on the dataset page: https://huggingface.co/datasets/sytelus/taocp_open_problems.tabularquestion-answeringn<1K1 likes33 downloads1mo agoHugging Face20CJJones /Elementary_Math_Word_Problems_LLM_Training_Short Dataset Card for Math Problem Generator Dataset Summary This dataset contains a subset of 100,000 procedurally generated math word problems, covering various mathematical concepts and difficulty levels. The problems were generated using a Java program that creates contextual word problems with detailed solutions and explanations. 🔗 Full dataset available on Gumroad The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more?… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Elementary_Math_Word_Problems_LLM_Training_Short.textquestion-answering10K<n<100K0 likes31 downloads7mo agoHugging Face21mamed0v /orca-math-word-problems-200k-turkmen Turkmen Orca Math Word Problems 200k Dataset Overview This dataset is a Turkmen translation of the original microsoft/orca-math-word-problems-200k dataset. The Orca Math Word Problems dataset contains 200,000 high-quality math word problems and their solutions. This Turkmen version aims to extend the accessibility of math problem-solving datasets to the Turkmen language community. Dataset Details Original Dataset: microsoft/orca-math-word-problems-200k… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/orca-math-word-problems-200k-turkmen.textquestion-answering100K<n<1M1 likes29 downloads2y agoHugging Face22MichaelAnthony /lemonseed-word-problems lemonseed-word-problems LemonSeed — 13-category arithmetic word problems with plan + scratchpad. Contents wordproblems.jsonl (2500 rows) Format JSON Lines (.jsonl), one example per line. Provenance Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella. textquestion-answering1K<n<10K0 likes24 downloads2mo agoHugging Face23wannaphong /orca-math-word-problems-200k Dataset Card This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of SLMs in Grade School Math for details about the dataset construction. Dataset Sources Repository: microsoft/orca-math-word-problems-200k Paper: Orca-Math: Unlocking the potential of SLMs in Grade School Math Direct Use This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/orca-math-word-problems-200k.textquestion-answering100K<n<1M0 likes18 downloads8mo agoHugging Face24Jasaxion /MathSmith-HC-ProblemsMathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy Overview This dataset is a collection of problems generated by the MathSmith-HC Problem-Synthesizer. Dataset Structure Each record is a JSON object with the following fields: { "problem": "<str>", // The generated math problem "rationale": "<str>" // The ratioanle process of question generation… See the full description on the dataset page: https://huggingface.co/datasets/Jasaxion/MathSmith-HC-Problems.textquestion-answering100K<n<1M0 likes17 downloads11mo agoHugging Face25farabi-lab /Problem-Solving-Insights-Based-on-Kazakh-Traditionsgated 🇰🇿 Problem-Solving Insights Based on Kazakh Traditions 📖 Overview Problem-Solving Insights Based on Kazakh Traditions is a instruction-tuning dataset designed to bridge the gap between ancient Kazakh wisdom and modern societal challenges. 📊 Dataset Statistics General Metrics Metric Count Total Samples 8,005 Total Words (approx.) 4,030,925 Avg. Words per Sample 503 Word Count Distribution (Per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Problem-Solving-Insights-Based-on-Kazakh-Traditions.texttext-generation1K<n<10K0 likes17 downloads2mo agoHugging Face26Shioniro /orca-math-word-problems-200k Dataset Card This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of SLMs in Grade School Math for details about the dataset construction. Dataset Sources Repository: microsoft/orca-math-word-problems-200k Paper: Orca-Math: Unlocking the potential of SLMs in Grade School Math Direct Use This dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/Shioniro/orca-math-word-problems-200k.textquestion-answering100K<n<1M0 likes16 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.