datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orca-math-word-problems-100k-en-zh-mix100k English and Chinese mixed version of microsoft/orca-math-word-problems-200k
adaption-math-word-problems-solutions
NuminaMath Worked Solutions
Competition and school maths problems with step-by-step solutions.
Rows
11,000
Domain
mathematics
Format
data.parquet, one row per example
Licence
apache-2.0
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
prompt
The prompt (user turn) as uploaded.
completion
The target response as uploaded.
enhanced_prompt
Empty in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-math-word-problems-solutions.Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedElementary_Math_Word_Problems_LLM_Training_Short
Dataset Card for Math Problem Generator
Dataset Summary
This dataset contains a subset of 100,000 procedurally generated math word problems, covering various mathematical concepts and difficulty levels. The problems were generated using a Java program that creates contextual word problems with detailed solutions and explanations.
🔗 Full dataset available on Gumroad
The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more?… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Elementary_Math_Word_Problems_LLM_Training_Short.word_problems
Word Problems
Public synthetic word-problem dataset generated with a lightweight CPU-first pipeline.
Snapshot
Rows: 21800
Format: JSONL
File: word_problems.jsonl
Generation date: 2026-03-17
Approximate size: 109.4 MB
Fields
Each row contains a word problem, a worked answer, metadata, and quality/provenance information from the generator.
