Team Ai
Modelpublic

Jasaxion/MathSmith-Hard-Problem-Synthesizer-Qwen3-8B

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
1likes30downloads
Model Card

MathSmith-Hard-Problem-Synthesizer-Qwen3-8B

MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy

![Paper](https://arxiv.org/abs/2508.05592) ![Project Page](https://jasaxion.github.io/MathSmith_ProjectPage/) ![GitHub](https://github.com/Jasaxion/MathSmith) ![License](LICENSE)

Overview

MathSmith is a framework for synthesizing challenging mathematical problems to enhance LLM reasoning. Rather than modifying existing problems, MathSmith constructs new ones from scratch by randomly sampling concept-explanation pairs from PlanetMath.

The model generates <rationale>–<problem> pairs, where:

  • —<rationale>: structured reasoning describing concept integration and difficulty design.
  • —<problem>: a single Olympiad-level mathematical question that admits a verifiable numeric or symbolic answer.

Compared with MathSmith-HC (complexity + consistency reward), MathSmith-Hard removes the consistency term to emphasize maximum reasoning depth and difficulty.


MathSmith Pipeline

The MathSmith framework consists of four main stages:

  1. 1.Concept Collection: Randomly sample concept–explanation pairs from PlanetMath to ensure data independence.
  1. 1.Supervised Fine-tuning (SFT): Train the model on collected concept–explanation pairs to establish foundational understanding.
  1. 1.Reinforcement Learning (RL): Optimize the model using GRPO with rewards based on:
  2. 2.Structural validity
  3. 3.Reasoning complexity
  4. 4.Answer consistency
  1. 1.Weakness-Focused Self-Improvement: Iteratively identify and address model weaknesses by generating targeted problem variants.

Dependence

  • —Transformers 4.52.4
  • —Pytorch 2.7.0+cu126
  • —Datasets 3.6.0
  • —Tokenizers 0.21.1

Citation

If you find this work useful, please cite:

bibtex
@article{zhan2025mathsmith,
  title={MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy},
  author={Zhan, Shaoxiong and Lai, Yanlin and Lu, Ziyu and Lin, Dahua and Yang, Ziqing and Tan, Fei},
  journal={arXiv preprint arXiv:2508.05592},
  year={2025}
}