Team Ai
Datasetpublic

saidutta69/unsolved-math-clean

๐Ÿง  Unsolved Math โ€” Clean 8,626 curated open research problems in mathematics and CS โ€” including 122 Millennium Prize Problems โ€” deduplicated, schema-flattened, and packaged as proper parquet configs with an eval-only benchmark view. A reasoning frontier dataset: every problem here is actually unsolved or partially solved โ€” ideal for honest capability probing instead of contaminated benchmarks. Clean derivative of ulamai/UnsolvedMath (8,785 problems). License unchanged:โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/unsolved-math-clean.

sourceHugging Facecc-by-4.0updated 28d agoView on Hugging Face
0likes64downloads
Dataset Card

๐Ÿง  Unsolved Math โ€” Clean

<div align="center"> <img src="https://photu.kashyalabanavli.site/racer-is-op.png" alt="RACER IS OP" width="100%"> </div>

<br>

8,626 curated open research problems in mathematics and CS โ€” including 122 Millennium Prize Problems โ€” deduplicated, schema-flattened, and packaged as proper parquet configs with an eval-only benchmark view.

A reasoning frontier dataset: every problem here is actually unsolved or partially solved โ€” ideal for honest capability probing instead of contaminated benchmarks.

Clean derivative of ulamai/UnsolvedMath (8,785 problems). License unchanged: CC-BY-4.0.

๐Ÿงน Quality Pipeline

StepRemovedReason
Raw input8,785Nested JSON
Duplicate titles153Same problem listed twice
Empty statements6No problem body
Final clean8,626Flattened schema, deterministic splits

๐Ÿ“Š Composition

By status:

StatusProblems
Open4,354
Partially solved3,580
Solved692

By difficulty:

LevelProblems
L1: Tractable908
L2: Intermediate516
L3: Advanced6,143
L4: Expert937
L5: Millennium Prize122

๐Ÿ“ฆ Configs

ConfigRowsDescription
problems (default)8,626Full metadata: statement, background, status, difficulty, category, provenance
open_problems_eval4,354Open problems only โ€” benchmark/eval view
research_prompts8,626Instruction-formatted: "analyze this open problem, discuss partial results and obstacles"

๐ŸŽฏ Usage

python
from datasets import load_dataset

ds = load_dataset("saidutta69/unsolved-math-clean", "problems", split="train")

# Honest eval: only truly open problems
ev = load_dataset("saidutta69/unsolved-math-clean", "open_problems_eval")

# Research-style SFT prompts
rp = load_dataset("saidutta69/unsolved-math-clean", "research_prompts")
print(rp["train"][0]["instruction"][:300])

โš ๏ธ Notes

  • โ€”These are unsolved problems โ€” there are no reference answers. Use for capability probing, calibration of hedging behavior, and research discussion; not for accuracy scoring.
  • โ€”LaTeX notation preserved verbatim.
  • โ€”Millennium-level problems are intentionally near-impossible; expect models to fail honestly.

๐Ÿ“œ Citation

bibtex
@misc{unsolved-math-clean,
  author = {Sai Dutta Abhishek Dash (clean); original by ulamai},
  title = {Unsolved Math โ€” Clean},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/saidutta69/unsolved-math-clean}}
}