eltociear/numpy-scipy-tasks-v1
numpy-scipy-tasks-v1 Task dataset for a numpy/scipy RL / eval environment, in the shape used by the Prime Intellect Environments Hub. 40 numerical computing tasks across 5 categories. Each task gives the model one or more input arrays and an instruction; the answer is the array left in result, graded with numpy.testing.assert_allclose against a reference. Grading is deterministic — no LLM judge, no external API. Category Tasks Covers array_ops 12 reshape, abs, sort… See the full description on the dataset page: https://huggingface.co/datasets/eltociear/numpy-scipy-tasks-v1.
numpy-scipy-tasks-v1
Task dataset for a numpy/scipy RL / eval environment, in the shape used by the Prime Intellect Environments Hub.
40 numerical computing tasks across 5 categories. Each task gives the model one or more input arrays and an instruction; the answer is the array left in result, graded with numpy.testing.assert_allclose against a reference. Grading is deterministic — no LLM judge, no external API.
Fields
Arrays serialise as {"data": <nested list>, "dtype": <numpy dtype>, "shape": [...]}.
How it was built, and why you can trust the answer key
Tasks are defined as (deterministic input arrays, instruction, reference solution). The expected output is computed by executing the reference solution, never written by hand, so the answer key cannot drift from the instruction.
Every task is then independently verified:
- the reference solution runs and returns a real-valued array
- it is deterministic (executed twice, compared exactly)
- the result is finite — no NaN or inf in the answer key
- the result is non-empty and not identical to its input (identity tasks carry no signal)
- both the result and every input array survive the serialisation round-trip exactly, dtype and shape included
All 40 tasks pass. Builder and verifier: `build_tasks.py`.
A note on grading tolerance
Grading uses assert_allclose (rtol 1e-6, atol 1e-8) rather than exact equality. Answers here come from eigensolvers, matrix exponentials, FFTs and optimisers, whose final bits legitimately differ across BLAS builds, CPU architectures and library versions; bit-exact grading would fail correct solutions for reasons unrelated to the model.
Shape is still compared exactly, and an integer-valued reference requires an integer answer, so the tolerance never launders a wrong-shaped or wrong-typed result.
