Team Ai
Datasetpublic

eltociear/numpy-scipy-tasks-v1

numpy-scipy-tasks-v1 Task dataset for a numpy/scipy RL / eval environment, in the shape used by the Prime Intellect Environments Hub. 40 numerical computing tasks across 5 categories. Each task gives the model one or more input arrays and an instruction; the answer is the array left in result, graded with numpy.testing.assert_allclose against a reference. Grading is deterministic — no LLM judge, no external API. Category Tasks Covers array_ops 12 reshape, abs, sort… See the full description on the dataset page: https://huggingface.co/datasets/eltociear/numpy-scipy-tasks-v1.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes72downloads
Dataset Card

numpy-scipy-tasks-v1

Task dataset for a numpy/scipy RL / eval environment, in the shape used by the Prime Intellect Environments Hub.

40 numerical computing tasks across 5 categories. Each task gives the model one or more input arrays and an instruction; the answer is the array left in result, graded with numpy.testing.assert_allclose against a reference. Grading is deterministic — no LLM judge, no external API.

CategoryTasksCovers
array_ops12reshape, abs, sort, cumsum, clip, boolean masking, argsort, transpose, axis reductions, standardisation, diff, rounding
linalg10matmul, inverse, symmetric eigenvalues, Cholesky, solve, SVD, QR, matrix exponential, pseudo-inverse, determinant
statistics8mean/median/std, percentiles, z-scores, correlation, covariance, chi-square test, linear regression, rank data
signal6real FFT magnitude, FFT frequencies, convolution, DCT-II, linear detrend, peak finding
optimization4polyfit, lstsq, BFGS minimisation, Brent root finding

Fields

FieldDescription
task_idstable id, e.g. npsp-017
categoryone of the five above
promptthe natural-language instruction shown to the model
input_dataJSON object mapping array name (a, m, s, g, sig, c) to a serialised array
expected_outputthe serialised reference result

Arrays serialise as {"data": <nested list>, "dtype": <numpy dtype>, "shape": [...]}.

How it was built, and why you can trust the answer key

Tasks are defined as (deterministic input arrays, instruction, reference solution). The expected output is computed by executing the reference solution, never written by hand, so the answer key cannot drift from the instruction.

Every task is then independently verified:

  • —the reference solution runs and returns a real-valued array
  • —it is deterministic (executed twice, compared exactly)
  • —the result is finite — no NaN or inf in the answer key
  • —the result is non-empty and not identical to its input (identity tasks carry no signal)
  • —both the result and every input array survive the serialisation round-trip exactly, dtype and shape included

All 40 tasks pass. Builder and verifier: `build_tasks.py`.

A note on grading tolerance

Grading uses assert_allclose (rtol 1e-6, atol 1e-8) rather than exact equality. Answers here come from eigensolvers, matrix exponentials, FFTs and optimisers, whose final bits legitimately differ across BLAS builds, CPU architectures and library versions; bit-exact grading would fail correct solutions for reasons unrelated to the model.

Shape is still compared exactly, and an integer-valued reference requires an integer answer, so the tolerance never launders a wrong-shaped or wrong-typed result.