Team Ai
Datasetpublic

TheJackBright/verisci-verified-science-math-code

VeriSci Verified Science Math Code Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage. Summary VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes65downloads
Dataset Card

VeriSci Verified Science Math Code

Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage.

Summary

VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python code-generation tasks, and explicit abstention when required variables are missing.

Every row is generated by a deterministic Python program and checked by a verifier. This makes the dataset suitable for supervised fine-tuning and for objective base-vs-adapted evaluation.

Fields

  • —id: stable row id.
  • —split: train, validation, or test.
  • —prompt: instruction for the model.
  • —completion: target response with reasoning and final JSON.
  • —reasoning_trace: intermediate reasoning steps.
  • —final_answer: machine-readable target answer.
  • —verifier: verifier type and parameters.
  • —variables: source variables used to generate the row.
  • —dedupe_signature: canonical hash used for de-duplication and split assignment.
  • —task_family: task family.
  • —difficulty: easy, medium, or hard.
  • —source: generation provenance.
  • —license: row license.

Task Families

  • —unit_checked_mechanics
  • —thermodynamics
  • —electric_circuits
  • —exponential_decay
  • —chemistry_stoichiometry
  • —chemistry_solutions
  • —numerical_ode
  • —numerical_integration
  • —finite_difference_pde
  • —unit_conversion
  • —vector_reasoning
  • —linear_modeling
  • —dimensional_analysis
  • —python_code_generation
  • —scientific_abstention
  • —code_abstention

Reproduction

bash
git clone <repository-url>
cd verisci-autoscientist
PYTHONPATH=src python3 -m verisci.generate \
  --rows 8000 \
  --out data/generated/verisci_8k.jsonl \
  --csv data/generated/verisci_8k.csv \
  --summary data/generated/verisci_8k_summary.json
PYTHONPATH=src python3 -m verisci.evaluate --data data/generated/verisci_8k.jsonl

Release Integrity

The current 8k public release is split-safe:

MetricValue
Rows8,000
Unique prompts8,000
Duplicate prompts0
Train/validation/test prompt leakage rows0
Dedupe signatures8,000
Gold verifier accuracy100%

Adaptive Data And AutoScientist

This dataset has been run through a low-credit Adaption pilot and is prepared for AutoScientist training.

Current platform evidence:

MetricValue
Pilot dataset ID09749657-7dde-4e20-8988-1d9ee53e9132
50-row Adaptive Data estimate1 credit, 11 minutes
Enhanced-completion auditFailed: 0/42 rows preserved Final: {...}
Source-column auditPassed: 42/42 rows preserved Final: {...}
12k source preflight dataset ID13b1c92a-a96b-4ce7-810b-4368ad6aa234
12k source-column auditPassed: 12,000/12,000 rows preserved Final: {...}
AutoScientist recommendationgoogle/gemma-3-4b-it, 1-epoch LoRA
Diagnostic 8k Llama run43b5486d-0bfc-4e9f-869e-a9892a679386; best win rate 49.48%; not final
Clean 8k candidate run7ea2c71b-94d7-472d-8288-795ab5e0a2c3 on mistralai/Mistral-7B-Instruct-v0.2
Clean 8k candidate dataset3103c6ac-7d61-4271-af62-41cb023de85e
Clean 8k latest metricSucceeded, 5/5 iterations, best win rate 52.45%, checkpoint packaged

Fill after final AutoScientist run:

MetricValue
Adaptive Data grade beforePending
Adaptive Data grade afterPending
Adaptive Data quality improvementPending
AutoScientist best win rate52.45%; below the 75% publish gate
Domain augmentation rows8,000 source rows in clean candidate
General augmentation rows0 in clean candidate

Limitations

VeriSci is synthetic and deliberately narrow. It is designed to test exact scientific and code reasoning patterns, not to replace expert review for high-stakes engineering, laboratory, medical, financial, or safety decisions.

License

Apache-2.0.