Team Ai
Datasetpublic

AdithyaSK/RetroEnv-RL

RetroEnv RL tasks Tasks for RetroEnv, a multi-turn tool-use environment for retrosynthesis. Given a target molecule, the agent plans a synthesis back to purchasable building blocks: it searches the stock and training precedents, checks proposed disconnections, and submits route trees with emit_routes. A deterministic verifier scores the trees against the route reported in the target's patent, with a reward in [0, 1] built from nine components. Code, the OpenEnv server… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/RetroEnv-RL.

sourceHugging Facecc-by-4.0updated 7d agoView on Hugging Face
0likes829downloads
Dataset Card

RetroEnv RL tasks

Tasks for RetroEnv, a multi-turn tool-use environment for retrosynthesis. Given a target molecule, the agent plans a synthesis back to purchasable building blocks: it searches the stock and training precedents, checks proposed disconnections, and submits route trees with emit_routes. A deterministic verifier scores the trees against the route reported in the target's patent, with a reward in [0, 1] built from nine components.

Code, the OpenEnv server, evaluation and training recipes: FineEnvs 09-retroenv (PR #30). SFT trajectories for the v3 train split: AdithyaSK/RetroEnv-SFT.

What is here

PathContents
tasks-public/v3 tasks as the policy sees them: target, step budget, routes asked for, public difficulty
tasks-private/the same tasks with their reference routes, the answer key the server grades against
stocks/paroutes-v2-n1.smithe 13,432 purchasable molecules; stock_retrieve is the policy's only access to it
difficulty.jsonlper-task heuristic tier and features (first reaction family, nearest train target, leaf rarity)
manifest.json, audit.json, checksums.jsonhow the splits were designed, the leakage audit, and SHA-256 of every file the server reads
normalized-routes.jsonl, baselines/report.jsonevery reference route, and oracle, one-route and empty-submission scores
retroeval-v2/the earlier 1,000-task benchmark (2- and 3-step routes), same layout, with the model board in results/
runs/v2-eval/the six models' v2 board episodes with full transcripts
media/retroenv-explainer.mp4an 83-second explainer of one episode

v3 splits

SplitTasks2 / 3 / 4 / 5 stepsTwo-routeHeavy atoms ≤20 / 21–30 / 31–40 / 41+Easy / medium / hard
train27,48913,376 / 8,878 / 3,724 / 1,5114711,134 / 11,386 / 4,705 / 26417,568 / 7,132 / 2,789
dev15042 / 42 / 36 / 301542 / 51 / 45 / 1260 / 56 / 34
eval25070 / 70 / 60 / 502570 / 85 / 75 / 2089 / 106 / 55
stress15042 / 42 / 36 / 301535 / 51 / 47 / 1758 / 56 / 36

Use train for RL, dev for checkpoint selection, eval for the board, and keep stress as a sealed second test. Eval covers all ten first-step reaction families (13 to 48 tasks each).

How the splits were made:

  • —The held-out splits were drawn first, against quotas on route length, first-step reaction family, target size and two-route tasks.
  • —No held-out task shares a target, scaffold, intermediate, reaction, patent or Morgan near-duplicate (Tanimoto ≥ 0.90) with train or with another held-out task.
  • —One-ring scaffolds, and the 33 multi-ring scaffolds found in at least 100 pool tasks (biphenyl, indole and others), group by exact structure. manifest.json lists them.
  • —A scripted solver that knows each route passes all 550 held-out tasks within the 16-turn, 32-call budget.

Use

Serve it with the OpenEnv server. This is the Docker image's default:

bash
RETROENV_TASKS_REPO=AdithyaSK/RetroEnv-RL python envs/retro_route/openenv/prepare.py
# RETROENV_TASKS_SUBDIR=retroeval-v2 serves v2 instead

Load the rows:

python
from datasets import load_dataset

tasks = load_dataset("AdithyaSK/RetroEnv-RL", "v3")   # what the policy sees
answers = load_dataset("AdithyaSK/RetroEnv-RL", "v3-with-references")

Get a full local copy for the evaluation scripts and the explorer:

bash
hf download AdithyaSK/RetroEnv-RL --repo-type dataset --local-dir benchmark/retroeval-v3 \
  --exclude "retroeval-v2/*" "runs/*" "media/*"

v2 model board

150 eval tasks, one attempt each, full toolset:

ModelPass@1 (95% CI)Exact routeRewardRefusedCost
Claude Opus 5.50.560 [0.48, 0.64]0.5930.75016%$13.94
Claude Sonnet 5.50.300 [0.23, 0.38]0.3400.6260%$9.76
GPT-5.6 Sol0.280 [0.21, 0.36]0.3200.7050%$23.25
DeepSeek V4.1 Flash0.107 [0.07, 0.17]0.1270.4360%$5.03
GPT-5.6 Luna0.073 [0.04, 0.13]0.1000.5650%$2.29
Qwen3.8 27B0.013 [0.00, 0.05]0.0070.1380%$8.27

A refusal ends the episode and scores the 0.05 floor. No models have been run on v3 yet.

Notes

  • —The answers to every split are public here, because PaRoutes is public and anyone can rebuild them. Say so when reporting results on a model that may have been trained on USPTO-derived data.
  • —validate_disconnection and reaction_class_lookup answer from the hidden reference. The server's unaided toolset removes them.

License and attribution

CC-BY-4.0. Routes and stock come from PaRoutes v2 (Genheden & Bjerrum, Digital Discovery 2022; Zenodo 7341155, CC-BY-4.0), which extracts reactions from USPTO patents.