AdithyaSK/RetroEnv-RL
RetroEnv RL tasks Tasks for RetroEnv, a multi-turn tool-use environment for retrosynthesis. Given a target molecule, the agent plans a synthesis back to purchasable building blocks: it searches the stock and training precedents, checks proposed disconnections, and submits route trees with emit_routes. A deterministic verifier scores the trees against the route reported in the target's patent, with a reward in [0, 1] built from nine components. Code, the OpenEnv server… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/RetroEnv-RL.
RetroEnv RL tasks
Tasks for RetroEnv, a multi-turn tool-use environment for retrosynthesis. Given a target molecule, the agent plans a synthesis back to purchasable building blocks: it searches the stock and training precedents, checks proposed disconnections, and submits route trees with emit_routes. A deterministic verifier scores the trees against the route reported in the target's patent, with a reward in [0, 1] built from nine components.
Code, the OpenEnv server, evaluation and training recipes: FineEnvs 09-retroenv (PR #30). SFT trajectories for the v3 train split: AdithyaSK/RetroEnv-SFT.
What is here
v3 splits
Use train for RL, dev for checkpoint selection, eval for the board, and keep stress as a sealed second test. Eval covers all ten first-step reaction families (13 to 48 tasks each).
How the splits were made:
- The held-out splits were drawn first, against quotas on route length, first-step reaction family, target size and two-route tasks.
- No held-out task shares a target, scaffold, intermediate, reaction, patent or Morgan near-duplicate (Tanimoto ≥ 0.90) with train or with another held-out task.
- One-ring scaffolds, and the 33 multi-ring scaffolds found in at least 100 pool tasks (biphenyl, indole and others), group by exact structure.
manifest.jsonlists them. - A scripted solver that knows each route passes all 550 held-out tasks within the 16-turn, 32-call budget.
Use
Serve it with the OpenEnv server. This is the Docker image's default:
RETROENV_TASKS_REPO=AdithyaSK/RetroEnv-RL python envs/retro_route/openenv/prepare.py
# RETROENV_TASKS_SUBDIR=retroeval-v2 serves v2 insteadLoad the rows:
from datasets import load_dataset
tasks = load_dataset("AdithyaSK/RetroEnv-RL", "v3") # what the policy sees
answers = load_dataset("AdithyaSK/RetroEnv-RL", "v3-with-references")Get a full local copy for the evaluation scripts and the explorer:
hf download AdithyaSK/RetroEnv-RL --repo-type dataset --local-dir benchmark/retroeval-v3 \
--exclude "retroeval-v2/*" "runs/*" "media/*"v2 model board
150 eval tasks, one attempt each, full toolset:
A refusal ends the episode and scores the 0.05 floor. No models have been run on v3 yet.
Notes
- The answers to every split are public here, because PaRoutes is public and anyone can rebuild them. Say so when reporting results on a model that may have been trained on USPTO-derived data.
validate_disconnectionandreaction_class_lookupanswer from the hidden reference. The server'sunaidedtoolset removes them.
License and attribution
CC-BY-4.0. Routes and stock come from PaRoutes v2 (Genheden & Bjerrum, Digital Discovery 2022; Zenodo 7341155, CC-BY-4.0), which extracts reactions from USPTO patents.
