datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RetroEnv-RL
RetroEnv RL tasks
Tasks for RetroEnv, a multi-turn tool-use environment for retrosynthesis. Given a target
molecule, the agent plans a synthesis back to purchasable building blocks: it searches the
stock and training precedents, checks proposed disconnections, and submits route trees with
emit_routes. A deterministic verifier scores the trees against the route reported in the
target's patent, with a reward in [0, 1] built from nine components.
Code, the OpenEnv server… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/RetroEnv-RL.RetroEnv
RetroEnv-RL
Multistep retrosynthesis tasks for agentic RL. Each task gives a target molecule, a depth
budget (longest linear sequence) and optionally a constraint; the agent plans routes with
tools and submits synthesis trees whose every leaf must be in the frozen stock. A
deterministic verifier judges each step against a frozen reaction library (known
reactions and frequent rdchiral retro-templates), never against a hidden answer; known
routes only add a similarity bonus.… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/RetroEnv.RetroKV-Fig-Datageometry-dash-retro-levelsFork of https://huggingface.co/datasets/yusp48/geometry-dash-levels.
Contains only retro levels with id < 11000000.
Use my gdparse library: pip install gdparse
repro-abc-bench-an-agentic-bio-capabilities-benchmark-for-biosecurity-traces
Agent traces
Agent sessions published from a Trackio Logbook.
2026-08-28-post-action-retrospection-716-coherent
Post-action retrospection 716 -- coherent rewrite (arm 1 of the PAR coherence experiment)
field
value
experiment
The exact 716 five-turn PAR rows that trained LASR-Callum/2026-08-26-qwen36-lora-table2-9284-post-action-retrospection-716-rank-64-dynbatch (mixture 2026-08-26-table2-9284-par716-train @ 42c8a74), with ONLY the trained turn (turn 4: private reasoning + reply) rewritten by Sonnet 5 so the reasoning ENDS on a first-person decision (what it won't do, per… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-28-post-action-retrospection-716-coherent.RetroDFM-R-inferenceRetroEnv-SFT
RetroEnv SFT trajectories
Multi-turn tool-calling episodes for RetroEnv,
a retrosynthesis environment. Each row plans a route for one v3 train task of
AdithyaSK/RetroEnv-RL back to
purchasable molecules and ends with emit_routes. Every row is an episode that the
environment's verifier passed, through the same OpenEnv server and agent loop used for
evaluation.
chemist (default)
plain (baseline)
Rows (train / validation)
31,897 / 323
27,222 / 267
Train tasks covered… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/RetroEnv-SFT.retro-weave-agent-editor-repair-diffs-v0.1
RetroInstruct Weave Agent Editor Repair Diffs
This component of RetroInstruct trains weave-agent to use the WeaveEditor to fix synthetic corruptions in the vein of
the Easy Prose Repair Diffs component.
Each row in the dataset provides the pieces you need to make a synthetic episode
demonstrating the agent:
Singling out one of three files as corrupted and in need of repair
Writing out a patch to the file as either a series of WeaveEditor edit() commands or a unidiff
Observing the… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-weave-agent-editor-repair-diffs-v0.1.retro-ascii-art-v1
RetroInstruct ASCII Art
This component of RetroInstruct trains language models to draw ASCII art. Many
advanced language models such as Microsoft Prometheus (Bing) and Claude 3 Opus
can draw impressive ASCII diagrams. Mistral-large on the other hand can't. Since
there should in principle be plenty of ASCII art in Common Crawl I suspect this
is caused by either Mistral's filters removing ASCII art from the pretraining
or instruction tuning data that doesn't reinforce the ability to… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-ascii-art-v1.InsightTokTokenTraining
InsightTok Token Training
This is a list of image prompts and token pairs for training.
2026-08-26-sonnet45-post-action-retrospection-natural-turn-design
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
date_generated
20260826_152715
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ c2fdee460e71fa28e9902edf1cc662db0d19cad8
models
per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-sonnet45-post-action-retrospection-natural-turn-design.retro-weave-eval-rubrics-v0.1
RetroInstruct Weave Evaluator Rubrics v0.1
This component of RetroInstruct trains the ability to break subjective weave rubric
items like "Is this good writing?" into parts which can be more objectively answered.
It is closery related to the word parts component
which is meant to train a similar skill. By making these rubrics the model gains
the ability to make in-context text classifiers and discriminators. These can be
used to drive a MCTS, filter language model
outputs to… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-weave-eval-rubrics-v0.1.adaptive-retro-gpt-1b-corpus
Adaptive-RETRO-GPT-1B Pretraining Corpus
Cleaned causal language modeling corpus for the Adaptive-RETRO-GPT-1B run.
Source: HuggingFaceFW/fineweb-edu / sample-10BT
Train rows: 80000
Validation rows: 4000
Format: JSONL with text and source
retroinstruct-mix-v0.2
RetroInstruct Mix v0.2
This is the first release of the RetroInstruct synthetic instruction dataset.
It is a mixture of 7 synthetic subsets:
RetroInstruct Weave Evaluator Questions: JDP - Answer questions about synthetic short form writing in the style of John David Pressman.
RetroInstruct Analogical Translations - Infer the generative process of bad faith reasoning by executing a bad faith process to generate arguments and reversing it.
RetroInstruct Part Lists For Dictionary… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retroinstruct-mix-v0.2.retroinstruct-agent-mix-v0.4DreadPoor__Mercury_In_Retrograde-8b-Model-Stock-details
Dataset Card for Evaluation run of DreadPoor/Mercury_In_Retrograde-8b-Model-Stock
Dataset automatically created during the evaluation run of model DreadPoor/Mercury_In_Retrograde-8b-Model-Stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Mercury_In_Retrograde-8b-Model-Stock-details.retro-easy-prose-repair-diffs-v0.1
RetroInstruct Easy Prose Repair Diffs
This component of RetroInstruct trains language models to repair prose by outputting
a diff that patches its flaws. The dataset is made through backtranslation by
running a synthetic corruption pass over prose. I use mostly syntactic
corruptions made with traditional programs, which makes them 'easy' compared to
more subtle semantic problems that could be introduced by a neural network. The
text I backtranslate from was generated by Mixtral… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-easy-prose-repair-diffs-v0.1.retro-weave-eval-jdp-v0.1
RetroInstruct Weave Evaluator Questions: JDP
This component of RetroInstruct trains the ability to answer yes-no questions such as "Does the current scene take place at a wedding party?". The logits from such questions can be taken to make in-context text classifiers and discriminators. These can be
used to drive a MCTS, filter language model
outputs to heighten the probability they satisfy certain properties, and validate
abstract properties of inputs. This set of questions is made… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-weave-eval-jdp-v0.1.RetroReasoner-dataretroinstruct-agent-mix-v0.1retroinstruct-agent-mix-v0.3retroinstruct-agent-mix-v0.5retroinstruct-agent-mix-v0.2gender_stereoset_rephrasedretro-weave-eval-analogical-translations-v0.1
RetroInstruct Analogical Translations
This component of RetroInstruct trains the weave evaluator
on analogical translations, a repeatable reasoning process for generating arguments
created for this dataset. I found that
trying to base good vs. poor arguments on individual named fallacies was both
tedious and failing to consistently produce flawed arguments. e.g. Asking
Mistral-large to generate arguments qualifying as an "appeal to possibility" would
generate many valid arguments… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-weave-eval-analogical-translations-v0.1.repro-retrofit-artifactsretroinstruct-mix-v0.1
