AryaWu/agentic-leetcode-trial-viewer
Agentic LeetCode trial viewer
Browse full agent trajectories (turn-by-turn actions, tool outputs, final code, pass/fail) from two training runs on the LeetCode-Python harbor benchmark. Pick a run and a problem in the sidebar to open a trial.
Runs included
- `tvgpo_v_no_kl_10turn` — TVGPO-V, no KL penalty, 10-turn ablation. W&B run `qp0tkrf3`.
- `grpo_kl_klfix_10turn` — plain GRPO with KL=0.01 ("klfix"), 10-turn. W&B run `3nz4y90h`.
Selection
Each run samples 8 rollouts per problem. For each run, the 100 problems with the highest reward variance across their 8-rollout group (i.e. most disagreement between passing and failing attempts, the most informative cases for GRPO-style advantage) were kept; everything else was dropped to keep this Space small. That's ~795–800 trials per run (a small number of problems had slightly fewer than 8 usable trials).
Trajectories are converted from harbor's own trial format (viewer/convert_harbor_trials.py in the training repo) and pre-rendered to static JSON — this Space is a static export of that repo's FastAPI trajectory viewer (viewer/view_trajectory.py), with every /api/* response baked to a file instead of computed live.
