Team Ai
Apppublic

AryaWu/agentic-leetcode-trial-viewer

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes
App README

Agentic LeetCode trial viewer

Browse full agent trajectories (turn-by-turn actions, tool outputs, final code, pass/fail) from two training runs on the LeetCode-Python harbor benchmark. Pick a run and a problem in the sidebar to open a trial.

Runs included

  • —`tvgpo_v_no_kl_10turn` — TVGPO-V, no KL penalty, 10-turn ablation. W&B run `qp0tkrf3`.
  • —`grpo_kl_klfix_10turn` — plain GRPO with KL=0.01 ("klfix"), 10-turn. W&B run `3nz4y90h`.

Selection

Each run samples 8 rollouts per problem. For each run, the 100 problems with the highest reward variance across their 8-rollout group (i.e. most disagreement between passing and failing attempts, the most informative cases for GRPO-style advantage) were kept; everything else was dropped to keep this Space small. That's ~795–800 trials per run (a small number of problems had slightly fewer than 8 usable trials).

Trajectories are converted from harbor's own trial format (viewer/convert_harbor_trials.py in the training repo) and pre-rendered to static JSON — this Space is a static export of that repo's FastAPI trajectory viewer (viewer/view_trajectory.py), with every /api/* response baked to a file instead of computed live.