Team Ai
Apppublic

keryszhan/agent-code-rl

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes
App README

Agent Code RL

Static project report for an evidence-backed code-generation Agent pipeline.

The Space presents the recovered SFT, Monte Carlo, PRM, and GRPO workflow, verified evaluation results, the final gated reward formula, and explicit limits on what the surviving artifacts can support.

Some plans, augmentations, rollouts, and responses in the linked dataset are AI-generated. See the Dataset Card and ATTRIBUTION.md before reuse.

  • —Model: <https://huggingface.co/keryszhan/qwen2.5-coder-7b-code-plan-sft>
  • —Dataset: <https://huggingface.co/datasets/keryszhan/agent-code-rl-artifacts>