keryszhan/agent-code-rl
0
Agent Code RL
Static project report for an evidence-backed code-generation Agent pipeline.
The Space presents the recovered SFT, Monte Carlo, PRM, and GRPO workflow, verified evaluation results, the final gated reward formula, and explicit limits on what the surviving artifacts can support.
Some plans, augmentations, rollouts, and responses in the linked dataset are AI-generated. See the Dataset Card and ATTRIBUTION.md before reuse.
- Model: <https://huggingface.co/keryszhan/qwen2.5-coder-7b-code-plan-sft>
- Dataset: <https://huggingface.co/datasets/keryszhan/agent-code-rl-artifacts>
