envsetup-rl-dl4c/shellcheck-reward-checkpoint
022
Reinforcement Fine-Tuned Qwen3-8B
This model is fine-tuned on the Qwen3-8B model using reinforcement learning for environment setup task. It's trained on zeroshot prompts from the EnvBench dataset.
This model is fine-tuned on the Qwen3-8B model using reinforcement learning for environment setup task. It's trained on zeroshot prompts from the EnvBench dataset.