code-critic
CodeCriticBench
CodeCriticBench: A Holistic Benchmark for Code Critique in LLMs
💥 Introduction
CodeCriticBench is a comprehensive benchmark designed to systematically evaluate the critique capabilities of large language models (LLMs) in both code generation and code-question answering tasks. Beyond focusing on code generation, this benchmark extends to code-related questions, offering multidimensional and fine-grained evaluation criteria to rigorously assess LLMs' reasoning and code… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CodeCriticBench.critic-sft-cwm-only
critic-sft-cwm-only
The CWM-only SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-CWM-only, the CWM-only arm of the corpus ablation in Table 3.
Each record is one critique point: a CWM-32B trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it under the paper's high-level prompt. The main corpus, critic-sft-cwm-qwen, is this set plus 1,915 records from Qwen3-Next trajectories.… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-only.critic-sft-cwm-only-detailed-prompt
critic-sft-cwm-only-detailed-prompt
The detailed-prompt SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-Detailed-Prompt, the comparison arm of the prompt ablation in Table 4.
Each record is one critique point: a CWM-32B trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it. The difference from critic-sft-cwm-only is the teacher prompt. Here the teacher used the detailed prompt… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-only-detailed-prompt.critic-sft-cwm-qwen
critic-sft-cwm-qwen
The main SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains both Qwen3-8B-Critic-SFT and Qwen3-4B-Critic-SFT.
Each record is one critique point: a coding agent's trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it. The teacher was prompted with the paper's high-level prompt, which asks for error detection and one or two sentences of guidance and forbids code and commands in the… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-qwen.code-feedback-deepseekv2-criticcode-feedback-10k-deepseekv2-critic
