Team Ai
Modelpublic

Rishubi/CodeRM-NT

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes17downloads
Model Card

CodeRM-NT

Paper | Github

Providing accurate reward signals for code generated by LLMs is a significant challenge in applying reinforcement learning (RL) to code generation. Existing methods rely on unit tests, which are expensive to curate and unreliable when automatically synthesized.

CodeRM-NT is a code reward model with no reliance on unit tests. Instead of executing test cases, it learns to estimate the functional correctness of generated Python code from rewards that are collected via Monte Carlo Tree Search (MCTS) guided by LLM-as-a-Judge.

Usage

python
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model = AutoModelForSequenceClassification.from_pretrained(
    "Rishubi/CodeRM-NT",
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("Rishubi/CodeRM-NT")

question = "Write a Python function `add(a, b)` that returns the sum of two integers."
response = "def add(a, b):\n    return a + b"

messages = [
    {"role": "user", "content": question},
    {"role": "assistant", "content": response},
]
input_ids = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
with torch.no_grad():
    reward = model(input_ids).logits.squeeze().float().item()
print(reward)  # higher is better

Results

Key Results

Training with CodeRM-NT consistently outperforms synthetic unit tests and other reward models across multiple code generation benchmarks:

ModelRewardHumanEvalHumanEval+MBPPMBPP+LCB-v5BCB-I-HardAvg.
Qwen2.5-Coder-1.5BUnit Tests73.267.770.961.15.16.147.4
CodeRM-NT75.069.572.060.85.57.448.4
Qwen2.5-Coder-3BUnit Tests86.682.374.964.613.015.556.2
CodeRM-NT88.482.375.966.113.614.256.8
Qwen2.5-Coder-7BUnit Tests90.987.885.473.017.318.262.1
CodeRM-NT90.286.086.874.617.518.262.2
GLM-4-9B-0414Unit Tests84.179.981.069.015.415.557.5
CodeRM-NT87.281.779.967.215.318.258.3
Qwen3-4B-ThinkingUnit Tests97.692.791.075.150.325.772.1
CodeRM-NT97.694.592.677.252.122.372.7

Citation

If you find our work helpful, please kindly cite our paper:

@inproceedings{xia-etal-2026-coderm,
    title = "{C}ode{RM}-{NT}: Reward Model for Code {RL} without Unit Tests",
    author = "Xia, Xiao  and
      Zhang, Dan  and
      Sun, Tianrui",
    editor = "Liakata, Maria  and
      Moreira, Viviane P.  and
      Zhang, Jiajun  and
      Jurgens, David",
    booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
    month = jul,
    year = "2026",
    address = "San Diego, California, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.findings-acl.2150/",
    pages = "43316--43333",
    ISBN = "979-8-89176-395-1"
}