Team Ai
Modelpublic

dipta007/atomicity-grounded-judge-balanced

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes17downloads
Model Card

DecomposeRL Tiny-Judge: Atomicity (grounded) Judge

<p align="center"> <a href="https://arxiv.org/abs/2605.27858v1"> <img src="https://img.shields.io/badge/%F0%9F%93%84_Paper-arXiv-b12a00?style=for-the-badge&labelColor=ffb300" alt="Paper"> </a> </p>

![Paper](https://arxiv.org/abs/2605.27858v1) ![Project Page](https://dipta007.github.io/DecomposeRL/) ![Dataset](https://huggingface.co/datasets/dipta007/decomposeRL-tiny-judge) ![Collection](https://huggingface.co/collections/dipta007/decomposerl) ![GitHub](https://github.com/dipta007/DecomposeRL)

A ModernBERT-large classifier that scores whether a generated sub-question is grounded in claim-specific entities — one of the five binary checks that make up the atomicity sub-signal of DecomposeRL's joint multiplicative quality reward.

It is part of the DecomposeRL tiny-judge stack — eight task-specific LoRA classifier heads on a shared ModernBERT-large backbone that distill a Qwen3-32B LLM judge into small, fast reward models. Swapping the 32B judge for this ~400M-parameter stack cuts GRPO judge compute by ~80% (240 → 48 GPU-hours) while retaining ~99% of in-domain accuracy.

Model Overview

PropertyValue
Model TypeModernBertForSequenceClassification (sequence classification)
Base Modelanswerdotai/ModernBERT-large (~400M params)
TrainingLoRA (r=64, α=128), merged into the base before release
Labels2-way: no / yes
Distilled fromQwen/Qwen3-32B judge labels
Dataset / config`dipta007/decomposeRL-tiny-judge` · atomicity_grounded
Train splittrain_balanced (class-balanced); selected on macro-F1
LanguageEnglish

What it judges

This head is one of five binary atomicity checks (is_question, single_focus, no_conjunctions, verifiable, grounded). At reward time the five yes/no predictions are averaged into the per-question atomicity score R_atom, which is then multiplied with the answerability (R_ans) and answer-correctness (R_corr) sub-signals to form the joint multiplicative quality reward (Eq. 7 in the paper).

Input format

Claim + candidate sub-question:

Claim: {claim}
Question: {question}

Label space

LabelNameMeaning
0nothe question references entities not present in the claim, or is generic
1yesthe question is anchored to entities that actually appear in the claim

Quickstart

python
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

repo = "dipta007/atomicity-grounded-judge-balanced"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSequenceClassification.from_pretrained(repo).eval()

text = (
    'Claim: Louis William Tomlinson is an English singer and songwriter who released "Back to You" with an American singer-songwriter who released her debut extended play on May 12, 2015, by Warner Bros. Records.\\n'
    'Question: Which single did Louis Tomlinson release with the American singer-songwriter?'
)

inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=8192)
with torch.no_grad():
    logits = model(**inputs).logits
pred = int(logits.argmax(-1))
print(pred, model.config.id2label[pred])
# expected: 1 -> yes

Training Data

Trained on the atomicity_grounded config of `dipta007/decomposeRL-tiny-judge`, whose labels are distilled from Qwen3-32B judge calls made during DecomposeRL reward computation. The model is fine-tuned with LoRA on the class-balanced train_balanced split, validated on the natural validation split, and the best checkpoint is chosen by macro-F1. LoRA adapters are merged into the backbone before release, so the model loads with a plain from_pretrained (no PEFT required).

Role in DecomposeRL

DecomposeRL trains a claim-verification policy with GRPO over a seven-reward ensemble. Five of those rewards are scored by an LLM judge, which dominates training-time GPU cost. The tiny-judge stack replaces that 32B judge with eight small distilled heads so reward scoring runs on the same single GPU as training. See the paper (tiny-judge ablation) and the DecomposeRL-7B model for the full reward design.

Intended Use

  • —In-scope: serving as a fast reward / scoring model inside the DecomposeRL training loop, or as a standalone classifier for the specific judgment above on claim-decomposition traces.
  • —Out-of-scope: general-purpose fact-checking, use on inputs that do not follow the input format above, or as a standalone end-to-end claim verifier (use DecomposeRL-7B for that).

Citation

bibtex
@article{dipta2025decomposerl,
  title={DecomposeRL: Learning to Ask Useful, Informative, and Diverse Questions for Semi-Supervised, Traceable Claim Verification},
  author={Shubhashis Roy Dipta and Ankur Padia and Francis Ferraro},
  year={2025},
  eprint={2605.27858},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2605.27858v1},
}

License

Released under the Apache 2.0 License.