rubric
Datasets
All datasets matching “rubric”gdpval_preference_rubricsnatural_reasoning_rubricsagent-cwm-rubrics-debug
agent-cwm rubric library + P/R debug bundle (large split: 27 mine / 37 held-out)
Layout
library/err__*.md — 139 error rubrics (frontmatter exception_class: = the class each commits to)
library/perf_rubrics/ — 209 performance rubrics (P1 skeleton); perf_rubrics_gated/ = 92 that passed the causal gate (own patch improved own source program above measured noise; gate_manifest.json has the strict list)
library/runtime_rubrics/ — 257 runtime-cost rubrics (not part of… See the full description on the dataset page: https://huggingface.co/datasets/EdwardoSunny/agent-cwm-rubrics-debug.RubricHub_v1
RubricHub
RubricHub is a large-scale (approximately 110K), multi-domain dataset that provides high-quality rubric-based supervision for open-ended generation tasks. It is constructed via an automated coarse-to-fine rubric generation framework, which integrates principle-guided synthesis, multi-model aggregation, and difficulty evolution to produce comprehensive and highly discriminative evaluation criteria, overcoming the supervision ceiling of… See the full description on the dataset page: https://huggingface.co/datasets/sojuL/RubricHub_v1.swe-agent-tool-rubrics-860
SWE Agent 逐 turn 工具调用评判数据集(860 个决策点)
本数据集来自 2026-08-06 的一次实验:**从真实 SWE agent 轨迹中归纳"怎么判断一次工具调用的好坏"**。
包含两个文件:
文件
行数
大小
内容
cases.jsonl
860
5.0 MB
决策点原始数据(题目、历史、两个候选命令、执行结果、现役判官打分)
map_io.jsonl
860
9.6 MB
每个决策点喂给 GPT-5.6 的完整 prompt 原文与完整回复
两个文件通过 case_id 一一对应。
背景:为什么是"按动作分类"而不是"按工具分类"
轨迹来自 slime 的 minimal harness,该 harness 只暴露一个工具 bash
(slime/agent/harness/minimal.py 里的 BASH_TOOL),全部 328,270 次调用的工具名都是 bash。
所以"不同工具用不同 rubric"无法按工具名实现,只能按命令在干什么分类。… See the full description on the dataset page: https://huggingface.co/datasets/MasterVito/swe-agent-tool-rubrics-860.RubricRM-Data
Link
GitHub: SKYLENAGE-AI/SKYLENAGE-JUDGER
Hugging Face Models:
skylenage-ai/SkyJM-Gen-4B
skylenage-ai/SkyJM-Gen-9B
skylenage-ai/SkyJM-Edit-4B
skylenage-ai/SkyJM-Edit-9B
Hugging Face Dataset: skylenage-ai/RubricRM-Data
ModelScope Models:
SKYLENAGE/SkyJM-Gen-4B
SKYLENAGE/SkyJM-Gen-9B
SKYLENAGE/SkyJM-Edit-4B
SKYLENAGE/SkyJM-Edit-9B
Citation
If you find this dataset useful, please cite our paper:
@misc{kan2026rubricrmgenerativerewardmodeling… See the full description on the dataset page: https://huggingface.co/datasets/skylenage-ai/RubricRM-Data.
