datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gdpval_preference_rubricsnatural_reasoning_rubricsagent-cwm-rubrics-debug
agent-cwm rubric library + P/R debug bundle (large split: 27 mine / 37 held-out)
Layout
library/err__*.md — 139 error rubrics (frontmatter exception_class: = the class each commits to)
library/perf_rubrics/ — 209 performance rubrics (P1 skeleton); perf_rubrics_gated/ = 92 that passed the causal gate (own patch improved own source program above measured noise; gate_manifest.json has the strict list)
library/runtime_rubrics/ — 257 runtime-cost rubrics (not part of… See the full description on the dataset page: https://huggingface.co/datasets/EdwardoSunny/agent-cwm-rubrics-debug.RubricHub_v1
RubricHub
RubricHub is a large-scale (approximately 110K), multi-domain dataset that provides high-quality rubric-based supervision for open-ended generation tasks. It is constructed via an automated coarse-to-fine rubric generation framework, which integrates principle-guided synthesis, multi-model aggregation, and difficulty evolution to produce comprehensive and highly discriminative evaluation criteria, overcoming the supervision ceiling of… See the full description on the dataset page: https://huggingface.co/datasets/sojuL/RubricHub_v1.swe-agent-tool-rubrics-860
SWE Agent 逐 turn 工具调用评判数据集(860 个决策点)
本数据集来自 2026-08-06 的一次实验:**从真实 SWE agent 轨迹中归纳"怎么判断一次工具调用的好坏"**。
包含两个文件:
文件
行数
大小
内容
cases.jsonl
860
5.0 MB
决策点原始数据(题目、历史、两个候选命令、执行结果、现役判官打分)
map_io.jsonl
860
9.6 MB
每个决策点喂给 GPT-5.6 的完整 prompt 原文与完整回复
两个文件通过 case_id 一一对应。
背景:为什么是"按动作分类"而不是"按工具分类"
轨迹来自 slime 的 minimal harness,该 harness 只暴露一个工具 bash
(slime/agent/harness/minimal.py 里的 BASH_TOOL),全部 328,270 次调用的工具名都是 bash。
所以"不同工具用不同 rubric"无法按工具名实现,只能按命令在干什么分类。… See the full description on the dataset page: https://huggingface.co/datasets/MasterVito/swe-agent-tool-rubrics-860.RubricRM-Data
Link
GitHub: SKYLENAGE-AI/SKYLENAGE-JUDGER
Hugging Face Models:
skylenage-ai/SkyJM-Gen-4B
skylenage-ai/SkyJM-Gen-9B
skylenage-ai/SkyJM-Edit-4B
skylenage-ai/SkyJM-Edit-9B
Hugging Face Dataset: skylenage-ai/RubricRM-Data
ModelScope Models:
SKYLENAGE/SkyJM-Gen-4B
SKYLENAGE/SkyJM-Gen-9B
SKYLENAGE/SkyJM-Edit-4B
SKYLENAGE/SkyJM-Edit-9B
Citation
If you find this dataset useful, please cite our paper:
@misc{kan2026rubricrmgenerativerewardmodeling… See the full description on the dataset page: https://huggingface.co/datasets/skylenage-ai/RubricRM-Data.llm-metric-tuluswerl-tmax-15k-rubric-gpt-5-6-sol
swerl-tmax-15k with task-quality rubric labels (gpt-5-6-sol)
hamishivi/swerl-tmax-15k, unchanged and unfiltered, with a per-task quality
label attached as extra columns.
This is not a verified or filtered dataset. Every one of the 14,601 original
records is present. Nothing has been dropped, repaired, or reordered. The labels
are one model's judgement about whether each task is sound enough to be useful RL
training data — an annotation layer, not a correctness guarantee.… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-rubric-gpt-5-6-sol.ai-vs-human-rubric-companion-data
Companion dataset for the AI-vs-human rubric study
This dataset is the data side of an anonymous submission. It pairs with a separate anonymous code repository that contains the runnable scripts, validators, and documentation. The two artifacts together reproduce every paper-facing headline number without re-running any API-backed stage. The code URL for review is https://anonymous.4open.science/r/codereviewer-47F3/README.md.
Paper sections and where their data are… See the full description on the dataset page: https://huggingface.co/datasets/forreview43/ai-vs-human-rubric-companion-data.v-rubrics-50k
V-Rubrics 50K
License notice — read before use: This aggregate package does not have a
single uniform data license. The V-Rubrics code license does not relicense
the upstream dataset records or embedded images. Every record remains
subject to its source dataset's license, terms, attribution requirements,
and usage restrictions. Review all applicable upstream terms before use,
redistribution, or commercial deployment.
V-Rubrics 50K is a training-only collection of 50,248… See the full description on the dataset page: https://huggingface.co/datasets/v-rubrics/v-rubrics-50k.amazon-c2-varied-rubrics
Amazon C2 varied-rubric distillation
This release exposes six balanced C2 SFT configurations: latent-state and non-diverse candidate panels at
K=1, K=2, and K=4 rubrics per retained reviewer. Each rubric-writer target is paired with one full-rubric
listwise judge target over the same variant's frozen 40-candidate panel. The K arms within a variant share one
reviewer cohort and are exact nested prefixes.
Config
Train rows
Validation
Test
Train reviewers… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c2-varied-rubrics.mR3-Dataset-Filtered3PolyGuardMix-en_prompt_en_thinking-filtered_correctrubric-rewards
Replication Data for "From Constitutions to Control: Interpretable Rewards for Aligning Language Models"
This repository contains all replication data for "From Constitutions to Control: Interpretable Rewards for Aligning Language Models"; replication code is available at
github.com/jgaeb/rubric-rewards, and the fine-tuned adapters are available from the model repository
jgaeb/rubric-rewards-adapters.
Data are available at three levels of processing:
raw/: Raw data stored as… See the full description on the dataset page: https://huggingface.co/datasets/jgaeb/rubric-rewards.mR3-Dataset-Filtered2R3-full-datasetrubricbench
Summary
RubricBench is a curated benchmark comprising 1,147 pairwise comparisons specifically designed to assess the reliability of rubric-guided evaluation.
It addresses the lack of a unified benchmark with both the discriminative complexity and the ground-truth rubric annotations required for rigorous analysis.
Each sample is augmented with expert-annotated, atomic rubrics derived strictly from instructions.
Dataset Structure & Domains
The dataset spans five… See the full description on the dataset page: https://huggingface.co/datasets/DonJoey/rubricbench.Rubrics
Dataset Card for Dataset Name
This dataset is from the paper "Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training."
Dataset Details
The dataset contains 5000 training prompts and the corresponding rubrics for the health domain and the generalist domains. Each domain also includes 1000 additional prompts for the end-to-end win rate comparison. Additionally, the finance domain contains 1145 training prompts with rubrics.… See the full description on the dataset page: https://huggingface.co/datasets/JunkaiZ/Rubrics.llm-metric-ultrafeedbackPolyGuardMix-tgt_prompt_tgt_thinkingllm-metric-ultrafeedback-newdeepwriting-rubrics-llama3-2mR3-Dataset-Filtered1-no-PolyGuardR3-full-dataset-no-gluerubric_rl_results
Rubric RL Evaluation Results
Evaluation data for rubric-based reward modeling experiments. Contains generated rubrics from multiple rubric generators and pairwise scoring results comparing rl-research/DR-Tulu-8B (RL, step_4000) vs rl-research/DR-Tulu-SFT-8B.
Data Structure
rubrics/ — Generated evaluation rubrics
Each JSONL file contains per-question rubrics with fields: prompt_id, question, generated_rubric, generated_rubric_raw, rubric_style, rubric_model.… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/rubric_rl_results.llm-metric-mrewardbenchmR3-Dataset-100K-EasyToHard
mR3 Dataset: Multilingual Rubric-Agnostic Reward Reasoning
Project Page | Paper | Code
This is the dataset used to train mR3, a massively multilingual, rubric-agnostic reward reasoning model.
Dataset Summary
The mR3 training dataset contains 100,000 high-quality samples curated from an initial pool of 4 million samples across 125 languages. It is designed to train reward models that can provide reasoning traces in both English and non-English settings, covering 72… See the full description on the dataset page: https://huggingface.co/datasets/rubricreward/mR3-Dataset-100K-EasyToHard.swe-rubrics-mined-1000rubric-grounded-faithfulness-eval
Rubric-Grounded Faithfulness Evaluation Resource
This repository hosts the anonymized evaluation resource accompanying the NeurIPS 2026 Evaluations and Datasets submission:
From Scores to Checks: Rubric-Grounded Faithfulness Evaluation for AI-Generated Images
The release contains the derived assets behind the paper's main claims: full-gold AIGCIQA2023 rubric labels, evidence-point and reviewed-counterfactual diagnostic subsets, a relabeled 2400-image T2I-CompBench human-eval… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1-afk-ops/rubric-grounded-faithfulness-eval.task-intents-and-rubrics
