Team Ai
20 results

rubric

cm2435-new /gdpval_preference_rubricsaudion<1K0 likes2.7k downloads6mo agoHugging Facedanikhan632 /natural_reasoning_rubricstext1K<n<10K0 likes1.2k downloads1y agoHugging FaceEdwardoSunny /agent-cwm-rubrics-debug agent-cwm rubric library + P/R debug bundle (large split: 27 mine / 37 held-out) Layout library/err__*.md — 139 error rubrics (frontmatter exception_class: = the class each commits to) library/perf_rubrics/ — 209 performance rubrics (P1 skeleton); perf_rubrics_gated/ = 92 that passed the causal gate (own patch improved own source program above measured noise; gate_manifest.json has the strict list) library/runtime_rubrics/ — 257 runtime-cost rubrics (not part of… See the full description on the dataset page: https://huggingface.co/datasets/EdwardoSunny/agent-cwm-rubrics-debug.0 likes1.2k downloads2mo agoHugging FacesojuL /RubricHub_v1 RubricHub RubricHub is a large-scale (approximately 110K), multi-domain dataset that provides high-quality rubric-based supervision for open-ended generation tasks. It is constructed via an automated coarse-to-fine rubric generation framework, which integrates principle-guided synthesis, multi-model aggregation, and difficulty evolution to produce comprehensive and highly discriminative evaluation criteria, overcoming the supervision ceiling of… See the full description on the dataset page: https://huggingface.co/datasets/sojuL/RubricHub_v1.texttext-generation100K<n<1M271 likes858 downloads8mo agoHugging FaceMasterVito /swe-agent-tool-rubrics-860 SWE Agent 逐 turn 工具调用评判数据集(860 个决策点) 本数据集来自 2026-08-06 的一次实验:**从真实 SWE agent 轨迹中归纳"怎么判断一次工具调用的好坏"**。 包含两个文件: 文件 行数 大小 内容 cases.jsonl 860 5.0 MB 决策点原始数据(题目、历史、两个候选命令、执行结果、现役判官打分) map_io.jsonl 860 9.6 MB 每个决策点喂给 GPT-5.6 的完整 prompt 原文与完整回复 两个文件通过 case_id 一一对应。 背景:为什么是"按动作分类"而不是"按工具分类" 轨迹来自 slime 的 minimal harness,该 harness 只暴露一个工具 bash (slime/agent/harness/minimal.py 里的 BASH_TOOL),全部 328,270 次调用的工具名都是 bash。 所以"不同工具用不同 rubric"无法按工具名实现,只能按命令在干什么分类。… See the full description on the dataset page: https://huggingface.co/datasets/MasterVito/swe-agent-tool-rubrics-860.tabular1K<n<10K1 likes503 downloads2mo agoHugging Faceskylenage-ai /RubricRM-Data Link GitHub: SKYLENAGE-AI/SKYLENAGE-JUDGER Hugging Face Models: skylenage-ai/SkyJM-Gen-4B skylenage-ai/SkyJM-Gen-9B skylenage-ai/SkyJM-Edit-4B skylenage-ai/SkyJM-Edit-9B Hugging Face Dataset: skylenage-ai/RubricRM-Data ModelScope Models: SKYLENAGE/SkyJM-Gen-4B SKYLENAGE/SkyJM-Gen-9B SKYLENAGE/SkyJM-Edit-4B SKYLENAGE/SkyJM-Edit-9B Citation If you find this dataset useful, please cite our paper: @misc{kan2026rubricrmgenerativerewardmodeling… See the full description on the dataset page: https://huggingface.co/datasets/skylenage-ai/RubricRM-Data.image10K<n<100K2 likes446 downloads1mo agoHugging Face