Team Ai
Datasetpublic

EdwardoSunny/agent-cwm-rubrics-debug

agent-cwm rubric library + P/R debug bundle (large split: 27 mine / 37 held-out) Layout library/err__*.md — 139 error rubrics (frontmatter exception_class: = the class each commits to) library/perf_rubrics/ — 209 performance rubrics (P1 skeleton); perf_rubrics_gated/ = 92 that passed the causal gate (own patch improved own source program above measured noise; gate_manifest.json has the strict list) library/runtime_rubrics/ — 257 runtime-cost rubrics (not part of… See the full description on the dataset page: https://huggingface.co/datasets/EdwardoSunny/agent-cwm-rubrics-debug.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes1.2kdownloads
Dataset Card

agent-cwm rubric library + P/R debug bundle (large split: 27 mine / 37 held-out)

Layout

  • —library/err__*.md — 139 error rubrics (frontmatter exception_class: = the class each commits to)
  • —library/perf_rubrics/ — 209 performance rubrics (P1 skeleton); perf_rubrics_gated/ = 92 that passed the causal gate (own patch improved own source program above measured noise; gate_manifest.json has the strict list)
  • —library/runtime_rubrics/ — 257 runtime-cost rubrics (not part of P/R by protocol)
  • —library/dupes/, library/lint_quarantine/ — removed near-duplicates and banned universal templates, with manifests
  • —library/candidates_perf_k27.jsonl — perf candidates incl. provenance (source competition/state, score delta, proposed_patch)
  • —results/pr_results_targeted_k27.json — the headline run: 145-err/209-perf judged blind on 347 held-out states
  • —results/pr_decisions_targeted_v1.jsonl — EVERY judge decision (rubricid, stateid, fires, evidence) — start debugging here
  • —results/truth_annotations_v5.jsonl — perf ground truth (annotator credits per weak state; explains_ids literal, explains_analogous_ids diagnostic)
  • —results/eval_states_v5.jsonl / eval_scored_v5.jsonl — the frozen eval corpus (held-out programs + recorded outcomes)
  • —results/gate_k27/ — causal-gate reports per competition

Headline numbers (frozen eval, 37 held-out comps)

familyprecisionrecallnotes
error (139)0.60 strict / 0.81 defect-real0.05 library-levelstrict = crashed with the named class; recall = failing states caught by >=1 rubric
perf gated (92)0.160.55truth = annotator credits criterion as literally explaining the weak-vs-strong measured gap
perf ungated (209)0.090.62controls (generic/shuffled) = 0.000 everywhere

Known diagnoses (see repo AGENTS.md for full history)

  • —Perf FPs concentrate on states whose gap NO criterion explains (coverage penalty), plus real-but-secondary patterns (judge can't see the gap; annotator can).
  • —Error recall is coverage-limited: exact-pattern rubrics vs a very diverse held-out failure pool; universal templates that fake recall are banned (see lint_quarantine).
  • —In flight: transfer gate (fix must improve programs it wasn't mined from) and a 27-comp high-volume error wave. Code: github.com/7peng/agent-cwm (rubric_gen/).