Team Ai
Datasetpublic

AmberTraceLabs/alignment-matrix-results

AmberTrace — Certified Alignment Matrix (results) Leaderboard data for the Certified Alignment Matrix Space — how faithfully open-weight models stay to a machine-checked decision policy as they reason. One row per model (20 models, 20 ranked) over the 1,350-item decision_eval_v1 corpus, scored against the proof-certified AmberTrace oracle (single sample, temperature 0). The headline is not accuracy but the direction of the errors — fail-open (under-restriction) on the… See the full description on the dataset page: https://huggingface.co/datasets/AmberTraceLabs/alignment-matrix-results.

sourceHugging Faceapache-2.0updated 25d agoView on Hugging Face
0likes55downloads
Dataset Card

AmberTrace — Certified Alignment Matrix (results)

Leaderboard data for the Certified Alignment Matrix Space — how faithfully open-weight models stay to a machine-checked decision policy as they reason. One row per model (20 models, 20 ranked) over the 1,350-item decision_eval_v1 corpus, scored against the proof-certified AmberTrace oracle (single sample, temperature 0). The headline is not accuracy but the direction of the errors — fail-open (under-restriction) on the safety-critical band is the failure a plain accuracy number hides.

This is scores, not a key. Per the AT = gold guardrail, no (features -> certified decision) pair is published — the certificate is obtained live from AmberTrace at eval time. This file records how models did against it.

Columns

columnmeaning
model, lab, paramsmodel, publisher, parameter count
reasoningthinking-enabled / reasoning-disabled / plain
cas_balancedcomposite alignment score, BALANCED scheme (headline, higher is better)
cas_safety_first, cas_capital_adequacyCAS under the other two penalty schemes
accuracyraw accuracy
fail_open_restrictivefail-open rate on the safety-critical band (lower is safer)
signed_bias(over-permit − over-deny)/n; negative = net cautious, positive = net fail-open
refusal_rate, parse_raterefusals; fraction parsed into an action
acc_by_structure, acc_by_action_countreasoning-complexity profile
capturelink to the capture behind the row

Reproduce / add your model

See the alignment matrix doc and the narrative writeup. Install ambertrace-rlvr (PyPI) and run your model with examples/run_alignment_matrix.py; scoring is live against AmberTrace, so a new row cannot be gamed by memorising a key.

  • —Repo: https://github.com/ambertrace-labs/ambertrace-rlvr
  • —PyPI: https://pypi.org/project/ambertrace-rlvr/