Team Ai
Datasetpublic

CompilingThings/compile-benchmark

CompilingThings Compile Benchmark for MQL5® This release evaluates compile success of generated MQL5 on two private held-out sets. The first is 300 Expert Advisor prompts, run on four arms: the base model, two tuned local models and one frontier API model. The second is 200 non-EA prompts (include files, custom indicators, scripts and services, 50 each), run on the three local arms, plus a stability re-run of one of them. The holdout results are attested, not fully verifiable:… See the full description on the dataset page: https://huggingface.co/datasets/CompilingThings/compile-benchmark.

sourceHugging Faceotherupdated 27d agoView on Hugging Face
1likes131downloads
verify_public_release.py638 linesDownload Raw Back to root
1#!/usr/bin/env python32"""Verify the public release, and ship with the package.3 4Needs Python 3.9 or later and nothing else - no third-party packages. Invoke it with5whichever launcher your system has: `py`, `python3` or `python`.6 7Most checks here are ones a third party can run on the published files. Checks 1 and 28are not: they need the evaluation pool, the frozen manifest and the id-to-name mapping,9none of which is distributed, and they report SKIPPED without them. Hashes are computed10with the semantics the harness used - sha256 over the UTF-8 encoding of the string, no11normalisation - so a mismatch is a real finding rather than an encoding artefact.12 13Checks:14  1  authority file hashes to the source_sha256 the frozen manifest recorded15  2  every published prompt is byte-identical to the evaluated prompt16  3  both published hashes recompute from the published prompt for every row, under17     each arm's rendering rule (local template for base/tuned; system+user for frontier)18  4  each item maps to exactly three result rows: one base, one tuned, one frontier19  5  the public rows reproduce every published table: 2/184, 170/184, 179/184, the20     base-vs-tuned contingency table and the tuned-vs-frontier paired table21  6  substituting into the published template reproduces prompt_sha256 for every22     local-arm row23  7  every SHA-256 quoted in the card is one the release stands behind24  8  every checksum-manifest entry recomputes from the shipped files25  9  the three v1.1 row files carry the row count, the arm ids and the per-arm row26     counts the card publishes27 10  every published headline count falls out of the rows under the pre-registered28     scoring rule: a pass is verdict_headline true AND truncated false29 11  every published per-arm truncation count falls out of the rows30 12  the arms of a row file cover one and the same item set, once each31 13  corpus_row_hashes.json carries 219,867 row hashes and agrees with its own count32  D  diagnostic: do the sealed spec_sha256 / prompt_sha256 reproduce from the prompt?33 34Checks 9 to 13 need nothing but the published files. They run on the holdout rows even35though the holdout prompts are private: the arithmetic of an attested result is still36arithmetic, and it is the only thing about the holdouts a reader can check.37 38Checks 1 and 2 need the evaluation pool, the frozen manifest and the id-to-name mapping.39None of those is distributed, so for a public reader they report SKIPPED and the run40continues. Checks 3 through 8 need nothing but the published files.41 42Check 6 is not redundant with check 3. Check 3 renders the prompt with this file's own43format_local_prompt; check 6 renders it the way a reader does, by substituting into the44template published in serving_template.json. They agree only if the published template is45correct - and it once was not.46 47Exit 0 all checks pass, 2 a check failed, 3 an input was missing or unreadable.48 49Usage:50    python3 verify_public_release.py --public .51 52With the undistributed inputs to hand, checks 1 and 2 also run:53    python3 verify_public_release.py --public <dir> --private <dir> \54        --authority <pool.jsonl> --manifest <frozen manifest.json>55"""56 57from __future__ import annotations58 59import argparse60import hashlib61import json62import re63import sys64from collections import defaultdict65from pathlib import Path66 67MANIFEST_NAME = "SHA256SUMS.txt"68 69SYSTEM_PROMPT = (70    "You are an expert MQL5 programmer. Write the complete MQL5 "71    "Expert Advisor code that implements the given specification exactly."72)73 74EXPECTED = {"n": 184, "base": 2, "tuned": 170, "a": 1, "b": 1, "c": 169, "d": 13}75# The frontier arm and its pairing against the tuned arm, added 2026-09-02.76EXPECTED_FRONTIER = {"n": 184, "frontier": 179,77                     "both_pass": 168, "tuned_only": 2,78                     "frontier_only": 11, "both_fail": 3}79 80CORPUS_NAME = "corpus_row_hashes.json"81EXPECTED_CORPUS_ROWS = 21986782 83# Every count the card publishes for the v1.1 row files, keyed by file and arm.84# `passes` is the headline figure under the pre-registered rule below; it is NOT a85# count of rows whose verdict_headline is true. On the 220k arm of the non-EA holdout86# the two differ by one - a row compiled despite being truncated - and that single87# item is why the rule is stated here rather than assumed.88EXPECTED_ROW_FILES = {89    "holdout_ea300_results.jsonl": {90        "rows": 1200, "items": 300, "item_key": "item_sha256",91        "arms": {92            "base": {"rows": 300, "passes": 1, "truncated": 3},93            "tuned83k": {"rows": 300, "passes": 281, "truncated": 0},94            "tuned220k": {"rows": 300, "passes": 282, "truncated": 0},95            "frontier": {"rows": 300, "passes": 286, "truncated": 4},96        },97    },98    "holdout_nonea200_results.jsonl": {99        "rows": 800, "items": 200, "item_key": "item_sha256",100        "arms": {101            "base": {"rows": 200, "passes": 56, "truncated": 14},102            "tuned83k": {"rows": 200, "passes": 135, "truncated": 13},103            "tuned220k": {"rows": 200, "passes": 168, "truncated": 12},104            "tuned220k_q8_replay": {"rows": 200, "passes": 169, "truncated": 11},105        },106    },107    "bridge_q8_184_results.jsonl": {108        "rows": 552, "items": 184, "item_key": "item_id",109        "arms": {110            "base": {"rows": 184, "passes": 2, "truncated": 2},111            "tuned83k": {"rows": 184, "passes": 175, "truncated": 0},112            "tuned220k": {"rows": 184, "passes": 169, "truncated": 0},113        },114    },115}116 117 118def scored_pass(row: dict) -> bool:119    """The pre-registered scoring rule, and the whole of it: a generation cut off at120    the token ceiling counts as a fail even if what it produced happened to compile."""121    return row.get("verdict_headline") is True and row.get("truncated") is not True122 123 124def format_api_prompt(system: str, spec: str) -> str:125    """Canonical rendering of the frontier arm's system+user pair. The vendor API126    takes the two as separate fields; the frontier rows' prompt_sha256 hashes this127    single-string rendering of them."""128    return f"{system}\n\n{spec}"129 130 131def sha256_text(text: str) -> str:132    """Byte-for-byte the harness's hash rule: sha256 over the UTF-8 encoding."""133    return hashlib.sha256(text.encode("utf-8")).hexdigest()134 135 136def sha256_file(path: Path) -> str:137    h = hashlib.sha256()138    with path.open("rb") as fh:139        for chunk in iter(lambda: fh.read(1 << 20), b""):140            h.update(chunk)141    return h.hexdigest()142 143 144def read_jsonl(path: Path):145    with path.open("r", encoding="utf-8") as fh:146        for lineno, line in enumerate(fh, 1):147            line = line.strip()148            if not line:149                continue150            try:151                yield json.loads(line)152            except json.JSONDecodeError as exc:153                raise SystemExit(f"{path}:{lineno}: {exc}")154 155 156def format_local_prompt(spec: str) -> str:157    """The serving template, rendered exactly as the harness rendered it."""158    return (159        f"<|system|>{SYSTEM_PROMPT}<|end|>\n"160        f"<|user|>{spec}<|end|>\n"161        f"<|assistant|>"162    )163 164 165def main(argv: list[str]) -> int:166    ap = argparse.ArgumentParser(description=__doc__)167    # Only --public is required. The authority pool, the frozen manifest and the id-to-name168    # mapping are not distributed, so a public reader has none of them; requiring any of169    # them made the script exit before check 3 and the release was not runnable by anyone170    # outside. Checks 3, 4 and 5 need nothing but the published files, and they are the171    # checks that matter to a reader: the hashes recompute, the pairing is one-to-one, and172    # the headline falls out of the rows.173    ap.add_argument("--public", type=Path, required=True)174    ap.add_argument("--private", type=Path,175                    help="holds item_id_mapping.jsonl; not distributed, so check 2 skips")176    ap.add_argument("--authority", type=Path,177                    help="the evaluation pool; not distributed, so check 1 skips")178    ap.add_argument("--manifest", type=Path,179                    help="the frozen manifest; not distributed, so check 1 skips")180    ap.add_argument("--sealed-rows", type=Path,181                    help="the sealed base-arm results file (not distributed), "182                         "for the reproduction diagnostic")183    args = ap.parse_args(argv)184 185    failures: list[str] = []186    notes: list[str] = []187 188    # ---- check 1 -----------------------------------------------------------------189    # Runs only when the pool and the frozen manifest are both to hand. Neither ships, so190    # for a public reader this reports SKIPPED and the run continues.191    have_authority = args.authority is not None and args.authority.is_file()192    have_manifest = args.manifest is not None and args.manifest.is_file()193    authority: dict[str, str] = {}194 195    if have_authority and have_manifest:196        manifest = json.loads(args.manifest.read_text(encoding="utf-8"))197        actual = sha256_file(args.authority)198        if actual == manifest["source_sha256"]:199            print(f"  [1] authority sha256 matches the frozen manifest  {actual[:16]}...")200        else:201            failures.append(202                f"[1] authority sha256 {actual} != manifest source_sha256 "203                f"{manifest['source_sha256']}"204            )205        # The authority is the only source of prompt text. Published prompts are checked206        # against it rather than the other way round.207        for rec in read_jsonl(args.authority):208            name = rec.get("ea_name")209            if name is not None:210                authority[name] = rec["prompt"]211    else:212        print("  [1] SKIPPED - needs the evaluation pool and the frozen manifest, "213              "neither of which is distributed")214 215    # Checks 3-8 verify the PUBLIC release, so they read the public prompts file216    # unconditionally. An earlier version fell back between the two, which meant a217    # present --private copy silently substituted for the release copy and the218    # verifier could pass a public file that no longer matched its results.219    public_prompts = args.public / "prompts.jsonl"220    if not public_prompts.is_file():221        print(f"MISSING: prompts.jsonl in {args.public}")222        return 3223 224    published = list(read_jsonl(public_prompts))225 226    # ---- check 2 -----------------------------------------------------------------227    # Needs the id-to-name mapping, which is not distributed. A public reader gets a228    # clear SKIPPED rather than a crash; internally the mapping is present and it runs.229    id_to_name: dict[str, str] = {}230    mapping_path = args.private / "item_id_mapping.jsonl" if args.private else None231    if mapping_path is not None and mapping_path.is_file():232        id_to_name = {r["item_id"]: r["ea_name"] for r in read_jsonl(mapping_path)}233 234    if not id_to_name or not authority:235        print("  [2] SKIPPED - needs the evaluation pool and the id-to-name mapping, "236              "neither of which is distributed")237    else:238        mismatched = [239            r["item_id"] for r in published240            if authority.get(id_to_name.get(r["item_id"], "")) != r["prompt"]241        ]242        if mismatched:243            failures.append(244                f"[2] {len(mismatched)} published prompt(s) differ from the evaluated "245                f"prompt"246            )247        else:248            print(f"  [2] all {len(published)} published prompts are byte-identical to "249                  f"the evaluated prompt")250 251    # ---- check 3 -----------------------------------------------------------------252    # The check any reader can run with nothing but the published files.253    public_hash = {r["item_id"]: sha256_text(r["prompt"]) for r in published}254    results_path = args.public / "per_item_results.jsonl"255    if not results_path.is_file():256        # The docstring's exit-3 contract for missing inputs, honoured here too —257        # without this a missing results file was an uncaught traceback.258        print(f"MISSING: per_item_results.jsonl in {args.public}")259        return 3260    rows = list(read_jsonl(results_path))261    prompt_by_id = {p["item_id"]: p["prompt"] for p in published}262    # The published prompt_sha256 per item for the LOCAL arms, for check 6 to compare a263    # reader-built template render to. Frontier rows hash a different rendering (the264    # API's system+user pair) and are checked under that rule in check 3.265    by_id_hash = {r["item_id"]: r.get("prompt_sha256") for r in rows266                  if r.get("arm") != "frontier"}267 268    hash_bad = [269        r["item_id"] for r in rows270        if public_hash.get(r["item_id"]) != r.get("prompt_sha256_public")271    ]272    tmpl_bad_rows = []273    for r in rows:274        if r["item_id"] not in prompt_by_id:275            continue276        prompt = prompt_by_id[r["item_id"]]277        want = (format_api_prompt(SYSTEM_PROMPT, prompt)278                if r.get("arm") == "frontier" else format_local_prompt(prompt))279        if sha256_text(want) != r.get("prompt_sha256"):280            tmpl_bad_rows.append(r["item_id"])281    if hash_bad:282        failures.append(f"[3] prompt_sha256_public does not recompute for "283                        f"{len(set(hash_bad))} item(s)")284    elif tmpl_bad_rows:285        failures.append(f"[3] prompt_sha256 does not recompute from the published "286                        f"serving template for {len(set(tmpl_bad_rows))} item(s)")287    else:288        print(f"  [3] both published hashes recompute from the published prompt for all "289              f"{len(public_hash)} items ({len(rows)} rows)")290 291    # ---- check 4 -----------------------------------------------------------------292    by_item: dict[str, list[dict]] = defaultdict(list)293    for r in rows:294        by_item[r["item_id"]].append(r)295    bad = {296        i: sorted(x["arm"] for x in v)297        for i, v in by_item.items()298        if sorted(x["arm"] for x in v) != ["base", "frontier", "tuned"]299    }300    if bad:301        failures.append(f"[4] {len(bad)} item(s) do not map to exactly one base, one "302                        f"tuned and one frontier row: {list(bad.items())[:3]}")303    else:304        print(f"  [4] all {len(by_item)} items map to exactly one base, one tuned and "305              f"one frontier row")306 307    # ---- check 5 -----------------------------------------------------------------308    # The headline is a claim about verdict_headline, so it is measured on that field.309    # The two verdict fields can and do diverge - warnings fail strict but not headline,310    # and some frontier rows show exactly that - so the divergence is counted and printed311    # below rather than assumed away, and the tables are measured on verdict_headline312    # only. An earlier version read verdict_strict and got away with it only while the313    # two happened to agree.314    disagree = [315        i for i, v in by_item.items()316        if any(x["verdict_headline"] != x["verdict_strict"] for x in v)317    ]318    if disagree:319        notes.append(f"[5] NOTE: verdict_headline and verdict_strict disagree on "320                     f"{len(disagree)} item-arm row(s); the headline below is measured on "321                     f"verdict_headline")322    else:323        notes.append(f"[5] verdict_headline and verdict_strict agree on all "324                     f"{len(rows)} rows")325 326    a = b = c = d = 0327    fp = f_both_pass = f_tuned_only = f_frontier_only = f_both_fail = 0328    for v in by_item.values():329        base = next(x for x in v if x["arm"] == "base")["verdict_headline"]330        tuned = next(x for x in v if x["arm"] == "tuned")["verdict_headline"]331        frontier = next(x for x in v if x["arm"] == "frontier")["verdict_headline"]332        if base and tuned:333            a += 1334        elif base and not tuned:335            b += 1336        elif tuned:337            c += 1338        else:339            d += 1340        if frontier:341            fp += 1342        if tuned and frontier:343            f_both_pass += 1344        elif tuned:345            f_tuned_only += 1346        elif frontier:347            f_frontier_only += 1348        else:349            f_both_fail += 1350    got = {"n": len(by_item), "base": a + b, "tuned": a + c,351           "a": a, "b": b, "c": c, "d": d}352    got_frontier = {"n": len(by_item), "frontier": fp,353                    "both_pass": f_both_pass, "tuned_only": f_tuned_only,354                    "frontier_only": f_frontier_only, "both_fail": f_both_fail}355    if got == EXPECTED and got_frontier == EXPECTED_FRONTIER:356        print(f"  [5] headline reproduces from the public rows: base {got['base']}/"357              f"{got['n']}, tuned {got['tuned']}/{got['n']}, frontier "358              f"{fp}/{got['n']}, a={a} b={b} c={c} d={d}; tuned-vs-frontier "359              f"{f_both_pass}/{f_tuned_only}/{f_frontier_only}/{f_both_fail}")360    else:361        failures.append(f"[5] headline does not reproduce. expected {EXPECTED} + "362                        f"{EXPECTED_FRONTIER}, got {got} + {got_frontier}")363    # ---- check 6 -----------------------------------------------------------------364    # The reader's path, and the reason this check exists. Checks 3 renders the prompt365    # with format_local_prompt, a function compiled into this script. A reader has no such366    # function: they parse serving_template.json, substitute the two placeholders, and367    # hash. Those are different routes to the same string only if the published template368    # is right. It once was not - the template field held the two-character sequence369    # backslash-n where a newline belonged, so every reader-side hash disagreed while370    # check 3 passed. A check that shares the harness's own definition cannot see that.371    template_path = args.public / "serving_template.json"372    if not template_path.is_file():373        failures.append("[6] serving_template.json is missing from the public release")374    else:375        tmpl = json.loads(template_path.read_text(encoding="utf-8"))376        sys_prompt = tmpl["system_prompt"]377        if sha256_text(sys_prompt) != tmpl["system_prompt_sha256"]:378            failures.append("[6] the published system prompt does not match its own "379                            "published hash")380        rendered_bad = [381            r["item_id"] for r in published382            if sha256_text(383                tmpl["template"].replace("{system_prompt}", sys_prompt)384                                .replace("{prompt}", r["prompt"])385            ) != by_id_hash.get(r["item_id"])386        ]387        if rendered_bad:388            failures.append(389                f"[6] substituting into the published template does not reproduce "390                f"prompt_sha256 for {len(set(rendered_bad))} item(s). A reader following "391                f"serving_template.json cannot reproduce our hashes."392            )393        else:394            print(f"  [6] substituting into the published template reproduces "395                  f"prompt_sha256 for all {len(published)} items")396 397 398    # ---- check 7 -----------------------------------------------------------------399    # Every 64-hex string in README.md must be a hash the release actually stands behind.400    # The card quotes file hashes in prose and in a table, and those are hand-carried from401    # the metadata when the card is written - so they go stale the moment any file is402    # regenerated. That happened: the card published a serving_template.json hash from an403    # earlier build while the manifest, the metadata and the projection report all agreed404    # on a different one. Nothing caught it, because nothing was comparing the two.405    readme = args.public / "README.md"406    manifest_file = args.public / MANIFEST_NAME407    if readme.is_file() and manifest_file.is_file():408        known: set[str] = set()409        for line in manifest_file.read_text(encoding="utf-8").splitlines():410            digest = line.split("  ")[0].strip()411            if len(digest) == 64:412                known.add(digest)413        # Hashes the release legitimately quotes that are not file hashes: model weights,414        # the corpus, the adapter, the evaluation pool, the normaliser source. Those live415        # in publication_metadata.json, so that file is the second authority.416        # serving_template.json is a third authority: it carries the system-prompt hash417        # and the worked example, which are legitimately quoted in the card and appear in418        # neither the manifest nor the metadata.419        authorities = "".join(420            (args.public / n).read_text(encoding="utf-8")421            for n in ("publication_metadata.json", "serving_template.json",422                      "PROJECTION_REPORT.json")423            if (args.public / n).is_file()424        )425        stale = sorted({426            h for h in re.findall(r"\b[0-9a-f]{64}\b", readme.read_text(encoding="utf-8"))427            if h not in known and h not in authorities428        })429        if stale:430            failures.append(431                f"[7] README.md quotes {len(stale)} sha256 value(s) that appear in none "432                f"of {MANIFEST_NAME}, publication_metadata.json, serving_template.json "433                f"or PROJECTION_REPORT.json, so the card disagrees with the release: "434                f"{stale[:3]}"435            )436        else:437            print("  [7] every sha256 quoted in the card is one the release stands behind")438    else:439        # A copy missing the card or the manifest must not verify. Without this branch440        # the check was silently skipped and a damaged copy printed ALL CHECKS PASSED.441        absent = [n for n, f in (("README.md", readme), (MANIFEST_NAME, manifest_file))442                  if not f.is_file()]443        failures.append(f"[7] cannot run: missing {', '.join(absent)}")444 445    # ---- check 8 -----------------------------------------------------------------446    # Recompute the manifest. Check 7 reads SHA256SUMS.txt to learn which hashes the447    # release stands behind, but never hashes a file - so a modified LICENSE or448    # CITATION.cff passed the verifier untouched while the card said the verifier checks449    # the release package. Third parties do not all have sha256sum, and telling them to450    # run a tool this script could run itself is not verification.451    if manifest_file.is_file():452        listed: dict[str, str] = {}453        for line in manifest_file.read_text(encoding="utf-8").splitlines():454            if not line.strip():455                continue456            digest, _, name = line.partition("  ")457            if len(digest.strip()) == 64 and name.strip():458                listed[name.strip()] = digest.strip()459 460        bad, missing = [], []461        for name, want in sorted(listed.items()):462            f = args.public / name463            if not f.is_file():464                missing.append(name)465            elif sha256_file(f) != want:466                bad.append(name)467 468        present = {p.name for p in args.public.iterdir() if p.is_file()}469        unlisted = sorted(present - set(listed) - {MANIFEST_NAME})470 471        if bad or missing:472            failures.append(473                f"[8] manifest reconciliation FAILED: {len(bad)} file(s) do not match "474                f"their recorded hash {bad[:3]}, {len(missing)} listed file(s) absent "475                f"{missing[:3]}"476            )477        elif unlisted:478            failures.append(479                f"[8] {len(unlisted)} file(s) in the release are not listed in "480                f"{MANIFEST_NAME}: {unlisted[:5]}"481            )482        else:483            print(f"  [8] all {len(listed)} manifest entries recompute, and the only "484                  f"unlisted file is {MANIFEST_NAME}, which cannot list itself")485    else:486        # Same rule as check 7: a copy with no manifest is unverifiable, not verified.487        failures.append(f"[8] cannot run: {MANIFEST_NAME} is missing")488 489    # ---- checks 9-12 -------------------------------------------------------------490    # The v1.1 row files. These ship, so a missing one is a failure and never a skip:491    # the card's headline tables are computed from them and a reader who cannot find492    # them cannot check a single holdout number.493    loaded: dict[str, list[dict]] = {}494    for fname in EXPECTED_ROW_FILES:495        path = args.public / fname496        if path.is_file():497            loaded[fname] = list(read_jsonl(path))498        else:499            failures.append(f"[9] {fname} is missing from the public release")500 501    shape_bad, pass_bad, trunc_bad, pair_bad = [], [], [], []502    for fname, want in EXPECTED_ROW_FILES.items():503        rows = loaded.get(fname)504        if rows is None:505            continue506        key = want["item_key"]507        by_arm: dict[str, list[dict]] = defaultdict(list)508        for r in rows:509            by_arm[r.get("arm")].append(r)510 511        if len(rows) != want["rows"]:512            shape_bad.append(f"{fname}: {len(rows)} rows, expected {want['rows']}")513        if sorted(by_arm) != sorted(want["arms"]):514            shape_bad.append(f"{fname}: arm ids {sorted(by_arm)}, expected "515                             f"{sorted(want['arms'])}")516        for arm, cell in want["arms"].items():517            got_rows = by_arm.get(arm, [])518            if len(got_rows) != cell["rows"]:519                shape_bad.append(f"{fname}/{arm}: {len(got_rows)} rows, expected "520                                 f"{cell['rows']}")521                continue522            got_pass = sum(1 for r in got_rows if scored_pass(r))523            if got_pass != cell["passes"]:524                pass_bad.append(f"{fname}/{arm}: {got_pass}/{len(got_rows)} passes, "525                                f"expected {cell['passes']}")526            got_trunc = sum(1 for r in got_rows if r.get("truncated") is True)527            if got_trunc != cell["truncated"]:528                trunc_bad.append(f"{fname}/{arm}: {got_trunc} truncated, expected "529                                 f"{cell['truncated']}")530 531        # One item set, once per arm. A duplicate key inside an arm and a key present532        # in one arm but not another are different faults, so both are named.533        reference = None534        for arm in sorted(want["arms"]):535            keys = [r.get(key) for r in by_arm.get(arm, [])]536            unique = set(keys)537            if len(unique) != len(keys):538                pair_bad.append(f"{fname}/{arm}: {len(keys) - len(unique)} duplicate "539                                f"{key} value(s)")540            if len(unique) != want["items"]:541                pair_bad.append(f"{fname}/{arm}: {len(unique)} distinct {key}, "542                                f"expected {want['items']}")543            if reference is None:544                reference = unique545            elif unique != reference:546                pair_bad.append(f"{fname}/{arm}: its {key} set differs from the first "547                                f"arm's by {len(unique ^ reference)} value(s)")548 549    for tag, label, problems in (550            (9, "row counts and arm ids", shape_bad),551            (10, "headline counts under the pre-registered scoring rule", pass_bad),552            (11, "per-arm truncation counts", trunc_bad),553            (12, "item pairing across arms", pair_bad)):554        if problems:555            failures.append(f"[{tag}] {label} do not reproduce: {problems[:3]}")556        elif loaded:557            print(f"  [{tag}] {label} reproduce for all {len(loaded)} v1.1 row files")558 559    # ---- check 13 ----------------------------------------------------------------560    corpus_path = args.public / CORPUS_NAME561    if not corpus_path.is_file():562        failures.append(f"[13] {CORPUS_NAME} is missing from the public release")563    else:564        corpus = json.loads(corpus_path.read_text(encoding="utf-8"))565        entries = corpus.get("row_content_hashes", [])566        declared = corpus.get("counts", {}).get("final")567        malformed = sum(1 for e in entries568                        if not re.fullmatch(r"[0-9a-f]{64}", str(e.get("sha256", ""))))569        if len(entries) != EXPECTED_CORPUS_ROWS:570            failures.append(f"[13] {CORPUS_NAME} carries {len(entries)} row hashes, "571                            f"expected {EXPECTED_CORPUS_ROWS}")572        elif declared != EXPECTED_CORPUS_ROWS:573            failures.append(f"[13] {CORPUS_NAME} declares counts.final {declared}, "574                            f"expected {EXPECTED_CORPUS_ROWS}")575        elif malformed:576            failures.append(f"[13] {CORPUS_NAME}: {malformed} entr(ies) do not carry a "577                            f"64-hex sha256")578        else:579            print(f"  [13] {CORPUS_NAME} carries {len(entries)} row hashes and agrees "580                  f"with its own counts.final")581 582    # ---- diagnostic --------------------------------------------------------------583    if args.sealed_rows and args.sealed_rows.is_file():584        sealed = {r["ea_name"]: r for r in read_jsonl(args.sealed_rows)}585        spec_ok = spec_bad = tmpl_ok = tmpl_bad = 0586        example = None587        for r in published:588            s = sealed.get(id_to_name.get(r["item_id"], ""))589            if not s:590                continue591            if sha256_text(r["prompt"]) == s.get("spec_sha256"):592                spec_ok += 1593            else:594                spec_bad += 1595                if example is None:596                    # Published rows carry item_id only; the name comes from the597                    # private mapping already resolved above.598                    example = (r["item_id"], id_to_name.get(r["item_id"], "?"),599                               sha256_text(r["prompt"]), s.get("spec_sha256"))600            if sha256_text(format_local_prompt(r["prompt"])) == s.get("prompt_sha256"):601                tmpl_ok += 1602            else:603                tmpl_bad += 1604        notes.append(605            f"[D] sealed spec_sha256 reproduces for {spec_ok}/{spec_ok + spec_bad} items; "606            f"sealed prompt_sha256 reproduces under the local serving template for "607            f"{tmpl_ok}/{tmpl_ok + tmpl_bad}"608        )609        if example:610            notes.append(f"[D] first divergence: {example[0]} ({example[1]})\n"611                         f"      sha256(published prompt) = {example[2]}\n"612                         f"      sealed spec_sha256       = {example[3]}")613        # A known-answer check on the hashing itself, so a zero above cannot be blamed614        # on this script's encoding.615        want = "820208ddccc935fe44222db63416e192f605183e1b77c8e833de0d8aa94c6377"616        got_sp = sha256_text(SYSTEM_PROMPT)617        notes.append(f"[D] instrument check: sha256(SYSTEM_PROMPT) "618                     f"{'MATCHES' if got_sp == want else 'DIFFERS FROM'} the sealed "619                     f"system_prompt_sha256")620        if got_sp != want:621            failures.append("[D] this script's hashing does not reproduce a known answer; "622                            "every hash result above is untrustworthy")623 624    for n in notes:625        print(n)626 627    if failures:628        print("\n=== FAILED ===")629        for f in failures:630            print(f"  {f}")631        return 2632    print("\n=== ALL CHECKS PASSED ===")633    return 0634 635 636if __name__ == "__main__":637    sys.exit(main(sys.argv[1:]))638