CompilingThings/compile-benchmark
CompilingThings Compile Benchmark for MQL5® This release evaluates compile success of generated MQL5 on two private held-out sets. The first is 300 Expert Advisor prompts, run on four arms: the base model, two tuned local models and one frontier API model. The second is 200 non-EA prompts (include files, custom indicators, scripts and services, 50 each), run on the three local arms, plus a stability re-run of one of them. The holdout results are attested, not fully verifiable:… See the full description on the dataset page: https://huggingface.co/datasets/CompilingThings/compile-benchmark.
1131
1#!/usr/bin/env python32"""Verify the public release, and ship with the package.3 4Needs Python 3.9 or later and nothing else - no third-party packages. Invoke it with5whichever launcher your system has: `py`, `python3` or `python`.6 7Most checks here are ones a third party can run on the published files. Checks 1 and 28are not: they need the evaluation pool, the frozen manifest and the id-to-name mapping,9none of which is distributed, and they report SKIPPED without them. Hashes are computed10with the semantics the harness used - sha256 over the UTF-8 encoding of the string, no11normalisation - so a mismatch is a real finding rather than an encoding artefact.12 13Checks:14 1 authority file hashes to the source_sha256 the frozen manifest recorded15 2 every published prompt is byte-identical to the evaluated prompt16 3 both published hashes recompute from the published prompt for every row, under17 each arm's rendering rule (local template for base/tuned; system+user for frontier)18 4 each item maps to exactly three result rows: one base, one tuned, one frontier19 5 the public rows reproduce every published table: 2/184, 170/184, 179/184, the20 base-vs-tuned contingency table and the tuned-vs-frontier paired table21 6 substituting into the published template reproduces prompt_sha256 for every22 local-arm row23 7 every SHA-256 quoted in the card is one the release stands behind24 8 every checksum-manifest entry recomputes from the shipped files25 9 the three v1.1 row files carry the row count, the arm ids and the per-arm row26 counts the card publishes27 10 every published headline count falls out of the rows under the pre-registered28 scoring rule: a pass is verdict_headline true AND truncated false29 11 every published per-arm truncation count falls out of the rows30 12 the arms of a row file cover one and the same item set, once each31 13 corpus_row_hashes.json carries 219,867 row hashes and agrees with its own count32 D diagnostic: do the sealed spec_sha256 / prompt_sha256 reproduce from the prompt?33 34Checks 9 to 13 need nothing but the published files. They run on the holdout rows even35though the holdout prompts are private: the arithmetic of an attested result is still36arithmetic, and it is the only thing about the holdouts a reader can check.37 38Checks 1 and 2 need the evaluation pool, the frozen manifest and the id-to-name mapping.39None of those is distributed, so for a public reader they report SKIPPED and the run40continues. Checks 3 through 8 need nothing but the published files.41 42Check 6 is not redundant with check 3. Check 3 renders the prompt with this file's own43format_local_prompt; check 6 renders it the way a reader does, by substituting into the44template published in serving_template.json. They agree only if the published template is45correct - and it once was not.46 47Exit 0 all checks pass, 2 a check failed, 3 an input was missing or unreadable.48 49Usage:50 python3 verify_public_release.py --public .51 52With the undistributed inputs to hand, checks 1 and 2 also run:53 python3 verify_public_release.py --public <dir> --private <dir> \54 --authority <pool.jsonl> --manifest <frozen manifest.json>55"""56 57from __future__ import annotations58 59import argparse60import hashlib61import json62import re63import sys64from collections import defaultdict65from pathlib import Path66 67MANIFEST_NAME = "SHA256SUMS.txt"68 69SYSTEM_PROMPT = (70 "You are an expert MQL5 programmer. Write the complete MQL5 "71 "Expert Advisor code that implements the given specification exactly."72)73 74EXPECTED = {"n": 184, "base": 2, "tuned": 170, "a": 1, "b": 1, "c": 169, "d": 13}75# The frontier arm and its pairing against the tuned arm, added 2026-09-02.76EXPECTED_FRONTIER = {"n": 184, "frontier": 179,77 "both_pass": 168, "tuned_only": 2,78 "frontier_only": 11, "both_fail": 3}79 80CORPUS_NAME = "corpus_row_hashes.json"81EXPECTED_CORPUS_ROWS = 21986782 83# Every count the card publishes for the v1.1 row files, keyed by file and arm.84# `passes` is the headline figure under the pre-registered rule below; it is NOT a85# count of rows whose verdict_headline is true. On the 220k arm of the non-EA holdout86# the two differ by one - a row compiled despite being truncated - and that single87# item is why the rule is stated here rather than assumed.88EXPECTED_ROW_FILES = {89 "holdout_ea300_results.jsonl": {90 "rows": 1200, "items": 300, "item_key": "item_sha256",91 "arms": {92 "base": {"rows": 300, "passes": 1, "truncated": 3},93 "tuned83k": {"rows": 300, "passes": 281, "truncated": 0},94 "tuned220k": {"rows": 300, "passes": 282, "truncated": 0},95 "frontier": {"rows": 300, "passes": 286, "truncated": 4},96 },97 },98 "holdout_nonea200_results.jsonl": {99 "rows": 800, "items": 200, "item_key": "item_sha256",100 "arms": {101 "base": {"rows": 200, "passes": 56, "truncated": 14},102 "tuned83k": {"rows": 200, "passes": 135, "truncated": 13},103 "tuned220k": {"rows": 200, "passes": 168, "truncated": 12},104 "tuned220k_q8_replay": {"rows": 200, "passes": 169, "truncated": 11},105 },106 },107 "bridge_q8_184_results.jsonl": {108 "rows": 552, "items": 184, "item_key": "item_id",109 "arms": {110 "base": {"rows": 184, "passes": 2, "truncated": 2},111 "tuned83k": {"rows": 184, "passes": 175, "truncated": 0},112 "tuned220k": {"rows": 184, "passes": 169, "truncated": 0},113 },114 },115}116 117 118def scored_pass(row: dict) -> bool:119 """The pre-registered scoring rule, and the whole of it: a generation cut off at120 the token ceiling counts as a fail even if what it produced happened to compile."""121 return row.get("verdict_headline") is True and row.get("truncated") is not True122 123 124def format_api_prompt(system: str, spec: str) -> str:125 """Canonical rendering of the frontier arm's system+user pair. The vendor API126 takes the two as separate fields; the frontier rows' prompt_sha256 hashes this127 single-string rendering of them."""128 return f"{system}\n\n{spec}"129 130 131def sha256_text(text: str) -> str:132 """Byte-for-byte the harness's hash rule: sha256 over the UTF-8 encoding."""133 return hashlib.sha256(text.encode("utf-8")).hexdigest()134 135 136def sha256_file(path: Path) -> str:137 h = hashlib.sha256()138 with path.open("rb") as fh:139 for chunk in iter(lambda: fh.read(1 << 20), b""):140 h.update(chunk)141 return h.hexdigest()142 143 144def read_jsonl(path: Path):145 with path.open("r", encoding="utf-8") as fh:146 for lineno, line in enumerate(fh, 1):147 line = line.strip()148 if not line:149 continue150 try:151 yield json.loads(line)152 except json.JSONDecodeError as exc:153 raise SystemExit(f"{path}:{lineno}: {exc}")154 155 156def format_local_prompt(spec: str) -> str:157 """The serving template, rendered exactly as the harness rendered it."""158 return (159 f"<|system|>{SYSTEM_PROMPT}<|end|>\n"160 f"<|user|>{spec}<|end|>\n"161 f"<|assistant|>"162 )163 164 165def main(argv: list[str]) -> int:166 ap = argparse.ArgumentParser(description=__doc__)167 # Only --public is required. The authority pool, the frozen manifest and the id-to-name168 # mapping are not distributed, so a public reader has none of them; requiring any of169 # them made the script exit before check 3 and the release was not runnable by anyone170 # outside. Checks 3, 4 and 5 need nothing but the published files, and they are the171 # checks that matter to a reader: the hashes recompute, the pairing is one-to-one, and172 # the headline falls out of the rows.173 ap.add_argument("--public", type=Path, required=True)174 ap.add_argument("--private", type=Path,175 help="holds item_id_mapping.jsonl; not distributed, so check 2 skips")176 ap.add_argument("--authority", type=Path,177 help="the evaluation pool; not distributed, so check 1 skips")178 ap.add_argument("--manifest", type=Path,179 help="the frozen manifest; not distributed, so check 1 skips")180 ap.add_argument("--sealed-rows", type=Path,181 help="the sealed base-arm results file (not distributed), "182 "for the reproduction diagnostic")183 args = ap.parse_args(argv)184 185 failures: list[str] = []186 notes: list[str] = []187 188 # ---- check 1 -----------------------------------------------------------------189 # Runs only when the pool and the frozen manifest are both to hand. Neither ships, so190 # for a public reader this reports SKIPPED and the run continues.191 have_authority = args.authority is not None and args.authority.is_file()192 have_manifest = args.manifest is not None and args.manifest.is_file()193 authority: dict[str, str] = {}194 195 if have_authority and have_manifest:196 manifest = json.loads(args.manifest.read_text(encoding="utf-8"))197 actual = sha256_file(args.authority)198 if actual == manifest["source_sha256"]:199 print(f" [1] authority sha256 matches the frozen manifest {actual[:16]}...")200 else:201 failures.append(202 f"[1] authority sha256 {actual} != manifest source_sha256 "203 f"{manifest['source_sha256']}"204 )205 # The authority is the only source of prompt text. Published prompts are checked206 # against it rather than the other way round.207 for rec in read_jsonl(args.authority):208 name = rec.get("ea_name")209 if name is not None:210 authority[name] = rec["prompt"]211 else:212 print(" [1] SKIPPED - needs the evaluation pool and the frozen manifest, "213 "neither of which is distributed")214 215 # Checks 3-8 verify the PUBLIC release, so they read the public prompts file216 # unconditionally. An earlier version fell back between the two, which meant a217 # present --private copy silently substituted for the release copy and the218 # verifier could pass a public file that no longer matched its results.219 public_prompts = args.public / "prompts.jsonl"220 if not public_prompts.is_file():221 print(f"MISSING: prompts.jsonl in {args.public}")222 return 3223 224 published = list(read_jsonl(public_prompts))225 226 # ---- check 2 -----------------------------------------------------------------227 # Needs the id-to-name mapping, which is not distributed. A public reader gets a228 # clear SKIPPED rather than a crash; internally the mapping is present and it runs.229 id_to_name: dict[str, str] = {}230 mapping_path = args.private / "item_id_mapping.jsonl" if args.private else None231 if mapping_path is not None and mapping_path.is_file():232 id_to_name = {r["item_id"]: r["ea_name"] for r in read_jsonl(mapping_path)}233 234 if not id_to_name or not authority:235 print(" [2] SKIPPED - needs the evaluation pool and the id-to-name mapping, "236 "neither of which is distributed")237 else:238 mismatched = [239 r["item_id"] for r in published240 if authority.get(id_to_name.get(r["item_id"], "")) != r["prompt"]241 ]242 if mismatched:243 failures.append(244 f"[2] {len(mismatched)} published prompt(s) differ from the evaluated "245 f"prompt"246 )247 else:248 print(f" [2] all {len(published)} published prompts are byte-identical to "249 f"the evaluated prompt")250 251 # ---- check 3 -----------------------------------------------------------------252 # The check any reader can run with nothing but the published files.253 public_hash = {r["item_id"]: sha256_text(r["prompt"]) for r in published}254 results_path = args.public / "per_item_results.jsonl"255 if not results_path.is_file():256 # The docstring's exit-3 contract for missing inputs, honoured here too —257 # without this a missing results file was an uncaught traceback.258 print(f"MISSING: per_item_results.jsonl in {args.public}")259 return 3260 rows = list(read_jsonl(results_path))261 prompt_by_id = {p["item_id"]: p["prompt"] for p in published}262 # The published prompt_sha256 per item for the LOCAL arms, for check 6 to compare a263 # reader-built template render to. Frontier rows hash a different rendering (the264 # API's system+user pair) and are checked under that rule in check 3.265 by_id_hash = {r["item_id"]: r.get("prompt_sha256") for r in rows266 if r.get("arm") != "frontier"}267 268 hash_bad = [269 r["item_id"] for r in rows270 if public_hash.get(r["item_id"]) != r.get("prompt_sha256_public")271 ]272 tmpl_bad_rows = []273 for r in rows:274 if r["item_id"] not in prompt_by_id:275 continue276 prompt = prompt_by_id[r["item_id"]]277 want = (format_api_prompt(SYSTEM_PROMPT, prompt)278 if r.get("arm") == "frontier" else format_local_prompt(prompt))279 if sha256_text(want) != r.get("prompt_sha256"):280 tmpl_bad_rows.append(r["item_id"])281 if hash_bad:282 failures.append(f"[3] prompt_sha256_public does not recompute for "283 f"{len(set(hash_bad))} item(s)")284 elif tmpl_bad_rows:285 failures.append(f"[3] prompt_sha256 does not recompute from the published "286 f"serving template for {len(set(tmpl_bad_rows))} item(s)")287 else:288 print(f" [3] both published hashes recompute from the published prompt for all "289 f"{len(public_hash)} items ({len(rows)} rows)")290 291 # ---- check 4 -----------------------------------------------------------------292 by_item: dict[str, list[dict]] = defaultdict(list)293 for r in rows:294 by_item[r["item_id"]].append(r)295 bad = {296 i: sorted(x["arm"] for x in v)297 for i, v in by_item.items()298 if sorted(x["arm"] for x in v) != ["base", "frontier", "tuned"]299 }300 if bad:301 failures.append(f"[4] {len(bad)} item(s) do not map to exactly one base, one "302 f"tuned and one frontier row: {list(bad.items())[:3]}")303 else:304 print(f" [4] all {len(by_item)} items map to exactly one base, one tuned and "305 f"one frontier row")306 307 # ---- check 5 -----------------------------------------------------------------308 # The headline is a claim about verdict_headline, so it is measured on that field.309 # The two verdict fields can and do diverge - warnings fail strict but not headline,310 # and some frontier rows show exactly that - so the divergence is counted and printed311 # below rather than assumed away, and the tables are measured on verdict_headline312 # only. An earlier version read verdict_strict and got away with it only while the313 # two happened to agree.314 disagree = [315 i for i, v in by_item.items()316 if any(x["verdict_headline"] != x["verdict_strict"] for x in v)317 ]318 if disagree:319 notes.append(f"[5] NOTE: verdict_headline and verdict_strict disagree on "320 f"{len(disagree)} item-arm row(s); the headline below is measured on "321 f"verdict_headline")322 else:323 notes.append(f"[5] verdict_headline and verdict_strict agree on all "324 f"{len(rows)} rows")325 326 a = b = c = d = 0327 fp = f_both_pass = f_tuned_only = f_frontier_only = f_both_fail = 0328 for v in by_item.values():329 base = next(x for x in v if x["arm"] == "base")["verdict_headline"]330 tuned = next(x for x in v if x["arm"] == "tuned")["verdict_headline"]331 frontier = next(x for x in v if x["arm"] == "frontier")["verdict_headline"]332 if base and tuned:333 a += 1334 elif base and not tuned:335 b += 1336 elif tuned:337 c += 1338 else:339 d += 1340 if frontier:341 fp += 1342 if tuned and frontier:343 f_both_pass += 1344 elif tuned:345 f_tuned_only += 1346 elif frontier:347 f_frontier_only += 1348 else:349 f_both_fail += 1350 got = {"n": len(by_item), "base": a + b, "tuned": a + c,351 "a": a, "b": b, "c": c, "d": d}352 got_frontier = {"n": len(by_item), "frontier": fp,353 "both_pass": f_both_pass, "tuned_only": f_tuned_only,354 "frontier_only": f_frontier_only, "both_fail": f_both_fail}355 if got == EXPECTED and got_frontier == EXPECTED_FRONTIER:356 print(f" [5] headline reproduces from the public rows: base {got['base']}/"357 f"{got['n']}, tuned {got['tuned']}/{got['n']}, frontier "358 f"{fp}/{got['n']}, a={a} b={b} c={c} d={d}; tuned-vs-frontier "359 f"{f_both_pass}/{f_tuned_only}/{f_frontier_only}/{f_both_fail}")360 else:361 failures.append(f"[5] headline does not reproduce. expected {EXPECTED} + "362 f"{EXPECTED_FRONTIER}, got {got} + {got_frontier}")363 # ---- check 6 -----------------------------------------------------------------364 # The reader's path, and the reason this check exists. Checks 3 renders the prompt365 # with format_local_prompt, a function compiled into this script. A reader has no such366 # function: they parse serving_template.json, substitute the two placeholders, and367 # hash. Those are different routes to the same string only if the published template368 # is right. It once was not - the template field held the two-character sequence369 # backslash-n where a newline belonged, so every reader-side hash disagreed while370 # check 3 passed. A check that shares the harness's own definition cannot see that.371 template_path = args.public / "serving_template.json"372 if not template_path.is_file():373 failures.append("[6] serving_template.json is missing from the public release")374 else:375 tmpl = json.loads(template_path.read_text(encoding="utf-8"))376 sys_prompt = tmpl["system_prompt"]377 if sha256_text(sys_prompt) != tmpl["system_prompt_sha256"]:378 failures.append("[6] the published system prompt does not match its own "379 "published hash")380 rendered_bad = [381 r["item_id"] for r in published382 if sha256_text(383 tmpl["template"].replace("{system_prompt}", sys_prompt)384 .replace("{prompt}", r["prompt"])385 ) != by_id_hash.get(r["item_id"])386 ]387 if rendered_bad:388 failures.append(389 f"[6] substituting into the published template does not reproduce "390 f"prompt_sha256 for {len(set(rendered_bad))} item(s). A reader following "391 f"serving_template.json cannot reproduce our hashes."392 )393 else:394 print(f" [6] substituting into the published template reproduces "395 f"prompt_sha256 for all {len(published)} items")396 397 398 # ---- check 7 -----------------------------------------------------------------399 # Every 64-hex string in README.md must be a hash the release actually stands behind.400 # The card quotes file hashes in prose and in a table, and those are hand-carried from401 # the metadata when the card is written - so they go stale the moment any file is402 # regenerated. That happened: the card published a serving_template.json hash from an403 # earlier build while the manifest, the metadata and the projection report all agreed404 # on a different one. Nothing caught it, because nothing was comparing the two.405 readme = args.public / "README.md"406 manifest_file = args.public / MANIFEST_NAME407 if readme.is_file() and manifest_file.is_file():408 known: set[str] = set()409 for line in manifest_file.read_text(encoding="utf-8").splitlines():410 digest = line.split(" ")[0].strip()411 if len(digest) == 64:412 known.add(digest)413 # Hashes the release legitimately quotes that are not file hashes: model weights,414 # the corpus, the adapter, the evaluation pool, the normaliser source. Those live415 # in publication_metadata.json, so that file is the second authority.416 # serving_template.json is a third authority: it carries the system-prompt hash417 # and the worked example, which are legitimately quoted in the card and appear in418 # neither the manifest nor the metadata.419 authorities = "".join(420 (args.public / n).read_text(encoding="utf-8")421 for n in ("publication_metadata.json", "serving_template.json",422 "PROJECTION_REPORT.json")423 if (args.public / n).is_file()424 )425 stale = sorted({426 h for h in re.findall(r"\b[0-9a-f]{64}\b", readme.read_text(encoding="utf-8"))427 if h not in known and h not in authorities428 })429 if stale:430 failures.append(431 f"[7] README.md quotes {len(stale)} sha256 value(s) that appear in none "432 f"of {MANIFEST_NAME}, publication_metadata.json, serving_template.json "433 f"or PROJECTION_REPORT.json, so the card disagrees with the release: "434 f"{stale[:3]}"435 )436 else:437 print(" [7] every sha256 quoted in the card is one the release stands behind")438 else:439 # A copy missing the card or the manifest must not verify. Without this branch440 # the check was silently skipped and a damaged copy printed ALL CHECKS PASSED.441 absent = [n for n, f in (("README.md", readme), (MANIFEST_NAME, manifest_file))442 if not f.is_file()]443 failures.append(f"[7] cannot run: missing {', '.join(absent)}")444 445 # ---- check 8 -----------------------------------------------------------------446 # Recompute the manifest. Check 7 reads SHA256SUMS.txt to learn which hashes the447 # release stands behind, but never hashes a file - so a modified LICENSE or448 # CITATION.cff passed the verifier untouched while the card said the verifier checks449 # the release package. Third parties do not all have sha256sum, and telling them to450 # run a tool this script could run itself is not verification.451 if manifest_file.is_file():452 listed: dict[str, str] = {}453 for line in manifest_file.read_text(encoding="utf-8").splitlines():454 if not line.strip():455 continue456 digest, _, name = line.partition(" ")457 if len(digest.strip()) == 64 and name.strip():458 listed[name.strip()] = digest.strip()459 460 bad, missing = [], []461 for name, want in sorted(listed.items()):462 f = args.public / name463 if not f.is_file():464 missing.append(name)465 elif sha256_file(f) != want:466 bad.append(name)467 468 present = {p.name for p in args.public.iterdir() if p.is_file()}469 unlisted = sorted(present - set(listed) - {MANIFEST_NAME})470 471 if bad or missing:472 failures.append(473 f"[8] manifest reconciliation FAILED: {len(bad)} file(s) do not match "474 f"their recorded hash {bad[:3]}, {len(missing)} listed file(s) absent "475 f"{missing[:3]}"476 )477 elif unlisted:478 failures.append(479 f"[8] {len(unlisted)} file(s) in the release are not listed in "480 f"{MANIFEST_NAME}: {unlisted[:5]}"481 )482 else:483 print(f" [8] all {len(listed)} manifest entries recompute, and the only "484 f"unlisted file is {MANIFEST_NAME}, which cannot list itself")485 else:486 # Same rule as check 7: a copy with no manifest is unverifiable, not verified.487 failures.append(f"[8] cannot run: {MANIFEST_NAME} is missing")488 489 # ---- checks 9-12 -------------------------------------------------------------490 # The v1.1 row files. These ship, so a missing one is a failure and never a skip:491 # the card's headline tables are computed from them and a reader who cannot find492 # them cannot check a single holdout number.493 loaded: dict[str, list[dict]] = {}494 for fname in EXPECTED_ROW_FILES:495 path = args.public / fname496 if path.is_file():497 loaded[fname] = list(read_jsonl(path))498 else:499 failures.append(f"[9] {fname} is missing from the public release")500 501 shape_bad, pass_bad, trunc_bad, pair_bad = [], [], [], []502 for fname, want in EXPECTED_ROW_FILES.items():503 rows = loaded.get(fname)504 if rows is None:505 continue506 key = want["item_key"]507 by_arm: dict[str, list[dict]] = defaultdict(list)508 for r in rows:509 by_arm[r.get("arm")].append(r)510 511 if len(rows) != want["rows"]:512 shape_bad.append(f"{fname}: {len(rows)} rows, expected {want['rows']}")513 if sorted(by_arm) != sorted(want["arms"]):514 shape_bad.append(f"{fname}: arm ids {sorted(by_arm)}, expected "515 f"{sorted(want['arms'])}")516 for arm, cell in want["arms"].items():517 got_rows = by_arm.get(arm, [])518 if len(got_rows) != cell["rows"]:519 shape_bad.append(f"{fname}/{arm}: {len(got_rows)} rows, expected "520 f"{cell['rows']}")521 continue522 got_pass = sum(1 for r in got_rows if scored_pass(r))523 if got_pass != cell["passes"]:524 pass_bad.append(f"{fname}/{arm}: {got_pass}/{len(got_rows)} passes, "525 f"expected {cell['passes']}")526 got_trunc = sum(1 for r in got_rows if r.get("truncated") is True)527 if got_trunc != cell["truncated"]:528 trunc_bad.append(f"{fname}/{arm}: {got_trunc} truncated, expected "529 f"{cell['truncated']}")530 531 # One item set, once per arm. A duplicate key inside an arm and a key present532 # in one arm but not another are different faults, so both are named.533 reference = None534 for arm in sorted(want["arms"]):535 keys = [r.get(key) for r in by_arm.get(arm, [])]536 unique = set(keys)537 if len(unique) != len(keys):538 pair_bad.append(f"{fname}/{arm}: {len(keys) - len(unique)} duplicate "539 f"{key} value(s)")540 if len(unique) != want["items"]:541 pair_bad.append(f"{fname}/{arm}: {len(unique)} distinct {key}, "542 f"expected {want['items']}")543 if reference is None:544 reference = unique545 elif unique != reference:546 pair_bad.append(f"{fname}/{arm}: its {key} set differs from the first "547 f"arm's by {len(unique ^ reference)} value(s)")548 549 for tag, label, problems in (550 (9, "row counts and arm ids", shape_bad),551 (10, "headline counts under the pre-registered scoring rule", pass_bad),552 (11, "per-arm truncation counts", trunc_bad),553 (12, "item pairing across arms", pair_bad)):554 if problems:555 failures.append(f"[{tag}] {label} do not reproduce: {problems[:3]}")556 elif loaded:557 print(f" [{tag}] {label} reproduce for all {len(loaded)} v1.1 row files")558 559 # ---- check 13 ----------------------------------------------------------------560 corpus_path = args.public / CORPUS_NAME561 if not corpus_path.is_file():562 failures.append(f"[13] {CORPUS_NAME} is missing from the public release")563 else:564 corpus = json.loads(corpus_path.read_text(encoding="utf-8"))565 entries = corpus.get("row_content_hashes", [])566 declared = corpus.get("counts", {}).get("final")567 malformed = sum(1 for e in entries568 if not re.fullmatch(r"[0-9a-f]{64}", str(e.get("sha256", ""))))569 if len(entries) != EXPECTED_CORPUS_ROWS:570 failures.append(f"[13] {CORPUS_NAME} carries {len(entries)} row hashes, "571 f"expected {EXPECTED_CORPUS_ROWS}")572 elif declared != EXPECTED_CORPUS_ROWS:573 failures.append(f"[13] {CORPUS_NAME} declares counts.final {declared}, "574 f"expected {EXPECTED_CORPUS_ROWS}")575 elif malformed:576 failures.append(f"[13] {CORPUS_NAME}: {malformed} entr(ies) do not carry a "577 f"64-hex sha256")578 else:579 print(f" [13] {CORPUS_NAME} carries {len(entries)} row hashes and agrees "580 f"with its own counts.final")581 582 # ---- diagnostic --------------------------------------------------------------583 if args.sealed_rows and args.sealed_rows.is_file():584 sealed = {r["ea_name"]: r for r in read_jsonl(args.sealed_rows)}585 spec_ok = spec_bad = tmpl_ok = tmpl_bad = 0586 example = None587 for r in published:588 s = sealed.get(id_to_name.get(r["item_id"], ""))589 if not s:590 continue591 if sha256_text(r["prompt"]) == s.get("spec_sha256"):592 spec_ok += 1593 else:594 spec_bad += 1595 if example is None:596 # Published rows carry item_id only; the name comes from the597 # private mapping already resolved above.598 example = (r["item_id"], id_to_name.get(r["item_id"], "?"),599 sha256_text(r["prompt"]), s.get("spec_sha256"))600 if sha256_text(format_local_prompt(r["prompt"])) == s.get("prompt_sha256"):601 tmpl_ok += 1602 else:603 tmpl_bad += 1604 notes.append(605 f"[D] sealed spec_sha256 reproduces for {spec_ok}/{spec_ok + spec_bad} items; "606 f"sealed prompt_sha256 reproduces under the local serving template for "607 f"{tmpl_ok}/{tmpl_ok + tmpl_bad}"608 )609 if example:610 notes.append(f"[D] first divergence: {example[0]} ({example[1]})\n"611 f" sha256(published prompt) = {example[2]}\n"612 f" sealed spec_sha256 = {example[3]}")613 # A known-answer check on the hashing itself, so a zero above cannot be blamed614 # on this script's encoding.615 want = "820208ddccc935fe44222db63416e192f605183e1b77c8e833de0d8aa94c6377"616 got_sp = sha256_text(SYSTEM_PROMPT)617 notes.append(f"[D] instrument check: sha256(SYSTEM_PROMPT) "618 f"{'MATCHES' if got_sp == want else 'DIFFERS FROM'} the sealed "619 f"system_prompt_sha256")620 if got_sp != want:621 failures.append("[D] this script's hashing does not reproduce a known answer; "622 "every hash result above is untrustworthy")623 624 for n in notes:625 print(n)626 627 if failures:628 print("\n=== FAILED ===")629 for f in failures:630 print(f" {f}")631 return 2632 print("\n=== ALL CHECKS PASSED ===")633 return 0634 635 636if __name__ == "__main__":637 sys.exit(main(sys.argv[1:]))638 