Team Ai
Datasetpublic

rootxhacker/patchbench

PatchBench A multi-language benchmark for evaluating whether an LLM can fix a bug correctly and without introducing new bugs. Each row presents a real issue report plus a large code context (~500–1500 lines); the model under test produces a patch; a per-row test suite then checks that (a) previously failing tests now pass and (b) the rest of the suite stays green. Status: v0 spec. Rows are being built. Row schema Field Type In tasks config Description… See the full description on the dataset page: https://huggingface.co/datasets/rootxhacker/patchbench.

sourceHugging Faceupdated 28d agoView on Hugging Face
0likes284downloads
verify.py508 linesDownload Raw Back to scripts
1"""PatchBench row verification + context expansion harness (v1.3).2 3For each candidate row:4  1. clone the repo at base_commit5  2. apply test_patch -> run rebuild_cmds -> run test_cmds (twice) -> parse per-test results6     verify: every FAIL_TO_PASS test FAILS pre-fix in both runs (a test that does not7     exist because its package fails to build counts as failing -- upstream semantics)8  3. apply gold_patch -> rebuild -> run test_cmds (twice) -> parse9     verify: every FAIL_TO_PASS PASSES and every PASS_TO_PASS still PASSES in both runs10  4. slice a 500-1500 line context window around the buggy files11  5. write a result record; verified rows additionally carry the full patchbench schema12 13P2P scoping: PASS_TO_PASS is evaluated over tests that PASS in BOTH pre-fix runs in14this environment. Tests upstream lists in P2P that fail pre-fix here (e.g. tests15needing docker/k8s) are recorded in p2p_env_dropped and excluded from grading --16they are not evaluable natively, not regressions.17 18Native (non-Docker) reproduction of the upstream env:19  - toolchain dirs are prepended to PATH per language20  - the upstream /testbed convention is rewritten to the local checkout path21  - after test_cmds, the harness collects the report files the frameworks wrote22    (surefire/gradle XML, gradle HTML, reports/* go-json dumps) and feeds them to23    the parser alongside console output; the row's print_cmds list is often empty24  - JS/TS: if package.json exists without node_modules, deps are installed with the25    repo's own package manager (detected from lockfile)26  - rust: cargo target dirs are wiped between rows AND all builds go to one shared27    CARGO_TARGET_DIR outside the repo, so disk usage is monitorable and reclaimable28 29v0.7: TimeoutExpired.stdout/stderr can be bytes (python quirk under capture_output) --30decode before concatenating. This crashed slow maven builds at the 2100s timeout.31v0.8: disk guard. Three rust jobs died on the 50G ephemeral-storage quota mid-row.32v0.9: java report files (surefire/gradle XML + gradle HTML) are collected and fed33to the parser; disk guard is du-based (container quota) and rust builds run under34a du watchdog so a runaway build kills the row, not the pod.35v1.0: used_bytes() timed out on 40G+ du scans (empty output parsed as 0 => guard blind); limit lowered to 32GiB, watchdog polls 20s and reclaims on kill; results pushed to the Hub incrementally after each row so evictions lose at most one row.36 37v1.1: one dev-profile cargo build grew the shared target dir 6.7G->43G and evicted the pod before the watchdog's du scan could catch it. Fixed both causes: cargo now strips debuginfo and disables incremental compilation in dev/test profiles (no effect on test outcomes, cuts target-dir size several-fold); the in-run watchdog monitors only /tmp/pb-cargo-target (fast du scans even under pressure) at a 20GiB limit instead of scanning all of /tmp at 32GiB.38 39v1.2: some rows' test_cmds use a nested repo/ layout (cd repo, --manifest-path40repo/Cargo.toml) from the upstream /testbed/repo convention -- a self-symlink now41makes both layouts resolve. Context expansion skips test files referenced by the42test_patch that don't exist on disk instead of crashing the row (eza 0070).43 44v1.3: the v1.2 repo-layout symlink was created in verify_row BEFORE clean_state, and45clean_state runs git clean -fdxq which deleted it -- noir rows ran with no repo/46self-link and parsed 0 tests. The symlink is now created inside run_tests, after47every clean_state, so both layouts resolve in every phase.48 49Usage (inside a job with the right language toolchain image):50  python verify.py --lang go --file pilot/candidates/go.jsonl --limit 651"""52import argparse, json, os, shlex, subprocess, sys, time, traceback53 54REPO_ID = "rootxhacker/patchbench"55CACHE = "/tmp/cache/repos"56CMD_TIMEOUT_S = 210057GIT_TIMEOUT_S = 90058MAX_CONTEXT_LINES = 150059# The 50G limit is a per-container ephemeral-storage quota. df/statvfs report the60# host filesystem and never see it, so disk accounting must be du-based.61CONTAINER_QUOTA_BYTES = 50 * 1024**362DU_LIMIT_BYTES = 32 * 1024**363RUST_SHARED_TARGET = "/tmp/pb-cargo-target"64RUST_WATCHDOG_LIMIT_BYTES = 20 * 1024**3  # watchdog limit on the target dir alone65 66LANG_PATH_PREPEND = {67    "go": ["/usr/local/go/bin"],68    "rust": ["/usr/local/cargo/bin", "/usr/local/rustup/bin"],69    "java": [],70    "javascript": [],71    "typescript": [],72}73EXTS = {"go": [".go"], "rust": [".rs"], "java": [".java"],74        "javascript": [".js", ".mjs"], "typescript": [".ts", ".tsx"]}75SKIP_DIRS = {".git", "node_modules", "vendor", "target", "dist", "build", "__pycache__", ".idea"}76 77TESTBED = "/testbed"78 79 80def run(cmd, cwd, timeout):81    t0 = time.time()82    try:83        p = subprocess.run(["bash", "-c", cmd], cwd=cwd, capture_output=True,84                           timeout=timeout, text=True, errors="replace")85        return p.returncode, (p.stdout or "") + (p.stderr or ""), time.time() - t086    except subprocess.TimeoutExpired as e:87        so, se = e.stdout or "", e.stderr or ""88        if isinstance(so, bytes):89            so = so.decode("utf-8", "replace")90        if isinstance(se, bytes):91            se = se.decode("utf-8", "replace")92        return 124, so + se + f"\n[TIMEOUT after {timeout}s]", time.time() - t093 94 95def sh(cmds, cwd, timeout=CMD_TIMEOUT_S):96    for c in (cmds or []):97        rc, out, _ = run(c, cwd, timeout)98        if rc != 0:99            return False, out[-4000:]100    return True, ""101 102 103def used_bytes():104    """Bytes used in the du-monitored dirs. df/statvfs see the host fs, not the105    per-container ephemeral-storage quota, so du is the only honest measure."""106    _, out, _ = run("du -sb /tmp /usr/local/cargo /root/.cache /var/tmp 2>/dev/null "107                    "| awk '{s+=$1} END {print s+0}'", "/tmp", 900)108    for line in reversed(out.splitlines()):109        s = line.strip()110        if s.isdigit():111            return int(s)112    return 0113 114 115def log_disk(tag):116    _, out, _ = run("du -sb /tmp/pb-cargo-target /tmp/cache /usr/local/cargo/registry "117                    "2>/dev/null", "/tmp", 120)118    print(f"[disk:{tag}] container_used={used_bytes() >> 30}G | {out.strip()[:400]}",119          flush=True)120 121 122def reclaim_disk():123    """Wipe build artifacts (not .git clones, not the registry) to free space."""124    run("rm -rf /tmp/pb-cargo-target", "/tmp", 600)125 126 127def disk_guard(tag):128    """Raise instead of blowing the container's ephemeral-storage quota -- an129    eviction loses all remaining rows."""130    u = used_bytes()131    if u > DU_LIMIT_BYTES:132        print(f"[disk:{tag}] used={u >> 30}G > limit={DU_LIMIT_BYTES >> 30}G, "133              "reclaiming shared target dir", flush=True)134        reclaim_disk()135        u = used_bytes()136    if u > CONTAINER_QUOTA_BYTES - 2 * 1024**3:137        raise RuntimeError(f"disk_pressure: used={u >> 30}G after reclaim "138                           f"(tag={tag}); aborting row to save the remaining rows")139 140 141def apply_patch(text, cwd):142    pf = os.path.join(cwd, ".pb_patch")143    with open(pf, "w") as f:144        f.write(text)145    rc, out, _ = run(f"git apply --whitespace=nowarn < {pf}", cwd, 120)146    if rc == 0:147        return None148    rc2, out2, _ = run(f"patch -p1 --forward --no-backup-if-mismatch < {pf}", cwd, 120)149    if rc2 == 0:150        return None151    return (out + out2)[-2000:]152 153 154def clean_state(repo_dir, lang=None):155    run("git checkout -f HEAD -- . && git clean -fdxq", repo_dir, 120)156    if lang == "rust":157        run("find . -name target -type d -prune -exec rm -rf {} +", repo_dir, 300)158        run(f"rm -rf {RUST_SHARED_TARGET}", "/tmp", 300)159 160 161def get_parser(row):162    ns = {}163    src = row.get("log_parser") or ""164    assert "def parser" in src, f"no log_parser for {row['id']}"165    exec(src, ns)166    return ns["parser"]167 168 169def ensure_js_deps(repo_dir):170    """Upstream JS/TS images ship node_modules; natively we must install them.171    Detect the package manager from the lockfile. Returns error string or None."""172    if not os.path.exists(os.path.join(repo_dir, "package.json")):173        return None174    if os.path.exists(os.path.join(repo_dir, "node_modules")):175        return None176    if os.path.exists(os.path.join(repo_dir, "pnpm-lock.yaml")):177        cmds = ["corepack enable 2>/dev/null; corepack prepare pnpm@latest --activate 2>/dev/null || npm install -g pnpm",178                "pnpm install --frozen-lockfile || pnpm install"]179    elif os.path.exists(os.path.join(repo_dir, "yarn.lock")):180        cmds = ["yarn install --frozen-lockfile || yarn install"]181    elif os.path.exists(os.path.join(repo_dir, "package-lock.json")):182        cmds = ["npm ci || npm install"]183    else:184        cmds = ["npm install"]185    ok, err = sh(cmds, repo_dir, 1800)186    return None if ok else f"js deps install failed: {err}"187 188 189DEFAULT_REPORT_FINDS = [r'for f in reports/*; do [ -f "$f" ] && cat "$f"; done']190REPORT_FINDS = {191    "java": DEFAULT_REPORT_FINDS + [192        r'find . \( -path "*surefire-reports*" -o -path "*failsafe-reports*" '193        r'-o -path "*build/test-results*" \) -name "*.xml" -type f '194        r'| sort | head -400 | while read -r f; do cat "$f"; done',195        r'find . -path "*reports/tests*" -name "*.html" -type f '196        r'| sort | head -200 | while read -r f; do cat "$f"; done',197    ],198}199 200 201def collect_reports(repo_dir, lang):202    """Cat the report files the test frameworks wrote (per-language find commands)203    into the output stream the log parser sees. Each output capped at 8M chars."""204    outs = []205    for cmd in REPORT_FINDS.get(lang, DEFAULT_REPORT_FINDS):206        _, o, _ = run(cmd, repo_dir, 300)207        if o.strip():208            outs.append(o[:8_000_000])209    return "\n".join(outs)210 211 212WATCHDOG_LANGS = {"rust"}213 214 215def watchdog_wrap(cmd):216    """Run cmd in its own process group under a du watchdog. setsid is critical:217    without it cargo survives `kill -9 -- -$pid` and keeps writing to the quota.218    Build runs in FOREGROUND; the du monitor polls in BACKGROUND; `wait $pid`219    decides the exit code (avoids the kill -0 zombie-hang on fast commands)."""220    limit = DU_LIMIT_BYTES221    q = shlex.quote(cmd)222    return (223        f"printf %s {q} > /tmp/pb_cmd.sh\n"224        "setsid bash /tmp/pb_cmd.sh & pid=$!\n"225        "( while kill -0 $pid 2>/dev/null; do\n"226        "    used=$(du -sb /tmp/pb-cargo-target 2>/dev/null "227        "| awk '{s+=$1} END {print s+0}')\n"228        f"    if [ \"$used\" -gt {RUST_WATCHDOG_LIMIT_BYTES} ]; then\n"229        "      kill -9 -- -$pid 2>/dev/null\n"230        "      echo \"[DISK_WATCHDOG] killed build: container=${used}B\"\n"231        "      rm -rf /tmp/pb-cargo-target 2>/dev/null\n"232        "      exit 0\n"233        "    fi\n"234        "    sleep 20 >/dev/null 2>&1\n"235        "  done ) & mon=$!\n"236        "wait $pid; rc=$?\n"237        "kill $mon 2>/dev/null\n"238        "exit $rc\n"239    )240 241 242def run_tests(row, repo_dir, lang, phase_tag=""):243    """Run rebuild_cmds, test_cmds, print_cmds, then collect the report files the244    test frameworks wrote (surefire/gradle XML, gradle HTML, reports/*) and feed245    them to the log parser. Returns (tests_dict, combined_output_tail)."""246    # v1.3: rows whose cmds reference a nested repo/ layout need a self-symlink,247    # and it must exist AFTER every clean_state (which git-cleans it away).248    _cmds = (row.get("test_cmds") or []) + (row.get("rebuild_cmds") or []) + (row.get("print_cmds") or [])249    if any(("repo/" in c) or c.rstrip().endswith("/repo") for c in _cmds):250        _sl = os.path.join(repo_dir, "repo")251        if not (os.path.islink(_sl) or os.path.exists(_sl)):252            os.symlink(repo_dir, _sl)253            print("repo-layout symlink created", flush=True)254    disk_guard(f"{phase_tag}-pre")255    parser = get_parser(row)256    dep_err = ensure_js_deps(repo_dir)257    wrap = watchdog_wrap if lang in WATCHDOG_LANGS else (lambda c: c)258    for c in (row.get("rebuild_cmds") or []):259        run(wrap(c), repo_dir, CMD_TIMEOUT_S)260    run("mkdir -p reports", repo_dir, 30)261    out = ""262    if dep_err:263        out += dep_err + "\n"264    for c in (row.get("test_cmds") or []):265        _, o, _ = run(wrap(c), repo_dir, CMD_TIMEOUT_S)266        out += o + "\n"267    for c in (row.get("print_cmds") or []):268        _, o, _ = run(c, repo_dir, 120)269        out += o + "\n"270    out += collect_reports(repo_dir, lang) + "\n"271    disk_guard(f"{phase_tag}-post")272    log_disk(f"{phase_tag}-post")273    return parser(out), out[-3000:]274 275 276def grade(tests, want, expect, allow_missing=False):277    """allow_missing=True for pre-fix 'fail' checks: a test that never ran (build278    failure in its package) is not passing, which is what upstream means by fail."""279    missing, wrong = [], []280    for t in want:281        s = tests.get(t)282        if s is None:283            if not allow_missing:284                missing.append(t)285        elif s != expect:286            wrong.append((t, s))287    return (not missing and not wrong), {"missing": len(missing), "wrong": wrong[:10],288                                         "missing_ids": missing[:5], "parsed_count": len(tests)}289 290 291def touched_paths(patch):292    paths = set()293    for line in patch.splitlines():294        if line.startswith("+++ b/"):295            paths.add(line[6:].split("\t")[0])296        elif line.startswith("--- a/"):297            paths.add(line[6:].split("\t")[0])298    return {p for p in paths if p != "/dev/null"}299 300 301def collect_context(repo_dir, buggy_paths, test_paths, lang, cap=MAX_CONTEXT_LINES):302    exts = tuple(EXTS[lang])303    prio = {"buggy": [], "same_dir": [], "tests": [], "other": []}304    buggy_dirs = {os.path.dirname(p) for p in buggy_paths}305    for root, dirs, files in os.walk(repo_dir):306        dirs[:] = [d for d in dirs if d not in SKIP_DIRS]307        for fn in files:308            if not fn.endswith(exts):309                continue310            p = os.path.relpath(os.path.join(root, fn), repo_dir)311            try:312                n = sum(1 for _ in open(os.path.join(repo_dir, p), errors="replace"))313            except OSError:314                continue315            if n > 600:316                continue317            if p in buggy_paths:318                prio["buggy"].append(p)319            elif p in test_paths:320                prio["tests"].append(p)321            elif os.path.dirname(p) in buggy_dirs:322                prio["same_dir"].append(p)323            else:324                prio["other"].append(p)325    chosen, total = [], 0326    for group in ["buggy", "same_dir", "tests", "other"]:327        for p in sorted(prio[group]):328            with open(os.path.join(repo_dir, p), errors="replace") as f:329                content = f.read()330            n = content.count("\n") + 1331            if total + n > cap:332                continue333            chosen.append({"path": p, "content": content})334            total += n335        if total > cap * 0.9:336            break337    return chosen, total338 339 340def verify_row(row, lang):341    t0 = time.time()342    rec = {"id": row["id"], "repo": row["repo"], "base_commit": row["base_commit"],343           "status": "env_failed", "verify_status": None, "error": None, "timing": {}}344    repo_dir = os.path.join(CACHE, row["repo"].replace("/", "__"))345    sha = row["base_commit"]346    if not os.path.isdir(repo_dir):347        os.makedirs(CACHE, exist_ok=True)348        rc, out, _ = run(f"git clone --filter=blob:none https://github.com/{row['repo']}.git {repo_dir}", CACHE, GIT_TIMEOUT_S)349        if rc != 0:350            rec["error"] = f"clone failed: {out[-500:]}"351            return rec352    for cmd in [f"git fetch origin {sha} --quiet", f"git checkout -f {sha} --quiet"]:353        rc, out, _ = run(cmd, repo_dir, GIT_TIMEOUT_S)354        if rc != 0:355            rec["error"] = f"checkout failed: {out[-500:]}"356            return rec357    log_disk(f"cloned:{row['id']}")358 359    if lang == "rust":360        os.environ["CARGO_TARGET_DIR"] = RUST_SHARED_TARGET361        # v1.1: a dev-profile build with debuginfo hit 43G on the 50G container362        # quota. Stripping debuginfo and disabling incremental compilation has363        # no effect on test outcomes, only on build artifacts.364        os.environ["CARGO_PROFILE_DEV_DEBUG"] = "0"365        os.environ["CARGO_PROFILE_TEST_DEBUG"] = "0"366        os.environ["CARGO_INCREMENTAL"] = "0"367    os.environ["PATH"] = ":".join(LANG_PATH_PREPEND.get(lang, [])) + ":" + os.environ["PATH"]368    def tb(cmds):369        return [c.replace(TESTBED, repo_dir) for c in (cmds or [])]370    row = {**row, "test_cmds": tb(row["test_cmds"]), "rebuild_cmds": tb(row.get("rebuild_cmds")),371           "print_cmds": tb(row.get("print_cmds"))}372    rec["test_cmds"], rec["rebuild_cmds"], rec["print_cmds"] = row["test_cmds"], row["rebuild_cmds"], row["print_cmds"]373    rec["timing"]["setup"] = round(time.time() - t0, 1)374 375    # ---- PRE-FIX (twice) ----376    clean_state(repo_dir, lang)377    err = apply_patch(row["test_patch"], repo_dir)378    if err:379        rec["error"] = f"test_patch failed to apply: {err}"380        return rec381    pre_runs = []382    for i in range(2):383        s = time.time()384        tests, tail = run_tests(row, repo_dir, lang, f"pre{i}:{row['id']}")385        pre_runs.append((tests, tail))386        rec["timing"][f"pre_test_{i}"] = round(time.time() - s, 1)387    pre_detail = grade(pre_runs[0][0], row["fail_to_pass"], "fail", allow_missing=True)[1]388    f2p_ok_pre = all(grade(t, row["fail_to_pass"], "fail", allow_missing=True)[0] for t, _ in pre_runs)389    rec["f2p_pre_ok"], rec["pre_detail"] = f2p_ok_pre, pre_detail390 391    # P2P scoped to this environment: keep only tests that pass in BOTH pre-fix runs392    p2p_scoped = [t for t in row["pass_to_pass"]393                  if pre_runs[0][0].get(t) == "pass" and pre_runs[1][0].get(t) == "pass"]394    rec["p2p_env_dropped_count"] = len(row["pass_to_pass"]) - len(p2p_scoped)395 396    # ---- POST-FIX (twice) ----397    clean_state(repo_dir, lang)398    err = apply_patch(row["test_patch"], repo_dir)399    if err:400        rec["error"] = f"test_patch re-apply failed: {err}"401        return rec402    err = apply_patch(row["gold_patch"], repo_dir)403    if err:404        rec["error"] = f"gold_patch failed to apply: {err}"405        return rec406    post_runs = []407    for i in range(2):408        s = time.time()409        tests, tail = run_tests(row, repo_dir, lang, f"post{i}:{row['id']}")410        post_runs.append((tests, tail))411        rec["timing"][f"post_test_{i}"] = round(time.time() - s, 1)412    f2p_ok_post = all(grade(t, row["fail_to_pass"], "pass")[0] for t, _ in post_runs)413    p2p_ok_post = all(grade(t, p2p_scoped, "pass")[0] for t, _ in post_runs)414    rec["f2p_post_ok"], rec["p2p_post_ok"] = f2p_ok_post, p2p_ok_post415    rec["post_detail"] = grade(post_runs[0][0], row["fail_to_pass"], "pass")[1]416 417    rec["timing"]["total"] = round(time.time() - t0, 1)418    rec["raw_tail"] = pre_runs[0][1][-1500:]419 420    if len(pre_runs[0][0]) == 0:421        rec["status"] = "no_tests_parsed"422    elif not f2p_ok_pre:423        rec["status"] = "f2p_not_failing_prefix"424    elif not f2p_ok_post:425        rec["status"] = "f2p_not_passing_postfix"426    elif not p2p_ok_post:427        rec["status"] = "p2p_regression"428    else:429        rec["status"] = "verified"430 431    if rec["status"] == "verified":432        buggy = touched_paths(row["gold_patch"])433        tfiles = touched_paths(row["test_patch"])434        ctx, nlines = collect_context(repo_dir, buggy, tfiles, lang)435        rec["verify_status"] = "verified"436        rec["context_files"] = ctx437        rec["context_lines"] = nlines438        rec["buggy_files"] = sorted(buggy - tfiles)439        rec["test_files"] = []440        for _p in sorted(tfiles):441            try:442                rec["test_files"].append({"path": _p, "content": open(os.path.join(repo_dir, _p), errors="replace").read()})443            except OSError:444                pass  # v1.2: missing files are silently skipped445        rec["eval"] = {"test_cmds": row["test_cmds"], "print_cmds": row.get("print_cmds"),446                       "log_parser": row["log_parser"], "rebuild_cmds": row.get("rebuild_cmds")}447    return rec448 449 450def main():451    ap = argparse.ArgumentParser()452    ap.add_argument("--lang", required=True)453    ap.add_argument("--file", required=True)454    ap.add_argument("--limit", type=int, default=6)455    ap.add_argument("--out", default=None)456    args = ap.parse_args()457 458    from huggingface_hub import hf_hub_download, HfApi459    token = os.environ.get("HF_TOKEN")460    cand_path = hf_hub_download(REPO_ID, args.file, repo_type="dataset", token=token)461    rows = [json.loads(l) for l in open(cand_path)]462    api = HfApi(token=token)463    out_path = args.out or f"/tmp/results_{args.lang}.jsonl"464    verified_path = f"/tmp/verified_{args.lang}.jsonl"465 466    done = 0467    with open(out_path, "w") as fo, open(verified_path, "w") as fv:468        for row in rows[:args.limit]:469            try:470                rec = verify_row(row, row["language"])471            except Exception as e:472                rec = {"id": row["id"], "status": "harness_crash", "error": repr(e)[:500],473                       "tb": traceback.format_exc()[-1500:]}474            fo.write(json.dumps(rec) + "\n")475            fo.flush()476            if rec.get("verify_status") == "verified":477                fv.write(json.dumps({**row, "result": {k: rec[k] for k in478                         ["timing", "context_lines", "buggy_files"]}}) + "\n")479                fv.flush()480            print(f"[{row['id']}] {rec['status']} "481                  f"pre_ok={rec.get('f2p_pre_ok')} post_ok={rec.get('f2p_post_ok')} "482                  f"p2p_ok={rec.get('p2p_post_ok')} parsed={rec.get('pre_detail', {}).get('parsed_count', '?')} "483                  f"err={(rec.get('error') or '')[:200]}", flush=True)484            if row["language"] == "rust":485                run(f"rm -rf {RUST_SHARED_TARGET}", "/tmp", 600)486            done += 1487            try:488                api.upload_file(path_or_fileobj=out_path, path_in_repo=f"pilot/results/{args.lang}.jsonl",489                                repo_id=REPO_ID, repo_type="dataset", commit_message=f"pilot {args.lang}: incremental {done} rows")490                if os.path.getsize(verified_path) > 0:491                    api.upload_file(path_or_fileobj=verified_path, path_in_repo=f"pilot/verified/{args.lang}.jsonl",492                                    repo_id=REPO_ID, repo_type="dataset", commit_message=f"pilot {args.lang}: verified {done}")493            except Exception as pe:494                print(f"[push] incremental push failed: {pe}", flush=True)495 496    for path, dest in [(out_path, f"pilot/results/{args.lang}.jsonl"),497                       (verified_path, f"pilot/verified/{args.lang}.jsonl")]:498        if os.path.getsize(path) > 0:499            api.upload_file(path_or_fileobj=path, path_in_repo=dest, repo_id=REPO_ID,500                            repo_type="dataset", commit_message=f"pilot {args.lang}: {dest.split('/')[-2]}")501    counts = {}502    for l in open(out_path):503        counts[json.loads(l)["status"]] = counts.get(json.loads(l)["status"], 0) + 1504    print(f"VERIFY_DONE lang={args.lang} rows={done} statuses={json.dumps(counts)}", flush=True)505 506 507if __name__ == "__main__":508    main()