rootxhacker/patchbench
PatchBench A multi-language benchmark for evaluating whether an LLM can fix a bug correctly and without introducing new bugs. Each row presents a real issue report plus a large code context (~500–1500 lines); the model under test produces a patch; a per-row test suite then checks that (a) previously failing tests now pass and (b) the rest of the suite stays green. Status: v0 spec. Rows are being built. Row schema Field Type In tasks config Description… See the full description on the dataset page: https://huggingface.co/datasets/rootxhacker/patchbench.
0284
1"""PatchBench row verification + context expansion harness (v1.3).2 3For each candidate row:4 1. clone the repo at base_commit5 2. apply test_patch -> run rebuild_cmds -> run test_cmds (twice) -> parse per-test results6 verify: every FAIL_TO_PASS test FAILS pre-fix in both runs (a test that does not7 exist because its package fails to build counts as failing -- upstream semantics)8 3. apply gold_patch -> rebuild -> run test_cmds (twice) -> parse9 verify: every FAIL_TO_PASS PASSES and every PASS_TO_PASS still PASSES in both runs10 4. slice a 500-1500 line context window around the buggy files11 5. write a result record; verified rows additionally carry the full patchbench schema12 13P2P scoping: PASS_TO_PASS is evaluated over tests that PASS in BOTH pre-fix runs in14this environment. Tests upstream lists in P2P that fail pre-fix here (e.g. tests15needing docker/k8s) are recorded in p2p_env_dropped and excluded from grading --16they are not evaluable natively, not regressions.17 18Native (non-Docker) reproduction of the upstream env:19 - toolchain dirs are prepended to PATH per language20 - the upstream /testbed convention is rewritten to the local checkout path21 - after test_cmds, the harness collects the report files the frameworks wrote22 (surefire/gradle XML, gradle HTML, reports/* go-json dumps) and feeds them to23 the parser alongside console output; the row's print_cmds list is often empty24 - JS/TS: if package.json exists without node_modules, deps are installed with the25 repo's own package manager (detected from lockfile)26 - rust: cargo target dirs are wiped between rows AND all builds go to one shared27 CARGO_TARGET_DIR outside the repo, so disk usage is monitorable and reclaimable28 29v0.7: TimeoutExpired.stdout/stderr can be bytes (python quirk under capture_output) --30decode before concatenating. This crashed slow maven builds at the 2100s timeout.31v0.8: disk guard. Three rust jobs died on the 50G ephemeral-storage quota mid-row.32v0.9: java report files (surefire/gradle XML + gradle HTML) are collected and fed33to the parser; disk guard is du-based (container quota) and rust builds run under34a du watchdog so a runaway build kills the row, not the pod.35v1.0: used_bytes() timed out on 40G+ du scans (empty output parsed as 0 => guard blind); limit lowered to 32GiB, watchdog polls 20s and reclaims on kill; results pushed to the Hub incrementally after each row so evictions lose at most one row.36 37v1.1: one dev-profile cargo build grew the shared target dir 6.7G->43G and evicted the pod before the watchdog's du scan could catch it. Fixed both causes: cargo now strips debuginfo and disables incremental compilation in dev/test profiles (no effect on test outcomes, cuts target-dir size several-fold); the in-run watchdog monitors only /tmp/pb-cargo-target (fast du scans even under pressure) at a 20GiB limit instead of scanning all of /tmp at 32GiB.38 39v1.2: some rows' test_cmds use a nested repo/ layout (cd repo, --manifest-path40repo/Cargo.toml) from the upstream /testbed/repo convention -- a self-symlink now41makes both layouts resolve. Context expansion skips test files referenced by the42test_patch that don't exist on disk instead of crashing the row (eza 0070).43 44v1.3: the v1.2 repo-layout symlink was created in verify_row BEFORE clean_state, and45clean_state runs git clean -fdxq which deleted it -- noir rows ran with no repo/46self-link and parsed 0 tests. The symlink is now created inside run_tests, after47every clean_state, so both layouts resolve in every phase.48 49Usage (inside a job with the right language toolchain image):50 python verify.py --lang go --file pilot/candidates/go.jsonl --limit 651"""52import argparse, json, os, shlex, subprocess, sys, time, traceback53 54REPO_ID = "rootxhacker/patchbench"55CACHE = "/tmp/cache/repos"56CMD_TIMEOUT_S = 210057GIT_TIMEOUT_S = 90058MAX_CONTEXT_LINES = 150059# The 50G limit is a per-container ephemeral-storage quota. df/statvfs report the60# host filesystem and never see it, so disk accounting must be du-based.61CONTAINER_QUOTA_BYTES = 50 * 1024**362DU_LIMIT_BYTES = 32 * 1024**363RUST_SHARED_TARGET = "/tmp/pb-cargo-target"64RUST_WATCHDOG_LIMIT_BYTES = 20 * 1024**3 # watchdog limit on the target dir alone65 66LANG_PATH_PREPEND = {67 "go": ["/usr/local/go/bin"],68 "rust": ["/usr/local/cargo/bin", "/usr/local/rustup/bin"],69 "java": [],70 "javascript": [],71 "typescript": [],72}73EXTS = {"go": [".go"], "rust": [".rs"], "java": [".java"],74 "javascript": [".js", ".mjs"], "typescript": [".ts", ".tsx"]}75SKIP_DIRS = {".git", "node_modules", "vendor", "target", "dist", "build", "__pycache__", ".idea"}76 77TESTBED = "/testbed"78 79 80def run(cmd, cwd, timeout):81 t0 = time.time()82 try:83 p = subprocess.run(["bash", "-c", cmd], cwd=cwd, capture_output=True,84 timeout=timeout, text=True, errors="replace")85 return p.returncode, (p.stdout or "") + (p.stderr or ""), time.time() - t086 except subprocess.TimeoutExpired as e:87 so, se = e.stdout or "", e.stderr or ""88 if isinstance(so, bytes):89 so = so.decode("utf-8", "replace")90 if isinstance(se, bytes):91 se = se.decode("utf-8", "replace")92 return 124, so + se + f"\n[TIMEOUT after {timeout}s]", time.time() - t093 94 95def sh(cmds, cwd, timeout=CMD_TIMEOUT_S):96 for c in (cmds or []):97 rc, out, _ = run(c, cwd, timeout)98 if rc != 0:99 return False, out[-4000:]100 return True, ""101 102 103def used_bytes():104 """Bytes used in the du-monitored dirs. df/statvfs see the host fs, not the105 per-container ephemeral-storage quota, so du is the only honest measure."""106 _, out, _ = run("du -sb /tmp /usr/local/cargo /root/.cache /var/tmp 2>/dev/null "107 "| awk '{s+=$1} END {print s+0}'", "/tmp", 900)108 for line in reversed(out.splitlines()):109 s = line.strip()110 if s.isdigit():111 return int(s)112 return 0113 114 115def log_disk(tag):116 _, out, _ = run("du -sb /tmp/pb-cargo-target /tmp/cache /usr/local/cargo/registry "117 "2>/dev/null", "/tmp", 120)118 print(f"[disk:{tag}] container_used={used_bytes() >> 30}G | {out.strip()[:400]}",119 flush=True)120 121 122def reclaim_disk():123 """Wipe build artifacts (not .git clones, not the registry) to free space."""124 run("rm -rf /tmp/pb-cargo-target", "/tmp", 600)125 126 127def disk_guard(tag):128 """Raise instead of blowing the container's ephemeral-storage quota -- an129 eviction loses all remaining rows."""130 u = used_bytes()131 if u > DU_LIMIT_BYTES:132 print(f"[disk:{tag}] used={u >> 30}G > limit={DU_LIMIT_BYTES >> 30}G, "133 "reclaiming shared target dir", flush=True)134 reclaim_disk()135 u = used_bytes()136 if u > CONTAINER_QUOTA_BYTES - 2 * 1024**3:137 raise RuntimeError(f"disk_pressure: used={u >> 30}G after reclaim "138 f"(tag={tag}); aborting row to save the remaining rows")139 140 141def apply_patch(text, cwd):142 pf = os.path.join(cwd, ".pb_patch")143 with open(pf, "w") as f:144 f.write(text)145 rc, out, _ = run(f"git apply --whitespace=nowarn < {pf}", cwd, 120)146 if rc == 0:147 return None148 rc2, out2, _ = run(f"patch -p1 --forward --no-backup-if-mismatch < {pf}", cwd, 120)149 if rc2 == 0:150 return None151 return (out + out2)[-2000:]152 153 154def clean_state(repo_dir, lang=None):155 run("git checkout -f HEAD -- . && git clean -fdxq", repo_dir, 120)156 if lang == "rust":157 run("find . -name target -type d -prune -exec rm -rf {} +", repo_dir, 300)158 run(f"rm -rf {RUST_SHARED_TARGET}", "/tmp", 300)159 160 161def get_parser(row):162 ns = {}163 src = row.get("log_parser") or ""164 assert "def parser" in src, f"no log_parser for {row['id']}"165 exec(src, ns)166 return ns["parser"]167 168 169def ensure_js_deps(repo_dir):170 """Upstream JS/TS images ship node_modules; natively we must install them.171 Detect the package manager from the lockfile. Returns error string or None."""172 if not os.path.exists(os.path.join(repo_dir, "package.json")):173 return None174 if os.path.exists(os.path.join(repo_dir, "node_modules")):175 return None176 if os.path.exists(os.path.join(repo_dir, "pnpm-lock.yaml")):177 cmds = ["corepack enable 2>/dev/null; corepack prepare pnpm@latest --activate 2>/dev/null || npm install -g pnpm",178 "pnpm install --frozen-lockfile || pnpm install"]179 elif os.path.exists(os.path.join(repo_dir, "yarn.lock")):180 cmds = ["yarn install --frozen-lockfile || yarn install"]181 elif os.path.exists(os.path.join(repo_dir, "package-lock.json")):182 cmds = ["npm ci || npm install"]183 else:184 cmds = ["npm install"]185 ok, err = sh(cmds, repo_dir, 1800)186 return None if ok else f"js deps install failed: {err}"187 188 189DEFAULT_REPORT_FINDS = [r'for f in reports/*; do [ -f "$f" ] && cat "$f"; done']190REPORT_FINDS = {191 "java": DEFAULT_REPORT_FINDS + [192 r'find . \( -path "*surefire-reports*" -o -path "*failsafe-reports*" '193 r'-o -path "*build/test-results*" \) -name "*.xml" -type f '194 r'| sort | head -400 | while read -r f; do cat "$f"; done',195 r'find . -path "*reports/tests*" -name "*.html" -type f '196 r'| sort | head -200 | while read -r f; do cat "$f"; done',197 ],198}199 200 201def collect_reports(repo_dir, lang):202 """Cat the report files the test frameworks wrote (per-language find commands)203 into the output stream the log parser sees. Each output capped at 8M chars."""204 outs = []205 for cmd in REPORT_FINDS.get(lang, DEFAULT_REPORT_FINDS):206 _, o, _ = run(cmd, repo_dir, 300)207 if o.strip():208 outs.append(o[:8_000_000])209 return "\n".join(outs)210 211 212WATCHDOG_LANGS = {"rust"}213 214 215def watchdog_wrap(cmd):216 """Run cmd in its own process group under a du watchdog. setsid is critical:217 without it cargo survives `kill -9 -- -$pid` and keeps writing to the quota.218 Build runs in FOREGROUND; the du monitor polls in BACKGROUND; `wait $pid`219 decides the exit code (avoids the kill -0 zombie-hang on fast commands)."""220 limit = DU_LIMIT_BYTES221 q = shlex.quote(cmd)222 return (223 f"printf %s {q} > /tmp/pb_cmd.sh\n"224 "setsid bash /tmp/pb_cmd.sh & pid=$!\n"225 "( while kill -0 $pid 2>/dev/null; do\n"226 " used=$(du -sb /tmp/pb-cargo-target 2>/dev/null "227 "| awk '{s+=$1} END {print s+0}')\n"228 f" if [ \"$used\" -gt {RUST_WATCHDOG_LIMIT_BYTES} ]; then\n"229 " kill -9 -- -$pid 2>/dev/null\n"230 " echo \"[DISK_WATCHDOG] killed build: container=${used}B\"\n"231 " rm -rf /tmp/pb-cargo-target 2>/dev/null\n"232 " exit 0\n"233 " fi\n"234 " sleep 20 >/dev/null 2>&1\n"235 " done ) & mon=$!\n"236 "wait $pid; rc=$?\n"237 "kill $mon 2>/dev/null\n"238 "exit $rc\n"239 )240 241 242def run_tests(row, repo_dir, lang, phase_tag=""):243 """Run rebuild_cmds, test_cmds, print_cmds, then collect the report files the244 test frameworks wrote (surefire/gradle XML, gradle HTML, reports/*) and feed245 them to the log parser. Returns (tests_dict, combined_output_tail)."""246 # v1.3: rows whose cmds reference a nested repo/ layout need a self-symlink,247 # and it must exist AFTER every clean_state (which git-cleans it away).248 _cmds = (row.get("test_cmds") or []) + (row.get("rebuild_cmds") or []) + (row.get("print_cmds") or [])249 if any(("repo/" in c) or c.rstrip().endswith("/repo") for c in _cmds):250 _sl = os.path.join(repo_dir, "repo")251 if not (os.path.islink(_sl) or os.path.exists(_sl)):252 os.symlink(repo_dir, _sl)253 print("repo-layout symlink created", flush=True)254 disk_guard(f"{phase_tag}-pre")255 parser = get_parser(row)256 dep_err = ensure_js_deps(repo_dir)257 wrap = watchdog_wrap if lang in WATCHDOG_LANGS else (lambda c: c)258 for c in (row.get("rebuild_cmds") or []):259 run(wrap(c), repo_dir, CMD_TIMEOUT_S)260 run("mkdir -p reports", repo_dir, 30)261 out = ""262 if dep_err:263 out += dep_err + "\n"264 for c in (row.get("test_cmds") or []):265 _, o, _ = run(wrap(c), repo_dir, CMD_TIMEOUT_S)266 out += o + "\n"267 for c in (row.get("print_cmds") or []):268 _, o, _ = run(c, repo_dir, 120)269 out += o + "\n"270 out += collect_reports(repo_dir, lang) + "\n"271 disk_guard(f"{phase_tag}-post")272 log_disk(f"{phase_tag}-post")273 return parser(out), out[-3000:]274 275 276def grade(tests, want, expect, allow_missing=False):277 """allow_missing=True for pre-fix 'fail' checks: a test that never ran (build278 failure in its package) is not passing, which is what upstream means by fail."""279 missing, wrong = [], []280 for t in want:281 s = tests.get(t)282 if s is None:283 if not allow_missing:284 missing.append(t)285 elif s != expect:286 wrong.append((t, s))287 return (not missing and not wrong), {"missing": len(missing), "wrong": wrong[:10],288 "missing_ids": missing[:5], "parsed_count": len(tests)}289 290 291def touched_paths(patch):292 paths = set()293 for line in patch.splitlines():294 if line.startswith("+++ b/"):295 paths.add(line[6:].split("\t")[0])296 elif line.startswith("--- a/"):297 paths.add(line[6:].split("\t")[0])298 return {p for p in paths if p != "/dev/null"}299 300 301def collect_context(repo_dir, buggy_paths, test_paths, lang, cap=MAX_CONTEXT_LINES):302 exts = tuple(EXTS[lang])303 prio = {"buggy": [], "same_dir": [], "tests": [], "other": []}304 buggy_dirs = {os.path.dirname(p) for p in buggy_paths}305 for root, dirs, files in os.walk(repo_dir):306 dirs[:] = [d for d in dirs if d not in SKIP_DIRS]307 for fn in files:308 if not fn.endswith(exts):309 continue310 p = os.path.relpath(os.path.join(root, fn), repo_dir)311 try:312 n = sum(1 for _ in open(os.path.join(repo_dir, p), errors="replace"))313 except OSError:314 continue315 if n > 600:316 continue317 if p in buggy_paths:318 prio["buggy"].append(p)319 elif p in test_paths:320 prio["tests"].append(p)321 elif os.path.dirname(p) in buggy_dirs:322 prio["same_dir"].append(p)323 else:324 prio["other"].append(p)325 chosen, total = [], 0326 for group in ["buggy", "same_dir", "tests", "other"]:327 for p in sorted(prio[group]):328 with open(os.path.join(repo_dir, p), errors="replace") as f:329 content = f.read()330 n = content.count("\n") + 1331 if total + n > cap:332 continue333 chosen.append({"path": p, "content": content})334 total += n335 if total > cap * 0.9:336 break337 return chosen, total338 339 340def verify_row(row, lang):341 t0 = time.time()342 rec = {"id": row["id"], "repo": row["repo"], "base_commit": row["base_commit"],343 "status": "env_failed", "verify_status": None, "error": None, "timing": {}}344 repo_dir = os.path.join(CACHE, row["repo"].replace("/", "__"))345 sha = row["base_commit"]346 if not os.path.isdir(repo_dir):347 os.makedirs(CACHE, exist_ok=True)348 rc, out, _ = run(f"git clone --filter=blob:none https://github.com/{row['repo']}.git {repo_dir}", CACHE, GIT_TIMEOUT_S)349 if rc != 0:350 rec["error"] = f"clone failed: {out[-500:]}"351 return rec352 for cmd in [f"git fetch origin {sha} --quiet", f"git checkout -f {sha} --quiet"]:353 rc, out, _ = run(cmd, repo_dir, GIT_TIMEOUT_S)354 if rc != 0:355 rec["error"] = f"checkout failed: {out[-500:]}"356 return rec357 log_disk(f"cloned:{row['id']}")358 359 if lang == "rust":360 os.environ["CARGO_TARGET_DIR"] = RUST_SHARED_TARGET361 # v1.1: a dev-profile build with debuginfo hit 43G on the 50G container362 # quota. Stripping debuginfo and disabling incremental compilation has363 # no effect on test outcomes, only on build artifacts.364 os.environ["CARGO_PROFILE_DEV_DEBUG"] = "0"365 os.environ["CARGO_PROFILE_TEST_DEBUG"] = "0"366 os.environ["CARGO_INCREMENTAL"] = "0"367 os.environ["PATH"] = ":".join(LANG_PATH_PREPEND.get(lang, [])) + ":" + os.environ["PATH"]368 def tb(cmds):369 return [c.replace(TESTBED, repo_dir) for c in (cmds or [])]370 row = {**row, "test_cmds": tb(row["test_cmds"]), "rebuild_cmds": tb(row.get("rebuild_cmds")),371 "print_cmds": tb(row.get("print_cmds"))}372 rec["test_cmds"], rec["rebuild_cmds"], rec["print_cmds"] = row["test_cmds"], row["rebuild_cmds"], row["print_cmds"]373 rec["timing"]["setup"] = round(time.time() - t0, 1)374 375 # ---- PRE-FIX (twice) ----376 clean_state(repo_dir, lang)377 err = apply_patch(row["test_patch"], repo_dir)378 if err:379 rec["error"] = f"test_patch failed to apply: {err}"380 return rec381 pre_runs = []382 for i in range(2):383 s = time.time()384 tests, tail = run_tests(row, repo_dir, lang, f"pre{i}:{row['id']}")385 pre_runs.append((tests, tail))386 rec["timing"][f"pre_test_{i}"] = round(time.time() - s, 1)387 pre_detail = grade(pre_runs[0][0], row["fail_to_pass"], "fail", allow_missing=True)[1]388 f2p_ok_pre = all(grade(t, row["fail_to_pass"], "fail", allow_missing=True)[0] for t, _ in pre_runs)389 rec["f2p_pre_ok"], rec["pre_detail"] = f2p_ok_pre, pre_detail390 391 # P2P scoped to this environment: keep only tests that pass in BOTH pre-fix runs392 p2p_scoped = [t for t in row["pass_to_pass"]393 if pre_runs[0][0].get(t) == "pass" and pre_runs[1][0].get(t) == "pass"]394 rec["p2p_env_dropped_count"] = len(row["pass_to_pass"]) - len(p2p_scoped)395 396 # ---- POST-FIX (twice) ----397 clean_state(repo_dir, lang)398 err = apply_patch(row["test_patch"], repo_dir)399 if err:400 rec["error"] = f"test_patch re-apply failed: {err}"401 return rec402 err = apply_patch(row["gold_patch"], repo_dir)403 if err:404 rec["error"] = f"gold_patch failed to apply: {err}"405 return rec406 post_runs = []407 for i in range(2):408 s = time.time()409 tests, tail = run_tests(row, repo_dir, lang, f"post{i}:{row['id']}")410 post_runs.append((tests, tail))411 rec["timing"][f"post_test_{i}"] = round(time.time() - s, 1)412 f2p_ok_post = all(grade(t, row["fail_to_pass"], "pass")[0] for t, _ in post_runs)413 p2p_ok_post = all(grade(t, p2p_scoped, "pass")[0] for t, _ in post_runs)414 rec["f2p_post_ok"], rec["p2p_post_ok"] = f2p_ok_post, p2p_ok_post415 rec["post_detail"] = grade(post_runs[0][0], row["fail_to_pass"], "pass")[1]416 417 rec["timing"]["total"] = round(time.time() - t0, 1)418 rec["raw_tail"] = pre_runs[0][1][-1500:]419 420 if len(pre_runs[0][0]) == 0:421 rec["status"] = "no_tests_parsed"422 elif not f2p_ok_pre:423 rec["status"] = "f2p_not_failing_prefix"424 elif not f2p_ok_post:425 rec["status"] = "f2p_not_passing_postfix"426 elif not p2p_ok_post:427 rec["status"] = "p2p_regression"428 else:429 rec["status"] = "verified"430 431 if rec["status"] == "verified":432 buggy = touched_paths(row["gold_patch"])433 tfiles = touched_paths(row["test_patch"])434 ctx, nlines = collect_context(repo_dir, buggy, tfiles, lang)435 rec["verify_status"] = "verified"436 rec["context_files"] = ctx437 rec["context_lines"] = nlines438 rec["buggy_files"] = sorted(buggy - tfiles)439 rec["test_files"] = []440 for _p in sorted(tfiles):441 try:442 rec["test_files"].append({"path": _p, "content": open(os.path.join(repo_dir, _p), errors="replace").read()})443 except OSError:444 pass # v1.2: missing files are silently skipped445 rec["eval"] = {"test_cmds": row["test_cmds"], "print_cmds": row.get("print_cmds"),446 "log_parser": row["log_parser"], "rebuild_cmds": row.get("rebuild_cmds")}447 return rec448 449 450def main():451 ap = argparse.ArgumentParser()452 ap.add_argument("--lang", required=True)453 ap.add_argument("--file", required=True)454 ap.add_argument("--limit", type=int, default=6)455 ap.add_argument("--out", default=None)456 args = ap.parse_args()457 458 from huggingface_hub import hf_hub_download, HfApi459 token = os.environ.get("HF_TOKEN")460 cand_path = hf_hub_download(REPO_ID, args.file, repo_type="dataset", token=token)461 rows = [json.loads(l) for l in open(cand_path)]462 api = HfApi(token=token)463 out_path = args.out or f"/tmp/results_{args.lang}.jsonl"464 verified_path = f"/tmp/verified_{args.lang}.jsonl"465 466 done = 0467 with open(out_path, "w") as fo, open(verified_path, "w") as fv:468 for row in rows[:args.limit]:469 try:470 rec = verify_row(row, row["language"])471 except Exception as e:472 rec = {"id": row["id"], "status": "harness_crash", "error": repr(e)[:500],473 "tb": traceback.format_exc()[-1500:]}474 fo.write(json.dumps(rec) + "\n")475 fo.flush()476 if rec.get("verify_status") == "verified":477 fv.write(json.dumps({**row, "result": {k: rec[k] for k in478 ["timing", "context_lines", "buggy_files"]}}) + "\n")479 fv.flush()480 print(f"[{row['id']}] {rec['status']} "481 f"pre_ok={rec.get('f2p_pre_ok')} post_ok={rec.get('f2p_post_ok')} "482 f"p2p_ok={rec.get('p2p_post_ok')} parsed={rec.get('pre_detail', {}).get('parsed_count', '?')} "483 f"err={(rec.get('error') or '')[:200]}", flush=True)484 if row["language"] == "rust":485 run(f"rm -rf {RUST_SHARED_TARGET}", "/tmp", 600)486 done += 1487 try:488 api.upload_file(path_or_fileobj=out_path, path_in_repo=f"pilot/results/{args.lang}.jsonl",489 repo_id=REPO_ID, repo_type="dataset", commit_message=f"pilot {args.lang}: incremental {done} rows")490 if os.path.getsize(verified_path) > 0:491 api.upload_file(path_or_fileobj=verified_path, path_in_repo=f"pilot/verified/{args.lang}.jsonl",492 repo_id=REPO_ID, repo_type="dataset", commit_message=f"pilot {args.lang}: verified {done}")493 except Exception as pe:494 print(f"[push] incremental push failed: {pe}", flush=True)495 496 for path, dest in [(out_path, f"pilot/results/{args.lang}.jsonl"),497 (verified_path, f"pilot/verified/{args.lang}.jsonl")]:498 if os.path.getsize(path) > 0:499 api.upload_file(path_or_fileobj=path, path_in_repo=dest, repo_id=REPO_ID,500 repo_type="dataset", commit_message=f"pilot {args.lang}: {dest.split('/')[-2]}")501 counts = {}502 for l in open(out_path):503 counts[json.loads(l)["status"]] = counts.get(json.loads(l)["status"], 0) + 1504 print(f"VERIFY_DONE lang={args.lang} rows={done} statuses={json.dumps(counts)}", flush=True)505 506 507if __name__ == "__main__":508 main()