alexshpunt/explicit-edit-benchmark
Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.
Explicit Edit Benchmark
226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.
Source code and benchmark runner: GitHub — Explicit Edit Benchmark
[Open the interactive Explorer](https://huggingface.co/spaces/alexshpunt/benchmark-explorer) to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.
Leaderboard by model route
Score v2 = coverage × quality, where quality is 75% first exact and 25% final exact. Repeated runs are averaged inside each configuration and task; configurations then have equal weight inside each task, and tasks have equal weight. The same numbers are in views.json, and the Explorer breaks them down by harness, version and reasoning mode.
The harness list behind each row is in data/models.jsonl.gz, and views.json holds the same aggregates for the other groupings: by harness, by agent and by reasoning mode.
Data contributors
Thank you to everyone who has shared benchmark observations. Your work makes this public comparison possible.
Accepted harnesses
Tables
The Dataset Viewer shows every config. dataset-index.json holds the source hashes, contracts, task sets, completeness and counts, and views.json, leaderboard.json and summary.json hold the aggregated rankings.
Source and contribution
Source code, run instructions and the contribution guide live at alexshpunt/explicit-edit-benchmark. Accepted bundles are kept under source/, the shards and summaries are views rebuilt from them, and each accepted harness family has a README badge under badges/.
Observations hold no prompts, arguments, commands, output, sessions, workspaces or credentials.
