Team Ai
Datasetpublic

alexshpunt/explicit-edit-benchmark

Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.

sourceHugging Facecc-by-4.0updated 1d agoView on Hugging Face
5likes19kdownloads
Dataset Card

Explicit Edit Benchmark

226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.

Source code and benchmark runner: GitHub — Explicit Edit Benchmark

[Open the interactive Explorer](https://huggingface.co/spaces/alexshpunt/benchmark-explorer) to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.

Leaderboard by model route

Score v2 = coverage × quality, where quality is 75% first exact and 25% final exact. Repeated runs are averaged inside each configuration and task; configurations then have equal weight inside each task, and tasks have equal weight. The same numbers are in views.json, and the Explorer breaks them down by harness, version and reasoning mode.

ModelProviderScoreCoverageTasksObservationsHarnesses
gpt-6.1-solopenai-codex100.0%100.0%2266783
muse-spark-1.3-contributoropencode-go100.0%100.0%22615825
deepseek-v4-prodeepseek99.7%100.0%2266782
gemini-3.8-flashgemini99.7%100.0%2262261
deepseek-v4.1-flashdeepseek99.4%100.0%226316611
deepseek-v4.1-flashopencode-go99.3%100.0%22613565
mimo-v2.6-proopencode-go98.6%100.0%2262261
mimo-v2.6-flashopencode-go98.3%100.0%22611304
qwen3.8-flashopencode-go98.3%100.0%22613564
deepseek-v4-flashdeepseek98.2%100.0%22624869
deepseek-v4-flashopencode-go97.3%100.0%2269043
gpt-6-solopenai-codex97.1%100.0%45211303
grok-4.6opencode-go97.1%100.0%2264522
deepseek-v4.1-flashrouterai-deepseek97.0%100.0%2262261
kimi-k2.7-codeopencode-go96.7%100.0%2264522
space-bunny-freeopencode-go96.3%100.0%2264522
glm-5.3-flashzai96.0%100.0%22618088
space-bunny-freeopencode95.8%100.0%2264522
gpt-5.6-solopenai-codex95.0%100.0%2266783
gpt-5.6-terraopenai-codex95.0%100.0%2266783
gpt-5.6-lunaopencode-go94.1%100.0%2262261
mimo-v2.6-proxiaomi93.8%100.0%2264522
gpt-5.6-lunaopenai-codex93.7%100.0%4521086035
gpt-6-lunaopencode-go93.0%100.0%2262261
glm-5.3-flashrouterai-z-ai92.8%100.0%2262261
qwen3.7-plusopencode-go92.4%100.0%2262261
glm-5.3-flashopencode-go92.0%100.0%2264522
minimax-m3opencode-go90.3%100.0%2266783
ling-3.0-flash-vlrouterai-deepinfra90.3%100.0%2262261
longcat-2.5-preview-freeopencode-go89.3%100.0%2264522
mimo-v2.5xiaomi89.2%100.0%226248610
gpt-6-lunaopenai-codex87.2%100.0%45213563
qwen3.6-35b-a3bunknown85.1%100.0%2262261
mercury-2.5routerai-inception76.4%100.0%2262261

The harness list behind each row is in data/models.jsonl.gz, and views.json holds the same aggregates for the other groupings: by harness, by agent and by reasoning mode.

Data contributors

Thank you to everyone who has shared benchmark observations. Your work makes this public comparison possible.

ContributorAccepted runsConfigurations
@ashokkumards11
@dirac-run44
@TreyThomasCodes77
@user222155
@xuankunv144
@Yugimob2323

Accepted harnesses

Tables

ConfigOne row per
profilesconfiguration that was run, with its agent, harness, model and exact versions
configurationsrecipe behind a configuration, safe to publish
trialstask and attempt, with the first and final exact result
roundsattempt, with timing, tokens, cost and timeout state
tool-callstool the agent used, with its category and outcome
submissionsaccepted run, with its owner, purpose and definitions

The Dataset Viewer shows every config. dataset-index.json holds the source hashes, contracts, task sets, completeness and counts, and views.json, leaderboard.json and summary.json hold the aggregated rankings.

Source and contribution

Source code, run instructions and the contribution guide live at alexshpunt/explicit-edit-benchmark. Accepted bundles are kept under source/, the shards and summaries are views rebuilt from them, and each accepted harness family has a README badge under badges/.

Observations hold no prompts, arguments, commands, output, sessions, workspaces or credentials.