Team Ai
Datasetpublic

zlab-princeton/SWEeper-Bench-traces

🧹 SWEeper-Bench agent traces Paper · Blog · Code · Benchmark This dataset has every run behind the SWEeper-Bench paper: 5,200 runs from 42 agent configurations. For each run you get the prompt the coding agent saw, its full trajectory, the shell commands it ran against the app, its patch, and the browser verifier's trajectories and verdicts. Main results All 15 agents ran on the 200 tasks with the open-ended prompt and the browser prompt template. A task… See the full description on the dataset page: https://huggingface.co/datasets/zlab-princeton/SWEeper-Bench-traces.

sourceHugging Facecc-by-4.0updated 3d agoView on Hugging Face
1likes4.3kdownloads
Dataset Card

🧹 SWEeper-Bench agent traces

Paper · Blog · Code · Benchmark

This dataset has every run behind the SWEeper-Bench paper: 5,200 runs from 42 agent configurations. For each run you get the prompt the coding agent saw, its full trajectory, the shell commands it ran against the app, its patch, and the browser verifier's trajectories and verdicts.

<p align="center"> <img src="https://raw.githubusercontent.com/zlab-princeton/sweeper-bench/main/assets/leaderboard.svg" alt="Pass rate versus mean cost per task for 15 agents. Grok 4.6 leads at 59.0%." width="90%"> </p>

Main results

All 15 agents ran on the 200 tasks with the open-ended prompt and the browser prompt template. A task passes only when both of its behavior tests pass. A task with no prediction or evaluation counts as a fail.

ModelHarnessEffortPass rate
grok-4.6Cursorxhigh-fast59.0%
qwen3.8-flashCodexxhigh57.5%
claude-fable-5-1Claude Codexhigh56.5%
claude-opus-5Claude Codexhigh56.5%
gpt-6-astraCodexxhigh56.5%
deepseek-v4.1-flashDeepSeek harnessxhigh51.5%
kimi-k3Kimi Codehigh48.5%
deepseek-v4-proDeepSeek harnessxhigh47.0%
gpt-5.6-solCodexxhigh47.0%
muse-spark-1.3Muse Codexhigh47.0%
glm-5.3Codexmax43.5%
gpt-5.6-lunaCodexxhigh42.0%
glm-5.3-flashCodexmax39.5%
gpt-5.6-terraCodexxhigh35.0%
qwen3.8-maxCodexnone28.0%

Ablations

`type`What changesRuns
ablation-promptThe agent gets a description of the bug (specified prompt). Five agents pass 96.0% to 96.5%.1,000
ablation-templateThe prompt no longer asks the agent to test in a browser (non-browser template).400
ablation-harnessgpt-5.6-luna in four other harnesses, on a 40-task subset.160
ablation-timeTime budgets of 20, 40, 60, and 80 minutes, on the same 40 tasks.640

Finding runs

runs.csv has one row per run, and the dataset viewer shows it as the runs split.

ColumnMeaning
run_idFolder name of the run
task_idsweeper-001 through sweeper-200
model, harness, effortThe agent
promptscoped (open-ended) or specified (bug described)
templatebrowser or non-browser
set200 (all tasks) or 40 (ablation subset)
time_budgetunlimited or minutes
typemain or one of the ablations above
target, preservationVerdict of each behavior test
overallpass only when both tests pass
hf_urlLink to the run's folder

Download a run

python
from huggingface_hub import hf_hub_download, snapshot_download
import pandas as pd

repo = "zlab-princeton/SWEeper-Bench-traces"
runs = pd.read_csv(hf_hub_download(repo, "runs.csv", repo_type="dataset"))

run = runs.query("model == 'gpt-6-astra' and task_id == 'sweeper-001' and type == 'main'").iloc[0]
path = snapshot_download(repo, repo_type="dataset", allow_patterns=f"{run.run_id}/*")

The whole dataset is about 5 GB in 65,000 files. To get all of it, git clone https://huggingface.co/datasets/zlab-princeton/SWEeper-Bench-traces is much faster than snapshot_download.

Files in a run

text
<run_id>/
  metadata.json                 Same fields as the runs.csv row
  prediction/
    prompt.txt                  Prompt given to the coding agent
    harness.stdout.txt          Agent output, including its trajectory
    harness.stderr.txt          Agent error output
    trajectory.jsonl            Claude Code only: stream events
    final.txt                   Claude Code only: final message
    status.shell.txt            Commands the agent ran in the app container
    patch                       The agent's patch
  evaluation/
    result.json                 Verdicts for both behavior tests
    <test-id>/
      prompt.txt                Verifier instructions
      trajectory.jsonl          Verifier trajectory in the browser
      last-message.txt          Verifier's final verdict

Some files are missing when a run stopped before that stage. The hf_url inside each metadata.json points to the original upload; use the one in runs.csv.

License

This dataset is released under CC BY 4.0. Traces quote source code from the benchmark applications, which keep their own licenses.

Citation

bibtex
@article{yao2026sweeperbench,
  title   = {SWEeper-Bench: Can Agents Discover Bugs in Interactive Software?},
  author  = {Yao, Yang and Chen, Haozhe and Kang, Bingyi and
             Narasimhan, Karthik R and Liu, Zhuang},
  journal = {arXiv preprint},
  year    = {2026}
}
zlab-princeton/SWEeper-Bench-traces · Team Ai