zlab-princeton/SWEeper-Bench-traces
🧹 SWEeper-Bench agent traces Paper · Blog · Code · Benchmark This dataset has every run behind the SWEeper-Bench paper: 5,200 runs from 42 agent configurations. For each run you get the prompt the coding agent saw, its full trajectory, the shell commands it ran against the app, its patch, and the browser verifier's trajectories and verdicts. Main results All 15 agents ran on the 200 tasks with the open-ended prompt and the browser prompt template. A task… See the full description on the dataset page: https://huggingface.co/datasets/zlab-princeton/SWEeper-Bench-traces.
🧹 SWEeper-Bench agent traces
Paper · Blog · Code · Benchmark
This dataset has every run behind the SWEeper-Bench paper: 5,200 runs from 42 agent configurations. For each run you get the prompt the coding agent saw, its full trajectory, the shell commands it ran against the app, its patch, and the browser verifier's trajectories and verdicts.
<p align="center"> <img src="https://raw.githubusercontent.com/zlab-princeton/sweeper-bench/main/assets/leaderboard.svg" alt="Pass rate versus mean cost per task for 15 agents. Grok 4.6 leads at 59.0%." width="90%"> </p>
Main results
All 15 agents ran on the 200 tasks with the open-ended prompt and the browser prompt template. A task passes only when both of its behavior tests pass. A task with no prediction or evaluation counts as a fail.
Ablations
Finding runs
runs.csv has one row per run, and the dataset viewer shows it as the runs split.
Download a run
from huggingface_hub import hf_hub_download, snapshot_download
import pandas as pd
repo = "zlab-princeton/SWEeper-Bench-traces"
runs = pd.read_csv(hf_hub_download(repo, "runs.csv", repo_type="dataset"))
run = runs.query("model == 'gpt-6-astra' and task_id == 'sweeper-001' and type == 'main'").iloc[0]
path = snapshot_download(repo, repo_type="dataset", allow_patterns=f"{run.run_id}/*")The whole dataset is about 5 GB in 65,000 files. To get all of it, git clone https://huggingface.co/datasets/zlab-princeton/SWEeper-Bench-traces is much faster than snapshot_download.
Files in a run
<run_id>/
metadata.json Same fields as the runs.csv row
prediction/
prompt.txt Prompt given to the coding agent
harness.stdout.txt Agent output, including its trajectory
harness.stderr.txt Agent error output
trajectory.jsonl Claude Code only: stream events
final.txt Claude Code only: final message
status.shell.txt Commands the agent ran in the app container
patch The agent's patch
evaluation/
result.json Verdicts for both behavior tests
<test-id>/
prompt.txt Verifier instructions
trajectory.jsonl Verifier trajectory in the browser
last-message.txt Verifier's final verdictSome files are missing when a run stopped before that stage. The hf_url inside each metadata.json points to the original upload; use the one in runs.csv.
License
This dataset is released under CC BY 4.0. Traces quote source code from the benchmark applications, which keep their own licenses.
Citation
@article{yao2026sweeperbench,
title = {SWEeper-Bench: Can Agents Discover Bugs in Interactive Software?},
author = {Yao, Yang and Chen, Haozhe and Kang, Bingyi and
Narasimhan, Karthik R and Liu, Zhuang},
journal = {arXiv preprint},
year = {2026}
}