Codeseys/composer-replication-framework
0
1---2title: '\bugs: Scaling Data for Software Engineering Agents'3id: bugs-scaling-data-for-software-engineering-agents4tags:5- deepread6created: '2026-06-10T00:23:06.342374Z'7source: https://arxiv.org/html/2504.217988source_domain: arxiv.org9fetched_at: '2026-06-10T00:23:06.342106Z'10fetch_provider: builtin11status: draft12type: note13tier: institutional14content_type: paper15deprecated: false16---17 18\bugs: Scaling Data for Software Engineering Agents19\bugs20: Scaling Data for Software Engineering Agents21John Yang22123,24Kilian Lieret25226,27Carlos E. Jimenez28229,30Alexander Wettig31232,33Kabir Khandpur34335,36\And37Yanzhe Zhang38139,40Binyuan Hui41442,43Ofir Press44245,46Ludwig Schmidt47148,49Diyi Yang50151\And52153Stanford University54255Princeton University56357Indepedent58459Alibaba Qwen60Abstract61Despite recent progress in Language Models (LMs) for software engineering, collecting training data remains a significant pain point.62Existing datasets are small, with at most63164,6500066167000681{,}000691 , 00070s of training instances from711172117311741175or fewer GitHub repositories.76The procedures to curate such datasets are often complex, necessitating hundreds of hours of human labor; companion execution environments also take up several terabytes of storage, severely limiting their scalability and usability.77To address this pain point, we introduce SWE-smith, a novel pipeline for generating software engineering training data at scale.78Given any Python codebase,79\bugs80constructs a corresponding execution environment, then automatically synthesizes8110082100831008410085s to86187,8800089190000911{,}000921 , 00093s of task instances that break existing test(s) in the codebase.94Using95\bugs96, we create a dataset of97509850995010050101k instances sourced from102128103128104128105128106GitHub repositories, an order of magnitude larger than all previous works.107We train108SWE-agent-LM-32B109, achieving11040.211140.211240.211340.2114% Pass@1 resolve rate on the SWE-bench Verified benchmark, state of the art among open source models.115We open source116\bugs117(collection procedure, task instances, trajectories, models) to lower the barrier of entry for research in LM systems for automated software engineering.118All assets available at119https://swesmith.com120.1211122Introduction123Figure 1:124Scaling task instances125(left) and126performance127(right) for SWE-agent’s with128\bugs129.130Using131\bugs132, we can create133100134100135100136100137s to1381000139100014010001411000142s of instances for any Python codebase, enabling us to train143SWE-agent-LM-32B144which achieves14540.214640.214740.214840.2149% on SWE-bench Verified.150Language Model (LM) agents, such as SWE-agent151(Yang et al.,1522024a153)154or OpenHands155(Wang et al.,1562024157)158,159have made remarkable progress towards automating software engineering (SE) tasks, as tracked by benchmarks such as SWE-bench160(Jimenez et al.,1612024b162)163.164However, the most effective agents still rely on proprietary LMs, as building open source LMs for SE remains bottlenecked by the lack of large-scale, high-quality training data.165To ensure that open research remains relevant in this field, it is critical to develop infrastructure for collecting software engineering training data at scale.166The current open-source ecosystem offers two kinds of data sources to train LMs on SE tasks.167One simple approach is to crawl pull requests (PRs) and issues from GitHub repositories.168However, without execution environments or tests, these instances offer no reliable way of validating generated solutions,169and LMs are limited to learning from the surface form of code170(Xie et al.,1712025a172)173or via rewards based on superficial string similarity174(Wei et al.,1752025176)177.178In contrast, SWE-bench provides reliable validation by running unit tests against proposed solutions.179Another line of work has simply extended the SWE-bench collection strategy to a new set of repositories for training purposes180(Pan et al.,1812024182)183.184This produces flexible environments for training and distilling LM agents, since we can generate agent trajectories and filter them based on the unit test results.185However, the scalability of this approach is severely limited by the challenges associated with SWE-bench’s collection strategy.186SWE-bench’s filtering process leaves only a small number of PRs that not only resolve a Github issue, but also make meaningful changes to unit tests.187Also, setting up execution environments for each instance requires a substantial amount of human intervention.188In this paper, we introduce the189\bugs190toolkit, which marries the flexible execution environments of SWE-bench with scalable instance collection (Figure1911192).193\bugs194features several techniques to automatically synthesize bugs in existing GitHub repositories,195such as (1) generating errant rewrites of functions with an LM, (2) procedurally modifying the abstract syntax tree (AST) of functions, (3) undoing PRs, and (4) combining bugs.196Our key insight is that execution-based validation can not only validate proposed solutions,197but also identify bug candidates which cause substantial software regression (i.e., break tests).198Figure 2:199\bugs200creates training data for software engineering agents by crafting bugs into real codebases.201Given a codebase, we employ several strategies to create task instances that break existing tests.202Using203\bugs204, we create20550206502075020850209k+ task instances with execution environments from210128211128212128213128214real world repositories.215In a nutshell,216\bugs217puts forth the following task creation workflow, as shown in Figure2182219.220Given a codebase, we automatically set up a corresponding environment using SWE-agent221(Yang et al.,2222024a223)224.225Within this environment, we then use the aforementioned techniques to synthesize226100227100228100229100230s to2311232,23300023412350002361,0002371 , 000238s of task instances.239Finally, we craft realistic issue descriptions automatically with LMs.240\bugs241’s design significantly reduces the amount of human labor and storage required for constructing execution environments.242Using243\bugs244, we create a dataset of24550246502475024850249k task instances across250128251128252128253128254real-world GitHub repositories.255Using the256\bugs257dataset, we achieve a new open-weight state of the art result on SWE-bench verified.258Using the SWE-smith task instances, we generate2595260,26101626252630162645{,}0162655 , 016266expert trajectories with Claude 3.7 Sonnet and fine-tune Qwen 2.5 Coder Instruct26732268322693227032271B.272The resulting LM,273SWE-agent-LM-32B274, achieves27540.2276%277percent27840.227940.2\%28040.2 %281(+28233.428333.428433.428533.4286%) on SWE-bench Verified in a single attempt, without inference-time scaling.287This sets a new state of the art for open-weight models.288The scale and diversity of the289\bugs290dataset enables us to begin establishing truths and investigate interesting phenomena about developing SWE-agents.291Training on more instances, bug types, and repositories helps.292LM generated issue text approximates real ones effectively.293Using294\bugs295, we find that it’s possible to optimize LMs to perform well for specific repositories while only suffering minor generalization loss.296We release297\bugs298as an open-source toolkit — including instances, environments, and trajectories — to catalyze the development of stronger open-source LM agents.2992300\bugs301: Software Task Generation at Scale302The core principle of303\bugs304’s collection strategy is to define an execution environment first, and then synthesize task instances within the environment.305Conceptually, this is a simple inversion of SWE-bench’s approach, which instead prioritizes identifying task instances, and then attempts to build an environment for each.306In this section, we describe the procedure in detail and show how, in practice,307\bugs308scales significantly better in terms of repositories, task instances, and storage.3092.1310Collection311Building execution environments for repositories with passing tests.312Given a repository, we run SWE-agent313(Yang et al.,3142024a315)316on the latest commit for at most317100318100319100320100321steps, instructing it to install the codebase and run the test suite.322We then manually verify the installation and testing instructions, check if more than32380324803258032680327% of existing tests pass, and finally create a Docker image for the repository.328We target repositories for the3295330,33100033253330003345,0003355 , 000336most downloaded packages listed in the Python Package Index (PyPI) as of November 18, 2024, sort the PyPI packages by GitHub stars, and then remove any PyPI package with less than3371338,33900034013410003421,0003431 , 000344stars, as well as all34512346123471234812349SWE-bench test repositories from consideration.350More in §351A.2352.353Creating task instance candidates.354Per repository, we employ four different strategies to create candidates.355As shown in Figure3562357, each strategy takes in a repository as input, then produces task instance candidates represented as358.diff359files.360Extensive details in §361B362.363•364LM Generation365: Per repository, we identify all programmatic entities (functions, classes), then take two approaches: (1) provide an LM with the function and prompt it to introduce errant366modifications367(henceforth referred to as “LM Modify”), and (2) given only the function header and docstring, ask the LM to368rewrite369it (“LM Rewrite”).370More in §371B.1372.373•374Procedural Modification375: Per function, we acquire an abstract syntax tree (AST) representation of the code, then randomly perform one or more transformations (e.g., remove a conditional/loop, change an operator, +37611377113781137911380more. See Table3818382).383More in §384B.2385.386•387Combine Bugs388: LM generation and Procedural Modification task instances exclusively edit one function or class.389To create more complex tasks that require editing multiple portions of the codebase, we devise a “Patch Combination” strategy that creates a task instance by aggregating candidates from the same file(s) or module(s).390More in §391B.3392.393•394Invert PRs395(or “PR Mirror”): Per repository, we collect all PRs that modify Python files.396Per PR, we attempt to397undo398its revisions in the current version of the repository.399To achieve this, we provide an LM with the PR’s code changes (a400.diff401plaintext) and prompt it to rewrite each affected file such that the PR edits are reverted.402Unlike SWE-bench, we do403not404check out the PR’s base commit, as the install specifications determined in the previous step may not be compatible with older versions of the repo.405More in §406B.4407.408Execution-based validation of candidates.409We apply each candidate patch to the corresponding repository, run the test suite, and only keep patches that break one or more existing, passing tests (referred to as410Fail-to-Pass411or412F2P413test(s)).414For efficiency purposes, we also limit testing runtime to two minutes; bug candidates that cause test runtimes in excess of this time limit are discarded.415Minor additional details in §416A.3417.418Generating problem statements.419The issue text associated with a bug can significantly alter the difficulty and feasibility of the task instance.420Detailed descriptions of “expected” vs. “observed” behavior or bug-reproduction code in issue text heavily affect an agent’s capacity to localize bugs or iterate on proposed solutions.421We explore several techniques covered fully in §422D423, and ultimately settle on a simple strategy.424Per task instance, we provide an LM with the425.diff426patch, source code of a random F2P test, and execution output from running the repository’s test suite with the bug patch applied.427We prompt the LM for GitHub issue-style text that includes reproduction code based on the F2P test.428What human labor remains?429The steps requiring manual effort are (1) parsing the correct installation setup procedures from the agent trajectory (430∼4317432similar-to433absent4347435\sim 7436∼ 7437min per repository),438and (2) implementing the parser for test outputs (439∼4401441similar-to442absent4431444\sim 1445∼ 1446min per repository).447Step two requires very little time because parsers can be reused for repositories with the same testing infrastructure (e.g.,448pytest449).450\bugs451removes the need for manual efforts to determine installation specifications for multiple versions of a codebase across time, the most costly step of SWE-bench collection.452Creating453\bugs454took one author455∼45620457similar-to458absent45920460\sim{}20461∼ 20462h of human labor.4632.2464Features465We apply466\bugs467to468128469128470128471128472Python repositories, generating a total of47350474504755047650477k instances.478Table4791480captures the key statistics.481On average, we generate482381483381484381485381486task instances per repository, with as many as4872277488227748922774902277491for492pandas-dev/pandas493.494We summarize the distribution of task instances per repository in Figure4951496, where repositories are grouped into one of six general categories.497\bugs498took $4991360500136050113605021360503to create ($5041000505100050610005071000508to generate bugs, $509160510160511160512160513for automatic repository installation with SWE-agent, $514200515200516200517200518to generate issues for51910520105211052210523K bugs).524Generating an issue costs5252.545262.545272.545282.54529¢ on average.530More dataset analyses in §531C532.533\captionof534figure535Distribution of instances per repo for536128537128538128539128540repo’s grouped into5416542654365446545categories.546Bug Type547Yield %548# Insts549Cost550F2P551Lines552Combine55396.955496.955596.955696.9557%55810559,5600925611056209256310,09256410 , 0925650.005660.005670.005680.00569¢5701557115572155731557411575115761157711578LM Modify57956.058056.058156.058256.0583%58417585,5868875871758888758917,88759017 , 8875910.385920.385930.385940.38595¢59645974598459946003601360236033604LM Rewrite60535.060635.060735.060835.0609%6104611,61217361346141736154,1736164 , 1736173.936183.936193.936203.93621¢622462346244625462624627246282462924630PR Mirror63133.863233.863333.863433.8635%6362637,63834463926403446412,3446422 , 3446435.536445.536455.536465.53647¢648364936503651365214653146541465514656Procedural65740.265840.265940.266040.2661%66215663,6646416651566664166715,64166815 , 6416690.006700.006710.006720.00673¢67476757676767776785679568056815682Total68350.168450,1376852.32¢68666875688Table 1:689Summary of690\bugs691statistics.692“Yield %” is the % of candidates generated by a strategy that break6931694+695limit-from69616971+6981 +699tests.700“Cost” is the average cost to generate one candidate.701“F2P” (Fail to Pass tests), “Lines [Edited]” are median values.702Bug generation strategies vary in cost and yield rate.703Of methods relying on LMs, PR Mirrors are more expensive because the task entails rewriting entire files, as opposed to individual functions for LM Modify and LM Rewrite.704Yield rates are limited by either lack of test coverage for the change or because the bug candidate did not actually introduce relevant issues.705For example, for LM Rewrite, the LM is asked to re-implement the function; it is706not707explicitly asked for bugs.708When requested outright (LM Modify), the yield is higher.709Dataset710# Tasks711# Repos712Exec?713Source714Env. Size715R2E716(Jain et al.,7172024718)7190.257200.257210.257220.25723k724137725137726137727137728Synth729270730270731270732270733GBs734R2E-gym (Subset)735(Jain et al.,7362025737)7384.67394.67404.67414.6742k74310744107451074610747Synth7484749475047514752TBs753SWE-bench-extra754(Badertdinov et al.,7552024756)7576.387586.387596.387606.38761k7622763276427652766k767Real768-769SWE-bench-train770(Jimenez et al.,7712024b772)77319774197751977619777k77837779377803778137782Real783-784SWE-fixer785(Xie et al.,7862025a787)788115789115790115791115792k793856794856795856796856797Real798-799SWE-gym800(Pan et al.,8012024802)8032.48042.48052.48062.4807k80811809118101181111812Real8136814681568166817TBs818\bugs81950820508215082250823k824128825128826128827128828Both829295830295831295832295833GBs834Table 2:835Comparison of open source training datasets for software engineering tasks.836Relative to existing datasets,837\bugs838has multiple times the number of task instances, repositories, and environments at a fraction of prior storage costs.839SWE-fixer and SWE-bench-train task instances do not have execution environments, so “Env. Size” is blank.840How difficult are841\bugs842task instances?843To determine whether task instances produced by844\bugs845are realistic and challenging, we train a Qwen8462.58472.58482.58492.585032851328523285332854B model on8551856,85769985818596998601{,}6998611 , 699862human-annotated (task, label) pairs from863Chowdhury et al. (8642024865)866to rate tasks as (867easy868,869medium870,871hard872) by training.873To quantify difficulty, each difficulty label corresponds to values of8741875187618771878/8795880588158825883/8849885988698879888.889The model achieves89075.389175.389275.389375.3894% test accuracy.895We then rate difficulty of task instances from both896\bugs897and prior SWE-bench style datasets898(Chowdhury et al.,8992024900; Jimenez et al.,9012024b902; Pan et al.,9032024904; Yang et al.,9052024b906)907.908\bugs909task instances span a broad range of difficulties, similar to SWE-bench and SWE-gym.910The average difficulty score for911\bugs912(9135.279145.279155.279165.27917–9185.729195.729205.729215.72922across bug generation strategies) is comparable to SWE-bench (9235.019245.019255.019265.01927) and SWE-gym (9285.629295.629305.629315.62932). This suggests SWE-smith enables realistic and appropriately challenging evaluation.933We discuss why bug strategies yield different levels of difficulty and visualize difficulty per dataset in §934E935.936Scaling execution environments.937Unlike SWE-bench which creates a Docker image per task instance,938\bugs939leverages a simpler design where tasks from the same repository share the same environment, reducing storage overhead significantly, as shown in Table9402941.942This approach not only makes scaling task instances more affordable, but also renders943\bugs944more accessible and maintainable than existing datasets.945We estimate that creating a similar quantity of task instances (94650947509485094950950k) using SWE-bench would require95150952509535095450955to956150957150958150959150960TBs of storage for environments, a961500962500963500964500965x difference.966Extended discussion in §967C.1968.9693970Experiments971To explore the utility of972\bugs973for training software engineering agents, we use rejection sampling fine-tuning974(Yuan et al.,9752023976)977as the primary procedure for improving a base LM with978\bugs979.980Our experiment workflow is as follows.981First, we curate a subset of982\bugs983task instances.984Next, we run an agent system with an expert model on this subset.985At this step, the trajectory corresponding to each run is recorded.986Then, we fine-tune the base (or “student”) model on the trajectories corresponding to resolved instances.987Finally, we evaluate the agent system run with the student model on a separate, test split.988Models.989For expert models, we use990claude-3-7-sonnet-20250219991(Anthropic,9922025993)994.995For fair comparisons with prior works996(Pan et al.,9972024998)999, we also use1000claude-3-5-sonnet-202406201001and1002gpt-4o-2024-08-061003.1004We use the1005Qwen-2.5-Coder-Instruct1006(Hui et al.,100720241008)100971010710117101271013B and1014321015321016321017321018B series as the base models.1019Training and hyperparameter details are in §1020F.11021.1022Agent system.1023We use SWE-agent1024(Yang et al.,10252024a1026)1027, an agent system for solving GitHub issues.1028SWE-agent provides a base LM with an Agent Computer Interface (ACI) that enables more effective interactions with a codebase.1029At each turn, SWE-agent prompts an LM to generate a ReAct1030(Yao et al.,10312023b1032)1033style (thought, action) pair, where the action either edits a file or executes a shell command.1034We choose SWE-agent because, at the time of writing, SWE-agent with Claude 3.7 Sonnet is the top open source solution on SWE-bench.1035When generating trajectories with expert models, we run SWE-agent for at most1036751037751038751039751040steps and $10412.0010422.0010432.0010442.001045cost limit.1046For inference of student models, we impose the same1047751048751049751050751051step maximum and fix temperature at10520.010530.010540.010550.01056.1057Full configuration details are in §1058F.11059.1060Evaluation metrics.1061We evaluate on the SWE-bench Lite and Verified1062(Chowdhury et al.,106320241064)1065subsets.1066SWE-bench evaluates AI systems on their ability to solve software issues from1067121068121069121070121071real world GitHub repositories.1072The Lite split is a subset of10733001074300107530010763001077instances, curated to be an easier evaluation set that’s less costly to run.1078The Verified split is a human-curated subset of10795001080500108150010825001083instances, selected for clearer problem statements and more reliable evaluation.1084To assess generalization beyond Python, we also evaluate on SWE-bench Multilingual, a new dataset introduced in this paper.1085SWE-Bench Multilingual consists of10863001087300108830010893001090task instances that cover109191092910939109491095additional programming languages.1096See §1097F.21098for more details.1099We report the1100% resolved1101metric, the proportion of successfully resolved instances.110241103Results1104Table110531106compares the performance of Qwen 2.5 Coder Instruct models (7B and 32B), fine-tuned on110751108,110901611105111101611125{,}01611135 , 0161114\bugs1115trajectories.1116We refer to them as1117SWE-agent-LM-7B1118and1119SWE-agent-LM-32B1120; the latter achieves state-of-the-art performance.1121Model1122System1123Train Size1124Lite1125Verified1126Closed Weight Models1127GPT-4o1128(OpenAI,11292024a1130)1131Agentless1132-113332.0113432.0113532.0113632.0113738.8113838.8113938.8114038.81141OpenHands1142-114322.0114422.0114522.0114622.01147-1148SWE-agent1149-115018.3115118.3115218.3115318.3115423.0115523.0115623.0115723.01158Claude 3.5 Sonnet1159(Anthropic,116020241161)1162Agentless1163-116440.7116540.7116640.7116740.7116850.8116950.8117050.8117150.81172AutoCodeRover1173-1174-117546.2117646.2117746.2117846.21179OpenHands1180-118141.7118241.7118341.7118441.7118553.0118653.0118753.0118853.01189SWE-agent1190-119123.0119223.0119323.0119423.0119533.6119633.6119733.6119833.61199Claude 3.7 Sonnet1200(Anthropic,