Team Ai
Modelpublic

Codeseys/composer-replication-framework

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes
introducing-composer-25-cursor.md179 linesDownload Raw Back to notes
1---2title: Introducing Composer 2.5 · Cursor3id: introducing-composer-25-cursor4tags:5- socratic-mcts-swe-worldmodel-8f6dea6created: '2026-06-09T04:19:33.431032Z'7source: https://cursor.com/blog/composer-2-58source_domain: cursor.com9fetched_at: '2026-06-09T04:19:32.954586Z'10fetch_provider: builtin11status: draft12type: note13deprecated: false14summary: Introducing Composer 2.5 · Cursor15---16 17Introducing Composer 2.5 · Cursor18Blog19/20research21Composer 2.5 is now available in Cursor.22It's a substantial improvement in intelligence and behavior over23Composer 224. It is better at sustained work on long-running tasks, follows complex instructions more reliably, and is more pleasant to collaborate with.25We improved Composer by scaling training, generating more complex RL environments, and introducing new learning methods.26In addition to training Composer 2.5 on more difficult tasks, we improved behavioral aspects of the model like communication style and effort calibration. These dimensions are not well captured by existing benchmarks, but we find that they matter for real-world usefulness.27Composer 2.5 is built on the same open-source checkpoint as Composer 2,28Moonshot's Kimi K2.529.30Together31with SpaceXAI32, we're training a significantly larger model from scratch, using 10x more total compute. With Colossus 2's million H100-equivalents and our combined data and training techniques, we expect this to be a major leap in model capability.33#34Training Composer 2.535Composer 2.5 contains several new improvements to our training stack. These changes target both model intelligence and usability.36#37Targeted RL with textual feedback38Credit assignment during RL is becoming an increasingly difficult challenge as rollouts can span hundreds of thousands of tokens. When a reward is computed over an entire rollout, it may be hard for the model to tell which specific decision helped or hurt the outcome. This is especially limiting when we want to discourage a localized behavior, such as a bad tool call, a confusing explanation, or a style violation. The final reward can tell us that something went wrong, but it is a noisy signal for39where40it went wrong.41To address this, we trained Composer 2.5 with targeted textual feedback.42143The idea is to provide feedback directly at the point in the trajectory where the model could have behaved better. For a target model message, we construct a short hint describing the desired improvement, insert that hint into the local context, and use the resulting model distribution as a teacher. We use the policy with the original context as the student and add an on-policy distillation KL loss that moves the student's token probabilities toward the teacher's. This gives us a localized training signal for the behavior we want to change, while still retaining the broader RL objective over the full trajectory.44As an illustration of the text feedback process, consider a long rollout that includes a tool call error where the model attempts to call a tool that is not available. During the rollout, the model will receive a “Tool not found” error and continue making additional valid tool calls. The fact that it hit one error in the process of hundreds of tool calls will have a minimal impact on its final reward.45With text feedback, we can target this specific mistake by inserting a hint in the context of the problematic turn, such as “Reminder: Available tools…” with a list of available tools. This hint changes the probabilities for the teacher, lowering those for the wrong tool and increasing those for a valid replacement. For that turn only, we then update the student weights towards to the new probabilities.46During the Composer 2.5 run, we applied this method to a variety of model behaviors, from coding style to model communication.47#48Synthetic data49During RL training, Composer's coding ability improves substantially to the point where it begins to get most training problems correct. To continue increasing intelligence, we both select for and create harder tasks dynamically throughout the run. Composer 2.5 is trained with 25x more synthetic tasks than Composer 2.50We use a range of approaches for creating synthetic tasks that are grounded in real codebases. For example, one synthetic approach is feature deletion. For these tasks the agent is given a codebase with a large set of tests, and asked to delete code and files in such a way that the codebase remains functional while specific testable features are removed. The synthetic task is to reimplement the feature, and the tests are used as a verifiable reward.51One downstream consequence of large scale synthetic task creation is that it can cause  unexpected reward hacking. As the model became more adept, Composer 2.5 was able to find increasingly sophisticated workarounds to solve the task at hand. In one example, the model found a leftover Python type-checking cache and reverse-engineered the format to find a deleted function signature. In another, it was able to find and decompile Java bytecode to reconstruct a third-party API. We were able to find and diagnose these problems using agentic monitoring tools, but they demonstrate the increasing care necessary for large scale RL.52#53Sharded Muon and dual mesh HSDP54For continued pretraining, we use Muon with distributed orthogonalization. After forming the momentum update, we run Newton-Schulz at the model's natural granularity: per attention head for attention projections, and per expert for stacked MoE weights.55The main cost is orthogonalizing expert weights. For sharded parameters, we batch same-shaped tensors, all-to-all shards into complete matrices, run Newton-Schulz, then all-to-all the result back to the original sharded layout. These transfers are asynchronous: while one task is waiting on communication, the optimizer runtime advances other Muon tasks, overlapping network and compute. This is equivalent to full-matrix Muon, but keeps the shard group busy; on the 1T model, optimizer step time is 0.2s.56This interacts closely with how we use HSDP for MoE models. HSDP forms multiple FSDP replicas and all-reduces gradients across corresponding shards. We use separate HSDP layouts for non-expert and expert weights: non-expert weights are comparatively small, so their FSDP groups can stay narrow, often within a node or rack, while expert weights hold most of the parameters and most of the Muon compute, so they use a wider expert sharding mesh.57Keeping these layouts separate also lets independent parallelism dimensions overlap: CP=2 and EP=8 can run on 8 GPUs instead of requiring 16 in a single shared mesh. This avoids wide communication for small non-expert state while spreading expert optimizer work over many GPUs.58#59Try Composer 2.560Composer 2.5 is priced at $0.50/M input and $2.50/M output tokens.61There's also a62faster variant with the same intelligence63at $3.00/M input and $15.00/M output tokens, a lower cost than the fast tiers of other frontier models. Similar to Composer 2, fast is the default option. See our64model docs65for full details.66Composer 2.5 includes double usage for the first week.67For more background on this approach see68Self-Distillation Enables Continual Learning69,70Reinforcement Learning via Self-Distillation71, and72Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models73.74↩75Related posts76Feb 9, 202677·78Research79Introducing Composer 1.580Cursor Team81·823 min read83Mar 19, 202684·85Research86Introducing Composer 287Cursor Team88·893 min read90May 6, 202691·92Research93Bootstrapping Composer with autoinstall94Shomil, Joshua & Andrew95·966 min read97View more posts98→99Blog100/101research102Composer 2.5 is now available in Cursor.103It's a substantial improvement in intelligence and behavior over104Composer 2105. It is better at sustained work on long-running tasks, follows complex instructions more reliably, and is more pleasant to collaborate with.106We improved Composer by scaling training, generating more complex RL environments, and introducing new learning methods.107In addition to training Composer 2.5 on more difficult tasks, we improved behavioral aspects of the model like communication style and effort calibration. These dimensions are not well captured by existing benchmarks, but we find that they matter for real-world usefulness.108Composer 2.5 is built on the same open-source checkpoint as Composer 2,109Moonshot's Kimi K2.5110.111Together112with SpaceXAI113, we're training a significantly larger model from scratch, using 10x more total compute. With Colossus 2's million H100-equivalents and our combined data and training techniques, we expect this to be a major leap in model capability.114#115Training Composer 2.5116Composer 2.5 contains several new improvements to our training stack. These changes target both model intelligence and usability.117#118Targeted RL with textual feedback119Credit assignment during RL is becoming an increasingly difficult challenge as rollouts can span hundreds of thousands of tokens. When a reward is computed over an entire rollout, it may be hard for the model to tell which specific decision helped or hurt the outcome. This is especially limiting when we want to discourage a localized behavior, such as a bad tool call, a confusing explanation, or a style violation. The final reward can tell us that something went wrong, but it is a noisy signal for120where121it went wrong.122To address this, we trained Composer 2.5 with targeted textual feedback.1231124The idea is to provide feedback directly at the point in the trajectory where the model could have behaved better. For a target model message, we construct a short hint describing the desired improvement, insert that hint into the local context, and use the resulting model distribution as a teacher. We use the policy with the original context as the student and add an on-policy distillation KL loss that moves the student's token probabilities toward the teacher's. This gives us a localized training signal for the behavior we want to change, while still retaining the broader RL objective over the full trajectory.125As an illustration of the text feedback process, consider a long rollout that includes a tool call error where the model attempts to call a tool that is not available. During the rollout, the model will receive a “Tool not found” error and continue making additional valid tool calls. The fact that it hit one error in the process of hundreds of tool calls will have a minimal impact on its final reward.126With text feedback, we can target this specific mistake by inserting a hint in the context of the problematic turn, such as “Reminder: Available tools…” with a list of available tools. This hint changes the probabilities for the teacher, lowering those for the wrong tool and increasing those for a valid replacement. For that turn only, we then update the student weights towards to the new probabilities.127During the Composer 2.5 run, we applied this method to a variety of model behaviors, from coding style to model communication.128#129Synthetic data130During RL training, Composer's coding ability improves substantially to the point where it begins to get most training problems correct. To continue increasing intelligence, we both select for and create harder tasks dynamically throughout the run. Composer 2.5 is trained with 25x more synthetic tasks than Composer 2.131We use a range of approaches for creating synthetic tasks that are grounded in real codebases. For example, one synthetic approach is feature deletion. For these tasks the agent is given a codebase with a large set of tests, and asked to delete code and files in such a way that the codebase remains functional while specific testable features are removed. The synthetic task is to reimplement the feature, and the tests are used as a verifiable reward.132One downstream consequence of large scale synthetic task creation is that it can cause  unexpected reward hacking. As the model became more adept, Composer 2.5 was able to find increasingly sophisticated workarounds to solve the task at hand. In one example, the model found a leftover Python type-checking cache and reverse-engineered the format to find a deleted function signature. In another, it was able to find and decompile Java bytecode to reconstruct a third-party API. We were able to find and diagnose these problems using agentic monitoring tools, but they demonstrate the increasing care necessary for large scale RL.133#134Sharded Muon and dual mesh HSDP135For continued pretraining, we use Muon with distributed orthogonalization. After forming the momentum update, we run Newton-Schulz at the model's natural granularity: per attention head for attention projections, and per expert for stacked MoE weights.136The main cost is orthogonalizing expert weights. For sharded parameters, we batch same-shaped tensors, all-to-all shards into complete matrices, run Newton-Schulz, then all-to-all the result back to the original sharded layout. These transfers are asynchronous: while one task is waiting on communication, the optimizer runtime advances other Muon tasks, overlapping network and compute. This is equivalent to full-matrix Muon, but keeps the shard group busy; on the 1T model, optimizer step time is 0.2s.137This interacts closely with how we use HSDP for MoE models. HSDP forms multiple FSDP replicas and all-reduces gradients across corresponding shards. We use separate HSDP layouts for non-expert and expert weights: non-expert weights are comparatively small, so their FSDP groups can stay narrow, often within a node or rack, while expert weights hold most of the parameters and most of the Muon compute, so they use a wider expert sharding mesh.138Keeping these layouts separate also lets independent parallelism dimensions overlap: CP=2 and EP=8 can run on 8 GPUs instead of requiring 16 in a single shared mesh. This avoids wide communication for small non-expert state while spreading expert optimizer work over many GPUs.139#140Try Composer 2.5141Composer 2.5 is priced at $0.50/M input and $2.50/M output tokens.142There's also a143faster variant with the same intelligence144at $3.00/M input and $15.00/M output tokens, a lower cost than the fast tiers of other frontier models. Similar to Composer 2, fast is the default option. See our145model docs146for full details.147Composer 2.5 includes double usage for the first week.148For more background on this approach see149Self-Distillation Enables Continual Learning150,151Reinforcement Learning via Self-Distillation152, and153Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models154.155↩156Related posts157Feb 9, 2026158·159Research160Introducing Composer 1.5161Cursor Team162·1633 min read164Mar 19, 2026165·166Research167Introducing Composer 2168Cursor Team169·1703 min read171May 6, 2026172·173Research174Bootstrapping Composer with autoinstall175Shomil, Joshua & Andrew176·1776 min read178View more posts179→