Codeseys/composer-replication-framework
0
1[2 {3 "id": "rec-1",4 "category": "break-paragraph",5 "severity": "high",6 "current": "RLVR-trained models systematically shortcut extensional verifiers, with shortcut prevalence *rising with task complexity and inference-time compute*; and monitors trained on synthetic hacks *fail to generalize* to in-the-wild hacking, so a `HackMonitor` validated on constructed examples is exactly the one likely to miss the real thing [29][30][31]. Cursor itself observed Composer 2.5 reverse-engineering a leftover type-check cache and decompiling Java bytecode to recover deleted signatures [1]. The oracle *bounds* the hack surface",7 "recommended": "RLVR-trained models systematically shortcut extensional verifiers, with shortcut prevalence *rising with task complexity and inference-time compute*; and monitors trained on synthetic hacks *fail to generalize* to in-the-wild hacking, so a `HackMonitor` validated on constructed examples is exactly the one likely to miss the real thing [29][30][31]. Cursor itself observed Composer 2.5 reverse-engineering a leftover type-check cache and decompiling Java bytecode to recover deleted signatures [1].\n\nThe oracle *bounds* the hack surface",8 "rationale": "Line 79 is the longest paragraph in the report (~1880 chars); breaking at the Cursor-example sentence boundary splits a dense wall into two scannable units."9 },10 {11 "id": "rec-2",12 "category": "break-paragraph",13 "severity": "high",14 "current": "Self-distillation in the inner loop is, in this configuration, a *stabilizer* and not only a collapse risk: SDFT shows on-policy self-distillation from demonstrations reduces catastrophic forgetting and lets a single model accumulate skills sequentially — the opposite of model collapse — and Channel-2 SDPO is exactly that on-policy, demonstration-conditioned regime, not the static-synthetic-data regime that collapses [36]. But the repo's own ADR-013 warns",15 "recommended": "Self-distillation in the inner loop is, in this configuration, a *stabilizer* and not only a collapse risk: SDFT shows on-policy self-distillation from demonstrations reduces catastrophic forgetting and lets a single model accumulate skills sequentially — the opposite of model collapse — and Channel-2 SDPO is exactly that on-policy, demonstration-conditioned regime, not the static-synthetic-data regime that collapses [36].\n\nBut the repo's own ADR-013 warns",16 "rationale": "Line 103 (~1670 chars) is a wall mixing the stabilizer claim, the amplification caveat, and the flywheel argument; breaking after the SDFT point isolates the first claim."17 },18 {19 "id": "rec-3",20 "category": "break-paragraph",21 "severity": "high",22 "current": "every working SWE flywheel optimizes a true execution oracle (Socratic-SWE +7.8 over three iters beating self-play at equal compute [10]; DeepSWE +20 Pass@1 in 200 RL steps on sparse 0/1 reward; SWE-RL 41% generalizing OOD [37]). Most collapse stories require a proxy or self-judged verifier",23 "recommended": "every working SWE flywheel optimizes a true execution oracle (Socratic-SWE +7.8 over three iters beating self-play at equal compute [10]; DeepSWE +20 Pass@1 in 200 RL steps on sparse 0/1 reward; SWE-RL 41% generalizing OOD [37]).\n\nMost collapse stories require a proxy or self-judged verifier",24 "rationale": "Further splits the remaining tail of the overlong line-103 paragraph at the working-flywheel / collapse-stories pivot so each half reads as one idea."25 },26 {27 "id": "rec-4",28 "category": "break-paragraph",29 "severity": "high",30 "current": "alignment-gated structure over naive train-on-all [55]. The content side is trainable:",31 "recommended": "alignment-gated structure over naive train-on-all [55].\n\nThe content side is trainable:",32 "rationale": "Line 27 (~1670 chars) packs four distinct facts into one paragraph; breaking before the trainable-content fact separates the anti-emergence evidence from the pro-training evidence."33 },34 {35 "id": "rec-5",36 "category": "break-paragraph",37 "severity": "high",38 "current": "branch factor × sandbox cold-start [6]. The layered posture: **gVisor",39 "recommended": "branch factor × sandbox cold-start [6].\n\nThe layered posture: **gVisor",40 "rationale": "Line 179 (~1616 chars) is a §8 wall; breaking before the layered-isolation discussion separates the framing sentence from the three-tier detail."41 },42 {43 "id": "rec-6",44 "category": "break-paragraph",45 "severity": "high",46 "current": "many small vLLM pods share a GPU [45]. One hosting fact feeds the platform choice:",47 "recommended": "many small vLLM pods share a GPU [45].\n\nOne hosting fact feeds the platform choice:",48 "rationale": "Splits the remaining tail of the overlong line-179 §8 paragraph at the hosting-fact pivot, separating sandbox/GPU sizing from the TRL-vs-VeRL engine choice."49 },50 {51 "id": "rec-7",52 "category": "break-paragraph",53 "severity": "high",54 "current": "the predicted `tool_error` kind — never reconstruct the full state, a high-entropy sea of irrelevant tokens [14][15]. The latent-motion line carries the same discipline into 2026:",55 "recommended": "the predicted `tool_error` kind — never reconstruct the full state, a high-entropy sea of irrelevant tokens [14][15].\n\nThe latent-motion line carries the same discipline into 2026:",56 "rationale": "Line 27 also runs the MuZero/Dreamer design discipline into the latent-motion result; breaking before the 2026 latent-motion sentence relieves the second half of this overlong paragraph."57 },58 {59 "id": "rec-8",60 "category": "break-paragraph",61 "severity": "medium",62 "current": "RL on the token's *placement* teaches the *governance* that is the real bottleneck [11].",63 "recommended": "RL on the token's *placement* teaches the *governance* that is the real bottleneck [11].\n",64 "rationale": "Line 33 (~1407 chars) runs the SDPO carrier and deliberate-token mechanisms together; inserting a break after the placement sentence separates the two mechanisms."65 },66 {67 "id": "rec-9",68 "category": "break-paragraph",69 "severity": "medium",70 "current": "that absence is the delta. So \"multi-model Monte-Carlo tree-of-work\" means, concretely:",71 "recommended": "that absence is the delta.\n\nSo \"multi-model Monte-Carlo tree-of-work\" means, concretely:",72 "rationale": "Line 19 (~1320 chars) is a §1 wall; breaking before the concrete restatement separates the repo-primitive mapping from the definitional summary."73 },74 {75 "id": "rec-10",76 "category": "break-paragraph",77 "severity": "medium",78 "current": "*structured/selective negatives beat both raw train-on-all and positives-only pruning.* The verdict: **train on all surviving branches, typed and routed by signal, never as raw negative policy gradient.**",79 "recommended": "*structured/selective negatives beat both raw train-on-all and positives-only pruning.*\n\nThe verdict: **train on all surviving branches, typed and routed by signal, never as raw negative policy gradient.**",80 "rationale": "Separates the bracketing observation from the verdict so the §4 headline verdict stands out before the numbered routing list."81 },82 {83 "id": "rec-11",84 "category": "break-paragraph",85 "severity": "medium",86 "current": "what makes divergence-gating mandatory [6]. The gating pays for itself:",87 "recommended": "what makes divergence-gating mandatory [6].\n\nThe gating pays for itself:",88 "rationale": "Line 191 (~1112 chars) is the dense Cost paragraph; breaking before the gating-savings sentence separates the cost problem from the mitigation."89 },90 {91 "id": "rec-12",92 "category": "break-paragraph",93 "severity": "medium",94 "current": "real for the *unguarded* version [8]. The escape is not better replay;",95 "recommended": "real for the *unguarded* version [8].\n\nThe escape is not better replay;",96 "rationale": "Line 17 (~1169 chars) is a §1 wall; breaking before the escape sentence separates the critique from the design response."97 },98 {99 "id": "rec-13",100 "category": "break-paragraph",101 "severity": "medium",102 "current": "for that turn only [1]. The frontier-variance curriculum is a homeostatic selection regulator,",103 "recommended": "for that turn only [1].\n\nThe frontier-variance curriculum is a homeostatic selection regulator,",104 "rationale": "Line 51 (~1316 chars) joins the mutation point and the curriculum point; breaking before the curriculum sentence separates two distinct GA-mapping claims."105 },106 {107 "id": "rec-14",108 "category": "break-paragraph",109 "severity": "low",110 "current": "and the trainer need *zero* changes, and `ModalSpawnExecutor` is the working existence proof [41].",111 "recommended": "and the trainer need *zero* changes, and `ModalSpawnExecutor` is the working existence proof [41].\n",112 "rationale": "Line 165 (~1021 chars) runs the ADR-005 framing and the AWS S3 mapping together; a trailing break after the existence-proof sentence eases the §8 lede paragraph."113 },114 {115 "id": "rec-15",116 "category": "bold-keyterms",117 "severity": "high",118 "current": "This upgrade from teacher-plurality to execution-oracle fitness is **the single most important change** and the one the corpus most strongly supports.",119 "recommended": "This upgrade from **teacher-plurality to execution-oracle fitness** is **the single most important change** and the one the corpus most strongly supports.",120 "rationale": "Bolds the load-bearing term \"teacher-plurality to execution-oracle fitness\" so a skimmer sees the report's central upgrade."121 },122 {123 "id": "rec-16",124 "category": "bold-keyterms",125 "severity": "high",126 "current": "A literal per-turn N-way tree is O(N^D) and economically fatal — ungated, a branching trace prices around $64 versus $0.98 flat [6].",127 "recommended": "A literal per-turn N-way tree is **O(N^D)** and economically fatal — ungated, a branching trace prices around **$64 versus $0.98 flat** [6].",128 "rationale": "Bolds the cost-blowup complexity and the headline price figures a skimmer needs to grasp why divergence-gating is mandatory."129 },130 {131 "id": "rec-17",132 "category": "bold-keyterms",133 "severity": "high",134 "current": "so collapse to a single rollout — turning O(N^D) into roughly O(N · decision-points) [6].",135 "recommended": "so collapse to a single rollout — turning O(N^D) into roughly **O(N · decision-points)** [6].",136 "rationale": "Bolds the target complexity after gating, the key quantitative payoff of the divergence-gated design."137 },138 {139 "id": "rec-18",140 "category": "bold-keyterms",141 "severity": "high",142 "current": "policy.\"** \"Prune versus train-on-all\" is a false binary.",143 "recommended": "policy.\"** **\"Prune versus train-on-all\" is a false binary.**",144 "rationale": "Bolds the §4 reframe conclusion so the skimmer catches the central thesis that the prune/train-on-all dichotomy is false."145 },146 {147 "id": "rec-19",148 "category": "bold-keyterms",149 "severity": "medium",150 "current": "and reaches 50.40% on SWE-bench Verified after three iterations [10].",151 "recommended": "and reaches **50.40% on SWE-bench Verified** after three iterations [10].",152 "rationale": "Bolds the headline Socratic-SWE pass-rate so the closest published analogue's result is scannable."153 },154 {155 "id": "rec-20",156 "category": "bold-keyterms",157 "severity": "medium",158 "current": "reaching 65.8% on SWE-bench Verified — crucially training *on all* trajectories for the world-model head",159 "recommended": "reaching **65.8% on SWE-bench Verified** — crucially training *on all* trajectories for the world-model head",160 "rationale": "Bolds the CWM existence-proof pass-rate, a key statistic supporting train-on-all for the world-model head."161 },162 {163 "id": "rec-21",164 "category": "bold-keyterms",165 "severity": "medium",166 "current": "across 2,695 networks mean causal fidelity is 0.49 (only 2.5% exceed 0.70), and at high dimension (N=100) the optimal encoder becomes causally blind (~1e-8) *while achieving 92% lower prediction error* [18].",167 "recommended": "across 2,695 networks **mean causal fidelity is 0.49** (only 2.5% exceed 0.70), and at high dimension (N=100) the optimal encoder becomes **causally blind (~1e-8) *while achieving 92% lower prediction error*** [18].",168 "rationale": "Bolds the two decisive predictive-causal-gap statistics that justify measuring foresight rather than next-state accuracy."169 },170 {171 "id": "rec-22",172 "category": "bold-keyterms",173 "severity": "medium",174 "current": "is the kill ablation: if it is ≈0, the token is a no-op and is cut",175 "recommended": "is **the kill ablation**: if it is ≈0, the token is a no-op and is cut",176 "rationale": "Bolds \"the kill ablation\" so the skimmer registers Foresight@k as the decisive cut criterion for the world-model head."177 },178 {179 "id": "rec-23",180 "category": "bold-keyterms",181 "severity": "medium",182 "current": "DeepSWE 42.2% Pass@1, 59% with test-time scaling, from pure outcome RL — stronger-teacher SFT *hurt* [43]",183 "recommended": "**DeepSWE 42.2% Pass@1**, 59% with test-time scaling, from pure outcome RL — stronger-teacher SFT *hurt* [43]",184 "rationale": "Bolds the DeepSWE headline figure, the incumbent baseline every later phase must beat at equal compute."185 },186 {187 "id": "rec-24",188 "category": "bold-keyterms",189 "severity": "medium",190 "current": "The strongest argument that this is buildable is that the substrate already exists — roughly nine-tenths of it.",191 "recommended": "The strongest argument that this is buildable is that the substrate already exists — **roughly nine-tenths of it**.",192 "rationale": "Bolds the reuse-fraction claim that anchors the entire §6 reuse-vs-build ledger."193 },194 {195 "id": "rec-25",196 "category": "bold-keyterms",197 "severity": "medium",198 "current": "whether the divergence-gated tree beats an equal-budget outcome-only GRPO baseline on long-horizon tasks — and it has never been run.",199 "recommended": "whether the **divergence-gated tree beats an equal-budget outcome-only GRPO baseline on long-horizon tasks** — and **it has never been run**.",200 "rationale": "Bolds the program's single most important unrun experiment so the skimmer catches the central open question of §7."201 },202 {203 "id": "rec-26",204 "category": "bold-keyterms",205 "severity": "medium",206 "current": "This is the single biggest architectural payoff of the object-store design on Kubernetes.",207 "recommended": "This is **the single biggest architectural payoff** of the object-store design on Kubernetes.",208 "rationale": "Bolds the headline architectural claim that gang scheduling is unneeded for inter-replica DiLoCo sync."209 },210 {211 "id": "rec-27",212 "category": "bold-keyterms",213 "severity": "low",214 "current": "First, `strip_thinking` must be `False`: ~67% of real Claude Code error-recovery turns are pure thinking, and stripping them yields empty SDPO masks that silently collapse two-thirds of the channel's supervision sites",215 "recommended": "First, `strip_thinking` must be `False`: **~67% of real Claude Code error-recovery turns are pure thinking**, and stripping them yields empty SDPO masks that silently collapse two-thirds of the channel's supervision sites",216 "rationale": "Bolds the 67%-thinking statistic that makes strip_thinking=False a load-bearing repo configuration fact."217 },218 {219 "id": "rec-28",220 "category": "bold-keyterms",221 "severity": "low",222 "current": "Calibration (ECE/Brier on the predicted-outcome head) is primary, because the documented failure is over-confidence; next-state accuracy is a secondary diagnostic.",223 "recommended": "**Calibration (ECE/Brier on the predicted-outcome head) is primary**, because the documented failure is over-confidence; next-state accuracy is a secondary diagnostic.",224 "rationale": "Bolds the primary-measurement decision (calibration over next-state accuracy), a load-bearing methodological choice in §2."225 },226 {227 "id": "rec-29",228 "category": "split-sentence",229 "severity": "medium",230 "current": "The divergence tree has a rigorous backbone: sibling A and B from a shared parent reaching different *executed* outcomes is a model-free Monte-Carlo counterfactual credit estimate, low-variance because the shared parent differences out the baseline — a group-relative/leave-one-out argument (Tree-GRPO [44]) — which the executed-sibling structure then approximates non-parametrically for the stronger, hindsight-conditioned variant that learned counterfactual-credit methods (CCA [33]) achieve with a learned hindsight model, and min-form/bottleneck-localized because the credit-bearing step is the earliest node where sibling subtrees separate [33].",231 "recommended": "The divergence tree has a rigorous backbone: sibling A and B from a shared parent reaching different *executed* outcomes is a model-free Monte-Carlo counterfactual credit estimate, low-variance because the shared parent differences out the baseline — a group-relative/leave-one-out argument (Tree-GRPO [44]). The executed-sibling structure then approximates non-parametrically the stronger, hindsight-conditioned variant that learned counterfactual-credit methods (CCA [33]) achieve with a learned hindsight model, and it is min-form/bottleneck-localized because the credit-bearing step is the earliest node where sibling subtrees separate [33].",232 "rationale": "This ~80-word run-on survived polish; splitting at the Tree-GRPO clause yields two readable sentences without losing any clause."233 }234]235 