Codeseys/composer-replication-framework
0
1# altered-minds × Composer Replication Framework2 3**Status**: Tie-in design doc.4**Date**: 2026-05-26 (Wave 13)5**Source workstream**: `llm-mental-alterations` (formerly Codeseys/llm-mental-alterations6on HF; user has indicated a rename to `altered-minds`)7 8## What altered-minds is studying9 10From the user's existing wiki notes (`~/wiki/projects/llm-mental-alterations.md`):11 12- Fine-tuning Llama-3.1-8B with **personality SFT** induces a depression/13 anxiety cognitive-distortion signature on MMLU `moral_scenarios`:14 - Class 3 ("both fine") collapses **−31.1pp**15 - Class 0 ("both wrong") improves **+4.6pp**16 - Multi-seed reproducible (4/4 seeds, n=895)17 - 18% of base-correct items broken18- Other domains affected: `high_school_chemistry +4.2pp`,19 `machine_learning +4.9pp` (reliably improved).20- H-3 Gemma-MoE hypothesis is deferred (Hopper-only).21- Spend so far: $9.75 / $400 budget.22 23The headline question driving the workstream is roughly:24**"What measurable cognitive alterations does personality-style SFT25introduce, and can we recover or sharpen them via downstream RL?"**26 27## Why this framework is the right second-stage workstream28 29altered-minds today is an **SFT-only** pipeline. A typical run:301. Take a base model (Llama-3.1-8B).312. Apply personality SFT.323. Evaluate on MMLU + alteration-specific probes.334. Document the alteration signature.34 35The Composer Replication Framework, by design, is a **post-SFT36reinforcement-learning framework**. It can take any HF model — including37an altered-minds-altered model — and apply:38- **GRPO** with verifiable rewards39- **SDPO/OPSD** self-distillation against the altered model's hint-40 conditioned forward passes41- **Trace-replay DPO** against N external teachers42 43That gives altered-minds three orthogonal axes of investigation it doesn't44currently have:45 46| Axis | What changes | What we learn |47|---|---|---|48| **GRPO with verifiable reward** | Train the altered model on math/code where ground truth is checkable | Does the alteration's "personality" persist under task-driven RL, or does it wash out? |49| **SDPO against the altered model's own hints** | Self-distillation — the altered model teaches itself with hint-conditioned forward passes | Can we **sharpen** the alteration without further SFT? |50| **Trace-replay DPO with frontier teachers** | The altered model rolls out, frontier teachers replay the same prompts, disagreement → DPO pairs | Where does the altered model **disagree** with frontier consensus? Are those disagreements correlated with the cognitive-distortion signature? |51 52The **third** axis is the most interesting for altered-minds specifically.53The framework's `replay_trace` + `extract_dpo_pairs` produce, by construction,54a dataset of "altered-model output" vs "frontier-consensus output" for any55prompt distribution. If the altered model's depression/anxiety signature56shows up in moral_scenarios, then the trace-replay output on57moral-scenario prompts is **a measurable corpus of the alteration**.58 59## Concrete plan: altered-minds-RL spike60 61### Phase 1 — model selection62Pick the altered-minds checkpoint that produced the strongest signature63(per the user's notes: the multi-seed Llama-3.1-8B personality-SFT run64where moral_scenarios class 3 collapsed −31.1pp).65 66### Phase 2 — domain-specific replaysim67 68Run `composer_replication.replaysim.replay_and_normalize_trace` against:69- A held-out moral_scenarios test set (the alteration locus)70- A held-out high_school_chemistry test set (where altered-minds *improved*)71- A held-out general MMLU baseline72 73Teachers: framework defaults (Claude Opus 4.7, GPT-5, DeepSeek V4 Pro).74This produces **three normalized DPO datasets** capturing where the75altered model disagrees with frontier consensus on each domain.76 77Cost estimate: ~$0.98/trace × 100 prompts × 3 domains ≈ **$300**.78Fits inside the user's existing $400 altered-minds budget.79 80### Phase 3 — GRPO with the framework81 82> **⚠️ SUPERSEDED by [ADR-013](adrs/ADR-013-lma-integration-channel-ladder.md).**83> The original all-channels-on combined recipe (α=0.2, β=0.4) is **not used**.84> A cross-family research critique (2026-05-29) found a combined-first run85> **scientifically uninterpretable**: it confounds four effects (task RL,86> self-distillation of altered reasoning, frontier-teacher imitation, KL87> anchoring), so any observed change in the alteration signature cannot be88> attributed to a channel. Worse, **SDPO against the altered model's own89> hint-conditioned forward pass is the channel most likely to AMPLIFY the90> distortion** (teacher == student-family; if hints add no independent91> information, the optimum is to imitate the altered conditional distribution,92> sharpening a soft bias into a hard preference). SDPO here is therefore an93> *experimental intervention*, not a benign stabilizer.94 95**Use the isolated-channel ladder (ADR-013) instead** — sweep arms A0–A4 with96identical seeds/prompts so each channel's effect is attributable:97 98| Arm | alpha_sdpo | beta_replay | Purpose |99|---|---|---|---|100| A0 | — | — | altered SFT, no RL (control) |101| A1 | 0.0 | 0.0 | GRPO-only baseline |102| A2 | **0.02** | 0.0 | +SDPO small (amplification probe) |103| A3 | 0.0 | **0.05** | +replay-DPO small (washout probe) |104| A4 | 0.02 | 0.05 | combined — only after A1–A3 interpretable |105 106`kl_beta=0.02` (KL-to-altered-init) on every RL arm, adaptive to 0.01–0.03107nats/token; hard-stop/LR-cut if KL > ~0.08. The framework provides the ladder108via `composer_replication.integrations.altered_minds.channel_ladder_configs()`,109the structured `MMLUFormatReward` (scores the final answer letter + format110only — never rationale style, so distorted-but-persuasive reasoning is not111rewarded), and `dual_kl_logger` (logs KL-to-altered-init **and** KL-to-base each112step — the washout-vs-amplification instrument).113 114Train for ~500 steps per arm on a single GPU. **Runnability today (2026-06):**115only **A1 (GRPO-only)** has a real Modal runner — [ADR-014](adrs/ADR-014-policy-optimization-objective-menu.md)116records that "the A1 run used `dr_grpo`" and that wiring the `objective=` menu through the117rest of the ladder runners is an open follow-up (Qwen-0.5B feasibility-test confirmed;118for Llama-8B use Modal + the framework's `ServerlessExecutor` per ADR-005 — local 5090119is too small). **A2 (SDPO) / A3 (replay-DPO) / A4 (combined) are scaffold + plan-builder120only**: running them on a real 8B checkpoint additionally needs a real error-trace SDPO121dataset, a replay-DPO preference corpus, and an A100 entrypoint that don't exist yet —122none of those is a closed artifact today. The real 8B/LMA-checkpoint run is *additionally*123**user-gated** (it spends grant budget). [ADR-013](adrs/ADR-013-lma-integration-channel-ladder.md)124ships the ladder scaffolding + the A1 capability, proven CPU-only on a small model125(`examples/altered_minds_channel_ladder/`); its sole remaining acceptance-gate box is that126user-gated real-spend go/no-go.127 128> **strip_thinking × SDPO foot-gun (A2/A4).** When the SDPO arms become runnable on real129> agent traces, SDPO REQUIRES `strip_thinking=False`: ~67% of error-recovery turns are130> pure thinking, so stripping them yields empty SDPO masks (the channel silently131> contributes nothing). Keep thinking tokens in the context for any SDPO-active arm.132 133### Phase 4 — re-evaluate134 135Re-run the same MMLU + alteration probes used originally on the136**post-RL** model. Three outcomes are possible:137 138| Outcome | Interpretation |139|---|---|140| Alteration signature persists at same magnitude | The alteration is robust to task-driven RL — useful as a lower bound on its "depth" |141| Alteration signature attenuates | Task-driven RL washes out personality-SFT — useful for understanding alteration brittleness |142| Alteration signature **amplifies** on channel-2-only ablation | SDPO is reinforcing the alteration; rare and significant — would be a publishable finding |143 144### Phase 5 — Decoupled DiLoCo for multi-personality experiments145 146Once a single altered-minds-RL run works, the framework's serverless147DiLoCo (ADR-005) lets us run **N personality-altered models in parallel148across Modal/HF Jobs**, with their pseudo-gradients pooled via object149storage. This becomes the natural sweep over personality types150(depression vs anxiety vs grandiose vs ...) at minimal incremental151infrastructure cost.152 153## Repo layout proposal154 155The Composer Replication Framework is intentionally generic. The156altered-minds-specific RL spike should live as a separate repo or157subdirectory **using** the framework, not inside it:158 159```160altered-minds/ # the renamed llm-mental-alterations repo161 composer_replication_runs/ # NEW162 moral_scenarios_replay.py # uses composer_replication.replaysim163 train_grpo.py # uses composer_replication.trainer164 eval_post_rl.py # standard altered-minds eval165 recipes/166 altered_minds.yaml # data-juicer recipe — symlinks/copies167 # composer_replication's default + adds168 # MMLU-format-aware ops169```170 171The framework provides the algorithm + infrastructure. The altered-minds172repo owns the experimental narrative + results.173 174## Open questions for the user175 176Before we proceed to Phase 1:177 1781. **Confirm the rename**: the wiki memory says `llm-mental-alterations`179 on HF; user wants `altered-minds` — should we rename the HF repo?1802. **Budget allocation**: the $300 trace-replay cost (Phase 2) eats most181 of the remaining $390 altered-minds budget. Is that acceptable, or182 should we use only one domain (moral_scenarios) for $100?1833. **GPU venue for Phase 3**: 8B-model RL on single-GPU is feasible on184 the user's RTX 5090 (32GB) for short runs, OR we use Modal A100s for185 a more aggressive run. Preference?186 187## References188 189- altered-minds workstream wiki: `~/wiki/projects/llm-mental-alterations.md`190- Framework ADRs: docs/adrs/ADR-001 through ADR-007191- Framework V1-V8 brief coverage: docs/V1_V8_COVERAGE.md192- Self-distillation landscape: docs/research/SELF_DISTILLATION_LANDSCAPE.md193 (relevant: TAID's annealed-teacher schedule could test "alteration194 recovery" by interpolating between altered-init and base-teacher)195 