Codeseys/composer-replication-framework
0
1"""altered_minds — framework-side, generic LMA integration glue (ADR-013).2 3This package is the *model-agnostic* scaffold that lets the Composer Replication4Framework drive the sister project llm-mental-alterations (LMA): take a5personality-altered SFT checkpoint and apply the framework's 3-channel RL to ask6whether task-driven RL washes out, preserves, or AMPLIFIES the alteration's7cognitive-distortion signature.8 9Nothing here loads an LMA checkpoint, calls Modal, or spends budget — that is10explicitly user-gated (ADR-013 "out of scope"). This package provides:11 12 - ``MMLUFormatReward`` : structured-answer reward (final letter + format13 only; never rationale style). Plus14 ``randomize_options`` and a logged option15 distribution so an "always C" exploit is16 detectable.17 - ``dual_kl_logger`` : logs KL(policy||altered_init) AND KL(policy||base)18 each step — the washout/amplification instrument.19 - ``channel_ladder_configs``: the A0-A4 isolated-channel ladder that REPLACES20 the old combined alpha=0.2/beta=0.4 recipe.21 22See docs/adrs/ADR-013-lma-integration-channel-ladder.md.23"""24from __future__ import annotations25 26from composer_replication.integrations.altered_minds.kl_logging import (27 dual_kl_logger,28 token_mean_kl,29)30from composer_replication.integrations.altered_minds.ladder import (31 LADDER_KL_BETA,32 channel_ladder_configs,33)34from composer_replication.integrations.altered_minds.reward import (35 MMLUFormatReward,36 parse_final_answer,37 randomize_options,38)39 40__all__ = [41 "MMLUFormatReward",42 "parse_final_answer",43 "randomize_options",44 "dual_kl_logger",45 "token_mean_kl",46 "channel_ladder_configs",47 "LADDER_KL_BETA",48]49 