Codeseys/composer-replication-framework
0
1# TROUBLESHOOTING — Wave 142 3This document catalogs every Wave-14-known failure mode in the Composer4Replication Framework, along with how to diagnose, fix, and verify each5one. It is intentionally surgical: the surface area added in Waves 12–146(SimPO/TAID/Entropy-OPD distillation kwargs, the PRIME-RL composer-loss7adapter, the serverless DiLoCo `MockManager` + `ObjectStoreAllReduce`8path, and the data-juicer-backed replaysim normalizer) introduced new9ways for users to trip themselves up. Each failure mode here is something10a maintainer has actually seen or anticipated during the cross-model11review of Wave 14.12 13If you hit something not covered below, jump to the14[How to file a bug report](#how-to-file-a-bug-report) section at the end —15the template there gives a maintainer everything they need to reproduce.16 17---18 19## Common things to check first20 21Before reading any further, run through this checklist. ~80% of "framework22broken" reports turn out to be one of these:23 241. **Python version.** The framework targets Python 3.10–3.12. The25 `pyproject.toml` `target-version` is `py310`. If you are on 3.13+,26 transitive deps (notably Ray, pulled in by data-juicer) may not yet27 ship wheels and will try to build from source. Run `python --version`.28 292. **Fresh virtual environment.** Mixing the framework into an existing30 environment that already has `torch`, `transformers`, `trl`, or31 `torchft` pinned to incompatible versions is the #1 source of import-32 time errors. Create a new venv: `python -m venv .venv && source33 .venv/bin/activate && pip install -e .[dev]`.34 353. **Editable install.** Most contributors run `pip install -e .` so36 that local edits to `composer_replication/` are picked up. If you37 `pip install composer-replication` from a registry instead, your38 edits to the source tree will be ignored. Confirm with39 `pip show composer-replication | grep Location`.40 414. **Optional extras.** Several modules are optional-dep gated:42 - `[replay]` — adds `httpx` (used for OpenRouter teacher calls).43 - `[train]` — adds TRL, peft, accelerate, datasets (production GRPO).44 - `[replaysim]` — adds `data-juicer` (and via it, Ray as a transitive).45 - `[serverless]` — adds `fsspec`. For non-local rendezvous URIs you46 also need a backend-specific fsspec adapter (see Failure Mode 5).47 - `[dev]` — adds `pytest`, `ruff`, etc.48 If you see `ModuleNotFoundError: No module named 'data_juicer'`, you49 forgot the extra. Install with `pip install -e .[replaysim]`.50 515. **Run the test suite first.** Before debugging anything, run the52 subset of tests touching the area you care about:53 ```54 pytest composer_replication/tests/ # core compose_loss55 pytest composer_replication/distillation/tests/ # SimPO / TAID / OPD56 pytest composer_replication/recipes/prime_rl/tests/ # PRIME-RL adapter57 pytest composer_replication/diloco/serverless/tests/ # MockManager + DiLoCo58 pytest composer_replication/replaysim/tests/ # data-juicer normalizer59 ```60 If any green test fails for you locally, the problem is environmental61 — fix that before digging into your own code.62 636. **Read the docstring of the symbol you're calling.** Wave 1464 docstrings are written to be the first line of documentation. The65 `compose_loss` docstring (`composer_replication/loss.py`) lists every66 required and optional input key. The `MockManager` docstring67 enumerates the torchft surface methods it implements.68 69---70 71## Failure modes72 73### 1. `pip install -e .[replaysim]` hangs or fails on Python 3.12 with a Ray-related path error74 75**SYMPTOM.** Installing the `[replaysim]` extra (which pulls76`data-juicer`) triggers a transitive install of Ray. On Python 3.12, the77first `import ray` (often during `pip` build hooks or the first time78data-juicer is loaded) fails with messages mentioning79`/tmp/ray/session_*` paths, missing `pyarrow` symbols, or `OSError:80[Errno 2] No such file or directory: '/dev/shm/ray-...'` inside Docker.81 82**DIAGNOSIS.** `data-juicer` declares `ray` as a transitive dependency.83On Python 3.12 the wheel matrix is incomplete for some Ray versions, and84Ray's first-import probes `/dev/shm` and `/tmp/ray` for its session85state. In a sandboxed container, restricted CI runner, or WSL86environment with a non-default `/tmp`, those probes fail. Wave 1487subagent T2 hit this in CI and worked around it by pinning Ray and by88making sure `/tmp` exists and is writable.89 90**FIX.**91- Prefer Python 3.11 if you're on 3.12+ and don't need 3.12 features.92- If you must stay on 3.12, ensure `/tmp` is writable and pre-create the93 session directory: `mkdir -p /tmp/ray && chmod 1777 /tmp/ray`.94- In Docker, mount a real tmpfs at `/dev/shm`:95 `docker run --shm-size=2g …`.96- If you don't need replaysim normalization, you can skip the extra97 entirely. The `DJNormalizer(skip_dj=True)` passthrough (see98 `composer_replication/replaysim/normalize.py:165`) does not import99 `data_juicer` and therefore does not import Ray.100 101**VERIFICATION.** The skip-dj passthrough is exercised by102`test_dj_normalizer_skip_dj_passthrough` and103`test_dj_normalizer_skip_dj_preserves_count` in104`composer_replication/replaysim/tests/test_replaysim.py`. Both run105without `data_juicer` installed:106 107```108pytest composer_replication/replaysim/tests/test_replaysim.py::test_dj_normalizer_skip_dj_passthrough -xvs109```110 111If that passes in your environment, your `[replaysim]`-less install is112healthy — only the full data-juicer code path requires Ray.113 114---115 116### 2. `compose_loss` produces wrong-looking numbers when combining new kwargs117 118**SYMPTOM.** You pass several Wave-14 distillation kwargs to119`compose_loss` (e.g. `dpo_variant="simpo"`, `sdpo_wrapper="taid"`,120`taid_schedule_step=0`, `simpo_beta=2.0`, `entropy_opd_h_max=…`), and121the loss curve looks wrong: NaNs, identically-zero `sdpo_jsd` channel,122or a `total` that is bit-different from your reference run with no123distillation kwargs at all.124 125**DIAGNOSIS.** `compose_loss` now has 13 keyword arguments and the126contract between them is non-trivial. Subagent T1's review identified127three combinations that look reasonable but are unsupported:128- Passing `taid_schedule_step` without `taid_total_steps` (or vice129 versa). The function raises `ValueError` clearly, but the message can130 scroll past in noisy logs.131- Passing `dpo_variant="simpo"` while still supplying132 `dpo_chosen_ref_logprobs`. Those keys are **silently ignored** —133 SimPO is reference-free.134- Passing `sdpo_wrapper="taid"` without supplying either135 `student_init_logits` OR `student_init_input_ids` in `inputs`. The136 function will fall back to a forward pass through the (possibly137 drifted) live model, which is a footgun late in training (see Failure138 Mode 8).139 140**FIX.** Read the docstring at the top of141`composer_replication/loss.py` (lines 25–39 list the three pluggable142losses and their preconditions). The general rule:143 144```python145from composer_replication import compose_loss146 147# Defaults (no distillation knobs) reproduce legacy 3-channel composition bit-exact.148out = compose_loss(model, inputs)149 150# To opt into SimPO, pass dpo_variant ONLY. Do not pass ref-logprob keys.151out = compose_loss(model, inputs, dpo_variant="simpo",152 simpo_beta=2.0, simpo_gamma=1.0)153 154# To opt into TAID, pass BOTH schedule_step AND total_steps, AND make sure155# inputs["student_init_logits"] is populated (see Failure Mode 8).156out = compose_loss(model, inputs, sdpo_wrapper="taid",157 taid_schedule_step=step, taid_total_steps=total_steps)158```159 160Setting all 13 kwargs to their defaults is **bit-exact equivalent** to161the pre-Wave-13 3-channel loss; if your defaults call gives different162numbers than your old code, file a bug.163 164**VERIFICATION.** The bit-exact equivalence and every supported165combination is locked in by the 11 integration tests in166`composer_replication/tests/test_compose_loss_integration.py`. The most167important ones:168- `test_defaults_bit_exact_with_legacy_kwargs` — passing the new kwargs169 at their defaults is identical to legacy.170- `test_simpo_does_not_require_ref_logprobs` — SimPO works with the171 ref-logprob keys absent from `inputs`.172- `test_taid_alpha_one_recovers_sdpo` — TAID with `alpha_min=alpha_max=1`173 reproduces standard SDPO.174- `test_taid_requires_schedule_step` / `test_taid_requires_total_steps` —175 the partial-config error path.176 177```178pytest composer_replication/tests/test_compose_loss_integration.py -xvs179```180 181---182 183### 3. `MockManager` works today but silently breaks after a torchft upgrade184 185**SYMPTOM.** Your serverless DiLoCo run starts, the first outer round186completes, and then `torchft.DiLoCo` raises an `AttributeError` on187something like `_use_async_quorum`, `should_commit`, or188`current_step` — or worse, it silently uses the wrong sync semantics.189 190**DIAGNOSIS.** `MockManager` is a duck-typed shim that mirrors191`torchft.Manager` rather than subclassing it. The surface it implements192is enumerated in the docstring at193`composer_replication/diloco/serverless/allreduce.py:215`:194 195> Methods/attributes DiLoCo touches: `allreduce`, `should_commit`,196> `start_quorum`, `current_step`, `disallow_state_dict_read`,197> `allow_state_dict_read`, `register_state_dict_fn`, `_use_async_quorum`198> (attribute), `num_participants`, `rank`.199 200The two **private** members in that list — `_use_async_quorum` and the201internal `current_step` counter — are private torchft API that may be202renamed without notice in any torchft minor release. Wave 14 subagent203T3 specifically called this out: "If torchft renames `_use_async_quorum`204to anything else, MockManager silently breaks because there is nothing205holding the contract beyond a string."206 207**FIX.**208- **Pin torchft.** In `pyproject.toml` keep your torchft version pinned209 to a known-good range (e.g. `torchft>=0.2,<0.4`). When you need to210 upgrade, do so deliberately and re-run the integration tests below211 before merging.212- **Watch the deprecation warning.** Wave 14 sets up a clear path to213 warn if `_use_async_quorum` is read on a fresh instance — see the214 comment at `allreduce.py:255`.215- **Don't pass an arbitrary torchft branch.** If you've patched torchft216 locally, the `MockManager` may need updating in lockstep. The217 surface-compatibility tests below will catch this in CI.218 219**VERIFICATION.** The full DiLoCo × MockManager surface is exercised by:220- `test_mock_manager_shape_compat` in221 `composer_replication/diloco/serverless/tests/test_serverless_local.py`222 — sanity check that all expected methods/attributes exist.223- `test_mockmanager_has_full_diloco_call_surface` in224 `composer_replication/diloco/serverless/tests/test_serverless_diloco_integration.py`225 — runs an end-to-end outer round through real torchft `DiLoCo`,226 hitting every method on the surface list above.227- `test_mockmanager_diloco_outer_round_completes` — full one-round228 smoke ending in a successful outer SGD step.229 230If any of these tests turn red after a torchft bump, **do not ship**:231inspect the new torchft Manager surface and update `MockManager`232to match.233 234```235pytest composer_replication/diloco/serverless/tests/test_serverless_diloco_integration.py -xvs236```237 238---239 240### 4. SimPO loss curve looks like noise241 242**SYMPTOM.** You wired in `dpo_variant="simpo"`, the run starts, and243the `trace_replay_dpo` channel either drifts to large negative values244(→ `total` blows up) or oscillates with much higher variance than245standard DPO. The loss curve "looks like noise."246 247**DIAGNOSIS.** SimPO uses **average per-token log-probability**248(`Σ logπ(c_t) / |c|`), not sum log-prob. From the SimPO docstring249(`composer_replication/distillation/simpo.py:11–18`):250 251> SimPO drops the reference-policy term, replaces it with a target252> margin γ, and uses **average sequence log-probability instead of253> sum**. […] L_SimPO = -log σ( β · [avg_logπ(c) - avg_logπ(r)] - γ )254 255If you compute `chosen_logprobs.sum()` (or any unmasked aggregation) and256hand it to SimPO as `chosen_avg_logprobs`, the loss is undefined: β=2.0257times a sum-log-prob is on a totally different scale than β=2.0 times an258average. The result looks plausible per-batch but the optimum is259nowhere near the dataset's true preference signal.260 261**FIX.** Use the helper262`composer_replication.distillation.simpo.avg_sequence_logprob`:263 264```python265from composer_replication.distillation.simpo import (266 simpo_loss, avg_sequence_logprob,267)268 269chosen_avg = avg_sequence_logprob(chosen_logprobs, chosen_response_mask)270rejected_avg = avg_sequence_logprob(rejected_logprobs, rejected_response_mask)271loss = simpo_loss(chosen_avg, rejected_avg, beta=2.0, gamma=1.0)272```273 274The mask is **1 on response tokens, 0 on prompt+padding** — same275convention as the rest of the framework. If you must roll your own276aggregation, divide by `response_mask.sum(dim=-1).clamp_min(1.0)`,277not by `response_mask.shape[-1]`.278 279**VERIFICATION.** The avg-vs-sum semantics are pinned by280`test_avg_sequence_logprob` in281`composer_replication/distillation/tests/test_distillation_losses.py`,282which constructs known per-token log-probs and asserts the helper283returns the correct per-sequence average. The end-to-end SimPO284loss-shape check is `test_simpo_loss_returns_scalar` in the same file.285 286```287pytest composer_replication/distillation/tests/test_distillation_losses.py::test_avg_sequence_logprob -xvs288pytest composer_replication/distillation/tests/test_distillation_losses.py::test_simpo_loss_lower_for_better_separation -xvs289```290 291---292 293### 5. `ObjectStoreAllReduce` works locally but fails on `s3://` at first allreduce294 295**SYMPTOM.** You construct296`ObjectStoreAllReduce(uri="s3://my-bucket/run42/", rank=0,297world_size=4)`. The constructor succeeds. The first call to298`allreduce(tensor, name="...")` raises `ImportError: Install s3fs to299access S3` or `botocore.exceptions.NoCredentialsError: Unable to locate300credentials`.301 302**DIAGNOSIS.** `ObjectStoreAllReduce` uses fsspec to reach the303backend, but **fsspec only ships protocol stubs, not adapters**. The304constructor doesn't know which protocol you'll use and doesn't305eagerly validate, so it accepts any URI. The `s3://` adapter requires:3061. The `s3fs` package (`pip install s3fs`), which is **not** in the307 default `[serverless]` extra.3082. Working AWS credentials (env vars, `~/.aws/credentials`, IAM role,309 or whatever your environment normally provides to boto3).310 311The same is true for `gs://` (`gcsfs`), `az://` (`adlfs`), and312`hf://` (`huggingface_hub`'s fsspec integration, which is included if313you have `huggingface_hub` installed).314 315**FIX.**316- Install the right adapter alongside the framework:317 ```318 pip install s3fs # for s3://319 pip install gcsfs # for gs://320 pip install adlfs # for az://321 ```322- Verify credentials work outside the framework first:323 ```324 python -c "import s3fs; print(s3fs.S3FileSystem().ls('my-bucket'))"325 ```326- If you're running on Modal/HF Jobs, set the credentials as Modal327 secrets / HF Jobs env vars in the executor config — not in your328 local shell.329 330The constructor could in principle perform an eager probe (e.g. a331`HEAD` on the rendezvous prefix) to fail fast at init time. Wave 14332deliberately did not add this because it adds a network round-trip on333every replica startup. If you want pre-flight validation in your334training script, call `fsspec.filesystem(protocol).ls(uri)` yourself335before constructing the manager.336 337**VERIFICATION.** The `file://` and bare-path code paths — the only338ones that don't need an extra adapter — are exercised by:339- `test_object_store_allreduce_local_paths_create_dir`340- `test_object_store_allreduce_world_size_1_passthrough`341- `test_object_store_allreduce_round_id_increments`342 343…all in344`composer_replication/diloco/serverless/tests/test_serverless_local.py`.345If those pass and your `s3://` URI fails, the framework is fine and346your fsspec adapter or credentials are the problem.347 348```349pytest composer_replication/diloco/serverless/tests/test_serverless_local.py -xvs350```351 352---353 354### 6. Custom replaysim recipe drops every record (or crashes data-juicer)355 356**SYMPTOM.** You wrote a custom replaysim YAML recipe modeled on357`composer_replication/recipes/replaysim/default.yaml`. It loads358without error, but every input DPO pair is dropped, OR data-juicer359raises `KeyError: 'text_key'`, OR it raises a complaint about360"expected str, got list" inside one of the filters.361 362**DIAGNOSIS.** Wave 14 fixed two related bugs in the *default* recipe363that custom-recipe authors will hit again. Both are documented in the364header comment at365`composer_replication/recipes/replaysim/default.yaml:21–35`:366 3671. **`text_keys` plural vs `text_key` singular.** The top-level368 dataset contract uses `text_keys: chosen` (plural). Each individual369 op uses `text_key: chosen` (singular). They are not interchangeable.370 data-juicer's dataset loader validates that the `text_keys` field371 exists on every record before any op runs; an op that uses372 `text_keys` instead of `text_key` is silently misconfigured.373 3742. **`chosen` / `rejected` as strings vs as list-of-dicts.**375 data-juicer ops like `text_length_filter`, `words_num_filter`,376 `special_characters_filter`, and `document_deduplicator` read a377 single string field. Pointing them at the chat-messages list378 (`chosen_messages`, `rejected_messages`) crashes or silently379 no-ops. The framework's `_dpo_pair_to_dj_record` keeps **both**380 shapes side-by-side: `chosen`/`rejected` (strings) for filter ops,381 and `chosen_messages`/`rejected_messages` (chat-messages list) for382 chat-aware ops + the `NormalizedDPOPair` round-trip.383 384**FIX.** Treat the default recipe as your starting template. Concretely:385- Always declare `text_keys: chosen` at the top.386- For every length/word/special-char op you add, duplicate it: once387 with `text_key: chosen`, once with `text_key: rejected`. (Each op388 takes only one `text_key` — see comment at lines 31–35 of389 `default.yaml`.)390- Never point a filter op at `chosen_messages` or `rejected_messages`.391 Those are list-of-dicts; only chat-aware ops accept that shape.392 393**VERIFICATION.** The two-shape contract is locked in by:394- `test_record_chosen_rejected_are_flat_strings_for_dj_text_ops` —395 asserts `chosen` and `rejected` are bare strings on every record396 produced by `_dpo_pair_to_dj_record`.397- `test_record_chosen_rejected_messages_carry_chat_shape` — asserts398 `chosen_messages` / `rejected_messages` exist as list-of-dicts.399- `test_dj_normalizer_e2e_default_recipe(tmp_path)` — runs the actual400 default recipe through real data-juicer end-to-end (skipped if401 `data_juicer` isn't importable).402 403…all in404`composer_replication/replaysim/tests/test_replaysim.py`. If those405pass and your custom recipe still drops everything, diff your YAML406against `default.yaml` until the two shapes align.407 408```409pytest composer_replication/replaysim/tests/test_replaysim.py -xvs410```411 412---413 414### 7. `ValueError: expected (seq,) shape, got (B, T)` from PRIME-RL composer_loss415 416**SYMPTOM.** You wired the PRIME-RL recipe into a training loop you417adapted from another framework (TRL, openrlhf, etc.), and on the very418first `loss_fn` call you get a `ValueError` mentioning shape419`(seq,)` versus `(B, T)`.420 421**DIAGNOSIS.** PRIME-RL calls its loss function **one sample at a422time**, with 1-D `(seq,)` tensors — not batched `(B, T)` tensors. The423recipe's docstring spells this out at424`composer_replication/recipes/prime_rl/composer_loss.py:16–30`:425 426> Note the **per-sample (seq,) shape** — PRIME-RL's runner calls the427> loss function one sample at a time, not on a batched (B, T) tensor.428 429Wave 14 fixed an earlier draft of the recipe that incorrectly assumed430`(B, T)`. The new version raises a clear `ValueError` if you hand it431the wrong shape, instead of silently broadcasting and producing432nonsense gradients. Users who are used to TRL or openrlhf — both of433which call the loss with batched tensors — see this on day one.434 435**FIX.**436- If you are running inside PRIME-RL via its `CustomLossConfig`, you437 don't need to do anything: PRIME-RL's runner produces `(seq,)`438 tensors and the recipe accepts them.439- If you are calling the recipe directly from your own runner, slice440 your batch into per-sample 1-D tensors before each call:441 ```python442 for b in range(B):443 inputs_b = LossInputs(444 trainer_logprobs=batched.trainer_logprobs[b],445 inference_logprobs=batched.inference_logprobs[b],446 advantages=batched.advantages[b],447 loss_mask=batched.loss_mask[b],448 teacher_logprobs=None if batched.teacher_logprobs is None449 else batched.teacher_logprobs[b],450 )451 loss = loss_fn(inputs_b, ...)452 ```453- If you genuinely need a batched API, write a thin wrapper around454 `loss_fn`. Don't patch the recipe — its shape contract is dictated455 by PRIME-RL, not by us.456 457**VERIFICATION.** The shape contract is pinned by two tests in458`composer_replication/recipes/prime_rl/tests/test_composer_loss.py`:459- `test_advantages_shape_validates_seq_accepted` — `(seq,)` succeeds.460- `test_advantages_shape_validates_bt_rejected` — `(B, T)` raises461 `ValueError`.462 463```464pytest composer_replication/recipes/prime_rl/tests/test_composer_loss.py -xvs465```466 467---468 469### 8. TAID can't run mid-training because `student_init_logits` is missing470 471**SYMPTOM.** You decide partway through a training run to enable472`sdpo_wrapper="taid"` (e.g. you read the TAID paper after step 2000473and want to retrofit). The next training step blows up — either with474a `KeyError` for `student_init_logits` / `student_init_input_ids`, or475with a strange-looking loss because the framework fell back to476re-running a forward pass through the *current* (drifted) model477instead of the init model.478 479**DIAGNOSIS.** TAID interpolates between the **student's distribution480at step 0** and the teacher's distribution. From the TAID docstring at481`composer_replication/distillation/taid.py:10–24`:482 483> TAID interpolates between an "identity" target (the student's own484> distribution at step 0) and the teacher's distribution, with the485> interpolation coefficient annealed from 0 → 1 over training.486 487That step-0 reference target has to come from somewhere. The framework488accepts it via either:4891. `inputs["student_init_logits"]` — a precomputed `(B, T, V)` tensor490 captured at training start (preferred for production), OR4912. `inputs["student_init_input_ids"]` — input ids for a frozen forward492 pass through `model`. **This assumes `model` has not yet drifted493 from init.** It is correct only at step 0 or in tests; in494 production it silently produces the wrong target.495 496If you forgot to capture the init logits at step 0, you cannot497faithfully use TAID mid-run.498 499**FIX.** Capture init logits at step 0 and persist them:500 501```python502# At step 0, before any optimizer.step() call:503with torch.no_grad():504 init_logits = model(input_ids=batch["input_ids"]).logits505 # Save to disk if you'll need them across restarts:506 torch.save(init_logits, "checkpoints/init_logits_batch0.pt")507 inputs["student_init_logits"] = init_logits508 509# Or, if you have a fixed eval probe set, capture init logits once510# for that fixed set and reuse them every step:511inputs["student_init_logits"] = cached_init_logits512```513 514If you genuinely have no step-0 snapshot, **TAID is not retrofittable**515to your run. Your options are:516- Restart from a checkpoint that *was* the step-0 model.517- Use a different distillation wrapper (`sdpo_wrapper="entropy_opd"`)518 that doesn't need init logits.519- Accept the bias from the live-model fallback path. Don't.520 521**VERIFICATION.** The precomputed-vs-live-fallback contract is exercised by:522- `test_taid_accepts_precomputed_student_init_logits` in523 `composer_replication/tests/test_compose_loss_integration.py` —524 passes precomputed logits and asserts the TAID-wrapped channel uses525 them.526- `test_taid_alpha_one_recovers_sdpo` — asserts that with527 `alpha_min=alpha_max=1.0` (i.e. pure teacher target, init logits528 ignored) TAID reproduces standard SDPO. If your training ignores529 init logits silently, *this* is the test that would have failed.530 531```532pytest composer_replication/tests/test_compose_loss_integration.py::test_taid_accepts_precomputed_student_init_logits -xvs533```534 535---536 537### 9. `ModalExecutor()` or `HFJobsExecutor()` raises `NotImplementedError` at construction538 539**SYMPTOM.** You write540`executor = ModalExecutor(app_name="my-app")` (or the HF Jobs541equivalent) in a production script and the constructor immediately542raises:543 544```545NotImplementedError: ModalExecutor is a v0 skeleton; full implementation pending.546Use LocalProcessExecutor for testing.547```548 549Same for `HFJobsExecutor`. This is at *init time*, not at the first550`launch_replicas` call.551 552**DIAGNOSIS.** Per ADR-005 the v0 release ships only the553`ServerlessExecutor` Protocol and the reference `LocalProcessExecutor`.554The Modal and HF Jobs implementations are **import-safe skeletons** —555the classes exist and you can `from … import ModalExecutor`, but556`__init__` raises `NotImplementedError` to prevent silent partial557behavior. See `modal.py:64` and `hf_jobs.py:64`.558 559This is intentional. We didn't want to ship a half-working Modal560executor that succeeds at `launch_replicas` and then silently fails561two-thirds of the way through `collect`.562 563**FIX.**564- Use `LocalProcessExecutor` for development, CI, and any single-host565 multi-process testing.566- For real cloud deployment in the v0 era, run your training script567 directly in Modal/HF Jobs by hand: write your own thin Modal568 function that constructs `MockManager(ObjectStoreAllReduce(uri,569 rank, world_size))` and runs the training loop. The skeleton570 docstrings at `modal.py:24–48` and `hf_jobs.py:26–49` show exactly571 the pattern.572- Watch the `BACKLOG.md` for v0 polish — the real implementations are573 scheduled.574 575**VERIFICATION.** That `LocalProcessExecutor` is fully functional and576correctly implements the Protocol is locked in by:577- `test_local_executor_runs_allreduce_across_replicas` in578 `composer_replication/diloco/serverless/tests/test_serverless_local.py`579 — runs N replicas locally, performs an allreduce across them.580- `test_local_executor_handles_multiple_rounds`581- `test_local_executor_reports_failed_replicas`582 583If those tests pass, your serverless DiLoCo machinery works — only the584specific cloud adapters are missing. The skeletons themselves are not585under test (raising in `__init__` is the contract).586 587```588pytest composer_replication/diloco/serverless/tests/test_serverless_local.py -xvs589```590 591---592 593### 10. DPPO mask drops every token — "loss became 0" or "no gradients"594 595**SYMPTOM.** You ported a PPO config from another framework (KL596penalty + clip ε=0.2 + value loss), wired it into the PRIME-RL recipe597with the default `dppo_mask_high=0.2` / `dppo_mask_low=0.2`, and the598training loss is suspiciously close to zero. Inspecting the recipe's599internal `keep_mask` shows nearly every token is being masked out.600 601**DIAGNOSIS.** PRIME-RL's "DPPO mask" is **not** the same as PPO602clipping, and not even the same as a log-ratio threshold. From the603recipe docstring at604`composer_replication/recipes/prime_rl/composer_loss.py` (mirroring605PRIME-RL upstream `prime_rl/trainer/rl/loss.py` lines 137-148):606 607> The mask gate is on **probability-space**608> `probs_diff = exp(trainer_lp) - exp(inference_lp)`, NOT on the609> log-ratio. A positive-advantage token is dropped iff610> `probs_diff > dppo_mask_high`; a negative-advantage token iff611> `probs_diff < -dppo_mask_low`. Masked tokens are **dropped from the612> policy-gradient term** but still contribute to the KL penalty.613 614The defaults `dppo_mask_high=dppo_mask_low=0.2` match PRIME-RL's615`DefaultLossConfig`. Because the gate is on probability-space, the616"in-band" zone is617`exp(trainer_lp) ∈ [exp(inference_lp) - 0.2, exp(inference_lp) + 0.2]`.618For a token with inference probability ~0.5 this is a fairly tight619band; for tokens at probability ~0.001 or ~0.999 the same threshold620behaves very differently from a log-ratio bound. This is by design —621PRIME-RL is bounding the absolute change in token probability, not the622multiplicative change.623 624The two failure modes:625 6261. **All tokens masked.** Trainer and inference engines disagree627 sharply (fp16 vs bf16, stale rollout cache, mismatched chat628 templates) and `probs_diff` exceeds 0.2 almost everywhere.6292. **No tokens masked.** Trainer ≈ inference (e.g. you forgot to step630 the optimizer between rollouts) so the bound is never binding and631 the policy never sees any DPPO regularization.632 633**FIX.** Inspect the empirical `probs_diff` distribution before634tuning:635 636```python637# In your training loop:638probs_diff = torch.exp(trainer_logprobs) - torch.exp(inference_logprobs)639print(torch.quantile(probs_diff.abs(), torch.tensor([0.5, 0.9, 0.99])))640```641 642For a healthy on-policy run with bf16 trainer + bf16 inference and643fresh rollouts, the central 99% of `|probs_diff|` should sit well644below `0.2`. If yours doesn't, the upstream divergence is the645problem, not the bound. Bumping `dppo_mask_high/low` to 0.5 or 1.0 is646a workaround but it disables the trust-region intent of DPPO.647 648**Do not** translate PPO ε=0.2 directly. PPO ε=0.2 is a multiplicative649log-ratio bound (`|log_ratio| < log(1.2) ≈ 0.18`); DPPO's 0.2 is an650**additive probability-space** bound. The semantics are different and651the defaults are deliberately tight in probability space.652 653If you genuinely want to disable the mask (e.g. for bug-isolation),654pass `dppo_mask_high=1e6, dppo_mask_low=1e6` (both are655`Field(..., ge=0)` upstream — negative values are rejected by656both PRIME-RL and our adapter). There is a regression test for657exactly this knob.658 659**VERIFICATION.**660- `test_dppo_mask_high_drops_positive_advantage_outliers` and661 `test_dppo_mask_low_drops_negative_advantage_outliers` in662 `composer_replication/recipes/prime_rl/tests/test_composer_loss.py`663 — assert that out-of-bound tokens are dropped from the664 policy-gradient term (with the upstream sign-of-advantage gate).665- `test_dppo_mask_sign_conditioned_on_advantage` — asserts that a666 positive-advantage token with a large *negative* probs_diff is NOT667 dropped (PRIME-RL only checks the upper bound for positive-advantage668 tokens).669- `test_dppo_bounds_can_be_disabled` — asserts that very wide bounds670 (`1e6`) pass every token through.671- `test_parity_with_prime_rl_default_loss_fn` — when `prime-rl` is672 installed, runs identical inputs through PRIME-RL upstream and our673 adapter and asserts the loss matches.674 675```676pytest composer_replication/recipes/prime_rl/tests/test_composer_loss.py -xvs677```678 679---680 681### 11. `compose_loss` runs but the GRPO channel doesn't behave like real GRPO682 683**SYMPTOM.** You read the README, saw the "3-channel composition: GRPO684+ SDPO + trace-replay DPO" tagline, called `compose_loss(model,685inputs)` directly in your training loop, and your reward curve never686moves the way it would in a real GRPO trainer. Or: you compared687against a TRL `GRPOTrainer` baseline and `compose_loss` produces688totally different numbers.689 690**DIAGNOSIS.** From the docstring at the top of691`composer_replication/loss.py:1–16`:692 693> This is a verification-harness mirror of694> `ComposerReplicationTrainer._compute_loss` that does NOT depend on695> TRL's GRPOTrainer parent. The GRPO channel is replaced with standard696> LM next-token-prediction cross-entropy, which is the limit GRPO697> converges to under deterministic rewards.698>699> Use it for: CPU smokes on real HF models, unit tests of loss700> composition without spinning up TRL, anywhere we want to verify701> gradient flow through the 3-channel sum without paying TRL's full702> machinery cost.703>704> **Do NOT use it as the production training loss.** Production =705> ComposerReplicationTrainer (a real GRPOTrainer subclass).706 707The `lm_ce` channel labelled "GRPO" in the LossComponents dataclass is708a **stub**: it is plain language-modeling cross-entropy. It is the709correct channel for verification (gradient flow, channel weighting,710distillation wiring), but it is not GRPO's surrogate objective and711will never produce the same numbers as real GRPO under stochastic712rewards.713 714Real GRPO requires:715- A reward model or rule-based reward,716- Per-prompt advantage estimation across G samples,717- An importance-sampling-ratio clip / mask.718 719Those live in TRL's `GRPOTrainer`, in our PRIME-RL recipe at720`composer_replication/recipes/prime_rl/composer_loss.py`, or (when721shipped) in a future VeRL recipe.722 723**FIX.**724- For production GRPO training, do **not** call `compose_loss` directly.725 Instead use one of:726 - `composer_replication.trainer.composer_trainer.ComposerReplicationTrainer`727 — TRL `GRPOTrainer` subclass, full machinery.728 - `composer_replication.recipes.prime_rl.composer_loss.loss_fn` —729 PRIME-RL's `CustomLossConfig` adapter (channel 1 is real DPPO-clipped GRPO).730- For ablations, smokes, and unit tests, `compose_loss` is the right731 tool — but log the `lm_ce` channel as `lm_ce`, not as `grpo`. The732 `LossComponents` dataclass already names the field correctly; if733 your wandb logger relabels it as "GRPO loss", fix the label.734 735**VERIFICATION.**736- The 11-test integration suite at737 `composer_replication/tests/test_compose_loss_integration.py` only738 asserts gradient flow + bit-exact composition; it deliberately does739 not assert any GRPO-specific property of `compose_loss`. That's the740 contract.741- The PRIME-RL recipe's real DPPO+KL behavior is asserted by742 `test_returns_finite_scalar`,743 `test_dppo_mask_high_drops_positive_advantage_outliers`,744 `test_dppo_mask_sign_conditioned_on_advantage`, and745 `test_parity_with_prime_rl_default_loss_fn` (skip-marked when746 `prime-rl` is not installed)747 in `composer_replication/recipes/prime_rl/tests/test_composer_loss.py`.748 Those tests verify a real importance-sampling-ratio gradient with749 PRIME-RL's advantage-conditioned mask, which `compose_loss` would750 not pass.751 752If you find yourself wanting `compose_loss` to behave like real GRPO,753that is the signal to switch to one of the production paths above.754 755```756pytest composer_replication/tests/test_compose_loss_integration.py::test_defaults_bit_exact_with_legacy_kwargs -xvs757pytest composer_replication/recipes/prime_rl/tests/test_composer_loss.py::test_returns_finite_scalar -xvs758```759 760---761 762### 10. `monarch` / `data-juicer` / `prime-rl` install (Wave 16)763 764**SYMPTOM.** `pip install -e ".[monarch]"`, `pip install -e ".[prime-rl]"`,765or `pip install -e ".[replaysim]"` fails immediately with a uv/pip766resolver error similar to:767 768```769× No solution found when resolving dependencies:770 ╰─▶ Because only monarch<=0.1.11 is available and771 composer-replication[monarch] depends on monarch>=0.4.1, we can772 conclude that composer-replication[monarch]'s requirements are773 unsatisfiable.774```775 776**DIAGNOSIS.** Three upstream packages the framework integrates with are777not currently pip-installable in their advertised versions:778 7791. **Meta's Monarch** is published on PyPI as780 `torchmonarch-nightly` (nightly wheels with platform constraints), not781 as `monarch`. The PyPI name `monarch` is unrelated to Meta's actor782 framework and tops out at `0.1.11`.7832. **Prime Intellect's prime-rl** is not registered on PyPI at all. It784 is published from source only.7853. **data-juicer** is not registered on PyPI under that exact name. The786 closest match (`py-data-juicer==1.0.0`) has broken transitive deps;787 newer `py-data-juicer` releases work but install ~150 transitive788 packages.789 790Wave 16 dropped all three extras from `pyproject.toml` rather than ship791unsatisfiable pins. The framework code paths that touch these libraries792import them lazily, so:793- `composer_replication.recipes.monarch` is a documentation skeleton794 that does NOT require monarch installed.795- `composer_replication.recipes.prime_rl.composer_loss` imports cleanly796 without prime-rl; the upstream parity test is `@skipif`-gated and the797 in-file shadow-parity test still verifies the loss formula798 independently.799- `composer_replication.replaysim.normalize.DJNormalizer(skip_dj=True)`800 works without `data_juicer`; only the full DJNormalizer code path801 needs it.802 803**FIX.** If you want any of these libraries' real functionality, install804from source alongside the framework:805 806```807# Meta Monarch (actor framework — see ADR-006)808pip install torchmonarch-nightly # OR install from source:809# git clone https://github.com/meta-pytorch/monarch && cd monarch && pip install -e .810 811# Prime Intellect prime-rl (Recipe C — see ADR-006)812git clone https://github.com/PrimeIntellect-ai/prime-rl813cd prime-rl && pip install -e .814 815# data-juicer (replaysim normalization — see ADR-004)816git clone https://github.com/modelscope/data-juicer817cd data-juicer && pip install -e .818```819 820**VERIFICATION.** A fresh checkout install with all surviving extras821should succeed:822 823```824uv venv --clear825uv pip install -e ".[diloco,replay,replaysim,train,dev]"826source .venv/bin/activate827python -m pytest -q # baseline 266 passed / 62 skipped (2026-06-09; varies by optional deps/Docker — see docs/V1_V8_COVERAGE.md)828```829 830If any of those extras fails to resolve, file a bug report — Wave 16831verified the full extras matrix installs from a clean venv on Python8323.11.833 834---835 836## How to file a bug report837 838If you've read the relevant section above and your problem persists,839file a bug. Include **all** sections of the template below — the most840common reason a maintainer can't repro is a missing piece of841environmental context.842 843```markdown844### What I expected vs what happened845(One paragraph.)846 847### Repro steps8481. ...8492. ...8503. ...851 852Minimal self-contained snippet (no `from my_local_thing import …`):853 854```python855# repro.py856from composer_replication import compose_loss857...858```859 860### Environment861- OS: (uname -a or `ver` on Windows)862- Python: (python --version)863- composer-replication: (pip show composer-replication | head -3)864- torch: (python -c "import torch; print(torch.__version__)")865- torchft: (python -c "import torchft; print(torchft.__version__)" || echo "n/a")866- transformers / trl: (versions, or "not installed")867- data-juicer / fsspec: (versions, or "not installed")868- s3fs / gcsfs / adlfs: (versions if relevant)869- GPU: (nvidia-smi -L or "CPU only")870- Install method: pip install -e . / wheel / other871- Extras installed: [replay] [replaysim] [serverless] [dev]872 873### What you've already tried874- [ ] Read the relevant Failure Mode section of docs/TROUBLESHOOTING.md875 (which one: ___)876- [ ] Ran `pytest <relevant test path>` and confirmed those tests pass877- [ ] Ran the repro snippet in a fresh venv878- [ ] Confirmed it reproduces on Python 3.11 (if you were on 3.12 / 3.13)879 880### Logs881(Full traceback. If it's a wrong-loss-curve rather than an exception,882paste loss values for the first 10 steps and link any wandb/tb run.)883 884### Hypothesis885(Optional. If you have a guess at where the bug is, name the file +886line number. We'll look there first.)887```888 889A few rules:890- **Do not** paste API keys, AWS credentials, or HuggingFace tokens.891- **Do** include the failing test name if you've narrowed it to one.892- **Do** distinguish "never worked" from "regressed between commit X893 and Y." A regression-bisect goes straight to the front of the queue.894- **One bug per issue.** Multi-headed reports lose items in triage.895 896The Wave-14 surface area is large, but the test suite covers it897densely — every section above corresponds to a green test that proves898the fix worked.899 