Codeseys/composer-replication-framework
0
1# RL Post-Training Frameworks Landscape & Meta PyTorch Stack Audit2 3> **Generated:** 2026-05-254> **Scope:** Audit of RL post-training frameworks beyond TRL+VeRL plus Meta's PyTorch agentic stack components, with a recommendation of two additions to the Composer Replication Framework.5> **Feeds:** ADR-006 (Algorithm-substrate selection)6> **Companion docs:** `~/wiki/research/post-training-framework/04-verl-trl.md`, `~/wiki/research/post-training-framework/03-monarch-torchforge-openenv.md`, `~/wiki/research/post-training-framework/02-diloco-family.md`7 8---9 10## TL;DR — Recommendation11 12| Slot | Pick | Why |13|---|---|---|14| **RL framework #3 (after TRL, VeRL)** | **PRIME-RL (PrimeIntellect-ai/prime-rl)** | First-class `CustomLossConfig` extension point (`trainer.loss.type=custom` + `import_path`) — the cleanest place we have to drop our **3-channel loss (RLVR + hint-distill + trace-replay)** without forking. Already uses the `verifiers` env protocol that bridges to OpenEnv. Async, decentralized substrate. Apache-2.0. INTELLECT-2 production receipts. |15| **Infra component (Meta stack)** | **Monarch (`meta-pytorch/monarch`)** as the actor-mesh control plane; **TorchTitan** is *also* tracked as the FSDP2/TP/PP training core but is already the trainer inside both PRIME-RL and TorchForge, so we adopt it transitively. The single net-new dependency is **Monarch**. | Monarch is the only Meta-stack component that is (a) actively shipped (v0.4 GA, v0.5 dev, weekly wheels), (b) decoupled from the now-paused TorchForge, and (c) able to host *any* SPMD trainer (TRL, VeRL, PRIME-RL) as an `ActorMesh`. BSD-3. Replaces Ray when our v0.2 lands. |16 17**What we do NOT add:**18- OpenRLHF — strong production framework (v0.9.10, 9.3K★, supports DAPO) but its custom-loss path requires modifying `openrlhf/models/loss.py` + a `Trainer` subclass. Strictly worse extension story than PRIME-RL for our specific need (3-channel loss).19- NeMo-Aligner — no GRPO, no DAPO, heavy NeMo/Megatron dependency. Wrong shape.20- Unsloth — TRL wrapper, RL kernels live in closed `unsloth_zoo`. We'd have to fork.21- LLaMA-Factory — TRL wrapper, no GRPO/DAPO (delegates to EasyR1).22- DeepSpeed-Chat — effectively unmaintained for new RL algos since Aug 2023; PPO/DPO only.23- TorchForge — Meta has marked the repo "development paused, consolidating into TorchTitan." Borrow patterns; do not depend on it.24- torchchat — inference / local deployment only; no training. Out of scope.25 26---27 28## Table of Contents29 301. [Audit Methodology](#1-audit-methodology)312. [RL Framework Audit](#2-rl-framework-audit)32 1. [OpenRLHF](#21-openrlhf)33 2. [PRIME-RL](#22-prime-rl)34 3. [NeMo-Aligner](#23-nemo-aligner)35 4. [Unsloth (RL)](#24-unsloth-rl)36 5. [LLaMA-Factory](#25-llama-factory)37 6. [DeepSpeed-Chat](#26-deepspeed-chat)383. [Meta PyTorch Agentic Stack — Infra vs Training Split](#3-meta-pytorch-agentic-stack)39 1. [Monarch (coordination/infra)](#31-monarch)40 2. [TorchTitan (training stack)](#32-torchtitan)41 3. [TorchForge (paused)](#33-torchforge)42 4. [torchchat (out of scope)](#34-torchchat)434. [Comparison Matrix](#4-comparison-matrix)445. [Recommendation Rationale](#5-recommendation-rationale)456. [Integration Sketches](#6-integration-sketches)467. [Sources](#7-sources)47 48---49 50## 1. Audit Methodology51 52For each framework, we capture five fields that determine whether it can host the Composer Replication Framework's three-channel loss (RLVR + hint-distill + trace-replay) on our existing OpenEnv-compatible TRL data path:53 541. **Repo + license + last commit + maturity** — primary GitHub source, license grade for redistribution, recency, and whether the project is *production*, *research*, or *archived*.552. **Algorithm coverage** — does it ship GRPO and DAPO out of the box? (DAPO matters because Composer-style training inherits its decoupled clip + dynamic sampling fixes for length and std biases.)563. **Custom-loss extension point** — concrete file/class/config where a custom 3-channel loss can be plugged. We strongly prefer a stable public hook over forking.574. **Integration cost** — rough lines of code needed for a `Recipe` doc + a skeleton `Trainer` subclass that runs end-to-end on a small env.585. **OpenEnv data-path fit** — does it already consume the OpenEnv contract (typed `reset`/`step`/`close`, MCP tool-calling) directly, or do we have to write a shim?59 60Primary sources: each repo's `README.md`, official releases page, and DeepWiki audits (where indexed). Secondary checks: PyPI release timelines for Meta packages.61 62---63 64## 2. RL Framework Audit65 66### 2.1 OpenRLHF67 68| Field | Value |69|---|---|70| **Repo** | https://github.com/OpenRLHF/OpenRLHF |71| **License** | Apache-2.0 |72| **Stars / contributors** | 9,312 ★ / 90 contributors |73| **Latest release** | v0.9.10, 2026-04-04 |74| **Last push** | 2026-04-05 |75| **Maturity** | **Production** — used in many public RLHF runs since 2023; tagline "An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & TIS & vLLM & Ray & Async RL)" |76| **Algorithms** | PPO, GRPO, **DAPO** (release notes; advertised as a primary feature in v0.9.x), REINFORCE++, REINFORCE++-baseline, RLOO, GSPO, Async RL, TIS (truncated importance sampling) |77| **Custom-loss extension point** | `openrlhf/models/loss.py` — `PolicyLoss`, `DPOLoss`, `SFTLoss`, `PairWiseLoss`, `LogExpLoss` are concrete `nn.Module`s. To add a 3-channel loss you would (a) add a new `nn.Module` (e.g. `ThreeChannelLoss`) here, then (b) subclass the relevant `Trainer` (e.g. `PPOTrainer` / a new GRPO-derived trainer) and replace `self.loss_fn`. There is **no config-driven custom-loss hook** equivalent to PRIME-RL's `CustomLossConfig` — you fork or vendor. |78| **Integration cost** | Higher than PRIME-RL. Estimated **~400–600 LOC**: ~150 LOC for a `ThreeChannelLoss` module, ~200 LOC for a `ComposerGRPOTrainer` subclass that routes the three signals (RLVR scalar, hint-distill teacher logprobs, trace-replay teacher logits), ~50 LOC for a `Recipe` doc, plus reward-fn glue. |79| **Data-path fit** | OpenRLHF's input is HF chat templates + a Python reward function or a remote reward URL (`--reward.remote_url`, `--train.agent_func_path`). It does **not** speak the OpenEnv `reset/step` protocol natively, but our existing OpenEnv→TRL adapter could be reused as a callable behind `agent_func_path`. **Medium** lift to wire OpenEnv. |80 81**Verdict:** Strong, mature, well-funded codebase with the *most* complete algorithm coverage of any candidate. Loses to PRIME-RL only because PRIME-RL has a first-class config-driven custom-loss hook that fits our exact need, and PRIME-RL already has the `verifiers`/OpenEnv shape baked into the orchestrator. We keep OpenRLHF on the radar as a fallback substrate if PRIME-RL's decentralized story is overkill for v0.1.82 83---84 85### 2.2 PRIME-RL86 87| Field | Value |88|---|---|89| **Repo** | https://github.com/PrimeIntellect-ai/prime-rl |90| **License** | Apache-2.0 |91| **Stars / contributors** | 1,398 ★ / 60 contributors |92| **Latest release** | v0.5.0, 2026-03-30 |93| **Last push** | 2026-05-25 (active today) |94| **Maturity** | **Production-research hybrid** — substrate behind INTELLECT-1/2 multi-DC runs; tagline "Async RL Training at Scale". Decentralized DiLoCo-shape compute is its differentiator. |95| **Algorithms** | **GRPO**, GSPO, on-policy distillation with a teacher model. `default_loss_fn` = DPPO + KL (a GRPO variant; similar lineage to DAPO's decoupled-clip idea but the upstream "DAPO" label is not used verbatim). |96| **Custom-loss extension point** | **Best in class.** `src/prime_rl/trainer/rl/loss.py` exposes a `LossInputs`/`LossOutputs` interface and `setup_loss_fn` resolves a config: `trainer.loss.type = "custom"` + `trainer.loss.import_path = "your_pkg.your_module.your_loss_fn"` + optional kwargs. The custom function receives `trainer_logprobs`, `inference_logprobs`, `teacher_logprobs`, `advantages`, `loss_mask` — i.e., the exact tensor inputs needed for a 3-channel loss (RLVR uses `advantages`, hint-distill uses `teacher_logprobs`, trace-replay can be threaded through `kwargs` as a precomputed reference). |97| **Integration cost** | **Lowest.** Estimated **~200–300 LOC total**: ~120 LOC for a `composer_three_channel_loss` function in our package + ~30 LOC of config (`recipes/composer_v0.toml`), ~80 LOC `Recipe` doc. No subclassing required for the loss. A small adapter is needed if we precompute the trace-replay teacher distribution outside the `LossInputs` struct. |98| **Data-path fit** | **Already aligned.** PRIME-RL's orchestrator consumes `verifiers` environments via `vf.EnvServer`. The OpenEnv ↔ verifiers shim is a known small adapter (the `verifiers` library is the Hub-side env runner that OpenEnv's TRL guide already uses). Our existing OpenEnv-compatible TRL data path drops in with a thin wrapper. |99 100**Verdict:** Best fit for the framework. The combination of (i) config-driven custom loss with the right tensor signatures already present, (ii) verifiers/OpenEnv shape, (iii) decentralized async training that maps to our DiLoCo plans, makes PRIME-RL the substrate of choice for v0.1. **Recommended addition #1.**101 102---103 104### 2.3 NeMo-Aligner105 106| Field | Value |107|---|---|108| **Repo** | https://github.com/NVIDIA/NeMo-Aligner |109| **License** | Apache-2.0 |110| **Maturity** | **Research-leaning production** — NVIDIA-maintained, tied to NeMo/Megatron-LM. Advertised as "early stages of development" in its own README. |111| **Algorithms** | PPO, REINFORCE, RS (Rejection Sampling), DPO, RPO. **No GRPO. No DAPO.** |112| **Custom-loss extension point** | `loss_func` method on Megatron model classes (e.g. `MegatronGPTDPOModel.loss_func`). Requires NeMo model-class subclassing and Megatron-LM familiarity. |113| **Integration cost** | High. Estimated **~800–1,200 LOC** including .nemo conversion of HF weights, Megatron model wrapping, custom Megatron `loss_func`, and a recipe. Plus the operational cost of running on Megatron-LM (Triton kernels, NeMo container). |114| **Data-path fit** | JSONL only; no OpenEnv. We'd write a full env adapter. |115 116**Verdict:** Wrong shape. No GRPO/DAPO and tightly bound to the NeMo ecosystem. Only relevant if we ever need NVIDIA-supported large-scale Megatron RL, which we don't for the Composer Replication v0.1/v0.2 horizon. **Reject.**117 118---119 120### 2.4 Unsloth (RL)121 122| Field | Value |123|---|---|124| **Repo** | https://github.com/unslothai/unsloth |125| **License** | Apache-2.0 (per public README; not surfaced by DeepWiki snapshot but well-known) |126| **Maturity** | **Production** for SFT and LoRA/QLoRA; **research/preview** for RL — RL support shipped in 2025 as a TRL patcher. |127| **Algorithms** | Wraps TRL → inherits TRL's GRPO; loss-type switch supports `"grpo"`, `"bnpo"`, `"dr_grpo"`, `"dapo"`, `"cispo"`. So **GRPO and DAPO are both available** through the patched-TRL path. |128| **Custom-loss extension point** | Problematic. The actual loss kernels live in `unsloth_zoo` (a *separate* compiled dependency). The patcher (`patch_trl_rl_trainers()`) generates modified TRL trainer classes via `exec()` from string templates. To add a new loss type you would have to (a) modify or fork `unsloth_zoo` to add a kernel, (b) extend `RL_REPLACEMENTS`, and (c) extend the `compute_loss()` switch in the patcher template. **There is no public Python subclass hook that survives the patching.** |129| **Integration cost** | Very high if we want our own loss. Forking `unsloth_zoo` defeats the purpose of using Unsloth (which is the optimized kernels). Estimated ~1,000+ LOC plus an external repo to maintain. |130| **Data-path fit** | TRL-shaped, so OpenEnv via TRL is fine — but only for *stock* TRL losses. Our 3-channel loss does not survive Unsloth's patching. |131 132**Verdict:** Excellent for memory-efficient SFT and stock-GRPO LoRA. Wrong tool for a custom loss. **Reject** as the substrate; we may still use it as an *optional* QLoRA accelerator inside a stock-GRPO ablation run.133 134---135 136### 2.5 LLaMA-Factory137 138| Field | Value |139|---|---|140| **Repo** | https://github.com/hiyouga/LLaMA-Factory |141| **License** | Apache-2.0 |142| **Maturity** | **Production** for breadth (50+ model families, SFT/DPO/PPO recipes), but RL is a thin TRL wrapper. |143| **Algorithms** | PPO, DPO, KTO, ORPO, SimPO via `Custom*Trainer` subclasses of the corresponding `trl.*Trainer` classes. **No GRPO. No DAPO** in the repo itself; the README points to **EasyR1** (an external GRPO framework) for those. |144| **Custom-loss extension point** | `compute_preference_loss` switch on `CustomDPOTrainer` (selects `sigmoid` / `hinge` / `ipo` / `kto_pair` / `orpo` / `simpo`). For PPO, you would subclass `CustomPPOTrainer` → which is `trl.PPOTrainer`. Effectively the same extension story as plain TRL, with a configuration layer on top. |145| **Integration cost** | Moderate, ~400 LOC, but you are essentially using TRL through one extra layer. |146| **Data-path fit** | Text/dataset-shaped, not OpenEnv-aware. Same OpenEnv-via-TRL story. |147 148**Verdict:** Useful as a multi-model SFT laboratory but does not move the ball for our RL-side requirements. **Reject** as substrate; we already have TRL.149 150---151 152### 2.6 DeepSpeed-Chat153 154| Field | Value |155|---|---|156| **Repo** | https://github.com/deepspeedai/DeepSpeedExamples (the `applications/DeepSpeed-Chat/` subtree) |157| **License** | Apache-2.0 |158| **Maturity** | **Effectively stale.** The README's "Latest News" cuts off in August 2023. CI patches in 2025 (e.g., #6982, #7015, #7052) are dependency-pinning fixes, not feature work. The roadmap to "generalize DeepSpeed-RLHF abstraction for a wider range of RL algorithms" has not landed. |159| **Algorithms** | PPO (3-stage RLHF) + DPO. **No GRPO. No DAPO.** |160| **Custom-loss extension point** | `DeepSpeedPPOTrainer.train_rlhf` / `actor_loss_fn` / `critic_loss_fn`. Editable but not config-hooked. |161| **Integration cost** | Moderate, but you inherit a frozen architecture. ~500 LOC. |162| **Data-path fit** | Prompt-dataset-shaped; no OpenEnv. |163 164**Verdict:** Pioneering for its time, no longer competitive on algorithm coverage. **Reject.**165 166---167 168## 3. Meta PyTorch Agentic Stack — Infra vs Training Split169 170The brief asked specifically to **distinguish coordination/infra from training-stack** components. The answer is:171 172| Component | Layer | Status (May 2026) | In our framework? |173|---|---|---|---|174| **Monarch** (`meta-pytorch/monarch`) | **Coordination / Infra** — actor mesh, RDMA data plane, supervision trees | **Active.** v0.4 GA (2026-03-26), v0.5 dev wheels daily, BSD-3 | **Yes — recommended addition.** |175| **TorchTitan** (`pytorch/torchtitan`) | **Training stack** — FSDP2 / TP / PP / CP / float8 / MXFP8 | **Active.** BSD-3, "extensive development". Has an experimental GRPO recipe (`experiments/rl/simple_grpo_sum_digits.py`) on Monarch. | **Indirectly** — already the trainer inside PRIME-RL and TorchForge. We adopt it transitively, not as a direct dependency. |176| **TorchForge** (`meta-pytorch/forge`) | RL post-training library | **Development paused** per the repo banner; consolidating into TorchTitan. ~685★. | **Pattern reference only.** Lift the Generator/Trainer/Rewarder *shape* but do not depend on the package. |177| **torchchat** (`pytorch/torchchat`) | **Inference / local deployment** | Active for its own scope, but: not a training framework; no RL surface. | **Out of scope.** |178| **OpenEnv** (`meta-pytorch/OpenEnv`) | Environment standard (covered separately) | Active. Already a v0 dependency of the framework. | Already adopted. |179 180### 3.1 Monarch181 182| Field | Value |183|---|---|184| **Repo** | https://github.com/meta-pytorch/monarch |185| **License** | BSD-3-Clause |186| **PyPI** | `torchmonarch`; v0.4.1 stable (2026-04-08), v0.5.0 dev wheels published daily through 2026-05-05 |187| **Maturity** | **Experimental but actively shipped.** "Currently in an experimental stage" per the repo's own status note, but with a functioning K8s operator, weekly wheels, ProcessMesh/ActorMesh APIs stable enough for VeRL backend experiments. |188| **Role in our stack** | **Pure coordination/infra.** It does not train models. It hosts whatever trainer you bring (TRL, VeRL, PRIME-RL, TorchTitan) as `Actor` subclasses on a `ProcMesh`. The `monarch.spmd.SPMDActor` automatically configures `RANK`/`LOCAL_RANK`/`WORLD_SIZE` for any PyTorch-distributed script — i.e., we can lift our existing TRL or PRIME-RL workers into Monarch with minimal change. |189| **Key abstractions** | `ProcMesh` (processes × hosts × GPUs), `ActorMesh` (typed actors with `@endpoint` methods), supervision trees, RDMA buffers, distributed tensors / DTensor integration. Underlying runtime: `hyperactor` (Rust). |190| **Why over Ray** | Tighter PyTorch/DTensor integration; explicit RDMA data plane (Ray uses object store + standard networking); single-controller mental model maps directly to RL post-training (one controller orchestrates Generator + Trainer + Rewarder + Env actors). |191| **Integration cost into Composer Replication** | **~300 LOC + ops**: (a) wrap our PRIME-RL trainer as an `SPMDActor`; (b) wrap our vLLM rollout server as an `Actor` with an `@endpoint generate(prompts)` method; (c) write a single controller script that creates a `ProcMesh`, spawns both meshes, and shuttles `DataProto`-shaped messages; (d) Recipe doc. The ops cost is the harder half — Monarch's K8s operator is new (v0.2.0+). |192| **Risk** | Pre-1.0; API churn possible (e.g., `KubernetesJob.add_mesh` signature changed in v0.5). Mitigation: pin to `torchmonarch==0.4.1` for v0.2 of our framework. |193 194### 3.2 TorchTitan195 196| Field | Value |197|---|---|198| **Repo** | https://github.com/pytorch/torchtitan |199| **License** | BSD-3-Clause |200| **Maturity** | **Active development** for pretraining; **experimental** for RL. The GRPO experiment (`torchtitan/experiments/rl/simple_grpo_sum_digits.py`) is in `experiments/`, which the repo explicitly disclaims as removable. |201| **Role** | **Training stack only.** Provides FSDP2 (per-parameter sharding), Tensor Parallel (incl. async TP), Pipeline Parallel (zero-bubble), Context Parallel (long-context), `torch.compile`, Float8, MXFP8, DDP, HSDP. |202| **OpenEnv-aware?** | No, but the experimental `RLTrainer` integrates `vLLM` + Monarch actors, which is the same shape PRIME-RL uses. |203| **Why we don't add it directly** | **PRIME-RL already uses TorchTitan-equivalent FSDP2 internals**, and TorchForge's training core was TorchTitan. Adding TorchTitan as a *direct* dependency would mean writing our own RL loop on top of it — that's TorchForge's job, and Meta paused exactly that effort. The right move is to depend on PRIME-RL, which has battle-tested distributed training patterns equivalent to TorchTitan's, and revisit TorchTitan directly only when we genuinely need its experimental zero-bubble PP or MXFP8 paths. |204 205### 3.3 TorchForge (Paused)206 207- Repo banner: **"Development paused — LLM training consolidating in TorchTitan."**208- ~685 ★, 100+ open issues, last meaningful release in early 2026.209- Patterns we should still copy:210 - Generator/Trainer/Rewarder ActorMesh decomposition211 - TorchStore-style RDMA weight broadcast212 - Async toggle between sync PPO-like and fully async off-policy213- **We do not add a TorchForge dependency.** Architectural reference only.214 215### 3.4 torchchat (Out of Scope)216 217- Inference / local deployment of LLMs (Eager / `torch.compile` / AOT Inductor / ExecuTorch / mobile).218- No training, no RL.219- Mentioned in the brief for completeness; ruled out cleanly.220 221---222 223## 4. Comparison Matrix224 225### 4.1 RL Frameworks226 227| Framework | License | Last release | Maturity | GRPO | DAPO | Custom-loss hook | OpenEnv fit | Est. integration LOC |228|---|---|---|---|---|---|---|---|---|229| **TRL** (baseline) | Apache-2.0 | Active | Production | ✅ | partial (tricks land per release) | Subclass `GRPOTrainer.compute_loss` | ✅ native (Oct 2025 OpenEnv guide) | already integrated |230| **VeRL** (baseline) | Apache-2.0 | Active | Production | ✅ | ✅ | `core_algos.py` + worker subclass | shim via Ray dataloader | already skeleton |231| **OpenRLHF** | Apache-2.0 | v0.9.10 (2026-04-04) | Production | ✅ | ✅ | `openrlhf/models/loss.py` + Trainer subclass; **no config hook** | shim via `agent_func_path` | ~400–600 |232| **PRIME-RL** ⭐ | Apache-2.0 | v0.5.0 (2026-03-30) | Prod-research | ✅ | partial (DPPO+KL variant; not labeled DAPO) | **`CustomLossConfig` import_path — first-class** | ✅ via `verifiers` (OpenEnv-compatible) | **~200–300** |233| **NeMo-Aligner** | Apache-2.0 | Active | Research-leaning | ❌ | ❌ | Megatron model `loss_func` | none; JSONL only | ~800–1,200 |234| **Unsloth (RL)** | Apache-2.0 | Active | Production (SFT) / preview (RL) | ✅ (via TRL patch) | ✅ (via TRL patch) | Loss kernels in closed `unsloth_zoo`; effectively unhookable | TRL-shaped | ~1,000+ (forking) |235| **LLaMA-Factory** | Apache-2.0 | Active | Production | ❌ (delegates to EasyR1) | ❌ | TRL `Custom*Trainer` subclass | TRL-shaped | ~400 |236| **DeepSpeed-Chat** | Apache-2.0 | Stale (Aug 2023 features; 2025 only CI fixes) | Effectively maintained-only | ❌ | ❌ | `DeepSpeedPPOTrainer` subclass | none | ~500 |237 238### 4.2 Meta PyTorch Stack239 240| Component | Layer | License | Status | In recommendation? |241|---|---|---|---|---|242| **Monarch** ⭐ | Coordination / actor mesh | BSD-3 | Active (v0.4 GA, v0.5 dev) | **Yes** |243| **TorchTitan** | Training stack | BSD-3 | Active; RL experimental | Indirect (via PRIME-RL) |244| **TorchForge** | RL library | BSD-3 | **Paused** | No — patterns only |245| **torchchat** | Inference / deployment | BSD-3 | Active | No — out of scope |246| **OpenEnv** | Environment standard | (Hub) | Active | Already adopted |247 248---249 250## 5. Recommendation Rationale251 252### 5.1 Why PRIME-RL, not OpenRLHF253 254OpenRLHF is in many ways the safer pick: more stars, more contributors, more algorithm coverage (it explicitly ships DAPO). The deciding factor is **the shape of our custom loss**.255 256The Composer Replication Framework's signature contribution is the **three-channel reward**:257 2581. **RLVR** — tests-pass scalar from the OpenEnv environment.2592. **Composer-style hint-distill (SDPO/OPSD)** — the model self-teaches against its own hint-conditioned roll-outs; needs `teacher_logprobs` aligned to the rollout token grid.2603. **Trace-replay multi-teacher PRM** (the novel bit) — N frozen external teachers' precomputed token-level distributions, replayed against the on-policy rollout.261 262PRIME-RL's `LossInputs` dataclass already exposes exactly the tensors we need:263```264trainer_logprobs, inference_logprobs, teacher_logprobs, advantages, loss_mask265```266A custom 3-channel loss is roughly:267```python268def composer_three_channel_loss(li: LossInputs, *, hint_weight, replay_weight, replay_logits) -> LossOutputs:269 rlvr = grpo_term(li.trainer_logprobs, li.inference_logprobs, li.advantages, li.loss_mask)270 hint = kl_term(li.trainer_logprobs, li.teacher_logprobs, li.loss_mask)271 replay = kl_term(li.trainer_logprobs, replay_logits, li.loss_mask)272 return LossOutputs(loss=rlvr + hint_weight * hint + replay_weight * replay, ...)273```274We register this with `trainer.loss.type = "custom"` + `import_path` and we're done. No subclassing, no `exec()`-patched template, no Megatron model wrapping.275 276OpenRLHF would require us to (a) add a `ThreeChannelLoss` `nn.Module` to `openrlhf/models/loss.py`, (b) subclass `PPOTrainer` (or equivalent GRPO trainer) to construct it with the right teacher-logprob plumbing, and (c) carry that fork forward. ~2× the LOC, plus a fork to maintain.277 278A second factor: PRIME-RL's `verifiers` env protocol is a direct precursor of OpenEnv's wire shape (HTTP/WebSocket env servers, typed observations). Our existing OpenEnv-compatible TRL data path translates with a thin adapter. OpenRLHF's `agent_func_path` is more of an escape hatch than a contract.279 280A third factor: PRIME-RL was *built for decentralized training* (INTELLECT-1/2). Even though our v0.1 stays on a single cluster, the v0.2 multi-DC story drops in cleanly. OpenRLHF is Ray-on-one-cluster by design.281 282### 5.2 Why Monarch, not TorchTitan or TorchForge283 284Among the four Meta-stack components in the brief, only one is both (a) ours to add and (b) genuinely new functionality:285 286- **TorchForge** is paused — depending on it now is a known dead end.287- **TorchTitan** is already inside PRIME-RL transitively (PRIME-RL uses FSDP2 plus a SHARDCAST weight-broadcast layer that is morally equivalent to what TorchTitan offers). Adding TorchTitan as a *direct* dependency means writing our own RL loop on top of it, which is exactly what TorchForge tried and paused. We get TorchTitan's benefits without owning the integration.288- **torchchat** is for local inference / mobile deployment — out of scope.289- **Monarch** is the unique value: a PyTorch-native actor mesh that lets us replace Ray (PRIME-RL's current orchestration substrate) with something that has explicit RDMA, supervision trees, and ProcMesh/ActorMesh primitives that map directly onto our (Generator, Trainer, Rewarder, EnvServer) topology.290 291The migration path is incremental:292- **v0.1:** PRIME-RL on Ray (current). Monarch listed as roadmap.293- **v0.2:** Wrap PRIME-RL's Trainer as a `monarch.spmd.SPMDActor`, vLLM Generator as an `Actor` with an `@endpoint generate()`. Switch the orchestrator from `ray.init()` to `this_host().spawn_procs()`.294- Risk-mitigation: pin to `torchmonarch==0.4.1` (the last GA release before v0.5 dev). Keep a Ray fallback path active until v0.2 is stable.295 296---297 298## 6. Integration Sketches299 300### 6.1 PRIME-RL Recipe skeleton301 302`recipes/composer_v0_prime_rl.toml` (~30 LOC):303 304```toml305# composer_v0_prime_rl.toml306[model]307name = "Qwen/Qwen3-32B" # or Kimi-K2.5 when MoE support lands308 309[data]310env = "swe_bench_lite" # via verifiers EnvServer; wraps our OpenEnv adapter311batch_size = 64312group_size = 16313 314[trainer]315algorithm = "grpo"316 317> **Realised in v0.1 (Wave 17 update):** Wave 14b shipped the PRIME-RL318> recipe at `composer_replication/recipes/prime_rl/prime_rl_config.yaml`319> as **YAML** with a different kwarg surface than the TOML sketch below.320> The actual recipe shape:321>322> ```yaml323> # composer_replication/recipes/prime_rl/prime_rl_config.yaml324> model:325> base: "Qwen/Qwen2.5-0.5B"326> attn_implementation: "flash_attention_2"327> dtype: "bfloat16"328> env:329> protocol: "verifiers"330> config: { name: "math/gsm8k", split: "train" }331> loss:332> custom:333> import_path: "composer_replication.recipes.prime_rl.composer_loss:loss_fn"334> kwargs:335> alpha_sdpo: 0.0 # channel 2 deferred in v0336> beta_dpo: 0.0 # channel 3 out-of-scope for PRIME-RL v0337> dppo_mask_high: 0.2 # PRIME-RL DPPO convention (NOT textbook PPO)338> dppo_mask_low: 0.2 # both must be >= 0 per Field(..., ge=0)339> adv_tau: 1.0 # advantage normalization340> kl_tau: 0.04 # KL coefficient341> ```342>343> The realised `loss_fn(inputs, **kwargs)` matches PRIME-RL's344> `LossInputs`/`LossOutputs` interface (read upstream `prime_rl/loss.py`345> for parity verification — Wave 14b's shadow-parity test independently346> restates the formula in347> `composer_replication/recipes/prime_rl/tests/test_composer_loss.py`).348>349> The pre-Wave-14b TOML/`hint_weight`/`replay_weight` sketch below is350> preserved as historical proposal context.351 352[trainer.loss]353type = "custom"354import_path = "composer_replication.recipes.prime_rl.composer_loss:loss_fn"355[trainer.loss.kwargs]356hint_weight = 0.5357replay_weight = 0.25358replay_logits_path = "/data/teachers/precomputed_replay.zarr"359 360[teacher]361model = "Qwen/Qwen3-32B" # same as policy = self-teacher for hint-distill362hint_template = "composer.hint_v1"363 364[orchestrator]365sync_mode = "async"366shardcast = true367```368 369`composer_replication/recipes/prime_rl/composer_loss.py` (~120 LOC; current Wave 14b370implementation defines `loss_fn(inputs, **kwargs)` rather than the371`composer_three_channel_loss(li, *, hint_weight, replay_weight, replay_logits)` 372signature sketched below):373 374```python375# composer_replication/recipes/prime_rl/composer_loss.py — sketch only;376# the actual signature evolved during Wave 14b. See module docstring for377# the current `loss_fn` contract.378from prime_rl.trainer.rl.loss import LossInputs, LossOutputs379 380def composer_three_channel_loss(381 li: LossInputs,382 *,383 hint_weight: float,384 replay_weight: float,385 replay_logits_handle: str,386) -> LossOutputs:387 # 1. RLVR via GRPO surrogate388 rlvr = grpo_surrogate(li.trainer_logprobs, li.inference_logprobs,389 li.advantages, li.loss_mask)390 391 # 2. Hint-distill: KL(policy || hint-conditioned teacher)392 hint = masked_kl(li.trainer_logprobs, li.teacher_logprobs, li.loss_mask)393 394 # 3. Trace-replay: KL(policy || precomputed multi-teacher mixture)395 replay = trace_replay_kl(li.trainer_logprobs, replay_logits_handle, li.loss_mask)396 397 total = rlvr + hint_weight * hint + replay_weight * replay398 return LossOutputs(399 loss=total,400 metrics={"rlvr": rlvr.item(), "hint": hint.item(), "replay": replay.item()},401 )402```403 404Plus `docs/recipes/composer_v0_prime_rl.md` (~50 LOC) describing data layout, teacher precomputation, and reproducibility hashes.405 406**Total: ~200 LOC of code + ~30 LOC config + ~50 LOC docs ≈ 280 LOC.**407 408### 6.2 Monarch wrap-up sketch (v0.2)409 410```python411# composer_replication/orchestrator/monarch_runner.py (~120 LOC)412from monarch.actor import Actor, endpoint413from monarch.proc_mesh import this_host, ProcMesh414 415class TrainerActor(Actor):416 @endpoint417 async def step(self, batch): ...418 419class GeneratorActor(Actor):420 @endpoint421 async def generate(self, prompts): ...422 423class RewarderActor(Actor):424 @endpoint425 async def score(self, traj): ...426 427async def main(cfg):428 train_mesh = await this_host().spawn_procs(TrainerActor, hosts=4, gpus=8)429 gen_mesh = await this_host().spawn_procs(GeneratorActor, hosts=2, gpus=8)430 rew_mesh = await this_host().spawn_procs(RewarderActor, hosts=1, gpus=2)431 432 async for step in range(cfg.steps):433 prompts = await env.batch()434 traj = await gen_mesh.generate.broadcast(prompts)435 rewards = await rew_mesh.score.broadcast(traj)436 await train_mesh.step.broadcast({"traj": traj, "rewards": rewards})437```438 439**Total: ~120 LOC controller + ~50 LOC ops (K8s operator manifest) + ~80 LOC recipe doc ≈ 250 LOC.**440 441---442 443## 7. Sources444 445### Primary446 447- **OpenRLHF** — https://github.com/OpenRLHF/OpenRLHF (README, Releases v0.9.10), Apache-2.0; DeepWiki: `openrlhf/models/loss.py`, `agent_func_path`.448- **PRIME-RL** — https://github.com/PrimeIntellect-ai/prime-rl (README, Releases v0.5.0), Apache-2.0; DeepWiki: `src/prime_rl/trainer/rl/loss.py`, `CustomLossConfig`, `LossInputs`/`LossOutputs`, `verifiers` integration.449- **NeMo-Aligner** — https://github.com/NVIDIA/NeMo-Aligner, Apache-2.0; DeepWiki: PPO/REINFORCE/DPO/RPO; `loss_func` on Megatron model classes.450- **Unsloth** — https://github.com/unslothai/unsloth, README RL section; DeepWiki: `patch_trl_rl_trainers()`, `unsloth_zoo` kernels, DAPO loss-type switch.451- **LLaMA-Factory** — https://github.com/hiyouga/LLaMA-Factory, Apache-2.0; DeepWiki: `CustomPPOTrainer`/`CustomDPOTrainer`, EasyR1 reference for GRPO.452- **DeepSpeed-Chat** — https://github.com/deepspeedai/DeepSpeedExamples (`applications/DeepSpeed-Chat/`), Apache-2.0; DeepWiki: 3-stage PPO, DPO; "Latest News" cutoff Aug 2023; 2025 PRs (#6982, #7015, #7052) confirming maintenance-only mode.453- **Monarch** — https://github.com/meta-pytorch/monarch, BSD-3; PyPI `torchmonarch` v0.4.1 (2026-04-08), v0.5.0 dev wheels through 2026-05-05; DeepWiki: `ProcMesh`, `ActorMesh`, `monarch.spmd.SPMDActor`.454- **TorchTitan** — https://github.com/pytorch/torchtitan, BSD-3; DeepWiki: FSDP2/TP/PP/CP, `torchtitan/experiments/rl/simple_grpo_sum_digits.py`, integration with vLLM and Monarch.455- **TorchForge** — https://github.com/meta-pytorch/forge, BSD-3, repo banner "development paused — consolidating in TorchTitan".456- **torchchat** — https://github.com/pytorch/torchchat, BSD-3; DeepWiki: inference-only (eager / `torch.compile` / AOT Inductor / ExecuTorch).457 458### Companion repository docs (already present)459 460- `~/wiki/research/post-training-framework/04-verl-trl.md` — VeRL vs TRL deep dive.461- `~/wiki/research/post-training-framework/03-monarch-torchforge-openenv.md` — full Meta-stack survey.462- `~/wiki/research/post-training-framework/02-diloco-family.md` — DiLoCo / OpenDiLoCo / PRIME-RL / INTELLECT-2.463- `~/wiki/projects/composer-replication-framework.md` — current TL;DR and stage plan.464 465### Notes on accuracy466 467- "DAPO" labeling: OpenRLHF and Unsloth both advertise DAPO as a first-class loss type; PRIME-RL implements a DAPO-equivalent (decoupled-clip + KL) but uses the internal name `DPPO+KL` in its default loss. For our purposes this is the same family.468- Last-commit dates and release versions are pulled from GitHub release pages (OpenRLHF, PRIME-RL) and PyPI release history (`torchmonarch`).469- Star counts and contributor counts reflect the snapshots returned by web search at the time of writing (May 2026) and will drift; the relative ordering is stable.470 