Codeseys/composer-replication-framework
0
1# ADR-003 — DiLoCo implementation choice for Spike 0082 3**Status**: Accepted4**Date**: 2026-05-265**Wave**: Phase 4 (deep work loop)6 7## Context8 9Spike 008 closes V2 of the vision validation: DiLoCo was tagged "deferred to10v0.2" in Wave 5 but the original brief said "combine with DiLoCo." We want a11real working integration, not a hand-rolled toy.12 13The integration target: take pseudo-gradient `δ = θ_local − θ_initial` after N14inner steps, apply Nesterov-momentum outer step across replicas. We need a15PyTorch-compatible reference implementation that runs in single-process for16unit tests AND scales out on torch.distributed when we eventually run real17multi-replica training.18 19## Options considered20 21| Repo | License | Last commit | Maturity | Streaming variant? | Single-process testable? |22|---|---|---|---|---|---|23| `meta-pytorch/torchft` | BSD-3 | 2026-04-03 | Library (PyPI, prebuilt wheels, real test suite, Meta-maintained) | Yes (DiLoCo class IS the Streaming generalization; vanilla = single fragment) | Yes (verified via `MagicMock(Manager)` + `_DummyWork` pattern in their own tests) |24| `OpenDiLoCo` (PrimeIntellect) | Apache 2.0 | 2024 | README says "no longer maintained"; replaced by `prime` | Partial | Hivemind dependency complicates testing |25| `prime` / INTELLECT-1 (PrimeIntellect) | Apache 2.0 | 2025 | Production framework (`ElasticDeviceMesh` etc.) | Yes | Heavy harness; not single-process friendly |26| `diloco_simple` | **No LICENSE file** | 2024-05-31 | 8 commits ever; pedagogical | No | NCCL-locked |27| DeepMind original (Douillard et al. arXiv:2311.08105) | — | — | No public reference impl | — | — |28 29Source: `docs/research/DILOCO_RECONNAISSANCE.md` (subagent recon, 2026-05-26).30 31## Decision32 33**`meta-pytorch/torchft` — `torchft.local_sgd.DiLoCo`** (BSD-3, active,34single-process testable).35 36Rationale:37 381. **Library, not research code.** Proper packaging on PyPI with prebuilt39 wheels (`pip install torchft-nightly`), real test suite, version history,40 maintained by Meta. The other live candidates are research codebases that41 break on torch version bumps.42 432. **Streaming DiLoCo is the generalization.** The `DiLoCo` class accepts44 `model_fragments` + `fragment_sync_delay` + `fragment_update_alpha`. Set45 `model_fragments=[model]` (single fragment, full-model sync) for vanilla46 DiLoCo. Add fragments + per-fragment delays for Streaming. We don't have47 to choose at the API level — both modes are one parameter apart.48 493. **Single-process unit-testable.** torchft's own tests use50 `MagicMock(Manager)` + `_DummyWork` to bypass NCCL. We can do the same:51 shared-buffer mock allreduce that does real averaging across two52 in-process replicas. Verified working pattern in the recon doc.53 544. **Pseudo-gradient computation is in `_save_grads` (line 324) and55 `perform_sync` (line 423).** Direct extension point — we can subclass or56 monkey-patch these to compose with our Composer trainer.57 58### Risks accepted (with mitigations)59 60| Risk | Mitigation |61|---|---|62| **Sign convention mismatch** — torchft computes `θ_initial − θ_local` (negation of our spec) | Explicit unit test: assert outer step direction matches DiLoCo paper sign. Document the convention in our `outer_optimizer.py`. |63| **Wheel brittleness for nightly** | Pin a specific dated nightly slug in our pyproject.toml; bump deliberately. |64| **`torch>=2.7` requirement** | Confirm our existing eidolon venv has it. Already does (verified). |65| **`fragment_sync_delay > 0` requires CUDA streams** | Spike 008 uses `fragment_sync_delay=0` (vanilla DiLoCo) for the smoke. Streaming with non-zero delay deferred to v0.2 (post-replication). |66 67## Consequences68 69### Accepted70 71- Spike 008 imports `torchft.local_sgd.DiLoCo` and runs the recon doc's72 ready-to-paste pytest pattern as the smoke:73 - 2 replicas, 4 inner steps, 2 outer rounds on a TinyMLP74 - shared-buffer mock allreduce (no NCCL)75 - assertions: replica equality after sync, params actually moved, Nesterov76 state populated, sync count matches expected77 78- `composer_replication.diloco` package wraps `torchft.local_sgd.DiLoCo`79 with our trainer's hooks. We DO NOT fork torchft — we depend on it as a80 versioned wheel.81 82- Our integration is "vanilla DiLoCo" (single fragment, full-model sync) for83 v0.1. Streaming DiLoCo is a configuration-flag away in v0.2.84 85- The sign-convention mismatch is **made explicit** in our wrapper code with86 a unit test that catches a sign flip if torchft ever inverts it.87 88### Rejected paths89 90- **Roll our own DiLoCo.** Tempting (the algorithm is short) but the test91 surface for distributed-correctness is large; reusing a Meta-maintained92 library cuts the audit burden.93- **`diloco_simple`.** Disqualified by the license absence alone.94- **`prime` / INTELLECT-1.** Right tool for production multi-node runs,95 wrong tool for a single-process unit test.96 97## Source98 99`docs/research/DILOCO_RECONNAISSANCE.md` (subagent recon, primary-sourced100from torchft repo cloned + read locally, 2026-05-26).101 