Codeseys/composer-replication-framework
0
1# ADR-002 — Trace source for Spike 007 (real LLM-application traces)2 3**Status**: Accepted4**Date**: 2026-05-265**Wave**: Phase 4 (deep work loop)6 7## Context8 9Spike 007 closes V5 of the vision validation: "real LLM-application traces."10Spike 001 used 50 hand-crafted synthetic states for the cost-floor measurement.11The framework's brief explicitly said *real traces*, so we owe Spike 007 a12primary-sourced ingestion path that converts a real, public, multi-turn agent13trace format into our existing `TraceState` TypedDict.14 15Existing schema (verified from `spikes/005-integrated-trainer-skeleton/teacher_replay.py`):16 17```python18class TraceState(TypedDict):19 state_id: str # unique within the trace20 messages: list[dict] # OpenAI-style conversation up to + incl this step21 student_action: str # what the student did at this step22```23 24(Earlier deep-work-loop notes called this `TraceExample` — that was a brain25glitch; the actual type is `TraceState` and there is no `TraceExample`.)26 27## Options considered28 29| Option | Schema | Acquisition | Signal density | License |30|---|---|---|---|---|31| (a) Claude Code session JSONL | Documented + 4 reverse-engineered schemas | **1,015 local sessions** zero-cost | per-step `tool_use` blocks = ideal teacher-correction sites | User-owned local files; framework MIT |32| (b) Cline VS Code extension | No stable export schema | Would need custom extraction | Unknown until extracted | Apache 2.0 (extension), trace data user-owned |33| (c) OpenHands trajectories | Documented (v0/v1 in flux) | Need to run OpenHands or download leaderboard submissions | Strong | MIT |34| (d) Aider chat history | Markdown chat (lossy for tool calls) | Local only if user runs Aider | Weak — collapses tool structure | Apache 2.0 |35| (e) SWE-bench leaderboard trajs | Heterogeneous, free-format | Public download | Strong but uneven | Per-submission (mostly permissive) |36| (f) SWE-smith-trajectories (HF) | Messages-only, structure collapsed | HF dataset download | Strong but lossy | MIT |37 38Source: `docs/research/TRACE_SOURCE_RECONNAISSANCE.md` (2026-05-26 subagent recon).39 40## Decision41 42**Option (a) — Claude Code session JSONL** at `~/.claude/projects/<encoded>/<sessionid>.jsonl`.43 44Wins on every axis we care about for Spike 007:45 461. **Acquisition cost: zero.** 1,015 real sessions already on this machine47 from the user's daily Claude Code use. No download, no consent48 negotiation, no rate limiting, no schema change risk during ingestion49 development.50 512. **Schema stability: empirically validated.** The subagent ran a programmatic52 audit on 8 real sessions; record types are stable across all of them.53 Anthropic publishes user-facing docs for the format; four independent54 community projects (claude-code-cli-tools, claudeflow, etc.) ship55 working parsers including one with a JSON Schema validated against56 ~50,000 real messages.57 583. **Signal density: maximal.** Every `tool_use` block is a candidate59 teacher-correction site. The 5 pre-selected sessions in the recon doc60 contain 6,762 tool_use messages (range 125 → 2,830 per session). That's61 100× the density of Spike 001's 50 synthetic states.62 634. **License: clean.** The trace files are user-owned files on the user's64 own machine. We don't redistribute them with the framework. The65 *ingester* code we write is MIT and ships in the framework. Anyone66 running the framework who wants real-trace ingestion uses their own67 local Claude Code sessions.68 69## Consequences70 71### Accepted72 73- Spike 007 implements `TraceIngester.ingest(path: Path) -> Iterator[TraceState]`74 for the Claude Code JSONL format.75- The TraceIngester ships as part of the package (Wave 10 packaging) under76 `composer_replication.ingestion.claude_code`.77- The recon doc's 5 pre-selected real sessions become the **smoke fixture**78 for Spike 007's tests. We pin to a known set of session IDs so the test79 is deterministic locally; CI users substitute their own.80- `ingestion/` directory pattern is established now to support adding81 ingesters for OpenHands and SWE-smith later if Spike 007 reveals82 signal-density gaps.83 84### Open questions resolved by ADR-00285 861. **Granularity** — One `TraceState` per assistant turn (not per `tool_use`).87 A single assistant turn often emits multiple `tool_use` blocks for one88 reasoning step; treating each tool_use as a separate state would89 over-fragment the conversation. Discussion in TRACE_SOURCE_RECONNAISSANCE90 §5.91 922. **`student_action` mapping** — The literal text of the assistant turn93 (concatenated `text` blocks of the Claude message) becomes94 `student_action`. The teacher-replay channel asks N teachers to produce95 their version of "what should the assistant do here?" given the96 `messages` history; we then DPO-compare teacher consensus vs literal97 student text.98 993. **Thinking blocks** — Strip `thinking` blocks from the message history100 passed to teachers (teachers don't have access to Claude's reasoning101 trace). KEEP them in the `student_action` for the student's own102 reproduction loop, since that's the actual generation we'd be RL-training.103 1044. **System prompt** — Inject a synthetic system prompt at message[0] of105 each `TraceState` describing "you are a coding agent" so teachers106 without their own coding-agent system prompt have a fair playing field.107 1085. **Subagent traces** — Skip them in v0.1; only ingest top-level sessions.109 Subagent traces have a different structure (parent task ID etc.) that110 would complicate the v0.1 ingester.111 112### Recon-flagged risk (not blocking)113 114- Anthropic doesn't publish a versioned schema. The TraceIngester pins to115 known record-types as of 2026-05-26 and gracefully degrades on unknown116 types. If Anthropic ships a breaking change to the JSONL format, we'd117 need to bump a `schema_version` constant in the ingester. Acceptable118 ongoing maintenance burden.119 120### Risk added 2026-05-26 by cross-model review (NOT BLOCKING but TO DOCUMENT)121 122- **Circularity / data-leakage in the teacher-replay channel.** Claude123 Code traces are produced by Claude. Our default teacher pool124 (`DEFAULT_TEACHERS`) includes `anthropic/claude-opus-4.7`. Training a125 student on Claude's outputs while Claude is one of the teachers126 voting on what the student should do produces a biased disagreement127 signal: Claude's vote is correlated with the trace's existing128 `student_action` (which Claude originally produced). This biases the129 multi-teacher consensus toward the existing answer.130 - **Mitigation**: when ingesting Claude Code traces, the user should131 drop Claude from the teacher pool and use a non-Claude consensus132 (Opus 4.7 → GPT-5 + DeepSeek V4-Pro, or any non-Claude pair).133 Documented here; not yet enforced in code.134 - **Open question for v0.2**: should `ClaudeCodeIngester` automatically135 annotate the source-model field on each trace and `replay_trace`136 automatically exclude same-family teachers? Defer the design until137 the post-replication phase reveals whether the bias is observable.138 139### Future ingesters140 141Open the door for two more ingesters in v0.2:142- `composer_replication.ingestion.openhands` — for users who run OpenHands143- `composer_replication.ingestion.swe_smith` — for users who download the HF dataset144 145Both follow the same `Iterator[TraceState]` contract.146 147## Source148 149`docs/research/TRACE_SOURCE_RECONNAISSANCE.md` (subagent recon, primary-sourced150including direct inspection of the user's local sessions, 2026-05-26).151 