Codeseys/composer-replication-framework
0
1# CITATION.cff — Citation File Format2# https://citation-file-format.github.io/3# Used by HF, GitHub, Zenodo to render a "Cite this repository" UI.4 5cff-version: 1.2.06message: "If you use this framework or its derivative artifacts, please cite as below."7type: software8title: "Composer 2.5 Replication Framework: Methodology and Integration Architecture for Open Replication of Cursor's Agentic Coding Recipe"9abstract: >10 An open-source methodology and integration architecture for replicating11 Cursor's Composer 2.5 recipe on a HuggingFace base model, plus a novel12 multi-teacher trace-replay distillation reward channel that complements13 the published SDPO/OPSD method (which Cursor's "Targeted RL with Textual14 Feedback" uses). Pre-experimental v0.0 release: methodology paper, audited15 recipe mapping, integration architecture across TRL/VeRL/OpenEnv,16 empirical economic-feasibility result for the novel channel ($0.98/trace),17 and a working code skeleton with 38 passing unit tests.18 19authors:20 - family-names: "Codeseys"21 given-names: ""22 affiliation: "Independent researcher"23 # Replace with real ORCID if available:24 # orcid: "https://orcid.org/0000-0000-0000-0000"25 26repository-code: "https://huggingface.co/Codeseys/composer-replication-framework"27url: "https://huggingface.co/Codeseys/composer-replication-framework"28date-released: "2026-05-25"29version: "0.0.0"30license: "MIT"31 32keywords:33 - reinforcement-learning34 - post-training35 - distillation36 - agentic-coding37 - composer-2.538 - cursor39 - kimi-k240 - grpo41 - dapo42 - sdpo43 - opsd44 - trl45 - verl46 - openenv47 - llm48 49# Primary upstream works this framework depends on / cites50references:51 - type: article52 title: "Introducing Composer 2.5"53 authors:54 - name: "Cursor Team"55 year: 202656 url: "https://cursor.com/blog/composer-2-5"57 58 - type: article59 title: "Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models"60 authors:61 - family-names: "Zhao"62 given-names: "Siyan"63 - family-names: "Xie"64 given-names: "Zhihui"65 - family-names: "Liu"66 given-names: "Mengchen"67 - family-names: "Huang"68 given-names: "Jing"69 - family-names: "Pang"70 given-names: "Guan"71 - family-names: "Chen"72 given-names: "Feiyu"73 - family-names: "Grover"74 given-names: "Aditya"75 year: 202676 url: "https://arxiv.org/abs/2601.18734"77 notes: "OPSD — single-LLM self-distillation; provides the reference loss implementation lifted by this framework."78 79 - type: article80 title: "Reinforcement Learning via Self-Distillation"81 authors:82 - family-names: "Hübotter"83 given-names: "Jonas"84 - family-names: "Lübeck"85 given-names: "Frederike"86 - family-names: "Behric"87 given-names: "Lejs"88 - family-names: "Baumann"89 given-names: "Anton"90 - family-names: "Bagatella"91 given-names: "Marco"92 - family-names: "Marta"93 given-names: "Daniel"94 - family-names: "Hakimi"95 given-names: "Ido"96 - family-names: "Shenfeld"97 given-names: "Idan"98 - family-names: "Buening"99 given-names: "Thomas Kleine"100 - family-names: "Guestrin"101 given-names: "Carlos"102 - family-names: "Krause"103 given-names: "Andreas"104 year: 2026105 url: "https://arxiv.org/abs/2601.20802"106 notes: "SDPO — formalizes the same mechanism as Cursor's Targeted RL with Textual Feedback. ICLR 2026 Scaling Post-training Workshop."107 