Codeseys/composer-replication-framework
0
1---2title: Group Sequence Policy Optimization3id: group-sequence-policy-optimization4tags:5- deepread6created: '2026-06-10T00:30:47.451043Z'7source: https://arxiv.org/html/2507.180718source_domain: arxiv.org9fetched_at: '2026-06-10T00:30:47.450852Z'10fetch_provider: builtin11status: draft12type: note13tier: institutional14content_type: paper15deprecated: false16---17 18Group Sequence Policy Optimization19Title:20Content selection saved. Describe the issue below:21Description:22License: arXiv.org perpetual non-exclusive license23arXiv:2507.18071v2 [cs.LG] 28 Jul 202524Group Sequence Policy Optimization25Chujie Zheng Shixuan Liu Mingze Li Xiong-Hui Chen Bowen Yu26†27†28footnotemark:29Chang Gao Kai Dang Yuqiong Liu Rui Men An Yang Jingren Zhou Junyang Lin30Qwen Team, Alibaba Inc31Corresponding authors.32Abstract33This paper introduces Group Sequence Policy Optimization (GSPO), our stable, efficient, and performant reinforcement learning algorithm for training large language models.34Unlike previous algorithms that adopt token-level importance ratios, GSPO defines the importance ratio based on sequence likelihood and performs sequence-level clipping, rewarding, and optimization.35We demonstrate that GSPO achieves superior training efficiency and performance compared to the GRPO algorithm, notably stabilizes Mixture-of-Experts (MoE) RL training, and has the potential for simplifying the design of RL infrastructure.36These merits of GSPO have contributed to the remarkable improvements in the latest Qwen3 models.37138Introduction39Reinforcement learning (RL) has emerged as a pivotal paradigm for scaling language models40(OpenAI,41202442; DeepSeek-AI,43202544; Qwen,452025b46;47a48)49.50Through large-scale RL, language models develop the capability to tackle sophisticated problems, such as competition-level mathematics and programming, by undertaking deeper and longer reasoning processes.51To successfully scale RL with greater computational investment, the foremost prerequisite is maintaining stable and robust training dynamics.52However, current state-of-the-art RL algorithms, exemplified by GRPO53(Shao et al.,54202455)56, exhibit severe stability issues when training gigantic language models, often resulting in catastrophic and irreversible model collapse57(Qwen,582025a59; MiniMax,60202561)62.63This instability hinders efforts to push the capability boundaries of language models through continued RL training.64In this paper, we identify that the instability of GRPO stems from the fundamental misapplication and invalidation of importance sampling weights in its algorithmic design.65This introduces high-variance training noise that progressively accumulates with increased response length and is further amplified by the clipping mechanism, ultimately precipitating model collapse.66To address these core limitations, we propose67Group Sequence Policy Optimization (GSPO)68, a new RL algorithm for training large language models.69The key innovation of GSPO lies in its theoretically grounded definition of importance ratio based on sequence likelihood70(Zheng et al.,71202372)73, aligning with the basic principle of importance sampling.74Additionally, GSPO computes the normalized rewards as the advantages of multiple responses to a query, ensuring the alignment between sequence-level rewarding and optimization.75Our empirical evaluation demonstrates the significant superiority of GSPO over GRPO in training stability, efficiency, and performance.76Critically, GSPO has inherently resolved the stability challenges in the RL training of large Mixture-of-Experts (MoE) models, eliminating the need for complex stabilization strategies, and shows the potential for simplifying RL infrastructure.77These merits of GSPO ultimately contributed to the exceptional performance improvements in the latest Qwen3 models.78We envision GSPO as a robust and scalable algorithmic foundation that will enable the continued advancement of large-scale RL training with language models.79280Preliminaries81Notation82In this paper, an autoregressive language model parameterized by83θ84\theta85is defined as a policy86π87θ88\pi_{\theta}89.90We use91x92x93to denote a query and94𝒟95\mathcal{D}96as the query set.97Given a response98y99y100to a query101x102x103, its likelihood under the policy104π105θ106\pi_{\theta}107is denoted as108π109θ110111(112y113|114x115)116=117∏118t119=1201121|122y123|124π125θ126127(128y129t130|131x132,133y134<135t136)137\pi_{\theta}(y|x)=\prod_{t=1}^{|y|}\pi_{\theta}(y_{t}|x,y_{<t})138where139|140y141|142|y|143denotes the number of tokens in144y145y146.147A query-response pair148(149x150,151y152)153(x,y)154can be scored by a verifier155r156r157, resulting in a reward158r159160(161x162,163y164)165∈166[1670168,1691170]171r(x,y)\in[0,1]172.173Proximal Policy Optimization (PPO)174Using samples generated from the old policy175π176θ177old178\pi_{\theta_{\text{old}}}179, PPO180(Schulman et al.,1812017182)183constrains the policy update within a proximal region of the old policy through the clipping mechanism.184Specifically, PPO employs the following objective for policy optimization (we omit the KL regularization term hereinafter for brevity, as it is not the focus of this paper):185𝒥186PPO187188(189θ190)191=192𝔼193x194∼195𝒟196,197y198∼199π200θ201old202(203⋅204|205x206)207208[2091210|211y212|213214∑215t216=2171218|219y220|221min222223(224w225t226227(228θ229)230231A232^233t234,235clip236237(238w239t240241(242θ243)244,2451246−247ε248,2491250+251ε252)253254A255^256t257)258]259,260\displaystyle\mathcal{J}_{\text{PPO}}(\theta)=\mathbb{E}_{x\sim\mathcal{D},\,y\sim\pi_{\theta_{\text{old}}}(\cdot|x)}\left[\frac{1}{|y|}\sum_{t=1}^{|y|}\min\left(w_{t}(\theta)\widehat{A}_{t},\,\mathrm{clip}\left(w_{t}(\theta),1-{\varepsilon},1+{\varepsilon}\right)\widehat{A}_{t}\right)\right],261(1)262where the importance ratio of the token263y264t265y_{t}266is defined as267w268t269270(271θ272)273=274π275θ276277(278y279t280|281x282,283y284<285t286)287π288θ289old290291(292y293t294|295x296,297y298<299t300)301w_{t}(\theta)=\frac{\pi_{\theta}(y_{t}|x,y_{<t})}{\pi_{\theta_{\text{old}}}(y_{t}|x,y_{<t})}302,303the advantage304A305^306t307\widehat{A}_{t}308of309y310t311y_{t}312is estimated by another value model, and313ε314\varepsilon315is the clipping range of importance ratios.316The core challenge of PPO in practice lies in its heavy reliance on the value model.317Specifically, the value model usually has a similar size to the policy model, introducing a considerable memory and computational burden.318Furthermore, the algorithmic effectiveness hinges on the reliability of its value estimate.319While acquiring a reliable value model is inherently challenging, ensuring its scalability to longer responses and more complex tasks presents an even greater challenge.320Group Relative Policy Optimization (GRPO)321GRPO322(Shao et al.,3232024324)325bypasses the need for the value model by computing the relative advantage of each response within a group of responses to the same query.326Specifically, GRPO optimizes the following objective:327𝒥328GRPO329330(331θ332)333=334𝔼335x336∼337𝒟338,339{340y341i342}343i344=3451346G347∼348π349θ350old351(352⋅353|354x355)356357[3581359G360361∑362i363=3641365G3661367|368y369i370|371372∑373t374=3751376|377y378i379|380min381382(383w384i385,386t387388(389θ390)391392A393^394i395,396t397,398clip399400(401w402i403,404t405406(407θ408)409,4101411−412ε413,4141415+416ε417)418419A420^421i422,423t424)425]426,427\displaystyle\mathcal{J}_{\text{GRPO}}(\theta)=\mathbb{E}_{x\sim\mathcal{D},\,\{y_{i}\}_{i=1}^{G}\sim\pi_{\theta_{\text{old}}}(\cdot|x)}\left[\frac{1}{G}\sum_{i=1}^{G}\frac{1}{|y_{i}|}\sum_{t=1}^{|y_{i}|}\min\left(w_{i,t}(\theta)\widehat{A}_{i,t},\,\mathrm{clip}\left(w_{i,t}(\theta),1-{\varepsilon},1+{\varepsilon}\right)\widehat{A}_{i,t}\right)\right],428(2)429where430G431G432is the number of generated responses to each query433x434x435(i.e., the group size), and the importance ratio436w437i438,439t440441(442θ443)444w_{i,t}(\theta)445and advantage446A447^448i449,450t451\widehat{A}_{i,t}452of token453y454i455,456t457y_{i,t}458are:459w460i461,462t463464(465θ466)467=468π469θ470471(472y473i474,475t476|477x478,479y480i481,482<483t484)485π486θ487old488489(490y491i492,493t494|495x496,497y498i499,500<501t502)503,504A505^506i507,508t509=510A511^512i513=514r515516(517x518,519y520i521)522−523mean524525(526{527r528529(530x531,532y533i534)535}536i537=5381539G540)541std542543(544{545r546547(548x549,550y551i552)553}554i555=5561557G558)559,560\displaystyle w_{i,t}(\theta)=\frac{\pi_{\theta}(y_{i,t}|x,y_{i,<t})}{\pi_{\theta_{\text{old}}}(y_{i,t}|x,y_{i,<t})},\quad\ \widehat{A}_{i,t}=\widehat{A}_{i}=\frac{r(x,y_{i})-\mathrm{mean}\left(\{r(x,y_{i})\}_{i=1}^{G}\right)}{\mathrm{std}\left(\{r(x,y_{i})\}_{i=1}^{G}\right)},561(3)562respectively, where all the tokens in563y564i565y_{i}566share the same advantage as567A568^569i570\widehat{A}_{i}571.5723573Motivation574The growth in model size, sparsity (e.g., in Mixture-of-Experts models), and response length necessitates a large rollout batch size to maximize hardware utilization during RL.575To improve sample efficiency, it is standard practice to partition a large batch of rollout data into multiple mini-batches for gradient updates.576This procedure inevitably introduces an off-policy learning setting, where responses577y578y579are sampled from an old policy580π581θ582old583\pi_{\theta_{\text{old}}}584rather than the current policy585π586θ587\pi_{\theta}588being optimized.589This also explains the necessity of the clipping mechanism in PPO and GRPO, which prevents overly “off-policy” samples from being involved in gradient estimation.590While mechanisms like clipping aim to manage this off-policy discrepancy, we identify a more fundamental issue in GRPO:591its objective is ill-posed592.593This problem becomes particularly acute when training large models on long-response tasks, leading to catastrophic model collapse.594The ill-posed nature of the GRPO objective stems from a misapplication of importance sampling weights.595The principle of importance sampling is to estimate the expectation of a function596f597f598under a target distribution599π600tar601\pi_{\text{tar}}602by re-weighting samples drawn from a behavior distribution603π604beh605\pi_{\text{beh}}606:607𝔼608z609∼610π611tar612613[614f615616(617z618)619]620=621𝔼622z623∼624π625beh626627[628π629tar630631(632z633)634π635beh636637(638z639)640641f642643(644z645)646]647.648\displaystyle\mathbb{E}_{z\sim\pi_{\text{tar}}}\left[f(z)\right]=\mathbb{E}_{z\sim\pi_{\text{beh}}}\left[\frac{\pi_{\text{tar}}(z)}{\pi_{\text{beh}}(z)}\,f(z)\right].649(4)650Crucially, this relies on averaging over multiple samples (651N652≫6531654N\gg 1655) from the behavior distribution656π657beh658\pi_{\text{beh}}659for the importance weight660π661tar662663(664z665)666π667beh668669(670z671)672\frac{\pi_{\text{tar}}(z)}{\pi_{\text{beh}}(z)}673to effectively correct for the distributional mismatch.674In contrast, GRPO applies the importance weight675π676θ677678(679y680i681,682t683|684x685,686y687i688,689<690t691)692π693θ694old695696(697y698i699,700t701|702x703,704y705i706,707<708t709)710\frac{\pi_{\theta}(y_{i,t}|x,y_{i,<t})}{\pi_{\theta_{\text{old}}}(y_{i,t}|x,y_{i,<t})}711at each token position712t713t714.715Since this weight is based on a single sample716y717i718,719t720y_{i,t}721from each next-token distribution722π723θ724old725(726⋅727|728x729,730y731i732,733<734t735)736\pi_{\theta_{\text{old}}}(\cdot|x,y_{i,<t})737, it fails to perform the intended distribution-correction role.738Instead, it introduces high-variance noise into the training gradients, which accumulates over long sequences and is exacerbated by the clipping mechanism.739We have empirically observed that this can lead to model collapse that is often irreversible.740Once the collapse occurs, resuming training is unavailing, even when reverting to a previous checkpoint and meticulously tuning hyperparameters (e.g., the clipping ranges), extending generation length, or switching the RL queries.741The above observation suggests a fundamental issue in GRPO’s design.742The failure of the token-level importance weight points to a core principle:743the unit of optimization objective should match the unit of reward744.745Since the reward is granted to the entire sequence, applying off-policy correction at the token level appears problematic.746This motivates us to forego the token-level objective and explore utilizing importance weights and performing optimization directly at the747sequence level748.7494750Algorithm7514.1752GSPO: Group Sequence Policy Optimization753While the token-level importance weight754π755θ756757(758y759i760,761t762|763x764,765y766i767,768<769t770)771π772θ773old774775(776y777i778,779t780|781x782,783y784i785,786<787t788)789\frac{\pi_{\theta}(y_{i,t}|x,y_{i,<t})}{\pi_{\theta_{\text{old}}}(y_{i,t}|x,y_{i,<t})}790is problematic in GRPO, we observe that in the context of language generation, the791sequence-level792importance weight793π794θ795796(797y798|799x800)801π802θ803old804805(806y807|808x809)810\frac{\pi_{\theta}(y|x)}{\pi_{\theta_{\text{old}}}(y|x)}811has a clear theoretical meaning: it reflects how far the response812y813y814sampled from815π816θ817old818(819⋅820|821x822)823\pi_{\theta_{\text{old}}}(\cdot|x)824deviates from825π826θ827(828⋅829|830x831)832\pi_{\theta}(\cdot|x)833, which naturally aligns with the sequence-level reward and can also serve as a meaningful indicator of the clipping mechanism.834Based on this straightforward observation, we propose the835Group Sequence Policy Optimization (GSPO)836algorithm.837GSPO employs the following sequence-level optimization objective:838𝒥839GSPO840841(842θ843)844=845𝔼846x847∼848𝒟849,850{851y852i853}854i855=8561857G858∼859π860θ861old862(863⋅864|865x866)867868[8691870G871872∑873i874=8751876G877min878879(880s881i882883(884θ885)886887A888^889i890,891clip892893(894s895i896897(898θ899)900,9011902−903ε904,9051906+907ε908)909910A911^912i913)914]915,916\displaystyle\mathcal{J}_{\text{GSPO}}(\theta)=\mathbb{E}_{x\sim\mathcal{D},\,\{y_{i}\}_{i=1}^{G}\sim\pi_{\theta_{\text{old}}}(\cdot|x)}\left[\frac{1}{G}\sum_{i=1}^{G}\min\left(s_{i}(\theta)\widehat{A}_{i},\,\mathrm{clip}\left(s_{i}(\theta),1-{\varepsilon},1+{\varepsilon}\right)\widehat{A}_{i}\right)\right],917(5)918where we adopt the group-based advantage estimation:919A920^921i922=923r924925(926x927,928y929i930)931−932mean933934(935{936r937938(939x940,941y942i943)944}945i946=9471948G949)950std951952(953{954r955956(957x958,959y960i961)962}963i964=9651966G967)968,969\displaystyle\widehat{A}_{i}=\frac{r(x,y_{i})-\mathrm{mean}\left(\{r(x,y_{i})\}_{i=1}^{G}\right)}{\mathrm{std}\left(\{r(x,y_{i})\}_{i=1}^{G}\right)},970(6)971and define the importance ratio972s973i974975(976θ977)978s_{i}(\theta)979based on sequence likelihood980(Zheng et al.,9812023982)983:984s985i986987(988θ989)990=991(992π993θ994995(996y997i998|999x1000)1001π1002θ1003old10041005(1006y1007i1008|1009x1010)1011)101211013|1014y1015i1016|1017=1018exp10191020(102111022|1023y1024i1025|10261027∑1028t1029=103011031|1032y1033i1034|1035log10361037π1038θ10391040(1041y1042i1043,1044t1045|1046x1047,1048y1049i1050,1051<1052t1053)1054π1055θ1056old10571058(1059y1060i1061,1062t1063|1064x1065,1066y1067i1068,1069<1070t1071)1072)1073.1074\displaystyle s_{i}(\theta)=\left(\frac{\pi_{\theta}(y_{i}|x)}{\pi_{\theta_{\text{old}}}(y_{i}|x)}\right)^{\frac{1}{|y_{i}|}}=\exp\left(\frac{1}{|y_{i}|}\sum_{t=1}^{|y_{i}|}\log\frac{\pi_{\theta}(y_{i,t}|x,y_{i,<t})}{\pi_{\theta_{\text{old}}}(y_{i,t}|x,y_{i,<t})}\right).1075(7)1076Therefore, GSPO applies clipping to entire responses instead of individual tokens to exclude the overly “off-policy” samples from gradient estimation, which matches both the sequence-level rewarding and optimization.1077Note that we adopt length normalization in1078s1079i10801081(1082θ1083)1084s_{i}(\theta)1085to reduce the variance and to control1086s1087i10881089(1090θ1091)1092s_{i}(\theta)1093within a unified numerical range.1094Otherwise, the likelihood changes of a few tokens can result in dramatic fluctuations of the sequence-level importance ratio, and the importance ratios of responses with different lengths will require varying clipping ranges.1095We also note that the clipping ranges in GSPO and in previous algorithms (e.g., GRPO) typically differ in order of magnitude due to the distinct definitions of importance ratios.10964.21097Gradient Analysis1098We can derive the gradient of the GSPO objective as follows (clipping is omitted for brevity):1099∇1100θ1101𝒥1102GSPO11031104(1105θ1106)1107=1108\displaystyle\nabla_{\theta}\mathcal{J}_{\text{GSPO}}(\theta)=1109∇1110θ1111𝔼1112x1113∼1114𝒟1115,1116{1117y1118i1119}1120i1121=112211123G1124∼1125π1126θ1127old1128(1129⋅1130|1131x1132)11331134[113511136G11371138∑1139i1140=114111142G1143s1144i11451146(1147θ1148)11491150A1151^1152i1153]1154\displaystyle\ \nabla_{\theta}\mathbb{E}_{x\sim\mathcal{D},\,\{y_{i}\}_{i=1}^{G}\sim\pi_{\theta_{\text{old}}}(\cdot|x)}\left[\frac{1}{G}\sum_{i=1}^{G}s_{i}(\theta)\widehat{A}_{i}\right]1155(8)1156=1157\displaystyle=1158𝔼1159x1160∼1161𝒟1162,1163{1164y1165i1166}1167i1168=116911170G1171∼1172π1173θ1174old1175(1176⋅1177|1178x1179)11801181[118211183G11841185∑1186i1187=118811189G1190s1191i11921193(1194θ1195)11961197A1198^1199i1200⋅