Training & Alignment
21 min
Why Token-Level RL Collapses: GSPO and Sequence-Level Importance Sampling
GRPO weights every token by its own importance ratio, and on long responses that single-sample estimator quietly poisons the gradient until the model collapses. GSPO moves the ratio up to the whole sequence, and Qwen3's largest m…