Reinforcement Learning from Verifiable Rewards (RLVR) with long Chain-of-Thought (CoT) trajectories powers frontier reasoning models such as DeepSeek-R1 and OpenAI o-series. However, live on-policy sampling consumes more than 80% of cluster compute during post-training. To control training costs, practitioners frequently reuse rollout trajectories across multiple optimization steps, introducing severe off-policy distribution mismatch between the active policy and the rollout-generating policy. Researchers from UCLA unveil a fundamental failure mode in clipped policy optimization (arXiv:2609.35433): 'Sign-Dependent Gradient Starvation'. Standard clipping mechanisms in PPO and GRPO suppress under-generated positive reasoning rollouts relegated to the low-importance-weight tail, while allowing severely over-generated negative trajectories to dominate high-weight updates. This causes models to discard promising reasoning paths during early training. The authors propose ReSPO (Reshaped Sequence Policy Optimization), replacing rigid clipping with a smooth, two-branch sequence-level kernel derived from an alpha-divergence variational objective and an exponential variance-control tilt. Across dense and MoE Qwen3 architectures, ReSPO rescues long positive reasoning rollouts, accelerates early convergence, and achieves superior benchmark accuracy under rollout reuse.

Key Takeaways

  • ✓UCLA identifies sign-dependent gradient starvation in off-policy RLVR, proposing ReSPO to replace clipped surrogate objectives
  • ✓Protects sparse positive reasoning rollouts at low-weight tails, boosting early effective gradient throughput by 2.4x on Qwen3 models
  • ✓Maintains monotonic convergence under 4x rollout reuse, elevating benchmark accuracy by 3.6 to 5.2 points while cutting sampling compute by 50%
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Background and the Problem

Reinforcement Learning from Verifiable Rewards (RLVR) represents the dominant paradigm for frontier reasoning models like DeepSeek-R1 and OpenAI o-series. However, rolling out long Chain-of-Thought (CoT) trajectories accounts for 80% to 90% of total compute in post-training clusters. To optimize throughput, teams implement rollout reuse across 2 to 4 policy gradient steps. This reuse inevitably introduces off-policy distribution mismatch. Standard algorithms rely on clipped surrogate objectives (PPO, GRPO), which introduce severe stability failures across long-horizon trajectories.

Architecture and How It Works

Researchers from UCLA analyze importance weight tails in off-policy reasoning and present ReSPO (arXiv:2609.35433):

  1. Sign-Dependent Gradient Starvation: In long reasoning traces, novel positive trajectories quickly slip into the low-importance-weight tail following slight parameter shifts. PPO clipping forces gradients for these under-generated positive samples to zero, causing models to abandon correct solutions. Concurrently, dominant negative trajectories at the high-weight tail deliver excessive punitive gradients, destabilizing convergence.
  2. Smooth Two-Branch ReSPO Kernel: Replaces rigid clipping with a two-branch kernel derived from an alpha-divergence objective and exponential variance-control tilt:
  • Positive Branch: Preserves non-zero gradient flow for low-importance positive rollouts, enabling early learning from sparse successes.
  • Negative Branch: Dampens high-weight negative trajectories to prevent gradient shock.
  1. Sequence-Level Variance Normalization: Aggregates token ratios into sequence-level importance metrics, curbing variance explosion across 32K+ context lengths.

Benchmarks and Measured Results

Evaluated on dense and MoE Qwen3 architectures across reasoning benchmarks:

  1. 2.4x Effective Early Gradient Throughput: During early training on MATH-500 and AIME, ReSPO accelerates reward convergence rates by 2.4x over clipped PPO and GRPO.
  2. Robust Under 4x Rollout Reuse: When trajectories are reused four times, standard PPO suffers policy degradation, whereas ReSPO achieves 3.6 to 5.2 percentage point accuracy improvements over baselines.
  3. Preserving Exploration Depth: Trajectory audits confirm that ReSPO prevents reasoning chain collapse, allowing models to maintain deep exploration and backtrack heuristics.

Getting Started for Developers

ReSPO provides a drop-in replacement for clipped policy optimization in post-training frameworks (e.g., OpenRLHF, verl). Platform engineers should replace PPO/GRPO clipping kernels with ReSPO's two-branch importance weighting. By increasing rollout reuse from 1 to 3-4 steps without degrading policy convergence, enterprise RL clusters can slash sampling compute expenses by over 50% while unlocking superior multi-step reasoning capabilities.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.