Reinforcement learning for code agents typically relies on executable tests for binary rewards, leading algorithms like GRPO to assign identical scalar advantages to all passing rollouts regardless of code bloat, extraneous modifications, or architectural hygiene. Researchers from ByteDance, Peking University, and Tsinghua introduced GAGAR, a quality-aware advantage redistribution framework. By staging rollouts in a shared workspace where an SFT agentic grader ranks passing candidates and dynamically shifts credit toward clean solutions, GAGAR successfully stabilizes code agent RL on 310B Flash and 1.02T Pro models.

Key Takeaways

  • ✓Overcoming Binary Verification Blind Spots: Conventional code agent RL treats all test-passing rollouts equally, causing policy optimization to reward bloated, fragile, or out-of-scope code as long as test assertions pass.
  • ✓Sum-Preserving Credit Redistribution: Stages rollout groups in a shared context where an SFT-trained agentic judge ranks solutions, downweighting sloppy passes while proportionally shifting advantages to clean, minimal-diff implementations without altering total group advantage magnitude.
  • ✓Industrial Validation on 1.02T Model: Evaluated at scale on MiMo-V2.6-Flash (310B) and MiMo-V2.6-Pro (1.02T), GAGAR curbs trajectory-length explosion, enhances training stability, and substantially elevates code cleanliness.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 / Background & Pain Points Reinforcement learning for software engineering agents predominantly leverages executable test suites for binary rewards (1 for passing, 0 for failure). Under standard Group Relative Policy Optimization (GRPO), all test-passing rollouts within a group receive identical positive scalar advantages. This design blind spot treats concise, maintainable 5-line patches identically to messy 100-line diffs that inject out-of-scope hacks, causing coding policies to degrade into generating bloated, fragile implementations over training iterations. ### 架构亮点与底层机制 / Architectural Highlights Researchers from ByteDance, Peking University, and Tsinghua developed GAGAR: 1. Dynamic Group Sampling: Focuses rollout retention on groups exhibiting diverse outcomes (both successes and failures) to maximize gradient quality; 2. Shared-Context Agentic Grader: Consolidates passing candidate implementations into a shared workspace where a specialized SFT agentic judge inspects code aesthetics, scope discipline, and structural simplicity, establishing an ordinal quality ranking; 3. Sum-Preserving Advantage Redistribution: Downweights advantages for lower-ranked implementations and redistributes the subtracted advantage proportionally onto top-ranked solutions, preserving the exact mathematical sum of group advantages to ensure stable policy convergence. ### 权威 Benchmark 与实测跑分对比 / Benchmark & Evaluation Evaluated at enterprise scale across pre-RL checkpoints of MiMo-V2.6-Flash (310B) and MiMo-V2.6-Pro (1.02T parameters): - Taming Trajectory Bloat: Restrains artificial trajectory-length inflation by 14.2% relative to standard GRPO baselines; - Clean Implementation Preference: Greatly improves in-scope modification rates, steering agents toward minimal, surgical diffs; - Trillion-Scale Stability: Successfully stabilizes RLVR on the 1.02T Pro checkpoint without training loss divergence or advantage collapse. ### 开发者实战落地与开箱指南 / Developer Practical Guide - Technical Reference: Detailed algorithms are formalized in arXiv preprint 2609.32577; - Training Pipeline Adoption: Coding agent post-training pipelines should insert an SFT judge to re-weight test-passing advantages before policy updates; - Lightweight Overhead: Grader inference can be executed asynchronously across decoupled evaluation nodes without throttling high-throughput GPU training clusters.