Addressing the structural vulnerability where standard GRPO broadcasts trajectory-level scalar advantages indiscriminately across heterogeneous tool calls and summary text, researchers introduce SLCA-GRPO. By establishing Segment-Locked Credit Assignment (SLCA), a Schema-Guided LLM Simulator (SGLS), and Hierarchical Rewards, the framework isolates execution advantages from preference noise without extra rollouts, driving a +9.15 pp gain on τ²-Bench and +1.36 pp on BFCL for 7B models.

Key Takeaways

  • ✓Eliminating CSCM Contamination: Fixes cross-segment credit misattribution (CSCM) where GRPO indiscriminately assigns trajectory-level advantages across heterogeneous text and tool segments within a single rollout group.
  • ✓Hierarchical Rewards & Zero-API Exploration: Introduces the Schema-Guided LLM Simulator (SGLS) and Hierarchical Rewards (HierR) to direct execution advantages to tool tokens and conversational preferences to summary tokens.
  • ✓Comprehensive Benchmark Gains: On a 7B backbone, outperforms standard GRPO, ToolPO, and RLTR by +2.53 pp in-domain, +1.36 pp on BFCL, and +9.15 pp on τ²-Bench while cutting tool redundancy and overhead.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 / Background & Pain Points Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with conversational natural language summaries. In standard on-policy Reinforcement Learning algorithms such as GRPO, trajectory-level scalar advantages are broadcast indiscriminately across all tokens. Consequently, gradient noise from free-form summary generation leaks into critical tool-selection tokens, creating Cross-Segment Credit Misattribution (CSCM) and optimization instability. ### 架构亮点与底层机制 / Architectural Highlights To resolve this structural flaw, researchers introduce SLCA-GRPO, defined by three primary components: 1. Segment-Locked Credit Assignment (SLCA): Decouples advantage estimation at structural segment boundaries within a single group of rollouts, eliminating the need for expensive intermediate branching rollouts; 2. Schema-Guided LLM Simulator (SGLS): Serves as deterministic foundational training infrastructure, bypassing expensive and non-deterministic real API invocations; 3. Hierarchical Rewards (HierR): Routes objective execution advantages exclusively to tool tokens and conversational preference advantages to summary tokens, completely cutting off the dominant cross-segment contamination channel. ### 权威 Benchmark 与实测跑分对比 / Benchmark & Evaluation Evaluated on a 7B backbone against standard GRPO, ToolPO, and RLTR under identical training budgets: - In-Domain Accuracy: Outperforms baselines by +2.53 pp with faster convergence; - Berkeley Function-Calling Leaderboard (BFCL): Improves generalization performance by +1.36 pp; - τ²-Bench: Achieves a dramatic +9.15 pp gain in multi-step nested agent environments; - Efficiency: Reduces redundant tool invocations and overall API execution overhead by 18.4%. ### 开发者实战落地与开箱指南 / Developer Practical Guide The complete codebase and datasets are fully open-source: - GitHub Repository: Accessible at https://github.com/SLCA-GRPO/SLCA-GRPO with PyTorch & TRL integration; - Dataset: Available at YanZhanPKU/SLCA-GRPO-Datasets on Hugging Face; - Best Practices: Enable segment decoupling and assign hierarchical loss scaling (0.7 tool / 0.3 summary) to guarantee stable convergence during agent post-training.