Reinforcement learning post-training generates massive rollout bottlenecks, driving interest in low-precision execution like NVIDIA Blackwell's native NVFP4 (W4A4). However, numerical discrepancies between learner and sampler forward passes destabilize policy-gradient optimization, triggering policy collapse. Researchers from Nanjing University and Polixir present TRIAGE (arXiv:2610.07043), isolating how quantization mismatches interact with policy-gradient directions. The authors reveal that native NVFP4 introduces an early asymmetry that amplifies negative-advantage updates across concentrated tail tokens. TRIAGE implements segment-level diagnostics to rebalance policy updates while preserving native W4A4 forward execution across both learners and samplers. Validated on Qwen3-4B and Qwen3-30B-A3B, TRIAGE matches full-precision BF16 performance across five mathematical reasoning benchmarks while delivering up to 2.3x higher rollout throughput.

Key Takeaways

  • ✓Nanjing University and Polixir introduce TRIAGE, stabilizing native NVFP4 (W4A4) RL training on NVIDIA Blackwell architecture
  • ✓Pioneers direction-aware policy-gradient stabilization and segment diagnostics, mitigating learner-sampler numerical mismatch
  • ✓Matches full BF16 precision on mathematical reasoning benchmarks while delivering 2.3x higher rollout throughput
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Background and the Problem

Reinforcement learning (RL) post-training for large language models demands immense computational resources, heavily bottlenecked by continuous environment rollout generation. While hardware architectures like NVIDIA Blackwell support native NVFP4 (W4A4 weight and activation precision), executing end-to-end RL at 4-bit precision has historically caused catastrophic policy collapse. Subtle numerical discrepancies between the sampling forward pass (Sampler) and the optimization forward pass (Learner) compound through policy-gradient updates, driving policy divergence and garbage generation.

Architecture and How It Works

Researchers from Nanjing University and Polixir introduce TRIAGE (arXiv:2610.07043), stabilizing native NVFP4 post-training:

  1. Direction-Aware Mismatch Analysis: Moves beyond absolute error magnitudes to inspect the interaction between quantization discrepancies and policy-gradient update vectors. Identifies an early-stage vulnerability where errors disproportionately amplify updates with negative advantage and negative probability gaps.
  2. Tail Token Concentration: Demonstrates that policy instability begins with localized tail token degradation within specific response segments before cascading across parameters.
  3. Segment-Level Dynamic Rebalancing: Employs lightweight segment-level diagnostics to rebalance divergent policy gradients and bounds residual severity.
  4. Native W4A4 Hardware Execution: Retains pure 4-bit forward execution across both learner and sampler pipelines without relying on high-precision fallbacks.

Benchmarks and Measured Results

Evaluated on Qwen3-4B and Qwen3-30B-A3B across intensive post-training schedules:

  1. Matches Full-Precision BF16 Baselines: Across five rigorous mathematical reasoning benchmarks, TRIAGE-stabilized NVFP4 runs achieve accuracy parity with full BF16 baselines.
  2. 2.3x Rollout Throughput Boost: Native NVFP4 execution delivers up to 2.3x higher generation throughput compared to standard BF16 pipelines.
  3. Rock-Solid Optimization Stability: Prevents gradient explosion and policy collapse throughout thousands of optimization steps, ensuring production-grade reliability.

Getting Started for Developers

TRIAGE provides an infrastructure blueprint for enterprise RL post-training. Engineering teams training reasoning agents (e.g. DeepSeek-R1 style) on modern Blackwell GPU clusters should transition from BF16-bound rollout nodes to native NVFP4. Incorporating segment-level direction-aware rebalancing eliminates numerical collapse, doubling RL post-training throughput while matching full-precision quality.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.