Reinforcement learning (RL) post-training is foundational to scaling reasoning capabilities in large language models. However, its memory footprint remains a prohibitive bottleneck: storing full-sized gradient buffers alongside multi-state Adam optimizer moments frequently triggers out-of-memory (OOM) failures when attempting to train large models (such as 27B parameters) on standard compute hardware. In newly released research (arXiv:2610.06647), researchers introduce LoGRA, an efficient RL post-training framework available within the open-source Molt library (github.com/skzhang1/labs-molt). LoGRA compresses training representations by retaining useful learning signals inside compact low-rank gradient sketches accumulated directly during backpropagation, facilitating both memory-efficient parameter updates and streamlined policy synchronization. To counteract instability from gradient compression, LoGRA introduces a Predicted-KL step control mechanism that forecasts policy change magnitudes prior to parameter updates, dynamically attenuating disruptive updates. Evaluated across reasoning benchmarks, LoGRA reduces average training memory by up to 45.7% with zero degradation in task performance, successfully sustaining stable RL training of a 27B model for over 1,100 steps on a single eight-GPU node where standard dense Adam crashes with OOM.

Key Takeaways

  • ✓LoGRA framework released, introducing low-rank gradient sketches to slash LLM RL training memory by up to 45.7%
  • ✓Breaks single-node 8-GPU memory barriers, sustaining stable RL training of a 27B model for over 1,100 steps
  • ✓Predicted-KL step control prevents policy drift, achieving full parity with standard AdamW on reasoning benchmarks
LoGRA: Scaling LLM Reinforcement Learning via Low-Rank Gradient Sketches, Slashing Memory by 45.7% for 27B Models
🖼️Official Media
Click to view high-res
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Background and the Problem

Reinforcement learning (RL) post-training (e.g., PPO, DPO, RLVR) is the engine behind advanced reasoning in models like DeepSeek-R1 and OpenAI o1. However, RL training requires substantially more GPU memory than supervised fine-tuning: storing multiple model copies alongside full-sized gradient tensors and multi-state Adam optimizer moments causes standard 8-GPU nodes to run out of memory (OOM) when training models larger than 20B parameters.

Architecture and How It Works

To overcome memory and communication bottlenecks, researchers introduce LoGRA (arXiv:2610.06647, code available in the Molt library at github.com/skzhang1/labs-molt):

  1. Low-Rank Gradient Sketches: Recognizing that RL gradients exhibit strong low-rank structure across backpropagation, LoGRA discards dense full-parameter gradient buffers. Gradients are dynamically projected and accumulated into compact low-rank sketches during backprop, dramatically reducing optimizer state footprint.
  2. Predicted-KL Step Control: To prevent gradient compression from causing destructive policy drift or collapse, LoGRA incorporates a predictive KL step controller that estimates the impending divergence in token probabilities before applying updates, dynamically scaling update magnitudes.
  3. Native Integration with Molt: Implemented cleanly in the PyTorch-native Molt RL library, fully interoperable with PyTorch FSDP and DeepSpeed ZeRO parallelism.

Benchmarks and Measured Results

Empirical verification across GSM8K, MATH, and HumanEval reasoning benchmarks demonstrates:

  1. Up to 45.7% Memory Reduction: Reduces average training memory by up to 45.7% compared to dense AdamW baselines while maintaining identical task learning dynamics.
  2. Stable 27B Model Training on a Single 8-GPU Node: While dense Adam crashes immediately with CUDA OOM on a standard single-node 8-GPU setup, LoGRA successfully sustains uninterrupted RL training of a 27B parameter model for over 1,100 steps.
  3. Parity in Task Accuracy: Post-training reasoning benchmarks match dense Adam baselines within ±0.3%, proving that low-rank compression preserves essential policy gradients without sacrificing downstream reasoning performance.

Getting Started for Developers

LoGRA is publicly accessible as part of the Molt library on GitHub (github.com/skzhang1/labs-molt). AI engineering teams and academic labs working on reasoning RL can activate LoGRA to scale model sizes from 7B to 27B+ on existing single-node clusters, democratizing frontier reinforcement learning without requiring multi-node cluster expansions.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.