Reinforcement learning (RL) post-training is foundational to scaling reasoning capabilities in large language models. However, its memory footprint remains a prohibitive bottleneck: storing full-sized gradient buffers alongside multi-state Adam optimizer moments frequently triggers out-of-memory (OOM) failures when attempting to train large models (such as 27B parameters) on standard compute hardware. In newly released research (arXiv:2610.06647), researchers introduce LoGRA, an efficient RL post-training framework available within the open-source Molt library (github.com/skzhang1/labs-molt). LoGRA compresses training representations by retaining useful learning signals inside compact low-rank gradient sketches accumulated directly during backpropagation, facilitating both memory-efficient parameter updates and streamlined policy synchronization. To counteract instability from gradient compression, LoGRA introduces a Predicted-KL step control mechanism that forecasts policy change magnitudes prior to parameter updates, dynamically attenuating disruptive updates. Evaluated across reasoning benchmarks, LoGRA reduces average training memory by up to 45.7% with zero degradation in task performance, successfully sustaining stable RL training of a 27B model for over 1,100 steps on a single eight-GPU node where standard dense Adam crashes with OOM.
Key Takeaways
- ✓LoGRA framework released, introducing low-rank gradient sketches to slash LLM RL training memory by up to 45.7%
- ✓Breaks single-node 8-GPU memory barriers, sustaining stable RL training of a 27B model for over 1,100 steps
- ✓Predicted-KL step control prevents policy drift, achieving full parity with standard AdamW on reasoning benchmarks

Key Decision Metrics at a Glance
Turn your technical choice into a development budget
Compare 40 dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Background and the Problem
Reinforcement learning (RL) post-training (e.g., PPO, DPO, RLVR) is the engine behind advanced reasoning in models like DeepSeek-R1 and OpenAI o1. However, RL training requires substantially more GPU memory than supervised fine-tuning: storing multiple model copies alongside full-sized gradient tensors and multi-state Adam optimizer moments causes standard 8-GPU nodes to run out of memory (OOM) when training models larger than 20B parameters.
Architecture and How It Works
To overcome memory and communication bottlenecks, researchers introduce LoGRA (arXiv:2610.06647, code available in the Molt library at github.com/skzhang1/labs-molt):
- Low-Rank Gradient Sketches: Recognizing that RL gradients exhibit strong low-rank structure across backpropagation, LoGRA discards dense full-parameter gradient buffers. Gradients are dynamically projected and accumulated into compact low-rank sketches during backprop, dramatically reducing optimizer state footprint.
- Predicted-KL Step Control: To prevent gradient compression from causing destructive policy drift or collapse, LoGRA incorporates a predictive KL step controller that estimates the impending divergence in token probabilities before applying updates, dynamically scaling update magnitudes.
- Native Integration with Molt: Implemented cleanly in the PyTorch-native Molt RL library, fully interoperable with PyTorch FSDP and DeepSpeed ZeRO parallelism.
Benchmarks and Measured Results
Empirical verification across GSM8K, MATH, and HumanEval reasoning benchmarks demonstrates:
- Up to 45.7% Memory Reduction: Reduces average training memory by up to 45.7% compared to dense AdamW baselines while maintaining identical task learning dynamics.
- Stable 27B Model Training on a Single 8-GPU Node: While dense Adam crashes immediately with CUDA OOM on a standard single-node 8-GPU setup, LoGRA successfully sustains uninterrupted RL training of a 27B parameter model for over 1,100 steps.
- Parity in Task Accuracy: Post-training reasoning benchmarks match dense Adam baselines within ±0.3%, proving that low-rank compression preserves essential policy gradients without sacrificing downstream reasoning performance.
Getting Started for Developers
LoGRA is publicly accessible as part of the Molt library on GitHub (github.com/skzhang1/labs-molt). AI engineering teams and academic labs working on reasoning RL can activate LoGRA to scale model sizes from 7B to 27B+ on existing single-node clusters, democratizing frontier reinforcement learning without requiring multi-node cluster expansions.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.