While looped transformers achieve high parameter efficiency by iteratively recycling layers for latent compute, fixed-depth looping expends redundant iterations on simple tokens and suffers from performance saturation. Tsinghua University's NICS lab developed TaH2, an adaptive looped framework that co-trains the backbone and an iteration decider via lookahead depth supervision. Evaluated on the rigorous AIME competition benchmarks, TaH2 accelerates the accuracy-compute scaling slope by 53% (2.74 vs 1.79) and surpasses non-looped baselines by +3.4 to +3.9 points at matched test-time compute.

Key Takeaways

  • ✓Adaptive Looping via Lookahead Depth Supervision: Discards rigid, uniform layer looping by training a lightweight iteration decider using online labels that predict whether subsequent recurrent passes genuinely improve token predictions.
  • ✓53% Steeper Test-Time Compute Scaling Slope: On challenging AIME competition benchmarks, improves the accuracy-to-compute slope from 1.79 to 2.74, outpacing non-looped baselines by +3.4 points at matched decoding FLOPs.
  • ✓Overcoming Looping Saturation: While existing recurrent transformers plateau beyond shallow depths, TaH2's performance margin over dense models expands continuously from +2.8 points at depth 2 to +3.9 points at depth 8.
  • ✓Complete Codebase Open-Sourced: Full post-training recipes, decider checkpoints, and inference evaluation harnesses released at github.com/thu-nics/TaH.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 / Background & Pain Points Test-time compute scaling is a vital mechanism for pushing frontier reasoning limits on competition mathematics and complex coding tasks. Looped transformers, which reuse parameters across recurrent steps, represent an elegant, parameter-efficient foundation for dynamic latent reasoning. However, mainstream looped architectures apply a rigid, fixed iteration depth to every token. This uniform compute expenditure wastes cycles on elementary tokens while rapidly hitting an accuracy ceiling as recurrent depths expand. ### 架构亮点与底层机制 / Architectural Highlights Researchers from Tsinghua University's NICS lab developed TaH2: 1. Lookahead Depth Supervision: Co-trains the foundation backbone alongside an iteration decider utilizing dynamic online supervision signals that ascertain whether additional layer loops genuinely enhance token prediction fidelity; 2. Dynamic Token Compute Allocation: Allows easily resolved tokens to exit early after shallow passes, concentrating recursive depth on critical mathematical branches and branching decisions; 3. Decoupled Training Regularization: Formulates a joint loss that aligns decider confidence with predictive correctness without destabilizing the shared recurrent representation. ### 权威 Benchmark 与实测跑分对比 / Benchmark & Evaluation Rigorous testing on the American Invitational Mathematics Examination (AIME) confirmed landmark test-time efficiency: - 53% Steeper Scaling Slope: Elevates the accuracy-compute scaling slope from 1.79 in dense baselines to 2.74 (a 53% improvement) per doubling of decoding FLOPs; - Higher Peak Accuracy: Outperforms dense non-looped baselines by +3.4 points at matched computational budgets; - Eliminating Plateau Saturation: While legacy recurrent models stagnate beyond shallow depths, TaH2's advantage over the baseline expands from +2.8 points at depth 2 to +3.9 points at depth 8. ### 开发者实战落地与开箱指南 / Developer Practical Guide - Open-Source Code: Available at https://github.com/thu-nics/TaH; - Paper Citation: Theoretical foundations are detailed in arXiv preprint 2609.35748; - Deployment Advice: Compatible with adaptive inference engines, making TaH2 an ideal architecture for running reasoning agents on resource-constrained consumer GPUs.