Developer 3s Key Decision Metrics
Endowing LLM agents with test-time self-improvement—the ability to iteratively refine solutions—is a vital milestone for autonomous engineering. Researchers from Renmin University and collaborators present AREX-2, an open-source framework leveraging long-horizon reflective trajectories. Decoupling self-improvement into reflection (generating superior alternatives) and long-horizon execution (sustaining iteration over multiple turns), AREX-2 trains on verified machine learning and algorithmic programming trajectories. Built on Qwen3.8-27B, AREX-2 achieves 81.8 on MLE-bench Lite and 70.7 on Frontier-CS, transferring effectively to deep research with 92.2 on GAIA and 93.8 on DeepSearchQA.
Key Takeaways
- ✓Decouples test-time self-improvement into domain-agnostic reflection and long-horizon execution
- ✓Achieves 81.8 on MLE-bench Lite and 70.7 on Frontier-CS with Qwen3.8-27B, transferring to 92.2 on GAIA
- ✓Breaks multi-turn degradation bottlenecks, showing monotonic score improvements as iteration budgets increase
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.