LLM-guided evolutionary algorithms (such as AlphaEvolve) have achieved breakthrough results in computational optimization, including circle packing and systems tuning. However, existing methods optimize strictly over iteration counts rather than fiscal expenditure, burning dozens of dollars per task. Researchers from NUS, Stanford, and HKUST present FrugalEvo (arXiv:2610.03675), a cost-aware program evolution framework maximizing gain per dollar. FrugalEvo implements an asymmetric dual-LLM division of labor: a powerful, higher-cost LLM explores high-level solution concepts, while a cheap, high-speed LLM generates and iteratively refines concrete implementations. Additionally, FrugalEvo engineers prompt harnesses to maximize prefix KV-cache reuse across evolutionary generations. Across 10 mathematical optimization problems and ALE-Bench-Lite suites, FrugalEvo matches or outperforms OpenEvolve, EvoX, and SwarmResearch. On circle packing, FrugalEvo achieves SOTA performance using GLM-5.3/Flash for only $0.55 and GPT-5.6 Terra/Luna for $1.68—slashing costs by up to 99% compared to conventional $50 multi-agent baselines.
Key Takeaways
- ✓Pioneers FrugalEvo, the first cost-aware LLM program evolution framework maximizing optimization gain per dollar spent
- ✓Pairs an asymmetric dual-LLM hierarchy with prefix KV-cache reuse, establishing the new Budget-Aware AUC (BA-AUC) benchmark
- ✓Achieves SOTA on circle packing for $0.55-$1.68, outperforming $50 multi-agent baselines like CORAL and SwarmResearch

Developer 3s Key Decision Metrics
Turn your technical choice into a development budget
Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
核心背景与行业痛点
LLM-guided evolutionary algorithms (such as AlphaEvolve and EvoX) have demonstrated potent capabilities for computational optimization, including mathematical packing, compiler tuning, and tensor kernel synthesis. However, conventional evolutionary pipelines suffer from a fundamental cost-blindness: frameworks optimize strictly across fixed generation iterations using expensive flagship models (e.g., GPT-4o, Claude 3.5 Sonnet), routinely running up bills of $50 to hundreds of dollars per single optimization run. Maximizing performance gain per dollar of budget represents the essential catalyst for deploying evolutionary coding agents into enterprise engineering pipelines.
架构亮点与底层机制
Researchers from NUS, Stanford University, and HKUST introduce FrugalEvo (arXiv:2610.03675), a cost-aware program evolution framework:
- Asymmetric Dual-LLM Hierarchy: Decouples strategic ideation from syntax implementation. A strong, higher-cost LLM is invoked sparingly to conceive novel conceptual strategies, while a high-throughput, low-cost LLM handles the voluminous tasks of code implementation, debugging, and iterative local refinements.
- Prefix-Cache-Aware Evolution Pipeline: Restructures evolutionary prompt harnesses to ensure shared token prefixes across successive generations, maximizing KV-cache reuse in modern inference engines (e.g., vLLM, SGLang, and commercial cache endpoints) and slashing input token charges by over 80%.
- Budget-Aware AUC (BA-AUC) Metric: Formulates Budget-Aware Area Under the Curve to measure performance gains as a direct function of cumulative expenditure, providing a standardized financial efficiency metric for agentic search.
权威 Benchmark 与实测跑分对比
Evaluated on 10 mathematical and systems optimization tasks and the ALE-Bench-Lite programming benchmark:
- Dominates BA-AUC Across 9 of 10 Tasks: Significantly outperforms OpenEvolve, ShinkaEvolve, AdaEvolve, and EvoX in budget-normalized convergence curves.
- $0.55 Solution Beats $50 Multi-Agent Baselines on Circle Packing: Reaches new state-of-the-art performance on circle packing using GLM-5.3 and its Flash variant for a mere $0.55, and using GPT-5.6 Terra and Luna for $1.68—matching or outperforming heavy multi-agent baselines (CORAL, SwarmResearch) that average ~$50 per run (a 97%–99% cost reduction).
- Superior Solution Quality on ALE-Bench-Lite: Achieved higher average optimization gains across all 10 real-world algorithmic challenges in ALE-Bench-Lite.
开发者实战落地与开箱指南
FrugalEvo is open-sourced on GitHub (chchenhui/frugalevo). Software engineers and researchers optimizing GPU kernels, algorithmic trading pipelines, and database query planners can immediately integrate FrugalEvo's dual-LLM orchestrator. By replacing monolithic frontier models with cache-aligned asymmetric model hierarchies, organizations can run continuous evolutionary code optimization at fractional computational expense.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.