News · Leaderboard

Latest news, beside the leaderboard

Read what just changed, and see who is ahead. Those are the two doors on this page.

Plans
R
Reflection AI@aicoder·1h ago
🔥 Trending

Reflection unveils Beam, its first open-weight model: 501B MoE (23B active) for coding and agents, 80.9% SWE-bench Verified, Apache 2.0 weights due this month

On Oct 5 Reflection AI introduced Beam, a sparse MoE with 501B total and 23B active parameters for coding, reasoning and agentic work. It was pretrained on 23.8T tokens, then put through a four-week high-compute RL run on 10.5K GB300 GPUs (100M+ rollouts), with 1M-token context. Reflection says it is competitive with GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks, matching GLM-5.2 on reasoning with 3–4x less inference compute. It is in final red-teaming with early-access signup only; weights, tech report and model card ship this month under Apache 2.0.

Reflection unveils Beam, its first open-weight model: 501B MoE (23B active) for coding and agents, 80.9% SWE-bench Verified, Apache 2.0 weights due this month
⚡ Key Takeaways
  • •Scale: sparse MoE, 501B total / 23B active, 52 layers, interleaved local+global attention, 1M-token context
  • •Self-reported scores: SWE-bench Verified 80.9, Terminal Bench v2.1 80.1, MCP Atlas 78.7, GPQA Diamond 90.5, AIME 2026 97.8, HLE (no tools) 36.2
  • •RL run: 10.5K NVIDIA GB300 GPUs for four weeks, 100M+ rollouts up to 256K context, ~1.3B sandboxes, ~1M environments
  • •Efficiency: GLM-5.2-level reasoning scores with 3–4x less inference compute; reasoning-effort parameter trades length for quality
  • •Availability: in final red-teaming with early-access signup; Apache 2.0 weights, tech report and fine-tuning/eval stack due this month
Read details→
H
Hugging Face Daily Papers@HuggingFace·1h ago
🔥 Trending

4DCodeBench: Stanford and MIT Benchmark Coding Agents on Inverse Graphics of Dynamic Physical Scenes

Inverse graphics—interpreting raw visual inputs into executable generative code representing scene geometry and dynamics—constitutes a foundational cognitive frontier for physical world models and embodied agents. Existing code generation benchmarks remain confined to static 2D/3D geometries or conversational software development, failing to assess an agent's grasp of physical causality. Researchers from Stanford University, MIT, and Johns Hopkins introduce 4DCodeBench (arXiv:2610.03715), the first benchmark evaluating agents on 4D inverse graphics through executable code synthesis. Agents must parse dynamic videos and implement compact abstractions such as physics simulators to reproduce observed continuum behaviors, including fluid dynamics, non-rigid deformation, and structural fractures. Evaluating frontier models reveals a stark capability gap: models with strong static 3D spatial reconstruction fail to generate executable simulations for complex time-evolving dynamics. 4DCodeBench establishes an open-source testbed and dataset to measure how coding agents interpret the physical dynamics of the physical universe.

4DCodeBench: Stanford and MIT Benchmark Coding Agents on Inverse Graphics of Dynamic Physical Scenes
⚡ Key Takeaways
  • •Stanford, MIT, and JHU introduce 4DCodeBench, the first inverse graphics benchmark synthesizing executable dynamic physics code
  • •Encompasses fluid dynamics, non-rigid deformations, and fractures with rigorous spatiotemporal and momentum-conservation evaluations
  • •Exposes that frontier models excelling at static 3D fail on dynamic 4D physics, with physically valid code rates dropping below 15%
Read details→
H
Hugging Face Daily Papers@HuggingFace·1h ago
🔥 Trending

RealCompanion: Auditing Human Understanding Over 120-Day Real-World Conversations Exposes Agent Memory Benchmarks

In long-term companion AI and memory-augmented agent research, standard evaluations rely on synthetic personas and generated queries that artificially predetermine relevance and leak retrieval cues. Researchers introduce RealCompanion (arXiv:2610.01780), releasing a landmark dataset derived from ten genuine human relationships with an AI companion spanning up to 120 days and 27,218 real messages, meticulously annotated with ground-truth reasoning traces. The empirical analysis yields three foundational findings: First, historical memory is rarely required and temporally distant—95.9% of user interactions are satisfied by immediate context, with only 2.2% genuinely necessitating historical memory retrieval; pooled evaluation scores create an illusion of competence where 96% of evidence gains come from queries needing no recall. Second, contemporary memory detectors fail entirely on unprompted human text; synthetic benchmarks leak intent cues, and merely tagging inputs as 'memories' inflates agent retrieval by 10 to 14 points. Third, distinct memory architectures reconstruct human personas with identical F1 accuracy across a staggering 31-fold disparity in computational cost.

RealCompanion: Auditing Human Understanding Over 120-Day Real-World Conversations Exposes Agent Memory Benchmarks
⚡ Key Takeaways
  • •Releases RealCompanion over 120 days and 27,218 real messages, puncturing synthetic benchmark illusions in agent memory
  • •Reveals extreme long-tail memory demands: only 2.2% of authentic messages require recall, while 95.9% rely strictly on recency
  • •Demonstrates that contemporary memory detectors fail on natural speech, and identical F1 persona accuracy spans a 31x cost gap
Read details→
H
Hugging Face Daily Papers@HuggingFace·1h ago
🔥 Trending

Latent-MOPD: Carnegie Mellon Pioneers Representation-Level Multi-Teacher On-Policy Distillation for LLMs

On-Policy Distillation (OPD) trains student LLMs on trajectories sampled from their own generation policies, providing robust alignment for mathematical and code reasoning. However, existing multi-teacher OPD frameworks restrict knowledge transfer purely to output token probability distributions, discarding rich semantic representations embedded in teachers' intermediate hidden states. Researchers from Carnegie Mellon University and Georgia Tech introduce Latent-MOPD (arXiv:2610.02381), the first representation-level multi-teacher OPD architecture for LLMs. Latent-MOPD integrates diverse specialists by coupling their output distributions with the internal representations used to generate them, without modifying teacher checkpoints. To coordinate disparate representations across specialists, Latent-MOPD selectively aligns late-layer targets, resolves channel mismatches through shared linear projections, and enforces domain-grouped gradient updates while gradually annealing hidden-state loss into token prediction. Across nine math, coding, and logic benchmarks, Latent-MOPD surpasses token-only and uniform-averaging baselines; matching the parameter budget of individual teachers, the distilled student outperforms the best domain teacher on a majority of benchmarks.

Latent-MOPD: Carnegie Mellon Pioneers Representation-Level Multi-Teacher On-Policy Distillation for LLMs
⚡ Key Takeaways
  • •Carnegie Mellon introduces Latent-MOPD, the first representation-level multi-teacher on-policy distillation framework for LLMs
  • •Pairs shared projections with domain-pure grouped updates and annealing schedules to eliminate latent representation collapse
  • •Dominates all 9 math, code, and logic benchmarks, with an equal-sized student outperforming domain-best teachers on most tasks
Read details→
ADSponsored
C
Cantina Security@AICoder·3h ago
🔥 Trending

Cantina open-sources apex-flash-1, a GLM-5.3-Flash GRPO post-train for security research: 66.7% pass@1 on 60 held-out vuln tasks at ~1/31 the cost of Opus 5 High

Cantina Security and Yeta Labs released apex-flash-1 (MIT), their first open-weights security research model: a GRPO post-train of GLM-5.3-Flash meant to work as a focused worker under a larger orchestrating agent, reading code, using tools, building exploits and verifying them against a running target. On 60 tasks from 20 held-out vulnerability cases it scores 66.7% pass@1 (40/60) vs 60.0% for base GLM-5.3-Flash and 71.7% for Claude Opus 5 High, at an estimated $2.38 per run vs $74.68 for Opus 5 High. An experimental abliterated variant ships alongside.

⚡ Key Takeaways
  • •Score: 66.7% pass@1 (40/60) on held-out tasks vs 60.0% for base GLM-5.3-Flash and 71.7% for Claude Opus 5 High
  • •Cost: ~$2.38 for the 60-task run vs $4.56 (GLM-5.3-Flash) and $74.68 (Opus 5 High) at provider pricing
  • •Training: 150 tasks from 50 real vulnerability cases in three views; rank-256 LoRA on all experts and routers plus full updates to 16 activation-selected experts, GRPO
  • •Data mix: 72% authorization/identity/scope-binding bugs, 18% numerical precision, remainder signature replay, payment rules and SSRF
  • •License: MIT open weights on Hugging Face plus an experimental abliterated variant; Codex harness recommended
Read details→
g
ggml-org@AICoder·3h ago
🚀 Release

llama.cpp v0.6.0 ships GLM-5.3-Flash (320B) and Qwen4Exp MTP speculative decoding, a new llama_batch_ext API, and up to ~3x faster few-row mat-mul on Apple GPUs

ggml-org released llama.cpp v0.6.0 on Oct 5 (v0.5.0 was Sep 23). It brings GLM-5.3-Flash (GLM5-Next, a 320B text+vision hybrid MoE) into a stable tag, adds MTP speculative decoding for Qwen4Exp (~1.5x decode on DGX Spark), introduces the llama_batch_ext / llama_process() API for mixed token and embedding batches, adds few-row MMA mat-mul kernels on Metal (up to ~3x for speculative and batched decoding), and bumps ggml to v0.26.0. Session file formats are bumped too.

⚡ Key Takeaways
  • •New models: GLM-5.3-Flash (GLM5-Next, 320B KDA/DSA hybrid text+vision MoE), Clef decision model, Ling 3.0 VL, LFM2.5-Encoder
  • •Speed: Qwen4Exp MTP speculative decoding ~1.5x decode on DGX Spark; Metal few-row MMA mat-mul up to ~3x
  • •API: llama_batch_ext + llama_process() for mixed token/embedding batches; server and examples migrated
  • •Compat: LLAMA_SESSION_VERSION 11 and LLAMA_STATE_SEQ_VERSION 4, so old session/state files must be regenerated
  • •Server: /v1/models reports modalities, /v1/embeddings takes typed vision/audio/video input, new /v1/systemone endpoint
Read details→
R
Reka Labs@RekaAILabs·5h ago
🔥 Trending

Reka releases Rho-1 research preview: a 19B omni-reasoning model trained from scratch that understands and generates text, images and video and emits robot actions in one network

On Oct 5 Reka Labs released a research preview of Rho-1, a 19B omni-reasoning model trained from scratch that unifies text, image, video and robot-action tokens in one network and one KV cache, with no tool calls or second model. Each transformer block has an understanding stream and a generation stream with shared attention, trained jointly with next-token prediction and flow matching. Reka reports base video generation at 0.79x real time (median) and a distilled 8-step variant (vs 99) returning a 5.3s clip in about a second. The checkpoint was trained on 320 H100s for three months. Research preview only; no public weights.

Reka releases Rho-1 research preview: a 19B omni-reasoning model trained from scratch that understands and generates text, images and video and emits robot actions in one network
⚡ Key Takeaways
  • •Scale: 19B parameters trained from scratch on just 320 H100s for three months
  • •Speed: base video at 0.79x real time (median), ~6s to a watchable stream; distilled 99→8 denoising steps returns a 5.3s clip in ~1s
  • •Architecture: dual expert streams (understanding/generation) per block with shared attention and one KV cache; next-token + flow-matching objectives
  • •Limits: native video capped at 672x384; long-horizon drift, temporal grounding and edit stability are admitted weak points
  • •Access: research preview only, no open weights or public API; partnerships via [email protected]
Read details→
H
Hugging Face Daily Papers@HuggingFace·5h ago
🔥 Trending

WEFT: Whole-System Evolution for Tool-Use Post-Training Scales Agentic Interaction with MegaMCP State Isolation

Recent attempts to scale tool-use post-training for autonomous agents have focused primarily on synthesizing standalone execution environments. However, environments constitute merely one component of a holistic agentic interaction system comprising environments, tasks, harnesses, and evaluators; scaling environments in isolation yields inconsistent training signals. Researchers from Fudan University introduce WEFT (arXiv:2609.36887), a whole-system evolution framework for tool-use post-training. WEFT coordinates environment breadth, task complexity, and interaction diversity, applying execution-driven self-evolution to attribute failures and iteratively revise flawed components. To ensure optimization stability and concurrency reliability, WEFT introduces prefix-preserving sampling, atomic-turn credit assignment, and MegaMCP—a Model Context Protocol infrastructure providing isolated, recoverable state management across concurrent rollouts. WEFT-8B and WEFT-14B outperform matched environment-scaling baselines across BFCL V4, tau^2-Bench, and Claw-Eval, with WEFT-14B beating Agent-World-14B by 6.41, 2.23, and 12.27 percentage points, while WEFT-35B-A3B excels on long-horizon benchmarks like Toolathlon-Verified.

WEFT: Whole-System Evolution for Tool-Use Post-Training Scales Agentic Interaction with MegaMCP State Isolation
⚡ Key Takeaways
  • •Pioneers WEFT, co-evolving environments, tasks, harnesses, and evaluators to overcome the limits of isolated environment synthesis
  • •Integrates prefix-preserving sampling, atomic-turn credit assignment, and MegaMCP isolation for resilient tool-use post-training
  • •WEFT-14B surges past Agent-World-14B by 6.41 pp on BFCL V4, 2.23 pp on tau^2-Bench, and 12.27 pp on Claw-Eval
Read details→
ADSponsored
H
Hugging Face Daily Papers@HuggingFace·5h ago
🔥 Trending

ProWAM: Meta FAIR and SJTU Advance World Action Modeling with Progressive Sparse Sub-Goal Visual Planning

World Action Models (WAMs) represent a foundational robotic paradigm by predicting future visual dynamics and Cartesian actions directly from initial observations. However, existing WAMs struggle with long-horizon execution: synthesizing dense video rollouts incurs prohibitive computational overhead and temporal drift, while predicting single terminal frames deprives the policy of progressive visual guidance. Meta FAIR and Shanghai Jiao Tong University introduce ProWAM (arXiv:2610.02508), an efficient architecture that jointly predicts actions and an ordered sequence of sparse visual sub-goals. Sub-goal prediction scales naturally by learning from large-scale action-free video datasets, decoupling complex visual foresight from low-level policy control. ProWAM executes a single video-backbone forward pass to cache sparse sub-goal features, enabling rapid closed-loop replanning through lightweight action denoising without full video re-generation. Evaluated across extensive robotic benchmarks, ProWAM sets SOTA records on LIBERO-Plus (85.8%, +35.9% relative gain) and randomized RoboTwin (75.7%). In real-world zero-shot physical manipulator trials, ProWAM achieves a 70.0% success rate in unseen environments, outpacing the strongest baseline by +15.0 percentage points (+27.3% relative gain).

ProWAM: Meta FAIR and SJTU Advance World Action Modeling with Progressive Sparse Sub-Goal Visual Planning
⚡ Key Takeaways
  • •Introduces ProWAM, replacing dense video rollouts with progressive sparse visual sub-goals for robust robotic foresight
  • •Executes a single video forward pass to cache sub-goal features, enabling rapid closed-loop action replanning
  • •Reaches 85.8% success on LIBERO-Plus (+35.9% relative gain) and lifts zero-shot physical robot success from 55% to 70%
Read details→
H
Hugging Face Daily Papers@HuggingFace·5h ago
🔥 Trending

FailBank: Self-Evolving Vision-Language-Action Models from Runtime Feedback with Guarded Policy Adaptation

Vision-Language-Action (VLA) architectures achieve broad generalization across robotics manipulation, but deploying them in unstructured physical environments requires balancing task progression with collision avoidance. Traditional safety approaches deploy external runtime shields (such as Control Barrier Functions, CBF) to override risky actions. However, external shields treat symptoms without altering the underlying policy weights, resulting in persistent policy-shield mismatches that freeze progress or induce mechanical oscillation. Researchers from the University of Notre Dame introduce FailBank (arXiv:2609.39820), a four-stage self-evolution framework converting runtime safety interventions into persistent policy improvements. During rollouts, a fixed CBF safety module acts as an observe-only teacher, generating counterfactual corrections while allowing the primary policy to maintain control. An outcome-aware admission filter selects viable corrections as positive targets while anchoring successful uncorrected trajectories for guarded LoRA fine-tuning. Evaluated on the VLA-Arena benchmark across multiple backbones, FailBank boosts task success rates by 6.9 to 8.5 percentage points while reducing collision costs by 23.8% to 35.6%. Crucially, compared to static runtime shielding, FailBank surges success rates by 9.5 to 25.4 percentage points, proving runtime feedback can provide persistent supervisory guidance.

FailBank: Self-Evolving Vision-Language-Action Models from Runtime Feedback with Guarded Policy Adaptation
⚡ Key Takeaways
  • •Open-sources FailBank, resolving persistent policy-shield mismatches where external runtime safety shields induce robot freezes
  • •Employs observe-only CBF modules for counterfactual corrections and guarded LoRA with quiet anchors to avoid catastrophic forgetting
  • •Boosts VLA-Arena success by 6.9-8.5 pp while cutting collision costs by 23.8%-35.6%, outperforming static shielding by 9.5-25.4 pp
Read details→
O
OpenAI@OpenAI·6h ago
🔥 Trending

OpenAI details textGrain text watermarking for the EU AI Act: invisible watermarks coming to ChatGPT and Codex output in the EU, opt-in for API customers worldwide

On Oct 5 OpenAI described its response to the EU AI Act's machine-readable text requirement: API customers worldwide can now opt in to text watermarking for select models (off by default), and in the coming weeks eligible ChatGPT and Codex text output in the EU will carry an invisible watermark across all plans. The technique, textGrain, embeds a statistical signal in word choices; OpenAI says it matched or beat SynthID for text and plans to open-source it. The detector is limited to approved researchers. At 1% false-positive rate, detection is ~80% on 200-token and ~95% on 400-token passages, falling to 17% after 25% synonym replacement; Astra benchmarks show no meaningful quality change.

OpenAI details textGrain text watermarking for the EU AI Act: invisible watermarks coming to ChatGPT and Codex output in the EU, opt-in for API customers worldwide
⚡ Key Takeaways
  • •Scope: invisible watermarks on eligible ChatGPT and Codex text in the EU in coming weeks (all plans); API opt-in worldwide, off by default
  • •Detection at 1% FPR: ~80% on 200-token passages, ~95% on 400-token; much lower for math-style text
  • •Robustness: 10% synonym replacement drops detection ~92%→66%; 25% drops it to 17%
  • •Quality: Astra (max) DeepSWE v1.1 72.80% vs 71.68%, Terminal-Bench 4.0 53.90% vs 56.06%, GPQA Diamond 94.44% vs 93.94% (unwatermarked vs watermarked)
  • •Openness: textGrain technical report published, open-source release planned; detector limited to approved researchers
Read details→
G
Google Bug Hunters (google/bughunters)@GoogleVRP·7h ago
🔥 Trending

Google pauses product-vulnerability payouts in its OSS VRP after a flood of automated AI submissions; $500–$7,500 band zeroed, update promised in Q1 2027

Google's Open Source Software Vulnerability Reward Program stopped accepting product-vulnerability reports on October 1, 2026, citing a significant rise in automated submissions that are mostly invalid. Commit f8bf23ad8 in google/bughunters (Sept 30) replaced the product-vulnerability reward bands ($500–$7,500 for flagship projects, $101–$3,133.7 for important ones) with dashes, while supply-chain compromise rewards stay at up to $31,337. Reports filed before Oct 1 are unaffected, some Google Cloud repos can route via Cloud VRP, and Google promises an update in Q1 2027.

⚡ Key Takeaways
  • •Effective 2026-10-01: OSS VRP no longer accepts product vulnerabilities (rules)
  • •Zeroed bands: $500–$7,500 (flagship) and $101–$3,133.7 (important) replaced with dashes in commit f8bf23ad8
  • •Unchanged: supply-chain compromises still pay $3,133.7–$31,337 (OT0) and $1,337–$13,337 (OT1)
  • •Pre-Oct 1 reports unaffected; some Google Cloud repos can go via Cloud VRP; Patch Rewards remains open
  • •Google commits to an update on the product-vulnerability track in Q1 2027
Read details→
T
TesterArmy@tester-army·9h ago
🛠️ Tooling

TesterArmy's open-source e2e tops GitHub daily trending: natural-language agent steps for web and mobile tests, replayed with zero model calls

e2e (tester-army/e2e), an Apache-2.0 AI end-to-end testing framework, gained 1,430 stars today to top GitHub Trending (3,828 total; 60,933 npm downloads last week). Tests mix natural-language agent.act/agent.assert steps with locator assertions; verified agent steps are recorded and replayed with no model calls until the app changes. v0.17.0 adds Copilot Responses-API models, OpenCode Console sign-in, and a decision-model executor for bounded actions.

TesterArmy's open-source e2e tops GitHub daily trending: natural-language agent steps for web and mobile tests, replayed with zero model calls
⚡ Key Takeaways
  • •#1 on GitHub daily trending: +1,430 stars today, 3,828 total (Apache-2.0)
  • •npm package e2e: 60,933 weekly downloads (Sep 28–Oct 4); latest 0.17.0 released Oct 4
  • •Record & replay: verified agent steps rerun with zero model calls until the app changes
  • •v0.16.0 cut install footprint from 117 to 29 packages (~36MB to ~31MB)
  • •7 packages: Playwright (Chromium/Firefox/WebKit), iOS/Android, PR reporter, hosted browsers/simulators, decision-model executor
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

CacheBack: Receiver-Conditioned Latent Communication Slashes Multi-Agent KV Cache by Up to 94% with 3.2x Lower Latency

Multi-agent systems solving long-context distributed tasks face a fundamental trade-off: textual communication is compact but introduces autoregressive decoding latency and token-level information loss, whereas latent communication transferring raw KV caches avoids token generation but causes explosive linear memory scaling across agents and context tokens. Researchers from Columbia University introduce CacheBack (arXiv:2609.32046), a robust, training-free realization of receiver-conditioned communication. The receiving agent transmits a concise description of its information requirements, which acts as a dynamic mask on the sender's attention weights to filter and compress the KV cache. Benchmarked on FanOutQA with Qwen 3, CacheBack prunes 75% to 94% of redundant cache states, boosts task accuracy by 14.7 percentage points, and slashes median completion latency by 3.2x compared to text messages across dense, Mamba-hybrid, and sliding-window architectures.

CacheBack: Receiver-Conditioned Latent Communication Slashes Multi-Agent KV Cache by Up to 94% with 3.2x Lower Latency
⚡ Key Takeaways
  • •Pioneers receiver-conditioned latent communication in CacheBack, overcoming linear KV cache scaling in multi-agent systems
  • •Training-free architecture filters 75% to 94% of redundant cache states using attention weights conditioned on receiver needs
  • •Lifts FanOutQA multi-hop accuracy by 14.7 percentage points and slashes median task latency by 3.2x over text messaging
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

VeriHarness: Google and Cambridge Rethink Agentic Verification for Long-Horizon Tasks with Over $100K in Open Datasets

As LLM agents tackle increasingly complex, long-horizon software engineering and workspace tasks, test-time output verification without ground-truth solutions or rubrics remains a critical bottleneck. Researchers from Google Cloud AI Research and the University of Cambridge uncover that classical consensus-based heuristics fail in deep multi-step workflows: disagreement frequently reveals correct alternative implementations, whereas consensus often masks shared blind spots. Motivated by this, they introduce VeriHarness (arXiv:2610.00972), transforming standard foundation models into agentic verifiers equipped with isolated workspaces, environmental execution tools, and modular verification skills. A disagreement resolver cross-checks competing claims against live environmental evidence, while a consensus challenger scrutinizes shared assertions for overlooked edge constraints. Guided by evidence-backed revisions, VeriHarness delivers substantial gains of +6.2 points on Gemini 3.5 Flash and +6.4 points on Claude Opus 4.8 across five workspace benchmarks, backed by an open-sourced pool of 26,000 rollouts produced at an evaluation cost exceeding $100,000.

VeriHarness: Google and Cambridge Rethink Agentic Verification for Long-Horizon Tasks with Over $100K in Open Datasets
⚡ Key Takeaways
  • •Discovers that consensus heuristics fail in long-horizon agent tasks: disagreement reveals solutions while consensus masks errors
  • •Introduces VeriHarness with sandbox-backed Disagreement Resolvers and Consensus Challengers, gaining +6.2 pts on Gemini and +6.4 pts on Claude
  • •Releases a landmark dataset of 26,000 long-horizon rollouts produced at an evaluation cost exceeding $100,000 under Google Research
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

FrugalEvo: Cost-Aware LLM-Guided Program Evolution Slashes Multi-Agent Optimization Costs by 97% with Prefix-Cache Reuse

LLM-guided evolutionary algorithms (such as AlphaEvolve) have achieved breakthrough results in computational optimization, including circle packing and systems tuning. However, existing methods optimize strictly over iteration counts rather than fiscal expenditure, burning dozens of dollars per task. Researchers from NUS, Stanford, and HKUST present FrugalEvo (arXiv:2610.03675), a cost-aware program evolution framework maximizing gain per dollar. FrugalEvo implements an asymmetric dual-LLM division of labor: a powerful, higher-cost LLM explores high-level solution concepts, while a cheap, high-speed LLM generates and iteratively refines concrete implementations. Additionally, FrugalEvo engineers prompt harnesses to maximize prefix KV-cache reuse across evolutionary generations. Across 10 mathematical optimization problems and ALE-Bench-Lite suites, FrugalEvo matches or outperforms OpenEvolve, EvoX, and SwarmResearch. On circle packing, FrugalEvo achieves SOTA performance using GLM-5.3/Flash for only $0.55 and GPT-5.6 Terra/Luna for $1.68—slashing costs by up to 99% compared to conventional $50 multi-agent baselines.

FrugalEvo: Cost-Aware LLM-Guided Program Evolution Slashes Multi-Agent Optimization Costs by 97% with Prefix-Cache Reuse
⚡ Key Takeaways
  • •Pioneers FrugalEvo, the first cost-aware LLM program evolution framework maximizing optimization gain per dollar spent
  • •Pairs an asymmetric dual-LLM hierarchy with prefix KV-cache reuse, establishing the new Budget-Aware AUC (BA-AUC) benchmark
  • •Achieves SOTA on circle packing for $0.55-$1.68, outperforming $50 multi-agent baselines like CORAL and SwarmResearch
Read details→
A
Anthropic / AWS Machine Learning Blog@AnthropicAI·11h ago
🔥 Trending

Anthropic brings in-country Claude inference to India: Opus 5, Sonnet 5 and Haiku 4.5 via Amazon Bedrock with data kept in India

On 2026-10-05 Anthropic announced Claude Opus 5, Sonnet 5 and Haiku 4.5 are live with in-country inference in India via an Amazon Bedrock India geographic cross-Region profile that routes only between Mumbai (ap-south-1) and Hyderabad (ap-south-2), delivering its August data-residency commitment for regulated Indian customers.

⚡ Key Takeaways
  • •Models: Claude Opus 5, Sonnet 5, Haiku 4.5 via the Bedrock India geographic inference profile
  • •Routing only between ap-south-1 (Mumbai) and ap-south-2 (Hyderabad); no data stored in destination Region; Bedrock zero data retention by default
  • •Model IDs use the in. prefix (e.g. in.anthropic.claude-sonnet-5); Messages API, InvokeModel, Converse, Guardrails and prompt routing supported
  • •Same AWS batch (Sep 29): in-region Opus 5/Sonnet 5 in Seoul and Sonnet 5 in Singapore
  • •Adoption per Economic Times: TCS rolling Claude to 50,000 staff; NPCI building its AiNxt agent platform on Claude
Read details→
G
Gökdeniz GülmezGoekdeniz-Guelmez·11h ago
🛠️ Tooling

MLX-LM-LoRA v5.8.6: Apple Silicon fine-tuning library jumps from 3.1.3, adds DSLA preference alignment and critic-free KLPO RL

Gökdeniz Gülmez shipped MLX-LM-LoRA v5.8.6 on Oct 5 (GitHub Release and PyPI), the first release after 3.1.3. It adds DSLA with DPO/ORPO/CPO objectives and critic-free KLPO with token/sequence routes and several KL estimators, improves GRPO stability and microbatching for Online DPO, XPO, RLHF REINFORCE and PPO, and adds fast VJP for gated-delta layers. Minimums rise to mlx>=0.32.3, mlx_lm>=0.32.0 and Python>=3.11; license standardized on Apache 2.0.

⚡ Key Takeaways
  • •Jump from v3.1.3 (Sep 16) straight to v5.8.6, merging nine PRs (#64, #94, #96–#102).
  • •Two new modes: --train-mode dsla (--dsla-loss dpo|orpo|cpo, default --latent-weight 0.1) and --train-mode klpo (--klpo-route token|sequence, estimators mc/binary/topk/full).
  • •KLPO skips GRPO group normalization, PPO ratio clipping and the reference model, reusing GRPO reward callbacks and generation/scoring.
  • •Raised minimums: mlx>=0.32.3, mlx_lm>=0.32.0, Python >=3.11.
  • •QAT now covers DSLA; new GitHub Pages site with a Python API reference.
Read details→
v
vLLMvllm-project·14h ago
🛠️ Tooling

vLLM v0.31.0: 717 commits, `vllm preload` fast restarts, DeepSeek-V4.1-Flash FlashMLA default, security gates and breaking changes

vLLM v0.31.0 (717 commits, 307 contributors, 96 new) adds the `vllm preload` weight-cache daemon for fast engine restarts and experimental CRIU engine snapshots, makes FlashMLA mega attention with NVFP4 compressed KV the SM100 default for DeepSeek-V4.1-Flash, brings draft-model speculative decoding to Model Runner V2, gates per-request multimodal kwargs behind a flag, and fixes prefix-cache key collisions between LoRA names and cache_salt. Several breaking changes ship too.

⚡ Key Takeaways
  • •Scale: 717 commits from 307 contributors (96 new).
  • •Fast restart: vllm preload keeps post-quantized weights resident in GPU memory across restarts (DP, MTP drafts, /health); experimental CRIU snapshots restore an initialized TP1 engine.
  • •Security: per-request multimodal kwargs rejected unless --trust-request-mm-kwargs; LoRA path now part of the block hash to stop prefix-cache collisions.
  • •Breaking: tokenizer_mode="slow" removed; online quantization="fp8" replaced by fp8_per_tensor; --enforce-eager also disables JIT warmup.
  • •Model paths: GLM-5.3-Flash metadata ops 1.6–4.8x faster and 3 GiB indexer workspace saved; fused KimiViT QK RoPE up to 29x; ~25 s faster Qwen3.8-Flash-Next weight loading on DGX Spark.
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

Recursive Self-Rewrite (RSR) Scales Complex Terminal Trajectories: Distilling Specialized Harnesses into General SFT Capabilities

Collecting successful trajectories for complex terminal tasks is essential for post-training coding agents. However, solving difficult tasks often relies on specialized harnesses that incorporate domain assumptions unavailable during real-world deployment; naive SFT on these traces causes verifier leakage and harness over-fitting. Researchers from the University of Maryland and Tencent AI Lab introduce Recursive Self-Rewrite (RSR, arXiv:2610.02826). Using a single foundation model (Qwen-3.8-27B), RSR pairs a planner (extracting operational runbooks), a critic (eliminating leakage), and an executor (replaying runbooks in fresh isolated sandboxes). RSR expands 2,001 specialized source rollouts into 11,094 clean general-harness trajectories, boosting Terminal-Bench 2 Pass@3 from 57.0% to 74.2% and surging Terminal-Bench Hard success from 39.0% to 63.0%.

Recursive Self-Rewrite (RSR) Scales Complex Terminal Trajectories: Distilling Specialized Harnesses into General SFT Capabilities
⚡ Key Takeaways
  • •Pioneers Recursive Self-Rewrite (RSR) to distill specialized harness experiences into clean general-harness SFT trajectories
  • •Expands 2,001 noisy exploration rollouts into 11,094 fully reproducible trajectories in pristine sandboxes
  • •Drives Terminal-Bench 2 Pass@3 from 57.0% to 74.2% and lifts Terminal-Bench Hard success from 39.0% to 63.0%
Read details→