⚡Trending:
I
InternLM / Shanghai AI Lab@InternLM·25m ago
🚀 Release

LMDeploy 0.18.0: Qwen3.5 dflash, SM90 quantized GEMMs, Ascend GLM-5.2; TurboMind W4A16 + piecewise CUDA Graph

Shanghai AI Lab’s LMDeploy shipped v0.18.0 on 2026-09-28: input logprobs, expanded SM90 quantized GEMMs with unified TurboMind linear paths, Ascend GLM-5.2, Qwen3.5 dflash, and MoE shared-expert/FFN sharding. PyTorch path prototypes TurboMind W4A16 (AWQ), piecewise CUDA Graph prefill, XTuner TileLang sparse MLA, checkpoint-engine weight updates, and request-only KV cache metrics.

LMDeploy 0.18.0: Qwen3.5 dflash, SM90 quantized GEMMs, Ascend GLM-5.2; TurboMind W4A16 + piecewise CUDA Graph
⚡ Key Takeaways
  • •Shipped v0.18.0 — Qwen3.5 dflash, Ascend GLM-5.2, SM90 quantized GEMMs
  • •TurboMind W4A16 AWQ prototype + piecewise CUDA Graph prefill + TileLang sparse MLA
  • •checkpoint-engine weight updates + request-only KV cache metrics + input logprobs
  • •Upgrade: pip install -U "lmdeploy>=0.18.0"
Read details→
M
Model Context Protocol@modelcontextprotocol·26m ago
🛠️ Tooling

MCP TypeScript SDK 2.2.0: M2M OAuth expectedIssuer, auto-paging list*, CJS jose types + Client.listen hang fixes

MCP’s TypeScript monorepo shipped v2.2.0 on 2026-09-28: client/server/core/server-legacy/codemod aligned at 2.2.0. M2M OAuth providers should pass `expectedIssuer`; `fetchToken()` throws `AuthorizationServerMismatchError` on issuer mismatch. Cursor-less listTools/listPrompts/listResources/listResourceTemplates now follow `nextCursor` (capped by `listMaxPages`). Fixes the 2.1.0 CJS jose types regression and Client.listen() unhandled rejection / send hang.

MCP TypeScript SDK 2.2.0: M2M OAuth expectedIssuer, auto-paging list*, CJS jose types + Client.listen hang fixes
⚡ Key Takeaways
  • •Shipped v2.2.0 — client/server/core/server-legacy/codemod aligned; node/express/hono/fastify unchanged
  • •M2M OAuth: pass expectedIssuer; fetchToken throws AuthorizationServerMismatchError on mismatch
  • •list* without cursor auto-follows nextCursor (listMaxPages cap)
  • •Fixes CJS jose types regression and Client.listen() rejection/hang
  • •Install: npm i @modelcontextprotocol/[email protected]
Read details→
A
Anthropic@AnthropicAI·26m ago
🛠️ Tooling

Anthropic Python SDK 1.9.0: claude-sonnet-5-5, between_tools thinking, cache diagnostics GA, tools while reply streams

Anthropic shipped Python SDK v1.9.0 on 2026-09-28: adds claude-sonnet-5-5, between_tools thinking, GA cache diagnostics on Message/MessageCreateParams, and optional tool execution while the reply streams. Managed Agents event filters and workspace rate-limit include_inherited/source land too, plus unnamed upload and stream()/parse() diagnostics fixes.

Anthropic Python SDK 1.9.0: claude-sonnet-5-5, between_tools thinking, cache diagnostics GA, tools while reply streams
⚡ Key Takeaways
  • •Shipped v1.9.0 — 31 commits vs v1.8.0; claude-sonnet-5-5 + between_tools thinking
  • •Cache diagnostics GA on Message / MessageCreateParams; stream()/parse() accept diagnostics
  • •Optional tool execution while the assistant reply streams
  • •Upgrade: pip install -U "anthropic>=1.9.0" (Python 3.10+)
Read details→
L
LangChain4j@langchain4j·5h ago
🛠️ Tooling

LangChain4j 1.20.2: A2A client tenant propagation—auto-attach tenant on messages without exposing it as an LLM tool arg

LangChain4j 1.20.2 (2026-09-28) focuses on Agent-to-Agent (A2A) client fixes: `tenant()` on `A2AClientInstance`/`A2AClientBuilder`, `@A2AClientAgent(tenant=…)` auto-attaches tenant on every `MessageSendParams` (or parses it from `/.well-known/{tenant}/agent-card.json`), and filters tenant method args so they are not exposed as LLM tool parameters. AgenticScope/ResultWithAgenticScope toString guards against circular refs.

LangChain4j 1.20.2: A2A client tenant propagation—auto-attach tenant on messages without exposing it as an LLM tool arg
⚡ Key Takeaways
  • •Shipped 1.20.2 — A2A tenant auto-propagation via annotation/builder or Agent Card URL
  • •Tenant args filtered from LLM tool surface; attached on MessageSendParams
  • •AgenticScope toString circular-ref guard
  • •Docs: a2a-protocol.org + docs.langchain4j.info/tutorials/agents
Read details→
ADSponsored
U
Unsloth@unslothai·5h ago
🛠️ Tooling

Unsloth v0.1.900-beta: local Laya Decision models + Skills editor; ~4.5× LTX-2.3 clips and up to 6.3× faster VAE decode

Unsloth shipped v0.1.900-beta on 2026-09-28 (PyPI unsloth 2026.9.12): Desktop runs Laya Decision models locally via a Jev-compatible `/v1/systemone` endpoint, adds a Skills CRUD editor and Library viewers for PDF/Office, and improves Apple Silicon with batched serving, structured outputs, and TurboQuant KV. Media pipelines claim ~4.5× LTX-2.3 clips, 1.7–6.3× VAE decode, and up to ~1 minute faster MiniMax-H3 first render with 25–29 GiB lower peak VRAM.

Unsloth v0.1.900-beta: local Laya Decision models + Skills editor; ~4.5× LTX-2.3 clips and up to 6.3× faster VAE decode
⚡ Key Takeaways
  • •Shipped v0.1.900-beta — local Laya Decision API + Skills editor + Library viewers
  • •LTX-2.3 ~4.5×; VAE decode 1.7–6.3×; MiniMax-H3 first render up to ~1 min faster, −25–29 GiB peak VRAM
  • •Apple Silicon: batched serving, structured outputs, TurboQuant KV; ModelScope downloads
  • •Docs: docs.unsloth.ai
Read details→
O
OpenAI@OpenAI·5h ago
🛠️ Tooling

OpenAI Python SDK 3.20.0: Agents credential/session options, Cyber access programs on Responses, opt-in WebSocket text/tool snapshots

OpenAI shipped Python SDK v3.20.0 on 2026-09-28: Agents credentials gain metadata and update-without-replace auth, sessions add ultrafast tier plus misalignment_policy_violation; Responses can select Cyber access programs (standard/daybreak_blue/daybreak_red); and an opt-in ResponsesWebSocketAccumulator yields immutable incremental text/tool snapshots. Live/Realtime WebSocket query/queue/TLS retry fixes land in the same cut.

OpenAI Python SDK 3.20.0: Agents credential/session options, Cyber access programs on Responses, opt-in WebSocket text/tool snapshots
⚡ Key Takeaways
  • •Shipped v3.20.0 — Agents credential metadata/session ultrafast + Cyber access programs
  • •Opt-in ResponsesWebSocketAccumulator for incremental text/tool snapshots (190+ WS tests)
  • •Live/Realtime WebSocket query/queue/TLS retry hardening
  • •Upgrade: pip install -U openai
Read details→
F
Firecrawl@firecrawl·9h ago
🛠️ Tooling

Firecrawl SDKs ship agent_hints: Python 4.45.0 / JS 4.42.0 preserve deterministic next-step guidance for agents

Firecrawl published firecrawl-py 4.45.0 and @mendable/firecrawl-js 4.42.0 on 2026-09-28, shipping `agent_hints` preservation from [#4641](https://github.com/firecrawl/firecrawl/pull/4641): with `X-Firecrawl-Agent-Hints: true`, Search/Scrape/Parse/Map responses can include deterministic next-step suggestions. Also includes Hangar browser/interact migration and `exchange.onTermsRequired` fields.

Firecrawl SDKs ship agent_hints: Python 4.45.0 / JS 4.42.0 preserve deterministic next-step guidance for agents
⚡ Key Takeaways
  • •Shipped firecrawl-py 4.45.0 + JS 4.42.0 with agent_hints preservation
  • •Opt-in via X-Firecrawl-Agent-Hints: true (off by default)
  • •Deterministic next-step rules for Search/Scrape/Parse/Map
  • •MCP pin still on 4.40.0 — bump needed for full hint survival
Read details→
O
OpenCode@anomalyco·9h ago
🛠️ Tooling

OpenCode 1.18.33: Cloudflare AI Gateway honors timeouts, MCP browser launch failures surface, debug config redacts credentials

OpenCode v1.18.33 (2026-09-28) makes Cloudflare AI Gateway models honor provider response/stream timeouts, reports MCP browser launcher exits, redacts credentials in debug config output, and aligns Gemini thinking defaults/effort options with supported controls across model generations.

OpenCode 1.18.33: Cloudflare AI Gateway honors timeouts, MCP browser launch failures surface, debug config redacts credentials
⚡ Key Takeaways
  • •Shipped v1.18.33 — CF AI Gateway timeouts honored
  • •MCP browser launcher immediate-exit failures reported
  • •Debug config redacts credentials/sensitive headers
  • •Gemini thinking defaults/effort aligned across generations
Read details→
ADSponsored
D
Docker@docker·9h ago
🛠️ Tooling

Docker cagent 1.145.0: ACP request traces via W3C traceparent/tracestate, server spans on 13 handlers, lean TUI scrollback fix

Docker shipped cagent (docker-agent) v1.145.0 on 2026-09-28: ACP requests now propagate W3C `traceparent`/`tracestate` with server spans for each of the 13 implemented protocol handlers, plus new lint cops and a lean-TUI fix that preserves scrollback during partial streaming tool calls.

Docker cagent 1.145.0: ACP request traces via W3C traceparent/tracestate, server spans on 13 handlers, lean TUI scrollback fix
⚡ Key Takeaways
  • •Shipped v1.145.0 — ACP W3C traceparent/tracestate + 13 handler spans
  • •Lean TUI keeps scrollback during partial streaming tool calls
  • •New FieldsSeqLookup / SlicesConcat lint cops
  • •Docs: docker.github.io/docker-agent
Read details→
H
Hugging Face Daily Papers@HuggingFace·11h ago
🔥 Trending

Continuous Depth Batching (CDB): Unlocking Depth-Adaptive Inference for Looped Language Models with 99% Speedup Realization

While looped language models enable depth-adaptive inference by executing fewer layer repetitions on easy tokens and more on complex tokens, variable loop counts break conventional batching engines like vLLM. Researchers from TUM and Imperial College introduced Continuous Depth Batching (CDB). By dynamically reorganizing batches between loop iterations, orchestrating looped KV-caches, and asynchronously predicting token exits, CDB achieves up to 99% of the theoretical maximum inference speedup across Ouro 1.4B and Huginn 3.5B architectures.

Continuous Depth Batching (CDB): Unlocking Depth-Adaptive Inference for Looped Language Models with 99% Speedup Realization
⚡ Key Takeaways
  • •Resolving the Looped Batching Impasse: Depth-adaptive inference allows tokens to exit recurrent transformer layers early, but differing loop counts prevent uniform forward passes in engines like vLLM; CDB introduces inter-step batch reformation.
  • •Asynchronous Exit Prediction & Looped KV Cache: Integrates a dynamic scheduler between recurrent blocks that manages recurrent KV-cache states and asynchronously prepares future batches based on exit-likelihood estimates.
  • •99% Theoretical Limit Realized: Validated on Ouro 1.4B and Huginn 3.5B, CDB achieves up to 99% of the upper-bound theoretical speedup, clearing the path for production deployment of adaptive-depth architectures.
Read details→
H
Hugging Face Daily Papers@HuggingFace·12h ago
🔥 Trending

InternW0-Δ Open-Sourced: Unified World Action Model Pretrained on 20,000+ Hours of Open Robotics Data

Shanghai AI Lab and OpenDataLab open-sourced InternW0-Δ, a unified World Action Model (WAM) pretrained on over 20,000 hours of curated robotic demonstrations, UMI data, and egocentric videos. Built on a Mixture-of-Transformers (MoT) architecture that binds visual predictive dynamics with robot actions, it introduces 'Causal Imprint' to inject future representations directly into the action expert without requiring costly future-video rollouts during inference.

InternW0-Δ Open-Sourced: Unified World Action Model Pretrained on 20,000+ Hours of Open Robotics Data
⚡ Key Takeaways
  • •20,000+ Hours Open-Source Corpus: Unifies robot manipulation, UMI dexterous data, and egocentric human demonstrations into the largest publicly released robotic training corpus.
  • •Mixture-of-Transformers (MoT): Coordinates video dynamics experts and action generation experts under semantic guidance from a frozen VLM, with 4D spatial priors distilled during pretraining.
  • •Causal Imprint Without Future Rollout: Eliminates the critical latency bottleneck of traditional WAMs that require rendering future video frames at test time, providing predictive latent cues directly for real-time control.
  • •Complete Open Source: Training code, model weights, data processing pipelines, and datasets released at internrobotics.github.io/InternW0-Delta/.
Read details→
G
Google DeepMind@GoogleDeepMind·12h ago
📊 Benchmark

Google DeepMind & Kaggle Launch Game Arena: Dynamic Competitive Benchmarks for LLM Strategic Evaluation

Addressing the rapid saturation and benchmark overfitting of static LLM evaluations, Google DeepMind and Kaggle unveiled Game Arena. By orchestrating head-to-head live agent matchups with dynamic Elo rating evolution, Game Arena prevents performance saturation across three pilot environments: Chess (perfect information), Texas Hold'em Poker (imperfect information), and Werewolf (multiplayer social deduction and deception), systematically benchmarking long-horizon strategic reasoning.

Google DeepMind & Kaggle Launch Game Arena: Dynamic Competitive Benchmarks for LLM Strategic Evaluation
⚡ Key Takeaways
  • •Dynamic Head-to-Head Evaluation: Replaces static, easily saturated benchmarks with real-time model-versus-model matchups and dynamic Elo tracking that scales naturally as LLM reasoning capabilities evolve.
  • •Comprehensive Game Theory Spectrum: Spans perfect information (Chess), imperfect information with strategic betting (Poker), and multi-turn multi-agent social deduction and deception (Werewolf).
  • •Reproducible Sandboxed Infrastructure: Jointly open-sourced by DeepMind and Kaggle, providing standardized Python tournament APIs, replay verification pipelines, and sandboxed anti-cheat guardrails.
Read details→
G
Google@google·14h ago
🚀 Release

Google ADK Python 2.10.0: experimental skill lifecycle, MongoDB vector/hybrid search, eval duration/token metrics

Google shipped ADK Python v2.10.0 on 2026-09-25 (`google-adk` 2.10.0): experimental skill lifecycle (`ADK_ENABLE_SKILL_LIFECYCLE=1`) with ephemeral lifecycles and active skill limits; MongoDB toolset for vector/hybrid search; evaluation metrics for duration, tokens, and model-call counts; better OpenAI reasoning model parameter adaptation and reasoning-token reporting. BigQuery protected write mode and instruction templating tighten with breaking changes.

Google ADK Python 2.10.0: experimental skill lifecycle, MongoDB vector/hybrid search, eval duration/token metrics
⚡ Key Takeaways
  • •Shipped v2.10.0 — pip install google-adk==2.10.0; docs at adk.dev
  • •Experimental skill lifecycle via ADK_ENABLE_SKILL_LIFECYCLE=1
  • •MongoDB toolset for in-flow vector/hybrid search
  • •Eval metrics: duration, tokens, model-call counts
  • •Breaking: tighter BigQuery protected writes; OpenAIResponsesLlm uses effort not thinking_config; empty eval sets raise ValueError
Read details→
B
BerriAI@BerriAI·14h ago
🚀 Release

LiteLLM 1.103.0: MCP delegated OAuth admission, Fuse/Capability routers, prompt-cache cost prediction

BerriAI shipped LiteLLM v1.103.0 on 2026-09-28: delegated MCP OAuth now requires admission; Capability classifier + Fuse V2 routers land; the proxy can predict prompt-cache costs across deployments, bind JWT claims to agents, add `tpd_limit`, bulk user/team APIs, and `/nvidia_nim` passthrough. Also: Bedrock S3 managed file delete/list, Friendli price auto-sync, Responses↔Chat reasoning mapping, and cosign image verification docs.

LiteLLM 1.103.0: MCP delegated OAuth admission, Fuse/Capability routers, prompt-cache cost prediction
⚡ Key Takeaways
  • •Shipped v1.103.0 — pip install litellm==1.103.0; docs.litellm.ai
  • •MCP delegated OAuth requires admission; cosign-verify Docker images
  • •Capability + Fuse V2 routers; cross-deployment prompt-cache cost prediction
  • •JWT→agent binding, tpd_limit, bulk user/team management APIs
  • •Bedrock S3 file delete/list, Friendli price sync, Codex model picker sync from proxy
Read details→
O
OpenAI@openai·14h ago
🛠️ Tooling

OpenAI Codex 0.158.0: MCP pre-registered OAuth, bearer-secured exec-server, fullscreen TUI paste & transparent image gen

OpenAI shipped Codex rust-v0.158.0 on 2026-09-28 (`@openai/codex` 0.158.0): MCP servers with pre-registered OAuth client secrets (`codex mcp add --oauth-client-secret`); bearer-token auth for exec-server WebSockets; fullscreen TUI copy-on-select/right-click paste with Markdown-preserving copies; explicit transparent backgrounds for image gen/edit. Elevated-permission terminal approvals default on; Windows/Linux/macOS sandbox and approval-retry fixes land too.

OpenAI Codex 0.158.0: MCP pre-registered OAuth, bearer-secured exec-server, fullscreen TUI paste & transparent image gen
⚡ Key Takeaways
  • •Shipped rust-v0.158.0 / @openai/codex 0.158.0 — docs: CLI + MCP
  • •MCP pre-registered OAuth via codex mcp add --oauth-client-secret
  • •Bearer-secured exec-server WebSockets; elevated terminal approvals on by default
  • •Fullscreen TUI copy/paste keeps Markdown; Mermaid quoted labels/& supported
  • •Sandbox fixes across Windows 10, Linux nested writable roots, macOS path aliases
Read details→
B
BerriAI@BerriAI·15h ago
🚀 Release

LiteLLM v1.103.0 Released: Hosted vLLM Batches, MCP 2.x Architecture, and Admin UI Web Search Interception

LiteLLM shipped v1.103.0, introducing native hosted_vllm offline batch processing and updating its MCP integration test suite to the modern MCP 2.x MCPServer architecture. The Admin UI now supports dynamic web search interception controls, while the proxy resolves reasoning-object to reasoning_effort translation for OpenAI o1/o3 series and eliminates Azure 400 errors on empty tool choices.

LiteLLM v1.103.0 Released: Hosted vLLM Batches, MCP 2.x Architecture, and Admin UI Web Search Interception
⚡ Key Takeaways
  • •Hosted vLLM Batching: Directly dispatches and tracks asynchronous batch inference across hosted vLLM clusters from unified LiteLLM client endpoints.
  • •MCP 2.x Alignment & Budget Guardrails: Migrates test harnesses to the modern MCP 2.x MCPServer API while hardening zero-budget project request blocking and team-linked JWT inheritance.
  • •Reasoning Effort & Azure Tool Fixes: Translates provider reasoning tokens into uniform reasoning_effort parameters and eliminates 400 errors on Azure tool_choice when no tools are attached.
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

MemoryAthena: Adaptive Routing Over Latent and Generated Memories for LLM Agents

Challenging the assumption that agent memory must rely solely on external explicit retrieval tables, researchers unveiled MemoryAthena. By establishing three memory pathways—direct Engram retrieval (E), cue-conditioned generation (GE), and causal backbone generation (GH)—governed by a lightweight 201M causal routing head trained on counterfactual token likelihood advantages, MemoryAthena raises five-task QA accuracy from 37.65 to 39.28 while boosting general NLP averages to 79.13.

MemoryAthena: Adaptive Routing Over Latent and Generated Memories for LLM Agents
⚡ Key Takeaways
  • •Three-Way Memory Architecture: Combines direct table retrieval (E), cue-guided generation (GE), and table-free causal backbone generation (GH) into a unified memory fabric.
  • •Counterfactual Causal Routing Head: Keeps the foundational LLM backbone and memory tables frozen, training a 201M routing head on counterfactual future-token likelihood advantages.
  • •Empirical Accuracy Gains: Boosts five-task long-horizon QA benchmarks from 37.65 to 39.28 and lifts six-task general NLP evaluations from 76.73 to 79.13.
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

RayOrch Open-Sourced: Lineage-Controlled Multi-Grain Dataflows Accelerate MinerU Scaling by 15.1x

Addressing the systemic bottlenecks in multimodal document processing (MinerU, Docling) where variable-length child generation breaks GPU batching and lineage tracking, Peking University and OpenDCAI open-sourced RayOrch. Operating under a lineage-controlled multi-grain dataflow model with per-call FIFO ready queues, RayOrch delivers a 15.14x speedup scaling MinerU from 4 to 64 NVIDIA H20 GPUs, cutting end-to-end execution time by 13.1%~16.0% compared to Ray Data and 29.0% compared to Daft.

RayOrch Open-Sourced: Lineage-Controlled Multi-Grain Dataflows Accelerate MinerU Scaling by 15.1x
⚡ Key Takeaways
  • •End-to-End Lineage Preservation: Unlike legacy data systems that flatten records and lose lineage, RayOrch preserves hierarchical parent-child relationships and deterministic order throughout multimodal expansion pipelines.
  • •Cross-Parent FIFO GPU Batching: Introduces per-call FIFO ready queues that aggregate ready child tokens across disparate parents for optimal GPU saturation, paired with typed parent-scoped failure suppression.
  • •Scalability & Latency Reductions: On NVIDIA H20 clusters, scales MinerU across 4 to 64 GPUs with 15.14x speedup, cutting end-to-end latency by 13.1%~16.0% versus Ray Data and 29.0% versus Daft on Docling.
Read details→
H
Hugging Face Daily Papers@HuggingFace·18h ago
🔥 Trending

PISA: Block Sparse Attention with O(N log N) Complexity and Hardware-Aware Triton Kernels

Researchers from SJTU and GAIR introduced PISA, a block-sparse attention mechanism that eliminates the quadratic bottleneck of conventional block selection in long-context LLMs. By adopting a coarse-to-fine pyramid Top-K selection hierarchy across O(log N) key levels, PISA reduces complexity to O(N log N). Accompanied by hardware-aware fused Triton kernels that avoid materializing query-key matrices, it matches dense baselines while substantially boosting long-context retrieval throughput.

PISA: Block Sparse Attention with O(N log N) Complexity and Hardware-Aware Triton Kernels
⚡ Key Takeaways
  • •Overcoming Quadratic Block Selection: Conventional block-sparse attention methods still suffer from O(N²) scoring overhead when selecting retained blocks; PISA achieves true O(N log N) complexity via pyramid Top-K selection.
  • •Coarse-to-Fine Key Hierarchy: Pools key vectors into an O(log N)-level hierarchy, applying LogSumExp scoring across bounded candidate pools to progressively narrow down candidate blocks.
  • •Hardware-Aware Triton Kernels: Custom-built fused Triton kernels for training and inference eliminate the need to materialize query-key score matrices in HBM, slashing memory footprint and memory bandwidth bottlenecks.
Read details→
H
Hugging Face Daily Papers@HuggingFace·18h ago
📊 Benchmark

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-Agent LLMs Across 50+ Rounds

Overcoming the limits of traditional multi-agent benchmarks that only test short-horizon (<20 steps) or competitive settings, UPenn researchers released AgentWorld. Spanning 50+ interaction rounds in an MMORPG sandbox across 100 human-annotated tasks with 3-20 asymmetric black-box agents, it introduces Causal Collaboration Effectiveness (CCE). Even frontier models achieve only 52.0% task success, exposing severe communication breakdown and shared plan decay.

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-Agent LLMs Across 50+ Rounds
⚡ Key Takeaways
  • •Long-Horizon Collaboration: Moves beyond <20-step toy benchmarks by simulating 50+ round interactions in an MMORPG sandbox with 3 to 20 asymmetric agents operating in black-box environments.
  • •Novel CCE Metric: Introduces Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions to measure effective team contribution.
  • •Frontier Model Failure Modes: Evaluated on Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B, top success rate reached only 52.0%, highlighting role confusion and plan erosion.
Read details→