News · Leaderboard

Latest news, beside the leaderboard

Read what just changed, and see who is ahead. Those are the two doors on this page.

Plans
H
Hugging Face Daily Papers@HuggingFace·22m ago
🔥 Trending

SafeActBench: NUS Investigates the Broken Evidence-to-Action Chain in Tool-Using Agents Across Six Operational Domains

Tool-using agents execute consequential modifications to external systems, yet nominal task success frequently conceals unverified actions executed without prior evidence. Researchers from the National University of Singapore (NUS) present a rigorous empirical study on the breakdowns along the evidence-to-action chain (arXiv:2610.07753). Evaluating ten model-harness configurations reveals a striking paradox: high static action assessment accuracy routinely masks brittle interactive execution. Failures overwhelmingly originate prior to execution: agents terminate prematurely before collecting requisite evidence, or initiate state-altering actions before foundational proofs are established. The authors formulate SafeActBench, spanning 656 rigorous cases across six operational domains and five protocols progressing from static judgment to complex multi-action workflows. Supported by a provenance-bound Evidence Ledger and deterministic trajectory evaluator, findings demonstrate that while single-step execution is stable once evidence is verified, multi-action workflows frequently collapse under unresolved prerequisites and truncated state transitions.

⚡ Key Takeaways
  • •NUS presents SafeActBench across 656 operational cases and 6 domains to evaluate where the evidence-to-action chain breaks
  • •Tests across 10 model-harness configurations show high static evaluation masks fragile execution, with over 70% of failures caused by premature actions
  • •Introduces the Provenance-bound Evidence Ledger, proving multi-action workflows suffer 64.7% failure rates from unresolved prerequisites
Read details→
H
Hugging Face Daily Papers@HuggingFace·22m ago
🔥 Trending

EVISKILL: Grounding Skill Evolution in Replayable Evidence Enables Continual Procedural Capability Accumulation Without Weight Updates

Continual skill evolution allows LLM agents to autonomously accumulate, refine, and reuse procedural knowledge across multi-turn interactions without updating model parameters. However, prevailing experience-driven reflection mechanisms suffer from two fundamental weaknesses: they decouple synthesized guidance from concrete behavioral traces, and they rely on monolithic global validation that blindly discards valid local corrections whenever an entire workflow fails. Researchers from Tsinghua University and partner institutions introduce EVISKILL (arXiv:2610.05030), an evidence-grounded framework anchoring procedural adaptation to deterministic execution traces. EVISKILL structures raw observations into Replayable Evidence Cards and synthesizes skill edits with explicit contextual lineage. A targeted replay engine verifies edits through isolated re-execution before cross-epoch refinement. Across three interactive benchmarks spanning six LLM backbones, EVISKILL consistently outperforms state-of-the-art reflective baselines, demonstrating durable procedural accumulation.

⚡ Key Takeaways
  • •Tsinghua University introduces EVISKILL, anchoring continual procedural skill evolution in Replayable Evidence Cards without parameter updates
  • •Decouples local edit verification from monolithic global validation, achieving a 3x increase in the retention of valid localized skills
  • •Validated across 3 interactive benchmarks and 6 foundation LLMs, consistently outperforming standard reflection architectures
Read details→
H
Hugging Face Daily Papers@HuggingFace·22m ago
🔥 Trending

NVIDIA Unveils NeMo-DCR: Bit-Exact Delta-Compressed Refit Slashes Trillion-Parameter Agentic RL Cross-Cluster Sync from 87.5 Min to 150s

Agentic reinforcement learning (RL) disaggregates centralized policy training from large-scale interactive environment rollouts, necessitating frequent cross-cluster policy weight synchronization (refit). Transferring a full 1T-parameter checkpoint across cloud regions currently consumes 87.5 minutes, idling massive rollout compute clusters. NVIDIA researchers introduce NeMo-DCR (Delta-Compressed Refit, arXiv:2610.08430), a bit-exact synchronization architecture that transmits only model deltas while guaranteeing receivers reconstruct identical parameter bits. Exploiting empirical measurements showing only ~1% of BF16 weights alter value bits per step, NeMo-DCR maps training shard changes into canonical coordinates via fixed affine projections and utilizes compressible XOR masks to guarantee exact bit-level parity. Streaming deltas over relay trees and committing changes via atomic joint commits, NeMo-DCR achieves 12x to 40x speedups for models spanning 30B to 1T parameters, slashing 1T cross-region refit latencies from 87.5 minutes to 150 seconds.

⚡ Key Takeaways
  • •NVIDIA introduces NeMo-DCR, delivering 12x to 40x faster weight synchronization for trillion-parameter agentic RL systems
  • •Capitalizes on 1% weight update sparsity in BF16 training using compressible XOR masks to guarantee bit-exact parity without drift
  • •Cuts 1T-parameter cross-region refit latency from 87.5 minutes to 150 seconds, unblocking massive multi-cluster agent training
Read details→
A
AtlassianAtlassian·1h ago
🚀 Release

Atlassian launches Agentic Multiplayer Protocol (AMP) and a rebuilt Atlassian MCP with 220+ tools, 15M+ daily calls and up to 25% fewer tokens

At Team '26 Europe, Atlassian introduced AMP, a framework giving agents an assigned identity, scoped authority, shared Teamwork Graph context and reviewable output, and made a rebuilt Atlassian MCP generally available with 220+ tools, OAuth 2.1/PKCE and tiered read/write/destructive scopes.

Atlassian launches Agentic Multiplayer Protocol (AMP) and a rebuilt Atlassian MCP with 220+ tools, 15M+ daily calls and up to 25% fewer tokens
⚡ Key Takeaways
  • •Atlassian MCP handles 15M+ calls/day; rebuilt from dozens of tools to 220+, with up to 25% token savings on the same Jira/Confluence work in internal Claude benchmarks
  • •Teamwork Graph (250B+ connections) now indexes Bitbucket/GitHub source down to functions, symbols and classes via Rovo Code Search, no clone needed
  • •Internal benchmark: agents grounded in Teamwork Graph were 44% more accurate with 48% fewer tokens
  • •Humans and agents already collaborate 10M+ times a month on Atlassian; Jira agent sessions link local and cloud agent runs back to work items
  • •MCP secured with OAuth 2.1 + PKCE by default, read/write/destructive scopes separated; EU-local AI inference added, non-human identity inventory coming soon
Read details→
ADSponsored
G
Google DeepMindGoogleDeepMind·2h ago
🚀 Release

Google officially ships Nano Banana 2.1 (gemini-nano-banana-2.1) GA: 1K/2K images now half the price of Nano Banana 2, and human-eval Elo 1050 beats Nano Banana Pro

On 2026-10-06 Google officially launched Nano Banana 2.1 and marked it GA in the Gemini API changelog as gemini-nano-banana-2.1. It is based on Gemini 3.6 Flash and is rolling out across the Gemini app, AI Studio, Search AI Mode, Flow, Stitch, Google Ads and Gemini Enterprise Agent Platform. Image output is priced at $30 per 1M tokens (Nano Banana 2: $60), which works out to $0.0336 per 1K image and $0.0504 per 2K image, about half the old price. The model card's side-by-side human eval puts text-to-image Elo at 1050, ahead of Nano Banana 2 (990) and Nano Banana Pro (935). gemini-3.1-flash-image is now deprecated, with no shutdown date announced. This confirms the silent Flow rollout we reported earlier.

Google officially ships Nano Banana 2.1 (gemini-nano-banana-2.1) GA: 1K/2K images now half the price of Nano Banana 2, and human-eval Elo 1050 beats Nano Banana Pro
⚡ Key Takeaways
  • •Per-image price: 1K $0.0336 (was $0.067, -50%), 2K $0.0504 (was $0.101, -50%), 4K $0.113 (was $0.151, -25%); Batch halves it again to $0.0168 per 1K image (pricing)
  • •Text side got pricier: input $1.50/1M tokens (was $0.50), text+thinking output $7.50 (was $3), so re-check costs for long prompts and many reference images
  • •Human-eval Elo (model card): T2I overall 1050 vs NB2 990 vs Pro 935; multi-character consistency 1106 vs 978 vs 1011; infographic factuality 0.521 vs 0.179 vs 0.265
  • •Specs: up to 14 reference images (4 characters + 10 objects), fixed 2K/4K tiling artifacts at 1:8/8:1, Google Web + Image Search grounding, thinking levels minimal/medium/high
  • •Migration: gemini-3.1-flash-image is deprecated (no shutdown date yet); 131,072-token input limit; no function calling, structured outputs or caching (model docs)
Read details→
G
GitHubgithub·2h ago
🚀 Release

GitHub stacked pull requests are generally available on all github.com plans: approvals survive rebases, whole stacks land via merge queue, and gh stack ships an AI-agent skill

On 2026-10-06 GitHub made stacked pull requests generally available on all github.com plans, with GitHub Enterprise Server support coming in a later release. Since the public preview, repos using stacks have merged 9% more code than their peers, and more than two-thirds of the top 1% of repos now use them, with a 5% faster time-to-merge. New in GA: approvals are kept for unchanged code after a rebase, replacement commits are signed automatically, a whole stack enters the merge queue as one merge group, stacks retarget automatically when their base branch is deleted, and stack auto-merge is rolling out over the coming weeks. The gh-stack CLI extension (v0.2.0) now supports Git worktrees, and `gh skill install github/gh-stack` teaches AI coding agents how to split large changes into stacked PRs.

GitHub stacked pull requests are generally available on all github.com plans: approvals survive rebases, whole stacks land via merge queue, and gh stack ships an AI-agent skill
⚡ Key Takeaways
  • •Impact (changelog): repos using stacks merge 9% more code; over 2/3 of the top 1% of repos use them, with 5% faster time-to-merge
  • •Approvals and signing: Rebase stack keeps approvals on unchanged code even when stale approvals are dismissed; replacement commits are signed and keep the original authorship
  • •Merge queue: a stack lands as 1 merge group; with the merge-commit method, each PR gets its own merge commit; stack auto-merge rolls out over the next few weeks
  • •Agents: gh-stack v0.2.0 (2026-10-02, 1.6k stars) supports worktrees and installs an agent skill via gh skill install github/gh-stack; the pull_request webhook gains a stacked action
  • •Coverage: all github.com plans now, GHES in an upcoming release; requires gh v2.0+ and Git 2.36+
Read details→
A
AnthropicAnthropicAI·4h ago
🔥 Trending

Anthropic expands its Cyber Verification Program into three access tiers with Opus 5.5, Sonnet 5.5 and Mythos 5.1; Red Team tier completes 34/50 CyScenarioBench tasks with zero blocks

On 2026-10-06 Anthropic merged Project Glasswing and the original Cyber Verification Program into one three-tier program: Defense Access (SOC, incident response, malware reversing, vuln validation), Red Team Access (authorized pentesting) and Specialized Access (safety-critical systems such as power grids, flight systems and telecom, vetted with the US government). All tiers get Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1 with reduced cyber blocking. On CyScenarioBench, GA models were blocked on the first prompt of every task, Defense tier blocked 46/50 trials, and Red Team tier had zero blocks and completed 34/50 (matching the 67.6% no-safeguard rate). Glasswing partners reported at least 129,000 verified vulnerabilities from April to July 2026.

Anthropic expands its Cyber Verification Program into three access tiers with Opus 5.5, Sonnet 5.5 and Mythos 5.1; Red Team tier completes 34/50 CyScenarioBench tasks with zero blocks
⚡ Key Takeaways
  • •3 tiers: Defense (decisions in days), Red Team (weeks, organizations only), Specialized (in-depth review with the US government); Glasswing members move to Specialized automatically (announcement)
  • •3 models in every tier: Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1, plus future models
  • •CyScenarioBench (10 tasks x 5 attempts): GA blocked on first prompt; Defense tier blocked 46/50; Red Team tier 0 blocks and 34/50 completed, matching the 67.6% no-safeguard rate
  • •Impact: at least 129,000 verified vulns found by Glasswing partners (Apr-Jul 2026) plus 5,500 from Anthropic's own OSS scanning (Apr-Oct), over 33,000 rated critical/high (Project Glasswing)
  • •Available on Claude Platform, Vertex AI and Microsoft Foundry; Bedrock only for Enterprise Frontier Safeguards customers; data retention required
Read details→
F
FeSensFeSens·4h ago
🔥 Trending

openTPU, an open-source AI accelerator developed by AI agents, goes viral: full RTL/ISA/compiler stack runs Qwen3.5, Gemma 4 and a 35B MoE on a Kintex-7 FPGA card at up to 85.8 tok/s

FeSens/openTPU (Apache-2.0) hit the Hacker News front page (250+ points, 300+ comments). Reusing the AI-driven method from auto-arch-tournament (AI-designed RISC-V cores), the author had AI agents iterate on an inference accelerator; one repo holds SystemVerilog RTL, the ISA, a bit-exact simulator, a kernel language/compiler and a profiler. On a Kintex-7 xc7k480t PCIe card it runs Qwen3, Qwen3.5, Gemma 4, Phi-4-mini and others with real weights, token-for-token identical to the simulator; LFM2.5-230M decodes at 85.8 tok/s in 4-bit and Qwen3.5-35B-A3B reaches 3.95 tok/s with host-streamed experts.

openTPU, an open-source AI accelerator developed by AI agents, goes viral: full RTL/ISA/compiler stack runs Qwen3.5, Gemma 4 and a 35B MoE on a Kintex-7 FPGA card at up to 85.8 tok/s
⚡ Key Takeaways
  • •Full stack in one repo: SystemVerilog RTL, 8x32-bit-word ISA, bit-exact Python simulator, @ol.jit kernel compiler and Lens profiler, Apache-2.0 (GitHub)
  • •Measured device decode (4-bit, int8 head): LFM2.5-230M 85.8 tok/s, Qwen3-0.6B 31.3, Qwen3.5-2B 12.09, Gemma 4 E2B 12.14 tok/s
  • •Uses 82-94% of the 17.1 GB/s DDR3-1066 peak while decoding; 133.33 MHz clock, four-column systolic array
  • •MoE offload: Qwen3.5-35B-A3B at 3.95 tok/s (153 MB of experts streamed per token, 62% slot hit), LFM2.5-8B-A1B at 10.6 tok/s (offload docs)
  • •4-bit weights (FP4 with two-level block scales, 4.25 bits/weight) decode 40-45% faster than int8, with per-model perplexity cost documented
Read details→
ADSponsored
H
Hugging Face Daily Papers@HuggingFace·4h ago
🔥 Trending

HERMES: Harness Engineering via Modular Executable Dev-Primitives Empowers Software Engineering Agents to Match Frontier Models

Large Language Model (LLM) coding agents equipped with bash terminals frequently collapse when deployed across massive enterprise codebases. Centralized agents struggle to reconstruct project states fragmented across multi-language source trees, build configs, and test suites, precipitating context window exhaustion and catastrophic semantic drift. Researchers from UIUC introduce Dev-Primitives and the HERMES harness engineering framework (arXiv:2610.07832). Dev-Primitives transform passive software components into active agentic nodes by pairing each repository artifact with a resident LLM endowed with natural-language reasoning, inter-component communication channels, and localized self-modification primitives. Coordinated via a dependency-aware dynamic activation scheduler and execution-guided fault localization, HERMES boosts task performance across four standard software engineering benchmarks by 12.4% over baseline harnesses. Crucially, powering HERMES with lightweight Qwen3-8B Dev-Primitives brings performance within 4.5% of homogeneous GPT-5.6 Sol configurations while slashing inference token overhead by 26.2% on Terminal-Bench 4.0.

⚡ Key Takeaways
  • •UIUC introduces Dev-Primitives and the HERMES framework, transforming passive software code artifacts into active LLM-powered nodes
  • •Dependency-aware dynamic activation and execution-guided bug diagnosis eliminate context explosion across multi-thousand-file repositories
  • •Outperforms baseline harnesses by 12.4% on average; lightweight Qwen3-8B primitives match GPT-5.6 Sol within 4.5% while cutting costs by 26.2%
Read details→
H
Hugging Face Daily Papers@HuggingFace·4h ago
🔥 Trending

Judged Useless, Queried Anyway: NTU Reveals Fatal Decoupling between Evidence Evaluation and Stopping in Tool-Using Agents

When an external tool or search endpoint repeatedly returns irrelevant noise, a rational agent should cease querying and act upon available knowledge. A comprehensive study by Nanyang Technological University (NTU) and A*STAR (arXiv:2610.06191) uncovers a profound behavioral failure mode: while LLM agents correctly identify tool outputs as useless 97% to 100% of the time, they systematically fail to translate those negative evaluations into stopping decisions, entering persistent query loops. Analyzing seven frontier open-source and proprietary models under controlled tool failure modes, the authors show that prompt warnings regarding token budgets, call costs, or explicit stopping guidelines fail to enforce rational early exits—often prompting smaller 7B-8B models to burn calls until hit by hard deadlines. Evidence-grounded stopping emerges reliably only when the execution harness enforces a programmatic policy: mandating final generation after five consecutive self-judged useless queries. This integration rule raises task completion rates across every evaluated model during source failures, confirmed via a 300-question pre-registered replication trial.

⚡ Key Takeaways
  • •NTU and A*STAR reveal a stark behavioral dissociation in tool-using agents: models judge results useless 97-100% of the time, yet query anyway
  • •Prompt-based cost warnings and token budgets fail to stop runaway loops, pushing 7B-8B models to exhaust limits blindly
  • •Enforcing an execution-layer stopping rule after five consecutive self-judged useless queries raises task success across all models
Read details→
H
Hugging Face Daily Papers@HuggingFace·4h ago
🔥 Trending

GUI-HARVEST: Self-Improving GUI Agents via Evidence-Driven Harness Evolution Lifts OSWorld by 12.3 Points and GPT-5 by 13.8%

The executable harness surrounding a GUI model governs how multimodal visual observations are assembled, how mouse/keyboard actions are executed, and how error recovery and termination gates operate. Automatically evolving this harness with frozen base models presents coupled challenges: grounding error diagnosis in visual UI transitions, attributing failures amid execution variability, and translating multi-task failure modes into reliable runtime code edits. Researchers introduce GUI-HARVEST (arXiv:2610.00948), an evidence-driven harness optimizer. GUI-HARVEST synchronizes model intentions with before-and-after screenshots, treats repeated rollouts as joint evidence units, and abstracts recurring failure patterns into bounded Python code edits applied directly to the harness. Evaluated on OSWorld-Verified across six general and GUI-specialized backbones, Qwen3-VL-32B-Instruct improves by 12.33 points. Remarkably, transferring the evolved harness to WindowsAgentArena without further optimization elevates GPT-5 by 13.87 percentage points at 50 steps, decisively outperforming Self-Harness and Meta-Harness. Code is available at github.com/GaryYang12345/GUI-HARVEST.

⚡ Key Takeaways
  • •Introduces GUI-HARVEST, an automatic evidence-driven harness optimizer enabling self-improvement for GUI agents with frozen base models
  • •Grounds error diagnosis in before-and-after UI screenshot deltas, translating multi-task failure modes into bounded Python source code edits
  • •Lifts Qwen3-VL-32B by 12.33 points on OSWorld-Verified and transfers zero-shot to WindowsAgentArena, boosting GPT-5 by 13.87 points
Read details→
S
Sierra@SierraPlatform·6h ago
🔥 Trending

Sierra and Meta announce Personal Agent Protocol, an open OAuth-based standard (with Shopify, Stripe, Walmart and more) for how personal AI agents work with businesses

On 2026-10-06 Sierra co-founders Bret Taylor and Clay Bavor announced Personal Agent Protocol, an open standard Meta and Sierra are developing with Genesys, Instinct, Rocket, Shopify, Stripe and Walmart. Built on OAuth, it lets a personal agent discover what a business offers, start a guest session, then act with user-granted read-only or write access via the company's website, APIs (MCP/OpenAPI) or its own agent. A v0.1 spec, design workshops and a reference implementation are planned for later in October; payments are a future extension.

Sierra and Meta announce Personal Agent Protocol, an open OAuth-based standard (with Shopify, Stripe, Walmart and more) for how personal AI agents work with businesses
⚡ Key Takeaways
  • •8 parties: Meta and Sierra co-develop; Genesys, Instinct, Rocket, Shopify, Stripe and Walmart are partners (announcement)
  • •3 access tiers: guest session, user-granted read-only, user-granted write, all on an OAuth session that carries across channels
  • •3 integration routes chosen by the business: website, APIs (MCP / OpenAPI), or the company's own agent
  • •Timeline: v0.1 spec, design workshops and a reference implementation planned for October 2026; payments, finer permissions and push notifications are future extensions
  • •No spec text or benchmarks are public yet
Read details→
O
OpenAI@OpenAI·8h ago
🔥 Trending

OpenAI publishes 722 math manuscripts from an unreleased internal frontier model on GitHub, claiming results on the quasi-Riemann hypothesis, the Hodge conjecture for CM abelian varieties and more

On 2026-10-06 OpenAI published [Sharing AI progress in mathematics](https://openai.com/index/sharing-ai-progress-in-mathematics/) and released [openai/math](https://github.com/openai/math) (Apache-2.0): 722 manuscripts in 372 result families produced by an unreleased internal frontier model, drawn from an evaluation of ~4,000 open problems at an average of ~3 hours of ChatGPT Pro thinking compute per result, with 10 abridged reasoning summaries and Lean formalizations for many results. OpenAI notes unformalized results may have issues; none are peer reviewed yet.

OpenAI publishes 722 math manuscripts from an unreleased internal frontier model on GitHub, claiming results on the quasi-Riemann hypothesis, the Hodge conjecture for CM abelian varieties and more
⚡ Key Takeaways
  • •Scale: 722 manuscripts in 372 result families, selected from an evaluation of ~4,000 open problems
  • •Compute: ~3 hours of ChatGPT Pro thinking per result on average, same unreleased internal model and fixed procedure
  • •Checkability: 235 of 372 families link to Lean; formalization.yaml lists 162 papers with a formalized main result, checkable via Comparator
  • •Transparency: 10 abridged reasoning summaries (irrationality exponent of pi, Mahler conjectures, free group factors, etc.) plus revision/citation protocols shaped with the IAS advisory group
  • •Caveat: OpenAI says some unformalized results could have issues; nothing is peer reviewed yet and the model is not released
Read details→
O
OpenAI@OpenAI·8h ago
🚀 Release

OpenAI opens the Decisions API public beta: POST /v1/decisions on gpt-6-luna returns typed predicate, choice and score answers ~10x faster than the Responses API, billed only for input at $0.10 per 1M tokens

Per the [OpenAI API changelog](https://developers.openai.com/api/docs/changelog), the [Decisions API](https://developers.openai.com/api/docs/guides/decisions) entered public beta on 2026-10-06 with gpt-6-luna: a dedicated POST /v1/decisions endpoint takes text/images plus a list of questions and returns typed predicate, choice or score answers with probabilities, about 10x faster than the Responses API. Billing is input-only at $0.10 per 1M tokens (no output or cache charges); ZDR and HIPAA are supported and GA is expected within weeks.

OpenAI opens the Decisions API public beta: POST /v1/decisions on gpt-6-luna returns typed predicate, choice and score answers ~10x faster than the Responses API, billed only for input at $0.10 per 1M tokens
⚡ Key Takeaways
  • •Speed: typed answers about 10x faster than the Responses API (Decisions guide)
  • •Price: input-only $0.10 per 1M tokens, no output or cache read/write fees; regular gpt-6-luna is $0.10 in / $0.50 out
  • •3 question types: predicate (0–1 probability), choice (value + probabilities + confidence), score (probability-weighted ordinal level)
  • •Compliance: ZDR and HIPAA eligible, US and Europe (EEA + Switzerland) data residency; GA expected in weeks
  • •Limits: gpt-6-luna only; images must be inline base64, no hosted URLs or file_id
Read details→
H
Hugging Face Daily Papers@HuggingFace·8h ago
🔥 Trending

Engram: Pre-Registered Agent Memory Trial Shows Raw Turn Selection Matches LLM Fact Extraction at 3,061x Lower Cost

Conversational memory for LLM agents traditionally divides into two opposing camps: expensive pipeline extraction that distills dialogues into structured factual memory stores versus raw conversational turn retrieval. Literature has reported contradictory findings. Researchers conducted a pre-registered, double-blind controlled study on held-out LoCoMo and LongMemEval benchmarks (arXiv:2609.34227). The investigation reveals that under constrained context budgets, selecting raw conversational turns via a lightweight typed decision model (Jev) is statistically non-inferior to full LLM fact extraction pipelines (95% one-sided bound -3.0 points vs. -5 margin)—while reducing write-time computational costs by a massive 3,061 times. Furthermore, the authors resolve literature contradictions by proving that reranking gains scale inversely with token budgets, while aggressive reranking degrades an agent's correct abstention reliability. All study protocols, data, and code are open-sourced at github.com/ris3abh/Engram.

⚡ Key Takeaways
  • •Presents Engram, the first pre-registered double-blind study testing raw-turn selection against LLM fact extraction for agent memory
  • •Shows raw turns selected via typed decision model Jev are statistically non-inferior to LLM fact extraction at 3,061x lower write cost
  • •Proves reranking gains decay sharply as token budgets expand, while aggressive reranking degrades an agent's correct abstention rate
Read details→
N
NVIDIA AI@NVIDIA·8h ago
🔥 Trending

NVIDIA NeMo Study: Agentic Retrieval Boosts Complex Search nDCG@10 by 8.7 Points but Incurs 107s Latency and 764K Tokens per Query

Dense retrieval serves as the default paradigm for searching unstructured enterprise corpora in modern RAG systems. However, dense retrieval is fundamentally bound to surface-level semantic similarity, failing when queries require multi-hop exploration, dynamic hypothesis reformulation, and causal synthesis. Researchers from NVIDIA NeMo-Retriever present an in-depth investigation into Agentic Retrieval (arXiv:2610.05750), coupling LLM ReAct loops with dense and lexical retrievers. Evaluating across challenging benchmarks including ViDoRe v3 and BRIGHT, agentic retrieval improves nDCG@10 by 8.7 points using the identical underlying embedding model. Crucially, NVIDIA transparently quantifies the operational trade-offs: agentic searches average 107.4 seconds per query (versus 0.67 seconds for conventional retrieval) and consume an average of 764.1K input tokens and 5.8K output tokens per query. This landmark study provides critical empirical benchmarks for enterprise architectures balancing retrieval depth against inference budgets (github.com/NVIDIA/NeMo-Retriever).

⚡ Key Takeaways
  • •NVIDIA NeMo-Retriever evaluates Agentic Retrieval coupling ReAct loops with dense retrievers for complex multi-hop search
  • •Improves nDCG@10 by 8.7 points using the identical base embedding model across ViDoRe v3 and BRIGHT benchmarks
  • •Transparently quantifies operational costs: average query latency reaches 107.4 seconds (vs 0.67s) consuming 764.1K input and 5.8K output tokens
Read details→
H
Hugging Face Daily Papers@HuggingFace·8h ago
🔥 Trending

OPPD: On-Policy Power Distillation Sharpens Sequence-Level Reasoning in a Single Pass, Matching 64-Candidate Search

Large language models frequently encounter a probability mass dilemma in multi-step reasoning: while the correct sequence might possess a higher individual probability than any single wrong candidate, the vast ocean of incorrect paths collectively dominates the posterior distribution. Sequence-level power sampling sharpens distributions toward the mode by raising probabilities to exponents greater than one, but requires generating and scoring dozens of rollout candidates at runtime. Researchers from USC and collaborators introduce On-Policy Power Distillation (OPPD, arXiv:2610.06804), an algorithm that teaches models to output power-sharpened solutions in a single forward pass without inference-time search. By running an on-policy Sequential Monte Carlo (SMC) teacher-student distillation loop, OPPD boosts single-generation accuracy by 23.0 points on MATH500 and 27.3 points on GSM8K without external reference solutions. Crucially, OPPD outperforms verified-reward GRPO on MATH500 (+3.8), GSM8K (+4.0), and AIME (+5.4) under identical compute budgets, while cross-generalizing to raise HumanEval code generation by 5.3 points. The repository is available at github.com/ArminAzizi98/OPPD.

⚡ Key Takeaways
  • •USC and collaborators introduce On-Policy Power Distillation (OPPD), distilling 64-candidate search distributions into a single pass
  • •Operates with zero reference answers, lifting MATH500 by 23.0 points and GSM8K by 27.3 points, beating verified-reward GRPO
  • •Cross-generalizes from mathematics to code, lifting HumanEval accuracy by 5.3 points, fully open-sourced on GitHub
Read details→
O
OpenAI@OpenAI·12h ago
🔥 Trending

OpenAI and Ironclad train GPT-6 Astra on real contracting software: 55.0% vs GPT-5.6 Sol's 41.6% on 11 workflows, ~48% less time per attempt

OpenAI's first vertical-SaaS computer-use collaboration: with Ironclad it defined 11 legal, procurement and commercial tasks, let models practice in hosted Ironclad environments, and used RL on synthetic tasks. GPT-6 Astra scores 55.0% mean rubric vs 41.6% for GPT-5.6 Sol, with estimated time per attempt down from 37.0 to 19.2 minutes; an internal model reaches 63.7%. OpenAI is inviting more software companies to partner.

OpenAI and Ironclad train GPT-6 Astra on real contracting software: 55.0% vs GPT-5.6 Sol's 41.6% on 11 workflows, ~48% less time per attempt
⚡ Key Takeaways
  • •11 real contracting tasks, each graded on 8–50 criteria; ~30–40 min per task for an experienced human user
  • •GPT-6 Astra (Max) 55.0% mean rubric vs GPT-5.6 Sol (High) 41.6%, ~32% relative gain
  • •Estimated time per attempt 37.0 → 19.2 minutes (~48% lower) — simulated, not measured customer savings
  • •Demo task: Astra ~94% of criteria in ~20 min vs Sol ~85% in ~32 min; an internal model hits 63.7%
  • •Synthetic training tasks built from public SEC EDGAR contracts; no OpenAI or Ironclad customer data used
Read details→
H
Hugging Face Daily Papers@HuggingFace·12h ago
🔥 Trending

ProgressCompass: Context-Aware Embodied Progress Reward Models Drive Long-Horizon Agent Planning

As embodied robotics agents tackle long-horizon tasks across household manipulation and industrial assembly, standard terminal outcome reward models (ORMs) fail to provide actionable learning signals along multi-step trajectories. While Process Reward Models (PRMs) evaluating step-by-step task progress percentages offer a viable solution, researchers from Northwestern University and CMU demonstrate in arXiv:2609.36684 that progress reward models collapse into random noise without appropriate context. The authors introduce ProgressCompass, formalizing the essential context dependencies required for embodied progress evaluation: task objectives, initial environmental baseline states, and continuous action history. Without grounding in preceding action sequences, PRMs routinely misjudge progress direction, confusing partial object states with completed goals. ProgressCompass provides an empirical testbed and modeling architectures to guide search and policy refinement in long-horizon robotic manipulation (andyzworks.github.io/progresscompass/).

⚡ Key Takeaways
  • •Northwestern and CMU present ProgressCompass, formalizing essential context dependencies for embodied Process Reward Models
  • •Demonstrates that progress estimation collapses without joint access to goal specs, initial baselines, and temporal action histories
  • •Boosts long-horizon robotic task success by 38.6% and reduces interaction overhead by 41.2% via context-aware PRM trajectory guidance
Read details→
H
Hugging Face Daily Papers@HuggingFace·12h ago
🔥 Trending

Video2Skill: Streaming Embodied Skill Discovery from Video Feeds Exposes Bottlenecks in Open-World Skill Expansion

Embodied manipulation policies generalize across diverse scenes by recombining a core library of reusable physical skills. However, current agents rely almost exclusively on human-curated, hard-coded skill primitives. To achieve autonomous lifelong learning, agents must perform Streaming Embodied Skill Discovery (SESD)—observing uncurated continuous video streams and incrementally building a persistent skill catalog. Researchers from Northwestern University and CMU introduce Video2Skill (arXiv:2609.36691), benchmarking SESD across robotic manipulation and human kitchen tasks under three core capabilities: temporal event localization, physical transformation grouping, and the meta-decision to reuse existing skills versus instantiating novel primitives. Evaluating 19 open-source Vision-Language Models (VLMs) reveals severe structural limitations: most models cluster manipulation events at near-chance accuracy, and scaling parameters fails to resolve error rates. Furthermore, trained models suffer from skill library stagnation, consolidating familiar primitives while failing to expand into unobserved transformations. The authors propose Counterfactual Library-State Rebalancing (CLaRe) to mitigate clustering collapse (andyzworks.github.io/video2skill/).

⚡ Key Takeaways
  • •Northwestern and CMU formulate Streaming Embodied Skill Discovery (SESD) and introduce Video2Skill benchmark across robot and human videos
  • •Evaluates 19 open VLMs showing near-chance clustering accuracy; joint perception merges distinct skills while text pipelines duplicate them
  • •Uncovers library stagnation: models fail to create novel skill primitives for unseen transformations, capping library size at <50% reference
Read details→