News · Leaderboard

Latest news, beside the leaderboard

Read what just changed, and see who is ahead. Those are the two doors on this page.

Plans
O
OpenAIOpenAI·1h ago
🔥 Trending

OpenAI and Ironclad train GPT-6 Astra on real contracting software: 55.0% vs GPT-5.6 Sol's 41.6% on 11 workflows, ~48% less time per attempt

OpenAI's first vertical-SaaS computer-use collaboration: with Ironclad it defined 11 legal, procurement and commercial tasks, let models practice in hosted Ironclad environments, and used RL on synthetic tasks. GPT-6 Astra scores 55.0% mean rubric vs 41.6% for GPT-5.6 Sol, with estimated time per attempt down from 37.0 to 19.2 minutes; an internal model reaches 63.7%. OpenAI is inviting more software companies to partner.

OpenAI and Ironclad train GPT-6 Astra on real contracting software: 55.0% vs GPT-5.6 Sol's 41.6% on 11 workflows, ~48% less time per attempt
⚡ Key Takeaways
  • •11 real contracting tasks, each graded on 8–50 criteria; ~30–40 min per task for an experienced human user
  • •GPT-6 Astra (Max) 55.0% mean rubric vs GPT-5.6 Sol (High) 41.6%, ~32% relative gain
  • •Estimated time per attempt 37.0 → 19.2 minutes (~48% lower) — simulated, not measured customer savings
  • •Demo task: Astra ~94% of criteria in ~20 min vs Sol ~85% in ~32 min; an internal model hits 63.7%
  • •Synthetic training tasks built from public SEC EDGAR contracts; no OpenAI or Ironclad customer data used
Read details→
H
Hugging Face Daily Papers@HuggingFace·1h ago
🔥 Trending

IGMWorld: Benchmarking Coding Agents on Executable World Editing across Minecraft and Terraria via Industry-Grade Modding

Interactive world models are progressing beyond passive simulation toward active world generation and agency. However, deliberately editing an existing, running executable world while preserving unedited systemic invariants remains an underexplored frontier. An international research consortium introduces IGMWorld and IGMBench (arXiv:2610.02331), framing world editing through industry-grade game modding across Minecraft (Java) and Terraria (C#/.NET). Spanning 110 tasks evaluated against 1,100+ deterministic executable criteria across property, entity, dynamics, and system-level intervention depths, IGMBench rigorously assesses coding agent precision. Frontier agents achieve 78.2% task-level success and 94.8% criterion-level pass rates. Crucially, failure rates scale monotonically with intervention depth—most failed edits build and load without syntax errors but diverge in runtime execution logic. Project assets are accessible at vinesmsuic.github.io/IGMWorld.

⚡ Key Takeaways
  • •Introduces IGMWorld and IGMBench, the first benchmark for executable world editing using Minecraft and Terraria game modding
  • •Defines four hierarchical intervention depths (property, entity, dynamics, system) with 110 tasks and 1,100+ deterministic executable criteria
  • •Frontier coding agents achieve 78.2% task-level and 94.8% criterion-level accuracy, identifying sharp drops in deep system interventions
Read details→
H
Hugging Face Daily Papers@HuggingFace·1h ago
🔥 Trending

ProgressCompass: Context-Aware Embodied Progress Reward Models Drive Long-Horizon Agent Planning

As embodied robotics agents tackle long-horizon tasks across household manipulation and industrial assembly, standard terminal outcome reward models (ORMs) fail to provide actionable learning signals along multi-step trajectories. While Process Reward Models (PRMs) evaluating step-by-step task progress percentages offer a viable solution, researchers from Northwestern University and CMU demonstrate in arXiv:2609.36684 that progress reward models collapse into random noise without appropriate context. The authors introduce ProgressCompass, formalizing the essential context dependencies required for embodied progress evaluation: task objectives, initial environmental baseline states, and continuous action history. Without grounding in preceding action sequences, PRMs routinely misjudge progress direction, confusing partial object states with completed goals. ProgressCompass provides an empirical testbed and modeling architectures to guide search and policy refinement in long-horizon robotic manipulation (andyzworks.github.io/progresscompass/).

⚡ Key Takeaways
  • •Northwestern and CMU present ProgressCompass, formalizing essential context dependencies for embodied Process Reward Models
  • •Demonstrates that progress estimation collapses without joint access to goal specs, initial baselines, and temporal action histories
  • •Boosts long-horizon robotic task success by 38.6% and reduces interaction overhead by 41.2% via context-aware PRM trajectory guidance
Read details→
H
Hugging Face Daily Papers@HuggingFace·1h ago
🔥 Trending

Video2Skill: Streaming Embodied Skill Discovery from Video Feeds Exposes Bottlenecks in Open-World Skill Expansion

Embodied manipulation policies generalize across diverse scenes by recombining a core library of reusable physical skills. However, current agents rely almost exclusively on human-curated, hard-coded skill primitives. To achieve autonomous lifelong learning, agents must perform Streaming Embodied Skill Discovery (SESD)—observing uncurated continuous video streams and incrementally building a persistent skill catalog. Researchers from Northwestern University and CMU introduce Video2Skill (arXiv:2609.36691), benchmarking SESD across robotic manipulation and human kitchen tasks under three core capabilities: temporal event localization, physical transformation grouping, and the meta-decision to reuse existing skills versus instantiating novel primitives. Evaluating 19 open-source Vision-Language Models (VLMs) reveals severe structural limitations: most models cluster manipulation events at near-chance accuracy, and scaling parameters fails to resolve error rates. Furthermore, trained models suffer from skill library stagnation, consolidating familiar primitives while failing to expand into unobserved transformations. The authors propose Counterfactual Library-State Rebalancing (CLaRe) to mitigate clustering collapse (andyzworks.github.io/video2skill/).

⚡ Key Takeaways
  • •Northwestern and CMU formulate Streaming Embodied Skill Discovery (SESD) and introduce Video2Skill benchmark across robot and human videos
  • •Evaluates 19 open VLMs showing near-chance clustering accuracy; joint perception merges distinct skills while text pipelines duplicate them
  • •Uncovers library stagnation: models fail to create novel skill primitives for unseen transformations, capping library size at <50% reference
Read details→
ADSponsored
G
Google DeepMind@GoogleDeepMind·3h ago
🚀 Release

Google DeepMind releases EmbeddingGemma 2: a 740M Apache-2.0 multimodal embedding model that maps text, code, images, audio and video into one space, +9.92 on MTEB Code

On 2026-10-06 Google DeepMind released [EmbeddingGemma 2](https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/), a 740M-parameter Apache-2.0 embedding model built on Gemma 4 that maps text, code, images, audio and video into one 768-dim space with an 8K context. MTEB Code rises from 68.76 to 78.68, and quantized text-only inference needs ~191MB RAM on a Pixel 11 Pro. Hugging Face Transformers v5.19.0 supports it natively.

Google DeepMind releases EmbeddingGemma 2: a 740M Apache-2.0 multimodal embedding model that maps text, code, images, audio and video into one space, +9.92 on MTEB Code
⚡ Key Takeaways
  • •Size: 740M params, modular 270M text backbone + 170M vision + 300M audio encoders, loadable on demand
  • •Code retrieval: MTEB Code 68.76 → 78.68 (+9.92); Google claims leading sub-1B multimodal scores on MTEB Code and MAEB
  • •Storage: Matryoshka embeddings truncate 768 → 512/256/128 dims, up to 6x smaller local vector stores
  • •On-device: ~191MB RAM text-only / ~567MB full multimodal (quantized, Pixel 11 Pro); 8K context (4x v1) covers 5.5 min audio, 29 images or 58 video frames
  • •Ecosystem: Apache 2.0 weights on Hugging Face and Kaggle; supported in Transformers v5.19.0, plus sentence-transformers, vLLM, llama.cpp, SGLang, Ollama and MLX per Google
Read details→
H
Hugging Face Daily Papers@HuggingFace·3h ago
🔥 Trending

DeskForge: Scalable Dense Supervision from Controllable Desktop Environments Propels Computer-Use Agent Grounding by 11.5%

Computer-use agents face fundamental grounding bottlenecks in real-world desktop environments due to layered windows, shifting layouts, and visually indistinguishable controls. Existing training datasets suffer from sparse labels and static capture conditions. Researchers from ETH Zurich and IBM Research present DeskForge (arXiv:2610.02320), a controllable synthetic desktop environment that orchestrates real-world software across varied window states, themes, and screen resolutions. DeskForge synthesizes dense annotations by fusing desktop screenshots, accessibility trees, and geometric window hierarchies, yielding DeskForge-1M—a corpus of 1.2M observations and 159.7M annotated UI element instances. Fine-tuning vision-language models on 200K DeskForge-1M examples improves zero-shot GUI grounding across five external benchmarks; Qwen3.5-4B gains 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G. In end-to-end task execution, task completions surge from 31 to 50 on WebArena-Infinity and from 3 to 15 on OpenApps. The dataset, models, and code are open-sourced at saidgurbuz.github.io/deskforge.

⚡ Key Takeaways
  • •ETH Zurich and IBM introduce DeskForge, a controllable environment synthesizing DeskForge-1M with 1.2M observations and 159.7M UI elements
  • •Fine-tuning Qwen3.5-4B boosts GUI grounding by 11.51 points on ScreenSpot-Pro and 10.11 points on OSWorld-G across five external benchmarks
  • •Multiplies long-horizon task completion: solves 50/119 tasks on WebArena-Infinity (up from 31) and 5x on OpenApps (from 3 to 15), fully open-sourced
Read details→
H
Hugging Face Daily Papers@HuggingFace·3h ago
🔥 Trending

LMBuild: Benchmarking LLM Agents on Generating Buildable and Functional 3D Structures with Physics and Kinematic Constraints

LLM agents have demonstrated striking dexterity in generating intricate 3D geometries, raising expectations for automated industrial design and physical manufacturing. However, producing visually aesthetic 3D meshes is fundamentally distinct from synthesizing objects that can be physically manufactured, assembled, and operate as intended. Contemporary benchmarks evaluate visual fidelity while ignoring physical realizability, joint kinematics, and assembly sequencing. Researchers from UIUC and Microsoft introduce LMBuild (arXiv:2610.04292), a comprehensive benchmark for evaluating LLM agents on generating buildable and functional 3D structures. LMBuild models generated objects as assembled physical systems encompassing discrete part decompositions, mechanical joints, physical materials, and procedural build sequences. Evaluating 30 frontier agent systems, LMBuild discovers that while basic geometric structural soundness is solved by top frontier models, kinematic functional affordance and physical operability remain profound bottlenecks; moreover, providing explicit functional specifications dramatically improves assembly completeness and physical kinematics.

⚡ Key Takeaways
  • •UIUC and Microsoft present LMBuild, a 64-page benchmark evaluating LLM agents on generating physically buildable, functional 3D assemblies
  • •Formalizes structures into part decompositions, kinematic joints, materials, and assembly sequences under four physical evaluation dimensions
  • •Benchmarking 30 agent systems reveals that while geometry alignment is solved, physical operability remains difficult, needing functional specifications
Read details→
H
Hugging Face Daily Papers@HuggingFace·3h ago
🔥 Trending

Self-Generated Feedback Destabilizes Test-Time Training: Causal Analysis and the Settlement Verification Protocol

Test-Time Training (TTT) enables language models to internalize streaming context directly into neural parameters during inference, emerging as a foundational architecture to transcend fixed context window bottlenecks. However, when autonomous agents learn continually from their own self-generated outputs, catastrophic representation degradation inevitably manifests over long deployment streams. Researchers from KAUST present a causal decomposition of this failure in arXiv:2610.05076. Across 128K-token sequences across three native TTT-E2E configurations (125M, 760M, 3B) as well as gradient-adapted Qwen3-4B models, updating weights on self-generated text severely degrades predictive capability on real human-written text. Three causal matching experiments pinpoint the culprit: a fundamental local alignment conflict where self-updates overfit idiosyncratic model outputs while actively corrupting real-world text distributions. To safeguard continual adaptation, the authors propose Settlement, an evidence-grounded verification protocol that assesses prospective parameter updates on independent real-world tokens prior to permanent weight commitment, eliminating over 98% of collapse damage.

⚡ Key Takeaways
  • •KAUST presents a causal decomposition of test-time training (TTT) failure under self-generated feedback across 128K-token sequences
  • •Demonstrates that updating weights on self-generated text improves self-fit while degrading general language modeling across diverse scales
  • •Introduces Settlement, an evidence-grounded verification protocol eliminating >98% of collapse damage by validating updates on independent text
Read details→
ADSponsored
G
GitHub Blog@github·5h ago
🚀 Release

GitHub open-sources ReviewBench, a 219-PR benchmark for AI code review: Copilot, Devin and Qodo lead; Cursor's grounded recall is 9.6%

GitHub's blog post [ReviewBench](https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/) (Oct 5, 2026) launches an offline AI code-review benchmark as a research preview. It has 219 PRs from 187 open-source repos in 19 languages, sampled to match the distribution of 103.9M GitHub PRs. The golden set combines human reviewers, several frontier LLMs and static analysis. Claude Sonnet 5 is the judge, and senior engineers agreed with its labels 96.6% of the time. On the [leaderboard](https://review-bench.ai/), grounded recall is 26.0% for Copilot Code Review (Balanced), 23.8% for Devin, 22.1% for Qodo and 9.6% for Cursor. The dataset, judge prompt and runner are MIT-licensed on [GitHub](https://github.com/review-bench/ReviewBench).

GitHub open-sources ReviewBench, a 219-PR benchmark for AI code review: Copilot, Devin and Qodo lead; Cursor's grounded recall is 9.6%
⚡ Key Takeaways
  • •Scale: 219 PRs from 187 open-source repos in 19 languages, sampled to match the distribution of 103.9M GitHub PRs (blog)
  • •Trust: senior engineers independently re-labeled every golden finding and agreed with the benchmark 96.6% of the time; Claude Sonnet 5 is the judge
  • •Leaderboard grounded recall: Copilot Balanced 26.0% (87.8% precision) > Devin 23.8% > Qodo 22.1% > Codex GPT-5.6 Sol Ultra 19.9% > Cursor 9.6%
  • •Offline matched online: in a multi-model ensemble experiment, the online A/B showed +8.0% precision, +13.6% recall and −8.0% cost per review; critical comments were predicted at +227% offline and measured at +262% online
  • •To submit: bring your own container image and model key, run the 25-PR test set, then the full 219 PRs over 3 rounds. The repo review-bench/ReviewBench is MIT-licensed
Read details→
M
Mistral AI@MistralAI·7h ago
🚀 Release

Mistral Large 4 "Le Chonk" public preview: 1.05T-total / 49B-active multimodal MoE, 1M context, open weights by month-end

On 2026-10-06 Mistral launched a public-preview API for [Mistral Large 4](https://mistral.ai/news/mistral-large-4/): a granular MoE with 1.05T total / 49B active parameters plus a 1.6B vision encoder, 1M context, trained from scratch on 3,800 Grace Blackwell GPUs in Mistral's own European datacenters. It posts 61.7% DeepSWE v1.1 and 93% Cybench; weights are promised by the end of October.

Mistral Large 4 "Le Chonk" public preview: 1.05T-total / 49B-active multimodal MoE, 1M context, open weights by month-end
⚡ Key Takeaways
  • •Scale: 1.05T total / 49B active MoE + 1.6B vision encoder, 1M context
  • •Coding (Artificial Analysis private eval): DeepSWE v1.1 61.7%, SWE-Atlas-QnA 59.4%, Terminal-Bench 4 28.3%, Coding Agent Index 49.8%
  • •Security: Cybench 93%; 82% on AA Cyber Index reproduce-and-patch test (highest of any model); Lakera B3 93.3% attack resistance
  • •Agents: AutomationBench 59.9%, AA-Briefcase 1,393 Elo; Surge AI blind coding eval 3.74, behind only Claude Opus 5 (4.22)
  • •Pricing per 1M tokens: $1.36 input / $0.14 cached / $4.18 output (docs also list a half-price tier); weights by end of month
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

UndoBench: Decoupling Nominal Task Competence from Fault Recovery in Tool-Using AI Agents via Counterfactual Mutation Paired Trials

Tool-using AI agents are rapidly assuming mission-critical automation responsibilities across enterprise software ecosystems. However, standard agent benchmarks focus exclusively on nominal, forward task completion—conflating initial planning competency with operational resilience and fault recovery. In real-world enterprise environments where network drops, idempotency failures, and partial commits are commonplace, unguided LLM retries routinely cause devastating side effects such as duplicate payments or corrupted databases. Researchers introduce UndoBench (arXiv:2610.05622), a comprehensive evaluation suite spanning 36 enterprise workflows and 36 production fault scenarios across 8 enterprise verticals. Utilizing counterfactual paired trials alongside wire-level effect-history and state oracles, UndoBench decouples task competence from recovery capability. Experiments reveal a severe vulnerability gap: while nominal task competence achieves 83.54%, conditional recovery success rates (CRSR) collapse to 46.72%, with naive retry producing duplicated external mutations in 53.33% of trials. The code is available at GitHub (tradertanmay/undobench).

⚡ Key Takeaways
  • •Introduces UndoBench, the first enterprise fault recovery and undo benchmark for tool-using AI agents across 8 domains and 36 fault scenarios
  • •Decouples nominal task competence from fault recovery: nominal task completion hits 83.54% while conditional recovery plunges to 46.72%, with naive retries causing duplicate mutations in 53.33% of trials
  • •Pinpoints phase-dependent execution vulnerabilities, urging agent frameworks to move from naive retry loops to deterministic compensation transactions
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

SHIFT: Dynamic Multi-Agent Harness Search Per-Query via Predictive MCTS Outperforms Static Baselines by 7.2%

Designing optimal multi-agent harnesses—specifying roles, prompts, tool registries, and communication graphs—is foundational to task performance. However, because different user requests demand wildly disparate cognitive workflows, static hand-crafted multi-agent topologies fail to generalize; conversely, running exhaustive trial-and-error executions at inference time incurs prohibitive latency and cost. Researchers from Google Cloud AI Research, Georgia Tech, and UC Davis introduce SHIFT (arXiv:2610.04137), an architecture that decouples physical execution from the per-query search loop. SHIFT trains a local LLM architect policy over harness-building actions paired with a learned value function predicting the utility frontier between accuracy and execution token cost. For every incoming query, Monte Carlo Tree Search (MCTS) constructs a customized multi-agent topology tailored to the problem. Evaluated across 9,193 tasks spanning six diverse benchmarks with a Gemini 3.5 Flash executor, SHIFT achieves an unprecedented ~80% mean accuracy, beating 17 competitive baselines by up to 7.2 percentage points while slashing execution tokens by 32%.

⚡ Key Takeaways
  • •Introduces SHIFT, a framework dynamically generating tailored multi-agent harnesses per-query via predictive Monte Carlo Tree Search
  • •Decouples execution from the search loop using an LLM architect policy and a learned value function balancing accuracy against token cost
  • •Achieves ~80% mean accuracy across 9,193 tasks in six benchmarks with Gemini 3.5 Flash, outperforming 17 baselines while cutting execution tokens by 32%
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

Foundations of Proactive Agents: 3T Principles and Proactivity-Gym Benchmark Reveal Trust Erosion from Uncalibrated Interventions

Contemporary LLM agents operate overwhelmingly in a reactive posture, awakening only when commanded by explicit user prompts. As continuous ambient compute becomes cost-effective, proactive agents that exploit idle cycles to anticipate needs and execute support before users ask represent the next frontier of intelligent personal assistants. However, proactivity is a double-edged sword: even flawlessly completed autonomous tasks can disrupt user focus, impose heavy cognitive verification overhead, and destroy human trust if initiated at misaligned moments. Researchers from KAIST and the University of Minnesota establish foundational theory for proactive agents in arXiv:2609.37267, framing design around three joint principles (3T): Task Capability, Temporal Allocation (aligning compute with cognitive availability), and Trust. They introduce PROACTIVITY-GYM, a stateful multi-day evaluation testbed with persona-conditioned simulated users. Evaluating 23 model-harness configurations alongside a 30-participant human study, the authors show that poorly timed interventions trigger severe trust collapse despite perfect task outcomes, whereas asynchronous assistance during user downtime (e.g., sleep-time compute) preserves deep user engagement.

⚡ Key Takeaways
  • •Establishes foundational theory for proactive agents around joint 3T principles: Task Capability, Temporal Allocation, and Trust
  • •Releases PROACTIVITY-GYM, a stateful multi-day evaluation testbed featuring persona-conditioned simulated users across continuous scenarios
  • •A 30-participant human study demonstrates that untimely interventions cause severe trust collapse despite correct outputs, validating sleep-time assistance as the optimal interaction paradigm
Read details→
K
Kandinsky Lab@AICoder·11h ago
🔥 Trending

Kandinsky 6.0 Video goes open source: 3B/29B joint audio-video diffusion models with 44 kHz lip-synced sound, MIT weights, runs on 16 GB GPUs

Kandinsky Lab open-sourced Kandinsky 6.0 Video on Oct 6: Lite (3B) and Pro (29B) text/image-to-audio-video diffusion models that generate 5-second clips with synchronized 44 kHz audio including lip-sync, plus built-in super-resolution to 1080p. Dual-stream CrossDiT; code, weights and Diffusers integration under MIT. The report claims it beats open LTX 2.5 on most VABench metrics and is competitive with proprietary systems on speech quality.

Kandinsky 6.0 Video goes open source: 3B/29B joint audio-video diffusion models with 44 kHz lip-synced sound, MIT weights, runs on 16 GB GPUs
⚡ Key Takeaways
  • •Scale: Lite 3B / Pro 29B; Pro = 19B video stream + 5B audio stream + 5B cross-attention (report)
  • •Output: 5 s clips with synchronized 44 kHz audio incl. lip-sync, T2AV and I2AV, built-in SR to 1920×1080
  • •Evals: beats open LTX 2.5 on most VABench metrics; trades wins with Veo 3.1 Fast (better artifacts/camera motion, behind on speech/sync); behind MiniMax H3 and Seedance 2.0 on visuals
  • •Speed: distilled to 10 NFE with 51% vs 49% human preference vs full model; non-distilled Pro Full HD ≈402 s on H100, ≈1247 s on RTX 4090
  • •Deploy: block offload cuts Pro SD peak memory 72.8 → 21.7 GiB, 16 GB preset; Diffusers, ComfyUI and vLLM-Omni support; MIT license
Read details→
H
Hugging Face Daily Papers@HuggingFace·13h ago
🔥 Trending

Code2Games: Enabling Coding Agents for Playable 3D Gaming World Generation via Blender-to-UE5 Coordinated Reconstruction

Synthesizing high-fidelity, interactive 3D gaming worlds directly from natural language game concepts requires holistic joint reasoning across scene geometry, spatial topology, gameplay objectives, and executable game mechanics. However, existing LLM coding agents typically generate isolated 3D assets or rudimentary scripts, resulting in broken dependencies and spatial-logical incoherence. To bridge this gap, researchers from Peking University and collaborators introduce Code2Games (arXiv:2610.05033), a dedicated agentic framework that generates structured 3D gaming worlds. Code2Games anchors world generation upon a procedural Blender foundation using a shared scene-gameplay representation that enforces persistent entity correspondence. It then transfers the assets into Unreal Engine 5 (UE5) and deploys an execution-guided reconstruction pipeline driven by compilation diagnostics, runtime traces, and automated playtesting feedback to dynamically rectify engine adaptation discrepancies. Evaluated on the newly introduced GameCode4D benchmark spanning ten tiers of gameplay complexity, Code2Games establishes new state-of-the-art marks in visual quality, interactive fidelity, and complete playable game stability.

⚡ Key Takeaways
  • •Introduces Code2Games, an agentic framework generating structured, playable 3D gaming worlds bridging Blender foundation to Unreal Engine 5
  • •Develops a shared scene-gameplay representation coupled with execution-guided reconstruction resolving engine adaptation defects via diagnostic loops
  • •Releases GameCode4D benchmark across ten complex game prompts, surpassing direct agent baselines in visual fidelity and playability
Read details→
H
Hugging Face Daily Papers@HuggingFace·13h ago
🔥 Trending

Periscope: Extending Frozen LLMs Beyond Context Windows via Training-Free Grid Factorization (4.5M Tokens on a Single 80GB GPU)

Standard language models read long texts through quadratic attention passes that bound context windows, deplete GPU memory, and suffer from length-induced accuracy degradation. Researchers from the Technology Innovation Institute (TII) introduce Periscope (arXiv:2610.04047), a training-free inference methodology that factorizes context reading across a 2D grid. Observing that long-context reasoning over discrete decision spaces is fundamentally an evidence-localization task, Periscope arranges N text chunks onto a K x K grid (where K = ceil(sqrt(N))). It probes a frozen LLM with K consecutive local spans and K strided spans across the document, scoring single-token log-odds to generate a comprehensive evidence map whose peak pinpoints the exact answer chunk. For a context of length s and chunk size c, a window W scales to W^2 / c tokens at s^1.5 cost. As each forward probe caches only a single span, a 27B model processes 4.5M tokens on a single 80GB GPU—whereas a conventional pass would demand 296GB of KV cache. On InfiniteBench, Periscope outperforms the best full-window reads by 5 points while dominating long-document retrieval benchmarks.

⚡ Key Takeaways
  • •Introduces training-free Periscope algorithm factorizing N chunks onto a K x K grid to scale physical window W to W^2 / c equivalent tokens
  • •Synthesizes evidence maps via local and strided log-odds, enabling a 27B model to digest 4.5M tokens on a single 80GB GPU vs 296GB KV cache
  • •Outperforms full-window reads on InfiniteBench by 5.2 points, matching 1M window benchmarks on LongBench v2 reading only 9k top-ranked tokens
Read details→
H
Hugging Face Daily Papers@HuggingFace·13h ago
🔥 Trending

MemAdapter: Counterfactual Adaptation Against Memory-Induced Sycophancy in Long-Term Memory LLM Agents

Persistent long-term memory empowers LLM-based agents to maintain personal preferences and contextual continuity across extensive multi-session interactions. However, accumulated memories frequently induce 'Memory-Induced Sycophancy'—compelling agents to slavishly conform to a user's historical misconceptions, obsolete beliefs, or flawed arguments at the expense of objective truth. Prior mitigation efforts operate under the premise that sycophancy stems exclusively from polluted or biased memories, attempting heuristic filtering at storage or retrieval stages. In reality, completely objective and valid memories can still trigger sycophancy when inappropriately prioritized across shifting conversational contexts. To overcome this systemic vulnerability, researchers from Jilin University and collaborators present MemAdapter (arXiv:2610.05162). MemAdapter adaptively modulates retrieved memory influence across three decoupled components: Counterfactual Induction, Context-Aware Reflection, and Evidence-Based Reasoning. Open-sourced on GitHub (DEEP-JLU/MemAdapter), the framework systematically restores objective grounding across three major sycophancy benchmarks while preserving benign personalization.

⚡ Key Takeaways
  • •Exposes memory-induced sycophancy in LLM agents and introduces MemAdapter for counterfactual memory calibration
  • •Decouples counterfactual induction, context-aware reflection, and evidence-based reasoning to segregate facts from subjective preferences
  • •Reduces sycophantic alignment rates by 45%-60% across three benchmarks while preserving over 98% benign personalization integrity
Read details→
O
OpenAI Codex (Thibault Sottiaux)@thsottiaux·17h ago
🚀 Release

Codex kicks off "28 days of improvements or a full reset": Day 1 makes GPT-6 Astra and GPT-6.1 Sol ~50% faster by default, including OpenCode, Pi, Amp and Devin via Sign in with ChatGPT

On Oct 5, Codex lead Thibault "Tibo" Sottiaux committed that for the next 28 days OpenAI will ship one clear improvement for most Codex/ChatGPT Work users each day, or give everyone a full usage reset. Day 1 is an inference speedup: GPT-6 Astra and GPT-6.1 Sol are ~50% faster by default on subscriptions (roughly 30 to 50 tokens/s), including partner tools using Sign in with ChatGPT. No API speed change was announced.

⚡ Key Takeaways
  • •28-day pledge: each day either one clear improvement or a full usage reset for everyone (OpenAI Developer Community thread)
  • •Day 1: GPT-6 Astra / GPT-6.1 Sol default speed ~+50%; Tibo cited roughly 30 to 50 tokens/s (closer to +67% if exact; both called approximate)
  • •Scope: OpenAI products plus every Sign in with ChatGPT partner (OpenCode, Pi, Amp, Devin), no config change, live within ~2 hours
  • •Subscription path only; no API speed change announced, so API-billed evals are unaffected
  • •One community user reported slower 6.1 Sol runs on a private benchmark after the change; unverified single data point
Read details→
Q
Q Labs (Samip Dahal, Bishwas Mandal, Serdar Gülbahar, Akshay Vegesna)@AICoder·17h ago
🔥 Trending

Dust: Q Labs pretrains transformers without backpropagation, matching or beating backprop at small scale and ~10³–10⁴× more efficient than weight-space ES, MIT-licensed code

Q Labs published Dust (Hacker News front page), a zeroth-order, forward-only method that perturbs activations independently per token so one forward pass evaluates thousands of virtual population members, then estimates gradients from loss changes. Pretraining GPT-style transformers on FineWeb, Dust beats tuned backprop at 100k and 1M tokens and closes the gap with population at 10M and 20M; the largest model tested is 243M parameters. The authors say it is not yet compute-efficient enough to replace backprop.

⚡ Key Takeaways
  • •At 1M tokens Dust's test loss drops below backprop from ~1k draws; at 100k from a few hundred (paper)
  • •20M tokens: power-law fitted limit 4.431 (95% CI 3.89–4.58) vs backprop 4.633; authors call it loosely constrained trend evidence
  • •Efficiency: ~10³–10⁴× better than weight-space ES (EGGROLL); EGGROLL at 256× the population still trails Dust at 64 draws
  • •Counterintuitive: at 10M tokens a 243M model beats a 120× smaller one at most population sizes
  • •qlabs-eng/dust (MIT): default population 16,384, 8-GPU torchrun for 1M/10M runs, plus a smaller single-GPU setup
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

ASCENT: Online Test-Time Training (OaTTT) for Long-Horizon Agents Enables Self-Evolution via Verified Experience Distillation

As autonomous LLM agents navigate complex long-horizon operational streams in production, enabling continuous test-time weight adaptation without catastrophic policy drift remains a foundational challenge. Conventional in-context adaptation stores reflections or skills as text, making reuse bottlenecked by fragile retrieval heuristics over a static, frozen policy; conversely, directly imitating or reinforcing tokens from single deployment rollouts destabilizes policy weights. Researchers formalize Online Agentic Test-Time Training (OaTTT) and introduce ASCENT (arXiv:2610.05303). The agent executes each task in a single streaming pass. A frozen copy of the initial model acts as a hindsight teacher, receiving verified execution trajectories as privileged context to compute calibrated next-token predictive distributions. Distilling these distributions into persistent LoRA fast weights continuously updates policy parameters across subsequent tasks without external teacher models or gold references. By pruning invalid-action turns to distill condensed execution trajectories, ASCENT systematically improves task success rates and interaction efficiency across ALFWorld, WebShop, and AppWorld, successfully transferring learned behaviors to held-out environments.

ASCENT: Online Test-Time Training (OaTTT) for Long-Horizon Agents Enables Self-Evolution via Verified Experience Distillation
⚡ Key Takeaways
  • •Formalizes Online Agentic Test-Time Training (OaTTT) and introduces ASCENT for teacher-free continual self-evolution
  • •Leverages a frozen initial model snapshot for privileged hindsight distillation, consolidating pruned traces into persistent LoRAs
  • •Demonstrates monotonic task success gains across ALFWorld, WebShop, and AppWorld with robust transfer to held-out environments
Read details→