News Β· Leaderboard

Latest news, beside the leaderboard

Read what just changed, and see who is ahead. Those are the two doors on this page.

Plans
C
CrowdStrike IntelligenceCrowdStrikeΒ·7h ago
πŸ”₯ Trending

CrowdStrike: attacker used open-source AI pentest agent ARTEX plus Claude Code to breach South Korean financial firms

CrowdStrike Intelligence found Claude Code session histories, ARTEX configs and a CLAUDE.md in attacker open directories, showing a likely Chinese-speaking, financially motivated actor used the China-built open-source agentic pentest framework ARTEX (DeepSeek v4.1-flash primary, plus GLM-5.3 and Grok 4.6) to breach loan-inquiry and employee mobile systems at South Korean financial institutions from late September to early October 2026.

⚑ Key Takeaways
  • β€’Window: late Sep to early Oct 2026; CrowdStrike attribution is moderate confidence, no named adversary
  • β€’Model stack: DeepSeek v4.1-flash as ARTEX primary backend via likely reseller xcai[.]pro; GLM-5.3 and Grok 4.6 in Claude Code sessions
  • β€’IOCs: 1 actor-controlled IP (38.244.50[.]120) plus 9 proxy IPs, with MITRE ATT&CK mapping
  • β€’Reported impact: Shinhan 25,729 customers, Hana 89, Yegaram ~40,000 (press reports)
  • β€’The upstream Autumn-27/ARTEX GitHub repo now returns 404
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·7h ago
πŸ”₯ Trending

Nanjing University LAMDA Unveils AgenticBBO-Bench: Benchmarking LLM Agents for Black-Box Optimization, Outperforming Top Numerical Solvers

Black-box optimization (BBO) poses foundational computational hurdles across chip design, database parameter tuning, hyperparameter optimization, and molecular engineering where objective function evaluations are prohibitively expensive and gradients are unavailable. Traditional Bayesian optimization and evolutionary algorithms rely strictly on numerical samples while discarding rich task semantics, whereas direct LLM prompting struggles with numerical precision and hallucinations. Researchers from Nanjing University's National Key Laboratory for Novel Software Technology (LAMDA Group) introduce AgenticBBO-Bench (arXiv:2610.12183), the first unified cross-domain benchmark for evaluating LLM agents in black-box optimization. Spanning synthetic functions, hyperparameter tuning, database configuration, chip layout, and molecular design under a standardized finite-budget protocol, Agentic BBO achieves superior family-averaged scores over direct LLM methods across all five domains and outperforms the strongest numerical optimizers in four. Evaluating seven frontier LLMs within the Codex agent harness establishes the Pareto frontier of performance and inference cost, with the benchmark and harness fully open-sourced.

⚑ Key Takeaways
  • β€’Nanjing University LAMDA releases AgenticBBO-Bench, benchmarking LLM agents across chip design, DB tuning, and molecular optimization
  • β€’Agentic BBO outperforms direct LLM prompting across all 5 domains and beats leading numerical optimizers in 4 out of 5 areas
  • β€’Maps Pareto frontier of frontier models within Codex harness, cutting required optimization steps by over 38% under finite budgets
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·7h ago
πŸ”₯ Trending

U-Space: TU Darmstadt Unveils Mechanistic Uncertainty Quantification in LLMs, Achieving Zero-Shot Token-Level Doubt Tracking Without Extra Sampling

Frontier language models generate incorrect conclusions with fluent, authoritative explanations, creating severe safety risks in autonomous systems where knowing when to defer to human review is essential. Traditional uncertainty quantification (UQ) methods rely either on compute-heavy repeated rollouts (e.g., semantic entropy across 10 to 20 samples) or supervised probes prone to confounding output length with true epistemic uncertainty. Researchers from TU Darmstadt and collaborating institutions introduce U-Space (arXiv:2610.09087), a mechanistic interpretability framework that uncovers evolving internal uncertainty within low-dimensional representations. By identifying semantic anchors of doubt and certainty in the unembedding matrix, U-Space constructs an orthogonal basis within the residual stream. The accompanying U-Lens projects intermediate token activations onto this basis, providing an interpretable, token-level uncertainty trajectory alongside an aggregated confidence score. Requiring zero training, zero labels, and zero repeated generations, U-Space outperforms established baselines under standard and length-controlled settings, with code fully open-sourced.

⚑ Key Takeaways
  • β€’TU Darmstadt open-sources U-Space, leveraging mechanistic interpretability to identify an orthogonal uncertainty subspace in the residual stream
  • β€’Eliminates labels, fine-tuning, and repeated sampling, reducing uncertainty estimation compute costs by over 90% compared to semantic entropy
  • β€’Robust against length confounding, providing real-time token-level doubt tracking that outperforms traditional probes on reasoning benchmarks
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·7h ago
πŸ”₯ Trending

HKUST Uncovers the Mechanistic Reality of Reasoning Distillation: On-Policy Distillation Teaches Compositional Skills, Not Factual Knowledge

The emergence of reasoning models such as OpenAI o1/o3 and DeepSeek-R1 sparked widespread efforts to distill dense reasoning traces into smaller 7B and 14B student models, unlocking striking performance leaps on math and coding benchmarks. However, a fundamental mechanistic question has lingered: does On-Policy Distillation (OPD) inject new parametric factual knowledge into the student, or does it merely impart compositional reasoning skills? Researchers from Hong Kong University of Science and Technology (HKUST) address this mystery in a landmark study (arXiv:2610.09639). Utilizing a controlled synthetic framework across four models from three distinct architectures, the authors independently isolate factual memory from multi-step compositional ability. The experiments establish that reverse-KL OPD reliably transfers compositional problem-solving skills across unseen task structures, but transfers minimal factual knowledge. Replacing reverse-KL with forward-KL restores factual memorization, while student rollouts specifically refine multi-step logical execution. These findings demonstrate that on-policy distillation does not expand a model's knowledge boundary, but rather teaches it to mobilize and organize facts it already possesses.

⚑ Key Takeaways
  • β€’HKUST demonstrates through controlled synthetic frameworks that on-policy distillation (OPD) transfers compositional skills but minimal factual knowledge
  • β€’Proves reverse-KL optimizes multi-step logical execution while forward-KL is required for parametric factual memorization
  • β€’Establishes a principled two-stage recipe for reasoning agents: forward-KL for factual ingestion, followed by reverse-KL for reasoning activation
Read details→
ADSponsored
B
Beethtbee_Β·9h ago
πŸš€ Release

Gemini API deprecates 3.7 Flash and silently routes it to 3.8 Flash; 3.5 Flash moves to 3.6, old Deep Research agent shuts down Oct 23

Google's Gemini API release notes (Oct 8 entry) deprecate gemini-3.7-flash, with all requests now auto-routed to gemini-3.8-flash; gemini-3.5-flash is likewise routed to gemini-3.6-flash. The deep-research-pro-preview-12-2025 agent shuts down on Oct 23 and must migrate to the 04-2026 versions. Developers began receiving email notices on Oct 10.

Gemini API deprecates 3.7 Flash and silently routes it to 3.8 Flash; 3.5 Flash moves to 3.6, old Deep Research agent shuts down Oct 23
⚑ Key Takeaways
  • β€’Two model IDs silently re-routed: gemini-3.7-flash β†’ gemini-3.8-flash, gemini-3.5-flash β†’ gemini-3.6-flash
  • β€’Same price: 3.8 Flash paid tier $0.75 in / $3.75 out per 1M tokens through 2026-12-31, then $1.50 / $7.50 from 2027-01-01
  • β€’deep-research-pro-preview-12-2025 shuts down 2026-10-23; migrate to deep-research-preview-04-2026 or deep-research-max-preview-04-2026
  • β€’Deprecations page lists 3.6 Flash with no shutdown date, so 3.7 was replaced before its predecessor
Read details→
D
Denodeno_landΒ·11h ago
πŸ”₯ Trending

Deno team joins Cloudflare: Deno runtime gets one more year of fixes, Deno Deploy shuts down in six months, celld folds into workerd for self-hosted agent infra

On Oct 9 Ryan Dahl announced the whole Deno team is joining Cloudflare. Deno runtime gets one more year of monthly bugfix/security releases, then official development ends (it stays open source); Deno Deploy runs six more months before shutting down, with migration help to Workers; JSR continues. Focus shifts to the Workers/Durable Objects model and merging self-hosted celld into workerd, pitched explicitly at agent harnesses.

Deno team joins Cloudflare: Deno runtime gets one more year of fixes, Deno Deploy shuts down in six months, celld folds into workerd for self-hosted agent infra
⚑ Key Takeaways
  • β€’Timeline: Deno runtime gets 12 months of monthly fixes then official development ends; Deno Deploy shuts down in 6 months with migration support for paying customers
  • β€’JSR keeps running on Cloudflare infra; rusty_v8 stays supported and heads into workerd
  • β€’celld (Apache-2.0, ~5.3k stars, v0.6.2 beta): a 58 MB static binary for self-hosted distributed Durable Objects backed by one S3-compatible bucket
  • β€’Vendor-reported celld numbers: 0.47 MB RAM per resident cell, 2,500 cells per 8 GB node, 0.2/0.3 ms p50/p99 stateless requests, ~20 s failover with zero lost acknowledged writes
  • β€’1,234 points on the Hacker News front page; Deno's announcement post drew 6.3k likes
Read details→
T
TypeSafe AItypesafeaiΒ·11h ago
πŸ”₯ Trending

Jev maker TypeSafe AI raises $870M Series A at $7.5B valuation led by a16z, less than a month after Jev's launch

TypeSafe AI announced an $870M Series A at a $7.5B valuation on Oct 9, led by a16z with Sequoia and existing investor DCVC; Martin Casado joins the board. Its 'System One' decision model Jev launched Sept 15 and returns calibrated, typed decisions instead of text. TypeSafe says a third of the Fortune 500 use it; a16z's post says 25% β€” the figures differ.

Jev maker TypeSafe AI raises $870M Series A at $7.5B valuation led by a16z, less than a month after Jev's launch
⚑ Key Takeaways
  • β€’$870M Series A at a $7.5B valuation, led by a16z with Sequoia and DCVC; Martin Casado joins the board
  • β€’Jev launched Sept 15, 2026 β€” roughly 3.5 weeks before the round (TechCrunch)
  • β€’a16z claims Jev generated 1 trillion tokens within 3 days of launch
  • β€’Enterprise adoption claims differ: TypeSafe says a third of the Fortune 500, a16z says 25% β€” both self-reported
  • β€’TypeSafe promises more machine-native models and enterprise features; docs at docs.typesafe.ai
Read details→
N
Nandakishor mNandakishorm1Β·13h ago
🌟 Open Source

Vega open-sources a physics-based typed decision model: frozen Qwen3.5-0.8B plus a 57 MB engine that rolls a ball into calibrated answers

Indie developer Nandakishor M open-sourced Vega (pip package vegaml, Apache 2.0): a frozen Qwen3.5-0.8B (4B also available) only reads, while a 57 MB trained engine models each allowed answer as a valley and rolls a damped ball into one, returning calibrated probabilities, conformal answer sets and an abstain flag. 73,728-token context and image input. The author's own paired tests beat hosted Jev 1.13.0 on tasks like phishing screening, while admitting it trails on reranking and multi-step reasoning.

Vega open-sources a physics-based typed decision model: frozen Qwen3.5-0.8B plus a 57 MB engine that rolls a ball into calibrated answers
⚑ Key Takeaways
  • β€’Size: 0.8B engine is 14.3M params / 57 MB with a 0.84M-param task adapter; the 4B engine is 124 MB
  • β€’Author-reported typed-decisions test (2,050 items): 0.763 (0.8B + adapter) and 0.803 (4B), at 22 ms / 28 ms per decision
  • β€’Phishing, 800 emails: Vega 75.4 acc, 252/400 caught vs Jev 1.13.0 61.9, 99/400 (author's paired run, McNemar p=3e-11)
  • β€’Median latency 267 ms on a T4 vs 591 ms hosted Jev; MMMU-Pro only 0.240, so not for university-level reasoning
  • β€’pip install vegaml; model card and a live Hugging Face Space are available
Read details→
ADSponsored
C
Cognition@cognitionΒ·13h ago
πŸ”₯ Trending

Devin now works with any personal ChatGPT plan: Go, Plus or Pro subscriptions cover GPT model usage instead of Devin quota

On Oct 9 Cognition said Devin can now link any personal ChatGPT plan (Go, Plus or Pro). The GPT share of each session is billed to the ChatGPT subscription instead of Devin quota or on-demand credits, across Devin Cloud, Desktop and CLI. It requires a self-serve Devin Pro, Max or Teams plan (not Free or Enterprise), and falls back to Devin quota when the ChatGPT allowance runs out.

Devin now works with any personal ChatGPT plan: Go, Plus or Pro subscriptions cover GPT model usage instead of Devin quota
⚑ Key Takeaways
  • β€’Eligible ChatGPT plans widened from Plus/Pro at the Sep 29 launch to Go, Plus and Pro; free ChatGPT accounts can't share usage
  • β€’Only self-serve Devin Pro ($20/mo), Max ($200/mo) and Teams ($80/mo minimum); not Free or Enterprise
  • β€’GPT usage bills only to ChatGPT, never double-counted against Devin quota; fast/priority variants excluded
  • β€’Vendor-cited Artificial Analysis Coding Agent Index v1.5: Fusion (GPT-6 Astra + SWE-2) scores 58.9 at $4.54/task vs Codex Astra (max) 61.6 at $7.47, 39% cheaper
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·13h ago
πŸ”₯ Trending

Ant Group Open-Sources Mara Chain: Rethinking Failure as a Stepping Stone for Autonomous Agent System Evolution

Modern AI agent systems and autonomous frameworks rely heavily on greedy search or reinforcement learning for self-evolution (such as harness synthesis, workflow routing, and prompt optimization). However, in non-convex and deceptive optimization landscapes, greedy pruning prematurely discards mutations that temporarily underperform, stranding systems in local optima. Researchers from Ant Group Research, Peking University, and Beijing Academy of Artificial Intelligence (BAAI) unveil Mara Chain (arXiv:2609.35855), an evolutionary optimization framework that rethinks failures as essential stepping stones. Mara Chain operates via three interconnected mechanisms: a stepping-stone archive combining semantic embeddings and execution traces for novelty preservation, a hierarchical mutator driven by backward failure attribution, and a Pareto-optimal non-dominated sorting selection module. On the AppWorld benchmark, Mara Chain achieves up to a 20.5% performance improvement over leading baselines (GEPA, ACE, SkillOpt-Lite) while requiring 65.5% fewer rollouts. On TerminalBench 2.1, it outstrips AHE and Meta-Harness by 20.2 and 22.5 percentage points respectively. The team has open-sourced the implementation under the AntOmniEvo repository.

⚑ Key Takeaways
  • β€’Ant Group, PKU, and BAAI unveil Mara Chain, rethinking failed mutations as stepping stones to bypass local optima in agent self-evolution
  • β€’Outperforms GEPA and ACE by up to 20.5% on AppWorld with 65.5% fewer rollouts, and beats AHE by 20.2 percentage points on TerminalBench 2.1
  • β€’Open-sources the complete AntOmniEvo framework with AST-aware stepping-stone buffering and hierarchical failure mutators
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·13h ago
πŸ”₯ Trending

REMORY: HKUST and Alibaba Introduce Residual Memory for Agent Context Compaction, Matching Full Context with Only 5.2% Token Overhead

As autonomous agents execute complex multi-step workflows across terminal consoles and web browsers, context window expansion and escalating token costs impose severe operational bottlenecks. Conventional compaction methods rely on text summarization or token pruning, both of which discard crucial state variables and induce semantic drift over extended sessions. Researchers from HKUST and Alibaba present REMORY (arXiv:2610.11287), introducing a residual memory paradigm designed for sequence-level context compaction. REMORY combines lossy natural language summaries with bounded sets of learned soft memory tokens, effectively forming a residual connection in the sequence dimension. While text summaries convey high-level narrative continuity, the soft memory tokens reconstruct lost fine-grained execution artifacts, API arguments, and intermediate environment states. On challenging benchmarks including BrowseComp and Terminal-Bench 2.1, REMORY retains 97% of full-context task performance while occupying only 5.2% of the original input token footprint, slashing redundant tool retries and loops by more than 40%.

⚑ Key Takeaways
  • β€’HKUST and Alibaba propose REMORY, establishing sequence-dimension residual memory connections to resolve agent context bloat
  • β€’Matches 97.2% of full-context capability using only 5.2% of token positions, slashing repetitive tool errors by 41.5% on Terminal-Bench
  • β€’Non-invasive plug-and-play architecture reduces TTFT latency by 73.8%, setting a new benchmark for long-horizon agent efficiency
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·13h ago
πŸ”₯ Trending

Incremental Open-Ended Deep Research (IOEDR): Tsinghua and RUC Introduce Structured Harness for Autonomous Research Agents, Slashing Searches by 61%

Autonomous deep research agents (such as OpenAI Deep Research and Perplexity) represent a powerful leap toward automated knowledge synthesis. However, on open-ended and comprehensive inquiries, existing agents rely primarily on unconstrained search-and-generate loops. This leads to three systemic failures: redundant query loops, conflicting claims across sections, and runaway token expansion. Researchers from Tsinghua University and Renmin University of China unveil IOEDR (arXiv:2610.11566), a framework for Incremental Open-Ended Deep Research with a Structured Harness. IOEDR models the evolving research report as a dynamic hierarchical state graph coupled with a persistent, deduplicated evidence pool. Rather than executing open-ended browsing, the agent uses structured harness constraints to incrementally plan sections, identify specific evidential gaps, and trigger precision queries with dynamic termination conditions. On DeepResearch Bench, IOEDR achieves superior report coherence and verifiable fact density while reducing total token consumption by 33.2% and slashing search API requests by 61.4%.

⚑ Key Takeaways
  • β€’Tsinghua and RUC introduce IOEDR, modeling deep research as an evolving hierarchical state graph coupled with a persistent evidence pool
  • β€’Slashes search API queries by 61.4% and token usage by 33.2% on DeepResearch Bench while eliminating cross-section factual drift
  • β€’Implements information-gain gating with a 92.4% noise rejection rate, establishing an efficient architectural pattern for deep research agents
Read details→
P
Prime Intellect@PrimeIntellectΒ·15h ago
πŸ”₯ Trending

Prime Intellect's Prime Agent used a 2,209-agent swarm to rewrite itself in Rust: time-to-type 738ms→56ms, 80%+ less memory

On Oct 9 Prime Intellect shipped a Rust rewrite of Prime Agent, its open-source coding agent harness. A root orchestrator ran 2,209 sub-agents across 10,000+ Prime Sandboxes and ~228.7B GLM-5.3 tokens over two weeks, with differential TUI, harness and protocol parity tests gating every merge. Vendor-measured cold time-to-type dropped from 737.8ms to 55.8ms and post-startup memory from 607MB to 106MB; Windows (beta) and Homebrew installs were added.

Prime Intellect's Prime Agent used a 2,209-agent swarm to rewrite itself in Rust: time-to-type 738ms→56ms, 80%+ less memory
⚑ Key Takeaways
  • β€’Scale: 2,209 agents, 16,758 agent-to-agent messages, 228.70B tokens (192.99B/1,981 agents for the port, 35.70B/228 for perf hillclimbing) across 10,000+ Prime Sandboxes
  • β€’Vendor-measured: cold time-to-type 737.8msβ†’55.8ms (13.22x), first paint 722.8msβ†’23.6ms, post-startup RSS 607.4MBβ†’106.0MB, install 172.1MBβ†’59.6MB
  • β€’Each task ran a Plannerβ†’Implementerβ†’adversarial Reviewer (different model)β†’Verifier state machine; failures loop back to the implementer
  • β€’A 3-day target-free hillclimb loop logged 144+ experiment/audit records and merged 69+ performance changes
  • β€’Codebase now 9 crates, largest file ~2,500 lines (was ~15,000); native Windows (beta) and Homebrew installs, still open source
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·16h ago
πŸ”₯ Trending

Xiaomi Unveils MiMo-V2.6: Scaling Reinforcement Learning Compute Up to 1M Context to Unlock End-to-End Self-Improvement

Reinforcement learning (RL) has emerged as the defining post-training paradigm driving frontier foundation models toward autonomous self-improvement, yet scaling RL across multi-step agent environments and million-token sequences introduces profound stability and infrastructure hurdles. Xiaomi's LLM-Core Team releases the comprehensive technical report for MiMo-V2.6 (arXiv:2610.11959), an omni-modal foundation model family that scales RL compute across three critical axes. First, it employs high-throughput asynchronous training processing 1,568 samples and 2.7 to 3.7 billion tokens per optimization step across context windows scaling up to 1M tokens. Second, it orchestrates training across diverse, multi-domain environments spanning code generation, computer vision, and cybersecurity under a heterogeneous mixture of agent harnesses. Third, it implements groupwise agentic grading to mitigate reward hacking and incentivize concise, token-efficient reasoning trajectories. By freezing the MoE router and establishing multi-layered defense barriers, MiMo-V2.6 guarantees rock-solid convergence at scale, while open-sourcing its RL framework, training dynamics, and interactive environments.

⚑ Key Takeaways
  • β€’Xiaomi releases MiMo-V2.6, scaling reinforcement learning compute across 1,568 samples, 3.7B tokens per step, and up to 1M context
  • β€’Introduces groupwise agentic grading to curb reward hacking and compress trajectory token lengths by over 35%
  • β€’Freezes MoE routers to eliminate routing collapse and open-sources the complete RL post-training infrastructure
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·16h ago
πŸ”₯ Trending

TestPrism: Nanjing University Exposes the Single-Reference Illusion in Coding Agent Test Generation, Slashing Pass Rates from 59.7% to 28%

Autonomous coding agents are increasingly leveraged to synthesize test suites for enterprise software. However, the standard practice of evaluating generated tests against a single reference solution overlooks alternative valid implementations and inflates perceived test quality. Researchers from Nanjing University and Hong Kong Polytechnic University unveil TestPrism (arXiv:2610.12289), exposing this systemic evaluation illusion. TestPrism comprises 300 test generation tasks sourced from 17 repositories and pairs them with 3,000 candidate implementations evenly divided into valid and invalid solutions. Its primary metric, the Joint Success Function, mandates that a test suite must simultaneously fail against buggy baselines, pass across all diverse valid implementations, and reject every invalid candidate. Evaluating fourteen frontier coding agent setups reveals that while models register a 59.67% success rate under legacy single-reference scoring, their actual Joint Success Function plunges to just 28.00%. To resolve these foundational flaws, the authors introduce TestHelix, an agentic synthesis harness leveraging peer cross-validation and recursive self-improvement to elevate Joint Success scores by up to 9.00 percentage points.

⚑ Key Takeaways
  • β€’Nanjing University introduces TestPrism, revealing single-reference evaluation artificially inflates coding agent test quality from 28.0% to 59.7%
  • β€’Evaluates 300 tasks against 3,000 candidate implementations (half valid, half invalid) under a strict Joint Success Function
  • β€’Presents TestHelix with peer cross-validation and recursive self-improvement, raising genuine test quality by up to 9.00 percentage points
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·16h ago
πŸ”₯ Trending

SparseEngine: Harbin Institute of Technology Open-Sources Sparse-First Inference Engine Delivering 2.5x Faster Decoding Than vLLM for Long-Horizon Agents

Long-horizon LLM agents accumulate expanding interaction histories that saturate GPU KV-cache memory and choke self-attention throughput. While sparse attention algorithms mitigate these bottlenecks mathematically, existing production serving engines (e.g. vLLM) are tightly coupled to dense tensor layouts, preventing heterogeneous sparse mechanisms from sharing infrastructure. Researchers from Harbin Institute of Technology (HIT) release SparseEngine (arXiv:2609.39068), an open-source, sparse-first inference engine built from the ground up for long-horizon agent workloads. SparseEngine establishes a shared lifecycle contract enabling 15 diverse sparse attention methods across four distinct families to control their custom KV representations while unifying memory management. It introduces Chain Cache to resume stateful KV-evicted contexts across multi-turn requests and Controllable Prefix-Cache Pruning to discard stale history without disrupting prefix matching. Empirical evaluations on long-horizon agent benchmarks show that SparseEngine delivers over 10x higher throughput under KV eviction, over 2.5x faster decoding throughput than vLLM at matched concurrency, and more than 2x end-to-end task speedups with zero degradation in model reasoning accuracy.

⚑ Key Takeaways
  • β€’Harbin Institute of Technology open-sources SparseEngine, a sparse-first inference engine natively supporting 15 sparse attention algorithms
  • β€’Introduces Chain Cache and Controllable Prefix-Cache Pruning to preserve prefix-cache benefits across multi-turn sparse agent sessions
  • β€’Delivers 2.5x faster decoding than vLLM, lifts throughput by 10x under KV eviction, and halves end-to-end agent task latency
Read details→
S
Satya Nadella@satyanadellaΒ·17h ago
πŸš€ Release

Microsoft launches Microsoft-Decision-1: a Qwen3.5-9B-based decision-scoring model at $0.042/M input tokens, free output, on Foundry and OpenRouter

On Oct 9 Microsoft released Microsoft-Decision-1, a decision-scoring model that returns calibrated probabilities over fixed options instead of generating text, aimed at routing, classification, verification, agent guardrails and AI judging. Microsoft reports top accuracy across 36 blind benchmarks, ~35x lower P50 latency than GPT-6 Sol and a 1.3% decision-flip rate under perturbation. Available on Microsoft Foundry and OpenRouter, 32K context, $0.042/M input, free output.

Microsoft launches Microsoft-Decision-1: a Qwen3.5-9B-based decision-scoring model at $0.042/M input tokens, free output, on Foundry and OpenRouter
⚑ Key Takeaways
  • β€’Pricing: $0.042/M input tokens, $0 output, 32,768-token context (OpenRouter)
  • β€’Vendor-reported: highest accuracy across 36 blind benchmarks (~150K questions); 2.5x faster than runner-up H2O-Lightning-4B v1.1 and 35x faster than GPT-6 Sol
  • β€’Robustness: 1.3% average decision flips across 8 perturbation types; zero flips when options are paraphrased, reversed or shuffled
  • β€’Base: post-trained from Qwen3.5-9B for single-pass scoring; Microsoft plans to rebase on MAI and OpenAI models
  • β€’Internal use: Xbox Research labeled 10K+ feedback items at near GPT-6 Sol quality, 14x+ faster and ~200x cheaper (vendor-reported)
Read details→
C
ClaudeDevs@ClaudeDevsΒ·21h ago
πŸš€ Release

Claude Code Projects opens to every Pro and Max user on the waitlist: one conversation coordinating parallel cloud threads

On Oct 9 ClaudeDevs said every Pro and Max user on the Claude Code Projects waitlist has been let in. A project is a long-running coordinator conversation that spins each task into its own thread, usually a parallel cloud session on its own branch that opens a PR when needed; threads can also run locally via Remote Control. Still public beta; not yet on Team or Enterprise.

⚑ Key Takeaways
  • β€’From Oct 9, every Pro/Max user on the waitlist has access; still public beta, not on Team/Enterprise.
  • β€’A project is one coordinator conversation plus N threads, each with its own context window; cloud threads work on their own branch and open PRs.
  • β€’Since Sep 23, threads can run on your own machine via Remote Control when a task needs local tools.
  • β€’New threads inherit project instructions and memory, each repo's CLAUDE.md and skills, and your claude.ai connectors.
  • β€’Projects draw on the same plan limits as other Claude Code sessions and use them faster; no benchmarks were published.
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·21h ago
πŸ”₯ Trending

Opera: A Verbal Critic Framework for Long-Horizon Coding Agents Tracks Diagnosed Flaws Until Resolution to Boost SWE-Bench by 15%

Long-horizon coding agents tackling complex software repositories depend on timely guidance, yet conventional LLM critics often degrade agent outcomes by misinterpreting in-progress exploratory commands or offering ephemeral advice that is never tracked. Researchers from Rutgers and Lehigh University unveil Opera (arXiv:2609.33987), an open-source verbal critic framework that treats diagnostic corrections as persistent notes tracked until root problems are definitively resolved. Opera orchestrates reviews through hybrid periodic and event-driven triggers, verifies proposed critiques against visible terminal artifacts, and monitors downstream trajectories to distinguish mere superficial compliance from genuine resolution. As a test-time intervention, Opera elevates baseline agent resolve rates by 12.4 percentage points on Terminal-Bench 2.1, 15.0 on SWE-Bench Pro, and 8.9 on DeepSWE v1.1 across four foundation backbones. Furthermore, fine-tuning Qwen3.5-9B on Opera-guided rollouts yields a 10.2 percentage point gain on held-out repositories without any test-time critic, preserving stability across heterogeneous execution harnesses (such as OpenHands to Terminus-2).

⚑ Key Takeaways
  • β€’Rutgers and Lehigh release Opera, introducing persistent diagnostic notes that track diagnosed faults until proven resolution in coding agents
  • β€’Elevates resolve rates by 12.4 points on Terminal-Bench 2.1 and 15.0 points on SWE-Bench Pro across multiple policy models
  • β€’Fine-tuning Qwen3.5-9B on Opera rollouts yields a 10.2 point gain on held-out repos without test-time critics, preserving cross-harness stability
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·21h ago
πŸ”₯ Trending

Memento 3: UCL and Huawei Noah's Ark Enable Recursive Self-Improvement for Frozen LLMs via Reflective Rulebooks, Clearing ARC-AGI-3

Endowing artificial agents with general intelligence requires inferring environment dynamics and refining world models continuously in the wild. Fine-tuning models directly is prohibitively expensive and precipitates catastrophic forgetting. Researchers from University College London (UCL) and Huawei Noah's Ark Lab UK present Memento 3 (arXiv:2610.11794), a framework enabling frozen LLM agents to achieve model-based Recursive Self-Improvement (RSI) purely via external memory. The agent maintains a natural-language rulebook representing its evolving hypotheses of environment physics, which is compiled into executable code for fast simulation and action planning. Governed by a continuous observation-reflection-revision-compilation-verification loop, the system leverages prediction errors to refine rules. Edits are merged only after passing LLM semantic alignment audits and deterministic cell-exact replay verification. On the rigorous ARC-AGI-3 abstraction benchmark, Memento 3 solves 100% of levels across all 25 public environments, scoring a perfect 100.0 Relative Human Action Efficiency (RHAE) while utilizing only 44% of human action counts. In Atari Pong, a synthesized feedback controller shuts out opponents 21:0 across three evaluation seeds with zero test-time LLM inference.

⚑ Key Takeaways
  • β€’UCL and Huawei Noah's Ark introduce Memento 3, unlocking model-based Recursive Self-Improvement for frozen LLMs via executable rulebooks
  • β€’Compiles natural-language physics hypotheses into code world models validated through cell-exact replay verification
  • β€’Clears all 25 games on ARC-AGI-3 using only 44% of human actions, and delivers 21:0 Pong shutouts with zero runtime LLM calls
Read details→