News · Leaderboard

Latest news, beside the leaderboard

Read what just changed, and see who is ahead. Those are the two doors on this page.

Plans
O
OpenAIOpenAI·1h ago
🔥 Trending

GPT-6 rolls out to all of ChatGPT with Intelligent UI: interactive charts and in-chat tools, Sol for paid tiers and Luna for free

On Oct 7 OpenAI began rolling GPT-6 out globally in ChatGPT together with Intelligent UI: the model decides when to answer with diagrams, interactive charts, forms, tappable buttons or small in-chat tools (calculators, games, bill splitters) instead of plain text. Plus, Pro, Business and Enterprise get it first on GPT-6 Sol; Go and free users follow on Oct 8 on GPT-6 Luna. OpenAI says GPT-6 can stream partial answers while still thinking, cutting wait times by 44%.

⚡ Key Takeaways
  • •Tiers: Plus / Pro / Business / Enterprise get GPT-6 Sol from Oct 7; Go and free get the more efficient GPT-6 Luna from Oct 8 (OpenAI announcement)
  • •Answers while thinking: partial answers stream before reasoning finishes; OpenAI says wait times drop 44% and GPT-6 beat GPT-5.6 on hard web searches in internal tests (The Decoder)
  • •Intelligent UI renders diagrams, interactive and editable charts, forms, tappable buttons and on-demand in-chat tools (savings calculator, retro game, bill splitter); users can dial visuals down (TechCrunch)
  • •Safety: OpenAI says GPT-6 resists attempts to bypass its safety training better (The Verge)
  • •Context: Google shipped a similar generative-interface feature for Gemini in May; ChatGPT now makes it a default for every user
Read details→
A
AnthropicAnthropicAI·3h ago
🔥 Trending

Anthropic launches Claude Haiku 5.5: $0.10/M input, 72.4% on OSWorld 2.1, about 75% cheaper than Haiku 4.5 on average

On Oct 7 Anthropic shipped Claude Haiku 5.5 (claude-haiku-5-5), completing the Claude 5.5 family. It targets high-volume, cost-sensitive work such as summaries, compaction, classification and coding subagents, and Anthropic calls it its fastest model at standard speed. Prompts up to 100K cost $0.10 input / $0.50 output per million tokens, and it is the first Haiku with effort levels. It scores 72.4% on the OSWorld 2.1 offline subset (Haiku 4.5: 15.7%) and 39.2% on Terminal-Bench 4.0. Sonnet 5.5 cache reads were cut 50% the same day, and Max/Team plans get monthly API credits.

Anthropic launches Claude Haiku 5.5: $0.10/M input, 72.4% on OSWorld 2.1, about 75% cheaper than Haiku 4.5 on average
⚡ Key Takeaways
  • •Pricing: prompts up to 100K cost $0.10 input / $0.50 output / $0.01 cache reads per million tokens ($0.50 / $2.50 above 100K) vs $1 / $5 for Haiku 4.5; Anthropic says about 75% cheaper on average (announcement)
  • •OSWorld 2.1 offline subset: 72.4% vs 15.7% (Haiku 4.5), 48.9% (GPT-6 Luna), 83.9% (Sonnet 5.5)
  • •Terminal-Bench 4.0: 39.2% (Haiku 4.5 0.0%, GPT-6 Luna 16.4%); FrontierCode 1.1 Main: 46.4% (Sonnet 5.5 xhigh 52.1%)
  • •First Haiku with low/medium/high/xhigh/max effort; Claude Code v2.1.293 makes it the default Haiku on the API with 1M context (release notes)
  • •Same day: Sonnet 5.5 cache reads cut from $0.20 to $0.10/M (about 20% cheaper on most agentic work); Max 5x / 20x / Team get $100 / $200 / up to $500 in monthly API credits
Read details→
X
X FreezeXFreeze·4h ago
🔥 Trending

FCC clears SpaceX Starlink Mobile next-gen D2D constellation: up to 15,000 VLEO sats, spectrum-lease waiver

On 2026-10-06 the FCC granted-in-part SpaceX’s D2D Constellation (S00735) under DA-26-1078: up to 15,000 sats at 326–335 km, with a §25.125 lease waiver plus PCS G and EchoStar-assigned bands. Milestone: 50% by 2032-10-07. ~150 Mbps/user is a company target, not an FCC measurement.

FCC clears SpaceX Starlink Mobile next-gen D2D constellation: up to 15,000 VLEO sats, spectrum-lease waiver
⚡ Key Takeaways
  • •Primary: FCC DA-26-1078 (released 2026-10-06) authorizes SpaceX D2D Constellation (S00735) for up to 15,000 NGSO satellites
  • •Orbits: nine shells at 326–335 km VLEO; US MSS+SCS, ex-US Direct-to-Cell, plus Ka/V/E/W feeder & TT&C
  • •Key waiver: §25.125(a)/(b) — SCS on AWS-3/AWS-H without a terrestrial spectrum lease; also PCS G Block + EchoStar-assigned bands (full ops after Step Two)
  • •Milestones: surety bond by 2026-11-04; 50% by 2032-10-07, remainder by 2035-10-07; 2020–2025 MHz and parts of ex-US bands deferred
  • •Speeds: no Mbps table in the Order; ~150 Mbps/user peak and ~650 sats/~4 Mbps V1 are company/press figures, not FCC measurements
Read details→
m
morlutomorluto·5h ago
🛠️ Tooling

REA goes viral: one MCP server lets Claude Code, Codex and Cursor reverse-engineer binaries, APKs and Electron apps (~13k GitHub stars)

REA (Reverse Engineer Anything, MIT) wraps Hopper, Ghidra and IDA Pro into an MCP server and CLI for coding agents: 41 native inspection tools and 14 investigation workflows across Mach-O/ELF/PE, .NET, Electron, websites and Android APKs, all running locally. Version 4.1.0 (Oct 6) adds headless-JADX APK analysis and Binwalk/Unblob firmware analysis. The repo has 12,962 stars and 1,375 forks.

REA goes viral: one MCP server lets Claude Code, Codex and Cursor reverse-engineer binaries, APKs and Electron apps (~13k GitHub stars)
⚡ Key Takeaways
  • •morluto/rea: 12,962 stars, 1,375 forks, MIT licence as of 2026-10-08
  • •Tool catalog: 41 native inspection tools, 14 investigation workflows, 11 browser-observation, 21 workspace/observation, 7 .NET, 5 APK tools
  • •4.1.0 (Oct 6) adds headless-JADX APK analysis, Binwalk/Unblob firmware, read-only IDA MCP providers, experimental Windows x64 Ghidra
  • •Works with 12 agents (Claude Code, Codex, Cursor, Gemini CLI, Windsurf, Devin, OpenCode, Antigravity, Copilot CLI and more) via npx rea-agents setup
  • •rea-agents npm: 1,291 downloads in the week of Sep 28–Oct 4; no official benchmark, but the DX-Ball showcase passes 3,205 original-x86 test cases
Read details→
ADSponsored
P
Perplexityperplexity_ai·5h ago
🔥 Trending

Perplexity open-sources pplx-embed-v2-late: 0.6B/9B multimodal ColBERT retrievers, OCR-free PDFs, small model queries the big model's index

On Oct 7 Perplexity released pplx-embed-v2-late, multimodal late-interaction (ColBERT) retrievers in 0.6B (340M active) and 9B (7.4B active) sizes, built on Qwen3.5 with bidirectional attention. Each token gets a 128-dim vector scored by MaxSim, covering text, images and rendered PDFs or slides without OCR. Both sizes share one embedding space, so the 0.6B model can query a 9B-built index. Public ViDoRe v3 image nDCG@10 is 62.3% / 65.2%; MIT-licensed weights are on Hugging Face.

⚡ Key Takeaways
  • •Public ViDoRe v3 nDCG@10: 0.6B 62.3% image / 61.2% markdown; 9B 65.2% image / 64.7% markdown (0.6B model card)
  • •Perplexity says the 0.6B model matches models with about 5x the active parameters on ViDoRe V3 (blog)
  • •Shared embedding space: index with 9B, query online with 0.6B to cut query-time cost
  • •Training: distilled from an internal 18B ColBERT teacher with a token-level LEAF-style objective; 9B fully fine-tunes the last 8 layers and uses LoRA elsewhere, including the vision encoder
  • •Setup: sentence-transformers >= 6.0.0 and transformers >= 5.4.0, MIT license; encode text-only and image-only batches separately (9B model card)
Read details→
P
Perplexityperplexity-ai·6h ago
🚀 Release

Perplexity open-sources pplx-decider-v1.1-27b: Decision Index rises from 56.4 to 61.56, beating Jev by 3.6 points

Perplexity released pplx-decider-v1.1-27b (Apache-2.0, fine-tuned from Qwen3.8-27B) on Hugging Face. It returns typed, probability-calibrated answers instead of text. Overall Decision Index is 61.56 vs 56.4 for v1 and 57.9 for Jev, and it is served via the Perplexity Decisions API.

Perplexity open-sources pplx-decider-v1.1-27b: Decision Index rises from 56.4 to 61.56, beating Jev by 3.6 points
⚡ Key Takeaways
  • •Overall Decision Index 61.56: +5.16 over v1 (56.4) and +3.66 over Jev (57.9) on the same Qwen3.8-27B backbone
  • •Categories: Language 69.45, Retrieval 61.26, Tools 78.88, Knowledge 48.18 (still below Jev's 51.4)
  • •Gains come mainly from lifting the causal mask on full-attention layers plus more training data incl. tasksource
  • •Decision head is a BF16 [255, 5120] readout, not a full-vocab lm_head; ~49 GiB of weights on a CUDA GPU
  • •Hosted via POST /v1/decisions with noul / choice / score questions, many per request (docs)
Read details→
E
Elon Muskelonmusk·7h ago
🚀 Release

Elon: Grok Bot will route each task to the best backend, including Claude Opus 5.5, Midjourney, Suno

On 2026-10-07 Elon Musk said Grok @Bot / @SpaceX will use the best backend model per task—naming Claude Opus 5.5, Midjourney, Suno and other leading APIs. Product context: [Introducing Grok Bot](https://x.ai/news/introducing-grok-bot) and [docs](https://docs.x.ai/grok-bot/overview). Routing rules, pricing, and timeline are still unspecified in official docs.

Elon: Grok Bot will route each task to the best backend, including Claude Opus 5.5, Midjourney, Suno
⚡ Key Takeaways
  • •Primary: @elonmusk 2026-10-07 — best backend per task incl. Opus 5.5, Midjourney, Suno
  • •Named backends: Claude Opus 5.5, Midjourney, Suno + other APIs; Grok models not said to be retired
  • •Product: persistent Bots + shared cloud computer — x.ai/news/introducing-grok-bot + docs.x.ai/grok-bot
  • •Unknown: routing map, user visibility/choice, third-party billing, data sharing with vendors
  • •No SWE/blind-eval scores in this announcement
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

UNREAL: NVIDIA Unifies Corpus-Scale Retrieval and 256K Long-Context Inference in a Single Frozen LLM with Under 500K Parameters

LLM systems handle evidence retrieval via two disjoint paradigms: external Retrieval-Augmented Generation (RAG) pipelines relying on auxiliary retriever-reranker stacks, and brute-force long-context inference suffering from quadratic attention overhead. NVIDIA and Technion introduce UNREAL (UNifying REtrieval And Long-Context with a Single Model, arXiv:2610.08463), a model-native framework unifying corpus retrieval and long-context filtering within a single frozen LLM. By deriving chunk representations and search queries directly from internal transformer activations, UNREAL introduces fewer than 500K parameters while keeping the backbone frozen. On a 3B-token, 21M-chunk Wikipedia corpus, UNREAL surpasses competitive retriever-reranker systems, lifting HotpotQA recall from 49.1% to 73.2% and 2WikiMultiHopQA from 31.7% to 60.1%. When applied to long-context sequences, UNREAL dynamically purges irrelevant distractors prior to generation, boosting 128K NoLiMa accuracy from 1.0% to 24.83% and 256K LV-Eval F1 to 54.66% while substantially curbing FLOPs and time-to-first-token (TTFT) from 32K context onward.

⚡ Key Takeaways
  • •NVIDIA and Technion present UNREAL, unifying corpus retrieval and long-context filtering in a single frozen LLM with under 500K parameters
  • •Surpasses SOTA retriever-reranker stacks on a 3B-token Wikipedia index, lifting HotpotQA recall from 49.1% to 73.2%
  • •Dynamic distractor pruning boosts 128K NoLiMa accuracy from 1.0% to 24.83% while significantly reducing TTFT and compute FLOPs
Read details→
ADSponsored
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

DAEDALUS: Bootstrapping Reusable Agent Memory via Self-Generated Tasks and Autonomous Dual-Agent Practice Without Human Oracles

LLM agents operating in novel API or terminal environments frequently repeat identical errors due to an absence of procedural memory. Existing memory bootstrapping approaches rely heavily on curated prompt playbooks or oracle verifiers that require extensive domain foresight. Illuin Technology releases DAEDALUS (arXiv:2610.08048), an autonomous dual-agent framework bootstrapping reusable memory entirely from self-generated practice without human oracles. DAEDALUS pairs an Explorer agent—which probes the execution environment to formulate curriculum-style tasks—with a Solver agent. When the solver stumbles, it extracts candidate heuristics that are committed into a persistent memory bank only after proving their efficacy across repeated verification trials. Across AppWorld, tau^2-bench, and AutomationBench, DAEDALUS boosts average success rates by up to 15.9 percentage points and raises Pass^5 rates by 2.2x over memory-less baselines, with generated procedural memories transferring seamlessly across heterogeneous foundation model families.

⚡ Key Takeaways
  • •Illuin Technology open-sources DAEDALUS, bootstrapping reusable procedural memory without human labels or oracle verifiers
  • •Employs an Explorer-Solver dual-agent self-play architecture to synthesize verified operational heuristics from failure traces
  • •Lifts task completion by up to 15.9 points and doubles Pass^5 by 2.2x across AppWorld, tau^2-bench, and AutomationBench
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

Harness-Aware Distillation: KAIST Focuses Agentic Model Distillation on Model-Unique Capabilities to Build Resilient SLM Agents

Deploying performant coding and OS agents on cost-efficient small language models (SLMs) is essential for enterprise scalability. In modern agent architectures, models operate within an execution harness that maintains workspace context, API state, and terminal feedback. Conventional distillation forces the student model to mimic the teacher's holistic sequence output, conflating harness-maintained facts with model reasoning and overloading the student's limited capacity. KAIST AI introduces Harness-Aware Distillation (HAD, arXiv:2610.02858). HAD decouples what the harness manages from the unique reasoning additions of the teacher model. By combining on-policy distillation with harness-conditioned action preference contrast and strict execution-record validity checks, HAD requires zero reward models or oracle success annotations. Across multiple long-horizon agent benchmarks, HAD-trained SLMs enter significantly fewer unproductive query loops and demonstrate substantially higher autonomous error-recovery rates.

⚡ Key Takeaways
  • •KAIST AI introduces Harness-Aware Distillation (HAD), optimizing small language model distillation for agent harnesses
  • •Decouples harness-managed state from model reasoning via action preference contrast and execution validation without reward labels
  • •Significantly suppresses unproductive infinite loops while doubling error-recovery rates, empowering 3B-8B SLMs for long-horizon agent tasks
Read details→
L
LiteLLM (BerriAI)@LiteLLM·11h ago
🚀 Release

LiteLLM open-sources Moyai, a self-hosted cloud coding agent: $101,872 on Devin vs ~$21,700 on Moyai over 31 days (79% cheaper)

On Oct 7 LiteLLM open-sourced Moyai (BerriAI/moyai), the self-hosted background coding agent its team runs daily: tasks from Slack or the browser run in isolated cloud workspaces that edit code, run tests and open PRs. Harnesses (Claude Agent SDK, Codex, Hermes, OpenCode, Deep Agents, Tool Loop) are switchable per session and inference routes through LiteLLM's 100+ providers. LiteLLM says replacing Devin cut a 31-day bill from $101,872 to about $21,700.

LiteLLM open-sources Moyai, a self-hosted cloud coding agent: $101,872 on Devin vs ~$21,700 on Moyai over 31 days (79% cheaper)
⚡ Key Takeaways
  • •Cost: $101,872 on Devin vs ~$21,700 on Moyai (~$700/day) over the same 31 days, a claimed 79% / ~$80,172 saving (launch post)
  • •Pluggable harnesses: Claude Agent SDK by default (prompt caching on), plus Hermes, Codex, OpenCode, Deep Agents and Tool Loop (docs)
  • •Inference via LiteLLM across 100+ providers; switch GPT-6 Astra / Claude Opus 5.5 / GLM-5.3 mid-session with per-teammate spend attribution
  • •Provider keys stay server-side and never reach the sandbox; each session gets its own terminal, filesystem and browser
  • •Caveat: ~58 stars at publish time and no LICENSE file detected by GitHub; confirm terms before commercial use
Read details→
H
Hugging Face Daily Papers@HuggingFace·13h ago
🔥 Trending

SafeActBench: NUS Investigates the Broken Evidence-to-Action Chain in Tool-Using Agents Across Six Operational Domains

Tool-using agents execute consequential modifications to external systems, yet nominal task success frequently conceals unverified actions executed without prior evidence. Researchers from the National University of Singapore (NUS) present a rigorous empirical study on the breakdowns along the evidence-to-action chain (arXiv:2610.07753). Evaluating ten model-harness configurations reveals a striking paradox: high static action assessment accuracy routinely masks brittle interactive execution. Failures overwhelmingly originate prior to execution: agents terminate prematurely before collecting requisite evidence, or initiate state-altering actions before foundational proofs are established. The authors formulate SafeActBench, spanning 656 rigorous cases across six operational domains and five protocols progressing from static judgment to complex multi-action workflows. Supported by a provenance-bound Evidence Ledger and deterministic trajectory evaluator, findings demonstrate that while single-step execution is stable once evidence is verified, multi-action workflows frequently collapse under unresolved prerequisites and truncated state transitions.

⚡ Key Takeaways
  • •NUS presents SafeActBench across 656 operational cases and 6 domains to evaluate where the evidence-to-action chain breaks
  • •Tests across 10 model-harness configurations show high static evaluation masks fragile execution, with over 70% of failures caused by premature actions
  • •Introduces the Provenance-bound Evidence Ledger, proving multi-action workflows suffer 64.7% failure rates from unresolved prerequisites
Read details→
H
Hugging Face Daily Papers@HuggingFace·13h ago
🔥 Trending

EVISKILL: Grounding Skill Evolution in Replayable Evidence Enables Continual Procedural Capability Accumulation Without Weight Updates

Continual skill evolution allows LLM agents to autonomously accumulate, refine, and reuse procedural knowledge across multi-turn interactions without updating model parameters. However, prevailing experience-driven reflection mechanisms suffer from two fundamental weaknesses: they decouple synthesized guidance from concrete behavioral traces, and they rely on monolithic global validation that blindly discards valid local corrections whenever an entire workflow fails. Researchers from Tsinghua University and partner institutions introduce EVISKILL (arXiv:2610.05030), an evidence-grounded framework anchoring procedural adaptation to deterministic execution traces. EVISKILL structures raw observations into Replayable Evidence Cards and synthesizes skill edits with explicit contextual lineage. A targeted replay engine verifies edits through isolated re-execution before cross-epoch refinement. Across three interactive benchmarks spanning six LLM backbones, EVISKILL consistently outperforms state-of-the-art reflective baselines, demonstrating durable procedural accumulation.

⚡ Key Takeaways
  • •Tsinghua University introduces EVISKILL, anchoring continual procedural skill evolution in Replayable Evidence Cards without parameter updates
  • •Decouples local edit verification from monolithic global validation, achieving a 3x increase in the retention of valid localized skills
  • •Validated across 3 interactive benchmarks and 6 foundation LLMs, consistently outperforming standard reflection architectures
Read details→
H
Hugging Face Daily Papers@HuggingFace·13h ago
🔥 Trending

NVIDIA Unveils NeMo-DCR: Bit-Exact Delta-Compressed Refit Slashes Trillion-Parameter Agentic RL Cross-Cluster Sync from 87.5 Min to 150s

Agentic reinforcement learning (RL) disaggregates centralized policy training from large-scale interactive environment rollouts, necessitating frequent cross-cluster policy weight synchronization (refit). Transferring a full 1T-parameter checkpoint across cloud regions currently consumes 87.5 minutes, idling massive rollout compute clusters. NVIDIA researchers introduce NeMo-DCR (Delta-Compressed Refit, arXiv:2610.08430), a bit-exact synchronization architecture that transmits only model deltas while guaranteeing receivers reconstruct identical parameter bits. Exploiting empirical measurements showing only ~1% of BF16 weights alter value bits per step, NeMo-DCR maps training shard changes into canonical coordinates via fixed affine projections and utilizes compressible XOR masks to guarantee exact bit-level parity. Streaming deltas over relay trees and committing changes via atomic joint commits, NeMo-DCR achieves 12x to 40x speedups for models spanning 30B to 1T parameters, slashing 1T cross-region refit latencies from 87.5 minutes to 150 seconds.

⚡ Key Takeaways
  • •NVIDIA introduces NeMo-DCR, delivering 12x to 40x faster weight synchronization for trillion-parameter agentic RL systems
  • •Capitalizes on 1% weight update sparsity in BF16 training using compressible XOR masks to guarantee bit-exact parity without drift
  • •Cuts 1T-parameter cross-region refit latency from 87.5 minutes to 150 seconds, unblocking massive multi-cluster agent training
Read details→
A
AtlassianAtlassian·14h ago
🚀 Release

Atlassian launches Agentic Multiplayer Protocol (AMP) and a rebuilt Atlassian MCP with 220+ tools, 15M+ daily calls and up to 25% fewer tokens

At Team '26 Europe, Atlassian introduced AMP, a framework giving agents an assigned identity, scoped authority, shared Teamwork Graph context and reviewable output, and made a rebuilt Atlassian MCP generally available with 220+ tools, OAuth 2.1/PKCE and tiered read/write/destructive scopes.

Atlassian launches Agentic Multiplayer Protocol (AMP) and a rebuilt Atlassian MCP with 220+ tools, 15M+ daily calls and up to 25% fewer tokens
⚡ Key Takeaways
  • •Atlassian MCP handles 15M+ calls/day; rebuilt from dozens of tools to 220+, with up to 25% token savings on the same Jira/Confluence work in internal Claude benchmarks
  • •Teamwork Graph (250B+ connections) now indexes Bitbucket/GitHub source down to functions, symbols and classes via Rovo Code Search, no clone needed
  • •Internal benchmark: agents grounded in Teamwork Graph were 44% more accurate with 48% fewer tokens
  • •Humans and agents already collaborate 10M+ times a month on Atlassian; Jira agent sessions link local and cloud agent runs back to work items
  • •MCP secured with OAuth 2.1 + PKCE by default, read/write/destructive scopes separated; EU-local AI inference added, non-human identity inventory coming soon
Read details→
G
Google DeepMind@GoogleDeepMind·15h ago
🚀 Release

Google officially ships Nano Banana 2.1 (gemini-nano-banana-2.1) GA: 1K/2K images now half the price of Nano Banana 2, and human-eval Elo 1050 beats Nano Banana Pro

On 2026-10-06 Google officially launched Nano Banana 2.1 and marked it GA in the Gemini API changelog as gemini-nano-banana-2.1. It is based on Gemini 3.6 Flash and is rolling out across the Gemini app, AI Studio, Search AI Mode, Flow, Stitch, Google Ads and Gemini Enterprise Agent Platform. Image output is priced at $30 per 1M tokens (Nano Banana 2: $60), which works out to $0.0336 per 1K image and $0.0504 per 2K image, about half the old price. The model card's side-by-side human eval puts text-to-image Elo at 1050, ahead of Nano Banana 2 (990) and Nano Banana Pro (935). gemini-3.1-flash-image is now deprecated, with no shutdown date announced. This confirms the silent Flow rollout we reported earlier.

Google officially ships Nano Banana 2.1 (gemini-nano-banana-2.1) GA: 1K/2K images now half the price of Nano Banana 2, and human-eval Elo 1050 beats Nano Banana Pro
⚡ Key Takeaways
  • •Per-image price: 1K $0.0336 (was $0.067, -50%), 2K $0.0504 (was $0.101, -50%), 4K $0.113 (was $0.151, -25%); Batch halves it again to $0.0168 per 1K image (pricing)
  • •Text side got pricier: input $1.50/1M tokens (was $0.50), text+thinking output $7.50 (was $3), so re-check costs for long prompts and many reference images
  • •Human-eval Elo (model card): T2I overall 1050 vs NB2 990 vs Pro 935; multi-character consistency 1106 vs 978 vs 1011; infographic factuality 0.521 vs 0.179 vs 0.265
  • •Specs: up to 14 reference images (4 characters + 10 objects), fixed 2K/4K tiling artifacts at 1:8/8:1, Google Web + Image Search grounding, thinking levels minimal/medium/high
  • •Migration: gemini-3.1-flash-image is deprecated (no shutdown date yet); 131,072-token input limit; no function calling, structured outputs or caching (model docs)
Read details→
G
GitHub@github·15h ago
🚀 Release

GitHub stacked pull requests are generally available on all github.com plans: approvals survive rebases, whole stacks land via merge queue, and gh stack ships an AI-agent skill

On 2026-10-06 GitHub made stacked pull requests generally available on all github.com plans, with GitHub Enterprise Server support coming in a later release. Since the public preview, repos using stacks have merged 9% more code than their peers, and more than two-thirds of the top 1% of repos now use them, with a 5% faster time-to-merge. New in GA: approvals are kept for unchanged code after a rebase, replacement commits are signed automatically, a whole stack enters the merge queue as one merge group, stacks retarget automatically when their base branch is deleted, and stack auto-merge is rolling out over the coming weeks. The gh-stack CLI extension (v0.2.0) now supports Git worktrees, and `gh skill install github/gh-stack` teaches AI coding agents how to split large changes into stacked PRs.

GitHub stacked pull requests are generally available on all github.com plans: approvals survive rebases, whole stacks land via merge queue, and gh stack ships an AI-agent skill
⚡ Key Takeaways
  • •Impact (changelog): repos using stacks merge 9% more code; over 2/3 of the top 1% of repos use them, with 5% faster time-to-merge
  • •Approvals and signing: Rebase stack keeps approvals on unchanged code even when stale approvals are dismissed; replacement commits are signed and keep the original authorship
  • •Merge queue: a stack lands as 1 merge group; with the merge-commit method, each PR gets its own merge commit; stack auto-merge rolls out over the next few weeks
  • •Agents: gh-stack v0.2.0 (2026-10-02, 1.6k stars) supports worktrees and installs an agent skill via gh skill install github/gh-stack; the pull_request webhook gains a stacked action
  • •Coverage: all github.com plans now, GHES in an upcoming release; requires gh v2.0+ and Git 2.36+
Read details→
A
AnthropicAnthropicAI·17h ago
🔥 Trending

Anthropic expands its Cyber Verification Program into three access tiers with Opus 5.5, Sonnet 5.5 and Mythos 5.1; Red Team tier completes 34/50 CyScenarioBench tasks with zero blocks

On 2026-10-06 Anthropic merged Project Glasswing and the original Cyber Verification Program into one three-tier program: Defense Access (SOC, incident response, malware reversing, vuln validation), Red Team Access (authorized pentesting) and Specialized Access (safety-critical systems such as power grids, flight systems and telecom, vetted with the US government). All tiers get Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1 with reduced cyber blocking. On CyScenarioBench, GA models were blocked on the first prompt of every task, Defense tier blocked 46/50 trials, and Red Team tier had zero blocks and completed 34/50 (matching the 67.6% no-safeguard rate). Glasswing partners reported at least 129,000 verified vulnerabilities from April to July 2026.

Anthropic expands its Cyber Verification Program into three access tiers with Opus 5.5, Sonnet 5.5 and Mythos 5.1; Red Team tier completes 34/50 CyScenarioBench tasks with zero blocks
⚡ Key Takeaways
  • •3 tiers: Defense (decisions in days), Red Team (weeks, organizations only), Specialized (in-depth review with the US government); Glasswing members move to Specialized automatically (announcement)
  • •3 models in every tier: Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1, plus future models
  • •CyScenarioBench (10 tasks x 5 attempts): GA blocked on first prompt; Defense tier blocked 46/50; Red Team tier 0 blocks and 34/50 completed, matching the 67.6% no-safeguard rate
  • •Impact: at least 129,000 verified vulns found by Glasswing partners (Apr-Jul 2026) plus 5,500 from Anthropic's own OSS scanning (Apr-Oct), over 33,000 rated critical/high (Project Glasswing)
  • •Available on Claude Platform, Vertex AI and Microsoft Foundry; Bedrock only for Enterprise Frontier Safeguards customers; data retention required
Read details→
F
FeSensFeSens·17h ago
🔥 Trending

openTPU, an open-source AI accelerator developed by AI agents, goes viral: full RTL/ISA/compiler stack runs Qwen3.5, Gemma 4 and a 35B MoE on a Kintex-7 FPGA card at up to 85.8 tok/s

FeSens/openTPU (Apache-2.0) hit the Hacker News front page (250+ points, 300+ comments). Reusing the AI-driven method from auto-arch-tournament (AI-designed RISC-V cores), the author had AI agents iterate on an inference accelerator; one repo holds SystemVerilog RTL, the ISA, a bit-exact simulator, a kernel language/compiler and a profiler. On a Kintex-7 xc7k480t PCIe card it runs Qwen3, Qwen3.5, Gemma 4, Phi-4-mini and others with real weights, token-for-token identical to the simulator; LFM2.5-230M decodes at 85.8 tok/s in 4-bit and Qwen3.5-35B-A3B reaches 3.95 tok/s with host-streamed experts.

openTPU, an open-source AI accelerator developed by AI agents, goes viral: full RTL/ISA/compiler stack runs Qwen3.5, Gemma 4 and a 35B MoE on a Kintex-7 FPGA card at up to 85.8 tok/s
⚡ Key Takeaways
  • •Full stack in one repo: SystemVerilog RTL, 8x32-bit-word ISA, bit-exact Python simulator, @ol.jit kernel compiler and Lens profiler, Apache-2.0 (GitHub)
  • •Measured device decode (4-bit, int8 head): LFM2.5-230M 85.8 tok/s, Qwen3-0.6B 31.3, Qwen3.5-2B 12.09, Gemma 4 E2B 12.14 tok/s
  • •Uses 82-94% of the 17.1 GB/s DDR3-1066 peak while decoding; 133.33 MHz clock, four-column systolic array
  • •MoE offload: Qwen3.5-35B-A3B at 3.95 tok/s (153 MB of experts streamed per token, 62% slot hit), LFM2.5-8B-A1B at 10.6 tok/s (offload docs)
  • •4-bit weights (FP4 with two-level block scales, 4.25 bits/weight) decode 40-45% faster than int8, with per-model perplexity cost documented
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

HERMES: Harness Engineering via Modular Executable Dev-Primitives Empowers Software Engineering Agents to Match Frontier Models

Large Language Model (LLM) coding agents equipped with bash terminals frequently collapse when deployed across massive enterprise codebases. Centralized agents struggle to reconstruct project states fragmented across multi-language source trees, build configs, and test suites, precipitating context window exhaustion and catastrophic semantic drift. Researchers from UIUC introduce Dev-Primitives and the HERMES harness engineering framework (arXiv:2610.07832). Dev-Primitives transform passive software components into active agentic nodes by pairing each repository artifact with a resident LLM endowed with natural-language reasoning, inter-component communication channels, and localized self-modification primitives. Coordinated via a dependency-aware dynamic activation scheduler and execution-guided fault localization, HERMES boosts task performance across four standard software engineering benchmarks by 12.4% over baseline harnesses. Crucially, powering HERMES with lightweight Qwen3-8B Dev-Primitives brings performance within 4.5% of homogeneous GPT-5.6 Sol configurations while slashing inference token overhead by 26.2% on Terminal-Bench 4.0.

⚡ Key Takeaways
  • •UIUC introduces Dev-Primitives and the HERMES framework, transforming passive software code artifacts into active LLM-powered nodes
  • •Dependency-aware dynamic activation and execution-guided bug diagnosis eliminate context explosion across multi-thousand-file repositories
  • •Outperforms baseline harnesses by 12.4% on average; lightweight Qwen3-8B primitives match GPT-5.6 Sol within 4.5% while cutting costs by 26.2%
Read details→