⚡Trending:
I
InclusionAI / Ant Group@AntLingAGI·19m ago
🚀 Release

InclusionAI Ling-3.1-flash: 560B/25B-active MoE; vendor SWE-Pro 65.39 & CyberGym 87.9; free on Vercel/OpenCode through Oct 13

On [2026-09-30](https://technode.com/2026/09/30/ant-group-launches-ling-3-1-flash-with-560-billion-parameters/) InclusionAI (Ant Group) announced **Ling-3.1-flash**: ~**560B** total / ~**25B** active MoE hybrid-reasoning model, designed for up to **1M** context (trial/gateway routes often expose **256K–262K**). Weights are **not** on Hugging Face yet. [Vercel AI Gateway](https://vercel.com/changelog/ling-3-1-flash-is-now-available-on-ai-gateway) and [OpenCode Zen](https://opencode.ai/docs/zen/) list limited-time free access (~through **2026-10-13**). Vendor-reported scores include SWE-Pro **65.39**, Terminal-Bench 4.0 **40.40**, CyberGym **87.90**—not independently reproduced.

InclusionAI Ling-3.1-flash: 560B/25B-active MoE; vendor SWE-Pro 65.39 & CyberGym 87.9; free on Vercel/OpenCode through Oct 13
⚡ Key Takeaways
  • •Scale: ~560B total / 25B active (vs Ling-3.0-flash 124B/5.1B); design target 1M context; trial/gateway often 256K–262K
  • •Architecture (vendor summary): 7×KDA : 1×Gated MLA; 512 routed experts pick 8 + 1 shared
  • •Vendor benches (unreproduced): SWE-Pro 65.39, TB4.0 40.40, CyberGym 87.90, BrowseComp 91.67
  • •Access: Vercel AI Gateway inclusionai/ling-3.1-flash(-free) via Novita free ~to Oct 13; OpenCode Zen ling-3.1-flash-free
  • •Caveats: no public weights yet; free routes may use data to improve the model; some scores are environment-qualified
Read details→
A
Anthropic@AnthropicAI·19m ago
🔥 Trending

Anthropic commits $100M to Claude Frontier Academy: train 10,000 Frontier Deployed Engineers by end of 2027

On [2026-10-02](https://www.anthropic.com/news/claude-frontier-academy) Anthropic launched **Claude Frontier Academy** with a **$100M** commitment to train **10,000** enterprise **Frontier Deployed Engineers (FDEs)** by end of **2027**. First cohorts include Accenture, Bain, Capgemini, CBA, Deloitte, McKinsey, Morgan Stanley, and Novo Nordisk. Path: multi-day in-person simulated deployment → **Claude Resident Engineer** badge → **12-week** residency on a named Claude use case → **Claude Frontier Deployed Engineer** badge (first expected early 2027). Nomination-only—not a public self-serve course.

Anthropic commits $100M to Claude Frontier Academy: train 10,000 Frontier Deployed Engineers by end of 2027
⚡ Key Takeaways
  • •Commitment: $100M to train 10,000 FDEs by end of 2027
  • •Path: in-person sim → Resident badge → 12-week residency → Frontier Deployed Engineer badge
  • •First cohorts: Accenture, Bain, Capgemini, CBA, Deloitte, McKinsey, Morgan Stanley, Novo Nordisk; SF/NY/London
  • •Not a new model API—enterprise talent program; nomination via Anthropic account teams
  • •Builds on Partner Network stats cited in the post (46k firms / 175k+ certs / ~4k Basecamp)
Read details→
K
Kyle Wiggers / Ai2@allen_ai·2h ago
🛠️ Tooling

Ai2 ships Olmo-core 3: open trillion-scale MoE training stack; 858 TFLOP/s/GPU at 1.2T on B300

On [2026-10-01](https://allenai.org/blog/olmocore3) Ai2 released **Olmo-core 3**, an open MoE training stack ([GitHub](https://github.com/allenai/OLMo-core), PyPI `ai2-olmo-core`). It moves from FSDP to **DDP + expert parallelism**; on 8×B300 a 47B MoE hits ~**2.7×** throughput (52k vs 19.4k tok/s/GPU). Systems benches include **1.2T** on 512 GPUs at up to **858** useful-model TFLOP/s/GPU and a DeepEP v2 capacity probe at **2.38T**. Report: [olmocore3](https://allenai.org/papers/olmocore3).

Ai2 ships Olmo-core 3: open trillion-scale MoE training stack; 858 TFLOP/s/GPU at 1.2T on B300
⚡ Key Takeaways
  • •Announced 2026-10-01: allenai.org/blog/olmocore3 + HF mirror + github.com/allenai/OLMo-core
  • •Stack: EP/PP/distributed optimizer; NVSHMEM rowwise EP + GPU-resident routing + grouped GEMM; optional DeepEP v2
  • •Throughput: expert pool 8→128 at ~3.2B active, <5% drop; 47B MoE on 8×B300 52k vs 19.4k tok/s/GPU (~2.7×)
  • •Scale: 1.2T / 58.36B active / 512 GPUs up to 858 TFLOP/s/GPU (random-routing systems ref); DeepEP v2 probe 2.38T
  • •MXFP8 ~+21% vs BF16, peak active mem 103→95 GiB; topology-agnostic checkpoints; next Olmo will be MoE
Read details→
C
Cohere Team@Cohere·2h ago
🔥 Trending

Cohere Embed 5 ships: Pro/Fast shared space; ViDoRe V3 Pro 85.8 avg; Fast at $0.08/M tokens

On [2026-09-30](https://cohere.com/blog/embed-5) Cohere launched **Embed 5** with **Pro** and **Fast** tiers that share one embedding space (index with Pro, query with Fast). Multimodal text/image/fused inputs, **128K** context, 100+ languages; API ids `embed-v5.0-pro` / `embed-v5.0-fast`. Blog: ViDoRe V3 Pro avg **85.8** (+8.8 vs Embed 4); text pricing **$0.12 / $0.08** per M tokens. Available on Cohere API, Model Vault, Azure Foundry, SageMaker; vLLM for private deploy.

Cohere Embed 5 ships: Pro/Fast shared space; ViDoRe V3 Pro 85.8 avg; Fast at $0.08/M tokens
⚡ Key Takeaways
  • •Shipped 2026-09-30: cohere.com/blog/embed-5 + docs.cohere.com/docs/cohere-embed
  • •Shared space: index Pro / query Fast; cross-model mean loss ~1.6–2.7% vs same-model on 40 dev sets
  • •ViDoRe V3: Pro 85.8 (+8.8 vs Embed 4) vs Voyage 4 Large 83.7 / Gemini Embedding 2 83.2; Fast 84.5
  • •128K ctx; dims 2048…256; float/int8/binary + Matryoshka; text $0.12/$0.08 per M (image $0.40/M both)
  • •API embed-v5.0-pro/fast; Foundry/SageMaker/vLLM; LangChain/Weaviate/Qdrant/Pinecone integrations
Read details→
ADSponsored
G
Gabriel Massadas / Nelson Duarte / Ashish Vinodkumar / Cloudflare@Cloudflare·4h ago
🛠️ Tooling

Cloudflare AI Search goes GA: native multimodal embeddings + PDF OCR; billing from Nov 1; hybrid search default

On [2026-10-01](https://blog.cloudflare.com/ai-search-ga/) Cloudflare made **AI Search generally available**: managed index/retrieval (Workers AI + Vectorize + R2 + Browser Run) ships native multimodal embeddings ([Qwen3-VL-Embedding](https://developers.cloudflare.com/ai-search/configuration/models/supported-models/)), OCR for scanned PDFs, 10 MiB text/OCR PDF limits, hybrid search by default, and [billing from 2026-11-01](https://developers.cloudflare.com/ai-search/platform/limits-pricing/) with a free allotment.

Cloudflare AI Search goes GA: native multimodal embeddings + PDF OCR; billing from Nov 1; hybrid search default
⚡ Key Takeaways
  • •GA 2026-10-01: changelog + blog.cloudflare.com/ai-search-ga/
  • •Multimodal embeddings: @cf/qwen/qwen3-vl-embedding-2b and google-ai-studio/gemini-embedding-2; image queries via REST/public endpoint
  • •OCR on all accounts; text/code + OCR PDFs up to 10 MiB (non-OCR PDFs stay 4 MiB)
  • •Hybrid search default; Workers AI embed/rerank folded into AI Search pricing (not Workers AI bill)
  • •Pricing: $0.75/1M ingest tokens, $2/GB-mo storage, $0.75/1k semantic, $0.10/1k full-text; free tier 5M ingest / 10 GB / 1k+1k queries; bill from Nov 1, 2026
Read details→
K
Kyle Wiggers / Ai2@allen_ai·4h ago
🔥 Trending

Ai2 open-sources AstaBrief 8B: Fast cited scientific reports; ~51s end-to-end vs Thinking ~178s

On [2026-10-02](https://huggingface.co/blog/allenai/astabrief) Ai2 open-sourced **AstaBrief 8B** (Qwen3-8B base): one-pass cited scientific reports from a question + retrieved snippets. Powers Asta **Fast** mode vs Claude Thinking. End-to-end report time averages **51.1s** vs **178.5s** (~3.5×). Weights: [allenai/AstaBrief_8B](https://huggingface.co/allenai/AstaBrief_8B) (updated 2026-10-02).

Ai2 open-sources AstaBrief 8B: Fast cited scientific reports; ~51s end-to-end vs Thinking ~178s
⚡ Key Takeaways
  • •Announced 2026-10-02: HF blog + https://huggingface.co/allenai/AstaBrief_8B
  • •Training: ~90K filtered real Asta queries → ~47K SFT; ~6K dual-judge-agree DPO; citation-density filtering
  • •System: one-pass full report (skip Thinking’s snippet summarize/cluster + section-by-section writing)
  • •Latency: Fast ~51.1s/report vs Thinking ~178.5s (~3.5×); authors did not re-run full eval vs today’s frontier
  • •Use in Asta Fast mode; self-host weights; example local PDF workflow + open training data
Read details→
M
Michelle Chen / Sam Else / Gabriel Massadas / Cloudflare@Cloudflare·6h ago
🛠️ Tooling

Cloudflare Web Search API open beta: Ceramic/Exa/Linkup via AI Gateway for structured agent grounding

On [2026-10-02](https://blog.cloudflare.com/introducing-web-search-api/) Cloudflare launched **Web Search API** (open beta): agents query the live web through [AI Gateway](https://developers.cloudflare.com/ai-gateway/) and get structured titles/URLs/snippets. Default [Ceramic.ai](https://developers.cloudflare.com/web-search/providers/) ($0.25/1k, ZDR); Exa ($7) and Linkup ($5) also available. Partners meet Verified-bot rules; REST, Workers `env.AI.websearch()`, BYOK, list pricing with no markup.

Cloudflare Web Search API open beta: Ceramic/Exa/Linkup via AI Gateway for structured agent grounding
⚡ Key Takeaways
  • •Announced 2026-10-02; docs https://developers.cloudflare.com/web-search/ (open beta)
  • •Providers: ceramic default $0.25/1k ZDR; exa $7/1k with highlights; linkup $5/1k fast mode
  • •Surfaces: REST /ai/websearch/ and Workers env.AI.websearch(); AI Gateway credits, no markup; BYOK supported
  • •Crawlers meet Verified-bot rules and return source links; Server Tools for web search coming
  • •Replaces guess-then-curl URL fetches for live grounding (e.g. Birthday Week docs)
Read details→
M
Marc Brooker / Mike Chambers / Fabio Nonato de Paula / AWS Strands@awscloud·6h ago
🔥 Trending

AWS Strands Decider 2B open-sourced: Qwen3.5-2B decision head, ~72% JevBench, ~115ms local routing/tool gates

On [2026-10-01](https://strandsagents.com/blog/introducing-strands-decider/) AWS Strands Labs released **Strands Decider 2B** ([`strands-labs/strands-decider`](https://github.com/strands-labs/strands-decider), Apache-2.0; weights [HF hobson-v19](https://huggingface.co/StrandsAgents/strands-decider-2B-hobson-v19)): LM head replaced by an option-scoring head. JevBench public **167/231 ≈72.3%**, Brier ~0.35; ~**115ms** median on RTX 3090. `pip install strands-decider` for routing, tool choice, and before_tool_call gates.

AWS Strands Decider 2B open-sourced: Qwen3.5-2B decision head, ~72% JevBench, ~115ms local routing/tool gates
⚡ Key Takeaways
  • •Blog 2026-10-01; github.com/strands-labs/strands-decider Apache-2.0 (~140★); HF StrandsAgents/strands-decider-2B-hobson-v19
  • •Arch: Qwen3.5-2B + LoRA r16 + ~1M pointer head; noul/choice/score + confidence (no free text)
  • •JevBench public 167/231 (~72.3%), Brier ~0.348; 3rd of 33 in 2B class per authors
  • •Latency ~115ms median RTX 3090; ~153ms small tasks on M3 MacBook
  • •Try: pip install strands-decider; ask/serve; wire into Strands before_tool_call interventions
Read details→
ADSponsored
P
Phillip Jones / Dan Carter / Cloudflare@Cloudflare·8h ago
🛠️ Tooling

Cloudflare OS open-sourced: company agent workspace with Gatekeeper capability security (10k+ GitHub stars)

On [2026-10-01](https://blog.cloudflare.com/cloudflare-os/) Cloudflare open-sourced **Cloudflare OS** ([github.com/cloudflare/cloudflare-os](https://github.com/cloudflare/cloudflare-os), Apache-2.0): a browser agent workspace with company context/skills, modifiable Gadgets, and a **Gatekeeper** capability-security layer (zero default access, observation-following policy, deferred simulated approvals). Deploy via [starter](https://github.com/cloudflare/cloudflare-os-starter) or https://os.cloudflare.app/deploy; inference through AI Gateway.

Cloudflare OS open-sourced: company agent workspace with Gatekeeper capability security (10k+ GitHub stars)
⚡ Key Takeaways
  • •Announced 2026-10-01; github.com/cloudflare/cloudflare-os Apache-2.0; >10k stars near launch (GitHub API)
  • •Trio: context-grounded agent workspace + Gatekeepers + per-user Gadgets (Dynamic Worker Facets + SQLite)
  • •Security: zero default access; credentials isolated; egress only via bindings; observation log follows sharing/egress
  • •Gatekeepers can simulate side effects and defer bulk approval vs sync HITL or auto-approve
  • •Try: pnpm run-local, os.cloudflare.app/deploy, cloudflare-os-starter; AI Gateway for models/budgets; early access
Read details→
T
Travis Beals / Google@Google·12h ago
🚀 Release

Google Falcon 9 Suncatcher: testing AI/TPU radiative cooling in vacuum

On 2026-10-01 Google’s [Project Suncatcher](https://blog.google/innovation-and-ai/models-and-research/google-research/project-suncatcher-prototype/) prototype, built with Planet, reached orbit on SpaceX Transporter-18 (Falcon 9). Contact confirmed. Focus: TPU survival plus **radiative/heat-pipe cooling in vacuum** (no airflow), plus radiation/launch stress data. Background: [Facts](https://blog.google/innovation-and-ai/models-and-research/google-research/google-project-suncatcher-facts/).

Google Falcon 9 Suncatcher: testing AI/TPU radiative cooling in vacuum
⚡ Key Takeaways
  • •Orbit: 2026-10-01 Transporter-18 / Falcon 9; Planet-built prototype; contact OK (Google blog)
  • •Cooling: no airflow in vacuum; heat pipes + radiators; thermal-vac tested (Facts)
  • •Payload: Google TPUs; Trillium TID >~5-year mission dose in proton-beam tests
  • •Press (not blog numbers): ~4 TPUs, ~1 kW solar, ~15 min duty cycles (Ars/IEEE)
  • •Next: 2027 dual-sat laser ISL; Research Blog + Joule for system design
Read details→
A
AWS News Blog@awscloud·12h ago
🛠️ Tooling

AWS Well-Architected Agent public preview: goal-aligned cloud optimization with IaC/CLI/console remediation packs

On [2026-10-01](https://aws.amazon.com/blogs/aws/announcing-aws-well-architected-agent-an-ai-powered-intelligence-to-optimize-your-cloud-environment-preview/) AWS announced public preview of the **Well-Architected Agent**: it correlates metrics, configs, and topology against Well-Architected practices across 65+ services, prioritizes by declared business goals, and ships console / updated IaC (Terraform, CDK, CloudFormation) / CLI remediation. Available in us-east-1, us-east-2, us-west-2; delivered via AWS Support.

AWS Well-Architected Agent public preview: goal-aligned cloud optimization with IaC/CLI/console remediation packs
⚡ Key Takeaways
  • •Launch: AWS News Blog 2026-10-01T20:04:04Z + What's New
  • •Goal-aligned three-level findings with console / updated IaC / CLI packs
  • •Cost/security/performance/resilience; 65+ services; pre-deploy IaC review (TF/CFN/CDK)
  • •Agent regions: us-east-1/2, us-west-2; needs AWS Support; ~24h after profile
  • •Docs: https://docs.aws.amazon.com/wellarchitected/latest/userguide/agent.html — evaluate GenAI output before applying
Read details→
N
NVIDIA Technical Blog@nvidia·12h ago
🛠️ Tooling

NVIDIA DOCA Agent Skills on GitHub: verified API/hardware contracts lift 65-prompt checklist pass from 19% to 100%

On [2026-10-01](https://developer.nvidia.com/blog/build-applications-on-nvidia-bluefield-faster-with-nvidia-doca-agent-skills/) NVIDIA published **DOCA AI agent skills** on [NVIDIA/skills](https://github.com/NVIDIA/skills): SKILL.md packs with real API signatures, hardware capability checks, and build constraints for Flow, GPUNetIO, PCC, RDMA, and more. Vendor 65-prompt eval: 19% checklist items without skills vs 100% with; RDMA demo used 73% less handwritten code and 46% fewer hardware commands. Installable into Claude Code, Codex, and other coding agents.

NVIDIA DOCA Agent Skills on GitHub: verified API/hardware contracts lift 65-prompt checklist pass from 19% to 100%
⚡ Key Takeaways
  • •Blog 2026-10-01 + https://github.com/NVIDIA/skills
  • •SKILL.md: real APIs, hardware capability manifests, build/preflight contracts
  • •Vendor 65-prompt eval: 19% → 100% checklist; top failures without skills documented in blog
  • •RDMA demo: −73% handwritten code, −46% hardware commands (vendor)
  • •Install into Claude Code/Codex-class agents; keep human gates on firmware writes
Read details→
D
DigitalOcean@digitalocean·14h ago
🛠️ Tooling

DigitalOcean Agent Droplets: harness + inference + 16k tools in one monthly plan from $50/$200

On [2026-10-01](https://www.digitalocean.com/blog/introducing-agent-droplets) DigitalOcean launched **Agent Droplets** for Managed Agents (public preview): Pro $50/mo (15% off) and Team $200/mo (20% off) bundling Harness Runtime microVMs, DO-hosted open-model inference (Kimi K3, GLM 5.3), session storage, and Action Gateway (16,000+ tools). Works with Claude Code, Codex CLI, OpenCode. Paired with prepaid Inference and Agents Balance.

DigitalOcean Agent Droplets: harness + inference + 16k tools in one monthly plan from $50/$200
⚡ Key Takeaways
Read details→
G
GitHub@github·16h ago
🛠️ Tooling

GitHub Copilot Dynamic Workflows public preview: code-defined multi-agent orchestration across CLI, app, and SDK

GitHub’s [2026-10-01 Changelog](https://github.blog/changelog/2026-10-01-dynamic-workflows-in-copilot-cli-and-the-copilot-app/) launched **Dynamic Workflows** in public preview for Copilot CLI, the Copilot app, and the SDK: **code-defined** reusable orchestration (sequential/parallel steps, pause checkpoints, subagent cross-checks)—unlike Autopilot or `/fleet`, where Copilot invents the plan each run. CLI needs `--experimental` or `/experimental on`; available on all Copilot plans.

⚡ Key Takeaways
  • •Anchors: Oct 1 Changelog + docs concepts + how-to pages linked above
  • •Code defines steps/conditions/handoffs; agents handle judgment. Distinct from Autopilot and /fleet
  • •Surfaces: CLI workflow run + /workflows; App Workflows button; SDK/extension canvases; share via personal or repo extensions
  • •Limits: max concurrent/total subagents, timeoutSeconds, approx maxAiCredits; prompt > workflow code > personal defaults
  • •Public preview; CLI experimental flag; no published agent-bench deltas—don’t conflate with HydraFusion or Computer Use
Read details→
C
Cursor@cursor_ai·18h ago
🛠️ Tooling

Cursor lists GLM 5.3 / Flash natively: 1M context, Other Models pool, from $1.4 / $0.15 per MTok

Cursor’s official [Models & Pricing](https://cursor.com/docs/models) and dedicated docs now list Z.ai **GLM 5.3** and **GLM 5.3 Flash** as addable third-party models: 1M context, Low/High/Max reasoning (High default), Other Models pool billing. Table rates: GLM 5.3 **$1.4/$0.26/$4.4**, Flash **$0.15/$0.029/$0.5** per MTok—same as 5.2, no long-context surcharge. Enable in Settings → Models (not the BYOK custom-endpoint path).

⚡ Key Takeaways
Read details→
A
AICoder Editorial@aicoder·20h ago
🔥 Trending

Developer 'A vs B' comparison intent surges: Sonnet 5.5, Sol, Argon, Air, and Copilot Computer Use collide in one week

Over ~7 days, Claude Sonnet 5.5, GPT-6.1 Sol, Gemini 4 Argon, JetBrains Air in IDEs EAP, and GitHub Copilot Computer Use landed in the same window. Search and HN discourse shifted from single-release chase to pairwise selection on price bands, terminal agents, IDE harnesses, and computer-use. Meta industry brief for readers without X.

⚡ Key Takeaways
  • •Same-week cascade: Sonnet 5.5 in Copilot (09-28) → Sol in Copilot (09-29) → Argon Fairwind preview (09-30) → Air EAP + Copilot Computer Use (10-01)
  • •Hot pairs: Sonnet↔Opus, Sonnet↔Sol, Argon↔Opus/Astra, Air↔Cursor, Copilot CU↔Claude CU, Codex↔Claude Code, Sol↔Astra same-family economics
  • •Selection axes: price band × terminal harness × IDE orchestration (Air ACP BYO) × computer-use for GUI-only legacy
  • •Compliance caveat: Anthropic subscription auth ToS constrains third-party harnesses; Air docs allow Claude subscription via terminal full-screen tab only
  • •Argon still Fairwind-gated (not GA); no public developer GA date—label comparison pages as preview-gated
Read details→
C
Cloudflare@Cloudflare·21h ago
🚀 Release

Cloudflare open-sources Clef & Clef-flash decision models on Workers AI — 209ms / 38.8ms median

[@Cloudflare](https://x.com/Cloudflare/status/2105747536510099540) (2026-10-01, #BirthdayWeek) launched its first in-house decision models **Clef** (27B) and **Clef-flash** (9B): hosted on Workers AI, [Apache 2.0](https://huggingface.co/Cloudflare/clef) weights, Jev System One–compatible; official blog cites ~**2.5× / 13×** median latency vs Jev and **2.2s vs 4.7s** in a threat-intel Browser Run workflow.

Cloudflare open-sources Clef & Clef-flash decision models on Workers AI — 209ms / 38.8ms median
⚡ Key Takeaways
  • •Primary: Cloudflare 2026-10-01 → Clef blog
  • •Median latency (43 runs): Clef 209.3 ms, Clef-flash 38.8 ms vs Jev 524.1 ms (~2.5× / 13×)
  • •Quality: BFCL 98.47/98.76; BANKING77 94.20; CLINC150+OOS 97.43; Clef family tops 7/10 decision benches
  • •Shape: Qwen backbone + non-AR schema scoring; 64K ctx; vision; @cf/cloudflare/clef / clef-flash; $0.24 / $0.09 per M input tokens
  • •Ship: Workers AI binding/REST; weights on HF; RL fine-tune via design-partner form
Read details→
Q
Qiao Zhang@zhangqiaorjc·22h ago
🔥 Trending

Gemini post-training lead: overnight fix for trainer–sampler mismatch, another 10% MFU; RSI in full swing

Gemini post-training/RL engineer Qiao Zhang ([@zhangqiaorjc](https://x.com/zhangqiaorjc/status/2105509406058463657), 2026-10-01), quoting the [Gemini 4 Argon](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/) launch, said mini-breakthroughs are accelerating: an overnight elimination of trainer–sampler mismatch and another ~10% MFU gain—RSI in full swing. Recipes are not in the official blog; access remains via [Fairwind](https://deepmind.google/fairwind-program/).

⚡ Key Takeaways
  • •Primary: @zhangqiaorjc 2026-10-01 quoting Argon launch — “RSI in full swing”
  • •Anecdotes: overnight elimination of trainer–sampler mismatch; another ~10% MFU — not official reproducible numbers
  • •Official anchors: Argon blog DeepSWE 77.9% / AutomationBench 51.3% / LVBench 91.7% / CWE-bench 68%; 1M output
  • •Access: Fairwind trusted defenders first; public API/AI Studio still gated
  • •Background only: TIM papers arXiv:2510.26788 / 2605.14220 — not claimed as Google’s method
Read details→
S
SuSu_酥酥@nft_chen·22h ago
🛠️ Tooling

Grok Bot proactive suggestions: offers help unprompted, asks before acting

Official [@bot](https://x.com/bot/status/2105713240701538538) (2026-10-01): Grok Bot can now suggest ways to help without you needing to ask. Fits the [Chief of Staff](https://cursor.com/docs/grok-bot/use-cases) pattern plus cloud computer, skills, and routines; consequential actions stay behind Approvals/Auto Review. Get the app at [x.ai/bot](https://x.ai/bot).

Grok Bot proactive suggestions: offers help unprompted, asks before acting
⚡ Key Takeaways
  • •Official: @bot 2026-10-01 — proactive help suggestions; CN brief: Tim @nft_chen/status/2105724531428233546
  • •Stack: persistent Bots, shared cloud computer, skills & routines — cursor.com/docs/grok-bot + x.ai/news/introducing-grok-bot
  • •Chief of Staff use case: calendar/Slack/inbox digest with source links; no send/reschedule without approval
  • •Trust: Approvals + Auto Review for consequential actions (cursor.com/docs/grok-bot/security)
  • •Install: x.ai/bot; included with paid Cursor / SuperGrok link; no public SWE-bench numbers
Read details→
G
GitHub@github·1d ago
🛠️ Tooling

GitHub Copilot Computer Use enters public preview in CLI and Copilot app (macOS/Windows)

On 2026-10-01 GitHub announced Computer Use in public preview for Copilot CLI and the Copilot app on macOS and Windows: agents can drive local desktop apps via accessibility tree/screenshots, clicks, typing, scroll/drag, and cross-app workflows—targeting GUI-only legacy software. Off by default; enable with `/computer on` or Settings; enterprise policy can block it.

GitHub Copilot Computer Use enters public preview in CLI and Copilot app (macOS/Windows)
⚡ Key Takeaways
Read details→