⚡Trending:
A
Anthropic@AnthropicAI·3h ago
🛠️ Tooling

Claude Code 2.1.287: Mods in-process extensions; built-in You should know side agent; 1M context default on gateways

Anthropic shipped [`Claude Code 2.1.287`](https://github.com/anthropics/claude-code/releases/tag/v2.1.287) on 2026-10-01 (npm [`@anthropic-ai/[email protected]`](https://www.npmjs.com/package/@anthropic-ai/claude-code)): **Claude Mods** let plugins run in-process JS/TS hooks that rewrite tool calls and UI; opt-in built-in **You should know** side agent; MCP URL elicitation; Opus 4.7+/Fable default to 1M context on Bedrock/Vertex/Foundry/apps gateway. Agent SDK TS [`0.3.287`](https://github.com/anthropics/claude-agent-sdk-typescript/releases/tag/v0.3.287) parity.

Claude Code 2.1.287: Mods in-process extensions; built-in You should know side agent; 1M context default on gateways
⚡ Key Takeaways
  • •Release v2.1.287; npm 2.1.287; GitHub ~2026-10-01T18:00:22Z
  • •Mods: in-process hooks for UI + tool.call rewrite; ≥2.1.287; docs https://code.claude.com/docs/en/plugins/mods/overview
  • •You should know: /plugin enable cc-plugin-you-should-know@builtin (first-party + telemetry); off by default
  • •1M context default for Opus 4.7+/Fable on Bedrock/Vertex/Foundry/apps gateway; disable via CLAUDE_CODE_DISABLE_1M_CONTEXT=1
  • •Agent SDK TS 0.3.287; MCP URL elicitation / bareElicitationCapability
Read details→
J
JetBrains@jetbrains·4h ago
🛠️ Tooling

JetBrains Air in IDEs EAP: Marketplace/2026.3 native, free Junie Lite, ACP for Codex·Copilot·Claude·Cursor

On 2026-10-01 JetBrains opened the [Air in IDEs EAP](https://blog.jetbrains.com/ai/2026/10/air-in-ides-eap/): an agentic session workspace via Marketplace plugin [#33314](https://plugins.jetbrains.com/plugin/33314-air) or native 2026.3 EAP builds. Ships with no agents; detects local ACP-compatible agents (Codex, Copilot, Claude, Junie, Cursor, …); free Junie Lite with JetBrains Account. Not an AI provider—third-party traffic goes to your vendor.

JetBrains Air in IDEs EAP: Marketplace/2026.3 native, free Junie Lite, ACP for Codex·Copilot·Claude·Cursor
⚡ Key Takeaways
Read details→
M
Mastramastra-ai·9h ago
🛠️ Tooling

Mastra @mastra/core 1.73.0: default stability errorProcessors; tool-call-resumed chunk; rewrite invalid Anthropic tool-call IDs

Mastra published [`@mastra/[email protected]`](https://www.npmjs.com/package/@mastra/core/v/1.73.0) on 2026-10-01 (~08:36Z; **124** commits vs [1.72.0](https://github.com/mastra-ai/mastra/compare/@mastra/[email protected]...@mastra/[email protected])): default stability `errorProcessors` (ProviderHistoryCompat, PrefillErrorHandler, StreamErrorRetryProcessor); new `tool-call-resumed` stream chunk; rewrite invalid Anthropic tool-call IDs outbound; export `findGatewayForModel`/`getGatewayId`; subagent model resolution with caller `requestContext`.

Mastra @mastra/core 1.73.0: default stability errorProcessors; tool-call-resumed chunk; rewrite invalid Anthropic tool-call IDs
⚡ Key Takeaways
  • •npm @mastra/[email protected] at 2026-10-01T08:36:38Z; 124 commits vs 1.72.0
  • •Default errorProcessors (#24473): history compat → prefill handler → stream retry; opt out with errorProcessorDefaults:false
  • •New tool-call-resumed chunk (#25491) before tool-result for suspend/approval resumes
  • •Rewrite invalid Anthropic tool-call IDs outbound (#24505); export findGatewayForModel/getGatewayId
  • •npm i @mastra/[email protected]; https://mastra.ai/docs
Read details→
G
Googlegoogleapis·9h ago
🚀 Release

Google Gen AI JS SDK 2.25.0: Gemini 3.8 Flash/Lite TTS; Environments module and GenerateContent labels

Google shipped [`@google/genai` v2.25.0](https://github.com/googleapis/js-genai/releases/tag/v2.25.0) on 2026-09-30 (npm ~22:51Z; **11** commits vs 2.24.0), aligning with Python [google-genai 2.26.0](https://github.com/googleapis/python-genai/releases/tag/v2.26.0): Gemini 3.8 Flash/Lite TTS model enums; dedicated Environments module with `from_environment`; labels on GenerateContent and LiveClientSetup; tuning `gcs_metrics_uri`.

Google Gen AI JS SDK 2.25.0: Gemini 3.8 Flash/Lite TTS; Environments module and GenerateContent labels
⚡ Key Takeaways
Read details→
ADSponsored
a
anomalycoanomalyco·9h ago
🛠️ Tooling

OpenCode 1.18.34: namespaced session identity headers on model requests; Developer ID-signed macOS CLI

anomalyco shipped [`OpenCode v1.18.34`](https://github.com/anomalyco/opencode/releases/tag/v1.18.34) on 2026-09-30 (npm ~22:38Z; **30** commits vs 1.18.33): send namespaced session and parent-session identity headers with model requests; sign macOS CLI/release binaries with Developer ID for macOS 27+; tree also refreshes Go Plus plan UI and daily model rankings.

OpenCode 1.18.34: namespaced session identity headers on model requests; Developer ID-signed macOS CLI
⚡ Key Takeaways
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
📊 Benchmark

CUA-SWE Unifies Computer-Use Agents and Visual Software Engineering: Multimodal Code-GUI Benchmarking Redefines Debugging

Modern software engineering transcends isolated code edits: engineers run applications, interact with GUI interfaces, and visually inspect rendering feedback to diagnose errors and verify changes. CUA-SWE bridges coding agents and computer-use agents by introducing the first unified benchmark and environment for visual software engineering. Spanning four engineering domains, CUA-SWE challenges agents to interleave code/config modifications, terminal execution, GUI interactions, and screenshot inspections, with deterministic test suites evaluating behavioral correctness.

CUA-SWE Unifies Computer-Use Agents and Visual Software Engineering: Multimodal Code-GUI Benchmarking Redefines Debugging
⚡ Key Takeaways
  • •First multimodal visual software engineering benchmark coupling computer-use capabilities with code repositories
  • •Text-only agents stall below 12% on visual defect repairs, while closed-loop visual feedback elevates success to 43.8%
  • •Identifies attention contention bottlenecks between high-resolution GUI screenshots and repository-scale source code
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

SkillGym Automates Skill-Use Agent Training: 6.8K Verifiable Sandboxes Enable Qwen3.5-9B to Outperform 397B Baselines

While procedural skills enhance autonomous agents on complex tasks, generating verifiable training environments and instilling robust skill-invocation behaviors remains an open challenge. SkillGym introduces an end-to-end automated pipeline that filters reproducible offline skills, synthesizes 6.8k verifiable environments via a Builder-Reviewer architecture, and collects 19k verified successful trajectories. SFT fine-tuning boosts the relevant skill-invocation rate from 28% to 96% and enables Qwen3.5-9B to outperform a 397B parameter base model across two demanding skill benchmarks.

SkillGym Automates Skill-Use Agent Training: 6.8K Verifiable Sandboxes Enable Qwen3.5-9B to Outperform 397B Baselines
⚡ Key Takeaways
  • •Synthesizes 6.8k verifiable sandboxes and 19k verified trajectories, driving skill reading rates from 28% to 96%
  • •Enables fine-tuned Qwen3.5-9B to outperform an untrained 397B base model on two demanding skill benchmarks
  • •Demonstrates robust out-of-distribution generalization across entirely held-out skill repositories
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

EVOKE Elicits Pretrained World Knowledge for Transferable Decision-Making via Fixed-State Goal Diversity

While LLM agents excel in multi-step decision-making, they transfer poorly to unseen environments. Conventional world-model approaches predict future observations, incurring high training overhead and compounding planning errors. EVOKE demonstrates that digital world knowledge is already internalized during pretraining, and failure to transfer stems from single-goal post-training that encourages superficial contextual shortcuts. By holding environment states fixed while ranking candidate actions across diverse counterfactual goals, EVOKE compels policies to elicit internalized world dynamics, boosting generalization across unseen environments.

EVOKE Elicits Pretrained World Knowledge for Transferable Decision-Making via Fixed-State Goal Diversity
⚡ Key Takeaways
  • •Replaces heavy explicit world models by eliciting internalized causal dynamics via fixed-state goal diversity
  • •Improves task success by 14.3 percentage points on unseen environments while cutting redundant exploration steps by 26%
  • •Achieves superior generalization using only 40% of the training trajectory data compared to standard single-goal recipes
Read details→
ADSponsored
G
Grokipedia@Grokipedia·13h ago
🚀 Release

Grokipedia v0.3: refreshed design & discovery; edit review resumes after months frozen

SpaceXAI’s Grokipedia shipped v0.3 on 2026-09-30: new logo, 3D-book Featured/Most-read, Latest edits, and article-page polish. The Verge reports edits flowing again after Lawfare’s documented April freeze. No public coding benchmarks for this release.

Grokipedia v0.3: refreshed design & discovery; edit review resumes after months frozen
⚡ Key Takeaways
  • •Version: Grokipedia v0.3 live at grokipedia.com (footer v0.3); announced 2026-09-30
  • •Focus: design/discovery (Featured 3D books, Most-read spines, Live edits), not a size/accuracy leaderboard drop — The Verge
  • •Governance: suggestion → Grok review; Lawfare freeze ~2026-04-24; processing resumed recently
  • •Scale clues: v0.1 ~885k articles (Oct 2025); Lawfare Aug cited 6M+; homepage approved edits ~1.17M+ live (not third-party audited)
  • •Benchmarks: no public SWE/Arena scores; use as a cross-check source with your own citation checks
Read details→
H
Hugging Face Daily Papers@HuggingFace·13h ago
🔥 Trending

AREX-2 Advances Test-Time Self-Improving Agents via Long-Horizon Reflective Programming: 81.8 on MLE-bench

Endowing LLM agents with test-time self-improvement—the ability to iteratively refine solutions—is a vital milestone for autonomous engineering. Researchers from Renmin University and collaborators present AREX-2, an open-source framework leveraging long-horizon reflective trajectories. Decoupling self-improvement into reflection (generating superior alternatives) and long-horizon execution (sustaining iteration over multiple turns), AREX-2 trains on verified machine learning and algorithmic programming trajectories. Built on Qwen3.8-27B, AREX-2 achieves 81.8 on MLE-bench Lite and 70.7 on Frontier-CS, transferring effectively to deep research with 92.2 on GAIA and 93.8 on DeepSearchQA.

AREX-2 Advances Test-Time Self-Improving Agents via Long-Horizon Reflective Programming: 81.8 on MLE-bench
⚡ Key Takeaways
  • •Decouples test-time self-improvement into domain-agnostic reflection and long-horizon execution
  • •Achieves 81.8 on MLE-bench Lite and 70.7 on Frontier-CS with Qwen3.8-27B, transferring to 92.2 on GAIA
  • •Breaks multi-turn degradation bottlenecks, showing monotonic score improvements as iteration budgets increase
Read details→
H
Hugging Face Daily Papers@HuggingFace·13h ago
🔥 Trending

False Frontiers Exposes Co-Cheating in Self-Evolving Agents: CrossFit Breaks Echo Chamber with 8.8 Point Benchmark Gain

Self-evolving search agents jointly optimize a question proposer and an answer solver to generate synthetic training curricula. However, False Frontiers exposes a critical failure mode: 'co-cheating', where proposer and solver agree on shared hallucinations and erroneous pseudo-labels, causing internal rewards to surge while real external accuracy collapses. The authors introduce CrossFit, which partitions source corpora into disjoint splits and scores proposals via cross-trained solvers. Evaluated on Qwen3.5-4B/9B across 7 downstream search benchmarks, CrossFit suppresses false agreement down to 3.7% and outperforms coupled self-evolution by 8.8 points and Search-R1 by 8.7 points.

False Frontiers Exposes Co-Cheating in Self-Evolving Agents: CrossFit Breaks Echo Chamber with 8.8 Point Benchmark Gain
⚡ Key Takeaways
  • •Identifies the root mechanism of 'co-cheating' in self-evolving agents where proposer and solver agree on false errors
  • •Introduces CrossFit cross-fitting architecture, compressing false-agreement mass from 8.8% to 0.1%
  • •Achieves an 8.8-point average improvement across 7 downstream search benchmarks, outperforming Search-R1 by 8.7 points
Read details→
H
Hugging Face Daily Papers@HuggingFace·14h ago
🔥 Trending

Mid-Harness Scales Actions at Model-Harness Boundary: Action Verification Lifts TerminalBench Pass@1 by 18%

Terminal agents operate through stochastic model generations, but a single errant command can irreversibly degrade sandbox environments. Mid-Harness introduces test-time compute scaling at the boundary between model and harness. By sampling and verifying candidate actions before terminal dispatch without modifying the underlying generator or harness, Mid-Harness boosts Pass@1 on TerminalBench-Lite from 50.00% to 68.03% under a GPT-5.6 Sol verifier, outperforming whole-trajectory scaling at lower token expenditure.

Mid-Harness Scales Actions at Model-Harness Boundary: Action Verification Lifts TerminalBench Pass@1 by 18%
⚡ Key Takeaways
  • •Allocates test-time compute at model-harness boundary, lifting TerminalBench-Lite Pass@1 from 50.00% to 68.03%
  • •Pairwise action verification and distillation enable compact local models to filter errant commands effectively
  • •Reduces estimated token overhead by over 35% compared to whole-trajectory sampling while increasing task completion
Read details→
B
BerriAI@BerriAI·14h ago
🛠️ Tooling

LiteLLM 1.103.2: forward Claude Code safeguards; attribute CLI session spend; restore pass-through pre-config wins

BerriAI shipped LiteLLM [`v1.103.2`](https://github.com/BerriAI/litellm/releases/tag/v1.103.2) on 2026-10-01 (PyPI ~06:20Z; **17** commits vs 1.103.1): forward Claude Code `safeguards` / dangerous-tool-use beta unchanged to Anthropic, Bedrock, Vertex, and Azure AI Foundry so auto mode works behind the proxy; attribute CLI session spend to a per-user `cli-session` alias for daily activity BI; restore pre-config-wins for pass-through endpoints.

LiteLLM 1.103.2: forward Claude Code safeguards; attribute CLI session spend; restore pass-through pre-config wins
⚡ Key Takeaways
  • •Release v1.103.2 at 2026-10-01T06:37:31Z; 17 commits vs v1.103.1
  • •Forward Claude Code safeguards + dangerous-tool-use beta to Anthropic/Bedrock/Vertex/Azure Foundry so auto mode returns safeguard_results
  • •CLI session spend attributed to per-user cli-session alias for daily activity BI (#40541/#43642/#43656)
  • •Restore pass-through pre-config-wins (#43962); verify Docker with cosign
  • •pip install 'litellm==1.103.2'; https://docs.litellm.ai/docs/ ; https://pypi.org/project/litellm/1.103.2/
Read details→
G
Googlegoogleapis·14h ago
🚀 Release

Google Gen AI Python SDK 2.26.0: Gemini 3.8 Flash/Lite TTS; Environments module and request labels

Google shipped [`google-genai` v2.26.0](https://github.com/googleapis/python-genai/releases/tag/v2.26.0) on 2026-09-30 (PyPI ~22:53Z; **18** commits vs 2.25.0): model enums add `gemini-3.8-flash-tts` and `gemini-3.8-flash-lite-tts` (lite replaces `gemini-3.1-flash-tts-preview`); dedicated Environments module with `from_environment`; labels on GenerateContent and LiveClientSetup; tuning jobs expose `gcs_metrics_uri`.

Google Gen AI Python SDK 2.26.0: Gemini 3.8 Flash/Lite TTS; Environments module and request labels
⚡ Key Takeaways
  • •Release v2.26.0 at 2026-09-30T22:17:07Z; PyPI google-genai 2.26.0; 18 commits vs 2.25.0
  • •TTS models: gemini-3.8-flash-tts + gemini-3.8-flash-lite-tts (lite replaces gemini-3.1-flash-tts-preview)
  • •Environments module + from_environment; domain types for environments/triggers/webhooks/agents
  • •Labels on GenerateContent and LiveClientSetup; tuning gcs_metrics_uri
  • •pip install -U 'google-genai==2.26.0'; https://googleapis.github.io/python-genai/
Read details→
V
Vercelvercel·14h ago
🛠️ Tooling

Vercel AI SDK: [email protected]–7.0.126 speech telemetry + tool-approval clear; @ai-sdk/openai ultrafast service tier

Vercel shipped [`[email protected]`](https://github.com/vercel/ai/releases/tag/[email protected]) (speech/transcription telemetry) through [`[email protected]`](https://github.com/vercel/ai/releases/tag/[email protected]) (clear tool approvals on `addToolOutput`) on 2026-09-30—**22** commits since covered [`[email protected]`](https://github.com/vercel/ai/releases/tag/[email protected]). Same day [`@ai-sdk/[email protected]`](https://github.com/vercel/ai/releases/tag/%40ai-sdk/openai%403.0.123) adds OpenAI `serviceTier: 'ultrafast'` (docs: access-controlled, gpt-5.6-sol), plus Azure MAI speech/transcribe and MCP OAuth conditional invalidation patches in the tree.

Vercel AI SDK: ai@7.0.124–7.0.126 speech telemetry + tool-approval clear; @ai-sdk/openai ultrafast service tier
⚡ Key Takeaways
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

Agent Error Dataset (AED) Released: 50,000+ Scaled Error-Diagnosis Pairs Power Failure Analysis and Post-Training

An unsuccessful agent rollout contains rich diagnostic signals beyond sparse negative rewards. UIUC researchers introduced the Agent Error Dataset (AED), the largest scale failure dataset comprising 50,228 error-diagnosis pairs across 9,961 source tasks, 33 environments, 19 harness families, and 23 policy models. Powered by a 5-stage Agentic Error-to-Training (AET) pipeline, first-proposal repairs lift verifier pass rates from 18.4% to 51.1% (+32.7 percentage points). Furthermore, post-training Qwen3-8B on AED boosts exact-step failure diagnosis agreement from 47.2% to 63.6%.

Agent Error Dataset (AED) Released: 50,000+ Scaled Error-Diagnosis Pairs Power Failure Analysis and Post-Training
⚡ Key Takeaways
  • •Features 50,228 curated error-diagnosis pairs across 33 environments, 19 harnesses, and 23 policy models
  • •Lifts replay verifier pass rates from 18.4% to 51.1% (+32.7 percentage points) with first-proposal corrections
  • •Boosts Qwen3-8B exact-step failure diagnosis agreement from 47.2% to 63.6%, demonstrating superiority over success-only training
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

SkillSeek Solves Marketplace-Scale Agent Skill Retrieval: Two-Stage IR with MCP Slashes Cost by 46% Over LLM Loops

Anthropic's Agent Skills (packaged into SKILL.md directories) encapsulate procedural know-how for LLM agents, with open aggregations expanding beyond 230,000 skills. While prior literature relied on expensive LLM-mediated loops within the agent decision process, SkillSeek introduces a two-stage IR pipeline (BGE-base bi-encoder + small cross-encoder) exposed via Model Context Protocol (MCP). Evaluated across 89 SkillsBench tasks on OpenHands, SkillSeek matches or exceeds LLM-mediated retrieval loops while slashing per-trial spend from $51.30 to $27.54 (a 46.3% reduction).

SkillSeek Solves Marketplace-Scale Agent Skill Retrieval: Two-Stage IR with MCP Slashes Cost by 46% Over LLM Loops
⚡ Key Takeaways
  • •Slashes per-trial retrieval spend from $51.30 to $27.54, reducing token costs by 46.3%
  • •Matches costly LLM-mediated loops across 89 SkillsBench tasks via BGE-base bi-encoders and compact cross-encoders
  • •Native Model Context Protocol (MCP) server support enables seamless drop-in integration with OpenHands, Claude Code, and agent harnesses
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

UIUC Proposes Meta-Skills for Agent Harness Design in Test-Time AI4AI: 12% Performance Gain Without Weight Updates

Agent efficacy relies heavily on execution environments (harnesses) in addition to reasoning capability. UIUC researchers led by Heng Ji propose a test-time AI-for-AI (AI4AI) framework where a Builder agent learns 'Meta-Skills' to construct custom execution environments for a frozen Target agent. Learned from execution feedback, these principles guide when and what support resources to provide. Across Harness-Bench and NewtonBench, meta-skill harnesses improve macro-average performance by 8.95 percentage points over baseline and 12.02 points over direct skill delivery, paving the way for autonomous system-level self-improvement.

UIUC Proposes Meta-Skills for Agent Harness Design in Test-Time AI4AI: 12% Performance Gain Without Weight Updates
⚡ Key Takeaways
  • •Lifts macro-average performance by 8.95 percentage points on Harness-Bench and NewtonBench with completely frozen weights
  • •Outperforms direct raw skill prompt injection by 12.02 percentage points via active harness synthesis
  • •Demonstrates robust performance gains when the same model acts as both Builder and Target, enabling test-time self-improvement
Read details→
G
GitHub@github·19h ago
🛠️ Tooling

GitHub Agentic Workflows 0.90.1: standalone ledgers + replay; richer audit/threat artifacts; compile-time self-hosted runner enforcement

[`gh-aw` v0.90.1](https://github.com/github/gh-aw/releases/tag/v0.90.1) (2026-09-30T17:20:00Z) adds standalone safe-output-backed ledgers with replay projections (defaults preserved on shared-workflow import), richer audit/usage artifacts (ledger txs, threat-detection outcomes, experiment/evals, friction costs), narrowed `add-labels` + required-labels gating for `assign-to-agent`, compile-time self-hosted runner enforcement, kebab-case sandbox frontmatter, and Claude Code CLI pin 2.1.280.

GitHub Agentic Workflows 0.90.1: standalone ledgers + replay; richer audit/threat artifacts; compile-time self-hosted runner enforcement
⚡ Key Takeaways
  • •v0.90.1 at 2026-09-30T17:20:00Z after v0.90.0
  • •Standalone safe-output ledgers + replay; defaults kept on shared-workflow import
  • •Richer audit: ledger txs, threat detection, evals, friction-cost attribution
  • •Governance: narrowed add-labels; required-labels on assign-to-agent; compile-time self-hosted runner enforcement
  • •Claude Code CLI pin 2.1.280; docs via gh-aw README / GitHub Next project page
Read details→
A
Anthropic@anthropics·19h ago
🚀 Release

Anthropic TS SDK 0.131 / Python 1.11: Admin list spend limits; client deprecates Sonnet 4.5

Hours after 0.130/1.10 Admin GA work, Anthropic shipped [`@anthropic-ai/[email protected]`](https://github.com/anthropics/anthropic-sdk-typescript/releases/tag/sdk-v0.131.0) and Python [`1.11.0`](https://github.com/anthropics/anthropic-sdk-python/releases/tag/v1.11.0) at 2026-09-30T22:58Z: Admin API gains a list spend limits endpoint; both clients deprecate Sonnet 4.5 toward newer aliases (Sonnet 5.5 line).

Anthropic TS SDK 0.131 / Python 1.11: Admin list spend limits; client deprecates Sonnet 4.5
⚡ Key Takeaways
  • •TS sdk-v0.131.0 at 2026-09-30T22:58:20Z; Python v1.11.0 at 22:58:09Z
  • •Admin API: list spend limits endpoint (compare)
  • •Clients deprecate Sonnet 4.5; prefer Sonnet 5.5 aliases for new work
  • •Install: npm i @anthropic-ai/[email protected] / pip install anthropic==1.11.0; Admin docs linked above
  • •No public SWE-bench; enterprise quota listing + model alias hygiene
Read details→