⚡Trending:
n
n8nn8n-io·4h ago
🛠️ Tooling

n8n 2.42.0: Agents on by default; AI Agent Tool nesting; MCP client/registry + Azure AI Foundry rename

n8n shipped [[email protected]](https://github.com/n8n-io/n8n/releases/tag/[email protected]) on 2026-09-29: Agents enabled by default (#39312); AI Agent Tool can use its own tools under a pre-v3 parent agent (#39160); MCP Client terminates sessions before close, MCP Registry flag removed from AI Assistant, editor adds Mistral Vibe to MCP picker; Azure OpenAI Chat Model renamed Azure AI Foundry with project/deployment listing and Entra app sign-in; plus shared Agent context and Agent Builder UX polish.

n8n 2.42.0: Agents on by default; AI Agent Tool nesting; MCP client/registry + Azure AI Foundry rename
⚡ Key Takeaways
  • •Release [email protected] on 2026-09-29; compare 2.41.0...2.42.0
  • •Agents enabled by default (#39312); shared Agent context (#38930)
  • •AI Agent Tool nests its own tools under pre-v3 parents (#39160)
  • •MCP: terminate session on client close; Registry ungated in AI Assistant; Mistral Vibe in MCP picker
  • •Azure OpenAI node renamed Azure AI Foundry Chat Model; docs AI Agents
Read details→
L
LangChainlangchain-ai·4h ago
🚀 Release

LangChain Anthropic 1.7.5 + Core 1.6.6: Claude Sonnet 5.5 compatibility, thinking/tool_choice guards

LangChain shipped `langchain-anthropic==1.7.5` and `langchain-core==1.6.6` on 2026-09-29, landing [#40882](https://github.com/langchain-ai/langchain/pull/40882): Claude Sonnet 5.5 model profile and capability discovery; reject unsupported forced tool_choice/thinking; allow mid-conversation system/tool changes; preserve computer-toolset namespaces, encrypted advisor blocks, and refusal details through streaming and replay. Invalid tool calls can fall back to unforced structured output or explicit `method="json_schema"`.

LangChain Anthropic 1.7.5 + Core 1.6.6: Claude Sonnet 5.5 compatibility, thinking/tool_choice guards
⚡ Key Takeaways
  • •Releases langchain-anthropic==1.7.5 + langchain-core==1.6.6
  • •Core PR #40882 merged 2026-09-29T13:05Z
  • •Validates thinking/tool_choice; mid-conversation system/tool changes; preserves streamed provider blocks
  • •Forced tool-choice fallback or explicit method="json_schema"
  • •Install: pip install -U "langchain-anthropic==1.7.5" "langchain-core==1.6.6"; docs Anthropic provider
Read details→
C
Charmbraceletcharmbracelet·4h ago
🛠️ Tooling

Charm Crush 0.97: toggle MCPs in UI; git branch in TUI; Claude Channels delivery + ctrl+b sidebar

Charm's terminal coding agent Crush shipped v0.97.0/v0.97.1 on 2026-09-29 (npm `@charmland/[email protected]`): toggle MCP servers on/off in the UI (global or local); show the current Git branch in the TUI; continue Claude Channels session injection with exactly-once routing; toggle the sidebar with ctrl+b; fix Anthropic cache-creation token counting, clipboard false failures, and OpenAI OAuth refresh via Catwalk. ~36 commits vs v0.96.0.

Charm Crush 0.97: toggle MCPs in UI; git branch in TUI; Claude Channels delivery + ctrl+b sidebar
⚡ Key Takeaways
  • •Releases v0.97.0 + v0.97.1; ~36 commits vs v0.96.0
  • •Toggle MCPs on/off in UI globally or locally (#3725)
  • •Git branch in TUI (#2437); sidebar toggle via ctrl+b (#3935)
  • •Claude Channels delivery part 2/3: session injection, exactly-once routing (#3346)
  • •Install: npm i -g @charmland/[email protected] or Homebrew; docs at charm.sh/crush
Read details→
C
Comet@comet-ml·8h ago
🛠️ Tooling

Comet Opik 2.2.84: online eval rules call Jev via OpenRouter Decisions API; queue automation + provenance

Comet Opik 2.2.84 (2026-09-29) lets online LLM-as-a-Judge rules call TypeSafe Jev end-to-end via OpenRouter `POST /api/alpha/decisions` (not chat/completions). Each Boolean metric becomes a yes/no question answered in one call; score 1 if probability ≥0.5, with `Probability: 0.93` in the reason. Backend adds `OpenRouterDecisionsClient` (workspace OpenRouter key, 429/5xx retries) for `~typesafe/jev-latest` / `typesafe/jev-1.13`. Also: Annotation Queue automation + item provenance. ~8 commits / 104 files vs 2.2.83.

Comet Opik 2.2.84: online eval rules call Jev via OpenRouter Decisions API; queue automation + provenance
⚡ Key Takeaways
  • •Release 2.2.84; ~8 commits / 104 files vs 2.2.83
  • •PR #8540: online judges call Jev via OpenRouter Decisions API
  • •Boolean metrics → yes/no questions; score 1 if p≥0.5; reason shows Probability
  • •Models: ~typesafe/jev-latest, typesafe/jev-1.13
  • •Queue automation + item provenance; trace stats / thread_id fixes
Read details→
ADSponsored
C
Cline@cline·8h ago
🚀 Release

Cline SDK 0.0.87: Windows proxy-safe hub startup; reasoning tokens not double-counted; Bedrock GPT-6 via inference profiles

Cline shipped SDK v0.0.87 on 2026-09-29: Windows hub startup works when `HTTP(S)_PROXY` is set—`ensureLoopbackProxyBypass()` adds loopback to `NO_PROXY` so Bun fetch no longer routes health probes through Clash/v2ray/corp proxies; `normalizeUsage()` excludes reasoning from `outputTokens` to stop double-counting; Gateway models keep per-model `apiProtocol`; bare GPT-6/GPT-5.6 on Bedrock use inference profiles (`@ai-sdk/amazon-bedrock` ^5.0.96); catalog 209→211 providers, 6,386→6,447 models; recommended list adds Claude Sonnet/Opus 5.5.

Cline SDK 0.0.87: Windows proxy-safe hub startup; reasoning tokens not double-counted; Bedrock GPT-6 via inference profiles
⚡ Key Takeaways
  • •Release sdk/v0.0.87 vs v0.0.86
  • •Windows proxy: ensureLoopbackProxyBypass() stops hub probes from going through HTTP(S)_PROXY
  • •outputTokens excludes reasoning subset (cost still on full output); no double-count in session totals
  • •Bedrock GPT-6/GPT-5.6 via inference profiles; @ai-sdk/amazon-bedrock ^5.0.96
  • •Catalog 209→211 providers, 6,386→6,447 models; Sonnet/Opus 5.5 on recommended list
Read details→
O
OpenAI@openai·8h ago
🚀 Release

OpenAI Codex 0.159.0: opt-in instant_interrupt steering; warning dismiss; richer Mermaid & thread pagination

OpenAI Codex shipped rust-v0.159.0 on 2026-09-29 (npm `@openai/[email protected]`): opt-in `instant_interrupt` steers mid-response or long code-mode calls; compact welcome/headers; warnings viewer dismisses reviewed items on close (`k` to keep); scroll transcript while reviewing a plan; richer native Mermaid; app-server can paginate thread history from an item anchor. ~89 commits / 300 files vs v0.158.0. Windows suppresses stray consoles for MCP/code-mode; `.aws` protected under writable roots; macOS sandbox TLS/proxy fixes.

OpenAI Codex 0.159.0: opt-in instant_interrupt steering; warning dismiss; richer Mermaid & thread pagination
⚡ Key Takeaways
  • •Release rust-v0.159.0 on 2026-09-29; ~89 commits / 300 files vs v0.158.0
  • •Opt-in instant_interrupt steers mid-response / long code-mode (#48135/#48141)
  • •Warnings dismiss on close (k to keep); scroll while reviewing plans; richer Mermaid
  • •App-server paginates thread history from an item anchor (#48151)
  • •Install: npm i -g @openai/[email protected]; removed tui.prompt_suggestions and bundled plugin-creator skill
Read details→
H
Hugging Face Daily Papers@HuggingFace·10h ago
🔥 Trending

GAGAR: Groupwise Agentic Grading and Advantage Redistribution Stabilizes Code Agent RL on 1.02T Models

Reinforcement learning for code agents typically relies on executable tests for binary rewards, leading algorithms like GRPO to assign identical scalar advantages to all passing rollouts regardless of code bloat, extraneous modifications, or architectural hygiene. Researchers from ByteDance, Peking University, and Tsinghua introduced GAGAR, a quality-aware advantage redistribution framework. By staging rollouts in a shared workspace where an SFT agentic grader ranks passing candidates and dynamically shifts credit toward clean solutions, GAGAR successfully stabilizes code agent RL on 310B Flash and 1.02T Pro models.

GAGAR: Groupwise Agentic Grading and Advantage Redistribution Stabilizes Code Agent RL on 1.02T Models
⚡ Key Takeaways
  • •Overcoming Binary Verification Blind Spots: Conventional code agent RL treats all test-passing rollouts equally, causing policy optimization to reward bloated, fragile, or out-of-scope code as long as test assertions pass.
  • •Sum-Preserving Credit Redistribution: Stages rollout groups in a shared context where an SFT-trained agentic judge ranks solutions, downweighting sloppy passes while proportionally shifting advantages to clean, minimal-diff implementations without altering total group advantage magnitude.
  • •Industrial Validation on 1.02T Model: Evaluated at scale on MiMo-V2.6-Flash (310B) and MiMo-V2.6-Pro (1.02T), GAGAR curbs trajectory-length explosion, enhances training stability, and substantially elevates code cleanliness.
Read details→
H
Hugging Face Daily Papers@HuggingFace·10h ago
📊 Benchmark

WideSWE: First Multi-Repository Benchmark for Coding Agents Spans 103 Ecosystems, Top Success Capped at 42.5%

Addressing the fundamental limitation of benchmarks like SWE-bench that assess coding agents strictly within isolated repositories, researchers from Zhejiang University introduced WideSWE. Mining changes across 103 real-world software ecosystems, WideSWE establishes 120 cross-repository tasks (balanced between 60 bug fixes and 60 feature developments) with comprehensive regression test suites. Across seven frontier agent configurations, full task success tops out at just 42.50%, exposing systemic failures in cross-repo dependency tracking and coordinated verification.

WideSWE: First Multi-Repository Benchmark for Coding Agents Spans 103 Ecosystems, Top Success Capped at 42.5%
⚡ Key Takeaways
  • •Cross-Repository Benchmark Evolution: Moves beyond single-codebase evaluations by capturing interdependent multi-repository modifications across 103 real-world open-source software ecosystems.
  • •Frontier Agents Top Out at 42.50%: Across seven leading agent setups, end-to-end task completion ranges from 10.83% to 42.50% (Codex CLI + GPT-5.6-sol achieving the high mark), exposing acute bottlenecks in cross-repo dependency coordination.
  • •Isolated vs. Joint Workflow Dynamics: Proves that treating repositories independently leads to omitted cross-system dependencies, whereas joint execution surfaces essential multi-repo type signatures and test harness context.
  • •Complete Open-Source Test Harness: Benchmark code, environmental harnesses, and evaluation suites released at github.com/ZJU-ACES-ISE/WideSWE.
Read details→
ADSponsored
A
Alibaba Qwen Team@Alibaba_Qwen·10h ago
🔥 Trending

Alibaba Unveils QwenGyre: Elastic Reinforcement Learning Framework for 700K-Token xLong-Horizon Agents

To overcome the severe GPU starvation and explosive trajectory redundancy of extreme-long (xlong) horizon agents spanning hours, hundreds of environment steps, and nearly 1M tokens per rollout, Alibaba's Qwen team introduced QwenGyre. By dynamically reallocating GPUs between rollout and training without pausing live interactions and pruning non-linear trajectory branches, QwenGyre enables 700K-token RL on the 2.4T-parameter Qwen 3.8 flagship, driving a 6.0% absolute gain on NL2RepoBench in 48 steps while delivering a 1.85x speedup.

Alibaba Unveils QwenGyre: Elastic Reinforcement Learning Framework for 700K-Token xLong-Horizon Agents
⚡ Key Takeaways
  • •Breakthrough Infrastructure for xLong-Horizon RL: First production-grade RL framework built to sustain single-rollout trajectories spanning up to 1M tokens and hundreds of interactive steps for software engineering agents.
  • •Elastic Live GPU Reallocation: Eradicates massive GPU idle bubbles caused by execution variance by dynamically migrating GPUs between rollout and backward training clusters without halting live environment containers.
  • •2.4T-Parameter Model Gains & 1.85x Speedup: Scaled to the 2.4-trillion-parameter Qwen 3.8 backbone with 700K tokens per rollout, driving an absolute 6.0% accuracy gain on NL2RepoBench (52.5% to 58.5%) in 48 steps with 1.85x higher training throughput.
Read details→
H
Hugging Face Daily Papers@HuggingFace·13h ago
🔥 Trending

Domain-Normalized MOPD: Rescaling Teacher Feedback Prevents Instruction Dominance in Multi-Expert Distillation

Multi-teacher on-policy distillation (MOPD) aims to synthesize diverse capabilities from specialized models (mathematics, coding, instruction-following) into a single versatile student. However, empirical studies reveal that standard MOPD students fail to beat single-specialist baselines and lose their math edge because instruction-following feedback is exponentially more dispersed and dominates student gradient updates. Researchers introduced Domain-Normalized MOPD (DN-MOPD), which rescales feedback by measured domain spread to recover math and coding expertise across six public benchmarks.

Domain-Normalized MOPD: Rescaling Teacher Feedback Prevents Instruction Dominance in Multi-Expert Distillation
⚡ Key Takeaways
  • •Diagnosing the Instruction Dominance Trap: Explains why traditional multi-teacher distillation falls short: instruction-following supervisory feedback exhibits significantly wider variance than mathematical tokens, disproportionately skewing parameter updates.
  • •Domain-Normalized Gradient Rescaling: Retains domain-specific routing while normalizing feedback magnitudes by empirical spread, constraining instruction drift and allowing specialized mathematical and programming signals to register effectively.
  • •Comprehensive Benchmark Validation: Evaluated across three Qwen3.5 parameter tiers on six public benchmarks under varied token limits, DN-MOPD consistently outperforms standard MOPD and successfully recaptures specialist-grade math reasoning.
Read details→
H
Hugging Face Daily Papers@HuggingFace·13h ago
🔥 Trending

KVCMAS: Low-Rank Online KV Cache Correction for Multi-Agent Systems Cuts TTFT by 2x and Peak VRAM by 3.7x

In collaborative multi-agent systems sharing a common foundation model, distinct system prefixes cause KV cache representations for the shared context to diverge, forcing each agent to redundantly re-prefill the entire conversation history. Researchers from Seoul National University and KAIST introduced KVCMAS, an online KV cache correction framework. By parameterizing cross-agent cache deviations as compact low-rank states and chaining corrections along agent pipelines without external reference prefills, KVCMAS accelerates TTFT by 2.0x while reducing peak GPU memory by up to 3.7x.

KVCMAS: Low-Rank Online KV Cache Correction for Multi-Agent Systems Cuts TTFT by 2x and Peak VRAM by 3.7x
⚡ Key Takeaways
  • •Resolving Multi-Agent Prefix Divergence: Addresses the fundamental impasse where distinct agent system prompts contaminate downstream KV representations for shared dialogue histories, defeating naive cache reuse.
  • •Low-Rank Chained Delta Corrections: Represents inter-agent cache deviations using compact low-rank projections, chaining updates dynamically along the agent pipeline without incurring out-of-band reference prefill overhead.
  • •2.0x TTFT Acceleration & 3.7x Memory Reduction: Achieves a 2.0x time-to-first-token speedup over non-shared multi-agent serving traces while slashing peak VRAM by up to 3.7x compared to state-of-the-art correction methods.
  • •Full Preservation of Output Fidelity: Maintains exact numerical cache states for the leading agent while preserving zero downstream accuracy degradation across complex multimodal and text agent suites.
Read details→
G
GitHub@github·13h ago
🛠️ Tooling

GitHub Agentic Workflows 0.90.0: friction cost in step summaries; two noop calls by default; five new trajectory graders

gh-aw v0.90.0 (2026-09-28) renders friction cost in step summaries, allows two `noop` safe-outputs by default, adds five trajectory graders with hardened generator contracts, and bumps `gh-aw-mcpg` to v0.4.27 / `gh-aw-firewall` to v0.28.27—shifting from 0.89.22’s sandbox focus toward observability and grading.

GitHub Agentic Workflows 0.90.0: friction cost in step summaries; two noop calls by default; five new trajectory graders
⚡ Key Takeaways
  • •Release v0.90.0 after 0.89.22
  • •Step summaries show friction cost (Cost Management)
  • •Default max two noop safe-outputs—skip the AI engine with zero AI Credits when work is empty
  • •Five new trajectory graders + hardened generator contracts
  • •Infra: gh-aw-mcpg v0.4.27, gh-aw-firewall v0.28.27; upload-assets summary formatting fix
Read details→
C
CrewAI@crewAIInc·13h ago
🛠️ Tooling

CrewAI 1.15.23: native Gemini 3.8 Flash; AMP eval of the last traced run

CrewAI 1.15.23 (2026-09-28) adds native Gemini 3.8 Flash support; `crewai eval` can score the last AMP-traced run (not just print it); task spans capture declared output format/results; LLM retries throttled providers and Bedrock `acall` falls back to sync. Tightens the observe→evaluate loop for multi-agent crews.

CrewAI 1.15.23: native Gemini 3.8 Flash; AMP eval of the last traced run
⚡ Key Takeaways
  • •Release 1.15.23 / PyPI
  • •Native Gemini 3.8 Flash via Google Gen AI SDK (LLM docs)
  • •crewai eval scores the last AMP-traced run; TUI Evaluate button + graded areas
  • •Task spans record declared output format and results
  • •Retries on throttled providers; Bedrock acall→sync fallback; SQLite/S3/Selenium cleanup fixes
Read details→
O
Ollama@ollama·13h ago
🚀 Release

Ollama Python 0.6.3: native systemone() for Decision Models; think accepts model-defined levels

Official Ollama Python client v0.6.3 shipped 2026-09-29: adds `ollama.systemone()` matching server v0.35.0 `/v1/systemone` Decision Models (choice/noul/score); relaxes `think` to string levels; fixes multi-tool examples and case-insensitive image extensions. Requires Ollama ≥0.35.0 plus local decision models such as nimble/tev1.

Ollama Python 0.6.3: native systemone() for Decision Models; think accepts model-defined levels
⚡ Key Takeaways
  • •Release v0.6.3 on 2026-09-29; ~6 commits vs v0.6.2 including examples/systemone.py
  • •New ollama.systemone(model, state, questions) batches choice/noul/score Decision Model queries
  • •Needs Ollama server v0.35.0+ and a local decision model (nimble/tev1)
  • •think accepts string levels for model-defined thinking (#697/#744)
  • •Install: pip install -U ollama==0.6.3
Read details→
H
Hugging Face Daily Papers@HuggingFace·13h ago
🔥 Trending

Tsinghua Open-Sources TaH2: Adaptive Looped Transformers Boost Test-Time Scaling Slope by 53% on AIME

While looped transformers achieve high parameter efficiency by iteratively recycling layers for latent compute, fixed-depth looping expends redundant iterations on simple tokens and suffers from performance saturation. Tsinghua University's NICS lab developed TaH2, an adaptive looped framework that co-trains the backbone and an iteration decider via lookahead depth supervision. Evaluated on the rigorous AIME competition benchmarks, TaH2 accelerates the accuracy-compute scaling slope by 53% (2.74 vs 1.79) and surpasses non-looped baselines by +3.4 to +3.9 points at matched test-time compute.

Tsinghua Open-Sources TaH2: Adaptive Looped Transformers Boost Test-Time Scaling Slope by 53% on AIME
⚡ Key Takeaways
  • •Adaptive Looping via Lookahead Depth Supervision: Discards rigid, uniform layer looping by training a lightweight iteration decider using online labels that predict whether subsequent recurrent passes genuinely improve token predictions.
  • •53% Steeper Test-Time Compute Scaling Slope: On challenging AIME competition benchmarks, improves the accuracy-to-compute slope from 1.79 to 2.74, outpacing non-looped baselines by +3.4 points at matched decoding FLOPs.
  • •Overcoming Looping Saturation: While existing recurrent transformers plateau beyond shallow depths, TaH2's performance margin over dense models expands continuously from +2.8 points at depth 2 to +3.9 points at depth 8.
  • •Complete Codebase Open-Sourced: Full post-training recipes, decider checkpoints, and inference evaluation harnesses released at github.com/thu-nics/TaH.
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

Resolving Training-Inference Mismatch in LLM RLVR: Calibrated Importance Sampling Eliminates Gradient Variance

In reinforcement learning with verifiable rewards (RLVR), rollout trajectories are generated by high-throughput inference engines while backward gradients are computed by training engines, introducing subtle floating-point probability discrepancies that trigger unbounded variance under exact importance sampling. Researchers identified the invariance of logit displacement distributions and introduced Calibrated Importance Sampling (CIS), featuring confidence-aware truncation that replaces unbounded second moments with constant bounds, establishing top scores across five mathematical benchmarks.

Resolving Training-Inference Mismatch in LLM RLVR: Calibrated Importance Sampling Eliminates Gradient Variance
⚡ Key Takeaways
  • •Exposing Engine Probability Discrepancies: Fast inference engines (vLLM) and training backends (Megatron) assign diverging probabilities to identical tokens due to kernel optimizations, inducing severe variance under standard importance sampling.
  • •Logit-Displacement Invariance: Formulates that per-logit perturbations before softmax remain approximately invariant to token confidence, enabling a principled confidence-aware importance weight cap.
  • •Theoretical Variance Bounds & Benchmark Leadership: Replaces unbounded second moments with strict constant bounds, outperforming existing baselines across three MoE backbones on five rigorous mathematical reasoning benchmarks.
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

ControlScope: Benchmarking Workflow Revision and Repair Granularity in Autonomous LLM Agents

Investigating how deeply an agent should revise an interrupted workflow, ControlScope systematically isolates three repair granularities: continuing generated code (KEEP), patching only tool call arguments (ARG), and rewriting the unfinished workflow (FULL). Across filesystem benchmarks, ALFWorld, and AppWorld, the study reveals that unconstrained full-workflow revisions regularly interrupt viable execution paths with self-doubt. Enforcing nested permissions and a 5-call budget cap saves 19.4% of model output tokens without degrading task success.

ControlScope: Benchmarking Workflow Revision and Repair Granularity in Autonomous LLM Agents
⚡ Key Takeaways
  • •The Revision Interruption Paradox: Demonstrates that unconstrained full-pipeline rewrites (FULL) frequently interrupt viable agent-written routines and trigger cascading logical regressions across multi-step environments.
  • •Granular Repair Scoping: Comparing KEEP, ARG (argument-only patching), and FULL reveals that narrow argument intervention delivers the optimal balance between repair precision and execution stability.
  • •19.4% Output Token Savings: A 5-call protection mechanism eliminates redundant agent hallucination loops, slashing logged output tokens by 19.4% while maintaining high completion rates.
  • •Nested Permission Boundaries: Recommends architectural firewalls separating the actions an agent can freely execute from the scope of workflow revisions it is permitted to initiate.
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

CompoWorld: Compositional Environment Scaling with 10,130 Tools and 448 Services Outperforms Frontier Agents

Addressing the limitation where synthetic training environments only generate tasks within isolated individual domains, researchers introduced CompoWorld. By scaling task spaces across 448 typed services exposing 10,130 tools and generating cross-service workflows via dependency graphs, paired with Completion-Focused Rubric Rewards in RL, a Qwen3.6-35B-A3B agent trained on just 3K SFT trajectories and 1K RL tasks gains +9.17 points across eight benchmarks and surpasses Claude Opus 4.6 on AutomationBench.

CompoWorld: Compositional Environment Scaling with 10,130 Tools and 448 Services Outperforms Frontier Agents
⚡ Key Takeaways
  • •Compositional Scaling Across 10,130 Tools: Breaks the single-environment toy benchmark barrier by deploying coding agents to implement 448 verified microservices exposing 10,130 tools under unified state schemas.
  • •Dependency Graph Random-Walk Workflows: Chains disparate services through dependency topologies to generate verifiable agent workflows where context and execution state must navigate multi-service boundaries.
  • •35B Backbone Surpasses Frontier Models: Lifts eight agent benchmark averages by +9.17 points, overtaking Claude Opus 4.6 on AutomationBench and setting a new efficiency record for open-weights agent models.
Read details→
L
LangChain@LangChainAI·18h ago
🛠️ Tooling

LangChain 1.4.3: Bedrock Mantle in init_chat_model, GPT-6 structured-output recognition, create_agent repairs invalid tool calls

LangChain shipped `langchain==1.4.3` on 2026-09-28 (on PyPI): `init_chat_model` supports AWS Bedrock Mantle chat models via langchain-aws Mantle interfaces; recognizes GPT-6 structured output without profiles; `create_agent` emits error `ToolMessage`s for invalid tool calls missing results on replay; sanitizes cache settings for fallback models. Key PRs: #40837, #40844, #40530, #40886 vs 1.4.2.

LangChain 1.4.3: Bedrock Mantle in init_chat_model, GPT-6 structured-output recognition, create_agent repairs invalid tool calls
⚡ Key Takeaways
  • •langchain==1.4.3 on PyPI (2026-09-28)
  • •init_chat_model + Bedrock Mantle (#40837) via langchain-aws Mantle chat wrappers
  • •GPT-6 structured output recognized without profiles (#40844)
  • •create_agent repairs invalid tool calls with error ToolMessages (#40530); sanitize fallback cache settings (#40886)
Read details→
A
Anthropic@AnthropicAI·18h ago
🚀 Release

Claude Code 2.1.284: defaults Sonnet to claude-sonnet-5-5 (1M ctx, $2/$10); /mcp reconnect all + spend limits in /usage

Anthropic shipped Claude Code v2.1.284 on 2026-09-28 (`@anthropic-ai/[email protected]`): `claude-sonnet-5-5` is now the default Sonnet on the Anthropic API—1M context, $2/$10 per MTok, $0.20/MTok cache reads. Adds “Yes, but ask again next time” for auto-mode reads, `/mcp reconnect all`, dollar spend limits in `/usage` and the status line (`used_usd`/`limit_usd`/`period`), and rebindable effort-slider keybindings. Stability fixes: damaged streams writing undefined, overloaded errors after thinking blocks, and double-compact when Prompt-too-long persists.

Claude Code 2.1.284: defaults Sonnet to claude-sonnet-5-5 (1M ctx, $2/$10); /mcp reconnect all + spend limits in /usage
⚡ Key Takeaways
  • •v2.1.284 makes claude-sonnet-5-5 the default Sonnet (1M ctx, $2/$10, $0.20 cache reads)
  • •/usage + status line show dollar spend limits; /mcp reconnect all retries failed MCP servers
  • •Auto-mode: Yes-but-ask-again for out-of-cwd reads; effort slider keybindings rebindable
  • •Stability: damaged-stream/undefined fixes; retry after thinking-block overload; second compact on Prompt-too-long
Read details→