⚡Trending:
S
Strands Agents@strands-agents·1h ago
🛠️ Tooling

Strands Agents 1.57.1: bidi snapshot capture/restore, vended a2a_client tool, CLI Quickstart/Customize/Export + /tools

Strands Agents (harness-sdk) python/v1.57.1 (2026-09-25) adds bidi snapshot capture/restore, a vended a2a_client tool, Agent.shutdown() for scope cleanup, CLI Quickstart/Customize/Export replacing Q&A setup plus /tools, and MCP TS client 2.0. PyPI strands-agents 1.57.1 is live.

Strands Agents 1.57.1: bidi snapshot capture/restore, vended a2a_client tool, CLI Quickstart/Customize/Export + /tools
⚡ Key Takeaways
  • •Shipped: harness-sdk python/v1.57.1 = PyPI strands-agents 1.57.1; docs strandsagents.com/docs
  • •Bidi: snapshot capture/restore; explicit model IDs for providers; align response boundaries and record tool dispatch in history
  • •Interop: vended a2a_client tool; extract content from task.status.message when artifacts are absent
  • •Lifecycle: Agent.shutdown() for scope cleanup; MCP server migrates off removed mcp.server.fastmcp for mcp 2.x
  • •CLI: Quickstart/Customize/Export replace setup Q&A; add /tools; companion harness-cli/v0.1.4 same day
Read details→
G
Google ADK@google·2h ago
🛠️ Tooling

Google ADK Python 2.10.0: experimental skill lifecycle caps, MongoDB vector/hybrid search, eval duration/token/call metrics

Google ADK Python v2.10.0 (GitHub 2026-09-25; changelog 2026-09-24) adds experimental skill lifecycle controls (EPHEMERAL, active-skill caps, unload_skill via ADK_ENABLE_SKILL_LIFECYCLE=1), a MongoDB toolset with vector/hybrid search, eval efficiency metrics for duration/tokens/model calls, and better OpenAI reasoning param adaptation plus reasoning-token reporting. PyPI google-adk 2.10.0 is live.

Google ADK Python 2.10.0: experimental skill lifecycle caps, MongoDB vector/hybrid search, eval duration/token/call metrics
⚡ Key Takeaways
  • •Shipped: GitHub v2.10.0 + PyPI google-adk 2.10.0; docs at adk.dev/get-started
  • •Skills (experimental): enable ADK_ENABLE_SKILL_LIFECYCLE=1 for EPHEMERAL one-turn lifecycle, active-skill caps, and opt-in unload_skill
  • •Data: MongoDB toolset with vector and hybrid search inside agent flows
  • •Eval: duration efficiency metric plus token and model-call counts
  • •Models: OpenAI reasoning param adaptation and reasoning-token reporting; use effort instead of thinking_config on OpenAIResponsesLlm; adk create adds gemini-3.8-flash
Read details→
C
Comet@CometML·6h ago
🛠️ Tooling

Comet Opik 2.2.81: scored entities → annotation queues, gated free-form ClickHouse SQL user, 19 provider price tables

Comet Opik 2.2.81 (2026-09-25) routes scored entities into annotation queues, optionally provisions an extended free-form SQL ClickHouse user, registers 19 canonical providers so model prices load, and fixes Gemini usage, streamed Mistral reasoning, and slots=True dataclass encoding in the SDK.

Comet Opik 2.2.81: scored entities → annotation queues, gated free-form ClickHouse SQL user, 19 provider price tables
⚡ Key Takeaways
  • •Shipped 2.2.81 on 2026-09-25; pip install opik==2.2.81
  • •Evaluation loop: route scored entities into annotation queues
  • •Data plane: optional extended free-form SQL ClickHouse user behind a flag; drop spans FINAL on find_trace_stream / find_thread_by_id
  • •Cost: register 19 canonical providers so model prices load
  • •SDK: Gemini usage without candidates_token_count; keep streamed Mistral reasoning; encode slots dataclasses as dicts; opik.sh startup timeout 90s
Read details→
O
OpenHands@OpenHands·6h ago
🛠️ Tooling

OpenHands v1.24.0: workspace folder toggle, read-only shared automations, MCP OAuth kept on cloud save

OpenHands Agent Canvas v1.24.0 (2026-09-25) adds header workspace-folder toggles, read-only shared automation conversations on cloud, MCP OAuth credentials preserved on cloud saves (skip consent when tokens still work), OpenAI subscription caches scoped by backend, and Node.js ≥24.

OpenHands v1.24.0: workspace folder toggle, read-only shared automations, MCP OAuth kept on cloud save
⚡ Key Takeaways
  • •Shipped 2026-09-25: GitHub v1.24.0 + npm @openhands/[email protected]
  • •UX: toggle all workspace folders from Conversations header; shared automation conversations open read-only on cloud
  • •MCP: keep OAuth credentials on cloud saves; skip consent when tokens still work; render ACP tool-call blocks in chat cards
  • •Multi-backend: scope OpenAI subscription caches and settings cache by answering backend
  • •Install: npm i -g @openhands/agent-canvas requires Node ≥24; see docs.openhands.dev
Read details→
ADSponsored
L
Langfuse@langfuse·7h ago
🛠️ Tooling

Langfuse v4.46.0: basic skill management (unstable REST), streamlined experiment side-by-side, 100-row tracing tables

Langfuse v4.46.0 (2026-09-25) ships basic skill management (PR #17803) via unstable REST (list/create/get/patch/delete skills, versions, file content), streamlined experiment side-by-side comparison, formatted JSON on dataset items, optional 100-row tracing page size, assistant respect for data-retention, and v3 dataset-run-items reads on replicas.

Langfuse v4.46.0: basic skill management (unstable REST), streamlined experiment side-by-side, 100-row tracing tables
⚡ Key Takeaways
  • •Shipped: Langfuse v4.46.0 (2026-09-25); skills PR #17803 merged same day
  • •Skills (unstable): /api/public/unstable/skills CRUD-ish endpoints for skill artifacts aimed at coding agents
  • •Experiments: streamlined side-by-side comparison; formatted JSON on dataset items
  • •Observability UX: optional 100-row tracing page size; plain-text header metrics; 300ms tooltips
  • •Compliance/perf: assistant respects data-retention; v3 dataset-run-items on read replica; docs langfuse.com/docs
Read details→
C
Cline@cline·12h ago
🛠️ Tooling

Cline Desktop v0.0.36: proxy-safe local hub, fit-to-screen UI, Bedrock GPT-6 without cross-region

Cline Desktop v0.0.36 (2026-09-25) skips system HTTP(S) proxies for local hub checks, fits oversized windows to small/scaled displays, loads plugin slash commands on workspace open, and runs Bedrock GPT-6/GPT-5.6 without cross-region inference (India via in. profiles).

Cline Desktop v0.0.36: proxy-safe local hub, fit-to-screen UI, Bedrock GPT-6 without cross-region
⚡ Key Takeaways
  • •Local hub checks bypass system HTTP(S) proxies (Clash/v2ray/corporate)
  • •Oversized windows shrink to fit and center on scaled 1080p displays
  • •Plugin slash commands load on workspace open with background retry
  • •Bedrock GPT-6/GPT-5.6 work without cross-region; India uses in. profiles
  • •Artifacts: deb/rpm/dmg/exe + signatures; Providers panel scrolls model lists
Read details→
D
DeepSeek@deepseek_ai·14h ago
🛠️ Tooling

Official DeepSeek Harness Desktop App Released: Bundles Cordis Plugin Core for Daemon-Free Local Agent Workflows

DeepSeek AI (@deepseek_ai) has officially launched the standalone desktop application for DeepSeek Harness (dsh) across macOS and Windows. Designed to eliminate the engineering overhead of managing CLI background daemons and local sandbox sockets, DeepSeek Harness Desktop integrates the web UI, the Cordis plugin execution engine, and isolated sandboxing into a native local client. It delivers out-of-the-box support for visual trajectory playback, arbitrary session branching, and multi-model sub-agent delegations.

Official DeepSeek Harness Desktop App Released: Bundles Cordis Plugin Core for Daemon-Free Local Agent Workflows
⚡ Key Takeaways
  • •Eliminates complex multi-process CLI terminal setups by packaging the Cordis agent engine into a standalone GUI.
  • •Integrates native visual trajectory playback, inspecting tool execution, file diffs, and chain-of-thought traces in milliseconds.
  • •Introduces deterministic session forking, enabling developers to branch and retry failed agent steps without invalidating global context.
  • •Supports heterogeneous sub-agent delegation, seamlessly piping sub-tasks to Claude Code, OpenCode, or local Ollama instances.
Read details→
P
Pydantic@pydantic·16h ago
🛠️ Tooling

Pydantic AI 2.50.0: DecisionModel protocol, named route selection, Gemini 3.8 Live realtime

Pydantic AI 2.50.0 (2026-09-25) adds DecisionModel (TypeSafeModel implements it), named route questions with optional decision_route_threshold, Gemini 3.8 Live realtime, durable-context hooks, and audio_seconds usage pricing.

Pydantic AI 2.50.0: DecisionModel protocol, named route selection, Gemini 3.8 Live realtime
⚡ Key Takeaways
  • •DecisionModel base + TypeSafeModel implementer; ask routes by name (#8696, #8733)
  • •Opt-in decision_route_threshold replaces tool-call lean; ModelSelectionContext.prompt (#8737/#8739)
  • •Gemini 3.8 Live realtime + RequestUsage.audio_seconds (#8393/#8427)
  • •decide spans + RunContext.in_durable_context (#8698/#8723)
  • •Upgrade: pip install -U 'pydantic-ai==2.50.0'
Read details→
ADSponsored
S
Stanford NLP / DSPy@stanfordnlp·17h ago
🚀 Release

DSPy 3.4.0: Jev via TypeSafe (Noul/Choice/Score), ReAnchor calibration, native lm15 engines

DSPy 3.4.0 (2026-09-25) integrates Jev through TypeSafe with experimental Noul/Choice/Score decision types and ReAnchor calibration; the LM layer prefers native lm15 engines, adds LocalInterpreter and async ReActV2. 3.4 is the LM transition release; 3.5 is the migration deadline.

DSPy 3.4.0: Jev via TypeSafe (Noul/Choice/Score), ReAnchor calibration, native lm15 engines
⚡ Key Takeaways
  • •Jev via TypeSafe: Noul/Choice/Score with probability evidence (#10463)
  • •ReAnchor fits thresholds/cuts/weights against your metric (#10475)
  • •engine=auto prefers native lm15; register_provider for compatible HTTP
  • •LocalInterpreter for trusted CPython (not a security sandbox); RLM interpreter_factory= kw-only
  • •3.4 transition → 3.5 migration deadline; see LM migration guide
Read details→
O
Ollama@ollama·18h ago
🚀 Release

Ollama v0.40 Released: Defaults to Apple Native MLX Runner on Apple Silicon for Doubled Local Throughput

Local LLM inference runtime Ollama (@ollama) officially rolled out v0.40.0. The standout architectural milestone is the default adoption of Apple's native MLX framework on all Apple Silicon Mac devices (M1 through M5), replacing the legacy CPU/Metal compute paths. Supporting major architectures like Qwen 3.8 and Llama 3 natively, the update harnesses Apple's unified memory bandwidth, slashing time-to-first-token by 42% and nearly doubling continuous generation throughput without manual configuration.

Ollama v0.40 Released: Defaults to Apple Native MLX Runner on Apple Silicon for Doubled Local Throughput
⚡ Key Takeaways
  • •Automatically routes model executions to Apple's native MLX framework by default on all Apple Silicon Mac hardware.
  • •Fully unlocks unified memory bandwidth, boosting Qwen 3.8 and Llama 3 throughput by 85% while cutting TTFT by 42%.
  • •Requires zero manual compiler flags or configuration switches; seamless upgrade via standard ollama run commands.
  • •Cross-platform binary distributions and complete Git release changelog available directly on GitHub.
Read details→
O
OpenAI@OpenAI·19h ago
🛠️ Tooling

Codex CLI 0.157.0: GPT-6 Sol/Luna (incl. Bedrock), fullscreen transcript + background daemon by default

OpenAI Codex CLI rust-v0.157.0 (2026-09-25) adds GPT-6 Sol/Luna (incl. Amazon Bedrock), enables fullscreen transcripts and auto background-server startup by default, adds an f shortcut to fork locked conversations, and enables /import in remote and local daemon sessions.

Codex CLI 0.157.0: GPT-6 Sol/Luna (incl. Bedrock), fullscreen transcript + background daemon by default
⚡ Key Takeaways
  • •GPT-6 Sol/Luna in catalog + Amazon Bedrock (#47332/#47347)
  • •Fullscreen transcript + auto background-server default (#47178/#47179/#47318)
  • •f forks locked conversations; /import in remote/daemon sessions (#47185/#47317)
  • •Proxy + network policy hardening for realtime/search (#47101/#47389)
  • •Upgrade: npm i -g @openai/[email protected]
Read details→
H
Hugging Face@huggingface·21h ago
🛠️ Tooling

Knowledge Pull Requests (KPRs) Released: Bringing Git PR Workflows to LLM Knowledge Bases

Researchers from Johns Hopkins University have introduced Knowledge Pull Requests (KPRs, arXiv: 2609.26634), an interpretable framework for continual document authoring. Treating knowledge curation analogously to software pull requests, KPRs decouple factual claim extraction from textual diffs and automate contradiction detection, outperforming blind text regeneration in enterprise RAG systems.

Knowledge Pull Requests (KPRs) Released: Bringing Git PR Workflows to LLM Knowledge Bases
⚡ Key Takeaways
  • •Translates code review paradigms into document curation, producing interpretable claim changelogs alongside textual diffs.
  • •Employs automated factual contradiction detectors, surfacing source discrepancies with 91.4% precision.
  • •Improves grounded knowledge retention by 48% over monolithic re-generation across multilingual benchmark corpora.
  • •Preprint, evaluation datasets, and pipeline code are accessible on arXiv and Hugging Face Papers.
Read details→
H
Hugging Face@huggingface·21h ago
🛠️ Tooling

Tencent and ZJU Open-Source IterSynth: Decoupled Deep-Search Agent Architecture Outperforms Existing 8B Baselines

Tencent and Zhejiang University's REAL Lab have open-sourced IterSynth (arXiv: 2609.29444, GitHub: Tencent/IterSynth), a groundbreaking paradigm for deep-search AI agents. Breaking free from the rigid single-agent ReAct framework where planning, evidence retrieval, and synthesis are tightly coupled—which typically causes severe context saturation and attention degradation—IterSynth alternates between a specialized Planner and a Synthesizer governed by an evolving summary state. Powered by Role-Decoupled Policy Optimization (RDPO), IterSynth-8B hits an average score of 50.7 across long-horizon benchmarks like BrowseComp and Xbench-DS, outperforming all prior <=8B search agents by +4.2%.

Tencent and ZJU Open-Source IterSynth: Decoupled Deep-Search Agent Architecture Outperforms Existing 8B Baselines
⚡ Key Takeaways
  • •Decouples single-policy ReAct agent architectures into specialized Planner and Synthesizer roles.
  • •Maintains an evolving summary state that slashes token context noise from multi-hop web scraping by 70%.
  • •Introduces Role-Decoupled Policy Optimization (RDPO) combining terminal rewards with fine-grained rubric evaluations.
  • •IterSynth-8B achieves a record 50.7 score on BrowseComp and Xbench-DS, lifting zero-shot reasoning on proprietary models.
  • •Complete PyTorch training harness, benchmarks, and checkpoints open-sourced on GitHub under Tencent.
Read details→
H
Hugging Face@huggingface·21h ago
🔥 Trending

AgentKernel Released: A Trust-Native Operating System for Autonomous AI Agents with Mandatory Security Boundaries

Security and systems researchers have introduced AgentKernel (arXiv: 2609.29647), a trust-native operating system designed specifically for autonomous AI agents. As modern agents routinely ingest untrusted external data and execute privileged tools, traditional application-level guardrails fail to prevent prompt injections, memory poisoning, and tool abuse. AgentKernel establishes a mandatory, non-bypassable OS substrate structured around four architectural pillars—Identity, Perception, Cognition, and Execution—bringing deterministic capability confinement and information flow control to the semantic plane.

AgentKernel Released: A Trust-Native Operating System for Autonomous AI Agents with Mandatory Security Boundaries
⚡ Key Takeaways
  • •Translates classical operating system process isolation and privilege rings into the semantic LLM plane.
  • •Structures security across four pillars: kernel-governed identity, graduated perception, information-flow-controlled memory, and semantic execution gates.
  • •Eradicates persistent memory poisoning and unauthorized shell tool escalation caused by indirect prompt injection.
  • •Achieves a 99.1% containment rate against red-team attack vectors while constraining runtime latency overhead to 3.2%.
  • •Preprint, architecture specifications, and security proofs published openly on arXiv and Hugging Face.
Read details→
H
Hugging Face@huggingface·21h ago
📊 Benchmark

Coding Agents Revolutionize Robot Manipulation: Synthesizing Generalized Task & Motion Planning (TAMP) Programs with 95% Success

A robotics research consortium led by Tom Silver and Matteo Merler has open-sourced groundbreaking findings titled 'Coding Agents for Generalized Task and Motion Planning Problems' (arXiv: 2609.30233, GitHub: tomsilver/robocode). Overcoming the fragility of human-engineered heuristics in Task and Motion Planning (TAMP), the authors deploy autonomous coding agents (Claude Code Opus 5, Codex) to synthesize generalized symbolic-geometric programs through closed-loop simulator interaction. Across 98,000 evaluation episodes spanning 28 benchmark environments, the synthesized programs achieved 56% to 95% success rates, outperforming hand-crafted planners while using an order of magnitude less compute.

Coding Agents Revolutionize Robot Manipulation: Synthesizing Generalized Task & Motion Planning (TAMP) Programs with 95% Success
⚡ Key Takeaways
  • •Replaces handcrafted robotic heuristics by deploying autonomous coding agents to synthesize generalized TAMP programs.
  • •Agents autonomously interact with physics simulators to stress-test edge cases, calibrate dynamics, and refine logic.
  • •Evaluated across 28 simulation environments and 98,000 episodes, lifting mean task success from 47% to up to 95%.
  • •Maintains superior robustness as object counts scale while requiring an order of magnitude less compute per instance.
  • •Full benchmark harness, agent interaction logs, and prompts open-sourced on GitHub under Apache-compatible license.
Read details→
H
Hugging Face@huggingface·21h ago
🔥 Trending

Yilun Du & Lvmin Zhang Unveil WROP & 16B PWM-WROP: Teaching Physical Object Permanence to World Models

A research collective including Yilun Du, Lvmin Zhang (creator of ControlNet), and Haotian Zhang has released 'Training Object Permanence in World Models' (arXiv: 2609.28654, Project: object-permanence.world). While video generative models produce photorealistic scenes, they routinely violate physical object permanence—causing occluded objects to magically vanish, morph, or teleport. The researchers introduce WROP, a cognitive benchmark spanning 150 tasks and 1.5 million training trajectories, alongside PWM-WROP, an open-source 16B world model trained on AWS Trainium2 that secured top rank among video continuation models in blind Elo evaluations.

Yilun Du & Lvmin Zhang Unveil WROP & 16B PWM-WROP: Teaching Physical Object Permanence to World Models
⚡ Key Takeaways
  • •Addresses foundational physical hallucinations in video generators where occluded objects disappear or deform.
  • •Releases WROP, a data infrastructure spanning 150 programmatic Blender cognitive tasks and 1.5M synthetic samples.
  • •Open-sources the 16B-parameter physical world model PWM-WROP alongside native PyTorch training recipes for AWS Trainium2.
  • •Outperforms 14 competitive video generation models, ranking first among video continuation architectures in blind Elo trials.
  • •Interactive project portal, evaluation exams, model checkpoints, and datasets released openly at object-permanence.world.
Read details→
M
Mastra@mastra_ai·22h ago
🛠️ Tooling

Mastra @mastra/core 1.71.0: eager tool execution by default, plus observability discovery/feedback features

Mastra shipped @mastra/core 1.71.0 on 2026-09-24: agent.stream starts each tool as soon as its args complete (eagerToolExecution on by default), cutting idle seconds on multi-tool steps; adds planTraceAggregate(), observability getFeatures discovery/feedback flags, and RE2-backed dataset schema regexes to close a DoS vector.

Mastra @mastra/core 1.71.0: eager tool execution by default, plus observability discovery/feedback features
⚡ Key Takeaways
  • •Default: tools start when their streamed args complete; opt out with `eagerToolExecution: false` (#25005)
  • •Observability stores declare discovery + feedback via `getFeatures()` (#25008, #25020)
  • •`planTraceAggregate()`: ≤365d range, ≤1000 buckets, ≤10,000 rows (#24868)
  • •Dataset schema regexes run on RE2; unsupported lookarounds/backrefs rejected (#25044)
  • •Upgrade: `npm i @mastra/[email protected]`
Read details→
H
Hugging Face@huggingface·22h ago
📊 Benchmark

Zhejiang University Open-Sources Spatial-Interactor: Closed-Loop Physical World Spatial Reasoning with LSI-108K

Researchers from Zhejiang University's OmniAI Lab have open-sourced Spatial-Interactor alongside the LSI-108K interactive spatial reasoning dataset (arXiv: 2609.23038). Addressing severe 3D hallucination in passive vision-language models, Spatial-Interactor introduces closed-loop active perception where an agent's physical camera-control actions inform subsequent spatial inferences, fully accessible under the Apache 2.0 license.

Zhejiang University Open-Sources Spatial-Interactor: Closed-Loop Physical World Spatial Reasoning with LSI-108K
⚡ Key Takeaways
  • •Replaces passive single-image inspection with active, multi-view camera manipulation to resolve spatial occlusions.
  • •Releases LSI-108K, the first large-scale interactive spatial dataset spanning 108,000 closed-loop trajectories.
  • •Improves 3D object localization, bounding box estimation, and obstacle avoidance accuracy by 33.8% on average.
  • •Complete PyTorch training scripts, evaluation benchmarks, and checkpoints open-sourced under Apache 2.0 on GitHub.
Read details→
A
AG2@ag2ai·22h ago
🛠️ Tooling

AG2 1.1.0: TypeSafe Jev decision-only agents + reserved context strip across transports

AG2 v1.1.0 (2026-09-24) adds TypeSafe Jev decision-only agents via TypeSafeConfig + response_schema, and strips reserved ag:/a2a: context keys on A2A/AG-UI/A2UI/NLIP to block remote pre-approval of gated tools.

AG2 1.1.0: TypeSafe Jev decision-only agents + reserved context strip across transports
⚡ Key Takeaways
  • •Decision agents: TypeSafeConfig + response_schema; bool/Enum/IntEnum as question shapes; probs on ModelMessage.metadata
  • •Jev rejects tools and streaming; invalid schemas raise locally before the wire
  • •Install: pip install "ag2[typesafe]"; see TypeSafe Jev blog + docs.ag2.ai
  • •Security: strip reserved ag:/a2a: prefixes both ways on A2A/AG-UI/A2UI/NLIP
  • •Also: MessageEnqueued, MCP required-arg validation + JSON-RPC codes, event retarget fix
Read details→
H
Hugging Face@huggingface·23h ago
📊 Benchmark

Breakthrough in Self-Organizing Agent Teams: Dynamic Role Emergence Boosts Complex Task Resolution by 27.4%

Researchers have published a seminal framework titled 'Self-Organizing Agent Teams Learn to Reason Together' (arXiv: 2609.22682). Overcoming the rigidity of hardcoded hierarchical agent structures, the architecture leverages graph rewriting to allow LLM agents to dynamically self-organize sub-teams and communication topologies, improving complex problem-solving rates by 27.4%.

Breakthrough in Self-Organizing Agent Teams: Dynamic Role Emergence Boosts Complex Task Resolution by 27.4%
⚡ Key Takeaways
  • •Replaces hardcoded managerial agent hierarchies with emergent, self-organizing dynamic role specialization.
  • •Employs utility-driven graph rewriting to dynamically prune noisy communication links, improving information flow by 61%.
  • •Achieves a 27.4% performance lift over static multi-agent baseline frameworks on complex engineering benchmarks.
  • •Full benchmark harness and topological simulation codebase open-sourced on GitHub and Hugging Face Papers.
Read details→