Nous Research (@NousResearch) reported that on Sep 2 @Teknium asked Hermes Agent to clean ~1.06M lines of Python; 1,393 parallel subagents finished in 19 hours, shrinking the repo to ~698k lines (−34.4%), collapsing 37 god files to 6, at about $19,302 run cost versus nearly $2M estimated eng hours. Details: nousresearch.com/refactoring-hermes-with-1393-agents.
⚡ Key Takeaways
•Hermes Agent spun 1,393 subagents and finished a million-line refactor in 19 hours.
•LOC fell about 34.4% (~1.06M to ~698k); 37 god files became 6.
•Run cost about $19,302 versus nearly $2M estimated engineering hours.
Artificial Analysis (@ArtificialAnlys) updated its Speech to Speech Index: Google Gemini 3.8 Live Extended Thinking (High) debuts at #1 with 82.6, ahead of GPT-Live-1 (Astra medium) at 81.5 and Grok Voice Think Fast 2.0 High at 81.3; standard Gemini 3.8 Live lands #5 at 76.0. The same drop covers Speech Agent Arena, Tau Voice (Extended Thinking leads at 68.6%), Big Bench Audio, time-to-first-audio, and input-audio cost ($0.84/hr standard, $3.50/hr Extended Thinking).
⚡ Key Takeaways
•Gemini 3.8 Live Extended Thinking (High) leads Speech to Speech Index at 82.6, above GPT-Live-1 Astra at 81.5.
•On Tau Voice, Extended Thinking leads at 68.6%; standard Gemini 3.8 Live scores 30.1%.
•Standard input audio is about $0.84/hr (cheapest in the Index); Extended Thinking about $3.50/hr, still below Astra at $5.83.
LLM evaluator Vals AI documented GPT-6 Astra completing an unusually long Minecraft task: building a semi-automatic blaze farm, collecting blaze rods, and reaching a warped forest for ender pearls. A creeper then destroyed its chest and bed, exposing practical weaknesses in critical-item protection and recovery.
⚡ Key Takeaways
•GPT-6 Astra built a semi-automatic blaze farm and collected six blaze rods
•It reached a warped forest, defeated multiple endermen, and collected three ender pearls
•The lost inventory highlights remaining gaps in persistent memory, risk management, and recovery planning
Arena updated the Image-to-WebDev leaderboard: GPT-6 Astra (Max) leads at 1733 points, 129 above GPT-5.6 Sol (xHigh); Claude Fable 5.1 (Max) is second at 1710. The board compares newly evaluated model releases on image-to-webdev delivery.
⚡ Key Takeaways
•GPT-6 Astra (Max) leads Image-to-WebDev at 1733 points.
•Claude Fable 5.1 (Max) ranks second at 1710.
•Four newly evaluated model releases landed on the board.
OpenAI engineer Charlie Marsh announced a renewal of Codex for OSS, offering open-source maintainers $100 Pro plans and doubling the grant pool from 5,000 to 10,000. The program includes AI code review and security tooling for critical projects, expanding agent-assisted engineering across open source.
⚡ Key Takeaways
•The Pro grant pool doubles from 5,000 to 10,000 places, and previous recipients may re-apply
•The $100 Pro plans target open-source maintainers building and maintaining critical projects
•The program includes AI code review and security tooling for practical software engineering
Claude Code Changelog (@ClaudeCodeLog) announced Claude Code 2.1.273 with 64 CLI changes. Highlights: fork Remote-Control sessions into local background sessions while preserving state; notify when an MCP server disconnects and reconnection fails with /mcp diagnostics; auto mode no longer pauses for Artifact uploads in cloud/Remote-Control sessions. npm latest confirms 2.1.273.
⚡ Key Takeaways
•2.1.273 ships 64 CLI changes and is live on npm.
•Remote-Control sessions can be forked into local background sessions with state preserved.
•MCP disconnect failures notify via /mcp; Artifact uploads no longer pause auto mode.
Keewano launched KeewanoDB alongside a $12M seed round. The database is designed for machine reasoning by preserving ordered event context for users, devices, and agents, so agents can read coherent history instead of reconstructing it from conventional tables.
⚡ Key Takeaways
•KeewanoDB treats ordered event context as a first-class data structure for agents
•Keewano announced a $12M seed round led by Hetz Ventures
•Sequence-aware storage aims to reduce the engineering cost of reconstructing agent state and context
OpenAI Developers published an official demo showing GPT-6 Astra helping Cognition's Devin support the claim that code works with tests before a team ships. The workflow puts test evidence into the coding agent's delivery process instead of relying only on generated output.
⚡ Key Takeaways
•GPT-6 Astra is paired with Devin for pre-ship code validation.
•The workflow uses test evidence to support that code works rather than relying only on generated output.
•The official OpenAI Developers post explicitly features Cognition's Devin.
Perplexity (@perplexity_ai) published CobbleDB: a key-value store built in two months by two engineers and hundreds of proactive always-on AI agents for Perplexity search web-content serving. After replacing DynamoDB, median batch-read latency fell from 31.4 ms to 5.60 ms and p99 from 123 ms to 24.2 ms; they plan to open-source it for other AI search teams.
⚡ Key Takeaways
•Two engineers plus hundreds of AI agents built production CobbleDB in two months.
TechCrunch reports WhatsApp Business launched an official MCP server so developers can use coding agents like Claude, Cursor, Codex, and ChatGPT for setup, messaging templates, testing, and troubleshooting—putting WhatsApp messaging behind the standard MCP tool interface.
⚡ Key Takeaways
•WhatsApp Business now has an official MCP server for coding agents.
•Claude, Cursor, Codex, and ChatGPT can handle setup, templates, testing, and troubleshooting.
•Customer messaging joins the same MCP toolchain as Slack/GitHub-style tools.
ChatGPT announced that on October 14, 2026, GPT-5.5 will leave ChatGPT, ChatGPT Work, and Codex across all plans; Codex users should switch to GPT-5.6 Sol or GPT-6 Astra. OpenAI Developers clarified that GPT-5.5 remains available via the OpenAI API Platform and in Codex sessions authenticated with an API key.
⚡ Key Takeaways
•Oct 14: GPT-5.5 exits ChatGPT, ChatGPT Work, and Codex across all plans.
•Codex plan users should move to GPT-5.6 Sol or GPT-6 Astra.
•GPT-5.5 stays on the API Platform and in API-key-authenticated Codex sessions.
Google DeepMind and Google AI Studio introduced two live dialogue audio models: Gemini 3.8 Live (scale, speed, cost efficiency) and Gemini 3.8 Live Extended Thinking (higher-complexity multi-step reasoning). Both add near real-time visual understanding, automatic detection across 97 languages, and background tool calling without breaking conversation flow. Developers can build in public preview via AI Studio and the Gemini API; consumer rollouts include Search Live and Gemini App, with enterprise private preview.
⚡ Key Takeaways
•Two live audio models: 3.8 Live (scale/speed/cost) and 3.8 Live Extended Thinking (complex multi-step reasoning).
•Near real-time vision, 97-language auto-detection, and background tool calling without breaking chat flow.
•Developers: public preview in AI Studio and Gemini API; consumer Search Live/Gemini App; enterprise private preview.
Factory (@FactoryAI) announced a $200M raise at a $5B valuation (over $400M total funding), more than tripling its $1.5B April valuation, to scale self-improving software development and autonomous coding agents (Droids) for the enterprise.
⚡ Key Takeaways
•Factory raised $200M at a $5B valuation.
•Total funding tops $400M; valuation more than triples the $1.5B April mark.
•Capital targets enterprise self-improving software development and coding agents.
Vercel v0 (@v0) said the product is now model-agnostic: through Vercel AI Gateway you can choose frontier, cheap, open, or fast models—including Claude, GPT, Kimi, GLM, Grok, and DeepSeek—to build apps in v0 without locking to one model vendor.
⚡ Key Takeaways
•v0 is model-agnostic through Vercel AI Gateway—no single-vendor lock-in.
•Choose Claude, GPT, Kimi, GLM, Grok, DeepSeek and other frontier/cheap/open/fast models.
•Announced by @v0 so the build workflow can swap models and cost profiles per task.
Cognition (@cognition) announced Devin now runs on its own Mac VM: it can build and test apps in an iOS Simulator, send a screen recording over Slack, and share a TestFlight link so humans can try the build. That extends the coding-agent loop from writing code to simulator validation and distribution.
⚡ Key Takeaways
•Devin gets its own Mac VM and can build/test in the iOS Simulator.
•It can Slack a screen recording and hand over a TestFlight link.
•Announced by @cognition, closing the mobile verify-and-distribute loop.
Gensyn released open-1b, a foundation model built around auditable and replayable training. Checkpoints can be rerun with Gensyn custom kernels across NVIDIA and Apple hardware, aiming to make model training verifiable rather than opaque.
⚡ Key Takeaways
•open-1b makes auditable and replayable training a first-class product goal
•Gensyn says checkpoints can be rerun with custom kernels across NVIDIA and Apple hardware
•Verifiable training could improve model supply-chain trust, research reproducibility, and enterprise adoption
TestingCatalog (@testingcatalog) reports Showly now lets users connect coding agents and turn agent-built pages into shareable links, with a private preview before public crawlable publish. It works with Claude Code, Codex, Cursor, OpenClaw, Hermes, and other MCP-compatible agents, and publishing requires human permission. Version history, deployment logs, and an agent activity log are included; the free plan offers 300 credits/month, unlimited live sites and custom domains, plus simple rollbacks.
⚡ Key Takeaways
•Connects Claude Code, Codex, Cursor, OpenClaw, Hermes, and other MCP-compatible coding agents.
•Private preview before public crawlable publish; human permission required to ship.
•Free plan: 300 credits/month, unlimited live sites and custom domains, with version history, deploy logs, and rollbacks.
Artificial Analysis (@ArtificialAnlys) published its Speech to Speech Index: OpenAI GPT-Live-1 with Astra (medium) as the delegated backend debuts at #1 with 81.5, while Sol (low) scores 80.1 at #3, with Grok Voice Think Fast 2.0 High between them at 81.3. GPT-Live-1 is a full-duplex speech model that can delegate reasoning and tool use to a backend text model mid-conversation. The same drop covers Speech Agent Arena, Tau Voice, Big Bench Audio, plus cost and time-to-first-audio.
⚡ Key Takeaways
•GPT-Live-1 (Astra medium) leads Speech to Speech Index at 81.5; Sol low is #3 at 80.1.
•On Tau Voice, Astra medium (67.9%) and Sol low (59.3%) take the top two agentic spots.
•Cost is about $5.83/$4.47 per hour of input audio incl. backend; TTFA about 1.34s/1.24s vs Grok Voice 0.70s.
Claude Code Changelog (@ClaudeCodeLog) announced Claude Code 2.1.272: the official changelog lists bug fixes and reliability improvements aimed at fewer crashes and more stable common workflows, shipping hours after the feature-heavy 2.1.271.
⚡ Key Takeaways
•2.1.272 is a CLI stability/reliability patch.
•Aimed at fewer crashes in common workflows.
•Follows the feature-heavy 2.1.271 by a few hours.
CoreSpeed (@CoreSpeedHQ) launched: connect apps through one MCP endpoint with built-in cross-agent memory, media generation, web and social research tools, plus agent identities, budgets, and activity logs. Start at https://corespeed.io.
⚡ Key Takeaways
•One MCP endpoint wires apps and toolchains.
•Built-in cross-agent memory, media gen, and web/social research.
•Includes agent identities, budgets, and activity logs.