O
OpenCode@opencode·3h ago
🛠️ Tooling

OpenCode adds GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 the same day

Open-source coding agent OpenCode (@opencode) said GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 are now available in OpenCode, wiring the same-day OpenAI and Anthropic launches into an open coding-agent workflow.

⚡ Key Takeaways
  • GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 are live in OpenCode.
  • Covers the same-day OpenAI and Anthropic model launches.
  • Same-day access for open-source coding-agent users.
Read details
O
OpenAI@OpenAI·4h ago
🚀 Release

OpenAI launches GPT-6 Sol and GPT-6 Luna: faster, cheaper, API prices 50% below GPT-5.6 promo

OpenAI launches GPT-6 Sol and GPT-6 Luna: faster, cheaper, API prices 50% below GPT-5.6 promo

OpenAI (@OpenAI) launched GPT-6 Sol and GPT-6 Luna in the GPT-6 family, carrying forward GPT-6 Astra advances in professional work, factuality, coding, computer use, and alignment into faster, more affordable models for work at scale. With more efficient caching and inference, Sol/Luna API prices are 50% lower than GPT-5.6 promotional pricing. OpenAI president Greg Brockman (@gdb) separately highlighted the coding and computer-use gains.

⚡ Key Takeaways
  • GPT-6 Sol and Luna join the GPT-6 lineup today—Astra-class strengths in faster, cheaper models for scale.
  • API prices are 50% below GPT-5.6 promotional pricing after caching and inference efficiency gains.
  • @gdb confirms coding, computer-use, factuality, and alignment gains carried from Astra.
Read details
C
Cursor@cursor_ai·5h ago
🚀 Release

Cursor ships Claude Opus 5.5: tops CursorBench at 57.8% (Max), ~40% cheaper per task than Opus 5

Cursor ships Claude Opus 5.5: tops CursorBench at 57.8% (Max), ~40% cheaper per task than Opus 5

Cursor (@cursor_ai) announced Claude Opus 5.5 is available in Cursor. Officially it is the new top model on CursorBench at 57.8% (Max) and costs about 40% less per task than Opus 5, wiring Anthropic’s same-day Opus 5.5 launch into a leading AI IDE.

⚡ Key Takeaways
  • Claude Opus 5.5 is live in Cursor.
  • Tops CursorBench at 57.8% (Max).
  • About 40% cheaper per task than Opus 5.
Read details
A
Anthropic@AnthropicAI·6h ago
🚀 Release

Anthropic says Claude Opus 5.5 is available today

Anthropic says Claude Opus 5.5 is available today

Anthropic (@AnthropicAI) confirmed Claude Opus 5.5 is available today; see the same-day @claudeai launch thread. First Claude 5.5 model, Fable 5.1-level on most tasks, ~40% cheaper than Opus 5 on typical workloads.

⚡ Key Takeaways
  • @AnthropicAI confirms Opus 5.5 is available today.
  • First model in the Claude 5.5 family; Fable 5.1-level on most tasks.
  • About 40% cheaper than Opus 5 on typical workloads.
Read details
ADSponsored
@
@AnthropicAI@AnthropicAI·6h ago
🚀 Release

Anthropic says Claude Opus 5.5 is available today

Anthropic confirmed Claude Opus 5.5 is available. It is the first model in the new Claude 5.5 family, performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.

⚡ Key Takeaways
  • @AnthropicAI confirms Opus 5.5 is available today.
  • First model in the Claude 5.5 family; Fable 5.1-level on most tasks.
  • About 40% cheaper than Opus 5 on typical workloads.
Read details
C
Claude@claudeai·6h ago
🚀 Release

Anthropic ships Claude Opus 5.5: Fable 5.1-level on most tasks, ~40% cheaper than Opus 5

Anthropic ships Claude Opus 5.5: Fable 5.1-level on most tasks, ~40% cheaper than Opus 5

Anthropic / Claude (@claudeai; confirmed by @AnthropicAI) launched Claude Opus 5.5 today, the first model in the new Claude 5.5 family. Officially it matches Claude Fable 5.1 on most tasks, and is a major step up from Opus 5 on agentic coding, computer use, and knowledge work. At default settings typical workloads cost about 40% less than Opus 5. Pro, Max, and Team five-hour usage limits are increased, with a bankable rate-limit reset for subscribers.

⚡ Key Takeaways
  • Opus 5.5 opens the Claude 5.5 family and is available today; most tasks match Claude Fable 5.1.
  • Leads Opus 5 on agentic coding, computer use, and knowledge work; ~40% cheaper on typical default workloads.
  • Higher Pro/Max/Team five-hour usage caps plus a bankable rate-limit reset for subscribers.
Read details
E
Elon Musk@elonmusk·7h ago
🛠️ Tooling

Elon: Grok 4.7 doing real Tesla engineering; Build harness gains self-validation and long-horizon monitoring

Elon Musk (@elonmusk) highlighted that Grok 4.7 is doing real engineering work at Tesla. The quoted post (@yunta_tsai) says Grok Build teams spent weeks tuning the 4.7 harness for cleaner prose, stronger self-validation, better thinking efficiency, long-horizon monitoring, and VLM/multimodal tasks—and that many FSD and Cybercab features were personally shipped via overnight agents.

⚡ Key Takeaways
  • Elon says Grok 4.7 is doing real engineering work at Tesla.
  • Build harness upgrades: prose quality, self-validation, thinking efficiency, long-horizon monitoring, VLM/multimodal.
  • Quoted claim: many FSD and Cybercab features shipped via overnight agents.
Read details
D
DHH@dhh·10h ago
🛠️ Tooling

DHH: Alibaba Cloud joins Omacom as founding patron with $3M; Omarchy coming to Qwen Book

DHH: Alibaba Cloud joins Omacom as founding patron with $3M; Omarchy coming to Qwen Book

DHH (@dhh) announced Alibaba Cloud as a Founding Corporate Patron of the Omacom Foundation: $3 million in funding, collaboration on Omarchy China, and bringing Omarchy to the newly announced Qwen Book. He argues agentic computers need a native agentic OS. Details: omarchy.org news post.

⚡ Key Takeaways
  • Alibaba Cloud commits $3M as an Omacom Foundation founding corporate patron.
  • Collaboration covers Omarchy China and shipping Omarchy onto the new Qwen Book.
  • Thesis: agentic computers need a native agentic OS (Omarchy).
Read details
ADSponsored
K
Kimi.ai@Kimi_Moonshot·11h ago
🛠️ Tooling

Moonshot ships Kimi Browser Extension: sidebar navigate/fill forms, record steps as Skills

Kimi.ai (@Kimi_Moonshot) launched the Kimi Browser Extension (formerly Kimi WebBridge): chat with Kimi in a browser sidebar to navigate sites, fill forms, and finish tasks. Record repetitive steps once as a Skill and let Kimi replay them next time. Available now at kimi.com/products/kimi-browser-extension and the Chrome Web Store.

⚡ Key Takeaways
  • Sidebar chat drives site navigation, form filling, and browser task completion.
  • Record repetitive flows once as Skills for Kimi to replay later.
  • Live on the product page and Chrome Web Store (upgrade from WebBridge).
Read details
T
Tencent Hy@TencentHunyuan·13h ago
📊 Benchmark

Tencent Hunyuan launches WebCraftBench: coding agents use live sites with coverage-guided scoring

Tencent Hunyuan (@TencentHunyuan) introduced WebCraftBench for the common failure mode where an agent claims a site is done while the homepage errors, buttons overlap, or an unrequested login flow appears. Agents operate the live app; coverage-guided exploration finds unreached paths; aesthetics, usability, and request match are scored. On 197 human-validated pairs it matches human preference 85.3%. Paper: https://arxiv.org/abs/2609.15387

⚡ Key Takeaways
  • Agents are evaluated on live apps, not only static screenshots or code diffs.
  • Coverage-guided exploration finds pages and flows that were never reached.
  • Matches human preference 85.3% on 197 validated pairs; paper arXiv:2609.15387.
Read details
X
X Freeze@XFreeze·17h ago
📉 Price Cut

Grok 4.7 Fast live in Grok Build and Cursor: ~2× output speed at ~2× token rates

Grok 4.7 Fast live in Grok Build and Cursor: ~2× output speed at ~2× token rates

@XFreeze reports Grok 4.7 Fast is live in Grok Build and Cursor: the same Grok 4.7 model on faster infra with roughly 2× output speed, at about 2× standard token rates. It is not in Grok Build’s free tier and is not on the public xAI API. Screenshots show Cursor’s “Grok 4.7 High Fast” picker and Grok Build’s Fast variant note.

⚡ Key Takeaways
  • Grok 4.7 Fast is available only in Grok Build and Cursor so far—not on the public xAI API.
  • Same model intelligence with ~2× output speed at ~2× standard Grok 4.7 token rates.
  • Excluded from Grok Build’s free tier; aimed at coding agents that trade higher token cost for faster inference.
Read details
K
Kimi.ai@Kimi_Moonshot·19h ago
🛠️ Tooling

Moonshot: Kimi K3 now on Amazon Bedrock with coding/agent workflows and explicit prompt caching

Moonshot: Kimi K3 now on Amazon Bedrock with coding/agent workflows and explicit prompt caching

@Kimi_Moonshot announced Kimi K3 is now on Amazon Bedrock for coding, document analysis, and extended agent workflows, with Bedrock access, encryption, and auditing controls, plus explicit prompt caching. AWS published a Bedrock model card so teams can call K3 inside the AWS stack.

⚡ Key Takeaways
  • Kimi K3 is live on Amazon Bedrock.
  • Targets coding, document analysis, and extended agent workflows with Bedrock access/encryption/auditing.
  • Explicit prompt caching supported; start via the AWS Bedrock model card.
Read details
Z
Z A D D Y@Zaddyzaddy·20h ago
📊 Benchmark

BugBunny VulnPR-100: Xiaomi MiMo-V2.6-Pro tops open weights (46/100 for $8.27), matches GPT-6 Astra at ~22x lower review cost

BugBunny VulnPR-100: Xiaomi MiMo-V2.6-Pro tops open weights (46/100 for $8.27), matches GPT-6 Astra at ~22x lower review cost

@Zaddyzaddy published @BugBunny_ai VulnPR-100 results: on 100 PRs with known vulns (up to one hour each), MiMo-V2.6-Pro found 46 for $8.27 — best open-weight score, matching GPT-6 Astra at roughly 22× lower review cost. MiMo-V2.6-Flash found 41 for $3.47. That is about $0.18 per vuln found for Pro and $0.08 for Flash.

⚡ Key Takeaways
  • MiMo-V2.6-Pro: 46/100 on VulnPR-100 for $8.27, best open-weight.
  • Matches GPT-6 Astra at ~22× lower review cost.
  • MiMo-V2.6-Flash: 41/100 for $3.47; ~$0.18/$0.08 per vuln (Pro/Flash).
Read details
A
Artificial Analysis@ArtificialAnlys·21h ago
🚀 Release

Artificial Analysis: StepFun Step 5 Preview scores 44 on Intelligence Index at ~2.8x lower cost than Kimi K3

Artificial Analysis: StepFun Step 5 Preview scores 44 on Intelligence Index at ~2.8x lower cost than Kimi K3

@ArtificialAnlys evaluated StepFun Step 5 Preview (600B-total / 27B-active MoE): Artificial Analysis Intelligence Index 44, matching Kimi K3 (max) at about $0.72 per Index task (~2.8x cheaper). Pricing is about $1/$2.70 per 1M input/output tokens with $0.05/M cached input; 1M context; open weights planned for October 15. It still lags peers on agentic evals such as GDPval-AA and Terminal-Bench 4.0.

⚡ Key Takeaways
  • AA Intelligence Index 44, tied with Kimi K3 (max); ~$0.72/task (~2.8x cheaper).
  • Pricing ~$1/$2.70 per 1M in/out with $0.05/M cached input; 600B/27B MoE; 1M context.
  • Open weights planned Oct 15; still trails peers on GDPval-AA and Terminal-Bench 4.0.
Read details
@
@elonmusk@elonmusk·21h ago
🔥 Trending

Elon Musk announces Grok 4.7

Elon Musk announces Grok 4.7

Elon Musk posted an announcement of Grok 4.7 with an image, signaling the latest Grok model from xAI.

⚡ Key Takeaways
  • Elon Musk posted an announcement of Grok 4.7 with an image, signaling the latest Grok model from xAI.
Read details
@
@elonmusk@elonmusk·21h ago
🔥 Trending

Grok 4.7 announced

Grok 4.7 announced

Elon Musk posted “Grok 4.7.” Same-day discussion said the model improves on Grok 4.6 for agentic coding and knowledge work, and that it works best with the xAI Build harness.

⚡ Key Takeaways
  • Elon Musk posted “Grok 4.7.” Same-day discussion said the model improves on Grok 4.6 for agentic coding and knowledge work, and that it works best with the xAI Build harness.
Read details
@
@elonmusk@elonmusk·21h ago
🔥 Trending

Grok 4.7 launches as a faster, lower-cost frontier model

Grok 4.7 launches as a faster, lower-cost frontier model

Elon Musk announced Grok 4.7. He described it as a strong mix of intelligence, speed, and low cost, said it works especially well with xAI's Build harness, and positioned it as a competitive everyday workhorse for agentic coding.

⚡ Key Takeaways
  • Elon Musk announced Grok 4.7. He described it as a strong mix of intelligence, speed, and low cost, said it works especially well with xAI's Build harness, and positioned it as a competitive everyday workhorse for agentic coding.
Read details
@
@elonmusk@elonmusk·21h ago
🔥 Trending

Grok 4.7 Launches as a Fast, Low-Cost Everyday Workhorse

Grok 4.7 Launches as a Fast, Low-Cost Everyday Workhorse

Elon Musk announced Grok 4.7. The model is a notable upgrade over Grok 4.6 at the same price and speed, working longer on hard tasks, checking its work more carefully, and shipping with the strongest safeguards to date. Musk said it places SpaceXAI third behind Anthropic and OpenAI for agentic coding, while remaining faster and cheaper. It is available in Cursor, Grok Build, and the Grok API.

⚡ Key Takeaways
  • Elon Musk announced Grok 4.7. The model is a notable upgrade over Grok 4.6 at the same price and speed, working longer on hard tasks, checking its work more carefully, and shipping with the strongest safeguards to date. Musk said it places SpaceXAI third behind Anthropic and OpenAI for agentic coding, while remaining faster and cheaper. It is available in Cursor, Grok Build, and the Grok API.
Read details
V
Vercel@vercel·23h ago
📉 Price Cut

Vercel: Grok 4.7 on AI Gateway with 40% off through September 27

@vercel announced Grok 4.7 is available on Vercel AI Gateway with 40% off through September 27, and can be tried with v0, eve, and fx. For developers building apps and coding agents on the Vercel stack, this lowers the cost of trying Grok 4.7.

⚡ Key Takeaways
  • Grok 4.7 is live on Vercel AI Gateway.
  • 40% off on AI Gateway through September 27.
  • Try it with v0, eve, and fx.
Read details
A
Artificial Analysis@ArtificialAnlys·23h ago
📊 Benchmark

Artificial Analysis: Grok 4.7 ranks just behind Anthropic on AA-Briefcase at ~50% Opus 5 cost/task

Artificial Analysis: Grok 4.7 ranks just behind Anthropic on AA-Briefcase at ~50% Opus 5 cost/task

@ArtificialAnlys published AA-Briefcase results: Grok 4.7 sits behind only Anthropic models, ranking just after Opus 5 at about half the cost per task. Versus Grok 4.6 on public AA-Briefcase-Lite due-diligence decks, Analytical Quality Elo rose 1698 to 1994 while Presentation Elo slipped 1531 to 1499. Example deck API cost: Grok 4.7 (xhigh) about $8 vs Grok 4.6 (xhigh) about $4.40.

⚡ Key Takeaways
  • AA-Briefcase: Grok 4.7 is behind only Anthropic, just after Opus 5, at ~50% cost/task.
  • Vs 4.6: Analytical Quality Elo 1698→1994; Presentation Elo 1531→1499.
  • Example decks: Grok 4.7 (xhigh) ~$8 vs Grok 4.6 (xhigh) ~$4.40.
Read details