AI Coding Live News

634Stories Tracked
99Peak Impact
63424h Fresh Stories
14+Linked Entities

Elon Musk Details Imminent Grok 4.7 Launch and RL Verification Tuning

Elon Musk shared that Grok 4.7 is in final polish, explaining that RL over-penalizing token length caused models to surrender prematurely on hard problems. The team is correcting length penalties and rigor before release in days.

Key Takeaways

  • Grok 4.7 completed base pretraining, now undergoing RL reward signal fine-tuning
  • Removes token length penalties causing premature model surrender on complex reasoning
  • Bolsters multi-step verification and reflection ahead of full rollout in days
Linked Models & Agents:

Abacus AI launches open-weights Smaug Flash for personal agents at ~$0.10/$0.40 per 1M tokens on RouteLLM

Abacus AI CEO @bindureddy announced open-weights Smaug Flash, a personal-agent fine-tune on RouteLLM API at about $0.10/M input and $0.40/M outputβ€”claiming DeepSeek Flash-level performance at roughly 4x lower price, with weights on Hugging Face.

Key Takeaways

  • Open-weights Smaug Flash tuned for personal agents.
  • RouteLLM API pricing ~$0.10/M in, $0.40/M out.
  • Claims DeepSeek Flash-level performance at ~4x lower price; HF weights.

OpenAI: GPT-Rosalind leaves research preview for eligible orgs via API, Codex, ChatGPT Enterprise

OpenAI Developers: GPT-Rosalind is out of research preview for eligible orgs via API, Codex, and ChatGPT Enterpriseβ€”stronger biological reasoning across papers and experiments. Codex Life Sciences plugins connect genomic/protein/translational workflows to 50+ scientific tools.

Key Takeaways

  • GPT-Rosalind GA for eligible orgs on API, Codex, and ChatGPT Enterprise.
  • Life-science focus: evidence weighing, target analysis, experiment planning.
  • Codex Life Sciences plugins link 50+ scientific tools and genomic/protein workflows.

Artificial Analysis: Devin Fusion joins Coding Agent Index; Fable scores 62, Astra cheaper/faster

Artificial Analysis independently benchmarked Devin Fusionβ€”the first multi-model coding agent on their Coding Agent Index. Claude Fable 5.1 (xhigh)+SWE-2 scores 62; GPT-6 Astra (xhigh)+SWE-2 scores 59 with ~43% lower cost and ~31% faster completion while retaining frontier-level performance.

Key Takeaways

  • First multi-model coding agent on Artificial Analysis Coding Agent Index.
  • Fable 5.1+SWE-2 scores 62; Astra+SWE-2 scores 59.
  • Astra config ~43% cheaper and ~31% faster task completion.
ADSponsored

Claude Code ships claude plugin eval: score plugins/skills vs a no-plugin baseline

Official @ClaudeDevs: Claude Code adds `claude plugin eval`β€”author test cases, score a plugin/skill against a no-plugin baseline (terminal + HTML), then `claude update`. Docs at code.claude.com/docs/en/plugin-evals. Evals burn tokens; MCP/hooks run as configured.

Key Takeaways

  • Official claude plugin eval command to quantify plugin/skill value.
  • Test cases with plugin vs no-plugin baseline; terminal + HTML reports.
  • Requires Claude Code update; evals use tokens and run MCP/hooks.

Elon Musk: Grok 4.7 needs a few more days

Elon Musk said Grok 4.7 still needs a few more days. He suggested RL may have penalized response length too much, causing the model to give up too early on hard tasks it can do and to check its work insufficiently.

Key Takeaways

  • Elon Musk said Grok 4.7 still needs a few more days. He suggested RL may have penalized response length too much, causing the model to give up too early on hard tasks it can do and to check its work insufficiently.
Linked Models & Agents:

Cognition launches Devin CLI Fusion: frontier planner + cheap executor, ~39% lower coding-bench cost

Cognition introduced Fusion in Devin CLIβ€”an efficient frontier harness for Fable & Astra claiming ~39% lower cost across coding benchmarks. Pick a preferred model for planning and a cost-effective model for execution.

Key Takeaways

  • Fusion multi-model harness ships in Devin CLI.
  • Frontier models (e.g. Fable/Astra) for planning; cheaper models for execution.
  • Claims ~39% lower cost across coding benchmarks.

Grok Bot Integrates with Salesforce, HubSpot, and Major GTM Tooling

xAI upgraded Grok Bot for sales teams with native integrations across Salesforce, HubSpot, Gong, Clay, and Granola, enabling autonomous customer research, meeting follow-ups, and CRM workflows.

Key Takeaways

  • Deep enterprise integrations across Salesforce, HubSpot, Gong, Clay, and Granola
  • Automates CRM lead enrichment, post-call action items, and pipeline hygiene
  • Evolves Grok Bot from chatbot into an autonomous enterprise workforce agent
ADSponsored

Warp adds built-in Grok Build CLI support: rich prompts, /remote-control sessions, code review panels

Warp now has built-in support for the Grok Build CLIβ€”rich input for longer pasted prompts and multi-cursor, /remote-control to share an agent session to another device, plus file explorer and code review panels.

Key Takeaways

  • Native Grok Build CLI support inside Warp.
  • Longer prompts, multi-cursor, and /remote-control cross-device sessions.
  • File explorer and code review panels available in-session.

Cursor Unveils Projects: Persistent Orchestrator Agent for Code Workflows

Cursor launched Projects, introducing a persistent orchestrator agent in a unified thread that dispatches specialized subagents, monitors PRs/CI, and retains deep project-level memory over time.

Key Takeaways

  • Persistent master thread solves context fatigue in large-scale multi-file development
  • Orchestrator agent actively dispatches and verifies specialized coding subagents
  • Autonomously monitors PRs and CI status with proactive self-healing capabilities
Showing 1–10 of 634 stories
Go to
…