Artificial Analysis (@ArtificialAnlys) shipped Intelligence Index v4.3: Terminal-Bench upgraded 2.1→4.0 (now mini-SWE-agent harness) and τ³-Banking replaced by AutomationBench-AA (657 private Zapier business workflows). Claude Fable 5.1 and GPT-6 Astra tie at 53; OpenAI holds most of the intelligence-vs-cost frontier across Astra reasoning efforts.

Key Takeaways

  • Benchmarks: Index v4.3 upgrades Terminal-Bench to 4.0 and adds AutomationBench-AA (private set).
  • Weighting: private-task/answer evals rise from 40% to 45%; category weights unchanged.
  • Results: Claude Fable 5.1 and GPT-6 Astra tie at 53; open-weights led by GLM-5.3 / Kimi K3 at 44.
ADSponsored