@ArtificialAnlys published AA-Briefcase results: Grok 4.7 sits behind only Anthropic models, ranking just after Opus 5 at about half the cost per task. Versus Grok 4.6 on public AA-Briefcase-Lite due-diligence decks, Analytical Quality Elo rose 1698 to 1994 while Presentation Elo slipped 1531 to 1499. Example deck API cost: Grok 4.7 (xhigh) about $8 vs Grok 4.6 (xhigh) about $4.40.
Key Takeaways
- ✓AA-Briefcase: Grok 4.7 is behind only Anthropic, just after Opus 5, at ~50% cost/task.
- ✓Vs 4.6: Analytical Quality Elo 1698→1994; Presentation Elo 1531→1499.
- ✓Example decks: Grok 4.7 (xhigh) ~$8 vs Grok 4.6 (xhigh) ~$4.40.

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.