Gemini 3.7 Flash
CursorBench #3 ยท $0.75 / $3.75 intro
Live coding leaderboards, verified API cost, and the model + agent stack that actually ships. Zero paid placements.
Enter your tech stack, task, and budget to calculate the optimal setup and estimated monthly cost in 30 seconds.
Subscribe to Cursor Pro ($20/mo), with API key for overflow
Plug direct API keys for heavy daily code review without financial anxiety
Real-time tracker of price cuts, new model drops, and context expansions across the AI landscape.
Cursor said Claude Fable 5.1 is now available in the editor and is the strongest model it has run on CursorBench 3.2, scoring 73.4% at max effort. The post highlights strong self-verification, which helps it finish hard coding tasks end to endโan IDE-side landing the same day as the Fable 5.1 launch.
ClaudeDevs said the Messages API will block editing Claude's context before thinking blocks in multi-turn chats, to harden against distillation that extracts chain-of-thought. The change applies first only to new accounts on Fable 5.1 (distillation often spins up many fake accounts), with plans to extend it to all users on future releases while looking for longer-term fixes.
Anthropic's developer account said Claude Fable 5.1 is live in Claude Code and on the Claude Platform at the same list price as Fable 5 with 75% cheaper API cache reads. In the same thread it reported Terminal-Bench 4.0 scores in Claude Code: Fable 5.1 at 55.8%, ahead of Fable 5 at 42% and Opus 5 at 52.3%, and said it leads both on agentic coding and computer-use evals.
Anthropic launched Claude Fable 5.1 on September 1 (API id claude-fable-5-1) for Pro, Max, Team, and Enterprise, plus the Claude API, AWS, Google Cloud, and Microsoft Foundry. List price stays $10 input / $50 output per 1M tokens, but cache reads drop to $0.25 (75% less than Fable 5), which Anthropic estimates cuts typical workloads about 25% and highly agentic ones up to about 45%. Claude Mythos 5.1 remains limited to trusted-access programs.
Instant recommendations for model pairings, top plans, and estimated monthly costs for your exact workflow.
Fast code completion, function refactors, and full multi-language daily development.
Complex codebase indexing, sandbox execution, cross-file refactors, and chain-of-thought reasoning.
Direct pay-as-you-go API calls, self-hosted gateways, cutting token bills by up to 80% without losing quality.
Centralized billing, multi-seat governance, zero data retention for training, and SLA backing.
From unbiased pass-rate benchmarks to transparent pricing matrices and copy-paste rules.
| Rank | Model / Lab | CursorBench Pass Rate | API Input / 1M | Context Window | Action |
|---|---|---|---|---|---|
| #1 | Claude Fable 5.1Anthropic | $10.00 | 1M | Compare โ | |
| #2 | Claude Fable 5Anthropic | $10.00 | 1M | Compare โ | |
| #3 | Grok 4.6xAI | $2.00 | 500K | Compare โ | |
| #4 | Gemini 3.7 FlashGoogle | $0.75 | 1.0M | Compare โ | |
| #5 | GPT-5.6 SolOpenAI | $5.00 | 1M | Compare โ | |
| #6 | Grok 4.5xAI | $2.00 | 500K | Compare โ |
New flagships, official numbers, source-linked. Not a press-release dump โ the rows you can actually call.
CursorBench #3 ยท $0.75 / $3.75 intro
+50% coding vs GLM-5.2 on Z.ai Code Bench
Tagged coding, sorted by input $ / 1M
Models tracked per lab
Simulate monthly token volume and compare flat-rate plans vs pay-as-you-go API costs
Every page is generated from a JSON file in our public git repo. Every row links back to the vendor's own pricing or release page.
We mirror LMArena, CursorBench, Aider, LiveCodeBench and more. Their ranking, our table โ no composite score.
Overseas and domestic Code / Agent / Token plans in one matrix. Wizard, filters, and a cost calculator.
Input / output cost per million tokens, context window, status. Auto-verified daily against vendor pages.
Machine-readable /api/v1 for models, pricing, plans, and leaderboards. Built for agents, not just browsers.
All data lives in a JSON file in our public git repo. Every change is a PR, fully audited.
A daily script fetches each vendor's pricing page and confirms our numbers still appear in it.
If a vendor changes their page and verification fails, the row gets a yellow badge and the failure reason is shown.
Vendors cannot pay to be added, removed, or re-ranked. There is no advertising. The site is funded by the operating company.