AI Coding Models Leaderboard
Aggregated official benchmark mirrors and real-world coding task scenarios. 100% verified, objective, and zero paid placement.
📌 7 official boards · mirrored 1:1
Each card previews a provider's Top 5. Click "View full" for the complete board, or jump to the official page. Data is 100% from each provider — we do no aggregation or math.
Agent Score· 30 models- #1Claude Fable 5.114.51
- #2GPT-6 Astra12.55
- #3Claude Opus 511.42
- #4Claude Fable 59.15
- #5GPT-5.6 Sol7.75
Pass rate%· 25 models- #1Claude Fable 5.175.6%
- #2Claude Fable 572.9%
- #3Gemini 3.8 Flash71.8%
- #4Grok 4.669.9%
- #5Gemini 3.7 Flash68.4%
Coding Index· 91 models- #1Claude Fable 5.153.4
- #2GPT-6 Astra52.8
- #3Claude Opus 550.7
- #4Claude Fable 549.7
- #5Muse Spark 1.348.2
Coding score%· 36 models- #1Claude Fable 5.174.2%
- #2Claude Fable 571.7%
- #3Claude Opus 571.7%
- #4Muse Spark 1.370.9%
- #5Kimi K369.9%
Pass rate%· 33 models- #1GPT-5.588.0%
- #2o3-pro84.9%
- #3Gemini 2.5 Pro83.1%
- #4o381.3%
- #5Grok 4.679.6%
Pass@1%· 15 models- #1DeepSeek-V3100.0%
- #2GPT-4o100.0%
- #3GPT-4o mini100.0%
- #4OpenAI o3-mini (High Reasoning Coding)100.0%
- #5o4-mini100.0%
HumanEval+%· 18 models- #1o189.0%
- #2o1-mini89.0%
- #3Qwen3.6-Plus87.2%
- #4GPT-4o87.2%
- #5DeepSeek-V386.6%
LMArena AgentAgent
| # | Model | Agent Score | Value Score |
|---|---|---|---|
| #1 | Claude Fable 5.1Anthropic | 14.51 | 7.9 |
| #2 | GPT-6 AstraOpenAI | 12.55 | 7.8 |
| #3 | Claude Opus 5Anthropic | 11.42 | 5.1 |
| #4 | Claude Fable 5Anthropic | 9.15 | 7.2 |
| #5 | GPT-5.6 SolOpenAI | 7.75 | 11.9 |
| #6 | Claude Opus 4.8Anthropic | 7.69 | 13.6 |
| #7 | Kimi K3Moonshot AI | 6.59 | 21.4 |
| #8 | Claude Sonnet 5Anthropic | 6.31 | 31.8 |
| #9 | Hunyuan-T1Tencent | 5.37 | 429 |
| #10 | GPT-5.5OpenAI | 5.33 | 16.7 |
| #11 | GLM-5.2Zhipu AI | 4.64 | 53.4 |
| #12 | DeepSeek-V4-ProDeepSeek | 4.43 | 117.8 |
| #13 | Gemini 3.8 FlashGoogle | 4.01 | 81.9 |
| #14 | Grok 4.5xAI | 3.93 | 39.5 |
| #15 | Qwen3.8-MaxQwen | 3.9 | 40.4 |
| #16 | Grok 4.6xAI | 3.24 | 42.9 |
| #17 | GLM-5.3Zhipu AI | 2.56 | 56.2 |
| #18 | DeepSeek-V4-FlashDeepSeek | 2.22 | 324.5 |
| #19 | GPT-5.6 TerraOpenAI | 1.47 | 26.2 |
| #20 | GPT-5.6 LunaOpenAI | 1.11 | 245.7 |
| #21 | Qwen3.8-27BQwen | 0.14 | 238.9 |
| #22 | Claude Sonnet 4.6Anthropic | -1.05 | 18.5 |
| #23 | Muse Spark 1.3Meta | -1.56 | 61.1 |
| #24 | Qwen3.7-MaxQwen | -3.09 | 59.2 |
| #25 | MiniMax-M3MiniMax | -5.26 | 156 |
| #26 | Gemini 3.6 FlashGoogle | -5.3 | 60.2 |
| #27 | Gemini 3.1 ProGoogle | -5.34 | 40.8 |
| #28 | Mistral Medium 3.5Mistral AI | -9.35 | 47.8 |
| #29 | MiniMax-M2.7MiniMax | -13.01 | 77.1 |
| #30 | Gemini 3.5 Flash-LiteGoogle | -13.43 | 77.7 |
The "Cross-board ranks" column shows this model's rank on the other boards at a glance (grey = not tested there). Boards use different methodologies, so raw values are not directly comparable.
LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.
📊 7 benchmark providers
LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.
Agent Score30 modelsCursorBench is Cursor's first-party agentic coding benchmark (v3.2). Tasks come from real Cursor sessions with ambiguous, multi-file requirements. Metric: pass rate (%). Mirrored from Cursor's official cursor.com/cursorbench snapshot.
Pass rate%25 modelsArtificial Analysis Coding Index: Focused specifically on code generation, reasoning, and synthesis, fully shifted from generic overall intelligence to coding performance.
Coding Index91 modelsLiveBench Coding: Extracts the five core coding benchmarks (code completion, code generation, JavaScript, Python, and TypeScript) for pure code performance.
Coding score%36 models133 real coding tasks (bug-fix / feature) across 6 languages. Metric: pass_rate_2.
Pass rate%33 modelsContinuously updated competitive-programming problems. Metric: Pass@1.
Pass@1%15 modelsAll rankings and numbers come directly from official sources; we only do name matching and aggregate presentation, no weighting or re-ranking. Boards use different methodologies, so do not compare raw values across boards.