AI Coding Models Leaderboard
Aggregated official benchmark mirrors and real-world coding task scenarios. 100% verified, objective, and zero paid placement.
📌 7 official boards · mirrored 1:1
Each card previews a provider's Top 5. Click "View full" for the complete board, or jump to the official page. Data is 100% from each provider — we do no aggregation or math.
Agent Score· 30 models- #1Claude Fable 5.113.85
- #2GPT-6 Astra12.39
- #3Claude Opus 511.06
- #4Claude Fable 59.09
- #5Claude Opus 4.87.57
Pass rate%· 25 models- #1Claude Fable 5.175.6%
- #2Claude Fable 572.9%
- #3Gemini 3.8 Flash71.8%
- #4Grok 4.669.9%
- #5Gemini 3.7 Flash68.4%
Coding Index· 92 models- #1Claude Fable 5.153.4
- #2GPT-6 Astra52.8
- #3Claude Opus 550.7
- #4Claude Fable 549.7
- #5Muse Spark 1.348.2
Coding score%· 37 models- #1DeepSeek-V4 (Next-Gen MoE)78.4%
- #2Claude Fable 5.174.2%
- #3Claude Fable 571.7%
- #4Claude Opus 571.7%
- #5Muse Spark 1.370.9%
Pass rate%· 33 models- #1GPT-5.588.0%
- #2o3-pro84.9%
- #3Gemini 2.5 Pro83.1%
- #4o381.3%
- #5Grok 4.679.6%
Pass@1%· 15 models- #1DeepSeek-V3100.0%
- #2GPT-4o100.0%
- #3GPT-4o mini100.0%
- #4OpenAI o3-mini (High Reasoning Coding)100.0%
- #5o4-mini100.0%
HumanEval+%· 18 models- #1o189.0%
- #2o1-mini89.0%
- #3Qwen3.6-Plus87.2%
- #4GPT-4o87.2%
- #5DeepSeek-V386.6%
LMArena AgentAgent
| # | Model | Agent Score | Value Score |
|---|---|---|---|
| #1 | Claude Fable 5.1Anthropic | 13.85 | 7.9 |
| #2 | GPT-6 AstraOpenAI | 12.39 | 7.8 |
| #3 | Claude Opus 5Anthropic | 11.06 | 5.1 |
| #4 | Claude Fable 5Anthropic | 9.09 | 7.2 |
| #5 | Claude Opus 4.8Anthropic | 7.57 | 13.6 |
| #6 | GPT-5.6 SolOpenAI | 7.51 | 12 |
| #7 | Kimi K3Moonshot AI | 6.46 | 21.5 |
| #8 | Claude Sonnet 5Anthropic | 5.88 | 31.9 |
| #9 | Hunyuan-T1Tencent | 5.23 | 434.2 |
| #10 | GPT-5.5OpenAI | 5.22 | 16.8 |
| #11 | GLM-5.2Zhipu AI | 4.56 | 53.7 |
| #12 | Muse Spark 1.3Meta | 4.23 | 67.9 |
| #13 | DeepSeek-V4-ProDeepSeek | 4.14 | 118.2 |
| #14 | Qwen3.8-MaxQwen | 3.97 | 40.8 |
| #15 | Gemini 3.8 FlashGoogle | 3.84 | 82.3 |
| #16 | Grok 4.5xAI | 3.83 | 39.7 |
| #17 | Grok 4.6xAI | 3.12 | 43 |
| #18 | GLM-5.3Zhipu AI | 2.62 | 56.6 |
| #19 | DeepSeek-V4-FlashDeepSeek | 2.17 | 326.7 |
| #20 | GPT-5.6 TerraOpenAI | 1.46 | 26.4 |
| #21 | GPT-5.6 LunaOpenAI | 0.73 | 246 |
| #22 | Qwen3.8-27BQwen | -0.08 | 240.1 |
| #23 | Claude Sonnet 4.6Anthropic | -1.27 | 18.5 |
| #24 | Qwen3.7-MaxQwen | -3.12 | 212.5 |
| #25 | MiniMax-M3MiniMax | -5.38 | 157.1 |
| #26 | Gemini 3.6 FlashGoogle | -5.41 | 60.7 |
| #27 | Gemini 3.1 ProGoogle | -5.6 | 40.9 |
| #28 | Mistral Medium 3.5Mistral AI | -10.12 | 46.1 |
| #29 | MiniMax-M2.7MiniMax | -13.22 | 77.8 |
| #30 | Gemini 3.5 Flash-LiteGoogle | -13.73 | 77.7 |
The "Cross-board ranks" column shows this model's rank on the other boards at a glance (grey = not tested there). Boards use different methodologies, so raw values are not directly comparable.
LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.
📊 7 benchmark providers
LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.
Agent Score30 modelsCursorBench is Cursor's first-party agentic coding benchmark (v3.2). Tasks come from real Cursor sessions with ambiguous, multi-file requirements. Metric: pass rate (%). Mirrored from Cursor's official cursor.com/cursorbench snapshot.
Pass rate%25 modelsArtificial Analysis Coding Index: Focused specifically on code generation, reasoning, and synthesis, fully shifted from generic overall intelligence to coding performance.
Coding Index92 modelsLiveBench Coding: Extracts the five core coding benchmarks (code completion, code generation, JavaScript, Python, and TypeScript) for pure code performance.
Coding score%37 models133 real coding tasks (bug-fix / feature) across 6 languages. Metric: pass_rate_2.
Pass rate%33 modelsContinuously updated competitive-programming problems. Metric: Pass@1.
Pass@1%15 modelsAll rankings and numbers come directly from official sources; we only do name matching and aggregate presentation, no weighting or re-ranking. Boards use different methodologies, so do not compare raw values across boards.