2026 Live Verified

AI Coding Models Leaderboard

Aggregated official benchmark mirrors and real-world coding task scenarios. 100% verified, objective, and zero paid placement.

7Boards
95Models ranked
777Evaluation Tasks
September 2026Period
2026-09-10Last update
We fetch raw rankings from 6 public benchmark providers directly — no re-weighting, no composite score. Click "Open official board" on each card to go to the source.

📌 7 official boards · mirrored 1:1

Each card previews a provider's Top 5. Click "View full" for the complete board, or jump to the official page. Data is 100% from each provider — we do no aggregation or math.

LMArena Agent
Human preference on real agentic coding sessions
Agent
Metric:Agent Score· 30 models
  1. #1Claude Fable 5.114.51
  2. #2GPT-6 Astra12.55
  3. #3Claude Opus 511.42
  4. #4Claude Fable 59.15
  5. #5GPT-5.6 Sol7.75
Official ↗
CursorBench
Cursor's first-party agentic coding benchmark (v3.2)
Agent
Metric:Pass rate%· 25 models
  1. #1Claude Fable 5.175.6%
  2. #2Claude Fable 572.9%
  3. #3Gemini 3.8 Flash71.8%
  4. #4Grok 4.669.9%
  5. #5Gemini 3.7 Flash68.4%
Official ↗
Artificial Analysis
Official Coding Index for software engineering
Coding
Metric:Coding Index· 91 models
  1. #1Claude Fable 5.153.4
  2. #2GPT-6 Astra52.8
  3. #3Claude Opus 550.7
  4. #4Claude Fable 549.7
  5. #5Muse Spark 1.348.2
Official ↗
LiveBench Coding
Contamination-resistant coding suite (completion, generation, JS, Python, TS)
Coding
Metric:Coding score%· 36 models
  1. #1Claude Fable 5.174.2%
  2. #2Claude Fable 571.7%
  3. #3Claude Opus 571.7%
  4. #4Muse Spark 1.370.9%
  5. #5Kimi K369.9%
Official ↗
Aider Polyglot
133 real coding tasks across 6 languages
Coding
Metric:Pass rate%· 33 models
  1. #1GPT-5.588.0%
  2. #2o3-pro84.9%
  3. #3Gemini 2.5 Pro83.1%
  4. #4o381.3%
  5. #5Grok 4.679.6%
Official ↗
LiveCodeBench
Continuously updated competitive-programming set
Coding
Metric:Pass@1%· 15 models
  1. #1DeepSeek-V3100.0%
  2. #2GPT-4o100.0%
  3. #3GPT-4o mini100.0%
  4. #4OpenAI o3-mini (High Reasoning Coding)100.0%
  5. #5o4-mini100.0%
Official ↗
EvalPlus
HumanEval+ strict test suite
Coding
Metric:HumanEval+%· 18 models
  1. #1o189.0%
  2. #2o1-mini89.0%
  3. #3Qwen3.6-Plus87.2%
  4. #4GPT-4o87.2%
  5. #5DeepSeek-V386.6%
Official ↗

LMArena AgentAgent

Metric: Agent Score·30 models·fetched 2026-09-10

Open official board ↗
#ModelAgent ScoreValue ScoreCross-board ranks
#114.51
7.9
C #1
A #1
L #1
ALE
#212.55
7.8
C
A #2
L #12
ALE
#3
Claude Opus 5Anthropic
11.42
5.1
C
A #3
L #3
ALE
#49.15
7.2
C #2
A #4
L #2
ALE
#57.75
11.9
C #6
A #6
L #9
ALE
#67.69
13.6
C #11
A #11
L #18
A #8
L #10
E
#7
Kimi K3Moonshot AI
6.59
21.4
C #14
A #9
L #5
ALE
#86.31
31.8
C #15
A #18
L #7
AL
E #8
#9
Hunyuan-T1Tencent
5.37
429
C
A #37
LALE
#10
GPT-5.5OpenAI
5.33
16.7
C #19
A #16
L #14
A #1
L
E #14
#11
GLM-5.2Zhipu AI
4.64
53.4
C #21
A #17
L #19
ALE
#124.43
117.8
C #12
A #20
L #17
ALE
#134.01
81.9
C #3
A #12
L #21
ALE
#143.93
39.5
C #7
A #15
L #22
ALE
#153.9
40.4
C #10
A #13
L #8
ALE
#163.24
42.9
C #4
A #8
L #15
A #5
LE
#17
GLM-5.3Zhipu AI
2.56
56.2
C #18
A #7
L #6
ALE
#182.22
324.5
C #17
A #23
L #25
ALE
#191.47
26.2
C #9
A #10
L #16
ALE
#201.11
245.7
C #16
A #19
L #20
ALE
#210.14
238.9
C
A #25
L #10
ALE
#22-1.05
18.5
C #23
A #28
L #27
A #13
L #9
E
#23-1.56
61.1
C
A #5
L #4
ALE
#24-3.09
59.2
C
A #30
L #32
A #15
L #13
E
#25
MiniMax-M3MiniMax
-5.26
156
C
A #31
L #34
ALE
#26-5.3
60.2
C
A #24
L #28
ALE
#27-5.34
40.8
C
A #29
L #29
AL
E #18
#28-9.35
47.8
C
A #52
LALE
#29-13.01
77.1
C
A #40
LALE
#30-13.43
77.7
C
A #41
L #26
ALE
Rank color:#1#2#3#4-5#6-10#11+

The "Cross-board ranks" column shows this model's rank on the other boards at a glance (grey = not tested there). Boards use different methodologies, so raw values are not directly comparable.

LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.

📊 7 benchmark providers

LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.

Agent Score30 models

CursorBench is Cursor's first-party agentic coding benchmark (v3.2). Tasks come from real Cursor sessions with ambiguous, multi-file requirements. Metric: pass rate (%). Mirrored from Cursor's official cursor.com/cursorbench snapshot.

Pass rate%25 models

Artificial Analysis Coding Index: Focused specifically on code generation, reasoning, and synthesis, fully shifted from generic overall intelligence to coding performance.

Coding Index91 models

LiveBench Coding: Extracts the five core coding benchmarks (code completion, code generation, JavaScript, Python, and TypeScript) for pure code performance.

Coding score%36 models

133 real coding tasks (bug-fix / feature) across 6 languages. Metric: pass_rate_2.

Pass rate%33 models

Continuously updated competitive-programming problems. Metric: Pass@1.

Pass@1%15 models

HumanEval+ strict test suite, 164 tasks.

HumanEval+%18 models

All rankings and numbers come directly from official sources; we only do name matching and aggregate presentation, no weighting or re-ranking. Boards use different methodologies, so do not compare raw values across boards.