2026 Live Verified

AI Coding Models Leaderboard

Aggregated official benchmark mirrors and real-world coding task scenarios. 100% verified, objective, and zero paid placement.

7Boards
95Models ranked
777Evaluation Tasks
September 2026Period
2026-09-11Last update
We fetch raw rankings from 6 public benchmark providers directly — no re-weighting, no composite score. Click "Open official board" on each card to go to the source.

📌 7 official boards · mirrored 1:1

Each card previews a provider's Top 5. Click "View full" for the complete board, or jump to the official page. Data is 100% from each provider — we do no aggregation or math.

LMArena Agent
Human preference on real agentic coding sessions
Agent
Metric:Agent Score· 30 models
  1. #1Claude Fable 5.113.85
  2. #2GPT-6 Astra12.39
  3. #3Claude Opus 511.06
  4. #4Claude Fable 59.09
  5. #5Claude Opus 4.87.57
Official ↗
CursorBench
Cursor's first-party agentic coding benchmark (v3.2)
Agent
Metric:Pass rate%· 25 models
  1. #1Claude Fable 5.175.6%
  2. #2Claude Fable 572.9%
  3. #3Gemini 3.8 Flash71.8%
  4. #4Grok 4.669.9%
  5. #5Gemini 3.7 Flash68.4%
Official ↗
Artificial Analysis
Official Coding Index for software engineering
Coding
Metric:Coding Index· 92 models
  1. #1Claude Fable 5.153.4
  2. #2GPT-6 Astra52.8
  3. #3Claude Opus 550.7
  4. #4Claude Fable 549.7
  5. #5Muse Spark 1.348.2
Official ↗
LiveBench Coding
Contamination-resistant coding suite (completion, generation, JS, Python, TS)
Coding
Metric:Coding score%· 37 models
  1. #1DeepSeek-V4 (Next-Gen MoE)78.4%
  2. #2Claude Fable 5.174.2%
  3. #3Claude Fable 571.7%
  4. #4Claude Opus 571.7%
  5. #5Muse Spark 1.370.9%
Official ↗
Aider Polyglot
133 real coding tasks across 6 languages
Coding
Metric:Pass rate%· 33 models
  1. #1GPT-5.588.0%
  2. #2o3-pro84.9%
  3. #3Gemini 2.5 Pro83.1%
  4. #4o381.3%
  5. #5Grok 4.679.6%
Official ↗
LiveCodeBench
Continuously updated competitive-programming set
Coding
Metric:Pass@1%· 15 models
  1. #1DeepSeek-V3100.0%
  2. #2GPT-4o100.0%
  3. #3GPT-4o mini100.0%
  4. #4OpenAI o3-mini (High Reasoning Coding)100.0%
  5. #5o4-mini100.0%
Official ↗
EvalPlus
HumanEval+ strict test suite
Coding
Metric:HumanEval+%· 18 models
  1. #1o189.0%
  2. #2o1-mini89.0%
  3. #3Qwen3.6-Plus87.2%
  4. #4GPT-4o87.2%
  5. #5DeepSeek-V386.6%
Official ↗

LMArena AgentAgent

Metric: Agent Score·30 models·fetched 2026-09-11

Open official board ↗
#ModelAgent ScoreValue ScoreCross-board ranks
#113.85
7.9
C #1
A #1
L #2
ALE
#212.39
7.8
C
A #2
L #13
ALE
#3
Claude Opus 5Anthropic
11.06
5.1
C
A #3
L #4
ALE
#49.09
7.2
C #2
A #4
L #3
ALE
#57.57
13.6
C #11
A #11
L #19
A #8
L #10
E
#67.51
12
C #6
A #6
L #10
ALE
#7
Kimi K3Moonshot AI
6.46
21.5
C #14
A #9
L #6
ALE
#85.88
31.9
C #15
A #19
L #8
AL
E #8
#9
Hunyuan-T1Tencent
5.23
434.2
C
A #38
LALE
#10
GPT-5.5OpenAI
5.22
16.8
C #19
A #17
L #15
A #1
L
E #14
#11
GLM-5.2Zhipu AI
4.56
53.7
C #21
A #18
L #20
ALE
#124.23
67.9
C
A #5
L #5
ALE
#134.14
118.2
C #12
A #21
L #18
ALE
#143.97
40.8
C #10
A #13
L #9
ALE
#153.84
82.3
C #3
A #12
L #22
ALE
#163.83
39.7
C #7
A #16
L #23
ALE
#173.12
43
C #4
A #8
L #16
A #5
LE
#18
GLM-5.3Zhipu AI
2.62
56.6
C #18
A #7
L #7
ALE
#192.17
326.7
C #17
A #24
L #26
ALE
#201.46
26.4
C #9
A #10
L #17
ALE
#210.73
246
C #16
A #20
L #21
ALE
#22-0.08
240.1
C
A #26
L #11
ALE
#23-1.27
18.5
C #23
A #29
L #28
A #13
L #9
E
#24-3.12
212.5
C
A #31
L #33
A #15
L #13
E
#25
MiniMax-M3MiniMax
-5.38
157.1
C
A #32
L #35
ALE
#26-5.41
60.7
C
A #25
L #29
ALE
#27-5.6
40.9
C
A #30
L #30
AL
E #18
#28-10.12
46.1
C
A #53
LALE
#29-13.22
77.8
C
A #41
LALE
#30-13.73
77.7
C
A #42
L #27
ALE
Rank color:#1#2#3#4-5#6-10#11+

The "Cross-board ranks" column shows this model's rank on the other boards at a glance (grey = not tested there). Boards use different methodologies, so raw values are not directly comparable.

LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.

📊 7 benchmark providers

LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.

Agent Score30 models

CursorBench is Cursor's first-party agentic coding benchmark (v3.2). Tasks come from real Cursor sessions with ambiguous, multi-file requirements. Metric: pass rate (%). Mirrored from Cursor's official cursor.com/cursorbench snapshot.

Pass rate%25 models

Artificial Analysis Coding Index: Focused specifically on code generation, reasoning, and synthesis, fully shifted from generic overall intelligence to coding performance.

Coding Index92 models

LiveBench Coding: Extracts the five core coding benchmarks (code completion, code generation, JavaScript, Python, and TypeScript) for pure code performance.

Coding score%37 models

133 real coding tasks (bug-fix / feature) across 6 languages. Metric: pass_rate_2.

Pass rate%33 models

Continuously updated competitive-programming problems. Metric: Pass@1.

Pass@1%15 models

HumanEval+ strict test suite, 164 tasks.

HumanEval+%18 models

All rankings and numbers come directly from official sources; we only do name matching and aggregate presentation, no weighting or re-ranking. Boards use different methodologies, so do not compare raw values across boards.