AI Coder Bench
Real-world coding benchmark rankings. 50 tasks across bug fixing, feature building, refactoring, system design, and debug & explain.
📌 7 official boards · mirrored 1:1
Each card previews a provider's Top 5. Click "View full" for the complete board, or jump to the official page. Data is 100% from each provider — we do no aggregation or math.
Agent Score· 27 models- #1Claude Fable 512.6
- #2Grok 4.610.5
- #3GPT-5.6 Sol10.11
- #4Kimi K39.75
- #5Claude Opus 4.89.5
Pass rate%· 17 models- #1Claude Fable 572.9%
- #2Grok 4.669.9%
- #3GPT-5.6 Sol67.2%
- #4Grok 4.566.7%
- #5GPT-5.6 Terra64.9%
AA Index· 49 models- #1Grok 4.661
- #2Claude Opus 560.7
- #3Claude Fable 559.9
- #4GPT-5.6 Sol58.9
- #5Kimi K357.1
Avg score%· 27 models- #1GPT-5.6 Sol82.9%
- #2Grok 4.682.1%
- #3Claude Fable 581.4%
- #4Claude Opus 580.7%
- #5GPT-5.580.5%
Pass rate%· 20 models- #1GPT-5.6 Sol88.0%
- #2Grok 4.686.5%
- #3GPT-5.584.9%
- #4Gemini 2.5 Pro83.1%
- #5Grok 4.579.6%
Pass@1%· 12 models- #1DeepSeek-V4-Pro100.0%
- #2GPT-5.4100.0%
- #3GPT-5.4 mini100.0%
- #4GPT-5.5100.0%
- #5Claude Sonnet 4.6100.0%
HumanEval+%· 18 models- #1GPT-5.489.0%
- #2GPT-5.4 mini89.0%
- #3Grok 4.688.4%
- #4Qwen3.6-Plus87.2%
- #5DeepSeek-V4-Pro86.6%
LMArena AgentAgent
| # | Model | Agent Score |
|---|---|---|
| #1 | Claude Fable 5Anthropic | 12.6 |
| #2 | Grok 4.6xAI | 10.5 |
| #3 | GPT-5.6 SolOpenAI | 10.11 |
| #4 | Kimi K3Moonshot AI | 9.75 |
| #5 | Claude Opus 4.8Anthropic | 9.5 |
| #6 | Claude Sonnet 5Anthropic | 8.61 |
| #7 | GPT-5.5OpenAI | 8.61 |
| #8 | GLM-5.2Zhipu AI | 7.12 |
| #9 | Grok 4.5xAI | 5.47 |
| #10 | GPT-5.4OpenAI | 5.38 |
| #11 | GPT-5.6 TerraOpenAI | 3.96 |
| #12 | GPT-5.6 LunaOpenAI | 3.28 |
| #13 | Claude Sonnet 4.6Anthropic | 3.14 |
| #14 | GLM-5.1Zhipu AI | 0.72 |
| #15 | Kimi K2.7 CodeMoonshot AI | 0.38 |
| #16 | Qwen3.7-MaxQwen | -0.17 |
| #17 | Gemini 3.1 ProGoogle | -0.31 |
| #18 | Gemini 3 ProGoogle | -0.61 |
| #19 | DeepSeek-V4-ProDeepSeek | -0.72 |
| #20 | Gemini 3.5 FlashGoogle | -0.94 |
| #21 | Kimi K2.6Moonshot AI | -1 |
| #22 | MiniMax-M3MiniMax | -2.55 |
| #23 | DeepSeek-V4-FlashDeepSeek | -2.95 |
| #24 | Grok 4.3xAI | -7.87 |
| #25 | Grok Build 0.1xAI | -7.97 |
| #26 | MiniMax-M2.7MiniMax | -11.11 |
| #27 | Nemotron-UltraNvidia | -13.08 |
The "Cross-board ranks" column shows this model's rank on the other boards at a glance (grey = not tested there). Boards use different methodologies, so raw values are not directly comparable.
LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.
📊 7 benchmark providers
LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.
Agent Score27 modelsCursorBench is Cursor's first-party agentic coding benchmark (v3.2). Tasks come from real Cursor sessions with ambiguous, multi-file requirements. Metric: pass rate (%). Mirrored from Cursor's official cursor.com/cursorbench snapshot.
Pass rate%17 modelsArtificial Analysis aggregates 20+ benchmarks (MMLU-Pro, GPQA, HLE, SciCode, IFBench, Terminal-bench, etc.) into a single Intelligence Index — the most comprehensive commercial evaluator for "overall capability".
AA Index49 modelsLiveBench refreshes its questions monthly to prevent contamination. Covers reasoning, math, coding, language, instruction following, and data analysis.
Avg score%27 models133 real coding tasks (bug-fix / feature) across 6 languages. Metric: pass_rate_2.
Pass rate%20 modelsContinuously updated competitive-programming problems. Metric: Pass@1.
Pass@1%12 modelsAll rankings and numbers come directly from official sources; we only do name matching and aggregate presentation, no weighting or re-ranking. Boards use different methodologies, so do not compare raw values across boards.