AI Coder Bench
Real-world coding benchmark rankings. 50 tasks across bug fixing, feature building, refactoring, system design, and debug & explain.
📌 4 lenses · 6 official boards
Each card previews a benchmark provider's Top 5. Click to expand full rankings, or jump straight to the official page.
Arena Elo· 12 models- #1Claude Fable 51631
- #2Claude Opus 4.81581
- #3Grok 4.51558
- #4Claude Sonnet 4.61544
- #5Kimi K2.61519
AA Index· 46 models- #1Claude Fable 559.9
- #2GPT-5.6 Sol58.9
- #3Kimi K357.1
- #4Claude Opus 4.855.7
- #5GPT-5.6 Terra55
Avg score%· 23 models- #1GPT-5.6 Sol82.9%
- #2Claude Fable 581.4%
- #3GPT-5.580.5%
- #4GPT-5.6 Terra80.3%
- #5Claude Opus 4.879.8%
Pass rate%· 19 models- #1GPT-5.6 Sol88.0%
- #2GPT-5.584.9%
- #3Gemini 2.5 Pro83.1%
- #4Grok 4.579.6%
- #5DeepSeek-V4-Pro74.2%
Pass{'@'}1%· 11 models- #1DeepSeek-V4-Pro100.0%
- #2GPT-5.4100.0%
- #3GPT-5.4 mini100.0%
- #4GPT-5.5100.0%
- #5Claude Sonnet 4.6100.0%
HumanEval+%· 17 models- #1GPT-5.489.0%
- #2GPT-5.4 mini89.0%
- #3Qwen3.6-Plus87.2%
- #4DeepSeek-V4-Pro86.6%
- #5DeepSeek-V4-Flash83.5%
LMArenaChat
| # | Model | Arena Elo |
|---|---|---|
| #1 | Claude Fable 5Anthropic | 1631 |
| #2 | Claude Opus 4.8Anthropic | 1581 |
| #3 | Grok 4.5xAI | 1558 |
| #4 | Claude Sonnet 4.6Anthropic | 1544 |
| #5 | Kimi K2.6Moonshot AI | 1519 |
| #6 | GPT-5.5OpenAI | 1488 |
| #7 | Gemini 3 ProGoogle | 1486 |
| #8 | GPT-5.6 SolOpenAI | 1486 |
| #9 | GPT-5.4OpenAI | 1472 |
| #10 | Gemini 2.5 ProGoogle | 1393 |
| #11 | Gemini 3.5 FlashGoogle | 1286 |
| #12 | Gemini 3.1 ProGoogle | 1211 |
The "Cross-board ranks" column shows this model's rank on the other 5 boards at a glance (grey = the model was not tested on that board).
Hundreds of thousands of human blind-vote conversations (Elo), from LMArena / Chatbot Arena. The most authoritative "user preference" board today.
📊 6 benchmark providers
Hundreds of thousands of human blind-vote conversations (Elo), from LMArena / Chatbot Arena. The most authoritative "user preference" board today.
Arena Elo12 modelsArtificial Analysis aggregates 20+ benchmarks (MMLU-Pro, GPQA, HLE, SciCode, IFBench, Terminal-bench, etc.) into a single Intelligence Index — the most comprehensive commercial evaluator for "overall capability".
AA Index46 modelsLiveBench refreshes its questions monthly to prevent contamination. Covers reasoning, math, coding, language, instruction following, and data analysis.
Avg score%23 models133 real coding tasks (bug-fix / feature) across 6 languages. Metric: pass_rate_2.
Pass rate%19 modelsContinuously updated competitive-programming problems. Metric: Pass@1.
Pass{'@'}1%11 modelsAll rankings and numbers come directly from official sources; we only do name matching and aggregate presentation, no weighting or re-ranking. Boards use different methodologies, so do not compare raw values across boards.