On 2 Oct 2026 OpenRouter shipped Model Router Benchmarks: seven model routers (Auto, Jev, NVIDIA Switchyard, Unbiased Pareto, Sakana Fugu, and variants) compared on GPQA Diamond, Tau3 Banking, MMMU Pro Vision, SWE Atlas QA, Deep SWE, and Terminal Bench 2.1. The Router Index defaults to quality 60% / time 20% / cost 20%. openrouter/pareto-code and Fusion are not on this board.

Key Takeaways

  • ✓Announced 2026-10-02 22:08 UTC; blog by Brian Thomas dated 10/2/2026; live page openrouter.ai/benchmarks/routers
  • ✓Seven router slugs in the public stats feed: openrouter/auto, typesafe/jev-router, nvidia/switchyard, unbiased/pareto, sakana/fugu-max, sakana/fugu-ultra, sakana/fugu-ultra-v2. pareto-code and Fusion were not benchmarked
  • ✓Router Index default is quality 60% / time 20% / cost 20% on a 0-10 scale; it is not raw accuracy
  • ✓Stats API 2026-10-04 SGT, Deep SWE n=113: Switchyard Flash+Opus 5.5 and Opus-only both 83/113 (~$279 / $285); Flash+GPT-6 Sol 74/113 at ~$55
  • ✓Terminal Bench 2.1 n=89: GPT-6 Astra 77/89 ($66.50), Opus 5.5 74/89 ($24.00); one Pareto run 71/89 ($27.20)
🧭

Turn your technical choice into a development budget

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Why it matters

Model routing is no longer only provider failover. OpenRouter's 2 Oct 2026 post says GPT-6 Astra's average price paid per token is over 48x DeepSeek v4 Flash, and an average 10-to-49-turn Codex session is about 21x more expensive. The best model for legal research, coding, planning, and rote execution also differs. Switching still rebuilds the prompt cache, misreads difficulty, and adds latency.

@OpenRouter announced Model Router Benchmarks on 2 Oct 2026, 22:08 UTC: 7 routers, 6 benchmarks, scored on quality, speed, and cost. The page will keep adding routers and refreshing scores.

How the routers differ

Provider routing (pick an inference host) is the old OpenRouter path. This leaderboard is model routing.

  • Blend, opaque membership, flat per-token price, sometimes with frontier escalation: Unbiased Pareto (unbiased/pareto) and Sakana Fugu.
  • One transparent model per turn, billed at that model's rate: Auto (openrouter/auto, trailing 7-day community spend) and Jev (typesafe/jev-router).
  • A declared cheap/strong pair: NVIDIA Switchyard (nvidia/switchyard).

Not on this board, per the post: openrouter/pareto-code, -latest family aliases, and Fusion. Unbiased Pareto is not Pareto Code.

Router Index is a 0-10 blend. Default weights: quality 60%, time per task 20%, cost 20%. That index is not raw accuracy.

Scores (raw API runs, not the index)

Public [/api/frontend/v1/stats/router-benchmarks](https://openrouter.ai/api/frontend/v1/stats/router-benchmarks) fetched 4 Oct 2026 SGT (truncated: false, 112 successful runs). The six suites matching the page columns: GPQA Diamond, Tau3 Banking, MMMU Pro Vision, SWE Atlas QA, Deep SWE, Terminal Bench 2.1. search_hle is in the feed but not one of those six. Compare only equal question counts.

Deep SWE, 113 questions: Switchyard GLM-5.3-Flash + Opus 5.5 83/113 ($278.71, 28 Sep); Switchyard Opus 5.5 only 83/113 ($285.01, 23 Sep); Switchyard GPT-6 Sol + Astra 80/113 ($265.41); Switchyard DeepSeek V4.1 Flash + GPT-6 Sol 74/113 ($54.61); Auto high 65/113 ($166.94, 29 Sep); Auto medium 62/113 ($90.35); Switchyard GPT-6 Luna only 60/113 ($7.50).

Terminal Bench 2.1, 89 questions (~30 Sep): GPT-6 Astra 77/89 ($66.50), Opus 5.5 74/89 ($24.00), GPT-6 Sol 71/89 ($16.02), one Pareto run 71/89 ($27.20; a second run is 145/178), GPT-6 Luna 63/89 ($1.22), Fugu Max 58/89 ($50.25). Jev Deep SWE is 159/226 and Pareto Deep SWE is 153/226 — different denominators. GPQA at 198 items: Switchyard Sol+Astra 187/198 ($6.71), Jev 186/198 ($3.15), Switchyard Flash+Sol 185/198 ($1.45).

Try them

Swap model on the existing chat API. Docs: blog, page, Auto, Jev, Switchyard. Coding-only openrouter/pareto-code (min_coding_score 0-1) is a different product and was not benchmarked here. Re-check Activity on your own repo; the post says the suites are general tasks, not your workload.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.