Comet Opik 2.2.84 (2026-09-29) lets online LLM-as-a-Judge rules call TypeSafe Jev end-to-end via OpenRouter POST /api/alpha/decisions (not chat/completions). Each Boolean metric becomes a yes/no question answered in one call; score 1 if probability ≥0.5, with Probability: 0.93 in the reason. Backend adds OpenRouterDecisionsClient (workspace OpenRouter key, 429/5xx retries) for ~typesafe/jev-latest / typesafe/jev-1.13. Also: Annotation Queue automation + item provenance. ~8 commits / 104 files vs 2.2.83.

Key Takeaways

  • ✓Release 2.2.84; ~8 commits / 104 files vs 2.2.83
  • ✓PR #8540: online judges call Jev via OpenRouter Decisions API
  • ✓Boolean metrics → yes/no questions; score 1 if p≥0.5; reason shows Probability
  • ✓Models: ~typesafe/jev-latest, typesafe/jev-1.13
  • ✓Queue automation + item provenance; trace stats / thread_id fixes
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Background Online LLM-as-a-Judge rules often force Boolean checks through chat completions. TypeSafe Jev is a decisions model: OpenRouter serves it on POST /api/alpha/decisions and rejects chat/completions—Opik needed a dedicated scoring path. ### What shipped 2.2.84 (PR #8540) routes Boolean online rules through OpenRouterDecisionsClient for ~typesafe/jev-latest / typesafe/jev-1.13. One Decisions call answers all yes/no questions; score 1 if p≥0.5. Also: queue automation + provenance; trace stats / thread_id fixes. ### Benchmarks No public judge-accuracy table. Signal: ~8 commits / 104 files vs 2.2.83. Calibrate the 0.5 threshold on your labels. ### Get started Upgrade to 2.2.84, set the workspace OpenRouter key, pick a Jev model in online eval rules. Docs: Opik, repo, TypeSafe.