Developer 3s Key Decision Metrics
@Cloudflare (2026-10-01, #BirthdayWeek) launched its first in-house decision models Clef (27B) and Clef-flash (9B): hosted on Workers AI, Apache 2.0 weights, Jev System One–compatible; official blog cites ~2.5× / 13× median latency vs Jev and 2.2s vs 4.7s in a threat-intel Browser Run workflow.
Key Takeaways
- ✓Primary: Cloudflare 2026-10-01 → Clef blog
- ✓Median latency (43 runs): Clef 209.3 ms, Clef-flash 38.8 ms vs Jev 524.1 ms (~2.5× / 13×)
- ✓Quality: BFCL 98.47/98.76; BANKING77 94.20; CLINC150+OOS 97.43; Clef family tops 7/10 decision benches
- ✓Shape: Qwen backbone + non-AR schema scoring; 64K ctx; vision;
@cf/cloudflare/clef/clef-flash; $0.24 / $0.09 per M input tokens - ✓Ship: Workers AI binding/REST; weights on HF; RL fine-tune via design-partner form
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Core Background & Industry Pain Points
@Cloudflare on 2026-10-01 (#BirthdayWeek) pointed to Introducing Clef: the Workers AI team’s first in-house decision models Clef and Clef-flash. Agent hot paths need cheap, fast, structured classify/route answers—not another autoregressive round. After Typesafe’s Jev popularized the pattern, Cloudflare ships a System One–compatible open alternative for its “agent cloud” thesis.
Architecture Highlights & Internals
A decision model takes state plus up to 64 typed questions (noul / choice / score) and returns per-option probabilities—no free-form text to parse. Per changelog: Clef 27B, Clef-flash 9B, 64K context, vision encoder (≤4 images). Backbone: frozen Qwen (Qwen3.8-27B / Qwen3.5-9B) + rank-256 LoRA and routing head; inference is prefill-only then parallel schema scoring (non-autoregressive). RL fine-tune path: AI Gateway datasets → Containers sandbox → Trainer → BYO Model on Workers AI.
Authoritative Benchmarks & Measured Scores
All figures are Cloudflare’s own (blog/changelog), not third-party replications: median latency Clef 209.3 ms, flash 38.8 ms, Jev 524.1 ms (~2.5× / 13×); p95 238.6 / 122.4 / 536.0 ms. Quality excerpts: BFCL 98.47 / 98.76 (vs Jev 95.75); BANKING77 macro-F1 94.20; CLINC150+OOS 97.43; Home appliances flash 97.73. Threat intel + Browser Run: Clef 2.2 s vs gpt-oss-120b 4.7 s. On Typesafe workflow evals, Clef wins 3/4 vs Jev. Full tables: blog + HF card.
Developer Hands-on Guide
- Workers:
env.AI.run("@cf/cloudflare/clef", { model: "clef", state, questions }); REST/ai/run/@cf/cloudflare/clef— see clef / clef-flash. - Pricing (docs): $0.24 / $0.09 per M input tokens; enterprise default: no read/store/train on your traffic unless fine-tuning.
- Apache 2.0 weights: Cloudflare/clef, Cloudflare/clef-flash.
- Domain fine-tune: RL design-partner form; self-serve platform still maturing. Treat the tweet as a pointer—blog/docs are the SLA.

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.