Two decision models that score options in one pass
Cloudflare released Clef and Clef-flash on 1 October 2026. Both are decision models: they take a state plus typed questions and return a probability for every allowed option.
The Clef model card says there is no free-form text generation and no output parsing. Question types match the System One style used by TypeSafe's Jev: noul, choice and score.
Clef is post-trained from Qwen3.8-27B. Clef-flash is post-trained from Qwen3.5-9B. Both keep a vision encoder. The Workers AI model page lists a 65,536 token context window and up to four images per request.
Hosted ids are @cf/cloudflare/clef and @cf/cloudflare/clef-flash, per the changelog.
Cloudflare joins the decision model race
Decision models rose after TypeSafe shipped Jev in September. AWS's Strands Labs followed with an open 2B model. See Strands Decider 2B: AWS's Strands Labs opens an Apache 2.0 decision model for agent routing and tool-call checks.
Cloudflare says Clef adds vision and a larger context window than Jev's 32k. It also says the API is Jev compatible, so callers can swap endpoints.
OpenAI put a hosted Decisions API into limited preview at DevDay. See OpenAI DevDay 2026 brings Ultrafast, a Decisions API preview, Codex cloud and plugin extensions.
Apache 2.0 weights plus a Workers AI price
Both weight repos carry Apache 2.0. When The Frontier checked at 19:07 SGT on 5 October, Cloudflare/clef had 1,302 likes on Hugging Face and clef-flash had 464.
Workers AI prices Clef at $0.24 per million input tokens and Clef-flash at $0.09. The model pages The Frontier read list input unit pricing only.
Cloudflare says it does not read, store or train on hosted requests unless the customer uses its fine-tuning product. Fine-tuning starts as a hands-on FDE service, with a self-serve platform planned later.
Pick Clef for multimodal gates on the edge
Cloudflare reports these Decision Index rows from its own harnesses. The Frontier found no independent test.
| Benchmark | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL, case exact | 98.47 | 98.76 | 95.75 |
| BANKING77, macro-F1 | 94.20 | 90.93 | 79.74 |
| When2Call, accuracy | 72.37 | 65.58 | 80.97 |
| Median latency, ms | 209.3 | 38.8 | 524.1 |
Jev leads When2Call in that table. Clef-flash leads BFCL and is far ahead on home appliance case exact in the blog's longer table (97.73 against Jev's 52.27).
Pick hosted Clef when you want edge GPUs, vision inputs and a System One shaped API. Pick the open weights when you can run an H200-class card, which is what the card's smoke test used.
Self-reported boards and a local GPU bill
All quality and latency figures above come from Cloudflare. Rows where rivals lead are printed in the same tables.
Local inference needs a large GPU. The Clef card says it was tested with torch 2.11 and transformers 5.10.2 on a single H200.
Cloudflare's blog pitches agents that can "decide" without a person in the loop. The model still returns probabilities only. Your code chooses the threshold and the action.
Blog, changelog, weights and Workers AI page
The announcement, changelog, both Hugging Face cards and the Workers AI model page are listed under Sources below.
