Overview
Clef is a 27B open decision model that Cloudflare released on October 1, 2026, alongside the smaller Clef-flash. A request carries a state (text, JSON or images) plus typed questions — yes/no (noul), multiple-choice (choice) and ordered rating (score) — and Clef returns a probability for every allowed option of every question, rather than free-form text. Cloudflare built it to be fully compatible with the API of Typesafe AI's Jev System One model, so the request format is the same, and serves it on Workers AI as @cf/cloudflare/clef. The weights are on Hugging Face under Apache 2.0.

Cloudflare names two differences from Jev: Clef has a vision encoder, so it can classify images (Workers AI accepts up to four embedded PNG, JPEG or WebP images per request), and it has a 64K context window against Jev's 32K. The Workers AI docs list a 65,536-token context window.
Under the hood, Clef runs Qwen3.8-27B for a single prefill pass and then scores the valid schema choices in parallel, so no intermediate text is generated token by token. A two-stage attention routing head lets each option pull the parts of the prompt relevant to it and lets fields cross-attend to each other before scoring. Cloudflare froze the Qwen backbone and trained the routing head together with rank-256 LoRA adapters, using label-smoothed cross-entropy plus a Brier loss for calibration on internal synthetic data, and a reinforcement-learning stage it calls RLCD (Reinforcement Learning for Calibrated Decisions).
In Cloudflare's launch table, Clef scores 98.47 on BFCL case exact, 94.20 macro-F1 on BANKING77 and 97.43 macro-F1 on CLINC150+OOS, ahead of Jev (95.75, 79.74 and 89.27) on all three; Jev leads on When2Call (80.97 against 72.37) and BRIGHT (47.52 against 45.91). On Typesafe's workflow evals it scores 64.7 on invoice processing against 61.8 for Jev. Its median latency across Cloudflare's 43 eval benchmarks is 209.3 ms against 524.1 ms for Jev, and Cloudflare's decision-index chart places it at 61.2 against 57.9 for Jev. Cloudflare marks its own Clef figures as self-reported.
Cloudflare announced a reinforcement-learning fine-tuning service for Clef at the same time, first run hands-on by its forward-deployed engineering team and later planned as a self-serve platform. It chains AI Gateway (to capture request data), Workers AI (rollouts against the base model), Containers (sandboxes that score actions), a new Trainer component that updates the weights, and Workers AI's bring-your-own-model support to redeploy the tuned checkpoint.

| Released | 2026-10-01 |
|---|---|
| License | Apache-2.0 |
| Weights | Open weights |
| Parameters | 27B |
| Context | 65,536 tokens (Workers AI) |
| Architecture | Frozen Qwen3.8-27B backbone (with its vision encoder) plus a jointly trained schema head and rank-256 LoRA adapters; a prefill-only pass, then every valid option is scored in parallel with no text generation |
| Modalities | Text, Vision |
Benchmarks

Clef against other decision models on benchmarks Cloudflare shortlisted from the Jev Decision Index, as published in Cloudflare's launch post (October 1, 2026).
| Benchmark | Clef | Clef-flash | Jev | DiffusionGemma Jev | Kev 9B | Laya |
|---|---|---|---|---|---|---|
| BFCL · case exact | 98.47 | 98.76 | 95.75 | 96.52 | 94.51 | 38.13 |
| ToolRet · nDCG@10 | 69.19 | 66.43 | 65.28 | 61.21 | 64.26 | 12.69 |
| API-Bank · accuracy | 91.93 | 93.11 | 88.19 | 83.66 | 56.3 | 11.41 |
| Home appliances · case exact | 82.95 | 97.73 | 52.27 | 42.05 | 25 | 0 |
| When2Call · accuracy | 72.37 | 65.58 | 80.97 | 75.44 | 49.62 | 11.94 |
| BANKING77 · macro-F1 | 94.2 | 90.93 | 79.74 | 74.28 | 84.83 | 14.29 |
| CLINC150+OOS · macro-F1 | 97.43 | 66.77 | 89.27 | 83.49 | 79.03 | 3.19 |
| BRIGHT · nDCG@10 | 45.91 | 39.26 | 47.52 | 42.94 | 38.53 | 19.9 |
| Amazon ESCI · macro-F1 | 57.48 | 57.39 | 55.21 | 53.37 | 49.22 | 24.4 |
| PhishNChips · accuracy | 79.6 | 75.05 | 62.55 | 85.35 | 50.75 | 50.15 |
| Median latency (ms, lower is better) | 209.3 ms | 38.8 ms | 524.1 ms | 84.4 ms | 51.4 ms | 5.8 ms |
| p95 latency (ms, lower is better) | 238.6 ms | 122.4 ms | 536 ms | 211.2 ms | 187.9 ms | 222.5 ms |
This model's scores
- BFCL (case exact)98.47%
- CLINC150+OOS (macro-F1)97.43%
- BANKING77 (macro-F1)94.2%
- API-Bank (accuracy)91.93%
- ToolRet (nDCG@10)69.19
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Pricing
| Input | $0.24 / 1M input tokens |
|---|
Workers AI unit price for @cf/cloudflare/clef (21,818 neurons per M input tokens); no separate output price is listed.
Strengths
- Ahead of Jev in Cloudflare's launch table on BFCL (98.47 against 95.75), BANKING77 (94.20 against 79.74) and CLINC150+OOS (97.43 against 89.27)
- 209.3 ms median latency across Cloudflare's 43 eval benchmarks, against 524.1 ms for Jev
- Reads images as well as text and JSON, which Jev does not
- 64K context window, double Jev's 32K
- Jev / System One API compatible, so an existing Jev request can be pointed at Workers AI
- Apache-2.0 weights on Hugging Face, plus hosted inference on Workers AI at $0.24 per M input tokens
Best for
- Reach for it to route, triage or score support tickets, emails and policy checks where code acts on calibrated probabilities instead of parsing generated text.
- Reach for it when the thing to classify includes screenshots or other images, such as categorising a rendered website, which Cloudflare's Threat Intelligence team uses it for.
- Reach for it when you already call Jev and want an open-weight, self-hostable or Workers AI-hosted drop-in.
- Reach for Clef-flash instead on latency-critical paths, where Cloudflare reports a 38.8 ms median.
How to access
| Provider | Model ID |
|---|---|
| Cloudflare Workers AI ↗ | @cf/cloudflare/clef |
| Hugging Face (weights) ↗ | Cloudflare/clef |
Clef (open decision models) — every version
The full lineage of the Clef (open decision models) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| Clefcurrent | 2026-10-01 | 64K | Apache-2.0 |
| Clef-flash | 2026-10-01 | 64K | Apache-2.0 |
FAQ
What is Clef?
Clef is a 27B open decision model from Cloudflare, released October 1, 2026. Given a state (text, JSON or images) and typed questions, it returns a probability for each allowed answer instead of generating text. It runs on Workers AI as @cf/cloudflare/clef and its weights are on Hugging Face under Apache 2.0.
Is Clef compatible with Jev?
Yes. Cloudflare describes Clef as fully Jev-API compatible: it takes the same state plus noul, choice and score questions as Typesafe AI's Jev System One model and returns strictly typed answers. Image input is a Clef extension to that API.
How does Clef compare with Jev?
In Cloudflare's launch table Clef leads Jev on BFCL (98.47 against 95.75), BANKING77 (94.20 against 79.74) and CLINC150+OOS (97.43 against 89.27), while Jev leads on When2Call (80.97 against 72.37) and BRIGHT (47.52 against 45.91). Clef's median latency is 209.3 ms against 524.1 ms for Jev. Cloudflare marks its Clef figures as self-reported.
What does Clef cost on Workers AI?
The Workers AI pricing page lists $0.24 per million input tokens (21,818 neurons per million input tokens) for @cf/cloudflare/clef. Clef-flash is $0.09 per million input tokens.
What is Clef built on?
A frozen Qwen3.8-27B backbone with its vision encoder, plus a schema head that routes evidence to each option and scores the options jointly. Cloudflare trained the head with rank-256 LoRA adapters using label-smoothed cross-entropy, a Brier loss for calibration and a reinforcement-learning stage called RLCD.