█

AI/TLDR

Clef

Cloudflare's open 27B multimodal decision model on Qwen3.8-27B that scores typed questions instead of writing text, behind a Jev-compatible API on Workers AI, released October 2026.

Clef (open decision models)Open weights
Released
1 Oct 2026
Context
65,536 tokens (Workers AI)
Parameters
27B
Input
$0.24 / 1M input tokens
License
Apache-2.0
Coverage
1 story

Overview

Clef is a 27B open decision model that Cloudflare released on October 1, 2026, alongside the smaller Clef-flash. A request carries a state (text, JSON or images) plus typed questions — yes/no (noul), multiple-choice (choice) and ordered rating (score) — and Clef returns a probability for every allowed option of every question, rather than free-form text. Cloudflare built it to be fully compatible with the API of Typesafe AI's Jev System One model, so the request format is the same, and serves it on Workers AI as @cf/cloudflare/clef. The weights are on Hugging Face under Apache 2.0.

Cloudflare diagram of a decision model: a support ticket reading "Our API started returning 500 errors 20 minutes ago" and two questions (which team: billing, technical or sales; is it urgent: yes or no) go into the decision model, which returns parallel outputs of Technical 100% and Urgent Yes 100%.
How a decision model answers several typed questions about one input in parallel.Cloudflare ↗

Cloudflare names two differences from Jev: Clef has a vision encoder, so it can classify images (Workers AI accepts up to four embedded PNG, JPEG or WebP images per request), and it has a 64K context window against Jev's 32K. The Workers AI docs list a 65,536-token context window.

Under the hood, Clef runs Qwen3.8-27B for a single prefill pass and then scores the valid schema choices in parallel, so no intermediate text is generated token by token. A two-stage attention routing head lets each option pull the parts of the prompt relevant to it and lets fields cross-attend to each other before scoring. Cloudflare froze the Qwen backbone and trained the routing head together with rank-256 LoRA adapters, using label-smoothed cross-entropy plus a Brier loss for calibration on internal synthetic data, and a reinforcement-learning stage it calls RLCD (Reinforcement Learning for Calibrated Decisions).

In Cloudflare's launch table, Clef scores 98.47 on BFCL case exact, 94.20 macro-F1 on BANKING77 and 97.43 macro-F1 on CLINC150+OOS, ahead of Jev (95.75, 79.74 and 89.27) on all three; Jev leads on When2Call (80.97 against 72.37) and BRIGHT (47.52 against 45.91). On Typesafe's workflow evals it scores 64.7 on invoice processing against 61.8 for Jev. Its median latency across Cloudflare's 43 eval benchmarks is 209.3 ms against 524.1 ms for Jev, and Cloudflare's decision-index chart places it at 61.2 against 57.9 for Jev. Cloudflare marks its own Clef figures as self-reported.

Cloudflare announced a reinforcement-learning fine-tuning service for Clef at the same time, first run hands-on by its forward-deployed engineering team and later planned as a self-serve platform. It chains AI Gateway (to capture request data), Workers AI (rollouts against the base model), Containers (sandboxes that score actions), a new Trainer component that updates the weights, and Workers AI's bring-your-own-model support to redeploy the tuned checkpoint.

Cloudflare diagram of its RL fine-tuning loop for Clef: AI Gateway captures production requests, your data pipeline prepares tasks and rewards, Workers AI generates rollouts starting from Clef, Containers run reward code, a new RL trainer updates the weights, and the selected checkpoint is deployed back to Workers AI with BYO Model.
The RL fine-tuning service Cloudflare announced alongside Clef.Cloudflare ↗
Released2026-10-01
LicenseApache-2.0
WeightsOpen weights
Parameters27B
Context65,536 tokens (Workers AI)
ArchitectureFrozen Qwen3.8-27B backbone (with its vision encoder) plus a jointly trained schema head and rank-256 LoRA adapters; a prefill-only pass, then every valid option is scored in parallel with no text generation
ModalitiesText, Vision

Benchmarks

Cloudflare's scatter chart of Decision Index 0.2.1 score against median latency: Clef at 61.2 and about 209 ms, Clef-flash at 57.1 and about 39 ms, both on the efficient frontier, against Jev at 57.9 and about 524 ms and open models such as Surogate Rune 26B-A4B v3 (57.4), AutoJev-27B (56.4), simple-jev Qwen3.8-27B (55.7) and Kev 9B (38.5).
Decision Index score against median latency, as published by Cloudflare (October 1, 2026). Clef and Clef-flash figures are self-reported. — Cloudflare

Clef against other decision models on benchmarks Cloudflare shortlisted from the Jev Decision Index, as published in Cloudflare's launch post (October 1, 2026).

BenchmarkClefClef-flashJevDiffusionGemma JevKev 9BLaya
BFCL · case exact98.4798.7695.7596.5294.5138.13
ToolRet · nDCG@1069.1966.4365.2861.2164.2612.69
API-Bank · accuracy91.9393.1188.1983.6656.311.41
Home appliances · case exact82.9597.7352.2742.05250
When2Call · accuracy72.3765.5880.9775.4449.6211.94
BANKING77 · macro-F194.290.9379.7474.2884.8314.29
CLINC150+OOS · macro-F197.4366.7789.2783.4979.033.19
BRIGHT · nDCG@1045.9139.2647.5242.9438.5319.9
Amazon ESCI · macro-F157.4857.3955.2153.3749.2224.4
PhishNChips · accuracy79.675.0562.5585.3550.7550.15
Median latency (ms, lower is better)209.3 ms38.8 ms524.1 ms84.4 ms51.4 ms5.8 ms
p95 latency (ms, lower is better)238.6 ms122.4 ms536 ms211.2 ms187.9 ms222.5 ms

Comparison source ↗

This model's scores

  1. BFCL (case exact)98.47%
  2. CLINC150+OOS (macro-F1)97.43%
  3. BANKING77 (macro-F1)94.2%
  4. API-Bank (accuracy)91.93%
  5. ToolRet (nDCG@10)69.19

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Pricing

Input$0.24 / 1M input tokens

Workers AI unit price for @cf/cloudflare/clef (21,818 neurons per M input tokens); no separate output price is listed.

Pricing source ↗

Strengths

  • Ahead of Jev in Cloudflare's launch table on BFCL (98.47 against 95.75), BANKING77 (94.20 against 79.74) and CLINC150+OOS (97.43 against 89.27)
  • 209.3 ms median latency across Cloudflare's 43 eval benchmarks, against 524.1 ms for Jev
  • Reads images as well as text and JSON, which Jev does not
  • 64K context window, double Jev's 32K
  • Jev / System One API compatible, so an existing Jev request can be pointed at Workers AI
  • Apache-2.0 weights on Hugging Face, plus hosted inference on Workers AI at $0.24 per M input tokens

Best for

  • Reach for it to route, triage or score support tickets, emails and policy checks where code acts on calibrated probabilities instead of parsing generated text.
  • Reach for it when the thing to classify includes screenshots or other images, such as categorising a rendered website, which Cloudflare's Threat Intelligence team uses it for.
  • Reach for it when you already call Jev and want an open-weight, self-hostable or Workers AI-hosted drop-in.
  • Reach for Clef-flash instead on latency-critical paths, where Cloudflare reports a 38.8 ms median.

How to access

ProviderModel ID
Cloudflare Workers AI ↗@cf/cloudflare/clef
Hugging Face (weights) ↗Cloudflare/clef

Clef (open decision models) — every version

The full lineage of the Clef (open decision models) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
Clefcurrent2026-10-0164KApache-2.0
Clef-flash2026-10-0164KApache-2.0

FAQ

What is Clef?

Clef is a 27B open decision model from Cloudflare, released October 1, 2026. Given a state (text, JSON or images) and typed questions, it returns a probability for each allowed answer instead of generating text. It runs on Workers AI as @cf/cloudflare/clef and its weights are on Hugging Face under Apache 2.0.

Is Clef compatible with Jev?

Yes. Cloudflare describes Clef as fully Jev-API compatible: it takes the same state plus noul, choice and score questions as Typesafe AI's Jev System One model and returns strictly typed answers. Image input is a Clef extension to that API.

How does Clef compare with Jev?

In Cloudflare's launch table Clef leads Jev on BFCL (98.47 against 95.75), BANKING77 (94.20 against 79.74) and CLINC150+OOS (97.43 against 89.27), while Jev leads on When2Call (80.97 against 72.37) and BRIGHT (47.52 against 45.91). Clef's median latency is 209.3 ms against 524.1 ms for Jev. Cloudflare marks its Clef figures as self-reported.

What does Clef cost on Workers AI?

The Workers AI pricing page lists $0.24 per million input tokens (21,818 neurons per million input tokens) for @cf/cloudflare/clef. Clef-flash is $0.09 per million input tokens.

What is Clef built on?

A frozen Qwen3.8-27B backbone with its vision encoder, plus a schema head that routes evidence to each option and scores the options jointly. Cloudflare trained the head with rank-256 LoRA adapters using label-smoothed cross-entropy, a Brier loss for calibration and a reinforcement-learning stage called RLCD.