█

AI/TLDR

Cloudflare · 2026-10-01 · major

Clef — Cloudflare's open decision models answer typed questions in milliseconds

Clef and Clef-flash are Cloudflare's Apache-2.0 decision models, built on Qwen 3.8-27B and Qwen 3.5-9B. They score typed choices instead of writing text, accept the Jev API, and run on Workers AI.

Cloudflare blog header for the Clef decision models announcement

Cloudflare's Clef models turn an input and a schema of typed questions into scored decisions, with open weights.

Key specs

Clef flash median latency38.8 ms
Bfcl (clef flash)98.76%

Quick facts

MakerCloudflare
ModelsClef (27B), Clef-flash (9B base)
Base modelsQwen 3.8-27B, Qwen 3.5-9B
LicenseApache-2.0
Context window64K tokens
AvailabilityHugging Face weights + Workers AI

Pricing

Clef input (Workers AI)$0.24 / 1M tokens
source ↗

What is it?

Clef is a family of two open-weight decision models Cloudflare released on 1 October 2026: Clef, built on Qwen 3.8-27B with rank-256 low-rank adapters, and the faster Clef-flash, built on Qwen 3.5-9B. A decision model does not write free text. Given text, JSON, images or video plus a schema of multiple-choice, true/false or ranked questions, Clef returns a probability for each allowed answer.

How does it work?

A two-stage attention routing step lets each valid choice pull the context it needs. At inference the model runs one prefill-only pass and then scores all valid schema choices in parallel, which is why there is no token-by-token generation. Cloudflare trained it with label-smoothed cross-entropy plus a Brier loss, and with a method it calls Reinforcement Learning for Calibrated Decisions (RLCD).

Why does it matter?

Agents make many small routing and classification calls, and an LLM is slow and inconsistent for those. On Cloudflare's numbers, Clef-flash's median latency is 38.8 ms against 524.1 ms for Jev, and both models double Jev's context to 64K. Because Clef speaks the Jev API, teams already using Jev-style tools can swap it in; Cloudflare is also starting an RL fine-tuning service, first with partners and later self-serve.

Who is it for?

developers building agent routing, triage and classification

Frequently asked questions

How much does Clef cost on Workers AI?
Clef is listed on Workers AI as @cf/cloudflare/clef at $0.24 per million input tokens, according to Cloudflare's model page. The weights for Clef and Clef-flash are also free to download from Hugging Face under the Apache-2.0 license, so teams can run the models on their own GPUs instead of paying for hosted inference.
How does Clef compare with Jev?
Cloudflare compares Clef with Jev on the Jev Decision Index. Clef and Clef-flash have a 64K context window against Jev's 32K, add vision input, and keep Jev API compatibility. Median latency across 43 benchmarks is 38.8 ms for Clef-flash and 209.3 ms for Clef, against 524.1 ms for Jev, in Cloudflare's own tests.
Which Clef model should I pick, Clef or Clef-flash?
Cloudflare positions Clef as the larger, stronger model and Clef-flash as the one tuned for latency. In Cloudflare's results Clef-flash scores slightly higher on BFCL (98.76% vs 98.47%) and API-Bank (93.11% vs 91.93%), while Clef posts the BANKING77 and CLINC150 numbers. Clef-flash is over five times faster at the median.
Can I fine-tune Clef with Cloudflare?
Cloudflare's RL fine-tuning offer for Clef starts as a hands-on partnership with its forward-deployed engineers, with an interest form on cloudflare.com. A self-serve platform for capturing data, fine-tuning and redeploying the model is announced as coming later, and Cloudflare has not published pricing for either phase.

Try it

@cf/cloudflare/clef on Workers AI

Sources · 4 outlets

Tags

  • clef
  • clef-flash
  • cloudflare
  • decision-model
  • jev
  • qwen
  • workers-ai
  • structured-output
  • open-weights
  • apache-2-0
  • reinforcement-learning

← All releases · Learn AI