Overview
Clef-flash is the 9B, latency-focused member of Cloudflare's Clef decision models, released on October 1, 2026 together with the 27B Clef. Like Clef, it takes a state (text, JSON or images) plus typed questions — yes/no (noul), multiple-choice (choice) and ordered rating (score) — and returns a probability for every allowed option instead of generating text, through an API compatible with Typesafe AI's Jev System One model. It is served on Workers AI as @cf/cloudflare/clef-flash, and the weights are on Hugging Face under Apache 2.0.

It shares Clef's design on a smaller backbone: Qwen3.5-9B, kept frozen, runs one prefill pass, and a routing head trained with rank-256 LoRA adapters scores all valid choices in parallel. The Workers AI docs list a 65,536-token context window and image input (up to four embedded PNG, JPEG or WebP images per request).
Speed is the point. Across the 43 eval benchmarks Cloudflare ran, Clef-flash has a 38.8 ms median and 122.4 ms p95 latency, against 209.3 ms and 238.6 ms for Clef and 524.1 ms and 536.0 ms for Jev. Cloudflare calls the larger Clef its precision model and Clef-flash the one for latency-critical decisions.
It holds up on quality in Cloudflare's launch table: 98.76 on BFCL case exact (Clef 98.47, Jev 95.75), 93.11 on API-Bank (Jev 88.19) and 97.73 on the home-appliances task (Jev 52.27). It trails both Clef and Jev on CLINC150+OOS (66.77 macro-F1 against 97.43 and 89.27). On Typesafe's workflow evals it scores 77 on customer service and 69.8 on agent-trace observability, against 76.0 and 71.6 for Jev, and Cloudflare's decision-index chart places it at 57.1 against 57.9 for Jev. Cloudflare marks its own Clef-flash figures as self-reported.
| Released | 2026-10-01 |
|---|---|
| License | Apache-2.0 |
| Weights | Open weights |
| Parameters | 9B |
| Context | 65,536 tokens (Workers AI) |
| Architecture | Frozen Qwen3.5-9B backbone (with its vision encoder) plus a jointly trained schema head and rank-256 LoRA adapters; a prefill-only pass, then every valid option is scored in parallel with no text generation |
| Modalities | Text, Vision |
Benchmarks

Clef-flash against other decision models on benchmarks Cloudflare shortlisted from the Jev Decision Index, as published in Cloudflare's launch post (October 1, 2026).
| Benchmark | Clef | Clef-flash | Jev | DiffusionGemma Jev | Kev 9B | Laya |
|---|---|---|---|---|---|---|
| BFCL · case exact | 98.47 | 98.76 | 95.75 | 96.52 | 94.51 | 38.13 |
| ToolRet · nDCG@10 | 69.19 | 66.43 | 65.28 | 61.21 | 64.26 | 12.69 |
| API-Bank · accuracy | 91.93 | 93.11 | 88.19 | 83.66 | 56.3 | 11.41 |
| Home appliances · case exact | 82.95 | 97.73 | 52.27 | 42.05 | 25 | 0 |
| When2Call · accuracy | 72.37 | 65.58 | 80.97 | 75.44 | 49.62 | 11.94 |
| BANKING77 · macro-F1 | 94.2 | 90.93 | 79.74 | 74.28 | 84.83 | 14.29 |
| CLINC150+OOS · macro-F1 | 97.43 | 66.77 | 89.27 | 83.49 | 79.03 | 3.19 |
| BRIGHT · nDCG@10 | 45.91 | 39.26 | 47.52 | 42.94 | 38.53 | 19.9 |
| Amazon ESCI · macro-F1 | 57.48 | 57.39 | 55.21 | 53.37 | 49.22 | 24.4 |
| PhishNChips · accuracy | 79.6 | 75.05 | 62.55 | 85.35 | 50.75 | 50.15 |
| Median latency (ms, lower is better) | 209.3 ms | 38.8 ms | 524.1 ms | 84.4 ms | 51.4 ms | 5.8 ms |
| p95 latency (ms, lower is better) | 238.6 ms | 122.4 ms | 536 ms | 211.2 ms | 187.9 ms | 222.5 ms |
This model's scores
- BFCL (case exact)98.76%
- Home appliances (case exact)97.73%
- API-Bank (accuracy)93.11%
- BANKING77 (macro-F1)90.93%
- ToolRet (nDCG@10)66.43
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Pricing
| Input | $0.09 / 1M input tokens |
|---|
Workers AI unit price for @cf/cloudflare/clef-flash (8,182 neurons per M input tokens); no separate output price is listed.
Strengths
- 38.8 ms median and 122.4 ms p95 latency across Cloudflare's 43 eval benchmarks, against 524.1 ms and 536.0 ms for Jev
- 98.76 on BFCL case exact and 93.11 on API-Bank in Cloudflare's launch table, ahead of Clef and Jev on both
- 97.73 on the home-appliances task, against 82.95 for Clef and 52.27 for Jev
- Reads images as well as text and JSON, with a 64K context window
- $0.09 per M input tokens on Workers AI, under half of Clef's $0.24
- Apache-2.0 weights on Hugging Face and a Jev-compatible request format
Best for
- Reach for it in the hot path of an agent, where a routing or yes/no decision has to come back in tens of milliseconds before an LLM acts.
- Reach for it for high-volume tool-selection and API-call classification, where it posts its best scores (BFCL, API-Bank).
- Reach for it when you want a small, self-hostable decision model that still accepts images.
- Reach for the 27B Clef instead for intent classification with many classes, where Clef-flash scores 66.77 on CLINC150+OOS against Clef's 97.43.
How to access
| Provider | Model ID |
|---|---|
| Cloudflare Workers AI ↗ | @cf/cloudflare/clef-flash |
| Hugging Face (weights) ↗ | Cloudflare/clef-flash |
Clef (open decision models) — every version
The full lineage of the Clef (open decision models) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| Clefcurrent | 2026-10-01 | 64K | Apache-2.0 |
| Clef-flash | 2026-10-01 | 64K | Apache-2.0 |
FAQ
What is Clef-flash?
Clef-flash is Cloudflare's 9B open decision model, released October 1, 2026 alongside the 27B Clef. It answers typed questions about a text, JSON or image input with a probability per option, runs on Workers AI as @cf/cloudflare/clef-flash, and its weights are on Hugging Face under Apache 2.0.
How is it different from Clef?
It uses a Qwen3.5-9B backbone instead of Qwen3.8-27B, so it is faster (38.8 ms median latency against 209.3 ms in Cloudflare's tests) and cheaper ($0.09 against $0.24 per million input tokens on Workers AI). Cloudflare positions Clef as the precision model and Clef-flash for latency-critical decisions.
How does it compare with Jev?
In Cloudflare's launch table Clef-flash beats Jev on BFCL (98.76 against 95.75), API-Bank (93.11 against 88.19) and the home-appliances task (97.73 against 52.27), and trails it on CLINC150+OOS (66.77 against 89.27) and When2Call (65.58 against 80.97). Its 38.8 ms median latency compares with 524.1 ms for Jev.
Can I run it myself?
Yes. The weights are published on Hugging Face as Cloudflare/clef-flash under Apache 2.0, and the model card includes Jev / System One compatible usage. It is also hosted on Workers AI.