Overview

Clef-omni is the multimodal member of Cloudflare's Clef decision models, released on October 9, 2026. One API call can carry text, a JSON state, images, audio (WAV or MP3) and video (MP4 or WebM), plus up to 64 typed questions. Instead of writing an answer, the model returns a probability for every allowed option. It runs on Workers AI as @cf/cloudflare/clef-omni, and the weights are on Hugging Face under Apache 2.0.
Cloudflare built it on Qwen3-Omni-30B-A3B-Instruct, a mixture-of-experts model, and removed the base model's text-to-speech output parts. The Qwen3 backbone stays frozen; training adds LoRA adapters and combines label-smoothed cross-entropy with a Brier score term so the probabilities are calibrated. Like the other Clef models it speaks the Jev API, so switching the model ID is enough to try it.
Cloudflare reports about 130 ms median latency for text-only decisions, about 150 ms with images, a few hundred milliseconds for audio clips and about 1.5 seconds for a 21-second video with sound. On its launch table Clef-omni posts the best Clef-family scores on BANKING77 (94.8), CLINC150+OOS (97.7) and Amazon ESCI (57.8), and trails Clef and Clef-flash on the home-appliances task (69.3).
| Released | 2026-10-09 |
|---|---|
| License | Apache-2.0 |
| Weights | Open weights |
| Parameters | 30B-A3B MoE (35B in BF16 safetensors) |
| Context | 64,000 tokens (Workers AI) |
| Architecture | Frozen Qwen3-Omni-30B-A3B-Instruct mixture-of-experts backbone with its text-to-speech output parts removed, plus LoRA adapters trained with label-smoothed cross-entropy and a Brier calibration loss; it returns one score per allowed option instead of generating text |
| Modalities | Text, Vision, Audio, Video |
Benchmarks
Clef-omni against the other Clef models and Jev, as published in Cloudflare's Clef-omni launch post (October 9, 2026).
| Benchmark | Clef-omni | Clef | Clef-flash | Jev |
|---|---|---|---|---|
| BFCL · case exact | 98.2 | 98.47 | 98.76 | 95.75 |
| ToolRet · nDCG@10 | 66.6 | 69.19 | 66.43 | 65.28 |
| API-Bank · accuracy | 92.7 | 91.93 | 93.11 | 88.19 |
| Home appliances · case exact | 69.3 | 82.95 | 97.73 | 52.27 |
| When2Call · accuracy | 63.3 | 72.37 | 65.58 | 80.97 |
| BANKING77 · macro-F1 | 94.8 | 94.2 | 90.93 | 79.74 |
| CLINC150+OOS · macro-F1 | 97.7 | 97.43 | 66.77 | 89.27 |
| BRIGHT · nDCG@10 | 42 | 45.91 | 39.26 | 47.52 |
| Amazon ESCI · macro-F1 | 57.8 | 57.48 | 57.39 | 55.21 |
| PhishNChips · accuracy | 73.2 | 79.6 | 75.05 | 62.55 |
This model's scores
- BFCL (case exact)98.2%
- CLINC150+OOS (macro-F1)97.7%
- BANKING77 (macro-F1)94.8%
- API-Bank (accuracy)92.7%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Pricing
| Input | $0.15 / 1M input tokens |
|---|
Workers AI price for @cf/cloudflare/clef-omni; Clef models do not charge for output tokens.
Strengths
- Text, images, audio and video in a single decision request
- Top Clef-family scores on BANKING77 (94.8) and CLINC150+OOS (97.7) in Cloudflare's launch table
- About 130 ms median latency for text-only decisions, 150 ms with images
- $0.15 per M input tokens on Workers AI, with no charge for output tokens
- Apache-2.0 weights on Hugging Face and a Jev-compatible request format
Best for
- Reach for it to route or triage support calls, voicemails or screen recordings without first transcribing them with a separate model.
- Reach for it for content checks on uploaded images and short videos, where each check is a yes/no or multiple-choice question.
- Reach for it for intent classification with many classes, where it posts its best scores (BANKING77, CLINC150+OOS).
- Reach for Clef or Clef-flash instead for text-only tool selection, where they score higher on the home-appliances task and When2Call.
How to access
| Provider | Model ID |
|---|---|
| Cloudflare Workers AI ↗ | @cf/cloudflare/clef-omni |
| Hugging Face (weights) ↗ | Cloudflare/clef-omni |
Clef (open decision models) — every version
The full lineage of the Clef (open decision models) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| Clef-omnicurrent | 2026-10-09 | 64K | Apache-2.0 |
| Clef | 2026-10-01 | 64K | Apache-2.0 |
| Clef-flash | 2026-10-01 | 24K | Apache-2.0 |
FAQ
What can Clef-omni read that Clef cannot?
Clef-omni adds audio (WAV or MP3) and video (MP4 or WebM) input to the text, JSON and images the other Clef models accept. On Workers AI a request can carry up to four images, four audio clips of up to 300 seconds each and two video clips of up to 60 seconds each, with all media embedded as base64.
How much does Clef-omni cost?
Clef-omni costs $0.15 per million input tokens on Workers AI, and Clef models do not charge for output tokens. That sits between Clef-flash at $0.038 and Clef at $0.24. The Apache-2.0 weights are also free to download from Hugging Face for self-hosting.
How fast is Clef-omni?
Cloudflare reports a median of about 130 ms for text-only decisions, about 150 ms with images, a few hundred milliseconds for audio clips and about 1.5 seconds for a 21-second video clip with sound.