█

AI/TLDR

Cloudflare · 2026-10-09 · major

Clef-omni — Cloudflare's open decision model now reads audio and video

Clef-omni is Cloudflare's new Apache-2.0 decision model on Qwen3-Omni-30B-A3B. It scores typed questions about text, images, audio and video in one call. Cloudflare also made Clef up to 2x faster and cut Clef-flash to $0.038 per 1M tokens.

Cloudflare blog header for the Clef-omni launch

Cloudflare's Clef-omni answers yes/no and multiple-choice questions about audio, video, images and text in one call.

Key specs

Text decision latency (median)~130 ms
Clinc150+oos (macro f1)97.7

Quick facts

MakerCloudflare
Base modelQwen3-Omni-30B-A3B-Instruct (MoE)
InputsText, JSON, images, audio, video
LicenseApache-2.0
Context window64K tokens on Workers AI
Model ID@cf/cloudflare/clef-omni

Pricing

Clef-omni input$0.15 / 1M tokens
Clef input$0.24 / 1M tokens
Clef-flash input · was $0.09$0.038 / 1M tokens
Output tokens · Clef models do not charge for output$0
source ↗

What is it?

Clef-omni adds audio (WAV, MP3) and video (MP4, WebM) to Cloudflare's Clef decision models, released on October 9, 2026. Like Clef and Clef-flash, it does not write text: it takes an input plus typed questions and returns a calibrated probability for each allowed answer, through the Jev-compatible API.

How does it work?

Under the hood is Qwen3-Omni-30B-A3B-Instruct, a mixture-of-experts model, with its text-to-speech output parts removed. Cloudflare keeps that backbone frozen and trains LoRA adapters with label-smoothed cross-entropy plus a Brier score loss, so the scores behave like real probabilities.

Why does it matter?

Agents that route calls, check uploads or triage screen recordings can now make a typed decision on raw media without a separate transcription or captioning step. The same launch cuts Clef-flash's price by more than half, making cheap high-volume routing on Workers AI cheaper still.

Who is it for?

developers building agent routing, moderation and triage on media

Frequently asked questions

How fast is Clef-omni on audio and video?
Cloudflare reports that Clef-omni answers text-only decisions in about 130 ms at the median and image decisions in about 150 ms. Audio clips take a few hundred milliseconds, and a 21-second video clip with sound takes about 1.5 seconds. Workers AI accepts up to two embedded video clips of up to 60 seconds each per request.
How does Clef-omni compare with Clef and Jev?
In Cloudflare's launch table Clef-omni has the best Clef-family scores on BANKING77 (94.8), CLINC150+OOS (97.7) and Amazon ESCI (57.8), ahead of Jev's 79.74, 89.27 and 55.21. It trails Clef and Clef-flash on the home-appliances task (69.3 against 82.95 and 97.73) and Jev on When2Call (63.3 against 80.97).
What changed for existing Clef and Clef-flash users?
Cloudflare moved hosted Clef to SGLang, which makes the same weights 1.7x to 2.0x faster at the median; ~800-token inputs dropped from 262 ms to 152 ms. Clef-flash fell from $0.09 to $0.038 per million input tokens, but its hosted context window shrank from 64K to 24K tokens. Cloudflare says only 0.24% of requests were longer than that.
Can I run Clef-omni on my own GPUs?
Yes. Cloudflare published the Clef-omni weights on Hugging Face as Cloudflare/clef-omni under Apache 2.0, and the model card includes Python code that loads the model and prints a probability for each option. Cloudflare's SGLang changes for Clef are also upstream in SGLang 0.5.22.

Try it

@cf/cloudflare/clef-omni on Workers AI

Sources · 3 outlets

Tags

  • clef-omni
  • clef
  • clef-flash
  • cloudflare
  • decision-model
  • multimodal
  • audio
  • video
  • jev
  • qwen
  • workers-ai
  • open-weights
  • apache-2-0
  • sglang

← All releases · Learn AI