blog.cloudflare.com web signal

Cloudflare Ships Multimodal Clef-omni, Cuts Flash to $0.038/M

TL;DR

  • Clef-flash input price drops from $0.09 to $0.038 per million tokens, while its hosted context window shrinks from 64k to 24k.
  • New multimodal Clef-omni accepts audio, video, image and text in a single call at $0.15 per million input tokens.
  • A move to SGLang yields median 1.7x to 2.0x speedups on Clef, with the integration landing as PR #42721 in SGLang 0.5.22.

Cloudflare dropped the input price of Clef-flash from $0.09 to $0.038 per million tokens and launched a multimodal sibling, Clef-omni, that scores a 21-second video clip with sound in about 1.5 seconds for $0.15 per million. The announcement, posted October 9 by Michelle Chen on the Cloudflare blog, bills itself as a sequel: "Following last week's release of Clef and Clef-flash, Cloudflare's open-weight decision models, we decided to bring forth more gifts."

Clef-omni accepts audio, video, image and text input in a single call, built on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts foundation with what the post calls a two-stage attention routing approach. Text-only decisions return in roughly 130ms median latency, images in about 150ms.

The hosted Clef-flash context window shrinks from 64k to 24k tokens, which Cloudflare justifies with one line of usage data: "From our usage data, we see that only 0.24% of requests exceed 24k input tokens." Teams past that ceiling can self-host; the weights are public.

The second gift is a serving stack. Cloudflare moved Clef to SGLang and reports median speedups of 1.7x to 2.0x from infrastructure changes alone, with the integration landing as PR #42721 in SGLang 0.5.22. It arrives the same week we logged Tsinghua's TokenRouter hitting 2x to 64x decode speedups for token-level routing, one of 94 inference stories we've tracked in the last 90 days.