OFICIAL Cloudflare Blog

Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash

What happened
Based on Cloudflare Blog · Oct 09, 2026

Cloudflare expands its open-weight decision models with multimodal Clef-omni, faster Clef, and cheaper Clef-flash, enhancing workflow flexibility and accessibility.

Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
Cloudflare Blog — Cloudflare
Key points
·
Clef-omni processes audio, video, text, and images in a single API call using a Qwen3-Omni-30B-A3B-Instruct foundation
·
Clef-flash is now cheaper than Jev with a hosted context window reduced to 24k tokens
·
Clef-omni delivers median response times of 130 ms for text and 1.5 seconds for a 21-second video clip
Key numbers
·
The model leverages a Qwen3-Omni-30B-A3B-Instruct foundation and eliminates the need for separate transcription or captioning steps, enabling unified decision-making across multiple input types.
·
Clef-omni delivers faster response times compared to prior models, with median processing times of 130 ms for text, 150 ms for images, and a few hundred milliseconds for audio clips.
·
A 21-second video clip with sound is processed in approximately 1.

Cloudflare has introduced Clef-omni, a new open-weight decision model capable of processing audio, video, text, and images in a single pipeline, marking a shift from traditional text-only models. The model leverages a Qwen3-Omni-30B-A3B-Instruct foundation and eliminates the need for separate transcription or captioning steps, enabling unified decision-making across multiple input types. Clef-omni is designed to handle synchronized media elements, such as video with embedded audio, by mapping them into a single sequence for joint processing and scoring.

Clef-omni delivers faster response times compared to prior models, with median processing times of 130 ms for text, 150 ms for images, and a few hundred milliseconds for audio clips. A 21-second video clip with sound is processed in approximately 1.5 seconds via a single API call, demonstrating significant efficiency gains. The model’s architecture includes a two-stage attention routing system that aggregates evidence from all input modalities before computing confidence scores, ensuring robust performance across diverse inputs.

Cloudflare also reduced the cost of Clef-flash, making it cheaper than Jev while maintaining high performance, though its hosted context window has been reduced to 24k tokens from 64k. The model weights remain unchanged for self-hosting, supporting a 256k context window, and the pricing adjustment targets broader accessibility for workflow integration. Clef, the higher-tier model, retains a 64k context window and benefits from infrastructure optimizations for faster speeds in the hosted version on Workers AI.

The company highlighted real-world applications of its Clef models, including spam detection in its public GitHub docs repo and use cases across Cloudflare’s teams for classification tasks. Cloudflare emphasized that Clef’s capabilities allow non-specialists to perform domain-agnostic detections without requiring extensive machine learning expertise or custom training data. The announcement underscores Cloudflare’s commitment to iterative innovation in open-weight decision models, with further optimizations and expansions planned.

Original source → Deals on Clipraptor.com →