blog.cloudflare.com via Hacker News

Cloudflare Open-Sources Clef, Beats Jev Latency on Edge GPUs

TL;DR

  • Clef and Clef-flash ship under Apache 2.0 with a 64k context window and a vision encoder; Jev is text-only.
  • Median latency comes in at 209.3ms for Clef versus 524.1ms for Jev; Clef-flash reaches p95 of 122.4ms.
  • A reinforcement-learning service lets customers capture, fine-tune, and redeploy their own Clef variants on Workers AI.

Cloudflare released Clef and Clef-flash, two Apache 2.0 decision models with a 64k context window and a vision encoder, aimed squarely at Typesafe AI's Jev. The October 1, 2026 announcement from Michelle Chen pitches the pair as a way to produce "bounded structured outputs cheaply, quickly and consistently that can be added into a workflow when a decision is required."

The latency numbers do most of the talking. Median response clocks 209.3ms for Clef against 524.1ms for Jev. Clef-flash, the smaller sibling built on Qwen 3.5-9B, hits p95 of 122.4ms where Jev reports 536.0ms. On quality, Cloudflare's post lists 94.20 macro-F1 on BANKING77 versus 79.74 for Jev, and 98.47 on BFCL case-exact versus 95.75.

Both Clef variants ship with image classification. Jev is text-only.

On the naming, the post is deadpan: "A decision model is analogous to a music clef because it helps define the domain of the context and the subsequent notes (actions) that follow it."

A paired reinforcement-learning offering runs first through a forward-deployed engineer team, then as a self-serve stack stitching AI Gateway, Workers AI, and Containers together so customers can "capture data, fine-tune, and redeploy the model, all on Cloudflare." The post names no launch customers for that service.

The timing is pointed. The same day, Amazon shipped Strands Decider 2B, its own open-source Jev rival, making this the second challenger to the closed decision-model category we've logged in a single day across the 307 open-source stories we've tracked over the past 90 days.