← Back to all posts
News

Cloudflare Releases Clef, Open-Weight Jev-Compatible Decision Models Under Apache 2.0

October 2, 2026 · 00:06 UTC · News
Cloudflare Releases Clef, Open-Weight Jev-Compatible Decision Models Under Apache 2.0

TL;DR

Cloudflare released Clef and Clef-flash on October 1: two decision models that answer typed questions with probabilities, accept the same request shape as TypeSafe's Jev, and add image input plus a 64k context window. Both are hosted on Workers AI and published as Apache 2.0 weights on Hugging Face. On Cloudflare's own benchmark run, Clef-flash posts a 38.8 ms median latency against Jev's 524.1 ms, and Clef beats Jev on most of the tasks in the table. The catch is price: Clef costs $0.24 per million input tokens and Clef-flash $0.09, against Jev's $0.042.


What shipped

Two models, one API, one fine-tuning pitch:

  • Clef: a frozen Qwen3.8-27B backbone plus a trained routing head and rank-256 LoRA adapters. Workers AI model id @cf/cloudflare/clef, 65,536-token context, $0.24 per million input tokens.
  • Clef-flash: the same recipe on Qwen3.5-9B, priced at $0.09 per million input tokens on Workers AI.
  • Inputs: text, JSON, and up to four images per request (4 MiB each), with 1 to 64 questions per call. Jev is text-only today with a 32k window, per Cloudflare.
  • An RL fine-tuning service, starting as a hands-on engagement with Cloudflare's forward-deployed engineers and later meant to become self-serve on AI Gateway, Containers, and a new Trainer component.

The Hugging Face repos went up on September 30 and both carry an Apache 2.0 license. The Clef checkpoint is 12 safetensors shards, roughly 27.4B parameters, plus a separate joint_head.safetensors and a custom joint_schema_model.py, so this is not a stock chat model you point Ollama at. Cloudflare links a vLLM pull request from the related DiffusionGemma work; plan on reading code before you serve it locally.

The API is the point

Cloudflare says Clef is "fully Jev-API compatible." The example request is the Jev shape almost verbatim: a state string plus a questions map with noul (yes/no/unknown), choice, and score types.

curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
  -d '{
    "model": "clef",
    "state": "Checkout has been failing for every customer for the last hour.",
    "questions": {
      "urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
      "severity": { "type": "score", "instructions": "How severe is the customer impact?",
        "criteria": ["No impact", "Minor", "Major", "Critical"] }
    }
  }'

That makes the decision-model layer swappable. If you built routing, triage, or guardrail code against Jev over the past few weeks, you now have a second hosted vendor and a self-hostable fallback without rewriting your schemas. A month ago "decision model" meant one company's product. It now means a wire format.

How it works

Clef runs the Qwen backbone in a single prefill pass over the state and questions, then scores every valid schema option in parallel. Nothing is generated token by token. A two-stage attention router lets each candidate option pull relevant evidence from the prompt, lets fields cross-attend to each other, and then a schema-bound head scores them. Training used label-smoothed cross-entropy plus a Brier loss for calibration on synthetic data that permutes field orders, prompts, and schemas, followed by an RL stage Cloudflare calls RLCD (Reinforcement Learning for Calibrated Decisions) that gives partial credit to adjacent ordinal answers.

If that sounds abstract: an LLM answering a multiple-choice form writes an essay and then circles a letter. Clef reads the form once and puts a probability next to every box at the same time. The frozen backbone does the reading; the small trained head does the circling.

state +schema frozen Qwenprefill only score all optionsin parallel no decode loop: probabilities come straight off the backbone
Clef never writes text: one prefill pass, then a trained head scores every allowed answer at once.

The numbers (Cloudflare's numbers)

Cloudflare scored six models on tasks it took from the Jev Decision Index: Clef, Clef-flash, Jev, a DiffusionGemma-based Jev clone, Kev 9B, and Convai's Laya. The full table is on a live benchmark site. Highlights:

  • Intent classification: Clef scores 94.20 macro-F1 on BANKING77 versus Jev's 79.74, and 97.43 on CLINC150+OOS versus 89.27.
  • Tool calling: 98.47 on BFCL case-exact for Clef and 98.76 for Clef-flash, versus 95.75 for Jev.
  • Where Jev wins: When2Call (80.97 vs 72.37), BRIGHT retrieval (47.52 vs 45.91), and agent-trace observability in TypeSafe's own workflow suite (71.6 vs 68.5). The DiffusionGemma Jev clone tops PhishNChips at 85.35, ahead of Clef's 79.60.
  • Flash is not a free lunch: Clef-flash drops to 66.77 on CLINC150+OOS, far below Jev's 89.27, even as it beats the big Clef on the home-appliances set (97.73 vs 82.95).
score (higher is better), Cloudflare-run evals Clef Jev BANKING77 94.20 79.74 CLINC150 97.43 89.27 When2Call 72.37 80.97 BRIGHT 45.91 47.52 BANKING77/CLINC150: macro-F1 · When2Call: acc · BRIGHT: nDCG@10
Clef wins big on intent classification and loses on knowing when to call a tool.

Treat all of this as a vendor benchmark. Cloudflare picked the subset, ran it, and hosts the leaderboard. Latency in particular depends on where you call from: TypeSafe notes its own published evals run from laptops on the US West Coast, and Cloudflare's models sit on its edge GPUs. A 524 ms median for Jev in Cloudflare's run is not the same claim as Jev being slow from your servers.

Latency vs price

Here is the trade in one picture. Clef-flash is the fastest of the serious contenders in Cloudflare's run, only Laya is quicker, and Laya scores in the low teens on most quality tasks.

median latency, ms (lower is better), 43 evals Clef-flash38.8 Kev 9B51.4 DiffGemma Jev84.4 Clef209.3 Jev524.1 hosted price, $ per million input tokens Jev0.042 Clef-flash0.09 Clef0.24
Clef-flash is about 13x faster than Jev in Cloudflare's run but about 2x the input price; Clef is about 5.7x.

Commenters on the Hacker News thread (around 400 points within hours) did the same math. At $0.24, the flagship is the expensive way to make a yes/no call. Clef-flash at $0.09 is the realistic default for hot-path routing. And because the weights are Apache 2.0, the price ceiling is really your own GPU bill: a 9B backbone (about 19 GB of bf16 weights) fits on a single 24 GB card, assuming the custom head ports cleanly.

Why this matters

Jev launched a category; within weeks, Kev, Laya, community clones, and now a public cloud with its own GPUs all shipped compatible alternatives. One HN commenter summed up TypeSafe's position as building a company on one model and getting cloned within a week. That is harsh but not wrong about the moat: the pattern (frozen backbone, scoring head, calibration loss) turned out to be reproducible, and Cloudflare published the recipe in a blog post.

For builders this is mostly good news. Decision calls are the cheap, boring glue between agents and actions: route this ticket, is this domain phishing, should this crawler get a challenge. That glue now has multiple vendors, a common request format, open weights you can audit, and a vision-capable option. Cloudflare's own example is a threat-intel workflow where Clef plus a headless browser classified a domain in 2.2 s versus 4.7 s for gpt-oss-120b.

Two caveats stay serious. First, "open-source" here means open weights: the training data, synthetic generators, and RLCD pipeline are not published. Second, nobody independent has yet reproduced the leaderboard, and the image-classification claims ship without any vision benchmark at all. Run your own eval set before you swap a production classifier.

Key Takeaways

  • Cloudflare released Clef (Qwen3.8-27B backbone) and Clef-flash (Qwen3.5-9B) as Apache 2.0 weights on Hugging Face and hosted on Workers AI.
  • Both accept the Jev request format, so existing Jev integrations can switch vendors with minimal code changes.
  • They add image input and a 64k context window, versus Jev's text-only 32k, per Cloudflare.
  • Cloudflare's benchmarks show big wins on intent classification and latency, and losses on When2Call, BRIGHT, and agent-trace observability.
  • Hosted input pricing is $0.24 per million tokens for Clef and $0.09 for Clef-flash, versus $0.042 for Jev.
  • Every number is vendor-run; the training pipeline is closed, and no vision benchmark was published.

Sources: Cloudflare blog: Introducing Clef, Hugging Face: Cloudflare/clef, Hugging Face: Cloudflare/clef-flash, Workers AI docs: Clef, Workers AI docs: Clef-flash, Clef live benchmark site, TypeSafe: Introducing System One Models and Jev, Hacker News discussion

AICloudflareClefDecision ModelsJevOpen WeightsWorkers AIQwen
CONSOLE
$