Skip to content

Clef Decision Models Guide: Use Cloudflare’s New RL Platform

Clef open-weight decision models just dropped on Workers AI with an RL fine-tuning platform. Here's how to call them, gate on confidence, and avoid the image token trap.

7 min readBeginner

Here’s the part that still feels weird a day after launch: Clef never writes a sentence. You hand it state plus a schema of typed questions, and it returns calibrated probabilities only – no JSON repair, no “let me think” tokens. Cloudflare dropped Clef open-weight decision models and a new RL fine-tuning platform on October 1, 2026, and the HN thread lit up within hours.

I needed a first-pass gate in front of a support inbox: is this toxic, which queue owns it, should a human see it before any LLM drafts a reply. General models were slow and chatty. Decision models are built for that exact fork in the road.

What actually shipped (and why agents care)

Clef and Clef-flash are the first models the Workers AI team trained in-house. Hosted IDs: @cf/cloudflare/clef (27B, post-trained from Qwen3.8-27B) and @cf/cloudflare/clef-flash (9B, from Qwen3.5-9B). Apache 2.0 weights sit on Hugging Face. The Cloudflare launch post is blunt about the surface area – vision encoder kept, text/JSON/images/video in, 65,536-token context on Workers AI, up to four images and 64 questions per request.

209.3 ms median for Clef. 38.8 ms for Clef-flash. Jev on the same 43-run suite: 524.1 ms. How? Non-autoregressive path – Qwen backbone does prefill only, then a joint schema head scores every allowed option of every question in one shot. Output tokens bill as zero because nothing is being written. BANKING77 macro-F1 on Cloudflare’s table: Clef 94.20 vs Jev 79.74. When2Call and a few reasoning-style rows still favor Jev; the hot path on the published suite favors Clef.

Three question types, System One / Jev shapes:

  • noul – yes/no; returns P(yes)
  • choice – one named option; choice + per-option probs + confidence
  • score – ordered rubric; probability-weighted level

Pro tip: Start every new workflow on Clef-flash. Promote individual questions to full Clef only when your labeled set shows Flash wrong or under-confident. Paying roughly 3× more per input token for every easy noul is how bills creep.

Practical setup: first call in under ten minutes

Grab an account ID and a Workers AI token from the dashboard (template with Workers AI Read + Edit). Export them, hit REST. Moderation + routing gate below – not the blog’s checkout ticket – so you can drop it in front of a real queue.

curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef-flash 
 -X POST 
 -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" 
 -d '{
 "model": "clef-flash",
 "state": {
 "message": "Your product is garbage and your CEO should be fired. Refund me or I post the chat logs.",
 "channel": "email",
 "account_age_days": 3
 },
 "questions": {
 "toxic": {
 "type": "noul",
 "instructions": "Does message contain personal attacks, threats, or severe harassment?"
 },
 "queue": {
 "type": "choice",
 "instructions": "Which queue should own this?",
 "criteria": {
 "trust_safety": "Abuse, threats, harassment",
 "billing": "Refunds, charges, invoices",
 "product": "Bugs or feature complaints without abuse",
 "human": "Ambiguous or high-risk"
 }
 },
 "priority": {
 "type": "score",
 "instructions": "How fast must a human respond?",
 "criteria": ["Can wait a day", "Same business day", "Within an hour", "Immediate"]
 }
 }
 }'

REST wraps the System One body under result with the usual success / errors envelope. The Workers AI binding returns the object directly. And model must be exactly clef or clef-flash – a leftover jev-latest string fails pattern 5006.

Inside a Worker: bind AI in wrangler.jsonc, call env.AI.run("@cf/cloudflare/clef-flash", { ... }). Thresholds live in your code. Example policy I actually use: page on-call only if toxic.noul > 0.85 and queue.confidence >= 0.55; else human. The model judges. You own the policy.

As of the October 2026 pricing page: Clef-flash $0.09 per million input tokens, Clef $0.24, output not billed. Free tier still 10,000 Neurons/day on Free and Paid Workers plans; paid overage $0.011 per 1,000 Neurons. Neuron burn depends on the model – enough to prototype hard before a card is required.

Advanced usage: confidence gates, images, and the RL path

Flash said 0.54. Full Clef said 0.04. Same vague question: “does this ticket include repro steps?” when the text only named a browser and a button. That split is the whole lesson. Tighten instructions until they name the evidence you require, then re-test. Low confidence on choice/score should short-circuit to a human – not to another LLM tool call.

Vision is the real gap versus text-only Jev. Four images max. The catch is tokenization order: Workers AI can bill a large PNG from base64 length before the vision encoder runs. Community testing on launch day (flaviocopes) blew past the 65,536 context limit on multi-hundred-KB originals; Workers AI effectively counted base64/4. Resize to ~1024px-wide JPEG at ~80% quality – under ~190KB – and billed image tokens drop into the low thousands. Keep each image small on purpose.

Model Best for Input $/M Median latency (CF suite) Self-host VRAM*
Clef-flash Hot path, high QPS $0.09 38.8 ms ~41 GB
Clef Precision / vision-heavy $0.24 209.3 ms ~85 GB

*Per Cloudflare PM comments to The Register, single concurrency, 64k context – this may have changed since.

Open weights mean local runs via the HF card’s load_release_model + systemone helpers when edge GPUs or data residency matter. Hosted still wins if you do not already own H200-class iron. At ~300 tokens per decision, napkin math on published rates puts a million calls near $12-13 on Jev versus ~$72 on full Clef; Flash narrows it but is still not the cheapest decision model. If pure $/decision dominates and you need neither vision nor 64k state, measure both.

RL fine-tuning is FDE / design-partner only today. Planned loop: AI Gateway captures traffic as a dataset → Workers AI rollouts → Containers as the RL sandbox → Trainer updates weights → Bring-Your-Own-Model redeploy on Workers AI. Interest form is linked from the changelog. No public self-serve date. Training data for base Clef stays private even though weights are Apache 2.0 – confirmed in press.

Honest limitations before you rewrite production

Coin-flip noul is the model saying the question is underspecified. Calibrated scores do that on purpose. There is no temperature knob to hide behind.

Some HN first-pass hate-speech tests called Clef slower or weaker than Jev. Benchmarks are not your workload. Run a 200-row labeled set before you swap endpoints. Day-one vision latency was spikier than plain text in independent checks even though Cloudflare’s threat-intel Browser Run story clocked ~2.2s end-to-end versus ~4.7s on gpt-oss-120b – re-measure vision after launch dust settles.

FAQ

Is Clef a drop-in for Jev?

Mostly. Same state/questions/answer shapes. Swap endpoint and auth, set model to clef or clef-flash, read result.answers on REST. Done.

When should I pick Clef-flash over Clef?

Default Flash on the request path – routing, bot allow/deny, “should I call this tool.” Promote when Flash confidence stays low on your labels, when multimodal judgment is heavy, or when one wrong queue is expensive. Remember the 0.54 vs 0.04 repro-steps split: your code can require model agreement or just escalate.

Can I fine-tune Clef on my own labels today?

Not self-serve. Design-partner FDE path only. Self-serve (Gateway datasets → Containers scoring → Trainer → BYO deploy) is the announced destination, not a button this week. If you only need light adaptation, tighten question instructions and confidence floors first – that fixed more of my false routes than swapping 9B for 27B.

Next action: copy the curl above, point it at ten real tickets from last week, and write down the confidence floor you’ll enforce before any LLM sees the message. Then check live scores on Cloudflare’s decision index demo and the Hugging Face Clef card if you want the local path.