Skip to content

Clef Decision Models Guide: Hands-On RL Fine-Tuning

Clef just dropped: open-source decision models plus RL fine-tuning on Workers AI. Here's how to call it, pick flash vs full, and avoid the real gotchas.

6 min readBeginner

Why agent hot-path decisions just got cheaper to get right

Your agent still waits on a full LLM to answer “is this urgent?” or “block this crawler?” then you parse half-broken JSON. That latency and cost pile up. On October 1, 2026 Cloudflare shipped Clef – open-source decision models plus an RL fine-tuning path – aimed at those bounded calls.

Clef and Clef-flash spit typed probabilities, not prose. No free-form text. No repair step. Jev-API compatible, Workers AI at the edge, vision included, Apache 2.0 weights on Hugging Face (Cloudflare’s launch post). First serious open-weight entry in the decision-model wave.

Quick context: what Clef actually returns

You send a state (string, JSON, images) plus up to 64 typed questions. Three shapes only:

  • noul – yes/no probability
  • choice – pick from your criteria map + per-option probs + confidence
  • score – ordered rubric, probability-weighted level

Clef = 27B precision (Qwen3.8-27B backbone). Clef-flash = 9B hot-path (Qwen3.5-9B). Both: 65,536-token window, vision encoder, up to four images. As of the October 2026 launch, the Workers AI model page lists Clef at $0.24 per million input tokens and flash at $0.09; output tokens billed at zero. Cloudflare’s 43-run set: flash median 38.8 ms (p95 122.4 ms), full Clef median 209.3 ms, Jev median 524.1 ms.

Hands-on: call Clef for a bot / content guardrail

Skip the usual checkout-ticket demo. Route a crawler or content flag – the call that belongs in the request path.

Grab a Workers AI token (Workers AI Read + Edit) and your account ID from the dashboard. Then:

export CLOUDFLARE_ACCOUNT_ID=your_id
export CLOUDFLARE_AUTH_TOKEN=your_token

curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef-flash 
 -X POST 
 -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" 
 -H "Content-Type: application/json" 
 -d '{
 "model": "clef-flash",
 "state": {
 "ua": "Mozilla/5.0 (compatible; ResearchBot/1.2)",
 "path": "/api/v1/export",
 "rate_last_min": 420,
 "note": "Hits export endpoint, ignores robots, residential ASN"
 },
 "questions": {
 "malicious": {
 "type": "noul",
 "instructions": "Is this crawler clearly abusive or scraping protected data?"
 },
 "action": {
 "type": "choice",
 "instructions": "What should the edge do next?",
 "criteria": {
 "allow": "Legitimate research or monitoring",
 "challenge": "Suspicious but not proven bad",
 "block": "Clear abuse or credential stuffing pattern"
 }
 },
 "severity": {
 "type": "score",
 "instructions": "How severe is the risk right now?",
 "criteria": ["None", "Low", "Medium", "High", "Critical"]
 }
 }
 }'

REST wraps answers under result. The Workers AI binding does not – you get the answers object straight. Expect shapes like answers.malicious.noul, answers.action.choice with per-option probs, plus a score you threshold yourself.

Inside a Worker the same body goes to env.AI.run("@cf/cloudflare/clef-flash", {...}). Page-on-call only if noul > 0.85 and confidence is solid. Policy lives in your code. The model never owns it.

Pro tip: start new questions on Clef-flash. Promote only where a labeled set shows flash drifting. Bill and latency stay in the cheap tier for most traffic.

Want vision? Images go in state (docs: up to four). Cloudflare’s threat-intel path already pairs this with Browser Run for domain screenshots – blog cites ~2.2 s end-to-end vs ~4.7 s for a large general LLM on a similar classify job.

False confidence feels worse than a blunt 0.5. If your criteria are empty, Clef will not invent a rubric for you – it just shrugs toward the middle.

Common pitfalls to avoid

Leave the happy path and three things bite fast.

  1. Image size vs context – File under the documented limit can still blow the 65k window. Launch-day community runs (see independent write-up) needed ~190 KB per image or smaller. Oversized inputs die with a context-window error even when the PNG “looks fine.”
  2. Vague instructions – “Does the ticket have repro steps?” lands near 0.5 noul. Rewrite to “Does the text name both the browser and the exact action that fails?” Calibration tracks your criteria, not vibes.
  3. Local hardware reality – Weights are free on Hugging Face, but full Clef wants ~85 GB VRAM and flash ~41 GB at 64k context / single concurrency (Cloudflare PM Michelle Chen, via The Register). Most laptops are out. Hosted Workers AI is the sane start.

Actually – watch the model field too. It must be exactly clef or clef-flash. A leftover jev-latest returns pattern error 5006.

Performance and the RL fine-tuning angle

Cloudflare’s benches say Clef leads seven of ten decision-index tasks. Headline number: BANKING77 macro-F1 94.20 vs Jev 79.74. Flash trades some accuracy for that sub-40 ms median. Day-one image calls in independent tests sometimes sat 13-30 s while text stayed snappy – park vision as async or offline until your own edge numbers look good.

Token list price sits above Jev. Pure high-volume text classification with no vision still favors the cheaper incumbent on dollars alone. Clef pulls ahead when you need vision, longer state, open weights, or the same edge network already serving your traffic. Free allocation: 10,000 Neurons/day on Workers AI (as of launch; confirm current dashboard).

The new piece: RL fine-tuning. Phase one is design-partner work with Forward-Deployed Engineers (interest form). Planned self-serve loop: AI Gateway captures production traffic → Workers AI rollouts → Containers as scoring sandbox → new Trainer updates weights → BYO Model redeploy. Internal Cloudflare already points this at support triage, bot classification, and Trust & Safety scoring. Exact self-serve date and dataset schema are still unstated. Training data stays private even though weights are Apache 2.0.

Is calibrated probability enough when the cost of a wrong block is a lawsuit? That’s the question only your risk model can answer.

When NOT to use Clef decision models

Skip it when you need free-form reasoning, multi-step tool planning, or creative generation – still an LLM job. Skip massive offline batches where every fraction of a cent wins and a classic classifier (or Jev) is enough. Skip local weights without datacenter GPUs. Don’t force vision into a user-facing hot path until your latency tests clear.

Already deep on Jev and happy with text-only? Migration is mostly endpoint + token + model name. Extra value is vision, context headroom, and the RL flywheel on the network that already carries your traffic.

FAQ

Is Clef free to try?

Yes, within Workers AI’s 10,000 Neurons/day. Past that you need Workers Paid. Input rates above apply; output isn’t billed.

Can I fine-tune Clef on my own labels today?

Not fully self-serve. Design-partner form + FDE team only. Stack is announced (Gateway capture → rollouts → Containers scoring → Trainer → redeploy); DIY schedule and data format aren’t public. Offline weight experiments are fine under Apache 2.0 – recipe and datasets stay closed.

Flash or full Clef for production agents?

Default flash on the request path. Take one afternoon: 50-100 labeled events, run both models, promote a question to full Clef only where flash confidence collapses or labels flip. Example: bot action.choice might stay on flash forever while a rare Trust & Safety score moves to 27B. Keep the LLM behind the decision, not in front of it – fire generation only after Clef says the path is worth it.

Next action: create the Workers AI token, paste the curl with one real event from your logs, and write down the three thresholds you’ll enforce. Then decide whether the design-partner RL form is worth filling this week.