Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 2 min read

Clef: Cloudflare's model that returns probabilities, not text

Clef answers typed questions about any input (text, JSON, images) with a probability per option in one forward pass. Here's how to use it for routing, triage and classification.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Most LLM “classification” means prompting a model, getting text back, and parsing it. Clef skips all of that. You give it a state (text, JSON, images or video) and a schema of typed questions, and it returns a probability for every allowed option in a single forward pass. There’s no generation and no parsing, and you get confidence scores for free.

It’s a post-train of Qwen3.8-27B with a small “joint schema head” on top. A faster sibling, Clef-Flash, has a median latency of ~39 ms, compared with ~209 ms for Clef.

Three question types

Type Answer Example
noul P(true) “Is a service down?”
choice distribution over named options department: billing / technical
score distribution over ordered options urgency: can wait → today

Support-ticket triage in one call

Clef ships its own loader code in the repo, so you snapshot_download it and import from there:

triage.py
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("Cloudflare/clef")
sys.path.insert(0, path)
from joint_schema_model import load_release_model, systemone
model, processor = load_release_model(path, device="cuda")
response = systemone(model, processor, {
"model": "clef",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
})
print(response["answers"])

Every answer includes confidence and probabilities. You can use those to send low-confidence tickets to a human. That’s a lot harder to do with free-text output.

Images work the same way

from PIL import Image
record = {
"state": {"task": "Review the attached receipt."},
"images": [Image.open("receipt.jpg")],
"questions": {"legible": {"type": "noul", "instructions": "Is the receipt total legible?"}},
}

Where it shines (and where it doesn’t)

On Cloudflare’s Decision Index it is excellent at intent classification (BANKING77 94.2, CLINC150 97.4), tool selection (BFCL 98.5) and hallucination detection (RAGTruth 79.4). It’s much weaker on hard reasoning (GPQA Diamond 48.0). Use Clef as a fast router and judge, and send hard questions to a reasoning model.

Good fits: ticket routing, content moderation queues, invoice and receipt checks, choosing which tool or agent to call, and RAG answer validation.

Related