Clef: Cloudflare's model that returns probabilities, not text
Clef answers typed questions about any input (text, JSON, images) with a probability per option in one forward pass. Here's how to use it for routing, triage and classification.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Most LLM “classification” means prompting a model, getting text back, and parsing it. Clef skips all of that. You give it a state (text, JSON, images or video) and a schema of typed questions, and it returns a probability for every allowed option in a single forward pass. There’s no generation and no parsing, and you get confidence scores for free.
It’s a post-train of Qwen3.8-27B with a small “joint schema head” on top. A faster sibling, Clef-Flash, has a median latency of ~39 ms, compared with ~209 ms for Clef.
Three question types
| Type | Answer | Example |
|---|---|---|
noul |
P(true) | “Is a service down?” |
choice |
distribution over named options | department: billing / technical |
score |
distribution over ordered options | urgency: can wait → today |
Support-ticket triage in one call
Clef ships its own loader code in the repo, so you snapshot_download it and import from there:
import sysfrom huggingface_hub import snapshot_download
path = snapshot_download("Cloudflare/clef")sys.path.insert(0, path)from joint_schema_model import load_release_model, systemone
model, processor = load_release_model(path, device="cuda")
response = systemone(model, processor, { "model": "clef", "state": "Our checkout started returning errors and orders are blocked.", "questions": { "department": { "type": "choice", "instructions": "Which team should handle the message?", "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}, }, "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]}, "outage": {"type": "noul", "instructions": "Is a service down?"}, },})print(response["answers"])Every answer includes confidence and probabilities. You can use those to send low-confidence tickets to a human. That’s a lot harder to do with free-text output.
Images work the same way
from PIL import Image
record = { "state": {"task": "Review the attached receipt."}, "images": [Image.open("receipt.jpg")], "questions": {"legible": {"type": "noul", "instructions": "Is the receipt total legible?"}},}Where it shines (and where it doesn’t)
On Cloudflare’s Decision Index it is excellent at intent classification (BANKING77 94.2, CLINC150 97.4), tool selection (BFCL 98.5) and hallucination detection (RAGTruth 79.4). It’s much weaker on hard reasoning (GPQA Diamond 48.0). Use Clef as a fast router and judge, and send hard questions to a reasoning model.
Good fits: ticket routing, content moderation queues, invoice and receipt checks, choosing which tool or agent to call, and RAG answer validation.
Related
Route Tickets and Intents Locally with GLiNER2.5-Decide
A 340M classifier from Fastino that scores intent, urgency, sentiment and routing labels you choose at call time, in one forward pass, on CPU or GPU.
fastino/GLiNER2.5-Decide
Run JEV-27B-VL for calibrated yes/no and choice decisions on images
A 27B vision decision model that returns a calibrated probability for every option in one forward pass. Here's how to serve it and use it.
autotrust/JEV-27B-VL
Laya: Typed Routing and Triage Decisions in One ~33 ms Pass
Laya answers typed questions about text with probabilities instead of generated text. It covers 100+ languages and is Apache 2.0. Fit temperatures before trusting the probabilities.
convaiinnovations/laya