Route Tickets and Intents Locally with GLiNER2.5-Decide
A 340M classifier from Fastino that scores intent, urgency, sentiment and routing labels you choose at call time, in one forward pass, on CPU or GPU.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
GLiNER2.5-Decide is an English classification model from Fastino, built on fastino/gliner2-large-v1 and the GLiNER2 schema-driven interface. You pass the labels when you call it. That covers intent, routing, sentiment, priority, moderation policy and multi-label tags, and it does not generate any tokens or need a prompt template. One call can score several label sets (the card calls them “heads”) over the same text. On Fastino’s own decision benchmark, the 340M model scores slightly higher than its 1B sibling (60.2% vs 59.6%) and a Qwen3.5-4B-based baseline (56.4%). You can run it locally and still pick your own labels.
Key specs
| Parameters | 340M |
| Encoder | DeBERTa-v3-large |
| Language | English |
| Runs on | CPU or GPU, through gliner2 |
| License | Apache 2.0 |
| Output | One string per single-label head, a list per multi-label head |
Exact-match accuracy on fastino/fast-decisions. The suite has 17 domains with 300 held-out examples each, and every model got the same text and candidate labels:
| Model | Avg |
|---|---|
| GLiNER2.5-Decide (340M) | 60.2% |
| GLiNER2.5-Decide-1B | 59.6% |
| JevK5 | 57.6% |
| GLiNER2.5-multi-Decide (287M) | 56.7% |
| SemIf (Qwen3.5-4B) | 56.4% |
| GLiFormer large-v1 | 49.0% |
| Laya Router | 46.6% |
Fastino published this benchmark, and it is the only evaluation on the card. The suite is English, and the card points to the dataset card for domain and label details. No variance is reported, so treat small gaps between models with caution. The 0.6-point gap between the 340M and 1B models is one example.
Install
pip install gliner2Run it
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")
result = model.classify_text( "My subscription renewed on April 15 for ¥5,400 after the service was already down. Can I get that charge refunded?", {"intent": [ "order_status", "refund_request", "cancel_subscription", "update_payment", "login_problem", "shipping_delay", "bug_report", "speak_to_human", "other", ]},)print(result) # e.g. {"intent": "refund_request"}The schema is a dict that maps a head name to its labels. A plain list makes a single-label head. For anything else, pass a dict with labels plus options such as multi_label, cls_threshold or prompt. You can also give labels as a dict of label to description. The model card says its outputs are “potential results”, so treat the sample outputs here as examples of the return format, not guaranteed answers.
Email and ticket triage in one call
A shared mailbox usually needs three answers per message: what the sender wants, how urgent it is, and which team owns it. The card’s email triage example scores all three heads in one call, so you run the model once per message instead of three times.
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")
TRIAGE_SCHEMA = { "intent": ["fyi", "request", "approval", "complaint", "newsletter", "security_alert"], "urgency": ["low", "normal", "high", "critical"], "route": ["support", "billing", "legal", "security", "finance", "archive"],}
def triage(email: str) -> dict: return model.classify_text(email, TRIAGE_SCHEMA)
def dispatch(decision: dict) -> str: if decision["route"] == "archive" or decision["intent"] == "newsletter": return "archive" if decision["urgency"] in ("high", "critical"): return f"page:{decision['route']}" return f"queue:{decision['route']}"
if __name__ == "__main__": email = ( "From: compliance@group.example\n" "Subject: Protocol update — action required today\n\n" "Please confirm the new retention rule is applied before Friday's audit." ) decision = triage(email) print(decision) # e.g. {"intent": "request", "urgency": "high", "route": "legal"} print(dispatch(decision)) # e.g. page:legalYour labels can be specific to your company. If a label name is ambiguous, give it a short description. The card shows this for a banking taxonomy:
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")
result = model.classify_text( "Please reset the card PIN. The new one never arrived and the old one is locked after three tries.", {"intent": { "labels": { "card_pin_change": "The customer wants a new PIN or the current PIN replaced", "card_lost": "The physical card is missing", "balance_inquiry": "The customer wants the current balance", }, }},)print(result) # e.g. {"intent": "card_pin_change"}Guardrails around a chat agent
Several of the card’s examples are small yes/no or policy checks, and they fit around an LLM agent. Use spam detection and moderation before the agent sees a message, a handoff check to decide when a human should take over, and a “finished” check to judge whether the agent completed its goal from the final state of the trace. The model generates no tokens, so you would expect these checks to be cheaper than another LLM call. The card gives no cost or latency numbers, though, so measure this yourself.
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")
def inbound_check(message: str) -> dict: return model.classify_text( message, { "label": ["spam", "ham"], "policy": ["allow", "personal_data", "harassment", "scam", "violence", "spam"], "handoff": ["yes", "no"], }, )
def agent_finished(goal: str, last_state: str) -> bool: text = f"Goal: {goal}\n{last_state}" result = model.classify_text(text, {"finished": ["yes", "no"]}) return result["finished"] == "yes"
if __name__ == "__main__": msg = "This is the third time I have explained the same missing refund. Stop the bot and get me a person." check = inbound_check(msg) print(check)
if check["label"] == "spam" or check["policy"] != "allow": print("blocked") elif check["handoff"] == "yes": print("send to human queue") else: print("send to agent")
done = agent_finished( "email the Q4 summary to every partner.", "Last action: draft saved in the hub.\n" "Send button is still disabled because two partners have no address.", ) print("finished" if done else "not finished, retry or escalate")In the card, spam, policy, handoff and finished are separate examples. Putting them into one schema here relies on the documented multi-head support. Check that combining heads doesn’t change your results compared with calling them separately.
Review analytics over a batch
The card’s sentiment example labels a review “mixed” but doesn’t say why. A multi-label “aspects” head returns each part of the product the review mentions. The card shows no batch API, so this example loops over the texts one at a time.
from collections import Counter
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")
SCHEMA = { "sentiment": ["positive", "negative", "mixed", "neutral"], "aspects": { "labels": ["battery", "keyboard", "screen", "camera", "price", "support"], "multi_label": True, "cls_threshold": 0.4, },}
reviews = [ "Battery dies before lunch, but the keyboard and the screen are the best I have used on a laptop.", "Support took a week to answer and the replacement was the wrong model.", "Great value for the price. The camera is fine for calls.",]
sentiments = Counter()aspect_counts = Counter()
for text in reviews: result = model.classify_text(text, SCHEMA) sentiments[result["sentiment"]] += 1 aspect_counts.update(result["aspects"]) print(result)
print("sentiment:", dict(sentiments))print("aspects:", aspect_counts.most_common())For a numeric score, pass the scale as strings, as the card does with ["0", "1", ..., "10"] for a rating. The model returns the label as a string, so convert it with int(...) before you store or sort by it.
Gotchas
- English only. Fastino points to
GLiNER2.5-multi-Decidefor multilingual input. That model scored 56.7% against this model’s 60.2% on the English suite. - It is not a general-purpose model. The card says it does not reason, explain or answer open questions. The
promptoption asks a question about a passage, but the answer is still a choice from your labels (the card’s example uses yes/no). The model doesn’t extract or generate text. - Single-label heads probably always return a label. The card says single-label tasks “return one string” but doesn’t say what happens when no label fits. Our reading is that you still get one of your labels. Add a catch-all such as the card’s
"other"so the model has somewhere to put text that fits nothing else. - Return types differ by head. Single-label heads return a string and multi-label heads return a list. Write the downstream code to handle both.
- The threshold is yours to tune. The card uses
cls_threshold: 0.4in its multi-label examples but gives no guidance on choosing it. Tune it on your own data. - About 60% exact match on the vendor’s benchmark. That is the average across 17 domains, so on that suite the model misses a large share of examples. Run your own evaluation before you route money or medical requests without a human review step.
- The card has no memory or latency numbers. It only says “CPU or GPU”. Benchmark throughput on your own hardware.
- The pipeline tag says token-classification, but the card only documents
classify_text. Load the model withgliner2’sAutoExtractor, not atransformerstoken-classification pipeline.
When to pick it
Pick GLiNER2.5-Decide when you need local, fixed-choice decisions over English text and you want to change the labels without retraining. Typical uses are intent routing, ticket queues, urgency, moderation policy, spam filtering, handoff gates and agent-completion checks. Running several heads in one pass is useful when one message needs intent, priority and routing together. On Fastino’s benchmark it edged out the 1B variant by 0.6 points and scored above a Qwen3.5-4B-based baseline, so a larger model isn’t clearly better for these tasks.
Skip it when the task needs an explanation, a free-text answer, extracted text or multi-step reasoning. Use an LLM for those. Skip it for non-English traffic too, and use the multilingual variant instead. If roughly 60% exact match on a vendor benchmark is too low for your risk level, use it as a first pass that a human or a larger model checks, not as the final decision.
Related
Laya: Typed Routing and Triage Decisions in One ~33 ms Pass
Laya answers typed questions about text with probabilities instead of generated text. It covers 100+ languages and is Apache 2.0. Fit temperatures before trusting the probabilities.
convaiinnovations/laya
Clef: Cloudflare's model that returns probabilities, not text
Clef answers typed questions about any input (text, JSON, images) with a probability per option in one forward pass. Here's how to use it for routing, triage and classification.
Cloudflare/clef
CLM-v0.1-8B: rank candidates and route agent decisions fast
A contrastive scorer on frozen Qwen3-8B embeddings. It ranks the candidates you give it and answers typed questions about a state.
Contrastive-LM/CLM-v0.1-8B