Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 7 min read

Route Tickets and Intents Locally with GLiNER2.5-Decide

A 340M classifier from Fastino that scores intent, urgency, sentiment and routing labels you choose at call time, in one forward pass, on CPU or GPU.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

GLiNER2.5-Decide is an English classification model from Fastino, built on fastino/gliner2-large-v1 and the GLiNER2 schema-driven interface. You pass the labels when you call it. That covers intent, routing, sentiment, priority, moderation policy and multi-label tags, and it does not generate any tokens or need a prompt template. One call can score several label sets (the card calls them “heads”) over the same text. On Fastino’s own decision benchmark, the 340M model scores slightly higher than its 1B sibling (60.2% vs 59.6%) and a Qwen3.5-4B-based baseline (56.4%). You can run it locally and still pick your own labels.

Key specs

Parameters 340M
Encoder DeBERTa-v3-large
Language English
Runs on CPU or GPU, through gliner2
License Apache 2.0
Output One string per single-label head, a list per multi-label head

Exact-match accuracy on fastino/fast-decisions. The suite has 17 domains with 300 held-out examples each, and every model got the same text and candidate labels:

Model Avg
GLiNER2.5-Decide (340M) 60.2%
GLiNER2.5-Decide-1B 59.6%
JevK5 57.6%
GLiNER2.5-multi-Decide (287M) 56.7%
SemIf (Qwen3.5-4B) 56.4%
GLiFormer large-v1 49.0%
Laya Router 46.6%

Fastino published this benchmark, and it is the only evaluation on the card. The suite is English, and the card points to the dataset card for domain and label details. No variance is reported, so treat small gaps between models with caution. The 0.6-point gap between the 340M and 1B models is one example.

Install

Terminal window
pip install gliner2

Run it

quickstart.py
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")
result = model.classify_text(
"My subscription renewed on April 15 for ¥5,400 after the service was already down. Can I get that charge refunded?",
{"intent": [
"order_status", "refund_request", "cancel_subscription", "update_payment",
"login_problem", "shipping_delay", "bug_report", "speak_to_human", "other",
]},
)
print(result) # e.g. {"intent": "refund_request"}

The schema is a dict that maps a head name to its labels. A plain list makes a single-label head. For anything else, pass a dict with labels plus options such as multi_label, cls_threshold or prompt. You can also give labels as a dict of label to description. The model card says its outputs are “potential results”, so treat the sample outputs here as examples of the return format, not guaranteed answers.

Email and ticket triage in one call

A shared mailbox usually needs three answers per message: what the sender wants, how urgent it is, and which team owns it. The card’s email triage example scores all three heads in one call, so you run the model once per message instead of three times.

triage.py
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")
TRIAGE_SCHEMA = {
"intent": ["fyi", "request", "approval", "complaint", "newsletter", "security_alert"],
"urgency": ["low", "normal", "high", "critical"],
"route": ["support", "billing", "legal", "security", "finance", "archive"],
}
def triage(email: str) -> dict:
return model.classify_text(email, TRIAGE_SCHEMA)
def dispatch(decision: dict) -> str:
if decision["route"] == "archive" or decision["intent"] == "newsletter":
return "archive"
if decision["urgency"] in ("high", "critical"):
return f"page:{decision['route']}"
return f"queue:{decision['route']}"
if __name__ == "__main__":
email = (
"From: compliance@group.example\n"
"Subject: Protocol update — action required today\n\n"
"Please confirm the new retention rule is applied before Friday's audit."
)
decision = triage(email)
print(decision) # e.g. {"intent": "request", "urgency": "high", "route": "legal"}
print(dispatch(decision)) # e.g. page:legal

Your labels can be specific to your company. If a label name is ambiguous, give it a short description. The card shows this for a banking taxonomy:

described_labels.py
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")
result = model.classify_text(
"Please reset the card PIN. The new one never arrived and the old one is locked after three tries.",
{"intent": {
"labels": {
"card_pin_change": "The customer wants a new PIN or the current PIN replaced",
"card_lost": "The physical card is missing",
"balance_inquiry": "The customer wants the current balance",
},
}},
)
print(result) # e.g. {"intent": "card_pin_change"}

Guardrails around a chat agent

Several of the card’s examples are small yes/no or policy checks, and they fit around an LLM agent. Use spam detection and moderation before the agent sees a message, a handoff check to decide when a human should take over, and a “finished” check to judge whether the agent completed its goal from the final state of the trace. The model generates no tokens, so you would expect these checks to be cheaper than another LLM call. The card gives no cost or latency numbers, though, so measure this yourself.

guards.py
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")
def inbound_check(message: str) -> dict:
return model.classify_text(
message,
{
"label": ["spam", "ham"],
"policy": ["allow", "personal_data", "harassment", "scam", "violence", "spam"],
"handoff": ["yes", "no"],
},
)
def agent_finished(goal: str, last_state: str) -> bool:
text = f"Goal: {goal}\n{last_state}"
result = model.classify_text(text, {"finished": ["yes", "no"]})
return result["finished"] == "yes"
if __name__ == "__main__":
msg = "This is the third time I have explained the same missing refund. Stop the bot and get me a person."
check = inbound_check(msg)
print(check)
if check["label"] == "spam" or check["policy"] != "allow":
print("blocked")
elif check["handoff"] == "yes":
print("send to human queue")
else:
print("send to agent")
done = agent_finished(
"email the Q4 summary to every partner.",
"Last action: draft saved in the hub.\n"
"Send button is still disabled because two partners have no address.",
)
print("finished" if done else "not finished, retry or escalate")

In the card, spam, policy, handoff and finished are separate examples. Putting them into one schema here relies on the documented multi-head support. Check that combining heads doesn’t change your results compared with calling them separately.

Review analytics over a batch

The card’s sentiment example labels a review “mixed” but doesn’t say why. A multi-label “aspects” head returns each part of the product the review mentions. The card shows no batch API, so this example loops over the texts one at a time.

reviews.py
from collections import Counter
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")
SCHEMA = {
"sentiment": ["positive", "negative", "mixed", "neutral"],
"aspects": {
"labels": ["battery", "keyboard", "screen", "camera", "price", "support"],
"multi_label": True,
"cls_threshold": 0.4,
},
}
reviews = [
"Battery dies before lunch, but the keyboard and the screen are the best I have used on a laptop.",
"Support took a week to answer and the replacement was the wrong model.",
"Great value for the price. The camera is fine for calls.",
]
sentiments = Counter()
aspect_counts = Counter()
for text in reviews:
result = model.classify_text(text, SCHEMA)
sentiments[result["sentiment"]] += 1
aspect_counts.update(result["aspects"])
print(result)
print("sentiment:", dict(sentiments))
print("aspects:", aspect_counts.most_common())

For a numeric score, pass the scale as strings, as the card does with ["0", "1", ..., "10"] for a rating. The model returns the label as a string, so convert it with int(...) before you store or sort by it.

Gotchas

  • English only. Fastino points to GLiNER2.5-multi-Decide for multilingual input. That model scored 56.7% against this model’s 60.2% on the English suite.
  • It is not a general-purpose model. The card says it does not reason, explain or answer open questions. The prompt option asks a question about a passage, but the answer is still a choice from your labels (the card’s example uses yes/no). The model doesn’t extract or generate text.
  • Single-label heads probably always return a label. The card says single-label tasks “return one string” but doesn’t say what happens when no label fits. Our reading is that you still get one of your labels. Add a catch-all such as the card’s "other" so the model has somewhere to put text that fits nothing else.
  • Return types differ by head. Single-label heads return a string and multi-label heads return a list. Write the downstream code to handle both.
  • The threshold is yours to tune. The card uses cls_threshold: 0.4 in its multi-label examples but gives no guidance on choosing it. Tune it on your own data.
  • About 60% exact match on the vendor’s benchmark. That is the average across 17 domains, so on that suite the model misses a large share of examples. Run your own evaluation before you route money or medical requests without a human review step.
  • The card has no memory or latency numbers. It only says “CPU or GPU”. Benchmark throughput on your own hardware.
  • The pipeline tag says token-classification, but the card only documents classify_text. Load the model with gliner2’s AutoExtractor, not a transformers token-classification pipeline.

When to pick it

Pick GLiNER2.5-Decide when you need local, fixed-choice decisions over English text and you want to change the labels without retraining. Typical uses are intent routing, ticket queues, urgency, moderation policy, spam filtering, handoff gates and agent-completion checks. Running several heads in one pass is useful when one message needs intent, priority and routing together. On Fastino’s benchmark it edged out the 1B variant by 0.6 points and scored above a Qwen3.5-4B-based baseline, so a larger model isn’t clearly better for these tasks.

Skip it when the task needs an explanation, a free-text answer, extracted text or multi-step reasoning. Use an LLM for those. Skip it for non-English traffic too, and use the multilingual variant instead. If roughly 60% exact match on a vendor benchmark is too low for your risk level, use it as a first pass that a human or a larger model checks, not as the final decision.

Related