Laya: Typed Routing and Triage Decisions in One ~33 ms Pass
Laya answers typed questions about text with probabilities instead of generated text. It covers 100+ languages and is Apache 2.0. Fit temperatures before trusting the probabilities.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Laya is a non-autoregressive decision model from Convai Innovations. You give it a state, which can be text, an email, a ticket or JSON, plus a set of typed questions. It returns typed answers with probabilities in a single forward pass. It never generates text, so you don’t parse output and it can’t invent answers outside the options you define. What sets it apart is the training method. Convai Innovations trained it with RLCD (Reinforcement Learning for Calibrated Decisions), where the reward is a strictly proper scoring rule. Under that reward, the model gets the highest expected score only when it reports honest probabilities. Even so, the card says the checkpoints ship over-confident, so you need to fit temperatures on your own data before you trust the numbers (see Gotchas). The repo holds three checkpoints: English, multilingual and one fine-tuned on a typed-decisions benchmark. A built-in Router picks the checkpoint for each input.
Key specs
| Checkpoint | Backbone | Params | Context | Best at |
|---|---|---|---|---|
convaiinnovations/laya (root) |
ModernBERT-large | 421M | 512 | English text, guardrails, email triage |
laya-multilingual |
mmBERT-base | 322M | 1024 (up to 8,192) | 100+ languages |
laya-typed-decisions |
ModernBERT-large | 421M | 1024 | the four typed-decisions workflows |
- Question types:
choice(pick one option),score(ordinal scale),noul(yes/no probability). You define the options at request time, so a new schema needs no retraining. - Latency on a T4: 39.5 ms (English) and 32.8 ms (multilingual) for one question. With 10 questions batched: 158.6 ms and 72.3 ms.
- Languages: 45 of 51 tested languages are usable (more than 3x random) with routing, against 23 of 51 for the English checkpoint alone.
- Measured results with routing: AG News 0.950, DAIR Emotion 0.595, XNLI English 0.860. On Banking77 it scores only 0.425. See Gotchas.
- Calibration: after fitting one temperature per question type and option count, mean ECE drops from 0.466 to 0.081 on
layaand from 0.314 to 0.106 onlaya-multilingual.
Install
pip install layaYou need Python 3.10 or newer. Optional extras: laya[serve] (HTTP server), laya[mcp], laya[langchain], laya[onnx] and laya[fast] (TileLang GPU fast path).
Run it
from laya import Router
router = Router() # downloads a checkpoint on first use
state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."questions = { "department": {"type": "choice", "instructions": "Which department should handle this?", "criteria": {"billing": "invoices, payments, refunds", "technical": "bugs, outages, system errors", "other": "everything else"}}, "urgency": {"type": "score", "instructions": "How urgent is this?", "criteria": ["not urgent", "soon", "blocking"]}, "churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},}
result = router.predict(state, questions)print(result["answers"]["department"]["choice"]) # billingprint(result["answers"]["urgency"]["score"])print(result["answers"]["churn_risk"]["noul"]) # probability the answer is yesprint(result["routing"]["model"]) # englishprint(result["routing"]) # includes a 'reason' for the routing choiceIf you only need one checkpoint, load it directly:
import laya
agent = laya.load("convaiinnovations/laya", subfolder="multilingual")result = agent.predict( {"body": "La aplicación se cierra cada vez que abro la configuración."}, {"department": {"type": "choice", "instructions": "Which department should handle this?", "criteria": {"billing": "invoices, payments, refunds", "technical": "bugs, outages, system errors", "other": "everything else"}}},)print(result["answers"]["department"]["choice"])To self-host it, laya-serve serves the Router behind the same POST /v1/systemone API that TypeSafe Jev uses:
pip install "laya[serve]"LAYA_API_KEY=change-me LAYA_DEVICE=cuda LAYA_PRELOAD=1 laya-servecurl -s localhost:8000/v1/systemone \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer change-me' \ -d '{ "state": {"document": "I was charged twice. Please fix this ASAP."}, "questions": {"billing": {"type": "noul", "instructions": "Is this ticket about billing?"}}}'Multilingual ticket triage in batch
This is the workflow the card spends the most time on. You send structured email fields, ask several questions in one pass, and let the Router send non-English tickets to the multilingual checkpoint. Router(preload=True) loads all three checkpoints into memory, so switching languages costs only language detection (under 1 ms).
The refund question below is a two-option choice with neutral keys rather than a noul. English tickets go to the English checkpoint, and the card says noul label bias is strongest there (see the guardrail section below).
from laya import Router
questions = { "department": { "type": "choice", "instructions": "Which department should handle this request?", "criteria": { "billing": "invoices, payments, refunds", "technical": "bugs, outages, system errors", "sales": "pricing, new contracts", "other": "everything else", }, }, "urgency": { "type": "score", "instructions": "How urgent is this request?", "criteria": ["not urgent", "soon", "critical deadline or blocking issue"], }, "refund_requested": { "type": "choice", "instructions": "Does the user explicitly request a refund?", "criteria": {"A": "yes, the user requests a refund", "B": "no, the user does not request a refund"}, },}
tickets = [ {"from": "user@acme.com", "subject": "Duplicate charge on invoice #4411", "body": "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."}, {"body": "मुझसे दो बार शुल्क लिया गया, कृपया पैसे वापस करें।"}, {"body": "La aplicación se cierra cada vez que abro la configuración."},]
with Router(preload=True) as router: for ticket in tickets: res = router.predict(ticket, questions) a = res["answers"] print(res["routing"]["model"], a["department"]["choice"], a["urgency"]["score"], a["refund_requested"]["choice"] == "A")If you already run a language-ID model, pass lang_guess="ro" to predict, or install a callable for every request with Router(preload=True, lang_guess=my_lid). The Router checks lang_guess after an explicit lang= and before its built-in detection. If a callable returns None, the Router falls back to detection.
A guardrail check that avoids the noul label bias
The card lists guardrails as a target use for the English checkpoint. It also warns that on this checkpoint, noul questions can follow their internal false:/true: labels instead of the input. When that happens, the model can return a confident “no” for clearly positive text. The card’s workaround is a two-option choice with neutral keys. The message here is English, so the Router sends it to the English checkpoint without an explicit override:
from laya import Router
router = Router()
checks = { "contains_pii": { "type": "choice", "instructions": "Does the message contain personal data such as a phone number, address or card number?", "criteria": {"A": "yes, it contains personal data", "B": "no, it contains no personal data"}, }, "abusive": { "type": "choice", "instructions": "Is the message abusive or harassing?", "criteria": {"A": "yes, it is abusive or harassing", "B": "no, it is not abusive"}, },}
message = "Call me on 555-0199 and stop wasting my time, you idiots."res = router.predict(message, checks)print(res["routing"]["model"]) # expect englishfor name, ans in res["answers"].items(): print(name, "flagged" if ans["choice"] == "A" else "ok")Check the answers against your own labelled data before blocking anything. Shipped checkpoints are over-confident until you fit temperatures (see Gotchas).
Classifying long documents
The multilingual checkpoint ships with a 1,024-token limit but can read up to 8,192 tokens. Set max_len=8192 and name the checkpoint explicitly, because long text that is mostly English would otherwise be routed to the English checkpoint.
from laya import Router
router = Router()
with open("contract.txt", encoding="utf-8") as f: long_document = f.read()
questions = { "doc_type": { "type": "choice", "instructions": "What kind of document is this?", "criteria": {"contract": "agreements, terms, NDAs", "invoice": "bills and payment requests", "other": "everything else"}, },}
result = router.predict(long_document, questions, model="multilingual", max_len=8192)print(result["answers"]["doc_type"]["choice"])Accuracy in the card’s test: 16 to 18 of 20 correct up to about 4,000 tokens, then anywhere from 8 to 17 of 20 beyond that. A 4,000-token input took about 1.7 s on an Apple GPU. Short inputs return the same answers with the higher limit.
Gotchas
- Zero-shot results on hard schemas are weak. On the typed-decisions benchmark, the base English checkpoint scores 0.362. The card gives two figures for multilingual: 0.342 in its measured table and 0.352 in its limits section. Both are below the 0.461 majority-class baseline. The 0.766 score comes from the checkpoint fine-tuned on that benchmark’s training split. The card describes Laya as “a fast base to specialise, not a zero-shot decision engine.” A fine-tuning notebook that runs on Kaggle’s free 2x T4 GPUs is linked from the card.
- Calibrate before trusting probabilities. Shipped checkpoints are over-confident. Fit one temperature per (question type, option count) on your own data. The card reports mean ECE dropping from 0.466 to 0.081 on
layaand from 0.314 to 0.106 onlaya-multilingualafter this. - Many options cause problems. All options share a fixed
head_max_lenbudget: 192 tokens on English, 256 on multilingual. With 77 labels on Banking77, each label gets about 3–4 tokens, and accuracy falls to 0.425. For questions with 50+ options, raiseagent.cfg["head_max_len"] = 512andagent.cfg["max_len"] = 1024or higher, or split the decision into a coarse step and a fine step. - Don’t route on non-Latin text with the English checkpoint. On Khmer it scores 0.000 accuracy at 0.952 confidence, so confidence gating won’t catch the failure. Use the Router or
laya-multilingual. scoreis the weakest question type (SST-5: 0.372).- Ignore
action.act_probability. It reads 1.0 for almost every input. Gate onconfidenceinstead, which reaches an AUROC of 0.77 in the card’s test. - Memory and reloads. The default lazy Router keeps two checkpoints in memory (English and multilingual).
preload=Trueloads all three.max_loaded=1reloads the model on every language switch, which takes 7–10 s. laya.load()can hang if TensorFlow is installed, because TF’s abseil runtime can deadlock model construction. Run withUSE_TF=0.- Server auth.
laya-servebinds to0.0.0.0with no authentication unless you setLAYA_API_KEY.
When to pick it
Pick Laya for routing, triage, classification, scoring and moderation decisions over a fixed set of options. It fits when you need low latency, structured output and open weights you can run on-premise under Apache 2.0. It is especially useful for multilingual traffic.
The card’s TypeSafe Jev figures come from third parties and weren’t measured in the same run. Sample sizes and prompts differ. Against those numbers, Laya scores higher on AG News (0.950 vs 0.910) and DAIR Emotion (0.595 vs 0.480), and the card puts it at roughly 6–8x faster for a single question. Jev leads in other places: Banking77 (0.870 vs 0.425), soft accuracy on typed-decisions (0.580 vs 0.471) and raw calibration before temperature scaling.
Skip it, or plan to fine-tune first, if:
- you have a complex domain schema and no labelled data;
- you need 50+ options in one question without tuning, since Jev scores 0.870 vs Laya’s 0.425 on Banking77;
- you mainly need fine-grained ordinal scores;
- your task requires generating text or extracting spans.
If you have labelled decisions from your own domain, the card’s fine-tuning results point to fine-tuning as the way to get high accuracy.
Related
Get calibrated yes/no and multiple-choice answers with GEV-26B-Decide
A LoRA adapter plus a small decision head on Gemma-4-26B-A4B that gives a calibrated probability for every option in about 45 ms, and can think when it is unsure.
autotrust/GEV-26B-Decide
Route Tickets and Intents Locally with GLiNER2.5-Decide
A 340M classifier from Fastino that scores intent, urgency, sentiment and routing labels you choose at call time, in one forward pass, on CPU or GPU.
fastino/GLiNER2.5-Decide
Clef: Cloudflare's model that returns probabilities, not text
Clef answers typed questions about any input (text, JSON, images) with a probability per option in one forward pass. Here's how to use it for routing, triage and classification.
Cloudflare/clef