Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 9 min read

Get calibrated yes/no and multiple-choice answers with GEV-26B-Decide

A LoRA adapter plus a small decision head on Gemma-4-26B-A4B that gives a calibrated probability for every option in about 45 ms, and can think when it is unsure.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

GEV-26B-Decide is an open-weights decision model from AutoTrust AI. It is a LoRA adapter and a small decision head on top of google/gemma-4-26B-A4B-it, a mixture-of-experts model with 26B parameters and about 4B active per token. You give it a typed question: yes/no, pick one of 2–256 options, or rate 0–5. It does not generate an answer. It returns a calibrated probability for every option in one forward pass, which takes about 45 ms on a B200. The notable part is opt-in “adaptive thinking”. When the fast pass (System 1) is less than 0.8 confident, the unmodified Gemma-4 backbone (System 2) reasons about the question, and its answer is averaged into the probabilities. The card reports that this keeps System 1’s calibration. Text and images both go through one set of weights and one vLLM engine. The model was previously published as autotrust/JEV-Gemma4-26B-A4B.

Key specs

Base google/gemma-4-26B-A4B-it (26B total, ≈4B active)
Decision kinds noul (yes/no), choice (2–256 options), score (0–5)
Context 262,144 tokens native; tested up to 128K
System 1 latency median 45 ms, 257 decisions/s with 64 concurrent clients (one B200)
Thinking speed ≈250 tokens/s, ≈438 tokens/s with speculative decoding
Calibration ECE 0.035 for System 1 and for adaptive mode (on the 1,754-question selection set)
Decision Index 0.2.1 62.48 (the authors’ own scoring, not a board entry)

The authors picked the threshold and the mix on 1,754 questions from six public sets that are not part of the Decision Index. On that same selection set, adaptive mode raised accuracy from 73.3% to 83.4%. It thought on 47.8% of the questions, and always thinking scored 83.8%. On the Decision Index’s Knowledge & Reasoning benchmarks, the biggest gains were GPQA Diamond (42.9 to 78.6), CRUXEval (67.5 to 90.7), MMLU-Pro (65.0 to 84.6) and BBH (75.0 to 92.0), with CLadder (71.0 to 86.6) close behind. Intent classification with every label offered at once scored 95.5% on CLINC150 (150 options) and 91.2% on MASSIVE (59 options). The CLINC150 training split is part of the model’s training data, so read the CLINC150 number with that in mind.

Install

The supported setup is the bundled vLLM server. It needs a vLLM build with Gemma-4 support plus patches/vllm-gemma4-lm-head-lora.patch. The authors tested it with a vLLM development build from September 2026.

Terminal window
hf download autotrust/GEV-26B-Decide --local-dir GEV-26B-Decide
bash GEV-26B-Decide/serve.sh # vLLM on :8000; one GPU with 80 GB or more

serve.sh starts the standard vLLM OpenAI server with an added POST /v1/decide route. Plain chat requests go to System 2, and decision requests use the jev-decision LoRA for System 1. To speed up thinking with speculative decoding, start it with MTP=1 bash GEV-26B-Decide/serve.sh.

Run it

Terminal window
curl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
"kind": "choice",
"state": "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball.",
"question": "How much does the ball cost?",
"options": ["$0.10", "$0.05", "$1.00", "$0.55"],
"thinking": "auto"}'

The response contains options, probabilities, choice, choice_index and usage. When you request thinking, it also contains a thinking object with used, think_tokens, think_seconds and finished_within_budget. GET /v1/decide/info lists the server defaults.

If you don’t want the server, the card also shows a transformers + peft path for System 1. It is shown in the confidence-gating section below.

Intent routing with many labels

This is the most direct production use: route support messages to one of many intents with no training. Keep thinking set to "off" here. The card shows thinking adds little to classification and lowered BANKING77 macro-F1 from 88.0 to 85.0. With more than 16 options, the server reads the options in groups of 16 and then runs a final round (strategy: "tournament", the default).

route_intents.py
import requests
from concurrent.futures import ThreadPoolExecutor
URL = "http://localhost:8000/v1/decide"
INTENTS = [
"card_lost_or_stolen", "refund_request", "change_delivery_address",
"cancel_order", "account_locked", "update_payment_method",
"track_parcel", "speak_to_human", "other",
]
def route(message):
body = {
"kind": "choice",
"state": message,
"question": "Which intent best describes this customer message?",
"options": INTENTS,
"thinking": "off",
}
r = requests.post(URL, json=body).json()
return message, r["choice"], max(r["probabilities"])
messages = [
"I can't log in, it says too many attempts.",
"Where is my package? It was due yesterday.",
"Please send the order to my office instead.",
"Someone took my wallet with my card in it.",
]
# System 1 is built for concurrency: 257 decisions/s with 64 clients on one B200
with ThreadPoolExecutor(max_workers=16) as pool:
for msg, intent, p in pool.map(route, messages):
print(f"{p:.2f} {intent:28s} {msg}")

Because the probabilities are calibrated, you can send low-confidence messages to a human queue rather than accept every top choice.

Reasoning questions with adaptive thinking

For questions where the answer can be checked step by step (math, logic, code, constraints), set "thinking": "auto". Thinking applies to noul and choice only. The card does not list it for score, so keep score requests at "thinking": "off". Easy questions still return in one pass. Hard ones can take many seconds. The card estimates, from token counts rather than measurements, a median of 13.4 s and a 90th percentile of 33 s over the Knowledge & Reasoning benchmarks. Those figures assume a single request on one idle B200 without speculative decoding (thinking at about 250 tokens/s). With speculative decoding the estimates fall to a 7.7 s median and a 19 s 90th percentile. You can cap the cost with think_budget and threshold.

reasoning.py
import requests
def decide(kind, state, question, options=None, thinking="auto", **extra):
# The card lists thinking for "noul" and "choice" only, not "score"
if kind == "score":
thinking = "off"
body = {"kind": kind, "state": state, "question": question,
"thinking": thinking, **extra}
if options:
body["options"] = options
r = requests.post("http://localhost:8000/v1/decide", json=body).json()
return dict(zip(r["options"], r["probabilities"])), r.get("thinking", {})
probs, info = decide(
"noul",
"John was born on 29 February 1996.",
"Was John's 7th birthday celebrated on a 29 February?",
)
print(probs, info)
probs, info = decide(
"choice",
"A train leaves at 09:40 and the trip takes 2 h 35 min.",
"When does it arrive?",
["11:75", "12:15", "12:05", "11:15"],
think_budget=4096, # cap thinking tokens
threshold=0.8, # think when System 1's top option is below this
)
print(probs, info.get("used"), info.get("think_seconds"))

The server also accepts return_reasoning to include System 2’s reasoning in the response. The card does not name the response field that holds it, so check a raw response before relying on it.

To use System 2 as an ordinary chat model, call the standard endpoint:

chat.py
import requests
r = requests.post("http://localhost:8000/v1/chat/completions", json={
"model": "autotrust/GEV-26B-Decide",
"messages": [{"role": "user", "content": "In one sentence, what is safety stock?"}],
"max_tokens": 200, "chat_template_kwargs": {"enable_thinking": False}})
print(r.json()["choices"][0]["message"]["content"])

Image checks and confidence gating

state can be a list that mixes text and images. The card reports 100% on simple synthetic checks such as colour, shape, printed digits and “is there a red object?”, and 78.4% on VL-RewardBench. Image decisions are zero-shot, because the head was trained on text.

image_check.py
import requests
body = {
"kind": "noul",
"state": ["Photo of the returned item: ", {"image": "https://example.com/return_1234.jpg"}],
"question": "Is the item visibly damaged?",
"thinking": "off",
}
r = requests.post("http://localhost:8000/v1/decide", json=body).json()
print(dict(zip(r["options"], r["probabilities"]))) # probabilities for ["false", "true"]

If you gate automatic actions on confidence, the card recommends the calibration_gold.json temperatures, which were calibrated against ground-truth answers. The transformers path lets you pick the file. The code below is the card’s code with the calibration filename changed. The card only shows the ["per_kind"] structure for calibration.json. We assume calibration_gold.json uses the same keys, but we have not checked, so open the file before you rely on this. The card also doesn’t say which transformers version provides Gemma4ForConditionalGeneration or how much GPU memory this path needs.

gated_decide.py
import json, torch
from huggingface_hub import snapshot_download
from peft import PeftModel
from safetensors.torch import load_file
from transformers import AutoTokenizer, Gemma4ForConditionalGeneration
d = snapshot_download("autotrust/GEV-26B-Decide")
tok = AutoTokenizer.from_pretrained(d)
base = Gemma4ForConditionalGeneration.from_pretrained(d, dtype=torch.bfloat16, device_map="cuda")
m = PeftModel.from_pretrained(base, f"{d}/adapter").merge_and_unload().eval()
backbone = m.model
jc = json.load(open(f"{d}/judge_config.json"))
# Assumption: calibration_gold.json has the same "per_kind" layout as calibration.json
T = json.load(open(f"{d}/calibration_gold.json"))["per_kind"] # ground-truth calibrated
head = load_file(f"{d}/head.safetensors"); W, b = head["proj.weight"].cuda(), head["proj.bias"].cuda()
@torch.no_grad()
def decide(kind, state, question, options):
lines = options if kind != "choice" else [f"{'ABCDEFGHIJKLMNOP'[i]}) {o}" for i, o in enumerate(options)]
text = f"[kind] {kind}\n[state] {state}\n[question] {question}\n[options]\n" + "\n".join(lines) + "\n[decision]:"
ids = torch.tensor([[tok.bos_token_id] + tok.encode(text, add_special_tokens=False)], device="cuda")
h = backbone(input_ids=ids, use_cache=False).last_hidden_state[0, -1].float()
z = 30.0 * torch.tanh((W @ h + b) / 30.0)
s, _ = jc["slots"]["ranges"][kind]
return dict(zip(options, torch.softmax(z[s:s + len(options)] / T[kind], 0).tolist()))
p = decide("noul", "Customer says the parcel arrived damaged and wants their money back.",
"Is the customer asking for a refund?", ["false", "true"])
if p["true"] >= 0.95:
print("auto-approve refund flow", p)
else:
print("send to agent", p)

The 0.95 threshold is an example. Choose yours on your own labelled data.

Gotchas

  • Patched vLLM required. The one-engine server needs Gemma-4 support plus the bundled lm_head LoRA patch. It was tested on a September 2026 vLLM development build.
  • Memory. serve.sh expects one GPU with 80 GB or more. The card gives no memory figure or transformers version for the transformers path.
  • Sampling defaults truncate probabilities. If you call /v1/completions for System 1 yourself, pass top_k: 0 and top_p: 1.0. Otherwise vLLM applies the generation config’s top_k=64 and top_p=0.95 as request defaults.
  • Thinking is noul and choice only. The card does not list thinking for score.
  • Thinking is not always better. It lowered BANKING77 macro-F1 (88.0 → 85.0) and CommonsenseQA accuracy (86.7 → 85.0). It did not help ChessBench, where two thirds of the thoughts hit the 8,192-token budget, or HLE, where System 1 is below chance.
  • Speculative decoding trades throughput. With MTP=1, thinking is about 1.8–1.9× faster, but System 1 throughput at 64 clients drops from 257 to 140 decisions per second.
  • The transformers path reads at most 16 options. For more, use the server.
  • noul and score accept only their canonical options.
  • Benchmark caveats. System 1 was trained on the BANKING77 and CLINC150 training splits, so the intent scores on those sets are not truly zero-shot. The 73.3% → 83.4% adaptive result and the 0.035 ECE come from the same 1,754 questions used to pick the threshold and mix. The Decision Index 62.48 is the authors’ own scoring, not a board result.
  • Untested ranges. Contexts above 128K tokens are untested, and decisions over images are zero-shot.
  • License. The adapter, head and calibration files are Apache-2.0, but the base weights in the repo fall under the Gemma 4 terms.

When to pick it

Pick it if you need fast classification, routing or yes/no checks with probabilities you can threshold. It fits especially well when you have many labels and no training data, or decisions inside a loop. The card reports 85 ms per click for computer use, which needs element text to work well, and 61 ms per robot-arm step. It could also serve for LLM-as-judge-style scoring, but the card tests that only on images: 78.4% on VL-RewardBench, a pairwise judge benchmark, with zero-shot image decisions. It reports no text judge benchmark. Try adaptive thinking for logic, math, science and code questions. There it gave large gains on GPQA Diamond, CRUXEval, CLadder, MMLU-Pro and BBH, and it runs only when System 1 is unsure. It added little on MuSR (67.3 → 67.6) and SATA-Bench (34.2 → 35.5), nothing on ChessBench, and only brought HLE up to chance.

Skip it if:

  • You can’t run a patched vLLM on an 80 GB-class GPU.
  • You need non-English decisions. The card calls the model English-centric.
  • You need free-form generation. System 2 is just the base Gemma-4 model, so use that directly.
  • You need fine-grained visual control. JEV-27B-VL completed 75% of the robot-arm scenes against this model’s 40%.
  • You need long structured inputs. The card says it is weaker there than JEV-27B.
  • The decisions are high-stakes and you have no confidence gating and human review in place.

Related