Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 11 min read

Run JEV-27B-VL for calibrated yes/no and choice decisions on images

A 27B vision decision model that returns a calibrated probability for every option in one forward pass. Here's how to serve it and use it.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

JEV-27B-VL from AutoTrust is Qwen3.8-27B with a “System 1” LoRA adapter and decision head added. You can ask it three kinds of question over text and images: yes/no, pick one of 2 to 256 options, or a 0–5 rating. It returns a probability for every option in a single forward pass. The card calls these probabilities calibrated. On the vision board it reports an ECE of 0.019, but the authors also say they haven’t systematically measured calibration on image tasks. The notable part is that the decision head was trained on text only, and the vision encoder is Qwen3.8-27B’s, unchanged. Even so, it ranks first of 20 models on the community Jev Decision Index 0.3 Vision board, where every image decision is zero-shot. The same weights also work as a normal vision chat model (“System 2”). That means a low-confidence decision can be passed to the chat model on the same server.

Key specs

Base Qwen/Qwen3.8-27B (weights unchanged) + JEV System 1 LoRA adapter and decision head
Decision types noul (yes/no), choice (2–256 options), score (0–5)
Context 262,144 tokens native; 20/20 correct on needle-style decisions up to 250K tokens
Weights / KV cache about 52 GB of weights; about 65 KB of KV cache per token
Decision Index Vision board Full 69.82, public 72.78, private 66.86, ECE 0.019 (rank 1 of 20)
Image calibration ECE 0.019 on the board; the authors say image-task calibration hasn’t been measured systematically
Reference on the same board Qwen3.5-397B-A17B stock VLM: Full 63.44
Accuracy on the board 77.8% over 13,695 rows, mean confidence 79.2%
Latency about 264 ms per row on an RTX PRO 6000
VL-RewardBench (multimodal judge) 78.3% overall
AgentRewardBench (agent judge) AUROC 0.91; precision 78.4 / recall 70.2 at threshold 0.5
System 2 HumanEval pass@1 78.0%, identical to Qwen3.8-27B

Install

The repository ships its own server script. serve.sh runs serve_decide.py, which is the standard vLLM OpenAI server plus a POST /v1/decide route for System 1.

Terminal window
hf download autotrust/JEV-27B-VL --local-dir JEV-27B-VL
bash JEV-27B-VL/serve.sh # vLLM on :8000; one GPU with 80 GB or more

If you want to set the flags yourself, this is the full command from the card:

Terminal window
python3 JEV-27B-VL/serve_decide.py --model JEV-27B-VL --served-model-name autotrust/JEV-27B-VL \
--enable-lora --max-lora-rank 32 --lora-modules jev-decision=JEV-27B-VL/adapter_vllm \
--logprobs-mode processed_logprobs --max-model-len 32768 --enable-prefix-caching --mamba-cache-mode align \
--limit-mm-per-prompt '{"image": 8}' --max-num-seqs 8 --trust-request-chat-template

For the full 256K context, set MAX_MODEL_LEN=262144 before you run serve.sh.

Run it

You can call the decision endpoint from any HTTP client:

Terminal window
curl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
"kind": "choice",
"state": "Customer: my card was charged twice for one coffee.",
"question": "Which team should handle this?",
"options": ["billing", "shipping", "tech support"]}'

The response includes options, probabilities, choice_index and choice. It also has an adaptation field, which is native for up to 16 options and wide-labels beyond that.

For images, make state a list that mixes strings and image objects. Each image goes where it appears in the list, and it can be a data URL or an https:// URL. The examples below share this small client, which is adapted from the card. It needs requests. The card’s version hard-codes image/jpeg. This one guesses the MIME type from the file extension, so PNG screenshots get the right data URL. It also checks the HTTP status, so a server error shows its real message instead of a KeyError.

jev.py
import base64, mimetypes, requests
URL = "http://localhost:8000"
def image(path):
mime = mimetypes.guess_type(path)[0] or "image/jpeg"
return {"image": f"data:{mime};base64," + base64.b64encode(open(path, "rb").read()).decode()}
def decide(kind, state, question, options=None):
body = {"kind": kind, "state": state, "question": question, **({"options": options} if options else {})}
resp = requests.post(f"{URL}/v1/decide", json=body)
resp.raise_for_status()
r = resp.json()
return dict(zip(r["options"], r["probabilities"]))
if __name__ == "__main__":
print(decide("noul", ["Photo: ", image("photo.jpg")],
"Is this scenario one where: the image shows food or cooking?"))
# e.g. {'false': 0.00, 'true': 1.00}

Classify marketplace listings from photo and title

Put all the categories in a single choice question rather than asking one yes/no question per category. On the card’s text intent-classification sets (MASSIVE, BANKING77, CLINC150), one choice question was more accurate than one yes/no question per option, and it costs one forward pass instead of one per option. The card tested several prompt changes on those same text sets, and only one clearly helped: giving similar options a one-line “use when” description that says what separates each option from its neighbours. None of this was measured on image tasks, so treat it as a starting point for a photo-plus-title task like this one. The descriptions below draw the boundary between electronics and household appliances explicitly, so a powered kitchen appliance fits only one category. In this example, the probability decides whether a listing is filed automatically or sent to a person.

listing_triage.py
import sys
from jev import decide, image
CATEGORIES = [
"electronics: use when the item is a personal device or gadget such as a phone, computer, audio gear, camera, cable or charger; not household or kitchen appliances.",
"clothing: use when the item is worn, including shoes and bags.",
"home and kitchen: use when the item is furniture, cookware, decor or a household appliance, including powered kitchen appliances.",
"toys: use when the item is made for children's play, including games, puzzles and electronic toys.",
]
THRESHOLD = 0.90 # validate this on your own labelled listings
def triage(photo_path, seller_title):
probs = decide(
"choice",
["Listing photo: ", image(photo_path), f"\nSeller title: {seller_title}"],
"Which category fits this listing?",
CATEGORIES,
)
best, p = max(probs.items(), key=lambda kv: kv[1])
label = best.split(":")[0]
return (label, p) if p >= THRESHOLD else ("needs_review", p)
if __name__ == "__main__":
print(triage(sys.argv[1], sys.argv[2]))

Keep newlines out of option text, because the prompt template puts one option on each line.

Judge whether a web agent finished its task

For its AgentRewardBench result, the card lists the inputs the model was given: the user’s goal, the agent’s actions and final message, and the final screenshot. From those it returns P(task completed). The card text does not give the prompt or question wording, so the wording below is ours. Don’t expect this exact code to reproduce the card’s numbers. If you want to reproduce them, the card links the authors’ code: JEV-27B-DEMO / 10-agent-judge. When you write the noul question, phrase it so that “true” is the outcome you want the probability of.

agent_judge.py
import json, sys
from jev import decide, image
def task_completed(goal, actions, final_message, screenshot_path):
probs = decide(
"noul",
[
f"User goal: {goal}\n",
"Agent actions:\n" + "\n".join(actions) + "\n",
f"Agent final message: {final_message}\n",
"Final screenshot: ", image(screenshot_path),
],
"Did the agent complete the user's goal?",
)
return probs["true"]
if __name__ == "__main__":
# trajectory.json: {"goal": ..., "actions": [...], "final_message": ..., "screenshot": "final.png"}
t = json.load(open(sys.argv[1]))
p = task_completed(t["goal"], t["actions"], t["final_message"], t["screenshot"])
print(f"P(completed) = {p:.3f}", "PASS" if p >= 0.5 else "FAIL")

At a threshold of 0.5, the card reports precision 78.4 and recall 70.2 on AgentRewardBench with its own prompt. Measure your prompt on labelled trajectories of your own before you trust a threshold. If false passes cost you more than false failures, raise the threshold.

Compare two answers about an image, escalating to System 2 when unsure

For pairwise judging, the card advises asking in both orders and averaging. About 7% of 16-option answers change with option order alone. If the averaged confidence is below 0.70, this script sends the question to System 2 on the same server. The prompt wording here is ours; the card links the code behind its VL-RewardBench result at JEV-27B-DEMO / 09-multimodal-judge.

The card’s Applied tasks table reports a result for this escalation policy, but it doesn’t name the task or dataset. The card also says its text results were measured with JEV-27B. Reported numbers: accuracy 0.792 with System 1 alone, 0.892 with escalation (70% of items answered by System 1), and 0.917 with thinking on everything. The card doesn’t say whether escalated items used thinking, though the “thinking on everything” baseline suggests they did. This script turns thinking off for System 2, so it differs from the reported setup in that way too. Escalation has also not been measured on image or pairwise tasks, so treat 0.70 as a starting point and check it on your own data.

The System 2 call matches the card’s example, which turns thinking off. The script then looks for a standalone “A” or “B” on the last line of the reply, so endings like **B** or (A). still parse.

pair_judge.py
import re, sys, requests
from jev import URL, decide, image
def judge(image_path, question, answer_a, answer_b):
def ask(first, second):
return decide(
"choice",
["Image: ", image(image_path),
f"\nQuestion: {question}\nAnswer 1: {first}\nAnswer 2: {second}"],
"Which answer is more accurate and helpful for this image?",
["Answer 1", "Answer 2"],
)
fwd = ask(answer_a, answer_b)
rev = ask(answer_b, answer_a)
p_a = (fwd["Answer 1"] + rev["Answer 2"]) / 2
return p_a
def system2(image_path, question, answer_a, answer_b):
img = {"type": "image_url", "image_url": {"url": image(image_path)["image"]}}
prompt = (f"Question: {question}\nAnswer A: {answer_a}\nAnswer B: {answer_b}\n"
"Which answer is better for this image? End with 'A' or 'B'.")
resp = requests.post(f"{URL}/v1/chat/completions", json={
"model": "autotrust/JEV-27B-VL",
"messages": [{"role": "user", "content": [img, {"type": "text", "text": prompt}]}],
"max_tokens": 1024, "chat_template_kwargs": {"enable_thinking": False}})
resp.raise_for_status()
text = resp.json()["choices"][0]["message"]["content"] or ""
lines = [l for l in text.strip().splitlines() if l.strip()]
found = re.findall(r"\b([AB])\b", lines[-1]) if lines else []
verdict = found[-1] if found else None
return verdict, text
if __name__ == "__main__":
img, q, a, b = sys.argv[1:5]
p_a = judge(img, q, a, b)
if max(p_a, 1 - p_a) >= 0.70:
print("A" if p_a >= 0.5 else "B", f"(System 1, P(A)={p_a:.3f})")
else:
verdict, text = system2(img, q, a, b)
print(verdict, "(System 2)" if verdict else f"(System 2, no clear verdict)\n{text}")

System 1 calls go through the jev-decision LoRA. System 2 calls use the served model name autotrust/JEV-27B-VL, which is the unmodified Qwen3.8-27B. The card says System 2 can optionally think step by step, but its example turns thinking off. If you turn it on, give it a larger max_tokens budget and strip the reasoning before you parse the answer.

Gotchas

  • --max-num-seqs 8 is required. If a batch has more than 8 sequences, vLLM’s LoRA path for this multimodal model class returns wrong System 1 probabilities. With the cap in place, extra requests queue and the results stay correct. The text-only JEV-27B does not have this problem.
  • You need a specific vLLM build. serve_decide.py was tested with a vLLM development build from September 2026 and uses logprob_token_ids. Stock vLLM may not run it.
  • Pass top_k: 0 and top_p: 1.0 if you skip /v1/decide. The model’s generation_config.json sets top_k=20 and top_p=0.95, and vLLM applies these as request defaults. With --logprobs-mode processed_logprobs, those defaults truncate the returned probabilities and zero out the less likely options. serve_decide.py handles this for you.
  • Keep --trust-request-chat-template. /v1/decide needs it to render the raw decision template for image decisions.
  • Memory. The weights are 52 GB, and a full 256K prompt adds about 17 GB of KV cache. On an 80 GB GPU, start with --max-model-len 131072.
  • Image size. The Qwen3.8 processor resizes images itself. Downscaling large images first, for example to at most 448 px, reduces the number of vision tokens and the latency.
  • Image calibration is unmeasured. The decision head was trained on text. The board’s ECE of 0.019 comes from an independent evaluation, but the authors say they have not systematically measured calibration on image tasks. Validate your thresholds on your own data.
  • More than 16 options. Beyond 16 options the model uses labels the head never saw in training. Accuracy holds up, but calibration has been checked on only one dataset.
  • Control loops. Ask simple visual questions, such as whether the target is left or right of the gripper. When asked to pick one of 8 motor commands directly, it completed 0 of 10 robot-arm scenes. For computer use, include the element text. With numbered boxes alone, task completion fell from 95% to 10%, often because the model declared the task complete too early.
  • License. Apache-2.0. That covers both the bundled Qwen3.8-27B weights and the adapter.

When to pick it

Pick JEV-27B-VL when you need a probability, not a paragraph. Good fits include routing, moderation, sorting images or documents into categories, judging agent trajectories or answers, and cold-start recommendation from thumbnails.

The card’s clearest size comparison is the Vision board. There it scores 6.4 Full-score points above the stock Qwen3.5-397B-A17B with about 15× fewer total parameters. On VL-RewardBench it reaches 78.3% overall, above every model on that benchmark’s 2025 leaderboard. The runner-up is Skywork-VL-Reward-7B at 73.3%, which beats JEV-27B-VL on the “general” category (65.6 vs 58.0). On MicroLens recommendation it scores AUC 0.727 from covers alone. That matches collaborative filtering trained on 59,045 users’ logs (0.728; the card puts the difference at 0.000, 95% interval −0.041 to +0.040). Note that the MicroLens result is one offline evaluation on one dataset with 200 users, ranking from covers only.

Because the output is a probability, you can set a threshold, act automatically above it, and send the rest to System 2 on the same server or to a person.

Skip it in these cases:

  • You don’t have an 80 GB GPU. The card says to use one GPU with 80 GB or more. It also lists JEV-9B, which matched this model on computer use with element text (95%), did better with numbered boxes only (37% vs 10%), and completed fewer robot-arm scenes (50% vs 75%). The card doesn’t give JEV-9B’s hardware requirements.
  • You can’t run the specific vLLM build that serve_decide.py needs.
  • You need high batch throughput on one server. The 8-sequence cap limits concurrency.
  • Your task is image editing evaluation. It is the model’s weakest MMRB2 task (57.8).
  • You need the best judge regardless of size. On MMRB2 its 63.8 average ranks 6th of the 9 judges in the card’s table. GPT-5, Gemini 3 Pro, Gemini 2.5 Pro, Qwen3-VL-32B and Gemini 2.5 Flash all score higher.
  • Your data is text only. The text-only JEV-27B gives nearly identical text decisions (mean probability difference 0.010 and 99.8% on the same side of 0.5 over 1,000 answer-checking decisions), and the card says it is not affected by the LoRA bug that hits batches of more than 8 sequences.

Related