Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 7 min read

Run Qwen3.8-27B for coding agents, documents and video

Qwen3.8-27B is an Apache-2.0 dense 27B vision-language model with adjustable reasoning. How to serve it and call it today.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Qwen3.8-27B is the Qwen team’s new dense 27B model. It reads images and video natively, and its thinking mode can be switched on or off for each request. It uses the Qwen3.5 architecture and is released under Apache 2.0. Its main selling point is agentic work: on the team’s own tables, this 27B model scores above Qwen3.7-Plus and the “Opus4.6 Max” baseline on several computer-use and agentic-coding benchmarks. The usual caveat applies, though: some of those benchmarks are in-house.

Key specs

Parameters 27B, dense
Layers 64, in a hybrid layout: 16 × (3 × Gated DeltaNet + 1 × Gated Attention), each followed by an FFN
Context 262,144 tokens natively, extendable to 1,000,000 with YaRN
Modalities Text, image, video input
Reasoning control enable_thinking, reasoning_effort (xhigh / medium / low), preserve_thinking
Extras Multi-token prediction (MTP) trained with multiple steps

These are the benchmarks most useful for deciding whether to try it (numbers from the card):

Benchmark Qwen3.8-27B Qwen3.6-27B Qwen3.7-Plus Opus4.6 Max
SWE-bench Pro 61.7 53.5 57.6 53.4*
Terminal Bench 2.1 73.0 63.4 64.0 78.2
OSWorld-Verified 84.3 63.9 73.3 72.7
AndroidWorld 81.9 70.3 81.0 62.0
OmniDocBench 1.5 91.1 89.4 91.4 86.6
GPQA Diamond 89.2 87.8 90.3 91.3
HLE 30.8 24.0 34.7 40.0

*The card notes that the Opus4.6 Max SWE-bench Pro score is the officially reported number. The other models were re-evaluated with the Claude Code harness on a corrected version of the benchmark.

Overall, it is strong on computer use and agentic coding. On knowledge-heavy reasoning (HLE, GPQA Diamond) it still trails Qwen3.7-Plus and Opus4.6 Max.

Install

The card recommends running the model behind an inference server and calling it through an OpenAI-compatible API. To serve it yourself, use one of the framework guides the card links:

Hosted access will be available through Qwen Cloud, which the card describes as “coming soon”.

Install the client and point it at your endpoint:

Terminal window
pip install -U openai
export OPENAI_BASE_URL='your-base-url'
export OPENAI_API_KEY='your-api-key'

Run it

Thinking is on by default. With streaming, the reasoning arrives in reasoning_content (or reasoning, depending on the server) and the answer arrives in content:

hello.py
from openai import OpenAI
client = OpenAI()
messages = [{"role": "user", "content": "Write a Python function to merge two sorted linked lists."}]
completion = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=messages,
reasoning_effort="xhigh",
stream=True,
stream_options={"include_usage": True},
)
reasoning, answer = "", ""
for chunk in completion:
if not chunk.choices:
print("\nUsage:", chunk.usage)
continue
delta = chunk.choices[0].delta
if getattr(delta, "reasoning_content", None) is not None:
reasoning += delta.reasoning_content
elif getattr(delta, "reasoning", None) is not None:
reasoning += delta.reasoning
if getattr(delta, "content", None):
print(delta.content, end="", flush=True)
answer += delta.content

Batch-read images in non-thinking mode

For extraction, captioning or quick visual Q&A you may not need a reasoning trace. Turn thinking off and use the instruct sampling settings from the card (temperature=0.7, top_p=0.8, top_k=20, presence_penalty=1.5). The example uses the card’s demo images. Replace them with your own scans or screenshots.

batch_images.py
from openai import OpenAI
client = OpenAI()
jobs = [
("https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/RealWorld/RealWorld-04.png",
"Where is this?"),
("https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/CI_Demo/mathv-1327.jpg",
"Describe the geometric figure in this image in one paragraph."),
]
for url, question in jobs:
response = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=[{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": url}},
{"type": "text", "text": question},
],
}],
temperature=0.7,
top_p=0.8,
presence_penalty=1.5,
extra_body={
"top_k": 20,
"chat_template_kwargs": {"enable_thinking": False},
},
)
print(url, "->", response.choices[0].message.content, "\n")

The card reports 91.1 on OmniDocBench 1.5 and 83.7 on CharXiv (RQ) in its “Without CI” setting. It doesn’t say whether thinking was on or off for those runs, so check non-thinking quality on your own documents before relying on it.

Ask questions about a video

Video goes in as a video_url content part. On vLLM you can also control frame sampling per request with mm_processor_kwargs. According to the card, this only works if the server was launched with --media-io-kwargs '{"video": {"num_frames": -1}}'. The defaults are fps=2 and do_sample_frames=True, so only pass it when you want a different fps. It is commented out below, as it is on the card.

video_qa.py
from openai import OpenAI
client = OpenAI()
messages = [{
"role": "user",
"content": [
{"type": "video_url",
"video_url": {"url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/video/N1cdUjctpG8.mp4"}},
{"type": "text",
"text": "How many porcelain jars were discovered in the niches located in the primary chamber of the tomb?"},
],
}]
response = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=messages,
# vLLM only, and only when launched with --media-io-kwargs '{"video": {"num_frames": -1}}'.
# Uncomment and change fps to set your own sampling rate (defaults: fps=2, do_sample_frames=True):
# extra_body={"mm_processor_kwargs": {"fps": 2, "do_sample_frames": True}},
)
print(response.choices[0].message.content)

For hour-long videos, the card recommends raising longest_edge in video_preprocessor_config.json (see Gotchas).

Multi-turn coding with tunable reasoning

reasoning_effort trades depth for speed. preserve_thinking keeps earlier reasoning in the context, which the card says helps agents stay consistent and improves KV-cache reuse. To carry a conversation forward, append the assistant turn together with its reasoning, as the card does:

coding_session.py
from openai import OpenAI
client = OpenAI()
def ask(messages, effort="medium"):
stream = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=messages,
reasoning_effort=effort,
temperature=1.0,
top_p=0.95,
extra_body={
"top_k": 20,
"chat_template_kwargs": {"enable_thinking": True, "preserve_thinking": True},
},
stream=True,
)
reasoning, answer = "", ""
for chunk in stream:
if not chunk.choices:
continue
d = chunk.choices[0].delta
if getattr(d, "reasoning_content", None) is not None:
reasoning += d.reasoning_content
elif getattr(d, "reasoning", None) is not None:
reasoning += d.reasoning
if getattr(d, "content", None):
answer += d.content
messages.append({
"role": "assistant",
"content": answer,
"reasoning_content": reasoning,
"reasoning": reasoning,
})
return answer
messages = [{"role": "user", "content": "Write a Python LRU cache class with get/put in O(1)."}]
print(ask(messages))
messages.append({"role": "user", "content": "Now make it thread-safe and add unit tests."})
print(ask(messages, effort="xhigh"))

Gotchas

  • Thinking is on by default. Every response starts with a <think> block unless you set enable_thinking: False. On Qwen Cloud, put "enable_thinking": False and "preserve_thinking": False directly in extra_body, not inside chat_template_kwargs.
  • There are two sampling profiles. Thinking mode uses temperature=1.0, top_p=0.95, presence_penalty=0.0. Non-thinking mode uses temperature=0.7, top_p=0.80, presence_penalty=1.5. Raising presence_penalty (range 0–2) reduces repetition, but it can cause language mixing and a small quality drop. Not every framework supports every parameter.
  • Lower effort isn’t always faster. The card warns that in multi-turn agent tasks, low effort can lead to more failures and retries, which raises total latency and token use.
  • Agentic tasks need large output budgets. On frameworks that support separate limits, the card suggests up to 262,144 tokens for reasoning and 131,072 for the final response. It frames these as “within the 1M context length”. Together they come to about 393K tokens, more than the native 262,144 context, so these budgets assume you have enabled the YaRN-extended context.
  • YaRN is static. To go past 262K tokens, either edit rope_parameters under text_config in config.json, or pass overrides at launch. The card’s commands need three things: an environment variable that allows the longer length (VLLM_ALLOW_LONG_MAX_MODEL_LEN=1, SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 or TOKENSPEED_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1), the rope override (--hf-overrides for vLLM and TokenSpeed, --json-model-override-args for SGLang), and the context flag (--max-model-len 1000000 for vLLM and TokenSpeed, --context-length 1000000 for SGLang). Copy the full commands from the card’s “Processing Ultra-Long Texts” section. The open frameworks apply a constant scaling factor, which can hurt short inputs. Only enable it when you need it, and size factor to your typical length (for example, 2.0 for about 524K tokens).
  • The video config is conservative by default. For hour-scale video, set {"longest_edge": 469762048, "shortest_edge": 4096} in video_preprocessor_config.json, which corresponds to about 224K video tokens.
  • There’s no Transformers code on the card. The weights are in Transformers format, but every example on the card goes through an API server. The 1M default context and built-in tools mentioned on the card belong to the hosted Qwen Cloud version, not the open weights.
  • Several headline numbers are in-house. QwenSWEBench (79.0), CoWorkBench (70.7) and RecreationBench (47.1) are the Qwen team’s own benchmarks. The card also gives no source details for JobBench or Agents’ Last Exam, so you can’t check those from the card either. Give more weight to the public ones.
  • The card doesn’t state hardware requirements. Plan for a GPU server with an inference engine. The card recommends SGLang, vLLM or TokenSpeed for production throughput.

When to pick it

Pick Qwen3.8-27B if you want an Apache-2.0, self-hostable model for agent loops that combine code with screens, documents or video. Its OSWorld-Verified (84.3), AndroidWorld (81.9) and SWE-bench Pro (61.7) scores are the strongest arguments. In our view, a single dense 27B model should also be simpler to deploy than a large MoE, though the card doesn’t make that comparison. It scores higher than Qwen3.6-27B on every benchmark the card lists, but some of the gaps are small (GPQA Diamond +1.4, OmniDocBench 1.5 +1.7, RealWorldQA +1.8).

Look elsewhere if your work is mostly hard knowledge reasoning: Qwen3.7-Plus and Opus4.6 Max score higher on HLE and GPQA Diamond. It also trails Opus4.6 Max on Terminal Bench 2.1 (73.0 vs 78.2) and NL2Repo-Bench (42.3 vs 47.6). And if you need a model that runs in-process with plain Transformers on a laptop, this card gives you no guidance for that setup.

Related