Run Qwen3.8-27B for coding agents, documents and video
Qwen3.8-27B is an Apache-2.0 dense 27B vision-language model with adjustable reasoning. How to serve it and call it today.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Qwen3.8-27B is the Qwen team’s new dense 27B model. It reads images and video natively, and its thinking mode can be switched on or off for each request. It uses the Qwen3.5 architecture and is released under Apache 2.0. Its main selling point is agentic work: on the team’s own tables, this 27B model scores above Qwen3.7-Plus and the “Opus4.6 Max” baseline on several computer-use and agentic-coding benchmarks. The usual caveat applies, though: some of those benchmarks are in-house.
Key specs
| Parameters | 27B, dense |
| Layers | 64, in a hybrid layout: 16 × (3 × Gated DeltaNet + 1 × Gated Attention), each followed by an FFN |
| Context | 262,144 tokens natively, extendable to 1,000,000 with YaRN |
| Modalities | Text, image, video input |
| Reasoning control | enable_thinking, reasoning_effort (xhigh / medium / low), preserve_thinking |
| Extras | Multi-token prediction (MTP) trained with multiple steps |
These are the benchmarks most useful for deciding whether to try it (numbers from the card):
| Benchmark | Qwen3.8-27B | Qwen3.6-27B | Qwen3.7-Plus | Opus4.6 Max |
|---|---|---|---|---|
| SWE-bench Pro | 61.7 | 53.5 | 57.6 | 53.4* |
| Terminal Bench 2.1 | 73.0 | 63.4 | 64.0 | 78.2 |
| OSWorld-Verified | 84.3 | 63.9 | 73.3 | 72.7 |
| AndroidWorld | 81.9 | 70.3 | 81.0 | 62.0 |
| OmniDocBench 1.5 | 91.1 | 89.4 | 91.4 | 86.6 |
| GPQA Diamond | 89.2 | 87.8 | 90.3 | 91.3 |
| HLE | 30.8 | 24.0 | 34.7 | 40.0 |
*The card notes that the Opus4.6 Max SWE-bench Pro score is the officially reported number. The other models were re-evaluated with the Claude Code harness on a corrected version of the benchmark.
Overall, it is strong on computer use and agentic coding. On knowledge-heavy reasoning (HLE, GPQA Diamond) it still trails Qwen3.7-Plus and Opus4.6 Max.
Install
The card recommends running the model behind an inference server and calling it through an OpenAI-compatible API. To serve it yourself, use one of the framework guides the card links:
- vLLM: https://recipes.vllm.ai/Qwen/Qwen3.8-27B
- SGLang: https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-27B
- TokenSpeed: https://lightseek.org/tokenspeed/recipes/models#qwen3-8
Hosted access will be available through Qwen Cloud, which the card describes as “coming soon”.
Install the client and point it at your endpoint:
pip install -U openai
export OPENAI_BASE_URL='your-base-url'export OPENAI_API_KEY='your-api-key'Run it
Thinking is on by default. With streaming, the reasoning arrives in reasoning_content (or reasoning, depending on the server) and the answer arrives in content:
from openai import OpenAI
client = OpenAI()
messages = [{"role": "user", "content": "Write a Python function to merge two sorted linked lists."}]
completion = client.chat.completions.create( model="Qwen/Qwen3.8-27B", messages=messages, reasoning_effort="xhigh", stream=True, stream_options={"include_usage": True},)
reasoning, answer = "", ""for chunk in completion: if not chunk.choices: print("\nUsage:", chunk.usage) continue delta = chunk.choices[0].delta if getattr(delta, "reasoning_content", None) is not None: reasoning += delta.reasoning_content elif getattr(delta, "reasoning", None) is not None: reasoning += delta.reasoning if getattr(delta, "content", None): print(delta.content, end="", flush=True) answer += delta.contentBatch-read images in non-thinking mode
For extraction, captioning or quick visual Q&A you may not need a reasoning trace. Turn thinking off and use the instruct sampling settings from the card (temperature=0.7, top_p=0.8, top_k=20, presence_penalty=1.5). The example uses the card’s demo images. Replace them with your own scans or screenshots.
from openai import OpenAI
client = OpenAI()
jobs = [ ("https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/RealWorld/RealWorld-04.png", "Where is this?"), ("https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/CI_Demo/mathv-1327.jpg", "Describe the geometric figure in this image in one paragraph."),]
for url, question in jobs: response = client.chat.completions.create( model="Qwen/Qwen3.8-27B", messages=[{ "role": "user", "content": [ {"type": "image_url", "image_url": {"url": url}}, {"type": "text", "text": question}, ], }], temperature=0.7, top_p=0.8, presence_penalty=1.5, extra_body={ "top_k": 20, "chat_template_kwargs": {"enable_thinking": False}, }, ) print(url, "->", response.choices[0].message.content, "\n")The card reports 91.1 on OmniDocBench 1.5 and 83.7 on CharXiv (RQ) in its “Without CI” setting. It doesn’t say whether thinking was on or off for those runs, so check non-thinking quality on your own documents before relying on it.
Ask questions about a video
Video goes in as a video_url content part. On vLLM you can also control frame sampling per request with mm_processor_kwargs. According to the card, this only works if the server was launched with --media-io-kwargs '{"video": {"num_frames": -1}}'. The defaults are fps=2 and do_sample_frames=True, so only pass it when you want a different fps. It is commented out below, as it is on the card.
from openai import OpenAI
client = OpenAI()
messages = [{ "role": "user", "content": [ {"type": "video_url", "video_url": {"url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/video/N1cdUjctpG8.mp4"}}, {"type": "text", "text": "How many porcelain jars were discovered in the niches located in the primary chamber of the tomb?"}, ],}]
response = client.chat.completions.create( model="Qwen/Qwen3.8-27B", messages=messages, # vLLM only, and only when launched with --media-io-kwargs '{"video": {"num_frames": -1}}'. # Uncomment and change fps to set your own sampling rate (defaults: fps=2, do_sample_frames=True): # extra_body={"mm_processor_kwargs": {"fps": 2, "do_sample_frames": True}},)print(response.choices[0].message.content)For hour-long videos, the card recommends raising longest_edge in video_preprocessor_config.json (see Gotchas).
Multi-turn coding with tunable reasoning
reasoning_effort trades depth for speed. preserve_thinking keeps earlier reasoning in the context, which the card says helps agents stay consistent and improves KV-cache reuse. To carry a conversation forward, append the assistant turn together with its reasoning, as the card does:
from openai import OpenAI
client = OpenAI()
def ask(messages, effort="medium"): stream = client.chat.completions.create( model="Qwen/Qwen3.8-27B", messages=messages, reasoning_effort=effort, temperature=1.0, top_p=0.95, extra_body={ "top_k": 20, "chat_template_kwargs": {"enable_thinking": True, "preserve_thinking": True}, }, stream=True, ) reasoning, answer = "", "" for chunk in stream: if not chunk.choices: continue d = chunk.choices[0].delta if getattr(d, "reasoning_content", None) is not None: reasoning += d.reasoning_content elif getattr(d, "reasoning", None) is not None: reasoning += d.reasoning if getattr(d, "content", None): answer += d.content messages.append({ "role": "assistant", "content": answer, "reasoning_content": reasoning, "reasoning": reasoning, }) return answer
messages = [{"role": "user", "content": "Write a Python LRU cache class with get/put in O(1)."}]print(ask(messages))
messages.append({"role": "user", "content": "Now make it thread-safe and add unit tests."})print(ask(messages, effort="xhigh"))Gotchas
- Thinking is on by default. Every response starts with a
<think>block unless you setenable_thinking: False. On Qwen Cloud, put"enable_thinking": Falseand"preserve_thinking": Falsedirectly inextra_body, not insidechat_template_kwargs. - There are two sampling profiles. Thinking mode uses
temperature=1.0, top_p=0.95, presence_penalty=0.0. Non-thinking mode usestemperature=0.7, top_p=0.80, presence_penalty=1.5. Raisingpresence_penalty(range 0–2) reduces repetition, but it can cause language mixing and a small quality drop. Not every framework supports every parameter. - Lower effort isn’t always faster. The card warns that in multi-turn agent tasks,
loweffort can lead to more failures and retries, which raises total latency and token use. - Agentic tasks need large output budgets. On frameworks that support separate limits, the card suggests up to 262,144 tokens for reasoning and 131,072 for the final response. It frames these as “within the 1M context length”. Together they come to about 393K tokens, more than the native 262,144 context, so these budgets assume you have enabled the YaRN-extended context.
- YaRN is static. To go past 262K tokens, either edit
rope_parametersundertext_configinconfig.json, or pass overrides at launch. The card’s commands need three things: an environment variable that allows the longer length (VLLM_ALLOW_LONG_MAX_MODEL_LEN=1,SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1orTOKENSPEED_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1), the rope override (--hf-overridesfor vLLM and TokenSpeed,--json-model-override-argsfor SGLang), and the context flag (--max-model-len 1000000for vLLM and TokenSpeed,--context-length 1000000for SGLang). Copy the full commands from the card’s “Processing Ultra-Long Texts” section. The open frameworks apply a constant scaling factor, which can hurt short inputs. Only enable it when you need it, and sizefactorto your typical length (for example,2.0for about 524K tokens). - The video config is conservative by default. For hour-scale video, set
{"longest_edge": 469762048, "shortest_edge": 4096}invideo_preprocessor_config.json, which corresponds to about 224K video tokens. - There’s no Transformers code on the card. The weights are in Transformers format, but every example on the card goes through an API server. The 1M default context and built-in tools mentioned on the card belong to the hosted Qwen Cloud version, not the open weights.
- Several headline numbers are in-house. QwenSWEBench (79.0), CoWorkBench (70.7) and RecreationBench (47.1) are the Qwen team’s own benchmarks. The card also gives no source details for JobBench or Agents’ Last Exam, so you can’t check those from the card either. Give more weight to the public ones.
- The card doesn’t state hardware requirements. Plan for a GPU server with an inference engine. The card recommends SGLang, vLLM or TokenSpeed for production throughput.
When to pick it
Pick Qwen3.8-27B if you want an Apache-2.0, self-hostable model for agent loops that combine code with screens, documents or video. Its OSWorld-Verified (84.3), AndroidWorld (81.9) and SWE-bench Pro (61.7) scores are the strongest arguments. In our view, a single dense 27B model should also be simpler to deploy than a large MoE, though the card doesn’t make that comparison. It scores higher than Qwen3.6-27B on every benchmark the card lists, but some of the gaps are small (GPQA Diamond +1.4, OmniDocBench 1.5 +1.7, RealWorldQA +1.8).
Look elsewhere if your work is mostly hard knowledge reasoning: Qwen3.7-Plus and Opus4.6 Max score higher on HLE and GPQA Diamond. It also trails Opus4.6 Max on Terminal Bench 2.1 (73.0 vs 78.2) and NL2Repo-Bench (42.3 vs 47.6). And if you need a model that runs in-process with plain Transformers on a laptop, this card gives you no guidance for that setup.
Related
Run Kimi K3: Moonshot's 2.8T Open MoE for Agents and Coding
Kimi K3 is a 2.8T-parameter open-weight multimodal MoE with a 1M-token context. What the model card says and how to call it.
moonshotai/Kimi-K3
Qwen3.8-Flash-Next: Serve a 6B-Active Multimodal Agent Model
Qwen's 125B MoE with 6B active params handles text, images and video. What it's good at, how to call it, and where it falls short.
Qwen/Qwen3.8-Flash-Next
Run Xing4.0-29B-A4B: a 4B-active MoE for coding and agent tasks
China Telecom's 29B MoE activates 4B parameters per token and has a 256K context. Here's how to call it through an OpenAI-compatible API.
XingChen-AGI/Xing4.0-29B-A4B