Qwen3.8-Flash-Next: Serve a 6B-Active Multimodal Agent Model
Qwen's 125B MoE with 6B active params handles text, images and video. What it's good at, how to call it, and where it falls short.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Qwen3.8-Flash-Next is an open-weight vision-language model from the Qwen team at Alibaba. The team calls it an experimental preview of the architecture that will power Qwen4. It is a 125B-parameter Mixture-of-Experts model with only 6B parameters active per token, plus a 51B n-gram embedding table, and it accepts text, images and video. Its main draw is agentic work. On the card’s own benchmarks it posts the top score among the compared models on SWE-bench Pro, Toolathlon Verified, AndroidWorld and the OSWorld 2.0 partial score. It does this with fewer activated parameters than every compared model whose size the card lists: Qwen3.8-27B (27B), Qwen3.7-Plus (17B active) and DeepSeek-V4-Flash-0731 (13B active). The card gives no parameter counts for Claude-Opus-4.6.
Key specs
| Spec | Value |
|---|---|
| Total parameters | 125B (6B activated), plus 51B n-gram embedding and 4B MTP |
| Layers | 48, laid out as 12 × (3 × Gated DeltaNet + 1 × Qwen Sparse Attention), each followed by MoE |
| Experts | 512 total, 10 routed + 1 shared active |
| Context | 262,144 tokens natively, up to 1,000,000 with YaRN |
| Inputs | Text, image, video |
| Default mode | Thinking (<think>...</think> before the answer) |
The architecture changes the card lists:
- Qwen Sparse Attention (QSA) picks micro-blocks rather than individual tokens, within a budget of 512 blocks or 2048 tokens. The goal is lower long-context latency.
- Gated Residual widens the residual stream into 4 branches with a bottleneck rank of 320. Information flowing through them is modulated by an element-wise, data-dependent read gate and a per-branch scalar write gate.
- N-gram embedding adds 20M bigram/trigram entries at layer 2. The card says these are cheaper to compute and easier to offload than MoE parameters.
Selected benchmark numbers from the card:
| Benchmark | Qwen3.8-Flash-Next | Qwen3.8-27B | DeepSeek-V4-Flash-0731 | Claude-Opus-4.6 (Max) |
|---|---|---|---|---|
| SWE-bench Pro | 62.5 | 61.7 | 56.0 | 53.4 |
| SWE-bench Multilingual | 81.0 | 73.8 | – | 77.5 |
| NL2Repo-Bench | 48.1 | 42.3 | 54.2 | 47.6 |
| Toolathlon Verified | 73.5 | 67.1 | 70.3 | – |
| HLE | 35.9 | 30.8 | 33.8 | 40.0 |
| AndroidWorld | 84.5 | 81.9 | not compared | 62.0 |
| OSWorld 2.0 (binary / partial) | 19.4 / 52.3 | 19.4 / 48.0 | not compared | – |
“–” means the card lists no score. DeepSeek-V4-Flash-0731 is not part of the card’s vision-language comparison, so it has no AndroidWorld or OSWorld 2.0 entries.
The model does not win every row. DeepSeek-V4-Flash leads on NL2Repo-Bench and Claude-Opus-4.6 leads on HLE. On the OSWorld 2.0 binary score, Qwen3.8-27B ties it at 19.4; it only leads on the partial score. The Claude SWE-bench Pro score is that model’s officially published number, while every other model in that row was run on a corrected version of the benchmark. CoWorkBench and RecreationBench, also listed on the card, are in-house benchmarks.
Install
The card’s quickstart uses an OpenAI-compatible server plus the openai client:
pip install -U openai
export OPENAI_BASE_URL="http://localhost:8000/v1"export OPENAI_API_KEY="EMPTY"To start the server, follow the model-specific guide for your engine: the vLLM recipe, the SGLang cookbook or the TokenSpeed recipe. The card recommends the latest framework versions and a dedicated serving engine for production.
The Chat Completions API also works with Qwen Cloud, but the scripts below need changes first. They are written for a self-hosted server. On Qwen Cloud you have to change model, and you pass enable_thinking and preserve_thinking directly instead of wrapping them in chat_template_kwargs.
Run it
This is a minimal non-streaming text request. Thinking is on by default, and reasoning_effort accepts xhigh (the default), medium and low.
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create( model="Qwen/Qwen3.8-Flash-Next", messages=[ {"role": "user", "content": "Write a Python function to merge two sorted linked lists."}, ], temperature=1.0, top_p=0.95, reasoning_effort="medium", extra_body={"top_k": 20},)print(response.choices[0].message.content)Extract data from screenshots and charts
For extraction you usually want direct answers without a reasoning trace. Set enable_thinking to False and use the card’s instruct-mode sampling settings. The script below runs a batch of chart or screenshot images through the model and asks for JSON in the prompt. The URLs are placeholders, so replace them with your own images. The card does not mention a structured-output mode, so validate what comes back.
import jsonfrom openai import OpenAI
client = OpenAI()
image_urls = [ "https://example.com/your-chart-01.png", # replace with your chart URL "https://example.com/your-screenshot-01.png", # replace with your screenshot URL]
PROMPT = ( "Describe this image as JSON with keys: " '"type" (chart, screenshot, photo or other), "title" (string or null), ' '"key_values" (list of short strings). Return only the JSON.')
results = []for url in image_urls: response = client.chat.completions.create( model="Qwen/Qwen3.8-Flash-Next", messages=[{ "role": "user", "content": [ {"type": "image_url", "image_url": {"url": url}}, {"type": "text", "text": PROMPT}, ], }], temperature=0.7, top_p=0.8, presence_penalty=1.5, extra_body={ "top_k": 20, "chat_template_kwargs": {"enable_thinking": False}, }, ) text = response.choices[0].message.content try: results.append({"url": url, "data": json.loads(text)}) except json.JSONDecodeError: results.append({"url": url, "raw": text})
print(json.dumps(results, indent=2))For reference, the card reports CharXiv (RQ) scores of 84.6 “without CI” and 90.6 “with CI” (the card does not define CI). In the “without CI” setting, Qwen3.7-Plus scores slightly higher at 85.8. The card does not say whether these scores were measured with thinking on or off, so don’t assume they carry over to the non-thinking setup above.
Multi-turn agent sessions with preserved thinking
By default the model keeps the thinking blocks from every earlier turn. The card says this helps keep decisions consistent across turns and improves KV cache utilization. To get that benefit, append each assistant reply to the history together with its reasoning, as the card’s streaming example does:
from openai import OpenAI
client = OpenAI()messages = []
def ask(user_text): messages.append({"role": "user", "content": user_text}) stream = client.chat.completions.create( model="Qwen/Qwen3.8-Flash-Next", messages=messages, extra_body={ "chat_template_kwargs": { "enable_thinking": True, "preserve_thinking": True, }, }, reasoning_effort="xhigh", stream=True, stream_options={"include_usage": True}, )
reasoning, answer = "", "" for chunk in stream: if not chunk.choices: print("\nUsage:", chunk.usage) continue delta = chunk.choices[0].delta if getattr(delta, "reasoning_content", None) is not None: reasoning += delta.reasoning_content elif getattr(delta, "reasoning", None) is not None: reasoning += delta.reasoning if getattr(delta, "content", None): print(delta.content, end="", flush=True) answer += delta.content
messages.append({ "role": "assistant", "content": answer, "reasoning_content": reasoning, "reasoning": reasoning, }) return answer
ask("Plan a refactor that splits a 2,000-line utils.py into modules. List the steps.")ask("Now write the code for step 1 only.")The card warns that in multi-turn agentic tasks, lowering reasoning_effort can make the whole task slower. Shallower analysis per turn can cause more failures and retries. Measure total task time before you lower it.
Ask questions about a video
The model also takes video input. On LVBench, the card’s long-video benchmark, it scores 76.6.
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create( model="Qwen/Qwen3.8-Flash-Next", messages=[{ "role": "user", "content": [ { "type": "video_url", "video_url": { "url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/video/N1cdUjctpG8.mp4" }, }, { "type": "text", "text": "How many porcelain jars were discovered in the niches located in the primary chamber of the tomb?", }, ], }],)print(response.choices[0].message.content)If vLLM was launched with --media-io-kwargs '{"video": {"num_frames": -1}}', you can control frame sampling by passing extra_body={"mm_processor_kwargs": {"fps": 2, "do_sample_frames": True}}. The card says this works only in vLLM.
Gotchas
-
Thinking is on by default. If you don’t need the reasoning, turn it off. On Qwen Cloud, pass
"enable_thinking": Falsedirectly. On self-hosted servers, pass it insidechat_template_kwargs. The same rule applies topreserve_thinking. -
Each mode has its own sampling settings. Thinking mode:
temperature=1.0, top_p=0.95, top_k=20, presence_penalty=0.0. Instruct mode:temperature=0.7, top_p=0.80, top_k=20, presence_penalty=1.5. On frameworks that support it, the card suggests adjustingpresence_penaltybetween 0 and 2 to reduce endless repetition. It warns that higher values can occasionally cause language mixing and a slight drop in model performance. -
Agent tasks need large output budgets. For engines that let you set reasoning and answer limits separately, the card suggests 262,144 tokens for reasoning and 131,072 for the final response.
-
YaRN is static. To go past 262,144 tokens you need to override
rope_parameters. For vLLM the card gives this command, where...stands for your usual serve arguments:Terminal window VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 vllm serve ... --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --max-model-len 1000000Static scaling can hurt quality on shorter inputs. Enable it only when you need long context, and pick a
factorthat fits your typical length. For example, the card suggests 2.0 if your inputs are usually around 524,288 tokens. -
Long-video defaults are conservative. For hour-long videos, set
longest_edgeto 469,762,048 invideo_preprocessor_config.json(the card pairs it withshortest_edge: 4096). -
The license is custom. It is
qwen-community-1.0, not Apache. Read the repo’s LICENSE file before using the model commercially. -
Speed depends on the engine. The card says inference efficiency varies a lot between frameworks.
When to pick it
Pick Qwen3.8-Flash-Next if you are building agents that write code, use tools or operate GUIs, and you can self-host a large MoE. On SWE-bench Pro, SWE-bench Multilingual, Toolathlon Verified, AndroidWorld and ClawEval-MM, its card scores beat every compared model that has a reported score, and it gets there with 6B active parameters. It also handles image and video understanding in the same deployment.
Skip it in these cases:
- Your hardware can’t hold the weights. Only 6B parameters are active, but you still have to store all 125B language-model parameters. The 51B of n-gram embeddings and 4B of MTP add to that, though the card says the n-gram embeddings are easier to offload than MoE parameters. The vision encoder adds more weights on top, and the card doesn’t give its size. The card gives no hardware requirements either.
- You need repo-level code generation or HLE-style multidisciplinary reasoning. On the card’s own numbers, DeepSeek-V4-Flash leads NL2Repo-Bench and Claude-Opus-4.6 leads HLE. On GPQA Diamond, Qwen3.8-Flash-Next leads (91.7 vs Claude’s 91.3).
- You need a stable, permissively licensed base. This is an experimental preview released under a custom license.
- You want 1M context out of the box. The card points to Qwen3.8-Flash on Qwen Cloud, which has 1M context by default and built-in tools.
Related
Self-Host MiMo-V2.6-Pro-RL, Xiaomi's 1T MoE Agent Model
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL
Run Qwen3.8-27B for coding agents, documents and video
Qwen3.8-27B is an Apache-2.0 dense 27B vision-language model with adjustable reasoning. How to serve it and call it today.
Qwen/Qwen3.8-27B
CLM-v0.1-8B: rank candidates and route agent decisions fast
A contrastive scorer on frozen Qwen3-8B embeddings. It ranks the candidates you give it and answers typed questions about a state.
Contrastive-LM/CLM-v0.1-8B