Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 7 min read

Qwen3.8-Flash-Next: Serve a 6B-Active Multimodal Agent Model

Qwen's 125B MoE with 6B active params handles text, images and video. What it's good at, how to call it, and where it falls short.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Qwen3.8-Flash-Next is an open-weight vision-language model from the Qwen team at Alibaba. The team calls it an experimental preview of the architecture that will power Qwen4. It is a 125B-parameter Mixture-of-Experts model with only 6B parameters active per token, plus a 51B n-gram embedding table, and it accepts text, images and video. Its main draw is agentic work. On the card’s own benchmarks it posts the top score among the compared models on SWE-bench Pro, Toolathlon Verified, AndroidWorld and the OSWorld 2.0 partial score. It does this with fewer activated parameters than every compared model whose size the card lists: Qwen3.8-27B (27B), Qwen3.7-Plus (17B active) and DeepSeek-V4-Flash-0731 (13B active). The card gives no parameter counts for Claude-Opus-4.6.

Key specs

Spec Value
Total parameters 125B (6B activated), plus 51B n-gram embedding and 4B MTP
Layers 48, laid out as 12 × (3 × Gated DeltaNet + 1 × Qwen Sparse Attention), each followed by MoE
Experts 512 total, 10 routed + 1 shared active
Context 262,144 tokens natively, up to 1,000,000 with YaRN
Inputs Text, image, video
Default mode Thinking (<think>...</think> before the answer)

The architecture changes the card lists:

  • Qwen Sparse Attention (QSA) picks micro-blocks rather than individual tokens, within a budget of 512 blocks or 2048 tokens. The goal is lower long-context latency.
  • Gated Residual widens the residual stream into 4 branches with a bottleneck rank of 320. Information flowing through them is modulated by an element-wise, data-dependent read gate and a per-branch scalar write gate.
  • N-gram embedding adds 20M bigram/trigram entries at layer 2. The card says these are cheaper to compute and easier to offload than MoE parameters.

Selected benchmark numbers from the card:

Benchmark Qwen3.8-Flash-Next Qwen3.8-27B DeepSeek-V4-Flash-0731 Claude-Opus-4.6 (Max)
SWE-bench Pro 62.5 61.7 56.0 53.4
SWE-bench Multilingual 81.0 73.8 – 77.5
NL2Repo-Bench 48.1 42.3 54.2 47.6
Toolathlon Verified 73.5 67.1 70.3 –
HLE 35.9 30.8 33.8 40.0
AndroidWorld 84.5 81.9 not compared 62.0
OSWorld 2.0 (binary / partial) 19.4 / 52.3 19.4 / 48.0 not compared –

“–” means the card lists no score. DeepSeek-V4-Flash-0731 is not part of the card’s vision-language comparison, so it has no AndroidWorld or OSWorld 2.0 entries.

The model does not win every row. DeepSeek-V4-Flash leads on NL2Repo-Bench and Claude-Opus-4.6 leads on HLE. On the OSWorld 2.0 binary score, Qwen3.8-27B ties it at 19.4; it only leads on the partial score. The Claude SWE-bench Pro score is that model’s officially published number, while every other model in that row was run on a corrected version of the benchmark. CoWorkBench and RecreationBench, also listed on the card, are in-house benchmarks.

Install

The card’s quickstart uses an OpenAI-compatible server plus the openai client:

Terminal window
pip install -U openai
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"

To start the server, follow the model-specific guide for your engine: the vLLM recipe, the SGLang cookbook or the TokenSpeed recipe. The card recommends the latest framework versions and a dedicated serving engine for production.

The Chat Completions API also works with Qwen Cloud, but the scripts below need changes first. They are written for a self-hosted server. On Qwen Cloud you have to change model, and you pass enable_thinking and preserve_thinking directly instead of wrapping them in chat_template_kwargs.

Run it

This is a minimal non-streaming text request. Thinking is on by default, and reasoning_effort accepts xhigh (the default), medium and low.

hello.py
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="Qwen/Qwen3.8-Flash-Next",
messages=[
{"role": "user", "content": "Write a Python function to merge two sorted linked lists."},
],
temperature=1.0,
top_p=0.95,
reasoning_effort="medium",
extra_body={"top_k": 20},
)
print(response.choices[0].message.content)

Extract data from screenshots and charts

For extraction you usually want direct answers without a reasoning trace. Set enable_thinking to False and use the card’s instruct-mode sampling settings. The script below runs a batch of chart or screenshot images through the model and asks for JSON in the prompt. The URLs are placeholders, so replace them with your own images. The card does not mention a structured-output mode, so validate what comes back.

extract_charts.py
import json
from openai import OpenAI
client = OpenAI()
image_urls = [
"https://example.com/your-chart-01.png", # replace with your chart URL
"https://example.com/your-screenshot-01.png", # replace with your screenshot URL
]
PROMPT = (
"Describe this image as JSON with keys: "
'"type" (chart, screenshot, photo or other), "title" (string or null), '
'"key_values" (list of short strings). Return only the JSON.'
)
results = []
for url in image_urls:
response = client.chat.completions.create(
model="Qwen/Qwen3.8-Flash-Next",
messages=[{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": url}},
{"type": "text", "text": PROMPT},
],
}],
temperature=0.7,
top_p=0.8,
presence_penalty=1.5,
extra_body={
"top_k": 20,
"chat_template_kwargs": {"enable_thinking": False},
},
)
text = response.choices[0].message.content
try:
results.append({"url": url, "data": json.loads(text)})
except json.JSONDecodeError:
results.append({"url": url, "raw": text})
print(json.dumps(results, indent=2))

For reference, the card reports CharXiv (RQ) scores of 84.6 “without CI” and 90.6 “with CI” (the card does not define CI). In the “without CI” setting, Qwen3.7-Plus scores slightly higher at 85.8. The card does not say whether these scores were measured with thinking on or off, so don’t assume they carry over to the non-thinking setup above.

Multi-turn agent sessions with preserved thinking

By default the model keeps the thinking blocks from every earlier turn. The card says this helps keep decisions consistent across turns and improves KV cache utilization. To get that benefit, append each assistant reply to the history together with its reasoning, as the card’s streaming example does:

agent_session.py
from openai import OpenAI
client = OpenAI()
messages = []
def ask(user_text):
messages.append({"role": "user", "content": user_text})
stream = client.chat.completions.create(
model="Qwen/Qwen3.8-Flash-Next",
messages=messages,
extra_body={
"chat_template_kwargs": {
"enable_thinking": True,
"preserve_thinking": True,
},
},
reasoning_effort="xhigh",
stream=True,
stream_options={"include_usage": True},
)
reasoning, answer = "", ""
for chunk in stream:
if not chunk.choices:
print("\nUsage:", chunk.usage)
continue
delta = chunk.choices[0].delta
if getattr(delta, "reasoning_content", None) is not None:
reasoning += delta.reasoning_content
elif getattr(delta, "reasoning", None) is not None:
reasoning += delta.reasoning
if getattr(delta, "content", None):
print(delta.content, end="", flush=True)
answer += delta.content
messages.append({
"role": "assistant",
"content": answer,
"reasoning_content": reasoning,
"reasoning": reasoning,
})
return answer
ask("Plan a refactor that splits a 2,000-line utils.py into modules. List the steps.")
ask("Now write the code for step 1 only.")

The card warns that in multi-turn agentic tasks, lowering reasoning_effort can make the whole task slower. Shallower analysis per turn can cause more failures and retries. Measure total task time before you lower it.

Ask questions about a video

The model also takes video input. On LVBench, the card’s long-video benchmark, it scores 76.6.

video_qa.py
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="Qwen/Qwen3.8-Flash-Next",
messages=[{
"role": "user",
"content": [
{
"type": "video_url",
"video_url": {
"url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/video/N1cdUjctpG8.mp4"
},
},
{
"type": "text",
"text": "How many porcelain jars were discovered in the niches located in the primary chamber of the tomb?",
},
],
}],
)
print(response.choices[0].message.content)

If vLLM was launched with --media-io-kwargs '{"video": {"num_frames": -1}}', you can control frame sampling by passing extra_body={"mm_processor_kwargs": {"fps": 2, "do_sample_frames": True}}. The card says this works only in vLLM.

Gotchas

  • Thinking is on by default. If you don’t need the reasoning, turn it off. On Qwen Cloud, pass "enable_thinking": False directly. On self-hosted servers, pass it inside chat_template_kwargs. The same rule applies to preserve_thinking.

  • Each mode has its own sampling settings. Thinking mode: temperature=1.0, top_p=0.95, top_k=20, presence_penalty=0.0. Instruct mode: temperature=0.7, top_p=0.80, top_k=20, presence_penalty=1.5. On frameworks that support it, the card suggests adjusting presence_penalty between 0 and 2 to reduce endless repetition. It warns that higher values can occasionally cause language mixing and a slight drop in model performance.

  • Agent tasks need large output budgets. For engines that let you set reasoning and answer limits separately, the card suggests 262,144 tokens for reasoning and 131,072 for the final response.

  • YaRN is static. To go past 262,144 tokens you need to override rope_parameters. For vLLM the card gives this command, where ... stands for your usual serve arguments:

    Terminal window
    VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 vllm serve ... --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --max-model-len 1000000

    Static scaling can hurt quality on shorter inputs. Enable it only when you need long context, and pick a factor that fits your typical length. For example, the card suggests 2.0 if your inputs are usually around 524,288 tokens.

  • Long-video defaults are conservative. For hour-long videos, set longest_edge to 469,762,048 in video_preprocessor_config.json (the card pairs it with shortest_edge: 4096).

  • The license is custom. It is qwen-community-1.0, not Apache. Read the repo’s LICENSE file before using the model commercially.

  • Speed depends on the engine. The card says inference efficiency varies a lot between frameworks.

When to pick it

Pick Qwen3.8-Flash-Next if you are building agents that write code, use tools or operate GUIs, and you can self-host a large MoE. On SWE-bench Pro, SWE-bench Multilingual, Toolathlon Verified, AndroidWorld and ClawEval-MM, its card scores beat every compared model that has a reported score, and it gets there with 6B active parameters. It also handles image and video understanding in the same deployment.

Skip it in these cases:

  • Your hardware can’t hold the weights. Only 6B parameters are active, but you still have to store all 125B language-model parameters. The 51B of n-gram embeddings and 4B of MTP add to that, though the card says the n-gram embeddings are easier to offload than MoE parameters. The vision encoder adds more weights on top, and the card doesn’t give its size. The card gives no hardware requirements either.
  • You need repo-level code generation or HLE-style multidisciplinary reasoning. On the card’s own numbers, DeepSeek-V4-Flash leads NL2Repo-Bench and Claude-Opus-4.6 leads HLE. On GPQA Diamond, Qwen3.8-Flash-Next leads (91.7 vs Claude’s 91.3).
  • You need a stable, permissively licensed base. This is an experimental preview released under a custom license.
  • You want 1M context out of the box. The card points to Qwen3.8-Flash on Qwen Cloud, which has 1M context by default and built-in tools.

Related