Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 6 min read

GLM-5.3-Flash: What You Can Run With Z.ai's 18B-Active Multimodal Model

Z.ai's first natively multimodal GLM-5 model: 320B total / 18B active parameters, hybrid sparse-linear attention, MIT license. What the card covers.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

GLM-5.3-Flash is an open-weight model from Z.ai (zai-org) and the first natively multimodal model in the GLM-5 series. It takes images and text as input and produces text. It has 320B total parameters and 18B active parameters. The card doesn’t name the architecture type, but a total/active split like this usually means a sparse mixture-of-experts design. The most notable change is in attention. This is the first GLM model to combine sparse and linear attention, and Z.ai says this “sharply” reduces the cost of serving long contexts while keeping long-context accuracy. It ships under the MIT license.

One thing to know before you read further: the model card gives almost no code. It links out to framework-specific guides instead of showing inference snippets. This post sticks to what the card says. Where the card stops, we point you to the guide it links rather than guess at an API.

Key specs

Total parameters 320B
Active parameters 18B
Modality Image + text in, text out
Languages (card metadata) English, Chinese
Attention Hybrid sparse + linear
Other architecture Manifold-Constrained Hyper-Connections (mHC)
Pre-training data 30T-token multimodal corpus, new base model
Thinking control reasoning_effort: low, high, max (default max)
License MIT

Z.ai says GLM-5.3-Flash beats GLM-5.2 “across benchmarks and real-world workloads at one-tenth the price” and comes close to Claude Opus 4.8 on coding and agentic benchmarks. The actual scores are only in a chart image on the card, so we don’t repeat any numbers here. Check the chart and the Z.ai blog post before you rely on those comparisons.

The card’s footnotes cover these benchmarks: HLE with tools, NL2Repo, DeepSWE, Terminal-Bench 2.1, Agent’s Last Exam, Toolathlon Verified, AutomationBench, GDPval-AA v2, and BabyVision. Not every footnote gives evaluation settings. The Agent’s Last Exam footnote is empty, and the GDPval-AA v2 footnote only says that models are evaluated by Artificial Analysis. The card doesn’t describe what each benchmark measures. Several of the setups it does describe involve coding harnesses and long agent runs, which is covered below.

Install

The card lists six supported serving stacks. Each one links to its own setup guide:

The card doesn’t give install commands, version pins or launch flags for any of these, so neither do we. The Transformers link points to a glm5_next doc page on the main branch. Check that page and the other guides for requirements before you install. If you don’t want to host the model yourself, it is also served on the Z.ai API Platform.

Run it

The card documents two controls that change behavior. You’ll want to set both on purpose:

  • reasoning_effort is a parameter that sets the thinking budget. Valid values are low, high and max. If you leave it out, or pass any other value, the model uses max. To use low or high, you have to pass them explicitly. The card doesn’t say what a smaller budget does to cost or latency. Our guess is that a smaller budget gives cheaper, faster responses, but that’s an inference, not a claim from the card.
  • clear_thinking is a chat-template setting that defaults to false. For chat use, the card says to pass clear_thinking=true explicitly.

The card doesn’t say how each framework passes these values through. Check the linked guide for the stack you choose. For sampling, the card gives settings only for its evaluations, not as general recommendations. The footnotes mostly use temperature=1.0 with top_p=0.95 (HLE, BabyVision) or top_p=1.0 (NL2Repo, Terminal-Bench, and DeepSWE at temperature=0.95). Our suggestion: if you want behavior close to the published evaluations, start from those values.

Coding agents and terminal work

Much of the card’s evaluation setup is agentic coding: DeepSWE run with the mini-swe-agent harness, Terminal-Bench 2.1 run inside Claude Code 2.1.207, and NL2Repo. If you’re thinking about using GLM-5.3-Flash as the backend for a coding agent, these are the settings Z.ai used:

  • Long generations. Terminal-Bench and NL2Repo used max_new_tokens of 65,536 and 64k respectively. The card doesn’t say how many tokens thinking uses at reasoning_effort=max, but evaluation limits this large suggest you shouldn’t set a small output limit.
  • Large context. DeepSWE ran with 400K context and NL2Repo with 1M context. Z.ai says the hybrid attention reduces long-context serving costs.
  • Long timeouts. DeepSWE and Terminal-Bench were each run with a 6-hour timeout. The card doesn’t say what the timeout applies to.

Z.ai also notes that for NL2Repo, “to prevent hacking,” it used rule-based and LLM-based judging against “malicious behaviors (e.g., unauthorized pip or curl operations)”. The card doesn’t report that the model actually did these things. Still, when you give this model a shell, sandbox it and restrict network and package-install access, the same way you would for any other agent.

Long-context tool use

The HLE-with-tools setup is a useful template for tool-using research agents. It allowed up to 163,840 generated tokens, a 300,000-token context window, and “a context management strategy.” The card doesn’t describe that strategy. So don’t assume you can fill 300K tokens with raw tool output and get the published results. You will need your own truncation or summarization of tool results.

Toolathlon Verified and AutomationBench are also on the list. For Toolathlon Verified, Z.ai used the official evaluation service and reports pass@1 averaged over 3 runs. For AutomationBench, it used v1.0.6 with the fix from PR #13. If your workload is workflow automation, look up what these benchmarks cover and check the chart for their scores.

Image understanding

This is the first GLM-5 model that accepts images natively. The card’s only vision-specific detail is in the BabyVision footnote: input images were resized so their shorter side was at least 1.5K pixels, “consistent with other baselines,” with a 164K-token context and temperature=1.0, top_p=0.95. That resizing was done to match the evaluation’s baselines, and the card doesn’t recommend it for general use. If you want to reproduce the evaluation setup, use the same resizing. As general guidance, not something the card says, larger images will likely use more context.

Gotchas

  • The default is maximum thinking. If you don’t set reasoning_effort, every request runs at max. Set low or high explicitly when you don’t need the full thinking budget.
  • Invalid values silently become max. A typo such as "medium" doesn’t change the budget to something in between. The card says any other value means max.
  • Chat needs clear_thinking=true. It defaults to false, and the card tells you to turn it on for chat.
  • 18B active parameters doesn’t mean 18B of memory. The full 320B parameter set still has to be available. The card gives no hardware requirements, so plan your capacity using the linked serving guides.
  • Benchmark numbers are image-only. The card’s text claims (“outperforms GLM-5.2”, “approaching Claude Opus 4.8”) have no inline scores. Read the chart, and test on your own tasks.
  • Language coverage. The card metadata lists only English and Chinese.

When to pick it

GLM-5.3-Flash is worth testing if you self-host and need agentic coding, long-context tool use, or a single model for both images and text, under a permissive MIT license. Z.ai says the hybrid attention reduces long-context serving costs, which matters for long agent sessions. Separately, the card says the model outperforms GLM-5.2 “at one-tenth the price”, but it doesn’t say whether that means API pricing or serving cost.

Don’t pick it if you need something that fits on one consumer GPU. Nothing on the card suggests that it does. If you work mainly in languages other than English and Chinese, note that only those two are listed, so test before you commit. Also skip it if you need copy-paste inference code today. The card leaves setup to external guides, and you’ll need to work through SGLang, vLLM or Transformers docs yourself. If you just want to try it, the hosted Z.ai API is the quickest way in.

Related