Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 7 min read

Run Xing4.0-29B-A4B: a 4B-active MoE for coding and agent tasks

China Telecom's 29B MoE activates 4B parameters per token and has a 256K context. Here's how to call it through an OpenAI-compatible API.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Xing4.0-29B-A4B is a mixture-of-experts language model from China Telecom Artificial Intelligence Technology Co., Ltd. It is the next model in the Xing series, which used to be called TeleChat. It has 29B parameters in total, but only 4B are active for each token. The native context length is 256K, and the card says it can be extended to 512K. The authors say it is the first model of this size trained entirely on Ascend NPUs with MindSpore. For most developers, the useful part is the agent and coding results. On the card’s own tables, it scores higher than Gemma4-26B-A4B and Qwen3.6-35B-A3B on Terminal-Bench 2.1, Claw-Eval and DeepresearchBII. The license is Apache 2.0.

Key specs

Xing4.0-29B-A4B
Parameters 29B total, 4B active
Layers 40
Attention MLA
Experts 64 routed, 4 active per token, 1 shared
Context 256K (extensible to 512K)
Architecture mHC + MLA + MTP
License Apache 2.0

Benchmarks from the card. All numbers are self-reported, and bold marks the highest score in each row:

Benchmark Xing4.0-29B-A4B Gemma4-26B-A4B Qwen3.6-35B-A3B
SWE-bench Verified 75.00 53.00 76.00
Terminal-Bench 2.1 57.50 30.00 51.50
Claw-Eval 76.55 71.49 74.54
DeepresearchBII 60.80 39.30 59.70
Tau3-Bench 64.63 58.90 67.20
SWE-bench Multilingual 66.00 51.00 67.20
AIME2026 90.00 88.30 92.70
IFBench 69.67 72.67 65.50
AA.LCR 61.00 66.00 62.00

The results are mixed. Xing leads clearly on Terminal-Bench 2.1, by 6 points over Qwen3.6 and 27.5 over Gemma4. It also leads on DeepresearchBII and Claw-Eval. Against Qwen3.6 it is roughly level or slightly behind on SWE-bench Verified, SWE-bench Multilingual, Tau3-Bench and AIME2026. It scores lowest of the three on AA.LCR. The card doesn’t say what AA.LCR measures, so look up the benchmark before you weigh that result. It matters most if long context is the main reason you are considering this model.

Install

The weights are in Hugging Face Transformers format. The card lists vLLM, SGLang and KTransformers for serving, but it does not include launch commands. Those are in the GitHub repository. Once you have a server running, the card’s client code uses the openai package:

Terminal window
pip install openai

The examples below read the endpoint from environment variables:

Terminal window
export XING_BASE_URL="http://your-server/v1"
export XING_API_KEY="your-api-key"

Run it

This is the card’s quickstart with the recommended settings for reasoning and general tasks:

quickstart.py
import os
from openai import OpenAI
client = OpenAI(
base_url=os.environ["XING_BASE_URL"],
api_key=os.environ["XING_API_KEY"],
)
completion = client.chat.completions.create(
model="Xing4.0-29B-A4B",
messages=[{"role": "user", "content": "Briefly explain the basic principles of quantum computing."}],
temperature=1.0,
top_p=0.95,
extra_body={
"repetition_penalty": 1.05,
"skip_special_tokens": False,
"spaces_between_special_tokens": False,
"chat_template_kwargs": {
"enable_thinking": True, # Set to False to disable thinking
},
},
)
print(completion.choices[0].message.content)

The card recommends two sets of sampling settings:

Scenario temperature top_p repetition_penalty
Complex reasoning / general tasks 1.0 0.95 1.05
Coding / agent tasks 0.8 0.95 1.05

Coding assistant

Terminal-Bench 2.1 is the model’s clearest lead on the card, and it scores 75.00 on SWE-bench Verified, just behind Qwen3.6. That makes coding a reasonable first test. This example uses the card’s coding settings (temperature=0.8) and keeps thinking on:

code_review.py
import os
from openai import OpenAI
client = OpenAI(
base_url=os.environ["XING_BASE_URL"],
api_key=os.environ["XING_API_KEY"],
)
buggy_code = '''
def moving_average(xs, k):
out = []
for i in range(len(xs)):
window = xs[i:i + k]
out.append(sum(window) / k)
return out
'''
completion = client.chat.completions.create(
model="Xing4.0-29B-A4B",
messages=[
{"role": "system", "content": "You are a careful senior Python reviewer."},
{
"role": "user",
"content": "Find the bug in this function, explain it in two sentences, "
"and return a corrected version with a short test.\n\n" + buggy_code,
},
],
temperature=0.8,
top_p=0.95,
extra_body={
"repetition_penalty": 1.05,
"skip_special_tokens": False,
"spaces_between_special_tokens": False,
"chat_template_kwargs": {"enable_thinking": True},
},
)
print(completion.choices[0].message.content)

The card also says the model has “targeted adaptation and format alignment” for agent frameworks including OpenCode, Claude Code, OpenClaw and Hermes. It doesn’t explain how to connect any of them, and not all of them speak the OpenAI-compatible API. Claude Code, for example, uses the Anthropic Messages API. Check each framework’s own docs and the GitHub repo before you wire one up.

Batch intent classification

The card names intent classification as a good target for fine-tuning. You can test the model on prompted classification before you decide to fine-tune it. For short, high-volume requests like these, turning thinking off should cut output length. The card doesn’t measure how much.

classify_batch.py
import os
from openai import OpenAI
client = OpenAI(
base_url=os.environ["XING_BASE_URL"],
api_key=os.environ["XING_API_KEY"],
)
LABELS = ["billing", "technical_issue", "cancel_service", "other"]
messages_in = [
"My internet has been dropping every evening since Monday.",
"Why was I charged twice this month?",
"I want to end my contract when it expires in March.",
"Do you have a store near the central station?",
]
def classify(text: str) -> str:
completion = client.chat.completions.create(
model="Xing4.0-29B-A4B",
messages=[
{
"role": "system",
"content": "Classify the customer message into exactly one label from: "
+ ", ".join(LABELS) + ". Reply with the label only.",
},
{"role": "user", "content": text},
],
temperature=0.8,
top_p=0.95,
extra_body={
"repetition_penalty": 1.05,
"skip_special_tokens": False,
"spaces_between_special_tokens": False,
"chat_template_kwargs": {"enable_thinking": False},
},
)
answer = (completion.choices[0].message.content or "").strip()
if answer in LABELS:
return answer
# Special tokens may surround the label, so accept a single label found in the text.
found = [label for label in LABELS if label in answer]
if len(found) == 1:
return found[0]
return f"UNPARSED: {answer!r}"
for msg in messages_in:
print(f"{classify(msg):<20} | {msg}")

Because skip_special_tokens is False, the raw output may contain template tokens as well as the label. The script first tries an exact match. If that fails, it accepts the reply when exactly one label appears in it. Anything else is marked unparsed so you can look at it, instead of quietly dropping it.

Question answering over a long document

With a 256K context, you can often put a whole contract, report or codebase file into the prompt and skip retrieval. Contract auditing and knowledge-based QA are two of the domains the card lists. The usable context depends on how your server was launched (for example vLLM’s --max-model-len), not just on the model. If the server isn’t set up for long inputs, the request will be rejected.

long_doc_qa.py
import os
import sys
from openai import OpenAI
client = OpenAI(
base_url=os.environ["XING_BASE_URL"],
api_key=os.environ["XING_API_KEY"],
)
doc_path = sys.argv[1] # e.g. python long_doc_qa.py contract.txt
with open(doc_path, encoding="utf-8") as f:
document = f.read()
question = "List every clause that allows either party to terminate early, with the clause number and notice period."
completion = client.chat.completions.create(
model="Xing4.0-29B-A4B",
messages=[
{
"role": "system",
"content": "Answer only from the provided document. Cite clause numbers. "
"If the document does not contain the answer, say so.",
},
{"role": "user", "content": f"<document>\n{document}\n</document>\n\n{question}"},
],
temperature=1.0,
top_p=0.95,
extra_body={
"repetition_penalty": 1.05,
"skip_special_tokens": False,
"spaces_between_special_tokens": False,
"chat_template_kwargs": {"enable_thinking": True},
},
)
print(completion.choices[0].message.content)

The card doesn’t publish a long-document QA result it explains, and Xing has the lowest AA.LCR score of the three models. Check answers on your own long documents before you depend on them.

Gotchas

  • No serving commands on the card. vLLM, SGLang and KTransformers are listed as supported, but launch flags, parallelism settings and dtype are only in the GitHub repo. The card has no direct transformers loading snippet either.
  • Hardware isn’t specified. The 4B active parameters reduce compute per token, but all 29B parameters still have to be in memory. Plan capacity for a 29B model, not a 4B one.
  • Non-standard sampling parameters. repetition_penalty, skip_special_tokens, spaces_between_special_tokens and chat_template_kwargs go through extra_body, as in the card. A hosted endpoint that doesn’t support them may reject or ignore them.
  • Special tokens are kept. The card’s examples set skip_special_tokens: False, so the output may contain template tokens. The card doesn’t describe how thinking output is formatted, so look at your raw responses before you write a parser.
  • Thinking can use a lot of tokens. The benchmark runs used large output budgets: max_tokens=131072 for AIME2026 and 64K for Terminal-Bench 2.1. Set enable_thinking: False when you don’t need reasoning.
  • 512K isn’t a default. The card calls 256K native and 512K an extension, and it doesn’t explain how to turn the extension on.
  • The benchmarks are the vendor’s. Each one used a specific harness and settings (for example SWE-agent with a 210K context, or terminus-2 with a 24-hour timeout). Your results with a different setup may differ.

When to pick it

Pick Xing4.0-29B-A4B if you are building terminal or shell agents, or deep-research agents. Its Terminal-Bench 2.1, Claw-Eval and DeepresearchBII scores are the rows it wins on the card. It is also a good fit if you want an Apache-2.0 model to fine-tune for classification, table understanding or contract review, since the card names LLaMA-Factory and MindFormers support. If you train on Ascend hardware, note that it was trained on Ascend 910C clusters with MindSpore/MindFormers. The card doesn’t cover serving on Ascend.

Look elsewhere if strict instruction following is your main need. On the card’s own numbers, Gemma4-26B-A4B scores higher on IFBench, and also on AA.LCR. For SWE-bench-style repo fixing, multilingual SWE tasks, Tau3-Bench and AIME-style math, Qwen3.6-35B-A3B is level or slightly ahead. Also skip it for now if you need copy-paste serving instructions from the model card itself, because the deployment details are only in the GitHub repo.

Related