Run Xing4.0-29B-A4B: a 4B-active MoE for coding and agent tasks
China Telecom's 29B MoE activates 4B parameters per token and has a 256K context. Here's how to call it through an OpenAI-compatible API.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Xing4.0-29B-A4B is a mixture-of-experts language model from China Telecom Artificial Intelligence Technology Co., Ltd. It is the next model in the Xing series, which used to be called TeleChat. It has 29B parameters in total, but only 4B are active for each token. The native context length is 256K, and the card says it can be extended to 512K. The authors say it is the first model of this size trained entirely on Ascend NPUs with MindSpore. For most developers, the useful part is the agent and coding results. On the card’s own tables, it scores higher than Gemma4-26B-A4B and Qwen3.6-35B-A3B on Terminal-Bench 2.1, Claw-Eval and DeepresearchBII. The license is Apache 2.0.
Key specs
| Xing4.0-29B-A4B | |
|---|---|
| Parameters | 29B total, 4B active |
| Layers | 40 |
| Attention | MLA |
| Experts | 64 routed, 4 active per token, 1 shared |
| Context | 256K (extensible to 512K) |
| Architecture | mHC + MLA + MTP |
| License | Apache 2.0 |
Benchmarks from the card. All numbers are self-reported, and bold marks the highest score in each row:
| Benchmark | Xing4.0-29B-A4B | Gemma4-26B-A4B | Qwen3.6-35B-A3B |
|---|---|---|---|
| SWE-bench Verified | 75.00 | 53.00 | 76.00 |
| Terminal-Bench 2.1 | 57.50 | 30.00 | 51.50 |
| Claw-Eval | 76.55 | 71.49 | 74.54 |
| DeepresearchBII | 60.80 | 39.30 | 59.70 |
| Tau3-Bench | 64.63 | 58.90 | 67.20 |
| SWE-bench Multilingual | 66.00 | 51.00 | 67.20 |
| AIME2026 | 90.00 | 88.30 | 92.70 |
| IFBench | 69.67 | 72.67 | 65.50 |
| AA.LCR | 61.00 | 66.00 | 62.00 |
The results are mixed. Xing leads clearly on Terminal-Bench 2.1, by 6 points over Qwen3.6 and 27.5 over Gemma4. It also leads on DeepresearchBII and Claw-Eval. Against Qwen3.6 it is roughly level or slightly behind on SWE-bench Verified, SWE-bench Multilingual, Tau3-Bench and AIME2026. It scores lowest of the three on AA.LCR. The card doesn’t say what AA.LCR measures, so look up the benchmark before you weigh that result. It matters most if long context is the main reason you are considering this model.
Install
The weights are in Hugging Face Transformers format. The card lists vLLM, SGLang and KTransformers for serving, but it does not include launch commands. Those are in the GitHub repository. Once you have a server running, the card’s client code uses the openai package:
pip install openaiThe examples below read the endpoint from environment variables:
export XING_BASE_URL="http://your-server/v1"export XING_API_KEY="your-api-key"Run it
This is the card’s quickstart with the recommended settings for reasoning and general tasks:
import osfrom openai import OpenAI
client = OpenAI( base_url=os.environ["XING_BASE_URL"], api_key=os.environ["XING_API_KEY"],)
completion = client.chat.completions.create( model="Xing4.0-29B-A4B", messages=[{"role": "user", "content": "Briefly explain the basic principles of quantum computing."}], temperature=1.0, top_p=0.95, extra_body={ "repetition_penalty": 1.05, "skip_special_tokens": False, "spaces_between_special_tokens": False, "chat_template_kwargs": { "enable_thinking": True, # Set to False to disable thinking }, },)
print(completion.choices[0].message.content)The card recommends two sets of sampling settings:
| Scenario | temperature | top_p | repetition_penalty |
|---|---|---|---|
| Complex reasoning / general tasks | 1.0 | 0.95 | 1.05 |
| Coding / agent tasks | 0.8 | 0.95 | 1.05 |
Coding assistant
Terminal-Bench 2.1 is the model’s clearest lead on the card, and it scores 75.00 on SWE-bench Verified, just behind Qwen3.6. That makes coding a reasonable first test. This example uses the card’s coding settings (temperature=0.8) and keeps thinking on:
import osfrom openai import OpenAI
client = OpenAI( base_url=os.environ["XING_BASE_URL"], api_key=os.environ["XING_API_KEY"],)
buggy_code = '''def moving_average(xs, k): out = [] for i in range(len(xs)): window = xs[i:i + k] out.append(sum(window) / k) return out'''
completion = client.chat.completions.create( model="Xing4.0-29B-A4B", messages=[ {"role": "system", "content": "You are a careful senior Python reviewer."}, { "role": "user", "content": "Find the bug in this function, explain it in two sentences, " "and return a corrected version with a short test.\n\n" + buggy_code, }, ], temperature=0.8, top_p=0.95, extra_body={ "repetition_penalty": 1.05, "skip_special_tokens": False, "spaces_between_special_tokens": False, "chat_template_kwargs": {"enable_thinking": True}, },)
print(completion.choices[0].message.content)The card also says the model has “targeted adaptation and format alignment” for agent frameworks including OpenCode, Claude Code, OpenClaw and Hermes. It doesn’t explain how to connect any of them, and not all of them speak the OpenAI-compatible API. Claude Code, for example, uses the Anthropic Messages API. Check each framework’s own docs and the GitHub repo before you wire one up.
Batch intent classification
The card names intent classification as a good target for fine-tuning. You can test the model on prompted classification before you decide to fine-tune it. For short, high-volume requests like these, turning thinking off should cut output length. The card doesn’t measure how much.
import osfrom openai import OpenAI
client = OpenAI( base_url=os.environ["XING_BASE_URL"], api_key=os.environ["XING_API_KEY"],)
LABELS = ["billing", "technical_issue", "cancel_service", "other"]
messages_in = [ "My internet has been dropping every evening since Monday.", "Why was I charged twice this month?", "I want to end my contract when it expires in March.", "Do you have a store near the central station?",]
def classify(text: str) -> str: completion = client.chat.completions.create( model="Xing4.0-29B-A4B", messages=[ { "role": "system", "content": "Classify the customer message into exactly one label from: " + ", ".join(LABELS) + ". Reply with the label only.", }, {"role": "user", "content": text}, ], temperature=0.8, top_p=0.95, extra_body={ "repetition_penalty": 1.05, "skip_special_tokens": False, "spaces_between_special_tokens": False, "chat_template_kwargs": {"enable_thinking": False}, }, ) answer = (completion.choices[0].message.content or "").strip() if answer in LABELS: return answer # Special tokens may surround the label, so accept a single label found in the text. found = [label for label in LABELS if label in answer] if len(found) == 1: return found[0] return f"UNPARSED: {answer!r}"
for msg in messages_in: print(f"{classify(msg):<20} | {msg}")Because skip_special_tokens is False, the raw output may contain template tokens as well as the label. The script first tries an exact match. If that fails, it accepts the reply when exactly one label appears in it. Anything else is marked unparsed so you can look at it, instead of quietly dropping it.
Question answering over a long document
With a 256K context, you can often put a whole contract, report or codebase file into the prompt and skip retrieval. Contract auditing and knowledge-based QA are two of the domains the card lists. The usable context depends on how your server was launched (for example vLLM’s --max-model-len), not just on the model. If the server isn’t set up for long inputs, the request will be rejected.
import osimport sysfrom openai import OpenAI
client = OpenAI( base_url=os.environ["XING_BASE_URL"], api_key=os.environ["XING_API_KEY"],)
doc_path = sys.argv[1] # e.g. python long_doc_qa.py contract.txtwith open(doc_path, encoding="utf-8") as f: document = f.read()
question = "List every clause that allows either party to terminate early, with the clause number and notice period."
completion = client.chat.completions.create( model="Xing4.0-29B-A4B", messages=[ { "role": "system", "content": "Answer only from the provided document. Cite clause numbers. " "If the document does not contain the answer, say so.", }, {"role": "user", "content": f"<document>\n{document}\n</document>\n\n{question}"}, ], temperature=1.0, top_p=0.95, extra_body={ "repetition_penalty": 1.05, "skip_special_tokens": False, "spaces_between_special_tokens": False, "chat_template_kwargs": {"enable_thinking": True}, },)
print(completion.choices[0].message.content)The card doesn’t publish a long-document QA result it explains, and Xing has the lowest AA.LCR score of the three models. Check answers on your own long documents before you depend on them.
Gotchas
- No serving commands on the card. vLLM, SGLang and KTransformers are listed as supported, but launch flags, parallelism settings and dtype are only in the GitHub repo. The card has no direct
transformersloading snippet either. - Hardware isn’t specified. The 4B active parameters reduce compute per token, but all 29B parameters still have to be in memory. Plan capacity for a 29B model, not a 4B one.
- Non-standard sampling parameters.
repetition_penalty,skip_special_tokens,spaces_between_special_tokensandchat_template_kwargsgo throughextra_body, as in the card. A hosted endpoint that doesn’t support them may reject or ignore them. - Special tokens are kept. The card’s examples set
skip_special_tokens: False, so the output may contain template tokens. The card doesn’t describe how thinking output is formatted, so look at your raw responses before you write a parser. - Thinking can use a lot of tokens. The benchmark runs used large output budgets:
max_tokens=131072for AIME2026 and 64K for Terminal-Bench 2.1. Setenable_thinking: Falsewhen you don’t need reasoning. - 512K isn’t a default. The card calls 256K native and 512K an extension, and it doesn’t explain how to turn the extension on.
- The benchmarks are the vendor’s. Each one used a specific harness and settings (for example SWE-agent with a 210K context, or terminus-2 with a 24-hour timeout). Your results with a different setup may differ.
When to pick it
Pick Xing4.0-29B-A4B if you are building terminal or shell agents, or deep-research agents. Its Terminal-Bench 2.1, Claw-Eval and DeepresearchBII scores are the rows it wins on the card. It is also a good fit if you want an Apache-2.0 model to fine-tune for classification, table understanding or contract review, since the card names LLaMA-Factory and MindFormers support. If you train on Ascend hardware, note that it was trained on Ascend 910C clusters with MindSpore/MindFormers. The card doesn’t cover serving on Ascend.
Look elsewhere if strict instruction following is your main need. On the card’s own numbers, Gemma4-26B-A4B scores higher on IFBench, and also on AA.LCR. For SWE-bench-style repo fixing, multilingual SWE tasks, Tau3-Bench and AIME-style math, Qwen3.6-35B-A3B is level or slightly ahead. Also skip it for now if you need copy-paste serving instructions from the model card itself, because the deployment details are only in the GitHub repo.
Related
Run K2-Horizon-MoVA-36B-A4B for Agents and 512K-Token Context
IFM's open MoE model runs 4B active parameters with a 512K context window. How to serve it, call it, and use it for agents and long documents.
IFM/K2-Horizon-MoVA-36B-A4B
Run Kimi K3: Moonshot's 2.8T Open MoE for Agents and Coding
Kimi K3 is a 2.8T-parameter open-weight multimodal MoE with a 1M-token context. What the model card says and how to call it.
moonshotai/Kimi-K3
DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash