Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 8 min read

Run K2-Horizon-MoVA-36B-A4B for Agents and 512K-Token Context

IFM's open MoE model runs 4B active parameters with a 512K context window. How to serve it, call it, and use it for agents and long documents.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

K2-Horizon-MoVA-36B-A4B is the sparse model in IFM’s K2-Horizon family. It is a Mixture-of-Experts model that uses Mixture-of-Values attention (MoVA). It has 36B parameters in total, but only 4B are active for each token. What sets the release apart is how much IFM published. Besides the weights, you get the training data, the training code, the training logs, and intermediate checkpoints from every stage of training. The model is released under Apache 2.0. The card does not state licenses for the data, the code repository or the logs, so check each one before you reuse it. In the card’s results it has the top score among the listed open models on tau3-Banking and Terminal-Bench 2.1, both of which test agentic use. It does not lead on most of the reasoning and knowledge benchmarks.

Key specs

Total / active params 36B / 4B
Architecture MoE with Mixture-of-Values attention
Context 524,288 tokens natively (512K)
Language English (en in the card’s metadata)
Training 22.9T pretraining tokens, then four midtraining stages and two SFT phases
License Apache 2.0 (model)

These are the benchmark scores that matter for choosing it. All numbers come from the card. The card says the baseline scores are from Artificial Analysis, with Muse Glimmer-30B at high reasoning effort and the other open models in their reasoning mode. The card does not say where K2-Horizon’s own scores come from. The rank column compares the seven open models in the card’s table.

Benchmark K2-Horizon-MoVA Rank (of 7) Best other open model in table
tau3-Banking (agentic tool use) 26.8 1st 23.5 (Muse Glimmer-30B)
Terminal-Bench 2.1 (agentic terminal) 58.6 1st 53.9 (Nemotron 3 Ultra, 550B)
Humanity’s Last Exam (no tools) 25.2 2nd 28.4 (Nemotron 3 Ultra)
GPQA Diamond 80.8 5th 86.7 (Nemotron 3 Ultra)
SciCode 38.9 4th 43.6 (Muse Glimmer-30B)
AA-LCR (long-context reasoning) 66.3 5th 80.0 (Muse Glimmer-30B)
AA-Omniscience Accuracy 18.8 5th (tied) 27.0 (Muse Glimmer-30B)
AA-Omniscience Non-Hallucination 69.2 3rd 87.0 (G9v3-39A5B)

On the two agentic benchmarks it scores higher than every model in the table, including Nemotron 3 Ultra, which has 550B parameters. The card files Terminal-Bench 2.1 under “Coding” but describes it as “agentic terminal use”. It is 2nd on Humanity’s Last Exam. On GPQA Diamond, SciCode and AA-LCR it is mid-table or lower. On non-hallucination it is 3rd, close behind Nemotron 3 Ultra (70.3) and well ahead of Qwen3.6-35B-A3B (49.5), Muse Glimmer-30B (18.1), Gemma 4 31B-it (15.0) and Nemotron 3 Super (13.0).

Every training stage has a checkpoint branch, such as pretrain_1100000, mid_4_10000 and sft_1_11000. The card’s serving commands use --revision main. The card does not say which training checkpoint main corresponds to.

Install

The card validates the Transformers path with these exact versions:

Terminal window
pip install "transformers==5.15.0" "torch==2.13.0" "safetensors==0.8.0"

The Transformers example below uses device_map="auto", which in Transformers normally requires the accelerate package. The card does not list or pin accelerate, so install it separately if you get an ImportError:

Terminal window
pip install accelerate

For serving, the card points to the vLLM recipe and the SGLang K2 Horizon cookbook. The client examples below also need pip install openai.

Run it

This is the SGLang server command from the card. It is the recipe the card validated on 2x H200, and the one its Best Practices recommend:

Terminal window
python3 -m sglang.launch_server \
--model-path IFM/K2-Horizon-MoVA-36B-A4B \
--revision main \
--tp 2 \
--ep 2 \
--dtype bfloat16 \
--attention-backend fa3 \
--json-model-override-args '{"xllm_source_router_gemm_partitions":2}' \
--reasoning-parser k2_horizon \
--tool-call-parser k2_horizon \
--host 0.0.0.0 --port 30000

The card also gives a vLLM command. It does not include the xllm_source_router_gemm_partitions override, which the card says preserves the checkpoint’s router numerics, so it is not a drop-in equivalent of the SGLang recipe:

Terminal window
vllm serve IFM/K2-Horizon-MoVA-36B-A4B \
--revision main \
--tensor-parallel-size 2 \
--enable-expert-parallel \
--trust-remote-code \
--dtype bfloat16 \
--reasoning-parser k2_horizon \
--tool-call-parser k2_horizon \
--enable-auto-tool-choice

Call the server with the OpenAI client. The card’s examples use port 30000, which is the SGLang server above. The vLLM command doesn’t set a port, so if you serve with vLLM, change base_url to match your server.

chat.py
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="IFM/K2-Horizon-MoVA-36B-A4B",
messages=[{"role": "user", "content": "Explain the result step by step."}],
temperature=1.0,
top_p=0.95,
max_tokens=32768,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high", "tool_call_format": "xml"}},
)
message = response.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print("Answer:", message.content)

You can also run it in plain Transformers (see the accelerate note under Install):

hf_generate.py
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "IFM/K2-Horizon-MoVA-36B-A4B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id, device_map="auto", dtype="bfloat16", low_cpu_mem_usage=True, trust_remote_code=True
)
inputs = tokenizer("Explain why long-context evaluation is difficult.", return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Agent tool calls

Its best results are on tau3-Banking and Terminal-Bench, so tool calling is the main reason to try it. Start the server with the k2_horizon tool-call parser, as in the commands above, and choose a call format with tool_call_format (json, xml or xml_typed; the default is xml). The card does not include a full tool-call request. The tools payload below uses the standard OpenAI-compatible tools schema, and the card does not confirm how either server handles it. Check it against your server. The script also sends the whole assistant message back with messages.append(msg), which may include reasoning_content. The card does not say whether reasoning should be passed back on tool-result turns, so treat that part as unverified too.

agent_tool.py
import json
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
MODEL = "IFM/K2-Horizon-MoVA-36B-A4B"
EXTRA = {"chat_template_kwargs": {"reasoning_effort": "high", "tool_call_format": "xml"}}
def get_balance(account_id: str) -> dict:
return {"account_id": account_id, "balance": 1520.75, "currency": "USD"}
tools = [{
"type": "function",
"function": {
"name": "get_balance",
"description": "Return the current balance for a bank account.",
"parameters": {
"type": "object",
"properties": {"account_id": {"type": "string"}},
"required": ["account_id"],
},
},
}]
messages = [{"role": "user", "content": "What's the balance on account ACC-1042?"}]
resp = client.chat.completions.create(
model=MODEL, messages=messages, tools=tools,
temperature=1.0, top_p=0.95, max_tokens=32768, extra_body=EXTRA,
)
msg = resp.choices[0].message
if msg.tool_calls:
messages.append(msg)
for call in msg.tool_calls:
args = json.loads(call.function.arguments)
result = get_balance(**args)
messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)})
resp = client.chat.completions.create(
model=MODEL, messages=messages, tools=tools,
temperature=1.0, top_p=0.95, max_tokens=32768, extra_body=EXTRA,
)
msg = resp.choices[0].message
print(msg.content)

Long-document Q&A

The context window is 524,288 tokens natively, and every stage from midtraining Stage 3 onward trained at 512K. That lets you send much longer documents in one prompt than short-context models allow. Long-context reasoning is not its strong point, though. On AA-LCR it scores 66.3, 5th of the 7 open models. Muse Glimmer-30B (80.0), Nemotron 3 Ultra (71.0), Gemma 4 31B-it (68.3) and Qwen3.6-35B-A3B (66.7) all score higher.

The card does not confirm that the default SGLang or vLLM configuration serves the full 512K tokens, or that a 512K prompt’s KV cache fits on 2x H200. Check your server’s max context length setting and memory before sending very long documents. The prompt plus max_tokens must also stay within 524,288 tokens.

long_doc_qa.py
import sys
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
with open(sys.argv[1], encoding="utf-8") as f:
document = f.read()
question = sys.argv[2]
response = client.chat.completions.create(
model="IFM/K2-Horizon-MoVA-36B-A4B",
messages=[{
"role": "user",
"content": f"<document>\n{document}\n</document>\n\nAnswer using only the document. "
f"Quote the passages you rely on.\n\nQuestion: {question}",
}],
temperature=1.0,
top_p=0.95,
max_tokens=32768,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(response.choices[0].message.content)

Run it with python long_doc_qa.py report.txt "What are the termination clauses?".

Batch extraction

The script below sends support tickets to the server in parallel and collects structured fields. It reads only content, because the reasoning parser returns the model’s reasoning separately in reasoning_content. The card gives no throughput, latency or cost figures. For measured H200 latency and throughput, see the SGLang cookbook and the vLLM recipe. Only 4B parameters are active per token, but all 36B still have to be loaded. High reasoning effort with max_tokens=32768 per request can also add up across a large batch.

batch_extract.py
import json
from concurrent.futures import ThreadPoolExecutor
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
tickets = [
"My card was charged twice for order 8812. Please refund one charge.",
"App crashes on login since the last update, iPhone 15.",
"How do I change the email on my account?",
]
PROMPT = (
"Extract fields from this support ticket. Reply with only a JSON object with keys "
'"category" (billing, bug, account, other), "urgency" (low, medium, high), and "summary".\n\n'
"Ticket: {ticket}"
)
def extract(ticket: str) -> dict:
resp = client.chat.completions.create(
model="IFM/K2-Horizon-MoVA-36B-A4B",
messages=[{"role": "user", "content": PROMPT.format(ticket=ticket)}],
temperature=1.0,
top_p=0.95,
max_tokens=32768,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
text = resp.choices[0].message.content or ""
try:
return json.loads(text.strip().removeprefix("```json").removesuffix("```"))
except json.JSONDecodeError:
return {"error": "unparsable", "raw": text}
with ThreadPoolExecutor(max_workers=8) as pool:
for ticket, result in zip(tickets, pool.map(extract, tickets)):
print(ticket[:40], "->", result)

Gotchas

  • Always use reasoning_effort: "high". The card reports every result at high effort and recommends medium and low only for research on reasoning effort. The card’s examples set max_tokens=32768.
  • Sampling. Use temperature=1.0 and top_p=0.95.
  • Router numerics on SGLang. Keep --json-model-override-args '{"xllm_source_router_gemm_partitions":2}'. The card says this override preserves the checkpoint’s router numerics. The card’s vLLM command does not include it.
  • Parsers. Turn on the k2_horizon reasoning parser for chat, and add the tool-call parser for agent use. Turn both off for plain completion-style generation.
  • trust_remote_code. The card’s Transformers and vLLM examples pass it. Its SGLang command does not.
  • accelerate for Transformers. The card’s Transformers example uses device_map="auto", which normally needs accelerate. The card’s validated versions don’t include it, so you may need to install it yourself.
  • Fine-tuning caveat. The bundled Hugging Face implementation can train differently from IFM’s native xLLM, including in the auxiliary load-balancing loss. You can fine-tune the HF checkpoints with SFT. To continue the original pretraining with matching behavior, use the native xLLM checkpoint and runtime.
  • Factual recall is mid-table. AA-Omniscience accuracy is 18.8, tied with Qwen3.6-35B-A3B for 5th of 7. Its non-hallucination rate (69.2) is 3rd of 7, behind G9v3-39A5B (87.0) and just behind Nemotron 3 Ultra (70.3). For factual questions, it is safer to ground answers in documents you provide.
  • Language. The card’s metadata lists only English (en). The card says nothing about other languages.
  • Technical report. As of the card’s 2026-09-28 update the report was listed as in progress, with an expected date of end of September 2026. Check whether it has been released.

When to pick it

Pick it if you are building agents for tool use or terminal tasks on your own hardware. In the card’s table it has the top tau3-Banking and Terminal-Bench 2.1 scores of any open model listed, while activating only 4B parameters per token. The 512K native context and Apache 2.0 model license also matter for on-prem long-document work, though its AA-LCR score is 5th of 7. Researchers have a further reason: the released data, code, logs and per-stage checkpoints let you study how capabilities change during training.

Skip it if your workload depends mainly on science QA or long-context reasoning. Nemotron 3 Ultra, Muse Glimmer-30B, Gemma 4 31B-it and Qwen3.6-35B-A3B each score higher than it on at least one of those benchmarks in the same table. The card’s metadata lists only English, so test it yourself before relying on it for other languages. The card does not state minimum hardware. Its only validated serving recipe is SGLang on 2x H200 in BF16, so plan for multi-GPU serving, or test smaller setups yourself.

Related