Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 8 min read

Run Nex-N2.5-mini on 2 H100s as a Tool-Calling Agent Model

Nex-AGI's smallest Nex-N2.5 agent model, served with SGLang on 2 H100s, with tool calling and per-request thinking modes.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Nex-N2.5-mini is the smallest model in Nex-AGI’s Nex-N2.5 family. The family is built for long, multi-step agent tasks: operating computers and browsers, and running and testing programs. Mini and Pro build on the multimodal base of Nex-N2. Nex-AGI uses vision to let the agent check what happened on screen and correct itself. Mini’s main selling point is size. It is the only Nex-N2.5 model with a published single-node recipe on two H100s. On OSWorld-G, the card’s figures put it 4th of the 8 models with scores. It is above Claude Opus 5 and GPT-5.6 Sol, and behind Nex-N2.5-Pro, Qwen3.8-Max and GLM-5.3-Flash. The model is licensed Apache-2.0. The card says the Nex-N2.5 weights “will be released as open source”, and it links the Hugging Face and ModelScope repos.

Key specs

The card doesn’t list a parameter count or context length for mini. Here is what it does give:

Nex-N2.5-mini
License Apache-2.0
Serving Nex-AGI’s SGLang fork, nexagi/sglang:v0.5.18-nex-patch
Reference hardware 1 node, 2 x H100, --tp 2
Reasoning parser qwen3
Tool-call parser qwen3_coder
Thinking control reasoning_effort: "none", "medium" (default), "high"
Recommended sampling temperature=0.7, top_p=0.95, top_k=40
Hosted option OpenRouter (nex-agi/nex-n2.5-mini)

Selected scores for mini from the card’s tables:

Benchmark Nex-N2.5-mini Nex-N2.5-Pro
OSWorld-G 82.9 87.4
OmniDoc 89.7 92.2
BrowseComp 83.4 89.7
OSWorld-Verified 71.2 82.2
WebArena-Verified 63.4 67.6
Terminal-Bench 2.1 73.4 82.7
Toolathlon Verified 54.6 68.5
SWE-Bench Pro 43.8 61.2

On OSWorld-G, mini scores 82.9. That is higher than the card’s figures for Claude Opus 5 (76.8) and GPT-5.6 Sol (77.7), but lower than Pro (87.4), Qwen3.8-Max (84.9) and GLM-5.3-Flash (83.3). OSWorld-Verified is weaker. Mini’s 71.2 beats GLM-5.3-Flash (62.3) but falls below every other model in that row, including MiniMax-M3 (75.2), DeepSeek-V4-Flash-Vision (76.7) and Qwen3.8-Max (86.1). On the coding benchmarks in the text table, mini is well behind every other model listed.

Keep in mind that these comparisons mix sources. Nex-AGI ran the coding evals for its own models with its NexAU harness. It ran the computer-use evals with its NexCUA harness, which hasn’t been released yet. According to the card, competitor scores come from official leaderboards and the providers’ own reports where those exist. Nex-AGI only ran its own evaluations for results with no public source. Because the setups may differ, the head-to-head numbers may not be directly comparable.

Install

Pull the prebuilt image that contains Nex-AGI’s patched SGLang:

Terminal window
docker pull nexagi/sglang:v0.5.18-nex-patch

Download the weights into a local directory from the Hugging Face repo or ModelScope. The card describes the open-source release in the future tense, so check that the weights are actually in the repo before you plan around them. You also need the chat template (see Gotchas).

The client examples below only need requests:

Terminal window
pip install requests

Run it

This launch command for mini on one node with 2 x H100 is based on the card’s. The one change is a second -v mount. The card’s command only mounts the model at /model, so its --chat-template path doesn’t exist inside the container. Here the template’s directory is mounted at /template and --chat-template points there. Replace both host paths with your own.

Terminal window
docker run --gpus all --shm-size 32g --ipc=host \
-p 30000:30000 \
-v /path/to/your/model:/model \
-v /path/to/nex-N2.5-mini:/template \
nexagi/sglang:v0.5.18-nex-patch \
python3 -m sglang.launch_server \
--model-path /model \
--tp 2 \
--host 0.0.0.0 --port 30000 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--chat-template /template/chat-template.jinja \
--mamba-scheduler-strategy extra_buffer

The server accepts OpenAI-compatible Chat Completions requests. Below is a minimal client that uses the card’s recommended sampling settings. Set MODEL to the name your server exposes.

hello.py
import requests
URL = "http://localhost:30000/v1/chat/completions"
MODEL = "<served-model-name>" # replace with the name your server exposes
resp = requests.post(URL, json={
"model": MODEL,
"messages": [{"role": "user", "content": "Explain how binary search works."}],
"reasoning_effort": "medium",
"temperature": 0.7,
"top_p": 0.95,
"top_k": 40,
}, timeout=600)
resp.raise_for_status()
print(resp.json()["choices"][0]["message"]["content"])

A note on top_k: the card lists top_k = 40 as a recommended sampling value, but its example request body doesn’t include it. top_k is also not part of the standard OpenAI Chat Completions API. SGLang’s OpenAI-compatible server accepts it as an extra parameter, but a strict OpenAI client or a different gateway may reject it. If yours does, drop it. This applies to every example in this post.

The server runs with --reasoning-parser qwen3, which separates the reasoning trace from the final answer. That leaves only the answer in content.

Agent tool call

Tool calling is turned on by the --tool-call-parser qwen3_coder flag in the launch command. The card doesn’t show a request body for tools. This example uses the standard OpenAI Chat Completions tools format, which an OpenAI-compatible server should accept. It runs one full cycle: the model asks for a tool, the script runs it, and the result goes back to the model. The script looks up each call by function name and returns an error to the model for unknown tools or bad arguments. It also stops after a fixed number of turns, so a model that keeps calling tools can’t loop forever.

agent_tool.py
import json
import requests
URL = "http://localhost:30000/v1/chat/completions"
MODEL = "<served-model-name>"
MAX_TURNS = 8
def get_ticket_status(ticket_id: str) -> dict:
# Stand-in for a real backend call
return {"ticket_id": ticket_id, "status": "waiting_on_customer", "age_days": 4}
HANDLERS = {"get_ticket_status": get_ticket_status}
TOOLS = [{
"type": "function",
"function": {
"name": "get_ticket_status",
"description": "Look up the current status of a support ticket.",
"parameters": {
"type": "object",
"properties": {"ticket_id": {"type": "string"}},
"required": ["ticket_id"],
},
},
}]
def chat(messages):
r = requests.post(URL, json={
"model": MODEL,
"messages": messages,
"tools": TOOLS,
"reasoning_effort": "medium",
"temperature": 0.7, "top_p": 0.95, "top_k": 40,
}, timeout=600)
r.raise_for_status()
return r.json()["choices"][0]["message"]
def run_tool(call) -> dict:
name = call["function"]["name"]
handler = HANDLERS.get(name)
if handler is None:
return {"error": f"unknown tool: {name}"}
try:
args = json.loads(call["function"]["arguments"] or "{}")
return handler(**args)
except (json.JSONDecodeError, TypeError) as e:
return {"error": f"bad arguments for {name}: {e}"}
messages = [{"role": "user", "content": "What's going on with ticket T-4821? Should I follow up?"}]
msg = chat(messages)
turns = 0
while msg.get("tool_calls"):
if turns >= MAX_TURNS:
raise RuntimeError(f"stopped after {MAX_TURNS} tool-call turns")
turns += 1
messages.append(msg)
for call in msg["tool_calls"]:
messages.append({
"role": "tool",
"tool_call_id": call["id"],
"content": json.dumps(run_tool(call)),
})
msg = chat(messages)
print(msg["content"])

Batch extraction with thinking turned off

High-volume, structured work like tagging, extraction or routing usually doesn’t need a reasoning trace for every row. With reasoning_effort: "none", the model answers directly. The example below sends requests concurrently instead of one at a time.

batch_extract.py
import json
from concurrent.futures import ThreadPoolExecutor
import requests
URL = "http://localhost:30000/v1/chat/completions"
MODEL = "<served-model-name>"
PROMPT = (
"Extract fields from the support email. Reply with JSON only, keys: "
'"product", "issue_type" (one of: billing, bug, how_to, other), "urgent" (true/false).\n\n'
"Email:\n{email}"
)
emails = [
"Hi, I was charged twice for my Pro plan this month. Please fix ASAP.",
"How do I export my dashboard to PDF in Analytics?",
"The mobile app crashes every time I open Settings on Android 15.",
]
def extract(email: str) -> dict:
r = requests.post(URL, json={
"model": MODEL,
"messages": [{"role": "user", "content": PROMPT.format(email=email)}],
"reasoning_effort": "none",
"temperature": 0.7, "top_p": 0.95, "top_k": 40,
}, timeout=300)
r.raise_for_status()
text = r.json()["choices"][0]["message"]["content"].strip()
try:
return json.loads(text)
except json.JSONDecodeError:
return {"error": "unparseable", "raw": text}
with ThreadPoolExecutor(max_workers=16) as pool:
for email, out in zip(emails, pool.map(extract, emails)):
print(out, "<-", email[:50])

The card doesn’t mention a JSON or structured-output mode, so validate the output yourself, as the script does.

Choosing a thinking mode per request

The card defines three thinking modes. A simple router can choose one per request based on the task type. Short lookups stay fast, and multi-step problems get full reasoning.

router.py
import requests
URL = "http://localhost:30000/v1/chat/completions"
MODEL = "<served-model-name>"
EFFORT = {
"lookup": "none", # respond directly
"general": "medium", # model decides whether to think
"planning": "high", # always think first
}
def ask(prompt: str, kind: str = "general") -> str:
r = requests.post(URL, json={
"model": MODEL,
"messages": [{"role": "user", "content": prompt}],
"reasoning_effort": EFFORT[kind],
"temperature": 0.7, "top_p": 0.95, "top_k": 40,
}, timeout=900)
r.raise_for_status()
return r.json()["choices"][0]["message"]["content"]
print(ask("What does HTTP 429 mean?", "lookup"))
print(ask("Plan a migration of a cron-based ETL job to an event-driven queue. List steps and risks.", "planning"))

Gotchas

  • Use the patched SGLang. The card deploys with Nex-AGI’s fork in nexagi/sglang:v0.5.18-nex-patch. The repo is tagged transformers, but the card doesn’t show a plain transformers inference path.
  • Make the chat template path visible inside the container. The card’s command mounts only the model at /model, but its --chat-template points to /path/to/nex-N2.5-mini/chat-template.jinja. That path has to exist inside the container. The command in “Run it” adds a second mount for this. If you use the card’s command as-is, mount the file’s directory or change the path.
  • Use reasoning_effort to control thinking. The chat template reads reasoning_effort. Flags like enable_thinking or thinking_mode won’t work unless a gateway translates them.
  • Use the right parsers. Mini needs --reasoning-parser qwen3 and --tool-call-parser qwen3_coder. The deepseek-r1 parser is for Max, not mini.
  • Expect to need two H100s. The card’s only mini recipe uses 2 x H100 with --tp 2. It gives no guidance on quantization or smaller GPUs.
  • Check that the weights are published. The card says the weights “will be released as open source” while linking the repos. Confirm the files are there before you count on self-hosting.
  • Your benchmark numbers may differ. Where no public source existed, Nex-AGI measured the scores itself with its own harnesses: NexAU for coding and NexCUA for computer use. NexCUA isn’t public yet. Competitor numbers come from leaderboards and provider reports where available, so they may have been measured under different setups. BrowseComp used context compaction once usage passed 60% of the context window. WebTest ran in oracle mode.
  • Image input isn’t documented. Mini is described as multimodal, but the card doesn’t show how to send images. Test this before you build a computer-use agent on it.

When to pick it

Pick Nex-N2.5-mini if you want an Apache-2.0 agent model that fits on one 2 x H100 node, and benchmarks like OSWorld-G, OmniDoc and BrowseComp are closest to your workload. Its OSWorld-G (82.9) and OmniDoc (89.7) scores are within a few points of Pro’s (87.4 and 92.2). BrowseComp is a wider gap: 83.4 against Pro’s 89.7. Pro’s reference setup uses 8 x H100. Mini also works as a self-hosted tool-calling backend with thinking you can switch per request.

Skip it for serious coding agents. On SWE-Bench Pro (43.8) and DeepSWE (36.1), it is far behind Pro and every other model in the text table. On the multimodal SWE-MM benchmark it scores 25.5, well behind Pro’s 38.2 but ahead of GLM-5.3-Flash’s 20.6. Skip it for long multi-step desktop tasks as well. It scores 30.5 on OSWorld-2 against Pro’s 56.4, and it is below most models in the OSWorld-Verified row. If you don’t have two H100s, try the hosted mini on OpenRouter before you commit to self-hosting.

Related