Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 8 min read

Run MiniCPM5-2B Locally for Tool Calls and Long Documents

MiniCPM5-2B is a 2.5B Llama-architecture model with 131K context and tool calling. Here's how to serve it and what it's good for.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

MiniCPM5-2B is a dense 2.5B-parameter chat model from OpenBMB, released under Apache-2.0 and aimed at local assistants, coding agents and tool-use workflows on modest hardware. On OpenBMB’s own benchmark set it averages 53.9. That puts it ahead of the 2B-class models it was compared against, and also ahead of every larger model in the table, the closest being Qwen3.5-4B at 51.1. It is also a standard LlamaForCausalLM, and the card says mainstream engines such as vLLM, SGLang, llama.cpp and Transformers load it with “no custom kernels, no model-code fork”. One exception: the card’s SGLang command for DSpark speculative decoding still passes --trust-remote-code (more on that below).

Key specs

Parameters 2,516,756,480 (1,981,982,720 non-embedding)
Architecture LlamaForCausalLM, 42 layers
Attention GQA, 16 query heads / 2 KV heads
Context length 131,072 tokens
Languages English, Chinese
Weights BF16. Also MLX 4-bit, GPTQ 4-bit, GGUF (the card’s example uses an F16 file) and LiteRT-LM builds
Post-training 400B tokens of deep-thinking SFT, then RL, then on-policy distillation from 16 RL expert models

Selected scores from the card. For each benchmark the table shows the best 2B-class rival and the best of the larger (4B-class) models:

Benchmark MiniCPM5-2B Best 2B-class rival Best larger model
LiveCodeBench v6 69.1 42.9 (Gemma-4-E2B-it) 58.9 (granite-4.2-3B)
AIME 2025 86.5 41.9 (LFM2.5-2.6B) 79.4 (granite-4.2-3B)
BFCL v4 66.6 61.1 (LFM2.5-2.6B) 56.8 (Qwen3.5-4B)
τ²-Bench Telecom 97.1 90.4 (LFM2.5-2.6B) 92.1† (Qwen3.5-4B)
SWE-bench Verified 46.4 6.0 (LFM2.5-2.6B) 36.8 (granite-4.2-3B)
NoLiMa (long context) 68.1 17.1 (Qwen3.5-2B) 43.5 (Qwen3.5-4B)
AA-LCR (long context) 59.0† 28.7† (Qwen3.5-2B) 61.0† (Qwen3.5-4B)
LongBench v2 (long context) 43.7 33.2 (Gemma-4-E2B-it) 47.3 (Qwen3.5-4B)
IFBench 66.3 59.0 (LFM2.5-2.6B) 73.0 (granite-4.2-3B)
IFEval 86.7 93.4 (LFM2.5-2.6B) 93.7 (granite-4.2-3B)
MMLU-Pro 70.8 65.2 (LFM2.5-2.6B) 78.0 (Qwen3.5-4B)

Scores marked † come from the official Artificial Analysis release. OpenBMB reproduced all the others internally.

Install

The card gives version floors for each backend:

Terminal window
# Python inference
pip install -U "transformers>=5.6" accelerate torch
# OpenAI-compatible servers (pick one)
pip install "vllm>=0.21"
pip install "sglang[srt]>=0.5.16"

Run it

With Transformers:

hello.py
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "openbmb/MiniCPM5-2B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
messages = [{"role": "user", "content": "Who are you? Please briefly introduce yourself."}]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
enable_thinking=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

As an OpenAI-compatible server with vLLM:

Terminal window
vllm serve openbmb/MiniCPM5-2B --port 8000
Terminal window
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "openbmb/MiniCPM5-2B",
"messages": [{"role": "user", "content": "Who are you? Please briefly introduce yourself."}],
"max_tokens": 128,
"temperature": 1.0
}'

For llama.cpp, use the GGUF build. The GGUF files are in openbmb/MiniCPM5-2B-GGUF, and the card’s command expects a local file called MiniCPM5-2B-F16.gguf, so check the repo’s file list and download that file first:

Terminal window
llama-server -m MiniCPM5-2B-F16.gguf -a MiniCPM5-2B --port 8080 -ngl 99 -c 8192 --jinja

This is the card’s example, and it is written for GPUs: -ngl 99 offloads the model’s layers to the GPU. An F16 file is also about as large as the BF16 weights, so it won’t save memory.

Tool calling with SGLang

OpenBMB recommends SGLang for tool calling. The model emits XML-style tool calls, and SGLang’s minicpm5 parser turns them into OpenAI-compatible tool_calls. Start the server with the parser turned on:

Terminal window
python -m sglang.launch_server --model-path openbmb/MiniCPM5-2B --port 30000 \
--tool-call-parser minicpm5

The request below uses a tools array in the standard OpenAI format, because the card says the server outputs OpenAI-compatible tool calls. The card doesn’t show a full tool-call request, so check the response shape yourself the first time.

tool_call.sh
curl http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "openbmb/MiniCPM5-2B",
"messages": [{"role": "user", "content": "What is the weather in Beijing right now?"}],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}],
"temperature": 1.0,
"top_p": 0.95,
"max_tokens": 1024
}'

If the call works, the response includes choices[0].message.tool_calls. You then run the function yourself. The card doesn’t explain how to pass the result back to this model. The usual OpenAI pattern is to append the result as a tool message and resend the conversation, but test that against this model’s chat template before you rely on it. The evidence for this use case is the τ²-Bench Telecom (97.1) and BFCL v4 (66.6) scores. τ³-Bench Banking is much lower at 20.8, even though that still leads the table, so don’t expect harder multi-step tool tasks to go smoothly.

Question answering over long documents

The model has 131,072 tokens of native context, so a long report can fit into a single prompt. Long context is a clear strength against 2B-class models. Against larger models the results are mixed:

  • NoLiMa: 68.1, against 17.1 for the best 2B rival and 43.5 for Qwen3.5-4B. This is the only long-context benchmark where it beats Qwen3.5-4B.
  • AA-LCR: 59.0, against 61.0 for Qwen3.5-4B.
  • LongBenchPro: 44.8, against 58.4 for Qwen3.5-4B. The best 2B rival, Gemma-4-E2B-it, is close behind at 42.2.
  • LongBench v2: 43.7, against 47.3 for Qwen3.5-4B.

Scores in the 40s on LongBench v2 and LongBenchPro mean you should check the answers. Don’t assume a long prompt will do everything a retrieval pipeline does.

long_doc_qa.py
import sys
from transformers import AutoModelForCausalLM, AutoTokenizer
MAX_CONTEXT = 131_072
MAX_NEW_TOKENS = 2048
model_id = "openbmb/MiniCPM5-2B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
doc_path, question = sys.argv[1], sys.argv[2]
with open(doc_path, encoding="utf-8") as f:
document = f.read()
messages = [{
"role": "user",
"content": f"Read the document below and answer the question.\n\n"
f"<document>\n{document}\n</document>\n\nQuestion: {question}",
}]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
enable_thinking=True,
return_dict=True,
return_tensors="pt",
)
prompt_len = inputs["input_ids"].shape[-1]
print(f"Prompt length: {prompt_len} tokens")
if prompt_len + MAX_NEW_TOKENS > MAX_CONTEXT:
sys.exit(f"Prompt too long: {prompt_len} + {MAX_NEW_TOKENS} exceeds {MAX_CONTEXT} tokens.")
inputs = inputs.to(model.device)
outputs = model.generate(**inputs, max_new_tokens=MAX_NEW_TOKENS)
print(tokenizer.decode(outputs[0][prompt_len:], skip_special_tokens=True))
Terminal window
python long_doc_qa.py annual_report.txt "What were the three biggest cost increases?"

MAX_NEW_TOKENS is set to 2048 because this is a deep-thinking model, and reasoning may come before the answer. The model itself is small, but the KV cache for a 100K-token prompt still takes a lot of memory. Check the printed prompt length before you run this on a small GPU.

Batch-classifying support tickets

To label a pile of text, serve the model with vLLM and send requests from a script. This example uses only the Python standard library and the sampling parameters the card recommends. It sends one request at a time in a plain loop, so it is simple but not fast. For real throughput, send requests concurrently (for example with a thread pool) so the server can batch them.

classify_batch.py
import json
import urllib.request
URL = "http://localhost:8000/v1/chat/completions"
LABELS = ["billing", "bug", "feature_request", "account", "other"]
tickets = [
"I was charged twice for my subscription this month.",
"The export button crashes the app on Android.",
"Could you add dark mode to the dashboard?",
"I can't reset my password, the email never arrives.",
]
def classify(text: str) -> str:
prompt = (
f"Classify this support ticket into exactly one of: {', '.join(LABELS)}.\n"
f"Reply with the label only.\n\nTicket: {text}"
)
body = json.dumps({
"model": "openbmb/MiniCPM5-2B",
"messages": [{"role": "user", "content": prompt}],
"temperature": 1.0,
"top_p": 0.95,
"min_p": 0.0,
"max_tokens": 2048,
}).encode()
req = urllib.request.Request(URL, data=body, headers={"Content-Type": "application/json"})
with urllib.request.urlopen(req) as resp:
content = json.load(resp)["choices"][0]["message"]["content"]
# Take the last line in case reasoning text comes before the label.
answer = content.strip().splitlines()[-1].strip().lower()
return answer if answer in LABELS else f"unparsed: {answer}"
for t in tickets:
print(f"{classify(t):<20} {t}")

max_tokens is set high because reasoning may come before the label. OpenBMB also ships openbmb/MiniCPM5-2B-DSpark, a draft model for speculative decoding. The card documents it for SGLang only and says it speeds up decoding without changing the target model’s output. It doesn’t give throughput figures for batch workloads.

Terminal window
python -m sglang.launch_server \
--model-path openbmb/MiniCPM5-2B \
--trust-remote-code \
--speculative-algorithm DSPARK \
--speculative-draft-model-path openbmb/MiniCPM5-2B-DSpark \
--speculative-dspark-block-size 7 \
--port 30000

This is the card’s command as written. Unlike the plain launch commands, it passes --trust-remote-code, which lets SGLang run code from the model repos. The card doesn’t say why the flag is needed here, so review what it allows before you use it. This command also starts SGLang on port 30000, not vLLM on port 8000. If you switch to it, change URL in classify_batch.py to http://localhost:30000/v1/chat/completions.

Gotchas

  • Sampling settings. The card recommends temperature=1.0, top_p=0.95, min_p=0.0. If outputs start repeating, add repetition_penalty=1.05. Support for these parameters varies by framework.
  • llama.cpp’s default min_p causes loops. The default min_p=0.05 can filter out the tokens the model needs to break out of a repetition loop. Set min_p to 0.0 explicitly in your requests.
  • The thinking output format isn’t documented. The Transformers example passes enable_thinking=True, but the card doesn’t describe how reasoning appears in the output. Look at raw responses before you write a parser.
  • Use SGLang for tool calls. The minicpm5 tool-call parser is documented for SGLang only. With other backends you may get the raw XML-style calls and have to parse them yourself.
  • Version floors. The card requires transformers>=5.6, vllm>=0.21 and sglang[srt]>=0.5.16. It doesn’t cover older versions.
  • llama.cpp context is set with -c. The card’s example uses -c 8192. Raise it if you want long-context behaviour, and budget the extra memory.
  • Default precision is BF16. torch_dtype="auto" loads the BF16 weights, about 5 GB. For a smaller footprint, use the MLX 4-bit or GPTQ 4-bit repos. The only GGUF file the card names is F16, which is about the same size as BF16.

When to pick it

Pick MiniCPM5-2B if you need a small, permissively licensed model to run on a laptop, a phone (through the LiteRT-LM build) or a cheap GPU, and your workload is math, competitive-style coding, tool calls, or questions over long documents. On those benchmarks it beats every 2B-class model in the card. On several of them (LiveCodeBench v6, AIME 2025, BFCL v4, SWE-bench Verified, NoLiMa) it also beats Qwen3.5-4B. On most long-context benchmarks other than NoLiMa, Qwen3.5-4B is still ahead.

Other options are better in these cases:

  • Strict instruction following. The results are mixed. LFM2.5-2.6B scores higher on IFEval (93.4 vs 86.7) and Multi-IF (76.8 vs 71.8), and its Multi-IF score is the best of any model in the card. MiniCPM5-2B leads it on IFBench (66.3 vs 59.0). granite-4.2-3B leads IFBench (73.0) and IFEval (93.7).
  • Broad knowledge or hard science questions. Qwen3.5-4B leads on MMLU-Pro (78.0 vs 70.8), GPQA-Diamond (77.1 vs 70.2) and SuperGPQA (52.8 vs 40.8).
  • Serious agentic coding. SWE-bench Pro is 14.4 and Terminal-Bench v2.1 is 8.6. Qwen3.5-4B scores 28.2 and 25.8 on those. Treat the model as a helper for small edits, not an autonomous engineer.
  • Languages other than English and Chinese. The card lists only those two.

The card also says the model’s answers on health, finance, law and politics are not expert-reviewed, so don’t use it for advice in those areas without your own safeguards.

Related