Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 2 min read

Kolibri 1: serving Aleph Alpha's 78B MoE with vLLM, reasoning and tool calls

Kolibri 1 is a German–English reasoning model with only 3.5B active parameters. How to serve it, control its thinking effort, and wire up tool calling.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Aleph Alpha released Kolibri 1 on October 3, 2026. It’s an Apache-2.0 mixture-of-experts reasoning model focused on German and English:

  • 78B total parameters, 3.46B active per token. It’s fast per token, but you still need memory for all 78B.
  • FP8 weights (~78 GB), so it fits on a single H200/B200 or two 80 GB GPUs.
  • Native context of 262K tokens, validated up to 1M.
  • Explicit reasoning effort levels and Hermes-style tool calling.

Why MoE matters here

With an MoE model you pay for memory as if it were a 78B model, but for compute as if it were a ~3.5B model. In practice you get high throughput and low latency per token once it’s loaded. It’s a good fit for a shared internal assistant that serves many users. It’s a poor fit for a single consumer GPU.

Serve it

Kolibri needs Aleph Alpha’s vLLM plugin:

Terminal window
pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice

For contexts beyond 262K, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'.

Query it with the OpenAI client

The recommended sampling settings are temperature=1.0, top_p=0.97 and top_k=128.

chat.py
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="Aleph-Alpha/Kolibri-1",
messages=[{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}],
temperature=1.0,
top_p=0.97,
extra_body={
"top_k": 128,
"chat_template_kwargs": {"reasoning_effort": "high", "enable_thinking": True},
},
)
print(response.choices[0].message.content)

reasoning_effort accepts low, medium, high or none. Use none (or enable_thinking: False) for snappy chat, and high for multi-step problems.

Tool calling

Pass the standard OpenAI tools schema, run the tool, send the result back, and ask again:

tools.py
import json
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
MODEL = "Aleph-Alpha/Kolibri-1"
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}},
"required": ["city"],
},
},
}]
def get_weather(city, unit="celsius"):
return {"city": city, "temperature": 18, "unit": unit, "conditions": "cloudy"}
messages = [{"role": "user", "content": "Wie ist das Wetter gerade in Heidelberg?"}]
msg = client.chat.completions.create(model=MODEL, messages=messages, tools=tools).choices[0].message
messages.append(msg)
for call in msg.tool_calls or []:
result = get_weather(**json.loads(call.function.arguments))
messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)})
final = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
print(final.choices[0].message.content)

When to pick it

Pick it when you need a self-hosted, Apache-licensed assistant for German documents, contracts or support, long-document RAG, or agent workflows inside the EU. For English-only use, compare it against other open MoE models of a similar size.

Related