Kolibri 1: serving Aleph Alpha's 78B MoE with vLLM, reasoning and tool calls
Kolibri 1 is a German–English reasoning model with only 3.5B active parameters. How to serve it, control its thinking effort, and wire up tool calling.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Aleph Alpha released Kolibri 1 on October 3, 2026. It’s an Apache-2.0 mixture-of-experts reasoning model focused on German and English:
- 78B total parameters, 3.46B active per token. It’s fast per token, but you still need memory for all 78B.
- FP8 weights (~78 GB), so it fits on a single H200/B200 or two 80 GB GPUs.
- Native context of 262K tokens, validated up to 1M.
- Explicit reasoning effort levels and Hermes-style tool calling.
Why MoE matters here
With an MoE model you pay for memory as if it were a 78B model, but for compute as if it were a ~3.5B model. In practice you get high throughput and low latency per token once it’s loaded. It’s a good fit for a shared internal assistant that serves many users. It’s a poor fit for a single consumer GPU.
Serve it
Kolibri needs Aleph Alpha’s vLLM plugin:
pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \ --reasoning-parser kolibri1 \ --tool-call-parser kolibri1 \ --enable-auto-tool-choiceFor contexts beyond 262K, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'.
Query it with the OpenAI client
The recommended sampling settings are temperature=1.0, top_p=0.97 and top_k=128.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create( model="Aleph-Alpha/Kolibri-1", messages=[{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}], temperature=1.0, top_p=0.97, extra_body={ "top_k": 128, "chat_template_kwargs": {"reasoning_effort": "high", "enable_thinking": True}, },)print(response.choices[0].message.content)reasoning_effort accepts low, medium, high or none. Use none (or enable_thinking: False) for snappy chat, and high for multi-step problems.
Tool calling
Pass the standard OpenAI tools schema, run the tool, send the result back, and ask again:
import jsonfrom openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")MODEL = "Aleph-Alpha/Kolibri-1"
tools = [{ "type": "function", "function": { "name": "get_weather", "description": "Get the current weather for a city.", "parameters": { "type": "object", "properties": {"city": {"type": "string"}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}}, "required": ["city"], }, },}]
def get_weather(city, unit="celsius"): return {"city": city, "temperature": 18, "unit": unit, "conditions": "cloudy"}
messages = [{"role": "user", "content": "Wie ist das Wetter gerade in Heidelberg?"}]msg = client.chat.completions.create(model=MODEL, messages=messages, tools=tools).choices[0].messagemessages.append(msg)
for call in msg.tool_calls or []: result = get_weather(**json.loads(call.function.arguments)) messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)})
final = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)print(final.choices[0].message.content)When to pick it
Pick it when you need a self-hosted, Apache-licensed assistant for German documents, contracts or support, long-document RAG, or agent workflows inside the EU. For English-only use, compare it against other open MoE models of a similar size.
Related
Run K2-Horizon-MoVA-36B-A4B for Agents and 512K-Token Context
IFM's open MoE model runs 4B active parameters with a 512K context window. How to serve it, call it, and use it for agents and long documents.
IFM/K2-Horizon-MoVA-36B-A4B
Self-Host MiMo-V2.6-Pro-RL, Xiaomi's 1T MoE Agent Model
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL
Run MiniCPM5-2B Locally for Tool Calls and Long Documents
MiniCPM5-2B is a 2.5B Llama-architecture model with 131K context and tool calling. Here's how to serve it and what it's good for.
openbmb/MiniCPM5-2B