Skip to content

Kolibri Open-Weight LLM: Hands-On Guide [2026]

Kolibri is an open-weight LLM from Aleph Alpha for German and English. Just dropped: how to run the 78B MoE, set reasoning effort, and avoid the hardware trap.

5 min readIntermediate

The #1 mistake with Kolibri right now

Kolibri is an open-weight LLM from Aleph Alpha for German and English. It landed 3 October 2026. People already read “3.46B active parameters” and assume laptop-friendly or stock-vLLM plug-and-play.

That’s the trap. Full FP8 weights still need about 78 GB resident memory. Active compute is cheap; keeping the whole MoE in VRAM is not. Skip the official inference plugin or the reasoning kwargs and you get a failed load – or long thinking traces you never asked for.

Think of it like packing a huge suitcase for a weekend: you only wear a few outfits, but the bag still needs the overhead bin. Wire the Aleph Alpha stack once, then dial effort per request. That is the day-one path that works.

Quick context

Release day was German Unity Day. Weights went out Apache 2.0 on Hugging Face as Aleph-Alpha/Kolibri-1 – 78.1B total MoE, 3.46B active per token (50 layers, 384 experts + 1 shared, 6 routed). German + English only. Knowledge cutoff 18 June 2026 for both. Trained on EU iron (768× B200s in Germany/Finland) per the tech report.

No major public hosted endpoint at launch. You bring the GPUs or you wait. As of 3 October 2026, community threads (HN / LocalLLaMA, summarized in trade press) were loud about sovereignty and about whether another EU MoE can outrun Chinese serving stacks on price.

Hands-on: serve it in under an hour

Hardware floor first – from the Kolibri product page and model card: ~78 GB FP8. Realistic minimum: 2× A100 80 GB, 2× H100, or one H200 / B200 / B300. One launch-day report saw ~170 tok/s on an RTX Pro 6000 FP8. Nice if you have that card. Do not plan on a single 24 GB consumer GPU.

Install the plugin that knows the arch:

pip install "aleph-alpha-inference>=1"

That pins a supported vLLM (0.29 at launch). Stock vLLM will not load Kolibri1ForCausalLM or the kolibri1 parsers – cryptic errors or silent failure.

Serve:

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 
 --reasoning-parser kolibri1 
 --tool-call-parser kolibri1 
 --enable-auto-tool-choice

Full 1M context is validated but optional. Add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}' only if you need it. Default native window is 262144; hybrid sliding-window attention (full attention every 5th layer) makes the jump to 1M costlier and weaker on hard tasks. Stay ≤262144 for complex work.

Client call (OpenAI-compatible):

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
 model="Aleph-Alpha/Kolibri-1",
 messages=[{"role": "user", "content": "Fasse die Vorteile eines MoE für deutsche Verwaltungstexte in drei Sätzen zusammen."}],
 temperature=1.0,
 top_p=0.97,
 extra_body={
 "chat_template_kwargs": {
 "reasoning_effort": "medium",
 "enable_thinking": True
 }
 },
)
print(response.choices[0].message.content)

Card sampling defaults: temperature 1.0, top_p 0.97, top_k 128. Efforts: none / low / medium / high. Thinking ships on – set reasoning_effort to "none" or enable_thinking: false when you want a straight reply.

German legal or admin text? Keep the prompt in German. The model reasons in the user’s language; English intermediate steps often hurt on compound-heavy domains.

Tool schemas work once the kolibri1 tool-call parser is on. Multi-turn tool calling is a weaker published spot – keep chains short and validate outputs.

Will “sovereign open weights” still matter if most teams cannot clear the 78 GB bar without a mini cluster? That tension is louder than the benchmark charts.

Common pitfalls

  • Reading 3.46B active as small-model VRAM.
  • Skipping aleph-alpha-inference or mixing vLLM minors.
  • Leaving high reasoning on for simple lookups (latency tax).
  • Pushing 1M context for RAG without measuring – cap complex jobs at 256k.
  • Expecting broad multilingual trivia. German+English depth was the bet.

What the numbers show

Self-reported use numbers (highest effort, same setup) from the launch post: overall EN 75.5 / DE 70.8 among the compared ~3B-active open MoEs. Headlines inside that table: AIME 2025 EN 96.9, AIME 2026 EN 96.0, GPQA Diamond EN 84.3, LiveCodeBench 85.9. They place it on a quality-vs-serving-cost front for both languages.

It trails some denser or newer Chinese MoEs on agentic multi-turn and closed-book knowledge. Third-party benches were still catching up on launch day – directional, not gospel. German-heavy pre-training share (~21-24% tokens) plus abstention when context is missing is what actually matters for DE document RAG.

When you should not use Kolibri

Skip it on consumer GPUs under ~80 GB total VRAM. Skip it for 10+ languages, for “top arena agentic coder today,” or for zero-setup API access this week. A dense ~27B that fits your box – or a hosted frontier model – will frustrate you less.

Pick it for on-prem German+English assistants, long-doc RAG that should say “I don’t know,” controllable reasoning cost, and Apache-2.0 weights you can ship. Public sector and regulated EU shops are who they built it for.

FAQ

Is Kolibri fully open source?

No. Weights and config are Apache 2.0 on Hugging Face. Training code stays with Aleph Alpha. Run, fine-tune, ship – without the full recipe.

How do I turn thinking off for fast answers?

Capital-city style prompt? Do not burn a high-effort trace. Pass chat_template_kwargs with reasoning_effort: "none" or enable_thinking: false. The kolibri1 reasoning parser still splits fields cleanly. Medium stays a sane default until you measure your own p95 latency.

Can I fine-tune it for my company’s German docs?

Yes under Apache 2.0. Start from the instruct checkpoint and keep the bilingual tokenizer. Full-parameter fine-tunes still see the 78B footprint (or heavy sharding), so many teams will try RAG plus light adapters first – especially because abstention already helps when retrieved context is thin. Aleph Alpha also sells enterprise specialization if you want them to own the loop. Either way, prove serve + medium-effort German admin prompts on your hardware before you rewrite pipelines.

Next action: run the serve command on a machine that actually clears the VRAM floor. One German admin prompt at medium effort. One English coding prompt at high. Compare latency and answer quality yourself.