Kolibri-1 matters because many practical applications need deep reasoning over very long documents while keeping inference cost reasonable for deployment. Instead of scaling a dense model, Kolibri uses a Mixture-of-Experts design to keep per-token compute low while preserving large model capacity, paired with explicit reasoning modes and structured tool-calling for safer agentic workflows.
Key Capabilities
- Long-context reasoning: trained and validated on very long sequences (native 262,144 tokens; quality validated up to 1,048,576 tokens), making it suitable for whole-document QA, policy or legal document review, and multi-document synthesis.
- Cost-efficient inference via MoE: 78B total parameters with ~3.46B active per token reduces runtime FLOPs for many tasks compared to similarly capable dense models, which lowers serving cost for interactive deployments.
- Instruction & agent-ready: post-training combined supervised fine-tuning and reinforcement learning that emphasize reasoning effort levels, tool calling, structured outputs and abstention behavior to reduce hallucinations in retrieval-augmented or tool-using settings.
- Practical deployment formats: weights published under Apache‑2.0 with FP8 weight formats and vLLM serving/plugins available for native integration into on-prem or sovereign stacks.
Who it's for and trade-offs
Great fit if you need a German/English assistant that reasons over long, domain documents (legal, public sector, aerospace, or large codebases) and you can host weights on dedicated hardware. It is designed to be integrated into systems where humans review outputs (RAG pipelines, orchestration layers calling APIs or code). Look elsewhere if you need the smallest possible memory footprint at inference (MoE requires holding full model in memory), require many-language breadth beyond DE/EN, or need turnkey cloud hosting without self-hosting effort.
Where it fits
Kolibri targets the middle ground between very large dense models and smaller efficient models: it aims to deliver high reasoning and long-context performance for German/English at lower serving FLOPs than dense counterparts, while trading higher memory requirements. It is positioned for organizations prioritizing sovereign deployment of open weights and specialized domain performance rather than maximally broad multilingual coverage.
How it works (brief)
Architecturally, Kolibri-1 is a 50-layer transformer MoE with many experts per layer (routing top‑6), RoPE positional scheme applied in sliding-window layers, and mixed precision (FP8 weights with bfloat16 for select components). Training combined large-scale pretraining, mid‑training and a long‑context phase, followed by SFT and RL focused on reasoning, tool use and long-context behavior.