Developer 3s Key Decision Metrics
On 3 Oct 2026, Aleph Alpha released open-weight Kolibri-1 (HF, Apache-2.0): 78B MoE / ~3.46B active, bilingual DE/EN, 262k native context (1M extrapolated), reasoning effort + tools. Official card: LiveCodeBench v6 85.9, SWE-Bench Verified 66.4, TerminalBench 2.1 27.7. Serve via aleph-alpha-inference + vLLM.
Key Takeaways
- ✓Shipped 2026-10-03; weights Aleph-Alpha/Kolibri-1 Apache-2.0; blog on aleph-alpha.com
- ✓78B total / ~3.46B active; ~78GB FP8; native 262k context (validated to 1M)
- ✓Official card: LiveCodeBench v6 85.9 · SWE-Verified 66.4 · TerminalBench 2.1 27.7 · HumanEval+ 92.7
- ✓Serve: aleph-alpha-inference +
vllm serve Aleph-Alpha/Kolibri-1with kolibri1 parsers - ✓Sovereign DE/EN on-prem focus — not a Terminal-Bench leader vs denser peers
Not enough VRAM? Compare cloud API and self-hosting costs
Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Why it matters
Aleph Alpha released Kolibri / Kolibri-1 on 3 Oct 2026 (German Unity Day): an open-weight bilingual (DE/EN) MoE under Apache-2.0 on Hugging Face. 78B total / ~3.46B active, built for regulated on-prem and agentic tool use—not a closed demo.
Architecture
50-layer MoE (384 experts, top-6 + shared), FP8 footprint ~78 GB. Native context 262k with validated extrapolation to 1M; mostly 512-token sliding-window attention. Reasoning effort none|low|medium|high plus Hermes-style tools via aleph-alpha-inference + vLLM (vllm serve Aleph-Alpha/Kolibri-1).
Benchmarks (official card only)
LiveCodeBench v6 85.9, SWE-Bench Verified 66.4, TerminalBench 2.1 27.7, HumanEval+ 92.7, AIME 2026 96.0, GPQA Diamond 84.3, Overall EN/DE 75.5/70.8. Strong DE/EN + grounding story; Terminal-Bench is not leadership vs denser peers. Full tables: tech report.
Get started
Weights on HF; pip install 'aleph-alpha-inference>=1' then the documented vLLM serve flags. Sampling: T=1.0, top_p=0.97, top_k=128. Blog: https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.