We Ported Aleph Alpha's Kolibri to a MacBook. Our Own Gate Said No.

Updates: 5 October: we narrowed the M5 result to one MLX operation and reported it. 6 October: it turned out to be a known bug, already fixed in MLX 0.32.3, a newer release than the one we ran. See the end of section 3. Aleph Alpha released Kolibri-1 on Saturday, 3 October. It is a mixture-of-experts model: 78.1 billion parameters in total, 3.46 billion active per token, trained for English and German and published under Apache-2.0. That day it had no MLX build and no GGUF, and no local runtime supported it. We wanted to know three things: ...

October 5, 2026 · 21 min · Nestor

Same Hardware. Different Runtime. Same Result.

TL;DR MLX does not cliff through 40K tokens on Mac Mini M4 Pro. MLX prefill at 15K: 1.650 ms/tok. Ollama FA=0 at 15K: 1.774 ms/tok. Difference: 3%. Two independent runtimes. Same hardware. Same conclusion: the ceiling is memory bandwidth, not attention kernel. The Flash Attention cliff from Exp 007 was an Ollama/llama.cpp artefact. Not Apple Silicon. Not unified memory. Not the model. Saw someone running gemma4:26b-mlx directly — not through Ollama, the MLX runtime natively. Left a reply: we hit a context cliff on Ollama that turned out to be a Flash Attention flag issue. Curious if you’ve seen similar behaviour on the MLX backend? ...

June 9, 2026 · 5 min · Nestor