We Stopped the Kolibri Benchmark Before Scoring One Answer
On Monday our own gate refused to certify our MLX port of Aleph Alpha’s Kolibri-1, and we published that instead of scores. Taking the failure apart left us two lessons: Force the routing. Kolibri sends each token to 6 of 384 experts in every layer. Two runs of our port that differed only in prefill chunk size picked different experts and drifted apart over a long context. Logits have to be compared with the expert choices forced. Update the runtime. The batched failure on our M5 Max was an MLX bug, already fixed in MLX 0.32.3. exp_037 is the experiment that post promised: the same questions and hypotheses, and most of the kit, behind a gate rebuilt on those lessons. We froze it on Wednesday, 7 October, before any gate run, with no amendment type that can change a threshold or a verdict rule. It ended this morning: ...