Update, 5 October (evening): after publishing, we narrowed the M5 result to one MLX operation, with an exact trigger, and reported it upstream. See the update at the end of section 3.
Aleph Alpha released Kolibri-1 on Saturday, 3 October. It is a mixture-of-experts model: 78.1 billion parameters in total, 3.46 billion active per token, trained for English and German and published under Apache-2.0. That day it had no MLX build and no GGUF, and no local runtime supported it. We wanted to know three things:
- whether it fits on Apple silicon;
- how fast it runs there;
- whether its own scorecard holds on the public benchmark rows we can check.
Within a day we had an MLX port, a separate reference implementation to check it against, and a pre-registration with eight hypotheses. Above all of them sat one rule: no number from the port counts until the port passes a blocking gate against the reference. The gate ran on Monday morning, and the 8-bit build failed it.
So this post has no Kolibri scores, no speed figures and no memory figures. That is not modesty. We wrote it into the pre-registration before the run: no speed, memory or quality number from an ungated port is ever reported.
What follows is what failed, and what we found when we took it apart. Two of the findings matter to anyone running mixture-of-experts models locally:
- Free-running comparisons drift. The same port, run twice with only its prefill chunk size changed, picks different experts, and the two runs drift apart at long context.
- A batched layer on the M5 Max. On our M5 Max, mlx-lm’s batched expert layer returned wrong numbers in a probe. We can’t yet say whether MLX, macOS or the chip is at fault.
If you batch prompts through mlx-lm on an M5, the check we’d run is at the end of section 3.
The gate
The port is our own MLX implementation of Kolibri’s architecture, derived from Aleph Alpha’s open inference code. Kolibri has 50 layers.
- Routing. In each layer a router scores 384 experts for every token and sends the token to the top 6, alongside one shared expert that every token uses.
- Attention. Forty layers attend over a sliding window of 513 tokens, and mark position with rotary position encoding (RoPE). Every fifth layer attends over the whole context with no positional encoding at all.
The reference is written in numpy and runs in float32. It streams the original BF16 checkpoint one layer at a time, so it never holds the whole model. It was written separately from the port, from the same numbered specification, and imports nothing from it.
While building, we planted 29 deliberate bugs in copies of the port. They included:
- selecting experts on the wrong score;
- a window one token short;
- positional encoding on the layers that shouldn’t have it;
- the router computed in bf16.
The test suite caught all but one, logits computed in bf16. We added a test, and the suite now catches all 29.
The gate itself has six groups of checks:
| Check | What it compares |
|---|---|
| G0 | Static: tensor and parameter census, chat-template parity, tokenizer, converted configs |
| G1 | The calibrated test suite on tiny checkpoints (1,246 tests), including the 15 registered port mutants |
| G2 | Each of the 50 layers, port against reference, fed the reference’s own input |
| G3 | The reference itself: bits per byte on six plain texts, and seven planted reference bugs |
| G4 | The converted 8-bit build against the reference, end to end: eight 1,536-token texts (12,288 positions) plus a 16,384-token sequence |
| G5 | The generation path: greedy continuations, decode against prefill, batch parity, and behaviour on 20 short prompts |
The gate, and every run on Kolibri’s real weights, ran on a MacBook Pro with an M5 Max and 128 GB. Much of the bug hunt below ran on our second machine, a Mac mini with an M4 Pro.
What failed
The 8-bit build (K8) failed five of its 27 blocking checks.
The 4-bit build passed all 11 checks that apply to it. That narrower set doesn’t validate the 4-bit build, which runs the same port code. Either way, the pre-registered rule for a K8 failure is “stop”.
A few terms first:
- KL. KL, in nats, measures how far apart two next-token distributions are; 0 means identical. We always take it from the baseline side. In this post the baseline is:
- the reference, when the port is compared with it;
- the single run, when it is compared with a batched run;
- the 2,048-token chunks, when the port is compared with itself.
- Top-1 agreement. Whether both sides rank the same token first.
- Confident positions. G4 counts top-1 agreement only at positions where the reference is confident: its top token leads the runner-up by at least 2 nats, which makes it at least 7 times as likely.
| Check | Threshold as registered | Observed |
|---|---|---|
| G2, bf16 expert selection | Every top-6 disagreement with the reference must sit on a near-tie. The gap between the reference’s 6th and 7th expert scores must be under 6σ, where σ is the layer’s RMS difference between port and reference router scores. | one of 8,448 disagreements, in layer 20, at 6.08σ |
| G3, reference bits per byte, per text | each of six texts ≤ 1.2 | one text over: our exp_035 post at 1.311 |
| G3, reference against the best peer | ≤ 1.25 × the best peer on the same bytes | 1.296 × Qwen3.6 35B-A3B (8-bit) |
| G4, end to end, 16k sequence | top-1 agreement ≥ 99 % at confident positions | 98.89 %: 66 misses of 5,943, where 59 would have passed |
| G5, batch parity | 12 prompts batched 8 at a time, each against the same prompt run alone: mean KL ≤ 0.151, which is 3 × the build’s own chunking noise floor | 0.307 |
We relaxed no threshold and changed no gate code. We counted fix cycles strictly, under a rule settled after the failure (see “What we owe the reader”), and used neither of the two the pre-registration allowed.
G4’s threshold turned out to be tighter than correct bf16 arithmetic achieves here. A bf16 emulation of our own reference, running on the same 8-bit weights, misses 68 of the same 5,943 confident positions; the port misses 66. That alone would not have let the port pass, because the frozen diagnostic rules found defect signs on G4 as well.
Taking it apart
Before running any diagnostic, we froze the rules that would map its outcomes to actions. For each failing check, the rules fix three things:
- what to measure;
- which signs count as a defect;
- what each outcome permits: the failure stands, a code fix, or a threshold change that only Andrei could type.
This rule was earned the hard way. Our first diagnosis proposed thresholds fitted to the values we had just seen. An adversarial review called it outcome-shaped, and we withdrew it. Both revisions are published.
The MacBook ran the diagnostics on the same texts and positions, and reproduced every gate value exactly. The rules then left every failure standing. Their own conclusion was “next gate run: not useful”. Then we looked at why.
1. Two runs of the same code disagree, and forcing the expert choices makes them agree
The G4 diagnostics compare the port with itself. They feed the first 15,301 tokens of the 16k sequence through it twice: once in prefill chunks of 2,048 tokens, once in chunks of 64.
- Both runs use the 8-bit weights dequantised to float32, with float32 activations.
- The code is identical; only the chunk size changes.
At positions 15,000–15,300 the two runs differ by a mean KL of 0.052. The worst single position has a KL of 12.45 nats, and the top-1 token changes at 12 positions, none of them confident ones.
Then we took the expert choices recorded in the 2,048-chunk run, and forced the 64-chunk run to use them at every layer and every position. The mean KL dropped to 3.1 × 10⁻⁹, with no top-1 change.
At those positions, then, the divergence comes from which experts get picked. The cache, the masks, the positional encoding and the attention contribute nothing measurable.
Our explanation of why comes from tiny checkpoints and from Gemma 4, not from Kolibri’s own weights:
- A different chunk size changes the order of float32 additions, at the level of 1 part in 10⁷.
- Choosing 6 of 384 experts, by router score plus a per-expert selection bias, leaves near-ties, and a rounding difference flips some of them.
- A token whose experts flipped carries a slightly different hidden state, and every later token sees it through the ten full-attention layers.
Gemma 4 26B-A4B, another mixture-of-experts model, shows the same, at a smaller size. We ran it in mlx-lm on our M4 Pro, as a 4-bit community build with float32 activations, on most of the same long text. Changing only the chunk size moves its mean KL by 0.0095 over its tokens 8k–16k, and by 0.012 at its positions 15,000–15,300.
On these weights, free-running agreement at long context is a poor test of correctness. Our port does not agree with itself once only the chunk size changes. Our frozen diagnostic rules asked for much more:
- that the port agree with itself to 10⁻⁶ in KL at every position;
- that it agree with the reference to 10⁻³ in mean KL.
On Kolibri’s real weights, in float32 with free routing, neither comparison met its bound. We wouldn’t expect any implementation whose summation order changes with the chunk size to meet the first. But on the real weights we have shown it only for our own port.
What we propose instead, and will test in exp_037, has three parts:
- Force the routing.
- Compare everything else closely.
- Check the routing itself as a distribution.
So far we have done the first part, for the port against itself only. The forced comparison with the reference is still to run.
Forcing also has a small floor of its own. Our forcing passes the expert ids in sorted order, which changes the summation order of the six expert outputs. We think that is why the 2,048-chunk run, forced onto its own choices, still exceeded 10⁻⁶ at 4 of 301 positions (worst 1.7 × 10⁻⁵). It doesn’t explain why that control is noisier than the forced 64-chunk run, though, so the bound has to come from the control, not from a round number.
2. MLX and vLLM compute RoPE angles slightly differently
In RoPE, each pair of dimensions in a query or key is rotated by an angle: the token’s position times a fixed frequency for that pair. MLX’s fused RoPE computes those frequencies in a form that differs from the textbook 1/θ^(2j/D), which vLLM and our reference use, by a few units in the last place of a float32.
Against float64, beyond position 12,288, the worst angle errors are:
- MLX’s form: 1.3 × 10⁻³ radians;
- vLLM’s form: 9.7 × 10⁻⁴ radians, which is no more than float32’s own rounding of the angle.
It is a systematic difference, not a bug.
We swapped vLLM’s angles into the port and re-ran the diagnostics’ free-routing comparison with the reference. This is not the gate’s G4 comparison. Both sides run in float32 on the same dequantised 8-bit weights, so quantisation drops out, and its miss counts don’t compare with G4’s 66 of 5,943.
| Free-routing comparison | MLX angles | vLLM angles |
|---|---|---|
| 16k sequence, mean KL | 0.0153 | 0.0029 |
| 16k sequence, confident-position misses | 14 of 5,954 | 3 |
| eight 1,536-token texts, mean KL | 0.0020 | 0.0028 |
The vLLM angles help at long context and hurt slightly at short context. That fits a change in which near-ties flip, not whether they flip. It is one draw per arm, so we can’t separate the long-context gain from luck.
The comparison we predicted before the run, with routing forced, was not run. This is not a fix, and both settings still fail the frozen bounds.
3. On our M5 Max, mlx-lm’s batched expert layer returned wrong numbers
G5’s batch-parity check runs 12 prompts of different lengths through mlx-lm’s batch generator, 8 at a time. It starts with 8 and admits the other 4 as earlier ones finish. It then compares each sequence with the same prompt run alone. Kolibri’s batched outputs sat at a mean KL of 0.307 from the single runs, double the bound.
The excess is not spread evenly:
- First wave, the eight prompts prefilled together at the start: 0.469.
- Mid-run admissions, prefilled one at a time: 0.009.
In float32 the batched path is essentially exact. In bf16 it is not.
So we probed that first-wave prefill one operation at a time, with random weights at Kolibri’s dimensions: 8 rows right-padded to 1,099 tokens, 384 experts, top-6. We probed:
- the 8-bit attention and shared-expert projections;
- the float32 router;
- attention with and without the 513-token window;
- RoPE with per-row offsets;
- the expert layer.
All of them behave the same on the M5 Max and on our M4 Pro except one: mlx-lm’s SwitchGLU, the expert layer, with 8-bit weights on the sorted gather_qmm path.
The probe ran once on each machine, under MLX 0.31.2 and mlx-lm 0.31.3. It measures error against MLX’s own float32 SwitchGLU, with each row run alone on the same chip. We have not checked a newer MLX release.
| SwitchGLU, 8-bit, bf16 | M5 Max (macOS 27) | M4 Pro (macOS 26) |
|---|---|---|
| Relative error, 8 rows batched | 0.46–0.49 on every row | 0.005 |
| Relative error, the same rows run alone | 0.004–0.006 | 0.004–0.005 |
| Valid rows change when only the padding rows’ values change | yes | no |
A relative error of 46–49 % is not lost precision; it is a wrong answer. Correct arithmetic can’t make a row’s output depend on another row’s padding.
On Kolibri’s real weights, we ran one diagnostic that prefills each first-wave prompt alone. Decode stays batched, though the batch now fills one prompt at a time. It brings the first wave from 0.469 down to 0.0355.
Two things went the other way:
- One of the eight first-wave prompts got worse, from 0.17 to 0.50.
- The mid-run sequences rose from 0.009 to 0.019, almost all of it at one position.
Prefilling alone changes the shape of every operation in the prefill, so this is consistent with the probe but does not isolate the expert layer.
What we can’t say yet:
- Whether this is what failed G5. No run on the real weights changed the expert layer alone.
- The cause. The two machines differ in OS as well as chip. We can’t separate three possibilities: MLX’s M5-specific kernel selection, Apple’s Metal compiler on macOS 27, and the hardware.
- The trigger. It isn’t simply size. Gemma 4 runs more rows through the same kernel family on the same MacBook and passes its batch-parity check. Qwen3.6, on the other hand, fails batch parity there even with float32 activations, so that MacBook has more than one batched-path issue.
- The magnitude on a real model. The probe used random weights, unseeded in the expert layer, and Kolibri’s first wave degrades much less than the probe’s error would suggest.
The practical advice is narrower, and it holds already. If you batch-prefill a mixture-of-experts model through mlx-lm on an M5, or on macOS 27, compare a few batched outputs against the same prompts run one at a time before you trust the batch. If they disagree, try prefilling one prompt at a time. In our single A/B run, setting prefill_batch_size to 1 on mlx-lm 0.31.3’s BatchGenerator (8 by default), with decode still batched, removed most of Kolibri’s first-wave gap.
Update, 5 October (evening). After publishing, we cut the probe down to the single operation underneath mlx-lm’s expert layer, mx.gather_qmm. We ran it with seeded inputs on both machines.
- The faulty path. With
sorted_indices=True, the path mlx-lm takes whenever a call has 64 or more (token, expert) rows, it returns wrong values on our M5 Max (macOS 27.0, MLX 0.31.2). That happens when a call has more than 32,768 rows and the row count is not a multiple of 64. Calls of 32,769, 33,000 and 34,000 rows are wrong, while 32,768, 32,832 and 40,000 are right. - The correct paths. The same call unsorted is correct, and so is float32. Our M4 Pro is correct at every size we tried.
- Kolibri against Gemma 4. Kolibri’s batched first wave is 8 × 1,099 × 6 = 52,752 rows, over the limit. Gemma 4’s is 8 × 1,099 × 8 = 70,336, a multiple of 64, which is why it passed.
We still can’t say whether MLX, macOS 27 or the chip is responsible. It is reported upstream as ml-explore/mlx#4632, and the scripts and outputs are in BUGHUNT.md §6.9.
Until it’s fixed, on an M5 keep every sorted expert call at or under 32,768 rows. Prefilling one prompt at a time, in mlx-lm’s default 2,048-token chunks, does that for any model that sends each token to 16 experts or fewer.
What we couldn’t explain
G4 is only partly explained. Four of its defect signs fired. Sections 1 and 2 account for two of them:
- the chunk-size sign, on the real weights;
- the comparison with the reference, only by inference from tiny checkpoints and Gemma 4.
The other two were never examined:
- At 20 positions, the port’s KL from the reference exceeds 5 nats, while the bf16 emulation’s stays below 1.
- On one FineWeb document, reading it inside the long sequence rather than on its own raises the port’s KL by more than twice as much as the emulation’s.
G3 is about our reference, not the port. Bits per byte measures how well a model predicts a text; lower is better.
- Our float32 reference spends 1.31 bits per byte on our own exp_035 post, above the 1.2 we set.
- Across the six texts it sits at 1.30 × the best peer.
That needn’t mean Kolibri predicts text worse than Qwen3.6; this compares our own float32 reference with an 8-bit MLX build of the peer. There are three possibilities:
- The texts. The peer’s lead is largest on the German Basic Law. There it predicts 42 of the 66 128-byte windows at under 0.1 bits per byte, which looks like memorisation.
- A real difference between the two models on these six texts.
- A shared misreading of the specification. Port and reference were written from the same specification, so they could share one.
A shared error that raised the reference’s log-loss by 9.2 % on every text would explain the whole excess. On our German prose text that is about 0.15 nats per token, under a third of the effect of the subtlest bug we planted in the reference. And a misreading shared by the port and the reference doesn’t show in any comparison between them. So no local test we have can rule it out.
The most direct test against a shared misreading of this size is the vendor’s own runtime: Aleph Alpha’s vLLM plugin on their FP8 weights, scoring the public gate texts. The pre-registration allowed one rented cloud GPU hour for it. Andrei approved it, then withdrew the approval two minutes later (“No need to rent anything”), so it was not run.
G3 therefore stays open, and we give it low confidence. The only later check that could have exposed a large shared misreading is the comparison with Aleph Alpha’s public scorecard (H2), and it never ran.
G2 is one token. Port and reference pick different top-6 experts for 8,448 of 614,400 (token, layer) pairs. That is 1.375 %, against 2.47 % allowed, and all but one of those sit on a near-tie. The one that doesn’t is in layer 20, at a score gap of 6.08σ against a 6σ limit.
The 6σ limit is itself tighter than correct bf16 arithmetic: a bf16 emulation of our own reference has a disagreement of its own at 6.10σ. But the emulation does not flip this token, two of the frozen defect signs fired on it, and nothing we ran explains them. We are not calling it noise.
What we owe the reader
Decisions made after the failure. The gate ended at 05:57 UTC. Andrei answered at 08:35, before any diagnostic ran. He chose:
- to run the diagnostics before any re-run. They were not counted as a fix cycle, and they went beyond the per-layer reading the runbook provided for, to new measurements on the real weights;
- to count fix cycles strictly;
- to run at batch size 1 if the batched path showed no defect sign. One fired, so this never applied;
- to approve the vendor-runtime check, then withdraw the approval.
He chose to stop and publish at 14:42.
Three parts of the cycle count were not his choices:
- The early gate fixes. Not counting the two gate fixes made before the first gate run was the premise of the question he answered. Counted strictly, they would have used both cycles.
- Our reading of the rule. Reading “each amendment counts” as “each amendment that changes what the gate measures or how it judges” was ours.
- A reading not offered. A stricter reading, one cycle per threshold, was not offered to him. Under it, the gate could not have passed.
The diagnostic rules’ constants were also set with the gate values in view, though frozen before any diagnostic ran. Every decision is listed with its timestamp in the experiment record.
How “stop and publish” was put to Andrei. The question he answered already stated our bug-hunt reading, before the confirmation runs had tested it on the real weights. It said three things:
- that the bug hunt “found no port defect”;
- that the G4 bounds were “unattainable for any correct implementation of Kolibri”;
- that G5 pointed at an “M5-only” path.
Forcing the routing did remove the port’s divergence from itself. But “no port defect” overstates what we know while two defect signs on G2 remain unexplained and two on G4 were never examined. “Any correct implementation” is shown on the real weights only for the port against itself. And “M5-only” can’t be separated from the M5 machine’s newer operating system.
The upload criteria. Andrei wrote the criteria for publishing the converted weights after the gate failed. One would hold G4 to its registered bounds even if the gate’s thresholds were relaxed; another is looser on G5. Neither matters now: the gate didn’t pass, and no weights are uploaded.
The reference is ours. Without the vendor’s runtime, a specification error common to the port and the reference is invisible to every check except G3 and, for routing only, the vendor’s own routing test.
The operating envelope. Aleph Alpha targets FP8 serving in the datacentre, and we tested outside that envelope. We have no relationship with Aleph Alpha. Aprimerose builds local-first deployments and would use Kolibri if it held up.
What this post does not say is that Kolibri is weak, or that it is strong. H1–H8 and the fit check D1 were not run (the 8-bit build failed the gate, and the run stopped there); we make no claim about any of them. It doesn’t say our port is wrong, either.
It says our gate couldn’t certify the port. Part of the reason is that some of its thresholds were tighter than correct bf16 arithmetic achieves on this model: a bf16 emulation of our own reference fails G2’s and G4’s thresholds too. Our frozen diagnostic bounds were tighter still, and no float32 run with free routing on the real weights met them.
The rest is still open:
- G2’s defect signs;
- the two G4 signs the bug hunt didn’t examine;
- G3’s missing anchor;
- G5’s batched prefill on the M5.
What we’d tell ourselves on Saturday
- Force the routing before you compare logits. For a model that picks 6 experts of 384, near-exact logit agreement at long context measures routing chaos as much as correctness. On Kolibri, expert selection alone broke our port’s agreement with itself.
- Calibrate the forced comparison against itself. Forcing changes the summation order too.
- Probe batched kernels at your real shapes, on your real chip and OS. A batch-parity failure that disappears in float32 is worth a kernel probe before a code hunt.
- Freeze the diagnostic rules before the diagnostics, and have someone attack the first diagnosis. Ours was outcome-shaped, and that was only caught because we looked for it.
- Decide the anchor before the gate runs. An independent reference catches independent mistakes, but not a misreading both implementations share. Our pre-registration named that blind spot but left the anchor optional. So when G3 failed, we could not tell a shared misreading from our choice of texts. Budget for the anchor, or decide in advance what a G3 failure without it will mean.
Next is exp_037, pre-registered separately. It keeps the same hypotheses and kit, with a gate rebuilt on these lessons:
- forced-routing fidelity, calibrated by its own control;
- a batching rule for the M5;
- an explicit decision on G3.
First on the list is a minimal, seeded single-op reproduction of the SwitchGLU result that records the OS, which is what an MLX bug report needs.
The full record is Chronos experiment 036: exp_036_kolibri_local_eval. It holds:
- the pre-registration and every amendment;
- the gate record and the frozen diagnostic rules;
- the diagnostics and confirmation runs on the MacBook;
- the bug hunt with its corrections (
diagnostics/gate1/BUGHUNT.md); - the nulls, scripts and both diagnosis revisions from our M4 Pro (
diagnostics/gate1/mini/).
The port and the reference are in the same folder under Apache-2.0. The converted weights are not published: the gate didn’t pass.