On Monday our own gate refused to certify our MLX port of Aleph Alpha’s Kolibri-1, and we published that instead of scores. Taking the failure apart left us two lessons:
- Force the routing. Kolibri sends each token to 6 of 384 experts in every layer. Two runs of our port that differed only in prefill chunk size picked different experts and drifted apart over a long context. Logits have to be compared with the expert choices forced.
- Update the runtime. The batched failure on our M5 Max was an MLX bug, already fixed in MLX 0.32.3.
exp_037 is the experiment that post promised: the same questions and hypotheses, and most of the kit, behind a gate rebuilt on those lessons. We froze it on Wednesday, 7 October, before any gate run, with no amendment type that can change a threshold or a verdict rule. It ended this morning:
- The port passed the gate in run 2, after one fix to a cache in our own kit. A pilot crash then forced a second fix, for a buffer leak in mlx-lm 0.32.0, and run 3, the last the rules allow, passed again. Neither fix touched the port, the reference it is checked against, a threshold or a gate rule.
- Then our plan rule said STOP. A pilot measured how long Kolibri’s answers really are, and the frozen rule turned those lengths into a schedule. The smallest plan needed 56.34 hours of machine time against a 31-hour budget. Andrei accepted the STOP.
- So there are still no Kolibri quality scores. There are four bench measurements from a gated port, of speed, memory, tokenizer efficiency and 4-bit fidelity, each beside its pre-registered threshold and without a verdict.
Two findings travel beyond Kolibri: a predictable crash in long batched generation on mlx-lm 0.32.0, for models with cache layers whose offset nothing reads, with a few-line workaround; and how much a reasoning model’s longest answers cost a local benchmark.
The result in Simplified Technical English
This section gives the result in short, simple sentences. It is for readers who do not use English as their first language.
What we did
- We made a port of the Kolibri-1 model to MLX. MLX is the machine-learning framework for Mac computers with Apple silicon.
- We wrote a reference implementation separately from the port.
- We compared the port with the reference. We call this comparison “the gate”.
- We wrote all the rules before the first gate run. We did not change these rules after that.
What occurred
- Gate run 1 failed. The cause was a defect in our own code, not in the port. We repaired the defect.
- Gate run 2 passed.
- Then the pilot run stopped with an error after 49,661 decode steps. The cause was a buffer leak in mlx-lm 0.32.0.
- We added a workaround to our runner code. We did not change the port.
- Gate run 3 passed. The rules permitted a maximum of three gate runs.
- The pilot showed that Kolibri gives very long answers. Some answers stopped at the token limit.
- Our rule used these lengths to calculate a test plan. The smallest plan needs 56.34 hours. The budget was 31 hours.
- Because of this, the rule gave the result STOP. Andrei accepted the STOP.
What we can tell you
- We have no quality scores for Kolibri.
- We have four measurements. We show each measurement next to its limit. We do not give a verdict.
| Item | Measurement | Limit |
|---|---|---|
| Decode speed of Kolibri 4-bit, as a ratio to Gemma 4 4-bit | 0.85 | at least 0.75 |
| Peak memory of Kolibri 4-bit with a 64k-token context | 43.37 GiB | at most 46.66 GiB |
| German text in each token, as a ratio to Gemma 4 and to Qwen3.6 | 1.18 and 1.17 | at least 1.15 |
| Difference between Kolibri 8-bit and 4-bit, as a ratio to the peer median | 0.43 | at most 1.5 |
What this does not tell you
- It does not tell you if Kolibri is good or bad.
- It does not tell you if Kolibri is fast or slow. The STOP is a calculation from the answer lengths, the step times and our budget.
If you use mlx-lm 0.32.0 for long batched generation
- Find the layers that use
BatchKVCacheand do not read the cache offset during decode. - If there are such layers, expect the error “Resource limit (499000) exceeded”.
- The error occurs after approximately (499,000 ÷ the number of these layers) decode steps.
- To prevent the error, evaluate the cache offsets one time at each decode step.
The gate, forced
exp_037’s gate tests the arithmetic and the routing separately:
- Forced. Our reference, a float32 numpy implementation written separately from the port, runs once with free routing and records its expert choices at every layer and position. Port and reference are then run with those choices forced, and compared tightly. Free-routing agreement is still checked, against a looser bound.
- Controls. Every new blocking check has a deliberately broken build that it must catch in the same run, on the real weights. A check that can’t see its planted bug fails the gate.
The port is exp_036’s with two changes: vLLM’s RoPE angles, decided on the specification before any forced real-weight value existed (the exp_036 post had tried them only as a diagnostic, and said they were not a fix), and two attention hooks that change no output. The port file’s sha256 was cd6153b8… in exp_036 and is 2c153357… in exp_037.
Two converted builds go through the gate, K8 (8-bit) and K4 (4-bit), and each record holds 43 checks, 36 of them blocking, plus 8 required controls. Everything on Kolibri’s real weights ran on a MacBook Pro with an M5 Max and 128 GB, under MLX 0.32.3 and mlx-lm 0.32.0. Both fixes were made on a Mac mini with an M4 Pro.
| Run | When (UTC) | Exit | K8 | K4 | Blocking checks passed | Controls caught |
|---|---|---|---|---|---|---|
| 1 | 7 Oct, 11:03–12:55 | 1 | FAIL (G1 only) | PASS | 35 of 36 | 8 of 8 |
| 2 | 8 Oct, 05:06–06:20 | 0 | PASS | PASS | 36 of 36 | 8 of 8 |
| 3 | 8 Oct, 18:06–19:21 | 0 | PASS | PASS | 36 of 36 | 8 of 8 |
Between runs only the runner, the code that drives generation, and the tests changed. The port, the reference, the gate code and both builds are byte-identical across the three records, and 41 of the 43 checks give byte-identical values.
KL, in nats, measures how far apart two next-token distributions are; 0 means identical. The reference’s bf16 emulation is the reference run with bf16 rounding wherever vLLM stores bf16: a calibrator for what correct bf16 arithmetic can reach. Run 3, from the closing block:
| Check | Threshold as registered | Run 3 |
|---|---|---|
| G4-F32: K8 in float32, routing forced, against the reference | mean KL ≤ 4.1 × 10⁻⁷ | 7.6 × 10⁻¹⁰ on eight 1,536-token texts; 5.5 × 10⁻¹¹ on a 16,384-token sequence |
| G4-F16: K8 in bf16, forced, as a multiple of the reference’s own bf16 emulation | ≤ 9 × | 1.009 × and 1.015 × |
| G4-N(i): K8 in bf16, free routing | mean KL ≤ 0.10 | 0.041 and 0.084 |
The exp_036 post called the forced comparison with the reference “still to run”. G4-F32 is that comparison: 7.6 × 10⁻¹⁰. G5-BP-lean, which decides the batch sizes Kolibri may run at, allowed every registered size, 1 to 16, for both builds.
What was loosened. Every rule we loosened had failed in exp_036, and each loosening was registered before any exp_037 run. Here is run 3 under exp_036’s rules. σ is a layer’s RMS router-logit error, so a disagreement at 6σ is an expert choice the reference made by a margin six times that error; bits per byte measures how well a model predicts a text, lower is better.
| exp_036’s rule (failed there) | exp_037’s rule | Run 3 | Under exp_036’s rule |
|---|---|---|---|
| G2: every expert-selection disagreement under 6σ | ≤ 2.474 %, and the worst pair ≤ 1.2 × the emulation’s worst, capped at 8 (here 7.32σ) | one pair at 6.08σ | fail |
| G4: top-1 agreement ≥ 99 % at confident positions (a 2-nat lead), 16k sequence | descriptive | 99.02 % (58 misses of 5,943) | pass |
| G3: our reference ≤ 1.2 bits per byte per text, and ≤ 1.25 × the best peer | descriptive | 1.3106; 1.2968 × Qwen3.6 | fail |
| G5 batch parity at batch 8, blocking | sets the allowed batch sizes; the runner at batch 1 is blocking | every size allowed, both builds | not computed |
The reference’s own bf16 emulation fails the old G2 and G4 rules as well (6.10σ in run 3; 68 misses of 5,943 in exp_036’s emulation). G3 has no such calibrator; it became descriptive by Andrei’s choice on 5 October. The flagged G2 pair is exp_036’s; its defect signs remain unexplained.
One caveat, which the pre-registration states itself: G2’s new constants, the 1.2 × and the cap of 8, were written on 5 October with exp_036’s port value, 6.08σ, in view. They rest on the emulation calibrator, not on that value, but they are not blind to it. Run 3’s worst pair sits at 6.08σ against the 7.32σ limit they produce.
In short: a mistake made identically in port and reference would pass every check here, and the one local sign that might point to one is unexplained. The blind spot, which the pre-registration requires verbatim (l.58):
Port and reference were written separately from one specification, so a misreading they share passes every comparison between them. Only the vendor’s own runtime could test that; Andrei chose on 2026-10-05 not to rent it (and had withdrawn the same anchor earlier that day, Amendment 8 D4, 08:37:33Z). exp_036 registered G3’s NLL sanity, beside the vendor’s routing test (routing only; it passed), as the local detectors of such a misreading (exp_036 HYPOTHESIS.md l.776), and G3 fired: our reference scores T1 at 1.31 bits per byte (the bound was 1.2) and 1.30 × Qwen3.6’s bits per byte on our six texts, and its secondary signs (G3-D) remain unexplained. A uniform +9.2 % NLL error would explain the first excess — about 0.17 nats per token on T3, under a third of the subtlest registered reference mutant — and no local test excludes it. The only downstream check is a tripwire on the vendor’s public scorecard. A misreading that moves Kolibri’s measured shortfall there by up to about 12–15 pp is more likely missed than caught; only beyond about 15 pp is detection near-certain. Every quality result is conditional on its absence. A tiny-checkpoint comparison of the vendor’s own model code, through vLLM 0.29.0, against our reference agreed to 1.1e-6 at every layer; it excludes a misreading in the wiring it exercises, not one in vLLM’s kernels, in FP8 serving, or in behaviour tiny random weights cannot show.
That tripwire belongs to a scored run, so it never ran. The local signal is unchanged: on run 3, G3 reads 1.3106 bits per byte on T1 and 1.2968 × Qwen3.6.
Run 1: a test that failed once
Run 1 failed K8 on G1 alone, 1 of 1,909 weight-free tests on tiny checkpoints; every comparison with the reference passed. The failing test simulates a crash partway through a cell, resumes, and asserts that the resumed records equal an uninterrupted run’s. Its failure text showed records q003 and q004 differing, q003 with equal answers; pytest’s truncation cut off the rest of q004 and hid the field that differed. The same test had passed on the same commit before the gate, and in every run on the Mac mini.
Andrei chose diagnostics over a fix or a stop, at 14:49 UTC. Before anything ran, we committed the diagnostic’s plan. It holds six experiments with exact commands and repeat counts, and a table that maps every outcome to one consequence: a gate fix, stop and publish, or a named follow-up. Two things in it mattered:
- A suspect, found by reading the code (section 2). Our scorers check that a tokenizer’s special-token ids match its family’s table, and remember every tokenizer that passed, keyed by
id()(its memory address in CPython), forever. Python can give a freed object’s address to a new one, so a test stub landing on a checked tokenizer’s address would be waved through. The plan named the field that would change and wrote the fix in advance. - A way out. If nothing reproduced it and no mechanism showed, the outcome was stop and publish.
Nothing reproduced it: not 114 pytest runs in four settings on the MacBook Pro, not 100 in-process run pairs, and not the Mac mini’s runs.
The probe aimed at the suspect hit. Earlier in G1’s order, one test checks a real Kolibri tokenizer and frees it, and the test stub has the same object layout. In a phase that kept its stubs alive, the second stub landed on the freed address and was taken as already checked. As the resumed run’s tokenizer, it gave records q003, q004 and q005 the scorers’ split label instead of the runner’s, with tokens and answers equal. That is what the gate showed.
The table classified it and named the fix written in advance. Amendment 1 makes one runner function compare a tokenizer’s ids with the family table on every call, without the cache. Real tokenizers match their tables, so no record field changes. A new test file adds six tests, one of which seeds the stale key on purpose: 5 of the 6 failed before the fix, and all 6 passed after. The probe shows our kit can produce exactly this difference; that G1’s process took this path is an inference.
Andrei gave the go for the fix and the re-run at 04:35 UTC on Thursday. Run 2 passed both builds on every check, using fix cycle 1 of 2. The defect was ours, in the test path. The fix takes the runner off the cache; the scorers’ cache itself remains, a disclosed residual risk outside the generation path.
The pilot that ran out of buffers, not memory
Run 2’s pass admitted the bench, from 06:52 to 08:22 UTC, and then the pilot: a short run of each model on its tasks, to measure answer lengths before the scored sessions are sized. At 09:17 UTC, in K8’s AIME cell, it crashed:
RuntimeError: [metal::malloc] Resource limit (499000) exceeded
The cell had admitted four problems together. Three finished, at 2,695, 3,582 and 5,478 tokens; the fourth reasoned on alone until the error, 49,661 decode steps into the cell. It was not a memory limit. We diagnosed it on tiny builds with Kolibri’s layer layout (Amendment 2):
- The line. At every decode step, mlx-lm 0.32.0’s
BatchKVCacherunsself.offset = self.offset + keys.shape[2](cache.py:929). MLX is lazy: this adds a graph node, computed only when something needs the value. - Ten layers never need it. Forty of Kolibri’s layers use a 513-token sliding window with rotary position encoding (RoPE). Every fifth layer, ten in all, attends over the whole context with no positional encoding (NoPE), through
BatchKVCache. There nothing in decode reads the offset, so the chain of pending additions grows by one link per step. - Each link holds a buffer. Every pending link keeps a live 4-byte Metal buffer. MLX caps the number of live buffers, not their bytes, at 499,000 on both our machines. Ten NoPE layers leave ten buffers per step.
- The sliding-window cache is guarded.
BatchRotatingKVCacheforces its offset withmx.depends(l.1172, l.1225).BatchKVCachehas no such line.
The arithmetic accounts for the crash. It was worked out after the crash, with the other live buffers, about 2,580, measured on tiny builds of Kolibri’s layout. The ceiling is then (499,000 − 2,580) / 10 ≈ 49,640 steps, less about 4 per sequence: about 49,624 for four. The pilot stopped after 49,661, 0.07 % later. On those tiny builds, a 50-layer model with 420,000, 440,000 and 490,000 buffers allocated in advance crashed after 7,642, 5,642 and 645 steps, exactly 10 per step, and ran 20,000 steps clean with the offsets evaluated.
The mechanism and the buffer count are confirmed; that the real K8 build behaves like a tiny build of its layout is an inference the step counts support. Our peer models aren’t exposed to this offset leak, because their cached layers read the offset for RoPE every step. We read that from mlx-lm’s code and checked it only on tiny random-initialised builds of their configurations, and that check covers the offset only. #1911, below, reports a different leak of cache state in the same loop on Qwen3.8-27B, one of our peer models, and #1731 the same error on a Qwen3.6-35B server driven through vllm-mlx. We did not test those paths. Qwen3.8 ran no pilot cell here; Gemma 4’s and Qwen3.6’s pilot cells all finished.
The gate couldn’t see it. Its generation checks are short and its comparisons cover positions up to 16,384, a limit the pre-registration disclosed.
Andrei’s decision. Offered an unchanged pilot re-run, a fix in the port instead, a runner fix without a gate run, or a stop, he chose at 14:32 UTC a gate fix followed by gate run 3, the last permitted. Unlike run 1’s, this diagnosis had no rules committed before it ran, and its probes are not committed.
The fix. A six-line helper passes every BatchKVCache offset to mx.async_eval once per step, outside the timed window. It runs the same exact integer additions earlier and changes no record field; on tiny builds, 4,000 steps took 47.9 s with it and without. Four new tests fail before it and pass after. On a 50-layer build with the buffer pool nearly full, for example, the old runner crashed after 1,242–1,247 steps, where the diagnosis had predicted about 1,240, and the new one ran 4,000.
Amendment 2 predicted run 3’s record: equal to run 2’s except G1’s test count and one check that reads the kit’s own source lines. Run 3 matched exactly. The next morning, the re-pilot’s K8 AIME cell ran 65,536 decode steps, past the old ceiling, and finished.
Upstream, checked on 9 October, outside the exp_037 record. mlx-lm’s latest release on PyPI is still 0.32.0, and cache.py on its main branch (at 9d8abd9) is byte-identical to the v0.32.0 tag, line 929 included. The hazard is known upstream, so we claim no novelty for the mechanism:
- An open pull request, ml-explore/mlx-lm#1911, followed the same error through
BatchGeneratoron Qwen3.8-27B, a hybrid model, and adds every cache’s state to the per-step evaluation. That state includesBatchKVCache’s offset, so, reading the code, it would cover Kolibri’s case; we haven’t tested it. A comment there reports a throughput cost at long context and suggests leaving the KV caches out, a variant that, by the same reading, would not cover Kolibri. - An earlier one, #1731, described the same unread-offset chain and proposed giving
BatchKVCachethe rotating cache’s guard. It was closed unmerged on 21 August by a maintainer citing review capacity.
What exp_037 adds is narrower: the offset leak inside mlx-lm 0.32.0’s own BatchGenerator, a loop #1731 had measured as bounded in August; a real model whose ten NoPE layers never read the offset; and a buffer-count ceiling that, worked out after the crash, matches its step count to 0.07 %. If you batch long generations through mlx-lm 0.32.0, count the layers whose BatchKVCache offset nothing reads in decode. If there are any, expect a ceiling of roughly MLX’s resource limit (499,000 on both our Macs), less the other live buffers, divided by their number, in decode steps per cell, less a few steps per sequence. Evaluating the offsets once per step removed it for us; our helper reads a private attribute of mlx-lm’s batch generator, so pin your version if you copy it. That covers the offset leak only; for recurrent cache state in hybrid models, see #1911.
Why we stopped
Run 3 passed on Thursday evening, and no gate run remained. On Friday morning the pilot ran again: 28 cells over four arms (two Kolibri builds, two peers), 2 h 38 min, complete, with no cell blocked and no parse failures. Answer lengths in tokens at high reasoning effort, mean and longest (summary):
| Model (build) | GPQA main EN | GPQA main DE | MMLU-ProX EN | MMLU-ProX DE | AIME EN |
|---|---|---|---|---|---|
| Kolibri (K8) | 10,115; max 32,768 (cap) | 11,657; 2 at the cap | 2,037 | 3,186 | 19,323; max 65,536 (cap) |
| Kolibri (K4) | 10,747; 1 at the cap | 12,076; 2 at the cap | 1,782 | 2,562 | — |
| Gemma 4 26B-A4B (8-bit) | 9,474; max 23,529 | 6,932; max 23,747 | 7,723 | 5,258 | — |
| Qwen3.6 35B-A3B (8-bit) | 7,075; max 16,792 | 4,351; max 13,708 | 2,505 | 2,714 | — |
The pilot is small, 8 questions per GPQA and MMLU-ProX cell and 4 AIME problems, so read it as tails, not averages:
- AIME. K8’s mean of 19.3k comes from one problem that ran to the 65,536-token cap; the other three took 2.7k–5.5k.
- GPQA. K8’s English mean is close to Gemma 4’s, but only Kolibri hit the 32,768 cap: 3 of 16 questions in each build, each still inside its reasoning.
- MMLU-ProX. Kolibri’s answers were shorter than Gemma 4’s.
- Quality. None of this is about quality. Pilot outputs are never scored.
The rule we froze on Wednesday turns these lengths into a plan, mechanically:
- Caps. A truncation, or an answer over half the cap (0.4 × for GPQA, whose pilot uses easier items), raises a task’s cap for every model: GPQA and MMLU-ProX to 65,536, AIME to 98,304. MMLU-ProX’s raise came from Gemma 4, not Kolibri.
- Batch size. A memory rule sets it. At a 65,536-token cap, K8’s GPQA batch halves from 8 to 4.
- Hours. A deterministic simulation replays every cell, with step costs fitted to the pilot’s decode steps and lengths resampled from the pilot, scaled up by one standard error (GPQA by a further 1.25 for the harder Diamond set). Answers cut off at the cap count at the raised cap, and the total is multiplied by 1.15.
- Budget. 31 hours of scored sessions (40 minus the 8.48 already spent, capped at 31), in two sessions of at most 16 hours.
- Ladder. The first of eleven plans, P0 (largest) to P10 (smallest), that fits the budget and the split is run. If none does: STOP, and no scored run unless Andrei registers a fourth session.
The plan record, in projected hours:
| P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 |
|---|---|---|---|---|---|---|---|---|---|---|
| 132.67 | 129.26 | 117.75 | 106.53 | 100.05 | 94.85 | 89.04 | 83.40 | 80.11 | 74.58 | 56.34 |
No rung fits. P10 is 25.34 hours over, and no rung splits into two sessions either: the cell at the head of every queue, K8 on all 198 English GPQA-Diamond questions, is projected at over 16 hours on its own. (The frozen projector, re-run post hoc on that cell alone, gives 16.24 h; it decides nothing.) Amendment 3, written by the rule’s code at 08:12 UTC, reads STOP.
How far off our planning was. The budget table put P10 at 9.9 hours nominal, 22.9 pessimistic and 30.3 adverse, assuming Kolibri’s GPQA answers averaged 5k, 8k or 10k tokens. The pilot measured 10.1k in English and 11.7k in German for K8, with some cut off at 32k.
Andrei accepted the STOP at 08:32 UTC this morning, without a fourth session.
The STOP is about our budget: 31 hours of one laptop, in sessions of up to 16 hours, at the sample sizes, caps and batch sizes we registered. It is a projection from Kolibri’s measured answer lengths and step times, not a speed result: it says neither that Kolibri is slow nor that it is fast. It says a scorecard with Kolibri’s long reasoning tails doesn’t fit in that box.
Four numbers without verdicts
The bench ran under run 2’s pass, run 3 re-certified the same port and builds, and neither fix touched the bench code. These are numbers from a gated port, which exp_036 couldn’t give.
Why no verdict. H1, H5 and H8 have registered one-sided tests, adjusted across the family of hypotheses; D1 has a three-way rule. Frozen code computes them, and it refuses to run without score files, which the STOP never produced. Andrei’s options were the values only, verdicts through a new amendment that switches that check off, or no claim. He chose the values, at 10:48 UTC. No rule was applied, so nothing below carries a verdict; for H1, H5 and H8, a point value beside a threshold isn’t a test either.
| What was measured | Pre-registered threshold | Measured | Predicted on 3 October | |
|---|---|---|---|---|
| H1 | Kolibri 4-bit’s decode speed at batch 1, as a multiple of Gemma 4 26B-A4B 4-bit’s | at least 0.75 | 0.8455 (110.52 tok/s; Gemma 4 130.3–131.7) | ≈ 0.85 |
| D1 | Kolibri 4-bit’s peak MLX memory with a 64k-token context, as the only model loaded | at most 46.66 GiB | 43.37 GiB, in each of 3 reps | ≈ 43.7 GiB |
| H5 | UTF-8 bytes per token on German web text, as a multiple of Gemma 4’s and Qwen3.6’s tokenizers | at least 1.15 against each | 1.1771 and 1.1685 | none registered (Aleph Alpha’s report implies 1.186 and 1.175) |
| H8 | KL(8-bit ‖ 4-bit) per byte on our six gate texts, as a multiple of the peer median | at most 1.5, and per text at most 1.5 on at least 5 of the 6 texts | 0.4320 | ≈ 1.0 |
All four were measured on the M5 Max under MLX 0.32.3 (bench files, tokenizer record, KL record). What each needs beside it:
- H1. The geometric mean over 10 paired blocks; the registered 95 % interval, recomputed, is [0.8430, 0.8479], which is not a test. Descriptively, K8 decodes at 86.13 tok/s at batch 1, against Gemma 4 8-bit’s 93.47.
- D1. Its threshold is 0.9 × the working-set limit measured on a 64 GB Mac mini. Carrying an M5 Max measurement to that chip is an assumption we registered and never tested. K8’s 77.4 GiB of weights can’t fit in 64 GB; that’s arithmetic, not a test.
- H5. Not blind: computed on the same corpus during exp_036’s build, as disclosed at registration. Kolibri packs 4.7925 bytes per token (Gemma 4 4.0714, Qwen3.6 4.1013). For that value we registered bands: “report-consistent if the estimate is in [4.80, 5.00]; card-consistent (≈ 4.7) if it is in [4.60, 4.80); neither otherwise.” We apply no label.
- H8. Gemma 4’s 4-bit build left the peer median by a registered rule (our peer check marked it speed-only), so the median is over Qwen3.6 and Qwen3.8. Per text, the ratio runs 0.257–0.651. Per token, the registered sensitivity, it is 0.4870 (per byte folds Kolibri’s denser German tokenizer into the ratio). Not blind, and the estimand is these six texts, two English and four German.
Before he chose, Andrei saw a scratch preview of the frozen verdict code run with its score-file check switched off. It printed CONFIRMED for all four. That preview is not a record, and we don’t publish it as a result.
One exploratory item, not analysed: control C1 ran K8 greedy, at effort none, at batch 8 against batch 1 on 100 MMLU-ProX questions. Six answers flipped, where we had predicted at most 2 %.
What this means if you want Kolibri on a Mac
- It runs, through our port, on a MacBook Pro with an M5 Max and 128 GB: under MLX 0.32.3, 110.52 tok/s at batch 1 for K4 and 86.13 for K8, and a 43.37 GiB peak for K4 at a 64k context. We measured nothing on a smaller machine.
- The port is public; converted weights are not. The port (
port/kolibri1.py) and the reference (reference/kolibri_ref.py) are Apache-2.0, as derived works of Aleph Alpha’s code; the rest of the kit is MIT (see NOTICE). - Energy, as a loose lower bound. powermetrics’ CPU + GPU + ANE estimate (an estimate; not the whole SoC, not wall power) was 21.74 Wh for gate run 3 (1 h 15 min) and 46.18 Wh for the re-pilot. That may be about half of the machine’s total (power records).
- Whether Kolibri is any good, we can’t tell you. Nothing was scored.
What we owe the reader
- Decisions made after results, all Andrei’s, each recorded to the second:
- 7 Oct, 14:49 UTC: diagnostics after run 1;
- 8 Oct, 04:35: the go for the cache fix and the re-run (an earlier “go”, given before he had read the classification, counts only as conditional);
- 8 Oct, 14:32: the fix and run 3 after the pilot crash, with “They stand”: if run 3 failed with K8 failing (exit 1 or 5), the four bench values would still be published. Read strictly, that would have bent the rule that no number from an ungated port is reported. Run 3 passed, so it never applied;
- 8 Oct, 17:02: “Registered rule”: on a K4 failure (exit 4), the registered consequence would apply instead (H1, H7, H8 and D1 not run; H5 stands);
- 8 Oct, 18:05: the go for gate run 3;
- 9 Oct, 08:32: accepting the STOP;
- 9 Oct, 10:48: values only for H1, D1, H5 and H8.
- The second diagnosis had no pre-committed rules, unlike the first. Its probes are not committed, and its one consequence was the second fix.
- Post hoc, deciding nothing: the re-pilot’s K8 records against the crashed pilot’s (equal except the timing fields), the 16.24-hour head-cell projection, H8 with Gemma 4 kept in the median (0.375), and the memory arithmetic that halves K8’s GPQA batch. The recipes are in the closing block.
- One re-pilot file is not public: K8’s AIME completions. Our frozen leak check, which keeps withheld benchmark text out of the repository, flagged 99 matches in the problem that ran to 65,536 tokens: runs of math notation built from one- and two-character tokens. None of it is withheld text, and AIME 2026 English outputs are publishable as registered, but the check is frozen, so the file stays out. The file’s sha256 is in the record (
5f9499a5…), and the step log and lengths the plan reads are committed. - No converted weights. The two upload criteria the gate decides hold in run 3. The scorecard criterion needs a scored run, the clean-environment check never ran, and the pilot, model-card and licence criteria were never assessed. Meeting them all would only have permitted an upload, on Andrei’s go.
- The teacher confound. Gemma 4 and Qwen3.8 are partly Kolibri’s teachers: Gemma 4 rephrased Kolibri’s English pre-training data, and Qwen3.8 generated supervised fine-tuning completions (Aleph Alpha’s report, pp. 25 and 53). Every comparison here, of speed, tokenizer efficiency, 4-bit fidelity or answer length, includes at least one of them.
- The operating envelope. Aleph Alpha targets FP8 serving in the datacentre, and we tested outside that envelope. We have no relationship with Aleph Alpha. Aprimerose builds local-first deployments and would use Kolibri if it held up.
What this post does not say is that Kolibri is good or bad. H2, H3, H4, H6 and H7 were not run (plan STOP: budget); we make no claim about them, so nothing here speaks to its scorecard, its English–German standing, its instruction following, its closed-book answers or what 4-bit costs on tasks. H1, D1, H5 and H8 carry no verdict either way. Nor does it say the port is right: “our port passed our gate” is the strongest statement we can make, and a misreading shared by port and reference would pass it too. The STOP is a projection under the caps, batch sizes and budget we registered, not a speed result.
What we’d tell ourselves on Wednesday
- Never key a cache on
id(). CPython reuses freed addresses, so test order becomes behaviour. Ours still serves the scorers, a disclosed risk. - Name the suspect before the diagnostics run. Every repetition came back clean; the probe that found the path existed because the plan named the suspect first.
- A lazy value nothing reads still holds a buffer, and MLX’s Metal limit counts buffers, not bytes.
- A short gate can’t see a long-run defect. Ours covers 16,384 positions; the leak needed about 49,600 decode steps.
- Size the budget from the tails. We registered P10 at 9.9–30.3 hours from assumed averages; the pilot projected 56.34.
- A flagged runtime change needs a check as long as the run. Our pre-pin review flagged mlx-lm 0.32.0’s cache changes, the offset update among them, for re-validation (Annex A), and the checks it named, G5-R1, G5-BP-lean and the G1 batching tests, were too short to see this. We haven’t checked whether the leak predates that change.
exp_037 ends here. Andrei declined the fourth session the rule offered, and nothing further is registered under this pre-registration. A Kolibri scorecard on one laptop would need a new pre-registration, with a budget sized from these pilot lengths; none exists yet.
The full record is Chronos experiment 037: exp_037_kolibri_forced_gate. The links above are pinned to commit 5df303c. It holds:
- the pre-registration, Amendments 1–3 and the closing block, with every decision and its UTC time;
- the three gate records, with per-layer tables and mutant records;
- the G1 diagnostics package, committed before it ran, with its outputs;
- the crashed pilot under
aborted/; - the bench records, the re-pilot’s cells and step logs, the plan record and the power records;
- the port and the reference (Apache-2.0) and the rest of the kit (MIT).
The converted weights are not published.