On Tuesday Andrei asked whether we should train Kolibri-1, Aleph Alpha’s English–German mixture-of-experts model, as his property agent. We advised against training. In exp_035, fine-tuning a small model on our study notes lowered its use of the retrieved facts in all three runs. The best build was an untrained model behind deterministic glue.
So exp_038 asks a narrow question. Put Kolibri into that glue as the answering model, hold everything else fixed, and see if it beats gemma4:26b, the model CasaSol already runs. German was the reason to try: many buyers on the Costa del Sol are German, and Kolibri was trained for German.
It doesn’t beat gemma:
- Lower on the main set. On the 60 sealed questions, Kolibri at 4-bit scored 0.10 lower than gemma4:26b on a 0–2 correctness scale (95 % CI −0.20 to 0.00). It was level in English (+0.03) and behind in Polish (−0.29) and Spanish (−0.17). The registered hypothesis, a gain of at least +0.20, is refuted.
- No edge in German. On 33 new German-buyer questions it was −0.03 (CI −0.15 to +0.09). Blind judges rated its German native on 13 of 36 answers, against gemma’s 30. Their notes point at English words carried over from the source text, which is English.
- More errors outside correctness. It made about four times as many unsupported claims (90 against 23). It misses four format floors and the harm guard.
- It fits a 64 GB Mac mini: 43.37 GiB at a 64k context.
- Effort medium closes the gap, 1.60 against gemma’s 1.62, but at 24 s a call, and 8 calls ran past 120 s.
The verdict, in the registered wording: “A +0.20 gain is not attainable on this set (in the A12 glue); German: not tested, interaction +0.061 [−0.182, +0.303]; ceiling-limited”. No live-bot trial. The answering model stays gemma4:26b.
There is one more finding, about our test set rather than Kolibri. gemma already sits 0.067 below the best score the set allows, so no model could have shown the +0.20 margin we registered. We registered a rule for exactly this case, and it fired.
The result in Simplified Technical English
This section gives the result in short, simple sentences. It is for readers who do not use English as their first language.
What we did
- CasaSol answers questions about Spanish property law. It finds text in our study guide and gives that text to a language model. The model writes the answer.
- We kept all parts the same. We changed only the model.
- We compared two models. The first model is Kolibri-1 at 4-bit. The second model is gemma4:26b. CasaSol uses gemma4:26b now.
- We wrote all the rules before the first answer. We did not change the rules after that.
- AI judges gave a score to each answer. The judges did not know which model wrote the answer.
What we found
- Kolibri got a lower score than gemma4:26b. The difference is 0.10 on a scale from 0 to 2.
- In English, the two models were almost equal. In Polish and in Spanish, Kolibri was worse.
- In German, Kolibri was not better than gemma4:26b.
- The judges said that the German of Kolibri was often not natural. Kolibri used English words in German sentences. The study guide is in English.
- Kolibri made more claims that had no support in the text.
- The 8-bit Kolibri was almost equal to the 4-bit Kolibri. Thus, the 4-bit compression is not the cause of the lower score.
- Kolibri fits in the memory of a Mac mini with 64 GB.
What this means
- Kolibri does not replace gemma4:26b in CasaSol.
- Our test cannot show a large gain for any model. gemma4:26b is already near the highest possible score.
What this does not tell you
- It does not tell you about Kolibri in other tasks.
- It does not tell you about Kolibri with German source text.
The setup
The glue. CasaSol answers Spanish property-law questions. exp_035 built and sealed the one answering path we can judge reproducibly. We call it A12, and it has four steps:
- BM25 retrieval over our study guide;
- a one-call scope gate, which decides if the question is in scope;
- one answer call;
- deterministic post-processing: citation stripping and a hand-off to a professional.
It is not the live bot’s production prompt or retrieval. We vendored the glue byte-identical from the CasaSol repository and never edited it. The harness injects the model backend and an additive German path (details).
The arms (table). All calls are greedy, and every arm gets identical retrieval and post-processing.
| Arm | Model | Where it ran | Role |
|---|---|---|---|
| K4 | Kolibri-1, our MLX 4-bit build, reasoning effort none | MacBook Pro, M5 Max, through exp_037’s gated port | primary |
| G26 | gemma4:26b, Q4_K_M | Mac mini, M4 Pro, Ollama 0.40.2 | the incumbent |
| K8 | Kolibri-1, 8-bit | MacBook Pro | reference: 8-bit against 4-bit |
| K4-MED | K4 with the answer call at effort medium | MacBook Pro | reference: what effort none costs |
| A12 | Qwen3-4B, untrained (exp_035’s reference build) | Mac mini, Ollama 0.33.0 | anchor: reproduces exp_035 byte for byte |
The sets (sealed by Andrei at 08:28 UTC this morning):
- v1, exp_035’s 60 sealed questions: English 31, Polish 17, Spanish 12. This is the primary set.
- The German-buyer slice: 36 new German questions (33 in scope, 3 out of scope), each with an English twin. Separate Fable instances wrote, reviewed and language-checked them. Each ran as a subagent of our Claude Code session, and each one’s file opens were audited from its own tool calls. No human native speaker checked them.
- 10 probes, checked by code only, built from the failure patterns of CasaSol’s first live users.
A German question is retrieved on its English twin, and the model gets the German question with English context. exp_038 measures a German answerer, not German retrieval. That choice matters below.
The judges. Twenty blind Fable instances (claude-fable-5-1) each scored one block of about nine rows. Within a block they scored every arm’s answers under random codes, against the gold answer and the context the model saw. Five blocks went to fresh judges a second time, as a reliability check. Every judge’s file opens were audited from its own tool calls. All 20 opened only their own block and their own output.
What a trial needed (the verdict map):
- a confirmed gain of at least +0.20 on v1;
- every format floor met in every language;
- the harm guard;
- a measured fit on a 64 GB Mac mini.
The result
| Test | Result |
|---|---|
| H1, Kolibri 4-bit − gemma4:26b on v1, n = 60 | −0.100, 95 % CI [−0.200, 0.000]: REFUTED |
| Headroom | best attainable 1.683; gemma 1.617; gap 0.067 < 0.30: the rule fires |
| H3, German, 33 in-scope rows | not tested (it runs only if H1 is confirmed); −0.030 [−0.152, +0.091] |
| German vs English twin, Kolibri − gemma | +0.061 [−0.182, +0.303] |
| Kolibri’s format floors | missed: out-of-scope declines in German and English (1 of 3 each); Spanish language 0.833; Spanish hand-off 0.909. Numeric and citation floors met; 0 fabricated citations kept |
| Harm guard | fails: one answer harmful for Kolibri and not for gemma |
| Fit on the Mac mini | 43.37 GiB at 64k, against a 46.66 GiB limit: CONFIRMED |
| Judge reliability, 199 answers judged twice | exact agreement 0.970, κ 0.959; H1 refuted under both judgments |
H1 is the only confirmatory test. Once it is refuted, the floors, the harm guard and the fit decide nothing. We report them as measured.
By arm (full table):
| gemma4:26b | Kolibri 4-bit | Kolibri 8-bit | Kolibri 4-bit, medium | Qwen3-4B (A12) | |
|---|---|---|---|---|---|
| v1 correctness (0–2) | 1.62 | 1.52 | 1.57 | 1.60 | 1.07 |
| English / Polish / Spanish | 1.55 / 1.65 / 1.75 | 1.58 / 1.35 / 1.58 | 1.55 / 1.53 / 1.67 | 1.68 / 1.41 / 1.67 | 1.06 / 0.94 / 1.25 |
| German rows / their English twins | 1.50 / 1.44 | 1.44 / 1.36 | 1.42 / 1.50 | 1.28 / 1.53 | — |
| German answers judged native | 30 of 36 | 13 of 36 | 10 of 36 | 10 of 36 | — |
| Unsupported claims | 23 | 90 | 89 | 39 | 100 (v1 only) |
| Answers stated from the model’s own weights | 2 | 10 | 9 | 3 | 1 |
| Median wall time per question | 11.4 s | 3.5 s | 5.0 s | 24.0 s | 6.9 s |
The wall times come from two machines: gemma and Qwen ran on the Mac mini, the Kolibri arms on the MacBook Pro. They are not a speed comparison.
gemma misses floors too. It declines only some out-of-scope questions in German and English, it misses the language floor in English, and it kept 2 fabricated citations. We had predicted it would hold every floor.
The ceiling we registered for
The judges follow rule X1. An answer to a question whose rule is not in the retrieved context can score at most 1, even if it is right from the model’s own memory. On v1, the rule is in context for 38 questions, and 3 questions are out of scope. A perfect answerer scores 2 on those 41 and at most 1 on the other 19: 101 / 60 = 1.683.
gemma scored 1.617. The headroom rule, registered before any answer, says that if the incumbent sits within 0.30 of the ceiling, a REFUTED H1 reads “a +0.20 gain is not attainable on this set”. It sits within 0.067.
This doesn’t rescue Kolibri. The point estimate is a loss, and the interval reaches 0.00 at best. It means v1 can no longer separate strong answerers in this glue. Much of the limit is retrieval: 19 of 60 questions get context without the rule. To separate strong models, a next test needs harder questions or better retrieval.
Kolibri’s German
Kolibri was trained for German, so we expected an edge in fluency, if not in correctness. We found the opposite: blind judges rated 13 of its 36 German answers native, against 30 for gemma.
We read the judges’ notes on those rows after the analysis ran. On most of Kolibri’s 23 non-native answers, the notes name English words and phrases dropped into German sentences, taken from the retrieved context. The context is English because our study guide is English and retrieval runs on the English twin. gemma translates them. Its 6 non-native answers are mostly single invented words. The notes stay private, because they quote the questions and answers.
Two limits apply:
- The German check is LLM-only. No human native reader checked the questions, the strings or the judgments.
- Our design feeds English context to a German question. That is how CasaSol works today, because it has only an English guide. Kolibri with German source text is untested.
Effort medium made Kolibri’s German worse, not better: its German rows scored 0.21 below their English twins, and only 10 of 36 were judged native.
Not the quantization
The obvious objection is that a quantized Kolibri should lose. Our data doesn’t point there:
- Both sides are 4-bit. gemma4:26b ran as a Q4_K_M build, Kolibri as our affine 4-bit MLX build. The two formats differ, but the bit width is the same.
- 8-bit barely moves it. Kolibri at 8-bit is +0.05 over its 4-bit build [−0.07, +0.18], and still −0.05 against gemma [−0.17, +0.07]. exp_037 had already measured Kolibri’s 4-bit loss as small: 0.43 × the Qwen peers’ median, in KL per byte.
- The loss follows language. Weighted by row count, English contributes +1.0 points to the difference, Polish −5.0 and Spanish −2.0. That is −6 over 60 questions, so the whole −0.10 comes from the Polish and Spanish rows. Kolibri was trained for English and German, and gemma for many languages.
- The sizes are close. Kolibri routes each token through about 3.5B active parameters, gemma4:26b through about 4B.
Our registered prior was a near tie, −0.10 to +0.15, and the result sits at its bottom. Nobody expected Kolibri to win in Polish or Spanish. German was the open question, and German is where it showed no edge.
Effort medium and the 8-bit build
- Effort medium. Kolibri’s vendor did not RL-train it at effort none, so K4-MED gives the answer call medium effort. It reached 1.60 against gemma’s 1.62 (−0.017 [−0.133, +0.100]). The cost was a 24 s median per question on the M5 Max, 46 s on German rows, and 8 calls over 120 s, the longest at 202.7 s. CasaSol’s probes require a reply within 120 s.
- The 8-bit build doesn’t fit 64 GB, so it was a reference only.
- Against exp_035’s 4B build, Kolibri at 4-bit is +0.45 [+0.30, +0.60]. That is a large gain over the model we use as a reproducibility anchor. It isn’t the bar, though. The bar is the model CasaSol already runs.
What we owe the reader
- The first judging pass was aborted unread (Amendment 1). The MacBook Pro took a copy of the private repository while gemma was still running on the mini, with 8 of its 60 answers written. At hand-back, our runbook copied the laptop’s whole results folder back, and the stale 8-row file overwrote gemma’s complete one. The judges scored 536 answers instead of 588. We saw it at collection, from counts alone, before reading any score. We restored the file from git, aborted the pass, rebuilt the bundles and had 20 fresh judges score everything again. The runbook now copies back only the folders the laptop produced.
- Two fixes were made without a typed amendment. After the push, the fit check could not find the 4-bit build, and the laptop’s first check refused to run without exact float32 pinned. We fixed both. Neither refusal had measured anything, and neither fix touches the frozen answer path. Each should have been a typed amendment. The result block records them as deviations.
- The result block left out two registered items, the judge drift check and word counts. An addendum adds them.
- The judge is stricter than exp_035’s. Qwen’s 60 answers are exp_035’s, byte for byte. Our judges agree exactly with exp_035’s on 85 % of them (κ 0.887). All 9 disagreements are one point lower here. Compare arms within this experiment, not absolute means across experiments.
- The limits (registered):
- Effort none is outside Kolibri’s training mix.
- The glue was tuned on Qwen and Gemma outputs in exp_035, which favours the incumbents.
- Gemma 4 rephrased part of Kolibri’s English pre-training data, so the two models’ outputs may be correlated.
- One judge family scored everything.
- The glue is not the live bot’s production path.
- The fit was measured on the mini. Kolibri’s quality was measured only on the MacBook Pro, because exp_037’s gate certified only that machine.
- What stays private: the glue’s prompts and strings, the questions and gold answers, the model outputs, the judge notes and the rights record. Every private file’s sha256 is published, and a leak check found no private text in the public files. The verdict can be re-derived from the per-row scores with the published analysis code.
- No relationship with Aleph Alpha. Aprimerose builds local-first deployments and would run Kolibri if it held up.
What we’d tell ourselves on Tuesday
- A copy-back is a write. Copy back only what the remote machine produced, never a whole folder it snapshotted earlier.
- A set the incumbent nearly saturates can only show losses. Check the headroom against the margin before sealing a test.
- The language of retrieval leaks into the answer. A German answerer fed English context writes English words into German.
- A fix after a refusal is still a change. Type it as an amendment when you make it, not in the closing block.
- Make a checklist of every registered item before writing the result. We missed two and had to add them in an addendum.
exp_038 ends here. Kolibri does not replace gemma4:26b in CasaSol. Nothing further is registered under this pre-registration.
The full record is Chronos experiment 038: exp_038_kolibri_casasol_dropin. The links above are pinned to commit ae45e4d. It holds:
- the pre-registration, Amendment 1, the result block and the addendum;
- per-row scores for every arm, the judge export with the code-to-arm map, per-arm summaries and the checks;
- the blindness audits as file lists;
- the kit: the MLX backend, the harness, the analysis and verdict code;
- the sha256 of every private file.