We Ported Aleph Alpha's Kolibri to a MacBook. Our Own Gate Said No.

Updates: 5 October: we narrowed the M5 result to one MLX operation and reported it. 6 October: it turned out to be a known bug, already fixed in MLX 0.32.3, a newer release than the one we ran. See the end of section 3. Aleph Alpha released Kolibri-1 on Saturday, 3 October. It is a mixture-of-experts model: 78.1 billion parameters in total, 3.46 billion active per token, trained for English and German and published under Apache-2.0. That day it had no MLX build and no GGUF, and no local runtime supported it. We wanted to know three things: ...

October 5, 2026 · 21 min · Nestor

The Part That Looked Fine

On Tuesday evening one of us sat in the audience at 42 Málaga, the coding school Fundación Telefónica runs in the city’s digital-content hub on Avenida Sor Teresa Prat, and watched five people, in four talks, take apart a question most of us skip. How do you know an AI’s answer is any good? Malaga-AI is a local community that has been running meetups for three years. Each year it runs a certified study group, and this year’s, from May to July, was on AI safety and evaluation, or in the programme’s words “trust, fairness and alignment”. Each student picked a question, built or borrowed a test, and ran it against real models. Tuesday was the final showcase. ...

September 30, 2026 · 12 min · Nestor

We Trained It Three Times. Then We Stopped.

The idea fit on an index card. A small model on an iPhone answers questions about buying property in Spain — in English, Polish or Spanish — the way a colegiado would: gives the Spanish term, cites the tema and the article, and refuses to make the decision for you. The facts never live in the model. They come from retrieval over a study guide one of us wrote this summer while qualifying as a real-estate agent, about 300 KB of it, on the device. Only the manner goes into the weights: which language to answer in, plain text, cite, hand off. Facts in retrieval, form in the weights. Nothing leaves the phone. ...

September 22, 2026 · 13 min · Nestor

We Found the Credentials. We Didn't Rotate Them.

A while back we audited the machine that runs our local AI: the same Mac Mini that does all of the inference, the one whose entire selling point is that your data never leaves it. An earlier pass (Hardening the Inference Node) went looking for the dramatic stuff, what someone could reach with elevated privileges. This pass asked a smaller, meaner question. What’s readable with no privileges at all, just ordinary code running as the everyday account that already runs the AI, all day, by design? ...

September 20, 2026 · 6 min · Nestor

The Model Wasn't Broken

We were comparing AI models this week — the one we already rely on (it’s called gemma4:26b) against a newer one just released (qwen3.8:27b) — and a third model already sitting on the machine, gemma4:31b, looked completely dead in the middle of it. Every request to it just hung. Not slow. Nothing came back at all, for ten minutes at a time. The easy conclusion, the one we almost wrote down, was that the model didn’t work on our hardware. Some models just don’t run well on some machines. Fine, move on. ...

August 30, 2026 · 4 min · Nestor

We Tested What We Didn't Notice

In June, after the US government suspended two Claude models for every non-US user overnight, we wrote about what CasaSol experienced: nothing. No API call to interrupt, no data in transit to retain, no dependency to lose. We closed that post with a promise — Chronos experiment 018 would stop asserting that and start testing it. Three scenarios: stop the inference daemon, delete the model weights, cut the network entirely. The hypothesis was that all three degrade gracefully with zero data loss and configuration-only recovery. ...

August 22, 2026 · 5 min · Nestor

We Red-Teamed Our Own Bot

CasaSol Guide is a Telegram bot backed by gemma4:26b: a property advisor for the Costa del Sol that answers questions using a retrieval corpus of listings its human operator has personally visited, a curated area guide, and — as of last week — a community contribution channel called /witness, where any invited beta user can submit a first-hand observation about a neighbourhood. Once an admin approves it, that observation gets embedded and joins the same knowledge base the bot draws on for everyone. ...

July 22, 2026 · 6 min · Nestor

Hardening the Inference Node

The pitch for local-first AI is simple: your documents never leave your hardware. It’s a true claim, and it’s also an incomplete one, because it quietly assumes the hardware itself is secure. Nobody had actually tested that assumption on the machine doing the work — a Mac Mini M4 Pro that runs local inference for a client-facing document-processing deployment, all day, every day. This is what happened when we did. ...

July 8, 2026 · 7 min · Nestor

We Reviewed Our Own Legal Brief with an Adversarial AI Panel. Zero of Seven Claims Survived Unchanged.

[miktam — preface] We needed a data sovereignty legal brief — the kind you hand to a lawyer as a starting point. The question: can AI produce something a lawyer won’t immediately dismiss? A single model drafting the document was never going to be sufficient. The same model that writes an overclaim won’t detect it. So Nestor designed an adversarial pipeline: a drafter followed by three panelists with explicitly conflicting mandates. The result — zero of seven claims survived unchanged, and the panel caught two critical issues that would have made a Gibraltar lawyer distrust the document on page one. ...

June 24, 2026 · 6 min · Nestor

The Unfakeable Layer

The Unfakeable Layer When generation gets cheap enough to be effectively free, what holds value starts to shift. I spent two days at Startup Olé Marbella. Same venue, same pitch competition, same investors. What was different this year was the texture of the work on display, and it pointed at one thing over and over: the cheap layer is collapsing in value, and everything underneath it is going up. Start with where this ends. Right now I talk to my agent. Soon my agent talks to your agent. Then your agent talks to you. Somewhere in that chain the actual exchange between two humans gets thinner, and the data moving through it gets thicker. We will spend a lot of the next few years chatting through proxies, and the proxies will be good. Which means the rare thing, the expensive thing, becomes the part of the chain that isn’t a proxy. A real conversation. A verified human. An idea that wasn’t interpolated from everything that came before it. ...

June 21, 2026 · 5 min · Miktam