Kai By Design

Every stranger who emails me is talking to my AI too

Three things are true about my setup at the same time, most days, and for a long time I only ever thought about them one at a time.

It holds things I can't afford to leak — client information, documents people trusted me with. It reads content from people I've never met, which isn't an edge case in my work, it is my work: inquiries, paperwork, attachments, all day. And it can send.

Any one of those is fine. All three at once is a different animal, and it took me embarrassingly long to see one shape instead of three features I happened to like. My AI reads my inbox. Anyone can write to my inbox. So anyone can write to my AI — and I had built exactly zero distinction between what a stranger said and what I asked for. Both arrive as text. Both get read by the same thing. To the machine they look identical, because they are.

The version that actually kept me up was the memory. I'd spent months making the system remember properly, which is the best work I've done on it. But in a system that remembers, a bad instruction doesn't misfire once and vanish. It gets written down. It becomes standing policy. The property that makes the thing valuable is the same one that turns a single bad afternoon into a permanent condition.

I went looking for how other people handle this. There's no shortage of writing about the attack — it has a name, it has taxonomies, careful papers going back years. What I couldn't find was the other half: what someone with one laptop and no security team does about it on a Tuesday, on a setup they built themselves.

So I did what I could and wrote the rules down. Where instructions are allowed to come from, what counts as data, what to do with a message that starts giving orders. Real work, and I stand behind it. But I'd also written, in the same document, the sentence that undid it: rules like these are advice to a model that is currently being manipulated. If it were reliably good at spotting the trick, there'd be no problem to solve. I documented the hole and left it there for a couple of weeks, which is its own kind of honest.

The fix, when it came, wasn't about making the reader smarter. It was about making the reader unable to do anything.

The part of my system that reads a stranger's message is now completely separate from the part that can act. It goes in a small room with the message and nothing else — no files, no ability to send, no way to touch the world. And when it comes back out, it isn't allowed to tell me what it thinks. It hands back one word off a short fixed list. Not a summary, not a recommendation, not a sentence it composed. If it can't decide, it fails to the safe answer rather than guessing.

Everything downstream — the part that files things, drafts replies, decides what reaches me — reads the word. It never reads the message.

I tested it against a set of deliberately nasty messages written to get past exactly this. The one I cared about is the oldest trick in the book, the one that empties real businesses' bank accounts every week: a friendly, plausible note saying payment details have changed, use these going forward. It came back as one word off the list. Not because it was clever enough to catch it — I don't know whether it noticed. It came back as one word because one word was the only thing it could produce. There was nowhere for the instruction to go.

That's the bit I keep coming back to. Being fooled isn't the failure. Everything that reads gets fooled eventually, people included. The failure is what the fooled thing is holding when it happens. Capability is the vulnerability. The rule I ended up with is simpler than all the writing that led to it: for any given job, break one of the three legs. If it's reading strangers, it doesn't get to act. If it's acting, it isn't reading strangers.

And it only works because of the rest. It leans on the habit of treating every clean-looking status report as a claim until it's checked — which is why I tested it instead of assuming. It leans on the memory work, which is what made the stakes obvious enough to bother. On its own it's a small piece of plumbing. It holds because of what it's sitting on.

If you've pointed any AI at your inbox, your messages, your support queue — and a lot of people have, mostly by clicking a button that promised to save time — you've already got all three legs. You just probably didn't decide to. Worth asking the question I should have asked years earlier: not is it smart enough to spot a trick, but what is it able to do the day it isn't.

kai min · kai by design · gta