◂ writeups // michaelz.dev

The 4 a.m. Intruder Was the Cat

At 4:52 one morning my phone received an alert that is designed to be the loudest thing my house can say: a person is at home while you are both away. Neither of us was home. I opened the recording it had saved. On the dining table, looking pleased with herself, was Pixel.

A false alarm from a cat isn't news. Anyone with a cheap camera has had one. What made this one worth writing up is what I found when I went looking for why: the system had correctly worked out that it was a cat, and then, through a perfectly reasonable chain of decisions, turned that into a sentence about a person.

Two detectors, one of them honest

The camera in my living room has its own built-in person detector. It's fast and it's free, and it's famously bad at exactly one thing: telling a cat from a human. That weakness is the entire reason I run a second opinion. When the camera fires, a small vision model on my own hardware looks at the same frame and writes one plain sentence about what it sees, with a confidence score for "is this a person?"

The alert that fired wasn't wired to the second opinion at all. It fired straight off the camera's own detector, the one component in the house already known to be unable to make this call. The one alert whose whole job is to get my attention at night was running on the weakest evidence I own.

The answer was right there

That was the first finding, and it's embarrassing but simple. The second one is the interesting one. The vision model had run on the same trigger, at the same moment. Its sentence was:

A cat is walking on the dining table.

Its person score was 0.157, which is a firm no. So far, perfect. But after a different incident a few weeks earlier, I'd added a cross-check. Before the narrator is allowed to claim an animal, a separate object detector has to agree. It scored the cat at 0.19 against a threshold of 0.25, so the claim was refused.

When a claim is refused, the narrator falls back to a cautious sentence. That fallback read: "Someone was detected, but the camera could not make them out."

Read that again with the facts next to it. The model said cat. Nothing in the pipeline thought it was a person. And the text the system produced when it was unsure was a person-shaped sentence, a stronger claim than the one it had just declined to make.

Prose is not a signal

This is the part I keep coming back to. Every piece behaved as designed. The guard refused an unconfirmed claim, which is exactly what guards are for. The fallback was written to sound careful. But "someone" is not a careful word. It's the alarming one, and it had been written as the polite thing to say when the system didn't know.

The lesson I wrote down is blunt: the sentence a system produces when it is uncertain is not a measurement. It's the uncertainty rendered as prose, and prose leans whichever way its author happened to lean. If you route alerts on it, you're routing on the author's mood. The things you can actually escalate on are the booleans and the numbers: is it a person, yes or no, and how sure.

Which direction is dangerous?

The obvious fix is to only alert when the model confirms a person. That would have been wrong in the opposite, and much worse, direction.

Think about the one case this alert exists for: someone who shouldn't be there, while we're away. Now imagine the vision model is down, or slow, or crashes on that frame. "Only alert on a confirmed person" quietly turns "I couldn't tell" into "nothing to report", on the single path that exists for an intruder. For deciding things, failing closed is usually right. Here, deciding no is the harm.

So the alert is now two tiers:

  • Tier one always fires and claims only what it knows: motion, while you're away. It saves the recording and a snapshot first, so the evidence survives even if the vision model is unavailable. It never waits for the AI.
  • Tier two is the verdict. It escalates loudly on a confirmed person. If the analysis failed, it escalates too, labelled unverified. Only an affirmative "analysed, and not a person" stays quiet, and even that writes a line in the log.

I tested it against the real 4:52 data: the cat case now stays silent, a confirmed person is loud, and a failed analysis escalates. I also ran it live, but with someone at home, so it correctly stopped before notifying anyone. I'll be honest about the gap: a real "away, and a real person" alert can't be produced without faking our presence, so the loud path is proven by replay and by structure, not by an actual intruder. I'm fine with that being the untested case.

And then I got the timeline wrong

One more confession, because it's the most useful part. While investigating, I searched the vision machine's logs between 04:00 and 06:00 for the event, found nothing, and confidently concluded that the narrator hadn't run at all.

It had. That machine logs in UTC. The event was at 02:52 in its logs, while my phone had shown 04:52 local time. One event can legitimately show up under two different clock times in two different systems, and an empty search window looks exactly like an event that never happened. I'd built a whole wrong theory on a timezone.

What I'd take from this

Three things, in order of how often they'll come up again:

  • Escalate on facts, not on wording. The fallback text is the system talking when it doesn't know. It's not evidence.
  • Decide which failure direction hurts. "Fail safe" means different things on a door lock and on an intruder alarm. Write down which one you're building before you pick a default.
  • Convert timezones before concluding nothing happened. An empty window is not an absence.

Pixel, for the record, has not been informed of any of this and continues to patrol the dining table at night.

Comments

◂ all writeups michaelz.dev ▸