◂ writeups // michaelz.dev

The Frame That Poisoned the Model

In Everything Was Green I wrote about the small vision model that narrates my living-room camera, and a memory fight on the graphics card that made it fail. After that was fixed, a second fault was left. Every so often the model would fail, the system would heal it, and it would go blind anyway for fifteen to thirty minutes. This is how I spent four days measuring the wrong thing, and what finally explained it.

The symptom

The narrator asks the model for a structured answer, using a grammar that forces valid JSON. Occasionally the model server would throw a deep internal error, roughly "unexpected empty grammar stack", and from then on every request failed. My self-healing restarted the model server and retried. Sometimes that worked. Often the next attempt failed too, and the camera stayed blind until a cooldown ran out.

That error message points straight at the grammar library, and for a while I believed it.

The benchmark that wasn't

I also had a daily scheduled restart of the model server, and after those restarts the fault never happened: zero out of eight. Restarts after a heal, by contrast, failed about four times in ten. That contrast looked like gold. Something about how a heal restarted the server (a readiness probe, a warm cache, timing) had to differ from the daily restart. I built theories on it for four days. I even removed the readiness probe to test one of them. The fault rate went from 43% to 44% to 42%. Nothing moved.

Here's what I eventually noticed. After the daily restart, the next request carried a fresh, random camera frame. After a heal, the retry carried the same frame that had just failed, by design. My own code had a comment explaining why: re-grabbing a new frame would change the input, so a success wouldn't prove the heal had worked.

So the two groups I'd been comparing didn't differ in the restart. They differed in the picture. My "control" was never a control.

Before you believe a comparison, list every way the two groups differ. Not the one you're interested in: all of them.

Testing the picture

The fix to the experiment was simple, and no earlier attempt had done it: restart the model server before every trial, and send a known-good control frame first, to prove the fresh process works before testing anything.

The result was not subtle. The frames that had started fault runs failed every single time: 12 out of 12, and they left the server broken 4 times out of 4. Known-good frames: 0 out of 6. Statistically, that's p = 0.0005. The fault is a property of the individual frame.

My very first attempt at this replay had, by the way, come back 36 out of 36 failures, controls included. I'd run it all against one process, the first bad frame broke it, and every row after that described the broken process rather than its own frame. I threw it away.

Not a grammar bug at all

With the grammar switched off entirely, the same frames don't throw an error. The model just answers ???????????????, the same token over and over. The model collapses on certain inputs. The grammar error is only what that collapse looks like when it's forced through a JSON shape.

The frames themselves look completely ordinary: a normal kitchen, normal light. The edge is numerical, not visual. Saved at JPEG quality 85, a poison frame fails. The same frame at quality 86 goes through cleanly. That made re-encoding look like a tempting fix, until I noticed the re-encoded versions answer differently. One saw nobody, another saw a cat, a cropped version saw a person. A fix that quietly changes the verdict is worse than the bug, so it was rejected.

Heal, then re-poison

Now the whole timeline made sense. A poison frame breaks the server. The heal restarts it, which genuinely fixes it. Then the retry sends the same poison frame and breaks it again. Over 69 hours of real data, the first heal of a fault run cured the problem 0 out of 5 times. The second heal came after a 15-minute cooldown, on a different frame, and cured it 5 out of 5.

The 15 to 30 minute blind periods were never how long the fault lasted. They were how long my own cooldown lasted.

The fix, and the test that lied

The retry now takes a fresh snapshot. About 1.4% of frames are poison, so a fault now costs one lost observation instead of half an hour of a blind camera. I watched it happen in production: poison frame, server broken, heal, fresh snapshot, success, five seconds total.

The twist is in the test suite. There was a test asserting the old behaviour, "the retry reuses the same frame", and it still passed against the new code. Its fake camera returned the same constant picture every time, so a fresh snapshot and no snapshot were indistinguishable. I rewrote it to hand out different frames and count how many were grabbed, then ran it against the old code to watch all seven new checks fail. A guard that can't fail isn't a guard. It's decoration that reads as coverage.

What I'd take from this

  • Name every difference between your two groups. Mine differed in the input, and four days of reasoning were about something else.
  • Reset state between trials, and prove the system with a control first. Otherwise one bad trial contaminates all the ones after it.
  • A retry that reuses its input can't recover from an input-caused fault. It re-triggers it on the very process you just cleaned.
  • Error messages describe where a failure surfaced, not where it started.

The underlying model bug is still out there. It's contained, not cured. But for the first time I can reproduce it on demand, which is the only real starting point for fixing it.

Comments

◂ all writeups michaelz.dev ▸