My main server has no remote console. If it hangs during a reboot, nothing on the network can see it or help it, and the only fix is someone physically pressing the power button. Months ago I set up a reboot watchdog for exactly that case, and my notes said, with some confidence, that a hung shutdown would now recover by itself in about two minutes. During a routine kernel update in September, it hung, and I found out what the watchdog was actually able to do.
The routine part
A kernel update on Proxmox means a reboot, and I'd learned to expect one quirk on these machines: the first reboot after switching kernels sometimes lands on the old one, even when every setting is verified correct. A second, identical reboot takes it. The backup server did exactly that and came up fine. The main server did the same thing on its first reboot, all 25 guests came back, and I rebooted it a second time to land the new kernel.
The hang
The server's own log ends cleanly: the journal recorded itself stopping, which is the last thing a normal shutdown writes. Eleven seconds later the switch saw its network link drop. Then, for about six minutes, the machine sat with its fans running and lights on, and no network link at all. Nothing came back until I pressed the power button myself. I'd only rebooted it because I was home, and that rule is the whole reason this was a six-minute story.
Why the watchdog did nothing
The watchdog was armed. I checked afterwards: the module was loaded and the two-minute reboot timeout was set. It was configured correctly. It just couldn't help.
The watchdog I'd set up is a software watchdog, a timer that lives inside the Linux kernel. If the kernel gets stuck while shutting down, that timer can fire and force a reset. But look at the evidence. The journal stopped cleanly and the network link dropped 11 seconds later. The kernel had finished its work and handed the machine over to the firmware to restart it. Whatever hung, hung after Linux was gone, and a timer inside Linux can't fire once Linux isn't running.
A watchdog can only rescue you from failures that happen inside the thing it lives in.
The rest of my notes had been written as if it covered the whole reboot. It had never actually been seen to rescue anything. Its configuration was verified; its effect never was.
What changed
Mostly what I believe and what I write down. The notes now say plainly that every reboot of that server needs someone physically near it, and they record the sentence that used to claim it self-recovers, so nobody trusts it again. A hardware watchdog in the motherboard's chipset could, in principle, cover the gap, because it runs outside the operating system. I haven't checked whether this board has a usable one. Until I do, the honest answer is: reboot it when you're home.
After the power-cycle the server booted the new kernel straight away. Everything I could compare came back: all 25 guests, the graphics card visible to the containers that use it, every storage pool online, the public services answering. The update itself was fine. Only my safety net was imaginary.
What I'd take from this
- Ask where a safety mechanism lives, and which failures happen outside it. A software watchdog can't see past the kernel.
- "Configured" isn't "proven". I'd verified the setting several times and never seen it work.
- Write the limits into the notes, not just the feature. The dangerous sentence was the confident one: "now self-recovers".
Comments