My orchestrator, the program that knows every service in my homelab and can check or restart them, checks itself too. One timer runs a health check every five minutes. Another runs a deeper wiring check every six hours, exercising the real HTTP and WebSocket paths end to end. In early September I noticed that both had last run on the thirteenth of August. Twenty-one days with no self-checking at all, and nothing had said a word.
Every glance said fine
The first thing anyone does with a suspicious timer on Linux is ask systemd whether it's enabled. Both were. systemctl is-enabled answered enabled the entire time, and that word is exactly the reassurance you're looking for.
It just doesn't mean what it sounds like. "Enabled" describes a symlink that tells systemd to start the timer at boot. It says nothing about whether the timer is running now. Ask the other question, is-active, and both said inactive. Ask when they'll next fire, and there was no answer at all.
Enabled is a statement about the next boot. It is not a statement about now.
How a restart killed them
Both timer files contained one line that looked responsible: Requires=orchestrator.service. The intent was obvious: don't run a check of the app if the app isn't there.
But Requires= doesn't only mean "needs". It also propagates stop. When the required unit stops, the requiring unit is stopped too. So the first ordinary restart of the orchestrator, the kind I do after any deploy, stopped both timers along with it.
And then nothing ever started them again. A timer that's "enabled" gets pulled in by timers.target at boot, and only at boot. The container hadn't rebooted since. The orchestrator came straight back up after the restart; its two self-checks didn't, and there was nothing to notice, because the thing that would have noticed was the self-check.
The fix is to ask less
The dependency was already in the right place. The service units that the timers start already declared that they need the orchestrator. So a check still refuses to run while the app is down, which was the whole point. The timers themselves never needed to know about the app at all.
So the fix was deleting that line from both timer files. A timer only needs to know when. What a run requires belongs on the thing that runs.
Proving it by doing the thing that broke it
After the change I restarted the orchestrator, the exact action that had killed both timers in August. This time both stayed active and armed, five minutes and six hours out. Then both checks ran and came back green. So the three-week gap hadn't hidden a real fault, but nothing except those checks could have told me that.
There was one more trap while verifying. I first checked the "next elapse" field and found it empty, and for a minute it looked like the fix had failed. These timers count from boot or from their last run, not from the wall clock, and for that kind of timer systemd reports the next run in a different field. The one I was reading is empty on a perfectly healthy timer of this type. systemctl list-timers shows the right one, and it showed both armed.
What I'd take from this
- Audit timers by when they'll next fire, never by whether they're enabled. "Enabled" was the most reassuring and least informative word on the screen.
- Put dependencies on the thing that runs, not on the thing that schedules it.
Requires=on a timer turns every restart of its target into a silent off switch. - A monitor needs a monitor, or at least a heartbeat someone else reads. A self-check that stops can't report that it stopped.
- Prove a fix by repeating the action that broke it. A check that would pass before and after the change proves nothing.
Comments