◂ writeups // michaelz.dev

Two Drives in a Drawer: an encrypted, offline copy of everything

My homelab had good backups and one weakness that every homelab backup has: they all lived in the same flat as the thing they were backing up. This week that changed. Two old 1 TB laptop drives now hold a full, encrypted, offline copy of everything I'd need to rebuild. One stays plugged in; the other goes in a drawer, and later to someone else's home. Getting there crashed my backup server once, ate a freshly written chunk store, and ended with a metric that honestly said "never" for 81 seconds.

What goes on the drives

I already run Proxmox Backup Server, which takes nightly backups of every container and VM. The cold drives get a pull sync of that whole datastore, plus three things I added for this: the NAS shares that existed as a single copy on one disk, the Home Assistant backups, and a small "rebuild kit", which is the configuration, keys and git bundles of my repositories I'd need on day one after a disaster. Media is left out on purpose: it can be re-acquired, and it wouldn't fit.

Both drives are encrypted. They open with a key file on the backup server, and with a recovery passphrase in a second slot, so a drive is still readable if the server and its key file are gone. That passphrase has to live outside the flat too, or an off-site drive is just an encrypted brick.

Plugging in should be the whole procedure

The goal was that rotating a drive means: unplug one, plug in the other, done. The backup server runs as a VM, so the drive has to travel from the host to the VM, get unlocked, mount, and sync, with nobody typing anything.

My first attempt passed the USB device straight into the VM. Under the write load of the first sync, the emulated USB controller hit an internal assertion and the hypervisor killed the whole backup VM. The host's USB controller logged that it was "probably busted", and the drive's firmware hung until I physically replugged it. About three and a half minutes of outage, with nothing scheduled at the time, and one useful lesson: don't pass a USB storage device through as USB. Now the host keeps its own driver and hands the drive to the VM as an ordinary virtual disk.

The crash had a second cost. One drive's freshly created chunk store, made five seconds before the VM died, was simply gone afterwards. The filesystem hadn't flushed it. The other drive's, created eleven seconds before, survived. Recreating it was quick, but it's a good reminder of what "written" means when the machine dies right after.

Then the unlock step didn't trigger the mount. Proxmox Backup Server waits for the disk to announce a filesystem ID, and an unlocked encrypted mapping never does. A small drop-in now starts the server's own attach step right after the unlock succeeds.

A copy that only ever adds

One setting matters more than the rest. The sync is configured to never delete on the cold drive what has disappeared from the source. If the main datastore is wiped, by me, by a bug or by ransomware, the next sync can't faithfully mirror that destruction onto the drive. Old backups on the cold drives are thinned by their own, longer retention schedule instead.

A backup that mirrors deletions is a replica, and a replica faithfully copies your worst day too.

Proving it with a restore, not an exit code

A sync that says "OK" proves the sync ran. So I restored from the cold drive and compared: a certificate archive came back as 137 files with a tree hash identical to the same archive from the main datastore; a database dump decompressed cleanly and ended with its "dump completed" line; the whole rebuild kit came back 606 of 606 files identical, and the git bundles cloned and passed an integrity check.

Two things looked like failures along the way and weren't. Right after the restore, du reported a fraction of the real size, because the filesystem hadn't caught up with its own accounting yet. And verifying a git bundle needs a repository to verify it in. Both cost me a minute of worry and taught me to compare file trees and hashes rather than sizes.

The 81 seconds of "never"

A monitor exports when each drive last synced successfully, and alerts if the newest cold copy is more than eight days old or a drive hasn't been home for 45 days. The second drive's first fill took seven and a half hours, finishing at 00:12. Ten minutes later the metric still said never.

It looked like a bug and wasn't one. The monitor runs every fifteen minutes, and its last run had happened 81 seconds before the sync finished. The next run picked it up and the pending alert cleared. "Never" was the correct answer at the moment it was given. I'd much rather have a monitor that's 15 minutes late than one that guesses.

What I'd take from this

  • Off-site starts with offline. A drawer is already better than nothing, as long as the passphrase isn't only on the drive.
  • Never mirror deletions to your last line of defence. Add-only sync, separate retention.
  • Pass storage to a VM as a disk, not as a USB device. The USB path held up until the first heavy write.
  • A backup is proven by a restore you compare, by hash and file count, not by the job's exit code and not by du.
  • A late monitor is better than a guessing one. Read when it last ran before calling a reading wrong.

Comments

◂ all writeups michaelz.dev ▸