Gauge Hartwell
← all write-ups

elimination-cascade, then five wrong guesses

Quorum Lost, Then Found: A Six-Week-Old Disabled Monitor, and Five Wrong Guesses Before the Real One

TL;DR: A routine command to check which hosts used systemd-timesyncd instead returned a fleet mostly unreachable. What followed was a genuine elimination cascade — SSH host keys, cloud-init timing, disk contention, Ceph version mismatches — each one a real, checkable finding, none of them the actual cause. The real problem had been sitting quietly for six weeks: a Ceph monitor, deliberately disabled at some point and never re-enabled, one failure away from taking the entire storage cluster's quorum down with it. Recovering it required correcting five separate mistakes made live, in sequence, before it finally worked.

The symptom

A single fleet-wide Ansible command — checking which hosts ran systemd-timesyncd versus chrony, entirely unrelated to storage — came back with most of the fleet UNREACHABLE, Connection refused on port 22, plus one host reporting the distinctly different No route to host.

First theory, tested and ruled out

The most recent change to the environment had been a fleet-wide Terraform apply adding an internal DNS server to every VM's cloud-init configuration, confirmed as an in-place update rather than a forced recreation. The obvious next check — had every affected VM's SSH host key actually changed, consistent with a full recreation despite what terraform plan reported — turned up a real, changed host key on one host. But it was the wrong host: the one that changed keys had never been on the unreachable list at all, and ansible.cfg's host_key_checking = False meant a changed key couldn't have caused Ansible's specific connection failures regardless. A real finding, not the cause.

Second theory, partially right

Checking the actual VMs in the Proxmox console (SSH being exactly what was failing, making it useless for diagnosis) showed the QEMU guest agent not running on every affected host. The theory: a cloud-init-triggered reboot request routed through an unresponsive guest agent could leave a VM half-rebooted — network torn down, never fully brought back — rather than cleanly succeeding or cleanly failing. Manually power-cycling the affected VMs fixed the immediate SSH problem, consistent with the theory, though the actual terraform apply log was never checked to fully confirm the mechanism.

Third theory, genuinely confirmed with real numbers

Asked whether a single spinning-disk Ceph OSD — already flagged as this build's top DR risk — could itself be the bottleneck during a fleet-wide simultaneous reboot. Checked directly:

ceph osd perf

returned 339ms apply/commit latency — more than ten times a healthy spinning disk's expected range. A live IO Pressure Stall graph showed the host sitting at 30-40% Full stall (not just Some — every non-idle process completely blocked) at the exact moment being investigated. Real, measured, serious evidence. But the timing didn't cleanly line up with the original incident window, and — critically — none of this explained why hosts stayed unreachable after being manually power-cycled. A genuine finding about this cluster's health, not the actual explanation for the SSH failures.

The detail that changed the whole investigation

Running ceph -s to check on the disk-contention theory surfaced something with no relationship to any of the above:

health: HEALTH_WARN
        1/3 mons down, quorum pve4,pve3

with out of quorum: pve2 (age 6w). This had been true for six weeks, discovered entirely by accident while checking something else. Ceph's quorum tolerates losing exactly one monitor out of three — this cluster was already at that limit, silently, the whole time.

Recovering it: five distinct mistakes, each one corrected before the next

systemctl status ceph-mon@pve2 showed disabled; preset: enabled — not a crash, a deliberate stop, consistent with an old maintenance window that was never closed out. Turning it back on immediately surfaced the real, specific error:

monitor data directory at '/var/lib/ceph/mon/ceph-pve2' does not exist:
have you run 'mkfs'?

Recovering a Ceph monitor from this state means bootstrapping a fresh data directory from the cluster's current map and keyring — not something to build by hand. What followed was five separate wrong turns, each one caught and corrected before compounding into the next:

  1. Wrong keyring path, assumed rather than confirmed — /etc/ceph/ceph.mon.keyring didn't exist. The fix wasn't guessing a second path; it was pulling the actual cluster keyring directly from a healthy monitor (ceph auth get mon. -o /tmp/mon.keyring), which is authoritative regardless of what any single host happens to have on disk.
  2. A file transfer blocked by the same bastion-style SSH restriction already established for lab VMs — direct SSH between Proxmox hosts themselves hung rather than failed, discovered only by testing a bare ssh connection after scp hung first. Routed through the one path confirmed working: the Windows management machine as a relay, same two-hop pattern already used earlier in this build for a Kubernetes CA cert.
  3. A Windows-side dead end unrelated to any of the above — scp failed with "Permission denied" writing into System32, because that's where the shell happened to be sitting, not because of anything wrong with the transfer itself.
  4. The actual recovery command run on the wrong physical host — every prompt in that session read root@pve3:~#, not pve2. The mistake created a stray, incorrectly-labeled monitor directory on pve3 while pve2 remained untouched. Caught by checking hostname explicitly before trusting any further output, and cleaned up before it could cause a second, unrelated problem.
  5. A one-character typo silently no-opping a cleanup command — rm -rf /car/lib/ceph/mon/ceph-pve2 (not /var/) ran without any error, because rm -rf on a path that doesn't exist simply succeeds, quietly. The next mkfs attempt still failed with "already exists," from the first attempt's leftover directory, since the intended cleanup had silently done nothing at all.

With the correct path finally cleared and mkfs run against the right host with the right keyring, systemctl start still reported failed — and that reported failure was, itself, a sixth misleading signal. journalctl showed nothing but Start request repeated too quickly on the two most recent attempts — no actual ceph-mon process line at all. Systemd's own restart rate-limiter, tripped by the burst of genuine failures earlier in the recovery, was silently refusing to even attempt starting the process anymore — meaning the real fix may well have already been correct, with nothing left to prove it. systemctl reset-failed cleared the lockout; the very next start succeeded cleanly, immediately confirmed by mon: 3 daemons, quorum pve2,pve4,pve3 (age 3s).

What this demonstrates

Three of the theories chased here were genuinely well-reasoned, produced real evidence, and were still wrong — a changed SSH host key, a guest-agent-stalled reboot, and severe measured disk contention were all true findings about this environment's actual state, discovered along the way to a root cause none of them were. That's not wasted effort; each one legitimately narrowed the search and surfaced separate, real issues worth having found regardless. The recovery itself is a study in a different discipline entirely: verifying every single assumption made under pressure — the keyring path, the network path, the write location, the hostname, the literal spelling of a cleanup command, and finally the meaning of a status report that was itself stale — because in a live recovery, each unverified assumption doesn't just fail cleanly, it compounds silently into the next mistake. The sharpest individual lesson is the last one: a service reporting failed doesn't always mean the fix failed. Sometimes it means the thing meant to test the fix never actually ran, and the failure being reported is an echo of a mistake that's already been corrected.