elimination-cascade, then five wrong guesses
Quorum Lost, Then Found: A Six-Week-Old Disabled Monitor, and Five Wrong Guesses Before the Real One
TL;DR: A routine command to check which hosts used systemd-timesyncd instead returned a fleet mostly unreachable. What followed was a genuine elimination cascade — SSH host keys, cloud-init timing, disk contention, Ceph version mismatches — each one a real, checkable finding, none of them the actual cause. The real problem had been sitting quietly for six weeks: a Ceph monitor, deliberately disabled at some point and never re-enabled, one failure away from taking the entire storage cluster's quorum down with it. Recovering it required correcting five separate mistakes made live, in sequence, before it finally worked.
The symptom
A single fleet-wide Ansible command — checking which hosts ran systemd-timesyncd
versus chrony, entirely unrelated to storage — came back with most of the fleet
UNREACHABLE, Connection refused on port 22, plus one host reporting the
distinctly different No route to host.
First theory, tested and ruled out
The most recent change to the environment had been a fleet-wide Terraform apply
adding an internal DNS server to every VM's cloud-init configuration, confirmed as
an in-place update rather than a forced recreation. The obvious next check —
had every affected VM's SSH host key actually changed, consistent with a full
recreation despite what terraform plan reported — turned up a real, changed host
key on one host. But it was the wrong host: the one that changed keys had never
been on the unreachable list at all, and ansible.cfg's host_key_checking =
False meant a changed key couldn't have caused Ansible's specific connection
failures regardless. A real finding, not the cause.
Second theory, partially right
Checking the actual VMs in the Proxmox console (SSH being exactly what was
failing, making it useless for diagnosis) showed the QEMU guest agent not running
on every affected host. The theory: a cloud-init-triggered reboot request routed
through an unresponsive guest agent could leave a VM half-rebooted — network torn
down, never fully brought back — rather than cleanly succeeding or cleanly
failing. Manually power-cycling the affected VMs fixed the immediate SSH problem,
consistent with the theory, though the actual terraform apply log was never
checked to fully confirm the mechanism.
Third theory, genuinely confirmed with real numbers
Asked whether a single spinning-disk Ceph OSD — already flagged as this build's top DR risk — could itself be the bottleneck during a fleet-wide simultaneous reboot. Checked directly:
ceph osd perf
returned 339ms apply/commit latency — more than ten times a healthy spinning disk's expected range. A live IO Pressure Stall graph showed the host sitting at 30-40% Full stall (not just Some — every non-idle process completely blocked) at the exact moment being investigated. Real, measured, serious evidence. But the timing didn't cleanly line up with the original incident window, and — critically — none of this explained why hosts stayed unreachable after being manually power-cycled. A genuine finding about this cluster's health, not the actual explanation for the SSH failures.
The detail that changed the whole investigation
Running ceph -s to check on the disk-contention theory surfaced something with
no relationship to any of the above:
health: HEALTH_WARN
1/3 mons down, quorum pve4,pve3
with out of quorum: pve2 (age 6w). This had been true for six weeks, discovered
entirely by accident while checking something else. Ceph's quorum tolerates losing
exactly one monitor out of three — this cluster was already at that limit, silently,
the whole time.
Recovering it: five distinct mistakes, each one corrected before the next
systemctl status ceph-mon@pve2 showed disabled; preset: enabled — not a crash,
a deliberate stop, consistent with an old maintenance window that was never closed
out. Turning it back on immediately surfaced the real, specific error:
monitor data directory at '/var/lib/ceph/mon/ceph-pve2' does not exist:
have you run 'mkfs'?
Recovering a Ceph monitor from this state means bootstrapping a fresh data directory from the cluster's current map and keyring — not something to build by hand. What followed was five separate wrong turns, each one caught and corrected before compounding into the next:
- Wrong keyring path, assumed rather than confirmed —
/etc/ceph/ceph.mon.keyringdidn't exist. The fix wasn't guessing a second path; it was pulling the actual cluster keyring directly from a healthy monitor (ceph auth get mon. -o /tmp/mon.keyring), which is authoritative regardless of what any single host happens to have on disk. - A file transfer blocked by the same bastion-style SSH restriction already
established for lab VMs — direct SSH between Proxmox hosts themselves hung
rather than failed, discovered only by testing a bare
sshconnection afterscphung first. Routed through the one path confirmed working: the Windows management machine as a relay, same two-hop pattern already used earlier in this build for a Kubernetes CA cert. - A Windows-side dead end unrelated to any of the above —
scpfailed with "Permission denied" writing intoSystem32, because that's where the shell happened to be sitting, not because of anything wrong with the transfer itself. - The actual recovery command run on the wrong physical host — every prompt
in that session read
root@pve3:~#, notpve2. The mistake created a stray, incorrectly-labeled monitor directory on pve3 while pve2 remained untouched. Caught by checkinghostnameexplicitly before trusting any further output, and cleaned up before it could cause a second, unrelated problem. - A one-character typo silently no-opping a cleanup command —
rm -rf /car/lib/ceph/mon/ceph-pve2(not/var/) ran without any error, becauserm -rfon a path that doesn't exist simply succeeds, quietly. The nextmkfsattempt still failed with "already exists," from the first attempt's leftover directory, since the intended cleanup had silently done nothing at all.
With the correct path finally cleared and mkfs run against the right host with
the right keyring, systemctl start still reported failed — and that reported
failure was, itself, a sixth misleading signal. journalctl showed nothing but
Start request repeated too quickly on the two most recent attempts — no actual
ceph-mon process line at all. Systemd's own restart rate-limiter, tripped by the
burst of genuine failures earlier in the recovery, was silently refusing to even
attempt starting the process anymore — meaning the real fix may well have already
been correct, with nothing left to prove it. systemctl reset-failed cleared the
lockout; the very next start succeeded cleanly, immediately confirmed by
mon: 3 daemons, quorum pve2,pve4,pve3 (age 3s).
What this demonstrates
Three of the theories chased here were genuinely well-reasoned, produced real
evidence, and were still wrong — a changed SSH host key, a guest-agent-stalled
reboot, and severe measured disk contention were all true findings about this
environment's actual state, discovered along the way to a root cause none of them
were. That's not wasted effort; each one legitimately narrowed the search and
surfaced separate, real issues worth having found regardless. The recovery itself
is a study in a different discipline entirely: verifying every single assumption
made under pressure — the keyring path, the network path, the write location, the
hostname, the literal spelling of a cleanup command, and finally the meaning of
a status report that was itself stale — because in a live recovery, each
unverified assumption doesn't just fail cleanly, it compounds silently into the
next mistake. The sharpest individual lesson is the last one: a service reporting
failed doesn't always mean the fix failed. Sometimes it means the thing meant to
test the fix never actually ran, and the failure being reported is an echo of a
mistake that's already been corrected.