state reconciliation
Three Systems That Don't Reconcile: A Stale Proxmox Lock and Orphaned Ceph Images
TL;DR: Multiple unrelated terraform apply runs started failing on a storage
lock timeout, alongside separate errors about VM resources Terraform believed
didn't exist yet but Proxmox and Ceph both disagreed with. The cause wasn't any one
broken system — it was three separate systems (Terraform's state, Proxmox's
cluster filesystem locking, and Ceph's RBD image lifecycle) left disagreeing with
each other after an earlier run was interrupted without cleaning up after itself.
The symptom
cfs-lock 'storage-vmdisks' error: got lock request timeout
alongside separate errors about VM config files and Ceph RBD disk images already existing for VM IDs that Terraform's own state file had no record of at all.
Investigation
First, checked whether the lock was actually protecting something real:
ls -la /etc/pve/priv/lock/
No active operation was holding it. The lock was stale — left over from an earlier
terraform apply or destroy that had been interrupted, either by an unrelated
error partway through that same run, or an earlier manual intervention, without
ever releasing the lock it had taken.
Separately, qm list and rbd ls vmdisks on the affected nodes turned up VM
config files and Ceph RBD images for IDs Terraform's state file had never heard
of — remnants of those same interrupted runs, now existing in Proxmox and Ceph
while being completely invisible to the tool meant to be managing them.
Root cause
Terraform, Proxmox's cluster filesystem locking, and Ceph's RBD image lifecycle are three genuinely separate systems, and none of them automatically reconcile with each other after a partial failure. An interrupted operation can leave all three internally consistent on their own terms — Terraform's state file is a coherent document, Proxmox's lock is a real lock, Ceph's RBD images are real images — while disagreeing with each other about what actually exists.
Fix
# Clear the stale lock
rm /etc/pve/priv/lock/storage-vmdisks
# Remove orphaned VM configs
qm destroy <id> --purge
# Remove orphaned RBD images if destroy alone doesn't clean them up
rbd rm vmdisks/vm-<id>-cloudinit
rbd rm vmdisks/vm-<id>-disk-0
# Reconcile Terraform's own state if it references something no longer real
terraform state rm proxmox_virtual_environment_vm.<name>
Also adopted terraform apply -target=<resource> as a general practice for
recovering one resource at a time while debugging a divergence like this, rather
than re-running the full plan and risking compounding the problem across unrelated
resources that were never actually broken.
What this demonstrates
Infrastructure-as-Code tools model intent, not ground truth — Terraform's state file is a record of what Terraform believes it created and still owns, not a live reflection of what actually exists on the cluster. When a run is interrupted partway through, that belief and reality can drift apart, and closing the gap requires manually reconciling every system actually involved — the orchestrator, its state file, and the underlying storage layer — not just retrying the same command and hoping it works this time. A clean retry only helps when the divergence is transient; a stale lock and orphaned storage images are not transient, they're a new, separate problem that the original failure left behind.