elimination-cascade
The Textbook Shape of a Hard Infrastructure Bug: Chasing DVWA's Vanishing NodePort
TL;DR: A deployed, working DVWA instance became completely unreachable from any
external client, with no single identified triggering change, while remaining
perfectly reachable from the K3s node it was actually running on. The eventual
cause was one mistyped word — tcp instead of udp — on a firewall rule for a
port that had nothing to do with the request path anyone was looking at.
The symptom
DVWA had been working. Then, with no specific change anyone could point to, it
stopped responding to every external client — browsers, curl from Windows,
curl from a separate Linux control node — while curl run directly from the K3s
master itself succeeded every time.
Narrowing scope before explaining cause
This took the most extensive elimination process of the build, and the order mattered as much as the individual checks:
- Ruled out DNS and client-side issues — tested from multiple clients and browsers; all external attempts failed identically, which argued against anything client-specific.
- Distinguished "refused" from "timed out."
curl -vshowed a timeout, not a refusal — a meaningful difference. A refusal means something is actively rejecting the connection; a timeout means packets go out into a void with no response at all, which points toward a silent drop somewhere in the path rather than an application-level decision. - Checked Proxmox's firewall, at both the Datacenter and VM level — confirmed disabled. Ruled out.
- Verified the Kubernetes application layer directly —
kubectl get pods,get svc, and criticallyget endpointsall showed a healthy pod correctly registered behind its Service. This ruled out the application and the Service definition entirely, not just made them look fine. - The pivotal test:
curlagainst the NodePort succeeded when run from the node actually hosting the DVWA pod, and failed with the identical timeout from every other node. This reframed the whole problem — not "NodePort is broken," but specifically "cross-node traffic to this pod is broken," a much narrower and more testable claim than where the investigation started. - Cross-node pod traffic in K3s runs over Flannel's VXLAN overlay, UDP port 8472.
Checking UFW directly:
sudo ufw status | grep 8472showed8472/tcp ALLOW— TCP, not UDP.
Root cause
The firewall rule for Flannel's overlay network had been written as TCP instead of UDP — the protocol Flannel actually uses. Same-node traffic never needed the overlay at all, so it kept working the whole time and masked the problem; cross-node traffic hit the firewall and was silently dropped before it ever reached kube-proxy or the pod.
Fix
ufw_allowed_ports:
- port: "8472"
proto: "udp"
comment: "Flannel VXLAN"
A second, unrelated bug found in the same investigation window
Once the VXLAN fix landed, NodePort access was still inconsistent across the
cluster's freshly built nodes — a different bug, discovered only because the first
one had just been fixed and the symptom hadn't fully gone away. K3s's embedded
kube-proxy programs its rules through the system's iptables command, and Ubuntu
26.04 defaults that command to the nft (nftables) backend rather than the
iptables-legacy semantics kube-proxy actually expects:
sudo update-alternatives --display iptables
showed iptables-nft active on all three nodes — silently producing a
non-functional rule set for NodePort routing, with nothing in K3s, kube-proxy, or
node status ever reporting unhealthy. Fixed per node:
sudo update-alternatives --set iptables /usr/sbin/iptables-legacy
sudo update-alternatives --set ip6tables /usr/sbin/ip6tables-legacy
sudo systemctl restart k3s # or k3s-agent on workers
and folded into the Ansible K3s role so it's applied automatically on any future node join or rebuild, rather than something to remember by hand.
What this demonstrates
This is the textbook shape of a hard infrastructure bug: a single mismatched value,
on a port number that looks unrelated to anything in the visible symptom, four
logical layers away from "a web page doesn't load." The breakthrough wasn't a
clever guess — it was methodically narrowing the scope of the failure (same-node
works, cross-node doesn't) before trying to explain its cause, which turned an
open-ended "why is this broken" into a specific, checkable claim. The second bug is
worth its own note: a cluster reporting Ready on every node says nothing about
whether the OS's actual networking primitives match what the orchestrator assumes
they are. Both bugs here share the same real lesson — a brand-new OS release
changing a quiet default that mature tooling wasn't yet written to detect.