Gauge Hartwell
← all write-ups

elimination-cascade

The Textbook Shape of a Hard Infrastructure Bug: Chasing DVWA's Vanishing NodePort

TL;DR: A deployed, working DVWA instance became completely unreachable from any external client, with no single identified triggering change, while remaining perfectly reachable from the K3s node it was actually running on. The eventual cause was one mistyped word — tcp instead of udp — on a firewall rule for a port that had nothing to do with the request path anyone was looking at.

The symptom

DVWA had been working. Then, with no specific change anyone could point to, it stopped responding to every external client — browsers, curl from Windows, curl from a separate Linux control node — while curl run directly from the K3s master itself succeeded every time.

Narrowing scope before explaining cause

This took the most extensive elimination process of the build, and the order mattered as much as the individual checks:

  1. Ruled out DNS and client-side issues — tested from multiple clients and browsers; all external attempts failed identically, which argued against anything client-specific.
  2. Distinguished "refused" from "timed out." curl -v showed a timeout, not a refusal — a meaningful difference. A refusal means something is actively rejecting the connection; a timeout means packets go out into a void with no response at all, which points toward a silent drop somewhere in the path rather than an application-level decision.
  3. Checked Proxmox's firewall, at both the Datacenter and VM level — confirmed disabled. Ruled out.
  4. Verified the Kubernetes application layer directly — kubectl get pods, get svc, and critically get endpoints all showed a healthy pod correctly registered behind its Service. This ruled out the application and the Service definition entirely, not just made them look fine.
  5. The pivotal test: curl against the NodePort succeeded when run from the node actually hosting the DVWA pod, and failed with the identical timeout from every other node. This reframed the whole problem — not "NodePort is broken," but specifically "cross-node traffic to this pod is broken," a much narrower and more testable claim than where the investigation started.
  6. Cross-node pod traffic in K3s runs over Flannel's VXLAN overlay, UDP port 8472. Checking UFW directly: sudo ufw status | grep 8472 showed 8472/tcp ALLOW — TCP, not UDP.

Root cause

The firewall rule for Flannel's overlay network had been written as TCP instead of UDP — the protocol Flannel actually uses. Same-node traffic never needed the overlay at all, so it kept working the whole time and masked the problem; cross-node traffic hit the firewall and was silently dropped before it ever reached kube-proxy or the pod.

Fix

ufw_allowed_ports:
  - port: "8472"
    proto: "udp"
    comment: "Flannel VXLAN"

A second, unrelated bug found in the same investigation window

Once the VXLAN fix landed, NodePort access was still inconsistent across the cluster's freshly built nodes — a different bug, discovered only because the first one had just been fixed and the symptom hadn't fully gone away. K3s's embedded kube-proxy programs its rules through the system's iptables command, and Ubuntu 26.04 defaults that command to the nft (nftables) backend rather than the iptables-legacy semantics kube-proxy actually expects:

sudo update-alternatives --display iptables

showed iptables-nft active on all three nodes — silently producing a non-functional rule set for NodePort routing, with nothing in K3s, kube-proxy, or node status ever reporting unhealthy. Fixed per node:

sudo update-alternatives --set iptables /usr/sbin/iptables-legacy
sudo update-alternatives --set ip6tables /usr/sbin/ip6tables-legacy
sudo systemctl restart k3s        # or k3s-agent on workers

and folded into the Ansible K3s role so it's applied automatically on any future node join or rebuild, rather than something to remember by hand.

What this demonstrates

This is the textbook shape of a hard infrastructure bug: a single mismatched value, on a port number that looks unrelated to anything in the visible symptom, four logical layers away from "a web page doesn't load." The breakthrough wasn't a clever guess — it was methodically narrowing the scope of the failure (same-node works, cross-node doesn't) before trying to explain its cause, which turned an open-ended "why is this broken" into a specific, checkable claim. The second bug is worth its own note: a cluster reporting Ready on every node says nothing about whether the OS's actual networking primitives match what the orchestrator assumes they are. Both bugs here share the same real lesson — a brand-new OS release changing a quiet default that mature tooling wasn't yet written to detect.