Gauge Hartwell
← all write-ups

elimination-cascade

Ten Layers Deep: Root-Causing a Silent VXLAN Checksum Bug Between K3s Nodes

TL;DR: A routine task — add kube-state-metrics and wire Prometheus into Grafana — turned into a multi-hour, ten-layer investigation that ended at a confirmed upstream QEMU/Proxmox regression silently corrupting VXLAN checksums between virtual machines. The eventual fix wasn't a patch at all: switching Flannel's backend from VXLAN encapsulation to direct routing (host-gw) sidestepped the entire bug class and turned out to be architecturally the right call regardless of the bug.

The setup

The K3s cluster in this lab spans a mix of virtualized and bare-metal nodes on the same local network. Nothing about the task looked unusual: kube-state-metrics needed to be deployed and pointed at as a new Prometheus data source in Grafana. What actually happened was a cascade of ten distinct, individually plausible theories, each ruled out in turn, before the real cause surfaced.

Layer 1 — A role that was never actually built

The first surprise had nothing to do with networking. Prometheus was returning "connection refused" against the monitoring host, and tracing that back led to discovering that node_exporter — checked off in the project checklist as already done — had no actual Ansible role behind it. It was built from scratch: install task, systemd unit, a ufw rule template. Rolling it out fleet-wide surfaced a separate, easy-to-miss Ansible gotcha: roles referenced with the shorthand - rolename list syntax silently ignore --tags. Only the - role: rolename dict form respects tags. Every role invoked with the shorthand form had been silently skipped whenever a tag-scoped run was used.

Layer 2 — Prometheus had been down for days, unnoticed

With node_exporter fixed, Prometheus itself turned out to have been crash-looping since several days earlier — invisible until something finally tried to query it. Three independent bugs were stacked on top of each other, each one only visible after the previous was fixed:

  1. A systemd unit flag typo — a hyphen where an equals sign belonged (--config.file-/... instead of --config.file=/...)
  2. A config path pointing at the wrong directory entirely
  3. A tabs-vs-spaces YAML syntax error a few lines into the config file

Each of these produced a distinct failure, and none of them were visible until the one before it was resolved. The real lesson here wasn't technical so much as operational: nothing had been alerting on Prometheus's own health, so a multi-day outage of the monitoring system went completely unnoticed. Silent monitoring failure is its own category of bug.

Layer 3 — Two out of three nodes wouldn't scrape

With Prometheus actually running, kube-state-metrics scraping still failed on two of the cluster's three nodes. Only the node the metrics pod happened to be scheduled on returned data; the other two timed out. The obvious first theory was a firewall gap — specifically, the VXLAN UDP port (8472) needed for Flannel's overlay network.

Layers 4–8 — A long run of plausible, wrong answers

What followed was a genuinely instructive elimination sequence, because each theory was reasonable given the evidence available at the time:

  • Ansible variable precedence. The VXLAN firewall rule existed in the worker group's variables but was missing from the parent group's — a real bug, and fixing it did get the rule onto the missing node. Traffic still failed.
  • Two competing firewall roles in the repo. An empty, leftover roles/firewall/ scaffold and the actual working role were both present, and the play responsible for the VXLAN rule was pointed at the empty one. Fixing that landed the rule correctly on all three nodes. Traffic still failed — so this wasn't the root cause either, just another real bug found along the way.
  • UFW's routed-traffic default policy. Since same-node traffic worked but cross-node forwarding didn't, the ufw route allow FORWARD chain looked like a good next suspect. Added an explicit allow rule for the Flannel-to-CNI path. No change.
  • iptables FORWARD chain / kube-router NetworkPolicy interaction. Verified IP forwarding was enabled at the kernel level, then inserted an explicit ACCEPT rule ahead of every Kubernetes-managed chain. Still no change — this definitively ruled out netfilter as the drop point.
  • MTU / Path MTU Discovery blackhole. A large, do-not-fragment ping succeeded cleanly, and plain ICMP and SSH between nodes worked fine. Only NodePort and pod-IP traffic specifically was affected — ruling out a fragmentation issue.

By this point, every layer a guest OS or standard Linux networking tool could observe had been checked and cleared.

Layer 9 — The test that actually proved something

The decisive step was running tcpdump directly on the affected node's flannel.1 interface while simultaneously watching the physical NIC beneath it. The physical NIC showed encapsulated VXLAN packets arriving — 87 of them, cleanly. The flannel.1 interface, which should have shown those same packets decapsulated, showed zero.

That's the detail that broke the case open: the packets weren't being dropped by a firewall rule or lost in transit. They were arriving intact at the network layer and then vanishing during kernel-level VXLAN decapsulation — a step that happens after tcpdump on the physical interface can observe anything, since packet capture happens before the kernel validates the encapsulated checksum.

Layer 10 — The actual root cause

The cause was a QEMU 10.2+ virtio-net feature-negotiation regression, specifically affecting Proxmox 9.2. When the physical NIC underneath a VM doesn't support UDP tunnel-checksum and tunnel-segmentation offload, QEMU incorrectly advertises those features to the guest's virtio-net driver anyway. The guest then believes it can offload VXLAN outer-UDP checksum computation to hardware that doesn't actually support it. The checksum gets computed incorrectly, the malformed packet looks structurally normal to any capture tool along the way, and the receiving host's kernel silently discards it as corrupt before it ever reaches the Flannel interface.

This is an actively tracked upstream bug in Proxmox's own issue tracker, staff-confirmed as reproducible, with no fix released at the time of this investigation. It affects any Kubernetes cluster with nodes split across Proxmox hosts running the affected QEMU/Proxmox version combination.

Fixes that looked right and weren't

Two fixes that should have worked, on paper, didn't:

  • Disabling checksum/segmentation offload from inside the guest (ethtool -K eth0 tx off rx off) — the bug lives in QEMU's feature negotiation, a layer below anything a guest-side tool can reach.
  • A Proxmox-forum-recommended fix explicitly disabling the tunnel-checksum arguments at the QEMU command-line level (-global virtio-net-pci.host_tunnel_csum=off, etc.). The arguments applied cleanly and were verified present in the VM's running config — the bug persisted anyway, for reasons that remain unclear.

The fix that actually worked

Rather than wait on an upstream QEMU/Proxmox fix, the resolution was architectural: switch Flannel's backend from VXLAN encapsulation to host-gw direct routing.

# /etc/rancher/k3s/config.yaml
flannel-backend: host-gw

Once applied and the cluster restarted (server first, then agents), the flannel.1 interfaces disappeared entirely, replaced by direct kernel routes to each node's pod CIDR over the physical network. kube-state-metrics scrape targets in Prometheus went from one of three healthy to three of three, immediately.

This isn't just a workaround. All three nodes in this cluster share the same Layer 2 network, meaning VXLAN encapsulation was providing zero actual benefit — there's no L3 boundary here for it to tunnel across. It was pure overhead and an unnecessary bug surface. host-gw is the standard choice for on-prem Kubernetes clusters where every node shares L2, offering higher throughput and lower CPU cost since there's no encapsulation step at all. VXLAN earns its complexity in cross-datacenter or cross-VPC topologies — not this one.

The fix was codified into the Ansible role that manages K3s configuration, with a comment linking back to this investigation so the reasoning survives the next person (or the next me) who looks at that file.

What this demonstrates

This investigation is really a case study in elimination-cascade debugging: every layer looked genuinely plausible when it was the current leading theory, and none of the rule-outs were wasted effort even though they weren't the answer — each one legitimately narrowed the search space and, in a couple of cases, fixed a real (if unrelated) bug along the way. The eventual breakthrough came from finding a vantage point — comparing the physical NIC's view against the virtual interface's view — that could actually distinguish "arrived and was silently discarded" from "never arrived at all," which no single-layer check up to that point could do.

It's also a reminder that the right fix isn't always the one that patches the reported bug. Here, the correct move was recognizing that the failing subsystem (VXLAN encapsulation) wasn't even architecturally necessary for this topology in the first place.