Gauge Hartwell
← all write-ups

sequential-failure

Two Unrelated Blockers Standing Up Falco: A Lying Test Tool and a CPU That Wasn't There

TL;DR: Getting Falco running hit two completely unrelated blockers back to back. The first looked like a network problem and turned out to be the diagnostic tool itself lying — a test pod's TLS client had a limitation the real workload didn't share, so every test "confirmed" a network failure that wasn't real. The second, right after clearing that, was a CPU compatibility crash several layers removed from anything network-related at all.

Act one: chasing a ghost

Falco's falcoctl-artifact-install init container failed fetching its rules index — with two different-looking errors on two different nodes: a DNS timeout on one, a TLS handshake failure on the other. Different symptoms, same init container, same fetch.

What followed was a long, genuinely reasonable elimination sequence:

  1. Checksum offload on the physical NIC (eth0) — disabled and retested, no change.
  2. Checksum offload on the VXLAN interface (flannel.1) — same, ruled out.
  3. Checksum offload on the Proxmox host's bridge (vmbr0) — same, ruled out.
  4. Raw MTU misconfiguration — checked directly; flannel.1 was already correctly set to 1450 to account for VXLAN overhead. Ruled out.
  5. A PMTU black hole (ICMP fragmentation-needed messages being silently dropped) — tested with a large ping from inside a pod. It succeeded cleanly near the MTU ceiling, directly contradicting the black-hole theory. Genuinely useful negative evidence, not just another "still failing."

Rather than reach for a sixth theory, the actual handshake was captured with tcpdump. The capture showed a clean TCP three-way handshake, a clean 120-byte ClientHello sent and received byte-for-byte intact across every hop, and a clean, correctly-formatted 7-byte TLS alert — fatal, handshake_failure — sent back by the remote server within milliseconds. Every layer of the network was healthy. The server was actively and correctly rejecting the handshake. This was never corruption.

The actual cause: the diagnostic pod used to reproduce the issue was built on bare busybox. BusyBox's built-in wget has famously incomplete TLS support — missing SNI in many builds — and the target domain was Cloudflare-fronted, which requires SNI to route to the correct backend at all. The test tool had the exact same limitation as the failure being diagnosed, so every test run "confirmed" a network problem that was actually just BusyBox being BusyBox.

Falco's own falcoctl binary is a proper Go program with full TLS support and shouldn't have hit this specific failure mode at all — but the underlying reachability of falcosecurity.github.io from this network was inconsistent enough to be a liability regardless of the TLS client involved. Rather than keep debugging a third-party dependency's flakiness, the fix used Falco's own documented Helm options to remove the dependency on that fetch entirely:

falcoctl.artifact.install.enabled=false
falcoctl.artifact.follow.enabled=false

Falco now runs on its baked-in default ruleset with no auto-update — a deliberate, documented trade-off, not a workaround that got forgotten about.

Act two: a CPU that wasn't there

With the network false lead cleared, Falco's main container crashed immediately with a completely different error:

Fatal glibc error: CPU does not support x86-64-v2

This read as a hardware compatibility problem, but worth verifying rather than assuming — "CPU doesn't support X" from inside a VM very often means the virtual CPU presented to the guest, not the physical one underneath it. That held up here: Proxmox's default VM CPU type (kvm64) emulates a lowest-common-denominator generic CPU that doesn't expose SSE4.2 — part of the x86-64-v2 baseline modern glibc builds increasingly assume is present. The physical host CPU supported it fine; the emulated CPU type presented to the guest did not.

Fixed by changing the affected VM's CPU type in Terraform to host, which passes through the physical CPU's real feature set instead of emulating a generic one. That trade-off was written down deliberately rather than treated as free: cpu: host ties a VM's exact feature set to whatever physical node it happens to be running on, which directly affects live migration reliability — a load-bearing assumption elsewhere in this build's HA plans. A follow-up task was added to verify migration still worked across the cluster's other nodes before relying on this operationally.

Worth flagging directly: a separate, later crash of this identical class (a NumPy build hitting the same missing-baseline error) was fixed differently — with a named CPU model, x86-64-v2-AES, specifically chosen over host because it's portable across nodes rather than pinned to one. Whether this Falco VM was ever revisited to match that better-considered fix, or is still running on the earlier host choice, is worth checking directly rather than assuming either way.

What this demonstrates

Two genuinely unrelated bugs, standing up the same service, in immediate succession. The first is really about tooling honesty: a "confirmed" result from a diagnostic tool with a hidden limitation of its own is worse than an inconclusive one, because it actively points the investigation in the wrong direction with false confidence — and the negative evidence gathered before that point (every checksum offload ruled out, MTU confirmed correct, large ICMP packets succeeding) wasn't wasted; it's what made the packet capture's contradiction obvious enough to actually catch. The second is a reminder that a fix resolving the symptom in front of you can quietly introduce a new constraint somewhere else in the system — worth writing that constraint down the moment it's introduced, while the reasoning is still fresh, rather than after it causes a confusing failure weeks later during an unrelated test.