sequential-failure
Two Unrelated Blockers Standing Up Falco: A Lying Test Tool and a CPU That Wasn't There
TL;DR: Getting Falco running hit two completely unrelated blockers back to back. The first looked like a network problem and turned out to be the diagnostic tool itself lying — a test pod's TLS client had a limitation the real workload didn't share, so every test "confirmed" a network failure that wasn't real. The second, right after clearing that, was a CPU compatibility crash several layers removed from anything network-related at all.
Act one: chasing a ghost
Falco's falcoctl-artifact-install init container failed fetching its rules
index — with two different-looking errors on two different nodes: a DNS timeout on
one, a TLS handshake failure on the other. Different symptoms, same init
container, same fetch.
What followed was a long, genuinely reasonable elimination sequence:
- Checksum offload on the physical NIC (
eth0) — disabled and retested, no change. - Checksum offload on the VXLAN interface (
flannel.1) — same, ruled out. - Checksum offload on the Proxmox host's bridge (
vmbr0) — same, ruled out. - Raw MTU misconfiguration — checked directly;
flannel.1was already correctly set to 1450 to account for VXLAN overhead. Ruled out. - A PMTU black hole (ICMP fragmentation-needed messages being silently dropped) — tested with a large ping from inside a pod. It succeeded cleanly near the MTU ceiling, directly contradicting the black-hole theory. Genuinely useful negative evidence, not just another "still failing."
Rather than reach for a sixth theory, the actual handshake was captured with
tcpdump. The capture showed a clean TCP three-way handshake, a clean 120-byte
ClientHello sent and received byte-for-byte intact across every hop, and a clean,
correctly-formatted 7-byte TLS alert — fatal, handshake_failure — sent back by
the remote server within milliseconds. Every layer of the network was healthy. The
server was actively and correctly rejecting the handshake. This was never
corruption.
The actual cause: the diagnostic pod used to reproduce the issue was built on bare
busybox. BusyBox's built-in wget has famously incomplete TLS support — missing
SNI in many builds — and the target domain was Cloudflare-fronted, which requires
SNI to route to the correct backend at all. The test tool had the exact same
limitation as the failure being diagnosed, so every test run "confirmed" a network
problem that was actually just BusyBox being BusyBox.
Falco's own falcoctl binary is a proper Go program with full TLS support and
shouldn't have hit this specific failure mode at all — but the underlying
reachability of falcosecurity.github.io from this network was inconsistent enough
to be a liability regardless of the TLS client involved. Rather than keep debugging
a third-party dependency's flakiness, the fix used Falco's own documented Helm
options to remove the dependency on that fetch entirely:
falcoctl.artifact.install.enabled=false
falcoctl.artifact.follow.enabled=false
Falco now runs on its baked-in default ruleset with no auto-update — a deliberate, documented trade-off, not a workaround that got forgotten about.
Act two: a CPU that wasn't there
With the network false lead cleared, Falco's main container crashed immediately with a completely different error:
Fatal glibc error: CPU does not support x86-64-v2
This read as a hardware compatibility problem, but worth verifying rather than
assuming — "CPU doesn't support X" from inside a VM very often means the virtual
CPU presented to the guest, not the physical one underneath it. That held up here:
Proxmox's default VM CPU type (kvm64) emulates a lowest-common-denominator
generic CPU that doesn't expose SSE4.2 — part of the x86-64-v2 baseline modern
glibc builds increasingly assume is present. The physical host CPU supported it
fine; the emulated CPU type presented to the guest did not.
Fixed by changing the affected VM's CPU type in Terraform to host, which passes
through the physical CPU's real feature set instead of emulating a generic one.
That trade-off was written down deliberately rather than treated as free: cpu:
host ties a VM's exact feature set to whatever physical node it happens to be
running on, which directly affects live migration reliability — a load-bearing
assumption elsewhere in this build's HA plans. A follow-up task was added to verify
migration still worked across the cluster's other nodes before relying on this
operationally.
Worth flagging directly: a separate, later crash of this identical class (a
NumPy build hitting the same missing-baseline error) was fixed differently — with a
named CPU model, x86-64-v2-AES, specifically chosen over host because it's
portable across nodes rather than pinned to one. Whether this Falco VM was ever
revisited to match that better-considered fix, or is still running on the earlier
host choice, is worth checking directly rather than assuming either way.
What this demonstrates
Two genuinely unrelated bugs, standing up the same service, in immediate succession. The first is really about tooling honesty: a "confirmed" result from a diagnostic tool with a hidden limitation of its own is worse than an inconclusive one, because it actively points the investigation in the wrong direction with false confidence — and the negative evidence gathered before that point (every checksum offload ruled out, MTU confirmed correct, large ICMP packets succeeding) wasn't wasted; it's what made the packet capture's contradiction obvious enough to actually catch. The second is a reminder that a fix resolving the symptom in front of you can quietly introduce a new constraint somewhere else in the system — worth writing that constraint down the moment it's introduced, while the reasoning is still fresh, rather than after it causes a confusing failure weeks later during an unrelated test.