Gauge Hartwell
← all write-ups

contradictory-evidence

When Static Config Lies: Migrating from Falco to Tetragon After a Plugin-Duplication Bug

TL;DR: Falco pods crash-looped with a plugin-duplication error that static configuration inspection insisted couldn't be happening. The actual cause was a Helm chart feature generating a duplicate plugin registration at runtime, in a file that never existed at rest and therefore never showed up in any static search. Rather than keep patching an under-documented interaction between three overlapping config subsystems, the decision was made to replace Falco outright with Tetragon.

The symptom

After resolving an unrelated CPU-compatibility crash earlier in the same session, Falco pods on two of the cluster's nodes came up and immediately crash-looped with:

Runtime error: cannot register plugin /usr/share/falco/plugins/libcontainer.so
in inspector: found another plugin with name "container". Aborting.. Exiting.

The obvious first theory — a duplicate plugin declared twice somewhere in config — was also, on inspection, wrong every time it was checked.

Chasing a bug that config said didn't exist

Three separate, reasonable theories were tested and ruled out in turn:

  1. Duplicate declaration in the rendered ConfigMap. Checked directly — the container plugin was declared exactly once. Ruled out.
  2. A self-restart race condition. Falco's watch_config_files: true default was a plausible candidate: a ConfigMap volume's atomic symlink swap colliding with Falco watching and reloading its own config mid-swap. Disabled the setting via helm upgrade and forced a pod restart. Pods crashed identically. Ruled out.
  3. A redundant plugin-install step. Falco's falcoctl component has a separate artifact-install step from its auto-update ("follow") step; the theory was that artifact-install was writing a redundant plugin config on every pod start. Checked — it was already disabled, same as the auto-update step. Ruled out.

A fourth attempt hit a dead end for an instructive reason: grepping every ConfigMap in the namespace for the specific file named in the crash log (falco.container_plugin.yaml) turned up nothing at all. That wasn't a failed search — it was the first real clue. The file wasn't missing; it simply didn't exist as a static Kubernetes object anywhere to be found, because it was being generated dynamically inside the pod's own ephemeral volume at startup.

There was also a more mundane obstacle worth naming directly: you cannot exec into a pod that's mid-CrashLoopBackOff — there's no running process to attach to during the backoff window. Live shell inspection simply wasn't available as a technique here, which forced a pivot to kubectl logs --previous and ConfigMap dumps instead. Adapting the debugging technique to a target that won't hold still turned out to matter as much as any individual theory.

The actual root cause

The breakthrough came from comparing two different Helm inspection commands against each other, because neither one told the full story alone:

  • helm get values -a (all computed values, including chart defaults) revealed collectors.containerEngine.enabled: true — a setting that hadn't been explicitly set, but was true by default.
  • helm get manifest (the fully rendered template output) showed load_plugins populated with container in the actual generated config — even though the same setting showed as an empty list under raw user-supplied values.

Put together, the picture was this: collectors.containerEngine.enabled independently does two things at template-render time. It generates a drop-in config file (/etc/falco/config.d/falco.container_plugin.yaml) that registers the container plugin, and it separately templates load_plugins: [container] directly into the main Falco config. Both paths register the same plugin, and the second registration attempt collides with the first — the actual cause of the crash, invisible to any single static inspection method because the collision only exists in the fully rendered output of two independent code paths inside the same chart.

The decision to replace rather than patch

At this point the choice was between continuing to patch around three overlapping, under-documented config subsystems (collectors.*, the manual plugins/load_plugins block, and falcoctl.artifact.*) or replacing Falco with something architecturally simpler. The decision was to migrate to Tetragon — Cilium/Isovalent's eBPF-native runtime security tool, which has no separate plugin or driver layer to collide with itself in the first place.

Migrating with parity, not a blind port

Before installing anything, kernel and runtime prerequisites were verified on all three nodes: BTF support, the specific kernel config flags Tetragon depends on (CONFIG_BPF, CONFIG_BPF_SYSCALL, CONFIG_BPF_LSM, CONFIG_CGROUPS, CONFIG_AUDITSYSCALL, CONFIG_BPF_KPROBE_OVERRIDE), architecture, and cgroup mode, with Tetragon's own tetra probe tool as the final authoritative check. One accepted limitation surfaced here: BPF LSM enforcement hooks weren't active at the kernel boot-parameter level, despite being compiled in. Since the migration's actual goal was detection parity with Falco rather than new enforcement capability, this was documented and accepted rather than chased further.

Falco's actual active ruleset was pulled directly from a stale-but-still-running reference pod that predated the same-day Helm changes — a useful reminder that when verifying "what's actually running now," it's worth explicitly excluding state that might not reflect current config, rather than assuming everything live is representative.

TracingPolicies were then authored to map to specific existing Falco rules, not ported 1:1:

TracingPolicy Falco rule replaced Mechanism
baseline-process-exec Baseline coverage requirement sys_execve kprobe
shell-in-container Terminal shell in container sys_execve + prefix match on shell binaries
privilege-escalation General privilege-change coverage sys_setuid + sys_setgid kprobes
file-monitoring-filtered Read sensitive file untrusted security_file_permission on /etc/shadow, /etc/passwd
outbound-connections Outbound connection rule family tcp_connect kprobe
kernel-module-injection Linux Kernel Module Injection Detected sys_init_module kprobe
release-agent-escape Detect release_agent File Container Escapes security_file_permission + prefix match on release_agent
memfd-fileless-exec Fileless execution via memfd_create sys_memfd_create kprobe

Authoring release-agent-escape surfaced a small but real gotcha worth flagging on its own: a case-sensitivity mismatch on the matching operator (PostFix where the correct value was Postfix) silently prevented the rule from matching anything until caught.

A handful of narrower Falco rules — AWS credential search, non-standard SSH ports, PTRACE anti-debug detection, log-clearing, symlink tricks — weren't individually ported. The authoring cost of a dedicated TracingPolicy for each was weighed against the marginal coverage they'd add on top of the broader policies already in place, and framed as accepted combined coverage rather than a gap.

A planned side-by-side trial period between Falco and Tetragon was deliberately skipped: Falco was still actively crash-looping, with restart counts climbing continuously, at the exact point Tetragon was fully validated with zero restarts across all three nodes. A tool that can't stay running isn't a meaningful baseline to compare anything against.

The monitoring pipeline that had to work before any of this was visible

Getting Tetragon's output into Grafana turned out to be its own multi-step debugging arc, distinct in character from the Falco investigation above. Alloy, the log-shipping agent, had never actually been deployed to the K3s nodes at all — an earlier fleet-wide rollout had only targeted the general Ansible host inventory, which the Terraform-provisioned K3s nodes were never part of. It was deployed properly as an in-cluster DaemonSet via Grafana's official Helm chart rather than retrofitted as a host-level service, which picks up every pod's output automatically with no per-host configuration.

Four smaller, independently diagnosable issues followed, each resolved with a single targeted check rather than a chain of false leads:

  1. A one-character typo (discover.kubernetes missing a letter) in the hand-written config, caught directly from Alloy's own reload-attempt logs.
  2. Loki briefly rejecting a batch of events as "too old" — not a bug at all, but Loki correctly refusing a stale backlog of pod logs that predated Alloy's own existence. Resolved on its own once the backlog drained.
  3. Logs ingested under a single combined instance label rather than the separate labels the original query assumed — the query wasn't wrong about the data being present, it was querying a label schema that didn't exist. The corrected query ({instance=~"kube-system/tetragon.*"}) matched immediately.
  4. Grafana's Explore UI double-quoting a value that already included quotes, producing a parse error — resolved by retyping the raw value or switching to the builder's code-input mode.

What this demonstrates

The Falco investigation and the Alloy pipeline work, done back to back in the same session, are genuinely different debugging skills. The Falco bug involved contradictory evidence across layers — static configuration insisted one thing while runtime behavior insisted another — and required systematically eliminating plausible-but-wrong theories until the contradiction itself pointed at the answer. The monitoring pipeline work was a chain of small, cleanly diagnosable issues, each one confirmed or ruled out with a single targeted check. Recognizing which kind of problem you're actually in changes how you should be spending your time on it.