contradictory-evidence
When Static Config Lies: Migrating from Falco to Tetragon After a Plugin-Duplication Bug
TL;DR: Falco pods crash-looped with a plugin-duplication error that static configuration inspection insisted couldn't be happening. The actual cause was a Helm chart feature generating a duplicate plugin registration at runtime, in a file that never existed at rest and therefore never showed up in any static search. Rather than keep patching an under-documented interaction between three overlapping config subsystems, the decision was made to replace Falco outright with Tetragon.
The symptom
After resolving an unrelated CPU-compatibility crash earlier in the same session, Falco pods on two of the cluster's nodes came up and immediately crash-looped with:
Runtime error: cannot register plugin /usr/share/falco/plugins/libcontainer.so
in inspector: found another plugin with name "container". Aborting.. Exiting.
The obvious first theory — a duplicate plugin declared twice somewhere in config — was also, on inspection, wrong every time it was checked.
Chasing a bug that config said didn't exist
Three separate, reasonable theories were tested and ruled out in turn:
- Duplicate declaration in the rendered ConfigMap. Checked directly — the
containerplugin was declared exactly once. Ruled out. - A self-restart race condition. Falco's
watch_config_files: truedefault was a plausible candidate: a ConfigMap volume's atomic symlink swap colliding with Falco watching and reloading its own config mid-swap. Disabled the setting viahelm upgradeand forced a pod restart. Pods crashed identically. Ruled out. - A redundant plugin-install step. Falco's
falcoctlcomponent has a separate artifact-install step from its auto-update ("follow") step; the theory was that artifact-install was writing a redundant plugin config on every pod start. Checked — it was already disabled, same as the auto-update step. Ruled out.
A fourth attempt hit a dead end for an instructive reason: grepping every ConfigMap in
the namespace for the specific file named in the crash log
(falco.container_plugin.yaml) turned up nothing at all. That wasn't a failed search —
it was the first real clue. The file wasn't missing; it simply didn't exist as a static
Kubernetes object anywhere to be found, because it was being generated dynamically
inside the pod's own ephemeral volume at startup.
There was also a more mundane obstacle worth naming directly: you cannot exec into a
pod that's mid-CrashLoopBackOff — there's no running process to attach to during the
backoff window. Live shell inspection simply wasn't available as a technique here, which
forced a pivot to kubectl logs --previous and ConfigMap dumps instead. Adapting the
debugging technique to a target that won't hold still turned out to matter as much as
any individual theory.
The actual root cause
The breakthrough came from comparing two different Helm inspection commands against each other, because neither one told the full story alone:
helm get values -a(all computed values, including chart defaults) revealedcollectors.containerEngine.enabled: true— a setting that hadn't been explicitly set, but was true by default.helm get manifest(the fully rendered template output) showedload_pluginspopulated withcontainerin the actual generated config — even though the same setting showed as an empty list under raw user-supplied values.
Put together, the picture was this: collectors.containerEngine.enabled independently
does two things at template-render time. It generates a drop-in config file
(/etc/falco/config.d/falco.container_plugin.yaml) that registers the container
plugin, and it separately templates load_plugins: [container] directly into the
main Falco config. Both paths register the same plugin, and the second registration
attempt collides with the first — the actual cause of the crash, invisible to any single
static inspection method because the collision only exists in the fully rendered output
of two independent code paths inside the same chart.
The decision to replace rather than patch
At this point the choice was between continuing to patch around three overlapping,
under-documented config subsystems (collectors.*, the manual plugins/load_plugins
block, and falcoctl.artifact.*) or replacing Falco with something architecturally
simpler. The decision was to migrate to Tetragon — Cilium/Isovalent's eBPF-native
runtime security tool, which has no separate plugin or driver layer to collide with
itself in the first place.
Migrating with parity, not a blind port
Before installing anything, kernel and runtime prerequisites were verified on all three
nodes: BTF support, the specific kernel config flags Tetragon depends on
(CONFIG_BPF, CONFIG_BPF_SYSCALL, CONFIG_BPF_LSM, CONFIG_CGROUPS,
CONFIG_AUDITSYSCALL, CONFIG_BPF_KPROBE_OVERRIDE), architecture, and cgroup mode,
with Tetragon's own tetra probe tool as the final authoritative check. One accepted
limitation surfaced here: BPF LSM enforcement hooks weren't active at the kernel
boot-parameter level, despite being compiled in. Since the migration's actual goal was
detection parity with Falco rather than new enforcement capability, this was documented
and accepted rather than chased further.
Falco's actual active ruleset was pulled directly from a stale-but-still-running reference pod that predated the same-day Helm changes — a useful reminder that when verifying "what's actually running now," it's worth explicitly excluding state that might not reflect current config, rather than assuming everything live is representative.
TracingPolicies were then authored to map to specific existing Falco rules, not ported 1:1:
| TracingPolicy | Falco rule replaced | Mechanism |
|---|---|---|
baseline-process-exec |
Baseline coverage requirement | sys_execve kprobe |
shell-in-container |
Terminal shell in container | sys_execve + prefix match on shell binaries |
privilege-escalation |
General privilege-change coverage | sys_setuid + sys_setgid kprobes |
file-monitoring-filtered |
Read sensitive file untrusted | security_file_permission on /etc/shadow, /etc/passwd |
outbound-connections |
Outbound connection rule family | tcp_connect kprobe |
kernel-module-injection |
Linux Kernel Module Injection Detected | sys_init_module kprobe |
release-agent-escape |
Detect release_agent File Container Escapes | security_file_permission + prefix match on release_agent |
memfd-fileless-exec |
Fileless execution via memfd_create | sys_memfd_create kprobe |
Authoring release-agent-escape surfaced a small but real gotcha worth flagging on its
own: a case-sensitivity mismatch on the matching operator (PostFix where the correct
value was Postfix) silently prevented the rule from matching anything until caught.
A handful of narrower Falco rules — AWS credential search, non-standard SSH ports, PTRACE anti-debug detection, log-clearing, symlink tricks — weren't individually ported. The authoring cost of a dedicated TracingPolicy for each was weighed against the marginal coverage they'd add on top of the broader policies already in place, and framed as accepted combined coverage rather than a gap.
A planned side-by-side trial period between Falco and Tetragon was deliberately skipped: Falco was still actively crash-looping, with restart counts climbing continuously, at the exact point Tetragon was fully validated with zero restarts across all three nodes. A tool that can't stay running isn't a meaningful baseline to compare anything against.
The monitoring pipeline that had to work before any of this was visible
Getting Tetragon's output into Grafana turned out to be its own multi-step debugging arc, distinct in character from the Falco investigation above. Alloy, the log-shipping agent, had never actually been deployed to the K3s nodes at all — an earlier fleet-wide rollout had only targeted the general Ansible host inventory, which the Terraform-provisioned K3s nodes were never part of. It was deployed properly as an in-cluster DaemonSet via Grafana's official Helm chart rather than retrofitted as a host-level service, which picks up every pod's output automatically with no per-host configuration.
Four smaller, independently diagnosable issues followed, each resolved with a single targeted check rather than a chain of false leads:
- A one-character typo (
discover.kubernetesmissing a letter) in the hand-written config, caught directly from Alloy's own reload-attempt logs. - Loki briefly rejecting a batch of events as "too old" — not a bug at all, but Loki correctly refusing a stale backlog of pod logs that predated Alloy's own existence. Resolved on its own once the backlog drained.
- Logs ingested under a single combined
instancelabel rather than the separate labels the original query assumed — the query wasn't wrong about the data being present, it was querying a label schema that didn't exist. The corrected query ({instance=~"kube-system/tetragon.*"}) matched immediately. - Grafana's Explore UI double-quoting a value that already included quotes, producing a parse error — resolved by retyping the raw value or switching to the builder's code-input mode.
What this demonstrates
The Falco investigation and the Alloy pipeline work, done back to back in the same session, are genuinely different debugging skills. The Falco bug involved contradictory evidence across layers — static configuration insisted one thing while runtime behavior insisted another — and required systematically eliminating plausible-but-wrong theories until the contradiction itself pointed at the answer. The monitoring pipeline work was a chain of small, cleanly diagnosable issues, each one confirmed or ruled out with a single targeted check. Recognizing which kind of problem you're actually in changes how you should be spending your time on it.