Gauge Hartwell
← all write-ups

silently dead configuration

The Variable That Was Never Read: A Firewall Rule That Only Looked Configured

TL;DR: group_vars/dns/vars.yml declared port 53 open for both TCP and UDP, using variables named ufw_allowed_ports_udp and ufw_allowed_ports_tcp. The firewall role never read either one. Every port listed under those two variables had been silently doing nothing since the file was written — a plausible-looking convention that simply wasn't the one the actual code expected.

The symptom

DNS queries to the internal server timed out from every other host, despite group_vars/dns/vars.yml clearly listing port 53 as open, under both protocols, in what looked like a reasonable, self-documenting structure:

ufw_allowed_ports:
  - "22"
  - "80"
  - "3000"
ufw_allowed_ports_udp:
  - "53"
ufw_allowed_ports_tcp:
  - "53"

This wasn't a first offense for this exact category of bug — a UFW port timeout had already been diagnosed and fixed twice earlier in this same session, for ports 3000 and (separately) an earlier pass at 53. Given "this has been a problem before," the instinct was to check the actual consuming code rather than rewrite the variable a third time and hope a different name stuck.

Root cause

roles/hardening/tasks/firewall.yml's port-opening task loops over exactly one variable:

- name: Allow configured ports
  community.general.ufw:
    rule: allow
    port: "{{ item.port if item is mapping else item }}"
    proto: "{{ item.proto | default('tcp') if item is mapping else 'tcp' }}"
  loop: "{{ ufw_allowed_ports }}"

ufw_allowed_ports_udp and ufw_allowed_ports_tcp never appear anywhere in that file. They aren't a valid alternate convention this role happens to also support — they're simply not read at all. A UDP port gets opened only by including it as a mapping with an explicit proto: udp key inside the one list the role actually loops over; there's no separate per-protocol variable in this role's design.

Fix

ufw_allowed_ports:
  - "22"
  - "80"
  - "3000"
  - { port: "53", proto: "tcp" }
  - { port: "53", proto: "udp" }

The two dead variables were deleted outright rather than left alongside the fix — a file with variables that look like they're doing something, and aren't, is worse than no file at all. Same reasoning already applied to removing the orphaned group_vars/k3s_masters.yml earlier in this build rather than leaving it as clutter.

The fleet-wide check that mattered more than the fix itself

Given this was already the second time a UFW rule had silently failed to apply, the fix wasn't considered complete until checking whether the same wrong convention had been copy-pasted anywhere else:

grep -rl "ufw_allowed_ports_udp\|ufw_allowed_ports_tcp" group_vars/

Any other host using the same dead variable names would be sitting on an identical, silent gap — reachable-looking on paper, actually wide open to the same timeout the DNS host just hit.

What this demonstrates

A variable name that looks reasonable, self-documenting, and internally consistent tells you nothing about whether any code actually reads it. This is the same category of failure as the orphaned k3s_masters.yml file from earlier in this build — a plausible convention, silently doing nothing, discovered only by reading the actual consuming code rather than trusting how the declaring file looked. The fact that this was the third UFW timeout in the same build is itself a signal worth acting on directly: once a specific failure shape repeats, checking the authoritative source before writing a fourth guess is faster than continuing to theorize about naming conventions from the declaring side alone.