seven short notes
Quick Hits: Seven Smaller Debugging Notes
Not every bug needs a full investigation to be worth writing down. These seven are each a single clean root cause and fix — smaller in scope than the longer write-ups, but each one taught something specific enough to keep.
A one-character typo that broke systemd entirely
node_exporter refused to enable via Ansible:
Unable to enable service node_exporter: Failed to enable unit: Invalid unit name /usr/local/bin/node_exporter
The error name was the actual clue — systemd was being asked to treat a file
path as a unit name. The templated .service file had:
WantedBy=/usr/local/bin/node_exporter
WantedBy expects a systemd target (multi-user.target), not a binary path — a
stray copy-paste artifact had left the binary's path where the target name
belonged. Fixed with WantedBy=multi-user.target. systemd error messages are
usually more literal than they first appear; this one meant exactly what it said,
once the instinct to suspect the Ansible module logic was set aside in favor of
reading the actual rendered file.
GitLab's Puma crash loop, several layers from the visible symptom
GitLab CE installed cleanly and gitlab-ctl reconfigure completed with no errors,
but the web UI returned a persistent 502. gitlab-ctl status showed every service,
including Puma, as run — which turned out to be misleading. gitlab-ctl tail
puma revealed Puma was actually in a restart loop on a roughly 15-second cycle,
and several screens into the log:
"message":"Unable to load application: ArgumentError: Invalid Timezone: PST"
gitlab_rails['time_zone'] had been set to 'PST' — a common system abbreviation,
but not a valid Rails ActiveSupport::TimeZone identifier. Because this failed
during application boot, Puma could never successfully initialize, and the process
supervisor kept dutifully restarting something destined to crash every time. Fixed
with gitlab_rails['time_zone'] = 'UTC'. A service manager reporting a process as
run doesn't mean the process is healthy — it may just mean it's running again,
moments after its last crash. The application's own log, not the supervisor's
status output, is where the real error was.
A silent variable override that looked like a version bug
Changing loki_version in group_vars/all/vars.yml had zero visible effect — the
same download error, with the same version string, no matter what the value was
changed to. Rather than keep iterating on the value, ansible -m debug -a
"var=loki_version" showed what Ansible actually resolved for the host, bypassing
task logic entirely. group_vars/monitoring/vars.yml still had its own
loki_version defined, silently shadowing every edit made to the less-specific
file — a direct consequence of Ansible's group_vars precedence rules. Fixed by
removing the duplicate and consolidating to one source of truth. "I changed the
value and nothing happened" is a strong signal the value isn't being read from
where it's assumed to be — not a reason to keep changing the value further.
When the bug wasn't a bug: Promtail had been removed
Promtail's download 404'd. Three different version pins were tried, each
independently confirmed to be a real, published tag — same 404 every time. Once the
version itself felt ruled out, the next hypothesis had to be structural: maybe the
asset stopped existing, not the version. Loki's release notes confirmed it —
Promtail was explicitly removed starting with Loki 3.7.3, deprecated since Loki 3.0
in favor of Grafana Alloy. Migrated the log-shipping role to Alloy using its
built-in alloy convert --source-format=promtail command, translating the existing
config automatically rather than hand-writing it from scratch. A 404 on a versioned
download can mean "wrong version" or "this doesn't exist in any version anymore" —
when several plausible versions all fail identically, that's the signal to stop
iterating on the version and check whether the artifact itself still ships at all.
Same root cause, three different symptoms
unzip was missing on the DNS host, breaking an Alloy install with a wall of
scary-looking tar fallback errors — Ansible's unarchive module tries tar when
it can't find unzip, and buries the actual cause at the bottom of a long error
listing every format tar tried and failed: Unable to find required 'unzip' or
'zipinfo' binary in the path. A day later, the identical failure class appeared on
the monitoring host installing Loki. Immediate fix: an apt: name=unzip task ahead
of each affected download. The real lesson is about where the fix belongs — a
missing base dependency will keep resurfacing on every new host until it's handled
at the fleet level (a shared bootstrap role run once per host), not the task level.
Consolidating this was deferred deliberately to save time mid-build, a known
trade-off made on purpose rather than by accident.
A blockinfile bug that silently broke a security tool
The Wazuh agent failed to start: Invalid element in the configuration:
'ossec_config' — meaning the config file's root element was somehow invalid.
xmllint --noout on the rendered file pinpointed it precisely, at a spot the file
looked structurally fine to the eye: Extra content at the end of the document —
the signature of multiple XML root elements in a file that permits exactly one. Two
separate blockinfile tasks — one for file integrity monitoring, one for SSH log
config — each wrapped their inserted content in a complete
<ossec_config>...</ossec_config> pair instead of just the inner element. Since
each task correctly used insertbefore: "</ossec_config>", the blocks landed in the
right position, but each one's content duplicated the root tag already there.
Fixed by stripping the redundant wrapper from each block, leaving only the actual
<syscheck>/<localfile> elements. blockinfile's idempotency only covers its own
marked section — it will not retroactively repair a file a previous buggy run
already corrupted. Fixing the task logic and fixing already-deployed instances of
the bug are two separate steps, and skipping the second leaves already-provisioned
hosts silently broken even after the fix ships.
The Jinja2 character that turned a variable into subtraction
"msg": "The task includes an option with an undefined variable.. 'prometheus' is undefined"
— but prometheus_version was defined correctly, right there in
group_vars/all/vars.yml. The error naming a shorter variable than the one
actually defined was the tell: something was splitting the name into two pieces.
The template had {{ prometheus-version }} — a hyphen instead of an underscore.
Jinja2 parsed that as the expression prometheus - version: subtracting one
undefined variable from another, rather than referencing a single variable. Fixed
with {{ prometheus_version }}. When an error names a variable shorter or
differently spelled than what's actually defined, check for a character doing
double duty as an operator — Jinja2 won't warn that a hyphen inside {{ }} means
something entirely different from a hyphen inside a plain string.