Gauge Hartwell
← all write-ups

seven short notes

Quick Hits: Seven Smaller Debugging Notes

Not every bug needs a full investigation to be worth writing down. These seven are each a single clean root cause and fix — smaller in scope than the longer write-ups, but each one taught something specific enough to keep.

A one-character typo that broke systemd entirely

node_exporter refused to enable via Ansible:

Unable to enable service node_exporter: Failed to enable unit: Invalid unit name /usr/local/bin/node_exporter

The error name was the actual clue — systemd was being asked to treat a file path as a unit name. The templated .service file had:

WantedBy=/usr/local/bin/node_exporter

WantedBy expects a systemd target (multi-user.target), not a binary path — a stray copy-paste artifact had left the binary's path where the target name belonged. Fixed with WantedBy=multi-user.target. systemd error messages are usually more literal than they first appear; this one meant exactly what it said, once the instinct to suspect the Ansible module logic was set aside in favor of reading the actual rendered file.

GitLab's Puma crash loop, several layers from the visible symptom

GitLab CE installed cleanly and gitlab-ctl reconfigure completed with no errors, but the web UI returned a persistent 502. gitlab-ctl status showed every service, including Puma, as run — which turned out to be misleading. gitlab-ctl tail puma revealed Puma was actually in a restart loop on a roughly 15-second cycle, and several screens into the log:

"message":"Unable to load application: ArgumentError: Invalid Timezone: PST"

gitlab_rails['time_zone'] had been set to 'PST' — a common system abbreviation, but not a valid Rails ActiveSupport::TimeZone identifier. Because this failed during application boot, Puma could never successfully initialize, and the process supervisor kept dutifully restarting something destined to crash every time. Fixed with gitlab_rails['time_zone'] = 'UTC'. A service manager reporting a process as run doesn't mean the process is healthy — it may just mean it's running again, moments after its last crash. The application's own log, not the supervisor's status output, is where the real error was.

A silent variable override that looked like a version bug

Changing loki_version in group_vars/all/vars.yml had zero visible effect — the same download error, with the same version string, no matter what the value was changed to. Rather than keep iterating on the value, ansible -m debug -a "var=loki_version" showed what Ansible actually resolved for the host, bypassing task logic entirely. group_vars/monitoring/vars.yml still had its own loki_version defined, silently shadowing every edit made to the less-specific file — a direct consequence of Ansible's group_vars precedence rules. Fixed by removing the duplicate and consolidating to one source of truth. "I changed the value and nothing happened" is a strong signal the value isn't being read from where it's assumed to be — not a reason to keep changing the value further.

When the bug wasn't a bug: Promtail had been removed

Promtail's download 404'd. Three different version pins were tried, each independently confirmed to be a real, published tag — same 404 every time. Once the version itself felt ruled out, the next hypothesis had to be structural: maybe the asset stopped existing, not the version. Loki's release notes confirmed it — Promtail was explicitly removed starting with Loki 3.7.3, deprecated since Loki 3.0 in favor of Grafana Alloy. Migrated the log-shipping role to Alloy using its built-in alloy convert --source-format=promtail command, translating the existing config automatically rather than hand-writing it from scratch. A 404 on a versioned download can mean "wrong version" or "this doesn't exist in any version anymore" — when several plausible versions all fail identically, that's the signal to stop iterating on the version and check whether the artifact itself still ships at all.

Same root cause, three different symptoms

unzip was missing on the DNS host, breaking an Alloy install with a wall of scary-looking tar fallback errors — Ansible's unarchive module tries tar when it can't find unzip, and buries the actual cause at the bottom of a long error listing every format tar tried and failed: Unable to find required 'unzip' or 'zipinfo' binary in the path. A day later, the identical failure class appeared on the monitoring host installing Loki. Immediate fix: an apt: name=unzip task ahead of each affected download. The real lesson is about where the fix belongs — a missing base dependency will keep resurfacing on every new host until it's handled at the fleet level (a shared bootstrap role run once per host), not the task level. Consolidating this was deferred deliberately to save time mid-build, a known trade-off made on purpose rather than by accident.

A blockinfile bug that silently broke a security tool

The Wazuh agent failed to start: Invalid element in the configuration: 'ossec_config' — meaning the config file's root element was somehow invalid. xmllint --noout on the rendered file pinpointed it precisely, at a spot the file looked structurally fine to the eye: Extra content at the end of the document — the signature of multiple XML root elements in a file that permits exactly one. Two separate blockinfile tasks — one for file integrity monitoring, one for SSH log config — each wrapped their inserted content in a complete <ossec_config>...</ossec_config> pair instead of just the inner element. Since each task correctly used insertbefore: "</ossec_config>", the blocks landed in the right position, but each one's content duplicated the root tag already there. Fixed by stripping the redundant wrapper from each block, leaving only the actual <syscheck>/<localfile> elements. blockinfile's idempotency only covers its own marked section — it will not retroactively repair a file a previous buggy run already corrupted. Fixing the task logic and fixing already-deployed instances of the bug are two separate steps, and skipping the second leaves already-provisioned hosts silently broken even after the fix ships.

The Jinja2 character that turned a variable into subtraction

"msg": "The task includes an option with an undefined variable.. 'prometheus' is undefined"

— but prometheus_version was defined correctly, right there in group_vars/all/vars.yml. The error naming a shorter variable than the one actually defined was the tell: something was splitting the name into two pieces. The template had {{ prometheus-version }} — a hyphen instead of an underscore. Jinja2 parsed that as the expression prometheus - version: subtracting one undefined variable from another, rather than referencing a single variable. Fixed with {{ prometheus_version }}. When an error names a variable shorter or differently spelled than what's actually defined, check for a character doing double duty as an operator — Jinja2 won't warn that a hyphen inside {{ }} means something entirely different from a hyphen inside a plain string.