Gauge Hartwell
← all write-ups

sequential-cascade, one wrong theory

Four Bugs Deep to Sync a Clock: Getting Every Host Onto the Internal NTP Server

TL;DR: Pointing the fleet's NTP clients at the new internal server should have been a one-line config change repeated across a role. It took four sequential bugs — each one only visible after fixing the one before it — including a wrong theory that was still a real, well-documented behavior, and ended with discovering that Ubuntu had quietly moved where a default configuration actually lives.

Bug one: a variable that existed, but not where the code expected

The first task checked which time-sync mechanism each host actually ran:

- name: Gather service facts
  ansible.builtin.service_facts:

- name: Point systemd-timesyncd at the internal NTP server
  ...
  when: "'systemd-timesyncd.service' in services"

which failed immediately: 'services' is undefined. service_facts had run successfully — the data existed — but this environment's ansible.cfg sets inject_facts_as_vars = false, meaning gathered facts never flatten into a bare top-level variable the way default Ansible behavior implies. The fix was referencing the real path: ansible_facts.services instead of services.

Bug two: the fix was written, but the old logic never got replaced

Adding a more precise check — querying systemctl show <unit> --property=LoadState directly rather than trusting service_facts — surfaced a specific, real gap on two hosts where the two sources disagreed. A new command task was added to register the real load state. But the task that actually mattered was never updated to use it:

- name: Point systemd-timesyncd at the internal NTP server
  ...
  when: "'systemd-timesyncd.service' in ansible_facts.services"   # never changed

The new check ran, registered a correct result, and was ignored — the real logic downstream was still governed by the exact condition already proven wrong. Fixed by actually pointing the when at the new variable: "'LoadState=loaded' in timesyncd_loadstate.stdout".

Bug three: a wrong theory that was still real

With both fixes in place, a --check --diff run showed every task skipping for every host — including the two check tasks, which had no when clause of their own at all. That specific detail ruled out the two fixes above; something else was suppressing everything uniformly.

The first theory was a well-documented, real Ansible behavior: attaching a when to a static import_tasks or a block copies that condition onto every task inside individually — exactly matching "every task skips, even ones with no condition of their own." This is a genuine, officially-documented gotcha, backed by an ansible-lint rule written specifically to warn against it. It was also not what was actually happening here.

The real cause was narrower and specific to the command module: ansible.builtin.command cannot meaningfully simulate its own output under --check — there's no way to know what a command would print without running it — so it doesn't execute at all, and the registered variable stays empty. Every downstream when referencing that empty variable correctly evaluated false, for every host, uniformly. The fix: check_mode: false on the two command tasks, since they're pure reads with no side effects, safe to always execute for real even during a check run.

Bug four: the fix reported success and changed nothing

With --check --diff finally showing a real, sensible diff, the play was run for real. It reported changed. chronyc tracking on the affected hosts still showed public Canonical NTP servers as the active source — not the internal server at all.

The lineinfile task had been searching for a line starting with pool inside chrony.conf, to replace it with the internal server. That line never matched, because current Ubuntu chrony packaging relocated the actual public-pool declaration entirely: chrony.conf itself now just says sourcedir /etc/chrony/sources.d, and the real pool directive lives in a separate file, /etc/chrony/sources.d/ubuntu-ntp-pools.sources. The regexp matched nothing, lineinfile fell back to appending the new server line at the end of chrony.conf — which is why it correctly reported changed — while the actual public source file sat completely untouched. Chrony, seeing both a proven public source and a brand-new unproven internal one, kept using the source it already trusted more.

Fixed by targeting the actual file Ubuntu now designates for this:

- name: Replace Ubuntu's default NTP pool sources with the internal server
  ansible.builtin.copy:
    dest: /etc/chrony/sources.d/ubuntu-ntp-pools.sources
    content: "server {{ ntp_server }} iburst\n"

Confirmed for real this time — chronyc tracking showing the internal server's actual IP as the Reference ID, not a hostname belonging to Canonical.

What this demonstrates

Three of these four bugs only became visible because the previous one was fixed first — a linear dependency chain where "it's still not working" kept meaning something different each time. The third bug is the one worth sitting with longest: a wrong theory (import_tasks/when scoping) was chased confidently, with real documentation and a real linter rule behind it, and it was still the wrong explanation for this specific symptom. Being well-researched and being correct aren't the same claim — the actual tell was checking what command specifically does under --check rather than trusting how convincingly the first theory fit the symptom on paper. And the final bug is a reminder that "changed: true" is a report about whether a task did something, not proof that the something it did was the right something — a lineinfile task can succeed completely, by its own definition of success, while missing the file that actually mattered.