sequential-cascade, one wrong theory
Four Bugs Deep to Sync a Clock: Getting Every Host Onto the Internal NTP Server
TL;DR: Pointing the fleet's NTP clients at the new internal server should have been a one-line config change repeated across a role. It took four sequential bugs — each one only visible after fixing the one before it — including a wrong theory that was still a real, well-documented behavior, and ended with discovering that Ubuntu had quietly moved where a default configuration actually lives.
Bug one: a variable that existed, but not where the code expected
The first task checked which time-sync mechanism each host actually ran:
- name: Gather service facts
ansible.builtin.service_facts:
- name: Point systemd-timesyncd at the internal NTP server
...
when: "'systemd-timesyncd.service' in services"
which failed immediately: 'services' is undefined. service_facts had run
successfully — the data existed — but this environment's ansible.cfg sets
inject_facts_as_vars = false, meaning gathered facts never flatten into a bare
top-level variable the way default Ansible behavior implies. The fix was
referencing the real path: ansible_facts.services instead of services.
Bug two: the fix was written, but the old logic never got replaced
Adding a more precise check — querying systemctl show <unit> --property=LoadState
directly rather than trusting service_facts — surfaced a specific, real gap on
two hosts where the two sources disagreed. A new command task was added to
register the real load state. But the task that actually mattered was never
updated to use it:
- name: Point systemd-timesyncd at the internal NTP server
...
when: "'systemd-timesyncd.service' in ansible_facts.services" # never changed
The new check ran, registered a correct result, and was ignored — the real logic
downstream was still governed by the exact condition already proven wrong. Fixed
by actually pointing the when at the new variable:
"'LoadState=loaded' in timesyncd_loadstate.stdout".
Bug three: a wrong theory that was still real
With both fixes in place, a --check --diff run showed every task skipping for
every host — including the two check tasks, which had no when clause of their
own at all. That specific detail ruled out the two fixes above; something else was
suppressing everything uniformly.
The first theory was a well-documented, real Ansible behavior: attaching a when
to a static import_tasks or a block copies that condition onto every task
inside individually — exactly matching "every task skips, even ones with no
condition of their own." This is a genuine, officially-documented gotcha, backed
by an ansible-lint rule written specifically to warn against it. It was also not
what was actually happening here.
The real cause was narrower and specific to the command module:
ansible.builtin.command cannot meaningfully simulate its own output under
--check — there's no way to know what a command would print without running
it — so it doesn't execute at all, and the registered variable stays empty. Every
downstream when referencing that empty variable correctly evaluated false, for
every host, uniformly. The fix: check_mode: false on the two command tasks,
since they're pure reads with no side effects, safe to always execute for real
even during a check run.
Bug four: the fix reported success and changed nothing
With --check --diff finally showing a real, sensible diff, the play was run for
real. It reported changed. chronyc tracking on the affected hosts still showed
public Canonical NTP servers as the active source — not the internal server at all.
The lineinfile task had been searching for a line starting with pool inside
chrony.conf, to replace it with the internal server. That line never matched,
because current Ubuntu chrony packaging relocated the actual public-pool
declaration entirely: chrony.conf itself now just says
sourcedir /etc/chrony/sources.d, and the real pool directive lives in a
separate file, /etc/chrony/sources.d/ubuntu-ntp-pools.sources. The regexp
matched nothing, lineinfile fell back to appending the new server line at the
end of chrony.conf — which is why it correctly reported changed — while the
actual public source file sat completely untouched. Chrony, seeing both a proven
public source and a brand-new unproven internal one, kept using the source it
already trusted more.
Fixed by targeting the actual file Ubuntu now designates for this:
- name: Replace Ubuntu's default NTP pool sources with the internal server
ansible.builtin.copy:
dest: /etc/chrony/sources.d/ubuntu-ntp-pools.sources
content: "server {{ ntp_server }} iburst\n"
Confirmed for real this time — chronyc tracking showing the internal server's
actual IP as the Reference ID, not a hostname belonging to Canonical.
What this demonstrates
Three of these four bugs only became visible because the previous one was fixed
first — a linear dependency chain where "it's still not working" kept meaning
something different each time. The third bug is the one worth sitting with longest:
a wrong theory (import_tasks/when scoping) was chased confidently, with real
documentation and a real linter rule behind it, and it was still the wrong
explanation for this specific symptom. Being well-researched and being correct
aren't the same claim — the actual tell was checking what command specifically
does under --check rather than trusting how convincingly the first theory fit the
symptom on paper. And the final bug is a reminder that "changed: true" is a report
about whether a task did something, not proof that the something it did was the
right something — a lineinfile task can succeed completely, by its own
definition of success, while missing the file that actually mattered.