layered failure, misleading tooling
Four Config Surfaces, One Symptom: Diagnosing a Cloudflare 1016 "Origin DNS Error"
TL;DR: After finally getting the deploy pipeline working, the site itself was still unreachable — Cloudflare returned a generic "Origin DNS error." The name suggests one problem. It's actually a catch-all covering four structurally different failure points, each living in a completely different configuration surface, and one early "success" during the investigation turned out to be a resolver artifact lying about a working DNS lookup.
Layer 1 — is the domain resolving at all, and a resolver that lied about it
dig gaugehartwell.com from the Ansible LXC returned completely empty output — no
error, no header, nothing at all. That alone was worth treating as suspicious;
dig normally prints something even for a failed lookup.
getent hosts gaugehartwell.com told a different story: two Cloudflare IPv6
addresses, apparent success. Except those were the exact same addresses returned
for hartwelg.com — a completely different domain. That's not a coincidence — the
resolver's search hartwelg.com directive was silently appending itself,
querying the nonsense gaugehartwell.com.hartwelg.com, and landing on a wildcard
record in the hartwelg.com zone. Every earlier "it resolves fine" check had
actually been testing the wrong domain the whole time.
Confirmed by forcing an absolute FQDN — a trailing dot disables search-domain substitution entirely:
getent hosts gaugehartwell.com.
Empty. That was the real, honest answer: the domain genuinely wasn't resolving, and the earlier apparent success was the resolver quietly substituting a different domain's answer.
Layer 2 — is the domain actually delegated to Cloudflare
Rather than keep trusting a resolver that had just been caught lying, this check bypassed it entirely with a direct DNS-over-HTTPS query:
curl "https://cloudflare-dns.com/dns-query?name=gaugehartwell.com&type=NS"
Returned alex.ns.cloudflare.com and brit.ns.cloudflare.com — delegation was
correct. One layer cleared, cleanly, with a tool that couldn't be fooled by local
search-domain configuration.
Layer 3 — does a CNAME record actually exist
Same DoH technique, different record type:
curl "https://cloudflare-dns.com/dns-query?name=gaugehartwell.com&type=CNAME"
Empty. Nobody had ever actually created the public DNS record binding
gaugehartwell.com to the tunnel endpoint. Added it via the Cloudflare dashboard.
Layer 4 — does the CNAME point at the right tunnel
With the record now in place, the site returned 1016 again — but this time with a
specific, useful reason: the CNAME's target UUID didn't match the currently-running
cloudflared tunnel. A single value correction fixed it.
Fix
Also added www.gaugehartwell.com as its own Public Hostname entry in the
Cloudflare Tunnel dashboard — which creates the matching DNS record automatically —
so both the bare domain and the www. subdomain resolve correctly.
What this demonstrates
"Origin DNS error" names one symptom for at least four structurally different causes — registrar delegation, zone record existence, CNAME target correctness, and tunnel-side route mapping — and debugging it means isolating which layer is actually broken rather than guessing at the most likely one. The more interesting lesson is buried in Layer 1: a resolver silently appending a search domain produced a completely plausible-looking success — real IP addresses, no error — for entirely the wrong query. When a diagnostic tool goes quiet in one attempt and reports a clean success in the next, the second answer isn't automatically the trustworthy one. The fix wasn't a louder version of the same tool; it was a more precise question — an absolute FQDN, then bypassing the local resolver chain altogether with a direct DoH query — that couldn't be fooled by local configuration in the same way.
A note on this build's own DNS pattern
This isn't the only time a search-domain-style resolution quirk has caused
confusion in this lab — hartwelg.com subdomains not reliably resolving to
internal addresses (needing AdGuard's still-pending split-horizon DNS, Phase 11)
has come up as a recurring friction point elsewhere in this build too. Worth
keeping in mind as one more data point for whether internal DNS is worth
prioritizing sooner rather than later.