Gauge Hartwell
← all write-ups

layered failure, misleading tooling

Four Config Surfaces, One Symptom: Diagnosing a Cloudflare 1016 "Origin DNS Error"

TL;DR: After finally getting the deploy pipeline working, the site itself was still unreachable — Cloudflare returned a generic "Origin DNS error." The name suggests one problem. It's actually a catch-all covering four structurally different failure points, each living in a completely different configuration surface, and one early "success" during the investigation turned out to be a resolver artifact lying about a working DNS lookup.

Layer 1 — is the domain resolving at all, and a resolver that lied about it

dig gaugehartwell.com from the Ansible LXC returned completely empty output — no error, no header, nothing at all. That alone was worth treating as suspicious; dig normally prints something even for a failed lookup.

getent hosts gaugehartwell.com told a different story: two Cloudflare IPv6 addresses, apparent success. Except those were the exact same addresses returned for hartwelg.com — a completely different domain. That's not a coincidence — the resolver's search hartwelg.com directive was silently appending itself, querying the nonsense gaugehartwell.com.hartwelg.com, and landing on a wildcard record in the hartwelg.com zone. Every earlier "it resolves fine" check had actually been testing the wrong domain the whole time.

Confirmed by forcing an absolute FQDN — a trailing dot disables search-domain substitution entirely:

getent hosts gaugehartwell.com.

Empty. That was the real, honest answer: the domain genuinely wasn't resolving, and the earlier apparent success was the resolver quietly substituting a different domain's answer.

Layer 2 — is the domain actually delegated to Cloudflare

Rather than keep trusting a resolver that had just been caught lying, this check bypassed it entirely with a direct DNS-over-HTTPS query:

curl "https://cloudflare-dns.com/dns-query?name=gaugehartwell.com&type=NS"

Returned alex.ns.cloudflare.com and brit.ns.cloudflare.com — delegation was correct. One layer cleared, cleanly, with a tool that couldn't be fooled by local search-domain configuration.

Layer 3 — does a CNAME record actually exist

Same DoH technique, different record type:

curl "https://cloudflare-dns.com/dns-query?name=gaugehartwell.com&type=CNAME"

Empty. Nobody had ever actually created the public DNS record binding gaugehartwell.com to the tunnel endpoint. Added it via the Cloudflare dashboard.

Layer 4 — does the CNAME point at the right tunnel

With the record now in place, the site returned 1016 again — but this time with a specific, useful reason: the CNAME's target UUID didn't match the currently-running cloudflared tunnel. A single value correction fixed it.

Fix

Also added www.gaugehartwell.com as its own Public Hostname entry in the Cloudflare Tunnel dashboard — which creates the matching DNS record automatically — so both the bare domain and the www. subdomain resolve correctly.

What this demonstrates

"Origin DNS error" names one symptom for at least four structurally different causes — registrar delegation, zone record existence, CNAME target correctness, and tunnel-side route mapping — and debugging it means isolating which layer is actually broken rather than guessing at the most likely one. The more interesting lesson is buried in Layer 1: a resolver silently appending a search domain produced a completely plausible-looking success — real IP addresses, no error — for entirely the wrong query. When a diagnostic tool goes quiet in one attempt and reports a clean success in the next, the second answer isn't automatically the trustworthy one. The fix wasn't a louder version of the same tool; it was a more precise question — an absolute FQDN, then bypassing the local resolver chain altogether with a direct DoH query — that couldn't be fooled by local configuration in the same way.

A note on this build's own DNS pattern

This isn't the only time a search-domain-style resolution quirk has caused confusion in this lab — hartwelg.com subdomains not reliably resolving to internal addresses (needing AdGuard's still-pending split-horizon DNS, Phase 11) has come up as a recurring friction point elsewhere in this build too. Worth keeping in mind as one more data point for whether internal DNS is worth prioritizing sooner rather than later.