Gauge Hartwell
← all write-ups

layered failure

Three Network Namespaces, Three Separate Fixes: The GitLab CI Networking Trilogy

TL;DR: A pipeline's clone step failed to resolve GitLab's own hostname. Fixing that surfaced a second, unrelated networking failure in the build stage. Fixing that surfaced a third. Each fix was correct and each one only solved its own layer — because a single GitLab CI job actually runs across three separate, non-communicating network namespaces stacked on top of each other.

Layer 1 — DNS resolution

The pipeline's clone step failed immediately:

Could not resolve host: gitlab.hartwelg.com

The obvious fix — add an entry to the runner VM's own /etc/hosts — did nothing. The reason became clear on inspection: pipeline job containers run in their own isolated network namespace and don't inherit the runner VM's host-level configuration at all. Docker's runner config needed to be told directly:

[runners.docker]
  extra_hosts = ["gitlab.hartwelg.com:10.0.1.10", "registry.hartwelg.com:10.0.1.10"]

Layer 2 — the wrong protocol

With DNS resolving correctly, the clone step immediately hit a different failure: a connection timeout on port 443. GitLab's actual external_url is plain HTTP internally, but the job container's default clone behavior assumed HTTPS on the resolved hostname. Fixed with an explicit override:

[[runners]]
  clone_url = "http://10.0.1.10"

Layer 3 — the inner Docker-in-Docker daemon

With both of those fixed, the pipeline's build stage specifically — not the clone stage, which was now working fine — still couldn't reach GitLab. The reason: Docker-based builds spin up a second, separate Docker daemon (docker:24-dind) as a service container, and that daemon has its own network namespace, entirely distinct from both the runner VM and the job container that Layer 1's fix already covered. The extra_hosts entry set at the runner level never propagated into this third, inner namespace. Fixed by applying the same override directly to the dind service definition, plus a feature flag that puts job and service containers on a shared network:

services:
  - name: docker:24-dind
    extra_hosts:
      - "gitlab.hartwelg.com:10.0.1.10"
variables:
  FF_NETWORK_PER_BUILD: "true"

Root cause

A single GitLab CI pipeline run isn't one network context — it's three, stacked: the runner VM itself, the isolated job container, and (for any Docker-based build) a second inner Docker daemon with its own namespace again. Each layer needed the identical fix applied separately, because none of them inherit configuration from the layer above.

What this demonstrates

The dangerous moment in this investigation wasn't any individual fix — it was between them. After Layer 1's fix, the clone step succeeded, and "it's fixed" was a genuinely reasonable conclusion to reach at that point. It just happened to be wrong, because the pipeline hadn't actually gotten anywhere near GitLab yet — the first success was DNS resolving, not the pipeline actually working end to end. Verifying all the way through a real, complete build and push — not stopping at the first green checkmark — was what surfaced the second and third layers. A fix that resolves one stage's networking says nothing about whether another, structurally separate stage has the same problem.