2 Commits
Author SHA1 Message Date
Stefan AgnerandClaude Fable 5 f8c7214617 Treat containers with corrupt storage metadata as missing (#7186)
* Treat containers with corrupt storage metadata as missing

Since Docker 29.4, containers whose RW layer fails to load while the
daemon restores its state at startup (e.g. after storage metadata was
corrupted by an unclean shutdown) are kept registered so they can still
be removed, instead of being dropped (moby/moby#51724). Every other
operation on such a container fails with a 500 "RWLayer of container
<id> is unexpectedly nil" error, including the inspect at the start of
nearly every Supervisor container operation. A single corrupt container
record therefore permanently breaks starting, restarting, rebuilding and
backing up of the affected add-on, plugin or Home Assistant Core: the
start path errors out on the initial state check, before reaching the
cleanup that would remove the broken record.

Older Docker versions dropped such containers at daemon startup, so
Supervisor saw them as missing and recreated them, healing the
installation as a side effect. Restore that behavior explicitly by
reporting a container as missing when a Docker API call fails with this
error. Recovery paths then recreate the container as before, with
stop_container's force delete (which no longer inspects first since
 #7175) removing the broken record along the way.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Broaden corrupt container detection and heal at start/restart

The detection matched only one shape of the nil-RW-layer message, the
inspect-time one built from the container id. Docker phrases the same
condition differently per call site: with the container name instead of
its id, with reversed word order on the containerd image store, and bare
on Windows. Loosen the pattern to cover them; "unexpectedly nil" appears
nowhere else in Docker, so it stays just as narrow.

With the containerd image store the inspect of such a container succeeds
entirely (the RW layer check in daemon/inspect.go is graph driver only),
so none of the inspect-time handling fires: the error only surfaces from
the actual start or restart call. Handle it there as well by removing
the broken record and reporting the container as missing, so the next
start takes the recreate path. Sentry shows this shape in the field
(SUPERVISOR-1B6Q).

Also soften the corrupt-record comment: the daemon marks the record once
while restoring state, for any error loading the RW layer, and a daemon
restart may recover it. Removal and recreation is still the safe
recovery for Supervisor-managed containers, which hold no state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Document why corrupt container removal stays at the job-locked sites

Review feedback asked to remove the corrupt record right where it is
detected instead of reporting the container as missing. The record
indeed cannot be used for anything until it is removed, but removal has
to go by name, and the read paths (is_running() and friends) run outside
the job locks that serialize a container's lifecycle. Names are only
unique at a given instant — the heal deletes the broken record and then
recreates the container under the same name — so a delete-by-name from
an unlocked path after a stale inspect could hit the freshly recreated
container rather than the broken one.

Removal therefore stays where the job lock is held: the inner start and
restart failures (added in the previous commit for the containerd image
store, where only those calls see the error) and stop_container's
cleanup, which every recreate flow reaches moments after detection. The
read paths keep reporting the container as missing without acting on it.

Say so in the code, and pin the split in tests: the read paths must not
remove the record, the locked sites must.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-10 11:48:01 +02:00
Stefan AgnerandClaude Opus 5 a14729745c Parse image references using Docker's domain splitting logic (#7139)
* Parse image references using Docker's domain splitting logic

The unsupported container evaluation split an image reference on its first
colon to strip the tag. For a registry with a port, that colon belongs to the
port, so `myregistry:5000/app:1.0` was reduced to `myregistry` and reported as
an unsupported image on the host.

Splitting an image reference correctly needs to know where the registry domain
ends, and the pieces for that were already there but wired up in a way that
could not be reused. IMAGE_REGISTRY_REGEX is a port of Docker's DomainRegexp,
which Docker itself uses to validate a domain, not to find one. Placing it in
front of the search meant get_registry_from_image() still had to re-derive the
answer with the dot/colon/localhost checks that follow the match, and callers
that needed the rest of the reference recovered it by slicing off
len(registry) + 1 characters.

Split the two jobs apart, mirroring Docker's reference implementation:

- split_docker_domain() finds the domain the way splitDockerDomain() does, by
  cutting at the first slash and testing the candidate. It returns the
  remainder as well, so callers no longer slice by length, and it canonicalizes
  index.docker.io to docker.io, which lets stored Docker Hub credentials apply
  to references using the legacy domain.
- is_registry_domain() validates a domain against DomainRegexp, which is what
  the regex is for. Image validation keeps rejecting malformed domains such as
  ".ghcr.io" through this check.
- get_registry_from_image() stays as a wrapper for the callers that only need
  the domain.

With the domain handled, splitting off the tag is a matter of taking the last
colon that has no slash after it, which is what Docker's TagRegexp allows.
split_image_tag() does that and drops any digest, so a digest-pinned image no
longer reads as unsupported either.

Move the image reference tests to tests/docker/test_utils.py next to the code
under test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Convert remaining image reference parsing to the new helpers

Canonicalizing index.docker.io to docker.io made _get_credentials() qualify
Docker Hub images from the raw reference, so a reference already carrying a
Docker Hub domain gained a second prefix: index.docker.io/org/app was pulled as
docker.io/index.docker.io/org/app. Before the canonicalization the legacy domain
matched no configured registry and the image was pulled anonymously under its
original name, so this only surfaced now, although the same doubling already
applied to an explicit docker.io/org/app reference. Qualify from the remainder
instead, which covers references without a domain and with either Docker Hub
domain.

Three more sites split an image reference on its first colon and hit the same
problem a registry with a port causes in the unsupported container evaluation:

- The image property of DockerInterface reported myreg:5000/supervisor:1.0 as
  myreg, and a digest reference as name@sha256.
- get_latest_version() read the tag of a RepoTags entry, which for
  myreg:5000/homeassistant:2026.8.0 yielded 5000/homeassistant:2026.8.0. That
  is not a known version strategy, so every tag was skipped and the lookup
  failed with "No version found". This is reachable with a user-overridden Core
  image or a plugin image on a registry with a port.
- The Supervisor start tag repair took the image name from a RepoTags entry the
  same way, leaving myreg as the name to tag.

Also align two Docker Hub details with normalize.go: the library/ prefix for
official images now applies whenever the resolved registry is Docker Hub, not
only when the reference carried no domain, and credential lookup falls back to
the legacy hub.docker.com key for an explicit Docker Hub domain as it already
did for references without one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Cover qualifying a Docker Hub image with either registry key

The credential test only covered an image without a domain and an image with a
Docker Hub domain against credentials stored under the official docker.io key.
Parametrize over both Docker Hub domains and both registry keys, so the pull
name is asserted for a reference carrying the legacy index.docker.io domain
while credentials are stored under hub.docker.com as well.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 09:50:00 +02:00