Files
supervisor/tests/docker
Stefan AgnerandClaude Fable 5 f8c7214617 Treat containers with corrupt storage metadata as missing (#7186)
* Treat containers with corrupt storage metadata as missing

Since Docker 29.4, containers whose RW layer fails to load while the
daemon restores its state at startup (e.g. after storage metadata was
corrupted by an unclean shutdown) are kept registered so they can still
be removed, instead of being dropped (moby/moby#51724). Every other
operation on such a container fails with a 500 "RWLayer of container
<id> is unexpectedly nil" error, including the inspect at the start of
nearly every Supervisor container operation. A single corrupt container
record therefore permanently breaks starting, restarting, rebuilding and
backing up of the affected add-on, plugin or Home Assistant Core: the
start path errors out on the initial state check, before reaching the
cleanup that would remove the broken record.

Older Docker versions dropped such containers at daemon startup, so
Supervisor saw them as missing and recreated them, healing the
installation as a side effect. Restore that behavior explicitly by
reporting a container as missing when a Docker API call fails with this
error. Recovery paths then recreate the container as before, with
stop_container's force delete (which no longer inspects first since
 #7175) removing the broken record along the way.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Broaden corrupt container detection and heal at start/restart

The detection matched only one shape of the nil-RW-layer message, the
inspect-time one built from the container id. Docker phrases the same
condition differently per call site: with the container name instead of
its id, with reversed word order on the containerd image store, and bare
on Windows. Loosen the pattern to cover them; "unexpectedly nil" appears
nowhere else in Docker, so it stays just as narrow.

With the containerd image store the inspect of such a container succeeds
entirely (the RW layer check in daemon/inspect.go is graph driver only),
so none of the inspect-time handling fires: the error only surfaces from
the actual start or restart call. Handle it there as well by removing
the broken record and reporting the container as missing, so the next
start takes the recreate path. Sentry shows this shape in the field
(SUPERVISOR-1B6Q).

Also soften the corrupt-record comment: the daemon marks the record once
while restoring state, for any error loading the RW layer, and a daemon
restart may recover it. Removal and recreation is still the safe
recovery for Supervisor-managed containers, which hold no state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Document why corrupt container removal stays at the job-locked sites

Review feedback asked to remove the corrupt record right where it is
detected instead of reporting the container as missing. The record
indeed cannot be used for anything until it is removed, but removal has
to go by name, and the read paths (is_running() and friends) run outside
the job locks that serialize a container's lifecycle. Names are only
unique at a given instant — the heal deletes the broken record and then
recreates the container under the same name — so a delete-by-name from
an unlocked path after a stale inspect could hit the freshly recreated
container rather than the broken one.

Removal therefore stays where the job lock is held: the inner start and
restart failures (added in the previous commit for the containerd image
store, where only those calls see the error) and stop_container's
cleanup, which every recreate flow reaches moments after detection. The
read paths keep reporting the container as missing without acting on it.

Say so in the code, and pin the split in tests: the read paths must not
remove the record, the locked sites must.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-10 11:48:01 +02:00
..