mirror of
https://github.com/home-assistant/supervisor.git
synced 2026-10-01 03:48:53 +01:00
* Treat containers with corrupt storage metadata as missing Since Docker 29.4, containers whose RW layer fails to load while the daemon restores its state at startup (e.g. after storage metadata was corrupted by an unclean shutdown) are kept registered so they can still be removed, instead of being dropped (moby/moby#51724). Every other operation on such a container fails with a 500 "RWLayer of container <id> is unexpectedly nil" error, including the inspect at the start of nearly every Supervisor container operation. A single corrupt container record therefore permanently breaks starting, restarting, rebuilding and backing up of the affected add-on, plugin or Home Assistant Core: the start path errors out on the initial state check, before reaching the cleanup that would remove the broken record. Older Docker versions dropped such containers at daemon startup, so Supervisor saw them as missing and recreated them, healing the installation as a side effect. Restore that behavior explicitly by reporting a container as missing when a Docker API call fails with this error. Recovery paths then recreate the container as before, with stop_container's force delete (which no longer inspects first since #7175) removing the broken record along the way. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Broaden corrupt container detection and heal at start/restart The detection matched only one shape of the nil-RW-layer message, the inspect-time one built from the container id. Docker phrases the same condition differently per call site: with the container name instead of its id, with reversed word order on the containerd image store, and bare on Windows. Loosen the pattern to cover them; "unexpectedly nil" appears nowhere else in Docker, so it stays just as narrow. With the containerd image store the inspect of such a container succeeds entirely (the RW layer check in daemon/inspect.go is graph driver only), so none of the inspect-time handling fires: the error only surfaces from the actual start or restart call. Handle it there as well by removing the broken record and reporting the container as missing, so the next start takes the recreate path. Sentry shows this shape in the field (SUPERVISOR-1B6Q). Also soften the corrupt-record comment: the daemon marks the record once while restoring state, for any error loading the RW layer, and a daemon restart may recover it. Removal and recreation is still the safe recovery for Supervisor-managed containers, which hold no state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Document why corrupt container removal stays at the job-locked sites Review feedback asked to remove the corrupt record right where it is detected instead of reporting the container as missing. The record indeed cannot be used for anything until it is removed, but removal has to go by name, and the read paths (is_running() and friends) run outside the job locks that serialize a container's lifecycle. Names are only unique at a given instant — the heal deletes the broken record and then recreates the container under the same name — so a delete-by-name from an unlocked path after a stale inspect could hit the freshly recreated container rather than the broken one. Removal therefore stays where the job lock is held: the inner start and restart failures (added in the previous commit for the containerd image store, where only those calls see the error) and stop_container's cleanup, which every recreate flow reaches moments after detection. The read paths keep reporting the container as missing without acting on it. Say so in the code, and pin the split in tests: the read paths must not remove the record, the locked sites must. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>