mirror of
https://github.com/home-assistant/supervisor.git
synced 2026-09-30 14:39:02 +01:00
main
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f8c7214617 |
Treat containers with corrupt storage metadata as missing (#7186)
* Treat containers with corrupt storage metadata as missing Since Docker 29.4, containers whose RW layer fails to load while the daemon restores its state at startup (e.g. after storage metadata was corrupted by an unclean shutdown) are kept registered so they can still be removed, instead of being dropped (moby/moby#51724). Every other operation on such a container fails with a 500 "RWLayer of container <id> is unexpectedly nil" error, including the inspect at the start of nearly every Supervisor container operation. A single corrupt container record therefore permanently breaks starting, restarting, rebuilding and backing up of the affected add-on, plugin or Home Assistant Core: the start path errors out on the initial state check, before reaching the cleanup that would remove the broken record. Older Docker versions dropped such containers at daemon startup, so Supervisor saw them as missing and recreated them, healing the installation as a side effect. Restore that behavior explicitly by reporting a container as missing when a Docker API call fails with this error. Recovery paths then recreate the container as before, with stop_container's force delete (which no longer inspects first since #7175) removing the broken record along the way. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Broaden corrupt container detection and heal at start/restart The detection matched only one shape of the nil-RW-layer message, the inspect-time one built from the container id. Docker phrases the same condition differently per call site: with the container name instead of its id, with reversed word order on the containerd image store, and bare on Windows. Loosen the pattern to cover them; "unexpectedly nil" appears nowhere else in Docker, so it stays just as narrow. With the containerd image store the inspect of such a container succeeds entirely (the RW layer check in daemon/inspect.go is graph driver only), so none of the inspect-time handling fires: the error only surfaces from the actual start or restart call. Handle it there as well by removing the broken record and reporting the container as missing, so the next start takes the recreate path. Sentry shows this shape in the field (SUPERVISOR-1B6Q). Also soften the corrupt-record comment: the daemon marks the record once while restoring state, for any error loading the RW layer, and a daemon restart may recover it. Removal and recreation is still the safe recovery for Supervisor-managed containers, which hold no state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Document why corrupt container removal stays at the job-locked sites Review feedback asked to remove the corrupt record right where it is detected instead of reporting the container as missing. The record indeed cannot be used for anything until it is removed, but removal has to go by name, and the read paths (is_running() and friends) run outside the job locks that serialize a container's lifecycle. Names are only unique at a given instant — the heal deletes the broken record and then recreates the container under the same name — so a delete-by-name from an unlocked path after a stale inspect could hit the freshly recreated container rather than the broken one. Removal therefore stays where the job lock is held: the inner start and restart failures (added in the previous commit for the containerd image store, where only those calls see the error) and stop_container's cleanup, which every recreate flow reaches moments after detection. The read paths keep reporting the container as missing without acting on it. Say so in the code, and pin the split in tests: the read paths must not remove the record, the locked sites must. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a14729745c |
Parse image references using Docker's domain splitting logic (#7139)
* Parse image references using Docker's domain splitting logic The unsupported container evaluation split an image reference on its first colon to strip the tag. For a registry with a port, that colon belongs to the port, so `myregistry:5000/app:1.0` was reduced to `myregistry` and reported as an unsupported image on the host. Splitting an image reference correctly needs to know where the registry domain ends, and the pieces for that were already there but wired up in a way that could not be reused. IMAGE_REGISTRY_REGEX is a port of Docker's DomainRegexp, which Docker itself uses to validate a domain, not to find one. Placing it in front of the search meant get_registry_from_image() still had to re-derive the answer with the dot/colon/localhost checks that follow the match, and callers that needed the rest of the reference recovered it by slicing off len(registry) + 1 characters. Split the two jobs apart, mirroring Docker's reference implementation: - split_docker_domain() finds the domain the way splitDockerDomain() does, by cutting at the first slash and testing the candidate. It returns the remainder as well, so callers no longer slice by length, and it canonicalizes index.docker.io to docker.io, which lets stored Docker Hub credentials apply to references using the legacy domain. - is_registry_domain() validates a domain against DomainRegexp, which is what the regex is for. Image validation keeps rejecting malformed domains such as ".ghcr.io" through this check. - get_registry_from_image() stays as a wrapper for the callers that only need the domain. With the domain handled, splitting off the tag is a matter of taking the last colon that has no slash after it, which is what Docker's TagRegexp allows. split_image_tag() does that and drops any digest, so a digest-pinned image no longer reads as unsupported either. Move the image reference tests to tests/docker/test_utils.py next to the code under test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Convert remaining image reference parsing to the new helpers Canonicalizing index.docker.io to docker.io made _get_credentials() qualify Docker Hub images from the raw reference, so a reference already carrying a Docker Hub domain gained a second prefix: index.docker.io/org/app was pulled as docker.io/index.docker.io/org/app. Before the canonicalization the legacy domain matched no configured registry and the image was pulled anonymously under its original name, so this only surfaced now, although the same doubling already applied to an explicit docker.io/org/app reference. Qualify from the remainder instead, which covers references without a domain and with either Docker Hub domain. Three more sites split an image reference on its first colon and hit the same problem a registry with a port causes in the unsupported container evaluation: - The image property of DockerInterface reported myreg:5000/supervisor:1.0 as myreg, and a digest reference as name@sha256. - get_latest_version() read the tag of a RepoTags entry, which for myreg:5000/homeassistant:2026.8.0 yielded 5000/homeassistant:2026.8.0. That is not a known version strategy, so every tag was skipped and the lookup failed with "No version found". This is reachable with a user-overridden Core image or a plugin image on a registry with a port. - The Supervisor start tag repair took the image name from a RepoTags entry the same way, leaving myreg as the name to tag. Also align two Docker Hub details with normalize.go: the library/ prefix for official images now applies whenever the resolved registry is Docker Hub, not only when the reference carried no domain, and credential lookup falls back to the legacy hub.docker.com key for an explicit Docker Hub domain as it already did for references without one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Cover qualifying a Docker Hub image with either registry key The credential test only covered an image without a domain and an image with a Docker Hub domain against credentials stored under the official docker.io key. Parametrize over both Docker Hub domains and both registry keys, so the pull name is asserted for a reference carrying the legacy index.docker.io domain while credentials are stored under hub.docker.com as well. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |