mirror of
https://github.com/home-assistant/supervisor.git
synced 2026-10-01 13:53:19 +01:00
multicast-plugin-enabled-option
105
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f8c7214617 |
Treat containers with corrupt storage metadata as missing (#7186)
* Treat containers with corrupt storage metadata as missing Since Docker 29.4, containers whose RW layer fails to load while the daemon restores its state at startup (e.g. after storage metadata was corrupted by an unclean shutdown) are kept registered so they can still be removed, instead of being dropped (moby/moby#51724). Every other operation on such a container fails with a 500 "RWLayer of container <id> is unexpectedly nil" error, including the inspect at the start of nearly every Supervisor container operation. A single corrupt container record therefore permanently breaks starting, restarting, rebuilding and backing up of the affected add-on, plugin or Home Assistant Core: the start path errors out on the initial state check, before reaching the cleanup that would remove the broken record. Older Docker versions dropped such containers at daemon startup, so Supervisor saw them as missing and recreated them, healing the installation as a side effect. Restore that behavior explicitly by reporting a container as missing when a Docker API call fails with this error. Recovery paths then recreate the container as before, with stop_container's force delete (which no longer inspects first since #7175) removing the broken record along the way. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Broaden corrupt container detection and heal at start/restart The detection matched only one shape of the nil-RW-layer message, the inspect-time one built from the container id. Docker phrases the same condition differently per call site: with the container name instead of its id, with reversed word order on the containerd image store, and bare on Windows. Loosen the pattern to cover them; "unexpectedly nil" appears nowhere else in Docker, so it stays just as narrow. With the containerd image store the inspect of such a container succeeds entirely (the RW layer check in daemon/inspect.go is graph driver only), so none of the inspect-time handling fires: the error only surfaces from the actual start or restart call. Handle it there as well by removing the broken record and reporting the container as missing, so the next start takes the recreate path. Sentry shows this shape in the field (SUPERVISOR-1B6Q). Also soften the corrupt-record comment: the daemon marks the record once while restoring state, for any error loading the RW layer, and a daemon restart may recover it. Removal and recreation is still the safe recovery for Supervisor-managed containers, which hold no state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Document why corrupt container removal stays at the job-locked sites Review feedback asked to remove the corrupt record right where it is detected instead of reporting the container as missing. The record indeed cannot be used for anything until it is removed, but removal has to go by name, and the read paths (is_running() and friends) run outside the job locks that serialize a container's lifecycle. Names are only unique at a given instant — the heal deletes the broken record and then recreates the container under the same name — so a delete-by-name from an unlocked path after a stale inspect could hit the freshly recreated container rather than the broken one. Removal therefore stays where the job lock is held: the inner start and restart failures (added in the previous commit for the containerd image store, where only those calls see the error) and stop_container's cleanup, which every recreate flow reaches moments after detection. The read paths keep reporting the container as missing without acting on it. Say so in the code, and pin the split in tests: the read paths must not remove the record, the locked sites must. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9adf131901 |
Skip redundant inspect calls when constructing Docker container handles (#7175)
* Skip redundant inspect calls when constructing Docker container handles Several places in the Docker layer called containers.get(id), which performs an inspect API call, purely to construct a DockerContainer handle immediately followed by another call (show(), exec(), attach()) that either performs its own inspect or doesn't need one at all. Switch these spots to containers.container(id), which builds the same handle locally with no I/O, removing the redundant round-trip. Affected call sites: container_is_initialized(), stop_container() and container_run_inside() in docker/manager.py, _get_container() and _attach() in docker/interface.py, attach()/retag()/update_start_tag() in docker/supervisor.py, and write_stdin() in docker/app.py. Since containers.container()-built handles don't carry the real container ID (unlike get(), which populates it from the inspect response), _attach() and DockerSupervisor.attach() now read the ID from the inspect payload (self._meta["Id"]) instead of the handle's .id property. This intentionally excludes container_stats(), which has its own one-shot inspect+stats opportunity being handled separately in #7168. Follow-up to review discussion in #7168: https://github.com/home-assistant/supervisor/pull/7168#discussion_r3883548019 * Fix tests broken by containers.container() switch Several tests simulated Docker errors/state by mocking containers.get() (the old call site), but the production code now builds container handles via the sync, no-I/O containers.container() instead and only calls get() directly from the small set of call sites intentionally left unconverted (DockerApp's legacy-name migration in attach(), and container_stats() which is handled by #7168). Retarget these tests to mock containers.container() (or both, where a single test exercises both an app's attach() and a plugin/core's attach()) so the simulated failures actually reach the code path they're meant to test. Also add the missing "Id" key to some mocked container inspect payloads in test_check_docker_config.py: _attach() now reads the container ID from the inspect response instead of the handle's .id, so tests need that field present. * Skip unnecessary inspect before stopping container A reviewer pointed out that calling show() before stop() in stop_container() is redundant I/O. Docker returns a 304 (not an error) if the container is already stopped, 404 if it does not exist, so we can call stop() directly and handle those responses without an extra inspect call. * Centralize container fetch error handling in DockerSupervisor attach(), retag() and update_start_tag() each duplicated a try/except block around containers.container(name).show(). Reuse the existing _get_container() helper from DockerInterface instead, raising DockerNotFound explicitly when the container is missing. |
||
|
|
2d153ac497 |
Add one-shot mode and docker-given totals to stats API (#7168)
* Add one-shot mode and docker-given totals to stats API - Add one-shot query option to container stats to skip the calculated CPU percentage window, returning docker-given totals directly. - Refactor container_stats() with translatable, structured Docker errors and timeout handling instead of generic Docker exceptions. - Translate low-level Docker errors into component-flavored errors (NotRunningError/StatsTimeoutError/UnknownError) for Home Assistant, Supervisor, and each plugin (Audio, Cli, CoreDNS, Multicast, Observer) as well as apps, mirroring the existing App.stats() pattern for consistent, informative API responses. - Add a shared api_return_stats() helper to remove duplicated stats handling logic across the 8 stats API endpoints. - DockerContainerNotFoundError now inherits from DockerNotFound. * Contain aiodocker _query access and split v1/v2 stats response model - Contain the reach into aiodocker's protected _query internals (needed for one-shot stats until aio-libs/aiodocker#1054 lands) to a single DockerAPI._query_one_shot_stats() helper. - Split the stats API response model between v1 (legacy, windowed CPU percentage still available via opt-in one_shot query string) and v2 (always one-shot, no query string, cpu_percent dropped since it would never carry a meaningful value). Both share the same api_return_stats(stats, legacy=...) helper so the response model stays in one place. - DockerAPI.container_stats() keeps its existing one_shot keyword only; the legacy/v1 vs v2 distinction is handled entirely in the API layer. * Detect stopped/restarting containers in one-shot stats response Docker returns a stub response containing only the container's id/name (no cpu_stats/memory_stats/networks data) for a stopped or restarting container instead of an error, even in one-shot mode - the not-running check in daemon.ContainerStats happens before the one-shot branch, so requesting one-shot stats doesn't change this. Detect the missing online_cpus field (only ever present while running) and raise DockerContainerNotRunningError instead of returning a payload that DockerStats would otherwise turn into a successful all-zero response. * Add tests for stats error paths introduced in one-shot/totals PR - Cover DockerAPI.container_stats/_query_one_shot_stats edge cases: empty one-shot response, empty windowed response, and the raw _query_one_shot_stats helper. - Add stats() error-translation tests (not-running, timeout, unknown) for the audio, cli, multicast and observer plugins, none of which had any coverage. - Add stats endpoint tests (v1 + v2) for the cli, audio, dns, multicast and observer REST APIs, plus a v1 success-path test for the apps stats endpoint. These close the codecov-flagged gaps for newly introduced code in this PR. * Unify one-shot and windowed container_stats() code paths Both modes now share the same error handling and not-running detection: - The windowed path no longer inspects the container first to check its running state. Docker returns the same id/name-only stub response (no online_cpus in cpu_stats) for a stopped or restarting container in windowed mode as it does in one-shot mode, so the existing online_cpus presence check now covers both, closing a race window where a container stopping between the inspect and the stats call would previously reach the DockerStats parser with an incomplete payload. - The windowed path builds its container handle via containers.container(), which constructs a DockerContainer from the name with no I/O, instead of containers.get(), which performs an unnecessary inspect call just to fetch the same information the stats call already returns. * Fix protected access pylint issue |
||
|
|
a14729745c |
Parse image references using Docker's domain splitting logic (#7139)
* Parse image references using Docker's domain splitting logic The unsupported container evaluation split an image reference on its first colon to strip the tag. For a registry with a port, that colon belongs to the port, so `myregistry:5000/app:1.0` was reduced to `myregistry` and reported as an unsupported image on the host. Splitting an image reference correctly needs to know where the registry domain ends, and the pieces for that were already there but wired up in a way that could not be reused. IMAGE_REGISTRY_REGEX is a port of Docker's DomainRegexp, which Docker itself uses to validate a domain, not to find one. Placing it in front of the search meant get_registry_from_image() still had to re-derive the answer with the dot/colon/localhost checks that follow the match, and callers that needed the rest of the reference recovered it by slicing off len(registry) + 1 characters. Split the two jobs apart, mirroring Docker's reference implementation: - split_docker_domain() finds the domain the way splitDockerDomain() does, by cutting at the first slash and testing the candidate. It returns the remainder as well, so callers no longer slice by length, and it canonicalizes index.docker.io to docker.io, which lets stored Docker Hub credentials apply to references using the legacy domain. - is_registry_domain() validates a domain against DomainRegexp, which is what the regex is for. Image validation keeps rejecting malformed domains such as ".ghcr.io" through this check. - get_registry_from_image() stays as a wrapper for the callers that only need the domain. With the domain handled, splitting off the tag is a matter of taking the last colon that has no slash after it, which is what Docker's TagRegexp allows. split_image_tag() does that and drops any digest, so a digest-pinned image no longer reads as unsupported either. Move the image reference tests to tests/docker/test_utils.py next to the code under test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Convert remaining image reference parsing to the new helpers Canonicalizing index.docker.io to docker.io made _get_credentials() qualify Docker Hub images from the raw reference, so a reference already carrying a Docker Hub domain gained a second prefix: index.docker.io/org/app was pulled as docker.io/index.docker.io/org/app. Before the canonicalization the legacy domain matched no configured registry and the image was pulled anonymously under its original name, so this only surfaced now, although the same doubling already applied to an explicit docker.io/org/app reference. Qualify from the remainder instead, which covers references without a domain and with either Docker Hub domain. Three more sites split an image reference on its first colon and hit the same problem a registry with a port causes in the unsupported container evaluation: - The image property of DockerInterface reported myreg:5000/supervisor:1.0 as myreg, and a digest reference as name@sha256. - get_latest_version() read the tag of a RepoTags entry, which for myreg:5000/homeassistant:2026.8.0 yielded 5000/homeassistant:2026.8.0. That is not a known version strategy, so every tag was skipped and the lookup failed with "No version found". This is reachable with a user-overridden Core image or a plugin image on a registry with a port. - The Supervisor start tag repair took the image name from a RepoTags entry the same way, leaving myreg as the name to tag. Also align two Docker Hub details with normalize.go: the library/ prefix for official images now applies whenever the resolved registry is Docker Hub, not only when the reference carried no domain, and credential lookup falls back to the legacy hub.docker.com key for an explicit Docker Hub domain as it already did for references without one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Cover qualifying a Docker Hub image with either registry key The credential test only covered an image without a domain and an image with a Docker Hub domain against credentials stored under the official docker.io key. Parametrize over both Docker Hub domains and both registry keys, so the pull name is asserted for a reference carrying the legacy index.docker.io domain while credentials are stored under hub.docker.com as well. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c119e0cff1 |
Rename Docker app container and builder names from addon_ to app_ (#7024)
* Rename Docker app container and builder names from addon_ to app_ * Update container state event tests to app_ names * Migrate legacy app container names on attach * Fix error handling * Fixes from feedback * Fix pytest --------- Co-authored-by: Stefan Agner <stefan@agner.ch> |
||
|
|
f49b81840e |
Add app-based mapping options for app configs (#6992)
* Rename addon map options to app equivalents * Add tests for apps/addons default mount targets * Fixes from feedback * Fix tests and clean up validation logic a bit * Update supervisor/apps/validate.py * Apply suggestions from code review Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> --------- Co-authored-by: Stefan Agner <stefan@agner.ch> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> |
||
|
|
5bc795d2f2 |
Migrate simple attrs classes to stdlib dataclasses (#7005)
* Migrate simple attrs classes to stdlib dataclasses Several plain data-holder classes still used the attrs library while newer code in the project uses stdlib dataclasses. Convert EventListener, Issue, Suggestion, HealthChanged, SupportedChanged, Device, Message, HostEntry, ServiceInfo, WhoamiData and the scheduler _Task to @dataclass, and switch the attr.evolve/attr.asdict call sites to dataclasses.replace/asdict. Field semantics are preserved: eq=False maps to compare=False, hash=False and default factories map directly, and the frozen/slots flags are kept, so equality, hashing and copy behavior are unchanged. The jobs module keeps using attrs for its validators and setter hooks, which have no stdlib dataclass equivalent. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Add slots to remaining dataclasses, freeze HostEntry Address review feedback: ServiceInfo, Message and HostEntry kept the flags of their attrs originals, which did not use slots. Enable slots on all three. HostEntry instances are never mutated after creation, so it can also be frozen. Message cannot be frozen because Discovery.send updates the config field of an existing message when an app re-sends a discovery message with changed configuration. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5c65da8a19 |
Refine Docker timeout handling and add regression tests (#6970)
* Refine Docker timeout handling and add regression tests * Add network exception path tests for Docker manager * Add timeout-path tests for Docker app and network * Remove unnecessary timeout=None overrides from pull calls and tests |
||
|
|
098867bbe2 |
Handle Docker status-check failures in Home Assistant API paths (#6955)
* Handle Docker status-check failures in Home Assistant API paths * Remove empty timeout msg from error |
||
|
|
81e235376e |
Fix typos repo-wide and add codespell pre-commit hook (#6949)
* Fix typos in comments, docstrings and log messages Correct 39 spelling mistakes across comments, docstrings and log/error message strings throughout the package (e.g. "conection" -> "connection", "Incomming" -> "Incoming", "Rasie" -> "Raise"). All changes are confined to human-readable text; no identifiers, attributes or D-Bus contracts are touched, so there is no behavior change. Found with codespell. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Fix typos in tests and CI workflow Correct spelling mistakes in test comments, docstrings and data, plus one in the builder workflow, so the whole tree is clean for the codespell hook added next. The assertion in test_network_manager.py is updated to match the corrected "Unknown error while processing" log message in the source. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Add codespell pre-commit hook Wire up codespell so spelling mistakes in comments, docstrings and strings are caught automatically. The vendored frontend panel is excluded, and "hass" and "astroid" are added to the ignore list as known false positives (the Home Assistant abbreviation and the pylint dependency package). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review feedback Improve grammar in several of the touched comments and docstrings: use the plural "ignore conditions" for the list-returning property, add the missing auxiliary verb and fix agreement in the timezone-filter comment, fix "backups ... use" agreement, and reword "underlay" to "underlying" in the arch module docstring. Also drop the "*.json" skip from the codespell hook. It was carried over from another project but is unnecessary here (all tracked JSON is clean), and skipping it would needlessly leave translation and data JSON unchecked. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Reword onboarding comment "overflight" was a literal calque of the German "überflogen"; use the idiomatic "skimmed through" instead. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
dc6a77507b |
fix(docker): restore add-on device access after USB re-enumeration (#6877)
* fix(docker): register hw listener and match by-id paths for options-based devices Two bugs caused a crash loop when a USB device re-enumerates to a different minor number (e.g. ttyACM0→ttyACM1) after a HAOS reboot: 1. _hw_listener was only registered when addon.static_devices was non-empty. Addons that expose a device via the options schema (e.g. Z-Wave JS `device:` option) never had the listener registered, so add_devices_allowed was never called when the device reappeared at a new minor. 2. _hardware_events matched only device.path and device.sysfs against static_devices. When static_devices (or the new options path) contains a by-id symlink, the match always failed because by-id paths live in device.links. Fix: extend the listener registration condition to also cover addon.devices (options-based), and expand the path-matching set to include device.links so by-id paths resolve correctly. For options-based devices, compare the incoming Device against addon.devices (which re-evaluates options.json against the live hardware list, picking up the new minor number automatically). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(docker): use option_device_paths for cheap by-id hw event matching Refactor _hardware_events to avoid per-event full options validation (including pwnd hashing). Introduce AppOptions.extract_device_paths and AppModel.option_device_paths to extract raw device paths from options without resolving against live hardware. Use set-intersection against {device.path, device.sysfs, *device.links} so by-id symlinks match correctly after re-enumeration for both static and options-based devices. Update test to use real schema/options setup and simulate a minor-number change (ttyACM0→ttyACM1) with a stable by-id symlink. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test: improve hw listener test coverage and add policy check Address PR review feedback: - Add hardware policy check in _hardware_events to prevent bypassing access restrictions on hotplug events (follows same pattern as startup cgroup setup) - Fix test_app_options_device_hw_listener to properly simulate USB re-enumeration with different minor numbers (166:0 → 166:1) - Add test_app_options_device_policy_check to verify policy enforcement for options-based devices - Update TEST_HW_DEVICE with realistic major/minor attributes (166:0 for tty) Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com> * test: mock HostFeature.OS_AGENT in hardware event tests The _hardware_events method has a @Job decorator with conditions=[JobCondition.OS_AGENT], which checks if HostFeature.OS_AGENT is in sys_host.features. Without mocking this, the job conditions fail and the hardware event handler is never invoked, causing add_devices_allowed to not be called and tests to fail. Add patch.object(type(coresys.host), "features", ...) to all four hardware event tests to ensure the OS_AGENT job condition is met. Fixes test failures: - test_app_new_device (all 6 parametrized cases) - test_app_new_device_no_haos - test_app_options_device_hw_listener - test_app_options_device_policy_check Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com> * test: fix TEST_DEV_PATH to match TEST_HW_DEVICE.path TEST_DEV_PATH was set to /dev/ttyAMA0 but TEST_HW_DEVICE.path is /dev/ttyACM0. This mismatch would cause the dev_path=TEST_DEV_PATH parametrized test cases to fail because the hardware event handler checks if the device path intersects with the app's allowed devices, and "/dev/ttyAMA0" != "/dev/ttyACM0". Update TEST_DEV_PATH from /dev/ttyAMA0 to /dev/ttyACM0 to match TEST_HW_DEVICE. Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com> * test: mock option_device_paths in test_app_options_device_hw_listener The test sets up schema and options but option_device_paths property may not be working as expected in the test environment. Add an explicit mock for option_device_paths to ensure it returns the by-id path, guaranteeing that: 1. The hardware listener is registered (checks option_device_paths at registration) 2. The device path matching works correctly in _hardware_events This ensures the test properly validates that hardware events are processed for options-based devices after re-enumeration. Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com> * test: make policy check test actually exercise the policy guard test_app_options_device_policy_check set the device option via persist["options"], but option_device_paths reads the merged options and does not pick that override up during the test, so it returned an empty set. The hardware event therefore failed the path-match guard and returned before reaching the allowed_for_access check. The assert_not_called() assertion then passed regardless of the policy outcome -- it would still pass if the policy guard were removed entirely. Mock option_device_paths to return the configured by-id path (mirroring test_app_options_device_hw_listener) so the event device matches and execution actually reaches the policy guard the test is meant to verify. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test: add unit test for AppOptions.extract_device_paths The integration tests exercise extract_device_paths only through a mocked option_device_paths property, so the schema-walking logic introduced for the hardware-event matching had no direct coverage. Add a unit test that drives every schema shape the recursion handles -- flat, optional, filtered, list, nested dict and list of dicts -- and asserts that non-device options, unset keys and empty values are skipped, without requiring the devices to exist in hardware. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Stefan Agner <stefan@agner.ch> |
||
|
|
7870cd786f |
Detect landingpage by io.hass.type label instead of version (#6935)
The landingpage image now stamps a real Core version into its io.hass.version label rather than the sentinel "landingpage" string. Supervisor decided whether Core was still just a landingpage by comparing that version against the LANDINGPAGE constant, so with the new label it mistook the landingpage for an installed Core. Restarting the Supervisor mid-install then stranded the system on the landingpage until the next OS reboot, with no install job scheduled and the watchdog hot-looping on the missing Core auth API. Override the version resolution for the Home Assistant container so a landingpage image (io.hass.type == "landingpage") always reports LANDINGPAGE, keeping every existing version check working unchanged while leaving io.hass.version free to carry the real version. Fixes #6934 |
||
|
|
5ea7908184 |
Drop obsolete UNIX_SOCKET_CORE_API feature flag (#6908)
Unix socket communication between Supervisor and Home Assistant Core was initially gated behind the UNIX_SOCKET_CORE_API feature flag, then enabled by default for Core 2026.5.1+ while older supported versions still required the flag. Now that the transport has settled and is on by default, the flag no longer serves a purpose. Remove the feature flag and the CORE_UNIX_SOCKET_DEFAULT_VERSION split, collapsing supports_unix_socket back to a single check: the Unix socket is used for any non-landingpage Core version at or above CORE_UNIX_SOCKET_MIN_VERSION. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c2b5482b22 |
Bump aiohttp from 3.13.5 to 3.14.0 (#6902)
* Bump aiohttp from 3.13.5 to 3.14.0 --- updated-dependencies: - dependency-name: aiohttp dependency-version: 3.14.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> * Adjust API layer for aiohttp 3.14 aiohttp 3.14 raises NotAppKeyWarning when a plain string is used as a request storage key. Convert REQUEST_FROM to a web.RequestKey instance so request[REQUEST_FROM] no longer triggers the warning (which the test suite escalates to an error, failing every authenticated endpoint). It is typed as RequestKey[Any] to preserve the current access semantics: handlers store different origins (App, Home Assistant, host, observer) and narrow the value to the concrete type they expect. aiohttp 3.14 also widened the request.post() return type to include bytearray. Update the _process_dict annotation accordingly to satisfy mypy. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Use encode_basic_auth for registry token requests aiohttp 3.14 deprecates the BasicAuth constructor and the client auth= request parameter (both removed in aiohttp 4.0), each emitting a DeprecationWarning at runtime. The registry manifest fetcher hit both when requesting a token with stored credentials. Switch to aiohttp.encode_basic_auth() and pass the result via an Authorization header instead. Add a test covering the credentials path, which the existing tests skipped by mocking _get_auth_token. It asserts the Authorization header is sent and, since the suite escalates warnings to errors, guards against the deprecations returning. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Stefan Agner <stefan@agner.ch> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
355396aeab |
Migrate addon/addons config paths and schema names to app/apps (#6865)
* Migrate config file and directory paths from addons to apps
- Rename addons.json -> apps.json (FILE_HASSIO_APPS constant)
- Rename addons/{core,data,local,git} -> apps/{core,data,local,git}
- Rename addon_configs -> app_configs
Backwards compatibility: on startup, Supervisor checks for legacy
paths and renames them if the new paths don't already exist.
- addons.json migration runs in AppManager.load_config (executor)
- Directory migrations run in bootstrap before initialize_system (executor)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Rename SCHEMA_ADDON(S)_* constants to SCHEMA_APP(S)_* in apps/validate.py
Update all references in supervisor and tests.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Fix remaining test references to legacy addons/* paths
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Opportunistic remove of addons dir since it should be empty post migration
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
||
|
|
d3028d7bfc |
Enable flake8-pyi, flake8-return, flake8-raise ruff rules (#6861)
Enable the PYI, RET and RSE ruff rule sets and fix the resulting violations across the codebase. The pylint no-else-* checks (RET505-508) are now disabled since ruff covers them. The cleanups are mechanical: - RSE102: drop empty parentheses from `raise Exception()` when no arguments are passed. - RET505-508: drop `else` branches that follow a `return`, `raise`, `continue` or `break`, flattening control flow. - RET502/504: add explicit return values and remove redundant assign-then-return patterns. - PYI030/032/041: tidy up type annotations (collapse literal unions, use `object` for `__eq__`/`__ne__`, drop redundant numeric unions). |
||
|
|
ed91b18c4b |
tests: enable flake8-pytest-style (PT) ruff rules (#6857)
* tests: enable flake8-pytest-style (PT) ruff rules Enable the `PT` ruff rule set and fix the resulting violations across the test suite: - PT006: pass parametrize argument names as tuples instead of a single comma-separated string. - PT022: switch fixtures that have no teardown from `yield` to `return` so the lack of cleanup is obvious at a glance. - PT011: add `match=` to broad `pytest.raises(ValueError)` blocks so the expected error is anchored to a specific message. - PT012: hoist setup (patches, branching) out of `pytest.raises()` blocks so only the call that is expected to raise remains inside. - PT013: replace `from pytest import X` with `import pytest` and access attributes via the module. - PT015: replace `try/except` + `assert False` patterns with `pytest.raises(...)`. - PT017: replace `assert` on exceptions inside `except` blocks with `pytest.raises(...) as exc_info` and assert on `exc_info.value`. No behavioral changes to the tests; the full suite still passes. * tests: address review feedback on PT ruff rule enablement - Fix fixture return-type annotations after switching `yield` to `return` in tests/conftest.py: drop the `Generator[...]`/`AsyncGenerator[...]` wrapper for `dns_manager_service`, `supervisor_internet`, `websession`, and `mock_update_data` so the annotation matches what the fixture actually returns. - Correct the return-type annotation of `fixture_ip6config_service` from `IP4ConfigService` to `IP6ConfigService`. - Fix recurring "excepiton" typo in tests/utils/test_exception_helper.py. * tests: verify backup cleanup on permission error After `test_new_backup_permission_error` raises `BackupPermissionError`, assert that no tarfile was left behind and `tmp_path` is empty. The previous version only checked that the exception was raised, which missed any regression where a partial tarfile would survive the failed create. * tests: rename DNS_GOOD_V6 to DNS_V6_UNSUPPORTED The constant was named "good" but its tests assert that the URLs are rejected by the DNS validator. The IPv6 URLs are well-formed but currently rejected because IPv6 doesn't work with the Docker network (see `dns_url` in supervisor/validate.py). Rename the constant and the related test to make the intent obvious. |
||
|
|
7ecfe42602 |
apps: log container exit code when app exits non-zero (#6848)
* apps: log container exit code when app exits non-zero Issue #6840 reports that stopping an app whose process exits 143 (SIGTERM default disposition) leaves the app in AppState.ERROR. ERROR is the right state for that — Docker itself treats any non-zero exit as a failure (e.g. `--restart on-failure`), and 143 specifically means the SIGTERM grace period was wasted because the app never installed a handler. But Supervisor previously logged nothing about it, leaving authors with no hint that their image is misbehaving. Plumb the exit code through DockerContainerStateEvent and log it from App.container_state_changed on transitions to FAILED: a warning for 143 nudging the author to trap SIGTERM and exit 0, and an error for any other non-zero code (crashes, SIGKILL after grace, app's own error exit). Refactor _container_state_from_model to return (state, exit_code) so the docker event monitor and DockerInterface.attach feed the same exit code through one code path instead of re-reading State.ExitCode in the caller. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * apps: address review feedback on exit-code logging - Replace bare 143 with EXIT_CODE_SIGTERM_DEFAULT (128 + signal.SIGTERM) in supervisor/docker/const.py so the reasoning is documented in code, not just in the log string. - Stop populating exit_code on STOPPED transitions. Previously the refactor made DockerInterface.attach emit exit_code=0 for cleanly stopped containers, while the monitor only emitted an exit code for abnormal exits. Align both paths so exit_code is only set on FAILED. - Add test_app_failed_logs_exit_code covering the three new branches (warning on 143, error on other non-zero, silent when None) and extend test_attach_existing_container to assert the event's exit_code field per state. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docker/monitor: flatten exit_code branch to satisfy pylint The previous if/else inside the `die` branch pushed the function over pylint's too-many-nested-blocks threshold (6/5). Collapse it back into a pair of conditional expressions: container_state via ternary on the exit code, exit_code via `die_exit_code or None` so 0 stays None. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Update supervisor/apps/app.py Co-authored-by: Mike Degatano <michael.degatano@gmail.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Co-authored-by: Mike Degatano <michael.degatano@gmail.com> |
||
|
|
f8880a72be |
Rename addon/addons to app/apps in filenames and imports (#6837)
* Rename addon/addons to app/apps in filenames and imports Continues the addon→app terminology migration (#6786). Renames all source files, test files, fixture files, and directories that contained 'addon'/'addons' in their names, and updates all imports accordingly. Resolution check files in supervisor/resolution/checks/ that were renamed override the slug property to preserve the existing API contract (slugs are exposed via the resolution info API and used to run checks by name). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Rename add-on.json fixture --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
c772a9bbb0 |
Replace fixed-duration sleeps after bus events with gather (#6803)
* Replace fixed-duration sleeps after bus events with gather Several tests use ``await asyncio.sleep(...)`` to "wait for the listener to run" after firing a bus event. The fixed duration is real wall-clock time and the wait can be indeterministic — if the handler chain happens to need slightly more time on a busy CI runner, the assertion races the handler. ``Bus.fire_event`` returns the listener tasks since #6252; capture and ``await asyncio.gather(*tasks)`` instead of sleeping. Touches test_bus.py (the bus tests were poking scheduling instead of verifying their assertions), test_home_assistant_watchdog.py, test_plugin_base.py, addons/test_manager.py, docker/test_addon.py, and test_store_execute_reload.py. Other cleanups in the same spirit: - ``_fire_test_event`` in addons/test_addon.py becomes ``async def`` and gathers the listener tasks itself, so its 17 call sites collapse to a single ``await _fire_test_event(...)``. - The two test_store_execute_reload.py sites that used the private ``_update_connectivity()`` helper are reworked to set the cached connectivity flag directly and fire the event themselves so they can gather the listener tasks the same way. - The two ``sleep(1)`` post-pull drains in docker/test_interface.py collapse to ``sleep(0)`` (handler tasks are already gathered inside pull_image), saving ~2s. - The ``sleep(0.01)`` waits inside ``container_events()`` task bodies (api/test_addons.py, api/test_store.py, backups/test_manager.py) are just one-yield-to-the-parent and become ``sleep(0)``. Switching to ``gather`` exposes a few latent test mocks that were silently swallowing TypeErrors as background-task failures before: - ``CGroup.add_devices_allowed`` is ``async def`` but was patched as a plain MagicMock in docker/test_addon.py — now patched via ``new_callable=AsyncMock``. - The watchdog does ``await (await self.start())`` / ``await (await self.restart())`` because ``App.start`` / ``App.restart`` return ``asyncio.Task``. The mocks in addons/test_addon.py (test_app_watchdog, test_watchdog_on_stop, test_watchdog_during_attach) needed ``AsyncMock(return_value=<settled future>)`` to mirror that shape rather than a plain MagicMock. * Factor bus.fire_event + gather pattern into a helper Per review feedback, the ``await asyncio.gather(*coresys.bus.fire_event(...))`` incantation was scattered across many call sites. Add ``tests.common.fire_bus_event`` that takes the coresys, event and data, fires the event and awaits the spawned listener tasks. Convert all matching sites to use it, including the ``_fire_test_event`` wrapper in addons/test_addon.py which now just builds the ``DockerContainerStateEvent`` and delegates. |
||
|
|
0de6d25fed |
Drop legacy test classes in favor of module-level functions (#6796)
Per CLAUDE.md, plain test_* functions are the project style; class- based test grouping is considered legacy. Convert the 24 test methods in test_pull_progress.py (TestLayerProgress, TestImagePullProgress) to module-level functions — none of them used self, so the rewrite is mechanical. Also rename three helper classes whose names accidentally matched pytest's Test* collection pattern, even though they are fakes/fixtures rather than test cases: - TestAddon -> FakeApp (data holder used as a fake App in pwned tests) - TestDockerInterface -> FakeDockerInterface (fixture/inner helper in docker tests) The two DBusServiceMock subclasses named TestInterface already had __test__ = False and are left alone. |
||
|
|
97bc19d4b3 |
Detect container registry rate limits uniformly (#6732)
* Detect container registry rate limits uniformly
Container registry rate limits reach Supervisor in three distinct shapes:
1. HTTP 429 from the daemon - recognised today, but the exception and
resolution issue are hardcoded to Docker Hub. Since Core/Supervisor/
plugin images all live on ghcr.io now, virtually every 429 we see in
the field is actually a GHCR throttle that we mislabel. The biggest
Sentry issue (SUPERVISOR-16BK) has >115k events / >93k users, all
pulling a ghcr.io image, yet each user is told to "log into
Docker Hub".
2. HTTP 500 with 'toomanyrequests' in the body - not recognised. Docker
daemons before 28.3.0 wrap upstream 429s as 500 (fixed upstream by
moby/moby 23fa0ae74a, "Cleanup http status error checks"). The large
fleet on older daemons still produces this shape.
3. JSON error event during a streaming pull - not recognised. Once the
daemon starts writing the 200 OK response body the status is locked
in, so rate limits that land during layer download arrive as plain
text in the pull stream. Happens on all recent daemon versions -
SUPERVISOR-13FQ (>16k events) and SUPERVISOR-13E0 (>8k events) are
two large examples.
Cases 2 and 3 propagate as plain DockerError, bypass the 429 detection in
install() entirely, never produce a DOCKER_RATELIMIT resolution issue, and
generate large amounts of Sentry noise. Case 1 is detected but routes
every GHCR 429 through Docker-Hub-specific messaging and suggestions.
Changes:
- Add DockerRegistryRateLimitExceeded as the common base class and
GithubContainerRegistryRateLimitExceeded alongside the existing
DockerHubRateLimitExceeded. All extend APITooManyRequests so callers
and retry logic can key off a single type.
- Add GITHUB_RATELIMIT IssueType so GHCR failures don't show the
"log in to Docker Hub" suggestion that DOCKER_RATELIMIT carries.
- PullLogEntry.exception now maps stream errors containing
'toomanyrequests' to DockerRegistryRateLimitExceeded (case 3).
- docker/interface.py:install() routes all three cases through a single
_registry_rate_limit_exception() helper that picks the right issue
type, suggestion and exception subclass based on the image's registry.
- utils/sentry.py filters APITooManyRequests (and anything wrapping it
via __cause__) in capture_exception / async_capture_exception. One
point of policy, every caller benefits.
Callers (supervisor.update(), plugin manager, homeassistant core) are
unchanged - UPDATE_FAILED issues still get created alongside the
registry-specific rate limit issue, giving users the full picture.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Consolidate Sentry noise filtering in one before_send hook
Move the APITooManyRequests filter from capture_exception /
async_capture_exception wrappers into the existing filter_data
before_send hook in supervisor/misc/filter.py, alongside the
AddonConfigurationError filter.
One isinstance tuple check instead of multiple layers, and every path
that reaches Sentry (including logging-integration and excepthook
captures, not just our explicit wrappers) now gets the same treatment.
The filter walks the __cause__ chain so wrapped rate-limit errors
(e.g. DockerHubRateLimitExceeded inside SupervisorUpdateError) still
get filtered. A debug log is emitted on each dropped event for
observability.
Review feedback from mdegat01 on #6732.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Drop GITHUB_RATELIMIT resolution issue
There is no actionable remediation for a GHCR rate limit - logging in
doesn't lift the quota the way it does for Docker Hub, and the cap is
on the authenticated account anyway. A resolution issue that just tells
the user "you were rate limited" adds UI noise without helping them.
Keep the GithubContainerRegistryRateLimitExceeded exception - retry
logic and the Sentry filter still key off it - but don't create a
resolution issue. A log entry from the exception constructor is
sufficient. Docker Hub still gets DOCKER_RATELIMIT + registry-login
suggestion since that is actionable.
Review feedback on #6732.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
7fb621234e |
Add Unix socket support for Core communication with feature flag (#6742)
* Use Unix socket for Supervisor to Core communication Reintroduce Unix socket support for Supervisor-to-Core communication (reverted in #6735) with the addition of a feature flag gate. The feature is now controlled by the `core_unix_socket` feature flag and disabled by default. When enabled and Core version supports it, Supervisor communicates with Core via a Unix socket at /run/os/core.sock instead of TCP. This eliminates the need for access token authentication on the socket path, as Core authenticates the peer by the socket connection itself. Key changes: - Add FeatureFlag.CORE_UNIX_SOCKET to gate the feature - HomeAssistantAPI: transport-aware session/url/websocket management - WSClient: separate connect() (Unix, no auth) and connect_with_auth() (TCP) class methods with proper error handling - APIProxy delegates websocket setup to api.connect_websocket() - Container state tracking for Unix session lifecycle - CI builder mounts /run/supervisor for integration tests Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Sort feature flags alphabetically * Drop per-call max_msg_size from WSClient Hardcode the WebSocket message size cap to 64 MB in WSClient and remove the parameter from WSClient.connect, connect_with_auth, _ws_connect, and HomeAssistantAPI.connect_websocket. This was only ever overridden by APIProxy, so threading it through four layers was unnecessary. max_msg_size is a cap, not a pre-allocation; aiohttp only grows buffers to the size of actual incoming messages. Supervisor's own control channel never approaches 64 MB, so unifying the limit has no runtime cost. Addresses review feedback on #6742. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
ba8c49935b |
Refactor internal addon references to app/apps (#6717)
* Rename addon→app in docstrings and comments Updates all docstrings and inline comments across supervisor/ and tests/ to use the new app/apps terminology. No runtime behaviour is changed by this commit. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Rename addon→app in code (variables, args, class names, functions) Renames all internal Python identifiers from addon/addons to app/apps: - Variable and argument names - Function and method names - Class names (Addon→App, AddonManager→AppManager, DockerAddon→DockerApp, all exception, check, and fixup classes, etc.) - String literals used as Python identifiers (pytest fixtures, parametrize param names, patch.object attribute strings, URL route match_info keys) External API contracts are preserved: JSON keys, error codes, discovery protocol fields, TypedDict/attr.s field names. Import module paths (supervisor/addons/) are also unchanged. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix partial backup/restore API to remap addons key to apps The external API accepts `addons` as the request body key (since ATTR_APPS = "addons"), but do_backup_partial and do_restore_partial now take an `apps` parameter after the rename. The **body expansion in both endpoints would pass `addons=...` causing a TypeError. Remap the key before expansion in both backup_partial and restore_partial: if ATTR_APPS in body: body["apps"] = body.pop(ATTR_APPS) Also adds test_restore_partial_with_addons_key to verify the restore path correctly receives apps= when addons is passed in the request body. This path had no existing test coverage. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix merge error * Adjust AppLoggerAdapter to use app_name Co-authored-by: Stefan Agner <stefan@agner.ch> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Stefan Agner <stefan@agner.ch> |
||
|
|
5c5428fde3 |
Revert "Use Unix socket for Supervisor to Core communication (#6590)" (#6735)
This reverts commit
|
||
|
|
28fa0b35bd |
Use Unix socket for Supervisor to Core communication (#6590)
* Use Unix socket for Supervisor to Core communication Switch internal Supervisor-to-Core HTTP and WebSocket communication from TCP (port 8123) to a Unix domain socket. The existing /run/supervisor directory on the host (already mounted at /run/os inside the Supervisor container) is bind-mounted into the Core container at /run/supervisor. Core receives the socket path via the SUPERVISOR_CORE_API_SOCKET environment variable, creates the socket there, and Supervisor connects to it via aiohttp.UnixConnector at /run/os/core.sock. Since the Unix socket is only reachable by processes on the same host, requests arriving over it are implicitly trusted and authenticated as the existing Supervisor system user. This removes the token round-trip where Supervisor had to obtain and send Bearer tokens on every Core API call. WebSocket connections are likewise authenticated implicitly, skipping the auth_required/auth handshake. Key design decisions: - Version-gated by CORE_UNIX_SOCKET_MIN_VERSION so older Core versions transparently continue using TCP with token auth - LANDINGPAGE is explicitly excluded (not a CalVer version) - Hard-fails with a clear error if the socket file is unexpectedly missing when Unix socket communication is expected - WSClient.connect() for Unix socket (no auth) and WSClient.connect_with_auth() for TCP (token auth) separate the two connection modes cleanly - Token refresh always uses the TCP websession since it is inherently a TCP/Bearer-auth operation - Logs which transport (Unix socket vs TCP) is being used on first request Closes #6626 Related Core PR: home-assistant/core#163907 Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Close WebSocket on handshake failure and validate auth_required Ensure the underlying WebSocket connection is closed before raising when the handshake produces an unexpected message. Also validate that the first TCP message is auth_required before sending credentials. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Fix pylint protected-access warnings in tests Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Check running container env before using Unix socket Split use_unix_socket into two properties to handle the Supervisor upgrade transition where Core is still running with a container started by the old Supervisor (without SUPERVISOR_CORE_API_SOCKET): - supports_unix_socket: version check only, used when creating the Core container to decide whether to set the env var - use_unix_socket: version check + running container env check, used for communication decisions This ensures TCP fallback during the upgrade transition while still hard-failing if the socket is missing after Supervisor configured Core to use it. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Improve Core API communication logging and error handling - Remove transport log from make_request that logged before Core container was attached, causing misleading connection logs - Log "Connected to Core via ..." once on first successful API response in get_api_state, when the transport is actually known - Remove explicit socket existence check from session property, let aiohttp UnixConnector produce natural connection errors during Core startup (same as TCP connection refused) - Add validation in get_core_state matching get_config pattern - Restore make_request docstring Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Guard Core API requests with container running check Add is_running() check to make_request and connect_websocket so no HTTP or WebSocket connection is attempted when the Core container is not running. This avoids misleading connection attempts during Supervisor startup before Core is ready. Also make use_unix_socket raise if container metadata is not available instead of silently falling back to TCP. This is a defensive check since is_running() guards should prevent reaching this state. Add attached property to DockerInterface to expose whether container metadata has been loaded. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Reset Core API connection state on container stop Listen for Core container STOPPED/FAILED events to reset the connection state: clear the _core_connected flag so the transport is logged again on next successful connection, and close any stale Unix socket session. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Only mount /run/supervisor if we use it * Fix pytest errors * Remove redundant is_running check from ingress panel update The is_running() guard in update_hass_panel is now redundant since make_request checks is_running() internally. Also mock is_running in the websession test fixture since tests using it need make_request to proceed past the container running check. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Bind mount /run/supervisor to Supervisor /run/os Home Assistant OS (as well as the Supervised run scripts) bind mount /run/supervisor to /run/os in Supervisor. Since we reuse this location for the communication socket between Supervisor and Core, we need to also bind mount /run/supervisor to Supervisor /run/os in CI. * Wrap WebSocket handshake errors in HomeAssistantAPIError Unexpected exceptions during the WebSocket handshake (KeyError, ValueError, TypeError from malformed messages) are now wrapped in HomeAssistantAPIError inside WSClient.connect/connect_with_auth. This means callers only need to catch HomeAssistantAPIError. Remove the now-unnecessary except (RuntimeError, ValueError, TypeError) from proxy _websocket_client and add a proper error message to the APIError per review feedback. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Narrow WebSocket handshake exception handling Replace broad `except Exception` with specific exception types that can actually occur during the WebSocket handshake: KeyError (missing dict keys), ValueError (bad JSON), TypeError (non-text WS message), aiohttp.ClientError (connection errors), and TimeoutError. This avoids silently wrapping programming errors into HomeAssistantAPIError. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Remove unused create_mountpoint from MountBindOptions The field was added but never used. The /run/supervisor host path is guaranteed to exist since HAOS creates it for the Supervisor container mount, so auto-creating the mountpoint is unnecessary. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Clear stale access token before raising on final retry Move token clear before the attempt check in connect_websocket so the stale token is always discarded, even when raising on the final attempt. Without this, the next call would reuse the cached bad token via _ensure_access_token's fast path, wasting a round-trip. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Add tests for Unix socket communication and Core API Add tests for the new Unix socket communication path and improve existing test coverage: - Version-based supports_unix_socket and env-based use_unix_socket - api_url/ws_url transport selection - Connection lifecycle: connected log after restart, ignoring unrelated container events - get_api_state/check_api_state parameterized across versions, responses, and error cases - make_request is_running guard and TCP flow with real token fetch - connect_websocket for both Unix and TCP (with token verification) - WSClient.connect/connect_with_auth handshake success, errors, cleanup on failure, and close with pending futures Consolidate existing tests into parameterized form and drop synthetic tests that covered very little. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
a30f2509a3 |
Enable IPv6 on Supervisor network by default for all installations (#6720)
* Improve Docker network test coverage and infrastructure Add test cases for enable_ipv6=None (no user setting) to test_network_recreation, verifying existing behavior where None leaves the network unchanged. Use pytest.param with descriptive IDs for better test readability. Add create_network_mock side_effect to the docker fixture so network creation returns realistic metadata built from the provided params. Remove redundant manual create mock setups from individual tests. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Enable IPv6 on Supervisor network by default for all installations Previously, IPv6 was only enabled by default for new installations (when enable_ipv6 config was None). Existing installations with IPv4-only networks were left unchanged unless the user explicitly set enable_ipv6 to true. Now, when no explicit IPv6 setting exists, the network is migrated to dual-stack on next boot. The same safety checks apply: migration is blocked if user containers are running and requires a reboot. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
a4a17a70a5 |
Add specific error message for registry authentication failures (#6678)
* Add specific error message for registry authentication failures
When a Docker image pull fails with 401 Unauthorized and registry
credentials are configured, raise DockerRegistryAuthError instead of
a generic DockerError. This surfaces a clear message to the user
("Docker registry authentication failed for <registry>. Check your
registry credentials") instead of "An unknown error occurred with
addon <name>".
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Add tests for registry authentication error handling
Test that a 401 during image pull raises DockerRegistryAuthError when
credentials are configured, and falls back to generic DockerError
when no credentials are present.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Add tests for addon install/update/rebuild auth failure handling
Test that DockerRegistryAuthError propagates correctly through
addon install, update, and rebuild paths without being wrapped
in a generic AddonUnknownError.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
1fd78dfc4e |
Fix Docker Hub registry auth for containerd image store (#6677)
aiodocker derives ServerAddress for X-Registry-Auth by doing
image.partition("/"). For Docker Hub images like
"homeassistant/amd64-supervisor", this extracts "homeassistant"
(the namespace) instead of "docker.io" (the registry).
With the classic graphdriver image store, ServerAddress was never
checked and credentials were sent regardless. With the containerd
image store (default since Docker v29 / HAOS 15), the resolver
compares ServerAddress against the actual registry host and silently
drops credentials on mismatch, falling back to anonymous access.
Fix by prefixing Docker Hub images with "docker.io/" when registry
credentials are configured, so aiodocker sets ServerAddress correctly.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
0ef71d1dd1 |
Drop unsupported architectures and machines, create issue for affected apps (#6607)
* Drop unsupported architectures and machines from Supervisor Since #5620 Supervisor no longer updates the version information on unsupported architectures and machines. This means users can no longer update to newer version of Supervisor since that PR got released. Furthermore since #6347 we also no longer build for these architectures. With this, any code related to these architectures becomes dead code and should be removed. This commit removes all refrences to the deprecated architectures and machines from Supervisor. This affects the following architectures: - armhf - armv7 - i386 And the following machines: - odroid-xu - qemuarm - qemux86 - raspberrypi - raspberrypi2 - raspberrypi3 - raspberrypi4 - tinker * Create issue if an app using a deprecated architecture is installed This adds a check to the resolution system to detect if an app is installed that uses a deprecated architecture. If so, it will show a warning to the user and recommend them to uninstall the app. * Formally deprecate machine add-on configs as well Not only deprecate add-on configs for unsupported architectures, but also for unsupported machines. * For installed add-ons architecture must always exist Fail hard in case of missing architecture, as this is a required field for installed add-ons. This will prevent the Supervisor from running with an unsupported configuration and causing further issues down the line. |
||
|
|
4a1c816b92 | Finish dockerpy to aiodocker migration (#6578) | ||
|
|
0cd668ec77 |
Fix typeguard errors by explicitly converting IP addresses to strings (#6531)
* Fix environment variable type errors by converting IP addresses to strings Environment variables must be strings, but IPv4Address and IPv4Network objects were being passed directly to container environment dictionaries, causing typeguard validation errors. Changes: - Convert IPv4Address objects to strings in homeassistant.py for SUPERVISOR and HASSIO environment variables - Convert IPv4Network object to string in observer.py for NETWORK_MASK environment variable - Update tests to expect string values instead of IP objects in environment dictionaries - Remove unused ip_network import from test_observer.py Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com> * Use explicit string conversion for extra_hosts IP addresses Use the !s format specifier in the f-string to explicitly convert IPv4Address objects to strings when building the ExtraHosts list. While f-strings implicitly convert objects to strings, using !s makes the conversion explicit and consistent with the environment variable fixes in the previous commit. Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com> |
||
|
|
d1a576e711 |
Fix Docker Hub manifest fetching by using correct registry API endpoint (#6525)
The manifest fetcher was using docker.io as the registry API endpoint, but Docker Hub's actual registry API is at registry-1.docker.io. When trying to access https://docker.io/v2/..., requests were being redirected to https://www.docker.com/ (the marketing site), which returned HTML instead of JSON, causing manifest fetching to fail. This matches exactly what Docker itself does internally - see daemon/pkg/registry/config.go:49 where Docker hardcodes DefaultRegistryHost = "registry-1.docker.io" for registry operations. Changes: - Add DOCKER_HUB_API constant for the actual API endpoint - Add _get_api_endpoint() helper to translate docker.io to registry-1.docker.io for HTTP API calls - Update _get_auth_token() and _fetch_manifest() to use the API endpoint - Keep docker.io as the registry identifier for naming and credentials - Add tests to verify the API endpoint translation Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com> |
||
|
|
a122b5f1e9 |
Migrate info, events and container logs to aiodocker (#6514)
* Migrate info and events to aiodocker * Migrate container logs to aiodocker * Fix dns plugin loop test * Fix mocking for docker info * Fixes from feedback * Harden monitor error handling * Deleted failing tests because they were not useful |
||
|
|
6957341c3e |
Refactor Docker pull progress with registry manifest fetcher (#6379)
* Use count-based progress for Docker image pulls Refactor Docker image pull progress to use a simpler count-based approach where each layer contributes equally (100% / total_layers) regardless of size. This replaces the previous size-weighted calculation that was susceptible to progress regression. The core issue was that Docker rate-limits concurrent downloads (~3 at a time) and reports layer sizes only when downloading starts. With size- weighted progress, large layers appearing late would cause progress to drop dramatically (e.g., 59% -> 29%) as the total size increased. The new approach: - Each layer contributes equally to overall progress - Per-layer progress: 70% download weight, 30% extraction weight - Progress only starts after first "Downloading" event (when layer count is known) - Always caps at 99% - job completion handles final 100% This simplifies the code by moving progress tracking to a dedicated module (pull_progress.py) and removing complex size-based scaling logic that tried to account for unknown layer sizes. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Exclude already-existing layers from pull progress calculation Layers that already exist locally should not count towards download progress since there's nothing to download for them. Only layers that need pulling are included in the progress calculation. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Add registry manifest fetcher for size-based pull progress Fetch image manifests directly from container registries before pulling to get accurate layer sizes upfront. This enables size-weighted progress tracking where each layer contributes proportionally to its byte size, rather than equal weight per layer. Key changes: - Add RegistryManifestFetcher that handles auth discovery via WWW-Authenticate headers, token fetching with optional credentials, and multi-arch manifest list resolution - Update ImagePullProgress to accept manifest layer sizes via set_manifest() and calculate size-weighted progress - Fall back to count-based progress when manifest fetch fails - Pre-populate layer sizes from manifest when creating layer trackers The manifest fetcher supports ghcr.io, Docker Hub, and private registries by using credentials from Docker config when available. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Clamp progress to 100 to prevent floating point precision issues Floating point arithmetic in weighted progress calculations can produce values slightly above 100 (e.g., 100.00000000000001). This causes validation errors when the progress value is checked. Add min(100, ...) clamping to both size-weighted and count-based progress calculations to ensure the result never exceeds 100. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Use sys_websession for manifest fetcher instead of creating new session Reuse the existing CoreSys websession for registry manifest requests instead of creating a new aiohttp session. This improves performance and follows the established pattern used throughout the codebase. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Make platform parameter required and warn on missing platform - Make platform a required parameter in get_manifest() and _fetch_manifest() since it's always provided by the calling code - Return None and log warning when requested platform is not found in multi-arch manifest list, instead of falling back to first manifest which could be the wrong architecture 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Log manifest fetch failures at warning level Users will notice degraded progress tracking when manifest fetch fails, so log at warning level to help diagnose issues. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Add pylint disable comments for protected access in manifest tests 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Separate download_current and total_size updates in pull progress Update download_current and total_size independently in the DOWNLOADING handler. This ensures download_current is updated even when total is not yet available. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Reject invalid platform format in manifest selection --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
a5c3781f9d | Migrate network interactions to aiodocker (#6505) | ||
|
|
2a4890e2b0 |
Bump aiodocker from 0.24.0 to 0.25.0 (#6448)
* Bump aiodocker from 0.24.0 to 0.25.0 Bumps [aiodocker](https://github.com/aio-libs/aiodocker) from 0.24.0 to 0.25.0. - [Release notes](https://github.com/aio-libs/aiodocker/releases) - [Changelog](https://github.com/aio-libs/aiodocker/blob/main/CHANGES.rst) - [Commits](https://github.com/aio-libs/aiodocker/compare/v0.24.0...v0.25.0) --- updated-dependencies: - dependency-name: aiodocker dependency-version: 0.25.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> * Update to new timeout configuration * Fix pytest failure --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Mike Degatano <michael.degatano@gmail.com> Co-authored-by: Stefan Agner <stefan@agner.ch> |
||
|
|
de02bc991a |
fix: pull missing images before running (#6500)
* fix: pull missing images before running * add tests for auto-pull behavior |
||
|
|
909a2dda2f |
Migrate (almost) all docker container interactions to aiodocker (#6489)
* Migrate all docker container interactions to aiodocker
* Remove containers_legacy since its no longer used
* Add back remove color logic
* Revert accidental invert of conditional in setup_network
* Fix typos found by copilot
* Apply suggestions from code review
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Revert "Apply suggestions from code review"
This reverts commit
|
||
|
|
753021d4d5 | Fix 'DockerMount is not JSON serializable' in DockerAPI.run_command (#6477) | ||
|
|
d23bc291d5 |
Migrate create container to aiodocker (#6415)
* Migrate create container to aiodocker * Fix extra hosts transformation * Env not Environment * Fix tests * Fixes from feedback --------- Co-authored-by: Jan Čermák <sairon@users.noreply.github.com> |
||
|
|
cdef1831ba |
Add option to Core settings to enable duplicated logs (#6400)
Introduce new option `duplicate_log_file` to HA Core configuration that will set an environment variable `HA_DUPLICATE_LOG_FILE=1` for the Core container if enabled. This will serve as a flag for Core to enable the legacy log file, along the standard logging which is handled by Systemd Journal. |
||
|
|
382f0e8aef |
Disable timeout for Docker image pull operations (#6391)
* Disable timeout for Docker image pull operations The aiodocker migration introduced a regression where image pulls could timeout during slow downloads. The session-level timeout (900s total) was being applied to pull operations, but docker-py explicitly sets timeout=None for pulls, allowing them to run indefinitely. When aiodocker receives timeout=None, it converts it to ClientTimeout(total=None), which aiohttp treats as "no timeout" (returns TimerNoop instead of enforcing a timeout). This fixes TimeoutError exceptions that could occur during installation on systems with slow network connections or when pulling large images. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Fix pytests --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
d220fa801f |
Await aiodocker import_image coroutine (#6378)
The aiodocker images.import_image() method returns a coroutine that needs to be awaited, but the code was iterating over it directly, causing "TypeError: 'coroutine' object is not iterable". Fixes SUPERVISOR-13D9 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
50d31202ae |
Use Docker's official registry domain detection logic (#6360)
* Use Docker's official registry domain detection logic Replace the custom IMAGE_WITH_HOST regex with a proper implementation based on Docker's reference parser (vendor/github.com/distribution/ reference/normalize.go). Changes: - Change DOCKER_HUB from "hub.docker.com" to "docker.io" (official default) - Add DOCKER_HUB_LEGACY for backward compatibility with "hub.docker.com" - Add IMAGE_DOMAIN_REGEX and get_domain() function that properly detects: - localhost (with optional port) - Domains with "." (e.g., ghcr.io, 127.0.0.1) - Domains with ":" port (e.g., myregistry:5000) - IPv6 addresses (e.g., [::1]:5000) - Update credential handling to support both docker.io and hub.docker.com - Add comprehensive tests for domain detection 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Refactor Docker domain detection to utils module Move get_domain function to supervisor/docker/utils.py and rename it to get_domain_from_image for consistency with get_registry_for_image. Use named group in the regex for better readability. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Rename domain to registry for consistency Use consistent "registry" terminology throughout the codebase: - Rename get_domain_from_image to get_registry_from_image - Rename IMAGE_DOMAIN_REGEX to IMAGE_REGISTRY_REGEX - Update named group from "domain" to "registry" - Update all related comments and variable names 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
6302c7d394 |
Fix progress when using containerd snapshotter (#6357)
* Fix progress when using containerd snapshotter * Add test for tiny image download under containerd-snapshotter * Fix API tests after progress allocation change * Fix test for auth changes * Apply suggestions from code review Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> --------- Co-authored-by: Stefan Agner <stefan@agner.ch> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> |
||
|
|
8a251e0324 |
Pass registry credentials to add-on build for private base images (#6356)
* Pass registry credentials to add-on build for private base images When building add-ons that use a base image from a private registry, the build would fail because credentials configured via the Supervisor API were not passed to the Docker-in-Docker build container. This fix: - Adds get_docker_config_json() to generate a Docker config.json with registry credentials for the base image - Creates a temporary config file and mounts it into the build container at /root/.docker/config.json so BuildKit can authenticate when pulling the base image - Cleans up the temporary file after build completes Fixes #6354 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Fix pylint errors * Apply suggestions from code review Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Refactor registry credential extraction into shared helper Extract duplicate logic for determining which registry matches an image into a shared `get_registry_for_image()` method in `DockerConfig`. This method is now used by both `DockerInterface._get_credentials()` and `AddonBuild.get_docker_config_json()`. Move `DOCKER_HUB` and `IMAGE_WITH_HOST` constants to `docker/const.py` to avoid circular imports between manager.py and interface.py. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Apply suggestions from code review Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * ruff format * Document raises --------- Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Mike Degatano <michael.degatano@gmail.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> |
||
|
|
ae7700f52c |
Fix private registry authentication for aiodocker image pulls (#6355)
* Fix private registry authentication for aiodocker image pulls After PR #6252 migrated image pulling from dockerpy to aiodocker, private registry authentication stopped working. The old _docker_login() method stored credentials in ~/.docker/config.json via dockerpy, but aiodocker doesn't read that file - it requires credentials passed explicitly via the auth parameter. Changes: - Remove unused _docker_login() method (dockerpy login was ineffective) - Pass credentials directly to pull_image() via new auth parameter - Add auth parameter to DockerAPI.pull_image() method - Add unit tests for Docker Hub and custom registry authentication Fixes #6345 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Ignore protected access in test * Fix plug-in pull test * Fix HA core tests --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
63a3dff118 |
Handle pull events with complete progress details only (#6320)
* Handle pull events with complete progress details only Under certain circumstances, Docker seems to send pull events with incomplete progress details (i.e., missing 'current' or 'total' fields). In practise, we've observed an empty dictionary for progress details as well as missing 'total' field (while 'current' was present). All events were using Docker 28.3.3 using the old, default Docker graph backend. * Fix docstring/comment |
||
|
|
30cc172199 |
Migrate images from dockerpy to aiodocker (#6252)
* Migrate images from dockerpy to aiodocker * Add missing coverage and fix bug in repair * Bind libraries to different files and refactor images.pull * Use the same socket again Try using the same socket again. * Fix pytest --------- Co-authored-by: Stefan Agner <stefan@agner.ch> |