mirror of
https://github.com/home-assistant/supervisor.git
synced 2026-10-08 07:14:31 +01:00
ab96eb313cffa8185caed410e2ee4d50dfffa63f
5933
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ab96eb313c | Pair astroid minor updates with pylint instead of disabling them | ||
|
|
050c129d6e | Keep offering astroid patch releases in Renovate | ||
|
|
de83dcabd1 | Move dependency updates from Dependabot to Renovate | ||
|
|
d0d259cf73 |
Bump debugpy from 1.8.21 to 1.8.22 (#7244)
Bumps [debugpy](https://github.com/microsoft/debugpy) from 1.8.21 to 1.8.22. - [Release notes](https://github.com/microsoft/debugpy/releases) - [Commits](https://github.com/microsoft/debugpy/compare/v1.8.21...v1.8.22) --- updated-dependencies: - dependency-name: debugpy dependency-version: 1.8.22 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
f484f5bb39 |
Bump urllib3 from 2.7.0 to 2.8.0 (#7243)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.7.0 to 2.8.0. - [Release notes](https://github.com/urllib3/urllib3/releases) - [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst) - [Commits](https://github.com/urllib3/urllib3/compare/2.7.0...2.8.0) --- updated-dependencies: - dependency-name: urllib3 dependency-version: 2.8.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
ec4b8c216d |
Bump codecov/codecov-action from 7.1.0 to 7.1.1 (#7242)
Bumps [codecov/codecov-action](https://github.com/codecov/codecov-action) from 7.1.0 to 7.1.1. - [Release notes](https://github.com/codecov/codecov-action/releases) - [Changelog](https://github.com/codecov/codecov-action/blob/main/CHANGELOG.md) - [Commits](https://github.com/codecov/codecov-action/compare/0b35c9ecc4f0529d0eb674914510c22f85b196b4...303a32d7a59b442fa8d48b6a1cc6825c09c847a5) --- updated-dependencies: - dependency-name: codecov/codecov-action dependency-version: 7.1.1 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
fa2a8f315b |
Bump sentry-sdk from 2.69.1 to 2.69.2 (#7241)
Bumps [sentry-sdk](https://github.com/getsentry/sentry-python) from 2.69.1 to 2.69.2. - [Release notes](https://github.com/getsentry/sentry-python/releases) - [Changelog](https://github.com/getsentry/sentry-python/blob/master/CHANGELOG.md) - [Commits](https://github.com/getsentry/sentry-python/compare/2.69.1...2.69.2) --- updated-dependencies: - dependency-name: sentry-sdk dependency-version: 2.69.2 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
89b7a1eb76 |
Bump ruff from 0.16.7 to 0.16.8 (#7240)
Bumps [ruff](https://github.com/astral-sh/ruff) from 0.16.7 to 0.16.8. - [Release notes](https://github.com/astral-sh/ruff/releases) - [Changelog](https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md) - [Commits](https://github.com/astral-sh/ruff/compare/0.16.7...0.16.8) --- updated-dependencies: - dependency-name: ruff dependency-version: 0.16.8 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
64ea3be432 |
Keep "!secret" references when persisting app options (#7237)
Since #6203 the set-options endpoint stores the value returned by the options validator. The validator replaces "!secret x" with the resolved secret before type-checking, so the plaintext secret ended up in the persisted app options instead of the reference (#7152). Make the validator non-destructive on request instead of restoring the references afterwards: AppOptions gains a resolve_secrets flag. When it is False, a "!secret x" value is still validated against the resolved secret (including the pwned hash and type coercion) but the reference is returned. The set-options endpoint validates with resolve_secrets=False, while write_options and the self options config endpoint keep resolving secrets for the container. Fixes #7152 Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>2026.09.3 |
||
|
|
a5d5c79a57 |
Add feature flag to drop MKNOD, AUDIT_WRITE and SETFCAP from apps (#7233)
containerd is introducing a "reduced" capability profile that removes NET_RAW, MKNOD, AUDIT_WRITE and SETFCAP from the default set, matching what Kubernetes' restricted profile has dropped for years. NET_RAW is handled by the app_drop_net_raw flag since it has a security implication of its own and known app fallout. This adds the remaining three under a separate flag. Add AUDIT_WRITE, MKNOD and SETFCAP to the Capabilities enum so apps can request them under privileged, and the app_reduced_capabilities development feature flag. When enabled, app containers are created with these capabilities in CapDrop unless the app explicitly declares them. The security rating is unchanged for them. None of the three has a consumer among the current core and community apps. Apps receive the host /dev bind-mounted, so MKNOD is not needed to see devices. The SSH apps and Debian's openssh-server are built without libaudit and PAM, so AUDIT_WRITE is unused. SETFCAP only matters for packages installed at runtime: Debian postinst scripts fall back to setuid with a warning and apk logs a failed xattr write, both without failing the install. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
b44d6aca86 |
Don't leave rejected AppArmor profile files in the store directory (#7229)
* Remove AppArmor profile file even when unloading fails OS Agent 1.14.0 (shipped with HAOS 18.3) validates AppArmor profiles with the actual apparmor_parser on both load and unload, and refuses to unload a profile file that defines unexpected or renamed profiles. remove_profile propagated that failure before unlinking the stored profile, so a file the agent rejects stayed in the AppArmor store directory, was retried on every startup, and could never be cleaned up through an uninstall. Log the unload failure and remove the stored profile file regardless. The profile may remain loaded in the kernel until the next reboot, but the file no longer lingers. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Remove copied AppArmor profile when the OS Agent rejects the load load_profile copies the profile into the AppArmor store directory before asking the OS Agent to load it, so a profile rejected by the agent's parser validation (added in OS Agent 1.14.0) was left behind on disk. On the next startup load() re-reads the store directory and retries the rejected file, and it could never be cleaned up unless the app was uninstalled. Drop the in-memory profile entry and remove the copied file again when the agent rejects the load, keeping the entry if the file cannot be removed so a later remove_profile can still reach it. Transient load failures are left untouched so a good profile is not discarded. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
92a8d1044a |
Add feature flag to drop NET_RAW from app containers (#7230)
Docker's default capability set includes CAP_NET_RAW, and Supervisor never set CapDrop, so every app container could open raw and packet sockets regardless of its manifest. On the shared internal network that allows ARP and MAC spoofing between apps and against Supervisor and Core traffic. Add the app_drop_net_raw development feature flag. When enabled, app containers are created with NET_RAW in CapDrop unless the app explicitly declares it under privileged. NET_RAW has been an accepted privileged entry since 2023, so apps that need raw sockets can opt in on any supported Supervisor. No capability is implied from another one: apps that rely on the Docker default, including ones declaring NET_ADMIN, must declare NET_RAW themselves. Docker's ping_group_range default lets ping work through ICMP datagram sockets without NET_RAW, so plain connectivity checks are unaffected. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
be3b5ea21e |
Bump codecov/codecov-action from 7.0.0 to 7.1.0 (#7235)
Bumps [codecov/codecov-action](https://github.com/codecov/codecov-action) from 7.0.0 to 7.1.0. - [Release notes](https://github.com/codecov/codecov-action/releases) - [Changelog](https://github.com/codecov/codecov-action/blob/main/CHANGELOG.md) - [Commits](https://github.com/codecov/codecov-action/compare/fb8b3582c8e4def4969c97caa2f19720cb33a72f...0b35c9ecc4f0529d0eb674914510c22f85b196b4) --- updated-dependencies: - dependency-name: codecov/codecov-action dependency-version: 7.1.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
11eb58e9ed |
Bump coverage from 7.16.0 to 7.16.1 (#7236)
Bumps [coverage](https://github.com/coveragepy/coveragepy) from 7.16.0 to 7.16.1. - [Release notes](https://github.com/coveragepy/coveragepy/releases) - [Changelog](https://github.com/coveragepy/coveragepy/blob/main/CHANGES.rst) - [Commits](https://github.com/coveragepy/coveragepy/compare/7.16.0...7.16.1) --- updated-dependencies: - dependency-name: coverage dependency-version: 7.16.1 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
c9543e8649 |
Slim down the agent instructions to essential guidance (#7228)
The instructions file carried product documentation on add-ons, the update system and backups, general Python advice, a directory tree and code samples. Every agent session loaded all of it. Current guidance for agent instruction files is a short list of commands and project specific conventions the agent cannot infer from the code. Restructure the file into the pull request template rule, development commands, Python 3.14 syntax notes, testing rules, good practices and the AI policy. Drop the coverage threshold, which nothing in CI enforces. The file shrinks from about 2,700 to 1,100 tokens. Keep the Supervisor specific conventions: relative imports, per-module const.py, CoreSysAttributes with the sys_* properties, the executor for blocking calls, api_process with api_validate, exceptions from exceptions.py and plain test_ functions. Replace add-on with app throughout and state that add-on is the legacy term. The code has moved to app naming, so agents must use it in new code and keep addon only where an existing API field, config key or class name still carries it. Adopt the section layout and several rules from the AGENTS.md in Home Assistant Core that were not in the old file: the pointer to the VS Code tasks, the note on lazily evaluated annotations, the test rules on type annotations, usefixtures, branching and parametrization, and the rules on small try clauses and comments. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1d184b2758 |
Model OS Agent AppArmor profile rejections as API errors (#7226)
* Model OS Agent AppArmor profile rejections as API errors When the OS Agent's AppArmor parser rejects a profile (for example an app profile that defines an unexpected profile name), the resulting HostAppArmorError surfaced through api_process as an "Unexpected error" with a full traceback and a Sentry event. The profile content is user-supplied through the app repository, so a parser rejection is a client error rather than a Supervisor fault. Add HostAppArmorLoadProfileError, a HostAppArmorError that also inherits APIError, and raise it from the D-Bus load path. It carries the profile name and the OS Agent's error message as extra fields so the reason is still relayed to the API caller, while the error is now answered with a plain 400 and no Sentry capture. Existing catch sites keep working since the class remains a HostAppArmorError. Extend the AppArmor D-Bus mock to inject load failures and add a test covering the relayed message and error metadata. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Limit the AppArmor API error to OS Agent parser rejections Only DBusFatalError, which from_dbus_error uses for service-specific failures, indicates the OS Agent rejected the profile. Transport failures such as timeouts or a missing reply keep raising the plain HostAppArmorError so they still surface as unexpected errors. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
0e7c4175c5 |
Block start/restart/rebuild/update APIs while Supervisor is not fully running (#7194)
* Block start/restart/rebuild/update APIs while Supervisor is not fully running Supervisor boots apps and Home Assistant Core in a specific order, and replays a similarly ordered sequence while restoring a backup (during which it is in the freeze state). Calling start/restart/rebuild/update via the API while one of those sequences is still in progress risks running it out of order, or blocking a scheduled start from that sequence outright via job concurrency. Add a require_running_system decorator that rejects these API calls unless Supervisor is fully started (CoreState.RUNNING), and apply it to: - App start/restart/rebuild (fixes #7189) and update - Home Assistant Core start/restart/rebuild/update - Plugin (audio/dns/multicast) restart/update, and cli/observer update Also add a translatable APISystemNotReadyError so the API returns a clean, user-facing message and error_key instead of leaking the technical JobConditionException message. * Address review: neutral error wording + require_running_system test coverage - Reword APISystemNotReadyError so it no longer implies the block is only due to startup, since the same guard also applies while frozen for a backup restore. - Add tests/api/test_require_running_system.py covering the rejection path (400 + system_not_ready_error, business logic not invoked) for app start/restart/rebuild, app update via the store API, Core start/restart/rebuild/update, and all plugin update/restart routes, parametrized across SETUP/STARTUP/FREEZE. * Simplify APISystemNotReadyError message wording * Switch status code to 503 |
||
|
|
cbeb340bde |
Evaluate synchronous job conditions before awaiting ones (#7210)
* Evaluate synchronous job conditions before awaiting ones Job.check_conditions grew by accretion: each new condition was inserted next to a related one, and the order was never a design decision. When the checks were introduced in #2250 every one of them was a cheap synchronous read, so the order did not matter. Since then FREE_SPACE became an executor call and the INTERNET_SYSTEM/INTERNET_HOST checks started awaiting a connectivity probe (#6765) without being moved. As a result a job that is going to be refused for an unsupported OS or architecture, disabled auto update, or a missing host feature first pays for a disk stat and a network probe. Move the three awaiting checks after all synchronous state reads, and move MOUNT_AVAILABLE out from behind PLUGINS_UPDATED into the synchronous section. PLUGINS_UPDATED stays last because it performs plugin updates rather than only checking. Document the rule so new conditions land in the right section. Besides avoiding needless I/O, this makes a refusal visible before the first suspension. Callers that eagerly start a job as a task can then rely on the task being done immediately when the job was refused, which the Supervisor auto update in #7201 depends on. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Add regression test for job condition evaluation order Refuse a job through a synchronous condition while it also carries FREE_SPACE and both INTERNET conditions, start it eagerly, and assert the task completes at once without the disk or connectivity probes being called. The test fails against the previous ordering because FREE_SPACE ran its probe before the synchronous refusal. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
94b4a0b77c |
Forward Core API proxy paths byte-for-byte and match deny rules on decoded path (#7225)
The Core API proxy compared its deny pattern against the once-decoded route capture and then formatted that string into the upstream URL, where yarl decoded percent-encoded unreserved characters a second time. A path such as hassio%255Fauth/password_reset passed both the middleware blacklist and the proxy's deny check but reached Core as /api/hassio_auth/password_reset, which the proxy executes as the Supervisor user. Forward the raw path exactly as received instead of the decoded capture, and build the upstream URL with encoded=True so the bytes checked are the bytes sent. This also stops lossy re-encoding of legitimate paths, e.g. an encoded slash no longer turns into a path separator. Match both the middleware blacklist and the proxy deny pattern against the recursively unquoted path so any encoding depth resolves to the same decision. The recursive unquote helper moves to module level so both call sites share it. Reported in GHSA-m2gm-724m-7rf9. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>2026.09.2 |
||
|
|
7f30791f5d |
Bump home-assistant/builder/actions/prepare-multi-arch-matrix from 2026.06.0 to 2026.09.0 (#7224)
Bumps [home-assistant/builder/actions/prepare-multi-arch-matrix](https://github.com/home-assistant/builder) from 2026.06.0 to 2026.09.0. - [Release notes](https://github.com/home-assistant/builder/releases) - [Commits](https://github.com/home-assistant/builder/compare/4de35182ce1e329181bffcbcc84d33db5e2c7e10...7412f0023ea9b6e58e8bb5059f1660f51376f49a) --- updated-dependencies: - dependency-name: home-assistant/builder/actions/prepare-multi-arch-matrix dependency-version: 2026.09.0 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
0a05c0f108 |
Bump home-assistant/builder/actions/publish-multi-arch-manifest from 2026.06.0 to 2026.09.0 (#7223)
Bumps [home-assistant/builder/actions/publish-multi-arch-manifest](https://github.com/home-assistant/builder) from 2026.06.0 to 2026.09.0. - [Release notes](https://github.com/home-assistant/builder/releases) - [Commits](https://github.com/home-assistant/builder/compare/4de35182ce1e329181bffcbcc84d33db5e2c7e10...7412f0023ea9b6e58e8bb5059f1660f51376f49a) --- updated-dependencies: - dependency-name: home-assistant/builder/actions/publish-multi-arch-manifest dependency-version: 2026.09.0 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
5ec2948ce2 |
Bump home-assistant/builder/actions/build-image from 2026.06.0 to 2026.09.0 (#7222)
Bumps [home-assistant/builder/actions/build-image](https://github.com/home-assistant/builder) from 2026.06.0 to 2026.09.0. - [Release notes](https://github.com/home-assistant/builder/releases) - [Commits](https://github.com/home-assistant/builder/compare/4de35182ce1e329181bffcbcc84d33db5e2c7e10...7412f0023ea9b6e58e8bb5059f1660f51376f49a) --- updated-dependencies: - dependency-name: home-assistant/builder/actions/build-image dependency-version: 2026.09.0 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
40e3ee7640 |
Fix supervisor config restore from encrypted backups (#7218)
* Fix supervisor config restore from encrypted backups The mount and Docker registry configuration stored in supervisor.tar was never restored from encrypted backups. securetar opens encrypted inner tars in streaming mode since it cannot seek in the ciphertext. Reading a member via getmember() followed by extractfile() requires seeking backwards: to find the member, tarfile reads all headers and thereby advances past the member data. This raises tarfile.StreamError, which was caught by the broad TarError handler and only logged as a warning, so the restore reported success while the mounts were silently missing. This affects every encrypted backup since the supervisor tar was introduced in 2026.03.3, and Core-created backups are encrypted by default. The data in those backups is intact, restoring them again with this fix recovers the configuration. Read the inner tar sequentially and pick up mounts.json and docker.json as they stream past. Add a test which stores and restores mounts and registries with a backup password set, as the existing tests only covered plain tars. Also document the archive layout in the Backup class docstring, noting that Core only encrypts/decrypts inner tars from a fixed list when rewriting a backup, so any new inner tar must be added there as well (see home-assistant/core#182137). Fixes #7213 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Add end-to-end test with encrypted supervisor config fixture Add backup_example_enc_supervisor.tar, a partial backup created by Supervisor with password "test123" that only contains supervisor.tar.gz with a CIFS mount and a registry. Restore it through the backup manager to cover the full encrypted restore path. Also cover supervisor tars with only one of the JSON files, as written by Supervisor 2026.03.x before docker.json was added, including unknown and non-file members which are skipped. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>2026.09.1 |
||
|
|
1a7853de9a |
Start the Supervisor auto update on every version reload (#7201)
* Start the Supervisor auto update on every version reload A pending Supervisor update blocks Core, OS and app updates through the SUPERVISOR_UPDATED job condition. Only the startup and the daily scheduled reload started the auto update. Every other reload, such as the one Core requests when the user opens Settings, made the new version known without installing it. The user then saw "supervisor needs to be updated first" until the daily task ran. The updater now starts the auto update task after every reload while the system is running. The scheduled task calls the updater reload directly. The Supervisor update job rejects a concurrent update request, so a user requested update during the auto update fails cleanly instead of creating an update failed issue. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC * Report success for an update request during a running Supervisor update The auto update can already run when Core or a user requests the update. The requested outcome is underway, so the API reports success and the caller sees the result through the Supervisor restart. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC * Trim comments around the Supervisor auto update Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC * Drop comment on the updater auto update trigger Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC * Raise specific errors when restoring a backup requires a Supervisor update Move the supervisor version check in do_restore_full/do_restore_partial into a shared _check_supervisor_version helper. If the backup requires a newer Supervisor version: - Raise the new translatable BackupSupervisorVersionError (400) with the existing message when auto update is disabled. - Otherwise kick off the Supervisor auto update in the background via sys_create_task(auto_update_supervisor()) and raise the new translatable BackupSupervisorUpdateInProgressError (503), telling the caller to retry once the update completes. * Recheck for Supervisor update before failing a version-mismatch restore When a restore requires a newer Supervisor version and auto update is enabled, refresh version info first in case a newer Supervisor was just released. If an update is already known to be needed, make sure it gets kicked off directly instead of waiting on a reload. Only fall back to the translatable version-mismatch error if no update is available after the recheck. Also move auto_update_supervisor from Tasks to the Supervisor class so it can be reused outside of the task scheduler. * Track auto update and version fetch tasks to close restore-check races Supervisor.auto_update_supervisor() now stores and returns its update task (created with eager_start so a synchronous failure is reflected immediately) instead of firing-and-forgetting it. This lets callers check the task's actual done state rather than guessing based on need_update, which could be wrong if the update job rejected the call for an unrelated reason. HomeAssistantCore's install retry loop now also catches SupervisorJobError when triggering the Supervisor update it depends on, so it keeps retrying while an update is already in progress instead of giving up. Add Updater.start_fetch_data(), which starts (or reuses) a version fetch and returns its task. fetch_data() sets its throttle timestamp before running, even on failure, so a caller relying on fetch_data() directly can't tell "someone else just refreshed" from "someone else's fetch just failed" - it would silently skip either way. Callers that need to know the true outcome should await the shared task instead. reload() and BackupManager._check_supervisor_version() are updated to use it: the latter awaits the fetch task directly (skipping it entirely if there's no connectivity to fetch with) before checking auto_update_supervisor()'s task, closing a race where a stale/failed throttle window could cause a 400 (version mismatch) response instead of a 503 (update in progress). * Use Job decorator's detach option for update and version fetch tasks #7211 added a `detach` option to the Job decorator, making the manual task-tracking previously added here redundant. Replace it with the decorator-native option: - `Supervisor.update()` is now `detach=True` with `concurrency=JobConcurrency.REJECT`. Calling it returns the task performing the update - which may already be in progress from any previous caller, manual or automatic - instead of raising `SupervisorJobError` or blocking the caller until it (and the Supervisor restart it triggers) completes. Errors are raised on the returned task, not to the immediate caller. - `Supervisor.auto_update_supervisor()` is now a thin, non-detached method that simply returns `update()`'s own task directly instead of tracking a separate task itself. This ensures a concurrent manual update and an auto update share the exact same underlying task regardless of which caller started it. It never awaits the task to completion, since update() restarts Supervisor and awaiting that here could drop the caller's connection before a response is sent. - `Updater.fetch_data()` is now `detach=True` with `concurrency=JobConcurrency.REJECT`, replacing the bespoke `start_fetch_data()`/`_fetch_task` tracking. `reload()` and `BackupManager._check_supervisor_version()` await the returned task directly to get the real outcome of a fetch instead of relying on the throttle window. - The Supervisor update API endpoint and the Home Assistant Core install retry loop now use `if task := await ...: await task` instead of catching `SupervisorJobError`, since a concurrent call no longer raises that error - it returns the shared in-progress task instead. Updated tests across supervisor.py, updater.py, the API and Home Assistant Core install retry loop to match the new return values and removed now-obsolete `SupervisorJobError`-based concurrency tests. * Fix too-many-nested-blocks pylint warning in HomeAssistantCore.install() Flatten the nested need_update/auto_update if/else into an if/elif so the Supervisor update retry try/except isn't nested one level deeper, which pylint 4.0.8 (unlike the previously installed version in this container) now flags as too-many-nested-blocks (6/5). Cache need_update in a local variable since it's now referenced twice in the same iteration and must be consistent between both checks. * Update supervisor/api/supervisor.py --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Mike Degatano <michael.degatano@gmail.com> Co-authored-by: Stefan Agner <stefan@agner.ch> |
||
|
|
24abbd2eba |
Consume the exception of a finished detached job task (#7217)
* Consume the exception of a finished detached job task With detach, errors raised by the job body surface on the returned task rather than to the caller. A caller that starts a detached job without awaiting the task, such as a scheduled reload starting the Supervisor auto update, would then leave an unretrieved exception that asyncio reports with a full traceback when the task is garbage collected. The wrapper has already captured the error on the job and logged it, so that report only adds noise. Retrieve the exception in the done callback the decorator already registers on every detached task. This only marks it as retrieved: callers that await the task still receive the exception. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Log a consumed detached job failure A HassioError is only logged when it is raised with a logger, so consuming the exception of an unawaited detached task could leave a failure without any trace once the job record is cleaned up. Log one warning line naming the job. A JobException is skipped because the wrapper already logged it with its traceback. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
fb84e69f29 |
Reinstall the stored Core version when its image is lost (#7204)
* Reinstall the stored Core version when its image is lost When the Docker storage is wiped, for example by the containerd snapshotter migration on Home Assistant OS 17 or a Docker storage reset, the image of the installed Home Assistant Core is gone. On the next start Supervisor treated this like a fresh installation: it installed the landingpage and then pulled the latest Core release. Users pinned to an older Core unexpectedly ended up on the newest release. Remember the stored version in load() and, if the image is missing, pull that version directly instead of going through the landingpage. Apps already behave this way through the repair path. Core then starts through the regular startup sequence once the image is present. A registry rate limit retries the pull like the landingpage install does. Any other pull failure, such as a version that is no longer published, logs a warning and falls back to the landingpage and the latest release. New installations, where no version is stored, are unaffected. Move the periodic progress logger into a method so the reinstall logs download progress as well. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Address review feedback for the Core reinstall Use the install_image property for the reinstall as well, so all Core install paths pick the image the same way. Let install_image fall back to the default image for this system when the updater has no image information yet. This makes the wait-for-updater loop in the landingpage install obsolete; the landingpage is preinstalled on Home Assistant OS anyway, so that loop only ran on Supervised systems. Drop the periodic progress logging from the reinstall. It was introduced for the first installation where the landingpage displays the Supervisor logs, which does not apply here. This keeps the progress logger local to install() again. Rework the new tests to mock at the Docker layer instead of patching methods on the class under test, so they assert which images are pulled and that no container gets created rather than that a method was not called. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
cd2368c5cd |
Bump sentry-sdk from 2.68.1 to 2.69.1 (#7215)
Bumps [sentry-sdk](https://github.com/getsentry/sentry-python) from 2.68.1 to 2.69.1. - [Release notes](https://github.com/getsentry/sentry-python/releases) - [Changelog](https://github.com/getsentry/sentry-python/blob/master/CHANGELOG.md) - [Commits](https://github.com/getsentry/sentry-python/compare/2.68.1...2.69.1) --- updated-dependencies: - dependency-name: sentry-sdk dependency-version: 2.69.1 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
1ab819a82f |
Bump ruff from 0.16.6 to 0.16.7 (#7216)
Bumps [ruff](https://github.com/astral-sh/ruff) from 0.16.6 to 0.16.7. - [Release notes](https://github.com/astral-sh/ruff/releases) - [Changelog](https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md) - [Commits](https://github.com/astral-sh/ruff/compare/0.16.6...0.16.7) --- updated-dependencies: - dependency-name: ruff dependency-version: 0.16.7 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
32545f0147 |
Bump time-machine from 3.5.0 to 3.5.1 (#7214)
Bumps [time-machine](https://github.com/adamchainz/time-machine) from 3.5.0 to 3.5.1. - [Changelog](https://github.com/adamchainz/time-machine/blob/main/docs/changelog.rst) - [Commits](https://github.com/adamchainz/time-machine/compare/3.5.0...3.5.1) --- updated-dependencies: - dependency-name: time-machine dependency-version: 3.5.1 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
b44b4acbc2 |
Add detach option to the Job decorator (#7211)
* Add detach option to the Job decorator A job decorated with detach=True awaits its conditions, concurrency control and throttling as before, but then runs the method in a separate task and returns that task to the caller. A refused job returns None (or raises on_condition, as today). With REJECT concurrency a call while the job is running returns the running task instead of raising. The task releases the concurrency lock and cleans up the job record when it completes. Group concurrency is not supported in this mode. This gives callers a definitive answer on whether a long running job was started, without having to await its completion. The Supervisor auto update needs this: a backup restore that requires a newer Supervisor has to know whether the update actually started before it tells the caller to wait for it, and the update itself stops the Supervisor, so it cannot be awaited from within an API request. Until now the only way to get that answer was to eagerly start the wrapped call as a task and inspect its done state, which silently depends on none of the job conditions suspending before the refusal. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Run detached jobs as root jobs and record throttled calls up front A detached job runs in a task created through sys_create_task, which clears the current job in the new context. A job created with the caller's job as parent therefore could not start there and raised JobStartException, which broke the motivating case of starting the Supervisor auto update from within the backup restore job. Create detached jobs as root jobs instead, as schedule_job already does. A detached job may outlive its caller, so it should not be a child of it anyway. The throttle bookkeeping happened inside the task, after the wrapper had returned. Two callers that were both runnable could pass the throttle check before either task ran and both start. Record the accepted call synchronously right after the throttle check in both modes; in the non-detached mode nothing suspends between the old and new place, so that behavior is unchanged. Add tests for a detached call from within a job and for concurrent calls to a throttled detached job. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Start detached job tasks eagerly so cancellation cannot skip cleanup A task cancelled before its first execution step never enters its coroutine, so the finally block in _run_detached would not run. The concurrency lock would stay held and the job would remain registered. Start the detached task eagerly so the runner is inside its try block before the task is handed to the caller. Add a test that cancels the returned task right away and checks the job is removed and a REJECT job can be started again. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Drop the reference to a finished detached job task The decorator lives for the process lifetime, so keeping the last detached task after it finished retained its result or exception traceback, along with the frames and arguments those reference, until the next call replaced it. Clear the reference from a done callback, guarded by identity so an older task cannot clear a newer one. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
f8c7214617 |
Treat containers with corrupt storage metadata as missing (#7186)
* Treat containers with corrupt storage metadata as missing Since Docker 29.4, containers whose RW layer fails to load while the daemon restores its state at startup (e.g. after storage metadata was corrupted by an unclean shutdown) are kept registered so they can still be removed, instead of being dropped (moby/moby#51724). Every other operation on such a container fails with a 500 "RWLayer of container <id> is unexpectedly nil" error, including the inspect at the start of nearly every Supervisor container operation. A single corrupt container record therefore permanently breaks starting, restarting, rebuilding and backing up of the affected add-on, plugin or Home Assistant Core: the start path errors out on the initial state check, before reaching the cleanup that would remove the broken record. Older Docker versions dropped such containers at daemon startup, so Supervisor saw them as missing and recreated them, healing the installation as a side effect. Restore that behavior explicitly by reporting a container as missing when a Docker API call fails with this error. Recovery paths then recreate the container as before, with stop_container's force delete (which no longer inspects first since #7175) removing the broken record along the way. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Broaden corrupt container detection and heal at start/restart The detection matched only one shape of the nil-RW-layer message, the inspect-time one built from the container id. Docker phrases the same condition differently per call site: with the container name instead of its id, with reversed word order on the containerd image store, and bare on Windows. Loosen the pattern to cover them; "unexpectedly nil" appears nowhere else in Docker, so it stays just as narrow. With the containerd image store the inspect of such a container succeeds entirely (the RW layer check in daemon/inspect.go is graph driver only), so none of the inspect-time handling fires: the error only surfaces from the actual start or restart call. Handle it there as well by removing the broken record and reporting the container as missing, so the next start takes the recreate path. Sentry shows this shape in the field (SUPERVISOR-1B6Q). Also soften the corrupt-record comment: the daemon marks the record once while restoring state, for any error loading the RW layer, and a daemon restart may recover it. Removal and recreation is still the safe recovery for Supervisor-managed containers, which hold no state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Document why corrupt container removal stays at the job-locked sites Review feedback asked to remove the corrupt record right where it is detected instead of reporting the container as missing. The record indeed cannot be used for anything until it is removed, but removal has to go by name, and the read paths (is_running() and friends) run outside the job locks that serialize a container's lifecycle. Names are only unique at a given instant — the heal deletes the broken record and then recreates the container under the same name — so a delete-by-name from an unlocked path after a stale inspect could hit the freshly recreated container rather than the broken one. Removal therefore stays where the job lock is held: the inner start and restart failures (added in the previous commit for the containerd image store, where only those calls see the error) and stop_container's cleanup, which every recreate flow reaches moments after detection. The read paths keep reporting the container as missing without acting on it. Say so in the code, and pin the split in tests: the read paths must not remove the record, the locked sites must. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
42502a1d0b |
Bump types-pyyaml from 6.0.12.20260815 to 6.0.12.20260906 (#7208)
Bumps [types-pyyaml](https://github.com/python/typeshed) from 6.0.12.20260815 to 6.0.12.20260906. - [Commits](https://github.com/python/typeshed/commits) --- updated-dependencies: - dependency-name: types-pyyaml dependency-version: 6.0.12.20260906 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
54a430f55e |
Hide new ntp host feature from v1 /host/info features field (#7207)
The python-supervisor-client library's HostInfo model types `features` as the strict `list[HostFeature]` instead of `list[HostFeature | str]`: https://github.com/home-assistant-libs/python-supervisor-client/blob/3459d25b7b223dfd2a49e5aa3f794ea822faab5c/aiohasupervisor/models/host.py#L49 #7181 added HostFeature.NTP, which now breaks the hassio integration on nightly since the client library raises on parsing it: Failed setup, will retry: Field "features" of type list[HostFeature] in HostInfo has invalid value ['reboot', 'shutdown', 'services', 'network', 'hostname', 'timedate', 'os_agent', 'haos', 'resolved', 'journal', 'disk', 'mount', 'ntp'] Work around this on the v1 /host/info endpoint by hiding "ntp" from the features field and exposing the full list via a new all_features field instead. The v2 API is unaffected since it has not shipped yet and its features field is already typed as list[HostFeature | str]. |
||
|
|
f5c3fa6ae1 |
Bump gitpython from 3.1.61 to 3.1.62 (#7209)
Bumps [gitpython](https://github.com/gitpython-developers/GitPython) from 3.1.61 to 3.1.62. - [Release notes](https://github.com/gitpython-developers/GitPython/releases) - [Changelog](https://github.com/gitpython-developers/GitPython/blob/main/CHANGES) - [Commits](https://github.com/gitpython-developers/GitPython/compare/3.1.61...3.1.62) --- updated-dependencies: - dependency-name: gitpython dependency-version: 3.1.62 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
2cd1fb514f |
Bump deepmerge from 3.0 to 3.0.1 (#7198)
Bumps [deepmerge](https://github.com/toumorokoshi/deepmerge) from 3.0 to 3.0.1. - [Release notes](https://github.com/toumorokoshi/deepmerge/releases) - [Commits](https://github.com/toumorokoshi/deepmerge/compare/v3.0...v3.0.1) --- updated-dependencies: - dependency-name: deepmerge dependency-version: 3.0.1 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
5b5f482cf8 |
Bump ruff from 0.16.5 to 0.16.6 (#7199)
Bumps [ruff](https://github.com/astral-sh/ruff) from 0.16.5 to 0.16.6. - [Release notes](https://github.com/astral-sh/ruff/releases) - [Changelog](https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md) - [Commits](https://github.com/astral-sh/ruff/compare/0.16.5...0.16.6) --- updated-dependencies: - dependency-name: ruff dependency-version: 0.16.6 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
0a7aa22405 |
Fix OOM/memory growth in tests/api by clearing aiohttp's middleware cache (#7196)
aiohttp caches built middleware chains in a module-level, process-wide functools.lru_cache(maxsize=1024) keyed on (handler, apps tuple). Our test suite builds a fresh RestAPI/web.Application (bound to a fresh CoreSys) for nearly every test, so almost every test call is a cache miss that inserts a new, never-reused entry - each one pinning an entire CoreSys object graph (docker mocks, D-Bus proxies, aiohttp routes, etc.) in memory until 1024 stale entries force LRU eviction. Over a full tests/api run this causes steady, effectively unbounded memory growth that reliably OOMs. Add an autouse fixture that clears aiohttp's internal _cached_build_middleware cache after every test. Verified with the full tests/api suite (1097 tests): now completes in ~2:15 with peak RSS ~490MB, versus growing past several GB and crashing before this change. |
||
|
|
8684f7e4cd |
Add /time API for setting NTP servers on HAOS (#7181)
* Add `/ntp` API for setting NTP servers on HAOS Add API for customizing NTP servers, using fixed DBus methods available since OS 18.3. The API is limited to HAOS only with OS Agent that sets the values in a daemon drop-in config. Changing the values triggers a systemd-timesyncd unit restart. The API currently exposes only the configuration overrides, not the currently applied value (like `timedatectl show-timesync --property=ServerName` does). If we want to add the effective NTP server, another field should be introduced in the API for this. Closes #6278 * Rename `/ntp` to `/time` * Nest options under `config` key Move the config options under a separate key to distinguish between runtime data that we plan to add. * Tighten NTP server regexp * Require systemd D-Bus for the NTP host feature * Always restart timesyncd when NTP servers are set |
||
|
|
c475bae6d8 |
Bump pylint from 4.0.7 to 4.0.8 (#7191)
Bumps [pylint](https://github.com/pylint-dev/pylint) from 4.0.7 to 4.0.8. - [Release notes](https://github.com/pylint-dev/pylint/releases) - [Commits](https://github.com/pylint-dev/pylint/compare/v4.0.7...v4.0.8) --- updated-dependencies: - dependency-name: pylint dependency-version: 4.0.8 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
fbab5e10a5 |
Report mount storage usage from the probe, not cached state (#7188)
* Report mount storage usage from the probe, not cached state The per-mount usage endpoint refused any mount whose cached state was not active. Since mounts became autofs-triggered, that state is only as fresh as the last reconcile probe, fifteen minutes apart, so a share that came back stayed unreportable for up to that long even though a probe would have found it healthy and mounted it. Drop the gate and let the probe decide. Its statvfs walks with LOOKUP_AUTOMOUNT, so it activates a dormant trigger and then reports what it found, which also means a usage request is no longer strictly read-only with respect to system state - it can cause the mount it measures. A path that still does not cross a filesystem boundary after that statvfs is one whose trigger is gone and which reverted to a plain directory; the existing boundary check catches it and its comment now says so. With the gate removed that is the only remaining not-mounted condition, so the mount_usage_not_active_error key, which named the cached state the endpoint no longer consults, gives way to mount_usage_not_mounted_error, which names what the probe observed. * Tighten comments for probe-authoritative mount usage. Cached state can be stale; the live probe is the authority, including dormant automount activation. |
||
|
|
75e47083b9 |
Rename Core port reservation units to drop the hassio prefix (#7173)
* Rename Core port reservation units to drop the hassio prefix The transient units introduced in #7131 were named hassio-port-reserve. Unlike most internal names, these are host visible: they show up in systemctl list-units, in journalctl and in any bug report from a boot where the reservation was involved. "hassio" is legacy branding we are moving away from, so avoid baking it into a user facing name. Rename both units to homeassistant-core-port-reserve, which says whose port is being held. The units are transient and only exist during the pre-Core boot window, so there is no upgrade or compatibility concern: a Supervisor update landing between reservation and release would at worst leave one inactive unit behind until the next host reboot. Also reflow the _reserve_core_port docstring so the reference to the hassio Docker bridge network reads as the network name it actually is. * Assert literal unit names in port reservation tests The tests imported _PORT_RESERVE_UNIT/_PORT_RESERVE_SERVICE and compared them against themselves, so every assertion held for any value the constants happened to have -- the rename in the previous commit passed without a single test failing. Use the literal unit names instead. That pins the two things systemd actually cares about: the .socket/.service suffixes it derives the unit type from, and the shared basename that tells it which service the socket activates.2026.09.0 |
||
|
|
30d6265768 |
Add Docker storage driver as a Sentry tag (#7187)
The storage driver is only available as event context today, which means Sentry cannot aggregate over it. Tag events with it in addition, so issues can be filtered and broken down by storage driver. This helps to spot problems specific to a driver (e.g. containerd snapshotter vs overlay2). Set the tag during setup as well, mirroring the context information which is already reported in that state. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c156937281 |
Bump coverage from 7.15.4 to 7.16.0 (#7183)
Bumps [coverage](https://github.com/coveragepy/coveragepy) from 7.15.4 to 7.16.0. - [Release notes](https://github.com/coveragepy/coveragepy/releases) - [Changelog](https://github.com/coveragepy/coveragepy/blob/main/CHANGES.rst) - [Commits](https://github.com/coveragepy/coveragepy/compare/7.15.4...7.16.0) --- updated-dependencies: - dependency-name: coverage dependency-version: 7.16.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
2de3f196ee |
Bump gitpython from 3.1.60 to 3.1.61 (#7184)
Bumps [gitpython](https://github.com/gitpython-developers/GitPython) from 3.1.60 to 3.1.61. - [Release notes](https://github.com/gitpython-developers/GitPython/releases) - [Changelog](https://github.com/gitpython-developers/GitPython/blob/main/CHANGES) - [Commits](https://github.com/gitpython-developers/GitPython/compare/3.1.60...3.1.61) --- updated-dependencies: - dependency-name: gitpython dependency-version: 3.1.61 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
9adf131901 |
Skip redundant inspect calls when constructing Docker container handles (#7175)
* Skip redundant inspect calls when constructing Docker container handles Several places in the Docker layer called containers.get(id), which performs an inspect API call, purely to construct a DockerContainer handle immediately followed by another call (show(), exec(), attach()) that either performs its own inspect or doesn't need one at all. Switch these spots to containers.container(id), which builds the same handle locally with no I/O, removing the redundant round-trip. Affected call sites: container_is_initialized(), stop_container() and container_run_inside() in docker/manager.py, _get_container() and _attach() in docker/interface.py, attach()/retag()/update_start_tag() in docker/supervisor.py, and write_stdin() in docker/app.py. Since containers.container()-built handles don't carry the real container ID (unlike get(), which populates it from the inspect response), _attach() and DockerSupervisor.attach() now read the ID from the inspect payload (self._meta["Id"]) instead of the handle's .id property. This intentionally excludes container_stats(), which has its own one-shot inspect+stats opportunity being handled separately in #7168. Follow-up to review discussion in #7168: https://github.com/home-assistant/supervisor/pull/7168#discussion_r3883548019 * Fix tests broken by containers.container() switch Several tests simulated Docker errors/state by mocking containers.get() (the old call site), but the production code now builds container handles via the sync, no-I/O containers.container() instead and only calls get() directly from the small set of call sites intentionally left unconverted (DockerApp's legacy-name migration in attach(), and container_stats() which is handled by #7168). Retarget these tests to mock containers.container() (or both, where a single test exercises both an app's attach() and a plugin/core's attach()) so the simulated failures actually reach the code path they're meant to test. Also add the missing "Id" key to some mocked container inspect payloads in test_check_docker_config.py: _attach() now reads the container ID from the inspect response instead of the handle's .id, so tests need that field present. * Skip unnecessary inspect before stopping container A reviewer pointed out that calling show() before stop() in stop_container() is redundant I/O. Docker returns a 304 (not an error) if the container is already stopped, 404 if it does not exist, so we can call stop() directly and handle those responses without an extra inspect call. * Centralize container fetch error handling in DockerSupervisor attach(), retag() and update_start_tag() each duplicated a try/except block around containers.container(name).show(). Reuse the existing _get_container() helper from DockerInterface instead, raising DockerNotFound explicitly when the container is missing. |
||
|
|
324d17e68b |
Use systemd automount units for network mounts (#7167)
* dbus: support aux units in start_transient_unit
Extend `Systemd.start_transient_unit` to accept the `aux` parameter
(`a(sa(sv))`), which has been hardcoded to `[]` since the wrapper was
introduced. Aux entries take the form `(unit_name, properties)` and
let systemd create multiple transient units atomically — most usefully
a `.mount` and its `.automount` companion in one D-Bus call.
Add the unit-property constants we'll need to drive that:
- `DBUS_ATTR_LAZY_UNMOUNT` ("LazyUnmount"): set on the `.mount` so
systemd umounts with MNT_DETACH. Pairs with `softerr`/`soft` to make
stop/restart sequences reliable even when the server is gone.
- `DBUS_ATTR_TIMEOUT_IDLE_USEC` ("TimeoutIdleUSec"): set on the
`.automount` to control how long the mount stays around after the
last access before autofs expires it.
- `DBUS_ATTR_WHERE` ("Where"): mount-point property; required on the
`.automount` unit (and useful explicitly on `.mount` units too).
No call sites change in this commit — pure plumbing.
* mounts: pair each network .mount with a .automount companion
Network mounts now get a transient `.automount` unit created
atomically alongside their `.mount`, via the aux parameter on
`StartTransientUnit`. The systemd-managed autofs trigger handles
activation lazily: the path exists in the VFS even before the
underlying network mount runs, and the first access that crosses the
trigger fires `mount.cifs`/`mount.nfs` on demand.
Why this matters:
- PID 1's path-walks (chase, daemon-reload, generator scans) stop
crossing dead NFS/CIFS lookups — autofs returns from kernel memory
rather than entering the network filesystem. The class of failure
fixed at the systemd-timeout layer in #6834 stops being reachable
in the first place.
- Failed accesses fail fast, bounded by the `.mount`'s `TimeoutSec=`
rather than hanging forever.
- Background reconnect (NFSv4 state-manager kthread, CIFS
delayed_work) recovers transparently when the server returns; no
remount needed.
The `.automount` carries `TimeoutIdleUSec=5min` so kernels can expire
the mount after inactivity and re-trigger on the next access. The
companion `.mount` carries `LazyUnmount=true` (MNT_DETACH on stop),
so umount returns immediately even when the server is unreachable;
existing fds drain in their soft/softerr timeout regime instead of
pinning the umount syscall.
`Mount.unmount()` now stops the `.automount` first so the autofs
trigger can't re-fire the underlying `.mount` during cleanup. The
automount stop is best-effort — if the unit is gone or refuses to
stop, we log and proceed to the `.mount` stop, which remains
authoritative.
Bind mounts opt out via a `creates_automount` class flag — they have
no server to wait on and the lazy semantics would only add
indirection. Network mounts (CIFS/NFS) opt in.
This is wiring only; the periodic mount reload, the reload/restart
escalation, and the inner bind-mount layer for media/share usage are
still in place. The next commit removes them now that automount
makes them unnecessary.
* mounts: drop bind layer, periodic reload, and reload/restart machinery
With the kernel's autofs trigger handling lazy activation and
transparent reconnect, almost all of the supervisor-side mount
choreography becomes unnecessary. Strip it out.
What goes away:
* The inner bind-mount layer. Media/share mounts used to live at
`path_extern_mounts/{name}` and have a second `.mount` unit bind
them into `path_extern_media/{name}` (resp. `share`). The
`.automount` companion can sit directly at the container-facing
path, and the parent-dir RSLAVE bind into add-on containers
surfaces the autofs trigger the same way. `NetworkMount.where`
now switches on usage:
- MEDIA → `path_extern_media/{name}`
- SHARE → `path_extern_share/{name}`
- BACKUP → `path_extern_mounts/{name}` (unchanged)
* `BindMount`, `BoundMount`, `MountManager._bound_mounts`, the
`_bind_mount`/`_bind_media`/`_bind_share` helpers, and the entire
emergency-fallback dance (`path_emergency`). The empty-read-only
dir trick was a workaround for the harder PID 1 wedge problem,
which autofs solves at the root. Failed shares now surface as
ETIMEDOUT/EHOSTDOWN on access plus a resolution issue.
* `MountManager.reload()` and the 15-minute `RUN_RELOAD_MOUNTS`
periodic task. autofs re-activates on access; we don't need to
poll, and not polling means we don't randomly trip stale-FH /
softreval / dead-server edge cases on a timer.
* `Mount.reload()` and `Mount._restart()`. The reload→restart
escalation existed to make `is_mounted` honest in the face of
systemd's local-only state; now `is_mounted` IS honest (probe-
based), and any recovery the kernel knows how to do happens
inside autofs without our involvement.
* The `RELOADING` safety net introduced by #6834. That whole class
of PID 1 wedge stops being reachable when path lookups don't
cross dead network mounts.
What changes shape:
* `NetworkMount.is_mounted()` no longer gates on systemd's
`ActiveState` before probing. The `.mount` unit is dormant
whenever autofs hasn't recently triggered it, so systemd state
is meaningless as a health signal. The probe (now via
`statvfs("/path/.")` so the trailing dot forces `LOOKUP_DIRECTORY`
and triggers autofs) is the source of truth, and after it returns
we set `self._state` to ACTIVE/INACTIVE so the API reports
reachability rather than autofs idle state.
* `MountManager.reload_mount()` becomes probe-only — no systemd
reload/restart calls. The user-facing semantics: "tell me if this
mount is reachable right now, and refresh the resolution issue
accordingly." If the mount is dead, autofs will re-trigger it on
the next consumer access; the supervisor doesn't need to force it.
* `Mount.load()` keeps `_update_state_await` for the no-job-dispatched
case where a previous supervisor left a unit in `activating`, then
probes once via `is_mounted` so state reflects reachability rather
than the lazy-mount idle state.
* `BackupManager` no longer walks `bound_mounts`. Backups skip
network-mount subdirectories during folder archive (so we don't
recurse into the share), and unmount/re-mount them around folder
restores (so writes target the local mount-point dir, not the
remote share).
Behavior tradeoffs accepted:
* Dead media/share access now ~30s ETIMEDOUT instead of empty
read-only dir. The empty-dir was a workaround for the PID 1
wedge that autofs eliminates at the root; if specific add-ons
rely on the old behavior we can revisit with a 2-stage automount
whose fallback target is an emergency dir.
* Resolution-issue lag: failed mounts no longer surface within
15 min on their own. They appear when the user probes via the
API, when `BackupManager.reload()` walks the location, or when
load-time activation fails.
* tests: adapt bind-layer-era tests after rebase
Main gained tests for the bind layer after this branch was written:
the #7013 rebind regression test and the #7072 bind-step rollback test
cover machinery that no longer exists, and the healthy-reload test
asserted the unconditional rebind. Drop the first two and reduce the
third to asserting that a healthy probe performs no systemd operations
at all.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: rearchitect automount setup from systemd/kernel review
A code-level review of systemd's automount implementation (verified
against v254.13, the version HAOS ships) and the kernel autofs/VFS
plumbing surfaced three correctness-critical flaws in the automount
design, plus several hardening gaps. See automount-rearchitecture.md
for the full analysis.
- Make the .automount the primary transient unit with the .mount as
aux, mirroring `systemd-mount --automount=yes`. Aux units get no
start job: with the .mount as primary the trigger was never armed
and the design silently degraded to eager mounting.
- Set StartLimitIntervalUSec=0 on the .mount. With the default start
rate limit (5 starts/10s, counting successful starts) a fast-failing
mount plus any polling consumer trips the limit within seconds, and
systemd then detaches the autofs trigger entirely
(AUTOMOUNT_FAILURE_MOUNT_START_LIMIT_HIT) — the path silently
becomes a plain writable local directory.
- Drop the 5-minute TimeoutIdleUSec (default 0 = never expire).
Kernel idle expiry decides busyness via may_umount_tree(), which
only counts the init-namespace mount instance — open files held by
container processes are invisible, so expiry would unmount shares
under actively writing add-ons.
- Probe with a plain statvfs: the statfs syscall walks with
LOOKUP_AUTOMOUNT and triggers by itself; the trailing-dot trick is
unnecessary. Classify ELOOP as a mount-propagation
misconfiguration in the probe error handling.
- Re-arm the trigger from reload_mount(): if the .automount unit is
failed or gone (e.g. autofs unmounted out-of-band), reset failure
state and re-create the pair before probing.
- Tear down legacy eager-mount units during load(). The old design's
bind unit for media/share occupies the exact unit name the network
.mount uses now; on a warm upgrade adoption would mistake the old
bind mount for the network mount and leak the legacy data mount.
- Honor the systemd job result when stopping the .mount during
unmount and reset failure state on both units afterwards so dead
transient units get garbage-collected.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: adapt local data repair to the autofs design
Port of the mount-target-not-empty repair (#7089) on top of the
automount rearchitecture:
- relocate_local_data() handles a single target directory — the mount
sits directly at its container-facing path, the bind layer and with
it the multi-directory case are gone
- a repair_trigger() failure caused by blocking local data raises the
mount failed issue with the move_local_data suggestion instead of
the plain variant
- local data at load is detected via the mount unit itself; the
bind-layer detection test is rewritten accordingly
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docker: use rslave propagation for execute_command share mount
The temporary container used for execute_command (e.g. core config
check) mounted /share without a propagation mode, unlike every other
share/media mount. With eager mounts this only meant missing mounts
made after container start; with lazy automount activation the shares
are routinely mounted after start, and accessing a not-yet-activated
automount from a private mount namespace fails with ELOOP.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: harden automount lifecycle from review findings
Address the findings of an adversarial review of the automount
rearchitecture:
- load() no longer adopts a dead automount trigger: a failed or
stopped .automount leaves the path a plain writable directory — the
silent local-write degradation this design must prevent. Tear down
and re-arm instead. An adopted pair whose probe fails now raises so
the manager surfaces the mount failed issue, same as a fresh mount.
- unmount() resolves the .mount unit only after stopping the
.automount (the lazy detach can garbage-collect the transient unit,
invalidating an earlier proxy), raises instead of warning when the
automount stop errors, and checks the stop job result so a failed
stop cannot leave an armed trigger behind a "successful" removal.
- repair_trigger() fully unmounts before re-arming: with the share
still attached, arming fails — or worse, mount() misreads the
mounted share's contents as blocking local data.
- Legacy unit teardown honors the stop job result for the same reason.
- reload_mount() escalates once to re-creating the unit pair when an
established mount is unreachable: a permanently dead session (e.g.
replaced server) keeps the path mounted, so the trigger can never
re-fire and kernel reconnection never succeeds. Teardown is safe now
(lazy unmount, no PID 1 path walks), unlike the removed reload →
restart escalation of the eager design.
- Reinstate the periodic mounts task as a probe-based reconcile: re-arm
dead triggers, refresh the reachability state reported by the API and
used by backup locations, and sync the mount failed issue in both
directions. No reload or restart of established mounts.
- Folder restore no longer fails after a successful restore when
re-mounting nested mounts cannot verify an unreachable server — the
trigger is armed, the share recovers on next access.
- Drop the now-unused Mount.update(), deduplicate unit name escaping,
hoist the backup exclusion set out of the per-file filter.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* tests: expect rslave propagation in core check container
The config check assertion missed the update for the rslave
propagation on the execute_command share mount.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* tests: cover reconcile trigger repair failure paths
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: arm automount before stopping the legacy data mount
Address review: the legacy teardown stopped the media/share bind unit
and the eager data mount in sequence, both strictly. When the bind
stop succeeded but the data mount stop failed — its unmount can time
out against an unreachable server, legacy units have no LazyUnmount —
load() raised with the container-facing path left a plain writable
directory and no trigger armed: the pollution mode this design is
meant to prevent. Reordering the stops would not help either, as
systemd stops the bind first anyway as a dependent of the data mount
(the implicit Requires= from RequiresMountsFor= on the bind's What=).
Instead, strictly stop only the path-conflicting unit (a failure
leaves the path covered, which is safe and retryable), arm the
automount right after, and stop the conflict-free legacy data mount
best-effort last — a failed stop logs a warning and leaves an orphaned
mount for the next Supervisor restart or a host reboot, with nothing
writable exposed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: discard dead session instead of re-creating units on reload
Address review: re-creating the unit pair on reload of an unreachable
established mount exposed the target as a plain writable directory
between the lazy unmount and arming the replacement — a continuously
writing add-on could block the re-arm or slip writes under the new
mount in the check-to-arm window.
Stop only the .mount unit instead, keeping the .automount armed:
systemd re-installs the autofs trigger over the path — the same
mechanism idle expiry uses, with the automount's Triggers= reference
keeping the transient .mount definition alive — so the path is never
locally writable. The re-probe then mounts fresh through the trigger,
which establishes a new session and thereby covers the permanently
dead session case (e.g. a replaced server) the escalation exists for.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: surface arming failures, keep restore teardown re-armed
Address review findings on the arming/re-arming error paths:
- mount() checks the StartTransientUnit job result: a failed start job
means the trigger never armed and the path is a plain writable
directory — a hard MountError before the probe, so it cannot be
mistaken for the armed-but-unreachable MountActivationError.
- Folder restore includes the nested-mount teardown in the try block:
an unmount failing halfway (trigger disarmed, share still attached)
previously exited before any re-arm, leaving the path unprotected.
The finally re-arms via repair_trigger(), which no-ops on a still
armed trigger and handles partially torn-down pairs.
- The post-restore re-arm suppresses only MountActivationError (armed,
recovers on next access). Any other failure — arming failed, local
data blocking the target — left the path unprotected and surfaces.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: fold the two legacy unit teardowns into one helper
The eager-mount-era cleanup queried the .mount unit a second time
although load() had just fetched it, and the strict teardown of the
unit occupying the automount's path and the best-effort teardown of
the mount at the mounts data directory were near identical. Reuse the
unit from load() and give the shared helper a strict flag, which drops
a D-Bus round trip per mount on every load.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: re-arm the trigger when unmount cannot stop the mount
Stopping the automount detaches the whole stack at the path, so an
unmount that then fails to stop the .mount leaves a plain writable
directory behind. Arm a fresh pair before raising, best effort — the
local data repair covers what lands there if that fails too.
The automount stop itself needs no such handling: systemd's
automount_stop() enters dead synchronously and the detach it performs
(MNT_DETACH, no server contact) only logs its errors, so a failure
there means systemd could not be reached at all.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: raise translatable errors for mount operations
Errors from setting up, unmounting and reloading a mount reach users
through the API, so describe what failed in a translatable message and
leave the systemd job result or D-Bus error to the log.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: remove the emergency folder of the eager-mount design
The read-only fallback directory has no place in the automount design.
Remove what is left of it on existing installations, next to the legacy
addons directory cleanup, keeping it if it holds anything but the empty
mount points.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: trim comments to the behavior they describe
Drop the narration of what changed and why from comments and
docstrings, keeping the notes that record why a tempting alternative
does not work. Two comments still described the reload to restart
escalation this branch removed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: correct stale claims in manager and reload test
The manager docstring credited the kernel with idle expiry, which is
deliberately disabled, and claimed no polling although the periodic
reconcile probes every 15 minutes. The reload test claimed systemd is
never contacted while the escalation stops the .mount unit; only reload
and restart of the unit are avoided.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: arm the trigger alone if re-creating the pair is rejected
Address review: the re-arm after a failed unmount submits the .mount as
an aux unit, but transient creation requires a pristine unit and the
definition is still loaded precisely when stopping it is what failed.
Fall back to creating the .automount on its own, which covers the path
and fires the surviving mount definition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* mounts: never arm the trigger over local data after a failed unmount
Address review: the fallback caught the target validation errors too,
so data written while the path was uncovered would end up beneath a
fresh trigger. Reconciliation would then find a healthy mount and the
repair skips mount points, leaving the data hidden for good. Leave the
path untouched instead, so the next reconcile offers to move it away.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
2d153ac497 |
Add one-shot mode and docker-given totals to stats API (#7168)
* Add one-shot mode and docker-given totals to stats API - Add one-shot query option to container stats to skip the calculated CPU percentage window, returning docker-given totals directly. - Refactor container_stats() with translatable, structured Docker errors and timeout handling instead of generic Docker exceptions. - Translate low-level Docker errors into component-flavored errors (NotRunningError/StatsTimeoutError/UnknownError) for Home Assistant, Supervisor, and each plugin (Audio, Cli, CoreDNS, Multicast, Observer) as well as apps, mirroring the existing App.stats() pattern for consistent, informative API responses. - Add a shared api_return_stats() helper to remove duplicated stats handling logic across the 8 stats API endpoints. - DockerContainerNotFoundError now inherits from DockerNotFound. * Contain aiodocker _query access and split v1/v2 stats response model - Contain the reach into aiodocker's protected _query internals (needed for one-shot stats until aio-libs/aiodocker#1054 lands) to a single DockerAPI._query_one_shot_stats() helper. - Split the stats API response model between v1 (legacy, windowed CPU percentage still available via opt-in one_shot query string) and v2 (always one-shot, no query string, cpu_percent dropped since it would never carry a meaningful value). Both share the same api_return_stats(stats, legacy=...) helper so the response model stays in one place. - DockerAPI.container_stats() keeps its existing one_shot keyword only; the legacy/v1 vs v2 distinction is handled entirely in the API layer. * Detect stopped/restarting containers in one-shot stats response Docker returns a stub response containing only the container's id/name (no cpu_stats/memory_stats/networks data) for a stopped or restarting container instead of an error, even in one-shot mode - the not-running check in daemon.ContainerStats happens before the one-shot branch, so requesting one-shot stats doesn't change this. Detect the missing online_cpus field (only ever present while running) and raise DockerContainerNotRunningError instead of returning a payload that DockerStats would otherwise turn into a successful all-zero response. * Add tests for stats error paths introduced in one-shot/totals PR - Cover DockerAPI.container_stats/_query_one_shot_stats edge cases: empty one-shot response, empty windowed response, and the raw _query_one_shot_stats helper. - Add stats() error-translation tests (not-running, timeout, unknown) for the audio, cli, multicast and observer plugins, none of which had any coverage. - Add stats endpoint tests (v1 + v2) for the cli, audio, dns, multicast and observer REST APIs, plus a v1 success-path test for the apps stats endpoint. These close the codecov-flagged gaps for newly introduced code in this PR. * Unify one-shot and windowed container_stats() code paths Both modes now share the same error handling and not-running detection: - The windowed path no longer inspects the container first to check its running state. Docker returns the same id/name-only stub response (no online_cpus in cpu_stats) for a stopped or restarting container in windowed mode as it does in one-shot mode, so the existing online_cpus presence check now covers both, closing a race window where a container stopping between the inspect and the stats call would previously reach the DockerStats parser with an incomplete payload. - The windowed path builds its container handle via containers.container(), which constructs a DockerContainer from the name with no I/O, instead of containers.get(), which performs an unnecessary inspect call just to fetch the same information the stats call already returns. * Fix protected access pylint issue |
||
|
|
5d6944a539 |
Bump ruff from 0.16.4 to 0.16.5 (#7179)
Signed-off-by: dependabot[bot] <support@github.com> |
||
|
|
3783aca1fe |
Bump gitpython from 3.1.59 to 3.1.60 (#7178)
Signed-off-by: dependabot[bot] <support@github.com> |
||
|
|
bc9c93791e |
Bump cryptography from 50.0.0 to 50.0.1 (#7180)
Signed-off-by: dependabot[bot] <support@github.com> |