Commit Graph
5911 Commits
Author SHA1 Message Date
Stefan AgnerandClaude Fable 5.1 9387c4ea65 Retry multicast cleanup after a failed disable
A failed container or image removal left the plugin persisted as
disabled but with its version and image still around, and every later
disable request returned early because the plugin was already disabled.

Move the removal into a dedicated step that stops the container, removes
the image using the stored version and only then forgets the version.
disable() always runs it, so a repeated request retries the cleanup, and
the disabled load path runs it as well to catch a disable interrupted by
a Supervisor exit. The step no longer relies on DockerInterface.remove,
which derives the version from container metadata that is not available
before an attach.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:56:35 +02:00
Stefan AgnerandClaude Fable 5.1 e950305d84 Silence pylint false positive on options schema
Add the same no-value-for-parameter disable the sibling API modules use
for module-level voluptuous schemas.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 14:05:50 +02:00
Stefan AgnerandClaude Fable 5.1 108ac2f15d Add enabled option to the multicast plugin
Since 2020 Supervisor unconditionally installs and runs the multicast
plugin on every installation. Users affected by its interference (mDNS
duplicates, multicast storms) have no supported way to turn it off. This
is phase 1 of the deprecation plan in architecture discussion #1457: an
opt-in disable switch.

The plugin config gains an `enabled` option (default true), persisted in
multicast.json and exposed via `GET /multicast/info` and a new
`POST /multicast/options` endpoint. Disabling stops and removes the
container, removes the image and forgets the installed version. Enabling
installs the current version and starts it, gated by the same job
conditions as a plugin update. While disabled, load skips attach, install
and start, the watchdog ignores container events, repair and auto-update
are skipped, and start/restart/update raise MulticastDisabledError
(HTTP 400).

The update endpoint now rejects a disabled plugin before comparing
versions, as the comparison raises on the None version left behind by
a disable.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 13:58:03 +02:00
Stefan AgnerandClaude Fable 5.1 fb84e69f29 Reinstall the stored Core version when its image is lost (#7204)
* Reinstall the stored Core version when its image is lost

When the Docker storage is wiped, for example by the containerd
snapshotter migration on Home Assistant OS 17 or a Docker storage reset,
the image of the installed Home Assistant Core is gone. On the next
start Supervisor treated this like a fresh installation: it installed
the landingpage and then pulled the latest Core release. Users pinned to
an older Core unexpectedly ended up on the newest release.

Remember the stored version in load() and, if the image is missing,
pull that version directly instead of going through the landingpage.
Apps already behave this way through the repair path. Core then starts
through the regular startup sequence once the image is present.

A registry rate limit retries the pull like the landingpage install
does. Any other pull failure, such as a version that is no longer
published, logs a warning and falls back to the landingpage and the
latest release. New installations, where no version is stored, are
unaffected.

Move the periodic progress logger into a method so the reinstall logs
download progress as well.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Address review feedback for the Core reinstall

Use the install_image property for the reinstall as well, so all Core
install paths pick the image the same way. Let install_image fall back
to the default image for this system when the updater has no image
information yet. This makes the wait-for-updater loop in the landingpage
install obsolete; the landingpage is preinstalled on Home Assistant OS
anyway, so that loop only ran on Supervised systems.

Drop the periodic progress logging from the reinstall. It was introduced
for the first installation where the landingpage displays the Supervisor
logs, which does not apply here. This keeps the progress logger local to
install() again.

Rework the new tests to mock at the Docker layer instead of patching
methods on the class under test, so they assert which images are pulled
and that no container gets created rather than that a method was not
called.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 10:46:22 +02:00
dependabot[bot] cd2368c5cd Bump sentry-sdk from 2.68.1 to 2.69.1 (#7215)
Bumps [sentry-sdk](https://github.com/getsentry/sentry-python) from 2.68.1 to 2.69.1.
- [Release notes](https://github.com/getsentry/sentry-python/releases)
- [Changelog](https://github.com/getsentry/sentry-python/blob/master/CHANGELOG.md)
- [Commits](https://github.com/getsentry/sentry-python/compare/2.68.1...2.69.1)

---
updated-dependencies:
- dependency-name: sentry-sdk
  dependency-version: 2.69.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 10:34:53 +02:00
dependabot[bot] 1ab819a82f Bump ruff from 0.16.6 to 0.16.7 (#7216)
Bumps [ruff](https://github.com/astral-sh/ruff) from 0.16.6 to 0.16.7.
- [Release notes](https://github.com/astral-sh/ruff/releases)
- [Changelog](https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md)
- [Commits](https://github.com/astral-sh/ruff/compare/0.16.6...0.16.7)

---
updated-dependencies:
- dependency-name: ruff
  dependency-version: 0.16.7
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 10:34:09 +02:00
dependabot[bot] 32545f0147 Bump time-machine from 3.5.0 to 3.5.1 (#7214)
Bumps [time-machine](https://github.com/adamchainz/time-machine) from 3.5.0 to 3.5.1.
- [Changelog](https://github.com/adamchainz/time-machine/blob/main/docs/changelog.rst)
- [Commits](https://github.com/adamchainz/time-machine/compare/3.5.0...3.5.1)

---
updated-dependencies:
- dependency-name: time-machine
  dependency-version: 3.5.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 10:21:18 +02:00
Stefan AgnerandClaude Fable 5.1 b44b4acbc2 Add detach option to the Job decorator (#7211)
* Add detach option to the Job decorator

A job decorated with detach=True awaits its conditions, concurrency
control and throttling as before, but then runs the method in a separate
task and returns that task to the caller. A refused job returns None (or
raises on_condition, as today). With REJECT concurrency a call while the
job is running returns the running task instead of raising. The task
releases the concurrency lock and cleans up the job record when it
completes. Group concurrency is not supported in this mode.

This gives callers a definitive answer on whether a long running job was
started, without having to await its completion. The Supervisor auto
update needs this: a backup restore that requires a newer Supervisor has
to know whether the update actually started before it tells the caller to
wait for it, and the update itself stops the Supervisor, so it cannot be
awaited from within an API request. Until now the only way to get that
answer was to eagerly start the wrapped call as a task and inspect its
done state, which silently depends on none of the job conditions
suspending before the refusal.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Run detached jobs as root jobs and record throttled calls up front

A detached job runs in a task created through sys_create_task, which
clears the current job in the new context. A job created with the
caller's job as parent therefore could not start there and raised
JobStartException, which broke the motivating case of starting the
Supervisor auto update from within the backup restore job. Create
detached jobs as root jobs instead, as schedule_job already does. A
detached job may outlive its caller, so it should not be a child of it
anyway.

The throttle bookkeeping happened inside the task, after the wrapper had
returned. Two callers that were both runnable could pass the throttle
check before either task ran and both start. Record the accepted call
synchronously right after the throttle check in both modes; in the
non-detached mode nothing suspends between the old and new place, so
that behavior is unchanged.

Add tests for a detached call from within a job and for concurrent calls
to a throttled detached job.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Start detached job tasks eagerly so cancellation cannot skip cleanup

A task cancelled before its first execution step never enters its
coroutine, so the finally block in _run_detached would not run. The
concurrency lock would stay held and the job would remain registered.
Start the detached task eagerly so the runner is inside its try block
before the task is handed to the caller.

Add a test that cancels the returned task right away and checks the job
is removed and a REJECT job can be started again.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Drop the reference to a finished detached job task

The decorator lives for the process lifetime, so keeping the last
detached task after it finished retained its result or exception
traceback, along with the frames and arguments those reference, until the
next call replaced it. Clear the reference from a done callback, guarded
by identity so an older task cannot clear a newer one.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-10 15:29:02 -04:00
Stefan AgnerandClaude Fable 5 f8c7214617 Treat containers with corrupt storage metadata as missing (#7186)
* Treat containers with corrupt storage metadata as missing

Since Docker 29.4, containers whose RW layer fails to load while the
daemon restores its state at startup (e.g. after storage metadata was
corrupted by an unclean shutdown) are kept registered so they can still
be removed, instead of being dropped (moby/moby#51724). Every other
operation on such a container fails with a 500 "RWLayer of container
<id> is unexpectedly nil" error, including the inspect at the start of
nearly every Supervisor container operation. A single corrupt container
record therefore permanently breaks starting, restarting, rebuilding and
backing up of the affected add-on, plugin or Home Assistant Core: the
start path errors out on the initial state check, before reaching the
cleanup that would remove the broken record.

Older Docker versions dropped such containers at daemon startup, so
Supervisor saw them as missing and recreated them, healing the
installation as a side effect. Restore that behavior explicitly by
reporting a container as missing when a Docker API call fails with this
error. Recovery paths then recreate the container as before, with
stop_container's force delete (which no longer inspects first since
 #7175) removing the broken record along the way.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Broaden corrupt container detection and heal at start/restart

The detection matched only one shape of the nil-RW-layer message, the
inspect-time one built from the container id. Docker phrases the same
condition differently per call site: with the container name instead of
its id, with reversed word order on the containerd image store, and bare
on Windows. Loosen the pattern to cover them; "unexpectedly nil" appears
nowhere else in Docker, so it stays just as narrow.

With the containerd image store the inspect of such a container succeeds
entirely (the RW layer check in daemon/inspect.go is graph driver only),
so none of the inspect-time handling fires: the error only surfaces from
the actual start or restart call. Handle it there as well by removing
the broken record and reporting the container as missing, so the next
start takes the recreate path. Sentry shows this shape in the field
(SUPERVISOR-1B6Q).

Also soften the corrupt-record comment: the daemon marks the record once
while restoring state, for any error loading the RW layer, and a daemon
restart may recover it. Removal and recreation is still the safe
recovery for Supervisor-managed containers, which hold no state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Document why corrupt container removal stays at the job-locked sites

Review feedback asked to remove the corrupt record right where it is
detected instead of reporting the container as missing. The record
indeed cannot be used for anything until it is removed, but removal has
to go by name, and the read paths (is_running() and friends) run outside
the job locks that serialize a container's lifecycle. Names are only
unique at a given instant — the heal deletes the broken record and then
recreates the container under the same name — so a delete-by-name from
an unlocked path after a stale inspect could hit the freshly recreated
container rather than the broken one.

Removal therefore stays where the job lock is held: the inner start and
restart failures (added in the previous commit for the containerd image
store, where only those calls see the error) and stop_container's
cleanup, which every recreate flow reaches moments after detection. The
read paths keep reporting the container as missing without acting on it.

Say so in the code, and pin the split in tests: the read paths must not
remove the record, the locked sites must.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-10 11:48:01 +02:00
dependabot[bot] 42502a1d0b Bump types-pyyaml from 6.0.12.20260815 to 6.0.12.20260906 (#7208)
Bumps [types-pyyaml](https://github.com/python/typeshed) from 6.0.12.20260815 to 6.0.12.20260906.
- [Commits](https://github.com/python/typeshed/commits)

---
updated-dependencies:
- dependency-name: types-pyyaml
  dependency-version: 6.0.12.20260906
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-10 11:46:02 +02:00
Mike Degatano 54a430f55e Hide new ntp host feature from v1 /host/info features field (#7207)
The python-supervisor-client library's HostInfo model types
`features` as the strict `list[HostFeature]` instead of
`list[HostFeature | str]`:
https://github.com/home-assistant-libs/python-supervisor-client/blob/3459d25b7b223dfd2a49e5aa3f794ea822faab5c/aiohasupervisor/models/host.py#L49

#7181 added HostFeature.NTP, which now breaks the hassio integration
on nightly since the client library raises on parsing it:

    Failed setup, will retry: Field "features" of type list[HostFeature] in HostInfo has invalid value ['reboot', 'shutdown', 'services', 'network', 'hostname', 'timedate', 'os_agent', 'haos', 'resolved', 'journal', 'disk', 'mount', 'ntp']

Work around this on the v1 /host/info endpoint by hiding "ntp" from
the features field and exposing the full list via a new all_features
field instead. The v2 API is unaffected since it has not shipped yet
and its features field is already typed as list[HostFeature | str].
2026-09-10 11:20:35 +02:00
dependabot[bot] f5c3fa6ae1 Bump gitpython from 3.1.61 to 3.1.62 (#7209)
Bumps [gitpython](https://github.com/gitpython-developers/GitPython) from 3.1.61 to 3.1.62.
- [Release notes](https://github.com/gitpython-developers/GitPython/releases)
- [Changelog](https://github.com/gitpython-developers/GitPython/blob/main/CHANGES)
- [Commits](https://github.com/gitpython-developers/GitPython/compare/3.1.61...3.1.62)

---
updated-dependencies:
- dependency-name: gitpython
  dependency-version: 3.1.62
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-10 10:28:45 +02:00
dependabot[bot] 2cd1fb514f Bump deepmerge from 3.0 to 3.0.1 (#7198)
Bumps [deepmerge](https://github.com/toumorokoshi/deepmerge) from 3.0 to 3.0.1.
- [Release notes](https://github.com/toumorokoshi/deepmerge/releases)
- [Commits](https://github.com/toumorokoshi/deepmerge/compare/v3.0...v3.0.1)

---
updated-dependencies:
- dependency-name: deepmerge
  dependency-version: 3.0.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-07 08:18:04 +02:00
dependabot[bot] 5b5f482cf8 Bump ruff from 0.16.5 to 0.16.6 (#7199)
Bumps [ruff](https://github.com/astral-sh/ruff) from 0.16.5 to 0.16.6.
- [Release notes](https://github.com/astral-sh/ruff/releases)
- [Changelog](https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md)
- [Commits](https://github.com/astral-sh/ruff/compare/0.16.5...0.16.6)

---
updated-dependencies:
- dependency-name: ruff
  dependency-version: 0.16.6
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-07 08:14:19 +02:00
Mike Degatano 0a7aa22405 Fix OOM/memory growth in tests/api by clearing aiohttp's middleware cache (#7196)
aiohttp caches built middleware chains in a module-level, process-wide
functools.lru_cache(maxsize=1024) keyed on (handler, apps tuple). Our test
suite builds a fresh RestAPI/web.Application (bound to a fresh CoreSys) for
nearly every test, so almost every test call is a cache miss that inserts a
new, never-reused entry - each one pinning an entire CoreSys object graph
(docker mocks, D-Bus proxies, aiohttp routes, etc.) in memory until 1024
stale entries force LRU eviction. Over a full tests/api run this causes
steady, effectively unbounded memory growth that reliably OOMs.

Add an autouse fixture that clears aiohttp's internal
_cached_build_middleware cache after every test. Verified with the full
tests/api suite (1097 tests): now completes in ~2:15 with peak RSS ~490MB,
versus growing past several GB and crashing before this change.
2026-09-04 16:23:42 -04:00
Jan Čermák 8684f7e4cd Add /time API for setting NTP servers on HAOS (#7181)
* Add `/ntp` API for setting NTP servers on HAOS

Add API for customizing NTP servers, using fixed DBus methods available
since OS 18.3. The API is limited to HAOS only with OS Agent that sets
the values in a daemon drop-in config. Changing the values triggers a
systemd-timesyncd unit restart.

The API currently exposes only the configuration overrides, not the
currently applied value (like `timedatectl show-timesync
--property=ServerName` does). If we want to add the effective NTP
server, another field should be introduced in the API for this.

Closes #6278

* Rename `/ntp` to `/time`

* Nest options under `config` key

Move the config options under a separate key to distinguish between
runtime data that we plan to add.

* Tighten NTP server regexp

* Require systemd D-Bus for the NTP host feature

* Always restart timesyncd when NTP servers are set
2026-09-02 17:22:20 +02:00
dependabot[bot] c475bae6d8 Bump pylint from 4.0.7 to 4.0.8 (#7191)
Bumps [pylint](https://github.com/pylint-dev/pylint) from 4.0.7 to 4.0.8.
- [Release notes](https://github.com/pylint-dev/pylint/releases)
- [Commits](https://github.com/pylint-dev/pylint/compare/v4.0.7...v4.0.8)

---
updated-dependencies:
- dependency-name: pylint
  dependency-version: 4.0.8
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-02 10:33:38 +02:00
Petar Petrov fbab5e10a5 Report mount storage usage from the probe, not cached state (#7188)
* Report mount storage usage from the probe, not cached state

The per-mount usage endpoint refused any mount whose cached state was not
active. Since mounts became autofs-triggered, that state is only as fresh
as the last reconcile probe, fifteen minutes apart, so a share that came
back stayed unreportable for up to that long even though a probe would
have found it healthy and mounted it.

Drop the gate and let the probe decide. Its statvfs walks with
LOOKUP_AUTOMOUNT, so it activates a dormant trigger and then reports what
it found, which also means a usage request is no longer strictly read-only
with respect to system state - it can cause the mount it measures.

A path that still does not cross a filesystem boundary after that statvfs
is one whose trigger is gone and which reverted to a plain directory; the
existing boundary check catches it and its comment now says so. With the
gate removed that is the only remaining not-mounted condition, so the
mount_usage_not_active_error key, which named the cached state the
endpoint no longer consults, gives way to mount_usage_not_mounted_error,
which names what the probe observed.

* Tighten comments for probe-authoritative mount usage.

Cached state can be stale; the live probe is the authority, including dormant automount activation.
2026-09-01 16:29:49 -04:00
Stefan Agner 75e47083b9 Rename Core port reservation units to drop the hassio prefix (#7173)
* Rename Core port reservation units to drop the hassio prefix

The transient units introduced in #7131 were named hassio-port-reserve.
Unlike most internal names, these are host visible: they show up in
systemctl list-units, in journalctl and in any bug report from a boot
where the reservation was involved. "hassio" is legacy branding we are
moving away from, so avoid baking it into a user facing name.

Rename both units to homeassistant-core-port-reserve, which says whose
port is being held. The units are transient and only exist during the
pre-Core boot window, so there is no upgrade or compatibility concern: a
Supervisor update landing between reservation and release would at worst
leave one inactive unit behind until the next host reboot.

Also reflow the _reserve_core_port docstring so the reference to the
hassio Docker bridge network reads as the network name it actually is.

* Assert literal unit names in port reservation tests

The tests imported _PORT_RESERVE_UNIT/_PORT_RESERVE_SERVICE and compared
them against themselves, so every assertion held for any value the
constants happened to have -- the rename in the previous commit passed
without a single test failing.

Use the literal unit names instead. That pins the two things systemd
actually cares about: the .socket/.service suffixes it derives the unit
type from, and the shared basename that tells it which service the
socket activates.
2026.09.0
2026-09-01 16:41:38 +02:00
Stefan AgnerandClaude Opus 5 30d6265768 Add Docker storage driver as a Sentry tag (#7187)
The storage driver is only available as event context today, which means
Sentry cannot aggregate over it. Tag events with it in addition, so issues
can be filtered and broken down by storage driver. This helps to spot
problems specific to a driver (e.g. containerd snapshotter vs overlay2).

Set the tag during setup as well, mirroring the context information which
is already reported in that state.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-01 16:28:04 +02:00
dependabot[bot] c156937281 Bump coverage from 7.15.4 to 7.16.0 (#7183)
Bumps [coverage](https://github.com/coveragepy/coveragepy) from 7.15.4 to 7.16.0.
- [Release notes](https://github.com/coveragepy/coveragepy/releases)
- [Changelog](https://github.com/coveragepy/coveragepy/blob/main/CHANGES.rst)
- [Commits](https://github.com/coveragepy/coveragepy/compare/7.15.4...7.16.0)

---
updated-dependencies:
- dependency-name: coverage
  dependency-version: 7.16.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-01 11:08:21 +02:00
dependabot[bot] 2de3f196ee Bump gitpython from 3.1.60 to 3.1.61 (#7184)
Bumps [gitpython](https://github.com/gitpython-developers/GitPython) from 3.1.60 to 3.1.61.
- [Release notes](https://github.com/gitpython-developers/GitPython/releases)
- [Changelog](https://github.com/gitpython-developers/GitPython/blob/main/CHANGES)
- [Commits](https://github.com/gitpython-developers/GitPython/compare/3.1.60...3.1.61)

---
updated-dependencies:
- dependency-name: gitpython
  dependency-version: 3.1.61
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-01 11:05:32 +02:00
Mike Degatano 9adf131901 Skip redundant inspect calls when constructing Docker container handles (#7175)
* Skip redundant inspect calls when constructing Docker container handles

Several places in the Docker layer called containers.get(id), which
performs an inspect API call, purely to construct a DockerContainer
handle immediately followed by another call (show(), exec(), attach())
that either performs its own inspect or doesn't need one at all. Switch
these spots to containers.container(id), which builds the same handle
locally with no I/O, removing the redundant round-trip.

Affected call sites: container_is_initialized(), stop_container() and
container_run_inside() in docker/manager.py, _get_container() and
_attach() in docker/interface.py, attach()/retag()/update_start_tag()
in docker/supervisor.py, and write_stdin() in docker/app.py.

Since containers.container()-built handles don't carry the real
container ID (unlike get(), which populates it from the inspect
response), _attach() and DockerSupervisor.attach() now read the ID from
the inspect payload (self._meta["Id"]) instead of the handle's .id
property.

This intentionally excludes container_stats(), which has its own
one-shot inspect+stats opportunity being handled separately in #7168.

Follow-up to review discussion in #7168:
https://github.com/home-assistant/supervisor/pull/7168#discussion_r3883548019

* Fix tests broken by containers.container() switch

Several tests simulated Docker errors/state by mocking
containers.get() (the old call site), but the production code now
builds container handles via the sync, no-I/O containers.container()
instead and only calls get() directly from the small set of call
sites intentionally left unconverted (DockerApp's legacy-name
migration in attach(), and container_stats() which is handled by
#7168). Retarget these tests to mock containers.container() (or
both, where a single test exercises both an app's attach() and a
plugin/core's attach()) so the simulated failures actually reach the
code path they're meant to test.

Also add the missing "Id" key to some mocked container inspect
payloads in test_check_docker_config.py: _attach() now reads the
container ID from the inspect response instead of the handle's .id,
so tests need that field present.

* Skip unnecessary inspect before stopping container

A reviewer pointed out that calling show() before stop() in
stop_container() is redundant I/O. Docker returns a 304 (not
an error) if the container is already stopped, 404 if it does
not exist, so we can call stop() directly and handle those
responses without an extra inspect call.

* Centralize container fetch error handling in DockerSupervisor

attach(), retag() and update_start_tag() each duplicated a
try/except block around containers.container(name).show().
Reuse the existing _get_container() helper from DockerInterface
instead, raising DockerNotFound explicitly when the container
is missing.
2026-09-01 00:52:00 +02:00
Stefan AgnerandClaude Fable 5 324d17e68b Use systemd automount units for network mounts (#7167)
* dbus: support aux units in start_transient_unit

Extend `Systemd.start_transient_unit` to accept the `aux` parameter
(`a(sa(sv))`), which has been hardcoded to `[]` since the wrapper was
introduced. Aux entries take the form `(unit_name, properties)` and
let systemd create multiple transient units atomically — most usefully
a `.mount` and its `.automount` companion in one D-Bus call.

Add the unit-property constants we'll need to drive that:

- `DBUS_ATTR_LAZY_UNMOUNT` ("LazyUnmount"): set on the `.mount` so
  systemd umounts with MNT_DETACH. Pairs with `softerr`/`soft` to make
  stop/restart sequences reliable even when the server is gone.
- `DBUS_ATTR_TIMEOUT_IDLE_USEC` ("TimeoutIdleUSec"): set on the
  `.automount` to control how long the mount stays around after the
  last access before autofs expires it.
- `DBUS_ATTR_WHERE` ("Where"): mount-point property; required on the
  `.automount` unit (and useful explicitly on `.mount` units too).

No call sites change in this commit — pure plumbing.

* mounts: pair each network .mount with a .automount companion

Network mounts now get a transient `.automount` unit created
atomically alongside their `.mount`, via the aux parameter on
`StartTransientUnit`. The systemd-managed autofs trigger handles
activation lazily: the path exists in the VFS even before the
underlying network mount runs, and the first access that crosses the
trigger fires `mount.cifs`/`mount.nfs` on demand.

Why this matters:

- PID 1's path-walks (chase, daemon-reload, generator scans) stop
  crossing dead NFS/CIFS lookups — autofs returns from kernel memory
  rather than entering the network filesystem. The class of failure
  fixed at the systemd-timeout layer in #6834 stops being reachable
  in the first place.
- Failed accesses fail fast, bounded by the `.mount`'s `TimeoutSec=`
  rather than hanging forever.
- Background reconnect (NFSv4 state-manager kthread, CIFS
  delayed_work) recovers transparently when the server returns; no
  remount needed.

The `.automount` carries `TimeoutIdleUSec=5min` so kernels can expire
the mount after inactivity and re-trigger on the next access. The
companion `.mount` carries `LazyUnmount=true` (MNT_DETACH on stop),
so umount returns immediately even when the server is unreachable;
existing fds drain in their soft/softerr timeout regime instead of
pinning the umount syscall.

`Mount.unmount()` now stops the `.automount` first so the autofs
trigger can't re-fire the underlying `.mount` during cleanup. The
automount stop is best-effort — if the unit is gone or refuses to
stop, we log and proceed to the `.mount` stop, which remains
authoritative.

Bind mounts opt out via a `creates_automount` class flag — they have
no server to wait on and the lazy semantics would only add
indirection. Network mounts (CIFS/NFS) opt in.

This is wiring only; the periodic mount reload, the reload/restart
escalation, and the inner bind-mount layer for media/share usage are
still in place. The next commit removes them now that automount
makes them unnecessary.

* mounts: drop bind layer, periodic reload, and reload/restart machinery

With the kernel's autofs trigger handling lazy activation and
transparent reconnect, almost all of the supervisor-side mount
choreography becomes unnecessary. Strip it out.

What goes away:

* The inner bind-mount layer. Media/share mounts used to live at
  `path_extern_mounts/{name}` and have a second `.mount` unit bind
  them into `path_extern_media/{name}` (resp. `share`). The
  `.automount` companion can sit directly at the container-facing
  path, and the parent-dir RSLAVE bind into add-on containers
  surfaces the autofs trigger the same way. `NetworkMount.where`
  now switches on usage:
    - MEDIA → `path_extern_media/{name}`
    - SHARE → `path_extern_share/{name}`
    - BACKUP → `path_extern_mounts/{name}` (unchanged)
* `BindMount`, `BoundMount`, `MountManager._bound_mounts`, the
  `_bind_mount`/`_bind_media`/`_bind_share` helpers, and the entire
  emergency-fallback dance (`path_emergency`). The empty-read-only
  dir trick was a workaround for the harder PID 1 wedge problem,
  which autofs solves at the root. Failed shares now surface as
  ETIMEDOUT/EHOSTDOWN on access plus a resolution issue.
* `MountManager.reload()` and the 15-minute `RUN_RELOAD_MOUNTS`
  periodic task. autofs re-activates on access; we don't need to
  poll, and not polling means we don't randomly trip stale-FH /
  softreval / dead-server edge cases on a timer.
* `Mount.reload()` and `Mount._restart()`. The reload→restart
  escalation existed to make `is_mounted` honest in the face of
  systemd's local-only state; now `is_mounted` IS honest (probe-
  based), and any recovery the kernel knows how to do happens
  inside autofs without our involvement.
* The `RELOADING` safety net introduced by #6834. That whole class
  of PID 1 wedge stops being reachable when path lookups don't
  cross dead network mounts.

What changes shape:

* `NetworkMount.is_mounted()` no longer gates on systemd's
  `ActiveState` before probing. The `.mount` unit is dormant
  whenever autofs hasn't recently triggered it, so systemd state
  is meaningless as a health signal. The probe (now via
  `statvfs("/path/.")` so the trailing dot forces `LOOKUP_DIRECTORY`
  and triggers autofs) is the source of truth, and after it returns
  we set `self._state` to ACTIVE/INACTIVE so the API reports
  reachability rather than autofs idle state.
* `MountManager.reload_mount()` becomes probe-only — no systemd
  reload/restart calls. The user-facing semantics: "tell me if this
  mount is reachable right now, and refresh the resolution issue
  accordingly." If the mount is dead, autofs will re-trigger it on
  the next consumer access; the supervisor doesn't need to force it.
* `Mount.load()` keeps `_update_state_await` for the no-job-dispatched
  case where a previous supervisor left a unit in `activating`, then
  probes once via `is_mounted` so state reflects reachability rather
  than the lazy-mount idle state.
* `BackupManager` no longer walks `bound_mounts`. Backups skip
  network-mount subdirectories during folder archive (so we don't
  recurse into the share), and unmount/re-mount them around folder
  restores (so writes target the local mount-point dir, not the
  remote share).

Behavior tradeoffs accepted:

* Dead media/share access now ~30s ETIMEDOUT instead of empty
  read-only dir. The empty-dir was a workaround for the PID 1
  wedge that autofs eliminates at the root; if specific add-ons
  rely on the old behavior we can revisit with a 2-stage automount
  whose fallback target is an emergency dir.
* Resolution-issue lag: failed mounts no longer surface within
  15 min on their own. They appear when the user probes via the
  API, when `BackupManager.reload()` walks the location, or when
  load-time activation fails.

* tests: adapt bind-layer-era tests after rebase

Main gained tests for the bind layer after this branch was written:
the #7013 rebind regression test and the #7072 bind-step rollback test
cover machinery that no longer exists, and the healthy-reload test
asserted the unconditional rebind. Drop the first two and reduce the
third to asserting that a healthy probe performs no systemd operations
at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: rearchitect automount setup from systemd/kernel review

A code-level review of systemd's automount implementation (verified
against v254.13, the version HAOS ships) and the kernel autofs/VFS
plumbing surfaced three correctness-critical flaws in the automount
design, plus several hardening gaps. See automount-rearchitecture.md
for the full analysis.

- Make the .automount the primary transient unit with the .mount as
  aux, mirroring `systemd-mount --automount=yes`. Aux units get no
  start job: with the .mount as primary the trigger was never armed
  and the design silently degraded to eager mounting.

- Set StartLimitIntervalUSec=0 on the .mount. With the default start
  rate limit (5 starts/10s, counting successful starts) a fast-failing
  mount plus any polling consumer trips the limit within seconds, and
  systemd then detaches the autofs trigger entirely
  (AUTOMOUNT_FAILURE_MOUNT_START_LIMIT_HIT) — the path silently
  becomes a plain writable local directory.

- Drop the 5-minute TimeoutIdleUSec (default 0 = never expire).
  Kernel idle expiry decides busyness via may_umount_tree(), which
  only counts the init-namespace mount instance — open files held by
  container processes are invisible, so expiry would unmount shares
  under actively writing add-ons.

- Probe with a plain statvfs: the statfs syscall walks with
  LOOKUP_AUTOMOUNT and triggers by itself; the trailing-dot trick is
  unnecessary. Classify ELOOP as a mount-propagation
  misconfiguration in the probe error handling.

- Re-arm the trigger from reload_mount(): if the .automount unit is
  failed or gone (e.g. autofs unmounted out-of-band), reset failure
  state and re-create the pair before probing.

- Tear down legacy eager-mount units during load(). The old design's
  bind unit for media/share occupies the exact unit name the network
  .mount uses now; on a warm upgrade adoption would mistake the old
  bind mount for the network mount and leak the legacy data mount.

- Honor the systemd job result when stopping the .mount during
  unmount and reset failure state on both units afterwards so dead
  transient units get garbage-collected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: adapt local data repair to the autofs design

Port of the mount-target-not-empty repair (#7089) on top of the
automount rearchitecture:

- relocate_local_data() handles a single target directory — the mount
  sits directly at its container-facing path, the bind layer and with
  it the multi-directory case are gone
- a repair_trigger() failure caused by blocking local data raises the
  mount failed issue with the move_local_data suggestion instead of
  the plain variant
- local data at load is detected via the mount unit itself; the
  bind-layer detection test is rewritten accordingly

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docker: use rslave propagation for execute_command share mount

The temporary container used for execute_command (e.g. core config
check) mounted /share without a propagation mode, unlike every other
share/media mount. With eager mounts this only meant missing mounts
made after container start; with lazy automount activation the shares
are routinely mounted after start, and accessing a not-yet-activated
automount from a private mount namespace fails with ELOOP.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: harden automount lifecycle from review findings

Address the findings of an adversarial review of the automount
rearchitecture:

- load() no longer adopts a dead automount trigger: a failed or
  stopped .automount leaves the path a plain writable directory — the
  silent local-write degradation this design must prevent. Tear down
  and re-arm instead. An adopted pair whose probe fails now raises so
  the manager surfaces the mount failed issue, same as a fresh mount.
- unmount() resolves the .mount unit only after stopping the
  .automount (the lazy detach can garbage-collect the transient unit,
  invalidating an earlier proxy), raises instead of warning when the
  automount stop errors, and checks the stop job result so a failed
  stop cannot leave an armed trigger behind a "successful" removal.
- repair_trigger() fully unmounts before re-arming: with the share
  still attached, arming fails — or worse, mount() misreads the
  mounted share's contents as blocking local data.
- Legacy unit teardown honors the stop job result for the same reason.
- reload_mount() escalates once to re-creating the unit pair when an
  established mount is unreachable: a permanently dead session (e.g.
  replaced server) keeps the path mounted, so the trigger can never
  re-fire and kernel reconnection never succeeds. Teardown is safe now
  (lazy unmount, no PID 1 path walks), unlike the removed reload →
  restart escalation of the eager design.
- Reinstate the periodic mounts task as a probe-based reconcile: re-arm
  dead triggers, refresh the reachability state reported by the API and
  used by backup locations, and sync the mount failed issue in both
  directions. No reload or restart of established mounts.
- Folder restore no longer fails after a successful restore when
  re-mounting nested mounts cannot verify an unreachable server — the
  trigger is armed, the share recovers on next access.
- Drop the now-unused Mount.update(), deduplicate unit name escaping,
  hoist the backup exclusion set out of the per-file filter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* tests: expect rslave propagation in core check container

The config check assertion missed the update for the rslave
propagation on the execute_command share mount.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* tests: cover reconcile trigger repair failure paths

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: arm automount before stopping the legacy data mount

Address review: the legacy teardown stopped the media/share bind unit
and the eager data mount in sequence, both strictly. When the bind
stop succeeded but the data mount stop failed — its unmount can time
out against an unreachable server, legacy units have no LazyUnmount —
load() raised with the container-facing path left a plain writable
directory and no trigger armed: the pollution mode this design is
meant to prevent. Reordering the stops would not help either, as
systemd stops the bind first anyway as a dependent of the data mount
(the implicit Requires= from RequiresMountsFor= on the bind's What=).

Instead, strictly stop only the path-conflicting unit (a failure
leaves the path covered, which is safe and retryable), arm the
automount right after, and stop the conflict-free legacy data mount
best-effort last — a failed stop logs a warning and leaves an orphaned
mount for the next Supervisor restart or a host reboot, with nothing
writable exposed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: discard dead session instead of re-creating units on reload

Address review: re-creating the unit pair on reload of an unreachable
established mount exposed the target as a plain writable directory
between the lazy unmount and arming the replacement — a continuously
writing add-on could block the re-arm or slip writes under the new
mount in the check-to-arm window.

Stop only the .mount unit instead, keeping the .automount armed:
systemd re-installs the autofs trigger over the path — the same
mechanism idle expiry uses, with the automount's Triggers= reference
keeping the transient .mount definition alive — so the path is never
locally writable. The re-probe then mounts fresh through the trigger,
which establishes a new session and thereby covers the permanently
dead session case (e.g. a replaced server) the escalation exists for.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: surface arming failures, keep restore teardown re-armed

Address review findings on the arming/re-arming error paths:

- mount() checks the StartTransientUnit job result: a failed start job
  means the trigger never armed and the path is a plain writable
  directory — a hard MountError before the probe, so it cannot be
  mistaken for the armed-but-unreachable MountActivationError.
- Folder restore includes the nested-mount teardown in the try block:
  an unmount failing halfway (trigger disarmed, share still attached)
  previously exited before any re-arm, leaving the path unprotected.
  The finally re-arms via repair_trigger(), which no-ops on a still
  armed trigger and handles partially torn-down pairs.
- The post-restore re-arm suppresses only MountActivationError (armed,
  recovers on next access). Any other failure — arming failed, local
  data blocking the target — left the path unprotected and surfaces.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: fold the two legacy unit teardowns into one helper

The eager-mount-era cleanup queried the .mount unit a second time
although load() had just fetched it, and the strict teardown of the
unit occupying the automount's path and the best-effort teardown of
the mount at the mounts data directory were near identical. Reuse the
unit from load() and give the shared helper a strict flag, which drops
a D-Bus round trip per mount on every load.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: re-arm the trigger when unmount cannot stop the mount

Stopping the automount detaches the whole stack at the path, so an
unmount that then fails to stop the .mount leaves a plain writable
directory behind. Arm a fresh pair before raising, best effort — the
local data repair covers what lands there if that fails too.

The automount stop itself needs no such handling: systemd's
automount_stop() enters dead synchronously and the detach it performs
(MNT_DETACH, no server contact) only logs its errors, so a failure
there means systemd could not be reached at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: raise translatable errors for mount operations

Errors from setting up, unmounting and reloading a mount reach users
through the API, so describe what failed in a translatable message and
leave the systemd job result or D-Bus error to the log.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: remove the emergency folder of the eager-mount design

The read-only fallback directory has no place in the automount design.
Remove what is left of it on existing installations, next to the legacy
addons directory cleanup, keeping it if it holds anything but the empty
mount points.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: trim comments to the behavior they describe

Drop the narration of what changed and why from comments and
docstrings, keeping the notes that record why a tempting alternative
does not work. Two comments still described the reload to restart
escalation this branch removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: correct stale claims in manager and reload test

The manager docstring credited the kernel with idle expiry, which is
deliberately disabled, and claimed no polling although the periodic
reconcile probes every 15 minutes. The reload test claimed systemd is
never contacted while the escalation stops the .mount unit; only reload
and restart of the unit are avoided.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: arm the trigger alone if re-creating the pair is rejected

Address review: the re-arm after a failed unmount submits the .mount as
an aux unit, but transient creation requires a pristine unit and the
definition is still loaded precisely when stopping it is what failed.
Fall back to creating the .automount on its own, which covers the path
and fires the surviving mount definition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: never arm the trigger over local data after a failed unmount

Address review: the fallback caught the target validation errors too,
so data written while the path was uncovered would end up beneath a
fresh trigger. Reconciliation would then find a healthy mount and the
repair skips mount points, leaving the data hidden for good. Leave the
path untouched instead, so the next reconcile offers to move it away.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-31 20:30:23 +02:00
Mike Degatano 2d153ac497 Add one-shot mode and docker-given totals to stats API (#7168)
* Add one-shot mode and docker-given totals to stats API

- Add one-shot query option to container stats to skip the calculated
  CPU percentage window, returning docker-given totals directly.
- Refactor container_stats() with translatable, structured Docker
  errors and timeout handling instead of generic Docker exceptions.
- Translate low-level Docker errors into component-flavored errors
  (NotRunningError/StatsTimeoutError/UnknownError) for Home Assistant,
  Supervisor, and each plugin (Audio, Cli, CoreDNS, Multicast,
  Observer) as well as apps, mirroring the existing App.stats()
  pattern for consistent, informative API responses.
- Add a shared api_return_stats() helper to remove duplicated stats
  handling logic across the 8 stats API endpoints.
- DockerContainerNotFoundError now inherits from DockerNotFound.

* Contain aiodocker _query access and split v1/v2 stats response model

- Contain the reach into aiodocker's protected _query internals (needed
  for one-shot stats until aio-libs/aiodocker#1054 lands) to a single
  DockerAPI._query_one_shot_stats() helper.
- Split the stats API response model between v1 (legacy, windowed CPU
  percentage still available via opt-in one_shot query string) and v2
  (always one-shot, no query string, cpu_percent dropped since it would
  never carry a meaningful value). Both share the same
  api_return_stats(stats, legacy=...) helper so the response model
  stays in one place.
- DockerAPI.container_stats() keeps its existing one_shot keyword only;
  the legacy/v1 vs v2 distinction is handled entirely in the API layer.

* Detect stopped/restarting containers in one-shot stats response

Docker returns a stub response containing only the container's id/name
(no cpu_stats/memory_stats/networks data) for a stopped or restarting
container instead of an error, even in one-shot mode - the not-running
check in daemon.ContainerStats happens before the one-shot branch, so
requesting one-shot stats doesn't change this. Detect the missing
online_cpus field (only ever present while running) and raise
DockerContainerNotRunningError instead of returning a payload that
DockerStats would otherwise turn into a successful all-zero response.

* Add tests for stats error paths introduced in one-shot/totals PR

- Cover DockerAPI.container_stats/_query_one_shot_stats edge cases: empty
  one-shot response, empty windowed response, and the raw _query_one_shot_stats
  helper.
- Add stats() error-translation tests (not-running, timeout, unknown) for the
  audio, cli, multicast and observer plugins, none of which had any coverage.
- Add stats endpoint tests (v1 + v2) for the cli, audio, dns, multicast and
  observer REST APIs, plus a v1 success-path test for the apps stats endpoint.

These close the codecov-flagged gaps for newly introduced code in this PR.

* Unify one-shot and windowed container_stats() code paths

Both modes now share the same error handling and not-running detection:
- The windowed path no longer inspects the container first to check its
  running state. Docker returns the same id/name-only stub response (no
  online_cpus in cpu_stats) for a stopped or restarting container in
  windowed mode as it does in one-shot mode, so the existing online_cpus
  presence check now covers both, closing a race window where a container
  stopping between the inspect and the stats call would previously reach
  the DockerStats parser with an incomplete payload.
- The windowed path builds its container handle via containers.container(),
  which constructs a DockerContainer from the name with no I/O, instead of
  containers.get(), which performs an unnecessary inspect call just to
  fetch the same information the stats call already returns.

* Fix protected access pylint issue
2026-08-31 17:11:38 +02:00
dependabot[bot] 5d6944a539 Bump ruff from 0.16.4 to 0.16.5 (#7179)
Signed-off-by: dependabot[bot] <support@github.com>
2026-08-31 09:12:39 +02:00
dependabot[bot] 3783aca1fe Bump gitpython from 3.1.59 to 3.1.60 (#7178)
Signed-off-by: dependabot[bot] <support@github.com>
2026-08-31 09:11:52 +02:00
dependabot[bot] bc9c93791e Bump cryptography from 50.0.0 to 50.0.1 (#7180)
Signed-off-by: dependabot[bot] <support@github.com>
2026-08-31 09:10:47 +02:00
Mike Degatano c5a5477c40 Reserve Core's host port before booting pre-Core apps (#7131)
* Reserve Core's host port before booting pre-Core apps

On a fresh system boot, Home Assistant Core and apps with startup: system
or startup: services run in sequence before Core. Core uses --network=host,
binding directly in the host network namespace. If an app claims Core's
HTTP port before Core starts, Core fails entirely with no recovery UI.

Supervisor itself runs on the hassio bridge and cannot bind in the host
network namespace via a plain socket call. Instead, before booting any
pre-Core app stages, start a minimal temporary Docker container
(hassio_port_reserve) from the Supervisor's own image with --network=host.
The container runs a Python one-liner that binds the port and blocks on
signal.pause(). It is killed and removed just before Core starts.

On Supervisor restart (Core already running), no container is created.
Stale containers from a crashed previous boot are detected and removed by
name before creating a new one.

Fixes https://github.com/home-assistant/supervisor/issues/6790

* Rework Core port reservation to use a systemd transient socket unit

Replace the Docker-container based port reservation with a transient
systemd .socket unit created over D-Bus. Supervisor runs on the hassio
bridge network and cannot bind Core's host port itself; asking systemd
to create and activate a transient socket unit lets systemd itself bind
the port in the host network namespace, with no process or container
involved.

- _reserve_core_port/_release_core_port are now systemd D-Bus calls
  instead of Docker container create/remove.
- A leftover unit from a previous crashed Supervisor boot (systemd
  refuses to redefine a transient unit that's still loaded, in any
  state, even with mode=replace) is cleaned up via a single
  _release_core_port() helper used both to release a successful
  reservation and to clean up a stale one before (re)creating it.
- reset_failed_unit() is only called when the unit doesn't cleanly
  reach INACTIVE, since systemd garbage collects a cleanly stopped
  transient unit on its own but keeps a FAILED one around.
- Rewrote tests/test_core.py to mock the systemd D-Bus interface
  instead of Docker, plus added coverage for the not-connected and
  unit-ends-up-failed edge cases.

* Address review feedback on port reservation

- Pair the transient socket unit with a trivial oneshot RemainAfterExit
  service created atomically via aux, since systemd needs a real
  activation target or it tears the socket down on the first incoming
  connection during boot.
- Reserve every configured http_server_host, not just the first one,
  defaulting to both 0.0.0.0 and :: when unset.
- Bracket IPv6 addresses in the socket's Listen directive.
- Catch HassioError instead of DBusError when releasing/reserving, since
  DBusNotConnectedError (raised on a mid-call bus disconnect) is not a
  DBusError subclass.
- _release_core_port now confirms both units actually reach
  INACTIVE/gone via direct state checks rather than trusting a
  wait_for_active_state timeout, retrying once and resetting FAILED
  units as needed. Releasing before Core starts remains best-effort:
  Core is always started regardless, since the reserved port only
  guards its own API bind and Core itself just retries that bind if it
  can't be confirmed released.

* Fixes from feedback and adjust to new default port
2026-08-28 13:38:46 +02:00
dependabot[bot] 442fe50e45 Bump time-machine from 3.4.0 to 3.5.0 (#7171)
Bumps [time-machine](https://github.com/adamchainz/time-machine) from 3.4.0 to 3.5.0.
- [Changelog](https://github.com/adamchainz/time-machine/blob/main/docs/changelog.rst)
- [Commits](https://github.com/adamchainz/time-machine/compare/3.4.0...3.5.0)

---
updated-dependencies:
- dependency-name: time-machine
  dependency-version: 3.5.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-28 10:32:20 +02:00
dependabot[bot] 5bba3465aa Bump sentry-sdk from 2.68.0 to 2.68.1 (#7170)
Bumps [sentry-sdk](https://github.com/getsentry/sentry-python) from 2.68.0 to 2.68.1.
- [Release notes](https://github.com/getsentry/sentry-python/releases)
- [Changelog](https://github.com/getsentry/sentry-python/blob/master/CHANGELOG.md)
- [Commits](https://github.com/getsentry/sentry-python/compare/2.68.0...2.68.1)

---
updated-dependencies:
- dependency-name: sentry-sdk
  dependency-version: 2.68.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-28 10:31:46 +02:00
Jan Čermák 68ed3b5dec Add endpoint to reset Docker storage (#7166)
* Add endpoint to reset Docker storage

Corrupted Docker image layers cannot be always fixed by removing and
re-pulling a single image, because layers are shared between images. The
only (easy) way to recover is wiping all of Docker's storage which is
currently cumbersome and requires OS shell access.

Home Assistant OS added a service that performs the Docker storage wipe
on boot when a flag file exists in home-assistant/operating-system#4982,
and home-assistant/os-agent#283 adds OS Agent DBus interface for
creating that file.

This adds a new API method POST /docker/reset-storage which calls OS
Agent's ScheduleDockerStorageReset and creates a reboot_required issue,
following the same pattern used for the Docker storage driver migration.

This is expected to ship in HAOS 18.3, so the endpoint checks for that
and returns 404 on older OS versions. Technically, it might be nicer to
gate on OS Agent version, but from user perspective it's better to
report the required OS version.

Refs #6555

* Add const for minimum OS version supporting the reset
2026-08-27 12:53:16 +02:00
Stefan Agner 210b49fb03 Add API to manage SSH authorized keys on Home Assistant OS (#7039)
* Add API to manage SSH authorized keys on Home Assistant OS

The OS Agent has long exposed AddSSHAuthKey and ClearSSHAuthKeys on its
io.hass.os System D-Bus object, but the Supervisor never wrapped them, so
there was no way to manage root's SSH authorized keys through the
Supervisor API. Add POST /os/ssh/authorized_keys, which replaces the
configured keys with the submitted list (an empty list just clears them).

Since OS Agent writes each key verbatim to /root/.ssh/authorized_keys as
root, the endpoint validates strictly before anything is written: plain
public keys only (no options, no certificates), a key type allowlist
matching what dropbear on Home Assistant OS can verify, no control
characters (one submitted key can never write more than one line), the
base64 blob must embed the declared key type, and entries are capped at
dropbear's 3000 byte per-line limit. The endpoint is admin-only for
add-on tokens.

Replacement clears the existing keys and then adds each key. OS Agent
releases up to 1.10.x return an error when clearing an already absent
file (inverted error check, since fixed), which is the state of every
first-time user, so this specific error is treated as the empty state it
reports.

dropbear on Home Assistant OS is gated by ConditionFileNotEmpty on the
authorized_keys file, which systemd only evaluates when the unit starts,
so the service is started after a non-empty key set is written. A running
dropbear re-reads the file on every authentication attempt and needs no
restart.

* Delegate SSH key validation to OS Agent

Review discussion questioned the full key validation (type allowlist,
base64 blob checks, canonicalization): OS Agent 1.10.0 validates
submitted keys itself and treats clearing an already absent
authorized_keys file as success, so the Supervisor can rely on it
instead of duplicating the logic.

Require OS Agent 1.10.0 or newer and reject requests on older releases
with 404, like the Raspberry Pi firmware endpoints do; this also makes
the missing-file compatibility shim for the clear call unnecessary. A
key rejected by OS Agent surfaces as an error response including its
validation message. The Supervisor keeps only a basic per-key sanity
check that runs before anything is written: no control characters (one
submitted key can never write more than one authorized_keys line) and
at most 3000 bytes (dropbear ignores longer lines, which would leave a
key that passes but never works).

* Support OS Agent releases before 1.10.0

Requiring OS Agent 1.10.0 would keep the feature unavailable until the
next OS update reaches users, while Supervisor updates roll out
independently. Drop the version requirement and accept that validation
is lossy on older releases: the Supervisor sanity check still prevents
writing more than one line per key or lines dropbear ignores, but
proper key validation only happens on OS Agent 1.10.0 or newer.

This brings back the need to tolerate the error OS Agent releases
before 1.10.0 return when clearing an already absent authorized_keys
file (inverted error check): match the complete os.Remove error message
for the authorized_keys path and treat it as the empty state clearing
aims for, on affected versions only.

* Split SSH authorized keys API into add and clear endpoints

Review feedback preferred endpoints mapping 1:1 onto the OS Agent D-Bus
methods over a single replace-the-set API, whose POST semantics were
also questioned (an idempotent full replacement would be PUT).

POST /os/ssh/authorized_keys now takes a single key ({"key": "..."})
and appends it via AddSSHAuthKey; DELETE /os/ssh/authorized_keys
removes all keys via ClearSSHAuthKeys. Each call maps to exactly one
OS Agent operation, so no request can partially succeed. Clients that
want to replace the configured set clear and re-add; a GET (which
needs an OS Agent extension first) and an idempotent PUT can be added
later.

The per-key sanity check, the dropbear service start after adding a
key, and the tolerance for the missing-file clear error of OS Agent
releases before 1.10.0 carry over unchanged.

* Add endpoint to list SSH authorized keys

With only add and clear operations the authorized_keys file is
write-only for API consumers: a user cannot audit which keys grant
access to the box, or whether any exist at all — including keys that
were imported from USB or written by add-ons.

OS Agent 1.11.0 added a ListSSHAuthKeys D-Bus method. Expose it as
GET /os/ssh/authorized_keys, returning the configured entries verbatim.
The endpoint requires OS Agent 1.11.0 and returns 404 on older
releases, like the Raspberry Pi firmware endpoints do; add and clear
keep working on all OS Agent releases.

* Restrict SSH authorized keys endpoints to Home Assistant Core

Review decision: manipulating root's SSH access is not a capability
add-ons should have, even with the admin role, so move the endpoints
from admin-only to the core_only middleware pattern. Only requests
authenticated with the Home Assistant Core token pass; add-on tokens of
any role, the CLI plugin, and the observer are rejected. This also
means the host shell (ha CLI) cannot use the endpoints for now — the
restriction can be opened up later.

The exclusion from the manager role allowlist is kept: if the path is
ever removed from core_only again, it falls back to admin-only rather
than becoming manager-accessible.

* Stop dropbear after clearing SSH authorized keys

Clearing all authorized keys is a revocation, but without stopping
dropbear the listener keeps running until reboot and established
sessions survive, as only a service stop terminates them. Stop the
service after a successful clear, mirroring the USB config import
(haos-config), which also stops dropbear when the imported
authorized_keys file is removed. Stopping an inactive unit is a no-op.

* Serialize SSH authorized keys jobs on a common lock

The add and clear jobs each perform a file operation followed by a
service operation, and nothing prevented them from running
concurrently. An interleaving like clear-file, add-key, start-dropbear,
stop-dropbear lets both requests succeed while the final service state
does not match the final key state (key configured, dropbear stopped).

Make OSManager a JobGroup and run both jobs with GROUP_QUEUE
concurrency, so each file-and-service operation completes before the
next starts. The regression test fails without the shared lock.
2026-08-27 10:12:44 +02:00
Stefan AgnerandClaude Fable 5 69c5a53864 Surface fixup failures to the caller applying a suggestion (#7145)
* Surface fixup failures to the caller applying a suggestion

Applying a suggestion whose fixup failed reported success: the fixup
base swallowed ResolutionFixupError, the API returned OK, and the
repair flow in Home Assistant completed and removed the repair from
the UI while the issue persisted in Supervisor.

Let the error propagate from the fixup instead. The autofix loop
already handles and logs per-fixup errors, bus-event triggered fixups
now get the same treatment in the event callback, and a user-applied
suggestion surfaces the failure as an API error so clients can show
that the fix did not apply. ResolutionFixupError becomes an APIError
with a translatable error key so it is reported as a client-visible
error instead of an unexpected one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Drop ResolutionFixupError, let fixup failures bubble

Per review: the generic wrapper existed for a time when suggestions
were only applied by autofix and the one job was separating fixup
failures from real bugs for Sentry. The suggestion API is in regular
use now and the wrapper actively hurts it — well-defined errors from
the underlying operations were caught and replaced with a generic
message.

Remove the exception type entirely and let the original errors reach
the caller. The autofix loop and the bus-event fixup path treat any
HassioError as an environmental/config failure: log and continue
without Sentry capture (the raise site reports to Sentry where
warranted); everything else is still captured. Direct raises in the
data disk fixups become HassOSDataDiskError, the could-not-start check
in the app start fixup becomes AppsError. ResolutionFixupJobError now
derives from ResolutionError and JobException.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Share unattended fixup error handling between autofix and bus events

Address review: the bus event listener only caught HassioError, so a
Supervisor bug in a bus-triggered fixup ended as an unretrieved task
exception without Sentry report, unlike the same bug during autofix.
Both unattended paths now use a common apply_fixup_safely() helper:
log HassioError (environmental, reported at the raise site if
warranted), capture everything else.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 23:14:44 +02:00
Copilotandagners f753e2c83c Fix security rating for apps without a web interface (#7165)
* Initial plan

* Fix security rating for apps without web UI

Co-authored-by: agners <34061+agners@users.noreply.github.com>

* Rate apps by network exposure

Co-authored-by: agners <34061+agners@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: agners <34061+agners@users.noreply.github.com>
2026-08-25 21:04:16 +02:00
Thibault MollemanandStefan Agner 6947d8b2b2 Strip whitespace from repository URLs in store validation (#7161)
* Strip whitespace from repository URLs in store validation

* Reject repository URLs containing internal whitespace

Tighten RE_REPOSITORY so the URL group rejects whitespace that
.strip() cannot fix, failing validation with a clear error instead of
an unclear git clone failure later. Also parametrize the existing add
repository API test instead of duplicating it and collapse the strip
test cases into one mixed-whitespace case.

---------

Co-authored-by: Stefan Agner <stefan@agner.ch>
2026-08-24 15:48:48 +02:00
dependabot[bot] fbf018f373 Bump ruff from 0.16.3 to 0.16.4 (#7163)
Bumps [ruff](https://github.com/astral-sh/ruff) from 0.16.3 to 0.16.4.
- [Release notes](https://github.com/astral-sh/ruff/releases)
- [Changelog](https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md)
- [Commits](https://github.com/astral-sh/ruff/compare/0.16.3...0.16.4)

---
updated-dependencies:
- dependency-name: ruff
  dependency-version: 0.16.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-24 09:02:02 +02:00
Mike Degatano 090ef31ec0 Refactor v2 API contracts to remove deprecated supervisor/apps/backups fields (#7111)
* tests(api): cover v1/v2 deprecated contract split

* api: drop deprecated fields from v2 supervisor/apps/backups contracts

* Proper fix for review finding

* Fixes from feedback
2026-08-21 11:55:28 +02:00
dependabot[bot] 6397faf22b Bump blockbuster from 1.5.26 to 1.5.27 (#7158)
Bumps [blockbuster](https://github.com/cbornet/blockbuster) from 1.5.26 to 1.5.27.
- [Release notes](https://github.com/cbornet/blockbuster/releases)
- [Commits](https://github.com/cbornet/blockbuster/compare/v1.5.26...v1.5.27)

---
updated-dependencies:
- dependency-name: blockbuster
  dependency-version: 1.5.27
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-21 11:49:55 +02:00
Petar Petrov f148020f3a Harden mount storage usage edge cases (#7154)
Follow-up to the per-mount storage usage endpoint, addressing three review
findings on the merged change.

Depths below 2 all yield the same totals-only response for a mount, so they
are normalized to one value before the probe registry is keyed. Concurrent
callers varying only the depth now share a single probe instead of each
parking an executor thread on the same mount.

A directory walk that races a deletion can count data the filesystem figure
no longer includes, producing children that sum past their parent. Such a
breakdown now gets dropped in favor of the totals, which are the consistent
part; the next request re-walks a settled tree.

The v1 legacy addon-id remap is now applied only to the system disk target.
A mount's children are real directory names, so one that happens to be
called apps_data must not be renamed - and a mount response dict is shared
with every concurrent waiter on the same probe, so remapping it in place
could leak renamed ids into v2 responses.
2026-08-20 15:13:35 -04:00
dependabot[bot] 99fb8e1295 Bump deepmerge from 2.1.0 to 3.0 (#7155)
Bumps [deepmerge](https://github.com/toumorokoshi/deepmerge) from 2.1.0 to 3.0.
- [Release notes](https://github.com/toumorokoshi/deepmerge/releases)
- [Commits](https://github.com/toumorokoshi/deepmerge/compare/v2.1.0...v3.0)

---
updated-dependencies:
- dependency-name: deepmerge
  dependency-version: '3.0'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-20 09:46:33 +02:00
Mike Degatano 0b5b0ce15f Fix built-in repository slugs and add origin URL migration support (#7114)
* Refactor built-in repository slugs

* Handle built-in repository origin URL migration

* Catch configparser errors when reading remote config

GitPython's Repo.remotes and Remote.url properties parse .git/config
directly (via GitConfigParser) rather than shelling out to git, so a
corrupted config file raises configparser.Error subclasses (e.g.
MissingSectionHeaderError), not git.CommandError or the other
GitPython exception types that were previously being caught. Also
widen the guarded section to cover the origin.url read, which is
reached via the same config-parsing code path and wasn't wrapped
before.
2026.08.0
2026-08-19 23:41:21 +02:00
Stefan AgnerandClaude Fable 5 ccdd8c1ef4 Add repair to move local data blocking a mount target (#7089)
* Add repair to move local data blocking a mount target

When an add-on writes into a media/share directory while its network
mount is not in place (#7037), the local data blocks re-creating the
mount: mounting over a non-empty directory is refused. Until now this
failed silently at Supervisor startup — the bind mounts were created as
fire-and-forget tasks — and the only way out was to remove the data
manually over SSH/Samba and re-create the mount via the API.

Surface the condition as a new mount_target_not_empty issue and offer a
move_local_data suggestion. The fixup moves the blocking data to a
<name>_local_recovery folder in a user-accessible location — media or
share for bind mount targets, local backup storage for backup mounts
(their data mount directory is not reachable for users) — then reloads
the mount. Nothing is deleted; users can inspect and clean up the
recovered data via the media browser or the share and backup folders.

Bind mount failures during load are now awaited and routed into
resolution issues instead of being swallowed as fire-and-forget tasks;
bind failures other than blocking local data create the existing
mount_failed issue. A successful mount reload dismisses a stale local
data issue.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Attach move local data suggestion to the mount failed issue

Review feedback on the repair: rather than introducing a separate
mount_target_not_empty issue type, keep a single mount failed issue
per mount and offer moving the blocking data as an additional
suggestion alongside reload and remove. Reload stays available for
users who prefer to clear the data themselves, and at most one repair
exists per mount. Adding is idempotent, so an already-raised mount
failed issue just gains the extra suggestion.

When re-creating the bind mount after a successful reload fails on
blocking local data, the mount failed issue is re-added together with
the move suggestion, since the reload already dismissed it.

This also resolves the reviewer note about not-a-directory conflicts
being reported under a not-empty issue type: the issue type no longer
encodes the filesystem detail, while the error messages keep the
distinction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Keep an empty directory in place after relocating local data

If remounting fails after the local data was moved aside (e.g. the
server is unreachable at that moment), the renamed directory left
nothing behind: media/share consumers saw the folder disappear
entirely. Recreate an empty directory right after the rename so the
path stays present regardless of whether the remount succeeds.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Drop move local data suggestion once the data was moved

When the remount after relocating local data fails (e.g. the server is
unreachable at that moment), the mount failed issue stays — but the
move suggestion stayed with it, offering to move data that is no
longer in the way. Dismiss the suggestion after the relocation step so
only reload and remove remain for the leftover failure. Detection
re-adds the move suggestion if local data blocks the target again.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Filter Core-facing suggestions by minimum Core version

The fix flow translations for a new suggestion ship with a Core
release. Older Core frontends render an unknown suggestion as an
empty, unlabeled menu entry in the repair fix flow. Filter such
suggestions from Core-facing output — the issue events sent over the
websocket and the resolution API responses when the caller is Home
Assistant — until the connected Core is new enough. Other API
consumers like the CLI always see the full suggestion list.

The move_local_data suggestion requires Core 2026.9.0b0, the release
its fix flow translations are targeted at.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Make fixup failure tests independent of error propagation behavior

Suppress a potential ResolutionFixupError from the failing fixup calls
so the tests pass both while fixup errors are swallowed and once they
propagate to the caller (#7150).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address review: one recovery folder, check OSError for known issues

Move all local data blocking a mount into a single recovery folder so
the user finds it as one fix: when more than one directory holds data,
later ones become subfolders named after their parent directory
instead of numbered sibling folders.

Also run OSError from the relocation through check_oserror to pick up
known filesystem issues like corruption (bad message).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Only filter Core-facing suggestions on v1 surfaces

Per review: no Core version predating the repair suggestion filtering
in its own fix flow (home-assistant/core#179540) supports the v2 API,
so the Supervisor-side compatibility filter is only needed where old
Core versions actually look. Filter the v1 resolution endpoints and
the legacy websocket issue payloads; the v2 endpoints and v2 event
payloads always carry the full suggestion list. The suggestions for
issue endpoint gets a v1 handler for this.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 23:40:47 +02:00
Petar Petrov 143da593b8 Report storage usage per mount (#7146)
Parameterize the disk usage endpoint as /host/disks/{disk}/usage so a
supervisor mount can be asked for its own storage figures. "default" still
means the system disk and keeps its existing response byte for byte; any
other value names a mount, which must exist and be active. The reserved
target wins over a mount of the same name, which the mount name pattern
permits.

Usage comes from the mount's own filesystem, with used derived as total
minus free so that reserved space counts as used, as it does for the system
disk.

A mount reports totals only by default. The directory walker recurses
regardless of max_depth and only gates whether children appear in its
output, so walking a whole network mount to report a single number would
cost many round trips for nothing; at depth 0 it is skipped outright rather
than called. Depth otherwise means what it means for the system disk, where
level 1 is the labeled known paths that a mount does not have, so a mount's
own subdirectories start at level 2. When a breakdown is produced, whatever
the walk cannot attribute is reported as an "other" child, which keeps every
node's children summing to its used_bytes.

Probes are deliberately not cut short. A caller showing a loader is better
served by a real answer than a fast failure, so the timeout is only a
backstop against a probe that never returns, and a slow one is confined to
its executor thread rather than blocking the rest of the API. Concurrent
callers asking for the same mount at the same depth share one probe instead
of each parking a thread on identical work.
2026-08-19 11:09:06 +03:00
Stefan AgnerandClaude Opus 5 a14729745c Parse image references using Docker's domain splitting logic (#7139)
* Parse image references using Docker's domain splitting logic

The unsupported container evaluation split an image reference on its first
colon to strip the tag. For a registry with a port, that colon belongs to the
port, so `myregistry:5000/app:1.0` was reduced to `myregistry` and reported as
an unsupported image on the host.

Splitting an image reference correctly needs to know where the registry domain
ends, and the pieces for that were already there but wired up in a way that
could not be reused. IMAGE_REGISTRY_REGEX is a port of Docker's DomainRegexp,
which Docker itself uses to validate a domain, not to find one. Placing it in
front of the search meant get_registry_from_image() still had to re-derive the
answer with the dot/colon/localhost checks that follow the match, and callers
that needed the rest of the reference recovered it by slicing off
len(registry) + 1 characters.

Split the two jobs apart, mirroring Docker's reference implementation:

- split_docker_domain() finds the domain the way splitDockerDomain() does, by
  cutting at the first slash and testing the candidate. It returns the
  remainder as well, so callers no longer slice by length, and it canonicalizes
  index.docker.io to docker.io, which lets stored Docker Hub credentials apply
  to references using the legacy domain.
- is_registry_domain() validates a domain against DomainRegexp, which is what
  the regex is for. Image validation keeps rejecting malformed domains such as
  ".ghcr.io" through this check.
- get_registry_from_image() stays as a wrapper for the callers that only need
  the domain.

With the domain handled, splitting off the tag is a matter of taking the last
colon that has no slash after it, which is what Docker's TagRegexp allows.
split_image_tag() does that and drops any digest, so a digest-pinned image no
longer reads as unsupported either.

Move the image reference tests to tests/docker/test_utils.py next to the code
under test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Convert remaining image reference parsing to the new helpers

Canonicalizing index.docker.io to docker.io made _get_credentials() qualify
Docker Hub images from the raw reference, so a reference already carrying a
Docker Hub domain gained a second prefix: index.docker.io/org/app was pulled as
docker.io/index.docker.io/org/app. Before the canonicalization the legacy domain
matched no configured registry and the image was pulled anonymously under its
original name, so this only surfaced now, although the same doubling already
applied to an explicit docker.io/org/app reference. Qualify from the remainder
instead, which covers references without a domain and with either Docker Hub
domain.

Three more sites split an image reference on its first colon and hit the same
problem a registry with a port causes in the unsupported container evaluation:

- The image property of DockerInterface reported myreg:5000/supervisor:1.0 as
  myreg, and a digest reference as name@sha256.
- get_latest_version() read the tag of a RepoTags entry, which for
  myreg:5000/homeassistant:2026.8.0 yielded 5000/homeassistant:2026.8.0. That
  is not a known version strategy, so every tag was skipped and the lookup
  failed with "No version found". This is reachable with a user-overridden Core
  image or a plugin image on a registry with a port.
- The Supervisor start tag repair took the image name from a RepoTags entry the
  same way, leaving myreg as the name to tag.

Also align two Docker Hub details with normalize.go: the library/ prefix for
official images now applies whenever the resolved registry is Docker Hub, not
only when the reference carried no domain, and credential lookup falls back to
the legacy hub.docker.com key for an explicit Docker Hub domain as it already
did for references without one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Cover qualifying a Docker Hub image with either registry key

The credential test only covered an image without a domain and an image with a
Docker Hub domain against credentials stored under the official docker.io key.
Parametrize over both Docker Hub domains and both registry keys, so the pull
name is asserted for a reference carrying the legacy index.docker.io domain
while credentials are stored under hub.docker.com as well.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 09:50:00 +02:00
Stefan AgnerandClaude Fable 5 f98117c35e Use store version's architecture when updating an app (#7144)
App.update() passed self.arch to DockerApp.update(), which matches the
system's supported architectures against the arch list of the *installed*
version. When the installed version no longer supports any of the system's
architectures (e.g. a 32-bit only app after 32-bit support was dropped)
while the new store version does, this raised an unhandled
HassioArchNotFound, surfacing as "Unexpected error during API call" (500).
The availability of the store version was already validated by the caller,
so the update itself is legitimate.

Use the cached store data's architecture list instead, which is also the
semantically correct source since it describes the image being pulled.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 08:58:09 +02:00
Mike Degatano 9ad1c198c2 Block supervisor/hassio WebSocket command types in the HA proxy (#7123)
* Block supervisor/hassio WebSocket command types in the HA proxy

An app with only homeassistant_api: true could send supervisor/api (or
other supervisor/* / hassio/* command type) frames through the Supervisor's
Core WebSocket proxy at /core/websocket. Core executes those via the hassio
integration's websocket_supervisor_api handler, which calls back into the
Supervisor using Core's own token. That token bypasses all role checks in
security.py, giving the app unrestricted Supervisor API access regardless of
its declared hassio_role or hassio_api flag.

Fix: filter command types on the app-to-Core direction of _proxy_message.
Any TEXT frame whose type field starts with 'supervisor/' or 'hassio/' is
rejected with a Core-shaped result/unauthorized response instead of being
forwarded. The connection stays open. Fail-open on unparseable frames (Core's
own validation already rejects those). The Core-to-app direction is
unfiltered.

* Remove the double json parse
2026-08-19 08:57:10 +02:00
Jan Čermák 384ed39b39 Use port 80 as default for Core/landingpage, omit scheme's default port in API URL (#7153)
Set default port of HA to 80, as this default is authoritative for
what's shown on the HA CLI banner when landing page is running - there's
no homeassistant.json at this time, so this value applies. We shouldn't
need to care about installs running an old landing page (without port 80
support at all), as Supervisor update results in the old default being
persisted homeassistant.json which is written on Supervisor restart, so
the banner still shows port 8123 if the old Supervisor (and hence old
landing page) were baked in an OS image.

Also append the port to the API URL only when it's not the scheme's
default to avoid :80/:443 suffixes when they're not necessary.

Closes #7151
2026-08-18 18:30:39 +02:00
dependabot[bot] 56c515b23b Bump types-pyyaml from 6.0.12.20260724 to 6.0.12.20260815 (#7149)
Bumps [types-pyyaml](https://github.com/python/typeshed) from 6.0.12.20260724 to 6.0.12.20260815.
- [Commits](https://github.com/python/typeshed/commits)

---
updated-dependencies:
- dependency-name: types-pyyaml
  dependency-version: 6.0.12.20260815
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 10:46:00 +02:00
dependabot[bot] a1865edec7 Bump orjson from 3.11.9 to 3.12.0 (#7148)
Bumps [orjson](https://github.com/ijl/orjson) from 3.11.9 to 3.12.0.
- [Release notes](https://github.com/ijl/orjson/releases)
- [Changelog](https://github.com/ijl/orjson/blob/master/CHANGELOG.md)
- [Commits](https://github.com/ijl/orjson/compare/3.11.9...3.12.0)

---
updated-dependencies:
- dependency-name: orjson
  dependency-version: 3.12.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 09:57:22 +02:00