Commit Graph
312 Commits
Author SHA1 Message Date
Stefan AgnerandClaude Fable 5.1 108ac2f15d Add enabled option to the multicast plugin
Since 2020 Supervisor unconditionally installs and runs the multicast
plugin on every installation. Users affected by its interference (mDNS
duplicates, multicast storms) have no supported way to turn it off. This
is phase 1 of the deprecation plan in architecture discussion #1457: an
opt-in disable switch.

The plugin config gains an `enabled` option (default true), persisted in
multicast.json and exposed via `GET /multicast/info` and a new
`POST /multicast/options` endpoint. Disabling stops and removes the
container, removes the image and forgets the installed version. Enabling
installs the current version and starts it, gated by the same job
conditions as a plugin update. While disabled, load skips attach, install
and start, the watchdog ignores container events, repair and auto-update
are skipped, and start/restart/update raise MulticastDisabledError
(HTTP 400).

The update endpoint now rejects a disabled plugin before comparing
versions, as the comparison raises on the None version left behind by
a disable.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 13:58:03 +02:00
Mike Degatano 54a430f55e Hide new ntp host feature from v1 /host/info features field (#7207)
The python-supervisor-client library's HostInfo model types
`features` as the strict `list[HostFeature]` instead of
`list[HostFeature | str]`:
https://github.com/home-assistant-libs/python-supervisor-client/blob/3459d25b7b223dfd2a49e5aa3f794ea822faab5c/aiohasupervisor/models/host.py#L49

#7181 added HostFeature.NTP, which now breaks the hassio integration
on nightly since the client library raises on parsing it:

    Failed setup, will retry: Field "features" of type list[HostFeature] in HostInfo has invalid value ['reboot', 'shutdown', 'services', 'network', 'hostname', 'timedate', 'os_agent', 'haos', 'resolved', 'journal', 'disk', 'mount', 'ntp']

Work around this on the v1 /host/info endpoint by hiding "ntp" from
the features field and exposing the full list via a new all_features
field instead. The v2 API is unaffected since it has not shipped yet
and its features field is already typed as list[HostFeature | str].
2026-09-10 11:20:35 +02:00
Jan Čermák 8684f7e4cd Add /time API for setting NTP servers on HAOS (#7181)
* Add `/ntp` API for setting NTP servers on HAOS

Add API for customizing NTP servers, using fixed DBus methods available
since OS 18.3. The API is limited to HAOS only with OS Agent that sets
the values in a daemon drop-in config. Changing the values triggers a
systemd-timesyncd unit restart.

The API currently exposes only the configuration overrides, not the
currently applied value (like `timedatectl show-timesync
--property=ServerName` does). If we want to add the effective NTP
server, another field should be introduced in the API for this.

Closes #6278

* Rename `/ntp` to `/time`

* Nest options under `config` key

Move the config options under a separate key to distinguish between
runtime data that we plan to add.

* Tighten NTP server regexp

* Require systemd D-Bus for the NTP host feature

* Always restart timesyncd when NTP servers are set
2026-09-02 17:22:20 +02:00
Petar Petrov fbab5e10a5 Report mount storage usage from the probe, not cached state (#7188)
* Report mount storage usage from the probe, not cached state

The per-mount usage endpoint refused any mount whose cached state was not
active. Since mounts became autofs-triggered, that state is only as fresh
as the last reconcile probe, fifteen minutes apart, so a share that came
back stayed unreportable for up to that long even though a probe would
have found it healthy and mounted it.

Drop the gate and let the probe decide. Its statvfs walks with
LOOKUP_AUTOMOUNT, so it activates a dormant trigger and then reports what
it found, which also means a usage request is no longer strictly read-only
with respect to system state - it can cause the mount it measures.

A path that still does not cross a filesystem boundary after that statvfs
is one whose trigger is gone and which reverted to a plain directory; the
existing boundary check catches it and its comment now says so. With the
gate removed that is the only remaining not-mounted condition, so the
mount_usage_not_active_error key, which named the cached state the
endpoint no longer consults, gives way to mount_usage_not_mounted_error,
which names what the probe observed.

* Tighten comments for probe-authoritative mount usage.

Cached state can be stale; the live probe is the authority, including dormant automount activation.
2026-09-01 16:29:49 -04:00
Stefan AgnerandClaude Fable 5 324d17e68b Use systemd automount units for network mounts (#7167)
* dbus: support aux units in start_transient_unit

Extend `Systemd.start_transient_unit` to accept the `aux` parameter
(`a(sa(sv))`), which has been hardcoded to `[]` since the wrapper was
introduced. Aux entries take the form `(unit_name, properties)` and
let systemd create multiple transient units atomically — most usefully
a `.mount` and its `.automount` companion in one D-Bus call.

Add the unit-property constants we'll need to drive that:

- `DBUS_ATTR_LAZY_UNMOUNT` ("LazyUnmount"): set on the `.mount` so
  systemd umounts with MNT_DETACH. Pairs with `softerr`/`soft` to make
  stop/restart sequences reliable even when the server is gone.
- `DBUS_ATTR_TIMEOUT_IDLE_USEC` ("TimeoutIdleUSec"): set on the
  `.automount` to control how long the mount stays around after the
  last access before autofs expires it.
- `DBUS_ATTR_WHERE` ("Where"): mount-point property; required on the
  `.automount` unit (and useful explicitly on `.mount` units too).

No call sites change in this commit — pure plumbing.

* mounts: pair each network .mount with a .automount companion

Network mounts now get a transient `.automount` unit created
atomically alongside their `.mount`, via the aux parameter on
`StartTransientUnit`. The systemd-managed autofs trigger handles
activation lazily: the path exists in the VFS even before the
underlying network mount runs, and the first access that crosses the
trigger fires `mount.cifs`/`mount.nfs` on demand.

Why this matters:

- PID 1's path-walks (chase, daemon-reload, generator scans) stop
  crossing dead NFS/CIFS lookups — autofs returns from kernel memory
  rather than entering the network filesystem. The class of failure
  fixed at the systemd-timeout layer in #6834 stops being reachable
  in the first place.
- Failed accesses fail fast, bounded by the `.mount`'s `TimeoutSec=`
  rather than hanging forever.
- Background reconnect (NFSv4 state-manager kthread, CIFS
  delayed_work) recovers transparently when the server returns; no
  remount needed.

The `.automount` carries `TimeoutIdleUSec=5min` so kernels can expire
the mount after inactivity and re-trigger on the next access. The
companion `.mount` carries `LazyUnmount=true` (MNT_DETACH on stop),
so umount returns immediately even when the server is unreachable;
existing fds drain in their soft/softerr timeout regime instead of
pinning the umount syscall.

`Mount.unmount()` now stops the `.automount` first so the autofs
trigger can't re-fire the underlying `.mount` during cleanup. The
automount stop is best-effort — if the unit is gone or refuses to
stop, we log and proceed to the `.mount` stop, which remains
authoritative.

Bind mounts opt out via a `creates_automount` class flag — they have
no server to wait on and the lazy semantics would only add
indirection. Network mounts (CIFS/NFS) opt in.

This is wiring only; the periodic mount reload, the reload/restart
escalation, and the inner bind-mount layer for media/share usage are
still in place. The next commit removes them now that automount
makes them unnecessary.

* mounts: drop bind layer, periodic reload, and reload/restart machinery

With the kernel's autofs trigger handling lazy activation and
transparent reconnect, almost all of the supervisor-side mount
choreography becomes unnecessary. Strip it out.

What goes away:

* The inner bind-mount layer. Media/share mounts used to live at
  `path_extern_mounts/{name}` and have a second `.mount` unit bind
  them into `path_extern_media/{name}` (resp. `share`). The
  `.automount` companion can sit directly at the container-facing
  path, and the parent-dir RSLAVE bind into add-on containers
  surfaces the autofs trigger the same way. `NetworkMount.where`
  now switches on usage:
    - MEDIA → `path_extern_media/{name}`
    - SHARE → `path_extern_share/{name}`
    - BACKUP → `path_extern_mounts/{name}` (unchanged)
* `BindMount`, `BoundMount`, `MountManager._bound_mounts`, the
  `_bind_mount`/`_bind_media`/`_bind_share` helpers, and the entire
  emergency-fallback dance (`path_emergency`). The empty-read-only
  dir trick was a workaround for the harder PID 1 wedge problem,
  which autofs solves at the root. Failed shares now surface as
  ETIMEDOUT/EHOSTDOWN on access plus a resolution issue.
* `MountManager.reload()` and the 15-minute `RUN_RELOAD_MOUNTS`
  periodic task. autofs re-activates on access; we don't need to
  poll, and not polling means we don't randomly trip stale-FH /
  softreval / dead-server edge cases on a timer.
* `Mount.reload()` and `Mount._restart()`. The reload→restart
  escalation existed to make `is_mounted` honest in the face of
  systemd's local-only state; now `is_mounted` IS honest (probe-
  based), and any recovery the kernel knows how to do happens
  inside autofs without our involvement.
* The `RELOADING` safety net introduced by #6834. That whole class
  of PID 1 wedge stops being reachable when path lookups don't
  cross dead network mounts.

What changes shape:

* `NetworkMount.is_mounted()` no longer gates on systemd's
  `ActiveState` before probing. The `.mount` unit is dormant
  whenever autofs hasn't recently triggered it, so systemd state
  is meaningless as a health signal. The probe (now via
  `statvfs("/path/.")` so the trailing dot forces `LOOKUP_DIRECTORY`
  and triggers autofs) is the source of truth, and after it returns
  we set `self._state` to ACTIVE/INACTIVE so the API reports
  reachability rather than autofs idle state.
* `MountManager.reload_mount()` becomes probe-only — no systemd
  reload/restart calls. The user-facing semantics: "tell me if this
  mount is reachable right now, and refresh the resolution issue
  accordingly." If the mount is dead, autofs will re-trigger it on
  the next consumer access; the supervisor doesn't need to force it.
* `Mount.load()` keeps `_update_state_await` for the no-job-dispatched
  case where a previous supervisor left a unit in `activating`, then
  probes once via `is_mounted` so state reflects reachability rather
  than the lazy-mount idle state.
* `BackupManager` no longer walks `bound_mounts`. Backups skip
  network-mount subdirectories during folder archive (so we don't
  recurse into the share), and unmount/re-mount them around folder
  restores (so writes target the local mount-point dir, not the
  remote share).

Behavior tradeoffs accepted:

* Dead media/share access now ~30s ETIMEDOUT instead of empty
  read-only dir. The empty-dir was a workaround for the PID 1
  wedge that autofs eliminates at the root; if specific add-ons
  rely on the old behavior we can revisit with a 2-stage automount
  whose fallback target is an emergency dir.
* Resolution-issue lag: failed mounts no longer surface within
  15 min on their own. They appear when the user probes via the
  API, when `BackupManager.reload()` walks the location, or when
  load-time activation fails.

* tests: adapt bind-layer-era tests after rebase

Main gained tests for the bind layer after this branch was written:
the #7013 rebind regression test and the #7072 bind-step rollback test
cover machinery that no longer exists, and the healthy-reload test
asserted the unconditional rebind. Drop the first two and reduce the
third to asserting that a healthy probe performs no systemd operations
at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: rearchitect automount setup from systemd/kernel review

A code-level review of systemd's automount implementation (verified
against v254.13, the version HAOS ships) and the kernel autofs/VFS
plumbing surfaced three correctness-critical flaws in the automount
design, plus several hardening gaps. See automount-rearchitecture.md
for the full analysis.

- Make the .automount the primary transient unit with the .mount as
  aux, mirroring `systemd-mount --automount=yes`. Aux units get no
  start job: with the .mount as primary the trigger was never armed
  and the design silently degraded to eager mounting.

- Set StartLimitIntervalUSec=0 on the .mount. With the default start
  rate limit (5 starts/10s, counting successful starts) a fast-failing
  mount plus any polling consumer trips the limit within seconds, and
  systemd then detaches the autofs trigger entirely
  (AUTOMOUNT_FAILURE_MOUNT_START_LIMIT_HIT) — the path silently
  becomes a plain writable local directory.

- Drop the 5-minute TimeoutIdleUSec (default 0 = never expire).
  Kernel idle expiry decides busyness via may_umount_tree(), which
  only counts the init-namespace mount instance — open files held by
  container processes are invisible, so expiry would unmount shares
  under actively writing add-ons.

- Probe with a plain statvfs: the statfs syscall walks with
  LOOKUP_AUTOMOUNT and triggers by itself; the trailing-dot trick is
  unnecessary. Classify ELOOP as a mount-propagation
  misconfiguration in the probe error handling.

- Re-arm the trigger from reload_mount(): if the .automount unit is
  failed or gone (e.g. autofs unmounted out-of-band), reset failure
  state and re-create the pair before probing.

- Tear down legacy eager-mount units during load(). The old design's
  bind unit for media/share occupies the exact unit name the network
  .mount uses now; on a warm upgrade adoption would mistake the old
  bind mount for the network mount and leak the legacy data mount.

- Honor the systemd job result when stopping the .mount during
  unmount and reset failure state on both units afterwards so dead
  transient units get garbage-collected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: adapt local data repair to the autofs design

Port of the mount-target-not-empty repair (#7089) on top of the
automount rearchitecture:

- relocate_local_data() handles a single target directory — the mount
  sits directly at its container-facing path, the bind layer and with
  it the multi-directory case are gone
- a repair_trigger() failure caused by blocking local data raises the
  mount failed issue with the move_local_data suggestion instead of
  the plain variant
- local data at load is detected via the mount unit itself; the
  bind-layer detection test is rewritten accordingly

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docker: use rslave propagation for execute_command share mount

The temporary container used for execute_command (e.g. core config
check) mounted /share without a propagation mode, unlike every other
share/media mount. With eager mounts this only meant missing mounts
made after container start; with lazy automount activation the shares
are routinely mounted after start, and accessing a not-yet-activated
automount from a private mount namespace fails with ELOOP.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: harden automount lifecycle from review findings

Address the findings of an adversarial review of the automount
rearchitecture:

- load() no longer adopts a dead automount trigger: a failed or
  stopped .automount leaves the path a plain writable directory — the
  silent local-write degradation this design must prevent. Tear down
  and re-arm instead. An adopted pair whose probe fails now raises so
  the manager surfaces the mount failed issue, same as a fresh mount.
- unmount() resolves the .mount unit only after stopping the
  .automount (the lazy detach can garbage-collect the transient unit,
  invalidating an earlier proxy), raises instead of warning when the
  automount stop errors, and checks the stop job result so a failed
  stop cannot leave an armed trigger behind a "successful" removal.
- repair_trigger() fully unmounts before re-arming: with the share
  still attached, arming fails — or worse, mount() misreads the
  mounted share's contents as blocking local data.
- Legacy unit teardown honors the stop job result for the same reason.
- reload_mount() escalates once to re-creating the unit pair when an
  established mount is unreachable: a permanently dead session (e.g.
  replaced server) keeps the path mounted, so the trigger can never
  re-fire and kernel reconnection never succeeds. Teardown is safe now
  (lazy unmount, no PID 1 path walks), unlike the removed reload →
  restart escalation of the eager design.
- Reinstate the periodic mounts task as a probe-based reconcile: re-arm
  dead triggers, refresh the reachability state reported by the API and
  used by backup locations, and sync the mount failed issue in both
  directions. No reload or restart of established mounts.
- Folder restore no longer fails after a successful restore when
  re-mounting nested mounts cannot verify an unreachable server — the
  trigger is armed, the share recovers on next access.
- Drop the now-unused Mount.update(), deduplicate unit name escaping,
  hoist the backup exclusion set out of the per-file filter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* tests: expect rslave propagation in core check container

The config check assertion missed the update for the rslave
propagation on the execute_command share mount.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* tests: cover reconcile trigger repair failure paths

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: arm automount before stopping the legacy data mount

Address review: the legacy teardown stopped the media/share bind unit
and the eager data mount in sequence, both strictly. When the bind
stop succeeded but the data mount stop failed — its unmount can time
out against an unreachable server, legacy units have no LazyUnmount —
load() raised with the container-facing path left a plain writable
directory and no trigger armed: the pollution mode this design is
meant to prevent. Reordering the stops would not help either, as
systemd stops the bind first anyway as a dependent of the data mount
(the implicit Requires= from RequiresMountsFor= on the bind's What=).

Instead, strictly stop only the path-conflicting unit (a failure
leaves the path covered, which is safe and retryable), arm the
automount right after, and stop the conflict-free legacy data mount
best-effort last — a failed stop logs a warning and leaves an orphaned
mount for the next Supervisor restart or a host reboot, with nothing
writable exposed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: discard dead session instead of re-creating units on reload

Address review: re-creating the unit pair on reload of an unreachable
established mount exposed the target as a plain writable directory
between the lazy unmount and arming the replacement — a continuously
writing add-on could block the re-arm or slip writes under the new
mount in the check-to-arm window.

Stop only the .mount unit instead, keeping the .automount armed:
systemd re-installs the autofs trigger over the path — the same
mechanism idle expiry uses, with the automount's Triggers= reference
keeping the transient .mount definition alive — so the path is never
locally writable. The re-probe then mounts fresh through the trigger,
which establishes a new session and thereby covers the permanently
dead session case (e.g. a replaced server) the escalation exists for.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: surface arming failures, keep restore teardown re-armed

Address review findings on the arming/re-arming error paths:

- mount() checks the StartTransientUnit job result: a failed start job
  means the trigger never armed and the path is a plain writable
  directory — a hard MountError before the probe, so it cannot be
  mistaken for the armed-but-unreachable MountActivationError.
- Folder restore includes the nested-mount teardown in the try block:
  an unmount failing halfway (trigger disarmed, share still attached)
  previously exited before any re-arm, leaving the path unprotected.
  The finally re-arms via repair_trigger(), which no-ops on a still
  armed trigger and handles partially torn-down pairs.
- The post-restore re-arm suppresses only MountActivationError (armed,
  recovers on next access). Any other failure — arming failed, local
  data blocking the target — left the path unprotected and surfaces.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: fold the two legacy unit teardowns into one helper

The eager-mount-era cleanup queried the .mount unit a second time
although load() had just fetched it, and the strict teardown of the
unit occupying the automount's path and the best-effort teardown of
the mount at the mounts data directory were near identical. Reuse the
unit from load() and give the shared helper a strict flag, which drops
a D-Bus round trip per mount on every load.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: re-arm the trigger when unmount cannot stop the mount

Stopping the automount detaches the whole stack at the path, so an
unmount that then fails to stop the .mount leaves a plain writable
directory behind. Arm a fresh pair before raising, best effort — the
local data repair covers what lands there if that fails too.

The automount stop itself needs no such handling: systemd's
automount_stop() enters dead synchronously and the detach it performs
(MNT_DETACH, no server contact) only logs its errors, so a failure
there means systemd could not be reached at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: raise translatable errors for mount operations

Errors from setting up, unmounting and reloading a mount reach users
through the API, so describe what failed in a translatable message and
leave the systemd job result or D-Bus error to the log.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: remove the emergency folder of the eager-mount design

The read-only fallback directory has no place in the automount design.
Remove what is left of it on existing installations, next to the legacy
addons directory cleanup, keeping it if it holds anything but the empty
mount points.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: trim comments to the behavior they describe

Drop the narration of what changed and why from comments and
docstrings, keeping the notes that record why a tempting alternative
does not work. Two comments still described the reload to restart
escalation this branch removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: correct stale claims in manager and reload test

The manager docstring credited the kernel with idle expiry, which is
deliberately disabled, and claimed no polling although the periodic
reconcile probes every 15 minutes. The reload test claimed systemd is
never contacted while the escalation stops the .mount unit; only reload
and restart of the unit are avoided.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: arm the trigger alone if re-creating the pair is rejected

Address review: the re-arm after a failed unmount submits the .mount as
an aux unit, but transient creation requires a pristine unit and the
definition is still loaded precisely when stopping it is what failed.
Fall back to creating the .automount on its own, which covers the path
and fires the surviving mount definition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: never arm the trigger over local data after a failed unmount

Address review: the fallback caught the target validation errors too,
so data written while the path was uncovered would end up beneath a
fresh trigger. Reconciliation would then find a healthy mount and the
repair skips mount points, leaving the data hidden for good. Leave the
path untouched instead, so the next reconcile offers to move it away.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-31 20:30:23 +02:00
Mike Degatano 2d153ac497 Add one-shot mode and docker-given totals to stats API (#7168)
* Add one-shot mode and docker-given totals to stats API

- Add one-shot query option to container stats to skip the calculated
  CPU percentage window, returning docker-given totals directly.
- Refactor container_stats() with translatable, structured Docker
  errors and timeout handling instead of generic Docker exceptions.
- Translate low-level Docker errors into component-flavored errors
  (NotRunningError/StatsTimeoutError/UnknownError) for Home Assistant,
  Supervisor, and each plugin (Audio, Cli, CoreDNS, Multicast,
  Observer) as well as apps, mirroring the existing App.stats()
  pattern for consistent, informative API responses.
- Add a shared api_return_stats() helper to remove duplicated stats
  handling logic across the 8 stats API endpoints.
- DockerContainerNotFoundError now inherits from DockerNotFound.

* Contain aiodocker _query access and split v1/v2 stats response model

- Contain the reach into aiodocker's protected _query internals (needed
  for one-shot stats until aio-libs/aiodocker#1054 lands) to a single
  DockerAPI._query_one_shot_stats() helper.
- Split the stats API response model between v1 (legacy, windowed CPU
  percentage still available via opt-in one_shot query string) and v2
  (always one-shot, no query string, cpu_percent dropped since it would
  never carry a meaningful value). Both share the same
  api_return_stats(stats, legacy=...) helper so the response model
  stays in one place.
- DockerAPI.container_stats() keeps its existing one_shot keyword only;
  the legacy/v1 vs v2 distinction is handled entirely in the API layer.

* Detect stopped/restarting containers in one-shot stats response

Docker returns a stub response containing only the container's id/name
(no cpu_stats/memory_stats/networks data) for a stopped or restarting
container instead of an error, even in one-shot mode - the not-running
check in daemon.ContainerStats happens before the one-shot branch, so
requesting one-shot stats doesn't change this. Detect the missing
online_cpus field (only ever present while running) and raise
DockerContainerNotRunningError instead of returning a payload that
DockerStats would otherwise turn into a successful all-zero response.

* Add tests for stats error paths introduced in one-shot/totals PR

- Cover DockerAPI.container_stats/_query_one_shot_stats edge cases: empty
  one-shot response, empty windowed response, and the raw _query_one_shot_stats
  helper.
- Add stats() error-translation tests (not-running, timeout, unknown) for the
  audio, cli, multicast and observer plugins, none of which had any coverage.
- Add stats endpoint tests (v1 + v2) for the cli, audio, dns, multicast and
  observer REST APIs, plus a v1 success-path test for the apps stats endpoint.

These close the codecov-flagged gaps for newly introduced code in this PR.

* Unify one-shot and windowed container_stats() code paths

Both modes now share the same error handling and not-running detection:
- The windowed path no longer inspects the container first to check its
  running state. Docker returns the same id/name-only stub response (no
  online_cpus in cpu_stats) for a stopped or restarting container in
  windowed mode as it does in one-shot mode, so the existing online_cpus
  presence check now covers both, closing a race window where a container
  stopping between the inspect and the stats call would previously reach
  the DockerStats parser with an incomplete payload.
- The windowed path builds its container handle via containers.container(),
  which constructs a DockerContainer from the name with no I/O, instead of
  containers.get(), which performs an unnecessary inspect call just to
  fetch the same information the stats call already returns.

* Fix protected access pylint issue
2026-08-31 17:11:38 +02:00
Jan Čermák 68ed3b5dec Add endpoint to reset Docker storage (#7166)
* Add endpoint to reset Docker storage

Corrupted Docker image layers cannot be always fixed by removing and
re-pulling a single image, because layers are shared between images. The
only (easy) way to recover is wiping all of Docker's storage which is
currently cumbersome and requires OS shell access.

Home Assistant OS added a service that performs the Docker storage wipe
on boot when a flag file exists in home-assistant/operating-system#4982,
and home-assistant/os-agent#283 adds OS Agent DBus interface for
creating that file.

This adds a new API method POST /docker/reset-storage which calls OS
Agent's ScheduleDockerStorageReset and creates a reboot_required issue,
following the same pattern used for the Docker storage driver migration.

This is expected to ship in HAOS 18.3, so the endpoint checks for that
and returns 404 on older OS versions. Technically, it might be nicer to
gate on OS Agent version, but from user perspective it's better to
report the required OS version.

Refs #6555

* Add const for minimum OS version supporting the reset
2026-08-27 12:53:16 +02:00
Stefan Agner 210b49fb03 Add API to manage SSH authorized keys on Home Assistant OS (#7039)
* Add API to manage SSH authorized keys on Home Assistant OS

The OS Agent has long exposed AddSSHAuthKey and ClearSSHAuthKeys on its
io.hass.os System D-Bus object, but the Supervisor never wrapped them, so
there was no way to manage root's SSH authorized keys through the
Supervisor API. Add POST /os/ssh/authorized_keys, which replaces the
configured keys with the submitted list (an empty list just clears them).

Since OS Agent writes each key verbatim to /root/.ssh/authorized_keys as
root, the endpoint validates strictly before anything is written: plain
public keys only (no options, no certificates), a key type allowlist
matching what dropbear on Home Assistant OS can verify, no control
characters (one submitted key can never write more than one line), the
base64 blob must embed the declared key type, and entries are capped at
dropbear's 3000 byte per-line limit. The endpoint is admin-only for
add-on tokens.

Replacement clears the existing keys and then adds each key. OS Agent
releases up to 1.10.x return an error when clearing an already absent
file (inverted error check, since fixed), which is the state of every
first-time user, so this specific error is treated as the empty state it
reports.

dropbear on Home Assistant OS is gated by ConditionFileNotEmpty on the
authorized_keys file, which systemd only evaluates when the unit starts,
so the service is started after a non-empty key set is written. A running
dropbear re-reads the file on every authentication attempt and needs no
restart.

* Delegate SSH key validation to OS Agent

Review discussion questioned the full key validation (type allowlist,
base64 blob checks, canonicalization): OS Agent 1.10.0 validates
submitted keys itself and treats clearing an already absent
authorized_keys file as success, so the Supervisor can rely on it
instead of duplicating the logic.

Require OS Agent 1.10.0 or newer and reject requests on older releases
with 404, like the Raspberry Pi firmware endpoints do; this also makes
the missing-file compatibility shim for the clear call unnecessary. A
key rejected by OS Agent surfaces as an error response including its
validation message. The Supervisor keeps only a basic per-key sanity
check that runs before anything is written: no control characters (one
submitted key can never write more than one authorized_keys line) and
at most 3000 bytes (dropbear ignores longer lines, which would leave a
key that passes but never works).

* Support OS Agent releases before 1.10.0

Requiring OS Agent 1.10.0 would keep the feature unavailable until the
next OS update reaches users, while Supervisor updates roll out
independently. Drop the version requirement and accept that validation
is lossy on older releases: the Supervisor sanity check still prevents
writing more than one line per key or lines dropbear ignores, but
proper key validation only happens on OS Agent 1.10.0 or newer.

This brings back the need to tolerate the error OS Agent releases
before 1.10.0 return when clearing an already absent authorized_keys
file (inverted error check): match the complete os.Remove error message
for the authorized_keys path and treat it as the empty state clearing
aims for, on affected versions only.

* Split SSH authorized keys API into add and clear endpoints

Review feedback preferred endpoints mapping 1:1 onto the OS Agent D-Bus
methods over a single replace-the-set API, whose POST semantics were
also questioned (an idempotent full replacement would be PUT).

POST /os/ssh/authorized_keys now takes a single key ({"key": "..."})
and appends it via AddSSHAuthKey; DELETE /os/ssh/authorized_keys
removes all keys via ClearSSHAuthKeys. Each call maps to exactly one
OS Agent operation, so no request can partially succeed. Clients that
want to replace the configured set clear and re-add; a GET (which
needs an OS Agent extension first) and an idempotent PUT can be added
later.

The per-key sanity check, the dropbear service start after adding a
key, and the tolerance for the missing-file clear error of OS Agent
releases before 1.10.0 carry over unchanged.

* Add endpoint to list SSH authorized keys

With only add and clear operations the authorized_keys file is
write-only for API consumers: a user cannot audit which keys grant
access to the box, or whether any exist at all — including keys that
were imported from USB or written by add-ons.

OS Agent 1.11.0 added a ListSSHAuthKeys D-Bus method. Expose it as
GET /os/ssh/authorized_keys, returning the configured entries verbatim.
The endpoint requires OS Agent 1.11.0 and returns 404 on older
releases, like the Raspberry Pi firmware endpoints do; add and clear
keep working on all OS Agent releases.

* Restrict SSH authorized keys endpoints to Home Assistant Core

Review decision: manipulating root's SSH access is not a capability
add-ons should have, even with the admin role, so move the endpoints
from admin-only to the core_only middleware pattern. Only requests
authenticated with the Home Assistant Core token pass; add-on tokens of
any role, the CLI plugin, and the observer are rejected. This also
means the host shell (ha CLI) cannot use the endpoints for now — the
restriction can be opened up later.

The exclusion from the manager role allowlist is kept: if the path is
ever removed from core_only again, it falls back to admin-only rather
than becoming manager-accessible.

* Stop dropbear after clearing SSH authorized keys

Clearing all authorized keys is a revocation, but without stopping
dropbear the listener keeps running until reboot and established
sessions survive, as only a service stop terminates them. Stop the
service after a successful clear, mirroring the USB config import
(haos-config), which also stops dropbear when the imported
authorized_keys file is removed. Stopping an inactive unit is a no-op.

* Serialize SSH authorized keys jobs on a common lock

The add and clear jobs each perform a file operation followed by a
service operation, and nothing prevented them from running
concurrently. An interleaving like clear-file, add-key, start-dropbear,
stop-dropbear lets both requests succeed while the final service state
does not match the final key state (key configured, dropbear stopped).

Make OSManager a JobGroup and run both jobs with GROUP_QUEUE
concurrency, so each file-and-service operation completes before the
next starts. The regression test fails without the shared lock.
2026-08-27 10:12:44 +02:00
Stefan AgnerandClaude Fable 5 69c5a53864 Surface fixup failures to the caller applying a suggestion (#7145)
* Surface fixup failures to the caller applying a suggestion

Applying a suggestion whose fixup failed reported success: the fixup
base swallowed ResolutionFixupError, the API returned OK, and the
repair flow in Home Assistant completed and removed the repair from
the UI while the issue persisted in Supervisor.

Let the error propagate from the fixup instead. The autofix loop
already handles and logs per-fixup errors, bus-event triggered fixups
now get the same treatment in the event callback, and a user-applied
suggestion surfaces the failure as an API error so clients can show
that the fix did not apply. ResolutionFixupError becomes an APIError
with a translatable error key so it is reported as a client-visible
error instead of an unexpected one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Drop ResolutionFixupError, let fixup failures bubble

Per review: the generic wrapper existed for a time when suggestions
were only applied by autofix and the one job was separating fixup
failures from real bugs for Sentry. The suggestion API is in regular
use now and the wrapper actively hurts it — well-defined errors from
the underlying operations were caught and replaced with a generic
message.

Remove the exception type entirely and let the original errors reach
the caller. The autofix loop and the bus-event fixup path treat any
HassioError as an environmental/config failure: log and continue
without Sentry capture (the raise site reports to Sentry where
warranted); everything else is still captured. Direct raises in the
data disk fixups become HassOSDataDiskError, the could-not-start check
in the app start fixup becomes AppsError. ResolutionFixupJobError now
derives from ResolutionError and JobException.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Share unattended fixup error handling between autofix and bus events

Address review: the bus event listener only caught HassioError, so a
Supervisor bug in a bus-triggered fixup ended as an unretrieved task
exception without Sentry report, unlike the same bug during autofix.
Both unattended paths now use a common apply_fixup_safely() helper:
log HassioError (environmental, reported at the raise site if
warranted), capture everything else.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 23:14:44 +02:00
Thibault MollemanandStefan Agner 6947d8b2b2 Strip whitespace from repository URLs in store validation (#7161)
* Strip whitespace from repository URLs in store validation

* Reject repository URLs containing internal whitespace

Tighten RE_REPOSITORY so the URL group rejects whitespace that
.strip() cannot fix, failing validation with a clear error instead of
an unclear git clone failure later. Also parametrize the existing add
repository API test instead of duplicating it and collapse the strip
test cases into one mixed-whitespace case.

---------

Co-authored-by: Stefan Agner <stefan@agner.ch>
2026-08-24 15:48:48 +02:00
Mike Degatano 090ef31ec0 Refactor v2 API contracts to remove deprecated supervisor/apps/backups fields (#7111)
* tests(api): cover v1/v2 deprecated contract split

* api: drop deprecated fields from v2 supervisor/apps/backups contracts

* Proper fix for review finding

* Fixes from feedback
2026-08-21 11:55:28 +02:00
Petar Petrov f148020f3a Harden mount storage usage edge cases (#7154)
Follow-up to the per-mount storage usage endpoint, addressing three review
findings on the merged change.

Depths below 2 all yield the same totals-only response for a mount, so they
are normalized to one value before the probe registry is keyed. Concurrent
callers varying only the depth now share a single probe instead of each
parking an executor thread on the same mount.

A directory walk that races a deletion can count data the filesystem figure
no longer includes, producing children that sum past their parent. Such a
breakdown now gets dropped in favor of the totals, which are the consistent
part; the next request re-walks a settled tree.

The v1 legacy addon-id remap is now applied only to the system disk target.
A mount's children are real directory names, so one that happens to be
called apps_data must not be renamed - and a mount response dict is shared
with every concurrent waiter on the same probe, so remapping it in place
could leak renamed ids into v2 responses.
2026-08-20 15:13:35 -04:00
Stefan AgnerandClaude Fable 5 ccdd8c1ef4 Add repair to move local data blocking a mount target (#7089)
* Add repair to move local data blocking a mount target

When an add-on writes into a media/share directory while its network
mount is not in place (#7037), the local data blocks re-creating the
mount: mounting over a non-empty directory is refused. Until now this
failed silently at Supervisor startup — the bind mounts were created as
fire-and-forget tasks — and the only way out was to remove the data
manually over SSH/Samba and re-create the mount via the API.

Surface the condition as a new mount_target_not_empty issue and offer a
move_local_data suggestion. The fixup moves the blocking data to a
<name>_local_recovery folder in a user-accessible location — media or
share for bind mount targets, local backup storage for backup mounts
(their data mount directory is not reachable for users) — then reloads
the mount. Nothing is deleted; users can inspect and clean up the
recovered data via the media browser or the share and backup folders.

Bind mount failures during load are now awaited and routed into
resolution issues instead of being swallowed as fire-and-forget tasks;
bind failures other than blocking local data create the existing
mount_failed issue. A successful mount reload dismisses a stale local
data issue.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Attach move local data suggestion to the mount failed issue

Review feedback on the repair: rather than introducing a separate
mount_target_not_empty issue type, keep a single mount failed issue
per mount and offer moving the blocking data as an additional
suggestion alongside reload and remove. Reload stays available for
users who prefer to clear the data themselves, and at most one repair
exists per mount. Adding is idempotent, so an already-raised mount
failed issue just gains the extra suggestion.

When re-creating the bind mount after a successful reload fails on
blocking local data, the mount failed issue is re-added together with
the move suggestion, since the reload already dismissed it.

This also resolves the reviewer note about not-a-directory conflicts
being reported under a not-empty issue type: the issue type no longer
encodes the filesystem detail, while the error messages keep the
distinction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Keep an empty directory in place after relocating local data

If remounting fails after the local data was moved aside (e.g. the
server is unreachable at that moment), the renamed directory left
nothing behind: media/share consumers saw the folder disappear
entirely. Recreate an empty directory right after the rename so the
path stays present regardless of whether the remount succeeds.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Drop move local data suggestion once the data was moved

When the remount after relocating local data fails (e.g. the server is
unreachable at that moment), the mount failed issue stays — but the
move suggestion stayed with it, offering to move data that is no
longer in the way. Dismiss the suggestion after the relocation step so
only reload and remove remain for the leftover failure. Detection
re-adds the move suggestion if local data blocks the target again.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Filter Core-facing suggestions by minimum Core version

The fix flow translations for a new suggestion ship with a Core
release. Older Core frontends render an unknown suggestion as an
empty, unlabeled menu entry in the repair fix flow. Filter such
suggestions from Core-facing output — the issue events sent over the
websocket and the resolution API responses when the caller is Home
Assistant — until the connected Core is new enough. Other API
consumers like the CLI always see the full suggestion list.

The move_local_data suggestion requires Core 2026.9.0b0, the release
its fix flow translations are targeted at.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Make fixup failure tests independent of error propagation behavior

Suppress a potential ResolutionFixupError from the failing fixup calls
so the tests pass both while fixup errors are swallowed and once they
propagate to the caller (#7150).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address review: one recovery folder, check OSError for known issues

Move all local data blocking a mount into a single recovery folder so
the user finds it as one fix: when more than one directory holds data,
later ones become subfolders named after their parent directory
instead of numbered sibling folders.

Also run OSError from the relocation through check_oserror to pick up
known filesystem issues like corruption (bad message).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Only filter Core-facing suggestions on v1 surfaces

Per review: no Core version predating the repair suggestion filtering
in its own fix flow (home-assistant/core#179540) supports the v2 API,
so the Supervisor-side compatibility filter is only needed where old
Core versions actually look. Filter the v1 resolution endpoints and
the legacy websocket issue payloads; the v2 endpoints and v2 event
payloads always carry the full suggestion list. The suggestions for
issue endpoint gets a v1 handler for this.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 23:40:47 +02:00
Petar Petrov 143da593b8 Report storage usage per mount (#7146)
Parameterize the disk usage endpoint as /host/disks/{disk}/usage so a
supervisor mount can be asked for its own storage figures. "default" still
means the system disk and keeps its existing response byte for byte; any
other value names a mount, which must exist and be active. The reserved
target wins over a mount of the same name, which the mount name pattern
permits.

Usage comes from the mount's own filesystem, with used derived as total
minus free so that reserved space counts as used, as it does for the system
disk.

A mount reports totals only by default. The directory walker recurses
regardless of max_depth and only gates whether children appear in its
output, so walking a whole network mount to report a single number would
cost many round trips for nothing; at depth 0 it is skipped outright rather
than called. Depth otherwise means what it means for the system disk, where
level 1 is the labeled known paths that a mount does not have, so a mount's
own subdirectories start at level 2. When a breakdown is produced, whatever
the walk cannot attribute is reported as an "other" child, which keeps every
node's children summing to its used_bytes.

Probes are deliberately not cut short. A caller showing a loader is better
served by a real answer than a fast failure, so the timeout is only a
backstop against a probe that never returns, and a slow one is confined to
its executor thread rather than blocking the rest of the API. Concurrent
callers asking for the same mount at the same depth share one probe instead
of each parking a thread on identical work.
2026-08-19 11:09:06 +03:00
Mike Degatano 9ad1c198c2 Block supervisor/hassio WebSocket command types in the HA proxy (#7123)
* Block supervisor/hassio WebSocket command types in the HA proxy

An app with only homeassistant_api: true could send supervisor/api (or
other supervisor/* / hassio/* command type) frames through the Supervisor's
Core WebSocket proxy at /core/websocket. Core executes those via the hassio
integration's websocket_supervisor_api handler, which calls back into the
Supervisor using Core's own token. That token bypasses all role checks in
security.py, giving the app unrestricted Supervisor API access regardless of
its declared hassio_role or hassio_api flag.

Fix: filter command types on the app-to-Core direction of _proxy_message.
Any TEXT frame whose type field starts with 'supervisor/' or 'hassio/' is
rejected with a Core-shaped result/unauthorized response instead of being
forwarded. The connection stays open. Fail-open on unparseable frames (Core's
own validation already rejects those). The Core-to-app direction is
unfiltered.

* Remove the double json parse
2026-08-19 08:57:10 +02:00
Jan Čermák 384ed39b39 Use port 80 as default for Core/landingpage, omit scheme's default port in API URL (#7153)
Set default port of HA to 80, as this default is authoritative for
what's shown on the HA CLI banner when landing page is running - there's
no homeassistant.json at this time, so this value applies. We shouldn't
need to care about installs running an old landing page (without port 80
support at all), as Supervisor update results in the old default being
persisted homeassistant.json which is written on Supervisor restart, so
the banner still shows port 8123 if the old Supervisor (and hence old
landing page) were baked in an OS image.

Also append the port to the API URL only when it's not the scheme's
default to avoid :80/:443 suffixes when they're not necessary.

Closes #7151
2026-08-18 18:30:39 +02:00
2578b1bf03 Hassio integration auth bypass from app with API access (#7122)
* Block Core hassio_auth endpoints from the add-on proxy

The API security blacklist is meant to stop add-ons from reaching Core's
"hassio" endpoints through the /core/api and /homeassistant/api proxy, but
the pattern only matched "hassio/" (with a trailing slash). Core's auth
endpoints are served at /api/hassio_auth and /api/hassio_auth/password_reset,
so they slipped past the blacklist and were passed through to the proxy.

The proxy authenticates upstream to Core as the Supervisor user, and Core's
HassIOPasswordReset only checks that the caller is the Supervisor user (no
owner check). As a result an add-on with homeassistant_api access could reach
the password-reset endpoint through the proxy and reset any user's password,
including the owner.

Widen the boundary after "hassio" to match both the loopback ("hassio/...")
and the auth endpoints ("hassio_auth...") so all hassio-prefixed Core
endpoints are blocked, and extend the blacklist test to cover them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XJ56EftdwXvqw2p3vjYS1z

* Refuse to proxy Core hassio endpoints (defense in depth)

Add a redundant guard in the Home Assistant API proxy so an add-on can never
reach Core's Supervisor-only "hassio" endpoints (hassio_auth,
hassio_auth/password_reset, the hassio loopback) through the proxy. These run
as the Supervisor user on Core, so forwarding them would let an add-on reset
arbitrary user passwords.

The security middleware blacklist already blocks these paths; this guard sits
at the proxy itself so the proxy cannot become a confused deputy if that
blacklist ever regresses. The two checks are independent by design.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XJ56EftdwXvqw2p3vjYS1z

* Check access before denylist

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-08-12 08:15:50 +02:00
c01efa05a5 Rename resolution addon issue types and check slugs to app-based naming (#7098)
* Rename resolution addon issue types and check slugs to app-based naming

Renames the following for consistency with the new apps terminology:
- Issue types: deprecated_addon -> deprecated_app,
  deprecated_arch_addon -> deprecated_arch_app,
  detached_addon_missing -> detached_app_missing,
  detached_addon_removed -> detached_app_removed
- Check slugs: addon_pwned -> app_pwned,
  deprecated_addon -> deprecated_app,
  deprecated_arch_addon -> deprecated_arch_app,
  detached_addon_missing -> detached_app_missing,
  detached_addon_removed -> detached_app_removed

Full backward compatibility is maintained:
- REST API V1 (root) returns legacy names and accepts legacy slugs
- REST API V2 (/v2) returns and accepts new names only
- WebSocket events use legacy names unless SUPERVISOR_WEBSOCKET_V2_API
  feature flag is enabled
- resolution.json files with legacy check slugs are automatically
  migrated to new names on load

Fixes #7029

* Apply suggestions from code review

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* Address review nits for resolution compatibility maps

* Update supervisor/resolution/const.py

---------

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Stefan Agner <stefan@agner.ch>
2026-08-04 10:57:42 +02:00
Stefan AgnerandClaude Fable 5 c5ea8e174e Gracefully handle lost systemd-journal-gatewayd connections (#7106)
* Gracefully end log stream when journal gateway connection is lost

When systemd-journal-gatewayd is stopped while a client follows logs
(e.g. on host reboot with the log viewer open), aiohttp raises
ClientPayloadError and advanced_logs_handler converted it to an
APIError. For the /supervisor/logs endpoints this got logged as an
unexpected error with a full traceback and captured to Sentry on every
occurrence (#7103, SUPERVISOR-1FHT).

Once the streaming response has started, an error response can no
longer be delivered anyway, so treat a lost connection to
systemd-journal-gatewayd like a client-side disconnect and end the
stream gracefully. The APIError is still raised when the connection is
lost before any data was sent to the client.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Treat unreachable systemd-journal-gatewayd as a known API error

Manual testing of the previous commit (killing systemd-journal-gatewayd
while Supervisor is running) showed that every log API request hitting
the dead gateway logs an "Unexpected error during API call" traceback
and captures HostServiceError to Sentry (SUPERVISOR-K8C), in addition
to the ERROR already logged at the raise site in journald_logs().

Make HostServiceError inherit from APIError as well, following the
HostContainerLogEpochError precedent, so the api_process decorators
return a plain 400 response without the redundant traceback and Sentry
capture. Also treat it like HostNotSupportedError in the supervisor
logs fallback wrapper: fall back to Docker container logs with a
warning instead of an exception log plus Sentry capture.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Use dedicated exception for journal gateway connection errors

HostServiceError is also raised by ServiceManager for systemd units
called through the API (e.g. /os/config/sync), where inheriting from
APIError would hide genuine service breakage from Sentry. Introduce
HostJournalGatewaydConnectionError subclassing HostServiceError and
APIError, and raise it for the gatewayd connection failure only, as
suggested in the PR review.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 10:12:23 +02:00
Mike Degatano 3dea234189 Add v1/v2 terminology shim for host disk usage response (#7100)
Implement API versioned response terminology for GET /host/disks/default/usage:
- V2 returns app terminology as-is (apps_data/apps_config)
- V1 applies compatibility shim and returns legacy addon terminology
  (addons_data/addons_config)

Also adds tests to verify:
- v1 keeps legacy IDs
- v2 keeps app IDs
- disk usage endpoint behavior remains stable across max_depth scenarios

Fixes #7030
2026-07-31 14:56:57 +02:00
Stefan AgnerandClaude Fable 5 f8ba166ed2 Enter STOPPING state before Supervisor restart API response (#7085)
* Enter STOPPING state before Supervisor restart API response

The Supervisor accepts new API requests in the window after responding
to /supervisor/restart: restart() schedules Core.stop() as a task and
returns immediately, so the state only changes to stopping after the
restart response has been sent. A request accepted in that window keeps
running while the Supervisor shuts down, and when stop() tears down the
API server after the 10 second stage 1 timeout the connection is
dropped without a response - the client sees EOF instead of an error.
This is what made the CI restore step flaky (see #7084 for the CI-side
fix): the restore request landed on the old, dying instance.

Make restart() transition to STOPPING before returning, via a new
Core.begin_stop() which contains the transition part previously at the
start of Core.stop(). The system validation middleware already rejects
requests outside of STARTUP/RUNNING/FREEZE, so any request arriving
after the restart response now gets a clear "System is not ready" error
instead of possibly being accepted and killed. Since the state can now
already be STOPPING when stop() runs, its re-entry guard is changed
from a state check to an explicit flag.

Supervisor.update() schedules the same stop task but is left unchanged:
entering STOPPING before update() returns would suppress the final job
progress event to Home Assistant (WebSocket messages are dropped in
CLOSING_STATES), regressing update progress reporting.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Allow stop() retry when begin_stop() fails

Address review feedback: stop() set _stop_initiated before awaiting
begin_stop(), so an exception or cancellation during the state
transition (e.g. a failing config write in _update_last_boot) would
leave the flag set and turn every later stop() call into a no-op,
with no way to retry the teardown.

Reset the flag and re-raise when begin_stop() fails. Nothing has been
torn down at that point, so a retry is safe.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Use a stopping_complete event instead of begin_stop()

Address review feedback questioning the extra infrastructure: instead
of splitting the state transition out of stop() (begin_stop()) and
guarding re-entry with an explicit flag, have stop() take an optional
stopping_complete event which is set once the STOPPING state is
entered, following the same pattern as backup/restore's
validation_complete. Core.stop() keeps its original structure and
re-entry semantics, and restart() waits for the event before
returning.

Should the stop task fail ahead of the state transition (only
possible through a failing config write in _update_last_boot()), the
event is never set and the restart request runs into the client
timeout - an accepted trade-off to keep restart() simple.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 13:14:55 +02:00
Stefan AgnerandClaude Fable 5 6bbb82aec0 Return a proper API error when adding a duplicate store repository (#7087)
Adding a repository that is already in the store raised a plain
StoreError. Since that is not an APIError, the api_process decorator
logs it as an unexpected error and reports it to Sentry, where it is
one of the most frequent issues (SUPERVISOR-1JYE, ~8000 affected
installations in 90 days). The events are ordinary client actions:
add-on setup flows and documentation links re-submitting a repository
the user already has, plus third-party automation re-adding its
repositories on every add-on start.

Introduce StoreRepositoryAlreadyAddedError, which is both a StoreError
and an APIConflict with an error key and message template, so the
request fails with a structured 409 Conflict response and no Sentry
report.

Fixes SUPERVISOR-1JYE

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 12:00:14 +02:00
Stefan AgnerandClaude Fable 5 f4ec256b94 Report real OS version to Core versions consuming version_pending (#7086)
* Report real OS version to Core versions consuming version_pending

The /os/info endpoint reports an installed update pending activation as
the current version so that Core update entities unaware of
version_pending don't offer the update again. Since
home-assistant/core#177155 Core uses version_pending to determine the
OS update state, so newer Core versions should get the real current
version again.

Limit the compat shim to Core versions predating that support. The Core
PR merged 2026-07-24 04:21 UTC, after that day's 02:00 UTC nightly
build, so the first nightly containing it is 2026-07-25's
(2026.8.0.dev202607250xxx).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Use exact first nightly version for the version_pending gate

The 2026-07-25 nightly (2026.8.0.dev202607250310) is the first build
containing home-assistant/core#177155: the 2026-07-24 nightly was built
from a commit predating the merge (verified via commit ancestry of the
builder workflow runs). Replace the midnight floor with the actual
published nightly version, matching the CORE_UNIX_SOCKET_MIN_VERSION
precedent of using the exact nightly stamp.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 17:34:53 +02:00
Mike Degatano 4152855713 Rename backup and restore job stages from addon(s) to app(s) (#7059)
* Mark backup/restore app stage names as breaking

Rename backup and restore job stages from addon/addons to app/apps, including delta and await restart stages, and update manager and tests accordingly.

* Add legacy backup stage mapping and fix enum typo

Map backup/restore job stage names to legacy websocket/REST v1 values while preserving v2 values, add API tests for both modes, and rename COPY_ADDITONAL_LOCATIONS to COPY_ADDITIONAL_LOCATIONS.
2026-07-24 12:59:58 +02:00
Mike Degatano 96798b5041 Migrate addon API error metadata to app naming (#7058) 2026-07-23 16:02:49 +02:00
Mike DegatanoandStefan Agner c119e0cff1 Rename Docker app container and builder names from addon_ to app_ (#7024)
* Rename Docker app container and builder names from addon_ to app_

* Update container state event tests to app_ names

* Migrate legacy app container names on attach

* Fix error handling

* Fixes from feedback

* Fix pytest

---------

Co-authored-by: Stefan Agner <stefan@agner.ch>
2026-07-23 15:43:33 +02:00
Mike Degatano a101e2eb5b Add websocket v2 jobs legacy-name toggle (#7051) 2026-07-17 14:07:36 +02:00
Brandon b977824507 Fix Core API proxy stripping the multipart boundary from Content-Type (#7050)
The proxy forwarded aiohttp's parsed request.content_type property, which
drops header parameters — including the multipart boundary. Core then
cannot parse any multipart body proxied from an add-on and raises
"boundary missed for Content-Type".

Forward the raw Content-Type header instead, matching what the stream()
proxy path already does.

Fixes #7049
2026-07-17 14:06:41 +02:00
Stefan AgnerandClaude Fable 5 dfe727410c Fix length validation of str and password options with bounds (#7044)
* Fix length validation of str and password options with bounds

App option schemas like str(1,32) validated the string value with
vol.Range, which compares the value itself against the numeric bounds.
Comparing a str with a float raises TypeError, which voluptuous reports
as "invalid value or type (must have a partial ordering)". As a result,
saving options always failed for any add-on using a bounded str or
password schema; unbounded str/password was unaffected since vol.Range
without bounds performs no comparison.

Use vol.Length to check the string length against the bounds, which is
the documented meaning of str(min,max). Also replace the str(value)
literal (a self-equality check) with the str type so non-string values
still fail validation, now with a clearer error message.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Update API test for changed non-string option error message

Validating a str option against the str type instead of the str(value)
literal changed the voluptuous error for non-string values from "not a
valid value" to "expected str". Update the expected message in the API
options error test accordingly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 13:31:28 +02:00
Mike Degatano 4a54a229b9 Migrate addon job names to app and scope legacy API compatibility (#7014)
* Migrate addon job names to app and trim legacy API aliases

* Limit job name compatibility to app_manager_update only
2026-07-14 13:45:40 +02:00
Stefan AgnerandClaude Fable 5 83d15aadcc Track installed OS update pending reboot activation (#7006)
* Track installed OS update pending reboot activation

Since #6982 the OS update no longer reboots the host automatically, so the
system keeps running the old version until the user reboots. In that window
the update kept being offered: need_update compares the running OS version,
which only changes on reboot. Requesting the same update again re-downloaded
and reinstalled the full image. Worse, a Supervisor restart in that window
dropped the in-memory REBOOT_REQUIRED issue and canceled the pending update
altogether, because mark_healthy marks the booted slot active on startup,
reverting the primary boot slot set by the rauc install.

Track the installed version awaiting a reboot as version_pending and expose
it in /os/info. A successful install sets it, updating again to that version
is rejected with a hint to reboot, and need_update no longer reports true for
an update that is already installed. On load, the pending state is recovered
from rauc by comparing the primary boot slot (via a new GetPrimary D-Bus
wrapper) with the booted slot, re-creating the REBOOT_REQUIRED issue as well.
mark_healthy now skips marking the booted slot active while an update is
pending.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Split pending update detection condition

Fix pylint R0916 (too-many-boolean-expressions) by splitting the guard in
_detect_pending_update into a slot data validity check and the actual
pending update check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Report pending OS update as current version in /os/info

The Core update entity derives update availability by comparing version
with version_latest from /os/info, so it keeps offering an update that is
already installed until the system is rebooted. Report an installed update
pending activation as the current version so existing Core releases
reflect update availability correctly. Once Core consumes version_pending,
this can be limited to Core versions predating that support.

The hassos field of the root /info endpoint keeps reporting the running
version, Core only uses it as a HAOS presence check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 18:45:18 +02:00
Mike Degatano 35542d5bcc Replace addon with app in docker container names for apps (#6997)
* Support dual app/addon log identifiers during migration

* Use single journald epoch query for dual app identifiers

* Add API-keyed container epoch error handling

* Template CONTAINER_LOG_EPOCH via extra fields
2026-07-03 15:44:22 +02:00
Mike Degatano a483d4a503 Make /os/update stop auto-rebooting after OS update (#6982)
* Make OS update require explicit reboot

* Avoid protected OS field access in API update tests
2026-06-30 14:49:24 +02:00
Stefan AgnerandClaude Opus 4.8 1857753e22 Redact app options in info for unprivileged apps (#6953)
The `/addons/{slug}/info` endpoint returned the target app's user options,
which can contain secrets such as passwords and API keys. The security
middleware grants every role (including the default role) access to any
`/.+/info` path, so an installed app with `hassio_api: true` and the default
role could read another app's options simply by requesting its info.

Redact the options field in info_data() unless the caller is entitled to see
it: Home Assistant Core (and other non-app internals), the app reading its
own info, or an app with the manager or admin role. Other apps reading a
different app's info now receive an empty options dict while all non-secret
metadata stays available for discovery. This mirrors the existing self-only
restriction on the dedicated /options/config endpoint.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 14:54:57 +02:00
Stefan AgnerandClaude Opus 4.8 81e235376e Fix typos repo-wide and add codespell pre-commit hook (#6949)
* Fix typos in comments, docstrings and log messages

Correct 39 spelling mistakes across comments, docstrings and log/error
message strings throughout the package (e.g. "conection" -> "connection",
"Incomming" -> "Incoming", "Rasie" -> "Raise"). All changes are confined
to human-readable text; no identifiers, attributes or D-Bus contracts are
touched, so there is no behavior change. Found with codespell.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix typos in tests and CI workflow

Correct spelling mistakes in test comments, docstrings and data, plus one
in the builder workflow, so the whole tree is clean for the codespell hook
added next. The assertion in test_network_manager.py is updated to match
the corrected "Unknown error while processing" log message in the source.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Add codespell pre-commit hook

Wire up codespell so spelling mistakes in comments, docstrings and strings
are caught automatically. The vendored frontend panel is excluded, and
"hass" and "astroid" are added to the ignore list as known false positives
(the Home Assistant abbreviation and the pylint dependency package).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Address review feedback

Improve grammar in several of the touched comments and docstrings: use the
plural "ignore conditions" for the list-returning property, add the missing
auxiliary verb and fix agreement in the timezone-filter comment, fix
"backups ... use" agreement, and reword "underlay" to "underlying" in the
arch module docstring.

Also drop the "*.json" skip from the codespell hook. It was carried over
from another project but is unnecessary here (all tracked JSON is clean),
and skipping it would needlessly leave translation and data JSON unchecked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Reword onboarding comment

"overflight" was a literal calque of the German "überflogen"; use the
idiomatic "skimmed through" instead.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 14:09:16 +02:00
Stefan AgnerandClaude Opus 4.8 dae48c62e4 Treat mount errors as API errors instead of unexpected failures (#6946)
Mount failures generally reflect user configuration or host conditions
(unreachable server, wrong credentials, ...) rather than a Supervisor
bug. As a plain HassioError, MountError reached the generic error branch
of api_process, which logged a full traceback as an "Unexpected error
during API call" and captured the exception to Sentry. The mount reload
path already encoded the opposite intent by explicitly skipping Sentry
for MountError.

Make MountError an APIError so mount failures are surfaced as client-side
errors with their explicit message, without a traceback or Sentry noise,
matching the existing JobException handling. MountNotFound additionally
inherits from APINotFound so it returns a 404 instead of a 400.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 09:32:22 +02:00
Jan Čermák ecf890a41c Fix blacklist for v2 endpoints in security middleware, add v2 security tests (#6933)
The blacklist in the security middleware didn't take the v2 prefix into
account, allowing to call routes that are supposed to be blacklisted for
apps with hassio and homeassistant APIs enabled in app config.

Add test checking that these routes are always blacklisted, and
parametrize other tests using v2 endpoints in the test_security module.
2026-06-15 17:55:58 +02:00
Jan Čermák 446a1aacd9 Return 404 for raspberrypi endpoints on boards without firmware (#6926)
* Return 404 for raspberrypi endpoints on boards without firmware

Return "not found" when board doesn't have the firmware update
available. Any other error may indicate it's worth retrying later (as
discussed in [1]), so 404 is more appropriate here.

[1] https://github.com/home-assistant/core/pull/172929#discussion_r3363706261

* Consistently return APINotFound with debug log
2026-06-09 17:15:19 +02:00
1e5d7812cd Bump aiohttp from 3.14.0 to 3.14.1 (#6922)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Stefan Agner <stefan@agner.ch>
Signed-off-by: dependabot[bot] <support@github.com>
2026-06-08 11:44:30 -05:00
Stefan Agner d910efa49e Surface Home Assistant update failures as translatable API errors (#6915)
Updating Core to a non-existent version (e.g. a mistyped beta tag) reported
success on the CLI: the image pull failed with a 404, but update() swallowed
the error via "with suppress(HomeAssistantError)" and then ran its post-update
health check against the still-running old Core, which passed.

Split the update routine into image install and start phases. A failed image
install leaves the running Core untouched, so there is nothing to health-check
or roll back; let it bubble out instead of masking it. Only failures after the
image is in place (e.g. the new container starting unhealthy) fall through to
the health check and rollback logic, where the error is now captured at debug
level rather than silently suppressed.

Model the update errors as client-facing APIError subclasses carrying an
error_key so the frontend can translate them and they no longer reach Sentry as
unexpected errors:

- HomeAssistantUpdateError: generic update failure
- HomeAssistantUpdateImageError: image download failed (includes the version)
- HomeAssistantUpdateAlreadyInstalledError: requested version already installed
2026-06-08 15:39:23 +02:00
Stefan AgnerandClaude Opus 4.8 8ec1c33aa4 Fix WebSocket transport None race condition in proxy (#6241)
Add a transport validity check before the WebSocket upgrade to handle
clients that disconnect during the handshake.

The connection can be lost between the Home Assistant API state check and
the server.prepare() call, leaving request.transport as None. aiohttp's
_pre_start() then raises ConnectionResetError (an AssertionError prior to
aiohttp 3.14.0), which propagates out of the handler as an unhandled
exception. The result is a 500 response and a Sentry report for what is
really just a client disconnect.

The fix detects the closed connection early and raises HTTPBadRequest
with a clear reason, turning the race into a clean 4xx response with a
warning log instead of error noise.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 09:48:47 +02:00
2f331aafa9 Add Raspberry Pi firmware update API (#6886)
* Add Raspberry Pi firmware update API

Expose `io.hass.os.Boards.RaspberryPi.Firmware` via a D-Bus proxy and
`GET/POST /os/boards/raspberrypi/firmware[/update]` REST endpoints for
Raspberry Pi 4 / 5 / CM4 (Yellow). Gated on OS Agent >= 1.9.0. The
update job raises a `REBOOT_REQUIRED` resolution issue on success and
rejects up front when the agent reports `update_blocked`.

The `blocked_reason` field currently returns only
`unsupported_boot_device` regardless of the underlying cause (CM4
without self-update, USB/NVMe boot, etc.), more reasons may be added
later if we need to make distinction.

Refs home-assistant/operating-system#4631

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Fix pylint issue in tests

* Return blocked_reason=None instead of empty string when update is not blocked

* Fix typo in update_raspberrypi_firmware docstring

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* Reject API call for update early if blocked

* Flip API availablility check in _check_rpi_firmware_available

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* Remove extra newline in docstring

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-06-02 17:37:25 +02:00
Stefan AgnerandClaude Opus 4.8 9f553c327c Only force a versioned Supervisor update in DEV mode (#6903)
Home Assistant Core now triggers a versionless Supervisor update during
onboarding to ensure Supervisor is current before it continues setup. It
treats a "no update available" response as the signal to proceed.

In DEV mode the update endpoint bypassed the need_update check entirely
and resolved a versionless request to the latest published version. So
Core's onboarding call made a freshly-built DEV Supervisor install the
latest published build and restart, which breaks the run_supervisor CI
job (the API disappears mid-test with "connection refused").

Tie the bypass to an explicit version instead: specifying a version is
still DEV-only, but a versionless request now always respects
need_update. Since need_update is always False in DEV, Core's onboarding
call becomes a no-op there, avoiding the update to the latest published
Supervisor on the dev channel.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 13:26:22 +02:00
Stefan AgnerandClaude Opus 4.8 a973d22e35 Derive App state from container state (#6890)
* Derive App state from container state

The App.state setter mixed two responsibilities: it both mutated a
private `_state` field and dispatched side effects (WebSocket events,
issue dismissal, startup_event signaling). On top of that, an installed
but never-started app stayed in AppState.UNKNOWN forever, because the
attach() image-only fallback never fires a container state-change event
and the AppState therefore kept its constructor default. Conceptually,
ContainerState.UNKNOWN ("container does not exist") and AppState.UNKNOWN
("nothing observed yet") happened to share a name but meant different
things, which made the distinction easy to lose.

Make App.state a pure derived property. The source of truth is the last
observed ContainerState (cached on the App), plus a sticky operation-
error flag for start/stop failures that the docker event stream cannot
reflect. When no container has been observed yet, the derivation falls
back to install signals: an attached instance (image present) is
STOPPED, otherwise UNKNOWN. As a side effect, an installed-but-never-
started app now correctly reports STOPPED instead of UNKNOWN.

container_state_changed updates the cached container state and routes
all side effects through a single _emit_state_change helper that diffs
old vs new derived state. The two start/stop failure paths route
through _set_operation_error. Uninstall resets the cached signals so
the derivation naturally returns UNKNOWN.

Tests use a new tests/common.force_app_state helper that pokes the
underlying signals directly; the production class no longer carries
test-only setters.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Fix App state drive to AppState.UNKNOWN

* Unify state mutation through _update_state

Previously, state-driving signal changes were spread across two helpers
(_set_operation_error, _emit_state_change(old_state)) and required each
caller to capture self.state before mutating a private field — leaking
implementation details to call sites and raising the "why am I emitting
the old state?" question pointed out in code review.

Replace both helpers with a single _update_state(*, container_state=,
operation_error=) entry point. Callers describe what changed via
keyword arguments (None leaves a signal untouched); the helper captures
the previous state, applies the updates, recomputes the derived state
and emits side effects if anything changed.

Diff against a tracked _last_state instead of a freshly derived
"current" state, so that an out-of-band mutation between updates does
not silently shift the comparison baseline. The concrete case is
App.uninstall: instance.remove() clears the docker meta mid-flow, which
would otherwise reshape the derivation (RUNNING with no healthcheck
becomes STARTED instead of STARTUP) and suppress the STARTUP transition
that resolves the start-wait task. As a side effect, the initial
UNKNOWN -> STOPPED transition on attach is also now reliably emitted.

Switch the uninstall path to ContainerState.UNKNOWN ("we know there is
no container") rather than the constructor sentinel None.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Cache app state instead of deriving on every read

Building on the previous commit, make App.state a plain read of a
cached _state field rather than re-deriving on every property access.
The derivation moves to _derive_state(), and _update_state() is the
sole place that recomputes and assigns _state, so the value consumers
read always matches what was last emitted to listeners.

This removes the _last_state bookkeeping introduced previously: with a
single cached value there is no longer a separate "derived now" vs
"last emitted" distinction to reconcile, and out-of-band mutations
(e.g. instance.remove() clearing _meta during uninstall) can no longer
silently shift what state returns between updates.

Call _update_state() at the end of load() so the cached state settles
once attach() has run. Image-only attaches do not fire a docker event,
so without this an installed app would stay in the constructor-default
UNKNOWN until first start; this also makes the initial transition on
attach observable to listeners.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Pass operation error to _derive_state instead of storing it

The two state-driving signals were not symmetric. _container_state is
genuinely persisted state ("the last thing docker told us") that
re-derivation legitimately reads across calls. _operation_error, on the
other hand, is a momentary "force ERROR for this transition" signal; the
persistence of an error condition already lives in the cached _state.

Storing it as an instance attribute implied a sticky cross-call behavior
that no call path actually exercised: every caller either set it
explicitly right before deriving (start/stop failures, container events)
or ran argless only at load time, where no failure has occurred.

Drop the _operation_error field and pass operation_error as a parameter
to _derive_state(), defaulting to False in _update_state(). A container
observation now supersedes a prior error implicitly via the default,
which lets the container-event and uninstall call sites drop their
explicit operation_error=False.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Settle load state synchronously from current_state

The argless _update_state() settle at the end of load() raced attach()'s
container-state event. attach() fires DOCKER_CONTAINER_STATE_CHANGE via
the bus, which schedules the container_state_changed listener as a task
rather than running it inline. In the deprecated-arch early-return path
there is no await between attach() and the settle, so the listener had
not run yet: _container_state was still None and the settle derived
STOPPED (instance attached) — emitting a transient UNKNOWN->STOPPED even
for a running container before the listener corrected it. The main path
only avoided this incidentally, by having awaits (check_image,
save_persist) in between for the listener to run.

Derive the load-time state synchronously from instance.current_state()
instead of relying on the asynchronously delivered event. current_state()
returns the real container state, or UNKNOWN when only an image is
present (which derives to STOPPED), so both paths settle correctly
without racing the event.

Add a regression test that loading a running container settles to
STARTED, and mock current_state() in the state-listener test which
relies on a clean UNKNOWN baseline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-01 19:50:06 +02:00
Mike DegatanoandCopilot 355396aeab Migrate addon/addons config paths and schema names to app/apps (#6865)
* Migrate config file and directory paths from addons to apps

- Rename addons.json -> apps.json (FILE_HASSIO_APPS constant)
- Rename addons/{core,data,local,git} -> apps/{core,data,local,git}
- Rename addon_configs -> app_configs

Backwards compatibility: on startup, Supervisor checks for legacy
paths and renames them if the new paths don't already exist.
- addons.json migration runs in AppManager.load_config (executor)
- Directory migrations run in bootstrap before initialize_system (executor)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Rename SCHEMA_ADDON(S)_* constants to SCHEMA_APP(S)_* in apps/validate.py

Update all references in supervisor and tests.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Fix remaining test references to legacy addons/* paths

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Opportunistic remove of addons dir since it should be empty post migration

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-05-27 09:47:32 +02:00
Stefan Agner 0f881d69fe Model client-state Apps* errors as APIError (#6856)
* apps: model client-state Apps* errors as APIError

Follow-up on #6739: with HassioError now logged and captured by Sentry
in api_process, a handful of Apps* exceptions raised from
AppManager.install/update/rebuild and AppModel._validate_availability
surfaced as "unexpected" 400s with a noisy log entry and a Sentry
event, even though they are all user/client-state errors (clicked
install on an already-installed add-on, "no update available", local
and store versions diverged, system architecture/machine/HA version
incompatible, etc.). The dominant offender is SUPERVISOR-1JVV
("No update available for app core_mosquitto", ~19k events / ~12k
users), but several siblings show the same shape.

Map these through properly so the API returns clean, structured 400s:

- Add modeled APIError subclasses in exceptions.py for the previously
  raw raises in apps/manager.py: AppAlreadyInstalledError,
  AppNotFoundError, AppNotInstalledError, AppNotInStoreError,
  AppNoUpdateAvailableError, AppRebuildVersionChangedError,
  AppRebuildImageBasedError. Each gets a stable error_key, a
  message_template, and an "addon" extra_field.
- Add APIError to AppNotSupportedArchitectureError,
  AppNotSupportedMachineTypeError and
  AppNotSupportedHomeAssistantVersionError so they behave the same as
  the other Apps* APIErrors instead of being treated as unexpected.
- Pass the app's display name (app.name from the add-on config)
  instead of the slug to extra_fields wherever an App or AppModel is
  available at the raise site, so users see "Mosquitto broker" rather
  than "core_mosquitto" in error messages. Slug is only used as a
  fallback when no app object exists (install of an unknown slug,
  update/rebuild of a slug that is not installed).
- Update raise sites in apps/manager.py and apps/model.py to use the
  new typed exceptions and the addon= keyword.

These are all runtime states users hit during normal interaction with
the apps UI, not Supervisor bugs worth paging on.

* apps: include slug alongside name in Apps* APIError extra_fields

Address review feedback from #6856: clients still need the slug to look
up additional add-on information (the name is for display only), and we
should be consistent about it across the Apps* errors touched by this
PR.

Every Apps* APIError raised with an App/AppModel available now carries
both `addon` (display name, used by the message_template) and `slug` in
extra_fields. Raise sites in apps/manager.py and apps/model.py pass
both. The two errors raised before an app object exists keep slug-only
extra_fields and use {slug} in their message:

- AppNotFoundError (install of an unknown slug)
- AppNotInstalledError (update/rebuild of a slug not in self.local)

Pre-existing Apps* APIErrors outside the scope of this PR
(AppUnknownError, AppConfigurationInvalidError, AppBootConfigCannot
ChangeError, AppNotRunningError, AppPortConflict, AppNotSupportedWrite
StdinError, AppBuild*) will be migrated in a follow-up.

* apps: introduce AppAPIError base for uniform addon/slug extra_fields

Address review follow-up on #6856: the addon/slug convention was
enforced only by hand-rolled __init__s, easy to drift on (forget slug,
use a different key, etc.). Promote it to a base class that owns the
shape of extra_fields for all App-related API errors.

- Add AppAPIError(AppsError, APIError). Its __init__ takes
  `app: AppModel | App | AppStore | str` and uniformly populates
  extra_fields with `addon` (display name) and `slug`. Pass a string
  when no app object exists; only `slug` is set in that case. Extra
  per-error fields flow through **extra_fields and merge with the
  defaults.
- Convert the new exceptions added in this PR
  (AppAlreadyInstalledError, AppNotFoundError, AppNotInstalledError,
  AppNotInStoreError, AppNoUpdateAvailableError,
  AppRebuildVersionChangedError) into thin subclasses that only declare
  error_key and message_template -- the __init__ is inherited.
- Migrate the AppNotSupported* errors (architecture, machine type, HA
  version) and AppRebuildImageBasedError to use AppAPIError too;
  their bespoke per-error fields go through **extra_fields. They keep
  inheriting AppNotSupportedError so `except AppNotSupportedError`
  callers (e.g., AppModel._available) still work; MRO routes __init__
  through AppAPIError.
- Update raise sites in apps/manager.py and apps/model.py to pass
  `app=<obj-or-slug>` instead of repeating `addon=...` and `slug=...`.

Pre-existing App* APIErrors outside this PR's scope
(AppUnknownError, AppConfigurationInvalidError,
AppBootConfigCannotChangeError, AppNotRunningError, AppPortConflict,
AppNotSupportedWriteStdinError, AppBuild*) will be migrated to
AppAPIError in a follow-up; the base class is in place for them.

* apps: tighten AppAPIError model per review

Address mdegat01's two follow-ups on #6856 (review approved as-is,
this is the cleanup):

- AppNotFoundError and AppNotInstalledError are raised before any
  App/AppModel object exists (unknown slug; not-installed slug). Pull
  them out of AppAPIError and inherit (AppsError, APIError) directly
  with a slug-only __init__ + {slug} message_template. Removes the
  conceptually-wrong str branch from AppAPIError.__init__: it now
  strictly requires an app-like object with .name and .slug.
- Restore per-class typed __init__ on AppNotSupportedArchitectureError,
  AppNotSupportedMachineTypeError and
  AppNotSupportedHomeAssistantVersionError so callers get an explicit
  signature for the bespoke architectures/machine_types/version params
  instead of dumping them through **extra_fields. Each override just
  delegates to AppAPIError.__init__, which keeps ownership of the
  addon/slug shape. The list-joining for architectures/machine_types
  moves back into the override (raise sites pass the raw list again).

AppRebuildImageBasedError takes no bespoke fields and stays a plain
AppAPIError subclass.
2026-05-22 16:54:59 +02:00
Stefan Agner d3028d7bfc Enable flake8-pyi, flake8-return, flake8-raise ruff rules (#6861)
Enable the PYI, RET and RSE ruff rule sets and fix the resulting
violations across the codebase. The pylint no-else-* checks
(RET505-508) are now disabled since ruff covers them.

The cleanups are mechanical:

- RSE102: drop empty parentheses from `raise Exception()` when no
  arguments are passed.
- RET505-508: drop `else` branches that follow a `return`, `raise`,
  `continue` or `break`, flattening control flow.
- RET502/504: add explicit return values and remove redundant
  assign-then-return patterns.
- PYI030/032/041: tidy up type annotations (collapse literal unions,
  use `object` for `__eq__`/`__ne__`, drop redundant numeric unions).
2026-05-22 11:04:49 +02:00
Stefan Agner ed91b18c4b tests: enable flake8-pytest-style (PT) ruff rules (#6857)
* tests: enable flake8-pytest-style (PT) ruff rules

Enable the `PT` ruff rule set and fix the resulting violations across the
test suite:

- PT006: pass parametrize argument names as tuples instead of a single
  comma-separated string.
- PT022: switch fixtures that have no teardown from `yield` to `return`
  so the lack of cleanup is obvious at a glance.
- PT011: add `match=` to broad `pytest.raises(ValueError)` blocks so the
  expected error is anchored to a specific message.
- PT012: hoist setup (patches, branching) out of `pytest.raises()`
  blocks so only the call that is expected to raise remains inside.
- PT013: replace `from pytest import X` with `import pytest` and access
  attributes via the module.
- PT015: replace `try/except` + `assert False` patterns with
  `pytest.raises(...)`.
- PT017: replace `assert` on exceptions inside `except` blocks with
  `pytest.raises(...) as exc_info` and assert on `exc_info.value`.

No behavioral changes to the tests; the full suite still passes.

* tests: address review feedback on PT ruff rule enablement

- Fix fixture return-type annotations after switching `yield` to `return`
  in tests/conftest.py: drop the `Generator[...]`/`AsyncGenerator[...]`
  wrapper for `dns_manager_service`, `supervisor_internet`, `websession`,
  and `mock_update_data` so the annotation matches what the fixture
  actually returns.
- Correct the return-type annotation of `fixture_ip6config_service` from
  `IP4ConfigService` to `IP6ConfigService`.
- Fix recurring "excepiton" typo in tests/utils/test_exception_helper.py.

* tests: verify backup cleanup on permission error

After `test_new_backup_permission_error` raises `BackupPermissionError`,
assert that no tarfile was left behind and `tmp_path` is empty. The
previous version only checked that the exception was raised, which
missed any regression where a partial tarfile would survive the failed
create.

* tests: rename DNS_GOOD_V6 to DNS_V6_UNSUPPORTED

The constant was named "good" but its tests assert that the URLs are
rejected by the DNS validator. The IPv6 URLs are well-formed but
currently rejected because IPv6 doesn't work with the Docker network
(see `dns_url` in supervisor/validate.py). Rename the constant and the
related test to make the intent obvious.
2026-05-20 22:17:54 +02:00
Stefan AgnerandClaude Opus 4.7 267fc6cd71 mounts: make is_mounted honest about server reachability (#6838)
* mounts: use softerr for NFS instead of soft

Switch the NFS mount option from `soft` to `softerr`. For HAOS-style
supervisor mounts (media, share, backup — not the root filesystem) the
error semantics matter:

* `softerr` returns `ETIMEDOUT` on timeout instead of `EIO` (`soft`).
  `EIO` is indistinguishable from "the disk is dying"; tools like
  SQLite, restic, rsync, ffmpeg tend to treat it as a hard storage
  failure (mark database corrupt, abort backup with a hard error,
  etc.). `ETIMEDOUT` is unambiguously "the network/server is gone,
  transient" and is more commonly handled as retry-later. Supervisor
  can also surface a clear "server unreachable" notification rather
  than a generic I/O error.

* `softerr` was added in kernel 5.10 precisely to give the fail-fast
  behavior of `soft` with a distinct errno so well-behaved apps can
  do the right thing.

* For writes-must-not-be-lost use cases (databases, paid storage,
  evidence-grade logging) one would want `hard,intr` and a different
  recovery story. HAOS NFS mounts are not that — they're add-on
  storage where "the share went offline, try again later" is the
  correct user-visible behavior.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* mounts: make is_mounted honest about server reachability

systemd's "active/mounted" state is derived from /proc/self/mountinfo
and doesn't reflect whether the backing server actually answers. For
CIFS in particular, smb3_reconfigure never contacts the server, so a
reload of a dead share returns active/mounted with no recovery
attempted. PR #4882 added a Path.is_mount() cross-check to catch this,
but Path.is_mount() relies on os.stat() of the mountpoint root — and
both NFS (softreval) and CIFS (cached root inode attrs) can serve that
stat from local state without going to the wire, so it lies in exactly
the dead-share scenario it was meant to detect.

The visible failure: an API reload of a backup mount whose server had
gone away "succeeded", supervisor then kicked off sys_backups.reload()
against the dead share, and the executor fanned out hundreds of
"Task exception was never retrieved" OSError(112) tracebacks as each
backup tarball's stat() parked in the kernel and failed.

Replace the local-state checks with a statvfs() probe in
NetworkMount.is_mounted(). os.statvfs() returns per-filesystem data
(free blocks, total blocks) that has no client-side cache in either
kernel — neither cifs_statfs() nor nfs_statfs() has an early-return on
cache freshness; both build and send a real FSSTAT / QUERY_FS_INFO
request. So the kernel either reaches the server or gives up with
ETIMEDOUT / EHOSTDOWN / ECONNABORTED. is_mounted() now reflects actual
reachability, finally fulfilling PR #4882's stated intent.

With is_mounted() honest, the reload/restart machinery falls out
naturally: Mount.reload()'s fast path is just `if await self.is_mounted()`,
the post-reload check is the same call, and the "reload succeeded
systemd-wise but probe failed" branch collapses into the existing
"not mounted after reload, try restart" branch. mount(), _restart()
and update() already call is_mounted(); they now get the probe for
free.

To make the probe meaningful, the network mount option strings get
explicit kernel-side timeouts:

* NFS switches from `soft` to `softerr` so timeouts surface as
  ETIMEDOUT rather than EIO. EIO is indistinguishable from
  "disk dying" and gets misinterpreted by SQLite/restic/rsync as a
  hard storage failure; ETIMEDOUT is unambiguously transient.
* CIFS gains `soft,echo_interval=10,retrans=0`, giving a ~30s
  per-operation detection budget (3 x echo_interval since last server
  response) that matches the NFS budget from `timeo=100,retrans=2`.
  Both protocols now fail bounded operations in roughly the same time.

The probe is intentionally not wrapped in an asyncio timeout: the
kernel-side bound is authoritative, and an asyncio timeout would only
orphan the executor thread without unblocking the syscall. The probe
emits debug logs with timing so the ~30s syscall wait on a dead share
is visible in LOGLEVEL=debug traces instead of appearing as a hang.

Tests:

* mock_is_mount fixture extended to also patch os.statvfs so existing
  tests that rely on a healthy mount don't need to know about the
  probe.
* New manager test split into _healthy_skips_systemd (probe succeeds,
  fast path, no systemd call) and _probe_failure_triggers_systemd_reload
  (probe fails, escalation runs). API reload test covers both paths.
* Existing tests that simulated "mount is down" via
  mock_is_mount.return_value=False updated to simulate probe failure
  via OSError(EHOSTDOWN), since is_mount is no longer the signal.

* mounts: drive systemd job waits via JobRemoved instead of state polling

`_update_state_await` had a race that has been present since network
mounts were introduced (#4269) and survived every subsequent rewrite
(#4733, c75f36305): the wait function reads the unit's ActiveState
*after* the systemd dispatch returns, then matches against a target
set that almost always includes the pre-call state (ACTIVE). If
systemd has not yet started dispatching the queued job — and for
fast operations like CIFS reload (smb3_reconfigure completes in
milliseconds with no server contact) it routinely has not — the wait
matches immediately on the pre-call state and returns "done" before
the operation has actually begun. The race was invisible until the
probe commit started doing honest network work after the wait
returned, at which point we spent a full ~30s NFS probe wait on a
mount that systemd was still in the middle of restarting.

Switch the job-dispatching paths (mount/unmount/reload/restart) to
systemd's `JobRemoved` Manager signal. The pattern:

    async with sys_dbus.systemd.job_removed() as jobs:
        job_path = await sys_dbus.systemd.restart_unit(...)
        await jobs.wait_for_job(job_path)

is structurally race-free: subscription is set up before dispatch,
the job_path is allocated synchronously inside the dispatch call,
and any JobRemoved for that path after subscription is queued. Even
for operations that complete faster than we can process the dispatch
return value (the CIFS reload case), the signal cannot be missed.

`_update_state_await` is kept for `load()` — that path observes
existing state without having dispatched a job, so the
PropertiesChanged-driven wait is the right primitive there.

* mounts: verify the path is a mount point before trusting statvfs

The userspace probe trusted statvfs() alone to declare a network mount
healthy. statvfs is uncacheable client-side for both NFS and CIFS, so
on a real mount it forces a wire RPC — but if the mount is gone from
the kernel mount table (e.g., after a restart cycle whose umount
succeeded but whose mount step failed, leaving the unit ACTIVE in
systemd's view), statvfs() operates on a plain directory on the
underlying root filesystem and happily returns the root fs's stats.
The mount is reported healthy when in fact it doesn't exist.

Add a pre-check using `Path.is_mount` — a parent-vs-path `st_dev`
comparison via stat() — to detect "this path is not actually a mount
point" before issuing statvfs. For the ghost-mount case (path on the
root fs) both stats are local and return immediately. For a real
mount the path-stat may cross into the filesystem driver, but the
result is correct in either case and the statvfs that follows
catches server-unreachable mounts that is_mount can't.

The full is_mounted contract is now: ACTIVE per systemd, present as
a mount point in our namespace, and server-reachable per statvfs.

* mounts: skip wasted probes when systemd already reported job failed

JobRemoved tells us whether the systemd job completed successfully —
we just weren't using it. The previous reload/restart paths ran their
post-job probe unconditionally, including the case where systemd had
already told us the operation failed. On a dead NFS share the kernel
is in transport-reconnect churn for tens of seconds after a killed
mount helper exits, so that probe takes 90+ seconds — and confirms
exactly what systemd already said. Plain wasted time.

* `Mount.reload()` now uses the systemd job result. On "done" we
  still probe (CIFS reload is local-only — smb3_reconfigure returns
  "done" even against a dead server, so the probe is the only check
  that catches the lying-CIFS case). On anything else (failed,
  timeout, canceled, dependency, skipped, or our own timeout
  returning None) we escalate directly to `_restart()` without
  probing.
* `Mount._restart()` skips the probe entirely on non-"done" results.
  The mount is definitively not active in that case and the probe
  would just spend another 30-90s on a dead share.

Also bump the JobRemoved wait timeout from UPDATE_STATE_TIMEOUT (40s,
sized for a single-helper invocation) to a new SYSTEMD_JOB_TIMEOUT
(90s). RestartUnit runs as stop + start, each bounded by the unit's
TimeoutSec (35s from #6834), so the worst case is ~70s plus systemd
queue dispatch. The previous 40s budget caused supervisor to time
out exactly one second before JobRemoved fired in the observed
dead-NFS restart cycle. UPDATE_STATE_TIMEOUT is kept at 40s for the
PropertiesChanged-driven wait in `load()`, where the layered timeout
invariant from #6827 still applies.

* dbus: filter signals at the wrapper level, drop JobRemovedSignal class

Mike's review on #6838: instead of carrying a bespoke JobRemovedSignal
context-manager class, push the filtering concept into the generic
DBusSignalWrapper. Any caller that subscribes to a broadcast signal
but only cares about specific payloads can pass a predicate.

* DBus.signal() gains an optional `message_filter: Callable[..., bool]`.
* DBusSignalWrapper.wait_for_signal() loops past messages where the
  filter returns False, returning the next match.
* JobRemovedSignal goes away. supervisor/dbus/systemd.py exposes a
  small factory `job_removed_filter(get_job_path)` that returns the
  matching predicate; the getter indirection lets callers subscribe
  before the dispatch returns the job path (which is required to be
  race-free — see the previous commit).

Mount._run_systemd_job() switches to:

    job_path: str | None = None
    async with systemd.connected_dbus.signal(
        DBUS_SIGNAL_SYSTEMD_JOB_REMOVED,
        job_removed_filter(lambda: job_path),
    ) as signal:
        job_path = await dispatch
        _id, _path, _unit, result = await signal.wait_for_signal()

The behavior is identical to the previous JobRemovedSignal-based
implementation. The signal queue still buffers messages received
after AddMatch is installed, so subscribing before dispatch keeps
the race-free guarantee.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 09:29:53 +02:00
Mike DegatanoandCopilot f8880a72be Rename addon/addons to app/apps in filenames and imports (#6837)
* Rename addon/addons to app/apps in filenames and imports

Continues the addon→app terminology migration (#6786).
Renames all source files, test files, fixture files, and
directories that contained 'addon'/'addons' in their names,
and updates all imports accordingly.

Resolution check files in supervisor/resolution/checks/ that were
renamed override the slug property to preserve the existing API
contract (slugs are exposed via the resolution info API and used
to run checks by name).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Rename add-on.json fixture

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-05-13 20:55:46 +02:00