Commit Graph
25 Commits
Author SHA1 Message Date
Stefan AgnerandClaude Fable 5 324d17e68b Use systemd automount units for network mounts (#7167)
* dbus: support aux units in start_transient_unit

Extend `Systemd.start_transient_unit` to accept the `aux` parameter
(`a(sa(sv))`), which has been hardcoded to `[]` since the wrapper was
introduced. Aux entries take the form `(unit_name, properties)` and
let systemd create multiple transient units atomically — most usefully
a `.mount` and its `.automount` companion in one D-Bus call.

Add the unit-property constants we'll need to drive that:

- `DBUS_ATTR_LAZY_UNMOUNT` ("LazyUnmount"): set on the `.mount` so
  systemd umounts with MNT_DETACH. Pairs with `softerr`/`soft` to make
  stop/restart sequences reliable even when the server is gone.
- `DBUS_ATTR_TIMEOUT_IDLE_USEC` ("TimeoutIdleUSec"): set on the
  `.automount` to control how long the mount stays around after the
  last access before autofs expires it.
- `DBUS_ATTR_WHERE` ("Where"): mount-point property; required on the
  `.automount` unit (and useful explicitly on `.mount` units too).

No call sites change in this commit — pure plumbing.

* mounts: pair each network .mount with a .automount companion

Network mounts now get a transient `.automount` unit created
atomically alongside their `.mount`, via the aux parameter on
`StartTransientUnit`. The systemd-managed autofs trigger handles
activation lazily: the path exists in the VFS even before the
underlying network mount runs, and the first access that crosses the
trigger fires `mount.cifs`/`mount.nfs` on demand.

Why this matters:

- PID 1's path-walks (chase, daemon-reload, generator scans) stop
  crossing dead NFS/CIFS lookups — autofs returns from kernel memory
  rather than entering the network filesystem. The class of failure
  fixed at the systemd-timeout layer in #6834 stops being reachable
  in the first place.
- Failed accesses fail fast, bounded by the `.mount`'s `TimeoutSec=`
  rather than hanging forever.
- Background reconnect (NFSv4 state-manager kthread, CIFS
  delayed_work) recovers transparently when the server returns; no
  remount needed.

The `.automount` carries `TimeoutIdleUSec=5min` so kernels can expire
the mount after inactivity and re-trigger on the next access. The
companion `.mount` carries `LazyUnmount=true` (MNT_DETACH on stop),
so umount returns immediately even when the server is unreachable;
existing fds drain in their soft/softerr timeout regime instead of
pinning the umount syscall.

`Mount.unmount()` now stops the `.automount` first so the autofs
trigger can't re-fire the underlying `.mount` during cleanup. The
automount stop is best-effort — if the unit is gone or refuses to
stop, we log and proceed to the `.mount` stop, which remains
authoritative.

Bind mounts opt out via a `creates_automount` class flag — they have
no server to wait on and the lazy semantics would only add
indirection. Network mounts (CIFS/NFS) opt in.

This is wiring only; the periodic mount reload, the reload/restart
escalation, and the inner bind-mount layer for media/share usage are
still in place. The next commit removes them now that automount
makes them unnecessary.

* mounts: drop bind layer, periodic reload, and reload/restart machinery

With the kernel's autofs trigger handling lazy activation and
transparent reconnect, almost all of the supervisor-side mount
choreography becomes unnecessary. Strip it out.

What goes away:

* The inner bind-mount layer. Media/share mounts used to live at
  `path_extern_mounts/{name}` and have a second `.mount` unit bind
  them into `path_extern_media/{name}` (resp. `share`). The
  `.automount` companion can sit directly at the container-facing
  path, and the parent-dir RSLAVE bind into add-on containers
  surfaces the autofs trigger the same way. `NetworkMount.where`
  now switches on usage:
    - MEDIA → `path_extern_media/{name}`
    - SHARE → `path_extern_share/{name}`
    - BACKUP → `path_extern_mounts/{name}` (unchanged)
* `BindMount`, `BoundMount`, `MountManager._bound_mounts`, the
  `_bind_mount`/`_bind_media`/`_bind_share` helpers, and the entire
  emergency-fallback dance (`path_emergency`). The empty-read-only
  dir trick was a workaround for the harder PID 1 wedge problem,
  which autofs solves at the root. Failed shares now surface as
  ETIMEDOUT/EHOSTDOWN on access plus a resolution issue.
* `MountManager.reload()` and the 15-minute `RUN_RELOAD_MOUNTS`
  periodic task. autofs re-activates on access; we don't need to
  poll, and not polling means we don't randomly trip stale-FH /
  softreval / dead-server edge cases on a timer.
* `Mount.reload()` and `Mount._restart()`. The reload→restart
  escalation existed to make `is_mounted` honest in the face of
  systemd's local-only state; now `is_mounted` IS honest (probe-
  based), and any recovery the kernel knows how to do happens
  inside autofs without our involvement.
* The `RELOADING` safety net introduced by #6834. That whole class
  of PID 1 wedge stops being reachable when path lookups don't
  cross dead network mounts.

What changes shape:

* `NetworkMount.is_mounted()` no longer gates on systemd's
  `ActiveState` before probing. The `.mount` unit is dormant
  whenever autofs hasn't recently triggered it, so systemd state
  is meaningless as a health signal. The probe (now via
  `statvfs("/path/.")` so the trailing dot forces `LOOKUP_DIRECTORY`
  and triggers autofs) is the source of truth, and after it returns
  we set `self._state` to ACTIVE/INACTIVE so the API reports
  reachability rather than autofs idle state.
* `MountManager.reload_mount()` becomes probe-only — no systemd
  reload/restart calls. The user-facing semantics: "tell me if this
  mount is reachable right now, and refresh the resolution issue
  accordingly." If the mount is dead, autofs will re-trigger it on
  the next consumer access; the supervisor doesn't need to force it.
* `Mount.load()` keeps `_update_state_await` for the no-job-dispatched
  case where a previous supervisor left a unit in `activating`, then
  probes once via `is_mounted` so state reflects reachability rather
  than the lazy-mount idle state.
* `BackupManager` no longer walks `bound_mounts`. Backups skip
  network-mount subdirectories during folder archive (so we don't
  recurse into the share), and unmount/re-mount them around folder
  restores (so writes target the local mount-point dir, not the
  remote share).

Behavior tradeoffs accepted:

* Dead media/share access now ~30s ETIMEDOUT instead of empty
  read-only dir. The empty-dir was a workaround for the PID 1
  wedge that autofs eliminates at the root; if specific add-ons
  rely on the old behavior we can revisit with a 2-stage automount
  whose fallback target is an emergency dir.
* Resolution-issue lag: failed mounts no longer surface within
  15 min on their own. They appear when the user probes via the
  API, when `BackupManager.reload()` walks the location, or when
  load-time activation fails.

* tests: adapt bind-layer-era tests after rebase

Main gained tests for the bind layer after this branch was written:
the #7013 rebind regression test and the #7072 bind-step rollback test
cover machinery that no longer exists, and the healthy-reload test
asserted the unconditional rebind. Drop the first two and reduce the
third to asserting that a healthy probe performs no systemd operations
at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: rearchitect automount setup from systemd/kernel review

A code-level review of systemd's automount implementation (verified
against v254.13, the version HAOS ships) and the kernel autofs/VFS
plumbing surfaced three correctness-critical flaws in the automount
design, plus several hardening gaps. See automount-rearchitecture.md
for the full analysis.

- Make the .automount the primary transient unit with the .mount as
  aux, mirroring `systemd-mount --automount=yes`. Aux units get no
  start job: with the .mount as primary the trigger was never armed
  and the design silently degraded to eager mounting.

- Set StartLimitIntervalUSec=0 on the .mount. With the default start
  rate limit (5 starts/10s, counting successful starts) a fast-failing
  mount plus any polling consumer trips the limit within seconds, and
  systemd then detaches the autofs trigger entirely
  (AUTOMOUNT_FAILURE_MOUNT_START_LIMIT_HIT) — the path silently
  becomes a plain writable local directory.

- Drop the 5-minute TimeoutIdleUSec (default 0 = never expire).
  Kernel idle expiry decides busyness via may_umount_tree(), which
  only counts the init-namespace mount instance — open files held by
  container processes are invisible, so expiry would unmount shares
  under actively writing add-ons.

- Probe with a plain statvfs: the statfs syscall walks with
  LOOKUP_AUTOMOUNT and triggers by itself; the trailing-dot trick is
  unnecessary. Classify ELOOP as a mount-propagation
  misconfiguration in the probe error handling.

- Re-arm the trigger from reload_mount(): if the .automount unit is
  failed or gone (e.g. autofs unmounted out-of-band), reset failure
  state and re-create the pair before probing.

- Tear down legacy eager-mount units during load(). The old design's
  bind unit for media/share occupies the exact unit name the network
  .mount uses now; on a warm upgrade adoption would mistake the old
  bind mount for the network mount and leak the legacy data mount.

- Honor the systemd job result when stopping the .mount during
  unmount and reset failure state on both units afterwards so dead
  transient units get garbage-collected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: adapt local data repair to the autofs design

Port of the mount-target-not-empty repair (#7089) on top of the
automount rearchitecture:

- relocate_local_data() handles a single target directory — the mount
  sits directly at its container-facing path, the bind layer and with
  it the multi-directory case are gone
- a repair_trigger() failure caused by blocking local data raises the
  mount failed issue with the move_local_data suggestion instead of
  the plain variant
- local data at load is detected via the mount unit itself; the
  bind-layer detection test is rewritten accordingly

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docker: use rslave propagation for execute_command share mount

The temporary container used for execute_command (e.g. core config
check) mounted /share without a propagation mode, unlike every other
share/media mount. With eager mounts this only meant missing mounts
made after container start; with lazy automount activation the shares
are routinely mounted after start, and accessing a not-yet-activated
automount from a private mount namespace fails with ELOOP.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: harden automount lifecycle from review findings

Address the findings of an adversarial review of the automount
rearchitecture:

- load() no longer adopts a dead automount trigger: a failed or
  stopped .automount leaves the path a plain writable directory — the
  silent local-write degradation this design must prevent. Tear down
  and re-arm instead. An adopted pair whose probe fails now raises so
  the manager surfaces the mount failed issue, same as a fresh mount.
- unmount() resolves the .mount unit only after stopping the
  .automount (the lazy detach can garbage-collect the transient unit,
  invalidating an earlier proxy), raises instead of warning when the
  automount stop errors, and checks the stop job result so a failed
  stop cannot leave an armed trigger behind a "successful" removal.
- repair_trigger() fully unmounts before re-arming: with the share
  still attached, arming fails — or worse, mount() misreads the
  mounted share's contents as blocking local data.
- Legacy unit teardown honors the stop job result for the same reason.
- reload_mount() escalates once to re-creating the unit pair when an
  established mount is unreachable: a permanently dead session (e.g.
  replaced server) keeps the path mounted, so the trigger can never
  re-fire and kernel reconnection never succeeds. Teardown is safe now
  (lazy unmount, no PID 1 path walks), unlike the removed reload →
  restart escalation of the eager design.
- Reinstate the periodic mounts task as a probe-based reconcile: re-arm
  dead triggers, refresh the reachability state reported by the API and
  used by backup locations, and sync the mount failed issue in both
  directions. No reload or restart of established mounts.
- Folder restore no longer fails after a successful restore when
  re-mounting nested mounts cannot verify an unreachable server — the
  trigger is armed, the share recovers on next access.
- Drop the now-unused Mount.update(), deduplicate unit name escaping,
  hoist the backup exclusion set out of the per-file filter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* tests: expect rslave propagation in core check container

The config check assertion missed the update for the rslave
propagation on the execute_command share mount.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* tests: cover reconcile trigger repair failure paths

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: arm automount before stopping the legacy data mount

Address review: the legacy teardown stopped the media/share bind unit
and the eager data mount in sequence, both strictly. When the bind
stop succeeded but the data mount stop failed — its unmount can time
out against an unreachable server, legacy units have no LazyUnmount —
load() raised with the container-facing path left a plain writable
directory and no trigger armed: the pollution mode this design is
meant to prevent. Reordering the stops would not help either, as
systemd stops the bind first anyway as a dependent of the data mount
(the implicit Requires= from RequiresMountsFor= on the bind's What=).

Instead, strictly stop only the path-conflicting unit (a failure
leaves the path covered, which is safe and retryable), arm the
automount right after, and stop the conflict-free legacy data mount
best-effort last — a failed stop logs a warning and leaves an orphaned
mount for the next Supervisor restart or a host reboot, with nothing
writable exposed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: discard dead session instead of re-creating units on reload

Address review: re-creating the unit pair on reload of an unreachable
established mount exposed the target as a plain writable directory
between the lazy unmount and arming the replacement — a continuously
writing add-on could block the re-arm or slip writes under the new
mount in the check-to-arm window.

Stop only the .mount unit instead, keeping the .automount armed:
systemd re-installs the autofs trigger over the path — the same
mechanism idle expiry uses, with the automount's Triggers= reference
keeping the transient .mount definition alive — so the path is never
locally writable. The re-probe then mounts fresh through the trigger,
which establishes a new session and thereby covers the permanently
dead session case (e.g. a replaced server) the escalation exists for.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: surface arming failures, keep restore teardown re-armed

Address review findings on the arming/re-arming error paths:

- mount() checks the StartTransientUnit job result: a failed start job
  means the trigger never armed and the path is a plain writable
  directory — a hard MountError before the probe, so it cannot be
  mistaken for the armed-but-unreachable MountActivationError.
- Folder restore includes the nested-mount teardown in the try block:
  an unmount failing halfway (trigger disarmed, share still attached)
  previously exited before any re-arm, leaving the path unprotected.
  The finally re-arms via repair_trigger(), which no-ops on a still
  armed trigger and handles partially torn-down pairs.
- The post-restore re-arm suppresses only MountActivationError (armed,
  recovers on next access). Any other failure — arming failed, local
  data blocking the target — left the path unprotected and surfaces.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: fold the two legacy unit teardowns into one helper

The eager-mount-era cleanup queried the .mount unit a second time
although load() had just fetched it, and the strict teardown of the
unit occupying the automount's path and the best-effort teardown of
the mount at the mounts data directory were near identical. Reuse the
unit from load() and give the shared helper a strict flag, which drops
a D-Bus round trip per mount on every load.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: re-arm the trigger when unmount cannot stop the mount

Stopping the automount detaches the whole stack at the path, so an
unmount that then fails to stop the .mount leaves a plain writable
directory behind. Arm a fresh pair before raising, best effort — the
local data repair covers what lands there if that fails too.

The automount stop itself needs no such handling: systemd's
automount_stop() enters dead synchronously and the detach it performs
(MNT_DETACH, no server contact) only logs its errors, so a failure
there means systemd could not be reached at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: raise translatable errors for mount operations

Errors from setting up, unmounting and reloading a mount reach users
through the API, so describe what failed in a translatable message and
leave the systemd job result or D-Bus error to the log.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: remove the emergency folder of the eager-mount design

The read-only fallback directory has no place in the automount design.
Remove what is left of it on existing installations, next to the legacy
addons directory cleanup, keeping it if it holds anything but the empty
mount points.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: trim comments to the behavior they describe

Drop the narration of what changed and why from comments and
docstrings, keeping the notes that record why a tempting alternative
does not work. Two comments still described the reload to restart
escalation this branch removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: correct stale claims in manager and reload test

The manager docstring credited the kernel with idle expiry, which is
deliberately disabled, and claimed no polling although the periodic
reconcile probes every 15 minutes. The reload test claimed systemd is
never contacted while the escalation stops the .mount unit; only reload
and restart of the unit are avoided.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: arm the trigger alone if re-creating the pair is rejected

Address review: the re-arm after a failed unmount submits the .mount as
an aux unit, but transient creation requires a pristine unit and the
definition is still loaded precisely when stopping it is what failed.
Fall back to creating the .automount on its own, which covers the path
and fires the surviving mount definition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: never arm the trigger over local data after a failed unmount

Address review: the fallback caught the target validation errors too,
so data written while the path was uncovered would end up beneath a
fresh trigger. Reconciliation would then find a healthy mount and the
repair skips mount points, leaving the data hidden for good. Leave the
path untouched instead, so the next reconcile offers to move it away.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-31 20:30:23 +02:00
Stefan AgnerandClaude Fable 5 69c5a53864 Surface fixup failures to the caller applying a suggestion (#7145)
* Surface fixup failures to the caller applying a suggestion

Applying a suggestion whose fixup failed reported success: the fixup
base swallowed ResolutionFixupError, the API returned OK, and the
repair flow in Home Assistant completed and removed the repair from
the UI while the issue persisted in Supervisor.

Let the error propagate from the fixup instead. The autofix loop
already handles and logs per-fixup errors, bus-event triggered fixups
now get the same treatment in the event callback, and a user-applied
suggestion surfaces the failure as an API error so clients can show
that the fix did not apply. ResolutionFixupError becomes an APIError
with a translatable error key so it is reported as a client-visible
error instead of an unexpected one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Drop ResolutionFixupError, let fixup failures bubble

Per review: the generic wrapper existed for a time when suggestions
were only applied by autofix and the one job was separating fixup
failures from real bugs for Sentry. The suggestion API is in regular
use now and the wrapper actively hurts it — well-defined errors from
the underlying operations were caught and replaced with a generic
message.

Remove the exception type entirely and let the original errors reach
the caller. The autofix loop and the bus-event fixup path treat any
HassioError as an environmental/config failure: log and continue
without Sentry capture (the raise site reports to Sentry where
warranted); everything else is still captured. Direct raises in the
data disk fixups become HassOSDataDiskError, the could-not-start check
in the app start fixup becomes AppsError. ResolutionFixupJobError now
derives from ResolutionError and JobException.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Share unattended fixup error handling between autofix and bus events

Address review: the bus event listener only caught HassioError, so a
Supervisor bug in a bus-triggered fixup ended as an unretrieved task
exception without Sentry report, unlike the same bug during autofix.
Both unattended paths now use a common apply_fixup_safely() helper:
log HassioError (environmental, reported at the raise site if
warranted), capture everything else.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 23:14:44 +02:00
Stefan AgnerandClaude Fable 5 ccdd8c1ef4 Add repair to move local data blocking a mount target (#7089)
* Add repair to move local data blocking a mount target

When an add-on writes into a media/share directory while its network
mount is not in place (#7037), the local data blocks re-creating the
mount: mounting over a non-empty directory is refused. Until now this
failed silently at Supervisor startup — the bind mounts were created as
fire-and-forget tasks — and the only way out was to remove the data
manually over SSH/Samba and re-create the mount via the API.

Surface the condition as a new mount_target_not_empty issue and offer a
move_local_data suggestion. The fixup moves the blocking data to a
<name>_local_recovery folder in a user-accessible location — media or
share for bind mount targets, local backup storage for backup mounts
(their data mount directory is not reachable for users) — then reloads
the mount. Nothing is deleted; users can inspect and clean up the
recovered data via the media browser or the share and backup folders.

Bind mount failures during load are now awaited and routed into
resolution issues instead of being swallowed as fire-and-forget tasks;
bind failures other than blocking local data create the existing
mount_failed issue. A successful mount reload dismisses a stale local
data issue.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Attach move local data suggestion to the mount failed issue

Review feedback on the repair: rather than introducing a separate
mount_target_not_empty issue type, keep a single mount failed issue
per mount and offer moving the blocking data as an additional
suggestion alongside reload and remove. Reload stays available for
users who prefer to clear the data themselves, and at most one repair
exists per mount. Adding is idempotent, so an already-raised mount
failed issue just gains the extra suggestion.

When re-creating the bind mount after a successful reload fails on
blocking local data, the mount failed issue is re-added together with
the move suggestion, since the reload already dismissed it.

This also resolves the reviewer note about not-a-directory conflicts
being reported under a not-empty issue type: the issue type no longer
encodes the filesystem detail, while the error messages keep the
distinction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Keep an empty directory in place after relocating local data

If remounting fails after the local data was moved aside (e.g. the
server is unreachable at that moment), the renamed directory left
nothing behind: media/share consumers saw the folder disappear
entirely. Recreate an empty directory right after the rename so the
path stays present regardless of whether the remount succeeds.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Drop move local data suggestion once the data was moved

When the remount after relocating local data fails (e.g. the server is
unreachable at that moment), the mount failed issue stays — but the
move suggestion stayed with it, offering to move data that is no
longer in the way. Dismiss the suggestion after the relocation step so
only reload and remove remain for the leftover failure. Detection
re-adds the move suggestion if local data blocks the target again.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Filter Core-facing suggestions by minimum Core version

The fix flow translations for a new suggestion ship with a Core
release. Older Core frontends render an unknown suggestion as an
empty, unlabeled menu entry in the repair fix flow. Filter such
suggestions from Core-facing output — the issue events sent over the
websocket and the resolution API responses when the caller is Home
Assistant — until the connected Core is new enough. Other API
consumers like the CLI always see the full suggestion list.

The move_local_data suggestion requires Core 2026.9.0b0, the release
its fix flow translations are targeted at.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Make fixup failure tests independent of error propagation behavior

Suppress a potential ResolutionFixupError from the failing fixup calls
so the tests pass both while fixup errors are swallowed and once they
propagate to the caller (#7150).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address review: one recovery folder, check OSError for known issues

Move all local data blocking a mount into a single recovery folder so
the user finds it as one fix: when more than one directory holds data,
later ones become subfolders named after their parent directory
instead of numbered sibling folders.

Also run OSError from the relocation through check_oserror to pick up
known filesystem issues like corruption (bad message).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Only filter Core-facing suggestions on v1 surfaces

Per review: no Core version predating the repair suggestion filtering
in its own fix flow (home-assistant/core#179540) supports the v2 API,
so the Supervisor-side compatibility filter is only needed where old
Core versions actually look. Filter the v1 resolution endpoints and
the legacy websocket issue payloads; the v2 endpoints and v2 event
payloads always carry the full suggestion list. The suggestions for
issue endpoint gets a v1 handler for this.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 23:40:47 +02:00
c01efa05a5 Rename resolution addon issue types and check slugs to app-based naming (#7098)
* Rename resolution addon issue types and check slugs to app-based naming

Renames the following for consistency with the new apps terminology:
- Issue types: deprecated_addon -> deprecated_app,
  deprecated_arch_addon -> deprecated_arch_app,
  detached_addon_missing -> detached_app_missing,
  detached_addon_removed -> detached_app_removed
- Check slugs: addon_pwned -> app_pwned,
  deprecated_addon -> deprecated_app,
  deprecated_arch_addon -> deprecated_arch_app,
  detached_addon_missing -> detached_app_missing,
  detached_addon_removed -> detached_app_removed

Full backward compatibility is maintained:
- REST API V1 (root) returns legacy names and accepts legacy slugs
- REST API V2 (/v2) returns and accepts new names only
- WebSocket events use legacy names unless SUPERVISOR_WEBSOCKET_V2_API
  feature flag is enabled
- resolution.json files with legacy check slugs are automatically
  migrated to new names on load

Fixes #7029

* Apply suggestions from code review

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* Address review nits for resolution compatibility maps

* Update supervisor/resolution/const.py

---------

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Stefan Agner <stefan@agner.ch>
2026-08-04 10:57:42 +02:00
Stefan Agner f8dbafe0bb Drop redundant @pytest.mark.asyncio decorators (#6795)
The pytest config sets ``asyncio_mode = "auto"``, which already
auto-marks every ``async def test_*`` as a coroutine test. The 38
``@pytest.mark.asyncio`` decorators sprinkled across the suite were
no-ops kept around from before that flag was set. Remove them along
with the now-unused ``import pytest`` lines they were the only
consumer of.

Pure mechanical cleanup; no test behavior changes.
2026-05-04 14:48:18 +02:00
bc24fb5449 Refactor API registration to support v1/v2 via shared methods (#6769)
* Refactor API registration to support v1/v2 via shared methods

- Add AppVersion StrEnum (V1, V2) to supervisor/api/const.py
- Replace self.v2_app with self._v2_app and expose a versions property
  (dict[AppVersion, web.Application]) computed dynamically so that test
  fixtures reassigning self.webapp are automatically reflected in V1
- All _register_* methods now accept a required app: web.Application
  parameter; version-specific routes are gated with
  "if app is self.versions[AppVersion.V1/V2]:"
- load() loops over enabled_versions (V1 always, V2 when feature-flagged)
  and calls each registration method once per version, no duplication
- Static resources are registered before webapp.add_subapp() to avoid
  registering into a frozen router
- add_subapp uses self.webapp directly for readability
- Fold _register_v2_apps/_register_v2_backups/_register_v2_store into
  their respective unified methods; remove the now-defunct _register_v2_*
  helpers and the _api_apps/_api_backups/_api_store instance vars
- _register_proxy and _register_ingress updated to accept app; legacy
  /homeassistant/* proxy routes gated behind V1 conditional

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Add dual v1/v2 parametrization to API tests

All 163 tests across 17 API modules that register identically on both
v1 and v2 now run against both versions via api_client_with_prefix.

- tests/api/conftest.py: advanced_logs_tester switched to
  api_client_with_prefix so log-endpoint tests are auto-parametrized;
  accepts optional v2_path_prefix kwarg for paths that differ by version
- tests/api/test_{auth,discovery,dns,docker,hardware,host,ingress,
  jobs,mounts,network,os,resolution,security,services,supervisor}.py:
  api_client -> api_client_with_prefix with path prefix unpacking
- supervisor/api/__init__.py: _register_panel() moved outside the
  version loop -- frontend static assets are V1-only
- tests/api/test_panel.py: kept on plain api_client (V1-only)

Tests intentionally kept V1-only:
- auth/discovery: use indirect api_client parametrize for addon context
- homeassistant: all tests call legacy /homeassistant/* paths (V1-only)
- jobs (4 tests): inner @Job-decorated classes register names into a
  module-level set; re-running the same test raises RuntimeError

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Extend dual v1/v2 parametrization to homeassistant and jobs tests

tests/api/conftest.py:
- Add core_api_client_with_root fixture parametrized over three paths:
    v1-core:   /core/...          (canonical v1 path)
    v1-legacy: /homeassistant/... (legacy v1 alias, same handlers)
    v2-core:   /v2/core/...       (canonical v2 path)

tests/api/test_homeassistant.py:
- Switch all 17 api_client tests to core_api_client_with_root so each
  test runs against all three access paths (v1 canonical, v1 legacy
  alias, v2 canonical), exercising every registered route

tests/api/test_jobs.py:
- Promote four inner TestClass definitions to module-level helpers
  (_JobsTreeTestHelper, _JobManualCleanupTestHelper,
  _JobsSortedTestHelper, _JobWithErrorTestHelper) so that @Job name
  registration into the global _JOB_NAMES set only happens once at
  import time rather than on each parametrized test run
- Replace closure references to outer-scope coresys with self.coresys
- Use api_client_with_prefix for dual-version coverage

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Fix typo

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-04-27 23:39:47 +02:00
Mike Degatano 9f00b6e34f Ensure uuid of dismissed suggestion/issue matches an existing one (#6582)
* Ensure uuid of dismissed suggestion/issue matches an existing one

* Fix lint, test and feedback issues

* Adjust existing tests and remove new ones for not found errors

* fix device access issue usage
2026-02-25 10:26:44 +01:00
Mike Degatano 0636e49fe2 Enable mypy part 1 (addons and api) (#5759)
* Fix mypy issues in addons

* Fix mypy issues in api

* fix docstring

* Brackets instead of get with default
2025-03-25 15:06:35 -04:00
Mike Degatano 324b059970 Move write of core state to executor (#5720) 2025-03-04 17:49:53 +01:00
Mike Degatano d8101ddba8 Use status 404 in more places when appropriate (#5480) 2024-12-17 11:18:32 +01:00
Stefan Agner f6faa18409 Bump pre-commit ruff to 0.5.7 and reformat (#5242)
It seems that the codebase is not formatted with the latest ruff
version. This PR reformats the codebase with ruff 0.5.7.
2024-08-13 20:53:56 +02:00
Mike Degatano 7fd6dce55f Migrate to Ruff for lint and format (#4852)
* Migrate to Ruff for lint and format

* Fix pylint issues

* DBus property sets into normal awaitable methods

* Fix tests relying on separate tasks in connect

* Fixes from feedback
2024-02-05 11:37:39 -05:00
Mike Degatano df030e6209 No bool conversion in resolution test (#3826) 2022-08-26 16:42:45 +02:00
Mike Degatano 2c5bb3f714 Fix suggestion test to be order agnostic (#3822) 2022-08-25 16:06:47 -04:00
Mike Degatano d5f9fcfdc7 Fire events on issue changes (#3818)
* Fire events on issue changes

* API includes if suggestion has autofix
2022-08-25 11:42:31 +02:00
Joakim Sørensen 419f603571 Rename snapshot -> backup (#2940) 2021-07-27 10:06:09 +02:00
Pascal Vizeli b1232c0d8d Resolution: API call for run check manual (#2719) 2021-03-15 10:33:06 +01:00
Pascal Vizeli 55382d000b Revert "Tweak check API path" (#2718)
This reverts commit e30171746b.
2021-03-15 10:16:32 +01:00
Pascal Vizeli e30171746b Tweak check API path (#2714) 2021-03-12 11:42:24 +01:00
73849b7468 Check management (#2703)
* Check management

* Add test

* Don't allow disable core_security

* options and decorator

* streamline config handling

* streamline v2

* fix logging

* Add tests

* Fix test

* cleanup v1

* fix api

* Add more test

* Expose option also for cli

* address comments from Paulus

* Address second comment

* Update supervisor/resolution/checks/base.py

Co-authored-by: Paulus Schoutsen <balloob@gmail.com>

* fix lint

* Fix black

Co-authored-by: Pascal Vizeli <pvizeli@syshack.ch>
Co-authored-by: Paulus Schoutsen <balloob@gmail.com>
2021-03-12 11:32:56 +01:00
Pascal Vizeli 2040102e21 Handle Unhealthy like Unsupported (#2255)
* Handle Unhealthy like Unsupported

* Add tests

* Add unhealthy to sentry

* Add test
2020-11-14 16:16:00 +01:00
Pascal Vizeli 7a9aac491e Use /info for resolution API (#2136) 2020-10-16 13:07:02 +02:00
Pascal VizeliandJoakim Sørensen d119e99001 Resolution: extend type and context (#2130)
* Resolution: extend type and context

* fix property

* add helper

* fix api

* fix tests

* Fix patch

* finish tests

* Update supervisor/resolution/const.py

Co-authored-by: Joakim Sørensen <joasoe@gmail.com>

* Update supervisor/resolution/const.py

Co-authored-by: Joakim Sørensen <joasoe@gmail.com>

* Fix type

* fix lint

* Update supervisor/api/resolution.py

Co-authored-by: Joakim Sørensen <joasoe@gmail.com>

* Update supervisor/resolution/__init__.py

Co-authored-by: Joakim Sørensen <joasoe@gmail.com>

* Update API & add more tests

* Update supervisor/api/resolution.py

Co-authored-by: Joakim Sørensen <joasoe@gmail.com>

* Update supervisor/resolution/__init__.py

Co-authored-by: Joakim Sørensen <joasoe@gmail.com>

* Update supervisor/resolution/__init__.py

Co-authored-by: Joakim Sørensen <joasoe@gmail.com>

* fix black

* remove azure ci

* fix test

* fix tests

* fix tests

* fix tests p2

Co-authored-by: Joakim Sørensen <joasoe@gmail.com>
2020-10-16 12:22:32 +02:00
Joakim SørensenandPascal Vizeli 02e72726a5 Add issues/suggestion to resolution center / start with diskspace (#2125)
Co-authored-by: Pascal Vizeli <pvizeli@syshack.ch>
2020-10-14 17:14:25 +02:00
Joakim Sørensen 028b170cff Add resolution manager and unsupported flags (#2124)
* Add unsupported reason flags

* Restore test_network.py

* Add Resolution manager object

* fix import
2020-10-13 12:54:17 +02:00