Files
supervisor/tests/api/test_homeassistant.py
Stefan AgnerandClaude Fable 5 324d17e68b Use systemd automount units for network mounts (#7167)
* dbus: support aux units in start_transient_unit

Extend `Systemd.start_transient_unit` to accept the `aux` parameter
(`a(sa(sv))`), which has been hardcoded to `[]` since the wrapper was
introduced. Aux entries take the form `(unit_name, properties)` and
let systemd create multiple transient units atomically — most usefully
a `.mount` and its `.automount` companion in one D-Bus call.

Add the unit-property constants we'll need to drive that:

- `DBUS_ATTR_LAZY_UNMOUNT` ("LazyUnmount"): set on the `.mount` so
  systemd umounts with MNT_DETACH. Pairs with `softerr`/`soft` to make
  stop/restart sequences reliable even when the server is gone.
- `DBUS_ATTR_TIMEOUT_IDLE_USEC` ("TimeoutIdleUSec"): set on the
  `.automount` to control how long the mount stays around after the
  last access before autofs expires it.
- `DBUS_ATTR_WHERE` ("Where"): mount-point property; required on the
  `.automount` unit (and useful explicitly on `.mount` units too).

No call sites change in this commit — pure plumbing.

* mounts: pair each network .mount with a .automount companion

Network mounts now get a transient `.automount` unit created
atomically alongside their `.mount`, via the aux parameter on
`StartTransientUnit`. The systemd-managed autofs trigger handles
activation lazily: the path exists in the VFS even before the
underlying network mount runs, and the first access that crosses the
trigger fires `mount.cifs`/`mount.nfs` on demand.

Why this matters:

- PID 1's path-walks (chase, daemon-reload, generator scans) stop
  crossing dead NFS/CIFS lookups — autofs returns from kernel memory
  rather than entering the network filesystem. The class of failure
  fixed at the systemd-timeout layer in #6834 stops being reachable
  in the first place.
- Failed accesses fail fast, bounded by the `.mount`'s `TimeoutSec=`
  rather than hanging forever.
- Background reconnect (NFSv4 state-manager kthread, CIFS
  delayed_work) recovers transparently when the server returns; no
  remount needed.

The `.automount` carries `TimeoutIdleUSec=5min` so kernels can expire
the mount after inactivity and re-trigger on the next access. The
companion `.mount` carries `LazyUnmount=true` (MNT_DETACH on stop),
so umount returns immediately even when the server is unreachable;
existing fds drain in their soft/softerr timeout regime instead of
pinning the umount syscall.

`Mount.unmount()` now stops the `.automount` first so the autofs
trigger can't re-fire the underlying `.mount` during cleanup. The
automount stop is best-effort — if the unit is gone or refuses to
stop, we log and proceed to the `.mount` stop, which remains
authoritative.

Bind mounts opt out via a `creates_automount` class flag — they have
no server to wait on and the lazy semantics would only add
indirection. Network mounts (CIFS/NFS) opt in.

This is wiring only; the periodic mount reload, the reload/restart
escalation, and the inner bind-mount layer for media/share usage are
still in place. The next commit removes them now that automount
makes them unnecessary.

* mounts: drop bind layer, periodic reload, and reload/restart machinery

With the kernel's autofs trigger handling lazy activation and
transparent reconnect, almost all of the supervisor-side mount
choreography becomes unnecessary. Strip it out.

What goes away:

* The inner bind-mount layer. Media/share mounts used to live at
  `path_extern_mounts/{name}` and have a second `.mount` unit bind
  them into `path_extern_media/{name}` (resp. `share`). The
  `.automount` companion can sit directly at the container-facing
  path, and the parent-dir RSLAVE bind into add-on containers
  surfaces the autofs trigger the same way. `NetworkMount.where`
  now switches on usage:
    - MEDIA → `path_extern_media/{name}`
    - SHARE → `path_extern_share/{name}`
    - BACKUP → `path_extern_mounts/{name}` (unchanged)
* `BindMount`, `BoundMount`, `MountManager._bound_mounts`, the
  `_bind_mount`/`_bind_media`/`_bind_share` helpers, and the entire
  emergency-fallback dance (`path_emergency`). The empty-read-only
  dir trick was a workaround for the harder PID 1 wedge problem,
  which autofs solves at the root. Failed shares now surface as
  ETIMEDOUT/EHOSTDOWN on access plus a resolution issue.
* `MountManager.reload()` and the 15-minute `RUN_RELOAD_MOUNTS`
  periodic task. autofs re-activates on access; we don't need to
  poll, and not polling means we don't randomly trip stale-FH /
  softreval / dead-server edge cases on a timer.
* `Mount.reload()` and `Mount._restart()`. The reload→restart
  escalation existed to make `is_mounted` honest in the face of
  systemd's local-only state; now `is_mounted` IS honest (probe-
  based), and any recovery the kernel knows how to do happens
  inside autofs without our involvement.
* The `RELOADING` safety net introduced by #6834. That whole class
  of PID 1 wedge stops being reachable when path lookups don't
  cross dead network mounts.

What changes shape:

* `NetworkMount.is_mounted()` no longer gates on systemd's
  `ActiveState` before probing. The `.mount` unit is dormant
  whenever autofs hasn't recently triggered it, so systemd state
  is meaningless as a health signal. The probe (now via
  `statvfs("/path/.")` so the trailing dot forces `LOOKUP_DIRECTORY`
  and triggers autofs) is the source of truth, and after it returns
  we set `self._state` to ACTIVE/INACTIVE so the API reports
  reachability rather than autofs idle state.
* `MountManager.reload_mount()` becomes probe-only — no systemd
  reload/restart calls. The user-facing semantics: "tell me if this
  mount is reachable right now, and refresh the resolution issue
  accordingly." If the mount is dead, autofs will re-trigger it on
  the next consumer access; the supervisor doesn't need to force it.
* `Mount.load()` keeps `_update_state_await` for the no-job-dispatched
  case where a previous supervisor left a unit in `activating`, then
  probes once via `is_mounted` so state reflects reachability rather
  than the lazy-mount idle state.
* `BackupManager` no longer walks `bound_mounts`. Backups skip
  network-mount subdirectories during folder archive (so we don't
  recurse into the share), and unmount/re-mount them around folder
  restores (so writes target the local mount-point dir, not the
  remote share).

Behavior tradeoffs accepted:

* Dead media/share access now ~30s ETIMEDOUT instead of empty
  read-only dir. The empty-dir was a workaround for the PID 1
  wedge that autofs eliminates at the root; if specific add-ons
  rely on the old behavior we can revisit with a 2-stage automount
  whose fallback target is an emergency dir.
* Resolution-issue lag: failed mounts no longer surface within
  15 min on their own. They appear when the user probes via the
  API, when `BackupManager.reload()` walks the location, or when
  load-time activation fails.

* tests: adapt bind-layer-era tests after rebase

Main gained tests for the bind layer after this branch was written:
the #7013 rebind regression test and the #7072 bind-step rollback test
cover machinery that no longer exists, and the healthy-reload test
asserted the unconditional rebind. Drop the first two and reduce the
third to asserting that a healthy probe performs no systemd operations
at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: rearchitect automount setup from systemd/kernel review

A code-level review of systemd's automount implementation (verified
against v254.13, the version HAOS ships) and the kernel autofs/VFS
plumbing surfaced three correctness-critical flaws in the automount
design, plus several hardening gaps. See automount-rearchitecture.md
for the full analysis.

- Make the .automount the primary transient unit with the .mount as
  aux, mirroring `systemd-mount --automount=yes`. Aux units get no
  start job: with the .mount as primary the trigger was never armed
  and the design silently degraded to eager mounting.

- Set StartLimitIntervalUSec=0 on the .mount. With the default start
  rate limit (5 starts/10s, counting successful starts) a fast-failing
  mount plus any polling consumer trips the limit within seconds, and
  systemd then detaches the autofs trigger entirely
  (AUTOMOUNT_FAILURE_MOUNT_START_LIMIT_HIT) — the path silently
  becomes a plain writable local directory.

- Drop the 5-minute TimeoutIdleUSec (default 0 = never expire).
  Kernel idle expiry decides busyness via may_umount_tree(), which
  only counts the init-namespace mount instance — open files held by
  container processes are invisible, so expiry would unmount shares
  under actively writing add-ons.

- Probe with a plain statvfs: the statfs syscall walks with
  LOOKUP_AUTOMOUNT and triggers by itself; the trailing-dot trick is
  unnecessary. Classify ELOOP as a mount-propagation
  misconfiguration in the probe error handling.

- Re-arm the trigger from reload_mount(): if the .automount unit is
  failed or gone (e.g. autofs unmounted out-of-band), reset failure
  state and re-create the pair before probing.

- Tear down legacy eager-mount units during load(). The old design's
  bind unit for media/share occupies the exact unit name the network
  .mount uses now; on a warm upgrade adoption would mistake the old
  bind mount for the network mount and leak the legacy data mount.

- Honor the systemd job result when stopping the .mount during
  unmount and reset failure state on both units afterwards so dead
  transient units get garbage-collected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: adapt local data repair to the autofs design

Port of the mount-target-not-empty repair (#7089) on top of the
automount rearchitecture:

- relocate_local_data() handles a single target directory — the mount
  sits directly at its container-facing path, the bind layer and with
  it the multi-directory case are gone
- a repair_trigger() failure caused by blocking local data raises the
  mount failed issue with the move_local_data suggestion instead of
  the plain variant
- local data at load is detected via the mount unit itself; the
  bind-layer detection test is rewritten accordingly

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docker: use rslave propagation for execute_command share mount

The temporary container used for execute_command (e.g. core config
check) mounted /share without a propagation mode, unlike every other
share/media mount. With eager mounts this only meant missing mounts
made after container start; with lazy automount activation the shares
are routinely mounted after start, and accessing a not-yet-activated
automount from a private mount namespace fails with ELOOP.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: harden automount lifecycle from review findings

Address the findings of an adversarial review of the automount
rearchitecture:

- load() no longer adopts a dead automount trigger: a failed or
  stopped .automount leaves the path a plain writable directory — the
  silent local-write degradation this design must prevent. Tear down
  and re-arm instead. An adopted pair whose probe fails now raises so
  the manager surfaces the mount failed issue, same as a fresh mount.
- unmount() resolves the .mount unit only after stopping the
  .automount (the lazy detach can garbage-collect the transient unit,
  invalidating an earlier proxy), raises instead of warning when the
  automount stop errors, and checks the stop job result so a failed
  stop cannot leave an armed trigger behind a "successful" removal.
- repair_trigger() fully unmounts before re-arming: with the share
  still attached, arming fails — or worse, mount() misreads the
  mounted share's contents as blocking local data.
- Legacy unit teardown honors the stop job result for the same reason.
- reload_mount() escalates once to re-creating the unit pair when an
  established mount is unreachable: a permanently dead session (e.g.
  replaced server) keeps the path mounted, so the trigger can never
  re-fire and kernel reconnection never succeeds. Teardown is safe now
  (lazy unmount, no PID 1 path walks), unlike the removed reload →
  restart escalation of the eager design.
- Reinstate the periodic mounts task as a probe-based reconcile: re-arm
  dead triggers, refresh the reachability state reported by the API and
  used by backup locations, and sync the mount failed issue in both
  directions. No reload or restart of established mounts.
- Folder restore no longer fails after a successful restore when
  re-mounting nested mounts cannot verify an unreachable server — the
  trigger is armed, the share recovers on next access.
- Drop the now-unused Mount.update(), deduplicate unit name escaping,
  hoist the backup exclusion set out of the per-file filter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* tests: expect rslave propagation in core check container

The config check assertion missed the update for the rslave
propagation on the execute_command share mount.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* tests: cover reconcile trigger repair failure paths

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: arm automount before stopping the legacy data mount

Address review: the legacy teardown stopped the media/share bind unit
and the eager data mount in sequence, both strictly. When the bind
stop succeeded but the data mount stop failed — its unmount can time
out against an unreachable server, legacy units have no LazyUnmount —
load() raised with the container-facing path left a plain writable
directory and no trigger armed: the pollution mode this design is
meant to prevent. Reordering the stops would not help either, as
systemd stops the bind first anyway as a dependent of the data mount
(the implicit Requires= from RequiresMountsFor= on the bind's What=).

Instead, strictly stop only the path-conflicting unit (a failure
leaves the path covered, which is safe and retryable), arm the
automount right after, and stop the conflict-free legacy data mount
best-effort last — a failed stop logs a warning and leaves an orphaned
mount for the next Supervisor restart or a host reboot, with nothing
writable exposed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: discard dead session instead of re-creating units on reload

Address review: re-creating the unit pair on reload of an unreachable
established mount exposed the target as a plain writable directory
between the lazy unmount and arming the replacement — a continuously
writing add-on could block the re-arm or slip writes under the new
mount in the check-to-arm window.

Stop only the .mount unit instead, keeping the .automount armed:
systemd re-installs the autofs trigger over the path — the same
mechanism idle expiry uses, with the automount's Triggers= reference
keeping the transient .mount definition alive — so the path is never
locally writable. The re-probe then mounts fresh through the trigger,
which establishes a new session and thereby covers the permanently
dead session case (e.g. a replaced server) the escalation exists for.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: surface arming failures, keep restore teardown re-armed

Address review findings on the arming/re-arming error paths:

- mount() checks the StartTransientUnit job result: a failed start job
  means the trigger never armed and the path is a plain writable
  directory — a hard MountError before the probe, so it cannot be
  mistaken for the armed-but-unreachable MountActivationError.
- Folder restore includes the nested-mount teardown in the try block:
  an unmount failing halfway (trigger disarmed, share still attached)
  previously exited before any re-arm, leaving the path unprotected.
  The finally re-arms via repair_trigger(), which no-ops on a still
  armed trigger and handles partially torn-down pairs.
- The post-restore re-arm suppresses only MountActivationError (armed,
  recovers on next access). Any other failure — arming failed, local
  data blocking the target — left the path unprotected and surfaces.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: fold the two legacy unit teardowns into one helper

The eager-mount-era cleanup queried the .mount unit a second time
although load() had just fetched it, and the strict teardown of the
unit occupying the automount's path and the best-effort teardown of
the mount at the mounts data directory were near identical. Reuse the
unit from load() and give the shared helper a strict flag, which drops
a D-Bus round trip per mount on every load.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: re-arm the trigger when unmount cannot stop the mount

Stopping the automount detaches the whole stack at the path, so an
unmount that then fails to stop the .mount leaves a plain writable
directory behind. Arm a fresh pair before raising, best effort — the
local data repair covers what lands there if that fails too.

The automount stop itself needs no such handling: systemd's
automount_stop() enters dead synchronously and the detach it performs
(MNT_DETACH, no server contact) only logs its errors, so a failure
there means systemd could not be reached at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: raise translatable errors for mount operations

Errors from setting up, unmounting and reloading a mount reach users
through the API, so describe what failed in a translatable message and
leave the systemd job result or D-Bus error to the log.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: remove the emergency folder of the eager-mount design

The read-only fallback directory has no place in the automount design.
Remove what is left of it on existing installations, next to the legacy
addons directory cleanup, keeping it if it holds anything but the empty
mount points.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: trim comments to the behavior they describe

Drop the narration of what changed and why from comments and
docstrings, keeping the notes that record why a tempting alternative
does not work. Two comments still described the reload to restart
escalation this branch removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: correct stale claims in manager and reload test

The manager docstring credited the kernel with idle expiry, which is
deliberately disabled, and claimed no polling although the periodic
reconcile probes every 15 minutes. The reload test claimed systemd is
never contacted while the escalation stops the .mount unit; only reload
and restart of the unit are avoided.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: arm the trigger alone if re-creating the pair is rejected

Address review: the re-arm after a failed unmount submits the .mount as
an aux unit, but transient creation requires a pristine unit and the
definition is still loaded precisely when stopping it is what failed.
Fall back to creating the .automount on its own, which covers the path
and fires the surviving mount definition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: never arm the trigger over local data after a failed unmount

Address review: the fallback caught the target validation errors too,
so data written while the path was uncovered would end up beneath a
fresh trigger. Reconciliation would then find a healthy mount and the
repair skips mount points, leaving the data hidden for good. Leave the
path untouched instead, so the next reconcile offers to move it away.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-31 20:30:23 +02:00

833 lines
30 KiB
Python

"""Test homeassistant api."""
import asyncio
from pathlib import Path
from unittest.mock import AsyncMock, Mock, PropertyMock, patch
from aiodocker.containers import DockerContainer
from aiohttp.test_utils import TestClient
from awesomeversion import AwesomeVersion
import pytest
from supervisor.backups.manager import BackupManager
from supervisor.const import DNS_SUFFIX, CoreState
from supervisor.coresys import CoreSys
from supervisor.docker.homeassistant import DockerHomeAssistant
from supervisor.docker.interface import DockerInterface
from supervisor.docker.manager import DockerAPI
from supervisor.exceptions import DockerError, HomeAssistantError
from supervisor.homeassistant.api import APIState, HomeAssistantAPI
from supervisor.homeassistant.const import WSEvent
import supervisor.homeassistant.core as ha_core
from supervisor.homeassistant.core import HomeAssistantCore
from supervisor.homeassistant.module import HomeAssistant
from supervisor.resolution.const import ContextType, IssueType
from supervisor.resolution.data import Issue
from supervisor.updater import Updater
from tests.common import AsyncIterator, load_json_fixture
@pytest.mark.parametrize("legacy_route", [True, False])
async def test_api_core_logs(
advanced_logs_tester: AsyncMock,
legacy_route: bool,
):
"""Test core logs."""
await advanced_logs_tester(
f"/{'homeassistant' if legacy_route else 'core'}",
"homeassistant",
v2_path_prefix="/core",
)
async def test_api_stats(
core_api_client_with_root: tuple[TestClient, str], container: DockerContainer
):
"""Test stats."""
api_client, root = core_api_client_with_root
container.show.return_value["State"]["Status"] = "running"
container.show.return_value["State"]["Running"] = True
if root.startswith("/v2"):
# V2 always requests one-shot stats
stats_fixture = load_json_fixture("container_stats.json")
del stats_fixture["precpu_stats"]
with patch.object(
DockerAPI,
"_query_one_shot_stats",
AsyncMock(return_value=stats_fixture),
):
resp = await api_client.get(f"{root}/stats")
else:
container.stats = AsyncMock(
return_value=[load_json_fixture("container_stats.json")]
)
resp = await api_client.get(f"{root}/stats")
assert resp.status == 200
result = await resp.json()
if root.startswith("/v2"):
assert "cpu_percent" not in result["data"]
else:
assert result["data"]["cpu_percent"] == 90.0
assert result["data"]["cpu_usage"] == 190
assert result["data"]["cpu_system_usage"] == 200
assert result["data"]["online_cpus"] == 24
assert result["data"]["memory_usage"] == 59700000
assert result["data"]["memory_limit"] == 4000000000
assert result["data"]["memory_percent"] == 1.49
async def test_api_stats_one_shot(
core_api_client_with_root: tuple[TestClient, str],
container: DockerContainer,
):
"""Test stats one-shot mode skips the calculated window.
V1 opts in via the ``one_shot`` query string; v2 always requests
one-shot stats and has no query string option (see test_api_stats).
"""
api_client, root = core_api_client_with_root
if root.startswith("/v2"):
pytest.skip("v2 always uses one-shot; covered by test_api_stats")
container.show.return_value["State"]["Status"] = "running"
container.show.return_value["State"]["Running"] = True
stats_fixture = load_json_fixture("container_stats.json")
del stats_fixture["precpu_stats"]
with patch.object(
DockerAPI, "_query_one_shot_stats", AsyncMock(return_value=stats_fixture)
):
resp = await api_client.get(f"{root}/stats?one_shot")
assert resp.status == 200
result = await resp.json()
assert result["data"]["cpu_percent"] is None
assert result["data"]["cpu_usage"] == 190
assert result["data"]["cpu_system_usage"] == 200
assert result["data"]["online_cpus"] == 24
container.stats.assert_not_called()
async def test_api_set_options(core_api_client_with_root: tuple[TestClient, str]):
"""Test setting options for homeassistant."""
api_client, root = core_api_client_with_root
resp = await api_client.get(f"{root}/info")
assert resp.status == 200
result = await resp.json()
assert result["data"]["watchdog"] is True
assert result["data"]["backups_exclude_database"] is False
with patch.object(HomeAssistant, "save_data") as save_data:
resp = await api_client.post(
f"{root}/options",
json={"backups_exclude_database": True, "watchdog": False},
)
assert resp.status == 200
save_data.assert_called_once()
resp = await api_client.get(f"{root}/info")
assert resp.status == 200
result = await resp.json()
assert result["data"]["watchdog"] is False
assert result["data"]["backups_exclude_database"] is True
async def test_api_set_image(
core_api_client_with_root: tuple[TestClient, str], coresys: CoreSys
):
"""Test changing the image for homeassistant."""
api_client, root = core_api_client_with_root
assert (
coresys.homeassistant.image == "ghcr.io/home-assistant/qemux86-64-homeassistant"
)
assert coresys.homeassistant.override_image is False
with patch.object(HomeAssistant, "save_data"):
resp = await api_client.post(
f"{root}/options",
json={"image": "test_image"},
)
assert resp.status == 200
assert coresys.homeassistant.image == "test_image"
assert coresys.homeassistant.override_image is True
with patch.object(HomeAssistant, "save_data"):
resp = await api_client.post(
f"{root}/options",
json={"image": "ghcr.io/home-assistant/qemux86-64-homeassistant"},
)
assert resp.status == 200
assert (
coresys.homeassistant.image == "ghcr.io/home-assistant/qemux86-64-homeassistant"
)
assert coresys.homeassistant.override_image is False
async def test_api_restart(
core_api_client_with_root: tuple[TestClient, str],
container: DockerContainer,
tmp_supervisor_data: Path,
):
"""Test restarting homeassistant."""
api_client, root = core_api_client_with_root
safe_mode_marker = tmp_supervisor_data / "homeassistant" / "safe-mode"
with patch.object(HomeAssistantCore, "_block_till_run"):
await api_client.post(f"{root}/restart")
container.restart.assert_called_once()
assert not safe_mode_marker.exists()
with patch.object(HomeAssistantCore, "_block_till_run"):
await api_client.post(f"{root}/restart", json={"safe_mode": True})
assert container.restart.call_count == 2
assert safe_mode_marker.exists()
@pytest.mark.usefixtures("path_extern")
async def test_api_rebuild(
core_api_client_with_root: tuple[TestClient, str],
coresys: CoreSys,
container: DockerContainer,
tmp_supervisor_data: Path,
):
"""Test rebuilding homeassistant."""
api_client, root = core_api_client_with_root
coresys.homeassistant.version = AwesomeVersion("2023.09.0")
safe_mode_marker = tmp_supervisor_data / "homeassistant" / "safe-mode"
with patch.object(HomeAssistantCore, "_block_till_run"):
await api_client.post(f"{root}/rebuild")
assert container.delete.call_count == 2
container.start.assert_called_once()
assert not safe_mode_marker.exists()
with patch.object(HomeAssistantCore, "_block_till_run"):
await api_client.post(f"{root}/rebuild", json={"safe_mode": True})
assert container.delete.call_count == 4
assert container.start.call_count == 2
assert safe_mode_marker.exists()
@pytest.mark.parametrize("action", ["rebuild", "restart", "stop", "update"])
async def test_migration_blocks_stopping_core(
core_api_client_with_root: tuple[TestClient, str], coresys: CoreSys, action: str
):
"""Test that an offline db migration in progress stops users from stopping/restarting core."""
api_client, root = core_api_client_with_root
coresys.homeassistant.api.get_api_state.return_value = APIState("NOT_RUNNING", True)
resp = await api_client.post(f"{root}/{action}")
assert resp.status == 503
result = await resp.json()
assert (
result["message"]
== "Offline database migration in progress, try again after it has completed"
)
async def test_force_rebuild_during_migration(
core_api_client_with_root: tuple[TestClient, str], coresys: CoreSys
):
"""Test force option rebuilds even during a migration."""
api_client, root = core_api_client_with_root
coresys.homeassistant.api.get_api_state.return_value = APIState("NOT_RUNNING", True)
with patch.object(HomeAssistantCore, "rebuild") as rebuild:
await api_client.post(f"{root}/rebuild", json={"force": True})
rebuild.assert_called_once()
async def test_force_restart_during_migration(
core_api_client_with_root: tuple[TestClient, str], coresys: CoreSys
):
"""Test force option restarts even during a migration."""
api_client, root = core_api_client_with_root
coresys.homeassistant.api.get_api_state.return_value = APIState("NOT_RUNNING", True)
with patch.object(HomeAssistantCore, "restart") as restart:
await api_client.post(f"{root}/restart", json={"force": True})
restart.assert_called_once()
async def test_force_stop_during_migration(
core_api_client_with_root: tuple[TestClient, str], coresys: CoreSys
):
"""Test force option stops even during a migration."""
api_client, root = core_api_client_with_root
coresys.homeassistant.api.get_api_state.return_value = APIState("NOT_RUNNING", True)
with patch.object(HomeAssistantCore, "stop") as stop:
await api_client.post(f"{root}/stop", json={"force": True})
stop.assert_called_once()
@pytest.mark.parametrize(
("make_backup", "backup_called", "update_called"),
[(True, True, False), (False, False, True)],
)
async def test_home_assistant_background_update(
core_api_client_with_root: tuple[TestClient, str],
coresys: CoreSys,
make_backup: bool,
backup_called: bool,
update_called: bool,
):
"""Test background update of Home Assistant."""
api_client, root = core_api_client_with_root
coresys.hardware.disk.get_disk_free_space = lambda x: 5000
event = asyncio.Event()
mock_update_called = mock_backup_called = False
# Mock backup/update as long-running tasks
async def mock_docker_interface_update(*args, **kwargs):
nonlocal mock_update_called
mock_update_called = True
await event.wait()
async def mock_partial_backup(*args, **kwargs):
nonlocal mock_backup_called
mock_backup_called = True
await event.wait()
with (
patch.object(DockerInterface, "update", new=mock_docker_interface_update),
patch.object(BackupManager, "do_backup_partial", new=mock_partial_backup),
patch.object(
DockerInterface,
"version",
new=PropertyMock(return_value=AwesomeVersion("2025.8.0")),
),
):
resp = await api_client.post(
f"{root}/update",
json={"background": True, "backup": make_backup, "version": "2025.8.3"},
)
assert mock_backup_called is backup_called
assert mock_update_called is update_called
assert resp.status == 200
body = await resp.json()
assert (job := coresys.jobs.get_job(body["data"]["job_id"]))
assert job.name == "home_assistant_core_update"
event.set()
async def test_background_home_assistant_update_fails_fast(
core_api_client_with_root: tuple[TestClient, str], coresys: CoreSys
):
"""Test background Home Assistant update returns error not job if validation doesn't succeed."""
api_client, root = core_api_client_with_root
coresys.hardware.disk.get_disk_free_space = lambda x: 5000
with (
patch.object(
DockerInterface,
"version",
new=PropertyMock(return_value=AwesomeVersion("2025.8.3")),
),
):
resp = await api_client.post(
f"{root}/update",
json={"background": True, "version": "2025.8.3"},
)
assert resp.status == 400
body = await resp.json()
assert body["message"] == "Home Assistant version 2025.8.3 is already installed"
assert body["error_key"] == "homeassistant_update_already_installed_error"
@pytest.mark.usefixtures("tmp_supervisor_data")
async def test_api_progress_updates_home_assistant_update(
core_api_client_with_root: tuple[TestClient, str],
coresys: CoreSys,
ha_ws_client: AsyncMock,
):
"""Test progress updates sent to Home Assistant for updates."""
api_client, root = core_api_client_with_root
coresys.hardware.disk.get_disk_free_space = lambda x: 5000
coresys.core.set_state(CoreState.RUNNING)
logs = load_json_fixture("docker_pull_image_log.json")
coresys.docker.images.pull.return_value = AsyncIterator(logs)
coresys.homeassistant.version = AwesomeVersion("2025.8.0")
with (
patch.object(
DockerHomeAssistant,
"version",
new=PropertyMock(return_value=AwesomeVersion("2025.8.0")),
),
patch.object(
HomeAssistantAPI,
"get_config",
return_value={"components": ["http", "frontend", "websocket_api"]},
),
patch.object(ha_core, "verify_frontend", AsyncMock(return_value=True)),
):
resp = await api_client.post(f"{root}/update", json={"version": "2025.8.3"})
assert resp.status == 200
events = [
{
"stage": evt.args[0]["data"]["data"]["stage"],
"progress": evt.args[0]["data"]["data"]["progress"],
"done": evt.args[0]["data"]["data"]["done"],
}
for evt in ha_ws_client.async_send_command.call_args_list
if "data" in evt.args[0]
and evt.args[0]["data"]["event"] == WSEvent.JOB
and evt.args[0]["data"]["data"]["name"] == "home_assistant_core_update"
]
# Count-based progress: 2 layers need pulling (each worth 50%)
# Layers that already exist are excluded from progress calculation
assert events[:5] == [
{
"stage": None,
"progress": 0,
"done": None,
},
{
"stage": None,
"progress": 0,
"done": False,
},
{
"stage": None,
"progress": 9.2,
"done": False,
},
{
"stage": None,
"progress": 25.6,
"done": False,
},
{
"stage": None,
"progress": 35.4,
"done": False,
},
]
assert events[-5:] == [
{
"stage": None,
"progress": 95.5,
"done": False,
},
{
"stage": None,
"progress": 96.9,
"done": False,
},
{
"stage": None,
"progress": 98.2,
"done": False,
},
{
"stage": None,
"progress": 100,
"done": False,
},
{
"stage": None,
"progress": 100,
"done": True,
},
]
@pytest.mark.usefixtures("path_extern")
async def test_config_check(
core_api_client_with_root: tuple[TestClient, str],
coresys: CoreSys,
container: DockerContainer,
):
"""Test config check API."""
api_client, root = core_api_client_with_root
coresys.homeassistant.version = AwesomeVersion("2025.1.0")
result = await api_client.post(f"{root}/check")
assert result.status == 200
coresys.docker.containers.create.assert_called_once_with(
{
"Image": "ghcr.io/home-assistant/qemux86-64-homeassistant:2025.1.0",
"Labels": {"supervisor_managed": ""},
"OpenStdin": False,
"StdinOnce": False,
"AttachStdin": False,
"AttachStdout": False,
"AttachStderr": False,
"HostConfig": {
"NetworkMode": "hassio",
"Init": True,
"Privileged": True,
"Mounts": [
{
"Type": "bind",
"Source": "/mnt/data/supervisor/homeassistant",
"Target": "/config",
"ReadOnly": False,
},
{
"Type": "bind",
"Source": "/mnt/data/supervisor/ssl",
"Target": "/ssl",
"ReadOnly": True,
},
{
"Type": "bind",
"Source": "/mnt/data/supervisor/share",
"Target": "/share",
"ReadOnly": False,
"BindOptions": {"Propagation": "rslave"},
},
],
"Dns": [str(coresys.docker.network.dns)],
"DnsSearch": [DNS_SUFFIX],
"DnsOptions": ["timeout:10"],
},
"Env": ["TZ=Etc/UTC"],
"Entrypoint": [],
"Cmd": [
"python3",
"-m",
"homeassistant",
"-c",
"/config",
"--script",
"check_config",
],
},
name=None,
)
container.start.assert_called_once()
@pytest.mark.usefixtures("path_extern")
async def test_config_check_error(
core_api_client_with_root: tuple[TestClient, str], container: DockerContainer
):
"""Test config check API strips color coding from log output on error."""
api_client, root = core_api_client_with_root
container.log.return_value = [
"\x1b[36mTest logs 1\x1b[0m\n",
"\x1b[36mTest logs 2\x1b[0m\n",
]
container.wait.return_value = {"StatusCode": 1}
result = await api_client.post(f"{root}/check")
assert result.status == 400
resp = await result.json()
assert resp["message"] == "Test logs 1\nTest logs 2\n"
async def test_update_frontend_check_success(
core_api_client_with_root: tuple[TestClient, str], coresys: CoreSys
):
"""Test that update succeeds when frontend check passes."""
api_client, root = core_api_client_with_root
coresys.hardware.disk.get_disk_free_space = lambda x: 5000
coresys.homeassistant.version = AwesomeVersion("2025.8.0")
with (
patch.object(DockerInterface, "is_running", AsyncMock(return_value=True)),
patch.object(
DockerHomeAssistant,
"version",
new=PropertyMock(return_value=AwesomeVersion("2025.8.0")),
),
patch.object(
HomeAssistantAPI,
"get_config",
return_value={"components": ["http", "frontend", "websocket_api"]},
),
patch.object(ha_core, "verify_frontend", AsyncMock(return_value=True)),
patch.object(DockerInterface, "cleanup") as mock_cleanup,
):
resp = await api_client.post(f"{root}/update", json={"version": "2025.8.3"})
assert resp.status == 200
mock_cleanup.assert_called_once()
async def test_update_frontend_check_fails_triggers_rollback(
core_api_client_with_root: tuple[TestClient, str],
coresys: CoreSys,
caplog: pytest.LogCaptureFixture,
tmp_supervisor_data: Path,
):
"""Test that update triggers rollback when health probes fail."""
api_client, root = core_api_client_with_root
coresys.hardware.disk.get_disk_free_space = lambda x: 5000
coresys.homeassistant.version = AwesomeVersion("2025.8.0")
# Mock successful first update, failed frontend check, then successful rollback
update_call_count = 0
async def mock_update(*args, **kwargs):
nonlocal update_call_count
update_call_count += 1
if update_call_count == 1:
# First update succeeds
coresys.homeassistant.version = AwesomeVersion("2025.8.3")
elif update_call_count == 2:
# Rollback succeeds
coresys.homeassistant.version = AwesomeVersion("2025.8.0")
with (
patch.object(DockerInterface, "update", new=mock_update),
patch.object(DockerInterface, "is_running", AsyncMock(return_value=True)),
patch.object(
DockerHomeAssistant,
"version",
new=PropertyMock(return_value=AwesomeVersion("2025.8.0")),
),
patch.object(
HomeAssistantAPI,
"get_config",
return_value={"components": ["http", "frontend", "websocket_api"]},
),
patch.object(ha_core, "verify_frontend", AsyncMock(return_value=False)),
patch.object(DockerInterface, "cleanup") as mock_cleanup,
):
resp = await api_client.post(f"{root}/update", json={"version": "2025.8.3"})
# Update should trigger rollback, which succeeds and returns 200
assert resp.status == 200
assert "HomeAssistant update failed -> rollback!" in caplog.text
# Should have called update twice (once for update, once for rollback)
assert update_call_count == 2
# An update_rollback issue should be created
assert (
Issue(IssueType.UPDATE_ROLLBACK, ContextType.CORE) in coresys.resolution.issues
)
# Old image should not be cleaned up so rollback doesn't need to re-download
mock_cleanup.assert_not_called()
async def test_update_websocket_api_missing_triggers_rollback(
core_api_client_with_root: tuple[TestClient, str],
coresys: CoreSys,
caplog: pytest.LogCaptureFixture,
tmp_supervisor_data: Path,
):
"""Test that update triggers rollback when websocket_api component is not loaded."""
api_client, root = core_api_client_with_root
coresys.hardware.disk.get_disk_free_space = lambda x: 5000
coresys.homeassistant.version = AwesomeVersion("2025.8.0")
update_call_count = 0
async def mock_update(*args, **kwargs):
nonlocal update_call_count
update_call_count += 1
if update_call_count == 1:
coresys.homeassistant.version = AwesomeVersion("2025.8.3")
elif update_call_count == 2:
coresys.homeassistant.version = AwesomeVersion("2025.8.0")
with (
patch.object(DockerInterface, "update", new=mock_update),
patch.object(DockerInterface, "is_running", AsyncMock(return_value=True)),
patch.object(
DockerHomeAssistant,
"version",
new=PropertyMock(return_value=AwesomeVersion("2025.8.0")),
),
patch.object(
HomeAssistantAPI,
"get_config",
return_value={"components": ["http", "frontend"]},
),
patch.object(
ha_core, "verify_frontend", AsyncMock(return_value=True)
) as mock_frontend_check,
patch.object(DockerInterface, "cleanup") as mock_cleanup,
):
resp = await api_client.post(f"{root}/update", json={"version": "2025.8.3"})
assert resp.status == 200
assert "API responds but websocket_api is not loaded" in caplog.text
assert "HomeAssistant update failed -> rollback!" in caplog.text
assert update_call_count == 2
mock_frontend_check.assert_not_called()
assert (
Issue(IssueType.UPDATE_ROLLBACK, ContextType.CORE) in coresys.resolution.issues
)
mock_cleanup.assert_not_called()
async def test_update_get_config_error_triggers_rollback(
core_api_client_with_root: tuple[TestClient, str],
coresys: CoreSys,
caplog: pytest.LogCaptureFixture,
tmp_supervisor_data: Path,
):
"""Test that update triggers rollback when get_config raises HomeAssistantError."""
api_client, root = core_api_client_with_root
coresys.hardware.disk.get_disk_free_space = lambda x: 5000
coresys.homeassistant.version = AwesomeVersion("2025.8.0")
update_call_count = 0
async def mock_update(*args, **kwargs):
nonlocal update_call_count
update_call_count += 1
if update_call_count == 1:
coresys.homeassistant.version = AwesomeVersion("2025.8.3")
elif update_call_count == 2:
coresys.homeassistant.version = AwesomeVersion("2025.8.0")
with (
patch.object(DockerInterface, "update", new=mock_update),
patch.object(DockerInterface, "is_running", AsyncMock(return_value=True)),
patch.object(
DockerHomeAssistant,
"version",
new=PropertyMock(return_value=AwesomeVersion("2025.8.0")),
),
patch.object(HomeAssistantAPI, "get_config", side_effect=HomeAssistantError),
patch.object(
ha_core, "verify_frontend", AsyncMock(return_value=True)
) as mock_frontend_check,
patch.object(DockerInterface, "cleanup") as mock_cleanup,
):
resp = await api_client.post(f"{root}/update", json={"version": "2025.8.3"})
assert resp.status == 200
assert "HomeAssistant update failed -> rollback!" in caplog.text
assert update_call_count == 2
mock_frontend_check.assert_not_called()
assert (
Issue(IssueType.UPDATE_ROLLBACK, ContextType.CORE) in coresys.resolution.issues
)
mock_cleanup.assert_not_called()
async def test_update_image_install_failure_surfaces_error(
core_api_client_with_root: tuple[TestClient, str],
coresys: CoreSys,
capture_exception: Mock,
caplog: pytest.LogCaptureFixture,
tmp_supervisor_data: Path,
):
"""Test that a failed image install surfaces a clean translated error.
When the target version does not exist the image pull fails before the
running Core is touched. The post-update health check would otherwise pass
against the still-healthy old version and mask the failure. The error is
modeled as a client-facing APIError so it does not reach Sentry.
"""
api_client, root = core_api_client_with_root
coresys.hardware.disk.get_disk_free_space = lambda x: 5000
coresys.homeassistant.version = AwesomeVersion("2025.8.0")
with (
patch.object(
DockerInterface, "update", side_effect=DockerError("404 not found")
),
patch.object(DockerInterface, "is_running", AsyncMock(return_value=True)),
patch.object(
DockerHomeAssistant,
"version",
new=PropertyMock(return_value=AwesomeVersion("2025.8.0")),
),
patch.object(HomeAssistantAPI, "get_config") as mock_get_config,
patch.object(ha_core, "verify_frontend", AsyncMock(return_value=True)),
patch.object(DockerInterface, "cleanup") as mock_cleanup,
):
resp = await api_client.post(f"{root}/update", json={"version": "2025.8.3"})
# The update reports a clean client error with a translatable key.
assert resp.status == 400
body = await resp.json()
assert body["result"] == "error"
assert body["error_key"] == "homeassistant_update_image_error"
assert "2025.8.3" in body["message"]
# A client error must not be captured as an unexpected error in Sentry.
capture_exception.assert_not_called()
assert "Unexpected error during API call" not in caplog.text
# No health check or rollback should run since the old Core is untouched.
mock_get_config.assert_not_called()
assert "HomeAssistant update failed -> rollback!" not in caplog.text
mock_cleanup.assert_not_called()
# Version is unchanged and no rollback issue is created.
assert coresys.homeassistant.version == AwesomeVersion("2025.8.0")
assert (
Issue(IssueType.UPDATE_ROLLBACK, ContextType.CORE)
not in coresys.resolution.issues
)
async def test_update_skips_health_check_when_core_not_running(
coresys: CoreSys,
caplog: pytest.LogCaptureFixture,
tmp_supervisor_data: Path,
):
"""Test that update skips health check and rollback when Core was stopped on entry.
Reproduces the backup-restore regression: the restore flow stops and
removes Core before calling core.update(); the post-update API check
must not fire because Core hasn't been started yet, otherwise it
triggers a spurious rollback that overwrites the restored image.
"""
coresys.hardware.disk.get_disk_free_space = lambda x: 5000
coresys.homeassistant.version = AwesomeVersion("2026.5.0b0")
coresys.homeassistant.set_image("ghcr.io/home-assistant/qemux86-64-homeassistant")
update_call_count = 0
async def mock_update(*args, **kwargs):
nonlocal update_call_count
update_call_count += 1
coresys.homeassistant.version = AwesomeVersion("2026.4.4")
with (
patch.object(DockerInterface, "update", new=mock_update),
patch.object(DockerInterface, "is_running", AsyncMock(return_value=False)),
patch.object(DockerInterface, "exists", AsyncMock(return_value=False)),
patch.object(
Updater,
"image_homeassistant",
new=PropertyMock(
return_value="ghcr.io/home-assistant/qemux86-64-homeassistant"
),
),
patch.object(
DockerHomeAssistant,
"version",
new=PropertyMock(return_value=AwesomeVersion("2026.4.4")),
),
patch.object(HomeAssistantAPI, "get_config") as mock_get_config,
patch.object(ha_core, "verify_frontend", AsyncMock()) as mock_frontend,
patch.object(DockerInterface, "cleanup") as mock_cleanup,
):
await coresys.homeassistant.core.update(AwesomeVersion("2026.4.4"))
# Only one update call: no rollback fired.
assert update_call_count == 1
assert "HomeAssistant update failed -> rollback!" not in caplog.text
# Health check must not run when Core wasn't running on entry.
mock_get_config.assert_not_called()
mock_frontend.assert_not_called()
# Caller (restore flow) is responsible for cleanup later.
mock_cleanup.assert_not_called()
assert (
Issue(IssueType.UPDATE_ROLLBACK, ContextType.CORE)
not in coresys.resolution.issues
)
assert coresys.homeassistant.version == AwesomeVersion("2026.4.4")