Files
supervisor/tests/test_supervisor.py
T
1a7853de9a Start the Supervisor auto update on every version reload (#7201)
* Start the Supervisor auto update on every version reload

A pending Supervisor update blocks Core, OS and app updates through the SUPERVISOR_UPDATED job condition. Only the startup and the daily scheduled reload started the auto update. Every other reload, such as the one Core requests when the user opens Settings, made the new version known without installing it. The user then saw "supervisor needs to be updated first" until the daily task ran.

The updater now starts the auto update task after every reload while the system is running. The scheduled task calls the updater reload directly. The Supervisor update job rejects a concurrent update request, so a user requested update during the auto update fails cleanly instead of creating an update failed issue.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC

* Report success for an update request during a running Supervisor update

The auto update can already run when Core or a user requests the update. The requested outcome is underway, so the API reports success and the caller sees the result through the Supervisor restart.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC

* Trim comments around the Supervisor auto update

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC

* Drop comment on the updater auto update trigger

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC

* Raise specific errors when restoring a backup requires a Supervisor update

Move the supervisor version check in do_restore_full/do_restore_partial into
a shared _check_supervisor_version helper. If the backup requires a newer
Supervisor version:
- Raise the new translatable BackupSupervisorVersionError (400) with the
  existing message when auto update is disabled.
- Otherwise kick off the Supervisor auto update in the background via
  sys_create_task(auto_update_supervisor()) and raise the new translatable
  BackupSupervisorUpdateInProgressError (503), telling the caller to retry
  once the update completes.

* Recheck for Supervisor update before failing a version-mismatch restore

When a restore requires a newer Supervisor version and auto update is
enabled, refresh version info first in case a newer Supervisor was just
released. If an update is already known to be needed, make sure it gets
kicked off directly instead of waiting on a reload. Only fall back to
the translatable version-mismatch error if no update is available after
the recheck.

Also move auto_update_supervisor from Tasks to the Supervisor class so
it can be reused outside of the task scheduler.

* Track auto update and version fetch tasks to close restore-check races

Supervisor.auto_update_supervisor() now stores and returns its update
task (created with eager_start so a synchronous failure is reflected
immediately) instead of firing-and-forgetting it. This lets callers
check the task's actual done state rather than guessing based on
need_update, which could be wrong if the update job rejected the call
for an unrelated reason.

HomeAssistantCore's install retry loop now also catches
SupervisorJobError when triggering the Supervisor update it depends on,
so it keeps retrying while an update is already in progress instead of
giving up.

Add Updater.start_fetch_data(), which starts (or reuses) a version
fetch and returns its task. fetch_data() sets its throttle timestamp
before running, even on failure, so a caller relying on fetch_data()
directly can't tell "someone else just refreshed" from "someone else's
fetch just failed" - it would silently skip either way. Callers that
need to know the true outcome should await the shared task instead.
reload() and BackupManager._check_supervisor_version() are updated to
use it: the latter awaits the fetch task directly (skipping it entirely
if there's no connectivity to fetch with) before checking
auto_update_supervisor()'s task, closing a race where a stale/failed
throttle window could cause a 400 (version mismatch) response instead
of a 503 (update in progress).

* Use Job decorator's detach option for update and version fetch tasks

#7211 added a `detach` option to the Job decorator, making the manual
task-tracking previously added here redundant. Replace it with the
decorator-native option:

- `Supervisor.update()` is now `detach=True` with
  `concurrency=JobConcurrency.REJECT`. Calling it returns the task
  performing the update - which may already be in progress from any
  previous caller, manual or automatic - instead of raising
  `SupervisorJobError` or blocking the caller until it (and the
  Supervisor restart it triggers) completes. Errors are raised on the
  returned task, not to the immediate caller.
- `Supervisor.auto_update_supervisor()` is now a thin, non-detached
  method that simply returns `update()`'s own task directly instead of
  tracking a separate task itself. This ensures a concurrent manual
  update and an auto update share the exact same underlying task
  regardless of which caller started it. It never awaits the task to
  completion, since update() restarts Supervisor and awaiting that here
  could drop the caller's connection before a response is sent.
- `Updater.fetch_data()` is now `detach=True` with
  `concurrency=JobConcurrency.REJECT`, replacing the bespoke
  `start_fetch_data()`/`_fetch_task` tracking. `reload()` and
  `BackupManager._check_supervisor_version()` await the returned task
  directly to get the real outcome of a fetch instead of relying on the
  throttle window.
- The Supervisor update API endpoint and the Home Assistant Core
  install retry loop now use `if task := await ...: await task` instead
  of catching `SupervisorJobError`, since a concurrent call no longer
  raises that error - it returns the shared in-progress task instead.

Updated tests across supervisor.py, updater.py, the API and Home
Assistant Core install retry loop to match the new return values and
removed now-obsolete `SupervisorJobError`-based concurrency tests.

* Fix too-many-nested-blocks pylint warning in HomeAssistantCore.install()

Flatten the nested need_update/auto_update if/else into an if/elif so
the Supervisor update retry try/except isn't nested one level deeper,
which pylint 4.0.8 (unlike the previously installed version in this
container) now flags as too-many-nested-blocks (6/5). Cache
need_update in a local variable since it's now referenced twice in the
same iteration and must be consistent between both checks.

* Update supervisor/api/supervisor.py

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Mike Degatano <michael.degatano@gmail.com>
Co-authored-by: Stefan Agner <stefan@agner.ch>
2026-09-14 15:56:21 +02:00

350 lines
12 KiB
Python

"""Test supervisor object."""
import asyncio
import errno
from unittest.mock import AsyncMock, MagicMock, Mock, patch
from aiohttp import ClientTimeout
from aiohttp.client_exceptions import ClientError
from awesomeversion import AwesomeVersion
import pytest
from supervisor.const import BusEvent, CoreState, UpdateChannel
from supervisor.coresys import CoreSys
from supervisor.docker.supervisor import DockerSupervisor
from supervisor.exceptions import (
DockerError,
SupervisorAppArmorError,
SupervisorUpdateError,
)
from supervisor.host.apparmor import AppArmorControl
from supervisor.resolution.const import ContextType, IssueType
from supervisor.resolution.data import Issue
from tests.common import MockResponse
@pytest.mark.parametrize(
("side_effect", "connectivity"), [(ClientError(), False), (None, True)]
)
async def test_connectivity_check(
coresys: CoreSys,
websession: MagicMock,
side_effect: Exception | None,
connectivity: bool,
):
"""Test connectivity check updates state based on probe outcome."""
assert coresys.supervisor.connectivity is True
websession.head = AsyncMock(side_effect=side_effect)
await coresys.supervisor.check_and_update_connectivity(force=True)
assert coresys.supervisor.connectivity is connectivity
async def test_connectivity_check_min_interval_when_connected(
coresys: CoreSys, websession: MagicMock
):
"""Non-forced checks within the min-interval use the cached state."""
websession.head = AsyncMock()
# First call runs the probe.
await coresys.supervisor.check_and_update_connectivity()
assert websession.head.call_count == 1
# Second call within the (10 min) window should not hit the network.
await coresys.supervisor.check_and_update_connectivity()
assert websession.head.call_count == 1
async def test_connectivity_check_force_bypasses_min_interval(
coresys: CoreSys, websession: MagicMock
):
"""force=True skips the min-interval short-circuit."""
websession.head = AsyncMock()
await coresys.supervisor.check_and_update_connectivity()
assert websession.head.call_count == 1
await coresys.supervisor.check_and_update_connectivity(force=True)
assert websession.head.call_count == 2
async def test_connectivity_check_coalesces_concurrent_callers(
coresys: CoreSys, websession: MagicMock
):
"""Concurrent callers await the same in-flight probe instead of each firing one."""
probe_started = asyncio.Event()
probe_release = asyncio.Event()
async def slow_head(*args, **kwargs):
probe_started.set()
await probe_release.wait()
websession.head = AsyncMock(side_effect=slow_head)
first = asyncio.create_task(
coresys.supervisor.check_and_update_connectivity(force=True)
)
await probe_started.wait()
# Kick off a pile of additional callers while the first probe is in flight.
concurrent = [
asyncio.create_task(coresys.supervisor.check_and_update_connectivity())
for _ in range(5)
]
# Let them all reach the in-flight await.
await asyncio.sleep(0)
probe_release.set()
await asyncio.gather(first, *concurrent)
assert websession.head.call_count == 1
async def test_connectivity_check_force_during_in_flight_triggers_rerun(
coresys: CoreSys, websession: MagicMock
):
"""A force signal arriving while a probe is in flight queues exactly one rerun."""
probe_started = asyncio.Event()
probe_release = asyncio.Event()
async def first_then_fast(*args, **kwargs):
if websession.head.call_count == 1:
probe_started.set()
await probe_release.wait()
websession.head = AsyncMock(side_effect=first_then_fast)
first = asyncio.create_task(
coresys.supervisor.check_and_update_connectivity(force=True)
)
await probe_started.wait()
# Forced call while a probe is in flight should set the rerun flag.
forced = asyncio.create_task(
coresys.supervisor.check_and_update_connectivity(force=True)
)
# Non-forced calls must NOT queue a rerun.
cheap = asyncio.create_task(coresys.supervisor.check_and_update_connectivity())
await asyncio.sleep(0)
probe_release.set()
await asyncio.gather(first, forced, cheap)
assert websession.head.call_count == 2
async def test_connectivity_check_owner_cancellation_cancels_probe(
coresys: CoreSys, websession: MagicMock
):
"""Owner cancellation propagates to the probe and skips updating last-check."""
probe_started = asyncio.Event()
probe_release = asyncio.Event()
async def slow_head(*args, **kwargs):
probe_started.set()
await probe_release.wait()
websession.head = AsyncMock(side_effect=slow_head)
last_check_before = coresys.supervisor._connectivity_last_check # pylint: disable=protected-access
owner = asyncio.create_task(
coresys.supervisor.check_and_update_connectivity(force=True)
)
await probe_started.wait()
owner.cancel()
with pytest.raises(asyncio.CancelledError):
await owner
# Owner cancellation must cancel the spawned probe, not orphan it,
# and the cached last-check timestamp must NOT advance.
assert coresys.supervisor._connectivity_check is None # pylint: disable=protected-access
assert coresys.supervisor._connectivity_last_check == last_check_before # pylint: disable=protected-access
# A subsequent non-forced call must therefore still run a probe.
websession.head = AsyncMock()
await coresys.supervisor.check_and_update_connectivity()
assert websession.head.call_count == 1
async def test_update_connectivity_fires_event_on_change(coresys: CoreSys):
"""SUPERVISOR_CONNECTIVITY_CHANGE fires only when the cached value changes."""
events: list[bool] = []
async def listener(state: bool) -> None:
events.append(state)
coresys.bus.register_event(BusEvent.SUPERVISOR_CONNECTIVITY_CHANGE, listener)
# Same value: no event.
coresys.supervisor._update_connectivity(True) # pylint: disable=protected-access
# Change to False: one event.
coresys.supervisor._update_connectivity(False) # pylint: disable=protected-access
# Change back to True: another event.
coresys.supervisor._update_connectivity(True) # pylint: disable=protected-access
await asyncio.sleep(0)
assert events == [False, True]
async def test_request_connectivity_check_is_fire_and_forget(
coresys: CoreSys, websession: MagicMock
):
"""request_connectivity_check schedules a check that runs asynchronously."""
websession.head = AsyncMock()
# Synchronous call must return without awaiting the HTTP probe.
result = coresys.supervisor.request_connectivity_check(force=True)
assert result is None
# Yield until the scheduled task has had a chance to complete.
for _ in range(5):
await asyncio.sleep(0)
assert websession.head.call_count == 1
async def test_update_failed(coresys: CoreSys, capture_exception: Mock):
"""Test update failure."""
# pylint: disable-next=protected-access
coresys.updater._data.setdefault("image", {})["supervisor"] = (
"ghcr.io/home-assistant/aarch64-hassio-supervisor"
)
err = DockerError()
with (
patch.object(DockerSupervisor, "install", side_effect=err),
patch.object(type(coresys.supervisor), "update_apparmor"),
):
task = await coresys.supervisor.update(AwesomeVersion("1.0"))
assert task
with pytest.raises(SupervisorUpdateError):
await task
capture_exception.assert_called_once_with(err)
assert (
Issue(IssueType.UPDATE_FAILED, ContextType.SUPERVISOR)
in coresys.resolution.issues
)
async def test_update_returns_running_task_for_concurrent_call(coresys: CoreSys):
"""Test a second update call while one is running gets back the same task."""
# pylint: disable-next=protected-access
coresys.updater._data.setdefault("image", {})["supervisor"] = (
"ghcr.io/home-assistant/aarch64-hassio-supervisor"
)
install_started = asyncio.Event()
install_finish = asyncio.Event()
async def install(*args, **kwargs):
install_started.set()
await install_finish.wait()
with (
patch.object(DockerSupervisor, "install", side_effect=install),
patch.object(DockerSupervisor, "update_start_tag"),
patch.object(type(coresys.supervisor), "update_apparmor"),
patch.object(type(coresys.core), "stop"),
):
first_task = await coresys.supervisor.update(AwesomeVersion("1.0"))
await install_started.wait()
assert first_task
assert not first_task.done()
second_task = await coresys.supervisor.update(AwesomeVersion("1.0"))
assert second_task is first_task
install_finish.set()
await first_task
@pytest.mark.parametrize(
"channel", [UpdateChannel.STABLE, UpdateChannel.BETA, UpdateChannel.DEV]
)
async def test_update_apparmor(
coresys: CoreSys, channel: UpdateChannel, websession: MagicMock, tmp_supervisor_data
):
"""Test updating apparmor."""
websession.get = Mock(return_value=MockResponse())
coresys.updater.channel = channel
with (
patch.object(AppArmorControl, "load_profile") as load_profile,
):
await coresys.supervisor.update_apparmor()
websession.get.assert_called_once_with(
f"https://version.home-assistant.io/apparmor_{channel}.txt",
timeout=ClientTimeout(total=10),
)
load_profile.assert_called_once()
async def test_update_apparmor_error(
coresys: CoreSys, websession: MagicMock, tmp_supervisor_data
):
"""Test error updating apparmor profile."""
websession.get = Mock(return_value=MockResponse())
with (
patch.object(AppArmorControl, "load_profile"),
patch("supervisor.supervisor.Path.write_text", side_effect=(err := OSError())),
):
err.errno = errno.EBUSY
with pytest.raises(SupervisorAppArmorError):
await coresys.supervisor.update_apparmor()
assert coresys.core.healthy is True
err.errno = errno.EBADMSG
with pytest.raises(SupervisorAppArmorError):
await coresys.supervisor.update_apparmor()
assert coresys.core.healthy is False
async def test_restart_returns_once_requests_rejected(coresys: CoreSys):
"""Test restart returns while stopping, after STOPPING state is entered.
The API responds to /supervisor/restart once restart() returns. The state
must already be STOPPING at that point so the system validation middleware
rejects requests sent after the response instead of accepting work that
the stop sequence kills mid-request.
"""
await coresys.core.set_state(CoreState.RUNNING)
teardown_release = asyncio.Event()
loop_stop_called = asyncio.Event()
async def blocked_api_stop():
await teardown_release.wait()
def blocked_loop_stop() -> None:
loop_stop_called.set()
coresys._websession = AsyncMock() # pylint: disable=protected-access
with (
patch.object(coresys.api, "stop", new=blocked_api_stop),
patch.object(coresys.scheduler, "shutdown", new=AsyncMock()),
patch.object(coresys.docker, "unload", new=AsyncMock()),
patch.object(coresys.homeassistant.api, "close", new=AsyncMock()),
patch.object(coresys.ingress, "unload", new=AsyncMock()),
patch.object(coresys.hardware, "unload", new=AsyncMock()),
patch.object(coresys.dbus, "unload", new=AsyncMock()),
patch.object(coresys.loop, "stop", side_effect=blocked_loop_stop),
):
await coresys.supervisor.restart()
# Restart returned while the teardown is still in flight
assert coresys.core.state == CoreState.STOPPING
assert coresys.core.exit_code == 100
# Let the stop sequence finish
teardown_release.set()
async with asyncio.timeout(1):
while coresys.core.state != CoreState.CLOSE:
await asyncio.sleep(0)
# Ensure the stop task reaches loop.stop while it is still patched,
# otherwise the real loop.stop can run after this context exits.
async with asyncio.timeout(1):
await loop_stop_called.wait()