mirror of
https://github.com/home-assistant/supervisor.git
synced 2026-09-30 05:55:03 +01:00
* Start the Supervisor auto update on every version reload A pending Supervisor update blocks Core, OS and app updates through the SUPERVISOR_UPDATED job condition. Only the startup and the daily scheduled reload started the auto update. Every other reload, such as the one Core requests when the user opens Settings, made the new version known without installing it. The user then saw "supervisor needs to be updated first" until the daily task ran. The updater now starts the auto update task after every reload while the system is running. The scheduled task calls the updater reload directly. The Supervisor update job rejects a concurrent update request, so a user requested update during the auto update fails cleanly instead of creating an update failed issue. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC * Report success for an update request during a running Supervisor update The auto update can already run when Core or a user requests the update. The requested outcome is underway, so the API reports success and the caller sees the result through the Supervisor restart. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC * Trim comments around the Supervisor auto update Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC * Drop comment on the updater auto update trigger Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC * Raise specific errors when restoring a backup requires a Supervisor update Move the supervisor version check in do_restore_full/do_restore_partial into a shared _check_supervisor_version helper. If the backup requires a newer Supervisor version: - Raise the new translatable BackupSupervisorVersionError (400) with the existing message when auto update is disabled. - Otherwise kick off the Supervisor auto update in the background via sys_create_task(auto_update_supervisor()) and raise the new translatable BackupSupervisorUpdateInProgressError (503), telling the caller to retry once the update completes. * Recheck for Supervisor update before failing a version-mismatch restore When a restore requires a newer Supervisor version and auto update is enabled, refresh version info first in case a newer Supervisor was just released. If an update is already known to be needed, make sure it gets kicked off directly instead of waiting on a reload. Only fall back to the translatable version-mismatch error if no update is available after the recheck. Also move auto_update_supervisor from Tasks to the Supervisor class so it can be reused outside of the task scheduler. * Track auto update and version fetch tasks to close restore-check races Supervisor.auto_update_supervisor() now stores and returns its update task (created with eager_start so a synchronous failure is reflected immediately) instead of firing-and-forgetting it. This lets callers check the task's actual done state rather than guessing based on need_update, which could be wrong if the update job rejected the call for an unrelated reason. HomeAssistantCore's install retry loop now also catches SupervisorJobError when triggering the Supervisor update it depends on, so it keeps retrying while an update is already in progress instead of giving up. Add Updater.start_fetch_data(), which starts (or reuses) a version fetch and returns its task. fetch_data() sets its throttle timestamp before running, even on failure, so a caller relying on fetch_data() directly can't tell "someone else just refreshed" from "someone else's fetch just failed" - it would silently skip either way. Callers that need to know the true outcome should await the shared task instead. reload() and BackupManager._check_supervisor_version() are updated to use it: the latter awaits the fetch task directly (skipping it entirely if there's no connectivity to fetch with) before checking auto_update_supervisor()'s task, closing a race where a stale/failed throttle window could cause a 400 (version mismatch) response instead of a 503 (update in progress). * Use Job decorator's detach option for update and version fetch tasks #7211 added a `detach` option to the Job decorator, making the manual task-tracking previously added here redundant. Replace it with the decorator-native option: - `Supervisor.update()` is now `detach=True` with `concurrency=JobConcurrency.REJECT`. Calling it returns the task performing the update - which may already be in progress from any previous caller, manual or automatic - instead of raising `SupervisorJobError` or blocking the caller until it (and the Supervisor restart it triggers) completes. Errors are raised on the returned task, not to the immediate caller. - `Supervisor.auto_update_supervisor()` is now a thin, non-detached method that simply returns `update()`'s own task directly instead of tracking a separate task itself. This ensures a concurrent manual update and an auto update share the exact same underlying task regardless of which caller started it. It never awaits the task to completion, since update() restarts Supervisor and awaiting that here could drop the caller's connection before a response is sent. - `Updater.fetch_data()` is now `detach=True` with `concurrency=JobConcurrency.REJECT`, replacing the bespoke `start_fetch_data()`/`_fetch_task` tracking. `reload()` and `BackupManager._check_supervisor_version()` await the returned task directly to get the real outcome of a fetch instead of relying on the throttle window. - The Supervisor update API endpoint and the Home Assistant Core install retry loop now use `if task := await ...: await task` instead of catching `SupervisorJobError`, since a concurrent call no longer raises that error - it returns the shared in-progress task instead. Updated tests across supervisor.py, updater.py, the API and Home Assistant Core install retry loop to match the new return values and removed now-obsolete `SupervisorJobError`-based concurrency tests. * Fix too-many-nested-blocks pylint warning in HomeAssistantCore.install() Flatten the nested need_update/auto_update if/else into an if/elif so the Supervisor update retry try/except isn't nested one level deeper, which pylint 4.0.8 (unlike the previously installed version in this container) now flags as too-many-nested-blocks (6/5). Cache need_update in a local variable since it's now referenced twice in the same iteration and must be consistent between both checks. * Update supervisor/api/supervisor.py --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Mike Degatano <michael.degatano@gmail.com> Co-authored-by: Stefan Agner <stefan@agner.ch>
350 lines
12 KiB
Python
350 lines
12 KiB
Python
"""Test supervisor object."""
|
|
|
|
import asyncio
|
|
import errno
|
|
from unittest.mock import AsyncMock, MagicMock, Mock, patch
|
|
|
|
from aiohttp import ClientTimeout
|
|
from aiohttp.client_exceptions import ClientError
|
|
from awesomeversion import AwesomeVersion
|
|
import pytest
|
|
|
|
from supervisor.const import BusEvent, CoreState, UpdateChannel
|
|
from supervisor.coresys import CoreSys
|
|
from supervisor.docker.supervisor import DockerSupervisor
|
|
from supervisor.exceptions import (
|
|
DockerError,
|
|
SupervisorAppArmorError,
|
|
SupervisorUpdateError,
|
|
)
|
|
from supervisor.host.apparmor import AppArmorControl
|
|
from supervisor.resolution.const import ContextType, IssueType
|
|
from supervisor.resolution.data import Issue
|
|
|
|
from tests.common import MockResponse
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
("side_effect", "connectivity"), [(ClientError(), False), (None, True)]
|
|
)
|
|
async def test_connectivity_check(
|
|
coresys: CoreSys,
|
|
websession: MagicMock,
|
|
side_effect: Exception | None,
|
|
connectivity: bool,
|
|
):
|
|
"""Test connectivity check updates state based on probe outcome."""
|
|
assert coresys.supervisor.connectivity is True
|
|
|
|
websession.head = AsyncMock(side_effect=side_effect)
|
|
await coresys.supervisor.check_and_update_connectivity(force=True)
|
|
|
|
assert coresys.supervisor.connectivity is connectivity
|
|
|
|
|
|
async def test_connectivity_check_min_interval_when_connected(
|
|
coresys: CoreSys, websession: MagicMock
|
|
):
|
|
"""Non-forced checks within the min-interval use the cached state."""
|
|
websession.head = AsyncMock()
|
|
|
|
# First call runs the probe.
|
|
await coresys.supervisor.check_and_update_connectivity()
|
|
assert websession.head.call_count == 1
|
|
|
|
# Second call within the (10 min) window should not hit the network.
|
|
await coresys.supervisor.check_and_update_connectivity()
|
|
assert websession.head.call_count == 1
|
|
|
|
|
|
async def test_connectivity_check_force_bypasses_min_interval(
|
|
coresys: CoreSys, websession: MagicMock
|
|
):
|
|
"""force=True skips the min-interval short-circuit."""
|
|
websession.head = AsyncMock()
|
|
|
|
await coresys.supervisor.check_and_update_connectivity()
|
|
assert websession.head.call_count == 1
|
|
|
|
await coresys.supervisor.check_and_update_connectivity(force=True)
|
|
assert websession.head.call_count == 2
|
|
|
|
|
|
async def test_connectivity_check_coalesces_concurrent_callers(
|
|
coresys: CoreSys, websession: MagicMock
|
|
):
|
|
"""Concurrent callers await the same in-flight probe instead of each firing one."""
|
|
probe_started = asyncio.Event()
|
|
probe_release = asyncio.Event()
|
|
|
|
async def slow_head(*args, **kwargs):
|
|
probe_started.set()
|
|
await probe_release.wait()
|
|
|
|
websession.head = AsyncMock(side_effect=slow_head)
|
|
|
|
first = asyncio.create_task(
|
|
coresys.supervisor.check_and_update_connectivity(force=True)
|
|
)
|
|
await probe_started.wait()
|
|
|
|
# Kick off a pile of additional callers while the first probe is in flight.
|
|
concurrent = [
|
|
asyncio.create_task(coresys.supervisor.check_and_update_connectivity())
|
|
for _ in range(5)
|
|
]
|
|
# Let them all reach the in-flight await.
|
|
await asyncio.sleep(0)
|
|
|
|
probe_release.set()
|
|
await asyncio.gather(first, *concurrent)
|
|
|
|
assert websession.head.call_count == 1
|
|
|
|
|
|
async def test_connectivity_check_force_during_in_flight_triggers_rerun(
|
|
coresys: CoreSys, websession: MagicMock
|
|
):
|
|
"""A force signal arriving while a probe is in flight queues exactly one rerun."""
|
|
probe_started = asyncio.Event()
|
|
probe_release = asyncio.Event()
|
|
|
|
async def first_then_fast(*args, **kwargs):
|
|
if websession.head.call_count == 1:
|
|
probe_started.set()
|
|
await probe_release.wait()
|
|
|
|
websession.head = AsyncMock(side_effect=first_then_fast)
|
|
|
|
first = asyncio.create_task(
|
|
coresys.supervisor.check_and_update_connectivity(force=True)
|
|
)
|
|
await probe_started.wait()
|
|
|
|
# Forced call while a probe is in flight should set the rerun flag.
|
|
forced = asyncio.create_task(
|
|
coresys.supervisor.check_and_update_connectivity(force=True)
|
|
)
|
|
# Non-forced calls must NOT queue a rerun.
|
|
cheap = asyncio.create_task(coresys.supervisor.check_and_update_connectivity())
|
|
await asyncio.sleep(0)
|
|
|
|
probe_release.set()
|
|
await asyncio.gather(first, forced, cheap)
|
|
|
|
assert websession.head.call_count == 2
|
|
|
|
|
|
async def test_connectivity_check_owner_cancellation_cancels_probe(
|
|
coresys: CoreSys, websession: MagicMock
|
|
):
|
|
"""Owner cancellation propagates to the probe and skips updating last-check."""
|
|
probe_started = asyncio.Event()
|
|
probe_release = asyncio.Event()
|
|
|
|
async def slow_head(*args, **kwargs):
|
|
probe_started.set()
|
|
await probe_release.wait()
|
|
|
|
websession.head = AsyncMock(side_effect=slow_head)
|
|
last_check_before = coresys.supervisor._connectivity_last_check # pylint: disable=protected-access
|
|
|
|
owner = asyncio.create_task(
|
|
coresys.supervisor.check_and_update_connectivity(force=True)
|
|
)
|
|
await probe_started.wait()
|
|
|
|
owner.cancel()
|
|
with pytest.raises(asyncio.CancelledError):
|
|
await owner
|
|
|
|
# Owner cancellation must cancel the spawned probe, not orphan it,
|
|
# and the cached last-check timestamp must NOT advance.
|
|
assert coresys.supervisor._connectivity_check is None # pylint: disable=protected-access
|
|
assert coresys.supervisor._connectivity_last_check == last_check_before # pylint: disable=protected-access
|
|
|
|
# A subsequent non-forced call must therefore still run a probe.
|
|
websession.head = AsyncMock()
|
|
await coresys.supervisor.check_and_update_connectivity()
|
|
assert websession.head.call_count == 1
|
|
|
|
|
|
async def test_update_connectivity_fires_event_on_change(coresys: CoreSys):
|
|
"""SUPERVISOR_CONNECTIVITY_CHANGE fires only when the cached value changes."""
|
|
events: list[bool] = []
|
|
|
|
async def listener(state: bool) -> None:
|
|
events.append(state)
|
|
|
|
coresys.bus.register_event(BusEvent.SUPERVISOR_CONNECTIVITY_CHANGE, listener)
|
|
|
|
# Same value: no event.
|
|
coresys.supervisor._update_connectivity(True) # pylint: disable=protected-access
|
|
# Change to False: one event.
|
|
coresys.supervisor._update_connectivity(False) # pylint: disable=protected-access
|
|
# Change back to True: another event.
|
|
coresys.supervisor._update_connectivity(True) # pylint: disable=protected-access
|
|
await asyncio.sleep(0)
|
|
|
|
assert events == [False, True]
|
|
|
|
|
|
async def test_request_connectivity_check_is_fire_and_forget(
|
|
coresys: CoreSys, websession: MagicMock
|
|
):
|
|
"""request_connectivity_check schedules a check that runs asynchronously."""
|
|
websession.head = AsyncMock()
|
|
|
|
# Synchronous call must return without awaiting the HTTP probe.
|
|
result = coresys.supervisor.request_connectivity_check(force=True)
|
|
assert result is None
|
|
|
|
# Yield until the scheduled task has had a chance to complete.
|
|
for _ in range(5):
|
|
await asyncio.sleep(0)
|
|
|
|
assert websession.head.call_count == 1
|
|
|
|
|
|
async def test_update_failed(coresys: CoreSys, capture_exception: Mock):
|
|
"""Test update failure."""
|
|
# pylint: disable-next=protected-access
|
|
coresys.updater._data.setdefault("image", {})["supervisor"] = (
|
|
"ghcr.io/home-assistant/aarch64-hassio-supervisor"
|
|
)
|
|
err = DockerError()
|
|
with (
|
|
patch.object(DockerSupervisor, "install", side_effect=err),
|
|
patch.object(type(coresys.supervisor), "update_apparmor"),
|
|
):
|
|
task = await coresys.supervisor.update(AwesomeVersion("1.0"))
|
|
assert task
|
|
with pytest.raises(SupervisorUpdateError):
|
|
await task
|
|
|
|
capture_exception.assert_called_once_with(err)
|
|
assert (
|
|
Issue(IssueType.UPDATE_FAILED, ContextType.SUPERVISOR)
|
|
in coresys.resolution.issues
|
|
)
|
|
|
|
|
|
async def test_update_returns_running_task_for_concurrent_call(coresys: CoreSys):
|
|
"""Test a second update call while one is running gets back the same task."""
|
|
# pylint: disable-next=protected-access
|
|
coresys.updater._data.setdefault("image", {})["supervisor"] = (
|
|
"ghcr.io/home-assistant/aarch64-hassio-supervisor"
|
|
)
|
|
install_started = asyncio.Event()
|
|
install_finish = asyncio.Event()
|
|
|
|
async def install(*args, **kwargs):
|
|
install_started.set()
|
|
await install_finish.wait()
|
|
|
|
with (
|
|
patch.object(DockerSupervisor, "install", side_effect=install),
|
|
patch.object(DockerSupervisor, "update_start_tag"),
|
|
patch.object(type(coresys.supervisor), "update_apparmor"),
|
|
patch.object(type(coresys.core), "stop"),
|
|
):
|
|
first_task = await coresys.supervisor.update(AwesomeVersion("1.0"))
|
|
await install_started.wait()
|
|
assert first_task
|
|
assert not first_task.done()
|
|
|
|
second_task = await coresys.supervisor.update(AwesomeVersion("1.0"))
|
|
assert second_task is first_task
|
|
|
|
install_finish.set()
|
|
await first_task
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"channel", [UpdateChannel.STABLE, UpdateChannel.BETA, UpdateChannel.DEV]
|
|
)
|
|
async def test_update_apparmor(
|
|
coresys: CoreSys, channel: UpdateChannel, websession: MagicMock, tmp_supervisor_data
|
|
):
|
|
"""Test updating apparmor."""
|
|
websession.get = Mock(return_value=MockResponse())
|
|
coresys.updater.channel = channel
|
|
with (
|
|
patch.object(AppArmorControl, "load_profile") as load_profile,
|
|
):
|
|
await coresys.supervisor.update_apparmor()
|
|
|
|
websession.get.assert_called_once_with(
|
|
f"https://version.home-assistant.io/apparmor_{channel}.txt",
|
|
timeout=ClientTimeout(total=10),
|
|
)
|
|
load_profile.assert_called_once()
|
|
|
|
|
|
async def test_update_apparmor_error(
|
|
coresys: CoreSys, websession: MagicMock, tmp_supervisor_data
|
|
):
|
|
"""Test error updating apparmor profile."""
|
|
websession.get = Mock(return_value=MockResponse())
|
|
with (
|
|
patch.object(AppArmorControl, "load_profile"),
|
|
patch("supervisor.supervisor.Path.write_text", side_effect=(err := OSError())),
|
|
):
|
|
err.errno = errno.EBUSY
|
|
with pytest.raises(SupervisorAppArmorError):
|
|
await coresys.supervisor.update_apparmor()
|
|
assert coresys.core.healthy is True
|
|
|
|
err.errno = errno.EBADMSG
|
|
with pytest.raises(SupervisorAppArmorError):
|
|
await coresys.supervisor.update_apparmor()
|
|
assert coresys.core.healthy is False
|
|
|
|
|
|
async def test_restart_returns_once_requests_rejected(coresys: CoreSys):
|
|
"""Test restart returns while stopping, after STOPPING state is entered.
|
|
|
|
The API responds to /supervisor/restart once restart() returns. The state
|
|
must already be STOPPING at that point so the system validation middleware
|
|
rejects requests sent after the response instead of accepting work that
|
|
the stop sequence kills mid-request.
|
|
"""
|
|
await coresys.core.set_state(CoreState.RUNNING)
|
|
|
|
teardown_release = asyncio.Event()
|
|
loop_stop_called = asyncio.Event()
|
|
|
|
async def blocked_api_stop():
|
|
await teardown_release.wait()
|
|
|
|
def blocked_loop_stop() -> None:
|
|
loop_stop_called.set()
|
|
|
|
coresys._websession = AsyncMock() # pylint: disable=protected-access
|
|
with (
|
|
patch.object(coresys.api, "stop", new=blocked_api_stop),
|
|
patch.object(coresys.scheduler, "shutdown", new=AsyncMock()),
|
|
patch.object(coresys.docker, "unload", new=AsyncMock()),
|
|
patch.object(coresys.homeassistant.api, "close", new=AsyncMock()),
|
|
patch.object(coresys.ingress, "unload", new=AsyncMock()),
|
|
patch.object(coresys.hardware, "unload", new=AsyncMock()),
|
|
patch.object(coresys.dbus, "unload", new=AsyncMock()),
|
|
patch.object(coresys.loop, "stop", side_effect=blocked_loop_stop),
|
|
):
|
|
await coresys.supervisor.restart()
|
|
|
|
# Restart returned while the teardown is still in flight
|
|
assert coresys.core.state == CoreState.STOPPING
|
|
assert coresys.core.exit_code == 100
|
|
|
|
# Let the stop sequence finish
|
|
teardown_release.set()
|
|
async with asyncio.timeout(1):
|
|
while coresys.core.state != CoreState.CLOSE:
|
|
await asyncio.sleep(0)
|
|
|
|
# Ensure the stop task reaches loop.stop while it is still patched,
|
|
# otherwise the real loop.stop can run after this context exits.
|
|
async with asyncio.timeout(1):
|
|
await loop_stop_called.wait()
|