Files
supervisor/tests/misc/test_tasks.py
T
1a7853de9a Start the Supervisor auto update on every version reload (#7201)
* Start the Supervisor auto update on every version reload

A pending Supervisor update blocks Core, OS and app updates through the SUPERVISOR_UPDATED job condition. Only the startup and the daily scheduled reload started the auto update. Every other reload, such as the one Core requests when the user opens Settings, made the new version known without installing it. The user then saw "supervisor needs to be updated first" until the daily task ran.

The updater now starts the auto update task after every reload while the system is running. The scheduled task calls the updater reload directly. The Supervisor update job rejects a concurrent update request, so a user requested update during the auto update fails cleanly instead of creating an update failed issue.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC

* Report success for an update request during a running Supervisor update

The auto update can already run when Core or a user requests the update. The requested outcome is underway, so the API reports success and the caller sees the result through the Supervisor restart.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC

* Trim comments around the Supervisor auto update

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC

* Drop comment on the updater auto update trigger

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SAWqYzDiFEvYjZwfBWoxkC

* Raise specific errors when restoring a backup requires a Supervisor update

Move the supervisor version check in do_restore_full/do_restore_partial into
a shared _check_supervisor_version helper. If the backup requires a newer
Supervisor version:
- Raise the new translatable BackupSupervisorVersionError (400) with the
  existing message when auto update is disabled.
- Otherwise kick off the Supervisor auto update in the background via
  sys_create_task(auto_update_supervisor()) and raise the new translatable
  BackupSupervisorUpdateInProgressError (503), telling the caller to retry
  once the update completes.

* Recheck for Supervisor update before failing a version-mismatch restore

When a restore requires a newer Supervisor version and auto update is
enabled, refresh version info first in case a newer Supervisor was just
released. If an update is already known to be needed, make sure it gets
kicked off directly instead of waiting on a reload. Only fall back to
the translatable version-mismatch error if no update is available after
the recheck.

Also move auto_update_supervisor from Tasks to the Supervisor class so
it can be reused outside of the task scheduler.

* Track auto update and version fetch tasks to close restore-check races

Supervisor.auto_update_supervisor() now stores and returns its update
task (created with eager_start so a synchronous failure is reflected
immediately) instead of firing-and-forgetting it. This lets callers
check the task's actual done state rather than guessing based on
need_update, which could be wrong if the update job rejected the call
for an unrelated reason.

HomeAssistantCore's install retry loop now also catches
SupervisorJobError when triggering the Supervisor update it depends on,
so it keeps retrying while an update is already in progress instead of
giving up.

Add Updater.start_fetch_data(), which starts (or reuses) a version
fetch and returns its task. fetch_data() sets its throttle timestamp
before running, even on failure, so a caller relying on fetch_data()
directly can't tell "someone else just refreshed" from "someone else's
fetch just failed" - it would silently skip either way. Callers that
need to know the true outcome should await the shared task instead.
reload() and BackupManager._check_supervisor_version() are updated to
use it: the latter awaits the fetch task directly (skipping it entirely
if there's no connectivity to fetch with) before checking
auto_update_supervisor()'s task, closing a race where a stale/failed
throttle window could cause a 400 (version mismatch) response instead
of a 503 (update in progress).

* Use Job decorator's detach option for update and version fetch tasks

#7211 added a `detach` option to the Job decorator, making the manual
task-tracking previously added here redundant. Replace it with the
decorator-native option:

- `Supervisor.update()` is now `detach=True` with
  `concurrency=JobConcurrency.REJECT`. Calling it returns the task
  performing the update - which may already be in progress from any
  previous caller, manual or automatic - instead of raising
  `SupervisorJobError` or blocking the caller until it (and the
  Supervisor restart it triggers) completes. Errors are raised on the
  returned task, not to the immediate caller.
- `Supervisor.auto_update_supervisor()` is now a thin, non-detached
  method that simply returns `update()`'s own task directly instead of
  tracking a separate task itself. This ensures a concurrent manual
  update and an auto update share the exact same underlying task
  regardless of which caller started it. It never awaits the task to
  completion, since update() restarts Supervisor and awaiting that here
  could drop the caller's connection before a response is sent.
- `Updater.fetch_data()` is now `detach=True` with
  `concurrency=JobConcurrency.REJECT`, replacing the bespoke
  `start_fetch_data()`/`_fetch_task` tracking. `reload()` and
  `BackupManager._check_supervisor_version()` await the returned task
  directly to get the real outcome of a fetch instead of relying on the
  throttle window.
- The Supervisor update API endpoint and the Home Assistant Core
  install retry loop now use `if task := await ...: await task` instead
  of catching `SupervisorJobError`, since a concurrent call no longer
  raises that error - it returns the shared in-progress task instead.

Updated tests across supervisor.py, updater.py, the API and Home
Assistant Core install retry loop to match the new return values and
removed now-obsolete `SupervisorJobError`-based concurrency tests.

* Fix too-many-nested-blocks pylint warning in HomeAssistantCore.install()

Flatten the nested need_update/auto_update if/else into an if/elif so
the Supervisor update retry try/except isn't nested one level deeper,
which pylint 4.0.8 (unlike the previously installed version in this
container) now flags as too-many-nested-blocks (6/5). Cache
need_update in a local variable since it's now referenced twice in the
same iteration and must be consistent between both checks.

* Update supervisor/api/supervisor.py

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Mike Degatano <michael.degatano@gmail.com>
Co-authored-by: Stefan Agner <stefan@agner.ch>
2026-09-14 15:56:21 +02:00

311 lines
11 KiB
Python

"""Test scheduled tasks."""
import asyncio
from collections.abc import AsyncGenerator
from shutil import copy
from unittest.mock import AsyncMock, Mock, PropertyMock, patch
from aiodocker.containers import DockerContainer
from awesomeversion import AwesomeVersion
import pytest
from supervisor.apps.app import App
from supervisor.const import ATTR_VERSION_TIMESTAMP, CoreState, FeatureFlag
from supervisor.coresys import CoreSys
from supervisor.exceptions import HomeAssistantError
from supervisor.homeassistant.api import HomeAssistantAPI
from supervisor.homeassistant.const import LANDINGPAGE
from supervisor.homeassistant.core import HomeAssistantCore
from supervisor.misc.tasks import Tasks
from supervisor.plugins.dns import PluginDns
from supervisor.supervisor import Supervisor
from tests.common import MockResponse, get_fixture_path
# pylint: disable=protected-access
@pytest.fixture(name="tasks")
async def fixture_tasks(
coresys: CoreSys, container: DockerContainer
) -> AsyncGenerator[Tasks]:
"""Return task manager."""
coresys.homeassistant.watchdog = True
coresys.homeassistant.version = AwesomeVersion("2023.12.0")
container.show.return_value["State"]["Status"] = "running"
container.show.return_value["State"]["Running"] = True
return Tasks(coresys)
async def test_watchdog_homeassistant_api(
tasks: Tasks, caplog: pytest.LogCaptureFixture
):
"""Test watchdog of homeassistant api."""
with (
patch.object(HomeAssistantAPI, "check_api_state", return_value=False),
patch.object(HomeAssistantCore, "restart") as restart,
):
await tasks._watchdog_homeassistant_api()
restart.assert_not_called()
assert "Watchdog missed an Home Assistant Core API response." in caplog.text
assert (
"Watchdog missed 2 Home Assistant Core API responses in a row. Restarting Home Assistant Core API!"
not in caplog.text
)
caplog.clear()
await tasks._watchdog_homeassistant_api()
restart.assert_called_once()
assert "Watchdog missed an Home Assistant Core API response." not in caplog.text
assert (
"Watchdog missed 2 Home Assistant Core API responses in a row. Restarting Home Assistant Core!"
in caplog.text
)
async def test_watchdog_homeassistant_api_off(tasks: Tasks, coresys: CoreSys):
"""Test watchdog of homeassistant api does not run when disabled."""
coresys.homeassistant.watchdog = False
with (
patch.object(HomeAssistantAPI, "check_api_state", return_value=False),
patch.object(HomeAssistantCore, "restart") as restart,
):
await tasks._watchdog_homeassistant_api()
await tasks._watchdog_homeassistant_api()
restart.assert_not_called()
async def test_watchdog_homeassistant_api_error_state(tasks: Tasks, coresys: CoreSys):
"""Test watchdog of homeassistant api does not restart when in error state."""
coresys.homeassistant.core._error_state = True
with (
patch.object(HomeAssistantAPI, "check_api_state", return_value=False),
patch.object(HomeAssistantCore, "restart") as restart,
):
await tasks._watchdog_homeassistant_api()
await tasks._watchdog_homeassistant_api()
restart.assert_not_called()
async def test_watchdog_homeassistant_api_landing_page(tasks: Tasks, coresys: CoreSys):
"""Test watchdog of homeassistant api does not monitor landing page."""
coresys.homeassistant.version = LANDINGPAGE
with (
patch.object(HomeAssistantAPI, "check_api_state", return_value=False),
patch.object(HomeAssistantCore, "restart") as restart,
):
await tasks._watchdog_homeassistant_api()
await tasks._watchdog_homeassistant_api()
restart.assert_not_called()
async def test_watchdog_homeassistant_api_not_running(
tasks: Tasks, container: DockerContainer
):
"""Test watchdog of homeassistant api does not monitor when home assistant not running."""
container.show.return_value["State"]["Status"] = "stopped"
container.show.return_value["State"]["Running"] = False
with (
patch.object(HomeAssistantAPI, "check_api_state", return_value=False),
patch.object(HomeAssistantCore, "restart") as restart,
):
await tasks._watchdog_homeassistant_api()
await tasks._watchdog_homeassistant_api()
restart.assert_not_called()
async def test_watchdog_homeassistant_api_reanimation_limit(
tasks: Tasks, caplog: pytest.LogCaptureFixture, capture_exception: Mock
):
"""Test watchdog of homeassistant api stops after max reanimation failures."""
with (
patch.object(HomeAssistantAPI, "check_api_state", return_value=False),
patch.object(
HomeAssistantCore, "restart", side_effect=(err := HomeAssistantError())
) as restart,
patch.object(HomeAssistantCore, "rebuild", side_effect=err) as rebuild,
):
for _ in range(5):
await tasks._watchdog_homeassistant_api()
restart.assert_not_called()
await tasks._watchdog_homeassistant_api()
restart.assert_called_once_with()
assert "Home Assistant watchdog reanimation failed!" in caplog.text
rebuild.assert_not_called()
restart.reset_mock()
capture_exception.assert_called_once_with(err)
# Next time it should try safe mode
caplog.clear()
await tasks._watchdog_homeassistant_api()
rebuild.assert_not_called()
await tasks._watchdog_homeassistant_api()
rebuild.assert_called_once_with(safe_mode=True)
restart.assert_not_called()
assert (
"Watchdog cannot reanimate Home Assistant Core, failed all 5 attempts. Restarting into safe mode"
in caplog.text
)
assert (
"Safe mode restart failed. Watchdog cannot bring Home Assistant online."
in caplog.text
)
# After safe mode has failed too, no more restart attempts
rebuild.reset_mock()
caplog.clear()
await tasks._watchdog_homeassistant_api()
assert "Watchdog missed an Home Assistant Core API response." in caplog.text
caplog.clear()
await tasks._watchdog_homeassistant_api()
assert not caplog.text
restart.assert_not_called()
rebuild.assert_not_called()
@pytest.mark.usefixtures("path_extern", "tmp_supervisor_data")
async def test_core_backup_cleanup(tasks: Tasks, coresys: CoreSys):
"""Test core backup task cleans up old backup files."""
await coresys.core.set_state(CoreState.RUNNING)
coresys.hardware.disk.get_disk_free_space = lambda x: 5000
# Put an old and new backup in folder
copy(get_fixture_path("backup_example.tar"), coresys.config.path_core_backup)
await coresys.backups.reload()
assert (old_backup := coresys.backups.get("7fed74c8"))
new_backup = await coresys.backups.do_backup_partial(
name="test", folders=["ssl"], location=".cloud_backup"
)
old_tar = old_backup.tarfile
new_tar = new_backup.tarfile
# pylint: disable-next=protected-access
await tasks._core_backup_cleanup()
assert coresys.backups.get(new_backup.slug)
assert not coresys.backups.get("7fed74c8")
assert new_tar.exists()
assert not old_tar.exists()
@pytest.mark.usefixtures("no_job_throttle")
async def test_update_dns_skipped_when_auto_update_disabled(
tasks: Tasks, coresys: CoreSys
):
"""Test plugin auto-update task is skipped when auto update is disabled."""
await coresys.core.set_state(CoreState.RUNNING)
coresys.hardware.disk.get_disk_free_space = lambda x: 5000
coresys.updater.auto_update = False
with patch.object(PluginDns, "update") as update:
await tasks._update_dns()
update.assert_not_called()
@pytest.mark.usefixtures("no_job_throttle", "supervisor_internet")
async def test_scheduled_reload_updater_triggers_one_supervisor_update(
tasks: Tasks, coresys: CoreSys, mock_update_data: MockResponse
):
"""Test the scheduled updater reload triggers exactly one supervisor update."""
coresys.hardware.disk.get_disk_free_space = lambda x: 5000
await coresys.core.set_state(CoreState.RUNNING)
# Make version data show a newer supervisor version
version_data = await mock_update_data.text()
mock_update_data.update_text(version_data.replace("2024.10.0", "2024.10.1"))
with (
patch.object(
Supervisor,
"version",
new=PropertyMock(return_value=AwesomeVersion("2024.10.0")),
),
patch.object(Supervisor, "update") as update,
):
await tasks.load()
update.assert_not_called()
# Advance the event loop clock by 24h+ so scheduled tasks fire.
# Patching loop.time makes all call_later callbacks appear due;
# a tiny real sleep lets _run_once re-evaluate and execute them.
loop = asyncio.get_event_loop()
original_time = loop.time
loop.time = lambda: original_time() + 86401
try:
# Busy-wait until call_later callbacks fire and create jobs
while not any(t.job and not t.job.done() for t in coresys.scheduler._tasks):
await asyncio.sleep(0)
# Wait for all scheduler-created tasks to finish
pending = [
t.job for t in coresys.scheduler._tasks if t.job and not t.job.done()
]
await asyncio.gather(*pending)
async with asyncio.timeout(5):
while not update.called:
await asyncio.sleep(0)
# Verify update was triggered exactly once
update.assert_called_once()
finally:
loop.time = original_time
await coresys.scheduler.shutdown()
@pytest.mark.usefixtures("tmp_supervisor_data")
@pytest.mark.parametrize(
("websocket_v2_enabled", "expected_message_type", "expected_slug_key"),
[(False, "hassio/update/addon", "addon"), (True, "hassio/update/app", "app")],
)
async def test_update_apps_auto_update_success(
tasks: Tasks,
ha_ws_client: AsyncMock,
install_app_example: App,
websocket_v2_enabled: bool,
expected_message_type: str,
expected_slug_key: str,
):
"""Test that an eligible app is auto-updated via websocket command."""
await tasks.sys_core.set_state(CoreState.RUNNING)
tasks.sys_config.set_feature_flag(
FeatureFlag.SUPERVISOR_WEBSOCKET_V2_API, websocket_v2_enabled
)
# Set up the app as eligible for auto-update
install_app_example.auto_update = True
install_app_example.data_store[ATTR_VERSION_TIMESTAMP] = 0
with patch.object(
App, "version", new=PropertyMock(return_value=AwesomeVersion("1.0"))
):
assert install_app_example.need_update is True
assert install_app_example.auto_update_available is True
# Make sure all job events from installing the app are cleared
ha_ws_client.async_send_command.reset_mock()
# pylint: disable-next=protected-access
await tasks._update_apps()
ha_ws_client.async_send_command.assert_any_call(
{
"type": expected_message_type,
expected_slug_key: install_app_example.slug,
"backup": True,
}
)