* Add one-shot mode and docker-given totals to stats API
- Add one-shot query option to container stats to skip the calculated
CPU percentage window, returning docker-given totals directly.
- Refactor container_stats() with translatable, structured Docker
errors and timeout handling instead of generic Docker exceptions.
- Translate low-level Docker errors into component-flavored errors
(NotRunningError/StatsTimeoutError/UnknownError) for Home Assistant,
Supervisor, and each plugin (Audio, Cli, CoreDNS, Multicast,
Observer) as well as apps, mirroring the existing App.stats()
pattern for consistent, informative API responses.
- Add a shared api_return_stats() helper to remove duplicated stats
handling logic across the 8 stats API endpoints.
- DockerContainerNotFoundError now inherits from DockerNotFound.
* Contain aiodocker _query access and split v1/v2 stats response model
- Contain the reach into aiodocker's protected _query internals (needed
for one-shot stats until aio-libs/aiodocker#1054 lands) to a single
DockerAPI._query_one_shot_stats() helper.
- Split the stats API response model between v1 (legacy, windowed CPU
percentage still available via opt-in one_shot query string) and v2
(always one-shot, no query string, cpu_percent dropped since it would
never carry a meaningful value). Both share the same
api_return_stats(stats, legacy=...) helper so the response model
stays in one place.
- DockerAPI.container_stats() keeps its existing one_shot keyword only;
the legacy/v1 vs v2 distinction is handled entirely in the API layer.
* Detect stopped/restarting containers in one-shot stats response
Docker returns a stub response containing only the container's id/name
(no cpu_stats/memory_stats/networks data) for a stopped or restarting
container instead of an error, even in one-shot mode - the not-running
check in daemon.ContainerStats happens before the one-shot branch, so
requesting one-shot stats doesn't change this. Detect the missing
online_cpus field (only ever present while running) and raise
DockerContainerNotRunningError instead of returning a payload that
DockerStats would otherwise turn into a successful all-zero response.
* Add tests for stats error paths introduced in one-shot/totals PR
- Cover DockerAPI.container_stats/_query_one_shot_stats edge cases: empty
one-shot response, empty windowed response, and the raw _query_one_shot_stats
helper.
- Add stats() error-translation tests (not-running, timeout, unknown) for the
audio, cli, multicast and observer plugins, none of which had any coverage.
- Add stats endpoint tests (v1 + v2) for the cli, audio, dns, multicast and
observer REST APIs, plus a v1 success-path test for the apps stats endpoint.
These close the codecov-flagged gaps for newly introduced code in this PR.
* Unify one-shot and windowed container_stats() code paths
Both modes now share the same error handling and not-running detection:
- The windowed path no longer inspects the container first to check its
running state. Docker returns the same id/name-only stub response (no
online_cpus in cpu_stats) for a stopped or restarting container in
windowed mode as it does in one-shot mode, so the existing online_cpus
presence check now covers both, closing a race window where a container
stopping between the inspect and the stats call would previously reach
the DockerStats parser with an incomplete payload.
- The windowed path builds its container handle via containers.container(),
which constructs a DockerContainer from the name with no I/O, instead of
containers.get(), which performs an unnecessary inspect call just to
fetch the same information the stats call already returns.
* Fix protected access pylint issue
Triaging SUPERVISOR-1JWK turned up a missed port conflict:
RE_PORT_CONFLICT_ERROR only matched one of the Docker daemon's
port-in-use message shapes. The two variants produced by current moby
— "Bind for <ip>:<port> failed: port is already allocated" from
portallocator and "failed to bind host port <ip>:<port>/<proto>:
address already in use" from osallocator — fell through to
DockerAPIError, got re-raised as AppUnknownError, and the watchdog
shipped them to Sentry as unknown errors.
Widen the regex to match all known shapes (including the older form
embedding the container endpoint, still observed from older daemons
and wrappers), anchored on the "failed to set up container networking"
prefix and one of the "address already in use" or "port is already
allocated" suffixes. Log the raw Docker message at debug level before
converting, so curious users can still see the exact upstream text
(host IP, container endpoint, protocol) when investigating which
process is holding the port.
The watchdog's _restart_after_problem now catches AppPortConflict
explicitly ahead of the generic AppsError handler: log a warning,
break the retry loop, do not call async_capture_exception. A port
conflict is an environment condition — another process grabbed the
port while the add-on was down — so retrying cannot make it succeed
and reporting to Sentry is noise.
With port conflicts now raised as typed APIError subclasses at the
detection site, the DockerAPIError → format_message() rewrite fallback
in api_return_error has no work left. Drop the fallback and delete
supervisor/utils/log_format.py along with its tests; the module only
ever handled port-conflict prose.
Fixes SUPERVISOR-1JWK
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The aiodocker 0.25.0 upgrade (PR #6448) changed how DockerError handles
the message parameter. The library now extracts the message string from
Docker API JSON responses before passing it to DockerError, rather than
passing the entire dict.
The port conflict detection tests were written before this change and
incorrectly passed dicts to DockerError. This caused TypeErrors when
the port conflict detection code tried to match err.message with a
regex, expecting a string but receiving a dict.
Update both test_addon_start_port_conflict_error and
test_observer_start_port_conflict to pass message strings directly,
matching the real aiodocker 0.25.0 behavior.
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
* Map port conflict on start error into a known error
* Apply suggestions from code review
* Run ruff format
---------
Co-authored-by: Stefan Agner <stefan@agner.ch>