Files
supervisor/tests/mounts
Stefan AgnerandClaude Fable 5 324d17e68b Use systemd automount units for network mounts (#7167)
* dbus: support aux units in start_transient_unit

Extend `Systemd.start_transient_unit` to accept the `aux` parameter
(`a(sa(sv))`), which has been hardcoded to `[]` since the wrapper was
introduced. Aux entries take the form `(unit_name, properties)` and
let systemd create multiple transient units atomically — most usefully
a `.mount` and its `.automount` companion in one D-Bus call.

Add the unit-property constants we'll need to drive that:

- `DBUS_ATTR_LAZY_UNMOUNT` ("LazyUnmount"): set on the `.mount` so
  systemd umounts with MNT_DETACH. Pairs with `softerr`/`soft` to make
  stop/restart sequences reliable even when the server is gone.
- `DBUS_ATTR_TIMEOUT_IDLE_USEC` ("TimeoutIdleUSec"): set on the
  `.automount` to control how long the mount stays around after the
  last access before autofs expires it.
- `DBUS_ATTR_WHERE` ("Where"): mount-point property; required on the
  `.automount` unit (and useful explicitly on `.mount` units too).

No call sites change in this commit — pure plumbing.

* mounts: pair each network .mount with a .automount companion

Network mounts now get a transient `.automount` unit created
atomically alongside their `.mount`, via the aux parameter on
`StartTransientUnit`. The systemd-managed autofs trigger handles
activation lazily: the path exists in the VFS even before the
underlying network mount runs, and the first access that crosses the
trigger fires `mount.cifs`/`mount.nfs` on demand.

Why this matters:

- PID 1's path-walks (chase, daemon-reload, generator scans) stop
  crossing dead NFS/CIFS lookups — autofs returns from kernel memory
  rather than entering the network filesystem. The class of failure
  fixed at the systemd-timeout layer in #6834 stops being reachable
  in the first place.
- Failed accesses fail fast, bounded by the `.mount`'s `TimeoutSec=`
  rather than hanging forever.
- Background reconnect (NFSv4 state-manager kthread, CIFS
  delayed_work) recovers transparently when the server returns; no
  remount needed.

The `.automount` carries `TimeoutIdleUSec=5min` so kernels can expire
the mount after inactivity and re-trigger on the next access. The
companion `.mount` carries `LazyUnmount=true` (MNT_DETACH on stop),
so umount returns immediately even when the server is unreachable;
existing fds drain in their soft/softerr timeout regime instead of
pinning the umount syscall.

`Mount.unmount()` now stops the `.automount` first so the autofs
trigger can't re-fire the underlying `.mount` during cleanup. The
automount stop is best-effort — if the unit is gone or refuses to
stop, we log and proceed to the `.mount` stop, which remains
authoritative.

Bind mounts opt out via a `creates_automount` class flag — they have
no server to wait on and the lazy semantics would only add
indirection. Network mounts (CIFS/NFS) opt in.

This is wiring only; the periodic mount reload, the reload/restart
escalation, and the inner bind-mount layer for media/share usage are
still in place. The next commit removes them now that automount
makes them unnecessary.

* mounts: drop bind layer, periodic reload, and reload/restart machinery

With the kernel's autofs trigger handling lazy activation and
transparent reconnect, almost all of the supervisor-side mount
choreography becomes unnecessary. Strip it out.

What goes away:

* The inner bind-mount layer. Media/share mounts used to live at
  `path_extern_mounts/{name}` and have a second `.mount` unit bind
  them into `path_extern_media/{name}` (resp. `share`). The
  `.automount` companion can sit directly at the container-facing
  path, and the parent-dir RSLAVE bind into add-on containers
  surfaces the autofs trigger the same way. `NetworkMount.where`
  now switches on usage:
    - MEDIA → `path_extern_media/{name}`
    - SHARE → `path_extern_share/{name}`
    - BACKUP → `path_extern_mounts/{name}` (unchanged)
* `BindMount`, `BoundMount`, `MountManager._bound_mounts`, the
  `_bind_mount`/`_bind_media`/`_bind_share` helpers, and the entire
  emergency-fallback dance (`path_emergency`). The empty-read-only
  dir trick was a workaround for the harder PID 1 wedge problem,
  which autofs solves at the root. Failed shares now surface as
  ETIMEDOUT/EHOSTDOWN on access plus a resolution issue.
* `MountManager.reload()` and the 15-minute `RUN_RELOAD_MOUNTS`
  periodic task. autofs re-activates on access; we don't need to
  poll, and not polling means we don't randomly trip stale-FH /
  softreval / dead-server edge cases on a timer.
* `Mount.reload()` and `Mount._restart()`. The reload→restart
  escalation existed to make `is_mounted` honest in the face of
  systemd's local-only state; now `is_mounted` IS honest (probe-
  based), and any recovery the kernel knows how to do happens
  inside autofs without our involvement.
* The `RELOADING` safety net introduced by #6834. That whole class
  of PID 1 wedge stops being reachable when path lookups don't
  cross dead network mounts.

What changes shape:

* `NetworkMount.is_mounted()` no longer gates on systemd's
  `ActiveState` before probing. The `.mount` unit is dormant
  whenever autofs hasn't recently triggered it, so systemd state
  is meaningless as a health signal. The probe (now via
  `statvfs("/path/.")` so the trailing dot forces `LOOKUP_DIRECTORY`
  and triggers autofs) is the source of truth, and after it returns
  we set `self._state` to ACTIVE/INACTIVE so the API reports
  reachability rather than autofs idle state.
* `MountManager.reload_mount()` becomes probe-only — no systemd
  reload/restart calls. The user-facing semantics: "tell me if this
  mount is reachable right now, and refresh the resolution issue
  accordingly." If the mount is dead, autofs will re-trigger it on
  the next consumer access; the supervisor doesn't need to force it.
* `Mount.load()` keeps `_update_state_await` for the no-job-dispatched
  case where a previous supervisor left a unit in `activating`, then
  probes once via `is_mounted` so state reflects reachability rather
  than the lazy-mount idle state.
* `BackupManager` no longer walks `bound_mounts`. Backups skip
  network-mount subdirectories during folder archive (so we don't
  recurse into the share), and unmount/re-mount them around folder
  restores (so writes target the local mount-point dir, not the
  remote share).

Behavior tradeoffs accepted:

* Dead media/share access now ~30s ETIMEDOUT instead of empty
  read-only dir. The empty-dir was a workaround for the PID 1
  wedge that autofs eliminates at the root; if specific add-ons
  rely on the old behavior we can revisit with a 2-stage automount
  whose fallback target is an emergency dir.
* Resolution-issue lag: failed mounts no longer surface within
  15 min on their own. They appear when the user probes via the
  API, when `BackupManager.reload()` walks the location, or when
  load-time activation fails.

* tests: adapt bind-layer-era tests after rebase

Main gained tests for the bind layer after this branch was written:
the #7013 rebind regression test and the #7072 bind-step rollback test
cover machinery that no longer exists, and the healthy-reload test
asserted the unconditional rebind. Drop the first two and reduce the
third to asserting that a healthy probe performs no systemd operations
at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: rearchitect automount setup from systemd/kernel review

A code-level review of systemd's automount implementation (verified
against v254.13, the version HAOS ships) and the kernel autofs/VFS
plumbing surfaced three correctness-critical flaws in the automount
design, plus several hardening gaps. See automount-rearchitecture.md
for the full analysis.

- Make the .automount the primary transient unit with the .mount as
  aux, mirroring `systemd-mount --automount=yes`. Aux units get no
  start job: with the .mount as primary the trigger was never armed
  and the design silently degraded to eager mounting.

- Set StartLimitIntervalUSec=0 on the .mount. With the default start
  rate limit (5 starts/10s, counting successful starts) a fast-failing
  mount plus any polling consumer trips the limit within seconds, and
  systemd then detaches the autofs trigger entirely
  (AUTOMOUNT_FAILURE_MOUNT_START_LIMIT_HIT) — the path silently
  becomes a plain writable local directory.

- Drop the 5-minute TimeoutIdleUSec (default 0 = never expire).
  Kernel idle expiry decides busyness via may_umount_tree(), which
  only counts the init-namespace mount instance — open files held by
  container processes are invisible, so expiry would unmount shares
  under actively writing add-ons.

- Probe with a plain statvfs: the statfs syscall walks with
  LOOKUP_AUTOMOUNT and triggers by itself; the trailing-dot trick is
  unnecessary. Classify ELOOP as a mount-propagation
  misconfiguration in the probe error handling.

- Re-arm the trigger from reload_mount(): if the .automount unit is
  failed or gone (e.g. autofs unmounted out-of-band), reset failure
  state and re-create the pair before probing.

- Tear down legacy eager-mount units during load(). The old design's
  bind unit for media/share occupies the exact unit name the network
  .mount uses now; on a warm upgrade adoption would mistake the old
  bind mount for the network mount and leak the legacy data mount.

- Honor the systemd job result when stopping the .mount during
  unmount and reset failure state on both units afterwards so dead
  transient units get garbage-collected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: adapt local data repair to the autofs design

Port of the mount-target-not-empty repair (#7089) on top of the
automount rearchitecture:

- relocate_local_data() handles a single target directory — the mount
  sits directly at its container-facing path, the bind layer and with
  it the multi-directory case are gone
- a repair_trigger() failure caused by blocking local data raises the
  mount failed issue with the move_local_data suggestion instead of
  the plain variant
- local data at load is detected via the mount unit itself; the
  bind-layer detection test is rewritten accordingly

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docker: use rslave propagation for execute_command share mount

The temporary container used for execute_command (e.g. core config
check) mounted /share without a propagation mode, unlike every other
share/media mount. With eager mounts this only meant missing mounts
made after container start; with lazy automount activation the shares
are routinely mounted after start, and accessing a not-yet-activated
automount from a private mount namespace fails with ELOOP.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: harden automount lifecycle from review findings

Address the findings of an adversarial review of the automount
rearchitecture:

- load() no longer adopts a dead automount trigger: a failed or
  stopped .automount leaves the path a plain writable directory — the
  silent local-write degradation this design must prevent. Tear down
  and re-arm instead. An adopted pair whose probe fails now raises so
  the manager surfaces the mount failed issue, same as a fresh mount.
- unmount() resolves the .mount unit only after stopping the
  .automount (the lazy detach can garbage-collect the transient unit,
  invalidating an earlier proxy), raises instead of warning when the
  automount stop errors, and checks the stop job result so a failed
  stop cannot leave an armed trigger behind a "successful" removal.
- repair_trigger() fully unmounts before re-arming: with the share
  still attached, arming fails — or worse, mount() misreads the
  mounted share's contents as blocking local data.
- Legacy unit teardown honors the stop job result for the same reason.
- reload_mount() escalates once to re-creating the unit pair when an
  established mount is unreachable: a permanently dead session (e.g.
  replaced server) keeps the path mounted, so the trigger can never
  re-fire and kernel reconnection never succeeds. Teardown is safe now
  (lazy unmount, no PID 1 path walks), unlike the removed reload →
  restart escalation of the eager design.
- Reinstate the periodic mounts task as a probe-based reconcile: re-arm
  dead triggers, refresh the reachability state reported by the API and
  used by backup locations, and sync the mount failed issue in both
  directions. No reload or restart of established mounts.
- Folder restore no longer fails after a successful restore when
  re-mounting nested mounts cannot verify an unreachable server — the
  trigger is armed, the share recovers on next access.
- Drop the now-unused Mount.update(), deduplicate unit name escaping,
  hoist the backup exclusion set out of the per-file filter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* tests: expect rslave propagation in core check container

The config check assertion missed the update for the rslave
propagation on the execute_command share mount.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* tests: cover reconcile trigger repair failure paths

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: arm automount before stopping the legacy data mount

Address review: the legacy teardown stopped the media/share bind unit
and the eager data mount in sequence, both strictly. When the bind
stop succeeded but the data mount stop failed — its unmount can time
out against an unreachable server, legacy units have no LazyUnmount —
load() raised with the container-facing path left a plain writable
directory and no trigger armed: the pollution mode this design is
meant to prevent. Reordering the stops would not help either, as
systemd stops the bind first anyway as a dependent of the data mount
(the implicit Requires= from RequiresMountsFor= on the bind's What=).

Instead, strictly stop only the path-conflicting unit (a failure
leaves the path covered, which is safe and retryable), arm the
automount right after, and stop the conflict-free legacy data mount
best-effort last — a failed stop logs a warning and leaves an orphaned
mount for the next Supervisor restart or a host reboot, with nothing
writable exposed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: discard dead session instead of re-creating units on reload

Address review: re-creating the unit pair on reload of an unreachable
established mount exposed the target as a plain writable directory
between the lazy unmount and arming the replacement — a continuously
writing add-on could block the re-arm or slip writes under the new
mount in the check-to-arm window.

Stop only the .mount unit instead, keeping the .automount armed:
systemd re-installs the autofs trigger over the path — the same
mechanism idle expiry uses, with the automount's Triggers= reference
keeping the transient .mount definition alive — so the path is never
locally writable. The re-probe then mounts fresh through the trigger,
which establishes a new session and thereby covers the permanently
dead session case (e.g. a replaced server) the escalation exists for.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: surface arming failures, keep restore teardown re-armed

Address review findings on the arming/re-arming error paths:

- mount() checks the StartTransientUnit job result: a failed start job
  means the trigger never armed and the path is a plain writable
  directory — a hard MountError before the probe, so it cannot be
  mistaken for the armed-but-unreachable MountActivationError.
- Folder restore includes the nested-mount teardown in the try block:
  an unmount failing halfway (trigger disarmed, share still attached)
  previously exited before any re-arm, leaving the path unprotected.
  The finally re-arms via repair_trigger(), which no-ops on a still
  armed trigger and handles partially torn-down pairs.
- The post-restore re-arm suppresses only MountActivationError (armed,
  recovers on next access). Any other failure — arming failed, local
  data blocking the target — left the path unprotected and surfaces.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: fold the two legacy unit teardowns into one helper

The eager-mount-era cleanup queried the .mount unit a second time
although load() had just fetched it, and the strict teardown of the
unit occupying the automount's path and the best-effort teardown of
the mount at the mounts data directory were near identical. Reuse the
unit from load() and give the shared helper a strict flag, which drops
a D-Bus round trip per mount on every load.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: re-arm the trigger when unmount cannot stop the mount

Stopping the automount detaches the whole stack at the path, so an
unmount that then fails to stop the .mount leaves a plain writable
directory behind. Arm a fresh pair before raising, best effort — the
local data repair covers what lands there if that fails too.

The automount stop itself needs no such handling: systemd's
automount_stop() enters dead synchronously and the detach it performs
(MNT_DETACH, no server contact) only logs its errors, so a failure
there means systemd could not be reached at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: raise translatable errors for mount operations

Errors from setting up, unmounting and reloading a mount reach users
through the API, so describe what failed in a translatable message and
leave the systemd job result or D-Bus error to the log.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: remove the emergency folder of the eager-mount design

The read-only fallback directory has no place in the automount design.
Remove what is left of it on existing installations, next to the legacy
addons directory cleanup, keeping it if it holds anything but the empty
mount points.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: trim comments to the behavior they describe

Drop the narration of what changed and why from comments and
docstrings, keeping the notes that record why a tempting alternative
does not work. Two comments still described the reload to restart
escalation this branch removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: correct stale claims in manager and reload test

The manager docstring credited the kernel with idle expiry, which is
deliberately disabled, and claimed no polling although the periodic
reconcile probes every 15 minutes. The reload test claimed systemd is
never contacted while the escalation stops the .mount unit; only reload
and restart of the unit are avoided.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: arm the trigger alone if re-creating the pair is rejected

Address review: the re-arm after a failed unmount submits the .mount as
an aux unit, but transient creation requires a pristine unit and the
definition is still loaded precisely when stopping it is what failed.
Fall back to creating the .automount on its own, which covers the path
and fires the surviving mount definition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mounts: never arm the trigger over local data after a failed unmount

Address review: the fallback caught the target validation errors too,
so data written while the path was uncovered would end up beneath a
fresh trigger. Reconciliation would then find a healthy mount and the
repair skips mount points, leaving the data hidden for good. Leave the
path untouched instead, so the next reconcile offers to move it away.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-31 20:30:23 +02:00
..