pmbootstrap shutdown unmounts porthole's own container binds, not just pmbootstrap's chroot mounts
scope: generic · severity: finding · confidence: proven · subsystem: build
The question — lib/porthole_image.py’s rootless assembler needs the
rootfs chroot’s /proc, /sys and /dev unmounted before mkfs.ext4 -d
recurses it (a live /proc makes mkfs.ext4 fail with “Permission denied
while opening auxv to copy” – see the assembler’s own commit). pmbootstrap shutdown is the unmount pmbootstrap itself uses between commands, verified
by hand to work rootless in this same workspace. Is it the right call to make
from _ph_assemble_image, once per image build?
The answer — no. pmbootstrap shutdown -> pmb.chroot.shutdown() ->
umount_all(...) walks /proc/self/mountinfo and unmounts every
mountpoint under pmbootstrap’s work dir (/pmb in this workspace), not just
the chroot passed to it. porthole’s sandbox bind-mounts its own paths –
/pmb/cache_git/pmaports among them – under that same /pmb, at container
creation, before pmbootstrap ever runs. umount_all cannot tell “a chroot
mount pmbootstrap made” from “a bind podman made that merely happens to sit
under the same work dir”, and unmounts both. The first build after a
shutdown call then fails at 0 seconds, naming an aport rather than a mount:
by the time pmbootstrap looks for pmaports/device/.../APKBUILD the
directory is an empty mountpoint again, having been unmounted out from under
the running container. Recovery is a container restart (sandbox down +
sandbox up), not anything pmbootstrap offers – there is no pmbootstrap
command that re-establishes a bind mount porthole set up, because
pmbootstrap does not know it exists.
What this rules out — pmbootstrap shutdown (or, by the same logic, any
pmb.chroot.shutdown/zap-family call with only_build_related unset or
scoped wider than one chroot) as a way to clean up before a filesystem
build in this workspace. It is not merely the /dev rbind hazard
what-a-rootless-workspace-cannot-do #5 already names – that finding is
about individual umount calls failing with “not mounted” inside the
recursive /dev bind, which is a correctness annoyance strict-mode zap dies
on. This is a blast-radius hazard: even a shutdown that itself reports
exit 0 and looks completely clean can still have unmounted infrastructure
that has nothing to do with the chroot it was meant to clean, because
umount_all scopes by “under the work dir”, and porthole’s own binds live
under the work dir too. The fix is to scope the unmount to the chroot’s own
path prefix (<chroot>/..., deepest mounts first, each umount best-effort
since the same /dev rbind hazard from #5 still applies at that finer
grain) and never call pmbootstrap shutdown, zap, or anything built on
umount_all against the whole work dir from inside a single build.
How it was established — measured directly on hardware during Task 2.4’s
rootless-assembler hardware gate: pmbootstrap shutdown run by hand, verified
clean (27 mounts, all gone), ruled safe on that basis; the very next build
failed naming an aport; /pmb/cache_git/pmaports found empty inside the
container; a sandbox down/sandbox up cycle was needed to recover, at which
point the same build succeeded again up to the point this finding’s fix
addresses. A second, separate measurement (mount count under the chroot
prefix vs. total mounts under /pmb, on a freshly recreated container)
confirmed the chroot-prefix scope is both sufficient (12 relevant mounts,
more mid-install) and safe (cache_git/pmaports is provably outside that
prefix). Overturned by: porthole’s sandbox setup moving its own binds outside
/pmb entirely, or a pmbootstrap release that scopes umount_all to the
chroot path itself rather than the whole work dir.
