aboutsummaryrefslogtreecommitdiff
path: root/docs
diff options
context:
space:
mode:
authorCraig Jennings <c@cjennings.net>2026-10-05 23:11:01 -0500
committerCraig Jennings <c@cjennings.net>2026-10-05 23:11:01 -0500
commit32008ac10409e1a6f8ddb8820c6e8ac30e575db4 (patch)
tree6f163aa8c1f3b25b7cfd0092461af45a0b7c6dc0 /docs
parentec5a021194c4ab3b9d9fd2b260b183ad0bdf347f (diff)
downloadarchsetup-32008ac10409e1a6f8ddb8820c6e8ac30e575db4.tar.gz
archsetup-32008ac10409e1a6f8ddb8820c6e8ac30e575db4.zip
docs(spec): review the guarded-upgrade spec and decompose its build
I ran three review rounds against the live code in archsetup and the dotfiles. They found 129 issues, and the blocking count fell from 13 to 4 to 0. Each one is dispositioned in the spec. The exact contract now lives only in Implementation phases, because the copies in Design, Decisions and Testing had drifted. The spec moves from DRAFT through READY to DOING, with nine non-blocking residuals left for the first build task. The parent task now carries 22 build tasks in commit-group order, ending with the IMPLEMENTED flip. The manual tests are under Manual testing and validation, and a stale-kernel age metric is filed as [#D].
Diffstat (limited to 'docs')
-rw-r--r--docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org5593
1 files changed, 5506 insertions, 87 deletions
diff --git a/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org b/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org
index d9ec8d4..e19ea6c 100644
--- a/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org
+++ b/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org
@@ -4,17 +4,25 @@
#+TODO: TODO | DONE
#+TODO: DRAFT READY DOING | IMPLEMENTED SUPERSEDED CANCELLED
-* DRAFT Guarded-upgrade completion
+* DOING Guarded-upgrade completion
:PROPERTIES:
:ID: 81cdfd72-db96-43d3-aa03-779878c99f3e
:END:
+- [2026-10-05 Mon @ 13:00 -0500] READY → DOING. Decomposed into build tasks under the archsetup todo.org parent "Topgrade guarded-upgrade build" (SPEC_ID 81cdfd72-db96-43d3-aa03-779878c99f3e), with the manual tests filed under Manual testing and validation.
+- [2026-10-05 Mon @ 12:10 -0500] DRAFT → READY. Three review rounds incorporated (48, 59 and 22 findings; blocking went 13, 4, 0), all 129 dispositioned. 9 non-blocking residuals from the last repair pass stay open under Review findings and are carried into the build tasks.
+- [2026-10-05 Mon @ 11:10 -0500] spec-review round 3 on the consolidated spec: no blocking findings. 22 non-blocking findings recorded and being resolved before READY.
+- [2026-10-05 Mon @ 10:40 -0500] round-2 review incorporated: all 59 findings dispositioned, the exact contract consolidated into Implementation phases, and machinery without a caller cut.
+- [2026-10-05 Mon @ 09:45 -0500] spec-review round 2 on the revised spec: Not ready. 59 new findings recorded, 4 of them blocking.
+- [2026-10-05 Mon @ 09:30 -0500] reworded the bodies of Decisions 2 to 7 to match the rewritten design; I reopened none of them, and Decision 1 is unchanged. Decision 2: the TTY fallback holds the kernel and DKMS sets. Decision 3: the TOML is a divergent list, rewritten to equal the hook and pinned by an archsetup test, and the rollout is three code commits. Decision 4: it names the record, which is always written (except on the EUID 0 and lock-held refusals), states the stamp predicate per form, gives each test suite a copy of one fixture, and its example is now 6 deferred. Decision 5: the DKMS set joins the hold, with =zfs-utils= joining through the closure; the gate covers every kernel whose modules the run changed, plus open failures; ratio has three kernels; the DKMS hook timing is corrected. Decision 6: the boot form is =upgrade-guarded --apply-armed=, which installs the flag's list from the cache that arming pre-downloaded, with no refresh and no network, and the armed everyday run's disarm cases are split by cause. Decision 7: the flag is root-owned, carries the armed =<name>=<version>= list, and =ExecStartPre= moves it off its path at the start of the attempt.
+- [2026-10-05 Mon @ 06:05 -0500] spec-review: Not ready. 48 findings recorded under Review findings, 13 of them blocking; status stays DRAFT until they are dispositioned.
+- [2026-10-05 Mon @ 05:14 -0500] named the split script =upgrade-guarded= and added =containers= to its topgrade =--disable= list (the containers step fails on locally built images, hit again on velox 2026-10-04).
- [2026-08-25 Tue @ 18:45 -0600] decisions closed 7/7. The kernel decision reversed on the velox DKMS failure chain: held on every everyday run, landed only in the dedicated session behind a DKMS/initramfs/snapshot gate.
- [2026-08-25 Tue @ 18:30 -0600] redirected: the everyday path is a live split upgrade (apply everything the guard would not block, defer the rest); the boot-time oneshot becomes the completion step for the deferred set. Decided while running exactly that by hand on ratio.
- [2026-08-25 Tue @ 06:39:42 -0600] drafted. Grounded in a live read of the maint engine, the pacman hooks, and the boot path on velox, not memory. The topgrade-freshness diagnosis that motivates it is in this session's log.
* Metadata
-| Status | draft |
+| Status | doing |
|----------+-------------------------------------------------------------|
| Owner | Craig Jennings |
|----------+-------------------------------------------------------------|
@@ -24,78 +32,127 @@
* Summary
-The waybar maintenance module shows topgrade freshness as permanently stale. The cause is a real one: on a machine running Hyprland, a full =topgrade= almost never exits 0, because its system step upgrades GPU/compositor libraries that the =hypr-live-update-guard= pacman hook correctly refuses to swap under a live session. The freshness stamp is gated on topgrade's exit code, so a correct, protective refusal reads as "you never run updates." This spec designs a safe path to actually complete a guarded upgrade, and makes that completion record the freshness stamp, so the metric tracks the true state of the system.
+The waybar maintenance module shows topgrade freshness as permanently stale. The cause is a real one: on a machine running Hyprland, a full =topgrade= almost never exits 0, because its system step upgrades GPU/compositor libraries that the =hypr-live-update-guard= pacman hook correctly refuses to swap under a live session. The freshness stamp is gated on topgrade's exit code, so a correct, protective refusal reads as "you never run updates."
+
+This spec designs a safe path to actually complete a guarded upgrade, and makes that completion record the freshness stamp, so the metric tracks the true state of the system. The everyday path is a live split run (=upgrade-guarded=) that applies everything outside its held set and records what it deferred: the GPU/compositor set while Hyprland runs and the kernel and DKMS sets always, together with any pending package that cannot move without them. The deferred set lands in a dedicated =upgrade-guarded --complete= session. The kernel and DKMS sets go in behind the =kernel-modules-check= gate, and the GPU/compositor set installs directly when Hyprland is stopped, or, when Hyprland is live, is pre-downloaded and finished by a boot-time oneshot from the package cache.
+
+Implementation phases is the one exact statement of behavior. Every other section explains the shape or the reasons and cites the phase that holds the detail.
* Problem / Context
-The metric reads one cache key, =topgrade_run= (=~/.local/state/maint/topgrade_run.json=). Absent, the probe (=maint/src/maint/probes/updates.py:114=) returns WARN, "no topgrade run recorded". Two writers stamp it: the =topgrade= PATH wrapper (=~/.dotfiles/hyprland/.local/bin/topgrade=) on =rc -eq 0=, and the panel's TOPGRADE lever (=doctor.py=), which returns before the stamp on any non-zero exit. The read path is sound (a sandboxed =maint stamp topgrade= writes the file and =maint status= then reads freshness 0); the file is simply never written.
+The waybar maintenance module's topgrade freshness comes from one cache key, =topgrade_run= (=~/.local/state/maint/topgrade_run.json=), read by =topgrade_freshness= in dotfiles =maint/src/maint/probes/updates.py:114=. It grades the stamp's age in days against =topgrade_warn_days = 14=. With no stamp at all it returns WARN, "no topgrade run recorded", a branch that only fires on a machine that has never stamped. Two writers stamp it today: the =topgrade= PATH wrapper (=~/.dotfiles/hyprland/.local/bin/topgrade=) when the real binary exits 0, and the panel's TOPGRADE lever (=doctor.py=), which returns before its stamp on any non-zero exit.
+
+The read path is sound: a sandboxed =maint stamp topgrade= writes the file, and =maint status= then reads freshness 0. The fault is that the file is almost never refreshed. On ratio it was last written 2026-07-08, the day the wrapper landed, so on 2026-10-05 the probe graded an age of just under 90 days and warned.
-It is never written because topgrade rarely exits 0 on this machine, and the reason is specific rather than flaky. =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook= is a =PreTransaction=/=AbortOnFail= hook that, when Hyprland is running and an upgrade changes the on-disk version of a GPU/compositor library, prints a BLOCKED banner and exits 1 — aborting the whole transaction before any file is swapped. Its trigger set is =mesa=, =mesa-*=, =wayland=, =libdrm=, =libglvnd=, =hyprland=, =aquamarine=, =hyprutils=, =hyprgraphics=, =vulkan-radeon=, =vulkan-intel=, =vulkan-mesa-layers=, =nvidia-utils=, =lib32-nvidia-utils=, =xorg-xwayland=. The guard exists for a proven failure: replacing those libraries under a live compositor makes the next GPU call hit a now-deleted mapping and SIGABRT, taking every Wayland client down (hit on ratio 2026-06-07).
+Topgrade rarely exits 0 on a machine running Hyprland, and the reason is specific rather than flaky. The =hypr-live-update-guard= pacman hook is a =PreTransaction=/=AbortOnFail= hook: when Hyprland is running and an upgrade changes the on-disk version of a GPU/compositor library, it prints a BLOCKED banner and exits 1, aborting the whole transaction before any file is swapped. Its 15 =Target= patterns (mesa and its subpackages, wayland, libdrm, libglvnd, hyprland and some of its core libraries, the Vulkan and NVIDIA userspace packages, xorg-xwayland) are the list =upgrade-guarded= later reads as its GPU patterns. The guard exists for a proven failure: replacing those libraries under a live compositor makes the next GPU call hit a now-deleted mapping and SIGABRT, taking every Wayland client down (hit on ratio 2026-06-07).
-So when any of those libraries has an update pending — a frequent event — topgrade's =system= step (it runs =yay=) aborts non-zero, topgrade returns non-zero, and neither writer stamps. The observed case: on 2026-08-24 topgrade ran at 17:49, hit the guard on =mesa= (26.1.7 → 26.2.1), and failed; the upgrade was then finished by hand with the guard's sentinel override, entirely outside the wrapper, so nothing stamped. The metric has read stale ever since.
+The installer writes the hook as =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook=, but both daily drivers still carry it under the legacy unprefixed name. Phase 4's rollout moves each machine to the =10-= name before its levers route through the script.
-Two framings of the fix are in tension, and choosing between them is the spec's central decision. Either the metric means "how recently did you run the sweep" (recency), so the stamp should decouple from topgrade's exit; or it means "is the system up to date" (state), so staying stale while a guarded upgrade is deferred is *correct* and the only real defect is that safely completing that upgrade doesn't stamp. This spec takes the state framing (see Decisions).
+So when any of those libraries has an update pending, which is often, topgrade's =system= step (it runs =yay=) aborts, topgrade exits non-zero, and neither writer stamps. The observed case: on 2026-08-24 topgrade ran at 17:49, hit the guard on =mesa= (26.1.7 → 26.2.1) and failed. The upgrade was then finished by hand with the guard's sentinel override, outside the wrapper, so nothing stamped. The metric had already gone stale in late July and stayed stale.
+
+The fix has two framings in tension. If the metric means "how recently did you run the sweep" (recency), the stamp should decouple from topgrade's exit. If it means "is the system up to date" (state), staying stale while an upgrade is deferred is correct, and the only real defect is that safely completing the deferred upgrade doesn't stamp. I take the state framing (Decision 1, "Metric means state, not recency"), so the rest of the spec builds a safe completion path and makes that completion the stamp.
* Goals and Non-Goals
** Goals
-- A safe, low-friction way to apply a guarded (GPU/compositor-library) upgrade, with Hyprland not live at swap time.
-- That completion records the =topgrade_run= freshness stamp, so the metric clears when the system is genuinely current.
+- An everyday UPDATE or TOPGRADE that completes live. It applies everything outside its held set and defers the GPU/compositor set (while Hyprland runs), the kernel and DKMS sets (always), and any pending package that can't move without them. A deferral alone never fails the run, and the deferred set stays visible in maint until it lands.
+- A safe, low-friction way to land the deferred set: a dedicated =upgrade-guarded --complete= session that lands the kernel and DKMS sets behind the =kernel-modules-check= gate, offers no reboot while that gate is open, and applies the GPU/compositor set with Hyprland not live at swap time (at the next boot, or directly from a console with Hyprland stopped).
+- =upgrade-guarded= writes the =topgrade_run= freshness stamp only when a run leaves nothing deferred and no gate open, so the metric clears exactly when the system is current and stays honestly stale while a protective deferral is outstanding. The per-mode predicate is in Phase 1.
- A boot-time upgrade path that can never lock the machine out of its session, however it fails.
-- The installer owns the durable pieces so a rebuilt machine has them without hand-setup.
+- The installer owns the durable pieces, so a rebuilt machine has them without hand-setup.
** Non-Goals
-- Weakening or bypassing the =hypr-live-update-guard= hook. It stays exactly as strict; this builds *around* it, not through it.
-- Making the full topgrade ecosystem sweep (git repos, vim, npm, ...) run at boot. Those never need a stopped compositor and are out of the boot path.
-- Changing how the kernel hazard is *guarded*. The hook stays silent on kernels; the split script holds them back on a live run as a second, separately-reasoned list (see Design), which is a deferral policy rather than a guard.
-- A general offline-update system for all of pacman. Scope is the guarded-library case.
+- Weakening or bypassing the =hypr-live-update-guard= hook. I keep it exactly as strict and build around it, not through it.
+- Running the topgrade ecosystem sweep (vim, npm, ...) at boot. None of it needs a stopped compositor, so it stays on the live path.
+- Changing how the kernel hazard is guarded. The hook stays silent on kernels. Holding the kernel and DKMS sets on every everyday run is a deferral policy inside =upgrade-guarded=, a separate list with its own reasons (Design), not a guard. An AUR upgrade that needs a newer kernel-set or DKMS-set package is the one ungated path; the everyday run detects it afterwards and opens the gate entry (Risks).
+- A general offline-update system for pacman. The boot form installs only the armed list (the GPU/compositor packages pre-downloaded at arming) from the package cache. Everything else lands live.
** Scope tiers
-- v1: the split-upgrade script (live: apply the non-blocked remainder, defer the rest, run the ecosystem sweep with the system step off, report the deferred set); maint's UPDATE/TOPGRADE levers route through it; an "apply on reboot" affordance that installs the held kernel live and arms the boot-time oneshot for the GPU/compositor set.
-- Out of scope: full-sweep-at-boot; touching the guard's policy.
-- vNext: none open — the kernel deferral that was vNext is now part of v1's held set.
+- v1: everything in Implementation phases. That is the two scripts and their installer step (Phase 1), maint's repointed levers with the deferred row and its APPLY action (Phase 2), the =archsetup-boot-upgrade.service= oneshot (Phase 3), and the docs and rollout on both daily drivers (Phase 4).
+- Out of scope: everything under Non-Goals.
+- vNext: a stale-kernel age in maint (how long the kernel set has been held, graded so an overdue dedicated session shows up), filed in =todo.org= as a [#D] task when the phases are decomposed.
* Design
-The shape follows one principle: the only part of topgrade that needs a stopped compositor is its =system= step when a guarded library is pending. Everything else runs fine live and rarely fails. So the safe path is small and targeted — apply the guarded system upgrade with Hyprland down, once, and stamp it — while the ordinary full sweep stays a normal live =topgrade= run.
+The shape follows one principle: the only packages that need a stopped compositor are the GPU/compositor libraries the guard protects, and the only ones that need a moment I chose are the kernel and its DKMS modules. Everything else runs fine live and rarely fails. So the safe path is small and targeted: land the kernel and DKMS sets behind a gate in a session I start on purpose, apply the GPU/compositor set once with Hyprland down, and stamp only when nothing is left deferred. Everything else, the ecosystem sweep included, runs live through the split script, which leaves those sets behind.
+
+Three pieces, but the everyday gesture is not a reboot. The split script's everyday run is a normal live update that leaves the dangerous few behind. Its dedicated =--complete= session lands them on a day I choose. A boot-time oneshot runs its =--apply-armed= form to finish the GPU/compositor part before anything maps those libraries. This section gives the shape and the reasons. Implementation phases holds the exact behavior, each piece below names the phase subsection that governs it, and where the two differ, the phase bullet is normative.
+
+*The sets.* Every set the script works with is derived from the system at run time, never hardcoded. The pending set is what pacman reports as upgradable, minus IgnorePkg rows. The GPU patterns are the =Target= lines of =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook=, taken verbatim with their globs (15 today), so the guard and the script share one source of truth. The script reads only the =10-= name, because the legacy unprefixed name sorts after =60-mkinitcpio-remove=, so a blocked bare =pacman -Syu= could delete the initramfs before the guard aborts. A missing hook, or one with no Target lines, makes every mode but the boot form refuse before any transaction rather than run with an empty list, which is deliberately unlike maint's =guard.trips=.
+
+The kernel set is the packages that own an installed kernel image, plus each kernel's headers where installed; firmware and API-header packages are never members. It is held whole, so a kernel and its headers move together, and only pending members ever move, so a foreign kernel such as ratio's =linux-lts-strix= (a local build with no headers and no zfs module, and the GRUB default) is never touched. The DKMS set is the packages that ship a =dkms.conf=, today =zfs-dkms=. Its exact pin, =zfs-utils=, isn't listed by name: the dependency closure adds it exactly when an upstream bump breaks the pin, and never for a pkgrel-only bump. I hold the kernel and DKMS sets on every everyday run, on both machines, because a failed DKMS rebuild on velox's ZFS root leaves the machine unbootable (see the kernel decision).
+
+Live means Hyprland is running, by the guard's own =pgrep -x Hyprland= test, so a Ctrl+Alt+F2 console beside a live session counts as live. Every held entry is kernel-kind (the kernel and DKMS sets and whatever the closure ties to them) or GPU-kind (the pending GPU-pattern matches and whatever the closure ties to them). The everyday run always holds the kernel side, and holds the GPU side only when live. The deferred set is the held entries still pending when a run ends. It is exactly the record's package list, and the record, the panel row and the armed list each cover the whole closure, not only the pattern matches. Exact behavior: Phase 1, Sets.
+
+*The split script.* A pacman =PreTransaction= hook can only abort or allow the transaction it is handed; it can't drop targets from it. So "upgrade everything except the held sets" can't live in the hook. It lives one layer up, in =upgrade-guarded= (archsetup =scripts/upgrade-guarded=), which the installer step that installs =hypr-live-update-guard= puts in =/usr/local/bin= along with its gate, =kernel-modules-check=. The script runs as the invoking user and refuses root, because the record and the stamp live in that user's =$HOME= and yay refuses root. Every root step the script itself runs goes through =sudo -n=, so a machine without its NOPASSWD rule fails the step with a reason rather than wedge on a password prompt the panel's runner can't answer. yay and the sweep escalate through their own sudo, but only after the refresh has succeeded under =sudo -n=. Each invocation runs exactly one mode: the everyday run, =--complete=, =--apply-armed= (the boot unit's alone) or =--dry-run=. The only option is =--no-topgrade=, and only in the everyday run.
+
+=--dry-run= is the preview. It prints what a run would hold, defer and execute, writes nothing, uses no sudo and takes no lock, and it resolves against a private sync db of its own, so it never touches the system db or the shared one another =checkupdates= caller may be using. Every other mode takes an exclusive lock beside the record, so a second run refuses instead of interleaving, and refuses while pacman's =db.lck= exists. The script never deletes =db.lck=: a stale lock can belong to a pacman that is still running, and only a person can tell. Exact behavior: Phase 1, Identity and privilege; Phase 1, Modes and options; Phase 1, Preconditions.
+
+*The everyday run.* Both panel levers call it. It captures the unread Arch news first, because yay judges news against installed build dates, and after the upgrade the news would read as old. It then clears informant's news hook, which would otherwise abort a =--noconfirm= transaction. The clear needs root, =--all= to run without prompting, and a timeout, because informant's feed fetch has none of its own. A failed clear is tolerated and leaves informant's hook live for that run, a residual I accept. Then come one refresh, the dependency closure over the held sets, and one =pacman -Su= against that same db with an =--ignore= per held entry and no second =-y=. yay upgrades the AUR packages next, and the ecosystem sweep follows unless =--no-topgrade= is given. The sweep calls =/usr/bin/topgrade= by absolute path, so the stamping =~/.local/bin/topgrade= PATH wrapper is never in the chain. It disables topgrade's =system= step (the script has already done that work), =git_repos= and =containers= (whose step fails on locally built images), each as its own argv element, because topgrade 17.12.2's parser rejects the comma-joined list, and =--no-ask-retry= keeps a failed step from waiting on stdin, since the panel's runner has no tty. A deferral alone never makes the exit non-zero; a failed step or a refusal does. Exact behavior: Phase 1, Everyday sequence.
+
+The guard hook stays installed, unchanged, as the backstop for a bare =pacman -Syu= typed at a shell. On the driven path it can still fire inside yay's own pacman call; Phase 1, Everyday sequence, says when, and Risks names the kernel-side case that yay can install ungated. Proof of concept: the guard-set half of this run, done by hand on ratio on 2026-08-25 while Hyprland was live, before the kernel and DKMS hold and the containers disable were added, resolved 724 of 730 pending packages (Emacs 31.1 among them) with the six guard hits deferred, after one unrelated fix, an orphaned =qemu-block-gluster= that had been dropped from the repo. That run upgraded linux, linux-lts, their headers, zfs-dkms and zfs-utils live, so under v1 the same day would have deferred at least twelve, and the orphan would have made the closure refuse with its name before any transaction.
+
+*The dependency closure.* =--ignore= alone isn't enough. pacman still enforces declared dependencies, so a pending package that needs a held package's new version, or whose upgrade breaks a versioned dependency of a held package, makes the transaction unresolvable, and under =--noconfirm= one unresolvable dependent fails the whole transaction, because pacman's skip prompt defaults to no. So before any transaction the script asks pacman what else has to wait, through a read-only print-mode probe against the run's one refreshed db, and holds those packages too until the probe resolves. When the answer isn't something it can safely hold around, such as an upgrade that would break an installed package that isn't pending (an orphan, foreign or AUR package, like =qemu-block-gluster= above), it refuses before any transaction and names what it found, rather than hold a package and its dependents indefinitely without saying why. A refusal never removes or forces anything. Exact behavior: Phase 1, Dependency closure.
+
+*The record and the stamp.* A deferred set nobody surfaces is a set that silently never lands, so the script records it durably in =${MAINT_STATE_DIR:-$HOME/.local/state/maint}/upgrade_deferred.json= (maint cache key =upgrade_deferred=), in maint's cache envelope, written atomically. =upgrade-guarded= is its only writer. maint's =upgrade_deferred= probe, =topgrade_freshness= and the script's own next run read it. Every mode but =--dry-run= writes it once as it finishes, refusals and failures included, so the panel always reflects the last run. =--complete= and =--apply-armed= also write it before their risky transactions, and the interrupt trap writes it too, so a kill or a power cut still leaves an accurate record with the remedy. A usage error, a root invocation and a held lock never touch it. The gate entry and the held snapshot carry forward through every mode, and only =--complete= changes them, except that an everyday run whose AUR step moved a kernel-side package opens the entry, holds its snapshot and writes the record at once (Phase 1, Invariants). "The gate is open" means the record holds a gate entry. Exact behavior: Phase 1, The record; Phase 1, Finish.
+
+=topgrade_run= keeps meaning "the system is current", so on the driven path =upgrade-guarded= is its only writer, and it stamps only when a run ends with nothing deferred and no gate open. The everyday run also needs the refresh, the upgrade, yay and the sweep all to have succeeded, so UPDATE, which skips the sweep, never stamps. =--complete= stamps only when it had something to complete, and never when it arms the boot unit, because the GPU set is still pending then. The stamp goes through maint by absolute path, because a system unit's PATH lacks =~/.local/bin=, and a missing maint or a failed stamp fails the run and is recorded, never silenced. doctor.py's lever stamp is deleted. The PATH wrapper keeps stamping a bare-shell =topgrade=, and the script never goes through it. Exact behavior: Phase 1, Stamp predicate.
+
+*For the user.* UPDATE runs =upgrade-guarded --no-topgrade= and TOPGRADE runs =upgrade-guarded=, and the =sysupgrade= alias runs the script wherever it is installed. On the common day both levers succeed. The panel's deferred row counts what the script held back, QUEUE tags the held names, and the strip's pending cell adds the held count. A deferral alone doesn't raise the row's severity or colour the waybar glyph, but freshness keeps aging until the deferred set lands, and after a TOPGRADE that left anything deferred the wall note says so and points at APPLY. A failed step, a refusal or an interruption shows on the row at WARN with its reason, and so do fixable advisories on deferred packages, which count against the row rather than against UPDATE. Arch news the run captured shows in the row's evidence until the row's DISMISS key clears it. While Hyprland isn't running (a TTY with no compositor, not a Ctrl+Alt+F2 console beside a live session), the everyday run holds only the kernel side, and the GPU/compositor set lands in the same run.
+
+Because the levers no longer arm a live swap, the guard's press-again refusal and the =--force= override go away. A tripped guard read only annotates the arm line with what the run will at least defer. The panel reads everything here from the record and the flag on its next probe, never from an exit code. Exact behavior: Phase 2 (Levers, Guard UX, The deferred row, DISMISS, Freshness and REBOOT, and CVE, QUEUE and the strip).
+
+*The dedicated session.* Landing the deferred set is =upgrade-guarded --complete=, a session I start on purpose in a terminal: the one APPLY opens, detached so closing the panel can't kill or orphan the kernel transaction, or any terminal or TTY. APPLY is offered on the deferred row while anything is deferred or the gate is open, on the freshness card while anything is deferred, and through REVIEW & FIX, and every route opens that same detached terminal. The session's first stage lands the kernel and DKMS sets live, with the desktop up for diagnosing, along with every pending package outside the GPU closure. Old modules are removed PreTransaction and rebuilt PostTransaction, so a failed DKMS build aborts nothing; the gate after stage 1 is what catches it. If the GPU closure would reach a kernel-side package, the session refuses before any transaction, because landing it would put a kernel behind the boot form or ahead of the gate.
+
+Two things happen before stage 1 touches anything, and on a ZFS root a third, the snapshot hold described below. Right after its refresh, the session removes any arm flag, and only its own GPU step can re-arm. Then, when a kernel-side package is pending or a gate is already open, it records the gate entry: which kernels are unverified, by pkgbase, and when the entry opened. A kill, a crash or a power cut mid-transaction therefore leaves the record naming the unverified kernels. From that moment until a gate passes, the deferred row is CRIT with REBOOT hidden, no arm flag exists, and the session never offers a reboot. The entry closes only when a gate passes, or when no kernel it names is still installed.
+
+The entry names every kernel whose modules the session could change: each kernel whose package or headers is pending, every kernel with a built DKMS module when a DKMS package is pending, and any kernel still unverified from an earlier session. Because it is keyed by pkgbase and resolved to module trees only after stage 1, a kernel version the session removed can never reach the gate, while the new version, and any sibling kernel whose DKMS module was rebuilt, is checked. The gate runs once, whatever stage 1's exit, against the time the entry opened. The exception is a session that opened the entry itself and whose stage 1 changed no kernel-side version, as when stage 1 fails on a download, signature or conflict error before committing anything. =/boot= is then intact and its initramfs predates the entry, so the time checks would hold the row at CRIT on a machine that boots fine. The gate runs its structural form instead, where a missing kernel image or initramfs, an unreadable image or a missing =zfs.ko= still fails it.
+
+A failed gate stops the session with the failing item named and exit 4. The machine keeps running the old kernel, the new one is installed but not booted, and the desktop is there for the fix; a later =--complete= re-runs the gate and closes the entry when it passes. Only a stage 1 that succeeded, with no gate left open, goes on to the GPU step. With Hyprland not running, the session applies the GPU/compositor set directly. With Hyprland live and the boot unit enabled, it pre-downloads that set while the network is up and arms the flag with its exact versions. With the unit not enabled, it says so and leaves the set for a console run with Hyprland stopped. It offers a reboot only from a tty, after a clean finish that armed or passed a time-checked gate. With nothing pending and no gate open, it reports that there is nothing to complete. Exact behavior: Phase 1, =--complete= sequence; Phase 1, Invariants.
+
+On a ZFS root, the fallback for a kernel that won't boot is the pre-pacman snapshot that ZFSBootMenu can boot. So on a ZFS root a session that opens the gate entry snapshots the root and places a ZFS user hold on that snapshot before stage 1. When no valid hold exists, a time-checked gate falls back to the oldest pre-pacman snapshot taken no earlier than a minute before the entry opened. The session records the held name. =zfs-pre-snapshot= prunes only unheld snapshots, so its routine prune can't destroy the one snapshot holding the old kernel before the new kernel has proven itself. The hold moves only when a later session, or an everyday run whose AUR step opens the entry, holds a newer snapshot, and a failed hold is reported without failing the run. The recorded name is the snapshot recovery boots and rollback releases. Exact behavior: Phase 1, The snapshot hold; Phase 4, Recovery.
-Three pieces, at two altitudes — but the everyday gesture is not a reboot. It is a normal live update that simply leaves the dangerous few behind.
+*The gate.* =kernel-modules-check= (archsetup =scripts/kernel-modules-check=) answers one question: would each named kernel boot with its modules? It takes the entry's pkgbases, writes nothing, makes no pacman call, and only =--complete= calls it (a by-hand run, as in Phase 4, only diagnoses). For each kernel it requires the module trees and the =/boot= image to exist, every DKMS module to be installed, and an initramfs newer than the image. On a ZFS root that initramfs must carry =zfs.ko=, read with =sudo -n lsinitcpio= because the images are root-only. The time-checked form also requires the initramfs to postdate the entry's opening and, on a ZFS root, a pre-pacman snapshot of the root dataset taken no more than a minute before it, the window =zfs-pre-snapshot='s 60 s skip leaves. The initramfs is never compared with the module tree's =vmlinuz=, whose mtime is its build date. Kernels outside the entry are never checked, so ratio's =linux-lts-strix=, where zfs is only 'added', never fails it. There is no standalone mode; the supported re-check is =upgrade-guarded --complete=. Exact behavior: Phase 1, The gate.
-*The split script.* A pacman =PreTransaction= hook can only abort or allow the transaction it is handed; it cannot drop targets from it. So "upgrade everything except the guarded set" cannot live in the hook — it lives one layer up, in a script the panel calls. On a live run the script: refreshes the sync db and reads the pending set (=checkupdates=); computes the *blocked set* = the guard's own trigger list (read from the installed hook's =Target= lines, so there is one source of truth, and version-aware the way the guard is — a same-version reinstall is not a swap) plus the *kernel set* (every installed kernel with its =-headers=, always as a set; held on every everyday run because a failed DKMS rebuild on velox's ZFS root leaves the machine unbootable — see the kernel decision); clears the news hook (=informant read=) where installed; runs =pacman -Syu --noconfirm --ignore=<blocked set>=; runs the AUR-only remainder (=yay -Sua --noconfirm=, AUR packages pinning a guarded version hold themselves back); then runs =topgrade --disable system,git_repos -y= so the other ecosystems still get their sweep and topgrade can actually exit 0. It writes the deferred set to a state file the panel reads, and exits 0 when the live part succeeded, whatever was deferred. The guard hook stays installed as the backstop for a bare =pacman -Syu= typed at a shell; on the driven path it never fires. Proof of concept: this exact sequence, run by hand on ratio on 2026-08-25 while Hyprland was live, resolved 724 of 730 pending packages (Emacs 31.1 among them) with the six guard hits deferred — after one unrelated fix, an orphaned =qemu-block-gluster= that had been dropped from the repo.
+*The arm flag.* The guard's =/run= sentinel is tmpfs and can't carry an intent across a reboot, so the arm flag lives at =/var/lib/archsetup/apply-upgrade-on-boot=, root-owned, and carries the exact =<name>=<version>= list to install. Arming first checks that the list resolves without pulling in a kernel-side package, so no kernel can ride in ungated as a dependency, then pre-downloads it with =pacman -Sw= while the network is up. A download-only run returns before PreTransaction hooks, so neither the guard nor informant's hook fires under a live Hyprland. Only =upgrade-guarded= writes the flag, atomically through =sudo -n=, and the boot unit consumes it. =--complete= clears it after its refresh and re-arms only from its GPU step. While armed, every everyday run whose refresh succeeded either re-arms on the current GPU-kind entries, so the armed versions follow the db that run synced, or disarms when there is nothing valid to arm, including on a machine whose boot unit isn't enabled, so no flag is left that nothing will consume. There is no =--disarm=, and the panel reads only the flag's presence and entries. Exact behavior: Phase 1, The arm flag; Phase 1, Everyday sequence.
-*For the user.* UPDATE and TOPGRADE on the panel run the split script; they succeed, and the panel shows "N deferred" when the script held anything back. Landing the deferred set is a dedicated session, chosen on purpose, run in the foreground from the panel's action or =guarded-upgrade --complete= in a terminal: first the kernel set, live, with the desktop still up; then the gate — every DKMS module built for the new kernel, a fresh initramfs, and on a ZFS root a pre-pacman snapshot to fall back on. If the gate fails the script stops there, names what failed, and does not reboot; the machine keeps running on the old kernel and the desktop is available for the fix. If it passes, the script arms a persistent flag for the GPU/compositor set and offers to reboot (or, from a TTY with no compositor, applies that set directly). On the next boot, before the autologin shell starts Hyprland, the deferred guarded upgrade runs in the console — the guard passes freely because nothing is live — the stamp is written, the flag is cleared, and boot continues into the session. No second reboot: the libraries are already current before anything maps them. If anything goes wrong, the machine still boots into Hyprland and the panel still shows the pending upgrade, so you are never worse off than before arming.
+*The boot unit.* =archsetup-boot-upgrade.service= is a system oneshot that runs only when the flag exists, so an unarmed boot skips it at no cost. It is ordered before =getty@tty1.service=, so it finishes before the tty1 login and so before =~/.profile.d/99-hyprland-autostart.sh= can start Hyprland. That ordering is mandatory: a parallel run would let Hyprland start mid-swap and reintroduce the exact crash the guard prevents. It is wanted by =multi-user.target= but required by nothing and not ordered after the network, so its failure or timeout fails no target the session needs, and boot proceeds past it regardless. It runs as the installing user, so =$HOME= points at that user's maint state, and a root =ExecStartPre= moves the flag into the unit's runtime directory before the attempt. Success, failure, a timeout, SIGKILL and power loss therefore all disarm: the intent survives exactly one reboot, and a failed attempt never wedges later boots. Output goes to tty1 rather than the console, because ratio's =/dev/console= is ttyS0. The 20-minute start timeout is sent as SIGINT, which pacman honors at a package boundary, releasing =db.lck=; the unit never deletes it. A failed or timed-out run leaves the unit failed, and maint's failed-units row names it. Exact behavior: Phase 3, The unit.
-*For the implementer.* A persistent arm flag (a file on a non-tmpfs path, e.g. =/var/lib/archsetup/apply-upgrade-on-boot=, so it survives the reboot the =/run= guard sentinel cannot). A system oneshot, =archsetup-boot-upgrade.service=, =ConditionPathExists= on the flag, ordered =Before=getty@tty1.service= so it completes before autologin execs Hyprland — this ordering is mandatory, because a parallel run would let Hyprland start mid-swap and reintroduce the exact crash the guard prevents. The unit is bounded (=TimeoutStartSec=) and best-effort: its failure or timeout must not fail any target the session needs, so boot proceeds past it regardless. Its =ExecStart= runs, as the user: =informant read= (clear the news hook that would otherwise abort the transaction), then =topgrade --only system= (or the equivalent =yay -Syu=), then =maint stamp topgrade= on success, then removes the flag unconditionally (a one-shot arm — a failed attempt disarms rather than retrying every boot). =sudo= works unattended (=%cjennings NOPASSWD: ALL=), so no password prompt wedges it.
+*The boot form.* =--apply-armed= runs only from the boot unit. It installs the armed list from the package cache in one transaction, with no refresh and no network, so it applies exactly the versions checked at arming, and the libraries are current before anything maps them, with no second reboot. Entries already installed at or above their armed version are dropped, so it never downgrades, and it refuses if the list would pull in a kernel-side package. Before the transaction it writes an interrupted record naming the remedy, so even SIGKILL or power loss leaves the next step on record, and its first tty1 line says an upgrade is running and not to power off. It also clears informant before the transaction: offline, the clear and informant's hook both fail fast and see an empty feed, so neither blocks. A network that comes up between the clear and the transaction can still let the hook abort it, which fails the attempt with the arm already consumed. If the run fails or is cut off, the tty1 login still appears and the record names what happened. An interrupted run can leave part of the set upgraded, so its remedy is =upgrade-guarded --complete= from a console with Hyprland stopped. There is no news capture, kernel step, gate, yay or sweep at boot. Exact behavior: Phase 1, =--apply-armed= sequence.
-The stamp also needs to happen when the upgrade is completed by other safe means — the by-hand sentinel-override path, or a =maint= command that does the same thing. The cleanest single home for the stamp is a small =maint apply-upgrade= (or a flag in the existing lever) that performs the guarded system upgrade and stamps on success, which both the boot unit and an interactive TTY run call. That keeps one code path that "completes a guarded upgrade and records it," rather than three writers that can drift.
+*Interruption and exit codes.* Until its finish writes the record, every mode traps INT, TERM and HUP, waits for the running child, records the interruption with its step and remedy (=--dry-run= records nothing), keeps any open gate entry as written, and exits 130. A signal after that write rewrites nothing: at =--complete='s reboot prompt it counts as no and the exit stays 0. A failed step, a refusal before any transaction, a failed gate and an interruption each have their own exit code, and on any non-zero exit the last stderr line names the reason. Exact behavior: Phase 1, Exit codes and interrupts.
+
+=upgrade-guarded= is the one code path that completes an upgrade and records it. A by-hand sentinel-override upgrade outside the script doesn't stamp, though the panel's row still clears, because the probe drops entries already installed. =upgrade-guarded --complete= from a console replaces that ritual.
* Alternatives Considered
** A. Decouple the stamp from topgrade's exit code (stamp on any real run)
- Good, because it is a one-line change to the wrapper and needs no boot machinery.
-- Bad, because it throws away honest signal: a topgrade that was blocked from applying a real upgrade would read as "fresh," so the metric stops meaning "up to date." On this machine the blocked case is the common case, so the metric would be fresh precisely when an upgrade is outstanding.
+- Bad, because it throws away honest signal: a topgrade that was blocked from applying a real upgrade would read as "fresh," so the metric stops meaning "up to date." On both daily drivers the blocked case is the common case, so the metric would be fresh precisely when an upgrade is outstanding.
- Neutral, because the failed steps still surface elsewhere (pending-updates count), so freshness would become redundant rather than wrong.
** B. Run the full topgrade live with the guard overridden, then reboot
-- Good, because it needs no new unit — arm the sentinel, run, reboot.
-- Bad, because the dangerous window is the whole rest of the run: mesa swaps early, then topgrade spends minutes on other ecosystems while the live compositor is one new GL context (a new window, the wallpaper daemon) away from SIGABRT. topgrade's own reboot-at-end is that window, not a fix for it.
-- Neutral, because it would stamp naturally on success — if it survived.
+- Good, because it needs no new unit: set the guard's =/run= override sentinel, run, reboot.
+- Bad, because the dangerous window is the whole rest of the run: mesa swaps early, then topgrade spends minutes on other ecosystems while the live compositor is one new GL context (a new window, the wallpaper daemon) away from SIGABRT. topgrade's own reboot-at-end is that window, not a fix for it. It would also land the kernel and DKMS sets with no module check.
+- Neutral, because it would stamp naturally on success, if it survived.
** C. Manual TTY ritual only (log out, run topgrade at the console, reboot), plus stamp
-- Good, because it is the safest path and needs almost no code — just make the completion stamp.
-- Bad, because it is all manual, every guarded-upgrade day; the friction is why it won't happen consistently, which is how the metric got stale in the first place.
-- Neutral, because it is exactly what the boot unit automates, so it is really "v1 minus the automation."
+- Good, because it is the safest path for the GPU/compositor set and needs almost no code, just the completion stamp.
+- Bad, because it is all manual, every guarded-upgrade day; the friction is why it won't happen consistently, which is how the metric got stale in the first place. A console topgrade also lands the kernel and DKMS sets with no module check.
+- Neutral, because it survives as the fallback: =upgrade-guarded --complete= from a console with Hyprland stopped is this ritual through the same script, gate included, and D automates its GPU/compositor half.
-** D. Boot-time armed oneshot, arch-only (this spec)
-- Good, because the risky swap happens with nothing live, the run is one bounded transaction with a tiny prompt surface, it stamps on success, and a failure degrades to "boots normally, try again."
-- Bad, because it puts a unit on the boot critical path, which must be bounded and non-fatal with care, and it is the most to build.
-- Neutral, because it composes with C: the same =maint apply-upgrade= path serves both an interactive TTY run and the boot unit.
+** D. Boot-time armed oneshot, pacman-only (the GPU/compositor half of this spec's completion step)
+- Good, because the GPU/compositor swap happens with nothing live, as one bounded transaction from the package cache with a tiny prompt surface. Arming pre-downloads the exact versions, so the boot run needs no network. It stamps when it leaves nothing deferred, and a failure or interruption degrades to a normal boot with the flag already consumed, finished by running =upgrade-guarded --complete= again.
+- Bad, because it puts a unit on the boot critical path, which must be bounded and non-fatal with care, and it is the most to build. The flag also has to follow the sync db between arming and the reboot.
+- Neutral, because it composes with C: the same =upgrade-guarded= script serves both, =--complete= from a console and =--apply-armed= from the boot unit. Kernels never ride it: they land in =--complete='s gated first stage, before anything arms, and both arming and the boot run refuse a target list that names a kernel-set or DKMS-set member.
+- Exact behavior: Phase 1 (the split-upgrade script) and Phase 3 (the boot-time unit).
** E. Split the live run: apply the non-blocked remainder now, defer the rest (this spec's everyday path)
-- Good, because it is what a careful operator does by hand anyway — and did, on ratio, the day this was decided. The live run succeeds on the common day, topgrade exits 0, the AUR and every other ecosystem stay current, and the guard's abort becomes the rare path rather than the default.
-- Bad, because Arch calls any =--ignore= run a partial upgrade. In practice pacman still enforces declared dependencies, so anything needing the newer mesa fails resolution instead of installing broken; the residual exposure is a package with an *unversioned* dependency built against a new ABI, which for mesa/wayland/libdrm is rare. Named, accepted.
-- Bad, because a deferred set nobody surfaces is a set that silently never lands — the same trap as the freshness stamp, one layer down. So the script must record the deferred set durably and the panel must show it; this is why D stays in the design as the completion step rather than being replaced.
-- Neutral, because it does not change the guard at all; it changes who decides the transaction's contents.
+- Good, because it is what a careful operator does by hand anyway, and did, on ratio, the day this was decided. The live run succeeds on the common day, the sweep exits 0, the AUR and every other ecosystem stay current, and the guard trips only in the rare case of an AUR upgrade that needs a held GPU/compositor package while Hyprland is live.
+- Bad, because Arch calls any =--ignore= run a partial upgrade. In practice pacman still enforces declared dependencies, so a package needing the newer mesa never installs against the old one. But under =--noconfirm= one unresolvable dependent fails the whole transaction (pacman's skip prompt defaults to no), which is why the script has pacman compute the dependency closure first, and refuses with the blocker named rather than force anything. The residual exposure is a package with an *unversioned* dependency built against a new ABI, which for mesa/wayland/libdrm is rare. Named, accepted.
+- Bad, because a deferred set nobody surfaces is a set that silently never lands, the same trap as the freshness stamp one layer down. So the script records the deferred set durably and the panel shows it, and this is why D stays in the design as the completion step rather than being replaced.
+- Neutral, because it does not change the guard at all; it changes who decides the transaction's contents: the script holds the GPU/compositor set while Hyprland is live, and the kernel and DKMS sets on every run.
+- Exact behavior: Phase 1 (the split-upgrade script).
* Decisions [7/7]
@@ -109,96 +166,5458 @@ CLOSED: [2026-08-25 Tue 18:45]
** DONE Everyday mechanism is the split live run (Alternative E); the boot oneshot (D) completes the deferred set
CLOSED: [2026-08-25 Tue 18:30]
- Owner / by-when: Craig / 2026-08-25
-- Context: the first draft made D the primary gesture, which means every guarded-library day is a reboot day. Craig's read while watching the ratio run: when the guard would trip, the rational move is to upgrade everything *except* the guarded and kernel items, then run the rest of topgrade without the yay piece — and that logic should be a script we can keep editing, not something baked into the panel.
-- Decision: I will build E as the path UPDATE and TOPGRADE always take on a live session, and keep D as the way the deferred set lands (arm + reboot). C remains the manual fallback through the same script from a TTY (no compositor → nothing blocked → a full run). B stays rejected on the live-swap risk.
-- Consequences: easier — the common day is one live run that succeeds; reboots are reserved for the days the deferred set is non-empty, and even then the machine keeps working until the reboot is convenient. Harder — two lists to maintain (the guard's, read from the hook; the kernel list, owned by the script), a state file the panel must render, and the partial-upgrade caveat above to keep an eye on.
+- Context: the first draft made D the primary gesture, which means every guarded-library day is a reboot day. My read while watching the ratio run: when the guard would trip, the rational move is to upgrade everything *except* the guarded and kernel items, then run the rest of topgrade without the yay piece — and that logic should be a script we can keep editing, not something baked into the panel.
+- Decision: I will build E as the path UPDATE and TOPGRADE always take, and keep D as the way the deferred set completes. =upgrade-guarded --complete= lands the kernel and DKMS sets behind the gate (the kernel decision below), then, while Hyprland is live, arms the boot oneshot to land the GPU/compositor entries on the next boot. C remains the manual fallback through the same script from a console with Hyprland stopped: an everyday run there holds only the kernel and DKMS sets, and =--complete= there applies the GPU/compositor set directly instead of arming. B stays rejected on the live-swap risk.
+- Consequences: easier — the common day is one live run that succeeds; reboots are reserved for the days I land the deferred set, and even then the machine keeps working until the reboot is convenient. Harder — two sources to keep straight (the guard's list, read from the hook; the kernel and DKMS sets, which the script derives from the installed system), a record the panel must render, and the partial-upgrade caveat in Alternative E, which the script meets by computing the dependency closure of everything it holds and refusing before any transaction when that closure can't be resolved.
+- Exact behavior: Phase 1, Everyday sequence and =--complete= sequence.
** DONE The script lives in archsetup beside the guard, and maint calls it
CLOSED: [2026-08-25 Tue 18:45]
- Owner / by-when: Craig / 2026-08-25
-- Context: the panel (dotfiles =maint=) and the guard (archsetup =scripts/hypr-live-update-guard=, installed to =/usr/local/bin=) live in different repos, and maint already carries its own copy of the trigger list as =[updates] guard_patterns= in the thresholds TOML.
-- Decision: ship the split script in archsetup next to the guard, installed by the same installer step, reading the blocked list from the installed hook so the guard and the script can never disagree. maint's UPDATE/TOPGRADE levers change their =argv= to the script; the TOML patterns stay as the panel's *display-side* mirror (the badge that says a run will defer) and gain a test asserting they match the hook.
-- Consequences: easier — one owner for "which libraries are dangerous," and a rebuilt machine gets the script with the guard. Harder — a cross-repo change (archsetup ships it, dotfiles wires it), so the rollout is two commits, archsetup first.
+- Context: the panel (dotfiles =maint=) and the guard (archsetup =scripts/hypr-live-update-guard=, installed to =/usr/local/bin=) live in different repos, and maint already carries its own, divergent pattern list as =[updates] guard_patterns= in the thresholds TOML (seeded from archsetup's =configs/maintenance-thresholds.toml=).
+- Decision: ship the split script, =upgrade-guarded=, in archsetup next to the guard, with its gate, =kernel-modules-check=, beside it; the installer step that installs =hypr-live-update-guard= installs both to =/usr/local/bin=. The script reads the GPU patterns from the installed hook's =Target= lines, so the guard and the script can never disagree, and refuses to run without them. maint's UPDATE and TOPGRADE levers change their =argv= to the script: UPDATE runs =upgrade-guarded --no-topgrade=, TOPGRADE runs =upgrade-guarded=. The TOML patterns become the panel's display-side mirror, read only for the arm line that warns a press will defer at least the matched packages. They are rewritten to equal the hook's 15 =Target= patterns and pinned by an archsetup test in =tests/installer-steps/= that compares the seed TOML with the hook heredoc.
+- Consequences: easier — one owner for "which libraries are dangerous," and a rebuilt machine gets the script with the guard. Harder — a cross-repo change (archsetup ships it, dotfiles wires it), so the rollout is three ordered commit groups: archsetup Phase 1 (both scripts and the TOML rewrite), dotfiles Phase 2 (the levers and the panel), then archsetup Phase 3 (the boot unit). Each machine installs the archsetup scripts and re-copies the rewritten TOML before any Phase 2 commit reaches it.
+- Exact behavior: the Implementation phases intro (commit groups and ordering invariants); Phase 2, Levers and Guard UX; Phase 4, Rollout.
** DONE What a split run stamps
CLOSED: [2026-08-25 Tue 18:45]
- Owner / by-when: Craig / 2026-08-25
- Context: under the state framing, a run that deferred six packages left the system *not* current, yet the sweep ran and every other ecosystem is fresh.
-- Decision: the script stamps =topgrade_run= only when the deferred set is empty. When it is non-empty it writes the deferred set to its own cache key, and the panel renders that as its own state ("6 deferred — apply on reboot") rather than as stale freshness. The boot oneshot stamps when it completes the deferred set. Freshness keeps meaning "current"; the deferred badge carries the other half.
-- Consequences: easier — no signal is thrown away, and the reboot nag has a precise count behind it. Harder — one more cache key and one more probe in maint.
+- Decision: the script stamps =topgrade_run= only when the system is current: nothing left deferred, no open kernel gate, and every step it ran succeeded. On the driven path it is the only writer. doctor.py's TOPGRADE-lever stamp is deleted, and the script calls =/usr/bin/topgrade= by absolute path, so the stamping PATH wrapper (which still stamps a bare =topgrade= typed at a shell) is never in its chain. UPDATE never stamps, and neither does a run that arms the boot oneshot or finds nothing to complete. A stamp that fails fails the run rather than passing silently. The deferred half goes to the script's own record, =${MAINT_STATE_DIR:-$HOME/.local/state/maint}/upgrade_deferred.json= (maint cache key =upgrade_deferred=), which it rewrites after every run, refusals, failures and interruptions included (Phase 1, The record). maint renders the record as its own state ("6 deferred", with APPLY to start the completion) rather than as stale freshness. Freshness keeps meaning "current"; the deferred row carries the other half.
+- Consequences: easier — no signal is thrown away, and the deferred row has a precise count behind it. Harder — one more cache key and one more probe in maint, and a record format two repos must agree on. archsetup holds the canonical fixture, bound to a fake run's actual output, and dotfiles holds a byte-identical copy whose test names its source, so a field change updates both in the same rollout. And because UPDATE never stamps, freshness clears only through a TOPGRADE with nothing deferred or a completion that ends empty.
+- Exact behavior: Phase 1, Stamp predicate and The record.
** DONE Kernel set is held on every everyday run and lands only in the dedicated session, gated on the DKMS result
CLOSED: [2026-08-25 Tue 18:45]
- Owner / by-when: Craig / 2026-08-25
-- Context: I first wrote this as "install the kernel live at apply-on-reboot," on the reasoning that a kernel swap crashes nothing and the modules-vanish window ends with the reboot. Craig asked what happens on velox when the DKMS rebuild fails, and the answer changed the decision. Velox is an encrypted ZFS root with =/boot= inside the root dataset, one kernel (=linux-lts=), and =zfs-dkms=. On a kernel upgrade the DKMS build runs PostTransaction, after the kernel is swapped and the old modules are deleted, so nothing can abort; a failed build leaves a new kernel beside an initramfs built for the old one, whose =zfs.ko= won't load, and the next boot can't import the pool. It is survivable — ZFSBootMenu can boot the pre-pacman snapshot, which holds the old kernel, initramfs, and modules — but it is a recovery session, not an update. Ratio (btrfs root, two kernels, zfs only for a data pool) is exposed only at the pool. The realistic triggers are a kernel major outrunning OpenZFS's supported range, a kernel upgraded without its headers, a toolchain regression, or a full disk.
-- Decision: the script holds the kernel set — every installed kernel with its =-headers=, moved as a set, never one without the other — on every everyday run, on both machines, so there is one rule rather than a per-host exception. The kernel set lands only in the dedicated session, live, while a working desktop exists for diagnosing, and the script gates what follows on the result: =dkms status= reports every DKMS module installed for the new kernel version, the initramfs is newer than the kernel image, and on a ZFS root a pre-pacman snapshot exists. A failed gate stops with the failure named and never reboots. The GPU/compositor set follows only after the gate passes — armed for the boot oneshot, or applied from a TTY. Kernels stay off the guard's list (the hook would block a TTY kernel upgrade for no reason). "Install the kernel live at apply-on-reboot" is withdrawn.
-- Consequences: easier — an everyday UPDATE can never put velox into the unbootable state, and the day the kernel moves is one Craig chose, sitting at the machine, expecting to handle issues. Harder — the kernel deferral is now standing, so the dedicated session has to happen on a cadence (security fixes ride the kernel), and the panel's deferred count carries a kernel most days; the gate is one more script to test, with fakes for =dkms status= and the image timestamps.
+- Context: I first wrote this as "install the kernel live at apply-on-reboot," on the reasoning that a kernel swap crashes nothing and the modules-vanish window ends with the reboot. Then I asked what happens on velox when the DKMS rebuild fails, and the answer changed the decision. Velox is an encrypted ZFS root with =/boot= inside the root dataset, one kernel (=linux-lts=), and =zfs-dkms=. On a kernel upgrade the old modules are removed PreTransaction (=70-dkms-upgrade=) and the rebuild runs PostTransaction (=70-dkms-install=), after the kernel is swapped; the alpm dkms script returns 0 even when a build fails, so nothing aborts. A failed build leaves a new kernel beside an initramfs that mkinitcpio wrote without =zfs.ko= (it only warns), and the next boot can't import the pool. The same chain runs with no kernel move at all when a DKMS package itself upgrades (today =zfs-dkms=, whose upgrade removes and rebuilds the module for every installed kernel). It is survivable — ZFSBootMenu can boot the pre-pacman snapshot, which holds the old kernel, initramfs, and modules — but it is a recovery session, not an update. Ratio (btrfs root, zfs only for a data pool) is exposed only at the pool; it has three kernels, including =linux-lts-strix=, a foreign local build with no headers and no zfs module, which is the GRUB default. The realistic triggers are a kernel major outrunning OpenZFS's supported range, a kernel upgraded without its headers, a toolchain regression, or a full disk.
+- Decision: the script holds the kernel set and the DKMS set on every everyday run, on both machines, whether or not Hyprland is running, so there is one rule rather than a per-host exception. Both sets are derived from what is installed, never from a fixed list. The kernel set is every installed kernel with its headers, held whole so a kernel and its headers always move together. The DKMS set is every package that ships a DKMS module (today =zfs-dkms=); its exact =zfs-utils= pin isn't listed by name, because the dependency closure adds =zfs-utils= exactly when an upstream bump breaks the pin. Only pending members ever move, so a foreign kernel (ratio's =linux-lts-strix=) is never touched. Both sets land only in the dedicated session, =upgrade-guarded --complete=, in its first stage, live on the running system (usually with the desktop up for diagnosing), and =kernel-modules-check= gates everything after that stage. Before the stage touches anything kernel-side, the session opens a gate entry in the record covering every kernel whose modules it could change, identified by pkgbase rather than by kernel version, so the gate checks whatever version stage 1 leaves installed. The gate checks that each of those kernels can boot with its modules: built, in a fresh initramfs, and on a ZFS root with a fresh pre-pacman snapshot behind it. The entry closes only when a gate passes or every kernel it names has been removed, so an interrupted or failed stage 1 leaves it open. While it is open no arm flag exists and nothing offers a reboot, and a failed gate exits 4 with the failure named. The GPU closure follows only after the gate passes: armed for the boot oneshot while Hyprland is running, or applied directly by a second stage when it isn't. If that closure would include a kernel-set or DKMS-set member, =--complete= refuses before any transaction. Kernels stay off the guard's list (the hook would block a TTY kernel upgrade for no reason). "Install the kernel live at apply-on-reboot" is withdrawn.
+- Consequences: easier — an everyday UPDATE can never put velox into the unbootable state, because no kernel and no DKMS module moves outside the gated session, and the day they move is one I chose, sitting at the machine, expecting to handle issues. Harder — the kernel deferral is now standing, so the dedicated session has to happen on a cadence (security fixes ride the kernel), and the panel's deferred count carries a kernel most days. From the start of a kernel-side stage 1 until a gate passes, the panel holds the deferred row at CRIT with REBOOT hidden, so a failed gate leaves the new kernel installed but not booted, with the desktop available for the fix, until a later =--complete= passes. On a ZFS root the session places a ZFS hold on the oldest pre-pacman snapshot taken since a minute before the entry opened, so the pre-snapshot prune can't destroy the fallback before it's needed; the hold moves only when a later session holds a newer one. And the gate is one more script to test.
+- Exact behavior: Phase 1, Sets, =--complete= sequence, Invariants, The gate and The snapshot hold.
** DONE Boot run applies exactly the deferred GPU/compositor set, nothing else
CLOSED: [2026-08-25 Tue 18:45]
- Owner / by-when: Craig / 2026-08-25
-- Context: the only packages that need a stopped compositor are the guard's trigger set; the rest run fine live and are what usually fail. The first draft phrased this as =topgrade --only system=; with the split script that wording is stale, and with the kernel decision above the kernel is not part of what boot applies either.
-- Decision: the boot oneshot runs the script's =--complete= form scoped to the deferred GPU/compositor set: one pacman transaction, no ecosystem sweep, no kernel. The full topgrade sweep stays a normal live run through the everyday path.
-- Consequences: easier — the boot path is fast, has a tiny interactive-prompt surface, and rarely fails. Harder — freshness after a boot run reflects the guarded set specifically, which is what the stamp decision above already accounts for.
+- Context: the only packages that need a stopped compositor are the guard's trigger set and what the dependency closure ties to it; the rest run fine live and are what usually fail. With the kernel decision above, the kernel and DKMS sets are not part of what boot applies either. A boot run can't count on the network, and a refresh at boot would install versions nobody checked when arming.
+- Decision: the boot oneshot runs the script's boot form, =upgrade-guarded --apply-armed=: exactly the armed GPU/compositor set, closure members included, in one pacman transaction, with no ecosystem sweep and no kernel. It installs the exact versions recorded at arming from the package cache, where arming pre-downloaded them while the network was up, so it needs no refresh and no network, and the unit has no network-online ordering. It skips news capture, the kernel step, the gate, yay and topgrade, never downgrades, and refuses rather than install anything that would bring in a kernel-set or DKMS-set member. The full topgrade sweep stays a normal live run through the everyday path.
+- Consequences: easier — the boot path is fast (one transaction from the cache, no network wait), has a tiny interactive-prompt surface (=--noconfirm=, no stdin, unread news cleared first with =informant read --all=), applies exactly the versions checked at arming, and rarely fails. Harder — the flag has to stay in step with the sync db, so every everyday run while armed whose refresh succeeds re-checks, re-downloads and rewrites it, or disarms; freshness after a boot run means the pacman side is current, not that the ecosystem sweep ran recently, which the stamp decision above accepts; and a network that comes up between the boot run's news clear and its transaction can let informant's hook abort the transaction, which the record names as a failed boot transaction with the arm already consumed.
+- Exact behavior: Phase 1, The arm flag and =--apply-armed= sequence; Phase 3, The unit.
** DONE Arm flag lives on a persistent path and is one-shot
CLOSED: [2026-08-25 Tue 18:45]
- Owner / by-when: Craig / 2026-08-25
-- Context: the guard's =/run= sentinel is tmpfs and cleared on reboot, so it cannot carry an intent across the reboot. A boot that retries forever on failure is its own outage.
-- Decision: We will use a persistent flag (=/var/lib/archsetup/=) that the boot unit removes unconditionally at the end of its attempt — success or failure disarms.
-- Consequences: easier — the intent survives exactly one reboot and a failed attempt never wedges subsequent boots. Harder — a failed attempt needs re-arming, which is correct (a human decides to try again) but is a manual step.
+- Context: the guard's =/run= sentinel is tmpfs and cleared on reboot, so it cannot carry an intent across the reboot. A boot that retries forever on failure is its own outage, and a flag removed at the end of the attempt is never removed when the attempt is killed or loses power.
+- Decision: We will use a persistent, root-owned flag, =/var/lib/archsetup/apply-upgrade-on-boot= (root:root 0644), that carries the armed list: one =<name>=<version>= line per GPU-kind upgrade target. Only =upgrade-guarded= writes it, atomically through =sudo -n=. =--complete= clears any flag right after its refresh and re-arms only from its GPU step, after the gate has passed, so no flag exists while a kernel gate is open. The boot unit moves the flag off its persistent path into its runtime directory at the start of its attempt, in =ExecStartPre=, before the boot form runs, so success, failure, a timeout, SIGKILL or power loss all disarm. The boot form reads the moved copy and never touches the flag.
+- Consequences: easier — the intent survives exactly one reboot, a failed or killed attempt never wedges subsequent boots, and the boot run applies exactly the versions checked at arming. Harder — a failed attempt needs re-arming by running =upgrade-guarded --complete= again, which is correct (a human decides to try again) but is a manual step. An attempt cut short by the unit's 20-minute timeout (sent as SIGINT, so pacman stops at a package boundary and releases =db.lck=) can leave part of the GPU/compositor set upgraded, so the boot form pre-writes an interrupted record naming the remedy (=upgrade-guarded --complete= from a console with Hyprland stopped) before its transaction. The panel only reads the flag's presence and entries.
+- Exact behavior: Phase 1, The arm flag and Invariants; Phase 3, The unit.
+
+* Review findings [129/138]
+Recorded by the 2026-10-05 spec-review (rubric: Not ready). Each finding names the spec passage, what it says or what the code does today, the risk, and the smallest change that closes it. =:blocking:= findings hold the rubric at Not ready until dispositioned. The evidence each finding was verified against is folded in its =EVIDENCE= drawer. Problem texts in the last nine (round 3 residual) items refer to earlier repair items as "gap N" or by short labels such as T11 or G6; those labels were working names and aren't defined elsewhere, so each item's title states the gap.
+
+** DONE Existing stamp writers mark a deferred run fresh :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Problem L32; Design L65; Decision 4 'What a split run stamps' L124-129; Phase 1 L155; Phase 2 L158; AC L168
+
+Decision 4 says only the script stamps topgrade_run, and only when the deferred set is empty. However, the script calls bare =topgrade --disable ... -y= and exits 0 'whatever was deferred'. Phase 2 only points the TOPGRADE/UPDATE levers' argv at the script. Neither phase retires the two existing writers that the Problem section names.
+
+Risk: On every run that defers guarded packages or the kernel, topgrade_run gets stamped twice, once by the wrapper and once by doctor. topgrade_age then reads fresh while the upgrade is still outstanding. That is the Alternative A outcome (L77) that Decisions 1 and 4 reject. The panel shows a deferred badge next to a fresh age, and AC L168 fails both through the panel and from a shell.
+
+Recommended change: Phase 1 (L155): the script invokes the packaged binary by absolute path (/usr/bin/topgrade, owned by the topgrade package), never =topgrade= through PATH, so the stamping wrapper is never in the chain. The script remains the only thing that runs =maint stamp topgrade=, and only when the deferred set is empty. Add a test: put a stamping fake =topgrade= ahead on PATH, defer a non-empty set, and assert that topgrade_run is untouched.
+
+Phase 2 (L158): delete doctor.py:169-170's =if rid == "topgrade"= stamp, because the script owns the stamp. Update the wrapper's header comment and test_topgrade_wrapper.py so they no longer describe the lever as stamping through the wrapper. Add a test: a zero-exit TOPGRADE lever whose script deferred something leaves topgrade_run unchanged.
+
+Acceptance criteria: extend L167 with "...and topgrade_run is not written when anything was deferred, whether the run starts from the panel or a shell."
+
+Blocking. Verification: confirmed.
+Disposition: accepted, folded into Design, Implementation phases, Acceptance criteria, Testing.
+:EVIDENCE:
+Spec (docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org):
+- L32 names the two writers: the PATH wrapper on rc 0, and the TOPGRADE lever in doctor.py.
+- L65 and L155: the script runs bare =topgrade --disable system,git_repos,containers -y= and "exits 0 when the live part succeeded, whatever was deferred".
+- L128 (Decision 4): "the script stamps topgrade_run only when the deferred set is empty".
+- L136: the deferred count "carries a kernel most days".
+- L158: Phase 2 only repoints argv and removes the sentinel wrap.
+
+Code and system:
+- ~/.dotfiles/hyprland/.local/bin/topgrade:28-43: stamp=1 unless -n/--dry-run/--version/-V/--help/-h; ="$real" "$@"=, then =maint stamp topgrade= on rc 0.
+- The wrapper header comment (L5-6) treats the two writers as intended: "The maint TOPGRADE lever resolves this wrapper too; doctor's own stamp is redundant but harmless."
+- =which -a topgrade= lists ~/.local/bin/topgrade before /usr/bin/topgrade. ~/.local/bin/topgrade is a symlink to ../../.dotfiles/hyprland/.local/bin/topgrade.
+- /proc/<pid>/environ PATH for waybar (2361) and the panel's sh (2086) starts ~/.local/bin:...:/usr/local/bin:/usr/bin.
+- ~/.dotfiles/maint/src/maint/doctor.py:166-170: any returncode 0 leads to =if rid == "topgrade": cache.put("topgrade_run", {"at": time.time()})=.
+- ~/.dotfiles/maint/src/maint/remedies.py:303-308: the "topgrade" lever, whose argv Phase 2 repoints at the script.
+- ~/.dotfiles/tests/maint/test_topgrade_wrapper.py:51-56 asserts the stamp on rc 0.
+- =pacman -Qo /usr/bin/topgrade= reports topgrade 17.12.2-1, which owns the real binary.
+- No bypass exists today: grep for upgrade-guarded, /usr/bin/topgrade and MAINT_NO_STAMP in maint/, hyprland/ and tests/ finds nothing.
+:END:
+
+** DONE topgrade --disable comma list is rejected by the installed topgrade :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L65; Phase 1 L155; Readiness External deps L189; history L11, L204
+
+The spec fixes the argv as =topgrade --disable system,git_repos,containers -y= and relies on it exiting 0. Line 189 still lists =topgrade --only system= as the verified dependency.
+
+Risk: Copied as written, the sweep fails argument parsing on every run before any step runs. No ecosystem is ever swept and topgrade never exits 0. A Phase 1 test with a faked binary that copies the spec's argv would pass and lock the bug in.
+
+Recommended change: 1. In Design L65 and Phase 1 L155, replace "topgrade --disable system,git_repos,containers -y" with "topgrade --disable system git_repos containers -y". Add a note that each step name is its own argv element, because the installed topgrade's clap parser rejects a comma-joined list with exit 2.
+2. Add to the Phase 1 tests: assert that the topgrade argv has each disabled step as a separate element. Add one check against the real binary: "/usr/bin/topgrade <the script's argv minus -y> --version" exits 0, so the faked harness cannot hide a parse error.
+3. In L65, say what a non-zero topgrade exit does. Either it counts against the script's exit and blocks the stamp, or the spec states that it doesn't, so a skipped sweep is never stamped silently.
+4. Replace the stale "topgrade --only system" in L189 and L193 (and the boot ExecStart in L69) with the script's --complete form, matching Decision L142.
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Design, Implementation phases, Readiness dimensions, Risks, Testing. Accepted, with two additions. First, the binary is called by absolute path, per F01 and my lever call. Second, the argv gains --no-ask-retry. The script ships from archsetup and runs under the panel's no-tty runner, so whether a failed step waits on stdin shouldn't depend on the user's topgrade.toml (no_retry = true today). Item 4's replacement for the stale --only system text names the boot form --apply-armed rather than --complete, per F08.
+:EVIDENCE:
+Commands run on ratio (read-only):
+- pacman -Q topgrade -> "topgrade 17.12.2-1"
+- /usr/bin/topgrade --disable system,git_repos,containers --version -> "error: invalid value 'system,git_repos,containers' for '--disable <STEP>...' ... tip: a similar value exists: 'system'", exit=2
+- /usr/bin/topgrade --disable system git_repos containers --version -> "topgrade 17.12.2", exit=0
+- /usr/bin/topgrade --disable system --disable git_repos --disable containers --version -> exit=0
+- /usr/bin/topgrade --disable system git_repos containers -y --version -> exit=0
+- topgrade --help line 40: "--disable <STEP>..." (the possible values include system, git_repos and containers)
+
+Spec (docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org):
+- L65 and L155 give the comma-joined argv.
+- L65: "exits 0 when the live part succeeded, whatever was deferred".
+- L128/L155: the stamp is gated only on an empty deferred set.
+- L141-142: the decision says the "--only system" wording is stale.
+- L69, L189 and L193 still say "topgrade --only system".
+
+Code:
+- ~/.dotfiles/maint/src/maint/remedies.py:308: ["topgrade", "--disable", "git_repos", "-y"]
+- todo.org (~L3129): "TOPGRADE exited 2 on every launch ... passed --disable git, rejected by clap before any step ran."
+:END:
+
+** DONE --noconfirm installs ignored packages, so soname cascades abort the split run :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L65 ('on the driven path it never fires'); Alternative E L97 ('fails resolution instead of installing broken'); Phase 1 L155; AC L167
+
+The spec assumes --ignore cleanly removes the blocked set, and that anything needing a newer blocked package fails resolution. On that basis it says the guard never fires on the driven path. It specifies no behaviour for two cases: a pending package that needs a held package's new version, and a held package that pins an old version of a pending one.
+
+Risk: Soname bumps in the hypr* family are routine on Hyprland release days. On those days the everyday UPDATE either re-pulls hyprutils (and the guard aborts) or fails with 'breaks dependency ... required by hyprland'. Nothing is applied, AC L167 can't be met, and the guard's abort becomes the common path again. If a prebuilt-module package like zfs-linux-lts is ever installed, the kernel hold is silently defeated.
+
+Recommended change: Phase 1 (L155): define the deferred set as a dependency closure rather than "hook targets ∪ kernel set". Start from the version-changing hook targets plus the kernel set, then repeat until nothing new is added:
+- Add any pending package whose new version needs a held package's new version (forward).
+- Add any pending package whose upgrade would break a versioned dependency of a held installed package (reverse).
+
+The simplest way to get this is to let pacman be the resolver. After the -Sy, run read-only =pacman -Sup --ignore=<set>=. Add each package named in "unable to satisfy dependency ... required by X" (add X) or in "installing X (...) breaks dependency ... required by <held>" (add X). Repeat until it exits clean. If it doesn't converge, or names a package outside the pending set (the L195 orphan case), refuse with the names. Never start a live -Syu that is known to fail.
+
+Related edits:
+- Change the L155 test to "the --ignore list is blocked ∪ kernel ∪ dependency closure". Add fixtures for both directions: hyprutils so=12→13 with hyprlang depending on =13 (forward), and an unguarded so bump that a held hyprland depends on (reverse).
+- State that the panel's deferred row (L128/L158) and the boot --complete transaction (L142) cover the whole closure in one transaction.
+- Correct L97: under --noconfirm, one unresolvable dependent fails the whole transaction (pacman's skip prompt defaults to no), which is why the closure exists.
+
+Drop the finding's "--noconfirm re-installs ignored packages" and zfs-linux-lts claims. pacman doesn't behave that way on this path.
+
+Blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Design, Alternatives, Decisions, Implementation phases, Risks, Testing. Partially confirmed. This follows the narrowed finding and my loop, and pins down what the recommendation left loose: locale and flags, both message forms, the iteration bound, the single refresh from F14, and how closure members get a kind (stage A is kernel-kind, stage B is GPU-kind).
+
+One rule is made exact. A reverse 'breaks dependency' line adds the upgrading package only when the requirer is already held. When the requirer is unheld and not pending, the line names a package outside the pending set: the orphan case I want refused. Adding the upgrading package there would silently hold it, and its dependents, indefinitely, instead of naming the orphan.
+:EVIDENCE:
+Spec: L65 (-Syu --noconfirm --ignore=<blocked set>; "on the driven path it never fires"; the 724/730 proof of concept on a mesa day). L97 ("fails resolution instead of installing broken"). L142 (boot run applies exactly the deferred set, one transaction). L155 ("the --ignore list is exactly blocked ∪ kernel set"). L167 (AC: applies everything else, exits 0). L195 (orphan case: detect, name, stop).
+
+pacman v7.1.0 source (lib/libalpm deps.c and sync.c are byte-identical at installed commit 54d9411):
+- deps.c:829: spkg = resolvedep(handle, missdep, handle->dbs_sync, *packages, 0). Prompt is 0, so ignored packages are skipped without the INSTALL_IGNOREPKG question (deps.c:662-667, 693-698).
+- deps.c:762: prompt=1 only in alpm_find_dbs_satisfier.
+- sync.c:439-468: unresolvable packages raise ALPM_QUESTION_REMOVE_PKGS. With no skip, GOTO_ERR(ALPM_ERR_UNSATISFIED_DEPS).
+- callback.c:509: q->skip = noyes(...). util.c:1736-1738: under noconfirm the preset (0) is returned, so the transaction fails.
+- sync.c:654-668 with deps.c:353-380: reverse-dependency break fails prepare, no prompt.
+- src/pacman/sync.c:750/754: messages "unable to satisfy dependency '%s' required by %s" and "installing %s (%s) breaks dependency '%s' required by %s".
+
+Live system and cache on ratio:
+- The installed hook /etc/pacman.d/hooks/hypr-live-update-guard.hook targets hyprutils and aquamarine but not hyprlang or hyprtoolkit.
+- bsdtar .PKGINFO: hyprutils-0.13.1-1 provides libhyprutils.so=12-64, 0.14.0-1 provides =13-64. hyprlang-0.6.8-4 depends on =12-64, 0.6.8-5 on =13-64.
+- aquamarine-0.14.0-1 provides libaquamarine.so=13-64, 0.15.0-2 provides =14-64. hyprtoolkit-0.5.4-4 depends on =13-64, 0.5.4-6 on =14-64.
+- pacman.log: the whole hypr family co-upgraded with pkgrel-only rebuilds on 2026-04-05, 05-03, 07-21 and 09-09.
+- expac -S: hyprlang, hyprlock, hyprwire, hyprpaper, hyprtoolkit, hyprland-guiutils and xdph all depend on libhyprutils.so=13-64. hyprland depends on libhyprlang.so=2-64, libhyprwire.so=3-64 and libhyprcursor.so=0-64 (all unguarded).
+- No mesa-family package carries a versioned dependency on a guarded package.
+- zfs-linux-lts is not installed (zfs-dkms 2.4.4-1 is).
+:END:
+
+** DONE zfs-dkms and its pinned zfs-utils bypass the kernel hold :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Decision 5 L134-136 (consequence: 'an everyday UPDATE can never put velox into the unbootable state'); Design L65; Phase 1 L155; Risks L196
+
+Only the kernel set (kernels plus -headers) is held on everyday runs, and the DKMS failure chain is treated only as a consequence of a kernel upgrade. Other DKMS-triggering packages go through the live --ignore run.
+
+Risk: An everyday run that moves zfs-dkms removes the working zfs.ko before the transaction and rebuilds it afterward, with nothing able to abort. A failed build (toolchain regression or full disk, both triggers Decision 5 lists) leaves velox with an initramfs that lacks zfs.ko. That is the unbootable state the kernel decision exists to prevent, reached on a run nobody chose. Holding zfs-dkms alone doesn't work either, because the exact zfs-utils pin makes the --ignore run fail dependency resolution.
+
+Recommended change: Recommended (needs my call because it widens Decision 5's held set): define a DKMS set and hold it next to the kernel set.
+- L135 (Decision 5) and L65 (Design): add "plus the DKMS set: every package owning /usr/src/*/dkms.conf (pacman -Qoq), together with any dependency it pins with '=' (today zfs-dkms + zfs-utils), held whole on every everyday run and applied in --complete with the kernel set, before kernel-modules-check runs."
+- L155 (Phase 1): change "the --ignore list is exactly blocked ∪ kernel set" to "... blocked ∪ kernel set ∪ DKMS set". Add a test that a pending zfs-dkms/zfs-utils upstream bump is held as a pair, and that a zfs-utils pkgrel-only bump is not held.
+- L168: change "everything but the kernel set" to "everything but the kernel and DKMS sets".
+
+If I decline to widen the set, make two smaller edits instead:
+- Rewrite L136 to "an everyday UPDATE never moves the kernel; a zfs-dkms move on an everyday run remains exposed, with the 05-zfs-snapshot/ZFSBootMenu fallback."
+- Have the everyday run call kernel-modules-check whenever it moved a dkms.conf-owning package, and surface any failure as a named, do-not-reboot warning.
+
+Either way, fix L196: "the DKMS hooks are PostTransaction" becomes "old modules are removed PreTransaction and the rebuild runs PostTransaction".
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Decisions, Design, Implementation phases, Acceptance criteria, Risks, Testing. Confirmed and accepted. This widens Decision 5's held set, recorded as a body rewording plus a history line, not a reopening. The '= pin' half comes through F03's closure instead of a separate parse. On ratio, installed zfs-dkms 2.4.4-1 depends on zfs-utils=2.4.4. With zfs-dkms held, pacman reports that an upstream zfs-utils bump breaks that pin, so the closure adds zfs-utils exactly then and leaves a pkgrel-only bump alone. One mechanism produces the pair and pkgrel test outcomes the recommendation asks for.
+:EVIDENCE:
+- Spec L65: blocked set = guard list "plus the kernel set (every installed kernel with its -headers ...)". Spec L155: kernel set = "every linux* kernel package and its -headers"; test: "the --ignore list is exactly blocked ∪ kernel set". L136: "an everyday UPDATE can never put velox into the unbootable state". L168: TTY run "applies everything but the kernel set". L196: "(the DKMS hooks are PostTransaction)".
+- /usr/share/libalpm/hooks/70-dkms-upgrade.hook: Operation=Upgrade, Target=usr/src/*/dkms.conf, When=PreTransaction, Exec=/usr/share/libalpm/scripts/dkms -D remove.
+- 70-dkms-install.hook: same targets, When=PostTransaction, dkms install.
+- 90-mkinitcpio-install.hook: Target=usr/src/*/dkms.conf, PostTransaction.
+- /usr/share/libalpm/scripts/dkms main(): collects ERROR_MESSAGES and ends with "return 0".
+- /usr/bin/mkinitcpio:409: warning "errors were encountered during the build. The image may not be complete." (the image is still written). /usr/lib/initcpio/install/zfs: map add_module ... zfs spl.
+- pacman -Qoq /usr/src/*/dkms.conf: zfs-dkms.
+- pacman -Si zfs-dkms: Repository archzfs, Depends On zfs-utils=2.4.4 lsb-release dkms. zfs-utils is installed at 2.4.4-3 (a pkgrel bump allowed by the pin).
+- /var/log/ratio-upgrade.log:704-705: "dkms remove --no-depmod zfs/2.4.3 -k ...". Lines ~1487-1488: "dkms install --no-depmod zfs/2.4.4 -k ...".
+- /var/log/pacman.log:20503-20504: zfs-utils and zfs-dkms upgraded 2.4.3 to 2.4.4 in the same run.
+- Mitigation that makes it survivable rather than fatal: archsetup:2421-2432 writes 05-zfs-snapshot.hook (PreTransaction, Target=*), which runs before the 70 dkms hooks.
+:END:
+
+** DONE kernel-modules-check fails permanently on ratio :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Phase 1 L155 ('for each kernel under /usr/lib/modules'); Decision 5 L134-135; Design L67; Risks L196; Phase 4 L164
+
+Phase 1's gate requires that, for each kernel under /usr/lib/modules, dkms status reports every registered module as installed. Decision 5 and Design scope the check to 'the new kernel version'. The spec describes ratio as having two kernels.
+
+Risk: Every --complete run on ratio stops at the gate, whatever it installed, so ratio can never arm and the GPU/compositor set never lands through the designed path. Phase 4's 'confirm the ratio path matches' cannot pass, and the deferred row and stale freshness become permanent on one of the two daily drivers.
+
+Recommended change: Edit L155 (Phase 1) to match Decision 5.
+
+Replace "for each kernel under =/usr/lib/modules=, =dkms status= reports every registered module =installed= for it, and its initramfs is newer than its =vmlinuz=" with:
+
+"for each kernel version this =--complete= run installed or changed (its =/usr/lib/modules/<ver>= directory, mapped to =/boot/vmlinuz-<pkgbase>= and its initramfs through =<ver>/pkgbase=), =dkms status -k <ver>= reports every registered module =installed=, and that kernel's initramfs is newer than its =vmlinuz=. Kernels the run did not touch are not checked. When the run changed no kernel, the DKMS and initramfs checks pass vacuously."
+
+Add one test to the Phase 1 test list: "an untouched kernel that is foreign, has no headers, and has DKMS modules only 'added' (ratio's =linux-lts-strix= shape) does not fail the gate."
+
+Correct "two kernels" at L134 and L196 to "three kernels, including =linux-lts-strix=, a foreign local build with no headers and no zfs module, which is the GRUB default."
+
+Related, optional: say in Phase 1 that the derived kernel set excludes packages absent from the sync repos, because "applied whole" cannot apply =linux-lts-strix= (pacman -Si: not found).
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Design, Decisions, Implementation phases, Risks, Testing. Confirmed. Scoping the gate to kernels this run changed is right, but on its own it leaves two holes, and both are closed here.
+- F04 puts the DKMS set into --complete's stage 1. A zfs-dkms-only move rebuilds modules for kernels that were already installed, so those kernels join the check list.
+- After a failed gate the new kernel is already installed. A second --complete changes no kernel, so it would pass vacuously, clear the CRIT row, and arm. The open failure's kernels are therefore re-checked against the recorded since.
+:EVIDENCE:
+Spec lines in docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L155: "for each kernel under =/usr/lib/modules=, =dkms status= reports every registered module =installed= for it"
+- L135: "=dkms status= reports every DKMS module installed for the new kernel version"
+- L67: "every DKMS module built for the new kernel"
+- L134 and L196: "Ratio (btrfs root, two kernels, ...)"
+- L164: "Confirm the ratio path matches."
+
+Live checks on ratio:
+- uname -r: 6.18.25-1-lts-strix
+- /proc/cmdline: BOOT_IMAGE=/@/boot/vmlinuz-linux-lts-strix
+- /etc/default/grub:2: GRUB_DEFAULT="gnulinux-linux-lts-strix-advanced-..."
+- ls /usr/lib/modules: 6.18.25-1-lts-strix, 6.18.54-1-lts, 7.2.7-arch1-1
+- pkgbase files: linux-lts-strix (no build/ dir), linux-lts (build/ present), linux (build/ present)
+- pacman -Qm: linux-lts-strix 6.18.25-1, Packager "Unknown Packager"
+- pacman -Si linux-lts-strix and pacman -Si linux-lts-strix-headers: "error: package ... was not found"
+- dkms status: "zfs/2.4.4, 6.18.54-1-lts, x86_64: installed" and "zfs/2.4.4, 7.2.7-arch1-1, x86_64: installed"
+- dkms status -k 6.18.25-1-lts-strix: "zfs/2.4.4: added"
+- modinfo -n zfs: "ERROR: Module zfs not found."
+- pacman -Qi zfs-dkms: 2.4.4-1 installed
+- /var/log/ratio-upgrade.log:1536: "Building image from preset: /etc/mkinitcpio.d/linux-lts-strix.preset", written 2026-08-25, the same day the decisions were closed.
+- docs/workflows/strix-soak-watch.org says the strix kernel is a standing local build, pinned as the GRUB default, until it is retired.
+
+Not a problem for the initramfs half: /boot/initramfs-linux-lts-strix.img (Sep 29) is newer than /boot/vmlinuz-linux-lts-strix (Apr 30), so that check passes.
+:END:
+
+** DONE Phase 1 and Decision 2 hold the kernel only when a compositor is live :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Phase 1 L155; Decision 2 L114; Non-Goals L51; Decision 5 L135; AC L168
+
+Phase 1 computes 'blocked set ... ∪ held-kernel set when a compositor is live'. Decision 2 calls a TTY run 'no compositor → nothing blocked → a full run', and Non-Goals L51 says 'holds them back on a live run'. Decision 5 and AC L168 hold the kernel on every everyday run, TTY included, and Phase 1's own test list says 'held whole on every everyday run'. The spec never says how liveness is detected.
+
+Risk: Read literally, the documented TTY fallback (C) moves velox's kernel with no DKMS gate. That is the unbootable path Decision 5 removes, and it would ship in v1 code paths and tests. If the script checks for a TTY instead of the process, a tty2 run with Hyprland still up skips the guarded patterns and the guard aborts.
+
+Recommended change: Five small text edits, no decision change:
+1. L155: replace "blocked set (hook =Target= lines, version-aware) ∪ held-kernel set when a compositor is live" with "blocked set (hook =Target= lines, version-aware; only when Hyprland is running, tested exactly as the guard tests it: =pgrep -x Hyprland=) ∪ kernel set (always, on every run except =--complete=)".
+2. L155 tests: add "with a liveness seam forced on and off (as =HYPR_GUARD_RUNNING= does for the guard), =--ignore= contains the kernel set on both branches and the blocked set only on the live branch."
+3. L114: change "(no compositor → nothing blocked → a full run)" to "(no compositor → no guarded library blocked → everything but the kernel set; the kernel still lands only through =--complete=)".
+4. L51: change "on a live run" to "on every run outside the dedicated =--complete= session".
+5. Optional, one clause at L67 or L168: "'no compositor' means Hyprland is not running, not merely a console; Ctrl+Alt+F2 leaves it live on tty1."
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Goals and Non-Goals, Decisions, Design, Implementation phases, Acceptance criteria, Testing. Confirmed. The DKMS set (F04) joins the kernel set wherever this finding says 'kernel set'.
+
+The seam is the script's own name. The guard reads HYPR_GUARD_RUNNING from pacman's environment, which sudo resets, so sharing the name would suggest one variable drives both.
+
+--complete's stage 1 holds the GPU closure even from a TTY, because Decision 5 puts the GPU set after the gate.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L155: "pending set → blocked set (hook =Target= lines, version-aware) ∪ held-kernel set when a compositor is live"; the same paragraph's tests say "held whole on every everyday run".
+- L114 (Decision 2, CLOSED 18:30): "C remains the manual fallback through the same script from a TTY (no compositor → nothing blocked → a full run)."
+- L51: "the split script holds them back on a live run".
+- L135 (Decision 5, CLOSED 18:45): "holds the kernel set ... on every everyday run, on both machines ... The kernel set lands only in the dedicated session".
+- L168 (AC): "The same run with no compositor live (a TTY) applies everything but the kernel set".
+- L12/L208: the 18:45 reversal to "held on every everyday run".
+- No line in the spec names a liveness test (grep for compositor/live/tty/pgrep).
+
+Guard scripts/hypr-live-update-guard:
+- L57-62: hyprland_running() uses the HYPR_GUARD_RUNNING override (documented at L32), else pgrep -x Hyprland.
+- L67: hyprland_running || exit 0.
+- L130-131: the banner says "from a TTY with Hyprland stopped: 1. Log out of Hyprland, or switch to a console (Ctrl+Alt+F2)". The Ctrl+Alt+F2 option leaves Hyprland running on tty1, so "TTY" and "no compositor" are not the same.
+:END:
+
+** DONE Phase 3 panel action is the withdrawn ungated kernel path :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Scope tiers v1 L55; Phase 3 L161; against Decision 5 L135, Design L67, Phase 1 L155, history L208
+
+L55 describes an 'apply on reboot' affordance that 'installs the held kernel live and arms the boot-time oneshot'. L161 says 'The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot', with 'the arm action's tests in maint'. Neither mentions kernel-modules-check.
+
+Risk: An implementer following Phase 3 builds a maint-side button that swaps the kernel, writes the flag and offers a reboot, with no DKMS, initramfs or snapshot check. On velox, a failed zfs-dkms build then reboots into a kernel whose initramfs can't import the pool. That is the exact failure Decision 5 was closed to prevent.
+
+Recommended change: L55: replace "an 'apply on reboot' affordance that installs the held kernel live and arms the boot-time oneshot for the GPU/compositor set" with "an 'apply on reboot' panel action that runs upgrade-guarded --complete (kernel set live, then the kernel-modules-check gate, then arm the boot-time oneshot for the GPU/compositor set)".
+
+L161: replace the first sentence with "The panel action runs upgrade-guarded --complete in the foreground (tty PROMPT mode). The script does every step: kernel set live, the gate, the arm flag and the reboot offer. maint never installs packages or writes the flag itself, and on a non-zero exit it shows the named gate failure and offers no reboot." Replace "the arm action's tests in maint" with "maint tests that the action's argv is upgrade-guarded --complete and that a non-zero exit leaves no reboot offer".
+
+L155: add one clause: "--complete exits non-zero when the gate fails."
+
+Optionally add an acceptance line after L169: "The panel's apply-on-reboot action, on a failed gate, neither arms nor offers a reboot."
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Goals and Non-Goals, Design, Implementation phases, Acceptance criteria, Testing. Confirmed, and reconciled with F40, which is also accepted. The action opens a detached terminal rather than using the lever runner's tty PROMPT mode. The panel then can't read --complete's exit code, so 'a non-zero exit leaves no reboot offer' becomes state-driven: F12's gate record hides REBOOT, and F34's arm flag shows it. foot gets --hold so the gate verdict stays readable after the script exits. The action moves into Phase 2 with its row (F16).
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L55 (v1 scope): "an 'apply on reboot' affordance that installs the held kernel live and arms the boot-time oneshot for the GPU/compositor set". No gate.
+- L161 (Phase 3): "The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot. ... the arm action's tests in maint". No --complete, no kernel-modules-check.
+- L134-135 (Decision 5, closed): "I first wrote this as 'install the kernel live at apply-on-reboot'... A failed gate stops with the failure named and never reboots. The GPU/compositor set follows only after the gate passes... 'Install the kernel live at apply-on-reboot' is withdrawn."
+- L208 (history, 18:45): "withdrew 'install live at apply-on-reboot'... added the kernel-modules-check gate to Phase 1 and the acceptance criteria".
+- L67 (Design): the dedicated session is "run in the foreground from the panel's action or upgrade-guarded --complete in a terminal: first the kernel set, live... then the gate... If the gate fails the script stops there... and does not reboot... If it passes, the script arms a persistent flag... and offers to reboot". This contradicts L161.
+- L155 (Phase 1): --complete runs the kernel set, then the gate, then arms; "--complete refuses to arm or reboot on that exit". Only kernel-modules-check is stated to exit non-zero. --complete's own exit code on a failed gate is never specified.
+- L169 (acceptance): the DKMS-failure criterion covers only "A --complete run", not the panel action.
+- Current maint (~/.dotfiles/maint/src/maint/doctor.py:161-176) runs lever steps by argv and has a tty PROMPT mode. Routing the action to a script argv fits the existing mechanism with no new maint-side package logic.
+:END:
+
+** DONE Boot run's ExecStart and scope contradict the boot-scope decision, and no boot form is defined :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design 'For the implementer' L69; Decision 4 L128; Decision 6 L138-143; Phase 1 --complete L155; Phase 3 L161; Readiness L189; Risks L193
+
+L69: ExecStart runs, as the user, 'informant read ..., then =topgrade --only system= (or the equivalent =yay -Syu=), then =maint stamp topgrade= on success, then removes the flag'. Phase 3 L161: ExecStart 'runs the script's --complete form'. Phase 1 defines --complete as 'apply the kernel set live, run the gate, then arm ... or, with no compositor live, apply it directly'. Decision 6 L142: the boot run applies exactly the deferred GPU/compositor set in one pacman transaction, with no sweep and no kernel. L69 stamps on any success, whereas Decision 4 stamps only when the deferred set is empty. No text says whether the boot run reads the set recorded at arm time or re-derives it with a sync refresh, or that it rewrites the deferred-set record.
+
+Risk: Whichever text the implementer follows, the boot run does one of two bad things. It may install a kernel (including one published since arming) at the console, with no gate and no desktop to diagnose from: the velox unbootable chain Decision 5 forbids. Or it runs the gate at boot and never applies the GPU set on ratio. A refresh at boot can also pull in packages that were never armed. The stamp can fire while a kernel is still deferred, and the panel keeps the pre-boot 'N deferred' count after a successful run.
+
+Recommended change: Define one boot-only form in Phase 1 (L155) and point every boot reference at it.
+
+1. Add to L155's flag list: =--apply-armed=, the boot form. It reads the GPU/compositor package list that --complete's arm step writes into the flag file. It runs =informant read= if installed, then a single =sudo pacman -S --needed --noconfirm <that list>=. It does not use -y. A refresh followed by a targeted -S is itself a partial upgrade and could pull versions that were never armed. It runs no kernel set, no kernel-modules-check, no yay, and no topgrade. On success it rewrites the deferred-set state file and stamps topgrade_run if the deferred set is now empty. Add one sentence to --complete: "arming writes the GPU/compositor package list into the flag file".
+2. Replace L69's ExecStart sentence, from "then =topgrade --only system=" through "on success", with: "runs =upgrade-guarded --apply-armed= as the user". Keep the unconditional flag removal.
+3. In L142 and L161, change "the script's =--complete= form (scoped to the deferred GPU/compositor set)" to "the script's =--apply-armed= form".
+4. In L71 and L93, replace "=maint apply-upgrade=" as the boot unit's path with "=upgrade-guarded=" (the archsetup script, per Decision 3). Its TTY form is --complete and its boot form is --apply-armed.
+5. In L189, drop "topgrade =--only system=". In L193, change "The =--only system= scope" to "The recorded-set scope".
+6. Add these Phase 1 tests: the --apply-armed transaction targets exactly the recorded list, never includes a kernel-set package, and never passes -y.
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Design, Decisions, Implementation phases, Readiness dimensions, Risks, Testing. Confirmed and adopted, reconciled with two of my calls.
+- Disarm-at-ExecStartPre means the boot form reads the copy that ExecStartPre moved into /run, not the flag path.
+- Exact versions in the flag mean the transaction uses versioned targets, filtered so it never downgrades.
+
+One safety check is added. pacman -S resolves dependencies from the sync db, so a GPU target needing a newer kernel-set or DKMS-set package would pull that package in at boot, ungated. A read-only, offline pacman -Sp of the exact targets runs before the transaction and refuses that case.
+:EVIDENCE:
+Spec text:
+- L69: "Its =ExecStart= runs, as the user: =informant read= …, then =topgrade --only system= (or the equivalent =yay -Syu=), then =maint stamp topgrade= on success, then removes the flag unconditionally".
+- L71: "a small =maint apply-upgrade= … which both the boot unit and an interactive TTY run call".
+- L93: "the same =maint apply-upgrade= path serves both an interactive TTY run and the boot unit".
+- L141: "The first draft phrased this as =topgrade --only system=; with the split script that wording is stale".
+- L142: "the boot oneshot runs the script's =--complete= form scoped to the deferred GPU/compositor set: one pacman transaction, no ecosystem sweep, no kernel".
+- L155: "=--complete= (the dedicated-session form: apply the kernel set live, run the gate, then arm the GPU/compositor set or, with no compositor live, apply it directly; stamp when the deferred set is empty)".
+- L161: "=ExecStart= runs the script's =--complete= form as the user".
+- L189: "External APIs & deps: topgrade =--only system=".
+- L193: "The =--only system= scope … shrink this to near zero".
+
+Git history:
+- =git show 77447d0:…spec.org= L64 has the identical ExecStart sentence, so it was never revised.
+- The uncommitted diff touches only the script name and the containers flag, not L69.
+
+Live system:
+- ~/.dotfiles/common/.config/topgrade.toml:131 sets =arch_package_manager = "yay"=, and L134 =# yay_arguments= is commented out.
+- /etc/pacman.conf:25 has =IgnorePkg = bridge-utils= and nothing else, so no kernel is ignored.
+- =pacman -Q= shows linux 7.2.7, linux-lts 6.18.54, and linux-lts-strix 6.18.25 (no headers).
+- =dkms status= shows zfs installed for 6.18.54-1-lts and 7.2.7-arch1-1 only.
+- =topgrade --help= confirms =--only <STEP>=.
+:END:
+
+** DONE Boot run's package source and network dependency are undefined :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L65-69; Decision 6 L138-143; Phase 3 L161; Readiness Performance L182
+
+The spec never says whether the boot run refreshes the sync db, downloads packages or waits for the network. Nothing pre-downloads the GPU set at arm time, and the unit declares no network ordering. The spec also never covers a stale arm (flag set, then days of further everyday runs before the reboot) or a flag that is still set when nothing is pending any more.
+
+Risk: If the boot run refreshes and downloads, an armed boot depends on pre-login network and mirrors, network-online delays the login prompt on every boot, and a refresh can pull new versions (including a kernel) into the boot transaction. If it doesn't, there is no rule for where packages come from. With no network the run fails every time and disarms (Decision 7), so the GPU set never lands.
+
+Recommended change: 1. In Design L69, replace "then =topgrade --only system= (or the equivalent =yay -Syu=)" with "then =upgrade-guarded --complete= in its boot form". In L189 and L193, drop the =--only system= references the same way.
+
+2. Add one paragraph to Decision 6 (or Phase 3):
+
+"The boot form never refreshes the sync db and has no network dependency; the unit declares no network-online ordering. When --complete arms, it runs =pacman -Sw --noconfirm <GPU set>= while the network is up and writes the exact pkg=version list into the arm flag. While armed, each everyday run re-runs -Sw for the deferred set after its own refresh and rewrites that list. At boot the run skips the kernel step and the gate. It installs the recorded list with =pacman -S --noconfirm= against the existing db, so pacman takes every file from the cache. A missing cache file fails fast and names the package. If every recorded package is already installed at or past its recorded version, the run is a no-op: it clears the flag, rewrites the deferred-set record, and stamps only if Decision 4's empty-deferred-set rule is met."
+
+3. Add a Phase 1 test that the boot form makes no -y/refresh call. Add an acceptance criterion that an armed boot with networking down still applies the set.
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Design, Decisions, Implementation phases, Acceptance criteria, Readiness dimensions, Testing. Confirmed; this matches my boot-source call, including the recommendation's re-download while armed. That keeps the flag's versions equal to whatever the system db each everyday run synced. Three gaps are closed:
+- An armed everyday run that can't compute the closure, pass the check, or re-download disarms. Otherwise it would leave a list the db no longer matches.
+- The boot form never downgrades. A recorded version below the installed one would downgrade under --noconfirm, so such entries are dropped.
+- F08's dependency check runs at arm time and again at boot, so no kernel can ride in as a dependency.
+:EVIDENCE:
+Spec:
+- L69 names =topgrade --only system= / =yay -Syu= as the boot ExecStart.
+- L141 (Decision 6 context) calls that wording "stale". L142 scopes the run to =--complete= with "no kernel".
+- L155 defines =--complete= as starting with "apply the kernel set live".
+- L189 and L193 still cite =--only system=.
+- L182 claims negligible cost when the flag is absent.
+- A grep of the spec for Sw|download|network|online|mirror|refresh|offline finds nothing about the boot run's package source or network. The only refresh mention is the live run at L65.
+
+Live system (ratio):
+- =checkupdates= shows guarded packages pending: mesa 1:26.2.4-1, vulkan-radeon, hyprland 0.56.2-4, and kernels linux 7.2.8 and linux-lts 6.18.55. None of them is in /var/cache/pacman/pkg.
+- maint uses plain =checkupdates= with no -d (maint/src/maint/probes/updates.py:152, doctor.py:83). /usr/bin/checkupdates downloads only with -d (L63, L173).
+- No timer pre-downloads. The only timers are snapper, man-db, tmpfiles, shadow, logrotate, keyring-wkd, paccache, reflector and btrfs-scrub.
+- Boot timing (monotonic µs): basic.target 7776379, NetworkManager-wait-online 8975960→14740538, network-online.target 14740879, getty@tty1 35133024.
+- NetworkManager-wait-online.service is enabled, with =ExecStart=/usr/bin/nm-online -s -q= and =NM_ONLINE_TIMEOUT=60=.
+- network-online.target is already pulled in by docker, minidlna, reflector and archlinux-keyring-wkd-sync.
+- systemd.unit(5), Conditions and Asserts: conditions "are checked at the time the queued start job is to be executed. The ordering dependencies are still respected... neither condition nor assertion expressions are suitable for conditionalizing unit dependencies."
+- reflector.timer: OnCalendar=weekly, Persistent=true (tangential).
+:END:
+
+** DONE Timeout is unspecified, and SIGTERM kills pacman mid-transaction and leaves db.lck :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L69; Phase 3 L161; AC L171; Readiness Errors L179, Performance L182, Config surface L185; Risks L192
+
+The spec treats a timeout as benign: 'worst case is a bounded delay, then a normal session' (L192), and AC L171 says a timeout lets the machine boot into Hyprland. The TimeoutStartSec value is left open (L185). Nothing defines the system's state after pacman is killed mid-commit.
+
+Risk: A timeout during extraction leaves /var/lib/pacman/db.lck behind and a half-extracted guarded set (mixed mesa/libdrm files). getty then starts and Hyprland launches against mixed libraries. Every later pacman, yay or upgrade-guarded run fails with 'unable to lock database' until someone removes the lock by hand, and the panel shows only the old deferred count.
+
+Recommended change: 1. Phase 3 (L161) and Design (L69): give archsetup-boot-upgrade.service KillSignal=SIGINT, so a timeout reaches pacman's interrupt handler, which stops at a package boundary and releases /var/lib/pacman/db.lck itself. Give TimeoutStartSec a concrete, generous value (e.g. 20min) so it fires only on a real hang.
+
+2. Move the disarm out of the tail of ExecStart. Remove the flag with ExecStartPre=+/usr/bin/rm -f /var/lib/archsetup/apply-upgrade-on-boot (or in ExecStopPost=), so success, failure and timeout all disarm. Reword Decision L149 to match ("removed at the start of the attempt" or "in ExecStopPost").
+
+3. Define the interrupted outcome. If the transaction did not complete, or db.lck survives, the unit writes "interrupted" plus the remedy into the deferred-set state file the panel already renders. It never deletes db.lck on its own.
+
+4. Correct the false claims. In Risks L194, replace "one atomic transaction, so there is no half-swapped library" with: pacman is atomic per package only, and only when allowed to stop at a package boundary, which is why the unit uses KillSignal=SIGINT. Drop "worst case is a bounded delay" from L192.
+
+5. AC L171: add that after a timeout no db.lck is left behind, the flag is cleared even though ExecStart was killed, and the panel names the interruption.
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Design, Decisions, Implementation phases, Acceptance criteria, Readiness dimensions, Risks. Confirmed; this matches my call: SIGINT, 20min, disarm at the start, an interrupted record, and never deleting db.lck. Two reconciliations:
+- The flag now carries the list, so ExecStartPre moves it into the unit's RuntimeDirectory instead of running rm. That disarms just as completely.
+- The interrupted record is pre-written before the boot transaction, not written only by a signal trap. A trap can't run on the SIGKILL that follows TimeoutStopSec, or on power loss.
+
+The ExecStopPost alternative is dropped.
+:EVIDENCE:
+Spec:
+- L69 / L161: "TimeoutStartSec bounded … ExecStart runs … then removes the flag unconditionally".
+- L149: the flag is removed "at the end of its attempt".
+- L171: the AC requires that on "a timeout … the flag is cleared" and the machine boots into Hyprland.
+- L179: "every failure path lands in … re-arm to retry".
+- L192: "worst case is a bounded delay, then a normal session".
+- L194: "each pacman run is one atomic transaction, so there is no half-swapped library".
+- grep of the spec for killsignal|db.lck|interrupt|SIGTERM|SIGINT|ExecStopPost|atomic finds nothing except the L194 atomicity claim.
+
+systemd (v262):
+- man systemd.service: TimeoutStartSec "Defaults to DefaultTimeoutStartSec= … except when Type=oneshot is used, in which case the timeout is disabled by default". TimeoutStartFailureMode "default to terminate … sending the signal specified in KillSignal= (defaults to SIGTERM)".
+- man systemd.kill: KillMode=control-group kills "all remaining processes in the control group".
+
+pacman 7.1.0.r9:
+- strings /usr/bin/pacman shows only "Interrupt signal received" and "Hangup signal received" handler messages. No SIGTERM handler is present.
+- libalpm.so.16 contains alpm_trans_interrupt and "transaction interrupted". So SIGINT has a graceful path: it stops at a package boundary during commit and releases the lock; during download it unlocks and exits.
+
+Live system: ratio's /etc/pacman.d/hooks holds hypr-live-update-guard.hook and 99-grub-sync-efi.hook, and /usr/share/libalpm/hooks/05-snap-pac-pre.hook runs PreTransaction, so a kill can land while the lock is held before commit even starts.
+:END:
+
+** DONE Bare =informant read= hangs unattended and can't clear the news hook when run as the user :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L65, L69; Phase 1 L155; AC L172; Reuse L183; Readiness External deps L189
+
+The live script and the boot oneshot (which runs 'as the user') both run a bare =informant read= to clear the news hook before the transaction. L189 lists informant as 'verified present'.
+
+Risk: The boot oneshot has no stdin, so the bare command either blocks until TimeoutStartSec or exits without marking anything read. Run as the user it can't save read state either, so informant's AbortOnFail hook aborts the transaction. The boot run then fails every time Arch has published news, consuming the arm. AC L172 fails, and the panel-driven live run can hang or abort the same way.
+
+Recommended change: L65, L155: replace "=informant read=" with "=sudo informant read --all=" (run only when =command -v informant= succeeds; failure tolerated, as at archsetup:1127-1129). Optionally log =informant list --unread= first so the news isn't dropped unseen.
+
+L69: change "Its =ExecStart= runs, as the user: =informant read= (...)" to say that the news-clear and pacman steps run as root via sudo, and the news-clear step is =sudo informant read --all= when installed. Add the reason: =/var/lib/informant.dat= is 0664 root:informant, the user isn't in that group, and a bare read prompts between items, so it crashes on null stdin or blocks on a tty.
+
+AC L172: change "(=informant read= precedes the transaction)" to "(=sudo informant read --all= precedes the transaction; verified with two or more unread items and no stdin)".
+
+L189 (optional): note that informant is present on velox only; ratio doesn't have it installed.
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Design, Implementation phases, Acceptance criteria, Readiness dimensions, Testing. Adopted, with one hardening. informant 0.6.0's feed fetch calls requests with no timeout. On a half-up network, from the panel or early in boot (the unit has no network ordering), the call could hang until the lever's 3600s timeout or the unit's 20-minute timeout. timeout 60 bounds it, and the run carries on either way. The optional 'log informant list --unread first' is superseded by F39's yay -Pwq capture.
+:EVIDENCE:
+- Spec L65 "clears the news hook (=informant read=) where installed"; L69 "Its =ExecStart= runs, as the user: =informant read= ..."; L155 "=informant read= if present →"; L172 AC "Unread Arch news does not wedge the boot run (=informant read= precedes the transaction)"; L189 "all verified present on velox".
+- informant 0.6.0-2 source, taken from the archangel build tree at ~/code/archangel/work/x86_64/airootfs/usr/lib/python3.14/site-packages/informant/:
+ - informant.py:132-145 is the bare-read loop, which calls =ui.prompt_yes_no('Read next item?')= between items.
+ - informant.py:146 is =fs.save_datfile()=, the only save, after the loop.
+ - ui.py:69 is =input(...)=, which raises EOFError on null stdin and blocks on a tty.
+ - informant.py:117-119 is =--all=, which marks everything read without prompting.
+ - config.py:19 sets the save file, =FILE_DEFAULT='/var/lib/informant.dat'=.
+ - file.py:45-53 catches PermissionError on that save, prints an error, and calls =sys.exit(255)=.
+- The rest of the package, from the same tree:
+ - usr/lib/tmpfiles.d/informant.conf: =f /var/lib/informant.dat 0664 root informant=
+ - usr/share/libalpm/hooks/00-informant.hook: =Target = *=, =When = PreTransaction=, =Exec = /usr/bin/informant check=, =AbortOnFail=
+ - informant.py:72-98: =check= exits with the unread count.
+- The installer's user setup (archsetup:1403-1405) adds =sys,adm,network,...,users= and does not add =informant=. The installer's own call, at archsetup:1121-1129 (commit cc0f7b5), is =informant read --all ... || true=, run as root, with the comment "a bare =informant read= is interactive and would hang an unattended run".
+- maint's =cmd.run= (maint/src/maint/cmd.py:22) uses =subprocess.run(..., capture_output=True)= with stdin inherited. A prompt there is either invisible or gets EOF.
+- On ratio, =pacman -Q informant= reports "package 'informant' was not found". =pacman -Si= shows informant 0.6.0-2 in [extra].
+:END:
+
+** DONE Panel offers REBOOT after a failed --complete gate :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design 'For the user' L67; Decision 5 L134-136; Phase 2 L158; Phase 3 L161; AC L169
+
+The spec says a failed gate 'stops there, names what failed, and does not reboot'. That holds only for the script. The spec never says what the panel shows after the kernel set has landed live and the gate has failed.
+
+Risk: On velox, the panel shows a red REBOOT key at the exact moment the gate has said a reboot can't import the pool. The message naming the failure is gone once the wall is cleared or the panel is reopened.
+
+Recommended change: Phase 1 (L155):
+- After the kernel set lands, if kernel-modules-check fails, --complete records the failure in the deferred-state file as gate: {ok: false, kernel: <ver>, failed: <item>, at: <ts>}.
+- It prints the failing item as its last stderr line.
+- The next --complete or kernel-modules-check run that passes clears that record.
+
+Phase 2 (L158) and Phase 3 (L161):
+- While gate.ok is false, the deferred probe renders a CRIT row: "kernel <ver> installed — gate failed: <item> — do not reboot". The usual "apply on reboot" wording is not shown.
+- The panel hides the REBOOT key, overriding reboot_required and offer_reboot.
+- Rewrite Phase 3's panel-action sentence to: runs --complete, which installs the kernel set, runs the gate, and arms and offers reboot only when the gate passes.
+
+Acceptance criteria: add "After a failed gate, the panel names the failing item persistently (it survives a wall dismiss and a panel reopen) and offers no REBOOT until a later gate run passes."
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Design, Implementation phases, Acceptance criteria, Testing. Confirmed, with two narrowings.
+- Only upgrade-guarded writes the record. A passing standalone kernel-modules-check therefore doesn't clear the failure, and the gate stays a stateless reader.
+- Clearing requires a real re-check of the recorded kernels against the recorded since. A --complete after a hand fix changes no kernel and would otherwise pass vacuously.
+
+APPLY stays visible while the failure is open, even when nothing else is deferred, so the re-check is reachable from the panel.
+:EVIDENCE:
+Spec:
+- L67: if the gate fails, the script "stops there, names what failed, and does not reboot".
+- L135: "A failed gate stops with the failure named and never reboots."
+- L155: --complete "refuses to arm or reboot on that exit".
+- L128/L158: the panel row reads "N deferred — apply on reboot".
+- L161: the Phase 3 panel action "installs the held-kernel set live, writes the persistent arm flag, and offers to reboot" (no gate, no failure branch).
+- L169: AC on a failed DKMS build.
+- No occurrence of reboot_required anywhere in the spec.
+
+Live system:
+- pacman -Qo /usr/lib/modules/6.18.25-1-lts-strix/modules.alias prints "No package owns".
+- /usr/share/libalpm/hooks/60-depmod.hook (PostTransaction, Remove/Upgrade on usr/lib/modules/*/) runs /usr/share/libalpm/scripts/depmod. That script does rm -f on modules.{alias,dep,...} and then rmdir --ignore-fail-on-non-empty "$f" for a kernel directory that no longer has modules.order.
+- 71-dkms-remove.hook is present.
+
+maint (in ~/.dotfiles/maint/src/maint):
+- probes/packages.py:131-146: required = not os.path.isdir(/usr/lib/modules/<uname -r>).
+- panel.py:506-510: reboot_key_visible() returns the reboot_required value.
+- gui.py:491-495: red REBOOT armed key when that is visible.
+- gui.py:1510-1511: offer_reboot only after an update or topgrade "ok".
+- remedies.py:314-320: the reboot remedy's metric_ids are ["reboot_required"].
+- doctor.py:166: return False, f"exit {proc.returncode}", stderr[-500:].
+- panel.py:308 and 389-396: the wall is a session-only log, and wall_clear() drops it.
+:END:
+
+** DONE UPDATE/TOPGRADE flag mapping and the stamp rule for partial runs are undefined :blocking:
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Scope tiers L55; Design L65, L67; Decision 4 L128; Phase 1 L155 (flags --no-topgrade, --no-aur; 'stamp only when nothing was deferred → exit 0 on a successful live part'); Phase 2 L158
+
+Both levers 'change their argv to the script', with no flags given. Today they are different remedies: UPDATE is a repo+AUR system update, and TOPGRADE is the ecosystem sweep that owns topgrade_age. Phase 1 defines --no-topgrade and --no-aur but names no consumer for them. The stamp predicate is only 'deferred set empty'. Nothing says what a run stamps or exits with when it skipped the sweep or the AUR, or when a sub-step failed (the pacman orphan stop in Risks L195, the AUR build, a topgrade step such as npm).
+
+Risk: There are two plausible implementations. In one, UPDATE silently gains the full ecosystem sweep and the two keys become duplicates. In the other, UPDATE runs --no-topgrade and the script stamps topgrade freshness for a run that never swept the ecosystems. Either way, a failed AUR or topgrade step with an empty deferred set reads 'current' under the state framing of Decision 1.
+
+Recommended change: Phase 2 (L158): state the lever mapping. UPDATE runs =upgrade-guarded --no-topgrade=; TOPGRADE runs =upgrade-guarded=. Add that on the driven path the script is the only thing that writes topgrade_run. That means removing doctor.py's rid=="topgrade" stamp, and having the script call the real topgrade binary (or tell the PATH wrapper not to stamp) so the inner sweep cannot stamp. Decision 4 (L128) and Phase 1 (L155): define the predicate for the everyday run. It stamps topgrade_run only when the sweep ran (no --no-topgrade or --no-aur), every step exited 0, and the deferred set is empty. If pacman fails, it skips the later steps, still writes the state file, exits non-zero, and does not stamp. --complete and the boot oneshot keep their own rule: stamp when their transaction succeeds and the deferred set is empty, with no sweep required. Add a Phase 1 test that a run with a non-empty deferred set leaves topgrade_run unwritten, through both the lever and the inner topgrade call.
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Decisions, Implementation phases, Testing. Accepted, with the completion-form rule tightened in two places.
+- A completion form stamps only when it started with something to complete. Otherwise a --complete that finds nothing would mint freshness after an UPDATE that never ran the sweep.
+- No form stamps while a gate failure is open, so TOPGRADE can't mark a do-not-reboot machine fresh.
+:EVIDENCE:
+- Spec L65 and L155: the script exits 0 "whatever was deferred" / "on a successful live part". L155: "stamp only when nothing was deferred". Flags --no-topgrade and --no-aur are defined with no consumer.
+- Spec L55, L67, L114, L121, L158: UPDATE and TOPGRADE both "change their argv to the script", with no per-lever flags.
+- Spec L128 (Decision 4): stamp topgrade_run only when the deferred set is empty. L136: the deferred count "carries a kernel most days".
+- Spec L32: names both existing writers. No phase removes or neutralizes either one.
+- ~/.dotfiles/maint/src/maint/remedies.py:294-312: update = ["yay","-Syu","--noconfirm"] with metric_ids updates_repo/updates_aur/cve_advisories; topgrade = ["topgrade","--disable","git_repos","-y"] with metric_ids topgrade_age.
+- ~/.dotfiles/maint/src/maint/doctor.py:169-170: =if rid == "topgrade": cache.put("topgrade_run", {"at": time.time()})= runs on any exit-0 step.
+- ~/.dotfiles/hyprland/.local/bin/topgrade:38-42: runs the real binary, then =maint stamp topgrade= when rc is 0. Its header says the maint lever resolves this wrapper too.
+- =which -a topgrade= lists ~/.local/bin/topgrade ahead of /usr/bin/topgrade.
+- Spec L142-143 (Decision 6): the boot/--complete path runs no sweep, and its stamp "reflects the guarded set specifically", so a predicate that requires the sweep would contradict it.
+:END:
+
+** DONE Blocked-set computation reads checkupdates' private db, not the db pacman -Syu uses
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L65 ('refreshes the sync db and reads the pending set (checkupdates)'; 'version-aware'); Phase 1 L155 and its tests
+
+The script 'refreshes the sync db and reads the pending set (checkupdates)'. It then computes the blocked set 'version-aware the way the guard is', by intersecting with the hook Targets, and passes it to =pacman -Syu --ignore==.
+
+Risk: A guard-style version check that runs before the real refresh leaves mesa, hyprland and vulkan-radeon out of --ignore. -Syu then pulls them in and the live guard aborts the whole transaction, so AC L167 fails in the state ratio is in right now. An update that appears in the second sync but not the first has the same effect. Treating exit 2 as a failure fails every day with nothing pending, and implementers will differ on how =mesa-*= becomes ignore entries.
+
+Recommended change: Design L65:
+- Replace "refreshes the sync db and reads the pending set (=checkupdates=); computes the *blocked set* = the guard's own trigger list (read from the installed hook's =Target= lines, ..., and version-aware the way the guard is — a same-version reinstall is not a swap)" with: "builds the *ignore list* = the installed hook's =Target= patterns passed verbatim to =--ignore= (IgnorePkg accepts shell globs, so =mesa-*= needs no expansion; a pattern with nothing pending is a no-op, and =-Syu= never reinstalls at the same version, so no version lookup is needed here) ∪ the kernel set".
+- After the yay step, add: "the *deferred set* is computed after the transaction from =pacman -Qu=, which reads the db =-Syu= just synced, filtered by the same patterns ∪ the kernel set."
+
+Phase 1 L155:
+- Change "blocked set (hook =Target= lines, version-aware)" to "ignore list (hook =Target= patterns verbatim) ∪ kernel set".
+- Change the test bullet "blocked-set computation against a fixture hook and version map" to "the =--ignore= list is exactly the fixture hook's Target patterns ∪ the kernel set; the deferred set is derived from a fake post-transaction =pacman -Qu=".
+
+Add one sentence: "=checkupdates= is used only by =--dry-run=; its exit 2 means an empty pending set, not a failure."
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Design, Implementation phases, Testing. Confirmed. The recommended text keeps pacman -Syu, which would sync a second time after F03's closure was computed against the first sync. That is the same race this finding names. Instead: one sudo pacman -Sy, the closure on that db, then sudo pacman -Su with no -y.
+:EVIDENCE:
+Commands below were run on ratio, 2026-10-05.
+- =checkupdates --help= (v1.13.1): the default db is =${TMPDIR:-/tmp}/checkup-db-${UID}=. In /usr/bin/checkupdates, :150 runs =pacman -Sy --dbpath "$CHECKUPDATES_DB"=, :155 drops =[...]= lines, and :181 is =exit 2= when nothing is pending.
+- =checkupdates -n= filtered to hook Targets and kernels lists: =hyprland 0.56.2-3 -> 0.56.2-4=, =mesa 1:26.2.3-1 -> 1:26.2.4-1=, =vulkan-radeon 1:26.2.3-1 -> 1:26.2.4-1=, plus linux, linux-headers, linux-lts and linux-lts-headers.
+- The guard-style lookup (=pacman -Q= vs =expac -S %v=, scripts/hypr-live-update-guard:78-82) reports mesa, hyprland and vulkan-radeon as installed = candidate, i.e. same version.
+- /var/lib/pacman/sync/extra.db is dated 29 Sep 08:21. =pacman -Qu= against the system db shows 0 updates; checkupdates shows 395.
+- tests/hypr-live-update-guard/test_hypr_live_update_guard.py:15-16,52: HYPR_GUARD_VERSIONS is a "pkg installed candidate" map. This is the "version map" the spec's Phase 1 test (L155) mirrors.
+- pacman.conf(5) on IgnorePkg: "Shell-style glob patterns are allowed." pacman(8) =--ignore= adds packages to the same ignore list.
+- /etc/pacman.d/hooks/hypr-live-update-guard.hook Targets: mesa, mesa-*, wayland, libdrm, libglvnd, hyprland, aquamarine, hyprutils, hyprgraphics, vulkan-radeon, vulkan-intel, vulkan-mesa-layers, nvidia-utils, lib32-nvidia-utils, xorg-xwayland. None use =!=.
+:END:
+
+** DONE Flag removal at the end of ExecStart never runs on timeout, kill or power loss
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L69; Decision 7 L145-150; Phase 3 L161; AC L171; Readiness L178
+
+L69 and L161: ExecStart runs the steps and then 'removes the flag unconditionally'. The flag directory /var/lib/archsetup is installer-owned (L178), and the unit runs as the user.
+
+Risk: Any timeout, kill, power loss or permissions slip leaves the flag in place, so every later boot re-runs the upgrade ahead of getty. That is the retry-every-boot outage Decision 7 exists to prevent, and it contradicts AC L171. The panel's RESTART key could also re-run the unit under a live session.
+
+Recommended change: In Design L69 and Phase 3 L161, replace "then removes the flag unconditionally" with: "the unit removes the flag before the transaction starts, privileged: ExecStartPre=+/usr/bin/rm -f /var/lib/archsetup/apply-upgrade-on-boot (the + runs it as root regardless of User=); ExecStart never touches the flag, so a timeout kill, SIGKILL, or power loss can't leave it armed."
+
+In Decision 7 L149, change "at the end of its attempt" to "at the start of its attempt". This is the same one-shot intent, and it now holds through a timeout.
+
+In L69 or L178, state that the flag is root-owned and that the panel's arm action creates it via sudo.
+
+In Phase 3 L161, add a forced-timeout case to the manual boot test: a short TimeoutStartSec plus a hung step, then check that the flag is gone, the next boot skips the unit, and Hyprland starts.
+
+If removal at the end is kept instead, move it to ExecStopPost=+/usr/bin/rm -f <flag> (that covers timeouts and failures, but not power loss).
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Design, Decisions, Implementation phases, Readiness dimensions, Testing. Partially confirmed; this follows the narrowing and my disarm-at-start call, with two reconciliations.
+- The flag now carries the package list (F09), so ExecStartPre moves it rather than deleting it.
+- maint never writes the flag (F07), so 'the panel's arm action creates it via sudo' becomes 'upgrade-guarded creates it via sudo'.
+:EVIDENCE:
+Spec L69: "Its ExecStart runs, as the user: informant read ..., then topgrade --only system ..., then maint stamp topgrade on success, then removes the flag unconditionally". L161: "ExecStart runs the script's --complete form as the user and removes the flag unconditionally". L149: "a persistent flag (/var/lib/archsetup/) that the boot unit removes unconditionally at the end of its attempt — success or failure disarms". L171: "A boot-upgrade failure (a failed step, a timeout, an aborted transaction) never blocks the session ... the flag is cleared".
+
+man systemd.service:
+- TimeoutStartSec: "the service will be considered failed and will be shut down again"
+- TimeoutStartFailureMode: the default "terminate" sends KillSignal (SIGTERM)
+- ExecStopPost: "commands specified with this setting are invoked when a service failed to start up correctly and is shut down again ... recommended ... for clean-up operations"
+
+Live repro: bash -c "sleep 30; rm -f flag", SIGTERM to the child and the shell, exit=143, and the flag still exists.
+
+Ownership: ls -ld /var/lib/archsetup gives drwxr-xr-x root (root-owned, 0755); the installer uses it at archsetup:257 (state_dir="/var/lib/archsetup/state").
+
+Guard backstop: scripts/hypr-live-update-guard:57-67 (hyprland_running uses pgrep -x Hyprland; it exits 0 only when Hyprland isn't running) blocks a live re-run.
+
+The panel paths exist as cited: maint/src/maint/probes/systemd.py:70-100 (failed_units WARN) and remedies.py:323-337 (unit_restart/unit_reset).
+:END:
+
+** DONE Deferred row severity, lever, band and bar impact are undefined
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Decision 4 L128; Decision 5 consequences L136; Phase 2 L158; ACs L166-174
+
+The spec only says the panel renders 'N deferred — apply on reboot' 'as its own row'. It gives no metric id, category or band, severity rule, lever or evidence shape. It also concedes that the count 'carries a kernel most days'.
+
+Risk: A WARN without a lever turns the waybar glyph gold almost every day, which just moves the 'permanently stale' symptom from the Summary (L28) somewhere else. A WARN with a lever adds a permanent ATTN. An OK never escalates the hold. The implementer has to invent all of this.
+
+Recommended change: Add two sentences to the end of Phase 2 (L158):
+
+"The row is metric upgrade_deferred, category updates (so it lands in the PACKAGES band). Its value is the deferred count and its evidence is the deferred rows (name, old, new). It grades OK when the set is empty and [WARN | OK — owner picks] when the set is non-empty. Phase 3's apply action is its lever, registered always:True so the action shows at any severity. A levered row never colours the waybar glyph. Age-based escalation stays the L197 follow-up."
+
+If WARN is chosen, also say that the row must not ship before its lever. Otherwise, in the window between Phase 2 and Phase 3, a lever-less WARN colours the glyph every day. Either land the row and the apply action in the same commit, or grade the row OK until Phase 3.
+
+Add one AC after L167: "With a non-empty deferred set and Hyprland live, the panel's deferred row offers the apply action and the waybar glyph colour is unchanged by the deferral."
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Acceptance criteria. Partially confirmed; this follows the narrowing. The open WARN-or-OK choice is resolved as OK for a plain deferral. The kernel is held most days, so WARN on mere presence would colour the glyph daily, which just moves the original stale-metric symptom onto another row; topgrade_age already carries the cadence nag.
+
+WARN is kept for states that need action (F10, F29, F36), and CRIT for a gate failure (F12). Because WARN is still reachable, the row ships in the same commit as its lever.
+:EVIDENCE:
+Spec L128: the deferred set is rendered "as its own state ('6 deferred — apply on reboot')". L158 (Phase 2): "A new probe reads the deferred-set state file and the panel renders 'N deferred — apply on reboot' as its own row". There is no severity, and no AC between L166 and L174 covers it.
+
+The lever does exist in the spec. L67 says "from the panel's action or upgrade-guarded --complete". L161 says "The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot." Escalation is explicitly deferred at L197: "a stale-kernel age in maint is a possible follow-up."
+
+Code:
+- ~/.dotfiles/maint/src/maint/remedies.py:535-545 (attach_levers): a lever is attached when the metric is WARN/CRIT, or when the remedy is marked "always".
+- remedies.py:294-312: the update and topgrade remedies are always:True. updates_repo is graded by grade_high(len, pending_warn=50) in probes/updates.py:53-61, so it reads OK below 50 with an always-lever.
+- ~/.dotfiles/maint/src/maint/indicator.py:41-57: diag is the metrics without levers, and only diag colours the glyph. Levered off-nominal metrics count as "actionable".
+- ~/.dotfiles/maint/src/maint/panel.py:36-47 (BAND_BY_CATEGORY maps updates to packages), 449-452 (band_lamp is the worst severity), 485-489 (attn_count counts every WARN/CRIT).
+
+Live data: ~/.local/state/maint/updates_repo.json (written 04:55 today, 395 rows) contains hyprland, linux, linux-headers, linux-lts, linux-lts-headers, mesa and vulkan-radeon. hyprland, mesa and vulkan-radeon match the /etc/pacman.d/hooks/hypr-live-update-guard.hook Target lines. vulkan-icd-loader and vulkan-mesa-implicit-layers are pending too, but they are not guard targets.
+:END:
+
+** DONE topgrade_age and the deferred row double-nag, and the TOPGRADE lever can't clear its metric
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Decision 4 L128; Decision 5 L136; Phase 2 L158; AC L170
+
+Decision 4 says a deferral renders 'rather than as stale freshness', but topgrade_age still counts days since the last stamp. With the kernel held on every everyday run, a stamp happens only after a dedicated session.
+
+Risk: Once the stamp is honest (see the stamp-writer finding), topgrade_age turns WARN 14 days after the last dedicated session, right next to the deferred row: two off-nominal rows for one fact. Pressing TOPGRADE runs a split run that can't clear it, and the wall reports the successful run as still warn.
+
+Recommended change: Add one sentence to Phase 2 (L158). Do not cap topgrade_age, because that conflicts with Decision 1.
+
+Suggested sentence: "topgrade_age grading is unchanged: per Decision 1 it ages from the last full-current stamp while anything is deferred. The deferred row is informational, carries the count, and hosts the apply-on-reboot / --complete action. While the deferred set is non-empty, the topgrade_age row offers that action instead of TOPGRADE, and the TOPGRADE wall note says 'N deferred — freshness clears on --complete' rather than a bare 're-probed topgrade_age: warn'."
+
+Add a matching acceptance criterion: "With only a kernel held, a successful TOPGRADE press leaves topgrade_age aging, and the wall names the deferred set and the --complete remedy."
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Acceptance criteria.
+:EVIDENCE:
+- Spec L107 (Decision 1): "Freshness stays stale while a guarded upgrade is genuinely un-applied."
+- Spec L77 (Alternative A): rejected because "a topgrade that was blocked from applying a real upgrade would read as 'fresh'."
+- Spec L128 (Decision 4): stamps only when the deferred set is empty; the deferred set is its own row; "Freshness keeps meaning 'current'."
+- Spec L136: "the panel's deferred count carries a kernel most days."
+- Spec L158 (Phase 2): no change to topgrade_age grading or to which lever hangs off it.
+- Spec L161 (Phase 3): the apply-on-reboot action is not bound to any metric.
+- Spec L197: a cadence nag for the kernel deferral is wanted.
+- maint/src/maint/probes/updates.py:114-131: grades days since topgrade_run against topgrade_warn_days.
+- configs/maintenance-thresholds.toml:60: topgrade_warn_days = 14.
+- maint/src/maint/remedies.py:303-312: the TOPGRADE lever has metric_ids and re_probe_ids ["topgrade_age"] and "always": True.
+- maint/src/maint/remedies.py:537-546: attach_levers shows "always" levers at any severity.
+- maint/src/maint/doctor.py:169-170: stamps topgrade_run on any successful rid=="topgrade" run; Phase 2 does not remove this, so the "can't clear" scenario only exists after the separate stamp-writer fix.
+- hyprland/.local/bin/topgrade: the wrapper also stamps on rc 0.
+- maint/src/maint/doctor.py:342-350: the re-probe note goes to the wall.
+:END:
+
+** DONE Kernel-set rule sweeps in firmware and api-headers, and ratio has three kernels
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Phase 1 L155 ('every linux* kernel package and its -headers'); Design L65; Decision 5 L134-135; Risks L196
+
+The kernel set is 'every linux* kernel package and its -headers', 'moved as a set, never one without the other'. Ratio is described as 'btrfs root, two kernels'.
+
+Risk: A literal glob holds linux-api-headers and all firmware, GPU firmware included, until the dedicated session, and inflates the 'N deferred' count. A substring match would hold archlinux-keyring indefinitely and start breaking signature checks. 'Never one without the other' can't hold for strix, and --complete can't pacman-install a locally built kernel, so the implementer has to invent the policy.
+
+Recommended change: L155: replace "(every =linux*= kernel package and its =-headers=)" with "(the packages owning =/usr/lib/modules/*/vmlinuz=, via =pacman -Qqo=, plus =<pkgbase>-headers= where that package is installed; =linux-firmware*=, =linux-api-headers= and other =linux*= names are not kernels)". Add: "the apply step moves only the pending members of the kernel set; a foreign kernel never appears in checkupdates and is left untouched."
+
+In the same paragraph, narrow the gate to match L135: "for each kernel version installed by this run" instead of "for each kernel under =/usr/lib/modules=". As written, the gate fails permanently on any kernel with no headers and no DKMS build.
+
+L134 and L196: change "two kernels" to "three kernels, one a locally built foreign package with no =-headers=".
+
+Phase 1 tests: add a ratio-shaped fixture with three vmlinuz owners, one of them foreign and headerless, linux-firmware* and linux-api-headers installed, and zfs built for two of the three kernels. Assert that firmware is never held and that the gate ignores kernels this run did not touch.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Design, Decisions, Implementation phases, Testing.
+:EVIDENCE:
+Spec text:
+- L155: "The kernel set is derived from what is installed (every =linux*= kernel package and its =-headers=)"
+- L155 (gate): "for each kernel under =/usr/lib/modules=, =dkms status= reports every registered module =installed= for it"
+- L65 and L135: "every installed kernel with its =-headers=" and "moved as a set, never one without the other"
+- L135 (gate): "for the new kernel version"
+- L134 and L196: "Ratio (btrfs root, two kernels ..."
+
+Live checks on ratio (uname -n = ratio):
+- pacman -Qq | grep -c '^linux' returns 24: linux, linux-api-headers, linux-firmware plus 17 linux-firmware-* split packages (amdgpu among them), linux-headers, linux-lts, linux-lts-headers, linux-lts-strix.
+- pacman -Qoq /usr/lib/modules/*/vmlinuz returns linux-lts-strix, linux-lts, linux. Each /usr/lib/modules/*/pkgbase file matches its owner.
+- pacman -Qm lists linux-lts-strix 6.18.25-1. pacman -Qi shows Packager "Unknown Packager". pacman -Si linux-lts-strix gives "error: package 'linux-lts-strix' was not found".
+- pacman -Qq | grep -- '-headers$' returns linux-api-headers, linux-headers, linux-lts-headers. There is no strix headers package.
+- dkms status shows "zfs/2.4.4, 6.18.54-1-lts, x86_64: installed" and "zfs/2.4.4, 7.2.7-arch1-1, x86_64: installed". Nothing is built for 6.18.25-1-lts-strix.
+- findmnt / shows btrfs, which confirms the btrfs-root part of L134.
+- checkupdates lists linux, linux-headers, linux-lts and linux-lts-headers as pending. No firmware and no linux-api-headers are pending today.
+:END:
+
+** DONE The gate's snapshot and initramfs checks pass vacuously
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Phase 1 L155 (kernel-modules-check); Decision 5 L135; AC L169; Risks L196
+
+The gate passes if each kernel's 'initramfs is newer than its vmlinuz' (each kernel under /usr/lib/modules) and, 'on a ZFS root, a pre-pacman_ snapshot of the root dataset exists'. L196 says the 05-zfs-snapshot hook snapshots before every transaction.
+
+Risk: Neither check can catch the failure it exists for. One is a snapshot that wasn't taken for this kernel change (skipped, or failed on a full pool, one of the DKMS triggers the spec lists); rolling back to an older retained snapshot then undoes more than the kernel transaction. The other is a stale initramfs beside a new kernel, which lets the gate wave through an unbootable reboot on velox. The 'fails on snapshot listings' test covers a state that never occurs.
+
+Recommended change: Phase 1, L155. Replace the gate sentence with:
+
+"The script records T0 before the kernel transaction. For each kernel, with pkgbase read from /usr/lib/modules/<ver>/pkgbase: dkms status reports every registered module =installed= for <ver>, and /boot/initramfs-<pkgbase>.img is newer than /boot/vmlinuz-<pkgbase>. For a kernel this run changed, the image must also be newer than T0. Never compare against /usr/lib/modules/<ver>/vmlinuz, whose mtime is the package build date. On a ZFS root, a =pre-pacman_= snapshot of the root dataset (findmnt -no SOURCE /) has a creation time of at least T0 minus 60s (zfs-pre-snapshot's MIN_INTERVAL). An older retained snapshot does not satisfy this."
+
+Phase 1 tests. Replace the snapshot and timestamp fixtures with:
+- snapshot listings where only pre-T0 snapshots exist (gate fails) and one is newer than T0 (gate passes)
+- an initramfs newer than the module-tree vmlinuz but older than T0 (gate fails)
+
+L196. Change "snapshots it before every transaction" to "snapshots it before each transaction (skipped within 60s of the previous snapshot; a failed snapshot only warns and does not abort)".
+
+Optional: on a ZFS root, also check with lsinitcpio that the image contains zfs.ko.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Decisions, Implementation phases, Risks, Testing. Confirmed, with the finding's corrections.
+- The gate finds the root dataset itself with findmnt. zfs-pre-snapshot hardcodes $POOL/ROOT/default, so a mismatch fails closed and the failure names the dataset.
+- The freshness and snapshot requirements apply only when the check list is non-empty, so a --complete that moves no kernel isn't failed for want of a snapshot.
+
+The optional lsinitcpio check is made required on a ZFS root. mkinitcpio writes an image without zfs.ko and only warns, and that image is exactly the velox failure.
+:EVIDENCE:
+Spec:
+- L135 (Decision 5): "the initramfs is newer than the kernel image, and on a ZFS root a pre-pacman snapshot exists"
+- L155 (Phase 1): "for each kernel under =/usr/lib/modules= ... its initramfs is newer than its =vmlinuz=; on a ZFS root, a =pre-pacman_= snapshot of the root dataset exists"
+- Phase 1 tests: "fails on ... snapshot listings"
+- L169: "the pre-pacman snapshot it required"
+- L196: "the =05-zfs-snapshot= hook snapshots it before every transaction"
+
+scripts/zfs-pre-snapshot:
+- :10-11 DATASET="${ZFS_PRE_DATASET:-$POOL/ROOT/default}", hardcoded, no findmnt
+- :13, :19-25 return early when the lock file is under MIN_INTERVAL=60s old
+- :14, :35-40 KEEP=10 prune
+- :30, :41-43 on failure, =echo "Warning: Failed to create snapshot"= with exit status 0
+
+archsetup:
+- :2421-2433 the 05-zfs-snapshot.hook heredoc has no AbortOnFail
+- :881-884 is_zfs_root uses =findmnt -n -o FSTYPE /= only
+
+Live system on ratio:
+- =pacman -Qi linux-lts= shows Build Date Fri 25 Sep 2026 11:39:08 and Install Date Tue 29 Sep 2026 08:15:55
+- /usr/lib/modules/6.18.54-1-lts/vmlinuz has mtime 2026-09-25 11:39:08, the build date
+- /boot/vmlinuz-linux-lts is 2026-09-29 08:17:18.09 and /boot/initramfs-linux-lts.img is 08:17:29
+- /usr/lib/modules/7.2.7-arch1-1/vmlinuz is 2026-09-21 13:51:14, matching linux's Build Date
+- /usr/share/libalpm/hooks/90-mkinitcpio-install.hook triggers on usr/lib/firmware/*, usr/lib/systemd/systemd, usr/bin/cryptsetup, usr/lib/modprobe.d/ and other paths
+- /usr/share/libalpm/scripts/mkinitcpio:148-151 runs =install -Dm644 -- "${line}" "${kernel}"= into /boot/vmlinuz-${pkgbase} before the preset build
+- /usr/lib/modules/*/pkgbase gives linux-lts-strix, linux-lts and linux
+:END:
+
+** DONE Ratio's guard hook still has the pre-migration name, and the missing-hook case is unspecified
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Problem L34; Design L65; Decision 3 L121; Phase 1 L155; Phase 2 L158; Phase 4 L164
+
+The spec names /etc/pacman.d/hooks/10-hypr-live-update-guard.hook as the installed hook the script reads, and relies on it as the backstop for a bare pacman -Syu. Phase 4's rollout installs only the new unit on existing machines.
+
+Risk: A script that hardcodes the 10- path finds no hook on ratio and computes an empty blocked set. The old-named guard still fires, so the live transaction aborts, and the Phase 2 equality test misreads or errors. On a machine with no hook at all, reading 'no patterns' would live-swap mesa with no backstop. Separately, on ratio today, a bare =pacman -Syu= under Hyprland with a kernel and a guarded lib both pending deletes /boot/vmlinuz and the initramfs for linux and linux-lts, then aborts, leaving those boot entries without images.
+
+Recommended change: Phase 1 (L155): after "blocked set (hook Target lines, version-aware)", add: "The hook is read from /etc/pacman.d/hooks/10-hypr-live-update-guard.hook. If that file is absent, or no Target lines parse from it, the live run refuses: it exits non-zero and names the path, and never computes an empty blocked set. This is deliberately unlike maint's guard.trips, which treats missing patterns as 'can't trip'." Add a matching test to the Phase 1 list: "a missing or Target-less hook fails closed and runs no pacman transaction."
+
+Phase 4 (L164): add: "The one-time install on existing machines also migrates the guard hook to 10-hypr-live-update-guard.hook and removes the unprefixed hypr-live-update-guard.hook (ratio still has the legacy name; velox's was hand-placed under it). The unprefixed name sorts after 60-mkinitcpio-remove, so a blocked bare pacman -Syu can delete the initramfs before the guard aborts. Confirm with ls /etc/pacman.d/hooks/ on both machines."
+
+Optional: add a post-rebuild-check assertion that the 10- name exists and the unprefixed name does not.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Design, Implementation phases, Testing. Confirmed. Under the fail-closed rule the script refuses on any machine that still carries the legacy hook name. Ratio does today (only hypr-live-update-guard.hook is in /etc/pacman.d/hooks), and velox's was hand-placed under it. So the migration can't wait for Phase 4; it moves to checklist step (a), before that machine's levers route through the script. The optional post-rebuild-check assertion is adopted.
+:EVIDENCE:
+- Live ratio: ls /etc/pacman.d/hooks/ shows only 99-grub-sync-efi.hook and hypr-live-update-guard.hook (543 bytes, 4 Jul, same content as the installer heredoc). No 10-hypr-live-update-guard.hook.
+- archsetup:2620-2625 migrates the name ("rm -f /etc/pacman.d/hooks/hypr-live-update-guard.hook", then writes 10-hypr-live-update-guard.hook). The comment at 2622-2624 says the guard must run before 60-mkinitcpio-remove. Commit be2277d (2026-07-18, "fix: order pacman safety hooks") added this.
+- archsetup:41-75: the arguments are only --config-file, --fresh, --status and similar flags. There is no way to run a single step, so updating an existing machine means placing files by hand.
+- /usr/share/libalpm/hooks/60-mkinitcpio-remove.hook is a PreTransaction hook on Path Remove usr/lib/modules/*/vmlinuz. man alpm-hooks says hooks run "in alphabetical order of their file name", so "hypr-live-update-guard" sorts after "60-mkinitcpio-remove". AbortOnFail aborts the transaction. pacman -Q shows linux 7.2.7 and linux-lts 6.18.54, both with headers.
+- tests/installer-steps/test_pacman_hook_order.py:54-76 checks the ordering and the cleanup, but only against the installer source, not a live machine. scripts/testing/tests/test_desktop.py:59 asserts the 10- path, but only in the VM harness. scripts/post-rebuild-check has no hook assertion (grep returned nothing).
+- ~/.dotfiles/maint/src/maint/guard.py:26-30 (trips) is the fail-open precedent: "Tolerant of a missing key — no patterns means the guard can't trip".
+- todo.org:3391-3396: the hook was placed on velox by hand as /etc/pacman.d/hooks/hypr-live-update-guard.hook on 2026-06-28.
+- Spec: L34 names the 10- path. L65 says the blocked set is "read from the installed hook's Target lines" and that the hook stays "the backstop for a bare pacman -Syu". L155 says "hook Target lines, version-aware" and lists no missing-hook test. L164 says "existing machines need the one-time install", with no hook migration. Grepping the spec for absent or missing finds them only for topgrade_run and the arm flag, never for the hook.
+:END:
+
+** DONE maint's guard_patterns is not a mirror of the hook, and the pin test is in the wrong repo
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Decision 3 L120-121; Phase 2 L158; Reuse L183
+
+The spec calls maint's =[updates] guard_patterns= its own copy of the hook's trigger list. It adds a Phase 2 test under the maint harness asserting that the TOML equals the installed hook's Target list.
+
+Risk: The equality test fails on its first run, and the spec doesn't say which side gets reconciled. A dotfiles unit test that reads /etc and ~/.config depends on the host and on how stale its installed copies are. Rewriting the TOML drops hyprlang and hyprcursor, which the --noconfirm closure finding shows belong to the soname family, and the badge under-predicts the deferral.
+
+Recommended change: Edits by location:
+
+- L120: replace "its own copy of the trigger list" with "its own, divergent pattern list".
+- L158: replace the test sentence with: "Rewrite archsetup configs/maintenance-thresholds.toml [updates] guard_patterns to exactly the hook heredoc's 15 Target patterns (the hook is the owner and stays unchanged). This drops hyprland-*, hyprlang, hyprcursor, *wayland*, wlroots*, lib32-mesa*, lib32-vulkan-radeon and lib32-vulkan-intel, and adds wayland, libdrm, libglvnd, vulkan-mesa-layers, nvidia-utils, lib32-nvidia-utils and xorg-xwayland. Update the TOML comment that still describes press-again arming. The pin test lives in archsetup tests/installer-steps/ and compares the sorted seed-TOML guard_patterns with the Target lines extracted from the hook heredoc in the archsetup script (the test_pacman_hook_order.py idiom). It reads neither /etc nor ~/.config."
+- L183: align with the L158 change.
+- L188 and Phase 4: note that existing machines need install_maintenance_thresholds re-run to pick up the new TOML.
+- L187 and L200: fix the phase numbering so each phase's tests are in one consistent place.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Decisions, Implementation phases, Readiness dimensions, Testing. Confirmed. Both files the change touches, configs/maintenance-thresholds.toml and the hook heredoc, belong to archsetup. The rewrite and the pin test therefore land in the Phase 1 archsetup commit, not in dotfiles Phase 2. The rewritten TOML has to reach existing machines, so the re-copy at checklist step (a) is mandatory, overriding F30's 'only if Phase 2 adds keys'.
+:EVIDENCE:
+- Spec L120: "maint already carries its own copy of the trigger list as [updates] guard_patterns".
+- Spec L121: "the TOML patterns stay as the panel's display-side mirror ... and gain a test asserting they match the hook".
+- Spec L158: "A test asserts the TOML guard_patterns equal the installed hook's Target list. Tests under the maint fake harness."
+- Spec L183 repeats the claim. L187 ("installer-step pytest for Phase 2") and L200 ("Phase-2 install under tests/installer-steps/") contradict L158.
+- configs/maintenance-thresholds.toml:65-72: mesa, mesa-*, lib32-mesa*, hyprland, hyprland-*, aquamarine, hyprutils, hyprlang, hyprcursor, hyprgraphics, *wayland*, wlroots*, vulkan-radeon, lib32-vulkan-radeon, vulkan-intel, lib32-vulkan-intel. The file is byte-identical to ~/.config/archsetup/maintenance-thresholds.toml (diff clean).
+- archsetup:2629-2643, the hook heredoc's 15 Target lines: mesa, mesa-*, wayland, libdrm, libglvnd, hyprland, aquamarine, hyprutils, hyprgraphics, vulkan-radeon, vulkan-intel, vulkan-mesa-layers, nvidia-utils, lib32-nvidia-utils, xorg-xwayland.
+- The installed hook's Target list matches the heredoc, but it sits at /etc/pacman.d/hooks/hypr-live-update-guard.hook. /etc/pacman.d/hooks/10-hypr-live-update-guard.hook does not exist on ratio (cat: No such file or directory).
+- archsetup:1640-1650: install_maintenance_thresholds copies the seed TOML to ~/.config/archsetup.
+- dotfiles maint/src/maint/thresholds.py:3-4: the TOML is "owned and installed by archsetup".
+- thresholds.py:18-19: env overrides "so tests and fixtures never touch real files".
+- thresholds.py:31-33: shipped_path.
+- maint/src/maint/guard.py:29: maint reads guard_patterns via fnmatchcase.
+- tests/maint/test_remedies_doctor.py:31-60 uses an inline SHIPPED_TOML. No dotfiles maint test references the archsetup checkout. Every test sets MAINT_THRESHOLDS to a temp file.
+- tests/installer-steps/test_pacman_hook_order.py already extracts hook data from the archsetup script source.
+:END:
+
+** DONE maint's guard refusal and press-again arm UX are left in place
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Phase 2 L158; Design 'For the user' L67; AC L167; Decision 3 L121
+
+Phase 2 removes only 'the press-again-to-force sentinel wrap'. Design and the AC say UPDATE/TOPGRADE run the script and succeed with a guarded library pending. The 'will defer' badge is unspecified.
+
+Risk: If only the wrap is removed, the panel still refuses on exactly the guarded days the script exists for, and AC L167 fails. The second press still says the run swaps live, which is no longer true. The badge under-predicts the deferral because it omits the kernel set.
+
+Recommended change: Replace the Phase 2 sentence "the press-again-to-force sentinel wrap goes away (the driven path never trips the guard)" with: "Drop guard: live_update from the update and topgrade remedies. iter_fix no longer refuses them. The --force sentinel wrap, the GUI's _update_force override, and the CLI --force meaning for these two go with it. Rewrite the tests that pin the tag, the refusal, and the wrap. guard.trips stays only as the arm-line annotation that Decision 3 calls the display-side mirror: on a tripped read, arm_line, _rearm_after_guard's text, and the doctor review suffix read 'UPDATE armed — will defer <guard matches> (kernels are always held) — press again to run $ upgrade-guarded', and nothing says live apply or REBOOT required." Don't have maint compute the pending kernel set itself. The post-run "N deferred" row, built from the script's state file, gives the exact count.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Testing. Partially confirmed; this follows the narrowing that maint does no kernel computation of its own. The arm line also names the DKMS set (F04), and it says 'at least', because the closure can hold more than the pattern match predicts.
+:EVIDENCE:
+Spec L158 (Phase 2): "UPDATE and TOPGRADE levers change their argv to the script; the press-again-to-force sentinel wrap goes away (the driven path never trips the guard)." Spec L121 (Decision 3): "the TOML patterns stay as the panel's display-side mirror (the badge that says a run will defer)". Spec L167 (AC): "UPDATE applies everything else, exits 0 ... the guard hook does not fire."
+
+~/.dotfiles/maint/src/maint/remedies.py:298 and :309: both update and topgrade carry "guard": "live_update".
+
+~/.dotfiles/maint/src/maint/doctor.py iter_fix: =if not dry_run and r.get("guard") == "live_update" and not force:= runs guard.trips(pending, th) and yields a "guard" event ("press again to apply live (REBOOT required after), or apply from a TTY"), then returns before any step. The wrap is a separate later line: =wrapped = bool(force and r.get("guard") == "live_update" and not dry_run)= -> _all_steps adds guard_sentinel_set/clear. The review suffix (doctor.py ~L415-421) reads "[live-update guard: ... — arms in the GUI; --force on the CLI]".
+
+viewmodel.py:99-107 arm_line(guard=...) gives "press again to apply live (REBOOT required after), or apply from a TTY".
+
+gui.py:1396-1405 _guard_arm_line sets self._update_force = bool(info["tripped"]). gui.py:1430-1437 fires with force=self._update_force for guarded rids, so the GUI bypasses the refusal on a warm tripped read. gui.py:1518-1535 _rearm_after_guard repeats the live-apply/REBOOT wording.
+
+panel.py:224-237: guarded()/update_guard() key on the same tag.
+
+cli.py:265-267: --force help "run a guarded remedy despite the live-update guard (the CLI's press-again)".
+
+Tests pinning the current behavior: tests/maint/test_remedies_doctor.py:437-438 (assert guard == "live_update" on both), :843-1003 (refusal, force, sentinel order); tests/maint/test_panel_levers.py:404 (asserts "press again to apply live (REBOOT required after)").
+
+configs/maintenance-thresholds.toml:65-72: guard_patterns has no linux*/kernel entries, so a guard-only badge omits the held kernel set the script always defers.
+:END:
+
+** DONE Deferred-set vocabulary, storage and arm-flag contents are undefined across repos
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L65, L67; Scope tiers L57; Decision 4 L128; Decision 6 L142; Phase 1 L155; Phase 2 L158; Phase 3 L161; Readiness Data model L178
+
+'blocked set' includes the kernel set in Design (=--ignore=<blocked set>=) but excludes it in Phase 1 ('blocked set ∪ held-kernel set'). The spec also uses 'held set', 'held-kernel set', 'guarded set' and 'GPU/compositor set'. Storage is called 'a state file' (L65, L155, L158) in some places and 'its own cache key' (L128) in another. Phase 1 (archsetup) writes the record and Phase 2 (dotfiles) reads it, but neither names the path, schema or writer. Data model (L178) lists only the arm flag and topgrade_run. The arm flag's contents (which packages the boot run applies) are unspecified, and no phase says that --complete or the boot run rewrites the record. The badge reads 'apply on reboot' even though the kernel portion lands live.
+
+Risk: Phases 1 and 2 land in different repos and sessions, so the implementers will invent incompatible paths and formats. A mismatch shows no deferred row and fails silently, because the 'state file round-trips' test only checks the script against itself. A sudo-invoked run writes to /root/.local/state/maint, which the panel never reads. The boot run has no defined source for which set to apply, and a record that is never cleared leaves a permanent 'N deferred' badge.
+
+Recommended change: 1. Definitions. Add one sentence to Design after the L65 set computation: "GPU/compositor set = pending hook Targets whose on-disk version changes, computed only when a compositor is live; kernel set = every installed linux* kernel with its -headers, held on every everyday run whether or not a compositor is live; deferred set = their union, which is exactly the --ignore list." Then:
+- In L65, change "--ignore=<blocked set>" to the deferred set.
+- In L155, rewrite "blocked set (...) ∪ held-kernel set when a compositor is live" as "GPU/compositor set (hook Target lines, version-aware, compositor live only) ∪ kernel set (always)".
+- Replace "held set" (L57), "held-kernel set" (L55, L161) and "guarded set" (L65, L143) with the defined terms.
+
+2. Name the record. In Decision 4 (L128), and as a new line in Data model (L178), add: "maint cache key upgrade_deferred: ~/.local/state/maint/upgrade_deferred.json (MAINT_STATE_DIR honoured), in cache.put's {written_at, data} envelope, data = {packages: [{name, old, new, kind: gpu|kernel}]}. upgrade-guarded is its only writer and runs as the user. It rewrites the record at the end of every mode (everyday, --complete, boot), with an empty list when nothing is deferred. maint's new probe is its only reader."
+
+3. Boot source. In Decision 6 (L142) and Phase 3 (L161), add: "the boot form applies the record's kind=gpu entries and never the kernel ones; the arm flag stays a presence-only file."
+
+4. Tests. Make Phase 1's "state file round-trips" test read the record in the envelope maint's cache.get expects, and point Phase 2's probe test at the same fixture JSON.
+
+Do not rename the badge, add SUDO_USER handling, or add gate/armed/boot fields to the schema.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Design, Decisions, Implementation phases, Readiness dimensions, Testing. Partially confirmed. This adopts the vocabulary and the named record, with two reconciliations.
+- My boot-source call puts the exact name=version list in the root-owned arm flag. The flag isn't presence-only: it is the boot form's source, and the record is the panel's.
+- Accepted findings F10, F12, F29 and F39 each need a durable field (interrupted, gate, AUR failure, news), so the schema carries result, failed_step, detail, gate and news.
+
+Arm state is still not stored, because the probe reads the flag.
+:EVIDENCE:
+Spec:
+- L65: "computes the *blocked set* = the guard's own trigger list ... plus the *kernel set* ...; runs pacman -Syu --noconfirm --ignore=<blocked set> ... It writes the deferred set to a state file the panel reads".
+- L155: "blocked set (hook Target lines, version-aware) ∪ held-kernel set when a compositor is live → ... deferred set written to a state file". Its tests: "the --ignore list is exactly blocked ∪ kernel set; the state file round-trips".
+- L128: "writes the deferred set to its own cache key"; L129: "one more cache key and one more probe in maint".
+- L158: "A new probe reads the deferred-set state file".
+- L178: Data model lists only the arm flag and the topgrade_run key.
+- L69: the arm flag is "a file on a non-tmpfs path".
+- L142: the boot run is "scoped to the deferred GPU/compositor set ... no kernel".
+- L155: --complete is "apply the kernel set live, run the gate, then arm ... or, with no compositor live, apply it directly".
+- L171, L181: the panel clears or still shows pending work after boot.
+- L135, L168: the kernel is held on every everyday run, including a TTY run.
+- L57: "held set". L143: "guarded set".
+
+Code:
+- ~/.dotfiles/maint/src/maint/cache.py:15-42: _dir() is MAINT_STATE_DIR or ~/.local/state/maint; put() writes {"written_at", "data"} to <name>.json atomically; get() returns (data, age).
+- cli.py:294: p_stamp choices=["topgrade"]. cli.py:228-231: cmd_stamp does cache.put(f"{what}_run", {"at": ...}).
+- remedies.py:294-311: the update and topgrade remedies are kind "user", not root.
+
+Adjacent defect, not part of this finding (flag separately): which -a topgrade resolves ~/.local/bin/topgrade first. That wrapper stamps topgrade_run on any rc 0 (wrapper L40-42), and doctor.py:169-170 also stamps on any exit-0 TOPGRADE lever. upgrade-guarded's own topgrade --disable system,... call, and a lever whose argv is the script (which exits 0 while deferring, per L65), would therefore stamp freshness on a deferring run, contradicting Decision 4 (L128).
+
+Also noticed: on ratio the installed hook is /etc/pacman.d/hooks/hypr-live-update-guard.hook, not 10-hypr-...; the installer renames it at archsetup:2621-2625. That affects "read the installed hook", not this finding.
+:END:
+
+** DONE Boot run is invisible on ratio's screen, and its outcome is never surfaced
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L67 ('runs in the console'); Decision 7 L149; AC L171; Readiness Errors L179, Observability L181, Performance L182, Config surface L185
+
+The spec assumes the boot run's output appears on the console and that failures are 'named', but it sets no StandardOutput or TTYPath. The flag is removed whatever the outcome, and success or failure lives only in the unit status and the journal. Nothing records the outcome where the panel reads it. AC L171 only requires that 'the panel still shows the pending work', and after a failed run the panel shows the same 'N deferred' as before arming.
+
+Risk: On ratio, a multi-minute armed boot shows a blank tty1, which invites a hard power-cycle mid-transaction (see the timeout finding). Any failure lives only in the journal. The panel gives no sign that the one-shot arm was spent, so the user either re-arms blind or never re-arms.
+
+Recommended change: This replaces the finding's larger proposal: no new outcome field and no "last boot apply failed" panel line.
+
+1. Phase 3 (L161) and Observability (L181). Name the unit's I/O: StandardInput=null, StandardOutput=tty, TTYPath=/dev/tty1. Not journal+console, because on ratio /dev/console is ttyS0. The script mirrors its output to the journal (for example through systemd-cat). Its first line is a banner on tty1: "applying N deferred GPU/compositor upgrades — do not power off".
+
+2. L69 and L161. Move the unconditional flag removal out of ExecStart into ExecStopPost=/usr/bin/rm -f <flag>. ExecStopPost runs after success, failure, and a TimeoutStartSec kill, and ExecStart's exit then carries the upgrade result. A failed or timed-out run leaves archsetup-boot-upgrade.service failed, and maint's existing failed_units row names it with its journalctl hint in the session that follows.
+
+3. AC L171. Append: "and a failed or timed-out run appears in maint's failed-units row as archsetup-boot-upgrade.service".
+
+4. The Phase 3 manual boot test. Add a check that the banner and progress are visible on ratio's monitor during the armed boot.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Design, Implementation phases, Acceptance criteria, Readiness dimensions, Testing. Partially confirmed; this follows the narrowing: no new outcome field, and maint's failed_units row is the surface. The proposed ExecStopPost disarm gives way to my ExecStartPre call. That keeps the property the finding wanted: ExecStart no longer ends in an rm, so its exit status is the unit's result, and failures and timeouts leave the unit failed.
+:EVIDENCE:
+Visibility half, all confirmed on ratio:
+- uname -n is ratio.
+- systemctl show reports DefaultStandardOutput=journal.
+- /proc/consoles shows "ttyS0 -W- (EC p a)" (C = preferred, the /dev/console target) and "tty0 -WU (E p )".
+- /etc/default/grub:6 has GRUB_CMDLINE_LINUX="console=tty0 console=ttyS0,115200", with quiet splash loglevel=2 at :5. /proc/cmdline matches.
+- pacman -Q plymouth: package not found.
+- /etc/systemd/system/getty@tty1.service.d/ does not exist (no autologin on ratio). systemctl cat getty@tty1 shows TTYReset=yes and TTYVTDisallocate=yes (getty@.service file lines 42 and 44).
+- ~/.dotfiles/hyprland/.profile.d/99-hyprland-autostart.sh:15 runs clear before start-hyprland.
+- The ttyS0 console is ratio-local. grep for ttyS0 finds nothing in the archsetup repo, so the hazard is ratio-specific and velox is unverified (offline).
+
+Surfacing half, refuted in part:
+- ~/.dotfiles/maint/src/maint/probes/systemd.py:70-101 failed_units emits one evidence entry per failed unit: unit, since, exit, and hint "journalctl -u {name} -b".
+- gui.py:1068-1074 renders it as a "failed units" section with "RESTART / RESET per unit".
+- status.py:50 includes it in maint status.
+
+Spec text:
+- L69: ExecStart "... then maint stamp topgrade on success, then removes the flag unconditionally".
+- L161: "ExecStart runs the script's --complete form as the user and removes the flag unconditionally".
+- L149: "removes unconditionally at the end of its attempt".
+- L171 AC: timeout leaves "the flag is cleared, and the panel still shows the pending work".
+- L181: "the boot run's output is on the console; its systemd unit status and journal record success/failure".
+:END:
+
+** DONE maint is not on the boot unit's PATH, so the boot stamp can silently not happen
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L69; Phase 1 L155; Phase 3 L161; AC L170; Readiness L189
+
+The boot oneshot runs the script as the user, and the script stamps with =maint stamp topgrade=, with no path or environment specified. 'maint stamp' is listed as verified present.
+
+Risk: The boot run applies the deferred set but writes no stamp. AC L170 (fresh topgrade_age after the armed reboot) then fails with no error anywhere.
+
+Recommended change: In Phase 3 (L161), after "=ExecStart= runs the script's =--complete= form as the user", add one sentence: "The unit sets =User== to the installing user (templated by the installer, as the BRIO udev rule is), so =HOME= points at that user's maint state. The script calls maint by absolute path (=$HOME/.local/bin/maint=), not through PATH, because a system unit's PATH is =/usr/local/bin:/usr/bin= and doesn't include =~/.local/bin=. If maint is missing or the stamp fails, the script logs a named failure to the journal instead of skipping silently (no =command -v ... || true=)." In L189, qualify the claim: =maint stamp= is present on the user's login PATH only (=~/.local/bin=, from dotfiles), not on the system PATH. Leave topgrade resolution out of this finding. The boot run doesn't invoke topgrade (Decision L142).
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Design, Implementation phases, Readiness dimensions, Testing. Partially confirmed; this follows the narrowing, so the topgrade-resolution clause is dropped (F01 covers the live path, and the boot form never runs topgrade). It keeps the recommended absolute-path call. maint stamp topgrade exists today (maint cli.py), and the boot unit already depends on the user's HOME for the record.
+
+'Logs a named failure' gets an exit code and a failed_step, so a failed stamp also reaches failed_units at boot.
+:EVIDENCE:
+- Spec L69: the boot unit's "ExecStart runs, as the user: ... then =maint stamp topgrade= on success". L155: --complete stamps when the deferred set is empty. L161: "=ExecStart= runs the script's =--complete= form as the user". L189: "=maint stamp= — all verified present on velox". The spec has no PATH or Environment wording anywhere (grep -n "PATH\|Environment\|absolute" matches only L32's unrelated "PATH wrapper").
+- =which -a maint= finds only ~/.local/bin/maint, a symlink to ../../.dotfiles/hyprland/.local/bin/maint (a python shim that loads maint/src from the dotfiles repo).
+- =systemd-path search-binaries-default= gives /usr/local/bin:/usr/bin.
+- ~/.dotfiles/common/.zprofile:4-12 says that without =export PATH="$HOME/.local/bin:$PATH"=, "maint, the topgrade wrapper, and every user script are unreachable".
+- ~/.dotfiles/hyprland/.local/bin/topgrade:40-42 has =command -v maint >/dev/null 2>&1; then maint stamp topgrade >/dev/null 2>&1 || true=, which skips the stamp silently.
+- The installer already templates the username into system files: archsetup:2595-2600 (BRIO rule: ARCHSETUP_USERNAME, then =sed -i "s/ARCHSETUP_USERNAME/${username}/"=).
+- maint's state path is ~ based (dotfiles maint/src/maint/cache.py:16-17, =os.path.expanduser("~/.local/state/maint")= unless MAINT_STATE_DIR is set), so the unit also needs the user's HOME. =User== provides it.
+- Counterpoint I checked: Decision L142 says the boot run is "one pacman transaction, no ecosystem sweep", so topgrade is not invoked at boot and wrapper resolution doesn't apply to this path.
+:END:
+
+** DONE Second-writer =maint apply-upgrade= path survives from the first draft
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design final paragraph L71; Alternative D Neutral L93
+
+L71: 'The cleanest single home for the stamp is a small maint apply-upgrade (or a flag in the existing lever) that performs the guarded system upgrade and stamps on success, which both the boot unit and an interactive TTY run call.' Alternative D: 'the same maint apply-upgrade path serves both'.
+
+Risk: It invites an implementer to build a third completion path and stamp writer in maint, which is the drift the paragraph itself warns against. Stamp ownership reads as contradictory.
+
+Recommended change: Three edits:
+- Replace L71 with: "=upgrade-guarded= is the one code path that completes an upgrade and records it. The everyday run stamps when nothing was deferred. =--complete= stamps when it lands the deferred set, whether the boot unit or a TTY invokes it. A by-hand sentinel-override upgrade outside the script does not stamp; =upgrade-guarded --complete= from a TTY replaces that ritual."
+- In Alternative D's Neutral (L93), change "the same =maint apply-upgrade= path" to "the same =upgrade-guarded --complete= path".
+- In L69, replace the ExecStart sentence ("=informant read= ... then =topgrade --only system= ... then =maint stamp topgrade= on success, then removes the flag") with "Its =ExecStart= runs =upgrade-guarded --complete= as the user (which clears news, applies the deferred GPU/compositor set, and stamps), then removes the flag unconditionally." This makes the unit match Decision 6 and Phase 3, and leaves the script as the only stamp writer.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Design, Alternatives. Confirmed. Its proposed boot ExecStart ('upgrade-guarded --complete ... then removes the flag unconditionally') conflicts with F08's distinct boot form and with my disarm-at-start call, so that sentence follows F08 and F10. The Design final paragraph and Alternative D rewrites stand, extended to name both completion forms.
+:EVIDENCE:
+- Spec L71 (current): "The cleanest single home for the stamp is a small =maint apply-upgrade= (or a flag in the existing lever) that performs the guarded system upgrade and stamps on success, which both the boot unit and an interactive TTY run call."
+- Spec L93: "the same =maint apply-upgrade= path serves both an interactive TTY run and the boot unit."
+- The contradicting sections:
+ - L121 (Decision 3): ship the split script in archsetup; "maint's UPDATE/TOPGRADE levers change their =argv= to the script".
+ - L142 (Decision 6): "the boot oneshot runs the script's =--complete= form".
+ - L155 (Phase 1): =scripts/upgrade-guarded=, with "=--complete= (... stamp when the deferred set is empty)".
+ - L161 (Phase 3): "=ExecStart= runs the script's =--complete= form as the user".
+ - L173 (AC): "via the Phase-1 path".
+- Draft provenance: =git show 77447d0:docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org= has the L71 text verbatim at its line 66 and the L93 text at its line 88. That draft's Phase 1 (its line 119) reads "A single =maint= entry point (=maint apply-upgrade=, ...)". Commit e777ddc rewrote the phases around the script but left these lines in place.
+- No implementation exists: =grep -rn "apply.upgrade" ~/.dotfiles/maint/src= finds nothing, and the only archsetup hits are in this spec.
+- Adjacent stale text: L69's ExecStart still lists "=topgrade --only system= ... then =maint stamp topgrade= on success". That conflicts with L142 and L161.
+:END:
+
+** DONE First-draft boot and dependency text remains in Risks, Readiness and Phase 4
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Risks L193; Readiness Errors L179, Performance L182, External deps L189; Phase 4 L164
+
+Risks: 'The --only system scope ... shrink this', and it lists 'AUR review' as a boot prompt. Readiness: 'one pacman/yay transaction at boot', and deps 'topgrade --only system, informant read, yay, maint stamp — all verified present on velox'. Errors: 'every failure path lands in boot normally ... re-arm'. Phase 4 flow: 'arm → reboot → console upgrade → session'.
+
+Risk: The implementer's dependency and failure checklists point at a removed command and miss the gate-failure state, so that state gets no observability or documentation.
+
+Recommended change: - L193 (Risks): replace "The =--only system= scope and =--noconfirm=-style flags shrink this to near zero" with "The boot form's single pacman transaction with =--noconfirm= shrinks this to near zero". Drop "AUR review" from the prompt list. Change "verified in Phase 2" to "verified in Phase 3".
+- L182: change "one pacman/yay transaction at boot" to "one pacman transaction at boot".
+- L189: replace the deps line with: "=checkupdates= (pacman-contrib, installer-owned), =yay=, =topgrade --disable system git_repos containers= (space-separated; topgrade 17.12.2 rejects the comma form), =informant read --all= where installed (velox; absent on ratio), =dkms status=, =zfs list -t snapshot= on a ZFS root, =maint stamp=."
+- Make the same two invocation fixes in L65 and L155: comma-separated =--disable= becomes space-separated, and =informant read= becomes =informant read --all=. Apply the =--all= fix in L69 and L172 too.
+- L179: append "A failed =--complete= gate leaves the new kernel installed, nothing armed and no reboot; the foreground run names the failure and the desktop stays up."
+- L164: change the flow to "kernel set live → gate → arm → reboot → console upgrade → session".
+- Optional, same leftover text:
+ - L187 and L200: Phase 1 tests are pytest beside the guard's tests, Phase 2 uses the maint fake harness, and Phase 3 uses tests/installer-steps/ plus maint tests.
+ - L69 and L71: replace the topgrade --only system / yay -Syu / maint apply-upgrade wording with "runs =upgrade-guarded --complete=".
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Design, Implementation phases, Readiness dimensions, Risks. Confirmed. The replacement deps line is completed with the commands other accepted findings add:
+- sudo and timeout for informant (F11);
+- yay -Pwq (F39);
+- pacman -Sp, -Sw and versioned -S (F08, F09);
+- vercmp;
+- systemctl is-enabled (F31);
+- lsinitcpio and findmnt (F19);
+- foot (F40).
+
+The Risks prompt-surface text is merged with F28's so the two edits don't collide.
+:EVIDENCE:
+Spec text (docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org):
+- L142: "one pacman transaction, no ecosystem sweep, no kernel"
+- L161: "ExecStart runs the script's --complete form"
+- L193: "provider choice, replace, AUR review ... The --only system scope ... verified in Phase 2"
+- L182: "one pacman/yay transaction at boot"
+- L189: "topgrade --only system, informant read, yay, maint stamp — all verified present on velox"
+- L179: "every failure path lands in boot normally"
+- L164: "arm → reboot → console upgrade → session"
+- L67, L135, L155, L169: gate failure stops, names the failure, does not reboot
+- L187: "installer-step pytest for Phase 2; maint unit tests for Phases 1 and 3". This inverts L155 (Phase 1 tests are pytest beside the guard's tests, tests/hypr-live-update-guard/) and L161 (Phase 3 tests go in tests/installer-steps/).
+- L69: "topgrade --only system (or the equivalent yay -Syu)"
+
+Commands run on ratio:
+- /usr/bin/topgrade --disable system,git_repos,containers --help → "error: invalid value 'system,git_repos,containers' for '--disable <STEP>...'". The space-separated form parses. topgrade --version reports 17.12.2. --help short-circuits any run, and the stamping wrapper was bypassed.
+- command -v: checkupdates=/usr/bin/checkupdates (owned by pacman-contrib 1.13.1-1, which the installer installs at archsetup:2175); dkms and zfs present; informant MISSING on ratio.
+
+Installer code:
+- archsetup:1124-1129: "--all marks without printing or prompting; a bare =informant read= is interactive and would hang an unattended run", then it runs informant read --all.
+:END:
+
+** DONE Test framework and phase-to-harness map contradict the phases
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Phase 1 L155 ('pytest'); Dev tooling L187; Risks L193; Testing L200
+
+L155 says 'pytest beside the guard's'. L200: 'Phase-1 and Phase-3 logic under the maint fake harness; Phase-2 install under tests/installer-steps/'. L187: 'installer-step pytest for Phase 2; maint unit tests for Phases 1 and 3'. L193 says boot non-interactivity is 'verified in Phase 2'.
+
+Risk: Following L200 puts the Phase 1 tests where they can't import the script. Following 'pytest' produces a red gate. The prompt-surface check is assigned to the wrong phase.
+
+Recommended change: Use one phase-to-test map and repeat it at L155, L187 and L200:
+- Phase 1 → archsetup tests/upgrade-guarded/ and tests/kernel-modules-check/. These use unittest + subprocess + env-var seams, with fake checkupdates, pacman, dkms and zfs on PATH (tests/zfs-pre-snapshot/fake-zfs can be reused).
+- Phase 2 → dotfiles tests/maint/ under the maint fake harness.
+- Phase 3 → archsetup tests/installer-steps/ for the install step, the arm action's tests in dotfiles tests/maint/, and the manual boot test in todo.org.
+
+Specific edits:
+- L155: change "pytest beside the guard's" to "unittest beside the guard's (tests/upgrade-guarded/, run by make test-unit)".
+- L187: change to "unittest for Phase 1 and the Phase 3 installer step in archsetup; maint unit tests for Phase 2 and the Phase 3 arm action; a manual boot test in todo.org".
+- L200: replace the first sentence to match this map.
+- L193: change "verified in Phase 2" to "the --noconfirm argv asserted in Phase 1's tests; no-stdin behavior verified by Phase 3's manual boot test".
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Readiness dimensions, Testing, Risks. Confirmed. The map follows the re-phasing. F16's same-commit rule moves the panel action into Phase 2, so Phase 3 is archsetup only, and F21's TOML pin test lands in Phase 1. The tests use unittest, which is what make test-unit runs: it globs tests/*/test_*.py, so new suites need no list edit.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L155: "Tests (pytest beside the guard's)".
+- L158: Phase 2 is maint wiring, "Tests under the maint fake harness".
+- L161: Phase 3: "Tests in =tests/installer-steps/= for the install step; the arm action's tests in maint".
+- L187: "installer-step pytest for Phase 2; maint unit tests for Phases 1 and 3".
+- L193: "verified in Phase 2".
+- L200: "Phase-1 and Phase-3 logic under the maint fake harness; Phase-2 install under =tests/installer-steps/=".
+
+The stale lines come from the first draft. In git show 77447d0, the phases are: Phase 1 "Stamp the safe-completion path", Phase 2 "The armed boot-time unit (archsetup)", Phase 3 "The arming affordance (dotfiles maint)". Its lines 148, 154 and 159 match the current L187, L193 and L200 word for word.
+
+Test conventions:
+- tests/hypr-live-update-guard/test_hypr_live_update_guard.py:1-21 is unittest with HYPR_GUARD_* env-var seams, and its docstring says "python3 -m unittest tests.hypr-live-update-guard...".
+- Makefile:58-68 (test-unit) loops over tests/*/test_*.py running python3 -m unittest "$mod" || fail=1.
+- grep for "import pytest" or bare "def test_" under archsetup tests/ finds nothing.
+- I ran a scratch pytest-style module through python3 -m unittest: "Ran 0 tests ... NO TESTS RAN", exit=5. pytest 9.1.1 is installed but no gate calls it.
+- ~/.dotfiles/tests/maint/test_remedies_doctor.py:18-27 sets REPO_ROOT to the dotfiles repo and puts only maint/src and panelkit/src on sys.path, with FAKE_TOOL = tests/maint/fake-tool.
+- ~/.dotfiles/Makefile:290-303 also globs tests/*/test_*.py through unittest.
+- tests/zfs-pre-snapshot/fake-zfs exists and could be reused.
+:END:
+
+** DONE yay -Sua doesn't hold AUR packages back, and its failure and privilege model is undefined
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L65 ('AUR packages pinning a guarded version hold themselves back'); Phase 1 L155 (=yay -Sua --noconfirm=; 'exit 0 on a successful live part')
+
+The script runs =yay -Sua --noconfirm= after the pacman run with --ignore. It assumes AUR packages that need a newer guarded library defer themselves.
+
+Risk: When it does fire, the spec doesn't say whether a failed AUR step counts as a failed live part, or what that does to the exit code, the stamp and the deferred badge. The implementer also has to invent the split between user and root steps.
+
+Recommended change: 1. L65: replace "(=yay -Sua --noconfirm=, AUR packages pinning a guarded version hold themselves back)" with "(=yay -Sua --noconfirm=; yay still resolves repo dependencies normally, so an AUR upgrade that needs a newer held package makes yay's own pacman call trip the guard and the AUR step fails)". Change "on the driven path it never fires" to "on the driven path it fires only in that AUR case".
+
+2. Phase 1 (L155): add one rule. Suggested default, consistent with the state framing: "A non-zero yay exit is a failed live part. The script still runs the topgrade sweep, records the AUR failure in the state file so the panel shows it beside the deferred count, does not stamp, and exits non-zero." Add a test for it.
+
+3. Phase 2 (L158): reword the parenthetical so the force-wrap removal no longer rests on "never trips the guard".
+
+4. L167: change "the guard hook does not fire" to "the guard hook does not fire from the pacman step". Add a criterion: "an AUR upgrade needing a held library fails the AUR step visibly without blocking the rest of the run".
+
+No change is needed for the privilege model; L69, L161 and L180 already cover it. If I want it explicit, add one clause to Phase 1: "runs as the user; pacman, informant and the arm-flag write go through sudo".
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Design, Implementation phases, Acceptance criteria, Risks, Testing. Partially confirmed; this follows the narrowing that privilege and the stamp are settled. The privilege clause is made explicit, including the EUID-0 refusal.
+
+Two Risks lines are added for the paths the pacman hold can't cover. yay's resolver can move a kernel-set or DKMS-set package when an AUR upgrade requires a newer one. And a DKMS package installed from the AUR would be upgraded by yay -Sua with no hold at all. Neither applies today on ratio, where zfs-dkms comes from the archzfs repo; checklist step (a) confirms velox.
+:EVIDENCE:
+- Spec L65: "runs the AUR-only remainder (=yay -Sua --noconfirm=, AUR packages pinning a guarded version hold themselves back)" and "on the driven path it never fires".
+- Spec L155: "=yay -Sua --noconfirm= → ... → exit 0 on a successful live part". There is no yay failure rule.
+- Spec L158: "the press-again-to-force sentinel wrap goes away (the driven path never trips the guard)".
+- Spec L167: "the guard hook does not fire".
+- man yay, -a/--aur: "Actions such as sysupgrade will only act on AUR packages. Note that dependency resolving will still act normally and include repository packages."
+- scripts/hypr-live-update-guard: L62 and L67 check pgrep for Hyprland, L86-105 split targets by version change, L121 prints the BLOCKED banner, L139 exits 1.
+- The installed hook (/etc/pacman.d/hooks/hypr-live-update-guard.hook) has AbortOnFail and targets mesa, mesa-*, wayland, libdrm, libglvnd and the others.
+- yay 13.0.1, yay -Pg: sudobin "sudo", sudoloop false.
+- /etc/pacman.conf:25 has IgnorePkg = bridge-utils only.
+- Foreign-package deps on guarded libraries are all unversioned or soname-only: claude-desktop (libdrm, mesa), insync (libglvnd), webkit2gtk (libdrm, mesa, wayland), zoom (mesa, libdrm), mpvpaper (libwayland-client.so=0-64, libwayland-egl.so=1-64). None of the AUR packages is itself a guard target.
+- Privilege is already stated: spec L69 ("Its =ExecStart= runs, as the user ... =sudo= works unattended"), L161 ("runs the script's =--complete= form as the user"), L180 ("the unit runs the upgrade as the user via sudo").
+- makepkg:1240-1243 refuses EUID 0.
+- The stamp is already covered: L128 says stamp only when the deferred set is empty, and in this scenario it is not empty.
+:END:
+
+** DONE No one-time install mechanism, and Phase 2 breaks the panel on an uninstalled machine
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Decision L122; Phase 1 L155; Phase 4 L164; Rollout L188; Testing L200
+
+The spec says 'existing machines need the one-time install' but gives no command. The ordering is stated only as commit order ('two commits, archsetup first'), and the install itself is deferred to Phase 4, after Phases 2 and 3.
+
+Risk: The commits land in order, but each machine picks them up independently. A machine that pulls Phase 2 before Phase 1 is hand-installed gets UPDATE/TOPGRADE failures, with the force path already removed. velox, which is offline, will hit this on its next pull.
+
+Recommended change: Replace the L164 sentence "existing machines need the one-time install" with an explicit per-machine checklist, and point L188 at it:
+
+"Existing machines: hyprland() runs inside the window_manager step, which is already marked complete in /var/lib/archsetup/state/, so re-running archsetup won't install anything new. On each machine, do these by hand, in this order:
+(a) After the Phase-1 commit: pull archsetup, then =install -m 755 scripts/upgrade-guarded scripts/kernel-modules-check /usr/local/bin/=.
+(b) Only after (a) on that machine: pull the Phase-2 dotfiles commit. The levers call upgrade-guarded and no longer have a force path, so pulling first breaks UPDATE and TOPGRADE on that machine.
+(c) After the Phase-3 commit: install archsetup-boot-upgrade.service, run =systemctl daemon-reload=, enable the unit, and confirm /var/lib/archsetup/ exists.
+Re-copy configs/maintenance-thresholds.toml only if Phase 2 adds keys to it.
+velox's dotfiles must not move past the Phase-2 commit until (a) is done there."
+
+Optional, in Phase 2: when upgrade-guarded is not on PATH, the lever reports "upgrade-guarded not installed (archsetup one-time install)" rather than the generic "tool missing or timed out". Don't fall back to the old yay/topgrade argv; that would contradict the decision at L114.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Readiness dimensions. Confirmed. The checklist gains items other resolutions make mandatory:
+- the hook migration (F20), which must precede first use because the script fails closed;
+- an unconditional TOML re-copy, because F21 rewrites guard_patterns;
+- a DKMS-origin check (F29).
+
+The checklist is keyed to three commits after the re-phasing. The optional missing-script lever message is adopted. Falling back to the old argv stays rejected, as the finding's check showed.
+:EVIDENCE:
+Spec:
+- L122: "the rollout is two commits, archsetup first".
+- L155: script "installed to /usr/local/bin by the step that installs the guard"; kernel-modules-check is a second new script.
+- L158: levers "change their argv to the script; the press-again-to-force sentinel wrap goes away".
+- L161: Phase 3 adds the unit, the installer step and the /var/lib/archsetup/ flag directory.
+- L164: "existing machines need the one-time install", with no steps given.
+- L188: same wording, no steps.
+- L200: "Roll to velox first, then ratio".
+
+archsetup (installer):
+- Line 257: state_dir="/var/lib/archsetup/state".
+- Lines 290-296: run_step prints "Skipping %s (already completed)" when the marker exists.
+- Lines 3994-3996: every step goes through run_step.
+- Lines 52-56 and 85-99: the only CLI flags are --fresh, --status, --config-file and --autologin. There is no per-step re-run.
+- Lines 2606-2654: guard binary cp plus hook heredoc, inline in hyprland() (defined at 2534).
+- Lines 2676-2683: window_manager() calls hyprland.
+- Lines 1640-1650: thresholds TOML installed with =install -m 0644=. That's a copy, not a stow link, and it's called from user_customizations (1449).
+
+Live system:
+- =ls /var/lib/archsetup/state/= lists window_manager, user_customizations, supplemental_software and others, so re-running archsetup on ratio skips the guard step.
+
+todo.org:
+- Lines 4006-4010: "supplemental_software is a completed step there and doesn't re-run". That is the same mechanism, already hit and fixed by hand per machine.
+
+dotfiles (maint):
+- remedies.py:294-312: update/topgrade levers have "always": True and fixed argv.
+- capability.py:65-81: which() checks only for zpool/snapper/docker/virsh.
+- cmd.py:21-24: OSError returns None.
+- doctor.py:161-163: None is reported as "tool missing or timed out".
+- ~/.local/bin/maint and ~/.local/bin/topgrade are symlinks into ~/.dotfiles/hyprland/.local/bin, so a dotfiles pull is live immediately.
+- No dotfiles auto-pull timer is installed (systemctl --user list-timers), so ordering per machine can be controlled by hand.
+:END:
+
+** DONE Phase 1 ships arming before any boot consumer exists
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Phase 1 L155 (--complete arms); Phase 3 L161 (unit + flag dir)
+
+Phase 1's --complete 'arm[s] the GPU/compositor set', and the gate refuses to 'arm or reboot'. The unit that consumes and removes the flag, and the /var/lib/archsetup flag directory, arrive only in Phase 3.
+
+Risk: A --complete run between Phase 1 and Phase 3 writes a flag that nothing removes, and its reboot offer applies nothing. When the unit lands, the stale flag fires on the next boot with a package set nobody chose for that day.
+
+Recommended change: In Phase 1 (L155), add a rule to the --complete description: --complete arms the GPU/compositor set only when archsetup-boot-upgrade.service is installed. Without the unit, it stops after a passing gate. The GPU/compositor set stays in the deferred state file, the script prints "apply from a TTY: upgrade-guarded --complete", and it offers no reboot for that set. Add a Phase 1 test: with the unit absent, --complete never writes the flag and never offers the GPU reboot. This covers both the window between Phase 1 and Phase 3 and the L188 rollback state. The alternative, moving the arm write into Phase 3, fixes the phase window but not rollback.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Testing. Partially confirmed; this follows the narrowing that the gap is not unsafe and the flag directory already exists. Detection is pinned to one command, systemctl is-enabled, which also covers a unit that is installed but not enabled (F41). Any stale flag is removed, so nothing is left for a later unit to consume unexpectedly.
+:EVIDENCE:
+- Spec L155 (Phase 1): "--complete (the dedicated-session form: apply the kernel set live, run the gate, then arm the GPU/compositor set or, with no compositor live, apply it directly ...)". Further on: "--complete refuses to arm or reboot on that exit. Usable from a TTY at once." Test: "--complete never reaches the arm step on a failed gate."
+- Spec L161 (Phase 3): "archsetup-boot-upgrade.service, installed by the installer ... Installer step + unit file + the /var/lib/archsetup/ flag directory."
+- Spec L67: once the gate passes, "the script arms a persistent flag ... and offers to reboot".
+- Spec L188: "removing the unit and flag reverts fully". The script stays installed, so it can still arm with no consumer.
+- Live system: systemctl cat archsetup-boot-upgrade.service returns "No files found", so the unit does not exist yet. /usr/local/bin has only hypr-live-update-guard and no upgrade-guarded.
+- Live system: /var/lib/archsetup exists and contains state/. The installer at archsetup:257 sets state_dir="/var/lib/archsetup/state", so the directory predates Phase 3 and does not block a Phase-1 flag write.
+- Spec L142 (Decision 6): the boot oneshot is "scoped to the deferred GPU/compositor set ... no kernel". A late-firing flag applies that set with nothing live, the designed safe path, so it is not an unchosen package set.
+:END:
+
+** DONE Acceptance criteria aren't mapped to verification, and the manual test covers only AC4
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Acceptance criteria L167-174; Phase 3 L161; Testing L200
+
+The only manual test is 'arm with a guarded lib pending, reboot, confirm', with no setup for producing a pending guarded lib. Nothing verifies AC3 (DKMS failure stops; snapshot bootable from ZBM), AC5 (boot failure or timeout never blocks the session) or AC6 (unread news). Rollout goes straight to velox.
+
+Risk: Goal 3's safety claim ('can never lock the machine out') and the ZBM fallback are never exercised before velox, where the failure mode is an unbootable machine. The happy-path test can only run on days when upstream happens to publish a guarded update.
+
+Recommended change: Replace the L200 paragraph with a list mapping each acceptance criterion to how it is verified, and correct L187 to match:
+
+- AC1, AC2, AC7: Phase 1 pytest beside the guard (ignore set equals blocked plus kernel set, deferred-state round-trip, stamp only on an empty deferred set) and the Phase 2 maint probe test.
+- AC3: Phase 1 gate tests cover stopping, naming the failure, and never arming. Either make the ZFSBootMenu snapshot-boot clause a documented recovery drill (in a =make test-keep FS_PROFILE=zfs= VM, or on velox) in todo.org, or move it from the AC into Risks as the fallback.
+- AC4: keep the manual test, but give it a trigger and a setup. Run it when =upgrade-guarded --dry-run= lists a GPU/compositor package, or in the VM create a pending guarded lib from a TTY by installing an older cached version.
+- AC5: add an installer-steps test that asserts the unit file statically. It should check Before=getty@tty1.service, a TimeoutStartSec, no Requires/BindsTo/RequiredBy tying it to getty or session targets, and flag removal placed where it also runs on timeout (ExecStopPost or remove-first). Also add a VM or manual test with a failing and a sleeping ExecStart test seam. Expected: boot reaches the Hyprland session, the flag is gone, and the panel still shows the deferred set.
+- AC6: add a Phase 1 pytest with a fake informant asserting that =informant read --all= runs before the pacman transaction in both the live and --complete forms. Also correct the bare =informant read= at L65, L69 and L172.
+- AC8: the existing tests/hypr-live-update-guard suite passes unchanged.
+
+Finally, state that the Phase 3 boot-unit checks run in =make test-keep FS_PROFILE=zfs= before the velox rollout.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Acceptance criteria, Testing, Risks, Readiness dimensions. Partially confirmed; this follows the narrowing. The map covers the consolidated acceptance list (glossary §18), which adds the criteria other findings introduce, and uses unittest, not pytest (F28). Of the two options for AC3's ZFSBootMenu clause, it moves to Risks as the fallback and is exercised by a non-gating recovery drill. No gate can check it before rollout, and the drill is where F45's pacman -U --dbonly step gets practiced. AC5's static unit test is the single source-inspection test owned by F41.
+:EVIDENCE:
+Spec:
+- L161: "a documented manual boot test (defer, arm, reboot, observe)"
+- L200: "scripted manual test ... (arm with a guarded lib pending, reboot, confirm ...). Roll to velox first, then ratio." No setup step for producing a pending guarded lib.
+- L155: Phase 1 tests include the gate pass/fail on fakes and "--complete never reaches the arm step on a failed gate", so AC3's stop half is covered. No test mentions informant, the unit's failure or timeout behavior, or the ZBM boot.
+- L184: the ordering is only "tested-by-inspection".
+- L69: ExecStart "... then removes the flag unconditionally", i.e. inside ExecStart.
+- L171: AC5 requires the flag to be cleared on a timeout.
+- L187 vs L155/L157/L200: the phase numbers do not match the phases.
+
+Code and system:
+- =pacman -Q informant= on ratio: "error: package 'informant' was not found".
+- archsetup:1121-1129: "If the base ships informant (e.g. an archangel-installed system) ... a bare =informant read= is interactive and would hang an unattended run", and the installer uses =informant read --all=.
+- Makefile:12-14, 35-36: FS_PROFILE ?= btrfs, with zfs available for test, test-keep and test-vm-base. Makefile:32/108: test-maint runs the break/fix scenarios.
+- scripts/testing/archsetup-test-zfs.conf exists.
+- scripts/testing/tests/test_boot.py:14-18 asserts /efi/EFI/ZBM/zfsbootmenu.efi on a ZFS root.
+- scripts/testing/maint-scenarios/ holds the 10-32 break/fix scripts.
+- =grep -rln 'Before=\|TimeoutStartSec\|\[Unit\]' tests/= returns nothing, so no existing static unit-file tests cover this.
+:END:
+
+** DONE Rollback is described as 'remove unit and flag', but the change isn't additive
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Readiness 'Rollout, compatibility & rollback' L188
+
+'additive; removing the unit and flag reverts fully.'
+
+Risk: Removing only the unit and flag leaves the panel calling the script with no force path. A partial rollback leaves the two repos disagreeing.
+
+Recommended change: Replace the "Rollout, compatibility & rollback" bullet at spec line 188 with:
+
+"Rollout, compatibility & rollback: not additive. Phase 2 repoints the UPDATE/TOPGRADE argv and removes the press-again sentinel wrap. Roll back in reverse rollout order:
+1. Revert the dotfiles Phase 2/3 commits first. That restores the yay/topgrade argv and the force path, and removes the deferred probe and the apply-on-reboot action.
+2. Then remove archsetup-boot-upgrade.service and run daemon-reload.
+3. Remove the /var/lib/archsetup/ arm flag, and /usr/local/bin/upgrade-guarded and kernel-modules-check.
+
+Removing only the unit and flag disables the boot path but leaves apply-on-reboot arming a flag nothing reads. Removing the archsetup scripts before reverting dotfiles leaves UPDATE/TOPGRADE pointing at a missing binary. The rewritten guard_patterns can stay. Existing machines need a one-time install, and a rebuild gets it from the installer. The guard hook is untouched throughout."
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Readiness dimensions. Partially confirmed; this follows the narrowing, adjusted for the three-commit rollout, the sysupgrade alias (F38), and the hook migration, which is not reverted. With F31 in place, removing only the unit no longer leaves --complete arming a flag that nothing reads.
+:EVIDENCE:
+- Spec line 188: "Rollout, compatibility & rollback: additive; removing the unit and flag reverts fully ... Rollback leaves the guard and manual TTY path intact."
+- Spec line 158 (Phase 2): "UPDATE and TOPGRADE levers change their argv to the script; the press-again-to-force sentinel wrap goes away."
+- Spec line 121: "the rollout is two commits, archsetup first."
+- Spec line 161 (Phase 3): the panel action "installs the held-kernel set live, writes the persistent arm flag, and offers to reboot."
+- Spec line 155: Phase 1 installs scripts/upgrade-guarded and kernel-modules-check.
+- Spec line 174: the hook is unchanged (already an acceptance criterion).
+- Dotfiles code I checked:
+ - ~/.dotfiles/maint/src/maint/remedies.py:282-293 holds guard_sentinel_set and guard_sentinel_clear.
+ - remedies.py:297 sets the update argv to ["yay","-Syu","--noconfirm"]. remedies.py:308 sets the topgrade argv to ["topgrade","--disable","git_repos","-y"]. Both have guard "live_update".
+ - ~/.dotfiles/maint/src/maint/doctor.py:174-188 is _all_steps, which adds the sentinel wrap. doctor.py:242 runs that wrap when force is set.
+ These are exactly what Phase 2 deletes or repoints, so the change is not additive.
+- Thresholds TOML:
+ - configs/maintenance-thresholds.toml:65-72 guard_patterns are mesa, mesa-*, lib32-mesa*, hyprland, hyprland-*, aquamarine, hyprutils, hyprlang, hyprcursor, hyprgraphics, *wayland*, wlroots*, vulkan-radeon, lib32-vulkan-radeon, vulkan-intel and lib32-vulkan-intel.
+ - The hook Target list in archsetup:2628-2642 adds libdrm, libglvnd, vulkan-mesa-layers, nvidia-utils, lib32-nvidia-utils and xorg-xwayland, and has no hyprlang, hyprcursor or wlroots*.
+ - So the Phase 2 pin test does force a TOML rewrite. The maint thresholds.py:31-33 reads this file as the shipped layer at ~/.config/archsetup/maintenance-thresholds.toml.
+- Cache location: maint cache.py:16-17 is ~/.local/state/maint, which is where the deferred-set "own cache key" from spec line 128 would land.
+:END:
+
+** DONE Armed-but-not-rebooted state is invisible, and REBOOT semantics are ambiguous
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L67 ('arms … and offers to reboot'); Decision 4 L128 ('apply on reboot'); Phase 2 L158; Phase 3 L161
+
+No phase says how the panel shows a pending arm, how the 'offer to reboot' appears in the panel, or how to disarm.
+
+Risk: A user reading 'N deferred — apply on reboot' may press the REBOOT key the strip already shows, and nothing is applied because nothing was armed. A user who arms on ratio gets no reboot affordance at all. On both machines, a pending arm is invisible.
+
+Recommended change: Phase 2 (L158): replace "N deferred — apply on reboot" with "N deferred". Pair it with a key that runs =upgrade-guarded --complete=, and say plainly that a plain reboot applies nothing until a --complete run has armed. Fix the example string at L128 the same way; this changes only the illustration, not the stamp decision. Phase 3 (L161): the deferred probe also reads the arm flag. While the flag is present, the row reads "armed — N apply at next boot". The panel action calls offer_reboot() when --complete exits armed, rather than relying on reboot_required, which stays false on ratio because the booted linux-lts-strix is a local build the kernel set never upgrades. --complete prompts for a reboot only when stdin is a tty; from the panel, the REBOOT key is the offer. Add one acceptance criterion: after arming from the panel, the row shows the armed state and REBOOT is offered, on both ratio and velox. A --disarm command is optional and can be left out.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Decisions, Implementation phases, Acceptance criteria. Partially confirmed; this follows the narrowing, so there is no --disarm. Under F40 the panel launches --complete in a detached terminal and never sees its exit. 'Call offer_reboot() when --complete exits armed' therefore becomes a state rule: REBOOT shows while the arm flag exists. That survives a panel close, and it works on ratio, where reboot_required stays false.
+:EVIDENCE:
+Spec:
+- L67: "arms a persistent flag ... and offers to reboot"
+- L128: "('6 deferred — apply on reboot')"
+- L135: the GPU set follows only after the gate passes
+- L158: "A new probe reads the deferred-set state file and the panel renders 'N deferred — apply on reboot' as its own row" (no arm-flag input)
+- L161: "The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot" (no display of the armed state, no mechanism for the offer)
+
+Code (~/.dotfiles/maint/src/maint):
+- gui.py:1509-1511: if rid in ("update","topgrade") and primary ok, call self.model.offer_reboot()
+- panel.py:501-510: reboot_key_visible() is reboot_offer OR the reboot_required value
+- gui.py:491-495: the REBOOT strip key is gated on reboot_key_visible()
+- packages.py:131-146: reboot_required is true only when /usr/lib/modules/$(uname -r) is missing
+- cmd.py:18-24: subprocess.run(cmd, capture_output=True, text=True, timeout=...)
+- doctor.py:149-161: user steps go through cmd.run
+- grep -rn "/var/lib/archsetup|deferred|upgrade-guarded" over maint/src/maint finds nothing
+
+Live checks on ratio:
+- uname -r: 6.18.25-1-lts-strix
+- pacman -Qo /usr/lib/modules/6.18.25-1-lts-strix: owned by linux-lts-strix 6.18.25-1
+- pacman -Si linux-lts-strix: "package 'linux-lts-strix' was not found"
+- pacman -Qm lists linux-lts-strix; Packager: Unknown Packager
+- No linux-lts-strix-headers is installed. linux, linux-lts, and both of their headers packages are separate packages.
+:END:
+
+** DONE Deferred row goes stale after a successful boot run or out-of-band completion
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Decision 4 L128; Design L71 (by-hand path); Decision 2 L114 (manual TTY path C); AC L167
+
+The boot run 'stamps when it completes the deferred set', but the spec never says it rewrites the deferred state. A by-hand sentinel run or a bare TTY pacman -Syu never touches the file at all.
+
+Risk: After a successful boot run, the panel shows a fresh topgrade_age next to 'N deferred', and the two contradict each other. AC L167 ('exact deferred set') fails after any out-of-band completion.
+
+Recommended change: Phase 1 (L155) and Decision L128:
+- Add: "Every mode, --complete and the boot run included, rewrites the deferred-set state file at the end of its run with what is still deferred, empty when nothing is. --complete drops the kernel set once it is installed and the GPU/compositor set once it is applied."
+- In L128, change "When it is non-empty it writes the deferred set" to "It always writes the deferred set (empty when nothing was held)".
+- Add a Phase 1 test: after a successful --complete or boot run, the state file is empty and the stamp is fresh.
+
+Phase 2 (L158), optional robustness for the out-of-band paths L71 names: the deferred probe drops any recorded entry whose installed version (pacman -Q, compared with vercmp) is at or above the recorded new version. A by-hand sentinel or TTY completion then clears the row without rerunning the script.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Decisions, Implementation phases, Testing.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L128: "the script stamps topgrade_run only when the deferred set is empty. When it is non-empty it writes the deferred set to its own cache key ... The boot oneshot stamps when it completes the deferred set." This line has no clear and no rewrite.
+- L155: the live pipeline includes "deferred set written to a state file". The --complete clause reads "apply the kernel set live, run the gate, then arm ... or ... apply it directly; stamp when the deferred set is empty", with no state-file write.
+- L142: the boot oneshot runs "the script's --complete form scoped to the deferred GPU/compositor set".
+- L158: the probe "reads the deferred-set state file".
+- L181: "the panel reflects the cleared or still-pending state after boot".
+- AC L170 checks only topgrade_age after the boot run, not the deferred row.
+
+Code: ~/.dotfiles/maint/src/maint/doctor.py:309-324 refreshes the pending cache (p_updates.scan_net) only after a successful update/topgrade lever run. ~/.dotfiles/maint/src/maint/probes/updates.py:145-160 shows scan_net writing updates_repo from checkupdates and updates_aur from yay -Qua, and nothing else deferral-related. The user unit maint-net-scan.timer runs it hourly (OnUnitActiveSec=1h). grep for "deferred" and "upgrade-guarded" across maint/src and archsetup/scripts returns nothing, so no existing code pins this behaviour.
+:END:
+
+** DONE CVE strip and card claim held packages are fixable via UPDATE
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Decision 5 consequences L136 ('security fixes ride the kernel'); Risks L197; Phase 2 L158
+
+The spec doesn't mention the CVE metric, which treats every queued package as something UPDATE will land.
+
+Risk: A fixable kernel or mesa advisory stays red after a successful UPDATE that, by design, cannot apply it. On the one signal that tracks security, the panel points at the wrong action.
+
+Recommended change: Add to Phase 2 (L158), after the deferred-row sentence: "Advisories on deferred packages count against the deferred row, not UPDATE. cve_queued (the red badge and strip), the CVE card's 'fixable via UPDATE' caption, and the UPDATE lever's cve_advisories binding all exclude names in the deferred set. The deferred row names those advisories ('N deferred · K fixable advisories') and goes to WARN when K > 0, so the kernel security fixes from the kernel decision have a visible reminder that points at --complete." Add a Phase 2 test: given a fixable advisory naming a deferred package, cve_queued returns 0, the card caption leaves it out, and the deferred row carries it at WARN. Optionally, extend acceptance criterion L167 with: "a fixable advisory on a deferred package shows on the deferred row, not as UPDATE-fixable."
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Acceptance criteria, Testing. Confirmed. This adds the finding's wider scope: REVIEW & FIX also stops offering UPDATE for an advisory on a deferred package. It also absorbs F48's cve_queued clause, so the exclusion has a single owner.
+:EVIDENCE:
+Spec: L65 and L135 hold the kernel and guard sets on every everyday run. L136 says "security fixes ride the kernel". L158 is the full Phase 2 display change, a deferred probe and the "N deferred — apply on reboot" row, with no CVE handling. L197 names the deferred row as the reminder for kernel security fixes. L167 (acceptance) covers only the deferred set.
+
+Code (~/.dotfiles/maint/src/maint/):
+- viewmodel.py:503-516: the cve_queued docstring says the strip "only turns red when running UPDATE would actually close an advisory". It counts any fixable advisory whose package is in updates_repo or updates_aur.
+- viewmodel.py:1012-1018: _card_cve returns "{fixable} fixable via UPDATE".
+- gui.py:641-672: _cve_queued feeds both the faceplate badge and the red strip class from _cache_list("updates_repo").
+- probes/updates.py:75-97: the cves() docstring says "the fix is one UPDATE away", and the metric is WARN whenever any advisory has =fixed= set.
+- remedies.py:294-302: the "update" lever has metric_ids ["updates_repo","updates_aur","cve_advisories"].
+- doctor.py:400-405: review() offers that lever whenever cve_advisories is WARN.
+- probes/updates.py ~160: updates_repo is written from checkupdates, so held packages stay in the queue after a split run.
+
+Live (ratio): checkupdates currently lists linux, linux-headers, linux-lts, linux-lts-headers, mesa, hyprland, vulkan-radeon and vulkan-mesa-implicit-layers. That is exactly the set a split run would hold, so every one of them would stay in updates_repo. arch-audit currently reports no fixable advisories, so the scenario is latent rather than active today.
+:END:
+
+** DONE Stale-kernel tracking is neither in v1 nor a tracked vNext item
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Scope tiers L57 ('vNext: none open'); Decision 5 consequences L136; Risks L197 ('a stale-kernel age … is a possible follow-up')
+
+The spec says the dedicated session 'has to happen on a cadence'. The only reminder is the deferred row, which is present most days, and the stale-kernel age is a 'possible follow-up' that no scope tier lists.
+
+Risk: The security half of the standing kernel hold has no owner. A row that shows nearly every day habituates the user, so an unpatched kernel ages silently.
+
+Recommended change: Change two spec lines, and keep v1 scope as it is:
+1. L57: replace "vNext: none open — …" with "vNext: a stale-kernel age in maint (how long the kernel set has been held, graded so an overdue dedicated session shows up); logged to todo.org as [#D]. The original live-kernel vNext is now part of v1's held set."
+2. L197: change "a stale-kernel age in maint is a possible follow-up" to "a stale-kernel age in maint is the vNext follow-up (see Scope tiers)".
+
+At decomposition, outside the spec, replace todo.org:740's stale "file the vNext [#D] kernel-reboot item" with that stale-kernel-age [#D] item. Do not add deferred_warn_days or hold-age grading to v1 unless I choose to.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Goals and Non-Goals, Risks.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L57: "vNext: none open — the kernel deferral that was vNext is now part of v1's held set."
+- L136: "the kernel deferral is now standing, so the dedicated session has to happen on a cadence (security fixes ride the kernel), and the panel's deferred count carries a kernel most days"
+- L197: "The panel's deferred row is the reminder; a stale-kernel age in maint is a possible follow-up."
+- git show 77447d0 L54, the original vNext: "fold the same arm-and-reboot pattern into a live-kernel-upgrade prompt (log to todo.org)". e777ddc, the decision-closing commit, added L57's "none open" and L197 together.
+
+todo.org:
+- :739-740: "...file the vNext =[#D]= kernel-reboot item, and commit the spec."
+- :759: the stale note covers only "the 'four open decisions' list above".
+- A grep for stale-kernel, kernel age or kernel-age finds no task.
+
+.ai/workflows/spec-review.org:
+- :33: "Deferred work is logged to todo.org (v1 = [#B], vNext/someday = [#D])"
+- :122: "Are deferrals captured (todo.org)?"
+
+~/.dotfiles/maint/src/maint/probes/updates.py:88-89:
+- =fixable = [a for a in data if a.get("fixed")]; sev = schema.WARN if fixable else schema.OK=
+
+~/.local/state/maint/cve.json, kernel entries:
+- AVG-2701 and AVG-2683 [linux-lts], fixed=None
+- AVG-1879, AVG-2345 and AVG-1594 [linux], fixed=None
+- So the CVE probe never flags a held kernel.
+
+There is no kernel or deferred threshold in configs/maintenance-thresholds.toml [updates] (lines 57-75).
+:END:
+
+** DONE Routine update entry points outside the panel still land the kernel ungated
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Non-Goals L51; Design L65 ('backstop for a bare pacman -Syu'); Decision 5 consequences L136; Phase 4 L164; Documentation plan L186; history L204-206
+
+The spec routes only the panel's UPDATE and TOPGRADE levers through upgrade-guarded. The documented routine update path is untouched: system-health-check Phase 3 step 6 runs plain topgrade, and the sysupgrade alias runs topgrade. topgrade's system step is yay -Syu, which upgrades every kernel and zfs-dkms. The guard hook has no kernel Target, so on kernels it isn't the backstop L65 calls it. The same workflow also tells the operator to run 'maint stamp topgrade' by hand. The containers fix lives only in the script's CLI --disable, so a plain topgrade still fails on containers, and that failure is what drives the hand stamp.
+
+Risk: Decision 5's consequence, 'an everyday UPDATE can never put velox into the unbootable state', holds only for the panel. The update velox actually gets, the health-check topgrade, still swaps linux-lts and zfs-dkms with no DKMS/initramfs/snapshot gate. That is the exact failure chain the decision was reversed to close. The hand-stamp instruction also writes freshness while a deferral is outstanding.
+
+Recommended change: Three edits.
+
+1. Add to Phase 4 (L164), after "Document the flow":
+
+"Update docs/workflows/system-health-check.org Phase 3 so the routine update goes through the script. Step 6 runs upgrade-guarded instead of plain topgrade. A day with a pending kernel or guarded library lands it through upgrade-guarded --complete as the dedicated session. Step 5's ratio strix addendum keys on --complete instead of 'before topgrade'. The 'maint stamp topgrade by hand' fallback becomes 'the script stamps; never hand-stamp while a deferred set is outstanding'. Append a resolution note to the 2026-09-12 containers KIL entry rather than rewriting it."
+
+2. Add the health-check workflow to the Documentation plan line (L186).
+
+3. Add one line to Risks, to spell out the consequence of L51:
+
+"Bare topgrade, yay and pacman -Syu (including the sysupgrade alias) still land the kernel set with no DKMS gate, because the hook is silent on kernels. The hold covers only the script's callers."
+
+Optional: point the sysupgrade alias at upgrade-guarded.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Readiness dimensions, Risks. Confirmed in substance; this follows the narrowing that the bare-command exposure needs one explicit Risks line. The optional alias change is taken. sysupgrade is the routine terminal path on both machines (alias sysupgrade="topgrade" at line 42 of both shell alias files). While it points at plain topgrade, it is exactly the ungated kernel and zfs-dkms route that Decision 5 closes. Repointing it is a one-line dotfiles change in the same commit as the levers, and it inherits the same step (a) ordering.
+:EVIDENCE:
+Spec facts:
+- L51: "The hook stays silent on kernels; the split script holds them back on a live run as a second, separately-reasoned list ... a deferral policy rather than a guard." This already concedes that bare commands do not gate kernels.
+- L65: the backstop sentence refers to the GPU guard.
+- L131 title: "lands only in the dedicated session". L136: "an everyday UPDATE can never put velox into the unbootable state".
+- L164 Phase 4 and L186 Documentation plan never mention docs/workflows/system-health-check.org. Grepping the spec for health, sysupgrade or workflow finds only L206: "Artifacts: velox health checks of 2026-09-12 and 2026-10-04".
+
+Code and docs:
+- docs/workflows/system-health-check.org:236, Phase 3 step 6: "Run =topgrade= for the actual update ... If the run happened outside the wrapper somehow, =maint stamp topgrade= records it by hand."
+- :235, step 5: the ratio strix addendum runs "before topgrade" when linux or linux-lts is pending.
+- :698: "If topgrade is present, Phase 3 uses it rather than the raw package manager."
+- :1067: the KIL entry tells the operator to run "=maint stamp topgrade= by hand".
+- todo.org:762-765 (2026-09-17): "velox has linux-lts 6.18.51 -> 6.18.52 pending today, which is the kernel-hold case this spec exists for, so no plain topgrade on velox until the hold is built or the kernel update runs as its own session."
+- .dotfiles/common/.zshrc.d/aliases.sh:42 and .bashrc.d/aliases.sh:42: alias sysupgrade="topgrade".
+- .dotfiles/common/.config/topgrade.toml:131: arch_package_manager = "yay". The misc disable list at :19 is ["emacs","poetry","gnome_shell_extensions","lensfun","uv"], with no containers, so a plain topgrade still fails at the containers step.
+- archsetup:2625-2649: the hook Targets list no linux* package.
+- /etc/pacman.conf IgnorePkg = bridge-utils only.
+- pacman -Q on ratio shows linux, linux-lts, linux-lts-strix and zfs-dkms 2.4.4 installed, all with nothing pinning them.
+- hyprland/.local/bin/topgrade wrapper comment: it stamps for "the system-health-check workflow, a plain shell".
+:END:
+
+** DONE The split run discards Arch news unread
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L65 ('clears the news hook (informant read) where installed'; topgrade --disable system); Design L69; AC L172; Reuse L183
+
+Each everyday run marks all Arch news read before the transaction and disables topgrade's system step. Nothing in the flow shows the news. Today the news gets attention two ways: on velox, informant's AbortOnFail hook blocks the transaction until the news is read, and topgrade's show_arch_news prints it during the system step. The split run removes both, and maint has no news surface.
+
+Risk: Manual-intervention notices are often about exactly the mesa, wayland and kernel transitions this design defers. With this change they're silently marked read, or never shown on ratio, right before an unattended 'pacman -Syu --noconfirm'. That removes the one interlock that forced a human to read them.
+
+Recommended change: Design L65 and Phase 1 L155: before the transaction, the script records the unread Arch news titles in its state file, using =yay -Pwq=, which needs no informant and works on both machines. It must run before =pacman -Syu=, because yay counts news as "new" relative to the build dates of installed packages. Only then does it clear the hook with =informant read --all= where informant is installed.
+
+Make the same =--all= correction to the boot ExecStart at L69, and to the parenthetical in AC L172. A bare =informant read= is interactive and hangs or fails with no tty.
+
+Phase 2 (L158): the deferred-set panel row also lists the recorded news titles until the user dismisses them.
+
+Risks: add one bullet. The driven path marks news read, so on velox informant's AbortOnFail hook no longer stops a run, and the panel row replaces it as the place news gets seen.
+
+Leave the auto-clear as designed. Whether unread news should instead stop the everyday run is an open call for the owner, not a review fix.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Design, Implementation phases, Risks, Testing. Partially confirmed; this follows the narrowing. The auto-clear stays, and whether unread news should stop a run stays my open call. The text also defines what the recommendation left open:
+- capture happens only in modes with network, never at boot;
+- titles carry forward until dismissed;
+- dismissal is a maint-owned key, so the script never reads panel state and stays the record's only writer.
+:EVIDENCE:
+Spec L65: "clears the news hook (=informant read=) where installed; runs =pacman -Syu --noconfirm --ignore=<blocked set>= ... then runs =topgrade --disable system,git_repos,containers -y=". Spec L69: the boot ExecStart runs "=informant read= (clear the news hook that would otherwise abort the transaction)". Spec L172 (AC): "Unread Arch news does not wedge the boot run". The spec never mentions surfacing news; the only hits for "news" are L65, L69, L172 and L183.
+
+archive/task-archive.org:604 says "informant: the base ships informant; its pacman PreTransaction hook (AbortOnFail) blocked archsetup's first transaction. Fix: informant read --all up front (guarded). PROVEN."
+
+archsetup:1121-1129 reads "registers a pacman PreTransaction hook (AbortOnFail) that blocks every package transaction while Arch news is unread ... --all marks without printing or prompting; a bare =informant read= is interactive and would hang an unattended run."
+
+~/.dotfiles/common/.config/topgrade.toml:144 sets show_arch_news = true, and :131 sets arch_package_manager = "yay". Both are under [linux], so they only take effect in the system step.
+
+=grep -rni 'news|informant' .dotfiles/maint/src/maint= finds no news probe (the only hit is the pacnew comment at gui.py:93).
+
+=pacman -Q informant= on ratio returns "error: package 'informant' was not found".
+
+~/.dotfiles/maint/src/maint/doctor.py:161-168: levers run via cmd.run(argv) and keep only =out[-1]= of stdout as the detail. cmd.py:18-24 shows cmd.run is subprocess.run(capture_output=True). So today's panel TOPGRADE path already drops topgrade's news output on ratio.
+
+=man yay=: "-w, --news Print new news ... News is considered new if it is newer than the build date of all native packages", and "-q Only show titles". That gives an informant-free title source on both machines, as long as it runs before the upgrade.
+:END:
+
+** DONE The panel's lever runner can't host the long --complete session
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L67 ('run in the foreground from the panel's action'); Phase 2 L158; Phase 3 L161 ('The panel action installs the held-kernel set live ... and offers to reboot'); Readiness Performance L182
+
+The spec runs the everyday split run and the dedicated kernel session through the maint lever path. That runner captures output, has no tty, and kills on a fixed timeout. It runs on a daemon thread that dies with the panel window. The spec specifies no timeout, detachment or log for the panel-hosted runs; the readiness items cover only the boot unit.
+
+Risk: A stray Escape, a waybar click, or the 3600 s ceiling kills the script partway through the kernel session, after the kernel transaction and before the gate verdict, the state-file write or the stamp. The panel reports 'tool missing or timed out' while pacman or DKMS may still be running orphaned. The 'offers to reboot' prompt has no stdin to read.
+
+Recommended change: Phase 3, L161: replace "The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot." with: "The panel action does not use the captured-output lever runner. Like MERGE, it opens a terminal running the dedicated session (foot -e upgrade-guarded --complete, via Popen with start_new_session=True), so closing the panel cannot kill or orphan the kernel transaction. The script prints the gate verdict there, writes the arm flag, and asks there whether to reboot. The panel picks up the outcome from the deferred-set state file on its next probe. The arm action's maint test asserts the terminal argv and the detach." Design L67: change "run in the foreground from the panel's action" to "run in a terminal the panel's action opens". Leave Phase 2 alone: the everyday levers already run through this runner with timeout 3600, and the argv change keeps that.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Design, Implementation phases, Testing. Partially confirmed; this follows the narrowing that the everyday levers keep the existing runner. foot gets --hold so the gate verdict stays readable after the script exits. The action lands in Phase 2 with its row (F16).
+:EVIDENCE:
+Spec:
+- L67: "run in the foreground from the panel's action or upgrade-guarded --complete in a terminal ... offers to reboot".
+- L161: "The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot ... the arm action's tests in maint".
+- L181-182: Observability and Performance cover only the boot unit.
+- L65: the script already writes the deferred set to a state file.
+
+maint (~/.dotfiles/maint/src/maint):
+- doctor.py:161: proc = cmd.run(step["argv"], timeout=r.get("timeout", 60)).
+- doctor.py:164-168: only the last stdout line (160 chars) and a 500-char stderr tail reach the wall; "tool missing or timed out" on None.
+- cmd.py:22: subprocess.run(cmd, capture_output=True, text=True, timeout=timeout), catching TimeoutExpired and returning None.
+- CPython subprocess.run source (checked live): except TimeoutExpired: process.kill(); process.wait(); raise. That kills the direct child only, and the Popen context exit closes the pipes.
+- remedies.py:294-313: update and topgrade are already kind "user", timeout 3600, so the everyday hazard is pre-existing.
+- gui.py:1452/1465: self.firing = True; threading.Thread(target=run, daemon=True).
+- gui.py:342-344: MaintPanel(Gtk.Application) with application_id and no hold(); gui.py:59 is GTK 4.0; no set_hide_on_close anywhere. So close() destroys the only ApplicationWindow and the app exits.
+- gui.py:358-361: activate on an already visible window calls close() (waybar on-click "maint-panel", waybar/config:84).
+- gui.py:1806-1810 + panel.py:246-254: Escape and q close the window. gui.py:421: the close button does too. No close path checks self.firing.
+- Detach precedent: gui.py:94-96 _MERGE_ARGV = ["foot","-e","sudo","pacdiff"]; gui.py:1672/1756 subprocess.Popen(..., start_new_session=True).
+- Kernel transaction weight: /var/log/ratio-upgrade.log:1487-1499 shows "(17/34) Install DKMS modules" (zfs for two kernels), then "(21/34) Updating linux initcpios".
+
+Adjacent fact, not part of this finding: doctor.py:169-170 stamps topgrade_run on any exit 0 when rid == "topgrade". If TOPGRADE's argv becomes the script (which exits 0 even when it deferred packages), that stamp contradicts the stamp-only-on-empty decision at L128 unless Phase 2 removes it.
+:END:
+
+** DONE Phase 3 install step isn't testable by the installer-steps harness, and unit enablement is unspecified
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Phase 3 L161; Phase 1 L155 ('the step that installs the guard'); Architecture L184 ('tested-by-inspection')
+
+Phase 3 says 'Installer step + unit file + the /var/lib/archsetup/ flag directory. Tests in tests/installer-steps/'. It names no function and no location for the unit source. Nothing says how the unit gets pulled into boot: there is no [Install]/WantedBy and no enable step.
+
+Risk: The 'tested-by-inspection' ordering claim has no harness to land in. A unit with no WantedBy never runs, so the manual test fails for reasons unrelated to the design.
+
+Recommended change: In Phase 3 (L161), after "non-fatal to every session target", add: "[Install] WantedBy=multi-user.target, enabled by the installer step (wanted, never required, so its failure can't fail the target)". Replace "Tests in =tests/installer-steps/= for the install step" with: "Tests in =tests/installer-steps/=: a source-inspection test in the style of test_pacman_hook_order.py asserting the unit text carries ConditionPathExists on the flag, Before=getty@tty1.service, an explicit TimeoutStartSec and WantedBy=multi-user.target, and that the installer enables the unit". At L184, change "tested-by-inspection" to "pinned by that installer-steps source-inspection test". Don't name the function and don't decide the network question in this edit. Separately, fix the stale phase numbers at L187 and L200 so the installer-step tests belong to Phase 3 and Phase 1's tests are the archsetup pytest beside the guard.
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Readiness dimensions, Testing. Partially confirmed; this follows the narrowing that the harness can source-inspect inline heredocs and no function is named. The assertion list grows to every directive the glossary settles. It replaces F32's separate static unit test, so there is one such test.
+:EVIDENCE:
+- Spec L69 lists ConditionPathExists, Before=getty@tty1.service, TimeoutStartSec and ExecStart. Spec L161 repeats them, adds "non-fatal to every session target", and says "Tests in tests/installer-steps/ for the install step". Neither line has an [Install] section or an enable step. Confirmed.
+- archsetup:2534-2655: the guard script copy and the 10-hypr-live-update-guard.hook heredoc are inline in hyprland(), next to pacman_install calls. Confirmed.
+- tests/installer-steps/test_pacman_hook_order.py:25-36 (written_hooks parses the hooks the source writes) and :42-45 (setUpClass reads the whole archsetup file) and :53-60 (asserts on the hook filename written inside hyprland()): a source-inspection test that already covers content inline in hyprland(). test_clone_user_repos.py, test_configure_service_discovery.py, test_install_camera_passthrough_rules.py and test_orchestrators.py also read the source text. This refutes "can't be extracted / no harness to land in".
+- More than 25 installer-steps tests use the sed -n '/^name() {/,/^}/p' extraction pattern, so convention already steers an implementer toward a top-level function. Spec L161 never mandates inline placement.
+- archsetup:2286-2299 (zfs-replicate.service: Type=oneshot, [Install] WantedBy=multi-user.target) and archsetup:2315 (installer runs systemctl enable on the unit set up in that block): an in-repo precedent for a oneshot unit written by heredoc and then enabled.
+- Live ratio: getty@.service has Before=getty.target. /etc/systemd/system/getty.target.wants/getty@tty1.service exists. multi-user.target Wants getty.target. So WantedBy=multi-user.target puts the oneshot in the same boot transaction as getty@tty1, and Before= takes effect.
+- Live ratio: NetworkManager-wait-online.service is enabled (matches the finding), but the network/pre-download question belongs to the separate package-source finding.
+- Spec L187 ("installer-step pytest for Phase 2; maint unit tests for Phases 1 and 3") and L200 ("Phase-2 install under tests/installer-steps/") contradict L155, L158 and L161 (Phase 1 = archsetup pytest beside the guard, Phase 2 = dotfiles maint, Phase 3 = installer step).
+:END:
+
+** DONE Documentation plan names no files, and current docs describe removed behaviour
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Readiness 'Documentation plan' L186; Phase 4 L164
+
+'a short "reboot to apply guarded upgrades" note in the maint docs; the installer step self-documents in-comment.'
+
+Risk: After Phase 2, the README and the TOML comment describe a force path that no longer exists and a stamping rule the decision reversed.
+
+Recommended change: Replace the Documentation plan bullet at L186 with: "Documentation plan: Phase 2 rewrites the 'The live-update guard' section of dotfiles maint/README.md. It drops press-again/--force, says UPDATE/TOPGRADE run upgrade-guarded (topgrade with --disable system,git_repos,containers), and describes the 'N deferred' row. The same phase updates the [updates] guard_patterns comment in archsetup configs/maintenance-thresholds.toml to say the patterns are the panel's display mirror of the hook, pinned by a test. Phase 4 adds the arm → reboot → console upgrade flow to that README section. The installer step self-documents in-comment." Add one clause to Phase 2 (L158): "and remove =maint fix --force=, which only served the guard override." Leave the guard banner and the wrapper paragraph alone in this finding. The wrapper-stamps-on-the-script's-own-topgrade-call conflict with L128 should be raised as its own correctness finding.
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Readiness dimensions. Partially confirmed; this follows the narrowing that the guard banner and the wrapper paragraph stay untouched. One detail is fixed: the README gives the space-separated --disable form (F02), because the recommended text quotes the comma form, which topgrade 17.12.2 rejects. The TOML comment edit happens with F21 in Phase 1, since both files are archsetup's.
+:EVIDENCE:
+Spec L158 (Phase 2): "the press-again-to-force sentinel wrap goes away". L155: topgrade runs as =--disable system,git_repos,containers -y=. L186: "Documentation plan: a short 'reboot to apply guarded upgrades' note in the maint docs". L164: "Document the flow".
+
+~/.dotfiles/maint/README.md:54-58:
+ "The panel arms with override wording (press again to force); the CLI equivalent is =--force= ... Topgrade always runs with =--disable git_repos=."
+Only README.md and src/ are present in ~/.dotfiles/maint, so the README is the only maint doc.
+
+configs/maintenance-thresholds.toml:63-64:
+ "a hit arms UPDATE/TOPGRADE instead of running ('press again to run anyway — or apply from a TTY')"
+
+Current code that Phase 2 removes:
+- doctor.py:216-242 (guard refusal, then force-wrapped sentinel)
+- remedies.py:308 (=topgrade --disable git_repos -y=)
+- cli.py:265-267 (--force help: "run a guarded remedy despite the live-update guard")
+
+Overstated parts:
+- Wrapper: the spec never changes it. Wrapper lines 40-42 stamp on rc 0, and README:60-64 matches that. ~/.local/bin/topgrade -> .dotfiles/hyprland/.local/bin/topgrade, and ~/.local/bin is first on PATH.
+- Banner: AC8 L174 says the hook is unchanged, and Non-goal L49 says the guard stays exactly as strict. The hook's Exec is /usr/local/bin/hypr-live-update-guard (archsetup:2648), so editing the banner edits what the hook runs.
+:END:
+
+** DONE Proof-of-concept 'exact sequence' did not hold kernels or zfs-dkms
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L65 (proof-of-concept sentence); history L213
+
+'Proof of concept: this exact sequence, run by hand on ratio on 2026-08-25 while Hyprland was live, resolved 724 of 730 pending packages ... with the six guard hits deferred.'
+
+Risk: This overstates the evidence: the kernel hold and the containers disable were never exercised. Under v1, the same day would defer 10 or more packages, not 6.
+
+Recommended change: In Design L65, replace "Proof of concept: this exact sequence, run by hand on ratio on 2026-08-25 while Hyprland was live, resolved 724 of 730 pending packages (Emacs 31.1 among them) with the six guard hits deferred" with "Proof of concept: the guard-set half of this sequence, run by hand on ratio on 2026-08-25 while Hyprland was live and before the kernel hold and the containers disable were added, resolved 724 of 730 pending packages (Emacs 31.1 among them) with the six guard hits deferred. That run upgraded linux, linux-lts and their headers live, so under v1 the same day would have deferred ten". Keep the trailing qemu-block-gluster clause as is. Leave history L213 unchanged.
+
+Optional, not blocking. Verification: confirmed.
+Disposition: modified, folded into Design. Confirmed. F04, which I accepted, puts zfs-dkms in v1's held set, and zfs-utils follows through the closure. Both moved 2.4.3 → 2.4.4 in that run, so the earlier 'drop the zfs-dkms clause' narrowing no longer holds. The count becomes 'at least twelve'; it is a lower bound because that day's closure can't be reconstructed.
+:EVIDENCE:
+- Spec L65: "Proof of concept: this exact sequence, run by hand on ratio on 2026-08-25 while Hyprland was live, resolved 724 of 730 pending packages (Emacs 31.1 among them) with the six guard hits deferred".
+- Spec L65 also defines that sequence as including "plus the *kernel set* ... held on every everyday run" and "topgrade --disable system,git_repos,containers -y".
+- Spec L135: the kernel set is held "on every everyday run, on both machines".
+- Spec L204: containers was added on 2026-10-05.
+- =/var/log/ratio-upgrade.log= L7-12: exactly six "ignoring package upgrade" lines (aquamarine, hyprland, hyprutils, mesa, vulkan-radeon, wayland).
+- =/var/log/ratio-upgrade.log= L16: "Packages (724)".
+- =/var/log/pacman.log= L20243: "upgraded linux (7.1.5.arch1-2 -> 7.1.9.arch1-2)".
+- =/var/log/pacman.log= L20261: "upgraded linux-headers".
+- =/var/log/pacman.log= L20262: "upgraded linux-lts (6.18.41-1 -> 6.18.46-1)".
+- =/var/log/pacman.log= L20263: "upgraded linux-lts-headers".
+- =/var/log/pacman.log= L20503-20504: zfs-utils and zfs-dkms 2.4.3 -> 2.4.4.
+- =/var/log/ratio-upgrade.log= L1488-1489: "dkms install --no-depmod zfs/2.4.4 -k 7.1.9-arch1-2" and "-k 6.18.46-1-lts".
+- =/var/log/ratio-upgrade.log= L1496 onward: the initramfs is rebuilt for the new kernels.
+- The log ends at the pacman post-hooks. =systemctl cat ratio-upgrade.service= returns "No files found" (it was a transient unit).
+:END:
+
+** DONE Summary and Design principle describe only the first draft
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Summary L28; Goals L43-46; Design opening L61, L63; Alternative D heading L90
+
+The Summary and Goals cover only the completion path. Design L61 says 'the ordinary full sweep stays a normal live topgrade run' and 'apply the guarded system upgrade with Hyprland down, once, and stamp it'. L63 says 'Three pieces, at two altitudes' but never lists them. D is labeled '(this spec)'.
+
+Risk: A reader who skims the Summary and the opening principle gets the superseded design before reaching the decisions.
+
+Recommended change: This is the smallest edit set; the first two items carry the substance.
+
+1. Summary L28: after "...record the freshness stamp", add: "The everyday path is a live split run (upgrade-guarded) that applies everything the guard would not block, holds the GPU/compositor and kernel sets, and reports what it deferred. Landing that deferred set is a dedicated, gated session that finishes at a boot-time oneshot."
+
+2. L61: replace "while the ordinary full sweep stays a normal live =topgrade= run" with "while everything else, the ecosystem sweep included, runs live through the split script, which defers the guarded and kernel sets".
+
+3. Goals: add one bullet: "An everyday UPDATE/TOPGRADE that succeeds live by deferring the guarded and kernel sets, and surfaces the deferred set durably."
+
+4. Optional: L63, drop "at two altitudes"; L90, relabel D "(this spec's completion step)".
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Summary, Goals and Non-Goals, Design, Alternatives. Partially confirmed; this follows the narrowing. The two optional edits are taken, and the wording uses the glossary's set names, so the DKMS set (F04) is named beside the kernel set.
+:EVIDENCE:
+- Spec L28 (Summary): "This spec designs a safe path to actually complete a guarded upgrade, and makes that completion record the freshness stamp". It does not mention the split live run.
+- Spec L43-46 (Goals): four bullets, all about completion, the boot path, or the installer. None covers the everyday live UPDATE.
+- Spec L61: "...while the ordinary full sweep stays a normal live =topgrade= run."
+- Spec L158 (Phase 2): "UPDATE and TOPGRADE levers change their =argv= to the script".
+- Spec L65: the script runs pacman -Syu --ignore=..., then yay -Sua, then topgrade --disable system,git_repos,containers -y.
+- git show 77447d0 has the identical Summary at draft L25, Goals at draft L40-43, and the L61 sentence at draft L58. Draft L60 is "Three pieces, at two altitudes." with nothing after it, but current L63 appends "— but the everyday gesture is not a reboot. It is a normal live update that simply leaves the dangerous few behind." So L63 has been edited and is not verbatim.
+- Current L65/L67/L69 are three bold-labelled paragraphs (split script, user, implementer).
+- Mitigations already in the spec: L13 (history line naming the live split upgrade as the everyday path), L51 and L55 (the split script named as v1), L167 (acceptance criterion for the live UPDATE).
+- Labels: L90 "D. ... (this spec)" vs L95 "E. ... (this spec's everyday path)".
+:END:
+
+** DONE ZFSBootMenu fallback leaves the separate pacman-db dataset out of sync
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Risks L196; Acceptance L169
+
+The spec treats booting the pre-pacman snapshot from ZFSBootMenu as a complete recovery for a failed kernel change on velox.
+
+Risk: After a fallback boot, /usr, /boot and the modules are back on the old kernel, but the pacman db still records the new kernel and zfs-dkms as installed. checkupdates then shows no pending kernel and the panel's deferred count drops it, so the machine silently runs a kernel the db doesn't know it has.
+
+Recommended change: Append one sentence to Risks L196, and carry it into the Phase 4 flow docs: "The pre-pacman snapshot covers zroot/ROOT/default only. The pacman db (zroot/var/lib/pacman) and the cache (zroot/var/cache) are separate datasets and stay current. So after booting the snapshot from ZFSBootMenu, re-register the old kernel set (and zfs-dkms, if that transaction moved it) with pacman -U --dbonly from /var/cache/pacman/pkg before the next upgrade-guarded run. Otherwise the db reports the new kernel as installed and the deferred row drops it." Do not offer "roll back zroot/var/lib/pacman to its matching snapshot" as a remedy: the hook never snapshots that dataset. A sturdier fix, outside this spec's scope, would have zfs-pre-snapshot snapshot both datasets atomically in one zfs snapshot call.
+
+Optional, not blocking. Verification: confirmed.
+Disposition: modified, folded into Risks, Implementation phases. Confirmed. zfs-utils is named beside zfs-dkms because F04 moves them as a pair.
+:EVIDENCE:
+- Spec L196: "the 05-zfs-snapshot hook snapshots it before every transaction, and ZFSBootMenu can boot that snapshot. The pacman cache also keeps the previous kernel and zfs-dkms for a downgrade." Spec L134: "It is survivable — ZFSBootMenu can boot the pre-pacman snapshot". Spec L169 requires only that the snapshot is bootable. Spec L197: "The panel's deferred row is the reminder". Nothing in the spec mentions the db (grep for snapshot, database, var/lib, downgrade).
+- scripts/zfs-pre-snapshot:11: DATASET="${ZFS_PRE_DATASET:-$POOL/ROOT/default}". Line 30: zfs snapshot "$DATASET@$SNAPSHOT_NAME", one dataset, no -r.
+- ~/code/archangel/installer/archangel:736-738: zfs create -o mountpoint=/var/cache ...; zfs create -o mountpoint=/var/lib -o canmount=off ...; zfs create -o mountpoint=/var/lib/pacman "$POOL_NAME/var/lib/pacman". Present since archangel's initial commit 2b691a0. archangel .ai/notes.org:107 also documents zroot/var/lib/pacman as the "Package database" dataset.
+- Velox was rebuilt from archangel on 2026-08-13/14 (.ai/sessions/2026-08-14-10-02-velox-mainboard-swap-reinstall-and-bringup.org:1079, "full reinstall via archangel+archsetup"). That makes the pre-reinstall claim at docs/design/2026-07-15-velox-boot-failure-handoff.org:21 and todo.org:1762 (no separate var/lib/pacman) stale.
+- /etc/pacman.conf:13: #DBPath = /var/lib/pacman/ (the default is in effect).
+- pacman -U --help lists "--dbonly only modify database entries, not package files".
+- Weak point in the original evidence: archsetup:2269 is sanoid config, not proof the dataset exists.
+:END:
+
+** DONE topgrade_run 'never written' is false on ratio
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Problem L32
+
+'Absent, the probe returns WARN ... the file is simply never written.'
+
+Risk: Cosmetic. The diagnosis still holds.
+
+Recommended change: Line 32: replace "the file is simply never written." with "the file is almost never refreshed: on ratio it was last written 2026-07-08, the day the wrapper landed, so the probe grades its age (88 days on 2026-10-05) against topgrade_warn_days = 14 and warns. The absent-file branch only fires on a machine that has never stamped." Line 34: change "It is never written because" to "It is rarely refreshed because". Line 36 (optional, same edit): change "The metric has read stale ever since." to "The metric had already gone stale in late July and stayed stale." Do not add a claim about velox, which cannot be checked while it is offline.
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Problem.
+:EVIDENCE:
+- Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:32: "...the file is simply never written."
+- Spec line 34: "It is never written because topgrade rarely exits 0..."
+- Spec line 36: "...so nothing stamped. The metric has read stale ever since."
+- On ratio, ls ~/.local/state/maint/ shows topgrade_run.json (69 bytes, dated 8 Jul 07:46). Its contents are {"written_at": 1783514783.122665, ...}. date -d @1783514783 gives Wed Jul 8 07:46:23 AM CDT 2026.
+- ~/.dotfiles/maint/src/maint/cache.py:35-42: get() returns (data, time.time() - written_at), or None only on OSError/ValueError/KeyError/TypeError. There is no TTL.
+- ~/.dotfiles/maint/src/maint/probes/updates.py:114-131: the None branch is the "no topgrade run recorded" WARN. Otherwise it computes days = age // 86400 and calls grade_high(days, topgrade_warn_days).
+- configs/maintenance-thresholds.toml:60 sets topgrade_warn_days = 14.
+- maint status (live) prints "warn topgrade_age Topgrade freshness: 88".
+- git log in ~/.dotfiles: commit 9b32a28 dated 2026-07-08 added hyprland/.local/bin/topgrade, the same day as the stamp.
+:END:
+
+** DONE The autologin claim holds only on velox
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Design L67, L69; Architecture L184
+
+The boot unit runs 'before the autologin shell starts Hyprland'.
+
+Risk: Low. Phase 4's 'confirm the ratio path matches' would find a password prompt, not autologin.
+
+Recommended change: Describe the tty1 login without depending on autologin, and keep the Before=getty@tty1.service ordering exactly as it is.
+
+- L67: replace "before the autologin shell starts Hyprland" with "before tty1's login shell can start Hyprland".
+- L69: replace "so it completes before autologin execs Hyprland" with "so it completes before the tty1 login (autologin where the installer enabled it, a password prompt otherwise), and so before ~/.profile.d/99-hyprland-autostart.sh can start Hyprland".
+- L184: replace "getty autologin ordering" with "getty@tty1 login ordering".
+- L192: replace "ahead of autologin" with "ahead of the tty1 login".
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Design, Readiness dimensions, Risks.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L67: "before the autologin shell starts Hyprland"
+- L69: "ordered Before=getty@tty1.service so it completes before autologin execs Hyprland"
+- L184: "getty autologin ordering"
+- L192: "the unit sits ahead of autologin"
+- L164: "Confirm the ratio path matches."
+
+Live on ratio (uname -n = ratio):
+- /etc/systemd/system/getty@tty1.service.d/ does not exist.
+- systemctl cat getty@tty1 shows the stock ExecStart=-/usr/bin/agetty --noreset --noclear - ${TERM}, with no --autologin.
+- /etc/systemd/system/getty.target.wants/getty@tty1.service is the stock getty@.service.
+- loginctl session 2 has TTY=tty1, Service=login, Type=wayland, Leader=1963.
+- Process chain: 1963 "login -- cjennings" → 2012 "-zsh" → 2033 /usr/bin/start-hyprland → 2038 Hyprland.
+- Hyprland is started by ~/.profile.d/99-hyprland-autostart.sh, which is guarded by [ "$XDG_VTNR" = "1" ] and runs start-hyprland. ~/.zshrc:10 sources ~/.profile, and ~/.profile loops over ~/.profile.d/*.sh.
+
+Installer archsetup configure_autologin (~L951-1004):
+- Writes getty@tty1.service.d/autologin.conf only when enable_autologin=true, or when the root is encrypted and the user answers yes.
+- A non-encrypted root returns 0 silently and writes no autologin.
+
+Velox autologin: could not be verified. The machine is offline, and no doc in the repo records it.
+:END:
+
+** DONE Strip 'pending' count and QUEUE don't distinguish held packages
+CLOSED: [2026-10-05 Mon 09:20]
+Where: Phase 2 L158; AC L167
+
+After a successful split run, the strip still counts the held set as pending, and QUEUE lists those packages without marking them.
+
+Risk: A successful UPDATE leaves 'N pending' on the strip next to 'N deferred', which reads as if the run failed. This is cosmetic, because updates_repo warns only at 50 or more.
+
+Recommended change: Append one sentence to Phase 2 (L158): "QUEUE (results wall and =maint queue=) tags any package in the deferred-set state file =[HELD]=, the strip's pending cell shows the held share (=N pending · M held=), and =cve_queued= excludes held packages so the CVE badge turns red only for advisories UPDATE can close. This is the 'run will defer' badge named in the script-ownership decision." No acceptance-criterion change is needed. Optionally extend L167 to read "...the panel shows the exact deferred set, marked held in QUEUE".
+
+Optional, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Acceptance criteria, Testing. Confirmed, with two changes. Its cve_queued clause is the same change as F36 and is carried there. And QUEUE's [HELD] tag is not Decision 3's 'run will defer' badge: that badge is F22's pre-run arm line, built from the TOML patterns, while [HELD] is a post-run mark built from the record.
+:EVIDENCE:
+- Spec L158 (Phase 2): the only panel change is a new probe plus the "N deferred — apply on reboot" row. L128: the deferred set is rendered "as its own state". L121: "the badge that says a run will defer", which no phase defines. L167 (AC): "the panel shows the exact deferred set".
+- ~/.dotfiles/maint/src/maint/doctor.py:309-316: =if rid in ("update", "topgrade") and swap_ok:= calls =p_updates.scan_net(th)= to refresh the pending cache after a fire.
+- probes/updates.py:150-158: scan_net writes checkupdates output to the updates_repo cache. --ignore'd packages are still listed by checkupdates.
+- probes/updates.py:53-67: the repo_updates count is len(updates_repo), graded against pending_warn. archsetup configs/maintenance-thresholds.toml:58 sets =pending_warn = 50=.
+- gui.py:447: the strip cell caption is ("updates_repo", "pending") with no held split.
+- viewmodel.py:479-483: queue row state is only "cve", "reboot" (from _needs_reboot: _REBOOT_EXACT includes mesa, wayland, linux, linux-lts; prefix nvidia) or "note". cli.py:189: tags = {"cve": "[CVE] ", "reboot": "[BOOT] ", "note": " "}. gui.py:468-486 (_on_show_queue) uses the same rows.
+- viewmodel.py:503-514 (cve_queued), whose docstring says the number turns red "only when running UPDATE would actually close an advisory". It counts every name in updates_repo/updates_aur, held ones included. It is used by gui.py:641-648 for both the faceplate badge and the strip.
+:END:
+
+** DONE F05: gate clause (b) checks kernels that stage 1 just removed, so a kernel and DKMS change in the same run always fails the gate :blocking:
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design 'The gate' L137 (b) and L140 (per-kver 'pkgbase exists'); Phase 1 gate list L2147 (b) and L2151; Phase 1 test L2197; Testing map L2489. Compare (c) at L138/L2148, which has the existence filter that (b) is missing.
+
+Clause (b) reads: 'if any DKMS-set package's version changed, every <kver> for which the pre-run dkms status showed a module installed'. Unlike (c), it has no 'that still have /usr/lib/modules/<kver>/vmlinuz' filter. The per-kver check then requires '/usr/lib/modules/<kver>/pkgbase exists'. The only test for (b) covers a DKMS-only change (L2197).
+
+Risk: Every --complete that moves a kernel together with zfs-dkms fails the gate on '<oldkver>: pkgbase', even when the new kernel and its modules are fine. The run exits 4 and nothing is armed. The row goes CRIT with 'do not reboot' and REBOOT is hidden. That is a false alarm on the exact path the gate exists for. A second --complete clears it only because (c) drops the vanished kver. The Phase 1 tests don't catch it.
+
+Recommended change: L137 and L2147: replace (b) with "(b) if any DKMS-set package's version changed, every <kver> for which the pre-run dkms status showed a module installed and that still has /usr/lib/modules/<kver>/vmlinuz after stage 1. A kernel whose version changed is covered by (a) under its new kver."
+
+L239 (Decision): change "or a DKMS-set package moved and that kernel had a module installed before the run" to "or a DKMS-set package moved and that kernel had a module installed before the run and is still installed after it".
+
+Add a test bullet after L2197 (Phase 1 Gate) and after L2489 (AC4 testing map): "one stage 1 moves a kernel-set package and zfs-dkms together. The old kver, which had zfs installed in the pre-run dkms status, is not on the check list. The new kver and every surviving kver with zfs installed are checked. With good fakes the gate passes."
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Decisions, Design, Testing. Both sheets fix this at the root instead of adding R01's existence filter to clause (b). The gate entry is keyed by pkgbase and resolved to the kvers present after stage 1. A kver the run removed then can't reach the gate by any path, on the first run or a re-check. The finding's test is adopted as written.
+:EVIDENCE:
+Spec L137 and L2147: "(b) if any DKMS-set package's version changed, every <kver> for which the pre-run dkms status showed a module installed". There is no existence filter. Compare (c) at L138 and L2148: "...that still have /usr/lib/modules/<kver>/vmlinuz".
+
+Spec L140 and L2151: the first per-kver requirement is "/usr/lib/modules/<kver>/pkgbase exists".
+
+Spec L130 and L2085: the dkms status is saved before stage 1. L2086: stage 1 lands the kernel set, the DKMS set and every other pending package except the GPU closure.
+
+Only the DKMS-only test exists: L2197 and L2489.
+
+Live output on ratio:
+- "pacman -Qqo" on /usr/lib/modules/7.2.7-arch1-1/pkgbase, /usr/lib/modules/7.2.7-arch1-1/vmlinuz, /usr/lib/modules/6.18.54-1-lts/pkgbase and /usr/lib/modules/6.18.54-1-lts/vmlinuz prints "linux linux linux-lts linux-lts". Each kernel package owns its kver's pkgbase, so an upgrade deletes it.
+- "dkms status" prints "zfs/2.4.4, 6.18.54-1-lts, x86_64: installed" and "zfs/2.4.4, 7.2.7-arch1-1, x86_64: installed".
+- "pacman -Q" shows linux 7.2.7.arch1-1 and linux-lts 6.18.54-1. The pkgrel is part of the kver.
+- /usr/share/libalpm/hooks/70-dkms-upgrade.hook removes the old modules PreTransaction (When = PreTransaction, Exec = dkms -D remove).
+
+Worked example: linux 7.2.7.arch1-1 moves to 7.2.8.arch1-1 and zfs-dkms moves in the same stage 1. (a) adds 7.2.8-arch1-1. (b) adds 7.2.7-arch1-1 and 6.18.54-1-lts. 7.2.7-arch1-1/pkgbase no longer exists, so the gate fails with "7.2.7-arch1-1: pkgbase" and exits 4.
+:END:
+
+** DONE F19: the required zfs.ko check runs lsinitcpio unprivileged against root-only 0600 initramfs images, so the gate fails on every ZFS-root --complete :blocking:
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L135, L140; Phase 1 L2027, L2144, L2154; sudo lists L80, L2027, Readiness Security L2385
+
+F19 made the lsinitcpio zfs.ko check required on a ZFS root (L140, L2154). kernel-modules-check is described as a stateless reader that --complete runs after stage 1. --complete runs as the invoking user and refuses EUID 0. Nothing gives the gate root: the sudo lists (L80, L2027, L2385) name only pacman -Sy/-Su/-S/-Sw, informant, the flag writes and systemctl reboot, and L2027 says ExecStartPre=+ is 'the only root step outside sudo'.
+
+Risk: On velox (ZFS root), the only machine where the gate matters most, every --complete with a non-empty check list fails the zfs.ko item. It exits 4 and writes an open gate failure; the row goes CRIT with REBOOT hidden, and nothing ever arms. Kernel and DKMS updates can never land through the designed path. The standalone 'no kver' TTY diagnostic fails the same way.
+
+Recommended change: Five edits:
+- L140 and L2154: replace "lsinitcpio of that image lists a zfs.ko" with "sudo lsinitcpio of that image (mkinitcpio writes images root 0600) lists a zfs.ko". Add: "if lsinitcpio itself fails, the item is 'cannot read <image>', not 'zfs.ko missing'". Add sudo to the gate's faked-on-PATH list on L140.
+- L80, L2027 and L2385: add "sudo lsinitcpio <image> in kernel-modules-check" to the sudo enumerations. The gate's standalone TTY diagnostic gets the same NOPASSWD sudo.
+- L2193 and L2480 (gate tests): add a case asserting the lsinitcpio call goes through the fake sudo, plus a case where an unreadable image fails as "cannot read", not as a missing zfs.ko.
+- L2412 (rollout step a): add "on a ZFS root, confirm kernel-modules-check <running kver> reads the real image without a permission error". Run it without --since, or with a since older than the image.
+- Optional alternative: have --complete run "sudo kernel-modules-check ..." instead. Sudo-ing only lsinitcpio inside the gate keeps the standalone diagnostic working the same way, so I'd do that.
+
+Blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Readiness dimensions, Testing.
+:EVIDENCE:
+- Spec L140 and L2154: "on a ZFS root (findmnt -no FSTYPE / is zfs), lsinitcpio of that image lists a zfs.ko". No sudo.
+- Spec L80, L2027 and L2385: the full sudo list is pacman -Sy/-Su/-S/-Sw, informant read --all, the flag writes and systemctl reboot. L2027 adds: "The boot unit's ExecStartPre=+ is the only root step outside sudo."
+- /usr/bin/mkinitcpio (pacman -Q: mkinitcpio 42.1-1) lines 1219-1220: "# Set umask to create initramfs images and unified kernel images as 600" / "umask 077".
+- ls -la /boot on ratio: initramfs-linux.img, initramfs-linux-lts.img, initramfs-linux-lts-strix.img and initramfs-linux-lts-fallback.img are all ".rw------- root". The vmlinuz-* files are 0644.
+- Running "lsinitcpio /boot/initramfs-linux.img" as uid=1000(cjennings) prints "==> ERROR: Unable to read file: '/boot/initramfs-linux.img'" and exits 1.
+- These work unprivileged: "dkms status" lists zfs/2.4.4 entries, and "stat -c %Y /boot/initramfs-linux.img" returns the mtime. Only the content read fails.
+- Spec L238: velox has /boot inside the encrypted ZFS root dataset. archsetup:3461 (tighten_efi_permissions) keeps boot images non-world-readable on purpose. A grep found no chmod or umask on initramfs anywhere in archsetup or scripts/.
+- Tests fake lsinitcpio (L2163, L2193, L2454, L2480). Rollout step (a) only confirms "upgrade-guarded --dry-run exits 0" (L2412), and --dry-run uses no sudo and does not run the gate (L80, L2041).
+- Velox was not checked directly because it is offline. The conclusion rests on the shared mkinitcpio umask.
+:END:
+
+** DONE Nothing marks a --complete kernel transaction as unverified before the gate runs, so in-progress, interrupted and aborted stage-1 runs can reboot velox with no initramfs :blocking:
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L128-133 (--complete order), L136 and Phase 1 L2146 (gate check list (a)), L168 and Phase 1 L2132 (interruption), Phase 1 L2083-2089, Phase 2 L2240 (REBOOT is state-driven), AC4 L2344
+
+Stage 1 runs =sudo pacman -Su --noconfirm --ignore=<GPU closure>=. The gate runs afterwards 'whatever stage 1's exit code', and its check list (a) holds only 'each kernel-set package whose installed version changed in this run'. The INT/TERM/HUP trap 'waits for the running child to exit, then writes result interrupted with the step and the remedy, and exits 130', so on that path no gate runs. The record's gate field is written only when a gate fails. REBOOT is hidden only while gate.ok is false, shown while the flag exists, and otherwise follows reboot_required. A stale arm flag is removed only after the gate or at the GPU step. A later --complete with 'nothing pending and no gate failure open' prints 'nothing to complete' and rewrites the record empty.
+
+Risk: On velox (ZFS root, one kernel), three paths reach an unbootable machine with no 'do not reboot':
+1. A normal stage 1: once 60-depmod removes the running kernel's module directory, the next 30 s panel refresh shows REBOOT. At that moment zfs is still building and /boot has no initramfs, because it was removed PreTransaction.
+2. Closing the APPLY window or logging out (SIGHUP), or a Ctrl+C that reaches the script: pacman stops, PostTransaction hooks are skipped, and /boot/vmlinuz-linux-lts and its initramfs stay deleted. The script exits 130 at WARN with no gate. The next --complete either wipes the record ('nothing to complete'), or, with GPU entries still pending, gets an empty check list (no version changed in this run), passes and arms. It then offers a reboot onto a kernel that was never gated.
+3. A commit abort, such as ENOSPC (the spec lists a full disk as a realistic trigger), before the kernel package extracts: the version is unchanged, so list (a) is empty, the gate is skipped, and the result is failed_step pacman at WARN while /boot is empty.
+
+Also, a flag left from an earlier arm keeps 'armed' and REBOOT visible through all of stage 1. Any of these leads to a ZFSBootMenu recovery.
+
+Recommended change: Smallest edit, in Design L128-139 and L168, Phase 1 L2087-2093, L2132 and L2146, Phase 2 L2240 and AC4 L2344:
+
+1. Stage 1 (L130 / L2087). When any kernel-set or DKMS-set member is pending, first =sudo rm -f= the arm flag. Then pre-write the record, in the same pattern as --apply-armed step 2:
+ - gate = {ok: false, pkgbases: [pkgbase of each pending kernel-set member; when a DKMS-set member is pending, every pkgbase the pre-run dkms status showed installed], failed: 'kernel transaction not yet verified', since: T0, at: now};
+ - add pkgbases to the gate schema at L108 / L2127.
+
+2. Check list (a) (L136 / L2146). Replace "whose installed version changed in this run" with "each pkgbase in the open gate entry (or pending at stage-1 start), resolved to the kver whose /usr/lib/modules/<kver>/pkgbase names it". Add to the per-kver requirements: /boot/vmlinuz-<pkgbase> exists. Re-check (c) uses the recorded pkgbases.
+
+3. Trap (L168 / L2132). In --complete, the INT/TERM/HUP trap leaves the open gate entry in place. Its detail reads '<step> interrupted — do not reboot; run upgrade-guarded --complete'. Only a gate pass clears it, which keeps 'nothing to complete' from firing and a follow-up run from arming without a re-check.
+
+4. Tests (Phase 1):
+ - during a fake stage 1, the record has gate.ok false and no flag;
+ - an INT or HUP during stage 1 exits 130 with the gate still open and any earlier flag gone;
+ - a stage-1 failure that leaves the kernel version unchanged but /boot/vmlinuz-<pkgbase> missing exits 4;
+ - a follow-up --complete after an interrupt re-checks and neither arms nor prompts until the check passes.
+
+5. AC4. Add: "From the start of a stage 1 that moves a kernel-set or DKMS-set package until a gate passes, including after an interrupted or failed stage 1, the deferred row is CRIT, REBOOT is hidden, no flag exists, and --complete never prompts for a reboot."
+
+Blocking. Verification: confirmed.
+Disposition: modified, folded into Design, Decisions, Implementation phases, Acceptance criteria, Testing, Risks. I adopted the pre-written unverified state keyed by pkgbases, the /boot/vmlinuz requirement, the trap rule, the CRIT and REBOOT rules, the tests and the AC4 outcome, with three changes. First, any flag is removed once, right after every successful --complete refresh (R48). That makes 'no flag while the gate is open' an invariant rather than a list of cases. Second, the entry is checked in one invocation with the recorded since, not in a separate re-check. Third, R03's rule as written raises a false CRIT when stage 1 fails before committing anything, as on a download, signature or conflict error. /boot is then intact, no snapshot was taken and the initramfs predates T0, so the since items fail. The row would say 'do not reboot' on a machine that boots fine, and keep saying it until a later --complete succeeded. So when this run opened the entry and stage 1 changed no kernel-set or DKMS-set version, the gate runs its structural form without --since. A missing vmlinuz or initramfs, an unreadable image or a missing zfs.ko still fails it, which covers R03's ENOSPC case.
+:EVIDENCE:
+Spec:
+- L130 / L2087-2089: stage 1 is a plain -Su with no prior state write.
+- L131 / L2090: "The gate, whatever stage 1's exit code".
+- L136 / L2146: "(a) for each kernel-set package whose installed version changed in this run".
+- L139: "An empty list skips the gate".
+- L129 / L2086: nothing pending and no gate failure open gives 'nothing to complete' and the record is rewritten empty.
+- L133 / L2099: the prompt fires when "the run armed or changed a kernel-set or DKMS-set package".
+- L168 / L2132: the trap "writes result interrupted ... exits 130", with no gate and no flag removal.
+- L2117: flag removal points are a gate failure, unit not enabled, or nothing GPU-side left.
+- L2240: REBOOT is hidden only while gate.ok is false and shown while the flag exists.
+- L2344-2348: AC4 covers only a failed DKMS build.
+- Phase 1 tests (L2183, L2196-2200) cover none of these cases.
+
+Live system:
+- /usr/share/libalpm/hooks/60-mkinitcpio-remove.hook: Type=Path, Operation=Remove, Target=usr/lib/modules/*/vmlinuz, When=PreTransaction. Its script's remove_kernel runs =rm -f= on the preset's ALL_kver and images (/etc/mkinitcpio.d/linux-lts.preset: ALL_kver=/boot/vmlinuz-linux-lts, default_image=/boot/initramfs-linux-lts.img).
+- man alpm-hooks CAVEATS: "PostTransaction hooks will not run if the transaction fails to complete for any reason."
+- 70-dkms-install and 90-mkinitcpio-install are PostTransaction. 60-depmod is PostTransaction and its script does =rmdir --ignore-fail-on-non-empty= on a kernel dir with no modules.order.
+- strings /usr/bin/pacman shows the "Interrupt signal received" and "Hangup signal received" fragments. libalpm.so.16 has "transaction interrupted", "transaction failed" and "problem occurred while upgrading %s". pacman is 7.1.0.
+- archsetup:2622-2624: "a blocked transaction can remove the current initramfs without reaching the PostTransaction hook that rebuilds it."
+
+maint:
+- probes/packages.py:131-146: required = not os.path.isdir(/usr/lib/modules/<release>).
+- panel.py:506-510: reboot_key_visible returns that value.
+- gui.py:92 _FULL_SECONDS=30. gui.py:1376-1384: _full_tick refreshes unless self.firing.
+- /usr/lib/modules/<running>/ holds unowned modules.alias and related files, so the old dir outlives the package until 60-depmod.
+:END:
+
+** DONE The gate's lsinitcpio check can't read the initramfs as the user, so the gate fails every time on a ZFS root :blocking:
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L80 (sudo step list) and L140 (gate: lsinitcpio of /boot/initramfs-<pkgbase>.img must list zfs.ko); Phase 1 L2027 (privilege list), L2154 (gate per-kver checks), L2163 and L2193 (tests fake lsinitcpio on PATH); Readiness Security L2385
+
+upgrade-guarded runs as the invoking user and refuses EUID 0. The steps that go through sudo are listed exhaustively: pacman -Sy/-Su/-S/-Sw, informant read --all, writing and removing the flag, and systemctl reboot. kernel-modules-check is not on that list, so --complete (and the standalone TTY diagnostic) runs it unprivileged. On a ZFS root it requires =lsinitcpio= of /boot/initramfs-<pkgbase>.img to list zfs.ko.
+
+Risk: On velox, every --complete whose check list is non-empty fails the gate at the zfs.ko item. That is every run that moves a kernel or zfs-dkms. It exits 4, records gate.ok false, and never arms. The CRIT row's remedy ('after fixing, run upgrade-guarded --complete') re-runs the same unreadable check, so the gate never closes, no form ever stamps, and the GPU/compositor set can only land by hand. The failure also blames a missing zfs.ko, which points me at a nonexistent DKMS problem on the do-not-reboot path. The standalone =kernel-modules-check= diagnostic fails the same way. This is the machine the gate exists to protect.
+
+Recommended change: Four edits to the spec.
+
+1. Phase 1 L2154 and the matching clause at Design L140: replace the zfs.ko bullet with: "on a ZFS root (findmnt -no FSTYPE / is zfs), sudo -n lsinitcpio of that image exits 0 and lists a zfs.ko with any compression suffix. The images are 0600 root (mkinitcpio sets umask 077), so the read goes through sudo. A non-zero lsinitcpio exit is its own failure item, '<kver>: cannot read /boot/initramfs-<pkgbase>.img', never reported as a missing zfs.ko."
+
+2. Add "lsinitcpio, inside kernel-modules-check" to the sudo step lists at L80, L2027 and Security L2385. Say that the standalone diagnostic uses the same sudo read.
+
+3. Phase 1 tests (L2193) and Testing (L2477-2480): add two gate cases.
+ - The fake lsinitcpio fails, as the real one does on a 0600 image, unless the fake sudo invoked it. The gate must read through sudo and pass.
+ - lsinitcpio exits 1 with no output. The gate fails with the "cannot read" item, not "zfs.ko missing".
+
+4. In the zfs-VM / velox manual check, add a step confirming that a passing --complete actually read the image (the gate output names zfs.ko as found).
+
+Blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Readiness dimensions, Testing.
+:EVIDENCE:
+Spec (docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org):
+- L80, L2027, L2385: the sudo lists leave out kernel-modules-check and lsinitcpio.
+- L131 and L2087: "run kernel-modules-check", with no sudo.
+- L140 and L2154: "lsinitcpio of that image lists a zfs.ko with any compression suffix".
+- L2163, L2193, L2454: lsinitcpio is faked on PATH.
+- A grep for 0600, "Unable to read" and "sudo lsinitcpio" in the spec finds nothing.
+
+Live system (ratio, uid 1000):
+- ls -l /boot shows all four initramfs-*.img as .rw------- root (linux, linux-lts, linux-lts-fallback, linux-lts-strix). vmlinuz-* are rw-r--r--.
+- lsinitcpio /boot/initramfs-linux-lts.img prints "==> ERROR: Unable to read file: '/boot/initramfs-linux-lts.img'" and exits 1.
+- /usr/bin/mkinitcpio:1219-1220 reads "# Set umask to create initramfs images and unified kernel images as 600" / "umask 077" (mkinitcpio 42.1-1).
+
+Repo knowledge of the velox case:
+- docs/workflows/system-health-check.org:238: "The sudo is load-bearing: the images are 0600, so an unprivileged lsinitcpio exits 1 with 'Unable to read file' ... looks exactly like a missing module (velox, 2026-09-12)."
+- todo.org:753-757: "Also for the kernel-modules-check gate: reading the initramfs needs root. ... The gate has to run as root and check the exit status, not just the count."
+:END:
+
+** DONE F03: closure parser counts every error line, but each failing probe also prints lines that match neither form, so the literal rule always refuses
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design 'The dependency closure' L96-L100 ('parses every error line, with or without the leading ::' and 'refuses when a line matches neither form'); Phase 1 Dependency closure L2047-L2050; Phase 1 test L2169; Testing AC2 map L2469-L2474
+
+The landed F03 text says the script 'parses every error line, with or without the leading ::'. Two shapes are classified (forward 'unable to satisfy dependency' and reverse 'installing X (v) breaks dependency'), and 'a line that matches neither form (a conflict, a missing target, a replace)' refuses with exit 3 and failed_step closure. The Phase 1 fake pacman emits only the two shapes.
+
+Risk: Under the literal contract, the error header alone matches neither form and refuses. The forward case's ':: ' skip-prompt lines also match neither form. So every probe that fails refuses, including the hyprutils/hyprlang and zfs-utils pin cases that AC2 and AC3 depend on, and the closure never grows. Every soname day becomes an exit-3 refusal. An implementer has to guess which lines to ignore. The Phase 1 fake emits only the two clean shapes, so the tests pass either way.
+
+Recommended change: L96-L100 and L2047-L2050: replace "parses every error line, with or without the leading ::" with this text:
+
+"reads the probe's stdout and parses each dependency line there, stripping an optional leading ':: '. Each line is classified by the two forms below. The probe's stderr preamble 'error: failed to prepare transaction (could not satisfy dependencies)' is expected on every failing iteration and is not classified. A preamble with any other reason in the parentheses refuses."
+
+Then change "a line matches neither form" to "a stdout dependency line matches neither form".
+
+L2169 and L2469: change the fake pacman so it replays pacman 7.1's verbatim failing -Sup transcript: the preamble on stderr, the "::" dependency lines on stdout, and exit 1. Add one assertion that a failing probe with the preamble and only well-formed lines adds names and does not refuse.
+
+Do not add the finding's ignore-list of "warning: ..." lines and skip-prompt lines. pacman suppresses all of them in print mode.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Design, Testing.
+:EVIDENCE:
+Spec text:
+- L96 and L2047: "parses every error line, with or without the leading ::".
+- L100 and L2050: refuses "when a line matches neither form".
+- L2169 and L2469: the fake pacman "emits both message shapes". No transcript, preamble or stream split is specified.
+- grep for "failed to prepare|neither form|skip the above|cannot resolve" finds only L96, L100 and L2050. The preamble is never mentioned.
+
+Installed pacman: 7.1.0.r9.g54d9411-2 (libalpm 16.0.1).
+
+pacman source, callback.c:424-440 (pacman 7.1.0 source): in print mode, cb_question answers INSTALL_IGNOREPKG and REPLACE_PKG with 1, answers every other question with 0, and returns before the REMOVE_PKGS branch at L493-512 that prints the "cannot be upgraded" and "Do you want to skip" lines.
+
+Empirical runs used a throwaway fixture db, with no system db and no sudo. The command was:
+LC_ALL=C pacman --config <fixture> --dbpath <fixture> -Sup --noconfirm --print-format '%n' ...
+
+- Forward case (--ignore=hyprutils, with hyprlang 2-1 depending on libhyprutils.so=13-64): exit 1.
+ - stdout: ":: unable to satisfy dependency 'libhyprutils.so=13-64' required by hyprlang"
+ - stderr: "error: failed to prepare transaction (could not satisfy dependencies)"
+ - No warning or skip-prompt lines, even though IgnorePkg=bridge-utils is pending and hyprutils is --ignore'd.
+- Reverse/orphan case: exit 1.
+ - stdout: ":: installing libbar (2-1) breaks dependency 'libbar.so=1-64' required by orphan"
+ - stderr: the same preamble.
+- Converged case: exit 0. stdout carries only the names; stderr is empty.
+:END:
+
+** DONE F11: the timeout 60 hardening bounds only the explicit informant call, and the 00-informant hook repeats the same unbounded fetch inside the transaction
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design everyday step 3 L85; boot form step 4 L160; Phase 1 everyday step 3 L2063 and --apply-armed step 4 L2102; Decision 6 consequences L250; AC7 L2357-2358; Risks 'Arch news' L2448 (F11 disposition L682 states the intent: 'timeout 60 bounds it, and the run carries on either way')
+
+Most of the F11 resolution is in the body and correct:
+- the call is timeout 60 sudo informant read --all, run only when command -v informant succeeds, with failure tolerated;
+- yay -Pwq captures the titles first;
+- the sudo/root reasons are stated;
+- the offline fallback is described;
+- AC7, the Readiness note (velox only) and the Phase 1 ordering test are all present.
+
+One gap remains. The body says the timeout 'bounds informant's feed fetch, which has no timeout of its own' and that 'failure, the timeout included, is tolerated' (L85, L2063). For --apply-armed it says 'the hook's informant check sees the same feed, so neither blocks the transaction' (L160, L2102). Risks L2448 says 'informant's AbortOnFail hook no longer stops a run'. Nothing covers what the hook does after the explicit call times out or fails.
+
+Risk: Take a slow or stalled archlinux.org fetch on velox while unread news exists. timeout 60 kills read --all before it saves, and the script carries on into the transaction. Then the 00-informant hook either:
+- stalls the transaction with pacman holding db.lck. At boot that lasts until TimeoutStartSec=20min, ending as interrupted. From the panel it lasts until the 3600 s runner timeout. This is exactly the hang the timeout was added to prevent, moved one step later.
+- or completes, sees one or more unread items, and exits non-zero, so AbortOnFail aborts the transaction.
+
+At boot the arm has already been consumed by ExecStartPre, so the GPU set doesn't land and the run records failed_step boot-transaction. On the everyday path it records failed_step pacman.
+
+On that path AC7 ('unread Arch news wedges no mode') fails, and L160/L2102 ('neither blocks the transaction') and L2448 ('no longer stops a run') are false. The failure is bounded and not unsafe: nothing is swapped, and pacman exits in the PREPARED state releasing db.lck.
+
+Recommended change: Smallest edit, option (b): name the residual and narrow the claims.
+
+1. L160 and L2102: replace "and the hook's informant check sees the same feed, so neither blocks the transaction" with: "When the fetch fails fast (no route or DNS), the hook's informant check also gets an empty feed and passes. When the clear times out or fails, nothing is marked read. Informant's hook then repeats the same unbounded fetch inside the transaction, so it can hold the transaction until TimeoutStartSec, or abort it on unread news (result failed, failed_step boot-transaction, arm consumed). The same abort can happen if the network comes up between the clear and the transaction."
+
+2. L85 and L2063: append: "A failed or timed-out clear leaves informant's hook live for this run's -Su, so the transaction can wait on the same fetch (bounded only by the lever runner's 3600 s) or abort on unread news (failed_step pacman, before any package changes)."
+
+3. L2448: change "no longer stops a run" to "no longer stops a run once the clear succeeds; a failed or timed-out clear leaves the hook able to stall or abort the transaction, and the record names it".
+
+4. AC7 (L2357): change to "Unread Arch news wedges no mode once the clear succeeds; a failed clear ends in a named failed_step, never a silent hang past the unit or runner timeout."
+
+5. Phase 1 tests (L2175 or L2499): add a case where a fake informant's read --all exits 124 and the fake pacman then fails as the hook abort would. Assert failed_step pacman (everyday) or boot-transaction (--apply-armed), no stamp, and for --apply-armed no re-arm.
+
+Option (a), if AC7 must hold unconditionally: when read --all exits non-zero, run that run's own -Su/-S with "--hookdir /etc/pacman.d/hooks --hookdir <per-run dir>". The per-run dir holds 00-informant.hook -> /dev/null, which keeps the guard hook active. In the everyday and --complete forms, do this only when yay -Pwq succeeded, so the news is never cleared unseen. Assert the extra --hookdir argv in the same new test.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Acceptance criteria, Risks, Testing.
+:EVIDENCE:
+Spec claims and gaps (docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org):
+- L85 and L2063: "The timeout bounds informant's feed fetch, which has no timeout of its own. Failure, the timeout included, is tolerated and logged."
+- L160: "offline-safe: informant 0.6.0 falls back to an empty feed and exits 0, and the hook's informant check sees the same feed, so neither blocks the transaction." L2102 makes the same claim.
+- L2357: "AC7: Unread Arch news wedges no mode." L2448: "on velox informant's AbortOnFail hook no longer stops a run".
+- L2175 and L2499 test only "a failing or absent informant is tolerated". There is no case where the clear fails and the hook then fires.
+- grep for hookdir, 00-informant and "informant check" finds no handling elsewhere.
+
+informant 0.6.0-2 (archangel airootfs, usr/lib/python3.14/site-packages/informant/):
+- feed.py:81-82: session.get(self.url) with no timeout. feed.py:83-86: on an exception it falls back to feedparser.parse(url), also with no timeout. feed.py:90-106: only a bozo URLError gives an empty feed.
+- informant.py:117-119: --all marks entries read in memory. informant.py:146: fs.save_datfile() runs after the fetch.
+- informant.py:165-167: an empty feed gives sys.exit() (0) with nothing marked.
+- informant.py:83-98: check marks and saves one unread item, but exits with the unread count either way.
+- usr/share/libalpm/hooks/00-informant.hook: Operation=Upgrade, Target=*, Target=!informant, When=PreTransaction, Exec=/usr/bin/informant check, AbortOnFail.
+
+Live system checks:
+- man alpm-hooks lists no timeout option, and says "Hooks may be disabled by overriding them with a symlink to /dev/null".
+- man pacman says --hookdir is "a alternative directory ... hooks in later directories taking precedence".
+- /etc/pacman.d/hooks holds hypr-live-update-guard.hook.
+- sysctl net.ipv4.tcp_syn_retries = 6.
+- archsetup:1801-1802 enables NetworkManager. The boot unit has no network ordering (spec L153).
+:END:
+
+** DONE F22: dropping guard: live_update also removes the only trigger for the arm-line annotation, _rearm_after_guard and the doctor review suffix that Phase 2 says to reword
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 2 Guard UX L2219-2220; Phase 2 tests L2262; Testing L2467
+
+L2219 drops guard: "live_update" from the update and topgrade remedies, along with _update_force and maint fix --force. L2220 then says guard.trips survives as the arm-line annotation, and that on a tripped read arm_line, _rearm_after_guard's text and the doctor review suffix read 'UPDATE armed — will defer at least <guard matches> ... — press again to run upgrade-guarded'.
+
+Risk: Implemented literally, the annotation never shows: guarded() is false, and no guard event is ever emitted. The implementer has to invent a new selector, such as a new remedy key or an rid check, or else keep rewording code that can no longer run. The L2262 test can pass against the pure arm_line function while the panel never calls it. The fixed text also says 'UPDATE armed' on the TOPGRADE key.
+
+Recommended change: In Phase 2 Guard UX, replace the second sentence of L2220 with this:
+
+"Dropping the tag removes the only selector, so panel.guarded() and doctor.py's review-suffix check key on rid in ('update', 'topgrade') instead, and test_panel_phase10.py:420-421 keeps asserting True for both. iter_fix no longer emits a 'guard' event, so delete _rearm_after_guard and the guard-event branch of _on_fired. Do not reword them. _guard_arm_line stops setting _update_force. On a tripped read, arm_line and the review suffix read '<label> armed — will defer at least <guard matches> (kernels and DKMS modules are always held) — press again to run upgrade-guarded'."
+
+Make matching edits elsewhere:
+- L2262 and L2467: change 'UPDATE armed' to '<label> armed'.
+- L2262: make the test drive the arm-press (_press_lever, or the strip-key handler) for both UPDATE and TOPGRADE with a tripped update_guard read, and assert the act line. Add one assertion that the review suffix carries the same annotation.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Testing. R07, R19 and R35 are one defect. I kept the rid selector (R07, R35) over R19's new arm_badge remedy key. I kept panel.guarded and rekeyed it rather than deleting it: _press_lever lives in gui.py, which needs GTK, so the arm press can only be tested GTK-free through panel.guarded and viewmodel.arm_line. The doctor review suffix is deleted rather than reworded (R35). The arm line is the display mirror Decision 3 names, and the roster's FIX reaches the same arm line. The label is generic (R07, R19) because TOPGRADE shares the path.
+:EVIDENCE:
+Spec:
+- L2219: "Drop =guard: "live_update"= from the update and topgrade remedies ... the GUI's =_update_force= override, and =maint fix --force=".
+- L2220: "=guard.trips= survives only as the arm-line annotation ... On a tripped read, =arm_line=, =_rearm_after_guard='s text and the doctor review suffix read 'UPDATE armed — will defer at least ...'".
+- L2262 and L2467: test the arm line text only.
+
+Code (dotfiles maint/src/maint), from a grep for '"guard"|guarded(|update_guard|live_update'. The only readers of the tag are:
+- remedies.py:298 and :309: the tag itself, on update and topgrade.
+- panel.py:228: guarded(rid) returns REMEDIES[rid].get("guard") == "live_update".
+- gui.py:1430-1433: only =if panel.guarded(rid):= reaches _guard_arm_line, which reaches viewmodel.arm_line(..., guard=tripped) via gui.py:1396-1405. Otherwise the plain arm_line runs.
+- gui.py:1496-1499: _rearm_after_guard runs only when a 'guard' event exists.
+- doctor.py:216-240: the 'guard' event is yielded only under =r.get("guard") == "live_update" and not force=.
+- doctor.py:415-421: the review suffix is gated on =r.get("guard") == "live_update"=.
+- gui.py:1525: _rearm_after_guard sets =self._update_force = True=.
+- viewmodel.py:99-107: arm_line formats {label}, not a fixed 'UPDATE'.
+
+Tests:
+- tests/maint/test_panel_phase10.py:420-421: assertTrue(panel.guarded("update")) and assertTrue(panel.guarded("topgrade")).
+- tests/maint/test_panel_levers.py:404: pins the live-apply wording.
+:END:
+
+** DONE F32: no test covers the everyday run's positive stamp or the open-gate no-stamp rule
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Stamp predicate L2137-2138 (Design L114-115); AC3 L2343; Phase 1 'Stamp and record' tests L2185-2191; Testing AC1/AC3 map L2465; gate tests L2195, L2486
+
+Every everyday stamp test is negative: a deferred set, an UPDATE-shaped run, a refresh, closure or pacman failure, a failing sweep (L2178), a failing yay (L2180). The only positive stamp tests are for --complete and --apply-armed (L2189, L2503, AC8). For an open gate failure, the tests only check that the gate field survives a later everyday run (L2195 'the gate field survives a following everyday run', L2486). None checks that the run doesn't stamp. AC3 (L2343) requires the everyday stamp 'only under the everyday rule ... and with no gate failure open', and the AC1/AC3 map (L2465) lists no case for either clause.
+
+Risk: Two broken implementations pass every listed test. One never stamps on the everyday path, which is the original permanently-stale defect. The other stamps fresh over an open gate failure on a CRIT 'do not reboot' day. AC3's stamp clause can't be signed off from the test map.
+
+Recommended change: This is the smallest edit. Phase 1 "Stamp and record", after L2187, add:
+- "a TOPGRADE-shaped everyday run (neither --no-topgrade nor --no-aur) whose refresh, -Su, yay and sweep fakes exit 0 and whose post-transaction pacman -Qu leaves nothing held writes topgrade_run; the same run with --no-aur doesn't stamp"
+
+At L2195, change "the gate field survives a following everyday run" to:
+- "a following everyday run that would otherwise stamp (TOPGRADE-shaped, every step exit 0, nothing deferred) carries the gate field forward unchanged and doesn't stamp"
+
+Mirror both in the Testing map:
+- At L2465 ("Phase 1, the stamp"), add the positive TOPGRADE-shaped case and the --no-aur case.
+- At L2486, apply the same "would otherwise stamp ... doesn't stamp" wording.
+
+Optionally, pin the L2190 and L2505 missing-maint case to that same positive everyday run, so the everyday path is shown to reach the stamp call.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Testing. Adopted except the '--no-aur doesn't stamp' variant, which goes away with R58.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L114-115 and L2137-2138: "No form stamps while gate.ok is false" and "everyday: stamps iff neither --no-topgrade nor --no-aur was given, the refresh, -Su, yay and topgrade all exited 0, and the record's packages list is empty".
+- L2343 (AC3): "It stamps only under the everyday rule ... and with no gate failure open."
+- L2185-2191 (Phase 1 "Stamp and record"): the only positive case is L2189 "after a successful --complete (stage 2 path) or --apply-armed ... the stamp is fresh". L2190 "a run that would stamp" names no mode.
+- L2195: "the gate field survives a following everyday run", with no stamp assertion. Same at L2486.
+- L2465 (AC1/AC3 map, "Phase 1, the stamp"): only the deferred-set, UPDATE-shaped and failure no-stamp cases.
+- L2503-2505 (AC8 map): the positive stamp and the missing-maint case, both completion forms.
+- L2253 and L2506 (Phase 2 / AC9): negative only.
+- Prior-round disposition (~L775): "No form stamps while a gate failure is open, so TOPGRADE can't mark a do-not-reboot machine fresh."
+- L124 and L2346: after a failed gate the new kernel is installed, not booted.
+
+Repo: tests/upgrade-guarded/ does not exist, so no test already covers this.
+:END:
+
+** DONE F29: the AUR wording says any held package trips the guard, but only GPU/compositor targets under a live Hyprland do
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design everyday step 7 L89; Phase 1 everyday step 7 L2067; Phase 2 Guard UX L2219 (contradicted by Risks L2447, Non-Goals L54, Risks L2446)
+
+L89 and L2067: 'yay still resolves repo dependencies normally, so an AUR upgrade that needs a newer held package makes yay's own pacman call trip the guard, and the AUR step fails.' L2219: 'on the driven path the guard fires only when an AUR upgrade makes yay's own pacman call reach a held package.' The everyday held set always contains the kernel-kind entries, and when Hyprland isn't live it contains only those (L75, L2041).
+
+Risk: F29's resolution existed to remove a false safety property, and the Design and Phase 1 text still state a narrower one. They imply that the AUR path can't move a held kernel or zfs-dkms. Risks says it can, and that path is the one that leaves velox unbootable. A reader who trusts the step 7 text would see no reason to guard the yay step.
+
+Recommended change: In L89 and L2067, replace the sentence "yay still resolves repo dependencies normally, so an AUR upgrade that needs a newer held package makes yay's own pacman call trip the guard, and the AUR step fails." with: "yay still resolves repo dependencies normally. An AUR upgrade that needs a newer held GPU/compositor package while Hyprland is live makes yay's own pacman call trip the guard, and the AUR step fails. One that needs a newer kernel-set or DKMS-set package installs it ungated (see Risks)."
+
+At L2219, changing "reach a held package" to "reach a held GPU/compositor package while Hyprland is live" is optional: the current wording is a true "only when" claim.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Design.
+:EVIDENCE:
+Spec text:
+- L89 and L2067 (identical): "yay still resolves repo dependencies normally, so an AUR upgrade that needs a newer held package makes yay's own pacman call trip the guard, and the AUR step fails."
+- L2041: "held set (everyday): ... It always contains the stage-A result, which is kernel-kind. Only when live does it add the GPU patterns..."
+- L54: "The hook stays silent on kernels."
+- L2447 (Risks): "yay's dependency resolution can move a kernel-set or DKMS-set package when an AUR upgrade requires a newer one; that path is ungated ... An AUR upgrade that needs a newer held GPU/compositor package trips the guard inside yay's own pacman call."
+- L2363 (AC10) is correctly scoped: "An AUR upgrade that needs a newer held GPU/compositor library under a live Hyprland fails the AUR step visibly."
+- L2219: "the guard fires only when an AUR upgrade makes yay's own pacman call reach a held package". This is a true necessary condition, not a defect.
+
+Live system:
+- =/etc/pacman.d/hooks/hypr-live-update-guard.hook:4-18= has Targets mesa, mesa-*, wayland, libdrm, libglvnd, hyprland, aquamarine, hyprutils, hyprgraphics, vulkan-radeon, vulkan-intel, vulkan-mesa-layers, nvidia-utils, lib32-nvidia-utils and xorg-xwayland. There is no linux* or zfs* target. On ratio the hook is still at the legacy unprefixed name; the spec's 10- rename is pending, and the Target content is the same.
+
+Code:
+- =scripts/hypr-live-update-guard:57-67= contains =hyprland_running || exit 0=, so the guard allows everything when Hyprland is not live.
+:END:
+
+** DONE F35: --complete's deferred-set definition still drops kernel-side entries that stage 1 failed to land
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design 'The sets' L77 and Phase 1 'Sets' L2043 (the deferred-set definition), against Decision 4 L227, Phase 1 record L2122, the --complete stage-1 failure path L2089 and APPLY visibility L2234
+
+The applied resolution text is present: L2122 and L227 say "--complete drops the kernel-side entries once stage 1 lands them", and L2120 says the record is rewritten "with an empty packages list when nothing is deferred". But the formal definition was never reconciled with it. L2043 (and L77) still says the deferred set is "the pending entries that remain after the run's last transaction ... and are in that run's held set (everyday) or GPU closure (--complete)", and that it "is exactly the record's packages list". So a --complete record can only ever hold GPU-closure entries.
+
+Risk: After a failed stage 1 (a conflict, a download error, a full disk), the kernel and DKMS entries that are still pending disappear from the record. On a day with nothing GPU-side pending, the row reads 0 deferred at WARN and APPLY is hidden. The retry can't be reached from the panel, and the row underreports the held kernel until the next everyday run recomputes stage A. An implementer also faces two contradictory rules (L2043 and L2122) and has to pick one.
+
+Recommended change: At Design L77 and Phase 1 L2043, change the --complete clause of the deferred-set definition from "or GPU closure (=--complete=)" to:
+
+"or, for =--complete=, its GPU closure (kind gpu) plus the still-pending members of the kernel set, the DKMS set and the previous record's kernel-kind entries (kind kernel), which is empty after a successful stage 1"
+
+Add one bullet to the Phase 1 tests: a --complete whose fake stage-1 =-Su= exits non-zero, with the gate passing or having nothing to check, ends with exit 1, result failed, failed_step pacman, and the kernel-set entries still in the record's packages with kind kernel. Add one bullet to the Phase 2 tests: from that fixture record, the deferred row shows N > 0 at WARN and offers APPLY.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Testing.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L77 (Design, the sets) and L2043 (Phase 1, Sets): "deferred set: the pending entries that remain after the run's last transaction ... and are in that run's held set (everyday) or GPU closure (--complete). ... It is exactly the record's packages list."
+- L227 (Decision 4) and L2122 (record, packages): "--complete drops the kernel-side entries once stage 1 lands them, and the GPU closure once it is applied". L2122 also says "A refusal before classification keeps the previous list", which covers refusals only.
+- L1585 (adopted resolution text): "rewrites the deferred-set state file ... with what is still deferred".
+- L131 and L2089: "Stage 1 failed while the gate passed or had nothing to check: failed_step pacman, exit 1, nothing armed." No packages rule.
+- L2224: "Value N: the record's packages minus the entries now installed at or above new".
+- L2234: "APPLY ... shown while N > 0 or a gate failure is open". L2236: APPLY on the topgrade_age row "While N > 0".
+- L2227-2229: WARN when result is failed; the row text adds detail.
+- L2085 and L132: stage 1 is =-Su --ignore=<GPU closure entries>=, and the GPU closure is refused if it contains a kernel-set or DKMS-set member, so on success no kernel-kind entry stays pending. That is why the extra clause is safe to apply unconditionally.
+:END:
+
+** DONE F40: REVIEW & FIX can still fire APPLY through the captured lever runner
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 2 APPLY L2234-2235; Phase 2 tests L2257; Testing AC4 L2490; deferred-row severities L2227-2228; CVE text L2243
+
+The resolution is present for the APPLY key. L2235 says APPLY opens =foot --hold -e upgrade-guarded --complete= through Popen(start_new_session=True) and never uses the captured runner. L2257 and L2490 test only that key's argv and detach. But L2234 makes APPLY "the lever of upgrade_deferred (always: True)", which is a REMEDIES entry. Nothing says how the REVIEW & FIX roster fires it.
+
+Risk: On exactly the CRIT gate-failure day, REVIEW & FIX lists APPLY, and its FIX runs =upgrade-guarded --complete= under iter_fix. Closing the panel (Escape, q, a waybar click) can then kill or SIGPIPE the script mid-kernel-transaction, between the DKMS and initcpio hooks. That is the velox failure F40 closed for the key path.
+
+Recommended change: Phase 2 APPLY, append to L2235: "APPLY is a confirm-tier remedy, so REVIEW & FIX lists it whenever the row is WARN or CRIT. The detached-terminal rule covers every panel entry point: the deferred row's key, the key on the topgrade_age row, and the roster's FIX. All three reach _on_lever/_press_lever, so dispatch on the rid there, and the panel never hands upgrade-guarded --complete to doctor.iter_fix." Phase 2 tests, extend L2257: "APPLY's argv and the detach, fired from the row key and from the REVIEW & FIX roster's FIX on a gate.ok-false fixture; in both cases Popen gets the foot argv with start_new_session=True and iter_fix is never called." Extend AC L2490 with the same roster clause.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Acceptance criteria, Testing.
+:EVIDENCE:
+Spec, docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L2234: "It is the lever of =upgrade_deferred= (always: True)".
+- L2235: "It doesn't use the captured-output lever runner. Like MERGE, it opens =foot --hold -e upgrade-guarded --complete= via =subprocess.Popen(..., start_new_session=True)=".
+- L2227-2228: the row is CRIT on gate.ok false and WARN on failed/refused/interrupted/advisories.
+- L2257 tests "APPLY's argv and the detach". L2490 is the key-path AC.
+- A grep for REVIEW & FIX/roster finds only the CVE exclusion lines (L2243, L2258, L2364, L2508). Nothing addresses APPLY in the roster.
+
+Code (~/.dotfiles/maint/src/maint):
+- remedies.py:23: "always marks free controls (CPU mode, charge limit, the update levers)".
+- remedies.py:534-546: attach_levers/levers_for derive levers from metric_ids.
+- update/topgrade (remedies.py:295-313) are tier "confirm" plus always: True, the template APPLY would follow.
+- doctor.py review() (~L389-424): =if r["tier"] != "confirm": continue= then yields every confirm remedy whose metric_ids hit WARN/CRIT.
+- viewmodel.py:182-199: review_rows gives every item-less remedy =_fix(ev["action"], "FIX", [])=.
+- gui.py:946-954: _digest_key kind "fix" -> self._on_lever. gui.py:781-785: card keys from panel.metric_keys -> self._on_lever. Both paths meet there.
+- gui.py:1407-1438: _press_lever -> _fire. gui.py:1468-1478: _fire -> doctor.iter_fix inside _stream's daemon thread (gui.py:1443-1460).
+- doctor.py:161: =cmd.run(step["argv"], timeout=r.get("timeout", 60))=.
+- cmd.py:18-23: subprocess.run(..., capture_output=True, timeout=timeout), which kills the child on TimeoutExpired.
+- gui.py:1665-1676: _on_merge is a separate Popen path, and MERGE has no REMEDIES entry.
+:END:
+
+** DONE F33: The rollback order leaves out the Phase 4 doc commits, so the health-check workflow points at a deleted binary
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Readiness 'Rollout, compatibility & rollback' L2418-2422; Implementation phases intro L2022; Phase 4 docs L2311-2313; Risks L2446
+
+The F33 text is present at L2404 and L2418-2422: not additive, revert the dotfiles Phase 2 commit, disable and remove the unit, then remove the flag and the scripts. But the list only covers Phase 2 and the machine state. L2022 adds one Phase 4 doc commit in each repo: the maint/README.md flow and recovery steps in dotfiles, and docs/workflows/system-health-check.org in archsetup. The F38 resolution then made the health-check workflow a third caller of upgrade-guarded (L2313, L2446).
+
+Risk: After a rollback that follows the spec exactly, the routine update in the health-check workflow tells the operator to run a binary that no longer exists. Its hand-stamp guidance also no longer matches the restored wrapper. The dotfiles README keeps describing APPLY and --complete, and the git revert of Phase 2 likely needs manual conflict resolution.
+
+Recommended change: Edit spec L2419-2422 as follows.
+
+Step 1 becomes: "1. Revert the Phase 4 doc commits (dotfiles =maint/README.md=, archsetup =docs/workflows/system-health-check.org=), then the dotfiles Phase 2 commit. This restores the yay/topgrade argv, the force path and the alias, removes the deferred probe and APPLY, and returns the health-check update step to plain topgrade. Revert the doc commit first because Phase 2 and Phase 4 edit the same README section."
+
+In L2422, change the second sentence to: "Removing the scripts before these reverts leaves UPDATE, TOPGRADE, the =sysupgrade= alias and the health-check workflow's update step pointing at a missing binary."
+
+Leave steps 2 and 3, and the "can stay" sentence, as they are.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Readiness dimensions.
+:EVIDENCE:
+- Spec L2022: "Phase 4 holds the docs, which land after Phase 3 as one doc commit in each repo (the =maint/README.md= flow and recovery steps in dotfiles, =docs/workflows/system-health-check.org= in archsetup)".
+- Spec L2313 (Phase 4 docs) and L2398 (Documentation plan): step 6 of system-health-check.org Phase 3 "runs =upgrade-guarded= instead of plain topgrade", and kernel, DKMS and GPU days go through --complete.
+- Spec L2446: "The hold covers only =upgrade-guarded='s callers: the panel levers, the =sysupgrade= alias and the system-health-check workflow's update step."
+- Spec L2419: "1. Revert the dotfiles Phase 2 commit." This is the first rollback step. No step reverts either Phase 4 doc commit.
+- Spec L2421: "3. Remove the arm flag, then =/usr/local/bin/upgrade-guarded= and =kernel-modules-check=."
+- Spec L2422: "Removing the scripts before reverting dotfiles leaves UPDATE and TOPGRADE pointing at a missing binary. The rewritten =guard_patterns=, the =10-= hook name and the record file can stay." The health-check update step is missing from both sentences.
+- Spec L2397: Phase 2 rewrites "the 'The live-update guard' section of dotfiles =maint/README.md=".
+- Spec L2311: Phase 4 "Add the flow to the 'The live-update guard' section of dotfiles =maint/README.md=". This is the same section, edited later.
+- ~/.dotfiles/maint/README.md:52 is "### The live-update guard". It is a short section of about 12 lines, so the two commits' hunks will sit next to each other.
+- docs/workflows/system-health-check.org:236 is today's step 6 ("Run =topgrade= for the actual update ..."), which Phase 4 rewrites. The file is tracked in archsetup (=git ls-files=), so a doc commit is the only thing that would undo the change.
+:END:
+
+** DONE F45: the recovery steps re-register only the kernel and zfs packages, but the snapshot boot reverts every package changed since the snapshot
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Risks 'Last-resort fallback on velox' L2444; Phase 4 Docs L2312; Readiness Documentation plan L2398; Testing recovery drill L2521
+
+All four places say: 'After booting the snapshot from ZFSBootMenu, re-register the old kernel set with pacman -U --dbonly from /var/cache/pacman/pkg before the next upgrade-guarded run, along with zfs-dkms and zfs-utils if that transaction moved them.' The resolution's text is present word for word. The problem is its scope: the pacman db (a separate dataset) also keeps every other package the reverted transactions changed, and these steps don't touch any of them.
+
+Risk: Someone who follows the documented recovery ends up with a db that records new versions, plus newly installed packages, for every non-kernel package those transactions moved, while the booted root holds the old files. pacman -Qu never lists them again, so neither upgrade-guarded nor the deferred row ever re-upgrades them. Later installs then resolve against library versions that aren't on disk. This is the same silent db/file divergence F45 was accepted to close; the resolution closes it only for the kernel and zfs packages.
+
+Recommended change: At L2444, L2312, L2398 and L2521, replace "re-register the old kernel set with pacman -U --dbonly from /var/cache/pacman/pkg ... along with zfs-dkms and zfs-utils if that transaction moved them" with:
+
+"reconcile the db with the booted root: for every package /var/log/pacman.log records as upgraded, downgraded or removed after the booted snapshot's creation time, re-register its old version with pacman -U --dbonly from /var/cache/pacman/pkg (yay's cache for an AUR package). For every package it records as installed after that time, drop it with pacman -Rdd --dbonly. Then confirm with pacman -Qkk. pacman.log lives on zroot/var/log, so it survives the snapshot boot. This covers the kernel set, zfs-dkms and zfs-utils, the non-kernel packages stage 1 landed, and any everyday runs made after the failed gate."
+
+Keep the existing sentence that rules out rolling back zroot/var/lib/pacman. In the L2521 drill, add one pending non-kernel package alongside the kernel update, so the drill actually exercises the reconcile step.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Risks, Readiness dimensions, Testing.
+:EVIDENCE:
+Spec text checked:
+- L2444 (Risks), L2312 (Phase 4 Docs), L2398 (Readiness) and L2521 (Testing drill) all say: "re-register the old kernel set with pacman -U --dbonly from /var/cache/pacman/pkg ... along with zfs-dkms and zfs-utils if that transaction moved them".
+- L2086: --complete stage 1 "lands the kernel set, the DKMS set and every other pending package except the GPU closure".
+- L2155: the gate needs a pre-pacman snapshot with creation >= since - 60 s.
+- L2057-2061: the everyday preconditions do not check =gate=.
+- L114 and L2137: an open gate failure only blocks stamping.
+- L2521: the drill pends only "a kernel update".
+- grep for pacman.log, Qkk and diverg finds no other recovery handling.
+
+Code and system checked:
+- scripts/zfs-pre-snapshot:11 is DATASET="${ZFS_PRE_DATASET:-$POOL/ROOT/default}", and :30 is a single =zfs snapshot= with no -r.
+- ~/code/archangel/installer/archangel:735 (var/log), :736 (var/cache) and :738 (var/lib/pacman) are separate datasets.
+- /etc/pacman.conf:14-15 leave CacheDir and LogFile at their defaults, so pacman.log lives on the var/log dataset.
+- archsetup:2180 configures paccache to keep 3 versions, so old package files are normally still in the cache.
+- /var/log/ratio-upgrade.log:16 is "Packages (724)", including glibc-2.44, gcc-libs-16.2.1, linux, linux-lts, zfs-dkms and zfs-utils.
+:END:
+
+** DONE --complete stamp rule is 'iff' / 'exactly when' in Phase 1 and AC8, but the nothing-to-complete path never stamps
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 1 stamp predicate L2139; AC8 L2360; Design L116, L129; Phase 1 --complete step 2 L2083; Decision 4 L229; Design L170
+
+Phase 1 L2139: --complete 'stamps iff it started with a non-empty record or an open gate failure, every step it ran succeeded, and the record ends empty'. AC8 L2360: topgrade_run 'is written exactly when' those three hold. But --complete step 2 (L2083, Design L129) says that when nothing is pending and no gate failure is open, it prints 'nothing to complete', rewrites the record empty, exits 0, 'No stamp.' Decision 4 L229 agrees: '...neither does one that finds nothing to complete.'
+
+Risk: Two implementers can disagree on whether freshness clears after a by-hand or out-of-band completion. Either way, one of the spec's own statements (the AC or the Decision) fails as written.
+
+Recommended change: Change L2139 to: "=--complete=: stamps iff it did not take step 2's nothing-to-complete exit, started with a non-empty record or an open gate failure, every step it ran succeeded, and the record ends empty (stage 2 applied the GPU closure, or there was none). An arming run never stamps, and the nothing-to-complete exit never stamps even when the record it started with was non-empty."
+
+Change AC8 at L2360 to: "...from a console with Hyprland stopped stamps freshness: =topgrade_run= is written exactly when the run got past the nothing-to-complete check, started with a non-empty record or an open gate failure, every step succeeded, and the record ends empty." This also drops "identically to the boot unit".
+
+Optionally append "never on the nothing-to-complete exit" to Design L116 for symmetry.
+
+Add a test next to L2504: a non-empty record whose packages are all already installed, nothing pending and no open gate failure. The run should print 'nothing to complete', rewrite the record empty, exit 0, and write no stamp.
+
+L229, L129 and L2083 need no change.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Acceptance criteria, Decisions, Testing.
+:EVIDENCE:
+- Spec L2139: "=--complete=: stamps iff it started with a non-empty record or an open gate failure, every step it ran succeeded, and the record ends empty".
+- Spec L2360 (AC8): "stamps freshness identically to the boot unit: =topgrade_run= is written exactly when the run started with a non-empty record or an open gate failure, every step succeeded, and the record ends empty."
+- Spec L2083: "If nothing is pending and no gate failure is open: print 'nothing to complete', remove any stale flag, rewrite the record (empty) and exit 0. No stamp."
+- Spec L129: the same text, ending "without stamping".
+- Spec L229 (Decision 4): "stamps only when it started with something to complete (a non-empty record or an open gate failure)... neither does one that finds nothing to complete". This is consistent with L2083.
+- Spec L116: "stamps only when..." This is a necessary condition and consistent.
+- Spec L170: "A by-hand sentinel-override upgrade outside the script doesn't stamp, though the panel's row still clears, because the probe drops entries already installed". So the record still lists the installed packages.
+- Spec L2446: "bare =topgrade=, =yay= and =pacman -Syu= still land the kernel and DKMS sets with no gate".
+- Spec L2140: "=--apply-armed=: stamps iff its transaction exited 0 or was a no-op, and the record ends empty". It has no started-non-empty clause, so it is not identical to --complete.
+- Spec L2504: "a no-op =--complete= (nothing pending, no open gate failure) doesn't stamp". There is no condition on the record, and no test pins the non-empty-record case.
+:END:
+
+** DONE Deferred-set definition drops kernel-kind entries that --complete stage 1 failed to land, contradicting 'drops the kernel-side entries once stage 1 lands them'
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design 'The sets' deferred set L77; Phase 1 deferred set L2043; Phase 1 record packages L2122; Decision 4 L227; --complete step 4 L2089; Design L103
+
+Deferred set (L77, L2043): 'the pending entries that remain after a run's last transaction and are in that run's held set (everyday) or GPU closure (--complete)... It is exactly the record's packages list.' Record packages (L2122, Decision 4 L227): '--complete drops the kernel-side entries once stage 1 lands them'. L103 adds 'A refusal before classification keeps the previous list', but 'classification' is never defined.
+
+Risk: After a failed stage 1, the panel row stops showing the still-held kernel and DKMS packages until the next everyday run re-adds them. Implementers will also disagree on what packages holds after a --complete refusal.
+
+Recommended change: Edit 1, the --complete deferred-set definition at L77 and L2043. After "or GPU closure (=--complete=)", add: "For =--complete=, the set also includes the pending entries outside the GPU closure that the previous record listed as kernel, or that are kernel-set or DKMS-set members. These are recorded as kind kernel. After a successful stage 1 this part is empty, so a failed stage 1 leaves the kernel side in the record."
+
+Edit 2, at L103 and L2122. Replace "A refusal before classification keeps the previous list." with "Every refusal (exit 3) keeps the previous list, since none runs a transaction."
+
+Edit 3, a new bullet in the Phase 1 tests (after L2201) and in Testing under "Phase 1, =--complete=" (after L2489): "a stage-1 =-Su= that fails before committing exits 1 with failed_step pacman and leaves the kernel-set and DKMS-set entries in the record's packages as kind kernel; a GPU closure that reaches a kernel member refuses and leaves packages unchanged."
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Testing.
+:EVIDENCE:
+Spec L77 and L2043 define the --complete deferred set as pending entries "in that run's ... GPU closure (--complete)" and say it "is exactly the record's packages list".
+
+Spec L129, L2053 and L2082: a GPU closure that contains a kernel-set or DKMS-set member refuses, so kernel and DKMS entries are never in the GPU closure.
+
+Spec L227 (Decision 4) and L2122: "--complete drops the kernel-side entries once stage 1 lands them".
+
+Spec L131 and L2089: "Stage 1 failed while the gate passed or had nothing to check: failed_step pacman, exit 1, nothing armed". Neither line says anything about packages.
+
+Spec L74 defines kernel-kind as "the stage-A closure result", and --complete has no stage A (L2051-2053).
+
+Spec L103 and L2122: "A refusal before classification keeps the previous list". grep -n -i classif finds only these two lines plus an unrelated use at L2379.
+
+Spec L2379: "Refusals happen before any transaction".
+
+Spec L2243-2244: cve_queued, the "fixable via UPDATE" caption, the UPDATE lever's cve_advisories and [HELD] "exclude names in the deferred set" and "read the record".
+
+Spec L2484-2489 and L2503 have no test for the packages list after a failed stage 1.
+
+The script doesn't exist yet: there is no upgrade-guarded or kernel-modules-check under scripts/. This is a pure contract inconsistency.
+:END:
+
+** DONE An open gate failure can never clear once its recorded kernel's vmlinuz is gone ('passes' vs 'had nothing to check')
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L108, L138, L140; Phase 1 L2089-L2090, L2127, L2137, L2148-L2149; Phase 1 test L2196; Testing L2488; Phase 2 CRIT row L2227
+
+The gate field is cleared only by 'a --complete whose re-check of those kernels passes' (L108, L2127), and Phase 1 step 4 sets gate to null only when the 'Gate passes' (L2090). L2089 treats 'passed' and 'had nothing to check' as separate outcomes. Check list (c) includes only recorded kernels 'that still have /usr/lib/modules/<kver>/vmlinuz' (L138, L2148). 'An empty list skips the gate' (L140, L2149). The test bullets say a recorded kernel whose modules directory (L2196) or vmlinuz (L2488) is gone 'drops out of the list'. Neither says the gate field then clears.
+
+Risk: The panel can be wedged at CRIT with freshness unstampable, with no supported way out short of editing the record by hand. Implementers must also choose whether a skipped gate clears an open failure.
+
+Recommended change: Edits to docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+
+1. L108, Design: after "...sets it to null;" add "a recorded kernel whose /usr/lib/modules/<kver>/vmlinuz is gone counts as re-checked and passed, so when none remain, a --complete whose own (a)/(b) check passed or was empty also sets it to null;".
+
+2. L2127, Phase 1 record field: append the same sentence.
+
+3. L2090, Phase 1 step 4: replace it with "Gate passes, or every recorded kernel dropped out of (c) and the (a)/(b) check passed or was empty: set =gate= to null."
+
+4. L2196: change "whose modules directory is gone" to "whose /usr/lib/modules/<kver>/vmlinuz is gone". Append "; when that leaves no recorded kernel and nothing else fails, the run sets =gate= to null and, with the record otherwise empty, stamps."
+
+5. L2488, Testing AC4: append "; when no recorded kernel remains, =--complete= sets =gate= to null".
+
+Leave the K-to-K' path as it is. Clearing when the new kernel passes is already correct.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Testing. With the entry keyed by pkgbase (R03), 'vmlinuz gone' becomes 'kernel package no longer installed'. That test is sturdier, because a module directory can outlive its vmlinuz. A recorded pkgbase drops out only when pacman reports the package not found, so a half-transacted kernel can't clear the entry by accident. When every recorded pkgbase drops out, the entry clears without invoking the gate, as the finding asks.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L108: "Only a =--complete= whose re-check of those kernels passes sets it to null".
+- L2127: "carries it forward unchanged until a =--complete= re-check of those kernels passes".
+- L138 and L2148, item (c): "the kernels of an open gate failure that still have =/usr/lib/modules/<kver>/vmlinuz=".
+- L140 and L2149: "An empty list skips the gate".
+- L2089: "Stage 1 failed while the gate passed or had nothing to check" (two separate outcomes).
+- L2090: "Gate passes: set =gate= to null".
+- L2193 and L2482: "an empty check list with no fresh snapshot passes" (supports the opposite reading).
+- L724, prior-round disposition: "Clearing requires a real re-check of the recorded kernels ... would otherwise pass vacuously".
+- L2137: "No form stamps while =gate.ok= is false".
+- L2227: the CRIT row with remedy "after fixing, run upgrade-guarded --complete".
+- L2240: REBOOT hidden while gate.ok is false.
+- L2196: "a recorded kernel whose modules directory is gone drops out of the list"; L2488 says vmlinuz instead.
+- L2224: probe value N filters packages but does nothing to the gate field.
+- Neither the L2196 test nor the L2488 test asserts that the gate field clears.
+
+Live system on ratio:
+- Kernels and owners from =pacman -Qqo /usr/lib/modules/*/vmlinuz=: 6.18.25-1-lts-strix is linux-lts-strix, 6.18.54-1-lts is linux-lts, 7.2.7-arch1-1 is linux.
+- =dkms status= shows zfs/2.4.4 installed for both 7.2.7-arch1-1 and 6.18.54-1-lts.
+- /usr/share/libalpm/scripts/depmod runs rmdir --ignore-fail-on-non-empty, so a modules directory can outlive its vmlinuz.
+- 71-dkms-remove.hook runs PreTransaction on Remove.
+:END:
+
+** DONE Phase 1 gate test bullet says 'only pre-T0 snapshots fail', contradicting the 60 s window it then states and the gate rule
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 1 tests L2193; Phase 1 gate L2155; Design L140; Decision 5 L239; Testing L2481; Risks L2444
+
+L2193: 'only pre-T0 snapshots fail; one created after T0 − 60 s passes'. The gate rule (L2155, Design L140) requires creation ≥ since − 60 s, and Testing L2481 says 'only snapshots created before since − 60 s fails'.
+
+Risk: A test written from L2193 would pin the opposite of the gate rule for exactly the case the 60 s window exists for, failing a correct gate or forcing a wrong one.
+
+Recommended change: At L2193, replace "only pre-T0 snapshots fail; one created after T0 − 60 s passes" with: "a listing whose newest root-dataset pre-pacman snapshot was created before T0 − 60 s fails, naming the dataset it looked for; one created at exactly T0 − 60 s, one created between T0 − 60 s and T0 (the snapshot hook's 60 s skip window), and one created after T0 each pass". Optionally, align L2481 the same way by adding the between-T0 − 60 s-and-T0 and exact-boundary cases, so that both test lists pin the ≥ comparison L2155 specifies.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Testing.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L2193: "snapshot listings: only pre-T0 snapshots fail; one created after T0 − 60 s passes"
+- L2155: "a =<rootds>@pre-pacman_*= snapshot with =creation= ≥ since − 60 s"
+- L140: "created no earlier than 60 s before since"
+- L2481: "only snapshots created before since − 60 s fails, naming the dataset it looked for, and one created after it passes"
+- L2444: "skipped within 60 s of the previous snapshot … which is why the gate requires one no more than 60 s older than the run's start"
+- L130 and L2085 set T0 = now before stage 1, and since = T0 (L108 and L2127, gate.since: <T0>).
+- L964 (prior review round's recommended test): "snapshot listings where only pre-T0 snapshots exist (gate fails) and one is newer than T0 (gate passes)". L2193 is the partially updated descendant of this line.
+
+Code: scripts/zfs-pre-snapshot:13 sets MIN_INTERVAL="${ZFS_PRE_MIN_INTERVAL:-60}", and :22 skips the snapshot when =now - last < MIN_INTERVAL=. This confirms that the snapshot covering a run can be up to 60 s older than T0.
+:END:
+
+** DONE APPLY is declared 'always: True' yet 'shown while N > 0 or a gate failure is open'; maint's lever model can't express either condition
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 2 APPLY L2234, L2236; Phase 2 deferred row L2229-L2230; AC9 L2361-L2362
+
+L2234: APPLY 'is the lever of upgrade_deferred (always: True), shown while N > 0 or a gate failure is open'. L2236: 'While N > 0, APPLY is also offered on the topgrade_age row in TOPGRADE's place.' L2229: a plain deferral (N > 0) grades OK.
+
+Risk: The implementer has to invent a conditional-lever mechanism, or ship APPLY permanently visible (or invisible in the common OK state), contradicting AC9 and the Phase 2 tests at L2255 and L2260.
+
+Recommended change: Keep "always: True" at L2234. It is required because L2229 grades N > 0 OK. Append: "The panel withholds the key through APPLY's panel.LEVER_KEYS items builder, which returns [] when the row's value is > 0 or the row is CRIT for a gate failure, and None otherwise (the _timer_items idiom). APPLY's metric_ids stay ['upgrade_deferred']."
+
+Replace L2236 with: "While N > 0, the topgrade_age card also shows an APPLY key. TOPGRADE is not a row key there today (it is absent from panel.LEVER_KEYS), so nothing is removed from the card, and it remains a strip key. The topgrade_age metric carries no deferred count, so topgrade_freshness adds an evidence row {deferred: N} read from the upgrade_deferred record, and APPLY's items builder reads it. In REVIEW & FIX, doctor.review lists APPLY instead of TOPGRADE for a WARN topgrade_age while that count is > 0."
+
+In L2260, L2362 and L2506, change "APPLY replaces/takes TOPGRADE's place on the topgrade_age row" to: "the topgrade_age card shows APPLY while N > 0 and not at N = 0, and doctor.review offers APPLY rather than TOPGRADE for a stale topgrade_age while N > 0".
+
+If a different channel is preferred (for example, metric_keys taking the envelope), name it in place of the evidence row.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Acceptance criteria, Testing. I kept the card key, as R18 recommends and as round 1's folded change asked ('the topgrade_age row offers that action instead of TOPGRADE'). Cutting it would partly undo that fold. One correction to the finding: attach_levers derives a metric's levers from metric_ids, so APPLY's metric_ids must include topgrade_age for that card to show the key. ['upgrade_deferred'] alone can't reach it. TOPGRADE was never a key on the topgrade_age card, because it is absent from panel.LEVER_KEYS. So 'in TOPGRADE's place' becomes 'the card also shows APPLY, and REVIEW & FIX swaps it in'.
+:EVIDENCE:
+Spec:
+- L2229: "OK otherwise, including N > 0".
+- L2234: "lever of upgrade_deferred (always: True), shown while N > 0 or a gate failure is open".
+- L2236: "While N > 0, APPLY is also offered on the topgrade_age row in TOPGRADE's place. TOPGRADE remains a strip key".
+- L2260, L2362 and L2506: "APPLY replaces/takes TOPGRADE's place on the topgrade_age row".
+- L2351 and L2514: "the deferred row then empties" (the row exists at N = 0).
+
+Code (dotfiles maint/src/maint/):
+- remedies.py:535-545, attach_levers: =levers = [rid for rid in levers_for(m["id"]) if off or REMEDIES[rid].get("always")]=. The always attachment is confirmed.
+- panel.py:155-160 (comment): "An items builder returns the remedy's item list from the metric's own evidence, or None to withhold the key (nothing to act on)".
+- panel.py:184-196, metric_keys: skips a rid not in LEVER_KEYS, or whose items_fn returns None.
+- panel.py:137-153: _timer_items and _battery_items return None conditionally, which is the precedent for the deferred row.
+- panel.py:161-181: LEVER_KEYS has no "topgrade" or "update" entry, so the topgrade_age card renders no TOPGRADE key.
+- gui.py:496-500: TOPGRADE is rendered only as a strip key.
+- doctor.py:389-425, review: selects remedies by =sev[mid] for mid in r["metric_ids"]= being WARN or CRIT. There is no per-metric condition and no swap.
+- probes/updates.py:114-131, topgrade_freshness: the metric carries days and cache_age only. No evidence, no deferred count.
+- indicator.py:42 and 56, and panel.py:474: lever-presence effects apply only to off-nominal rows, so APPLY attached at N = 0 (OK) has no side effect.
+:END:
+
+** DONE Removing the live_update guard tag silently disables the arm-line annotation the spec keeps, and _rearm_after_guard becomes unreachable superseded code
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 2 Guard UX L2219-L2220; Phase 2 test L2262; Testing L2467; Decision 3 L220
+
+L2219: 'Drop guard: "live_update" from the update and topgrade remedies, so iter_fix no longer refuses them.' L2220: 'guard.trips survives only as the arm-line annotation... On a tripped read, arm_line, _rearm_after_guard's text and the doctor review suffix read 'UPDATE armed — will defer at least <guard matches> ... — press again to run upgrade-guarded''. Test L2262 pins that arm line.
+
+Risk: If implemented literally, the display-side mirror that Decision 3 keeps never renders and the L2262 test can't pass. The implementer has to invent a new display-only marker and decide what happens to the dead re-arm path.
+
+Recommended change: In Phase 2 Guard UX, replace L2220 with:
+
+"guard.trips survives only as the pre-run arm-line badge that Decision 3 calls the display-side mirror. The update and topgrade remedies swap the guard tag for a display-only key, "arm_badge": True. panel.guarded() becomes panel.arm_badge(rid), reading that key, and gates both _guard_arm_line in _press_lever and the doctor review suffix (doctor.py:415). Nothing gates execution on it. On a tripped read, the arm line reads '<title> armed — will defer at least <guard matches> (kernels and DKMS modules are always held) — press again to run $ <argv>'. The review suffix reads '[will defer at least <guard matches> — kernels and DKMS modules are always held]'. Nothing says live apply or REBOOT required. With the refusal gone, nothing emits a "guard" event any more, so delete _rearm_after_guard, its call at gui.py:1496-1498, and the "guard" branches at panel.py:368 and cli.py:114, along with their tests."
+
+Then change the test at L2262, and the matching sentence at L2467, to: "an UPDATE press with a tripped guard.trips read, driven through _press_lever, shows 'UPDATE ... armed — will defer at least <guard matches> ...', and a TOPGRADE press shows its own label. A remedy without arm_badge never calls guard.trips."
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Testing. Resolved with R07: the rid selector instead of a new arm_badge remedy key, and the review suffix deleted rather than reworded. The deletions R19 lists are adopted: _rearm_after_guard and the guard branches.
+:EVIDENCE:
+Spec text:
+- L2219 says: Drop guard: "live_update" from the update and topgrade remedies, so iter_fix no longer refuses them.
+- L2220 says: On a tripped read, arm_line, _rearm_after_guard's text and the doctor review suffix read 'UPDATE armed — will defer at least ...'.
+- L2262 and L2467 pin that arm line.
+- L1090, from the prior review round: "panel.py:224-237: guarded()/update_guard() key on the same tag".
+- No body section names any replacement gate. grep for panel.guarded, update_guard and "display-only" finds only L1090 and L2394; L2394 says the TOML guard_patterns is display-only.
+
+Code (all paths under ~/.dotfiles/maint/src/maint):
+- remedies.py:298 and :309: update and topgrade both carry "guard": "live_update". No other remedy has the tag.
+- panel.py:224-228: def guarded(rid): return remedies.REMEDIES.get(rid, {}).get("guard") == "live_update"
+- gui.py:1430-1433: if panel.guarded(rid): self._act(self._guard_arm_line(title, argv, th)) else: self._act(viewmodel.arm_line(title, argv)). This is the only call site of _guard_arm_line (gui.py:1396-1405), which is the only caller passing guard= to viewmodel.arm_line (viewmodel.py:99-107).
+- doctor.py:415: if r.get("guard") == "live_update": wraps the review suffix built from guard.trips.
+- doctor.py:216-240: the only site that emits a "guard" event, gated by the same tag, which is also the refusal L2219 removes. Its consumers are gui.py:1496-1498 (calls _rearm_after_guard, defined at gui.py:1518-1535), panel.py:368 and cli.py:114. All of them become unreachable.
+- gui.py:1430-1431 serves both rids, so the literal 'UPDATE armed' would also appear on a TOPGRADE press.
+:END:
+
+** DONE Phase 1 tests need a kernel-set and DKMS fixture, but the named seams don't reach upgrade-guarded's /usr/lib/modules, /usr/src or db.lck reads
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 1 seams L2163; Testing L2453-L2454; Readiness L2401; kernel and DKMS sets L2038-L2039 (Design L71-L72); test L2167; db.lck precondition L2060; --apply-armed step 3 L2101
+
+upgrade-guarded's seams are UPGRADE_GUARDED_HOOK, _HYPR_RUNNING, _ARM_FLAG, _ARMED_LIST, _TOPGRADE, _MAINT, _UNIT_ENABLED and MAINT_STATE_DIR. KMC_MODULES_DIR and KMC_BOOT_DIR are 'Its [the gate's] test seams' (Design L140). The kernel set is derived from the owners of /usr/lib/modules/*/vmlinuz plus /usr/lib/modules/<kver>/pkgbase, and the DKMS set from /usr/src/*/dkms.conf.
+
+Risk: The implementer has to invent seams (or reuse the gate's KMC_* variables, which the spec assigns only to the gate), or the kernel-set tests end up host-dependent.
+
+Recommended change: 1. Add three seams to the seam list at L2163, and mirror them at L2401 and L2453:
+ - UPGRADE_GUARDED_MODULES_DIR (default /usr/lib/modules). Alternatively, state that upgrade-guarded honours KMC_MODULES_DIR and passes it through to the gate.
+ - UPGRADE_GUARDED_SRC_DIR (default /usr/src).
+ - UPGRADE_GUARDED_DB_LCK (default /var/lib/pacman/db.lck).
+2. Name the matching seam in parentheses after the paths at L71-L72 and L2038-L2039, and after db.lck at L83 and L2060.
+3. Add systemd-cat to the fake lists at L2163, L2401 and L2454.
+4. Add one bullet under "Steps and ordering", after L2182: "a file at UPGRADE_GUARDED_DB_LCK makes the everyday run, --complete and --apply-armed each exit 3 naming the file and the manual remedy, run no transaction, leave the file in place and write the record".
+5. Extend L2206 and L2496 to read "a missing, empty or unparseable list, a live Hyprland, a held lock or a present db.lck each exit 3".
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Testing, Readiness dimensions. This is the same gap as R29. upgrade-guarded reads its module tree through the gate's KMC_MODULES_DIR rather than a second UPGRADE_GUARDED_MODULES_DIR. One fixture tree then drives the kernel set, the pkgbase resolution and the gate. R20's db.lck and systemd-cat items are adopted through R29.
+:EVIDENCE:
+Spec text:
+- L140: "Its test seams are =KMC_MODULES_DIR= and =KMC_BOOT_DIR=" (said of kernel-modules-check only).
+- L2163, L2401 and L2453 list exactly UPGRADE_GUARDED_HOOK, HYPR_RUNNING, ARM_FLAG, ARMED_LIST, TOPGRADE, MAINT, UNIT_ENABLED, MAINT_STATE_DIR, plus "the gate's KMC_MODULES_DIR and KMC_BOOT_DIR".
+- Fakes (L2163, L2454): pacman, sudo, yay, informant, systemctl, vercmp, checkupdates, dkms, zfs, findmnt, lsinitcpio. systemd-cat is not among them.
+- L2488, under "Phase 1, --complete": "a recorded kernel whose /usr/lib/modules/<kver>/vmlinuz is gone drops out of the check list".
+- L2060 and L2099: the db.lck refusals. L2181-L2182, L2206 and L2496: precondition tests that cover the held upgrade-guarded.lock but never db.lck.
+
+Live system (ratio):
+- ls /usr/lib/modules lists 6.18.25-1-lts-strix, 6.18.54-1-lts and 7.2.7-arch1-1, with pkgbase files reading linux-lts-strix, linux-lts and linux.
+- /usr/src/zfs-2.4.4/dkms.conf exists.
+- pacman -Qqo on those paths returns linux-lts-strix, linux-lts, linux and zfs-dkms.
+- /usr/bin/systemd-cat is owned by systemd 262-1.
+- /var/lib/pacman/db.lck is absent.
+
+Repo idiom:
+- tests/hypr-live-update-guard/test_hypr_live_update_guard.py:46-52 and tests/zfs-pre-snapshot/test_zfs_pre_snapshot.py:49-57 give every path the script touches its own env seam (HYPR_GUARD_SENTINEL, ZFS_PRE_LOCKFILE, and so on).
+:END:
+
+** DONE The record's 'detail' is defined as one line but specified as multi-line pacman output
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L107 vs L91, L100; Phase 1 L2126 vs L2051, L2076; Phase 2 L2228
+
+'detail: the last stderr line, or the remedy text' (L107, L2126). Closure refusals put 'pacman's lines and the package names in detail' (L100, L2051). The armed closure refusal prints 'disarmed: closure refused' before the refusal's lines 'so detail keeps pacman's lines' (L91, L2076).
+
+Risk: The writer and the panel renderer may disagree on shape (string vs lines), and so may the shared fixture.
+
+Recommended change: Replace the detail bullet at L107 and again at L2126 with:
+
+"=detail=: a string naming the reason. Usually it is the last stderr line or the remedy text. On a closure refusal it is pacman's error lines followed by one line naming the refused packages, joined with newlines, and that names line is also the run's last stderr line. The row's WARN text uses only detail's last line, cut to 160 characters as doctor.py does for lever details. The row's evidence shows detail whole."
+
+At L91, L248, L2076 and L2382, change "so detail keeps pacman's lines" to "so the disarm line stays out of detail".
+
+With those edits, the definition, the test at L2473, the shared fixture and the panel renderer all agree on one shape.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Testing.
+:EVIDENCE:
+Spec text that conflicts:
+
+- L107 and L2126: "=detail=: the last stderr line, or the remedy text."
+- L100 and L2051: "Refusing means exit 3, result refused, failed_step closure, pacman's lines and the package names in detail".
+- L2473 (AC2 test): "orphan: ... refuses (exit 3, result refused, failed_step closure, pacman's lines and the names in detail)".
+- L91, L2076, L2382: "prints 'disarmed: closure refused' before the refusal's own lines, so detail keeps pacman's lines".
+- L2228: "WARN when result is failed, refused or interrupted (the text adds detail)".
+- L2225: "Evidence: ... the record's detail".
+- L233: "both test suites read copies of one fixture".
+
+Code: ~/.dotfiles/maint/src/maint/doctor.py:168 has =detail = out[-1].strip()[:160] if out else "done"=. That is maint's existing convention: detail is one line, capped at 160 characters.
+
+I found no other spec line, including the Review findings section around L1116-1123, that defines whether detail is a single line or a newline-joined block, or what the row renders from it.
+:END:
+
+** DONE The KEEP=10 prune can destroy the only old-kernel snapshot before the gated kernel first boots
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Risks L2444 (Last-resort fallback on velox); Decision 5 context L238 ('It is survivable — ZFSBootMenu can boot the pre-pacman snapshot'); gate snapshot rule L140 and Phase 1 L2155; armed everyday rewrite L2073; Decision 2 consequences L214; gate-failure window L2195
+
+The gate requires a pre-pacman snapshot created no earlier than since − 60 s, and Risks relies on ZFSBootMenu booting that snapshot 'with the old kernel, initramfs and modules' after a failed gate or a failure the gate misses. Nothing protects that snapshot afterwards. Meanwhile the design expects everyday runs to continue while armed ('the machine keeps working until the reboot is convenient'; every armed everyday run rewrites the flag) and during an open gate failure (the gate field survives everyday runs).
+
+Risk: Roughly ten transactions between stage 1 and the first boot of the new kernel prune the old-kernel snapshot. That is a few days of armed everyday runs, or installing build dependencies to fix a gate failure. After that, a failure the gate misses at first boot, or an unplanned reboot during a gate failure (battery, power), leaves velox recoverable only from live media. The spec presents that case as survivable through ZFSBootMenu.
+
+Recommended change: Smallest edit that actually protects the snapshot. It touches Phase 1, Risks and one test.
+
+1. Phase 1, --complete step 3. After the gate evaluates on a ZFS root with a non-empty check list, take the oldest <rootds>@pre-pacman_* snapshot whose creation is ≥ since − 60 s. That is the snapshot the gate accepted, holding the pre-stage-1 state. Run =sudo zfs hold upgrade-guarded <snap>= on it whether the gate passed or failed, and record its name in the record (for example a =held_snapshot= field; maint ignores unknown fields).
+
+2. Release. An everyday run or --complete runs =sudo zfs release upgrade-guarded <snap>= and clears the field in either case:
+- =uname -r= equals a kver that passed the gate, meaning a boot into the gated kernel happened;
+- a later --complete holds a newer snapshot.
+A missing snapshot is tolerated.
+
+3. scripts/zfs-pre-snapshot. Skip held snapshots when pruning, for example by listing =-o name,userrefs= and pruning only rows with userrefs 0. Otherwise the EBUSY destroy prints an error on every transaction while the hold is in place.
+
+4. Tests. Add a case in tests/upgrade-guarded/ where a fake zfs logs =hold= and =release=:
+- the hold is placed after a passing gate and after a failing gate;
+- it is released on the first run whose fake =uname -r= is the gated kver;
+- it is never released while gate.ok is false.
+Add a case in tests/zfs-pre-snapshot/ showing that a held snapshot is not destroyed.
+
+5. Risks L2444. Add one sentence: "zfs-pre-snapshot keeps only the newest 10 pre-pacman snapshots, so --complete holds the snapshot the gate accepted until a boot into the gated kernel is seen; without that hold, about ten later transactions (armed everyday runs, or fixes during an open gate failure) would prune the only snapshot with the old kernel."
+
+Documentation-only alternative, if the hold is out of scope: say in Risks L2444 and in the Phase 4 README recovery steps that the fallback lasts only about 10 snapshot-producing transactions (at least 60 s apart) after --complete. Then tell the user to reboot promptly after arming, and during an open gate failure to hold that snapshot by hand (=sudo zfs hold keep <snap>=) before running further everyday upgrades.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Acceptance criteria, Risks, Testing, Readiness dimensions. I took the hold, which is what actually protects the fallback, but dropped the uname-based release. The hold moves only when a later --complete holds a newer snapshot. So exactly one extra snapshot per machine stays pinned, and there is no boot-detection logic to specify or test. I kept a held_snapshot record field rather than scanning snapshots for the tag. The field gives the README recovery a named snapshot to boot, and lets rollback release it. Holding happens only after a --since-form gate. The structural form has no since to pick a snapshot by, and it runs only when no kernel-side version changed. Hold and release failures are tolerated and logged, because the hold protects the fallback rather than gating anything.
+:EVIDENCE:
+Spec:
+- L238: "It is survivable — ZFSBootMenu can boot the pre-pacman snapshot, which holds the old kernel, initramfs, and modules".
+- L2444: "an unplanned reboot after a failed gate, or a failure the gate misses, is recovered from ZFSBootMenu ... can boot the pre-pacman snapshot with the old kernel, initramfs and modules".
+- L214: "even then the machine keeps working until the reboot is convenient".
+- L2073-2077: step 9, the everyday run while armed.
+- L2195 / L2486: "the gate field survives a following everyday run".
+- L140 / L2155: "creation ≥ since − 60 s ... An older retained snapshot doesn't count". This is a lower bound only.
+- L2148: (c) re-checks "with that record's since".
+- Searching the spec for KEEP, retain, prune, "zfs hold" and release turns up only L988, where a prior finding's evidence lists ":14, :35-40 KEEP=10 prune". That finding's disposition (L1000-1006) never addressed retention.
+
+Code:
+- scripts/zfs-pre-snapshot:14 =KEEP="${ZFS_PRE_KEEP:-10}"=. Lines 35-40 list the snapshots sorted by creation, then =grep '@pre-pacman_' | head -n -"$KEEP"=, then =zfs destroy "$old"=. Lines 19-25 skip a snapshot within 60 s of the previous one.
+- archsetup:2421-2433: the 05-zfs-snapshot.hook uses Target = * and PreTransaction, with no AbortOnFail.
+- tests/zfs-pre-snapshot/fake-zfs handles only snapshot, destroy and list. Hold and release would need adding.
+
+Live system (ratio):
+- grep 'transaction started' /var/log/pacman.log gives 585 transactions. Per day: 6 on 2026-07-21, 5 on 07-22, 6 on 08-01, 5 on 09-29. That is roughly 10 within a week or two.
+- man zfs-hold: "If a hold exists on a snapshot, attempts to destroy that snapshot by using the zfs destroy command return EBUSY."
+:END:
+
+** DONE The ZFSBootMenu recovery re-registers only the kernel set, but the rollback reverts every package landed since the snapshot
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 4 Docs L2312; Risks L2444; Readiness Documentation plan L2398; recovery drill L2521
+
+'After booting the snapshot from ZFSBootMenu, re-register the old kernel set with pacman -U --dbonly from /var/cache/pacman/pkg before the next upgrade-guarded run, along with zfs-dkms and zfs-utils if that transaction moved them.'
+
+Risk: After the fallback boot, every package other than the kernel set that changed after the booted snapshot (stage 1's extras and the armed everyday runs) has old files on disk while the un-rolled-back db records the new version. pacman never re-extracts them, and -Qu and the deferred row don't list them. Later upgrades satisfy dependencies the filesystem doesn't actually have, such as a new soname. That silent db/filesystem split is what the documented steps leave behind.
+
+Recommended change: In Phase 4 Docs L2312, the Readiness documentation plan L2398 and Risks L2444, replace "re-register the old kernel set with =pacman -U --dbonly= from =/var/cache/pacman/pkg= before the next =upgrade-guarded= run, along with =zfs-dkms= and =zfs-utils= if that transaction moved them" with:
+
+"reconcile the db to the booted snapshot before the next =upgrade-guarded= run. Read =/var/log/pacman.log= (on =zroot/var/log=, which stays current), take every =[ALPM] upgraded=, =downgraded=, =installed= and =removed= line timestamped after the snapshot's =creation=, and undo them newest first:
+- upgraded or downgraded: =pacman -U --dbonly= the old version from =/var/cache/pacman/pkg=
+- installed: =pacman -Rdd --dbonly=
+- removed: =pacman -U --dbonly= the removed version
+
+Then =pacman -Qkk= over those names reports no mismatches. This covers the kernel and DKMS sets (so the next =--complete= lands them again behind the gate), every other package stage 1 landed, and anything later everyday runs landed."
+
+Keep the existing "Otherwise the db reports the new kernel as installed…" sentence.
+
+In the todo.org recovery drill at L2521, add one package outside the kernel set to stage 1. After the reconcile, confirm that package shows its old version in =pacman -Q= and that =pacman -Qkk= is clean for it.
+
+Optional, and a larger change: snapshot =<pool>/var/lib/pacman= atomically with the root dataset (one =zfs snapshot= call) so the two can be rolled back as a pair.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Risks, Readiness dimensions, Testing.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L130 / L2086: Stage 1 "lands the kernel set, the DKMS set and every other pending package except the GPU closure."
+- L2155: the gate requires the =<rootds>@pre-pacman_*= snapshot with creation ≥ since − 60 s, which is the snapshot taken before stage 1.
+- L2444: the fallback covers "an unplanned reboot after a failed gate, or a failure the gate misses". The recovery text re-registers "the old kernel set ... along with zfs-dkms and zfs-utils if that transaction moved them". L2312, L2398 and L2521 say the same.
+- L2191: "the gate field survives a following everyday run". L2073-2077 (everyday step 9): armed everyday runs also transact. Both land packages after stage 1.
+- L1912-1930: the prior finding addressed only the kernel and zfs-dkms. Its disposition adds zfs-utils and nothing else.
+
+Code:
+- ~/code/archangel/installer/archangel:720 creates ROOT/default. Lines 735, 736 and 738 create var/log, var/cache and var/lib/pacman as separate datasets under the pool, not under ROOT.
+- scripts/zfs-pre-snapshot:11 sets =DATASET="${ZFS_PRE_DATASET:-$POOL/ROOT/default}"=. Line 30 runs =zfs snapshot "$DATASET@$SNAPSHOT_NAME"=: a single dataset, no -r.
+- archsetup:2412 and 2421-2432 install the hook with =Exec = /usr/local/bin/zfs-pre-snapshot= and no env override.
+- ratio's /var/log/pacman.log has lines like "[2026-01-25T21:01:39-0600] [ALPM] upgraded gimp (3.0.6-2 -> 3.0.8-1)" and "[ALPM] installed iana-etc (20251215-1)". The log has the old version and a timestamp for every package, which is enough to reconcile the db to the snapshot.
+- Velox was offline, so I didn't check its dataset layout directly. I relied on archangel's installer and the prior finding's reinstall evidence.
+:END:
+
+** DONE Forward closure rule refuses when the unsatisfiable chain runs through a new dependency
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L97 ('unable to satisfy ... required by <X>' adds X; 'If X isn't in the pending set, the script refuses'); Phase 1 L2048; tests L2169 and L2469-2474 (AC2)
+
+Every 'unable to satisfy dependency '<d>' required by <X>' line adds X when X is pending. If X is not pending, the script refuses with exit 3 and failed_step closure.
+
+Risk: The refusal branch of the forward rule can fire only on a chain through a new dependency, and in that case the same iteration also names the pending ancestor (ptool) that should be held. On a release day when an unguarded pending package gains a new dependency built against a held soname, the everyday run and --complete refuse instead of holding the ancestor. They exit 3, name the new package rather than the culprit, and upgrade nothing, so AC2 fails. A fake pacman written to the spec's rule would lock this in.
+
+Recommended change: At L97 and L2048, replace "If X isn't in the pending set, the script refuses." with: "If X isn't in the pending set, X is a dependency the transaction would newly install (pacman prints this form only for transaction targets and the packages they pull in). The line adds nothing, because the pending package that pulls X in is named on its own line in the same output. If no line in an iteration adds a name, the existing no-progress rule refuses."
+
+At L2169 (Phase 1 tests) and L2470 (AC2 forward), add a fixture: "forward through a new dependency: the fake emits 'unable to satisfy dependency 'libhyprutils.so=13-64' required by newlib' and 'unable to satisfy dependency 'newlib' required by ptool', where newlib is neither pending nor installed. ptool is held with its stage's kind, the run exits 0 and the -Su runs."
+
+Leave L346 as it is, because it is the earlier review's historical text.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Testing.
+:EVIDENCE:
+Spec lines:
+- L97 and L2048: "unable to satisfy dependency '<d>' required by <X>" adds X; "If X isn't in the pending set, the script refuses."
+- L69/L2036: the pending set is the LC_ALL=C pacman -Qu rows.
+- L2338 (AC2): "the dependency closure holds that package too, and the run exits 0".
+- L2169 and L2470: the only forward fixtures are hyprutils and hyprlang, with no chain through a new dependency.
+
+First reproduction (pacman v7.1.0, fake db in scratch via --dbpath/--config):
+- Installed: hyprutils 1-1 (provides libhyprutils.so=12-64) and ptool 1-1.
+- Sync: hyprutils 2-1 (=13-64); ptool 2-1, which depends on newlib; newlib 1-1, which depends on libhyprutils.so=13-64.
+- pacman -Qu lists only hyprutils and ptool.
+- -Sup --noconfirm --print-format %n --ignore=hyprutils exits 1. Stdout: ":: unable to satisfy dependency 'libhyprutils.so=13-64' required by newlib" and ":: unable to satisfy dependency 'newlib' required by ptool".
+- Adding --ignore=ptool exits 0.
+- Adding --ignore=newlib instead still fails on "'newlib' required by ptool".
+
+Second reproduction (chain ptool -> n1 -> n2 -> held soname, plus qtool -> n2): a single probe prints lines naming n2, n1, ptool, n2 and qtool. Every pending ancestor shows up in the same iteration. --ignore=hyprutils --ignore=ptool --ignore=qtool exits 0 and leaves "other" to upgrade.
+
+Live system: hyprpaper (unguarded) depends on hyprwire and hyprtoolkit (expac -Q %N). hyprtoolkit depends on libhyprutils.so=13-64 and libaquamarine.so=14-64. hyprutils and aquamarine are both Targets in /etc/pacman.d/hooks/hypr-live-update-guard.hook.
+:END:
+
+** DONE The closure probe's output split (stdout detail vs stderr header) is unspecified, and the generic header matches neither form
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L96 ('parses every error line, with or without the leading ::') and L100 ('refuses when a line matches neither form'); Phase 1 L2047 and L2050; tests L2169 and L2469
+
+The script parses 'every error line' of the -Sup probe and refuses on any line that matches neither the 'unable to satisfy' form nor the 'breaks dependency' form.
+
+Risk: Read literally, the only stderr line, and the first line of the combined output, is the generic header, which matches neither form. Every non-zero probe would then refuse and the closure would never add anything. An implementer who captures only stderr as the 'error lines' never sees the dependency lines at all. Either way the soname-day behaviour in AC2 is lost. The Phase 1 fake pacman emits whatever split the implementer assumed, so the tests can't catch it.
+
+Recommended change: Replace the "parse every error line…" sentence at L96 and at L2047 with this output contract:
+
+"The probe's output contract (LC_ALL=C, stdout not a tty):
+- On exit 0, stdout is the %n target list.
+- On a non-zero exit, stderr carries one header, 'error: failed to prepare transaction (<reason>)', and stdout carries the ':: '-prefixed detail lines.
+- The script classifies the stdout ':: ' lines against the two forms below, and accepts the header only when <reason> is 'could not satisfy dependencies'."
+
+Change L100 and L2050 to read: "It also refuses when the header has any other reason, when any other stderr line appears, or when a ':: ' line matches neither form, …". Keep the rest of the sentence as it is.
+
+At L2169 and in the AC2 bullets at L2469-2475, add: "the fake pacman reproduces the split: the detail lines on stdout and the 'could not satisfy dependencies' header on stderr. A header with any other reason refuses, and so does a stray stderr line."
+
+Don't name conflicts as an example of a header the probe will report. In -p mode pacman suppresses warnings, and its conflict checking under print mode hasn't been verified here.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Design, Testing.
+:EVIDENCE:
+I reproduced this read-only on ratio against the existing checkupdates db (pacman v7.1.0, 397 pending), running each command under bash with the streams captured separately.
+
+Forward case: LC_ALL=C pacman -Sup --dbpath /tmp/checkup-db-1000 --noconfirm --print-format %n --ignore=gdb-common
+- exit 1
+- STDOUT, exactly one line: ":: unable to satisfy dependency 'gdb-common=18.1' required by gdb"
+- STDERR: "error: failed to prepare transaction (could not satisfy dependencies)"
+
+Reverse case: the same command with --ignore=gdb
+- exit 1
+- STDOUT: ":: installing gdb-common (18.1-1) breaks dependency 'gdb-common=17.2' required by gdb"
+- STDERR: the same "error: failed to prepare transaction (could not satisfy dependencies)" header
+
+Other observations:
+- On failure, stdout carries no %n target list, only the ":: " lines.
+- cat -A showed no colour escapes even though /etc/pacman.conf:33 sets Color, because stdout is not a tty.
+- --ignore=aws-c-common, aws-c-io, aws-crt-cpp and harfbuzz each exited 0 (the success path).
+
+Spec references:
+- L96 and L2047: "parse every error line, with or without the =::= prefix"
+- L100 and L2050: refuse "when a line matches neither form (conflict, missing target, replace)"
+- L2169 and L2469-2475: the fake pacman only "emits both message shapes"
+- L346: the prior finding defines the two forms but no stream or header contract
+
+No implementation exists yet (no upgrade-guarded in scripts/ or tests/), so the spec is the only contract.
+:END:
+
+** DONE pacman exit 1 means both 'empty' and 'error' for -Qu and -Qqo, and the empty DKMS-set case is undefined
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L69 and Phase 1 L2036 (-Qu 'Exit 1 means nothing is pending, not a failure'); deferred set from post-transaction -Qu at L77 and L2043; DKMS set at L72 and L2039 and kernel set at L71 and L2038 (pacman -Qqo of globs)
+
+Exit 1 from pacman -Qu is read as an empty pending set. The DKMS and kernel sets are 'the owners of /usr/src/*/dkms.conf' and '/usr/lib/modules/*/vmlinuz' via =pacman -Qqo=, and the spec doesn't say what a no-match or unowned path does.
+
+Risk: A post-transaction -Qu that errors would be read as 'nothing deferred'. The record would then be empty, and a TOPGRADE run could stamp fresh while packages are still held. A Hyprland machine with no DKMS package gets upgrade-guarded from the installer, and there deriving the DKMS set errors; depending on set -e handling the run aborts or proceeds with an undefined set.
+
+Recommended change: Three spec edits:
+
+1. L69 and L2036: replace "Exit 1 means nothing is pending, not a failure" (L2036's wording is "Exit 1 means empty, not failure") with: "Exit 1 with empty stdout and no =error:= line on stderr means nothing is pending (warnings are tolerated); any other non-zero exit is a failure. Before the transaction it is failed_step refresh, exit 1, with no transaction. After the transaction it is failed_step pacman, exit 1: the record keeps the pre-transaction pending ∩ held set as packages, and nothing stamps."
+
+2. L71/L72 and L2038/L2039: add "Only existing paths are passed to =pacman -Qqo= (nullglob). No match is an empty set, not a failure, and a vmlinuz with no owner is left out like a foreign kernel."
+
+3. Testing (L2165-2172): add two bullets:
+ - "with no /usr/src/*/dkms.conf the DKMS set is empty and the everyday run proceeds";
+ - "a fake pacman -Qu that prints an error: line and exits 1 fails the step, and the record does not empty or stamp."
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Testing. Adopted, plus one case the finding exposes but doesn't name. kernel-modules-check runs dkms, and a Hyprland machine with no DKMS module (a btrfs install without NVIDIA) has no dkms binary. 'Command not found' would then fail every kernel gate there. So the module item passes when dkms isn't installed. A KMC_DKMS seam, where empty means absent, lets the tests exercise that on machines that do have dkms.
+:EVIDENCE:
+Spec lines:
+- L69 and L2036: "Exit 1 means nothing is pending, not a failure" / "Exit 1 means empty, not failure".
+- L77 and L2043: the deferred set comes from a post-transaction pacman -Qu, and it is the record's packages list.
+- L115: the everyday run stamps when "the record's packages list is empty".
+- L129: --complete with nothing pending "rewrites the record empty and exits 0".
+- L72 and L2039: DKMS set = owners of /usr/src/*/dkms.conf (pacman -Qqo).
+- L2165-2172: the Testing list has no case for a -Qu error or an empty DKMS set.
+- L2524: rollout step (a) runs --dry-run only on the two daily drivers.
+
+Commands run on ratio:
+- LC_ALL=C pacman -Qu --dbpath /nonexistent-db prints "error: failed to resolve path ..." and exits 1.
+- The real LC_ALL=C pacman -Qu prints nothing to stdout or stderr and exits 1, so exit 1 alone can't tell the two apart.
+- pacman --config <copy of pacman.conf plus a repo with no db> -Qu leaves stdout empty, prints "warning: database file for 'nonexistentrepo' does not exist" to stderr, and exits 1. The "stderr empty" test would wrongly fail this empty case.
+- bash -c 'pacman -Qqo /usr/src/*/nonexist.conf' prints "error: No package owns /usr/src/*/nonexist.conf" and exits 1.
+- pacman -Qqo /usr/lib/modules/*/vmlinuz prints linux-lts-strix, linux-lts and linux and exits 0. No unowned vmlinuz exists today.
+- pacman -Qqo /usr/src/*/dkms.conf prints zfs-dkms and exits 0 on ratio.
+
+Code:
+- archsetup:2614-2653 installs the guard inside hyprland() on every Hyprland install.
+- grep zfs-dkms in archsetup finds nothing; it isn't installed by the installer.
+- scripts/testing/lib/vm-utils.sh:23 sets FS_PROFILE="${FS_PROFILE:-btrfs}".
+:END:
+
+** DONE The deferred row's state with no record on disk is undefined, and every machine reaches Phase 2 in that state
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 2 'The deferred row' L2222-2230; Phase 1 --dry-run L2033; Rollout step (a) L2322/L2411
+
+Phase 2 defines N, evidence and severity only from fields of an existing record (L2223-2229). Nothing says what upgrade_deferred reports when ~/.local/state/maint/upgrade_deferred.json is absent or unreadable. Rollout step (a) runs only 'upgrade-guarded --dry-run', which 'writes nothing (no record, stamp or flag)' (L2033), so every existing machine pulls the Phase 2 commit (step b) with no record. Every fresh rebuild is in the same state until its first run.
+
+Risk: The implementer has to invent the first state every machine passes through after rollout (b), including the row's severity, whether APPLY shows, and whether the strip reddens. The two machines could also end up on different readings if they're handled ad hoc.
+
+Recommended change: In Phase 2 'The deferred row' (after L2229), add one bullet:
+
+"No record (cache.get returns None: never written yet, since --dry-run writes none): OK, value 0, text '0 deferred', no APPLY. Nothing has run, so nothing is deferred; unlike topgrade_age, this row doesn't treat unknown as stale. A record whose data lacks the documented fields reads UNPROBED 'deferred-set record malformed — rerun upgrade-guarded' (the _cached idiom), with no APPLY, and never crashes the envelope."
+
+Because the record is written atomically, the probe doesn't need to tell absent from unreadable apart. That matches cache.get's single None.
+
+In the Phase 2 test list (after L2254), add: "with no record file: the row reads OK 0, offers no APPLY, and the topgrade_age row keeps TOPGRADE; a record whose data lacks packages reads unprobed."
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Acceptance criteria, Testing.
+:EVIDENCE:
+Spec:
+- L2222-2229: deferred row. Value comes from "the record's packages". Severity is CRIT on gate.ok false, WARN on result failed/refused/interrupted or K > 0, and "OK otherwise". There is no absent-record clause.
+- L2234: APPLY is "shown while N > 0 or a gate failure is open".
+- L2033 and L2184: --dry-run "writes nothing (no record, stamp or flag)".
+- L2316-2323 and L2405-2412: rollout step (a) ends with "confirm upgrade-guarded --dry-run exits 0", then (b) pulls Phase 2.
+- L2120: the record is written only by upgrade-guarded, at the end of a non-dry-run mode, atomically (tmp then rename).
+- L2250-2262: Phase 2 tests have no no-record or malformed-record case.
+
+Code:
+- ~/.dotfiles/maint/src/maint/cache.py:35-42: get() returns None on OSError/ValueError/KeyError/TypeError, so absent and unreadable can't be told apart.
+- maint/src/maint/probes/updates.py:114-126: topgrade_freshness turns None into schema.WARN, "no topgrade run recorded".
+- updates.py:22-39: _cached turns None or a wrong shape into schema.unprobed. Its docstring says "a wrong-shape payload must degrade here, never crash the envelope".
+- remedies.py:535-545: attach_levers adds always-remedies "regardless of severity".
+- indicator.py:42: diag = metrics without levers. Only these colour the glyph.
+:END:
+
+** DONE The everyday split transaction is never run against real pacman before the levers are repointed, and AC1's hook-silent bullet has no test
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): AC1 L2335; Testing map for AC1/AC3 L2461-2468; Phase 1 tests L2163; Rollout step (a) L2322/L2411; manual tests L2512-2519
+
+Every Phase 1 test drives upgrade-guarded against a fake pacman on PATH (L2163). Rollout step (a)'s only live check is that 'upgrade-guarded --dry-run' exits 0, and dry-run never transacts and uses checkupdates' private db (L2033). The written manual tests cover only the boot path (L2512: 'What no unit test can cover is the boot path itself'). The Testing map assigns AC1 only to the fake-pacman tests and to Phase 2 rendering.
+
+Risk: The first real exercise of the core feature is a press of UPDATE after rollout (b). A divergence between the fakes and real pacman shows up there as a guard abort or a closure refusal, while AC1 has already been ticked on fake evidence.
+
+Recommended change: Three edits:
+
+1. In Phase 4 rollout step (a), at L2316-2322 and its Readiness copy at L2405-2411, add a last bullet. On ratio: "with Hyprland live, run upgrade-guarded --no-topgrade --no-aur from a terminal in the session, right after a --dry-run. Expected: exit 0, no BLOCKED banner, the record's packages equal dry-run's predicted deferred set, and a following pacman -Qu lists only those names plus [ignored] rows. Pull the Phase 2 commit only after this passes."
+
+2. In Testing, at L2512-2519, add two manual entries for todo.org under Manual testing and validation. Each gets its own what-it-verifies line, steps and Expected line.
+ - "Everyday split run, live": run the step (a) command above. If --dry-run lists a GPU-kind entry, that entry is the hook-silence check.
+ - "Everyday split run, Hyprland stopped": the same command from a TTY with Hyprland stopped. Expected: the GPU-kind entries land, and only kernel-kind entries remain in the record and in pacman -Qu.
+ Neither entry waits on a GPU-kind update being pending before (b). If none is pending at (a), run the GPU-kind live variant on the first day one is, before ticking AC1.
+
+3. In the AC1/AC3 map at L2461, add a sub-bullet. It maps AC1's "hook silent" bullet and its exit-0/applies-everything bullet, and AC3's held-set claim, to these two manual entries.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Testing, Acceptance criteria. I took the live check before the levers move, on both machines. I moved the Hyprland-stopped variant into the make test-keep FS_PROFILE=zfs VM, which is headless and so naturally not live, rather than asking for a logged-out session on a daily driver. The command is 'upgrade-guarded --no-topgrade', because R58 removes --no-aur. It also exercises yay -Sua, which UPDATE runs anyway.
+:EVIDENCE:
+Spec:
+- L2322 and L2411: rollout step (a) ends with "confirm upgrade-guarded --dry-run exits 0".
+- L2323 and L2412: step (b) repoints the levers and the alias, and "the force path is gone".
+- L2033 and L80: dry-run "never touches the system db"; its pending set comes from checkupdates and its closure runs pacman -Sup --dbpath against checkupdates' db.
+- L2036: the everyday pending set is "LC_ALL=C pacman -Qu ... minus rows marked [ignored]".
+- L2043 and L77: the deferred set is read from a post-transaction pacman -Qu.
+- L2163: every Phase 1 test runs "with fakes on PATH for pacman, sudo, yay ...".
+- L2335: AC1 says "leave the hypr-live-update-guard hook silent during the pacman step".
+- L2460: "Each acceptance criterion maps to the cases that verify it."
+- L2461-2468: the AC1/AC3 map has no case that observes the hook.
+- L2512: "What no unit test can cover is the boot path itself"; the manual tests at L2513-2519 are all boot tests.
+- L2030 and L80: --no-aur "exists for manual runs and tests", but no manual run is scheduled.
+- L94: the proof of concept ran "before the kernel/DKMS hold and the containers disable were added".
+
+System (ratio):
+- /usr/bin/checkupdates:155: mapfile -t updates < <(pacman -Qu --dbpath "$CHECKUPDATES_DB" ... | grep -v '\[.*\]'). Dry-run never sees an [ignored] row.
+- /etc/pacman.conf:25: IgnorePkg = bridge-utils.
+- /var/log/ratio-upgrade.log: six "warning: <pkg>: ignoring package upgrade" lines, then "Packages (724)", with linux, linux-lts, zfs-dkms and zfs-utils inside the transaction. This was a real --ignore run, but without the hold or this script's argv.
+- scripts/hypr-live-update-guard:121: " BLOCKED: live GPU/compositor library upgrade while Hyprland is running". The banner the recommended check looks for exists.
+- pacman 7.1.0.r9.g54d9411-2 and pacman-contrib 1.13.1-1 are installed.
+:END:
+
+** DONE upgrade-guarded reads several real system paths with no named seam, so the specified kernel-set fixture can't be built and tests depend on the host
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 1 seam list L2163; kernel set and DKMS set L2038-2039; gate check list (c) L2148; preconditions L2060, L2099; informant L2063; systemd-cat L2101; kernel-set test L2167; informant test L2175
+
+The seam list gives upgrade-guarded only UPGRADE_GUARDED_* and MAINT_STATE_DIR. KMC_MODULES_DIR and KMC_BOOT_DIR are the gate's seams (L140, L2401). Yet upgrade-guarded itself globs /usr/lib/modules/*/vmlinuz, reads /usr/lib/modules/<kver>/pkgbase and /usr/src/*/dkms.conf, checks /usr/lib/modules/<kver>/vmlinuz for open-gate kernels, refuses on /var/lib/pacman/db.lck, detects informant with 'command -v', and pipes through systemd-cat.
+
+Risk: The kernel-set and DKMS-set tests see the host's real kernels: three on ratio, one on velox. They either can't build their fixture or pass on one daily driver and fail on the other. A pacman running in another terminal turns every test that reaches the precondition into a refusal. 'Absent informant' can't be expressed on velox. The --apply-armed tests write into the real journal.
+
+Recommended change: Edit L2163 and its mirrors at L2401 and L2453, and add a matching line to Readiness Dev tooling:
+"upgrade-guarded reads its module tree from ${KMC_MODULES_DIR:-/usr/lib/modules}, the same seam as the gate, so one fixture tree drives the kernel set, pkgbase, the vmlinuz-to-kver mapping and check list (c). It also honours UPGRADE_GUARDED_SRC_DIR (default /usr/src, for the dkms.conf glob), UPGRADE_GUARDED_DB_LCK (default /var/lib/pacman/db.lck) and UPGRADE_GUARDED_INFORMANT (the command to probe, default informant; empty means absent)."
+
+Add systemd-cat to the fakes-on-PATH list.
+
+Under "Steps and ordering" (near L2181-2182), and in the verification line at L2510, add:
+"a file at UPGRADE_GUARDED_DB_LCK makes the everyday run, --complete and --apply-armed each exit 3 before any transaction, naming the file and the remedy, and the file is still present afterwards."
+
+At L2175, change "absent informant" to "absent informant (UPGRADE_GUARDED_INFORMANT empty)".
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Readiness dimensions, Testing.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L2163: "...the env seams UPGRADE_GUARDED_HOOK, ..._HYPR_RUNNING, ..._ARM_FLAG, ..._ARMED_LIST, ..._TOPGRADE, ..._MAINT, ..._UNIT_ENABLED and maint's MAINT_STATE_DIR, plus KMC_MODULES_DIR and KMC_BOOT_DIR, with fakes on PATH for pacman, sudo, yay, informant, systemctl, vercmp, checkupdates, dkms, zfs, findmnt and lsinitcpio". systemd-cat isn't in the fakes list.
+- L140: "Its test seams are KMC_MODULES_DIR and KMC_BOOT_DIR". L2401 and L2453 both say "the gate's KMC_MODULES_DIR and KMC_BOOT_DIR".
+- Hardcoded paths: L2038 (/usr/lib/modules/*/vmlinuz, /usr/lib/modules/<kver>/pkgbase), L2039 (/usr/src/*/dkms.conf), L2148 (c) (/usr/lib/modules/<kver>/vmlinuz), L2060 and L2099 (db.lck), L2063 (command -v informant), L2101 (systemd-cat -t archsetup-boot-upgrade).
+- L2167 is the three-owner fixture, which asserts on "the gate's check list". L2175 says "a failing or absent informant is tolerated". L2427 says informant is installed on velox and absent on ratio.
+- db.lck refusal: the AC at L2379 requires it, but the Phase 1 precondition tests at L2181-2182 and L2510 cover only EUID 0 and the held flock.
+
+Live system (ratio):
+- ls /usr/lib/modules/*/vmlinuz returns three files: 6.18.25-1-lts-strix (linux-lts-strix), 6.18.54-1-lts (linux-lts) and 7.2.7-arch1-1 (linux).
+- /usr/src/zfs-2.4.4/dkms.conf exists.
+- stat /var/lib/pacman shows root:root 755.
+- informant is absent.
+- /usr/bin/systemd-cat is present.
+
+Repo precedent:
+- tests/zfs-pre-snapshot/test_zfs_pre_snapshot.py:50 prepends the fake bin to the real PATH (env["PATH"] = self.bin + os.pathsep + env["PATH"]), and :52 seams the lock path with ZFS_PRE_LOCKFILE.
+- tests/hypr-live-update-guard/test_hypr_live_update_guard.py:51 seams the sentinel path with HYPR_GUARD_SENTINEL.
+:END:
+
+** DONE APPLY's place in maint's remedy model and its CLI form are undefined
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 2 APPLY L2234-2236; Phase 2 tests L2257
+
+APPLY 'is the lever of upgrade_deferred (always: True)' but 'doesn't use the captured-output lever runner'. Like MERGE, it Popens foot, and 'maint's CLI form of APPLY runs upgrade-guarded --complete in the foreground'. No remedy id, remedy kind or CLI invocation is named.
+
+Risk: The implementer has to invent a new remedy kind or key kind, its id, and a non-capturing CLI execution path. Built as an ordinary 'user' remedy, the CLI form runs the whole kernel session with its output captured and under a timeout. The tty reboot prompt then never appears, and the gate verdict is cut to a 160-character last line.
+
+Recommended change: In the Phase 2 APPLY block (L2234-2235), replace the first sentence and the last sentence of the second bullet with this:
+
+"APPLY is the REMEDIES entry =apply_deferred=:
+- tier confirm and a new kind =terminal=;
+- argv =['upgrade-guarded', '--complete']=;
+- metric_ids =['upgrade_deferred']=, always True;
+- no timeout;
+- not a strip key, and in no macro.
+
+Every GUI press of a terminal-kind remedy arms on the first press and Popens =['foot', '--hold', '-e', *argv]= with start_new_session=True on the second, never calling _fire. That covers the row lever, the =topgrade_age= row's APPLY, and REVIEW & FIX's FIX key. =maint fix apply_deferred= is the CLI form, and it is the usage line REVIEW & FIX prints. It runs argv with inherited stdin, stdout and stderr and no timeout, never through cmd.run, and returns the script's exit code."
+
+Add two test bullets at L2257:
+- iter_fix and =maint fix apply_deferred= never call cmd.run, and run the argv with inherited stdio.
+- A REVIEW & FIX FIX press on apply_deferred Popens the foot argv and never streams through _fire.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Testing. Adopted, with APPLY's metric_ids widened to include topgrade_age so that R18's card key attaches through attach_levers.
+:EVIDENCE:
+Spec:
+- L2234: "It is the lever of =upgrade_deferred= (always: True)".
+- L2235: "doesn't use the captured-output lever runner. Like MERGE, it opens =foot --hold -e upgrade-guarded --complete= via subprocess.Popen(..., start_new_session=True) ... maint's CLI form of APPLY runs =upgrade-guarded --complete= in the foreground". No id, no kind, no CLI command.
+- L2257: tests cover only "APPLY's argv and the detach". L2490 covers the GUI path only.
+- L2096 and L2201: --complete prompts 'Reboot now? [y/N]' whenever stdin is a tty.
+- L697 (prior evidence): "maint's cmd.run ... capture_output=True with stdin inherited. A prompt there is either invisible or gets EOF."
+
+Code (~/.dotfiles/maint/src/maint):
+- remedies.py:535-545: always is honored only for REMEDIES entries.
+- Remedy kinds: priv (35), user (7) and macro (1); no other kind exists.
+- viewmodel.py:259: MERGE is a row key {"kind": "merge"}, not a remedy.
+- gui.py:986: the merge key is dispatched to _on_merge.
+- gui.py:1665-1672: _on_merge Popens _MERGE_ARGV with start_new_session=True.
+- gui.py:1408-1445: _press_lever is "the arm-or-fire state machine behind every remedy key". It leads to _fire and the doctor stream.
+- gui.py:740-743: the REVIEW & FIX roster gives every item-less remedy an armed FIX key.
+- doctor.py:389-426: review() lists =maint fix <rid>= for confirm-tier remedies on WARN or CRIT metrics.
+- doctor.py:161: non-priv steps run cmd.run(step["argv"], timeout=r.get("timeout", 60)).
+- doctor.py:164-168: on success the detail is the last stdout line, cut to 160 characters; on failure, a 500-character stderr tail.
+- cmd.py:22: subprocess.run(cmd, capture_output=True, text=True, timeout=timeout).
+- cli.py:136-148: cmd_fix routes through doctor.iter_fix only.
+- grep for stdin, inherit or capture_output=False in maint finds no foreground execution path.
+:END:
+
+** DONE Phase 2's GUI-level tests have no harness, and an unhandled key kind silently becomes MERGE
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 2 tests L2250-2262 (L2257 'APPLY's argv and the detach'); DISMISS key L2225; AC7 L2357-2359
+
+Phase 2 requires a test of APPLY's argv and detach, placed 'like MERGE'. It adds a DISMISS key that writes upgrade_news_dismissed. The only DISMISS test listed checks the read side ('leaves dismissed news titles off the row'). Nothing says whether DISMISS clears one title or all of them.
+
+Risk: Put where the spec places it, the APPLY detach test can't be written without GTK. A new 'dismiss' or 'apply' key that _digest_key misses opens pacdiff instead, and no test would catch it. AC7's 'until they are dismissed' has no test of the write.
+
+Recommended change: Change L2225 to: "...which the row's DISMISS key writes. DISMISS is one unarmed row key, shown while any undismissed title is listed. It adds every title currently shown to upgrade_news_dismissed through a GTK-free helper in panel.py, and gui.py's _digest_key handles the new kind explicitly."
+
+Change L2257 to: "APPLY's argv and the detach: the argv and a detached launcher live in panel.py, and gui.py calls them. The test patches subprocess.Popen there and asserts the foot --hold argv and start_new_session=True."
+
+Add a test bullet after L2254: "DISMISS writes every shown title to upgrade_news_dismissed, after which the row lists none of them, and a title first recorded by a later run shows."
+
+Optional: a test that each key kind viewmodel emits is one gui.py dispatches explicitly.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Testing. I adopted the GTK-free placement and the explicit dispatch. DISMISS follows R38's resolution: arm-then-fire, overwriting the key with the record's whole news list. That replaces one unarmed press that adds the shown titles, which would grow the key without bound. The optional 'every key kind is dispatched' test is not taken. It needs gui.py imported, which needs GTK, and handling the new kind explicitly already closes the risk this change introduces.
+:EVIDENCE:
+Spec L2225: "the news titles minus those in maint key =upgrade_news_dismissed=, which the row's DISMISS key writes". L2254: the only DISMISS-related test, read side only. L2257: "APPLY's argv and the detach". L2235: "Like MERGE, it opens =foot --hold -e upgrade-guarded --complete= via =subprocess.Popen(..., start_new_session=True)=". L2359 (AC7): "the deferred row shows them until they are dismissed". L2500: the Testing map for AC7 names only the read side.
+
+Code in ~/.dotfiles/maint/src/maint/:
+- gui.py:57-60 imports cairo and gi and requires Gtk 4.0.
+- gui.py:945-986: _digest_key's final unconditional return sends any unhandled kind to self._on_merge(a[0]).
+- gui.py:94: _MERGE_ARGV = ["foot", "-e", "sudo", "pacdiff"].
+- gui.py:1665-1677: _on_merge Popens it with start_new_session=True.
+- viewmodel.py:204-206 lists key kinds as fix or keep/unkeep/merge. viewmodel.py:259 emits kind "merge", which relies on the fall-through.
+- grep -i dismiss across maint/src/maint/*.py finds no news-dismiss handler today.
+
+Tests in ~/.dotfiles/tests/maint/:
+- test_panel.py:3-7: "The GTK view (maint/gui.py) is never imported here — every decision it renders lives in the GTK-free presenter (maint/panel.py)".
+- grep for gui imports in tests/maint finds only that docstring.
+
+On ratio, python3 with gi.require_version('Gtk','4.0') imports fine.
+:END:
+
+** DONE The Phase 3 boot checks gating velox can't be observed in the VM the spec names, and that VM gets the published dotfiles
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Testing L2513-2514, L2526; Readiness Dev tooling L2403
+
+'The Phase 3 boot checks pass in the make test-keep FS_PROFILE=zfs VM before the unit reaches velox.' The checks expect a visible tty1 banner and login, and the 'Armed boot' check starts with 'arm from APPLY and confirm armed... and REBOOT'. The same paragraph also says 'Where Hyprland isn't live in the VM, UPGRADE_GUARDED_HYPR_RUNNING=1 forces the arm path'.
+
+Risk: The gate that has to pass before velox can't produce its expected observations as written, so it either stalls or gets waived. Unless Phase 2 has been pushed, the VM's maint has no deferred row, so 'the deferred row empty / at WARN' can't be checked either. Pushing Phase 2 early, on the other hand, exposes it to a velox pull before step (a).
+
+Recommended change: Replace the Trigger bullet's last sentence at L2513 with a VM-procedure bullet:
+
+"In the VM (make test-keep FS_PROFILE=zfs runs QEMU headless with no compositor), arm over SSH from a console with 'UPGRADE_GUARDED_HYPR_RUNNING=1 upgrade-guarded --complete' instead of APPLY. Confirm the armed state from the flag and the record, not the panel. Observe tty1 through the QEMU monitor's screendump on vm-images/qemu-monitor-zfs.sock; a guest reboot keeps the same QEMU process and SSH. Confirm the banner and the before-login ordering from 'journalctl -u archsetup-boot-upgrade' and getty@tty1's start timestamp. Arming from APPLY, the panel's armed row and REBOOT are checked on ratio and velox (AC5)."
+
+Change L2526 (and the matching clause at L2403) to: "The Phase 3 boot checks other than Unread news (velox) and Banner on ratio pass in the make test-keep FS_PROFILE=zfs VM before the unit reaches velox."
+
+Optionally add "run it with DOTFILES_SOURCE=$HOME/.dotfiles if the Phase 2 commit isn't on the published remote yet."
+
+Do not adopt the finding's debug-vm.sh step. debug-vm.sh:92-94 restores the clean-install snapshot on the kept base disk and would erase the install under test.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Testing.
+:EVIDENCE:
+Spec L2526: "The Phase 3 boot checks pass in the =make test-keep FS_PROFILE=zfs= VM before the unit reaches velox." L2403: same gate under Dev tooling. L2514: "Armed boot, networking down: arm from APPLY and confirm 'armed — N apply at next boot' and REBOOT; ... the banner ... on tty1, the set installed from the cache before the tty1 login". L2513: "Where Hyprland isn't live in the VM, =UPGRADE_GUARDED_HYPR_RUNNING=1= forces the arm path." L2515: "Unread news: an armed boot on velox". L2519: "Banner on ratio". L2235: APPLY opens "foot --hold -e upgrade-guarded --complete" from the panel.
+
+Code:
+- scripts/testing/run-test.sh:150: start_qemu "$DISK_PATH" "disk" "" "none"
+- scripts/testing/lib/vm-utils.sh:226-229: gtk gets virtio-vga-gl and -display gtk, anything else gets -display none
+- scripts/testing/tests/test_desktop.py:7: "(the headless test VM has none)"
+- scripts/testing/archsetup-vm.conf:15: AUTOLOGIN=yes
+- dotfiles maint/src/maint/cli.py:32-40: _human prints severity, id, label and value only
+- run-test.sh:197: DOTFILES_SOURCE defaults to https://git.cjennings.net/dotfiles.git
+- run-test.sh:176: archsetup bundled from HEAD
+- scripts/testing/debug-vm.sh:92-94: 'if [ -z "$OVERLAY_DISK" ] && snapshot_exists "$VM_DISK" "clean-install"; then ... restore_snapshot "$VM_DISK" "clean-install"'
+- vm-utils.sh:63 and 73: DISK_PATH=archsetup-base-zfs.qcow2 and MONITOR_SOCK=qemu-monitor-zfs.sock under FS_PROFILE=zfs
+- vm-utils.sh:198-215: the qemu argv has no -no-reboot and no -vga none
+:END:
+
+** DONE The routine-update workflow keeps running bare topgrade on velox until the Phase 4 doc commit, which waits on Phase 3
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Implementation phases L2022 (docs land after Phase 3); Phase 4 L2313; Risks 'Entry points outside the hold' L2446; Testing L2526
+
+The docs/workflows/system-health-check.org change, which makes step 6 run upgrade-guarded and stop hand-stamping, lands in Phase 4 'after Phase 3'. Phase 3 can't reach velox until the VM boot checks pass. Risks says the hold covers 'the system-health-check workflow's update step'.
+
+Risk: Between Phase 2 and Phase 4, an agent-run health check on velox lands the kernel and zfs-dkms ungated through bare topgrade. That's the unbootable chain the kernel decision exists to prevent, while the spec claims the workflow is covered. It may also hand-stamp freshness while the record lists deferred entries.
+
+Recommended change: Change the L2022 intro sentence to: "The system-health-check.org edit lands as an archsetup doc commit right after Phase 1. Phase 4 holds the remaining docs (the maint/README.md flow and recovery steps in dotfiles), which land after Phase 3, and the per-machine rollout."
+
+Move the L2313 bullet out of Phase 4 into Phase 1, worded as: "Update docs/workflows/system-health-check.org Phase 3, step 6:
+- Where upgrade-guarded --dry-run exits 0 (step (a) done on this machine), run upgrade-guarded instead of plain topgrade.
+- A day with a pending kernel or DKMS package lands it through upgrade-guarded --complete as the dedicated session. Until step (c), a pending GPU/compositor set finishes from a console with Hyprland stopped (L2095).
+- Where step (a) isn't done and a kernel or DKMS package is pending, don't run plain topgrade.
+Step 5's strix addendum keys on --complete. The hand-stamp fallback becomes 'the script stamps; never hand-stamp while a deferred set is outstanding'. Append a resolution note to the 2026-09-12 containers KIL entry."
+
+Leave L2446 unchanged.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Readiness dimensions.
+:EVIDENCE:
+- Spec L2022: "Phase 4 holds the docs, which land after Phase 3 as one doc commit in each repo (... =docs/workflows/system-health-check.org= in archsetup)".
+- Spec L2313: the health-check edit (step 6 runs upgrade-guarded; kernel/DKMS/GPU days go through --complete; never hand-stamp) sits under Phase 4.
+- Spec L2095: --complete with the unit not enabled prints 'apply from a console with Hyprland stopped' and writes no flag. That path works after Phase 1 alone ("covers the window between Phases 1 and 3").
+- Spec L2133-2134: --complete stamps through ~/.local/bin/maint stamp topgrade. dotfiles maint/src/maint/cli.py:292 already has the stamp subparser, so stamping needs no Phase 2 code.
+- Spec L2446: bare topgrade, yay and pacman -Syu still land the kernel and DKMS sets ungated. The hold covers "the system-health-check workflow's update step", which describes the end state.
+- Spec L2526: the VM boot checks gate only "the unit reaches velox". They don't gate the Phase 3 commit.
+- docs/workflows/system-health-check.org:236, step 6: "Run =topgrade= ... If the run happened outside the wrapper somehow, =maint stamp topgrade= records it by hand." Also :1067, the KIL entry with the hand stamp after the containers failure.
+- .ai/project-workflows/system-health-check.org symlinks to that file, and ~/projects/home's copy is only a pointer to it. So this file is the canonical workflow.
+- todo.org:763-765 (2026-09-17): "no plain topgrade on velox until the hold is built". todo.org:778-781: "The 2026-10-04 velox health check ran a plain topgrade that moved the kernel 6.18.51 → 6.18.55, against the note above."
+:END:
+
+** DONE Repointing the sysupgrade alias in common/ breaks it on dwm and none installs, which never get upgrade-guarded
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 2 'Alias and README' L2247; Rollout L2315 ('A rebuilt machine gets every piece from the installer')
+
+Phase 2 points the sysupgrade alias in common/.zshrc.d/aliases.sh and common/.bashrc.d/aliases.sh at upgrade-guarded.
+
+Risk: On a dwm or none rebuild, sysupgrade becomes 'command not found'. That contradicts 'a rebuilt machine gets every piece from the installer'.
+
+Recommended change: Replace the spec L2247 bullet with:
+
+"- The =sysupgrade= alias in =common/.zshrc.d/aliases.sh= and =common/.bashrc.d/aliases.sh= (today plain =topgrade=) becomes conditional, because common/ is also stowed on dwm installs, where =upgrade-guarded= and its hook are never installed: 'if command -v upgrade-guarded >/dev/null 2>&1; then alias sysupgrade=upgrade-guarded; else alias sysupgrade=topgrade; fi'. minimal/'s alias (DESKTOP_ENV=none) stays plain =topgrade=."
+
+Add a matching Phase 2 test bullet, for example a source check that both common alias files carry the conditional and that the minimal ones are unchanged.
+
+Optionally, reword Rollout L2315 to "A rebuilt Hyprland machine gets every piece from the installer."
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Testing.
+:EVIDENCE:
+- Spec L2246-2247: "The =sysupgrade= alias in =common/.zshrc.d/aliases.sh= and =common/.bashrc.d/aliases.sh=, today an alias for plain =topgrade=, points at =upgrade-guarded=."
+- Spec L2025: "both installed to =/usr/local/bin= by the installer step that installs =hypr-live-update-guard= (inside =hyprland()=)". Spec L80 and L220 say the same.
+- Spec L2037: a missing hook file makes every mode except --apply-armed refuse with exit 3.
+- archsetup:1497-1503: the 'dwm|hyprland)' branch stows common, then "$desktop_env".
+- archsetup:1518-1521: the 'none)' branch stows only minimal.
+- archsetup:2678-2690: window_manager runs dwm for dwm, hyprland for hyprland, and skips for none.
+- archsetup:2614-2653: the guard binary and the 10- hook are installed inside hyprland().
+- archsetup:189-190: DESKTOP_ENV is validated as dwm, hyprland or none.
+- archsetup:3429: 'aur_install topgrade' is in supplemental_software, which runs for every environment.
+- ~/.dotfiles/common/.zshrc.d/aliases.sh:42 and common/.bashrc.d/aliases.sh:42: alias sysupgrade="topgrade".
+- ~/.dotfiles/minimal/.zshrc.d/aliases.sh:42 and minimal/.bashrc.d/aliases.sh:42: alias sysupgrade="topgrade". The spec doesn't touch these, which refutes the none claim.
+- The hyprland stow package has no .zshrc.d or .bashrc.d (ls -a: .config .gnupg .local .profile.d). common/.zshrc:176-177 and common/.bashrc:57-58 glob-source *.sh from those directories.
+- dwm installs are exercised in tests/installer-steps/test_orchestrators.py and scripts/testing/tests/test_desktop.py.
+:END:
+
+** DONE Phase 2 keeps the guard arm-line wording in three code paths that its own tag removal makes unreachable
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 2 Guard UX L2219-2220; Phase 2 tests L2252, L2262; Testing L2467
+
+L2219 removes guard: "live_update" from the update and topgrade remedies, along with the iter_fix refusal, the --force sentinel wrap and _update_force. L2220 then requires arm_line, _rearm_after_guard's text and the doctor review suffix to read 'UPDATE armed — will defer at least <guard matches> (kernels and DKMS modules are always held) — press again to run upgrade-guarded' on a tripped read.
+
+Risk: Implemented as written, Decision 3's display badge never renders, which leaves the rewritten TOML mirror and its pin test guarding a display that nothing shows. The alternative is that the implementer invents a new gating key. Either way, at least two of the three reworded sites are dead code whose text tests still pin.
+
+Recommended change: Replace L2220 with the following:
+
+"=guard.trips= survives only as the arm-line annotation that Decision 3 calls the display-side mirror. The arm press for UPDATE and TOPGRADE (strip keys and the roster FIX key alike) is gated on rid in ('update', 'topgrade'), not on the dropped tag. =panel.guarded= is deleted, or rekeyed to that rid set. On a tripped read, that arm line reads 'UPDATE armed — will defer at least <guard matches> (kernels and DKMS modules are always held) — press again to run upgrade-guarded'. Nothing says live apply or REBOOT required. With the refusal gone, doctor emits no 'guard' event, so =_rearm_after_guard=, the 'guard' branch of =_on_fired=, the 'guard' event handling at cli.py:114 and panel.py:368, and the doctor review suffix (doctor.py:415) are deleted with it. maint computes no kernel set of its own; the post-run row built from the record gives the exact set."
+
+Make two matching test edits:
+- In the Phase 2 test list (L2262) and in Testing (L2467), replace the arm-line bullet with: "the arm press for UPDATE and for TOPGRADE, on a tripped =guard.trips= read with no guard tag on either remedy, shows 'UPDATE armed — will defer at least <guard matches> (kernels and DKMS modules are always held) — press again to run upgrade-guarded'."
+- Add to the rewritten-tests bullet: "test_panel_levers.py:397-406 and test_panel_phase10.py:417-422 rewritten to the rid-gated annotation."
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Testing. I adopted the rid gating and the deletions, the review suffix included. panel.guarded is rekeyed rather than deleted, so the arm press stays testable without GTK. The label is generic rather than a fixed 'UPDATE', because TOPGRADE shares the path (R07, R19).
+:EVIDENCE:
+The arm-line gate:
+- gui.py:1430-1433 reads "if panel.guarded(rid): self._act(self._guard_arm_line(title, argv, th)) else: self._act(viewmodel.arm_line(title, argv))".
+- panel.py:228 reads "return remedies.REMEDIES.get(rid, {}).get("guard") == "live_update"".
+- remedies.py:298 and 309 are the only carriers of the tag.
+
+The guard event and its consumers:
+- doctor.py:216, "if not dry_run and r.get("guard") == "live_update" and not force:", is the only place that yields _ev("guard", ...) (doctor.py:235).
+- gui.py:1496-1498, guard_ev -> _rearm_after_guard.
+- gui.py:1525, "self._update_force = True".
+- cli.py:114 and panel.py:368 also consume kind == "guard".
+
+The review suffix:
+- doctor.py:399 iterates remedies.REMEDIES.
+- doctor.py:415, "if r.get("guard") == "live_update":", appends the suffix.
+
+The spec text:
+- L2219: "Drop guard: "live_update" from the update and topgrade remedies, so iter_fix no longer refuses them ... the GUI's _update_force override".
+- L2220: "On a tripped read, arm_line, _rearm_after_guard's text and the doctor review suffix read 'UPDATE armed — will defer at least ...'".
+- No other spec line defines a new gate. The grep hits are only L2219, L2220, L2262, L2467 and the history at L1075-1094.
+
+Tests pinning the old behavior:
+- tests/maint/test_panel_levers.py:397-406 pins "press again to apply live (REBOOT required after)".
+- tests/maint/test_panel_phase10.py:417-422 pins panel.guarded("update") and panel.guarded("topgrade") as True.
+:END:
+
+** DONE The script contract is restated two to four times, and the copies have already drifted
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L68-170 vs Phase 1 L2024-2208; Decisions 4, 6 and 7 L227-257; Readiness L2371-2434; Testing L2451-2526
+
+Each of these appears in full in two to four places:
+- the everyday sequence: L82-92 and L2056-2078
+- --complete: L128-133 and L2080-2096
+- --apply-armed: L156-166, L2098-2107 and L249
+- the record schema: L102-111, L227, L2118-2129 and L2372
+- the stamp predicate: L113-118, L228-232 and L2135-2141
+- the arm flag: L142-147, L256, L2109-2116 and L2373
+- the rollout checklist: L2315-2328, L2404-2417 and L2524
+- the test lists: L2163-2208, L2250-2262 and L2292-2307, again as L2460-2510
+
+Risk: Every later edit has to land in every copy. An implementer or test author who reads the wrong copy builds the wrong test or skips the Targets check. The wrong gate test would reject exactly the snapshot window that the 05-zfs-snapshot hook's 60 s skip produces (L2444).
+
+Recommended change: Smallest edit, in three parts:
+
+1. Change L2193's snapshot clause to: "snapshots created before since − 60 s fail, naming the dataset; one created at or after since − 60 s passes (including one in [T0 − 60 s, T0), the window the hook's 60 s skip leaves)". That matches L2155 and L2481.
+
+2. Bring L2408 in line with L2319 by replacing it with: "write =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook= from the installer heredoc, confirm its Targets match the heredoc's, and remove the unprefixed name (confirm with =ls /etc/pacman.d/hooks/=; ratio still has the legacy name, and velox's was hand-placed under it)". Better still, replace L2405-2417 with "(a) to (c) as Phase 4's Rollout checklist" so only one checklist remains.
+
+3. Add one sentence to the Implementation phases intro (L2022): "Where Design, Decisions, Readiness or Testing restate a phase bullet, the phase bullet is normative." That way a later disagreement resolves one way.
+
+The broader consolidation (Design down to set definitions and the 'For the user' narrative, Decisions down to choice plus a pointer, Readiness and Testing citing phase bullets by AC) is an optional simplicity follow-up, not needed to close this finding.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Design, Decisions, Implementation phases, Acceptance criteria, Readiness dimensions, Risks, Testing.
+:EVIDENCE:
+- Spec L2193: "only pre-T0 snapshots fail; one created after T0 − 60 s passes".
+- Spec L2155: "a <rootds>@pre-pacman_* snapshot with creation ≥ since − 60 s".
+- Spec L140: "created no earlier than 60 s before since".
+- Spec L2481: "only snapshots created before since − 60 s fails ... and one created after it passes".
+- Spec L2444: the hook skips within 60 s, "which is why the gate requires one no more than 60 s older than the run's start".
+- scripts/zfs-pre-snapshot:13 sets MIN_INTERVAL="${ZFS_PRE_MIN_INTERVAL:-60}", and lines 19-24 exit 0 without a snapshot when the lockfile is younger than 60 s. So the newest pre-pacman snapshot can predate T0 by up to 60 s.
+- Spec L2319: "write ... from the installer heredoc, confirm its Targets match the heredoc's, remove the unprefixed ...".
+- Spec L2524: "the hook migrated to the 10- name with its Targets matching the installer heredoc's".
+- Spec L2408: "migrate the guard hook to 10-hypr-live-update-guard.hook and remove the unprefixed name". There is no Targets check.
+- ls /etc/pacman.d/hooks/ shows 99-grub-sync-efi.hook and hypr-live-update-guard.hook (legacy name on ratio).
+- Diffing the Target lines of ratio's live hook against the archsetup heredoc (archsetup:2625-2650): 15 vs 15, identical. Velox wasn't checked (offline).
+- The duplicated ranges were re-read and match the finding's citations.
+- No precedence statement among the copies anywhere in the spec (Implementation phases intro, L2022).
+:END:
+
+** DONE kernel-modules-check's no-argument 'standalone diagnostic' mode has no caller and no test, and needs foreign-kernel logic of its own
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L140 (final sentences), L108; Phase 1 L2156, L2127
+
+With no kver given, the gate checks every installed kernel whose vmlinuz owner isn't foreign, without the since requirements ('the standalone TTY diagnostic'). L108 and L2127 add that a passing standalone run never clears the record's gate.
+
+Risk: An extra untested branch and package-ownership query inside a safety script. A regression there would only show when someone runs it by hand during a recovery.
+
+Recommended change: 1. Make at least one kver mandatory. In the gate paragraph at L140 and the gate list at L2156, replace the no-kver sentence with: "At least one =<kver>= is required; with none it exits 2 (usage). For manual diagnosis, run =kernel-modules-check <kver>=, using the kver the CRIT row names."
+2. Spell out the empty-list case. At L140/L2149 and step 4 (L2087), say: "=upgrade-guarded= skips the invocation when a check list is empty, so an empty list means no gate call, snapshot check included."
+3. Move the test "an empty check list with no fresh snapshot passes" (L2193, L2482) out of the =kernel-modules-check= tests and into the =--complete= tests, reworded as: "a =--complete= whose check list is empty never invokes =kernel-modules-check= and passes with no fresh snapshot."
+4. Add one gate test: "no kver exits 2."
+5. Simplify the record wording. At L108, L2127 and AC4 (L2347), replace "a passing standalone =kernel-modules-check= never clears it" with "only a =--complete= re-check clears it."
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Acceptance criteria, Testing. I cut the no-argument diagnostic and made at least one argument mandatory, as recommended, with two changes. The arguments are pkgbases, not kvers, because the entry is keyed by pkgbase (R03). The gate then owns every failure item, 'no module tree' included, and stays pacman-free. And --since stays optional, because it now has a caller: the structural check after a stage 1 that changed nothing (R03). The supported manual re-check is upgrade-guarded --complete.
+:EVIDENCE:
+Spec, docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L135: =kernel-modules-check [--since <epoch>] [<kver>...]= … "Only =--complete= runs it, after stage 1".
+- L140: "An empty list skips the gate, snapshot check included. Kernels outside the list are never checked" … "With no kver given it checks every installed kernel whose vmlinuz owner isn't foreign, without the since requirements; that is the standalone TTY diagnostic. Its test seams are =KMC_MODULES_DIR= and =KMC_BOOT_DIR=, with dkms, zfs, findmnt and lsinitcpio faked on PATH." No pacman fake is listed.
+- L2087: "Gate: whatever stage 1's exit code, build the check list and run =kernel-modules-check= (below)." There is no skip-when-empty instruction.
+- L2149 and L2156 repeat L140's two sentences.
+- The test lists (L2193 and L2482, the latter under "Phase 1, =kernel-modules-check= on fake…") include "an empty check list with no fresh snapshot passes". Every other gate test passes kvers.
+- L108, L2127 and L2347 exist only to say a passing standalone run doesn't clear the record.
+- Phase 4 docs and recovery (L2309-2330, L2398) and the recovery drill (L2521) never use the no-kver form.
+
+Live system (ratio):
+- =uname -n= prints ratio, and =findmnt -no FSTYPE /= prints btrfs.
+- =pacman -Qm= lists linux-lts-strix 6.18.25-1.
+- =dkms status -k 6.18.25-1-lts-strix= prints "zfs/2.4.4: added".
+- 6.18.54-1-lts and 7.2.7-arch1-1 both print "installed".
+- The foreign-owner filter is therefore the only thing that keeps the standalone mode from failing on ratio.
+:END:
+
+** DONE upgrade_news_dismissed is a new state key that is specified only by its reader
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 2 L2225; Phase 2 tests L2254; Readiness L2375; Risks L2448; Testing L2500
+
+The deferred row's DISMISS key writes the maint key upgrade_news_dismissed, and the probe hides any news title found in it.
+
+Risk: The implementer has to invent the write semantics. An append-only key would grow for the life of the machine. And the one user action in the news feature goes untested.
+
+Recommended change: Replace the Phase 2 Evidence bullet clause at L2225 ("which the row's DISMISS key writes") with this text:
+
+"which the row's DISMISS key writes: one press, shown only while undismissed titles exist, overwrites the maint cache key upgrade_news_dismissed (cache.put, beside the record) with the news titles the row is showing, so the key never holds more than the record's 10."
+
+Make the matching edit to Readiness L2375: "upgrade_news_dismissed (maint cache key, cache.put envelope), overwritten by the deferred row's DISMISS key with the row's shown titles; never appended, so it holds at most 10."
+
+Add one Phase 2 test after L2254: "pressing DISMISS on a row showing news writes exactly the shown titles to upgrade_news_dismissed, replacing any prior value, and the re-probed row shows no titles and no DISMISS key."
+
+If one press is the wrong friction for hiding a manual-intervention notice, say arm-then-fire there instead, as MARK KNOWN does. Either way, state the choice in the spec.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Readiness dimensions, Testing. I adopted the bounded key and the write-side test, with two changes. First, DISMISS overwrites the key with the record's whole news list rather than the titles the row is showing. Overwriting with only the shown lines brings earlier dismissals back once a new line arrives: dismiss A and B, C arrives, DISMISS stores only C, and A and B return. The record already caps news at 10, so the key stays bounded. Second, DISMISS is arm-then-fire, which the finding offers as the alternative. Hiding a manual-intervention notice is silencing a signal, and MARK KNOWN, the panel's other signal-silencing key, is arm-then-fire for exactly that reason (gui.py:1687).
+:EVIDENCE:
+Spec mentions of upgrade_news_dismissed and DISMISS (all reader-side):
+- L2225: "the news titles minus those in maint key =upgrade_news_dismissed=, which the row's DISMISS key writes"
+- L2375: "=upgrade_news_dismissed= (maint key), written by the deferred row's DISMISS key."
+- L2254: the test only "leaves dismissed news titles off the row"
+- L2500: the row "lists the record's news minus the titles in =upgrade_news_dismissed="
+- L1742 (review disposition): "dismissal is a maint-owned key". No shape, store or write rule is given anywhere.
+
+Record news cap: L109 and L2128 say news is "merged with the previous record's, deduplicated, keeping the newest 10." The dismissed key has no cap.
+
+Store naming: the spec calls the record a "maint cache key" (L106, L2372) and topgrade_run a "maint-owned cache key" (L2374). The dismissed key is only called a "maint key" (L2375).
+
+maint code:
+- cache.py:24-32: put() writes ~/.local/state/maint/<name>.json in the {written_at, data} envelope.
+- thresholds.py:36-38: the curation store is ~/.config/maint/curation.toml.
+- curation.py:164-190: mark() appends to a table's add list, dedups and never trims.
+- gui.py:1653-1663 (KEEP) and gui.py:1686-1700 (MARK KNOWN) are the existing row-level persisted-decision keys. MARK KNOWN is arm-then-fire, per its docstring "silencing signal deserves the same two-press deliberateness as a remedy".
+- grep -i dismiss in maint/src/maint finds only panel.py:247, panel.py:390 and gui.py:603/616, none of which is row-level.
+
+Phase 2 test list (L2250-2262): APPLY gets "APPLY's argv and the detach", but no item presses DISMISS.
+:END:
+
+** DONE F19: the Phase 1 snapshot test bullet contradicts the gate's 60 s tolerance
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 1 Gate tests L2193 vs gate rule L2155/L140 and Testing L2481
+
+L2193: 'only pre-T0 snapshots fail; one created after T0 − 60 s passes'. The gate rule (L2155) passes any snapshot with creation ≥ since − 60 s. The Testing map (L2481) says 'only snapshots created before since − 60 s fails'.
+
+Risk: A test written from L2193 with a fixture at T0 − 30 s would assert a failure the gate must not produce. Either the test or the gate ends up wrong.
+
+Recommended change: At L2193, replace "only pre-T0 snapshots fail; one created after T0 − 60 s passes" with "only snapshots created before T0 − 60 s fail, naming the dataset; one created at or after T0 − 60 s passes, including one taken up to 60 s before T0". Optionally, at L2481, fix the agreement ("fails" → "fail") and change "one created after it passes" to "one created at or after it passes", so the boundary matches the ≥ in L2155.
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Testing.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L140 (Design): "requires a <rootds>@pre-pacman_* snapshot created no earlier than 60 s before since"
+- L2155 (Phase 1 gate rule): "a <rootds>@pre-pacman_* snapshot with creation ≥ since − 60 s"
+- L2193 (Phase 1 Gate tests): "only pre-T0 snapshots fail; one created after T0 − 60 s passes"
+- L2481 (Testing, AC4): "only snapshots created before since − 60 s fails, naming the dataset it looked for, and one created after it passes"
+- L964 (Review findings, the origin of the L2193 wording): "snapshot listings where only pre-T0 snapshots exist (gate fails) and one is newer than T0 (gate passes)"
+- L2444: "(skipped within 60 s of the previous snapshot ...) which is why the gate requires one no more than 60 s older than the run's start"
+Code: scripts/zfs-pre-snapshot:13 has MIN_INTERVAL="${ZFS_PRE_MIN_INTERVAL:-60}", and lines 16-24 exit 0 without snapshotting when now - last < MIN_INTERVAL.
+Counterexample: a snapshot created at T0 − 30 s is pre-T0, so L2193's first clause makes it fail. It is also after T0 − 60 s, so L2193's second clause, L2155 and L2481 make it pass.
+:END:
+
+** DONE F23: the shared record fixture became two unsynchronized copies with no canonical source
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Decision 4 Consequences L233; Phase 1 tests L2191; Phase 2 tests L2254; Testing L2464
+
+The resolution said Phase 2's probe test reads the same fixture as Phase 1. The spec has 'both test suites read copies of one fixture' (L233) and 'Phase 2's probe test reads a copy of it' (L2191, L2254, L2464). It names neither the canonical file nor how the copies stay in step.
+
+Risk: A later change to the record format, such as a new field or a renamed kind, passes both suites while the panel misreads the live record. That is the silent cross-repo mismatch F23 was meant to close.
+
+Recommended change: Replace L2191's bullet with: "the record round-trips through the canonical fixture tests/upgrade-guarded/fixtures/upgrade_deferred.json in cache.put's envelope. Phase 2's probe test reads a verbatim copy at dotfiles tests/maint/fixtures/upgrade_deferred.json. That test's module docstring names the archsetup path as its source, and any change to the record schema updates both files in the same rollout." Then make L2254 and L2464 refer to "the copy of Phase 1's canonical fixture (see Phase 1 tests)". Don't add a cross-repo comparison test or a provenance key inside the JSON.
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Decisions, Testing. Merged into R56, which names both paths and binds the canonical file to the script's real output. The copy's test docstring names its source, as R40 asks. There is no cross-repo diff check and no provenance key in the JSON, because each copy is already bound to its own side's tests.
+:EVIDENCE:
+Spec:
+- L233: "both test suites read copies of one fixture in maint's envelope"
+- L2191: "the record round-trips through a shared fixture JSON in cache.put's envelope; Phase 2's probe test reads a copy of it."
+- L2254: "the probe reads a copy of Phase 1's fixture record" (says which side is canonical)
+- L2464: "Phase 2's probe test reads a copy of the same fixture"
+- L2163 and L2401: Phase 1 tests live in archsetup tests/upgrade-guarded/
+- L2119-2120 and L2372: full record schema (path, envelope, data fields)
+- L2022: rollout order is archsetup Phase 1, then dotfiles Phase 2, then archsetup Phase 3
+
+The spec gives no fixture filename and no rule for keeping the copies in sync. Grepping the spec for 'copy|copies|re-copy|in step|drift' finds nothing about the fixture beyond the lines above.
+
+Prior review note, L1063: "No dotfiles maint test references the archsetup checkout."
+
+Code:
+- ~/.dotfiles/maint/src/maint/cache.py:35-42 reads the envelope as entry["data"] and entry["written_at"]. The file is plain JSON with no comment syntax.
+- ~/.dotfiles/tests/maint/gen_fixtures.py:4-5 shows the existing way of citing an archsetup source: in a docstring.
+- dotfiles shipped fixtures live in maint/src/maint/fixtures/ (bad.json, good.json). There is no tests/maint/fixtures/ directory yet.
+:END:
+
+** DONE F31: removing the unit while a flag exists leaves it dangling, because the everyday rewrite has no is-enabled check
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design everyday step 9 L91 and arm-flag paragraph L147; Phase 1 everyday step 9 L2073-2077; Readiness Rollout L2422; Testing rollout L2525; Phase 2 row L2229 and REBOOT L2240
+
+Only --complete checks =systemctl is-enabled= (L132, L2094-2095). Everyday step 9 runs whenever the flag is present: it re-runs the arm procedure (resolve check, =sudo pacman -Sw=) and rewrites the flag, with no unit check. L2422 says 'Removing only the unit leaves --complete unable to arm, because it checks is-enabled, so nothing is left dangling.' L2525 says 'Disabling and removing the unit later falls back the same way.'
+
+Risk: If the unit is disabled or removed while armed, every everyday run keeps re-downloading and re-arming a flag nothing consumes. The panel then shows 'armed' and offers a REBOOT that applies nothing, which is F31's misleading-reboot outcome in the rollback state, and it lasts until someone runs --complete by hand. It isn't unsafe, but L2422's 'nothing is left dangling' is false in that state.
+
+Recommended change: Preferred (behavior fix):
+- Phase 1 step 9 (L2073): add a first case. "Unit not enabled (systemctl is-enabled --quiet archsetup-boot-upgrade.service fails; seam UPGRADE_GUARDED_UNIT_ENABLED): sudo rm -f the flag and print 'disarmed: boot unit not enabled'. No failed_step; the exit code is unaffected; the other cases are skipped."
+- Mirror it in one clause each:
+ - Design step 9 (L91).
+ - The arm-flag paragraph (L147).
+ - The writer contract (L2116): "an everyday run removes it when the unit isn't enabled".
+ - Decision "While armed" (L248).
+- Add a test bullet beside L2202 and L2493: "with the flag present and the unit reported not enabled, an everyday run removes the flag, prints the disarmed line and never runs -Sp/-Sw." With that in place, L2422 and L2525 are accurate as written.
+
+Docs-only alternative: reword L2422 to: "Removing only the unit while a flag exists leaves the flag, and everyday runs keep rewriting it, until rollback step 3 or a --complete run removes it, so remove the flag with the unit." Drop "Disabling and removing the unit later falls back the same way" from L2525, or add "after removing any flag" to it.
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Testing.
+:EVIDENCE:
+- Spec L91 and L2073-2077: the step-9 cases are "no GPU-kind entries remain" (remove), "otherwise" (arm procedure and rewrite), closure refused, and resolve/-Sw failure. None of them checks the unit.
+- Spec L147: "While armed, every everyday run whose refresh succeeded rewrites or removes the flag". L2116 says "Everyday runs rewrite or remove it while armed", and L248 has the same wording. None of these mentions is-enabled.
+- Spec L132 and L2094-2095: only --complete gates on "systemctl is-enabled --quiet archsetup-boot-upgrade.service" (seam UPGRADE_GUARDED_UNIT_ENABLED), and it removes any stale flag when the unit is off.
+- Spec L2229: the deferred row reads "'armed — N apply at next boot' while the flag exists". L2240: "REBOOT ... is shown while the flag exists".
+- Spec L2419-2422: rollback runs in this order: (1) revert dotfiles, (2) disable and remove the unit, (3) remove the flag. Then: "Removing only the unit leaves --complete unable to arm, because it checks is-enabled, so nothing is left dangling."
+- Spec L2525: "Disabling and removing the unit later falls back the same way", which describes --complete only.
+- Spec L1519 (prior review): "With F31 in place, removing only the unit no longer leaves --complete arming a flag that nothing reads". That disposition covers new arming, not a flag that already exists.
+- Live (ratio): systemctl cat archsetup-boot-upgrade.service prints "No files found", so the unit doesn't exist yet. Nothing in the code contradicts the spec reading.
+:END:
+
+** DONE F25: a failed stamp has no defined path into the record, because every mode writes the record before stamping
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Stamp predicate L113 and L2136; everyday step 10 L92 and L2078; --complete step L133 and L2096; --apply-armed step L164 and L2107; Phase 1 test L2190 and L2505; Phase 2 row severity L2228
+
+Every mode's last step is 'Write the record, stamp per the stamp predicate, and exit' or 'Write the record and stamp', and the predicate itself reads the written record ('the record ends empty'). A missing maint or a failed stamp then 'prints stamp failed: <reason>, sets failed_step stamp and exits 1'. failed_step is a record field, and nothing says the record is rewritten after the stamp, or which result value goes with failed_step stamp.
+
+Risk: An implementer who follows the stated order has to decide whether to rewrite the record. If it isn't rewritten, the run exits 1 but the record says result ok with failed_step null. The deferred row then stays OK while freshness never clears, which is the panel half of the silent stamp failure F25 set out to close. The exit code, the stderr line and, at boot, the failed unit still show it.
+
+Recommended change: In the stamp predicate's first bullet at L2136, and in the matching sentence at L113, replace "sets =failed_step= stamp and exits 1" with "rewrites the record with result failed, =failed_step= stamp and that line as detail (packages, gate and news unchanged), and exits 1". In the L2190 test bullet and in AC8 at L2505, add "and the record reads result failed, failed_step stamp". No change to the sequence steps is needed.
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Testing. Rather than a second record write after a failed stamp, every mode's finish computes the record, stamps when the predicate holds, folds a stamp failure into the computed record, and writes once. The outcome the finding asks for is unchanged with one write, and the predicate reads the computed record rather than the file.
+:EVIDENCE:
+- Spec L113 and L2136: "A missing maint, or a failed stamp, prints 'stamp failed: <reason>', sets =failed_step= stamp and exits 1."
+- Spec L106 and L2125: failed_step is a record data field, and its enum includes stamp.
+- Spec L2378 (Readiness dimensions, Errors): "Exit codes: 0 ok, 1 a step failed, ... On any non-zero exit the last stderr line names the reason, and the record's =result=, =failed_step= and =detail= carry it to the panel."
+- Spec L2380 repeats the stamp-failure rule: it sets failed_step stamp and exits 1.
+- Spec L2078 "10. Write the record, stamp per the predicate below, and exit." L2096 "6. Write the record and stamp." L2107 "rewrite the record ... Stamp per the predicate below and exit". These sequence lines say nothing about rewriting the record after a stamp failure.
+- Spec L2072 and L2081 show the same implicit convention for other steps: "On failure: =failed_step= refresh, exit 1", with no result value given. AC10 at L2507 expects failed_step aur to grade WARN, which relies on result failed being implied. The WARN rule at L2228 grades on result failed, refused or interrupted.
+- Spec L2190 and L2505: the stamp test asserts failed_step stamp but not result.
+- ~/.dotfiles/maint/src/maint/cli.py:228 and 292-295: maint stamp topgrade exists and can fail like any other subprocess.
+:END:
+
+** DONE F33: 'Removing only the unit ... nothing is left dangling' is false while a flag is armed
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Readiness rollback L2422; everyday step 9 L91 and L2073-2077; arm-flag rules L147; REBOOT rule L2240; armed row text L2229
+
+L2422 says: "Removing only the unit leaves --complete unable to arm, because it checks is-enabled, so nothing is left dangling."
+
+Risk: If the unit is removed while armed (a partial rollback, or between steps 2 and 3 of the full one), the panel keeps advertising 'armed — N apply at next boot' and offering REBOOT. Every everyday run keeps re-downloading the set, and a reboot applies nothing. That is the misleading-REBOOT state F34 closed. Nothing is unsafe, and a later --complete or an empty GPU set removes the flag.
+
+Recommended change: Smallest edit is to reword L2422's first sentence to: "Removing only the unit leaves --complete unable to arm, because it checks is-enabled. Remove any present arm flag at the same time: everyday runs keep an existing flag in step without checking the unit, so the panel would keep showing 'armed — N apply at next boot' and REBOOT with no unit left to consume it." An alternative that also covers a unit disabled by any other means is to add a first bullet to everyday step 9 (L2073, mirrored at L91): "the unit not enabled (systemctl is-enabled, seam UPGRADE_GUARDED_UNIT_ENABLED): sudo rm -f the flag and print 'disarmed: boot unit not enabled'". It would come with a matching Phase 1 test. Then the L2422 sentence would be true as written.
+
+Optional, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Readiness dimensions. Resolved by R41's behaviour fix, which is the finding's own alternative. It makes the existing rollback sentence true, so no reword is needed.
+:EVIDENCE:
+- Spec L2422: "Removing only the unit leaves =--complete= unable to arm, because it checks =is-enabled=, so nothing is left dangling."
+- Spec L2073-2077 (Phase 1 everyday step 9): "While armed (flag present) and step 4 succeeded: no GPU-kind entries remain: sudo rm -f the flag; otherwise: run the arm procedure (below) on the current GPU-kind entries and rewrite the flag". There is no is-enabled condition. L91 has the same wording in Design.
+- Spec L2112-2115: the arm procedure is (a) a resolve check, (b) -Sw, (c) an atomic write. It has no unit check.
+- Spec L2116: "--complete's arm step writes the flag, only when the unit is enabled. Everyday runs rewrite or remove it while armed." The is-enabled gate is scoped to --complete only.
+- Spec L2229: "'armed — N apply at next boot' while the flag exists". L2240: REBOOT "is shown while the flag exists".
+- Spec L150: the unit's ConditionPathExists is the only consumer of the flag. Once the unit is removed or disabled, nothing reads it at boot.
+- Mitigating factors: L2418-2421 (the full rollback reverts the dotfiles first and removes the flag in step 3), and L2095 (--complete with the unit not enabled removes any stale flag).
+:END:
+
+** DONE F38: Keying the strix addendum on --complete misses linux-firmware-only days
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 4 docs L2313 ("Step 5's ratio strix addendum keys on --complete instead of 'before topgrade'"); Design 'The sets' L71 and Phase 1 L2038 (linux-firmware* is not a kernel)
+
+The F38 resolution is present at L2313, including step 6 -> upgrade-guarded, the --complete days, the hand-stamp rewrite and the KIL note. Step 5's trigger is moved to --complete for every trigger.
+
+Risk: On a day when only linux-firmware is pending, the addendum keyed on --complete never runs before that package lands. The watch then misses its most likely signal until the next kernel day. The addendum is research rather than a gate, so nothing is unsafe.
+
+Recommended change: Spec L2313: replace "Step 5's ratio strix addendum keys on =--complete= instead of \"before topgrade\"." with "Step 5's ratio strix addendum keeps its triggers and runs before whichever =upgrade-guarded= invocation lands each one: before =--complete= for linux, linux-lts or a held mesa, and before step 6's everyday run for linux-firmware, which is never held."
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases.
+:EVIDENCE:
+- Spec L2313 (Phase 4 docs): "Step 5's ratio strix addendum keys on =--complete= instead of \"before topgrade\"." It has no per-trigger split.
+- Spec L71 and L2038: "=linux-firmware*=, =linux-api-headers= and other =linux*= names are not kernels."
+- Spec L2167 and L2463: tests assert "firmware and api-headers are never held".
+- Spec L1708 (prior review evidence) says step 5 runs "before topgrade" when linux or linux-lts is pending. It omits linux-firmware and mesa.
+- docs/workflows/system-health-check.org:235: step 5 fires "if =linux=, =linux-lts=, =linux-firmware=, or a major =mesa= bump is pending ... before topgrade".
+- docs/workflows/strix-soak-watch.org:39 says linux-firmware is "the most likely vector for the real fix". At :68-72 the watch reads the SMC version "BEFORE topgrade" and examines the delta "After topgrade installs the new linux-firmware".
+- /etc/pacman.d/hooks/hypr-live-update-guard.hook and archsetup:2629-2643: the Targets are mesa, mesa-*, wayland, libdrm, libglvnd, the hypr* packages, vulkan-*, nvidia-utils, lib32-nvidia-utils and xorg-xwayland. No linux-firmware.
+- pacman -Qi linux / linux-lts: Depends On is coreutils, initramfs and kmod. linux-firmware is required only by mkinitcpio-firmware, so the closure never adds it.
+:END:
+
+** DONE The failed_step vocabulary has no value for several refusals that still write the record
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design record L106; Phase 1 record L2125; everyday precondition L2060 (Design L83); --apply-armed step 1 L2099 (Design L157); Readiness L2379
+
+failed_step is 'null, hook, refresh, closure, pacman, aur, topgrade, gate, download, boot-transaction or stamp'. Only the EUID 0 and lock-held refusals skip the record write (L102, L2120). Every other refusal writes result refused.
+
+Risk: Implementers will choose null, invent a value, or reuse boot-transaction differently. The panel and the shared fixture record that both repos read (Decision 4 L233) can then disagree on the vocabulary.
+
+Recommended change: L106 and L2125: append "precondition" to the failed_step list. Then add one sentence after the list in each place: "A precondition refusal that writes the record (db.lck present; for --apply-armed, a missing, empty or unparseable list, a live Hyprland or db.lck) sets failed_step precondition. --complete's refusal for a kernel-set or DKMS-set member in the GPU closure sets failed_step closure. An interruption during news capture or informant records failed_step precondition."
+
+Testing, Phase 1 (L2494 and the preconditions line L2510): add "db.lck present exits 3 with result refused and failed_step precondition, and the file is not deleted". Change "each exit 3" for the --apply-armed list and Hyprland cases to "each exit 3 with result refused and failed_step precondition". The held-lock case keeps its existing "record unchanged" assertion.
+
+A smaller alternative: one sentence saying "any refusal not assigned a step above records failed_step null, with the reason in detail".
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases, Acceptance criteria, Testing. I adopted the precondition value and folded the existing 'hook' value into it. A missing or Target-less hook is a precondition refusal like db.lck, and the detail names the path. Nothing in either repo keys on 'hook', so the vocabulary shrinks instead of growing.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L106 and L2125: "failed_step: null, hook, refresh, closure, pacman, aur, topgrade, gate, download, boot-transaction or stamp."
+- L102 and L2120: the record is rewritten at the end of every mode "refusals, failures and interruptions included", with only the EUID 0 and lock-held refusals exempt.
+- L2060: db.lck refusal, exit 3 with detail. No failed_step. L2081: --complete reuses the everyday preconditions.
+- L2082: --complete refuses if the GPU closure holds a kernel-set or DKMS-set member. No failed_step (contrast L2051, where the closure-loop refusal is named closure).
+- L2099: --apply-armed exits 3 for a missing, empty or unparseable list, a live Hyprland, a held lock, or db.lck. No failed_step. L2100 pre-writes boot-transaction only after this step. L2105 names boot-transaction for the resolve-check refusal only.
+- L2132 and L2183: an interrupt records "the step". News (step 2) and informant (step 3) have no enum value.
+- L2494: the Phase 1 test for the --apply-armed preconditions asserts only "each exit 3", not the record. No automated test covers the db.lck refusal (db.lck appears only at L2354 and L2517).
+- Phase 2 severity, L2234-2237: WARN when result is failed, refused or interrupted. L2261 is the only Phase 2 test that keys on failed_step (aur).
+- No implementation exists yet: there is no scripts/upgrade-guarded in archsetup, and grep finds no failed_step in ~/.dotfiles/maint/src.
+:END:
+
+** DONE 'N' means three different counts across the deferred row, the armed text and the strip, and the armed count reads the flag's raw line count
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L122, L142; Phase 2 L2224, L2229, L2234, L2244; AC5 L2349
+
+Row value N is the record's packages minus installed entries (L2224). In the armed state, 'armed — N apply at next boot' uses 'N the flag's line count' (L2229). The strip reads 'N pending · M held' (Design L122, L2244), where N is the pending count and M is the deferred count. APPLY is 'shown while N > 0' (L2234). The flag parser ignores '#' lines (L142), but the probe 'reads only its presence and line count'.
+
+Risk: Small UI miscounts, and it is unclear which N gates APPLY.
+
+Recommended change: At L2229 (mirror at L124 and L2495), replace "with N the flag's line count" with "with A the number of flag entries (non-'#' lines) not yet installed at or above their version (=pacman -Q= plus =vercmp=, as for N); when the row's N exceeds A, append ' · <N−A> more deferred'". At L2223, L257 and L2373, replace "presence and line count" with "presence and entries". At L2234, write "shown while the row's value N > 0". Optionally, at L2244, define M as the row's value N, and say whether 'N pending' includes the held share.
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Acceptance criteria, Testing.
+:EVIDENCE:
+Spec L142 and L2111: the flag carries one <name>=<version> line per GPU-kind entry and '#' lines are ignored. L2116: "=upgrade-guarded= is the only writer". The arm procedure (L2112-2115) writes no comment lines.
+
+L2224: "Value N: the record's =packages= minus the entries now installed at or above =new=".
+L2229: "'armed — N apply at next boot' while the flag exists, with N the flag's line count".
+L2234: APPLY "shown while N > 0 or a gate failure is open".
+L2231: "A kernel sits in the deferred set most days".
+L91 and L2073-2075: an everyday run while armed rewrites the flag from "the current GPU-kind entries" only.
+L2438: a bare pacman -Sy or yay can refresh without re-arming.
+L2104: the boot form drops entries already at or above the recorded version, so the flag's line count can exceed what actually applies.
+L2244: "the strip's pending cell reads 'N pending · M held'. Both read the record."
+maint doctor.py:56 pending_updates() is the existing source of the strip's pending count, which is distinct from the record.
+:END:
+
+** DONE The wall note 'freshness clears on --complete' contradicts the rule that an arming --complete never stamps
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L122; Phase 2 L2239; AC9 L2362; Phase 2 test L2260 vs Decision 4 L229, Design L116
+
+After a TOPGRADE that left anything deferred, the wall reads 'N deferred — freshness clears on --complete'.
+
+Risk: The note promises a result the spec's own predicate withholds in the most common case.
+
+Recommended change: Change the note in all five places that state or assert it (L122, L2239, L2260, L2362, L2506) to one string that holds for both kinds of deferral: 'N deferred — freshness clears once they land (APPLY)'. If you'd rather keep the --complete wording for the common case, use two strings instead:
+- When the record holds no gpu-kind entry: 'N deferred — freshness clears on --complete'.
+- When it holds any gpu-kind entry: 'N deferred — freshness clears after APPLY and the next boot'.
+
+In that case, add one Phase 2 test line next to L2260 asserting the gpu-kind variant.
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Design, Acceptance criteria, Testing.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L116: "=--complete= stamps only when ... the record ends empty ... An arming run never stamps."
+- L117: "=--apply-armed= stamps when its transaction exited 0 or was a no-op, and the record ends empty."
+- L126: on the boot run, "The stamp is written when nothing is left deferred."
+- L132 (step 4): in the live case with the unit enabled, --complete runs the arm procedure. With the unit not enabled it removes any stale flag and writes none, so the record stays non-empty.
+- L229: "A run that arms the boot oneshot never stamps."
+- The note is attached to any deferral at L122 ("after a TOPGRADE that left anything deferred") and L2239 ("After a TOPGRADE press while anything is deferred").
+- L240: "the panel's deferred count carries a kernel most days". The common case is kernel-only, where --complete does stamp.
+- L2362 (AC9) and L2506 (AC9 test) pin the note only for the kernel-only case. L2260 asserts it for any non-empty record.
+- L2234: the row reads 'armed — N apply at next boot' while the flag exists, which corrects the user after an arm.
+:END:
+
+** DONE Flag handling is unspecified when --complete refreshes the db and then refuses or fails stage 1
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L129, L131, L147; Phase 1 L2082, L2089, L2116; Risks L2438
+
+--complete runs its own 'sudo pacman -Sy' (L2082). It removes the flag on a gate failure, an unenabled unit, nothing GPU-side left, or nothing to complete, and it re-arms when live and enabled. For a GPU-closure kernel refusal or a closure-loop refusal it says nothing about the flag, and for a stage-1 failure it says only 'nothing armed' (L2089).
+
+Risk: Low, because the boot form fails before any package changes when a version is gone, but the spec is internally inconsistent about who keeps the flag in step.
+
+Recommended change: Make --complete mirror everyday step 9 after a successful refresh:
+- Append to --complete step 2 (L2082, and its twin in the Design text at L129): "A refusal here removes any existing flag with sudo rm -f and prints 'disarmed: closure refused' before the refusal's own lines, as everyday step 9 does."
+- At L2089 and L131, change "nothing armed" to "removes any stale flag and arms nothing".
+- Extend the --complete removal list at L147 and L2101 to "on a refusal after its refresh, a failed stage 1, a gate failure, when the unit isn't enabled, or when nothing GPU-side is left".
+- Add a Phase 1 test bullet next to L2474/L2492: "with a flag present, a --complete closure refusal and a failed stage 1 each remove it."
+
+If the intent is to leave the flag in place, add "a refusing --complete" to L2438's list of refreshes that don't re-arm instead, and say "leaves any existing flag untouched" at L2082/L2089.
+
+Optional, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Risks, Testing. Rather than adding refusal and stage-1-failure cases to --complete's list of flag removals, --complete removes any flag once, right after its refresh succeeds, and only its arm step writes a new one. That one rule covers a closure refusal, a failed stage 1, a failed gate, a unit that isn't enabled, nothing GPU-side, and nothing to complete. With R03 it also guarantees that no flag exists while the gate is open.
+:EVIDENCE:
+- Spec L91 and L2076: everyday step 9, "the closure refused: remove the flag and print 'disarmed: closure refused' before the refusal's own lines ... exits 3".
+- Spec L129 and L2082: --complete runs "sudo pacman -Sy", then "compute the pending set and the GPU closure, refusing if the closure holds a kernel-set or DKMS-set member". Nothing about the flag on refusal. Only the nothing-pending branch (L2083) says "remove any stale flag".
+- Spec L131 and L2089: "Stage 1 failed while the gate passed or had nothing to check: failed_step pacman, exit 1, nothing armed."
+- Spec L147: "While armed, every everyday run whose refresh succeeded rewrites or removes the flag (step 9 above), so the armed versions follow the db that run synced. --complete removes the flag on a gate failure, when the unit isn't enabled, or when nothing GPU-side is left ... Re-arming means running --complete again."
+- Spec L2101 repeats the same three-case list.
+- Spec L2438: "what breaks the agreement is a manual cache clean, or a db refresh that didn't re-arm (a bare pacman -Sy or yay)". --complete is not named.
+- Spec L2474: the --complete GPU-closure refusal is tested to "refuse the same way", with no flag assertion. The flag-on-refusal tests at L2202 and L2493 are everyday-only.
+- Why it is safe: L2111 (the boot form's resolve check refuses a kernel ride-in) and L2112 ("A version the db no longer carries fails as target not found ... both fail before any package changes").
+:END:
+
+** DONE Phase 2 removes the lever stamp and --force but leaves code and doc text that still describes them
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 2 L2216, L2219, L2248; Documentation plan L2397
+
+Phase 2 rewrites only the PATH wrapper's header comment, tests/maint/test_topgrade_wrapper.py, and the 'The live-update guard' README section.
+
+Risk: Leftover text from the replaced design tells the user that a successful TOPGRADE stamps, and documents a flag that no longer exists.
+
+Recommended change: Append one bullet to Phase 2's Levers or Guard UX list (after L2219):
+
+"Also reword the leftover text from the replaced design:
+- the topgrade_freshness docstring and its 'no topgrade run recorded' error in probes/updates.py:115-127, so they name upgrade-guarded (current runs only) and the PATH wrapper (bare-shell runs) as the writers, with no doctor or TOPGRADE-lever stamp;
+- the doctor.py module docstring (L28-30) and guard.py's docstring (L5-7), so neither describes a refusal or --force;
+- the =maint fix= usage line at maint/README.md:23, dropping [--force]."
+
+In the Documentation plan's Phase 2 bullet (L2397), add: "and drops [--force] from the =maint fix= usage line at README.md:23."
+
+Optional, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Readiness dimensions. Adopted, plus two items that went stale the same way. The topgrade PATH wrapper's header says the maint TOPGRADE lever resolves the wrapper and that the system-health-check workflow runs through it. test_topgrade_wrapper.py's docstring repeats the workflow claim. After the change, the lever runs upgrade-guarded, which calls /usr/bin/topgrade directly, and the workflow's update step runs upgrade-guarded.
+:EVIDENCE:
+- Spec L2216 names only "the PATH wrapper ~/.local/bin/topgrade and tests/maint/test_topgrade_wrapper.py" for rewording. L2248 and L2397 limit the README rewrite to "the 'The live-update guard' section".
+- ~/.dotfiles/maint/src/maint/probes/updates.py:116-118: "The topgrade PATH wrapper stamps the cache on any successful run ... and the TOPGRADE lever stamps via doctor"
+- updates.py:125-126: error="no topgrade run recorded — any successful run stamps it (PATH wrapper / TOPGRADE lever)"
+- ~/.dotfiles/maint/README.md:23: "maint fix <id> [items...] [--dry-run] [--force]". The guard section starts at README.md:52 ("### The live-update guard").
+- ~/.dotfiles/maint/src/maint/doctor.py:28-30: "the live-update guard refuses UPDATE / TOPGRADE when the pending set trips it (`=--force=` is the CLI's "press again")"
+- ~/.dotfiles/maint/src/maint/guard.py:5-7: "tells the UPDATE/TOPGRADE key to arm with "press again to run anyway — or apply from a TTY" (the CLI mirrors that with `=--force=`)"
+- ~/.dotfiles/maint/src/maint/doctor.py:169-170 is the stamp that L2216 deletes, so after Phase 2 "stamps via doctor" is false.
+:END:
+
+** DONE Rollback is labelled 'reverse order' but runs Phase 2, then 3, then 1
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Readiness, Rollout L2418-L2421
+
+'Roll back in reverse order: 1. Revert the dotfiles Phase 2 commit. 2. disable and remove the unit... 3. Remove the arm flag, then the scripts.'
+
+Risk: Cosmetic, but someone following the label instead of the list would disable the unit before reverting dotfiles.
+
+Recommended change: At L2418, replace "Roll back in reverse order:" with "Roll back in this order (dotfiles first, so nothing calls the scripts once they are removed; steps 1 and 2 are independent of each other):". Leave the numbered steps unchanged.
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Readiness dimensions.
+:EVIDENCE:
+Spec file: docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org
+- L2405-L2417: rollout order (a) after the Phase 1 commit (install the scripts), (b) only after (a), pull the Phase 2 dotfiles commit, (c) after the Phase 3 commit (write, daemon-reload and enable archsetup-boot-upgrade.service).
+- L2418-L2421: "Roll back in reverse order: 1. Revert the dotfiles Phase 2 commit ... 2. sudo systemctl disable archsetup-boot-upgrade.service, remove the unit ... 3. Remove the arm flag, then /usr/local/bin/upgrade-guarded and kernel-modules-check." That order is b, c, a, not c, b, a.
+- L2422: "Removing only the unit leaves --complete unable to arm, because it checks is-enabled, so nothing is left dangling. Removing the scripts before reverting dotfiles leaves UPDATE and TOPGRADE pointing at a missing binary." So disabling the unit before reverting dotfiles is explicitly safe, which refutes the finding's stated risk.
+- L2382: "With the boot unit not enabled, --complete prints the console remedy and writes no flag."
+- L1511-L1514 (prior finding's recommended text): "Roll back in reverse rollout order: 1. Revert the dotfiles Phase 2/3 commits first ..." The label dates from when Phase 3 was still grouped with dotfiles.
+:END:
+
+** DONE =flock -n <lockfile>= as written is a usage error and holds no lock
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L83 (everyday precondition); Phase 1 L2059; --apply-armed L157 and L2099 ('the lock is held'); test L2182
+
+Precondition: '=flock -n ${MAINT_STATE_DIR:-$HOME/.local/state/maint}/upgrade-guarded.lock= succeeds, else 'another upgrade-guarded run is in progress''.
+
+Risk: Copied literally, the precondition never succeeds, so every run refuses as 'in progress'. A working one-shot =flock file true= would release the lock immediately and guard nothing. The spec doesn't say the lock is held for the whole run, which is what makes the lock-held refusals and the 'record untouched' rule meaningful. The tests use the real flock, so the literal form would fail on the first test run, but the lifetime still has to be invented.
+
+Recommended change: Make the same edit at L83 and L2059, and have L2099 refer to it. Replace "=flock -n <path>/upgrade-guarded.lock= succeeds" with:
+
+"take an exclusive, non-blocking flock on =${MAINT_STATE_DIR:-$HOME/.local/state/maint}/upgrade-guarded.lock= (creating the directory and the file if missing) through a descriptor held open until the process exits, so the lock covers the whole run (in sh: =exec 9>>\"$lock\"; flock -n 9=; in Python: =fcntl.flock(fd, LOCK_EX|LOCK_NB)=); if the lock is unavailable, refuse with 'another upgrade-guarded run is in progress'".
+
+At L2376, add "held for the life of each non-dry-run invocation". Optionally, make the L2182 test take the lock with a real flock held by the test process. It should then assert both exit 3 and that a normal run with the lock free gets past this precondition.
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Readiness dimensions, Testing.
+:EVIDENCE:
+Command run: bash -c 'flock -n <scratch>/t.lock; echo rc=$?' printed "flock: bad file descriptor: '<scratch>/t.lock'" and rc=64.
+
+flock --version reports util-linux 2.42.4. Its usage lines accept only "flock [options] <file>|<directory> <command> [<argument>...]", "flock [options] <file>|<directory> -c <command>" or "flock [options] <file descriptor number>".
+
+Spec lines:
+- L83 and L2059: "flock -n ${MAINT_STATE_DIR:-$HOME/.local/state/maint}/upgrade-guarded.lock succeeds, else 'another upgrade-guarded run is in progress'".
+- L2079: --complete reuses everyday steps 1 to 3.
+- L2099: --apply-armed refuses if "the lock is held".
+- L2182 and L2510: "a held upgrade-guarded.lock exits 3 and leaves the record unchanged".
+- L2376: "upgrade-guarded.lock, beside the record, is a flock file and holds no data".
+- L2430: external deps include "flock and pgrep".
+
+A grep of the spec for lock-lifetime wording (held for/until/through, lifetime, for the whole/life/duration, exec 9, file descriptor, fcntl) returns no hits. The spec states no implementation language. The sibling script scripts/hypr-live-update-guard is #!/bin/sh and the tests are Python unittest.
+:END:
+
+** DONE yay -Pwq returns date-prefixed lines, oldest first, not bare titles
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L84 and L109 (news: 'titles merged ... deduplicated, keeping the newest 10'); Phase 1 L2062 and L2128; Phase 2 DISMISS (L2225)
+
+The record's news field holds the 'unread Arch news titles' from =yay -Pwq=, merged with the previous record's, deduplicated, with the newest 10 kept.
+
+Risk: 'Newest 10' depends on knowing the order, and dedup and DISMISS matching depend on whether the date prefix is kept. Phase 1 writes the shared fixture and Phase 2 reads a copy, so the two repos can disagree on the entry shape.
+
+Recommended change: At L109 and L2128, replace the news bullet with: "=news=: this run's =yay -Pwq= lines, each kept verbatim as =YYYY-MM-DD <title>= (yay prints them oldest first and the date is part of the line), merged with the previous record's, deduplicated on the whole line, sorted by the date prefix, keeping the 10 latest. An empty output with exit 0 means no new news."
+
+At L2225 and L2500, add: "DISMISS stores the same verbatim line, and the row filters on whole-line equality."
+
+At L2499, change "cap at the newest 10" to "cap at the 10 latest by date prefix". Also make the fake yay emit date-prefixed lines, oldest first.
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Testing.
+:EVIDENCE:
+Spec: L109 and L2128, "=news=: this run's =yay -Pwq= titles merged with the previous record's, deduplicated, keeping the newest 10." L2499: "titles merge with the previous record's, deduplicate, and cap at the newest 10." L2225: "the news titles minus those in maint key =upgrade_news_dismissed=". There is no other definition of a news entry anywhere in the spec.
+
+Live check on ratio. yay --version reports "yay v13.0.1 - libalpm v16.0.1". The command yay -Pwwq printed 10 lines, oldest first, from "2025-11-05 waydroid >= 1.5.4-3 update may require manual intervention" through "2026-09-22 Mkinitcpio >=42 requires manual intervention for TPM2-based unlocking of LUKS devices", with rc=0. The command yay -Pwq printed nothing, also with rc=0.
+
+The yay(8) man page reads "-q, --quiet Only show titles when printing news" and "News is considered new if it is newer than the build date of all native packages."
+
+maint has no news code yet: grep -rn news over ~/.dotfiles/maint/src/maint finds only a comment at gui.py:93.
+:END:
+
+** DONE --dry-run should own its checkupdates db rather than share the default path with maint
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L80 and Phase 1 L2036 (--dry-run: pending set from checkupdates, closure via =pacman -Sup --dbpath <checkupdates' db>=); test L2184
+
+--dry-run takes its pending set from =checkupdates= and runs its closure against 'checkupdates' db', without saying how the script learns that path.
+
+Risk: The script would have to re-derive checkupdates' default path from TMPDIR and UID. A maint probe that refreshes the shared db between the dry-run's checkupdates read and its -Sup probes can make the pending set and the probe db disagree, and the preview would then show a spurious 'X not in pending set' refusal.
+
+Recommended change: L80 and L2036: replace "runs its closure with =pacman -Sup --dbpath <checkupdates' db>=" with "creates a private directory (=mktemp -d=, removed on exit), exports =CHECKUPDATES_DB= to it, runs =checkupdates= against it, and passes that same path to every closure probe's =pacman -Sup --dbpath=; it never uses the shared default =${TMPDIR:-/tmp}/checkup-db-${UID}=". L2184 and L2475: change "its closure passes =--dbpath= with checkupdates' db" to "every closure probe passes =--dbpath= equal to the =CHECKUPDATES_DB= the fake checkupdates received, and that path is not the shared default".
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Testing.
+:EVIDENCE:
+- Spec L80 and L2036: "runs its closure with pacman -Sup --dbpath <checkupdates' db>", with no mechanism for obtaining the path. L2184 and L2475 assert "--dbpath with checkupdates' db", with no defined value. L2048 refuses when X is not in the pending set.
+- /usr/bin/checkupdates:132-133: CHECKUPDATES_DB defaults to "${TMPDIR:-/tmp}/checkup-db-${UID}/" only when unset, so it is overridable from the environment.
+- /usr/bin/checkupdates:136: trap 'rm -f $CHECKUPDATES_DB/db.lck' INT TERM EXIT. Each run deletes the lock in the shared dir on exit.
+- /usr/bin/checkupdates:150: fakeroot pacman -Sy --dbpath "$CHECKUPDATES_DB". Line 181: exit 2 when nothing is pending.
+- maint calls bare checkupdates against the same default db at ~/.dotfiles/maint/src/maint/probes/updates.py:152 and ~/.dotfiles/maint/src/maint/doctor.py:83.
+- Live system: /tmp/checkup-db-1000 has mtime 09:09. maint-net-scan.timer last ran at 09:08:45 and runs hourly, which confirms the default db is shared and refreshed on a schedule.
+:END:
+
+** DONE The new post-rebuild-check assertion has no seam, so the Phase 1 commit's own test gate goes red on ratio until the rollout step that follows it
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 1 L2161 (post-rebuild-check hook-name assertion), L2208 (its test); Implementation phases L2022; Rollout step (a) L2315-2322
+
+post-rebuild-check gains a check keyed on whether /usr/local/bin/hypr-live-update-guard exists, asserting that /etc/pacman.d/hooks/10-hypr-live-update-guard.hook exists and the unprefixed name doesn't. The spec names no seam for either path. Ratio's hook migration is rollout step (a), which the spec orders after the Phase 1 commit.
+
+Risk: The unseamed check reads ratio's real /etc and fails all 32 all-clean tests. So 'make test-unit' can't pass on ratio before the commit, and ratio can only be fixed after it. Either the commit lands red or the machine is migrated out of order. Adding a ninth check also moves every 'check N/8' assertion.
+
+Recommended change: At the end of the L2161 bullet, add: "It follows the script's seam convention. PRC_GUARD_BIN overrides the guard path (set-but-empty = not installed), and PRC_PACMAN_HOOKS_DIR overrides /etc/pacman.d/hooks. Every test env, run_check's defaults included, sets PRC_GUARD_BIN empty so the suite never reads the real /etc. It is check 9, so the header, the usage text and the 'check N/8' counters become 9."
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Testing.
+:EVIDENCE:
+- Spec L2161 is the assertion. It names no seam. A grep for PRC_ in the spec returns nothing.
+- Spec L2163 names UPGRADE_GUARDED_* and KMC_MODULES_DIR / KMC_BOOT_DIR for the other two scripts.
+- Spec L2208: "tests/post-rebuild-check/ covers the hook-name assertion, and the check is skipped when the guard isn't installed". The guard is installed on both daily drivers, so this test needs a guard-path seam.
+- scripts/post-rebuild-check:55-82 documents an env seam for every probe ("for each, set-but-empty means ... unset means run the real probe"). PRC_IDLE_DAEMON at 72-73 already uses the installed-or-not MISSING shape.
+- tests/post-rebuild-check/test_post_rebuild_check.py:54-77: run_check sets all nine PRC_ seams to clean defaults. Other tests build their own env dicts at lines 167, 213, 368, 912, 930, 955, 983, 1008 and 1065. 32 lines assert returncode 0, and 7 assert '/8' strings.
+- Live ratio (uname -n = ratio): /usr/local/bin/hypr-live-update-guard exists. /etc/pacman.d/hooks holds 99-grub-sync-efi.hook and hypr-live-update-guard.hook, with no 10- file.
+- scripts/post-rebuild-check: "eight checks" at lines 3 and 95; report "check N/8" at 218, 367, 412, 475, 502, 579, 617 and 667.
+- Makefile:58-64: test-unit globs tests/*/test_*.py, so the post-rebuild-check suite runs on every test-unit.
+:END:
+
+** DONE Phase 1 and Phase 2 are each too large to build and verify in one session, and both can be split without a broken intermediate state
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Implementation phases L2022; Phase 1 L2024-2208; Phase 2 L2210-2262; Decision 3 (three code commits)
+
+Phase 1 is one commit with two new scripts, four modes, the closure loop, the arm procedure, the record, the stamp predicate, signal handling, the installer step, the TOML rewrite and the post-rebuild-check change. Its test list runs to 38 bullets, many compound (L2163-2208). Phase 2 is one dotfiles commit.
+
+Risk: A session that can't finish a phase ends with a half-verified commit or one large unreviewed diff. In Phase 1 the risky logic (the closure loop, the gate, the arm flag) shares a single verification pass with everything else.
+
+Recommended change: L2022: replace "Three code commits across two repos, in this order: archsetup (Phase 1), dotfiles (Phase 2), archsetup (Phase 3)." with "Three ordered commit groups across two repos: archsetup (Phase 1), dotfiles (Phase 2), archsetup (Phase 3). A phase may land as several commits if each one is green on its suite (make test-unit, or tests/maint/). The constraints are ordering and coupling, not commit count: Phase 4 step (a) follows Phase 1's last commit; no Phase 2 commit reaches a machine before step (a) there; the deferred row and APPLY land in one commit; the lever repoint, the doctor.py stamp deletion, the guard UX and --force removal, and the sysupgrade alias land in one commit."
+
+L2211: change "One dotfiles commit, pulled" to "The Phase 2 commits, pulled".
+
+L221 (Decision 3): change "the rollout is three code commits" to "the rollout is three ordered commit groups".
+
+L2316, Readiness (a) and L2523: change "After the Phase 1 commit" to "After Phase 1's last commit", and "three code commits" to "three commit groups".
+
+Readiness rollback step 1: change "Revert the dotfiles Phase 2 commit" to "Revert the dotfiles Phase 2 commits, newest first".
+
+Leave L2159 and L2231 as they are; the new L2022 sentence restates them as invariants. Don't hard-code the 1a-1d / 2a-2b split; leave the decomposition to the implementer's approach step.
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Decisions.
+:EVIDENCE:
+Spec (docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org):
+- L2022: "Three code commits across two repos, in this order".
+- L221 (Decision 3): "so the rollout is three code commits".
+- L2211: "One dotfiles commit, pulled on a machine only after Phase 4's step (a)".
+- L2231: "The row and APPLY ship in this same commit."
+- L2159: TOML and hook name "change in the same commit".
+- L2523: "Rollout follows the three code commits in order".
+- L2316 and Readiness step (a): "After the Phase 1 commit".
+- Readiness rollback step 1: "Revert the dotfiles Phase 2 commit."
+
+Size:
+- sed -n 2024,2208p | wc -w gives 4575 words for Phase 1, and Phase 2 (L2210-2262) gives 1231.
+- L2163-2208 has 45 '- ' lines; minus the 5 group headers and the 2 installer-steps/post-rebuild-check bullets, that leaves 38 test bullets, matching the claim.
+
+Coupling the split must keep:
+- dotfiles doctor.py:169-170: =if rid == "topgrade": cache.put("topgrade_run", ...)=. Repointing TOPGRADE without deleting this stamps a deferred run fresh (prior blocking finding at L262).
+- remedies.py:298/308: "guard": "live_update" on update/topgrade, and doctor.py:216/242 refuse or force-wrap guarded remedies. The guard removal has to land with the repoint.
+- L860/L867: the row ships with its lever (APPLY).
+
+Dotfiles touch points confirmed:
+- remedies.py:294-313, panel.py:224-237, viewmodel.py:99-104, gui.py:352/1396-1436/1496 (_update_force, _guard_arm_line), cli.py:265 (--force).
+- tests/maint/test_remedies_doctor.py has 55 case-insensitive guard/force matches (the finding said 37; same direction, not material).
+
+No hazard between sub-commits:
+- L1378 evidence: dotfiles pulls go live immediately (symlinks), and there is no auto-pull timer.
+- L2095: --complete with the unit not enabled arms nothing.
+- Nothing calls upgrade-guarded before Phase 2.
+:END:
+
+** DONE The cross-repo record fixture has no named paths and nothing keeps its two copies in sync
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 1 L2191; Testing L2464; Decision 4 consequences L233
+
+'The record round-trips through a shared fixture JSON in cache.put's envelope; Phase 2's probe test reads a copy of it.' Neither copy's path is named, and nothing compares the copies.
+
+Risk: A later change to the writer in archsetup leaves the dotfiles copy passing against a stale shape. The first sign is the live panel misreading the record.
+
+Recommended change: Replace the L2191 bullet (and the matching sentence at L2464) with: "a fixed fake everyday run's record equals tests/upgrade-guarded/fixtures/upgrade_deferred.json in every field except written_at; Phase 2's probe test reads a byte-identical copy at tests/maint/fixtures/upgrade_deferred.json, and any change to the record's fields updates both copies in the same rollout." This binds the fixture to the script's real output and names both paths. A cross-repo diff check is optional beyond that. If one is wanted, add it to Phase 4 step (a) as "diff the two fixture copies".
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases, Decisions, Testing.
+:EVIDENCE:
+- Spec L2191: "the record round-trips through a shared fixture JSON in cache.put's envelope; Phase 2's probe test reads a copy of it." No path is named.
+- Spec L2254: "the probe reads a copy of Phase 1's fixture record". No path is named.
+- Spec L2464: "The record round-trips through a fixture JSON in cache.put's {written_at, data} envelope, and Phase 2's probe test reads a copy of the same fixture."
+- Spec L233: "a record format two repos must agree on, so both test suites read copies of one fixture in maint's envelope".
+- Spec L2117-2125 (Phase 1 'The record') and L2372 (Data model) fully specify the path, envelope, atomic write and every data field, so an implementer has the shape without the fixture.
+- Spec L1105/L1116: the earlier finding's risk was that "the 'state file round-trips' test only checks the script against itself", and its fix was to point Phase 2 at the same fixture.
+- Rollout L2318-2323 and L2540-2544 have no fixture comparison step.
+- archsetup: =find tests -name '*.json'= returns nothing, so there's no existing fixture convention.
+- dotfiles: maint/src/maint/fixtures/ contains bad.json, good.json and gen_fixtures.py, which are panel board fixtures loaded through panel.load_fixture (tests/maint/test_panel_levers.py:721).
+- dotfiles maint/src/maint/cache.py:31 writes {"written_at": time.time(), "data": data}, and cache.py:40 reads entry["data"] and entry["written_at"].
+:END:
+
+** DONE topgrade_age's absent-stamp text still says the TOPGRADE lever stamps
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Phase 2 Levers L2216
+
+Phase 2's stamp-removal edits name doctor.py:169-170, the wrapper's header comment and tests/maint/test_topgrade_wrapper.py.
+
+Risk: On a fresh machine the row tells the user that any successful TOPGRADE clears it. That's the very misreading the deferred row and the wall note exist to correct.
+
+Recommended change: At spec L2216, after "...and =tests/maint/test_topgrade_wrapper.py=", insert: ", and the =topgrade_freshness= docstring and absent-stamp error in =maint/src/maint/probes/updates.py:115-126=,". Then change the closing clause to "so none of them says the lever stamps through the wrapper or through doctor". Also give the replacement error text: 'no topgrade run recorded — upgrade-guarded stamps it once nothing is left deferred; a bare-shell topgrade stamps through the PATH wrapper'. Optionally, add updates.py to the file list in the Phase 2 done criterion at L2467.
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases.
+:EVIDENCE:
+- Spec L2216: "Delete doctor.py:169-170 ... Rewrite the header comment of the PATH wrapper =~/.local/bin/topgrade= and =tests/maint/test_topgrade_wrapper.py= so neither says the lever stamps through the wrapper." It doesn't mention updates.py.
+- Spec L113-118: the stamp predicate. The everyday run stamps only when the record's packages list is empty, and UPDATE never stamps. L120: "The panel's TOPGRADE lever no longer stamps on its own."
+- Spec L2253 (Phase 2 test): "a zero-exit TOPGRADE lever whose record lists deferred packages leaves =topgrade_run= unchanged".
+- Spec L2467 and L2506: the Phase 2 done criterion and AC9 also leave out the probe text.
+- ~/.dotfiles/maint/src/maint/probes/updates.py:115-118, docstring: "The topgrade PATH wrapper stamps the cache on any successful run (TTY, workflow, shell), and the TOPGRADE lever stamps via doctor".
+- updates.py:122-126, error text: "no topgrade run recorded — any successful run stamps it (PATH wrapper / TOPGRADE lever)".
+- ~/.dotfiles/maint/src/maint/doctor.py:169-170: the =if rid == "topgrade": cache.put("topgrade_run", ...)= branch, which Phase 2 deletes.
+- grep over tests/maint finds no test that pins the updates.py error string.
+:END:
+
+** DONE --no-aur has no caller, yet it widens the stamp predicate and the usage rules
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L80, L89, L115; Phase 1 L2029-2030, L2067, L2138; tests L2181; AC3 L2343; Testing L2510
+
+--no-aur skips the yay step. 'No lever passes it, and it exists for manual runs and tests' (L80, L2030). It is a usage error outside the everyday mode, and the everyday stamp requires that neither --no-topgrade nor --no-aur was given.
+
+Risk: It is a second skip flag that the stamp predicate, AC3, the usage-error rule and the mode tests each have to carry, for a manual convenience with no owner.
+
+Recommended change: Delete --no-aur from the body sections. Leave the historical text in Review findings (L752-771) unchanged.
+- L80 and L2029: change "passing =--no-topgrade= or =--no-aur= outside the everyday run/mode" to "passing =--no-topgrade= outside the everyday run/mode". Delete the sentence "=--no-aur= skips the yay step; no lever passes it, and it exists for manual runs and tests." from L80 and L2030.
+- L89 and L2067: change "7. Unless =--no-aur=: =yay -Sua --noconfirm=." to "7. =yay -Sua --noconfirm=."
+- L115, L228 and L2138: change "neither =--no-topgrade= nor =--no-aur= was given" to "=--no-topgrade= was not given".
+- L2181 and L2510: change "combining modes, or =--no-topgrade= or =--no-aur= outside the everyday mode, exits 2" to "combining modes, or =--no-topgrade= outside the everyday mode, exits 2".
+- L2343 (AC3): change "a TOPGRADE run (neither =--no-topgrade= nor =--no-aur=)" to "a TOPGRADE run (no =--no-topgrade=)".
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases, Acceptance criteria, Testing, Design, Decisions.
+:EVIDENCE:
+Spec references (docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org):
+- L80 and L2030: "=--no-aur= skips the yay step; no lever passes it, and it exists for manual runs and tests."
+- L89 and L2067: "7. Unless =--no-aur=: =yay -Sua --noconfirm=. ... On failure: failed_step aur, and step 8 still runs."
+- L115, L228 (Decision 4, not in the finding's list) and L2138: the everyday run stamps only when neither --no-topgrade nor --no-aur was given.
+- L2163: the fakes on PATH include yay.
+- L2180: "a failing fake yay: the sweep still runs, the record says failed_step aur, nothing stamps, and the exit is 1".
+- L2181 and L2510: the only test that mentions --no-aur is the usage-error test.
+- L2343 (AC3): "a TOPGRADE run (neither --no-topgrade nor --no-aur)".
+- L2247 (sysupgrade alias) and L2313 (system-health-check step 6) pass no flags.
+- L752-760: the prior review asked for a consumer for --no-topgrade and --no-aur. Only --no-topgrade got one: UPDATE.
+
+Code: grep -rn 'no-aur' over the dotfiles maint/ tree, hyprland/.local/bin, and archsetup's scripts/, tests/, the archsetup installer and docs/workflows/ found nothing. Nothing exists today that would pass the flag.
+:END:
+
+** DONE Record fields mode and gate.at have no reader, and gate.ok is always false when present
+CLOSED: [2026-10-05 Mon 10:40]
+Where (round 2): Design L104, L108; Phase 1 L2123, L2127; Readiness L2372
+
+data.mode is everyday, complete or apply-armed. gate is either null or {ok: false, kernels, failed, since, at}.
+
+Risk: Every field belongs to the cross-repo fixture that both test suites share (L233, L2191). Each one is a contract obligation with no consumer.
+
+Recommended change: This is the smallest edit:
+- Delete the "- =mode=: everyday, complete or apply-armed." bullet at L104 and L2123.
+- At L108 and L2127, change "since: <T0>, at: <epoch>}" to "since: <T0>}".
+- At L2372, change "{packages: [...], mode, result, failed_step, detail, gate, news}" to "{packages: [...], result, failed_step, detail, gate, news}".
+
+Optional further step: drop ok as well. Define gate as "null, or {kernels: [<kver>...], failed: <item>, since: <T0>}" at L108 and L2127. Then reword "gate.ok is false" to "gate is non-null" at L232, L2137, L2227, L2240, L2255 and L2490.
+
+If I want mode kept for human diagnosis instead, keep it but add "informational; no reader depends on it" to L104 and L2123.
+
+Optional, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases, Decisions, Acceptance criteria, Readiness dimensions, Testing. Adopted, including the optional drop of ok, and one step further: the result value gate-failed goes too. A gate failure is result failed with failed_step gate and exit 4. The panel grades CRIT from the gate field, not from result. So no reader loses anything, and the cross-repo fixture contract gets smaller.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L104 and L2123 define the mode values: everyday, complete or apply-armed.
+- L108 and L2127 define gate as null or {ok: false, kernels, failed, since: <T0>, at: <epoch>}.
+- L2372 lists data as {packages, mode, result, failed_step, detail, gate, news}.
+
+Readers:
+- L2223-2229: the probe reads the record, the flag's presence and line count, packages (minus installed), detail, news, gate.ok (CRIT), result (WARN) and advisories.
+- L2240: REBOOT reads gate.ok and the flag.
+- L2148 and L2487: the re-check reads the recorded kernels and since.
+- L2137 and L232: the stamp reads gate.ok.
+
+A grep of the spec for the record's mode field and for gate.at turns up only the definitions at L104, L108, L2123, L2127 and L2372. The other line it hits, L710, is the earlier finding's {ok: false, ..., at: <ts>}, which is where ok and at came from.
+
+Tests: L2186-2196 and L2479-2490 assert packages, failed_step, the gate's survival and clearing, and gate.ok false. None asserts mode or gate.at.
+
+Code: ~/.dotfiles/maint/src/maint has no upgrade_deferred probe yet, and its only "mode" read (viewmodel.py:821) is the storage probe's evidence. The archsetup scripts/ and tests/ do not mention upgrade_deferred or upgrade-guarded.
+:END:
+
+** DONE R10: P1-complete case (i) drops 'result failed', and nothing replaces the pre-written result on the stage-1-failed, gate-passed path
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 1 --complete step 3 L4166 (pre-write), step 5 last bullet L4175, Finish L4198, The record L4212, Exit codes L4271; P1-complete case (i) L4351
+
+R10 resolved P1-complete (i) as: the run 'ends with exit 1, result failed, failed_step pacman, gate null, and the kernel-side entries still in packages as kind kernel'. L4351 asserts exit 1, failed_step pacman, gate null and the kernel kind kernel entries, but not result failed. In the body, step 3 pre-writes 'result interrupted, failed_step pacman' (L4166). Step 5's fail branch overrides it explicitly ('result failed, failed_step gate, exit 4', L4173). The branch R10 covers, 'if stage 1 failed and the gate didn't fail ... failed_step pacman, exit 1' (L4175), sets no result. Neither does the pass branch ('gate = null', L4172). Finish step 1 is only 'Compute the record' (L4198), and no rule anywhere maps the exit code to result (L4271 lists exit codes only).
+
+Risk: Take a mirror or signature failure in stage 1, the common case R03's structural form was added for. An implementation that carries the pre-written record forward ends with result interrupted and fails no test. The trap's 'interrupted' (exit 130) then becomes indistinguishable from a download failure (exit 1). The same gap leaves a successful --complete's result unspecified after the pre-write.
+
+Recommended change: Three edits, all small:
+
+1. L4351: change "exits 1 with failed_step pacman" to "exits 1 with result failed, failed_step pacman". This restores the wording R10 agreed.
+
+2. Finish, L4198: change "1. Compute the record." to:
+
+ "1. Compute the record. result is ok on exit 0, failed on exit 1 or 4, and refused on exit 3. On exit 0, failed_step is null. These replace any value pre-written by --complete step 3 or --apply-armed step 2, and detail is replaced likewise. Only the interrupt trap writes interrupted."
+
+ This one rule covers the L4175 branch, both success paths, and everyday's failure steps, which also depend on the unstated mapping today.
+
+3. L4334: append "and the record reads result ok with failed_step null", so a stale pre-write after a successful --complete or --apply-armed fails a test.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases (the contract) and the sections that cite it.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- --complete pre-write:
+ - L4166: "result interrupted, failed_step pacman;"
+ - L4167: detail '--complete interrupted during the kernel transaction — do not reboot; ...'
+- L4173, gate fail: "result failed, failed_step gate, exit 4" (explicit override).
+- L4175: "if stage 1 failed and the gate didn't fail (...): failed_step pacman, exit 1." No result.
+- L4172, gate pass: "gate = null". No result.
+- L4186, --apply-armed pre-write: "result interrupted, failed_step boot-transaction ...". L4190 and L4191 override it explicitly on failure. L4192, "Finish, exiting with the transaction's status (0 on a no-op)", sets no result.
+- L4198: "1. Compute the record." Finish gives no rule for result.
+- L4212: result is "ok, failed, refused or interrupted". L4271 lists exit codes only.
+- L4351: (i) asserts "exits 1 with failed_step pacman, gate null and the kernel-side entries still in packages as kind kernel". result is not asserted.
+- L4334: the success tests for --complete and --apply-armed assert only "packages is empty and the stamp is fresh".
+- L4402: Phase 2 grades WARN when result is failed, refused or interrupted.
+- L4643: "the record's result, failed_step and detail carry it to the panel". This implies the mapping without stating it.
+- Grep of Phase 1 for "result": every value is set at one of three kinds of site: an explicit refusal or failure (L4089, 4090, 4106, 4129, 4159, 4173, 4185, 4190, 4191, 4199), a pre-write (L4166, 4186), or the trap (L4272). There is no site for "result ok" and none on L4175.
+
+The agreed round-2 resolution for this case: "(i) A stage 1 that exits non-zero before committing, with the structural gate passing, ends with exit 1, result failed, failed_step pacman, gate null, ..."
+:END:
+
+** DONE R30: the roster-FIX routing test became an assertion that can't fail, and the terminal-kind branch has no GTK-free seam
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 2, APPLY L4412 (_press_lever routes terminal-kind remedies to the launcher); Phase 2 tests intro L4438 ('No test imports gui.py or GTK'); P2-apply L4458
+
+L4412 puts the routing decision inside gui.py's _press_lever: 'On the second, _press_lever calls a GTK-free panel.py launcher ... never _fire or doctor.iter_fix. That covers the row key, the topgrade_age card key and REVIEW & FIX's FIX.' The R30 resolution's test bullet was 'A roster FIX press reaches the launcher and never streams through _fire.' P2-apply L4458 tests something weaker. It checks that the row key and the roster's FIX 'resolve to apply_deferred with kind terminal', that the launcher calls the patched Popen with the foot argv, and that 'iter_fix is never called'. No test drives the press path, so nothing checks that a press on a terminal-kind rid reaches the launcher instead of _fire. Calling the launcher directly never calls iter_fix, so that assertion can't fail.
+
+Risk: The defect R30 and the earlier F40 exist to prevent can come back with the whole suite green. Suppose an implementer adds the kind, the launcher and the items builder but misses the branch in _press_lever, or wires only the row key. REVIEW & FIX's FIX on APPLY then runs 'upgrade-guarded --complete' through _fire and iter_fix inside the panel process. The kernel transaction gets no terminal, the gate verdict and reboot prompt can't be seen, and the transaction is tied to the panel's lifetime.
+
+Recommended change: Move the routing decision into panel.py itself rather than adding a predicate.
+
+Phase 2, APPLY, L4412: replace "On the second, =_press_lever= calls a GTK-free panel.py launcher that runs =Popen(['foot', '--hold', '-e', *argv], start_new_session=True)=, never =_fire= or =doctor.iter_fix=." with:
+
+"On the second, =_press_lever= makes one call, =panel.fire_press(rid, items, stream, fixture)=, and nothing else. For a terminal-kind rid, fire_press runs the launcher =Popen(['foot', '--hold', '-e', *argv], start_new_session=True)=, or on a fixture board returns 'fixture board, not applied' the way MERGE does. For any other rid it calls stream, which is gui.py's =_fire=. It never calls =doctor.iter_fix= for a terminal-kind rid."
+
+Phase 2 tests, P2-apply, L4458: replace "; =iter_fix= is never called" with:
+
+"; =panel.fire_press= given the row key's rid and items, and the roster FIX's (token ='apply_deferred:'=), calls the patched Popen once and never the stream callable; for =update= it calls the stream callable and never Popen; and with fixture set it calls neither."
+
+Optionally, add "a detached foot terminal opens and the wall streams nothing" to M-boot-armed's Expected (L4517). That makes the manual APPLY press check the branch explicitly.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases and the sections that cite it. Partially confirmed. Merged with T15, which reports the same defect: the routing branch sits in gui.py, which no test imports. I took T02's shape: _press_lever makes one GTK-free panel call that owns the choice, and that call has a fixture-board branch like MERGE's. I also took T15's interface, a zero-argument fire callable in place of a stream/_fire signature, so the call doesn't depend on _fire's arguments after Phase 2 removes force. T15's control case (update calls fire, never Popen) is in the test. The optional M-boot-armed Expected clause is included, because it is the only check that gui.py actually makes the call.
+:EVIDENCE:
+Spec (line numbers checked):
+- L4412: "On the second, =_press_lever= calls a GTK-free panel.py launcher ... never =_fire= or =doctor.iter_fix=. That covers the row key, the =topgrade_age= card key and REVIEW & FIX's FIX."
+- L4438: "No test imports gui.py or GTK." L4695 says the same.
+- L4458: "the row key and the roster's FIX both resolve to =apply_deferred= with kind terminal, and the panel.py launcher calls the patched =subprocess.Popen= ...; =iter_fix= is never called"
+- L4392: panel.guarded is kept as the GTK-free seam for the arm line.
+- L4605 (AC4) and L4702 (AC4 is verified by P1-gate, P1-complete, P2-row and P2-apply).
+- L4517: M-boot-armed "arms from APPLY" on ratio and velox, but no Expected clause names the foot terminal.
+- L3121: the earlier round's adopted bullet was "A REVIEW & FIX FIX press on apply_deferred Popens the foot argv and never streams through _fire."
+
+Code (~/.dotfiles/maint/src/maint):
+- gui.py:1408-1438: _press_lever calls model.press(token). On "armed" it builds the arm line; otherwise it calls self._fire(...).
+- gui.py:1440-1442: _on_lever calls _press_lever.
+- gui.py:784-785 and 835-836 (card keys) and gui.py:950-954 (_digest_key, the roster FIX) all call _on_lever, so the entry points share one branch.
+- gui.py:1474-1485: _fire streams doctor.iter_fix through _stream's daemon thread.
+- gui.py:1477: dry = bool(self.fixture).
+- gui.py:1665-1670: _on_merge refuses on a fixture board.
+- panel.py:200-202: key_token is rid + ":" + items, so item-less row and roster keys both carry the token "apply_deferred:".
+- panel.py:224-228: guarded() is a pure predicate.
+- panel.py:314-321: PanelModel.press returns "armed" or "fire" and does no dispatch.
+
+Tests:
+- tests/maint/test_panel.py:3-4 says the GTK view is never imported.
+- A grep of tests/maint for _press_lever or a gui import finds nothing.
+- tests/maint/panel_smoke.py drives presses over AT-SPI, but it isn't part of make test and the spec doesn't name it.
+:END:
+
+** DONE R34: minimal/'s alias files are symlinks to common/'s, so 'stays plain topgrade' and 'minimal files are unchanged' can't both hold
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 2, Alias and README, L4434 ('minimal/'s alias stays plain topgrade'); Phase 2 tests, P2-text, L4468 ('the minimal files are unchanged')
+
+The spec makes the sysupgrade alias in common/.zshrc.d/aliases.sh and common/.bashrc.d/aliases.sh conditional. It says minimal/'s alias stays plain topgrade, and P2-text adds a source check that both common files carry the conditional and the minimal files are unchanged.
+
+Risk: Once common's files change, the minimal 'files' carry the conditional too. A content check that minimal still reads plain topgrade fails. The only way to keep minimal at plain topgrade is to replace the symlinks with regular files, and that turns test_every_shared_path_is_a_symlink_in_minimal red and brings back the drift the dedup was built to stop. The implementer has to choose which rule to break. The runtime outcome is fine either way, since the conditional falls back to topgrade where upgrade-guarded is absent, but the spec's statement and its test are false as written.
+
+Recommended change: L4434: replace "minimal/'s alias stays plain topgrade." with "minimal/'s alias files are symlinks to these (tests/tier-dedup pins that), so a none install gets the same conditional, which falls back to topgrade because upgrade-guarded is never installed there."
+
+L4468: replace "and the minimal files are unchanged" with "and minimal/'s alias files are still symlinks resolving to them".
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases (the contract) and the sections that cite it.
+:EVIDENCE:
+Spec L4434 (Phase 2, Alias and README): "...common/ is also stowed on dwm installs, which never get the scripts. minimal/'s alias stays plain topgrade."
+Spec L4468 (P2-text): "a source check that both common alias files carry the conditional and the minimal files are unchanged;"
+
+Output of git ls-files -s in ~/.dotfiles:
+ 100644 4a59203... common/.bashrc.d/aliases.sh
+ 100644 4a59203... common/.zshrc.d/aliases.sh
+ 120000 fd2654f... minimal/.bashrc.d/aliases.sh
+ 120000 f40d468... minimal/.zshrc.d/aliases.sh
+
+Output of ls -la:
+ minimal/.bashrc.d/aliases.sh -> ../../common/.bashrc.d/aliases.sh
+ minimal/.zshrc.d/aliases.sh -> ../../common/.zshrc.d/aliases.sh
+
+grep shows common/.{bash,zsh}rc.d/aliases.sh:42 is alias sysupgrade="topgrade".
+
+~/.dotfiles/tests/tier-dedup/test_tier_dedup.py:
+- The docstring says minimal's duplicates "are now relative symlinks into common/, making drift structurally impossible".
+- test_every_shared_path_is_a_symlink_in_minimal (L53-58) fails on any regular-file copy.
+- test_every_symlink_resolves_to_commons_copy (L60-71) requires each link to resolve to common's file.
+
+archsetup:1518-1520: the none) branch stows only minimal.
+
+A grep of the spec body (outside Review findings) finds minimal only at L4434 and L4468.
+:END:
+
+** DONE R42: the single Finish write isn't the last thing --complete does, so a HUP or Ctrl+C at the reboot prompt overwrites the finished record
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 1 Finish L4196-4200; --complete step 7 L4181; Exit codes and interrupts L4272; --complete pre-write detail L4167; The record writers L4220; Phase 2 WARN L4402, REBOOT L4426, APPLY launcher L4412
+
+R42 is otherwise present as agreed. Finish computes the record, folds a stamp failure into it, and then says '3. Write the record once and exit.' (L4200). L4330 and AC8 L4619 assert result failed and failed_step stamp. But --complete step 7 (L4181) reads 'Finish. Then prompt "Reboot now? [y/N]"', so the process keeps running after its single write. The trap (L4272) is not scoped. 'Every mode traps INT, TERM and HUP' and 'writes result interrupted with the step ... and the remedy (the pre-written detail where one exists)'. L4220 lists the trap as a record writer with no exception for after the finish.
+
+Risk: Two parts of the spec conflict: 'and exit' in Finish and 'Then prompt' in step 7. An implementer who follows the trap literally ships a record that contradicts a successful run. It tells the user 'do not reboot' while the panel offers REBOOT, and it breaks the single-write outcome R42 was meant to guarantee. Nothing is unsafe, because the message errs toward not rebooting and a later run rewrites it. But it is misleading after the most common APPLY outcome.
+
+Recommended change: Phase 1 Finish, L4200: replace "3. Write the record once and exit." with "3. Write the record once. From then on the trap writes nothing and only exits 130. Then exit, except that =--complete= goes on to its step 7 prompt."
+Exit codes and interrupts, L4272: after "and exits 130." add "After Finish's write it exits 130 without writing."
+P1-arm-boot, after L4363: add "- an INT or HUP at the reboot prompt exits 130 and leaves the record byte-identical to the one Finish wrote."
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases and the sections that cite it. Confirmed, and it overlaps T06: both say Finish 'exits' before step 7's prompt and the trap is still live there. They conflict on the outcome. T04 wants exit 130 with no rewrite; T06 wants an exit of 0. I took T06's semantics. When the prompt shows, the run has already succeeded and the record reads ok, so a signal or EOF at the prompt is a 'no' and the exit stays 0. Exit 130 would contradict the record. Finish step 3's half of this fix is carried in T01's Finish edit. The step 7, trap and test edits are here, and T06 adds nothing separate.
+:EVIDENCE:
+- Spec L4200: "3. Write the record once and exit."
+- Spec L4181: "7. Finish. Then prompt 'Reboot now? [y/N]' only when the exit is 0, stdin is a tty (the APPLY terminal is one), and the run armed or its gate passed in the =--since= form."
+- Spec L4272: "Every mode traps INT, TERM and HUP. The trap waits for the running child ..., writes result interrupted with the step ... and the remedy (the pre-written detail where one exists), keeps gate as written, and exits 130." It does not say what happens after Finish.
+- Spec L4220: the record is "Written by every mode's finish ..., the =--complete= pre-write, the =--apply-armed= pre-write and the interrupt trap." There is no post-finish exception.
+- Spec L4164-4167: the pre-write detail is '--complete interrupted during the kernel transaction — do not reboot; run upgrade-guarded --complete'.
+- Spec L4402: the row goes WARN on result interrupted and shows detail's last line. L4426: REBOOT is shown while the flag exists and gate is null.
+- Spec L4412: APPLY runs Popen(['foot', '--hold', '-e', *argv]). On this machine "foot --help" lists "-H,--hold remain open after child process exits", and "pacman -Q foot" reports foot 1.28.0-2.
+- Spec L4360-4369 (P1-arm-boot): only L4363 tests the prompt ("prompts for a reboot only when the exit is 0 and stdin is a tty"). No case covers a signal arriving at the prompt.
+:END:
+
+** DONE R53: --dry-run's private db reaches only the closure probes; the normative Sets bullet still takes the pending set from the stale system sync db
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 1, Sets L4093 (pending set); Phase 1, Modes and options L4082 (--dry-run); Dependency closure L4117, L4125, L4128; Rollout step (a) L4559 and M-split-live L4562; Design L85
+
+L4082: --dry-run 'creates a private CHECKUPDATES_DB with mktemp -d (removed on exit), runs checkupdates against it (exit 2 means empty), and passes that path as every closure probe's --dbpath.' L4093, which has no mode exception: 'Pending set: the rows of LC_ALL=C pacman -Qu (name old -> new) against the system sync db, minus rows marked [ignored].' No body line says that --dry-run's pending set is checkupdates' output. R53's 'current' text quoted the earlier spec as saying '--dry-run takes its pending set from checkupdates', and that clause did not survive the rewrite.
+
+Risk: An implementer who follows the normative Sets bullet builds a preview that resolves its pending set and its closure against two different dbs. The preview can then refuse when no refusal is due, or predict the wrong deferred set. That undermines the rollout's dry-run gate and the M-split-live comparison, which is the defect R53 was accepted to remove.
+
+Recommended change: At L4093, after "...minus rows marked =[ignored]= (pacman.conf IgnorePkg, such as bridge-utils on ratio)", add: "Under =--dry-run= it is instead the lines =checkupdates= prints against the run's private =CHECKUPDATES_DB= (Modes and options), which already drop =[ignored]= rows. Exit 2 means empty, and any other non-zero exit fails the preview with exit 1 and no record."
+
+Optionally, extend the --dry-run case at L4325 so the test pins the source: give the fake checkupdates and the fake pacman -Qu different rows, and assert that the printed pending set is the checkupdates rows.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases (the contract) and the sections that cite it.
+:EVIDENCE:
+- Spec L4093: "Pending set: the rows of =LC_ALL=C pacman -Qu= (=name old -> new=) against the system sync db, minus rows marked =[ignored]=". It has no --dry-run exception.
+- Spec L4082: "creates a private =CHECKUPDATES_DB= with =mktemp -d= ..., runs =checkupdates= against it (exit 2 means empty), and passes that path as every closure probe's =--dbpath=". It does not say the pending set is taken from this output.
+- Spec L4117 does have the carve-out: "runs against the run's one refreshed db (=--dry-run= passes its private db as =--dbpath=)". So the Sets bullet is the odd one out.
+- Spec L4052: phase bullets are normative over Design. L85 ("resolves against a private sync db of its own") therefore yields to L4093.
+- Spec L4081: dry-run "uses no sudo", so there is no -Sy and the system db stays stale. L754 records the system db showing 0 updates while checkupdates showed 395.
+- Spec L4125 / L4128: a probe line naming a package that is not pending "adds nothing", and "an iteration that adds no new name" refuses.
+- Spec L4325: the test checks only the probes' --dbpath, not the source of the pending set.
+- Spec L4559, L4562, L4592: the rollout gate and M-split-live compare against dry-run's predicted deferred set.
+- /usr/bin/checkupdates (pacman-contrib 1.13.1-1): L150 runs "fakeroot -- pacman -Sy ... --dbpath "$CHECKUPDATES_DB"". L155 runs "pacman -Qu --dbpath "$CHECKUPDATES_DB" ... | grep -v '\[.*\]'", so [ignored] rows are already dropped. L181 is "exit 2".
+:END:
+
+** DONE --complete's reboot prompt comes after a Finish that already exits, and the trap is live at the prompt
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 1, --complete sequence step 7 (line 4181); Phase 1, Finish (lines 4196-4200); Exit codes and interrupts (line 4272); The record failed_step list (line 4213); P1-arm-boot (line 4363)
+
+Step 7 reads 'Finish. Then prompt Reboot now? [y/N] ...', but Finish ends with 'Write the record once and exit.' The trap covers INT, TERM and HUP in every mode with no exception for the prompt, and failed_step has no value for an interruption there.
+
+Risk: An implementer has to decide where the prompt sits relative to the record write and the trap. If the prompt follows the write with the trap still armed, Ctrl+C or closing the window at 'Reboot now?' (a natural way to say no) overwrites a clean arming or a passed-gate record with result interrupted, sets an undefined failed_step, and exits 130. The row then shows WARN after a session that succeeded. No P1 case covers an interrupt at the prompt.
+
+Recommended change: Three edits.
+
+L4181: replace the step with "7. Finish writes the record. Then, only when the exit is 0, stdin is a tty (the APPLY terminal is one), and the run armed or its gate passed in the --since form, prompt 'Reboot now? [y/N]' before exiting. The interrupt trap no longer applies at the prompt: INT, TERM, HUP, EOF or any answer but y count as no, the record isn't rewritten, and the exit stays 0. Yes runs sudo -n systemctl reboot."
+
+L4200: change it to "3. Write the record once, then exit (for --complete, after step 7's prompt)."
+
+L4272: insert "until finish has written the record" after "traps INT, TERM and HUP".
+
+Then add a P1-arm-boot case after L4363: "an INT or HUP at the reboot prompt leaves the record byte-identical and exits 0."
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases and the sections that cite it. Confirmed. This is the same defect as T04, merged there. I adopted T06's semantics in T04's edits: the prompt comes after Finish's single write, a signal or EOF at the prompt counts as no, and the exit stays 0. T01 E1 carries Finish step 3 ('then exit, except that --complete first runs its step 7 prompt'). Nothing separate here, so no edit is applied twice.
+:EVIDENCE:
+The spec lines:
+- L4181: "7. Finish. Then prompt 'Reboot now? [y/N]' only when the exit is 0, stdin is a tty ..."
+- L4200: "3. Write the record once and exit."
+- L4272: "Every mode traps INT, TERM and HUP. The trap waits for the running child ..., writes result interrupted with the step ... and the remedy (the pre-written detail where one exists), keeps gate as written, and exits 130." It makes no exception for the prompt.
+- L4164-4167: the --complete pre-write detail '... — do not reboot; run upgrade-guarded --complete'.
+- L4402: the row is "WARN when result is failed, refused or interrupted (the text adds detail's last line ...)".
+- L4412: APPLY runs Popen(['foot', '--hold', '-e', ...]), so closing that window sends HUP.
+- L4213: failed_step has no value for an interrupt at the prompt.
+- L4363: the only P1 prompt case is "prompts for a reboot only when the exit is 0 and stdin is a tty".
+
+Shell check: bash -c 'trap "echo trapped-at-prompt; exit 130" INT HUP; (sleep 0.5; kill -HUP $$) & read -r ans' with stdin a FIFO that has no data printed "trapped-at-prompt", and the exit was 130.
+:END:
+
+** DONE The lever runner's 3600 s timeout SIGKILLs upgrade-guarded past its trap, so it bounds nothing and leaves the record stale
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 1, Everyday sequence step 3 sub-bullet (line 4136); Phase 2, Levers (line 4380); Exit codes and interrupts (line 4272); Design 'The record and the stamp' (line 93); Risks, Arch news (line 4692); AC7 (line 4614)
+
+Phase 2 keeps UPDATE and TOPGRADE on the existing lever runner with timeout 3600. Phase 1 calls a stuck informant-hook wait 'bounded by the lever runner's 3600 s on a lever run'. Design says the panel 'always reflects the last run' and that the trap writes the record on a kill. The everyday run has no pre-write, and the trap catches only INT, TERM and HUP.
+
+Risk: When a lever run passes 3600 s (an informant fetch that hangs inside 00-informant, or a large AUR build plus the sweep), upgrade-guarded is SIGKILLed with no trap and no record write. The panel keeps showing the previous run's result. Its sudo, pacman, yay or topgrade descendants keep running unsupervised, and a pacman still in the hook keeps db.lck, so later runs refuse. 'Bounded by 3600 s' bounds only the panel's wait.
+
+Recommended change: This is option (a), using SIGINT only, to match the boot unit.
+
+1. Phase 2, Levers (L4380): after "stay on the existing lever runner", add:
+
+ "except for the timeout: _execute_step starts a remedy marked long with start_new_session=True. Once its timeout passes, it sends SIGINT to the process group, waits up to 120 s, then SIGKILLs the group and reports 'timed out — interrupted'. Never SIGTERM, because pacman has no SIGTERM handler. upgrade-guarded's trap then records result interrupted (exit 130)."
+
+2. P2-levers (L4439): add a case. A fake argv that sleeps past a 1 s timeout receives SIGINT, its trap writes its marker, and no process in its group survives.
+
+3. Phase 1 step 3 (L4136) and Risks, Arch news (L4692): change "bounded by the lever runner's 3600 s" to:
+
+ "on a lever run, ended by the runner's 3600 s timeout as a SIGINT to the run's process group, which stops pacman at a package boundary and which the trap records as interrupted"
+
+4. Design L93 can then stand as written.
+
+If maint isn't to change, take the alternative. Replace those two passages with the true residual: a lever timeout SIGKILLs only the script, so no trap runs and the record stays at the previous run, while sudo/pacman continue orphaned holding db.lck and may die on SIGPIPE mid-package. Then name that path in Risks, Partial ecosystem state, and drop "the lever runner's timeout" from AC7's list of endings that swap nothing.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases (the contract) and the sections that cite it.
+:EVIDENCE:
+- ~/.dotfiles/maint/src/maint/doctor.py:161 runs =proc = cmd.run(step["argv"], timeout=r.get("timeout", 60))=. doctor.py:163 returns "tool missing or timed out".
+- ~/.dotfiles/maint/src/maint/cmd.py:22 runs =subprocess.run(cmd, capture_output=True, text=True, timeout=timeout)= and catches TimeoutExpired.
+- CPython 3.14.7 subprocess.run, from inspect.getsource on this machine: on TimeoutExpired it calls =process.kill()= then =process.wait()=, and Popen.kill sends SIGKILL to the direct child only.
+- remedies.py:301 and :312 set ="timeout": 3600= for update and topgrade.
+- grep for start_new_session/killpg/os.setsid in maint finds only gui.py:1672/1756/1766 (MERGE/journal/net Popen). The lever runner has no process-group handling.
+- Spec L4272: the trap covers INT, TERM and HUP only. L4136 and L4692: "bounded by the lever runner's 3600 s". L93: "so the panel always reflects the last run". The everyday sequence (L4133-4148) has no pre-write; the only pre-writes are L4164 (--complete) and L4186 (--apply-armed).
+- Spec L4498 and L4684: pacman has no SIGTERM handler, and the unit uses SIGINT to the cgroup.
+- =bash -c 'pacman -Ql | head -1 >/dev/null; echo ${PIPESTATUS[0]}'= prints 141: pacman dies on SIGPIPE, at least in query mode.
+:END:
+
+** DONE 'minimal/'s alias stays plain topgrade' and 'the minimal files are unchanged' are impossible: they are symlinks to the common files
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 2, Alias and README (line 4434); P2-text (line 4468)
+
+Phase 2 rewrites the sysupgrade alias in common/.zshrc.d/aliases.sh and common/.bashrc.d/aliases.sh, says 'minimal/'s alias stays plain topgrade', and P2-text asserts 'the minimal files are unchanged'.
+
+Risk: P2-text as written either fails (reading through the symlink shows the new conditional) or is vacuous (the symlink target is unchanged), so an implementer has to guess its intent. At runtime minimal gets the same conditional. That is harmless, because it falls back to topgrade where upgrade-guarded is absent, but the spec describes a separate file that doesn't exist.
+
+Recommended change: Line 4434: replace "minimal/'s alias stays plain topgrade." with "minimal/'s aliases.sh files are symlinks into common/ (tests/tier-dedup pins this), so they carry the same conditional, which resolves to topgrade where upgrade-guarded isn't installed. Don't replace them with regular files."
+Line 4468: replace "and the minimal files are unchanged" with "and minimal/.zshrc.d/aliases.sh and minimal/.bashrc.d/aliases.sh are still symlinks resolving to them". Alternatively, drop the minimal clause, since the tier-dedup test already enforces this.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases and the sections that cite it. Confirmed. This is the same finding as T03 and is merged there. T03's edits say the minimal files are symlinks that carry the conditional, and they replace the 'unchanged' check with a symlink check. That covers this finding's recommended wording.
+:EVIDENCE:
+Spec line 4434: "...common/ is also stowed on dwm installs, which never get the scripts. minimal/'s alias stays plain topgrade."
+Spec line 4468: "a source check that both common alias files carry the conditional and the minimal files are unchanged;"
+Running git ls-files -s in ~/.dotfiles shows:
+ 100644 ... common/.bashrc.d/aliases.sh
+ 100644 ... common/.zshrc.d/aliases.sh
+ 120000 ... minimal/.bashrc.d/aliases.sh
+ 120000 ... minimal/.zshrc.d/aliases.sh
+ls -la shows minimal/.bashrc.d/aliases.sh -> ../../common/.bashrc.d/aliases.sh and minimal/.zshrc.d/aliases.sh -> ../../common/.zshrc.d/aliases.sh.
+grep sysupgrade finds only common/.zshrc.d/aliases.sh:42 and common/.bashrc.d/aliases.sh:42, both alias sysupgrade="topgrade".
+~/.dotfiles/tests/tier-dedup/test_tier_dedup.py: test_every_shared_path_is_a_symlink_in_minimal fails on "regular-file duplicates in minimal/ — convert to symlinks into common/".
+archsetup:1518-1521: the none) branch runs stow on minimal only.
+Spec line 4690: the hold covers "the sysupgrade alias where upgrade-guarded is installed".
+:END:
+
+** DONE Some acceptance-criterion clauses map to test groups whose cases don't check them
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Testing mapping (lines 4702, 4705, 4710, 4711); AC4 (line 4603); AC7 (line 4616); AC12 (line 4628); AC13 (line 4629); P1-steps (line 4320); P1-complete (line 4359); P1-installer (lines 4371-4373)
+
+AC12's 'installer writes the hook as 10-hypr-live-update-guard.hook' is mapped to 'P1-installer (TOML pin, hook name)'. AC13's selection rule (the oldest pre-pacman snapshot no earlier than a minute before the entry) and 'held_snapshot names it' are mapped to P1-complete. AC4's 'across a wall dismiss and a panel reopen' is mapped to P2-row and P2-apply. AC7's 'record the unread items before the clear' is mapped to P1-steps.
+
+Risk: These clauses can be ticked with nothing having checked them, and the target-selection rule is the one that decides whether the right fallback snapshot stays pinned.
+
+Recommended change: In P1-complete (spec L4359), after "and a structural-form gate places no hold.", add: "A fake =zfs list -H -p -t snapshot -o name,creation= offers four root-dataset pre-pacman snapshots: one created at since − 61 s, one at exactly since − 60 s, one inside the window, and one after a later everyday run. The hold targets the since − 60 s snapshot, and the record's held_snapshot equals that name. A follow-up --since gate with the same recorded since holds nothing new and releases nothing." The four-snapshot fixture is the whole fix; nothing else needs to change. Optionally, change the AC12 mapping text at L4710 from "P1-installer (TOML pin, hook name)" to "P1-installer (TOML pin, which locates the heredoc by its exact =10-hypr-live-update-guard.hook= path)". Make no change for the AC4 or AC7 parts.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases and the sections that cite it. Partially confirmed, and the finding was narrowed to the AC13 selection rule. The four-snapshot selection case goes into P1-complete. Because T11 changes when the post-gate hold applies, the case is folded into T11's single rewrite of the P1-complete hold bullet (T11 E6), with its fixture shaped to the new rule (held_snapshot older than since − 60 s). The AC12 hook-name gap is closed in the phase, not in Testing: the TOML pin test now locates the heredoc by its exact 10- path. No change for the AC4 or AC7 parts, which the check found already correct.
+:EVIDENCE:
+- Spec L4248: "the target is the oldest =<rootds>@pre-pacman_*= with creation ≥ gate.since − 60 s". L4252: "held_snapshot = target".
+- Spec L4359 (P1-complete, the only hold case): "a fake zfs logs hold and release. The hold is placed after a passing and after a failing =--since= gate, the previous hold is released, the same target is never re-held, and a structural-form gate places no hold." It asserts neither the selection nor the held_snapshot value.
+- Spec L4354: "the gate field survives a following everyday run". L4355: a follow-up --complete "runs the =--since= form against the recorded since".
+- Spec L4711: "AC13 …: P1-complete and P1-installer". L4373 (P1-installer): only the zfs-pre-snapshot prune of held snapshots.
+- Spec L4575 (D-zbm, "non-gating", L4573): "records the held snapshot in held_snapshot".
+- archsetup:2625 =cat > /etc/pacman.d/hooks/10-hypr-live-update-guard.hook << 'HOOKEOF'=. tests/installer-steps/test_pacman_hook_order.py:47-61 asserts one written hook ending hypr-live-update-guard.hook that sorts before 60-mkinitcpio-remove.hook, plus the ^\d{2}- prefix. scripts/testing/tests/test_desktop.py:59 asserts the exact 10- path in the VM.
+- Spec L4134: "News: =yay -Pwq=, before the first transaction, because yay judges news against installed build dates. It needs no informant." =man yay=: "-w, --news Print new news … News is considered new if it is newer than the build date of all native packages."
+- Spec L4410: "the panel reads the outcome from the record and the flag on its next probe". L4427: "The panel never infers arm or gate state from an exit code". L4447 (P2-row): "a gate-open fixture gives the CRIT text, REBOOT hidden and APPLY shown".
+:END:
+
+** DONE After an unfinished stage 1, the stated remedy (--complete) can never clear the gate
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 1, --complete sequence steps 3–5 (lines 4162–4175, detail at 4167); Phase 2 CRIT text (4401); Phase 4 README flow and recovery (4537)
+
+The pre-written detail says 'do not reboot; run upgrade-guarded --complete'. The CRIT row says 'run upgrade-guarded --complete after any fix'. The only manual step documented is diagnosis with kernel-modules-check (4537). The spec names no repair.
+
+Risk: The machine sits at CRIT 'do not reboot' indefinitely, and the one remedy the spec gives can't clear it. On velox, an operator with no documented way out eventually reboots into a /boot with no kernel.
+
+Recommended change: 1. Phase 1, --complete step 4 (spec 4168). Append:
+ "When the previous gate is non-null, stage 1 also names reinstall targets after the --ignore list, without --needed:
+ - for each pkgbase still in gate.pkgbases, the owner (pacman -Qqo) of its ${KMC_MODULES_DIR}/<kver>/vmlinuz, plus <pkgbase>-headers where installed;
+ - every DKMS-set member.
+ A same-version reinstall is an Upgrade path operation, so 90-mkinitcpio-install (install_kernel, which restores the preset and /boot/vmlinuz-<pkgbase>) and 70-dkms-upgrade/70-dkms-install fire again. 60-mkinitcpio-remove fires only on Remove. Without this, a stage 1 interrupted after the kernel committed, which skips every PostTransaction hook, can never pass a later gate."
+
+2. Step 3's detail (4167). Keep the wording; it becomes true once item 1 lands.
+
+3. P1-complete. Add:
+ "after an interrupted stage 1 whose kernel landed (fake /boot/vmlinuz-<pkgbase> missing, nothing kernel-side pending), the follow-up stage 1 argv names the gate's kernel package, its headers and the DKMS-set members as targets without --needed; with the fake hooks restoring the image, the gate passes and sets gate null."
+
+4. Phase 4 README bullet (4537). Add the manual equivalent:
+ "if the check reports a missing /boot image or initramfs, sudo pacman -S <pkgbase> <pkgbase>-headers <dkms packages> re-runs the post-transaction hooks; then run upgrade-guarded --complete."
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases and the sections that cite it. Confirmed. A stage 1 cut off after the kernel committed skips 90-mkinitcpio-install and the DKMS hooks. A follow-up -Su then has nothing kernel-side pending, so the gate's /boot item fails forever. I kept the reinstall-on-open-gate fix, with two narrowings. First, targets resolve with pacman -Q <pkgbase>, the same check step 5 already uses, rather than a new pacman -Qqo over vmlinuz, which finds nothing once the module tree is gone. Second, I dropped the README manual-pacman step, because --complete now does the reinstall itself and the README already names --complete as the only clearing path. The pre-write detail stays as it is, because it is now true.
+:EVIDENCE:
+Spec citations:
+- 4165–4167: pre-write sets gate.failed 'kernel transaction not finished' and the detail '... do not reboot; run upgrade-guarded --complete'.
+- 4168: stage 1 is a plain -Su with no named targets.
+- 4171: a follow-up run (previous gate non-null) uses --since.
+- 4173: a failed gate exits 4.
+- 4238: the gate requires /boot/vmlinuz-<pkgbase>.
+- 4401: CRIT text 'run upgrade-guarded --complete after any fix'.
+- 4537: the README gives diagnosis only.
+- 4354: P1-complete expects the follow-up --complete to clear the gate after an interrupt.
+- 4347 and 4603: interrupted stage 1 is a planned-for case.
+
+Live system (ratio):
+- /usr/share/libalpm/hooks/60-mkinitcpio-remove.hook: 'Operation = Remove', 'Target = usr/lib/modules/*/vmlinuz', 'When = PreTransaction'.
+- 90-mkinitcpio-install.hook: second [Trigger] is 'Operation = Install', 'Operation = Upgrade', 'Target = usr/lib/modules/*/vmlinuz'; 'When = PostTransaction'.
+- 70-dkms-install.hook: PostTransaction, Install/Upgrade on usr/src/*/dkms.conf, usr/lib/modules/*/build/include/ and usr/lib/modules/*/modules.order.
+- /usr/share/libalpm/scripts/mkinitcpio:
+ - remove_kernel does 'rm -f -- "${filelist[@]}"' (the kernel copy plus the images), then remove_preset moves a non-template preset to .pacsave;
+ - install_kernel (install_preset, then 'install -Dm644 -- "${line}" "${kernel}"') runs only for a */vmlinuz install target;
+ - every other trigger only sets package=1, which becomes 'mkinitcpio -P'.
+- /etc/mkinitcpio.d/linux-lts.preset: PRESETS=(default fallback), ALL_kver="/boot/vmlinuz-linux-lts". It differs from /usr/share/mkinitcpio/hook.preset, so it goes to .pacsave.
+- /usr/bin/mkinitcpio:133: error "specified kernel image does not exist".
+- strings on libalpm: 'transaction started' / 'transaction interrupted' / 'transaction completed'.
+- Versions: pacman 7.1.0.r9.g54d9411-2, mkinitcpio 42.1-1, dkms 3.4.3-2.
+:END:
+
+** DONE Power loss or a kill during stage 1 gets no snapshot hold, and Recovery points at the wrong snapshot
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 1, --complete step 5 (4174); The gate (4242); The snapshot hold (4248, 4255); Exit codes and interrupts (4272); Phase 4 Recovery (4541–4548); D-zbm (4572–4577)
+
+The hold is placed only after a --since gate runs. The trap exits 130 without running the gate, and power loss runs nothing. Both the gate's snapshot item and the hold target are 'creation ≥ since − 60 s', with no upper bound. Recovery boots 'the snapshot the record's held_snapshot names', and D-zbm drills only the failed-gate path.
+
+Risk: Power loss in the kernel transaction is the case that most needs ZFSBootMenu recovery. In that case Recovery names a stale snapshot or none, its =pacman -U --dbonly= steps fail on db.lck, and the kernel package is left unregistered. Separately, the snapshot check can pass on, and pin, a snapshot that cannot fall back to the old kernel.
+
+Recommended change: 1. Root-step list (4064-4070): add =zfs snapshot=.
+2. --complete step 3 (4162-4167): on a ZFS root, when this run opens the entry (the previous gate is null), after the pre-write and before stage 1:
+ - run =sudo -n zfs snapshot <rootds>@pre-pacman_<YYYY-MM-DD_HH-MM-SS>=, then =sudo -n zfs hold upgrade-guarded <it>=;
+ - release the previous held_snapshot, tolerating a missing one, set held_snapshot to the new snapshot, and write the record again;
+ - a failure prints 'snapshot hold failed: <reason>' and stage 1 still runs.
+ When the gate is already open, keep held_snapshot.
+3. The snapshot hold (4248): apply the post-gate hold only when held_snapshot is null or was created before gate.since − 60. The held snapshot then can't be pruned, so the gate item at 4242 can stay unchanged and can no longer pass on a snapshot taken after the kernel landed.
+4. Recovery (4541-4548), before the reconcile:
+ - "Remove /var/lib/pacman/db.lck if a power cut left it; nothing is running after a reboot."
+ - After the reconcile: "A power cut mid-transaction leaves the package being extracted with no [ALPM] line and a db entry without desc that pacman -Q still lists at the new version. Find it with =pacman -Dk= ('description file is missing'), delete that entry's directory under /var/lib/pacman/local, and =pacman -U --dbonly= its old version from the cache."
+ - End the reconcile with "=pacman -Dk= and =pacman -Qkk= are clean."
+5. D-zbm (4572-4577): add a variant that powers off the VM from the QEMU monitor during stage 1's extraction, boots held_snapshot, and runs Recovery. Expected: the old kernel set, and clean =pacman -Dk= and =pacman -Qkk=.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases and the sections that cite it. Partially confirmed; the recommended change reflects the narrowed finding. The hold is placed only after a --since gate. A kill or power cut in stage 1 therefore leaves no hold and a stale held_snapshot, and Recovery boots that stale name. I kept the fix: when --complete opens the entry, it takes and holds its own pre-stage-1 snapshot, and the post-gate hold applies only when that hold is missing. I also kept the Recovery additions (db.lck, the missing-desc entry, a clean pacman -Dk) and the non-gating D-zbm power-off variant. I verified that zfs-pre-snapshot's 60 s skip keys on its own lockfile, not on existing snapshots. The hook may therefore still cut its own snapshot seconds later, or fail on a same-second name, which it only warns about. Both are harmless, so no extra rule is needed. Merges: T09's selection case and T18's hold-failure case go into one rewrite of the P1-complete hold bullet (E6). T18 owns the failure bullet of The snapshot hold. T21 owns Recovery's opening. Design and Risks get pointer fixes. AC13 still holds as written.
+:EVIDENCE:
+Spec lines checked:
+- 4174 and 4248-4255: the hold happens only "after a --since-form gate on a ZFS root".
+- 4272: the trap "writes result interrupted ... keeps gate as written, and exits 130". It runs no gate and places no hold.
+- 4242 and 4248: the only bound is "creation ≥ since − 60 s".
+- 4206 and 4217: held_snapshot is carried forward.
+- 4541: "After booting the snapshot the record's held_snapshot names from ZFSBootMenu".
+- 109: "The recorded name is the snapshot recovery boots".
+- 4688: "an unplanned reboot while the gate is open ... is recovered from ZFSBootMenu".
+- 4572-4577: D-zbm covers only the failed gate.
+- 4132-4152: the everyday sequence never refuses while the gate is open.
+- 4685: power loss can leave "a half-extracted package and a stale db.lck".
+
+Code and system:
+- scripts/zfs-pre-snapshot:14 sets KEEP=10. Lines 35-40 prune on every snapshot it creates.
+- archsetup:2421-2433: 05-zfs-snapshot.hook runs on every Upgrade, Install and Remove transaction (Target = *), PreTransaction.
+- /usr/share/libalpm/hooks/60-mkinitcpio-remove.hook is a PreTransaction Path Remove trigger on usr/lib/modules/*/vmlinuz that runs "mkinitcpio remove". It removes the /boot image during a kernel upgrade, so a power cut in that window leaves velox with no /boot/vmlinuz-linux-lts.
+- Installed pacman is 7.1.0.
+
+Scratch db test (a throwaway =--dbpath=, local/linux-lts-6.18.30-1/ holding only mtree):
+- =pacman -Q= printed "linux-lts 6.18.30-1" and exited 0.
+- =pacman -Qi linux-lts= printed "error: could not open file .../desc".
+- =pacman -Dk= printed "error: 'linux-lts-6.18.30-1': description file is missing".
+:END:
+
+** DONE KillMode=control-group sends SIGINT to the output mirror too, so pacman can die of SIGPIPE mid-package
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 1, --apply-armed step 3 (4187); Phase 3, The unit (4489, 4491) and Disarm and failure (4498); AC6 (4611); M-boot-midtx (4520)
+
+All boot-run output is mirrored to the journal with systemd-cat while it also goes to tty1. The unit's KillSignal=SIGINT is sent with KillMode left at control-group. The spec relies on pacman stopping at a package boundary and releasing db.lck.
+
+Risk: A boot-run timeout can leave a half-extracted GPU/compositor package and a stale db.lck. That is the outcome KillSignal=SIGINT exists to prevent, and AC6's 'no db.lck remains' claim fails intermittently because it is a race.
+
+Recommended change: Phase 1, --apply-armed step 3 (4187): after the systemd-cat sentence, add: 'The mirror must outlive the unit's SIGINT, because KillMode=control-group signals it too, and pacman dies of SIGPIPE if the reader of its output is gone. Start the mirror asynchronously, for example =exec > >(tee >(systemd-cat -t archsetup-boot-upgrade)) 2>&1=, which bash starts with SIGINT ignored; otherwise run =trap "" INT= before exec. Never use a foreground pipeline.'
+
+Phase 3, Disarm and failure (4498): add 'the output mirror ignores SIGINT (Phase 1, --apply-armed step 3)'.
+
+P1-arm-boot (after 4368): add a case. Run a fake --apply-armed under setsid with a fake pacman that traps INT, writes to stderr after the signal, then removes a fake lock file and exits 130. Send SIGINT to the whole process group to emulate control-group. Assert that the fake pacman was not killed by SIGPIPE, the fake lock is gone, and the record reads interrupted.
+
+Don't adopt the KillMode=mixed alternative unless the trap at 4272 is also changed to forward SIGINT to its child.
+
+Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases and the sections that cite it. Partially confirmed; the recommended change reflects the narrowed finding. KillMode=control-group SIGINTs a foreground-pipeline mirror, and pacman dies of SIGPIPE writing its interrupt message. I widened the fix from the boot form to a rule in Exit codes and interrupts. The same group-wide SIGINT also reaches every mode: Ctrl+C in a terminal --complete during the kernel transaction, and T07's lever-timeout SIGINT. Any in-script tee or stderr capture would kill pacman mid-package the same way. The --apply-armed step points at the rule with the bash example. KillMode stays control-group, and T12's own advice is not to switch to mixed.
+:EVIDENCE:
+Spec:
+- 4187: 'All output is mirrored to the journal with systemd-cat -t archsetup-boot-upgrade'. No construction is given.
+- 4489: 'KillSignal=SIGINT, with KillMode left at control-group'.
+- 4491: StandardInput=null, StandardOutput=tty, TTYPath=/dev/tty1.
+- 4498: 'A timeout sends SIGINT to the cgroup. pacman ... stops at a package boundary and releases db.lck'.
+- 4272: the trap 'waits for the running child'. It does not forward the signal.
+- 4368: the P1 test only checks that output goes 'through the fake systemd-cat'.
+- 4611 (AC6) and 4684 (Risks) claim no db.lck remains after a SIGINT-ended timeout.
+
+Live checks on ratio:
+- man systemd.kill: control-group means 'all remaining processes in the control group of this unit will be killed'. Default is control-group.
+- man systemd.exec: 'If the TTY is used for output only, the executed process will not become the controlling process of the terminal'.
+- man sudoers, use_pty: 'If the sudo process is not attached to a terminal, use_pty has no effect'. sudo is 1.9.17p2.
+- =bash -c 'pacman -Ql glibc | head -n1 >/dev/null; echo ${PIPESTATUS[*]}'= printed =141 0=, so pacman 7.1.0 dies of SIGPIPE.
+- strings /usr/bin/pacman shows the 'Interrupt signal received' and 'Hangup signal' messages beside alpm_trans_interrupt. nm shows sigaction only.
+
+SIGINT disposition by mirror construction, bash 5.3.20, run from a parent with SIGINT/SIGQUIT reset to SIG_DFL, reading /proc/<pid>/status SigIgn:
+
+| Construction | SigIgn | SIGINT |
+|---|---|---|
+| foreground pipeline element (=echo \| sh -c ...=) | 0x0 | dies |
+| =exec > >(tee ...)= child | 0x6 | ignored |
+| =<(...)= / =>(...)= | 0x6 | ignored |
+| pipeline inside a procsub | 0x6 | ignored |
+| tee with a nested =>(...)= | 0x6 | ignored |
+| explicit =&= | 0x6 | ignored |
+
+Sibling scripts: scripts/hypr-live-update-guard starts with #!/bin/sh; scripts/zfs-pre-snapshot starts with #!/bin/bash. The spec names no language for upgrade-guarded.
+:END:
+
+** DONE The everyday run never detects a kernel or DKMS change made by yay, so REBOOT stays offered
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 1, Everyday sequence step 7 (4142); Invariants (4206); Phase 2, Freshness and REBOOT (4426); Risks (4691)
+
+An AUR upgrade that needs a newer kernel-set or DKMS-set package 'installs it ungated', and Risks accepts this. Only --complete writes gate, so the everyday run opens no entry and records no kernel-side change.
+
+Risk: This reproduces the velox failure chain the spec exists to prevent (a new kernel, an unverified zfs build, an initramfs possibly without zfs.ko), with the panel offering REBOOT and no row saying anything.
+
+Recommended change: 1. Everyday step 7 (4140-4142): replace the 'installs it ungated' sub-bullet with this:
+
+"Before yay, record =pacman -Q= of the kernel-set and DKMS-set members and T7 = now. After yay, whatever its exit, recompute the kernel set and compare. If any member's version changed, or a new kernel-set member appeared, open gate as --complete step 3 does: {pkgbases: the previous pkgbases ∪ the pkgbases of the changed kernels ∪ (on a DKMS change) every kver with an =installed= line, since: the previous since or T7, failed: 'kernel-side package moved by the AUR step'}. Then set failed_step aur, detail '<names> moved by yay -Sua — do not reboot; run upgrade-guarded --complete', exit 1."
+
+2. Everyday step 9 (4147): add a first case, "gate is non-null: =sudo -n rm -f= the flag". This keeps invariant 4203.
+
+3. Invariants (4206): change it to "Only =--complete= and everyday step 7's kernel-change check write gate; only =--complete= closes it. Every other mode carries gate and held_snapshot forward unchanged." --complete then takes the --since gate form on its own (4171, previous gate non-null) and never prints 'nothing to complete' (4161).
+
+4. Risks (4691), Non-Goals (63) and Design (89): say the change is detected and gated rather than silent. Reconcile 4687's "land only in --complete's stage 1".
+
+5. P1-steps: add a case. A fake yay that bumps a kernel-set member's version opens gate, removes an existing flag, exits 1 with failed_step aur, and doesn't stamp. A following --complete runs kernel-modules-check in the --since form.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases (the contract) and the sections that cite it.
+:EVIDENCE:
+Spec lines:
+- 63: "An AUR upgrade that needs a newer kernel-set or DKMS-set package is the one ungated path, named in Risks."
+- 4142: "One that needs a newer kernel-set or DKMS-set package installs it ungated (Risks names this)."
+- 4691: "installs it ungated; no installed foreign package has such a dependency today."
+- 4206: "Only =--complete= writes gate."
+- 4108: everyday deferred set = post-transaction pending ∩ held. A package yay upgraded is no longer pending.
+- 4227: the everyday stamp needs only that every step exited 0 and packages is empty.
+- 4426: "REBOOT is hidden while gate is non-null ... Otherwise it is shown while the flag exists ... and is otherwise unchanged."
+- 4161: --complete with "nothing is pending and gate is null: print 'nothing to complete' ... exit 0".
+- 4164: the pre-write and gate entry happen only "when a kernel-set or DKMS-set member is pending, or gate is non-null".
+- 4687: "they land only in --complete's stage 1, behind the kernel-modules-check gate ... While it is open, nothing arms or offers a reboot."
+- 4203: "No arm flag exists while gate is non-null."
+- 4147-4151: step 9's 'otherwise' case re-arms.
+- 4558: rollout step (a) checks only "pacman -Qmq lists no DKMS-set package".
+
+Code:
+- ~/.dotfiles/maint/src/maint/gui.py:1509-1511 calls self.model.offer_reboot() when rid is update or topgrade and the primary event is ok.
+- ~/.dotfiles/maint/src/maint/probes/packages.py:131-146: reboot_required is true when /usr/lib/modules/$(uname -r) is gone.
+
+Commands:
+- /usr/share/libalpm/hooks has 70-dkms-install.hook, 70-dkms-upgrade.hook and 90-mkinitcpio-install.hook.
+- On ratio, the dependencies of the foreign packages show nothing kernel- or DKMS-related. linux-lts-strix depends on coreutils, initramfs and kmod; mkinitcpio-firmware depends on linux-firmware*, which is never a kernel-set member. So there is no instance today, as the spec and the finding both say.
+:END:
+
+** DONE REBOOT can stay pressable for up to 30 s after the gate opens
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 2, APPLY (4412) and Freshness and REBOOT (4426)
+
+REBOOT is hidden while gate is non-null, but only when the panel re-renders from a probe. APPLY's launcher triggers no re-probe, and the reboot remedy doesn't check the gate when it fires.
+
+Risk: If REBOOT was showing before APPLY (after a successful UPDATE, or while armed), it stays visible and can be confirmed during stage 1, until the next full tick. A shutdown sends pacman SIGTERM, which it doesn't handle.
+
+Recommended change: Two edits, both in the Phase 2 body.
+
+1. In Phase 2 "Freshness and REBOOT" (4426), append: "The reboot remedy re-reads the record (cache.get upgrade_deferred) when it fires. While gate is non-null it runs nothing and yields a fail event: 'gate open — do not reboot: <gate.failed>; run upgrade-guarded --complete'. This closes the window between --complete's pre-write and the panel's next probe, and covers a --complete started from a shell or by maint fix reboot." The check goes in doctor.iter_fix (GTK-free) so the panel key, REVIEW & FIX and the CLI all get it. Amend Invariants 4204 to say the same: "the panel hides REBOOT once it has probed, and the reboot remedy refuses at fire time while gate is non-null."
+
+2. Under P2-row (around 4447), add a test: "On a gate-open fixture, iter_fix('reboot') yields fail with the gate text and never calls the priv reboot argv, even with reboot_offer set."
+
+Optional: have the APPLY launcher also clear reboot_offer, so a stale post-UPDATE latch doesn't come back after a successful --complete that didn't need a reboot. An immediate re-probe at launch isn't needed, because the gate hasn't been written yet at that point.
+
+Should-fix, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases (the contract) and the sections that cite it.
+:EVIDENCE:
+Spec 4204: "No reboot is offered while gate is non-null: the panel hides REBOOT".
+Spec 4601-4602 (AC4): the gate is open "from the start of its stage 1" and "REBOOT is hidden whatever the arm flag, reboot_required or offer_reboot say", "whether the run started from the panel or a shell".
+Spec 4162-4168: the gate pre-write is --complete step 3, after step 1 (news capture, informant clear) and step 2 (-Sy, flag removal, closure probe), immediately before stage 1.
+Spec 4412-4413: the launcher only runs Popen(['foot','--hold','-e',*argv], start_new_session=True), "never _fire or doctor.iter_fix"; "the panel reads the outcome ... on its next probe".
+Spec 4426: the gate hides REBOOT, but the offer_reboot latch is otherwise kept.
+
+dotfiles maint/src/maint/gui.py:
+- 92: _FULL_SECONDS = 30.
+- 1317: the full refresh runs on a 30 s timer.
+- 1376-1384: _full_tick refreshes unless self.firing.
+- 1325-1327: refresh_full returns early while busy.
+- 491: the REBOOT key renders from model.reboot_key_visible().
+- 1510-1511: offer_reboot() runs after any ok update/topgrade lever.
+
+panel.py:
+- 306: reboot_offer is reset only in __init__.
+- 501-504: offer_reboot sets it True.
+- 506-510: reboot_key_visible returns the latch or reboot_required.
+
+The reboot remedy has no fire-time check:
+- remedies.py:314-320: kind priv, verb reboot, no record read.
+- priv.py:340-341, 376: argv ["systemctl", "reboot"].
+- doctor.py:192-252: iter_fix has no reboot precondition.
+
+The spec's own statements on what a SIGTERM does to stage 1 (rather than the SIGINT pacman honors):
+- 115: pacman honors SIGINT at a package boundary.
+- 2082-2086: 60-mkinitcpio-remove runs PreTransaction, and PostTransaction hooks don't run if the transaction fails.
+:END:
+
+** DONE R11: P2-apply never fires a key, so the roster-FIX routing it was meant to pin goes untested
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 2 APPLY L4412; Phase 2 tests intro L4438; P2-apply L4458; Guard UX rationale L4392
+
+L4412 puts the terminal-kind branch inside _press_lever: 'On the second, _press_lever calls a GTK-free panel.py launcher ... never _fire or doctor.iter_fix'. L4438 says 'No test imports gui.py or GTK'. So L4458 can only check that the row key and the roster's FIX 'resolve to apply_deferred with kind terminal', and separately that the launcher Popens the foot argv. Its 'iter_fix is never called' clause has no fire path to observe. The R11 resolution asked for a test 'fired from the row key and from the roster's FIX ... and iter_fix is never called'.
+
+Risk: If _press_lever's terminal branch is missing or wrong, APPLY from either the row key or the roster's FIX runs upgrade-guarded --complete through iter_fix inside the panel's daemon stream thread (gui.py:1443-1460). Closing the panel can then kill it mid-kernel-transaction, which is the failure F40/R11 closed. Every automated test would still pass, because L4458 never exercises the routing.
+
+Recommended change: Phase 2 APPLY, L4412: replace "On the second, =_press_lever= calls a GTK-free panel.py launcher that runs =Popen([...], start_new_session=True)=, never =_fire= or =doctor.iter_fix=." with "On the second, =_press_lever= delegates to a GTK-free =panel.press_fire(rid, items, fire)=. For a terminal-kind remedy it calls the panel.py launcher, which runs =Popen(['foot', '--hold', '-e', *argv], start_new_session=True)=. Otherwise it calls =fire= (gui's =_fire=). For the reason given at L4392, this is the only GTK-free seam for the choice, so it never reaches =_fire= or =doctor.iter_fix= for a terminal kind."
+
+Phase 2 tests, L4458: replace "...and the panel.py launcher calls the patched =subprocess.Popen= with ... =start_new_session=True=; =iter_fix= is never called" with "driving =panel.press_fire= with the row key's and the roster FIX key's rid and items calls the patched =subprocess.Popen= with =['foot', '--hold', '-e', 'upgrade-guarded', '--complete']= and =start_new_session=True=, and never calls the passed =fire= stub or a patched =doctor.iter_fix=. Control case: =panel.press_fire('update', [], fire)= calls =fire= and never calls Popen."
+
+Optional, not blocking. Verification: confirmed.
+Disposition: modified, folded into Implementation phases and the sections that cite it. Confirmed. This is the same defect as T02 and is merged there. T02's edit moves the terminal-kind choice into GTK-free panel.fire_press(rid, items, fire, fixture), with a zero-argument fire callable. That is T15's interface. The P2-apply rewrite drives fire_press with both the row key's and the roster FIX's rid and items, asserts that neither the fire stub nor iter_fix is called, and includes T15's update control case.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L4412: "On the second, =_press_lever= calls a GTK-free panel.py launcher ... never =_fire= or =doctor.iter_fix=."
+- L4438: "No test imports gui.py or GTK."
+- L4458: "the row key and the roster's FIX both resolve to =apply_deferred= with kind terminal, and the panel.py launcher calls the patched =subprocess.Popen= ...; =iter_fix= is never called"
+- L4392: "=_press_lever= lives in gui.py, which needs GTK, so this is the only GTK-free seam for the arm line."
+- L4414: iter_fix itself runs apply_deferred's argv with inherited stdio for the CLI form.
+- L4605 (AC4): "APPLY, from the row or from REVIEW & FIX's FIX, opens ... in a detached =foot= terminal and never runs it inside the panel." L4708: AC4 is verified by P2-apply among others.
+
+Code ~/.dotfiles/maint/src/maint/gui.py:
+- 1408-1438: _press_lever, where the second press calls self._fire(...) unconditionally.
+- 1474-1484: _fire, which calls doctor.iter_fix inside _stream's daemon thread (1443-1460).
+- 785, 954, 1440-1445: card keys, roster fix keys and strip keys all reach _press_lever. _fire has no other caller.
+- 94 and 1672: the existing MERGE Popen lives in gui.py and is untested.
+
+Tests: ~/.dotfiles/tests/maint/test_panel.py:3-7 says "The GTK view (maint/gui.py) is never imported here". gui appears only in that docstring.
+
+Mitigation: Phase 3 M-boot-armed says "On ratio and velox the test arms from APPLY".
+:END:
+
+** DONE R19: the tripped arm line drops the argv, so UPDATE names TOPGRADE's command
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 2 Guard UX L4393; Levers L4380
+
+L4393 reads: 'On a tripped read, the arm line for either key reads "<label> armed — will defer at least <guard matches> (kernels and DKMS modules are always held) — press again to run upgrade-guarded"'. UPDATE's argv is ['upgrade-guarded', '--no-topgrade'] (L4380), so UPDATE's tripped arm line names the bare 'upgrade-guarded', which is TOPGRADE's argv, and shows no argv at all.
+
+Risk: With a tripped guard read, UPDATE's act line tells the user a press runs the bare upgrade-guarded, the full sweep, when it actually runs upgrade-guarded --no-topgrade. That breaks maint's existing arm-line parity criterion on exactly the two levers this spec repoints. An implementer has to choose between the spec's literal string and the parity rule.
+
+Recommended change: At L4393, change the tail of the tripped string from "press again to run upgrade-guarded'" to "press again to run $ <argv>'", where <argv> is panel.arm_argv's string. That is 'upgrade-guarded --no-topgrade' for UPDATE and 'upgrade-guarded' for TOPGRADE, the same as the untripped line. At L4442, extend the bullet to read "=viewmodel.arm_line= with a tripped read gives the arm-line string with each label, ending '$ upgrade-guarded --no-topgrade' for UPDATE and '$ upgrade-guarded' for TOPGRADE".
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases (the contract) and the sections that cite it.
+:EVIDENCE:
+- Spec L4380: "UPDATE's argv is =['upgrade-guarded', '--no-topgrade']=; TOPGRADE's is =['upgrade-guarded']=."
+- Spec L4393: "On a tripped read, the arm line for either key reads '<label> armed — will defer at least <guard matches> (kernels and DKMS modules are always held) — press again to run upgrade-guarded'."
+- Spec L4442: "=viewmodel.arm_line= with a tripped read gives the arm-line string with each label". This pins the literal string.
+- ~/.dotfiles/maint/src/maint/viewmodel.py:99-107. The docstring says "The argv is the exact dry-run string (the parity criterion)". The guarded form ends "or apply from a TTY $ {argv}" and the plain form ends "press again to run $ {argv}".
+- ~/.dotfiles/maint/src/maint/panel.py:205-207. arm_argv is "The exact command string the arm-press shows — identical to the dry-run's argv (the parity acceptance criterion)".
+- ~/.dotfiles/maint/src/maint/gui.py:1396-1433. _guard_arm_line passes argv through to viewmodel.arm_line for guarded rids, and the plain path calls arm_line(title, argv).
+- ~/.dotfiles/tests/maint/panel_smoke.py:396-398 checks "act line names the exact argv (dry-run parity)". That check covers REMOVE, not UPDATE, but it shows the convention.
+:END:
+
+** DONE R27: only the row handles a malformed record; topgrade_freshness reads the same record through the shared helper with no defined behavior
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 2, The deferred row L4397 (shared helper) and L4399 (malformed record reads UNPROBED, 'never crashes the envelope'); Phase 2, APPLY L4409-4411 (topgrade_age items builder and doctor.review key on evidence deferred); The record L4221 (topgrade_freshness is a reader)
+
+The R27 resolution covers a missing or malformed record only for the upgrade_deferred row. N is computed 'through a helper shared with topgrade_freshness' (L4397), and topgrade_freshness 'adds evidence {deferred: N}, computed by the shared helper' (L4410). For a malformed record, the spec doesn't say what that helper returns or what topgrade_freshness reports. A missing record is covered (L4452).
+
+Risk: An implementer has to decide how topgrade_freshness handles a record whose data lacks packages. If the shared helper raises there, build_status raises, and every metric on the board is lost, not just the deferred row. The topgrade_age card's APPLY key and REVIEW & FIX's choice between APPLY and TOPGRADE also depend on that evidence value, and nothing defines it.
+
+Recommended change: At L4397, append one sentence: "The helper returns None for a missing record or one whose data lacks the documented fields (packages a list of {name, old, new, kind}); the row maps None-from-malformed to its UNPROBED text, and =topgrade_freshness= grades unchanged with evidence ={deferred: 0}=, so the =topgrade_age= card shows no APPLY and REVIEW & FIX keeps TOPGRADE."
+
+At L4453, change the test bullet to: "a record without packages reads UNPROBED, =topgrade_age= still grades from its stamp with evidence deferred 0 and no APPLY, and build_status returns an ok envelope."
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases (the contract) and the sections that cite it.
+:EVIDENCE:
+Spec: L4397 "through a helper shared with =topgrade_freshness="; L4399 the malformed-record UNPROBED rule is scoped to the row only; L4409 the items builder keys on "=topgrade_age= while its evidence deferred > 0"; L4410 "=topgrade_freshness= adds evidence ={deferred: N}=, computed by the shared helper"; L4411 doctor.review drops topgrade on evidence deferred > 0; L4452 covers only the no-record case for topgrade_age; L4453 "a record without packages reads UNPROBED" asserts the row only.
+
+Code:
+- ~/.dotfiles/maint/src/maint/status.py:41: topgrade_freshness is called bare inside the metrics list; the only try/except is at L22-24, around thresholds.load.
+- ~/.dotfiles/maint/src/maint/cache.py:35-42: get returns (entry["data"], age) with no shape validation, catching only OSError, ValueError, KeyError and TypeError on the envelope.
+- ~/.dotfiles/maint/src/maint/probes/__init__.py:5-6: "collectors never raise past their own metric".
+- ~/.dotfiles/maint/src/maint/probes/updates.py:115-133: the current topgrade_freshness has no evidence and no validation path.
+:END:
+
+** DONE R22: the snapshot hold runs after the gate line is re-emitted as the last stderr line, so a hold failure takes over the last line and detail
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): --complete step 5, L4173-4174; The snapshot hold L4253; The record, detail L4214; AC4 L4604; P1-complete L4353
+
+Step 5 lists 'fail: ... exit 4, with that line re-emitted as the run's last stderr line' and then 'after a --since-form gate on a ZFS root, pass or fail: the snapshot hold'. A hold or release failure 'prints snapshot hold failed: <reason> and changes neither result nor exit' (L4253), but the spec doesn't say which stream it goes to or where it falls relative to the re-emitted gate line. detail is defined as the run's last stderr line (L4214).
+
+Risk: Take a failing --since gate on a ZFS root where 'sudo -n zfs hold' also fails, for example when the NOPASSWD rule is missing. Implemented in the listed order, 'snapshot hold failed: …' (and zfs's own error) becomes the run's last stderr line. detail and the WARN fallback text then name the hold failure, not the failing gate item. That contradicts AC4 and the claim that a hold failure changes nothing but a warning.
+
+Recommended change: Three small edits:
+1. L4173: replace "with that line re-emitted as the run's last stderr line" with "detail = that line, and the finish re-emits it as the run's last stderr line, after the snapshot hold".
+2. L4253: append "to stderr, before the finish prints the run's reason line".
+3. At the end of the L4359 test bullet, add: "A fake zfs whose hold exits 1 prints 'snapshot hold failed', and result and exit are unchanged. After a failing gate the run still exits 4, with the gate item as its last stderr line and as detail."
+To cover the L4175 exit-1 path as well, add to Finish step 3 (L4200): "On a non-zero exit, print detail's last line to stderr last."
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: modified, folded into Implementation phases and the sections that cite it. Partially confirmed; the recommended change reflects the narrowed finding. In the listed order, a 'snapshot hold failed' line after the re-emitted gate line becomes the last stderr line and the detail. I kept the fix with three placement changes. 'Print detail last on a non-zero exit' lives in T01's single Finish edit, so it also covers the stage-1-failed exit-1 path. The hold-failure test lives in T11's single P1-complete hold-bullet rewrite. The failure bullet here also names a snapshot failure, for T11's pre-stage-1 snapshot.
+:EVIDENCE:
+Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- L4173: "fail: gate.failed = its last stderr line, result failed, failed_step gate, exit 4, with that line re-emitted as the run's last stderr line;"
+- L4174: "after a --since-form gate on a ZFS root, pass or fail: the snapshot hold;"
+- L4175: stage 1 failed and gate didn't fail → failed_step pacman, exit 1. The hold also runs on this path.
+- L4253: "A hold or release failure prints 'snapshot hold failed: <reason>' and changes neither result nor exit." No stream is given and no position relative to the reason line.
+- L4214: detail "is the run's last stderr line".
+- L4271: "On any non-zero exit, the last stderr line names the reason."
+- L4149: the closure-refusal precedent: "the refusal's own lines, which the finish prints last".
+- L4196-4200: Finish (compute, stamp, write) says nothing about printing the reason last.
+- L4353: P1-complete: a failed gate "exits 4 with that line last on stderr".
+- L4359: the hold test uses a fake zfs that only logs, with no failing-hold case.
+- L4604: AC4 requires exit 4 "with the failing item as its last stderr line and in the CRIT text".
+- grep: 'hold failed' appears only at L4253.
+No implementation exists yet: there is no upgrade-guarded or kernel-modules-check in scripts/ or tests/.
+:END:
+
+** DONE --apply-armed step 7's 'target not found' case can't be reached, because step 6's -Sp refuses it first with a different result and exit
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 1, --apply-armed sequence steps 6-7 (lines 4190-4191); Risks, Boot package source (line 4682)
+
+Step 6 runs pacman -Sp on the versioned list, and any failure exits 3 with result refused. Step 7 says 'A version the db no longer carries fails as target not found ... On failure: result failed, failed_step boot-transaction, exit 1.'
+
+Risk: The stale-flag-after-a-bare-refresh case in Risks ends as refused/3, not failed/1. A test written from step 7 expects the wrong result and exit code. The unit fails either way, so the impact is small.
+
+Recommended change: Line 4190: after "...exits 3: result refused, failed_step boot-transaction." insert "A version the db no longer carries (a refresh since arming that didn't re-arm) fails here as target not found."
+
+Line 4191: replace "A version the db no longer carries fails as target not found, and a missing cache file fails on download; both fail before any package changes." with "A missing cache file fails on download, before any package changes."
+
+Leave the rest of step 7 ("On failure: result failed, failed_step boot-transaction, exit 1.") unchanged. No edit is needed at line 4682.
+
+Optional, not blocking. Verification: confirmed.
+Disposition: accepted, folded into Implementation phases (the contract) and the sections that cite it.
+:EVIDENCE:
+Spec line 4190: "=LC_ALL=C pacman -Sp --print-format '%n' '<n>=<v>'...=, offline and read-only. A failure, or a kernel-set or DKMS-set member in its output, exits 3: result refused, failed_step boot-transaction."
+
+Spec line 4191: "=sudo -n pacman -S --noconfirm --needed '<n>=<v>'...= from the cache. A version the db no longer carries fails as target not found, and a missing cache file fails on download; both fail before any package changes. On failure: result failed, failed_step boot-transaction, exit 1."
+
+Spec line 4193: "There is no refresh, news capture, kernel step, gate, yay or sweep at boot, and no network dependency."
+
+Spec line 4682 (Risks, Boot package source): a bare pacman -Sy or yay between arming and reboot "breaks that; the boot run then fails before any package changes, the record names it".
+
+Ratio, read-only:
+- LC_ALL=C pacman -Sp --print-format '%n' 'zlib=0:0.0-1' prints "error: target not found: zlib=0:0.0-1" and exits 1.
+- LC_ALL=C pacman -Sp --print-format '%n %v' 'zlib=1:1.3.2-3' (the current sync version) prints "zlib 1:1.3.2-3" and exits 0.
+
+So the stale-version case fails at -Sp, and the stage-6 rule sends it to refused/3.
+
+grep shows "target not found" / "no longer carries" appears in the body only at line 4191. None of the test bullets at lines 4366-4369 asserts the step-7 outcome for a stale version.
+:END:
+
+** DONE 'The only list of root steps' and 'every root step is sudo -n' leave out yay's own sudo in step 7
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 1, Identity and privilege (lines 4064-4071); Everyday sequence step 7 (line 4140); Design (line 83); Readiness, Security (line 4650)
+
+Phase 1 says every root step is sudo -n, and lists them as 'the only list of root steps'. Everyday step 7 runs yay -Sua --noconfirm, which escalates its own pacman calls.
+
+Risk: The completeness claim that Readiness relies on is false. Without the NOPASSWD rule, a shell run prompts inside yay rather than failing with a reason as Design line 83 promises. A lever run fails only because there's no tty.
+
+Recommended change: This is the smallest edit. It's a wording fix, not a behavior change.
+- Line 4064: change "This is the only list of root steps:" to "This is the only list of the script's own root steps:".
+- Add one bullet after line 4070: "yay (Everyday step 7) and the topgrade sweep (step 8, through topgrade.toml's pre_sudo =sudo -v=) escalate on their own through plain sudo, not sudo -n. Both run only after steps 4 and 6 have succeeded under sudo -n, so a machine without the NOPASSWD rule fails at the refresh, with a reason, before either one runs."
+- Line 4650: change "is the only list of root steps" to "is the only list of the script's own root steps; yay and the sweep escalate through their own sudo (Phase 1, Identity and privilege)".
+- Line 83 can stay as written, because the refresh failing first keeps its promise true.
+
+Passing --sudoflags=-n to yay is a reasonable alternative. On its own it doesn't close the gap, because topgrade's pre_sudo still runs plain =sudo -v=.
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases (the contract) and the sections that cite it.
+:EVIDENCE:
+Spec lines in docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:
+- 4064: "Every root step is =sudo -n= ... This is the only list of root steps:" The listed items cover pacman, informant, the flag writes, zfs, reboot and lsinitcpio. Neither yay nor topgrade is there.
+- 4650: "Phase 1, Identity and privilege, is the only list of root steps."
+- 4137: step 4 is "One =sudo -n pacman -Sy=. On failure: failed_step refresh, exit 1."
+- 4139: step 6 is =sudo -n pacman -Su=, and "On failure ... steps 7 and 8 are skipped."
+- 4140: step 7 is "=yay -Sua --noconfirm=, always."
+
+Commands run on ratio:
+- =yay --help | grep sudo= lists --sudo, --sudoflags and --sudoloop.
+- ~/.config/yay/ doesn't exist.
+- =yay -Pg= shows "sudobin": "sudo", "sudoflags": "", "sudoloop": false. Installed version is yay 13.0.1-1.
+- ~/.config/topgrade.toml (symlink to .dotfiles/common/.config/topgrade.toml), lines 12 and 15: =pre_sudo = true= and =sudo_command = "sudo"=.
+- =topgrade --dry-run --no-ask-retry --disable system git_repos containers= prints "―― Sudo ――" and then "Dry running: /usr/bin/sudo -v". Installed version is topgrade 17.12.2-1.
+:END:
+
+** DONE Recovery depends on held_snapshot, but nothing the user sees ever shows it
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Phase 4, Recovery (line 4541); D-zbm (line 4575); Phase 1, The snapshot hold (line 4253); Phase 2, The deferred row evidence (line 4405); Readiness, Observability (line 4654)
+
+Recovery starts 'After booting the snapshot the record's held_snapshot names from ZFSBootMenu', and Design line 109 calls it 'the snapshot recovery boots'. The hold prints only on failure, and the row's evidence lists name, old, new, kind, detail and news, with no held_snapshot.
+
+Risk: At the ZFSBootMenu prompt after a failed boot, the person has to pick the right pre-pacman snapshot without having been shown its name. A newer pre-pacman snapshot holds the new, unverified kernel.
+
+Recommended change: Smallest edit, in Phase 4, Recovery (L4541): replace the opening clause with:
+
+"From ZFSBootMenu, boot the snapshot the hold pins. It is the record's held_snapshot. When the root won't boot, find it from the ZFSBootMenu recovery shell as the =<rootds>@pre-pacman_*= row with userrefs above 0 in =zfs list -H -t snapshot -o name,userrefs zroot/ROOT/default=, and confirm the tag with =zfs holds <snap>=. Never pick the newest pre-pacman snapshot, which can already hold the new kernel. Then, before the next =upgrade-guarded= run:"
+
+Optional, if a visible copy is wanted: in Phase 2, The deferred row, add held_snapshot to the evidence list (L4405), "...; held_snapshot when set; and the news lines...". The success-path print and the CRIT text change aren't needed.
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases (the contract) and the sections that cite it.
+:EVIDENCE:
+Spec lines:
+- L4253: "A hold or release failure prints 'snapshot hold failed: <reason>'" (nothing prints on success).
+- L4405: the evidence is "name, old, new and kind rows; detail in full; and the news lines", with no held_snapshot.
+- L4654: Observability lists the deferred set, result, detail, gate, armed state and news, with no snapshot.
+- L4217: held_snapshot exists only as a record field.
+- L4541: "After booting the snapshot the record's held_snapshot names from ZFSBootMenu".
+- L4688: Recovery is for an unplanned reboot while the gate is open, or a failure the gate misses (machine not booted).
+- L4575: in the drill, the record is read while the system is still up, so the drill is unaffected.
+
+Code:
+- scripts/zfs-pre-snapshot:27-28: names are pre-pacman_$(date +%Y-%m-%d_%H-%M-%S), so the name carries no marker of which snapshot is held.
+- docs/design/2026-07-15-velox-boot-failure-handoff.org:16-18: earlier velox triage was done from the ZBM recovery shell, so that shell is the realistic place where the name is needed.
+:END:
+
+** DONE After power loss during the boot transaction, the next boot surfaces nothing outside the panel
+CLOSED: [2026-10-05 Mon 12:10]
+Where (round 3): Readiness, Errors (4649); Phase 3, Disarm and failure (4496–4497); Phase 1, --apply-armed step 2 (4186)
+
+Readiness says every failure path ends at the tty1 login 'with the record naming the failure and the unit failed'.
+
+Risk: On velox with autologin, Hyprland may fail to start and drop to a tty1 shell that shows only 'Hyprland session ended', with no hint about db.lck or the remedy.
+
+Recommended change: Smallest edit, spec line 4649. Replace "Every failure path ends at the tty1 login, with the record naming the failure and the unit failed." with: "A failure, timeout or SIGKILL ends at the tty1 login in the same boot, with the record naming the failure and the unit failed. After power loss, the next boot skips the unit on its condition and no failed state survives, so the pre-written interrupted record (the deferred row, or maint status from a console) and the db.lck refusal are what remain."
+
+Optional: append to Risks 4684 after "names it with the manual remedy": "; maint's failed-units row does not show it, because the next boot skips the unit."
+
+A login-time notice or /etc/issue line is not needed. Add one only if I want a proactive tty surface.
+
+Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding).
+Disposition: accepted, folded into Implementation phases (the contract) and the sections that cite it.
+:EVIDENCE:
+- Spec 4649: "The boot unit is bounded and self-disarming ... success, failure, timeout, SIGKILL and power loss all disarm ... Every failure path ends at the tty1 login, with the record naming the failure and the unit failed."
+- Spec 4475: ConditionPathExists=/var/lib/archsetup/apply-upgrade-on-boot.
+- Spec 4484-4485: RuntimeDirectory=archsetup-boot-upgrade; ExecStartPre=+/usr/bin/mv -f <flag> /run/archsetup-boot-upgrade/armed.list. /run is tmpfs, so neither the flag nor the moved list survives power loss.
+- Spec 4497 and 4653: the failed_units promise covers "a failed or timed-out run" only.
+- ~/.dotfiles/maint/src/maint/probes/systemd.py:70-76: failed_units runs systemctl list-units --state=failed, which only sees units failed in the current boot. A unit skipped by its condition is inactive, not failed.
+- Spec 4186: the record is pre-written as interrupted / boot-transaction with the --complete remedy, so it survives power loss.
+- Spec 4089: db.lck present means exit 3 with detail 'pacman database locked (<path>): confirm no pacman is running, then remove it by hand'.
+- Spec 4684 (Risks) already accepts that "Power loss ... can still leave a half-extracted package and a stale db.lck ... every later transacting run refuses and names it with the manual remedy."
+- Spec 4517 and 4608: maint status is used as a CLI read of maint state, so the deferred row is not panel-only.
+- ~/.dotfiles/hyprland/.profile.d/99-hyprland-autostart.sh: on Hyprland exit it prints only "Hyprland session ended. Type 'start-hyprland' to restart." That part of the finding holds, but it is the generic behavior of any failed session.
+:END:
+
+** TODO Stale snapshot-hold clause in Decision 5's Consequences
+Where: round 3 residual from the final repair pass; the problem text names the passages.
+
+Gaps 3 and 4 replaced the 'only a later --complete or session moves the hold' clause at The snapshot hold (L5003), AC13 (L5390) and Design (L111). The same clause is still in the kernel Decision's Consequences at L193: 'the hold moves only when a later session holds a newer one'. That contradicts everyday step 7 (L4888), the first and last bullets of The snapshot hold (L4996, L5003), Invariants (L4954) and Readiness (L5401). The same sentence also still names the hold target as 'the oldest pre-pacman snapshot taken since a minute before the entry opened'. That is only the --since-gate fallback. Since T11, --complete step 3 (L4915) takes its own snapshot and holds it. The precedence line at L4797 makes Decisions non-normative, so no implementer is misled. That is why this is should-fix and not blocking.
+
+Recommended change: Decisions, 'Kernel set is held on every everyday run...', Consequences (L193). OLD: "On a ZFS root the session places a ZFS hold on the oldest pre-pacman snapshot taken since a minute before the entry opened, so the pre-snapshot prune can't destroy the fallback before it's needed; the hold moves only when a later session holds a newer one." NEW: "On a ZFS root, a session that opens the gate entry snapshots the root and places a ZFS hold on that snapshot before stage 1, and an everyday run whose AUR step opens the entry holds the oldest pre-pacman snapshot taken no earlier than a minute before that step, so the pre-snapshot prune can't destroy the fallback before it's needed; the hold moves only when a later session, or an everyday run whose AUR step opens the entry, holds a newer one."
+
+Should-fix, not blocking. Verified against the code in the final repair pass; left open because the review loop was bounded there.
+
+** TODO Untested everyday-run hold and release behind AC13
+Where: round 3 residual from the final repair pass; the problem text names the passages.
+
+AC13 (L5390) makes two claims about the everyday run. An everyday run whose AUR step opens the entry holds the oldest pre-pacman snapshot taken no earlier than a minute before that step. And, from the gap-3 edit, such a run 'holds a newer one and releases it'. Testing maps AC13 only to P1-complete and P1-installer (L5472), and neither group runs everyday step 7. The step-7 hold is checked only in P1-steps (L5075), which is not mapped to AC13. No case anywhere checks that step 7 releases the previous held_snapshot. So an implementation whose step 7 places the new hold but skips the release passes every case mapped to AC13. That leaves two snapshots with an upgrade-guarded hold and only one named in the record. It breaks 'exactly one extra snapshot per machine stays pinned' (L5003). Rollback step 3 (L5344) then releases only the named one, and Recovery's 'row with userrefs above 0' (L5299) finds two.
+
+Recommended change: (1) Phase 1 tests, P1-steps (L5075). OLD: "and on a ZFS root holds the oldest pre-pacman snapshot created at or after that time − 60 s and names it in held_snapshot;" NEW: "and on a ZFS root, starting from a record with gate null and held_snapshot naming an older snapshot, holds the oldest pre-pacman snapshot created at or after that time − 60 s, releases the previous held_snapshot and names the new one in held_snapshot;" (2) Testing / Verification / Rollout (L5472). OLD: "- AC13 (the held snapshot on a ZFS root): P1-complete and P1-installer." NEW: "- AC13 (the held snapshot on a ZFS root): P1-complete, P1-steps and P1-installer. P1-steps carries the everyday AUR-step hold and its release of the previous snapshot."
+
+Should-fix, not blocking. Verified against the code in the final repair pass; left open because the review loop was bounded there.
+
+** TODO First-failure precedence for failed_step and detail when several steps fail
+Where: round 3 residual from the final repair pass; the problem text names the passages.
+
+This ambiguity predates T01, but it sits in the rule T01 rewrote. Finish step 1 (L4946) takes detail from 'the failing step', in the singular. In one everyday run, three steps can each fail. Step 7 sets failed_step aur, and step 8 still runs (L4886). Step 8 sets failed_step topgrade (L4889). Step 9's arm sets failed_step download with detail 'disarmed: <reason>' (L4898). Nothing says which failed_step and detail the record keeps when two of them fail. An implementer who assigns in order keeps the last failure. Then a failed sweep overwrites step 7's '<names> moved by yay -Sua — do not reboot; run upgrade-guarded --complete' detail. It also makes AC10's 'failed_step is aur' false whenever the sweep fails too. No machine is put at risk: when step 7 opens the gate, gate.failed keeps the row CRIT and REBOOT refused.
+
+Recommended change: Finish, step 1 (L4946). After "On a non-zero exit, detail is the failing step's own reason, never the pre-written text." insert: "When more than one step fails in one run (everyday steps 7, 8 and 9 can), failed_step and detail are the first failing step's; each step's own failed_step is what the record holds when that step is the first failure."
+
+Optional, not blocking. Verified against the code in the final repair pass; left open because the review loop was bounded there.
+
+** TODO Ambiguous version rule in the power-cut recovery step
+Where: round 3 residual from the final repair pass; the problem text names the passages.
+
+Gap 6's fix landed, and the libalpm facts in it hold. In add.c commit_single_pkg, _alpm_remove_single_package(oldpkg, newpkg) runs at L485 and ends in _alpm_local_db_remove (remove.c L742). That comes before _alpm_local_db_prepare (L493) and extraction (L500 on). The 'upgraded/installed/reinstalled' logaction runs only after _alpm_local_db_write (L599-630). pacman -Dk's 'description file is missing' matches a prepared directory with no desc. local_db_populate skips such an entry as 'corrupted database entry', so the earlier reconcile -U --dbonly calls still work while the entry exists. The new version rule has two holes, though. (1) 'the one named by the last [ALPM] line' is ambiguous when that line is 'upgraded P (a -> b)' or 'downgraded P (a -> b)', since the line names two versions and the root holds the right-hand one. (2) The rule assumes an upgrade. A cut during a fresh install has no old entry: stage 1 lands every pending package outside the GPU closure, so new dependencies are realistic. Then there is no earlier line, or the last earlier line is 'removed P (x)'. Read literally, the rule then gives no version at all, or registers x, a package the booted root doesn't have. The closing pacman -Qkk over re-registered packages would then report missing files, and nothing tells the user what to do.
+
+Recommended change: Phase 4, Recovery, power-cut bullet (L5306). Old: "then =sudo pacman -U --dbonly= from the cache the version the booted root holds: the one named by the last =[ALPM]= line for that package timestamped before the snapshot's creation." New: "then =sudo pacman -U --dbonly= from the cache the version the booted root holds: the version that the last =[ALPM]= line for that package timestamped before the snapshot's creation left installed (on an upgraded or downgraded line, the one after =->=). When that line is a removed line, or there is no such line, the booted root doesn't have the package, so register nothing."
+
+Should-fix, not blocking. Verified against the code in the final repair pass; left open because the review loop was bounded there.
+
+** TODO Overbroad bare-wait qualifier in Exit codes and interrupts
+Where: round 3 residual from the final repair pass; the problem text names the passages.
+
+Gap 7's fix landed and its core claim matches the bash 5.3.20 man page: a bare wait waits for running background jobs and for the last-executed process substitution only when its PID still equals $!. The informant-clear example is also right, because --apply-armed starts the mirror at step 3 and runs the informant clear in the foreground at step 4. The qualifier 'which is true during any foreground step after the mirror starts' overclaims. $! changes as soon as the script starts another background job or process substitution. 'The trap waits on the transacting child's PID' means step 7's pacman runs as a background child, so during Finish's foreground stamp, with the trap still armed, $! is the transaction child and a bare wait returns at once. I checked this with two small test scripts under bash 5.3.20. A bare wait in the trap after a background child exited 130 without hanging. A later '< <(...)' moved $! off the mirror. A test written from 'any foreground step' would fail for any step after step 7. The rule itself is correct.
+
+Recommended change: Phase 1, Exit codes and interrupts (L5022). Old: "which is true during any foreground step after the mirror starts (the informant clear, for one)." New: "which stays true from the moment the mirror starts until the script next starts a background job or another process substitution, so during =--apply-armed='s informant clear, for one."
+
+Optional, not blocking. Verified against the code in the final repair pass; left open because the review loop was bounded there.
+
+** TODO ZFSBootMenu snapshot boot instead of a MOD+R rollback in Recovery and D-zbm
+Where: round 3 residual from the final repair pass; the problem text names the passages.
+
+The gap 8 fix landed at line 5338 and matches the harness. qemu_cmd in scripts/testing/lib/vm-utils.sh has no -no-reboot or -no-shutdown, so system_reset keeps QEMU running, and run-test.sh:139-141 does restore the clean-install snapshot. The problem is that the variant, the base drill (5336) and Recovery (5299) all say 'boot held_snapshot from ZFSBootMenu', but ZFSBootMenu can't boot a snapshot in place. In its snapshot view, ENTER duplicates the snapshot into a new boot environment (zfs send | recv), MOD+X clones and promotes, and MOD+C clones. Only MOD+R ('roll back snapshot') leaves zroot/ROOT/default as the root. Two concrete failures follow. (1) With MOD+C, the next --complete has the gate already open, so step 3 takes no snapshot, and the --since gate looks for <rootds>@pre-pacman_* (rootds from findmnt). The clone has no snapshots of its own, and zfs-pre-snapshot:11 hardcodes $POOL/ROOT/default. So every later --complete exits 4 on the snapshot item and the gate never closes, which is the mismatch-fails-closed case at line 926. ENTER and MOD+X leave the same hook/rootds split, so any entry everyday step 7 opens later fails the same way. (2) With MOD+X, the promote renames the held snapshot to <newBE>@pre-pacman_*. The variant's Expected ('zfs holds shows the upgrade-guarded tag ... and held_snapshot names it') then fails, and later releases of held_snapshot miss the real hold, so it stays pinned. Also, make test-keep starts QEMU with -display none and -serial file:, so the VM gives no keyboard input to ZFSBootMenu, and its zbm.timeout is 3 s (archangel installer:1115).
+
+Recommended change: In Phase 4, Recovery, replace "Boot the held snapshot from ZFSBootMenu: the record's held_snapshot, which the deferred row's evidence shows." with "Roll zroot/ROOT/default back to the held snapshot from ZFSBootMenu, then boot zroot/ROOT/default. The held snapshot is the record's held_snapshot, which the deferred row's evidence shows. Open the snapshot list (MOD+S), import read/write first if asked (MOD+W), select it, and press MOD+R ('roll back snapshot'). Never use ENTER (duplicate), MOD+X (clone and promote) or MOD+C (clone). Each one boots a new boot environment while zfs-pre-snapshot keeps snapshotting zroot/ROOT/default, so later gates fail closed on the snapshot item, and clone and promote also renames the held snapshot away from held_snapshot." In D-zbm line 5336, replace "Boot that snapshot from ZFSBootMenu, then run the Recovery reconcile." with "Roll back to that snapshot from ZFSBootMenu as Recovery says, sending the keys with the QEMU monitor's sendkey and watching with screendump (the VM has no display or serial input), then run the Recovery reconcile." In the variant (line 5338), replace "Boot held_snapshot from ZFSBootMenu and run Recovery." with "Roll back to held_snapshot from ZFSBootMenu the same way, and run Recovery."
+
+Should-fix, not blocking. Verified against the code in the final repair pass; left open because the review loop was bounded there.
+
+** TODO REBOOT-hidden claim in Design and Decision 5 that ignores the probe window
+Where: round 3 residual from the final repair pass; the problem text names the passages.
+
+The gap 10 fix landed in AC4 (line 5363). Its wording matches Phase 2, Freshness and REBOOT (5182-5183), and the strip REBOOT key really does reach doctor.iter_fix (gui.py _on_strip_key, then _press_lever, then _fire, then iter_fix at 1483), and P2-row covers the refusal. But the claim T14 showed false is still in two other places. Design line 105 says 'From that moment until a gate passes, the deferred row is CRIT with REBOOT hidden', with 'that moment' being the pre-stage-1 record write. Decision 'Kernel set is held...' Consequences (line 193) says 'From the start of a kernel-side stage 1 until a gate passes, the panel holds the deferred row at CRIT with REBOOT hidden'. The panel shows neither until its next probe (gui.py _FULL_SECONDS = 30). The fire-time refusal is what closes that window, and neither paragraph mentions it.
+
+Recommended change: Line 105: replace "From that moment until a gate passes, the deferred row is CRIT with REBOOT hidden, no arm flag exists, and the session never offers a reboot." with "From that moment until a gate passes, no arm flag exists and the session never offers a reboot. The panel shows the deferred row at CRIT with REBOOT hidden from its next probe, and a reboot fired before that probe refuses (Phase 2, Freshness and REBOOT)." Line 193: replace "From the start of a kernel-side stage 1 until a gate passes, the panel holds the deferred row at CRIT with REBOOT hidden, so" with "From the start of a kernel-side stage 1 until a gate passes, the panel holds the deferred row at CRIT with REBOOT hidden from its next probe, and maint's reboot remedy refuses at fire time, so".
+
+Should-fix, not blocking. Verified against the code in the final repair pass; left open because the review loop was bounded there.
+
+** TODO Reset cue and missed-extraction handling in the D-zbm power-cut variant
+Where: round 3 residual from the final repair pass; the problem text names the passages.
+
+The variant's cue to reset is when 'stage 1 prints its '(n/m) upgrading' lines'. pacman prints that form only when stdout is a tty, because it turns off the progress bar otherwise. Without one it prints 'upgrading <pkg>...' instead (pacman 7.1.0 has both format strings, '(%*zu/%*zu) %ls' and 'upgrading %s...'). A plain 'ssh host upgrade-guarded --complete' has no tty, so the cue as written never appears. The Expected also assumes a package was 'cut off mid-extraction', but a hand-sent system_reset can easily land between packages or during the PostTransaction hooks. Nothing tells the operator how to recognize that or what to do then.
+
+Recommended change: In the variant (line 5338), replace "run =upgrade-guarded --complete= over SSH. While stage 1 prints its '(n/m) upgrading' lines," with "run =upgrade-guarded --complete= over =ssh -t= (pacman prints '(n/m) upgrading <pkg>' only on a tty; without one it prints 'upgrading <pkg>...'). As the kernel headers package's upgrading line appears,". After "and run Recovery." add "If pacman.log shows the reset landed outside extraction (every target has an =[ALPM]= line, or 'running post-transaction hooks' is logged), repeat on a fresh =make test-keep FS_PROFILE=zfs= VM."
+
+Optional, not blocking. Verified against the code in the final repair pass; left open because the review loop was bounded there.
+
+** TODO yay -Pwq versus yay -Sua scope in the sudo-refusal test and Readiness
+Where: round 3 residual from the final repair pass; the problem text names the passages.
+
+The fixes for gaps 12 and 13 landed and match the code and tools: sudo 1.9.17p2's sudoers.so carries 'a password is required', the installer writes '%$username ALL=(ALL) NOPASSWD: ALL' (archsetup:1414), common/.config/topgrade.toml has pre_sudo = true and sudo_command = "sudo", and the boot form's first sudo step that isn't tolerated is step 7. The gap 11 wording is also correct: reboot_required() in maint/src/maint/probes/packages.py:131-146 is WARN only when /usr/lib/modules/$(uname -r) is gone, and the 60-depmod hook script rmdirs that tree once the package files leave. One new contradiction came in with the gap 13 test. The new P1-steps bullet asserts 'neither yay nor the sweep is invoked' when the fake sudo refuses -n at the refresh. Everyday step 2 runs =yay -Pwq= before the refresh (step 4), and another P1-steps bullet requires that it runs before the first transaction. So the fake yay is always invoked in that run, and an implementer who asserts on the fake yay's call log gets a test that can't pass. The gap 12 Readiness parenthetical says the same thing less sharply: 'before yay or the sweep runs', though yay -Pwq has already run. Identity and privilege ('yay (everyday step 7)') and Design ('escalate') are already scoped correctly.
+
+Recommended change: (1) Implementation phases > Phase 1 > Phase 1 tests > P1-steps, the fake-sudo bullet. OLD: "and as the run's last stderr line, and neither yay nor the sweep is invoked;" NEW: "and as the run's last stderr line, and neither =yay -Sua= nor the sweep is invoked (=yay -Pwq=, step 2, has already run);". (2) Readiness dimensions > Security & privacy. OLD: "(at the refresh in the everyday run and =--complete=, before yay or the sweep runs;" NEW: "(at the refresh in the everyday run and =--complete=, before =yay -Sua= or the sweep runs;".
+
+Should-fix, not blocking. Verified against the code in the final repair pass; left open because the review loop was bounded there.
* Implementation phases
+Three ordered commit groups across two repos: archsetup Phase 1, dotfiles Phase 2, then archsetup Phase 3. Phase 1's group holds both scripts, the installer step, the thresholds TOML rewrite, post-rebuild-check's check 9, the userrefs-aware prune in =scripts/zfs-pre-snapshot=, and the =docs/workflows/system-health-check.org= edit as a doc commit after the code. Phase 4 is the dotfiles =maint/README.md= flow-and-recovery commit plus the per-machine rollout. A phase may land as several commits if each is green on its suite (=make test-unit=, or =tests/maint/=).
+
+Ordering invariants:
+- Rollout step (a) on a machine follows Phase 1's last commit.
+- No Phase 2 commit reaches a machine before step (a) there.
+- The deferred row and APPLY land in one commit.
+- The lever repoint, the deletion of doctor.py's stamp, the guard-UX and =--force= removal, and the =sysupgrade= alias land in one commit.
+
+Where Design, Decisions, Acceptance criteria, Readiness, Risks or Testing restate a phase bullet, the phase bullet is normative.
+
+Every automated test uses unittest, never pytest. archsetup's =make test-unit= globs =tests/*/test_*.py=, so the new suites need no list edit. Each phase lists its named test groups once, and Testing maps the acceptance criteria to them. Manual tests are entries in =todo.org= under Manual testing and validation, each with what it verifies, one action per step and an Expected line.
+
** Phase 1 — The split-upgrade script (archsetup)
-=scripts/guarded-upgrade= (name open), installed to =/usr/local/bin= by the step that installs the guard. Behaviour as in Design: pending set → blocked set (hook =Target= lines, version-aware) ∪ held-kernel set when a compositor is live → =informant read= if present → =pacman -Syu --noconfirm --ignore=…= → =yay -Sua --noconfirm= → =topgrade --disable system,git_repos -y= → deferred set written to a state file → stamp only when nothing was deferred → exit 0 on a successful live part. The kernel set is derived from what is installed (every =linux*= kernel package and its =-headers=), never a hardcoded pair, and is always held or applied whole. Flags: =--dry-run= (print the plan and the deferred set, change nothing), =--no-topgrade=, =--no-aur=, =--complete= (the dedicated-session form: apply the kernel set live, run the gate, then arm the GPU/compositor set or, with no compositor live, apply it directly; stamp when the deferred set is empty). The gate is its own small script, =kernel-modules-check=: for each kernel under =/usr/lib/modules=, =dkms status= reports every registered module =installed= for it, and its initramfs is newer than its =vmlinuz=; on a ZFS root, a =pre-pacman_= snapshot of the root dataset exists. It exits non-zero with the failing item named, and =--complete= refuses to arm or reboot on that exit. Usable from a TTY at once. Tests (pytest beside the guard's): blocked-set computation against a fixture hook and version map; the kernel set is derived from the installed kernels and held whole on every everyday run; the =--ignore= list is exactly blocked ∪ kernel set; the state file round-trips; stamps only on an empty deferred set; =--dry-run= is IO-free; the gate passes and fails on fake =dkms status= output, image timestamps, and snapshot listings, and =--complete= never reaches the arm step on a failed gate.
+
+*** Files and install
+- =scripts/upgrade-guarded= and =scripts/kernel-modules-check= install to =/usr/local/bin= through the =hyprland()= installer step that installs =hypr-live-update-guard=. Both are usable from a TTY as soon as they're installed.
+- The guard hook is =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook=. The legacy unprefixed name is never read: it sorts after =60-mkinitcpio-remove=, so a blocked bare =pacman -Syu= can delete the initramfs before the guard aborts.
+
+*** Identity and privilege
+- =upgrade-guarded= runs as the invoking user, because the record and the stamp live in that user's =$HOME= and yay refuses root. EUID 0 exits 3 and leaves the record untouched.
+- Every root step the script itself runs is =sudo -n=, so a missing NOPASSWD rule (today =%cjennings NOPASSWD: ALL=) fails instead of prompting. This is the only list of the script's own root steps:
+ - pacman =-Sy=, =-Su=, =-S= and =-Sw=;
+ - =timeout 60 sudo -n informant read --all=, because =/var/lib/informant.dat= is 0664 root:informant and the user isn't in that group;
+ - the arm flag's =install=, =mv= and =rm=;
+ - =zfs snapshot=, =zfs hold= and =zfs release=;
+ - =systemctl reboot=;
+ - inside =kernel-modules-check=, =lsinitcpio <image>=, because mkinitcpio writes the images root 0600.
+- The boot unit's =ExecStartPre=+= move of the flag is the only root step outside sudo.
+- yay (everyday step 7) and the sweep (step 8, through =topgrade.toml='s =pre_sudo= =sudo -v=) escalate through their own plain sudo, not =sudo -n=. Both run only after steps 4 and 6 succeeded under =sudo -n=, so a machine without the NOPASSWD rule fails at the refresh, with a reason, before either runs.
+- Every pacman call whose output is parsed runs under =LC_ALL=C=.
+
+*** Modes and options
+- Exactly one mode per invocation: everyday (no mode flag), =--complete=, =--apply-armed= or =--dry-run=.
+- =--no-topgrade= is the only option and is valid only in everyday mode; it skips the sweep. Any other flag or combination exits 2 and writes no record.
+- everyday is the live split run. UPDATE runs =upgrade-guarded --no-topgrade=, TOPGRADE runs =upgrade-guarded=, and the =sysupgrade= alias runs =upgrade-guarded= where it is installed (Phase 2).
+- =--complete= is the dedicated session, run from any terminal or TTY, or from the detached terminal the panel's APPLY opens.
+- =--apply-armed= is the boot form. Only the boot unit runs it.
+- =--dry-run= prints the pending set, the held set with kinds, the GPU closure, the predicted deferred set and every argv.
+ - It writes nothing (no record, stamp or flag), uses no sudo and takes no lock.
+ - It creates a private =CHECKUPDATES_DB= with =mktemp -d= (removed on exit), runs =checkupdates --nocolor= against it (exit 2 means empty), and passes that path as every closure probe's =--dbpath=. =--nocolor= is required: without it, checkupdates passes =--color always= to its =pacman -Qu= whenever its stderr is a tty and pacman.conf sets =Color= (archsetup enables it), so a preview run from a terminal would read ANSI escapes into every package name. It never uses the shared =${TMPDIR:-/tmp}/checkup-db-${UID}=.
+
+*** Preconditions
+Checked in this order, before anything else runs. =--dry-run= writes no record on any of them.
+- Usage, per Modes and options: exit 2, no record write.
+- EUID 0: exit 3, record untouched.
+- Lock (every mode but =--dry-run=): an exclusive, non-blocking flock on =upgrade-guarded.lock= beside the record, creating the directory and the file if missing, through a descriptor held until the process exits (sh: =exec 9>>"$lock"; flock -n 9=). If it is held: 'another upgrade-guarded run is in progress', exit 3, record untouched.
+- =${UPGRADE_GUARDED_DB_LCK:-/var/lib/pacman/db.lck}= present (everyday, =--complete= and =--apply-armed=): exit 3, result refused, failed_step precondition, detail 'pacman database locked (<path>): confirm no pacman is running, then remove it by hand'. The script never deletes it.
+- The hook (seam =UPGRADE_GUARDED_HOOK=) unreadable or with no =Target= lines (every mode but =--apply-armed=): exit 3, result refused, failed_step precondition, the path named. It never proceeds with an empty pattern list, which is deliberately unlike maint's =guard.trips=.
+
+*** Sets
+- Pending set: the rows of =LC_ALL=C pacman -Qu= (=name old -> new=) against the system sync db, minus rows marked =[ignored]= (pacman.conf IgnorePkg, such as bridge-utils on ratio). Under =--dry-run= it is instead the lines =checkupdates --nocolor= prints against the run's private =CHECKUPDATES_DB= (Modes and options), which already drop =[ignored]= rows; exit 2 means empty, and any other non-zero exit ends the preview with exit 1 and writes nothing.
+ - For =pacman -Qu= (every mode but =--dry-run=), exit 1 with empty stdout and no =error:= line on stderr means empty. Warnings are tolerated. checkupdates' exit 1 (such as '==> ERROR: Cannot fetch updates') is never read as empty.
+ - Any other non-zero =pacman -Qu= exit is a failure. Before the transaction: failed_step refresh, exit 1, no transaction. After it: failed_step pacman, exit 1, packages = the pre-transaction pending set ∩ what the run held, and no stamp.
+- GPU patterns: the hook's =Target= lines verbatim, 15 today. Globs such as =mesa-*= stay globs: pacman accepts globs in =--ignore=, a pattern with nothing pending is a no-op, and =-Su= never reinstalls at the same version, so no version lookup is needed.
+- Kernel set: for each existing =${KMC_MODULES_DIR:-/usr/lib/modules}/<kver>/vmlinuz= (the glob expands with nullglob), its owner from its own =pacman -Qqo=, plus =<pkgbase>-headers= where installed, with pkgbase read from =<kver>/pkgbase=.
+ - A path no package owns is left out, and no match is an empty set.
+ - =linux-firmware*= and =linux-api-headers= are never members.
+ - The set is held whole by name, so a kernel and its headers move together. Only pending members move, so a foreign kernel (ratio's =linux-lts-strix=, a local build with no headers and no zfs module, the GRUB default) is never touched.
+ - Ratio has three kernels today (linux and linux-lts with their headers, plus linux-lts-strix); velox has linux-lts and its headers.
+- DKMS set: the owners of each existing =${UPGRADE_GUARDED_SRC_DIR:-/usr/src}/*/dkms.conf= (nullglob, one =pacman -Qqo= per path). No match is an empty set, not a failure. Today it is =zfs-dkms=, which comes from the archzfs repo on ratio. Its exact pin (=zfs-utils=2.4.4=) isn't listed by name: =zfs-utils= joins only through the closure, exactly when an upstream bump breaks the pin and never for a pkgrel-only bump.
+- live: =pgrep -x Hyprland= succeeds, the guard's own test (seam =UPGRADE_GUARDED_HYPR_RUNNING=1|0=). A console reached with Ctrl+Alt+F2 while Hyprland runs on tty1 is live.
+- Kinds: every held entry is kind kernel or kind gpu. kernel is the stage-A closure result (the kernel set, the DKMS set and what the closure adds to them). gpu is the pending GPU-pattern matches plus the stage-B additions, and under =--complete= the GPU closure.
+- Held set (everyday): the =--ignore= list. It always holds the stage-A result. Only when live does it add the GPU patterns, verbatim, and the stage-B result. Each entry is its own =--ignore=<entry>= argv element.
+- GPU closure (=--complete=): the closure seeded with the GPU patterns, whether or not live. If it contains a kernel-set or DKMS-set member, the run refuses (exit 3, result refused, failed_step closure, no transaction), because landing that member would put a kernel behind the boot form or ahead of the gate.
+- Deferred set, which is exactly the record's packages and covers the whole closure, not just the pattern matches:
+ - everyday: the post-transaction pending set ∩ the held set;
+ - =--complete=: the post-transaction pending set ∩ the GPU closure (kind gpu), plus the kernel-set members, DKMS-set members and previous kind-kernel entries still pending after the transaction (kind kernel). The kernel part is empty after a stage 1 that landed it;
+ - =--apply-armed=: the previous packages minus the entries installed at or above their =new= (=pacman -Q= plus =vercmp=);
+ - every refusal, and any run that ends before its first transaction (a refresh failure included), keeps the previous packages;
+ - 'nothing to complete' writes it empty;
+ - an IgnorePkg row never appears in it, because the pending set drops =[ignored]= rows.
+
+*** Dependency closure
+- Under =--noconfirm= one unresolvable dependent fails the whole =-Su= (pacman's skip prompt defaults to no), so the script has pacman compute what else has to wait before it transacts.
+- Probe: =LC_ALL=C pacman -Sup --noconfirm --print-format '%n' --ignore=<entry>...=. It is read-only, unprivileged, lock-free and download-free, and runs against the run's one refreshed db (=--dry-run= passes its private db as =--dbpath=).
+- Seeds:
+ - everyday stage A: the kernel set ∪ the DKMS set;
+ - everyday stage B, only when live: the stage-A result ∪ the GPU patterns;
+ - =--complete=: the GPU patterns.
+- Loop: run the probe with the seed plus every name added so far.
+ - Exit 0: stdout is the target list, and the loop ends.
+ - Non-zero exit: stderr must be exactly the one line 'error: failed to prepare transaction (could not satisfy dependencies)'. Each stdout line, with an optional leading =::= and its space stripped, is classified:
+ - "unable to satisfy dependency '<d>' required by <X>": add X if X is pending. Otherwise X is a dependency the transaction would newly install, and the line adds nothing, because the pending package that pulls X in is named on its own line in the same output;
+ - "installing <X> (<v>) breaks dependency '<d>' required by <Y>": add X if Y is held, else refuse. An unheld Y is installed and not pending, so it is an orphan, foreign or AUR package (the 2026-08-25 =qemu-block-gluster= case), and holding X would hold it and its dependents indefinitely instead of naming Y.
+ - There is no ignore list for warning or skip-prompt lines, because pacman suppresses them in print mode.
+- Refuse on another header reason, any other stderr line, a stdout line matching neither form (a conflict, a missing target, a replace), an iteration that adds no new name, or more than |pending|+1 iterations.
+- A refusal exits 3: result refused, failed_step closure, no transaction, never =-Rdd= and never force. detail holds pacman's stdout lines, then one line naming the refused packages. The db stays refreshed and nothing is upgraded.
+- The transaction that follows is =sudo -n pacman -Su --noconfirm --ignore=<entry>...= against the same db, with no second =-y=.
+
+*** Everyday sequence
+1. Preconditions.
+2. News: =yay -Pwq=, before the first transaction, because yay judges news against installed build dates. It needs no informant. Failure is tolerated.
+3. Informant clear: when =${UPGRADE_GUARDED_INFORMANT-informant}= is non-empty and found, =timeout 60 sudo -n informant read --all=. It needs =--all= because a bare read prompts between items, which crashes on null stdin or blocks on a tty, and the timeout because informant 0.6.0's feed fetch has none of its own. Failure, the timeout included, is tolerated and logged.
+ - A failed or timed-out clear leaves informant's =00-informant= hook live for this run's transaction. The transaction can then wait on the same unbounded fetch (on a lever run, ended at the runner's 3600 s timeout by a SIGINT to the run's process group (Phase 2, Levers), which the trap records as interrupted; a terminal run, APPLY's =--complete= included, is bounded only by Ctrl+C, which the trap records the same way) or abort on unread news before any package changes (failed_step pacman). This is the named residual.
+4. One =sudo -n pacman -Sy=. On failure: failed_step refresh, exit 1.
+5. The sets, closure stage A and, when live, stage B, then the held set. A refusal skips steps 6 to 8 and exits 3 after steps 9 and 10.
+6. =sudo -n pacman -Su --noconfirm --ignore=<entry>...=, with no =-y=. On failure: failed_step pacman, and steps 7 and 8 are skipped.
+7. =yay -Sua --noconfirm=, always. On failure: failed_step aur, and step 8 still runs.
+ - yay still resolves repo dependencies normally. An AUR upgrade that needs a newer held GPU/compositor package while Hyprland is live makes yay's own pacman call trip the guard, so the step fails.
+ - One that needs a newer kernel-set or DKMS-set package installs it ungated. So before yay the run records =pacman -Q= of the kernel-set and DKMS-set members, the kvers with an =installed= line in =${KMC_DKMS-dkms} status=, and T7 = now, and after yay, whatever its exit, it recomputes the kernel set and compares. If a member's version changed or a new kernel-set member appeared, it opens gate as =--complete= step 3 does: {pkgbases: the previous gate's pkgbases ∪ the pkgbases of kernels whose package or headers changed ∪ (on a DKMS-set change) the pkgbase of every kver that had an =installed= line before yay, since: the previous since or T7, failed: 'kernel-side package moved by the AUR step'}. The step then counts as failed: failed_step aur, detail '<names> moved by yay -Sua — do not reboot; run upgrade-guarded --complete', exit 1. When this opens the entry (the previous gate was null) on a ZFS root, it also runs the hold steps of The snapshot hold on the oldest =<rootds>@pre-pacman_*= with creation ≥ T7 − 60 s, which predates yay's transaction; a failure only prints 'snapshot hold failed: <reason>'. Before step 8 the run removes any flag with =sudo -n rm -f=, printing 'disarmed: gate open', then writes the record at once, as =--complete='s pre-write does, with packages carried forward and that gate, result failed, failed_step aur and that detail. So an interrupt, a kill or a power cut during step 8 leaves the entry on disk with its remedy, and no flag beside it. Step 8 still runs.
+8. The sweep, unless =--no-topgrade=: =/usr/bin/topgrade --no-ask-retry --disable system git_repos containers -y= (seam =UPGRADE_GUARDED_TOPGRADE=), each its own argv element. On failure: failed_step topgrade.
+ - It is called by absolute path, so the stamping =~/.local/bin/topgrade= wrapper is never in the chain.
+ - topgrade 17.12.2's clap parser rejects the comma-joined list with exit 2, and =-y= stays last because it takes optional step values.
+ - =--no-ask-retry= keeps a failed step from waiting on stdin whatever =topgrade.toml= says, because the panel's runner has no tty. containers is off because its step fails on locally built images.
+9. While the flag exists and step 4 succeeded, the first matching case applies:
+ - gate is non-null (step 7 opened it): =sudo -n rm -f= the flag and print 'disarmed: gate open';
+ - the unit isn't enabled (=systemctl is-enabled --quiet archsetup-boot-upgrade.service= fails; seam =UPGRADE_GUARDED_UNIT_ENABLED=): =sudo -n rm -f= the flag and print 'disarmed: boot unit not enabled'. It sets no failed_step and leaves the exit unaffected;
+ - the closure refused: remove the flag and print 'disarmed: closure refused' ahead of the refusal's own lines, which the finish prints last, so the disarm line stays out of detail. The run exits 3;
+ - no GPU-kind entries remain: remove the flag;
+ - otherwise: run the arm procedure on the current GPU-kind entries, so the flag's versions follow the db this run synced. On failure: remove the flag, failed_step download, detail 'disarmed: <reason>', exit 1.
+10. Finish. Exit 3 on a refusal; otherwise 0 iff every step that ran exited 0. A deferral alone never makes the exit non-zero.
+
+*** =--complete= sequence
+1. Everyday steps 1 to 3.
+2. Refresh, then clear the flag:
+ - =sudo -n pacman -Sy=. On failure: failed_step refresh, exit 1, and the flag is kept;
+ - then =sudo -n rm -f= the flag. The arm in step 6 is the only later writer;
+ - if the removal fails: result failed, failed_step precondition, detail 'cannot remove arm flag: <reason>', exit 1, before the pre-write and stage 1, so no gate entry opens while a flag exists;
+ - compute the sets and the GPU closure. A refusal exits 3;
+ - if nothing is pending and gate is null: print 'nothing to complete', write packages empty, exit 0, no stamp.
+3. Pre-stage-1 state:
+ - record =pacman -Q= of the kernel-set and DKMS-set members, and T0 = now in epoch seconds;
+ - when a kernel-set or DKMS-set member is pending, or gate is non-null, pre-write the record with packages carried forward:
+ - gate = {pkgbases: the previous gate's pkgbases ∪ the pkgbases of kernels whose package or =<pkgbase>-headers= is pending ∪ (when a DKMS-set member is pending) the pkgbase of every kver with an =installed= line in =${KMC_DKMS-dkms} status=, since: the previous since or T0, failed: 'kernel transaction not finished'};
+ - result interrupted, failed_step pacman;
+ - detail '--complete interrupted during the kernel transaction — do not reboot; run upgrade-guarded --complete'.
+ - after that pre-write, on a ZFS root and only when the previous gate was null (this run opens the entry): =sudo -n zfs snapshot <rootds>@pre-pacman_<YYYY-MM-DD_HH-MM-SS>= (=zfs-pre-snapshot='s name form, so its prune reclaims the snapshot once released), then the hold steps of The snapshot hold on it, then write the record again. A snapshot or hold failure prints 'snapshot hold failed: <reason>' and stage 1 still runs. When the gate was already open, held_snapshot is kept.
+4. Stage 1: =sudo -n pacman -Su --noconfirm --ignore=<GPU closure entry>...=. It lands the kernel set, the DKMS set and every other pending package outside the GPU closure. Old modules are removed PreTransaction and rebuilt PostTransaction, so a failed DKMS build aborts nothing; the gate is what catches it. When the previous gate was non-null, stage 1 also names as targets, after the =--ignore= list and without =--needed=, each recorded pkgbase that =pacman -Q= finds, its =<pkgbase>-headers= where installed, and every DKMS-set member. Their same-version reinstall fires =90-mkinitcpio-install= (which restores the preset and =/boot/vmlinuz-<pkgbase>=) and the DKMS install hook again, so a stage 1 cut off after the kernel committed, which skipped every PostTransaction hook, can still pass a later gate.
+5. Gate, whatever stage 1's exit, when gate is non-null:
+ - drop each recorded pkgbase for which =LC_ALL=C pacman -Q <pkgbase>= fails with 'was not found'. If none remain, gate = null and the gate is not invoked;
+ - otherwise invoke =kernel-modules-check= once over the remaining pkgbases: in the structural form (no =--since=) only when this run opened the entry (the previous gate was null) and stage 1 changed no kernel-set or DKMS-set version against step 3's =pacman -Q=; otherwise with =--since <gate.since>=;
+ - pass: gate = null;
+ - fail: gate.failed and detail = its last stderr line, result failed, failed_step gate, exit 4; the finish prints detail last, after the snapshot hold;
+ - after a =--since=-form gate on a ZFS root, pass or fail: the snapshot hold;
+ - if stage 1 failed and the gate didn't fail (it passed, every pkgbase dropped out, or gate was null): failed_step pacman, exit 1.
+6. GPU step, only when stage 1 succeeded and gate is null:
+ - nothing GPU-side pending: nothing;
+ - not live: stage 2, =sudo -n pacman -Su --noconfirm= on the same db with no ignore list, which lands exactly the remaining closure. On failure: failed_step pacman, exit 1;
+ - live and the unit enabled (the test and seam of everyday step 9): the arm procedure. On failure: failed_step download, exit 1;
+ - live and the unit not enabled: print 'boot unit not enabled — apply from a console with Hyprland stopped: upgrade-guarded --complete'. The GPU-kind entries stay in the record. This covers the time before rollout step (c), and a rollback that removes the unit.
+7. Finish, which writes the record. Then, only when the exit is 0, stdin is a tty (the APPLY terminal is one), and the run armed or its gate passed in the =--since= form, prompt 'Reboot now? [y/N]' before exiting. The default is no; yes runs =sudo -n systemctl reboot=. INT, TERM, HUP or EOF at the prompt count as no: the record isn't rewritten and the exit stays 0.
+
+*** =--apply-armed= sequence
+1. Preconditions (there is no hook check), then read =${UPGRADE_GUARDED_ARMED_LIST:-/run/archsetup-boot-upgrade/armed.list}=, the copy the unit's =ExecStartPre= moved off the flag path. The boot form never touches the flag path.
+ - A missing, empty or unparseable list, or a live Hyprland: exit 3, result refused, failed_step precondition, with the remedy.
+2. Pre-write the record: result interrupted, failed_step boot-transaction, detail 'boot upgrade interrupted — finish from a console with Hyprland stopped: upgrade-guarded --complete', everything else carried forward. It is written before the transaction, so SIGKILL and power loss leave it too.
+3. The first tty1 line is 'archsetup: applying N deferred GPU/compositor upgrades — do not power off', N being the list's entry count. All output is mirrored to the journal with =systemd-cat -t archsetup-boot-upgrade=, through a reader that ignores SIGINT (Exit codes and interrupts), for example =exec > >(tee >(systemd-cat -t archsetup-boot-upgrade)) 2>&1=.
+4. Informant clear, as everyday step 3. Offline, the clear and the hook both fail fast and see an empty feed, so neither blocks. A network that comes up between the clear and the transaction can let the hook abort it: failed_step boot-transaction, with the arm already consumed.
+5. Drop the entries installed at or above their recorded version (=vercmp=), so the boot form never downgrades; an entry installed below it, or not installed, stays. If none are left, the run is a no-op and goes to step 8.
+6. =LC_ALL=C pacman -Sp --print-format '%n' '<n>=<v>'...=, offline and read-only. A failure, or a kernel-set or DKMS-set member in its output, exits 3: result refused, failed_step boot-transaction. A version the db no longer carries (a refresh since arming that didn't re-arm) fails here as target not found. =pacman -S= resolves dependencies from the sync db, so this check is what keeps a kernel from riding in ungated.
+7. =sudo -n pacman -S --noconfirm --needed '<n>=<v>'...= from the cache. A missing cache file fails on download, before any package changes. On failure: result failed, failed_step boot-transaction, exit 1.
+8. Finish, exiting with the transaction's status (0 on a no-op).
+
+There is no refresh, news capture, kernel step, gate, yay or sweep at boot, and no network dependency.
+
+*** Finish
+Every mode but =--dry-run= ends here:
+1. Compute the record. result follows the exit: ok on 0, failed on 1 or 4, refused on 3. On exit 0, failed_step is null and detail is empty. On a non-zero exit, detail is the failing step's own reason, never the pre-written text. Where the step itself specifies a detail (Preconditions, Dependency closure, everyday steps 7 and 9, =--complete= steps 2 and 5), detail is that text. For any other failed pacman call, detail is its last =error:= line, or its last stderr line when it printed none, as when =sudo -n= refuses for want of the NOPASSWD rule and prints only its own message, which starts 'sudo:'. Any other failure gets one line naming the reason, such as the kernel-side member =--apply-armed= step 6 refused. These replace whatever the =--complete= or =--apply-armed= pre-write left, detail included. Finish never writes interrupted and never keeps the pre-written detail; only the interrupt trap does.
+2. If the stamp predicate holds, stamp. A missing maint or a failed stamp prints 'stamp failed: <reason>', sets result failed, failed_step stamp and detail 'stamp failed: <reason>' in the computed record, leaves packages, gate, news and held_snapshot unchanged, and makes the exit 1.
+3. Write the record once. On a non-zero exit, print detail to stderr last, so its last line is the run's last stderr line. From this write on, the interrupt trap writes nothing. Then exit, except that =--complete= first runs its step 7 prompt.
+
+*** Invariants
+- No arm flag exists while gate is non-null.
+- No reboot is offered while gate is non-null: the panel hides REBOOT once it has probed, the reboot remedy refuses at fire time (Phase 2, Freshness and REBOOT), and the prompt needs exit 0.
+- The gate entry opens before =--complete='s stage 1, or after an everyday AUR step that moved a kernel-side package (everyday step 7), and closes only when a gate passes or every recorded pkgbase drops out.
+- Only =--complete= and everyday step 7 write gate or move held_snapshot, and only =--complete= closes gate. Otherwise every mode carries gate and held_snapshot forward unchanged.
+
+*** The record
+- =${MAINT_STATE_DIR:-$HOME/.local/state/maint}/upgrade_deferred.json= (maint cache key =upgrade_deferred=), in maint cache.put's ={written_at: <epoch float>, data: {...}}= envelope, written atomically (=upgrade_deferred.json.tmp.<pid>=, then rename).
+- =data= fields:
+ - packages: =[{name, old, new, kind: gpu|kernel}]=, the deferred set;
+ - result: ok, failed, refused or interrupted;
+ - failed_step: null, precondition, refresh, closure, pacman, aur, topgrade, gate, download, boot-transaction or stamp;
+ - detail: empty on exit 0; otherwise a string whose last line names the reason and is the run's last stderr line. A closure refusal puts pacman's stdout lines first, then one line naming the refused packages;
+ - gate: null, or ={pkgbases: [...], since: <epoch>, failed: <item>}=. 'Gate open' means gate is non-null;
+ - news: up to 10 verbatim 'YYYY-MM-DD <title>' lines from =yay -Pwq= (which prints them oldest first), merged with the previous record's, deduplicated on the whole line, sorted, the latest 10 kept. Empty output with exit 0 means no new news;
+ - held_snapshot: null, or the name of the snapshot the hold protects.
+- There are no mode, gate.ok, gate.at, armed or boot fields. Armed means the flag exists, and a boot run's outcome goes in result.
+- failed_step precondition covers a missing or Target-less hook, db.lck, =--apply-armed='s list and Hyprland refusals, =--complete='s failed arm-flag removal, and an interruption during news capture or the informant clear.
+- Written by every mode's finish except =--dry-run= (refusals, failures and interruptions included), the =--complete= pre-write, the =--apply-armed= pre-write, everyday step 7's gate write and the interrupt trap. Never written by a usage error, EUID 0 or a held lock.
+- Writer: =upgrade-guarded= only. Readers: maint's =upgrade_deferred= probe, =topgrade_freshness= (its deferred count) and =upgrade-guarded= itself.
+- Fixtures: the canonical copy is archsetup =tests/upgrade-guarded/fixtures/upgrade_deferred.json=, equal to a fixed fake everyday run's record except written_at. A byte-identical copy sits at dotfiles =tests/maint/fixtures/upgrade_deferred.json=, and its test module's docstring names the archsetup path as its source. A field change updates both in one rollout.
+
+*** Stamp predicate
+- The stamp is =topgrade_run=, written through =${UPGRADE_GUARDED_MAINT:-$HOME/.local/bin/maint} stamp topgrade=, by absolute path because a system unit's PATH is =/usr/local/bin:/usr/bin=. It is never wrapped in =command -v ... || true=.
+- No form stamps while gate is non-null at finish.
+- everyday: =--no-topgrade= wasn't given; the refresh, =-Su=, yay and the sweep all exited 0; and packages is empty. UPDATE therefore never stamps.
+- =--complete=: it got past the nothing-to-complete check, started with non-empty packages or a non-null gate, every step it ran succeeded, and it finishes with packages empty and gate null. An arming run and the nothing-to-complete exit never stamp.
+- =--apply-armed=: the transaction exited 0 or was a no-op, and packages ends empty.
+- =--dry-run=: never.
+- On the driven path =upgrade-guarded= is the only writer. The =~/.local/bin/topgrade= PATH wrapper stamps bare-shell topgrade runs, and Phase 2 deletes doctor.py's lever stamp.
+
+*** The gate
+- Usage: =kernel-modules-check [--since <epoch>] <pkgbase>...=. At least one pkgbase is required; none, or an unknown option, exits 2.
+- It is a stateless reader: it writes nothing and makes no pacman call. Of the scripts, only =--complete= calls it, never with an empty list. Phase 4 also runs it by hand (rollout step (a) and CRIT-row diagnosis); a by-hand run writes nothing and never clears gate. The only supported re-check, and the only thing that clears gate, is =upgrade-guarded --complete=.
+- Per pkgbase it requires:
+ - at least one kver whose =${KMC_MODULES_DIR:-/usr/lib/modules}/<kver>/pkgbase= names it and whose =vmlinuz= exists, else '<pkgbase>: no module tree';
+ - =${KMC_BOOT_DIR:-/boot}/vmlinuz-<pkgbase>= exists;
+ - =initramfs-<pkgbase>.img= exists there and is newer than =vmlinuz-<pkgbase>= and, with =--since=, than since. It is never compared with the module-tree =vmlinuz=, whose mtime is the build date;
+ - for each such kver, =${KMC_DKMS-dkms} status -k <kver>= exits 0 and every line reads =installed=. When =KMC_DKMS= is empty or dkms isn't found (a machine with no DKMS module), the item passes;
+ - on a ZFS root (=findmnt -no FSTYPE /= is =zfs=): =sudo -n lsinitcpio <image>= exits 0, else '<pkgbase>: cannot read <image>', never reported as a missing module; and its listing has =zfs.ko= with any compression suffix. mkinitcpio writes an image without it and only warns, and that image is the velox failure.
+- With =--since= on a ZFS root, once per invocation: a =<rootds>@pre-pacman_*= snapshot (rootds is =findmnt -no SOURCE /=) with creation ≥ since − 60 s, read from =zfs list -H -p -t snapshot -o name,creation=. The 60 s matches the snapshot hook's skip window. On failure the item names the dataset.
+- Without =--since=, the newer-than-since and snapshot items are skipped.
+- Exit codes: 0 pass; 1 fail, with the last stderr line '<pkgbase>: <item>' or 'snapshot: <item>'; 2 usage.
+- A kver whose dkms status shows a module only 'added' (ratio's =linux-lts-strix=) never enters gate.pkgbases through the DKMS clause, and a foreign kernel is never pending, so the gate never checks one.
+
+*** The snapshot hold
+- =--complete= step 3 holds its own pre-stage-1 snapshot when it opens the entry, and everyday step 7 holds its pre-yay one when it opens the entry. After a =--since=-form gate on a ZFS root, pass or fail, and only when held_snapshot is null, no longer exists, or was created before gate.since − 60 s, the target is the oldest =<rootds>@pre-pacman_*= with creation ≥ gate.since − 60 s.
+- If the target exists and differs from held_snapshot:
+ 1. =sudo -n zfs hold upgrade-guarded <target>=;
+ 2. =sudo -n zfs release upgrade-guarded <held_snapshot>=, tolerating a missing snapshot;
+ 3. held_snapshot = target.
+- A snapshot, hold or release failure prints 'snapshot hold failed: <reason>' to stderr and changes neither result, exit nor detail; the finish's reason line still comes after it (Finish).
+- =scripts/zfs-pre-snapshot= lists =-o name,userrefs= and prunes only rows with userrefs 0, so a held snapshot is never destroyed and doesn't count toward KEEP.
+- I hold the snapshot because about ten later transactions would otherwise prune the only one with the old kernel, initramfs and modules before the gated kernel first boots. I take and hold the snapshot before stage 1, so a kill or a power cut mid-transaction already leaves the right snapshot pinned and named. The hold moves only when a later run holds a newer snapshot by the rules above (=--complete= step 3 or everyday step 7 opening an entry, or a =--since= gate that finds no valid hold), and each move releases the previous one, so exactly one extra snapshot per machine stays pinned. The structural form places no hold: it has no since to pick a snapshot by, and runs only when no kernel-side version changed.
+
+*** The arm flag
+- =/var/lib/archsetup/apply-upgrade-on-boot= (seam =UPGRADE_GUARDED_ARM_FLAG=), root:root 0644, on a persistent path so it survives the reboot the guard's =/run= sentinel cannot. The directory exists today, and Phase 3's installer step also creates it.
+- One =<name>=<version>= line per GPU-kind upgrade target, the version as pacman prints it, epoch included (=mesa=1:26.2.4-1=). Lines starting with =#= are ignored. Only upgrade targets are listed; new dependencies come from the cache at boot, so they keep their dependency install reason.
+- Arm procedure:
+ - (a) =LC_ALL=C pacman -Sp --print-format '%n' <names>= succeeds and names no kernel-set or DKMS-set member;
+ - (b) =sudo -n pacman -Sw --noconfirm <names>= while the network is up, which downloads the dependencies too. A download-only run returns before PreTransaction hooks, so neither the guard nor informant's hook fires under a live Hyprland;
+ - (c) a user temp file, then =sudo -n install -m 0644 -o root -g root <tmp> <flag>.new=, then =sudo -n mv -f <flag>.new <flag>=.
+- Writers:
+ - =--complete= removes it once, right after its refresh, and only its GPU step re-arms;
+ - everyday step 9 rewrites or removes it;
+ - the unit's =ExecStartPre= consumes it.
+- There is no =--disarm=; re-arming means running =--complete= again. The panel reads the flag's presence and entries only.
+
+*** Exit codes and interrupts
+- Exit codes: 0 ok; 1 a step failed; 2 usage; 3 refused before any transaction; 4 gate failed; 130 interrupted. On any non-zero exit, the last stderr line names the reason.
+- Every mode traps INT, TERM and HUP until Finish writes the record; a signal after that write never rewrites it (=--complete= step 7 covers the reboot prompt). The trap waits for the running child (systemd SIGKILLs whatever is left after =TimeoutStopSec=), writes result interrupted with the step (precondition during news capture or the informant clear) and the remedy (the pre-written detail where one exists), keeps gate as written, and exits 130. An everyday run interrupted after yay started and before step 7's record write first runs step 7's comparison. On a kernel-side move it also runs step 7's snapshot hold (when this opens the entry) and its =sudo -n rm -f= of any flag. The record the trap writes therefore holds the entry, names the held pre-yay snapshot, and has no flag beside it. Under =--dry-run= it writes nothing.
+- =db.lck= is never deleted.
+- No process that carries pacman's output (a journal mirror, a tee, a capture of stderr) may die of SIGINT. A SIGINT sent to the whole group (Ctrl+C in a terminal, the lever runner's timeout, the boot unit's KillMode=control-group) would otherwise leave pacman writing to a closed pipe, and pacman dies of SIGPIPE mid-package. Such a reader starts asynchronously with SIGINT ignored (in bash, a process substitution, which starts that way), never as a foreground pipeline element. The trap waits on the transacting child's PID, never with a bare =wait=: in bash a bare =wait= also waits for the last process substitution whenever its PID is still =$!=, which is true during any foreground step after the mirror starts (the informant clear, for one). The mirror can't exit while the script holds its write end, so the trap would never return.
+
+*** Other Phase 1 changes
+- Rewrite =[updates] guard_patterns= in =configs/maintenance-thresholds.toml= to exactly the hook heredoc's 15 Target patterns; the hook is the owner and stays unchanged. That drops =hyprland-*=, =hyprlang=, =hyprcursor=, =*wayland*=, =wlroots*=, =lib32-mesa*=, =lib32-vulkan-radeon= and =lib32-vulkan-intel=, and adds =wayland=, =libdrm=, =libglvnd=, =vulkan-mesa-layers=, =nvidia-utils=, =lib32-nvidia-utils= and =xorg-xwayland=. The comment's press-again text becomes 'display mirror of the hook's Target list, pinned by a test'.
+- =scripts/post-rebuild-check= gains check 9: where the guard is installed (=PRC_GUARD_BIN= overrides =/usr/local/bin/hypr-live-update-guard=, and set but empty means not installed), =10-hypr-live-update-guard.hook= exists in the hooks directory (=PRC_PACMAN_HOOKS_DIR= overrides =/etc/pacman.d/hooks=) and the unprefixed =hypr-live-update-guard.hook= doesn't; elsewhere the check is skipped. The header, the usage text and the 'check N/8' counters become 9.
+- =scripts/zfs-pre-snapshot='s prune becomes userrefs-aware, per The snapshot hold.
+
+*** The system-health-check edit
+A doc commit after Phase 1's code, to Phase 3 of =docs/workflows/system-health-check.org=:
+- Step 6 runs =upgrade-guarded= instead of plain topgrade on a machine where rollout step (a) is done.
+- A pending kernel or DKMS package lands through =upgrade-guarded --complete= as the dedicated session.
+- Until rollout step (c), a pending GPU/compositor set finishes from a console with Hyprland stopped.
+- Where step (a) isn't done and a kernel or DKMS package is pending, don't run plain topgrade.
+- Step 5's ratio strix addendum keeps its triggers and runs before whichever =upgrade-guarded= invocation lands each one: before =--complete= for linux, linux-lts or a held mesa, and before step 6's everyday run for linux-firmware, which is never held.
+- The "maint stamp topgrade by hand" fallback becomes 'the script stamps; never hand-stamp while a deferred set is outstanding'.
+- The 2026-09-12 containers KIL entry gets a resolution note appended rather than a rewrite.
+
+*** Seams and fakes
+- =upgrade-guarded= seams: =UPGRADE_GUARDED_HOOK=, =UPGRADE_GUARDED_HYPR_RUNNING=, =UPGRADE_GUARDED_ARM_FLAG=, =UPGRADE_GUARDED_ARMED_LIST=, =UPGRADE_GUARDED_TOPGRADE=, =UPGRADE_GUARDED_MAINT=, =UPGRADE_GUARDED_UNIT_ENABLED=, =UPGRADE_GUARDED_SRC_DIR=, =UPGRADE_GUARDED_DB_LCK= and =UPGRADE_GUARDED_INFORMANT= (empty means absent); maint's =MAINT_STATE_DIR=; and the gate's =KMC_MODULES_DIR= and =KMC_DKMS=, shared so one fixture tree drives the sets, the pkgbase resolution and the gate.
+- =kernel-modules-check= seams: =KMC_MODULES_DIR=, =KMC_BOOT_DIR= and =KMC_DKMS= (empty means absent).
+- =post-rebuild-check= seams: =PRC_GUARD_BIN= and =PRC_PACMAN_HOOKS_DIR=. Every test env, run_check's defaults included, sets =PRC_GUARD_BIN= empty, so the suite never reads the real =/etc=.
+- Fakes on PATH: pacman, sudo (honouring =-n=), yay (date-prefixed news lines, oldest first), informant, systemctl, systemd-cat, vercmp, checkupdates, dkms, zfs (=tests/zfs-pre-snapshot/fake-zfs=, which gains hold, release and userrefs), findmnt and lsinitcpio.
+- The tests drive both scripts through =subprocess=.
+
+*** Phase 1 tests
+In =tests/upgrade-guarded/= and =tests/kernel-modules-check/= unless a group names another directory.
+- P1-sets:
+ - the =--ignore= list is exactly the fixture hook's patterns plus the kernel and DKMS sets plus the closure names, each its own element;
+ - with =UPGRADE_GUARDED_HYPR_RUNNING= forced each way, the everyday list holds the kernel and DKMS sets on both branches and the GPU patterns only on the live one, and =--complete='s stage 1 holds the GPU closure on both;
+ - a kernel-set fixture under =KMC_MODULES_DIR= has three vmlinuz owners (one foreign and headerless), =linux-firmware*= and =linux-api-headers= installed, an unowned vmlinuz path, and zfs built for two of the three kernels. Firmware and api-headers are never held, the unowned path is left out, the kernel set is held whole, and the foreign kernel never enters the pending set or gate.pkgbases;
+ - a pending upstream =zfs-dkms=/=zfs-utils= bump is held as a pair; a =zfs-utils= pkgrel-only bump with =zfs-dkms= unchanged is not held;
+ - with no dkms.conf under =UPGRADE_GUARDED_SRC_DIR=, the DKMS set is empty and the run proceeds;
+ - a fake =-Qu= exiting 1 with empty stdout and a warning line reads as empty; one that prints an =error:= line and exits 1 fails the step (failed_step refresh), keeps packages and doesn't stamp; under =--dry-run=, a fake checkupdates that exits 1 with empty stdout and no =error:= line ends the preview with exit 1, prints no pending set and writes nothing;
+ - the deferred set comes from a fake post-transaction =-Qu= intersected with the held set, and excludes an =[ignored]= row;
+ - a missing or Target-less hook fails closed: exit 3, result refused, failed_step precondition, the path named, no transaction.
+- P1-closure:
+ - the fake pacman replays pacman 7.1's failing =-Sup= transcript (the header on stderr, the =::= lines on stdout, exit 1). A failing probe carrying only well-formed lines adds names and doesn't refuse, and lines parse alike with and without =::=;
+ - forward: =hyprutils= (soname 12 → 13) is held, and =hyprlang=, which requires the new soname, is added as kind gpu;
+ - forward through a new dependency: the fake emits the newlib line and the ptool line, with newlib neither pending nor installed. ptool is held with its stage's kind, the run exits 0, and =-Su= runs;
+ - reverse: an unguarded soname bump that a held =hyprland= depends on is added;
+ - orphan: a reverse break whose requirer is unheld and not pending refuses (exit 3, result refused, failed_step closure), no =-Su= runs, and detail holds pacman's stdout lines, then one line naming the refused packages;
+ - a header with another reason, a stray stderr line, an unparseable stdout line, an iteration that adds nothing, and more than |pending|+1 iterations each refuse;
+ - the run calls =sudo -n pacman -Sy= once, and the =-Su= call carries no =-y=.
+- P1-steps:
+ - EUID 0 refuses with exit 3 and leaves the record untouched (run under =unshare --map-root-user=); combining modes, an unknown option, or =--no-topgrade= outside the everyday mode exits 2 and writes no record;
+ - while the test holds a real flock on the lock file, the run exits 3 and the record is unchanged; with the lock free, the run gets past the precondition;
+ - a file at =UPGRADE_GUARDED_DB_LCK= makes the everyday run, =--complete= and =--apply-armed= each exit 3 before any transaction, with result refused and failed_step precondition, naming the file and the remedy; the file is still present afterwards;
+ - a fake sudo that refuses =-n= on the refresh, as a missing NOPASSWD rule does (exit 1, one 'sudo: a password is required' line, no =error:= line), gives exit 1, result failed and failed_step refresh, with that line as detail's last line and as the run's last stderr line, and neither yay nor the sweep is invoked;
+ - =yay -Pwq= runs before the first pacman transaction in the everyday run and =--complete=, and =--apply-armed= never calls it. With the fake yay's date-prefixed lines, the merge with the previous record's lines, the whole-line dedup and the 10-latest cap are asserted;
+ - =timeout 60 sudo -n informant read --all= runs before the first transaction in the everyday, =--complete= and =--apply-armed= forms. A failing informant, and =UPGRADE_GUARDED_INFORMANT= set empty, are tolerated;
+ - a fake =informant read --all= that exits 124, followed by a fake pacman that fails as the hook's abort would, gives failed_step pacman (everyday) or boot-transaction (=--apply-armed=), no stamp and no re-arm;
+ - the sweep argv is asserted element by element. One real-binary check, skipped when =/usr/bin/topgrade= is absent, runs =/usr/bin/topgrade --no-ask-retry --disable system git_repos containers -y --version= and expects exit 0. A failing fake sweep gives failed_step topgrade, exit 1 and no stamp;
+ - =yay -Sua= runs on every everyday run. A failing fake yay still lets the sweep run, records failed_step aur, doesn't stamp, and exits 1;
+ - a fake yay that bumps a kernel-set member's version opens gate (failed 'kernel-side package moved by the AUR step', since the time recorded before yay), removes an existing flag and writes the entry to the record before the fake sweep starts, exits 1 with failed_step aur and doesn't stamp, and on a ZFS root holds the oldest pre-pacman snapshot created at or after that time − 60 s and names it in held_snapshot; an INT during that sweep, or during a fake yay that has already bumped the member, leaves the record with gate open and result interrupted and no flag present, and on a ZFS root with the pre-yay snapshot held and named in held_snapshot; a following =--complete= runs =kernel-modules-check= in the =--since= form; a fake yay that bumps =zfs-dkms= and leaves the fake dkms status showing zfs only 'added' afterwards opens gate with the pkgbase of each kver that showed =installed= before yay, so gate.pkgbases is never empty;
+ - =--dry-run= writes no record, stamp or flag, calls no sudo and takes no lock. Every probe's =--dbpath= equals the =CHECKUPDATES_DB= the fake checkupdates received, and that path is not the shared default. With the fake checkupdates and the fake =pacman -Qu= given different rows, the printed pending set is the checkupdates rows, and the fake checkupdates received =--nocolor=;
+ - an INT while a fake step runs leaves result interrupted with that step and the remedy, and exits 130. An INT during news capture or the informant clear records failed_step precondition.
+- P1-record-stamp:
+ - with a stamping fake =topgrade= first on PATH and =UPGRADE_GUARDED_TOPGRADE= pointing at a non-stamping fake, a run that defers a non-empty set leaves =topgrade_run.json= untouched and never invokes the PATH fake;
+ - a TOPGRADE-shaped everyday run whose refresh, =-Su=, yay and sweep fakes exit 0, and whose post-transaction =-Qu= leaves nothing held, writes =topgrade_run=;
+ - that same run with =UPGRADE_GUARDED_MAINT= pointing at a missing file exits 1, and the record reads result failed, failed_step stamp, detail 'stamp failed: <reason>', with packages unchanged;
+ - that same run started from a record with gate non-null carries gate and held_snapshot forward unchanged and doesn't stamp;
+ - an UPDATE-shaped run (=--no-topgrade=) with nothing deferred doesn't stamp;
+ - a refresh failure (exit 1), a closure refusal (exit 3) or a pacman failure (exit 1) skips the AUR and sweep steps, still writes the record, and doesn't stamp. The refresh failure and the refusal keep the previous packages;
+ - after a successful =--complete= on the stage-2 path that started with a kernel-set member pending (so its pre-write ran), or a successful =--apply-armed=, packages is empty, the record reads result ok with failed_step null and detail empty, and the stamp is fresh;
+ - from a non-empty record whose packages are all installed, with nothing pending and gate null, =--complete= prints 'nothing to complete', writes an empty record, exits 0 and doesn't stamp; an arming =--complete= doesn't stamp;
+ - a fixed fake everyday run's record equals =tests/upgrade-guarded/fixtures/upgrade_deferred.json= in every field except written_at.
+- P1-gate (=tests/kernel-modules-check/=):
+ - on fake dkms status output, image timestamps, lsinitcpio listings and snapshot listings, a good fixture passes and each item fails on its own: a module not =installed= for a kver (last stderr line '<pkgbase>: <item>'), a missing =vmlinuz-<pkgbase>= in the boot dir, an initramfs older than =vmlinuz-<pkgbase>=, with =--since= an initramfs newer than the module-tree vmlinuz but older than since, and on a ZFS root an image without =zfs.ko=;
+ - a pkgbase that no kver names fails as '<pkgbase>: no module tree';
+ - the fake lsinitcpio fails unless invoked through the fake sudo, and the gate passes reading through sudo. lsinitcpio exiting 1 with no output fails as '<pkgbase>: cannot read <image>';
+ - with =--since=, a listing whose newest root-dataset pre-pacman snapshot was created before since − 60 s fails, naming the dataset. One created at exactly since − 60 s, one in [since − 60 s, since) (the snapshot hook's 60 s skip window) and one after since each pass;
+ - without =--since=, an image older than any since passes and no snapshot is required;
+ - =KMC_DKMS= set empty passes the module item;
+ - no pkgbase, or an unknown option, exits 2.
+- P1-complete:
+ - during a fake stage 1, the record's gate is non-null (failed 'kernel transaction not finished', result interrupted, failed_step pacman) and no flag exists;
+ - INT or HUP during stage 1 exits 130, with the entry kept and any earlier flag gone;
+ - a single stage 1 that moves linux and zfs-dkms together: the old linux kver is never checked, the new linux kver and the untouched linux-lts kver are, and with good fakes the gate passes;
+ - a DKMS-set-only change puts in gate.pkgbases the pkgbase of every kver with an =installed= dkms status line, and never a strix-shaped kver whose zfs is only 'added';
+ - a stage 1 that changes no version but leaves =vmlinuz-<pkgbase>= missing exits 4;
+ - a stage 1 that fails before committing, with =/boot= intact, runs the structural form, clears the gate and exits 1 with result failed, failed_step pacman, gate null and the kernel-side entries still in packages as kind kernel;
+ - a stage 1 that exits non-zero after the kernel landed, with a passing gate, exits 1 with result failed, failed_step pacman, gate null, the kernel gone from packages, and a detail and last stderr line that name the stage 1 failure, not the pre-written interrupted text;
+ - a failed gate sets result failed, failed_step gate and gate.failed to the gate's last stderr line, exits 4 with that line last on stderr, and reaches neither the arm step nor the prompt; the gate field survives a following everyday run;
+ - a follow-up =--complete= after an interrupt or a failed gate runs the =--since= form against the recorded since, and neither arms nor prompts until it passes, which sets gate null;
+ - after an interrupted stage 1 whose kernel landed (=vmlinuz-<pkgbase>= missing from the fake boot dir, nothing kernel-side pending), the follow-up's stage 1 argv names the recorded kernel package, its headers and the DKMS-set members as targets without =--needed=; with a fake pacman that restores the image on that reinstall, the gate passes and sets gate null;
+ - a recorded pkgbase whose package was removed drops out. With none left, =--complete= sets gate null without invoking the gate and, with packages otherwise empty, stamps. A recorded pkgbase still installed but with no module tree fails the gate with exit 4;
+ - a =--complete= with nothing kernel-side pending and no open gate never invokes =kernel-modules-check=;
+ - with a flag present, a =--complete= closure refusal, a failed stage 1 and a failed gate each leave no flag; a refresh failure leaves it in place;
+ - a GPU closure that reaches a kernel-set member refuses (exit 3, failed_step closure) before any transaction and leaves packages unchanged;
+ - a fake zfs logs snapshot, hold and release. On a ZFS root, a =--complete= that opens the entry snapshots and holds =<rootds>@pre-pacman_<ts>= after the pre-write and before stage 1, releases the previous held_snapshot, and the record read during the fake stage 1 already names the new snapshot in held_snapshot; a follow-up with the gate already open takes no snapshot before stage 1;
+ - from a record whose open gate has since S and whose held_snapshot predates S − 60 s, with a listing of four root-dataset pre-pacman snapshots (created at S − 61 s, at exactly S − 60 s, inside the window, and after a later everyday run), a passing and a failing =--since= gate each hold the S − 60 s snapshot, release the previous one and leave held_snapshot naming it; a held_snapshot created at or after S − 60 s is kept and nothing is re-held; a structural-form gate adds no hold after the gate, so in that run the fake zfs logs only step 3's snapshot and hold;
+ - a fake zfs whose snapshot or hold exits 1 prints 'snapshot hold failed' and changes neither result nor exit, stage 1 still runs, and after a failing gate the run still exits 4 with the gate item as detail and as its last stderr line.
+- P1-arm-boot:
+ - with the unit reported enabled and Hyprland live, =--complete='s arm runs the resolve check, =sudo -n pacman -Sw= on the GPU-kind names, =sudo -n install= and =sudo -n mv=, and writes their =<name>=<version>= lines to the flag;
+ - with the unit reported not enabled, =--complete= never writes the flag, prints the 'boot unit not enabled' line and never prompts;
+ - =--complete= prompts for a reboot only when the exit is 0 and stdin is a tty, and an INT or HUP at the prompt runs no reboot, leaves the record byte-identical to the one Finish wrote, and exits 0;
+ - an everyday run with the flag present rewrites it to the current GPU-kind entries and removes it when none remain. On a closure refusal it removes the flag and prints 'disarmed: closure refused', with failed_step closure and exit 3. On a resolve-check or =-Sw= failure it removes the flag, with failed_step download, detail 'disarmed: <reason>' and exit 1;
+ - with the flag present and the unit reported not enabled, an everyday run removes the flag, prints 'disarmed: boot unit not enabled', and never runs =-Sp= or =-Sw=;
+ - =--apply-armed= targets exactly the recorded list, never passes =-y= or refreshes (the fake pacman asserts the argv), and drops an entry already at or above its recorded version without downgrading;
+ - =--apply-armed= never includes a kernel-set or DKMS-set package, and refuses (exit 3, failed_step boot-transaction) when =pacman -Sp= would pull one in as a dependency or reports a stale version as target not found;
+ - =--apply-armed= pre-writes the interrupted record before the transaction, prints the banner as its first line, and sends its output through the fake =systemd-cat=;
+ - run under =setsid=, a SIGINT to the whole process group during a fake =--apply-armed= transaction (a fake pacman that traps INT, writes to stderr after the signal, removes a fake lock file and exits 130) leaves the fake pacman not killed by SIGPIPE, the fake lock gone, and the record interrupted;
+ - a missing, empty or unparseable list, or a live Hyprland, makes =--apply-armed= exit 3 with result refused and failed_step precondition; a held lock exits 3 with the record unchanged.
+- P1-installer:
+ - =tests/installer-steps/=: a source-inspection test that the step installing =hypr-live-update-guard= also installs both scripts to =/usr/local/bin=, and the TOML pin test, which compares the sorted seed-TOML patterns with the Target lines parsed from the heredoc the archsetup script writes to =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook=, located by that exact path (the =test_pacman_hook_order.py= idiom), and reads neither =/etc= nor =~/.config=;
+ - =tests/post-rebuild-check/=: check 9 with the guard present and absent, and with the legacy and =10-= hook names, through the =PRC_= seams; the '/8' assertions become '/9';
+ - =tests/zfs-pre-snapshot/=: a held snapshot (userrefs above 0) is never destroyed and doesn't count toward KEEP.
+- M-split-console (manual, in the =make test-keep FS_PROFILE=zfs= VM, with pending packages created per Phase 3's VM procedure): run =upgrade-guarded --no-topgrade= from a console with Hyprland not running. Expected: exit 0; the pending GPU-pattern packages land; the record and =pacman -Qu= list only kernel-kind entries (plus =[ignored]= rows in =-Qu=).
** Phase 2 — Wire maint to it (dotfiles)
-UPDATE and TOPGRADE levers change their =argv= to the script; the press-again-to-force sentinel wrap goes away (the driven path never trips the guard). A new probe reads the deferred-set state file and the panel renders "N deferred — apply on reboot" as its own row. A test asserts the TOML =guard_patterns= equal the installed hook's =Target= list. Tests under the maint fake harness.
+The Phase 2 commits are pulled on a machine only after Phase 4's step (a) there: the levers and the alias call =upgrade-guarded= and the force path is gone, so pulling them first breaks UPDATE and TOPGRADE on that machine.
+
+*** Levers
+- UPDATE's argv is =['upgrade-guarded', '--no-topgrade']=; TOPGRADE's is =['upgrade-guarded']=. Both keep kind user, timeout 3600 and always True, stay on the existing lever runner, and stay strip keys (gui.py:496).
+- The runner's timeout handling changes for user-kind remedies marked long, today only these two: =_execute_step= starts them with =start_new_session=True= and keeps reading their output. When the timeout passes, it sends SIGINT to the process group, waits up to 120 s, then SIGKILLs the group and reports 'timed out — interrupted'. It never sends SIGTERM, which pacman doesn't handle. pacman stops at a package boundary, and =upgrade-guarded='s trap records the run as interrupted (exit 130).
+- When =upgrade-guarded= isn't on PATH, the lever reads 'upgrade-guarded not installed (archsetup one-time install)'. There is no fallback to the old argv.
+- Delete doctor.py:169-170, the branch that stamps =topgrade_run= for the topgrade remedy, so on the driven path the script is the only writer. The PATH wrapper is otherwise unchanged: it keeps stamping bare-shell runs and is never in the script's chain.
+- Reword the stale text:
+ - the =topgrade_freshness= docstring and its absent-stamp error in =maint/src/maint/probes/updates.py:115-127=, so neither says the lever stamps. The error reads 'no topgrade run recorded — upgrade-guarded stamps it once nothing is left deferred; a bare-shell topgrade stamps through the PATH wrapper';
+ - the doctor.py module docstring (L28-30) and guard.py's docstring (L5-7), so neither describes a refusal or =--force=;
+ - the =maint fix= usage line at =maint/README.md:23=, which drops =[--force]=;
+ - the header comment of =hyprland/.local/bin/topgrade= and the docstring of =tests/maint/test_topgrade_wrapper.py=, so they name bare-shell runs as the wrapper's job and no longer claim the lever or the health-check workflow.
+
+*** Guard UX
+- Drop =guard: "live_update"= from the update and topgrade remedies, so =iter_fix= no longer refuses them. The removal rests on the levers no longer arming a live swap: the script holds the GPU patterns whenever Hyprland is live, so on the driven path the guard fires only when an AUR upgrade reaches a held GPU/compositor package while Hyprland is live.
+- Deleted with the tag: =iter_fix='s refusal and its 'guard' event; the =--force= sentinel wrap for these two remedies; the GUI's =_update_force=; =maint fix --force=; =_rearm_after_guard=; the guard branch of =_on_fired=; the 'guard' handling at cli.py:114 and panel.py:368; and the doctor review suffix (doctor.py:415-421).
+- =panel.guarded(rid)= becomes =rid in {'update', 'topgrade'}=, with no remedy tag. I keep it rather than delete it because =_press_lever= lives in gui.py, which needs GTK, so this is the only GTK-free seam for the arm line.
+- =guard.trips=, which reads the TOML patterns Decision 3 calls the display-side mirror, survives only as the arm-line annotation. On a tripped read, the arm line for either key reads '<label> armed — will defer at least <guard matches> (kernels and DKMS modules are always held) — press again to run $ <argv>', where <argv> is =panel.arm_argv='s string ('upgrade-guarded --no-topgrade' for UPDATE, 'upgrade-guarded' for TOPGRADE), as on the untripped line. Nothing gates execution on it, and nothing says live apply or REBOOT required. maint computes no kernel set of its own; the post-run row gives the exact set.
-** Phase 3 — "Apply on reboot" and the boot-time unit (archsetup + maint)
-The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot. =archsetup-boot-upgrade.service=, installed by the installer: =ConditionPathExists= the flag, =Before=getty@tty1.service=, =TimeoutStartSec= bounded, non-fatal to every session target, =ExecStart= runs the script's =--complete= form as the user and removes the flag unconditionally. Installer step + unit file + the =/var/lib/archsetup/= flag directory. Tests in =tests/installer-steps/= for the install step; the arm action's tests in maint; a documented manual boot test (defer, arm, reboot, observe) in =todo.org= under Manual testing and validation.
+*** The deferred row
+- Probe =upgrade_deferred=, category updates (PACKAGES band). It reads the record through cache.get, and the arm flag's presence and entries, never a raw line count. It never writes the flag.
+- N is the record's packages minus the entries installed at or above =new= (=pacman -Q= plus =vercmp=), through a helper shared with =topgrade_freshness=, so a by-hand or console completion clears the row without rerunning the script. The helper never raises: for a missing record, or one whose data lacks the documented fields (packages a list of ={name, old, new, kind}=), it returns None, and =topgrade_freshness= then grades unchanged with evidence ={deferred: 0}=, so the =topgrade_age= card shows no APPLY and REVIEW & FIX keeps TOPGRADE.
+- No record (cache.get returns None): OK, value 0, '0 deferred', with no APPLY and no DISMISS.
+- A record whose data lacks the documented fields: UNPROBED 'deferred-set record malformed — rerun upgrade-guarded' (the =_cached= idiom), with no APPLY. It never crashes the envelope.
+- Severity, first match wins:
+ - CRIT while gate is non-null: 'do not reboot — <pkgbases> not verified: <gate.failed>; run upgrade-guarded --complete after any fix';
+ - WARN when result is failed, refused or interrupted (the text adds detail's last line, cut to 160 characters), or when K > 0 fixable advisories name deferred packages ('N deferred · K fixable advisories');
+ - OK otherwise: 'N deferred', or while the flag exists 'armed — A apply at next boot', where A is the flag entries (non-=#= lines) not installed at or above their version, with ' · <N−A> more deferred' appended when N > A.
+- The row never colours the waybar glyph. A kernel sits in the deferred set most days, so a plain deferral stays OK rather than moving the stale-metric symptom onto another row; =topgrade_age= already carries the cadence nag. Age-based escalation is the vNext follow-up in Scope tiers.
+- Evidence: name, old, new and kind rows; detail in full; held_snapshot when set; and the news lines not in =upgrade_news_dismissed=. The evidence says plainly that a plain reboot applies nothing until =--complete= has armed.
+
+*** APPLY
+- The REMEDIES entry =apply_deferred=: tier confirm, a new kind =terminal=, argv =['upgrade-guarded', '--complete']=, metric_ids =['upgrade_deferred', 'topgrade_age']=, always True, no timeout, not a strip key, and in no macro.
+- =panel.LEVER_KEYS= gets an =apply_deferred= items builder (the =_timer_items= idiom): =[]= on =upgrade_deferred= while N > 0 or gate is non-null, =[]= on =topgrade_age= while its evidence deferred > 0, and None otherwise.
+- =topgrade_freshness= adds evidence ={deferred: N}=, computed by the shared helper.
+- =doctor.review= offers =apply_deferred= only where the panel would show its key, and drops topgrade for a =topgrade_age= whose evidence deferred > 0. TOPGRADE stays a strip key, so the sweep stays one press away.
+- Every GUI press of a terminal-kind remedy arms on the first press. On the second, =_press_lever= makes one call, =panel.fire_press(rid, items, fire, fixture)=, where =fire= is a callable that runs gui.py's =_fire= for this press. For a terminal-kind rid, fire_press runs the panel.py launcher, =Popen(['foot', '--hold', '-e', *argv], start_new_session=True)=, or on a fixture board launches nothing and returns 'fixture board, not applied' as MERGE does; it never calls =fire=. For any other rid it calls =fire=. So a terminal-kind press never reaches =_fire= or =doctor.iter_fix=, and the choice lives in GTK-free code. That covers the row key, the =topgrade_age= card key and REVIEW & FIX's FIX. The detach means closing the panel can't kill or orphan the kernel transaction, and =--hold= keeps the gate verdict readable after the script exits.
+- The script does every step in that terminal: stage 1, the gate, the arm flag and the reboot prompt. maint never installs packages or writes the flag; the panel reads the outcome from the record and the flag on its next probe.
+- =maint fix apply_deferred= is the CLI form, and it is the usage line REVIEW & FIX prints. It and =iter_fix= run the argv with inherited stdin, stdout and stderr and no timeout, never through =cmd.run=, and return the script's exit code.
+
+*** DISMISS
+- A row key of a new kind, =dismiss=, which gui.py's =_digest_key= dispatches explicitly, with no fall-through to MERGE.
+- It is shown while any record news line is absent from =upgrade_news_dismissed=.
+- It is arm-then-fire, the MARK KNOWN idiom (gui.py:1687), because hiding a manual-intervention notice silences a signal.
+- On fire, a GTK-free panel.py helper overwrites the maint cache key =upgrade_news_dismissed= (cache.put, beside the record) with the record's whole news list. Overwriting with the whole list keeps earlier dismissals hidden when a new line arrives, and the record's 10-line cap keeps the key bounded.
+- The row hides news lines by whole-line equality.
+
+*** Freshness and REBOOT
+- =topgrade_age= grading is unchanged: per Decision 1 it ages from the last full-current stamp while anything is deferred.
+- After a TOPGRADE that left N > 0, the wall note is 'N deferred — freshness clears once they land (APPLY)' rather than a bare re-probed =topgrade_age= line.
+- REBOOT is hidden while gate is non-null, overriding the flag, =reboot_required= and =offer_reboot=. Otherwise it is shown while the flag exists, which covers ratio, where =reboot_required= stays false, and is otherwise unchanged. The panel never infers arm or gate state from an exit code.
+- The reboot remedy re-reads the record (=cache.get('upgrade_deferred')=) when it fires: while gate is non-null, =doctor.iter_fix= runs nothing and yields a fail event 'gate open — do not reboot: <gate.failed>; run upgrade-guarded --complete'. The check lives in iter_fix, so it covers the panel key before the next probe hides it, REVIEW & FIX and =maint fix reboot=. =doctor.review= also omits =reboot= while gate is non-null, as it omits =apply_deferred= where the panel shows no key, because =reboot_required= goes WARN whenever the running kernel's package has moved (its =/usr/lib/modules/$(uname -r)= is gone), which is the usual gate-open state.
+
+*** CVE, QUEUE and the strip
+- Advisories on deferred packages count against the deferred row, not UPDATE: =cve_queued= (the badge and the strip), the CVE card's 'fixable via UPDATE' caption, UPDATE's =cve_advisories= binding and REVIEW & FIX all exclude deferred names.
+- QUEUE (the results wall and =maint queue=) tags deferred names [HELD]. [HELD] is a post-run mark; the pre-run mark is the arm-line annotation built from the TOML patterns (Guard UX).
+- The strip's pending cell reads 'P pending · N held', where P is the existing pending count (held included) and N is the row's value.
+
+*** Alias and README
+- The =sysupgrade= alias in =common/.zshrc.d/aliases.sh= and =common/.bashrc.d/aliases.sh= becomes =if command -v upgrade-guarded >/dev/null 2>&1; then alias sysupgrade=upgrade-guarded; else alias sysupgrade=topgrade; fi=. common/ is also stowed on dwm installs, which never get the scripts. minimal/'s two =aliases.sh= files are symlinks into common/ (=tests/tier-dedup= pins that), so a none install gets the same conditional, which falls back to topgrade there; they stay symlinks.
+- Rewrite the 'The live-update guard' section of =maint/README.md=: drop press-again and =--force=; say UPDATE runs =upgrade-guarded --no-topgrade= and TOPGRADE runs =upgrade-guarded= (whose sweep is =/usr/bin/topgrade --no-ask-retry --disable system git_repos containers -y=); describe the deferred row, APPLY, DISMISS and the armed state. The guard banner and the wrapper paragraph stay as they are.
+
+*** Phase 2 tests
+In dotfiles =tests/maint/= under the maint fake harness. No test imports gui.py or GTK.
+- P2-levers: the two lever argvs, with kind user, timeout 3600, always True and strip keys; the missing-binary message; a zero-exit TOPGRADE lever whose record lists deferred packages leaves =topgrade_run= unchanged; and a fake long user-kind remedy whose argv traps INT and sleeps past a 1 s timeout receives SIGINT, its trap writes its marker, the step reports 'timed out — interrupted', and no process in its group survives.
+- P2-guard-ux:
+ - =panel.guarded('update')= and =panel.guarded('topgrade')= stay True, with no guard key on either remedy;
+ - =viewmodel.arm_line= with a tripped read gives the arm-line string with each label, ending '$ upgrade-guarded --no-topgrade' for UPDATE and '$ upgrade-guarded' for TOPGRADE;
+ - =review()= output carries no suffix, =iter_fix= emits no guard event, and =maint fix= has no =--force=;
+ - test_panel_levers.py:397-406 and test_panel_phase10.py:417-422 are rewritten to match, along with the tests that pinned the tag, the refusal and the wrap.
+- P2-row, from the fixture copy and variants of it:
+ - N drops an entry installed at or above its =new=, and dismissed news lines stay off the row;
+ - a gate-open fixture gives the CRIT text, REBOOT hidden and APPLY shown, and =doctor.iter_fix('reboot', ...)= on it yields a fail event naming gate.failed and never calls the priv reboot, and =doctor.review= on it, with =reboot_required= at WARN, lists no =reboot= entry;
+ - with the flag present, the armed text, and REBOOT shown whatever =reboot_required= says;
+ - an armed fixture plus a kind-kernel entry renders the appended count, and A drops below the flag's line count once an entry is installed;
+ - a record with failed_step aur grades WARN with detail's last line;
+ - a fixture with kernel-kind entries kept after a failed stage 1 and gate null shows N > 0 at WARN and offers APPLY;
+ - with no record file, the row is OK 0 with no APPLY, the =topgrade_age= card has no APPLY, and REVIEW & FIX keeps TOPGRADE for a stale =topgrade_age=;
+ - a record without packages reads UNPROBED, =topgrade_age= still grades from its stamp with evidence deferred 0 and no APPLY, and =build_status= still returns every other metric;
+ - a non-empty record leaves the waybar glyph colour unchanged;
+ - a fixture with held_snapshot set lists that name in the row's evidence, and one with held_snapshot null lists none.
+- P2-apply:
+ - the =topgrade_age= card shows APPLY at N > 0 and not at N = 0;
+ - REVIEW & FIX offers APPLY rather than TOPGRADE for a stale =topgrade_age= while N > 0, and TOPGRADE at N = 0;
+ - on a gate-open fixture, the row key and the roster's FIX both resolve to =apply_deferred= with kind terminal, and =panel.fire_press= given each one's rid and items calls the patched =subprocess.Popen= once with =['foot', '--hold', '-e', 'upgrade-guarded', '--complete']= and =start_new_session=True=, and never calls the =fire= stub or a patched =doctor.iter_fix=. =panel.fire_press('update', [], fire, None)= calls =fire= and never Popen. With a fixture name, the =apply_deferred= presses call neither Popen nor =fire= and return 'fixture board, not applied', while =panel.fire_press('update', [], fire, '<fixture>')= still calls =fire=, because =_fire= dry-runs there;
+ - =iter_fix= and =maint fix apply_deferred= never call =cmd.run=, and run the argv with inherited stdio.
+- P2-dismiss, with =cache.put= patched:
+ - the first press arms and writes nothing;
+ - the second writes exactly the record's news list, replacing any prior value;
+ - the re-probed row shows no lines and no DISMISS key;
+ - a line first recorded by a later run shows, and the earlier dismissed lines don't.
+- P2-cve-queue: given a fixable advisory on a deferred package, =cve_queued= is 0, the CVE caption omits it, neither UPDATE's =cve_advisories= binding nor REVIEW & FIX offers UPDATE for it, and the row carries it at WARN as 'N deferred · K fixable advisories'. A fixture record's names are tagged [HELD] in QUEUE's rows, and the strip reads 'P pending · N held'.
+- P2-text:
+ - after a zero-exit TOPGRADE that left N > 0, the wall note reads 'N deferred — freshness clears once they land (APPLY)';
+ - a source check that both common alias files carry the conditional and that =minimal/.zshrc.d/aliases.sh= and =minimal/.bashrc.d/aliases.sh= are still symlinks resolving to them;
+ - the =topgrade_age= absent-stamp error equals the new string, and the README's =maint fix= usage line carries no =--force=.
+
+** Phase 3 — The boot-time unit (archsetup)
+
+*** Installer step
+- An installer step writes the system unit =archsetup-boot-upgrade.service= from a heredoc, with =ARCHSETUP_USERNAME= replaced by sed (the BRIO rule precedent), runs =install -d -m 0755 /var/lib/archsetup=, and enables the unit with =systemctl enable=, without =--now=. The step documents itself in comments.
+
+*** The unit
+- [Unit]:
+ - =ConditionPathExists=/var/lib/archsetup/apply-upgrade-on-boot=, so a boot with no flag skips the unit at negligible cost;
+ - =Before=getty@tty1.service=. This holds the tty1 login (autologin where enabled, a password prompt otherwise) until the unit ends, and so holds off =~/.profile.d/99-hyprland-autostart.sh=. The ordering is load-bearing: a parallel run would let Hyprland start mid-swap and reintroduce the crash the guard prevents;
+ - default dependencies stay on (After=basic.target), so local and ZFS mounts are up, =/home= and =/var/cache/pacman= included;
+ - no Requires=, BindsTo= or RequiredBy=, and no network-online ordering.
+- [Service]:
+ - =Type=oneshot=;
+ - =User=ARCHSETUP_USERNAME=, substituted with the installing user, so =HOME= points at that user's maint state;
+ - =RuntimeDirectory=archsetup-boot-upgrade=, created when the unit starts and removed when it stops;
+ - =ExecStartPre=+/usr/bin/mv -f /var/lib/archsetup/apply-upgrade-on-boot /run/archsetup-boot-upgrade/armed.list=. The =+= runs it as root regardless of User=;
+ - =ExecStart=/usr/local/bin/upgrade-guarded --apply-armed=, with nothing after it;
+ - =TimeoutStartSec=20min=;
+ - =KillSignal=SIGINT=, with KillMode left at control-group;
+ - =TimeoutStopSec= left at the 90 s default, as the SIGKILL backstop;
+ - =StandardInput=null=, =StandardOutput=tty= and =TTYPath=/dev/tty1=, with stderr inheriting. Not journal+console, because ratio's =/dev/console= is ttyS0; the script mirrors its own output to the journal;
+ - no Restart=.
+- [Install]: =WantedBy=multi-user.target=. The unit is wanted, never required, so its failure can't fail the target.
+
+*** Disarm and failure
+- =ExecStartPre= consumes the flag before =ExecStart= runs, so success, failure, a timeout, SIGKILL and power loss all disarm, and a failed attempt never wedges later boots.
+- =ExecStart='s exit is the unit's result, so a failed or timed-out run leaves =archsetup-boot-upgrade.service= failed, and maint's existing =failed_units= row names it with its journalctl hint.
+- A timeout sends SIGINT to the cgroup. pacman, which has no SIGTERM handler, stops at a package boundary and releases =db.lck=, and the script records interrupted. The output mirror ignores SIGINT (Phase 1, Exit codes and interrupts), so pacman's output never loses its reader. The unit never deletes =db.lck=.
+- The boot form never refreshes and has no network dependency (Phase 1, =--apply-armed= sequence), which is why the unit has no network-online ordering.
+
+*** Phase 3 tests
+- P3-unit (=tests/installer-steps/=): a source-inspection test in the =test_pacman_hook_order.py= idiom. It asserts:
+ - =ConditionPathExists=/var/lib/archsetup/apply-upgrade-on-boot=;
+ - =Before=getty@tty1.service=;
+ - =Type=oneshot=;
+ - =User=ARCHSETUP_USERNAME= and its sed substitution;
+ - =RuntimeDirectory=archsetup-boot-upgrade=;
+ - the =ExecStartPre=+= mv line;
+ - =ExecStart=/usr/local/bin/upgrade-guarded --apply-armed=, with nothing after it;
+ - =TimeoutStartSec=20min=;
+ - =KillSignal=SIGINT=;
+ - =StandardInput=null=, =StandardOutput=tty= and =TTYPath=/dev/tty1=;
+ - =WantedBy=multi-user.target=;
+ - the absence of Requires=, BindsTo=, RequiredBy=, network-online.target and Restart=;
+ - the =install -d= and the enable call.
+- Manual boot tests. On a daily driver, run them when =upgrade-guarded --dry-run= lists a GPU-kind package; in the VM, follow the VM procedure below.
+ - M-boot-armed (armed boot, networking down): arm, confirm the armed state, disable networking, reboot. Expected: arming fired neither the guard nor informant's hook; the banner 'archsetup: applying N deferred GPU/compositor upgrades — do not power off' shows on tty1; the set installs from the cache before the tty1 login; the flag is gone; the record's packages are empty; and =topgrade_run= is freshly stamped when nothing else is deferred. In the VM, the arming =--complete= also has a kernel-set package pending, so it passes the real gate (=sudo -n lsinitcpio= included) on the ZFS root before it arms. On ratio and velox the test arms from APPLY (a detached =foot= terminal opens and the panel's wall streams nothing), confirms the row reads 'armed — A apply at next boot' with REBOOT shown before the reboot, and confirms =maint status= reads a fresh =topgrade_age= afterwards.
+ - M-boot-news (velox): an armed boot with two or more unread news items. Expected: the run completes with nothing waiting on stdin.
+ - M-boot-timeout: add a temporary drop-in with a short =TimeoutStartSec= and an =Environment=PATH= that puts the Phase 1 fake sudo and fake pacman, the pacman set to sleep, ahead of =/usr/bin=; then reboot armed. Expected: the tty1 login appears once the timeout fires, the flag is gone, the next boot skips the unit, the record says interrupted with the remedy and still lists the set, and =systemctl is-failed archsetup-boot-upgrade.service= reports failed (on a daily driver, the deferred row lists the set at WARN and maint's failed-units row names the unit). Remove the drop-in afterwards.
+ - M-boot-midtx: a real armed set with a =TimeoutStartSec= short enough to land inside the transaction. Expected: no =/var/lib/pacman/db.lck= remains, the record says interrupted, and =upgrade-guarded --complete= from a console then finishes the set.
+ - M-boot-fail: the M-boot-timeout drop-in with the fake pacman answering =-Q= and =-Sp= normally and exiting 1 on =-S=. Expected: the login appears, the flag is gone, the record names the failure (failed_step boot-transaction) and still lists the set, and the unit is failed. Remove the drop-in afterwards.
+ - M-boot-ratio: an armed boot on ratio, whose =/dev/console= is ttyS0. Expected: the banner and pacman's progress stay visible on the monitor through the transaction, and REBOOT was offered while armed even though =reboot_required= stays false.
+- VM procedure:
+ - =make test-keep FS_PROFILE=zfs= runs QEMU headless with no compositor, so Hyprland is never live there;
+ - create pending packages from a console by installing an older version of a GPU-pattern package (and, for M-boot-armed and M-split-console, of a kernel-set package) from the cache or the Arch Linux Archive;
+ - arm over SSH with =UPGRADE_GUARDED_HYPR_RUNNING=1 upgrade-guarded --complete= instead of APPLY, and confirm the armed state from the flag and the record;
+ - observe tty1 with the QEMU monitor's screendump on =vm-images/qemu-monitor-zfs.sock=; a guest reboot keeps the QEMU process and SSH;
+ - confirm the banner and the before-login ordering from =journalctl -u archsetup-boot-upgrade= and getty@tty1's start timestamp, and a failed run with =systemctl is-failed archsetup-boot-upgrade.service=;
+ - never use =debug-vm.sh= here: it restores the clean-install snapshot and would erase the install under test.
+- VM gate: =make test-keep FS_PROFILE=zfs= passes M-boot-armed, M-boot-timeout, M-boot-midtx and M-boot-fail, plus M-split-console, before the unit reaches velox. M-boot-news and M-boot-ratio run on their machines after step (c).
** Phase 4 — Docs, rollout, and both daily drivers
-Document the flow (arm → reboot → console upgrade → session). Roll the unit to velox and ratio (installer already covers a rebuild; existing machines need the one-time install). Confirm the ratio path matches.
+
+*** README flow and recovery
+One dotfiles commit to the 'The live-update guard' section of =maint/README.md=:
+- The flow: UPDATE defers → APPLY / =--complete=: kernel and DKMS sets live → gate → pre-download and arm → reboot → tty1 console upgrade → login.
+- For manual diagnosis of a CRIT row, =kernel-modules-check <pkgbase>= with the pkgbase the row shows. Only =upgrade-guarded --complete= clears the gate.
+- The recovery steps below, which are the only copy.
+
+*** Recovery
+Boot the held snapshot from ZFSBootMenu: the record's held_snapshot, which the deferred row's evidence shows. When the root won't boot, find it from the ZFSBootMenu recovery shell as the =zroot/ROOT/default@pre-pacman_*= row with userrefs above 0 in =zfs list -H -t snapshot -o name,userrefs zroot/ROOT/default=, and confirm the =upgrade-guarded= tag with =zfs holds <snap>=. Pick by the hold, never by age: a pre-pacman snapshot newer than the held one can already hold the new kernel. Then, before the next =upgrade-guarded= run:
+- The pre-pacman snapshot covers =zroot/ROOT/default= only. The pacman db (=zroot/var/lib/pacman=), the cache (=zroot/var/cache=) and the log (=zroot/var/log=) are separate datasets and stay current.
+- If a power cut left =/var/lib/pacman/db.lck=, remove it by hand; nothing is running after a reboot.
+- Reconcile the pacman db to the booted root from =/var/log/pacman.log=. Take every =[ALPM]= upgraded, downgraded, installed and removed line timestamped after the snapshot's creation, and undo them newest first:
+ - upgraded or downgraded: =sudo pacman -U --dbonly= the old version from =/var/cache/pacman/pkg= (yay's cache for an AUR package);
+ - installed: =sudo pacman -Rdd --dbonly=;
+ - removed: =sudo pacman -U --dbonly= the removed version.
+- A power cut mid-package leaves that package with no =[ALPM]= line, and can leave a local db entry that =pacman -Dk= reports as 'description file is missing'. pacman removes the old entry before it extracts the new one, so the db no longer records the old version. Delete that entry's directory under =/var/lib/pacman/local=, then =sudo pacman -U --dbonly= from the cache the version the booted root holds: the one named by the last =[ALPM]= line for that package timestamped before the snapshot's creation.
+- Then =pacman -Dk= is clean, and =pacman -Qkk= over those names and over any package the previous step re-registered reports no mismatch.
+- Never roll back =zroot/var/lib/pacman=, which the snapshot hook never snapshots.
+
+*** Rollout
+A rebuilt Hyprland machine gets every piece from the installer. Existing machines don't: the guard install runs inside =window_manager=, which is already marked complete in =/var/lib/archsetup/state/=, so re-running archsetup installs nothing. On each existing machine, ratio and velox, by hand, in this order:
+- (a) After Phase 1's last commit:
+ - pull archsetup;
+ - =sudo install -m 755 scripts/upgrade-guarded scripts/kernel-modules-check /usr/local/bin/=, and on a ZFS root also =sudo install -m 755 scripts/zfs-pre-snapshot /usr/local/bin/=;
+ - write =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook= from the installer heredoc and confirm its Targets match the heredoc's;
+ - remove the unprefixed =hypr-live-update-guard.hook= and confirm with =ls /etc/pacman.d/hooks/=. Ratio still has the legacy name, and velox's was hand-placed under it; the script refuses without the =10-= name;
+ - re-copy =configs/maintenance-thresholds.toml= to =~/.config/archsetup/= after diffing for local edits;
+ - confirm =pacman -Qmq= lists no DKMS-set package, because a DKMS package from the AUR would sit outside the hold (=yay -Sua= upgrades it ungated);
+ - confirm =upgrade-guarded --dry-run= exits 0;
+ - on a ZFS root, confirm =kernel-modules-check "$(cat /usr/lib/modules/$(uname -r)/pkgbase)"= exits 0. This structural run proves the =sudo -n lsinitcpio= read against the real image;
+ - run M-split-live.
+- M-split-live (manual): from a terminal in the live session, right after =--dry-run=, run =upgrade-guarded --no-topgrade=. Expected: exit 0, no BLOCKED banner, the record's packages equal dry-run's predicted deferred set, and =pacman -Qu= then lists only those names plus =[ignored]= rows. When dry-run lists a GPU-kind entry, this run is the hook-silence check; if none is pending, the hook-silence part runs on the first day one is, before AC1 is ticked.
+- (b) Only after (a) on that machine: pull the Phase 2 commits. No Phase 2 commit reaches a machine's dotfiles before (a) is done there, velox included.
+- (c) After Phase 3, and on velox only after the VM gate passes:
+ - write =archsetup-boot-upgrade.service= from the installer heredoc, with the username substituted;
+ - =sudo install -d -m 0755 /var/lib/archsetup=;
+ - =sudo systemctl daemon-reload=;
+ - =sudo systemctl enable archsetup-boot-upgrade.service=.
+- Until step (c), =--complete= reports the unit not enabled and arms nothing, so the GPU/compositor set lands from a console instead.
+- Confirm the ratio path: tty1 shows a password prompt rather than autologin, and the =Before=getty@tty1.service= ordering holds that prompt; the banner shows on its monitor rather than on ttyS0 (M-boot-ratio); and of its three kernels, =linux-lts-strix= is never moved or gated.
+
+*** Recovery drill (D-zbm)
+A non-gating drill in =todo.org=, run in the =make test-keep FS_PROFILE=zfs= VM; no criterion waits on it.
+- With a kernel update and one non-kernel package pending, make the =zfs-dkms= build fail for the new kernel.
+- Run =upgrade-guarded --complete=, and confirm it exits 4 naming the failure, arms nothing, and records the held snapshot in held_snapshot.
+- Boot that snapshot from ZFSBootMenu, then run the Recovery reconcile.
+- Expected: =pacman -Q= shows the old kernel set and the non-kernel package's old version, and =pacman -Qkk= over the reconciled names reports no mismatch.
+- Variant, on a VM whose record has gate null (a fresh =make test-keep FS_PROFILE=zfs= run, not the VM the drill above left with its gate open), so that =--complete= step 3 takes and holds its own snapshot: with a kernel update pending, run =upgrade-guarded --complete= over SSH. While stage 1 prints its '(n/m) upgrading' lines, hard-reset the guest with the QEMU monitor's =system_reset=. That reset is the power-cut equivalent and keeps the QEMU process. Never use =system_powerdown=, a graceful ACPI shutdown that runs the trap instead, and never use =quit=: both end QEMU, and =make test-keep= then restores the clean-install snapshot. Boot held_snapshot from ZFSBootMenu and run Recovery. Expected: =zfs holds= shows the =upgrade-guarded= tag on the snapshot step 3 took before stage 1, and held_snapshot names it; =pacman -Q= shows the old kernel set; =pacman -Dk= is clean; and =pacman -Qkk= over the reconciled names, including the package cut off mid-extraction, reports no mismatch.
+
+*** Rollback
+Roll back in this order (dotfiles and docs first, so nothing calls the scripts once they are removed; steps 1 and 2 are independent):
+1. Revert the dotfiles Phase 4 README commit, then the Phase 2 commits newest first, then the archsetup =system-health-check.org= doc commit. The README goes first because Phase 2 and Phase 4 edit the same section. Reverting Phase 2 restores the old argv, the force path and the alias, and removes the deferred row and APPLY.
+2. =sudo systemctl disable archsetup-boot-upgrade.service=, remove the unit, =sudo systemctl daemon-reload=, and remove the flag.
+3. Release held_snapshot if set (=sudo zfs release upgrade-guarded <name>=), then remove both scripts from =/usr/local/bin=.
+
+Removing the scripts before step 1 leaves UPDATE, TOPGRADE, the =sysupgrade= alias and the health-check workflow's update step pointing at a missing binary. Removing only the unit leaves =--complete= unable to arm, and the next everyday run removes any flag (Phase 1, everyday step 9), so nothing is left dangling. The rewritten =guard_patterns=, the =10-= hook name, the userrefs-aware prune, post-rebuild-check's check 9 and the record file can stay. The guard hook's behavior is untouched throughout.
* Acceptance criteria
-- [ ] With a guarded library pending and Hyprland live, UPDATE applies everything else, exits 0, and the panel shows the exact deferred set; the guard hook does not fire.
-- [ ] The same run with no compositor live (a TTY) applies everything but the kernel set and stamps only if nothing was deferred.
-- [ ] A =--complete= run whose DKMS build fails stops before arming or rebooting, names the failure, and leaves the machine running on the old kernel; on velox the pre-pacman snapshot it required is bootable from ZFSBootMenu.
-- [ ] With a guarded library pending, arming and rebooting applies it in the console before Hyprland starts, and =maint status= then reads a fresh =topgrade_age=.
-- [ ] A boot-upgrade failure (a failed step, a timeout, an aborted transaction) never blocks the session: the machine boots into Hyprland, the flag is cleared, and the panel still shows the pending work.
-- [ ] Unread Arch news does not wedge the boot run (=informant read= precedes the transaction).
-- [ ] A guarded upgrade completed from a TTY via the Phase-1 path stamps freshness identically to the boot unit.
-- [ ] The =hypr-live-update-guard= hook is unchanged and still blocks a live guarded swap.
+Each criterion is an observable outcome. Implementation phases hold the exact behavior behind it, and Testing names the tests that check it. "Gate open" means the record's =gate= is non-null, and a "gate check" is a =kernel-modules-check= run by =--complete=.
+- [ ] AC1: With a GPU/compositor package pending and Hyprland live, UPDATE (=upgrade-guarded --no-topgrade=) and TOPGRADE (=upgrade-guarded=) each:
+ - apply everything outside the held set and, when every step succeeds, exit 0, because a deferral alone never makes the exit non-zero;
+ - leave the =hypr-live-update-guard= hook silent during the pacman step;
+ - leave =pacman -Qu= listing only the deferred set a =--dry-run= run just before predicted, plus =[ignored]= rows, with the panel's =upgrade_deferred= row showing that set and QUEUE tagging its names =[HELD]=;
+ - never write =topgrade_run= while anything is deferred, whether the run starts from the panel or a shell.
+- [ ] AC2: On a soname day, when a pending package needs a held package's new version, the dependency closure holds that package too, and the run exits 0 having applied everything else.
+ - An upstream =zfs-utils= bump that breaks =zfs-dkms='s exact =zfs-utils= pin is held together with =zfs-dkms=. A pkgrel-only =zfs-utils= bump is not held.
+ - When the chain runs through a dependency the transaction would newly install, the pending package that pulls it in is held and the run doesn't refuse.
+ - When the requirer is installed but not pending (an orphan, foreign or AUR package, like ratio's =qemu-block-gluster= on 2026-08-25), the run refuses before any transaction (exit 3, the record's =failed_step= is =closure=), names the refused packages in its last stderr line, and never forces or removes a package.
+- [ ] AC3: The same UPDATE or TOPGRADE run with Hyprland not running applies everything but the kernel and DKMS sets and their closure, so only entries of kind =kernel= remain in the record and in =pacman -Qu= (besides =[ignored]= rows).
+ - "Not running" means =pgrep -x Hyprland= fails. A console reached with Ctrl+Alt+F2 while Hyprland runs on tty1 is live, so a run there holds the GPU patterns as in AC1.
+ - It writes =topgrade_run= only when it is a TOPGRADE run (no =--no-topgrade=) whose refresh, pacman, AUR and sweep steps all exited 0, nothing is left deferred and no gate is open. UPDATE never stamps.
+- [ ] AC4: A =--complete= with a kernel-set or DKMS-set member pending opens the gate from the start of its stage 1 until a gate check passes, whether the run started from the panel or a shell.
+ - While the gate is open, the deferred row is CRIT ('do not reboot — <pkgbases> not verified: <gate.failed>; run upgrade-guarded --complete after any fix'), no arm flag exists, and no =--complete= prompts to reboot. REBOOT is hidden from the panel's first probe after the gate opens, whatever the arm flag, =reboot_required= or =offer_reboot= say. A reboot fired before that probe, from the panel key, REVIEW & FIX or =maint fix reboot=, runs nothing and fails with 'gate open — do not reboot: <gate.failed>; run upgrade-guarded --complete'.
+ - The gate stays open while stage 1 runs, after an interrupted stage 1 (exit 130), and across a wall dismiss and a panel reopen. It closes only when a gate check passes or every recorded pkgbase has been uninstalled.
+ - A DKMS build that fails for the new kernel makes the run exit 4 with the failing item as its last stderr line and in the CRIT text, and the record reads result =failed= with =failed_step= =gate=. The machine keeps running the old kernel, with the new one installed but not booted.
+ - While anything is deferred or the gate is open, the row offers APPLY. APPLY, from the row or from REVIEW & FIX's FIX, opens =upgrade-guarded --complete= in a detached =foot= terminal and never runs it inside the panel.
+- [ ] AC5: With a GPU/compositor package pending, Hyprland live and the boot unit enabled, APPLY arms the boot upgrade on ratio and velox. The deferred row then reads 'armed — A apply at next boot', where A counts the armed entries not yet installed, with ' · <N−A> more deferred' appended while other entries stay deferred, and REBOOT shows.
+ - A reboot with networking down applies the armed set from the pacman cache on tty1 before the tty1 login, with the banner 'archsetup: applying N deferred GPU/compositor upgrades — do not power off' visible.
+ - Afterwards the flag is gone, the deferred row drops the entries that landed, and =maint status= reads a fresh =topgrade_age= when nothing else is deferred.
+- [ ] AC6: A boot-upgrade failure (a failed step, a timeout or an aborted transaction) never blocks the session:
+ - the tty1 login appears and the flag is gone, even when the timeout killed =ExecStart=;
+ - no =/var/lib/pacman/db.lck= remains after a timeout that SIGINT ends (the SIGKILL backstop and power loss are the residual Risks names);
+ - the record's result and =failed_step= name what went wrong, its detail names the reason (and, for an interruption, the remedy), and the deferred row still lists what did not land, at WARN;
+ - =archsetup-boot-upgrade.service= is left failed and shows in maint's failed-units row.
+- [ ] AC7: With two or more unread Arch news items and no stdin, unread news wedges no mode once the informant clear succeeds. A failed clear ends in a named =failed_step=, the lever runner's or the unit's timeout, or, in a terminal run (APPLY's =--complete= or a shell), a wait the person can end with Ctrl+C (exit 130, recorded as interrupted); none of these swaps anything.
+ - Where informant is installed, =timeout 60 sudo -n informant read --all= runs before the first transaction in all three transacting modes (everyday, =--complete= and =--apply-armed=).
+ - The everyday and =--complete= runs record the unread items (=yay -Pwq=) before the clear, and the deferred row lists them with a DISMISS key until they are dismissed.
+- [ ] AC8: A =--complete= from a console with Hyprland stopped lands the deferred GPU/compositor set and stamps freshness. It writes =topgrade_run= exactly when it got past the nothing-to-complete check, started with something deferred or the gate open, every step it ran succeeded, and it finishes with nothing deferred and no gate open.
+ - An arming run never stamps, and neither does a run that prints 'nothing to complete', even when the record it started from listed packages that have since been installed.
+ - A failed stamp makes the exit 1 and leaves the record at result =failed= with =failed_step= =stamp=.
+- [ ] AC9: While anything is deferred or the gate is open, the deferred row offers APPLY, and the deferral never changes the waybar glyph colour.
+ - After a successful TOPGRADE that left N > 0 entries deferred, =topgrade_age= keeps aging, its card also offers APPLY, REVIEW & FIX offers APPLY rather than TOPGRADE for it when stale, and the wall note reads 'N deferred — freshness clears once they land (APPLY)'.
+ - Once nothing is deferred, the =topgrade_age= card drops APPLY and REVIEW & FIX offers TOPGRADE for it again.
+ - With no record file at all, the row reads OK '0 deferred' with no APPLY or DISMISS, and REVIEW & FIX keeps TOPGRADE for a stale =topgrade_age=.
+- [ ] AC10: An AUR upgrade that needs a newer held GPU/compositor package under a live Hyprland fails the AUR step visibly without blocking the topgrade sweep. The guard aborts yay's own pacman call, the record's =failed_step= is =aur=, the row is WARN, the run exits 1, and nothing stamps.
+- [ ] AC11: A fixable advisory on a deferred package shows on the deferred row at WARN ('N deferred · K fixable advisories'), not as UPDATE-fixable: =cve_queued=, the CVE card's 'fixable via UPDATE' caption, UPDATE's =cve_advisories= binding and REVIEW & FIX all exclude it.
+- [ ] AC12: The =hypr-live-update-guard= hook is unchanged and still blocks a live GPU/compositor swap.
+ - Every mode but =--apply-armed= refuses before any transaction (exit 3, the path named; outside =--dry-run= the record's =failed_step= is =precondition=) when =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook= is unreadable or has no =Target= lines.
+ - The installer writes the hook as =10-hypr-live-update-guard.hook=, and the seed TOML's =guard_patterns= equal its =Target= list.
+- [ ] AC13: On a ZFS root, a =--complete= that opens the gate entry snapshots the root dataset and puts an =upgrade-guarded= hold on that snapshot before its stage 1, and the record's =held_snapshot= names it. An everyday run whose AUR step opens the entry holds, the same way, the oldest =pre-pacman_*= snapshot taken no earlier than a minute before that step. When no such hold exists (=held_snapshot= is null, gone, or taken more than a minute before the entry opened), a gate check in the =--since= form, pass or fail, holds the oldest =pre-pacman_*= snapshot taken no earlier than a minute before the entry opened, and =held_snapshot= names that one. =zfs-pre-snapshot='s prune never destroys the held snapshot or counts it toward its keep limit, and it stays held until a later =--complete=, or an everyday run whose AUR step opens the entry, holds a newer one and releases it.
+
+I treat booting the held snapshot from ZFSBootMenu as the fallback for a kernel that won't boot, not as a gating criterion. AC13 keeps that snapshot from being pruned, Phase 4 holds the recovery steps and the non-gating D-zbm drill, and Risks weighs the cost.
* Readiness dimensions
Answer each, or write "N/A because…".
-- Data model & ownership: the arm flag (=/var/lib/archsetup/=, installer-owned) and the =topgrade_run= cache key (maint-owned). No user-authored data.
-- Errors, empty states & failure: the boot unit is best-effort and self-disarming; every failure path lands in "boot normally, metric stays stale, re-arm to retry." Named, non-silent.
-- Security & privacy: relies on the existing =%cjennings NOPASSWD: ALL=; the unit runs the upgrade as the user via sudo, adds no new privilege. Note the NOPASSWD breadth as a pre-existing fact, not introduced here.
-- Observability: the boot run's output is on the console; its systemd unit status and journal record success/failure; the panel reflects the cleared or still-pending state after boot.
-- Performance & scale: one pacman/yay transaction at boot; bounded by =TimeoutStartSec=. Negligible boot-time cost when the flag is absent (=ConditionPathExists= skips the unit).
-- Reuse & lost opportunities: reuses the guard's trigger list by reading the installed hook (single source of truth for "which libs are dangerous"), =informant=, =maint stamp=, =checkupdates=, and topgrade's own step switches. The one duplicate that exists today — maint's TOML =guard_patterns= — is kept as a display mirror and pinned to the hook by a test rather than removed.
-- Architecture fit & weak points: integration points are the pacman hook set, getty autologin ordering, and the maint cache. Weak point: the =Before=getty@tty1= ordering is load-bearing for safety; a parallel run reintroduces the live-swap crash. Mitigated by making the ordering explicit and tested-by-inspection.
-- Config surface: the arm flag path and the timeout. Defaults safe (absent flag = no-op).
-- Documentation plan: a short "reboot to apply guarded upgrades" note in the maint docs; the installer step self-documents in-comment.
-- Dev tooling: installer-step pytest for Phase 2; maint unit tests for Phases 1 and 3; a manual boot test in =todo.org=.
-- Rollout, compatibility & rollback: additive; removing the unit and flag reverts fully. Existing machines need a one-time install; a rebuild gets it from the installer. Rollback leaves the guard and manual TTY path intact.
-- External APIs & deps: topgrade =--only system=, =informant read=, =yay=, =maint stamp= — all verified present on velox this session. No external service.
+- Data model & ownership: five pieces of state, none of it user-authored, plus a lock file. Each has one writer on the driven path (=topgrade_run= also has the PATH wrapper for bare-shell runs), and its exact shape lives in the phase subsection cited.
+ - The record, =${MAINT_STATE_DIR:-$HOME/.local/state/maint}/upgrade_deferred.json= (maint cache key =upgrade_deferred=). =upgrade-guarded= is its only writer; maint's =upgrade_deferred= probe, =topgrade_freshness= and the script itself read it. It carries no mode, armed or boot state: armed means the flag exists, and a gate is open while =gate= is non-null. A fixture in each repo pins its shape, and a field change updates both in one rollout. Exact fields, writes and fixtures: Phase 1, The record.
+ - The arm flag, =/var/lib/archsetup/apply-upgrade-on-boot=, root-owned. =upgrade-guarded= writes and removes it through sudo, the boot unit's =ExecStartPre= consumes it, and the panel reads only its presence and entries. The =/var/lib/archsetup/= directory is installer-owned. Exact rules: Phase 1, The arm flag.
+ - =topgrade_run= (maint cache key), written only through =maint stamp topgrade=: by =upgrade-guarded= under its stamp predicate, and by the =~/.local/bin/topgrade= PATH wrapper for bare-shell runs. doctor.py's lever stamp is deleted. Exact predicate: Phase 1, Stamp predicate.
+ - =upgrade_news_dismissed= (maint cache key, beside the record), written only by the deferred row's DISMISS key, which overwrites it with the record's whole news list, so it stays bounded by the record's cap of 10 lines. Exact behavior: Phase 2, DISMISS.
+ - On a ZFS root, one =upgrade-guarded= user hold on a pre-pacman snapshot, named by the record's =held_snapshot=. Only =--complete=, and an everyday run whose AUR step opens the gate entry, place or move it, the =zfs-pre-snapshot= prune never destroys a held snapshot, and rollback releases it. Exact behavior: Phase 1, The snapshot hold.
+ - =upgrade-guarded.lock=, beside the record: a flock file that holds no data, held for the life of every invocation except =--dry-run=. Exact rule: Phase 1, Preconditions.
+- Errors, empty states & failure: every failure is named and non-silent, and leaves a state the next run can recover from.
+ - On any non-zero exit the last stderr line names the reason, and the record's =result=, =failed_step= and =detail= carry it to the panel. A deferral alone never makes the exit non-zero. Exit codes and the interrupt trap: Phase 1, Exit codes and interrupts.
+ - Refusals exit 3 before any transaction and leave nothing upgraded. The script never deletes =db.lck= and never proceeds on an empty hook pattern list. Exact refusal rules: Phase 1, Preconditions and Dependency closure.
+ - News capture and the informant clear never stop a run; Risks (Arch news) names the residual when the clear fails. A failed AUR step still runs the sweep, and nothing stamps. A stamp that can't be written exits 1 with =failed_step= stamp and leaves the rest of the record as computed (Phase 1, Finish).
+ - A failed gate leaves the new kernel installed, nothing armed and no reboot offered. The terminal run names the failing item, the record keeps the gate open, and the panel shows a CRIT row with REBOOT hidden until a later =--complete= closes the gate. Exact rules: Phase 1, Invariants and The gate; Phase 2, The deferred row.
+ - Arming that can't resolve or download arms nothing. While armed, an everyday run whose refresh succeeds rewrites or removes the flag, and removes it outright when the boot unit isn't enabled, so a flag never outlives its unit (Phase 1, Everyday sequence, step 9).
+ - Empty states: an empty pending set is a normal run that writes an empty record. A =--complete= with nothing pending and no open gate prints 'nothing to complete' and exits 0. A boot run whose armed entries are all installed already is a no-op. With no record the deferred row reads OK '0 deferred', and a malformed record reads UNPROBED (Phase 2, The deferred row).
+ - The boot unit is bounded and self-disarming. =ExecStartPre= moves the flag off its persistent path before the attempt, so success, failure, timeout, SIGKILL and power loss all disarm. The 20-minute start timeout sends SIGINT, which stops pacman at a package boundary so it releases =db.lck= itself. A failure, timeout or SIGKILL ends at the tty1 login in the same boot, with the record naming the failure and the unit failed. After power loss the next boot skips the unit on its condition and no failed state survives, so the pre-written interrupted record (the deferred row, or =maint status= from a console) and, if a =db.lck= was left, the refusal that names it are what remain. Retrying means running =upgrade-guarded --complete=. Exact unit: Phase 3, The unit.
+- Security & privacy: =upgrade-guarded= runs as the invoking user and refuses EUID 0, so the record and the stamp live in that user's =$HOME= (yay refuses root anyway). Every root step the script itself runs is =sudo -n= on the existing =%cjennings NOPASSWD: ALL= rule, so a missing rule fails with a reason instead of prompting (at the refresh in the everyday run and =--complete=, before yay or the sweep runs; at the transaction in the boot form), and the unit's =ExecStartPre=+= move is the only root step outside sudo. The gate stays unprivileged apart from one =sudo -n lsinitcpio= read, because mkinitcpio writes the initramfs images root 0600. Phase 1, Identity and privilege, is the only list of the script's own root steps; yay and the sweep escalate through their own sudo. The flag is root-owned, so arming a boot takes sudo. I add no new privilege; the NOPASSWD breadth is a pre-existing fact, not introduced here. The record holds package names and versions, news lines, failure detail and a snapshot name, and nothing leaves the machine.
+- Observability:
+ - The boot run writes to tty1, not journal+console, because ratio's =/dev/console= is ttyS0. It opens with a do-not-power-off banner naming the count, and mirrors all output to the journal under the =archsetup-boot-upgrade= tag (Phase 1, --apply-armed sequence; Phase 3, The unit).
+ - =ExecStart= is the script with nothing after it, so its exit is the unit's result. A failed or timed-out run leaves =archsetup-boot-upgrade.service= failed, and maint's existing =failed_units= row names it with its journalctl hint in the session that follows.
+ - Every mode except =--dry-run= writes the record, so after any run (panel, shell or boot unit) the deferred row shows the deferred set, the last result and detail, any open gate, the armed state read from the flag, the held snapshot's name when one is set, and the undismissed news (Phase 2, The deferred row).
+ - =--dry-run= prints the sets, the GPU closure and every argv, and writes nothing (Phase 1, Modes and options).
+- Performance & scale: one pacman transaction at boot, installed from the cache that arming pre-downloaded, so there's no network wait. It's bounded by the 20-minute start timeout, and a boot with no flag skips the unit at negligible cost. An everyday run is one refresh, read-only and download-free closure probes (at most |pending|+1 per stage), one =-Su=, then yay and the sweep; the informant clear is bounded by =timeout 60=. The =--complete= session is long by nature (a kernel transaction plus DKMS builds), so APPLY opens it in a detached =foot= terminal rather than under the lever runner's 3600 s timeout.
+- Reuse & lost opportunities: reuses the guard's own =Target= list by reading the installed hook (the single source of truth for which libraries are dangerous) and the guard's liveness test (=pgrep -x Hyprland=); =maint stamp topgrade= and maint's cache.put envelope; maint's =failed_units= row for boot failures; =informant read --all= where installed, with =yay -Pwq= capturing the news first so clearing the hook never discards it unread; =checkupdates= for =--dry-run=; topgrade's own step switches; the MERGE precedent for APPLY's detached terminal; the MARK KNOWN arm-then-fire idiom for DISMISS; the BRIO rule's sed templating for the unit's =User==; =zfs-pre-snapshot='s pre-pacman snapshots, which the gate requires and the hold pins; and =tests/zfs-pre-snapshot/fake-zfs= in the tests. maint's TOML =guard_patterns= is today a divergent list, not a copy of the hook. Phase 1 rewrites it to exactly the hook's 15 =Target= patterns and keeps it as the panel's display-side mirror, pinned by a =tests/installer-steps/= test that reads neither =/etc= nor =~/.config=.
+- Architecture fit & weak points: integration points are the guard hook file (read for its =Target= lines), the pacman sync db and package cache, sudo, getty@tty1 login ordering, the pre-pacman snapshots on a ZFS root, and the maint cache and panel. The first weak point is the =Before=getty@tty1.service= ordering, which is load-bearing for safety. It holds the tty1 login, and with it =~/.profile.d/99-hyprland-autostart.sh=, until the unit ends; a parallel run would reintroduce the live-swap crash. I made the ordering explicit and pinned it with the unit's source-inspection test (P3-unit). The second is that the panel (dotfiles) calls a binary that existing machines get by a one-time hand install from archsetup, so the per-machine order in Phase 4, Rollout, is load-bearing too.
+- Config surface: nothing to tune. The GPU patterns come from the hook, the flag path is fixed, and the boot timeout is fixed at 20 minutes with SIGINT, keeping the 90 s default stop timeout as the SIGKILL backstop (Phase 3, The unit). =MAINT_STATE_DIR= is honoured as maint honours it; the =UPGRADE_GUARDED_*= and =KMC_*= variables are test seams, not configuration. The TOML =guard_patterns= is display-only. Defaults are safe: with no flag, the unit is skipped.
+- Documentation plan:
+ - Phase 1 rewrites the =guard_patterns= comment in =configs/maintenance-thresholds.toml= to 'display mirror of the hook's Target list, pinned by a test', and lands the =docs/workflows/system-health-check.org= edit as a doc commit right after its code. The exact edit is in Phase 1.
+ - Phase 2 rewrites the 'The live-update guard' section of dotfiles =maint/README.md= for the new levers, the deferred row, APPLY and the armed state, dropping press-again and =--force=. Its stale-text bullet lists the docstrings, usage line and wrapper header it rewords.
+ - Phase 4 adds the =--complete= flow and the recovery steps to that README section. Phase 4, Recovery, is the only copy of the recovery steps.
+ - The installer steps self-document in comments.
+- Dev tooling: unittest only (Testing). Phase 1, Seams and fakes, is the only list of seams and fakes; each phase names its test groups, and Testing maps them to the criteria and states the VM gate.
+- Rollout, compatibility & rollback: not additive, because Phase 2 repoints UPDATE, TOPGRADE and the =sysupgrade= alias and removes the force path. Existing machines need the one-time steps, in order, in Phase 4, Rollout; Phase 4, Rollback, reverts dotfiles and docs first and lists what can stay. The guard hook's behavior is untouched throughout.
+- External APIs & deps:
+ - pacman (=-Sy=, =-Su=, =-S= with versioned targets, =-Sw=, the =-Sup= and =-Sp= print modes, =-Q=, =-Qu=, =-Qqo=), =vercmp= and =sudo=;
+ - =yay -Sua= and =-Pwq=;
+ - topgrade at =/usr/bin/topgrade=, whose disabled steps go as separate argv elements because topgrade 17.12.2 rejects the comma form with exit 2 (exact argv: Phase 1, Everyday sequence, step 8);
+ - =informant read --all= where installed (velox; absent on ratio), under =timeout=;
+ - =checkupdates= (pacman-contrib; =--dry-run= only);
+ - =dkms status=, plus =findmnt=, =lsinitcpio= and =zfs list=, =snapshot=, =hold= and =release= on a ZFS root;
+ - =flock=, =pgrep= and =timeout=;
+ - =systemctl is-enabled=, =systemctl reboot= and =systemd-cat=;
+ - =foot= (APPLY's terminal);
+ - =$HOME/.local/bin/maint stamp topgrade=. maint is on the user's login PATH only (=~/.local/bin=, from dotfiles), not on a system unit's PATH, so the script calls it by absolute path;
+ - no external service beyond the Arch mirrors and news feed that pacman, yay and informant already use.
* Risks, Rabbit Holes, and Drawbacks
-- Boot critical path: the unit sits ahead of autologin, so a hang would delay boot. Mitigated by =TimeoutStartSec= and non-fatal wiring; worst case is a bounded delay, then a normal session.
-- Interactive prompts under no stdin: =yay=/pacman can still prompt (provider choice, replace, AUR review) even with =assume_yes=. The =--only system= scope and =--noconfirm=-style flags shrink this to near zero, but a prompt with no stdin fails the run (benign) — needs a genuinely non-interactive invocation, verified in Phase 2.
-- Partial ecosystem state: N/A for the GPU hazard — each pacman run is one atomic transaction, so there is no half-swapped library. The =--ignore= run is a partial upgrade in Arch's sense; pacman's dependency resolution is the safety net, and the residual unversioned-ABI exposure is accepted in Alternative E.
-- Orphans that block resolution: a package dropped from the repo but still pinning an old version (ratio's =qemu-block-gluster= on 2026-08-25) fails the whole transaction. The script should detect the "could not satisfy dependencies" case, name the foreign package, and stop with the remedy — never =-Rdd= on its own.
-- The kernel on a DKMS ZFS root: a failed =zfs-dkms= build after the kernel swap cannot be aborted (the DKMS hooks are PostTransaction) and leaves velox unbootable on the new kernel. Mitigated by holding the kernel set on every everyday run, landing it only in the dedicated session behind the gate, and by the standing fallback: =/boot= lives in the root dataset, the =05-zfs-snapshot= hook snapshots it before every transaction, and ZFSBootMenu can boot that snapshot. The pacman cache also keeps the previous kernel and =zfs-dkms= for a downgrade. Ratio's exposure is its data pool only (btrfs root, two kernels).
-- Standing kernel deferral: because the everyday run never moves the kernel, the dedicated session has to happen on a cadence or kernel security fixes sit unapplied. The panel's deferred row is the reminder; a stale-kernel age in maint is a possible follow-up.
+- Boot critical path: the unit runs ahead of the tty1 login, so a hang delays boot. A fixed 20-minute start timeout ends the run with SIGINT, the default stop timeout is the SIGKILL backstop, and nothing the session needs depends on the unit. The worst case is a 20-minute delay (plus 90 s if SIGINT doesn't end the run), then the login, with the record and maint's failed-units row both showing it. The flag leaves its path before the attempt starts, so every outcome, power loss included, disarms and no later boot retries. See Phase 3, The unit.
+- Boot package source: the boot form installs the flag's exact versions from the cache that arming filled, with no refresh and no network, so the flag, the cache and the system sync db must still agree at boot. A manual cache clean, or a refresh that didn't re-arm (a bare =pacman -Sy= or =yay=), between arming and reboot breaks that; the boot run then fails before any package changes, the record names it, and =upgrade-guarded --complete= re-arms. =paccache -k3= keeps the newest downloads, an armed everyday run re-arms against its own refresh, and =--complete= removes any flag right after its refresh, so only its own arm step leaves one. See Phase 1, The arm flag.
+- Interactive prompts under no stdin: pacman and yay can still reach a question (provider choice, replace, conflict) even with =--noconfirm=, which takes the default; a default of no fails the step, benignly. The boot form's single versioned transaction shrinks this to near zero, and the sweep's =-y= and =--no-ask-retry= keep a failed topgrade step from waiting on stdin under the panel's no-tty runner.
+- Partial ecosystem state: pacman is atomic per package only, and only when allowed to stop at a package boundary. It has no SIGTERM handler, so the unit stops it with SIGINT, which ends the transaction at the next boundary and releases =db.lck=. An interrupted boot run can leave part of the GPU/compositor set upgraded; the pre-written record already names the remedy. Power loss, or the SIGKILL backstop landing mid-package, can still leave a half-extracted package and a stale =db.lck=, which nothing here deletes: every later transacting run refuses and names it with the manual remedy. The =--ignore= run is itself a partial upgrade; the closure probe is its safety net, and I accept the residual unversioned-ABI exposure in Alternative E. See Phase 1, Preconditions and Exit codes and interrupts.
+- Orphans that block resolution: a package dropped from the repo but still pinning an old version (ratio's =qemu-block-gluster= on 2026-08-25) would fail the whole transaction. The closure probe sees it first, as a "breaks dependency" line whose requirer isn't held, and the script refuses with the names before any transaction, never with =-Rdd= or force. Removing or rebuilding the orphan stays manual. See Phase 1, Dependency closure.
+- Refreshed db after a refusal: a refused closure leaves the system sync db refreshed and nothing upgraded. Until an =upgrade-guarded= run succeeds again, don't install single packages with =pacman -S=: against the newer db that is a partial upgrade with no closure behind it. The deferred row shows the refusal at WARN until then.
+- The kernel and DKMS modules on a ZFS root: DKMS removes the old modules before the transaction and rebuilds after it (=70-dkms-upgrade=, =70-dkms-install=), so a failed =zfs-dkms= build aborts nothing, and mkinitcpio still writes an initramfs without =zfs.ko=, warning only. Velox then fails its next boot, whether the transaction moved the kernel or only =zfs-dkms=. So every everyday run holds the kernel and DKMS sets (and what the closure pulls in, such as =zfs-utils=), and they land only in =--complete='s stage 1 (or through yay, which then opens the gate entry), behind the =kernel-modules-check= gate. The gate entry is written before stage 1 starts, so a crashed or interrupted stage 1 already reads "do not reboot", and it closes only when a gate passes or the kernels it names are gone. While it is open, nothing arms or offers a reboot and the row is CRIT, so the machine keeps running the old kernel. When the entry is the run's own and stage 1 changed nothing kernel-side, the gate runs in its structural form, so a transaction that failed before committing raises no false alarm. Ratio's exposure is its data pool only (btrfs root); its GRUB default, the foreign =linux-lts-strix=, has no headers and is never pending, so neither stage 1 nor the gate touches it. See Phase 1, --complete sequence, The gate and Invariants.
+- Last-resort fallback on velox: an unplanned reboot while the gate is open, or a failure the gate misses, is recovered from ZFSBootMenu. =/boot= lives in the root dataset, which the =05-zfs-snapshot= hook snapshots before each transaction, so ZFSBootMenu can boot the old kernel, initramfs and modules. The hook skips a snapshot within 60 s of the last and only warns on failure, which is why the =--since= gate requires one taken no earlier than 60 s before its since time. =zfs-pre-snapshot= keeps only the newest ten, so =--complete= holds a pre-pacman snapshot from before its stage 1 and names it in the record (Phase 1, The snapshot hold), and the prune skips it. The cost is one pinned snapshot per machine, growing until a later =--complete=, or an everyday run whose AUR step opens the entry, holds a newer one; a failed hold only warns. The pacman db and cache are separate datasets that stay current, so booting the snapshot reverts every package changed since it was taken while the db still lists the new versions. Before the next =upgrade-guarded= run, reconcile the db to the booted root from =/var/log/pacman.log= (on =zroot/var/log=, also current); never roll back =zroot/var/lib/pacman=. See Phase 1, The snapshot hold, and Phase 4 for the recovery steps and the non-gating D-zbm drill.
+- Standing kernel deferral: because the everyday run never moves the kernel or DKMS sets, the dedicated session (=upgrade-guarded --complete=) has to happen on a cadence or kernel security fixes sit unapplied. The deferred row is the reminder, and it goes to WARN when a fixable advisory names a deferred package; a stale-kernel age in maint is the vNext follow-up (see Scope tiers).
+- Entry points outside the hold: bare =topgrade=, =yay= and =pacman -Syu= still land the kernel and DKMS sets with no gate, because the guard hook is silent on kernels. The hold covers only =upgrade-guarded='s callers: the panel levers, the =sysupgrade= alias where =upgrade-guarded= is installed, and the system-health-check workflow's update step.
+- AUR paths around the hold: yay resolves repo dependencies itself, so an AUR upgrade that needs a newer kernel-set or DKMS-set package installs it ungated. The everyday run detects that after yay and opens the gate entry, so the row goes CRIT and REBOOT hides until =--complete= verifies it (Phase 1, Everyday sequence, step 7); no installed foreign package has such a dependency today. A DKMS package from the AUR would bypass the hold entirely through =yay -Sua=, so the DKMS set must come from the sync repos (Phase 4's step (a) checks it). An AUR upgrade that needs a newer held GPU/compositor package while Hyprland is live trips the guard inside yay's own pacman call: the AUR step fails (the row shows WARN), the sweep still runs, and nothing stamps. See Phase 1, Everyday sequence.
+- Arch news: the driven path clears informant's feed under a 60 s timeout before any transaction, so once the clear succeeds, informant's AbortOnFail hook no longer stops a run. A failed or timed-out clear leaves the =00-informant= hook live for that run's transaction, which can then wait on the same unbounded fetch (bounded by the lever runner's 3600 s or the unit's 20 minutes, and in a terminal run only by Ctrl+C) or abort on unread news before any package changes, with a named failed step; I accept that residual. At boot, offline, the clear and the hook both fail fast against an empty feed, but a network that comes up in between can let the hook abort the transaction after the arm is consumed; =upgrade-guarded --complete= then finishes it. The deferred row is now where news gets seen: the everyday run and =--complete= capture unread titles with =yay -Pwq= before clearing, and the row shows them until DISMISS hides them. A manual-intervention notice waits on the row instead of blocking the transaction it warns about. See Phase 1, Everyday sequence and --apply-armed sequence, and Phase 2, The deferred row.
* Testing / Verification / Rollout
-Phase-1 and Phase-3 logic under the maint fake harness; Phase-2 install under =tests/installer-steps/=. The one thing no unit test can cover — that an armed reboot actually applies the upgrade pre-session and stamps — is a scripted manual test in =todo.org= (arm with a guarded lib pending, reboot, confirm the console run, the fresh metric, and a normal session). Roll to velox first, then ratio.
+Every automated test uses unittest, never pytest. archsetup's =make test-unit= globs =tests/*/test_*.py=, so the new suites run without a list edit, and the Phase 2 groups run in dotfiles =tests/maint/= without GTK.
+Each group and manual test named here is defined once, with its cases, in its phase: the P1 groups and M-split-console in Phase 1, the P2 groups in Phase 2, P3-unit, the M-boot tests and the VM procedure in Phase 3, and M-split-live and D-zbm in Phase 4. Every manual test is its own entry in =todo.org= under Manual testing and validation.
+
+Each acceptance criterion is verified by:
+- AC1 (the live everyday run): P1-sets, P1-closure, P1-steps, P1-record-stamp, P2-levers, P2-row, P2-cve-queue and M-split-live. M-split-live carries the hook-silent and exit-0 outcomes on both daily drivers.
+- AC2 (soname day and the closure refusal): P1-closure and P1-sets.
+- AC3 (the run with Hyprland not running): P1-sets, P1-record-stamp and M-split-console. P1-record-stamp carries the stamp clause, and M-split-console the held-set claim.
+- AC4 (the kernel-side session and its gate): P1-gate, P1-complete, P2-row, P2-apply and M-boot-armed. P2-apply pins the routing inside =panel.fire_press=, and M-boot-armed's APPLY press on ratio and velox carries the clause that the panel calls it and opens the detached =foot= terminal, since no automated test imports gui.py.
+- AC5 (arming and the armed boot): P1-arm-boot, P2-row, P3-unit, M-boot-armed and M-boot-ratio.
+- AC6 (a boot-upgrade failure never blocks the session): P1-arm-boot, P3-unit, M-boot-timeout, M-boot-midtx and M-boot-fail.
+- AC7 (Arch news): P1-steps, P2-levers (the lever timeout's SIGINT), P2-dismiss and M-boot-news.
+- AC8 (the =--complete= stamp): P1-record-stamp.
+- AC9 (APPLY, the wall note and the no-record row): P2-row, P2-apply and P2-text.
+- AC10 (a failed AUR step): P1-steps and P2-row.
+- AC11 (advisories on deferred packages): P2-cve-queue.
+- AC12 (the guard hook unchanged): the existing =tests/hypr-live-update-guard= suite, unchanged; P1-sets (fail closed); and P1-installer (TOML pin, hook name).
+- AC13 (the held snapshot on a ZFS root): P1-complete and P1-installer.
+P2-guard-ux gates no criterion. It pins Phase 2's guard-UX changes and runs with the rest of that suite.
+
+No unit test can drive the boot path itself: an armed reboot that applies the set on tty1 before the login, then lets boot continue however the run ends. That is why the M-boot tests are manual.
+I hold the boot unit back from velox until the =make test-keep FS_PROFILE=zfs= VM passes every M-boot test except M-boot-news (run on velox) and M-boot-ratio, plus M-split-console. The VM runs follow Phase 3's VM procedure. The VM arms over SSH rather than from the panel, so arming from APPLY, the armed row and REBOOT get checked on ratio and velox under AC5.
+Two runs show the real gate reading a root-only initramfs. The first is M-boot-armed's =--complete=, which must pass the gate on the VM's ZFS root. The second is the structural =kernel-modules-check= that Phase 4 step (a) runs on each ZFS root.
+
+Rollout, rollback, recovery and the non-gating D-zbm drill live only in Phase 4.
* Review and iteration history
+** 2026-10-05 Mon @ 12:10:00 -0500 — Craig Jennings — responder + reviewer
+- What: responded to three review rounds and moved the spec to READY. Round 1 (48 findings, 13 blocking): 7 accepted, 41 accepted with reconciling detail, none rejected. The design calls accepted were the pacman-resolved dependency closure, a boot form (=--apply-armed=) that installs a pre-downloaded set from cache with no network, a SIGINT/20-minute boot timeout that disarms at the start, and the UPDATE/TOPGRADE lever mapping with the script as the only writer of =topgrade_run=. Round 2 re-reviewed the revised spec (59 findings, 4 blocking: 36 accepted, 23 modified) and consolidated the exact contract into Implementation phases, which every other section now cites. It also cut machinery with no caller: =--no-aur=, unread record fields, kernel-modules-check's standalone mode, and the three-clause gate list, replaced by one pkgbase-keyed entry written before stage 1. Round 3 (22 findings, none blocking: 12 accepted, 10 modified) closed safety gaps around interrupts, the pre-stage-1 snapshot hold and power-cut recovery. A bounded repair pass followed; its 9 remaining non-blocking residuals stay open under Review findings.
+- Why: the decisions closed on 2026-08-25 left the first draft's text in place, and several assumptions about installed tools didn't hold (topgrade's argv, informant run as the user, root-only initramfs reads, checkupdates' db). Restating the contract in several sections caused drift between the copies, so I made one section authoritative.
+- Artifacts: the Review findings section; tool versions checked: topgrade 17.12.2, pacman 7.1.0, bash 5.3.20, sudo 1.9.17p2.
+** 2026-10-05 Mon @ 06:05:00 -0500 — Craig Jennings — reviewer
+- What: spec-review, rubric Not ready. Six independent passes read the spec against the live code in archsetup and the dotfiles (factual accuracy, internal consistency, boot-path safety, pacman/yay/topgrade mechanics, phasing and tests, panel UX), each finding was checked by an adversarial verifier that tried to refute it, and a completeness pass looked for gaps. 89 raw findings merged to 45 plus 3 from the completeness pass; 48 survived verification (13 blocking, 27 should-fix, 8 optional) and 2 were refuted (velox facts unverifiable from ratio, and a battery condition for the boot run). The Phases section is present and decomposable once the blockers close.
+- Why: the decisions closed on 2026-08-25 left text from the first draft in Design, Readiness and Risks, and several tool assumptions (topgrade's argv, informant run as the user, checkupdates' db, the kernel-set glob) didn't hold against the installed versions. The blockers cluster in four places: existing stamp writers that would mark a deferred run fresh, the split run's dependency handling, the kernel gate on ratio, and an undefined boot-run form (package source, timeout, news clearing).
+- Artifacts: the Review findings section above; tool versions checked this pass: topgrade 17.12.2, pacman and yay as installed on ratio.
+** 2026-10-05 Mon @ 05:14:17 -0500 — Craig Jennings — author
+- What: named the split script =upgrade-guarded= (was "name open") and added =containers= to the topgrade =--disable= list in Design and Phase 1.
+- Why: the containers step exits non-zero whenever local-only images exist (=cj/telega-server= and friends), which alone keeps topgrade from exiting 0. Pulling newer images under running containers isn't an upgrade path I use, and an ignore list would rot as images come and go.
+- Artifacts: velox health checks of 2026-09-12 and 2026-10-04.
** 2026-08-25 Tue @ 18:45 -0600 — Craig Jennings — author
- What: closed all seven decisions. Reversed the kernel decision (hold on every everyday run; land only in the dedicated session, gated on DKMS built, initramfs fresh, snapshot present; withdrew "install live at apply-on-reboot"), reworded the boot-scope decision for the split design, added the =kernel-modules-check= gate to Phase 1 and the acceptance criteria, and wrote the velox failure chain and the ZFSBootMenu fallback into Risks.
- Why: on velox a failed =zfs-dkms= rebuild after a kernel swap is unabortable and unbootable; that belongs in a session I chose, not in an update I expected to touch applications.