aboutsummaryrefslogtreecommitdiff
path: root/todo.org
diff options
context:
space:
mode:
authorCraig Jennings <c@cjennings.net>2026-10-05 23:11:01 -0500
committerCraig Jennings <c@cjennings.net>2026-10-05 23:11:01 -0500
commit32008ac10409e1a6f8ddb8820c6e8ac30e575db4 (patch)
tree6f163aa8c1f3b25b7cfd0092461af45a0b7c6dc0 /todo.org
parentec5a021194c4ab3b9d9fd2b260b183ad0bdf347f (diff)
downloadarchsetup-32008ac10409e1a6f8ddb8820c6e8ac30e575db4.tar.gz
archsetup-32008ac10409e1a6f8ddb8820c6e8ac30e575db4.zip
docs(spec): review the guarded-upgrade spec and decompose its build
I ran three review rounds against the live code in archsetup and the dotfiles. They found 129 issues, and the blocking count fell from 13 to 4 to 0. Each one is dispositioned in the spec. The exact contract now lives only in Implementation phases, because the copies in Design, Decisions and Testing had drifted. The spec moves from DRAFT through READY to DOING, with nine non-blocking residuals left for the first build task. The parent task now carries 22 build tasks in commit-group order, ending with the IMPLEMENTED flip. The manual tests are under Manual testing and validation, and a stale-kernel age metric is filed as [#D].
Diffstat (limited to 'todo.org')
-rw-r--r--todo.org1655
1 files changed, 1653 insertions, 2 deletions
diff --git a/todo.org b/todo.org
index ed72255..015f8b0 100644
--- a/todo.org
+++ b/todo.org
@@ -711,7 +711,7 @@ Grading: feature, no hard date, real improvement to the install = [#B].
:solo: — build path (installer + checker + tests) and verify path (pytest;
the live proof already exists on velox) with no open decision.
-** TODO [#A] Topgrade guarded-upgrade spec — decisions, review, decomposition :feature:maint:dotfiles:
+** TODO [#A] Topgrade guarded-upgrade build :feature:maint:dotfiles:
SCHEDULED: <2026-09-23 Wed>
:PROPERTIES:
:CREATED: [2026-08-25 Tue]
@@ -725,7 +725,7 @@ xorg-xwayland under a live Hyprland) plus any failing ecosystem step makes that
exit almost unreachable. Diagnosed 2026-08-24/25; the fix is specced, not
hacked, because it spans two repos and the design is contested.
-Spec: [[file:docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org][2026-08-25-topgrade-guarded-upgrade-spec.org]] (DRAFT).
+Spec: [[file:docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org][2026-08-25-topgrade-guarded-upgrade-spec.org]] (DOING).
Four open decisions, all mine to make before the spec can move:
1. Freshness means *state* (a guarded upgrade still un-applied stays stale),
@@ -763,6 +763,629 @@ DRAFT → READY, and spec-response into build tasks. velox has
=linux-lts 6.18.51 → 6.18.52= pending today, which is the kernel-hold case this
spec exists for, so no plain topgrade on velox until the hold is built or the
kernel update runs as its own session.
+*** 2026-10-05 Mon @ 06:08:00 -0500 Spec-review ran: Not ready, 13 blocking findings
+Before the review I named the script =upgrade-guarded= and added =containers= to
+its topgrade =--disable= list. The review recorded 48 findings in the spec's
+=Review findings= section (13 blocking, 27 should-fix, 8 optional); the spec stays
+DRAFT until spec-response dispositions them. The blockers cluster in four places:
+the existing stamp writers (wrapper and doctor.py) would mark a deferred run
+fresh; the split run's =--ignore= needs a dependency closure; =kernel-modules-check=
+as written fails forever on ratio's headerless =linux-lts-strix=; and the boot run
+has no defined form (package source without network, timeout that can kill pacman
+mid-commit, =informant read= run as the user). Also: the installed topgrade
+rejects a comma-joined =--disable= list, so each step is its own argument.
+
+The 2026-10-04 velox health check ran a plain topgrade that moved the kernel
+6.18.51 → 6.18.55, against the note above. I checked the gate by hand before the
+reboot (=dkms status= installed for 6.18.55, =zfs.ko= in both initramfs images
+read with sudo) and the reboot was clean, but the hold this spec builds would
+have kept the kernel out of that run.
+
+Next: spec-response on the findings, then DRAFT → READY and decomposition into
+build tasks here.
+
+*** 2026-10-05 Mon @ 12:30:00 -0500 Spec READY; build tasks below
+The spec went DRAFT → READY at 12:10 after three review rounds. Its
+Implementation phases section is the contract every task below points at,
+and it moves to DOING with these tasks (the first task confirms that). Nine
+non-blocking residual findings stayed open, and the first task closes them
+before any code. The four-decisions list in the body above is history.
+
+The order follows the spec's commit groups: archsetup Phase 1, rollout step
+(a) on both daily drivers, dotfiles Phase 2 and step (b), archsetup Phase 3
+and step (c) (velox only after the VM gate), the Phase 4 README, then the
+flip. While Phase 2 sits unpushed, I push nothing from that dotfiles
+checkout, because any push carries it.
+
+On velox nothing runs plain topgrade, =yay -Syu= or =pacman -Syu=: not a
+shell, not the =sysupgrade= alias, and not the panel's UPDATE or TOPGRADE,
+which run =yay -Syu= and topgrade until step (b) lands Phase 2 there. Before
+step (a), the kernel and zfs-dkms move only in a session I start on purpose,
+with the gate checked by hand as on 2026-10-04. From step (a) until step
+(b), I update velox only with =upgrade-guarded= from a terminal.
+*** TODO Close the nine residual spec findings :chore:maint:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable: the spec, with all nine "round 3 residual" findings under
+Review findings applied and completed in place, the way the other findings
+were: =DONE=, a =CLOSED:= line, and a =Disposition:= line (accepted, or
+modified with the reason). Repo: archsetup, the spec file only.
+
+First the status heading. spec-response flips it READY → DOING when it files
+these tasks. If it still reads READY when I start, I flip it before anything
+else, in three lines: the keyword DOING, a newest-first history line ("READY
+→ DOING. Decomposed into build tasks under the archsetup todo.org parent,
+SPEC_ID 81cdfd72-db96-43d3-aa03-779878c99f3e."), and the Metadata Status
+cell set to doing.
+
+The line numbers inside the findings predate the round-3 consolidation, so I
+locate each edit by its quoted OLD text, never by line number:
+- Decision 5 Consequences: the snapshot-hold clause ("the hold moves only
+ when a later session holds a newer one").
+- AC13's everyday claims: P1-steps' step-7 bullet gains the release of the
+ previous held_snapshot, and Testing maps AC13 to P1-steps too.
+- Finish step 1: when several steps fail, the first one's failed_step and
+ detail win.
+- Recovery's power-cut bullet: register the version after =->=, and nothing
+ when the line is a removed line or there is none.
+- Exit codes and interrupts: narrow the bare-=wait= qualifier.
+- ZFSBootMenu: roll back with MOD+R (never ENTER, MOD+X or MOD+C) in
+ Recovery, D-zbm and its variant.
+- Design and Decision 5: REBOOT hides from the panel's next probe, and the
+ fire-time refusal covers the window before it.
+- The D-zbm variant: =ssh -t=, the headers' upgrading line as the reset cue,
+ and a repeat on a fresh VM when the reset missed extraction.
+- P1-steps' fake-sudo bullet and Readiness: =yay -Sua= is what never runs,
+ since =yay -Pwq= already ran.
+The snapshot-hold and REBOOT findings both rewrite Decision 5's
+Consequences, so I apply those two together. I also reword the three places
+that still say ZFSBootMenu boots the held snapshot (Design's ZFS fallback
+paragraph, the note after AC13, and Risks' last-resort bullet) to a MOD+R
+rollback, so no section points a reader at ENTER.
+
+This task goes first because several findings change test bullets (P1-steps)
+and the manual procedures (Recovery, D-zbm) that later tasks build against.
+
+Verified by re-reading. The findings quote their own OLD text, so a
+whole-file search always matches; I check outside Review findings only. For
+each replaced OLD string this prints 0, and for each inserted or replacing
+NEW sentence (the first-failure rule in Finish is an insertion, so its anchor
+stays) it prints 1:
+: sed '/^\* Review findings/,/^\* Implementation phases/d' SPEC | grep -cF 'TEXT'
+Each residual reads DONE with a CLOSED line and a disposition. The Review
+findings heading cookie reads 138 of 138, updated by hand or with =emacs
+--batch= and =org-update-statistics-cookies=, because an edit outside Emacs
+doesn't recompute it. The status heading reads DOING. It's a doc edit, so
+there's no suite to run.
+Contract: Review findings, the nine "round 3 residual" entries.
+
+*** TODO Thresholds mirror, check 9 and hold-aware prune :feature:installer:zfs:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (archsetup): Phase 1's three standalone changes, one commit
+each, each green on =make test-unit=:
+- =configs/maintenance-thresholds.toml='s =[updates] guard_patterns=
+ rewritten to exactly the hook heredoc's 15 Target patterns, with the
+ comment reworded; the TOML pin test in =tests/installer-steps/=.
+- =scripts/post-rebuild-check= check 9 (the =10-= hook name wherever the
+ guard is installed) behind =PRC_GUARD_BIN= and =PRC_PACMAN_HOOKS_DIR=, with
+ the header, usage and counters moved to 9 and every test env setting
+ =PRC_GUARD_BIN= empty.
+- =scripts/zfs-pre-snapshot='s prune skips rows with userrefs above 0, and
+ =tests/zfs-pre-snapshot/fake-zfs= gains hold, release and userrefs, which
+ the =--complete= tests reuse.
+Tests: P1-installer's TOML pin, post-rebuild-check and zfs-pre-snapshot
+cases.
+
+None of these is live on a machine until step (a) installs or re-copies it,
+so they land first. Check 9 flags a legacy hook name if post-rebuild-check
+runs before step (a); that's the check working.
+Verified by =make test-unit= green before and after, plus a mutation spot
+check: hand-edit one TOML pattern and the pin test goes red.
+Contract: Phase 1, Other Phase 1 changes; Phase 1 tests, P1-installer.
+
+*** TODO kernel-modules-check gate :feature:zfs:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (archsetup): =scripts/kernel-modules-check=, the stateless
+reader that says whether each named kernel would boot with its modules:
+usage and exit codes, the per-pkgbase items, the =--since= initramfs and
+snapshot items, the =sudo -n lsinitcpio= read on a ZFS root, and the seams
+=KMC_MODULES_DIR=, =KMC_BOOT_DIR= and =KMC_DKMS=. It doesn't depend on
+=upgrade-guarded=, so it lands first.
+Tests: P1-gate in =tests/kernel-modules-check/=, with fakes for dkms,
+findmnt, =zfs list= and an lsinitcpio that fails unless run through the
+fake sudo.
+Verified by =make test-unit= green, plus a read-only real run from the repo
+on ratio against =linux-lts=, expecting exit 0 (btrfs root, so no sudo
+read). I don't point it at the running =linux-lts-strix=: its zfs is only
+'added', and the gate never checks it. The first real =sudo -n lsinitcpio=
+read is step (a)'s structural run on velox.
+Contract: Phase 1, The gate; Phase 1 tests, P1-gate.
+
+*** TODO upgrade-guarded foundation: modes, preconditions, sets, record :feature:maint:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (archsetup): =scripts/upgrade-guarded= up to, but not
+including, the closure and the transaction steps:
+- mode and option parsing (exit 2, no record);
+- the preconditions in order: usage, EUID 0, the flock held on fd 9,
+ =db.lck=, and the =10-= hook's Target lines;
+- the sets: pending (with =pacman -Qu='s exit-1-empty rule and its
+ =[ignored]= rows dropped), GPU patterns, kernel set, DKMS set and live;
+- the record (maint's cache envelope, tmp file then rename) and Finish's
+ result, failed_step and detail rules, with the reason as the last stderr
+ line;
+- the exit codes, and the INT/TERM/HUP trap that waits on the child's PID;
+- =--dry-run= on its private =CHECKUPDATES_DB= with =checkupdates --nocolor=;
+- the harness: every seam, the fakes on PATH, and subprocess-driven
+ unittest suites.
+Nothing calls the script until Phase 2, so a partial script is harmless.
+Tests: P1-sets' cases that need neither the closure nor a transaction (the
+three-kernel fixture with the foreign headerless kernel, firmware and
+api-headers, and the unowned vmlinuz, minus its gate.pkgbases clause; the
+empty DKMS set; the =-Qu= exit handling; the fail-closed hook), plus a
+supporting case that the pending set drops =[ignored]= rows; and P1-steps'
+usage, EUID 0, lock, =db.lck= and =--dry-run= cases (the probe =--dbpath=
+clause lands with the closure). P1-sets' post-transaction deferred-set case
+waits for the everyday-run task, and the gate.pkgbases clause for the
+=--complete= task.
+Verified by =make test-unit= green, plus a real =--dry-run= from the repo on
+ratio. It refuses with exit 3 naming the missing =10-= hook, because ratio
+keeps the legacy name until step (a). With =UPGRADE_GUARDED_HOOK= pointed at
+the legacy file, it prints ratio's real sets, and =linux-lts-strix= isn't
+pending.
+Contract: Phase 1, Identity and privilege; Modes and options;
+Preconditions; Sets; The record; Finish; Exit codes and interrupts; Seams
+and fakes.
+
+*** TODO upgrade-guarded dependency closure :feature:maint:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (archsetup): the closure loop, meaning the read-only =-Sup=
+print-mode probe, the everyday stage-A seed, the live-only stage-B seed,
+the =--complete= GPU seed, the two line forms, every refusal rule with its
+detail, the iteration bound, the kinds each stage assigns, and the held set
+built from them. =--dry-run= then prints the GPU closure, the predicted
+deferred set and the would-be =-Su= argv. I keep this apart from the
+transaction steps because it's the riskiest logic in Phase 1, and it gets
+its own verification pass.
+Tests: P1-closure, except its single =-Sy= case and the "=-Su= runs" clause
+of its new-dependency case, which need the everyday transaction (next
+task); P1-sets' held-set cases (the full =--ignore= list with each entry its
+own element, the everyday half of the live and not-live case, the
+=zfs-dkms=/=zfs-utils= pair held together, a pkgrel-only bump not held),
+asserted on =--dry-run='s printed =-Su= argv until the next task re-points
+them at the fake pacman's real call; P1-steps' =--dry-run= probe =--dbpath=
+clause. The orphan case's "no =-Su= runs" clause is vacuous until the next
+task adds the transaction, and bites from then on.
+Verified by =make test-unit= green, plus the real =--dry-run= on ratio
+(hook seam as before), expecting a closure without a refusal. A refusal
+naming a real orphan is itself a finding I clear by hand before step (a).
+Contract: Phase 1, Dependency closure; Sets (Kinds, Held set, GPU closure).
+
+*** TODO upgrade-guarded everyday run, arm procedure and stamp :feature:maint:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (archsetup): the everyday sequence, steps 1 to 10, except step
+7's kernel-side detection, which lands with the snapshot-hold task:
+- =yay -Pwq= news capture and the record's merge, dedup and cap;
+- the informant clear under =timeout 60 sudo -n=;
+- one =sudo -n pacman -Sy=, then =-Su= with the =--ignore= list and no =-y=;
+- =yay -Sua= on every run, and the sweep argv by absolute path;
+- step 9's flag cases, with the arm procedure (resolve check, =-Sw=, the
+ =install= and =mv=) and the unit-enabled seam;
+- Finish's stamp step through =UPGRADE_GUARDED_MAINT=, the everyday stamp
+ predicate, and the first-failure rule the first task adds to Finish;
+- the canonical fixture
+ =tests/upgrade-guarded/fixtures/upgrade_deferred.json= from a fixed fake
+ run.
+Tests: P1-steps (the =sudo -n= refusal, news, informant, the exit-124
+abort, the sweep argv and its real-binary =--version= check, the failing
+yay, the INT cases); P1-closure's single =-Sy= case and the "=-Su= runs"
+clause of its new-dependency case; P1-sets' post-transaction deferred-set
+case (a fake post-transaction =-Qu= intersected with the held set, excluding
+an =[ignored]= row), with its held-set cases re-pointed at the real =-Su=
+call; P1-record-stamp's everyday cases and the fixture equality;
+P1-arm-boot's two everyday flag cases; and one case for the first-failure
+rule: a failing fake yay and a failing fake sweep leave failed_step aur and
+yay's detail.
+Verified by =make test-unit= green on ratio, where =/usr/bin/topgrade=
+exists, so the real-binary check runs instead of skipping. The first run
+against real pacman is M-split-live at step (a).
+Contract: Phase 1, Everyday sequence; The arm flag; Stamp predicate;
+Finish; The record (Fixtures).
+
+*** TODO upgrade-guarded --complete and its gate entry :feature:maint:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (archsetup): the =--complete= sequence, steps 1 to 7, built and
+tested on fixtures that aren't a ZFS root (the fake findmnt reports another
+filesystem, so no snapshot or hold runs):
+- refresh then flag removal, 'nothing to complete', the pre-stage-1 record
+ and gate entry, stage 1 (with the reinstall targets when a gate is
+ already open), the gate in its structural or =--since= form, the GPU step
+ (stage 2, arm, or the unit-not-enabled line), and the tty-only reboot
+ prompt after Finish;
+- the =--complete= stamp predicate and the Invariants.
+Nothing calls or installs the script until Phase 1 is done, so a
+=--complete= without its hold is harmless; the next task adds the hold.
+Tests: P1-complete except its three snapshot-hold bullets; the
+gate.pkgbases clause of P1-sets' three-kernel case; P1-record-stamp's
+=--complete= cases; P1-arm-boot's =--complete= arm, unit-not-enabled and
+prompt cases; the =--complete= half of P1-sets' live and not-live case; the
+=--complete= side of P1-steps' news and informant cases.
+Verified by =make test-unit= green.
+Contract: Phase 1, =--complete= sequence; Invariants; Stamp predicate;
+Exit codes and interrupts (the prompt).
+
+*** TODO Snapshot hold and the everyday AUR kernel-side check :feature:maint:zfs:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (archsetup), after the =--complete= task:
+- =--complete= step 3's own snapshot and hold on a ZFS root, the post-gate
+ =--since= target rule, and the release of the previous hold;
+- everyday step 7's kernel-side detection after yay: the gate entry, the
+ pre-yay snapshot hold, the flag removal and the immediate record write;
+- the trap's matching step-7 path.
+Tests: P1-complete's three snapshot-hold bullets; P1-steps' AUR kernel-side
+case, with its INT cases and the release-of-the-previous-hold clause the
+first task adds. Its "a following =--complete= runs the =--since= form"
+clause needs the previous task, which is why this one comes after it.
+Verified by =make test-unit= green. The first real snapshot and hold is
+M-boot-armed's =--complete= in the VM, which opens the entry on a ZFS root.
+Contract: Phase 1, The snapshot hold; Everyday sequence (step 7); Exit
+codes and interrupts; Invariants.
+
+*** TODO upgrade-guarded --apply-armed boot form :feature:maint:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (archsetup): =--apply-armed=, steps 1 to 8. That means reading
+the list the unit moved off the flag path (refusing a missing, empty or
+unparseable list, or a live Hyprland), the interrupted pre-write, the tty1
+banner, the journal mirror through a SIGINT-ignoring process substitution,
+the informant clear, the =vercmp= drop of entries already at or above
+their version, the offline =-Sp= check that refuses kernel-side members and
+stale versions, the =-S --needed= transaction from the cache, and Finish
+with its stamp predicate.
+Tests: P1-arm-boot's =--apply-armed= cases, including the =setsid=
+process-group SIGINT case (no SIGPIPE death, the fake lock released, the
+record interrupted); P1-steps' clause that =--apply-armed= never calls
+=yay -Pwq=, and its =--apply-armed= informant and exit-124 cases;
+P1-record-stamp's =--apply-armed= success case.
+Verified by =make test-unit= green. The real boot path is the VM gate.
+Contract: Phase 1, =--apply-armed= sequence; Exit codes and interrupts;
+Stamp predicate.
+
+*** TODO Script install step and health-check workflow edit :feature:installer:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (archsetup), two commits that close Phase 1's group:
+1. The =hyprland()= step that installs =hypr-live-update-guard= also installs
+ =upgrade-guarded= and =kernel-modules-check= to =/usr/local/bin=, with
+ P1-installer's source-inspection test.
+2. A doc commit after the code: =docs/workflows/system-health-check.org=
+ Phase 3, per the contract's seven bullets. Step 6 runs =upgrade-guarded=
+ where step (a) is done. Kernel and DKMS days go through =--complete=.
+ Until step (c), the GPU set finishes from a console. Where step (a)
+ isn't done and a kernel or DKMS package is pending, plain topgrade
+ doesn't run. The strix addendum is keyed to each invocation. The script
+ stamps, so there's no hand stamp. The 2026-09-12 containers entry gets a
+ resolution note.
+After this, Phase 1 is complete and step (a) can start.
+Verified by =make test-unit= green, and the workflow edit re-read against
+the contract's bullets.
+Contract: Phase 1, Files and install; The system-health-check edit; Phase
+1 tests, P1-installer.
+
+*** TODO Rollout step (a) on ratio :chore:ratio:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+After Phase 1's last commit, by hand on ratio, in the contract's order:
+- pull archsetup, then =sudo install -m 755= both scripts to
+ =/usr/local/bin/=;
+- write =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook= from the
+ installer heredoc, confirm its Targets, remove the legacy unprefixed hook
+ and check =ls /etc/pacman.d/hooks/=;
+- diff, then re-copy =configs/maintenance-thresholds.toml= to
+ =~/.config/archsetup/=;
+- confirm =pacman -Qmq= lists no DKMS-set package;
+- confirm =upgrade-guarded --dry-run= exits 0;
+- run M-split-live on ratio (Manual testing and validation).
+Ratio's root is btrfs, so the =zfs-pre-snapshot= install and the structural
+gate run belong to velox's step (a).
+Verified by M-split-live on ratio passing, and post-rebuild-check's check 9
+clean here (the =10-= hook present, the legacy name gone). If I build Phase
+2 on ratio, its dotfiles checkout is the stowed, live one, so each Phase 2
+commit is live on ratio as it's written; that's why Phase 2 starts only
+after step (a) on the build machine.
+Contract: Phase 4, Rollout (a) and M-split-live.
+
+*** TODO Rollout step (a) on velox :chore:velox:zfs:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+After Phase 1's last commit, by hand on velox: the same list as ratio's
+step (a), plus the ZFS-root items:
+- =sudo install -m 755 scripts/zfs-pre-snapshot /usr/local/bin/= for the
+ hold-aware prune;
+- =kernel-modules-check "$(cat /usr/lib/modules/$(uname -r)/pkgbase)"= exits
+ 0, which proves the =sudo -n lsinitcpio= read against the real image;
+- velox's guard hook was hand-placed under the legacy name, so the =10-=
+ rewrite and the legacy removal matter here too;
+- run M-split-live on velox.
+On velox nothing runs plain topgrade, =yay -Syu= or =pacman -Syu=: not a
+shell, not the =sysupgrade= alias, and not the panel's UPDATE or TOPGRADE,
+which run =yay -Syu= and topgrade until step (b) lands Phase 2 there. Before
+this task, the kernel and zfs-dkms move only in a session I start on
+purpose, with the gate checked by hand: =dkms status= installed for the new
+kernel, and =zfs.ko= in each initramfs, read with sudo. From here until
+step (b), I update velox only with =upgrade-guarded= from a terminal, and
+the health-check workflow's step 6 runs it.
+Phase 2 can be built on ratio while this waits, but I don't push any Phase 2
+commit until this is done, because any dotfiles pull on velox would bring it
+in.
+Verified by M-split-live on velox, the structural gate run exiting 0, and
+post-rebuild-check's check 9 clean here (the =10-= hook present, the legacy
+name gone).
+Contract: Phase 4, Rollout (a) and M-split-live.
+
+*** TODO maint deferred row, APPLY and gate-aware REBOOT :feature:maint:dotfiles:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (dotfiles), one commit, because the row and APPLY land together:
+- the =upgrade_deferred= probe and the helper it shares with
+ =topgrade_freshness= (N, never raises, None on a missing or malformed
+ record), the severity order, the armed text from the flag, and the
+ evidence (rows, detail, held_snapshot, and news lines not in
+ =upgrade_news_dismissed=, which stays empty until DISMISS lands);
+- =apply_deferred= (tier confirm, kind terminal), its =LEVER_KEYS= builder,
+ =topgrade_freshness='s =deferred= evidence, =doctor.review='s routing,
+ =panel.fire_press= launching =foot --hold -e upgrade-guarded --complete=
+ detached, and =maint fix apply_deferred= on inherited stdio;
+- Freshness and REBOOT: the wall note, REBOOT hidden while a gate is open
+ and shown while armed, =iter_fix='s fire-time refusal, and =review=
+ omitting =reboot= while a gate is open;
+- the byte-identical fixture copy at
+ =tests/maint/fixtures/upgrade_deferred.json=, with its test module's
+ docstring naming the archsetup source.
+I land this before the lever repoint, so gate-aware REBOOT already exists
+the first time a lever runs =upgrade-guarded=, whose AUR step can open a
+gate. Until the repoint the row reflects only shell runs, and with no
+record it reads OK '0 deferred'.
+
+Repo: =~/.dotfiles=, a separate project. Confirm the crossing before
+starting, or run it from a dotfiles session. This starts after step (a) on
+the machine I build Phase 2 on (=uname -n=): its dotfiles checkout is stowed
+and live, so each commit reaches that machine as it's written. Baseline:
+=make test= in =~/.dotfiles=; the red screen-lock suite on ratio is the
+known failure, tracked in "screen-lock test suite red on ratio".
+Tests: P2-row, P2-apply, and P2-text's wall-note case. No test imports
+gui.py.
+Verified by the =tests/maint= suites and =make test= green apart from that
+known red; =cmp= on the two fixture copies; and =make test-panel-maint=, the
+AT-SPI smoke over the fixture boards. =make test= never runs it, and it's
+the only check that drives gui.py, which this commit edits. Ratio has no
+weston or sway, so the smoke uses the live compositor and opens a panel
+window; I run it on a spare workspace, and a SKIP is not a pass. The real
+APPLY press is M-boot-armed on ratio and velox.
+Contract: Phase 2, The deferred row; APPLY; Freshness and REBOOT; Phase 1,
+The record (Fixtures).
+
+*** TODO maint levers repointed to upgrade-guarded :feature:maint:dotfiles:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (dotfiles), the one commit the spec ties together:
+- UPDATE and TOPGRADE argv and the missing-binary message;
+- doctor.py's =topgrade_run= stamp deleted;
+- the guard UX removed: the =live_update= tag, =iter_fix='s refusal and
+ guard event, =--force= and its sentinel wrap, =_update_force=,
+ =_rearm_after_guard=, the guard branches in =_on_fired=, cli.py and
+ panel.py, and the review suffix; =panel.guarded= becomes a set test, and
+ the tripped arm line becomes the annotation;
+- the =sysupgrade= conditional in both common alias files, with minimal's
+ files left as symlinks;
+- the stale text: the =topgrade_freshness= docstring and absent-stamp
+ error, the doctor.py and guard.py docstrings, the README =maint fix=
+ usage line, and the wrapper's header and its test docstring;
+- =tests/maint/panel_smoke.py='s "arm line shows the exact update argv"
+ check, which pins =yay -Syu --noconfirm=, now matching
+ =upgrade-guarded --no-topgrade=. The spec doesn't list the file, but the
+ repoint breaks it.
+The runner's long-remedy timeout (new session, SIGINT to the group, 120 s,
+then SIGKILL, 'timed out — interrupted') can land as its own commit just
+before this one. UPDATE and TOPGRADE are already user-kind and marked long,
+so it applies to today's =yay -Syu= and topgrade argv the moment it lands.
+That's safe: today a timeout SIGKILLs only the direct child
+(=subprocess.run='s timeout), and a SIGINT to the whole group stops pacman
+at a package boundary instead.
+
+Repo: =~/.dotfiles=, a separate project. Confirm the crossing before
+starting, or run it from a dotfiles session.
+Tests: P2-levers, P2-guard-ux (including the rewritten
+test_panel_levers.py and test_panel_phase10.py cases), and P2-text's
+alias, absent-stamp and usage-line cases.
+Verified by the =tests/maint= suites and =make test= green apart from the
+known screen-lock red, and =make test-panel-maint= passing on a spare
+workspace (a SKIP is not a pass). After it, the first UPDATE press on the
+build machine runs the same command M-split-live already ran there.
+Contract: Phase 2, Levers; Guard UX; Alias and README (the alias).
+
+*** TODO maint DISMISS, held-package CVE/QUEUE/strip, README :feature:maint:dotfiles:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (dotfiles), three commits after the repoint:
+1. DISMISS: the =dismiss= key kind, dispatched explicitly in =_digest_key=,
+ arm-then-fire, and the GTK-free panel.py helper that overwrites
+ =upgrade_news_dismissed= with the record's whole news list. Tests:
+ P2-dismiss.
+2. CVE, QUEUE and the strip: deferred names out of =cve_queued=, the CVE
+ caption, UPDATE's binding and REVIEW & FIX; =[HELD]= tags in QUEUE; and
+ the strip's 'P pending · N held'. Tests: P2-cve-queue.
+3. The 'The live-update guard' section of =maint/README.md= rewritten for
+ the new levers, the sweep argv, the deferred row, APPLY, DISMISS and the
+ armed state, with no press-again or =--force=. The guard banner and the
+ wrapper paragraph stay.
+Then I push Phase 2, and only once step (a) is done on velox. Until then I
+push nothing from this dotfiles checkout, because any push carries it.
+
+Repo: =~/.dotfiles=, a separate project. Confirm the crossing before
+starting, or run it from a dotfiles session.
+Verified by the =tests/maint= suites and =make test= green apart from the
+known red, =make test-panel-maint= passing on a spare workspace (the strip's
+pending caption is in its path; a SKIP is not a pass), and the README
+section re-read against the contract.
+Contract: Phase 2, DISMISS; CVE, QUEUE and the strip; Alias and README.
+
+*** TODO Rollout step (b) on the other daily driver :chore:dotfiles:velox:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Pull the Phase 2 commits into the other daily driver's dotfiles, only after
+step (a) is done there. The build machine already has them through its live
+checkout, so this step is for the other one: velox when I build Phase 2 on
+ratio. If I build it on velox, this is ratio's step (b) and the topic tag
+moves to =:ratio:=.
+Verified by that machine's dotfiles HEAD matching the remote, the
+=tests/maint= suites green there, and =maint status= there listing the
+=upgrade_deferred= row.
+Contract: Phase 4, Rollout (b); the Phase 2 intro.
+
+*** TODO archsetup-boot-upgrade.service installer step :feature:installer:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (archsetup): an installer step that writes the unit from a
+heredoc with =ARCHSETUP_USERNAME= substituted by sed, runs
+=install -d -m 0755 /var/lib/archsetup=, and enables the unit without
+=--now=, documented in its own comments. The spec names no function. It
+can sit beside the guard install in =hyprland()=, or in a top-level
+function called from there, and the source-inspection idiom reads either.
+Tests: P3-unit in =tests/installer-steps/=: every directive the contract
+lists, plus the absence of Requires=, BindsTo=, RequiredBy=, network-online
+and Restart=.
+Verified by =make test-unit= green. The boot behavior itself is the VM gate.
+Contract: Phase 3, Installer step; The unit; Disarm and failure; Phase 3
+tests, P3-unit.
+
+*** TODO Rollout step (c) on ratio :chore:ratio:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+After Phase 3's commit, by hand on ratio: write the unit from the installer
+heredoc with the username substituted, then run
+=sudo install -d -m 0755 /var/lib/archsetup=, =sudo systemctl daemon-reload=
+and =sudo systemctl enable archsetup-boot-upgrade.service=. The VM gate
+holds back only velox, so ratio can go first.
+Then, on the first day =upgrade-guarded --dry-run= lists a GPU-kind
+package, run M-boot-armed on ratio and M-boot-ratio on the same armed boot.
+Run the strix check on the first =--complete= that lands linux or
+linux-lts. All three are under Manual testing and validation.
+Verified by those three manual tests.
+Contract: Phase 4, Rollout (c) and the ratio-path confirmation.
+
+*** TODO VM gate for the boot unit :test:installer:zfs:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Run the five VM-gate tests in the =make test-keep FS_PROFILE=zfs= VM,
+following Phase 3's VM procedure, in this order: M-split-console,
+M-boot-armed, then M-boot-midtx on the same VM, then M-boot-timeout and
+M-boot-fail. M-boot-midtx needs the VM M-boot-armed leaves, because it
+times its cut from that boot's journal, so nothing reinstalls between them.
+Each test is a child of Manual testing and validation.
+=make test-keep= bundles archsetup from HEAD, so Phase 1 and Phase 3 must be
+committed first. The VM clones the published dotfiles, so this runs after
+the Phase 2 push; M-boot-timeout and M-boot-fail then read maint's
+failed-units and deferred rows inside the VM, which AC6 asks for. I never
+use =debug-vm.sh= here: it restores the clean-install snapshot and erases
+the install under test.
+M-boot-timeout and M-boot-fail use small inline stand-in fakes in place of
+the Phase 1 fake sudo and pacman: they pass every query to the real pacman
+and sleep or fail only on the transaction, so the VM tests don't depend on
+how the unit-test fakes are laid out.
+The unit reaches velox only after all five pass.
+Contract: Phase 3, VM procedure and VM gate; Testing / Verification /
+Rollout.
+
+*** TODO Rollout step (c) on velox :chore:velox:zfs:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Only after the VM gate passes: the same four steps as ratio's step (c).
+Then, on the first day =upgrade-guarded --dry-run= lists a GPU-kind
+package, run M-boot-armed on velox and M-boot-news.
+Verified by those two manual tests.
+Contract: Phase 4, Rollout (c).
+
+*** TODO maint README flow and recovery steps :chore:maint:dotfiles:solo:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+Deliverable (dotfiles), one commit to the 'The live-update guard' section
+of =maint/README.md=, after Phase 3. It adds the flow line, the
+=kernel-modules-check <pkgbase>= diagnosis note for a CRIT row (only
+=--complete= clears the gate), and the recovery steps as their only copy.
+The steps come from the spec's Recovery as the first task amends it: the
+MOD+R rollback and the version-after-the-arrow rule.
+Repo: =~/.dotfiles=, a separate project. Confirm the crossing before
+starting, or run it from a dotfiles session.
+The non-gating D-zbm drill and its power-cut variant exercise these steps
+(Manual testing and validation), and nothing waits on them.
+Verified by =make test= green apart from the known screen-lock red, and the
+section re-read against the contract.
+Contract: Phase 4, README flow and recovery; Recovery; Recovery drill.
+
+*** TODO Flip the spec to IMPLEMENTED (+ dated history line + Metadata mirror) :chore:maint:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+This waits until every task above is done and the gating manual tests have
+passed: M-split-live on both machines, the five VM-gate tests,
+M-boot-armed on ratio and velox, M-boot-news, M-boot-ratio and the strix
+check. Then tick each acceptance criterion against the evidence that
+Testing / Verification / Rollout maps to it. AC1 stays unticked until a live
+run with a GPU-kind entry pending has shown the hook silent. Last, the spec
+edits: the status keyword DOING → IMPLEMENTED, a dated history line that
+says why, and the Metadata Status mirror set to implemented. D-zbm is
+non-gating and doesn't hold this.
+Contract: the spec's status heading and Metadata table.
+
+** TODO [#D] Stale-kernel age in maint :feature:maint:dotfiles:
+:PROPERTIES:
+:LAST_REVIEWED: 2026-10-05
+:END:
+The guarded-upgrade build holds the kernel and DKMS sets on every everyday
+run, so kernel security fixes wait for an =upgrade-guarded --complete= I
+start on purpose. The deferred row is the only reminder, and it shows a
+kernel most days, so it stops reading as a nag. This is the spec's vNext: a
+maint age for how long the kernel set has been held, graded so an overdue
+dedicated session shows up.
+
+The design is still open. The record has no held-since field, so the age
+needs either a new field (both fixture copies change in one rollout) or a
+derivation from pacman.log. maint's CVE probe doesn't cover it: as of
+2026-10-05 the kernel advisories in its cache carry no fixed version, so it
+never flags a held kernel.
+
+I kept this out of v1 on purpose. It replaces the "vNext kernel-reboot
+item" in the build parent's original body.
+Spec: [[file:docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org][guarded-upgrade spec]], Scope tiers and Risks (Standing kernel
+deferral).
** TODO [#C] post-rebuild-check: probe that Emacs frames come up Wayland-native :feature:emacs:velox:solo:quick:
:PROPERTIES:
@@ -3529,6 +4152,1034 @@ minutes is the entry freeze: hold power for 10 s, and write down the cycle
number and whether the keyboard backlight was lit. A cycle that instead comes
straight back with "Cannot allocate memory" is the ARC task, not a freeze.
+*** M-split-console: an everyday run with no compositor defers only kernel-kind entries
+What we're verifying: with Hyprland not running, =upgrade-guarded
+--no-topgrade= lands the pending GPU/compositor packages on real pacman and
+leaves only kernel-kind entries deferred (AC3's held-set claim). It's the
+first test of the VM gate.
+- Make sure Phase 1's commits are in HEAD (=make test-keep= bundles
+ archsetup from HEAD).
+#+begin_src sh :results output
+# VM helpers; every later block in this test sources them. Root logs in by
+# key only after the install, so vmroot uses the key the harness left with
+# the newest test results.
+cat > /tmp/ug-vm.sh <<'EOF'
+cd ~/code/archsetup || exit 1
+o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222"
+key=$(ls -td test-results/*/root_key 2>/dev/null | head -1)
+vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; }
+vmroot() { ssh $o -i "$key" root@localhost "$@"; }
+mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; }
+shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1
+ magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; }
+shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d"
+ for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done
+ for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done
+ echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; }
+EOF
+echo "helpers written to /tmp/ug-vm.sh"
+#+end_src
+#+begin_src sh :results output
+# The full install runs first; this takes a while, and Emacs waits on it.
+cd ~/code/archsetup && make test-keep FS_PROFILE=zfs 2>&1 | tail -6
+#+end_src
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'pacman -Q wayland libdrm $(pacman -Qqo /usr/lib/modules/*/vmlinuz)'
+#+end_src
+- On https://archive.archlinux.org/packages/, find the release before the
+ installed one for =wayland= or =libdrm= (whichever downgrades without
+ breaking a dependency), and for the VM's kernel and its headers.
+- Paste the three URLs into the next block.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+gpu_url='PASTE-URL'; kernel_url='PASTE-URL'; headers_url='PASTE-URL'
+vmroot "pacman -U --noconfirm $gpu_url $kernel_url $headers_url" | tail -3
+vm 'LC_ALL=C pacman -Qu'
+#+end_src
+- Confirm Hyprland isn't running in the VM. If the block prints LIVE, stop
+ and record it under this test: the VM procedure's premise has failed.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'pgrep -x Hyprland >/dev/null && echo "Hyprland LIVE: the test premise fails" || echo "Hyprland not running"'
+#+end_src
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'upgrade-guarded --no-topgrade > /tmp/ug-console.log 2>&1; echo "exit: $?"
+tail -3 /tmp/ug-console.log
+jq -r ".data.packages[] | \"\(.kind) \(.name)\"" ~/.local/state/maint/upgrade_deferred.json
+echo "-- pacman -Qu:"; LC_ALL=C pacman -Qu'
+#+end_src
+Expected: Hyprland was not running; the run printed exit 0; the GPU-pattern
+package is back at its current version; every record line reads kernel;
+and =pacman -Qu= lists only the kernel and its headers, plus any
+=[ignored]= row.
+
+*** M-split-live on ratio: the live everyday run leaves the guard silent
+What we're verifying: the everyday split transaction on real pacman with
+Hyprland live, before any lever calls it. It should exit 0, print no
+BLOCKED banner, and defer exactly what =--dry-run= predicted (AC1). It's
+the last item of step (a) on ratio.
+- Run it in ratio's live Hyprland session, right after the rest of step
+ (a).
+#+begin_src sh :results output
+upgrade-guarded --dry-run > /tmp/ug-dry.txt 2>&1; echo "dry-run exit: $?"
+cat /tmp/ug-dry.txt
+#+end_src
+- Note the predicted deferred set, and whether it lists a GPU-kind entry.
+#+begin_src sh :results output
+# The real run: refresh, -Su with the held set ignored, then yay -Sua. It
+# takes minutes, and Emacs waits on it.
+upgrade-guarded --no-topgrade > /tmp/ug-live.log 2>&1; echo "exit: $?"
+echo "BLOCKED lines: $(grep -c BLOCKED /tmp/ug-live.log)"
+jq -r '.data.packages[] | "\(.kind) \(.name)"' ~/.local/state/maint/upgrade_deferred.json | sort
+echo "-- pacman -Qu:"; LC_ALL=C pacman -Qu
+#+end_src
+- If the dry-run listed no GPU-kind entry, leave this test open and repeat
+ it on the first day one is pending. The hook-silence half needs one.
+Expected: exit 0; zero BLOCKED lines; the record's names equal the
+dry-run's predicted deferred set; =pacman -Qu= lists only those names plus
+the =[ignored]= bridge-utils row; and on a run with a GPU-kind entry
+pending, that entry is in the record and the guard never printed its
+banner.
+
+*** M-split-live on velox: the live everyday run leaves the guard silent
+What we're verifying: the same everyday split transaction on velox's ZFS
+root with Hyprland live, before any lever calls it. It should exit 0, print
+no BLOCKED banner, and defer exactly what =--dry-run= predicted, with the
+kernel and zfs-dkms among the deferred entries when pending (AC1). It's the
+last item of step (a) on velox.
+- Run it in velox's live Hyprland session, right after the rest of step
+ (a).
+#+begin_src sh :results output
+uname -r
+upgrade-guarded --dry-run > /tmp/ug-dry.txt 2>&1; echo "dry-run exit: $?"
+cat /tmp/ug-dry.txt
+#+end_src
+- Note the predicted deferred set, and whether it lists a GPU-kind entry.
+#+begin_src sh :results output
+# The real run: refresh, -Su with the held set ignored, then yay -Sua. It
+# takes minutes, and Emacs waits on it.
+upgrade-guarded --no-topgrade > /tmp/ug-live.log 2>&1; echo "exit: $?"
+echo "BLOCKED lines: $(grep -c BLOCKED /tmp/ug-live.log)"
+jq -r '.data.packages[] | "\(.kind) \(.name)"' ~/.local/state/maint/upgrade_deferred.json | sort
+echo "-- pacman -Qu:"; LC_ALL=C pacman -Qu
+pacman -Q $(pacman -Qqo /usr/lib/modules/*/vmlinuz)
+#+end_src
+- If the dry-run listed no GPU-kind entry, leave this test open and repeat
+ it on the first day one is pending. The hook-silence half needs one.
+Expected: exit 0; zero BLOCKED lines; the record's names equal the
+dry-run's predicted deferred set; =pacman -Qu= lists only those names plus
+any =[ignored]= rows; the kernel package's version is unchanged; and on a
+run with a GPU-kind entry pending, that entry is in the record and the
+guard never printed its banner.
+
+*** M-boot-armed in the ZFS VM: the armed set lands from cache on tty1 before login, offline
+What we're verifying: the completion path on a ZFS root with no network at
+boot. =--complete= should pass the real gate (=sudo -n lsinitcpio=
+included) and arm without firing the guard or informant's hook. The boot
+unit should then apply the armed set before getty@tty1 starts (AC5, and
+AC4's gate on a real initramfs). Second test of the VM gate.
+- Make sure Phase 1's and Phase 3's commits are in HEAD.
+#+begin_src sh :results output
+# VM helpers; every later block in this test sources them. Root logs in by
+# key only after the install, so vmroot uses the key the harness left with
+# the newest test results.
+cat > /tmp/ug-vm.sh <<'EOF'
+cd ~/code/archsetup || exit 1
+o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222"
+key=$(ls -td test-results/*/root_key 2>/dev/null | head -1)
+vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; }
+vmroot() { ssh $o -i "$key" root@localhost "$@"; }
+mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; }
+shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1
+ magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; }
+shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d"
+ for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done
+ for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done
+ echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; }
+EOF
+echo "helpers written to /tmp/ug-vm.sh"
+#+end_src
+#+begin_src sh :results output
+# The full install runs first; this takes a while, and Emacs waits on it.
+cd ~/code/archsetup && make test-keep FS_PROFILE=zfs 2>&1 | tail -6
+#+end_src
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'pacman -Q wayland libdrm $(pacman -Qqo /usr/lib/modules/*/vmlinuz)'
+#+end_src
+- On https://archive.archlinux.org/packages/, find the release before the
+ installed one for =wayland= or =libdrm=, and for the VM's kernel and its
+ headers.
+- Paste the three URLs into the next block.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+gpu_url='PASTE-URL'; kernel_url='PASTE-URL'; headers_url='PASTE-URL'
+vmroot "pacman -U --noconfirm $gpu_url $kernel_url $headers_url" | tail -3
+vm 'LC_ALL=C pacman -Qu'
+#+end_src
+- Arm over SSH with the live branch forced, since Hyprland never runs in
+ the VM.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'UPGRADE_GUARDED_HYPR_RUNNING=1 upgrade-guarded --complete > /tmp/ug-arm.log 2>&1; echo "exit: $?"
+echo "pre-transaction hook runs: $(grep -c "Running pre-transaction hooks" /tmp/ug-arm.log)"
+echo "BLOCKED lines: $(grep -c BLOCKED /tmp/ug-arm.log)"
+echo "-- flag:"; cat /var/lib/archsetup/apply-upgrade-on-boot
+echo "-- record:"; jq -c ".data | {result, gate, held_snapshot, n: (.packages | length)}" ~/.local/state/maint/upgrade_deferred.json'
+#+end_src
+- Reboot the guest and take its network link down straight away.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'sudo systemctl reboot'; sleep 3; mon set_link net0 off >/dev/null; echo "link off"
+#+end_src
+- Capture tty1 for three minutes while it boots.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+shots 90
+#+end_src
+- Open the frames in order (=imv /tmp/ug-shots/=).
+- Bring the link back.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon set_link net0 on >/dev/null; sleep 20; echo "link on"
+#+end_src
+- Read the boot's outcome.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'journalctl -b -t archsetup-boot-upgrade --no-pager | head -2
+systemctl show archsetup-boot-upgrade -p Result -p ExecMainExitTimestampMonotonic
+systemctl show getty@tty1 -p ExecMainStartTimestampMonotonic
+ls /var/lib/archsetup/apply-upgrade-on-boot 2>&1
+jq -c ".data | {result, failed_step, n: (.packages | length)}" ~/.local/state/maint/upgrade_deferred.json
+echo "stamp age: $(( $(date +%s) - $(jq -r ".written_at | floor" ~/.local/state/maint/topgrade_run.json) )) s"
+echo "uptime: $(cut -d. -f1 /proc/uptime) s"'
+#+end_src
+Expected: the arm exited 0 with exactly one pre-transaction hook run (stage
+1's; the arm is forced live, so no stage 2 runs, and the arm's =-Sw= ran no
+hooks) and zero BLOCKED lines. The flag listed the GPU entries as
+=name=version= lines, and the record had gate null with held_snapshot
+naming a pre-pacman snapshot. A frame shows 'archsetup: applying N deferred
+GPU/compositor upgrades — do not power off' on tty1 before any frame shows
+the login, and the journal's first line is that banner. Result=success, and
+the unit's exit timestamp is below getty@tty1's start timestamp. Afterwards
+the flag is gone, the record reads result ok with n 0, and the stamp's age
+is below the uptime.
+
+*** M-boot-midtx in the ZFS VM: a timeout inside a real transaction leaves no db.lck
+What we're verifying: when the start timeout's SIGINT lands in the middle
+of a real transaction, pacman should stop at a package boundary and
+release =db.lck=. The record should name the interruption, and
+=--complete= from a console should finish the set (AC6). Third test of the
+VM gate, on the VM M-boot-armed leaves.
+- Use the VM M-boot-armed just left. Nothing reinstalls between the two
+ tests, so its journal still times a real boot transaction.
+#+begin_src sh :results output
+# VM helpers; every later block in this test sources them. Root logs in by
+# key only after the install, so vmroot uses the key the harness left with
+# the newest test results.
+cat > /tmp/ug-vm.sh <<'EOF'
+cd ~/code/archsetup || exit 1
+o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222"
+key=$(ls -td test-results/*/root_key 2>/dev/null | head -1)
+vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; }
+vmroot() { ssh $o -i "$key" root@localhost "$@"; }
+mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; }
+shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1
+ magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; }
+shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d"
+ for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done
+ for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done
+ echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; }
+EOF
+echo "helpers written to /tmp/ug-vm.sh"
+#+end_src
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'journalctl -t archsetup-boot-upgrade -o short-monotonic --no-pager | grep -E "applying|upgrading|installing|transaction" | tail -12'
+#+end_src
+- On https://archive.archlinux.org/packages/, find the previous releases of
+ two or more GPU-pattern packages, mesa among them, so the transaction is
+ long enough to cut.
+- Paste their URLs into the next block.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+gpu_urls='PASTE-URL PASTE-URL'
+vmroot "pacman -U --noconfirm $gpu_urls" | tail -2; vm 'LC_ALL=C pacman -Qu'
+#+end_src
+- Arm over SSH with the live branch forced.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'UPGRADE_GUARDED_HYPR_RUNNING=1 upgrade-guarded --complete > /tmp/ug-arm.log 2>&1; echo "exit: $?"
+cat /var/lib/archsetup/apply-upgrade-on-boot'
+#+end_src
+- Work out a start timeout that lands inside the transaction: the seconds
+ from the banner to about halfway through the 'upgrading' lines above,
+ scaled up for the larger set.
+- Put it in the next block.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+secs=PASTE-SECONDS
+vmroot "install -d /etc/systemd/system/archsetup-boot-upgrade.service.d
+printf '[Service]\nTimeoutStartSec=%ss\n' $secs > /etc/systemd/system/archsetup-boot-upgrade.service.d/test.conf
+systemctl daemon-reload && cat /etc/systemd/system/archsetup-boot-upgrade.service.d/test.conf"
+#+end_src
+- Reboot the guest.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'sudo systemctl reboot'; echo rebooting
+#+end_src
+- Capture tty1 for three minutes.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+shots 90
+#+end_src
+- Open the frames in order (=imv /tmp/ug-shots/=).
+- Read the outcome once the login shows.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'ls /var/lib/pacman/db.lck 2>&1
+jq -c ".data | {result, failed_step, detail, n: (.packages | length)}" ~/.local/state/maint/upgrade_deferred.json
+journalctl -b -t archsetup-boot-upgrade --no-pager | tail -4'
+#+end_src
+- If the journal shows the transaction finished, or never started, change
+ the seconds and repeat from the downgrade block.
+- Confirm Hyprland isn't running, so the next =--complete= applies the set
+ rather than arming it. If the block prints LIVE, record it under this
+ test.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'pgrep -x Hyprland >/dev/null && echo "Hyprland LIVE: the test premise fails" || echo "Hyprland not running"'
+#+end_src
+- Finish the set from a console.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'upgrade-guarded --complete > /tmp/ug-finish.log 2>&1; echo "exit: $?"
+LC_ALL=C pacman -Qu
+jq -c ".data | {result, n: (.packages | length)}" ~/.local/state/maint/upgrade_deferred.json'
+#+end_src
+- Remove the drop-in.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vmroot 'rm -rf /etc/systemd/system/archsetup-boot-upgrade.service.d && systemctl daemon-reload && echo removed'
+#+end_src
+Expected: no =db.lck= after the cut-off boot, and the journal stopped
+between packages. The record read result interrupted with failed_step
+boot-transaction and the =--complete= remedy. Then the console
+=--complete= exited 0, =pacman -Qu= no longer lists the armed packages, and
+the record reads result ok with n 0.
+
+*** M-boot-timeout in the ZFS VM: a hung boot run times out to the login and disarms
+What we're verifying: a boot transaction that never ends is cut off by the
+start timeout (shortened here to 90 s). The login should still appear, the
+flag should be gone, the failure should be visible in the record and in
+maint, and the next boot should skip the unit (AC6). Fourth test of the VM
+gate.
+#+begin_src sh :results output
+# VM helpers; every later block in this test sources them. Root logs in by
+# key only after the install, so vmroot uses the key the harness left with
+# the newest test results.
+cat > /tmp/ug-vm.sh <<'EOF'
+cd ~/code/archsetup || exit 1
+o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222"
+key=$(ls -td test-results/*/root_key 2>/dev/null | head -1)
+vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; }
+vmroot() { ssh $o -i "$key" root@localhost "$@"; }
+mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; }
+shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1
+ magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; }
+shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d"
+ for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done
+ for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done
+ echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; }
+EOF
+echo "helpers written to /tmp/ug-vm.sh"
+#+end_src
+- Run the next block for a fresh VM, or skip it to reuse the VM
+ M-boot-midtx left (the block reinstalls from clean).
+#+begin_src sh :results output
+# The full install runs first; this takes a while, and Emacs waits on it.
+cd ~/code/archsetup && make test-keep FS_PROFILE=zfs 2>&1 | tail -6
+#+end_src
+- On https://archive.archlinux.org/packages/, find the release before the
+ installed =wayland= or =libdrm=.
+- Paste its URL into the next block.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+gpu_url='PASTE-URL'
+vmroot "pacman -U --noconfirm $gpu_url" | tail -2; vm 'LC_ALL=C pacman -Qu'
+#+end_src
+- Arm over SSH with the live branch forced.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'UPGRADE_GUARDED_HYPR_RUNNING=1 upgrade-guarded --complete > /tmp/ug-arm.log 2>&1; echo "exit: $?"
+cat /var/lib/archsetup/apply-upgrade-on-boot'
+#+end_src
+- Install the stand-in fakes and the drop-in. The fake pacman hands every
+ query to the real one and sleeps on the transaction, and the fake sudo
+ drops =-n=. They do the job the Phase 1 fakes do.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vmroot 'set -e
+d=/usr/local/lib/ug-fakes
+install -d "$d" /etc/systemd/system/archsetup-boot-upgrade.service.d
+cat > "$d/sudo" <<"EOF"
+#!/bin/sh
+# Drop -n and run the command as the caller.
+[ "$1" = -n ] && shift
+exec "$@"
+EOF
+cat > "$d/pacman" <<"EOF"
+#!/bin/sh
+# Real pacman for every query; the transaction (-S) sleeps or fails.
+if [ "$1" = -S ]; then
+ [ "$UG_FAKE_MODE" = fail ] && { echo "error: failed to commit transaction (fake)" >&2; exit 1; }
+ exec sleep 3600
+fi
+exec /usr/bin/pacman "$@"
+EOF
+chmod 755 "$d/sudo" "$d/pacman"
+cat > /etc/systemd/system/archsetup-boot-upgrade.service.d/test.conf <<"EOF"
+[Service]
+TimeoutStartSec=90s
+Environment=PATH=/usr/local/lib/ug-fakes:/usr/local/bin:/usr/bin
+Environment=UG_FAKE_MODE=sleep
+EOF
+systemctl daemon-reload && echo installed'
+#+end_src
+- Reboot the guest.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'sudo systemctl reboot'; echo rebooting
+#+end_src
+- Capture tty1 for four minutes.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+shots 120
+#+end_src
+- Open the frames in order (=imv /tmp/ug-shots/=).
+- Read the outcome once the login shows.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'ls /var/lib/archsetup/apply-upgrade-on-boot 2>&1
+systemctl is-failed archsetup-boot-upgrade.service
+systemctl show archsetup-boot-upgrade -p ExecMainStartTimestampMonotonic
+systemctl show getty@tty1 -p ExecMainStartTimestampMonotonic
+jq -c ".data | {result, failed_step, detail, n: (.packages | length)}" ~/.local/state/maint/upgrade_deferred.json
+ls /var/lib/pacman/db.lck 2>&1
+~/.local/bin/maint status 2>&1 | grep -i -E "failed|deferred"'
+#+end_src
+- Reboot the guest once more.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'sudo systemctl reboot'; echo rebooting
+#+end_src
+- About a minute later, read whether the unit ran.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'systemctl show archsetup-boot-upgrade -p ConditionResult -p ActiveState'
+#+end_src
+- Remove the drop-in and the fakes.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vmroot 'rm -rf /etc/systemd/system/archsetup-boot-upgrade.service.d /usr/local/lib/ug-fakes && systemctl daemon-reload && echo removed'
+#+end_src
+Expected: the login appeared about 90 s into the unit: getty@tty1's start
+timestamp is about 90,000,000 µs above the unit's
+ExecMainStartTimestampMonotonic. After it, the flag was gone, =is-failed=
+printed failed, and no =db.lck= existed. The record read result
+interrupted with failed_step boot-transaction, a detail naming
+=upgrade-guarded --complete=, and n equal to the armed count. maint's
+failed-units row names archsetup-boot-upgrade.service, and the deferred row
+lists the armed set at WARN. The next boot shows ConditionResult=no.
+
+*** M-boot-fail in the ZFS VM: a failed boot transaction still reaches the login and disarms
+What we're verifying: a boot transaction that fails outright should still
+end at the tty1 login. The flag should be gone, the record should name the
+failed step, maint should show the failed unit and the deferred set, and
+the unit should be left failed (AC6). Last test of the VM gate.
+#+begin_src sh :results output
+# VM helpers; every later block in this test sources them. Root logs in by
+# key only after the install, so vmroot uses the key the harness left with
+# the newest test results.
+cat > /tmp/ug-vm.sh <<'EOF'
+cd ~/code/archsetup || exit 1
+o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222"
+key=$(ls -td test-results/*/root_key 2>/dev/null | head -1)
+vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; }
+vmroot() { ssh $o -i "$key" root@localhost "$@"; }
+mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; }
+shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1
+ magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; }
+shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d"
+ for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done
+ for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done
+ echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; }
+EOF
+echo "helpers written to /tmp/ug-vm.sh"
+#+end_src
+- Run the next block for a fresh VM, or skip it to reuse the VM the
+ previous M-boot test left (the block reinstalls from clean).
+#+begin_src sh :results output
+# The full install runs first; this takes a while, and Emacs waits on it.
+cd ~/code/archsetup && make test-keep FS_PROFILE=zfs 2>&1 | tail -6
+#+end_src
+- On https://archive.archlinux.org/packages/, find the release before the
+ installed =wayland= or =libdrm=.
+- Paste its URL into the next block.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+gpu_url='PASTE-URL'
+vmroot "pacman -U --noconfirm $gpu_url" | tail -2; vm 'LC_ALL=C pacman -Qu'
+#+end_src
+- Arm over SSH with the live branch forced.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'UPGRADE_GUARDED_HYPR_RUNNING=1 upgrade-guarded --complete > /tmp/ug-arm.log 2>&1; echo "exit: $?"
+cat /var/lib/archsetup/apply-upgrade-on-boot'
+#+end_src
+- Install the stand-in fakes and the drop-in, with the fake pacman set to
+ answer =-Q= and =-Sp= normally and exit 1 on =-S=.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vmroot 'set -e
+d=/usr/local/lib/ug-fakes
+install -d "$d" /etc/systemd/system/archsetup-boot-upgrade.service.d
+cat > "$d/sudo" <<"EOF"
+#!/bin/sh
+# Drop -n and run the command as the caller.
+[ "$1" = -n ] && shift
+exec "$@"
+EOF
+cat > "$d/pacman" <<"EOF"
+#!/bin/sh
+# Real pacman for every query; the transaction (-S) sleeps or fails.
+if [ "$1" = -S ]; then
+ [ "$UG_FAKE_MODE" = fail ] && { echo "error: failed to commit transaction (fake)" >&2; exit 1; }
+ exec sleep 3600
+fi
+exec /usr/bin/pacman "$@"
+EOF
+chmod 755 "$d/sudo" "$d/pacman"
+cat > /etc/systemd/system/archsetup-boot-upgrade.service.d/test.conf <<"EOF"
+[Service]
+TimeoutStartSec=90s
+Environment=PATH=/usr/local/lib/ug-fakes:/usr/local/bin:/usr/bin
+Environment=UG_FAKE_MODE=fail
+EOF
+systemctl daemon-reload && echo installed'
+#+end_src
+- Reboot the guest.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'sudo systemctl reboot'; echo rebooting
+#+end_src
+- Capture tty1 for two minutes.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+shots 60
+#+end_src
+- Open the frames in order (=imv /tmp/ug-shots/=).
+- Read the outcome once the login shows.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'ls /var/lib/archsetup/apply-upgrade-on-boot 2>&1
+systemctl is-failed archsetup-boot-upgrade.service
+jq -c ".data | {result, failed_step, detail, n: (.packages | length)}" ~/.local/state/maint/upgrade_deferred.json
+~/.local/bin/maint status 2>&1 | grep -i -E "failed|deferred"'
+#+end_src
+- Remove the drop-in and the fakes.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vmroot 'rm -rf /etc/systemd/system/archsetup-boot-upgrade.service.d /usr/local/lib/ug-fakes && systemctl daemon-reload && echo removed'
+#+end_src
+Expected: the login appeared without waiting out the timeout, the flag is
+gone, and =is-failed= prints failed. The record reads result failed with
+failed_step boot-transaction, a detail whose last line is the fake's
+'error: failed to commit transaction (fake)', and n equal to the armed
+count. maint's failed-units row names archsetup-boot-upgrade.service, and
+the deferred row lists the armed set at WARN.
+
+*** M-boot-armed on ratio: APPLY arms, the row shows it, and the offline reboot applies the set
+What we're verifying: the panel path the VM can't cover. APPLY should open
+=upgrade-guarded --complete= in a detached foot terminal, the row should
+read armed with REBOOT shown, and an offline armed boot should apply the
+set before the login (AC4's APPLY clause, AC5).
+- Run it after step (c) on ratio, on a day the dry-run below lists a
+ GPU-kind entry.
+#+begin_src sh :results output
+upgrade-guarded --dry-run 2>&1 | tail -20
+#+end_src
+- Open the maint panel.
+- Find the deferred row.
+- Press APPLY on the deferred row. The first press arms.
+- Press APPLY again to fire it.
+- Watch the panel's wall and the foot terminal that opens, until the
+ terminal shows its 'Reboot now?' prompt.
+- Press Enter at the prompt, which answers no.
+- Wait up to 30 s for the panel's next probe.
+- Read the deferred row and the REBOOT key.
+- Read the flag.
+#+begin_src sh :results output
+cat /var/lib/archsetup/apply-upgrade-on-boot
+#+end_src
+- Take networking down.
+#+begin_src sh :results output
+nmcli networking off && echo "networking off"
+#+end_src
+- Press REBOOT on the panel.
+- Confirm the reboot the way the panel asks.
+- Watch the monitor through the boot until the tty1 password prompt.
+- Log in.
+- Bring networking back.
+#+begin_src sh :results output
+nmcli networking on && echo "networking on"
+#+end_src
+#+begin_src sh :results output
+ls /var/lib/archsetup/apply-upgrade-on-boot 2>&1
+jq -c '.data | {result, failed_step, n: (.packages | length)}' ~/.local/state/maint/upgrade_deferred.json
+maint status 2>&1 | grep -i -E 'topgrade|deferred'
+#+end_src
+Expected: the second APPLY press opened a foot terminal running
+=upgrade-guarded --complete=, and the panel's wall streamed nothing. The
+terminal ended at the prompt with no BLOCKED banner. Before the reboot, the
+row read 'armed — A apply at next boot' (with ' · <N−A> more deferred' if
+anything else stayed deferred) and REBOOT was shown. The banner showed on
+the monitor before the password prompt. Afterwards the flag is gone, the
+record reads result ok with n 0, and =maint status= shows topgrade_age
+fresh when nothing else is deferred.
+
+*** M-boot-ratio: the banner and progress show on ratio's monitor, and REBOOT shows while armed
+What we're verifying: ratio's own boot path. Its =/dev/console= is ttyS0,
+so the unit's tty1 output has to reach the monitor. Its tty1 shows a
+password prompt that the =Before=getty@tty1.service= ordering holds. And
+REBOOT shows from the flag alone, since reboot_required stays false there
+(AC5, and Phase 4's ratio-path confirmation).
+- Run it on the boot M-boot-armed on ratio arms, doing the steps up to the
+ REBOOT press before that test's own REBOOT press; or arm the same way.
+#+begin_src sh :results output
+ls -l /var/lib/archsetup/apply-upgrade-on-boot
+maint status 2>&1 | grep -i reboot
+#+end_src
+- Read the panel's REBOOT key.
+- Press REBOOT on the panel.
+- Confirm the reboot the way the panel asks.
+- Watch the monitor from the firmware screen to the tty1 password prompt.
+- Log in.
+#+begin_src sh :results output
+systemctl show archsetup-boot-upgrade -p ExecMainExitTimestampMonotonic
+systemctl show getty@tty1 -p ExecMainStartTimestampMonotonic
+#+end_src
+Expected: before the reboot, reboot_required read clear and REBOOT was
+shown anyway. The banner and pacman's per-package progress stayed on the
+monitor for the whole transaction, and the password prompt appeared only
+after it. The unit's exit timestamp is below getty@tty1's start timestamp.
+
+*** ratio's linux-lts-strix is never moved or gated
+What we're verifying: the foreign, headerless GRUB-default kernel stays out
+of the pending set and the gate entry on a real =--complete= (Phase 4's
+ratio-path confirmation).
+- Run it on ratio the first time a =--complete= lands linux or linux-lts.
+#+begin_src sh :results output
+pacman -Q linux-lts-strix | tee /tmp/ug-strix-before.txt
+LC_ALL=C pacman -Qu | grep strix || echo "strix not pending"
+#+end_src
+- Start =upgrade-guarded --complete= with APPLY, or in a terminal.
+- While stage 1 runs, read the gate entry from a second terminal.
+#+begin_src sh :results output
+jq -c '.data.gate' ~/.local/state/maint/upgrade_deferred.json
+#+end_src
+- Let the session finish.
+#+begin_src sh :results output
+pacman -Q linux-lts-strix | diff /tmp/ug-strix-before.txt - && echo "strix unchanged"
+grep "$(date +%Y-%m-%d)" /var/log/pacman.log | grep linux-lts-strix || echo "no strix lines today"
+#+end_src
+Expected: strix wasn't pending; gate.pkgbases during stage 1 listed linux
+and/or linux-lts and never linux-lts-strix; strix's version is unchanged;
+and no pacman.log line from today names it.
+
+*** M-boot-armed on velox: APPLY arms, the row shows it, and the offline reboot applies the set
+What we're verifying: the same panel path on velox's ZFS root with
+autologin. APPLY's detached terminal should run the real gate when a
+kernel is pending, the row should read armed with REBOOT shown, and an
+offline armed boot should finish before autologin starts Hyprland (AC4's
+APPLY clause, AC5).
+- Run it after step (c) on velox, on a day the dry-run below lists a
+ GPU-kind entry.
+#+begin_src sh :results output
+upgrade-guarded --dry-run 2>&1 | tail -20
+#+end_src
+- Open the maint panel.
+- Find the deferred row.
+- Press APPLY on the deferred row. The first press arms.
+- Press APPLY again to fire it.
+- Watch the panel's wall and the foot terminal that opens, until the
+ terminal shows its 'Reboot now?' prompt.
+- Press Enter at the prompt, which answers no.
+- Wait up to 30 s for the panel's next probe.
+- Read the deferred row and the REBOOT key.
+- Read the flag.
+#+begin_src sh :results output
+cat /var/lib/archsetup/apply-upgrade-on-boot
+#+end_src
+- Take networking down.
+#+begin_src sh :results output
+nmcli networking off && echo "networking off"
+#+end_src
+- Press REBOOT on the panel.
+- Confirm the reboot the way the panel asks.
+- Watch tty1 from the ZFSBootMenu countdown until autologin starts
+ Hyprland.
+- Bring networking back.
+#+begin_src sh :results output
+nmcli networking on && echo "networking on"
+#+end_src
+#+begin_src sh :results output
+ls /var/lib/archsetup/apply-upgrade-on-boot 2>&1
+jq -c '.data | {result, failed_step, gate, n: (.packages | length)}' ~/.local/state/maint/upgrade_deferred.json
+journalctl -b -t archsetup-boot-upgrade --no-pager | head -2
+maint status 2>&1 | grep -i -E 'topgrade|deferred'
+#+end_src
+Expected: the second APPLY press opened a foot terminal running
+=upgrade-guarded --complete=, and the panel's wall streamed nothing. If a
+kernel was pending, the terminal showed the gate passing before the arm.
+Before the reboot the row read 'armed — A apply at next boot' and REBOOT
+was shown. The banner showed on tty1 before Hyprland started, and the
+journal's first line is that banner. Afterwards the flag is gone, the
+record reads result ok with gate null and n 0, and =maint status= shows
+topgrade_age fresh when nothing else is deferred.
+
+*** M-boot-news on velox: unread Arch news doesn't stall the boot run
+What we're verifying: with two or more unread news items at boot, the boot
+form's informant clear should keep the transaction from waiting on stdin
+(AC7 at boot).
+- Run it after step (c) on velox, on a day =upgrade-guarded --dry-run=
+ lists a GPU-kind entry.
+- Open the maint panel at its deferred row.
+- Press APPLY on the deferred row. The first press arms.
+- Press APPLY again to fire it.
+- Press Enter at the terminal's 'Reboot now?' prompt, which answers no.
+ Arming cleared the news, so the news has to become unread after this.
+#+begin_src sh :results output
+informant list --unread 2>&1 | head
+#+end_src
+- If fewer than two items are unread, move informant's state aside so the
+ whole feed reads unread. The last step puts it back, and no pacman
+ transaction should run before the reboot.
+#+begin_src sh :results output
+sudo mv /var/lib/informant.dat /var/lib/informant.dat.bak && informant list --unread 2>&1 | head -3
+#+end_src
+- Press REBOOT on the panel.
+- Confirm the reboot the way the panel asks.
+- Watch tty1 until autologin starts Hyprland.
+#+begin_src sh :results output
+systemctl show archsetup-boot-upgrade -p Result -p ExecMainStatus
+journalctl -b -t archsetup-boot-upgrade --no-pager | tail -5
+jq -c '.data | {result, failed_step, n: (.packages | length)}' ~/.local/state/maint/upgrade_deferred.json
+#+end_src
+- If you moved informant's state aside, restore it.
+#+begin_src sh :results output
+[ -e /var/lib/informant.dat.bak ] && sudo mv -f /var/lib/informant.dat.bak /var/lib/informant.dat && echo restored
+#+end_src
+Expected: the boot run ended with Result=success well inside its timeout,
+the journal runs straight from the banner to the end of the transaction
+with nothing waiting on input, and the record reads result ok.
+
+*** D-zbm: recover from a failed zfs-dkms build by rolling back to the held snapshot
+What we're verifying: the README's recovery steps work end to end on a ZFS
+root. A =--complete= whose zfs build fails should exit 4 with its snapshot
+held and named. A ZFSBootMenu rollback (MOD+R) plus the pacman-log
+reconcile should then leave a consistent system. This is non-gating; no
+criterion waits on it.
+- Follow the recovery steps in =maint/README.md= (Phase 4), which carry the
+ first build task's MOD+R and version fixes.
+#+begin_src sh :results output
+# VM helpers; every later block in this test sources them. Root logs in by
+# key only after the install, so vmroot uses the key the harness left with
+# the newest test results.
+cat > /tmp/ug-vm.sh <<'EOF'
+cd ~/code/archsetup || exit 1
+o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222"
+key=$(ls -td test-results/*/root_key 2>/dev/null | head -1)
+vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; }
+vmroot() { ssh $o -i "$key" root@localhost "$@"; }
+mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; }
+shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1
+ magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; }
+shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d"
+ for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done
+ for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done
+ echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; }
+EOF
+echo "helpers written to /tmp/ug-vm.sh"
+#+end_src
+#+begin_src sh :results output
+# The full install runs first; this takes a while, and Emacs waits on it.
+cd ~/code/archsetup && make test-keep FS_PROFILE=zfs 2>&1 | tail -6
+#+end_src
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'pacman -Q less $(pacman -Qqo /usr/lib/modules/*/vmlinuz)'
+#+end_src
+- On https://archive.archlinux.org/packages/, find the previous releases of
+ the VM's kernel, its headers and one small leaf package (=less=).
+- Paste the three URLs into the next block.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+kernel_url='PASTE-URL'; headers_url='PASTE-URL'; leaf_url='PASTE-URL'
+vmroot "pacman -U --noconfirm $kernel_url $headers_url $leaf_url" | tail -3
+vm 'LC_ALL=C pacman -Qu'
+#+end_src
+- Make the zfs build fail for any kernel built from now on. The VM is
+ disposable after this drill.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vmroot 'for f in /usr/src/zfs-*/dkms.conf; do echo "MAKE[0]=\"false\"" >> "$f"; tail -1 "$f"; done'
+#+end_src
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'upgrade-guarded --complete > /tmp/ug-zbm.log 2>&1; echo "exit: $?"; tail -2 /tmp/ug-zbm.log
+ls /var/lib/archsetup/apply-upgrade-on-boot 2>&1
+snap=$(jq -r .data.held_snapshot ~/.local/state/maint/upgrade_deferred.json); echo "held: $snap"
+zfs holds "$snap"'
+#+end_src
+- Write down the held snapshot's name; a rolled-back root may not keep it
+ in the record.
+- Reboot the guest and catch ZFSBootMenu inside its 3 s countdown by
+ sending Escape every half second (the VM takes no keyboard input any
+ other way).
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'sudo systemctl reboot'; sleep 5
+for i in $(seq 60); do mon sendkey esc >/dev/null; sleep 0.5; done
+shot
+#+end_src
+- If the shot shows anything but the boot-environment list, change the
+ =sleep= before the loop and run the block again; if it shows the
+ firmware's setup screen, which also answers Escape, run
+ =mon system_reset= first.
+- Open the snapshot list with MOD+S (Ctrl+S).
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey ctrl-s >/dev/null; sleep 1; shot
+#+end_src
+- If ZFSBootMenu asks to import the pool read/write, press MOD+W.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey ctrl-w >/dev/null; sleep 1; shot
+#+end_src
+- Move the selection onto the held snapshot one row at a time, checking
+ each shot.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey down >/dev/null; sleep 1; shot
+#+end_src
+- Roll back with MOD+R. Never press ENTER (duplicate), MOD+X (clone and
+ promote) or MOD+C (clone) here.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey ctrl-r >/dev/null; sleep 1; shot
+#+end_src
+- Answer the rollback confirmation the screen shows; for a y/N prompt, the
+ next block sends y.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey y >/dev/null; sleep 2; shot
+#+end_src
+- Go back to the boot-environment list.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey esc >/dev/null; sleep 1; shot
+#+end_src
+- Boot =zroot/ROOT/default=.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey ret >/dev/null; sleep 40; shot
+#+end_src
+- List the pacman.log lines to undo, newest first, with the held name
+ pasted in.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+snap='PASTE-HELD-SNAPSHOT'
+vmroot "ls /var/lib/pacman/db.lck 2>&1
+t=\$(date -d @\$(zfs get -Hp -o value creation $snap) +%Y-%m-%dT%H:%M:%S); echo \"since \$t\"
+awk -v t=\"\$t\" 'substr(\$1, 2, 19) >= t && \$2 == \"[ALPM]\" && \$3 ~ /^(upgraded|downgraded|installed|removed)\$/' /var/log/pacman.log | tac"
+#+end_src
+- Undo each listed line, newest first, as the README's recovery steps say.
+- Paste the reconciled names into the check block and run it.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+names='PASTE-RECONCILED-NAMES'
+vmroot "pacman -Q $names; pacman -Dk && echo 'Dk clean'; pacman -Qkk $names 2>&1 | tail -4"
+#+end_src
+Expected: =--complete= exited 4 with the zfs module item as its last line,
+left no flag, and named a held snapshot that =zfs holds= shows with the
+=upgrade-guarded= tag. After the rollback and the reconcile, =pacman -Q=
+shows the old kernel, headers and =less= versions, =pacman -Dk= is clean,
+and =pacman -Qkk= over the reconciled names reports no mismatch.
+
+*** D-zbm power-cut variant: a reset during stage 1 leaves the right snapshot held and a reconcilable db
+What we're verifying: a hard reset during =--complete='s stage 1, on a VM
+with no open gate, should leave the snapshot step 3 took held and named.
+Recovery's power-cut steps, including the half-registered package, should
+then reconcile the db. This is non-gating.
+- Use a fresh =make test-keep FS_PROFILE=zfs= VM, not the one D-zbm left
+ with its gate open.
+#+begin_src sh :results output
+# VM helpers; every later block in this test sources them. Root logs in by
+# key only after the install, so vmroot uses the key the harness left with
+# the newest test results.
+cat > /tmp/ug-vm.sh <<'EOF'
+cd ~/code/archsetup || exit 1
+o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222"
+key=$(ls -td test-results/*/root_key 2>/dev/null | head -1)
+vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; }
+vmroot() { ssh $o -i "$key" root@localhost "$@"; }
+mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; }
+shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1
+ magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; }
+shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d"
+ for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done
+ for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done
+ echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; }
+EOF
+echo "helpers written to /tmp/ug-vm.sh"
+#+end_src
+#+begin_src sh :results output
+# The full install runs first; this takes a while, and Emacs waits on it.
+cd ~/code/archsetup && make test-keep FS_PROFILE=zfs 2>&1 | tail -6
+#+end_src
+- On https://archive.archlinux.org/packages/, find the previous releases of
+ the VM's kernel and its headers.
+- Paste the two URLs into the next block.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+kernel_url='PASTE-URL'; headers_url='PASTE-URL'
+vmroot "pacman -U --noconfirm $kernel_url $headers_url" | tail -3
+vm 'LC_ALL=C pacman -Qu'
+#+end_src
+- In a separate terminal, start the session with a tty. pacman prints
+ '(n/m) upgrading <pkg>' only on a tty; without one it prints 'upgrading
+ <pkg>...'.
+#+begin_src sh :eval no
+sshpass -p archsetup ssh -t -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -p 2222 cjennings@localhost upgrade-guarded --complete
+#+end_src
+- As soon as stage 1's first upgrading line appears, read the held
+ snapshot's name.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vm 'jq -r .data.held_snapshot ~/.local/state/maint/upgrade_deferred.json'
+#+end_src
+- As the kernel headers package's upgrading line appears, run the next
+ block. It hard-resets the guest, then sends Escape every half second to
+ catch ZFSBootMenu's 3 s countdown. Never use =system_powerdown= or
+ =quit=: both end QEMU, and =make test-keep= then restores the clean
+ install.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon system_reset >/dev/null; sleep 3
+for i in $(seq 60); do mon sendkey esc >/dev/null; sleep 0.5; done
+shot
+#+end_src
+- If the shot shows anything but the boot-environment list, run
+ =mon system_reset= and the Escape loop again with a changed delay.
+- Open the snapshot list with MOD+S (Ctrl+S).
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey ctrl-s >/dev/null; sleep 1; shot
+#+end_src
+- If ZFSBootMenu asks to import the pool read/write, press MOD+W.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey ctrl-w >/dev/null; sleep 1; shot
+#+end_src
+- Move the selection onto the held snapshot one row at a time, checking
+ each shot. If its name wasn't read in time, it's the
+ =zroot/ROOT/default@pre-pacman_*= row taken just before stage 1.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey down >/dev/null; sleep 1; shot
+#+end_src
+- Roll back with MOD+R. Never press ENTER (duplicate), MOD+X (clone and
+ promote) or MOD+C (clone) here.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey ctrl-r >/dev/null; sleep 1; shot
+#+end_src
+- Answer the rollback confirmation the screen shows; for a y/N prompt, the
+ next block sends y.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey y >/dev/null; sleep 2; shot
+#+end_src
+- Go back to the boot-environment list.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey esc >/dev/null; sleep 1; shot
+#+end_src
+- Boot =zroot/ROOT/default=.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+mon sendkey ret >/dev/null; sleep 40; shot
+#+end_src
+- Read the hold, the lock and the db's state, with the held name pasted in.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+snap='PASTE-HELD-SNAPSHOT'
+vmroot "zfs holds $snap; ls /var/lib/pacman/db.lck 2>&1; pacman -Dk 2>&1 | tail -3; tail -8 /var/log/pacman.log"
+#+end_src
+- If pacman.log shows the reset landed outside extraction (every target
+ has an =[ALPM]= line, or 'running post-transaction hooks' is logged),
+ stop and repeat on a fresh VM.
+- Remove =db.lck= if the read above showed it.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+vmroot 'rm -fv /var/lib/pacman/db.lck'
+#+end_src
+- Delete the local db entry =pacman -Dk= reported as 'description file is
+ missing', with its directory name pasted in.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+entry='PASTE-ENTRY-DIR'
+vmroot "rm -rv /var/lib/pacman/local/$entry"
+#+end_src
+- Find the version the booted root holds for that package: the last
+ =[ALPM]= line for it before the snapshot's creation (on an upgraded or
+ downgraded line, the version after =->=; on a removed line, or with no
+ line, register nothing and skip the next step).
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+snap='PASTE-HELD-SNAPSHOT'; pkg='PASTE-PACKAGE'
+vmroot "t=\$(date -d @\$(zfs get -Hp -o value creation $snap) +%Y-%m-%dT%H:%M:%S)
+awk -v t=\"\$t\" -v p=$pkg 'substr(\$1, 2, 19) < t && \$2 == \"[ALPM]\" && \$4 == p' /var/log/pacman.log | tail -1"
+#+end_src
+- Register that version from the cache, with its package file pasted in.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+pkgfile='PASTE-CACHE-FILE'
+vmroot "pacman -U --dbonly --noconfirm /var/cache/pacman/pkg/$pkgfile"
+#+end_src
+- List the later =[ALPM]= lines to undo, newest first.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+snap='PASTE-HELD-SNAPSHOT'
+vmroot "t=\$(date -d @\$(zfs get -Hp -o value creation $snap) +%Y-%m-%dT%H:%M:%S); echo \"since \$t\"
+awk -v t=\"\$t\" 'substr(\$1, 2, 19) >= t && \$2 == \"[ALPM]\" && \$3 ~ /^(upgraded|downgraded|installed|removed)\$/' /var/log/pacman.log | tac"
+#+end_src
+- Undo each listed line, newest first, as the README's recovery steps say.
+- Paste the reconciled names into the check block and run it.
+#+begin_src sh :results output
+. /tmp/ug-vm.sh
+names='PASTE-RECONCILED-NAMES'
+vmroot "pacman -Q $names; pacman -Dk && echo 'Dk clean'; pacman -Qkk $names 2>&1 | tail -4"
+#+end_src
+Expected: =zfs holds= showed the =upgrade-guarded= tag on the snapshot step
+3 took before stage 1, and held_snapshot named it. After the reconcile,
+=pacman -Q= shows the old kernel set, =pacman -Dk= is clean, and =pacman
+-Qkk= over the reconciled names reports no mismatch, including the package
+cut off mid-extraction.
+
** DOING [#B] Prepare for GitHub open-source release
:PROPERTIES:
:LAST_REVIEWED: 2026-08-17