diff options
Diffstat (limited to 'todo.org')
| -rw-r--r-- | todo.org | 1655 |
1 files changed, 1653 insertions, 2 deletions
@@ -711,7 +711,7 @@ Grading: feature, no hard date, real improvement to the install = [#B]. :solo: — build path (installer + checker + tests) and verify path (pytest; the live proof already exists on velox) with no open decision. -** TODO [#A] Topgrade guarded-upgrade spec — decisions, review, decomposition :feature:maint:dotfiles: +** TODO [#A] Topgrade guarded-upgrade build :feature:maint:dotfiles: SCHEDULED: <2026-09-23 Wed> :PROPERTIES: :CREATED: [2026-08-25 Tue] @@ -725,7 +725,7 @@ xorg-xwayland under a live Hyprland) plus any failing ecosystem step makes that exit almost unreachable. Diagnosed 2026-08-24/25; the fix is specced, not hacked, because it spans two repos and the design is contested. -Spec: [[file:docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org][2026-08-25-topgrade-guarded-upgrade-spec.org]] (DRAFT). +Spec: [[file:docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org][2026-08-25-topgrade-guarded-upgrade-spec.org]] (DOING). Four open decisions, all mine to make before the spec can move: 1. Freshness means *state* (a guarded upgrade still un-applied stays stale), @@ -763,6 +763,629 @@ DRAFT → READY, and spec-response into build tasks. velox has =linux-lts 6.18.51 → 6.18.52= pending today, which is the kernel-hold case this spec exists for, so no plain topgrade on velox until the hold is built or the kernel update runs as its own session. +*** 2026-10-05 Mon @ 06:08:00 -0500 Spec-review ran: Not ready, 13 blocking findings +Before the review I named the script =upgrade-guarded= and added =containers= to +its topgrade =--disable= list. The review recorded 48 findings in the spec's +=Review findings= section (13 blocking, 27 should-fix, 8 optional); the spec stays +DRAFT until spec-response dispositions them. The blockers cluster in four places: +the existing stamp writers (wrapper and doctor.py) would mark a deferred run +fresh; the split run's =--ignore= needs a dependency closure; =kernel-modules-check= +as written fails forever on ratio's headerless =linux-lts-strix=; and the boot run +has no defined form (package source without network, timeout that can kill pacman +mid-commit, =informant read= run as the user). Also: the installed topgrade +rejects a comma-joined =--disable= list, so each step is its own argument. + +The 2026-10-04 velox health check ran a plain topgrade that moved the kernel +6.18.51 → 6.18.55, against the note above. I checked the gate by hand before the +reboot (=dkms status= installed for 6.18.55, =zfs.ko= in both initramfs images +read with sudo) and the reboot was clean, but the hold this spec builds would +have kept the kernel out of that run. + +Next: spec-response on the findings, then DRAFT → READY and decomposition into +build tasks here. + +*** 2026-10-05 Mon @ 12:30:00 -0500 Spec READY; build tasks below +The spec went DRAFT → READY at 12:10 after three review rounds. Its +Implementation phases section is the contract every task below points at, +and it moves to DOING with these tasks (the first task confirms that). Nine +non-blocking residual findings stayed open, and the first task closes them +before any code. The four-decisions list in the body above is history. + +The order follows the spec's commit groups: archsetup Phase 1, rollout step +(a) on both daily drivers, dotfiles Phase 2 and step (b), archsetup Phase 3 +and step (c) (velox only after the VM gate), the Phase 4 README, then the +flip. While Phase 2 sits unpushed, I push nothing from that dotfiles +checkout, because any push carries it. + +On velox nothing runs plain topgrade, =yay -Syu= or =pacman -Syu=: not a +shell, not the =sysupgrade= alias, and not the panel's UPDATE or TOPGRADE, +which run =yay -Syu= and topgrade until step (b) lands Phase 2 there. Before +step (a), the kernel and zfs-dkms move only in a session I start on purpose, +with the gate checked by hand as on 2026-10-04. From step (a) until step +(b), I update velox only with =upgrade-guarded= from a terminal. +*** TODO Close the nine residual spec findings :chore:maint:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable: the spec, with all nine "round 3 residual" findings under +Review findings applied and completed in place, the way the other findings +were: =DONE=, a =CLOSED:= line, and a =Disposition:= line (accepted, or +modified with the reason). Repo: archsetup, the spec file only. + +First the status heading. spec-response flips it READY → DOING when it files +these tasks. If it still reads READY when I start, I flip it before anything +else, in three lines: the keyword DOING, a newest-first history line ("READY +→ DOING. Decomposed into build tasks under the archsetup todo.org parent, +SPEC_ID 81cdfd72-db96-43d3-aa03-779878c99f3e."), and the Metadata Status +cell set to doing. + +The line numbers inside the findings predate the round-3 consolidation, so I +locate each edit by its quoted OLD text, never by line number: +- Decision 5 Consequences: the snapshot-hold clause ("the hold moves only + when a later session holds a newer one"). +- AC13's everyday claims: P1-steps' step-7 bullet gains the release of the + previous held_snapshot, and Testing maps AC13 to P1-steps too. +- Finish step 1: when several steps fail, the first one's failed_step and + detail win. +- Recovery's power-cut bullet: register the version after =->=, and nothing + when the line is a removed line or there is none. +- Exit codes and interrupts: narrow the bare-=wait= qualifier. +- ZFSBootMenu: roll back with MOD+R (never ENTER, MOD+X or MOD+C) in + Recovery, D-zbm and its variant. +- Design and Decision 5: REBOOT hides from the panel's next probe, and the + fire-time refusal covers the window before it. +- The D-zbm variant: =ssh -t=, the headers' upgrading line as the reset cue, + and a repeat on a fresh VM when the reset missed extraction. +- P1-steps' fake-sudo bullet and Readiness: =yay -Sua= is what never runs, + since =yay -Pwq= already ran. +The snapshot-hold and REBOOT findings both rewrite Decision 5's +Consequences, so I apply those two together. I also reword the three places +that still say ZFSBootMenu boots the held snapshot (Design's ZFS fallback +paragraph, the note after AC13, and Risks' last-resort bullet) to a MOD+R +rollback, so no section points a reader at ENTER. + +This task goes first because several findings change test bullets (P1-steps) +and the manual procedures (Recovery, D-zbm) that later tasks build against. + +Verified by re-reading. The findings quote their own OLD text, so a +whole-file search always matches; I check outside Review findings only. For +each replaced OLD string this prints 0, and for each inserted or replacing +NEW sentence (the first-failure rule in Finish is an insertion, so its anchor +stays) it prints 1: +: sed '/^\* Review findings/,/^\* Implementation phases/d' SPEC | grep -cF 'TEXT' +Each residual reads DONE with a CLOSED line and a disposition. The Review +findings heading cookie reads 138 of 138, updated by hand or with =emacs +--batch= and =org-update-statistics-cookies=, because an edit outside Emacs +doesn't recompute it. The status heading reads DOING. It's a doc edit, so +there's no suite to run. +Contract: Review findings, the nine "round 3 residual" entries. + +*** TODO Thresholds mirror, check 9 and hold-aware prune :feature:installer:zfs:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (archsetup): Phase 1's three standalone changes, one commit +each, each green on =make test-unit=: +- =configs/maintenance-thresholds.toml='s =[updates] guard_patterns= + rewritten to exactly the hook heredoc's 15 Target patterns, with the + comment reworded; the TOML pin test in =tests/installer-steps/=. +- =scripts/post-rebuild-check= check 9 (the =10-= hook name wherever the + guard is installed) behind =PRC_GUARD_BIN= and =PRC_PACMAN_HOOKS_DIR=, with + the header, usage and counters moved to 9 and every test env setting + =PRC_GUARD_BIN= empty. +- =scripts/zfs-pre-snapshot='s prune skips rows with userrefs above 0, and + =tests/zfs-pre-snapshot/fake-zfs= gains hold, release and userrefs, which + the =--complete= tests reuse. +Tests: P1-installer's TOML pin, post-rebuild-check and zfs-pre-snapshot +cases. + +None of these is live on a machine until step (a) installs or re-copies it, +so they land first. Check 9 flags a legacy hook name if post-rebuild-check +runs before step (a); that's the check working. +Verified by =make test-unit= green before and after, plus a mutation spot +check: hand-edit one TOML pattern and the pin test goes red. +Contract: Phase 1, Other Phase 1 changes; Phase 1 tests, P1-installer. + +*** TODO kernel-modules-check gate :feature:zfs:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (archsetup): =scripts/kernel-modules-check=, the stateless +reader that says whether each named kernel would boot with its modules: +usage and exit codes, the per-pkgbase items, the =--since= initramfs and +snapshot items, the =sudo -n lsinitcpio= read on a ZFS root, and the seams +=KMC_MODULES_DIR=, =KMC_BOOT_DIR= and =KMC_DKMS=. It doesn't depend on +=upgrade-guarded=, so it lands first. +Tests: P1-gate in =tests/kernel-modules-check/=, with fakes for dkms, +findmnt, =zfs list= and an lsinitcpio that fails unless run through the +fake sudo. +Verified by =make test-unit= green, plus a read-only real run from the repo +on ratio against =linux-lts=, expecting exit 0 (btrfs root, so no sudo +read). I don't point it at the running =linux-lts-strix=: its zfs is only +'added', and the gate never checks it. The first real =sudo -n lsinitcpio= +read is step (a)'s structural run on velox. +Contract: Phase 1, The gate; Phase 1 tests, P1-gate. + +*** TODO upgrade-guarded foundation: modes, preconditions, sets, record :feature:maint:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (archsetup): =scripts/upgrade-guarded= up to, but not +including, the closure and the transaction steps: +- mode and option parsing (exit 2, no record); +- the preconditions in order: usage, EUID 0, the flock held on fd 9, + =db.lck=, and the =10-= hook's Target lines; +- the sets: pending (with =pacman -Qu='s exit-1-empty rule and its + =[ignored]= rows dropped), GPU patterns, kernel set, DKMS set and live; +- the record (maint's cache envelope, tmp file then rename) and Finish's + result, failed_step and detail rules, with the reason as the last stderr + line; +- the exit codes, and the INT/TERM/HUP trap that waits on the child's PID; +- =--dry-run= on its private =CHECKUPDATES_DB= with =checkupdates --nocolor=; +- the harness: every seam, the fakes on PATH, and subprocess-driven + unittest suites. +Nothing calls the script until Phase 2, so a partial script is harmless. +Tests: P1-sets' cases that need neither the closure nor a transaction (the +three-kernel fixture with the foreign headerless kernel, firmware and +api-headers, and the unowned vmlinuz, minus its gate.pkgbases clause; the +empty DKMS set; the =-Qu= exit handling; the fail-closed hook), plus a +supporting case that the pending set drops =[ignored]= rows; and P1-steps' +usage, EUID 0, lock, =db.lck= and =--dry-run= cases (the probe =--dbpath= +clause lands with the closure). P1-sets' post-transaction deferred-set case +waits for the everyday-run task, and the gate.pkgbases clause for the +=--complete= task. +Verified by =make test-unit= green, plus a real =--dry-run= from the repo on +ratio. It refuses with exit 3 naming the missing =10-= hook, because ratio +keeps the legacy name until step (a). With =UPGRADE_GUARDED_HOOK= pointed at +the legacy file, it prints ratio's real sets, and =linux-lts-strix= isn't +pending. +Contract: Phase 1, Identity and privilege; Modes and options; +Preconditions; Sets; The record; Finish; Exit codes and interrupts; Seams +and fakes. + +*** TODO upgrade-guarded dependency closure :feature:maint:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (archsetup): the closure loop, meaning the read-only =-Sup= +print-mode probe, the everyday stage-A seed, the live-only stage-B seed, +the =--complete= GPU seed, the two line forms, every refusal rule with its +detail, the iteration bound, the kinds each stage assigns, and the held set +built from them. =--dry-run= then prints the GPU closure, the predicted +deferred set and the would-be =-Su= argv. I keep this apart from the +transaction steps because it's the riskiest logic in Phase 1, and it gets +its own verification pass. +Tests: P1-closure, except its single =-Sy= case and the "=-Su= runs" clause +of its new-dependency case, which need the everyday transaction (next +task); P1-sets' held-set cases (the full =--ignore= list with each entry its +own element, the everyday half of the live and not-live case, the +=zfs-dkms=/=zfs-utils= pair held together, a pkgrel-only bump not held), +asserted on =--dry-run='s printed =-Su= argv until the next task re-points +them at the fake pacman's real call; P1-steps' =--dry-run= probe =--dbpath= +clause. The orphan case's "no =-Su= runs" clause is vacuous until the next +task adds the transaction, and bites from then on. +Verified by =make test-unit= green, plus the real =--dry-run= on ratio +(hook seam as before), expecting a closure without a refusal. A refusal +naming a real orphan is itself a finding I clear by hand before step (a). +Contract: Phase 1, Dependency closure; Sets (Kinds, Held set, GPU closure). + +*** TODO upgrade-guarded everyday run, arm procedure and stamp :feature:maint:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (archsetup): the everyday sequence, steps 1 to 10, except step +7's kernel-side detection, which lands with the snapshot-hold task: +- =yay -Pwq= news capture and the record's merge, dedup and cap; +- the informant clear under =timeout 60 sudo -n=; +- one =sudo -n pacman -Sy=, then =-Su= with the =--ignore= list and no =-y=; +- =yay -Sua= on every run, and the sweep argv by absolute path; +- step 9's flag cases, with the arm procedure (resolve check, =-Sw=, the + =install= and =mv=) and the unit-enabled seam; +- Finish's stamp step through =UPGRADE_GUARDED_MAINT=, the everyday stamp + predicate, and the first-failure rule the first task adds to Finish; +- the canonical fixture + =tests/upgrade-guarded/fixtures/upgrade_deferred.json= from a fixed fake + run. +Tests: P1-steps (the =sudo -n= refusal, news, informant, the exit-124 +abort, the sweep argv and its real-binary =--version= check, the failing +yay, the INT cases); P1-closure's single =-Sy= case and the "=-Su= runs" +clause of its new-dependency case; P1-sets' post-transaction deferred-set +case (a fake post-transaction =-Qu= intersected with the held set, excluding +an =[ignored]= row), with its held-set cases re-pointed at the real =-Su= +call; P1-record-stamp's everyday cases and the fixture equality; +P1-arm-boot's two everyday flag cases; and one case for the first-failure +rule: a failing fake yay and a failing fake sweep leave failed_step aur and +yay's detail. +Verified by =make test-unit= green on ratio, where =/usr/bin/topgrade= +exists, so the real-binary check runs instead of skipping. The first run +against real pacman is M-split-live at step (a). +Contract: Phase 1, Everyday sequence; The arm flag; Stamp predicate; +Finish; The record (Fixtures). + +*** TODO upgrade-guarded --complete and its gate entry :feature:maint:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (archsetup): the =--complete= sequence, steps 1 to 7, built and +tested on fixtures that aren't a ZFS root (the fake findmnt reports another +filesystem, so no snapshot or hold runs): +- refresh then flag removal, 'nothing to complete', the pre-stage-1 record + and gate entry, stage 1 (with the reinstall targets when a gate is + already open), the gate in its structural or =--since= form, the GPU step + (stage 2, arm, or the unit-not-enabled line), and the tty-only reboot + prompt after Finish; +- the =--complete= stamp predicate and the Invariants. +Nothing calls or installs the script until Phase 1 is done, so a +=--complete= without its hold is harmless; the next task adds the hold. +Tests: P1-complete except its three snapshot-hold bullets; the +gate.pkgbases clause of P1-sets' three-kernel case; P1-record-stamp's +=--complete= cases; P1-arm-boot's =--complete= arm, unit-not-enabled and +prompt cases; the =--complete= half of P1-sets' live and not-live case; the +=--complete= side of P1-steps' news and informant cases. +Verified by =make test-unit= green. +Contract: Phase 1, =--complete= sequence; Invariants; Stamp predicate; +Exit codes and interrupts (the prompt). + +*** TODO Snapshot hold and the everyday AUR kernel-side check :feature:maint:zfs:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (archsetup), after the =--complete= task: +- =--complete= step 3's own snapshot and hold on a ZFS root, the post-gate + =--since= target rule, and the release of the previous hold; +- everyday step 7's kernel-side detection after yay: the gate entry, the + pre-yay snapshot hold, the flag removal and the immediate record write; +- the trap's matching step-7 path. +Tests: P1-complete's three snapshot-hold bullets; P1-steps' AUR kernel-side +case, with its INT cases and the release-of-the-previous-hold clause the +first task adds. Its "a following =--complete= runs the =--since= form" +clause needs the previous task, which is why this one comes after it. +Verified by =make test-unit= green. The first real snapshot and hold is +M-boot-armed's =--complete= in the VM, which opens the entry on a ZFS root. +Contract: Phase 1, The snapshot hold; Everyday sequence (step 7); Exit +codes and interrupts; Invariants. + +*** TODO upgrade-guarded --apply-armed boot form :feature:maint:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (archsetup): =--apply-armed=, steps 1 to 8. That means reading +the list the unit moved off the flag path (refusing a missing, empty or +unparseable list, or a live Hyprland), the interrupted pre-write, the tty1 +banner, the journal mirror through a SIGINT-ignoring process substitution, +the informant clear, the =vercmp= drop of entries already at or above +their version, the offline =-Sp= check that refuses kernel-side members and +stale versions, the =-S --needed= transaction from the cache, and Finish +with its stamp predicate. +Tests: P1-arm-boot's =--apply-armed= cases, including the =setsid= +process-group SIGINT case (no SIGPIPE death, the fake lock released, the +record interrupted); P1-steps' clause that =--apply-armed= never calls +=yay -Pwq=, and its =--apply-armed= informant and exit-124 cases; +P1-record-stamp's =--apply-armed= success case. +Verified by =make test-unit= green. The real boot path is the VM gate. +Contract: Phase 1, =--apply-armed= sequence; Exit codes and interrupts; +Stamp predicate. + +*** TODO Script install step and health-check workflow edit :feature:installer:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (archsetup), two commits that close Phase 1's group: +1. The =hyprland()= step that installs =hypr-live-update-guard= also installs + =upgrade-guarded= and =kernel-modules-check= to =/usr/local/bin=, with + P1-installer's source-inspection test. +2. A doc commit after the code: =docs/workflows/system-health-check.org= + Phase 3, per the contract's seven bullets. Step 6 runs =upgrade-guarded= + where step (a) is done. Kernel and DKMS days go through =--complete=. + Until step (c), the GPU set finishes from a console. Where step (a) + isn't done and a kernel or DKMS package is pending, plain topgrade + doesn't run. The strix addendum is keyed to each invocation. The script + stamps, so there's no hand stamp. The 2026-09-12 containers entry gets a + resolution note. +After this, Phase 1 is complete and step (a) can start. +Verified by =make test-unit= green, and the workflow edit re-read against +the contract's bullets. +Contract: Phase 1, Files and install; The system-health-check edit; Phase +1 tests, P1-installer. + +*** TODO Rollout step (a) on ratio :chore:ratio: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +After Phase 1's last commit, by hand on ratio, in the contract's order: +- pull archsetup, then =sudo install -m 755= both scripts to + =/usr/local/bin/=; +- write =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook= from the + installer heredoc, confirm its Targets, remove the legacy unprefixed hook + and check =ls /etc/pacman.d/hooks/=; +- diff, then re-copy =configs/maintenance-thresholds.toml= to + =~/.config/archsetup/=; +- confirm =pacman -Qmq= lists no DKMS-set package; +- confirm =upgrade-guarded --dry-run= exits 0; +- run M-split-live on ratio (Manual testing and validation). +Ratio's root is btrfs, so the =zfs-pre-snapshot= install and the structural +gate run belong to velox's step (a). +Verified by M-split-live on ratio passing, and post-rebuild-check's check 9 +clean here (the =10-= hook present, the legacy name gone). If I build Phase +2 on ratio, its dotfiles checkout is the stowed, live one, so each Phase 2 +commit is live on ratio as it's written; that's why Phase 2 starts only +after step (a) on the build machine. +Contract: Phase 4, Rollout (a) and M-split-live. + +*** TODO Rollout step (a) on velox :chore:velox:zfs: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +After Phase 1's last commit, by hand on velox: the same list as ratio's +step (a), plus the ZFS-root items: +- =sudo install -m 755 scripts/zfs-pre-snapshot /usr/local/bin/= for the + hold-aware prune; +- =kernel-modules-check "$(cat /usr/lib/modules/$(uname -r)/pkgbase)"= exits + 0, which proves the =sudo -n lsinitcpio= read against the real image; +- velox's guard hook was hand-placed under the legacy name, so the =10-= + rewrite and the legacy removal matter here too; +- run M-split-live on velox. +On velox nothing runs plain topgrade, =yay -Syu= or =pacman -Syu=: not a +shell, not the =sysupgrade= alias, and not the panel's UPDATE or TOPGRADE, +which run =yay -Syu= and topgrade until step (b) lands Phase 2 there. Before +this task, the kernel and zfs-dkms move only in a session I start on +purpose, with the gate checked by hand: =dkms status= installed for the new +kernel, and =zfs.ko= in each initramfs, read with sudo. From here until +step (b), I update velox only with =upgrade-guarded= from a terminal, and +the health-check workflow's step 6 runs it. +Phase 2 can be built on ratio while this waits, but I don't push any Phase 2 +commit until this is done, because any dotfiles pull on velox would bring it +in. +Verified by M-split-live on velox, the structural gate run exiting 0, and +post-rebuild-check's check 9 clean here (the =10-= hook present, the legacy +name gone). +Contract: Phase 4, Rollout (a) and M-split-live. + +*** TODO maint deferred row, APPLY and gate-aware REBOOT :feature:maint:dotfiles:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (dotfiles), one commit, because the row and APPLY land together: +- the =upgrade_deferred= probe and the helper it shares with + =topgrade_freshness= (N, never raises, None on a missing or malformed + record), the severity order, the armed text from the flag, and the + evidence (rows, detail, held_snapshot, and news lines not in + =upgrade_news_dismissed=, which stays empty until DISMISS lands); +- =apply_deferred= (tier confirm, kind terminal), its =LEVER_KEYS= builder, + =topgrade_freshness='s =deferred= evidence, =doctor.review='s routing, + =panel.fire_press= launching =foot --hold -e upgrade-guarded --complete= + detached, and =maint fix apply_deferred= on inherited stdio; +- Freshness and REBOOT: the wall note, REBOOT hidden while a gate is open + and shown while armed, =iter_fix='s fire-time refusal, and =review= + omitting =reboot= while a gate is open; +- the byte-identical fixture copy at + =tests/maint/fixtures/upgrade_deferred.json=, with its test module's + docstring naming the archsetup source. +I land this before the lever repoint, so gate-aware REBOOT already exists +the first time a lever runs =upgrade-guarded=, whose AUR step can open a +gate. Until the repoint the row reflects only shell runs, and with no +record it reads OK '0 deferred'. + +Repo: =~/.dotfiles=, a separate project. Confirm the crossing before +starting, or run it from a dotfiles session. This starts after step (a) on +the machine I build Phase 2 on (=uname -n=): its dotfiles checkout is stowed +and live, so each commit reaches that machine as it's written. Baseline: +=make test= in =~/.dotfiles=; the red screen-lock suite on ratio is the +known failure, tracked in "screen-lock test suite red on ratio". +Tests: P2-row, P2-apply, and P2-text's wall-note case. No test imports +gui.py. +Verified by the =tests/maint= suites and =make test= green apart from that +known red; =cmp= on the two fixture copies; and =make test-panel-maint=, the +AT-SPI smoke over the fixture boards. =make test= never runs it, and it's +the only check that drives gui.py, which this commit edits. Ratio has no +weston or sway, so the smoke uses the live compositor and opens a panel +window; I run it on a spare workspace, and a SKIP is not a pass. The real +APPLY press is M-boot-armed on ratio and velox. +Contract: Phase 2, The deferred row; APPLY; Freshness and REBOOT; Phase 1, +The record (Fixtures). + +*** TODO maint levers repointed to upgrade-guarded :feature:maint:dotfiles:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (dotfiles), the one commit the spec ties together: +- UPDATE and TOPGRADE argv and the missing-binary message; +- doctor.py's =topgrade_run= stamp deleted; +- the guard UX removed: the =live_update= tag, =iter_fix='s refusal and + guard event, =--force= and its sentinel wrap, =_update_force=, + =_rearm_after_guard=, the guard branches in =_on_fired=, cli.py and + panel.py, and the review suffix; =panel.guarded= becomes a set test, and + the tripped arm line becomes the annotation; +- the =sysupgrade= conditional in both common alias files, with minimal's + files left as symlinks; +- the stale text: the =topgrade_freshness= docstring and absent-stamp + error, the doctor.py and guard.py docstrings, the README =maint fix= + usage line, and the wrapper's header and its test docstring; +- =tests/maint/panel_smoke.py='s "arm line shows the exact update argv" + check, which pins =yay -Syu --noconfirm=, now matching + =upgrade-guarded --no-topgrade=. The spec doesn't list the file, but the + repoint breaks it. +The runner's long-remedy timeout (new session, SIGINT to the group, 120 s, +then SIGKILL, 'timed out — interrupted') can land as its own commit just +before this one. UPDATE and TOPGRADE are already user-kind and marked long, +so it applies to today's =yay -Syu= and topgrade argv the moment it lands. +That's safe: today a timeout SIGKILLs only the direct child +(=subprocess.run='s timeout), and a SIGINT to the whole group stops pacman +at a package boundary instead. + +Repo: =~/.dotfiles=, a separate project. Confirm the crossing before +starting, or run it from a dotfiles session. +Tests: P2-levers, P2-guard-ux (including the rewritten +test_panel_levers.py and test_panel_phase10.py cases), and P2-text's +alias, absent-stamp and usage-line cases. +Verified by the =tests/maint= suites and =make test= green apart from the +known screen-lock red, and =make test-panel-maint= passing on a spare +workspace (a SKIP is not a pass). After it, the first UPDATE press on the +build machine runs the same command M-split-live already ran there. +Contract: Phase 2, Levers; Guard UX; Alias and README (the alias). + +*** TODO maint DISMISS, held-package CVE/QUEUE/strip, README :feature:maint:dotfiles:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (dotfiles), three commits after the repoint: +1. DISMISS: the =dismiss= key kind, dispatched explicitly in =_digest_key=, + arm-then-fire, and the GTK-free panel.py helper that overwrites + =upgrade_news_dismissed= with the record's whole news list. Tests: + P2-dismiss. +2. CVE, QUEUE and the strip: deferred names out of =cve_queued=, the CVE + caption, UPDATE's binding and REVIEW & FIX; =[HELD]= tags in QUEUE; and + the strip's 'P pending · N held'. Tests: P2-cve-queue. +3. The 'The live-update guard' section of =maint/README.md= rewritten for + the new levers, the sweep argv, the deferred row, APPLY, DISMISS and the + armed state, with no press-again or =--force=. The guard banner and the + wrapper paragraph stay. +Then I push Phase 2, and only once step (a) is done on velox. Until then I +push nothing from this dotfiles checkout, because any push carries it. + +Repo: =~/.dotfiles=, a separate project. Confirm the crossing before +starting, or run it from a dotfiles session. +Verified by the =tests/maint= suites and =make test= green apart from the +known red, =make test-panel-maint= passing on a spare workspace (the strip's +pending caption is in its path; a SKIP is not a pass), and the README +section re-read against the contract. +Contract: Phase 2, DISMISS; CVE, QUEUE and the strip; Alias and README. + +*** TODO Rollout step (b) on the other daily driver :chore:dotfiles:velox: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Pull the Phase 2 commits into the other daily driver's dotfiles, only after +step (a) is done there. The build machine already has them through its live +checkout, so this step is for the other one: velox when I build Phase 2 on +ratio. If I build it on velox, this is ratio's step (b) and the topic tag +moves to =:ratio:=. +Verified by that machine's dotfiles HEAD matching the remote, the +=tests/maint= suites green there, and =maint status= there listing the +=upgrade_deferred= row. +Contract: Phase 4, Rollout (b); the Phase 2 intro. + +*** TODO archsetup-boot-upgrade.service installer step :feature:installer:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (archsetup): an installer step that writes the unit from a +heredoc with =ARCHSETUP_USERNAME= substituted by sed, runs +=install -d -m 0755 /var/lib/archsetup=, and enables the unit without +=--now=, documented in its own comments. The spec names no function. It +can sit beside the guard install in =hyprland()=, or in a top-level +function called from there, and the source-inspection idiom reads either. +Tests: P3-unit in =tests/installer-steps/=: every directive the contract +lists, plus the absence of Requires=, BindsTo=, RequiredBy=, network-online +and Restart=. +Verified by =make test-unit= green. The boot behavior itself is the VM gate. +Contract: Phase 3, Installer step; The unit; Disarm and failure; Phase 3 +tests, P3-unit. + +*** TODO Rollout step (c) on ratio :chore:ratio: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +After Phase 3's commit, by hand on ratio: write the unit from the installer +heredoc with the username substituted, then run +=sudo install -d -m 0755 /var/lib/archsetup=, =sudo systemctl daemon-reload= +and =sudo systemctl enable archsetup-boot-upgrade.service=. The VM gate +holds back only velox, so ratio can go first. +Then, on the first day =upgrade-guarded --dry-run= lists a GPU-kind +package, run M-boot-armed on ratio and M-boot-ratio on the same armed boot. +Run the strix check on the first =--complete= that lands linux or +linux-lts. All three are under Manual testing and validation. +Verified by those three manual tests. +Contract: Phase 4, Rollout (c) and the ratio-path confirmation. + +*** TODO VM gate for the boot unit :test:installer:zfs: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Run the five VM-gate tests in the =make test-keep FS_PROFILE=zfs= VM, +following Phase 3's VM procedure, in this order: M-split-console, +M-boot-armed, then M-boot-midtx on the same VM, then M-boot-timeout and +M-boot-fail. M-boot-midtx needs the VM M-boot-armed leaves, because it +times its cut from that boot's journal, so nothing reinstalls between them. +Each test is a child of Manual testing and validation. +=make test-keep= bundles archsetup from HEAD, so Phase 1 and Phase 3 must be +committed first. The VM clones the published dotfiles, so this runs after +the Phase 2 push; M-boot-timeout and M-boot-fail then read maint's +failed-units and deferred rows inside the VM, which AC6 asks for. I never +use =debug-vm.sh= here: it restores the clean-install snapshot and erases +the install under test. +M-boot-timeout and M-boot-fail use small inline stand-in fakes in place of +the Phase 1 fake sudo and pacman: they pass every query to the real pacman +and sleep or fail only on the transaction, so the VM tests don't depend on +how the unit-test fakes are laid out. +The unit reaches velox only after all five pass. +Contract: Phase 3, VM procedure and VM gate; Testing / Verification / +Rollout. + +*** TODO Rollout step (c) on velox :chore:velox:zfs: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Only after the VM gate passes: the same four steps as ratio's step (c). +Then, on the first day =upgrade-guarded --dry-run= lists a GPU-kind +package, run M-boot-armed on velox and M-boot-news. +Verified by those two manual tests. +Contract: Phase 4, Rollout (c). + +*** TODO maint README flow and recovery steps :chore:maint:dotfiles:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +Deliverable (dotfiles), one commit to the 'The live-update guard' section +of =maint/README.md=, after Phase 3. It adds the flow line, the +=kernel-modules-check <pkgbase>= diagnosis note for a CRIT row (only +=--complete= clears the gate), and the recovery steps as their only copy. +The steps come from the spec's Recovery as the first task amends it: the +MOD+R rollback and the version-after-the-arrow rule. +Repo: =~/.dotfiles=, a separate project. Confirm the crossing before +starting, or run it from a dotfiles session. +The non-gating D-zbm drill and its power-cut variant exercise these steps +(Manual testing and validation), and nothing waits on them. +Verified by =make test= green apart from the known screen-lock red, and the +section re-read against the contract. +Contract: Phase 4, README flow and recovery; Recovery; Recovery drill. + +*** TODO Flip the spec to IMPLEMENTED (+ dated history line + Metadata mirror) :chore:maint: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +This waits until every task above is done and the gating manual tests have +passed: M-split-live on both machines, the five VM-gate tests, +M-boot-armed on ratio and velox, M-boot-news, M-boot-ratio and the strix +check. Then tick each acceptance criterion against the evidence that +Testing / Verification / Rollout maps to it. AC1 stays unticked until a live +run with a GPU-kind entry pending has shown the hook silent. Last, the spec +edits: the status keyword DOING → IMPLEMENTED, a dated history line that +says why, and the Metadata Status mirror set to implemented. D-zbm is +non-gating and doesn't hold this. +Contract: the spec's status heading and Metadata table. + +** TODO [#D] Stale-kernel age in maint :feature:maint:dotfiles: +:PROPERTIES: +:LAST_REVIEWED: 2026-10-05 +:END: +The guarded-upgrade build holds the kernel and DKMS sets on every everyday +run, so kernel security fixes wait for an =upgrade-guarded --complete= I +start on purpose. The deferred row is the only reminder, and it shows a +kernel most days, so it stops reading as a nag. This is the spec's vNext: a +maint age for how long the kernel set has been held, graded so an overdue +dedicated session shows up. + +The design is still open. The record has no held-since field, so the age +needs either a new field (both fixture copies change in one rollout) or a +derivation from pacman.log. maint's CVE probe doesn't cover it: as of +2026-10-05 the kernel advisories in its cache carry no fixed version, so it +never flags a held kernel. + +I kept this out of v1 on purpose. It replaces the "vNext kernel-reboot +item" in the build parent's original body. +Spec: [[file:docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org][guarded-upgrade spec]], Scope tiers and Risks (Standing kernel +deferral). ** TODO [#C] post-rebuild-check: probe that Emacs frames come up Wayland-native :feature:emacs:velox:solo:quick: :PROPERTIES: @@ -3529,6 +4152,1034 @@ minutes is the entry freeze: hold power for 10 s, and write down the cycle number and whether the keyboard backlight was lit. A cycle that instead comes straight back with "Cannot allocate memory" is the ARC task, not a freeze. +*** M-split-console: an everyday run with no compositor defers only kernel-kind entries +What we're verifying: with Hyprland not running, =upgrade-guarded +--no-topgrade= lands the pending GPU/compositor packages on real pacman and +leaves only kernel-kind entries deferred (AC3's held-set claim). It's the +first test of the VM gate. +- Make sure Phase 1's commits are in HEAD (=make test-keep= bundles + archsetup from HEAD). +#+begin_src sh :results output +# VM helpers; every later block in this test sources them. Root logs in by +# key only after the install, so vmroot uses the key the harness left with +# the newest test results. +cat > /tmp/ug-vm.sh <<'EOF' +cd ~/code/archsetup || exit 1 +o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222" +key=$(ls -td test-results/*/root_key 2>/dev/null | head -1) +vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; } +vmroot() { ssh $o -i "$key" root@localhost "$@"; } +mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; } +shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1 + magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; } +shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d" + for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done + for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done + echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; } +EOF +echo "helpers written to /tmp/ug-vm.sh" +#+end_src +#+begin_src sh :results output +# The full install runs first; this takes a while, and Emacs waits on it. +cd ~/code/archsetup && make test-keep FS_PROFILE=zfs 2>&1 | tail -6 +#+end_src +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'pacman -Q wayland libdrm $(pacman -Qqo /usr/lib/modules/*/vmlinuz)' +#+end_src +- On https://archive.archlinux.org/packages/, find the release before the + installed one for =wayland= or =libdrm= (whichever downgrades without + breaking a dependency), and for the VM's kernel and its headers. +- Paste the three URLs into the next block. +#+begin_src sh :results output +. /tmp/ug-vm.sh +gpu_url='PASTE-URL'; kernel_url='PASTE-URL'; headers_url='PASTE-URL' +vmroot "pacman -U --noconfirm $gpu_url $kernel_url $headers_url" | tail -3 +vm 'LC_ALL=C pacman -Qu' +#+end_src +- Confirm Hyprland isn't running in the VM. If the block prints LIVE, stop + and record it under this test: the VM procedure's premise has failed. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'pgrep -x Hyprland >/dev/null && echo "Hyprland LIVE: the test premise fails" || echo "Hyprland not running"' +#+end_src +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'upgrade-guarded --no-topgrade > /tmp/ug-console.log 2>&1; echo "exit: $?" +tail -3 /tmp/ug-console.log +jq -r ".data.packages[] | \"\(.kind) \(.name)\"" ~/.local/state/maint/upgrade_deferred.json +echo "-- pacman -Qu:"; LC_ALL=C pacman -Qu' +#+end_src +Expected: Hyprland was not running; the run printed exit 0; the GPU-pattern +package is back at its current version; every record line reads kernel; +and =pacman -Qu= lists only the kernel and its headers, plus any +=[ignored]= row. + +*** M-split-live on ratio: the live everyday run leaves the guard silent +What we're verifying: the everyday split transaction on real pacman with +Hyprland live, before any lever calls it. It should exit 0, print no +BLOCKED banner, and defer exactly what =--dry-run= predicted (AC1). It's +the last item of step (a) on ratio. +- Run it in ratio's live Hyprland session, right after the rest of step + (a). +#+begin_src sh :results output +upgrade-guarded --dry-run > /tmp/ug-dry.txt 2>&1; echo "dry-run exit: $?" +cat /tmp/ug-dry.txt +#+end_src +- Note the predicted deferred set, and whether it lists a GPU-kind entry. +#+begin_src sh :results output +# The real run: refresh, -Su with the held set ignored, then yay -Sua. It +# takes minutes, and Emacs waits on it. +upgrade-guarded --no-topgrade > /tmp/ug-live.log 2>&1; echo "exit: $?" +echo "BLOCKED lines: $(grep -c BLOCKED /tmp/ug-live.log)" +jq -r '.data.packages[] | "\(.kind) \(.name)"' ~/.local/state/maint/upgrade_deferred.json | sort +echo "-- pacman -Qu:"; LC_ALL=C pacman -Qu +#+end_src +- If the dry-run listed no GPU-kind entry, leave this test open and repeat + it on the first day one is pending. The hook-silence half needs one. +Expected: exit 0; zero BLOCKED lines; the record's names equal the +dry-run's predicted deferred set; =pacman -Qu= lists only those names plus +the =[ignored]= bridge-utils row; and on a run with a GPU-kind entry +pending, that entry is in the record and the guard never printed its +banner. + +*** M-split-live on velox: the live everyday run leaves the guard silent +What we're verifying: the same everyday split transaction on velox's ZFS +root with Hyprland live, before any lever calls it. It should exit 0, print +no BLOCKED banner, and defer exactly what =--dry-run= predicted, with the +kernel and zfs-dkms among the deferred entries when pending (AC1). It's the +last item of step (a) on velox. +- Run it in velox's live Hyprland session, right after the rest of step + (a). +#+begin_src sh :results output +uname -r +upgrade-guarded --dry-run > /tmp/ug-dry.txt 2>&1; echo "dry-run exit: $?" +cat /tmp/ug-dry.txt +#+end_src +- Note the predicted deferred set, and whether it lists a GPU-kind entry. +#+begin_src sh :results output +# The real run: refresh, -Su with the held set ignored, then yay -Sua. It +# takes minutes, and Emacs waits on it. +upgrade-guarded --no-topgrade > /tmp/ug-live.log 2>&1; echo "exit: $?" +echo "BLOCKED lines: $(grep -c BLOCKED /tmp/ug-live.log)" +jq -r '.data.packages[] | "\(.kind) \(.name)"' ~/.local/state/maint/upgrade_deferred.json | sort +echo "-- pacman -Qu:"; LC_ALL=C pacman -Qu +pacman -Q $(pacman -Qqo /usr/lib/modules/*/vmlinuz) +#+end_src +- If the dry-run listed no GPU-kind entry, leave this test open and repeat + it on the first day one is pending. The hook-silence half needs one. +Expected: exit 0; zero BLOCKED lines; the record's names equal the +dry-run's predicted deferred set; =pacman -Qu= lists only those names plus +any =[ignored]= rows; the kernel package's version is unchanged; and on a +run with a GPU-kind entry pending, that entry is in the record and the +guard never printed its banner. + +*** M-boot-armed in the ZFS VM: the armed set lands from cache on tty1 before login, offline +What we're verifying: the completion path on a ZFS root with no network at +boot. =--complete= should pass the real gate (=sudo -n lsinitcpio= +included) and arm without firing the guard or informant's hook. The boot +unit should then apply the armed set before getty@tty1 starts (AC5, and +AC4's gate on a real initramfs). Second test of the VM gate. +- Make sure Phase 1's and Phase 3's commits are in HEAD. +#+begin_src sh :results output +# VM helpers; every later block in this test sources them. Root logs in by +# key only after the install, so vmroot uses the key the harness left with +# the newest test results. +cat > /tmp/ug-vm.sh <<'EOF' +cd ~/code/archsetup || exit 1 +o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222" +key=$(ls -td test-results/*/root_key 2>/dev/null | head -1) +vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; } +vmroot() { ssh $o -i "$key" root@localhost "$@"; } +mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; } +shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1 + magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; } +shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d" + for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done + for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done + echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; } +EOF +echo "helpers written to /tmp/ug-vm.sh" +#+end_src +#+begin_src sh :results output +# The full install runs first; this takes a while, and Emacs waits on it. +cd ~/code/archsetup && make test-keep FS_PROFILE=zfs 2>&1 | tail -6 +#+end_src +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'pacman -Q wayland libdrm $(pacman -Qqo /usr/lib/modules/*/vmlinuz)' +#+end_src +- On https://archive.archlinux.org/packages/, find the release before the + installed one for =wayland= or =libdrm=, and for the VM's kernel and its + headers. +- Paste the three URLs into the next block. +#+begin_src sh :results output +. /tmp/ug-vm.sh +gpu_url='PASTE-URL'; kernel_url='PASTE-URL'; headers_url='PASTE-URL' +vmroot "pacman -U --noconfirm $gpu_url $kernel_url $headers_url" | tail -3 +vm 'LC_ALL=C pacman -Qu' +#+end_src +- Arm over SSH with the live branch forced, since Hyprland never runs in + the VM. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'UPGRADE_GUARDED_HYPR_RUNNING=1 upgrade-guarded --complete > /tmp/ug-arm.log 2>&1; echo "exit: $?" +echo "pre-transaction hook runs: $(grep -c "Running pre-transaction hooks" /tmp/ug-arm.log)" +echo "BLOCKED lines: $(grep -c BLOCKED /tmp/ug-arm.log)" +echo "-- flag:"; cat /var/lib/archsetup/apply-upgrade-on-boot +echo "-- record:"; jq -c ".data | {result, gate, held_snapshot, n: (.packages | length)}" ~/.local/state/maint/upgrade_deferred.json' +#+end_src +- Reboot the guest and take its network link down straight away. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'sudo systemctl reboot'; sleep 3; mon set_link net0 off >/dev/null; echo "link off" +#+end_src +- Capture tty1 for three minutes while it boots. +#+begin_src sh :results output +. /tmp/ug-vm.sh +shots 90 +#+end_src +- Open the frames in order (=imv /tmp/ug-shots/=). +- Bring the link back. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon set_link net0 on >/dev/null; sleep 20; echo "link on" +#+end_src +- Read the boot's outcome. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'journalctl -b -t archsetup-boot-upgrade --no-pager | head -2 +systemctl show archsetup-boot-upgrade -p Result -p ExecMainExitTimestampMonotonic +systemctl show getty@tty1 -p ExecMainStartTimestampMonotonic +ls /var/lib/archsetup/apply-upgrade-on-boot 2>&1 +jq -c ".data | {result, failed_step, n: (.packages | length)}" ~/.local/state/maint/upgrade_deferred.json +echo "stamp age: $(( $(date +%s) - $(jq -r ".written_at | floor" ~/.local/state/maint/topgrade_run.json) )) s" +echo "uptime: $(cut -d. -f1 /proc/uptime) s"' +#+end_src +Expected: the arm exited 0 with exactly one pre-transaction hook run (stage +1's; the arm is forced live, so no stage 2 runs, and the arm's =-Sw= ran no +hooks) and zero BLOCKED lines. The flag listed the GPU entries as +=name=version= lines, and the record had gate null with held_snapshot +naming a pre-pacman snapshot. A frame shows 'archsetup: applying N deferred +GPU/compositor upgrades — do not power off' on tty1 before any frame shows +the login, and the journal's first line is that banner. Result=success, and +the unit's exit timestamp is below getty@tty1's start timestamp. Afterwards +the flag is gone, the record reads result ok with n 0, and the stamp's age +is below the uptime. + +*** M-boot-midtx in the ZFS VM: a timeout inside a real transaction leaves no db.lck +What we're verifying: when the start timeout's SIGINT lands in the middle +of a real transaction, pacman should stop at a package boundary and +release =db.lck=. The record should name the interruption, and +=--complete= from a console should finish the set (AC6). Third test of the +VM gate, on the VM M-boot-armed leaves. +- Use the VM M-boot-armed just left. Nothing reinstalls between the two + tests, so its journal still times a real boot transaction. +#+begin_src sh :results output +# VM helpers; every later block in this test sources them. Root logs in by +# key only after the install, so vmroot uses the key the harness left with +# the newest test results. +cat > /tmp/ug-vm.sh <<'EOF' +cd ~/code/archsetup || exit 1 +o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222" +key=$(ls -td test-results/*/root_key 2>/dev/null | head -1) +vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; } +vmroot() { ssh $o -i "$key" root@localhost "$@"; } +mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; } +shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1 + magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; } +shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d" + for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done + for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done + echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; } +EOF +echo "helpers written to /tmp/ug-vm.sh" +#+end_src +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'journalctl -t archsetup-boot-upgrade -o short-monotonic --no-pager | grep -E "applying|upgrading|installing|transaction" | tail -12' +#+end_src +- On https://archive.archlinux.org/packages/, find the previous releases of + two or more GPU-pattern packages, mesa among them, so the transaction is + long enough to cut. +- Paste their URLs into the next block. +#+begin_src sh :results output +. /tmp/ug-vm.sh +gpu_urls='PASTE-URL PASTE-URL' +vmroot "pacman -U --noconfirm $gpu_urls" | tail -2; vm 'LC_ALL=C pacman -Qu' +#+end_src +- Arm over SSH with the live branch forced. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'UPGRADE_GUARDED_HYPR_RUNNING=1 upgrade-guarded --complete > /tmp/ug-arm.log 2>&1; echo "exit: $?" +cat /var/lib/archsetup/apply-upgrade-on-boot' +#+end_src +- Work out a start timeout that lands inside the transaction: the seconds + from the banner to about halfway through the 'upgrading' lines above, + scaled up for the larger set. +- Put it in the next block. +#+begin_src sh :results output +. /tmp/ug-vm.sh +secs=PASTE-SECONDS +vmroot "install -d /etc/systemd/system/archsetup-boot-upgrade.service.d +printf '[Service]\nTimeoutStartSec=%ss\n' $secs > /etc/systemd/system/archsetup-boot-upgrade.service.d/test.conf +systemctl daemon-reload && cat /etc/systemd/system/archsetup-boot-upgrade.service.d/test.conf" +#+end_src +- Reboot the guest. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'sudo systemctl reboot'; echo rebooting +#+end_src +- Capture tty1 for three minutes. +#+begin_src sh :results output +. /tmp/ug-vm.sh +shots 90 +#+end_src +- Open the frames in order (=imv /tmp/ug-shots/=). +- Read the outcome once the login shows. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'ls /var/lib/pacman/db.lck 2>&1 +jq -c ".data | {result, failed_step, detail, n: (.packages | length)}" ~/.local/state/maint/upgrade_deferred.json +journalctl -b -t archsetup-boot-upgrade --no-pager | tail -4' +#+end_src +- If the journal shows the transaction finished, or never started, change + the seconds and repeat from the downgrade block. +- Confirm Hyprland isn't running, so the next =--complete= applies the set + rather than arming it. If the block prints LIVE, record it under this + test. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'pgrep -x Hyprland >/dev/null && echo "Hyprland LIVE: the test premise fails" || echo "Hyprland not running"' +#+end_src +- Finish the set from a console. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'upgrade-guarded --complete > /tmp/ug-finish.log 2>&1; echo "exit: $?" +LC_ALL=C pacman -Qu +jq -c ".data | {result, n: (.packages | length)}" ~/.local/state/maint/upgrade_deferred.json' +#+end_src +- Remove the drop-in. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vmroot 'rm -rf /etc/systemd/system/archsetup-boot-upgrade.service.d && systemctl daemon-reload && echo removed' +#+end_src +Expected: no =db.lck= after the cut-off boot, and the journal stopped +between packages. The record read result interrupted with failed_step +boot-transaction and the =--complete= remedy. Then the console +=--complete= exited 0, =pacman -Qu= no longer lists the armed packages, and +the record reads result ok with n 0. + +*** M-boot-timeout in the ZFS VM: a hung boot run times out to the login and disarms +What we're verifying: a boot transaction that never ends is cut off by the +start timeout (shortened here to 90 s). The login should still appear, the +flag should be gone, the failure should be visible in the record and in +maint, and the next boot should skip the unit (AC6). Fourth test of the VM +gate. +#+begin_src sh :results output +# VM helpers; every later block in this test sources them. Root logs in by +# key only after the install, so vmroot uses the key the harness left with +# the newest test results. +cat > /tmp/ug-vm.sh <<'EOF' +cd ~/code/archsetup || exit 1 +o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222" +key=$(ls -td test-results/*/root_key 2>/dev/null | head -1) +vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; } +vmroot() { ssh $o -i "$key" root@localhost "$@"; } +mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; } +shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1 + magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; } +shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d" + for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done + for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done + echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; } +EOF +echo "helpers written to /tmp/ug-vm.sh" +#+end_src +- Run the next block for a fresh VM, or skip it to reuse the VM + M-boot-midtx left (the block reinstalls from clean). +#+begin_src sh :results output +# The full install runs first; this takes a while, and Emacs waits on it. +cd ~/code/archsetup && make test-keep FS_PROFILE=zfs 2>&1 | tail -6 +#+end_src +- On https://archive.archlinux.org/packages/, find the release before the + installed =wayland= or =libdrm=. +- Paste its URL into the next block. +#+begin_src sh :results output +. /tmp/ug-vm.sh +gpu_url='PASTE-URL' +vmroot "pacman -U --noconfirm $gpu_url" | tail -2; vm 'LC_ALL=C pacman -Qu' +#+end_src +- Arm over SSH with the live branch forced. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'UPGRADE_GUARDED_HYPR_RUNNING=1 upgrade-guarded --complete > /tmp/ug-arm.log 2>&1; echo "exit: $?" +cat /var/lib/archsetup/apply-upgrade-on-boot' +#+end_src +- Install the stand-in fakes and the drop-in. The fake pacman hands every + query to the real one and sleeps on the transaction, and the fake sudo + drops =-n=. They do the job the Phase 1 fakes do. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vmroot 'set -e +d=/usr/local/lib/ug-fakes +install -d "$d" /etc/systemd/system/archsetup-boot-upgrade.service.d +cat > "$d/sudo" <<"EOF" +#!/bin/sh +# Drop -n and run the command as the caller. +[ "$1" = -n ] && shift +exec "$@" +EOF +cat > "$d/pacman" <<"EOF" +#!/bin/sh +# Real pacman for every query; the transaction (-S) sleeps or fails. +if [ "$1" = -S ]; then + [ "$UG_FAKE_MODE" = fail ] && { echo "error: failed to commit transaction (fake)" >&2; exit 1; } + exec sleep 3600 +fi +exec /usr/bin/pacman "$@" +EOF +chmod 755 "$d/sudo" "$d/pacman" +cat > /etc/systemd/system/archsetup-boot-upgrade.service.d/test.conf <<"EOF" +[Service] +TimeoutStartSec=90s +Environment=PATH=/usr/local/lib/ug-fakes:/usr/local/bin:/usr/bin +Environment=UG_FAKE_MODE=sleep +EOF +systemctl daemon-reload && echo installed' +#+end_src +- Reboot the guest. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'sudo systemctl reboot'; echo rebooting +#+end_src +- Capture tty1 for four minutes. +#+begin_src sh :results output +. /tmp/ug-vm.sh +shots 120 +#+end_src +- Open the frames in order (=imv /tmp/ug-shots/=). +- Read the outcome once the login shows. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'ls /var/lib/archsetup/apply-upgrade-on-boot 2>&1 +systemctl is-failed archsetup-boot-upgrade.service +systemctl show archsetup-boot-upgrade -p ExecMainStartTimestampMonotonic +systemctl show getty@tty1 -p ExecMainStartTimestampMonotonic +jq -c ".data | {result, failed_step, detail, n: (.packages | length)}" ~/.local/state/maint/upgrade_deferred.json +ls /var/lib/pacman/db.lck 2>&1 +~/.local/bin/maint status 2>&1 | grep -i -E "failed|deferred"' +#+end_src +- Reboot the guest once more. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'sudo systemctl reboot'; echo rebooting +#+end_src +- About a minute later, read whether the unit ran. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'systemctl show archsetup-boot-upgrade -p ConditionResult -p ActiveState' +#+end_src +- Remove the drop-in and the fakes. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vmroot 'rm -rf /etc/systemd/system/archsetup-boot-upgrade.service.d /usr/local/lib/ug-fakes && systemctl daemon-reload && echo removed' +#+end_src +Expected: the login appeared about 90 s into the unit: getty@tty1's start +timestamp is about 90,000,000 µs above the unit's +ExecMainStartTimestampMonotonic. After it, the flag was gone, =is-failed= +printed failed, and no =db.lck= existed. The record read result +interrupted with failed_step boot-transaction, a detail naming +=upgrade-guarded --complete=, and n equal to the armed count. maint's +failed-units row names archsetup-boot-upgrade.service, and the deferred row +lists the armed set at WARN. The next boot shows ConditionResult=no. + +*** M-boot-fail in the ZFS VM: a failed boot transaction still reaches the login and disarms +What we're verifying: a boot transaction that fails outright should still +end at the tty1 login. The flag should be gone, the record should name the +failed step, maint should show the failed unit and the deferred set, and +the unit should be left failed (AC6). Last test of the VM gate. +#+begin_src sh :results output +# VM helpers; every later block in this test sources them. Root logs in by +# key only after the install, so vmroot uses the key the harness left with +# the newest test results. +cat > /tmp/ug-vm.sh <<'EOF' +cd ~/code/archsetup || exit 1 +o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222" +key=$(ls -td test-results/*/root_key 2>/dev/null | head -1) +vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; } +vmroot() { ssh $o -i "$key" root@localhost "$@"; } +mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; } +shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1 + magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; } +shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d" + for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done + for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done + echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; } +EOF +echo "helpers written to /tmp/ug-vm.sh" +#+end_src +- Run the next block for a fresh VM, or skip it to reuse the VM the + previous M-boot test left (the block reinstalls from clean). +#+begin_src sh :results output +# The full install runs first; this takes a while, and Emacs waits on it. +cd ~/code/archsetup && make test-keep FS_PROFILE=zfs 2>&1 | tail -6 +#+end_src +- On https://archive.archlinux.org/packages/, find the release before the + installed =wayland= or =libdrm=. +- Paste its URL into the next block. +#+begin_src sh :results output +. /tmp/ug-vm.sh +gpu_url='PASTE-URL' +vmroot "pacman -U --noconfirm $gpu_url" | tail -2; vm 'LC_ALL=C pacman -Qu' +#+end_src +- Arm over SSH with the live branch forced. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'UPGRADE_GUARDED_HYPR_RUNNING=1 upgrade-guarded --complete > /tmp/ug-arm.log 2>&1; echo "exit: $?" +cat /var/lib/archsetup/apply-upgrade-on-boot' +#+end_src +- Install the stand-in fakes and the drop-in, with the fake pacman set to + answer =-Q= and =-Sp= normally and exit 1 on =-S=. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vmroot 'set -e +d=/usr/local/lib/ug-fakes +install -d "$d" /etc/systemd/system/archsetup-boot-upgrade.service.d +cat > "$d/sudo" <<"EOF" +#!/bin/sh +# Drop -n and run the command as the caller. +[ "$1" = -n ] && shift +exec "$@" +EOF +cat > "$d/pacman" <<"EOF" +#!/bin/sh +# Real pacman for every query; the transaction (-S) sleeps or fails. +if [ "$1" = -S ]; then + [ "$UG_FAKE_MODE" = fail ] && { echo "error: failed to commit transaction (fake)" >&2; exit 1; } + exec sleep 3600 +fi +exec /usr/bin/pacman "$@" +EOF +chmod 755 "$d/sudo" "$d/pacman" +cat > /etc/systemd/system/archsetup-boot-upgrade.service.d/test.conf <<"EOF" +[Service] +TimeoutStartSec=90s +Environment=PATH=/usr/local/lib/ug-fakes:/usr/local/bin:/usr/bin +Environment=UG_FAKE_MODE=fail +EOF +systemctl daemon-reload && echo installed' +#+end_src +- Reboot the guest. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'sudo systemctl reboot'; echo rebooting +#+end_src +- Capture tty1 for two minutes. +#+begin_src sh :results output +. /tmp/ug-vm.sh +shots 60 +#+end_src +- Open the frames in order (=imv /tmp/ug-shots/=). +- Read the outcome once the login shows. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'ls /var/lib/archsetup/apply-upgrade-on-boot 2>&1 +systemctl is-failed archsetup-boot-upgrade.service +jq -c ".data | {result, failed_step, detail, n: (.packages | length)}" ~/.local/state/maint/upgrade_deferred.json +~/.local/bin/maint status 2>&1 | grep -i -E "failed|deferred"' +#+end_src +- Remove the drop-in and the fakes. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vmroot 'rm -rf /etc/systemd/system/archsetup-boot-upgrade.service.d /usr/local/lib/ug-fakes && systemctl daemon-reload && echo removed' +#+end_src +Expected: the login appeared without waiting out the timeout, the flag is +gone, and =is-failed= prints failed. The record reads result failed with +failed_step boot-transaction, a detail whose last line is the fake's +'error: failed to commit transaction (fake)', and n equal to the armed +count. maint's failed-units row names archsetup-boot-upgrade.service, and +the deferred row lists the armed set at WARN. + +*** M-boot-armed on ratio: APPLY arms, the row shows it, and the offline reboot applies the set +What we're verifying: the panel path the VM can't cover. APPLY should open +=upgrade-guarded --complete= in a detached foot terminal, the row should +read armed with REBOOT shown, and an offline armed boot should apply the +set before the login (AC4's APPLY clause, AC5). +- Run it after step (c) on ratio, on a day the dry-run below lists a + GPU-kind entry. +#+begin_src sh :results output +upgrade-guarded --dry-run 2>&1 | tail -20 +#+end_src +- Open the maint panel. +- Find the deferred row. +- Press APPLY on the deferred row. The first press arms. +- Press APPLY again to fire it. +- Watch the panel's wall and the foot terminal that opens, until the + terminal shows its 'Reboot now?' prompt. +- Press Enter at the prompt, which answers no. +- Wait up to 30 s for the panel's next probe. +- Read the deferred row and the REBOOT key. +- Read the flag. +#+begin_src sh :results output +cat /var/lib/archsetup/apply-upgrade-on-boot +#+end_src +- Take networking down. +#+begin_src sh :results output +nmcli networking off && echo "networking off" +#+end_src +- Press REBOOT on the panel. +- Confirm the reboot the way the panel asks. +- Watch the monitor through the boot until the tty1 password prompt. +- Log in. +- Bring networking back. +#+begin_src sh :results output +nmcli networking on && echo "networking on" +#+end_src +#+begin_src sh :results output +ls /var/lib/archsetup/apply-upgrade-on-boot 2>&1 +jq -c '.data | {result, failed_step, n: (.packages | length)}' ~/.local/state/maint/upgrade_deferred.json +maint status 2>&1 | grep -i -E 'topgrade|deferred' +#+end_src +Expected: the second APPLY press opened a foot terminal running +=upgrade-guarded --complete=, and the panel's wall streamed nothing. The +terminal ended at the prompt with no BLOCKED banner. Before the reboot, the +row read 'armed — A apply at next boot' (with ' · <N−A> more deferred' if +anything else stayed deferred) and REBOOT was shown. The banner showed on +the monitor before the password prompt. Afterwards the flag is gone, the +record reads result ok with n 0, and =maint status= shows topgrade_age +fresh when nothing else is deferred. + +*** M-boot-ratio: the banner and progress show on ratio's monitor, and REBOOT shows while armed +What we're verifying: ratio's own boot path. Its =/dev/console= is ttyS0, +so the unit's tty1 output has to reach the monitor. Its tty1 shows a +password prompt that the =Before=getty@tty1.service= ordering holds. And +REBOOT shows from the flag alone, since reboot_required stays false there +(AC5, and Phase 4's ratio-path confirmation). +- Run it on the boot M-boot-armed on ratio arms, doing the steps up to the + REBOOT press before that test's own REBOOT press; or arm the same way. +#+begin_src sh :results output +ls -l /var/lib/archsetup/apply-upgrade-on-boot +maint status 2>&1 | grep -i reboot +#+end_src +- Read the panel's REBOOT key. +- Press REBOOT on the panel. +- Confirm the reboot the way the panel asks. +- Watch the monitor from the firmware screen to the tty1 password prompt. +- Log in. +#+begin_src sh :results output +systemctl show archsetup-boot-upgrade -p ExecMainExitTimestampMonotonic +systemctl show getty@tty1 -p ExecMainStartTimestampMonotonic +#+end_src +Expected: before the reboot, reboot_required read clear and REBOOT was +shown anyway. The banner and pacman's per-package progress stayed on the +monitor for the whole transaction, and the password prompt appeared only +after it. The unit's exit timestamp is below getty@tty1's start timestamp. + +*** ratio's linux-lts-strix is never moved or gated +What we're verifying: the foreign, headerless GRUB-default kernel stays out +of the pending set and the gate entry on a real =--complete= (Phase 4's +ratio-path confirmation). +- Run it on ratio the first time a =--complete= lands linux or linux-lts. +#+begin_src sh :results output +pacman -Q linux-lts-strix | tee /tmp/ug-strix-before.txt +LC_ALL=C pacman -Qu | grep strix || echo "strix not pending" +#+end_src +- Start =upgrade-guarded --complete= with APPLY, or in a terminal. +- While stage 1 runs, read the gate entry from a second terminal. +#+begin_src sh :results output +jq -c '.data.gate' ~/.local/state/maint/upgrade_deferred.json +#+end_src +- Let the session finish. +#+begin_src sh :results output +pacman -Q linux-lts-strix | diff /tmp/ug-strix-before.txt - && echo "strix unchanged" +grep "$(date +%Y-%m-%d)" /var/log/pacman.log | grep linux-lts-strix || echo "no strix lines today" +#+end_src +Expected: strix wasn't pending; gate.pkgbases during stage 1 listed linux +and/or linux-lts and never linux-lts-strix; strix's version is unchanged; +and no pacman.log line from today names it. + +*** M-boot-armed on velox: APPLY arms, the row shows it, and the offline reboot applies the set +What we're verifying: the same panel path on velox's ZFS root with +autologin. APPLY's detached terminal should run the real gate when a +kernel is pending, the row should read armed with REBOOT shown, and an +offline armed boot should finish before autologin starts Hyprland (AC4's +APPLY clause, AC5). +- Run it after step (c) on velox, on a day the dry-run below lists a + GPU-kind entry. +#+begin_src sh :results output +upgrade-guarded --dry-run 2>&1 | tail -20 +#+end_src +- Open the maint panel. +- Find the deferred row. +- Press APPLY on the deferred row. The first press arms. +- Press APPLY again to fire it. +- Watch the panel's wall and the foot terminal that opens, until the + terminal shows its 'Reboot now?' prompt. +- Press Enter at the prompt, which answers no. +- Wait up to 30 s for the panel's next probe. +- Read the deferred row and the REBOOT key. +- Read the flag. +#+begin_src sh :results output +cat /var/lib/archsetup/apply-upgrade-on-boot +#+end_src +- Take networking down. +#+begin_src sh :results output +nmcli networking off && echo "networking off" +#+end_src +- Press REBOOT on the panel. +- Confirm the reboot the way the panel asks. +- Watch tty1 from the ZFSBootMenu countdown until autologin starts + Hyprland. +- Bring networking back. +#+begin_src sh :results output +nmcli networking on && echo "networking on" +#+end_src +#+begin_src sh :results output +ls /var/lib/archsetup/apply-upgrade-on-boot 2>&1 +jq -c '.data | {result, failed_step, gate, n: (.packages | length)}' ~/.local/state/maint/upgrade_deferred.json +journalctl -b -t archsetup-boot-upgrade --no-pager | head -2 +maint status 2>&1 | grep -i -E 'topgrade|deferred' +#+end_src +Expected: the second APPLY press opened a foot terminal running +=upgrade-guarded --complete=, and the panel's wall streamed nothing. If a +kernel was pending, the terminal showed the gate passing before the arm. +Before the reboot the row read 'armed — A apply at next boot' and REBOOT +was shown. The banner showed on tty1 before Hyprland started, and the +journal's first line is that banner. Afterwards the flag is gone, the +record reads result ok with gate null and n 0, and =maint status= shows +topgrade_age fresh when nothing else is deferred. + +*** M-boot-news on velox: unread Arch news doesn't stall the boot run +What we're verifying: with two or more unread news items at boot, the boot +form's informant clear should keep the transaction from waiting on stdin +(AC7 at boot). +- Run it after step (c) on velox, on a day =upgrade-guarded --dry-run= + lists a GPU-kind entry. +- Open the maint panel at its deferred row. +- Press APPLY on the deferred row. The first press arms. +- Press APPLY again to fire it. +- Press Enter at the terminal's 'Reboot now?' prompt, which answers no. + Arming cleared the news, so the news has to become unread after this. +#+begin_src sh :results output +informant list --unread 2>&1 | head +#+end_src +- If fewer than two items are unread, move informant's state aside so the + whole feed reads unread. The last step puts it back, and no pacman + transaction should run before the reboot. +#+begin_src sh :results output +sudo mv /var/lib/informant.dat /var/lib/informant.dat.bak && informant list --unread 2>&1 | head -3 +#+end_src +- Press REBOOT on the panel. +- Confirm the reboot the way the panel asks. +- Watch tty1 until autologin starts Hyprland. +#+begin_src sh :results output +systemctl show archsetup-boot-upgrade -p Result -p ExecMainStatus +journalctl -b -t archsetup-boot-upgrade --no-pager | tail -5 +jq -c '.data | {result, failed_step, n: (.packages | length)}' ~/.local/state/maint/upgrade_deferred.json +#+end_src +- If you moved informant's state aside, restore it. +#+begin_src sh :results output +[ -e /var/lib/informant.dat.bak ] && sudo mv -f /var/lib/informant.dat.bak /var/lib/informant.dat && echo restored +#+end_src +Expected: the boot run ended with Result=success well inside its timeout, +the journal runs straight from the banner to the end of the transaction +with nothing waiting on input, and the record reads result ok. + +*** D-zbm: recover from a failed zfs-dkms build by rolling back to the held snapshot +What we're verifying: the README's recovery steps work end to end on a ZFS +root. A =--complete= whose zfs build fails should exit 4 with its snapshot +held and named. A ZFSBootMenu rollback (MOD+R) plus the pacman-log +reconcile should then leave a consistent system. This is non-gating; no +criterion waits on it. +- Follow the recovery steps in =maint/README.md= (Phase 4), which carry the + first build task's MOD+R and version fixes. +#+begin_src sh :results output +# VM helpers; every later block in this test sources them. Root logs in by +# key only after the install, so vmroot uses the key the harness left with +# the newest test results. +cat > /tmp/ug-vm.sh <<'EOF' +cd ~/code/archsetup || exit 1 +o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222" +key=$(ls -td test-results/*/root_key 2>/dev/null | head -1) +vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; } +vmroot() { ssh $o -i "$key" root@localhost "$@"; } +mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; } +shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1 + magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; } +shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d" + for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done + for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done + echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; } +EOF +echo "helpers written to /tmp/ug-vm.sh" +#+end_src +#+begin_src sh :results output +# The full install runs first; this takes a while, and Emacs waits on it. +cd ~/code/archsetup && make test-keep FS_PROFILE=zfs 2>&1 | tail -6 +#+end_src +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'pacman -Q less $(pacman -Qqo /usr/lib/modules/*/vmlinuz)' +#+end_src +- On https://archive.archlinux.org/packages/, find the previous releases of + the VM's kernel, its headers and one small leaf package (=less=). +- Paste the three URLs into the next block. +#+begin_src sh :results output +. /tmp/ug-vm.sh +kernel_url='PASTE-URL'; headers_url='PASTE-URL'; leaf_url='PASTE-URL' +vmroot "pacman -U --noconfirm $kernel_url $headers_url $leaf_url" | tail -3 +vm 'LC_ALL=C pacman -Qu' +#+end_src +- Make the zfs build fail for any kernel built from now on. The VM is + disposable after this drill. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vmroot 'for f in /usr/src/zfs-*/dkms.conf; do echo "MAKE[0]=\"false\"" >> "$f"; tail -1 "$f"; done' +#+end_src +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'upgrade-guarded --complete > /tmp/ug-zbm.log 2>&1; echo "exit: $?"; tail -2 /tmp/ug-zbm.log +ls /var/lib/archsetup/apply-upgrade-on-boot 2>&1 +snap=$(jq -r .data.held_snapshot ~/.local/state/maint/upgrade_deferred.json); echo "held: $snap" +zfs holds "$snap"' +#+end_src +- Write down the held snapshot's name; a rolled-back root may not keep it + in the record. +- Reboot the guest and catch ZFSBootMenu inside its 3 s countdown by + sending Escape every half second (the VM takes no keyboard input any + other way). +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'sudo systemctl reboot'; sleep 5 +for i in $(seq 60); do mon sendkey esc >/dev/null; sleep 0.5; done +shot +#+end_src +- If the shot shows anything but the boot-environment list, change the + =sleep= before the loop and run the block again; if it shows the + firmware's setup screen, which also answers Escape, run + =mon system_reset= first. +- Open the snapshot list with MOD+S (Ctrl+S). +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey ctrl-s >/dev/null; sleep 1; shot +#+end_src +- If ZFSBootMenu asks to import the pool read/write, press MOD+W. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey ctrl-w >/dev/null; sleep 1; shot +#+end_src +- Move the selection onto the held snapshot one row at a time, checking + each shot. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey down >/dev/null; sleep 1; shot +#+end_src +- Roll back with MOD+R. Never press ENTER (duplicate), MOD+X (clone and + promote) or MOD+C (clone) here. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey ctrl-r >/dev/null; sleep 1; shot +#+end_src +- Answer the rollback confirmation the screen shows; for a y/N prompt, the + next block sends y. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey y >/dev/null; sleep 2; shot +#+end_src +- Go back to the boot-environment list. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey esc >/dev/null; sleep 1; shot +#+end_src +- Boot =zroot/ROOT/default=. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey ret >/dev/null; sleep 40; shot +#+end_src +- List the pacman.log lines to undo, newest first, with the held name + pasted in. +#+begin_src sh :results output +. /tmp/ug-vm.sh +snap='PASTE-HELD-SNAPSHOT' +vmroot "ls /var/lib/pacman/db.lck 2>&1 +t=\$(date -d @\$(zfs get -Hp -o value creation $snap) +%Y-%m-%dT%H:%M:%S); echo \"since \$t\" +awk -v t=\"\$t\" 'substr(\$1, 2, 19) >= t && \$2 == \"[ALPM]\" && \$3 ~ /^(upgraded|downgraded|installed|removed)\$/' /var/log/pacman.log | tac" +#+end_src +- Undo each listed line, newest first, as the README's recovery steps say. +- Paste the reconciled names into the check block and run it. +#+begin_src sh :results output +. /tmp/ug-vm.sh +names='PASTE-RECONCILED-NAMES' +vmroot "pacman -Q $names; pacman -Dk && echo 'Dk clean'; pacman -Qkk $names 2>&1 | tail -4" +#+end_src +Expected: =--complete= exited 4 with the zfs module item as its last line, +left no flag, and named a held snapshot that =zfs holds= shows with the +=upgrade-guarded= tag. After the rollback and the reconcile, =pacman -Q= +shows the old kernel, headers and =less= versions, =pacman -Dk= is clean, +and =pacman -Qkk= over the reconciled names reports no mismatch. + +*** D-zbm power-cut variant: a reset during stage 1 leaves the right snapshot held and a reconcilable db +What we're verifying: a hard reset during =--complete='s stage 1, on a VM +with no open gate, should leave the snapshot step 3 took held and named. +Recovery's power-cut steps, including the half-registered package, should +then reconcile the db. This is non-gating. +- Use a fresh =make test-keep FS_PROFILE=zfs= VM, not the one D-zbm left + with its gate open. +#+begin_src sh :results output +# VM helpers; every later block in this test sources them. Root logs in by +# key only after the install, so vmroot uses the key the harness left with +# the newest test results. +cat > /tmp/ug-vm.sh <<'EOF' +cd ~/code/archsetup || exit 1 +o="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=10 -p 2222" +key=$(ls -td test-results/*/root_key 2>/dev/null | head -1) +vm() { sshpass -p archsetup ssh $o cjennings@localhost "$@"; } +vmroot() { ssh $o -i "$key" root@localhost "$@"; } +mon() { echo "$*" | socat - UNIX-CONNECT:vm-images/qemu-monitor-zfs.sock; } +shot() { mon screendump /tmp/ug-tty1.ppm >/dev/null; sleep 1 + magick /tmp/ug-tty1.ppm /tmp/ug-tty1.png && echo /tmp/ug-tty1.png; } +shots() { d=/tmp/ug-shots; rm -rf "$d"; mkdir -p "$d" + for i in $(seq -w 1 "${1:-90}"); do mon "screendump $d/$i.ppm" >/dev/null; sleep 2; done + for f in "$d"/*.ppm; do magick "$f" "${f%.ppm}.png" && rm -f "$f"; done + echo "$(ls "$d" | wc -l) frames, 2 s apart, in $d"; } +EOF +echo "helpers written to /tmp/ug-vm.sh" +#+end_src +#+begin_src sh :results output +# The full install runs first; this takes a while, and Emacs waits on it. +cd ~/code/archsetup && make test-keep FS_PROFILE=zfs 2>&1 | tail -6 +#+end_src +- On https://archive.archlinux.org/packages/, find the previous releases of + the VM's kernel and its headers. +- Paste the two URLs into the next block. +#+begin_src sh :results output +. /tmp/ug-vm.sh +kernel_url='PASTE-URL'; headers_url='PASTE-URL' +vmroot "pacman -U --noconfirm $kernel_url $headers_url" | tail -3 +vm 'LC_ALL=C pacman -Qu' +#+end_src +- In a separate terminal, start the session with a tty. pacman prints + '(n/m) upgrading <pkg>' only on a tty; without one it prints 'upgrading + <pkg>...'. +#+begin_src sh :eval no +sshpass -p archsetup ssh -t -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -p 2222 cjennings@localhost upgrade-guarded --complete +#+end_src +- As soon as stage 1's first upgrading line appears, read the held + snapshot's name. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vm 'jq -r .data.held_snapshot ~/.local/state/maint/upgrade_deferred.json' +#+end_src +- As the kernel headers package's upgrading line appears, run the next + block. It hard-resets the guest, then sends Escape every half second to + catch ZFSBootMenu's 3 s countdown. Never use =system_powerdown= or + =quit=: both end QEMU, and =make test-keep= then restores the clean + install. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon system_reset >/dev/null; sleep 3 +for i in $(seq 60); do mon sendkey esc >/dev/null; sleep 0.5; done +shot +#+end_src +- If the shot shows anything but the boot-environment list, run + =mon system_reset= and the Escape loop again with a changed delay. +- Open the snapshot list with MOD+S (Ctrl+S). +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey ctrl-s >/dev/null; sleep 1; shot +#+end_src +- If ZFSBootMenu asks to import the pool read/write, press MOD+W. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey ctrl-w >/dev/null; sleep 1; shot +#+end_src +- Move the selection onto the held snapshot one row at a time, checking + each shot. If its name wasn't read in time, it's the + =zroot/ROOT/default@pre-pacman_*= row taken just before stage 1. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey down >/dev/null; sleep 1; shot +#+end_src +- Roll back with MOD+R. Never press ENTER (duplicate), MOD+X (clone and + promote) or MOD+C (clone) here. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey ctrl-r >/dev/null; sleep 1; shot +#+end_src +- Answer the rollback confirmation the screen shows; for a y/N prompt, the + next block sends y. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey y >/dev/null; sleep 2; shot +#+end_src +- Go back to the boot-environment list. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey esc >/dev/null; sleep 1; shot +#+end_src +- Boot =zroot/ROOT/default=. +#+begin_src sh :results output +. /tmp/ug-vm.sh +mon sendkey ret >/dev/null; sleep 40; shot +#+end_src +- Read the hold, the lock and the db's state, with the held name pasted in. +#+begin_src sh :results output +. /tmp/ug-vm.sh +snap='PASTE-HELD-SNAPSHOT' +vmroot "zfs holds $snap; ls /var/lib/pacman/db.lck 2>&1; pacman -Dk 2>&1 | tail -3; tail -8 /var/log/pacman.log" +#+end_src +- If pacman.log shows the reset landed outside extraction (every target + has an =[ALPM]= line, or 'running post-transaction hooks' is logged), + stop and repeat on a fresh VM. +- Remove =db.lck= if the read above showed it. +#+begin_src sh :results output +. /tmp/ug-vm.sh +vmroot 'rm -fv /var/lib/pacman/db.lck' +#+end_src +- Delete the local db entry =pacman -Dk= reported as 'description file is + missing', with its directory name pasted in. +#+begin_src sh :results output +. /tmp/ug-vm.sh +entry='PASTE-ENTRY-DIR' +vmroot "rm -rv /var/lib/pacman/local/$entry" +#+end_src +- Find the version the booted root holds for that package: the last + =[ALPM]= line for it before the snapshot's creation (on an upgraded or + downgraded line, the version after =->=; on a removed line, or with no + line, register nothing and skip the next step). +#+begin_src sh :results output +. /tmp/ug-vm.sh +snap='PASTE-HELD-SNAPSHOT'; pkg='PASTE-PACKAGE' +vmroot "t=\$(date -d @\$(zfs get -Hp -o value creation $snap) +%Y-%m-%dT%H:%M:%S) +awk -v t=\"\$t\" -v p=$pkg 'substr(\$1, 2, 19) < t && \$2 == \"[ALPM]\" && \$4 == p' /var/log/pacman.log | tail -1" +#+end_src +- Register that version from the cache, with its package file pasted in. +#+begin_src sh :results output +. /tmp/ug-vm.sh +pkgfile='PASTE-CACHE-FILE' +vmroot "pacman -U --dbonly --noconfirm /var/cache/pacman/pkg/$pkgfile" +#+end_src +- List the later =[ALPM]= lines to undo, newest first. +#+begin_src sh :results output +. /tmp/ug-vm.sh +snap='PASTE-HELD-SNAPSHOT' +vmroot "t=\$(date -d @\$(zfs get -Hp -o value creation $snap) +%Y-%m-%dT%H:%M:%S); echo \"since \$t\" +awk -v t=\"\$t\" 'substr(\$1, 2, 19) >= t && \$2 == \"[ALPM]\" && \$3 ~ /^(upgraded|downgraded|installed|removed)\$/' /var/log/pacman.log | tac" +#+end_src +- Undo each listed line, newest first, as the README's recovery steps say. +- Paste the reconciled names into the check block and run it. +#+begin_src sh :results output +. /tmp/ug-vm.sh +names='PASTE-RECONCILED-NAMES' +vmroot "pacman -Q $names; pacman -Dk && echo 'Dk clean'; pacman -Qkk $names 2>&1 | tail -4" +#+end_src +Expected: =zfs holds= showed the =upgrade-guarded= tag on the snapshot step +3 took before stage 1, and held_snapshot named it. After the reconcile, +=pacman -Q= shows the old kernel set, =pacman -Dk= is clean, and =pacman +-Qkk= over the reconciled names reports no mismatch, including the package +cut off mid-extraction. + ** DOING [#B] Prepare for GitHub open-source release :PROPERTIES: :LAST_REVIEWED: 2026-08-17 |
