From 32008ac10409e1a6f8ddb8820c6e8ac30e575db4 Mon Sep 17 00:00:00 2001 From: Craig Jennings Date: Mon, 5 Oct 2026 23:11:01 -0500 Subject: docs(spec): review the guarded-upgrade spec and decompose its build I ran three review rounds against the live code in archsetup and the dotfiles. They found 129 issues, and the blocking count fell from 13 to 4 to 0. Each one is dispositioned in the spec. The exact contract now lives only in Implementation phases, because the copies in Design, Decisions and Testing had drifted. The spec moves from DRAFT through READY to DOING, with nine non-blocking residuals left for the first build task. The parent task now carries 22 build tasks in commit-group order, ending with the IMPLEMENTED flip. The manual tests are under Manual testing and validation, and a stale-kernel age metric is filed as [#D]. --- .../2026-08-25-topgrade-guarded-upgrade-spec.org | 5611 +++++++++++++++++++- todo.org | 1655 +++++- 2 files changed, 7168 insertions(+), 98 deletions(-) diff --git a/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org b/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org index d9ec8d4..e19ea6c 100644 --- a/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org +++ b/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org @@ -4,17 +4,25 @@ #+TODO: TODO | DONE #+TODO: DRAFT READY DOING | IMPLEMENTED SUPERSEDED CANCELLED -* DRAFT Guarded-upgrade completion +* DOING Guarded-upgrade completion :PROPERTIES: :ID: 81cdfd72-db96-43d3-aa03-779878c99f3e :END: +- [2026-10-05 Mon @ 13:00 -0500] READY → DOING. Decomposed into build tasks under the archsetup todo.org parent "Topgrade guarded-upgrade build" (SPEC_ID 81cdfd72-db96-43d3-aa03-779878c99f3e), with the manual tests filed under Manual testing and validation. +- [2026-10-05 Mon @ 12:10 -0500] DRAFT → READY. Three review rounds incorporated (48, 59 and 22 findings; blocking went 13, 4, 0), all 129 dispositioned. 9 non-blocking residuals from the last repair pass stay open under Review findings and are carried into the build tasks. +- [2026-10-05 Mon @ 11:10 -0500] spec-review round 3 on the consolidated spec: no blocking findings. 22 non-blocking findings recorded and being resolved before READY. +- [2026-10-05 Mon @ 10:40 -0500] round-2 review incorporated: all 59 findings dispositioned, the exact contract consolidated into Implementation phases, and machinery without a caller cut. +- [2026-10-05 Mon @ 09:45 -0500] spec-review round 2 on the revised spec: Not ready. 59 new findings recorded, 4 of them blocking. +- [2026-10-05 Mon @ 09:30 -0500] reworded the bodies of Decisions 2 to 7 to match the rewritten design; I reopened none of them, and Decision 1 is unchanged. Decision 2: the TTY fallback holds the kernel and DKMS sets. Decision 3: the TOML is a divergent list, rewritten to equal the hook and pinned by an archsetup test, and the rollout is three code commits. Decision 4: it names the record, which is always written (except on the EUID 0 and lock-held refusals), states the stamp predicate per form, gives each test suite a copy of one fixture, and its example is now 6 deferred. Decision 5: the DKMS set joins the hold, with =zfs-utils= joining through the closure; the gate covers every kernel whose modules the run changed, plus open failures; ratio has three kernels; the DKMS hook timing is corrected. Decision 6: the boot form is =upgrade-guarded --apply-armed=, which installs the flag's list from the cache that arming pre-downloaded, with no refresh and no network, and the armed everyday run's disarm cases are split by cause. Decision 7: the flag is root-owned, carries the armed === list, and =ExecStartPre= moves it off its path at the start of the attempt. +- [2026-10-05 Mon @ 06:05 -0500] spec-review: Not ready. 48 findings recorded under Review findings, 13 of them blocking; status stays DRAFT until they are dispositioned. +- [2026-10-05 Mon @ 05:14 -0500] named the split script =upgrade-guarded= and added =containers= to its topgrade =--disable= list (the containers step fails on locally built images, hit again on velox 2026-10-04). - [2026-08-25 Tue @ 18:45 -0600] decisions closed 7/7. The kernel decision reversed on the velox DKMS failure chain: held on every everyday run, landed only in the dedicated session behind a DKMS/initramfs/snapshot gate. - [2026-08-25 Tue @ 18:30 -0600] redirected: the everyday path is a live split upgrade (apply everything the guard would not block, defer the rest); the boot-time oneshot becomes the completion step for the deferred set. Decided while running exactly that by hand on ratio. - [2026-08-25 Tue @ 06:39:42 -0600] drafted. Grounded in a live read of the maint engine, the pacman hooks, and the boot path on velox, not memory. The topgrade-freshness diagnosis that motivates it is in this session's log. * Metadata -| Status | draft | +| Status | doing | |----------+-------------------------------------------------------------| | Owner | Craig Jennings | |----------+-------------------------------------------------------------| @@ -24,78 +32,127 @@ * Summary -The waybar maintenance module shows topgrade freshness as permanently stale. The cause is a real one: on a machine running Hyprland, a full =topgrade= almost never exits 0, because its system step upgrades GPU/compositor libraries that the =hypr-live-update-guard= pacman hook correctly refuses to swap under a live session. The freshness stamp is gated on topgrade's exit code, so a correct, protective refusal reads as "you never run updates." This spec designs a safe path to actually complete a guarded upgrade, and makes that completion record the freshness stamp, so the metric tracks the true state of the system. +The waybar maintenance module shows topgrade freshness as permanently stale. The cause is a real one: on a machine running Hyprland, a full =topgrade= almost never exits 0, because its system step upgrades GPU/compositor libraries that the =hypr-live-update-guard= pacman hook correctly refuses to swap under a live session. The freshness stamp is gated on topgrade's exit code, so a correct, protective refusal reads as "you never run updates." + +This spec designs a safe path to actually complete a guarded upgrade, and makes that completion record the freshness stamp, so the metric tracks the true state of the system. The everyday path is a live split run (=upgrade-guarded=) that applies everything outside its held set and records what it deferred: the GPU/compositor set while Hyprland runs and the kernel and DKMS sets always, together with any pending package that cannot move without them. The deferred set lands in a dedicated =upgrade-guarded --complete= session. The kernel and DKMS sets go in behind the =kernel-modules-check= gate, and the GPU/compositor set installs directly when Hyprland is stopped, or, when Hyprland is live, is pre-downloaded and finished by a boot-time oneshot from the package cache. + +Implementation phases is the one exact statement of behavior. Every other section explains the shape or the reasons and cites the phase that holds the detail. * Problem / Context -The metric reads one cache key, =topgrade_run= (=~/.local/state/maint/topgrade_run.json=). Absent, the probe (=maint/src/maint/probes/updates.py:114=) returns WARN, "no topgrade run recorded". Two writers stamp it: the =topgrade= PATH wrapper (=~/.dotfiles/hyprland/.local/bin/topgrade=) on =rc -eq 0=, and the panel's TOPGRADE lever (=doctor.py=), which returns before the stamp on any non-zero exit. The read path is sound (a sandboxed =maint stamp topgrade= writes the file and =maint status= then reads freshness 0); the file is simply never written. +The waybar maintenance module's topgrade freshness comes from one cache key, =topgrade_run= (=~/.local/state/maint/topgrade_run.json=), read by =topgrade_freshness= in dotfiles =maint/src/maint/probes/updates.py:114=. It grades the stamp's age in days against =topgrade_warn_days = 14=. With no stamp at all it returns WARN, "no topgrade run recorded", a branch that only fires on a machine that has never stamped. Two writers stamp it today: the =topgrade= PATH wrapper (=~/.dotfiles/hyprland/.local/bin/topgrade=) when the real binary exits 0, and the panel's TOPGRADE lever (=doctor.py=), which returns before its stamp on any non-zero exit. + +The read path is sound: a sandboxed =maint stamp topgrade= writes the file, and =maint status= then reads freshness 0. The fault is that the file is almost never refreshed. On ratio it was last written 2026-07-08, the day the wrapper landed, so on 2026-10-05 the probe graded an age of just under 90 days and warned. + +Topgrade rarely exits 0 on a machine running Hyprland, and the reason is specific rather than flaky. The =hypr-live-update-guard= pacman hook is a =PreTransaction=/=AbortOnFail= hook: when Hyprland is running and an upgrade changes the on-disk version of a GPU/compositor library, it prints a BLOCKED banner and exits 1, aborting the whole transaction before any file is swapped. Its 15 =Target= patterns (mesa and its subpackages, wayland, libdrm, libglvnd, hyprland and some of its core libraries, the Vulkan and NVIDIA userspace packages, xorg-xwayland) are the list =upgrade-guarded= later reads as its GPU patterns. The guard exists for a proven failure: replacing those libraries under a live compositor makes the next GPU call hit a now-deleted mapping and SIGABRT, taking every Wayland client down (hit on ratio 2026-06-07). -It is never written because topgrade rarely exits 0 on this machine, and the reason is specific rather than flaky. =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook= is a =PreTransaction=/=AbortOnFail= hook that, when Hyprland is running and an upgrade changes the on-disk version of a GPU/compositor library, prints a BLOCKED banner and exits 1 — aborting the whole transaction before any file is swapped. Its trigger set is =mesa=, =mesa-*=, =wayland=, =libdrm=, =libglvnd=, =hyprland=, =aquamarine=, =hyprutils=, =hyprgraphics=, =vulkan-radeon=, =vulkan-intel=, =vulkan-mesa-layers=, =nvidia-utils=, =lib32-nvidia-utils=, =xorg-xwayland=. The guard exists for a proven failure: replacing those libraries under a live compositor makes the next GPU call hit a now-deleted mapping and SIGABRT, taking every Wayland client down (hit on ratio 2026-06-07). +The installer writes the hook as =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook=, but both daily drivers still carry it under the legacy unprefixed name. Phase 4's rollout moves each machine to the =10-= name before its levers route through the script. -So when any of those libraries has an update pending — a frequent event — topgrade's =system= step (it runs =yay=) aborts non-zero, topgrade returns non-zero, and neither writer stamps. The observed case: on 2026-08-24 topgrade ran at 17:49, hit the guard on =mesa= (26.1.7 → 26.2.1), and failed; the upgrade was then finished by hand with the guard's sentinel override, entirely outside the wrapper, so nothing stamped. The metric has read stale ever since. +So when any of those libraries has an update pending, which is often, topgrade's =system= step (it runs =yay=) aborts, topgrade exits non-zero, and neither writer stamps. The observed case: on 2026-08-24 topgrade ran at 17:49, hit the guard on =mesa= (26.1.7 → 26.2.1) and failed. The upgrade was then finished by hand with the guard's sentinel override, outside the wrapper, so nothing stamped. The metric had already gone stale in late July and stayed stale. -Two framings of the fix are in tension, and choosing between them is the spec's central decision. Either the metric means "how recently did you run the sweep" (recency), so the stamp should decouple from topgrade's exit; or it means "is the system up to date" (state), so staying stale while a guarded upgrade is deferred is *correct* and the only real defect is that safely completing that upgrade doesn't stamp. This spec takes the state framing (see Decisions). +The fix has two framings in tension. If the metric means "how recently did you run the sweep" (recency), the stamp should decouple from topgrade's exit. If it means "is the system up to date" (state), staying stale while an upgrade is deferred is correct, and the only real defect is that safely completing the deferred upgrade doesn't stamp. I take the state framing (Decision 1, "Metric means state, not recency"), so the rest of the spec builds a safe completion path and makes that completion the stamp. * Goals and Non-Goals ** Goals -- A safe, low-friction way to apply a guarded (GPU/compositor-library) upgrade, with Hyprland not live at swap time. -- That completion records the =topgrade_run= freshness stamp, so the metric clears when the system is genuinely current. +- An everyday UPDATE or TOPGRADE that completes live. It applies everything outside its held set and defers the GPU/compositor set (while Hyprland runs), the kernel and DKMS sets (always), and any pending package that can't move without them. A deferral alone never fails the run, and the deferred set stays visible in maint until it lands. +- A safe, low-friction way to land the deferred set: a dedicated =upgrade-guarded --complete= session that lands the kernel and DKMS sets behind the =kernel-modules-check= gate, offers no reboot while that gate is open, and applies the GPU/compositor set with Hyprland not live at swap time (at the next boot, or directly from a console with Hyprland stopped). +- =upgrade-guarded= writes the =topgrade_run= freshness stamp only when a run leaves nothing deferred and no gate open, so the metric clears exactly when the system is current and stays honestly stale while a protective deferral is outstanding. The per-mode predicate is in Phase 1. - A boot-time upgrade path that can never lock the machine out of its session, however it fails. -- The installer owns the durable pieces so a rebuilt machine has them without hand-setup. +- The installer owns the durable pieces, so a rebuilt machine has them without hand-setup. ** Non-Goals -- Weakening or bypassing the =hypr-live-update-guard= hook. It stays exactly as strict; this builds *around* it, not through it. -- Making the full topgrade ecosystem sweep (git repos, vim, npm, ...) run at boot. Those never need a stopped compositor and are out of the boot path. -- Changing how the kernel hazard is *guarded*. The hook stays silent on kernels; the split script holds them back on a live run as a second, separately-reasoned list (see Design), which is a deferral policy rather than a guard. -- A general offline-update system for all of pacman. Scope is the guarded-library case. +- Weakening or bypassing the =hypr-live-update-guard= hook. I keep it exactly as strict and build around it, not through it. +- Running the topgrade ecosystem sweep (vim, npm, ...) at boot. None of it needs a stopped compositor, so it stays on the live path. +- Changing how the kernel hazard is guarded. The hook stays silent on kernels. Holding the kernel and DKMS sets on every everyday run is a deferral policy inside =upgrade-guarded=, a separate list with its own reasons (Design), not a guard. An AUR upgrade that needs a newer kernel-set or DKMS-set package is the one ungated path; the everyday run detects it afterwards and opens the gate entry (Risks). +- A general offline-update system for pacman. The boot form installs only the armed list (the GPU/compositor packages pre-downloaded at arming) from the package cache. Everything else lands live. ** Scope tiers -- v1: the split-upgrade script (live: apply the non-blocked remainder, defer the rest, run the ecosystem sweep with the system step off, report the deferred set); maint's UPDATE/TOPGRADE levers route through it; an "apply on reboot" affordance that installs the held kernel live and arms the boot-time oneshot for the GPU/compositor set. -- Out of scope: full-sweep-at-boot; touching the guard's policy. -- vNext: none open — the kernel deferral that was vNext is now part of v1's held set. +- v1: everything in Implementation phases. That is the two scripts and their installer step (Phase 1), maint's repointed levers with the deferred row and its APPLY action (Phase 2), the =archsetup-boot-upgrade.service= oneshot (Phase 3), and the docs and rollout on both daily drivers (Phase 4). +- Out of scope: everything under Non-Goals. +- vNext: a stale-kernel age in maint (how long the kernel set has been held, graded so an overdue dedicated session shows up), filed in =todo.org= as a [#D] task when the phases are decomposed. * Design -The shape follows one principle: the only part of topgrade that needs a stopped compositor is its =system= step when a guarded library is pending. Everything else runs fine live and rarely fails. So the safe path is small and targeted — apply the guarded system upgrade with Hyprland down, once, and stamp it — while the ordinary full sweep stays a normal live =topgrade= run. +The shape follows one principle: the only packages that need a stopped compositor are the GPU/compositor libraries the guard protects, and the only ones that need a moment I chose are the kernel and its DKMS modules. Everything else runs fine live and rarely fails. So the safe path is small and targeted: land the kernel and DKMS sets behind a gate in a session I start on purpose, apply the GPU/compositor set once with Hyprland down, and stamp only when nothing is left deferred. Everything else, the ecosystem sweep included, runs live through the split script, which leaves those sets behind. + +Three pieces, but the everyday gesture is not a reboot. The split script's everyday run is a normal live update that leaves the dangerous few behind. Its dedicated =--complete= session lands them on a day I choose. A boot-time oneshot runs its =--apply-armed= form to finish the GPU/compositor part before anything maps those libraries. This section gives the shape and the reasons. Implementation phases holds the exact behavior, each piece below names the phase subsection that governs it, and where the two differ, the phase bullet is normative. + +*The sets.* Every set the script works with is derived from the system at run time, never hardcoded. The pending set is what pacman reports as upgradable, minus IgnorePkg rows. The GPU patterns are the =Target= lines of =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook=, taken verbatim with their globs (15 today), so the guard and the script share one source of truth. The script reads only the =10-= name, because the legacy unprefixed name sorts after =60-mkinitcpio-remove=, so a blocked bare =pacman -Syu= could delete the initramfs before the guard aborts. A missing hook, or one with no Target lines, makes every mode but the boot form refuse before any transaction rather than run with an empty list, which is deliberately unlike maint's =guard.trips=. + +The kernel set is the packages that own an installed kernel image, plus each kernel's headers where installed; firmware and API-header packages are never members. It is held whole, so a kernel and its headers move together, and only pending members ever move, so a foreign kernel such as ratio's =linux-lts-strix= (a local build with no headers and no zfs module, and the GRUB default) is never touched. The DKMS set is the packages that ship a =dkms.conf=, today =zfs-dkms=. Its exact pin, =zfs-utils=, isn't listed by name: the dependency closure adds it exactly when an upstream bump breaks the pin, and never for a pkgrel-only bump. I hold the kernel and DKMS sets on every everyday run, on both machines, because a failed DKMS rebuild on velox's ZFS root leaves the machine unbootable (see the kernel decision). + +Live means Hyprland is running, by the guard's own =pgrep -x Hyprland= test, so a Ctrl+Alt+F2 console beside a live session counts as live. Every held entry is kernel-kind (the kernel and DKMS sets and whatever the closure ties to them) or GPU-kind (the pending GPU-pattern matches and whatever the closure ties to them). The everyday run always holds the kernel side, and holds the GPU side only when live. The deferred set is the held entries still pending when a run ends. It is exactly the record's package list, and the record, the panel row and the armed list each cover the whole closure, not only the pattern matches. Exact behavior: Phase 1, Sets. + +*The split script.* A pacman =PreTransaction= hook can only abort or allow the transaction it is handed; it can't drop targets from it. So "upgrade everything except the held sets" can't live in the hook. It lives one layer up, in =upgrade-guarded= (archsetup =scripts/upgrade-guarded=), which the installer step that installs =hypr-live-update-guard= puts in =/usr/local/bin= along with its gate, =kernel-modules-check=. The script runs as the invoking user and refuses root, because the record and the stamp live in that user's =$HOME= and yay refuses root. Every root step the script itself runs goes through =sudo -n=, so a machine without its NOPASSWD rule fails the step with a reason rather than wedge on a password prompt the panel's runner can't answer. yay and the sweep escalate through their own sudo, but only after the refresh has succeeded under =sudo -n=. Each invocation runs exactly one mode: the everyday run, =--complete=, =--apply-armed= (the boot unit's alone) or =--dry-run=. The only option is =--no-topgrade=, and only in the everyday run. + +=--dry-run= is the preview. It prints what a run would hold, defer and execute, writes nothing, uses no sudo and takes no lock, and it resolves against a private sync db of its own, so it never touches the system db or the shared one another =checkupdates= caller may be using. Every other mode takes an exclusive lock beside the record, so a second run refuses instead of interleaving, and refuses while pacman's =db.lck= exists. The script never deletes =db.lck=: a stale lock can belong to a pacman that is still running, and only a person can tell. Exact behavior: Phase 1, Identity and privilege; Phase 1, Modes and options; Phase 1, Preconditions. + +*The everyday run.* Both panel levers call it. It captures the unread Arch news first, because yay judges news against installed build dates, and after the upgrade the news would read as old. It then clears informant's news hook, which would otherwise abort a =--noconfirm= transaction. The clear needs root, =--all= to run without prompting, and a timeout, because informant's feed fetch has none of its own. A failed clear is tolerated and leaves informant's hook live for that run, a residual I accept. Then come one refresh, the dependency closure over the held sets, and one =pacman -Su= against that same db with an =--ignore= per held entry and no second =-y=. yay upgrades the AUR packages next, and the ecosystem sweep follows unless =--no-topgrade= is given. The sweep calls =/usr/bin/topgrade= by absolute path, so the stamping =~/.local/bin/topgrade= PATH wrapper is never in the chain. It disables topgrade's =system= step (the script has already done that work), =git_repos= and =containers= (whose step fails on locally built images), each as its own argv element, because topgrade 17.12.2's parser rejects the comma-joined list, and =--no-ask-retry= keeps a failed step from waiting on stdin, since the panel's runner has no tty. A deferral alone never makes the exit non-zero; a failed step or a refusal does. Exact behavior: Phase 1, Everyday sequence. + +The guard hook stays installed, unchanged, as the backstop for a bare =pacman -Syu= typed at a shell. On the driven path it can still fire inside yay's own pacman call; Phase 1, Everyday sequence, says when, and Risks names the kernel-side case that yay can install ungated. Proof of concept: the guard-set half of this run, done by hand on ratio on 2026-08-25 while Hyprland was live, before the kernel and DKMS hold and the containers disable were added, resolved 724 of 730 pending packages (Emacs 31.1 among them) with the six guard hits deferred, after one unrelated fix, an orphaned =qemu-block-gluster= that had been dropped from the repo. That run upgraded linux, linux-lts, their headers, zfs-dkms and zfs-utils live, so under v1 the same day would have deferred at least twelve, and the orphan would have made the closure refuse with its name before any transaction. + +*The dependency closure.* =--ignore= alone isn't enough. pacman still enforces declared dependencies, so a pending package that needs a held package's new version, or whose upgrade breaks a versioned dependency of a held package, makes the transaction unresolvable, and under =--noconfirm= one unresolvable dependent fails the whole transaction, because pacman's skip prompt defaults to no. So before any transaction the script asks pacman what else has to wait, through a read-only print-mode probe against the run's one refreshed db, and holds those packages too until the probe resolves. When the answer isn't something it can safely hold around, such as an upgrade that would break an installed package that isn't pending (an orphan, foreign or AUR package, like =qemu-block-gluster= above), it refuses before any transaction and names what it found, rather than hold a package and its dependents indefinitely without saying why. A refusal never removes or forces anything. Exact behavior: Phase 1, Dependency closure. + +*The record and the stamp.* A deferred set nobody surfaces is a set that silently never lands, so the script records it durably in =${MAINT_STATE_DIR:-$HOME/.local/state/maint}/upgrade_deferred.json= (maint cache key =upgrade_deferred=), in maint's cache envelope, written atomically. =upgrade-guarded= is its only writer. maint's =upgrade_deferred= probe, =topgrade_freshness= and the script's own next run read it. Every mode but =--dry-run= writes it once as it finishes, refusals and failures included, so the panel always reflects the last run. =--complete= and =--apply-armed= also write it before their risky transactions, and the interrupt trap writes it too, so a kill or a power cut still leaves an accurate record with the remedy. A usage error, a root invocation and a held lock never touch it. The gate entry and the held snapshot carry forward through every mode, and only =--complete= changes them, except that an everyday run whose AUR step moved a kernel-side package opens the entry, holds its snapshot and writes the record at once (Phase 1, Invariants). "The gate is open" means the record holds a gate entry. Exact behavior: Phase 1, The record; Phase 1, Finish. + +=topgrade_run= keeps meaning "the system is current", so on the driven path =upgrade-guarded= is its only writer, and it stamps only when a run ends with nothing deferred and no gate open. The everyday run also needs the refresh, the upgrade, yay and the sweep all to have succeeded, so UPDATE, which skips the sweep, never stamps. =--complete= stamps only when it had something to complete, and never when it arms the boot unit, because the GPU set is still pending then. The stamp goes through maint by absolute path, because a system unit's PATH lacks =~/.local/bin=, and a missing maint or a failed stamp fails the run and is recorded, never silenced. doctor.py's lever stamp is deleted. The PATH wrapper keeps stamping a bare-shell =topgrade=, and the script never goes through it. Exact behavior: Phase 1, Stamp predicate. + +*For the user.* UPDATE runs =upgrade-guarded --no-topgrade= and TOPGRADE runs =upgrade-guarded=, and the =sysupgrade= alias runs the script wherever it is installed. On the common day both levers succeed. The panel's deferred row counts what the script held back, QUEUE tags the held names, and the strip's pending cell adds the held count. A deferral alone doesn't raise the row's severity or colour the waybar glyph, but freshness keeps aging until the deferred set lands, and after a TOPGRADE that left anything deferred the wall note says so and points at APPLY. A failed step, a refusal or an interruption shows on the row at WARN with its reason, and so do fixable advisories on deferred packages, which count against the row rather than against UPDATE. Arch news the run captured shows in the row's evidence until the row's DISMISS key clears it. While Hyprland isn't running (a TTY with no compositor, not a Ctrl+Alt+F2 console beside a live session), the everyday run holds only the kernel side, and the GPU/compositor set lands in the same run. -Three pieces, at two altitudes — but the everyday gesture is not a reboot. It is a normal live update that simply leaves the dangerous few behind. +Because the levers no longer arm a live swap, the guard's press-again refusal and the =--force= override go away. A tripped guard read only annotates the arm line with what the run will at least defer. The panel reads everything here from the record and the flag on its next probe, never from an exit code. Exact behavior: Phase 2 (Levers, Guard UX, The deferred row, DISMISS, Freshness and REBOOT, and CVE, QUEUE and the strip). -*The split script.* A pacman =PreTransaction= hook can only abort or allow the transaction it is handed; it cannot drop targets from it. So "upgrade everything except the guarded set" cannot live in the hook — it lives one layer up, in a script the panel calls. On a live run the script: refreshes the sync db and reads the pending set (=checkupdates=); computes the *blocked set* = the guard's own trigger list (read from the installed hook's =Target= lines, so there is one source of truth, and version-aware the way the guard is — a same-version reinstall is not a swap) plus the *kernel set* (every installed kernel with its =-headers=, always as a set; held on every everyday run because a failed DKMS rebuild on velox's ZFS root leaves the machine unbootable — see the kernel decision); clears the news hook (=informant read=) where installed; runs =pacman -Syu --noconfirm --ignore==; runs the AUR-only remainder (=yay -Sua --noconfirm=, AUR packages pinning a guarded version hold themselves back); then runs =topgrade --disable system,git_repos -y= so the other ecosystems still get their sweep and topgrade can actually exit 0. It writes the deferred set to a state file the panel reads, and exits 0 when the live part succeeded, whatever was deferred. The guard hook stays installed as the backstop for a bare =pacman -Syu= typed at a shell; on the driven path it never fires. Proof of concept: this exact sequence, run by hand on ratio on 2026-08-25 while Hyprland was live, resolved 724 of 730 pending packages (Emacs 31.1 among them) with the six guard hits deferred — after one unrelated fix, an orphaned =qemu-block-gluster= that had been dropped from the repo. +*The dedicated session.* Landing the deferred set is =upgrade-guarded --complete=, a session I start on purpose in a terminal: the one APPLY opens, detached so closing the panel can't kill or orphan the kernel transaction, or any terminal or TTY. APPLY is offered on the deferred row while anything is deferred or the gate is open, on the freshness card while anything is deferred, and through REVIEW & FIX, and every route opens that same detached terminal. The session's first stage lands the kernel and DKMS sets live, with the desktop up for diagnosing, along with every pending package outside the GPU closure. Old modules are removed PreTransaction and rebuilt PostTransaction, so a failed DKMS build aborts nothing; the gate after stage 1 is what catches it. If the GPU closure would reach a kernel-side package, the session refuses before any transaction, because landing it would put a kernel behind the boot form or ahead of the gate. -*For the user.* UPDATE and TOPGRADE on the panel run the split script; they succeed, and the panel shows "N deferred" when the script held anything back. Landing the deferred set is a dedicated session, chosen on purpose, run in the foreground from the panel's action or =guarded-upgrade --complete= in a terminal: first the kernel set, live, with the desktop still up; then the gate — every DKMS module built for the new kernel, a fresh initramfs, and on a ZFS root a pre-pacman snapshot to fall back on. If the gate fails the script stops there, names what failed, and does not reboot; the machine keeps running on the old kernel and the desktop is available for the fix. If it passes, the script arms a persistent flag for the GPU/compositor set and offers to reboot (or, from a TTY with no compositor, applies that set directly). On the next boot, before the autologin shell starts Hyprland, the deferred guarded upgrade runs in the console — the guard passes freely because nothing is live — the stamp is written, the flag is cleared, and boot continues into the session. No second reboot: the libraries are already current before anything maps them. If anything goes wrong, the machine still boots into Hyprland and the panel still shows the pending upgrade, so you are never worse off than before arming. +Two things happen before stage 1 touches anything, and on a ZFS root a third, the snapshot hold described below. Right after its refresh, the session removes any arm flag, and only its own GPU step can re-arm. Then, when a kernel-side package is pending or a gate is already open, it records the gate entry: which kernels are unverified, by pkgbase, and when the entry opened. A kill, a crash or a power cut mid-transaction therefore leaves the record naming the unverified kernels. From that moment until a gate passes, the deferred row is CRIT with REBOOT hidden, no arm flag exists, and the session never offers a reboot. The entry closes only when a gate passes, or when no kernel it names is still installed. -*For the implementer.* A persistent arm flag (a file on a non-tmpfs path, e.g. =/var/lib/archsetup/apply-upgrade-on-boot=, so it survives the reboot the =/run= guard sentinel cannot). A system oneshot, =archsetup-boot-upgrade.service=, =ConditionPathExists= on the flag, ordered =Before=getty@tty1.service= so it completes before autologin execs Hyprland — this ordering is mandatory, because a parallel run would let Hyprland start mid-swap and reintroduce the exact crash the guard prevents. The unit is bounded (=TimeoutStartSec=) and best-effort: its failure or timeout must not fail any target the session needs, so boot proceeds past it regardless. Its =ExecStart= runs, as the user: =informant read= (clear the news hook that would otherwise abort the transaction), then =topgrade --only system= (or the equivalent =yay -Syu=), then =maint stamp topgrade= on success, then removes the flag unconditionally (a one-shot arm — a failed attempt disarms rather than retrying every boot). =sudo= works unattended (=%cjennings NOPASSWD: ALL=), so no password prompt wedges it. +The entry names every kernel whose modules the session could change: each kernel whose package or headers is pending, every kernel with a built DKMS module when a DKMS package is pending, and any kernel still unverified from an earlier session. Because it is keyed by pkgbase and resolved to module trees only after stage 1, a kernel version the session removed can never reach the gate, while the new version, and any sibling kernel whose DKMS module was rebuilt, is checked. The gate runs once, whatever stage 1's exit, against the time the entry opened. The exception is a session that opened the entry itself and whose stage 1 changed no kernel-side version, as when stage 1 fails on a download, signature or conflict error before committing anything. =/boot= is then intact and its initramfs predates the entry, so the time checks would hold the row at CRIT on a machine that boots fine. The gate runs its structural form instead, where a missing kernel image or initramfs, an unreadable image or a missing =zfs.ko= still fails it. -The stamp also needs to happen when the upgrade is completed by other safe means — the by-hand sentinel-override path, or a =maint= command that does the same thing. The cleanest single home for the stamp is a small =maint apply-upgrade= (or a flag in the existing lever) that performs the guarded system upgrade and stamps on success, which both the boot unit and an interactive TTY run call. That keeps one code path that "completes a guarded upgrade and records it," rather than three writers that can drift. +A failed gate stops the session with the failing item named and exit 4. The machine keeps running the old kernel, the new one is installed but not booted, and the desktop is there for the fix; a later =--complete= re-runs the gate and closes the entry when it passes. Only a stage 1 that succeeded, with no gate left open, goes on to the GPU step. With Hyprland not running, the session applies the GPU/compositor set directly. With Hyprland live and the boot unit enabled, it pre-downloads that set while the network is up and arms the flag with its exact versions. With the unit not enabled, it says so and leaves the set for a console run with Hyprland stopped. It offers a reboot only from a tty, after a clean finish that armed or passed a time-checked gate. With nothing pending and no gate open, it reports that there is nothing to complete. Exact behavior: Phase 1, =--complete= sequence; Phase 1, Invariants. + +On a ZFS root, the fallback for a kernel that won't boot is the pre-pacman snapshot that ZFSBootMenu can boot. So on a ZFS root a session that opens the gate entry snapshots the root and places a ZFS user hold on that snapshot before stage 1. When no valid hold exists, a time-checked gate falls back to the oldest pre-pacman snapshot taken no earlier than a minute before the entry opened. The session records the held name. =zfs-pre-snapshot= prunes only unheld snapshots, so its routine prune can't destroy the one snapshot holding the old kernel before the new kernel has proven itself. The hold moves only when a later session, or an everyday run whose AUR step opens the entry, holds a newer snapshot, and a failed hold is reported without failing the run. The recorded name is the snapshot recovery boots and rollback releases. Exact behavior: Phase 1, The snapshot hold; Phase 4, Recovery. + +*The gate.* =kernel-modules-check= (archsetup =scripts/kernel-modules-check=) answers one question: would each named kernel boot with its modules? It takes the entry's pkgbases, writes nothing, makes no pacman call, and only =--complete= calls it (a by-hand run, as in Phase 4, only diagnoses). For each kernel it requires the module trees and the =/boot= image to exist, every DKMS module to be installed, and an initramfs newer than the image. On a ZFS root that initramfs must carry =zfs.ko=, read with =sudo -n lsinitcpio= because the images are root-only. The time-checked form also requires the initramfs to postdate the entry's opening and, on a ZFS root, a pre-pacman snapshot of the root dataset taken no more than a minute before it, the window =zfs-pre-snapshot='s 60 s skip leaves. The initramfs is never compared with the module tree's =vmlinuz=, whose mtime is its build date. Kernels outside the entry are never checked, so ratio's =linux-lts-strix=, where zfs is only 'added', never fails it. There is no standalone mode; the supported re-check is =upgrade-guarded --complete=. Exact behavior: Phase 1, The gate. + +*The arm flag.* The guard's =/run= sentinel is tmpfs and can't carry an intent across a reboot, so the arm flag lives at =/var/lib/archsetup/apply-upgrade-on-boot=, root-owned, and carries the exact === list to install. Arming first checks that the list resolves without pulling in a kernel-side package, so no kernel can ride in ungated as a dependency, then pre-downloads it with =pacman -Sw= while the network is up. A download-only run returns before PreTransaction hooks, so neither the guard nor informant's hook fires under a live Hyprland. Only =upgrade-guarded= writes the flag, atomically through =sudo -n=, and the boot unit consumes it. =--complete= clears it after its refresh and re-arms only from its GPU step. While armed, every everyday run whose refresh succeeded either re-arms on the current GPU-kind entries, so the armed versions follow the db that run synced, or disarms when there is nothing valid to arm, including on a machine whose boot unit isn't enabled, so no flag is left that nothing will consume. There is no =--disarm=, and the panel reads only the flag's presence and entries. Exact behavior: Phase 1, The arm flag; Phase 1, Everyday sequence. + +*The boot unit.* =archsetup-boot-upgrade.service= is a system oneshot that runs only when the flag exists, so an unarmed boot skips it at no cost. It is ordered before =getty@tty1.service=, so it finishes before the tty1 login and so before =~/.profile.d/99-hyprland-autostart.sh= can start Hyprland. That ordering is mandatory: a parallel run would let Hyprland start mid-swap and reintroduce the exact crash the guard prevents. It is wanted by =multi-user.target= but required by nothing and not ordered after the network, so its failure or timeout fails no target the session needs, and boot proceeds past it regardless. It runs as the installing user, so =$HOME= points at that user's maint state, and a root =ExecStartPre= moves the flag into the unit's runtime directory before the attempt. Success, failure, a timeout, SIGKILL and power loss therefore all disarm: the intent survives exactly one reboot, and a failed attempt never wedges later boots. Output goes to tty1 rather than the console, because ratio's =/dev/console= is ttyS0. The 20-minute start timeout is sent as SIGINT, which pacman honors at a package boundary, releasing =db.lck=; the unit never deletes it. A failed or timed-out run leaves the unit failed, and maint's failed-units row names it. Exact behavior: Phase 3, The unit. + +*The boot form.* =--apply-armed= runs only from the boot unit. It installs the armed list from the package cache in one transaction, with no refresh and no network, so it applies exactly the versions checked at arming, and the libraries are current before anything maps them, with no second reboot. Entries already installed at or above their armed version are dropped, so it never downgrades, and it refuses if the list would pull in a kernel-side package. Before the transaction it writes an interrupted record naming the remedy, so even SIGKILL or power loss leaves the next step on record, and its first tty1 line says an upgrade is running and not to power off. It also clears informant before the transaction: offline, the clear and informant's hook both fail fast and see an empty feed, so neither blocks. A network that comes up between the clear and the transaction can still let the hook abort it, which fails the attempt with the arm already consumed. If the run fails or is cut off, the tty1 login still appears and the record names what happened. An interrupted run can leave part of the set upgraded, so its remedy is =upgrade-guarded --complete= from a console with Hyprland stopped. There is no news capture, kernel step, gate, yay or sweep at boot. Exact behavior: Phase 1, =--apply-armed= sequence. + +*Interruption and exit codes.* Until its finish writes the record, every mode traps INT, TERM and HUP, waits for the running child, records the interruption with its step and remedy (=--dry-run= records nothing), keeps any open gate entry as written, and exits 130. A signal after that write rewrites nothing: at =--complete='s reboot prompt it counts as no and the exit stays 0. A failed step, a refusal before any transaction, a failed gate and an interruption each have their own exit code, and on any non-zero exit the last stderr line names the reason. Exact behavior: Phase 1, Exit codes and interrupts. + +=upgrade-guarded= is the one code path that completes an upgrade and records it. A by-hand sentinel-override upgrade outside the script doesn't stamp, though the panel's row still clears, because the probe drops entries already installed. =upgrade-guarded --complete= from a console replaces that ritual. * Alternatives Considered ** A. Decouple the stamp from topgrade's exit code (stamp on any real run) - Good, because it is a one-line change to the wrapper and needs no boot machinery. -- Bad, because it throws away honest signal: a topgrade that was blocked from applying a real upgrade would read as "fresh," so the metric stops meaning "up to date." On this machine the blocked case is the common case, so the metric would be fresh precisely when an upgrade is outstanding. +- Bad, because it throws away honest signal: a topgrade that was blocked from applying a real upgrade would read as "fresh," so the metric stops meaning "up to date." On both daily drivers the blocked case is the common case, so the metric would be fresh precisely when an upgrade is outstanding. - Neutral, because the failed steps still surface elsewhere (pending-updates count), so freshness would become redundant rather than wrong. ** B. Run the full topgrade live with the guard overridden, then reboot -- Good, because it needs no new unit — arm the sentinel, run, reboot. -- Bad, because the dangerous window is the whole rest of the run: mesa swaps early, then topgrade spends minutes on other ecosystems while the live compositor is one new GL context (a new window, the wallpaper daemon) away from SIGABRT. topgrade's own reboot-at-end is that window, not a fix for it. -- Neutral, because it would stamp naturally on success — if it survived. +- Good, because it needs no new unit: set the guard's =/run= override sentinel, run, reboot. +- Bad, because the dangerous window is the whole rest of the run: mesa swaps early, then topgrade spends minutes on other ecosystems while the live compositor is one new GL context (a new window, the wallpaper daemon) away from SIGABRT. topgrade's own reboot-at-end is that window, not a fix for it. It would also land the kernel and DKMS sets with no module check. +- Neutral, because it would stamp naturally on success, if it survived. ** C. Manual TTY ritual only (log out, run topgrade at the console, reboot), plus stamp -- Good, because it is the safest path and needs almost no code — just make the completion stamp. -- Bad, because it is all manual, every guarded-upgrade day; the friction is why it won't happen consistently, which is how the metric got stale in the first place. -- Neutral, because it is exactly what the boot unit automates, so it is really "v1 minus the automation." +- Good, because it is the safest path for the GPU/compositor set and needs almost no code, just the completion stamp. +- Bad, because it is all manual, every guarded-upgrade day; the friction is why it won't happen consistently, which is how the metric got stale in the first place. A console topgrade also lands the kernel and DKMS sets with no module check. +- Neutral, because it survives as the fallback: =upgrade-guarded --complete= from a console with Hyprland stopped is this ritual through the same script, gate included, and D automates its GPU/compositor half. -** D. Boot-time armed oneshot, arch-only (this spec) -- Good, because the risky swap happens with nothing live, the run is one bounded transaction with a tiny prompt surface, it stamps on success, and a failure degrades to "boots normally, try again." -- Bad, because it puts a unit on the boot critical path, which must be bounded and non-fatal with care, and it is the most to build. -- Neutral, because it composes with C: the same =maint apply-upgrade= path serves both an interactive TTY run and the boot unit. +** D. Boot-time armed oneshot, pacman-only (the GPU/compositor half of this spec's completion step) +- Good, because the GPU/compositor swap happens with nothing live, as one bounded transaction from the package cache with a tiny prompt surface. Arming pre-downloads the exact versions, so the boot run needs no network. It stamps when it leaves nothing deferred, and a failure or interruption degrades to a normal boot with the flag already consumed, finished by running =upgrade-guarded --complete= again. +- Bad, because it puts a unit on the boot critical path, which must be bounded and non-fatal with care, and it is the most to build. The flag also has to follow the sync db between arming and the reboot. +- Neutral, because it composes with C: the same =upgrade-guarded= script serves both, =--complete= from a console and =--apply-armed= from the boot unit. Kernels never ride it: they land in =--complete='s gated first stage, before anything arms, and both arming and the boot run refuse a target list that names a kernel-set or DKMS-set member. +- Exact behavior: Phase 1 (the split-upgrade script) and Phase 3 (the boot-time unit). ** E. Split the live run: apply the non-blocked remainder now, defer the rest (this spec's everyday path) -- Good, because it is what a careful operator does by hand anyway — and did, on ratio, the day this was decided. The live run succeeds on the common day, topgrade exits 0, the AUR and every other ecosystem stay current, and the guard's abort becomes the rare path rather than the default. -- Bad, because Arch calls any =--ignore= run a partial upgrade. In practice pacman still enforces declared dependencies, so anything needing the newer mesa fails resolution instead of installing broken; the residual exposure is a package with an *unversioned* dependency built against a new ABI, which for mesa/wayland/libdrm is rare. Named, accepted. -- Bad, because a deferred set nobody surfaces is a set that silently never lands — the same trap as the freshness stamp, one layer down. So the script must record the deferred set durably and the panel must show it; this is why D stays in the design as the completion step rather than being replaced. -- Neutral, because it does not change the guard at all; it changes who decides the transaction's contents. +- Good, because it is what a careful operator does by hand anyway, and did, on ratio, the day this was decided. The live run succeeds on the common day, the sweep exits 0, the AUR and every other ecosystem stay current, and the guard trips only in the rare case of an AUR upgrade that needs a held GPU/compositor package while Hyprland is live. +- Bad, because Arch calls any =--ignore= run a partial upgrade. In practice pacman still enforces declared dependencies, so a package needing the newer mesa never installs against the old one. But under =--noconfirm= one unresolvable dependent fails the whole transaction (pacman's skip prompt defaults to no), which is why the script has pacman compute the dependency closure first, and refuses with the blocker named rather than force anything. The residual exposure is a package with an *unversioned* dependency built against a new ABI, which for mesa/wayland/libdrm is rare. Named, accepted. +- Bad, because a deferred set nobody surfaces is a set that silently never lands, the same trap as the freshness stamp one layer down. So the script records the deferred set durably and the panel shows it, and this is why D stays in the design as the completion step rather than being replaced. +- Neutral, because it does not change the guard at all; it changes who decides the transaction's contents: the script holds the GPU/compositor set while Hyprland is live, and the kernel and DKMS sets on every run. +- Exact behavior: Phase 1 (the split-upgrade script). * Decisions [7/7] @@ -109,96 +166,5458 @@ CLOSED: [2026-08-25 Tue 18:45] ** DONE Everyday mechanism is the split live run (Alternative E); the boot oneshot (D) completes the deferred set CLOSED: [2026-08-25 Tue 18:30] - Owner / by-when: Craig / 2026-08-25 -- Context: the first draft made D the primary gesture, which means every guarded-library day is a reboot day. Craig's read while watching the ratio run: when the guard would trip, the rational move is to upgrade everything *except* the guarded and kernel items, then run the rest of topgrade without the yay piece — and that logic should be a script we can keep editing, not something baked into the panel. -- Decision: I will build E as the path UPDATE and TOPGRADE always take on a live session, and keep D as the way the deferred set lands (arm + reboot). C remains the manual fallback through the same script from a TTY (no compositor → nothing blocked → a full run). B stays rejected on the live-swap risk. -- Consequences: easier — the common day is one live run that succeeds; reboots are reserved for the days the deferred set is non-empty, and even then the machine keeps working until the reboot is convenient. Harder — two lists to maintain (the guard's, read from the hook; the kernel list, owned by the script), a state file the panel must render, and the partial-upgrade caveat above to keep an eye on. +- Context: the first draft made D the primary gesture, which means every guarded-library day is a reboot day. My read while watching the ratio run: when the guard would trip, the rational move is to upgrade everything *except* the guarded and kernel items, then run the rest of topgrade without the yay piece — and that logic should be a script we can keep editing, not something baked into the panel. +- Decision: I will build E as the path UPDATE and TOPGRADE always take, and keep D as the way the deferred set completes. =upgrade-guarded --complete= lands the kernel and DKMS sets behind the gate (the kernel decision below), then, while Hyprland is live, arms the boot oneshot to land the GPU/compositor entries on the next boot. C remains the manual fallback through the same script from a console with Hyprland stopped: an everyday run there holds only the kernel and DKMS sets, and =--complete= there applies the GPU/compositor set directly instead of arming. B stays rejected on the live-swap risk. +- Consequences: easier — the common day is one live run that succeeds; reboots are reserved for the days I land the deferred set, and even then the machine keeps working until the reboot is convenient. Harder — two sources to keep straight (the guard's list, read from the hook; the kernel and DKMS sets, which the script derives from the installed system), a record the panel must render, and the partial-upgrade caveat in Alternative E, which the script meets by computing the dependency closure of everything it holds and refusing before any transaction when that closure can't be resolved. +- Exact behavior: Phase 1, Everyday sequence and =--complete= sequence. ** DONE The script lives in archsetup beside the guard, and maint calls it CLOSED: [2026-08-25 Tue 18:45] - Owner / by-when: Craig / 2026-08-25 -- Context: the panel (dotfiles =maint=) and the guard (archsetup =scripts/hypr-live-update-guard=, installed to =/usr/local/bin=) live in different repos, and maint already carries its own copy of the trigger list as =[updates] guard_patterns= in the thresholds TOML. -- Decision: ship the split script in archsetup next to the guard, installed by the same installer step, reading the blocked list from the installed hook so the guard and the script can never disagree. maint's UPDATE/TOPGRADE levers change their =argv= to the script; the TOML patterns stay as the panel's *display-side* mirror (the badge that says a run will defer) and gain a test asserting they match the hook. -- Consequences: easier — one owner for "which libraries are dangerous," and a rebuilt machine gets the script with the guard. Harder — a cross-repo change (archsetup ships it, dotfiles wires it), so the rollout is two commits, archsetup first. +- Context: the panel (dotfiles =maint=) and the guard (archsetup =scripts/hypr-live-update-guard=, installed to =/usr/local/bin=) live in different repos, and maint already carries its own, divergent pattern list as =[updates] guard_patterns= in the thresholds TOML (seeded from archsetup's =configs/maintenance-thresholds.toml=). +- Decision: ship the split script, =upgrade-guarded=, in archsetup next to the guard, with its gate, =kernel-modules-check=, beside it; the installer step that installs =hypr-live-update-guard= installs both to =/usr/local/bin=. The script reads the GPU patterns from the installed hook's =Target= lines, so the guard and the script can never disagree, and refuses to run without them. maint's UPDATE and TOPGRADE levers change their =argv= to the script: UPDATE runs =upgrade-guarded --no-topgrade=, TOPGRADE runs =upgrade-guarded=. The TOML patterns become the panel's display-side mirror, read only for the arm line that warns a press will defer at least the matched packages. They are rewritten to equal the hook's 15 =Target= patterns and pinned by an archsetup test in =tests/installer-steps/= that compares the seed TOML with the hook heredoc. +- Consequences: easier — one owner for "which libraries are dangerous," and a rebuilt machine gets the script with the guard. Harder — a cross-repo change (archsetup ships it, dotfiles wires it), so the rollout is three ordered commit groups: archsetup Phase 1 (both scripts and the TOML rewrite), dotfiles Phase 2 (the levers and the panel), then archsetup Phase 3 (the boot unit). Each machine installs the archsetup scripts and re-copies the rewritten TOML before any Phase 2 commit reaches it. +- Exact behavior: the Implementation phases intro (commit groups and ordering invariants); Phase 2, Levers and Guard UX; Phase 4, Rollout. ** DONE What a split run stamps CLOSED: [2026-08-25 Tue 18:45] - Owner / by-when: Craig / 2026-08-25 - Context: under the state framing, a run that deferred six packages left the system *not* current, yet the sweep ran and every other ecosystem is fresh. -- Decision: the script stamps =topgrade_run= only when the deferred set is empty. When it is non-empty it writes the deferred set to its own cache key, and the panel renders that as its own state ("6 deferred — apply on reboot") rather than as stale freshness. The boot oneshot stamps when it completes the deferred set. Freshness keeps meaning "current"; the deferred badge carries the other half. -- Consequences: easier — no signal is thrown away, and the reboot nag has a precise count behind it. Harder — one more cache key and one more probe in maint. +- Decision: the script stamps =topgrade_run= only when the system is current: nothing left deferred, no open kernel gate, and every step it ran succeeded. On the driven path it is the only writer. doctor.py's TOPGRADE-lever stamp is deleted, and the script calls =/usr/bin/topgrade= by absolute path, so the stamping PATH wrapper (which still stamps a bare =topgrade= typed at a shell) is never in its chain. UPDATE never stamps, and neither does a run that arms the boot oneshot or finds nothing to complete. A stamp that fails fails the run rather than passing silently. The deferred half goes to the script's own record, =${MAINT_STATE_DIR:-$HOME/.local/state/maint}/upgrade_deferred.json= (maint cache key =upgrade_deferred=), which it rewrites after every run, refusals, failures and interruptions included (Phase 1, The record). maint renders the record as its own state ("6 deferred", with APPLY to start the completion) rather than as stale freshness. Freshness keeps meaning "current"; the deferred row carries the other half. +- Consequences: easier — no signal is thrown away, and the deferred row has a precise count behind it. Harder — one more cache key and one more probe in maint, and a record format two repos must agree on. archsetup holds the canonical fixture, bound to a fake run's actual output, and dotfiles holds a byte-identical copy whose test names its source, so a field change updates both in the same rollout. And because UPDATE never stamps, freshness clears only through a TOPGRADE with nothing deferred or a completion that ends empty. +- Exact behavior: Phase 1, Stamp predicate and The record. ** DONE Kernel set is held on every everyday run and lands only in the dedicated session, gated on the DKMS result CLOSED: [2026-08-25 Tue 18:45] - Owner / by-when: Craig / 2026-08-25 -- Context: I first wrote this as "install the kernel live at apply-on-reboot," on the reasoning that a kernel swap crashes nothing and the modules-vanish window ends with the reboot. Craig asked what happens on velox when the DKMS rebuild fails, and the answer changed the decision. Velox is an encrypted ZFS root with =/boot= inside the root dataset, one kernel (=linux-lts=), and =zfs-dkms=. On a kernel upgrade the DKMS build runs PostTransaction, after the kernel is swapped and the old modules are deleted, so nothing can abort; a failed build leaves a new kernel beside an initramfs built for the old one, whose =zfs.ko= won't load, and the next boot can't import the pool. It is survivable — ZFSBootMenu can boot the pre-pacman snapshot, which holds the old kernel, initramfs, and modules — but it is a recovery session, not an update. Ratio (btrfs root, two kernels, zfs only for a data pool) is exposed only at the pool. The realistic triggers are a kernel major outrunning OpenZFS's supported range, a kernel upgraded without its headers, a toolchain regression, or a full disk. -- Decision: the script holds the kernel set — every installed kernel with its =-headers=, moved as a set, never one without the other — on every everyday run, on both machines, so there is one rule rather than a per-host exception. The kernel set lands only in the dedicated session, live, while a working desktop exists for diagnosing, and the script gates what follows on the result: =dkms status= reports every DKMS module installed for the new kernel version, the initramfs is newer than the kernel image, and on a ZFS root a pre-pacman snapshot exists. A failed gate stops with the failure named and never reboots. The GPU/compositor set follows only after the gate passes — armed for the boot oneshot, or applied from a TTY. Kernels stay off the guard's list (the hook would block a TTY kernel upgrade for no reason). "Install the kernel live at apply-on-reboot" is withdrawn. -- Consequences: easier — an everyday UPDATE can never put velox into the unbootable state, and the day the kernel moves is one Craig chose, sitting at the machine, expecting to handle issues. Harder — the kernel deferral is now standing, so the dedicated session has to happen on a cadence (security fixes ride the kernel), and the panel's deferred count carries a kernel most days; the gate is one more script to test, with fakes for =dkms status= and the image timestamps. +- Context: I first wrote this as "install the kernel live at apply-on-reboot," on the reasoning that a kernel swap crashes nothing and the modules-vanish window ends with the reboot. Then I asked what happens on velox when the DKMS rebuild fails, and the answer changed the decision. Velox is an encrypted ZFS root with =/boot= inside the root dataset, one kernel (=linux-lts=), and =zfs-dkms=. On a kernel upgrade the old modules are removed PreTransaction (=70-dkms-upgrade=) and the rebuild runs PostTransaction (=70-dkms-install=), after the kernel is swapped; the alpm dkms script returns 0 even when a build fails, so nothing aborts. A failed build leaves a new kernel beside an initramfs that mkinitcpio wrote without =zfs.ko= (it only warns), and the next boot can't import the pool. The same chain runs with no kernel move at all when a DKMS package itself upgrades (today =zfs-dkms=, whose upgrade removes and rebuilds the module for every installed kernel). It is survivable — ZFSBootMenu can boot the pre-pacman snapshot, which holds the old kernel, initramfs, and modules — but it is a recovery session, not an update. Ratio (btrfs root, zfs only for a data pool) is exposed only at the pool; it has three kernels, including =linux-lts-strix=, a foreign local build with no headers and no zfs module, which is the GRUB default. The realistic triggers are a kernel major outrunning OpenZFS's supported range, a kernel upgraded without its headers, a toolchain regression, or a full disk. +- Decision: the script holds the kernel set and the DKMS set on every everyday run, on both machines, whether or not Hyprland is running, so there is one rule rather than a per-host exception. Both sets are derived from what is installed, never from a fixed list. The kernel set is every installed kernel with its headers, held whole so a kernel and its headers always move together. The DKMS set is every package that ships a DKMS module (today =zfs-dkms=); its exact =zfs-utils= pin isn't listed by name, because the dependency closure adds =zfs-utils= exactly when an upstream bump breaks the pin. Only pending members ever move, so a foreign kernel (ratio's =linux-lts-strix=) is never touched. Both sets land only in the dedicated session, =upgrade-guarded --complete=, in its first stage, live on the running system (usually with the desktop up for diagnosing), and =kernel-modules-check= gates everything after that stage. Before the stage touches anything kernel-side, the session opens a gate entry in the record covering every kernel whose modules it could change, identified by pkgbase rather than by kernel version, so the gate checks whatever version stage 1 leaves installed. The gate checks that each of those kernels can boot with its modules: built, in a fresh initramfs, and on a ZFS root with a fresh pre-pacman snapshot behind it. The entry closes only when a gate passes or every kernel it names has been removed, so an interrupted or failed stage 1 leaves it open. While it is open no arm flag exists and nothing offers a reboot, and a failed gate exits 4 with the failure named. The GPU closure follows only after the gate passes: armed for the boot oneshot while Hyprland is running, or applied directly by a second stage when it isn't. If that closure would include a kernel-set or DKMS-set member, =--complete= refuses before any transaction. Kernels stay off the guard's list (the hook would block a TTY kernel upgrade for no reason). "Install the kernel live at apply-on-reboot" is withdrawn. +- Consequences: easier — an everyday UPDATE can never put velox into the unbootable state, because no kernel and no DKMS module moves outside the gated session, and the day they move is one I chose, sitting at the machine, expecting to handle issues. Harder — the kernel deferral is now standing, so the dedicated session has to happen on a cadence (security fixes ride the kernel), and the panel's deferred count carries a kernel most days. From the start of a kernel-side stage 1 until a gate passes, the panel holds the deferred row at CRIT with REBOOT hidden, so a failed gate leaves the new kernel installed but not booted, with the desktop available for the fix, until a later =--complete= passes. On a ZFS root the session places a ZFS hold on the oldest pre-pacman snapshot taken since a minute before the entry opened, so the pre-snapshot prune can't destroy the fallback before it's needed; the hold moves only when a later session holds a newer one. And the gate is one more script to test. +- Exact behavior: Phase 1, Sets, =--complete= sequence, Invariants, The gate and The snapshot hold. ** DONE Boot run applies exactly the deferred GPU/compositor set, nothing else CLOSED: [2026-08-25 Tue 18:45] - Owner / by-when: Craig / 2026-08-25 -- Context: the only packages that need a stopped compositor are the guard's trigger set; the rest run fine live and are what usually fail. The first draft phrased this as =topgrade --only system=; with the split script that wording is stale, and with the kernel decision above the kernel is not part of what boot applies either. -- Decision: the boot oneshot runs the script's =--complete= form scoped to the deferred GPU/compositor set: one pacman transaction, no ecosystem sweep, no kernel. The full topgrade sweep stays a normal live run through the everyday path. -- Consequences: easier — the boot path is fast, has a tiny interactive-prompt surface, and rarely fails. Harder — freshness after a boot run reflects the guarded set specifically, which is what the stamp decision above already accounts for. +- Context: the only packages that need a stopped compositor are the guard's trigger set and what the dependency closure ties to it; the rest run fine live and are what usually fail. With the kernel decision above, the kernel and DKMS sets are not part of what boot applies either. A boot run can't count on the network, and a refresh at boot would install versions nobody checked when arming. +- Decision: the boot oneshot runs the script's boot form, =upgrade-guarded --apply-armed=: exactly the armed GPU/compositor set, closure members included, in one pacman transaction, with no ecosystem sweep and no kernel. It installs the exact versions recorded at arming from the package cache, where arming pre-downloaded them while the network was up, so it needs no refresh and no network, and the unit has no network-online ordering. It skips news capture, the kernel step, the gate, yay and topgrade, never downgrades, and refuses rather than install anything that would bring in a kernel-set or DKMS-set member. The full topgrade sweep stays a normal live run through the everyday path. +- Consequences: easier — the boot path is fast (one transaction from the cache, no network wait), has a tiny interactive-prompt surface (=--noconfirm=, no stdin, unread news cleared first with =informant read --all=), applies exactly the versions checked at arming, and rarely fails. Harder — the flag has to stay in step with the sync db, so every everyday run while armed whose refresh succeeds re-checks, re-downloads and rewrites it, or disarms; freshness after a boot run means the pacman side is current, not that the ecosystem sweep ran recently, which the stamp decision above accepts; and a network that comes up between the boot run's news clear and its transaction can let informant's hook abort the transaction, which the record names as a failed boot transaction with the arm already consumed. +- Exact behavior: Phase 1, The arm flag and =--apply-armed= sequence; Phase 3, The unit. ** DONE Arm flag lives on a persistent path and is one-shot CLOSED: [2026-08-25 Tue 18:45] - Owner / by-when: Craig / 2026-08-25 -- Context: the guard's =/run= sentinel is tmpfs and cleared on reboot, so it cannot carry an intent across the reboot. A boot that retries forever on failure is its own outage. -- Decision: We will use a persistent flag (=/var/lib/archsetup/=) that the boot unit removes unconditionally at the end of its attempt — success or failure disarms. -- Consequences: easier — the intent survives exactly one reboot and a failed attempt never wedges subsequent boots. Harder — a failed attempt needs re-arming, which is correct (a human decides to try again) but is a manual step. +- Context: the guard's =/run= sentinel is tmpfs and cleared on reboot, so it cannot carry an intent across the reboot. A boot that retries forever on failure is its own outage, and a flag removed at the end of the attempt is never removed when the attempt is killed or loses power. +- Decision: We will use a persistent, root-owned flag, =/var/lib/archsetup/apply-upgrade-on-boot= (root:root 0644), that carries the armed list: one === line per GPU-kind upgrade target. Only =upgrade-guarded= writes it, atomically through =sudo -n=. =--complete= clears any flag right after its refresh and re-arms only from its GPU step, after the gate has passed, so no flag exists while a kernel gate is open. The boot unit moves the flag off its persistent path into its runtime directory at the start of its attempt, in =ExecStartPre=, before the boot form runs, so success, failure, a timeout, SIGKILL or power loss all disarm. The boot form reads the moved copy and never touches the flag. +- Consequences: easier — the intent survives exactly one reboot, a failed or killed attempt never wedges subsequent boots, and the boot run applies exactly the versions checked at arming. Harder — a failed attempt needs re-arming by running =upgrade-guarded --complete= again, which is correct (a human decides to try again) but is a manual step. An attempt cut short by the unit's 20-minute timeout (sent as SIGINT, so pacman stops at a package boundary and releases =db.lck=) can leave part of the GPU/compositor set upgraded, so the boot form pre-writes an interrupted record naming the remedy (=upgrade-guarded --complete= from a console with Hyprland stopped) before its transaction. The panel only reads the flag's presence and entries. +- Exact behavior: Phase 1, The arm flag and Invariants; Phase 3, The unit. -* Implementation phases +* Review findings [129/138] +Recorded by the 2026-10-05 spec-review (rubric: Not ready). Each finding names the spec passage, what it says or what the code does today, the risk, and the smallest change that closes it. =:blocking:= findings hold the rubric at Not ready until dispositioned. The evidence each finding was verified against is folded in its =EVIDENCE= drawer. Problem texts in the last nine (round 3 residual) items refer to earlier repair items as "gap N" or by short labels such as T11 or G6; those labels were working names and aren't defined elsewhere, so each item's title states the gap. -** Phase 1 — The split-upgrade script (archsetup) -=scripts/guarded-upgrade= (name open), installed to =/usr/local/bin= by the step that installs the guard. Behaviour as in Design: pending set → blocked set (hook =Target= lines, version-aware) ∪ held-kernel set when a compositor is live → =informant read= if present → =pacman -Syu --noconfirm --ignore=…= → =yay -Sua --noconfirm= → =topgrade --disable system,git_repos -y= → deferred set written to a state file → stamp only when nothing was deferred → exit 0 on a successful live part. The kernel set is derived from what is installed (every =linux*= kernel package and its =-headers=), never a hardcoded pair, and is always held or applied whole. Flags: =--dry-run= (print the plan and the deferred set, change nothing), =--no-topgrade=, =--no-aur=, =--complete= (the dedicated-session form: apply the kernel set live, run the gate, then arm the GPU/compositor set or, with no compositor live, apply it directly; stamp when the deferred set is empty). The gate is its own small script, =kernel-modules-check=: for each kernel under =/usr/lib/modules=, =dkms status= reports every registered module =installed= for it, and its initramfs is newer than its =vmlinuz=; on a ZFS root, a =pre-pacman_= snapshot of the root dataset exists. It exits non-zero with the failing item named, and =--complete= refuses to arm or reboot on that exit. Usable from a TTY at once. Tests (pytest beside the guard's): blocked-set computation against a fixture hook and version map; the kernel set is derived from the installed kernels and held whole on every everyday run; the =--ignore= list is exactly blocked ∪ kernel set; the state file round-trips; stamps only on an empty deferred set; =--dry-run= is IO-free; the gate passes and fails on fake =dkms status= output, image timestamps, and snapshot listings, and =--complete= never reaches the arm step on a failed gate. +** DONE Existing stamp writers mark a deferred run fresh :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Problem L32; Design L65; Decision 4 'What a split run stamps' L124-129; Phase 1 L155; Phase 2 L158; AC L168 -** Phase 2 — Wire maint to it (dotfiles) -UPDATE and TOPGRADE levers change their =argv= to the script; the press-again-to-force sentinel wrap goes away (the driven path never trips the guard). A new probe reads the deferred-set state file and the panel renders "N deferred — apply on reboot" as its own row. A test asserts the TOML =guard_patterns= equal the installed hook's =Target= list. Tests under the maint fake harness. +Decision 4 says only the script stamps topgrade_run, and only when the deferred set is empty. However, the script calls bare =topgrade --disable ... -y= and exits 0 'whatever was deferred'. Phase 2 only points the TOPGRADE/UPDATE levers' argv at the script. Neither phase retires the two existing writers that the Problem section names. -** Phase 3 — "Apply on reboot" and the boot-time unit (archsetup + maint) -The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot. =archsetup-boot-upgrade.service=, installed by the installer: =ConditionPathExists= the flag, =Before=getty@tty1.service=, =TimeoutStartSec= bounded, non-fatal to every session target, =ExecStart= runs the script's =--complete= form as the user and removes the flag unconditionally. Installer step + unit file + the =/var/lib/archsetup/= flag directory. Tests in =tests/installer-steps/= for the install step; the arm action's tests in maint; a documented manual boot test (defer, arm, reboot, observe) in =todo.org= under Manual testing and validation. +Risk: On every run that defers guarded packages or the kernel, topgrade_run gets stamped twice, once by the wrapper and once by doctor. topgrade_age then reads fresh while the upgrade is still outstanding. That is the Alternative A outcome (L77) that Decisions 1 and 4 reject. The panel shows a deferred badge next to a fresh age, and AC L168 fails both through the panel and from a shell. -** Phase 4 — Docs, rollout, and both daily drivers -Document the flow (arm → reboot → console upgrade → session). Roll the unit to velox and ratio (installer already covers a rebuild; existing machines need the one-time install). Confirm the ratio path matches. +Recommended change: Phase 1 (L155): the script invokes the packaged binary by absolute path (/usr/bin/topgrade, owned by the topgrade package), never =topgrade= through PATH, so the stamping wrapper is never in the chain. The script remains the only thing that runs =maint stamp topgrade=, and only when the deferred set is empty. Add a test: put a stamping fake =topgrade= ahead on PATH, defer a non-empty set, and assert that topgrade_run is untouched. -* Acceptance criteria -- [ ] With a guarded library pending and Hyprland live, UPDATE applies everything else, exits 0, and the panel shows the exact deferred set; the guard hook does not fire. -- [ ] The same run with no compositor live (a TTY) applies everything but the kernel set and stamps only if nothing was deferred. -- [ ] A =--complete= run whose DKMS build fails stops before arming or rebooting, names the failure, and leaves the machine running on the old kernel; on velox the pre-pacman snapshot it required is bootable from ZFSBootMenu. -- [ ] With a guarded library pending, arming and rebooting applies it in the console before Hyprland starts, and =maint status= then reads a fresh =topgrade_age=. -- [ ] A boot-upgrade failure (a failed step, a timeout, an aborted transaction) never blocks the session: the machine boots into Hyprland, the flag is cleared, and the panel still shows the pending work. -- [ ] Unread Arch news does not wedge the boot run (=informant read= precedes the transaction). -- [ ] A guarded upgrade completed from a TTY via the Phase-1 path stamps freshness identically to the boot unit. -- [ ] The =hypr-live-update-guard= hook is unchanged and still blocks a live guarded swap. +Phase 2 (L158): delete doctor.py:169-170's =if rid == "topgrade"= stamp, because the script owns the stamp. Update the wrapper's header comment and test_topgrade_wrapper.py so they no longer describe the lever as stamping through the wrapper. Add a test: a zero-exit TOPGRADE lever whose script deferred something leaves topgrade_run unchanged. -* Readiness dimensions -Answer each, or write "N/A because…". -- Data model & ownership: the arm flag (=/var/lib/archsetup/=, installer-owned) and the =topgrade_run= cache key (maint-owned). No user-authored data. -- Errors, empty states & failure: the boot unit is best-effort and self-disarming; every failure path lands in "boot normally, metric stays stale, re-arm to retry." Named, non-silent. -- Security & privacy: relies on the existing =%cjennings NOPASSWD: ALL=; the unit runs the upgrade as the user via sudo, adds no new privilege. Note the NOPASSWD breadth as a pre-existing fact, not introduced here. -- Observability: the boot run's output is on the console; its systemd unit status and journal record success/failure; the panel reflects the cleared or still-pending state after boot. -- Performance & scale: one pacman/yay transaction at boot; bounded by =TimeoutStartSec=. Negligible boot-time cost when the flag is absent (=ConditionPathExists= skips the unit). -- Reuse & lost opportunities: reuses the guard's trigger list by reading the installed hook (single source of truth for "which libs are dangerous"), =informant=, =maint stamp=, =checkupdates=, and topgrade's own step switches. The one duplicate that exists today — maint's TOML =guard_patterns= — is kept as a display mirror and pinned to the hook by a test rather than removed. -- Architecture fit & weak points: integration points are the pacman hook set, getty autologin ordering, and the maint cache. Weak point: the =Before=getty@tty1= ordering is load-bearing for safety; a parallel run reintroduces the live-swap crash. Mitigated by making the ordering explicit and tested-by-inspection. -- Config surface: the arm flag path and the timeout. Defaults safe (absent flag = no-op). -- Documentation plan: a short "reboot to apply guarded upgrades" note in the maint docs; the installer step self-documents in-comment. -- Dev tooling: installer-step pytest for Phase 2; maint unit tests for Phases 1 and 3; a manual boot test in =todo.org=. -- Rollout, compatibility & rollback: additive; removing the unit and flag reverts fully. Existing machines need a one-time install; a rebuild gets it from the installer. Rollback leaves the guard and manual TTY path intact. -- External APIs & deps: topgrade =--only system=, =informant read=, =yay=, =maint stamp= — all verified present on velox this session. No external service. +Acceptance criteria: extend L167 with "...and topgrade_run is not written when anything was deferred, whether the run starts from the panel or a shell." -* Risks, Rabbit Holes, and Drawbacks -- Boot critical path: the unit sits ahead of autologin, so a hang would delay boot. Mitigated by =TimeoutStartSec= and non-fatal wiring; worst case is a bounded delay, then a normal session. -- Interactive prompts under no stdin: =yay=/pacman can still prompt (provider choice, replace, AUR review) even with =assume_yes=. The =--only system= scope and =--noconfirm=-style flags shrink this to near zero, but a prompt with no stdin fails the run (benign) — needs a genuinely non-interactive invocation, verified in Phase 2. -- Partial ecosystem state: N/A for the GPU hazard — each pacman run is one atomic transaction, so there is no half-swapped library. The =--ignore= run is a partial upgrade in Arch's sense; pacman's dependency resolution is the safety net, and the residual unversioned-ABI exposure is accepted in Alternative E. -- Orphans that block resolution: a package dropped from the repo but still pinning an old version (ratio's =qemu-block-gluster= on 2026-08-25) fails the whole transaction. The script should detect the "could not satisfy dependencies" case, name the foreign package, and stop with the remedy — never =-Rdd= on its own. -- The kernel on a DKMS ZFS root: a failed =zfs-dkms= build after the kernel swap cannot be aborted (the DKMS hooks are PostTransaction) and leaves velox unbootable on the new kernel. Mitigated by holding the kernel set on every everyday run, landing it only in the dedicated session behind the gate, and by the standing fallback: =/boot= lives in the root dataset, the =05-zfs-snapshot= hook snapshots it before every transaction, and ZFSBootMenu can boot that snapshot. The pacman cache also keeps the previous kernel and =zfs-dkms= for a downgrade. Ratio's exposure is its data pool only (btrfs root, two kernels). -- Standing kernel deferral: because the everyday run never moves the kernel, the dedicated session has to happen on a cadence or kernel security fixes sit unapplied. The panel's deferred row is the reminder; a stale-kernel age in maint is a possible follow-up. +Blocking. Verification: confirmed. +Disposition: accepted, folded into Design, Implementation phases, Acceptance criteria, Testing. +:EVIDENCE: +Spec (docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org): +- L32 names the two writers: the PATH wrapper on rc 0, and the TOPGRADE lever in doctor.py. +- L65 and L155: the script runs bare =topgrade --disable system,git_repos,containers -y= and "exits 0 when the live part succeeded, whatever was deferred". +- L128 (Decision 4): "the script stamps topgrade_run only when the deferred set is empty". +- L136: the deferred count "carries a kernel most days". +- L158: Phase 2 only repoints argv and removes the sentinel wrap. -* Testing / Verification / Rollout -Phase-1 and Phase-3 logic under the maint fake harness; Phase-2 install under =tests/installer-steps/=. The one thing no unit test can cover — that an armed reboot actually applies the upgrade pre-session and stamps — is a scripted manual test in =todo.org= (arm with a guarded lib pending, reboot, confirm the console run, the fresh metric, and a normal session). Roll to velox first, then ratio. +Code and system: +- ~/.dotfiles/hyprland/.local/bin/topgrade:28-43: stamp=1 unless -n/--dry-run/--version/-V/--help/-h; ="$real" "$@"=, then =maint stamp topgrade= on rc 0. +- The wrapper header comment (L5-6) treats the two writers as intended: "The maint TOPGRADE lever resolves this wrapper too; doctor's own stamp is redundant but harmless." +- =which -a topgrade= lists ~/.local/bin/topgrade before /usr/bin/topgrade. ~/.local/bin/topgrade is a symlink to ../../.dotfiles/hyprland/.local/bin/topgrade. +- /proc//environ PATH for waybar (2361) and the panel's sh (2086) starts ~/.local/bin:...:/usr/local/bin:/usr/bin. +- ~/.dotfiles/maint/src/maint/doctor.py:166-170: any returncode 0 leads to =if rid == "topgrade": cache.put("topgrade_run", {"at": time.time()})=. +- ~/.dotfiles/maint/src/maint/remedies.py:303-308: the "topgrade" lever, whose argv Phase 2 repoints at the script. +- ~/.dotfiles/tests/maint/test_topgrade_wrapper.py:51-56 asserts the stamp on rc 0. +- =pacman -Qo /usr/bin/topgrade= reports topgrade 17.12.2-1, which owns the real binary. +- No bypass exists today: grep for upgrade-guarded, /usr/bin/topgrade and MAINT_NO_STAMP in maint/, hyprland/ and tests/ finds nothing. +:END: + +** DONE topgrade --disable comma list is rejected by the installed topgrade :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L65; Phase 1 L155; Readiness External deps L189; history L11, L204 + +The spec fixes the argv as =topgrade --disable system,git_repos,containers -y= and relies on it exiting 0. Line 189 still lists =topgrade --only system= as the verified dependency. + +Risk: Copied as written, the sweep fails argument parsing on every run before any step runs. No ecosystem is ever swept and topgrade never exits 0. A Phase 1 test with a faked binary that copies the spec's argv would pass and lock the bug in. + +Recommended change: 1. In Design L65 and Phase 1 L155, replace "topgrade --disable system,git_repos,containers -y" with "topgrade --disable system git_repos containers -y". Add a note that each step name is its own argv element, because the installed topgrade's clap parser rejects a comma-joined list with exit 2. +2. Add to the Phase 1 tests: assert that the topgrade argv has each disabled step as a separate element. Add one check against the real binary: "/usr/bin/topgrade --version" exits 0, so the faked harness cannot hide a parse error. +3. In L65, say what a non-zero topgrade exit does. Either it counts against the script's exit and blocks the stamp, or the spec states that it doesn't, so a skipped sweep is never stamped silently. +4. Replace the stale "topgrade --only system" in L189 and L193 (and the boot ExecStart in L69) with the script's --complete form, matching Decision L142. + +Blocking. Verification: confirmed. +Disposition: modified, folded into Design, Implementation phases, Readiness dimensions, Risks, Testing. Accepted, with two additions. First, the binary is called by absolute path, per F01 and my lever call. Second, the argv gains --no-ask-retry. The script ships from archsetup and runs under the panel's no-tty runner, so whether a failed step waits on stdin shouldn't depend on the user's topgrade.toml (no_retry = true today). Item 4's replacement for the stale --only system text names the boot form --apply-armed rather than --complete, per F08. +:EVIDENCE: +Commands run on ratio (read-only): +- pacman -Q topgrade -> "topgrade 17.12.2-1" +- /usr/bin/topgrade --disable system,git_repos,containers --version -> "error: invalid value 'system,git_repos,containers' for '--disable ...' ... tip: a similar value exists: 'system'", exit=2 +- /usr/bin/topgrade --disable system git_repos containers --version -> "topgrade 17.12.2", exit=0 +- /usr/bin/topgrade --disable system --disable git_repos --disable containers --version -> exit=0 +- /usr/bin/topgrade --disable system git_repos containers -y --version -> exit=0 +- topgrade --help line 40: "--disable ..." (the possible values include system, git_repos and containers) + +Spec (docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org): +- L65 and L155 give the comma-joined argv. +- L65: "exits 0 when the live part succeeded, whatever was deferred". +- L128/L155: the stamp is gated only on an empty deferred set. +- L141-142: the decision says the "--only system" wording is stale. +- L69, L189 and L193 still say "topgrade --only system". + +Code: +- ~/.dotfiles/maint/src/maint/remedies.py:308: ["topgrade", "--disable", "git_repos", "-y"] +- todo.org (~L3129): "TOPGRADE exited 2 on every launch ... passed --disable git, rejected by clap before any step ran." +:END: + +** DONE --noconfirm installs ignored packages, so soname cascades abort the split run :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L65 ('on the driven path it never fires'); Alternative E L97 ('fails resolution instead of installing broken'); Phase 1 L155; AC L167 + +The spec assumes --ignore cleanly removes the blocked set, and that anything needing a newer blocked package fails resolution. On that basis it says the guard never fires on the driven path. It specifies no behaviour for two cases: a pending package that needs a held package's new version, and a held package that pins an old version of a pending one. + +Risk: Soname bumps in the hypr* family are routine on Hyprland release days. On those days the everyday UPDATE either re-pulls hyprutils (and the guard aborts) or fails with 'breaks dependency ... required by hyprland'. Nothing is applied, AC L167 can't be met, and the guard's abort becomes the common path again. If a prebuilt-module package like zfs-linux-lts is ever installed, the kernel hold is silently defeated. + +Recommended change: Phase 1 (L155): define the deferred set as a dependency closure rather than "hook targets ∪ kernel set". Start from the version-changing hook targets plus the kernel set, then repeat until nothing new is added: +- Add any pending package whose new version needs a held package's new version (forward). +- Add any pending package whose upgrade would break a versioned dependency of a held installed package (reverse). + +The simplest way to get this is to let pacman be the resolver. After the -Sy, run read-only =pacman -Sup --ignore==. Add each package named in "unable to satisfy dependency ... required by X" (add X) or in "installing X (...) breaks dependency ... required by " (add X). Repeat until it exits clean. If it doesn't converge, or names a package outside the pending set (the L195 orphan case), refuse with the names. Never start a live -Syu that is known to fail. + +Related edits: +- Change the L155 test to "the --ignore list is blocked ∪ kernel ∪ dependency closure". Add fixtures for both directions: hyprutils so=12→13 with hyprlang depending on =13 (forward), and an unguarded so bump that a held hyprland depends on (reverse). +- State that the panel's deferred row (L128/L158) and the boot --complete transaction (L142) cover the whole closure in one transaction. +- Correct L97: under --noconfirm, one unresolvable dependent fails the whole transaction (pacman's skip prompt defaults to no), which is why the closure exists. + +Drop the finding's "--noconfirm re-installs ignored packages" and zfs-linux-lts claims. pacman doesn't behave that way on this path. + +Blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Design, Alternatives, Decisions, Implementation phases, Risks, Testing. Partially confirmed. This follows the narrowed finding and my loop, and pins down what the recommendation left loose: locale and flags, both message forms, the iteration bound, the single refresh from F14, and how closure members get a kind (stage A is kernel-kind, stage B is GPU-kind). + +One rule is made exact. A reverse 'breaks dependency' line adds the upgrading package only when the requirer is already held. When the requirer is unheld and not pending, the line names a package outside the pending set: the orphan case I want refused. Adding the upgrading package there would silently hold it, and its dependents, indefinitely, instead of naming the orphan. +:EVIDENCE: +Spec: L65 (-Syu --noconfirm --ignore=; "on the driven path it never fires"; the 724/730 proof of concept on a mesa day). L97 ("fails resolution instead of installing broken"). L142 (boot run applies exactly the deferred set, one transaction). L155 ("the --ignore list is exactly blocked ∪ kernel set"). L167 (AC: applies everything else, exits 0). L195 (orphan case: detect, name, stop). + +pacman v7.1.0 source (lib/libalpm deps.c and sync.c are byte-identical at installed commit 54d9411): +- deps.c:829: spkg = resolvedep(handle, missdep, handle->dbs_sync, *packages, 0). Prompt is 0, so ignored packages are skipped without the INSTALL_IGNOREPKG question (deps.c:662-667, 693-698). +- deps.c:762: prompt=1 only in alpm_find_dbs_satisfier. +- sync.c:439-468: unresolvable packages raise ALPM_QUESTION_REMOVE_PKGS. With no skip, GOTO_ERR(ALPM_ERR_UNSATISFIED_DEPS). +- callback.c:509: q->skip = noyes(...). util.c:1736-1738: under noconfirm the preset (0) is returned, so the transaction fails. +- sync.c:654-668 with deps.c:353-380: reverse-dependency break fails prepare, no prompt. +- src/pacman/sync.c:750/754: messages "unable to satisfy dependency '%s' required by %s" and "installing %s (%s) breaks dependency '%s' required by %s". + +Live system and cache on ratio: +- The installed hook /etc/pacman.d/hooks/hypr-live-update-guard.hook targets hyprutils and aquamarine but not hyprlang or hyprtoolkit. +- bsdtar .PKGINFO: hyprutils-0.13.1-1 provides libhyprutils.so=12-64, 0.14.0-1 provides =13-64. hyprlang-0.6.8-4 depends on =12-64, 0.6.8-5 on =13-64. +- aquamarine-0.14.0-1 provides libaquamarine.so=13-64, 0.15.0-2 provides =14-64. hyprtoolkit-0.5.4-4 depends on =13-64, 0.5.4-6 on =14-64. +- pacman.log: the whole hypr family co-upgraded with pkgrel-only rebuilds on 2026-04-05, 05-03, 07-21 and 09-09. +- expac -S: hyprlang, hyprlock, hyprwire, hyprpaper, hyprtoolkit, hyprland-guiutils and xdph all depend on libhyprutils.so=13-64. hyprland depends on libhyprlang.so=2-64, libhyprwire.so=3-64 and libhyprcursor.so=0-64 (all unguarded). +- No mesa-family package carries a versioned dependency on a guarded package. +- zfs-linux-lts is not installed (zfs-dkms 2.4.4-1 is). +:END: + +** DONE zfs-dkms and its pinned zfs-utils bypass the kernel hold :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Decision 5 L134-136 (consequence: 'an everyday UPDATE can never put velox into the unbootable state'); Design L65; Phase 1 L155; Risks L196 + +Only the kernel set (kernels plus -headers) is held on everyday runs, and the DKMS failure chain is treated only as a consequence of a kernel upgrade. Other DKMS-triggering packages go through the live --ignore run. + +Risk: An everyday run that moves zfs-dkms removes the working zfs.ko before the transaction and rebuilds it afterward, with nothing able to abort. A failed build (toolchain regression or full disk, both triggers Decision 5 lists) leaves velox with an initramfs that lacks zfs.ko. That is the unbootable state the kernel decision exists to prevent, reached on a run nobody chose. Holding zfs-dkms alone doesn't work either, because the exact zfs-utils pin makes the --ignore run fail dependency resolution. + +Recommended change: Recommended (needs my call because it widens Decision 5's held set): define a DKMS set and hold it next to the kernel set. +- L135 (Decision 5) and L65 (Design): add "plus the DKMS set: every package owning /usr/src/*/dkms.conf (pacman -Qoq), together with any dependency it pins with '=' (today zfs-dkms + zfs-utils), held whole on every everyday run and applied in --complete with the kernel set, before kernel-modules-check runs." +- L155 (Phase 1): change "the --ignore list is exactly blocked ∪ kernel set" to "... blocked ∪ kernel set ∪ DKMS set". Add a test that a pending zfs-dkms/zfs-utils upstream bump is held as a pair, and that a zfs-utils pkgrel-only bump is not held. +- L168: change "everything but the kernel set" to "everything but the kernel and DKMS sets". + +If I decline to widen the set, make two smaller edits instead: +- Rewrite L136 to "an everyday UPDATE never moves the kernel; a zfs-dkms move on an everyday run remains exposed, with the 05-zfs-snapshot/ZFSBootMenu fallback." +- Have the everyday run call kernel-modules-check whenever it moved a dkms.conf-owning package, and surface any failure as a named, do-not-reboot warning. + +Either way, fix L196: "the DKMS hooks are PostTransaction" becomes "old modules are removed PreTransaction and the rebuild runs PostTransaction". + +Blocking. Verification: confirmed. +Disposition: modified, folded into Decisions, Design, Implementation phases, Acceptance criteria, Risks, Testing. Confirmed and accepted. This widens Decision 5's held set, recorded as a body rewording plus a history line, not a reopening. The '= pin' half comes through F03's closure instead of a separate parse. On ratio, installed zfs-dkms 2.4.4-1 depends on zfs-utils=2.4.4. With zfs-dkms held, pacman reports that an upstream zfs-utils bump breaks that pin, so the closure adds zfs-utils exactly then and leaves a pkgrel-only bump alone. One mechanism produces the pair and pkgrel test outcomes the recommendation asks for. +:EVIDENCE: +- Spec L65: blocked set = guard list "plus the kernel set (every installed kernel with its -headers ...)". Spec L155: kernel set = "every linux* kernel package and its -headers"; test: "the --ignore list is exactly blocked ∪ kernel set". L136: "an everyday UPDATE can never put velox into the unbootable state". L168: TTY run "applies everything but the kernel set". L196: "(the DKMS hooks are PostTransaction)". +- /usr/share/libalpm/hooks/70-dkms-upgrade.hook: Operation=Upgrade, Target=usr/src/*/dkms.conf, When=PreTransaction, Exec=/usr/share/libalpm/scripts/dkms -D remove. +- 70-dkms-install.hook: same targets, When=PostTransaction, dkms install. +- 90-mkinitcpio-install.hook: Target=usr/src/*/dkms.conf, PostTransaction. +- /usr/share/libalpm/scripts/dkms main(): collects ERROR_MESSAGES and ends with "return 0". +- /usr/bin/mkinitcpio:409: warning "errors were encountered during the build. The image may not be complete." (the image is still written). /usr/lib/initcpio/install/zfs: map add_module ... zfs spl. +- pacman -Qoq /usr/src/*/dkms.conf: zfs-dkms. +- pacman -Si zfs-dkms: Repository archzfs, Depends On zfs-utils=2.4.4 lsb-release dkms. zfs-utils is installed at 2.4.4-3 (a pkgrel bump allowed by the pin). +- /var/log/ratio-upgrade.log:704-705: "dkms remove --no-depmod zfs/2.4.3 -k ...". Lines ~1487-1488: "dkms install --no-depmod zfs/2.4.4 -k ...". +- /var/log/pacman.log:20503-20504: zfs-utils and zfs-dkms upgraded 2.4.3 to 2.4.4 in the same run. +- Mitigation that makes it survivable rather than fatal: archsetup:2421-2432 writes 05-zfs-snapshot.hook (PreTransaction, Target=*), which runs before the 70 dkms hooks. +:END: + +** DONE kernel-modules-check fails permanently on ratio :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Phase 1 L155 ('for each kernel under /usr/lib/modules'); Decision 5 L134-135; Design L67; Risks L196; Phase 4 L164 + +Phase 1's gate requires that, for each kernel under /usr/lib/modules, dkms status reports every registered module as installed. Decision 5 and Design scope the check to 'the new kernel version'. The spec describes ratio as having two kernels. + +Risk: Every --complete run on ratio stops at the gate, whatever it installed, so ratio can never arm and the GPU/compositor set never lands through the designed path. Phase 4's 'confirm the ratio path matches' cannot pass, and the deferred row and stale freshness become permanent on one of the two daily drivers. + +Recommended change: Edit L155 (Phase 1) to match Decision 5. + +Replace "for each kernel under =/usr/lib/modules=, =dkms status= reports every registered module =installed= for it, and its initramfs is newer than its =vmlinuz=" with: + +"for each kernel version this =--complete= run installed or changed (its =/usr/lib/modules/= directory, mapped to =/boot/vmlinuz-= and its initramfs through =/pkgbase=), =dkms status -k = reports every registered module =installed=, and that kernel's initramfs is newer than its =vmlinuz=. Kernels the run did not touch are not checked. When the run changed no kernel, the DKMS and initramfs checks pass vacuously." + +Add one test to the Phase 1 test list: "an untouched kernel that is foreign, has no headers, and has DKMS modules only 'added' (ratio's =linux-lts-strix= shape) does not fail the gate." + +Correct "two kernels" at L134 and L196 to "three kernels, including =linux-lts-strix=, a foreign local build with no headers and no zfs module, which is the GRUB default." + +Related, optional: say in Phase 1 that the derived kernel set excludes packages absent from the sync repos, because "applied whole" cannot apply =linux-lts-strix= (pacman -Si: not found). + +Blocking. Verification: confirmed. +Disposition: modified, folded into Design, Decisions, Implementation phases, Risks, Testing. Confirmed. Scoping the gate to kernels this run changed is right, but on its own it leaves two holes, and both are closed here. +- F04 puts the DKMS set into --complete's stage 1. A zfs-dkms-only move rebuilds modules for kernels that were already installed, so those kernels join the check list. +- After a failed gate the new kernel is already installed. A second --complete changes no kernel, so it would pass vacuously, clear the CRIT row, and arm. The open failure's kernels are therefore re-checked against the recorded since. +:EVIDENCE: +Spec lines in docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org: +- L155: "for each kernel under =/usr/lib/modules=, =dkms status= reports every registered module =installed= for it" +- L135: "=dkms status= reports every DKMS module installed for the new kernel version" +- L67: "every DKMS module built for the new kernel" +- L134 and L196: "Ratio (btrfs root, two kernels, ...)" +- L164: "Confirm the ratio path matches." + +Live checks on ratio: +- uname -r: 6.18.25-1-lts-strix +- /proc/cmdline: BOOT_IMAGE=/@/boot/vmlinuz-linux-lts-strix +- /etc/default/grub:2: GRUB_DEFAULT="gnulinux-linux-lts-strix-advanced-..." +- ls /usr/lib/modules: 6.18.25-1-lts-strix, 6.18.54-1-lts, 7.2.7-arch1-1 +- pkgbase files: linux-lts-strix (no build/ dir), linux-lts (build/ present), linux (build/ present) +- pacman -Qm: linux-lts-strix 6.18.25-1, Packager "Unknown Packager" +- pacman -Si linux-lts-strix and pacman -Si linux-lts-strix-headers: "error: package ... was not found" +- dkms status: "zfs/2.4.4, 6.18.54-1-lts, x86_64: installed" and "zfs/2.4.4, 7.2.7-arch1-1, x86_64: installed" +- dkms status -k 6.18.25-1-lts-strix: "zfs/2.4.4: added" +- modinfo -n zfs: "ERROR: Module zfs not found." +- pacman -Qi zfs-dkms: 2.4.4-1 installed +- /var/log/ratio-upgrade.log:1536: "Building image from preset: /etc/mkinitcpio.d/linux-lts-strix.preset", written 2026-08-25, the same day the decisions were closed. +- docs/workflows/strix-soak-watch.org says the strix kernel is a standing local build, pinned as the GRUB default, until it is retired. + +Not a problem for the initramfs half: /boot/initramfs-linux-lts-strix.img (Sep 29) is newer than /boot/vmlinuz-linux-lts-strix (Apr 30), so that check passes. +:END: + +** DONE Phase 1 and Decision 2 hold the kernel only when a compositor is live :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Phase 1 L155; Decision 2 L114; Non-Goals L51; Decision 5 L135; AC L168 + +Phase 1 computes 'blocked set ... ∪ held-kernel set when a compositor is live'. Decision 2 calls a TTY run 'no compositor → nothing blocked → a full run', and Non-Goals L51 says 'holds them back on a live run'. Decision 5 and AC L168 hold the kernel on every everyday run, TTY included, and Phase 1's own test list says 'held whole on every everyday run'. The spec never says how liveness is detected. + +Risk: Read literally, the documented TTY fallback (C) moves velox's kernel with no DKMS gate. That is the unbootable path Decision 5 removes, and it would ship in v1 code paths and tests. If the script checks for a TTY instead of the process, a tty2 run with Hyprland still up skips the guarded patterns and the guard aborts. + +Recommended change: Five small text edits, no decision change: +1. L155: replace "blocked set (hook =Target= lines, version-aware) ∪ held-kernel set when a compositor is live" with "blocked set (hook =Target= lines, version-aware; only when Hyprland is running, tested exactly as the guard tests it: =pgrep -x Hyprland=) ∪ kernel set (always, on every run except =--complete=)". +2. L155 tests: add "with a liveness seam forced on and off (as =HYPR_GUARD_RUNNING= does for the guard), =--ignore= contains the kernel set on both branches and the blocked set only on the live branch." +3. L114: change "(no compositor → nothing blocked → a full run)" to "(no compositor → no guarded library blocked → everything but the kernel set; the kernel still lands only through =--complete=)". +4. L51: change "on a live run" to "on every run outside the dedicated =--complete= session". +5. Optional, one clause at L67 or L168: "'no compositor' means Hyprland is not running, not merely a console; Ctrl+Alt+F2 leaves it live on tty1." + +Blocking. Verification: confirmed. +Disposition: modified, folded into Goals and Non-Goals, Decisions, Design, Implementation phases, Acceptance criteria, Testing. Confirmed. The DKMS set (F04) joins the kernel set wherever this finding says 'kernel set'. + +The seam is the script's own name. The guard reads HYPR_GUARD_RUNNING from pacman's environment, which sudo resets, so sharing the name would suggest one variable drives both. + +--complete's stage 1 holds the GPU closure even from a TTY, because Decision 5 puts the GPU set after the gate. +:EVIDENCE: +Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org: +- L155: "pending set → blocked set (hook =Target= lines, version-aware) ∪ held-kernel set when a compositor is live"; the same paragraph's tests say "held whole on every everyday run". +- L114 (Decision 2, CLOSED 18:30): "C remains the manual fallback through the same script from a TTY (no compositor → nothing blocked → a full run)." +- L51: "the split script holds them back on a live run". +- L135 (Decision 5, CLOSED 18:45): "holds the kernel set ... on every everyday run, on both machines ... The kernel set lands only in the dedicated session". +- L168 (AC): "The same run with no compositor live (a TTY) applies everything but the kernel set". +- L12/L208: the 18:45 reversal to "held on every everyday run". +- No line in the spec names a liveness test (grep for compositor/live/tty/pgrep). + +Guard scripts/hypr-live-update-guard: +- L57-62: hyprland_running() uses the HYPR_GUARD_RUNNING override (documented at L32), else pgrep -x Hyprland. +- L67: hyprland_running || exit 0. +- L130-131: the banner says "from a TTY with Hyprland stopped: 1. Log out of Hyprland, or switch to a console (Ctrl+Alt+F2)". The Ctrl+Alt+F2 option leaves Hyprland running on tty1, so "TTY" and "no compositor" are not the same. +:END: + +** DONE Phase 3 panel action is the withdrawn ungated kernel path :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Scope tiers v1 L55; Phase 3 L161; against Decision 5 L135, Design L67, Phase 1 L155, history L208 + +L55 describes an 'apply on reboot' affordance that 'installs the held kernel live and arms the boot-time oneshot'. L161 says 'The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot', with 'the arm action's tests in maint'. Neither mentions kernel-modules-check. + +Risk: An implementer following Phase 3 builds a maint-side button that swaps the kernel, writes the flag and offers a reboot, with no DKMS, initramfs or snapshot check. On velox, a failed zfs-dkms build then reboots into a kernel whose initramfs can't import the pool. That is the exact failure Decision 5 was closed to prevent. + +Recommended change: L55: replace "an 'apply on reboot' affordance that installs the held kernel live and arms the boot-time oneshot for the GPU/compositor set" with "an 'apply on reboot' panel action that runs upgrade-guarded --complete (kernel set live, then the kernel-modules-check gate, then arm the boot-time oneshot for the GPU/compositor set)". + +L161: replace the first sentence with "The panel action runs upgrade-guarded --complete in the foreground (tty PROMPT mode). The script does every step: kernel set live, the gate, the arm flag and the reboot offer. maint never installs packages or writes the flag itself, and on a non-zero exit it shows the named gate failure and offers no reboot." Replace "the arm action's tests in maint" with "maint tests that the action's argv is upgrade-guarded --complete and that a non-zero exit leaves no reboot offer". + +L155: add one clause: "--complete exits non-zero when the gate fails." + +Optionally add an acceptance line after L169: "The panel's apply-on-reboot action, on a failed gate, neither arms nor offers a reboot." + +Blocking. Verification: confirmed. +Disposition: modified, folded into Goals and Non-Goals, Design, Implementation phases, Acceptance criteria, Testing. Confirmed, and reconciled with F40, which is also accepted. The action opens a detached terminal rather than using the lever runner's tty PROMPT mode. The panel then can't read --complete's exit code, so 'a non-zero exit leaves no reboot offer' becomes state-driven: F12's gate record hides REBOOT, and F34's arm flag shows it. foot gets --hold so the gate verdict stays readable after the script exits. The action moves into Phase 2 with its row (F16). +:EVIDENCE: +Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org: +- L55 (v1 scope): "an 'apply on reboot' affordance that installs the held kernel live and arms the boot-time oneshot for the GPU/compositor set". No gate. +- L161 (Phase 3): "The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot. ... the arm action's tests in maint". No --complete, no kernel-modules-check. +- L134-135 (Decision 5, closed): "I first wrote this as 'install the kernel live at apply-on-reboot'... A failed gate stops with the failure named and never reboots. The GPU/compositor set follows only after the gate passes... 'Install the kernel live at apply-on-reboot' is withdrawn." +- L208 (history, 18:45): "withdrew 'install live at apply-on-reboot'... added the kernel-modules-check gate to Phase 1 and the acceptance criteria". +- L67 (Design): the dedicated session is "run in the foreground from the panel's action or upgrade-guarded --complete in a terminal: first the kernel set, live... then the gate... If the gate fails the script stops there... and does not reboot... If it passes, the script arms a persistent flag... and offers to reboot". This contradicts L161. +- L155 (Phase 1): --complete runs the kernel set, then the gate, then arms; "--complete refuses to arm or reboot on that exit". Only kernel-modules-check is stated to exit non-zero. --complete's own exit code on a failed gate is never specified. +- L169 (acceptance): the DKMS-failure criterion covers only "A --complete run", not the panel action. +- Current maint (~/.dotfiles/maint/src/maint/doctor.py:161-176) runs lever steps by argv and has a tty PROMPT mode. Routing the action to a script argv fits the existing mechanism with no new maint-side package logic. +:END: + +** DONE Boot run's ExecStart and scope contradict the boot-scope decision, and no boot form is defined :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Design 'For the implementer' L69; Decision 4 L128; Decision 6 L138-143; Phase 1 --complete L155; Phase 3 L161; Readiness L189; Risks L193 + +L69: ExecStart runs, as the user, 'informant read ..., then =topgrade --only system= (or the equivalent =yay -Syu=), then =maint stamp topgrade= on success, then removes the flag'. Phase 3 L161: ExecStart 'runs the script's --complete form'. Phase 1 defines --complete as 'apply the kernel set live, run the gate, then arm ... or, with no compositor live, apply it directly'. Decision 6 L142: the boot run applies exactly the deferred GPU/compositor set in one pacman transaction, with no sweep and no kernel. L69 stamps on any success, whereas Decision 4 stamps only when the deferred set is empty. No text says whether the boot run reads the set recorded at arm time or re-derives it with a sync refresh, or that it rewrites the deferred-set record. + +Risk: Whichever text the implementer follows, the boot run does one of two bad things. It may install a kernel (including one published since arming) at the console, with no gate and no desktop to diagnose from: the velox unbootable chain Decision 5 forbids. Or it runs the gate at boot and never applies the GPU set on ratio. A refresh at boot can also pull in packages that were never armed. The stamp can fire while a kernel is still deferred, and the panel keeps the pre-boot 'N deferred' count after a successful run. + +Recommended change: Define one boot-only form in Phase 1 (L155) and point every boot reference at it. + +1. Add to L155's flag list: =--apply-armed=, the boot form. It reads the GPU/compositor package list that --complete's arm step writes into the flag file. It runs =informant read= if installed, then a single =sudo pacman -S --needed --noconfirm =. It does not use -y. A refresh followed by a targeted -S is itself a partial upgrade and could pull versions that were never armed. It runs no kernel set, no kernel-modules-check, no yay, and no topgrade. On success it rewrites the deferred-set state file and stamps topgrade_run if the deferred set is now empty. Add one sentence to --complete: "arming writes the GPU/compositor package list into the flag file". +2. Replace L69's ExecStart sentence, from "then =topgrade --only system=" through "on success", with: "runs =upgrade-guarded --apply-armed= as the user". Keep the unconditional flag removal. +3. In L142 and L161, change "the script's =--complete= form (scoped to the deferred GPU/compositor set)" to "the script's =--apply-armed= form". +4. In L71 and L93, replace "=maint apply-upgrade=" as the boot unit's path with "=upgrade-guarded=" (the archsetup script, per Decision 3). Its TTY form is --complete and its boot form is --apply-armed. +5. In L189, drop "topgrade =--only system=". In L193, change "The =--only system= scope" to "The recorded-set scope". +6. Add these Phase 1 tests: the --apply-armed transaction targets exactly the recorded list, never includes a kernel-set package, and never passes -y. + +Blocking. Verification: confirmed. +Disposition: modified, folded into Design, Decisions, Implementation phases, Readiness dimensions, Risks, Testing. Confirmed and adopted, reconciled with two of my calls. +- Disarm-at-ExecStartPre means the boot form reads the copy that ExecStartPre moved into /run, not the flag path. +- Exact versions in the flag mean the transaction uses versioned targets, filtered so it never downgrades. + +One safety check is added. pacman -S resolves dependencies from the sync db, so a GPU target needing a newer kernel-set or DKMS-set package would pull that package in at boot, ungated. A read-only, offline pacman -Sp of the exact targets runs before the transaction and refuses that case. +:EVIDENCE: +Spec text: +- L69: "Its =ExecStart= runs, as the user: =informant read= …, then =topgrade --only system= (or the equivalent =yay -Syu=), then =maint stamp topgrade= on success, then removes the flag unconditionally". +- L71: "a small =maint apply-upgrade= … which both the boot unit and an interactive TTY run call". +- L93: "the same =maint apply-upgrade= path serves both an interactive TTY run and the boot unit". +- L141: "The first draft phrased this as =topgrade --only system=; with the split script that wording is stale". +- L142: "the boot oneshot runs the script's =--complete= form scoped to the deferred GPU/compositor set: one pacman transaction, no ecosystem sweep, no kernel". +- L155: "=--complete= (the dedicated-session form: apply the kernel set live, run the gate, then arm the GPU/compositor set or, with no compositor live, apply it directly; stamp when the deferred set is empty)". +- L161: "=ExecStart= runs the script's =--complete= form as the user". +- L189: "External APIs & deps: topgrade =--only system=". +- L193: "The =--only system= scope … shrink this to near zero". + +Git history: +- =git show 77447d0:…spec.org= L64 has the identical ExecStart sentence, so it was never revised. +- The uncommitted diff touches only the script name and the containers flag, not L69. + +Live system: +- ~/.dotfiles/common/.config/topgrade.toml:131 sets =arch_package_manager = "yay"=, and L134 =# yay_arguments= is commented out. +- /etc/pacman.conf:25 has =IgnorePkg = bridge-utils= and nothing else, so no kernel is ignored. +- =pacman -Q= shows linux 7.2.7, linux-lts 6.18.54, and linux-lts-strix 6.18.25 (no headers). +- =dkms status= shows zfs installed for 6.18.54-1-lts and 7.2.7-arch1-1 only. +- =topgrade --help= confirms =--only =. +:END: + +** DONE Boot run's package source and network dependency are undefined :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L65-69; Decision 6 L138-143; Phase 3 L161; Readiness Performance L182 + +The spec never says whether the boot run refreshes the sync db, downloads packages or waits for the network. Nothing pre-downloads the GPU set at arm time, and the unit declares no network ordering. The spec also never covers a stale arm (flag set, then days of further everyday runs before the reboot) or a flag that is still set when nothing is pending any more. + +Risk: If the boot run refreshes and downloads, an armed boot depends on pre-login network and mirrors, network-online delays the login prompt on every boot, and a refresh can pull new versions (including a kernel) into the boot transaction. If it doesn't, there is no rule for where packages come from. With no network the run fails every time and disarms (Decision 7), so the GPU set never lands. + +Recommended change: 1. In Design L69, replace "then =topgrade --only system= (or the equivalent =yay -Syu=)" with "then =upgrade-guarded --complete= in its boot form". In L189 and L193, drop the =--only system= references the same way. + +2. Add one paragraph to Decision 6 (or Phase 3): + +"The boot form never refreshes the sync db and has no network dependency; the unit declares no network-online ordering. When --complete arms, it runs =pacman -Sw --noconfirm = while the network is up and writes the exact pkg=version list into the arm flag. While armed, each everyday run re-runs -Sw for the deferred set after its own refresh and rewrites that list. At boot the run skips the kernel step and the gate. It installs the recorded list with =pacman -S --noconfirm= against the existing db, so pacman takes every file from the cache. A missing cache file fails fast and names the package. If every recorded package is already installed at or past its recorded version, the run is a no-op: it clears the flag, rewrites the deferred-set record, and stamps only if Decision 4's empty-deferred-set rule is met." + +3. Add a Phase 1 test that the boot form makes no -y/refresh call. Add an acceptance criterion that an armed boot with networking down still applies the set. + +Blocking. Verification: confirmed. +Disposition: modified, folded into Design, Decisions, Implementation phases, Acceptance criteria, Readiness dimensions, Testing. Confirmed; this matches my boot-source call, including the recommendation's re-download while armed. That keeps the flag's versions equal to whatever the system db each everyday run synced. Three gaps are closed: +- An armed everyday run that can't compute the closure, pass the check, or re-download disarms. Otherwise it would leave a list the db no longer matches. +- The boot form never downgrades. A recorded version below the installed one would downgrade under --noconfirm, so such entries are dropped. +- F08's dependency check runs at arm time and again at boot, so no kernel can ride in as a dependency. +:EVIDENCE: +Spec: +- L69 names =topgrade --only system= / =yay -Syu= as the boot ExecStart. +- L141 (Decision 6 context) calls that wording "stale". L142 scopes the run to =--complete= with "no kernel". +- L155 defines =--complete= as starting with "apply the kernel set live". +- L189 and L193 still cite =--only system=. +- L182 claims negligible cost when the flag is absent. +- A grep of the spec for Sw|download|network|online|mirror|refresh|offline finds nothing about the boot run's package source or network. The only refresh mention is the live run at L65. + +Live system (ratio): +- =checkupdates= shows guarded packages pending: mesa 1:26.2.4-1, vulkan-radeon, hyprland 0.56.2-4, and kernels linux 7.2.8 and linux-lts 6.18.55. None of them is in /var/cache/pacman/pkg. +- maint uses plain =checkupdates= with no -d (maint/src/maint/probes/updates.py:152, doctor.py:83). /usr/bin/checkupdates downloads only with -d (L63, L173). +- No timer pre-downloads. The only timers are snapper, man-db, tmpfiles, shadow, logrotate, keyring-wkd, paccache, reflector and btrfs-scrub. +- Boot timing (monotonic µs): basic.target 7776379, NetworkManager-wait-online 8975960→14740538, network-online.target 14740879, getty@tty1 35133024. +- NetworkManager-wait-online.service is enabled, with =ExecStart=/usr/bin/nm-online -s -q= and =NM_ONLINE_TIMEOUT=60=. +- network-online.target is already pulled in by docker, minidlna, reflector and archlinux-keyring-wkd-sync. +- systemd.unit(5), Conditions and Asserts: conditions "are checked at the time the queued start job is to be executed. The ordering dependencies are still respected... neither condition nor assertion expressions are suitable for conditionalizing unit dependencies." +- reflector.timer: OnCalendar=weekly, Persistent=true (tangential). +:END: + +** DONE Timeout is unspecified, and SIGTERM kills pacman mid-transaction and leaves db.lck :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L69; Phase 3 L161; AC L171; Readiness Errors L179, Performance L182, Config surface L185; Risks L192 + +The spec treats a timeout as benign: 'worst case is a bounded delay, then a normal session' (L192), and AC L171 says a timeout lets the machine boot into Hyprland. The TimeoutStartSec value is left open (L185). Nothing defines the system's state after pacman is killed mid-commit. + +Risk: A timeout during extraction leaves /var/lib/pacman/db.lck behind and a half-extracted guarded set (mixed mesa/libdrm files). getty then starts and Hyprland launches against mixed libraries. Every later pacman, yay or upgrade-guarded run fails with 'unable to lock database' until someone removes the lock by hand, and the panel shows only the old deferred count. + +Recommended change: 1. Phase 3 (L161) and Design (L69): give archsetup-boot-upgrade.service KillSignal=SIGINT, so a timeout reaches pacman's interrupt handler, which stops at a package boundary and releases /var/lib/pacman/db.lck itself. Give TimeoutStartSec a concrete, generous value (e.g. 20min) so it fires only on a real hang. + +2. Move the disarm out of the tail of ExecStart. Remove the flag with ExecStartPre=+/usr/bin/rm -f /var/lib/archsetup/apply-upgrade-on-boot (or in ExecStopPost=), so success, failure and timeout all disarm. Reword Decision L149 to match ("removed at the start of the attempt" or "in ExecStopPost"). + +3. Define the interrupted outcome. If the transaction did not complete, or db.lck survives, the unit writes "interrupted" plus the remedy into the deferred-set state file the panel already renders. It never deletes db.lck on its own. + +4. Correct the false claims. In Risks L194, replace "one atomic transaction, so there is no half-swapped library" with: pacman is atomic per package only, and only when allowed to stop at a package boundary, which is why the unit uses KillSignal=SIGINT. Drop "worst case is a bounded delay" from L192. + +5. AC L171: add that after a timeout no db.lck is left behind, the flag is cleared even though ExecStart was killed, and the panel names the interruption. + +Blocking. Verification: confirmed. +Disposition: modified, folded into Design, Decisions, Implementation phases, Acceptance criteria, Readiness dimensions, Risks. Confirmed; this matches my call: SIGINT, 20min, disarm at the start, an interrupted record, and never deleting db.lck. Two reconciliations: +- The flag now carries the list, so ExecStartPre moves it into the unit's RuntimeDirectory instead of running rm. That disarms just as completely. +- The interrupted record is pre-written before the boot transaction, not written only by a signal trap. A trap can't run on the SIGKILL that follows TimeoutStopSec, or on power loss. + +The ExecStopPost alternative is dropped. +:EVIDENCE: +Spec: +- L69 / L161: "TimeoutStartSec bounded … ExecStart runs … then removes the flag unconditionally". +- L149: the flag is removed "at the end of its attempt". +- L171: the AC requires that on "a timeout … the flag is cleared" and the machine boots into Hyprland. +- L179: "every failure path lands in … re-arm to retry". +- L192: "worst case is a bounded delay, then a normal session". +- L194: "each pacman run is one atomic transaction, so there is no half-swapped library". +- grep of the spec for killsignal|db.lck|interrupt|SIGTERM|SIGINT|ExecStopPost|atomic finds nothing except the L194 atomicity claim. + +systemd (v262): +- man systemd.service: TimeoutStartSec "Defaults to DefaultTimeoutStartSec= … except when Type=oneshot is used, in which case the timeout is disabled by default". TimeoutStartFailureMode "default to terminate … sending the signal specified in KillSignal= (defaults to SIGTERM)". +- man systemd.kill: KillMode=control-group kills "all remaining processes in the control group". + +pacman 7.1.0.r9: +- strings /usr/bin/pacman shows only "Interrupt signal received" and "Hangup signal received" handler messages. No SIGTERM handler is present. +- libalpm.so.16 contains alpm_trans_interrupt and "transaction interrupted". So SIGINT has a graceful path: it stops at a package boundary during commit and releases the lock; during download it unlocks and exits. + +Live system: ratio's /etc/pacman.d/hooks holds hypr-live-update-guard.hook and 99-grub-sync-efi.hook, and /usr/share/libalpm/hooks/05-snap-pac-pre.hook runs PreTransaction, so a kill can land while the lock is held before commit even starts. +:END: + +** DONE Bare =informant read= hangs unattended and can't clear the news hook when run as the user :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L65, L69; Phase 1 L155; AC L172; Reuse L183; Readiness External deps L189 + +The live script and the boot oneshot (which runs 'as the user') both run a bare =informant read= to clear the news hook before the transaction. L189 lists informant as 'verified present'. + +Risk: The boot oneshot has no stdin, so the bare command either blocks until TimeoutStartSec or exits without marking anything read. Run as the user it can't save read state either, so informant's AbortOnFail hook aborts the transaction. The boot run then fails every time Arch has published news, consuming the arm. AC L172 fails, and the panel-driven live run can hang or abort the same way. + +Recommended change: L65, L155: replace "=informant read=" with "=sudo informant read --all=" (run only when =command -v informant= succeeds; failure tolerated, as at archsetup:1127-1129). Optionally log =informant list --unread= first so the news isn't dropped unseen. + +L69: change "Its =ExecStart= runs, as the user: =informant read= (...)" to say that the news-clear and pacman steps run as root via sudo, and the news-clear step is =sudo informant read --all= when installed. Add the reason: =/var/lib/informant.dat= is 0664 root:informant, the user isn't in that group, and a bare read prompts between items, so it crashes on null stdin or blocks on a tty. + +AC L172: change "(=informant read= precedes the transaction)" to "(=sudo informant read --all= precedes the transaction; verified with two or more unread items and no stdin)". + +L189 (optional): note that informant is present on velox only; ratio doesn't have it installed. + +Blocking. Verification: confirmed. +Disposition: modified, folded into Design, Implementation phases, Acceptance criteria, Readiness dimensions, Testing. Adopted, with one hardening. informant 0.6.0's feed fetch calls requests with no timeout. On a half-up network, from the panel or early in boot (the unit has no network ordering), the call could hang until the lever's 3600s timeout or the unit's 20-minute timeout. timeout 60 bounds it, and the run carries on either way. The optional 'log informant list --unread first' is superseded by F39's yay -Pwq capture. +:EVIDENCE: +- Spec L65 "clears the news hook (=informant read=) where installed"; L69 "Its =ExecStart= runs, as the user: =informant read= ..."; L155 "=informant read= if present →"; L172 AC "Unread Arch news does not wedge the boot run (=informant read= precedes the transaction)"; L189 "all verified present on velox". +- informant 0.6.0-2 source, taken from the archangel build tree at ~/code/archangel/work/x86_64/airootfs/usr/lib/python3.14/site-packages/informant/: + - informant.py:132-145 is the bare-read loop, which calls =ui.prompt_yes_no('Read next item?')= between items. + - informant.py:146 is =fs.save_datfile()=, the only save, after the loop. + - ui.py:69 is =input(...)=, which raises EOFError on null stdin and blocks on a tty. + - informant.py:117-119 is =--all=, which marks everything read without prompting. + - config.py:19 sets the save file, =FILE_DEFAULT='/var/lib/informant.dat'=. + - file.py:45-53 catches PermissionError on that save, prints an error, and calls =sys.exit(255)=. +- The rest of the package, from the same tree: + - usr/lib/tmpfiles.d/informant.conf: =f /var/lib/informant.dat 0664 root informant= + - usr/share/libalpm/hooks/00-informant.hook: =Target = *=, =When = PreTransaction=, =Exec = /usr/bin/informant check=, =AbortOnFail= + - informant.py:72-98: =check= exits with the unread count. +- The installer's user setup (archsetup:1403-1405) adds =sys,adm,network,...,users= and does not add =informant=. The installer's own call, at archsetup:1121-1129 (commit cc0f7b5), is =informant read --all ... || true=, run as root, with the comment "a bare =informant read= is interactive and would hang an unattended run". +- maint's =cmd.run= (maint/src/maint/cmd.py:22) uses =subprocess.run(..., capture_output=True)= with stdin inherited. A prompt there is either invisible or gets EOF. +- On ratio, =pacman -Q informant= reports "package 'informant' was not found". =pacman -Si= shows informant 0.6.0-2 in [extra]. +:END: + +** DONE Panel offers REBOOT after a failed --complete gate :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Design 'For the user' L67; Decision 5 L134-136; Phase 2 L158; Phase 3 L161; AC L169 + +The spec says a failed gate 'stops there, names what failed, and does not reboot'. That holds only for the script. The spec never says what the panel shows after the kernel set has landed live and the gate has failed. + +Risk: On velox, the panel shows a red REBOOT key at the exact moment the gate has said a reboot can't import the pool. The message naming the failure is gone once the wall is cleared or the panel is reopened. + +Recommended change: Phase 1 (L155): +- After the kernel set lands, if kernel-modules-check fails, --complete records the failure in the deferred-state file as gate: {ok: false, kernel: , failed: , at: }. +- It prints the failing item as its last stderr line. +- The next --complete or kernel-modules-check run that passes clears that record. + +Phase 2 (L158) and Phase 3 (L161): +- While gate.ok is false, the deferred probe renders a CRIT row: "kernel installed — gate failed: — do not reboot". The usual "apply on reboot" wording is not shown. +- The panel hides the REBOOT key, overriding reboot_required and offer_reboot. +- Rewrite Phase 3's panel-action sentence to: runs --complete, which installs the kernel set, runs the gate, and arms and offers reboot only when the gate passes. + +Acceptance criteria: add "After a failed gate, the panel names the failing item persistently (it survives a wall dismiss and a panel reopen) and offers no REBOOT until a later gate run passes." + +Blocking. Verification: confirmed. +Disposition: modified, folded into Design, Implementation phases, Acceptance criteria, Testing. Confirmed, with two narrowings. +- Only upgrade-guarded writes the record. A passing standalone kernel-modules-check therefore doesn't clear the failure, and the gate stays a stateless reader. +- Clearing requires a real re-check of the recorded kernels against the recorded since. A --complete after a hand fix changes no kernel and would otherwise pass vacuously. + +APPLY stays visible while the failure is open, even when nothing else is deferred, so the re-check is reachable from the panel. +:EVIDENCE: +Spec: +- L67: if the gate fails, the script "stops there, names what failed, and does not reboot". +- L135: "A failed gate stops with the failure named and never reboots." +- L155: --complete "refuses to arm or reboot on that exit". +- L128/L158: the panel row reads "N deferred — apply on reboot". +- L161: the Phase 3 panel action "installs the held-kernel set live, writes the persistent arm flag, and offers to reboot" (no gate, no failure branch). +- L169: AC on a failed DKMS build. +- No occurrence of reboot_required anywhere in the spec. + +Live system: +- pacman -Qo /usr/lib/modules/6.18.25-1-lts-strix/modules.alias prints "No package owns". +- /usr/share/libalpm/hooks/60-depmod.hook (PostTransaction, Remove/Upgrade on usr/lib/modules/*/) runs /usr/share/libalpm/scripts/depmod. That script does rm -f on modules.{alias,dep,...} and then rmdir --ignore-fail-on-non-empty "$f" for a kernel directory that no longer has modules.order. +- 71-dkms-remove.hook is present. + +maint (in ~/.dotfiles/maint/src/maint): +- probes/packages.py:131-146: required = not os.path.isdir(/usr/lib/modules/). +- panel.py:506-510: reboot_key_visible() returns the reboot_required value. +- gui.py:491-495: red REBOOT armed key when that is visible. +- gui.py:1510-1511: offer_reboot only after an update or topgrade "ok". +- remedies.py:314-320: the reboot remedy's metric_ids are ["reboot_required"]. +- doctor.py:166: return False, f"exit {proc.returncode}", stderr[-500:]. +- panel.py:308 and 389-396: the wall is a session-only log, and wall_clear() drops it. +:END: + +** DONE UPDATE/TOPGRADE flag mapping and the stamp rule for partial runs are undefined :blocking: +CLOSED: [2026-10-05 Mon 09:20] +Where: Scope tiers L55; Design L65, L67; Decision 4 L128; Phase 1 L155 (flags --no-topgrade, --no-aur; 'stamp only when nothing was deferred → exit 0 on a successful live part'); Phase 2 L158 + +Both levers 'change their argv to the script', with no flags given. Today they are different remedies: UPDATE is a repo+AUR system update, and TOPGRADE is the ecosystem sweep that owns topgrade_age. Phase 1 defines --no-topgrade and --no-aur but names no consumer for them. The stamp predicate is only 'deferred set empty'. Nothing says what a run stamps or exits with when it skipped the sweep or the AUR, or when a sub-step failed (the pacman orphan stop in Risks L195, the AUR build, a topgrade step such as npm). + +Risk: There are two plausible implementations. In one, UPDATE silently gains the full ecosystem sweep and the two keys become duplicates. In the other, UPDATE runs --no-topgrade and the script stamps topgrade freshness for a run that never swept the ecosystems. Either way, a failed AUR or topgrade step with an empty deferred set reads 'current' under the state framing of Decision 1. + +Recommended change: Phase 2 (L158): state the lever mapping. UPDATE runs =upgrade-guarded --no-topgrade=; TOPGRADE runs =upgrade-guarded=. Add that on the driven path the script is the only thing that writes topgrade_run. That means removing doctor.py's rid=="topgrade" stamp, and having the script call the real topgrade binary (or tell the PATH wrapper not to stamp) so the inner sweep cannot stamp. Decision 4 (L128) and Phase 1 (L155): define the predicate for the everyday run. It stamps topgrade_run only when the sweep ran (no --no-topgrade or --no-aur), every step exited 0, and the deferred set is empty. If pacman fails, it skips the later steps, still writes the state file, exits non-zero, and does not stamp. --complete and the boot oneshot keep their own rule: stamp when their transaction succeeds and the deferred set is empty, with no sweep required. Add a Phase 1 test that a run with a non-empty deferred set leaves topgrade_run unwritten, through both the lever and the inner topgrade call. + +Blocking. Verification: confirmed. +Disposition: modified, folded into Decisions, Implementation phases, Testing. Accepted, with the completion-form rule tightened in two places. +- A completion form stamps only when it started with something to complete. Otherwise a --complete that finds nothing would mint freshness after an UPDATE that never ran the sweep. +- No form stamps while a gate failure is open, so TOPGRADE can't mark a do-not-reboot machine fresh. +:EVIDENCE: +- Spec L65 and L155: the script exits 0 "whatever was deferred" / "on a successful live part". L155: "stamp only when nothing was deferred". Flags --no-topgrade and --no-aur are defined with no consumer. +- Spec L55, L67, L114, L121, L158: UPDATE and TOPGRADE both "change their argv to the script", with no per-lever flags. +- Spec L128 (Decision 4): stamp topgrade_run only when the deferred set is empty. L136: the deferred count "carries a kernel most days". +- Spec L32: names both existing writers. No phase removes or neutralizes either one. +- ~/.dotfiles/maint/src/maint/remedies.py:294-312: update = ["yay","-Syu","--noconfirm"] with metric_ids updates_repo/updates_aur/cve_advisories; topgrade = ["topgrade","--disable","git_repos","-y"] with metric_ids topgrade_age. +- ~/.dotfiles/maint/src/maint/doctor.py:169-170: =if rid == "topgrade": cache.put("topgrade_run", {"at": time.time()})= runs on any exit-0 step. +- ~/.dotfiles/hyprland/.local/bin/topgrade:38-42: runs the real binary, then =maint stamp topgrade= when rc is 0. Its header says the maint lever resolves this wrapper too. +- =which -a topgrade= lists ~/.local/bin/topgrade ahead of /usr/bin/topgrade. +- Spec L142-143 (Decision 6): the boot/--complete path runs no sweep, and its stamp "reflects the guarded set specifically", so a predicate that requires the sweep would contradict it. +:END: + +** DONE Blocked-set computation reads checkupdates' private db, not the db pacman -Syu uses +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L65 ('refreshes the sync db and reads the pending set (checkupdates)'; 'version-aware'); Phase 1 L155 and its tests + +The script 'refreshes the sync db and reads the pending set (checkupdates)'. It then computes the blocked set 'version-aware the way the guard is', by intersecting with the hook Targets, and passes it to =pacman -Syu --ignore==. + +Risk: A guard-style version check that runs before the real refresh leaves mesa, hyprland and vulkan-radeon out of --ignore. -Syu then pulls them in and the live guard aborts the whole transaction, so AC L167 fails in the state ratio is in right now. An update that appears in the second sync but not the first has the same effect. Treating exit 2 as a failure fails every day with nothing pending, and implementers will differ on how =mesa-*= becomes ignore entries. + +Recommended change: Design L65: +- Replace "refreshes the sync db and reads the pending set (=checkupdates=); computes the *blocked set* = the guard's own trigger list (read from the installed hook's =Target= lines, ..., and version-aware the way the guard is — a same-version reinstall is not a swap)" with: "builds the *ignore list* = the installed hook's =Target= patterns passed verbatim to =--ignore= (IgnorePkg accepts shell globs, so =mesa-*= needs no expansion; a pattern with nothing pending is a no-op, and =-Syu= never reinstalls at the same version, so no version lookup is needed here) ∪ the kernel set". +- After the yay step, add: "the *deferred set* is computed after the transaction from =pacman -Qu=, which reads the db =-Syu= just synced, filtered by the same patterns ∪ the kernel set." + +Phase 1 L155: +- Change "blocked set (hook =Target= lines, version-aware)" to "ignore list (hook =Target= patterns verbatim) ∪ kernel set". +- Change the test bullet "blocked-set computation against a fixture hook and version map" to "the =--ignore= list is exactly the fixture hook's Target patterns ∪ the kernel set; the deferred set is derived from a fake post-transaction =pacman -Qu=". + +Add one sentence: "=checkupdates= is used only by =--dry-run=; its exit 2 means an empty pending set, not a failure." + +Should-fix, not blocking. Verification: confirmed. +Disposition: modified, folded into Design, Implementation phases, Testing. Confirmed. The recommended text keeps pacman -Syu, which would sync a second time after F03's closure was computed against the first sync. That is the same race this finding names. Instead: one sudo pacman -Sy, the closure on that db, then sudo pacman -Su with no -y. +:EVIDENCE: +Commands below were run on ratio, 2026-10-05. +- =checkupdates --help= (v1.13.1): the default db is =${TMPDIR:-/tmp}/checkup-db-${UID}=. In /usr/bin/checkupdates, :150 runs =pacman -Sy --dbpath "$CHECKUPDATES_DB"=, :155 drops =[...]= lines, and :181 is =exit 2= when nothing is pending. +- =checkupdates -n= filtered to hook Targets and kernels lists: =hyprland 0.56.2-3 -> 0.56.2-4=, =mesa 1:26.2.3-1 -> 1:26.2.4-1=, =vulkan-radeon 1:26.2.3-1 -> 1:26.2.4-1=, plus linux, linux-headers, linux-lts and linux-lts-headers. +- The guard-style lookup (=pacman -Q= vs =expac -S %v=, scripts/hypr-live-update-guard:78-82) reports mesa, hyprland and vulkan-radeon as installed = candidate, i.e. same version. +- /var/lib/pacman/sync/extra.db is dated 29 Sep 08:21. =pacman -Qu= against the system db shows 0 updates; checkupdates shows 395. +- tests/hypr-live-update-guard/test_hypr_live_update_guard.py:15-16,52: HYPR_GUARD_VERSIONS is a "pkg installed candidate" map. This is the "version map" the spec's Phase 1 test (L155) mirrors. +- pacman.conf(5) on IgnorePkg: "Shell-style glob patterns are allowed." pacman(8) =--ignore= adds packages to the same ignore list. +- /etc/pacman.d/hooks/hypr-live-update-guard.hook Targets: mesa, mesa-*, wayland, libdrm, libglvnd, hyprland, aquamarine, hyprutils, hyprgraphics, vulkan-radeon, vulkan-intel, vulkan-mesa-layers, nvidia-utils, lib32-nvidia-utils, xorg-xwayland. None use =!=. +:END: + +** DONE Flag removal at the end of ExecStart never runs on timeout, kill or power loss +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L69; Decision 7 L145-150; Phase 3 L161; AC L171; Readiness L178 + +L69 and L161: ExecStart runs the steps and then 'removes the flag unconditionally'. The flag directory /var/lib/archsetup is installer-owned (L178), and the unit runs as the user. + +Risk: Any timeout, kill, power loss or permissions slip leaves the flag in place, so every later boot re-runs the upgrade ahead of getty. That is the retry-every-boot outage Decision 7 exists to prevent, and it contradicts AC L171. The panel's RESTART key could also re-run the unit under a live session. + +Recommended change: In Design L69 and Phase 3 L161, replace "then removes the flag unconditionally" with: "the unit removes the flag before the transaction starts, privileged: ExecStartPre=+/usr/bin/rm -f /var/lib/archsetup/apply-upgrade-on-boot (the + runs it as root regardless of User=); ExecStart never touches the flag, so a timeout kill, SIGKILL, or power loss can't leave it armed." + +In Decision 7 L149, change "at the end of its attempt" to "at the start of its attempt". This is the same one-shot intent, and it now holds through a timeout. + +In L69 or L178, state that the flag is root-owned and that the panel's arm action creates it via sudo. + +In Phase 3 L161, add a forced-timeout case to the manual boot test: a short TimeoutStartSec plus a hung step, then check that the flag is gone, the next boot skips the unit, and Hyprland starts. + +If removal at the end is kept instead, move it to ExecStopPost=+/usr/bin/rm -f (that covers timeouts and failures, but not power loss). + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Design, Decisions, Implementation phases, Readiness dimensions, Testing. Partially confirmed; this follows the narrowing and my disarm-at-start call, with two reconciliations. +- The flag now carries the package list (F09), so ExecStartPre moves it rather than deleting it. +- maint never writes the flag (F07), so 'the panel's arm action creates it via sudo' becomes 'upgrade-guarded creates it via sudo'. +:EVIDENCE: +Spec L69: "Its ExecStart runs, as the user: informant read ..., then topgrade --only system ..., then maint stamp topgrade on success, then removes the flag unconditionally". L161: "ExecStart runs the script's --complete form as the user and removes the flag unconditionally". L149: "a persistent flag (/var/lib/archsetup/) that the boot unit removes unconditionally at the end of its attempt — success or failure disarms". L171: "A boot-upgrade failure (a failed step, a timeout, an aborted transaction) never blocks the session ... the flag is cleared". + +man systemd.service: +- TimeoutStartSec: "the service will be considered failed and will be shut down again" +- TimeoutStartFailureMode: the default "terminate" sends KillSignal (SIGTERM) +- ExecStopPost: "commands specified with this setting are invoked when a service failed to start up correctly and is shut down again ... recommended ... for clean-up operations" + +Live repro: bash -c "sleep 30; rm -f flag", SIGTERM to the child and the shell, exit=143, and the flag still exists. + +Ownership: ls -ld /var/lib/archsetup gives drwxr-xr-x root (root-owned, 0755); the installer uses it at archsetup:257 (state_dir="/var/lib/archsetup/state"). + +Guard backstop: scripts/hypr-live-update-guard:57-67 (hyprland_running uses pgrep -x Hyprland; it exits 0 only when Hyprland isn't running) blocks a live re-run. + +The panel paths exist as cited: maint/src/maint/probes/systemd.py:70-100 (failed_units WARN) and remedies.py:323-337 (unit_restart/unit_reset). +:END: + +** DONE Deferred row severity, lever, band and bar impact are undefined +CLOSED: [2026-10-05 Mon 09:20] +Where: Decision 4 L128; Decision 5 consequences L136; Phase 2 L158; ACs L166-174 + +The spec only says the panel renders 'N deferred — apply on reboot' 'as its own row'. It gives no metric id, category or band, severity rule, lever or evidence shape. It also concedes that the count 'carries a kernel most days'. + +Risk: A WARN without a lever turns the waybar glyph gold almost every day, which just moves the 'permanently stale' symptom from the Summary (L28) somewhere else. A WARN with a lever adds a permanent ATTN. An OK never escalates the hold. The implementer has to invent all of this. + +Recommended change: Add two sentences to the end of Phase 2 (L158): + +"The row is metric upgrade_deferred, category updates (so it lands in the PACKAGES band). Its value is the deferred count and its evidence is the deferred rows (name, old, new). It grades OK when the set is empty and [WARN | OK — owner picks] when the set is non-empty. Phase 3's apply action is its lever, registered always:True so the action shows at any severity. A levered row never colours the waybar glyph. Age-based escalation stays the L197 follow-up." + +If WARN is chosen, also say that the row must not ship before its lever. Otherwise, in the window between Phase 2 and Phase 3, a lever-less WARN colours the glyph every day. Either land the row and the apply action in the same commit, or grade the row OK until Phase 3. + +Add one AC after L167: "With a non-empty deferred set and Hyprland live, the panel's deferred row offers the apply action and the waybar glyph colour is unchanged by the deferral." + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Implementation phases, Acceptance criteria. Partially confirmed; this follows the narrowing. The open WARN-or-OK choice is resolved as OK for a plain deferral. The kernel is held most days, so WARN on mere presence would colour the glyph daily, which just moves the original stale-metric symptom onto another row; topgrade_age already carries the cadence nag. + +WARN is kept for states that need action (F10, F29, F36), and CRIT for a gate failure (F12). Because WARN is still reachable, the row ships in the same commit as its lever. +:EVIDENCE: +Spec L128: the deferred set is rendered "as its own state ('6 deferred — apply on reboot')". L158 (Phase 2): "A new probe reads the deferred-set state file and the panel renders 'N deferred — apply on reboot' as its own row". There is no severity, and no AC between L166 and L174 covers it. + +The lever does exist in the spec. L67 says "from the panel's action or upgrade-guarded --complete". L161 says "The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot." Escalation is explicitly deferred at L197: "a stale-kernel age in maint is a possible follow-up." + +Code: +- ~/.dotfiles/maint/src/maint/remedies.py:535-545 (attach_levers): a lever is attached when the metric is WARN/CRIT, or when the remedy is marked "always". +- remedies.py:294-312: the update and topgrade remedies are always:True. updates_repo is graded by grade_high(len, pending_warn=50) in probes/updates.py:53-61, so it reads OK below 50 with an always-lever. +- ~/.dotfiles/maint/src/maint/indicator.py:41-57: diag is the metrics without levers, and only diag colours the glyph. Levered off-nominal metrics count as "actionable". +- ~/.dotfiles/maint/src/maint/panel.py:36-47 (BAND_BY_CATEGORY maps updates to packages), 449-452 (band_lamp is the worst severity), 485-489 (attn_count counts every WARN/CRIT). + +Live data: ~/.local/state/maint/updates_repo.json (written 04:55 today, 395 rows) contains hyprland, linux, linux-headers, linux-lts, linux-lts-headers, mesa and vulkan-radeon. hyprland, mesa and vulkan-radeon match the /etc/pacman.d/hooks/hypr-live-update-guard.hook Target lines. vulkan-icd-loader and vulkan-mesa-implicit-layers are pending too, but they are not guard targets. +:END: + +** DONE topgrade_age and the deferred row double-nag, and the TOPGRADE lever can't clear its metric +CLOSED: [2026-10-05 Mon 09:20] +Where: Decision 4 L128; Decision 5 L136; Phase 2 L158; AC L170 + +Decision 4 says a deferral renders 'rather than as stale freshness', but topgrade_age still counts days since the last stamp. With the kernel held on every everyday run, a stamp happens only after a dedicated session. + +Risk: Once the stamp is honest (see the stamp-writer finding), topgrade_age turns WARN 14 days after the last dedicated session, right next to the deferred row: two off-nominal rows for one fact. Pressing TOPGRADE runs a split run that can't clear it, and the wall reports the successful run as still warn. + +Recommended change: Add one sentence to Phase 2 (L158). Do not cap topgrade_age, because that conflicts with Decision 1. + +Suggested sentence: "topgrade_age grading is unchanged: per Decision 1 it ages from the last full-current stamp while anything is deferred. The deferred row is informational, carries the count, and hosts the apply-on-reboot / --complete action. While the deferred set is non-empty, the topgrade_age row offers that action instead of TOPGRADE, and the TOPGRADE wall note says 'N deferred — freshness clears on --complete' rather than a bare 're-probed topgrade_age: warn'." + +Add a matching acceptance criterion: "With only a kernel held, a successful TOPGRADE press leaves topgrade_age aging, and the wall names the deferred set and the --complete remedy." + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: accepted, folded into Implementation phases, Acceptance criteria. +:EVIDENCE: +- Spec L107 (Decision 1): "Freshness stays stale while a guarded upgrade is genuinely un-applied." +- Spec L77 (Alternative A): rejected because "a topgrade that was blocked from applying a real upgrade would read as 'fresh'." +- Spec L128 (Decision 4): stamps only when the deferred set is empty; the deferred set is its own row; "Freshness keeps meaning 'current'." +- Spec L136: "the panel's deferred count carries a kernel most days." +- Spec L158 (Phase 2): no change to topgrade_age grading or to which lever hangs off it. +- Spec L161 (Phase 3): the apply-on-reboot action is not bound to any metric. +- Spec L197: a cadence nag for the kernel deferral is wanted. +- maint/src/maint/probes/updates.py:114-131: grades days since topgrade_run against topgrade_warn_days. +- configs/maintenance-thresholds.toml:60: topgrade_warn_days = 14. +- maint/src/maint/remedies.py:303-312: the TOPGRADE lever has metric_ids and re_probe_ids ["topgrade_age"] and "always": True. +- maint/src/maint/remedies.py:537-546: attach_levers shows "always" levers at any severity. +- maint/src/maint/doctor.py:169-170: stamps topgrade_run on any successful rid=="topgrade" run; Phase 2 does not remove this, so the "can't clear" scenario only exists after the separate stamp-writer fix. +- hyprland/.local/bin/topgrade: the wrapper also stamps on rc 0. +- maint/src/maint/doctor.py:342-350: the re-probe note goes to the wall. +:END: + +** DONE Kernel-set rule sweeps in firmware and api-headers, and ratio has three kernels +CLOSED: [2026-10-05 Mon 09:20] +Where: Phase 1 L155 ('every linux* kernel package and its -headers'); Design L65; Decision 5 L134-135; Risks L196 + +The kernel set is 'every linux* kernel package and its -headers', 'moved as a set, never one without the other'. Ratio is described as 'btrfs root, two kernels'. + +Risk: A literal glob holds linux-api-headers and all firmware, GPU firmware included, until the dedicated session, and inflates the 'N deferred' count. A substring match would hold archlinux-keyring indefinitely and start breaking signature checks. 'Never one without the other' can't hold for strix, and --complete can't pacman-install a locally built kernel, so the implementer has to invent the policy. + +Recommended change: L155: replace "(every =linux*= kernel package and its =-headers=)" with "(the packages owning =/usr/lib/modules/*/vmlinuz=, via =pacman -Qqo=, plus =-headers= where that package is installed; =linux-firmware*=, =linux-api-headers= and other =linux*= names are not kernels)". Add: "the apply step moves only the pending members of the kernel set; a foreign kernel never appears in checkupdates and is left untouched." + +In the same paragraph, narrow the gate to match L135: "for each kernel version installed by this run" instead of "for each kernel under =/usr/lib/modules=". As written, the gate fails permanently on any kernel with no headers and no DKMS build. + +L134 and L196: change "two kernels" to "three kernels, one a locally built foreign package with no =-headers=". + +Phase 1 tests: add a ratio-shaped fixture with three vmlinuz owners, one of them foreign and headerless, linux-firmware* and linux-api-headers installed, and zfs built for two of the three kernels. Assert that firmware is never held and that the gate ignores kernels this run did not touch. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: accepted, folded into Design, Decisions, Implementation phases, Testing. +:EVIDENCE: +Spec text: +- L155: "The kernel set is derived from what is installed (every =linux*= kernel package and its =-headers=)" +- L155 (gate): "for each kernel under =/usr/lib/modules=, =dkms status= reports every registered module =installed= for it" +- L65 and L135: "every installed kernel with its =-headers=" and "moved as a set, never one without the other" +- L135 (gate): "for the new kernel version" +- L134 and L196: "Ratio (btrfs root, two kernels ..." + +Live checks on ratio (uname -n = ratio): +- pacman -Qq | grep -c '^linux' returns 24: linux, linux-api-headers, linux-firmware plus 17 linux-firmware-* split packages (amdgpu among them), linux-headers, linux-lts, linux-lts-headers, linux-lts-strix. +- pacman -Qoq /usr/lib/modules/*/vmlinuz returns linux-lts-strix, linux-lts, linux. Each /usr/lib/modules/*/pkgbase file matches its owner. +- pacman -Qm lists linux-lts-strix 6.18.25-1. pacman -Qi shows Packager "Unknown Packager". pacman -Si linux-lts-strix gives "error: package 'linux-lts-strix' was not found". +- pacman -Qq | grep -- '-headers$' returns linux-api-headers, linux-headers, linux-lts-headers. There is no strix headers package. +- dkms status shows "zfs/2.4.4, 6.18.54-1-lts, x86_64: installed" and "zfs/2.4.4, 7.2.7-arch1-1, x86_64: installed". Nothing is built for 6.18.25-1-lts-strix. +- findmnt / shows btrfs, which confirms the btrfs-root part of L134. +- checkupdates lists linux, linux-headers, linux-lts and linux-lts-headers as pending. No firmware and no linux-api-headers are pending today. +:END: + +** DONE The gate's snapshot and initramfs checks pass vacuously +CLOSED: [2026-10-05 Mon 09:20] +Where: Phase 1 L155 (kernel-modules-check); Decision 5 L135; AC L169; Risks L196 + +The gate passes if each kernel's 'initramfs is newer than its vmlinuz' (each kernel under /usr/lib/modules) and, 'on a ZFS root, a pre-pacman_ snapshot of the root dataset exists'. L196 says the 05-zfs-snapshot hook snapshots before every transaction. + +Risk: Neither check can catch the failure it exists for. One is a snapshot that wasn't taken for this kernel change (skipped, or failed on a full pool, one of the DKMS triggers the spec lists); rolling back to an older retained snapshot then undoes more than the kernel transaction. The other is a stale initramfs beside a new kernel, which lets the gate wave through an unbootable reboot on velox. The 'fails on snapshot listings' test covers a state that never occurs. + +Recommended change: Phase 1, L155. Replace the gate sentence with: + +"The script records T0 before the kernel transaction. For each kernel, with pkgbase read from /usr/lib/modules//pkgbase: dkms status reports every registered module =installed= for , and /boot/initramfs-.img is newer than /boot/vmlinuz-. For a kernel this run changed, the image must also be newer than T0. Never compare against /usr/lib/modules//vmlinuz, whose mtime is the package build date. On a ZFS root, a =pre-pacman_= snapshot of the root dataset (findmnt -no SOURCE /) has a creation time of at least T0 minus 60s (zfs-pre-snapshot's MIN_INTERVAL). An older retained snapshot does not satisfy this." + +Phase 1 tests. Replace the snapshot and timestamp fixtures with: +- snapshot listings where only pre-T0 snapshots exist (gate fails) and one is newer than T0 (gate passes) +- an initramfs newer than the module-tree vmlinuz but older than T0 (gate fails) + +L196. Change "snapshots it before every transaction" to "snapshots it before each transaction (skipped within 60s of the previous snapshot; a failed snapshot only warns and does not abort)". + +Optional: on a ZFS root, also check with lsinitcpio that the image contains zfs.ko. + +Should-fix, not blocking. Verification: confirmed. +Disposition: modified, folded into Decisions, Implementation phases, Risks, Testing. Confirmed, with the finding's corrections. +- The gate finds the root dataset itself with findmnt. zfs-pre-snapshot hardcodes $POOL/ROOT/default, so a mismatch fails closed and the failure names the dataset. +- The freshness and snapshot requirements apply only when the check list is non-empty, so a --complete that moves no kernel isn't failed for want of a snapshot. + +The optional lsinitcpio check is made required on a ZFS root. mkinitcpio writes an image without zfs.ko and only warns, and that image is exactly the velox failure. +:EVIDENCE: +Spec: +- L135 (Decision 5): "the initramfs is newer than the kernel image, and on a ZFS root a pre-pacman snapshot exists" +- L155 (Phase 1): "for each kernel under =/usr/lib/modules= ... its initramfs is newer than its =vmlinuz=; on a ZFS root, a =pre-pacman_= snapshot of the root dataset exists" +- Phase 1 tests: "fails on ... snapshot listings" +- L169: "the pre-pacman snapshot it required" +- L196: "the =05-zfs-snapshot= hook snapshots it before every transaction" + +scripts/zfs-pre-snapshot: +- :10-11 DATASET="${ZFS_PRE_DATASET:-$POOL/ROOT/default}", hardcoded, no findmnt +- :13, :19-25 return early when the lock file is under MIN_INTERVAL=60s old +- :14, :35-40 KEEP=10 prune +- :30, :41-43 on failure, =echo "Warning: Failed to create snapshot"= with exit status 0 + +archsetup: +- :2421-2433 the 05-zfs-snapshot.hook heredoc has no AbortOnFail +- :881-884 is_zfs_root uses =findmnt -n -o FSTYPE /= only + +Live system on ratio: +- =pacman -Qi linux-lts= shows Build Date Fri 25 Sep 2026 11:39:08 and Install Date Tue 29 Sep 2026 08:15:55 +- /usr/lib/modules/6.18.54-1-lts/vmlinuz has mtime 2026-09-25 11:39:08, the build date +- /boot/vmlinuz-linux-lts is 2026-09-29 08:17:18.09 and /boot/initramfs-linux-lts.img is 08:17:29 +- /usr/lib/modules/7.2.7-arch1-1/vmlinuz is 2026-09-21 13:51:14, matching linux's Build Date +- /usr/share/libalpm/hooks/90-mkinitcpio-install.hook triggers on usr/lib/firmware/*, usr/lib/systemd/systemd, usr/bin/cryptsetup, usr/lib/modprobe.d/ and other paths +- /usr/share/libalpm/scripts/mkinitcpio:148-151 runs =install -Dm644 -- "${line}" "${kernel}"= into /boot/vmlinuz-${pkgbase} before the preset build +- /usr/lib/modules/*/pkgbase gives linux-lts-strix, linux-lts and linux +:END: + +** DONE Ratio's guard hook still has the pre-migration name, and the missing-hook case is unspecified +CLOSED: [2026-10-05 Mon 09:20] +Where: Problem L34; Design L65; Decision 3 L121; Phase 1 L155; Phase 2 L158; Phase 4 L164 + +The spec names /etc/pacman.d/hooks/10-hypr-live-update-guard.hook as the installed hook the script reads, and relies on it as the backstop for a bare pacman -Syu. Phase 4's rollout installs only the new unit on existing machines. + +Risk: A script that hardcodes the 10- path finds no hook on ratio and computes an empty blocked set. The old-named guard still fires, so the live transaction aborts, and the Phase 2 equality test misreads or errors. On a machine with no hook at all, reading 'no patterns' would live-swap mesa with no backstop. Separately, on ratio today, a bare =pacman -Syu= under Hyprland with a kernel and a guarded lib both pending deletes /boot/vmlinuz and the initramfs for linux and linux-lts, then aborts, leaving those boot entries without images. + +Recommended change: Phase 1 (L155): after "blocked set (hook Target lines, version-aware)", add: "The hook is read from /etc/pacman.d/hooks/10-hypr-live-update-guard.hook. If that file is absent, or no Target lines parse from it, the live run refuses: it exits non-zero and names the path, and never computes an empty blocked set. This is deliberately unlike maint's guard.trips, which treats missing patterns as 'can't trip'." Add a matching test to the Phase 1 list: "a missing or Target-less hook fails closed and runs no pacman transaction." + +Phase 4 (L164): add: "The one-time install on existing machines also migrates the guard hook to 10-hypr-live-update-guard.hook and removes the unprefixed hypr-live-update-guard.hook (ratio still has the legacy name; velox's was hand-placed under it). The unprefixed name sorts after 60-mkinitcpio-remove, so a blocked bare pacman -Syu can delete the initramfs before the guard aborts. Confirm with ls /etc/pacman.d/hooks/ on both machines." + +Optional: add a post-rebuild-check assertion that the 10- name exists and the unprefixed name does not. + +Should-fix, not blocking. Verification: confirmed. +Disposition: modified, folded into Design, Implementation phases, Testing. Confirmed. Under the fail-closed rule the script refuses on any machine that still carries the legacy hook name. Ratio does today (only hypr-live-update-guard.hook is in /etc/pacman.d/hooks), and velox's was hand-placed under it. So the migration can't wait for Phase 4; it moves to checklist step (a), before that machine's levers route through the script. The optional post-rebuild-check assertion is adopted. +:EVIDENCE: +- Live ratio: ls /etc/pacman.d/hooks/ shows only 99-grub-sync-efi.hook and hypr-live-update-guard.hook (543 bytes, 4 Jul, same content as the installer heredoc). No 10-hypr-live-update-guard.hook. +- archsetup:2620-2625 migrates the name ("rm -f /etc/pacman.d/hooks/hypr-live-update-guard.hook", then writes 10-hypr-live-update-guard.hook). The comment at 2622-2624 says the guard must run before 60-mkinitcpio-remove. Commit be2277d (2026-07-18, "fix: order pacman safety hooks") added this. +- archsetup:41-75: the arguments are only --config-file, --fresh, --status and similar flags. There is no way to run a single step, so updating an existing machine means placing files by hand. +- /usr/share/libalpm/hooks/60-mkinitcpio-remove.hook is a PreTransaction hook on Path Remove usr/lib/modules/*/vmlinuz. man alpm-hooks says hooks run "in alphabetical order of their file name", so "hypr-live-update-guard" sorts after "60-mkinitcpio-remove". AbortOnFail aborts the transaction. pacman -Q shows linux 7.2.7 and linux-lts 6.18.54, both with headers. +- tests/installer-steps/test_pacman_hook_order.py:54-76 checks the ordering and the cleanup, but only against the installer source, not a live machine. scripts/testing/tests/test_desktop.py:59 asserts the 10- path, but only in the VM harness. scripts/post-rebuild-check has no hook assertion (grep returned nothing). +- ~/.dotfiles/maint/src/maint/guard.py:26-30 (trips) is the fail-open precedent: "Tolerant of a missing key — no patterns means the guard can't trip". +- todo.org:3391-3396: the hook was placed on velox by hand as /etc/pacman.d/hooks/hypr-live-update-guard.hook on 2026-06-28. +- Spec: L34 names the 10- path. L65 says the blocked set is "read from the installed hook's Target lines" and that the hook stays "the backstop for a bare pacman -Syu". L155 says "hook Target lines, version-aware" and lists no missing-hook test. L164 says "existing machines need the one-time install", with no hook migration. Grepping the spec for absent or missing finds them only for topgrade_run and the arm flag, never for the hook. +:END: + +** DONE maint's guard_patterns is not a mirror of the hook, and the pin test is in the wrong repo +CLOSED: [2026-10-05 Mon 09:20] +Where: Decision 3 L120-121; Phase 2 L158; Reuse L183 + +The spec calls maint's =[updates] guard_patterns= its own copy of the hook's trigger list. It adds a Phase 2 test under the maint harness asserting that the TOML equals the installed hook's Target list. + +Risk: The equality test fails on its first run, and the spec doesn't say which side gets reconciled. A dotfiles unit test that reads /etc and ~/.config depends on the host and on how stale its installed copies are. Rewriting the TOML drops hyprlang and hyprcursor, which the --noconfirm closure finding shows belong to the soname family, and the badge under-predicts the deferral. + +Recommended change: Edits by location: + +- L120: replace "its own copy of the trigger list" with "its own, divergent pattern list". +- L158: replace the test sentence with: "Rewrite archsetup configs/maintenance-thresholds.toml [updates] guard_patterns to exactly the hook heredoc's 15 Target patterns (the hook is the owner and stays unchanged). This drops hyprland-*, hyprlang, hyprcursor, *wayland*, wlroots*, lib32-mesa*, lib32-vulkan-radeon and lib32-vulkan-intel, and adds wayland, libdrm, libglvnd, vulkan-mesa-layers, nvidia-utils, lib32-nvidia-utils and xorg-xwayland. Update the TOML comment that still describes press-again arming. The pin test lives in archsetup tests/installer-steps/ and compares the sorted seed-TOML guard_patterns with the Target lines extracted from the hook heredoc in the archsetup script (the test_pacman_hook_order.py idiom). It reads neither /etc nor ~/.config." +- L183: align with the L158 change. +- L188 and Phase 4: note that existing machines need install_maintenance_thresholds re-run to pick up the new TOML. +- L187 and L200: fix the phase numbering so each phase's tests are in one consistent place. + +Should-fix, not blocking. Verification: confirmed. +Disposition: modified, folded into Decisions, Implementation phases, Readiness dimensions, Testing. Confirmed. Both files the change touches, configs/maintenance-thresholds.toml and the hook heredoc, belong to archsetup. The rewrite and the pin test therefore land in the Phase 1 archsetup commit, not in dotfiles Phase 2. The rewritten TOML has to reach existing machines, so the re-copy at checklist step (a) is mandatory, overriding F30's 'only if Phase 2 adds keys'. +:EVIDENCE: +- Spec L120: "maint already carries its own copy of the trigger list as [updates] guard_patterns". +- Spec L121: "the TOML patterns stay as the panel's display-side mirror ... and gain a test asserting they match the hook". +- Spec L158: "A test asserts the TOML guard_patterns equal the installed hook's Target list. Tests under the maint fake harness." +- Spec L183 repeats the claim. L187 ("installer-step pytest for Phase 2") and L200 ("Phase-2 install under tests/installer-steps/") contradict L158. +- configs/maintenance-thresholds.toml:65-72: mesa, mesa-*, lib32-mesa*, hyprland, hyprland-*, aquamarine, hyprutils, hyprlang, hyprcursor, hyprgraphics, *wayland*, wlroots*, vulkan-radeon, lib32-vulkan-radeon, vulkan-intel, lib32-vulkan-intel. The file is byte-identical to ~/.config/archsetup/maintenance-thresholds.toml (diff clean). +- archsetup:2629-2643, the hook heredoc's 15 Target lines: mesa, mesa-*, wayland, libdrm, libglvnd, hyprland, aquamarine, hyprutils, hyprgraphics, vulkan-radeon, vulkan-intel, vulkan-mesa-layers, nvidia-utils, lib32-nvidia-utils, xorg-xwayland. +- The installed hook's Target list matches the heredoc, but it sits at /etc/pacman.d/hooks/hypr-live-update-guard.hook. /etc/pacman.d/hooks/10-hypr-live-update-guard.hook does not exist on ratio (cat: No such file or directory). +- archsetup:1640-1650: install_maintenance_thresholds copies the seed TOML to ~/.config/archsetup. +- dotfiles maint/src/maint/thresholds.py:3-4: the TOML is "owned and installed by archsetup". +- thresholds.py:18-19: env overrides "so tests and fixtures never touch real files". +- thresholds.py:31-33: shipped_path. +- maint/src/maint/guard.py:29: maint reads guard_patterns via fnmatchcase. +- tests/maint/test_remedies_doctor.py:31-60 uses an inline SHIPPED_TOML. No dotfiles maint test references the archsetup checkout. Every test sets MAINT_THRESHOLDS to a temp file. +- tests/installer-steps/test_pacman_hook_order.py already extracts hook data from the archsetup script source. +:END: + +** DONE maint's guard refusal and press-again arm UX are left in place +CLOSED: [2026-10-05 Mon 09:20] +Where: Phase 2 L158; Design 'For the user' L67; AC L167; Decision 3 L121 + +Phase 2 removes only 'the press-again-to-force sentinel wrap'. Design and the AC say UPDATE/TOPGRADE run the script and succeed with a guarded library pending. The 'will defer' badge is unspecified. + +Risk: If only the wrap is removed, the panel still refuses on exactly the guarded days the script exists for, and AC L167 fails. The second press still says the run swaps live, which is no longer true. The badge under-predicts the deferral because it omits the kernel set. + +Recommended change: Replace the Phase 2 sentence "the press-again-to-force sentinel wrap goes away (the driven path never trips the guard)" with: "Drop guard: live_update from the update and topgrade remedies. iter_fix no longer refuses them. The --force sentinel wrap, the GUI's _update_force override, and the CLI --force meaning for these two go with it. Rewrite the tests that pin the tag, the refusal, and the wrap. guard.trips stays only as the arm-line annotation that Decision 3 calls the display-side mirror: on a tripped read, arm_line, _rearm_after_guard's text, and the doctor review suffix read 'UPDATE armed — will defer (kernels are always held) — press again to run $ upgrade-guarded', and nothing says live apply or REBOOT required." Don't have maint compute the pending kernel set itself. The post-run "N deferred" row, built from the script's state file, gives the exact count. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Implementation phases, Testing. Partially confirmed; this follows the narrowing that maint does no kernel computation of its own. The arm line also names the DKMS set (F04), and it says 'at least', because the closure can hold more than the pattern match predicts. +:EVIDENCE: +Spec L158 (Phase 2): "UPDATE and TOPGRADE levers change their argv to the script; the press-again-to-force sentinel wrap goes away (the driven path never trips the guard)." Spec L121 (Decision 3): "the TOML patterns stay as the panel's display-side mirror (the badge that says a run will defer)". Spec L167 (AC): "UPDATE applies everything else, exits 0 ... the guard hook does not fire." + +~/.dotfiles/maint/src/maint/remedies.py:298 and :309: both update and topgrade carry "guard": "live_update". + +~/.dotfiles/maint/src/maint/doctor.py iter_fix: =if not dry_run and r.get("guard") == "live_update" and not force:= runs guard.trips(pending, th) and yields a "guard" event ("press again to apply live (REBOOT required after), or apply from a TTY"), then returns before any step. The wrap is a separate later line: =wrapped = bool(force and r.get("guard") == "live_update" and not dry_run)= -> _all_steps adds guard_sentinel_set/clear. The review suffix (doctor.py ~L415-421) reads "[live-update guard: ... — arms in the GUI; --force on the CLI]". + +viewmodel.py:99-107 arm_line(guard=...) gives "press again to apply live (REBOOT required after), or apply from a TTY". + +gui.py:1396-1405 _guard_arm_line sets self._update_force = bool(info["tripped"]). gui.py:1430-1437 fires with force=self._update_force for guarded rids, so the GUI bypasses the refusal on a warm tripped read. gui.py:1518-1535 _rearm_after_guard repeats the live-apply/REBOOT wording. + +panel.py:224-237: guarded()/update_guard() key on the same tag. + +cli.py:265-267: --force help "run a guarded remedy despite the live-update guard (the CLI's press-again)". + +Tests pinning the current behavior: tests/maint/test_remedies_doctor.py:437-438 (assert guard == "live_update" on both), :843-1003 (refusal, force, sentinel order); tests/maint/test_panel_levers.py:404 (asserts "press again to apply live (REBOOT required after)"). + +configs/maintenance-thresholds.toml:65-72: guard_patterns has no linux*/kernel entries, so a guard-only badge omits the held kernel set the script always defers. +:END: + +** DONE Deferred-set vocabulary, storage and arm-flag contents are undefined across repos +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L65, L67; Scope tiers L57; Decision 4 L128; Decision 6 L142; Phase 1 L155; Phase 2 L158; Phase 3 L161; Readiness Data model L178 + +'blocked set' includes the kernel set in Design (=--ignore==) but excludes it in Phase 1 ('blocked set ∪ held-kernel set'). The spec also uses 'held set', 'held-kernel set', 'guarded set' and 'GPU/compositor set'. Storage is called 'a state file' (L65, L155, L158) in some places and 'its own cache key' (L128) in another. Phase 1 (archsetup) writes the record and Phase 2 (dotfiles) reads it, but neither names the path, schema or writer. Data model (L178) lists only the arm flag and topgrade_run. The arm flag's contents (which packages the boot run applies) are unspecified, and no phase says that --complete or the boot run rewrites the record. The badge reads 'apply on reboot' even though the kernel portion lands live. + +Risk: Phases 1 and 2 land in different repos and sessions, so the implementers will invent incompatible paths and formats. A mismatch shows no deferred row and fails silently, because the 'state file round-trips' test only checks the script against itself. A sudo-invoked run writes to /root/.local/state/maint, which the panel never reads. The boot run has no defined source for which set to apply, and a record that is never cleared leaves a permanent 'N deferred' badge. + +Recommended change: 1. Definitions. Add one sentence to Design after the L65 set computation: "GPU/compositor set = pending hook Targets whose on-disk version changes, computed only when a compositor is live; kernel set = every installed linux* kernel with its -headers, held on every everyday run whether or not a compositor is live; deferred set = their union, which is exactly the --ignore list." Then: +- In L65, change "--ignore=" to the deferred set. +- In L155, rewrite "blocked set (...) ∪ held-kernel set when a compositor is live" as "GPU/compositor set (hook Target lines, version-aware, compositor live only) ∪ kernel set (always)". +- Replace "held set" (L57), "held-kernel set" (L55, L161) and "guarded set" (L65, L143) with the defined terms. + +2. Name the record. In Decision 4 (L128), and as a new line in Data model (L178), add: "maint cache key upgrade_deferred: ~/.local/state/maint/upgrade_deferred.json (MAINT_STATE_DIR honoured), in cache.put's {written_at, data} envelope, data = {packages: [{name, old, new, kind: gpu|kernel}]}. upgrade-guarded is its only writer and runs as the user. It rewrites the record at the end of every mode (everyday, --complete, boot), with an empty list when nothing is deferred. maint's new probe is its only reader." + +3. Boot source. In Decision 6 (L142) and Phase 3 (L161), add: "the boot form applies the record's kind=gpu entries and never the kernel ones; the arm flag stays a presence-only file." + +4. Tests. Make Phase 1's "state file round-trips" test read the record in the envelope maint's cache.get expects, and point Phase 2's probe test at the same fixture JSON. + +Do not rename the badge, add SUDO_USER handling, or add gate/armed/boot fields to the schema. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Design, Decisions, Implementation phases, Readiness dimensions, Testing. Partially confirmed. This adopts the vocabulary and the named record, with two reconciliations. +- My boot-source call puts the exact name=version list in the root-owned arm flag. The flag isn't presence-only: it is the boot form's source, and the record is the panel's. +- Accepted findings F10, F12, F29 and F39 each need a durable field (interrupted, gate, AUR failure, news), so the schema carries result, failed_step, detail, gate and news. + +Arm state is still not stored, because the probe reads the flag. +:EVIDENCE: +Spec: +- L65: "computes the *blocked set* = the guard's own trigger list ... plus the *kernel set* ...; runs pacman -Syu --noconfirm --ignore= ... It writes the deferred set to a state file the panel reads". +- L155: "blocked set (hook Target lines, version-aware) ∪ held-kernel set when a compositor is live → ... deferred set written to a state file". Its tests: "the --ignore list is exactly blocked ∪ kernel set; the state file round-trips". +- L128: "writes the deferred set to its own cache key"; L129: "one more cache key and one more probe in maint". +- L158: "A new probe reads the deferred-set state file". +- L178: Data model lists only the arm flag and the topgrade_run key. +- L69: the arm flag is "a file on a non-tmpfs path". +- L142: the boot run is "scoped to the deferred GPU/compositor set ... no kernel". +- L155: --complete is "apply the kernel set live, run the gate, then arm ... or, with no compositor live, apply it directly". +- L171, L181: the panel clears or still shows pending work after boot. +- L135, L168: the kernel is held on every everyday run, including a TTY run. +- L57: "held set". L143: "guarded set". + +Code: +- ~/.dotfiles/maint/src/maint/cache.py:15-42: _dir() is MAINT_STATE_DIR or ~/.local/state/maint; put() writes {"written_at", "data"} to .json atomically; get() returns (data, age). +- cli.py:294: p_stamp choices=["topgrade"]. cli.py:228-231: cmd_stamp does cache.put(f"{what}_run", {"at": ...}). +- remedies.py:294-311: the update and topgrade remedies are kind "user", not root. + +Adjacent defect, not part of this finding (flag separately): which -a topgrade resolves ~/.local/bin/topgrade first. That wrapper stamps topgrade_run on any rc 0 (wrapper L40-42), and doctor.py:169-170 also stamps on any exit-0 TOPGRADE lever. upgrade-guarded's own topgrade --disable system,... call, and a lever whose argv is the script (which exits 0 while deferring, per L65), would therefore stamp freshness on a deferring run, contradicting Decision 4 (L128). + +Also noticed: on ratio the installed hook is /etc/pacman.d/hooks/hypr-live-update-guard.hook, not 10-hypr-...; the installer renames it at archsetup:2621-2625. That affects "read the installed hook", not this finding. +:END: + +** DONE Boot run is invisible on ratio's screen, and its outcome is never surfaced +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L67 ('runs in the console'); Decision 7 L149; AC L171; Readiness Errors L179, Observability L181, Performance L182, Config surface L185 + +The spec assumes the boot run's output appears on the console and that failures are 'named', but it sets no StandardOutput or TTYPath. The flag is removed whatever the outcome, and success or failure lives only in the unit status and the journal. Nothing records the outcome where the panel reads it. AC L171 only requires that 'the panel still shows the pending work', and after a failed run the panel shows the same 'N deferred' as before arming. + +Risk: On ratio, a multi-minute armed boot shows a blank tty1, which invites a hard power-cycle mid-transaction (see the timeout finding). Any failure lives only in the journal. The panel gives no sign that the one-shot arm was spent, so the user either re-arms blind or never re-arms. + +Recommended change: This replaces the finding's larger proposal: no new outcome field and no "last boot apply failed" panel line. + +1. Phase 3 (L161) and Observability (L181). Name the unit's I/O: StandardInput=null, StandardOutput=tty, TTYPath=/dev/tty1. Not journal+console, because on ratio /dev/console is ttyS0. The script mirrors its output to the journal (for example through systemd-cat). Its first line is a banner on tty1: "applying N deferred GPU/compositor upgrades — do not power off". + +2. L69 and L161. Move the unconditional flag removal out of ExecStart into ExecStopPost=/usr/bin/rm -f . ExecStopPost runs after success, failure, and a TimeoutStartSec kill, and ExecStart's exit then carries the upgrade result. A failed or timed-out run leaves archsetup-boot-upgrade.service failed, and maint's existing failed_units row names it with its journalctl hint in the session that follows. + +3. AC L171. Append: "and a failed or timed-out run appears in maint's failed-units row as archsetup-boot-upgrade.service". + +4. The Phase 3 manual boot test. Add a check that the banner and progress are visible on ratio's monitor during the armed boot. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Design, Implementation phases, Acceptance criteria, Readiness dimensions, Testing. Partially confirmed; this follows the narrowing: no new outcome field, and maint's failed_units row is the surface. The proposed ExecStopPost disarm gives way to my ExecStartPre call. That keeps the property the finding wanted: ExecStart no longer ends in an rm, so its exit status is the unit's result, and failures and timeouts leave the unit failed. +:EVIDENCE: +Visibility half, all confirmed on ratio: +- uname -n is ratio. +- systemctl show reports DefaultStandardOutput=journal. +- /proc/consoles shows "ttyS0 -W- (EC p a)" (C = preferred, the /dev/console target) and "tty0 -WU (E p )". +- /etc/default/grub:6 has GRUB_CMDLINE_LINUX="console=tty0 console=ttyS0,115200", with quiet splash loglevel=2 at :5. /proc/cmdline matches. +- pacman -Q plymouth: package not found. +- /etc/systemd/system/getty@tty1.service.d/ does not exist (no autologin on ratio). systemctl cat getty@tty1 shows TTYReset=yes and TTYVTDisallocate=yes (getty@.service file lines 42 and 44). +- ~/.dotfiles/hyprland/.profile.d/99-hyprland-autostart.sh:15 runs clear before start-hyprland. +- The ttyS0 console is ratio-local. grep for ttyS0 finds nothing in the archsetup repo, so the hazard is ratio-specific and velox is unverified (offline). + +Surfacing half, refuted in part: +- ~/.dotfiles/maint/src/maint/probes/systemd.py:70-101 failed_units emits one evidence entry per failed unit: unit, since, exit, and hint "journalctl -u {name} -b". +- gui.py:1068-1074 renders it as a "failed units" section with "RESTART / RESET per unit". +- status.py:50 includes it in maint status. + +Spec text: +- L69: ExecStart "... then maint stamp topgrade on success, then removes the flag unconditionally". +- L161: "ExecStart runs the script's --complete form as the user and removes the flag unconditionally". +- L149: "removes unconditionally at the end of its attempt". +- L171 AC: timeout leaves "the flag is cleared, and the panel still shows the pending work". +- L181: "the boot run's output is on the console; its systemd unit status and journal record success/failure". +:END: + +** DONE maint is not on the boot unit's PATH, so the boot stamp can silently not happen +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L69; Phase 1 L155; Phase 3 L161; AC L170; Readiness L189 + +The boot oneshot runs the script as the user, and the script stamps with =maint stamp topgrade=, with no path or environment specified. 'maint stamp' is listed as verified present. + +Risk: The boot run applies the deferred set but writes no stamp. AC L170 (fresh topgrade_age after the armed reboot) then fails with no error anywhere. + +Recommended change: In Phase 3 (L161), after "=ExecStart= runs the script's =--complete= form as the user", add one sentence: "The unit sets =User== to the installing user (templated by the installer, as the BRIO udev rule is), so =HOME= points at that user's maint state. The script calls maint by absolute path (=$HOME/.local/bin/maint=), not through PATH, because a system unit's PATH is =/usr/local/bin:/usr/bin= and doesn't include =~/.local/bin=. If maint is missing or the stamp fails, the script logs a named failure to the journal instead of skipping silently (no =command -v ... || true=)." In L189, qualify the claim: =maint stamp= is present on the user's login PATH only (=~/.local/bin=, from dotfiles), not on the system PATH. Leave topgrade resolution out of this finding. The boot run doesn't invoke topgrade (Decision L142). + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Design, Implementation phases, Readiness dimensions, Testing. Partially confirmed; this follows the narrowing, so the topgrade-resolution clause is dropped (F01 covers the live path, and the boot form never runs topgrade). It keeps the recommended absolute-path call. maint stamp topgrade exists today (maint cli.py), and the boot unit already depends on the user's HOME for the record. + +'Logs a named failure' gets an exit code and a failed_step, so a failed stamp also reaches failed_units at boot. +:EVIDENCE: +- Spec L69: the boot unit's "ExecStart runs, as the user: ... then =maint stamp topgrade= on success". L155: --complete stamps when the deferred set is empty. L161: "=ExecStart= runs the script's =--complete= form as the user". L189: "=maint stamp= — all verified present on velox". The spec has no PATH or Environment wording anywhere (grep -n "PATH\|Environment\|absolute" matches only L32's unrelated "PATH wrapper"). +- =which -a maint= finds only ~/.local/bin/maint, a symlink to ../../.dotfiles/hyprland/.local/bin/maint (a python shim that loads maint/src from the dotfiles repo). +- =systemd-path search-binaries-default= gives /usr/local/bin:/usr/bin. +- ~/.dotfiles/common/.zprofile:4-12 says that without =export PATH="$HOME/.local/bin:$PATH"=, "maint, the topgrade wrapper, and every user script are unreachable". +- ~/.dotfiles/hyprland/.local/bin/topgrade:40-42 has =command -v maint >/dev/null 2>&1; then maint stamp topgrade >/dev/null 2>&1 || true=, which skips the stamp silently. +- The installer already templates the username into system files: archsetup:2595-2600 (BRIO rule: ARCHSETUP_USERNAME, then =sed -i "s/ARCHSETUP_USERNAME/${username}/"=). +- maint's state path is ~ based (dotfiles maint/src/maint/cache.py:16-17, =os.path.expanduser("~/.local/state/maint")= unless MAINT_STATE_DIR is set), so the unit also needs the user's HOME. =User== provides it. +- Counterpoint I checked: Decision L142 says the boot run is "one pacman transaction, no ecosystem sweep", so topgrade is not invoked at boot and wrapper resolution doesn't apply to this path. +:END: + +** DONE Second-writer =maint apply-upgrade= path survives from the first draft +CLOSED: [2026-10-05 Mon 09:20] +Where: Design final paragraph L71; Alternative D Neutral L93 + +L71: 'The cleanest single home for the stamp is a small maint apply-upgrade (or a flag in the existing lever) that performs the guarded system upgrade and stamps on success, which both the boot unit and an interactive TTY run call.' Alternative D: 'the same maint apply-upgrade path serves both'. + +Risk: It invites an implementer to build a third completion path and stamp writer in maint, which is the drift the paragraph itself warns against. Stamp ownership reads as contradictory. + +Recommended change: Three edits: +- Replace L71 with: "=upgrade-guarded= is the one code path that completes an upgrade and records it. The everyday run stamps when nothing was deferred. =--complete= stamps when it lands the deferred set, whether the boot unit or a TTY invokes it. A by-hand sentinel-override upgrade outside the script does not stamp; =upgrade-guarded --complete= from a TTY replaces that ritual." +- In Alternative D's Neutral (L93), change "the same =maint apply-upgrade= path" to "the same =upgrade-guarded --complete= path". +- In L69, replace the ExecStart sentence ("=informant read= ... then =topgrade --only system= ... then =maint stamp topgrade= on success, then removes the flag") with "Its =ExecStart= runs =upgrade-guarded --complete= as the user (which clears news, applies the deferred GPU/compositor set, and stamps), then removes the flag unconditionally." This makes the unit match Decision 6 and Phase 3, and leaves the script as the only stamp writer. + +Should-fix, not blocking. Verification: confirmed. +Disposition: modified, folded into Design, Alternatives. Confirmed. Its proposed boot ExecStart ('upgrade-guarded --complete ... then removes the flag unconditionally') conflicts with F08's distinct boot form and with my disarm-at-start call, so that sentence follows F08 and F10. The Design final paragraph and Alternative D rewrites stand, extended to name both completion forms. +:EVIDENCE: +- Spec L71 (current): "The cleanest single home for the stamp is a small =maint apply-upgrade= (or a flag in the existing lever) that performs the guarded system upgrade and stamps on success, which both the boot unit and an interactive TTY run call." +- Spec L93: "the same =maint apply-upgrade= path serves both an interactive TTY run and the boot unit." +- The contradicting sections: + - L121 (Decision 3): ship the split script in archsetup; "maint's UPDATE/TOPGRADE levers change their =argv= to the script". + - L142 (Decision 6): "the boot oneshot runs the script's =--complete= form". + - L155 (Phase 1): =scripts/upgrade-guarded=, with "=--complete= (... stamp when the deferred set is empty)". + - L161 (Phase 3): "=ExecStart= runs the script's =--complete= form as the user". + - L173 (AC): "via the Phase-1 path". +- Draft provenance: =git show 77447d0:docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org= has the L71 text verbatim at its line 66 and the L93 text at its line 88. That draft's Phase 1 (its line 119) reads "A single =maint= entry point (=maint apply-upgrade=, ...)". Commit e777ddc rewrote the phases around the script but left these lines in place. +- No implementation exists: =grep -rn "apply.upgrade" ~/.dotfiles/maint/src= finds nothing, and the only archsetup hits are in this spec. +- Adjacent stale text: L69's ExecStart still lists "=topgrade --only system= ... then =maint stamp topgrade= on success". That conflicts with L142 and L161. +:END: + +** DONE First-draft boot and dependency text remains in Risks, Readiness and Phase 4 +CLOSED: [2026-10-05 Mon 09:20] +Where: Risks L193; Readiness Errors L179, Performance L182, External deps L189; Phase 4 L164 + +Risks: 'The --only system scope ... shrink this', and it lists 'AUR review' as a boot prompt. Readiness: 'one pacman/yay transaction at boot', and deps 'topgrade --only system, informant read, yay, maint stamp — all verified present on velox'. Errors: 'every failure path lands in boot normally ... re-arm'. Phase 4 flow: 'arm → reboot → console upgrade → session'. + +Risk: The implementer's dependency and failure checklists point at a removed command and miss the gate-failure state, so that state gets no observability or documentation. + +Recommended change: - L193 (Risks): replace "The =--only system= scope and =--noconfirm=-style flags shrink this to near zero" with "The boot form's single pacman transaction with =--noconfirm= shrinks this to near zero". Drop "AUR review" from the prompt list. Change "verified in Phase 2" to "verified in Phase 3". +- L182: change "one pacman/yay transaction at boot" to "one pacman transaction at boot". +- L189: replace the deps line with: "=checkupdates= (pacman-contrib, installer-owned), =yay=, =topgrade --disable system git_repos containers= (space-separated; topgrade 17.12.2 rejects the comma form), =informant read --all= where installed (velox; absent on ratio), =dkms status=, =zfs list -t snapshot= on a ZFS root, =maint stamp=." +- Make the same two invocation fixes in L65 and L155: comma-separated =--disable= becomes space-separated, and =informant read= becomes =informant read --all=. Apply the =--all= fix in L69 and L172 too. +- L179: append "A failed =--complete= gate leaves the new kernel installed, nothing armed and no reboot; the foreground run names the failure and the desktop stays up." +- L164: change the flow to "kernel set live → gate → arm → reboot → console upgrade → session". +- Optional, same leftover text: + - L187 and L200: Phase 1 tests are pytest beside the guard's tests, Phase 2 uses the maint fake harness, and Phase 3 uses tests/installer-steps/ plus maint tests. + - L69 and L71: replace the topgrade --only system / yay -Syu / maint apply-upgrade wording with "runs =upgrade-guarded --complete=". + +Should-fix, not blocking. Verification: confirmed. +Disposition: modified, folded into Design, Implementation phases, Readiness dimensions, Risks. Confirmed. The replacement deps line is completed with the commands other accepted findings add: +- sudo and timeout for informant (F11); +- yay -Pwq (F39); +- pacman -Sp, -Sw and versioned -S (F08, F09); +- vercmp; +- systemctl is-enabled (F31); +- lsinitcpio and findmnt (F19); +- foot (F40). + +The Risks prompt-surface text is merged with F28's so the two edits don't collide. +:EVIDENCE: +Spec text (docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org): +- L142: "one pacman transaction, no ecosystem sweep, no kernel" +- L161: "ExecStart runs the script's --complete form" +- L193: "provider choice, replace, AUR review ... The --only system scope ... verified in Phase 2" +- L182: "one pacman/yay transaction at boot" +- L189: "topgrade --only system, informant read, yay, maint stamp — all verified present on velox" +- L179: "every failure path lands in boot normally" +- L164: "arm → reboot → console upgrade → session" +- L67, L135, L155, L169: gate failure stops, names the failure, does not reboot +- L187: "installer-step pytest for Phase 2; maint unit tests for Phases 1 and 3". This inverts L155 (Phase 1 tests are pytest beside the guard's tests, tests/hypr-live-update-guard/) and L161 (Phase 3 tests go in tests/installer-steps/). +- L69: "topgrade --only system (or the equivalent yay -Syu)" + +Commands run on ratio: +- /usr/bin/topgrade --disable system,git_repos,containers --help → "error: invalid value 'system,git_repos,containers' for '--disable ...'". The space-separated form parses. topgrade --version reports 17.12.2. --help short-circuits any run, and the stamping wrapper was bypassed. +- command -v: checkupdates=/usr/bin/checkupdates (owned by pacman-contrib 1.13.1-1, which the installer installs at archsetup:2175); dkms and zfs present; informant MISSING on ratio. + +Installer code: +- archsetup:1124-1129: "--all marks without printing or prompting; a bare =informant read= is interactive and would hang an unattended run", then it runs informant read --all. +:END: + +** DONE Test framework and phase-to-harness map contradict the phases +CLOSED: [2026-10-05 Mon 09:20] +Where: Phase 1 L155 ('pytest'); Dev tooling L187; Risks L193; Testing L200 + +L155 says 'pytest beside the guard's'. L200: 'Phase-1 and Phase-3 logic under the maint fake harness; Phase-2 install under tests/installer-steps/'. L187: 'installer-step pytest for Phase 2; maint unit tests for Phases 1 and 3'. L193 says boot non-interactivity is 'verified in Phase 2'. + +Risk: Following L200 puts the Phase 1 tests where they can't import the script. Following 'pytest' produces a red gate. The prompt-surface check is assigned to the wrong phase. + +Recommended change: Use one phase-to-test map and repeat it at L155, L187 and L200: +- Phase 1 → archsetup tests/upgrade-guarded/ and tests/kernel-modules-check/. These use unittest + subprocess + env-var seams, with fake checkupdates, pacman, dkms and zfs on PATH (tests/zfs-pre-snapshot/fake-zfs can be reused). +- Phase 2 → dotfiles tests/maint/ under the maint fake harness. +- Phase 3 → archsetup tests/installer-steps/ for the install step, the arm action's tests in dotfiles tests/maint/, and the manual boot test in todo.org. + +Specific edits: +- L155: change "pytest beside the guard's" to "unittest beside the guard's (tests/upgrade-guarded/, run by make test-unit)". +- L187: change to "unittest for Phase 1 and the Phase 3 installer step in archsetup; maint unit tests for Phase 2 and the Phase 3 arm action; a manual boot test in todo.org". +- L200: replace the first sentence to match this map. +- L193: change "verified in Phase 2" to "the --noconfirm argv asserted in Phase 1's tests; no-stdin behavior verified by Phase 3's manual boot test". + +Should-fix, not blocking. Verification: confirmed. +Disposition: modified, folded into Implementation phases, Readiness dimensions, Testing, Risks. Confirmed. The map follows the re-phasing. F16's same-commit rule moves the panel action into Phase 2, so Phase 3 is archsetup only, and F21's TOML pin test lands in Phase 1. The tests use unittest, which is what make test-unit runs: it globs tests/*/test_*.py, so new suites need no list edit. +:EVIDENCE: +Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org: +- L155: "Tests (pytest beside the guard's)". +- L158: Phase 2 is maint wiring, "Tests under the maint fake harness". +- L161: Phase 3: "Tests in =tests/installer-steps/= for the install step; the arm action's tests in maint". +- L187: "installer-step pytest for Phase 2; maint unit tests for Phases 1 and 3". +- L193: "verified in Phase 2". +- L200: "Phase-1 and Phase-3 logic under the maint fake harness; Phase-2 install under =tests/installer-steps/=". + +The stale lines come from the first draft. In git show 77447d0, the phases are: Phase 1 "Stamp the safe-completion path", Phase 2 "The armed boot-time unit (archsetup)", Phase 3 "The arming affordance (dotfiles maint)". Its lines 148, 154 and 159 match the current L187, L193 and L200 word for word. + +Test conventions: +- tests/hypr-live-update-guard/test_hypr_live_update_guard.py:1-21 is unittest with HYPR_GUARD_* env-var seams, and its docstring says "python3 -m unittest tests.hypr-live-update-guard...". +- Makefile:58-68 (test-unit) loops over tests/*/test_*.py running python3 -m unittest "$mod" || fail=1. +- grep for "import pytest" or bare "def test_" under archsetup tests/ finds nothing. +- I ran a scratch pytest-style module through python3 -m unittest: "Ran 0 tests ... NO TESTS RAN", exit=5. pytest 9.1.1 is installed but no gate calls it. +- ~/.dotfiles/tests/maint/test_remedies_doctor.py:18-27 sets REPO_ROOT to the dotfiles repo and puts only maint/src and panelkit/src on sys.path, with FAKE_TOOL = tests/maint/fake-tool. +- ~/.dotfiles/Makefile:290-303 also globs tests/*/test_*.py through unittest. +- tests/zfs-pre-snapshot/fake-zfs exists and could be reused. +:END: + +** DONE yay -Sua doesn't hold AUR packages back, and its failure and privilege model is undefined +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L65 ('AUR packages pinning a guarded version hold themselves back'); Phase 1 L155 (=yay -Sua --noconfirm=; 'exit 0 on a successful live part') + +The script runs =yay -Sua --noconfirm= after the pacman run with --ignore. It assumes AUR packages that need a newer guarded library defer themselves. + +Risk: When it does fire, the spec doesn't say whether a failed AUR step counts as a failed live part, or what that does to the exit code, the stamp and the deferred badge. The implementer also has to invent the split between user and root steps. + +Recommended change: 1. L65: replace "(=yay -Sua --noconfirm=, AUR packages pinning a guarded version hold themselves back)" with "(=yay -Sua --noconfirm=; yay still resolves repo dependencies normally, so an AUR upgrade that needs a newer held package makes yay's own pacman call trip the guard and the AUR step fails)". Change "on the driven path it never fires" to "on the driven path it fires only in that AUR case". + +2. Phase 1 (L155): add one rule. Suggested default, consistent with the state framing: "A non-zero yay exit is a failed live part. The script still runs the topgrade sweep, records the AUR failure in the state file so the panel shows it beside the deferred count, does not stamp, and exits non-zero." Add a test for it. + +3. Phase 2 (L158): reword the parenthetical so the force-wrap removal no longer rests on "never trips the guard". + +4. L167: change "the guard hook does not fire" to "the guard hook does not fire from the pacman step". Add a criterion: "an AUR upgrade needing a held library fails the AUR step visibly without blocking the rest of the run". + +No change is needed for the privilege model; L69, L161 and L180 already cover it. If I want it explicit, add one clause to Phase 1: "runs as the user; pacman, informant and the arm-flag write go through sudo". + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Design, Implementation phases, Acceptance criteria, Risks, Testing. Partially confirmed; this follows the narrowing that privilege and the stamp are settled. The privilege clause is made explicit, including the EUID-0 refusal. + +Two Risks lines are added for the paths the pacman hold can't cover. yay's resolver can move a kernel-set or DKMS-set package when an AUR upgrade requires a newer one. And a DKMS package installed from the AUR would be upgraded by yay -Sua with no hold at all. Neither applies today on ratio, where zfs-dkms comes from the archzfs repo; checklist step (a) confirms velox. +:EVIDENCE: +- Spec L65: "runs the AUR-only remainder (=yay -Sua --noconfirm=, AUR packages pinning a guarded version hold themselves back)" and "on the driven path it never fires". +- Spec L155: "=yay -Sua --noconfirm= → ... → exit 0 on a successful live part". There is no yay failure rule. +- Spec L158: "the press-again-to-force sentinel wrap goes away (the driven path never trips the guard)". +- Spec L167: "the guard hook does not fire". +- man yay, -a/--aur: "Actions such as sysupgrade will only act on AUR packages. Note that dependency resolving will still act normally and include repository packages." +- scripts/hypr-live-update-guard: L62 and L67 check pgrep for Hyprland, L86-105 split targets by version change, L121 prints the BLOCKED banner, L139 exits 1. +- The installed hook (/etc/pacman.d/hooks/hypr-live-update-guard.hook) has AbortOnFail and targets mesa, mesa-*, wayland, libdrm, libglvnd and the others. +- yay 13.0.1, yay -Pg: sudobin "sudo", sudoloop false. +- /etc/pacman.conf:25 has IgnorePkg = bridge-utils only. +- Foreign-package deps on guarded libraries are all unversioned or soname-only: claude-desktop (libdrm, mesa), insync (libglvnd), webkit2gtk (libdrm, mesa, wayland), zoom (mesa, libdrm), mpvpaper (libwayland-client.so=0-64, libwayland-egl.so=1-64). None of the AUR packages is itself a guard target. +- Privilege is already stated: spec L69 ("Its =ExecStart= runs, as the user ... =sudo= works unattended"), L161 ("runs the script's =--complete= form as the user"), L180 ("the unit runs the upgrade as the user via sudo"). +- makepkg:1240-1243 refuses EUID 0. +- The stamp is already covered: L128 says stamp only when the deferred set is empty, and in this scenario it is not empty. +:END: + +** DONE No one-time install mechanism, and Phase 2 breaks the panel on an uninstalled machine +CLOSED: [2026-10-05 Mon 09:20] +Where: Decision L122; Phase 1 L155; Phase 4 L164; Rollout L188; Testing L200 + +The spec says 'existing machines need the one-time install' but gives no command. The ordering is stated only as commit order ('two commits, archsetup first'), and the install itself is deferred to Phase 4, after Phases 2 and 3. + +Risk: The commits land in order, but each machine picks them up independently. A machine that pulls Phase 2 before Phase 1 is hand-installed gets UPDATE/TOPGRADE failures, with the force path already removed. velox, which is offline, will hit this on its next pull. + +Recommended change: Replace the L164 sentence "existing machines need the one-time install" with an explicit per-machine checklist, and point L188 at it: + +"Existing machines: hyprland() runs inside the window_manager step, which is already marked complete in /var/lib/archsetup/state/, so re-running archsetup won't install anything new. On each machine, do these by hand, in this order: +(a) After the Phase-1 commit: pull archsetup, then =install -m 755 scripts/upgrade-guarded scripts/kernel-modules-check /usr/local/bin/=. +(b) Only after (a) on that machine: pull the Phase-2 dotfiles commit. The levers call upgrade-guarded and no longer have a force path, so pulling first breaks UPDATE and TOPGRADE on that machine. +(c) After the Phase-3 commit: install archsetup-boot-upgrade.service, run =systemctl daemon-reload=, enable the unit, and confirm /var/lib/archsetup/ exists. +Re-copy configs/maintenance-thresholds.toml only if Phase 2 adds keys to it. +velox's dotfiles must not move past the Phase-2 commit until (a) is done there." + +Optional, in Phase 2: when upgrade-guarded is not on PATH, the lever reports "upgrade-guarded not installed (archsetup one-time install)" rather than the generic "tool missing or timed out". Don't fall back to the old yay/topgrade argv; that would contradict the decision at L114. + +Should-fix, not blocking. Verification: confirmed. +Disposition: modified, folded into Implementation phases, Readiness dimensions. Confirmed. The checklist gains items other resolutions make mandatory: +- the hook migration (F20), which must precede first use because the script fails closed; +- an unconditional TOML re-copy, because F21 rewrites guard_patterns; +- a DKMS-origin check (F29). + +The checklist is keyed to three commits after the re-phasing. The optional missing-script lever message is adopted. Falling back to the old argv stays rejected, as the finding's check showed. +:EVIDENCE: +Spec: +- L122: "the rollout is two commits, archsetup first". +- L155: script "installed to /usr/local/bin by the step that installs the guard"; kernel-modules-check is a second new script. +- L158: levers "change their argv to the script; the press-again-to-force sentinel wrap goes away". +- L161: Phase 3 adds the unit, the installer step and the /var/lib/archsetup/ flag directory. +- L164: "existing machines need the one-time install", with no steps given. +- L188: same wording, no steps. +- L200: "Roll to velox first, then ratio". + +archsetup (installer): +- Line 257: state_dir="/var/lib/archsetup/state". +- Lines 290-296: run_step prints "Skipping %s (already completed)" when the marker exists. +- Lines 3994-3996: every step goes through run_step. +- Lines 52-56 and 85-99: the only CLI flags are --fresh, --status, --config-file and --autologin. There is no per-step re-run. +- Lines 2606-2654: guard binary cp plus hook heredoc, inline in hyprland() (defined at 2534). +- Lines 2676-2683: window_manager() calls hyprland. +- Lines 1640-1650: thresholds TOML installed with =install -m 0644=. That's a copy, not a stow link, and it's called from user_customizations (1449). + +Live system: +- =ls /var/lib/archsetup/state/= lists window_manager, user_customizations, supplemental_software and others, so re-running archsetup on ratio skips the guard step. + +todo.org: +- Lines 4006-4010: "supplemental_software is a completed step there and doesn't re-run". That is the same mechanism, already hit and fixed by hand per machine. + +dotfiles (maint): +- remedies.py:294-312: update/topgrade levers have "always": True and fixed argv. +- capability.py:65-81: which() checks only for zpool/snapper/docker/virsh. +- cmd.py:21-24: OSError returns None. +- doctor.py:161-163: None is reported as "tool missing or timed out". +- ~/.local/bin/maint and ~/.local/bin/topgrade are symlinks into ~/.dotfiles/hyprland/.local/bin, so a dotfiles pull is live immediately. +- No dotfiles auto-pull timer is installed (systemctl --user list-timers), so ordering per machine can be controlled by hand. +:END: + +** DONE Phase 1 ships arming before any boot consumer exists +CLOSED: [2026-10-05 Mon 09:20] +Where: Phase 1 L155 (--complete arms); Phase 3 L161 (unit + flag dir) + +Phase 1's --complete 'arm[s] the GPU/compositor set', and the gate refuses to 'arm or reboot'. The unit that consumes and removes the flag, and the /var/lib/archsetup flag directory, arrive only in Phase 3. + +Risk: A --complete run between Phase 1 and Phase 3 writes a flag that nothing removes, and its reboot offer applies nothing. When the unit lands, the stale flag fires on the next boot with a package set nobody chose for that day. + +Recommended change: In Phase 1 (L155), add a rule to the --complete description: --complete arms the GPU/compositor set only when archsetup-boot-upgrade.service is installed. Without the unit, it stops after a passing gate. The GPU/compositor set stays in the deferred state file, the script prints "apply from a TTY: upgrade-guarded --complete", and it offers no reboot for that set. Add a Phase 1 test: with the unit absent, --complete never writes the flag and never offers the GPU reboot. This covers both the window between Phase 1 and Phase 3 and the L188 rollback state. The alternative, moving the arm write into Phase 3, fixes the phase window but not rollback. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Implementation phases, Testing. Partially confirmed; this follows the narrowing that the gap is not unsafe and the flag directory already exists. Detection is pinned to one command, systemctl is-enabled, which also covers a unit that is installed but not enabled (F41). Any stale flag is removed, so nothing is left for a later unit to consume unexpectedly. +:EVIDENCE: +- Spec L155 (Phase 1): "--complete (the dedicated-session form: apply the kernel set live, run the gate, then arm the GPU/compositor set or, with no compositor live, apply it directly ...)". Further on: "--complete refuses to arm or reboot on that exit. Usable from a TTY at once." Test: "--complete never reaches the arm step on a failed gate." +- Spec L161 (Phase 3): "archsetup-boot-upgrade.service, installed by the installer ... Installer step + unit file + the /var/lib/archsetup/ flag directory." +- Spec L67: once the gate passes, "the script arms a persistent flag ... and offers to reboot". +- Spec L188: "removing the unit and flag reverts fully". The script stays installed, so it can still arm with no consumer. +- Live system: systemctl cat archsetup-boot-upgrade.service returns "No files found", so the unit does not exist yet. /usr/local/bin has only hypr-live-update-guard and no upgrade-guarded. +- Live system: /var/lib/archsetup exists and contains state/. The installer at archsetup:257 sets state_dir="/var/lib/archsetup/state", so the directory predates Phase 3 and does not block a Phase-1 flag write. +- Spec L142 (Decision 6): the boot oneshot is "scoped to the deferred GPU/compositor set ... no kernel". A late-firing flag applies that set with nothing live, the designed safe path, so it is not an unchosen package set. +:END: + +** DONE Acceptance criteria aren't mapped to verification, and the manual test covers only AC4 +CLOSED: [2026-10-05 Mon 09:20] +Where: Acceptance criteria L167-174; Phase 3 L161; Testing L200 + +The only manual test is 'arm with a guarded lib pending, reboot, confirm', with no setup for producing a pending guarded lib. Nothing verifies AC3 (DKMS failure stops; snapshot bootable from ZBM), AC5 (boot failure or timeout never blocks the session) or AC6 (unread news). Rollout goes straight to velox. + +Risk: Goal 3's safety claim ('can never lock the machine out') and the ZBM fallback are never exercised before velox, where the failure mode is an unbootable machine. The happy-path test can only run on days when upstream happens to publish a guarded update. + +Recommended change: Replace the L200 paragraph with a list mapping each acceptance criterion to how it is verified, and correct L187 to match: + +- AC1, AC2, AC7: Phase 1 pytest beside the guard (ignore set equals blocked plus kernel set, deferred-state round-trip, stamp only on an empty deferred set) and the Phase 2 maint probe test. +- AC3: Phase 1 gate tests cover stopping, naming the failure, and never arming. Either make the ZFSBootMenu snapshot-boot clause a documented recovery drill (in a =make test-keep FS_PROFILE=zfs= VM, or on velox) in todo.org, or move it from the AC into Risks as the fallback. +- AC4: keep the manual test, but give it a trigger and a setup. Run it when =upgrade-guarded --dry-run= lists a GPU/compositor package, or in the VM create a pending guarded lib from a TTY by installing an older cached version. +- AC5: add an installer-steps test that asserts the unit file statically. It should check Before=getty@tty1.service, a TimeoutStartSec, no Requires/BindsTo/RequiredBy tying it to getty or session targets, and flag removal placed where it also runs on timeout (ExecStopPost or remove-first). Also add a VM or manual test with a failing and a sleeping ExecStart test seam. Expected: boot reaches the Hyprland session, the flag is gone, and the panel still shows the deferred set. +- AC6: add a Phase 1 pytest with a fake informant asserting that =informant read --all= runs before the pacman transaction in both the live and --complete forms. Also correct the bare =informant read= at L65, L69 and L172. +- AC8: the existing tests/hypr-live-update-guard suite passes unchanged. + +Finally, state that the Phase 3 boot-unit checks run in =make test-keep FS_PROFILE=zfs= before the velox rollout. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Acceptance criteria, Testing, Risks, Readiness dimensions. Partially confirmed; this follows the narrowing. The map covers the consolidated acceptance list (glossary §18), which adds the criteria other findings introduce, and uses unittest, not pytest (F28). Of the two options for AC3's ZFSBootMenu clause, it moves to Risks as the fallback and is exercised by a non-gating recovery drill. No gate can check it before rollout, and the drill is where F45's pacman -U --dbonly step gets practiced. AC5's static unit test is the single source-inspection test owned by F41. +:EVIDENCE: +Spec: +- L161: "a documented manual boot test (defer, arm, reboot, observe)" +- L200: "scripted manual test ... (arm with a guarded lib pending, reboot, confirm ...). Roll to velox first, then ratio." No setup step for producing a pending guarded lib. +- L155: Phase 1 tests include the gate pass/fail on fakes and "--complete never reaches the arm step on a failed gate", so AC3's stop half is covered. No test mentions informant, the unit's failure or timeout behavior, or the ZBM boot. +- L184: the ordering is only "tested-by-inspection". +- L69: ExecStart "... then removes the flag unconditionally", i.e. inside ExecStart. +- L171: AC5 requires the flag to be cleared on a timeout. +- L187 vs L155/L157/L200: the phase numbers do not match the phases. + +Code and system: +- =pacman -Q informant= on ratio: "error: package 'informant' was not found". +- archsetup:1121-1129: "If the base ships informant (e.g. an archangel-installed system) ... a bare =informant read= is interactive and would hang an unattended run", and the installer uses =informant read --all=. +- Makefile:12-14, 35-36: FS_PROFILE ?= btrfs, with zfs available for test, test-keep and test-vm-base. Makefile:32/108: test-maint runs the break/fix scenarios. +- scripts/testing/archsetup-test-zfs.conf exists. +- scripts/testing/tests/test_boot.py:14-18 asserts /efi/EFI/ZBM/zfsbootmenu.efi on a ZFS root. +- scripts/testing/maint-scenarios/ holds the 10-32 break/fix scripts. +- =grep -rln 'Before=\|TimeoutStartSec\|\[Unit\]' tests/= returns nothing, so no existing static unit-file tests cover this. +:END: + +** DONE Rollback is described as 'remove unit and flag', but the change isn't additive +CLOSED: [2026-10-05 Mon 09:20] +Where: Readiness 'Rollout, compatibility & rollback' L188 + +'additive; removing the unit and flag reverts fully.' + +Risk: Removing only the unit and flag leaves the panel calling the script with no force path. A partial rollback leaves the two repos disagreeing. + +Recommended change: Replace the "Rollout, compatibility & rollback" bullet at spec line 188 with: + +"Rollout, compatibility & rollback: not additive. Phase 2 repoints the UPDATE/TOPGRADE argv and removes the press-again sentinel wrap. Roll back in reverse rollout order: +1. Revert the dotfiles Phase 2/3 commits first. That restores the yay/topgrade argv and the force path, and removes the deferred probe and the apply-on-reboot action. +2. Then remove archsetup-boot-upgrade.service and run daemon-reload. +3. Remove the /var/lib/archsetup/ arm flag, and /usr/local/bin/upgrade-guarded and kernel-modules-check. + +Removing only the unit and flag disables the boot path but leaves apply-on-reboot arming a flag nothing reads. Removing the archsetup scripts before reverting dotfiles leaves UPDATE/TOPGRADE pointing at a missing binary. The rewritten guard_patterns can stay. Existing machines need a one-time install, and a rebuild gets it from the installer. The guard hook is untouched throughout." + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Readiness dimensions. Partially confirmed; this follows the narrowing, adjusted for the three-commit rollout, the sysupgrade alias (F38), and the hook migration, which is not reverted. With F31 in place, removing only the unit no longer leaves --complete arming a flag that nothing reads. +:EVIDENCE: +- Spec line 188: "Rollout, compatibility & rollback: additive; removing the unit and flag reverts fully ... Rollback leaves the guard and manual TTY path intact." +- Spec line 158 (Phase 2): "UPDATE and TOPGRADE levers change their argv to the script; the press-again-to-force sentinel wrap goes away." +- Spec line 121: "the rollout is two commits, archsetup first." +- Spec line 161 (Phase 3): the panel action "installs the held-kernel set live, writes the persistent arm flag, and offers to reboot." +- Spec line 155: Phase 1 installs scripts/upgrade-guarded and kernel-modules-check. +- Spec line 174: the hook is unchanged (already an acceptance criterion). +- Dotfiles code I checked: + - ~/.dotfiles/maint/src/maint/remedies.py:282-293 holds guard_sentinel_set and guard_sentinel_clear. + - remedies.py:297 sets the update argv to ["yay","-Syu","--noconfirm"]. remedies.py:308 sets the topgrade argv to ["topgrade","--disable","git_repos","-y"]. Both have guard "live_update". + - ~/.dotfiles/maint/src/maint/doctor.py:174-188 is _all_steps, which adds the sentinel wrap. doctor.py:242 runs that wrap when force is set. + These are exactly what Phase 2 deletes or repoints, so the change is not additive. +- Thresholds TOML: + - configs/maintenance-thresholds.toml:65-72 guard_patterns are mesa, mesa-*, lib32-mesa*, hyprland, hyprland-*, aquamarine, hyprutils, hyprlang, hyprcursor, hyprgraphics, *wayland*, wlroots*, vulkan-radeon, lib32-vulkan-radeon, vulkan-intel and lib32-vulkan-intel. + - The hook Target list in archsetup:2628-2642 adds libdrm, libglvnd, vulkan-mesa-layers, nvidia-utils, lib32-nvidia-utils and xorg-xwayland, and has no hyprlang, hyprcursor or wlroots*. + - So the Phase 2 pin test does force a TOML rewrite. The maint thresholds.py:31-33 reads this file as the shipped layer at ~/.config/archsetup/maintenance-thresholds.toml. +- Cache location: maint cache.py:16-17 is ~/.local/state/maint, which is where the deferred-set "own cache key" from spec line 128 would land. +:END: + +** DONE Armed-but-not-rebooted state is invisible, and REBOOT semantics are ambiguous +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L67 ('arms … and offers to reboot'); Decision 4 L128 ('apply on reboot'); Phase 2 L158; Phase 3 L161 + +No phase says how the panel shows a pending arm, how the 'offer to reboot' appears in the panel, or how to disarm. + +Risk: A user reading 'N deferred — apply on reboot' may press the REBOOT key the strip already shows, and nothing is applied because nothing was armed. A user who arms on ratio gets no reboot affordance at all. On both machines, a pending arm is invisible. + +Recommended change: Phase 2 (L158): replace "N deferred — apply on reboot" with "N deferred". Pair it with a key that runs =upgrade-guarded --complete=, and say plainly that a plain reboot applies nothing until a --complete run has armed. Fix the example string at L128 the same way; this changes only the illustration, not the stamp decision. Phase 3 (L161): the deferred probe also reads the arm flag. While the flag is present, the row reads "armed — N apply at next boot". The panel action calls offer_reboot() when --complete exits armed, rather than relying on reboot_required, which stays false on ratio because the booted linux-lts-strix is a local build the kernel set never upgrades. --complete prompts for a reboot only when stdin is a tty; from the panel, the REBOOT key is the offer. Add one acceptance criterion: after arming from the panel, the row shows the armed state and REBOOT is offered, on both ratio and velox. A --disarm command is optional and can be left out. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Decisions, Implementation phases, Acceptance criteria. Partially confirmed; this follows the narrowing, so there is no --disarm. Under F40 the panel launches --complete in a detached terminal and never sees its exit. 'Call offer_reboot() when --complete exits armed' therefore becomes a state rule: REBOOT shows while the arm flag exists. That survives a panel close, and it works on ratio, where reboot_required stays false. +:EVIDENCE: +Spec: +- L67: "arms a persistent flag ... and offers to reboot" +- L128: "('6 deferred — apply on reboot')" +- L135: the GPU set follows only after the gate passes +- L158: "A new probe reads the deferred-set state file and the panel renders 'N deferred — apply on reboot' as its own row" (no arm-flag input) +- L161: "The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot" (no display of the armed state, no mechanism for the offer) + +Code (~/.dotfiles/maint/src/maint): +- gui.py:1509-1511: if rid in ("update","topgrade") and primary ok, call self.model.offer_reboot() +- panel.py:501-510: reboot_key_visible() is reboot_offer OR the reboot_required value +- gui.py:491-495: the REBOOT strip key is gated on reboot_key_visible() +- packages.py:131-146: reboot_required is true only when /usr/lib/modules/$(uname -r) is missing +- cmd.py:18-24: subprocess.run(cmd, capture_output=True, text=True, timeout=...) +- doctor.py:149-161: user steps go through cmd.run +- grep -rn "/var/lib/archsetup|deferred|upgrade-guarded" over maint/src/maint finds nothing + +Live checks on ratio: +- uname -r: 6.18.25-1-lts-strix +- pacman -Qo /usr/lib/modules/6.18.25-1-lts-strix: owned by linux-lts-strix 6.18.25-1 +- pacman -Si linux-lts-strix: "package 'linux-lts-strix' was not found" +- pacman -Qm lists linux-lts-strix; Packager: Unknown Packager +- No linux-lts-strix-headers is installed. linux, linux-lts, and both of their headers packages are separate packages. +:END: + +** DONE Deferred row goes stale after a successful boot run or out-of-band completion +CLOSED: [2026-10-05 Mon 09:20] +Where: Decision 4 L128; Design L71 (by-hand path); Decision 2 L114 (manual TTY path C); AC L167 + +The boot run 'stamps when it completes the deferred set', but the spec never says it rewrites the deferred state. A by-hand sentinel run or a bare TTY pacman -Syu never touches the file at all. + +Risk: After a successful boot run, the panel shows a fresh topgrade_age next to 'N deferred', and the two contradict each other. AC L167 ('exact deferred set') fails after any out-of-band completion. + +Recommended change: Phase 1 (L155) and Decision L128: +- Add: "Every mode, --complete and the boot run included, rewrites the deferred-set state file at the end of its run with what is still deferred, empty when nothing is. --complete drops the kernel set once it is installed and the GPU/compositor set once it is applied." +- In L128, change "When it is non-empty it writes the deferred set" to "It always writes the deferred set (empty when nothing was held)". +- Add a Phase 1 test: after a successful --complete or boot run, the state file is empty and the stamp is fresh. + +Phase 2 (L158), optional robustness for the out-of-band paths L71 names: the deferred probe drops any recorded entry whose installed version (pacman -Q, compared with vercmp) is at or above the recorded new version. A by-hand sentinel or TTY completion then clears the row without rerunning the script. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: accepted, folded into Decisions, Implementation phases, Testing. +:EVIDENCE: +Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org: +- L128: "the script stamps topgrade_run only when the deferred set is empty. When it is non-empty it writes the deferred set to its own cache key ... The boot oneshot stamps when it completes the deferred set." This line has no clear and no rewrite. +- L155: the live pipeline includes "deferred set written to a state file". The --complete clause reads "apply the kernel set live, run the gate, then arm ... or ... apply it directly; stamp when the deferred set is empty", with no state-file write. +- L142: the boot oneshot runs "the script's --complete form scoped to the deferred GPU/compositor set". +- L158: the probe "reads the deferred-set state file". +- L181: "the panel reflects the cleared or still-pending state after boot". +- AC L170 checks only topgrade_age after the boot run, not the deferred row. + +Code: ~/.dotfiles/maint/src/maint/doctor.py:309-324 refreshes the pending cache (p_updates.scan_net) only after a successful update/topgrade lever run. ~/.dotfiles/maint/src/maint/probes/updates.py:145-160 shows scan_net writing updates_repo from checkupdates and updates_aur from yay -Qua, and nothing else deferral-related. The user unit maint-net-scan.timer runs it hourly (OnUnitActiveSec=1h). grep for "deferred" and "upgrade-guarded" across maint/src and archsetup/scripts returns nothing, so no existing code pins this behaviour. +:END: + +** DONE CVE strip and card claim held packages are fixable via UPDATE +CLOSED: [2026-10-05 Mon 09:20] +Where: Decision 5 consequences L136 ('security fixes ride the kernel'); Risks L197; Phase 2 L158 + +The spec doesn't mention the CVE metric, which treats every queued package as something UPDATE will land. + +Risk: A fixable kernel or mesa advisory stays red after a successful UPDATE that, by design, cannot apply it. On the one signal that tracks security, the panel points at the wrong action. + +Recommended change: Add to Phase 2 (L158), after the deferred-row sentence: "Advisories on deferred packages count against the deferred row, not UPDATE. cve_queued (the red badge and strip), the CVE card's 'fixable via UPDATE' caption, and the UPDATE lever's cve_advisories binding all exclude names in the deferred set. The deferred row names those advisories ('N deferred · K fixable advisories') and goes to WARN when K > 0, so the kernel security fixes from the kernel decision have a visible reminder that points at --complete." Add a Phase 2 test: given a fixable advisory naming a deferred package, cve_queued returns 0, the card caption leaves it out, and the deferred row carries it at WARN. Optionally, extend acceptance criterion L167 with: "a fixable advisory on a deferred package shows on the deferred row, not as UPDATE-fixable." + +Should-fix, not blocking. Verification: confirmed. +Disposition: modified, folded into Implementation phases, Acceptance criteria, Testing. Confirmed. This adds the finding's wider scope: REVIEW & FIX also stops offering UPDATE for an advisory on a deferred package. It also absorbs F48's cve_queued clause, so the exclusion has a single owner. +:EVIDENCE: +Spec: L65 and L135 hold the kernel and guard sets on every everyday run. L136 says "security fixes ride the kernel". L158 is the full Phase 2 display change, a deferred probe and the "N deferred — apply on reboot" row, with no CVE handling. L197 names the deferred row as the reminder for kernel security fixes. L167 (acceptance) covers only the deferred set. + +Code (~/.dotfiles/maint/src/maint/): +- viewmodel.py:503-516: the cve_queued docstring says the strip "only turns red when running UPDATE would actually close an advisory". It counts any fixable advisory whose package is in updates_repo or updates_aur. +- viewmodel.py:1012-1018: _card_cve returns "{fixable} fixable via UPDATE". +- gui.py:641-672: _cve_queued feeds both the faceplate badge and the red strip class from _cache_list("updates_repo"). +- probes/updates.py:75-97: the cves() docstring says "the fix is one UPDATE away", and the metric is WARN whenever any advisory has =fixed= set. +- remedies.py:294-302: the "update" lever has metric_ids ["updates_repo","updates_aur","cve_advisories"]. +- doctor.py:400-405: review() offers that lever whenever cve_advisories is WARN. +- probes/updates.py ~160: updates_repo is written from checkupdates, so held packages stay in the queue after a split run. + +Live (ratio): checkupdates currently lists linux, linux-headers, linux-lts, linux-lts-headers, mesa, hyprland, vulkan-radeon and vulkan-mesa-implicit-layers. That is exactly the set a split run would hold, so every one of them would stay in updates_repo. arch-audit currently reports no fixable advisories, so the scenario is latent rather than active today. +:END: + +** DONE Stale-kernel tracking is neither in v1 nor a tracked vNext item +CLOSED: [2026-10-05 Mon 09:20] +Where: Scope tiers L57 ('vNext: none open'); Decision 5 consequences L136; Risks L197 ('a stale-kernel age … is a possible follow-up') + +The spec says the dedicated session 'has to happen on a cadence'. The only reminder is the deferred row, which is present most days, and the stale-kernel age is a 'possible follow-up' that no scope tier lists. + +Risk: The security half of the standing kernel hold has no owner. A row that shows nearly every day habituates the user, so an unpatched kernel ages silently. + +Recommended change: Change two spec lines, and keep v1 scope as it is: +1. L57: replace "vNext: none open — …" with "vNext: a stale-kernel age in maint (how long the kernel set has been held, graded so an overdue dedicated session shows up); logged to todo.org as [#D]. The original live-kernel vNext is now part of v1's held set." +2. L197: change "a stale-kernel age in maint is a possible follow-up" to "a stale-kernel age in maint is the vNext follow-up (see Scope tiers)". + +At decomposition, outside the spec, replace todo.org:740's stale "file the vNext [#D] kernel-reboot item" with that stale-kernel-age [#D] item. Do not add deferred_warn_days or hold-age grading to v1 unless I choose to. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: accepted, folded into Goals and Non-Goals, Risks. +:EVIDENCE: +Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org: +- L57: "vNext: none open — the kernel deferral that was vNext is now part of v1's held set." +- L136: "the kernel deferral is now standing, so the dedicated session has to happen on a cadence (security fixes ride the kernel), and the panel's deferred count carries a kernel most days" +- L197: "The panel's deferred row is the reminder; a stale-kernel age in maint is a possible follow-up." +- git show 77447d0 L54, the original vNext: "fold the same arm-and-reboot pattern into a live-kernel-upgrade prompt (log to todo.org)". e777ddc, the decision-closing commit, added L57's "none open" and L197 together. + +todo.org: +- :739-740: "...file the vNext =[#D]= kernel-reboot item, and commit the spec." +- :759: the stale note covers only "the 'four open decisions' list above". +- A grep for stale-kernel, kernel age or kernel-age finds no task. + +.ai/workflows/spec-review.org: +- :33: "Deferred work is logged to todo.org (v1 = [#B], vNext/someday = [#D])" +- :122: "Are deferrals captured (todo.org)?" + +~/.dotfiles/maint/src/maint/probes/updates.py:88-89: +- =fixable = [a for a in data if a.get("fixed")]; sev = schema.WARN if fixable else schema.OK= + +~/.local/state/maint/cve.json, kernel entries: +- AVG-2701 and AVG-2683 [linux-lts], fixed=None +- AVG-1879, AVG-2345 and AVG-1594 [linux], fixed=None +- So the CVE probe never flags a held kernel. + +There is no kernel or deferred threshold in configs/maintenance-thresholds.toml [updates] (lines 57-75). +:END: + +** DONE Routine update entry points outside the panel still land the kernel ungated +CLOSED: [2026-10-05 Mon 09:20] +Where: Non-Goals L51; Design L65 ('backstop for a bare pacman -Syu'); Decision 5 consequences L136; Phase 4 L164; Documentation plan L186; history L204-206 + +The spec routes only the panel's UPDATE and TOPGRADE levers through upgrade-guarded. The documented routine update path is untouched: system-health-check Phase 3 step 6 runs plain topgrade, and the sysupgrade alias runs topgrade. topgrade's system step is yay -Syu, which upgrades every kernel and zfs-dkms. The guard hook has no kernel Target, so on kernels it isn't the backstop L65 calls it. The same workflow also tells the operator to run 'maint stamp topgrade' by hand. The containers fix lives only in the script's CLI --disable, so a plain topgrade still fails on containers, and that failure is what drives the hand stamp. + +Risk: Decision 5's consequence, 'an everyday UPDATE can never put velox into the unbootable state', holds only for the panel. The update velox actually gets, the health-check topgrade, still swaps linux-lts and zfs-dkms with no DKMS/initramfs/snapshot gate. That is the exact failure chain the decision was reversed to close. The hand-stamp instruction also writes freshness while a deferral is outstanding. + +Recommended change: Three edits. + +1. Add to Phase 4 (L164), after "Document the flow": + +"Update docs/workflows/system-health-check.org Phase 3 so the routine update goes through the script. Step 6 runs upgrade-guarded instead of plain topgrade. A day with a pending kernel or guarded library lands it through upgrade-guarded --complete as the dedicated session. Step 5's ratio strix addendum keys on --complete instead of 'before topgrade'. The 'maint stamp topgrade by hand' fallback becomes 'the script stamps; never hand-stamp while a deferred set is outstanding'. Append a resolution note to the 2026-09-12 containers KIL entry rather than rewriting it." + +2. Add the health-check workflow to the Documentation plan line (L186). + +3. Add one line to Risks, to spell out the consequence of L51: + +"Bare topgrade, yay and pacman -Syu (including the sysupgrade alias) still land the kernel set with no DKMS gate, because the hook is silent on kernels. The hold covers only the script's callers." + +Optional: point the sysupgrade alias at upgrade-guarded. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Implementation phases, Readiness dimensions, Risks. Confirmed in substance; this follows the narrowing that the bare-command exposure needs one explicit Risks line. The optional alias change is taken. sysupgrade is the routine terminal path on both machines (alias sysupgrade="topgrade" at line 42 of both shell alias files). While it points at plain topgrade, it is exactly the ungated kernel and zfs-dkms route that Decision 5 closes. Repointing it is a one-line dotfiles change in the same commit as the levers, and it inherits the same step (a) ordering. +:EVIDENCE: +Spec facts: +- L51: "The hook stays silent on kernels; the split script holds them back on a live run as a second, separately-reasoned list ... a deferral policy rather than a guard." This already concedes that bare commands do not gate kernels. +- L65: the backstop sentence refers to the GPU guard. +- L131 title: "lands only in the dedicated session". L136: "an everyday UPDATE can never put velox into the unbootable state". +- L164 Phase 4 and L186 Documentation plan never mention docs/workflows/system-health-check.org. Grepping the spec for health, sysupgrade or workflow finds only L206: "Artifacts: velox health checks of 2026-09-12 and 2026-10-04". + +Code and docs: +- docs/workflows/system-health-check.org:236, Phase 3 step 6: "Run =topgrade= for the actual update ... If the run happened outside the wrapper somehow, =maint stamp topgrade= records it by hand." +- :235, step 5: the ratio strix addendum runs "before topgrade" when linux or linux-lts is pending. +- :698: "If topgrade is present, Phase 3 uses it rather than the raw package manager." +- :1067: the KIL entry tells the operator to run "=maint stamp topgrade= by hand". +- todo.org:762-765 (2026-09-17): "velox has linux-lts 6.18.51 -> 6.18.52 pending today, which is the kernel-hold case this spec exists for, so no plain topgrade on velox until the hold is built or the kernel update runs as its own session." +- .dotfiles/common/.zshrc.d/aliases.sh:42 and .bashrc.d/aliases.sh:42: alias sysupgrade="topgrade". +- .dotfiles/common/.config/topgrade.toml:131: arch_package_manager = "yay". The misc disable list at :19 is ["emacs","poetry","gnome_shell_extensions","lensfun","uv"], with no containers, so a plain topgrade still fails at the containers step. +- archsetup:2625-2649: the hook Targets list no linux* package. +- /etc/pacman.conf IgnorePkg = bridge-utils only. +- pacman -Q on ratio shows linux, linux-lts, linux-lts-strix and zfs-dkms 2.4.4 installed, all with nothing pinning them. +- hyprland/.local/bin/topgrade wrapper comment: it stamps for "the system-health-check workflow, a plain shell". +:END: + +** DONE The split run discards Arch news unread +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L65 ('clears the news hook (informant read) where installed'; topgrade --disable system); Design L69; AC L172; Reuse L183 + +Each everyday run marks all Arch news read before the transaction and disables topgrade's system step. Nothing in the flow shows the news. Today the news gets attention two ways: on velox, informant's AbortOnFail hook blocks the transaction until the news is read, and topgrade's show_arch_news prints it during the system step. The split run removes both, and maint has no news surface. + +Risk: Manual-intervention notices are often about exactly the mesa, wayland and kernel transitions this design defers. With this change they're silently marked read, or never shown on ratio, right before an unattended 'pacman -Syu --noconfirm'. That removes the one interlock that forced a human to read them. + +Recommended change: Design L65 and Phase 1 L155: before the transaction, the script records the unread Arch news titles in its state file, using =yay -Pwq=, which needs no informant and works on both machines. It must run before =pacman -Syu=, because yay counts news as "new" relative to the build dates of installed packages. Only then does it clear the hook with =informant read --all= where informant is installed. + +Make the same =--all= correction to the boot ExecStart at L69, and to the parenthetical in AC L172. A bare =informant read= is interactive and hangs or fails with no tty. + +Phase 2 (L158): the deferred-set panel row also lists the recorded news titles until the user dismisses them. + +Risks: add one bullet. The driven path marks news read, so on velox informant's AbortOnFail hook no longer stops a run, and the panel row replaces it as the place news gets seen. + +Leave the auto-clear as designed. Whether unread news should instead stop the everyday run is an open call for the owner, not a review fix. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Design, Implementation phases, Risks, Testing. Partially confirmed; this follows the narrowing. The auto-clear stays, and whether unread news should stop a run stays my open call. The text also defines what the recommendation left open: +- capture happens only in modes with network, never at boot; +- titles carry forward until dismissed; +- dismissal is a maint-owned key, so the script never reads panel state and stays the record's only writer. +:EVIDENCE: +Spec L65: "clears the news hook (=informant read=) where installed; runs =pacman -Syu --noconfirm --ignore== ... then runs =topgrade --disable system,git_repos,containers -y=". Spec L69: the boot ExecStart runs "=informant read= (clear the news hook that would otherwise abort the transaction)". Spec L172 (AC): "Unread Arch news does not wedge the boot run". The spec never mentions surfacing news; the only hits for "news" are L65, L69, L172 and L183. + +archive/task-archive.org:604 says "informant: the base ships informant; its pacman PreTransaction hook (AbortOnFail) blocked archsetup's first transaction. Fix: informant read --all up front (guarded). PROVEN." + +archsetup:1121-1129 reads "registers a pacman PreTransaction hook (AbortOnFail) that blocks every package transaction while Arch news is unread ... --all marks without printing or prompting; a bare =informant read= is interactive and would hang an unattended run." + +~/.dotfiles/common/.config/topgrade.toml:144 sets show_arch_news = true, and :131 sets arch_package_manager = "yay". Both are under [linux], so they only take effect in the system step. + +=grep -rni 'news|informant' .dotfiles/maint/src/maint= finds no news probe (the only hit is the pacnew comment at gui.py:93). + +=pacman -Q informant= on ratio returns "error: package 'informant' was not found". + +~/.dotfiles/maint/src/maint/doctor.py:161-168: levers run via cmd.run(argv) and keep only =out[-1]= of stdout as the detail. cmd.py:18-24 shows cmd.run is subprocess.run(capture_output=True). So today's panel TOPGRADE path already drops topgrade's news output on ratio. + +=man yay=: "-w, --news Print new news ... News is considered new if it is newer than the build date of all native packages", and "-q Only show titles". That gives an informant-free title source on both machines, as long as it runs before the upgrade. +:END: + +** DONE The panel's lever runner can't host the long --complete session +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L67 ('run in the foreground from the panel's action'); Phase 2 L158; Phase 3 L161 ('The panel action installs the held-kernel set live ... and offers to reboot'); Readiness Performance L182 + +The spec runs the everyday split run and the dedicated kernel session through the maint lever path. That runner captures output, has no tty, and kills on a fixed timeout. It runs on a daemon thread that dies with the panel window. The spec specifies no timeout, detachment or log for the panel-hosted runs; the readiness items cover only the boot unit. + +Risk: A stray Escape, a waybar click, or the 3600 s ceiling kills the script partway through the kernel session, after the kernel transaction and before the gate verdict, the state-file write or the stamp. The panel reports 'tool missing or timed out' while pacman or DKMS may still be running orphaned. The 'offers to reboot' prompt has no stdin to read. + +Recommended change: Phase 3, L161: replace "The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot." with: "The panel action does not use the captured-output lever runner. Like MERGE, it opens a terminal running the dedicated session (foot -e upgrade-guarded --complete, via Popen with start_new_session=True), so closing the panel cannot kill or orphan the kernel transaction. The script prints the gate verdict there, writes the arm flag, and asks there whether to reboot. The panel picks up the outcome from the deferred-set state file on its next probe. The arm action's maint test asserts the terminal argv and the detach." Design L67: change "run in the foreground from the panel's action" to "run in a terminal the panel's action opens". Leave Phase 2 alone: the everyday levers already run through this runner with timeout 3600, and the argv change keeps that. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Design, Implementation phases, Testing. Partially confirmed; this follows the narrowing that the everyday levers keep the existing runner. foot gets --hold so the gate verdict stays readable after the script exits. The action lands in Phase 2 with its row (F16). +:EVIDENCE: +Spec: +- L67: "run in the foreground from the panel's action or upgrade-guarded --complete in a terminal ... offers to reboot". +- L161: "The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot ... the arm action's tests in maint". +- L181-182: Observability and Performance cover only the boot unit. +- L65: the script already writes the deferred set to a state file. + +maint (~/.dotfiles/maint/src/maint): +- doctor.py:161: proc = cmd.run(step["argv"], timeout=r.get("timeout", 60)). +- doctor.py:164-168: only the last stdout line (160 chars) and a 500-char stderr tail reach the wall; "tool missing or timed out" on None. +- cmd.py:22: subprocess.run(cmd, capture_output=True, text=True, timeout=timeout), catching TimeoutExpired and returning None. +- CPython subprocess.run source (checked live): except TimeoutExpired: process.kill(); process.wait(); raise. That kills the direct child only, and the Popen context exit closes the pipes. +- remedies.py:294-313: update and topgrade are already kind "user", timeout 3600, so the everyday hazard is pre-existing. +- gui.py:1452/1465: self.firing = True; threading.Thread(target=run, daemon=True). +- gui.py:342-344: MaintPanel(Gtk.Application) with application_id and no hold(); gui.py:59 is GTK 4.0; no set_hide_on_close anywhere. So close() destroys the only ApplicationWindow and the app exits. +- gui.py:358-361: activate on an already visible window calls close() (waybar on-click "maint-panel", waybar/config:84). +- gui.py:1806-1810 + panel.py:246-254: Escape and q close the window. gui.py:421: the close button does too. No close path checks self.firing. +- Detach precedent: gui.py:94-96 _MERGE_ARGV = ["foot","-e","sudo","pacdiff"]; gui.py:1672/1756 subprocess.Popen(..., start_new_session=True). +- Kernel transaction weight: /var/log/ratio-upgrade.log:1487-1499 shows "(17/34) Install DKMS modules" (zfs for two kernels), then "(21/34) Updating linux initcpios". + +Adjacent fact, not part of this finding: doctor.py:169-170 stamps topgrade_run on any exit 0 when rid == "topgrade". If TOPGRADE's argv becomes the script (which exits 0 even when it deferred packages), that stamp contradicts the stamp-only-on-empty decision at L128 unless Phase 2 removes it. +:END: + +** DONE Phase 3 install step isn't testable by the installer-steps harness, and unit enablement is unspecified +CLOSED: [2026-10-05 Mon 09:20] +Where: Phase 3 L161; Phase 1 L155 ('the step that installs the guard'); Architecture L184 ('tested-by-inspection') + +Phase 3 says 'Installer step + unit file + the /var/lib/archsetup/ flag directory. Tests in tests/installer-steps/'. It names no function and no location for the unit source. Nothing says how the unit gets pulled into boot: there is no [Install]/WantedBy and no enable step. + +Risk: The 'tested-by-inspection' ordering claim has no harness to land in. A unit with no WantedBy never runs, so the manual test fails for reasons unrelated to the design. + +Recommended change: In Phase 3 (L161), after "non-fatal to every session target", add: "[Install] WantedBy=multi-user.target, enabled by the installer step (wanted, never required, so its failure can't fail the target)". Replace "Tests in =tests/installer-steps/= for the install step" with: "Tests in =tests/installer-steps/=: a source-inspection test in the style of test_pacman_hook_order.py asserting the unit text carries ConditionPathExists on the flag, Before=getty@tty1.service, an explicit TimeoutStartSec and WantedBy=multi-user.target, and that the installer enables the unit". At L184, change "tested-by-inspection" to "pinned by that installer-steps source-inspection test". Don't name the function and don't decide the network question in this edit. Separately, fix the stale phase numbers at L187 and L200 so the installer-step tests belong to Phase 3 and Phase 1's tests are the archsetup pytest beside the guard. + +Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Implementation phases, Readiness dimensions, Testing. Partially confirmed; this follows the narrowing that the harness can source-inspect inline heredocs and no function is named. The assertion list grows to every directive the glossary settles. It replaces F32's separate static unit test, so there is one such test. +:EVIDENCE: +- Spec L69 lists ConditionPathExists, Before=getty@tty1.service, TimeoutStartSec and ExecStart. Spec L161 repeats them, adds "non-fatal to every session target", and says "Tests in tests/installer-steps/ for the install step". Neither line has an [Install] section or an enable step. Confirmed. +- archsetup:2534-2655: the guard script copy and the 10-hypr-live-update-guard.hook heredoc are inline in hyprland(), next to pacman_install calls. Confirmed. +- tests/installer-steps/test_pacman_hook_order.py:25-36 (written_hooks parses the hooks the source writes) and :42-45 (setUpClass reads the whole archsetup file) and :53-60 (asserts on the hook filename written inside hyprland()): a source-inspection test that already covers content inline in hyprland(). test_clone_user_repos.py, test_configure_service_discovery.py, test_install_camera_passthrough_rules.py and test_orchestrators.py also read the source text. This refutes "can't be extracted / no harness to land in". +- More than 25 installer-steps tests use the sed -n '/^name() {/,/^}/p' extraction pattern, so convention already steers an implementer toward a top-level function. Spec L161 never mandates inline placement. +- archsetup:2286-2299 (zfs-replicate.service: Type=oneshot, [Install] WantedBy=multi-user.target) and archsetup:2315 (installer runs systemctl enable on the unit set up in that block): an in-repo precedent for a oneshot unit written by heredoc and then enabled. +- Live ratio: getty@.service has Before=getty.target. /etc/systemd/system/getty.target.wants/getty@tty1.service exists. multi-user.target Wants getty.target. So WantedBy=multi-user.target puts the oneshot in the same boot transaction as getty@tty1, and Before= takes effect. +- Live ratio: NetworkManager-wait-online.service is enabled (matches the finding), but the network/pre-download question belongs to the separate package-source finding. +- Spec L187 ("installer-step pytest for Phase 2; maint unit tests for Phases 1 and 3") and L200 ("Phase-2 install under tests/installer-steps/") contradict L155, L158 and L161 (Phase 1 = archsetup pytest beside the guard, Phase 2 = dotfiles maint, Phase 3 = installer step). +:END: + +** DONE Documentation plan names no files, and current docs describe removed behaviour +CLOSED: [2026-10-05 Mon 09:20] +Where: Readiness 'Documentation plan' L186; Phase 4 L164 + +'a short "reboot to apply guarded upgrades" note in the maint docs; the installer step self-documents in-comment.' + +Risk: After Phase 2, the README and the TOML comment describe a force path that no longer exists and a stamping rule the decision reversed. + +Recommended change: Replace the Documentation plan bullet at L186 with: "Documentation plan: Phase 2 rewrites the 'The live-update guard' section of dotfiles maint/README.md. It drops press-again/--force, says UPDATE/TOPGRADE run upgrade-guarded (topgrade with --disable system,git_repos,containers), and describes the 'N deferred' row. The same phase updates the [updates] guard_patterns comment in archsetup configs/maintenance-thresholds.toml to say the patterns are the panel's display mirror of the hook, pinned by a test. Phase 4 adds the arm → reboot → console upgrade flow to that README section. The installer step self-documents in-comment." Add one clause to Phase 2 (L158): "and remove =maint fix --force=, which only served the guard override." Leave the guard banner and the wrapper paragraph alone in this finding. The wrapper-stamps-on-the-script's-own-topgrade-call conflict with L128 should be raised as its own correctness finding. + +Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Implementation phases, Readiness dimensions. Partially confirmed; this follows the narrowing that the guard banner and the wrapper paragraph stay untouched. One detail is fixed: the README gives the space-separated --disable form (F02), because the recommended text quotes the comma form, which topgrade 17.12.2 rejects. The TOML comment edit happens with F21 in Phase 1, since both files are archsetup's. +:EVIDENCE: +Spec L158 (Phase 2): "the press-again-to-force sentinel wrap goes away". L155: topgrade runs as =--disable system,git_repos,containers -y=. L186: "Documentation plan: a short 'reboot to apply guarded upgrades' note in the maint docs". L164: "Document the flow". + +~/.dotfiles/maint/README.md:54-58: + "The panel arms with override wording (press again to force); the CLI equivalent is =--force= ... Topgrade always runs with =--disable git_repos=." +Only README.md and src/ are present in ~/.dotfiles/maint, so the README is the only maint doc. + +configs/maintenance-thresholds.toml:63-64: + "a hit arms UPDATE/TOPGRADE instead of running ('press again to run anyway — or apply from a TTY')" + +Current code that Phase 2 removes: +- doctor.py:216-242 (guard refusal, then force-wrapped sentinel) +- remedies.py:308 (=topgrade --disable git_repos -y=) +- cli.py:265-267 (--force help: "run a guarded remedy despite the live-update guard") + +Overstated parts: +- Wrapper: the spec never changes it. Wrapper lines 40-42 stamp on rc 0, and README:60-64 matches that. ~/.local/bin/topgrade -> .dotfiles/hyprland/.local/bin/topgrade, and ~/.local/bin is first on PATH. +- Banner: AC8 L174 says the hook is unchanged, and Non-goal L49 says the guard stays exactly as strict. The hook's Exec is /usr/local/bin/hypr-live-update-guard (archsetup:2648), so editing the banner edits what the hook runs. +:END: + +** DONE Proof-of-concept 'exact sequence' did not hold kernels or zfs-dkms +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L65 (proof-of-concept sentence); history L213 + +'Proof of concept: this exact sequence, run by hand on ratio on 2026-08-25 while Hyprland was live, resolved 724 of 730 pending packages ... with the six guard hits deferred.' + +Risk: This overstates the evidence: the kernel hold and the containers disable were never exercised. Under v1, the same day would defer 10 or more packages, not 6. + +Recommended change: In Design L65, replace "Proof of concept: this exact sequence, run by hand on ratio on 2026-08-25 while Hyprland was live, resolved 724 of 730 pending packages (Emacs 31.1 among them) with the six guard hits deferred" with "Proof of concept: the guard-set half of this sequence, run by hand on ratio on 2026-08-25 while Hyprland was live and before the kernel hold and the containers disable were added, resolved 724 of 730 pending packages (Emacs 31.1 among them) with the six guard hits deferred. That run upgraded linux, linux-lts and their headers live, so under v1 the same day would have deferred ten". Keep the trailing qemu-block-gluster clause as is. Leave history L213 unchanged. + +Optional, not blocking. Verification: confirmed. +Disposition: modified, folded into Design. Confirmed. F04, which I accepted, puts zfs-dkms in v1's held set, and zfs-utils follows through the closure. Both moved 2.4.3 → 2.4.4 in that run, so the earlier 'drop the zfs-dkms clause' narrowing no longer holds. The count becomes 'at least twelve'; it is a lower bound because that day's closure can't be reconstructed. +:EVIDENCE: +- Spec L65: "Proof of concept: this exact sequence, run by hand on ratio on 2026-08-25 while Hyprland was live, resolved 724 of 730 pending packages (Emacs 31.1 among them) with the six guard hits deferred". +- Spec L65 also defines that sequence as including "plus the *kernel set* ... held on every everyday run" and "topgrade --disable system,git_repos,containers -y". +- Spec L135: the kernel set is held "on every everyday run, on both machines". +- Spec L204: containers was added on 2026-10-05. +- =/var/log/ratio-upgrade.log= L7-12: exactly six "ignoring package upgrade" lines (aquamarine, hyprland, hyprutils, mesa, vulkan-radeon, wayland). +- =/var/log/ratio-upgrade.log= L16: "Packages (724)". +- =/var/log/pacman.log= L20243: "upgraded linux (7.1.5.arch1-2 -> 7.1.9.arch1-2)". +- =/var/log/pacman.log= L20261: "upgraded linux-headers". +- =/var/log/pacman.log= L20262: "upgraded linux-lts (6.18.41-1 -> 6.18.46-1)". +- =/var/log/pacman.log= L20263: "upgraded linux-lts-headers". +- =/var/log/pacman.log= L20503-20504: zfs-utils and zfs-dkms 2.4.3 -> 2.4.4. +- =/var/log/ratio-upgrade.log= L1488-1489: "dkms install --no-depmod zfs/2.4.4 -k 7.1.9-arch1-2" and "-k 6.18.46-1-lts". +- =/var/log/ratio-upgrade.log= L1496 onward: the initramfs is rebuilt for the new kernels. +- The log ends at the pacman post-hooks. =systemctl cat ratio-upgrade.service= returns "No files found" (it was a transient unit). +:END: + +** DONE Summary and Design principle describe only the first draft +CLOSED: [2026-10-05 Mon 09:20] +Where: Summary L28; Goals L43-46; Design opening L61, L63; Alternative D heading L90 + +The Summary and Goals cover only the completion path. Design L61 says 'the ordinary full sweep stays a normal live topgrade run' and 'apply the guarded system upgrade with Hyprland down, once, and stamp it'. L63 says 'Three pieces, at two altitudes' but never lists them. D is labeled '(this spec)'. + +Risk: A reader who skims the Summary and the opening principle gets the superseded design before reaching the decisions. + +Recommended change: This is the smallest edit set; the first two items carry the substance. + +1. Summary L28: after "...record the freshness stamp", add: "The everyday path is a live split run (upgrade-guarded) that applies everything the guard would not block, holds the GPU/compositor and kernel sets, and reports what it deferred. Landing that deferred set is a dedicated, gated session that finishes at a boot-time oneshot." + +2. L61: replace "while the ordinary full sweep stays a normal live =topgrade= run" with "while everything else, the ecosystem sweep included, runs live through the split script, which defers the guarded and kernel sets". + +3. Goals: add one bullet: "An everyday UPDATE/TOPGRADE that succeeds live by deferring the guarded and kernel sets, and surfaces the deferred set durably." + +4. Optional: L63, drop "at two altitudes"; L90, relabel D "(this spec's completion step)". + +Optional, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: modified, folded into Summary, Goals and Non-Goals, Design, Alternatives. Partially confirmed; this follows the narrowing. The two optional edits are taken, and the wording uses the glossary's set names, so the DKMS set (F04) is named beside the kernel set. +:EVIDENCE: +- Spec L28 (Summary): "This spec designs a safe path to actually complete a guarded upgrade, and makes that completion record the freshness stamp". It does not mention the split live run. +- Spec L43-46 (Goals): four bullets, all about completion, the boot path, or the installer. None covers the everyday live UPDATE. +- Spec L61: "...while the ordinary full sweep stays a normal live =topgrade= run." +- Spec L158 (Phase 2): "UPDATE and TOPGRADE levers change their =argv= to the script". +- Spec L65: the script runs pacman -Syu --ignore=..., then yay -Sua, then topgrade --disable system,git_repos,containers -y. +- git show 77447d0 has the identical Summary at draft L25, Goals at draft L40-43, and the L61 sentence at draft L58. Draft L60 is "Three pieces, at two altitudes." with nothing after it, but current L63 appends "— but the everyday gesture is not a reboot. It is a normal live update that simply leaves the dangerous few behind." So L63 has been edited and is not verbatim. +- Current L65/L67/L69 are three bold-labelled paragraphs (split script, user, implementer). +- Mitigations already in the spec: L13 (history line naming the live split upgrade as the everyday path), L51 and L55 (the split script named as v1), L167 (acceptance criterion for the live UPDATE). +- Labels: L90 "D. ... (this spec)" vs L95 "E. ... (this spec's everyday path)". +:END: + +** DONE ZFSBootMenu fallback leaves the separate pacman-db dataset out of sync +CLOSED: [2026-10-05 Mon 09:20] +Where: Risks L196; Acceptance L169 + +The spec treats booting the pre-pacman snapshot from ZFSBootMenu as a complete recovery for a failed kernel change on velox. + +Risk: After a fallback boot, /usr, /boot and the modules are back on the old kernel, but the pacman db still records the new kernel and zfs-dkms as installed. checkupdates then shows no pending kernel and the panel's deferred count drops it, so the machine silently runs a kernel the db doesn't know it has. + +Recommended change: Append one sentence to Risks L196, and carry it into the Phase 4 flow docs: "The pre-pacman snapshot covers zroot/ROOT/default only. The pacman db (zroot/var/lib/pacman) and the cache (zroot/var/cache) are separate datasets and stay current. So after booting the snapshot from ZFSBootMenu, re-register the old kernel set (and zfs-dkms, if that transaction moved it) with pacman -U --dbonly from /var/cache/pacman/pkg before the next upgrade-guarded run. Otherwise the db reports the new kernel as installed and the deferred row drops it." Do not offer "roll back zroot/var/lib/pacman to its matching snapshot" as a remedy: the hook never snapshots that dataset. A sturdier fix, outside this spec's scope, would have zfs-pre-snapshot snapshot both datasets atomically in one zfs snapshot call. + +Optional, not blocking. Verification: confirmed. +Disposition: modified, folded into Risks, Implementation phases. Confirmed. zfs-utils is named beside zfs-dkms because F04 moves them as a pair. +:EVIDENCE: +- Spec L196: "the 05-zfs-snapshot hook snapshots it before every transaction, and ZFSBootMenu can boot that snapshot. The pacman cache also keeps the previous kernel and zfs-dkms for a downgrade." Spec L134: "It is survivable — ZFSBootMenu can boot the pre-pacman snapshot". Spec L169 requires only that the snapshot is bootable. Spec L197: "The panel's deferred row is the reminder". Nothing in the spec mentions the db (grep for snapshot, database, var/lib, downgrade). +- scripts/zfs-pre-snapshot:11: DATASET="${ZFS_PRE_DATASET:-$POOL/ROOT/default}". Line 30: zfs snapshot "$DATASET@$SNAPSHOT_NAME", one dataset, no -r. +- ~/code/archangel/installer/archangel:736-738: zfs create -o mountpoint=/var/cache ...; zfs create -o mountpoint=/var/lib -o canmount=off ...; zfs create -o mountpoint=/var/lib/pacman "$POOL_NAME/var/lib/pacman". Present since archangel's initial commit 2b691a0. archangel .ai/notes.org:107 also documents zroot/var/lib/pacman as the "Package database" dataset. +- Velox was rebuilt from archangel on 2026-08-13/14 (.ai/sessions/2026-08-14-10-02-velox-mainboard-swap-reinstall-and-bringup.org:1079, "full reinstall via archangel+archsetup"). That makes the pre-reinstall claim at docs/design/2026-07-15-velox-boot-failure-handoff.org:21 and todo.org:1762 (no separate var/lib/pacman) stale. +- /etc/pacman.conf:13: #DBPath = /var/lib/pacman/ (the default is in effect). +- pacman -U --help lists "--dbonly only modify database entries, not package files". +- Weak point in the original evidence: archsetup:2269 is sanoid config, not proof the dataset exists. +:END: + +** DONE topgrade_run 'never written' is false on ratio +CLOSED: [2026-10-05 Mon 09:20] +Where: Problem L32 + +'Absent, the probe returns WARN ... the file is simply never written.' + +Risk: Cosmetic. The diagnosis still holds. + +Recommended change: Line 32: replace "the file is simply never written." with "the file is almost never refreshed: on ratio it was last written 2026-07-08, the day the wrapper landed, so the probe grades its age (88 days on 2026-10-05) against topgrade_warn_days = 14 and warns. The absent-file branch only fires on a machine that has never stamped." Line 34: change "It is never written because" to "It is rarely refreshed because". Line 36 (optional, same edit): change "The metric has read stale ever since." to "The metric had already gone stale in late July and stayed stale." Do not add a claim about velox, which cannot be checked while it is offline. + +Optional, not blocking. Verification: confirmed. +Disposition: accepted, folded into Problem. +:EVIDENCE: +- Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org:32: "...the file is simply never written." +- Spec line 34: "It is never written because topgrade rarely exits 0..." +- Spec line 36: "...so nothing stamped. The metric has read stale ever since." +- On ratio, ls ~/.local/state/maint/ shows topgrade_run.json (69 bytes, dated 8 Jul 07:46). Its contents are {"written_at": 1783514783.122665, ...}. date -d @1783514783 gives Wed Jul 8 07:46:23 AM CDT 2026. +- ~/.dotfiles/maint/src/maint/cache.py:35-42: get() returns (data, time.time() - written_at), or None only on OSError/ValueError/KeyError/TypeError. There is no TTL. +- ~/.dotfiles/maint/src/maint/probes/updates.py:114-131: the None branch is the "no topgrade run recorded" WARN. Otherwise it computes days = age // 86400 and calls grade_high(days, topgrade_warn_days). +- configs/maintenance-thresholds.toml:60 sets topgrade_warn_days = 14. +- maint status (live) prints "warn topgrade_age Topgrade freshness: 88". +- git log in ~/.dotfiles: commit 9b32a28 dated 2026-07-08 added hyprland/.local/bin/topgrade, the same day as the stamp. +:END: + +** DONE The autologin claim holds only on velox +CLOSED: [2026-10-05 Mon 09:20] +Where: Design L67, L69; Architecture L184 + +The boot unit runs 'before the autologin shell starts Hyprland'. + +Risk: Low. Phase 4's 'confirm the ratio path matches' would find a password prompt, not autologin. + +Recommended change: Describe the tty1 login without depending on autologin, and keep the Before=getty@tty1.service ordering exactly as it is. + +- L67: replace "before the autologin shell starts Hyprland" with "before tty1's login shell can start Hyprland". +- L69: replace "so it completes before autologin execs Hyprland" with "so it completes before the tty1 login (autologin where the installer enabled it, a password prompt otherwise), and so before ~/.profile.d/99-hyprland-autostart.sh can start Hyprland". +- L184: replace "getty autologin ordering" with "getty@tty1 login ordering". +- L192: replace "ahead of autologin" with "ahead of the tty1 login". + +Optional, not blocking. Verification: confirmed. +Disposition: accepted, folded into Design, Readiness dimensions, Risks. +:EVIDENCE: +Spec docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org: +- L67: "before the autologin shell starts Hyprland" +- L69: "ordered Before=getty@tty1.service so it completes before autologin execs Hyprland" +- L184: "getty autologin ordering" +- L192: "the unit sits ahead of autologin" +- L164: "Confirm the ratio path matches." + +Live on ratio (uname -n = ratio): +- /etc/systemd/system/getty@tty1.service.d/ does not exist. +- systemctl cat getty@tty1 shows the stock ExecStart=-/usr/bin/agetty --noreset --noclear - ${TERM}, with no --autologin. +- /etc/systemd/system/getty.target.wants/getty@tty1.service is the stock getty@.service. +- loginctl session 2 has TTY=tty1, Service=login, Type=wayland, Leader=1963. +- Process chain: 1963 "login -- cjennings" → 2012 "-zsh" → 2033 /usr/bin/start-hyprland → 2038 Hyprland. +- Hyprland is started by ~/.profile.d/99-hyprland-autostart.sh, which is guarded by [ "$XDG_VTNR" = "1" ] and runs start-hyprland. ~/.zshrc:10 sources ~/.profile, and ~/.profile loops over ~/.profile.d/*.sh. + +Installer archsetup configure_autologin (~L951-1004): +- Writes getty@tty1.service.d/autologin.conf only when enable_autologin=true, or when the root is encrypted and the user answers yes. +- A non-encrypted root returns 0 silently and writes no autologin. + +Velox autologin: could not be verified. The machine is offline, and no doc in the repo records it. +:END: + +** DONE Strip 'pending' count and QUEUE don't distinguish held packages +CLOSED: [2026-10-05 Mon 09:20] +Where: Phase 2 L158; AC L167 + +After a successful split run, the strip still counts the held set as pending, and QUEUE lists those packages without marking them. + +Risk: A successful UPDATE leaves 'N pending' on the strip next to 'N deferred', which reads as if the run failed. This is cosmetic, because updates_repo warns only at 50 or more. + +Recommended change: Append one sentence to Phase 2 (L158): "QUEUE (results wall and =maint queue=) tags any package in the deferred-set state file =[HELD]=, the strip's pending cell shows the held share (=N pending · M held=), and =cve_queued= excludes held packages so the CVE badge turns red only for advisories UPDATE can close. This is the 'run will defer' badge named in the script-ownership decision." No acceptance-criterion change is needed. Optionally extend L167 to read "...the panel shows the exact deferred set, marked held in QUEUE". + +Optional, not blocking. Verification: confirmed. +Disposition: modified, folded into Implementation phases, Acceptance criteria, Testing. Confirmed, with two changes. Its cve_queued clause is the same change as F36 and is carried there. And QUEUE's [HELD] tag is not Decision 3's 'run will defer' badge: that badge is F22's pre-run arm line, built from the TOML patterns, while [HELD] is a post-run mark built from the record. +:EVIDENCE: +- Spec L158 (Phase 2): the only panel change is a new probe plus the "N deferred — apply on reboot" row. L128: the deferred set is rendered "as its own state". L121: "the badge that says a run will defer", which no phase defines. L167 (AC): "the panel shows the exact deferred set". +- ~/.dotfiles/maint/src/maint/doctor.py:309-316: =if rid in ("update", "topgrade") and swap_ok:= calls =p_updates.scan_net(th)= to refresh the pending cache after a fire. +- probes/updates.py:150-158: scan_net writes checkupdates output to the updates_repo cache. --ignore'd packages are still listed by checkupdates. +- probes/updates.py:53-67: the repo_updates count is len(updates_repo), graded against pending_warn. archsetup configs/maintenance-thresholds.toml:58 sets =pending_warn = 50=. +- gui.py:447: the strip cell caption is ("updates_repo", "pending") with no held split. +- viewmodel.py:479-483: queue row state is only "cve", "reboot" (from _needs_reboot: _REBOOT_EXACT includes mesa, wayland, linux, linux-lts; prefix nvidia) or "note". cli.py:189: tags = {"cve": "[CVE] ", "reboot": "[BOOT] ", "note": " "}. gui.py:468-486 (_on_show_queue) uses the same rows. +- viewmodel.py:503-514 (cve_queued), whose docstring says the number turns red "only when running UPDATE would actually close an advisory". It counts every name in updates_repo/updates_aur, held ones included. It is used by gui.py:641-648 for both the faceplate badge and the strip. +:END: + +** DONE F05: gate clause (b) checks kernels that stage 1 just removed, so a kernel and DKMS change in the same run always fails the gate :blocking: +CLOSED: [2026-10-05 Mon 10:40] +Where (round 2): Design 'The gate' L137 (b) and L140 (per-kver 'pkgbase exists'); Phase 1 gate list L2147 (b) and L2151; Phase 1 test L2197; Testing map L2489. Compare (c) at L138/L2148, which has the existence filter that (b) is missing. + +Clause (b) reads: 'if any DKMS-set package's version changed, every for which the pre-run dkms status showed a module installed'. Unlike (c), it has no 'that still have /usr/lib/modules//vmlinuz' filter. The per-kver check then requires '/usr/lib/modules//pkgbase exists'. The only test for (b) covers a DKMS-only change (L2197). + +Risk: Every --complete that moves a kernel together with zfs-dkms fails the gate on ': pkgbase', even when the new kernel and its modules are fine. The run exits 4 and nothing is armed. The row goes CRIT with 'do not reboot' and REBOOT is hidden. That is a false alarm on the exact path the gate exists for. A second --complete clears it only because (c) drops the vanished kver. The Phase 1 tests don't catch it. + +Recommended change: L137 and L2147: replace (b) with "(b) if any DKMS-set package's version changed, every for which the pre-run dkms status showed a module installed and that still has /usr/lib/modules//vmlinuz after stage 1. A kernel whose version changed is covered by (a) under its new kver." + +L239 (Decision): change "or a DKMS-set package moved and that kernel had a module installed before the run" to "or a DKMS-set package moved and that kernel had a module installed before the run and is still installed after it". + +Add a test bullet after L2197 (Phase 1 Gate) and after L2489 (AC4 testing map): "one stage 1 moves a kernel-set package and zfs-dkms together. The old kver, which had zfs installed in the pre-run dkms status, is not on the check list. The new kver and every surviving kver with zfs installed are checked. With good fakes the gate passes." + +Blocking. Verification: confirmed. +Disposition: modified, folded into Implementation phases, Decisions, Design, Testing. Both sheets fix this at the root instead of adding R01's existence filter to clause (b). The gate entry is keyed by pkgbase and resolved to the kvers present after stage 1. A kver the run removed then can't reach the gate by any path, on the first run or a re-check. The finding's test is adopted as written. +:EVIDENCE: +Spec L137 and L2147: "(b) if any DKMS-set package's version changed, every for which the pre-run dkms status showed a module installed". There is no existence filter. Compare (c) at L138 and L2148: "...that still have /usr/lib/modules//vmlinuz". + +Spec L140 and L2151: the first per-kver requirement is "/usr/lib/modules//pkgbase exists". + +Spec L130 and L2085: the dkms status is saved before stage 1. L2086: stage 1 lands the kernel set, the DKMS set and every other pending package except the GPU closure. + +Only the DKMS-only test exists: L2197 and L2489. + +Live output on ratio: +- "pacman -Qqo" on /usr/lib/modules/7.2.7-arch1-1/pkgbase, /usr/lib/modules/7.2.7-arch1-1/vmlinuz, /usr/lib/modules/6.18.54-1-lts/pkgbase and /usr/lib/modules/6.18.54-1-lts/vmlinuz prints "linux linux linux-lts linux-lts". Each kernel package owns its kver's pkgbase, so an upgrade deletes it. +- "dkms status" prints "zfs/2.4.4, 6.18.54-1-lts, x86_64: installed" and "zfs/2.4.4, 7.2.7-arch1-1, x86_64: installed". +- "pacman -Q" shows linux 7.2.7.arch1-1 and linux-lts 6.18.54-1. The pkgrel is part of the kver. +- /usr/share/libalpm/hooks/70-dkms-upgrade.hook removes the old modules PreTransaction (When = PreTransaction, Exec = dkms -D remove). + +Worked example: linux 7.2.7.arch1-1 moves to 7.2.8.arch1-1 and zfs-dkms moves in the same stage 1. (a) adds 7.2.8-arch1-1. (b) adds 7.2.7-arch1-1 and 6.18.54-1-lts. 7.2.7-arch1-1/pkgbase no longer exists, so the gate fails with "7.2.7-arch1-1: pkgbase" and exits 4. +:END: + +** DONE F19: the required zfs.ko check runs lsinitcpio unprivileged against root-only 0600 initramfs images, so the gate fails on every ZFS-root --complete :blocking: +CLOSED: [2026-10-05 Mon 10:40] +Where (round 2): Design L135, L140; Phase 1 L2027, L2144, L2154; sudo lists L80, L2027, Readiness Security L2385 + +F19 made the lsinitcpio zfs.ko check required on a ZFS root (L140, L2154). kernel-modules-check is described as a stateless reader that --complete runs after stage 1. --complete runs as the invoking user and refuses EUID 0. Nothing gives the gate root: the sudo lists (L80, L2027, L2385) name only pacman -Sy/-Su/-S/-Sw, informant, the flag writes and systemctl reboot, and L2027 says ExecStartPre=+ is 'the only root step outside sudo'. + +Risk: On velox (ZFS root), the only machine where the gate matters most, every --complete with a non-empty check list fails the zfs.ko item. It exits 4 and writes an open gate failure; the row goes CRIT with REBOOT hidden, and nothing ever arms. Kernel and DKMS updates can never land through the designed path. The standalone 'no kver' TTY diagnostic fails the same way. + +Recommended change: Five edits: +- L140 and L2154: replace "lsinitcpio of that image lists a zfs.ko" with "sudo lsinitcpio of that image (mkinitcpio writes images root 0600) lists a zfs.ko". Add: "if lsinitcpio itself fails, the item is 'cannot read ', not 'zfs.ko missing'". Add sudo to the gate's faked-on-PATH list on L140. +- L80, L2027 and L2385: add "sudo lsinitcpio in kernel-modules-check" to the sudo enumerations. The gate's standalone TTY diagnostic gets the same NOPASSWD sudo. +- L2193 and L2480 (gate tests): add a case asserting the lsinitcpio call goes through the fake sudo, plus a case where an unreadable image fails as "cannot read", not as a missing zfs.ko. +- L2412 (rollout step a): add "on a ZFS root, confirm kernel-modules-check reads the real image without a permission error". Run it without --since, or with a since older than the image. +- Optional alternative: have --complete run "sudo kernel-modules-check ..." instead. Sudo-ing only lsinitcpio inside the gate keeps the standalone diagnostic working the same way, so I'd do that. + +Blocking. Verification: confirmed. +Disposition: accepted, folded into Implementation phases, Readiness dimensions, Testing. +:EVIDENCE: +- Spec L140 and L2154: "on a ZFS root (findmnt -no FSTYPE / is zfs), lsinitcpio of that image lists a zfs.ko". No sudo. +- Spec L80, L2027 and L2385: the full sudo list is pacman -Sy/-Su/-S/-Sw, informant read --all, the flag writes and systemctl reboot. L2027 adds: "The boot unit's ExecStartPre=+ is the only root step outside sudo." +- /usr/bin/mkinitcpio (pacman -Q: mkinitcpio 42.1-1) lines 1219-1220: "# Set umask to create initramfs images and unified kernel images as 600" / "umask 077". +- ls -la /boot on ratio: initramfs-linux.img, initramfs-linux-lts.img, initramfs-linux-lts-strix.img and initramfs-linux-lts-fallback.img are all ".rw------- root". The vmlinuz-* files are 0644. +- Running "lsinitcpio /boot/initramfs-linux.img" as uid=1000(cjennings) prints "==> ERROR: Unable to read file: '/boot/initramfs-linux.img'" and exits 1. +- These work unprivileged: "dkms status" lists zfs/2.4.4 entries, and "stat -c %Y /boot/initramfs-linux.img" returns the mtime. Only the content read fails. +- Spec L238: velox has /boot inside the encrypted ZFS root dataset. archsetup:3461 (tighten_efi_permissions) keeps boot images non-world-readable on purpose. A grep found no chmod or umask on initramfs anywhere in archsetup or scripts/. +- Tests fake lsinitcpio (L2163, L2193, L2454, L2480). Rollout step (a) only confirms "upgrade-guarded --dry-run exits 0" (L2412), and --dry-run uses no sudo and does not run the gate (L80, L2041). +- Velox was not checked directly because it is offline. The conclusion rests on the shared mkinitcpio umask. +:END: + +** DONE Nothing marks a --complete kernel transaction as unverified before the gate runs, so in-progress, interrupted and aborted stage-1 runs can reboot velox with no initramfs :blocking: +CLOSED: [2026-10-05 Mon 10:40] +Where (round 2): Design L128-133 (--complete order), L136 and Phase 1 L2146 (gate check list (a)), L168 and Phase 1 L2132 (interruption), Phase 1 L2083-2089, Phase 2 L2240 (REBOOT is state-driven), AC4 L2344 + +Stage 1 runs =sudo pacman -Su --noconfirm --ignore==. The gate runs afterwards 'whatever stage 1's exit code', and its check list (a) holds only 'each kernel-set package whose installed version changed in this run'. The INT/TERM/HUP trap 'waits for the running child to exit, then writes result interrupted with the step and the remedy, and exits 130', so on that path no gate runs. The record's gate field is written only when a gate fails. REBOOT is hidden only while gate.ok is false, shown while the flag exists, and otherwise follows reboot_required. A stale arm flag is removed only after the gate or at the GPU step. A later --complete with 'nothing pending and no gate failure open' prints 'nothing to complete' and rewrites the record empty. + +Risk: On velox (ZFS root, one kernel), three paths reach an unbootable machine with no 'do not reboot': +1. A normal stage 1: once 60-depmod removes the running kernel's module directory, the next 30 s panel refresh shows REBOOT. At that moment zfs is still building and /boot has no initramfs, because it was removed PreTransaction. +2. Closing the APPLY window or logging out (SIGHUP), or a Ctrl+C that reaches the script: pacman stops, PostTransaction hooks are skipped, and /boot/vmlinuz-linux-lts and its initramfs stay deleted. The script exits 130 at WARN with no gate. The next --complete either wipes the record ('nothing to complete'), or, with GPU entries still pending, gets an empty check list (no version changed in this run), passes and arms. It then offers a reboot onto a kernel that was never gated. +3. A commit abort, such as ENOSPC (the spec lists a full disk as a realistic trigger), before the kernel package extracts: the version is unchanged, so list (a) is empty, the gate is skipped, and the result is failed_step pacman at WARN while /boot is empty. + +Also, a flag left from an earlier arm keeps 'armed' and REBOOT visible through all of stage 1. Any of these leads to a ZFSBootMenu recovery. + +Recommended change: Smallest edit, in Design L128-139 and L168, Phase 1 L2087-2093, L2132 and L2146, Phase 2 L2240 and AC4 L2344: + +1. Stage 1 (L130 / L2087). When any kernel-set or DKMS-set member is pending, first =sudo rm -f= the arm flag. Then pre-write the record, in the same pattern as --apply-armed step 2: + - gate = {ok: false, pkgbases: [pkgbase of each pending kernel-set member; when a DKMS-set member is pending, every pkgbase the pre-run dkms status showed installed], failed: 'kernel transaction not yet verified', since: T0, at: now}; + - add pkgbases to the gate schema at L108 / L2127. + +2. Check list (a) (L136 / L2146). Replace "whose installed version changed in this run" with "each pkgbase in the open gate entry (or pending at stage-1 start), resolved to the kver whose /usr/lib/modules//pkgbase names it". Add to the per-kver requirements: /boot/vmlinuz- exists. Re-check (c) uses the recorded pkgbases. + +3. Trap (L168 / L2132). In --complete, the INT/TERM/HUP trap leaves the open gate entry in place. Its detail reads ' interrupted — do not reboot; run upgrade-guarded --complete'. Only a gate pass clears it, which keeps 'nothing to complete' from firing and a follow-up run from arming without a re-check. + +4. Tests (Phase 1): + - during a fake stage 1, the record has gate.ok false and no flag; + - an INT or HUP during stage 1 exits 130 with the gate still open and any earlier flag gone; + - a stage-1 failure that leaves the kernel version unchanged but /boot/vmlinuz- missing exits 4; + - a follow-up --complete after an interrupt re-checks and neither arms nor prompts until the check passes. + +5. AC4. Add: "From the start of a stage 1 that moves a kernel-set or DKMS-set package until a gate passes, including after an interrupted or failed stage 1, the deferred row is CRIT, REBOOT is hidden, no flag exists, and --complete never prompts for a reboot." + +Blocking. Verification: confirmed. +Disposition: modified, folded into Design, Decisions, Implementation phases, Acceptance criteria, Testing, Risks. I adopted the pre-written unverified state keyed by pkgbases, the /boot/vmlinuz requirement, the trap rule, the CRIT and REBOOT rules, the tests and the AC4 outcome, with three changes. First, any flag is removed once, right after every successful --complete refresh (R48). That makes 'no flag while the gate is open' an invariant rather than a list of cases. Second, the entry is checked in one invocation with the recorded since, not in a separate re-check. Third, R03's rule as written raises a false CRIT when stage 1 fails before committing anything, as on a download, signature or conflict error. /boot is then intact, no snapshot was taken and the initramfs predates T0, so the since items fail. The row would say 'do not reboot' on a machine that boots fine, and keep saying it until a later --complete succeeded. So when this run opened the entry and stage 1 changed no kernel-set or DKMS-set version, the gate runs its structural form without --since. A missing vmlinuz or initramfs, an unreadable image or a missing zfs.ko still fails it, which covers R03's ENOSPC case. +:EVIDENCE: +Spec: +- L130 / L2087-2089: stage 1 is a plain -Su with no prior state write. +- L131 / L2090: "The gate, whatever stage 1's exit code". +- L136 / L2146: "(a) for each kernel-set package whose installed version changed in this run". +- L139: "An empty list skips the gate". +- L129 / L2086: nothing pending and no gate failure open gives 'nothing to complete' and the record is rewritten empty. +- L133 / L2099: the prompt fires when "the run armed or changed a kernel-set or DKMS-set package". +- L168 / L2132: the trap "writes result interrupted ... exits 130", with no gate and no flag removal. +- L2117: flag removal points are a gate failure, unit not enabled, or nothing GPU-side left. +- L2240: REBOOT is hidden only while gate.ok is false and shown while the flag exists. +- L2344-2348: AC4 covers only a failed DKMS build. +- Phase 1 tests (L2183, L2196-2200) cover none of these cases. + +Live system: +- /usr/share/libalpm/hooks/60-mkinitcpio-remove.hook: Type=Path, Operation=Remove, Target=usr/lib/modules/*/vmlinuz, When=PreTransaction. Its script's remove_kernel runs =rm -f= on the preset's ALL_kver and images (/etc/mkinitcpio.d/linux-lts.preset: ALL_kver=/boot/vmlinuz-linux-lts, default_image=/boot/initramfs-linux-lts.img). +- man alpm-hooks CAVEATS: "PostTransaction hooks will not run if the transaction fails to complete for any reason." +- 70-dkms-install and 90-mkinitcpio-install are PostTransaction. 60-depmod is PostTransaction and its script does =rmdir --ignore-fail-on-non-empty= on a kernel dir with no modules.order. +- strings /usr/bin/pacman shows the "Interrupt signal received" and "Hangup signal received" fragments. libalpm.so.16 has "transaction interrupted", "transaction failed" and "problem occurred while upgrading %s". pacman is 7.1.0. +- archsetup:2622-2624: "a blocked transaction can remove the current initramfs without reaching the PostTransaction hook that rebuilds it." + +maint: +- probes/packages.py:131-146: required = not os.path.isdir(/usr/lib/modules/). +- panel.py:506-510: reboot_key_visible returns that value. +- gui.py:92 _FULL_SECONDS=30. gui.py:1376-1384: _full_tick refreshes unless self.firing. +- /usr/lib/modules// holds unowned modules.alias and related files, so the old dir outlives the package until 60-depmod. +:END: + +** DONE The gate's lsinitcpio check can't read the initramfs as the user, so the gate fails every time on a ZFS root :blocking: +CLOSED: [2026-10-05 Mon 10:40] +Where (round 2): Design L80 (sudo step list) and L140 (gate: lsinitcpio of /boot/initramfs-.img must list zfs.ko); Phase 1 L2027 (privilege list), L2154 (gate per-kver checks), L2163 and L2193 (tests fake lsinitcpio on PATH); Readiness Security L2385 + +upgrade-guarded runs as the invoking user and refuses EUID 0. The steps that go through sudo are listed exhaustively: pacman -Sy/-Su/-S/-Sw, informant read --all, writing and removing the flag, and systemctl reboot. kernel-modules-check is not on that list, so --complete (and the standalone TTY diagnostic) runs it unprivileged. On a ZFS root it requires =lsinitcpio= of /boot/initramfs-.img to list zfs.ko. + +Risk: On velox, every --complete whose check list is non-empty fails the gate at the zfs.ko item. That is every run that moves a kernel or zfs-dkms. It exits 4, records gate.ok false, and never arms. The CRIT row's remedy ('after fixing, run upgrade-guarded --complete') re-runs the same unreadable check, so the gate never closes, no form ever stamps, and the GPU/compositor set can only land by hand. The failure also blames a missing zfs.ko, which points me at a nonexistent DKMS problem on the do-not-reboot path. The standalone =kernel-modules-check= diagnostic fails the same way. This is the machine the gate exists to protect. + +Recommended change: Four edits to the spec. + +1. Phase 1 L2154 and the matching clause at Design L140: replace the zfs.ko bullet with: "on a ZFS root (findmnt -no FSTYPE / is zfs), sudo -n lsinitcpio of that image exits 0 and lists a zfs.ko with any compression suffix. The images are 0600 root (mkinitcpio sets umask 077), so the read goes through sudo. A non-zero lsinitcpio exit is its own failure item, ': cannot read /boot/initramfs-.img', never reported as a missing zfs.ko." + +2. Add "lsinitcpio, inside kernel-modules-check" to the sudo step lists at L80, L2027 and Security L2385. Say that the standalone diagnostic uses the same sudo read. + +3. Phase 1 tests (L2193) and Testing (L2477-2480): add two gate cases. + - The fake lsinitcpio fails, as the real one does on a 0600 image, unless the fake sudo invoked it. The gate must read through sudo and pass. + - lsinitcpio exits 1 with no output. The gate fails with the "cannot read" item, not "zfs.ko missing". + +4. In the zfs-VM / velox manual check, add a step confirming that a passing --complete actually read the image (the gate output names zfs.ko as found). + +Blocking. Verification: confirmed. +Disposition: accepted, folded into Implementation phases, Readiness dimensions, Testing. +:EVIDENCE: +Spec (docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org): +- L80, L2027, L2385: the sudo lists leave out kernel-modules-check and lsinitcpio. +- L131 and L2087: "run kernel-modules-check", with no sudo. +- L140 and L2154: "lsinitcpio of that image lists a zfs.ko with any compression suffix". +- L2163, L2193, L2454: lsinitcpio is faked on PATH. +- A grep for 0600, "Unable to read" and "sudo lsinitcpio" in the spec finds nothing. + +Live system (ratio, uid 1000): +- ls -l /boot shows all four initramfs-*.img as .rw------- root (linux, linux-lts, linux-lts-fallback, linux-lts-strix). vmlinuz-* are rw-r--r--. +- lsinitcpio /boot/initramfs-linux-lts.img prints "==> ERROR: Unable to read file: '/boot/initramfs-linux-lts.img'" and exits 1. +- /usr/bin/mkinitcpio:1219-1220 reads "# Set umask to create initramfs images and unified kernel images as 600" / "umask 077" (mkinitcpio 42.1-1). + +Repo knowledge of the velox case: +- docs/workflows/system-health-check.org:238: "The sudo is load-bearing: the images are 0600, so an unprivileged lsinitcpio exits 1 with 'Unable to read file' ... looks exactly like a missing module (velox, 2026-09-12)." +- todo.org:753-757: "Also for the kernel-modules-check gate: reading the initramfs needs root. ... The gate has to run as root and check the exit status, not just the count." +:END: + +** DONE F03: closure parser counts every error line, but each failing probe also prints lines that match neither form, so the literal rule always refuses +CLOSED: [2026-10-05 Mon 10:40] +Where (round 2): Design 'The dependency closure' L96-L100 ('parses every error line, with or without the leading ::' and 'refuses when a line matches neither form'); Phase 1 Dependency closure L2047-L2050; Phase 1 test L2169; Testing AC2 map L2469-L2474 + +The landed F03 text says the script 'parses every error line, with or without the leading ::'. Two shapes are classified (forward 'unable to satisfy dependency' and reverse 'installing X (v) breaks dependency'), and 'a line that matches neither form (a conflict, a missing target, a replace)' refuses with exit 3 and failed_step closure. The Phase 1 fake pacman emits only the two shapes. + +Risk: Under the literal contract, the error header alone matches neither form and refuses. The forward case's ':: ' skip-prompt lines also match neither form. So every probe that fails refuses, including the hyprutils/hyprlang and zfs-utils pin cases that AC2 and AC3 depend on, and the closure never grows. Every soname day becomes an exit-3 refusal. An implementer has to guess which lines to ignore. The Phase 1 fake emits only the two clean shapes, so the tests pass either way. + +Recommended change: L96-L100 and L2047-L2050: replace "parses every error line, with or without the leading ::" with this text: + +"reads the probe's stdout and parses each dependency line there, stripping an optional leading ':: '. Each line is classified by the two forms below. The probe's stderr preamble 'error: failed to prepare transaction (could not satisfy dependencies)' is expected on every failing iteration and is not classified. A preamble with any other reason in the parentheses refuses." + +Then change "a line matches neither form" to "a stdout dependency line matches neither form". + +L2169 and L2469: change the fake pacman so it replays pacman 7.1's verbatim failing -Sup transcript: the preamble on stderr, the "::" dependency lines on stdout, and exit 1. Add one assertion that a failing probe with the preamble and only well-formed lines adds names and does not refuse. + +Do not add the finding's ignore-list of "warning: ..." lines and skip-prompt lines. pacman suppresses all of them in print mode. + +Should-fix, not blocking. Verification: partially confirmed (the recommended change reflects the narrowed finding). +Disposition: accepted, folded into Implementation phases, Design, Testing. +:EVIDENCE: +Spec text: +- L96 and L2047: "parses every error line, with or without the leading ::". +- L100 and L2050: refuses "when a line matches neither form". +- L2169 and L2469: the fake pacman "emits both message shapes". No transcript, preamble or stream split is specified. +- grep for "failed to prepare|neither form|skip the above|cannot resolve" finds only L96, L100 and L2050. The preamble is never mentioned. + +Installed pacman: 7.1.0.r9.g54d9411-2 (libalpm 16.0.1). + +pacman source, callback.c:424-440 (pacman 7.1.0 source): in print mode, cb_question answers INSTALL_IGNOREPKG and REPLACE_PKG with 1, answers every other question with 0, and returns before the REMOVE_PKGS branch at L493-512 that prints the "cannot be upgraded" and "Do you want to skip" lines. + +Empirical runs used a throwaway fixture db, with no system db and no sudo. The command was: +LC_ALL=C pacman --config --dbpath -Sup --noconfirm --print-format '%n' ... + +- Forward case (--ignore=hyprutils, with hyprlang 2-1 depending on libhyprutils.so=13-64): exit 1. + - stdout: ":: unable to satisfy dependency 'libhyprutils.so=13-64' required by hyprlang" + - stderr: "error: failed to prepare transaction (could not satisfy dependencies)" + - No warning or skip-prompt lines, even though IgnorePkg=bridge-utils is pending and hyprutils is --ignore'd. +- Reverse/orphan case: exit 1. + - stdout: ":: installing libbar (2-1) breaks dependency 'libbar.so=1-64' required by orphan" + - stderr: the same preamble. +- Converged case: exit 0. stdout carries only the names; stderr is empty. +:END: + +** DONE F11: the timeout 60 hardening bounds only the explicit informant call, and the 00-informant hook repeats the same unbounded fetch inside the transaction +CLOSED: [2026-10-05 Mon 10:40] +Where (round 2): Design everyday step 3 L85; boot form step 4 L160; Phase 1 everyday step 3 L2063 and --apply-armed step 4 L2102; Decision 6 consequences L250; AC7 L2357-2358; Risks 'Arch news' L2448 (F11 disposition L682 states the intent: 'timeout 60 bounds it, and the run carries on either way') + +Most of the F11 resolution is in the body and correct: +- the call is timeout 60 sudo informant read --all, run only when command -v informant succeeds, with failure tolerated; +- yay -Pwq captures the titles first; +- the sudo/root reasons are stated; +- the offline fallback is described; +- AC7, the Readiness note (velox only) and the Phase 1 ordering test are all present. + +One gap remains. The body says the timeout 'bounds informant's feed fetch, which has no timeout of its own' and that 'failure, the timeout included, is tolerated' (L85, L2063). For --apply-armed it says 'the hook's informant check sees the same feed, so neither blocks the transaction' (L160, L2102). Risks L2448 says 'informant's AbortOnFail hook no longer stops a run'. Nothing covers what the hook does after the explicit call times out or fails. + +Risk: Take a slow or stalled archlinux.org fetch on velox while unread news exists. timeout 60 kills read --all before it saves, and the script carries on into the transaction. Then the 00-informant hook either: +- stalls the transaction with pacman holding db.lck. At boot that lasts until TimeoutStartSec=20min, ending as interrupted. From the panel it lasts until the 3600 s runner timeout. This is exactly the hang the timeout was added to prevent, moved one step later. +- or completes, sees one or more unread items, and exits non-zero, so AbortOnFail aborts the transaction. + +At boot the arm has already been consumed by ExecStartPre, so the GPU set doesn't land and the run records failed_step boot-transaction. On the everyday path it records failed_step pacman. + +On that path AC7 ('unread Arch news wedges no mode') fails, and L160/L2102 ('neither blocks the transaction') and L2448 ('no longer stops a run') are false. The failure is bounded and not unsafe: nothing is swapped, and pacman exits in the PREPARED state releasing db.lck. + +Recommended change: Smallest edit, option (b): name the residual and narrow the claims. + +1. L160 and L2102: replace "and the hook's informant check sees the same feed, so neither blocks the transaction" with: "When the fetch fails fast (no route or DNS), the hook's informant check also gets an empty feed and passes. When the clear times out or fails, nothing is marked read. Informant's hook then repeats the same unbounded fetch inside the transaction, so it can hold the transaction until TimeoutStartSec, or abort it on unread news (result failed, failed_step boot-transaction, arm consumed). The same abort can happen if the network comes up between the clear and the transaction." + +2. L85 and L2063: append: "A failed or timed-out clear leaves informant's hook live for this run's -Su, so the transaction can wait on the same fetch (bounded only by the lever runner's 3600 s) or abort on unread news (failed_step pacman, before any package changes)." + +3. L2448: change "no longer stops a run" to "no longer stops a run once the clear succeeds; a failed or timed-out clear leaves the hook able to stall or abort the transaction, and the record names it". + +4. AC7 (L2357): change to "Unread Arch news wedges no mode once the clear succeeds; a failed clear ends in a named failed_step, never a silent hang past the unit or runner timeout." + +5. Phase 1 tests (L2175 or L2499): add a case where a fake informant's read --all exits 124 and the fake pacman then fails as the hook abort would. Assert failed_step pacman (everyday) or boot-transaction (--apply-armed), no stamp, and for --apply-armed no re-arm. + +Option (a), if AC7 must hold unconditionally: when read --all exits non-zero, run that run's own -Su/-S with "--hookdir /etc/pacman.d/hooks --hookdir ". The per-run dir holds 00-informant.hook -> /dev/null, which keeps the guard hook active. In the everyday and --complete forms, do this only when yay -Pwq succeeded, so the news is never cleared unseen. Assert the extra --hookdir argv in the same new test. + +Should-fix, not blocking. Verification: confirmed. +Disposition: accepted, folded into Implementation phases, Acceptance criteria, Risks, Testing. +:EVIDENCE: +Spec claims and gaps (docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org): +- L85 and L2063: "The timeout bounds informant's feed fetch, which has no timeout of its own. Failure, the timeout included, is tolerated and logged." +- L160: "offline-safe: informant 0.6.0 falls back to an empty feed and exits 0, and the hook's informant check sees the same feed, so neither blocks the transaction." L2102 makes the same claim. +- L2357: "AC7: Unread Arch news wedges no mode." L2448: "on velox informant's AbortOnFail hook no longer stops a run". +- L2175 and L2499 test only "a failing or absent informant is tolerated". There is no case where the clear fails and the hook then fires. +- grep for hookdir, 00-informant and "informant check" finds no handling elsewhere. + +informant 0.6.0-2 (archangel airootfs, usr/lib/python3.14/site-packages/informant/): +- feed.py:81-82: session.get(self.url) with no timeout. feed.py:83-86: on an exception it falls back to feedparser.parse(url), also with no timeout. feed.py:90-106: only a bozo URLError gives an empty feed. +- informant.py:117-119: --all marks entries read in memory. informant.py:146: fs.save_datfile() runs after the fetch. +- informant.py:165-167: an empty feed gives sys.exit() (0) with nothing marked. +- informant.py:83-98: check marks and saves one unread item, but exits with the unread count either way. +- usr/share/libalpm/hooks/00-informant.hook: Operation=Upgrade, Target=*, Target=!informant, When=PreTransaction, Exec=/usr/bin/informant check, AbortOnFail. + +Live system checks: +- man alpm-hooks lists no timeout option, and says "Hooks may be disabled by overriding them with a symlink to /dev/null". +- man pacman says --hookdir is "a alternative directory ... hooks in later directories taking precedence". +- /etc/pacman.d/hooks holds hypr-live-update-guard.hook. +- sysctl net.ipv4.tcp_syn_retries = 6. +- archsetup:1801-1802 enables NetworkManager. The boot unit has no network ordering (spec L153). +:END: + +** DONE F22: dropping guard: live_update also removes the only trigger for the arm-line annotation, _rearm_after_guard and the doctor review suffix that Phase 2 says to reword +CLOSED: [2026-10-05 Mon 10:40] +Where (round 2): Phase 2 Guard UX L2219-2220; Phase 2 tests L2262; Testing L2467 + +L2219 drops guard: "live_update" from the update and topgrade remedies, along with _update_force and maint fix --force. L2220 then says guard.trips survives as the arm-line annotation, and that on a tripped read arm_line, _rearm_after_guard's text and the doctor review suffix read 'UPDATE armed — will defer at least ... — press again to run upgrade-guarded'. + +Risk: Implemented literally, the annotation never shows: guarded() is false, and no guard event is ever emitted. The implementer has to invent a new selector, such as a new remedy key or an rid check, or else keep rewording code that can no longer run. The L2262 test can pass against the pure arm_line function while the panel never calls it. The fixed text also says 'UPDATE armed' on the TOPGRADE key. + +Recommended change: In Phase 2 Guard UX, replace the second sentence of L2220 with this: + +"Dropping the tag removes the only selector, so panel.guarded() and doctor.py's review-suffix check key on rid in ('update', 'topgrade') instead, and test_panel_phase10.py:420-421 keeps asserting True for both. iter_fix no longer emits a 'guard' event, so delete _rearm_after_guard and the guard-event branch of _on_fired. Do not reword them. _guard_arm_line stops setting _update_force. On a tripped read, arm_line and the review suffix read '