aboutsummaryrefslogtreecommitdiff
path: root/docs
diff options
context:
space:
mode:
Diffstat (limited to 'docs')
-rw-r--r--docs/2026-08-13-velox-reinstall-runbook.org184
-rw-r--r--docs/2026-08-15-velox-uefi-boot-entry-reference.org108
-rw-r--r--docs/design/2026-07-10-net-bt-failure-taxonomy.org4
-rw-r--r--docs/design/2026-07-15-velox-boot-failure-handoff.org61
-rw-r--r--docs/design/2026-08-14-velox-reinstall-gaps-1.org137
-rw-r--r--docs/design/2026-08-14-velox-reinstall-gaps-2.org63
-rw-r--r--docs/design/2026-08-14-velox-reinstall-gaps-3.org98
-rw-r--r--docs/post-install-checklist.org39
-rw-r--r--docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org213
-rw-r--r--docs/workflows/system-health-check.org35
10 files changed, 937 insertions, 5 deletions
diff --git a/docs/2026-08-13-velox-reinstall-runbook.org b/docs/2026-08-13-velox-reinstall-runbook.org
new file mode 100644
index 0000000..2d99fc9
--- /dev/null
+++ b/docs/2026-08-13-velox-reinstall-runbook.org
@@ -0,0 +1,184 @@
+#+TITLE: Velox Reinstall Runbook — DR Test of archangel + archsetup
+#+AUTHOR: Craig Jennings
+#+DATE: 2026-08-13
+
+Context: velox's mainboard swapped Intel → AMD (Ryzen AI 9 HX 370, Radeon
+890M, 96GB RAM). Old SSD intact but the new board's NVRAM has no boot entry,
+and velox is ZFS root + ZFSBootMenu, so a stock Arch USB can't even read the
+pool. Decision: full reinstall via archangel + archsetup, run deliberately as
+a disaster-recovery test of the ISO and scripts before the Sunday flight.
+Recent backup in hand; ratio available as the working machine.
+
+* Outcome (recorded 2026-09-13 Sun)
+
+The drill ran on 2026-08-13 and 14 and velox came back as a working daily
+driver: fresh install from the archangel ISO, keys and data restored from
+the salvage backup, 23 repos re-cloned, rsyncshot reinstalled, hibernate
+proven end to end. The checklist below was the live plan; it was not ticked
+as the phases ran, so read it as the plan, not a log of each step.
+
+What the drill found, each filed as its own task rather than fixed in place:
+velox's truenas backups had silently stopped on 2026-07-06 (found on 08-13
+before partitioning, which is what made the salvage pass required); four phantom
+reboots were a ribbon disturbed by the board swap; a fresh install never
+clones rulesets, never links the .emacs.d systemd user units, ships no
+brightness udev rule, and loses gcalcli and the signal-cli registration.
+Those live in archsetup's todo as the post-rebuild verification pass and its
+siblings. The one commit that existed only on the old disk (emacs-wttrin
+bf0457f) was rescued as a bundle and has its own task.
+
+Companion documents: the UEFI boot-entry recovery reference
+([[file:2026-08-15-velox-uefi-boot-entry-reference.org][2026-08-15-velox-uefi-boot-entry-reference.org]])
+and the three gap reports under docs/design (2026-08-14-velox-reinstall-gaps-1
+to 3).
+
+Fallback ordering if the test finds a real gap:
+- Before partitioning starts: the old system is intact — the ZBM repair
+ route (efibootmgr entry pointing at the ZBM loader on the ESP, then
+ amd-ucode swap in a chroot) is still available.
+- After partitioning: the floor is a manual Arch install; the backup makes
+ that survivable.
+
+* Phase 0 — Preflight on ratio (agent-driven, done before you leave the desk)
+
+- [ ] Rebuild the ISO with archsetup baked in: the 2026-08-02 ISO predates
+ the microcode vendor-detection fix (archsetup, 2026-08-08) and was built
+ without ARCHSETUP_DIR at all.
+ #+begin_src sh
+ cd ~/code/archangel && sudo ARCHSETUP_DIR="$HOME/code/archsetup" ./build.sh
+ #+end_src
+ Two traps in that one line, and either alone silently produces a bare ISO
+ with archsetup absent (archangel, 2026-08-20). The =VAR=value= form is
+ required because sudo's =env_reset= discards an exported variable. And
+ =$HOME= is required because zsh does not expand a tilde on the right-hand
+ side of an assignment — the earlier =ARCHSETUP_DIR=~/code/archsetup= here
+ passed the literal string. build.sh now warns and reports baked/not-baked
+ in its closing summary, so the failure is visible rather than silent.
+- [ ] build.sh fixes before the final rebuild (archangel repo):
+ - rsync exclude for =.ai= (keeps =archsetup/.ai/private-design/= — the
+ credential audit — off the portable USB stick).
+ - copy =installer/velox-*.conf= to =airootfs/root/= so the machine profile
+ is on the ISO at =/root/velox-zfs.conf=.
+- [ ] Verify the ISO carries: =/code/archsetup= (with =install_cpu_microcode=),
+ =/root/velox-zfs.conf=, no =.ai/private-design=. Loop-mount or unsquashfs
+ spot-check.
+- [X] USB ready (done 2026-08-13 15:25): the new ISO was copied to the Ventoy
+ drive, sha256-verified against the source, and the 2026-04-09 + 2026-06-16
+ archangel ISOs removed. Boot the stick and pick
+ =archangel-2026-08-13-vmlinuz-6.18.43-lts-x86_64.iso= from the Ventoy menu.
+
+* Phase 1 — UEFI setup on velox (BIOS screen, before any boot)
+
+- [ ] Disable Secure Boot. Mandatory — the ZFS kernel modules are unsigned;
+ the new board ships with it enforced by factory default.
+- [ ] Set the system clock. The board swap reset the RTC to 2025-01-01;
+ a wrong clock breaks TLS and pacman signature checks in the live env.
+ Rough accuracy is fine — NTP tightens it once networked.
+- [ ] While you're in setup: check boot-order UI shows the USB.
+
+* Phase 2 — Salvage pass (live ISO, BEFORE running the installer) — REQUIRED
+
+NOT optional insurance. Verified 2026-08-13: velox's newest truenas backup is
+DAILY.0 = 2026-07-06 — five weeks stale. The backup timer on velox broke
+around Jul 6 (truenas itself only went dark Jul 24, and it's back now; ratio
+and mybitch backed up today). Everything since Jul 6 exists only on the old
+SSD — including =wolf.conf.gpg= (created Jul 29), which is therefore in NO
+backup at all. This pass also keeps the repair fallback alive until
+partitioning starts.
+
+- [ ] Network up (=nmtui= or ethernet), then confirm clock: =timedatectl=.
+- [ ] Import the old pool read-only and unlock:
+ #+begin_src sh
+ zpool import -N -o readonly=on -R /mnt zroot
+ zfs load-key zroot # passphrase prompt
+ zfs mount zroot/ROOT/default
+ zfs mount -a 2>/dev/null # home datasets etc.; ignore failures
+ #+end_src
+- [ ] Push a full fresh backup to truenas over the LAN — mirror the layout
+ the backup job uses (etc + home), into a clearly-named one-off dir:
+ #+begin_src sh
+ rsync -aHAX --info=progress2 /mnt/etc /mnt/home \
+ truenas:/mnt/vault/backups/velox/pre-reinstall-2026-08-13/
+ #+end_src
+ (=/usr= is in the regular backups but is all reinstallable — skip unless
+ paranoid. The 96GB-RAM board will not be the bottleneck; the LAN is.)
+- [ ] Spot-check the copy landed: =wolf.conf.gpg=, =.ssh=, =.gnupg=, newest
+ files in =~/documents= and =~/downloads=.
+- [ ] Check for uncommitted repo work and either push or note it:
+ =~/.emacs.d= (known: the auto-dim-other-buffers.el unresolved merge),
+ =~/.dotfiles=, anything under =~/code=.
+- [ ] Export cleanly: =cd /; zfs unmount -a; zpool export zroot=.
+
+* Phase 3 — Install (the actual DR test)
+
+- [ ] Review the profile, then run the installer:
+ #+begin_src sh
+ less /root/velox-zfs.conf # FILESYSTEM=zfs, HOSTNAME=velox, single nvme
+ archangel --config-file /root/velox-zfs.conf
+ #+end_src
+ Note: the profile's ZFS_PASSPHRASE / ROOT_PASSWORD are the =welcome=
+ placeholders — fine for install; both change post-install (=zfs change-key
+ zroot= for the pool, =passwd= for root).
+- [ ] Record every rough edge as a DR-test finding — that's the point of
+ running it this way. Anything that needs a manual nudge gets a todo entry
+ in archangel or archsetup afterward.
+- [ ] Reboot into ZBM → boot the new environment.
+
+* Phase 4 — archsetup (first boot of the installed system)
+
+- [ ] Log in as root, network up, then verify the clock synced.
+- [ ] Get archsetup — two paths, test the offline one since this is a DR
+ drill (the online curl path is the everyday alternative):
+ #+begin_src sh
+ # offline: mount the install USB and copy the baked tree
+ mount /dev/disk/by-label/ARCHANGEL* /mnt 2>/dev/null || mount /dev/sdX1 /mnt
+ cp -r /mnt/code/archsetup /root/archsetup && cd /root/archsetup
+ ./archsetup
+ #+end_src
+- [ ] Expected on the new board: =install_cpu_microcode= detects
+ AuthenticAMD and installs amd-ucode (verified 2026-08-13, 7/7 tests).
+ Podman socket, camera udev rule, tlp radio state, ZFS /tmp mask are all
+ in the installer now — none need manual application afterward.
+- [ ] archsetup clones + stows dotfiles. The velox host tier has no Intel
+ assumptions (swept 2026-08-13); maint's capability probe runtime-detects
+ amd-pstate.
+
+* Phase 5 — Post-install restore + verification
+
+- [ ] Restore from backup (credentials, ssh keys, gpg, user data). The
+ secrets-bundle-in-ISO design is not built yet — manual restore is the
+ known gap, not a test failure.
+- [ ] WireGuard: decrypt + re-place =wolf.conf.gpg= at =~/.config/wireguard/=;
+ re-import the NM profile (autoconnect off, as before).
+- [ ] Change the placeholder passwords: =passwd=, =zfs change-key zroot=.
+- [ ] PSR workaround — REQUIRED on this board. The Ryzen AI 300 has a known
+ idle instability (Panel Self Refresh hangs/reboots the machine; hit during
+ the live session 2026-08-13). Add =amdgpu.dcdebugmask=0x610= to the
+ installed system's kernel command line — velox boots via ZBM, so set it on
+ the pool: =zfs set org.zfsbootmenu:commandline="... amdgpu.dcdebugmask=0x610" zroot/ROOT/default=
+ (keep the existing args; append). Revisit after a BIOS update ≥3.05 or a
+ kernel that fixes PSR on Strix Point — track via the Framework issue
+ tracker (SoftwareFirmwareIssueTracker #110).
+- [ ] New-hardware spot-checks:
+ - =journalctl -k | grep -i microcode= — amd-ucode applied.
+ - =cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_driver= — expect
+ amd-pstate(-epp).
+ - wifi + bluetooth up (new board radios), touchpad behavior, camera.
+ - =glxinfo -B= / =vulkaninfo --summary= — Radeon 890M on RADV.
+- [ ] Fresh clones automatically carry the post-purge rewritten git history —
+ closes the clone-reconcile rider from 2026-08-11 without action.
+- [ ] Fix and verify the backup timer on the fresh install — it was silently
+ broken since ~Jul 6. After the first manual run succeeds, confirm a new
+ DAILY.0 appears under =truenas:/mnt/vault/backups/velox/=. Diagnose why it
+ broke (timer unit dead? mount failure? credential?) if the old journal
+ survives in the salvage copy.
+- [ ] Update the machine-identity memory: velox is now AMD (amd-pstate),
+ both daily drivers AMD. Fix the stale =intel_pstate= comment in
+ =airplane-mode= line 6 while at it (cosmetic).
+- [ ] File every DR-test finding in the owning project's todo.
+
+* Timing
+
+Today is Thursday; the flight is Sunday. Target: Phases 0–4 tonight or
+Friday, leaving Saturday as pure buffer. If the install stalls past Friday
+evening, cut losses to the manual-install floor.
diff --git a/docs/2026-08-15-velox-uefi-boot-entry-reference.org b/docs/2026-08-15-velox-uefi-boot-entry-reference.org
new file mode 100644
index 0000000..1eacfe5
--- /dev/null
+++ b/docs/2026-08-15-velox-uefi-boot-entry-reference.org
@@ -0,0 +1,108 @@
+#+TITLE: Velox UEFI Boot Entry — Recovery Reference
+#+AUTHOR: Craig Jennings
+#+DATE: 2026-08-15
+
+Captured 2026-08-15 before a BIOS update (03.05 → 04.02) as insurance against
+the update clearing NVRAM. Velox's mainboard swap on 2026-08-13 left exactly
+this kind of empty NVRAM, which is what forced the reinstall — so a cleared
+boot entry is the specific failure worth being able to undo in one command
+rather than reconstruct.
+
+* Before recreating anything: check Secure Boot first
+
+The 04.02 update landed on 2026-09-12 (staged via fwupd, flashed on the next
+reboot). It did NOT clear NVRAM: Boot0001 survived with its command line
+intact. What it did was re-enable Secure Boot (Enforce Secure Boot =
+Enabled), so the unsigned ZBM loader was rejected and the Framework BIOS
+reported it as "Default Boot Device Missing / no bootable drive" rather than
+a security violation. The symptom is indistinguishable from the NVRAM wipe
+this document was written for.
+
+The tell: booting the Ventoy stick shows shim's MOK management screen.
+
+Fix: F2 → Security → Secure Boot → Enforce Secure Boot = Disabled, F10. The
+machine then boots straight into ZBM. No efibootmgr needed.
+
+So when velox says no bootable device after a firmware update, check Secure
+Boot before touching the boot entries. Only if Secure Boot is already off
+and =efibootmgr -v= (from the stick) shows Boot0001 gone does the recreate
+below apply.
+
+Before any future firmware update, record both =efibootmgr -v= and the
+Secure Boot state so the post-reboot diagnosis is a comparison, not a guess.
+The pre-firmware-update checklist in
+[[file:workflows/system-health-check.org][docs/workflows/system-health-check.org]]
+(Phase 3) carries the steps.
+
+Still open: velox's ESP has no removable-media fallback (=/efi/EFI/BOOT=
+does not exist), so a real NVRAM wipe would still need the stick. Copying
+=zfsbootmenu.efi= to =/efi/EFI/BOOT/BOOTX64.EFI= would let it boot unaided.
+That belongs to archangel's ZBM install, where it is filed as [#C]
+"Installed systems have no removable-media boot fallback on the ESP"
+(2026-09-12) and ships with the next ISO rebuild after it lands.
+
+* State at capture
+
+- BIOS: 03.05 (2025-10-30)
+- BootCurrent: 0001
+- BootOrder: 2001,0001,2002,2003 (USB ahead of ZBM — why the Ventoy stick
+ boots when it's inserted)
+- Timeout: 0 seconds
+
+* The entry that matters
+
+=Boot0001* ZFSBootMenu=
+
+| field | value |
+|----------------+----------------------------------------------|
+| ESP part GUID | 8e51b680-f90a-444f-8da5-7e4f93625775 |
+|----------------+----------------------------------------------|
+| partition | 1 (GPT), start 0x800, size 0x100000 |
+|----------------+----------------------------------------------|
+| loader path | =\EFI\ZBM\zfsbootmenu.efi= |
+|----------------+----------------------------------------------|
+| cmdline (data) | =spl_hostid=0x22f8a7a1 zbm.timeout=3= |
+| | =zbm.prefer=zroot zbm.import_policy=hostid= |
+|----------------+----------------------------------------------|
+
+The =data= field is that command line in UTF-16LE, which is how efibootmgr
+passes it as optional data. Recreate with =-u= and the plain string; efibootmgr
+does the encoding.
+
+* Recreating it
+
+From a booted system (or the archangel ISO), with the ESP identified as
+=/dev/nvme0n1p1= or whatever it enumerates as:
+
+#+begin_src bash
+efibootmgr --create \
+ --disk /dev/nvme0n1 --part 1 \
+ --label "ZFSBootMenu" \
+ --loader '\EFI\ZBM\zfsbootmenu.efi' \
+ --unicode 'spl_hostid=0x22f8a7a1 zbm.timeout=3 zbm.prefer=zroot zbm.import_policy=hostid'
+#+end_src
+
+Confirm the disk/part against =lsblk -o NAME,PARTUUID,PARTTYPENAME= first —
+the partition GUID above is the authoritative identifier, not the device name,
+which can enumerate differently.
+
+Then set the order so ZBM is reachable:
+
+#+begin_src bash
+efibootmgr --bootorder 0001,2001,2002,2003
+#+end_src
+
+(The original order put USB first. Keep whichever you prefer; what matters is
+that the ZBM entry exists and is in the list.)
+
+* Other entries (firmware-generated, recreate themselves)
+
+| Boot2001 | EFI USB Device |
+|----------+----------------|
+| Boot2002 | EFI DVD/CDROM |
+|----------+----------------|
+| Boot2003 | EFI Network |
+|----------+----------------|
+
+These are stock firmware entries and come back on their own. Only Boot0001
+carries anything unique.
diff --git a/docs/design/2026-07-10-net-bt-failure-taxonomy.org b/docs/design/2026-07-10-net-bt-failure-taxonomy.org
index 74790c6..4d57b86 100644
--- a/docs/design/2026-07-10-net-bt-failure-taxonomy.org
+++ b/docs/design/2026-07-10-net-bt-failure-taxonomy.org
@@ -96,6 +96,7 @@ Six layers, mirroring the net doctor's probe ladder (link → IP/DHCP → gatewa
- Another daemon overwrites resolv.conf (yes). DNS works then breaks (or breaks after VPN up/down) as dhcpcd/openvpn/openresolv rewrites resolv.conf. Multiple tools claim it with no coordination. Fix: pick one manager (openresolv =resolvconf=NO=, dhcpcd =nohook resolv.conf=), point resolv.conf at the stub, restart resolved. [[https://github.com/adrienverge/openfortivpn/issues/674][openfortivpn 674]]
- nsswitch.conf hosts line broken (yes). All resolution fails, or LAN/mDNS names never resolve; the hosts line lacks =resolve=/=dns= in the right order or references an uninstalled nss module. Fix: set =hosts: mymachines resolve [!UNAVAIL=return] files myhostname dns=. [[https://man.archlinux.org/man/nss-resolve.8.en][nss-resolve]]
- Avahi/.local mDNS not resolving (yes). *.local names don't resolve though unicast DNS works. nss-mdns not wired in, or resolved's built-in mDNS collides with avahi. Fix: install nss-mdns, add =mdns_minimal [NOTFOUND=return]= before =resolve=, enable avahi-daemon, disable resolved MulticastDNS if both run. [[https://wiki.archlinux.org/title/Avahi][archwiki avahi]]
+- Clock skew breaks DNS itself, and NTP cannot recover it (yes; field-observed 2026-08-19, velox, not from the 2026-07-10 sweep). Nothing resolves at all — not a slow lookup, a dead one — after a boot with a wrong clock. =DNSSEC=yes= validates RRSIG inception/expiry windows against the wall clock, so a clock weeks off fails every query before it leaves the machine. Measured on velox 2026-08-19 with the clock wound back 27 days: resolved logged =signature-expired= against the root DNSKEY and every DS beneath it, and resolution died outright. =DNSOverTLS=yes= is *not* what bites, despite being the obvious suspect — the DoT handshake to =1.1.1.1:853= verified clean at that same clock, because a resolver certificate is good for about a year while an RRSIG window is days to weeks. A skew large enough to break DNSSEC normally leaves the certificate valid. The trap is the recovery path: NTP daemons name their servers by hostname (=pool 2.arch.pool.ntp.org=, =NTP=time.cloudflare.com=), so the daemon that would fix the clock needs the DNS that the clock is breaking. Neither side moves and the machine cannot self-heal — diagnosis needs a second device. Distinguish from the plain clock-skew entry in the egress layer by where it bites: that one has working DNS and failing HTTPS, this one has no DNS at all. Confirm with =dig @1.1.1.1 example.com +short=, which goes out plain UDP/53 and bypasses resolved entirely; an answer there with resolved still failing puts the fault in the validation layer, not the network. Fix: set the clock by hand (=timedatectl set-time=), then =resolvectl flush-caches=. =DNSSEC=allow-downgrade= does *not* help here, which is worth knowing because it is the obvious reach: resolved downgrades when a server lacks DNSSEC support, and a signature-window failure is a validation failure rather than a support failure, so no downgrade fires. Measured on velox: six retries over eighteen seconds, plus =resolvectl reset-server-features=, all dead. The only cure is correcting the clock, which is why the NTP source has to be reachable without DNS. Prevent by giving the NTP daemon at least one source addressed by IP, which needs neither DNS nor a certificate — =server 162.159.200.1 iburst= in a chrony drop-in. Note =timedatectl set-ntp true= is *not* a fix here: it starts a daemon that still cannot resolve its pool.
** Egress / captive portal / MTU / proxy / clock / upstream
@@ -107,7 +108,7 @@ Six layers, mirroring the net doctor's probe ladder (link → IP/DHCP → gatewa
- PPPoE / VPN link with a lower MTU not clamped (no). Browsing works but big transfers / some HTTPS hang. A PPPoE (1492) or VPN path has a smaller MTU and the too-large segments get dropped. Fix: set the tunnel/link MTU down (=.mtu 1420= for VPN, 1492 for PPPoE) or MSS-clamp on the gateway. [[https://thelineman.ca/articles/article-8-mtu-vpn-mss][vpn mtu/mss]]
- Stale http_proxy env var points at a dead proxy (no). Every curl/wget/pacman fails though the network is fine; browsers may work. A leftover =http_proxy= points at an offline/off-network proxy. Fix: unset the vars, remove the export from =~/.profile= / =/etc/environment=. [[https://everything.curl.dev/usingcurl/proxies/env.html][curl proxy env]]
- Unreachable PAC file off the corporate network hangs everything (no). Away from the office the browser stalls with no error. A system proxy set to "automatic" with a PAC URL that only resolves on the corporate LAN blocks waiting instead of falling back to DIRECT. Fix: switch system proxy to None (=gsettings … org.gnome.system.proxy mode 'none'=) or clear the PAC URL. [[https://bugzilla.mozilla.org/show_bug.cgi?id=1121800][ff pac hang]]
-- Clock skew breaks every TLS handshake (yes). "Your connection is not private" on every HTTPS site though ping/DNS work; the clock is hours/years off. A dual-boot Windows RTC-localtime, unsynced NTP, or a dead CMOS battery leaves the clock wrong. Fix: =timedatectl set-ntp true= (=set-local-rtc 0= on dual-boot), replace the CMOS battery if it recurs. [[https://wiki.archlinux.org/title/System_time][archwiki system time]]
+- Clock skew breaks every TLS handshake (yes). "Your connection is not private" on every HTTPS site though ping/DNS work; the clock is hours/years off. A dual-boot Windows RTC-localtime, unsynced NTP, or a dead CMOS battery leaves the clock wrong. Fix: =timedatectl set-ntp true= (=set-local-rtc 0= on dual-boot), replace the CMOS battery if it recurs. This entry assumes DNS still works; when the resolver runs DoT or DNSSEC the same skew kills DNS first and =set-ntp true= cannot recover it — see the clock/DNS deadlock in the DNS layer. [[https://wiki.archlinux.org/title/System_time][archwiki system time]]
- Firewall default-deny drops all egress (yes). No traffic leaves right after enabling a firewall, or after both ufw and firewalld are on; even DNS fails. A default outgoing-deny policy, or two firewalls fighting over nftables. Fix: allow egress (=ufw default allow outgoing=) and run only one firewall. [[https://wiki.archlinux.org/title/Uncomplicated_Firewall][archwiki ufw]]
- VPN kill-switch / leftover iptables rule strangles egress after VPN drops (yes; distinct from the route-capture case). Internet dies the moment the VPN disconnects and never returns until reboot. A kill-switch rule pinned traffic to tun0 and the leftover rule keeps dropping everything on the real interface. Fix: flush the stale rules (=iptables -F; iptables -P OUTPUT ACCEPT=, or restart the firewall), reconnect. [[https://bbs.archlinux.org/viewtopic.php?id=300104][arch ufw killswitch]]
- IPv6 egress broken while IPv4 works (no; the egress angle of the broken-v6 family). Pages load slowly/intermittently; IPv4-only hosts are fine. The network advertises IPv6 with no working route and Happy Eyeballs keeps trying the dead AAAA path. Fix: =nmcli con modify <con> ipv6.method disabled= until the network's IPv6 is fixed. [[https://help.ubuntu.com/community/WebBrowsingSlowIPv6IPv4][ubuntu slow ipv6]]
@@ -309,6 +310,7 @@ Probe: dns-config + resolver-health + dns-resolve + the doctor's dns-test (which
- VPN split-DNS not applied :: AUTO — =resolvectl domain/default-route= on the VPN link.
- IPv6 AAAA lookups stall :: AUTO — disable IPv6 on the link (or the single-request option). Also cluster 8.
- Another daemon overwrites resolv.conf :: PRIV — pick one manager, point resolv.conf at the stub.
+- Clock skew breaks DNSSEC validation, NTP deadlocked behind it :: PRIV — set the clock by hand, flush caches; prevent with an IP-addressed NTP source. The doctor must reach this verdict *before* any resolved restart, which cannot help and reads as a loop.
- nsswitch.conf hosts line / avahi mDNS broken :: PRIV — fix the hosts line, install nss-mdns.
** Cluster 6 — names resolve, egress blocked
diff --git a/docs/design/2026-07-15-velox-boot-failure-handoff.org b/docs/design/2026-07-15-velox-boot-failure-handoff.org
new file mode 100644
index 0000000..5ec996f
--- /dev/null
+++ b/docs/design/2026-07-15-velox-boot-failure-handoff.org
@@ -0,0 +1,61 @@
+#+TITLE: Velox boot failure — ZBM found no bootable kernel; diagnosis in progress, recovery plan attached
+#+AUTHOR: Craig Jennings
+#+DATE: 2026-07-15
+
+* Why this is coming to archsetup
+
+Velox fails to boot: ZFSBootMenu reports it can't find a bootable environment with a kernel. Craig reports the last working velox session was an archsetup health-check run that included the pacman upgrade — so the breakage most likely happened inside archsetup's own workflow, and Craig wants the diagnosis + retrospective to continue here with full context. The .emacs.d session (where this was triaged, only because that's where Craig was sitting) hands off everything below.
+
+A phone photo of the zfs list output from velox's ZBM recovery shell accompanies this note in the inbox.
+
+* Timeline
+
+- 2026-07-13 ~23:50 CDT — velox last seen on the tailnet (per tailscale status read 2026-07-14 ~17:50).
+- During that last session: archsetup health-check workflow ran, including a pacman upgrade (Craig's recollection — pacman.log will confirm exact times).
+- 2026-07-14 late evening — Craig boots velox; ZBM: no bootable environment with a kernel.
+- 2026-07-14/15 — triage from the ZBM recovery shell, Craig driving, guided from the .emacs.d session.
+
+* Facts established so far (from the ZBM recovery shell)
+
+- zroot imported, health ONLINE. Every dataset's keystatus is "available" — encryption unlocked, not a key problem.
+- Layout confirmed from zfs list: zroot/ROOT/default (mountpoint /), separate datasets for home, home/root, media, var, var/cache, var/lib, var/lib/docker plus many docker layer children (legacy mountpoints). NOTE: no separate zroot/var/log dataset — /var/log lives inside zroot/var. That differs from the sanoid dataset list in archsetup's configure_zfs_snapshots (which configures zroot/var/log and zroot/var/lib/pacman as their own datasets) — worth reconciling in the retrospective.
+- Mounted the BE read-only style: mkdir -p /mnt/be && mount -t zfs -o zfsutil zroot/ROOT/default /mnt/be.
+- THE FINDING: /mnt/be/boot contains ONLY intel-ucode.img. vmlinuz-linux, initramfs-linux.img, and initramfs-linux-fallback.img are all gone.
+
+* Working hypothesis
+
+A kernel upgrade during the health-check run removed the old kernel files and never completed installing the new ones (interrupted transaction, mkinitcpio failure, or a /boot shadowing issue), and the machine was powered off with /boot empty. Arch's upgrade removes the running kernel's files at package-replace time, so a failure between "remove old" and "install new + mkinitcpio" leaves exactly this state: microcode present, kernel and initramfs absent.
+
+* Remaining diagnosis steps (not yet run — velox is sitting at the ZBM shell)
+
+1. Read pacman's log (on the zroot/var dataset):
+ #+begin_src sh
+ mkdir -p /mnt/var
+ mount -t zfs -o zfsutil zroot/var /mnt/var
+ tail -60 /mnt/var/log/pacman.log
+ #+end_src
+ Expect the failed/interrupted kernel transaction near the end; note its timestamp.
+2. List recovery candidates:
+ #+begin_src sh
+ zfs list -t snapshot zroot/ROOT/default | tail -20
+ #+end_src
+ Sanoid is configured for hourly=6/daily=7 on the ROOT dataset, so a pre-damage snapshot should exist. Check whether any pre-pacman_* snapshots appear — that tells us whether the 2026-06-29 pre-pacman hook design is actually installed on velox.
+
+* Recovery plan (agreed with Craig, pending the log read)
+
+1. Pick the newest zroot/ROOT/default snapshot that predates the failed transaction.
+2. If the pool is imported read-only (zpool get readonly zroot): zpool export zroot && zpool import -f -N zroot.
+3. zfs rollback -r zroot/ROOT/default@<snapshot> (the -r discards snapshots newer than the target; home/var/media are separate datasets and untouched).
+4. zpool export zroot, reboot — ZBM should now see the kernel.
+5. After first boot: re-run pacman -Syu attended, and confirm /boot holds vmlinuz-linux + initramfs-linux.img before any shutdown.
+
+* Retrospective candidates for archsetup
+
+- Does the health-check / upgrade flow verify /boot contents (kernel + initramfs present, mkinitcpio exit status) after a kernel upgrade? This failure would have been caught by a one-line post-upgrade assertion.
+- Is the pre-pacman snapshot hook (2026-06-29 design, zroot/ROOT/default@pre-pacman_<ts>) installed on velox? The snapshot listing in step 2 above answers this empirically.
+- The sanoid config vs actual dataset layout mismatch (var/log, var/lib/pacman) noted above.
+- Whether the upgrade step should refuse to end the session (or page Craig) when a kernel transaction errors.
+
+* Related loose end already in your inbox
+
+A separate note (2026-07-14-1751) asks to add inetutils to the install base; velox also still needs that package installed once it boots again.
diff --git a/docs/design/2026-08-14-velox-reinstall-gaps-1.org b/docs/design/2026-08-14-velox-reinstall-gaps-1.org
new file mode 100644
index 0000000..cf0d723
--- /dev/null
+++ b/docs/design/2026-08-14-velox-reinstall-gaps-1.org
@@ -0,0 +1,137 @@
+#+TITLE: What the velox reinstall left behind — four gaps the install could close
+#+AUTHOR: Craig Jennings
+
+* Heads-up: this was found from a .emacs.d session
+
+I opened a .emacs.d session on velox this morning, two days after the fresh
+Arch install, and the first thing it did was fail: there was no =.ai/=
+directory to read. Chasing that turned up four separate things the reinstall
+did not restore. Three I repaired from the session; one needs me at my phone.
+
+None of this is a .emacs.d bug. They are all install-side gaps, which is why
+they are landing in your inbox. Machine is velox; ratio was the reference for
+every comparison below.
+
+* Gap 1 — the gitignored tooling layer does not survive a reinstall
+
+=~/.emacs.d= was re-cloned on 2026-08-13. Git brought back every tracked file
+and none of the agent tooling, because =.gitignore= deliberately excludes it:
+=.ai/=, =.claude/=, =CLAUDE.md=, =todo.org=, and =inbox/= were all simply
+absent. That is the correct ignore policy — this repo relays to a public
+mirror — but it means a reinstall silently drops the entire working state of
+every gitignore-mode project.
+
+The damage on velox was total rather than partial: 374 files, 4.5 MB,
+including =todo.org= (556 KB) and 184 archived session files. Nothing carries
+it. Not git, not stow, not the bootstrap.
+
+I recovered it by rsyncing the set from ratio over the tailnet. Ratio was
+authoritative and velox held nothing, so there was no merge to adjudicate —
+which is luck, not design. Had velox held a few days of divergent state, this
+would have been a hand reconciliation. It has been one before: 2026-07-31, when
+the two machines' =.ai/= trees had forked to zero files in common.
+
+Worth knowing: this is fleet-general. Every project on the box that gitignores
+its =.ai/= has the same hole, not just =.emacs.d=.
+
+What the install could do: after cloning a project, check whether a sibling
+daily driver holds a =.ai/= for it, and offer to pull it across. Or at minimum,
+list the projects whose tooling layer is missing so the gap is visible on day
+one instead of at the first session that trips over it.
+
+* Gap 2 — stowed user timers come back linked but not enabled
+
+The unit files all arrived correctly through the dotfiles stow, symlinked into
+=~/.config/systemd/user/= and resolving fine. But being present is not being
+enabled, and the reinstall enabled only some of them:
+
+| unit | velox after reinstall | ratio |
+|---------------------------+-----------------------+----------|
+| calendar-sync.timer | enabled, active | enabled |
+| agenda-render-cache.timer | enabled, active | enabled |
+| roam-sync.timer | *linked, inactive* | enabled |
+| signal-receive.timer | *linked, inactive* | enabled |
+| emacs.service | linked, inactive | linked |
+
+=emacs.service= reads the same on both machines, so I take that one as
+intentional and left it alone. The other two are real drift: =systemctl --user
+enable= writes a =timers.target.wants= symlink into =~/.config/systemd/user/=,
+and that symlink is not stow-managed, so nothing in the dotfiles repo carries
+it. A stowed unit file is inert until something enables it.
+
+I enabled both with =systemctl --user enable --now=. Both fired immediately and
+exited clean, and both now show a next elapse.
+
+What the install could do: enable the units it stows, explicitly, as a named
+step. The inconsistency is the tell — two of four came back enabled, which
+suggests something enables a subset and nothing enumerates the rest.
+
+* Gap 3 — the roam clone was stale, and held a diff that would have destroyed data
+
+This one has an ordering constraint, so it matters more than its size suggests.
+
+velox's =~/org/roam= was ten commits behind ratio, stuck at the 2026-08-04
+auto-sync while ratio was at 2026-08-14 — a direct consequence of gap 2, since
+=roam-sync.timer= was never enabled here.
+
+The dangerous part: velox's clone also carried an *uncommitted* =inbox.org=
+that had been emptied. Seventeen deletions, file down to zero bytes, holding a
+pre-2026-08-04 state whose captures were long since processed on ratio.
+
+So the naive repair — enable =roam-sync.timer= and let it catch up — would have
+committed that emptying and pushed it, deleting the four live inbox items on
+ratio. The timer is the repo's only committer and it commits whatever it finds.
+
+I checked ratio's =inbox.org= first and confirmed it was a strict superset of
+velox's HEAD version (same three items plus an 2026-08-09 capture), which made
+the local change provably worthless. Then discarded it, fast-forwarded to
+=a411b43=, and only then enabled the timer. Clone is clean and current, first
+sync ran green.
+
+What the install could do: if it ever enables =roam-sync= on a rebuilt machine,
+reconcile the clone *before* enabling, not after. An auto-committing timer
+pointed at a stale dirty clone is a data-loss path, and the failure is silent
+and remote — it lands on the *other* machine.
+
+* Gap 4 — signal-cli lost its registration, and that breaks the whole fleet
+
+=signal-receive.service= ran for the first time and reported:
+
+: signal-receive: +15045173983 not registered on this machine — nothing to do
+
+velox's signal-cli data dir holds a 39-byte empty =accounts.json=. Ratio still
+has both numbers. So the reinstall wiped the registration, and per the design
+notes velox was supposed to be the *primary* — ratio is the linked device.
+
+The effect is wider than velox, because of how =agent-text= dispatches: if the
+local signal-cli holds the account it sends directly, otherwise it ssh-relays to
+a hardcoded velox. Velox no longer holds it, so a send from here relays to
+itself and fails; a send from any third machine relays to velox and fails the
+same way. Only ratio still works, and only via the direct branch. The error text
+blames "velox down or unreachable", which is misleading — velox is up and on the
+tailnet, it just is not registered.
+
+This is the one I could not repair from the session: re-linking needs me at my
+phone (Signal → Settings → Linked Devices, scanning the QR from =signal-cli
+link -n velox=). Filed in .emacs.d's todo.org as [#B].
+
+What the install could do: verify =signal-cli listAccounts= is non-empty after a
+rebuild and say so loudly if it is not. Silent loss of the phone channel is
+exactly the kind of thing nobody notices until the page that mattered never
+arrives.
+
+* Summary of what I changed on velox
+
+- Restored =.ai/=, =.claude/=, =CLAUDE.md=, =todo.org=, =inbox/= to
+ =~/.emacs.d= by rsync from ratio.
+- Discarded the stale local =inbox.org= diff in =~/org/roam= and fast-forwarded
+ the clone to current.
+- Enabled and started =roam-sync.timer= and =signal-receive.timer=.
+
+Left alone, deliberately: =emacs.service= (matches ratio), and velox's Signal
+registration (needs the phone).
+
+One unrelated thing I noticed while comparing the machines: ratio's signal-cli
+warns its messages were last received twelve days ago, even though its
+=signal-receive.timer= is enabled and active. That may be nothing, but the
+receive cadence there is worth a look.
diff --git a/docs/design/2026-08-14-velox-reinstall-gaps-2.org b/docs/design/2026-08-14-velox-reinstall-gaps-2.org
new file mode 100644
index 0000000..95842ac
--- /dev/null
+++ b/docs/design/2026-08-14-velox-reinstall-gaps-2.org
@@ -0,0 +1,63 @@
+#+TITLE: Fifth reinstall gap — machine-local .local.el config, and a general shape
+#+AUTHOR: Craig Jennings
+
+* Follow-up to this morning's handoff
+
+Sent you four gaps an hour ago
+([[file:2026-08-14-velox-reinstall-gaps-1.org][the first report]]). Here is a fifth,
+found straight afterwards when I noticed calendar sync was dead on velox.
+
+* What was broken
+
+=calendar-sync.timer= was enabled and firing every fifteen minutes, and failing
+every time with exit 255:
+
+: calendar-sync: No calendars configured (set calendar-sync-calendars)
+
+The three output files sat at zero bytes. The cause is that
+=~/.emacs.d/calendar-sync.local.el= is gitignored, so the reinstall deleted it
+along with everything else untracked, and the module's loader treats a missing
+file as a *silent* no-op. So the config vanished quietly and the only symptom
+was a failing unit nobody was watching.
+
+Cheap to fix once found: the repo tracks =calendar-sync.local.el.example=, and
+that template already encodes the shape velox uses — feeds resolved by
+=:secret-host= against =authinfo.gpg= rather than inlined. The authinfo entries
+had survived, because =~/.authinfo.gpg= is a stow symlink into the dotfiles repo.
+So rebuilding was one copy, and all three feeds now sync clean and land
+byte-identical to ratio's.
+
+* The general shape, which is the part worth acting on
+
+This is the same failure as gap 1, one layer down, and it is worth stating
+generally because the install can act on it:
+
+- A tracked =*.local.el.example= template plus a gitignored =*.local.el= is a
+ deliberate pattern in this config, not a one-off. =.gitignore= lines 56-58
+ list three of them: =calendar-sync.local.el=, =signal-config.local.el=,
+ =google-keep.local.el=. Every one of those is gone on velox right now. I have
+ only repaired the calendar one.
+- Secrets held *by reference* survive a rebuild; secrets held *inline* do not.
+ The calendar config came back for free because the tokens were in
+ =authinfo.gpg=, which is stow-managed and therefore travels. Ratio's copy of
+ the same file inlines its URLs, and had ratio been the machine rebuilt, those
+ three feed tokens would simply have been gone.
+- The failure was silent by design. A missing local config is a no-op, which is
+ right for a machine that never configured the feature and wrong for one that
+ just lost it.
+
+* What the install could do
+
+- After a rebuild, enumerate every tracked =*.local.el.example= in a project and
+ report which have no corresponding =*.local.el=. That is a one-line find and it
+ turns a silent no-op into a visible checklist item.
+- Same for any =*.local.*= convention elsewhere in the fleet — the pattern is not
+ specific to Emacs.
+- Worth pairing with gap 2: a unit that is enabled and failing every fifteen
+ minutes for two days is its own signal. A post-rebuild pass over
+ =systemctl --user list-units --state=failed= would have caught this one
+ without knowing anything about calendars.
+
+That last one generalizes best. Of the five gaps I have sent you, three were
+things that *looked* fine — a stowed unit file, an enabled timer, a present
+clone — and were not.
diff --git a/docs/design/2026-08-14-velox-reinstall-gaps-3.org b/docs/design/2026-08-14-velox-reinstall-gaps-3.org
new file mode 100644
index 0000000..9675973
--- /dev/null
+++ b/docs/design/2026-08-14-velox-reinstall-gaps-3.org
@@ -0,0 +1,98 @@
+#+TITLE: Reinstall gaps, part three — per-install certs and credentials, and one failure that hid the others
+#+AUTHOR: Craig Jennings
+
+* Third handoff today
+
+Two earlier notes covered five gaps
+([[file:2026-08-14-velox-reinstall-gaps-1.org][the first report]] and
+[[file:2026-08-14-velox-reinstall-gaps-2.org][the follow-up]]).
+Email was the last thing broken on velox after the 2026-08-13 rebuild, and it
+turned up two more — both the same shape, and one of them with a property worth
+generalizing.
+
+Email is fully working now: three accounts, 21,853 messages, 4.0 GB indexed.
+
+* Gap 6 — the Proton Bridge TLS cert is per-install, and its absence disabled every account
+
+=~/.mbsyncrc= carries =CertificateFile /home/cjennings/.config/protonbridge.pem=.
+That file did not exist after the rebuild, and it cannot be restored from backup
+or copied from the other machine: Proton Bridge generates a fresh self-signed
+cert per installation. Velox's is issued 2026-08-13 23:44 with a different
+fingerprint from ratio's 2026-01-30 one.
+
+*The part worth acting on is the blast radius.* mbsync parses its entire config
+before doing any work, so a missing =CertificateFile= referenced by *one* account
+aborts the run for *all* of them. Gmail and dmail need no bridge and no cert, and
+both were dead anyway. The error names only the missing pem, so the symptom
+("no mail at all") and the message ("this one file is missing") look unrelated.
+
+Recovery does not need the bridge GUI. The running bridge presents the cert on
+its own IMAP port, so it can be pulled straight off the handshake:
+
+: openssl s_client -connect 127.0.0.1:1143 -starttls imap -showcerts </dev/null \
+: | sed -n '/BEGIN CERTIFICATE/,/END CERTIFICATE/p' > ~/.config/protonbridge.pem
+
+That is a two-second, fully scriptable step, which makes it a good candidate for
+the install rather than a runbook line.
+
+* Gap 7 — the bridge password is per-install too, and reports a stale value misleadingly
+
+=~/.mbsyncrc= resolves the cmail password with =cat ~/.config/.cmailpass=. That
+file is plaintext and, unusually for my setup, a real file rather than a stow
+symlink — so it is not in the dotfiles repo, not encrypted, and not carried to a
+new machine.
+
+The file survived the rebuild but held the *previous* install's password, because
+the bridge regenerates it per installation. Ratio's and velox's differ by sha256,
+confirmed today.
+
+*The diagnostic trap:* Proton Bridge answers a wrong password with =no such
+user=. I read that as "the bridge has no account signed in" and went looking for
+a login problem. The account was configured the whole time. If the install ever
+validates bridge connectivity, it should not treat =no such user= as evidence
+about account state.
+
+* The generalization
+
+Gaps 6 and 7 are the same as 1 through 5, sharpened. Everything that broke in
+this rebuild was *generated on the machine by an application* rather than carried
+by git, stow, or the dotfiles repo:
+
+| gap | artifact | why it did not travel |
+| 1 | =.ai/=, =todo.org=, =CLAUDE.md= | gitignored |
+| 2 | =timers.target.wants= symlinks | written by systemctl enable |
+| 3 | roam clone state | local working tree |
+| 4 | signal-cli registration | per-device identity |
+| 5 | =*.local.el= configs | gitignored |
+| 6 | bridge TLS cert | per-install, regenerated |
+| 7 | bridge password | per-install, regenerated |
+
+Gaps 6 and 7 add a distinction the earlier note missed. For 1, 3 and 5 the old
+value is still correct, so *restoring* fixes them. For 4, 6 and 7 the old value is
+*worthless* — the application has generated a new one, and only *re-deriving*
+from the live system fixes them. An install that tries to restore these will
+produce exactly what happened here: a file that exists, looks right, and
+authenticates against nothing.
+
+So the install's post-rebuild checklist wants two columns, not one: what to
+restore, and what to re-derive.
+
+* What the install could do
+
+- Re-derive the bridge cert from the running bridge with the =openssl s_client=
+ line above. Scriptable, no GUI, no secrets.
+- Re-derive the bridge password from the bridge rather than expecting the file to
+ be right, and rewrite =.cmailpass=. (I have filed a task on my side to make
+ =PassCmd= ask the bridge directly, which would remove the file entirely.)
+- Add a cheap post-rebuild validation that =mbsync --list= parses. Config-parse
+ failures disable every account at once and say nothing about mail, so they are
+ worth catching explicitly rather than via "no new mail" hours later.
+- More generally: keep the restore list and the re-derive list separate, per the
+ table above.
+
+* Unrelated, but noticed while comparing the machines
+
+=~/.config/.gmailpass.gpg= and =~/.config/.dmailpass.gpg= resolve to mode 777 in
+the dotfiles repo, on both machines. They are gpg-encrypted so the contents are
+safe, but world-writable is wrong for a credential file. That is a dotfiles fix,
+not an archsetup one — noting it here only because it surfaced in the same pass.
diff --git a/docs/post-install-checklist.org b/docs/post-install-checklist.org
index 97fc0d5..8c48938 100644
--- a/docs/post-install-checklist.org
+++ b/docs/post-install-checklist.org
@@ -18,6 +18,32 @@ bluetooth pairing landed below.
* Checklist
+** Run the post-rebuild check first
+
+Before working through the manual steps below, run:
+
+#+begin_src sh
+~/code/archsetup/scripts/post-rebuild-check
+#+end_src
+
+It runs the five checks a rebuilt machine actually needs — failed units,
+user units that are present but never enabled, =*.example= configs whose
+real sibling is missing, gitignore-mode projects missing the working state
+their own =.gitignore= names, and the signal-cli registration. Each prints
+a line whether or not it finds anything; exit 1 means something needs
+attention.
+
+These are the gaps velox hit within two days of its 2026-08-13 reinstall,
+and three of the five looked fine on casual inspection: a stowed unit file,
+an enabled-looking timer, a present git clone. Run it again a day or two
+after the install, once timers have had a chance to fail.
+
+It normally finishes in a second or two. On a machine whose user systemd is
+wedged it takes a couple of minutes instead, because every =systemctl= call
+is bounded at five seconds and check 2 makes one per unit. That is the slow
+case working as intended: it reports what it could not read rather than
+hanging. Set =PRC_SYSTEMCTL_TIMEOUT= lower to cut the wait.
+
** Pair bluetooth peripherals
Pairing is inherently interactive (scan, pick the device, confirm), so it
@@ -71,7 +97,12 @@ needs doing.
The installer's completion message carries the steps; recorded here too so
the checklist is complete:
-1. Clone claude-templates to =~/projects/claude-templates= if missing.
-2. Run =protonmail-bridge --cli=, log in, then quit.
-3. Run =~/code/archsetup/scripts/cmail-setup-finish.sh=.
-4. First mail sync: =mbsync cmail && mu index=.
+1. Run =protonmail-bridge --cli=, log in, then quit.
+2. Run =~/code/archsetup/scripts/cmail-setup-finish.sh=.
+3. First mail sync: =mbsync cmail && mu index=.
+
+Sending mail also needs =cmail-action= on PATH, which rulesets owns: clone it
+to =~/code/rulesets= and run =make install=. That is not a prerequisite for the
+steps above — the setup script warns and carries on — but =mbsync= is the first
+thing that wants it. An agent session runs =make install= at startup, so on a
+machine that runs them the link appears on its own.
diff --git a/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org b/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org
new file mode 100644
index 0000000..d9ec8d4
--- /dev/null
+++ b/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org
@@ -0,0 +1,213 @@
+#+TITLE: Guarded-Upgrade Completion — keeping topgrade freshness honest
+#+AUTHOR: Craig Jennings
+#+DATE: 2026-08-25
+#+TODO: TODO | DONE
+#+TODO: DRAFT READY DOING | IMPLEMENTED SUPERSEDED CANCELLED
+
+* DRAFT Guarded-upgrade completion
+:PROPERTIES:
+:ID: 81cdfd72-db96-43d3-aa03-779878c99f3e
+:END:
+- [2026-08-25 Tue @ 18:45 -0600] decisions closed 7/7. The kernel decision reversed on the velox DKMS failure chain: held on every everyday run, landed only in the dedicated session behind a DKMS/initramfs/snapshot gate.
+- [2026-08-25 Tue @ 18:30 -0600] redirected: the everyday path is a live split upgrade (apply everything the guard would not block, defer the rest); the boot-time oneshot becomes the completion step for the deferred set. Decided while running exactly that by hand on ratio.
+- [2026-08-25 Tue @ 06:39:42 -0600] drafted. Grounded in a live read of the maint engine, the pacman hooks, and the boot path on velox, not memory. The topgrade-freshness diagnosis that motivates it is in this session's log.
+
+* Metadata
+
+| Status | draft |
+|----------+-------------------------------------------------------------|
+| Owner | Craig Jennings |
+|----------+-------------------------------------------------------------|
+| Reviewer | Craig Jennings |
+|----------+-------------------------------------------------------------|
+| Related | maint =topgrade_age= metric; =hypr-live-update-guard= hook |
+
+* Summary
+
+The waybar maintenance module shows topgrade freshness as permanently stale. The cause is a real one: on a machine running Hyprland, a full =topgrade= almost never exits 0, because its system step upgrades GPU/compositor libraries that the =hypr-live-update-guard= pacman hook correctly refuses to swap under a live session. The freshness stamp is gated on topgrade's exit code, so a correct, protective refusal reads as "you never run updates." This spec designs a safe path to actually complete a guarded upgrade, and makes that completion record the freshness stamp, so the metric tracks the true state of the system.
+
+* Problem / Context
+
+The metric reads one cache key, =topgrade_run= (=~/.local/state/maint/topgrade_run.json=). Absent, the probe (=maint/src/maint/probes/updates.py:114=) returns WARN, "no topgrade run recorded". Two writers stamp it: the =topgrade= PATH wrapper (=~/.dotfiles/hyprland/.local/bin/topgrade=) on =rc -eq 0=, and the panel's TOPGRADE lever (=doctor.py=), which returns before the stamp on any non-zero exit. The read path is sound (a sandboxed =maint stamp topgrade= writes the file and =maint status= then reads freshness 0); the file is simply never written.
+
+It is never written because topgrade rarely exits 0 on this machine, and the reason is specific rather than flaky. =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook= is a =PreTransaction=/=AbortOnFail= hook that, when Hyprland is running and an upgrade changes the on-disk version of a GPU/compositor library, prints a BLOCKED banner and exits 1 — aborting the whole transaction before any file is swapped. Its trigger set is =mesa=, =mesa-*=, =wayland=, =libdrm=, =libglvnd=, =hyprland=, =aquamarine=, =hyprutils=, =hyprgraphics=, =vulkan-radeon=, =vulkan-intel=, =vulkan-mesa-layers=, =nvidia-utils=, =lib32-nvidia-utils=, =xorg-xwayland=. The guard exists for a proven failure: replacing those libraries under a live compositor makes the next GPU call hit a now-deleted mapping and SIGABRT, taking every Wayland client down (hit on ratio 2026-06-07).
+
+So when any of those libraries has an update pending — a frequent event — topgrade's =system= step (it runs =yay=) aborts non-zero, topgrade returns non-zero, and neither writer stamps. The observed case: on 2026-08-24 topgrade ran at 17:49, hit the guard on =mesa= (26.1.7 → 26.2.1), and failed; the upgrade was then finished by hand with the guard's sentinel override, entirely outside the wrapper, so nothing stamped. The metric has read stale ever since.
+
+Two framings of the fix are in tension, and choosing between them is the spec's central decision. Either the metric means "how recently did you run the sweep" (recency), so the stamp should decouple from topgrade's exit; or it means "is the system up to date" (state), so staying stale while a guarded upgrade is deferred is *correct* and the only real defect is that safely completing that upgrade doesn't stamp. This spec takes the state framing (see Decisions).
+
+* Goals and Non-Goals
+
+** Goals
+- A safe, low-friction way to apply a guarded (GPU/compositor-library) upgrade, with Hyprland not live at swap time.
+- That completion records the =topgrade_run= freshness stamp, so the metric clears when the system is genuinely current.
+- A boot-time upgrade path that can never lock the machine out of its session, however it fails.
+- The installer owns the durable pieces so a rebuilt machine has them without hand-setup.
+
+** Non-Goals
+- Weakening or bypassing the =hypr-live-update-guard= hook. It stays exactly as strict; this builds *around* it, not through it.
+- Making the full topgrade ecosystem sweep (git repos, vim, npm, ...) run at boot. Those never need a stopped compositor and are out of the boot path.
+- Changing how the kernel hazard is *guarded*. The hook stays silent on kernels; the split script holds them back on a live run as a second, separately-reasoned list (see Design), which is a deferral policy rather than a guard.
+- A general offline-update system for all of pacman. Scope is the guarded-library case.
+
+** Scope tiers
+- v1: the split-upgrade script (live: apply the non-blocked remainder, defer the rest, run the ecosystem sweep with the system step off, report the deferred set); maint's UPDATE/TOPGRADE levers route through it; an "apply on reboot" affordance that installs the held kernel live and arms the boot-time oneshot for the GPU/compositor set.
+- Out of scope: full-sweep-at-boot; touching the guard's policy.
+- vNext: none open — the kernel deferral that was vNext is now part of v1's held set.
+
+* Design
+
+The shape follows one principle: the only part of topgrade that needs a stopped compositor is its =system= step when a guarded library is pending. Everything else runs fine live and rarely fails. So the safe path is small and targeted — apply the guarded system upgrade with Hyprland down, once, and stamp it — while the ordinary full sweep stays a normal live =topgrade= run.
+
+Three pieces, at two altitudes — but the everyday gesture is not a reboot. It is a normal live update that simply leaves the dangerous few behind.
+
+*The split script.* A pacman =PreTransaction= hook can only abort or allow the transaction it is handed; it cannot drop targets from it. So "upgrade everything except the guarded set" cannot live in the hook — it lives one layer up, in a script the panel calls. On a live run the script: refreshes the sync db and reads the pending set (=checkupdates=); computes the *blocked set* = the guard's own trigger list (read from the installed hook's =Target= lines, so there is one source of truth, and version-aware the way the guard is — a same-version reinstall is not a swap) plus the *kernel set* (every installed kernel with its =-headers=, always as a set; held on every everyday run because a failed DKMS rebuild on velox's ZFS root leaves the machine unbootable — see the kernel decision); clears the news hook (=informant read=) where installed; runs =pacman -Syu --noconfirm --ignore=<blocked set>=; runs the AUR-only remainder (=yay -Sua --noconfirm=, AUR packages pinning a guarded version hold themselves back); then runs =topgrade --disable system,git_repos -y= so the other ecosystems still get their sweep and topgrade can actually exit 0. It writes the deferred set to a state file the panel reads, and exits 0 when the live part succeeded, whatever was deferred. The guard hook stays installed as the backstop for a bare =pacman -Syu= typed at a shell; on the driven path it never fires. Proof of concept: this exact sequence, run by hand on ratio on 2026-08-25 while Hyprland was live, resolved 724 of 730 pending packages (Emacs 31.1 among them) with the six guard hits deferred — after one unrelated fix, an orphaned =qemu-block-gluster= that had been dropped from the repo.
+
+*For the user.* UPDATE and TOPGRADE on the panel run the split script; they succeed, and the panel shows "N deferred" when the script held anything back. Landing the deferred set is a dedicated session, chosen on purpose, run in the foreground from the panel's action or =guarded-upgrade --complete= in a terminal: first the kernel set, live, with the desktop still up; then the gate — every DKMS module built for the new kernel, a fresh initramfs, and on a ZFS root a pre-pacman snapshot to fall back on. If the gate fails the script stops there, names what failed, and does not reboot; the machine keeps running on the old kernel and the desktop is available for the fix. If it passes, the script arms a persistent flag for the GPU/compositor set and offers to reboot (or, from a TTY with no compositor, applies that set directly). On the next boot, before the autologin shell starts Hyprland, the deferred guarded upgrade runs in the console — the guard passes freely because nothing is live — the stamp is written, the flag is cleared, and boot continues into the session. No second reboot: the libraries are already current before anything maps them. If anything goes wrong, the machine still boots into Hyprland and the panel still shows the pending upgrade, so you are never worse off than before arming.
+
+*For the implementer.* A persistent arm flag (a file on a non-tmpfs path, e.g. =/var/lib/archsetup/apply-upgrade-on-boot=, so it survives the reboot the =/run= guard sentinel cannot). A system oneshot, =archsetup-boot-upgrade.service=, =ConditionPathExists= on the flag, ordered =Before=getty@tty1.service= so it completes before autologin execs Hyprland — this ordering is mandatory, because a parallel run would let Hyprland start mid-swap and reintroduce the exact crash the guard prevents. The unit is bounded (=TimeoutStartSec=) and best-effort: its failure or timeout must not fail any target the session needs, so boot proceeds past it regardless. Its =ExecStart= runs, as the user: =informant read= (clear the news hook that would otherwise abort the transaction), then =topgrade --only system= (or the equivalent =yay -Syu=), then =maint stamp topgrade= on success, then removes the flag unconditionally (a one-shot arm — a failed attempt disarms rather than retrying every boot). =sudo= works unattended (=%cjennings NOPASSWD: ALL=), so no password prompt wedges it.
+
+The stamp also needs to happen when the upgrade is completed by other safe means — the by-hand sentinel-override path, or a =maint= command that does the same thing. The cleanest single home for the stamp is a small =maint apply-upgrade= (or a flag in the existing lever) that performs the guarded system upgrade and stamps on success, which both the boot unit and an interactive TTY run call. That keeps one code path that "completes a guarded upgrade and records it," rather than three writers that can drift.
+
+* Alternatives Considered
+
+** A. Decouple the stamp from topgrade's exit code (stamp on any real run)
+- Good, because it is a one-line change to the wrapper and needs no boot machinery.
+- Bad, because it throws away honest signal: a topgrade that was blocked from applying a real upgrade would read as "fresh," so the metric stops meaning "up to date." On this machine the blocked case is the common case, so the metric would be fresh precisely when an upgrade is outstanding.
+- Neutral, because the failed steps still surface elsewhere (pending-updates count), so freshness would become redundant rather than wrong.
+
+** B. Run the full topgrade live with the guard overridden, then reboot
+- Good, because it needs no new unit — arm the sentinel, run, reboot.
+- Bad, because the dangerous window is the whole rest of the run: mesa swaps early, then topgrade spends minutes on other ecosystems while the live compositor is one new GL context (a new window, the wallpaper daemon) away from SIGABRT. topgrade's own reboot-at-end is that window, not a fix for it.
+- Neutral, because it would stamp naturally on success — if it survived.
+
+** C. Manual TTY ritual only (log out, run topgrade at the console, reboot), plus stamp
+- Good, because it is the safest path and needs almost no code — just make the completion stamp.
+- Bad, because it is all manual, every guarded-upgrade day; the friction is why it won't happen consistently, which is how the metric got stale in the first place.
+- Neutral, because it is exactly what the boot unit automates, so it is really "v1 minus the automation."
+
+** D. Boot-time armed oneshot, arch-only (this spec)
+- Good, because the risky swap happens with nothing live, the run is one bounded transaction with a tiny prompt surface, it stamps on success, and a failure degrades to "boots normally, try again."
+- Bad, because it puts a unit on the boot critical path, which must be bounded and non-fatal with care, and it is the most to build.
+- Neutral, because it composes with C: the same =maint apply-upgrade= path serves both an interactive TTY run and the boot unit.
+
+** E. Split the live run: apply the non-blocked remainder now, defer the rest (this spec's everyday path)
+- Good, because it is what a careful operator does by hand anyway — and did, on ratio, the day this was decided. The live run succeeds on the common day, topgrade exits 0, the AUR and every other ecosystem stay current, and the guard's abort becomes the rare path rather than the default.
+- Bad, because Arch calls any =--ignore= run a partial upgrade. In practice pacman still enforces declared dependencies, so anything needing the newer mesa fails resolution instead of installing broken; the residual exposure is a package with an *unversioned* dependency built against a new ABI, which for mesa/wayland/libdrm is rare. Named, accepted.
+- Bad, because a deferred set nobody surfaces is a set that silently never lands — the same trap as the freshness stamp, one layer down. So the script must record the deferred set durably and the panel must show it; this is why D stays in the design as the completion step rather than being replaced.
+- Neutral, because it does not change the guard at all; it changes who decides the transaction's contents.
+
+* Decisions [7/7]
+
+** DONE Metric means state, not recency
+CLOSED: [2026-08-25 Tue 18:45]
+- Owner / by-when: Craig / 2026-08-25
+- Context: the stamp gate can mean "ran the sweep" or "system is current." The whole fix differs by which.
+- Decision: We will keep the state meaning. Freshness stays stale while a guarded upgrade is genuinely un-applied, and the fix is to make *safe completion* stamp — not to loosen the gate.
+- Consequences: easier — the metric stays trustworthy as an is-current signal, and Alternative A is off the table. Harder — completion now needs a real safe path (the rest of this spec) rather than a one-line wrapper change.
+
+** DONE Everyday mechanism is the split live run (Alternative E); the boot oneshot (D) completes the deferred set
+CLOSED: [2026-08-25 Tue 18:30]
+- Owner / by-when: Craig / 2026-08-25
+- Context: the first draft made D the primary gesture, which means every guarded-library day is a reboot day. Craig's read while watching the ratio run: when the guard would trip, the rational move is to upgrade everything *except* the guarded and kernel items, then run the rest of topgrade without the yay piece — and that logic should be a script we can keep editing, not something baked into the panel.
+- Decision: I will build E as the path UPDATE and TOPGRADE always take on a live session, and keep D as the way the deferred set lands (arm + reboot). C remains the manual fallback through the same script from a TTY (no compositor → nothing blocked → a full run). B stays rejected on the live-swap risk.
+- Consequences: easier — the common day is one live run that succeeds; reboots are reserved for the days the deferred set is non-empty, and even then the machine keeps working until the reboot is convenient. Harder — two lists to maintain (the guard's, read from the hook; the kernel list, owned by the script), a state file the panel must render, and the partial-upgrade caveat above to keep an eye on.
+
+** DONE The script lives in archsetup beside the guard, and maint calls it
+CLOSED: [2026-08-25 Tue 18:45]
+- Owner / by-when: Craig / 2026-08-25
+- Context: the panel (dotfiles =maint=) and the guard (archsetup =scripts/hypr-live-update-guard=, installed to =/usr/local/bin=) live in different repos, and maint already carries its own copy of the trigger list as =[updates] guard_patterns= in the thresholds TOML.
+- Decision: ship the split script in archsetup next to the guard, installed by the same installer step, reading the blocked list from the installed hook so the guard and the script can never disagree. maint's UPDATE/TOPGRADE levers change their =argv= to the script; the TOML patterns stay as the panel's *display-side* mirror (the badge that says a run will defer) and gain a test asserting they match the hook.
+- Consequences: easier — one owner for "which libraries are dangerous," and a rebuilt machine gets the script with the guard. Harder — a cross-repo change (archsetup ships it, dotfiles wires it), so the rollout is two commits, archsetup first.
+
+** DONE What a split run stamps
+CLOSED: [2026-08-25 Tue 18:45]
+- Owner / by-when: Craig / 2026-08-25
+- Context: under the state framing, a run that deferred six packages left the system *not* current, yet the sweep ran and every other ecosystem is fresh.
+- Decision: the script stamps =topgrade_run= only when the deferred set is empty. When it is non-empty it writes the deferred set to its own cache key, and the panel renders that as its own state ("6 deferred — apply on reboot") rather than as stale freshness. The boot oneshot stamps when it completes the deferred set. Freshness keeps meaning "current"; the deferred badge carries the other half.
+- Consequences: easier — no signal is thrown away, and the reboot nag has a precise count behind it. Harder — one more cache key and one more probe in maint.
+
+** DONE Kernel set is held on every everyday run and lands only in the dedicated session, gated on the DKMS result
+CLOSED: [2026-08-25 Tue 18:45]
+- Owner / by-when: Craig / 2026-08-25
+- Context: I first wrote this as "install the kernel live at apply-on-reboot," on the reasoning that a kernel swap crashes nothing and the modules-vanish window ends with the reboot. Craig asked what happens on velox when the DKMS rebuild fails, and the answer changed the decision. Velox is an encrypted ZFS root with =/boot= inside the root dataset, one kernel (=linux-lts=), and =zfs-dkms=. On a kernel upgrade the DKMS build runs PostTransaction, after the kernel is swapped and the old modules are deleted, so nothing can abort; a failed build leaves a new kernel beside an initramfs built for the old one, whose =zfs.ko= won't load, and the next boot can't import the pool. It is survivable — ZFSBootMenu can boot the pre-pacman snapshot, which holds the old kernel, initramfs, and modules — but it is a recovery session, not an update. Ratio (btrfs root, two kernels, zfs only for a data pool) is exposed only at the pool. The realistic triggers are a kernel major outrunning OpenZFS's supported range, a kernel upgraded without its headers, a toolchain regression, or a full disk.
+- Decision: the script holds the kernel set — every installed kernel with its =-headers=, moved as a set, never one without the other — on every everyday run, on both machines, so there is one rule rather than a per-host exception. The kernel set lands only in the dedicated session, live, while a working desktop exists for diagnosing, and the script gates what follows on the result: =dkms status= reports every DKMS module installed for the new kernel version, the initramfs is newer than the kernel image, and on a ZFS root a pre-pacman snapshot exists. A failed gate stops with the failure named and never reboots. The GPU/compositor set follows only after the gate passes — armed for the boot oneshot, or applied from a TTY. Kernels stay off the guard's list (the hook would block a TTY kernel upgrade for no reason). "Install the kernel live at apply-on-reboot" is withdrawn.
+- Consequences: easier — an everyday UPDATE can never put velox into the unbootable state, and the day the kernel moves is one Craig chose, sitting at the machine, expecting to handle issues. Harder — the kernel deferral is now standing, so the dedicated session has to happen on a cadence (security fixes ride the kernel), and the panel's deferred count carries a kernel most days; the gate is one more script to test, with fakes for =dkms status= and the image timestamps.
+
+** DONE Boot run applies exactly the deferred GPU/compositor set, nothing else
+CLOSED: [2026-08-25 Tue 18:45]
+- Owner / by-when: Craig / 2026-08-25
+- Context: the only packages that need a stopped compositor are the guard's trigger set; the rest run fine live and are what usually fail. The first draft phrased this as =topgrade --only system=; with the split script that wording is stale, and with the kernel decision above the kernel is not part of what boot applies either.
+- Decision: the boot oneshot runs the script's =--complete= form scoped to the deferred GPU/compositor set: one pacman transaction, no ecosystem sweep, no kernel. The full topgrade sweep stays a normal live run through the everyday path.
+- Consequences: easier — the boot path is fast, has a tiny interactive-prompt surface, and rarely fails. Harder — freshness after a boot run reflects the guarded set specifically, which is what the stamp decision above already accounts for.
+
+** DONE Arm flag lives on a persistent path and is one-shot
+CLOSED: [2026-08-25 Tue 18:45]
+- Owner / by-when: Craig / 2026-08-25
+- Context: the guard's =/run= sentinel is tmpfs and cleared on reboot, so it cannot carry an intent across the reboot. A boot that retries forever on failure is its own outage.
+- Decision: We will use a persistent flag (=/var/lib/archsetup/=) that the boot unit removes unconditionally at the end of its attempt — success or failure disarms.
+- Consequences: easier — the intent survives exactly one reboot and a failed attempt never wedges subsequent boots. Harder — a failed attempt needs re-arming, which is correct (a human decides to try again) but is a manual step.
+
+* Implementation phases
+
+** Phase 1 — The split-upgrade script (archsetup)
+=scripts/guarded-upgrade= (name open), installed to =/usr/local/bin= by the step that installs the guard. Behaviour as in Design: pending set → blocked set (hook =Target= lines, version-aware) ∪ held-kernel set when a compositor is live → =informant read= if present → =pacman -Syu --noconfirm --ignore=…= → =yay -Sua --noconfirm= → =topgrade --disable system,git_repos -y= → deferred set written to a state file → stamp only when nothing was deferred → exit 0 on a successful live part. The kernel set is derived from what is installed (every =linux*= kernel package and its =-headers=), never a hardcoded pair, and is always held or applied whole. Flags: =--dry-run= (print the plan and the deferred set, change nothing), =--no-topgrade=, =--no-aur=, =--complete= (the dedicated-session form: apply the kernel set live, run the gate, then arm the GPU/compositor set or, with no compositor live, apply it directly; stamp when the deferred set is empty). The gate is its own small script, =kernel-modules-check=: for each kernel under =/usr/lib/modules=, =dkms status= reports every registered module =installed= for it, and its initramfs is newer than its =vmlinuz=; on a ZFS root, a =pre-pacman_= snapshot of the root dataset exists. It exits non-zero with the failing item named, and =--complete= refuses to arm or reboot on that exit. Usable from a TTY at once. Tests (pytest beside the guard's): blocked-set computation against a fixture hook and version map; the kernel set is derived from the installed kernels and held whole on every everyday run; the =--ignore= list is exactly blocked ∪ kernel set; the state file round-trips; stamps only on an empty deferred set; =--dry-run= is IO-free; the gate passes and fails on fake =dkms status= output, image timestamps, and snapshot listings, and =--complete= never reaches the arm step on a failed gate.
+
+** Phase 2 — Wire maint to it (dotfiles)
+UPDATE and TOPGRADE levers change their =argv= to the script; the press-again-to-force sentinel wrap goes away (the driven path never trips the guard). A new probe reads the deferred-set state file and the panel renders "N deferred — apply on reboot" as its own row. A test asserts the TOML =guard_patterns= equal the installed hook's =Target= list. Tests under the maint fake harness.
+
+** Phase 3 — "Apply on reboot" and the boot-time unit (archsetup + maint)
+The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot. =archsetup-boot-upgrade.service=, installed by the installer: =ConditionPathExists= the flag, =Before=getty@tty1.service=, =TimeoutStartSec= bounded, non-fatal to every session target, =ExecStart= runs the script's =--complete= form as the user and removes the flag unconditionally. Installer step + unit file + the =/var/lib/archsetup/= flag directory. Tests in =tests/installer-steps/= for the install step; the arm action's tests in maint; a documented manual boot test (defer, arm, reboot, observe) in =todo.org= under Manual testing and validation.
+
+** Phase 4 — Docs, rollout, and both daily drivers
+Document the flow (arm → reboot → console upgrade → session). Roll the unit to velox and ratio (installer already covers a rebuild; existing machines need the one-time install). Confirm the ratio path matches.
+
+* Acceptance criteria
+- [ ] With a guarded library pending and Hyprland live, UPDATE applies everything else, exits 0, and the panel shows the exact deferred set; the guard hook does not fire.
+- [ ] The same run with no compositor live (a TTY) applies everything but the kernel set and stamps only if nothing was deferred.
+- [ ] A =--complete= run whose DKMS build fails stops before arming or rebooting, names the failure, and leaves the machine running on the old kernel; on velox the pre-pacman snapshot it required is bootable from ZFSBootMenu.
+- [ ] With a guarded library pending, arming and rebooting applies it in the console before Hyprland starts, and =maint status= then reads a fresh =topgrade_age=.
+- [ ] A boot-upgrade failure (a failed step, a timeout, an aborted transaction) never blocks the session: the machine boots into Hyprland, the flag is cleared, and the panel still shows the pending work.
+- [ ] Unread Arch news does not wedge the boot run (=informant read= precedes the transaction).
+- [ ] A guarded upgrade completed from a TTY via the Phase-1 path stamps freshness identically to the boot unit.
+- [ ] The =hypr-live-update-guard= hook is unchanged and still blocks a live guarded swap.
+
+* Readiness dimensions
+Answer each, or write "N/A because…".
+- Data model & ownership: the arm flag (=/var/lib/archsetup/=, installer-owned) and the =topgrade_run= cache key (maint-owned). No user-authored data.
+- Errors, empty states & failure: the boot unit is best-effort and self-disarming; every failure path lands in "boot normally, metric stays stale, re-arm to retry." Named, non-silent.
+- Security & privacy: relies on the existing =%cjennings NOPASSWD: ALL=; the unit runs the upgrade as the user via sudo, adds no new privilege. Note the NOPASSWD breadth as a pre-existing fact, not introduced here.
+- Observability: the boot run's output is on the console; its systemd unit status and journal record success/failure; the panel reflects the cleared or still-pending state after boot.
+- Performance & scale: one pacman/yay transaction at boot; bounded by =TimeoutStartSec=. Negligible boot-time cost when the flag is absent (=ConditionPathExists= skips the unit).
+- Reuse & lost opportunities: reuses the guard's trigger list by reading the installed hook (single source of truth for "which libs are dangerous"), =informant=, =maint stamp=, =checkupdates=, and topgrade's own step switches. The one duplicate that exists today — maint's TOML =guard_patterns= — is kept as a display mirror and pinned to the hook by a test rather than removed.
+- Architecture fit & weak points: integration points are the pacman hook set, getty autologin ordering, and the maint cache. Weak point: the =Before=getty@tty1= ordering is load-bearing for safety; a parallel run reintroduces the live-swap crash. Mitigated by making the ordering explicit and tested-by-inspection.
+- Config surface: the arm flag path and the timeout. Defaults safe (absent flag = no-op).
+- Documentation plan: a short "reboot to apply guarded upgrades" note in the maint docs; the installer step self-documents in-comment.
+- Dev tooling: installer-step pytest for Phase 2; maint unit tests for Phases 1 and 3; a manual boot test in =todo.org=.
+- Rollout, compatibility & rollback: additive; removing the unit and flag reverts fully. Existing machines need a one-time install; a rebuild gets it from the installer. Rollback leaves the guard and manual TTY path intact.
+- External APIs & deps: topgrade =--only system=, =informant read=, =yay=, =maint stamp= — all verified present on velox this session. No external service.
+
+* Risks, Rabbit Holes, and Drawbacks
+- Boot critical path: the unit sits ahead of autologin, so a hang would delay boot. Mitigated by =TimeoutStartSec= and non-fatal wiring; worst case is a bounded delay, then a normal session.
+- Interactive prompts under no stdin: =yay=/pacman can still prompt (provider choice, replace, AUR review) even with =assume_yes=. The =--only system= scope and =--noconfirm=-style flags shrink this to near zero, but a prompt with no stdin fails the run (benign) — needs a genuinely non-interactive invocation, verified in Phase 2.
+- Partial ecosystem state: N/A for the GPU hazard — each pacman run is one atomic transaction, so there is no half-swapped library. The =--ignore= run is a partial upgrade in Arch's sense; pacman's dependency resolution is the safety net, and the residual unversioned-ABI exposure is accepted in Alternative E.
+- Orphans that block resolution: a package dropped from the repo but still pinning an old version (ratio's =qemu-block-gluster= on 2026-08-25) fails the whole transaction. The script should detect the "could not satisfy dependencies" case, name the foreign package, and stop with the remedy — never =-Rdd= on its own.
+- The kernel on a DKMS ZFS root: a failed =zfs-dkms= build after the kernel swap cannot be aborted (the DKMS hooks are PostTransaction) and leaves velox unbootable on the new kernel. Mitigated by holding the kernel set on every everyday run, landing it only in the dedicated session behind the gate, and by the standing fallback: =/boot= lives in the root dataset, the =05-zfs-snapshot= hook snapshots it before every transaction, and ZFSBootMenu can boot that snapshot. The pacman cache also keeps the previous kernel and =zfs-dkms= for a downgrade. Ratio's exposure is its data pool only (btrfs root, two kernels).
+- Standing kernel deferral: because the everyday run never moves the kernel, the dedicated session has to happen on a cadence or kernel security fixes sit unapplied. The panel's deferred row is the reminder; a stale-kernel age in maint is a possible follow-up.
+
+* Testing / Verification / Rollout
+Phase-1 and Phase-3 logic under the maint fake harness; Phase-2 install under =tests/installer-steps/=. The one thing no unit test can cover — that an armed reboot actually applies the upgrade pre-session and stamps — is a scripted manual test in =todo.org= (arm with a guarded lib pending, reboot, confirm the console run, the fresh metric, and a normal session). Roll to velox first, then ratio.
+
+* Review and iteration history
+** 2026-08-25 Tue @ 18:45 -0600 — Craig Jennings — author
+- What: closed all seven decisions. Reversed the kernel decision (hold on every everyday run; land only in the dedicated session, gated on DKMS built, initramfs fresh, snapshot present; withdrew "install live at apply-on-reboot"), reworded the boot-scope decision for the split design, added the =kernel-modules-check= gate to Phase 1 and the acceptance criteria, and wrote the velox failure chain and the ZFSBootMenu fallback into Risks.
+- Why: on velox a failed =zfs-dkms= rebuild after a kernel swap is unabortable and unbootable; that belongs in a session I chose, not in an update I expected to touch applications.
+- Artifacts: this session's log (velox boot layout verified live: ZBM on the ESP, =/boot= in =zroot/ROOT/default=, one kernel, =zfs-dkms 2.4.4=).
+** 2026-08-25 Tue @ 18:30 -0600 — Craig Jennings — author
+- What: made the split live run (E) the everyday path and the boot oneshot (D) the completion step; added the script-ownership, stamp-semantics, and kernel-hold decisions; rewrote the phases around the script; added the orphan-blocks-resolution risk.
+- Why: watching a 724-of-730 guarded run succeed by hand on ratio made it obvious the guard's abort should be the rare path, and that the logic belongs in an editable script the panel calls rather than in the panel.
+- Artifacts: this session's log; the ratio run (=ratio-upgrade.service=, =/var/log/ratio-upgrade.log=).
+** 2026-08-25 Tue @ 06:39:42 -0600 — Craig Jennings — author
+- What: initial draft.
+- Why: the topgrade-freshness metric reads permanently stale because the guard blocks the arch step; designing a safe completion path rather than loosening the gate.
+- Artifacts: this session's log; =hypr-live-update-guard= hook; maint =topgrade_age= probe.
diff --git a/docs/workflows/system-health-check.org b/docs/workflows/system-health-check.org
index b4f34a5..43aeff5 100644
--- a/docs/workflows/system-health-check.org
+++ b/docs/workflows/system-health-check.org
@@ -235,6 +235,16 @@ Updates are separate from issue investigation. After all issues are addressed (o
5. *Host-specific kernel watches.* On ratio: if =linux=, =linux-lts=, =linux-firmware=, or a major =mesa= bump is pending, run the addendum at [[file:strix-soak-watch.org][docs/workflows/strix-soak-watch.org]] before topgrade. Retire the addendum (delete the file + this bullet) when the strix-lts custom kernel is retired.
6. Run =topgrade= for the actual update (config at =~/.config/topgrade.toml=). On maint hosts (ratio, velox) plain =topgrade= resolves to the dotfiles PATH wrapper, which stamps the console's topgrade-freshness metric on success — no extra step. If the run happened outside the wrapper somehow, =maint stamp topgrade= records it by hand.
7. If linux-firmware, kernel, or Mesa were updated, recommend a reboot
+8. On a ZFS-root host, after a kernel bump confirm the new initramfs carries the zfs module before rebooting: =sudo lsinitcpio /boot/initramfs-linux-lts.img | grep -c 'zfs.ko'=. The =sudo= is load-bearing: the images are 0600, so an unprivileged =lsinitcpio= exits 1 with "Unable to read file" on stderr and nothing on stdout, and once piped into =grep -c= that empty stdout reads as a count of 0 and looks exactly like a missing module (velox, 2026-09-12).
+
+*** Before a firmware (BIOS) update
+
+Firmware stays a manual step (=topgrade.toml= keeps =[firmware] upgrade = false=); =fwupdmgr update= stages it and the next reboot flashes it. Before staging, capture the two things a bad reboot will make you guess at:
+
+1. =sudo efibootmgr -v= — every boot entry with its loader path and command line, pasted into the session's context file.
+2. Secure Boot state — =bootctl status 2>/dev/null | grep -i 'secure boot'=.
+
+After the flash, if the machine reports no bootable device, check Secure Boot *first*. The Framework 04.02 update on velox re-enabled it, which rejects the unsigned ZFSBootMenu loader and reads as "Default Boot Device Missing" rather than a security violation; the boot entries were untouched (see the Known Issues Log, 2026-09-12). Only when Secure Boot is off and =efibootmgr -v= from a stick shows the entry gone does the boot-entry recreate apply (velox: =docs/2026-08-15-velox-uefi-boot-entry-reference.org= in archsetup).
*** Two-Stage Reboot Pattern (MANDATORY if Phase 3 installed kernel / iproute2 / systemd / NetworkManager)
@@ -1036,3 +1046,28 @@ Each entry is scoped to one host (or =any=). When Phase 1 cross-references findi
- Functional status at 2026-06-13 check: no coredumps yet that day; Telegram scans still worked from cached chat state. Treat as an app/server-container crash, not a machine-health fault.
- Classification: KNOWN — annotate future =telega-server= coredumps on ratio as =KNOWN — dockerized telega-server musl SIGSEGV= if the signature matches =tdat_plist_value= / unexpected plist value or otherwise stays inside the Telega container. Escalate only if crashes become continuous, break Telegram workflows, or appear after moving off the Docker musl build.
- Deferred remediation options, in order of least disruption: update the Emacs =telega= package, rebuild/pull a newer =telega-server= image, pin a known-good pre-2026-06 image digest, build =telega-server= natively, or report upstream with =coredumpctl= and log evidence.
+
+** 2026-09-12: velox — Framework BIOS update re-enabled Secure Boot (reads as "no bootable device")
+:host: velox
+- Symptom: after fwupd staged system firmware 0.0.3.5 → 0.0.4.2 and the reboot flashed it, the BIOS reported "Default Boot Device Missing / no bootable drive". Indistinguishable from the NVRAM wipe that forced the 2026-08-13 reinstall.
+- Actual cause: the update set Enforce Secure Boot = Enabled. The unsigned ZFSBootMenu loader was rejected and the firmware reported it as a missing device, not a security violation. Boot0001 ZFSBootMenu survived intact with its command line; no =efibootmgr= was needed.
+- The tell: booting the Ventoy stick showed shim's MOK management screen.
+- Fix: F2 → Security → Secure Boot → Enforce Secure Boot = Disabled, F10. Boots straight into ZBM. Firmware confirmed at 04.02, pools healthy.
+- Prevention: the pre-firmware-update checklist in Phase 3 (record =efibootmgr -v= and the Secure Boot state before staging). Check Secure Boot before assuming NVRAM loss.
+- Still open: velox's ESP has no removable-media fallback (=/efi/EFI/BOOT= absent), so a real NVRAM wipe would still need the stick. Filed in archangel, which owns the ZBM install, as [#C] "Installed systems have no removable-media boot fallback on the ESP" (2026-09-12); it ships with the next ISO rebuild after it lands.
+
+** 2026-09-12: any — fwupdmgr activates passim, a public LAN listener
+:host: any
+- Symptom: running =fwupdmgr= (refresh, update) D-Bus-activates =passim.service=, fwupd's LAN metadata-sharing daemon, which listens on =0.0.0.0:27500= and trips the maint listeners check to crit.
+- The unit is static (no =[Install]= section), so =systemctl disable= is a no-op and it comes back on the next fwupdmgr run. Masking is what holds: =systemctl mask passim.service=. Ratio has been masked since 2026-07-21; velox was stopped and disabled on 2026-09-12 (the disable being the no-op) and got the mask on 2026-09-13. The installer masks it as part of installing fwupd. =P2pPolicy=nothing= under =[fwupd]= in =/etc/fwupd/fwupd.conf= also works, but that file is pacman-owned and invites pacnew churn, so the mask is the form in use.
+- Classification: KNOWN — a passim listener means a machine that predates the mask or lost it; mask it, don't allowlist it.
+
+** 2026-09-12: velox — topgrade containers step fails on locally built images
+:host: velox (ratio has the same shape with its own local images)
+- Symptom: topgrade exits 1 after a clean package run because the containers step tries to =docker pull= images that were built locally (=cj/telega-server=, =telega-server-glycin=) and gets "pull access denied". The wrapper then never writes the topgrade-freshness stamp; =maint stamp topgrade= by hand after confirming the package steps succeeded.
+- Classification: KNOWN — the step cannot succeed while local-only images exist. The fix (disable the containers step, or list the images under =ignored_containers=) is tracked on the topgrade guarded-upgrade task in archsetup's todo.
+
+** 2026-09-12: velox — mkinitcpio "Possibly missing firmware" for xhci_pci_renesas and qat_6xxx
+:host: velox
+- Stock Arch mkinitcpio noise on a kernel rebuild: =xhci_pci_renesas= wants the Renesas USB controller blob (AUR =upd72020x-fw=) and =qat_6xxx= is Intel QuickAssist firmware. Neither is hardware this machine has.
+- Classification: KNOWN — harmless; annotate and move on unless the named hardware appears.