diff options
52 files changed, 14069 insertions, 1760 deletions
diff --git a/archive/task-archive.org b/archive/task-archive.org index 2b5d26e..f2edf58 100644 --- a/archive/task-archive.org +++ b/archive/task-archive.org @@ -1568,3 +1568,2293 @@ Live verification on ratio replaced the planned VM run for remedies 1, 4 and 5, Left open, not a v1 gap: mpv played silently while the doctor read healthy. The stack was genuinely fine, and per-application stream routing is an explicit spec Non-Goal. Worth a task only if it recurs. Second sighting, 2026-07-10: Chrome stopped recognizing the microphone while the stack was healthy (Shure MV7+ default, unmuted, 82%, PTT off). Restarting Chrome fixed it; nothing on the machine was touched. Both sightings share a shape the doctor cannot see: the graph is fine and one client cannot use it. A third sighting turns this into a design question, namely whether the doctor should say "the stack is fine, the fault is in the application" rather than a bare healthy. +** DONE [#C] Net panel: Enterprise error never dismisses :bug:dotfiles:network: +CLOSED: [2026-07-12 Sun] +Fixed in dotfiles =a157bed=. Root cause: error toasts are sticky by design (so background refreshes can't wipe an unread error), but the enterprise join hint's flow posts no follow-up status and row clicks post none either, so nothing ever replaced it. Fix: a window-wide capture-phase click gesture dismisses a sticky toast on the user's next interaction; policy in =viewmodel.toast_action_plan= (unit-tested), timed toasts and background clears unchanged. Panel smoke run confirms launch/doctor/close with the gesture installed. Pointer-level dismiss is a manual-testing child (AT-SPI can't drive pointer gestures). Repro screenshot: =~/pictures/screenshots/2026-07-10_195911.png=. +** DONE [#C] Net diagnostics leak connection names + SSIDs into copyable report and --json :bug:dotfiles:network:solo: +CLOSED: [2026-07-12 Sun] +Resolved in dotfiles =df1543a=: the =redact_ssid= toggle now scrubs saved profile names, active SSIDs, envelope-carried names, and =.nmconnection= keyfile basenames from the copyable report and the diag/doctor =--json= envelopes (one systemic pass in =redact.py=; MAC/IP scrub applies to those envelopes too). On-screen output and functional envelopes (status/list) unchanged. 15 new tests; live-verified on ratio (toggle on removes the active connection name from =diagnose --json=, default unchanged). +The net doctor's copyable report (=report.py=, =scrub_text=) scrubs only MAC/IP, and =net diag/doctor --json= (=cli.py=) dumps the raw dict with no redaction. SSID redaction lives only in the event log (=redact_event=, gated on =redact_ssid=, default off). So a connection name (usually the SSID) appears in the clear in the link-step evidence and in every =--json= consumer — the copyable report is exactly the text a user pastes into a bug report. Secrets (PSK/password/token/portal URL) are already stripped, so this is names, not credentials: Minor severity, graded on severity alone per the privacy carve-out. + +Split out of the 2026-07-11 net-doctor-expansion spec review: that spec's new rival-manager/keyfile-perms verdicts keep parity with this pre-existing behavior rather than half-solve it. Fix shape: extend redaction to cover the connection name + keyfile basename across the copyable report and =--json= (one systemic pass, not per-verdict), with a redaction test. Engine-wide, so it wants one coherent change rather than being bolted onto the expansion work. +** DONE [#B] Bt doctor expansion v1 — build the READY spec :feature:dotfiles:bluetooth: +CLOSED: [2026-07-12 Sun] +:PROPERTIES: +:SPEC_ID: 3d4d61c4-e5df-44e9-b8e0-40b31452c3f7 +:END: +Build the [[file:docs/specs/2026-07-11-bt-doctor-expansion-spec.org][bt doctor expansion]] (IMPLEMENTED). Adds a dmesg firmware-hint probe (names the missing blob on a no-adapter fault) and a boot-enablement probe (catches an adapter disabled at boot) to the shipped bt doctor (=~/.dotfiles/bluetooth/=). Archsetup owns the dotfiles work end to end. All phases shipped and fake-verified (d19fdca, f05a9b4, d7d859f); the live reboot-persistence half is on the manual-testing checklist. +*** 2026-07-11 Sat @ 03:06:32 -0500 Built the two read-only probes +New module =bluetooth/src/bt/probes.py= plus =doctor.py= wiring, on dotfiles main (=d19fdca=, pushed). Two reads the diagnose chain never did: =firmware_hint()= scans the current boot's kernel log for per-vendor firmware-load failures (Intel ibt-*.sfi, MediaTek BT_RAM_CODE, Realtek rtl_bt, Broadcom .hcd, Qualcomm QCA), returning the named blob via a bounded =cmd.run(journalctl -k)= that reuses the =doctor.py:84= precedent; =boot_enablement()= reads three boot-persistence signals (bluez AutoEnable from main.conf [Policy], =systemctl is-enabled bluetooth=, whether TLP lists bluetooth in =DEVICES_TO_DISABLE_ON_STARTUP=). =diagnose()= gates the firmware read to the no-adapter branch and the boot read to the soft-blocked/powered-off branch, so a healthy run reads neither; the raw signals ride a new =probes= key that =doctor()= carries into =--json=. Detection only: no verdict, formatter, or repair change (that's Phase 1). Every read degrades to None on an unreadable tool/file, so a probe that can't see never invents a fault. AutoEnable absent/unset reads None, not false, matching bluez's compiled default of true, so only an explicit =AutoEnable=false= is the fault. New env roots for tests (=BT_MAIN_CONF=, =BT_TLP_CONF=, defaulting to absent temp paths in the Sandbox base so no test reads real /etc); =fake-journalctl= branches on =-k=, =fake-systemctl= answers =is-enabled bluetooth=. 125 bt tests, full =make test= green; =/review-code= approved (no Critical/Important; one Minor noting the firmware read also covers the btctl-unavailable branch, harmless). Inbox note sent to dotfiles. +*** 2026-07-11 Sat @ 03:16:59 -0500 Built the firmware-hint Guide verdict +On dotfiles main (=f05a9b4=, pushed). The no-adapter step now names the blob: a new =_no_adapter_step= consults =probes.firmware_hint()= on a genuine no-adapter fault and, on a per-vendor signature match, sets =evidence= to "no Bluetooth adapter found — <Vendor> firmware <blob> failed to load" and =next_action= to "update linux-firmware and reboot", tagged with a new =code="no-adapter-firmware"= so a =--json= consumer can branch without string-matching. A clean log keeps the generic hardware/driver verdict. A Guide, not a repair: the step carries no =repair= action, so =--fix= never touches it, and it needs no privilege model (so Phase 1 lands independently of the shared cross-panel model). A missing =bluetoothctl= (=BtctlError=) short-circuits before the firmware read, so the verdict fires only on a real no-adapter fault, not a broken install — this also tightened Phase 0 (which read the log on both None branches) to the genuine no-adapter case. =_mk= gained a uniform =code= key (default None) added to every diagnose step, mirroring the existing =repair= key; no test asserts an exact step key-set, verified. =format_doctor_human= already renders =evidence=/=next_action=, so no formatter change. 133 bt tests (+8), full =make test= green; =/review-code= clean. Inbox note sent to dotfiles. +*** 2026-07-11 Sat @ 06:48:21 -0500 Built the persistent-power verdict + fix +On dotfiles main (=d7d859f=, pushed). =_powered_step= consumes the bt Phase 0 boot-enablement probe: AutoEnable explicitly false, service disabled at boot, or TLP listing bluetooth → =powered-off-persistent= (code + evidence naming the cause) carrying the =persist-power= repair; absent/unset config reads as auto-enable-on (bluez default) → the plain =power-on=, so healthy machines can't false-positive. The =persist-power= repair fixes only the causes set (AutoEnable=true, =systemctl enable bluetooth=, drop bluetooth from the TLP list), powers the adapter on now, and verifies each cause cleared. Config edits are pure idempotent text transforms in a new =bootconf= module (comments preserved), staged and installed via a fixed-destination =cp= verb so the write needs root but the mutation is unit-testable. =priv.py= gained three narrow verbs (=enable-bluetooth=, =write-main-conf=, =write-tlp-conf=). The repair is Privileged and resolves through =panelkit= before running: can't-elevate degrades to the guide. The bt shim gained =panelkit= on its path. 287 bt tests + 65 make-test suites green, review-code Approve, voice. Inbox note sent. Live half (real reboot persistence) is the VM/manual checklist. +*** 2026-07-12 Sun @ 09:14:00 -0500 Flipped the bt spec to IMPLEMENTED and logged the vNext items +Spec status heading now IMPLEMENTED (dated history line + Status mirror); all four phase headings DONE. vNext items (stale-bond signature, connection-parameter hints, bt-audio-profile expansion) logged as the "Bt doctor vNext" task. The live half (real reboot persistence) remains on the manual-testing checklist — findings there come back as bugs. +** DONE [#B] One copy + close control pair on every output wall :feature:dotfiles:solo: +CLOSED: [2026-07-12 Sun] +Resolved in dotfiles =dccd744=: every wall carries the o-copy/o-clear overlay pair. bluetooth gained the copy key (transcript via the new =viewmodel.step_copy_line=, CLI-shaped); maint traded its header COPY key for the overlay pair, kept HIDE, and its ✕ clears the session log via =PanelModel.wall_clear=; the four hand-rolled wl-copy calls collapsed into =panelkit.clipboard.copy_text= (PANELKIT_WLCOPY test seam), moving maint off the GTK clipboard so its copies survive the panel closing too. 15 new tests (8 clipboard, 4 step_copy_line, 3 wall_clear); full suite 66 green; all four panel smokes run live off-workspace — maint + audio fully pass, net + bt fail only the pre-existing state-word startup race. + +Converge all four instrument-console output walls on the net panel's well controls: a copy glyph and a ✕ close, as an overlay at the top right, hidden until content lands. Craig's call, 2026-07-10, while reviewing the audio doctor's wall: "network panels as standard across all others, make it consistent." + +Where they stand today, no two alike. net has copy + ✕ (=net/src/net/gui.py=, the =o-copy= / =o-clear= overlay). bluetooth has ✕ but no copy. maint has COPY + HIDE as keys in a header row, and no ✕. audio just gained copy + ✕ (dotfiles =bd33440=). + +Work: give bluetooth a copy key, give maint the overlay pair, and lift the four hand-rolled =_copy_output= implementations into one shared helper rather than a fifth copy. maint keeps HIDE alongside close, because its wall is a persistent session action log you collapse and keep, where net's, bt's, and audio's are per-run results you dismiss. + +Copy text is per panel but one rule: it pastes as that panel's CLI prints, so the paste lines up with the terminal a user is already looking at. audio's =viewmodel.wall_copy_text()= is the worked example. + +Consistent with the 2026-07-07 scope note on the sibling task above: net's compact glyph overlay is what standardizes, not maint's wide COPY-key header row. +** DONE [#C] Org-capture float popup grows too large :bug:hyprland:quick:solo: +CLOSED: [2026-07-14 Tue] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-13 +:END: +Craig answered the pre-flight (2026-07-14): cap at 120 wide, height proportional. Applied as 120 Emacs columns (11 px/col measured from the live daemon) = 1320 px wide, height 653 px from the old rule's aspect. Ratio's size rule shrank from the 1892x936 scratchpad match and both hosts gained a max_size growth cap (the field is max_size — bare "maxsize" is invalid and hyprctl reload won't say so; check hyprctl configerrors). Verified live: config clean, a probe window floats at exactly 1320x653. Dotfiles 9c4dc2f. +** DONE [#C] Panel smoke: faceplate state-word assertion fails on the live compositor :bug:dotfiles:test: +CLOSED: [2026-07-14 Tue] +Diagnosed and fixed within the session: not a race — test drift. Dotfiles b581d5d (2026-07-05) made the faceplate word the static subsystem identity (NETWORKING / BLUETOOTH / AUDIO) and updated the audio smoke, but the net and bt smokes kept asserting the retired live-state words and had failed on every run since. Both now assert the identity word like audio's (dotfiles 32cd99f); both smokes run RESULT: OK end to end, which also green-gates the doctor-streaming change. +** DONE [#C] Realtime lamp output for the net + bt doctors :feature:solo: +CLOSED: [2026-07-14 Tue] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-09 +:END: +Shipped in dotfiles 0318a91. Both doctors stream: diagnose() emits each step as it completes (bt streams the first diagnosis only — the fix loop's re-diagnoses would replay the chain), and a repair's row goes up amber at attempt start and settles green/red with narration + evidence at completion, so the lamp blinks for the repair's real duration. The 3.5-entry height cap turned out to already be in both wells (it landed with the doctor expansions), so only the streaming half needed building. 5 new tests across net + bt; both suites green; AT-SPI smokes at parity with HEAD (one pre-existing state-word failure, filed separately). +Retrofit the net doctor (=~/.dotfiles/net/src/net/doctor.py=) and bluetooth doctor (=~/.dotfiles/bluetooth/src/bt/doctor.py=) to stream results as a live output wall — one lamp per escalation step, amber while running, green on success, red on failure — instead of a final summary. Matches the maintenance-console doctor design (see [[file:docs/design/maintenance-console-design-ideas.org][maintenance-console-design-ideas.org]], "Doctor = live output wall"). Goal: every doctor in the system reads the same way. Both doctors already step through an escalation chain re-probing after each, so the steps are natural lamp boundaries. + +Scope note (Craig, 2026-07-07): realtime lamp *behavior* only. The maintenance console's wider results-wall layout (date+time stamp column, COPY, persistent history) does NOT backport — the net/bt panels are ~400px wide and lack the horizontal real estate. Their existing output wells keep their compact layout; this task just makes them stream live. + +Addendum (Craig, 2026-07-07): DO backport the 3.5-entry height convention — every panel's output well caps at 3.5 visible entries, the half-visible entry being the scroll cue, with the dark slate-on-black scrollbar. Layout stays compact per above; only the height cap + scroll affordance carries over. +** DONE [#B] Absorb the clock-panel project into the dotfiles :feature:waybar:dotfiles: +CLOSED: [2026-07-18 Sat] +Absorbed into =~/.dotfiles= (commit 3fab11d): package =clock/src/clock/= (renamed from clock_panel), the six PNG watchface layers packaged inside the module at =clock/src/clock/assets/=, a stowed =clock-panel= shell shim (LD_PRELOADs gtk4-layer-shell), waybar left-click now =clock-panel toggle= with the absolute path dropped, tests converted pytest→unittest into =tests/clock/= plus an asset-load guard. Kept the layer-shell overlay and the socket toggle. The standalone repo is archived (ARCHIVED.md), kept for its design history. Verified live: the bar click renders the polished watchface. +** DONE [#A] Velox boot recovery — no kernel in BE :bug:velox:zfs: +CLOSED: [2026-07-19 Sun] +Recovered. Velox boots linux-lts 6.18.38 and is back on the tailnet (up 1d+, /boot holds initramfs-linux-lts.img). The pre-pacman ZFS snapshot rollback restored the kernel from the ZBM recovery shell. +Velox won't boot: ZBM prompts for the passphrase, unlocks, then reports no bootable environment with a kernel. Cause: an interrupted kernel =-Syu= removed the old kernel and never installed the new one — /mnt/be/boot (from zroot/ROOT/default) holds ONLY intel-ucode.img; vmlinuz-linux + both initramfs are gone. /boot lives inside zroot/ROOT/default (no separate boot dataset), so root-dataset snapshots capture it. + +Status 2026-07-15: a first rollback attempt did NOT fix it (square zero after reboot) — suspected typo in the snapshot name, so the rollback likely errored and did nothing. NOT verified. Next session: verify state in the ZBM recovery shell BEFORE any reboot. + +Recovery lever: the pre-pacman ZFS snapshot hook (live on velox since 2026-06-29) snapshots zroot/ROOT/default@pre-pacman_<ts> before every pacman transaction. The newest =pre-pacman_<ts>= predating the failed upgrade holds the intact old kernel — roll back to it. + +Morning steps (Craig at velox ZBM → recovery shell, Ctrl+R): +#+begin_src sh +# 1. pool writable + key loaded +zpool get readonly zroot +zfs get -H -o value keystatus zroot/ROOT/default +# if readonly=on: zpool export zroot && zpool import -f -N zroot +# if keystatus=unavailable: zfs load-key zroot + +# 2. list snapshots — COPY THE EXACT NAME (the typo bit here last time) +zfs list -t snapshot -o name,creation zroot/ROOT/default | grep pre-pacman + +# 3. see current /boot state (read-only mount) +umount /mnt/be 2>/dev/null; mkdir -p /mnt/be +mount -t zfs -o zfsutil,ro zroot/ROOT/default /mnt/be +ls -la /mnt/be/boot + +# 4. if /boot still shows only intel-ucode.img: redo rollback with the exact name +umount /mnt/be 2>/dev/null +zfs rollback -r zroot/ROOT/default@pre-pacman_<EXACT-TS> # -r, NOT -R + +# 5. VERIFY before reboot — remount RO, confirm the kernel is back +mount -t zfs -o zfsutil,ro zroot/ROOT/default /mnt/be +ls -la /mnt/be/boot # MUST show vmlinuz-linux + initramfs-linux.img +umount /mnt/be + +# 6. only once /boot shows a kernel: +zpool export zroot && reboot +#+end_src +Scope: only zroot/ROOT/default reverts; /home, /var, /media are separate datasets, untouched. After boot: =pacman -Syu= attended, confirm /boot holds vmlinuz-linux + initramfs before any shutdown. Full diagnosis: =inbox/PROCESSED-2026-07-15-0002-from-.emacs.d-velox-boot-failure-handoff.org=; ZBM photo: =inbox/PROCESSED-2026-07-15-0002-from-.emacs.d-PXL_20260715_043758976.jpg= (local on ratio; inbox is gitignored). +** DONE [#C] Restore date-format scrolling on the waybar date module :feature:waybar:dotfiles:quick: +CLOSED: [2026-07-19 Sun] +Shipped dotfiles 9dfe082: date-only ring (ordinal/full/longdate), on-scroll rewired, layout guard flipped. UTC/time stay on the time module. +Date and time are separate fixed-position controls. The time display cycles its +own formats, including UTC; the date/calendar control cycles date-only formats +and never displays a second time. Implement the dedicated format rings, +tooltip behavior, and tests together in the dotfiles Waybar configuration. +Reference material for the compact clock/chronograph treatment is filed in +[[file:working/clock-display-references/][working/clock-display-references/]]. + +*** 2026-07-19 Sun @ 04:36:26 -0500 Folded clock-panel interaction direction +The clock-panel handoff settled the prior open question: UTC belongs only to +the time ring, while the date ring is date-only. The existing task is therefore +a focused follow-up, not a two-line restoration of the old combined ring. +** DONE [#C] Notification sound loudness :chore:audio:quick:solo: +CLOSED: [2026-07-19 Sun] +Shipped dotfiles 808ca23: NOTIFY_VOLUME default 65536->39322 (0.6 gain) in both notify copies. +Reduce notification-sound playback loudness by 40% (0.6 gain, approximately +-4.4 dB). Change the =NOTIFY_VOLUME= playback control rather than re-encoding +the normalized sound files; verify each notification type still plays clearly. +** DONE [#C] Show the active wired interface in the Waybar network module :feature:waybar:network: +CLOSED: [2026-07-19 Sun] +Shipped dotfiles 22867f9: select_device prefers connected wifi -> connected ethernet -> wifi fallback, so a live cable shows the wired glyph+iface instead of Offline. +When Ethernet is active, replace the offline-WiFi presentation with the wired +interface glyph and interface name. +** DONE [#C] Let the clock panel dismiss itself on right click :feature:clock:waybar: +CLOSED: [2026-07-19 Sun] +Shipped dotfiles fc9a2b7: secondary-button gesture -> ClockApplication._dismiss hides the open panel. Live-verified with Craig 2026-07-19. +Make a right click inside the open clock panel toggle it closed. Preserve left +click for its established interaction; the Waybar time module remains the +explicit way to reopen the panel. +** DONE [#C] Make the WiFi toggle connect the best available profile :feature:network: +CLOSED: [2026-07-19 Sun] +Shipped dotfiles 9105361: manage.wifi_radio -> _connect_best_saved activates the strongest in-range saved profile on enable; nothing in range falls back to NM autoconnect. +When enabling WiFi, automatically connect to the highest-priority available +saved network instead of requiring a panel selection first. +** DONE [#A] Tracked WireGuard private keys in repo — public leak, resolved :bug:security:network: +CLOSED: [2026-07-20 Mon] +Confirmed a live public leak, not just at-risk: git.cjennings.net runs cgit (scan-path=/var/git), so archsetup.git was anonymously cloneable over https. An unauthenticated clone pulled the configs with intact PrivateKeys. Exposed 2026-07-05 (c7b7d16) to 2026-07-20. Regraded to P1/[#A] (public credential exposure, severity-alone carve-out) from the initial [#B]. +Scope was wider than first found: the current 3 configs (assets/wireguard-config/wg-*.conf) plus 7 older ones at the pre-reorg path assets/wireguard/ (switzerland x2, USCALA/USCASF/USDC/USGAAT/USNY) — 10 config files, all with real keys. +Resolution: Craig expired all the Proton WireGuard configs (keys dead). Purged all 10 from every commit with git filter-repo, force-pushed main + v0.5, and ran git gc --prune=now on the server bare repo. Verified via anonymous clone: zero real-key blobs reachable, all old exposed commits gone. Stopped tracking plaintext (gitignore + README, out-of-band configs only). +Follow-ups filed below: harden cgit exposure; installer no longer ships configs. +** DONE [#C] Installer chpasswd unguarded — unloggable primary user :bug:solo:quick: +CLOSED: [2026-07-20 Mon] +Fixed (fa3135a): extracted set_user_password, which guards the chpasswd with error_fatal so a failure aborts loudly instead of silently leaving no password. Fake-chpasswd test pins the guard fires on failure and stays quiet on success. +Grading: Major severity (fresh system's primary user can't log in) x rare edge case (chpasswd seldom fails) = P3 = [#C]. +archsetup:1168 runs =echo "$user:$pass" | chpasswd= with no guard, then unsets the password next line; set -e is off (line 21), so a silent failure leaves no password and no log entry. Fix: guard with error_fatal (report + "set it by hand: passwd $user") before unsetting. See findings doc (S2). +** DONE [#C] Installer nvme early module never built into initramfs :bug:solo: +CLOSED: [2026-07-20 Mon] +Fixed in e0d22bd: extracted ensure_nvme_early_module, which rebuilds the initramfs whenever it changed the conf (regardless of ZFS root) and scopes the presence check to the MODULES line. TDD via tests/installer-steps/test_ensure_nvme_early_module.py. +Grading: Minor severity (module autoload still boots the system) x most-machines (all Craig's ZFS-root boxes) = P3 = [#C]. +archsetup:2910 writes MODULES=(nvme) but the only mkinitcpio -P in boot_ux runs =if ! is_zfs_root=, so on ZFS-root non-Framework machines the early-load hardening is never compiled in. Also archsetup:2918 greps the whole file for "nvme" (not the MODULES line). Fix: rebuild initramfs after the MODULES edit regardless of ZFS; scope the presence grep to =^MODULES=(=. See findings doc (S3). +** DONE [#C] Installer disk-space pre-flight check is fragile :bug:solo:quick: +CLOSED: [2026-07-20 Mon] +Fixed in aef074f: extracted check_disk_space using df -P (wrap-safe) and a KB comparison (no truncation bias); non-numeric df output falls back to zero so a malformed read aborts loudly. TDD via tests/installer-steps/test_check_disk_space.py. +Grading: Major severity (aborts a valid install) x some (df wraps long device names on a live ISO / device-mapper root) = P3 = [#C]. +archsetup:487 parses =df / | awk 'NR==2'=, which reads the device-name line (empty $4 -> 0 GB) when df wraps; archsetup:488 also integer-truncates the GB compare against the 20 GB floor. Fix: =df -P /= (single-line) or =df --output=avail=; compare in KB to avoid the rounding bias. See findings doc (S1). +** DONE [#C] Installer run_step state + exit-code handling :bug:solo: +CLOSED: [2026-07-20 Mon] +Fixed in 6de55d2: run_step records the state marker whenever the step function returns (a return past error_fatal's exit means only a non-fatal warning is left), added local to run_step/show_status, and captured pacman's real exit in the refresh loop. TDD via tests/installer-steps/test_run_step.py. +Grading: Major severity (resume re-runs steps and can abort on a survivable warning) x some (a step whose last action is a non-fatal failure) = P3 = [#C]. +archsetup:298 marks a step complete only when its function returns 0, but error_warn/run_task return 1, so a non-fatal-failing step never writes its marker and re-runs on resume. Also archsetup:1034 reports =$?= of the =false= test, not pacman's real exit code; and run_step locals (290/318) leak to global scope. Fix: step functions =return 0= explicitly (or gate run_step on a per-step error flag); capture the real exit code; add =local=. See findings doc (S1). +** DONE [#C] cmail password decrypted world-readable before chmod :bug:security:solo:quick:cmail: +CLOSED: [2026-07-20 Mon] +Already fixed in dffecf5 (before this session): decrypt_to_secure wraps the gpg decrypt in a 0077-umask subshell so the file is 0600 from creation, with tests/cmail/ verifying the umask at write time. The task was stale; verified green and closed. +Grading: security carve-out — brief local plaintext exposure of the mail password, requires a concurrent local shell during install; narrow window = low severity = P3 = [#C]. +scripts/cmail-setup-finish.sh:52 gpg-decrypts to ~/.config/.cmailpass at the process umask (often 0644), then chmod 600 on the next line. Fix: =(umask 077; gpg ... --output ...)= or decrypt to a mktemp 0600 file and mv into place (mirror the import-wireguard mktemp -d 0700 pattern). See findings doc (S4). +** DONE [#C] Installer sudoers.pacnew blind copy risks lockout :bug:solo:quick: +CLOSED: [2026-07-20 Mon] +Fixed in c80e855: extracted replace_sudoers_pacnew, which runs visudo -cf on the pacnew and only copies a validated file (warns and keeps the working sudoers otherwise). TDD via tests/installer-steps/test_replace_sudoers_pacnew.py. +Grading: Major severity (a malformed sudoers locks out privilege escalation) x rare edge case = P3 = [#C]. +archsetup:1146 does =[ -f /etc/sudoers.pacnew ] && cp /etc/sudoers.pacnew /etc/sudoers= with no validation, right before the NOPASSWD rule at 1183. Fix: =visudo -cf /etc/sudoers.pacnew && cp ... || error_warn=. See findings doc (S2). +** DONE [#C] WireGuard import leaves full-tunnel VPN live on failure :bug:solo:network: +CLOSED: [2026-07-20 Mon] +Fixed in 36daf76: the down now runs before the rename modify (targets the stable UUID), so a failed modify under set -e can't leave a live full-tunnel VPN. Added a connection-down case to fake-nmcli and two ordering tests. +Grading: Major severity (all traffic silently routed through Proton until manual cleanup) x rare (nmcli modify failure) = P3 = [#C]. +scripts/import-wireguard-configs.sh:51-62 imports (which brings the 0.0.0.0/0 tunnel up), renames, then deactivates; under set -e a failed modify aborts before the down, leaving the tunnel live. Fix: bring the connection down right after parsing the UUID, before the rename. See findings doc (S4). +** DONE [#C] net-scenarios diagnose failure exits green :bug:test:solo: +CLOSED: [2026-07-20 Mon] +Fixed in cf211cd: a diagnose miss sets a per-scenario rc carried to the subshell exit, so the run fails honestly while still running fix + assert. New harness at tests/net-scenarios/ drives the real script with stubbed ssh/rsync/jq. +Grading: Major severity (a net-doctor diagnosis regression is reported as a passing run — false green on a diagnostic tool) x rare edge case (only when a diagnosis regresses and this first-draft harness is relied on) = P3 = [#C]. +scripts/testing/run-net-scenarios.sh:103 — the scenario_diagnose_expect else-branch prints fail "...diagnose did NOT name it" but never forces a non-zero subshell exit, so ( ... ) || fails=... leaves fails unincremented and the script prints "all scenarios passed" + exit 0. Fix: exit 1 in that branch like the other two checks. See findings doc (S5). +** DONE [#C] pacman-hook-order test is a tautology :test:solo:quick: +CLOSED: [2026-07-20 Mon] +Fixed in 1b7236b: the test now extracts the hook filenames the installer writes and compares them against the stock 60-mkinitcpio-remove name (pacman's filename ordering is the real invariant, not source position). Mutation-verified: a 05->70 rename fails the new compare where the old literal compare stayed true. +Grading: Major severity (guards boot-critical hook ordering — a reorder that removes the current initramfs without a rebuild is unbootable, and this test would ship it green) x rare (hook order rarely changes) = P3 = [#C]. +tests/installer-steps/test_pacman_hook_order.py:20 — the two assertLess calls compare string literals ("05..." < "60..."), a constant ASCII fact always true regardless of file content; the ordering the test exists to protect is never measured. Only the assertIn presence checks do real work. Fix: assert on positions — text.index("05-zfs-snapshot.hook") < text.index("60-mkinitcpio-remove.hook") (and the guard hook). See findings doc (S6). +** DONE [#C] Add inetutils to install base :feature:solo:quick:network: +CLOSED: [2026-07-20 Mon] +Already done in 1115543 (earlier today): inetutils sits in install_required_software, with tests/installer-steps/test_required_software.py pinning it (test_installs_inetutils_for_ftp, green). The task was stale; verified and closed. The next full VM run covers the install-path verification. +Original context: TRAMP's /ftp: method needs =/usr/bin/ftp= (GNU inetutils); dirvish has an FTP quick-access entry. Installed manually on ratio 2026-07-14. From .emacs.d handoff 2026-07-14-1751. +** DONE [#D] Installer resume-idempotency cluster :bug:solo: +CLOSED: [2026-07-20 Mon] +Fixed in 8917f2f: extracted crontab_append_once (dedup guard), zfs_scrub_timer_units (one timer per pool, warn on none instead of @.timer), and enable_user_service (wants-symlink; gamemode now uses it and syncthing folds into the shared helper). TDD via tests/installer-steps/test_idempotency_cluster.py. +Grading: Minor severity x rare edge case (re-run after a mid-step failure) = P4 = [#D]. Group of small non-idempotent / wrong-target spots. +crontab log-cleanup line duplicates on resume (archsetup:1713 — guard on absence); zfs scrub timer picks an arbitrary pool via =head -1= and yields =@.timer= when empty (archsetup:1857); gamemode enabled via =systemctl --user= which the script itself documents fails at install time (archsetup:2419 — use the manual wants-symlink like syncthing). See findings doc (S2, S3). +** DONE [#D] Installer unguarded chmod/cp after non-fatal ops :bug:solo:quick: +CLOSED: [2026-07-20 Mon] +Fixed in dd41036: extracted install_executable (guarded cp + chmod +x) for the two zfs scripts; guarded the two hypr-live-update-guard chmods inline with error_warn. TDD via tests/installer-steps/test_install_executable.py. +Grading: Minor severity x rare edge case (only when a preceding non-fatal cp/clone failed) = P4 = [#D]. +With set -e off, unguarded chmod/cp hit missing/partial files silently: hypr-live-update-guard chmods (archsetup:2108/2144), zfs-replicate cp (archsetup:1820) leaving a service with a dead ExecStart, zfs-pre-snapshot cp (archsetup:1943) leaving a broken pacman hook. Fix: wrap each in =(...) >> log 2>&1 || error_warn=. See findings doc (S2, S3). +** DONE [#D] normalize-notify-sounds temp/atomicity can corrupt tracked file :bug:solo:quick: +CLOSED: [2026-07-20 Mon] +Fixed in a29769e: resolves the real target via readlink -f, stages the temp beside it, guards on a non-empty encode, and atomically mv's into place (preserving the stow symlink); an EXIT trap cleans a leaked temp. TDD via tests/normalize-notify/ with fake ffmpeg. +Grading: Minor severity (corrupts a repo-tracked sound file, recoverable via git) x rare (ffmpeg failure/interrupt) = P4 = [#D]. +scripts/normalize-notify-sounds.sh:39-46 has no EXIT trap on the mktemp and does =cat "$tmp" > "$f"= (truncate-first) where $f is a stow symlink into the repo; a zero-byte/failed encode writes a corrupt file. Fix: EXIT trap; =[ -s "$tmp" ]= guard; write $f.tmp and overwrite on success. See findings doc (S4). +** DONE [#D] VM test-framework robustness cluster :bug:test:solo: +CLOSED: [2026-07-20 Mon] +Fixed in 866d327: profile-suffixed PID/monitor/serial paths, kill_qemu reaps-or-polls to death before the snapshot restore, debug-vm uses DISK_PATH, and both runners report an honest ARCHSETUP_COMPLETED marker instead of a fake exit code. TDD via tests/vm-framework/test_vm_utils.py (suffix red->green; kill_qemu as a contract pin). +Grading: Minor severity x rare edge case (each fires only in a narrow test-harness path) = P4 = [#D]. Group of four small framework bugs from the S5 audit. +scripts/testing/debug-vm.sh:49 hardcodes the btrfs base disk, ignoring the profile-correct DISK_PATH from init_vm_paths (FS_PROFILE=zfs boots the wrong base or fatals); lib/vm-utils.sh:284 kill_qemu -9's and deletes the PID file without waiting, so a force-kill restore races the dying qemu's qcow2 lock and silently leaves the base image dirty (fix: wait for the PID); lib/vm-utils.sh:69 leaves PID_FILE/MONITOR_SOCK/SERIAL_LOG un-suffixed so parallel btrfs+zfs runs collide (fix: suffix by FS_PROFILE like DISK_PATH); run-test.sh:287 (and run-test-baremetal.sh:234) reports a completion-marker grep as ARCHSETUP_EXIT_CODE, not the installer's real exit — misleading since the installer runs set -e off and can error then still write the marker (fix: rename + capture the true status). Testinfra remains the real pass/fail backstop. See findings doc (S5). +** DONE [#D] Gallery-widget prototype elisp bugs :bug:design:solo:quick: +CLOSED: [2026-07-20 Mon] +Fixed in 552736e: shared clamp feeds needle + readout (150 renders 100%), explicit cl-lib require, and gallery-widget--source-dir with a default-directory fallback. TDD: 3 new ERT tests (clamp red->green; the other two land as pins since svg.el transitively loads cl-lib). +Grading: Minor severity x rare edge case (out-of-range input / cold byte-compile / interactive re-eval) = P4 = [#D]. Prototype code, all three Minor. +docs/prototypes/gallery-widget.el:139 renders the readout from the unclamped value while the needle clamps 0-100, so at value 150 the needle pins at +60 degrees but the text reads "150%" (fix: clamp once, format both from it); :69 calls cl-loop without (require 'cl-lib) — works only via the autoload cookie, bites on a cold byte-compile (fix: add the require); :29 computes its dir from (or load-file-name buffer-file-name), both nil on interactive re-eval outside a load/file buffer (fix: fall back to default-directory). See findings doc (S7). +** DONE [#D] Audit test-quality cluster (Python + elisp) :test:solo: +CLOSED: [2026-07-20 Mon] +Fixed in 179fbd5 (plus 552736e for the gauge-level clamp test): socket check via find -type s, gen_tokens degenerate case pinned exactly as characterization, tick count as direct occurrences, and write-svg covered. All five items dispositioned. +Grading: no runtime behavior change; test-suite quality. Group of five weak/missing tests from the S6/S7 audit. +scripts/testing/tests/test_desktop.py:96 passes a shell glob to `test -S`, which breaks on zero or multiple sockets (masked today because the test always skips); tests/gallery-tokens/test_gen_tokens.py:181 asserts properties too weak to notice the marker output is garbled (impossible input, so low); tests/gallery-widgets/test-gallery-widget.el:77 counts ticks via split-string + cl-count-if :start 1 (a coincidence of split semantics, not a match count); :47 tests the needle-angle helper's clamp but never the rendered readout at an out-of-range value (exactly why the S7 readout/needle bug ships green — add a gauge-level boundary case); :159 leaves gallery-widget-write-svg uncovered (add a Normal write-to-temp case). See findings doc (S6, S7). +** DONE [#B] Installer GRUB_CMDLINE overwrite drops boot params :bug:solo: +CLOSED: [2026-07-21 Tue] +Fixed in f9da097: update_grub_cmdline merges the current value with archsetup's tokens (existing tokens survive, same-key conflicts resolve to archsetup's value) behind a refuse-to-write safety check, via awk + mv with a backup_system_file first. TDD via tests/installer-steps/test_grub_cmdline.py (8 cases incl. cryptdevice/resume/zfs survival and idempotence). +Grading: Critical severity (unbootable) x some-users-sometimes (machines whose base install set a cryptdevice=/resume=/zfs= cmdline param) = P2 = [#B]. +archsetup:3054 rewrites the whole GRUB_CMDLINE_LINUX_DEFAULT line with a fixed string; nothing re-adds a pre-existing cryptdevice/resume/zfs token, so grub-mkconfig (3059) can bake an unbootable config. Fix: read the current value and append only the missing tokens; assert any pre-existing boot-critical token survives before grub-mkconfig. See [[file:docs/design/2026-07-19-sentry-code-findings.org][sentry code findings]] (S3). +** DONE [#C] Maint status wall copy buttons :feature:maint:dotfiles: +CLOSED: [2026-07-21 Tue] +Shipped in dotfiles 8bc79ba per Craig's calls (one global button, rendered text): COPY on the doctor row serializes every category band via the same card_spec the GUI renders, through panelkit clipboard. TDD tests/maint/test_status_copy.py, full dotfiles make test green, inbox note sent. Live check pending: open the maint panel, press COPY, paste. +Craig's roam capture 2026-07-20, routed via .emacs.d sentry inbox-zero as archsetup-owned UI work. Dotfiles maint panel work; archsetup drives it end-to-end per the standing rule. +** DONE [#B] Build: desktop-settings panel :feature:hyprland:dotfiles: +CLOSED: [2026-07-22 Wed] +:PROPERTIES: +:SPEC_ID: d6bb1e73-ec90-4327-85ee-bfa762da5bce +:END: +The GTK build of the desktop-settings panel per the spec (docs/specs/2026-07-02-desktop-settings-panel-spec.org, DOING; normative reference: prototype 37). Work happens in dotfiles settings/ — archsetup drives the lifecycle. Two non-blocking build-time picks live in the spec's Review findings (wallpaper setter tool; store location/format) — decide in phase 1 and record there. +*** 2026-07-22 Wed @ 13:14:01 -0500 Built the backings engine (phase 1) — dotfiles 7a15237 +Landed as dotfiles settings/src/settings (10 modules) + tests/settings (118 tests against fake binaries, auto-discovered by make test — 81 suites green). Covers brightness/kbd (5% floor, x10 drum), toggles (dim, pointer cycle via toggle-touchpad, caffeine), DND class-split (dunst pause level 60, close-all before unpause, alarms punch through live), powerprofilesctl, nightlight (resident gammastep), hypridle.conf renderer + symlink-safe write + caffeine-respecting reload + hyprlock grace, suntimes (pure NOAA math), and the wallpaper engine (awww/mpvpaper/projector adapters, galleries, random draw, atomic JSON store). All three build-time picks recorded as DONE findings in the spec (setter=awww, store=state.json, nightlight=gammastep). Handoff note in ~/.dotfiles/inbox/. +*** 2026-07-22 Wed @ 15:26:44 -0500 Built the presenters (phase 2) — dotfiles 5172289 +Three GTK-free models per prototype 37, all at 100% line coverage (tests/settings/test_presenters.py, 100 tests; full repo suite green before and after). programs.py: the matrix — eight complete programs (Craig's four factory scenes drafted here per the pre-flight pick, slots 1-4 first-class), pin rows + power radio row, activate returns the full sets, member writes return apply/updated with active-is-live surviving. bench.py: drum mapping (screen never reads 0, floor 5%; kbd floors at 0), idle rail order clamping between enabled neighbors, park/unpark with re-clamp, caffeine bypass, view-state builder tolerant of no-backlight None. channels.py: the eight-channel bank, per-mode sources visibility, alpha/recency sort (unlabeled last), the shared mint/edit/delete grammar for pairs/sets/colors (press arm-cycle, two-picture set minimum, dup rejection, selection clamping), sources guardrails, interval wheel, previews. Handoff note in ~/.dotfiles/inbox/. +*** 2026-07-22 Wed @ 16:03:52 -0500 Ported prototype 37's instruments to GTK (phase 3) — dotfiles 33d82eb +The panel renders P37 end to end. New instruments.py carries the three Cairo instruments as clock-free humble objects: ProgramMatrix (glyph/numbered heads over jewel pins + CPU POWER paper letter wheels), DrumRoller (paper drums, drag-to-set, dimmed n/a on no-backlight machines), TripDial (sqrt 300° scale, colored stage tabs, OFF-notch parking, exact-minutes drag counter, BYPASSED · CAFFEINE stamp, bottom legend). gui.py rebuilt to P37's layout with the wallpaper sub-view: channel bank with drawn faces, minted pair/color/set trays (alpha/time sort, edit/delete chip feet), the three presses (pair arm-cycle, color picker, set press + interval wheel), sources with a folder picker. New GTK-free glue all unit-tested (test_panel_glue.py, 33 tests): dial geometry in bench, matrix/idle/wallpaper wiring in panel, presenter-vocabulary channels (pair/solid/random-from-set) in wallpaper.apply. AT-SPI smoke (make test-panel-settings) drives the real wiring against faked backings + a sandboxed store, pinned to its own child pid so it can never fire a live panel's backings. Visually verified on a headless output against P37 captures (main + pair/single/solid/random). Adaptations recorded in the handoff: five-stage dial (WATCH gets its own green — the engine runs watch separately, P37 merged the label), DESKTOP_SETTINGS_START_VIEW test seam. Full suite 84 suites green; window rule widened for the 540px panel. Handoff note in ~/.dotfiles/inbox/. +*** 2026-07-22 Wed @ 16:47:54 -0500 Integrated phase 4 — dotfiles 680b50d +Bar consolidation had landed early (74f723e); this pass shipped the rest. settings-project hosts the watch/clock/world channels as HTML faces (settings/faces/) on a gtk-layer-shell background window over WebKit2 — all three visually verified on a headless output, world reading the waybar worldclock roster via query param. settings-watch is the hypridle watch-stage host: throwaway-profile chrome kiosk that reveals only after its window maps behind the lock and relocks before teardown — a failed face degrades to the plain lock, never a bare desktop (unlocked lifecycle verified live; the locked swap goes to the manual checklist). Sun-pair location reads whereami live per transition with last-good cache in state.json (verified live: 9.5s first beat, New Orleans coords, Gogh day side applied); desktop-settings-tick.timer (2 min, enabled on ratio, added to the installer) drives flips and random draws — 23ms no-op beats. dunstrc history_length 100 protects held alarms (full DND cycle verified against live dunst; wtimer alarms already CRITICAL via the notify wrapper, no promotion rule needed). Live hypridle rewrite verified — five-stage regime rendered through the stow symlink, caffeine respected (found engaged, daemon correctly left stopped). Refresh signals needed no rewiring (touchpad signals itself via toggle-touchpad). 45 new tests; suite 84 suites green; smoke 13/13. Handoff note in ~/.dotfiles/inbox/. +Velox one-time steps (sync doesn't carry): mpvpaper (AUR), optionally power-profiles-daemon (service off), and systemctl --user enable --now desktop-settings-tick.timer. +*** 2026-07-22 Wed @ 17:05:58 -0500 Landed the 17-point end-to-end pass — dotfiles 9038eee +Prototype 37's 17-point suite re-derived against the real panel (the original Playwright script wasn't preserved; the functional surface in the spec's Final prototype section is the source). tests/settings/panel_e2e.py + run-panel-e2e.sh + =make test-panel-e2e=: points 1-14 drive the running panel over AT-SPI (program recall with per-backing verification across FOCUS/BATTERY/slot1, pointer console keys, all eight wallpaper channels including projected watch/world stop/start ordering, close); points 15-17 cover the Cairo instruments (drums, tripper dial clamp/park/render/reload, matrix pins + letter wheels with active-is-live) at the backing layer, since AT-SPI can't reach a DrawingArea's hit-tests. Same safety posture as the smoke: sandboxed store, faked backings, pid-pinned a11y node. 17/17 green on ratio's live compositor; full suite 85 green; smoke 13/13; ruff clean. The drag gestures go to the manual checklist below. Handoff note in ~/.dotfiles/inbox/. +*** 2026-07-22 Wed @ 17:05:58 -0500 Flipped the spec to IMPLEMENTED +docs/specs/2026-07-02-desktop-settings-panel-spec.org DOING → IMPLEMENTED with a dated history line naming the shipping commits (dotfiles 7a15237 / 74f723e / 5172289 / 33d82eb / 680b50d / 9038eee) and the verification evidence (85 suites, smoke 13/13, e2e 17/17). The four panel drag-gesture checks and the locked-path night-watch swap live under "Manual testing and validation" — human-eye checks, not implementation blockers. +** CANCELLED [#B] Hyprland layoutmsg crash — bad_variant_access (upstream) :bug:hyprland: +CLOSED: [2026-07-21 Tue] +Dropped 2026-07-21 (Craig's call) — not tracking the upstream report. The crash evidence (both reports + tmpfs log excerpts) and the voice-passed issue draft stay preserved in [[file:working/hyprland-layoutmsg-crash/][working/hyprland-layoutmsg-crash/]] if it recurs and is worth reviving. +Grading: Critical severity (SIGSEGV kills the whole desktop session; every GUI app's unsaved state lost) x rare edge case (twice in ~4.5 months: 2026-03-07 on v0.54.1, 2026-07-20 on v0.55.4) = P2 = [#B]. Upstream Hyprland bug, not this repo's code — the task tracks reporting it and picking up the fix. +A layoutmsg mfact dispatch (layout-resize, mod+H/L) throws std::bad_variant_access inside Layout::CAlgorithm::layoutMsg, uncaught, SIGSEGV. Both crashes fired from the layout-resize mfact path (keycode 104 shrink today, 108 grow in March). Layout at crash was master and the identical mfact had worked seconds earlier; the pre-crash window held monocle<->master toggles, two window closes dropping focus to "[Window nullptr]", and togglefloating x2. Monocle is a registered v0.55 layout (log shows graceful "Unknown monocle layoutmsg" rejects), so the config is not at fault; related edges are guarded ("mfact -> no window") while this path misses its variant guard. Repo has no newer build (0.55.4-1 installed and repo). +Evidence preserved in [[file:working/hyprland-layoutmsg-crash/][working/hyprland-layoutmsg-crash/]] (both crash reports + excerpts from the tmpfs session log, extracted before reboot loses it). +Next: Craig posts the issue himself (2026-07-20 decision) — the voice-passed draft is [[file:working/hyprland-layoutmsg-crash/issue-draft.md][issue-draft.md]], with both crash reports and the log excerpts beside it for attaching. Watch the repo for a fixed release and close on confirmation. The layout-resize script guard was declined (a script can't observe the internal desync). +** DONE [#C] WireGuard import is now config-less — decide feature fate :feature:network: +CLOSED: [2026-07-21 Tue] +Decided 2026-07-21 (Craig): KEEP the import feature. The out-of-band flow is already in place — =assets/wireguard-config/= carries a README documenting "drop plaintext =*.conf= locally at install time (gitignored); ship encrypted =*.conf.gpg= to track", and its =.gitignore= enforces it (=*.conf= blocked, =!*.conf.gpg= allowed). The script also already no-ops gracefully on an empty dir (=shopt -s nullglob= + a =found= flag), so nothing ships and nothing errors when no configs are present. Nothing to build; the fate decision was the whole task. +scripts/import-wireguard-configs.sh reads assets/wireguard-config/*.conf, but no configs ship in the repo anymore (removed as a public-leak fix; .gitignore blocks plaintext). +** DONE [#C] Dupre theme waybar.css drifted from live style.css :bug:dotfiles:waybar: +CLOSED: [2026-07-21 Tue] +Fixed in dotfiles 3e4e7ff (2026-07-20, "fix(theme): sync dupre waybar.css with the live weather rules") — dupre/waybar.css is byte-identical to live again, restoring the =#custom-weather= selectors/hover/gold divider, so =tests/theme-css= is green. +Grading: Minor severity (cosmetic, reverts only on a theme switch) × rare edge case (dupre is already the active theme) = P4 = [#D] on user impact, bumped to [#C] because the dotfiles =make test= stays RED until synced, poisoning the green baseline for every future commit. +The weather-kit work added =#custom-weather= selectors to =hyprland/.config/waybar/style.css= but never mirrored them into =hyprland/.config/themes/dupre/waybar.css=. =tests/theme-css= asserts the two files are identical (set-theme copies the theme file over the live one), so switching to dupre would silently revert the weather chip styling. Fix: sync the theme file to live. Pre-existing; found 2026-07-19 during an unrelated commit's green-baseline run. +** DONE [#B] Dotfiles tests leak state across files :bug:test:dotfiles:solo: +CLOSED: [2026-07-23 Thu] +Resolved 2026-07-23 as dotfiles =c333598=. The polluter was =tests/weather/test_weather.py=, and it accounted for all 38 failures on its own. + +The mechanism was not the env leak the body below guessed at — tests/weather never writes =os.environ=. Its whereami fake did =weather.subprocess.run = ...= on a freshly-loaded module object. The fresh module isolated the weather code, but =weather.subprocess= is the one shared stdlib module object every module in the process holds, so the assignment replaced =subprocess.run= process-wide and never restored it. Every later test file got weather's fake result back from =subprocess.run=; the tell was wtimer asserting on =r.returncode= and getting "'R' object has no attribute 'returncode'", where =R= is weather's fake result class. + +Triage: TEST HYGIENE, not production global state. The weather script reads env at import and never writes, so no long-lived-process caching defect sits behind it. A scan for the same pattern (patching a stdlib module attribute reached through another module's namespace) finds exactly one instance in the suite — the three other =setattr= sites all snapshot and restore. So the planned shared env helper across 28 files was aimed at the wrong target and wasn't needed. + +Fix: rebind the loaded module's own =subprocess= name to a stub namespace, so nothing outside that module changes and there is nothing to restore. + +Gate: =make test= now runs two gates per the add-don't-replace decision — =test-forked= (one process per file, catches order dependence) and the new =test-shared= (every suite in one process, catches leakage). Built on stdlib unittest rather than pytest, since pytest was only the diagnostic tool and isn't a project dependency. Verified as a real gate, not just green today: with the defect deliberately reintroduced it goes red, and green once restored. A focused test in tests/weather pins the invariant on the culprit as well, because the shared gate alone blames the three victim files. + +Verification: 3500 tests, both gates, exit 0. + +Original finding follows. + +Found 2026-07-23 during the speedrun. =make test= is green, but it runs each test file in its own =python3 -m unittest= process, which hides cross-file state leakage. A single-process whole-tree run (=python3 -m pytest tests/ -p no:randomly=) fails 38: 22 in =tests/wtimer/test_wtimer.py=, 10 in =tests/zoom-web/test_zoom_web.py=, 6 in =tests/wlogout-menu/test_wlogout_menu.py=. + +Not a regression — a worktree at the pre-speedrun commit produces the identical 22/10/6 profile, so this predates tonight's work. Those three files also pass cleanly when run together (170 passed), so the polluter is a fourth file somewhere in the tree that mutates global state (env var, cwd, or a module-level patch) without restoring it. 28 test files write =os.environ= directly. + +Why it matters: the green gate can't see this class of bug, so a real isolation defect — or a genuine failure that only appears under a different order — passes CI silently. Bisect by running the tree with subsets until the polluter is identified (pytest's =-p no:randomly= keeps the order stable while bisecting), fix its cleanup, then decide whether =make test= should gain a single-process pass so the gate covers it. +** DONE [#B] Wallpaper view freezes the panel — thumbnail decode :bug:dotfiles:solo: +CLOSED: [2026-07-23 Thu] +Craig reported 2026-07-23: selecting the wallpaper button freezes the module and the compositor asks whether to kill it. Root cause proven: =_Thumb._draw= decoded each source image with =new_from_file_at_scale= on the GTK main thread. Measured on Craig's 78 wallpapers — a viewport of the 8 largest takes 3.7s, the whole set 13s. That block trips Hyprland's "not responding" watchdog. + +Grading: Critical severity (panel unusable, watchdog kill) × every user every time the wallpaper view opens = P1 = [#A] by the matrix. Held at [#B] because step 1 already shipped and removes the user-visible freeze; the remainder is a latency enhancement, not a showstopper. + +*** 2026-07-23 Thu @ 15:40 Step 1 — async decode (dotfiles f45f321) +Moved the decode to a worker thread via a new =settings/thumbcache.py= (pure, injected decode/scheduler/thread; 6 tests). The thumb shows its dark ground until the pixbuf lands, then redraws. Verified live on a headless output: worst main-loop stall opening the pair view dropped from multi-second to 68ms; the cache filled with 81 decoded pixbufs (the one miss is a .webm, correctly falling back to the ▶ glyph). Full suite 3512, both gates, smoke OK. This alone fixes the reported freeze. + +*** 2026-07-23 Thu @ 16:30 Step 2 — persistent on-disk cache (dotfiles 463cc4f) +Built the persistent layer: =settings/thumbstore.py= decodes each source once to a 512px PNG under =~/.cache/settings/thumbs=, keyed by path + mtime so an edited wallpaper self-invalidates. The hot-path decode reads that PNG and scales in-memory. Warming rides the existing =settings tick= CLI verb (the 2-min timer already runs it), building up to =WARM_PER_BEAT=8= missing thumbnails per beat — best-effort, journals a line on failure, never blocks the wallpaper flip. thumbstore is pure (stat/decode/load/save injected); 10 tests. + +Went with incremental warming (8/beat, ~10 beats to full) as the safe default rather than full-warm-on-change — the per-beat cap is a one-line flip if Craig wants it faster. Measured: hot-path decode of a viewport dropped from 3.7s cold to 47ms warm. No installer change (the tick service already runs =settings tick=); cache lives outside the repo. Full suite 3522, both gates, smoke OK, live panel verified (81 pixbufs render, 48ms worst stall warm). +** DONE [#C] Panel scrollbars too short :bug:dotfiles:quick:solo: +CLOSED: [2026-07-23 Thu] +Shipped 2026-07-23 as dotfiles =0d64837= (22px scrollbar, 16px trough, 14px slider thickness with a 48px floor along the travel axis). Left open by oversight during the speedrun; closing now. + +Follow-on, and my own regression: enlarging the bar to 22px is what made it start covering the thumbnails, because nothing grew the tray to match. Craig reported it the same day ("scrollbars that obscure the images") and it's fixed in =c0ddf57= — the tray now reserves a 22px lane for the bar as a margin on the scrolled box, so the bar sits below the images instead of across them. Measured before: tray 68px, content 68px, a visible 14px bar inside the same 68px. After: tray 90, content 68, bar clear. The lane is a constant under the scrollbar CSS with a note to keep the two in step, since the coupling between bar thickness and tray height is exactly what broke. + +From the roam inbox (Craig, claimed 2026-07-23): all scrollbars need to be much taller than before. The always-visible scrollbars shipped in 7e8eb4a set =min-height: 10px; min-width: 10px= on the slider (=settings/src/settings/gui.py=, the =.dupre-panel scrollbar slider= rule) — that's the floor for a short slider, and the trough itself is thin. Raise both the slider floor and the trough thickness so the bar is comfortably grabbable. Cosmetic × every glance at the wallpaper trays = P3 = [#C]. +** DONE [#C] Velox refresh sweep :chore:maint: +CLOSED: [2026-07-23 Thu] +From the roam inbox (Craig, claimed 2026-07-23): velox needs bringing up to date, the mouse/touchpad module is still there, investigate what else didn't move over. + +Resolved 2026-07-23 by a full sweep over tailscale. The touchpad module was already gone — velox's running waybar (started 01:05, after the reboot) and its tracked config both carry zero =custom/touchpad= entries; what Craig saw was the pre-restow waybar process from before the reboot, and the reboot cleared it. Sweep results: both machines at dotfiles f9b6404 (all three hyprland lock/exit fixes live on velox, config errors clean, =allow_session_lock_restore= reads true); stow restow clean, only the expected skip-worktree files; rulesets pulled to 50fc7ca and =make install= run (agent-text verified working by invoking it — an earlier "MISSING" reading was a PATH artifact of the non-interactive ssh shell, not a real gap); desktop-settings tick timer active; mpvpaper, power-profiles-daemon, gtk4-layer-shell, webkit2gtk all present. + +Genuine remaining differences, all per-machine installs rather than sync failures: =cmail-action=, =gcalcli=, and =playwright= aren't installed on velox, and =obsbot-wb-guard.service= isn't enabled there (the OBSBOT lives on ratio). None block anything; file separately if velox should send mail or drive browser tests. +*** 2026-08-20 Thu @ 09:53:04 -0700 One of those three closed itself; two still stand +=cmail-action= is on both daily drivers now, and nothing did it deliberately. It moved into rulesets at =claude-templates/bin/=, and rulesets' =make install= links that whole directory into =~/.local/bin= at every session start — so velox picked it up on its own. Verified here: the symlink was written 05:44 this morning by this session's own startup, and the tool runs. + +=gcalcli= and =playwright= are still absent on velox, which stays correct until I say velox should send calendar invites or drive browser tests. =obsbot-wb-guard= is still right to be off here; the camera is on ratio. + +Leaving the paragraph above as written rather than striking it (rulesets suggested striking). It is the resolution note of a task closed 2026-07-23 and it was accurate that day. Editing a closed record to match today makes it a worse record, and the useful correction is this dated entry, not a redaction. +** DONE [#C] Weather tooltip sunrise and sunset :feature:waybar:weather:quick:solo: +CLOSED: [2026-07-23 Thu] +Shipped 2026-07-23 as dotfiles =de62e9d=. The two rows sit directly below Humidity in the current-conditions block, rendered in the footer's 12-hour format (=%-I:%M %p=) so the tooltip reads one way throughout. + +Confirmed the no-extra-round-trip premise held: =sunrise,sunset= joined the existing =&daily== block. Split =forecast_url= and =reading_from= out of =fetch= so both the request and the reading are testable without network — that's what let the new cases cover a payload missing the fields. Six tests (Normal/Boundary/Error): row placement and format, a pre-change cache with no sun fields, an unparseable stamp, today's pair picked out of the six-day arrays, and the API omitting them. Reused the existing =_at= helper rather than adding a near-duplicate =_first=. + +Live-verified against the real API: sunrise 6:14 AM, sunset 7:59 PM for today in New Orleans, rendering in the actual tooltip. Full suite 3506 tests, both gates, exit 0. + +Open, not blocking: every other header row carries a glyph (thermometer, droplet, wind arrow) and the sun rows are plain text. The file's glyphs are marked font-confirmed codepoints, and I haven't verified a sunrise/sunset glyph renders rather than showing tofu, so I left them bare. Craig's call. + +From the roam inbox (Craig, claimed 2026-07-23): in the weather module's hover text, the section immediately after the location ends with the current humidity. Add the sunrise and sunset times for the current location directly below it. + +Cheap to source: the module already calls Open-Meteo with a =&daily== block (=common/.local/bin/weather=, the forecast URL around line 336), so =sunrise,sunset= joins that same request with no extra round trip — normalise_daily already parses the daily arrays. Times arrive as local ISO strings; render in Craig's canonical clock format rather than re-deriving one. The settings package's =suntimes.py= (pure NOAA math, no network) stays the offline fallback path if the API field is ever absent — don't duplicate its math here. +** DONE [#C] Maint doctor-row copy button :refactor:maint:quick:solo: +CLOSED: [2026-07-23 Thu] +Shipped 2026-07-23 as dotfiles =761fa5c=, "fix(maint): drop the COPY key from the doctor row" — the key and its orphaned handler removed from =maint/src/maint/gui.py=. =viewmodel.status_copy_text= stays: it's a tested pure serializer and the obvious source if a copy surface returns somewhere better placed. + +Correction to the body below: it describes a per-row button and a separate global one. There is only one COPY key, and it IS the global one Craig added in 8bc79ba two days earlier. He tried it and wanted it gone, so the row now reads DOCTOR · CLEAN UP · REVIEW & FIX. + +From the roam inbox (Craig, claimed 2026-07-23): remove the per-doctor-row copy button (next to REVIEW and FIX) from the maint status wall. The global COPY key (dotfiles 8bc79ba, "one global button copying rendered text") stays the one copy surface — the per-row button turned out to be clutter next to it. +** DONE [#C] WiFi tooltip signal strength :feature:waybar:network: +CLOSED: [2026-07-22 Wed] +From the roam inbox (Craig, claimed 2026-07-22): add signal strength to the WiFi tooltip. + +Resolved 2026-07-22: the tooltip's signal line existed but never fired on ratio — the mt7925 driver leaves /proc/net/wireless empty (legacy WEXT procfs unimplemented), so the dBm read returned None and the bar glyph fell to the weakest tier. Fix in dotfiles net/: an iw-dev-link nl80211 fallback (only spawns when procfs is empty), a signal_percent mapping, and an enriched line — Signal: ▂▄▆█ 100% · -32 dBm (excellent) — bars by band, percent, raw dBm, band word. The bar icon tier fixed itself as a side effect. +** DONE [#B] Desktop-settings dropdown panel :feature:waybar: +CLOSED: [2026-07-22 Wed] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-22 +:END: +Resolved 2026-07-22: shipped end to end via the "Build: desktop-settings panel" task (dotfiles 7a15237 → 9038eee; spec IMPLEMENTED, 85 suites + smoke 13/13 + e2e 17/17). Every open question below got settled in the spec: bar consolidation landed (74f723e), the wallpaper manager became the in-panel sub-view, and the format pickers split into their own sibling spec ([[file:docs/specs/2026-07-19-display-format-single-source-of-truth-spec.org]], DRAFT stub). Remaining human-eye checks live under "Manual testing and validation". + +Original body follows as the record. + +Initial spec written 2026-07-02: [[file:docs/specs/2026-07-02-desktop-settings-panel-spec.org]] (DRAFT — four decisions await Craig's review before build; architecture updated to the net panel's Blueprint/GTK4 stack). + +One waybar dropdown gathering the desktop toggles and sliders into a single settings panel, opened from a gear/settings glyph on the bar. Incorporate: +- *Auto-dim* toggle (the =custom/dim= feature just shipped — fold in here, or keep the standalone indicator and mirror it). +- *Brightness* slider (backlight, via brightnessctl). +- *Keyboard-backlight* brightness slider (brightnessctl on the kbd_backlight class). +- *Mouse* enable/disable toggle — shown only when a mouse is connected. +- *Trackpad* enable/disable toggle — shown only when a trackpad is connected (mirror =toggle-touchpad= / =touchpad-auto=). +- *Idle inhibitor* (the =custom/idle= module that replaced the built-in =idle_inhibitor= 2026-06-24 — toggles the hypridle daemon, state-synced icon). +- *Airplane mode* (the existing =airplane-mode= toggle; laptop-only). + +The conditional rows (mouse, trackpad, airplane) appear only when their hardware/context applies — reuse the laptop/device detection the airplane and touchpad indicators already do. + +Design / open questions (propose before building): +- Panel tech: sliders need a real toolkit (waybar can't host a slider), so a GTK4 + gtk4-layer-shell app like pocketbook is the likely shape. +- Which existing standalone bar modules (dim, touchpad, airplane, idle_inhibitor) collapse INTO this panel vs. stay on the bar as quick-access indicators. Craig's call. + +Implementation notes: a small GTK layer-shell app (mirror pocketbook's structure: src-layout Python package, pytest, Makefile) talking to brightnessctl / hyprctl / the touchpad + airplane helpers. Lives in the dotfiles repo or in-tree like pocketbook. TDD the backing toggle/slider logic. Sizable — worth a design doc first. + +Home handoff 2026-07-19 (inbox, resolving the open "few other things" decision — fold into the spec, close the open decision, extend the controls table, then run spec-review, may flip DRAFT→READY). Ownership: home drives the build (dotfiles settings/), archsetup keeps the canonical spec. Full reconciliation in home docs/design/2026-07-19-desktop-settings-module-brainstorm.org. +- ADD controls: night-light / color temperature; Do Not Disturb / notifications (dunst); lock / suspend quick actions; power profile (performance/balanced/saver); scenes/profiles — one control flipping several toggles at once (Focus, Presentation, Battery-saver, Night). Scenes are the payoff of consolidating everything. +- OUT (record reasons): volume / master-mute stays with the audio panel (no mirror here); theme light/dark goes to the theme-studio task. +- FORMAT PICKERS pulled to their own future sibling spec — time/date/weather format is out of THIS panel. Rationale: format settings live in many programs, so the design problem is a single source of truth for the canonical format. Track a future sibling-spec stub in docs/specs (time/date/weather format single-source-of-truth); Craig thinking it through separately, not started. +- STILL OPEN (spec already flags): wallpaper manager confirmed in scope, but row-that-opens-a-sub-view vs its own sub-spec undecided — resolve at spec-review. +** DONE [#C] Gallery probe: the fader-drag check is flaky :bug:test:design:quick:solo: +CLOSED: [2026-07-23 Thu] +Fixed 2026-07-23. Root cause confirmed rather than suspected: =panel-widget-gallery.html= line 74 sets =html{scroll-behavior:smooth}=, so =scrollIntoView= animates and the fixed 200ms sleep sometimes read =getBoundingClientRect= mid-scroll. The drag then dispatched at stale coordinates, the press missed the fader, and the check reported a dead widget. + +Fix: scroll with =behavior:'instant'=. The probe never needed the animation, so this removes the race instead of waiting it out. Also added a =settledRect= guard (rect stable across two reads AND on-screen) for zoom/column relayout, and a =hits()= assertion that the press actually lands on the fader before the drag goes out. + +Applied to the toggle-click check too — it shares the same fixed-sleep shape, and it failed for this exact reason during the diagnosis, so fixing only the fader would have left half the defect. + +Worth recording: my FIRST fix was wrong and made it worse. Polling until the rect stopped changing returned pre-scroll coordinates every time, because two identical samples are also what you get before the animation starts — an intermittent failure became a consistent one. The new hit-test assertion is what caught it, printing the press point at y=1326 against a 1200px window. That's the argument for asserting the press landed rather than only asserting the readout moved. + +Verified against the measured 1-in-6 failure rate: 8 consecutive runs, all three checks passing, with the press point identical every run (429,480) — deterministic, not lucky. Full probe 96 PASS, 0 FAIL, exit 0. + +=probe.mjs= check 3 ("fader drag tracks at 3x") intermittently reports =level 68 -> level 68=, i.e. the synthetic drag never registers. It has presumably been doing this all along unnoticed, since the suite is normally run once per batch. + +Grading: *Minor* severity (a false FAIL costs a re-run and a few minutes, and never ships a defect) x *most users, frequently* = P3 = =[#C]=. + +Frequency measured 2026-07-16, not estimated: 1 failure in 6 consecutive runs, having already fired twice in about fifteen that afternoon. The first grading guessed "some users, sometimes" (~1 in 10); at ~1 in 6, both people who run this suite hit it most sessions, so the row is "most users, frequently". The letter lands on =[#C]= either way, but the input was wrong and the matrix is only worth anything if its inputs are measured. + +Suspected cause: the check clicks the 3x size chip, calls =scrollIntoView=, waits a fixed 200ms, then reads =getBoundingClientRect= and dispatches the drag against those coordinates. If the zoom relayout or the smooth scroll hasn't settled, the rect is stale and the press lands off the fader — so the drag is a no-op and the readout never moves. The other timing-sensitive checks share the same fixed-sleep shape. + +*Do not fix this by raising the sleep.* That hides the race rather than removing it and leaves the check failing again on a slower run. Wait on the actual condition instead: poll until the rect stops changing between frames, or assert the press landed on the fader before dispatching the drag (the probes' own README already warns that a =find()= miss dispatches into nothing and reports as a widget bug). + +Why it matters beyond the annoyance: a gate that cries wolf gets its real failures ignored, and this suite is the only thing standing between the gallery and a silent regression. + +Recurrences: 2026-07-18 batch-6 gate (first run, passed 3 reruns); 2026-07-18 batch-9 gate (first cold run, =level 68 -> level 68=, passed 2 reruns); 2026-07-21 double-speedrun run (flashed one RED mid-run, passed on rerun). All were a session's first/early probe run — consistent with the stale-rect theory (cold-start relayout settling slower than the fixed 200ms sleep). +** DONE [#C] Dotfiles stow conflicts: first-launch risk + restow directory handling :bug:dotfiles:quick:solo: +CLOSED: [2026-07-23 Thu] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-14 +:END: +Closed 2026-07-23. The last open item was the velox check, and velox is reachable again. Pulled it from =02df01a= to =de62e9d= (clean tree, fast-forward), then ran =make conflicts hyprland=, which dry-runs every tier: "No stow conflicts", exit 0. Its old conflict copy had already cleared against the updated repo, so there was nothing for =make reset= to do. + +Verified the pull is live through the symlinks rather than just present in the repo: =~/.local/bin/weather= resolves into the dotfiles tree and returns today's sun times on velox. + +Note for the record: =make conflicts common= is not a valid invocation — =check-de= rejects it, because common and the host tier are auto-included in the DE-scoped run. =make conflicts hyprland= is the whole check. +*** 2026-07-14 Tue @ 00:51:51 -0500 Ratio calibre check passed; waypaper canonical decided (dark-lion) +Ratio's ~/.config/calibre is a directory symlink into the dotfiles repo (stow folded the whole dir), so the first-launch gap never existed there — check closed. Craig decided dark-lion.jpg is the canonical waypaper wallpaper; the repo config.ini updated from the that-one-up-there.jpg placeholder (the file is skip-worktree volatile, unskipped for the commit and re-flagged). Remaining: when velox is back online, run make conflicts / make reset there so its old conflict copy clears against the updated repo. +*** 2026-07-02 Thu @ 17:30:00 -0400 Shipped the Makefile hardening + first-launch guard (dotfiles 42a82d2) +The solo-able subset landed in the speedrun. =make conflicts <de>= is the loud first-launch guard: dry-runs all tiers, parses all four stow error shapes (plain file conflict, foreign symlink, dir-over-file, and restow's unstow_contents non-directory ERROR), lists each blocker with a directory/foreign-symlink marker, exits 1 when any exist. =make reset= now pre-clears the directory and foreign-symlink blockers =--adopt= aborts atomically on (removals printed; repo version wins per the target's contract), then adopts + git-checkouts as before. =make restow='s overwrite path switched rm -f → rm -rf so directory conflicts clear. 8 sandbox tests drive the real Makefile against a throwaway HOME (44 suites green). Also verified on velox: the whereami and mpd-playlists conflicts noted in this task were already hand-converted 2026-06-29 — =make conflicts hyprland= reports clean live. REMAINING (deferred per Craig's speedrun pre-flight): the waypaper canonical decision (live velox dark-lion.jpg vs repo that-one-up-there.jpg) and the ratio calibre-symlink check (ratio paused). +From the velox calibre incident (2026-06-27, note in ~/.dotfiles/inbox/processed/): calibre was launched before =make stow= ran, wrote its own default config into =~/.config/calibre/=, and silently blocked its own stow — it ran on factory defaults while the rest of common/ stowed fine. General pattern: any GUI app that auto-creates config on first run, launched before stow, blocks its own stow the same way. Velox was repaired by hand (=ln -srf= symlinks byte-identical to =stow --no-folding= output). + +Remaining work (re-graded C 2026-07-02 — the first-launch risk and the Makefile handling shipped in the speedrun; what's left is a paused-machine check): +- Waypaper canonical decision (Craig): RESOLVED 2026-07-14 — dark-lion.jpg is canonical (dotfiles fea3e93), repo config.ini updated off the that-one-up-there.jpg placeholder. +- Ratio check: RESOLVED 2026-07-14 — ratio's =~/.config/calibre= is a directory symlink into the repo (stow folded the dir), so the first-launch gap never existed there. +- When velox is back online: run =make conflicts= / =make reset= there so its old conflict copy clears against the updated repo. (velox carries a separate boot-recovery task; check once it's reachable.) +** DONE [#A] Hyprlock lockout: AMD-iGPU DPMS invalidates the lock, session wedges :bug:hyprland:installer:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as archsetup =a9391c9= + dotfiles =3046c9c=, both pushed; applied live to ratio and velox. Reboot ratio to activate the root fix (=amdgpu.runpm=0=); the watchdog covers until then. + +WHAT HAPPENED. Ratio's screen idle-locked, then wedged: hyprlock gone, the compositor still holding the ext-session-lock, no password prompt, recoverable only from a console. Recovered live with =hyprctl dispatch exec hyprlock= (=allow_session_lock_restore=true= was already set, so a replacement client adopted the dead lock). + +ROOT CAUSE (evidence, not the first guess). My first read was "hyprlock crashed on its screenshot buffer" — WRONG. Coredumps are captured here (two telega SIGSEGVs the same afternoon) and there is NO hyprlock coredump, so it did not segfault; memory was fine, so not OOM. The hyprland log shows the real chain: =Modesetting DP-4= / =Restoring crtc 86= (a display modeset) → =color management protocol is enabled and outputs changed= → =SessionLock.cpp:50 SessionLockSurface object remains but surface is being destroyed=. A display power cycle tore down the lock surface. Online research confirms it's a documented AMD-integrated-Radeon issue (hyprlock#953, Hyprland#5822): the GPU resources the lock client holds become invalid when the display powers down and back up. Ratio is a Strix Halo Radeon 8060S — exactly that hardware, and its cmdline already carried =amdgpu.dcdebugmask=0x10= + =no_vpe_idle_pg=1= display workarounds, a history of the same fragility. + +THE FIX, four layers, research-validated: +1. Root cause: =amdgpu.runpm=0= on the kernel cmdline (AMD only, added in =update_grub_cmdline= behind =detect_gpu_vendors=). Keeps GPU runtime PM from invalidating the resources on a display cycle. Live in ratio's grub.cfg; effective next boot. +2. Separate crash cause: =configure_hyprlock_pam= writes a complete =/etc/pam.d/hyprlock= (auth/account/session). The package default is =auth include login= only, so pam_end() crashes on uninitialised handles. Applied live to both machines. +3. Recovery net: the =screen-lock= watchdog (dotfiles) relaunches hyprlock on a non-zero exit; hypridle's =lock_cmd= routes through it. Independently the same shape as the community's watchdog layer. +4. NOT done, deliberately: the =dpms off= listener stays in the committed hypridle — =runpm=0= makes it safe on AMD, and it's wanted on Intel/velox for idle display-off. Ratio's test rail already removed it as a local choice. + +REVERTED a wrong turn: I'd first built a screenshot-to-file change (grim the desktop, point hyprlock at the file) on the theory the live screencopy buffer crashed. The research showed the cause is GPU runtime PM, not the background source, so I dropped it and reverted hyprlock.conf to =path = screenshot=. + +PROCESS NOTE — I hit the pathspec-commit trap AGAIN (the one the =Two agent sessions sharing one repo= VERIFY documents). After surgically staging only the =lock_cmd= line via =git update-index=, I ran =git commit <path> -m ...=, which commits the WORKING TREE of that path, not the index — so it committed ratio's test rail (dpms-off removed, timeout 450) with a message claiming dpms-off stays. Caught it before push, =git reset --soft=, re-verified. The rule: after =update-index=, commit with =git commit= (no pathspec), never =git commit <path>=. + +Tests: archsetup 372 (test_grub_cmdline AMD-runpm cases + test_hyprlock_pam, both call sites in CALL_SITES); dotfiles 3687 incl. tests/screen-lock. Each guard proven by deletion. + +Grading: Critical severity (full session lockout, console-only recovery) x rare edge case (needs an idle lock plus a display modeset on the AMD iGPU) = P2 = [#B by the matrix]. Raised to [#A] here because it stranded a live machine and the root fix needs a reboot to arm — worth Craig seeing at the top until he reboots ratio. +** DONE [#B] Adversarial review of the sentry run — six fixes reworked :bug:test:tooling:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Craig asked for a skeptical review of every sentry change. Eight agents covered all 23 code commits, each told to disbelieve by default and to answer three questions per commit: does the problem exist and is it reachable, is the fix correct or is there a better one, would each test fail with the fix reverted. Every finding below was re-verified by hand before acting on it. + +SIX COMMITS NEEDED WORK, now fixed: archsetup =1207ca5= (wipedisk), =96e12b5= (firmware trim), =560e1dd= (autologin), =3c2155d= (initramfs tabs); dotfiles =ec7229b= (tunnel import), =a81aa0e= (thumbnail sweep), =56807e5= (three residual guards), =c90ee34= (event-log isolation). Both suites green: archsetup 341, dotfiles 3687 on both gates. + +THE ONE THAT MATTERED MOST. =wipedisk= ran =blkdiscard -f= BEFORE the busy check. =-f= disables the exclusive open util-linux has used since 2.36, so on the exact case the round-11 commit reasoned about — the user picked the wrong disk — it discarded a live filesystem and only then let sgdisk fail, printing "could not clear the partition table ... run this again". Data gone, user told nothing happened. The ordering predates the sentry commit, but round 11 wrote reasoning about the busy-disk case into the comment and error text while leaving the discard first, which made the misreport worse in the one direction that costs something. Dropping =-f= makes the kernel's own O_EXCL the gate. + +THREE PATTERNS WORTH MORE THAN THE INDIVIDUAL FIXES: + +1. CALL SITES WENT UNTESTED IN FIVE SUITES. Every helper had thorough tests; not one proved it was called. Deleting the call left everything green — including the guard on a =pacman -Rdd= of twelve firmware packages, whose removal would have run the trim on ratio. Closed with =CALL_SITES= in =test_orchestrators= (nine pairs, static) and a wiring assertion in the settings suite. Static on purpose: the behavioural harness runs un-stubbed bodies for real, which is fine for an orchestrator and not for a leaf that removes packages. + +2. A NEW OUTCOME VALUE NEEDS EVERY CONSUMER WALKED, EVERY TIME. Done for the portal enum in round 3, skipped for the tunnel-import one in round 4 — where =import_configs= folded a disarm failure into "none imported (N failed)", the opposite of what happened, in the multi-select flow the GUI actually uses. + +3. MY FIXTURES TWICE CLAIMED A FIDELITY THEY DID NOT HAVE. The wipedisk fixture used this machine's real disk names, so five of six tests passed with the seam removed. The mkplaylist fake does a full =cat > /dev/null= drain while its docstring says it "drains stdin exactly when the real one would" — which is what let the wrong failure mode survive. + +AND ONE FINDING WAS DISPROVED OUTRIGHT: round 1's =a57c443= claimed ffmpeg drains the read loop so only the first track is processed. Measured under strace and driven end to end with real ffmpeg (three runs of three, four 120s mp3s), the loop never truncates. The hazard is real and =-nostdin= is right; the symptom was reasoned from shellcheck SC2095 and never run. Corrected in =a30741a=, along with the OpenVPN autoconnect claim and the "four consumers" undercount. + +ALL THREE NOW CLOSED, in dotfiles =c7cb40d= (pushed). =_restore_dot='s =noop= split into =already-on= and =not-managed=, so the step stops claiming a restore that never happened. =_disable_dot= checks its restart as well as its move, since moving the drop-in aside does nothing until resolved reloads. + +The thumbnail one could not be built as described, and that is worth recording. The cache name is a SHA-1 of realpath plus mtime, so no filename says which source it came from; per-source sweeping would mean changing the key format and invalidating every cached thumbnail. Bounding the growth gets the same result for less: a deferred sweep now trims to a 500-file ceiling, oldest first, because eviction is safe exactly where sweeping is not (an evicted thumbnail is rebuilt on the next warm pass, costing one decode and never a file). WHEN A FIX CANNOT BE BUILT AS SPECIFIED, SAY SO AND SOLVE THE ACTUAL HAZARD — the hazard here was unbounded growth, not imprecise attribution. +** DONE [#D] Repair tiers call an unverifiable service restart a failed one :bug:network:bluetooth:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as dotfiles =041d6b9= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). 7 new tests across =tests/bt/test_bt.py= and =tests/net/test_net.py=; dotfiles suite 3665 -> 3672, =make test= exit 0 on both gates. Each of the three guards proven a real gate by deleting it and watching the suite go red. + +Found in the 2026-07-24 sentry bug-hunt, round 14, on the cross-package =repair.py= diff that rounds 5-13 had left unspent. + +=cmd.service_active= is tri-state in both the net and bt packages, and its docstring says so outright: True, False, or None when systemctl itself can't answer (absent binary, or a timeout). Six callers. Three rule on it correctly — =bt/doctor._service_step= branches on None with "systemctl unavailable — can't check the service", and =net/diag= compares =is False= at both its call sites. Three tested it with plain truthiness: + +- =bt/repair.py= =repair_service_restart= +- =net/repair.py= =_service_restart= (the nm-restart and resolved-restart tiers) +- =net/repair.py= =repair_unmask_nm= + +So an unanswerable systemctl was reported as "bluetooth.service is still not active" / "NetworkManager still isn't running after a restart" — a statement about the service made on no evidence at all. Each then pointed the user at =journalctl -u <unit>=, which is the same systemd client stack that had just failed to answer. That last part is round 10's read again: an error message advertising a remedy it cannot honour. + +All three now report =warn= on None, with evidence naming the verification rather than the service, and a next action of checking systemd is reachable and re-running the doctor. Control flow is unchanged: =warn= was already a status both packages emit, both CLIs already exit non-zero on anything but =pass=, and =net/doctor= only inspects a repair step's status for the =dns-test= tier — every consumer was checked before the change, not after. (An adversarial re-review counted twelve, not four; all twelve handle =warn= correctly, so the conclusion held while the claim understated the work.) + +THE SEAM FOR THE TESTS, worth reusing: both suites already carry an exec-failure harness that plants a non-executable file on an emptied PATH, which is exactly what makes =cmd.run= return None. So the None case is reachable through the real code path with no mocking at all. Each test class asserts that premise first (=service_active= really is None in the sandbox) rather than assuming it. + +Grading: Minor severity (the claim is wrong but errs pessimistic — it says a repair failed when it may have worked, rather than falsely reassuring; nothing is damaged) x rare edge case = P4 = [#D]. Fixed rather than filed because the change is three branches and it completes a class — leaving two of three sites collapsed is the failure mode the round-6 =c2eb3e1= commit exists to remember. + +NOT PART OF THIS CLASS, checked and left alone: =settings/toggles.dim_state= is the only other genuine True/False/None helper in the tree, and both its callers pass the value through to the viewmodel rather than collapsing it. Every other "or None" in the packages is two-state (a value or nothing), where falsy handling is correct. +** DONE [#B] Firmware trim gated on a DMI field that never carries the vendor :bug:tooling:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as archsetup =2e228f7= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). New =tests/installer-steps/test_framework_firmware_trim.py=, 12 tests carrying the real DMI strings off both daily drivers. Each of the three conditions proven load-bearing by deleting it and watching the suite go red, and the old gate proven wrong by restoring it (4 failures). + +Found in the 2026-07-24 sentry bug-hunt, round 13, reading archsetup's remaining state-mutating steps. =trim_firmware= gated on =grep -qi "framework" /sys/class/dmi/id/product_name= and no Framework machine has "framework" in =product_name= — it lives in =sys_vendor=. Read live: velox is =Framework= / ="Laptop (13th Gen Intel Core)"=, ratio is =Framework= / ="Desktop (AMD Ryzen AI Max 300 Series)"=. The gate returns false on both, so the step has been a silent no-op on the exact hardware it was written for. velox IS trimmed today (=linux-firmware-{atheros,intel,realtek,whence}= and nothing else) but not by this code path. + +THE REPAIR IS WHERE THE DANGER IS, which is why this is worth reading twice. Swapping =product_name= for =sys_vendor= is the obvious one-word fix and it is wrong: ratio is a Framework Desktop, and =trim_firmware= runs =pacman -Rdd linux-firmware-amdgpu=, which takes the firmware its Ryzen AI Max iGPU needs to bring up a display. Today only the =grep -qi intel /proc/cpuinfo= second gate stands between ratio and that. So =is_framework_intel_laptop= wants three DMI facts — vendor Framework, and a model naming both Laptop and Intel — and the cpuinfo read stays as an independent second gate rather than the only one. + +Verified live after the change: velox TRIM=yes, ratio TRIM=no, where the old gate said no to both. + +Grading: Minor severity (the trim never happens; nothing breaks, the machine just carries ~550MB it was meant to shed) x every user, every time (every Framework Intel install, which is the whole population the step targets) = P2 = [#B]. The AMD-firmware removal is not graded separately because it never shipped — it is the hazard the fix is shaped to avoid. +** DONE [#B] Fresh install leaves the dotfiles repo permanently dirty :bug:tooling:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as archsetup =c3b3617= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). New =tests/installer-steps/test_mark_volatile_configs.py=, 8 tests against a fixture git repo with =sudo= stubbed on PATH. Every guard proven a real gate by deletion. A note went to =~/.dotfiles/inbox/= because =skip-volatile= now has an outside caller. + +Found in the 2026-07-24 sentry bug-hunt, round 13, diffing archsetup's =stow_dotfiles= against the dotfiles Makefile's =stow= target — two implementations of one operation, which is round 5's read applied across repos rather than across packages. + +The Makefile's =stow= target ends with =$(MAKE) skip-volatile=, setting git's skip-worktree bit on the four configs their apps rewrite in place (=btop=, =qalculate=, =calibre=, =waypaper=; the list is =volatile-configs=). archsetup stows inline with raw =stow= calls and never ran that step. So a machine archsetup installed goes dirty the first time one of those apps writes its config, and every later =git pull --ff-only= trips over paths the user never edited. Confirmed by grep: archsetup contains no =skip-volatile=, no =volatile=, and no =make stow= — yet both daily drivers carry the bits, so they came from a hand-run =make stow=, not the installer. ratio in fact carries seven, three more than =volatile-configs= lists, which is evidence the churn is real and ongoing. + +The fix calls the dotfiles target rather than copying its logic, so the volatile list stays single-source. Two details that are load-bearing: it runs *after* =git restore .= so the bit lands on a pristine tree, and it runs as the user, because root writing =.git/index= leaves it root-owned and the user's next git command then cannot update the index at all. A checkout with no Makefile is a quiet no-op — nothing to delegate to is not an error. + +DELIBERATELY NOT DONE: replacing the whole inline stow with =make -C "$dotfiles_dir" stow "$desktop_env"=. The Makefile stows =--target=$(HOME)=, which during an install is root's home, and it carries interactive conflict handling; archsetup stows =--target=/home/$username --adopt= as root on purpose. =skip-volatile= is the one target with no such coupling — it works on the repo through =git -C= and never reads HOME. + +Grading: Minor severity (a repo that reads dirty forever and pulls that need a stash; the workaround is one command) x every user, every time (every fresh install that stows dotfiles) = P2 = [#B]. +** DONE [#C] Unattended install blocks on an interactive prompt :bug:tooling:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as archsetup =cbcb53f= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). New =tests/installer-steps/test_configure_autologin.py= (11) and =tests/installer-steps/test_select_locale.py= (11). Every guard proven a real gate by breaking it and watching the suite go red: dropping the autologin unattended branch fails 1 (on the leftover-stdin assertion, which is the real gate — the drop-in still gets written because the read swallows the sentinel and treats it as "yes"); dropping the locale unattended branch fails 1; breaking either precedence rule fails 2. + +Found in the 2026-07-24 sentry bug-hunt, round 12, continuing through archsetup's own installer. Two members of one class, which is the point: round 10 fixed the third member and left these. + +THE CLASS: an advisory prompt — one that carries its own default — still reading stdin under =--config-file=, the documented unattended mode. Round 10 ruled on it for =nvidia_preflight='s rc-10 prompt. Two sites never got the ruling. + +1. =configure_autologin=. When =enable_autologin= is unset (=AUTOLOGIN= is optional, and =archsetup.conf.example= line 31 ships it commented out) and the root is encrypted, it prompted =Enable automatic console login for $username? [Y/n]= on a bare =read=. It runs from =configure_encrypted_autologin=, inside =boot_ux=, the last entry in =STEPS= — so an unattended install of an encrypted machine works for 40-60 minutes and then sits at a prompt nobody is watching. Under =curl | bash= it is worse: stdin is the script itself, so the read eats a line of source. + +2. =select_locale= (extracted from =preflight_checks= by this commit). The =Choice [1]:= menu fired whenever =/etc/locale.conf= carried no =LANG== and =LOCALE= was unset — also commented out in the example config. archsetup does not require an archangel install, and =configure_build_environment='s own "no LANG=" branch is proof it expects that state. + +Both now take the prompt's own default under =--config-file= and print an =[OK] ... (unattended, --config-file)= line saying so. An explicit =AUTOLOGIN=yes/no= or =LOCALE== still wins; the default only answers a question nobody can. + +WHAT MADE THEM TESTABLE, which is round 10's read (d) applied again: =configure_autologin= hardcoded =/etc/systemd/system/getty@tty1.service.d= and =select_locale= hardcoded =/etc/locale.conf=, so neither could run against a fixture — while their siblings =replace_sudoers_pacnew= and =ensure_nvme_early_module= both take a defaulted path argument for exactly that reason. Both now do. Zero shellcheck delta against HEAD; =make test-unit= 276 -> 298, exit 0. + +Grading: Major severity (unattended installation, a documented feature, does not complete; recoverable by pressing a key, no data loss) x some users, sometimes (needs unattended mode plus an omitted key) = P3 = [#C]. + +THE PROMPTS DELIBERATELY LEFT ALONE, because the class is "prompts with a default", not "all prompts": username (line 636) and password (648/650) have no default to take — there is no sane fallback for either, and =archsetup.conf.example= documents both as "If not set, you will be prompted". They also fire in =preflight_checks=, in the first second of the run, where a blocked prompt is visible rather than silent. The "Enter locale" sub-prompt is reachable only from menu choice 9, which unattended never picks. +** DONE [#C] wipedisk says "Disk erased." when it erased nothing :bug:tooling:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as archsetup =HEAD= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). New =tests/wipedisk/test_wipedisk.py=, 6 tests running the real script against a fixture device directory with fake blkdiscard/sgdisk on PATH. All four guards proven real by deleting each and watching the suite go red. + +Found in the 2026-07-24 sentry bug-hunt, round 11, reading =scripts/= — 30 lines, no tests, and the most destructive script in the repo. Not installed by the installer; it is run by hand from the checkout, which is why the frequency axis stays low. + +Three defects, all of which make the script's final word untrue: + +1. =sgdisk --zap-all= had its result discarded, and "Disk erased." printed unconditionally. sgdisk refuses a busy device — a mounted filesystem or a live md/LVM/ZFS holder — which is exactly what a user hits after picking the wrong disk. So the tool announced an erase it had not performed and exited 0. + +2. "Disk erased." overstates what the tool does even on success. =sgdisk --zap-all= destroys partition tables, not data, and =blkdiscard -f ... || true= deliberately tolerates a device that cannot discard. On a disk without discard support the script cleared the partition table and left every byte readable, while telling the user the disk was erased. That is the one path where the wrong belief has a privacy consequence — someone trusting the message before disposing of a drive. + +3. The prompt says "Select the disk id to use" and then listed every entry in =/dev/disk/by-id=. On this machine that is 18 entries of which 12 are =-partN= partitions (verified by listing it). The menu promised disks and offered partitions. + +Fix: whole disks only (globbed rather than =ls | grep=, so a name with whitespace cannot split into two menu entries); the zap's result is checked and a failure exits 1 naming the busy-device cause; the closing message reports what actually happened, and when discard was unsupported it says the data is still recoverable and points at =nvme format= / =hdparm= for a disposal-grade wipe. + +Grading: Major severity (the tool reports an outcome it did not achieve; in the disposal case that is a data-exposure consequence) × rare edge case (a hand-run helper the installer does not install, and defect 1 additionally needs sgdisk to fail) = P3 = [#C]. + +Worth recording about the tests rather than the code: two of the six passed against the unmodified script for the wrong reason. Without the =WIPEDISK_BY_ID= override the script read the real =/dev/disk/by-id=, so the harness was driving a menu of this machine's actual disks (harmless — the fake blkdiscard/sgdisk shadowed the real ones on PATH — but it was not testing the fixture). And =test_empty_by_id_directory= was not a gate at first: with the guard deleted the empty select menu still falls through to the confirm prompt, reads EOF and declines, so exit code and call log alone pass either way. It now asserts the message. +** DONE [#B] zfs-replicate reports success when every backup failed :bug:backup:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as archsetup =HEAD= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). Diagnostics moved to stderr; the loop counts failures and exits 1 when any dataset failed. New =tests/zfs-replicate/test_zfs_replicate.py=, 9 tests driving the real script with a fake syncoid and a fake ping on PATH (the =tests/zfs-pre-snapshot/fake-zfs= pattern). Both fixes proven real gates by reverting them: dropping the counter fails 3, putting =error()= back on stdout fails 1. + +Found in the 2026-07-24 sentry bug-hunt, round 11, reading =scripts/= — 73 lines with no test file, installed by =configure_zfs_snapshots= as =/usr/local/bin/zfs-replicate= and run by =zfs-replicate.service=, a =Type=oneshot= on a nightly timer. Its exit code and its journal output are the only signals anyone ever sees. + +Two defects, both verified by running the script rather than argued: + +1. The full-replication loop caught each =syncoid= failure, warned, carried on, then printed "Replication complete." and exited 0 regardless. Driven with a fake syncoid failing all four datasets: four =[WARN] Failed= lines, then "Replication complete.", exit code 0. systemd records =Result=success=. A backup that has not run for months is indistinguishable from a working one — and the whole point of the tool is to have a copy when the primary is gone. + +2. =determine_host= runs inside a command substitution (=TRUENAS_HOST=$(determine_host)=) and its =error()= wrote to stdout. On an unreachable TrueNAS the message was captured into =TRUENAS_HOST= and discarded, and =set -e= then killed the script. Driven with both hosts unreachable: exit 1 and completely empty output. A nightly service failing with nothing in the journal to say why. + +Same class as three bugs already fixed this session — =_restore_dot= claiming "DNS-over-TLS restored" without checking, =portal_restore_watch= discarding its outcome, =import_config= returning ok on an unchecked modify. A mutating operation that reports a success it did not get. + +Grading: Critical severity (a backup system that reports success while backing nothing up; the failure surfaces only when the backup is needed — graded on the harm once in the failure state, not on how rarely it is entered) × rare edge case (needs a ZFS root, a reachable TrueNAS, and the user enabling the timer by hand — archsetup deliberately does not enable it, and =findmnt -n -o FSTYPE /= on this machine says btrfs, so it is latent here) = P2 = [#B]. + +Left alone: =BACKUP_PATH="backups" # TODO: Configure actual path= is still an unresolved TODO in the destination, and single-dataset mode relies on =set -e= to propagate a syncoid failure rather than reporting it. Neither is a defect in the sense above; the TODO is Craig's call. +** DONE [#D] Wireless regdom is silently unset for a three-letter-language locale :bug:installer:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as archsetup =249bb93=. =locale_country= matches the =_CC= group instead of counting characters, and =set_wireless_regdom= verifies the substitution landed rather than trusting sed's exit code. 16 tests. +=configure_networking= derives the wireless regulatory domain by fixed offset: =wireless_region="${current_lang:3:2}"=, with a comment reading "extract country code (positions 3-4)". That is correct only for a two-letter language code. + +=validate_config= accepts =^[a-z]{2,3}(_[A-Z]{2})?...=, so a three-letter language is a legal =LOCALE=, and glibc ships 75 of them (=agr_PE=, =ast_ES=, =ber_DZ=, =ayc_PE=, ...). Verified by running the expansion: =ber_DZ.UTF-8= yields =_D=, =ayc_PE.UTF-8= yields =_P=, =C= yields the empty string, =POSIX= yields =IX=. + +The sed that follows only uncomments an existing =#WIRELESS_REGDOM="XX"= line in =/etc/conf.d/wireless-regdom= (176 of them, owned by wireless-regdb). A garbage region matches nothing, sed exits 0, and the =|| error_warn= never fires — so the regdom is never set and nothing says so. The task line does print the garbage region ("configuring wireless regulatory domain (_D)"), so it is visible in the log rather than fully silent. + +Confirmed the mechanism itself works for the normal case: line 168 of this machine's =/etc/conf.d/wireless-regdom= reads =WIRELESS_REGDOM="US"= uncommented, which is archsetup's own edit. + +Grading: Minor severity (WiFi falls back to the conservative "00" regdomain — fewer channels and lower tx power, but WiFi works) × rare edge case (one of 75 three-letter-language locales, or a =LOCALE= with no country) = P4 = [#D]. + +Fix when it comes up: derive the country from the =_CC= group by pattern rather than by offset, and warn when it cannot be derived or when the sed changed nothing. Worth doing together with the sibling gap — nothing in the installer verifies that a =sed -i= uncomment actually matched, so a distro reshuffling one of these config files would fail the same silent way. All 22 =sed -i= sites share that stance, so it is a uniform design choice rather than an odd one out. +** DONE [#B] Initramfs hook swap can leave a LUKS machine unbootable :bug:installer:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as archsetup =HEAD= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). The swap moved into =switch_udev_hook_to_systemd=, which declines when =hooks_need_busybox_init= sees a standalone =encrypt= token, and the caller now rebuilds the initramfs only when the conf actually changed. New =tests/installer-steps/test_switch_udev_hook.py=, 10 tests; both guards proven real by breaking them (removing the refusal: 4 failures; loosening the token match to a bare =encrypt= substring: 1 failure). + +Found in the 2026-07-24 sentry bug-hunt, round 10. =configure_initramfs_hook= ran =sed -i '/^HOOKS=/ s/\budev\b/systemd/'= on any non-ZFS root, then =mkinitcpio -P=. Its only guard was =is_zfs_root=. + +Why that breaks a LUKS machine, verified against the installed mkinitcpio rather than argued: +- =/usr/lib/initcpio/install/systemd= line 70 is =add_symlink /init usr/lib/systemd/systemd=, so the systemd hook replaces the busybox init outright. +- =/usr/lib/initcpio/hooks/encrypt= is an =#!/usr/bin/ash= script whose entire body is a =run_hook()= function — the busybox init's mechanism. Under systemd init nothing calls it. +- =mkinitcpio= carries no conflict check for the pairing (grepped; nothing), so the rebuild succeeds and archsetup reports success. +- This machine's own =/etc/mkinitcpio.conf= documents the two valid pairings as separate examples: =udev= + =encrypt= (line 45) and =systemd= + =sd-encrypt= (line 51). The sed converted half of the first pairing and produced neither. + +Effect: on a LUKS root using the standard busybox =encrypt= hook, archsetup rewrites HOOKS to =systemd= while leaving =encrypt= behind, rebuilds the initramfs, and exits cleanly. At the next boot the root is never unlocked. The machine needs live media and manual mkinitcpio surgery to recover. + +The sibling asymmetry: =is_encrypted_root()= already exists in this script and =configure_autologin= uses it to branch on exactly this condition. The initramfs step consulted neither it nor HOOKS. =merge_grub_cmdline='s own comment names =cryptdevice== as a boot-critical parameter to preserve — and =cryptdevice== is read only by the =encrypt= hook, so archsetup explicitly anticipates the configuration that another of its steps then breaks. + +Grading: Critical severity (the machine will not boot and recovery needs external media — graded on the harm once in the failure state, not on how rarely it is entered) × some users, sometimes (LUKS-encrypted non-ZFS root using the busybox =encrypt= hook; deterministic for those machines, absent everywhere else) = P2 = [#B]. + +Deliberately not attempted: migrating =encrypt= to =sd-encrypt=. That means rewriting the kernel cmdline from =cryptdevice== to =rd.luks.name== against the volume's UUID, which is a real migration and not a mechanical edit. Refusing the cosmetic swap keeps a working machine working, which is the right trade against quieter fsck output. +** DONE [#D] keymap and consolefont hooks are inert under the systemd initramfs :bug:installer:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as archsetup =249bb93=. The swap rewrites both to =sd-vconsole=, collapsing them into one entry and never duplicating an existing one. The open question is answered: this machine is KEYMAP=us with no encrypt hook, but the function runs on LUKS machines where a non-US layout at the passphrase prompt is exactly what sd-vconsole restores. 7 tests. +Same class as the =encrypt= bug above, but cosmetic rather than boot-critical, so it was filed rather than bundled into that fix. + +Enumerating the busybox-only hooks on this machine (every hook under =/usr/lib/initcpio/hooks/= defining =run_hook=/=run_earlyhook=/=run_latehook=) gives: btrfs, consolefont, encrypt, grub-btrfs-overlayfs, keymap, memdisk, resume, sleep, udev, usr. All go inert once =/init= is systemd. Of those, =encrypt= is the only boot-critical one — =resume= is handled natively by systemd's hibernate-resume generator, and =btrfs= by udev rules (this machine runs =btrfs= alongside =systemd= and boots fine). + +=keymap= and =consolefont= are the live leftovers. Run =grep '^HOOKS=' /etc/mkinitcpio.conf= on this machine: the line carries =systemd= plus =keymap consolefont= and no =udev=, so archsetup's swap has already run here and both hooks are installed into the image and never executed. The systemd equivalent is the single =sd-vconsole= hook, which is what the distro's own systemd example on line 51 of =/etc/mkinitcpio.conf= uses. + +Effect: the early-boot console keeps the default font and keymap until =systemd-vconsole-setup= runs in the real root. =add_nvme_early_module= sets =FONT=ter-132n= in =/etc/vconsole.conf= expecting it to apply at that stage, so the configured font is briefly not what archsetup asked for. + +Grading: Cosmetic severity (a few seconds of default console font on a machine that boots normally) × some users, sometimes = P4 = [#D]. + +Fix when it comes up: have =switch_udev_hook_to_systemd= also rewrite =keymap consolefont= to =sd-vconsole= when it performs the swap, and add the fixture cases to =tests/installer-steps/test_switch_udev_hook.py=. Worth confirming first whether a non-US keymap is ever needed at the initramfs prompt on a machine that reaches this path. +** DONE [#B] NVIDIA Wayland preflight blocks dwm and headless installs :bug:installer:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as archsetup =HEAD= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). The NVIDIA block moved out of =preflight_checks= into a new =nvidia_preflight= function that returns early unless =desktop_env= is =hyprland= and archsetup is the one installing drivers. New =tests/nvidia-preflight/test_nvidia_preflight_gate.py=, 11 tests; each of the three guards was proven a real gate by deleting it and watching the suite go red (3, 1, and 1 failures respectively). + +Found in the 2026-07-24 sentry bug-hunt, round 10, reading archsetup's own installer. =preflight_checks= called =nvidia_preflight_report= unconditionally and exited 1 on rc 11 (repo driver below the 535 Wayland floor, or =pacman -Si nvidia-utils= unable to answer). The check is Wayland-specific — every line it prints names Wayland/Hyprland — but it ran before any =desktop_env= branch and consulted neither =desktop_env= nor =skip_gpu_drivers=. + +Effect, proven empirically rather than argued (three scenarios driven against the extracted block): =DESKTOP_ENV=dwm= plus =--no-gpu-drivers= on an NVIDIA machine with an old repo driver aborts the install; so does =DESKTOP_ENV=none=. Neither install ever runs a compositor, and =--no-gpu-drivers= means the user installs the driver themselves. Worse, the abort's own fix hint reads "install with DESKTOP_ENV=dwm (X11) instead" — the one remedy it prints is the one it refuses to honor, so the user has no working workaround short of editing the script. + +The sibling asymmetry that makes it an oversight rather than a decision: =install_gpu_drivers= returns early on =skip_gpu_drivers=, and =display_server= / =window_manager= both branch on =desktop_env= with a =none= arm that skips outright. The preflight gate applied neither ruling. + +Second defect at the same site, fixed in the same commit: the rc-10 path (card detected, driver fine) prompts with a bare =read=. =--config-file= is documented as "unattended installation", and =aur_install= already rules that a prompt not covered by =--noconfirm= "blocks forever waiting for input" on a headless install. The rc-10 prompt is advisory, so it now answers itself with its own =[Y/n]= default when a config file was supplied. rc 11 stays a hard stop either way. + +Grading: Critical severity (archsetup cannot be run at all on that machine, and the printed workaround does not work — graded on the harm once in the failure state, not on how rarely it is entered) × rare edge case (needs an NVIDIA card, a repo driver below the floor or an unsynced pacman db, and a non-hyprland =desktop_env=; hyprland is the default and Craig's own machines are AMD and Intel) = P2 = [#B]. + +Noted, not fixed: =display_server= and =window_manager= both point their unknown-value hint at a =--desktop-env= flag that the argument parser does not implement. Both arms are unreachable today (=validate_config= rejects a bad =DESKTOP_ENV=, and without a config file the value is always the default), so it is a stale string rather than a live defect. +** DONE [#B] mkplaylist retags only the first file :bug:music:quick:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as dotfiles =a57c443= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). =ffmpeg -nostdin= on the conversion call. New =tests/mkplaylist= suite, 12 tests; removing the flag turns the suite red (verified by reverting: 5 failures, green on restore). NOTE: the fake ffmpeg does a full =cat > /dev/null= drain, which the real one does not do — so the suite gates the flag's presence, not the production failure mode. The docstring claiming the fake "drains stdin exactly when the real one would" is false and should be corrected. +Found in the 2026-07-24 sentry bug-hunt (shellcheck SC2095). =common/.local/bin/mkplaylist=: =generate_music_m3u= pipes the file list into =tag_music_file= (line 130), which consumes it with =while IFS= read -r file=. Inside that loop, =ffmpeg -i "$file" -vn -c:a flac "$outputfile"= (line 46) reads stdin by default for its interactive keyboard controls, so it consumes bytes the loop is relying on. + +CORRECTION (2026-07-24, from an adversarial re-review): the failure mode stated above — "the loop sees EOF and exits after the first file" — is WRONG, and this task originally asserted it. Measured under strace, ffmpeg polls fd 0 and reads roughly one byte per half-second of transcode wall time; flac encoding runs about 2000x realtime, so a ten-minute mp3 converts in ~0.28s and yields zero or one stolen byte, never a drain. Driven end to end with real ffmpeg against four 120s mp3s, three runs of three: all four were converted and retagged every time. The loop never truncated. + +What is real is the hazard, not the observed symptom: one stolen byte mangles a path, which makes mid3v2/metaflac fail and =set -e= abort the run loudly. =-nostdin= is still the right fix and the commit still stands. The original finding came from shellcheck SC2095 plus reasoning, and was never run — which is exactly what "verify before filing" exists to prevent. + +Effect: on a directory of non-flac audio, only the first file is converted and retagged. Files 2..N are silently skipped — no error, no output, and the playlist itself still generates (a separate =find=), so nothing signals that the retagging stopped. + +Grading: Major severity (the retagging feature is broken past the first file, and it fails silently) × most users frequently (the script exists to batch-process a directory, so more than one non-flac file is the normal case) = P2 = [#B]. + +Fix: =ffmpeg -nostdin= (or =< /dev/null= on the call). Verifiable with a fake =ffmpeg= on PATH asserting it is invoked once per input file. +** DONE [#C] timezone-change prints command-not-found instead of its help :bug:tooling:quick:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as dotfiles =15d2b63= (committed locally, deliberately NOT pushed — held for Craig's morning review), together with the Portugal-zone defect below. New =tests/timezone-change= suite, 12 tests. +Found in the 2026-07-24 sentry bug-hunt (shellcheck SC2288). =common/.local/bin/timezone-change=, default =*)= case (lines 63-67): =echo= sits alone on its own line, so the following quoted string runs as a *command* rather than as its argument. + +#+begin_src sh +*) + echo + "Invalid option chosen." + echo + "Some valid options are: eastern, central, pacific, rome, london, st_lucia, italy, france, spain ." + ;; +#+end_src + +The user gets two blank lines and two =command not found= errors; the list of valid options never prints. The timezone is correctly left unchanged, so this is an output defect only. + +Grading: Minor severity (wrong output on an error path, nothing corrupted) × some users sometimes (only on an unrecognized option) = P3 = [#C]. + +Fix: fold each string into its =echo=. Verifiable by running the script with a bogus argument and asserting the option list appears on stdout. +** DONE [#C] Thumbnail sweep wipes the whole cache when a wallpaper source is unreadable :bug:settings:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as dotfiles =0bd8c67= (committed locally, deliberately NOT pushed — held for Craig's morning review). 8 new tests. +Found in the 2026-07-24 sentry bug-hunt, reviewing the orphan sweep shipped the night before (dotfiles =e752a16=). =os.walk= stays silent about a directory it cannot enter, so =wallpaper.scan_sources= returns =[]= for a source that is missing, renamed, or permission-denied — the same answer it gives for a gallery the user emptied on purpose. =settings/cli.py= tick then hands that empty list to =thumbstore.sweep_orphans=, =live_names= comes back empty, and every cache-shaped file is classified an orphan. + +Proven empirically rather than reasoned: seeding three well-formed thumbnails plus a stray README, then sweeping against a nonexistent source directory, deleted all three (the README survived, so the cache-name regex guard works — it just doesn't help here). + +Effect once entered: the entire persistent thumbnail cache is deleted, so the next wallpaper-view open pays the cold-decode cost the cache was built to remove (measured at 3.7s for a viewport of Craig's largest 8, which is what tripped the compositor's kill prompt), and the tick needs roughly ten idle beats — about twenty minutes — to rewarm at =WARM_PER_BEAT= 8. + +Grading: Major severity (grading the being-in-it, per the don't-double-count-rarity rule: the cache is gone, the original freeze returns, and recovery is unattended and slow) × rare edge case (both configured sources — =~/videos/wallpaper= and =~/pictures/wallpaper= — are local directories, so this needs one deleted, renamed, or made unreadable while a beat fires; a removable or network source would hit it routinely) = P3 = [#C]. + +Fixed in this session: new =wallpaper.sources_available(sources)= tells "readable and empty" apart from "could not read", and =sweep_orphans= grew a =sources_ok= parameter that declines to sweep when it is False. Deferring a sweep costs only some stale files; sweeping wrongly costs the whole cache. +** DONE [#C] timezone-change sets a nonexistent zone for Portugal :bug:tooling:quick:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as dotfiles =15d2b63= (committed locally, deliberately NOT pushed — held for Craig's morning review). =Europe/Lisbon=. The suite also pins the general invariant: every zone the script can emit must exist in tzdata, so a future bad entry fails at test time rather than in Craig's hands. +Found in the 2026-07-24 sentry bug-hunt, validating every zone the script sets against =/usr/share/zoneinfo=. =common/.local/bin/timezone-change= line 39 maps =portugal= / =lisbon= to =Europe/Portugal=, which is not a tzdata identifier — the real one is =Europe/Lisbon= (a bare =Portugal= legacy alias also exists at the top level, but not under =Europe/=). =timedatectl set-timezone "Europe/Portugal"= fails, so the timezone is never changed. + +The other 17 zones the script sets all resolve correctly, so this is the single bad entry. + +Grading: Major severity (the option is wholly broken — the zone is not set and the command errors) × rare edge case (one option of eighteen, hit only when actually switching to Portugal) = P3 = [#C]. + +Fix: =Europe/Lisbon=. Verifiable by asserting the argument handed to a fake =timedatectl=, plus a suite-wide check that every zone the script names exists in the tzdata database. +** DONE [#C] settings-project stop() can SIGTERM an unrelated process :bug:settings:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Found in the 2026-07-24 sentry bug-hunt, round 2, reviewing =settings/src/settings/project.py=. =stop()= read a pid out of =$XDG_RUNTIME_DIR/settings-project.pid= and SIGTERMed it with no check that the pid still belonged to the projection. A projection that dies without running =stop()= (crash, OOM, a failed =execvpe= on the clock path — that last one was already noted as tolerated residue) leaves the file behind, so once the kernel wraps its pid counter that pid can name something else entirely, and the next =start= or =stop= kills it. + +This is a hazard the codebase had already ruled on elsewhere and simply hadn't applied here: =maint/src/maint/doctor.py= revalidates =/proc/<pid>/comm= against the expected name before its KILL remedy fires, explicitly to refuse recycled pids. + +Grading: Major severity (grading the being-in-it — an arbitrary user process takes a SIGTERM, and an editor with unsaved work is a plausible victim) × rare edge case (needs an unclean exit *and* pid reuse; =pid_max= here is 4194304, so wrap-around takes a very long time) = P3 = [#C]. + +Fixed as dotfiles =722994e= (committed locally, deliberately NOT pushed — held for Craig's morning review). The pidfile now records the process start time from =/proc/<pid>/stat= next to the pid, and =stop()= fires only when the recorded value still matches the live process. Start time is the right token rather than =comm=: it is mode-independent (the clock channel execs into =python3=, so comm changes while comm-matching would have needed per-mode knowledge) and it is exec-stable, verified directly — pid and start time were identical either side of an =execvpe=. A recycled pid cannot reproduce it. Legacy bare-pid pidfiles keep the old unconditional behavior so the upgrade never strands a live projection. +** DONE [#C] wtimer alarms fire an hour off on the eve of a DST change :bug:timer:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Found in the 2026-07-24 sentry bug-hunt, round 3, reading =timer/src/timer/engine.py=. =parse_alarm= resolves a bare wall-clock time ("07:00") to its next occurrence: it builds today's instant, and when that is already past it rolled forward with =epoch += 86400=. A DST day is 23 or 25 hours long, so a fixed 86400 lands on the wrong wall time whenever tomorrow crosses a transition. + +Reproduced against America/Chicago and the two 2026 US transitions. Asking for =07:00= at 08:00 on Sat 2026-03-07 (spring forward that Sunday) gave 08:00 Sunday — an hour late. Asking for =07:00= at 08:00 on Sat 2026-10-31 (fall back that Sunday) gave 06:00 Sunday — an hour early. + +The recurring path was never affected, which is what makes this an oversight rather than a design choice: =next_alarm= walks candidate days and rebuilds =datetime(y, m, d, hh, mm)= per day, so it is already DST-correct. Only the one-shot rollover took the shortcut. Both were pinned by the new tests. + +Grading: Major severity (grading the being-in-it — an alarm that fires an hour off has wholly failed at the one thing an alarm does, and the fall-back direction wakes you early while the spring-forward direction lets you oversleep) × rare edge case (two nights a year, and only when the requested wall time has already passed today) = P3 = [#C]. + +Fixed as dotfiles =9b6c2c9= (committed locally, deliberately NOT pushed — held for Craig's morning review). The rollover now rebuilds the local time on tomorrow's calendar date, the same construction =next_alarm= uses. Eight tests pin =TZ=America/Chicago= (saved and restored around each case), covering both transitions, the twelve-hour form, an ordinary-day control, a DST eve where the requested time is still ahead, and two characterization cases asserting the recurring path stays DST-safe. +** DONE [#B] net portal-restore claims encrypted DNS is back without checking :bug:net:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Found in the 2026-07-24 sentry bug-hunt, round 3, reading =net/src/net/repair.py=. A captive-portal login moves the DNS-over-TLS drop-in aside so plain DNS can reach the venue's login page, and =_restore_dot()= moves it back afterwards. It fired both privileged steps — the =mv= and the =systemctl restart systemd-resolved= — and returned ="restored"= without reading either result. =repair_portal_restore()= then rendered a pass step reading "DNS-over-TLS restored". + +So a declined or failed =sudo -n mv= left DNS-over-TLS off while the tool told the user it was back on. The same for a resolved restart that fails: the drop-in is on disk but the running resolver is still serving plain DNS. + +The asymmetry is what makes it an oversight rather than a decision. The sibling =_disable_dot()=, twenty lines up, checks its own move with =_ok()= and returns False rather than claiming a success it did not get. The restore half simply never got the same treatment, and it is the half where the failure is silent — the disable path's failure is visible immediately because the portal page won't load. + +Grading: graded on severity alone under the privacy carve-out. DNS queries continue in cleartext to the venue resolver on an untrusted network, and the affirmative "restored" message is what removes the user's reason to check. Bounded by =net diagnose='s =encrypted-dns= step, which exists precisely to catch a portal run that never restored, so the exposure ends at the next diagnose rather than persisting unseen forever. Major severity = P2 = [#B]. + +Fixed as dotfiles =018c0c5= (committed locally, deliberately NOT pushed — held for Craig's morning review). Both privileged steps are now checked, with two new outcomes: ="failed"= when the move back fails (encrypted DNS still off, rendered as a fail step) and ="unapplied"= when the drop-in is back but resolved would not restart (rendered as a warn step). Each names the command to run by hand. Four tests cover both failures at the =_restore_dot()= and step levels, mirroring the existing declined-move test on the disable side. +** DONE [#B] the portal restore watcher fails silently, so DNS stays in the clear :bug:net:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Found in the 2026-07-24 sentry bug-hunt, round 4, reading the rest of =net/src/net/repair.py= after the round-3 fix above. =portal_restore_watch()= polls until the link comes back online, calls =_restore_dot()=, and discards the outcome entirely. + +Three things compound into a silent failure. The watcher is spawned detached with =stdin=, =stdout=, and =stderr= all on =/dev/null=, so nothing it could print reaches anyone. It runs outside the =repair()= dispatch, so unlike every other mutating tier it never wrote an event-log line either. And =repair_portal_login= tells the user "encrypted DNS restores itself once you're online", which is precisely what removes their reason to check. A ="failed"=, ="unapplied"=, or ="ambiguous"= restore therefore left the machine on plain DNS on a venue network with no signal at any level. + +This is the round-3 finding one layer out, and the asymmetry is the tell: =018c0c5= taught =repair_portal_restore()= — the *manual fallback* — to stop claiming a success it did not get, while the *automatic* path, the one that actually runs in the normal flow, kept dropping the same result on the floor. Fixing the fallback and leaving the primary silent is a worse split than the original bug. + +Grading: graded on severity alone under the privacy carve-out, exactly as the round-3 sibling. Same exposure (cleartext DNS to an untrusted venue resolver), same bound (=net diagnose='s =encrypted-dns= step catches the stranded state), and the same affirmative promise removing the reason to look. Major severity = P2 = [#B]. + +Fixed as dotfiles =601c5b4= (committed locally, deliberately NOT pushed — held for Craig's morning review). The watcher now returns the outcome, appends a =portal-restore-watch= event with it, and fires a persistent =notify security= alert on each of the three failing outcomes, each naming the command to run by hand. A clean restore stays silent. Five tests: one per failing outcome, one pinning the silence on a clean restore, and one on the event-log line. The whole =TestPortalLogin= class now shadows =notify= with a logging fake, so no future watcher test can fire a real desktop notification mid-suite. +** DONE [#D] dns-override failure path says "reverted" without checking :bug:net:quick:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as dotfiles =2cf3fb3=. The revert is checked; a declined one now says 1.1.1.1 is still set and names =resolvectl revert <iface>=. +Found in the 2026-07-24 sentry bug-hunt, round 3, sweeping for siblings of the portal-restore finding above. =net/src/net/repair.py=, =repair_dns_override()= failure path: when the 1.1.1.1 override doesn't restore resolution, it calls =priv.run("dns-revert", iface)=, discards the result, and returns evidence reading "override didn't restore resolution — reverted". A failed revert leaves 1.1.1.1 set on the link while the step says it was removed. + +Same defect class as the portal-restore bug, three hundred lines up in the same file, and it survived the sweep only because the consequence is much smaller. Every other mutating repair in this file verifies by re-measuring afterwards rather than by reading an exit code, which is the stronger pattern and is why the sweep otherwise came back dry. + +Grading: Minor severity (a stale per-link override sends DNS to Cloudflare instead of the venue resolver, it dies on the next reconnect, and =net diagnose='s =dns-override-present= step exists specifically to catch it) × rare edge case (needs the override to fail *and* the revert to fail) = P4 = [#D]. + +Fix: the same idiom the portal-restore fix now uses. Wrap the revert in =_ok()= and drop the "— reverted" claim (or say the revert failed and name =resolvectl revert <iface>=) when it returns False. The existing =RepairHarness= makes the privileged call fail with =NET_SUDO="false"=, so the test is a near-copy of =test_restore_reports_failure_when_the_move_back_is_declined=. +** DONE [#B] a timezone-less Date header crashes the whole net diagnose run :bug:net:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Found in the 2026-07-24 sentry bug-hunt, round 4, reading =net/src/net/diag.py=. =_clock_skew_s()= fetches the probe server's =Date= header with =curl -sI=, parses it with =parsedate_to_datetime=, and subtracts it from a timezone-aware =datetime.now(timezone.utc)=. RFC 5322 allows a =Date= to carry =-0000=, which means UTC while explicitly claiming no local zone, and a =Date= with no zone at all parses leniently as well. Both come back *naive*, and subtracting a naive datetime from an aware one raises =TypeError=. + +The =try= wraps only the =parsedate_to_datetime= call, so the =TypeError= from the line below it is uncaught. It escapes =_clock_skew_s=, escapes =_steps_egress_edges=, and takes down the entire =diagnose()= run — no report, no steps, a Python traceback. =net doctor= runs diagnose first, so the panel's doctor button dies with it. + +Verified against Python 3.14.6 before writing the fix: =parsedate_to_datetime("Thu, 01 Jan 2020 00:00:00 -0000")= returns =tzinfo=None=, and the subtraction raises. The zoneless form behaves the same. Only the =GMT= form (which the well-behaved probe host sends) comes back aware, which is why this never showed up in normal use. + +What makes it more than a curiosity is *when* the code runs. =_steps_egress_edges= fires only after the http-probe has already failed, so the server answering that =HEAD= is frequently a captive portal's interception appliance rather than the real probe host — and a minimal embedded HTTP stack is exactly the kind that emits a non-GMT =Date=. The one path guaranteed to be talking to a non-standard server is the one that can't survive a non-standard header. + +Grading: Major severity (grading the being-in-it — the diagnostic tool produces no report at all, and =net doctor= goes with it, on precisely the broken network it exists to diagnose) × rare edge case (needs a failing probe *and* a portal appliance that omits a numeric offset) = P2 = [#B]. + +Fixed as dotfiles =8933500= (committed locally, deliberately NOT pushed — held for Craig's morning review). A naive parse is now read as UTC, which is what =-0000= means. Two tests, and the second is the one that matters: it drives a *current* =-0000= timestamp and asserts no clock row, so a lazy "catch =TypeError= and return None" fix would fail it while the correct reading passes. +** DONE [#C] a tunnel import that can't be disarmed still reports success :bug:net:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Found in the 2026-07-24 sentry bug-hunt, round 4, reading =net/src/net/manage.py=. =import_config()= imports a WireGuard or OpenVPN config as an NM profile, then fires =nmcli connection modify <uuid> connection.id <name> connection.autoconnect no= — and discarded the result, returning =ok=True= regardless. + +That modify is the whole safety of the feature, and the module's own docstring says so: =nmcli connection import= *auto-activates* the profile it creates, "which nobody asked for by picking a file", so "every import here ends with the profile deactivated and autoconnect off". A failed modify inverts that. For WireGuard — a device-type connection — autoconnect stays on, so the tunnel re-arms itself at the next boot and takes the default route with it, and the profile keeps the transient staged interface name (=wgpvpn=) while the envelope reports the config's real name, so the panel names a profile that isn't there. + +CORRECTION (2026-07-24, from an adversarial re-review): the blanket claim originally written here — that a failed disarm re-arms the tunnel at boot — is wrong for OpenVPN. =man 5 nm-settings-nmcli= states autoconnect is not implemented for VPN profiles, and an OpenVPN import is an NM VPN profile, so the modify is near-cosmetic on that half. The bug is real and security-relevant for WireGuard, which is the primary case; the severity as stated overreached to cover both. + +The asymmetry, again the tell: =_nmcli_import()=, twenty lines up in the same file, checks its own =returncode= and raises rather than return a UUID it did not get. The modify below it never got the same treatment. + +Grading: Major severity (grading the being-in-it — a full-tunnel VPN the user never asked to connect arms on every boot and carries all their egress, it persists across reboots rather than self-healing, and the affirmative "imported X" is what removes the reason to check) × rare edge case (needs the modify to fail after the import succeeded) = P3 = [#C]. + +Fixed as dotfiles =e0d4d8a= (committed locally, deliberately NOT pushed — held for Craig's morning review). New =_disarm()= returns whether the modify took. On failure the profile is still deactivated first — the import already brought it up, and the verdict shouldn't decide whether it keeps running — and then a =disarm-failed= envelope names the UUID and the exact command to finish the job. Three tests: the failing verdict, =import_configs= counting it as failed rather than imported, and a characterization test pinning that the deactivate still runs on the failure path. +** DONE [#C] a binary that can't be exec'd crashes the panels instead of degrading :bug:net:bluetooth:audio:maint:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Found in the 2026-07-24 sentry bug-hunt, round 5, comparing the four panel packages' subprocess wrappers against each other. + +Every wrapper in the panels states the same contract: an unusable tool becomes a degraded result, never an exception. =cmd.run= returns None; =nmcli.run=, =btctl.run= and =pactl.run= raise their own domain error, which every caller already guards on; =speedtest.run_speedtest= returns an error envelope. All of them caught only =FileNotFoundError=, so they kept the contract for a tool that is *absent* and broke it for a tool that is *present but unusable*. + +Verified against Python 3.14.6 rather than argued. =subprocess.run= raises =PermissionError= for a file without its execute bit, =OSError= (ENOEXEC, "Exec format error") for an executable file that is neither a binary nor a script with a shebang, =NotADirectoryError= when a path component is a plain file, and =OSError= when a fork is refused under memory or PID pressure. None of the four is =FileNotFoundError=, so each escapes the guard: waybar's net/bt/audio modules die rather than dimming, and a maint probe takes the whole envelope with it — in exactly the machine state maint exists to report on. + +The asymmetry, and this codebase had already ruled on it three separate times: =net/iw.py='s =signal_dbm= and =settings/spawn.py='s =detached= both catch =(OSError, subprocess.TimeoutExpired)=, and =audio/cmd.py='s doctor-tier =probe()= enumerates =FileNotFoundError=, =NotADirectoryError= and =PermissionError= as "absent" under a docstring promising it never raises. Its sibling =run()=, twenty lines up in the same file, kept the narrow catch — as did all five copies of =run()= and all three tool wrappers. =audio/status.py='s docstring records that this same class already bit once ("the bar's audio module died rather than dimming"); that fix widened the guard's *scope* and left its *exception set* alone. + +Grading: Major severity (grading the being-in-it — the status surface is dead while the condition holds, and for maint the tool that reports the fault is the one that dies of it; no data loss, and it clears when the tool or the pressure does) × rare edge case (needs a binary with wrong permissions, a lost shebang, or a fork refused under pressure) = P3 = [#C]. + +Fixed as dotfiles =44fdae1= (committed locally, deliberately NOT pushed — held for Craig's morning review). Widened to =OSError= across net, bt, audio, maint and panelkit — five =cmd.run= helpers, the three tool wrappers, =probe._curl= and =speedtest.run_speedtest=. The domain-error wrappers keep their "<tool> not found" message for a genuinely absent binary and add a second arm naming the errno for an unusable one, so the report can still tell the two apart. 28 tests, one class per package, driving all three exec failures against real files on a temp PATH; each was watched failing against unmodified production code first (27 red). Audio's class carries a characterization case pinning =cmd.probe='s existing behavior, so the sibling that got this right can't regress into the one that didn't. +** DONE [#C] a failed pty-backed spawn strands both ends of the pty :bug:net:bluetooth:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Found in the 2026-07-24 sentry bug-hunt, round 6, auditing the =subprocess.Popen= sites the round-5 fix didn't reach. + +Two spawns open a pty before launching and catch only =FileNotFoundError= around the =Popen=: =bt/pairing.py='s =pair_interactive= (bluetoothctl under a pty so the passkey agent is interactive) and =net/speedtest.py='s =run_speedtest_stream= (speedtest-go under a pty because it buffers everything to exit when piped). Both are the same exec-failure class as =44fdae1= — a binary present but not executable raises =PermissionError=, a lost shebang raises =OSError= — and neither is =FileNotFoundError=. + +What makes these worse than the =run= wrappers is where the cleanup lives. =os.close(master)= and =os.close(slave)= sit *inside* the =FileNotFoundError= arm, so an escaping =OSError= skips them: every failed attempt strands two descriptors. Both call sites are buttons in a long-lived panel process — the pairing flow and the console's SPEED key — and a user who gets no feedback presses again, so the leak accumulates under exactly the conditions that caused it. + +Grading: Major severity (grading the being-in-it — a descriptor leak in a process meant to run for days, on a path the user retries, plus the exception escaping a documented "(ok, detail)" / error-envelope contract) × rare edge case (needs an unusable bluetoothctl or speedtest-go) = P3 = [#C]. + +Fixed as dotfiles =c2eb3e1= (committed locally, deliberately NOT pushed — held for Craig's morning review). An =OSError= arm on each closes both ends and returns the module's own failure shape, naming the errno. Four tests: two pin the return contract, two count =/proc/self/fd= across three attempts — the fd count is what actually fails against unmodified code, and it was watched failing before the fix. + +The wider sweep this came from is recorded so it isn't repeated: every =except FileNotFoundError= in production was enumerated. The other exec sites were already correct (=maint/gui.py= x3, =net/kick.py=, =timer/engine.py= x2, =timer/gui.py=, =net/repair.py= x2, =audio/peak.py= all catch =OSError=), and the remaining hits are file-open catches, not exec. =clock/__main__.py='s =toggle()= has no guard at all but spawns =sys.executable=, which is by definition runnable; not filed. +** DONE [#C] one impatient client kills the clock panel's toggle listener for good :bug:clock:waybar:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Found in the 2026-07-24 sentry bug-hunt, round 6, sweeping every acquired resource (pty, socket, mkstemp, tempdir) for cleanup that isn't in a =finally=. + +=clock/src/clock/app.py='s =_listen()= guards =accept()= with =except OSError: return= and leaves the request body — =recv=, =runtime_log=, =sendall= — outside any guard. =send_toggle()= in =__main__.py= gives the panel 0.25s to acknowledge, then closes. An ack later than that hits a dead peer and raises =BrokenPipeError=, which escapes the =while= loop and ends the listener thread. + +Verified empirically, not argued: a client that connects, sends, and gives up after 250ms makes the server's =sendall= raise =BrokenPipeError= (errno 32) and the listener thread exits. + +What makes it Major rather than a nuisance is that it neither self-heals nor announces itself. The socket file stays bound, so every later =clock toggle= still *connects* — then stalls the full 250ms, gets no reply, and falls through to spawning =clock serve=. GTK's single-instance forwarding turns that into =do_activate= on the running service, and =do_activate= calls =show_clock()=, not =toggle()=. So from the first bad client onward, clicking the waybar time module opens the panel every time and never closes it; the only ways out are the right-click dismiss inside the panel or restarting the service. Nothing logs it. + +Grading: Major severity (grading the being-in-it — the toggle is one-way from then on, it persists for the life of the service, and there is no signal it happened) × rare edge case (needs a reply to miss the 250ms budget: a busy main loop mid-redraw, a slow runtime-log write, or an interrupted =clock toggle=) = P3 = [#C]. + +Fixed as dotfiles =7c02614= (committed locally, deliberately NOT pushed — held for Craig's morning review). An =OSError= arm around the request body scopes a dead peer to its own request, mirroring the guard =accept()= already had. =GLib.idle_add= runs before the ack, so the user's click still takes effect — only the acknowledgement is lost. New =tests/clock/test_socket.py=, 3 tests driving the real =_listen= against a stand-in owner (it touches only =self._socket= and =self.toggle=, so no Gtk.Application is needed). The gate is the second toggle after an impatient first: it times out on unmodified code because no listener is left. The other two pin what the fix must preserve — the toggle fires even when the ack can't be delivered, and an unknown command is still answered without toggling. + +Left alone deliberately: =do_activate= calling =show_clock()= rather than =toggle()=. Changing it would alter what a cold =clock toggle= does on first launch, which is a design call for Craig rather than part of this defect. Worth raising if he ever wants the spawn path to toggle too. +** DONE [#B] fuzzel breaks the pinentry protocol loop on every passphrase :bug:security:gpg:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Found in the 2026-07-24 sentry bug-hunt, round 7 — from the live journal rather than from reading. Grepping this boot for tracebacks turned up four instances of =pinentry-fuzzel: line 36: read: 0: read error: Resource temporarily unavailable=, and every one sits 4-7 seconds after a =GETPIN= (the time it takes to type a passphrase). The =BYE= handler's log line never appears once. + +=hyprland/.local/bin/pinentry-fuzzel= speaks the Assuan pinentry protocol on a pipe gpg-agent keeps open, reading one command per iteration of =while read cmd rest=. The =GETPIN= arm shells out to fuzzel, which *inherits that pipe as its stdin*. fuzzel runs an event loop over its own input, so it sets =O_NONBLOCK= on fd 0 — and =--dmenu= would read the pipe as menu items besides. The flag lands on the shared open file description and outlives fuzzel, so the shell's next =read= fails with =EAGAIN= and the loop ends mid-protocol. + +Grading: Minor severity (the passphrase is delivered *before* the break, so decrypts still succeed and nothing is corrupted — what's lost is everything after: =BYE= is never acknowledged, and gpg-agent's same-connection retry after a wrong passphrase, =SETERROR= then =GETPIN= again, can't be served; that retry is what the script's "reenter" label exists for, and it has never once been reachable) × every user, every time (four for four in the journal, and the test reproduces it deterministically) = P2 = [#B]. + +Fixed as dotfiles =e727dcd= (committed locally, deliberately NOT pushed — held for Craig's morning review). =< /dev/null= on the fuzzel call, so the non-blocking flag lands somewhere harmless; =--lines 0= was already there, so no menu input was ever wanted. =ENABLE_LOGGING= became env-overridable as a test seam — the script logs through an absolute =/usr/bin/logger= that PATH can't shadow, so without it every test run would write ten lines into the real journal. + +New =tests/pinentry-fuzzel/=, 8 tests driving the real script over a live pipe the way gpg-agent does. The fake fuzzel sets =O_NONBLOCK= on whatever fd 0 it is handed, exactly as the real one does, which is what makes them a gate rather than a restatement of the fix. Four fail against unmodified code — one reproducing the journal's message verbatim — and one records the fd fuzzel was given, pinning the cause rather than the symptom. + +THE CALIBRATION NOTE, and it is about my own earlier sweep. This is the same shape as round 1's =a57c443= (ffmpeg draining the pipe a =while read= loop was consuming). Round 1 swept both repos for siblings of that bug and came back empty — because it searched for the *mechanism* (a child that drains stdin) rather than the *shape* (a child that inherits stdin at all inside a read loop). Two different mechanisms, one shape, and the narrower search missed a live daily-use instance. Scope a class sweep by shape, not by the mechanism of the first instance found. +** DONE [#C] a truncated webcam record strands every camera off :bug:settings:privacy:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Found in the 2026-07-24 sentry bug-hunt, round 8, sweeping production for non-atomic file writes. + +=settings/src/settings/webcam.py='s =_record()= wrote =~/.local/state/settings/webcam.json= with a plain truncate-in-place =open(path, "w")=. That record is the only route back on, and the module docstring says so: deauthorizing a camera removes its video4linux nodes, so =usb_devices()= returns nothing afterward and =_recorded()= becomes the sole source of the paths to re-authorize. A write that truncated and then failed left an empty file; =_recorded()= caught the resulting =JSONDecodeError= and returned =[]=; =_known_devices()= then had nothing; and =set_power(True)= returned None without re-authorizing anything. Every camera stranded off, with no way back through the panel until a replug or a reboot. + +The asymmetry, seventh instance of this read: six other state writers in the tree already write through a temp file and a rename — =maint/cache=, =net/cache=, =audio/ptt=, =timer/engine=, =settings/store=, =maint/curation=. The one whose loss is most expensive was the one that didn't. + +Grading: Major severity (grading the being-in-it — the privacy switch becomes one-way, the panel offers no route back, and the user has to know to replug the camera or write sysfs by hand; bounded by the fact that a reboot re-enumerates USB and restores authorized=1) × rare edge case (needs a crash or ENOSPC inside a microsecond-wide write window) = P3 = [#C]. + +Fixed as dotfiles =8b40b79= (committed locally, deliberately NOT pushed — held for Craig's morning review). =_record= now mirrors =store.save=: =mkstemp= in the target directory, write, =os.replace=, unlink the temp on any failure. Four tests; the gate is a =_record= whose =json.dump= raises, after which the previous record must still be readable — it isn't on the old code. The other three pin what the fix must preserve: no temp-file residue, the =_recorded()= round trip, and the end-to-end power-off/power-on with the class symlinks removed, which is the scenario the record exists for. + +HOW IT WAS FOUND, and it confirms round 7's lesson twice over. Round 4 ran an atomic-write sweep and reported "nine sites, six unique-per-writer, three sharing a fixed =.tmp=" — it enumerated the writers that *were* atomic and compared their temp-file naming, and never asked which state writers aren't atomic at all. Same narrowing that made round 1's stdin sweep miss the pinentry bug: the sweep was scoped to a property of the instances already found rather than to the shape of the hazard. +** DONE [#D] a failed wallpaper apply reports "nothing to apply" :bug:settings:quick:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as dotfiles =2cf3fb3=. The decision moved to =gui.wallpaper_apply_toast=, a module-level pure helper, because the callback lives inside a GTK widget where no test can reach it. 5 tests. +Found in the 2026-07-24 sentry bug-hunt, round 7, sweeping the settings panel's worker callbacks. + +=settings/gui.py='s =_async= passes an exception through as the *result* rather than as a separate error argument, so every =done= callback has to test =isinstance(res, Exception)=. Five do — =_mx_pin=, =_mx_letter=, =_after_matrix=, =_set_pointer=, the drum/dial/gallery/refresh callbacks. =_wp_apply= is the one that doesn't: + +#+begin_src python +def _wp_apply(self, note="Wallpaper set"): + self._async(lambda: panel.wallpaper_apply(self.state), + lambda ok: self._toast( + note if ok is True else "nothing to apply", + good=ok is True)) +#+end_src + +=panel.wallpaper_apply= calls =store.save=, which can raise =OSError= (disk full, a permissions change on the config dir). The exception then arrives as =ok=, =ok is True= is False, and the toast reads "nothing to apply" — describing a no-op when the apply actually failed. The toast is at least marked =good=False= (red), so the user gets a negative signal; what's lost is the reason, which every sibling callback surfaces via =str(res)=. + +Grading: Minor severity (wrong text on an error path, correctly marked as a failure, nothing corrupted) × rare edge case (needs =store.save= or =wallpaper.apply= to raise rather than return False) = P4 = [#D]. + +Fix: give it the same =isinstance(res, Exception)= arm its five siblings have — toast =str(res)= on an exception, keep the current two-way message otherwise. One callback, three lines. +** DONE [#D] two manage.py nmcli reads sit outside their own error conversion :bug:net:quick:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as dotfiles =2cf3fb3=. =_key_mgmt= converts both nmcli exceptions to "", which both call sites already treat as neither wpa-eap nor sae. 2 tests, including one driving =_classify_up_failure= end to end. +Found in the 2026-07-24 sentry bug-hunt, round 4, reading =net/src/net/manage.py=. =nmcli.run()= raises =NmcliTimeout= on timeout and =NmcliError= on a missing binary, and every mutation in this module is written to convert both into a result envelope. Two calls escape that conversion because they run through =_key_mgmt()=, which wraps =nmcli.get_value= and catches nothing: + +- =edit()= line 243 calls =_key_mgmt(uuid)= for the enterprise-profile refusal *before* its own =try=, while the next four lines catch exactly those two exceptions around =nmcli.run=. +- =_classify_up_failure()= calls it on =up()='s failure path, so a slow =connection show= turns a classifiable activation failure into an exception. + +Consequence is a leaked exception where the caller expected an envelope. The panel absorbs it — =gui.bg()= catches =Exception= and renders =str(e)= — so there it degrades to a worse message rather than a crash. =net edit= from the CLI has no such catch and prints a traceback. + +Grading: Minor severity (the operation fails either way; what's lost is the classified message, and only the CLI path shows a traceback) × rare edge case (=connection show= has a 2s timeout and nmcli's presence is already established by the time either site runs) = P4 = [#D]. + +Fix: give =_key_mgmt= the same conversion its callers use — catch =(nmcli.NmcliError, nmcli.NmcliTimeout)= and return "", which both call sites already handle correctly (neither "wpa-eap" nor "sae"). One =try= in one helper covers both sites. +** DONE [#D] three atomic writers share one fixed .tmp name :bug:quick:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Fixed as dotfiles =2cf3fb3=. All three carry =.tmp.$(getpid)=, matching the six writers that already did. 6 tests across audio and maint. +Found in the 2026-07-24 sentry bug-hunt, round 4, sweeping both repos for the temp-file half of the atomic-write idiom. The tree writes state atomically in nine places, and six of them make the temp path unique per writer: =net/cache.py= and =timer/engine.py= both use =f"{path}.tmp.{os.getpid()}"=, and =settings/store.py=, =settings/idle.py=, =bt/repair.py=, =net/probe.py= all use =tempfile.mkstemp=/=NamedTemporaryFile=. Three use a bare =path + ".tmp"=: + +- =audio/src/audio/ptt.py= =write_state= (the lead carried over from round 3's Next Steps) +- =maint/src/maint/cache.py= =put= +- =maint/src/maint/curation.py= =_write_user= + +=os.replace= makes the *rename* atomic, but a shared temp name is not: two writers open the same path, the second truncates under the first, and the file that gets renamed into place is a blend of both. The loser's own =os.replace= then raises =FileNotFoundError=, because the winner already renamed the name out from under it. + +Real concurrent-writer pairs exist for two of the three. =maint/cache.py= =updates_repo= is written by =maint-net-scan.timer= hourly and again by =doctor._fresh_pending()= at UPDATE fire time. =audio/ptt.py= has three writers by design (the CLI toggle bound to a key, the waybar right-click, and the GTK panel) — its module docstring says so. =curation.py= is written by panel key presses and CLI verbs. + +Grading: Minor severity (every reader degrades rather than crashes — =cache.get= catches =ValueError= and reports no data, =read_state= reads a torn file as disarmed, and both recover on the next write; the sharpest edge is the loser's =FileNotFoundError= aborting the rest of =scan_net=, which the next hourly run repairs) × rare edge case (the write window is a millisecond or two, and the overlapping writers are an hourly timer against a human keypress) = P4 = [#D]. + +Fix: give all three the =f"{path}.tmp.{os.getpid()}"= form the two careful siblings already use. It is three one-line changes and needs no new abstraction. Note this closes the torn-file half only — the read-modify-write in =ptt.toggle_plan= and =curation.set_preference= can still lose an update between two writers, which wants a lock rather than a temp-name change and should stay a separate decision. +** DONE [#C] dmenuexitmenu word-splits its menu so no entry matches :bug:dwm:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +Found in the 2026-07-24 sentry bug-hunt (shellcheck SC2128). =dwm/.local/bin/dmenuexitmenu= line 4 expands the menu unquoted: =choice=$(echo -e $menuitems | dmenu ...)=. Word-splitting collapses the runs of spaces the labels carry, so dmenu shows =Lock= where the =case= arm expects =Lock = (two spaces) and =Logout = where the arm expects a trailing space. No arm matches, so choosing an entry does nothing at all. + +CORRECTION (2026-07-24): THE BUG AS FILED DOES NOT EXIST. Ran it. =echo= rejoins the words split off the unquoted expansion with single spaces, and no label carries two spaces, so quoted and unquoted produce byte-identical output — verified against the exact literals from git rather than a retyped copy. Every =case= arm matches and every menu action works. + +What is real is latent. An unquoted expansion collapses a double space and glob-expands a =*=; the second was demonstrated turning a label into a directory listing. No current label triggers either. + +Hardened anyway in dotfiles =2cf3fb3= as robustness, not as a bug fix: the expansion is quoted and the bogus one-element array is now a plain string. Output confirmed unchanged byte-for-byte. New =tests/dmenuexitmenu/= (10 tests) pins the working behaviour, and shellcheck on the file drops from three findings to one. + +SECOND SENTRY FILING DISPROVED BY RUNNING IT, after =a57c443= (mkplaylist). Both came from a shellcheck hit plus reasoning, neither was executed. A static-analysis finding says a construct is unsafe, not that it currently misbehaves, and both filings treated the first as the second. +** DONE [#C] Timer module hero hierarchy :feature:waybar:timer:quick:solo: +CLOSED: [2026-07-24 Fri] +From the roam inbox (Craig, claimed 2026-07-22). Which display ("hero") wins the waybar timer module when several timer modes run simultaneously: pomodoro wins over everything (the user is actively working; it's likely their main focus). The rest rank in chronological order of when they would ring. Worked example: with a just-started 15-min timer, a 1-hr timer at 10 minutes left, a pomodoro, and an alarm ringing in 12 minutes — show the pomodoro; when it completes, the 1-hr timer (rings first), then the alarm, then the 15-min timer. Feeds the timer-panel spec (docs/specs/2026-07-02-timer-panel-spec.org). + +Shipped as dotfiles =9eedb39=. Pomodoro wins the hero, then soonest-to-ring, in both selectors (=engine.select_primary= for the bar, =panel.primary_id= for the GTK hero). Craig's worked example is a test. FLAGGED FOR CRAIG: the two selectors diverge on a *ringing* alarm (the bar excludes it, the panel gives it the hero) and I left that as-is rather than reverse a deliberate choice. Whether to unify them is your call. +** DONE [#C] Timer module: drop RING message, persistent notifications :bug:waybar:timer:quick:solo: +CLOSED: [2026-07-24 Fri] +From the roam inbox (Craig, claimed 2026-07-22). Remove the RING message from the timer module display; verify all timer and alarm notifications are persistent; the icon returns to normal once the notification has fired. Rationale: keeps timers and pomodoros from interfering with one another's displays (pairs with the hero-hierarchy task above). + +Shipped as dotfiles =9eedb39=. The tooltip no longer prints RING or a (ringing) suffix; a fired alarm shows its clock time and its persistent notification carries the alert. Verified the timer and alarm completion notes already set persist=True. +** DONE [#C] PTT icon outline removal :bug:waybar:quick:solo: +CLOSED: [2026-07-24 Fri] +From the roam inbox (Craig, claimed 2026-07-22): the waybar PTT icon should not have an outline. Cosmetic × every-glance = P3 = [#C]. + +Shipped as dotfiles =e63c0cf= (live style.css + dupre theme source). Removed the amber/green text-shadow glow from the armed/talk states, the only outline-like effect on the icon. FLAGGED FOR CRAIG: this is my read of "outline" (the glow). If you meant the glyph shape itself, it's a one-line revert. Confirm live by pressing PTT. +** DONE [#C] Video wallpapers don't fit the desktop :bug:dotfiles:solo: +CLOSED: [2026-07-24 Fri] +From the roam inbox (Craig, claimed 2026-07-23): videos don't fit the desktop in desktop-settings. The video channel drives mpvpaper (=settings/src/settings/wallpaper.py=); mpvpaper passes options through to mpv, so the fit is a =--panscan=/=--video-unscaled=/keepaspect question rather than a layout one. Reproduce with a video whose aspect differs from the output, pick the mode that fills without distorting (cover, matching how the image channels behave), and cover it in the wallpaper tests. Minor severity × whenever the video channel is selected = P3 = [#C]. + +Shipped as dotfiles =04d1489=. =set_video= now passes =panscan=1.0=, so mpvpaper fills the output and crops the overflow instead of letterboxing; keepaspect stays on so nothing stretches. Tested against the mpvpaper arg log. +** DONE [#C] World-clock wallpaper arrangement :feature:dotfiles: +CLOSED: [2026-07-24 Fri] +Shipped 2026-07-24 as dotfiles =6afbe09=, iterated live with Craig. The grid of boxed mini-clocks became a centered vertical clock line: cities down a spine, west (Honolulu) top to east (Wellington) bottom, labels alternating both sides, no boxes. Each shows city / time (12h) / day+date / timezone region name ("US Central"). Day/night dimming + amber home carried over, title dropped, cursor restored over the desktop. Prototypes archived in archsetup 40216e7. The face is parameterized (=?layout=vertical|horizontal=, =?hour12=1|0=) so the panel pickers below can drive it. +** DONE [#C] Floating layout — should we? :feature:hyprland: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +From the roam inbox (Craig, claimed 2026-07-23): consider whether Hyprland should offer a floating layout — how it would work, the benefits, and the complexity. A brainstorm/spike, not a build: the deliverable is an assessment Craig reads and decides on, not a shipped layout. Not :solo:. When picked up, run it as a brainstorm — how a floating mode coexists with the current tiling binds (toggle keybind, per-workspace vs global, window-rule interactions), what it buys over the existing =togglefloating=, and the config/muscle-memory cost — then bring Craig the recommendation. + +CONCRETE PROPOSAL from a second roam item (Craig, 2026-07-24 via work) — "floating mode as the easiest mode": +- Can't select floating until at least one window is displayed. +- Entering floating freezes each window's position and floats it exactly where it is. +- During floating, drag windows with mod+mouse-drag. +- Exiting floating switches to tiling or monocle and lets that layout take over. +Craig's note: "simple, could be useful for different reasons." This is the design the brainstorm should evaluate first — assess feasibility against Hyprland's actual float/tile transitions (does freezing current geometry survive the tiling↔floating switch, does re-tiling on exit reflow cleanly) before recommending. + +ASSESSED, dotfiles =8cf4728=: =docs/2026-07-24-floating-layout-assessment.org=. Verdict: buildable and worth building on a capture-then-restore of window geometry (=hyprctl clients -j= gives at/size), which is a real gesture plain =togglefloating= can't express. Craig's four-rule proposal is folded in and each rule assessed. One taste call flagged (exit to previous layout vs always monocle). Ready to file a build task on Craig's go. +** DONE [#C] World clock wallpaper: bold the city names :feature:dotfiles:quick:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +From the roam inbox (Craig, 2026-07-24 via work): bold the city names on the world-clock wallpaper face (=settings/faces/world.html=, shipped =6afbe09=). Cosmetic × every glance at the world face = P3 = [#C]. Solo — a CSS weight change, screenshot-verifiable — but it's a visual call, so build it and show the render rather than close off a green suite. Pairs with the open world-face picker task. + +Shipped as dotfiles =e63c0cf=. =.lbl .city= is now =font-weight:700=. Rendered offscreen and confirmed the bold reads well over the time/zone lines; home city stays amber. Comparison render was on ws5 for Craig. +** DONE [#C] Floating clock toggles on control+mod+c :feature:dotfiles:hyprland:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +From the roam inbox (Craig, 2026-07-24 via work): a control+mod+c keychord should toggle the floating clock, the same as clicking the time waybar module. + +This answers the design question the round-6 clock-toggle fix deliberately left open (see the =clock toggle listener= DONE task above): =do_activate= calls =show_clock()= rather than =toggle()=, and the note there flagged "worth raising if he ever wants the spawn path to toggle too." He does. Build: a hyprland keybind bound to =clock toggle=, and confirm the toggle path (not show-only) fires whether the service is cold or warm. Solo — buildable and locally verifiable. + +Shipped as dotfiles =e73a70e=. =bind = $mod CONTROL, C, exec, clock-panel toggle= reuses the exact command the time module's click runs, so it toggles identically. Registered clean on reload. Live keypress is Craig's to confirm. +** DONE [#C] Calculator scratchpad won't toggle closed on mod+x :bug:hyprland:solo: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +From the roam inbox (Craig, 2026-07-24 via work): =mod+x= opens the calculator scratchpad but doesn't close it — Craig has to kill the window by hand. A second =mod+x= should toggle it shut. Almost certainly a =togglespecialworkspace= vs plain =exec= binding in the hyprland config, or a scratchpad window-rule mismatch. Minor severity (a workaround exists: kill the window) × every time the calc scratchpad is used = P3 = [#C]. Solo — a keybind/window-rule fix, locally verifiable. + +Shipped as dotfiles =e73a70e=. New =calc-toggle= script (mirrors fuzzel-toggle: pgrep -x, pkill or launch), and =mod+X= now points at it, so a second press closes the calculator. 3 tests in tests/calc-toggle. +** DONE [#C] Saving and recalling window configurations :feature:hyprland: +CLOSED: [2026-07-24 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +From the roam inbox (Craig, 2026-07-24 via work), a research idea: Craig wants to save a specific window+app arrangement and have it reappear on demand. What has to be known and built to make that happen — is there prior art (another WM or OS that does session/layout save-restore), what information do those need (app identity, geometry, workspace, launch command), and what are their rules. Explore how far Hyprland can get (hyprctl clients + dispatch, exec rules, window rules by class/title), document thoroughly, and review with Craig next time. Not :solo: — the deliverable is an assessment he reads and decides on, and it may spawn a build task once the shape is clear. Offer to file the build separately if part of it turns out urgent. + +RESEARCHED, dotfiles =8cf4728=: =docs/2026-07-24-window-config-save-recall-assessment.org=. Prior art surveyed (i3/sway =append_layout= swallow, KDE window rules, macOS Moom). Three tiers from cheapest: (1) reposition open windows — buildable + testable now; (2) relaunch + place by class rule; (3) full swallow-by-title, which hits the same-class ambiguity every tool hands back to the user. Recommends shipping tier 1; tiers 2-3 need Craig's call on how much manual disambiguation he'll accept. +** DONE [#B] Weather tooltip caching :feature:waybar:weather:solo: +CLOSED: [2026-07-25 Sat 10:53] +From the roam inbox (Craig, claimed 2026-07-22): retrieve the weather tooltip data once per hour and cache it. If the network is unavailable, display the cached tooltip with explanatory text saying so. Dotfiles-side work (archsetup owns the lifecycle); touches common/.local/bin/weather. +Verified complete in the 2026-07-25 batch: the weather CLI already had the hourly default TTL, fresh-cache no-fetch path, stale fallback, and explicit offline footer. Its 33-test suite and the full dotfiles suite pass. +** DONE [#B] Settings gear becomes four device toggles :feature:waybar:dotfiles:solo: +CLOSED: [2026-07-25 Sat 10:53] +From the roam inbox (Craig, claimed 2026-07-23): the waybar gear should become four icons — touchpad, mouse, webcam, and a notification bubble. Clicking each toggles that setting directly. The first three turn red when disabled; the bubble turns red when DND is enabled. + +Today =custom/settings= (=hyprland/.config/waybar/config=) is one gear glyph () whose only job is =on-click: settings-panel=. The toggles themselves already exist and are tested — the settings package owns touchpad, mouse, and webcam (=webcam.py= is the USB-authorized kill switch from 2026-07-22), so this is a bar-side surface over existing backends rather than new capability. + +Note the state-polarity split when wiring the colors: three read "red = off" and DND reads "red = on". That asymmetry is deliberate (red means "something is disabled that normally isn't, or suppressed that normally isn't"), so encode it per-icon rather than deriving one rule. + +Decided 2026-07-23 (Craig): the gear STAYS alongside the four toggles as the panel launcher. So the bar's right side grows from 12 modules to 16 — the four toggles are net-new, the gear keeps its =on-click: settings-panel=. Open sub-question for build time, not blocking: whether the four toggles are four separate waybar modules or one custom module rendering four glyphs (fewer layout entries, one exec). Pick at build; the four-module shape is simplest and matches how mic/net already sit as individual modules. +Shipped in the 2026-07-25 batch as four independent JSON modules over the existing verified settings backends. Touchpad, mouse, and webcam turn terracotta when disabled; DND uses the deliberate inverse polarity; unavailable hardware dims. The gear remains the panel launcher. The live and Dupre theme CSS copies stay byte-identical. +** DONE [#C] Wallpaper panel selection and scroll state :feature:dotfiles:solo: +CLOSED: [2026-07-25 Sat 10:53] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-25 +:END: +From the roam inbox (Craig, 2026-07-25). Screenshot: =~/pictures/screenshots/2026-07-25_013041.png=. Three related behaviors in the settings wallpaper panel (=settings/src/settings/wallpaper.py=): +1. Open at the wallpaper currently displayed, not the top of the list. +2. Highlight that wallpaper as selected in the scrollable pane while it shows in the preview. +3. Keep the scroll position when a picture is selected. Today selecting a picture snaps the scroll back to the top, which is the bug half of this. +Grade: minor scroll-reset defect x every panel selection = P3 = [#C]; the open-at-current and select-current behaviors are enhancements at the same level. One type tag, so filed =:feature:= with the scroll-reset called out as the bug. Solo: buildable in the settings GTK panel, agent-verifiable via headless capture plus the wallpaper.py tests, no design call — swww query gives the current wallpaper, and scroll-position preservation and row selection are standard GTK. +Shipped in the 2026-07-25 batch. The panel queries =awww query= off the UI thread, prefers the actually displayed image over stale stored state, highlights it, scrolls it into view on first open, and remembers the horizontal adjustment across selection-triggered rebuilds. +** DONE [#C] Net tooltip IPs and line order :feature:waybar:network:solo: +CLOSED: [2026-07-25 Sat 10:53] +From the roam inbox (Craig, claimed 2026-07-23): in the wifi hover, add the internal IP, external IP, and gateway IP just below the Interface line; move the Signal line to just above the keyboard-shortcuts line. Design constraint: the bar's hot path does no network I/O (status.py deliberately skips _address_facts on the 2s beat) — internal IP + gateway can ride cheap local reads, but the external IP must come from a cache the connectivity probe refreshes, never a live lookup in waybar-net. +Shipped in the 2026-07-25 batch. The slow connectivity probe caches local addressing and a validated external IP with the network identity; the Waybar hot path only reads that valid cache. Tooltip order is Interface, internal/external/gateway IPs, connectivity detail, throughput, Signal, shortcut. +** DONE [#B] Dupre Kit merge — casting additions :feature:tooling:solo: +CLOSED: [2026-07-25 Sat 10:53] +Fold docs/prototypes/dupre-kit-additions.js back into the kit proper: detentFader (NEW — multi-detent slide attenuator with speedbump drag physics: magnet + escape hysteresis, parked tick glow) and the drumRoller redefinition (UPGRADE — 1..N channels and min/max range; stock hardcodes two drums and throws on one, defaults reproduce stock exactly) and the guardedToggle redefinition (UPGRADE — lever throws with rotateX so it flips toward the viewer instead of the stock 180° planar spin that sweeps sideways mid-transition; contract unchanged). Merge means: builders into widgets.js, the additions CSS into DUPRE_CSS, additions-scoped gradients into the shared defs plate, gallery cards for both in panel-widget-gallery.html, and POLICY entries. Origin: the desktop-settings casting sitting 2026-07-21 — Craig's direction is that components get finished by being needed ("the ones needed most will have had the most attention"), so more additions may accrue here before the merge; batch them. +Shipped in the 2026-07-25 batch. =widgets.js= now owns all three builders, shared gradients/CSS, contracts, and policy records; additions no longer redefines them when older casting pages load it. The gallery has a three-detent fader card and a three-channel 0–12 drum demonstration (112 cards total). Static ownership tests, JS syntax checks, and the complete headless interaction probe pass. +** DONE [#C] Maint live-refresh hairline replacement :feature:maint:solo: +CLOSED: [2026-07-25 Sat 10:53] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-14 +:END: +From the roam inbox (routed 2026-07-13): the memory-killer section seemed to update too often, and "it's a bit unclear what the line is doing; consider something else." Diagnosis (2026-07-14): the data cadence is already the requested 3s (gui live tier, _LIVE_SECONDS); the perceived churn is the live-refresh hairline — the 2px bar under the live sections that drains full-to-empty over each 3s window, redrawn at 150ms (gui._hair_tick, viewmodel.refresh_fraction). It exists to tell a stale board from a frozen one (2026-07-09), but it reads as constant unexplained motion. Design call for Craig: replace the draining line with something whose meaning is legible — candidates: a dot that blinks once per refresh, a "3s" age caption that only appears when refresh is overdue, slowing the drain redraw, or dropping the indicator on live tiers and keeping it only when data goes stale. Keep the stale-vs-frozen distinguishability that motivated the hairline. +*** 2026-07-21 Tue @ 08:35:00 -0500 Decided (Craig): silent-until-stale age caption +Replace the draining 2px hairline with an age caption that shows ONLY when refresh is overdue (e.g. "3s", "8s" once past the expected window) and shows nothing while the board is healthy. This keeps the stale-vs-frozen signal — a frozen board surfaces a growing age number, a live one stays clean — while removing the constant motion the hairline created. Implementation (dotfiles, archsetup-owned): drop =gui._hair_tick= / the hairline draw, add an overdue-age caption driven off =viewmodel.refresh_fraction= (or the last-refresh timestamp) rendered only past the live window. Now unblocked; needs a live visual check on the panel after. +Shipped in the 2026-07-25 batch. The animated draw area and 150ms timer are gone; the memory section header stays silent through the healthy three-second window, then shows a once-per-second growing age caption. Pure boundary tests and the full maintenance suite pass. +** DONE [#D] Test-framework + prototype refactor cluster :refactor:solo: +CLOSED: [2026-07-25 Sat 10:53] +Grading: no behavior change; parking lot. Refactors from the S5-S7 audit, distinct from the installer refactor rollup above. +scripts/testing/run-test.sh + run-test-baremetal.sh duplicate the run/poll/report skeleton and have drifted (VM uses setsid + copy helpers, baremetal uses nohup + hand-rolled sshpass scp) — extract the shared core so baremetal inherits the sturdier paths; run-maint-nspawn.sh:66 + run-maint-scenarios.sh:78 duplicate the transport-independent _scenario_var/_validate_scenario/run_scenario (a sourced lib/maint-scenario.sh); run-test.sh:251,265 uses two different mechanisms (pgrep vs ps|grep) for the same liveness check; docs/prototypes/gen_tokens.py:78 repeats the section-iteration skeleton across four emitters; gallery-widget.el:95,136 hardcodes SVG arc/hub path strings that duplicate the cx/cy/radius geometry (dial desyncs silently on a constant change); gallery-widget.el:72,84 leans on the private svg--append. See findings doc (S5, S6, S7). +Completed test-first in the 2026-07-25 batch. QEMU and bare-metal runners share liveness/report helpers; maintenance transports share scenario validation/execution; token emitters share ordered section traversal; and the Emacs SVG gauge shares semicircle geometry and uses the public DOM append API. Every fast Python/ERT suite passes. +** DONE [#B] Two agent sessions sharing one git repo :chore:tooling: +CLOSED: [2026-07-26 Sun] +Craig approved the shared-rules-layer solution on 2026-07-26. + +Use one repository-scoped publish lock for every session and worktree sharing a clone. Derive the lock name from the real Git common-directory path; hold it across reconcile, stage, staged review, and commit; track the owning session and reviewed staged-tree fingerprint; refresh it after conversational waits; and repeat the staged review if ownership or the fingerprint changed. Ordinary working-tree edits remain concurrent. + +An approval waiver never waives the staged review, because that review is the gate that reads the actual hunks entering the commit. Rulesets owns the implementation in =commits.md=, =agent-lock=, and its Bats coverage; archsetup sent the approved implementation package through the rulesets inbox. +** DONE [#A] Reboot ratio to activate amdgpu.runpm=0 :bug:hyprland:ratio: +CLOSED: [2026-07-28 Tue] DEADLINE: <2026-07-28 Tue> +:PROPERTIES: +:CREATED: [2026-07-28 Tue] +:LAST_REVIEWED: 2026-07-28 +:END: +Craig's plan: close everything down, run topgrade, then reboot. Alarm set for 08:00 (=at= job 56, persistent desktop notify). + +=amdgpu.runpm=0= sits in =/etc/default/grub= and in the generated =/boot/grub/grub.cfg= (5 occurrences, so the reboot will actually apply it) but is absent from =/proc/cmdline=. The box has been up since 2026-07-22 21:10 and the fix landed 2026-07-24, so the running kernel predates it. The GPU is AMD Strix Halo (Radeon 8060S, =1002:1586=), exactly what the parameter targets: runtime power management invalidates the GPU resources hyprlock holds across a display power-cycle, so hyprlock exits without unlocking. + +That is the root cause under the 2026-07-27 lockdead screen. The screen-lock flock fix (dotfiles =ec18fd7=) stops one dead client from becoming a lockdead screen, but it treats the symptom -- this reboot treats the cause. + +Not :solo: — Craig closes his own session and runs topgrade first. + +Rebooted 2026-07-28 08:59. =amdgpu.runpm=0= confirmed present in =/proc/cmdline= afterward, so the parameter is finally live. + +Correction, 2026-07-29: the claim above and in the body that this is "the root cause under the 2026-07-27 lockdead screen" is wrong, and superseded. hyprlock was never crashing. Every logged exit is =rc=143=, SIGTERM, from =settings-watch= killing it by design. See =[#B] Night watch and the lock watchdog fight each other=. The reboot was still worth doing (the parameter is a genuine mitigation for a real AMD defect) but it did not fix this, and the lockdead screens continued after it. +** DONE [#B] Caffeine state is unreadable on both surfaces :bug:dotfiles:design:solo: +CLOSED: [2026-07-28 Tue] +:PROPERTIES: +:CREATED: [2026-07-28 Tue] +:LAST_REVIEWED: 2026-07-28 +:END: +Neither surface that reports caffeine tells the truth reliably, so there is no way to know at a glance whether the screen will lock. Found while investigating the 2026-07-27 lockout, where Craig believed caffeine was on and the screen locked anyway. + +Defect 1 — the settings panel shows a frozen value. =gui.py= calls =_refresh_async()= once during window construction (line 454) and again only after the user's own actions (=_after_matrix=, line 663). The only two =GLib.timeout_add= calls are one-shots (the 2400ms toast hide and a 350ms fire), so nothing re-reads state on a timer. An open panel therefore displays the caffeine value from the moment it opened, forever. Any external flip -- the waybar click, Super+I, the =caffeine-toggle= script -- leaves it stale with no self-correction. The panel is the only surface in the repo carrying a caffeine control (=panel.py:25=); maint has none and does not embed these toggles, so this is the display Craig read. + +Defect 2 — there is no caffeine indicator on the bar at all. =custom/caffeine= appears in neither the stowed =hyprland/.config/waybar/config= nor the live generated =/run/user/1000/waybar/config=, and no =custom/caffeine= block is defined anywhere in the waybar config dir. The =waybar-caffeine= script exists, works, and has its own passing test suite, but nothing displays it. So the bar has never been a source of caffeine state, and the keybind and script have been signalling (=pkill -RTMIN+8 waybar=) a module that isn't there. + +(An earlier read of this task said the bar showed two near-identical glyphs. That was wrong: the module is absent, not merely unstyled. The script's class names are still backwards -- =active= when caffeine is OFF, =inhibited= when ON -- and neither class is styled, but both points are moot until the module is actually in the bar.) + +Grading: Major severity (the panel reports state wrongly while it is open, and the only other surface does not exist, so there is no reliable source for a setting Craig actively manages) x most-of-the-time (any external toggle while the panel is open; the bar never shows it) = P2 = [#B]. + +Fix all three. Wire =custom/caffeine= into the bar, rename its classes so they describe caffeine rather than idle, and style them from the existing palette. Give the panel's toggle row a re-read on a timer or on focus-in. Solo -- buildable and testable, and the direction is settled by the defects rather than a taste call, though the bar color is worth a glance from Craig once it renders. + +All three shipped as dotfiles =033076c=, pushed. =custom/caffeine= now sits in the bar between DND and settings on =interval: 2=; classes renamed =on=/=off= and both styled, caffeine-ON in the theme's gold =#dab53d=; the panel re-reads live state every 3s while visible. Verified live in the stowed config and the generated =/run/user/1000/waybar/config=. The full suite caught a theme-copy regression (=themes/dupre/waybar.css= out of sync with =waybar/style.css=) that the focused suites missed. +** DONE [#B] hyprlock still exits mid-lock; the watchdog relaunch is silent :bug:hyprland:dotfiles: +CLOSED: [2026-07-29 Wed] +:PROPERTIES: +:CREATED: [2026-07-28 Tue] +:LAST_REVIEWED: 2026-07-28 +:END: +Craig, 2026-07-28 ~15:00: saw the Hyprland lockdead/error text blurred *behind* a working lock screen; it vanished when he authenticated. + +That ordering is the diagnosis. hyprlock's blur samples what the compositor is currently rendering, so the compositor was already showing lockdead when the new hyprlock attached. Sequence: hyprlock exits non-zero (no coredump, so it exits rather than crashing), Hyprland renders lockdead because the client is gone while the session stays locked, =screen-lock='s watchdog relaunches within =LOCK_RELAUNCH_DELAY= (0.5s), and the new client draws over the lockdead frame and blurs it. + +*The recovery worked.* On 2026-07-27 this same hyprlock exit produced two contending clients and a session recoverable only from another console. It now self-heals in half a second, and the residue is cosmetic. Both the flock guard (dotfiles =ec18fd7=) and the watchdog did their jobs — verified in this session's compositor log, where all four lock events created exactly one =sessionLock= and one =sessionLockSurface= each, against two of each on 2026-07-27. + +Two things remain. + +*Why hyprlock exits.* The wrapper's header blames GPU-resource invalidation across a display power-cycle (hyprlock#953), which =amdgpu.runpm=0= targets — and that parameter is live as of the 2026-07-28 08:59 reboot, confirmed in =/proc/cmdline=. There is also no DPMS idle rule any more (=e900903=), so idling never power-cycles the display. Yet hyprlock still exited. Strongest untested candidate: a screen recording (=wf-recorder= into =~/sync/recordings/2026-07-28-12-53-57.mkv=, running 12:53 until Craig killed it) held screencopy sessions on DP-4 across the lock. The compositor log carries 2454 screenshare sessions and a =CScreencopyProtocol= bind in the window between the last two locks. A screencopy client churning dmabufs alongside hyprlock's own is a plausible way to invalidate them, and it was the one large new variable that day. + +*The relaunch is silent.* The watchdog loop re-runs hyprlock and logs nothing, so there is no record of how often this fires, when, or with what exit code — which is exactly why the frequency couldn't be established from the logs. Log the exit code and a timestamp on each relaunch. + +Grading: Major severity (the lock client dies mid-lock, and the pre-fix version of this wedged a session unrecoverably) x most users frequently (twice in three days, and this is a single-user machine, so every occurrence lands on the only user) = P2 = [#B]. Downgraded from the 2026-07-27 [#A] because the wedge is fixed and the failure now self-heals. + +An earlier draft of this grading said "some users sometimes", which the matrix maps to P3 = [#C], not the [#B] written beside it. The frequency row was the wrong input rather than the letter: on a one-user machine a fault hitting twice in three days is frequent, not occasional. Corrected the input per the rule that a disputed grade is fixed at its inputs. + +Solo for the instrumentation half only: adding the relaunch logging is buildable, testable against the existing =tests/screen-lock= suite, and needs no decision. Diagnosing the exit is not solo — it needs a reproduction, and the likely trigger is Craig recording his screen. + +Next step when picked up: land the relaunch logging first so the next occurrence produces evidence, then try to reproduce by locking with =wf-recorder= running. + +Superseded 2026-07-29 by =[#A] Night watch and the lock watchdog fight each other=. The logging landed (dotfiles =5bbe2c3=) and answered it within hours: three =rc=143= entries, SIGTERM, from =settings-watch= killing hyprlock by design. Nothing was crashing, so both the AMD-iGPU and the screen-recorder hypotheses in this task are wrong. Kept closed rather than deleted because the reasoning that led here is worth the record. +** DONE [#A] Idle commits silently drop the screen-lock wrapper :bug:hyprland:dotfiles:security: +CLOSED: [2026-08-04 Tue] DEADLINE: <2026-07-29 Wed> +:PROPERTIES: +:CREATED: [2026-07-29 Wed] +:LAST_REVIEWED: 2026-07-29 +:END: +Caught live 2026-07-29 05:30, seconds after it happened, while verifying that Craig's watch-stage change had landed. + +=idle.py= renders the *whole* hypridle.conf, including a hardcoded =GENERAL= block. That block said =lock_cmd = pidof hyprlock || hyprlock=. The live config said =|| screen-lock=. So every idle-stage commit through the panel rewrote =lock_cmd= and dropped the wrapper out of the chain. + +The wrapper is not incidental. It carries the flock duplicate guard (the fix for the 2026-07-27 unrecoverable wedge), the crash-relaunch watchdog, and the relaunch log. Parking one stage removed all three in a single write, and nothing said so. + +=tests/settings/test_settings.py:562= asserted the bare =|| hyprlock= form, so the suite *enforced* the regression. That is why 3845 tests stayed green through a day of work on exactly this subsystem. A test can pin the bug as readily as the fix. + +The false-negative this sets up is worth naming: with the wrapper gone the relaunch log stops receiving entries, and an empty log reads as "the problem is fixed" when it means "the instrument was removed". The =screen-lock= header already warns that an empty file is not proof; this is the mechanism that would have produced one. + +Fixed by dotfiles =ab059fb= (2026-07-29 05:59). Verified 2026-08-04 against the tree rather than the commit message: =idle.py:47= renders =lock_cmd = pidof hyprlock || screen-lock || hyprlock=, the live =hypridle.conf= matches, and =test_settings.py= now asserts the wrapper is in the chain plus a second test for the bare-hyprlock fallback. The test that used to pin the bug now pins the fix. + +The evidence that matters is the one this task named: =~/.local/var/log/screen-lock.log= is *receiving entries*, so the instrument is present. An empty log was the false negative to fear, and it did not happen. + +FIXED here, TDD, in the working tree pending commit: +- =idle.py= =GENERAL= now names =screen-lock=, with a comment saying why the line is load-bearing. +- The test now pins the wrapper form. Red first against the old template. +- Live config rewritten through the panel's own path and hypridle restarted; =lock_cmd= confirmed back to =screen-lock=, one hypridle running. + +Grading: Critical severity (=write_conf= truncates, so any hypridle key the renderer does not model is silently deleted rather than preserved — that is configuration data loss, and the =lock_cmd= case proved it happens in the field) x some users sometimes (only when an idle stage is committed, which is rare) = P2 = [#B]. + +An earlier draft graded this [#A] on a "security carve-out". That was wrong: disarming the guard is an availability problem, not a leak, and the carve-out is for privacy, security, compliance and safety. The severity band is what carries the weight here, and silent deletion of configuration is the =Critical= band's data-loss case. + +Two further fixes came out of an independent review of the first one: + +- *Fail-open restored.* =pidof hyprlock || screen-lock= made the wrapper the end of the chain, and =screen-lock= is a stow symlink in =~/.local/bin=, not a system binary. An unstowed tree, or a hypridle started without =~/.local/bin= on PATH, resolves it to 127 — so the screen would never lock *at all*. That is worse than the duplicate client the wrapper prevents. The chain now ends =|| hyprlock=, matching the wrapper's own fail-open discipline. +- *The file now says it is generated.* Three comment lines at the top of the rendered output name the renderer and warn that edits are overwritten. The absence of that header is how the divergence survived unnoticed. + +Still open, and why this stays a task rather than closing with the fixes: the header warns, but nothing *prevents* the next divergence, and the exposure is wider than =lock_cmd= alone. The review enumerated it: + +- =before_sleep_cmd= and =after_sleep_cmd= sit in the same hardcoded block, at identical risk. +- Every stage command is hardcoded in =_stage_commands= (brightness level, lock, watch, dpms, suspend), same one-way overwrite. +- =write_conf= *truncates* rather than merges, so any hypridle key the renderer does not know about (=ignore_dbus_inhibit=, =ignore_systemd_inhibit=, =inhibit_sleep=, =on-lock=, =on-unlock=) is deleted rather than preserved. That is the largest hole: a key nobody has added yet would vanish the first time a stage is parked. + +Options: have the renderer preserve the existing general block and unknown keys instead of emitting its own, or accept the template as the single source and move every hypridle setting into the panel. A design call for Craig, and the truncation half is the part that will bite next. +** DONE [#A] Comet KVM setup for truenas :feature:infra:truenas: +CLOSED: [2026-08-08 Sat] +:PROPERTIES: +:CREATED: [2026-07-27 Mon] +:LAST_REVIEWED: 2026-07-27 +:END: +Resolved: Craig wired up and configured the Comet himself, confirmed working +2026-08-08. The ATX power-board follow-up (hard power-cycle for a truly wedged +box) remains unfiled — raise it if the next outage shows the KVM alone isn't +enough. +Wire up the GL.iNet Comet (GL-RM1) IP KVM against truenas. It was bought 2026-01-14 for exactly this job and its KB node still reads "Arrived, not yet set up." + +Why now: truenas went dark 2026-07-24 and stayed unreachable. Diagnosis from ratio on 2026-07-27 — no tailnet contact for 3 days, 100% packet loss on 192.168.86.5, ARP entry FAILED (nothing answers ARP for the address, so the NIC is down at layer 2), every service port closed, while the gateway and a dozen other LAN hosts stayed reachable. Wake-on-LAN to 70:85:c2:db:9d:94 drew no response. With no console and no out-of-band power control there was no remote remedy at all, so recovery needed hands on the box. The Comet closes exactly that gap: BIOS/UEFI console, Wake-on-LAN, and browser access over its native Tailscale integration. + +Not :solo: — the physical cabling is Craig's, and the Tailscale enrollment needs his account. + +Steps, from the KB node ([[id:67bc5994-a763-48e2-926f-4ac0d1bad3db][GL.iNet Comet (GL-RM1) - KVM]]): +1. HDMI from truenas video out to the Comet's HD IN. +2. USB-A-to-USB-C from the Comet to a truenas USB port (keyboard/mouse emulation). +3. Ethernet to the network. +4. Power via USB-C (5V/2A). +5. Reach the web interface and enroll it in Tailscale, so it's usable when the LAN side of truenas is the thing that's broken. + +Then verify while truenas is healthy, rather than discovering the gaps during the next outage: confirm the console shows POST and the BIOS, that keyboard input reaches the box, and that Wake-on-LAN from the Comet actually powers it on. Enable WOL in the truenas BIOS if that last check fails — this outage never established whether it was on. + +Worth considering as a follow-up: the ATX power board accessory gives hard power-cycle control for a truly wedged box, which the KVM alone can't do. +** DONE [#A] Review post-archsetup laptop setup steps (velox 2026-04-10) +CLOSED: [2026-08-08 Sat] +:PROPERTIES: +:LAST_REVIEWED: 2026-08-08 +:END: +Closed at the 2026-08-08 session: every open item got its automate-vs-document +call and the work landed the same night (tests green, committed). Residual: +velox itself still needs the new tlp.d radio line and a dotfiles pull — folded +into the [#A] sleep/suspend task, which works the same files on velox anyway. +Items discovered during velox setup that needed manual intervention after archsetup. +Decide which should be automated in archsetup vs documented as post-install steps. + +*** 2026-08-08 Sat @ 04:43:42 -0500 Automated radio enable via TLP (rfkill boot soft-block) +Root cause sharpened during triage: archsetup masks systemd-rfkill on laptops +(it fights TLP), so nothing restored radio state at boot — the "unblock once +should stick" premise was wrong under the mask. Fix in the TLP custom conf: +=DEVICES_TO_ENABLE_ON_STARTUP="bluetooth wifi"=, the TLP-native mechanism. +configure_tlp_power parametrized for tests; covered by +tests/installer-steps/test_configure_tlp_power.py. + +*** 2026-07-04 Sat @ 11:48:24 -0500 Automated /efi restrictive mount permissions in fstab generation +archsetup:2827-2836 now rewrites the /efi fstab line to =fmask=0177,dmask=0077= (idempotent), so fresh installs no longer land the world-accessible =fmask=0022,dmask=0022= default. Confirmed via the 2026-07-04 task audit. (Original velox note: default vfat mount had =fmask=0022,dmask=0022=, hand-fixed to restrictive; bootctl warned about a world-accessible random-seed file.) + +*** 2026-08-08 Sat @ 04:43:42 -0500 Automated tmp.mount mask for ZFS /tmp +New mask_tmp_mount_for_zfs, called from configure_snapshots' ZFS branch: +masks tmp.mount only when the pool actually carries a dataset mounted at +/tmp (exact match), silent no-op without zfs or without the dataset. Covered +by tests/installer-steps/test_mask_tmp_mount_for_zfs.py; the orchestrator +dispatch pin updated. + +*** 2026-08-08 Sat @ 04:43:42 -0500 Automated CPU microcode install by vendor +New install_cpu_microcode, first in boot_ux so grub-mkconfig and mkinitcpio's +microcode hook both see the installed /boot/<vendor>-ucode.img: vendor_id from +/proc/cpuinfo → intel-ucode / amd-ucode, error_warn on unknown vendor. +Covered by tests/installer-steps/test_install_cpu_microcode.py; boot_ux +sequence pin updated. + +*** 2026-07-04 Sat @ 11:48:24 -0500 Automated syncthing user-service enable in archsetup +archsetup:2263-2271 now installs syncthing and enables the user service (via symlink), so fresh installs no longer leave it installed-but-disabled. Confirmed via the 2026-07-04 task audit. (Original velox note: package installed but service not enabled; hand-fixed with =systemctl enable --now syncthing@cjennings=.) + +*** 2026-08-08 Sat @ 04:43:42 -0500 Closed the awww-daemon crash watch — no recurrence +The April boot crash never recurred across four months of daily use on both +machines (and the wallpaper stack has since been reworked). Reopen as its own +bug with fresh evidence if it ever comes back. + +*** 2026-08-08 Sat @ 04:43:42 -0500 Automated touchpad device detection in the pointer scripts +The scripts were already in stowed dotfiles with binds — the open half was the +hardcoded Framework device name. Both touchpad-auto and toggle-touchpad now +auto-detect the touchpad (first pointer named *touchpad*, pixa fallback) and +derive the internal-pointer exclusion set from the detected name, so they +agree on any machine. Test seams added (--detect / --has-external-mouse); +tests/touchpad-auto/ new, toggle-touchpad suite still green. Dotfiles commit; +velox picks it up on its next pull. + +*** 2026-08-08 Sat @ 04:43:42 -0500 Documented bluetooth pairing in the post-install checklist +Inherently interactive, so it can't ride the installer. Documented in the new +[[file:docs/post-install-checklist.org][docs/post-install-checklist.org]] along +with the Proton Bridge steps — the standing home for manual post-install work. +Consider: document as post-install step. No automation possible. + +*** 2026-05-26 Tue @ 13:32:31 -0500 pocketbook install concern moot — pulled from publication, folded in-tree +Resolved by removing pocketbook from archsetup's provisioning entirely. It's nowhere near ready, so the github mirror + cjennings.net repo were deleted and the project was folded into the archsetup tree at =pocketbook/=. Dropped the =gtk4-layer-shell= dep + =pip_install= from =archsetup= and the clone from =scripts/post-install.sh=. No fresh install pulls pocketbook now, so "not installed on velox" no longer applies. Re-wiring the install is tracked in the new pocketbook development backlog. + +*** TODO Review: Tailscale needs login after install +~tailscaled~ service was enabled but needed ~tailscale up~ for interactive auth. +Old machine entry needed cleanup in admin console. +Consider: document as post-install step. + +*** TODO Review: docs/ directories need manual sync from existing machine +docs/ dirs (gitignored) for ~/code and ~/projects repos needed scp/rsync from ratio. +Same for ~/.emacs.d/docs/. Not in git, so not available after clone. +Consider: document as post-install step or create a sync script. +** DONE [#C] Waybar modules run together — need subtle separators :bug:dotfiles:waybar: +CLOSED: [2026-08-08 Sat] +Closed at the 2026-08-08 task review: Craig confirms the separator work landed +a while back and the bar reads correctly now. +Craig misreads where one module ends and the next begins — the wind (weather) value runs straight into the date with no visual stop, so he reads the wind figure as the start of the date. Add a light, subtle separator or spacing between adjacent Waybar modules. +Grading: Minor severity (legibility, nothing broken) x frequent (every glance at the bar) = P3 = [#C]. +Not fully :solo: — needs Craig's eye on the result (separator style is a taste call, plus a live visual check). Prior work added a date-facing divider (dotfiles 103cccb); evidently not enough, so revisit the whole inter-module treatment rather than just the weather/date seam. From .emacs.d handoff 2026-07-20-1114 (roam capture; waybar is archsetup-owned per the dotfiles standing rule). +** CANCELLED [#C] Add a whole-display dim mode :feature:hyprland: +CLOSED: [2026-08-08 Sat] +Killed at the 2026-08-08 task review: the July auto-dim work covers the actual +need; no separate dim-everything mode wanted. +Extend auto-dim with an explicit “dim everything” setting for bright +non-dark-mode contexts, with a security/usability review of its scope. +** DONE [#C] Fix install errors surfaced by the 2026-05-11 VM test run +CLOSED: [2026-08-08 Sat] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-06 +:END: +Closed at the 2026-08-08 task review: every archsetup-attributable error was +fixed and verified (fontconfig, dconf x2, emacs-stow, AUR exit-0 logging at +the root); the residual four reproduce unchanged and are diagnosed +environment/non-critical, with two 2026-06-28 full runs attributing zero +issues to archsetup. Residual thread: confirm the firewall nf_tables pair on +bare metal at the next real install — no container task needed to carry it. +*** 2026-06-28 Sun @ 13:29:29 -0400 Audit reconcile: 2026-06-28 btrfs+zfs runs reproduce the same residual set +Newer full runs landed since the 2026-06-11 reconcile below: the 2026-06-25 zfs run (Testinfra 96/0) and the 2026-06-28 btrfs+zfs runs (97/0, "zero attributed issues"). The residual four were NOT fixed and reproduce unchanged: =enabling firewall= (archsetup:1496-1498, carries a VM-kernel note), =enabling gamemode for user= (archsetup:2221, non-critical), and =tidaler (AUR)=. Zero archsetup-attributed Testinfra issues across both profiles confirms these are environment / non-critical, not archsetup bugs. Bare-metal confirmation of the firewall pair is still the open thread. + +*** 2026-06-15 Mon @ 23:53:21 -0500 Audit reconcile: latest VM run (2026-06-11) confirms the surviving error set +The most recent VM run (=test-results/20260611-113904/=) carries four error-summary entries: =enabling firewall= + =verifying firewall is active= (the iptables/nf_tables "Could not fetch rule set generation id" pair, still unconfirmed on bare metal), =enabling gamemode for user= (non-critical), and =tidaler (AUR)=. The earlier fontconfig/dconf fixes held — none reappear. So the count is down from the 7→6 anchor below to four, all of them the known-residual items already itemized. +Errors logged during the VM install. Status as of the 2026-05-11 18:36 run (=test-results/20260511-183643/archsetup-output.log=) after the =48c9439= fontconfig/dconf fix: 7 → 6. +- refreshing font cache — RESOLVED in =48c9439= (now installs =fontconfig= before calling =fc-cache=). +- configuring GTK file chooser — RESOLVED in =ecab29f= (switched to a system-wide dconf db at =/etc/dconf/db/site.d/=; needs no session bus during install). +- configuring GNOME interface settings in dconf — RESOLVED in =ecab29f= (same fix as the GTK file chooser above). +- enabling firewall — exit 1: =iptables v1.8.13 (nf_tables): Could not fetch rule set generation id: Invalid argument=. Still present in the 18:36 run; likely a VM-kernel/nf_tables artifact — confirm on bare metal before treating as an archsetup bug. +- verifying firewall is active — exit 1 (follow-on from the firewall-enable error). +- enabling gamemode for user — exit 1 → step "gaming" FAILED — non-critical. +- tidaler (AUR) — logged in the error summary with exit code 0 (odd; logging quirk or transient AUR build noise?). +Also seen in the 18:36 run's log-diff (post-install systemd noise, probably VM-environment): =pam_systemd … CreateSession failed= / =logind: Failed to start session scope … Permission denied=, and =Failed to start Proton VPN Daemon= (no VPN config in the test VM). + +*** 2026-05-19 Tue @ 13:18:56 -0500 Fixed AUR exit-0 logging bug at the root +Root cause was in =retry_install=: =last_exit_code=$?= ran AFTER =if eval ...; then return 0; fi=. Bash defines an if-compound's exit status as zero when no condition tested true, so a failing eval's exit code got overwritten with 0 before reaching =error_warn=. Fix in =8221c54=: capture =$?= from =eval= directly into a local var, then compare against the captured value in the if. VM-verified in =test-results/20260519-115318/=: =mkinitcpio-firmware (AUR)= and =tidaler (AUR)= now report =error code: 1= (yay's actual exit) instead of the misleading =error code: 0=. The same packages still appear in the summary because yay returns non-zero when sub-deps fail to build (e.g. =aic94xx-firmware=), but the codes are accurate now. If the underlying sub-dep failures stay noisy, that's a separate concern — open a new task. + +*** 2026-05-16 Sat @ 09:00:41 -0500 AI Response: Surfaced the expanded AUR-exit-0 pattern +2026-05-16 07:40 VM run passed (52/0/5) with the same warning profile as the 2026-05-11 18:36 run. Error count went 7 → 13: 5 fixed/unchanged, +5 new AUR-exit-0 entries (broadens the existing tidaler item into the dedicated =[#B]= subtask above), +1 genuinely new error in =setting up emacs configuration files= (=git pull= ran in =~/.emacs.d= which existed from stow but had no =.git=). Patched =archsetup:1932-1945= with a three-branch check: clone if missing/empty, pull if =.git= exists, =git init=/=fetch=/=checkout= in place if the dir came from stow. + +*** 2026-05-19 Tue @ 01:25:26 -0500 Verified the b9907c7 emacs-stow fix end-to-end +=make test= 21:44 → 22:29 (42 min), =test-results/20260518-214516/=. 52/0/5, =ArchSetup Exit Code: 0=. The third-branch path fired correctly — install log =archsetup-2026-05-18-21-45-46.log:14358-14365= shows =From https://git.cjennings.net/dotemacs= → =[new branch] main -> origin/main= → =Reset branch 'main'= → =branch 'main' set up to track 'origin/main'=. No exit-128, no =fatal: not a git repository=. Error Summary down to 7 (was 13 on 2026-05-16); the emacs entry is gone. AUR exit-0 logging triggered for 2 packages this run (mkinitcpio-firmware, tidaler) vs 6 on 2026-05-16 — same bug class, fewer triggers, still tracked under =[#B] AUR exit-0 logged as error=. Issue Attribution: 1 ARCHSETUP entry (Proton VPN Daemon failed — known VM-no-VPN-config artifact). Cleanup ran clean via the normal path. +** CANCELLED [#C] Review current tool pain points annually +CLOSED: [2026-08-08 Sat] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-06 +:END: +Killed at the 2026-08-08 task review: an undated annual intention that never +fired — pain points get surfaced organically as they bite. +Once-yearly systematic inventory of known deficiencies and friction points in current toolset +** DONE [#B] Podman API socket and camera-passthrough udev rule :feature:solo: +CLOSED: [2026-08-09 Sun] +:PROPERTIES: +:CREATED: [2026-08-07 Fri] +:LAST_REVIEWED: 2026-08-07 +:END: +Shipped 2026-08-09: the installer enables the rootless podman socket at +install time (enable_user_service grew a wants-target arg so socket units +land in sockets.target.wants) and ships +=72-usb-passthrough-cameras.rules= — numbered below 73 per the winvm +rule-ordering correction, GROUP/MODE as the verified grant, uaccess tag kept. +Applied live on ratio (socket enabled+active, 99- file retired, udev +reloaded); velox apply rides the velox-return riders on the sleep/suspend +task. The uaccess-alone hypothesis stays untested until a camera is attached. +From winvm 2026-08-07 (ratio). Two one-time machine-level setups, both live on +ratio and absent on velox; full evidence and rationale in +[[file:docs/design/2026-08-07-podman-socket-and-camera-udev.md]]. + +- Enable the rootless podman socket at install time + (=systemctl --user enable --now podman.socket=). Socket-activated, zero idle + cost; every podman GUI/API client needs it, and its absence fails silently + (Pods opens to an empty window). The installer already carries the + "=systemctl --user enable= fails during install" workaround pattern + (=archsetup:1270=, =:2722=) — use it. +- Ship a udev rule granting GROUP="video", MODE="0660" on the OBSBOT + (3564:ff02) and BRIO (046d:085e) USB nodes so =usbredirect= can claim them + for VM passthrough. CORRECTED (winvm, 2026-08-08): the original "uaccess + can't ACL raw USB nodes" claim was wrong — the mechanism is rule ordering. + The ACL is applied by =73-seat-late.rules=, so a =99-= rule adds the tag + after that already ran; distro rules that add the tag all sort at or below + 70. [@70] So number our file below 73 (e.g. =72-usb-passthrough-cameras.rules=), + keep the verified GROUP/MODE grant, and keep the tag — correctly ordered it + may make uaccess work on its own (untested hypothesis; a tighter grant if + it holds, needs the camera plugged in to verify). Reconcile ratio's + existing =99-= file (winvm installed it) when the installer version lands. + +Scope: installer step + rule file + tests per existing shapes, and apply both +live to velox over tailscale (daily-driver sync — neither exists there today). +** DONE [#B] Velox touchpad interrupt line is dead — needs a part or a BIOS fix :bug:velox:hardware: +CLOSED: [2026-08-15 Sat] +:PROPERTIES: +:CREATED: [2026-08-15 Sat] +:LAST_REVIEWED: 2026-08-15 +:END: +*Fixed 2026-08-15 23:05 by reseating the correct connector* — a seating fault +all along, no part needed. Verified at the kernel level on the 23:05 boot: the +=did not ack reset within 1000 ms= message is gone (clean handshake), and the +interrupt count went 0 → 1795. Power-key events also zero, so both faults from +the mainboard swap are closed. + +What made this take three attempts is worth keeping: two of the connectors on +that board were decoys. The input-cover ribbon looked like the obvious suspect +and fixing it *did* resolve the power button, which made it look like the whole +answer. Then the 4-pin connector next to the printed =TOUCHPAD= label looked +like the touchpad's own — and its cable is silkscreened =PIN 1-2 - GND / +PIN 3-4 - VCC=, four contacts of pure power, incapable of carrying i2c or an +interrupt. Reading that silkscreen off the photo is what ruled it out and sent +the search to the ribbon that actually crosses to the mainboard. + +The ordered touchpad becomes a spare, which is what Craig wanted from it anyway. +The diagnostic path below is left intact — it is the reusable part: =dmesg= +for the i2c-HID reset message and the interrupt count in =/proc/interrupts= +together separate "device absent" from "device present but its interrupt line is +open", and a live USB separates hardware from software in two minutes. +Split from the ribbon-reseat task 2026-08-15 once the reseat fixed the power +button and left this untouched — they are two faults, not one. + +*Diagnosed to the interrupt line specifically, with software eliminated.* +- The i2c *data* path works. =i2c_hid_acpi= read the HID descriptor, returned + the right product ID (=093A:0274=), =hid-multitouch= bound, and input6/7/8/9 + were created. A descriptor read is a real bus transaction, so the device is + electrically present and answering. +- The *interrupt* path never fires. IRQ 81, =amd_gpio= hwirq 8, level-triggered, + =actions=PIXA3854:00= — the handler is correctly registered on the pin the + firmware names. Count is 0 across all 24 CPUs, including during active + swiping. +- =dmesg=: =i2c_hid_acpi i2c-PIXA3854:00: device did not ack reset within 1000 ms=. + The i2c-HID reset handshake is acknowledged *by the device asserting the + interrupt*, so the first operation needing that line already failed at boot, + before anything touched the pad. That is why the fault reproduces on any boot + in ten seconds. +- *Software ruled out by live USB.* Same "did not ack reset" message and no + pointer movement under Ubuntu's kernel (2026-08-15). Not a driver, not + libinput, not Hyprland, not this install. + +Three candidates remain, all needing a part or firmware: +1. Open conductor on the touchpad's own cable or a bad contact at either end. + Framework sells "Touchpad Cable" as a discrete spare, so it is separately + replaceable — and the input-cover ribbon reseat would not have touched it. +2. The touchpad module's interrupt output is dead while its i2c slave still + answers. Indistinguishable from 1 without swapping parts. +3. Firmware naming the wrong GPIO. The DSDT says =amd_gpio= pin 8; if this + board revision routes the interrupt elsewhere, the kernel watches a pin that + never toggles. Plausible because the mainboard is days old to this machine + and its firmware already needed the PSR workaround. BIOS is 03.05 + (2025-10-30); kernel 6.18.44-1-lts. + +*The connector that was reseated is NOT the touchpad's — confirmed from the +board photo.* Craig reseated the 4-pin connector near the printed word +=TOUCHPAD=. Its cable is silkscreened =PIN 1-2 - GND / PIN 3-4 - VCC= — four +contacts, all of them power. No clock, no data, no interrupt; almost certainly +the keyboard backlight feed. An i2c-HID touchpad cannot run through it, so that +reseat could never have fixed this, and *the free retry remains untried*. +Photo: [[file:working/velox-touchpad-interrupt/touchpad-module-underside-2026-08-15.jpg][working/velox-touchpad-interrupt/touchpad-module-underside-2026-08-15.jpg]]. + +Visible on that board: the controller IC marked =PCT3854= (matching the kernel's +=PIXA3854=), a larger =CON3= carrying a blue-backed ribbon with "26" marked +beside it, a white ZIF past the Framework QR label, and a further connector at +the board's end. The one that matters is whichever ribbon physically *leaves the +input cover and reaches the mainboard* — that is the touchpad cable, and its far +end is the press-fit connector at the board. Reseat both ends of that one before +fitting any new part. + +*BIOS 04.02 exists but does not look relevant.* Checked 2026-08-15 with velox +on AC at 90%: fwupd offers 0.0.3.5 → 0.0.4.2. Read the changelog — the only +touchpad line is haptic-touchpad support for the Laptop 13 *Pro* chassis, and +this machine has a conventional PixArt =PIXA3854=. The rest is BIOS Setup +layout, option naming, TPM behavior, iGPU defaults, PMF slider. Nothing about +GPIO routing or interrupt configuration. So candidate 3's cheap test is weaker +than it looked when it was filed sight-unseen; still worth doing (unlisted +fixes happen, and ACPI tables change), just no longer the front-runner. +Deliberately deferred past the flight — a cleared NVRAM is the failure that +started this whole rebuild. Boot-entry recovery reference captured at +[[file:docs/2026-08-15-velox-uefi-boot-entry-reference.org][docs/2026-08-15-velox-uefi-boot-entry-reference.org]]. + +Order of attack on return, cheapest first: reseat the touchpad's *own* press +connector at the mainboard (free, untried) → BIOS 04.02 → fit the replacement +touchpad. Craig's call 2026-08-15: order the parts now anyway, since they are +worth holding as spares regardless of which candidate wins. + +*What to order.* The replacement *Touchpad* ships with the Touchpad Cable +pre-installed, so that single part covers candidates 1 and 2 together — no need +to buy both to cover both. A bare Touchpad Cable is worth adding only as a cheap +spare. The *Input Cover* is a different and more expensive part, and nothing +points at it: the keyboard works, so the input-cover ribbon is carrying signal. +Framework's marketplace renders its catalogue in JavaScript, so prices could not +be read programmatically — search "Touchpad" under Laptop 13 parts. + +*Also worth a Framework support ticket* — the touchpad died coincident with +their mainboard swap, which may put it inside whatever recourse that carries. + +Grading: Major severity (a laptop's built-in pointer is entirely dead — the +counter-argument is that an external mouse is a complete workaround, which +would make it Minor; I took Major because losing the integrated pointer degrades +the machine's portability, which is the whole point of the laptop) x every user, +every time = P1 = [#A]. Filed [#B] rather than [#A] only because an [#A] must +carry a date and Craig's return date isn't known yet — date it and raise it to +[#A] when it is. + +Workaround in the meantime: Bluetooth mouse, already in use. +** DONE [#C] hypridle.conf is generated per-machine but tracked :refactor:dotfiles: +CLOSED: [2026-08-14 Fri] +:PROPERTIES: +:CREATED: [2026-08-14 Fri] +:LAST_REVIEWED: 2026-08-14 +:END: +Fixed in dotfiles 83ae7aa. The render is untracked and gitignored; the +store is the only source of truth; hypridle-start renders at session start +and owns the fallback ordering. Both machines reconciled and re-stowed: +velox's tree is clean for the first time today and still carries its own +policy (dim 5, lock 10, suspend-then-hibernate 30), ratio's config is +byte-unchanged. +=hyprland/.config/hypr/hypridle.conf= is rendered by the settings panel +from each machine's own stage config, and it is also a tracked file stowed +to every machine. So a machine whose idle policy differs from the +committed default carries permanent working-tree dirt, and every pull +there needs a stash/pop dance (velox, twice on 2026-08-14). Worse, the +committed copy is whichever machine last committed it, which is how a +desktop ended up tracking a laptop's suspend-then-hibernate line. +Options to weigh: gitignore the rendered file and track only a template or +the stage defaults; render to a non-stowed path and have hypridle read +that; or keep it tracked but commit a machine-neutral render. The first +looks right — the store already holds the real source of truth, and the +rendered file is a build artifact. +** DONE [#A] Reseat velox input-cover ribbon — phantom power button :bug:velox:hardware: +CLOSED: [2026-08-26 Wed] DEADLINE: <2026-08-26 Wed> +:PROPERTIES: +:CREATED: [2026-08-13 Thu] +:LAST_REVIEWED: 2026-08-13 +:END: +Machine off, lift the input cover (Framework QR-guided procedure, 5 +fasteners), reseat its ribbon connector to the mainboard — disturbed in the +2026-08-13 board swap. Root cause of every "mystery reboot" that day: +chassis flex (flash-drive touch, ethernet bump, lid partially lowered) +fired phantom power-button presses — journalctl -b -1 showed "Power key +pressed short." → orderly logind poweroff, then the glitching button +powered it back on. While in there, reseat the USB expansion cards too — +the flaky slot (two hard resets, one no-enumeration) is likely the same +flex problem. +THIRD SYMPTOM (2026-08-13 evening): touchpad delivers ZERO input events — +15s synchronized libinput debug-events capture while swiping caught +nothing, though i2c enumeration and a driver rebind handshake are clean. +Signature of a dead interrupt line on the same ribbon. Keyboard + power +LED lines work; BT mouse is the interim pointer. +ESCALATED 2026-08-13 21:00: a fourth event killed the machine THROUGH the +shield. Previous boot's journal ends mid-line (tailscaled chatter) with no +shutdown sequence at all — a hard power cut, not logind acting. So the +glitch now reaches the EC/hardware power path, which no software setting +can intercept. The reseat is the only fix, and this is a +lose-work-without-warning failure mode, not an inconvenience. +Interim shield (already live): /etc/systemd/logind.conf.d/powerkey.conf +sets HandlePowerKey=ignore — phantom presses log but do nothing; EC-level +10s hold still force-cuts. Consider keeping it even after the repair. +Verify after reseat: flex the chassis edges + partially lower the lid, then +grep the journal for new "Power key pressed" lines — zero means fixed. +Must be done before the Sunday flight — a phantom press mid-travel with the +shield on is survivable, but the connector should not be trusted at 30,000 +feet on the loose setting. + +*** 2026-08-19 Wed @ 14:40:00 -0700 Retracted: the RTC reset is not this task's, and I should not have filed it here +I attributed the 2026-08-19 network outage to this ribbon earlier today. Craig +pushed back — he reseated it before the trip to get the touchpad working — and +he is right. The evidence does not support the attribution and some of it points +the other way. + +What actually holds. Boot -3 ended at 01:33:18 with no shutdown sequence: no +power-off target, no unmounting. The next boot's kernel line reads =rtc_cmos +00:01: setting system clock to 2025-01-01T00:00:16 UTC=, a firmware default, so +the RTC was reset rather than drifted. No firmware update was applied +(=fwupdmgr get-history= is empty) and the battery is fine. + +What refutes the ribbon. This boot logged *zero* =Power key pressed= events, and +so did the four boots before it. The phantom-press symptom had genuinely stopped +after 08-15, exactly as the 08-16 session recorded. The earlier events logged a +power-key press and an orderly poweroff; this logged neither, which makes it a +different signature, not a worse version of the same one. + +What I got wrong methodologically: I anchored on the most salient open hardware +task and read association as evidence. I even wrote "I can't prove it is the +same connector" and then filed it here anyway, which is the tell. + +Two things I checked and can rule out. There were no OOM kills — the 3,433 +matching lines are a systemd unit named "Periodically re-score Claude Code +processes for the OOM-killer" firing on a timer, not memory pressure, and there +is not a single "Killed process" line. Thermal is clean; the only mentions are +boot-time zone registration at 34C and 45C. + +One real thing the same window did surface, tracked separately: a python3 crash +loop, 251 core dumps in the final ten minutes, SIGABRT with =XFreeThreads= and +=PyEval_RestoreThread= in the trace. It does not explain the RTC, because +software cannot clear it, but it is its own problem. + +The open question that would settle the RTC is for Craig, not the journal: a +long power-button hold on a Framework triggers an EC-level reset that clears the +RTC, which fits a wedged machine being forced off. A 4-second hold would not. + +*** 2026-08-17 Mon @ 19:57:42 -0700 Not done, and the two symptoms now disagree +The reseat did not happen before the flight, and velox is travelling. The +deadline blew past on 08-14. + +The two symptoms have separated, which is worth recording because it changes +what the evidence proves. The phantom presses have stopped: fifteen "Power key +pressed" entries between 08-14 04:29 and 08-15 20:04, then nothing at all +across five boots including today's. The touchpad has not — there is still no +touchpad node under =/dev/input/by-path/=, which is the same dead interrupt +line the body describes. + +So the quiet power button is not evidence the connector reseated itself. The +interrupt line is the symptom that cannot be masked in software, and it is +still dead, so the ribbon is still unseated. The most likely reason the +presses stopped is that the machine has been sitting on hotel surfaces instead +of being carried and flexed. + +The interim shield is still live (=HandlePowerKey=ignore=), and the escalation +note stands: an EC-level glitch cuts power below systemd regardless of it. +*** 2026-08-15 Sat @ 23:05:00 -0500 The reseat did happen, and the touchpad came back — this contradicts the 08-17 read +Recording this because a parallel session concluded on 08-17 that the reseat had +not happened and the touchpad was still dead. Both halves were done and verified +that night, so the two accounts disagree and the disagreement should be visible +rather than silently resolved by whichever session committed last. + +What was done: the input-cover ribbon was reseated first, which fixed the +phantom power button — the 22:09 boot logged zero =Power key pressed= lines +after Craig flexed the chassis, against nine on the boot before. The touchpad +did not change, because the input-cover ribbon is not its connector. The 4-pin +connector beside the printed =TOUCHPAD= label is silkscreened =PIN 1-2 GND / +PIN 3-4 VCC= — pure power, so it cannot carry i2c or an interrupt. Reseating the +ribbon that actually crosses to the mainboard fixed it. + +Measured, not assumed: the touchpad interrupt (=amd_gpio= pin 8) went from 0 +counts across all 24 CPUs to 1795, and =i2c_hid_acpi ... did not ack reset +within 1000 ms= disappeared from the boot log. Craig confirmed the pointer moved. + +*Why the 08-17 probe likely misread it:* it checked for a node under +=/dev/input/by-path/=. i2c-HID touchpads frequently get no =by-path= symlink +even when fully working, so its absence is not evidence of a dead interrupt +line. The falsifiable check is the interrupt count in =/proc/interrupts= while +the pad is being touched, or the reset message in =dmesg=. + +*Left open rather than closed* — velox was refusing ssh at merge time on 08-20, +so the current state could not be re-verified, and a later regression cannot be +ruled out. One second of Craig's time settles it: move the pointer. If it works, +close this; if it does not, the interrupt line went back down and that is new +information. + +*** 2026-08-26 Wed @ 22:30:46 -0600 Closed: the reseat was done on 08-15 and the task was never marked +I reseated the ribbon on 2026-08-15 and never closed this. The 08-15 entry +above already records the verification: zero =Power key pressed= lines on the +22:09 boot after flexing the chassis, the touchpad interrupt count back up +once the right connector was reseated. This boot shows zero presses as well. +The interim shield (=HandlePowerKey=ignore= in +=/etc/systemd/logind.conf.d/powerkey.conf=) is still live; I'm leaving it in +place, since a phantom press with it on costs nothing and without it costs +the session. +** CANCELLED [#B] agent-text relay reports success for a message that went nowhere :bug: +CLOSED: [2026-08-19 Wed] +:PROPERTIES: +:CREATED: [2026-08-19 Wed] +:LAST_REVIEWED: 2026-08-19 +:END: + +Not a defect. rulesets refuted it with measurements and I reproduced theirs +before accepting: on velox, whose account store is empty, +=signal-cli -a +15550000000 send= exits 1 with "User +15550000000 is not +registered", and =ssh 100.71.182.1 'exit 7'= returns 7, so a non-zero code +propagates faithfully back through the relay. The loop's +=[ "$rc" -eq 0 ] && break= therefore advances to the next host exactly as +intended. signal-cli fails closed. + +I filed this off a conditional in their handoff — ".emacs.d raised a case +neither of you tested ... *if* signal-cli send exits zero against an empty +account store" — and turned the "if" into a graded [#B] with a =:blocked:= tag +on another project, without running the one command that settles it. The +machine that proves it was in front of me the whole time. Their ask is fair and +I am recording it rather than the outcome alone: verify before filing a defect +against someone else's work, especially one carrying a blocking tag. +** DONE [#B] Clock/DNS bootstrap deadlock — recovery needs a second device :bug:velox: +CLOSED: [2026-08-19 Wed] +:PROPERTIES: +:CREATED: [2026-08-19 Wed] +:LAST_REVIEWED: 2026-08-19 +:END: + +The installer wrote both halves of a deadlock. =configure_dns= pins +=DNSOverTLS=yes= with =DNSSEC=yes=, and both validate against the wall clock; +the chrony step enables chronyd without writing a config, so the machine runs +Arch's stock one whose only source is =pool 2.arch.pool.ntp.org= — a hostname. +Boot with a wrong clock and DoT certificate validation fails, so nothing +resolves; chrony then cannot resolve its pool, so the clock stays wrong. +Neither side moves. It caught velox on the road 2026-08-19 and had to be +diagnosed from a phone. + +Fixed at the root: the installer now writes +=/etc/chrony.d/10-bootstrap-ip-ntp.conf= with two IP-addressed Cloudflare +sources and points stock chrony.conf at the drop-in. An address needs no DNS +and carries no certificate, so the escape hatch holds whatever broke the clock. +velox has the same drop-in applied live, verified with =chronyc -n sources= +(=162.159.200.1= selected) and =timedatectl= reporting synchronized. + +What is left here is the part I could not verify: the decisive test is a full +power-down and cold boot, confirming the clock corrects itself untouched. See +the manual-testing entry. Until that runs, the fix is sound by construction +rather than demonstrated. + +Grading: Critical severity (total loss of network — no DNS means no egress, and +recovery needs a second device) x some users sometimes (only machines that boot +with a wrong clock, which is any RTC fault, BIOS reset, or drained cell) = P2 = +[#B]. Graded on the being-in-it, not the getting-into-it: once the machine is in +this state it is fully offline with no local path out. + +*** 2026-08-19 Wed @ 12:25:00 -0700 Reproduced it, and the mechanism was not what either of us said +I wound velox's clock back 27 days with chronyd stopped and watched it fail. +Resolution died outright, and plain UDP/53 to 1.1.1.1 kept answering throughout +— the discriminator the doctor keys on, confirmed live rather than reasoned. + +The cause is DNSSEC, not DNS-over-TLS. resolved logged =signature-expired= +against the root DNSKEY and every DS beneath it. The DoT handshake to +=1.1.1.1:853= verified clean at that same clock, and the Cloudflare certificate +runs Dec 2025 to Dec 2026, so it was never outside its window. An RRSIG window +is days to weeks and a certificate is good for a year, so a skew that breaks +DNSSEC normally leaves DoT untouched. The phone session blamed the certificate +and I carried that forward into the first commit; both were wrong. + +=DNSSEC=allow-downgrade= does not rescue it either, which matters because it is +the obvious reach and it is what ratio runs. resolved downgrades when a server +lacks DNSSEC support, and a signature-window failure is a validation failure, so +no downgrade fires. Six retries over eighteen seconds plus +=resolvectl reset-server-features=, all dead. I briefly believed otherwise off a +test whose success was a cache hit (=Data from: cache network=). + +So ratio was exposed after all, and I have given it the same drop-in. Its +=162.159.200.1= is selected and its clock is synchronized. + +The fix itself is verified end to end: with the clock wound back and no DNS at +all, chronyd reached the IP-addressed source and stepped the clock from +2026-07-23 straight back to 2026-08-19. That is the whole claim, demonstrated +rather than argued. + +Also settled: the clock landed on 2026-07-23 because that is systemd 261.2's +build date to the minute (=/usr/lib/systemd/systemd=, 10:43:59), and systemd +advances a garbage RTC to its own build epoch at boot. Not timesyncd's +last-good-sync timestamp, which cannot be it — timesyncd is disabled here. That +also confirms the RTC really was reading earlier than that, so the coin cell +stays the prime suspect. + +*** 2026-08-19 Wed @ 10:12:00 -0700 Root fix, doctor verdict, and taxonomy entry landed +The installer carries the drop-in; =post-rebuild-check= grew a sixth check that +fails a machine whose every NTP source is a hostname; the net failure taxonomy +gained the mode in its DNS layer plus a cluster 5 triage line, and its existing +egress-layer clock entry now says outright that its remedy does not apply when +DoT or DNSSEC is on. + +The doctor half is in dotfiles: =classify.py= reached "DNS not resolving → net +repair dns-test" here, which cannot help, because every public resolver fails +the same clock-sensitive validation — so the doctor sent you round a loop. It +now emits a =clock-dns= row ahead of the generic DNS verdict. Detection is +deliberately DNS-free: a local =timedatectl= read for sync state, and a bypass +query addressed by IP over plain UDP/53 to tell "resolved is refusing to +validate" apart from "DNS is genuinely dead". +** DONE [#C] DNSSEC strictness on the travelling laptop :velox: +CLOSED: [2026-08-19 Wed] +:PROPERTIES: +:CREATED: [2026-08-19 Wed] +:LAST_REVIEWED: 2026-08-19 +:END: + +Craig chose =allow-downgrade= everywhere. Applied to velox, ratio, and the +installer, and ratio's =DNSOverTLS= tightened from =opportunistic= to =yes= in +the same pass, so all three now agree: encrypted DNS always, validation +best-effort. + +The reasoning that settled it: the deadlock is fixed by the IP-addressed NTP +source, and =allow-downgrade= was measured not to help with it at all. What +=allow-downgrade= does buy is the venue-resolver case the taxonomy documents, +where =yes= turns a resolver that mangles DNSSEC records into no answer at all. +That is a hotel and airport problem, so it is velox's problem, and the +encryption is the half worth being strict about. +** CANCELLED [#C] Branch network policy on laptop vs desktop in the installer :feature: +CLOSED: [2026-08-19 Wed] +:PROPERTIES: +:CREATED: [2026-08-19 Wed] +:LAST_REVIEWED: 2026-08-19 +:END: + +Cancelled because the decision above emptied it. All three motivating cases now +want the same value on every machine: =DNSSEC=allow-downgrade=, a stable +per-network wifi MAC, and an IP-addressed NTP source. A branch with nothing to +put on either side is machinery built for a divergence that does not exist, and +it would be the kind of scaffolding that rots unread. + +Worth keeping the observation, which is the part with a shelf life: when a +network default does need to differ by machine class, the test already exists. +=ls /sys/class/power_supply/BAT*= is what =prune_waybar_battery=, the ppd mask, +and the TLP config all key on. Reopen this then rather than building it now. +** DONE [#C] Automate the clock/DNS deadlock repair in the net doctor :feature: +CLOSED: [2026-08-19 Wed] +:PROPERTIES: +:CREATED: [2026-08-19 Wed] +:LAST_REVIEWED: 2026-08-19 +:END: + +Shipped as the =clock-ip-ntp= repair, and the verdict is =fixable= rather than +terminal. Both open questions got answered by driving a real deadlock instead of +reasoning about it: =chronyc add server= returns =200 OK= against a running +chronyd, and =makestep= needs a sample to land, so it took four calls and about +eight seconds rather than working on the first. The repair retries accordingly. + +Verified end to end on velox against a genuine deadlock (wrong clock, chronyd +running with only an unresolvable hostname source, DNS dead): the repair +corrected the clock in 6.1 seconds and DNS came back. + +The live run also caught a defect no unit test would have. The doctor reported +"Saved password for SpectrumSetup-3C was rejected" — because +=_recent_auth_failure= greps =journalctl --since -5min=, and a clock weeks off +windows onto a different incident's entries. It would have sent Craig to +re-enter a password that was never wrong. Fixed twice over: the journal half is +now skipped when the clock is untrustworthy, and the clock verdict is ordered +above the auth verdict, since everything below it reasons over timestamps that +only mean something once the clock is right. Airplane mode and hard rfkill stay +above, being physical states the clock has no bearing on. +** DONE [#D] net-scenarios harness times out under back-to-back suite runs :test:tooling: +CLOSED: [2026-08-19 Wed] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-24 +:END: +=tests/net-scenarios/test_run_net_scenarios.py= errored on all 5 tests twice during round 12, each time =subprocess.TimeoutExpired= after its 20s budget on =scripts/testing/run-net-scenarios.sh --target root@fake-vm=. Both occurrences were in =make test-unit= runs launched immediately after a previous full run. It then passed 6 runs in a row (3 on a pristine tree, 3 with the round-12 change), and standalone it finishes in 0.09s, so this is not a regression from any code change. + +The harness stubs =ssh=, =rsync= and =jq= onto =PATH=, so nothing should touch the network at all — which is what makes a 20s timeout suspicious rather than merely slow. Worth reproducing under load before deciding whether the fix is a larger timeout or a real hang in the script. Evidence logs from the round: =/tmp/tu.log= and =/tmp/tu2.log= (tmpfs, gone after reboot). + +Not graded on the bug matrix: it is test infrastructure, not the shipped codebase. + +*** 2026-08-08 Sat @ 05:05:00 -0500 Recurred under concurrent load, same signature +All 5 tests hit the 20s TimeoutExpired again during a =make test-unit= run +that overlapped two review subagents running their own suites on the box. +Standalone immediately after: 0.095s, all pass; the following quiet-machine +full run was clean. Confirms the load-sensitivity read — reproduce under +deliberate load before choosing between a bigger budget and a real hang. + + +*** 2026-08-19 Wed @ 14:50:00 -0700 Root-caused and fixed: inherited stdin, not load +Not load, and not the network. The harness stubs ssh as =cat >/dev/null=, which +drains stdin to EOF. With no explicit stdin the stub inherits whatever the test +runner had, so it returned instantly when stdin was redirected and blocked +forever when it was a terminal or a live pipe. All five tests then burned their +20-second budget. + +That is why it looked like a load effect: a run launched immediately after +another inherited a different stdin than a standalone invocation. A/B measured +today — =make test-unit </dev/null= exits 0, the same target with an open pipe +on stdin hangs on all five. The note above guessed at "a larger timeout or a +real hang in the script" and it was neither. + +Fixed by pinning =stdin=subprocess.DEVNULL= in =run_script=. Verified both ways: +the previously-failing open-pipe case and the redirected case both pass in +0.08s, and a full =make test-unit= under a live pipe is clean across 50 suites. +** DONE [#A] powerprofilesctl crashes on a loop since ppd was masked :bug:velox:dotfiles: +CLOSED: [2026-08-21 Fri] +:PROPERTIES: +:CREATED: [2026-08-17 Mon] +:LAST_REVIEWED: 2026-08-17 +:END: +Something polls power state every 10-30 seconds, and each poll runs +=powerprofilesctl get=, which SIGABRTs. 47 coredumps on velox on 2026-08-17 +alone, the earliest at 08:34, four in one minute while I was watching. + +Cause is the 2026-08-16 fix that masked =power-profiles-daemon= so TLP +survives on laptops. That fix is right and stays. What it did not account for +is the settings module's power backing +(=~/.dotfiles/settings/src/settings/power.py=), which shells out to +=powerprofilesctl=. Against a masked unit the D-Bus activation fails with +=NameHasNoOwner ... unit is masked=, and the caller aborts rather than +degrading. + +Run by hand the same command exits 0 and prints the error, so the abort is +context-dependent and the caller needs finding before the fix is written. +Ratio does not mask ppd, which is why this is velox-only and why it appeared +the day after the masking. + +Costs: journal spam, coredump disk churn, and repeated failed D-Bus +activations on a travelling laptop's battery. It is also the leading suspect +for the wedged user manager filed below. + +Fix shape: =power.py= should treat a masked or unavailable ppd as a +first-class "no profile control here" state rather than an error path, and +the poller should stop retrying a unit it has been told is masked. The +machine-level half is already correct. + +Grading: Major severity (a crash loop burning battery and filling the +journal, silently) x every user every time on any laptop with the TLP fix +applied = P1 = [#A]. + +*** 2026-08-17 Mon @ 19:57:42 -0700 The loop stopped at the reboot; the defect did not +velox rebooted at 16:04 and there have been zero coredumps since, against 47 +in the twelve hours before it. So the loop is not currently burning anything. + +That is not a fix, and the distinction matters for whoever picks this up. +=powerprofilesctl get= still fails exactly as recorded — =NameHasNoOwner ... +unit is masked= — so every precondition for the loop is intact and it returns +whenever the caller next polls. What the reboot cleared is the caller's state, +not the bug. + +Narrowed the search the body asks for: =power.py= is the *only* file in +dotfiles that shells out to =powerprofilesctl= (=SETTINGS_POWERPROFILESCTL=, +line 14), so the caller is inside the settings module rather than waybar or a +timer. Worth knowing that the coredumps are =powerprofilesctl= itself aborting +— it is a python script, which is why they log as =/usr/bin/python3.14= +SIGABRT rather than under its own name. + +Grade unchanged. The matrix inputs did not move: the severity is what happens +while the machine is in that state, and the frequency row is every laptop +carrying the TLP fix. A quiet interval since a reboot is not a frequency +change. + +Fixed in dotfiles =e89d9db=. The caller was =waybar.py=, using =panel.read_state()= (the full snapshot of every control) to read one boolean, four bar modules deep on a 2-second interval. Two fixes, each needed alone: =panel.read_control()= reads a single control's backing, and =power.masked()= checks the mask symlink before shelling out. Verified with a logging stub: full snapshot unmasked calls powerprofilesctl once, masked calls it zero, and a waybar poll calls it zero even unmasked. +** DONE [#A] The installer clones my two working repos shallow and read-only :bug:velox: +CLOSED: [2026-08-21 Fri] +:PROPERTIES: +:CREATED: [2026-08-17 Mon] +:LAST_REVIEWED: 2026-08-17 +:END: +=archsetup:1432= clones the user's archsetup repo and =archsetup:1445= clones +dotfiles, both with =--depth 1=. Those are not build directories. They are the +two repos I actively develop in, and on velox they came back from the +2026-08-13 rebuild with 7 commits of history each instead of 851. + +Found 2026-08-17, and found the worst way: I ran the credential-file history +check that the GitHub-release task asks for, and it reported all five files +absent from history with a clean exit. The real answer is that this clone +cannot see the history those files live in. A shallow clone does not error on +=git log -- <path>=, it answers "no commits" — so a security question came back +falsely clean, and nothing about the output said otherwise. + +Everything else it breaks is quieter: =git log=, =blame=, =bisect=, and any +archaeology past the boundary. The tree looks completely normal, which is why +this survived four days on the machine. + +The right shape is already in the codebase. =scripts/post-install.sh:42-51= +takes depth as a per-repo argument and defaults to a full clone, so wallpaper +gets =--depth 1= and org does not. The AUR build clones (=archsetup:855=, +=:1673=, =:1677=) are correctly shallow and stay that way. Only the two +user-repo sites change. + +*Second defect, same two lines, found 2026-08-17 while pushing:* the dotfiles +clone could not push at all. =archsetup:245= defaults =dotfiles_repo= to +=https://git.cjennings.net/dotfiles.git=, the public read-only endpoint, so +=git push= returned 403. Ratio uses =git@cjennings.net:dotfiles.git= and +archsetup's own clone uses the matching ssh form, so velox was the odd one out +purely because it was the machine rebuilt by the installer. Repointed velox's +remote and pushed. + +That half needs a decision rather than a fix, which is why this task is no +longer =:solo:=. The https default is *correct for a stranger* installing +archsetup, who has no ssh key on the server, and this repo is being prepared +for public release. It is wrong for my own machines, which need to push. The +override already exists (=DOTFILES_REPO=, documented in +=archsetup.conf.example=), so the question is only where my personal value +lives: a config the personal ISO bakes in, a post-install step, or a detection +that prefers ssh when a key is present. Craig's call. + +*Decided 2026-08-19: the ISO bakes the value, and a check nets the rest.* +=archsetup:240= has the identical default for =archsetup_repo=, so this was +always two repos rather than one. I ruled out detection — archsetup never +restores =~/.ssh=, so key-presence at clone time depends on ordering it +doesn't control, and "any key means ssh" would break a stranger who has an +unrelated one. I ruled out a bare post-install step for the reason this whole +class of bug exists: manual steps don't get run, which is why this sat four +days. So the personal ISO carries =ARCHSETUP_REPO= / =DOTFILES_REPO= in the +ssh form (noted on the secrets/ISO task), and =post-rebuild-check= check 8 +flags any working repo still on the read-only endpoint — covering curl|bash +and stock-ISO installs, which the ISO value cannot reach. + +Repair on a machine already built: =git fetch --unshallow= in each repo, and +=git remote set-url origin git@cjennings.net:<repo>.git= for dotfiles. + +Grading: Major severity (two working repos silently missing their history on +the machine I develop on, and it returns confidently wrong answers to history +questions rather than failing) x every user every time (every fresh install, +both daily drivers) = P1 = [#A]. + +Not :solo:. The depth half is (two lines plus tests in the existing +=tests/installer-steps/= shape, verifiable by asserting the clone command +carries no =--depth= for these two repos). The remote-URL half needs the +decision above, so the task as a whole waits on it. Split it in two if the +depth fix is wanted sooner. +*** 2026-08-19 Wed @ 23:05:00 -0700 Dropped --depth from both user-repo clones +=archsetup:1462= and =:1475= now clone full history; +=tests/installer-steps/test_clone_user_repos.py= covers it with 8 cases, and +one of them asserts the AUR build clones still carry =--depth 1= so the fix +can't be over-applied by a careless repo-wide sed. Both my repos on velox were +already unshallowed by hand last session, so this is prevention rather than +repair. +*** 2026-08-19 Wed @ 23:05:00 -0700 Settled the remote-URL half and netted it +See the decision recorded above. The ISO half is a note on the secrets/ISO +task; the net is =post-rebuild-check= check 8, which ships now. + +Both halves resolved. Depth: =a028aa5= drops =--depth 1= from both user-repo clones, with 8 tests including one asserting the AUR build clones stay shallow. Remote URL: decided 2026-08-19 (see above) — the personal ISO carries the ssh form, and =post-rebuild-check= check 8 (=87ff0b7=) flags any working repo still on the read-only endpoint, covering the install paths the ISO cannot reach. +** DONE [#D] Worldclock tooltip blanks on one bad timezone row :bug:dotfiles:waybar:quick:solo: +CLOSED: [2026-08-21 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-25 +:END: +Found by sentry (2026-07-25), verified by exercising. =hyprland/.local/bin/waybar-worldclock= builds each zone with =ZoneInfo(tz)= inside the loop (line ~99) with no guard, so a single malformed timezone row in =worldclock.conf= raises =ZoneInfoNotFoundError= and crashes the whole python pass. The tooltip then renders empty and *every* zone is lost, not just the bad row; the traceback only reaches stderr, where waybar never surfaces it. +Repro: a conf with =America/Chicago|Home=, =Not/AZone|Bad=, =Europe/London|London= renders =tooltip: ""= (Home and London gone too). +Grade: minor severity (one module's tooltip blanks, no data loss) x rare edge case (a malformed conf row) = P4 = [#D]. +Fix: wrap the per-row =ZoneInfo=/=datetime= in a try/except and =continue=, so a typo drops only that row and the valid zones still render. Solo + quick: the script already has an env-override test harness (=WAYBAR_TIME_EPOCH=, =WAYBAR_WORLDCLOCK_CONF=), so a red-first test is cheap. + +Fixed in dotfiles =8f692f5=. The per-row =ZoneInfo= is guarded, so a malformed row drops itself and the valid zones still render. Five cases, including a bad row first — the ordering that looks least like one typo and most like the module being broken. Caught the broad =except Exception= rather than =ZoneInfoNotFoundError=, because the row also parses floats and calls strftime and the contract wanted is "a bad row costs only itself". +** DONE [#C] obsbot-wb-guard polls forever on machines with no OBSBOT :bug:dotfiles:quick:solo: +CLOSED: [2026-08-21 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-08-16 +:END: +=obsbot-wb-guard.service= is =WantedBy=graphical-session.target= and lives in the shared =common/= stow tier, so it starts on every machine. Its main path is =while :; do check_once; sleep 2; done=, and =check_once= returns early when the camera node is absent. On a machine with no OBSBOT attached that is a process waking every two seconds forever to do nothing, which on a laptop is battery spend for zero benefit. No restart loop, though: the loop never exits, so =Restart=on-failure= never fires. + +Found 2026-08-16 on velox, after enabling it to match ratio and then having to disable it again by hand. A per-machine disable is the wrong shape, because it drifts velox from ratio permanently and a re-stow or a future audit will just put it back. + +Fix: give the unit =ConditionPathExists= on the camera node (=/dev/v4l/by-id/usb-Remo_Tech_Co.__Ltd._OBSBOT_PW106-video-index0=, the same default the script uses) so systemd skips it on any machine without the camera and starts it normally on ratio. Then re-enable it on velox, where it will simply be skipped. Note the limit: a camera plugged in later will not start it until the next login, which is the right trade against a permanent poll. + +Careful when disabling by hand in the meantime: =systemctl --user disable= on a *linked* unit deletes the unit symlink, and that symlink is stow-managed, so a bare disable silently removes a file from the dotfiles stow tree. Restore the link afterward or re-stow. + +Grade: minor severity (wasted wakeups and battery, no data loss, no failure) x every boot on any machine without the camera = P3 = [#C]. + +Solo: buildable here (archsetup owns dotfiles end-to-end), verifiable by the agent (assert the unit is skipped on velox and still active on ratio), and no design call left open. + +Fixed in dotfiles =566dd14=. =ConditionPathExists= on the camera node, so systemd skips the unit where the camera is absent. velox is now =enabled= like ratio and reports =ConditionResult=no=; the stow symlink is untouched. Found while doing it: ratio has a Logitech BRIO and no OBSBOT on USB at all, so the 2-second poll was pointless on the desktop too, not merely costing laptop battery. A test asserts the unit's condition path and the script's =OBSBOT_WB_DEVICE= default stay equal, since drift there is invisible in both directions. +** DONE [#C] Spine face tests decay against the wall clock :bug:test:dotfiles:solo: +CLOSED: [2026-08-21 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-08-02 +:END: +=settings/faces/timeline-face-spine.test.mjs= has thirteen =SP.spineRows(g, h)= calls that omit the third argument, so =ref= falls back to its =new Date()= default while the file's events fixture is pinned to =JUL= (2026-07-31 18:30 UTC). Any assertion that depends on how much room the day needs is then measured against today's clock, and rots as the fixture recedes. + +One of them, "spacing is uniform everywhere except the gap home opens", had already rotted: green on 07-31 because that was the fixture's own date, red by 08-02. Fixed in place on 2026-08-02 by pinning =JUL=; the remaining thirteen pass today by luck. The measurement, for whoever picks this up — with =ref=now= the even step is 85.21 and home's gaps are 129.10 / 65.40 (the lower one collapses below a plain gap); with =ref=JUL= the step is 78.54 and the gaps are 129.10 / 145.46. Only the lower gap moves, because =up= does not depend on events and =down= does. + +Six other calls in the same file already pass =JUL= explicitly, so the convention exists and this is a miss, not a gap in the design. Fix: pass =JUL= at every call whose assertion reads geometry. Leave the call around line 747 alone — it sweeps =new Date(t0)= deliberately. + +Grade: minor severity (dev-facing only; no product behavior is wrong, the face itself is fine) x some users, sometimes (each call rots independently, whenever the fixture drifts far enough) = P3 = [#C]. Not merely cosmetic though: a suite that goes red for no real reason is how a genuine regression gets waved through. + +Solo — mechanical, an existing convention to copy, and verifiable by running the suite plus re-running it under a faked clock to prove the determinism actually holds. + + +Fixed in dotfiles =c96a216=. All thirteen bare calls now pass =JUL=. The task's "line 747" was stale (the deliberate =t0= sweep is at 893 and already passed its own ref, so it was never at risk), and the continuation-form call closes its arguments on the next line, which is why a naive grep counts fourteen. Added a guard that reads the file and fails with the offending line numbers, and verified it bites by stripping =JUL= from one call and confirming it went red naming that line. +** DONE [#A] Velox still carries the install placeholder passwords :bug:security:velox: +CLOSED: [2026-08-23 Sun] SCHEDULED: <2026-08-20 Thu> +:PROPERTIES: +:CREATED: [2026-08-20 Thu] +:LAST_REVIEWED: 2026-08-20 +:END: +Closed 2026-08-23: I'd already rotated all three on the 08-14 bringup day, so +this task was never live. Verified on velox before closing — =chage -l= puts the +last password change for both =cjennings= and =root= at Aug 14 2026, and +=/etc/zfs/zroot.key= was rewritten 2026-08-14 05:29 and no longer holds the +placeholder (checked with a =grep -qx= that returns a yes/no without reading the +key into a transcript). + +The premise below was wrong, and it's worth naming how. Nothing ever tested the +credentials: the claim came from an unticked runbook item plus the archangel +session handing back the values the *installer* had set, which reads as "these +are current" only if you assume nobody changed them in between. An inference +about a security exposure got recorded in the same voice as a measurement. The +one command that settles it costs a second. + +Original body follows. + +The 2026-08-13 reinstall set placeholder credentials and the runbook's Phase 5 +item to replace them (=passwd=, =zfs change-key zroot=) was never ticked. +Believed still live 2026-08-20 via the archangel handoff, which had to hand +them back to Craig to get into the machine: =welcome1= for the pool, =welcome= +for the accounts. + +So velox's full-disk encryption is currently protected by a dictionary word +with a digit, on the machine that travels. Anyone who picks it up owns the pool +and every account on it — the encryption is doing no work at all. + +Two commands, both on velox: +- =passwd= for each account. +- =zfs change-key zroot= for the pool passphrase. Note this is the ZBM unlock + passphrase, so get it right before rebooting. + +Grading: *severity-alone carve-out* — this is a security exposure, so the +frequency row does not discount it (=todo-format.md=). Critical severity: total +compromise of an encrypted-at-rest laptop from a guessable string, with the +device leaving the house. = P1 = [#A]. + +Distinct from the =VERIFY [#A] Rotate the credentials exposed by the 2026-08-09 +dotfiles leak= under the cgit audit — that one covers credentials a crawler +already took from a public repo. This one is a local default never changed. Both +are rotation work; neither substitutes for the other. +** CANCELLED [#B] Consistent keybinding family for the panel console :feature:hyprland: +CLOSED: [2026-08-21 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-09 +:END: +Merged into =[#B] Reconcile panel keybindings around Super+N=, which now carries +this body's detail: the collision list, the velox plain-keyboard constraint, and +maintenance-M as the priority chord. Cancelled rather than done — the work is +still open, just tracked in one place instead of two. + +Consider putting every panel (net, bluetooth, audio, timer, and the coming maintenance console) on one consistent chord family — a shared modifier set (Super+Shift, Control+Alt, or similar) plus a mnemonic letter per panel (N/B/A/T/M). Today the panels open via waybar clicks only; a uniform chord family makes them keyboard-reachable and predictable. Watch for collisions with existing binds: Super+Shift+A is already PTT toggle, and the hold-to-talk grave bind is load-bearing. Decide the family, audit current hyprland binds for conflicts, wire via the dotfiles hyprland config, and document in the keybind reference. Both machines (velox can't QMK-remap, so chords must work on a plain laptop keyboard). +*** 2026-07-14 Tue @ 00:31:36 -0500 Folded Craig's ask for a maintenance-panel keybinding; bumped [#C] → [#B] +Craig asked (in session, 2026-07-14) for a maintenance keybinding specifically — the panel he's reaching for without one. Maintenance (M) is the priority chord when this task gets worked. The capture graduated the task from parking lot to active backlog. +** DONE [#B] Waybar network module — custom/net :feature:waybar:network: +CLOSED: [2026-08-21 Fri] +:PROPERTIES: +:LAST_REVIEWED: 2026-07-09 +:END: +Closed 2026-08-21: the module shipped and is in daily use. Phases 1-4 all landed +in dotfiles, and the tunnels track absorbed most of what Phase 5 originally +covered. The one piece genuinely left — the =net vpn= CLI subcommand — is now its +own task below, so the residual is tracked at its real size instead of holding a +finished umbrella open. +Unifies the old wifi-no-internet indicator (was =[#C]=) and the network-manager +dropdown (was =[#B]=) into one =custom/net= module: a tested Python =net= engine +(nmcli + diagnostics), a thin bar indicator, and a GTK4 layer-shell panel. Code +lives in the dotfiles repo (hyprland tier + a =net/= package like pocketbook); +archsetup only installs deps. Secrets stay in NetworkManager's own store (no +separate credential store). The =captive= script becomes the diagnostics engine. +Full design, acceptance criteria, and the failure-mode coverage table: +[[file:docs/design/2026-06-29-waybar-network-module-spec.org][2026-06-29-waybar-network-module-spec.org]]. + +Phases below, dependency order. Engine/unit work is agent-verifiable (=unittest= ++ fakes on PATH, coverage via venv); the live-network and visual states need real +conditions, filed under "Manual testing and validation". + +*** 2026-06-29 Mon @ 20:19:11 -0400 Phase 1 shipped — indicator + console recovery +Shipped to the dotfiles repo (10 commits, =5254bd8=..=c095a22=, pushed to main). +The =net= engine is a src-layout Python package in-tree, imported by a bin shim +that resolves the stow symlink back to the repo — so it runs from a bare TTY with +no install, which the recovery path depends on. + +Landed: =net status= (fast path, one nmcli call + sysfs, degraded fallback in +budget) + =net probe= (native captive probe, single-flight flock, atomic cache, +fresh/stale/expired/unknown classes, iface/SSID/UUID invalidation); =waybar-net= +replacing =custom/netspeed=, throughput → tooltip, CSS states in both themes + +live; =net diagnose= (read-only steps) + =net repair= (rfkill/reset/bounce/ +dns-test, cleanup-verified) + =net doctor [--fix]= with the four terminal +classifications; =net portal= + the =captive --probe-json= refactor; redacted +JSONL event log; Makefile recovery targets (=make online= etc.); =~/.config/net/ +config=. Verified live: =make net-status= reads the real wlp170s0 / @Hyatt_WiFi. + +Airplane (Craig's call, option 1): =custom/net= absorbs only the *display* — net +reads the airplane-mode state file and shows an airplane state/glyph. The +airplane-mode toggle stays (it's a low-power mode — radios + CPU + brightness + +services — not a radio switch), now on =custom/net='s right-click + signal 15. +Deleted: =waybar-airplane=, =waybar-netspeed=, =custom/airplane=, their tests + +css. =airplane-mode= kept. + +Tests: 160 in =tests/net/= (fake nmcli/curl/rfkill/resolvectl/ping/getent/ +systemctl on a temp PATH; doctor-classification fixtures; degraded-under-slow- +nmcli benchmark) + the =captive= probe-mode tests; full dotfiles suite green (32 +suites). Coverage-gap pass via throwaway venv: pure modules ≥90% branch +(classify 100%), IO-error branches excused in the test docstring. +Deferred to Phase 2/3: archsetup deps (gtk4-layer-shell/python-gobject Phase 2, +speedtest-go-bin Phase 3 — not added before the code that needs them). +Verify (manual, live): see Manual testing and validation. + +*** 2026-06-29 Mon @ 22:19:25 -0400 Phase 2 shipped — panel shell + connection management +Shipped to dotfiles (commits =4e7740f=..=24bcac5=, pushed). Engine: =net list= (saved +MRU + in-range wifi scan, infrastructure types filtered), =net up/down= (UUID-keyed, +mutation safety — keep prior link until target activates, classify wrong-password vs +generic, report auto-reactivation), =net add/edit/remove/rescan= (open + WPA-PSK; +enterprise activate-only; secret to NM's store, never our JSON/log — tested). + +Panel: a GTK-free PanelModel (selection, four state machines, the UX-flow enable +rules, terminal states) + a GTK4 gtk4-layer-shell window (=net panel=) anchored +top-right under the bar — Connections section with MRU list, active marked, signal +glyph, row-click select, Connect/Add/Forget/Rescan, confirm-on-forget, worker-thread +engine calls via GLib.idle_add. GTK imported lazily so the CLI/tests stay GTK-free. + +Bar interactions (settled with Craig over live iteration): left = =net-panel= toggle, +middle = =net portal=, right = =net-fix= (notify the doctor result when one-way; open +a terminal only when the outcome is fixable — the sudo/interactive case). Airplane on +Super+Shift+A. archsetup adds =gtk4-layer-shell= + =python-gobject= (this commit); +already on velox. + +Tests: 204 in tests/net (merge ordering/dedup, up/down mutation safety, no-secret-leak +on add/edit, panel model + state machines, gui row-format helpers). Full dotfiles suite +green (32 suites). Live-verified on velox: panel opens/toggles, list shows real 24 +profiles, right-click notification delivers (Craig confirmed). Phase 3 (diagnose/repair/ +speedtest IN the panel) is next; the engine for it already exists from Phase 1. + +*** 2026-06-29 Mon @ 22:43:40 -0400 Phase 3 shipped — diagnostics + speed test in the panel +Shipped to dotfiles (=91277cf=..=691abcb=) + archsetup (=48052d6=, speedtest-go-bin), +pushed. Engine: =net speedtest= (parses speedtest-go --json → ping from latency ns, +down/up from per-server byte rates; missing-backend / offline / malformed → error +envelope per the failure table). Panel grew a section switcher with four pages: +- Connections (Phase 2). +- Diagnose: =net diagnose= on a worker thread, each step a row (✓/✗/… glyph + title + + redacted evidence), read-only; Open-portal button when captive. +- Repair: "Get me online" (=net doctor --fix=) + tiers (rfkill/reset/bounce/dns-test) + + force portal. Confirmations in-panel with the spec's exact wording; the privileged + tiers run via =net-popup= terminal (where the sudo prompt + step output, incl. + cleanup-verified, show) — a panel has no tty, and pkexec would mean a prompt per op. +- Speed test: in-process =net speedtest= (no privilege → inline result: ↓/↑ Mbps + ping + + server), Run/Cancel (Cancel pkills the child), error envelope shown. + +213 net tests; pure helpers (step_indicator, format_speedtest) unit-tested. Full +dotfiles suite green (32 suites). One unverified assumption: speedtest-go's dl/ul unit +(taken as bytes/s; =BYTES_PER_SEC= flips it) — needs one real run vs a reference. The +in-panel repair streaming (vs terminal) is a named future polish once the GUI-privilege +story settles. + +The waybar network module ([#B] parent) is now COMPLETE through Phase 3. Phase 4 +(in-app help + user guide) and Phase 5 (VPN/WireGuard) remain as future work; the core +feature (indicator + recovery + panel + diagnostics + speed test) is done. +Verify (manual, live): see Manual testing and validation. + +*** 2026-07-09 Thu @ 16:32:54 -0500 Audit reconcile: Phase 4 is filed on the dotfiles side, waiting on them +The dotfiles project accepted the Phase 4 handoff and filed it as a =[#C]= task in their own =todo.org= (their note, 2026-07-08 16:56): the help-text audit + panel help affordance, the user-guide/README, and the ratio rollout doc. Not started there. They ping when it lands, and this task's Phase 4 child closes then. Nothing to do here meanwhile. + +*** 2026-08-17 Mon @ 19:57:42 -0700 Landed on the dotfiles side; the block is cleared +dotfiles shipped it as =138da7b= and closed its own task, so this one closes +with it and the =:blocked:= tag comes off. Found by checking their =todo.org= +rather than waiting for the ping — their close-out note says "archsetup pinged +so its Phase 4 task can close", so the handoff worked and only this end was +left open. + +All three acceptance criteria are met on their side: the help audit found and +fixed a stale =net repair= action list (nine of nineteen actions were named; +both the CLI help and =repair.py='s docstring now generate from the ACTIONS +registry), =net/README.md= covers every command plus the recovery targets, and +the ratio rollout is documented with both daily drivers verified current. + +They split the panel help affordance out rather than inventing it — no sibling +panel has one, so its shape is a design call. It is tracked on their side, not +here. + +Original deliverable, for the record: in-app help (=net --help= + per-command, +panel help affordance); README/user-guide; archsetup Hyprland dep install +(=gtk4-layer-shell=, =python-gobject=, =speedtest-go-bin=); ratio manual dep + +stow step. Handed off 2026-07-04 with the archsetup deps already confirmed +installed. + +*** 2026-08-21 Fri @ 14:18:03 -0700 Promoted the Phase 5 residual out to its own task +Rescoped 2026-07-04 (audit): the tunnels track already shipped most of the original Phase 5. Panel tunnel bring-up/down and detection landed (dotfiles 2d9d060 probes tailscale/NM-wireguard/Proton; 21db05a brings overlays up/down from the panel's Tunnels sub-view; 31ba056 diagnose/doctor understand tunnel routes; archsetup 2e40781 wireguard config import; the net-panel-other-interfaces spec is IMPLEMENTED). What remains for Phase 5 is only the =net vpn ...= CLI subcommand — cli.py still has no vpn/tunnel parser. Fold the panel's existing tunnel operations into a CLI surface; spec separately when picked up. +** DONE [#A] Ratio: pull .emacs.d before upgrading Emacs to 31.1 :chore:ratio:emacs: +CLOSED: [2026-08-25 Tue] +:PROPERTIES: +:CREATED: [2026-08-25 Tue] +:LAST_REVIEWED: 2026-08-25 +:END: +Emacs 31.1's warnings.el defers daemon-startup warnings into a closure holding +the =*Warnings*= buffer; the config's dashboard-only sweep killed that buffer, +so the first client frame of every fresh 31.1 daemon failed on Wayland and +emacsclient silently fell back to =$DISPLAY= (XWayland, pgtk warning dialog). +Fixed in =.emacs.d= commit =63831060= (2026-08-25, velox verified live: +=GdkWaylandDisplay=). Ratio is still on 30.2, which lacks the deferring code, +so it is fine until it upgrades — then it hits the same trap once per daemon +start unless the fix is pulled first. + +Order on ratio: =git -C ~/.emacs.d pull= (the push from velox is the telega +session's; confirm =63831060= is on origin first), then the =pacman -Syu= that +brings =emacs-wayland 31.1=, then restart the daemon. Check afterwards: +=emacsclient -e '(pgtk-backend-display-class)'= → =GdkWaylandDisplay=. + +*** 2026-08-25 18:10 — pull already landed; the upgrade half remains +Checked ratio over tailscale: =~/.emacs.d= is clean at =91fbac72= (= =origin/main=), +and =63831060= is an ancestor of HEAD — =modules/undead-buffers.el= carries the +=*Warnings*= entry. Ratio is on =emacs-wayland 30.2-3= with =31.1-1= pending among +720 updates (last full upgrade 2026-08-01; kernel 7.1.5 → 7.1.9 also pending, +btrfs root, uptime 3.5 weeks). The daemon is a plain =emacs --daemon= (not a +user unit) holding 2 live frames, so the restart step will drop those frames. +What remains: the =pacman -Syu= on ratio, the daemon restart, and the +=(pgtk-backend-display-class)= check. + +*** 2026-08-25 Tue @ 18:35:00 -0600 Upgraded ratio to Emacs 31.1 and verified the Wayland backend +Ran the upgrade over tailscale as a transient unit (=ratio-upgrade.service=, +log at =/var/log/ratio-upgrade.log=): 714 packages, =--ignore= on the six +packages the live-update guard would have blocked (aquamarine, hyprland, +hyprutils, mesa, vulkan-radeon, wayland — still pending, apply from a TTY +before the reboot). One orphan cleared first: =qemu-block-gluster= had been +dropped from the repo and pinned =qemu-common=; the new =qemu-full= no +longer needs it. Killed the plain =emacs --daemon= (no modified buffers, no +graphical frames), started =emacs.service= instead so the daemon carries the +systemd user environment, and probed from a throwaway frame: +=(pgtk-backend-display-class)= → =GdkWaylandDisplay=, =*Warnings*= alive. +Ratio still wants a reboot for =linux 7.1.9=. Pacnews to review there: +=/etc/ssh/sshd_config.pacnew= and two =/etc/tpm2-tss/fapi-profiles/*.json=. @@ -1200,6 +1200,36 @@ configure_build_environment() { echo 'OPTIONS=""' > /etc/sysconfig/chronyd systemctl enable chronyd.service >> "$logfile" 2>&1 || error_warn "$action" "$?" + # Bootstrap NTP sources addressed by IP, never by hostname. + # + # Arch's stock chrony.conf names its pool by hostname, and the DNS this + # installer configures later runs DNSSEC=yes, which validates signature + # windows against the wall clock, so a machine that boots with a wrong + # clock resolves nothing: chrony cannot reach the pool, so the clock stays + # wrong, so DNS stays dead. Neither side moves, and recovery needs a second + # device to look up an NTP address by hand. An IP-addressed source needs no + # DNS and no certificate, so it breaks the deadlock unattended. I would + # rather carry two extra server lines than lose a laptop's network to any + # RTC fault. See the clock/DNS deadlock entry in the net failure taxonomy + # under docs/design/. + action="adding IP-addressed NTP bootstrap sources" && display "task" "$action" + mkdir -p /etc/chrony.d + cat << 'EOF' > /etc/chrony.d/10-bootstrap-ip-ntp.conf +# Reachable without DNS, so a wrong clock can always correct itself. +server 162.159.200.1 iburst +server 162.159.200.123 iburst +EOF + # Stock chrony.conf reads no drop-in directory, so point it at one. + if [ -f /etc/chrony.conf ]; then + backup_system_file /etc/chrony.conf + if ! grep -qE '^[[:space:]]*confdir[[:space:]]+/etc/chrony\.d' /etc/chrony.conf; then + printf '\n# Read drop-ins (archsetup owns /etc/chrony.d).\nconfdir /etc/chrony.d\n' \ + >> /etc/chrony.conf || error_warn "$action" "$?" + fi + else + error_warn "$action (no /etc/chrony.conf to point at /etc/chrony.d)" 1 + fi + action="configuring compiler to use all processor cores" && display "task" "$action" backup_system_file /etc/makepkg.conf sed -i "s/-j2/-j$(nproc)/;s/^#MAKEFLAGS/MAKEFLAGS/" /etc/makepkg.conf >> "$logfile" 2>&1 @@ -1428,8 +1458,13 @@ clone_user_repos() { # Without this, symlinks could point to /root or a tmpfs that disappears. user_archsetup_dir="/home/$username/code/archsetup" action="cloning archsetup to user's home directory" && display "task" "$action" + # Full history, deliberately. This is a working repo, not a build tree, and + # a shallow clone degrades silently: `git log -- <path>` answers "no + # commits" past the graft point rather than failing, so history questions + # come back confidently wrong. The AUR clones stay shallow; they're + # discarded after the build. (mkdir -p "$(dirname "$user_archsetup_dir")" && \ - git clone --depth 1 "$archsetup_repo" "$user_archsetup_dir" && \ + git clone "$archsetup_repo" "$user_archsetup_dir" && \ chown -R "$username": "/home/$username/code") \ >> "$logfile" 2>&1 || error_warn "$action" "$?" @@ -1442,7 +1477,8 @@ clone_user_repos() { # leaves /home/$username root-owned — so a clone running as the user fails with # "Permission denied" creating ~/.dotfiles. Cloning as root sidesteps that, and # chown -R gives the user the working tree. Mirrors the archsetup clone above. - (git clone --depth 1 --branch "$dotfiles_branch" "$dotfiles_repo" "$dotfiles_dir" \ + # Full history for the same reason as the archsetup clone above. + (git clone --branch "$dotfiles_branch" "$dotfiles_repo" "$dotfiles_dir" \ && chown -R "$username": "$dotfiles_dir") >> "$logfile" 2>&1 || error_warn "$action" "$?" # Q5: the --adopt/restore conflict handling below needs a real git checkout. @@ -1773,8 +1809,12 @@ configure_networking() { wifi.scan-rand-mac-address=yes [connection-mac-randomization] -# Random MAC for each WiFi connection (prevents tracking) -wifi.cloned-mac-address=random +# Stable per-network MAC, not a fresh one per connection. Both hide the real +# address from a venue; random also hides this machine from itself, so every +# reconnect at a hotel or airport looks like a new device and the portal login +# starts over. stable derives a different address per network, so the privacy +# across networks is unchanged and the reconnect friction goes away. +wifi.cloned-mac-address=stable # Stable MAC for ethernet (avoids issues with MAC-based DHCP reservations) ethernet.cloned-mac-address=stable EOF @@ -1791,7 +1831,17 @@ EOF DNS=1.1.1.1#cloudflare-dns.com 9.9.9.9#dns.quad9.net FallbackDNS=1.0.0.1#cloudflare-dns.com 149.112.112.112#dns.quad9.net DNSOverTLS=yes -DNSSEC=yes +# allow-downgrade, not yes. Venue resolvers that mangle DNSSEC records are +# common on hotel and airport wifi, and yes turns that into no answer at all +# rather than an unauthenticated one. The encryption is the part worth being +# strict about, so DNSOverTLS stays yes. +# +# This is not what fixes the clock deadlock, despite being the obvious reach. +# Resolved downgrades when a server lacks DNSSEC support, and a clock-skew +# signature failure is a validation failure, so no downgrade fires. Measured on +# velox 2026-08-19: dead across six retries and a reset-server-features. The +# IP-addressed NTP source above is what breaks that deadlock. +DNSSEC=allow-downgrade # Disable mDNS in resolved - avahi handles .local resolution exclusively MulticastDNS=no EOF @@ -1802,6 +1852,8 @@ EOF dns=systemd-resolved EOF + configure_tunnel_dns_over_tls + # Note: If Docker containers have DNS issues, systemd-resolved's stub resolver # (127.0.0.53) may be the cause. Fix: configure Docker to use direct DNS, or # disable systemd-resolved and use /etc/resolv.conf directly. (2026-01-18) @@ -1811,6 +1863,43 @@ EOF run_task "linking resolv.conf to systemd-resolved" ln -sf /run/systemd/resolve/stub-resolv.conf /etc/resolv.conf } +configure_tunnel_dns_over_tls() { + # The resolved drop-in above pins DNSOverTLS=yes for every link. Proton + # VPN (proton0) and the static Proton WireGuard profiles (wgpvpn) push + # an in-tunnel resolver, 10.2.0.1, that answers plain port 53 and never + # completes TLS on 853, so with strict DoT every lookup through the + # tunnel hangs (2026-09-09). The Proton client deletes and recreates + # its NM profile on every connect, so a per-profile dns-over-tls + # setting can't stick and the net doctor's per-link repair is + # session-only there. A NetworkManager [connection-*] default matched + # on the interface names is what NM pushes to resolved on every + # activation, with no script and no race; wifi and everything else keep + # the strict setting. Verified live on ratio 2026-09-10 and velox + # 2026-09-12. + # + # $1 is the conf.d directory, defaulting to the system's so tests can + # run against a temp dir. + local confdir="${1:-/etc/NetworkManager/conf.d}" + local dropin="$confdir/tunnel-dns-over-tls.conf" + + action="turning DNS over TLS off for the Proton tunnel links" && display "task" "$action" + + mkdir -p "$confdir" 2>> "$logfile" || { error_warn "$action" "$?"; return 1; } + cat > "$dropin" << 'NMEOF' 2>> "$logfile" || { error_warn "$action" "$?"; return 1; } +# Proton VPN (proton0) and the static Proton WireGuard profiles (wgpvpn) push +# an in-tunnel resolver (10.2.0.1) that answers plain port 53 and never +# completes TLS on 853. The global resolved drop-in pins DNSOverTLS=yes, and +# the Proton client recreates its profile on every connect and sets no +# per-link mode, so every lookup through the tunnel fails. Default DNS over +# TLS off for those links only; wifi and everything else keep the strict +# setting. Diagnosed 2026-09-09/10. +[connection-tunnel-dot] +match-device=interface-name:proton0,interface-name:wgpvpn +connection.dns-over-tls=0 +NMEOF + chmod 644 "$dropin" 2>> "$logfile" || error_warn "$action" "$?" +} + configure_backlight_access() { # Screen backlight and keyboard-LED brightness, writable by the video # group. Arch ships brightnessctl with no udev rules: it relies on @@ -2014,8 +2103,18 @@ configure_service_discovery() { run_task "enabling avahi for mDNS discovery" systemctl enable avahi-daemon.service fi - pacman_install wsdd - run_task "enabling wsdd for Windows network discovery" systemctl enable wsdd.service + # wsdd.service used to be enabled here "for Windows network discovery". + # That is the host daemon: it advertises THIS machine as a Samba host to + # Windows clients, and nothing here runs Samba, so it advertised a + # share server that doesn't exist while listening on every interface + # including the VPN and tailscale links (2026-09-12). Browsing Windows + # shares is the other direction and is gvfs-wsdd's job; it spawns its + # own wsdd in discovery mode (see supplemental_software). A machine set + # up before this change still has the unit enabled, so a re-run turns it + # off; a fresh install has no unit yet and the guard skips quietly. + if systemctl is-enabled --quiet wsdd.service 2>/dev/null; then + run_task "disabling wsdd.service (no Samba host to advertise)" systemctl disable --now wsdd.service + fi pacman_install geoclue # geolocation service for location-aware apps run_task "enabling geoclue geolocation service" systemctl enable geoclue.service @@ -2472,6 +2571,9 @@ hyprland() { # enables TLP on battery machines, and the two daemons fight over # platform profiles. On TLP machines the panel's power control reads # as unavailable, which the settings engine handles. + # Not enabling it here is necessary but NOT sufficient on those machines: + # ppd is D-Bus activated, so configure_tlp_power masks it outright. Without + # that mask the panel starts ppd on demand and systemd kills TLP. pacman_install power-profiles-daemon if ! ls /sys/class/power_supply/BAT* &>/dev/null; then run_task "enabling power-profiles-daemon" systemctl enable power-profiles-daemon.service @@ -2965,6 +3067,7 @@ install_programming_languages() { # Shell pacman_install shellcheck # Shell script linter pacman_install shfmt # Shell script formatter + pacman_install bash-language-server # Bash language server; prog-shell.el warns at every start without it # Go pacman_install delve # Go programming language debugger @@ -3184,6 +3287,17 @@ EOF } ### Supplemental Software +mask_fwupd_passim() { + # fwupd depends on passim, a daemon that shares firmware metadata with + # other machines on the LAN by listening publicly on 0.0.0.0:27500. Any + # fwupdmgr run D-Bus-activates it, and the unit is static (no [Install] + # section), so `systemctl disable` is a no-op: it came back on velox the + # next time fwupdmgr ran (2026-09-12). Masking is what holds, and it is + # how ratio has carried it since 2026-07-21. fwupd itself is unaffected; + # it just stops offering metadata to the LAN. + run_task "masking passim (fwupd's LAN metadata sharing daemon)" systemctl mask passim.service +} + supplemental_software() { display "title" "Supplemental Software" @@ -3206,6 +3320,7 @@ supplemental_software() { pacman_install fdupes # identify binary duplicates pacman_install filezilla # ftp gui pacman_install gimp # image editor + pacman_install git-lfs # large-file storage; repos tracking LFS globs fail checkout without it pacman_install gparted # disk partition utility pacman_install gst-plugin-pipewire # gstreamer audio plugin for pipewire pacman_install gst-plugins-base # gstreamer base audio plugins @@ -3217,9 +3332,11 @@ supplemental_software() { pacman_install gucharmap # gui display of character maps pacman_install gzip # compression tool pacman_install handbrake # video transcoder + pacman_install imv # wayland image viewer (gui-open --image execs it) pacman_install libconfig # library for processing structured config files pacman_install libmad # mpeg audio decoder pacman_install libmpeg2 # library for decoding mpeg video streams + pacman_install libreoffice-fresh # office suite; the mimeapps.list defaults resolve to its .desktop files pacman_install maim # screenshot utility pacman_install mosh # alt SSH terminal with roaming and responsiveness support pacman_install odt2txt # converts from open document to text @@ -3243,6 +3360,9 @@ supplemental_software() { else aur_install slack-desktop # team messaging fi + # Claude desktop app (repackaged official .deb). Hard-depends on + # qemu-system-x86, edk2-ovmf and virtiofsd for its Cowork VM. + aur_install claude-desktop # zoom retired 2026-07-02: zoom-web (dotfiles) opens meetings in the browser pacman_install iperf3 # network bandwidth testing pacman_install bind # DNS utilities (dig, host, nslookup) @@ -3250,6 +3370,7 @@ supplemental_software() { pacman_install smartmontools # monitors hard drives pacman_install lm_sensors # temperature sensors (maintenance console) pacman_install fwupd # firmware update checks (maintenance console) + mask_fwupd_passim pacman_install lynis # security auditing tool pacman_install telegram-desktop # messenger application # LaTeX - minimal set for document compilation with latexmk @@ -3303,7 +3424,7 @@ supplemental_software() { aur_install nsxiv # image viewer aur_install snore-git # sleep with feedback pacman_install gvfs-smb # SMB network share browsing in Nautilus - pacman_install wsdd # WS-Discovery daemon (Windows network discovery) + pacman_install wsdd # WS-Discovery client, spawned by gvfs-wsdd; wsdd.service stays off (no Samba here) pacman_install gvfs-wsdd # WS-Discovery backend for gvfs (browse Windows shares) aur_install topgrade # upgrade everything utility aur_install ueberzug # allows for displaying images in terminals @@ -3550,6 +3671,46 @@ EOF run_task "enabling TLP service" systemctl enable tlp.service systemctl mask systemd-rfkill.service systemd-rfkill.socket >> "$logfile" 2>&1 || \ error_warn "masking systemd-rfkill for TLP" "$?" + # Masking systemd-rfkill leaves the resume edge with no owner. TLP's own + # sleep hook runs `tlp resume`, but DEVICES_TO_ENABLE_ON_STARTUP means + # startup and TLP has no ON_RESUME, so radio state is not restored after + # a sleep cycle. WiFi survives because NetworkManager unblocks itself; + # bluetooth stays soft-blocked, and after a hibernate its controller + # comes back wedged as well. This hook closes both, and it belongs here + # rather than beside the other installs because the mask above is what + # creates the gap it fills. + # Arch does not ship /etc/systemd/system-sleep, and install_executable + # is a plain cp, so without this the install warns and leaves no hook. + mkdir -p /etc/systemd/system-sleep >> "$logfile" 2>&1 || \ + error_warn "creating /etc/systemd/system-sleep" "$?" + install_executable "$user_archsetup_dir/scripts/zz-bluetooth-resume" \ + /etc/systemd/system-sleep/zz-bluetooth-resume + # power-profiles-daemon.service declares + # "Conflicts=tuned.service tlp.service auto-cpufreq.service ..." (note + # the direction: the line is in ppd's unit, NOT tlp's — grepping + # tlp.service for it finds nothing). So systemd TERMs TLP the moment ppd + # starts. Leaving ppd merely disabled does not prevent that: it ships + # D-Bus activation files, and the desktop-settings panel's own + # powerprofilesctl call activates it on demand. Velox ran that way from + # its 2026-08-13 rebuild until 2026-08-16 — TLP failed at every boot and + # none of its battery policy applied, while the machine looked correctly + # configured. Masking blocks D-Bus activation too, so TLP survives and + # the panel's power control reads as unavailable, which the settings + # engine handles. + systemctl mask power-profiles-daemon.service >> "$logfile" 2>&1 || \ + error_warn "masking power-profiles-daemon for TLP" "$?" + # Mask first, then stop: masking blocks any re-activation in the gap, and + # a mask alone leaves an already-running ppd running. This script runs on + # a booted system (a repair or re-run is normal), so without the stop TLP + # stays dead until the next reboot with nothing saying so. + # The order is load-bearing, not cosmetic, so don't "tidy" it: only + # hyprland() installs ppd, so on a battery machine running dwm or no + # desktop env the unit does not exist. Masking first creates the + # /dev/null fragment, so the unit loads as masked and the stop exits 0. + # Reversed, the stop would hit an unloaded unit, exit 5, and fire + # error_warn on every such install. + systemctl stop power-profiles-daemon.service >> "$logfile" 2>&1 || \ + error_warn "stopping power-profiles-daemon for TLP" "$?" fi } @@ -3813,10 +3974,12 @@ outro() { printf "\n" printf "If you use Proton Mail Bridge for cmail triage, finish the setup\n" printf "after reboot:\n" - printf " 1. Clone claude-templates to ~/projects/claude-templates if missing.\n" - printf " 2. Run 'protonmail-bridge --cli', log in, then quit.\n" - printf " 3. Run ~/code/archsetup/scripts/cmail-setup-finish.sh\n" - printf " 4. First mail sync: mbsync cmail && mu index\n" + printf " 1. Run 'protonmail-bridge --cli', log in, then quit.\n" + printf " 2. Run ~/code/archsetup/scripts/cmail-setup-finish.sh\n" + printf " 3. First mail sync: mbsync cmail && mu index\n" + printf "\n" + printf "Sending mail also needs cmail-action, which rulesets owns:\n" + printf "clone it to ~/code/rulesets and run 'make install'.\n" printf "\n" printf "Please reboot before working with your new workstation.\n\n" diff --git a/working/velox-reinstall/velox-reinstall-runbook.org b/docs/2026-08-13-velox-reinstall-runbook.org index 02671e2..2d99fc9 100644 --- a/working/velox-reinstall/velox-reinstall-runbook.org +++ b/docs/2026-08-13-velox-reinstall-runbook.org @@ -9,6 +9,29 @@ pool. Decision: full reinstall via archangel + archsetup, run deliberately as a disaster-recovery test of the ISO and scripts before the Sunday flight. Recent backup in hand; ratio available as the working machine. +* Outcome (recorded 2026-09-13 Sun) + +The drill ran on 2026-08-13 and 14 and velox came back as a working daily +driver: fresh install from the archangel ISO, keys and data restored from +the salvage backup, 23 repos re-cloned, rsyncshot reinstalled, hibernate +proven end to end. The checklist below was the live plan; it was not ticked +as the phases ran, so read it as the plan, not a log of each step. + +What the drill found, each filed as its own task rather than fixed in place: +velox's truenas backups had silently stopped on 2026-07-06 (found on 08-13 +before partitioning, which is what made the salvage pass required); four phantom +reboots were a ribbon disturbed by the board swap; a fresh install never +clones rulesets, never links the .emacs.d systemd user units, ships no +brightness udev rule, and loses gcalcli and the signal-cli registration. +Those live in archsetup's todo as the post-rebuild verification pass and its +siblings. The one commit that existed only on the old disk (emacs-wttrin +bf0457f) was rescued as a bundle and has its own task. + +Companion documents: the UEFI boot-entry recovery reference +([[file:2026-08-15-velox-uefi-boot-entry-reference.org][2026-08-15-velox-uefi-boot-entry-reference.org]]) +and the three gap reports under docs/design (2026-08-14-velox-reinstall-gaps-1 +to 3). + Fallback ordering if the test finds a real gap: - Before partitioning starts: the old system is intact — the ZBM repair route (efibootmgr entry pointing at the ZBM loader on the ESP, then @@ -22,8 +45,15 @@ Fallback ordering if the test finds a real gap: the microcode vendor-detection fix (archsetup, 2026-08-08) and was built without ARCHSETUP_DIR at all. #+begin_src sh - cd ~/code/archangel && sudo ARCHSETUP_DIR=~/code/archsetup ./build.sh + cd ~/code/archangel && sudo ARCHSETUP_DIR="$HOME/code/archsetup" ./build.sh #+end_src + Two traps in that one line, and either alone silently produces a bare ISO + with archsetup absent (archangel, 2026-08-20). The =VAR=value= form is + required because sudo's =env_reset= discards an exported variable. And + =$HOME= is required because zsh does not expand a tilde on the right-hand + side of an assignment — the earlier =ARCHSETUP_DIR=~/code/archsetup= here + passed the literal string. build.sh now warns and reports baked/not-baked + in its closing summary, so the failure is visible rather than silent. - [ ] build.sh fixes before the final rebuild (archangel repo): - rsync exclude for =.ai= (keeps =archsetup/.ai/private-design/= — the credential audit — off the portable USB stick). diff --git a/docs/2026-08-15-velox-uefi-boot-entry-reference.org b/docs/2026-08-15-velox-uefi-boot-entry-reference.org new file mode 100644 index 0000000..1eacfe5 --- /dev/null +++ b/docs/2026-08-15-velox-uefi-boot-entry-reference.org @@ -0,0 +1,108 @@ +#+TITLE: Velox UEFI Boot Entry — Recovery Reference +#+AUTHOR: Craig Jennings +#+DATE: 2026-08-15 + +Captured 2026-08-15 before a BIOS update (03.05 → 04.02) as insurance against +the update clearing NVRAM. Velox's mainboard swap on 2026-08-13 left exactly +this kind of empty NVRAM, which is what forced the reinstall — so a cleared +boot entry is the specific failure worth being able to undo in one command +rather than reconstruct. + +* Before recreating anything: check Secure Boot first + +The 04.02 update landed on 2026-09-12 (staged via fwupd, flashed on the next +reboot). It did NOT clear NVRAM: Boot0001 survived with its command line +intact. What it did was re-enable Secure Boot (Enforce Secure Boot = +Enabled), so the unsigned ZBM loader was rejected and the Framework BIOS +reported it as "Default Boot Device Missing / no bootable drive" rather than +a security violation. The symptom is indistinguishable from the NVRAM wipe +this document was written for. + +The tell: booting the Ventoy stick shows shim's MOK management screen. + +Fix: F2 → Security → Secure Boot → Enforce Secure Boot = Disabled, F10. The +machine then boots straight into ZBM. No efibootmgr needed. + +So when velox says no bootable device after a firmware update, check Secure +Boot before touching the boot entries. Only if Secure Boot is already off +and =efibootmgr -v= (from the stick) shows Boot0001 gone does the recreate +below apply. + +Before any future firmware update, record both =efibootmgr -v= and the +Secure Boot state so the post-reboot diagnosis is a comparison, not a guess. +The pre-firmware-update checklist in +[[file:workflows/system-health-check.org][docs/workflows/system-health-check.org]] +(Phase 3) carries the steps. + +Still open: velox's ESP has no removable-media fallback (=/efi/EFI/BOOT= +does not exist), so a real NVRAM wipe would still need the stick. Copying +=zfsbootmenu.efi= to =/efi/EFI/BOOT/BOOTX64.EFI= would let it boot unaided. +That belongs to archangel's ZBM install, where it is filed as [#C] +"Installed systems have no removable-media boot fallback on the ESP" +(2026-09-12) and ships with the next ISO rebuild after it lands. + +* State at capture + +- BIOS: 03.05 (2025-10-30) +- BootCurrent: 0001 +- BootOrder: 2001,0001,2002,2003 (USB ahead of ZBM — why the Ventoy stick + boots when it's inserted) +- Timeout: 0 seconds + +* The entry that matters + +=Boot0001* ZFSBootMenu= + +| field | value | +|----------------+----------------------------------------------| +| ESP part GUID | 8e51b680-f90a-444f-8da5-7e4f93625775 | +|----------------+----------------------------------------------| +| partition | 1 (GPT), start 0x800, size 0x100000 | +|----------------+----------------------------------------------| +| loader path | =\EFI\ZBM\zfsbootmenu.efi= | +|----------------+----------------------------------------------| +| cmdline (data) | =spl_hostid=0x22f8a7a1 zbm.timeout=3= | +| | =zbm.prefer=zroot zbm.import_policy=hostid= | +|----------------+----------------------------------------------| + +The =data= field is that command line in UTF-16LE, which is how efibootmgr +passes it as optional data. Recreate with =-u= and the plain string; efibootmgr +does the encoding. + +* Recreating it + +From a booted system (or the archangel ISO), with the ESP identified as +=/dev/nvme0n1p1= or whatever it enumerates as: + +#+begin_src bash +efibootmgr --create \ + --disk /dev/nvme0n1 --part 1 \ + --label "ZFSBootMenu" \ + --loader '\EFI\ZBM\zfsbootmenu.efi' \ + --unicode 'spl_hostid=0x22f8a7a1 zbm.timeout=3 zbm.prefer=zroot zbm.import_policy=hostid' +#+end_src + +Confirm the disk/part against =lsblk -o NAME,PARTUUID,PARTTYPENAME= first — +the partition GUID above is the authoritative identifier, not the device name, +which can enumerate differently. + +Then set the order so ZBM is reachable: + +#+begin_src bash +efibootmgr --bootorder 0001,2001,2002,2003 +#+end_src + +(The original order put USB first. Keep whichever you prefer; what matters is +that the ZBM entry exists and is in the list.) + +* Other entries (firmware-generated, recreate themselves) + +| Boot2001 | EFI USB Device | +|----------+----------------| +| Boot2002 | EFI DVD/CDROM | +|----------+----------------| +| Boot2003 | EFI Network | +|----------+----------------| + +These are stock firmware entries and come back on their own. Only Boot0001 +carries anything unique. diff --git a/docs/design/2026-07-10-net-bt-failure-taxonomy.org b/docs/design/2026-07-10-net-bt-failure-taxonomy.org index 74790c6..4d57b86 100644 --- a/docs/design/2026-07-10-net-bt-failure-taxonomy.org +++ b/docs/design/2026-07-10-net-bt-failure-taxonomy.org @@ -96,6 +96,7 @@ Six layers, mirroring the net doctor's probe ladder (link → IP/DHCP → gatewa - Another daemon overwrites resolv.conf (yes). DNS works then breaks (or breaks after VPN up/down) as dhcpcd/openvpn/openresolv rewrites resolv.conf. Multiple tools claim it with no coordination. Fix: pick one manager (openresolv =resolvconf=NO=, dhcpcd =nohook resolv.conf=), point resolv.conf at the stub, restart resolved. [[https://github.com/adrienverge/openfortivpn/issues/674][openfortivpn 674]] - nsswitch.conf hosts line broken (yes). All resolution fails, or LAN/mDNS names never resolve; the hosts line lacks =resolve=/=dns= in the right order or references an uninstalled nss module. Fix: set =hosts: mymachines resolve [!UNAVAIL=return] files myhostname dns=. [[https://man.archlinux.org/man/nss-resolve.8.en][nss-resolve]] - Avahi/.local mDNS not resolving (yes). *.local names don't resolve though unicast DNS works. nss-mdns not wired in, or resolved's built-in mDNS collides with avahi. Fix: install nss-mdns, add =mdns_minimal [NOTFOUND=return]= before =resolve=, enable avahi-daemon, disable resolved MulticastDNS if both run. [[https://wiki.archlinux.org/title/Avahi][archwiki avahi]] +- Clock skew breaks DNS itself, and NTP cannot recover it (yes; field-observed 2026-08-19, velox, not from the 2026-07-10 sweep). Nothing resolves at all — not a slow lookup, a dead one — after a boot with a wrong clock. =DNSSEC=yes= validates RRSIG inception/expiry windows against the wall clock, so a clock weeks off fails every query before it leaves the machine. Measured on velox 2026-08-19 with the clock wound back 27 days: resolved logged =signature-expired= against the root DNSKEY and every DS beneath it, and resolution died outright. =DNSOverTLS=yes= is *not* what bites, despite being the obvious suspect — the DoT handshake to =1.1.1.1:853= verified clean at that same clock, because a resolver certificate is good for about a year while an RRSIG window is days to weeks. A skew large enough to break DNSSEC normally leaves the certificate valid. The trap is the recovery path: NTP daemons name their servers by hostname (=pool 2.arch.pool.ntp.org=, =NTP=time.cloudflare.com=), so the daemon that would fix the clock needs the DNS that the clock is breaking. Neither side moves and the machine cannot self-heal — diagnosis needs a second device. Distinguish from the plain clock-skew entry in the egress layer by where it bites: that one has working DNS and failing HTTPS, this one has no DNS at all. Confirm with =dig @1.1.1.1 example.com +short=, which goes out plain UDP/53 and bypasses resolved entirely; an answer there with resolved still failing puts the fault in the validation layer, not the network. Fix: set the clock by hand (=timedatectl set-time=), then =resolvectl flush-caches=. =DNSSEC=allow-downgrade= does *not* help here, which is worth knowing because it is the obvious reach: resolved downgrades when a server lacks DNSSEC support, and a signature-window failure is a validation failure rather than a support failure, so no downgrade fires. Measured on velox: six retries over eighteen seconds, plus =resolvectl reset-server-features=, all dead. The only cure is correcting the clock, which is why the NTP source has to be reachable without DNS. Prevent by giving the NTP daemon at least one source addressed by IP, which needs neither DNS nor a certificate — =server 162.159.200.1 iburst= in a chrony drop-in. Note =timedatectl set-ntp true= is *not* a fix here: it starts a daemon that still cannot resolve its pool. ** Egress / captive portal / MTU / proxy / clock / upstream @@ -107,7 +108,7 @@ Six layers, mirroring the net doctor's probe ladder (link → IP/DHCP → gatewa - PPPoE / VPN link with a lower MTU not clamped (no). Browsing works but big transfers / some HTTPS hang. A PPPoE (1492) or VPN path has a smaller MTU and the too-large segments get dropped. Fix: set the tunnel/link MTU down (=.mtu 1420= for VPN, 1492 for PPPoE) or MSS-clamp on the gateway. [[https://thelineman.ca/articles/article-8-mtu-vpn-mss][vpn mtu/mss]] - Stale http_proxy env var points at a dead proxy (no). Every curl/wget/pacman fails though the network is fine; browsers may work. A leftover =http_proxy= points at an offline/off-network proxy. Fix: unset the vars, remove the export from =~/.profile= / =/etc/environment=. [[https://everything.curl.dev/usingcurl/proxies/env.html][curl proxy env]] - Unreachable PAC file off the corporate network hangs everything (no). Away from the office the browser stalls with no error. A system proxy set to "automatic" with a PAC URL that only resolves on the corporate LAN blocks waiting instead of falling back to DIRECT. Fix: switch system proxy to None (=gsettings … org.gnome.system.proxy mode 'none'=) or clear the PAC URL. [[https://bugzilla.mozilla.org/show_bug.cgi?id=1121800][ff pac hang]] -- Clock skew breaks every TLS handshake (yes). "Your connection is not private" on every HTTPS site though ping/DNS work; the clock is hours/years off. A dual-boot Windows RTC-localtime, unsynced NTP, or a dead CMOS battery leaves the clock wrong. Fix: =timedatectl set-ntp true= (=set-local-rtc 0= on dual-boot), replace the CMOS battery if it recurs. [[https://wiki.archlinux.org/title/System_time][archwiki system time]] +- Clock skew breaks every TLS handshake (yes). "Your connection is not private" on every HTTPS site though ping/DNS work; the clock is hours/years off. A dual-boot Windows RTC-localtime, unsynced NTP, or a dead CMOS battery leaves the clock wrong. Fix: =timedatectl set-ntp true= (=set-local-rtc 0= on dual-boot), replace the CMOS battery if it recurs. This entry assumes DNS still works; when the resolver runs DoT or DNSSEC the same skew kills DNS first and =set-ntp true= cannot recover it — see the clock/DNS deadlock in the DNS layer. [[https://wiki.archlinux.org/title/System_time][archwiki system time]] - Firewall default-deny drops all egress (yes). No traffic leaves right after enabling a firewall, or after both ufw and firewalld are on; even DNS fails. A default outgoing-deny policy, or two firewalls fighting over nftables. Fix: allow egress (=ufw default allow outgoing=) and run only one firewall. [[https://wiki.archlinux.org/title/Uncomplicated_Firewall][archwiki ufw]] - VPN kill-switch / leftover iptables rule strangles egress after VPN drops (yes; distinct from the route-capture case). Internet dies the moment the VPN disconnects and never returns until reboot. A kill-switch rule pinned traffic to tun0 and the leftover rule keeps dropping everything on the real interface. Fix: flush the stale rules (=iptables -F; iptables -P OUTPUT ACCEPT=, or restart the firewall), reconnect. [[https://bbs.archlinux.org/viewtopic.php?id=300104][arch ufw killswitch]] - IPv6 egress broken while IPv4 works (no; the egress angle of the broken-v6 family). Pages load slowly/intermittently; IPv4-only hosts are fine. The network advertises IPv6 with no working route and Happy Eyeballs keeps trying the dead AAAA path. Fix: =nmcli con modify <con> ipv6.method disabled= until the network's IPv6 is fixed. [[https://help.ubuntu.com/community/WebBrowsingSlowIPv6IPv4][ubuntu slow ipv6]] @@ -309,6 +310,7 @@ Probe: dns-config + resolver-health + dns-resolve + the doctor's dns-test (which - VPN split-DNS not applied :: AUTO — =resolvectl domain/default-route= on the VPN link. - IPv6 AAAA lookups stall :: AUTO — disable IPv6 on the link (or the single-request option). Also cluster 8. - Another daemon overwrites resolv.conf :: PRIV — pick one manager, point resolv.conf at the stub. +- Clock skew breaks DNSSEC validation, NTP deadlocked behind it :: PRIV — set the clock by hand, flush caches; prevent with an IP-addressed NTP source. The doctor must reach this verdict *before* any resolved restart, which cannot help and reads as a loop. - nsswitch.conf hosts line / avahi mDNS broken :: PRIV — fix the hosts line, install nss-mdns. ** Cluster 6 — names resolve, egress blocked diff --git a/docs/design/2026-07-15-velox-boot-failure-handoff.org b/docs/design/2026-07-15-velox-boot-failure-handoff.org new file mode 100644 index 0000000..5ec996f --- /dev/null +++ b/docs/design/2026-07-15-velox-boot-failure-handoff.org @@ -0,0 +1,61 @@ +#+TITLE: Velox boot failure — ZBM found no bootable kernel; diagnosis in progress, recovery plan attached +#+AUTHOR: Craig Jennings +#+DATE: 2026-07-15 + +* Why this is coming to archsetup + +Velox fails to boot: ZFSBootMenu reports it can't find a bootable environment with a kernel. Craig reports the last working velox session was an archsetup health-check run that included the pacman upgrade — so the breakage most likely happened inside archsetup's own workflow, and Craig wants the diagnosis + retrospective to continue here with full context. The .emacs.d session (where this was triaged, only because that's where Craig was sitting) hands off everything below. + +A phone photo of the zfs list output from velox's ZBM recovery shell accompanies this note in the inbox. + +* Timeline + +- 2026-07-13 ~23:50 CDT — velox last seen on the tailnet (per tailscale status read 2026-07-14 ~17:50). +- During that last session: archsetup health-check workflow ran, including a pacman upgrade (Craig's recollection — pacman.log will confirm exact times). +- 2026-07-14 late evening — Craig boots velox; ZBM: no bootable environment with a kernel. +- 2026-07-14/15 — triage from the ZBM recovery shell, Craig driving, guided from the .emacs.d session. + +* Facts established so far (from the ZBM recovery shell) + +- zroot imported, health ONLINE. Every dataset's keystatus is "available" — encryption unlocked, not a key problem. +- Layout confirmed from zfs list: zroot/ROOT/default (mountpoint /), separate datasets for home, home/root, media, var, var/cache, var/lib, var/lib/docker plus many docker layer children (legacy mountpoints). NOTE: no separate zroot/var/log dataset — /var/log lives inside zroot/var. That differs from the sanoid dataset list in archsetup's configure_zfs_snapshots (which configures zroot/var/log and zroot/var/lib/pacman as their own datasets) — worth reconciling in the retrospective. +- Mounted the BE read-only style: mkdir -p /mnt/be && mount -t zfs -o zfsutil zroot/ROOT/default /mnt/be. +- THE FINDING: /mnt/be/boot contains ONLY intel-ucode.img. vmlinuz-linux, initramfs-linux.img, and initramfs-linux-fallback.img are all gone. + +* Working hypothesis + +A kernel upgrade during the health-check run removed the old kernel files and never completed installing the new ones (interrupted transaction, mkinitcpio failure, or a /boot shadowing issue), and the machine was powered off with /boot empty. Arch's upgrade removes the running kernel's files at package-replace time, so a failure between "remove old" and "install new + mkinitcpio" leaves exactly this state: microcode present, kernel and initramfs absent. + +* Remaining diagnosis steps (not yet run — velox is sitting at the ZBM shell) + +1. Read pacman's log (on the zroot/var dataset): + #+begin_src sh + mkdir -p /mnt/var + mount -t zfs -o zfsutil zroot/var /mnt/var + tail -60 /mnt/var/log/pacman.log + #+end_src + Expect the failed/interrupted kernel transaction near the end; note its timestamp. +2. List recovery candidates: + #+begin_src sh + zfs list -t snapshot zroot/ROOT/default | tail -20 + #+end_src + Sanoid is configured for hourly=6/daily=7 on the ROOT dataset, so a pre-damage snapshot should exist. Check whether any pre-pacman_* snapshots appear — that tells us whether the 2026-06-29 pre-pacman hook design is actually installed on velox. + +* Recovery plan (agreed with Craig, pending the log read) + +1. Pick the newest zroot/ROOT/default snapshot that predates the failed transaction. +2. If the pool is imported read-only (zpool get readonly zroot): zpool export zroot && zpool import -f -N zroot. +3. zfs rollback -r zroot/ROOT/default@<snapshot> (the -r discards snapshots newer than the target; home/var/media are separate datasets and untouched). +4. zpool export zroot, reboot — ZBM should now see the kernel. +5. After first boot: re-run pacman -Syu attended, and confirm /boot holds vmlinuz-linux + initramfs-linux.img before any shutdown. + +* Retrospective candidates for archsetup + +- Does the health-check / upgrade flow verify /boot contents (kernel + initramfs present, mkinitcpio exit status) after a kernel upgrade? This failure would have been caught by a one-line post-upgrade assertion. +- Is the pre-pacman snapshot hook (2026-06-29 design, zroot/ROOT/default@pre-pacman_<ts>) installed on velox? The snapshot listing in step 2 above answers this empirically. +- The sanoid config vs actual dataset layout mismatch (var/log, var/lib/pacman) noted above. +- Whether the upgrade step should refuse to end the session (or page Craig) when a kernel transaction errors. + +* Related loose end already in your inbox + +A separate note (2026-07-14-1751) asks to add inetutils to the install base; velox also still needs that package installed once it boots again. diff --git a/docs/design/2026-08-14-velox-reinstall-gaps-1.org b/docs/design/2026-08-14-velox-reinstall-gaps-1.org new file mode 100644 index 0000000..cf0d723 --- /dev/null +++ b/docs/design/2026-08-14-velox-reinstall-gaps-1.org @@ -0,0 +1,137 @@ +#+TITLE: What the velox reinstall left behind — four gaps the install could close +#+AUTHOR: Craig Jennings + +* Heads-up: this was found from a .emacs.d session + +I opened a .emacs.d session on velox this morning, two days after the fresh +Arch install, and the first thing it did was fail: there was no =.ai/= +directory to read. Chasing that turned up four separate things the reinstall +did not restore. Three I repaired from the session; one needs me at my phone. + +None of this is a .emacs.d bug. They are all install-side gaps, which is why +they are landing in your inbox. Machine is velox; ratio was the reference for +every comparison below. + +* Gap 1 — the gitignored tooling layer does not survive a reinstall + +=~/.emacs.d= was re-cloned on 2026-08-13. Git brought back every tracked file +and none of the agent tooling, because =.gitignore= deliberately excludes it: +=.ai/=, =.claude/=, =CLAUDE.md=, =todo.org=, and =inbox/= were all simply +absent. That is the correct ignore policy — this repo relays to a public +mirror — but it means a reinstall silently drops the entire working state of +every gitignore-mode project. + +The damage on velox was total rather than partial: 374 files, 4.5 MB, +including =todo.org= (556 KB) and 184 archived session files. Nothing carries +it. Not git, not stow, not the bootstrap. + +I recovered it by rsyncing the set from ratio over the tailnet. Ratio was +authoritative and velox held nothing, so there was no merge to adjudicate — +which is luck, not design. Had velox held a few days of divergent state, this +would have been a hand reconciliation. It has been one before: 2026-07-31, when +the two machines' =.ai/= trees had forked to zero files in common. + +Worth knowing: this is fleet-general. Every project on the box that gitignores +its =.ai/= has the same hole, not just =.emacs.d=. + +What the install could do: after cloning a project, check whether a sibling +daily driver holds a =.ai/= for it, and offer to pull it across. Or at minimum, +list the projects whose tooling layer is missing so the gap is visible on day +one instead of at the first session that trips over it. + +* Gap 2 — stowed user timers come back linked but not enabled + +The unit files all arrived correctly through the dotfiles stow, symlinked into +=~/.config/systemd/user/= and resolving fine. But being present is not being +enabled, and the reinstall enabled only some of them: + +| unit | velox after reinstall | ratio | +|---------------------------+-----------------------+----------| +| calendar-sync.timer | enabled, active | enabled | +| agenda-render-cache.timer | enabled, active | enabled | +| roam-sync.timer | *linked, inactive* | enabled | +| signal-receive.timer | *linked, inactive* | enabled | +| emacs.service | linked, inactive | linked | + +=emacs.service= reads the same on both machines, so I take that one as +intentional and left it alone. The other two are real drift: =systemctl --user +enable= writes a =timers.target.wants= symlink into =~/.config/systemd/user/=, +and that symlink is not stow-managed, so nothing in the dotfiles repo carries +it. A stowed unit file is inert until something enables it. + +I enabled both with =systemctl --user enable --now=. Both fired immediately and +exited clean, and both now show a next elapse. + +What the install could do: enable the units it stows, explicitly, as a named +step. The inconsistency is the tell — two of four came back enabled, which +suggests something enables a subset and nothing enumerates the rest. + +* Gap 3 — the roam clone was stale, and held a diff that would have destroyed data + +This one has an ordering constraint, so it matters more than its size suggests. + +velox's =~/org/roam= was ten commits behind ratio, stuck at the 2026-08-04 +auto-sync while ratio was at 2026-08-14 — a direct consequence of gap 2, since +=roam-sync.timer= was never enabled here. + +The dangerous part: velox's clone also carried an *uncommitted* =inbox.org= +that had been emptied. Seventeen deletions, file down to zero bytes, holding a +pre-2026-08-04 state whose captures were long since processed on ratio. + +So the naive repair — enable =roam-sync.timer= and let it catch up — would have +committed that emptying and pushed it, deleting the four live inbox items on +ratio. The timer is the repo's only committer and it commits whatever it finds. + +I checked ratio's =inbox.org= first and confirmed it was a strict superset of +velox's HEAD version (same three items plus an 2026-08-09 capture), which made +the local change provably worthless. Then discarded it, fast-forwarded to +=a411b43=, and only then enabled the timer. Clone is clean and current, first +sync ran green. + +What the install could do: if it ever enables =roam-sync= on a rebuilt machine, +reconcile the clone *before* enabling, not after. An auto-committing timer +pointed at a stale dirty clone is a data-loss path, and the failure is silent +and remote — it lands on the *other* machine. + +* Gap 4 — signal-cli lost its registration, and that breaks the whole fleet + +=signal-receive.service= ran for the first time and reported: + +: signal-receive: +15045173983 not registered on this machine — nothing to do + +velox's signal-cli data dir holds a 39-byte empty =accounts.json=. Ratio still +has both numbers. So the reinstall wiped the registration, and per the design +notes velox was supposed to be the *primary* — ratio is the linked device. + +The effect is wider than velox, because of how =agent-text= dispatches: if the +local signal-cli holds the account it sends directly, otherwise it ssh-relays to +a hardcoded velox. Velox no longer holds it, so a send from here relays to +itself and fails; a send from any third machine relays to velox and fails the +same way. Only ratio still works, and only via the direct branch. The error text +blames "velox down or unreachable", which is misleading — velox is up and on the +tailnet, it just is not registered. + +This is the one I could not repair from the session: re-linking needs me at my +phone (Signal → Settings → Linked Devices, scanning the QR from =signal-cli +link -n velox=). Filed in .emacs.d's todo.org as [#B]. + +What the install could do: verify =signal-cli listAccounts= is non-empty after a +rebuild and say so loudly if it is not. Silent loss of the phone channel is +exactly the kind of thing nobody notices until the page that mattered never +arrives. + +* Summary of what I changed on velox + +- Restored =.ai/=, =.claude/=, =CLAUDE.md=, =todo.org=, =inbox/= to + =~/.emacs.d= by rsync from ratio. +- Discarded the stale local =inbox.org= diff in =~/org/roam= and fast-forwarded + the clone to current. +- Enabled and started =roam-sync.timer= and =signal-receive.timer=. + +Left alone, deliberately: =emacs.service= (matches ratio), and velox's Signal +registration (needs the phone). + +One unrelated thing I noticed while comparing the machines: ratio's signal-cli +warns its messages were last received twelve days ago, even though its +=signal-receive.timer= is enabled and active. That may be nothing, but the +receive cadence there is worth a look. diff --git a/docs/design/2026-08-14-velox-reinstall-gaps-2.org b/docs/design/2026-08-14-velox-reinstall-gaps-2.org new file mode 100644 index 0000000..95842ac --- /dev/null +++ b/docs/design/2026-08-14-velox-reinstall-gaps-2.org @@ -0,0 +1,63 @@ +#+TITLE: Fifth reinstall gap — machine-local .local.el config, and a general shape +#+AUTHOR: Craig Jennings + +* Follow-up to this morning's handoff + +Sent you four gaps an hour ago +([[file:2026-08-14-velox-reinstall-gaps-1.org][the first report]]). Here is a fifth, +found straight afterwards when I noticed calendar sync was dead on velox. + +* What was broken + +=calendar-sync.timer= was enabled and firing every fifteen minutes, and failing +every time with exit 255: + +: calendar-sync: No calendars configured (set calendar-sync-calendars) + +The three output files sat at zero bytes. The cause is that +=~/.emacs.d/calendar-sync.local.el= is gitignored, so the reinstall deleted it +along with everything else untracked, and the module's loader treats a missing +file as a *silent* no-op. So the config vanished quietly and the only symptom +was a failing unit nobody was watching. + +Cheap to fix once found: the repo tracks =calendar-sync.local.el.example=, and +that template already encodes the shape velox uses — feeds resolved by +=:secret-host= against =authinfo.gpg= rather than inlined. The authinfo entries +had survived, because =~/.authinfo.gpg= is a stow symlink into the dotfiles repo. +So rebuilding was one copy, and all three feeds now sync clean and land +byte-identical to ratio's. + +* The general shape, which is the part worth acting on + +This is the same failure as gap 1, one layer down, and it is worth stating +generally because the install can act on it: + +- A tracked =*.local.el.example= template plus a gitignored =*.local.el= is a + deliberate pattern in this config, not a one-off. =.gitignore= lines 56-58 + list three of them: =calendar-sync.local.el=, =signal-config.local.el=, + =google-keep.local.el=. Every one of those is gone on velox right now. I have + only repaired the calendar one. +- Secrets held *by reference* survive a rebuild; secrets held *inline* do not. + The calendar config came back for free because the tokens were in + =authinfo.gpg=, which is stow-managed and therefore travels. Ratio's copy of + the same file inlines its URLs, and had ratio been the machine rebuilt, those + three feed tokens would simply have been gone. +- The failure was silent by design. A missing local config is a no-op, which is + right for a machine that never configured the feature and wrong for one that + just lost it. + +* What the install could do + +- After a rebuild, enumerate every tracked =*.local.el.example= in a project and + report which have no corresponding =*.local.el=. That is a one-line find and it + turns a silent no-op into a visible checklist item. +- Same for any =*.local.*= convention elsewhere in the fleet — the pattern is not + specific to Emacs. +- Worth pairing with gap 2: a unit that is enabled and failing every fifteen + minutes for two days is its own signal. A post-rebuild pass over + =systemctl --user list-units --state=failed= would have caught this one + without knowing anything about calendars. + +That last one generalizes best. Of the five gaps I have sent you, three were +things that *looked* fine — a stowed unit file, an enabled timer, a present +clone — and were not. diff --git a/docs/design/2026-08-14-velox-reinstall-gaps-3.org b/docs/design/2026-08-14-velox-reinstall-gaps-3.org new file mode 100644 index 0000000..9675973 --- /dev/null +++ b/docs/design/2026-08-14-velox-reinstall-gaps-3.org @@ -0,0 +1,98 @@ +#+TITLE: Reinstall gaps, part three — per-install certs and credentials, and one failure that hid the others +#+AUTHOR: Craig Jennings + +* Third handoff today + +Two earlier notes covered five gaps +([[file:2026-08-14-velox-reinstall-gaps-1.org][the first report]] and +[[file:2026-08-14-velox-reinstall-gaps-2.org][the follow-up]]). +Email was the last thing broken on velox after the 2026-08-13 rebuild, and it +turned up two more — both the same shape, and one of them with a property worth +generalizing. + +Email is fully working now: three accounts, 21,853 messages, 4.0 GB indexed. + +* Gap 6 — the Proton Bridge TLS cert is per-install, and its absence disabled every account + +=~/.mbsyncrc= carries =CertificateFile /home/cjennings/.config/protonbridge.pem=. +That file did not exist after the rebuild, and it cannot be restored from backup +or copied from the other machine: Proton Bridge generates a fresh self-signed +cert per installation. Velox's is issued 2026-08-13 23:44 with a different +fingerprint from ratio's 2026-01-30 one. + +*The part worth acting on is the blast radius.* mbsync parses its entire config +before doing any work, so a missing =CertificateFile= referenced by *one* account +aborts the run for *all* of them. Gmail and dmail need no bridge and no cert, and +both were dead anyway. The error names only the missing pem, so the symptom +("no mail at all") and the message ("this one file is missing") look unrelated. + +Recovery does not need the bridge GUI. The running bridge presents the cert on +its own IMAP port, so it can be pulled straight off the handshake: + +: openssl s_client -connect 127.0.0.1:1143 -starttls imap -showcerts </dev/null \ +: | sed -n '/BEGIN CERTIFICATE/,/END CERTIFICATE/p' > ~/.config/protonbridge.pem + +That is a two-second, fully scriptable step, which makes it a good candidate for +the install rather than a runbook line. + +* Gap 7 — the bridge password is per-install too, and reports a stale value misleadingly + +=~/.mbsyncrc= resolves the cmail password with =cat ~/.config/.cmailpass=. That +file is plaintext and, unusually for my setup, a real file rather than a stow +symlink — so it is not in the dotfiles repo, not encrypted, and not carried to a +new machine. + +The file survived the rebuild but held the *previous* install's password, because +the bridge regenerates it per installation. Ratio's and velox's differ by sha256, +confirmed today. + +*The diagnostic trap:* Proton Bridge answers a wrong password with =no such +user=. I read that as "the bridge has no account signed in" and went looking for +a login problem. The account was configured the whole time. If the install ever +validates bridge connectivity, it should not treat =no such user= as evidence +about account state. + +* The generalization + +Gaps 6 and 7 are the same as 1 through 5, sharpened. Everything that broke in +this rebuild was *generated on the machine by an application* rather than carried +by git, stow, or the dotfiles repo: + +| gap | artifact | why it did not travel | +| 1 | =.ai/=, =todo.org=, =CLAUDE.md= | gitignored | +| 2 | =timers.target.wants= symlinks | written by systemctl enable | +| 3 | roam clone state | local working tree | +| 4 | signal-cli registration | per-device identity | +| 5 | =*.local.el= configs | gitignored | +| 6 | bridge TLS cert | per-install, regenerated | +| 7 | bridge password | per-install, regenerated | + +Gaps 6 and 7 add a distinction the earlier note missed. For 1, 3 and 5 the old +value is still correct, so *restoring* fixes them. For 4, 6 and 7 the old value is +*worthless* — the application has generated a new one, and only *re-deriving* +from the live system fixes them. An install that tries to restore these will +produce exactly what happened here: a file that exists, looks right, and +authenticates against nothing. + +So the install's post-rebuild checklist wants two columns, not one: what to +restore, and what to re-derive. + +* What the install could do + +- Re-derive the bridge cert from the running bridge with the =openssl s_client= + line above. Scriptable, no GUI, no secrets. +- Re-derive the bridge password from the bridge rather than expecting the file to + be right, and rewrite =.cmailpass=. (I have filed a task on my side to make + =PassCmd= ask the bridge directly, which would remove the file entirely.) +- Add a cheap post-rebuild validation that =mbsync --list= parses. Config-parse + failures disable every account at once and say nothing about mail, so they are + worth catching explicitly rather than via "no new mail" hours later. +- More generally: keep the restore list and the re-derive list separate, per the + table above. + +* Unrelated, but noticed while comparing the machines + +=~/.config/.gmailpass.gpg= and =~/.config/.dmailpass.gpg= resolve to mode 777 in +the dotfiles repo, on both machines. They are gpg-encrypted so the contents are +safe, but world-writable is wrong for a credential file. That is a dotfiles fix, +not an archsetup one — noting it here only because it surfaced in the same pass. diff --git a/docs/post-install-checklist.org b/docs/post-install-checklist.org index 97fc0d5..8c48938 100644 --- a/docs/post-install-checklist.org +++ b/docs/post-install-checklist.org @@ -18,6 +18,32 @@ bluetooth pairing landed below. * Checklist +** Run the post-rebuild check first + +Before working through the manual steps below, run: + +#+begin_src sh +~/code/archsetup/scripts/post-rebuild-check +#+end_src + +It runs the five checks a rebuilt machine actually needs — failed units, +user units that are present but never enabled, =*.example= configs whose +real sibling is missing, gitignore-mode projects missing the working state +their own =.gitignore= names, and the signal-cli registration. Each prints +a line whether or not it finds anything; exit 1 means something needs +attention. + +These are the gaps velox hit within two days of its 2026-08-13 reinstall, +and three of the five looked fine on casual inspection: a stowed unit file, +an enabled-looking timer, a present git clone. Run it again a day or two +after the install, once timers have had a chance to fail. + +It normally finishes in a second or two. On a machine whose user systemd is +wedged it takes a couple of minutes instead, because every =systemctl= call +is bounded at five seconds and check 2 makes one per unit. That is the slow +case working as intended: it reports what it could not read rather than +hanging. Set =PRC_SYSTEMCTL_TIMEOUT= lower to cut the wait. + ** Pair bluetooth peripherals Pairing is inherently interactive (scan, pick the device, confirm), so it @@ -71,7 +97,12 @@ needs doing. The installer's completion message carries the steps; recorded here too so the checklist is complete: -1. Clone claude-templates to =~/projects/claude-templates= if missing. -2. Run =protonmail-bridge --cli=, log in, then quit. -3. Run =~/code/archsetup/scripts/cmail-setup-finish.sh=. -4. First mail sync: =mbsync cmail && mu index=. +1. Run =protonmail-bridge --cli=, log in, then quit. +2. Run =~/code/archsetup/scripts/cmail-setup-finish.sh=. +3. First mail sync: =mbsync cmail && mu index=. + +Sending mail also needs =cmail-action= on PATH, which rulesets owns: clone it +to =~/code/rulesets= and run =make install=. That is not a prerequisite for the +steps above — the setup script warns and carries on — but =mbsync= is the first +thing that wants it. An agent session runs =make install= at startup, so on a +machine that runs them the link appears on its own. diff --git a/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org b/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org new file mode 100644 index 0000000..d9ec8d4 --- /dev/null +++ b/docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org @@ -0,0 +1,213 @@ +#+TITLE: Guarded-Upgrade Completion — keeping topgrade freshness honest +#+AUTHOR: Craig Jennings +#+DATE: 2026-08-25 +#+TODO: TODO | DONE +#+TODO: DRAFT READY DOING | IMPLEMENTED SUPERSEDED CANCELLED + +* DRAFT Guarded-upgrade completion +:PROPERTIES: +:ID: 81cdfd72-db96-43d3-aa03-779878c99f3e +:END: +- [2026-08-25 Tue @ 18:45 -0600] decisions closed 7/7. The kernel decision reversed on the velox DKMS failure chain: held on every everyday run, landed only in the dedicated session behind a DKMS/initramfs/snapshot gate. +- [2026-08-25 Tue @ 18:30 -0600] redirected: the everyday path is a live split upgrade (apply everything the guard would not block, defer the rest); the boot-time oneshot becomes the completion step for the deferred set. Decided while running exactly that by hand on ratio. +- [2026-08-25 Tue @ 06:39:42 -0600] drafted. Grounded in a live read of the maint engine, the pacman hooks, and the boot path on velox, not memory. The topgrade-freshness diagnosis that motivates it is in this session's log. + +* Metadata + +| Status | draft | +|----------+-------------------------------------------------------------| +| Owner | Craig Jennings | +|----------+-------------------------------------------------------------| +| Reviewer | Craig Jennings | +|----------+-------------------------------------------------------------| +| Related | maint =topgrade_age= metric; =hypr-live-update-guard= hook | + +* Summary + +The waybar maintenance module shows topgrade freshness as permanently stale. The cause is a real one: on a machine running Hyprland, a full =topgrade= almost never exits 0, because its system step upgrades GPU/compositor libraries that the =hypr-live-update-guard= pacman hook correctly refuses to swap under a live session. The freshness stamp is gated on topgrade's exit code, so a correct, protective refusal reads as "you never run updates." This spec designs a safe path to actually complete a guarded upgrade, and makes that completion record the freshness stamp, so the metric tracks the true state of the system. + +* Problem / Context + +The metric reads one cache key, =topgrade_run= (=~/.local/state/maint/topgrade_run.json=). Absent, the probe (=maint/src/maint/probes/updates.py:114=) returns WARN, "no topgrade run recorded". Two writers stamp it: the =topgrade= PATH wrapper (=~/.dotfiles/hyprland/.local/bin/topgrade=) on =rc -eq 0=, and the panel's TOPGRADE lever (=doctor.py=), which returns before the stamp on any non-zero exit. The read path is sound (a sandboxed =maint stamp topgrade= writes the file and =maint status= then reads freshness 0); the file is simply never written. + +It is never written because topgrade rarely exits 0 on this machine, and the reason is specific rather than flaky. =/etc/pacman.d/hooks/10-hypr-live-update-guard.hook= is a =PreTransaction=/=AbortOnFail= hook that, when Hyprland is running and an upgrade changes the on-disk version of a GPU/compositor library, prints a BLOCKED banner and exits 1 — aborting the whole transaction before any file is swapped. Its trigger set is =mesa=, =mesa-*=, =wayland=, =libdrm=, =libglvnd=, =hyprland=, =aquamarine=, =hyprutils=, =hyprgraphics=, =vulkan-radeon=, =vulkan-intel=, =vulkan-mesa-layers=, =nvidia-utils=, =lib32-nvidia-utils=, =xorg-xwayland=. The guard exists for a proven failure: replacing those libraries under a live compositor makes the next GPU call hit a now-deleted mapping and SIGABRT, taking every Wayland client down (hit on ratio 2026-06-07). + +So when any of those libraries has an update pending — a frequent event — topgrade's =system= step (it runs =yay=) aborts non-zero, topgrade returns non-zero, and neither writer stamps. The observed case: on 2026-08-24 topgrade ran at 17:49, hit the guard on =mesa= (26.1.7 → 26.2.1), and failed; the upgrade was then finished by hand with the guard's sentinel override, entirely outside the wrapper, so nothing stamped. The metric has read stale ever since. + +Two framings of the fix are in tension, and choosing between them is the spec's central decision. Either the metric means "how recently did you run the sweep" (recency), so the stamp should decouple from topgrade's exit; or it means "is the system up to date" (state), so staying stale while a guarded upgrade is deferred is *correct* and the only real defect is that safely completing that upgrade doesn't stamp. This spec takes the state framing (see Decisions). + +* Goals and Non-Goals + +** Goals +- A safe, low-friction way to apply a guarded (GPU/compositor-library) upgrade, with Hyprland not live at swap time. +- That completion records the =topgrade_run= freshness stamp, so the metric clears when the system is genuinely current. +- A boot-time upgrade path that can never lock the machine out of its session, however it fails. +- The installer owns the durable pieces so a rebuilt machine has them without hand-setup. + +** Non-Goals +- Weakening or bypassing the =hypr-live-update-guard= hook. It stays exactly as strict; this builds *around* it, not through it. +- Making the full topgrade ecosystem sweep (git repos, vim, npm, ...) run at boot. Those never need a stopped compositor and are out of the boot path. +- Changing how the kernel hazard is *guarded*. The hook stays silent on kernels; the split script holds them back on a live run as a second, separately-reasoned list (see Design), which is a deferral policy rather than a guard. +- A general offline-update system for all of pacman. Scope is the guarded-library case. + +** Scope tiers +- v1: the split-upgrade script (live: apply the non-blocked remainder, defer the rest, run the ecosystem sweep with the system step off, report the deferred set); maint's UPDATE/TOPGRADE levers route through it; an "apply on reboot" affordance that installs the held kernel live and arms the boot-time oneshot for the GPU/compositor set. +- Out of scope: full-sweep-at-boot; touching the guard's policy. +- vNext: none open — the kernel deferral that was vNext is now part of v1's held set. + +* Design + +The shape follows one principle: the only part of topgrade that needs a stopped compositor is its =system= step when a guarded library is pending. Everything else runs fine live and rarely fails. So the safe path is small and targeted — apply the guarded system upgrade with Hyprland down, once, and stamp it — while the ordinary full sweep stays a normal live =topgrade= run. + +Three pieces, at two altitudes — but the everyday gesture is not a reboot. It is a normal live update that simply leaves the dangerous few behind. + +*The split script.* A pacman =PreTransaction= hook can only abort or allow the transaction it is handed; it cannot drop targets from it. So "upgrade everything except the guarded set" cannot live in the hook — it lives one layer up, in a script the panel calls. On a live run the script: refreshes the sync db and reads the pending set (=checkupdates=); computes the *blocked set* = the guard's own trigger list (read from the installed hook's =Target= lines, so there is one source of truth, and version-aware the way the guard is — a same-version reinstall is not a swap) plus the *kernel set* (every installed kernel with its =-headers=, always as a set; held on every everyday run because a failed DKMS rebuild on velox's ZFS root leaves the machine unbootable — see the kernel decision); clears the news hook (=informant read=) where installed; runs =pacman -Syu --noconfirm --ignore=<blocked set>=; runs the AUR-only remainder (=yay -Sua --noconfirm=, AUR packages pinning a guarded version hold themselves back); then runs =topgrade --disable system,git_repos -y= so the other ecosystems still get their sweep and topgrade can actually exit 0. It writes the deferred set to a state file the panel reads, and exits 0 when the live part succeeded, whatever was deferred. The guard hook stays installed as the backstop for a bare =pacman -Syu= typed at a shell; on the driven path it never fires. Proof of concept: this exact sequence, run by hand on ratio on 2026-08-25 while Hyprland was live, resolved 724 of 730 pending packages (Emacs 31.1 among them) with the six guard hits deferred — after one unrelated fix, an orphaned =qemu-block-gluster= that had been dropped from the repo. + +*For the user.* UPDATE and TOPGRADE on the panel run the split script; they succeed, and the panel shows "N deferred" when the script held anything back. Landing the deferred set is a dedicated session, chosen on purpose, run in the foreground from the panel's action or =guarded-upgrade --complete= in a terminal: first the kernel set, live, with the desktop still up; then the gate — every DKMS module built for the new kernel, a fresh initramfs, and on a ZFS root a pre-pacman snapshot to fall back on. If the gate fails the script stops there, names what failed, and does not reboot; the machine keeps running on the old kernel and the desktop is available for the fix. If it passes, the script arms a persistent flag for the GPU/compositor set and offers to reboot (or, from a TTY with no compositor, applies that set directly). On the next boot, before the autologin shell starts Hyprland, the deferred guarded upgrade runs in the console — the guard passes freely because nothing is live — the stamp is written, the flag is cleared, and boot continues into the session. No second reboot: the libraries are already current before anything maps them. If anything goes wrong, the machine still boots into Hyprland and the panel still shows the pending upgrade, so you are never worse off than before arming. + +*For the implementer.* A persistent arm flag (a file on a non-tmpfs path, e.g. =/var/lib/archsetup/apply-upgrade-on-boot=, so it survives the reboot the =/run= guard sentinel cannot). A system oneshot, =archsetup-boot-upgrade.service=, =ConditionPathExists= on the flag, ordered =Before=getty@tty1.service= so it completes before autologin execs Hyprland — this ordering is mandatory, because a parallel run would let Hyprland start mid-swap and reintroduce the exact crash the guard prevents. The unit is bounded (=TimeoutStartSec=) and best-effort: its failure or timeout must not fail any target the session needs, so boot proceeds past it regardless. Its =ExecStart= runs, as the user: =informant read= (clear the news hook that would otherwise abort the transaction), then =topgrade --only system= (or the equivalent =yay -Syu=), then =maint stamp topgrade= on success, then removes the flag unconditionally (a one-shot arm — a failed attempt disarms rather than retrying every boot). =sudo= works unattended (=%cjennings NOPASSWD: ALL=), so no password prompt wedges it. + +The stamp also needs to happen when the upgrade is completed by other safe means — the by-hand sentinel-override path, or a =maint= command that does the same thing. The cleanest single home for the stamp is a small =maint apply-upgrade= (or a flag in the existing lever) that performs the guarded system upgrade and stamps on success, which both the boot unit and an interactive TTY run call. That keeps one code path that "completes a guarded upgrade and records it," rather than three writers that can drift. + +* Alternatives Considered + +** A. Decouple the stamp from topgrade's exit code (stamp on any real run) +- Good, because it is a one-line change to the wrapper and needs no boot machinery. +- Bad, because it throws away honest signal: a topgrade that was blocked from applying a real upgrade would read as "fresh," so the metric stops meaning "up to date." On this machine the blocked case is the common case, so the metric would be fresh precisely when an upgrade is outstanding. +- Neutral, because the failed steps still surface elsewhere (pending-updates count), so freshness would become redundant rather than wrong. + +** B. Run the full topgrade live with the guard overridden, then reboot +- Good, because it needs no new unit — arm the sentinel, run, reboot. +- Bad, because the dangerous window is the whole rest of the run: mesa swaps early, then topgrade spends minutes on other ecosystems while the live compositor is one new GL context (a new window, the wallpaper daemon) away from SIGABRT. topgrade's own reboot-at-end is that window, not a fix for it. +- Neutral, because it would stamp naturally on success — if it survived. + +** C. Manual TTY ritual only (log out, run topgrade at the console, reboot), plus stamp +- Good, because it is the safest path and needs almost no code — just make the completion stamp. +- Bad, because it is all manual, every guarded-upgrade day; the friction is why it won't happen consistently, which is how the metric got stale in the first place. +- Neutral, because it is exactly what the boot unit automates, so it is really "v1 minus the automation." + +** D. Boot-time armed oneshot, arch-only (this spec) +- Good, because the risky swap happens with nothing live, the run is one bounded transaction with a tiny prompt surface, it stamps on success, and a failure degrades to "boots normally, try again." +- Bad, because it puts a unit on the boot critical path, which must be bounded and non-fatal with care, and it is the most to build. +- Neutral, because it composes with C: the same =maint apply-upgrade= path serves both an interactive TTY run and the boot unit. + +** E. Split the live run: apply the non-blocked remainder now, defer the rest (this spec's everyday path) +- Good, because it is what a careful operator does by hand anyway — and did, on ratio, the day this was decided. The live run succeeds on the common day, topgrade exits 0, the AUR and every other ecosystem stay current, and the guard's abort becomes the rare path rather than the default. +- Bad, because Arch calls any =--ignore= run a partial upgrade. In practice pacman still enforces declared dependencies, so anything needing the newer mesa fails resolution instead of installing broken; the residual exposure is a package with an *unversioned* dependency built against a new ABI, which for mesa/wayland/libdrm is rare. Named, accepted. +- Bad, because a deferred set nobody surfaces is a set that silently never lands — the same trap as the freshness stamp, one layer down. So the script must record the deferred set durably and the panel must show it; this is why D stays in the design as the completion step rather than being replaced. +- Neutral, because it does not change the guard at all; it changes who decides the transaction's contents. + +* Decisions [7/7] + +** DONE Metric means state, not recency +CLOSED: [2026-08-25 Tue 18:45] +- Owner / by-when: Craig / 2026-08-25 +- Context: the stamp gate can mean "ran the sweep" or "system is current." The whole fix differs by which. +- Decision: We will keep the state meaning. Freshness stays stale while a guarded upgrade is genuinely un-applied, and the fix is to make *safe completion* stamp — not to loosen the gate. +- Consequences: easier — the metric stays trustworthy as an is-current signal, and Alternative A is off the table. Harder — completion now needs a real safe path (the rest of this spec) rather than a one-line wrapper change. + +** DONE Everyday mechanism is the split live run (Alternative E); the boot oneshot (D) completes the deferred set +CLOSED: [2026-08-25 Tue 18:30] +- Owner / by-when: Craig / 2026-08-25 +- Context: the first draft made D the primary gesture, which means every guarded-library day is a reboot day. Craig's read while watching the ratio run: when the guard would trip, the rational move is to upgrade everything *except* the guarded and kernel items, then run the rest of topgrade without the yay piece — and that logic should be a script we can keep editing, not something baked into the panel. +- Decision: I will build E as the path UPDATE and TOPGRADE always take on a live session, and keep D as the way the deferred set lands (arm + reboot). C remains the manual fallback through the same script from a TTY (no compositor → nothing blocked → a full run). B stays rejected on the live-swap risk. +- Consequences: easier — the common day is one live run that succeeds; reboots are reserved for the days the deferred set is non-empty, and even then the machine keeps working until the reboot is convenient. Harder — two lists to maintain (the guard's, read from the hook; the kernel list, owned by the script), a state file the panel must render, and the partial-upgrade caveat above to keep an eye on. + +** DONE The script lives in archsetup beside the guard, and maint calls it +CLOSED: [2026-08-25 Tue 18:45] +- Owner / by-when: Craig / 2026-08-25 +- Context: the panel (dotfiles =maint=) and the guard (archsetup =scripts/hypr-live-update-guard=, installed to =/usr/local/bin=) live in different repos, and maint already carries its own copy of the trigger list as =[updates] guard_patterns= in the thresholds TOML. +- Decision: ship the split script in archsetup next to the guard, installed by the same installer step, reading the blocked list from the installed hook so the guard and the script can never disagree. maint's UPDATE/TOPGRADE levers change their =argv= to the script; the TOML patterns stay as the panel's *display-side* mirror (the badge that says a run will defer) and gain a test asserting they match the hook. +- Consequences: easier — one owner for "which libraries are dangerous," and a rebuilt machine gets the script with the guard. Harder — a cross-repo change (archsetup ships it, dotfiles wires it), so the rollout is two commits, archsetup first. + +** DONE What a split run stamps +CLOSED: [2026-08-25 Tue 18:45] +- Owner / by-when: Craig / 2026-08-25 +- Context: under the state framing, a run that deferred six packages left the system *not* current, yet the sweep ran and every other ecosystem is fresh. +- Decision: the script stamps =topgrade_run= only when the deferred set is empty. When it is non-empty it writes the deferred set to its own cache key, and the panel renders that as its own state ("6 deferred — apply on reboot") rather than as stale freshness. The boot oneshot stamps when it completes the deferred set. Freshness keeps meaning "current"; the deferred badge carries the other half. +- Consequences: easier — no signal is thrown away, and the reboot nag has a precise count behind it. Harder — one more cache key and one more probe in maint. + +** DONE Kernel set is held on every everyday run and lands only in the dedicated session, gated on the DKMS result +CLOSED: [2026-08-25 Tue 18:45] +- Owner / by-when: Craig / 2026-08-25 +- Context: I first wrote this as "install the kernel live at apply-on-reboot," on the reasoning that a kernel swap crashes nothing and the modules-vanish window ends with the reboot. Craig asked what happens on velox when the DKMS rebuild fails, and the answer changed the decision. Velox is an encrypted ZFS root with =/boot= inside the root dataset, one kernel (=linux-lts=), and =zfs-dkms=. On a kernel upgrade the DKMS build runs PostTransaction, after the kernel is swapped and the old modules are deleted, so nothing can abort; a failed build leaves a new kernel beside an initramfs built for the old one, whose =zfs.ko= won't load, and the next boot can't import the pool. It is survivable — ZFSBootMenu can boot the pre-pacman snapshot, which holds the old kernel, initramfs, and modules — but it is a recovery session, not an update. Ratio (btrfs root, two kernels, zfs only for a data pool) is exposed only at the pool. The realistic triggers are a kernel major outrunning OpenZFS's supported range, a kernel upgraded without its headers, a toolchain regression, or a full disk. +- Decision: the script holds the kernel set — every installed kernel with its =-headers=, moved as a set, never one without the other — on every everyday run, on both machines, so there is one rule rather than a per-host exception. The kernel set lands only in the dedicated session, live, while a working desktop exists for diagnosing, and the script gates what follows on the result: =dkms status= reports every DKMS module installed for the new kernel version, the initramfs is newer than the kernel image, and on a ZFS root a pre-pacman snapshot exists. A failed gate stops with the failure named and never reboots. The GPU/compositor set follows only after the gate passes — armed for the boot oneshot, or applied from a TTY. Kernels stay off the guard's list (the hook would block a TTY kernel upgrade for no reason). "Install the kernel live at apply-on-reboot" is withdrawn. +- Consequences: easier — an everyday UPDATE can never put velox into the unbootable state, and the day the kernel moves is one Craig chose, sitting at the machine, expecting to handle issues. Harder — the kernel deferral is now standing, so the dedicated session has to happen on a cadence (security fixes ride the kernel), and the panel's deferred count carries a kernel most days; the gate is one more script to test, with fakes for =dkms status= and the image timestamps. + +** DONE Boot run applies exactly the deferred GPU/compositor set, nothing else +CLOSED: [2026-08-25 Tue 18:45] +- Owner / by-when: Craig / 2026-08-25 +- Context: the only packages that need a stopped compositor are the guard's trigger set; the rest run fine live and are what usually fail. The first draft phrased this as =topgrade --only system=; with the split script that wording is stale, and with the kernel decision above the kernel is not part of what boot applies either. +- Decision: the boot oneshot runs the script's =--complete= form scoped to the deferred GPU/compositor set: one pacman transaction, no ecosystem sweep, no kernel. The full topgrade sweep stays a normal live run through the everyday path. +- Consequences: easier — the boot path is fast, has a tiny interactive-prompt surface, and rarely fails. Harder — freshness after a boot run reflects the guarded set specifically, which is what the stamp decision above already accounts for. + +** DONE Arm flag lives on a persistent path and is one-shot +CLOSED: [2026-08-25 Tue 18:45] +- Owner / by-when: Craig / 2026-08-25 +- Context: the guard's =/run= sentinel is tmpfs and cleared on reboot, so it cannot carry an intent across the reboot. A boot that retries forever on failure is its own outage. +- Decision: We will use a persistent flag (=/var/lib/archsetup/=) that the boot unit removes unconditionally at the end of its attempt — success or failure disarms. +- Consequences: easier — the intent survives exactly one reboot and a failed attempt never wedges subsequent boots. Harder — a failed attempt needs re-arming, which is correct (a human decides to try again) but is a manual step. + +* Implementation phases + +** Phase 1 — The split-upgrade script (archsetup) +=scripts/guarded-upgrade= (name open), installed to =/usr/local/bin= by the step that installs the guard. Behaviour as in Design: pending set → blocked set (hook =Target= lines, version-aware) ∪ held-kernel set when a compositor is live → =informant read= if present → =pacman -Syu --noconfirm --ignore=…= → =yay -Sua --noconfirm= → =topgrade --disable system,git_repos -y= → deferred set written to a state file → stamp only when nothing was deferred → exit 0 on a successful live part. The kernel set is derived from what is installed (every =linux*= kernel package and its =-headers=), never a hardcoded pair, and is always held or applied whole. Flags: =--dry-run= (print the plan and the deferred set, change nothing), =--no-topgrade=, =--no-aur=, =--complete= (the dedicated-session form: apply the kernel set live, run the gate, then arm the GPU/compositor set or, with no compositor live, apply it directly; stamp when the deferred set is empty). The gate is its own small script, =kernel-modules-check=: for each kernel under =/usr/lib/modules=, =dkms status= reports every registered module =installed= for it, and its initramfs is newer than its =vmlinuz=; on a ZFS root, a =pre-pacman_= snapshot of the root dataset exists. It exits non-zero with the failing item named, and =--complete= refuses to arm or reboot on that exit. Usable from a TTY at once. Tests (pytest beside the guard's): blocked-set computation against a fixture hook and version map; the kernel set is derived from the installed kernels and held whole on every everyday run; the =--ignore= list is exactly blocked ∪ kernel set; the state file round-trips; stamps only on an empty deferred set; =--dry-run= is IO-free; the gate passes and fails on fake =dkms status= output, image timestamps, and snapshot listings, and =--complete= never reaches the arm step on a failed gate. + +** Phase 2 — Wire maint to it (dotfiles) +UPDATE and TOPGRADE levers change their =argv= to the script; the press-again-to-force sentinel wrap goes away (the driven path never trips the guard). A new probe reads the deferred-set state file and the panel renders "N deferred — apply on reboot" as its own row. A test asserts the TOML =guard_patterns= equal the installed hook's =Target= list. Tests under the maint fake harness. + +** Phase 3 — "Apply on reboot" and the boot-time unit (archsetup + maint) +The panel action installs the held-kernel set live, writes the persistent arm flag, and offers to reboot. =archsetup-boot-upgrade.service=, installed by the installer: =ConditionPathExists= the flag, =Before=getty@tty1.service=, =TimeoutStartSec= bounded, non-fatal to every session target, =ExecStart= runs the script's =--complete= form as the user and removes the flag unconditionally. Installer step + unit file + the =/var/lib/archsetup/= flag directory. Tests in =tests/installer-steps/= for the install step; the arm action's tests in maint; a documented manual boot test (defer, arm, reboot, observe) in =todo.org= under Manual testing and validation. + +** Phase 4 — Docs, rollout, and both daily drivers +Document the flow (arm → reboot → console upgrade → session). Roll the unit to velox and ratio (installer already covers a rebuild; existing machines need the one-time install). Confirm the ratio path matches. + +* Acceptance criteria +- [ ] With a guarded library pending and Hyprland live, UPDATE applies everything else, exits 0, and the panel shows the exact deferred set; the guard hook does not fire. +- [ ] The same run with no compositor live (a TTY) applies everything but the kernel set and stamps only if nothing was deferred. +- [ ] A =--complete= run whose DKMS build fails stops before arming or rebooting, names the failure, and leaves the machine running on the old kernel; on velox the pre-pacman snapshot it required is bootable from ZFSBootMenu. +- [ ] With a guarded library pending, arming and rebooting applies it in the console before Hyprland starts, and =maint status= then reads a fresh =topgrade_age=. +- [ ] A boot-upgrade failure (a failed step, a timeout, an aborted transaction) never blocks the session: the machine boots into Hyprland, the flag is cleared, and the panel still shows the pending work. +- [ ] Unread Arch news does not wedge the boot run (=informant read= precedes the transaction). +- [ ] A guarded upgrade completed from a TTY via the Phase-1 path stamps freshness identically to the boot unit. +- [ ] The =hypr-live-update-guard= hook is unchanged and still blocks a live guarded swap. + +* Readiness dimensions +Answer each, or write "N/A because…". +- Data model & ownership: the arm flag (=/var/lib/archsetup/=, installer-owned) and the =topgrade_run= cache key (maint-owned). No user-authored data. +- Errors, empty states & failure: the boot unit is best-effort and self-disarming; every failure path lands in "boot normally, metric stays stale, re-arm to retry." Named, non-silent. +- Security & privacy: relies on the existing =%cjennings NOPASSWD: ALL=; the unit runs the upgrade as the user via sudo, adds no new privilege. Note the NOPASSWD breadth as a pre-existing fact, not introduced here. +- Observability: the boot run's output is on the console; its systemd unit status and journal record success/failure; the panel reflects the cleared or still-pending state after boot. +- Performance & scale: one pacman/yay transaction at boot; bounded by =TimeoutStartSec=. Negligible boot-time cost when the flag is absent (=ConditionPathExists= skips the unit). +- Reuse & lost opportunities: reuses the guard's trigger list by reading the installed hook (single source of truth for "which libs are dangerous"), =informant=, =maint stamp=, =checkupdates=, and topgrade's own step switches. The one duplicate that exists today — maint's TOML =guard_patterns= — is kept as a display mirror and pinned to the hook by a test rather than removed. +- Architecture fit & weak points: integration points are the pacman hook set, getty autologin ordering, and the maint cache. Weak point: the =Before=getty@tty1= ordering is load-bearing for safety; a parallel run reintroduces the live-swap crash. Mitigated by making the ordering explicit and tested-by-inspection. +- Config surface: the arm flag path and the timeout. Defaults safe (absent flag = no-op). +- Documentation plan: a short "reboot to apply guarded upgrades" note in the maint docs; the installer step self-documents in-comment. +- Dev tooling: installer-step pytest for Phase 2; maint unit tests for Phases 1 and 3; a manual boot test in =todo.org=. +- Rollout, compatibility & rollback: additive; removing the unit and flag reverts fully. Existing machines need a one-time install; a rebuild gets it from the installer. Rollback leaves the guard and manual TTY path intact. +- External APIs & deps: topgrade =--only system=, =informant read=, =yay=, =maint stamp= — all verified present on velox this session. No external service. + +* Risks, Rabbit Holes, and Drawbacks +- Boot critical path: the unit sits ahead of autologin, so a hang would delay boot. Mitigated by =TimeoutStartSec= and non-fatal wiring; worst case is a bounded delay, then a normal session. +- Interactive prompts under no stdin: =yay=/pacman can still prompt (provider choice, replace, AUR review) even with =assume_yes=. The =--only system= scope and =--noconfirm=-style flags shrink this to near zero, but a prompt with no stdin fails the run (benign) — needs a genuinely non-interactive invocation, verified in Phase 2. +- Partial ecosystem state: N/A for the GPU hazard — each pacman run is one atomic transaction, so there is no half-swapped library. The =--ignore= run is a partial upgrade in Arch's sense; pacman's dependency resolution is the safety net, and the residual unversioned-ABI exposure is accepted in Alternative E. +- Orphans that block resolution: a package dropped from the repo but still pinning an old version (ratio's =qemu-block-gluster= on 2026-08-25) fails the whole transaction. The script should detect the "could not satisfy dependencies" case, name the foreign package, and stop with the remedy — never =-Rdd= on its own. +- The kernel on a DKMS ZFS root: a failed =zfs-dkms= build after the kernel swap cannot be aborted (the DKMS hooks are PostTransaction) and leaves velox unbootable on the new kernel. Mitigated by holding the kernel set on every everyday run, landing it only in the dedicated session behind the gate, and by the standing fallback: =/boot= lives in the root dataset, the =05-zfs-snapshot= hook snapshots it before every transaction, and ZFSBootMenu can boot that snapshot. The pacman cache also keeps the previous kernel and =zfs-dkms= for a downgrade. Ratio's exposure is its data pool only (btrfs root, two kernels). +- Standing kernel deferral: because the everyday run never moves the kernel, the dedicated session has to happen on a cadence or kernel security fixes sit unapplied. The panel's deferred row is the reminder; a stale-kernel age in maint is a possible follow-up. + +* Testing / Verification / Rollout +Phase-1 and Phase-3 logic under the maint fake harness; Phase-2 install under =tests/installer-steps/=. The one thing no unit test can cover — that an armed reboot actually applies the upgrade pre-session and stamps — is a scripted manual test in =todo.org= (arm with a guarded lib pending, reboot, confirm the console run, the fresh metric, and a normal session). Roll to velox first, then ratio. + +* Review and iteration history +** 2026-08-25 Tue @ 18:45 -0600 — Craig Jennings — author +- What: closed all seven decisions. Reversed the kernel decision (hold on every everyday run; land only in the dedicated session, gated on DKMS built, initramfs fresh, snapshot present; withdrew "install live at apply-on-reboot"), reworded the boot-scope decision for the split design, added the =kernel-modules-check= gate to Phase 1 and the acceptance criteria, and wrote the velox failure chain and the ZFSBootMenu fallback into Risks. +- Why: on velox a failed =zfs-dkms= rebuild after a kernel swap is unabortable and unbootable; that belongs in a session I chose, not in an update I expected to touch applications. +- Artifacts: this session's log (velox boot layout verified live: ZBM on the ESP, =/boot= in =zroot/ROOT/default=, one kernel, =zfs-dkms 2.4.4=). +** 2026-08-25 Tue @ 18:30 -0600 — Craig Jennings — author +- What: made the split live run (E) the everyday path and the boot oneshot (D) the completion step; added the script-ownership, stamp-semantics, and kernel-hold decisions; rewrote the phases around the script; added the orphan-blocks-resolution risk. +- Why: watching a 724-of-730 guarded run succeed by hand on ratio made it obvious the guard's abort should be the rare path, and that the logic belongs in an editable script the panel calls rather than in the panel. +- Artifacts: this session's log; the ratio run (=ratio-upgrade.service=, =/var/log/ratio-upgrade.log=). +** 2026-08-25 Tue @ 06:39:42 -0600 — Craig Jennings — author +- What: initial draft. +- Why: the topgrade-freshness metric reads permanently stale because the guard blocks the arch step; designing a safe completion path rather than loosening the gate. +- Artifacts: this session's log; =hypr-live-update-guard= hook; maint =topgrade_age= probe. diff --git a/docs/workflows/system-health-check.org b/docs/workflows/system-health-check.org index b4f34a5..43aeff5 100644 --- a/docs/workflows/system-health-check.org +++ b/docs/workflows/system-health-check.org @@ -235,6 +235,16 @@ Updates are separate from issue investigation. After all issues are addressed (o 5. *Host-specific kernel watches.* On ratio: if =linux=, =linux-lts=, =linux-firmware=, or a major =mesa= bump is pending, run the addendum at [[file:strix-soak-watch.org][docs/workflows/strix-soak-watch.org]] before topgrade. Retire the addendum (delete the file + this bullet) when the strix-lts custom kernel is retired. 6. Run =topgrade= for the actual update (config at =~/.config/topgrade.toml=). On maint hosts (ratio, velox) plain =topgrade= resolves to the dotfiles PATH wrapper, which stamps the console's topgrade-freshness metric on success — no extra step. If the run happened outside the wrapper somehow, =maint stamp topgrade= records it by hand. 7. If linux-firmware, kernel, or Mesa were updated, recommend a reboot +8. On a ZFS-root host, after a kernel bump confirm the new initramfs carries the zfs module before rebooting: =sudo lsinitcpio /boot/initramfs-linux-lts.img | grep -c 'zfs.ko'=. The =sudo= is load-bearing: the images are 0600, so an unprivileged =lsinitcpio= exits 1 with "Unable to read file" on stderr and nothing on stdout, and once piped into =grep -c= that empty stdout reads as a count of 0 and looks exactly like a missing module (velox, 2026-09-12). + +*** Before a firmware (BIOS) update + +Firmware stays a manual step (=topgrade.toml= keeps =[firmware] upgrade = false=); =fwupdmgr update= stages it and the next reboot flashes it. Before staging, capture the two things a bad reboot will make you guess at: + +1. =sudo efibootmgr -v= — every boot entry with its loader path and command line, pasted into the session's context file. +2. Secure Boot state — =bootctl status 2>/dev/null | grep -i 'secure boot'=. + +After the flash, if the machine reports no bootable device, check Secure Boot *first*. The Framework 04.02 update on velox re-enabled it, which rejects the unsigned ZFSBootMenu loader and reads as "Default Boot Device Missing" rather than a security violation; the boot entries were untouched (see the Known Issues Log, 2026-09-12). Only when Secure Boot is off and =efibootmgr -v= from a stick shows the entry gone does the boot-entry recreate apply (velox: =docs/2026-08-15-velox-uefi-boot-entry-reference.org= in archsetup). *** Two-Stage Reboot Pattern (MANDATORY if Phase 3 installed kernel / iproute2 / systemd / NetworkManager) @@ -1036,3 +1046,28 @@ Each entry is scoped to one host (or =any=). When Phase 1 cross-references findi - Functional status at 2026-06-13 check: no coredumps yet that day; Telegram scans still worked from cached chat state. Treat as an app/server-container crash, not a machine-health fault. - Classification: KNOWN — annotate future =telega-server= coredumps on ratio as =KNOWN — dockerized telega-server musl SIGSEGV= if the signature matches =tdat_plist_value= / unexpected plist value or otherwise stays inside the Telega container. Escalate only if crashes become continuous, break Telegram workflows, or appear after moving off the Docker musl build. - Deferred remediation options, in order of least disruption: update the Emacs =telega= package, rebuild/pull a newer =telega-server= image, pin a known-good pre-2026-06 image digest, build =telega-server= natively, or report upstream with =coredumpctl= and log evidence. + +** 2026-09-12: velox — Framework BIOS update re-enabled Secure Boot (reads as "no bootable device") +:host: velox +- Symptom: after fwupd staged system firmware 0.0.3.5 → 0.0.4.2 and the reboot flashed it, the BIOS reported "Default Boot Device Missing / no bootable drive". Indistinguishable from the NVRAM wipe that forced the 2026-08-13 reinstall. +- Actual cause: the update set Enforce Secure Boot = Enabled. The unsigned ZFSBootMenu loader was rejected and the firmware reported it as a missing device, not a security violation. Boot0001 ZFSBootMenu survived intact with its command line; no =efibootmgr= was needed. +- The tell: booting the Ventoy stick showed shim's MOK management screen. +- Fix: F2 → Security → Secure Boot → Enforce Secure Boot = Disabled, F10. Boots straight into ZBM. Firmware confirmed at 04.02, pools healthy. +- Prevention: the pre-firmware-update checklist in Phase 3 (record =efibootmgr -v= and the Secure Boot state before staging). Check Secure Boot before assuming NVRAM loss. +- Still open: velox's ESP has no removable-media fallback (=/efi/EFI/BOOT= absent), so a real NVRAM wipe would still need the stick. Filed in archangel, which owns the ZBM install, as [#C] "Installed systems have no removable-media boot fallback on the ESP" (2026-09-12); it ships with the next ISO rebuild after it lands. + +** 2026-09-12: any — fwupdmgr activates passim, a public LAN listener +:host: any +- Symptom: running =fwupdmgr= (refresh, update) D-Bus-activates =passim.service=, fwupd's LAN metadata-sharing daemon, which listens on =0.0.0.0:27500= and trips the maint listeners check to crit. +- The unit is static (no =[Install]= section), so =systemctl disable= is a no-op and it comes back on the next fwupdmgr run. Masking is what holds: =systemctl mask passim.service=. Ratio has been masked since 2026-07-21; velox was stopped and disabled on 2026-09-12 (the disable being the no-op) and got the mask on 2026-09-13. The installer masks it as part of installing fwupd. =P2pPolicy=nothing= under =[fwupd]= in =/etc/fwupd/fwupd.conf= also works, but that file is pacman-owned and invites pacnew churn, so the mask is the form in use. +- Classification: KNOWN — a passim listener means a machine that predates the mask or lost it; mask it, don't allowlist it. + +** 2026-09-12: velox — topgrade containers step fails on locally built images +:host: velox (ratio has the same shape with its own local images) +- Symptom: topgrade exits 1 after a clean package run because the containers step tries to =docker pull= images that were built locally (=cj/telega-server=, =telega-server-glycin=) and gets "pull access denied". The wrapper then never writes the topgrade-freshness stamp; =maint stamp topgrade= by hand after confirming the package steps succeeded. +- Classification: KNOWN — the step cannot succeed while local-only images exist. The fix (disable the containers step, or list the images under =ignored_containers=) is tracked on the topgrade guarded-upgrade task in archsetup's todo. + +** 2026-09-12: velox — mkinitcpio "Possibly missing firmware" for xhci_pci_renesas and qat_6xxx +:host: velox +- Stock Arch mkinitcpio noise on a kernel rebuild: =xhci_pci_renesas= wants the Renesas USB controller blob (AUR =upd72020x-fw=) and =qat_6xxx= is Intel QuickAssist firmware. Neither is hardware this machine has. +- Classification: KNOWN — harmless; annotate and move on unless the named hardware appears. diff --git a/scripts/cmail-setup-finish.sh b/scripts/cmail-setup-finish.sh index 949023f..8c27eda 100755 --- a/scripts/cmail-setup-finish.sh +++ b/scripts/cmail-setup-finish.sh @@ -1,32 +1,37 @@ #!/usr/bin/env bash # SPDX-License-Identifier: GPL-3.0-or-later -# cmail-setup-finish.sh — finish Proton Mail Bridge + cmail-action setup after -# Bridge first-run. Idempotent; safe to re-run after a Bridge cert rotation or -# a claude-templates re-clone. +# cmail-setup-finish.sh — finish Proton Mail Bridge setup after Bridge +# first-run. Idempotent; safe to re-run after a Bridge cert rotation. # # Pre-reqs (the script aborts if any are missing): # - protonmail-bridge installed (archsetup handles it) # - You have run 'protonmail-bridge --cli', logged in, and quit at least once # (the script looks for state at ~/.config/protonmail/bridge-v3/) -# - claude-templates cloned at ~/projects/claude-templates # - dotfiles stowed (~/.config/.cmailpass.gpg present) # +# Not a pre-req, but checked and warned about: cmail-action on PATH. rulesets' +# `make install` links it, and session start runs that, so on a machine that +# runs agent sessions it arrives without anyone asking. On one that doesn't, +# it needs the command by hand. The script never invokes it either way. +# # What it does: # 1. Decrypts ~/.config/.cmailpass.gpg → ~/.config/.cmailpass (mode 0600) # 2. Copies Bridge's self-signed cert → ~/.config/protonbridge.pem -# 3. Symlinks ~/projects/claude-templates/.ai/scripts/cmail-action.py -# → ~/.local/bin/cmail-action -# 4. Removes the leftover ~/.config/autostart/Proton Mail Bridge.desktop +# 3. Removes the leftover ~/.config/autostart/Proton Mail Bridge.desktop # stub (it double-launches Bridge alongside the systemd user service # and throws an "orphan instance" dialog every login) -# 5. Installs a wait-for-dns drop-in so Bridge doesn't spam +# 4. Installs a wait-for-dns drop-in so Bridge doesn't spam # name-resolution errors during the early-boot DNS race -# 6. Enables + starts the protonmail-bridge user service -# 7. Verifies Bridge is listening on 127.0.0.1:1143 / :1025 +# 5. Enables + starts the protonmail-bridge user service +# 6. Verifies Bridge is listening on 127.0.0.1:1143 / :1025 +# +# It no longer installs cmail-action. That moved to rulesets +# (claude-templates/bin/), whose `make install` owns the symlink. set -euo pipefail err() { printf 'error: %s\n' "$*" >&2; exit 1; } +warn() { printf 'warning: %s\n' "$*" >&2; } info() { printf '==> %s\n' "$*"; } ok() { printf ' %s\n' "$*"; } @@ -47,9 +52,20 @@ bridge_state="$HOME/.config/protonmail/bridge-v3" [ -d "$bridge_state" ] \ || err "Bridge has no state at $bridge_state — run 'protonmail-bridge --cli' and log in first" -cmail_action_src="$HOME/projects/claude-templates/.ai/scripts/cmail-action.py" -[ -f "$cmail_action_src" ] \ - || err "cmail-action.py not found at $cmail_action_src — clone claude-templates first" +# cmail-action is no longer this script's to install. It lives in rulesets at +# claude-templates/bin/, and rulesets' `make install` links everything there +# into ~/.local/bin. Session start runs that, so on a machine that runs agent +# sessions the symlink arrives on its own; on one that doesn't, it needs the +# command below. +# +# A warning rather than an abort, because this script never invokes the tool. +# Its job is to leave Bridge working, and it can finish that whether or not a +# mail client has been linked yet. Aborting here would make Bridge setup +# depend on rulesets being cloned and installed first, an ordering neither +# repo otherwise needs, and would strand a fresh machine with Bridge ready and +# the script refusing to configure it. +command -v cmail-action >/dev/null 2>&1 \ + || warn "cmail-action not on PATH — run 'make -C ~/code/rulesets install' before sending mail" cmailpass_enc="$HOME/.config/.cmailpass.gpg" [ -f "$cmailpass_enc" ] \ @@ -69,13 +85,7 @@ cert_dst="$HOME/.config/protonbridge.pem" cp "$cert_src" "$cert_dst" ok "copied $cert_src → $cert_dst" -# 4. Symlink cmail-action -info "symlinking cmail-action" -mkdir -p "$HOME/.local/bin" -ln -sf "$cmail_action_src" "$HOME/.local/bin/cmail-action" -ok "linked $HOME/.local/bin/cmail-action → $cmail_action_src" - -# 5. Remove leftover XDG autostart stub +# 4. Remove leftover XDG autostart stub # The systemd --user service is the canonical launcher. The autostart .desktop # starts a second Bridge instance that can't get the lock and pops up an # "orphan instance" dialog every login. @@ -88,7 +98,7 @@ else ok "no autostart stub present" fi -# 6. Install wait-for-dns drop-in +# 5. Install wait-for-dns drop-in # User-instance systemd doesn't carry network-online.target / nss-lookup.target, # so the packaged unit's After=network.target doesn't imply DNS readiness. # Bridge starts before the resolver is up and its first API calls all fail @@ -107,7 +117,7 @@ ok "wrote $dropin_file" systemctl --user daemon-reload ok "reloaded systemd user units" -# 7. Enable + start systemd user service +# 6. Enable + start systemd user service info "enabling protonmail-bridge user service" was_active=0 systemctl --user is-active --quiet protonmail-bridge.service && was_active=1 @@ -119,7 +129,7 @@ else ok "service active" fi -# 8. Verify +# 7. Verify info "verifying Bridge is listening" listening="$(ss -ltn 2>/dev/null || true)" missing="" diff --git a/scripts/post-rebuild-check b/scripts/post-rebuild-check new file mode 100755 index 0000000..aa7ef83 --- /dev/null +++ b/scripts/post-rebuild-check @@ -0,0 +1,676 @@ +#!/bin/sh +# SPDX-License-Identifier: GPL-3.0-or-later +# post-rebuild-check - the eight checks a rebuilt machine actually needs. +# +# A rebuilt machine looks finished and isn't. Five gaps surfaced on velox +# within two days of the 2026-08-13 reinstall, and three of them LOOKED +# fine: a stowed unit file, an enabled timer, a present git clone. Each +# check below is cheap and turns a silent no-op into a visible line: +# +# 1. failed systemd units, user and system scope (calendar-sync failed +# every 15 minutes for two days with nobody watching) +# 2. user unit files present but not enabled (roam-sync and +# signal-receive came back linked and inert -- a unit file being +# present is not the same as running) +# 3. tracked *.example files whose real sibling is missing (three +# *.local.el were gone on velox; the .example survives in git, +# the real file never does) +# 4. gitignore-mode projects missing tooling paths their own .gitignore +# names (a reinstall drops every such project's untracked working +# state -- 374 files in .emacs.d's case -- and nothing carries it) +# 5. signal-cli holds no registered account (velox lost its registration +# in the rebuild; at the time agent-text relayed into velox, so that +# silently broke paging for the WHOLE fleet. agent-text now walks +# AGENT_TEXT_RELAYS in order and skips itself, so an unregistered +# machine is only fatal when no relay host is registered either) +# 6. every NTP source is named by hostname (a wrong clock fails the +# DoT/DNSSEC validation this machine's DNS runs on, so nothing +# resolves -- including the NTP pool that would fix the clock; velox +# deadlocked exactly this way 2026-08-19 and needed a second device) +# 7. hypridle installed but not running (nothing then triggers idle lock +# or suspend, so a laptop runs until its battery is gone -- which is +# how velox reset the RTC that caused check 6's deadlock in the first +# place; a caffeine remembered from an earlier boot is the known cause) +# 8. a working repo cloned from the read-only https endpoint (correct +# for a stranger with no key on the server, wrong for this machine, +# which finds out at the first push with a 403 -- velox's dotfiles +# remote sat that way for four days after its rebuild) +# +# The .gitignore rule in check 4 is what scopes it: a tooling path is only +# expected where the project's own .gitignore names it, so a project that +# never had a todo.org never flags. The ignore file is the project's own +# record of what it is supposed to hold untracked. +# +# EVERY PROBE FAILS CLOSED. A check that cannot run reports a finding, never +# a pass. This matters more here than anywhere else in the script: the whole +# point is catching silent no-ops, so a silent no-op in the checker would be +# the worst possible defect. `systemctl --user` exits 1 with empty output +# when there is no user bus -- over ssh, from cron, under sudo, on a TTY +# before the graphical session starts -- and reading that as "no failed +# units" would report a machine as healthy exactly when nothing was checked. +# +# Exit 0 when every check is clean, 1 when any check found something, +# 2 on usage error. +# +# Test seams (env; for each, set-but-empty means "the probe ran and found +# nothing", unset means "run the real probe"): +# PRC_FAILED_UNITS newline list of "scope:unit" (scope user|system) +# PRC_UNIT_STATES newline list of "unit-file state" replacing the +# user-unit-dir enumeration + is-enabled calls +# PRC_LOCAL_SCAN_ROOTS NEWLINE-separated roots for the *.example scan +# (default: ~/.emacs.d ~/.dotfiles) +# PRC_PROJECT_ROOTS NEWLINE-separated project dirs for check 4 +# (default: ~/code/* ~/projects/* ~/.emacs.d +# ~/.dotfiles) +# PRC_SIGNAL_ACCOUNTS signal-cli listAccounts output; "" = no account, +# the special value MISSING = binary absent +# PRC_NTP_SOURCES newline list of configured NTP server addresses; +# the special value MISSING = no NTP daemon active +# PRC_CHRONY_CONF path to chrony.conf (a fixture, under test) -- the +# confdir it names is what decides which drop-ins count +# PRC_IDLE_DAEMON pgrep output for hypridle; "" = installed but not +# running, the special value MISSING = not installed +# PRC_REPO_REMOTES newline list of "path origin-url"; an empty URL +# means origin could not be read +# PRC_UNITS_EXPECTED_DISABLED +# newline list of units whose not-enabled state is +# deliberate here, replacing the file below +# PRC_UNITS_EXPECTED_DISABLED_FILE +# path to that list (default: +# $XDG_CONFIG_HOME/post-rebuild-check/units-expected-disabled). +# One unit per line, # starts a comment. Machine-local +# on purpose: the same unit is correctly enabled on one +# box and not another +# PRC_SYSTEMCTL path to the systemctl binary (a fake, under test) +# PRC_SYSTEMCTL_TIMEOUT seconds to allow each systemctl call (default 5) +# +# Roots are newline-separated, not space-separated, because a POSIX +# `for root in $var` splits on spaces and turns one real directory into +# several imaginary missing ones. + +usage() { + cat <<'EOF' +post-rebuild-check - verify a rebuilt machine is actually finished + +Runs the eight checks that caught velox's 2026-08 reinstall gaps: failed +units, present-but-inert user units, orphaned *.example configs, missing +per-project tooling state, the signal-cli registration, whether time sync +can recover from a wrong clock without DNS, whether anything still +triggers idle lock and suspend, and whether the working repos can push. + +Usage: post-rebuild-check [--help] + +Exit 0 when every check is clean, 1 when any check found something. +Every probe fails closed: a check that cannot run is a finding, not a pass. +EOF +} + +case "${1:-}" in + --help|-h) usage; exit 0 ;; + "") ;; + *) echo "post-rebuild-check: unknown argument: $1" >&2; usage >&2; exit 2 ;; +esac + +# Own the internal flags rather than inheriting them, so a caller's unrelated +# variable of the same name cannot manufacture or mask a finding. +TOTAL_FINDINGS=0 +CHECK_FINDINGS=0 +FINDING_LINES="" +signal_missing="" +ntp_missing="" +idle_absent="" + +# Every systemctl call is bounded. A wedged user manager spins and answers +# nothing -- seen live on velox 2026-08-17, where `is-enabled`, `cat`, and +# `list-unit-files` all hung while `list-units` still returned. Unbounded, this +# script would hang on the first unit and never reach the remaining checks, +# which is a worse failure than reporting nothing: a check that hangs is its +# own outage, and the machine most in need of checking is the one it hangs on. +# A timeout yields empty output and a non-zero status, and both are already +# handled as findings, so bounding the call is all that is needed to fail closed. +CHRONY_CONF=${PRC_CHRONY_CONF:-/etc/chrony.conf} +SCTL_TIMEOUT=${PRC_SYSTEMCTL_TIMEOUT:-5} +SYSTEMCTL=${PRC_SYSTEMCTL:-systemctl} + +sctl() { + if command -v timeout >/dev/null 2>&1; then + timeout "$SCTL_TIMEOUT" "$SYSTEMCTL" "$@" + else + # Say so rather than dropping the bound silently: without timeout a + # wedged manager hangs this run indefinitely, and the whole point of + # the bound is that a check which hangs reports nothing at all. + [ -n "${sctl_unbounded_warned:-}" ] || { + echo "post-rebuild-check: timeout(1) not found — systemctl calls are UNBOUNDED and may hang" >&2 + sctl_unbounded_warned=1 + } + "$SYSTEMCTL" "$@" + fi +} + +WORK=${TMPDIR:-/tmp}/.post-rebuild-check.$$ +if ! mkdir "$WORK" 2>/dev/null; then + # Every check stages its input through a file in here. Without it each + # loop would read nothing and every check would come back clean, which is + # the one failure this script must never produce. + echo "post-rebuild-check: cannot create a work directory under ${TMPDIR:-/tmp}" >&2 + echo " nothing was checked; this is not a pass" >&2 + exit 1 +fi +trap 'rm -rf "$WORK"' EXIT HUP INT TERM + +STAGE="$WORK/stage" + +finding() { + CHECK_FINDINGS=$((CHECK_FINDINGS + 1)) + TOTAL_FINDINGS=$((TOTAL_FINDINGS + 1)) + FINDING_LINES="${FINDING_LINES} DEVIATION: $1 +" +} + +# Print the check's one visible line, then its findings. The visible line +# is the point: a silent no-op is exactly what let the gaps sit unseen. +report() { + if [ "$CHECK_FINDINGS" -eq 0 ]; then + echo "$1 — ok" + else + echo "$1 — $CHECK_FINDINGS finding(s)" + printf '%s' "$FINDING_LINES" + fi + CHECK_FINDINGS=0 + FINDING_LINES="" +} + +# Stage a value into $STAGE for the read loops. A failed write is fatal for +# the same reason a missing work directory is. +stage() { + if ! printf '%s\n' "$1" > "$STAGE" 2>/dev/null; then + echo "post-rebuild-check: cannot write $STAGE" >&2 + echo " nothing was checked; this is not a pass" >&2 + exit 1 + fi +} + +# --- 1. failed units ------------------------------------------------------ + +if [ -n "${PRC_FAILED_UNITS+set}" ]; then + failed=$PRC_FAILED_UNITS +else + failed="" + if user_out=$(sctl --user list-units --state=failed --no-legend --plain 2>/dev/null); then + failed=$(printf '%s' "$user_out" | awk 'NF {print "user:"$1}') + else + finding "could not query user units (no user bus?) — nothing was checked in this scope" + fi + if sys_out=$(sctl list-units --state=failed --no-legend --plain 2>/dev/null); then + failed="$failed +$(printf '%s' "$sys_out" | awk 'NF {print "system:"$1}')" + else + finding "could not query system units — nothing was checked in this scope" + fi +fi +stage "$failed" +while IFS= read -r line; do + [ -n "$line" ] || continue + scope=${line%%:*} + unit=${line#*:} + finding "$scope unit failed: $unit" +done < "$STAGE" +report "check 1/8: failed units" + +# --- 2. user unit files present but not enabled --------------------------- +# +# Units nothing intends to enable here are read from a machine-local list. +# "Enabled" is this check's proxy for "will actually run", and the proxy is +# wrong for a unit nobody means to enable on this box. velox carries four, for +# four different reasons: geoclue-agent is redundant because hyprland's +# exec-once starts the binary directly, emacs is started on demand by +# emacsclient, obs-record-watchdog only matters while recording, and +# obsbot-wb-guard needs an OBSBOT the machine does not have. Left unexempted +# they report at every run, and four permanent lines in front of every real one +# teach you to skim the output -- the same argument check 4 makes about +# CLAUDE.md. +# +# Machine-local rather than a marker in the shared unit file, because +# obsbot-wb-guard is correctly ENABLED on ratio. One unit, a different right +# answer per machine, so the shared file cannot hold the answer. +# +# An entry that turns out to be enabled after all is still a finding. Without +# that the list rots into somewhere real findings go to die, which is worse +# than the noise it removes. + +EXPECT_DISABLED_FILE="${PRC_UNITS_EXPECTED_DISABLED_FILE:-${XDG_CONFIG_HOME:-$HOME/.config}/post-rebuild-check/units-expected-disabled}" +if [ -n "${PRC_UNITS_EXPECTED_DISABLED+set}" ]; then + expect_disabled=$PRC_UNITS_EXPECTED_DISABLED +elif [ -f "$EXPECT_DISABLED_FILE" ]; then + expect_disabled=$(cat "$EXPECT_DISABLED_FILE" 2>/dev/null) +else + expect_disabled="" +fi +# Strip comments and blanks once, here, so the membership test below is a +# plain word match. The reason a unit is exempt is the most useful thing about +# the entry, so the format has to carry one. +expect_disabled=$(printf '%s\n' "$expect_disabled" \ + | sed 's/#.*//' | awk 'NF {print $1}') + +if [ -n "${PRC_UNIT_STATES+set}" ]; then + states=$PRC_UNIT_STATES +else + states="" + unit_dir="${XDG_CONFIG_HOME:-$HOME/.config}/systemd/user" + if [ ! -d "$unit_dir" ]; then + finding "no user unit directory at $unit_dir — nothing was checked" + else + for f in "$unit_dir"/*.timer "$unit_dir"/*.service; do + # -L as well as -e: a stow symlink whose target moved in the + # rebuild is exactly the "looked fine" case this check is for, + # and -e is false for a broken link. + [ -e "$f" ] || [ -L "$f" ] || continue + name=$(basename "$f") + # A link with nothing behind it is its own finding, decided on the + # filesystem rather than from systemd. `is-enabled` calls a + # dangling link "not-found" -- the same answer it gives for a unit + # that was never installed -- so routing this through the state + # table below would drop it silently. + if [ -L "$f" ] && [ ! -e "$f" ]; then + finding "stowed unit file points at a missing target: $name" + continue + fi + # is-enabled exits non-zero AND prints a state for disabled and + # linked, so the exit code cannot distinguish "this unit is + # disabled" from "the query failed". The output can: a real answer + # is always a word. Empty means no answer, which is a finding + # rather than a silent skip -- with no user bus (ssh, cron, sudo, + # a TTY before the graphical session) every unit answers empty, + # and treating that as unknown-so-ignore would pass the machine + # while reading nothing at all. + # + # No separate bus probe: `is-system-running` and + # `show-environment` both block here, and a check that can hang is + # its own outage. + state=$(sctl --user is-enabled "$name" 2>/dev/null) + if [ -z "$state" ]; then + finding "could not read the enablement state of $name — it was not checked" + continue + fi + states="${states}${name} ${state} +" + done + fi +fi +stage "$states" +# A second copy for the sibling-timer lookup below, so the awk that reads it +# is never the same open file as the loop reading it. +cp "$STAGE" "$WORK/states" 2>/dev/null || { + echo "post-rebuild-check: cannot write $WORK/states" >&2 + echo " nothing was checked; this is not a pass" >&2; exit 1; } +while read -r name state; do + [ -n "$name" ] || continue + case "$state" in + disabled|linked) ;; + *) continue ;; + esac + # A timer-activated service is SUPPOSED to sit linked-and-not-enabled: + # the timer owns activation, and enabling the service as well would run + # it at boot on top of its schedule. So a service is suppressed only when + # its sibling timer can actually start it (enabled), or when the timer is + # itself inert and therefore the finding already -- reporting both would + # name one gap twice. A masked, static, or not-found timer starts + # nothing, so the service beneath it is as dead as one with no timer. + case "$name" in + *.service) + timer="${name%.service}.timer" + tstate=$(awk -v t="$timer" '$1 == t {print $2; exit}' "$WORK/states") + # enabled-runtime (enabled until reboot) and generated (something + # produced and installed it) are live activation paths, so the + # service under one is being started and is not a finding. + # disabled and linked suppress for a different reason: the timer + # is then the finding itself, reported in its own right. + # + # "indirect" deliberately does NOT suppress. It means the unit + # file itself is not enabled, only that some Also= relative might + # be, so nothing here is known to start the service. The + # fail-closed rule says the uncertain case flags. + case "$tstate" in + enabled|enabled-runtime|generated) continue ;; + disabled|linked) continue ;; + esac + ;; + esac + # Deliberately not enabled on this machine. Checked last, so it suppresses + # only this finding and never the dangling-link one decided above on the + # filesystem. + case " +$expect_disabled +" in + *" +$name +"*) continue ;; + esac + finding "unit file present but not enabled: $name ($state)" +done < "$STAGE" +# The exemption list, checked in the other direction. An entry whose unit is +# enabled after all suppresses nothing, and leaving it there is how the list +# turns into a place real findings go to die. The loop above cannot catch this: +# it skips any state that is not disabled or linked, so an enabled unit never +# reaches it. +printf '%s\n' "$expect_disabled" > "$WORK/expect" 2>/dev/null || { + echo "post-rebuild-check: cannot write $WORK/expect" >&2 + echo " nothing was checked; this is not a pass" >&2; exit 1; } +while IFS= read -r name; do + [ -n "$name" ] || continue + estate=$(awk -v u="$name" '$1 == u {print $2; exit}' "$WORK/states") + case "$estate" in + enabled|enabled-runtime) + finding "$name is listed as expected-disabled but is $estate — drop the stale exemption" ;; + esac +done < "$WORK/expect" +report "check 2/8: unit files" + +# --- 3. *.example files whose real sibling is missing --------------------- + +if [ -n "${PRC_LOCAL_SCAN_ROOTS+set}" ]; then + scan_roots=$PRC_LOCAL_SCAN_ROOTS +else + scan_roots="$HOME/.emacs.d +$HOME/.dotfiles" +fi +printf '%s\n' "$scan_roots" > "$WORK/roots" 2>/dev/null || { + echo "post-rebuild-check: cannot write $WORK/roots" >&2; exit 1; } +while IFS= read -r root; do + [ -n "$root" ] || continue + if [ ! -d "$root" ]; then + finding "scan root missing: $root" + continue + fi + # Vendored package trees ship their own .example docs; those belong to + # the package, not to this machine, so they are noise in front of the + # real findings this check exists for. + # + # -prune, not -not -path: the latter filters find's OUTPUT while still + # descending, so an unreadable directory inside a tree we deliberately + # ignore would set find's exit status and be reported as an unscanned + # part of the root. Pruning means those trees are never entered, so the + # exit status only reflects places this check actually wanted to read. + # + # That status matters: find exits non-zero when it cannot descend + # somewhere, having printed only what it could reach. Discarding it would + # hide every orphan under an unreadable directory behind a clean "ok", + # which is the defect this script exists to catch. + if ! find "$root" \ + \( -name .git -o -name elpa -o -name straight \ + -o -name node_modules -o -name .venv \) -prune \ + -o -name '*.example' -print > "$WORK/examples" 2>/dev/null; then + finding "could not fully scan $root — part of it was not checked" + fi + while IFS= read -r ex; do + [ -n "$ex" ] || continue + # -e, so a sibling that exists only as a dangling symlink counts as + # missing. It is not a config the machine can read. + [ -e "${ex%.example}" ] || finding "example without its real file: $ex" + done < "$WORK/examples" +done < "$WORK/roots" +report "check 3/8: local files" + +# --- 4. gitignore-mode projects missing their tooling --------------------- + +if [ -n "${PRC_PROJECT_ROOTS+set}" ]; then + projects=$PRC_PROJECT_ROOTS +else + projects=$(ls -d "$HOME"/code/*/ "$HOME"/projects/*/ 2>/dev/null; \ + printf '%s\n%s\n' "$HOME/.emacs.d" "$HOME/.dotfiles") +fi +printf '%s\n' "$projects" > "$WORK/projects" 2>/dev/null || { + echo "post-rebuild-check: cannot write $WORK/projects" >&2; exit 1; } +# CLAUDE.md is deliberately absent from this set. It is seed-only -- +# install-lang writes it once and the project owns it afterward -- so most +# projects legitimately never have one, and ratio shows the identical +# absences in the identical projects. That match is what proves it is the +# steady state rather than reinstall drift, and flagging it would put nine +# standing findings in front of every real one. +# +# .claude/ is absent for the same reason and proven the same way. The +# bootstrap and the gitignore sweep write it into the ignore set of every +# gitignore-mode project whether or not one ever exists there, so the entry is +# aspirational rather than a promise -- pearl, rsyncshot and yt-sync each name +# it and none of the three has ever had one, on velox or on ratio. Dropping it +# loses no real signal either: a project that genuinely carries a .claude/ +# (rules and hooks from a language bundle) has it re-synced by +# sync-language-bundle.sh at every session start, so a true absence heals +# itself before this check would run. +# +# The list is fed to the inner loop straight from a heredoc rather than +# staged through a file. It is a constant, so a file bought nothing and cost +# a fifth unguarded write: had it failed (a full tmpfs, say) the inner loop +# would read nothing and every project would pass silently, which is the one +# outcome this script must never produce. The heredoc is the inner loop's own +# stdin and leaves the outer loop's redirect alone. +while IFS= read -r proj; do + [ -n "$proj" ] || continue + proj=${proj%/} + # -e not -d: in a worktree or submodule .git is a file naming the real + # gitdir, and a -d test would skip those projects silently. + [ -e "$proj/.git" ] || continue + [ -f "$proj/.gitignore" ] || continue + while read -r disk pattern; do + # Both the anchored (/.ai/) and unanchored (.ai/) ignore styles exist + # across the fleet; the sweep-gitignore audit hit exactly that split. + # + # grep exits 1 for no-match and 2 for an error, so the two are told + # apart rather than both read as "the ignore file does not name this". + # An unreadable .gitignore would otherwise pass the whole project. + grep -Eq "^/?${pattern}/?\$" "$proj/.gitignore" 2>/dev/null + case $? in + 0) [ -e "$proj/$disk" ] \ + || finding "$proj: .gitignore names $disk but it is missing on disk" ;; + 1) ;; + *) finding "$proj: could not read .gitignore — the project was not checked" + break ;; + esac + done <<'EOF' +.ai \.ai +todo.org todo\.org +inbox inbox +EOF +done < "$WORK/projects" +report "check 4/8: project tooling" + +# --- 5. signal-cli registration ------------------------------------------- + +if [ -n "${PRC_SIGNAL_ACCOUNTS+set}" ]; then + accounts=$PRC_SIGNAL_ACCOUNTS + if [ "$accounts" = "MISSING" ]; then + accounts="" + signal_missing=1 + fi +else + if command -v signal-cli >/dev/null 2>&1; then + if ! accounts=$(signal-cli listAccounts 2>/dev/null); then + accounts="" + finding "signal-cli listAccounts failed — the registration was not checked" + signal_missing=skip + fi + else + accounts="" + signal_missing=1 + fi +fi +if [ "$signal_missing" = 1 ]; then + finding "signal-cli is not installed — paging relies on it fleet-wide" +elif [ -z "$signal_missing" ] && [ -z "$accounts" ]; then + finding "no signal account registered — this machine can only page by relaying to one that has an account; if no host in AGENT_TEXT_RELAYS is registered either, the whole fleet loses paging" +fi +report "check 5/8: signal registration" + +# --- 6. NTP can recover a wrong clock without DNS ------------------------- +# +# The clock/DNS bootstrap deadlock. This machine resolves through DNSOverTLS +# with DNSSEC, and both validate against the wall clock, so a boot with a +# wrong clock resolves nothing at all. If every configured NTP source is named +# by hostname, the daemon that would correct the clock needs the DNS the clock +# is breaking, and the machine cannot recover without a second device -- +# which is exactly what happened on velox 2026-08-19. One source addressed by +# IP breaks the cycle, so that is what this check looks for. + +# True when the argument is an address rather than a name. An address needs no +# resolver, which is the whole property being checked. +is_ip_literal() { + case "$1" in + "") return 1 ;; + *:*) case "$1" in *[!0-9A-Fa-f:]*) return 1 ;; esac + return 0 ;; + *[!0-9.]*) return 1 ;; + *.*) return 0 ;; + esac + return 1 +} + +if [ -n "${PRC_NTP_SOURCES+set}" ]; then + ntp_sources=$PRC_NTP_SOURCES + if [ "$ntp_sources" = "MISSING" ]; then + ntp_sources="" + ntp_missing=1 + fi +elif sctl is-active chronyd >/dev/null 2>&1; then + # The main file, plus any drop-in directory chrony.conf actually names. + # + # The confdir read is the load-bearing part. A drop-in is inert unless + # chrony.conf points at its directory, and Arch's stock chrony.conf points + # at none -- so globbing /etc/chrony.d unconditionally would find the + # IP-addressed source, report the machine healthy, and be describing a file + # chrony never opens. That is a false pass on exactly the misconfiguration + # this check exists to catch, so the sources are read only from files + # chrony is actually told to read. + ntp_conf_files=$CHRONY_CONF + for ntp_dir in $(awk '$1 == "confdir" || $1 == "sourcedir" { print $2 }' \ + "$CHRONY_CONF" 2>/dev/null); do + for ntp_f in "$ntp_dir"/*.conf "$ntp_dir"/*.sources; do + [ -f "$ntp_f" ] && ntp_conf_files="$ntp_conf_files $ntp_f" + done + done + # Unquoted on purpose: the accumulated list is several paths, and none of + # this script's own paths contain spaces. + ntp_sources=$(cat $ntp_conf_files 2>/dev/null \ + | awk '$1 == "server" || $1 == "pool" { print $2 }') +elif sctl is-active systemd-timesyncd >/dev/null 2>&1; then + ntp_sources=$(awk -F= '/^[[:space:]]*NTP=/ { print $2 }' \ + /etc/systemd/timesyncd.conf 2>/dev/null | tr ' ' '\n') +else + ntp_sources="" + ntp_missing=1 +fi + +if [ "$ntp_missing" = 1 ]; then + finding "no NTP implementation is active — nothing corrects the clock, and a wrong clock takes DNS down with it" +elif [ -z "$ntp_sources" ]; then + finding "no NTP sources are configured — nothing was checked, and nothing corrects the clock" +else + ntp_has_literal="" + stage "$ntp_sources" + while IFS= read -r src_addr; do + [ -z "$src_addr" ] && continue + if is_ip_literal "$src_addr"; then + ntp_has_literal=1 + fi + done < "$STAGE" + if [ -z "$ntp_has_literal" ]; then + finding "every NTP source is named by hostname — a wrong clock breaks DNS, so nothing can resolve them and the clock stays wrong" + fi +fi +report "check 6/8: NTP bootstrap" + +# --- 7. the idle daemon survives session start ---------------------------- +# +# A laptop that never sleeps has no symptom until the battery is gone, so +# nothing surfaces this without being asked. On velox 2026-08-19 hypridle +# started cleanly at 15:29:48 and `settings restore` killed it six seconds +# later, replaying a caffeine stored in an earlier boot. The machine ran +# 11h40m fully awake on battery, died when it flattened, and reset its RTC -- +# which took DNS down with it, the same deadlock check 6 exists for. The +# desktop looked correct throughout. +# +# Behavioural on purpose: this asks whether the daemon is alive, not why it +# might not be, so a stale caffeine, a crash, and a broken config all surface +# the same way. Gated on hypridle being installed, because that is what marks +# a machine as using it -- archsetup installs it only for Hyprland, so a +# headless or dwm box would otherwise report a finding on every run. + +if [ -n "${PRC_IDLE_DAEMON+set}" ]; then + idle_pids=$PRC_IDLE_DAEMON + if [ "$idle_pids" = "MISSING" ]; then + idle_pids="" + idle_absent=1 + fi +elif command -v hypridle >/dev/null 2>&1; then + # pgrep exits non-zero with no match, which is the not-running case rather + # than a probe failure, so the || keeps `set -e`-style callers out of it. + idle_pids=$(pgrep -x hypridle 2>/dev/null) || idle_pids="" +else + idle_pids="" + idle_absent=1 +fi + +if [ "$idle_absent" = 1 ]; then + : # hypridle is not part of this machine -- nothing to check +elif [ -z "$idle_pids" ]; then + finding "hypridle is installed but not running — nothing triggers idle lock or suspend, so this machine stays awake until its battery is gone; a caffeine remembered from an earlier boot is the known cause" +fi +report "check 7/8: idle daemon" + +# --- 8. working repos cloned from the read-only endpoint ------------------ +# +# archsetup clones the user's own archsetup and dotfiles from +# https://git.cjennings.net/..., which serves anonymous clones and refuses +# pushes. That default is correct for a stranger installing archsetup -- they +# have no key on the server -- and wrong for this machine, which has to push. +# ARCHSETUP_REPO / DOTFILES_REPO override it, but only where they are +# configured: a curl|bash install, or a rebuild from a stock ISO, takes the +# default straight back. +# +# Nothing about the tree shows it. The clone is complete and ordinary, and the +# machine finds out at the first push, with a 403 -- which is how velox's +# dotfiles remote was found on 2026-08-17, four days after its rebuild, by +# which time the same rebuild's shallow clone had already answered a +# credential-history question wrongly. +# +# Only the read-only endpoint is flagged. An https remote elsewhere may be +# perfectly pushable through a credential helper, and guessing about hosts +# this machine does not own would stand noise in front of the real findings. + +if [ -n "${PRC_REPO_REMOTES+set}" ]; then + repo_remotes=$PRC_REPO_REMOTES +else + repo_remotes="" + for repo in "$HOME/code/archsetup" "$HOME/.dotfiles"; do + # -e not -d: a worktree or submodule .git is a file naming the gitdir. + [ -e "$repo/.git" ] || continue + # A repo with no origin still gets a line, with an empty URL, so the + # loop below reports it rather than skipping it into a silent pass. + repo_url=$(git -C "$repo" remote get-url origin 2>/dev/null) + repo_remotes="${repo_remotes}${repo} ${repo_url} +" + done +fi + +stage "$repo_remotes" +while IFS= read -r repo_line; do + [ -n "$repo_line" ] || continue + repo_path=${repo_line%% *} + repo_url=${repo_line#"$repo_path"} + repo_url=${repo_url# } + case "$repo_url" in + "") + finding "$repo_path: origin could not be read — the remote was not checked" ;; + https://git.cjennings.net/*|https://cjennings.net/*) + finding "$repo_path: origin is the read-only endpoint ($repo_url) — git push returns 403; set the ssh form, or ARCHSETUP_REPO/DOTFILES_REPO before installing" ;; + esac +done < "$STAGE" +report "check 8/8: repo remotes" + +# --- summary -------------------------------------------------------------- + +if [ "$TOTAL_FINDINGS" -eq 0 ]; then + echo "all checks clean" + exit 0 +fi +echo "$TOTAL_FINDINGS finding(s) across 8 checks" +exit 1 diff --git a/scripts/testing/tests/test_config_applied.py b/scripts/testing/tests/test_config_applied.py index 00c410e..08ffc1b 100644 --- a/scripts/testing/tests/test_config_applied.py +++ b/scripts/testing/tests/test_config_applied.py @@ -40,7 +40,8 @@ def test_makepkg_options_trimmed(host): @pytest.mark.attribution("archsetup") -@pytest.mark.parametrize("rel", ["dns.conf", "wifi-privacy.conf"]) +@pytest.mark.parametrize("rel", ["dns.conf", "wifi-privacy.conf", + "tunnel-dns-over-tls.conf"]) def test_networkmanager_dropin(host, rel): assert host.file("/etc/NetworkManager/conf.d/%s" % rel).exists diff --git a/scripts/zz-bluetooth-resume b/scripts/zz-bluetooth-resume new file mode 100755 index 0000000..4273339 --- /dev/null +++ b/scripts/zz-bluetooth-resume @@ -0,0 +1,86 @@ +#!/bin/sh +# SPDX-License-Identifier: GPL-3.0-or-later +# zz-bluetooth-resume - put bluetooth back after a sleep cycle. +# +# A systemd-sleep hook. Two things break bluetooth across sleep on a TLP +# laptop, and nothing else on the machine fixes either one. +# +# 1. The rfkill soft-block is not restored. systemd-rfkill would do it, and +# it is masked here deliberately -- it fights TLP's radio handling, so +# configure_tlp_power masks it and TLP owns radios instead. TLP's own +# sleep hook runs `tlp resume`, but its setting is +# DEVICES_TO_ENABLE_ON_STARTUP: startup, not resume. TLP has no ON_RESUME +# at all, so the resume edge has no owner. WiFi survives only because +# NetworkManager unblocks itself; bluetooth has no equivalent. +# +# 2. The controller comes back wedged from a hibernate. It reports powered +# and unblocked while scanning finds nothing whatever -- zero devices +# where the same room gave seventeen a minute later -- and bluetoothd +# logs "Failed to set mode" and "Failed to add device <mac>" at the +# instant of resume. Reloading btusb clears it. +# +# Both observed on velox 2026-08-21, on the first suspend-then-hibernate cycle +# after hibernate was switched back on. The second symptom is why unblocking +# alone is not enough: rfkill was cleared by hand and scanning still returned +# nothing until the driver was reloaded. +# +# The hook re-asserts TLP's own declared intent rather than inventing a policy. +# A machine whose TLP config does not ask for bluetooth keeps it off, which is +# what stops this from overriding a deliberate block at every wakeup. +# +# The zz- prefix orders it after TLP's own hook, so `tlp resume` has finished +# before this runs. +# +# Test seams: BTR_RFKILL, BTR_MODPROBE, BTR_TLP_CONF, BTR_TLP_CONF_DIR, +# BTR_SETTLE (seconds to wait between driver unload and load). + +set -u + +RFKILL="${BTR_RFKILL:-rfkill}" +MODPROBE="${BTR_MODPROBE:-modprobe}" +TLP_CONF="${BTR_TLP_CONF:-/etc/tlp.conf}" +TLP_CONF_DIR="${BTR_TLP_CONF_DIR:-/etc/tlp.d}" +SETTLE="${BTR_SETTLE:-1}" + +# post only. The pre phase has nothing to do, and acting there would fight the +# suspend it is about to run. +[ "${1:-}" = "post" ] || exit 0 + +# Does TLP ask for bluetooth on this machine? Comments are stripped first, so a +# commented-out example in the stock config cannot be read as a policy. Both +# the main file and any drop-in count, and the last assignment wins the same +# way TLP itself resolves them. +wants_bluetooth() { + cat "$TLP_CONF" "$TLP_CONF_DIR"/*.conf 2>/dev/null \ + | sed 's/#.*//' \ + | awk -F= '/DEVICES_TO_ENABLE_ON_STARTUP/ { v = $2 } END { print v }' \ + | tr -d '"' \ + | tr ' ' '\n' \ + | grep -qx "bluetooth" +} + +wants_bluetooth || exit 0 + +# The wedge follows a hibernate, which reinitialises the controller from a +# saved image. A plain suspend brings USB back intact, so reloading there would +# tear down a working adapter for nothing. +# +# suspend-then-hibernate reports that name whether or not it reached the +# hibernate stage, so this reloads on a cycle that only suspended. That is the +# cheap side of the trade: a couple of seconds against an adapter that answers +# nothing until someone notices and reloads it by hand. +case "${2:-}" in + hibernate|suspend-then-hibernate) + "$MODPROBE" -r btusb 2>/dev/null || true + [ "$SETTLE" = "0" ] || sleep "$SETTLE" + "$MODPROBE" btusb 2>/dev/null || true + ;; +esac + +# After the reload, not before: a freshly loaded btusb can come up soft-blocked +# and would undo an earlier unblock. +"$RFKILL" unblock bluetooth 2>/dev/null || true + +# Never fail. systemd-sleep logs a failing hook, and that noise outlives the +# cause it describes; nothing here is worth alarming a resume over. +exit 0 diff --git a/tests/bluetooth-resume/test_bluetooth_resume.py b/tests/bluetooth-resume/test_bluetooth_resume.py new file mode 100644 index 0000000..6d8ed87 --- /dev/null +++ b/tests/bluetooth-resume/test_bluetooth_resume.py @@ -0,0 +1,139 @@ +"""Tests for scripts/zz-bluetooth-resume. + +Two things break bluetooth across a sleep cycle on a TLP laptop, and nothing +else on the machine fixes either. + +The rfkill soft-block is not restored. systemd-rfkill would do it, but it is +masked deliberately -- it fights TLP's radio handling, so TLP owns radios +instead. TLP's own sleep hook runs `tlp resume`, and its setting is +DEVICES_TO_ENABLE_ON_STARTUP: startup, not resume. There is no ON_RESUME in +TLP's vocabulary, so the resume edge has no owner at all. WiFi survives only +because NetworkManager unblocks itself; bluetooth has no equivalent. + +The controller also comes back wedged from a hibernate. It reports powered and +unblocked while scanning finds nothing whatever -- zero devices where the same +room gave seventeen a minute later. bluetoothd logs "Failed to set mode" and +"Failed to add device <mac>" at the instant of resume. Reloading btusb clears +it. + +Both observed on velox 2026-08-21, on its first suspend-then-hibernate cycle +after hibernate was switched back on. + +The hook re-asserts TLP's own declared intent rather than inventing a policy, +so a machine that deliberately blocks bluetooth keeps it blocked. + +Run from repo root: + python3 -m unittest tests.bluetooth-resume.test_bluetooth_resume +""" + +import os +import stat +import subprocess +import tempfile +import unittest + +REPO_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")) +HOOK = os.path.join(REPO_ROOT, "scripts", "zz-bluetooth-resume") + +TLP_WANTS_BT = 'DEVICES_TO_ENABLE_ON_STARTUP="bluetooth wifi"\n' +TLP_WIFI_ONLY = 'DEVICES_TO_ENABLE_ON_STARTUP="wifi"\n' + + +def run(phase="post", kind="suspend-then-hibernate", tlp_conf=TLP_WANTS_BT, + conf_present=True): + """Drive the hook with rfkill and modprobe faked, and read back the calls.""" + with tempfile.TemporaryDirectory() as d: + calls = os.path.join(d, "calls.log") + bindir = os.path.join(d, "bin") + os.makedirs(bindir) + for tool in ("rfkill", "modprobe"): + p = os.path.join(bindir, tool) + with open(p, "w") as fh: + fh.write(f'#!/bin/sh\necho "{tool} $*" >> "{calls}"\nexit 0\n') + os.chmod(p, 0o755) + conf = os.path.join(d, "tlp.conf") + if conf_present: + with open(conf, "w") as fh: + fh.write(tlp_conf) + env = dict(os.environ) + env.update({ + "BTR_RFKILL": os.path.join(bindir, "rfkill"), + "BTR_MODPROBE": os.path.join(bindir, "modprobe"), + "BTR_TLP_CONF": conf, + "BTR_TLP_CONF_DIR": os.path.join(d, "tlp.d"), + "BTR_SETTLE": "0", + }) + r = subprocess.run(["sh", HOOK, phase, kind], env=env, + capture_output=True, text=True, timeout=20) + log = "" + if os.path.exists(calls): + with open(calls) as fh: + log = fh.read() + return r, log + + +class BluetoothResume(unittest.TestCase): + # --- Normal --------------------------------------------------------- + def test_hibernate_reloads_the_driver_and_unblocks(self): + _, log = run(kind="suspend-then-hibernate") + self.assertIn("modprobe -r btusb", log) + self.assertIn("modprobe btusb", log) + self.assertIn("rfkill unblock bluetooth", log) + + def test_the_unblock_comes_after_the_reload(self): + # A freshly loaded btusb can come up soft-blocked, so unblocking first + # would be undone by the reload that follows it. + _, log = run() + self.assertLess(log.index("modprobe btusb"), + log.index("rfkill unblock")) + + def test_plain_suspend_unblocks_without_reloading(self): + # The wedge was seen coming out of hibernate, which reinitialises the + # controller from a saved image. A plain suspend restores USB intact, + # so reloading there would cost a working adapter for nothing. + _, log = run(kind="suspend") + self.assertIn("rfkill unblock bluetooth", log) + self.assertNotIn("btusb", log) + + # --- Boundary ------------------------------------------------------- + def test_the_pre_phase_does_nothing(self): + _, log = run(phase="pre") + self.assertEqual(log, "") + + def test_a_tlp_policy_without_bluetooth_is_left_alone(self): + # The hook re-asserts TLP's stated intent. It must not invent one, or + # a machine that deliberately keeps bluetooth off gets it turned on at + # every wakeup. + _, log = run(tlp_conf=TLP_WIFI_ONLY) + self.assertEqual(log, "") + + def test_a_commented_out_policy_does_not_count(self): + _, log = run(tlp_conf='#DEVICES_TO_ENABLE_ON_STARTUP="bluetooth"\n') + self.assertEqual(log, "") + + def test_hibernate_proper_also_reloads(self): + _, log = run(kind="hibernate") + self.assertIn("modprobe -r btusb", log) + + # --- Error ---------------------------------------------------------- + def test_a_missing_tlp_config_is_left_alone(self): + # No declared policy means no intent to re-assert. Failing safe here + # means doing nothing, not guessing. + _, log = run(conf_present=False) + self.assertEqual(log, "") + + def test_the_hook_always_exits_zero(self): + # systemd-sleep logs a failing hook and the noise outlives the cause. + # Nothing here is worth delaying or alarming a resume over. + for kind in ("suspend", "hibernate", "suspend-then-hibernate"): + with self.subTest(kind=kind): + r, _ = run(kind=kind) + self.assertEqual(r.returncode, 0, r.stderr) + + def test_it_is_executable(self): + self.assertTrue(os.stat(HOOK).st_mode & stat.S_IXUSR, + "systemd-sleep only runs executables") + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/installer-steps/test_clone_user_repos.py b/tests/installer-steps/test_clone_user_repos.py new file mode 100644 index 0000000..51d8434 --- /dev/null +++ b/tests/installer-steps/test_clone_user_repos.py @@ -0,0 +1,155 @@ +"""Test clone_user_repos: the two user repos are cloned with full history. + +archsetup and dotfiles are not build directories. They are the two repos I +actively develop in on every machine this installer builds, so a shallow clone +is wrong for both. Velox came back from its 2026-08-13 rebuild with 7 commits +of history in each instead of 851, and nothing about the tree said so. + +The quiet failure is what makes this worth a test rather than a one-line fix. +`git log -- <path>` against a shallow clone does not error; it answers "no +commits". So a credential-history check run on that machine reported five +sensitive files absent from history and exited clean, when the real answer was +that the clone could not see the history they live in. A security question came +back falsely reassuring. Everything else it breaks — blame, bisect, any +archaeology past the graft point — is merely annoying by comparison. + +The AUR build clones are a different case and stay shallow: they are throwaway +build trees, cloned to run `make install` and then discarded, where history has +no value and the download cost is real. So this suite asserts both halves — +full history for the two user repos, and depth still pinned on the AUR path — +because a fix applied with too broad a brush would regress the build clones +without failing any test that only looked at the user repos. + +Method: sed-extract clone_user_repos from the real `archsetup`, fake git / +mkdir / chown / display / error_warn / error_fatal, and read back the git +command lines the function issued. + +Run from repo root: + python3 -m unittest tests.installer-steps.test_clone_user_repos +""" + +import os +import re +import subprocess +import tempfile +import textwrap +import unittest + +REPO_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")) +ARCHSETUP = os.path.join(REPO_ROOT, "archsetup") + + +def run(clone_fails=False, make_git_dir=True): + """Drive clone_user_repos with every side effect faked. + + dotfiles_dir is pre-created with a .git so the function's "is this a real + checkout?" guard passes on the happy path; make_git_dir=False exercises the + guard itself. + """ + with tempfile.TemporaryDirectory() as d: + dotfiles_dir = os.path.join(d, "dotfiles") + os.makedirs(dotfiles_dir) + if make_git_dir: + os.makedirs(os.path.join(dotfiles_dir, ".git")) + clone_rc = 1 if clone_fails else 0 + script = textwrap.dedent(f"""\ + logfile=/dev/null + action="" + username=testuser + archsetup_repo="https://example.invalid/archsetup.git" + dotfiles_repo="https://example.invalid/dotfiles.git" + dotfiles_branch=main + dotfiles_dir="{dotfiles_dir}" + display() {{ :; }} + mkdir() {{ echo "MKDIR: $*" >> "{d}/calls.log"; return 0; }} + chown() {{ echo "CHOWN: $*" >> "{d}/calls.log"; return 0; }} + git() {{ + echo "GIT: $*" >> "{d}/calls.log" + case "$1" in + clone) return {clone_rc} ;; + *) return 0 ;; + esac + }} + error_warn() {{ echo "WARN: $1" >> "{d}/calls.log"; return 1; }} + error_fatal() {{ echo "FATAL: $1" >> "{d}/calls.log"; exit 1; }} + source <(sed -n '/^clone_user_repos() {{/,/^}}/p' "{ARCHSETUP}") + clone_user_repos + echo "RC=$?" >> "{d}/calls.log" + exit 0 + """) + subprocess.run( + ["bash", "-c", script], capture_output=True, text=True, timeout=10, + ) + with open(os.path.join(d, "calls.log")) as fh: + return fh.read() + + +def clone_lines(log): + return [ln for ln in log.splitlines() if ln.startswith("GIT: clone")] + + +class CloneUserRepos(unittest.TestCase): + # ------------------------------------------------------------ normal ---- + def test_both_user_repos_are_cloned(self): + lines = clone_lines(run()) + self.assertEqual(len(lines), 2, + f"expected an archsetup clone and a dotfiles clone, got: {lines}") + self.assertTrue(any("archsetup.git" in ln for ln in lines)) + self.assertTrue(any("dotfiles.git" in ln for ln in lines)) + + def test_archsetup_clone_carries_full_history(self): + """A shallow archsetup clone answers history questions wrongly.""" + line = next(ln for ln in clone_lines(run()) if "archsetup.git" in ln) + self.assertNotIn("--depth", line, + "archsetup is a working repo, not a build tree — a shallow " + "clone makes `git log -- <path>` answer 'no commits' instead " + "of failing, which is how a credential-history check came " + "back falsely clean on velox") + + def test_dotfiles_clone_carries_full_history(self): + line = next(ln for ln in clone_lines(run()) if "dotfiles.git" in ln) + self.assertNotIn("--depth", line, + "dotfiles is a working repo, not a build tree") + + def test_dotfiles_clone_still_pins_the_branch(self): + """Dropping --depth must not disturb the --branch argument beside it.""" + line = next(ln for ln in clone_lines(run()) if "dotfiles.git" in ln) + self.assertIn("--branch main", line) + + # ---------------------------------------------------------- boundary ---- + def test_no_user_repo_clone_is_shallow_by_any_spelling(self): + """--depth, --depth=N and -depth are all shallow; catch the lot.""" + for line in clone_lines(run()): + self.assertNotRegex(line, r"(^|\s)-{1,2}depth(\s|=)", + f"user-repo clone must be full: {line}") + + def test_aur_build_clones_stay_shallow(self): + """The fix must not over-apply — build trees are throwaway. + + Read against the real file rather than the extracted function, because + these clones live in a different function entirely and the risk being + guarded is a careless repo-wide sed. + """ + with open(ARCHSETUP) as fh: + source = fh.read() + build_clones = re.findall(r"^.*git clone.*build_dir.*$", source, re.M) + self.assertTrue(build_clones, "expected AUR build clones to exist") + for line in build_clones: + self.assertIn("--depth 1", line, + f"AUR build clone should stay shallow: {line.strip()}") + + # ------------------------------------------------------------- error ---- + def test_clone_failure_is_reported_not_swallowed(self): + log = run(clone_fails=True) + self.assertIn("WARN:", log, + "a failed clone must surface through error_warn") + + def test_dotfiles_clone_producing_no_checkout_is_fatal(self): + """The stow/restore steps downstream need a real checkout.""" + log = run(make_git_dir=False) + self.assertIn("FATAL:", log) + self.assertNotIn("RC=", log, "error_fatal must halt, not fall through") + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/installer-steps/test_configure_service_discovery.py b/tests/installer-steps/test_configure_service_discovery.py new file mode 100644 index 0000000..90ce041 --- /dev/null +++ b/tests/installer-steps/test_configure_service_discovery.py @@ -0,0 +1,85 @@ +"""Pin configure_service_discovery's source: the WS-Discovery host daemon +is never enabled. + +wsdd.service advertises this machine as a Samba host to Windows clients. +Nothing the installer sets up runs Samba, so enabled it advertised a share +server that doesn't exist while listening on every interface, VPN and +tailscale links included (2026-09-12). Browsing Windows shares is the other +direction: gvfs-wsdd spawns its own wsdd in discovery mode and needs only +the package. + +Method: the step writes straight to /etc (geoclue.conf, the dbus-broker +drop-in) with no directory parameter, so unlike the other step tests this +one can't run the function in a temp dir. It sed-extracts the source and +asserts on the calls it contains. + + python3 -m unittest tests.installer-steps.test_configure_service_discovery +""" + +import os +import re +import unittest + +REPO_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")) +ARCHSETUP = os.path.join(REPO_ROOT, "archsetup") + + +def function_source(name): + with open(ARCHSETUP) as f: + text = f.read() + m = re.search(r"^%s\(\) \{\n.*?^\}\n" % re.escape(name), text, re.S | re.M) + assert m, "function %s not found in archsetup" % name + return m.group(0) + + +ENABLE_WSDD = r"(systemctl\s+enable|enable_service)\s+(--now\s+)?wsdd\b" + + +def calls(src): + """Non-comment lines — what the function actually runs.""" + return [l for l in src.splitlines() if l.strip() and not l.strip().startswith("#")] + + +class ConfigureServiceDiscovery(unittest.TestCase): + # ------------------------------------------------------------ normal ---- + def test_does_not_enable_the_wsdd_host_daemon(self): + # Both spellings the script uses: a raw systemctl call and the + # enable_service helper (which takes the bare unit name). + body = "\n".join(calls(function_source("configure_service_discovery"))) + self.assertNotRegex(body, ENABLE_WSDD) + + def test_turns_off_a_previously_enabled_wsdd_on_rerun(self): + # Machines installed before this change still have the unit + # enabled; a re-run converges them. The guard keeps a fresh install + # (no unit yet) quiet. + body = "\n".join(calls(function_source("configure_service_discovery"))) + self.assertRegex(body, r"systemctl\s+is-enabled\s+(--quiet\s+)?wsdd\.service") + self.assertRegex(body, r"systemctl\s+disable\s+--now\s+wsdd\.service") + + def test_still_enables_the_discovery_it_does_want(self): + # Characterization: removing wsdd must not have taken avahi (mDNS) + # or geoclue with it. + body = "\n".join(calls(function_source("configure_service_discovery"))) + self.assertIn("systemctl enable avahi-daemon.service", body) + self.assertIn("systemctl enable geoclue.service", body) + + # ---------------------------------------------------------- boundary ---- + def test_wsdd_package_still_installed_for_gvfs(self): + # The package stays: gvfs-wsdd depends on it and spawns its own + # discovery-mode instance. Only the host service is gone. + body = "\n".join(calls(function_source("supplemental_software"))) + self.assertRegex(body, r"pacman_install\s+wsdd\b") + self.assertRegex(body, r"pacman_install\s+gvfs-wsdd\b") + + # ------------------------------------------------------------- error ---- + def test_no_other_step_enables_wsdd_either(self): + # A regression that re-enables it from a different step is the same + # bug; scan the whole script, not just the one function. + with open(ARCHSETUP) as f: + lines = [l for l in f if l.strip() and not l.strip().startswith("#")] + offenders = [l.rstrip() for l in lines if re.search(ENABLE_WSDD, l)] + self.assertEqual(offenders, []) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/installer-steps/test_configure_tlp_power.py b/tests/installer-steps/test_configure_tlp_power.py index c88e0c2..1ddff72 100644 --- a/tests/installer-steps/test_configure_tlp_power.py +++ b/tests/installer-steps/test_configure_tlp_power.py @@ -1,4 +1,4 @@ -"""Test configure_tlp_power's radio-enable line and laptop gating. +"""Test configure_tlp_power's radio-enable line, daemon masking, and laptop gating. systemd-rfkill is masked on laptops because it fights TLP's radio handling — which means nothing restores radio state at boot unless TLP is told to. The @@ -6,6 +6,20 @@ velox 2026-04-10 setup found wifi and bluetooth soft-blocked on first boot for exactly this reason. The conf written here must carry DEVICES_TO_ENABLE_ON_STARTUP so a fresh install comes up with radios on. +power-profiles-daemon is masked and stopped on laptops for the same class of +reason. power-profiles-daemon.service declares "Conflicts=tuned.service +tlp.service auto-cpufreq.service ..." — the line is in ppd's unit, not tlp's — +so systemd TERMs TLP the instant ppd starts. Leaving ppd merely disabled does +not prevent that: ppd ships D-Bus activation files, and the desktop-settings +panel's own powerprofilesctl call activates it on demand. Velox ran that way +from its 2026-08-13 rebuild until 2026-08-16, with TLP failing at every boot and +none of its battery policy applied, while the machine looked correctly +configured. Masking blocks D-Bus activation too, which both keeps TLP alive and +makes the panel's power control read as unavailable, the behavior the +package-install site in `archsetup` already documents as intended. The stop is +what makes a repair re-run take effect on a booted machine, where a mask alone +would leave a running ppd running. + Method: sed-extract configure_tlp_power from the real `archsetup`, point it at a temp tlp.d dir and a temp power-supply dir, and fake pacman_install / run_task / display / error_warn / systemctl. @@ -25,7 +39,7 @@ REPO_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")) ARCHSETUP = os.path.join(REPO_ROOT, "archsetup") -def run(battery=True, bat_name="BAT0", unwritable_tlpd=False): +def run(battery=True, bat_name="BAT0", unwritable_tlpd=False, systemctl_fails=False): with tempfile.TemporaryDirectory() as d: psdir = os.path.join(d, "power_supply") os.makedirs(psdir) @@ -37,13 +51,14 @@ def run(battery=True, bat_name="BAT0", unwritable_tlpd=False): os.chmod(tlpd, stat.S_IRUSR | stat.S_IXUSR) # The real mask call redirects stdout into $logfile, so the fake # systemctl records to a side file the test reads back instead. + sysrc = 1 if systemctl_fails else 0 script = textwrap.dedent(f"""\ logfile=/dev/null action="" display() {{ :; }} pacman_install() {{ echo "INSTALL: $1"; }} run_task() {{ echo "TASK: $1"; }} - systemctl() {{ echo "SYSTEMCTL: $*" >> "{d}/systemctl.log"; }} + systemctl() {{ echo "SYSTEMCTL: $*" >> "{d}/systemctl.log"; return {sysrc}; }} error_warn() {{ echo "WARN: $1"; return 1; }} source <(sed -n '/^configure_tlp_power() {{/,/^}}/p' "{ARCHSETUP}") configure_tlp_power "{tlpd}" "{psdir}" @@ -74,6 +89,44 @@ class ConfigureTlpPower(unittest.TestCase): r.stdout) self.assertIn("TASK: enabling TLP service", r.stdout) + def test_laptop_masks_power_profiles_daemon(self): + r = run(battery=True) + self.assertIn("SYSTEMCTL: mask power-profiles-daemon.service", r.stdout, + "ppd's unit declares Conflicts=...tlp.service..., so ppd " + "must be masked or it TERMs TLP whenever it is activated") + + def test_laptop_stops_running_power_profiles_daemon(self): + """Masking alone leaves an already-running ppd running. + + The installer runs on a booted system, so a repair re-run would + otherwise mask ppd, leave it live, and let it keep TLP dead until the + next reboot with nothing reporting it. + """ + r = run(battery=True) + self.assertIn("SYSTEMCTL: stop power-profiles-daemon.service", r.stdout) + + def test_ppd_is_masked_before_it_is_stopped(self): + """Order matters: stopping first leaves a window to re-activate in.""" + calls = [line for line in run(battery=True).stdout.splitlines() + if line.startswith("SYSTEMCTL:") and "power-profiles-daemon" in line] + verbs = [line.split()[1] for line in calls] + self.assertEqual(verbs, ["mask", "stop"]) + + def test_power_profiles_daemon_is_masked_not_merely_disabled(self): + """Disabling ppd is not enough — D-Bus activation ignores it. + + This is the whole point of the mask, so assert the verb directly. A + `disable` here would pass a naive "ppd is handled" check while leaving + the panel's powerprofilesctl call free to start ppd and kill TLP. + """ + r = run(battery=True) + ppd_calls = [line for line in r.stdout.splitlines() + if line.startswith("SYSTEMCTL:") and "power-profiles-daemon" in line] + self.assertTrue(ppd_calls, "configure_tlp_power must act on ppd at all") + for line in ppd_calls: + self.assertNotIn(" disable ", line, + "disable leaves D-Bus activation live; only mask blocks it") + def test_radio_line_is_active_not_commented(self): r = run(battery=True) conf = r.stdout.split("CONF:[")[1].split("]")[0] @@ -97,7 +150,35 @@ class ConfigureTlpPower(unittest.TestCase): self.assertIn("INSTALL: tlp", r.stdout) self.assertIn('DEVICES_TO_ENABLE_ON_STARTUP', r.stdout) + def test_desktop_keeps_power_profiles_daemon(self): + """A batteryless machine must NOT get ppd masked. + + There is no TLP on a desktop to conflict with it, and the package-install + site enables ppd precisely so the settings panel's three-way power + control works there. Masking it here would break that control for no gain. + """ + r = run(battery=False) + self.assertNotIn("power-profiles-daemon", r.stdout) + # ------------------------------------------------------------- error ---- + def test_failed_ppd_mask_warns_and_does_not_crash(self): + """A masking failure must surface, not pass silently. + + Silence is the exact failure mode being fixed: velox looked configured + while TLP was dead. If the mask cannot be applied, say so. + + Assert on the harness's own RC= line, not on r.returncode. The harness + script ends in a literal `exit 0`, so r.returncode is 0 no matter what + configure_tlp_power does — asserting it can never fail, which would make + this test the same silent no-op it exists to catch. + """ + r = run(battery=True, systemctl_fails=True) + self.assertIn("WARN: masking power-profiles-daemon for TLP", r.stdout) + self.assertIn("WARN: stopping power-profiles-daemon for TLP", r.stdout) + self.assertIn("RC=", r.stdout, + "the function must return so the install continues, " + "not exit and take the script down with it") + @unittest.skipUnless(os.geteuid() != 0, "root ignores directory write bits") def test_unwritable_tlpd_warns_and_does_not_crash(self): r = run(battery=True, unwritable_tlpd=True) diff --git a/tests/installer-steps/test_configure_tunnel_dns_over_tls.py b/tests/installer-steps/test_configure_tunnel_dns_over_tls.py new file mode 100644 index 0000000..89677d8 --- /dev/null +++ b/tests/installer-steps/test_configure_tunnel_dns_over_tls.py @@ -0,0 +1,147 @@ +"""Test configure_tunnel_dns_over_tls — per-link DoT off for VPN tunnels. + +The resolved drop-in pins DNSOverTLS=yes globally. Proton VPN (proton0) and +the static Proton WireGuard profiles (wgpvpn) push an in-tunnel resolver, +10.2.0.1, that answers plain port 53 and never completes TLS on 853, so +every lookup through the tunnel hangs. The Proton client recreates its NM +profile on each connect, so a per-profile setting can't stick; a +NetworkManager [connection-*] default matched on the interface names is +what NM pushes to resolved on every activation. Diagnosed 2026-09-09/10, +verified live on ratio and velox. + +Method: sed-extract configure_tunnel_dns_over_tls from the real +`archsetup`, point it at a temp conf.d, and assert on the file it writes. + + python3 -m unittest tests.installer-steps.test_configure_tunnel_dns_over_tls +""" + +import os +import re +import stat +import subprocess +import tempfile +import textwrap +import unittest + +REPO_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")) +ARCHSETUP = os.path.join(REPO_ROOT, "archsetup") +FILENAME = "tunnel-dns-over-tls.conf" + + +def run(confdir): + script = textwrap.dedent(f"""\ + logfile=/dev/null + action="" + display() {{ :; }} + error_warn() {{ echo "WARN: $1"; return 1; }} + source <(sed -n '/^configure_tunnel_dns_over_tls() {{/,/^}}/p' "{ARCHSETUP}") + configure_tunnel_dns_over_tls "{confdir}" + echo "RC=$?" + exit 0 + """) + return subprocess.run( + ["bash", "-c", script], capture_output=True, text=True, timeout=10, + ) + + +def rc_of(r): + m = re.search(r"^RC=(\d+)$", r.stdout, re.M) + assert m, "no RC line in output: %r / %r" % (r.stdout, r.stderr) + return int(m.group(1)) + + +def body(confdir): + with open(os.path.join(confdir, FILENAME)) as f: + return f.read() + + +def directives(text): + """The non-comment, non-blank lines — what NM actually parses.""" + return [l.strip() for l in text.splitlines() + if l.strip() and not l.lstrip().startswith("#")] + + +class ConfigureTunnelDnsOverTls(unittest.TestCase): + # ------------------------------------------------------------ normal ---- + def test_writes_a_connection_default_that_turns_dot_off(self): + with tempfile.TemporaryDirectory() as d: + r = run(d) + self.assertEqual(rc_of(r), 0) + lines = directives(body(d)) + self.assertIn("[connection-tunnel-dot]", lines) + self.assertIn("connection.dns-over-tls=0", lines) + + def test_matches_both_proton_interfaces_and_nothing_else(self): + # proton0 is the Proton client; wgpvpn is the static WireGuard + # profiles. A match-device line without both leaves one tunnel + # broken; one with a wildcard would turn DoT off for wifi too. + with tempfile.TemporaryDirectory() as d: + run(d) + match = [l for l in directives(body(d)) if l.startswith("match-device=")] + self.assertEqual(len(match), 1) + devices = match[0].split("=", 1)[1].split(",") + self.assertEqual(sorted(devices), + ["interface-name:proton0", "interface-name:wgpvpn"]) + + def test_explains_why_the_global_setting_is_overridden_per_link(self): + # The file contradicts the resolved drop-in next to it, so it has + # to carry the reason: the in-tunnel resolver that never completes + # TLS, and why the Proton client can't hold a per-profile setting. + with tempfile.TemporaryDirectory() as d: + run(d) + text = body(d).lower() + self.assertIn("10.2.0.1", text) + self.assertIn("853", text) + self.assertIn("recreates", text) + + def test_drop_in_is_world_readable_not_writable(self): + with tempfile.TemporaryDirectory() as d: + run(d) + mode = stat.S_IMODE(os.stat(os.path.join(d, FILENAME)).st_mode) + self.assertEqual(mode, 0o644) + + # ---------------------------------------------------------- boundary ---- + def test_running_twice_leaves_one_identical_file(self): + with tempfile.TemporaryDirectory() as d: + run(d) + first = body(d) + r = run(d) + self.assertEqual(rc_of(r), 0) + self.assertEqual(first, body(d)) + self.assertEqual(os.listdir(d), [FILENAME]) + + def test_absent_directory_is_created(self): + with tempfile.TemporaryDirectory() as d: + nested = os.path.join(d, "etc", "NetworkManager", "conf.d") + r = run(nested) + self.assertEqual(rc_of(r), 0) + self.assertIn("connection.dns-over-tls=0", directives(body(nested))) + + def test_leaves_sibling_drop_ins_alone(self): + # dns.conf, wifi-privacy.conf and wifi-powersave-off.conf share the + # directory; this step must add a file, never rewrite the dir. + with tempfile.TemporaryDirectory() as d: + sibling = os.path.join(d, "dns.conf") + with open(sibling, "w") as f: + f.write("[main]\ndns=systemd-resolved\n") + run(d) + with open(sibling) as f: + self.assertEqual(f.read(), "[main]\ndns=systemd-resolved\n") + + # ------------------------------------------------------------- error ---- + @unittest.skipUnless(os.geteuid() != 0, "root ignores directory write bits") + def test_unwritable_directory_warns_and_does_not_crash(self): + with tempfile.TemporaryDirectory() as d: + confdir = os.path.join(d, "ro") + os.mkdir(confdir) + os.chmod(confdir, 0o500) + try: + r = run(confdir) + self.assertIn("WARN:", r.stdout) + self.assertNotEqual(rc_of(r), 0) + finally: + os.chmod(confdir, 0o700) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/installer-steps/test_dotfiles_dependency_packages.py b/tests/installer-steps/test_dotfiles_dependency_packages.py new file mode 100644 index 0000000..58b7275 --- /dev/null +++ b/tests/installer-steps/test_dotfiles_dependency_packages.py @@ -0,0 +1,119 @@ +"""Pin the packages the dotfiles assume are present. + +Three packages are not optional extras: something the dotfiles ship depends +on each one, and when the package is missing the dependent silently does the +wrong thing rather than failing. + +- libreoffice-fresh :: common/.config/mimeapps.list maps presentations, + documents and spreadsheets to libreoffice-impress/-writer/-calc. With the + package absent those .desktop files don't exist, so xdg-mime falls through + to the next application claiming the type — on a machine with the winvm + dotfiles that is powerpoint.desktop, which boots a Windows VM to open a + deck (2026-09-18). +- imv :: gui-open --image execs imv, and without it every agent-side image + render fails with "required application is unavailable" (2026-09-18). +- git-lfs :: a repo tracking globs in LFS fails every checkout and merge + with "smudge filter lfs failed" (2026-09-20). + +All three were installed before the 2026-08-13 rebuild and absent after it, +which is the regression this pins: they are dependencies of shipped defaults, +so the installer has to declare them rather than leave them to whatever a +machine happens to carry. + +Method mirrors test_required_software: sed-extract supplemental_software from +the real `archsetup`, stub pacman_install as a recorder, run it, and assert +against what it actually invoked. Running the function beats matching its +source text, because the property under test is "every machine installs this" +and only a run resolves the conditionals that could make that false. The +function is run once per desktop environment for the same reason. + +Run from repo root: + python3 -m unittest tests.installer-steps.test_dotfiles_dependency_packages +""" + +import os +import subprocess +import textwrap +import unittest + +REPO_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")) +ARCHSETUP = os.path.join(REPO_ROOT, "archsetup") + +DEPENDENCY_PACKAGES = ("libreoffice-fresh", "imv", "git-lfs") + +# Every desktop environment the installer branches on inside this function. +DESKTOP_ENVS = ("dwm", "hyprland") + + +def declared_packages(desktop_env): + """Return (exit_code, [pacman package, ...]) for one desktop environment. + + Everything the function calls besides pacman_install is stubbed to a no-op, + so the run records package declarations and nothing else. aur_install is + deliberately separate: these three are pacman packages, and folding the two + recorders together would let an AUR declaration satisfy the pin. + """ + script = textwrap.dedent(f"""\ + desktop_env={desktop_env} + display() {{ :; }} + aur_install() {{ :; }} + mask_fwupd_passim() {{ :; }} + run_task() {{ :; }} + error_warn() {{ :; }} + pacman_install() {{ echo "$1"; }} + source <(sed -n '/^supplemental_software() {{/,/^}}/p' "{ARCHSETUP}") + supplemental_software + """) + result = subprocess.run( + ["bash", "-c", script], capture_output=True, text=True, timeout=30, + ) + return result.returncode, result.stdout.split() + + +class DotfilesDependencyPackages(unittest.TestCase): + # ------------------------------------------------------------ normal ---- + def test_each_dependency_package_is_installed(self): + rc, pkgs = declared_packages("hyprland") + self.assertEqual(rc, 0) + for package in DEPENDENCY_PACKAGES: + with self.subTest(package=package): + self.assertIn( + package, pkgs, + f"{package} is a dependency of a shipped dotfiles default " + "and must be declared", + ) + + # ---------------------------------------------------------- boundary ---- + def test_each_is_declared_exactly_once(self): + # A second declaration is dead weight and drifts out of sync with the + # first when one of them is edited. + rc, pkgs = declared_packages("hyprland") + self.assertEqual(rc, 0) + for package in DEPENDENCY_PACKAGES: + with self.subTest(package=package): + self.assertEqual(pkgs.count(package), 1) + + # ------------------------------------------------------------- error ---- + def test_none_is_gated_behind_a_desktop_environment(self): + # The dotfiles defaults that need these apply on every DE, so a + # declaration reachable under only one of them would leave the same + # hole on the other. + for env in DESKTOP_ENVS: + rc, pkgs = declared_packages(env) + self.assertEqual(rc, 0) + for package in DEPENDENCY_PACKAGES: + with self.subTest(desktop_env=env, package=package): + self.assertIn(package, pkgs) + + def test_harness_observes_desktop_environment_gating(self): + # The control for the test above: ranger IS gated to dwm on purpose, so + # if this run can't see that, the DE-independence assertion is vacuous + # and would pass against a gated package too. + _, dwm = declared_packages("dwm") + _, hyprland = declared_packages("hyprland") + self.assertIn("ranger", dwm) + self.assertNotIn("ranger", hyprland) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/installer-steps/test_mask_fwupd_passim.py b/tests/installer-steps/test_mask_fwupd_passim.py new file mode 100644 index 0000000..8977504 --- /dev/null +++ b/tests/installer-steps/test_mask_fwupd_passim.py @@ -0,0 +1,79 @@ +"""Test mask_fwupd_passim — keep fwupd's LAN metadata daemon off. + +fwupd pulls in passim, a daemon that shares firmware metadata with other +machines on the LAN by listening publicly on 0.0.0.0:27500. Any fwupdmgr +run D-Bus-activates it, and because the unit is static (no [Install] +section) `systemctl disable` is a no-op: it came back on velox the next +time fwupdmgr ran (2026-09-12). Masking is what holds, and it is how +ratio has carried it since 2026-07-21. + +Method: sed-extract mask_fwupd_passim from the real `archsetup`; fake +run_task (minus error_warn, so the failure test sees the function's own +return path) and systemctl, and assert on the call. + + python3 -m unittest tests.installer-steps.test_mask_fwupd_passim +""" + +import os +import re +import subprocess +import textwrap +import unittest + +REPO_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")) +ARCHSETUP = os.path.join(REPO_ROOT, "archsetup") + + +def run(systemctl_body='echo "SYSTEMCTL: $*";'): + script = textwrap.dedent(f"""\ + logfile=/dev/null + action="" + display() {{ :; }} + run_task() {{ echo "TASK: $1"; shift; "$@"; }} + systemctl() {{ {systemctl_body} }} + error_warn() {{ echo "WARN: $1"; return 1; }} + source <(sed -n '/^mask_fwupd_passim() {{/,/^}}/p' "{ARCHSETUP}") + mask_fwupd_passim + echo "RC=$?" + exit 0 + """) + return subprocess.run( + ["bash", "-c", script], capture_output=True, text=True, timeout=10, + ) + + +def rc_of(r): + m = re.search(r"^RC=(\d+)$", r.stdout, re.M) + assert m, "no RC line in output: %r / %r" % (r.stdout, r.stderr) + return int(m.group(1)) + + +class MaskFwupdPassim(unittest.TestCase): + # ------------------------------------------------------------ normal ---- + def test_masks_the_passim_unit(self): + r = run() + self.assertIn("SYSTEMCTL: mask passim.service", r.stdout) + self.assertEqual(rc_of(r), 0) + + def test_masks_rather_than_disables(self): + # disable is a no-op on a static unit and is the mistake this step + # exists to avoid; a mask is the only call that sticks. + r = run() + calls = [l for l in r.stdout.splitlines() if l.startswith("SYSTEMCTL:")] + self.assertEqual(calls, ["SYSTEMCTL: mask passim.service"]) + + # ---------------------------------------------------------- boundary ---- + # Re-running is idempotent because `systemctl mask` on an already-masked + # unit exits 0 and changes nothing; that property lives in systemctl, so + # there is no stateless test here that could tell it apart from a pass. + + # ------------------------------------------------------------- error ---- + def test_failed_mask_reports_nonzero(self): + # The fake run_task returns the command's status without the real + # error_warn wiring, so this checks the function's own return path. + r = run(systemctl_body='echo "SYSTEMCTL: $*"; return 1;') + self.assertNotEqual(rc_of(r), 0) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/net-scenarios/test_run_net_scenarios.py b/tests/net-scenarios/test_run_net_scenarios.py index 1d92185..a9cd275 100644 --- a/tests/net-scenarios/test_run_net_scenarios.py +++ b/tests/net-scenarios/test_run_net_scenarios.py @@ -63,6 +63,12 @@ class RunNetScenarios(unittest.TestCase): return subprocess.run( ["bash", SCRIPT, "--target", "root@fake-vm"], capture_output=True, text=True, timeout=20, env=env, + # The stubbed ssh is `cat >/dev/null`, which drains stdin to EOF. + # Without this the stub inherits whatever stdin the test runner + # had, so `make test-unit` passed when stdin was redirected and + # hung on all five tests when it was a terminal or a live pipe -- + # which is how it gets run by hand. + stdin=subprocess.DEVNULL, ) def test_all_checks_pass_exits_zero(self): diff --git a/tests/post-rebuild-check/test_post_rebuild_check.py b/tests/post-rebuild-check/test_post_rebuild_check.py new file mode 100644 index 0000000..bc887c6 --- /dev/null +++ b/tests/post-rebuild-check/test_post_rebuild_check.py @@ -0,0 +1,1145 @@ +"""Tests for the post-rebuild-check script. + +A rebuilt machine looks finished and isn't: on velox 2026-08-13 five gaps +surfaced within two days, three of which LOOKED fine (a stowed unit file, an +enabled timer, a present git clone). The script runs the checks from the +post-rebuild task and turns each silent no-op into a visible line: + + 1. failed systemd units (user and system scope) + 2. user unit files present but not enabled (linked-and-inert timers) + 3. tracked *.example files whose real sibling is missing + 4. gitignore-mode projects missing tooling paths their own .gitignore names + 5. signal-cli holds no registered account + 6. every NTP source named by hostname (a wrong clock takes DNS with it) + 7. hypridle installed but not running (nothing triggers idle suspend) + 8. a working repo cloned read-only (push returns 403) + +Exit 0 with every check clean, 1 when any check found something. + +Test seams (env vars the production script honors; for each, SET-BUT-EMPTY +means "the real probe ran and found nothing", UNSET means "run the real +probe"): + PRC_FAILED_UNITS newline list of "scope:unit" (scope user|system) + PRC_UNIT_STATES newline list of "unit-file state" for the user unit dir + PRC_LOCAL_SCAN_ROOTS newline-separated roots to scan for *.example orphans + PRC_PROJECT_ROOTS newline-separated project dirs for the tooling check + PRC_SIGNAL_ACCOUNTS signal-cli listAccounts output ("" = no accounts); + PRC_NTP_SOURCES newline list of configured NTP server addresses + ("MISSING" = no NTP daemon active) + the special value MISSING means the binary is absent + PRC_IDLE_DAEMON pgrep output for hypridle ("" = installed but not + running; "MISSING" = not installed on this machine) + PRC_REPO_REMOTES newline list of "path<space>origin-url" for the + push-capability check ("" = no repos to check) + PRC_UNITS_EXPECTED_DISABLED + newline list of units whose not-enabled state is + deliberate on this machine ("" = no exemptions) + +Run from repo root: + python3 -m unittest tests.post-rebuild-check.test_post_rebuild_check +""" + +import os +import shutil +import subprocess +import tempfile +import time +import unittest + + +REPO_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")) +CHECK = os.path.join(REPO_ROOT, "scripts", "post-rebuild-check") + + +def run_check(failed_units="", unit_states="", local_roots="", + project_roots="", signal_accounts="+15045551234", + ntp_sources="162.159.200.1\npool.ntp.org", + idle_daemon="4242", repo_remotes="", + units_expected_disabled=""): + """Run the script with every probe stubbed; defaults are all-clean. + + Roots are newline-separated. Empty means "the seam is set and names no + roots" -- the script tests with ${VAR+set}, so an empty value is still + set and never falls through to the real probe. + """ + env = dict(os.environ) + env["PRC_FAILED_UNITS"] = failed_units + env["PRC_UNIT_STATES"] = unit_states + env["PRC_LOCAL_SCAN_ROOTS"] = local_roots + env["PRC_PROJECT_ROOTS"] = project_roots + env["PRC_SIGNAL_ACCOUNTS"] = signal_accounts + env["PRC_NTP_SOURCES"] = ntp_sources + env["PRC_IDLE_DAEMON"] = idle_daemon + env["PRC_REPO_REMOTES"] = repo_remotes + env["PRC_UNITS_EXPECTED_DISABLED"] = units_expected_disabled + return subprocess.run( + ["sh", CHECK], capture_output=True, text=True, timeout=30, env=env, + ) + + +class NtpBootstrap(unittest.TestCase): + """Check 6 — the clock/DNS bootstrap deadlock. + + A wrong clock fails the DoT certificate and DNSSEC signature checks this + machine's DNS runs on, so nothing resolves; and an NTP daemon whose every + source is a hostname then cannot resolve the servers that would correct + the clock. One source addressed by IP is what makes the machine able to + recover on its own. + """ + + # --- Normal cases --------------------------------------------------- + + def test_an_ip_addressed_source_is_clean(self): + r = run_check(ntp_sources="162.159.200.1\npool.ntp.org") + self.assertIn("check 6/8: NTP bootstrap — ok", r.stdout) + self.assertEqual(r.returncode, 0, r.stdout) + + def test_all_hostname_sources_is_a_finding(self): + # The velox 2026-08-19 shape exactly: stock Arch chrony.conf, whose + # only source is a pool hostname. + r = run_check(ntp_sources="2.arch.pool.ntp.org") + self.assertIn("every NTP source is named by hostname", r.stdout) + self.assertEqual(r.returncode, 1) + + def test_an_ipv6_addressed_source_counts(self): + r = run_check(ntp_sources="2606:4700:f1::1") + self.assertIn("check 6/8: NTP bootstrap — ok", r.stdout) + + # --- Boundary cases ------------------------------------------------- + + def test_the_literal_may_sit_anywhere_in_the_list(self): + # Order must not matter; the property is "at least one", and the + # drop-in that carries it is read after the main config. + r = run_check(ntp_sources="a.pool.ntp.org\nb.pool.ntp.org\n162.159.200.1") + self.assertIn("check 6/8: NTP bootstrap — ok", r.stdout) + + def test_blank_lines_between_sources_are_ignored(self): + r = run_check(ntp_sources="\n\n162.159.200.1\n\n") + self.assertIn("check 6/8: NTP bootstrap — ok", r.stdout) + + def test_a_hostname_containing_digits_and_dots_is_not_an_address(self): + # The trap in any naive "looks like an IP" test: these resolve through + # DNS like any other name, so counting one as an address would hand a + # deadlocked machine a clean bill. + for host in ("0.arch.pool.ntp.org", "3.us.pool.ntp.org", "time1.google.com"): + with self.subTest(host=host): + r = run_check(ntp_sources=host) + self.assertIn("every NTP source is named by hostname", r.stdout) + + # --- Error cases ---------------------------------------------------- + + def test_no_ntp_daemon_is_a_finding(self): + r = run_check(ntp_sources="MISSING") + self.assertIn("no NTP implementation is active", r.stdout) + self.assertEqual(r.returncode, 1) + + def test_no_sources_configured_is_a_finding(self): + # Fails closed: an empty list proves nothing about the machine, and + # reporting ok would be a false pass on a box with no time sync at all. + r = run_check(ntp_sources="") + self.assertIn("no NTP sources are configured", r.stdout) + self.assertEqual(r.returncode, 1) + + # --- the confdir false pass ------------------------------------------ + # + # A drop-in is inert unless chrony.conf names its directory, and Arch's + # stock chrony.conf names none. Reading the drop-in without checking for + # confdir would find the IP-addressed source, call the machine healthy, and + # be describing a file chrony never opens — a false pass on exactly the + # misconfiguration this check exists to catch. + + def _chrony_fixture(self, main_lines, dropin_lines=None): + """Write a chrony.conf (plus an adjacent drop-in dir) and return its path.""" + d = tempfile.mkdtemp(prefix="prc-chrony-") + self.addCleanup(shutil.rmtree, d, True) + dropin_dir = os.path.join(d, "chrony.d") + os.makedirs(dropin_dir) + if dropin_lines is not None: + with open(os.path.join(dropin_dir, "10-bootstrap-ip-ntp.conf"), "w") as f: + f.write(dropin_lines) + conf = os.path.join(d, "chrony.conf") + with open(conf, "w") as f: + f.write(main_lines.replace("@DROPIN@", dropin_dir)) + return conf + + def _run_real_probe(self, chrony_conf): + """Run with PRC_NTP_SOURCES unset so the real chrony reader runs.""" + env = dict(os.environ) + env.update({"PRC_FAILED_UNITS": "", "PRC_UNIT_STATES": "", + "PRC_LOCAL_SCAN_ROOTS": "", "PRC_PROJECT_ROOTS": "", + "PRC_SIGNAL_ACCOUNTS": "+15045551234", + "PRC_IDLE_DAEMON": "4242", + "PRC_REPO_REMOTES": "", + "PRC_UNITS_EXPECTED_DISABLED": "", + "PRC_CHRONY_CONF": chrony_conf}) + env.pop("PRC_NTP_SOURCES", None) + return subprocess.run(["sh", CHECK], capture_output=True, text=True, + timeout=30, env=env) + + def test_dropin_without_confdir_does_not_count(self): + # The regression. The IP-addressed source is present on disk but + # chrony.conf never points at it, so the machine is still deadlock-prone + # and the check has to say so. + conf = self._chrony_fixture("pool 2.arch.pool.ntp.org iburst\n", + "server 162.159.200.1 iburst\n") + r = self._run_real_probe(conf) + if "no NTP implementation is active" in r.stdout: + self.skipTest("no chronyd on this host — the reader branch can't run") + self.assertIn("every NTP source is named by hostname", r.stdout) + + def test_dropin_with_confdir_counts(self): + # The same two files, with chrony.conf actually naming the directory. + conf = self._chrony_fixture( + "pool 2.arch.pool.ntp.org iburst\nconfdir @DROPIN@\n", + "server 162.159.200.1 iburst\n") + r = self._run_real_probe(conf) + if "no NTP implementation is active" in r.stdout: + self.skipTest("no chronyd on this host — the reader branch can't run") + self.assertIn("check 6/8: NTP bootstrap — ok", r.stdout) + + def test_confdir_naming_an_empty_directory_is_not_a_pass(self): + # confdir present, nothing behind it: the sources are the hostname-only + # main file, so the finding stands. + conf = self._chrony_fixture( + "pool 2.arch.pool.ntp.org iburst\nconfdir @DROPIN@\n", None) + r = self._run_real_probe(conf) + if "no NTP implementation is active" in r.stdout: + self.skipTest("no chronyd on this host — the reader branch can't run") + self.assertIn("every NTP source is named by hostname", r.stdout) + + def test_unset_seam_falls_through_to_the_real_probe(self): + # Same contract as every other seam: unset means "really look", so a + # caller who forgets the variable cannot silently skip the check. + env = dict(os.environ) + env.update({"PRC_FAILED_UNITS": "", "PRC_UNIT_STATES": "", + "PRC_LOCAL_SCAN_ROOTS": "", "PRC_PROJECT_ROOTS": "", + "PRC_SIGNAL_ACCOUNTS": "+15045551234"}) + env.pop("PRC_NTP_SOURCES", None) + r = subprocess.run(["sh", CHECK], capture_output=True, text=True, + timeout=30, env=env) + self.assertIn("check 6/8: NTP bootstrap", r.stdout) + + +class IdleDaemon(unittest.TestCase): + """Check 7 — whether anything still triggers idle lock and suspend. + + A laptop that never sleeps has no symptom until the battery is gone, which + is why this needs a check rather than trusting the desktop to look right. + On velox 2026-08-19 hypridle started cleanly at 15:29:48 and `settings + restore` killed it six seconds later, replaying a caffeine stored in an + earlier boot. The machine then ran 11h40m fully awake on battery, died when + it flattened, and reset its RTC — which took DNS down with it, the very + deadlock check 6 exists for. Nothing looked wrong at any point. + + The check is behavioural: it asks whether the daemon is alive, not why it + might not be, so a stale caffeine, a crash and a bad config all surface the + same way. + """ + + # --- Normal cases --------------------------------------------------- + + def test_running_daemon_is_clean(self): + r = run_check(idle_daemon="4242") + self.assertEqual(r.returncode, 0, r.stdout + r.stderr) + self.assertNotIn("DEVIATION", r.stdout) + + def test_installed_but_not_running_flags(self): + r = run_check(idle_daemon="") + self.assertEqual(r.returncode, 1) + self.assertIn("hypridle", r.stdout) + self.assertIn("DEVIATION", r.stdout) + + def test_the_finding_names_the_consequence_not_just_the_process(self): + # "hypridle is not running" reads as a detail. The reason it matters is + # that the machine stays awake until the battery is gone, and that is + # what has to be in the line someone skims at 1am. + r = run_check(idle_daemon="") + self.assertIn("awake", r.stdout.lower()) + + def test_the_finding_names_the_known_cause(self): + # Behavioural checks are cheap to write and expensive to act on. Naming + # the one cause already seen saves the reader the investigation this + # session had to do from scratch. + r = run_check(idle_daemon="") + self.assertIn("caffeine", r.stdout.lower()) + + # --- Boundary cases ------------------------------------------------- + + def test_not_installed_is_not_a_finding(self): + # A headless or dwm machine never installs hypridle — archsetup pulls + # it in only for Hyprland. Flagging its absence there would be noise on + # every run, and noise is how a real finding gets skimmed past. + r = run_check(idle_daemon="MISSING") + self.assertEqual(r.returncode, 0, r.stdout + r.stderr) + self.assertNotIn("DEVIATION", r.stdout) + + def test_several_pids_still_read_as_running(self): + # pgrep prints one pid per line. More than one is its own problem (five + # concurrent daemons wedged a velox session on 2026-07-22) but it is + # not *this* check's, and it must not read as "not running". + r = run_check(idle_daemon="4242\n4243") + self.assertEqual(r.returncode, 0, r.stdout + r.stderr) + + def test_the_check_always_prints_its_line(self): + for pids in ("4242", "", "MISSING"): + with self.subTest(pids=pids): + self.assertIn("idle daemon", + run_check(idle_daemon=pids).stdout.lower()) + + +class UnitsExpectedDisabled(unittest.TestCase): + """Check 2 — units nothing intends to enable on this machine. + + "Enabled" is the check's proxy for "will actually run", and the proxy is + wrong for a unit nobody means to enable here. velox carries four such + units, for four different reasons: geoclue-agent is redundant because + hyprland's exec-once starts the binary directly, emacs is started on demand + by emacsclient, obs-record-watchdog only matters while recording, and + obsbot-wb-guard needs an OBSBOT the machine doesn't have. + + Left unexempted they report at every run, which is the standing-findings + problem check 4's own comment already argues against: four permanent lines + in front of every real one teach you to skim the output. + + The exemption is machine-local rather than a marker in the shared unit + file, because obsbot-wb-guard is correctly ENABLED on ratio. Same unit, + different right answer per machine. + """ + + # --- Normal cases --------------------------------------------------- + + def test_an_exempt_unit_is_not_flagged(self): + r = run_check(unit_states="emacs.service linked", + units_expected_disabled="emacs.service") + self.assertEqual(r.returncode, 0, r.stdout) + self.assertNotIn("emacs.service", r.stdout) + + def test_a_non_exempt_unit_still_flags(self): + r = run_check(unit_states="roam-sync.timer linked", + units_expected_disabled="emacs.service") + self.assertEqual(r.returncode, 1) + self.assertIn("roam-sync.timer", r.stdout) + + def test_several_exemptions_all_apply(self): + r = run_check( + unit_states=("emacs.service linked\n" + "geoclue-agent.service linked\n" + "obsbot-wb-guard.service linked"), + units_expected_disabled=("emacs.service\n" + "geoclue-agent.service\n" + "obsbot-wb-guard.service")) + self.assertEqual(r.returncode, 0, r.stdout) + + # --- Boundary cases ------------------------------------------------- + + def test_an_exemption_that_is_actually_enabled_is_a_finding(self): + # A stale exemption must surface rather than sit there suppressing + # nothing. Otherwise the list rots into a place real findings go to + # die, which is worse than the noise it was added to remove. + r = run_check(unit_states="obsbot-wb-guard.service enabled", + units_expected_disabled="obsbot-wb-guard.service") + self.assertEqual(r.returncode, 1) + self.assertIn("obsbot-wb-guard.service", r.stdout) + + def test_comments_and_blank_lines_are_ignored(self): + # The reason a unit is exempt is the most useful thing about the + # entry, so the format has to hold a comment next to it. + r = run_check(unit_states="emacs.service linked", + units_expected_disabled=("# started on demand\n" + "\n" + "emacs.service # not by systemd\n")) + self.assertEqual(r.returncode, 0, r.stdout) + + def test_no_exemptions_flags_everything_as_before(self): + r = run_check(unit_states="emacs.service linked", + units_expected_disabled="") + self.assertEqual(r.returncode, 1) + self.assertIn("emacs.service", r.stdout) + + def test_an_exemption_does_not_suppress_a_dangling_link(self): + # A stowed unit pointing at a missing target is a different finding, + # decided on the filesystem. Exempting the name must not hide that. + d = tempfile.mkdtemp(prefix="prc-units-") + self.addCleanup(shutil.rmtree, d, True) + unit_dir = os.path.join(d, "systemd", "user") + os.makedirs(unit_dir) + link = os.path.join(unit_dir, "emacs.service") + os.symlink(os.path.join(d, "gone.service"), link) + env = dict(os.environ) + env.update({"PRC_FAILED_UNITS": "", "PRC_LOCAL_SCAN_ROOTS": "", + "PRC_PROJECT_ROOTS": "", + "PRC_SIGNAL_ACCOUNTS": "+15045551234", + "PRC_NTP_SOURCES": "162.159.200.1", + "PRC_IDLE_DAEMON": "4242", + "PRC_REPO_REMOTES": "", + "PRC_UNITS_EXPECTED_DISABLED": "emacs.service", + "XDG_CONFIG_HOME": d}) + env.pop("PRC_UNIT_STATES", None) + r = subprocess.run(["sh", CHECK], capture_output=True, text=True, + timeout=30, env=env) + self.assertIn("points at a missing target", r.stdout) + self.assertEqual(r.returncode, 1) + + +class RepoPushCapability(unittest.TestCase): + """Check 8 — a working repo cloned from the read-only endpoint. + + archsetup clones the user's own archsetup and dotfiles from + https://git.cjennings.net/..., the anonymous read-only endpoint. That is + the right default for a stranger installing archsetup, who has no key on + the server, and the wrong one for this machine, which has to push. The + override exists (ARCHSETUP_REPO / DOTFILES_REPO) but only applies where it + is configured — a curl|bash install, or a rebuild from a stock ISO, picks + the default straight back up. + + Nothing about the tree shows it. The clone is complete and ordinary, and + the machine finds out at the first push, with a 403. That is how velox's + dotfiles remote was found on 2026-08-17, four days after its rebuild. + + Only the read-only endpoint is flagged. An https remote to some other host + may well be pushable with a credential helper, and guessing about hosts + this machine does not own would put standing noise in front of the real + findings. + """ + + RO = "https://git.cjennings.net/dotfiles.git" + RW = "git@cjennings.net:dotfiles.git" + + # --- Normal cases --------------------------------------------------- + + def test_an_ssh_remote_is_clean(self): + r = run_check(repo_remotes=f"/home/x/.dotfiles {self.RW}") + self.assertEqual(r.returncode, 0, r.stdout) + self.assertNotIn("DEVIATION", r.stdout) + + def test_the_read_only_endpoint_flags(self): + r = run_check(repo_remotes=f"/home/x/.dotfiles {self.RO}") + self.assertEqual(r.returncode, 1) + self.assertIn("/home/x/.dotfiles", r.stdout) + + def test_the_finding_names_the_consequence(self): + # "the remote is https" is a detail. That pushing fails is the point. + r = run_check(repo_remotes=f"/home/x/.dotfiles {self.RO}") + self.assertIn("push", r.stdout.lower()) + + def test_every_offending_repo_is_named(self): + r = run_check(repo_remotes=(f"/home/x/.dotfiles {self.RO}\n" + f"/home/x/code/archsetup {self.RO}")) + self.assertIn("/home/x/.dotfiles", r.stdout) + self.assertIn("/home/x/code/archsetup", r.stdout) + + # --- Boundary cases ------------------------------------------------- + + def test_a_mixed_set_flags_only_the_read_only_one(self): + r = run_check(repo_remotes=(f"/home/x/.dotfiles {self.RW}\n" + f"/home/x/code/archsetup {self.RO}")) + self.assertEqual(r.returncode, 1) + self.assertIn("/home/x/code/archsetup", r.stdout) + self.assertNotIn("/home/x/.dotfiles", r.stdout) + + def test_an_https_remote_to_another_host_is_not_flagged(self): + # GitHub over https is pushable with a credential helper. Flagging it + # would be a guess about a host this machine does not own. + r = run_check(repo_remotes="/home/x/code/thing https://github.com/a/b.git") + self.assertEqual(r.returncode, 0, r.stdout) + + def test_no_repos_is_not_a_finding(self): + r = run_check(repo_remotes="") + self.assertEqual(r.returncode, 0, r.stdout) + + def test_blank_lines_are_ignored(self): + r = run_check(repo_remotes=f"\n\n/home/x/.dotfiles {self.RW}\n\n") + self.assertEqual(r.returncode, 0, r.stdout) + + def test_the_check_always_prints_its_line(self): + for remotes in ("", f"/home/x/.dotfiles {self.RW}", + f"/home/x/.dotfiles {self.RO}"): + with self.subTest(remotes=remotes): + self.assertIn("repo remotes", + run_check(repo_remotes=remotes).stdout.lower()) + + # --- Error cases ---------------------------------------------------- + + def test_a_repo_with_no_origin_is_a_finding(self): + # Fails closed. A repo whose origin could not be read was not checked, + # and reporting it clean is the false pass this script exists to avoid. + r = run_check(repo_remotes="/home/x/.dotfiles") + self.assertEqual(r.returncode, 1) + self.assertIn("/home/x/.dotfiles", r.stdout) + + +class AllClean(unittest.TestCase): + # --- Normal cases --------------------------------------------------- + + def test_all_clean_exits_zero(self): + r = run_check() + self.assertEqual(r.returncode, 0, r.stdout + r.stderr) + + def test_all_clean_prints_one_line_per_check(self): + # The visible line per check is the point of the script: a silent + # no-op is exactly what let the velox gaps sit unseen for two days. + r = run_check() + for label in ("failed units", "unit files", "local files", + "project tooling", "signal"): + self.assertIn(label, r.stdout.lower()) + + def test_all_clean_summary_says_clean(self): + r = run_check() + self.assertIn("all checks clean", r.stdout.lower()) + + def test_all_clean_no_deviation_lines(self): + r = run_check() + self.assertNotIn("DEVIATION", r.stdout) + + +class FailedUnits(unittest.TestCase): + # --- Normal cases --------------------------------------------------- + + def test_failed_user_unit_flags(self): + r = run_check(failed_units="user:calendar-sync.service") + self.assertEqual(r.returncode, 1) + self.assertIn("calendar-sync.service", r.stdout) + self.assertIn("DEVIATION", r.stdout) + + def test_failed_system_unit_flags(self): + r = run_check(failed_units="system:tlp.service") + self.assertEqual(r.returncode, 1) + self.assertIn("tlp.service", r.stdout) + + def test_multiple_failed_units_each_reported(self): + r = run_check( + failed_units="user:calendar-sync.service\nsystem:tlp.service") + self.assertIn("calendar-sync.service", r.stdout) + self.assertIn("tlp.service", r.stdout) + + # --- Boundary cases ------------------------------------------------- + + def test_blank_lines_in_seam_ignored(self): + r = run_check(failed_units="\n\nuser:a.service\n\n") + self.assertEqual(r.returncode, 1) + self.assertIn("a.service", r.stdout) + + +class UnitFilesNotEnabled(unittest.TestCase): + # --- Normal cases --------------------------------------------------- + + def test_disabled_timer_flags(self): + r = run_check(unit_states="roam-sync.timer disabled") + self.assertEqual(r.returncode, 1) + self.assertIn("roam-sync.timer", r.stdout) + + def test_linked_timer_flags(self): + # The exact velox case: a unit symlinked into the user dir by hand, + # never enabled — present, inert, and it LOOKS installed. + r = run_check(unit_states="signal-receive.timer linked") + self.assertEqual(r.returncode, 1) + self.assertIn("signal-receive.timer", r.stdout) + + def test_enabled_timer_passes(self): + r = run_check(unit_states="roam-sync.timer enabled") + self.assertEqual(r.returncode, 0, r.stdout) + + def test_static_service_passes(self): + # A service with no [Install] section is pulled in by its timer; + # "static" is its healthy state, not a gap. + r = run_check(unit_states="roam-sync.service static") + self.assertEqual(r.returncode, 0, r.stdout) + + def test_disabled_service_flags(self): + r = run_check(unit_states="obsbot-wb-guard.service disabled") + self.assertEqual(r.returncode, 1) + self.assertIn("obsbot-wb-guard.service", r.stdout) + + # --- Boundary cases ------------------------------------------------- + + def test_mixed_states_only_inert_reported(self): + r = run_check(unit_states="a.timer enabled\nb.timer disabled\n" + "c.service static\nd.service linked") + self.assertEqual(r.returncode, 1) + self.assertNotIn("a.timer", r.stdout) + self.assertIn("b.timer", r.stdout) + self.assertNotIn("c.service", r.stdout) + self.assertIn("d.service", r.stdout) + + def test_masked_unit_passes(self): + # Masking is a deliberate act (ppd on laptops), not rebuild rot. + r = run_check(unit_states="power-profiles-daemon.service masked") + self.assertEqual(r.returncode, 0, r.stdout) + + def test_service_whose_timer_is_enabled_passes(self): + # A timer-activated service is SUPPOSED to sit linked-not-enabled: + # the timer owns activation, and enabling the service too would run + # it at boot as well. Six of velox's units are this shape, and + # flagging them is the noise that gets a check ignored. + r = run_check(unit_states="roam-sync.service linked\n" + "roam-sync.timer enabled") + self.assertEqual(r.returncode, 0, r.stdout) + + def test_service_whose_timer_is_inert_flags_the_timer_only(self): + # When the timer itself never got enabled, the timer is the finding. + # Naming the service too would double-count one gap. + r = run_check(unit_states="obs-record-watchdog.service linked\n" + "obs-record-watchdog.timer linked") + self.assertEqual(r.returncode, 1) + self.assertIn("obs-record-watchdog.timer", r.stdout) + self.assertNotIn("obs-record-watchdog.service", r.stdout) + + def test_service_without_a_timer_still_flags(self): + # Nothing else can start it, so linked-not-enabled means dead. + r = run_check(unit_states="emacs.service linked") + self.assertEqual(r.returncode, 1) + self.assertIn("emacs.service", r.stdout) + + def test_a_runtime_enabled_timer_suppresses_its_service(self): + # enabled-runtime is a live activation path (enabled until reboot) + # and generated means something produced and installed it, so the + # service beneath either is being started and is not a finding. + for state in ("enabled-runtime", "generated"): + with self.subTest(timer=state): + r = run_check(unit_states=f"foo.service linked\n" + f"foo.timer {state}") + self.assertEqual(r.returncode, 0, + f"a {state} timer failed to suppress") + + def test_an_indirect_timer_does_not_suppress_its_service(self): + # "indirect" means the unit file itself is NOT enabled -- only that + # some Also= relative might be. Under this script's own fail-closed + # rule the uncertain case flags, so suppressing here would be the + # masked blind spot again in a narrower form. + r = run_check(unit_states="foo.service linked\nfoo.timer indirect") + self.assertEqual(r.returncode, 1, + "an indirect timer suppressed a service nothing starts") + self.assertIn("foo.service", r.stdout) + + def test_a_service_whose_timer_cannot_start_it_still_flags(self): + # Suppression is earned by a timer that can actually run the service. + # A masked, static, or absent timer starts nothing, so the service is + # as dead as one with no timer at all -- and suppressing on the mere + # presence of a timer line hides exactly that. + for state in ("masked", "static", "not-found"): + with self.subTest(timer=state): + r = run_check(unit_states=f"foo.service linked\n" + f"foo.timer {state}") + self.assertEqual(r.returncode, 1, + f"a {state} timer suppressed a dead service") + self.assertIn("foo.service", r.stdout) + + +class LocalExampleOrphans(unittest.TestCase): + # --- Normal cases --------------------------------------------------- + + def test_example_without_sibling_flags(self): + with tempfile.TemporaryDirectory() as root: + open(os.path.join(root, "auth.local.el.example"), "w").close() + r = run_check(local_roots=root) + self.assertEqual(r.returncode, 1) + self.assertIn("auth.local.el.example", r.stdout) + + def test_a_real_file_that_is_a_dangling_symlink_still_flags(self): + # A sibling that exists only as a broken link is not a config the + # machine can read, so it is the same gap as an absent one. + with tempfile.TemporaryDirectory() as root: + open(os.path.join(root, "auth.local.el.example"), "w").close() + os.symlink("/nonexistent/stow/target", + os.path.join(root, "auth.local.el")) + r = run_check(local_roots=root) + self.assertEqual(r.returncode, 1, + "a dangling sibling counted as present") + self.assertIn("auth.local.el.example", r.stdout) + + def test_example_with_sibling_passes(self): + with tempfile.TemporaryDirectory() as root: + open(os.path.join(root, "auth.local.el.example"), "w").close() + open(os.path.join(root, "auth.local.el"), "w").close() + r = run_check(local_roots=root) + self.assertEqual(r.returncode, 0, r.stdout) + + def test_nested_example_found(self): + with tempfile.TemporaryDirectory() as root: + sub = os.path.join(root, "modules") + os.makedirs(sub) + open(os.path.join(sub, "mail.local.el.example"), "w").close() + r = run_check(local_roots=root) + self.assertEqual(r.returncode, 1) + self.assertIn("mail.local.el.example", r.stdout) + + # --- Boundary cases ------------------------------------------------- + + def test_two_roots_both_scanned(self): + with tempfile.TemporaryDirectory() as a, \ + tempfile.TemporaryDirectory() as b: + open(os.path.join(a, "one.example"), "w").close() + open(os.path.join(b, "two.example"), "w").close() + r = run_check(local_roots=a + "\n" + b) + self.assertIn("one.example", r.stdout) + self.assertIn("two.example", r.stdout) + + def test_vendored_package_dirs_not_scanned(self): + # elpa/ and friends hold third-party packages that ship their own + # .example docs. Those are the package's business, not this machine's, + # and one of them (dirvish's) was the only finding check 3 produced on + # velox — a standing false positive in front of any real one. + with tempfile.TemporaryDirectory() as root: + for vendor in ("elpa", "node_modules", ".venv", "straight"): + d = os.path.join(root, vendor, "pkg-1.0", "docs") + os.makedirs(d) + open(os.path.join(d, "config.example"), "w").close() + r = run_check(local_roots=root) + self.assertEqual(r.returncode, 0, r.stdout) + + def test_an_unreadable_vendored_dir_is_not_a_finding(self): + # The vendored trees are excluded by design, so failing to descend + # into one is not a gap in what this check covers. Filtering find's + # output without pruning its descent turns a package directory + # nobody wanted read into a standing "could not fully scan". + with tempfile.TemporaryDirectory() as root: + locked = os.path.join(root, "elpa", "pkg-1.0") + os.makedirs(locked) + os.chmod(locked, 0o000) + try: + r = run_check(local_roots=root) + finally: + os.chmod(locked, 0o755) + self.assertEqual(r.returncode, 0, r.stdout) + + def test_git_dir_not_scanned(self): + # .git holds hooks' sample files; those are git's, not the tree's. + with tempfile.TemporaryDirectory() as root: + g = os.path.join(root, ".git", "hooks") + os.makedirs(g) + open(os.path.join(g, "pre-commit.example"), "w").close() + r = run_check(local_roots=root) + self.assertEqual(r.returncode, 0, r.stdout) + + # --- Error cases ---------------------------------------------------- + + def test_an_unreadable_subdirectory_is_a_finding_not_a_pass(self): + # find exits non-zero when it cannot descend somewhere, and prints + # what it did reach. Discarding that status hides every orphan under + # the unreadable directory behind a clean "ok" -- the same defect + # class as a probe that cannot run reading as a pass. + with tempfile.TemporaryDirectory() as root: + locked = os.path.join(root, "locked") + os.makedirs(locked) + open(os.path.join(locked, "auth.local.el.example"), "w").close() + os.chmod(locked, 0o000) + try: + r = run_check(local_roots=root) + finally: + os.chmod(locked, 0o755) + self.assertEqual(r.returncode, 1, + "an unreadable directory read as nothing to check") + self.assertIn("could not", r.stdout.lower()) + + def test_missing_root_is_its_own_finding(self): + # A scan root that's gone is a rebuild gap too, not a pass. + r = run_check(local_roots="/nonexistent/scan-root") + self.assertEqual(r.returncode, 1) + self.assertIn("/nonexistent/scan-root", r.stdout) + + def test_a_path_with_spaces_is_one_root_not_three(self): + # Roots arrive newline-separated for this reason: splitting on spaces + # turns one real directory into several imaginary missing ones. + with tempfile.TemporaryDirectory() as base: + root = os.path.join(base, "a dir with spaces") + os.makedirs(root) + open(os.path.join(root, "orphan.example"), "w").close() + r = run_check(local_roots=root) + self.assertEqual(r.stdout.count("DEVIATION"), 1, r.stdout) + self.assertIn("orphan.example", r.stdout) + + +class ProjectTooling(unittest.TestCase): + def project(self, root, gitignore_lines, present=()): + os.makedirs(os.path.join(root, ".git")) + with open(os.path.join(root, ".gitignore"), "w") as f: + f.write("\n".join(gitignore_lines) + "\n") + for p in present: + path = os.path.join(root, p) + if p.endswith("/"): + os.makedirs(path, exist_ok=True) + else: + open(path, "w").close() + + # --- Normal cases --------------------------------------------------- + + def test_ignored_but_absent_tooling_flags(self): + with tempfile.TemporaryDirectory() as root: + self.project(root, [".ai/", "todo.org"]) + r = run_check(project_roots=root) + self.assertEqual(r.returncode, 1) + for missing in (".ai", "todo.org"): + self.assertIn(missing, r.stdout) + + def test_claude_dir_absence_never_flags(self): + # Same shape as CLAUDE.md below, and proven the same way. The bootstrap + # and the gitignore sweep write `.claude/` into the ignore set of every + # gitignore-mode project whether or not one ever exists there, so the + # entry is aspirational rather than a promise. Three projects tripped + # this on velox, and ratio is missing the identical directory in the + # identical three, which is what proves it is the steady state and not + # reinstall drift. + # + # Nor does dropping it lose a real signal. A project that genuinely + # carries one (rules and hooks from a language bundle) has it re-synced + # by sync-language-bundle.sh at every session start, so a true absence + # heals itself before this check would ever run. + with tempfile.TemporaryDirectory() as root: + self.project(root, [".ai/", ".claude/"], present=(".ai/",)) + r = run_check(project_roots=root) + self.assertEqual(r.returncode, 0, r.stdout) + self.assertNotIn(".claude", r.stdout) + + def test_ignored_and_present_tooling_passes(self): + with tempfile.TemporaryDirectory() as root: + self.project(root, [".ai/", "CLAUDE.md"], + present=(".ai/", "CLAUDE.md")) + r = run_check(project_roots=root) + self.assertEqual(r.returncode, 0, r.stdout) + + def test_claude_md_absence_never_flags(self): + # CLAUDE.md is seed-only: install-lang writes it once and the project + # owns it afterward, so most projects legitimately never have one. + # Ratio shows the identical absences in the identical projects, which + # is what proves it is the steady state and not reinstall drift. + with tempfile.TemporaryDirectory() as root: + self.project(root, [".ai/", "CLAUDE.md"], present=(".ai/",)) + r = run_check(project_roots=root) + self.assertEqual(r.returncode, 0, r.stdout) + + def test_missing_ai_dir_still_flags(self): + # The one that carries real working state — 374 files in .emacs.d's + # case — and that nothing restores: not git, not stow, not bootstrap. + with tempfile.TemporaryDirectory() as root: + self.project(root, [".ai/", "CLAUDE.md"]) + r = run_check(project_roots=root) + self.assertEqual(r.returncode, 1) + self.assertIn(".ai", r.stdout) + self.assertNotIn("CLAUDE.md", r.stdout) + + def test_unignored_tooling_never_expected(self): + # A project that never gitignored todo.org never had one to lose; + # the project's own .gitignore is the record of what it should hold. + with tempfile.TemporaryDirectory() as root: + self.project(root, [".ai/"], present=(".ai/",)) + r = run_check(project_roots=root) + self.assertEqual(r.returncode, 0, r.stdout) + + # --- Boundary cases ------------------------------------------------- + + def test_anchored_ignore_style_recognized(self): + # Both /.ai/ (anchored) and .ai/ (unanchored) styles exist across + # the fleet; the sweep-gitignore audit hit exactly this split. + with tempfile.TemporaryDirectory() as root: + self.project(root, ["/.ai/", "/CLAUDE.md"]) + r = run_check(project_roots=root) + self.assertEqual(r.returncode, 1) + self.assertIn(".ai", r.stdout) + + def test_a_worktree_or_submodule_is_still_a_project(self): + # In a worktree or submodule, .git is a file pointing at the real + # gitdir rather than a directory, so a -d test skips the project + # silently. + with tempfile.TemporaryDirectory() as root: + with open(os.path.join(root, ".git"), "w") as f: + f.write("gitdir: /somewhere/else\n") + with open(os.path.join(root, ".gitignore"), "w") as f: + f.write(".ai/\n") + r = run_check(project_roots=root) + self.assertEqual(r.returncode, 1) + self.assertIn(".ai", r.stdout) + + def test_an_unreadable_gitignore_is_a_finding_not_a_pass(self): + # grep exits 2 on error and 1 on no-match, so treating both as + # "nothing named" lets an unreadable ignore file pass the project + # silently. + with tempfile.TemporaryDirectory() as root: + self.project(root, [".ai/"]) + os.chmod(os.path.join(root, ".gitignore"), 0o000) + try: + r = run_check(project_roots=root) + finally: + os.chmod(os.path.join(root, ".gitignore"), 0o644) + self.assertEqual(r.returncode, 1, + "an unreadable .gitignore read as nothing to check") + self.assertIn("could not", r.stdout.lower()) + + def test_non_git_dir_skipped(self): + with tempfile.TemporaryDirectory() as root: + with open(os.path.join(root, ".gitignore"), "w") as f: + f.write(".ai/\n") + r = run_check(project_roots=root) + self.assertEqual(r.returncode, 0, r.stdout) + + def test_project_without_gitignore_skipped(self): + with tempfile.TemporaryDirectory() as root: + os.makedirs(os.path.join(root, ".git")) + r = run_check(project_roots=root) + self.assertEqual(r.returncode, 0, r.stdout) + + def test_unrelated_ignore_lines_no_flags(self): + with tempfile.TemporaryDirectory() as root: + self.project(root, ["*.pyc", "node_modules/", "dist/"]) + r = run_check(project_roots=root) + self.assertEqual(r.returncode, 0, r.stdout) + + +class SignalAccount(unittest.TestCase): + # --- Normal cases --------------------------------------------------- + + def test_registered_account_passes(self): + r = run_check(signal_accounts="Number: +15045173983 ...") + self.assertEqual(r.returncode, 0, r.stdout) + + def test_no_account_flags(self): + # The velox case: a wiped registration silently breaks paging for + # the whole fleet, because agent-text relays into this machine. + r = run_check(signal_accounts="") + self.assertEqual(r.returncode, 1) + self.assertIn("signal", r.stdout.lower()) + + # --- Error cases ---------------------------------------------------- + + def test_missing_binary_flags(self): + r = run_check(signal_accounts="MISSING") + self.assertEqual(r.returncode, 1) + self.assertIn("signal-cli", r.stdout) + + def test_a_stray_signal_missing_in_the_environment_is_ignored(self): + # The script's own internal flag must not be settable from outside, + # or a caller's unrelated variable turns a registered account into a + # "signal-cli is not installed" finding. + env = dict(os.environ) + env.update({"PRC_FAILED_UNITS": "", "PRC_UNIT_STATES": "", + "PRC_LOCAL_SCAN_ROOTS": "", "PRC_PROJECT_ROOTS": "", + "PRC_NTP_SOURCES": "162.159.200.1", + "PRC_SIGNAL_ACCOUNTS": "+15045551234", + "PRC_IDLE_DAEMON": "4242", + "PRC_REPO_REMOTES": "", + "PRC_UNITS_EXPECTED_DISABLED": "", + "signal_missing": "1"}) + r = subprocess.run(["sh", CHECK], capture_output=True, text=True, + timeout=30, env=env) + self.assertEqual(r.returncode, 0, r.stdout) + + def test_a_stray_idle_absent_in_the_environment_is_ignored(self): + # Same class as the flag above, and the dangerous direction: an + # inherited idle_absent=1 would make check 7 skip silently and report + # ok on a machine that cannot sleep, which is the exact false pass the + # check exists to prevent. + env = dict(os.environ) + env.update({"PRC_FAILED_UNITS": "", "PRC_UNIT_STATES": "", + "PRC_LOCAL_SCAN_ROOTS": "", "PRC_PROJECT_ROOTS": "", + "PRC_NTP_SOURCES": "162.159.200.1", + "PRC_SIGNAL_ACCOUNTS": "+15045551234", + "PRC_IDLE_DAEMON": "", + "idle_absent": "1"}) + r = subprocess.run(["sh", CHECK], capture_output=True, text=True, + timeout=30, env=env) + self.assertEqual(r.returncode, 1, r.stdout) + self.assertIn("hypridle is installed but not running", r.stdout) + + +class ProbeFailure(unittest.TestCase): + """A probe that could not run must never read as a clean check. + + This is the defect the whole script exists to catch, so it would be the + worst possible place to have it. `systemctl --user` exits 1 with empty + output when there is no user bus -- over ssh, from cron, under sudo, or on + a TTY before the graphical session starts. Reading that as "no failed + units" reports a machine as healthy precisely when nothing can be checked. + """ + + def unset(self, *names): + """Run with the named seams unset, so the real probes execute.""" + env = dict(os.environ) + env.update({"PRC_FAILED_UNITS": "", "PRC_UNIT_STATES": "", + "PRC_LOCAL_SCAN_ROOTS": "", "PRC_PROJECT_ROOTS": "", + "PRC_NTP_SOURCES": "162.159.200.1", + "PRC_SIGNAL_ACCOUNTS": "+15045551234"}) + for n in names: + env.pop(n, None) + env["XDG_RUNTIME_DIR"] = "/nonexistent-runtime-dir" + return subprocess.run(["sh", CHECK], capture_output=True, text=True, + timeout=30, env=env) + + # --- Error cases ---------------------------------------------------- + + def test_unreachable_user_bus_is_a_finding_not_a_pass(self): + r = self.unset("PRC_FAILED_UNITS") + self.assertEqual(r.returncode, 1, + "a failed probe reported the machine as clean") + self.assertIn("could not", r.stdout.lower()) + + def test_unreachable_user_bus_fails_the_unit_state_check_too(self): + r = self.unset("PRC_UNIT_STATES") + self.assertEqual(r.returncode, 1, + "a failed probe reported the machine as clean") + + def test_an_unusable_tmpdir_is_a_finding_not_a_pass(self): + # Every check stages its input through a temp file. If that write + # fails, each loop reads nothing and every check comes back clean -- + # with real findings passed in. + env = dict(os.environ) + env.update({"PRC_FAILED_UNITS": "user:calendar-sync.service", + "PRC_UNIT_STATES": "", "PRC_LOCAL_SCAN_ROOTS": "", + "PRC_PROJECT_ROOTS": "", + "PRC_SIGNAL_ACCOUNTS": "+15045551234", + "PRC_IDLE_DAEMON": "4242", + "PRC_REPO_REMOTES": "", + "PRC_UNITS_EXPECTED_DISABLED": "", + "TMPDIR": "/nonexistent-tmp-dir"}) + r = subprocess.run(["sh", CHECK], capture_output=True, text=True, + timeout=30, env=env) + self.assertEqual(r.returncode, 1, + "an unwritable TMPDIR swallowed a real finding") + + +class RealUnitDirEnumeration(unittest.TestCase): + """Check 2's unseamed path, where the unit files are read off disk. + + The PRC_UNIT_STATES seam skips this enumeration entirely, so a defect in + it survives every seamed test. That is where the dangling-stow-link case + lives, and a dangling stow link is precisely the requirement's headline + example of a unit file that LOOKED fine. + """ + + def run_real(self, config_home): + env = dict(os.environ) + env.update({"PRC_FAILED_UNITS": "", "PRC_LOCAL_SCAN_ROOTS": "", + "PRC_PROJECT_ROOTS": "", + "PRC_SIGNAL_ACCOUNTS": "+15045551234", + "PRC_IDLE_DAEMON": "4242", + "PRC_REPO_REMOTES": "", + "PRC_UNITS_EXPECTED_DISABLED": "", + "XDG_CONFIG_HOME": config_home}) + env.pop("PRC_UNIT_STATES", None) + return subprocess.run(["sh", CHECK], capture_output=True, text=True, + timeout=30, env=env) + + # --- Boundary cases ------------------------------------------------- + + def test_a_dangling_stow_link_is_enumerated_not_skipped(self): + with tempfile.TemporaryDirectory() as home: + unit_dir = os.path.join(home, "systemd", "user") + os.makedirs(unit_dir) + os.symlink("/nonexistent/stow/roam-sync.timer", + os.path.join(unit_dir, "roam-sync.timer")) + r = self.run_real(home) + self.assertEqual(r.returncode, 1, + "a dangling stow link read as nothing to check") + self.assertIn("roam-sync.timer", r.stdout) + self.assertIn("missing target", r.stdout) + + # --- Error cases ---------------------------------------------------- + + def test_a_missing_unit_directory_is_a_finding(self): + with tempfile.TemporaryDirectory() as home: + r = self.run_real(home) + self.assertEqual(r.returncode, 1) + self.assertIn("no user unit directory", r.stdout) + + +class WedgedSystemctl(unittest.TestCase): + """A systemd manager that never answers must not hang the check. + + Seen live on velox 2026-08-17: the user manager spun at 96% CPU with + `is-enabled`, `cat`, and `list-unit-files` all hanging while `list-units` + still returned. Unbounded, the check stops at the first unit and never + runs checks 3 through 5, so the machine most in need of checking is the + one it reports nothing about. + """ + + def run_with_fake(self, script_body, timeout_s="1"): + """Run against a fake systemctl, with check 2's unit dir empty. + + Pointing XDG_CONFIG_HOME at an empty directory keeps check 2 from + making one call per real unit, so the test measures the bound rather + than the size of this machine's unit directory. + """ + with tempfile.TemporaryDirectory() as d: + fake = os.path.join(d, "systemctl") + with open(fake, "w") as f: + f.write(script_body) + os.chmod(fake, 0o755) + env = dict(os.environ) + env.update({"PRC_LOCAL_SCAN_ROOTS": "", "PRC_PROJECT_ROOTS": "", + "PRC_SIGNAL_ACCOUNTS": "+15045551234", + "PRC_IDLE_DAEMON": "4242", + "PRC_REPO_REMOTES": "", + "PRC_UNITS_EXPECTED_DISABLED": "", + "PRC_SYSTEMCTL": fake, + "PRC_SYSTEMCTL_TIMEOUT": timeout_s, + "XDG_CONFIG_HOME": d}) + env.pop("PRC_FAILED_UNITS", None) + env.pop("PRC_UNIT_STATES", None) + start = time.monotonic() + r = subprocess.run(["sh", CHECK], capture_output=True, + text=True, timeout=60, env=env) + return r, time.monotonic() - start + + # --- Error cases ---------------------------------------------------- + + def test_a_hanging_systemctl_is_bounded_and_reported(self): + r, _ = self.run_with_fake("#!/bin/sh\nsleep 300\n") + self.assertEqual(r.returncode, 1) + self.assertIn("could not query user units", r.stdout) + # The run must reach the end rather than stopping at the first call. + self.assertIn("check 8/8", r.stdout) + + def test_a_hanging_systemctl_does_not_stall_the_whole_run(self): + # The fake sleeps 8s against a 1s bound, so a bounded run lands near + # 2s (two calls) and an unbounded one near 16s. Deliberately short + # enough that losing the bound fails this assertion in seconds rather + # than hitting the subprocess ceiling a minute later -- a regression + # nobody waits out is a regression nobody catches. + _, elapsed = self.run_with_fake("#!/bin/sh\nsleep 8\n") + self.assertLess(elapsed, 6, + "the run was not bounded by PRC_SYSTEMCTL_TIMEOUT") + + +class Reporting(unittest.TestCase): + # --- Normal cases --------------------------------------------------- + + def test_findings_counted_in_summary(self): + # Assert the count in the summary line specifically. A bare + # assertIn("2") passes on the always-present "check 2/5" text, so it + # stays green even when the counter is arithmetically wrong. + r = run_check(failed_units="user:a.service\nuser:b.service", + unit_states="c.timer disabled") + self.assertEqual(r.returncode, 1) + summary = r.stdout.strip().splitlines()[-1] + self.assertEqual(summary, "3 finding(s) across 8 checks") + + def test_the_summary_count_tracks_every_check(self): + # One finding from each of the five, so a counter that drops or + # double-counts any single check shows up here. + with tempfile.TemporaryDirectory() as scan, \ + tempfile.TemporaryDirectory() as proj: + open(os.path.join(scan, "orphan.example"), "w").close() + os.makedirs(os.path.join(proj, ".git")) + with open(os.path.join(proj, ".gitignore"), "w") as f: + f.write(".ai/\n") + r = run_check(failed_units="user:a.service", + unit_states="b.timer disabled", + local_roots=scan, project_roots=proj, + signal_accounts="") + summary = r.stdout.strip().splitlines()[-1] + self.assertEqual(summary, "5 finding(s) across 8 checks") + + def test_help_exits_zero(self): + r = subprocess.run(["sh", CHECK, "--help"], + capture_output=True, text=True, timeout=10) + self.assertEqual(r.returncode, 0) + self.assertIn("post-rebuild-check", r.stdout) + + # --- Error cases ---------------------------------------------------- + + def test_unknown_flag_errors(self): + r = subprocess.run(["sh", CHECK, "--bogus"], + capture_output=True, text=True, timeout=10) + self.assertNotEqual(r.returncode, 0) + + +if __name__ == "__main__": + unittest.main() @@ -1,6 +1,7 @@ #+TITLE: ArchSetup Tasks #+AUTHOR: Craig Jennings #+DATE: 2026-02-14 +#+PRIORITIES: A D D * Archsetup Priority Scheme @@ -45,58 +46,944 @@ below): input-side-spec.org (DRAFT, four decisions open). * Archsetup Open Work -** TODO [#A] Reseat velox input-cover ribbon — phantom power button :bug:velox:hardware: -DEADLINE: <2026-08-14 Fri> +** TODO [#C] Closed tasks pinned to the agenda by a live planning line :chore:quick:solo: :PROPERTIES: -:CREATED: [2026-08-13 Thu] -:LAST_REVIEWED: 2026-08-13 +:LAST_REVIEWED: 2026-09-25 :END: -Machine off, lift the input cover (Framework QR-guided procedure, 5 -fasteners), reseat its ribbon connector to the mainboard — disturbed in the -2026-08-13 board swap. Root cause of every "mystery reboot" that day: -chassis flex (flash-drive touch, ethernet bump, lid partially lowered) -fired phantom power-button presses — journalctl -b -1 showed "Power key -pressed short." → orderly logind poweroff, then the glitching button -powered it back on. While in there, reseat the USB expansion cards too — -the flaky slot (two hard resets, one no-enumeration) is likely the same -flex problem. -THIRD SYMPTOM (2026-08-13 evening): touchpad delivers ZERO input events — -15s synchronized libinput debug-events capture while swiping caught -nothing, though i2c enumeration and a driver rebind handshake are clean. -Signature of a dead interrupt line on the same ribbon. Keyboard + power -LED lines work; BT mouse is the interim pointer. -ESCALATED 2026-08-13 21:00: a fourth event killed the machine THROUGH the -shield. Previous boot's journal ends mid-line (tailscaled chatter) with no -shutdown sequence at all — a hard power cut, not logind acting. So the -glitch now reaches the EC/hardware power path, which no software setting -can intercept. The reseat is the only fix, and this is a -lose-work-without-warning failure mode, not an inconvenience. -Interim shield (already live): /etc/systemd/logind.conf.d/powerkey.conf -sets HandlePowerKey=ignore — phantom presses log but do nothing; EC-level -10s hold still force-cuts. Consider keeping it even after the repair. -Verify after reseat: flex the chassis edges + partially lower the lid, then -grep the journal for new "Power key pressed" lines — zero means fixed. -Must be done before the Sunday flight — a phantom press mid-travel with the -shield on is survivable, but the connector should not be trusted at 30,000 -feet on the loose setting. -** DOING [#A] Velox reinstall — DR test of archangel + archsetup :velox:chore: -DEADLINE: <2026-08-15 Sat> +Four closed tasks in the Resolved section kept the SCHEDULED or DEADLINE they +carried while open, so org still renders them on the agenda as weeks overdue. Find +them with the grep below rather than by line number — an insertion anywhere above +shifts every number, and the first version of this task carried four that were +already stale when I wrote them: + +#+begin_src sh :results output +grep -nE '^(CLOSED|SCHEDULED|DEADLINE):.*(CLOSED|SCHEDULED|DEADLINE):' todo.org +#+end_src + +My org config sets =org-agenda-skip-scheduled-if-done= to nil, so a terminal +keyword doesn't suppress them — only removing the planning line does. + +This is the top-level counterpart to the rule that strips the planning line from +a dated sub-task entry. An interactive close stamps =CLOSED:= and leaves any +pre-existing =SCHEDULED:= in place, which is how all four survived. + +Fix: delete the SCHEDULED/DEADLINE token from each line the grep returns. At =**= +keep the =CLOSED:= cookie; at =***= and deeper delete the whole planning line, +CLOSED included, since a dated log heading carries its date in the heading. All +four current instances are =**=, but the task is built to be run later, which is +when a deeper one could appear. Verify by re-running the grep and getting no +output. + +The pattern is order-independent on purpose. Matching =CLOSED:.*SCHEDULED:= would +miss a planning line written the other way round and then report clean over an +instance it never looked at, and the =^= anchor is what stops the grep matching +its own source line. + +Found 2026-09-25 while fixing a fifth instance that a review caught in the +then-uncommitted diff. That one is fixed; these four predate it and were left +out so the commit didn't grow a second concern. + +** TODO [#B] velox's USB hub tears down and rebuilds under the Jabra :bug:velox:hardware:audio: +:PROPERTIES: +:LAST_REVIEWED: 2026-09-25 +:END: +The hub at usb 3-2 and whatever hangs off 3-2.1 disconnect and re-enumerate +together, near-daily, and it takes the call audio with it. When the Jabra +Speak2 75 comes back as a new device number, Zoom doesn't follow it, so output +goes nowhere while the device name and every setting still look correct. + +Grading: Major severity (audio dies mid-call and needs a manual re-pick; no +data loss, and a workaround exists) x most-users-frequently (near-daily across +the whole retained journal, on the primary call path) = P2 = [#B]. + +Evidence, 2026-09-25 from velox's persistent journal: +- Today: 11:06:01 the Jabra (0b0e:24ef) enumerates on 3-2.1 as device 5; + 11:06:02 both it and parent hub 3-2 (device 4) disconnect; 11:06:03 the hub + returns as device 6 and the Jabra as 7. One second after plug-in. +- The same 3-2 / 3-2.1 pair drops on 09-18, 09-20, 09-21 (five times between + 08:49 and 08:53), 09-21 18:08, 09-22 (four), 09-23, 09-24 (two), 09-25. +- Companion lines in today's window: "5:0: failed to get current value for + ch 0 (-22)", "cannot get min/max values for control 2 (id 5)", and + "ucsi_acpi USBC000:00: unknown error 256". + +The parent hub is 0a12:4010 — a dock or dongle, not the headset — so the hub is +the likelier fault and the Jabra the casualty. Some of the listed drops are +plausibly me unplugging the dock at the end of a day; the 09-21 cluster of five +inside five minutes and today's one-second-after-plug-in re-enumeration are not. + +Not :solo: — splitting hub from headset needs hardware I have to move: the Jabra +on a direct port with no dock, a different dock or cable, and the dock with +something else on it. The log work is done. + +*** 2026-09-25 Fri @ 12:30:00 -0400 Work's read: bypass the dock, which runs the experiment for free +Work landed the same conclusion about the hub being the fault rather than the +Jabra, and drew the better practical consequence from it. Both call paths now +have a measured failure mode — flaky hub on wired, flaky HFP on bluetooth — so +choosing between them is the wrong frame. Plugging the Jabra straight into a +laptop port avoids both. + +That also collapses the hardware isolation this task is waiting on. If the direct +port holds for a few days of real calls, the hub is implicated and the headset is +cleared, with no deliberate test to run — just using it is the experiment. If it +drops anyway, the fault is downstream of the hub and this task's scope changes. +Try the direct port first and read the result off the journal. + +Cross-boot queries need =journalctl _TRANSPORT=kernel= with no =-b=. Plain +=journalctl -k= implies =-b= and silently scopes to the current boot, which is +what made this look like a single event with no baseline. + +** TODO [#B] xhci on velox refuses D3hot thousands of times per boot :bug:velox: +:PROPERTIES: +:LAST_REVIEWED: 2026-09-25 +:END: +=xhci_hcd 0000:c3:00.0: Refused to change power state from D0 to D3hot=, at a +roughly fixed rate all session, every boot. The controller never reaches D3hot, +so it holds D0 for the life of the boot — a power-management failure and a +plausible battery cost on a laptop. + +Grading: Minor severity (nothing the user does fails; log noise plus a suspected +but unmeasured power cost) x every-boot-every-time = P2 = [#B]. Regrade to Major +if the drain turns out to be measurable — grading the being-in-it, a controller +pinned in D0 costs continuously rather than in a bounded trickle. + +Counts across the five retained boots (oldest to newest): 15772, 1859, 16644, +9987, 5747. Today's 5747 is the lowest of the set, not an anomaly. + +Lead, not a conclusion: the installer puts TLP on every battery machine and TLP +owns USB power policy. Check its USB autosuspend handling against both these +refusals and the hub instability in the task above — they may share a cause. + +Not :solo: — the diagnosis half is mine to run, but deciding whether to change +velox's power policy is a preference call about battery versus device stability, +so it needs your answer before anything is written. + +** TODO [#B] The installer never makes the journal persistent :bug:solo:quick: +:PROPERTIES: +:LAST_REVIEWED: 2026-09-25 +:END: +=configure_encrypted_autologin= writes =/etc/systemd/journald.conf.d/ +retention.conf= with =SystemMaxUse=500M= and stops there (archsetup:3624). +=Storage== is left at its default of =auto=, which is persistent only when +=/var/log/journal= already exists — and on a fresh Arch install it does not. So +a machine this installer builds keeps no journal across a reboot, and any +post-incident question that spans a boot is unanswerable on it. + +Grading: Minor severity (the machine works; what's lost is the ability to +diagnose across a reboot) x every-fresh-install = P2 = [#B]. Not [#A] because no +live machine is impaired — see below. + +Both daily drivers already carry a hand-written =persistent.conf= with +=Storage=persistent=, so this bites only future installs. velox's is dated +2026-08-13 17:48, an hour before the installer's own journald block ran at +18:51 on rebuild day, which is how I know the installer didn't write it. Ratio +has the same pair of files. + +Fix: add =Storage=persistent= to the block that already writes retention.conf, +with a test pinning it beside the SystemMaxUse assertion. One file, and it +belongs in the config the installer already owns rather than a second drop-in. + +Found while verifying a work handoff that had concluded velox kept no baseline — +it did, because of the hand-written file. A fresh machine wouldn't have. +** TODO [#C] Airplane panel key follow-ups from the f977418 commit :refactor:dotfiles:quick: +:PROPERTIES: +:LAST_REVIEWED: 2026-09-23 +:END: +Four items left standing when the airplane keybind moved into the net panel +(dotfiles f977418, 2026-09-23). None blocked the commit. archsetup drives the +dotfiles work end to end per the standing rule in notes.org. + +- The AIRPLANE console key is shown on desktops (net/src/net/gui.py:576). On + ratio the flow is: confirm the prompt, then "airplane mode isn't available + on this machine". PanelModel already carries has_wifi and has_speedtest + capability flags; a battery or laptop flag could desensitize the key before + the question is asked. This one is a design call, which is why the task + isn't :solo:. +- "LEAVE AIRPLANE" is hardcoded in net/src/net/classify.py:52 and diag.py:84 + instead of shared from viewmodel.AIRPLANE_LEAVE_KEY. Tests pin the coupling, + so a rename would surface, but one source is cleaner. +- tests/net/panel_smoke.py:57 checks the DOCTOR and SPEED TEST console keys + only. Add AIRPLANE, and consider a smoke step that opens the confirm dialog + and cancels. Needs a compositor to run. +- Two pre-existing comments on untouched lines still say "keybind": + hyprland/.local/bin/airplane-mode:38 and + tests/airplane-mode/test_airplane_mode.py:353. + +** TODO [#C] Post-install check that every mimeapps.list handler exists :feature:quick:solo: +:PROPERTIES: +:CREATED: [2026-09-22 Tue] +:LAST_REVIEWED: 2026-09-22 +:END: +Work's optional ask from the 2026-09-18 libreoffice handoff, kept because the +failure it catches is silent. + +The dotfiles =common/.config/mimeapps.list= "[Default Applications]" block +names a .desktop file per type. When the package behind one is missing, +xdg-mime does not error — it falls through to the next application claiming +that type. On velox that meant every .pptx opened PowerPoint inside the +Windows VM for a month, and the only symptom was that it felt slow. + +Add a post-install check that reads every .desktop named in that block and +reports the ones absent from the machine. Declaring the packages (done +2026-09-22 for libreoffice-fresh, imv and git-lfs) fixes today's instance; this +catches the next one, including a handler the dotfiles add later. + +Natural home is =scripts/post-rebuild-check=, which already runs this shape of +verification. + +** TODO [#B] Own the meeting transcription service install :feature:velox:ratio:tooling: +:PROPERTIES: +:CREATED: [2026-09-22 Tue] +:LAST_REVIEWED: 2026-09-22 +:END: +Work built a self-hosted meeting transcription service (whisper.cpp plus +pyannote diarization) that has run on ratio and velox since 2026-09-17, and +handed the service side here on 2026-09-19 because it is machine setup rather +than application work. Installed by hand on both machines today; nothing +reinstalls it. + +The bundle is in [[file:working/meeting-transcription-service/][working/meeting-transcription-service/]], with work's handoff note +beside it. Scanned for credentials on arrival: clean. + +What the install has to provide per machine: +- =~/.local/share/pyannote-diarize/.venv= — Python 3.12, CPU torch, + pyannote.audio 4.0.7, about 1.3 GB, built with uv. +- =~/.local/share/whisper-models/ggml-large-v3-turbo-q5_0.bin=, plus + whisper-cpp itself. +- The three =src/= scripts where the units expect them, and both user units + enabled with linger on so the path unit fires without a login session. + +Open decisions before this is buildable, which is why it isn't =:solo:=: +- Where the code lives in this repo — a new top-level dir, or under =scripts/=. +- How the Hugging Face step is handled. Accepting the pyannote model terms and + caching the model is one-time, online, and interactive. The token is a + credential and this repo is anonymously cloneable, so it cannot be committed + here; the installer can only prompt for it or read it from the private + secrets path. + +Known rough edge work flagged: when two runs overlap the second finds the lock +held, the worker returns silently, and the client reports "finished without +producing a transcript". Rerunning works. The message should name the lock. + +** TODO [#C] Orchestrator sequence pin misses an added step :test:quick:solo: +:PROPERTIES: +:CREATED: [2026-09-17 Thu] +:LAST_REVIEWED: 2026-09-17 +:END: +Noticed during the 2026-08-08 pre-vacation sweep and never filed; confirmed +still open 2026-09-17. + +=tests/installer-steps/test_orchestrators.py= defines recorder stubs only for +the sub-steps it expects. A step added to an orchestrator without updating the +pin calls an undefined function: bash prints "command not found" to stderr, +nothing reaches stdout, and the recorded sequence still matches. The +=returncode= assertion sees only the last call's status, so the test passes +unless the new step happens to be last. The module docstring claims it catches +"a dropped, added, or reordered" call; it catches drops and reorders. + +Fix: define =command_not_found_handle() { echo "UNSTUBBED:$1"; }= in the +generated script so an unstubbed call lands in stdout and fails the equality, +plus a test proving it (the file already has one of those for guarded helpers, +=test_the_check_would_notice_a_missing_call=). + +Grading: Minor severity (the VM harness still runs the real steps; only the +fast pin is blind) x some developers sometimes (fires only when a step is added +without updating the pin) = P3 = [#C]. +** TODO [#B] Swap velox's MT7925 for an Intel AX210 :chore:velox:hardware: +:PROPERTIES: +:CREATED: [2026-09-16 Wed] +:LAST_REVIEWED: 2026-09-16 +:END: +Ordered from the Framework Marketplace on 2026-09-16; waiting on delivery. +An Intel AX210 will replace velox's MediaTek MT7925 +(RZ717, Filogic 360) in the M.2 2230 slot. The ordered part is AX210.NGWG.NV +(the .NV suffix means no vPro). A vPro card won't work in the Framework, so +check the part number on the card when it arrives. + +Why: HFP call audio on the MT7925 fails in firmware. It sends zero-filled SCO +frames on sentinel handle 0x0E00, and the kernel logs "SCO packet for unknown +connection handle 3584". It's pending upstream with no fix. Three headsets fail +on velox. The MT7925 is also step 4 of the hibernate-freeze mitigations. Losing +WiFi 7 is fine; the AX210 does WiFi 6E and BT 5.3. linux-firmware-intel is +already installed on velox, and the installer's firmware trim only runs on +Intel-CPU Framework 13s, so archsetup needs no change. See the Bluetooth wedge +bug's 2026-09-16 entries. + +Before the swap: note the saved WiFi profiles (NetworkManager keeps them, they +aren't tied to the card), and power off fully (not hibernate). Keep the MT7925 +in case the AX210 is a dud. + +After the swap, verify: +- lspci -k shows the AX210 on iwlwifi; the Bluetooth controller enumerates as + Intel (btintel) and hci0 is up. +- WiFi joins a saved network, and net status/probe/diagnose read it correctly. + The net code has an nl80211 path written for the mt7925, so check that the + signal line still shows. +- Re-pair the Sonys and any other Bluetooth devices, since the bonds belong to + the old controller's address. Use the agent-backed pairing recipe in the KB, + then count key sections to confirm a real bond. +- A real call load of several minutes in HFP, counting "Failure in Bluetooth + audio transport" and kernel "unknown connection handle" errors. Target zero. + Note the Intel risk: AX201/AX211 had an eSCO handle-reuse bug with WirePlumber + 0.5.17 after repeated profile switching. +- Suspend/resume and one hibernate cycle with the new card. +Then close the Bluetooth wedge bug, or regrade it, and update the MT7925 step +in the hibernate mitigations. + +** DOING [#C] Net doctor takes the DNS ladder on a DoT-blocking captive portal :bug:dotfiles:network: +:PROPERTIES: +:CREATED: [2026-09-14 Mon] +:LAST_REVIEWED: 2026-09-14 +:END: +An open captive-portal network that filters TCP 853 before login, 2026-09-14. +The network's resolver +172.20.0.1 answers plain UDP 53, but TCP 853 is filtered until you log in. With +resolved pinned to DNSOverTLS=yes, every lookup on the WiFi link times out. +The probe read "no-internet" instead of "captive", and the doctor ran +repair:dns-test (cleanup-unverified), then repair:dns-override (fail, +reverted), twice (08:32, 08:35). It never offered portal-login. Craig had to +tether to his phone. A dns-override to 1.1.1.1 can't pass a walled garden +anyway. The same flow worked on a different portal on 09-13. + +Once DNS resolves (through the tether), the pinned probe sees the portal and +diagnose recommends net portal. + +Also unexplained: the probe log reads "online" on wlp192s0 from 08:43 to +09:50, then flaps captive/online until 09:52, while a pinned curl shows the +portal still intercepting. Check whether the probe follows the default route +(the tether) while labeling the WiFi interface. + +Remaining before close: a live doctor run on that network with the tether +unplugged (the fix is committed and live on velox; ratio gets it on its next +dotfiles pull). + +Follow-ups the review found in existing code (not in this fix): +- repair.py _extract_portal_url (~603-619) rejects any URL containing a + detection-host name anywhere, query string included. This portal echoes the + requested URL in its OS= parameter, so the portal URL the probe found gets + thrown away, and repair_portal_login (~687) falls back to opening a trigger + page (neverssl.com). The login should still work through interception, but + the found URL is lost. Fix: match detection hosts against the hostname only, + or pass the probe's portal_url through. +- probe.py CLOUDFLARE_HOSTS: a controller that hosts its login page on 1.1.1.1 + itself (older Cisco wireless, https://1.1.1.1/login.html) reads as + reachable, in both the followed-redirect and body branches. + +Grading: Major severity (no path online on such a network without knowing to +run net portal by hand; the doctor's own repairs can't succeed) × some users, +sometimes (networks that filter 853 before login; another portal worked) = +P3 = [#C]. +*** 2026-09-14 Mon @ 12:12:48 -0400 Found the root cause in the probe's IP fallback, fixed it test-first +My first guess (extend the tunnel-dot shape test to WiFi) was wrong. With DNS +dead, run_probe falls back to http://1.1.1.1/ over the WiFi. The portal answers that +literal with HTTP 200 and a meta refresh, so curl's effective URL stays on +1.1.1.1. classify_ip only checked the effective host, and called the page +"reachable". So the probe said no-internet, diagnose emitted no portal row, and +the classifier fell through to dns-test. The hostname path already reads +body-level redirects; the IP path didn't. + +Fix in dotfiles net/src/net/probe.py: classify_ip reads the body with +extract_portal_url and calls it captive when the target host isn't Cloudflare. +Tests in tests/net/test_net.py, all using a placeholder-scrubbed portal +fixture: classify_ip normal, boundary, and error cases (Cloudflare body +redirect, v6 literal, refresh with no URL, javascript: target, followed +redirect wins); probe_and_cache end to end; diagnose emitting the portal row; +doctor --fix running portal-login instead of dns-test/dns-override. Red with +four failures, green at 991. The live IP probe over that network now +classifies it captive and finds the portal URL. +*** 2026-09-14 Mon @ 12:22:23 -0400 Committed and pushed the fix as dotfiles b2688e4 +An isolated review approved it. I added its two Minor test gaps (JS redirect, +malformed target) and made the comments vendor-generic. Full make test: forked +and shared runs each 4401 OK, faces 171 pass. I sent an FYI to the dotfiles +inbox. The net CLI runs from the repo, so velox has the fix now. + +** TODO [#B] Bluetooth audio link drops wedge the PipeWire graph on velox :bug:velox:audio: +:PROPERTIES: +:CREATED: [2026-09-14 Mon] +:LAST_REVIEWED: 2026-09-14 +:END: +First seen 2026-09-14 with the Jabra Speak2 55 MS on velox. +There are two faults, and the second makes the first much worse. + +1. The link drops. "Failure in Bluetooth audio transport" hit six times between + 10:09 and 11:10: on A2DP (sep1/fd0) and HFP (fd60), at 10:09, three times + around 10:10, 10:54, and 11:10. Each one matches a kernel "Bluetooth: hci0: + ACL (or SCO) packet for unknown connection handle" line. The controller is + the MediaTek MT7925 (0e8d:7925, btusb; WiFi firmware build 20260813). + Suspects I haven't separated yet: MT7925 btusb firmware or driver, range or + 2.4 GHz coexistence with the WiFi on the same chip, or A2DP/HFP profile + switching when the mic opens. +2. The graph wedges. After the drops, the Jabra's nodes sat in error and every + client round trip hung: pactl info, wpctl status/inspect, pw-dump, and the + mic-mute key (four stuck wpctl set-mute processes). pw-cli info 0 still + answered, and no thread was spinning. Only a restart of pipewire, + pipewire-pulse, and wireplumber cleared it. Versions: pipewire 1.6.8, + wireplumber 0.5.17, bluez 5.87, kernel 6.18 LTS. + +Grading: Major severity (all audio control is dead until a manual service +restart; the restart is the only workaround) × most users, frequently (six +drops in about an hour of use on the one Bluetooth speaker in play) = P2 = +[#B]. One session of data; re-read the frequency row if it doesn't recur. + +Next: on a recurrence, capture btmon and the kernel log across a drop, and +check whether the wedge follows only HFP use. Try the same speaker on ratio to +split controller from device. Look upstream for MT7925 "unknown connection +handle" reports and for wireplumber bluez nodes stuck in error. + +*** 2026-09-15 Tue @ 14:12:11 -0400 Reproduced the link drop with a second speaker, a Speak2 75 +The replacement Jabra Speak2 75 dropped the same way while +paired straight to velox's MT7925: "Failure in Bluetooth audio transport" at +12:47:20 (A2DP, sep3/fd0) and 12:47:49 (HFP, fd62), with a kernel "hci0: ACL +packet for unknown connection handle" line at 12:47:21. The graph didn't wedge +this time; pactl, wpctl, and pw-dump all kept answering. A second device with +the same signature points further toward the controller side than the speaker. +I removed velox's pairings for both Jabras, and the 75 now runs over its USB +cable (0b0e:24ef), which works for playback and mic. Its Link 390 dongle +(0b0e:2e56) enumerates but was never linked to the speaker, and pairing one +needs Jabra Direct (Windows/Mac only). +*** 2026-09-16 Wed @ 13:29:34 -0400 A pairing that never bonded looks like this bug, so rule it out first +A one-shot bluetoothctl pair, with the adapter at Pairable: no (velox's normal +state), can leave a bond file with no key section. It shows up as transport and +AVDTP failures and the device dropping the link. The 09-15 log records the +Speak2 75 as Bonded: yes, so this doesn't explain the 75's drops. The 55 MS's +bond state was never checked, and its pairing is gone now. On a recurrence, +count the device's key sections first (the one-liner in the node). +[[id:f23b7085-c35e-43c1-ae02-9e9c67e3848b][bluetoothctl one-shot pair can complete without bonding]] +*** 2026-09-16 Wed @ 13:43:40 -0400 Third device, same kernel handle errors; pairing and settling ruled out +The Sony WF-1000XM6 (bonded, one LinkKey) failed the same way today, and the +evidence now points at the MT7925's SCO path. + +Work's session: transport failures from 10:44 to 10:50, about 90 clean seconds, +then six more in two minutes (11:01 to 11:04) during a Meet call, with the buds +beeping on each disconnect. The kernel logged 23 "SCO packet for unknown +connection handle" and 38 "ACL packet for unknown connection handle" errors in +exactly those two windows. mpv played over A2DP for twenty minutes with no +failures. Every failure came after the card switched to a headset profile. SCO +carries HFP voice. In work's run, a 35 s full-duplex hold of each HFP codec +(CVSD, mSBC, LC3-SWB) passed with zero failures while real calls failed. The +earlier "link settling after a fresh pair" reading was wrong: the quiet stretch +was just quiet. + +Mine, from 13:23: the card churned between profiles for six minutes, then +pipewire-pulse refused about 200 connections as "too many client application +connections" (13:30:07 to 13:30:10). pactl, wpctl, and pw-cli info 0 all hung, +the same wedge as 09-14, and restarting the three user units cleared it. After +the restart, a 20 s LC3-SWB hold failed: the mic was pure digital zero, both +bluez nodes went to error, and the kernel logged SCO and ACL handle errors. +10 s holds of CVSD and mSBC passed. Then at 13:40:00, :17, and :22, three +headset-profile switches failed the same way while mpv played, each time an +app opened the mic. WirePlumber then had LC3-SWB saved as the headset profile. + +What this rules out: a bad bond (the Sonys are bonded), settling (the failures +came back under load). The codec isn't ruled out. The only failure with a +known codec was LC3-SWB (which passed work's hold), CVSD and mSBC have only +passes on record, and the codec in play during the three 13:40 failures wasn't +captured. Three devices share the signature, and the handle errors come from +the controller. + +Still untested: WiFi/Bluetooth coexistence on the shared MT7925. velox was on a +busy shared WiFi network through every failure. Next test: a real call load +with WiFi off and the network over a USB tether, counting transport failures +and kernel handle errors. Then a btmon capture across a failure. + +Workaround until then: take calls on the Speak2 75 over USB. With autoswitch on, +any app that opens a mic pulls Bluetooth music into HFP and hits the fault. +[[id:eaa85ba0-ad4f-4a38-a27f-c39b879c341f][Measure Bluetooth audio dropouts across the load that fails, not a quiet window]] +*** 2026-09-16 Wed @ 14:42:50 -0400 Codec-filter experiment and a saved-profile trap +Work tried pinning the call codec at about 14:30 with Craig's go-ahead: it left +lc3_swb out of monitor.bluez.properties bluez5.codecs and restarted wireplumber. +The filter does reach HFP. LC3-SWB disappeared, and headset-head-unit became +MSBC. But every A2DP profile, AAC included, vanished too, even though all the +A2DP codec names were listed. Work reverted within a minute and all six profiles +came back. It's unclear whether the config dropped A2DP or the three-second wait +was just too short for re-enumeration. If codec pinning is still wanted after +the AX210 swap, retest on a day without calls and wait longer before judging. +The wireplumber.conf.d directory is empty now, so the revert is on disk. + +The saved-profile trap: WirePlumber saves profiles by name, and the name +headset-head-unit maps to whichever codec is best at load time (LC3-SWB with +the default config, MSBC under the filter). At 14:42 both the card's +default-profile and saved-headset-profile read headset-head-unit, so calls +autoswitch to LC3-SWB. And because default-profile is a headset profile, the +card may come back in HFP (call-quality music) after a reconnect or a +wireplumber restart, until something switches it to a2dp-sink. Nothing changed: +work asked for a hold on velox audio for the rest of 09-16. +*** 2026-09-16 Wed @ 15:08:18 -0400 Matched to an upstream MT7925 firmware report; fix is a new card +The kernel's "SCO packet for unknown connection handle 3584" is handle 0x0E00, +the sentinel handle in an August 2026 linux-bluetooth report: "MT7925 +(0e8d:0717): HFP microphone unusable — firmware delivers zero-filled (e)SCO +frames on sentinel handle 0x0E00" (bluez/bluetooth-next PR #744). It's the same +chip under a different USB ID; velox enumerates as 0e8d:7925. That report +used PipeWire 1.6.8 and MT7925 BT firmware from 20260622 and 20260810, and its +frames were all zeros, matching the zero-filled LC3-SWB mic hold here. The +reporter traced it to the firmware's transparent-mode receive path. A +handle-rewrite patch only proved the payload is zeros. There's no maintainer +reply, nothing merged, and no fix. The reporter also found it depends on the +headset. velox at the time: kernel 6.18.51-lts, linux-firmware-mediatek +20260910, BT firmware built 20260813. + +A related report on Intel AX201/AX211 shows handle reuse after repeated +A2DP/HFP switching with WirePlumber 0.5.17 (velox's version), and the bluetooth +maintainer called that a firmware bug too. Autoswitch churn makes these faults +more likely. + +So the leading explanation is MediaTek firmware, not the headsets and not our +config. Coexistence stays untested, but the upstream match means the WiFi-off +test no longer decides the fix. The fix is replacing the card: an Intel AX210 +is ordered (see the swap task at the top of Open Work). The dotfiles net code +also calls ratio's card an mt7925, so testing on ratio wouldn't separate +controller from headset. That's unconfirmed, because ratio didn't answer ssh. + +Once the AX210 is in, re-run the call-load test there. Close this bug if it's +clean, or regrade it if Intel shows the same signature. + +*** 2026-09-22 Tue @ 01:47:56 -0400 A fourth failure on 09-17, and restarting pipewire-pulse is not a full recovery +From the work session's 09-19 handoff, read back from velox's journal. + +The failure recurred Thursday 2026-09-17 at 10:59:32 EDT: kernel "ACL packet +for unknown connection handle 3837" on boot a440a2a0, with wireplumber logging +"Failure in Bluetooth audio transport" for 58:18:62:AA:62:9D at 10:59:11 and +10:59:32. That boot carried five unknown-handle lines. Daily transport-failure +counts now run 13 on 09-14, 3 on 09-15, 44 on 09-16, and 11 on 09-17 (eight of +them between 08:56 and 08:58). The handle differs each time, which is what the +upstream sentinel-handle report predicts, so this is the same firmware fault +rather than a new one. + +Recovery gap worth knowing before the AX210 lands: restarting pipewire-pulse +strands every Chromium and Electron audio helper process. Each of those apps +stays silent until it is itself restarted, so the service restart alone leaves +the browser and any Electron app on a call dead. Restart the apps too. + +** TODO [#B] Visual separator between adjacent waybar modules :feature:waybar:dotfiles:quick: +:PROPERTIES: +:CREATED: [2026-09-13 Sun] +:LAST_REVIEWED: 2026-09-13 +:END: +Captured by Craig 2026-07-20 and routed here from the roam inbox, then lost +in the processed pile: the wind (weather) value runs straight into the date +with no visual stop, so the wind figure reads as the start of the date. Add +a light separator or spacing between adjacent modules so each one's edge is +unmistakable. Check the current bar first (the mic/PTT merge and the weather +chip grouping landed after the capture), then decide the form: a thin rule, +a dot glyph, or just margin. That choice is mine, so not solo; the CSS itself +is a quick change in the dotfiles waybar stylesheet. + +** TODO [#C] Saving and restoring a window configuration :feature:hyprland:research: +:PROPERTIES: +:CREATED: [2026-09-13 Sun] +:LAST_REVIEWED: 2026-09-13 +:END: +Research idea captured by Craig 2026-07-24 (the one item of that batch that +never got filed): when I want a specific window orientation, I indicate it +and the window-plus-app configuration reappears. What would we need to know +or store to make that happen? Is there another desktop or OS that does it, +what information do they keep, and what are their rules? Explore how far +Hyprland can get, document thoroughly, and review the findings with me before +building anything. The deliverable is a research note under docs/design, so +this is not solo. + +** TODO [#C] maint backup_freshness probe blind to backup_run remedy runs :bug:maint:dotfiles:solo: +:PROPERTIES: +:CREATED: [2026-09-12 Sat] +:LAST_REVIEWED: 2026-09-12 +:END: +From home's 2026-09-12 velox health check. The backup_freshness probe +(dotfiles =maint/src/maint/probes/services.py=) parses the tail of +=/var/log/rsyncshot.log=, which only the root crontab lines append to. The +=backup_run= remedy (=remedies.py=) runs rsyncshot without that redirect, so +right after a successful =maint fix backup_run= (DAILY.0 rotated on truenas +at 21:29) the probe still said "daily 272h ago" and stayed CRITICAL until +the next cron daily landed. + +Fix: make the remedy log the way the cron lines do (append the run to +=/var/log/rsyncshot.log=), so the probe and the remedy read and write the +same record. Reading the snapshot directories instead is the other route, +but they live on truenas, so the probe can't see them without a network +call on the status path. + +Grading: Minor (a false CRITICAL that hides nothing, but persists for days) +x some users sometimes (only after a by-hand remedy) = P3 = [#C]. Solo: the +remedy-logs fix is mechanical and the maint suite covers remedies. Dotfiles +work, which this project carries end to end. + +** TODO [#B] Loopback-only binds in maint's unexpected-listener set :bug:maint:dotfiles: +:PROPERTIES: +:CREATED: [2026-09-12 Sat] +:LAST_REVIEWED: 2026-09-12 +:END: +From home's 2026-09-12 velox health check. The slack-mcp-deepsat container's +=docker-proxy= on =127.0.0.1:13080= is a persistent unexpected-listener warn +on velox (ratio runs the same container). A loopback bind is not a LAN +exposure, and the firewall digest already distinguishes wildcard-bound +listeners from the rest. + +Two routes: treat loopback-only listeners as informational (drop them from +the unexpected set, keep them in the rows view), or curate =docker-proxy= as +expected per machine from the panel's MARK EXPECTED. The first is the durable +fix; the second is the workaround that works today. + +Not solo: whether a loopback port belongs in the signal at all is my call +(a local port is still reachable from a browser), so the design question +comes first. Grading: Minor (a permanent warn that trains me to skim past +the listeners row) x every user every time (both daily drivers run the +container, and the warn never clears) = P2 = [#B]. + +** TODO [#B] Speedtest button cancels an in-flight run :feature:dotfiles:network: +:PROPERTIES: +:CREATED: [2026-09-01 Tue] +:LAST_REVIEWED: 2026-09-01 +:END: + +From the roam inbox (routed 2026-09-01), Craig's words: "pressing the +speedtest button on network admin panel when speedtest is already running +should cancel the speedtest. However, we should leave any numbers on the +display as if the speedtest completed successfully." + +Net panel work lives in ~/.dotfiles; archsetup owns it end-to-end per the +standing rule in notes.org. Distinct from the [#C] speedtest-history task +(that one persists results over time; this one is in-flight cancel +semantics). Behavior is fully specified: second press kills the running +test, display keeps whatever numbers are already shown as a completed +result. [#B]: real improvement to the active panel family, no hard date. + +** TODO [#B] gcalcli in the installer, token carried from the other daily driver :feature:velox:tooling:solo: +:PROPERTIES: +:CREATED: [2026-08-25 Tue] +:LAST_REVIEWED: 2026-08-25 +:END: +From home's 2026-08-25 handoff: velox's 08-13 reinstall left it without +gcalcli, and on 08-21 both calendar write paths on velox were down at once +(the google-calendar MCP with expired tokens, and no gcalcli), so a booked lab +appointment sat uncalendared for three days. The daily-drivers one-time-setup +drift, exactly. + +Part 1 is done (2026-08-25, this session): =pipx install gcalcli==4.5.1= on +velox to match ratio, then ratio's =~/.local/share/gcalcli/{oauth,cache}= +copied over tailscale (=oauth= is a 1 KB pickled google-auth credential, +=chmod 600=). =gcalcli list= on velox returned all six calendars with no +re-consent, so the token is portable between the daily drivers and the OAuth +click-through is not needed when the other machine is reachable. + +Part 2, this task: make the installer do it. +- =pip_install gcalcli= in the tool set beside =pip_install yt-dlp= (archsetup + ~line 3109; =pip_install= wraps =pipx install= as =$username=). Pin or not: + ratio and velox are both 4.5.1; unpinned matches how yt-dlp is installed. +- The credential can't be installed: add a named post-install manual step + "copy =~/.local/share/gcalcli/oauth= from the other daily driver + (=scp <other>:.local/share/gcalcli/oauth ~/.local/share/gcalcli/=, + =chmod 600=), or run =gcalcli init= per + =assets/2026-02-01-gcalcli-setup.org= when neither machine has it." +- A =post-rebuild-check= item: gcalcli on PATH and the oauth file present, so + the drift is caught by the checker rather than by a missed appointment. +- Tests: an installer-steps pytest asserting the tool set carries + =pip_install gcalcli=; a post-rebuild-check test for the new item, both + states. +- When it lands, confirm back to home (=inbox-send home=) so it can retire + its "gcalcli is not installed on velox" notes. + +Grading: feature, no hard date, real improvement to the install = [#B]. +:solo: — build path (installer + checker + tests) and verify path (pytest; +the live proof already exists on velox) with no open decision. + +** TODO [#A] Topgrade guarded-upgrade spec — decisions, review, decomposition :feature:maint:dotfiles: +SCHEDULED: <2026-09-23 Wed> +:PROPERTIES: +:CREATED: [2026-08-25 Tue] +:LAST_REVIEWED: 2026-09-17 +:SPEC_ID: 81cdfd72-db96-43d3-aa03-779878c99f3e +:END: +The waybar maint module's "topgrade freshness" warning never clears: the stamp +is written only when topgrade exits 0, and the =hypr-live-update-guard= +PreTransaction hook (mesa, wayland, hyprland, vulkan-*, nvidia-utils, +xorg-xwayland under a live Hyprland) plus any failing ecosystem step makes that +exit almost unreachable. Diagnosed 2026-08-24/25; the fix is specced, not +hacked, because it spans two repos and the design is contested. + +Spec: [[file:docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org][2026-08-25-topgrade-guarded-upgrade-spec.org]] (DRAFT). + +Four open decisions, all mine to make before the spec can move: +1. Freshness means *state* (a guarded upgrade still un-applied stays stale), + not recency (any run stamps). +2. Primary mechanism is alternative D: an armed boot-time oneshot ordered + before =getty@tty1= (no display manager to order against). +3. The boot run is arch-only (=topgrade --only system=), not the full sweep. +4. The arm flag lives on a persistent path and is one-shot. + +Then: flip the decisions DONE, run spec-review (DRAFT → READY), run +spec-response to decompose the four phases into build tasks here, file the +vNext =[#D]= kernel-reboot item, and commit the spec. + +*** 2026-09-12 Sat @ 23:26:50 -0500 Containers step and initramfs-read findings from the 2026-09-12 velox run +home's velox health check hit another route to the unreachable exit 0 this +spec is about: topgrade exited 1 only because its containers step tried to +=docker pull= locally built images (=cj/telega-server=, =telega-server-glycin=; +ratio has the same shape with its own local images) and got "pull access +denied", so the stamp was applied by hand with =maint stamp topgrade=. Two +fixes: add =containers= to the =--disable= list in the guarded-upgrade's +topgrade invocation, or list the local images under =[containers] +ignored_containers= in topgrade.toml. I lean to disabling the step: pulling +newer images underneath running containers isn't an upgrade path I use, and +an enumerated ignore list rots as images come and go. + +Also for the =kernel-modules-check= gate: reading the initramfs needs root. +The images are 0600, so an unprivileged =lsinitcpio= exits 1 with "Unable +to read file" on stderr and nothing on stdout; piped into =grep -c=, that +empty stdout reads as 0 and looks like a missing module. The gate has to +run as root and check the exit status, not just the count. +*** 2026-09-17 Thu @ 08:59:59 -0400 Re-dated to 09-23; the "four open decisions" list above is stale +The spec is still DRAFT with =Decisions [7/7]=: every decision is made, +including the four listed above. What's left is the isolated spec-review, +DRAFT → READY, and spec-response into build tasks. velox has +=linux-lts 6.18.51 → 6.18.52= pending today, which is the kernel-hold case this +spec exists for, so no plain topgrade on velox until the hold is built or the +kernel update runs as its own session. + +** TODO [#C] post-rebuild-check: probe that Emacs frames come up Wayland-native :feature:emacs:velox:solo:quick: +:PROPERTIES: +:CREATED: [2026-08-25 Tue] +:LAST_REVIEWED: 2026-08-25 +:END: +On 2026-08-24, the first Emacs 31.1 start on velox opened its first frame on +XWayland (=:0=) with the pgtk "unsupported under X" dialog, while every later +frame went to =wayland-1=. The cause was in the Emacs config (a startup buffer +sweep killed =*Warnings*= while 31.1's warnings.el held it for a deferred +display, so the first =make-frame= failed and emacsclient fell back to +=$DISPLAY=); fixed in =.emacs.d= the same night. The trap generalizes: any +first-frame error on a PGTK daemon silently lands the session on X, and nothing +in the post-rebuild pass would notice. + +Design, under the script's fail-closed contract (a probe that cannot run +reports a finding, never a pass): +- Bound every =emacsclient= call with =timeout=, as the =systemctl= calls are. + A daemon stuck in a prompt is a finding, not a hang. +- No daemon (=emacsclient= cannot connect): a visible finding, "not checked: + no Emacs daemon, start Emacs and re-run". Emacs is started on demand here, + so this is the common state right after a rebuild, and the line is the point. +- Daemon up but no GUI frame yet: never skip, and never call + =pgtk-backend-display-class= with no frame (it errors with "Frames are not + in use"). Request one invisible frame through a waiting client in the + background, =timeout 20 emacsclient -c -F '((visibility . nil) (name . + "prc-probe"))'=, so the probe walks the same first-frame path that failed on + 2026-08-24. Then wait, bounded (poll for a frame named =prc-probe= for up to + the same 20 s), before inspecting: the =-e= must not run before the =-c= + has connected. Zero pgtk frames after the request is a finding in its own + right, because a frame request that produced nothing is the first-frame + failure this check hunts. Delete the probe frame after reading. +- Inspect every pgtk frame, not the selected display, and only pgtk frames: + a tty client frame (=emacsclient -t= in tmux) carries =$DISPLAY= as its + display parameter and its terminal is not a display, so it would both trip + the =:0= rule and make =pgtk-backend-display-class= error. + #+begin_src sh + emacsclient -e '(mapcar (lambda (f) (list (frame-parameter f (quote display)) (pgtk-backend-display-class (frame-terminal f)))) (seq-filter (lambda (f) (eq (framep f) (quote pgtk))) (frame-list)))' + #+end_src + Expected: at least one entry, every display equal to =$WAYLAND_DISPLAY=, + every class =GdkWaylandDisplay=. An empty list, a =:0= entry, or a + =GdkX11Display= is a finding. A build without =pgtk-backend-display-class= + is a finding too: the installer installs =emacs-wayland=. +- Limit, stated in the check's output: it sees live frames only. A first X + frame that was already closed is invisible, so this reports the machine's + current state, not its history. The invisible probe frame is created and + never mapped, so it exercises =make-frame= (where 2026-08-24 failed), not + the window-show path. + +Tests alongside the other checks, one per state: no daemon, no frame (probe +frame requested and waited for), probe frame never appears, Wayland-only, a +=:0= frame present, a tty client frame present alongside Wayland frames, hung +daemon, an =*ERROR*= reply from =emacsclient -e= (exit 1, a finding), non-pgtk +build. + +** TODO [#B] Timeline spine test picks the wrong "next" event off Denver :bug:dotfiles:test: +:PROPERTIES: +:CREATED: [2026-08-24 Mon] +:LAST_REVIEWED: 2026-08-24 +:END: +=make test= in dotfiles is red before any of this session's work. Two failures, +both in =settings/faces/timeline-face-spine.test.mjs=: "event bars never leave +the plot" and "exactly one event is marked as next, and it is the soonest ahead". + +NOT the bug =c96a216= fixed. Every =spineRows= call in that file is pinned to +=JUL=, and =scene()= and =EVENTS()= both default to it, so the fixture side is +already clean and the file's own guard test passes. + +TWO THINGS TO SETTLE, and they may be one bug or two: + +1. =timeline-face-spine.js:466= — =const next = timedOnly(events).find((e) => e.s + >= refMs)= takes the first array element starting at or after now, which is + the *soonest* only if =events= is sorted by start time. The test's failure + message is exactly that it is not: a bar ahead of the spine starts at x=1651.2 + while the one marked =event-next= sits at x=2132.8. Either sort before the + find, or use a min-by rather than a find. + +2. Why it is red *here* and presumably green on ratio. The most recent commit to + =timeline-face-spine.js= is =ffe43ab feat(settings): draw home where the + machine is, not where its zone is=. This machine is =America/Denver= (Craig + travelling); the tests pass =home("New Orleans")= explicitly. If a + machine-resolved home overrides the explicit argument, the geometry drifts + and the test is machine-dependent — which makes it useless as a gate, since it + would only ever fail on the machine nobody runs it on. Confirm by running the + faces suite with =TZ=America/Chicago= and again with =TZ=America/Denver=. + +If item 2 confirms, the design question is whether machine-resolved home belongs +in the pure geometry layer at all, or whether the host should resolve it and pass +it in — which is what the test already assumes. + +Not blocking the Lua port: =make test-faces= is disjoint from the hypr config and +the three suites that work touches. + +** TODO [#B] Qt apps render oversized on velox :bug:velox:solo: +:PROPERTIES: +:CREATED: [2026-08-19 Wed] +:LAST_REVIEWED: 2026-08-19 +:END: + +From the roam inbox, Craig's words: "qt apps look huge on velox. how do we make +it look better on this particular machine, and not change ratio. it seems they +should have different QT configs." + +The shape is per-machine Qt scaling. velox is a high-DPI Framework panel and +ratio drives ordinary-DPI monitors, so one global Qt scale factor cannot suit +both. The fix has to be host-scoped rather than a value written into the shared +config, which is the same tier split the dotfiles already use. + +Grading: Minor severity (apps work, they are just the wrong size) x every user +every time (every Qt app launch on velox) = P2 = [#B]. + +*** 2026-08-19 Wed @ 15:05:00 -0700 Root cause found and fixed; needs a logout to take effect +velox's =conf.d/local.conf= scaled the panel twice. The monitor line sets +=1.566667= and the same file exported =QT_SCALE_FACTOR,1.5= and =GDK_SCALE,1.5=, +and Qt 6 on Wayland already takes its scale from the compositor, so the two +multiplied. Measured rather than reasoned: with the override Qt reports a +960x640 logical screen, without it 1440x960, and 2256/1.566667 is exactly 1440. +That is 1.5x too large, which matches "huge" precisely. + +Those env lines were not careless. The comment above them explains they existed +to compensate for =xwayland:force_zero_scaling = true= in the shared +hyprland.conf, which makes XWayland clients render unscaled and tiny. The +approach was what failed: an env var reaches every app, so fixing XWayland broke +every native Wayland client. Removing the vars alone would have traded "Qt huge" +for "Zoom tiny", so velox now turns =force_zero_scaling= off for itself instead. +XWayland scales through the compositor there, coming out correctly sized and +slightly soft. ratio is untouched and needs nothing, its monitor being scale 1. + +=force_zero_scaling= took effect on =hyprctl reload=. The env removal will not: +Hyprland applies =env== lines with setenv at parse time and never unsets them, +so the running compositor still hands 1.5 to everything it spawns. Craig has to +log out and back in. + +*** VERIFY Is CALIBRE_OVERRIDE_DPI still needed after the scaling fix? +The same file pins =CALIBRE_OVERRIDE_DPI,96= with the comment "calibre renders +oversized at the 1.57 compositor scale". Calibre is a Qt app, so that was almost +certainly this same double-scaling seen through one application and worked +around per-app rather than at the root. With the multiplier gone, the pin is +probably redundant and may now render calibre too small. + +Left in place rather than removed on a guess, since it was validated at 96 on +2026-06-27 and calibre has its own DPI handling. Worth opening calibre after the +next login and deciding by eye. + +The cursor entry in the same file records this identical failure a third time: +"Pre-scaling it (the old 36 = 24 x 1.5) double-applied on top of the +compositor's scale." Three instances of one mistake in one file, two previously +fixed in isolation without anyone naming the pattern. + +** TODO [#C] Waybar panels launch expanded instead of collapsed :bug:dotfiles:waybar: +:PROPERTIES: +:CREATED: [2026-08-19 Wed] +:LAST_REVIEWED: 2026-08-19 +:END: + +From the roam inbox, Craig's words: "waybar panels should start up collapsed. +currently both the left and the right waybar panels launch expanded." + +Panel source is =~/.dotfiles=. Its heading in the roam inbox read "archsetup." +with a period rather than a colon, so the routing prefix did not match cleanly; +claimed on the plain reading of the text. + +Grading: Cosmetic severity (presentation only, nothing is lost) x every user +every time (every session start) = P3 = [#C]. + +** VERIFY [#C] The visible analog clock avoids being dragged :velox: +:PROPERTIES: +:CREATED: [2026-08-19 Wed] +:LAST_REVIEWED: 2026-08-19 +:END: + +From the roam inbox, captured verbatim: "the visible analog clock avoids being +dragged. ask me about this." + +Filed as a VERIFY because the capture asks for a conversation rather than +describing a defect. What is the clock avoiding being dragged by, and is the +avoidance the bug or the intended behaviour? + +** TODO [#C] A failed hostname lookup takes seven seconds :bug: :PROPERTIES: -:CREATED: [2026-08-13 Thu] -:LAST_REVIEWED: 2026-08-13 +:CREATED: [2026-08-19 Wed] +:LAST_REVIEWED: 2026-08-19 :END: -Mainboard swapped Intel→AMD (Ryzen AI 9 HX 370); new NVRAM has no boot entry. -Decision: full reinstall via archangel+archsetup, run deliberately as a -disaster-recovery drill before the Sunday flight. Runbook (live checklist): -[[file:working/velox-reinstall/velox-reinstall-runbook.org][working/velox-reinstall/velox-reinstall-runbook.org]] -Done 2026-08-13: ISO rebuilt (archangel-2026-08-13, archsetup baked with AMD -microcode detection, velox profiles at /root/, .ai/inbox excluded — build.sh -edits pending commit in archangel), contents verified, dotfiles swept clean of -Intel assumptions. -Finding folded in: velox's truenas backups silently stopped ~Jul 6 (newest is -DAILY.0 Jul 6; wolf.conf.gpg from Jul 29 is in NO backup). Salvage pass in the -runbook is therefore REQUIRED before partitioning, and the fresh install must -fix + verify the backup timer (runbook Phase 5). + +=getent hosts fake-vm= takes about 7.2 seconds to return not-found on velox. +Measured repeatedly with the cache flushed between runs. Anything that looks up +a name that does not exist pays it: an ssh typo, shell completion, a script +probing for a host. + +Not caused by the DNSSEC change. A/B measured today, cache flushed each time: +7691ms and 7232ms on =allow-downgrade= against 6804ms and 7482ms on =yes=, so +the setting makes no difference and this predates it. The likely shape is the +tailnet search domain (=search tailf3bb8c.ts.net=) being tried first, then the +two DoT upstreams, each with its own timeout, before NXDOMAIN comes back. + +Found because it blew a 20-second timeout in +=tests.net-scenarios.test_run_net_scenarios=, which shells out to ssh a +deliberately-bogus =root@fake-vm=. That suite passes on its own and the failure +did not recur, so the timeout needed this latency plus the DNS disruption from +the clock testing running alongside it. Worth knowing that the suite sits close +enough to the edge for a slow resolver to tip it. + +Grading: Minor severity (nothing behaves wrong, it just waits) x some users +sometimes (every failed lookup, which is occasional rather than constant) = P3 = +[#C]. + +** TODO [#B] Signal tray icon invisible under waybar (Electron 43 well-known-name SNI) :bug:waybar:velox: +:PROPERTIES: +:CREATED: [2026-08-25 Tue] +:LAST_REVIEWED: 2026-08-25 +:END: +Since signal-desktop 8.24.0 (Electron 43.4.0) Signal's tray item registers +under a well-known bus name (=org.freedesktop.StatusNotifierItem-<pid>-1=) +and answers Properties.Get/GetAll only when addressed by that name. waybar's +GDBus proxy addresses the owning unique name instead, reads back no Id or +Category, and logs "Invalid Status Notifier Item", so the icon never shows. +With =--start-in-tray= that leaves Signal running with no window and no icon; +launching it again from fuzzel raises the existing window. Slack (older +Electron, unique-name registration) is unaffected. Measured 2026-08-25 with a +bus monitor: the same connection returns the value for the well-known name +and "error occurred in Get" for its own unique name. + +Grading: Major severity (the app is unreachable from the desktop while +"running") × every user every time on velox = P1 by the matrix, held at +[#B] because the workaround (relaunch to raise the window) is cheap and the +fix is upstream. + +Upstream: [[https://github.com/Alexays/Waybar/issues/5240][Waybar #5240]] (open, proposes a raw-call fallback in item.cpp +proxyReady) and [[https://github.com/signalapp/Signal-Desktop/issues/7992][Signal-Desktop #7992]] (open, "Upstream Change Needed"). +Downgrading to 8.23.0 is closed off: 8.24.x migrated the SQLCipher schema to +1770 and 8.23.0 quits with DBVersionFromFutureError (tried and reverted +2026-08-25). + +Re-test after a waybar or signal-desktop upgrade, from the repo root: +#+begin_src sh :results output +grep -c 'Invalid Status Notifier Item' "$(\ls -t ~/.local/var/log/waybar-*.log | head -1)" +n=$(busctl --user list --no-legend | awk '$1 ~ /StatusNotifierItem-/ && $3=="signal-desktop"{print $1}') +busctl --user call "$n" /StatusNotifierItem org.freedesktop.DBus.Properties Get ss org.kde.StatusNotifierItem Id +busctl --user call "$(busctl --user call org.freedesktop.DBus /org/freedesktop/DBus org.freedesktop.DBus GetNameOwner s "$n" | cut -d'"' -f2)" /StatusNotifierItem org.freedesktop.DBus.Properties Get ss org.kde.StatusNotifierItem Id +#+end_src +Expected when fixed: 0 "Invalid" lines in a fresh waybar log, or both Get +calls returning the Id (either side fixing it clears the icon). + +Ratio is on 8.21.0 and unaffected until its next upgrade brings 8.24.x. +Alternatives if it drags on: change Signal's tray setting so it keeps a +window (=~/.config/Signal/ephemeral.json= =system-tray-setting=), or run a +waybar carrying the #5240 fallback. + ** TODO [#B] Truenas session-host VM for long-running agent sessions :feature:tooling: :PROPERTIES: :CREATED: [2026-08-13 Thu] @@ -120,6 +1007,160 @@ hand-copied key sprawl; a firm RAM carve-out so builds don't fight the ZFS ARC; headless only — desktop-coupled sessions stay on ratio/velox. Build deliberately AFTER the vacation, not before Sunday. Companion idea (cheaper, complementary): put ratio on the UPS. +** TODO [#B] post-rebuild-check: route every probe through one guarded helper :refactor:solo: +:PROPERTIES: +:CREATED: [2026-08-17 Mon] +:LAST_REVIEWED: 2026-08-17 +:END: +The script works and is well tested, but its shape keeps producing the same +bug. Across three review rounds the reviewer found FOUR separate instances of +"the probe failed and the check reported ok", each in a different place: +=systemctl= in check 1, the enablement read in check 2, =find= in check 3, and +=grep= in check 4. A fifth was latent in an unguarded staged write. Every one +was individually fixed, and I only stopped finding more because someone kept +looking. + +That is a design problem rather than four bugs. The script has five +hand-written probes, and each one has to remember to branch on its own exit +status. Nothing enforces it, nothing fails a review that forgets it, and the +failure is invisible because the wrong behaviour is a clean "ok". + +Shape: one helper every probe must go through, which cannot return a value +without an explicit success, so that "I could not read this" is +unrepresentable as "nothing to report". Roughly: + +: probe "<what>" <command...> # sets a value on success, records a finding otherwise + +Then each check consumes the helper's result rather than a raw command +substitution, and a new check written later inherits the discipline instead of +having to re-derive it. Worth pairing with a test that asserts no check can +report ok when its probe exits non-zero, generically, so the fifth instance is +caught by the suite rather than by a reviewer. + +Not urgent: the current version is correct as far as anyone has found, ships +with 58 tests, and proved itself on a genuinely wedged machine. This is +prevention. + +Grading: Minor severity (no known live defect, the risk is future) x +most-users-frequently (every future edit to this script) = P3 = [#C]... except +the failure mode is silent and the script's whole job is catching silent +failures, so a regression here is uniquely undetectable. P2 = [#B]. + +:solo: — the surface is one script and its suite, the refactor is +behaviour-preserving, and the existing 58 tests plus a mutation battery are +the objective check that it stayed so. +** TODO [#B] velox's systemd --user spins at 96% and cannot resolve unit files :bug:velox: +:PROPERTIES: +:CREATED: [2026-08-17 Mon] +:LAST_REVIEWED: 2026-08-17 +:END: +Live on velox 2026-08-17 from about 10:29. =systemd --user= (pid 2235) sits +in state R at 96% CPU, measured over a 3-second sample rather than taken from +the lifetime average. It stopped logging at 10:29, so its timers appear to +have stopped firing too. + +The split is the diagnostic: =systemctl --user list-units= still returns +instantly, while =is-enabled=, =cat=, =show=, and =list-unit-files= all hang +indefinitely. So the manager answers from its in-memory unit list and wedges +on anything that has to resolve unit files. It is spinning in userspace, not +blocked on I/O (=/proc/2235/wchan= is 0, no syscall pending). + +Remedies tried, neither worked: =systemctl --user daemon-reexec= hangs like +every other unit-file call, and the signal form (=kill -59=, SIGRTMIN+25) +was accepted but changed nothing. The next step is a logout/login or reboot, +which is Craig's call because it closes his running session. I deliberately +did not kill the manager: that would tear down the graphical session and +everything under it. + +Suspected cause is the powerprofilesctl crash loop filed above, whose +repeated activation attempts against a masked unit are the only new load on +this machine. I cannot prove it, and I have to name the other candidate +honestly: my own =post-rebuild-check= runs called =systemctl --user +is-enabled= roughly thirty times per run over several runs, and the wedge +appeared during that window. The crash loop predates those runs by an hour +and a half, which is why it is the leading suspect rather than the certain +one. + +What it costs: unit-file operations are unavailable, user timers appear +stopped, and a core is pinned on a laptop running on battery. + +Grading: Major severity (a pinned core and stopped user timers, invisible +unless you look) x rare edge case (one machine, specific conditions) = P2 = +[#B]... except that this is a live, ongoing drain on a travelling machine +rather than a latent defect, so it takes [#A] until the machine is back to +normal. Re-grade to [#B] once resolved and the question is only prevention. + +*** 2026-08-17 Mon @ 19:57:42 -0700 The reboot cleared it; re-graded [#A] to [#B] as the task instructed +velox rebooted at 16:04. The wedge is gone: =systemctl --user is-enabled +roam-sync.timer= now answers =enabled= in well under a second, where every +unit-file call hung indefinitely before, and =list-timers= shows +calendar-sync, roam-sync and agenda-render-cache all firing on schedule +again. So the remedy the task named — a logout or reboot — was taken and +worked. + +Nothing here was diagnosed further, which means the cause is still unproven +and both candidates in the body stand. What is left is prevention, and the +task's own grading says that is [#B]: the live-drain argument was the only +thing holding it at [#A], and the drain has stopped. Re-graded per that +instruction rather than by a fresh judgment. + +Reproducing it deliberately is the open question, and it is not obviously +worth doing — it costs a wedged session to learn something the crash-loop fix +may make moot. +** TODO [#B] post-rebuild-check needs a reference-host mode :feature:velox:solo: +:PROPERTIES: +:CREATED: [2026-08-17 Mon] +:LAST_REVIEWED: 2026-08-17 +:END: +=scripts/post-rebuild-check= ships and works, but its first live run on velox +2026-08-17 showed the output is mostly steady state rather than drift. Of the +8 findings that survived three rounds of false-positive removal, comparing +against ratio says exactly ONE is real: =obsbot-wb-guard= is enabled on ratio +and merely linked on velox, which is the deliberate deferral recorded +2026-08-16. The other three unit findings (=emacs=, =geoclue-agent=, +=obs-record-watchdog.timer=) are linked on ratio too, and the three =.claude= +absences are absent on ratio too. + +So the signal-to-noise is about 1:7, and the thing that separates them is a +comparison against the other daily driver — the same discipline that kept the +2026-08-16 session honest when check 4 read as nine projects missing +=CLAUDE.md= and ratio turned out to be missing the identical files. + +Shape: =--reference-host <host>= runs the same five checks on the far machine +over tailscale (ssh, read-only) and reports only the *differences*. Findings +present on both machines are steady state and get summarized as a count rather +than listed. Falls back to the current standalone behavior when the reference +host is unreachable, and says so. + +Grading: Minor severity (the tool works and its findings are accurate; they +are just buried) x every use = P3 = [#C]... except that a check nobody reads +is a check that isn't run, which is the failure mode the whole task existed to +close. Most-users-frequently x Major = P2 = [#B]. + +:solo: — the checks exist, the ssh path is proven (the 2026-08-17 session ran +exactly this comparison by hand), and correctness is verifiable locally by +diffing the two reports. +*** 2026-08-21 Fri @ 07:10:00 -0700 The premise moved: velox now reports 1 finding, not 8 +Re-scope before building. The 1:7 ratio this task argues from is gone, and two +of the three things it cites as noise are fixed at the source rather than +filtered. + +=87ff0b7= gave check 2 a machine-local expected-disabled list, so the four unit +findings are declared intent rather than noise, and an entry whose unit turns +out to be enabled is itself reported so the list cannot rot. =3fbf3e0= dropped +=.claude= from check 4's expected set, since the gitignore sweep writes that +line into every project whether or not one exists. velox went 8 findings to 1. + +So the open question is no longer "how do we cut the noise" but whether a live +reference-host diff still earns its place against a static declaration of +intent. They are different tools: the list is offline, explicit, and states +what a machine means; the diff is automatic and catches drift nobody declared. +The reference-host comparison is still what *found* all of this, twice, by +hand. That is an argument for it and not against. + +Worth knowing this task already contained the whole 8-to-1 analysis when it was +filed 2026-08-17, and a session on 2026-08-20 re-derived it from scratch without +reading it. Not :solo: any more — the design call above is Craig's. ** TODO [#C] screen-lock test suite red on ratio :bug:test:dotfiles: :PROPERTIES: :CREATED: [2026-08-13 Thu] @@ -147,17 +1188,174 @@ default, and a udev rule granting the video/input group write access so it works without sudo (a bare ssh session got EPERM). Check whether Fn+Space (EC-handled on Frameworks) already cycles it — if so the bind is a complement, not the only path. Ratio: n/a (desktop). +** TODO [#B] Post-rebuild verification pass :feature:velox:solo: +:PROPERTIES: +:CREATED: [2026-08-14 Fri] +:LAST_REVIEWED: 2026-09-17 +:END: +A rebuilt machine looks finished and isn't. Five gaps surfaced on velox +within two days of the 2026-08-13 reinstall, and three of them LOOKED +fine: a stowed unit file, an enabled timer, a present git clone. From the +.emacs.d handoffs 2026-08-14 (inbox, both PROCESSED) plus what this +session found independently. +The generalizable fix is one pass the installer runs at the end, or a +=post-rebuild-check= script the checklist points at. Each item is cheap +and turns a silent no-op into a visible line: +1. =systemctl --user list-units --state=failed= — calendar-sync had been + failing every 15 minutes for two days with nobody watching. +2. Every stowed/linked user unit that is NOT enabled. roam-sync and + signal-receive came back linked and inert; two others were never + linked at all. A unit file being present is not the same as running. +3. Every tracked =*.local.el.example= (or =*.local.*=) with no sibling + real file. Three exist in .emacs.d; all three were gone on velox. +4. Every gitignore-mode project missing its =.ai/=, =.claude/=, + =CLAUDE.md=, =todo.org=, =inbox/=. A reinstall drops the entire + working state of every such project — 374 files and 4.5 MB in + .emacs.d's case — and nothing carries it: not git, not stow, not the + bootstrap. +5. =signal-cli listAccounts= non-empty. velox lost its registration, and + because agent-text relays to a hardcoded velox, that breaks the phone + channel for the WHOLE FLEET, not just this machine. +6. =mbsync --list= parses. The Proton Bridge TLS cert + (=~/.config/protonbridge.pem=, referenced by =~/.mbsyncrc=) is generated + per *installation*, so it cannot be restored or copied between machines. + Its absence aborts the config parse, which kills *every* account — gmail + and dmail need no bridge and died anyway. The error names only the missing + pem, so "no mail at all" and "this one file is missing" look unrelated. + Re-derive it off the running bridge's own handshake, no GUI, no secrets: + =openssl s_client -connect 127.0.0.1:1143 -starttls imap -showcerts </dev/null | sed -n '/BEGIN CERTIFICATE/,/END CERTIFICATE/p' > ~/.config/protonbridge.pem= +7. The bridge password (=~/.config/.cmailpass=) is per-install too. It is a + real file rather than a stow symlink, so it survived the rebuild holding + the *previous* install's value — worse than absent, because it looks + right. Diagnostic trap: the bridge answers a wrong password with =no such + user=, which reads as "no account signed in" and sends you hunting a login + problem that doesn't exist. Never treat =no such user= as evidence about + account state. + +*The distinction that organizes all seven* (from the .emacs.d handoff +2026-08-14, inbox): every artifact that broke was generated on the machine by +an application rather than carried by git, stow, or dotfiles. But they split +two ways, and conflating them is what produces a file that exists, looks +right, and authenticates against nothing: +- *Restore* — the old value is still correct: gitignored tooling (1), + roam clone state (3), =*.local.el= configs (5). +- *Re-derive* — the old value is worthless because the application minted a + new one: signal-cli registration (4), bridge cert (6), bridge password (7). +So the checklist wants two columns, not one. + +Originally graded [#A] because item 5 was live and silently disabling paging, with +a flight that Sunday forcing the date. Both inputs have expired: the flight was +2026-08-17 and item 5 no longer gates anything time-boxed, so the 2026-09-17 +review dropped this to [#B] and unscheduled it. Kept here because the reason the +grade was ever [#A] is worth knowing; it is not the current read. + +*** 2026-09-13 Sun @ 07:14:54 -0500 Moved the three gap reports into the reinstall working dir +The 2026-08-14 reports that define the five gaps moved out of inbox/ into +[[file:docs/design/2026-08-14-velox-reinstall-gaps-1.org][gaps 1-4]], +[[file:docs/design/2026-08-14-velox-reinstall-gaps-2.org][gap 5]] and +[[file:docs/design/2026-08-14-velox-reinstall-gaps-3.org][the two email-side gaps]] (the Bridge cert and the Bridge password). +They file with the rest of the reinstall artifacts when that task closes. +*** 2026-09-17 Thu @ 08:59:59 -0400 Re-graded [#A] → [#B]; gaps 6 and 7 are the remainder +=scripts/post-rebuild-check= ships and covers gaps 1-5, with two [#B] +follow-ups filed (the guarded probe helper and the reference-host mode). Both +reasons for [#A] are gone: gap 5 is a check now, and the flight has passed. +What the script still lacks is the two email-side gaps: =mbsync --list= +parsing (the per-install Bridge cert) and the Bridge password (a real IMAP +login against =127.0.0.1:1143=, never reading =no such user= as account +state). Add those two, then close. + +** TODO [#B] Restoring a git repo from backup can resurrect a dangerous diff :bug: +:PROPERTIES: +:CREATED: [2026-08-14 Fri] +:LAST_REVIEWED: 2026-08-14 +:END: +My 2026-08-14 restore of =~/org= from the salvage brought back roam's +=.git= deliberately ("simpler, preserves everything exactly"). It also +brought back a clone ten commits stale AND an uncommitted =inbox.org= +emptied to zero bytes. roam-sync is the repo's only committer and commits +whatever it finds, so enabling that timer would have committed the +emptying and pushed it — deleting the live inbox items ON RATIO. A +.emacs.d session caught it, verified ratio's copy was a strict superset, +discarded the local diff, fast-forwarded, and only then enabled the timer. +Lesson to encode somewhere durable: restoring a git repo from a backup is +not the safe option it looks like. For any repo with a live remote, +re-clone and carry only proven-needed work; where a backup copy is +restored anyway, reconcile it against the remote BEFORE any +auto-committing timer is enabled. The failure here would have been silent +and landed on a different machine. +** TODO [#B] Nothing installs the .emacs.d systemd user units :bug:velox: +:PROPERTIES: +:CREATED: [2026-08-14 Fri] +:LAST_REVIEWED: 2026-08-14 +:END: +=~/.emacs.d/systemd/= ships four user units (agenda-render-cache +service+timer, calendar-sync service+timer). On ratio they are symlinked +into =~/.config/systemd/user/= by hand. Nothing does that on a fresh +machine: they are not stowed (they live in .emacs.d, not dotfiles) and +archsetup does not link them. +Consequence found on velox 2026-08-14: the world wallpaper face drew +nothing, because it reads =~/.cache/settings/agenda.json= and the timer +that exports it was never installed. Calendar sync was silently dead for +the same reason — which is the second time that particular timer has gone +missing (see the 2026-08-01 session, where its auto-start was the bug). +Linked and enabled by hand on velox; export verified (316 bytes, 1 event). +Fix belongs in whichever owns the seam: either .emacs.d gains an install +step for its own units, or archsetup links them alongside the dotfiles +stow. Prefer the former — the repo that ships a unit should install it. +Grading: Major severity (two background services silently absent, and the +failure looks like a data problem rather than a missing timer) x every +fresh install = P2 = [#B]. +** TODO [#C] Panel can leave a channel selected with nothing to show :bug:dotfiles: +:PROPERTIES: +:CREATED: [2026-08-14 Fri] +:LAST_REVIEWED: 2026-08-14 +:END: +velox's store carries channel "pair" with pair_sel unset, so +channels.selected_pair() returns None and wallpaper.apply() fails every +time. Found 2026-08-14 when the new session-start restore reported +"unavailable" and fell through to the waypaper fallback — the machine +still showed dark-lion, which looks exactly like the bug that was just +fixed. ratio is fine (channel world, pair_sel 0). +Two candidate fixes, needs a call: either the panel refuses to switch to a +channel whose selection is empty, or apply() falls back to the first +minted pair/set when the index is unset. The second is friendlier and +matches "the store is the source of truth" — a channel with exactly one +plausible reading should not be a dead end. +Grading: Minor severity (one fallback still puts a wallpaper up) x some +users sometimes = P3 = [#C]. +** TODO [#B] archsetup doesn't clone rulesets :bug:velox: +:PROPERTIES: +:CREATED: [2026-08-14 Fri] +:LAST_REVIEWED: 2026-09-17 +:END: +A fresh install has claude but no =ai=, no skills, no rules, no hooks, +because =~/code/rulesets= is never cloned. Found on velox 2026-08-14 when +=ai= wasn't on PATH. archsetup clones dotemacs, dotfiles, the suckless +tools and itself, so rulesets is the one workstation repo it misses, and +without it the whole agent tooling layer is absent on a rebuilt machine. +Fix: clone it alongside the others (=RULESETS_REPO=, defaulting to +git@cjennings.net:rulesets.git) and run =make install= afterwards, which +is what links the 54 symlinks into ~/.claude and ~/.local/bin. Graded +Major severity (a rebuilt machine silently loses every agent workflow) +x most-users-frequently = P2 = [#B]. Worked around by hand on velox +already; this is the durable half. ** TODO [#B] Hibernate in the settings dial power actions :feature:dotfiles: :PROPERTIES: :CREATED: [2026-08-13 Thu] :LAST_REVIEWED: 2026-08-13 :END: Add hibernate alongside suspend/lock in the settings module's dial power -actions. Sequencing (Craig confirmed the dial placement 2026-08-13): -1. Prove hibernate on velox first — systemctl hibernate through a real - resume; the chain (LUKS swap p3, keyfile-in-initramfs, encrypt+resume - hooks, resume= on the ZBM cmdline) went live with the 2026-08-13 - reinstall but is untested on this AMD board. +actions. The wlogout exit menu already carries it (keybind h) and needs no +work; the dial is the remaining surface. Sequencing (Craig confirmed the +dial placement 2026-08-13): +1. DONE 2026-08-14 00:14 — hibernate proven end to end on velox, driven + from the exit menu so the wiring was exercised too. Evidence: the boot + id was unchanged across the cycle (f14152f9…) and uptime kept counting + 3h15m → 3h18m, so it genuinely resumed rather than rebooting; the + journal carries "PM: hibernation: hibernation exit" and the + HibernateLocation EFI variable being cleared. Took 9.3s wall. The whole + chain works: suspend-to-disk into the LUKS-encrypted swap, resume via + the keyfile embedded in the initramfs, one passphrase at ZBM. 2. Then consider suspend-then-hibernate as the default lid behavior (systemd sleep.conf HibernateDelaySec) — hibernate's savings with no button at all; possibly a "deep sleep" toggle in the module. @@ -165,6 +1363,7 @@ actions. Sequencing (Craig confirmed the dial placement 2026-08-13): Note: ratio has no swap partition, so hibernate stays velox-only until ratio gets one; the dial entry should degrade gracefully where there's no resume target. +** TODO [#A] Move secrets out of public dotfiles → private repo + combined personal ISO :feature:security:dotfiles: :PROPERTIES: :CREATED: [2026-08-11 Tue] :LAST_REVIEWED: 2026-08-11 @@ -184,86 +1383,30 @@ archangel+archsetup ISO that's already ~80% built. Two ISO modes: generic don't start the migration until the credentials are rotated. Not started. Not :solo: — repo standup and history rewrite are Craig's calls; promote to a real spec (spec-create) when work resumes. -** VERIFY [#A] Pre-vacation fix list — morning review -SCHEDULED: <2026-08-08 Sat> -:PROPERTIES: -:LAST_REVIEWED: 2026-08-08 -:END: -The full todo.org sweep you asked for before sleeping, ranked by what I'd fix -before departure (~2026-08-15, velox travels). Approve, reorder, or strike; -items needing your call say so. - -1. Velox reliability (the anchor — [#A] sleep/suspend, rescheduled Wed - 2026-08-12). Velox is out for repair/upgrade until Tuesday or Wednesday - (Craig, 2026-08-08), so every velox item waits for its return — a tight - but workable window before the ~08-15 departure. Riders already folded - in: the tlp.d radio-enable line, a dotfiles pull, the touchpad-detection - spot-check. -2. Velox machine health for travel (NEW — filed nowhere else): resolve the - ~/code/auto-dim-other-buffers.el merge conflict (literal conflict markers - in a loaded .el; its emacs suite has been red since 2026-08-01), clear the - stale password prompt sitting on its screen since 2026-07-31, and run a - maint doctor pass. -3. Remote access verified from OUTSIDE the LAN while you're still home: - tailscale to ratio, truenas, and truenas-kvm from a phone hotspot. - DECIDED (Craig, 2026-08-08): the wolf WireGuard profile gets set up on - velox when it returns Tue/Wed — added to the velox-return riders. Cheap - at home, expensive to debug from a hotel. -4. The cgit secrets/privacy audit ([#B] below): a world-readable secret - standing while you're away is the worst timing. The repo-by-repo scan is - mine to run; the public-vs-private call per repo is yours. The archsetup - cgit move can wait unless the audit finds something. -5. Already scheduled today: osbot camera (needs the camera plugged in). - Buildable any time: the podman socket + camera udev task (:solo:). -6. Optional travel niceties blocked on upfront-answerable design calls in - their bodies (two for hotspot/metered WiFi in amber, one for network-panel - ordering by availability). Answer the calls and I can build both. -7. Deliberately left off: offline LLM (you declined the vacation track), - night-watch/lock-watchdog (ratio stays home with no user to relock; say so - if you disagree). Found tonight, low priority: the orchestrator sequence - pin can't see an added-but-unstubbed call (it caught drops only) — worth a - harness hardening pass someday. -** DONE [#B] Podman API socket and camera-passthrough udev rule :feature:solo: -CLOSED: [2026-08-09 Sun] -:PROPERTIES: -:CREATED: [2026-08-07 Fri] -:LAST_REVIEWED: 2026-08-07 -:END: -Shipped 2026-08-09: the installer enables the rootless podman socket at -install time (enable_user_service grew a wants-target arg so socket units -land in sockets.target.wants) and ships -=72-usb-passthrough-cameras.rules= — numbered below 73 per the winvm -rule-ordering correction, GROUP/MODE as the verified grant, uaccess tag kept. -Applied live on ratio (socket enabled+active, 99- file retired, udev -reloaded); velox apply rides the velox-return riders on the sleep/suspend -task. The uaccess-alone hypothesis stays untested until a camera is attached. -From winvm 2026-08-07 (ratio). Two one-time machine-level setups, both live on -ratio and absent on velox; full evidence and rationale in -[[file:docs/design/2026-08-07-podman-socket-and-camera-udev.md]]. - -- Enable the rootless podman socket at install time - (=systemctl --user enable --now podman.socket=). Socket-activated, zero idle - cost; every podman GUI/API client needs it, and its absence fails silently - (Pods opens to an empty window). The installer already carries the - "=systemctl --user enable= fails during install" workaround pattern - (=archsetup:1270=, =:2722=) — use it. -- Ship a udev rule granting GROUP="video", MODE="0660" on the OBSBOT - (3564:ff02) and BRIO (046d:085e) USB nodes so =usbredirect= can claim them - for VM passthrough. CORRECTED (winvm, 2026-08-08): the original "uaccess - can't ACL raw USB nodes" claim was wrong — the mechanism is rule ordering. - The ACL is applied by =73-seat-late.rules=, so a =99-= rule adds the tag - after that already ran; distro rules that add the tag all sort at or below - 70. So number our file below 73 (e.g. =72-usb-passthrough-cameras.rules=), - keep the verified GROUP/MODE grant, and keep the tag — correctly ordered it - may make uaccess work on its own (untested hypothesis; a tighter grant if - it holds, needs the camera plugged in to verify). Reconcile ratio's - existing =99-= file (winvm installed it) when the installer version lands. - -Scope: installer step + rule file + tests per existing shapes, and apply both -live to velox over tailscale (daily-driver sync — neither exists there today). + +*Also bake the push-capable repo URLs into the personal ISO* (decided +2026-08-19). =archsetup:240= and =:245= default =archsetup_repo= and +=dotfiles_repo= to =https://git.cjennings.net/...=, the anonymous read-only +endpoint. That default is right for a stranger installing archsetup — no key on +the server — and wrong for my machines, which have to push: velox came back +from its rebuild unable to push either repo, and I only found out at a 403 four +days later. I decided against detecting an ssh key in the installer, because +archsetup never restores =~/.ssh= (I do that by hand), so key-presence at clone +time depends on ordering the installer doesn't control, and a naive "any key +means ssh" would break a stranger who happens to have one. The override already +exists and is documented — =ARCHSETUP_REPO= / =DOTFILES_REPO= in +=archsetup.conf.example= — so the personal ISO just needs to carry the ssh +form of both, alongside the secrets bundle. The generic ISO keeps the https +default untouched. + +The gap that leaves is a curl|bash or stock-ISO install, which takes the https +default straight back. =post-rebuild-check= check 8 covers that path — it flags +a working repo whose origin is the read-only endpoint — so the ISO value is the +fix and the check is the net under it. +** TODO [#B] Settings toggles reset silently at session start :bug:dotfiles: :PROPERTIES: :CREATED: [2026-07-28 Tue] -:LAST_REVIEWED: 2026-07-28 +:LAST_REVIEWED: 2026-09-17 :END: Craig, from the roam inbox 2026-07-28: "launching into wayland doesn't honor previous caffeine settings ...or I expect any other settings in the desktop settings module." Captured right after the 08:59 reboot. @@ -294,7 +1437,6 @@ Related: =[#B] Caffeine state is unreadable on both surfaces= covers display acc Launched by hand afterward it runs fine and survives, so gammastep is not broken — it loses a race against compositor readiness at session start. Nothing relaunches it, so night light is simply off for the whole session, silently. (An earlier read of this said night light "has likely never worked from the config". That was wrong: the failure is a startup race, not a permanent break.) Worth its own task — the fix is a readiness wait or a retry around that exec-once, not a persistence change. Filed here for now because it surfaced during this investigation. - ** TODO [#C] Re-apply the active program at session start :refactor:dotfiles:hyprland: :PROPERTIES: :CREATED: [2026-07-30 Thu] @@ -333,11 +1475,10 @@ An earlier draft graded this Minor, arrived at [#C], and then wrote [#B] beside Not :solo: — naming is Craig's taste call, and the rail's vocabulary should be decided as a set. -** TODO [#A] Night watch and the lock watchdog fight each other :bug:hyprland:dotfiles: -DEADLINE: <2026-07-31 Fri> +** TODO [#B] Night watch and the lock watchdog fight each other :bug:hyprland:dotfiles: :PROPERTIES: :CREATED: [2026-07-29 Wed] -:LAST_REVIEWED: 2026-07-29 +:LAST_REVIEWED: 2026-09-17 :END: ROOT CAUSE of the lockdead screens, found 2026-07-29 00:50 within minutes of the relaunch logging going live. hyprlock is not crashing. It is being killed on purpose, by us. @@ -379,11 +1520,15 @@ Not :solo: — the fix is a design decision between two subsystems, both of whic Option 2 is the one I would argue for, but it is Craig's call. Previous title and framing of this task, kept for the record: "hyprlock still exits mid-lock; the watchdog relaunch is silent". The instrumentation that closed that gap is dotfiles =5bbe2c3=, and it paid for itself in about six hours. +*** 2026-09-17 Thu @ 08:59:59 -0400 Re-graded [#A] → [#B] and dropped the 07-31 deadline +The grading paragraph above already concluded [#B]; the heading never followed +it. The watch stage was parked in =db5ac60= on 2026-07-29, so the collision is +mitigated and the choice between the three fixes has no date pressure. ** TODO [#B] The wireguard gpg convention is inert; the installer can't read it :bug:security:network: :PROPERTIES: :CREATED: [2026-07-28 Tue] -:LAST_REVIEWED: 2026-07-28 +:LAST_REVIEWED: 2026-09-17 :END: Found 2026-07-28 by an independent review, while deciding whether to commit a newly-encrypted =wolf.conf.gpg=. @@ -404,10 +1549,10 @@ Until it is resolved, do not commit any =*.conf.gpg=. The encrypted =wolf.conf.g Related: =[#B] Move archsetup off cgit= and =[#B] Audit cgit-published repos for secrets and privacy=. Both are still open, and both argue for keeping new secrets out of this repo until they land. -** TODO [#B] Timer presets should start in one click :feature:dotfiles:timer: +** TODO [#C] Timer presets should start in one click :feature:dotfiles:timer: :PROPERTIES: :CREATED: [2026-07-28 Tue] -:LAST_REVIEWED: 2026-07-28 +:LAST_REVIEWED: 2026-09-17 :END: Craig, captured 2026-07-28, routed here by home's inbox-zero pass from the shared roam inbox. @@ -420,38 +1565,14 @@ Panel source is =~/.dotfiles/timer/src/timer/= (=gui.py= for the view, =panel.py Not :solo: — the rearrangement is described but not settled, and the result is a visual judgment Craig has to see. The one-click behaviour is buildable on its own; the layout wants a pass in front of him. Per the UI-prototyping rule, sketch the arrangement before touching production code. Related: =[#C] Add a time selector to the timer panel= covers a duration picker for the same input area. Design them together when either is picked up. - -** DONE [#A] Comet KVM setup for truenas :feature:infra:truenas: -CLOSED: [2026-08-08 Sat] -:PROPERTIES: -:CREATED: [2026-07-27 Mon] -:LAST_REVIEWED: 2026-07-27 -:END: -Resolved: Craig wired up and configured the Comet himself, confirmed working -2026-08-08. The ATX power-board follow-up (hard power-cycle for a truly wedged -box) remains unfiled — raise it if the next outage shows the KVM alone isn't -enough. -Wire up the GL.iNet Comet (GL-RM1) IP KVM against truenas. It was bought 2026-01-14 for exactly this job and its KB node still reads "Arrived, not yet set up." - -Why now: truenas went dark 2026-07-24 and stayed unreachable. Diagnosis from ratio on 2026-07-27 — no tailnet contact for 3 days, 100% packet loss on 192.168.86.5, ARP entry FAILED (nothing answers ARP for the address, so the NIC is down at layer 2), every service port closed, while the gateway and a dozen other LAN hosts stayed reachable. Wake-on-LAN to 70:85:c2:db:9d:94 drew no response. With no console and no out-of-band power control there was no remote remedy at all, so recovery needed hands on the box. The Comet closes exactly that gap: BIOS/UEFI console, Wake-on-LAN, and browser access over its native Tailscale integration. - -Not :solo: — the physical cabling is Craig's, and the Tailscale enrollment needs his account. - -Steps, from the KB node ([[id:67bc5994-a763-48e2-926f-4ac0d1bad3db][GL.iNet Comet (GL-RM1) - KVM]]): -1. HDMI from truenas video out to the Comet's HD IN. -2. USB-A-to-USB-C from the Comet to a truenas USB port (keyboard/mouse emulation). -3. Ethernet to the network. -4. Power via USB-C (5V/2A). -5. Reach the web interface and enroll it in Tailscale, so it's usable when the LAN side of truenas is the thing that's broken. - -Then verify while truenas is healthy, rather than discovering the gaps during the next outage: confirm the console shows POST and the BIOS, that keyboard input reaches the box, and that Wake-on-LAN from the Comet actually powers it on. Enable WOL in the truenas BIOS if that last check fails — this outage never established whether it was on. - -Worth considering as a follow-up: the ATX power board accessory gives hard power-cycle control for a truly wedged box, which the KVM alone can't do. +*** 2026-09-17 Thu @ 08:59:59 -0400 Re-graded [#B] → [#C] +Untouched for seven weeks, and the layout wants a prototype pass in front of me +before any production code. ** TODO [#C] Post-upgrade hooks: compositor restart reminder + font cache rebuild :feature:infra:ratio:solo: :PROPERTIES: :CREATED: [2026-07-25 Sat] -:LAST_REVIEWED: 2026-07-25 +:LAST_REVIEWED: 2026-09-17 :END: Handoff from home (2026-07-25), originally combining the 2026-06-07 stale-compositor incident and 2026-06-08 fontconfig crash diagnosis. Add two reproducible pacman =PostTransaction= hooks through archsetup; do not make one-off =/etc= edits: @@ -459,23 +1580,10 @@ Handoff from home (2026-07-25), originally combining the 2026-06-07 stale-compos 2. On =Upgrade= of =fontconfig=, =freetype2=, or =harfbuzz=, run =/usr/bin/fc-cache -f= after the transaction. The fontconfig 2.17→2.18 cache-format change left stale cache-9 files that crashed Qt6 apps in =FcCharSetHasChar= until the system font cache was rebuilt. Acceptance: hook files are source-controlled and installed by archsetup; package/operation/action fields are asserted from the generated hook text; the reminder is print-only and exits successfully; the font hook runs only after successful matching upgrades and invokes the absolute =fc-cache= path. Validate with the fast installer tests plus a disposable pacman-hook parser/install check when practical. -** TODO [#D] net-scenarios harness times out under back-to-back suite runs :test:tooling: -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -=tests/net-scenarios/test_run_net_scenarios.py= errored on all 5 tests twice during round 12, each time =subprocess.TimeoutExpired= after its 20s budget on =scripts/testing/run-net-scenarios.sh --target root@fake-vm=. Both occurrences were in =make test-unit= runs launched immediately after a previous full run. It then passed 6 runs in a row (3 on a pristine tree, 3 with the round-12 change), and standalone it finishes in 0.09s, so this is not a regression from any code change. - -The harness stubs =ssh=, =rsync= and =jq= onto =PATH=, so nothing should touch the network at all — which is what makes a 20s timeout suspicious rather than merely slow. Worth reproducing under load before deciding whether the fix is a larger timeout or a real hang in the script. Evidence logs from the round: =/tmp/tu.log= and =/tmp/tu2.log= (tmpfs, gone after reboot). - -Not graded on the bug matrix: it is test infrastructure, not the shipped codebase. - -*** 2026-08-08 Sat @ 05:05:00 -0500 Recurred under concurrent load, same signature -All 5 tests hit the 20s TimeoutExpired again during a =make test-unit= run -that overlapped two review subagents running their own suites on the box. -Standalone immediately after: 0.095s, all pass; the following quiet-machine -full run was clean. Confirms the load-sensitivity read — reproduce under -deliberate load before choosing between a bigger budget and a real hang. - +*** 2026-09-17 Thu @ 08:59:59 -0400 Lead, not diagnosed: ratio's Hyprland aborted on 08-28 +ratio logged a Hyprland SIGABRT coredump at 2026-08-28 15:44, three days after +the 714-package upgrade of 08-25. Check whether it's the stale-compositor shape +from 06-04 before citing it as a second occurrence. ** VERIFY Should coredump entries group as one journal-digest row per binary? :maint: :PROPERTIES: :LAST_REVIEWED: 2026-07-24 @@ -507,22 +1615,14 @@ Not =:quick:= despite being small: four pieces with tests is a sitting rather th From the roam inbox (Craig, claimed 2026-07-23): the wallpaper channel switches on sunrise/sunset today (the sun-pair mode, =settings/src/settings/wallpaper.py=, location read live via whereami with a state.json cache). Add a timed-schedule mode as an alternative: fixed clock times drive the transitions rather than the solar calc. Not :solo: — the capture itself flags the missing inputs ("we'll need to know the transition times, and how many of them there are"). The count and the times are a design decision Craig owes: is it a two-image day/night flip at fixed hours, an N-way ring across the day, per-image dwell vs shared interval? The =set= channel already does fixed-interval cycling through a set, so the new part is specifically clock-anchored transition points, not just "a timer". Ask for the schedule shape at pickup, then build against the existing wallpaper.apply presenter vocabulary. -** TODO [#D] Worldclock tooltip blanks on one bad timezone row :bug:dotfiles:waybar:quick:solo: -:PROPERTIES: -:LAST_REVIEWED: 2026-07-25 -:END: -Found by sentry (2026-07-25), verified by exercising. =hyprland/.local/bin/waybar-worldclock= builds each zone with =ZoneInfo(tz)= inside the loop (line ~99) with no guard, so a single malformed timezone row in =worldclock.conf= raises =ZoneInfoNotFoundError= and crashes the whole python pass. The tooltip then renders empty and *every* zone is lost, not just the bad row; the traceback only reaches stderr, where waybar never surfaces it. -Repro: a conf with =America/Chicago|Home=, =Not/AZone|Bad=, =Europe/London|London= renders =tooltip: ""= (Home and London gone too). -Grade: minor severity (one module's tooltip blanks, no data loss) x rare edge case (a malformed conf row) = P4 = [#D]. -Fix: wrap the per-row =ZoneInfo=/=datetime= in a try/except and =continue=, so a typo drops only that row and the valid zones still render. Solo + quick: the script already has an env-override test harness (=WAYBAR_TIME_EPOCH=, =WAYBAR_WORLDCLOCK_CONF=), so a red-first test is cheap. -** TODO [#C] Auto-dim status forgotten on layout change :bug:dotfiles: +** TODO [#C] Auto-dim status forgotten on layout change :bug:dotfiles:solo: :PROPERTIES: -:LAST_REVIEWED: 2026-07-25 +:LAST_REVIEWED: 2026-09-17 :END: From the roam inbox (Craig, 2026-07-25). If auto-dim is toggled off and the layout then changes, auto-dim silently comes back on. A layout switch should not touch the auto-dim state. Likely related to the 2026-07-25 =layout-cycle= rebuild (floating ring) or a hook it fires — check whether the layout-change path resets the dim toggle, and where auto-dim state lives. Grade: minor severity (dim re-enables unexpectedly, no data loss) x every layout change made while dim is off = P3 = [#C]. -** TODO [#C] Maint queue button status wall needs a copy button :feature:dotfiles:maint: +** TODO [#C] Maint queue button status wall needs a copy button :feature:dotfiles:maint:quick:solo: :PROPERTIES: -:LAST_REVIEWED: 2026-07-25 +:LAST_REVIEWED: 2026-09-17 :END: From the roam inbox (Craig, 2026-07-25). The maintenance queue button shows the status wall but has no copy button. Add one, following the global COPY key pattern already on the maint doctor wall (dotfiles =8bc79ba=). Grade: cosmetic/feature = [#C]. ** TODO [#B] Panel family: unify the look across net/bt/maint/audio and desktop-settings :feature:design:dotfiles: @@ -535,20 +1635,6 @@ From the roam inbox (Craig, claimed 2026-07-24): the network, bt, maint, and aud :LAST_REVIEWED: 2026-08-02 :END: From the roam inbox (Craig, claimed 2026-07-24): panel labels look cut off; a few more pixels of space fixes it. He named the audio and bt panels, but his "before" capture is the networking panel (=~/pictures/screenshots/2026-07-23_202419.png=; "after" resizing =~/pictures/screenshots/2026-07-23_202458.png=), so the whole panel family likely shares the tight spacing. Confirm which panels clip at pickup, then add the padding/width. Grade: cosmetic × every glance at the affected panels = P3 = [#C]. Solo — buildable (CSS/size tweak) and screenshot-verifiable, no design call once the clipping panels are identified. -** TODO [#C] Spine face tests decay against the wall clock :bug:test:dotfiles:solo: -:PROPERTIES: -:LAST_REVIEWED: 2026-08-02 -:END: -=settings/faces/timeline-face-spine.test.mjs= has thirteen =SP.spineRows(g, h)= calls that omit the third argument, so =ref= falls back to its =new Date()= default while the file's events fixture is pinned to =JUL= (2026-07-31 18:30 UTC). Any assertion that depends on how much room the day needs is then measured against today's clock, and rots as the fixture recedes. - -One of them, "spacing is uniform everywhere except the gap home opens", had already rotted: green on 07-31 because that was the fixture's own date, red by 08-02. Fixed in place on 2026-08-02 by pinning =JUL=; the remaining thirteen pass today by luck. The measurement, for whoever picks this up — with =ref=now= the even step is 85.21 and home's gaps are 129.10 / 65.40 (the lower one collapses below a plain gap); with =ref=JUL= the step is 78.54 and the gaps are 129.10 / 145.46. Only the lower gap moves, because =up= does not depend on events and =down= does. - -Six other calls in the same file already pass =JUL= explicitly, so the convention exists and this is a miss, not a gap in the design. Fix: pass =JUL= at every call whose assertion reads geometry. Leave the call around line 747 alone — it sweeps =new Date(t0)= deliberately. - -Grade: minor severity (dev-facing only; no product behavior is wrong, the face itself is fine) x some users, sometimes (each call rots independently, whenever the fixture drifts far enough) = P3 = [#C]. Not merely cosmetic though: a suite that goes red for no real reason is how a genuine regression gets waved through. - -Solo — mechanical, an existing convention to copy, and verifiable by running the suite plus re-running it under a faked clock to prove the determinism actually holds. - ** TODO [#C] Night-watch live telemetry :feature:maint: :PROPERTIES: :LAST_REVIEWED: 2026-08-02 @@ -560,13 +1646,14 @@ Craig, 2026-07-21 ("mind. blown."): drive the Dupre Night Watch screensaver (doc :END: Craig, 2026-07-21: the wlogout window (Super+Shift+Q — lock/reboot/shutdown/logout/suspend/hibernate) "isn't great and has bugs." Review it end to end: catalogue the specific bugs, then assess the design against the Dupre instrument-console family (it predates the panel aesthetic). Config lives in dotfiles; the bind is hyprland.conf:428 (=pgrep -x wlogout || wlogout-menu=). Context: the desktop-settings panel spec withdrew lock/suspend in favor of this screen (2026-07-21 amendment), so it's now the sole owner of session-exit actions — worth being good. Grade each bug found via the severity×frequency matrix; this parent stays a [#C] review until specifics emerge. ** TODO [#A] Audit cgit-published repos for secrets and privacy :bug:security: +SCHEDULED: <2026-09-24 Thu> :PROPERTIES: -:LAST_REVIEWED: 2026-08-09 +:LAST_REVIEWED: 2026-09-17 :END: Grading: security carve-out — cgit at git.cjennings.net serves every repo under scan-path=/var/git over unauthenticated https (any repo is anonymously cloneable). Raised [#B] → [#A] on 2026-08-09: the scan found a real live-credential leak (below), so this is now a confirmed exposure with an open rotation blocking, not a hypothetical. Drops back to [#B] once rotation is done and the visibility rulings are made. Not :solo: — needs Craig's decisions and the credential rotation. Steps: list repos under /var/git; for each, decide intended public vs private; scan each for secrets (done, below); for any meant-to-be-private repo, actually restrict access (cgit repo.hide only hides from the index — a known repo name is still cloneable; use http auth or move it off the public scan-path); for public repos, confirm no secrets and add a pre-receive/CI secret scan. archsetup's own move is decided and tracked separately below. *** VERIFY [#A] Rotate the credentials exposed by the 2026-08-09 dotfiles leak -SCHEDULED: <2026-08-10 Mon> +SCHEDULED: <2026-09-24 Thu> A plaintext credential file was briefly public in the dotfiles repo and was confirmed pulled by an external crawler before the purge, so every credential in it must be rotated. Full list, forensic detail, and the remediation record @@ -585,31 +1672,87 @@ doc above (not published, since they map the setup). Follow-ons: the rotation VERIFY above, velox reconcile on return, the secrets-repo split (top of Open Work), the wireguard =.gitignore= bug (line ~191), the cgit move (below), and a pre-receive secret-scan hook so this can't recur. -*** TODO [#B] velox: reconcile its clones after the history rewrite -velox was offline for repair during the 2026-08-09 purge, so its clones still -hold the pre-rewrite history and are diverged from the rewritten remotes. On -its return: force-fetch + rebase local work onto the rewritten main in both -repos (or re-clone), force-update the local tag, local-gc, before its next -push. Also on the velox riders on the sleep/suspend task. +*** 2026-08-17 Mon @ 19:57:42 -0700 Moot — the 08-13 wipe re-cloned velox from the rewritten remotes +This asked velox to reconcile clones that no longer exist. The machine was +wiped and reinstalled on 2026-08-13, so every repo on it was cloned fresh +*after* the purge and never held the pre-rewrite history at all. The runbook +anticipated this ("fresh clones automatically carry the post-purge rewritten +git history"); nobody closed the task once the reinstall took that route. + +Verified rather than assumed: both repos are level with =origin/main= today — +archsetup at =6faa31c=, dotfiles at =65940f2=, both trees clean. + +One thing the reinstall did leave, and it is filed separately: the installer +cloned both repos =--depth 1=, so the history was present-but-truncated until +today's =git fetch --unshallow= (see the shallow-clone =[#A]=). A reconcile +against the rewritten remote was still unnecessary — a shallow clone of the +right history is not a diverged clone of the wrong one. ** TODO [#B] Move archsetup off cgit to cjennings@cjennings.net :chore:security: :PROPERTIES: -:LAST_REVIEWED: 2026-07-21 +:LAST_REVIEWED: 2026-08-17 :END: Decided (Craig, 2026-07-20): move the archsetup repo off the public cgit host (git@cjennings.net, scan-path /var/git) to Craig's private account remote cjennings@cjennings.net, so it is no longer world-cloneable. This is the archsetup-specific fix for the cgit-exposure finding above. Plan: create a bare repo under cjennings's control off the cgit scan-path (e.g. =~cjennings/git/archsetup.git=); push current main + tags there; migrate the post-receive hook that publishes the installer to =/var/www/cjennings/archsetup= so curl-install keeps working (the single published file stays public by design; only the repo goes private); update the origin remote on ratio and velox to =cjennings@cjennings.net:git/archsetup.git=; remove =/var/git/archsetup.git= so cgit no longer serves it. Verify: anonymous =git clone https://git.cjennings.net/archsetup.git= fails, the new private clone works from both machines, and the curl-install URL still returns the installer. Keep the two daily drivers' remotes in sync (daily-drivers rule). + +*** 2026-08-17 Mon @ 19:57:42 -0700 Re-checked: unstarted, and the exposure is confirmed live +Ran the task's own verification step as it stands today, which is the honest +way to check an unstarted task rather than reading its body back. Anonymous +=git ls-remote https://git.cjennings.net/archsetup.git= succeeded with no +credentials and returned =6faa31c= — this afternoon's HEAD. So the repo is +still world-cloneable and current to the commit, not a stale published +snapshot. + +=origin= on this machine is still =git@cjennings.net:archsetup.git=, the cgit +account, so nothing has moved. Everything in the plan stands unchanged. + +*** 2026-08-21 Fri @ 14:12:46 -0700 Recorded the publication mechanism: placement is the only control +The work project verified its own repo reads "not served" against a control +repo that reads PUBLIC, and reported the mechanism back: the host publishes via +=GIT_HTTP_EXPORT_ALL= over =GIT_PROJECT_ROOT=/var/git=, so *publication is +directory placement and nothing else* — there is no per-repo marker, no +=git-daemon-export-ok= file, no opt-in flag to check. A repo is public because +of where it sits. + +That is the durable hazard for this task's plan, and it cuts both ways. It +confirms the approach — a bare repo created outside the scan-path is private by +construction, which is exactly what the plan already specifies. It also means +nothing in a repo itself records whether it is exposed, so any future move +*into* =/var/git= publishes silently, with no local artifact to notice. Their +own repo is private for this reason alone: it lives under =/var/cjennings/git/=, +outside the served root. + +Caveat they raised and I agree with: any enumeration of the served set is a +snapshot, not a standing fact. The set moved twice while three projects were +measuring it. Verify placement at the time of the move rather than trusting a +recorded list. ** TODO [#B] Velox boot-failure retrospective — upgrade guard gaps :bug:zfs:maint: :PROPERTIES: -:LAST_REVIEWED: 2026-07-21 +:LAST_REVIEWED: 2026-08-26 :END: Post-mortem for the 2026-07-15 velox no-kernel boot failure, from the archsetup/maint code review: - maint's UPDATE remedy runs a plain =yay -Syu --noconfirm= (remedies.py:297). The live-update guard (guard.py) only matches mesa/hyprland (the 2026-06-07 live-swap class) — it never checks /boot, kernel, initramfs, or mkinitcpio exit. No post-upgrade /boot assertion exists. An interrupted kernel transaction slips straight through. - Add a post-upgrade /boot assertion: after a transaction touching linux/linux-*, confirm vmlinuz-* + initramfs-*.img present and mkinitcpio exit 0; refuse to end the run (or page Craig) otherwise. Would have caught this. - Sanoid-vs-actual dataset drift: configure_zfs_snapshots configures zroot/var/log + zroot/var/lib/pacman as separate datasets; velox's actual layout has neither separate (/var/log sits inside zroot/var). Reconcile. - CONFIRMED (2026-07-21): the pre-pacman snapshot hook fired on velox — the 2026-07-15 no-kernel boot was recovered via the pre-pacman ZFS snapshot rollback, and velox is back on the tailnet running linux-lts 6.18.38 with initramfs present (2026-07-19 session). Root-cause hook-ordering fix shipped separately. Still open: the post-upgrade /boot assertion in guard.py and the sanoid-vs-actual dataset drift reconcile (the two bullets above). - -** TODO [#B] Assess a Hyprland left-drag window gesture :feature:hyprland: -:PROPERTIES: -:LAST_REVIEWED: 2026-07-21 +*** 2026-08-26 Wed @ 22:35:01 -0600 The /boot assertion now lives in the topgrade spec; the dataset drift is what remains here +The post-upgrade /boot assertion is covered by the kernel-modules-check gate in +[[file:docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org][the topgrade guarded-upgrade spec]] +(dkms built for the new kernel, initramfs newer than vmlinuz, pre-pacman +snapshot on a ZFS root), which ships with that spec's Phase 1 rather than here. +What this task still owns is the sanoid-vs-actual dataset drift: whether to +split zroot/var/log and zroot/var/lib/pacman out as configure_zfs_snapshots +assumes, or change the config to match the layout velox actually has. That is +a call I have not made, so the task stays [#B] and not solo. + +*** 2026-09-13 Sun @ 07:14:54 -0500 Filed the original boot-failure handoff under docs/design +The 2026-07-15 diagnosis and recovery plan that opened this task sat in +inbox/ as a processed file; it now lives at +[[file:docs/design/2026-07-15-velox-boot-failure-handoff.org][docs/design/2026-07-15-velox-boot-failure-handoff.org]] +so the timeline survives the inbox sweep. + +** TODO [#C] Assess a Hyprland left-drag window gesture :feature:hyprland: +:PROPERTIES: +:LAST_REVIEWED: 2026-08-26 :END: Evaluate whether a global left-click drag can move ordinary windows without breaking application selection, text interaction, or Wayland security @@ -618,35 +1761,94 @@ any binding. ** TODO [#B] Reconcile panel keybindings around Super+N :feature:hyprland: :PROPERTIES: -:LAST_REVIEWED: 2026-07-21 +:LAST_REVIEWED: 2026-08-21 :END: -Swap the notification and networking bindings so primary panels are one -Super-plus-letter chord away, audit the other exceptions, and bring the -proposed family to Craig for a final mapping decision. +Put every panel on one consistent chord family — net, bluetooth, audio, timer, +and the maintenance console — as a shared modifier set plus a mnemonic letter +per panel (N/B/A/T/M). Today they open by waybar click only, so a uniform +family is what makes them keyboard-reachable and predictable. The immediate +move is swapping the notification and networking bindings so the primary panels +sit one Super-plus-letter chord away. + +Maintenance (M) is the chord I want first — it is the panel I keep reaching for +without one. + +Constraints: +- Super+Shift+A is already the PTT toggle, and the hold-to-talk grave bind is + load-bearing. Audit every current hyprland bind for conflicts before + proposing a family, and treat these two as fixed. +- Both machines have to work the same way. Velox can't QMK-remap, so the chords + have to be typable on a plain laptop keyboard. + +Steps: settle the modifier family, audit the existing binds for collisions, +wire it through the dotfiles hyprland config, and document it in the keybind +reference. + +The family itself is the one call I haven't made — the audit and the wiring +follow from it, so that decision comes first rather than last. + +*** 2026-08-21 Fri @ 14:15:22 -0700 Merged the duplicate keybinding-family task into this one +Two tasks were carrying one job: this one and =[#B] Consistent keybinding family +for the panel console=, filed separately and both stalled. This one had the +tighter framing and the more recent review; that one had the better body — the +specific collisions, the velox plain-keyboard constraint, and maintenance-M as +the priority chord. Folded its detail in here and cancelled it, since two +half-specified tasks for one decision is plausibly why neither moved. + +Not =:solo:= and not =:quick:=: the modifier family is a preference call I have +to make, judging what's load-bearing among the existing binds needs me too, and +the audit plus wiring plus docs runs past thirty minutes on its own. ** TODO [#B] Add storage-capacity signals to the maintenance module :feature:maint: :PROPERTIES: -:LAST_REVIEWED: 2026-07-21 +:LAST_REVIEWED: 2026-08-26 :END: Investigate capacity and growth diagnostics for full disks, identify the appropriate remedies, and incorporate a clear storage signal into the maintenance console. -** TODO [#B] Add per-channel controls to the audio panel :feature:audio: +** TODO [#C] Add per-channel controls to the audio panel :feature:audio: :PROPERTIES: -:LAST_REVIEWED: 2026-07-21 +:LAST_REVIEWED: 2026-08-26 :END: Expose channel-level input and output volume controls without losing the existing device-level workflow. ** DOING [#B] Widget gallery upgrades :feature:design: :PROPERTIES: -:LAST_REVIEWED: 2026-07-13 +:LAST_REVIEWED: 2026-08-23 :END: -Usability + documentation pass over the [[file:docs/prototypes/panel-widget-gallery.html][panel widget gallery]], orthogonal to the component-generation spec work, so it runs on the =gallery-upgrades= branch (squash merge to main after Craig's UI confirmation + tweaks). Items 1-4 run as a no-approvals speedrun (Craig authorized 2026-07-12); item 5 is a joint brainstorm after the merge. +Usability + documentation pass over the [[file:docs/prototypes/panel-widget-gallery.html][panel widget gallery]], orthogonal to the component-generation spec work. Items 1-4 run as a no-approvals speedrun (Craig authorized 2026-07-12); item 5 is a joint brainstorm. + +The =gallery-upgrades= branch this originally described is gone — no local or remote ref, and every gallery commit since has landed straight on main. Whether it was squash-merged or abandoned, the branch stopped describing how this work runs, so the line came out at the 2026-08-23 review rather than being left to mislead. Work on main. *** TODO Extraction-readiness bar for every gallery component :refactor:design: Craig's standing directive (2026-07-18, set while finishing the split-flap): every =DUPRE.*= builder should meet the bar the split-flap now sets, since these become regular components. The bar: a contract comment documenting every opt and the full handle surface; no page globals touched (page owns cadence via handles/callbacks, e.g. =onSettle=); all component CSS in one named =DUPRE_CSS= block; refactored until no opportunity worth doing remains (small named helpers, no duplication); construction axes declared via =STYLES= where the component has them. Sweep the existing builders against that list, fix the gaps, and make the bar a stated convention in the widgets.js header or README so new builders inherit it. Overlaps the component-generation spec's extraction phase — reconcile there rather than doing the work twice. +*Audited 2026-08-23 — the sweep covered the bulk and stopped short.* Ten commits +on 2026-07-18 (=43725ff= … =1dd929d=) carried ~104 of the 112 builders over. Two +criteria are fully met: 105 builders carry a real contract comment naming opts +and the handle surface, and component CSS is wholly consolidated — the gallery's +own =<style>= block holds nothing but page chrome (masthead, grid, toc, card +frames, validation lamps, palette). Three gaps remain, and they are the whole of +what's left: + +1. *Seven builders were never swept*: =telegraphIndicator=, =radarSweep=, + =dotMatrix=, =flipDisc=, =dekatron=, =gearIndicator=, =blinkenlights=. Each + carries a one-line comment describing what the widget does, with no opts, no + handle surface, no CSS statement. They sit at the tail of =widgets.js= after + =responseGraph=, and the last batch was "well-through-response" — the sweep + stopped one builder short of the end and never came back. Not a clean cut: + =dayDateCal= is in that tail and does have a contract. +2. *The bar was never written down.* The README documents the API shape (Builder + contract) and =DUPRE_CSS= (Styling) — two of the five criteria. The checklist + itself appears nowhere, so a new builder inherits nothing and the sweep has to + be re-derived from this task every time. +3. *One page-global reach survived*: =patchBay= does + =window.addEventListener('resize', draw)= and never removes it. Fails the + no-page-globals criterion and leaks besides — a torn-down instance keeps + redrawing on every resize. The other two =document= reaches are benign + (=indexPlate= guards a shared SVG def; the other is the CSS injector). + *** 2026-07-18 Sat @ 04:32:20 -0500 Made the N20 split-flap an honest Solari mechanism =GW.splitFlap= rebuilt from the drop-fade fake: charset-as-drum stepping (=opts.chars= is the flap order, one flip at a time through intermediates, staggered arrival), re-aim-not-queue retargeting, and the real two-half-panel fold (WAAPI, backfaces hidden), with =animate:false= collapsing to instant jump for reduced motion. Handle grew =setText=/=chars=/=reading()=; default width 3 → 4 cells. Nine probe checks written red-first (arrival, intermediates, one-flap stepping, re-aim discriminator, instant path); technique studied from HotFX and re-derived — no license on their repo, nothing copied (reference filed in =working/retro-stereo-widgets/references/=). Review: sound; its two test-strength notes addressed in the same change. @@ -803,11 +2005,25 @@ The cross's cell assignments were read off card names and spec sheets, not audit ** DOING [#B] Retro widget catalogue :feature:design: :PROPERTIES: :SPEC_ID: 3ac0d42c-db1a-4d21-bce4-e63785fef0ba -:LAST_REVIEWED: 2026-07-13 +:LAST_REVIEWED: 2026-08-23 :END: The panel widget gallery ([[file:docs/prototypes/panel-widget-gallery.html][docs/prototypes/panel-widget-gallery.html]]) grows into a retro-instrument component catalogue: reference photos of period hardware → gallery cards (the visual + behavioral spec) → reusable components for three targets (emacs svg.el, web/React, waybar). Tokens single-sourced in [[file:docs/prototypes/tokens.json][tokens.json]] (gen_tokens.py emits web/waybar/elisp); svg.el proof widget shipped (gallery-widget.el, needle gauge). Reference photos live in [[file:working/retro-stereo-widgets/][working/retro-stereo-widgets/]]. Collection converged at R56, then reopened at R57 as the taxonomy found empty cells (110 cards, all behaviorally verified; probes in [[file:tests/gallery-probes/][tests/gallery-probes/]]). Build runs per the [[file:docs/specs/2026-07-12-component-generation-spec.org][component-generation spec]] (DOING; reviewed + decomposed 2026-07-12): web extraction first (ungated, lossless), then demand-gated Emacs/waybar ports. Banked variant/composition ledger lives in the 2026-07-11/12 session archive. + +*The eight open subtasks are really one decision plus three builds* (noted at the +2026-08-23 review, because "8 open" reads as more contested than it is). Phase 1 +shipped. Phases 3, 4 and 5 and the spec flip are each gated, directly or +transitively, on *Phase 2 — the demand inventory*, which is my matrix to write +and nobody else's. That single artifact has been the whole chain's blocker since +2026-07-12. The three genuinely independent items are the magic-eye rebuild, the +wind-direction rose, and weather kit integration. + +Weather kit integration may already be unblocked: its note says live panel +verification "awaits a stowed desktop with a private weather location +configured", and both daily drivers are stowed now. Check whether +=$WEATHER_LAT=/=$WEATHER_LON= or =~/.config/weather/config.json= is set before +treating it as still waiting. *** TODO [#B] Rebuild the magic-eye tube component :feature:design: Reinstate the magic-eye tuning/level indicator (EM34/EM84/6E5 family), but replace the earlier weak UI rather than reviving it unchanged. The component @@ -869,10 +2085,18 @@ Restyle the audio panel's GTK CSS onto =tokens-waybar.css= + the banked composit After ~5 hand ports, weigh widget-level codegen with evidence (mechanical duplication vs judgment per port). Recorded as a dated decision in the spec; go spawns its own spec. *** TODO Flip the spec to IMPLEMENTED When the phases above close: status heading keyword → =IMPLEMENTED=, dated history line with the reason, Metadata =Status= mirror. Three lines, one file. -** TODO [#B] Net doctor expansion v1 — VM live verification :feature:dotfiles:network: + +*** 2026-09-13 Sun @ 07:21:18 -0500 Filed the two Maeda applets into the clock display references +The Line (C2, 1997) and Cosmos (C1, 1995) standalone applets sent from the +website project on 2026-07-30 sat in inbox/ as processed files; they and +their notes now live beside the other references as +=2026-07-30-maeda-{line,cosmos}-standalone.html= and =-notes.org=. +Reference only, regenerate rather than edit. + +** TODO [#B] Net doctor expansion v1 — VM live verification :feature:dotfiles:network:solo: :PROPERTIES: :SPEC_ID: ce29b103-ed9d-4f56-bf8c-9ed8fe680ff3 -:LAST_REVIEWED: 2026-07-13 +:LAST_REVIEWED: 2026-08-25 :END: Build the [[file:docs/specs/2026-07-11-net-doctor-expansion-spec.org][net doctor expansion]] (IMPLEMENTED). Adds the control-plane cluster (rival-manager / nm-masked / keyfile-perms) and a sharper auth verdict to the shipped net doctor (=~/.dotfiles/net/=). Archsetup owns the dotfiles work end to end — edit, test, commit, and push in =~/.dotfiles=, then drop an inbox note. All build phases shipped and fake-verified; the one open piece is the VM live verification below. *** 2026-07-11 Sat @ 02:47:47 -0500 Built the read-only control-plane probe @@ -892,23 +2116,15 @@ On dotfiles main (=12e3e76=, pushed). =gather_context= derives an auth cause on *** 2026-07-12 Sun @ 09:14:00 -0500 Flipped the net spec to IMPLEMENTED and logged the vNext items Spec status heading now IMPLEMENTED (dated history line + Status mirror); all four phase headings DONE. vNext items (flaky/drops cluster, DoT/DNSSEC verdict, profile hygiene) logged as the "Net doctor vNext" task. The privileged-fix live halves remain with the VM live-verification sub-task and the manual-testing checklist — findings there come back as bugs. -** TODO [#B] Consistent keybinding family for the panel console :feature:hyprland: -:PROPERTIES: -:LAST_REVIEWED: 2026-07-09 -:END: -Consider putting every panel (net, bluetooth, audio, timer, and the coming maintenance console) on one consistent chord family — a shared modifier set (Super+Shift, Control+Alt, or similar) plus a mnemonic letter per panel (N/B/A/T/M). Today the panels open via waybar clicks only; a uniform chord family makes them keyboard-reachable and predictable. Watch for collisions with existing binds: Super+Shift+A is already PTT toggle, and the hold-to-talk grave bind is load-bearing. Decide the family, audit current hyprland binds for conflicts, wire via the dotfiles hyprland config, and document in the keybind reference. Both machines (velox can't QMK-remap, so chords must work on a plain laptop keyboard). -*** 2026-07-14 Tue @ 00:31:36 -0500 Folded Craig's ask for a maintenance-panel keybinding; bumped [#C] → [#B] -Craig asked (in session, 2026-07-14) for a maintenance keybinding specifically — the panel he's reaching for without one. Maintenance (M) is the priority chord when this task gets worked. The capture graduated the task from parking lot to active backlog. - ** DOING [#B] Run-time privilege model, standard across every panel doctor :feature:dotfiles: :PROPERTIES: -:LAST_REVIEWED: 2026-07-13 +:LAST_REVIEWED: 2026-08-25 :END: The audio input/output doctor is gaining a run-time privilege model (see [[file:docs/specs/2026-07-10-audio-doctor-input-side-spec.org][docs/specs/2026-07-10-audio-doctor-input-side-spec.org]], decision "The doctor may use sudo, resolved by context at run time"). Craig's call, 2026-07-10: make it a standard, "revise the other panels to be consistent with these changes." The model: a doctor resolves its privilege at run time from three signals — passwordless sudo available (=sudo -n true=, which never hangs), a tty to prompt at, and whether it is the GUI panel. Four remedy classes: Auto (user-scope, reversible), Privileged (needs sudo — runs where passwordless, prompts on a CLI tty, degrades to Guide in a GUI with neither), Reboot-tail (run the applicable part, then instruct the reboot), and Guide (physical/BIOS/wait-for-upstream, nothing to run). Safety floor: every Privileged and Reboot-tail remedy defaults to Confirm or Arm tier, never silent Auto, because passwordless sudo is not consequence-free. -The shared helper is built (see the dated entry below). What remains is per-panel adoption: wire each doctor's remedies through =panelkit.privmodel.resolve()= and audit them against the Confirm/Arm floor, and reconcile maint's =priv.py= build/fire table with the model rather than leaving its implicit always-passwordless assumption. That wiring lives in the per-panel fix phases (net Phase 1, bt Phase 2, audio input/output), each needing a real privileged host to verify =--fix= end to end, so none is agent-solo. +The shared helper is built (see the dated entries below), maint is reconciled onto it, and net is wired: =classify.py= carries the =remedy_class= on its privileged verdicts and =doctor.py= resolves each through =panelkit.privmodel.resolve()= (net Phase 1, shipped 07-11). Adoption is the gate only — every panel's repair actions already exist; what adoption changes is whether and how an existing privileged action is allowed to run (RUN where passwordless, PROMPT on a CLI tty, GUIDE in a GUI), under the Confirm/Arm floor. What remains, checked against the tree 2026-08-25: bluetooth is part-wired (=bt/doctor.py= makes one =resolve(PRIVILEGED, ...)= call, no per-remedy classes yet — audit its individual fixes against the floor), and audio has nothing on the doctor side (pending the input-side spec). Each needs a real privileged host to verify =--fix= end to end, so not agent-solo. Craig's decision, 2026-07-12: maint's harmless-reclaim privileged remedies (the silent CLEAN UP set — paccache keep3, journal vacuum, coredump clean) STAY silent-auto. The reconciliation gives that class a sanctioned, documented exception to the confirm floor rather than forcing Confirm/Arm; the value of the floor holds for everything else. Where sudo is not passwordless, maint should degrade per the model (prompt on a tty, guide in a GUI) instead of hard-failing. @@ -933,7 +2149,7 @@ The support machinery was deliberately kept for this task: =layout-navigate= and ** TODO [#B] Audit dotfiles/common directory :chore:dotfiles: :PROPERTIES: -:LAST_REVIEWED: 2026-07-14 +:LAST_REVIEWED: 2026-08-25 :END: Refiled from the archsetup task audit (2026-06-28), landed via ~/.dotfiles/inbox; the dotfiles content split into its own repo 2026-06-16 but the task tracking stays here per Craig (2026-07-02). Three parts: - Review all 50+ scripts in =~/.local/bin= and remove unused ones. @@ -945,121 +2161,32 @@ ACTION before the kill pass: redo the reference scan to grep all invocation sour *** 2026-07-14 Tue @ 01:40:48 -0500 Built the audit evidence report Shipped as =docs/2026-07-14-bin-audit-evidence.org= in the dotfiles repo (260752a). 150 scripts bucketed: 69 keep (referenced or cron-driven), 74 kill candidates (zero references in the tree), 7 flagged (all dwm-tier, expected on a hyprland host). Config sweep: audacious and wofi configs are orphan candidates, ranger needs an install-vs-delete call (declared in archsetup but not installed on ratio). Shell history was too shallow (~700 lines) to prove by-hand disuse either way — the kill pass stays Craig's call in the parent task. -** TODO [#B] Waybar network module — custom/net :feature:waybar:network: -:PROPERTIES: -:LAST_REVIEWED: 2026-07-09 -:END: -Unifies the old wifi-no-internet indicator (was =[#C]=) and the network-manager -dropdown (was =[#B]=) into one =custom/net= module: a tested Python =net= engine -(nmcli + diagnostics), a thin bar indicator, and a GTK4 layer-shell panel. Code -lives in the dotfiles repo (hyprland tier + a =net/= package like pocketbook); -archsetup only installs deps. Secrets stay in NetworkManager's own store (no -separate credential store). The =captive= script becomes the diagnostics engine. -Full design, acceptance criteria, and the failure-mode coverage table: -[[file:docs/design/2026-06-29-waybar-network-module-spec.org][2026-06-29-waybar-network-module-spec.org]]. - -Phases below, dependency order. Engine/unit work is agent-verifiable (=unittest= -+ fakes on PATH, coverage via venv); the live-network and visual states need real -conditions, filed under "Manual testing and validation". - -*** 2026-06-29 Mon @ 20:19:11 -0400 Phase 1 shipped — indicator + console recovery -Shipped to the dotfiles repo (10 commits, =5254bd8=..=c095a22=, pushed to main). -The =net= engine is a src-layout Python package in-tree, imported by a bin shim -that resolves the stow symlink back to the repo — so it runs from a bare TTY with -no install, which the recovery path depends on. - -Landed: =net status= (fast path, one nmcli call + sysfs, degraded fallback in -budget) + =net probe= (native captive probe, single-flight flock, atomic cache, -fresh/stale/expired/unknown classes, iface/SSID/UUID invalidation); =waybar-net= -replacing =custom/netspeed=, throughput → tooltip, CSS states in both themes + -live; =net diagnose= (read-only steps) + =net repair= (rfkill/reset/bounce/ -dns-test, cleanup-verified) + =net doctor [--fix]= with the four terminal -classifications; =net portal= + the =captive --probe-json= refactor; redacted -JSONL event log; Makefile recovery targets (=make online= etc.); =~/.config/net/ -config=. Verified live: =make net-status= reads the real wlp170s0 / @Hyatt_WiFi. - -Airplane (Craig's call, option 1): =custom/net= absorbs only the *display* — net -reads the airplane-mode state file and shows an airplane state/glyph. The -airplane-mode toggle stays (it's a low-power mode — radios + CPU + brightness + -services — not a radio switch), now on =custom/net='s right-click + signal 15. -Deleted: =waybar-airplane=, =waybar-netspeed=, =custom/airplane=, their tests + -css. =airplane-mode= kept. - -Tests: 160 in =tests/net/= (fake nmcli/curl/rfkill/resolvectl/ping/getent/ -systemctl on a temp PATH; doctor-classification fixtures; degraded-under-slow- -nmcli benchmark) + the =captive= probe-mode tests; full dotfiles suite green (32 -suites). Coverage-gap pass via throwaway venv: pure modules ≥90% branch -(classify 100%), IO-error branches excused in the test docstring. -Deferred to Phase 2/3: archsetup deps (gtk4-layer-shell/python-gobject Phase 2, -speedtest-go-bin Phase 3 — not added before the code that needs them). -Verify (manual, live): see Manual testing and validation. - -*** 2026-06-29 Mon @ 22:19:25 -0400 Phase 2 shipped — panel shell + connection management -Shipped to dotfiles (commits =4e7740f=..=24bcac5=, pushed). Engine: =net list= (saved -MRU + in-range wifi scan, infrastructure types filtered), =net up/down= (UUID-keyed, -mutation safety — keep prior link until target activates, classify wrong-password vs -generic, report auto-reactivation), =net add/edit/remove/rescan= (open + WPA-PSK; -enterprise activate-only; secret to NM's store, never our JSON/log — tested). - -Panel: a GTK-free PanelModel (selection, four state machines, the UX-flow enable -rules, terminal states) + a GTK4 gtk4-layer-shell window (=net panel=) anchored -top-right under the bar — Connections section with MRU list, active marked, signal -glyph, row-click select, Connect/Add/Forget/Rescan, confirm-on-forget, worker-thread -engine calls via GLib.idle_add. GTK imported lazily so the CLI/tests stay GTK-free. - -Bar interactions (settled with Craig over live iteration): left = =net-panel= toggle, -middle = =net portal=, right = =net-fix= (notify the doctor result when one-way; open -a terminal only when the outcome is fixable — the sudo/interactive case). Airplane on -Super+Shift+A. archsetup adds =gtk4-layer-shell= + =python-gobject= (this commit); -already on velox. - -Tests: 204 in tests/net (merge ordering/dedup, up/down mutation safety, no-secret-leak -on add/edit, panel model + state machines, gui row-format helpers). Full dotfiles suite -green (32 suites). Live-verified on velox: panel opens/toggles, list shows real 24 -profiles, right-click notification delivers (Craig confirmed). Phase 3 (diagnose/repair/ -speedtest IN the panel) is next; the engine for it already exists from Phase 1. - -*** 2026-06-29 Mon @ 22:43:40 -0400 Phase 3 shipped — diagnostics + speed test in the panel -Shipped to dotfiles (=91277cf=..=691abcb=) + archsetup (=48052d6=, speedtest-go-bin), -pushed. Engine: =net speedtest= (parses speedtest-go --json → ping from latency ns, -down/up from per-server byte rates; missing-backend / offline / malformed → error -envelope per the failure table). Panel grew a section switcher with four pages: -- Connections (Phase 2). -- Diagnose: =net diagnose= on a worker thread, each step a row (✓/✗/… glyph + title + - redacted evidence), read-only; Open-portal button when captive. -- Repair: "Get me online" (=net doctor --fix=) + tiers (rfkill/reset/bounce/dns-test) - + force portal. Confirmations in-panel with the spec's exact wording; the privileged - tiers run via =net-popup= terminal (where the sudo prompt + step output, incl. - cleanup-verified, show) — a panel has no tty, and pkexec would mean a prompt per op. -- Speed test: in-process =net speedtest= (no privilege → inline result: ↓/↑ Mbps + ping - + server), Run/Cancel (Cancel pkills the child), error envelope shown. - -213 net tests; pure helpers (step_indicator, format_speedtest) unit-tested. Full -dotfiles suite green (32 suites). One unverified assumption: speedtest-go's dl/ul unit -(taken as bytes/s; =BYTES_PER_SEC= flips it) — needs one real run vs a reference. The -in-panel repair streaming (vs terminal) is a named future polish once the GUI-privilege -story settles. - -The waybar network module ([#B] parent) is now COMPLETE through Phase 3. Phase 4 -(in-app help + user guide) and Phase 5 (VPN/WireGuard) remain as future work; the core -feature (indicator + recovery + panel + diagnostics + speed test) is done. -Verify (manual, live): see Manual testing and validation. - -*** 2026-07-09 Thu @ 16:32:54 -0500 Audit reconcile: Phase 4 is filed on the dotfiles side, waiting on them -The dotfiles project accepted the Phase 4 handoff and filed it as a =[#C]= task in their own =todo.org= (their note, 2026-07-08 16:56): the help-text audit + panel help affordance, the user-guide/README, and the ratio rollout doc. Not started there. They ping when it lands, and this task's Phase 4 child closes then. Nothing to do here meanwhile. - -*** TODO Phase 4 — docs + rollout :network:blocked: -Deliverable: in-app help (=net --help= + per-command, panel help affordance); -README/user-guide (commands, indicator states, panel, config keys, make targets, -troubleshooting from the failure table, rollback); archsetup Hyprland dep install -(=gtk4-layer-shell=, =python-gobject=, =speedtest-go-bin=); ratio manual dep + -stow step. -Verify: =net --help= and each subcommand complete; user-guide covers every command -+ the recovery targets. -Build handed off to the dotfiles project 2026-07-04 (=~/.dotfiles/inbox/2026-07-04-1305-from-archsetup-phase4-handoff.md=): archsetup deps confirmed installed, the remaining help/user-guide/rollout-doc work is in the net package. dotfiles pings back when it lands. - -*** TODO Phase 5 — VPN / WireGuard CLI fold (vNext) :network: -Rescoped 2026-07-04 (audit): the tunnels track already shipped most of the original Phase 5. Panel tunnel bring-up/down and detection landed (dotfiles 2d9d060 probes tailscale/NM-wireguard/Proton; 21db05a brings overlays up/down from the panel's Tunnels sub-view; 31ba056 diagnose/doctor understand tunnel routes; archsetup 2e40781 wireguard config import; the net-panel-other-interfaces spec is IMPLEMENTED). What remains for Phase 5 is only the =net vpn ...= CLI subcommand — cli.py still has no vpn/tunnel parser. Fold the panel's existing tunnel operations into a CLI surface; spec separately when picked up. +** TODO [#C] net vpn CLI subcommand :feature:network:dotfiles:solo: +:PROPERTIES: +:LAST_REVIEWED: 2026-08-21 +:END: +=cli.py= in the dotfiles =net/= package has no vpn/tunnel parser, so everything +the panel can already do with tunnels has no command-line equivalent. Fold the +panel's existing tunnel operations into a =net vpn ...= surface mirroring what +the Tunnels sub-view does — bring an overlay up, take it down, report status. + +The operations themselves already exist and are tested: dotfiles =2d9d060= +probes tailscale / NM-wireguard / Proton, =21db05a= brings overlays up and down +from the panel, =31ba056= taught diagnose and doctor to understand tunnel +routes, and archsetup =2e40781= imports wireguard configs. This is a CLI surface +over shipped behavior, not new capability. + +Graded [#C] rather than [#B]: the panel already does the job, so this is +convenience rather than a gap. It earns a bump if I find myself wanting tunnel +control from a bare TTY — which is the same recovery-path argument that made the +rest of =net= worth having as a CLI. + +=:solo:= — the subcommand shape is obvious (it mirrors the panel), the =net= +package's fake-based harness covers the build and verify path, and archsetup owns +the dotfiles work end to end. Not =:quick:=: every prior net phase landed with +twenty-odd new tests and a review pass, so this runs past a spare moment. + +Carved out of the =custom/net= umbrella when that closed on 2026-08-21. ** TODO [#B] Local offline LLM runtime + per-host model cache :tooling:llm: :PROPERTIES: @@ -1099,82 +2226,9 @@ Boot the configured endpoint and send a short prompt; surface success/failure + Acceptance: fresh VM install of the ratio profile reaches an endpoint on =:8081= that answers a smoke prompt; velox profile gets Q4_K_M + 8B and answers a prompt within reasonable laptop latency; network-down install completes successfully with the pending-models warning surfaced. -** DONE [#A] Review post-archsetup laptop setup steps (velox 2026-04-10) -CLOSED: [2026-08-08 Sat] -:PROPERTIES: -:LAST_REVIEWED: 2026-08-08 -:END: -Closed at the 2026-08-08 session: every open item got its automate-vs-document -call and the work landed the same night (tests green, committed). Residual: -velox itself still needs the new tlp.d radio line and a dotfiles pull — folded -into the [#A] sleep/suspend task, which works the same files on velox anyway. -Items discovered during velox setup that needed manual intervention after archsetup. -Decide which should be automated in archsetup vs documented as post-install steps. - -*** 2026-08-08 Sat @ 04:43:42 -0500 Automated radio enable via TLP (rfkill boot soft-block) -Root cause sharpened during triage: archsetup masks systemd-rfkill on laptops -(it fights TLP), so nothing restored radio state at boot — the "unblock once -should stick" premise was wrong under the mask. Fix in the TLP custom conf: -=DEVICES_TO_ENABLE_ON_STARTUP="bluetooth wifi"=, the TLP-native mechanism. -configure_tlp_power parametrized for tests; covered by -tests/installer-steps/test_configure_tlp_power.py. - -*** 2026-07-04 Sat @ 11:48:24 -0500 Automated /efi restrictive mount permissions in fstab generation -archsetup:2827-2836 now rewrites the /efi fstab line to =fmask=0177,dmask=0077= (idempotent), so fresh installs no longer land the world-accessible =fmask=0022,dmask=0022= default. Confirmed via the 2026-07-04 task audit. (Original velox note: default vfat mount had =fmask=0022,dmask=0022=, hand-fixed to restrictive; bootctl warned about a world-accessible random-seed file.) - -*** 2026-08-08 Sat @ 04:43:42 -0500 Automated tmp.mount mask for ZFS /tmp -New mask_tmp_mount_for_zfs, called from configure_snapshots' ZFS branch: -masks tmp.mount only when the pool actually carries a dataset mounted at -/tmp (exact match), silent no-op without zfs or without the dataset. Covered -by tests/installer-steps/test_mask_tmp_mount_for_zfs.py; the orchestrator -dispatch pin updated. - -*** 2026-08-08 Sat @ 04:43:42 -0500 Automated CPU microcode install by vendor -New install_cpu_microcode, first in boot_ux so grub-mkconfig and mkinitcpio's -microcode hook both see the installed /boot/<vendor>-ucode.img: vendor_id from -/proc/cpuinfo → intel-ucode / amd-ucode, error_warn on unknown vendor. -Covered by tests/installer-steps/test_install_cpu_microcode.py; boot_ux -sequence pin updated. - -*** 2026-07-04 Sat @ 11:48:24 -0500 Automated syncthing user-service enable in archsetup -archsetup:2263-2271 now installs syncthing and enables the user service (via symlink), so fresh installs no longer leave it installed-but-disabled. Confirmed via the 2026-07-04 task audit. (Original velox note: package installed but service not enabled; hand-fixed with =systemctl enable --now syncthing@cjennings=.) - -*** 2026-08-08 Sat @ 04:43:42 -0500 Closed the awww-daemon crash watch — no recurrence -The April boot crash never recurred across four months of daily use on both -machines (and the wallpaper stack has since been reworked). Reopen as its own -bug with fresh evidence if it ever comes back. - -*** 2026-08-08 Sat @ 04:43:42 -0500 Automated touchpad device detection in the pointer scripts -The scripts were already in stowed dotfiles with binds — the open half was the -hardcoded Framework device name. Both touchpad-auto and toggle-touchpad now -auto-detect the touchpad (first pointer named *touchpad*, pixa fallback) and -derive the internal-pointer exclusion set from the detected name, so they -agree on any machine. Test seams added (--detect / --has-external-mouse); -tests/touchpad-auto/ new, toggle-touchpad suite still green. Dotfiles commit; -velox picks it up on its next pull. - -*** 2026-08-08 Sat @ 04:43:42 -0500 Documented bluetooth pairing in the post-install checklist -Inherently interactive, so it can't ride the installer. Documented in the new -[[file:docs/post-install-checklist.org][docs/post-install-checklist.org]] along -with the Proton Bridge steps — the standing home for manual post-install work. -Consider: document as post-install step. No automation possible. - -*** 2026-05-26 Tue @ 13:32:31 -0500 pocketbook install concern moot — pulled from publication, folded in-tree -Resolved by removing pocketbook from archsetup's provisioning entirely. It's nowhere near ready, so the github mirror + cjennings.net repo were deleted and the project was folded into the archsetup tree at =pocketbook/=. Dropped the =gtk4-layer-shell= dep + =pip_install= from =archsetup= and the clone from =scripts/post-install.sh=. No fresh install pulls pocketbook now, so "not installed on velox" no longer applies. Re-wiring the install is tracked in the new pocketbook development backlog. - -*** TODO Review: Tailscale needs login after install -~tailscaled~ service was enabled but needed ~tailscale up~ for interactive auth. -Old machine entry needed cleanup in admin console. -Consider: document as post-install step. - -*** TODO Review: docs/ directories need manual sync from existing machine -docs/ dirs (gitignored) for ~/code and ~/projects repos needed scp/rsync from ratio. -Same for ~/.emacs.d/docs/. Not in git, so not available after clone. -Consider: document as post-install step or create a sync script. - ** TODO [#B] Test + CI infrastructure :test: :PROPERTIES: -:LAST_REVIEWED: 2026-07-13 +:LAST_REVIEWED: 2026-08-25 :END: Umbrella for the test-harness and CI-automation buildout. Consolidated from the 2026-06-28 task audit: these were scattered top-level tasks circling one effort, re-homed as children so the work reads as a unit. Each child ships independently and keeps the priority it carried before. No CI runner exists yet, so the CI/CD-pipeline child gates several of the others. @@ -1246,7 +2300,7 @@ Keep test runs performant as installs and post-install tests grow (target < 2 ho :LAST_REVIEWED: 2026-05-21 :END: Proactive monitoring integrated with testing -*** TODO [#B] Fix VM cloning machine-ID conflicts for parallel testing +*** TODO [#C] Fix VM cloning machine-ID conflicts for parallel testing :no-sync: :PROPERTIES: :LAST_REVIEWED: 2026-05-21 :END: @@ -1257,9 +2311,9 @@ Need to investigate proper machine-ID regeneration that doesn't break networking Would enable parallel test execution in CI/CD Priority C because snapshot-based testing meets current needs -** TODO [#B] Review undeclared ratio packages for installer inclusion :chore: +** TODO [#C] Review undeclared ratio packages for installer inclusion :chore: :PROPERTIES: -:LAST_REVIEWED: 2026-07-09 +:LAST_REVIEWED: 2026-08-21 :END: Triggered by the 2026-06-14 =make package-diff= run on ratio: 62 packages are installed but not declared in archsetup. Stripped of the structural buckets — pacstrap base/boot/kernel (base, linux*, grub, efibootmgr, sudo, btrfs-progs, fwupd, logrotate, ex-vi-compat, linux-lts-strix, zram-generator), the =make deps= VM set (qemu-full, virt-manager, virt-viewer, libguestfs, bridge-utils, dnsmasq, archiso), and the yay bootstrap — these 40 remain. Check the ones to add to the installer, then rerun =make package-diff= to confirm they clear. @@ -1267,6 +2321,10 @@ Evidence report (2026-07-14, count now 64): [[file:docs/design/2026-07-14-undecl Some entries are libraries likely pulled in as dependencies (blas-openblas, openblas, eigen, tk, lib32-openal, pkcs11-helper, gtk4-layer-shell, webkit2gtk, sane, freerdp, rust-bindgen) — check those only if you want them declared explicitly rather than left to dependency resolution. +git-lfs, imv and libreoffice-fresh are ticked: the installer declares all three +as of 2026-09-25, so a later walk of this list should skip them rather than +re-deriving the case for each. The rest of the list is untouched. + - [ ] aws-cli-v2 - [ ] bats - [ ] blas-openblas @@ -1276,16 +2334,16 @@ Some entries are libraries likely pulled in as dependencies (blas-openblas, open - [ ] flatpak - [ ] freerdp - [ ] geeqie -- [ ] git-lfs +- [X] git-lfs - [ ] github-cli - [ ] gtk4-layer-shell - [ ] hugo -- [ ] imv +- [X] imv - [ ] lc0 - [ ] lc0-network-sm - [ ] ledger - [ ] lib32-openal -- [ ] libreoffice-fresh +- [X] libreoffice-fresh - [ ] minidlna - [ ] openai-codex - [ ] openblas @@ -1308,6 +2366,30 @@ Some entries are libraries likely pulled in as dependencies (blas-openblas, open - [ ] webkit2gtk - [ ] whisper.cpp +*** 2026-08-21 Fri @ 14:28:18 -0700 Dropped to [#C], and the sharper measurement is on velox now +Re-graded [#B] → [#C]. Not a change of mind about the value — nothing has been +ticked since I filed it on 2026-06-14, across two full cycles, and by my own +scheme [#B] means "this cycle" while [#C] is the parking lot. The grade should +say where it actually sits. + +The premise also moved. This list measures ratio, which carries years of +accumulated manual installs tangled up with whatever archsetup put there, so a +package being undeclared says little about whether it matters. Velox is the +better instrument now: rebuilt from archsetup on 2026-08-13 and working, so a +=make package-diff= there compares what the installer declares against what a +machine actually needs, with only days of drift on top. Re-run it on velox +before walking these forty by hand. + +Worth noting the empirical result already came in. The gaps that actually hurt +after that rebuild — rulesets never cloned, the .emacs.d systemd units never +linked, the missing =*.local.*= configs — surfaced on their own, and not one of +them is on this list. That is evidence about what this kind of list catches. + +Related but distinct: =[#B] Installed-package drift audit= below builds the tool +for the opposite direction (declared but missing, and provider substitutions). +If that lands first it plausibly subsumes the detection half of this one, leaving +only the include/ignore judgment. + ** TODO [#B] Installed-package drift audit :chore:packages:solo: :PROPERTIES: :LAST_REVIEWED: 2026-08-02 @@ -1330,7 +2412,7 @@ machine state. ** TODO [#B] Security hardening + audit :security: :PROPERTIES: -:LAST_REVIEWED: 2026-07-14 +:LAST_REVIEWED: 2026-08-25 :END: Umbrella for the security-hardening and audit effort. Consolidated from the 2026-06-28 task audit, re-homing the scattered security tasks as children so the work reads as a unit. Each child ships independently and keeps its prior priority. @@ -1346,12 +2428,12 @@ Umbrella for the security-hardening and audit effort. Consolidated from the 2026 **** TODO [#B] Implement port scanning check **** TODO [#B] Create security posture verification script **** TODO [#B] Set up intrusion detection monitoring -*** TODO [#B] Document threat model and mitigations within 6 months +*** TODO [#B] Document threat model and mitigations :PROPERTIES: :LAST_REVIEWED: 2026-05-21 :END: Identify attack vectors, what's mitigated, what remains -*** TODO [#B] Complete security education within 3 months +*** TODO [#B] Security education :PROPERTIES: :LAST_REVIEWED: 2026-06-24 :END: @@ -1363,9 +2445,9 @@ Read recommended resources to make informed security decisions (see metrics for Practical guidelines for working in public spaces ** TODO [#A] Ensure sleep/suspend works on laptops -SCHEDULED: <2026-08-12 Wed> +SCHEDULED: <2026-09-25 Fri> :PROPERTIES: -:LAST_REVIEWED: 2026-08-08 +:LAST_REVIEWED: 2026-09-17 :END: Raised [#B] → [#A] and scheduled at the 2026-08-08 review: Craig leaves on vacation ~2026-08-15 and velox is the travel machine — suspend and battery @@ -1391,14 +2473,418 @@ Add kernel parameter: ~rtc_cmos.use_acpi_alarm=1~ (will become systemd default) Consider: ~acpi_mask_gpe=0x1A~ for battery drain, suspend-then-hibernate config See Framework community notes on logind.conf and sleep.conf settings +*** 2026-08-17 Mon @ 19:57:42 -0700 Four of the five riders are done; WireGuard is the one left +The riders were written for "when velox returns from repair". It came back as +a full reinstall instead, and the installer carried most of them, so I checked +each on the live machine rather than reading the list back: + +- tlp radio-enable — done. =/etc/tlp.d/01-custom.conf:10= carries + =DEVICES_TO_ENABLE_ON_STARTUP="bluetooth wifi"=, written by the installer. +- touchpad auto-detection — the dotfiles half is done: =touchpad-auto + --detect= prints =pixa3854:00-093a:0274-touchpad=. Read that carefully + though — it names the device the config expects, not a device delivering + events. The touchpad is still dead on the ribbon fault, so this rider is + satisfied and the hardware still is not. +- podman socket — done, =podman.socket= is enabled. +- camera udev — done, =72-usb-passthrough-cameras.rules= is installed. +- *wolf WireGuard — not done, and it is the one that was time-critical.* No + =~/.config/wireguard/wolf.conf.gpg= and no WireGuard profile in + NetworkManager. The 08-08 decision set this up specifically so velox could + reach home from the road, on the argument that it is cheap at home and + expensive from a hotel. velox is now in the hotel. + +The suspend work itself is untouched — no kernel parameter, no drain +measurement. Only the riders moved. + +*** 2026-08-25 Tue @ 11:57:48 -0600 Logged the two Aug 23 hibernate-leg failures +suspend-then-hibernate failed its hibernate leg twice on 2026-08-23 (21:06 and +23:29): "Failed to put system to sleep. System resumed again: Device or +resource busy". Noticed during the 08-24 Lua-port session and parked there; +filed here at Craig's direction so the sleep task carries it. Nothing +diagnosed yet — first step is =journalctl -b -1 -u systemd-suspend-then-hibernate= +around those timestamps to see which device reported busy. + +*** 2026-08-26 Wed @ 16:16:08 -0600 Diagnosed the hibernate battery drain: three separate faults, one task each +Craig hibernated twice in ten days and found the battery dead both times. Read +all 39 boots since the 08-13 reinstall, upower's charge history +(=/var/lib/upower/history-charge-Framewo-55-03F5.dat=, root-only, starts +08-19), sysfs, and the scripts inside =/efi/EFI/ZBM/zfsbootmenu.efi=. + +Hibernate is configured right and has worked: five hibernate+resume cycles +since reinstall (08-13, 08-17 14:31, and three suspend-then-hibernate cycles on +08-20/21). The two fatal events are the two overnight explicit +=systemctl hibernate= runs, 08-17 22:20 and 08-21 19:58. Both journals end at +"PM: hibernation: hibernation entry"; the next power-ons (08-18 10:13, 08-22 +17:22) were fresh boots whose resume hook found no image, no later swapon +reported a leftover suspend signature, and on 08-22 the battery read 2% at +power-on. The 08-23/24 night was on AC and not a battery death (three suspends +failed to enter, machine awake all night at the charge limit; the 09:57 end was +three power-key presses and a hard cut at 63%). The 08-19 death was the +caffeine/hypridle one already diagnosed. + +Three faults, tracked as the children below: +- Hibernate hard-freezes on entry (documented on Framework 13 AMD incl. Ryzen + AI 300: black screen, never powers off, intermittent, amdgpu-side). Fits + everything: the freeze precedes the swap signature, so the next boot is + fresh, and a frozen laptop at ~5 W empties 44.7 Wh in ~8 h. Unprovable from + logs by nature; the alternative (completed hibernate, unattended power-on to + the ZBM passphrase prompt) predicts a surviving image, which neither boot + had — see the VERIFY. +- ZFS ARC starves the hibernate image: "Image allocation is 8118265 pages + short" today 14:09, "390678 pages short" 08-20 09:14. ARC 58 GB of 93, + =zfs_arc_max=0= so =c_max= = RAM − 1 GiB; the kernel must free RAM − + =image_size= (37.4 GB) ≈ 56 GB. systemd falls back to s2idle and retries + every 90 min, so suspend-then-hibernate never actually hibernates. +- The SD card reader (090c:3350, =sda=, no media) can block suspend entirely: + "Freezing remaining freezable tasks failed after 20s (wq_busy=1)", pending + =disk_events_workfn= on =events_freezable_pwr_efficient=, three times on + 08-23/24. On battery that is a dead laptop by morning. + +Mistake worth remembering: =journalctl --since … -k= silently limits itself to +the current boot (=-k= implies =-b=); cross-boot kernel facts need +=_TRANSPORT=kernel= or an explicit =-b=. + +*** TODO Hibernate entry freeze — confirm under observation, then mitigate :bug:velox:hibernate: +Interim rule until this closes: do not hibernate unattended on battery. Shut +down, or suspend on AC. + +What is known: the two dead-battery hibernates match the Framework 13 AMD +"hard freeze on hibernate entry" reports (community threads 69516 and 53860, +Arch bbs 293242): screen black, power LED on, never powers off; intermittent +(one report: every 6–7 cycles); TTM/amdgpu warnings; improved by newer +=linux-firmware=; no confirmed fix. Board A9, BIOS 03.05, linux-lts 6.18.46, +=amdgpu.dcdebugmask=0x610= already on the cmdline. + +Confirm first: the "Hibernate entry freeze: five observed cycles on AC" test +under Manual testing and validation. A failed cycle there is the proof the +journal cannot give. + +Mitigations to try in order once confirmed, one at a time, re-running the +cycles after each: (1) =linux-firmware= at current, then =linux-firmware-git= +if the freeze persists; (2) =/sys/power/disk= = =shutdown= instead of +=platform= (a systemd =HibernateMode=shutdown= drop-in), which skips the ACPI +S4 path some Framework users found hanging; (3) a newer kernel (=linux= vs +=linux-lts=) for the amdgpu delta; (4) unload =mt7925e= in a pre-sleep hook +if the freeze survives the first three. Not =:solo:=: each cycle needs a +person watching the power LED. + +*** TODO ZFS ARC starves the hibernate image — cap it or shrink it pre-hibernate :bug:zfs:velox:solo: +The arithmetic: the kernel preallocates RAM − =image_size= pages before +snapshotting; with 93 GB RAM and the default =image_size= (2/5 of RAM, +37.4 GB) that is ~56 GB, and only free memory plus what shrinkers give back +counts. ARC was 58 GB today and the ZFS shrinker released little inside the +preallocation window, so it came up 31 GiB short. Nothing in +=/etc/modprobe.d/= sets =zfs_arc_max=. + +Two fixes, either or both: +- Cap the ARC: =options zfs zfs_arc_max=<bytes>= in =/etc/modprobe.d/zfs.conf= + (16 GiB leaves ~70 GB reclaimable) plus =echo <bytes> > + /sys/module/zfs/parameters/zfs_arc_max= for the running system. +- Or a =/usr/lib/systemd/system-sleep/= pre hook for the hibernate class that + lowers =zfs_arc_max=, waits for =size= in + =/proc/spl/kstat/zfs/arcstats= to fall, and restores it post-sleep. Keeps + the big ARC while awake. +- Raising =image_size= toward the kernel's ceiling (about half of RAM) also + shrinks the demand; combine with the cap. +Install it through archsetup so the next rebuild carries it (velox-only: ratio +has no swap partition). + +Verify: after the change =arcstats size= drops below the cap within seconds; +then one live suspend-then-hibernate cycle on AC with the delay temporarily +short shows "hibernation exit" and no "Image allocation … short" line in the +journal. That live cycle rides the entry-freeze test above; the ARC half is +checkable without it. + +*** TODO SD card reader media polling can block suspend :bug:velox:solo: +The reader (USB 090c:3350 Silicon Motion, =sda=, "Media removed, stopped +polling" at boot yet =events_poll_msecs= = −1 → default 2000 ms) left a +=disk_events_workfn= item pending on the freezable workqueue three times on +08-23/24, and the freezer gives up after 20 s: "Failed to put system to +sleep … Device or resource busy". Same symptom as the flaky expansion slot in +the ribbon task; a stalled poll never completes. + +Fix: a udev rule for that vendor/product setting +=ATTR{events_poll_msecs}="0"= (or =block.events_dfl_poll_msecs=0= on the +cmdline if every removable disk should stop polling), shipped by archsetup. +Verify with =rtcwake -m mem -s 20= on AC: journal shows "PM: suspend entry" +and "PM: suspend exit" with no "Freezing remaining freezable tasks failed", +and =/sys/block/sda/events_poll_msecs= reads 0 after a replug. Pulling the +card before sleeping is the manual workaround meanwhile. + +*** VERIFY After the 08-17 and 08-21 dead batteries, did the first power-on hang, or boot straight to a fresh login? +Decides between the two mechanisms. An entry freeze leaves no image, so the +next power-on boots straight through. A completed hibernate followed by an +unattended power-on (phantom power button, ZBM passphrase prompt until dead) +leaves the image in place, so the next power-on would try to resume — and the +only way that ends in the fresh boots the journal shows is a hung resume that +got force-cut. If both power-ons went straight to a fresh login, the freeze +is the answer. + +** TODO [#A] Port Hyprland config to Lua before 0.57 drops .conf support :hyprland:dotfiles: +SCHEDULED: <2026-09-24 Thu> +:PROPERTIES: +:LAST_REVIEWED: 2026-09-17 +:END: +Hyprland prints "You are using the .conf config format, support for which will be +removed in Hyprland 0.57" at every start. Installed and in =extra= is 0.56.2-1, so +the *next* release breaks the config. Craig's call 2026-08-24: port now, under no +time pressure, rather than pin the package or wait for the upgrade to force it. + +STATE (2026-08-24): built and verified in a nested compositor, *not deployed*. +I deployed it to the dotfiles tree this afternoon and rolled it back the same +hour on Craig's call — the switch had not been checked on real hardware and the +machine has to stay usable. The dotfiles repo is untouched at =8f692f5= and the +live config is the original =hyprland.conf=; =hyprctl reload= after the rollback +returned zero configerrors and the velox host override is applied +(=xwayland:force_zero_scaling= false), so the per-host chain is intact. + +Everything needed to redeploy is in =working/hyprland-lua-port/= with a README +carrying the step-by-step: the three =.lua= deliverables, plus the two reader +changes saved as patches (=reader-changes-for-lua.patch= for dotfiles' +=dotfiles-validate= and three test suites, =test-desktop-for-lua.patch= for +archsetup's post-install checks). Both patches were verified to apply clean +against their repos, and every assertion in them was mutation-tested — each one +confirmed to go red when the property it guards is removed. Replay them rather +than rewriting the assertions. + +Also settled along the way: =themes/dupre/hyprland.conf= is dead. Nothing sources +it, no apply script exists, and it has silently drifted from the live config +(=dab53dff= / =2c2f32ff= against =daa520ff= / =444444ff=). It needs no porting. + +HOW IT WAS BUILT. =hyprlang2lua= (github.com/EIonTusk/hyprlang2lua, AUR 0.7.1-1) +converts hyprlang to the 0.55+ Lua format and preserves comments. Built from +source into a scratchpad with the local Go rather than installing the AUR package +— one dependency, and no PKGBUILD executed. Run with =--no-merge=, which emits +each config section as its own =hl.config()= call at its original position; the +default merges them into one hoisted call, which both scrambles comment placement +and puts the =conf.d= source glob BEFORE the config it must override. + +THREE DEFECTS THE CONVERTER INTRODUCED, all fixed by hand: +1. Source glob emitted before the merged config block, silently reversing every + per-host override — including velox's =force_zero_scaling = false=, which is + the 2026-08-19 Qt scaling fix. =--no-merge= plus moving the glob to the last + line fixes it. +2. =bind = CTRL $mod, S= became ="CTRL" .. mod .. " + S"= → ="CTRLSUPER + S"=; + same for =CTRL ALT $mod, K=. Proven fatal, not merely odd: Hyprland answers + =hl.bind: failed to parse key string: Unknown keysym: "CTRLSUPER"=. Two dead + keybinds. +3. Super+RETURN is an =exec= of a shell pipeline; the converter pattern-matched + the =hyprctl dispatch layoutmsg= prefix and swallowed =&& sleep 0.05 && ...= + as the layoutmsg argument. That sleep is the one proven load-bearing 19/19. + +Also rebuilt the autostart section: the generator hoists every =exec-once= into +one block at the end and leaves the comments stranded where the commands were. +Each command now sits under its own comment again via =at_start= / =at_shutdown= +/ =at_reload= collectors that the =hl.on()= handlers at the bottom replay. + +VERIFIED by running both configs in a nested Hyprland on a headless output and +diffing runtime state, not by reading: 38 config keys identical (only type +*rendering* differs — =bool: true= where hyprlang prints =int: 1=); +=xwayland:force_zero_scaling= false on both sides, so the host override still +wins; 103 binds registered on both sides with 100 of 103 matching exactly; +autostart list byte-identical to the 16 =exec-once= lines in order; zero +=configerrors=; and no ERR/WARN line in the ported run that is absent from the +original run. + +KNOWN DELTA — the three =bindm= binds. hyprlang reports =mouse: true=, the Lua +path reports =mouse: false=, and the raw bind struct confirms the flag is unset +rather than merely unreported. Not a transcription error: the wiki documents +exactly the spelling used (=hl.bind("ALT + mouse:272", hl.dsp.window.drag(), +{ mouse = true })=), and all four candidate spellings were tested — none sets it. +Reads as a gap in 0.56.2's Lua config manager. Kept the documented spelling: it +is correct upstream, harmless now, and starts working when the gap closes. Per +the wiki that flag is what makes the action fire *while held*, so Super+drag to +move a window may fire once instead of tracking. Settle it in five seconds after +switching; worth an upstream report if it survives 0.57. + +WHAT REMAINS: +1. *Craig decides when to switch.* The port is ready to go in; it has not been + run on real hardware. The gate is the "Hyprland Lua config" test under Manual + testing and validation, which is written to be run right after the switch. +2. *Four review gates travel with the redeploy*, all written up in that README: + exclude any in-repo =retired/= dir from =dotfiles-validate= (its find globs + across slashes and would validate the dead config); add tests for the new + =dotfiles-validate= Lua branch (25 lines, currently zero coverage — proven + vacuous, since stubbing both regexes to =NEVERMATCHES= still passes 15 tests); + guard that =hl_source_glob= stays the last statement (the one invariant the + per-host layer rests on, and archsetup's VM cannot catch it); and sweep the + ~15 prose comments still naming =hyprland.conf=, of which + =waybar-reserve:12= is the load-bearing one. +3. *Redeploy per =working/hyprland-lua-port/README.org=*, which carries the eight + steps and the two traps that bit on 2026-08-24: =make restow hyprland= aborts + on the pre-existing =obsbot-wb-guard.service= conflict in =common= (restow + =hyprland= and the host package individually), and the running Hyprland + rewrites a stub =hyprland.conf= within a second of the symlink vanishing + (silence it with =hyprctl keyword misc:disable_autoreload 1=, stow, set back + to 0). +4. *Push dotfiles before committing archsetup.* One-directional and load-bearing: + archsetup's post-install suite asserts =~/.config/hypr/hyprland.lua= and the + installer clones the dotfiles *remote* (=archsetup:1481=), so a local commit is + not enough. Confirm with =git ls-tree -r origin/main --name-only | grep + hypr/hyprland.lua=. +5. *Do not leave the =.conf= beside the =.lua= as a rollback.* With both present + Hyprland 0.56.2 loads the =.lua= — proven in a nested instance with a fixture + whose =.conf= set =gaps_in=11= and =.lua= set =77=; the result was 77. A + =.conf= left in place buys nothing and only obscures which file is live. Move + it out of the stow package instead. +6. *=bindm='s missing =mouse= flag is worth an upstream report* if it survives + 0.57. Documented spelling, four variants tested, flag never set. +*** 2026-09-17 Thu @ 08:59:59 -0400 Re-dated to 09-24; 0.57 still isn't in the repos +velox runs =hyprland 0.56.2-3= with no Hyprland update pending, so nothing has +broken yet. The switch still needs me at the keyboard for one restart and the +manual test. + ** TODO [#B] Manual testing and validation :test: :PROPERTIES: -:LAST_REVIEWED: 2026-07-09 +:LAST_REVIEWED: 2026-08-23 :END: -Craig's standing checklist of everything that isn't agent-verifiable. Each child is one test in the =verification.md= shape (title, what we're verifying, steps, Expected). A child that fails gets its actual behavior written under it and is promoted to a top-level TODO. 44 checks pending as of the 2026-07-09 audit. +Craig's standing checklist of everything that isn't agent-verifiable. Each child is one test in the =verification.md= shape (title, what we're verifying, steps, Expected). A child that fails gets its actual behavior written under it and is promoted to a top-level TODO. 62 checks pending as of the 2026-08-23 review — up from 44 at the 2026-07-09 audit, so the queue has gained 18 in six weeks and nothing has drained it. A checklist that only grows is on its way to being where tests get filed rather than run; if the next review finds it higher again, the container needs a scheduled sweep rather than another re-stamp. Priority and type tag added by that audit: the task carried neither, which kept the project's largest live container out of the agenda entirely. +*** Airplane console key in the net panel: does it engage, confirm, and let you back out? +What we're verifying: that the new AIRPLANE key actually drives the mode both +ways, that engaging asks first, and that the way out is visible from inside +airplane mode. The GTK widget layer has no unit coverage and the AT-SPI smoke +can't run on velox (no weston/sway, no at-spi-bus-launcher), so this is the +only check that exercises the real wiring. + +Do it on AC, and not while you need the network — it stops tailscale, the VPN, +syncthing, avahi, cups and inbound ssh, and dims the screen. + +- Open the net panel (Super+Shift+N). The CONSOLE row should now show three + keys: DOCTOR, SPEED TEST, AIRPLANE. +- Click AIRPLANE. +Expected: a dialog naming what it will do — wifi off, services stopped, screen +dimmed, CPU to power — with Cancel and an AIRPLANE button. +- Press Cancel. +Expected: nothing happens. Wifi stays up, the key still reads AIRPLANE. +- Click AIRPLANE again, then confirm. +Expected: the key's lamp flashes while it runs, then the faceplate shows the +AIRPLANE badge with a gold lamp, and the key's label changes to LEAVE AIRPLANE. +Wifi is off and the screen is dimmer. +- While engaged, try the faceplate wifi switch. +Expected: it refuses and the status line says to press LEAVE AIRPLANE. It must +NOT name a keyboard shortcut — the old message said Super+Shift+A, which is +push-to-talk. +- Click LEAVE AIRPLANE. +Expected: no confirmation this time, it just runs. Wifi comes back, brightness +returns to where it was, and the key reads AIRPLANE again. +#+begin_src sh :results output +# The services it stopped should be back. Anything listed here is still down. +for s in tailscaled.service avahi-daemon.service cups.service sshd.service fail2ban.service; do + systemctl is-active --quiet "$s" || echo "still stopped: $s" +done +systemctl --user is-active --quiet syncthing.service || echo "still stopped: syncthing (user)" +echo "airplane state: $(cat "${XDG_RUNTIME_DIR}/airplane-state" 2>/dev/null | head -1)" +#+end_src +Expected: no "still stopped" lines, and the state reads mode=off. A service +that was already stopped before you engaged is correctly left alone, so check +it was running first if one shows up. + +*** Hyprland Lua config: does the real desktop come up, and does Super+drag track? +What we're verifying: that the Lua port drives a real Hyprland session the way +the .conf did, and specifically whether the one known delta — the three =bindm= +binds losing their =mouse= flag — actually costs anything. A nested compositor +proved 38 config keys, 103 binds and the host-override chain identical, but it +cannot test real input devices or a real DRM display. + +PRECONDITION: run this *only after* redeploying the port per +=working/hyprland-lua-port/README.org=. As of 2026-08-24 the port is rolled back +and the live config is the original =hyprland.conf=, so running this now just +confirms the old config — which is not what it is for. Run it on velox; the stub +check in the last block is velox-specific. +- Restart Hyprland (log out and back in, or =hyprctl dispatch exit= from a TTY). +- Confirm the desktop comes up: waybar present and not off-screen, wallpaper + restored, dunst notifications working. +#+begin_src sh :results output +# Which config did it actually load, and did anything fail to parse? The log is +# per-instance under the runtime dir, not in ~/.local/share. Resolve the newest +# instance dir rather than reading $HYPRLAND_INSTANCE_SIGNATURE: Emacs runs as a +# daemon that survives the logout in step 1, so a block run from it can still be +# carrying the PREVIOUS session's signature. That path is gone after the restart, +# grep prints nothing, and an empty result under "Expected: names hyprland.lua" +# reads as "the port failed" when it in fact succeeded -- the worst possible +# wrong answer at exactly the wrong moment. +log="$(\ls -td "$XDG_RUNTIME_DIR"/hypr/*/ | head -1)hyprland.log" +echo "reading: $log" +grep -iE '\[cfg\].*(lua|legacy)' "$log" | tail -3 +hyprctl configerrors +#+end_src +Expected: the log names hyprland.lua, and configerrors is empty. +- Hold Super and drag a window with the left mouse button. +Expected: the window tracks the pointer continuously while Super is held. If it +jumps once and stops, the =bindm= =mouse= flag gap is real and costs the drag — +write that here, promote to a top-level TODO, and report upstream. +- Hold Super and drag with the right mouse button (resize), same check. +- Walk the keymap: the launcher, terminal, browser, screenshot chords, the panel + family (Super+Shift+B for bluetooth), workspace switching, layout cycling. +Expected: every chord does what it did before the port. +#+begin_src sh :results output +# The stub .conf should stay gone now that Hyprland started from the .lua. +# Refuse to touch a symlink: on a host that has not been through this port yet, +# ~/.config/hypr/hyprland.conf is still the stow link to the real config, and +# deleting it would report "stays gone" as a pass while having broken the desktop. +f=~/.config/hypr/hyprland.conf +if [ -L "$f" ]; then + echo "REFUSING: $f is a symlink (a live stowed config), not the stub." +elif [ -f "$f" ]; then + rm -f "$f"; sleep 2 + [ -e "$f" ] && echo "REGENERATED — still stubbing" || echo "stays gone" +else + echo "already absent — nothing to do" +fi +#+end_src +Expected: "stays gone". If it regenerates, Hyprland is still resolving its config +to the .conf path and the port is not actually live — stop and investigate. + +*** 2026-09-13 Sun @ 07:57:32 -0500 Retired the lock-screen clock check: fixed, per Craig +Craig reported on 2026-09-13 Sun that the stale clock after a real sleep no longer +happens on velox, so the three-outcome check never needed running. The +parent bug task is closed with the same note. + +*** Lock keybind has no crash or hang recovery +What we're verifying: that a hand-lock is as recoverable as an idle lock. +=hyprland.conf:495= is =bind = $mod, ESCAPE, exec, hyprlock=, which runs the +binary directly. hypridle's =lock_cmd= routes through =screen-lock=, which +relaunches a hyprlock that exits non-zero; the keybind bypasses that entirely. +=~/.local/var/log/screen-lock.log= does not exist on velox, so the watchdog has +never recorded a relaunch — consistent with it rarely being in the path at all. + +- Lock with Super+Escape. +- From another tty (ctrl+alt+F3), log in and run: =pkill -x hyprlock= +- Return to the graphical tty. +Expected: with the keybind as written, the session is left locked with no client +and Hyprland draws its "lockscreen app died" screen. If instead a fresh password +prompt appears, something is already relaunching it and the gap is closed. + +*** Clock/DNS deadlock: does the next abrupt power loss strand velox again? +What we're verifying: that the machine survives an RTC reset unattended. Not the +coin cell, which is new with the 2026-08-13 mainboard and is ruled out. The RTC +did not drift on 2026-08-19, it was reset to exactly 2025-01-01T00:00:16 by an +abrupt power loss at 01:33:18 that left no shutdown sequence in the journal. + +This one can't be scheduled. Run the block the next time velox comes up after an +unexpected power loss, before touching the clock. +#+begin_src sh :results output +echo "--- what did the RTC read at this boot? ---" +journalctl -b 0 | grep -m1 'rtc_cmos.*setting system clock' +echo "--- did systemd have to advance the clock to its build epoch? ---" +journalctl --list-boots | tail -3 +echo "--- sources: is an IP-addressed one selected? ---" +chronyc -n sources +echo "--- clock + DNS ---" +timedatectl | grep -iE 'Local time|RTC time|synchronized' +getent hosts gnu.org || echo "DNS DEAD" +#+end_src +Expected: even if the RTC came up at 2025-01-01 and systemd advanced the clock +to 2026-07-23, chrony reached 162.159.200.1 without DNS, stepped the clock to +now, and names resolve. You did nothing. + +If instead the clock is still wrong or DNS is dead, the fix did not hold in the +field despite holding under a simulated skew. Capture that whole block and +promote this to a top-level TODO. + *** Floating layout: freeze positions, border flash, glyph, exit to master What we're verifying: the rebuilt floating mode (Super+Shift+F) floats every window on the workspace via per-window setfloating (the old workspaceopt allfloat was deprecated and no-op'd, which is why nothing floated), freezes each in place, flashes the border gold on entry and exit, flips the waybar glyph to the floating icon, and exits to master. Live-verified on a headless output already (windows floated in place, dragged to overlap, glyph read Floating, toggled back clean); this is the on-your-own-monitor confirmation. - Go to a workspace with 2-3 tiled windows in master. @@ -1990,9 +3476,35 @@ NOTE (2026-07-04 audit): the "four-tab panel" framing predates the instrument-co - Expected: ↓/↑ Mbps + ping + server shown inline. - Byte-rate→Mbps unit: VERIFIED 2026-06-30 (velox). Raw =speedtest-go --json= dl_speed read ~3.66M, unambiguously bytes/s (29 down / 80 up Mbps); =net speedtest= reported 33.62 / 77.99 through the wired path. =BYTES_PER_SEC = True= + =* 8 / 1e6= are correct, no flip needed. Remaining here is only that the panel renders the inline result. +*** Hibernate entry freeze: five observed hibernate cycles on AC +What we're verifying: whether velox hard-freezes on hibernate entry (black +screen, power LED on, never powers off), the documented Framework 13 AMD +failure that fits both dead-battery events. The journal cannot show it; a +person watching the LED can. +- Plug in AC, lid open, nothing important unsaved. +- Note the cycle number, then hibernate from a terminal: +#+begin_src sh :results output +date; systemctl hibernate +#+end_src +- Watch: the screen goes black; within about two minutes the power LED goes + off and the fans stop. +- Press power, enter the ZBM passphrase, and confirm the same session comes + back (windows still open). +- Check that the cycle was a real hibernate and not a fallback: +#+begin_src sh :results output +journalctl -b -o short-iso | grep -E "systemd-sleep|hibernation (entry|exit)|Image allocation|Failed to put" | tail -6 +#+end_src +- Repeat until five cycles are logged. +Expected: all five cycles power off within two minutes and resume into the +same session, with "hibernation exit" and no "Image allocation … short" line. +A cycle where the screen stays black with the power LED on for more than five +minutes is the entry freeze: hold power for 10 s, and write down the cycle +number and whether the keyboard backlight was lit. A cycle that instead comes +straight back with "Cannot allocate memory" is the ARC task, not a freeze. + ** DOING [#B] Prepare for GitHub open-source release :PROPERTIES: -:LAST_REVIEWED: 2026-07-09 +:LAST_REVIEWED: 2026-08-17 :END: Remove personal info, credentials, and code quality issues before publishing. *** 2026-07-21 Tue @ 08:00:00 -0500 Audit reconcile: the four assets/ "& Claude" author lines are fixed @@ -2049,6 +3561,19 @@ Recommend: fresh repo for GitHub (keep cjennings.net remote with full history). History is now 589 commits (the 2026-05-11 note's "275" is stale). Only the calendar-feed file has been filter-repo'd so far (2026-05-20). The five credential files remain in history at their pre-=b10cba5= paths: =.tidal-dl.token.json= (5 commits), =calibre/smtp.py.json= (6), =transmission/settings.json= (5), =.msmtprc= (8), =.mbsyncrc= (9). None are tracked in the current tree. The scrub-or-fresh-repo decision still stands. ***** 2026-07-04 Sat @ 11:48:24 -0500 Count refresh — history now 565 commits; re-verify the 5-file claim before scrubbing The 2026-07-04 audit found the history is now 565 commits, down from the 589 recorded above. Because the count dropped, re-verify that the five credential files are still present in history (re-run the per-file =git log --all -- <path>= check) before relying on the scrub scope — the earlier count is stale and the file set may have moved. +***** 2026-08-17 Mon @ 10:20:00 -0700 Corrected the paths — every prior check has been querying paths that never existed +The five filenames recorded above are not the paths these files live at, and =git log -- <path>= answers "no commits" for a path it has never seen rather than erroring. So the checks return a clean result and mean nothing. The real paths, from =git log --all --name-only --diff-filter=A= over the full history, all sit under the pre-migration =dotfiles/= tree: + +- =dotfiles/system/.msmtprc= (3 commits) +- =dotfiles/system/.mbsyncrc= (2) +- =dotfiles/system/.config/calibre/smtp.py.json= (2) +- =dotfiles/system/.config/transmission/settings.json= (2) +- =dotfiles/system/.config/.tidal-dl.token.json= (2) +- =dotfiles/system/.config/.tidal-dl.json= (2) — a *sixth* file, never recorded here + +Use those paths for any future check, not the bare filenames. History is 891 commits; none of the six are in the current tree. The scrub-or-fresh-repo decision still stands and its scope is six files, not five. + +This surfaced while re-verifying on velox, where the check ALSO returned a false clean for a second, unrelated reason: the clone was shallow (7 commits), so it could not see the history either way. Both failures produce the same confident zero. Filed as =[#A] The installer shallow-clones the two repos I develop in=. ***** 2026-07-21 Tue @ 08:00:00 -0500 Re-verified: history now 851 commits; five files still present, per-file counts dropped 2026-07-21 audit re-verification. History is now 851 commits (=git rev-list --all --count=). The five credential files are still in history but at fewer commits each than the 2026-06-28 record: =.tidal-dl.token.json= 3 (was 5), =calibre/smtp.py.json= 4 (was 6), =transmission/settings.json= 3 (was 5), =.msmtprc= 5 (was 8), =.mbsyncrc= 6 (was 9). None are in the current tree. The scrub-or-fresh-repo decision still stands; the scope is smaller than recorded. @@ -2107,9 +3632,9 @@ Rewrote the bare =if $var= boolean conditionals (=show_status_only=, =fresh_inst *** 2026-05-26 Tue @ 15:27:09 -0500 eval task moot — the line-434 eval is gone, the survivor is deliberate Verified: the only =eval= left in =archsetup= is line 578 in =retry_install=, and it's intentional and documented — it captures =$?= directly from =eval "$cmd"= to dodge the if-compound-swallows-exit-code trap. Replacing it with an array would reintroduce that bug. The line-434 eval this task pointed at no longer exists. Nothing to change. -** TODO [#B] The audio doctor never checks the microphone :bug:audio: +** TODO [#C] The audio doctor never checks the microphone :bug:audio: :PROPERTIES: -:LAST_REVIEWED: 2026-07-13 +:LAST_REVIEWED: 2026-08-25 :END: The classifier is output-only. =diag.probe_semantic= already collects =default_source= and =default_source_present=, and =classify.py= reads neither: the word "source" appears once in the whole module, in the graph row that counts them. So a muted mic, a default source naming an unplugged device, or a mic at zero volume all classify as =healthy=, and the verdict prints "the default output is present and audible" while the input side goes unexamined. Found 2026-07-10 while asking whether the doctor would have caught Chrome losing the mic. It would not have. @@ -2119,20 +3644,15 @@ Work: mirror the sink rules onto the source. =probe_semantic= gains =default_sou Two things not to get wrong. An absent microphone is legitimate on a desktop, so "no input devices" must never be a fault the way =no-output-devices= is. And a monitor source is a legitimate default source (recording desktop audio), which is why =probe_semantic= passes =include_monitors=True= — inheriting the panel's display filter here would call a working setup broken. -Specced 2026-07-10 after discussion with Craig, and the design grew past the original gap: [[file:docs/specs/2026-07-10-audio-doctor-input-side-spec.org][docs/specs/2026-07-10-audio-doctor-input-side-spec.org]] (DRAFT, four decisions open). A doctor key per direction, a kernel-level capture probe below PipeWire, PTT-aware muting, and a direction-aware guard. The precedence question the build would have faced is gone: a doctor per direction means the user's press says which side they came to fix. +Specced 2026-07-10 after discussion with Craig, and the design grew past the original gap: [[file:docs/specs/2026-07-10-audio-doctor-input-side-spec.org][docs/specs/2026-07-10-audio-doctor-input-side-spec.org]] (DRAFT, three decisions open as of 2026-08-25). A doctor key per direction, a kernel-level capture probe below PipeWire, PTT-aware muting, and a direction-aware guard. The precedence question the build would have faced is gone: a doctor per direction means the user's press says which side they came to fix. Parent spec: [[file:docs/specs/2026-07-09-audio-doctor-spec.org][docs/specs/2026-07-09-audio-doctor-spec.org]] (IMPLEMENTED). This is a v1 gap found after the fact, not a phase of it. -** DONE [#C] Waybar modules run together — need subtle separators :bug:dotfiles:waybar: -CLOSED: [2026-08-08 Sat] -Closed at the 2026-08-08 task review: Craig confirms the separator work landed -a while back and the bar reads correctly now. -Craig misreads where one module ends and the next begins — the wind (weather) value runs straight into the date with no visual stop, so he reads the wind figure as the start of the date. Add a light, subtle separator or spacing between adjacent Waybar modules. -Grading: Minor severity (legibility, nothing broken) x frequent (every glance at the bar) = P3 = [#C]. -Not fully :solo: — needs Craig's eye on the result (separator style is a taste call, plus a live visual check). Prior work added a date-facing divider (dotfiles 103cccb); evidently not enough, so revisit the whole inter-module treatment rather than just the weather/date seam. From .emacs.d handoff 2026-07-20-1114 (roam capture; waybar is archsetup-owned per the dotfiles standing rule). +Grading (2026-08-25 review): Major severity — the doctor's verdict is silently wrong for a whole direction, workaround is checking the mic by hand — × "some users, sometimes" (mic faults are occasional) = P3 = [#C]. Was held at [#B] ungraded; regraded by the matrix. + ** TODO [#C] Weather chip color signals unclear + unenforced :bug:dotfiles:waybar:weather: :PROPERTIES: -:LAST_REVIEWED: 2026-07-21 +:LAST_REVIEWED: 2026-08-26 :END: From the roam inbox (2026-07-20): the shipped Waybar weather chip's comfort coloring reads as noise — it shows amber for no clear reason, and some items are bolded, which isn't a legible signal. Craig's intended scheme (every item except the arrow key colored by whether the weather is comfortable; NO bold or italic anywhere): - Normal — all text white: temp in 60-85; condition sunny/clear/etc. @@ -2179,36 +3699,31 @@ already consumes, so the indicator pays nothing new. Where NM can't guess panel affordance for it can come later. Live phone-hotspot check is Craig's manual-testing entry; everything else verifies with fakes. -** CANCELLED [#C] Add a whole-display dim mode :feature:hyprland: -CLOSED: [2026-08-08 Sat] -Killed at the 2026-08-08 task review: the July auto-dim work covers the actual -need; no separate dim-everything mode wanted. -Extend auto-dim with an explicit “dim everything” setting for bright -non-dark-mode contexts, with a security/usability review of its scope. - ** TODO [#C] Net panel speedtest history :feature:dotfiles:network: :PROPERTIES: -:LAST_REVIEWED: 2026-07-14 +:LAST_REVIEWED: 2026-08-25 :END: From the roam inbox (routed 2026-07-13): the networking panel should track speedtests over time with appropriate info. Shape: persist each SPEED TEST result (timestamp, down/up, latency, server) to a small local store and surface history in the net panel. Design questions for work time: retention window, which fields matter, and presentation within the panel's ~400px width (recent-results list vs trend readout). Point-in-time results exist today; the gap is comparison across days and venues. ** TODO [#C] zfs base VM image build failure: ZFS DKMS module missing :bug:zfs: :PROPERTIES: -:LAST_REVIEWED: 2026-07-09 +:LAST_REVIEWED: 2026-08-17 :END: =FS_PROFILE=zfs make test-vm-base= fails inside the VM at initramfs time: archangel reports "ZFS module not found! DKMS build may have failed" against the installed kernel (linux-lts 6.18.38 at the 2026-07-08 attempt). Consequences: the maint scenario harness's zfs lane (Phase 12) is filtered but unexercised, and a real zfs bare-metal install via archangel would plausibly hit the same wall. Priority per the bug matrix: Major severity (zfs install path broken) × some-users-sometimes = P3. When fixed, run =FS_PROFILE=zfs bash scripts/testing/run-maint-scenarios.sh --list= and add zfs scenario files (zpool scrub / autotrim / snapshot destroy) to the harness. +*** 2026-08-17 Mon @ 10:08:51 -0700 Rechecked: archzfs still on 2.3.3, still blocked +Ran the unblock check from the diagnosis below: archzfs' x86_64 index still serves only =zfs-dkms-2.3.3=. The first release supporting 6.18 is 2.4.0, so the blocking condition is unchanged and there is still nothing on our side to fix. Recheck again with the same one-liner. *** 2026-07-14 Tue @ 01:40:48 -0500 Diagnosed: OpenZFS/kernel version skew, blocked on archzfs Reproduced in ~1 minute of install: =dkms install zfs/2.3.3 -k 6.18.38-2-lts= exits 1 during pacstrap. Root cause confirmed: OpenZFS 2.3.3's META declares Linux-Maximum 6.15, and the VM installs linux-lts 6.18.38. The first release supporting 6.18 is 2.4.0 (2.4.1 covers 6.19), and archzfs currently serves only zfs-dkms 2.3.3-1 — nothing on our side to fix. Unblock condition: archzfs publishes zfs-dkms ≥2.4.0; recheck with =curl -s https://archzfs.com/archzfs/x86_64/ | grep -o 'zfs-dkms-[0-9.]*'=, then rerun =FS_PROFILE=zfs make test-vm-base=. ** TODO [#C] Waybar collapse control: replace the triangle glyph :feature:waybar: :PROPERTIES: -:LAST_REVIEWED: 2026-07-14 +:LAST_REVIEWED: 2026-08-26 :END: From the 2026-07-04 roam capture. The waybar collapse mechanism (click the triangle, the bar sections redisplay shortened) works, but the triangle glyph doesn't match the instrument-console aesthetic the panels now use. Replace it with something in keeping with the console look. Aesthetic decision — bring Craig two or three concrete glyph/style options (a machined chevron, a console-key style expander, an engraved caret) before wiring. Dotfiles waybar config (handled per the archsetup-owns-dotfiles rule). Raised alongside the net-panel/audio speedrun; deferred from it because the glyph choice is a taste call. ** TODO [#C] Net panel: driver-health diagnostic tier :feature:network: :PROPERTIES: -:LAST_REVIEWED: 2026-07-14 +:LAST_REVIEWED: 2026-08-26 :END: Follow-up from the 2026-07-04 net-panel hardening speedrun (Craig's cj question on the no-WiFi item). The shipped no-wifi-hardware verdict covers "no adapter at all." This tier covers "adapter present but the driver is wedged": read-only health signals — =ip link= (device present but no-carrier / down), =dmesg= / =journalctl -k= for firmware-load failures, =rfkill= for a hard block, =modinfo= / =lsmod= for the driver module — classified before a generic reset. Remedy actions: a privileged =modprobe -r <mod> && modprobe <mod>= reload of the wifi driver, and a firmware-package pointer when the failure is a missing/failed firmware load. Dotfiles net-package work (handled per the archsetup-owns-dotfiles rule). Design pass first to decide whether it's worth a repair tier vs a needs-user-action pointer. @@ -2222,45 +3737,9 @@ Tool choice is the open decision (needs Craig): =nerd-dictation= (Vosk, lighter, *** 2026-07-21 Tue @ 08:40:00 -0500 Decided (Craig): whisper.cpp + wtype, system-wide STT engine = =whisper.cpp= (accurate offline, optional GPU on ratio's Radeon). Typing backend = =wtype= (Wayland-native virtual-keyboard injection into the focused window, no root/daemon), with =ydotool= (uinput) held as a fallback only if a specific app — some XWayland/Electron surface — won't accept wtype's synthetic input. One system-wide path that also covers Emacs buffers and the Claude Code prompt; the Emacs-native =whisper.el= route was NOT chosen. Build scope: whisper.cpp + a model (start with a mid-size English model, tune later), a Hyprland push-to-talk keybind driving a record→transcribe→wtype pipeline, and an autostart/service entry, folded into archsetup so it lands on ratio + velox. Now unblocked (agent-buildable; verification includes a live dictation check). -** DONE [#C] Fix install errors surfaced by the 2026-05-11 VM test run -CLOSED: [2026-08-08 Sat] +** TODO [#C] Osbot camera configuration :chore:quick: :PROPERTIES: -:LAST_REVIEWED: 2026-07-06 -:END: -Closed at the 2026-08-08 task review: every archsetup-attributable error was -fixed and verified (fontconfig, dconf x2, emacs-stow, AUR exit-0 logging at -the root); the residual four reproduce unchanged and are diagnosed -environment/non-critical, with two 2026-06-28 full runs attributing zero -issues to archsetup. Residual thread: confirm the firewall nf_tables pair on -bare metal at the next real install — no container task needed to carry it. -*** 2026-06-28 Sun @ 13:29:29 -0400 Audit reconcile: 2026-06-28 btrfs+zfs runs reproduce the same residual set -Newer full runs landed since the 2026-06-11 reconcile below: the 2026-06-25 zfs run (Testinfra 96/0) and the 2026-06-28 btrfs+zfs runs (97/0, "zero attributed issues"). The residual four were NOT fixed and reproduce unchanged: =enabling firewall= (archsetup:1496-1498, carries a VM-kernel note), =enabling gamemode for user= (archsetup:2221, non-critical), and =tidaler (AUR)=. Zero archsetup-attributed Testinfra issues across both profiles confirms these are environment / non-critical, not archsetup bugs. Bare-metal confirmation of the firewall pair is still the open thread. - -*** 2026-06-15 Mon @ 23:53:21 -0500 Audit reconcile: latest VM run (2026-06-11) confirms the surviving error set -The most recent VM run (=test-results/20260611-113904/=) carries four error-summary entries: =enabling firewall= + =verifying firewall is active= (the iptables/nf_tables "Could not fetch rule set generation id" pair, still unconfirmed on bare metal), =enabling gamemode for user= (non-critical), and =tidaler (AUR)=. The earlier fontconfig/dconf fixes held — none reappear. So the count is down from the 7→6 anchor below to four, all of them the known-residual items already itemized. -Errors logged during the VM install. Status as of the 2026-05-11 18:36 run (=test-results/20260511-183643/archsetup-output.log=) after the =48c9439= fontconfig/dconf fix: 7 → 6. -- refreshing font cache — RESOLVED in =48c9439= (now installs =fontconfig= before calling =fc-cache=). -- configuring GTK file chooser — RESOLVED in =ecab29f= (switched to a system-wide dconf db at =/etc/dconf/db/site.d/=; needs no session bus during install). -- configuring GNOME interface settings in dconf — RESOLVED in =ecab29f= (same fix as the GTK file chooser above). -- enabling firewall — exit 1: =iptables v1.8.13 (nf_tables): Could not fetch rule set generation id: Invalid argument=. Still present in the 18:36 run; likely a VM-kernel/nf_tables artifact — confirm on bare metal before treating as an archsetup bug. -- verifying firewall is active — exit 1 (follow-on from the firewall-enable error). -- enabling gamemode for user — exit 1 → step "gaming" FAILED — non-critical. -- tidaler (AUR) — logged in the error summary with exit code 0 (odd; logging quirk or transient AUR build noise?). -Also seen in the 18:36 run's log-diff (post-install systemd noise, probably VM-environment): =pam_systemd … CreateSession failed= / =logind: Failed to start session scope … Permission denied=, and =Failed to start Proton VPN Daemon= (no VPN config in the test VM). - -*** 2026-05-19 Tue @ 13:18:56 -0500 Fixed AUR exit-0 logging bug at the root -Root cause was in =retry_install=: =last_exit_code=$?= ran AFTER =if eval ...; then return 0; fi=. Bash defines an if-compound's exit status as zero when no condition tested true, so a failing eval's exit code got overwritten with 0 before reaching =error_warn=. Fix in =8221c54=: capture =$?= from =eval= directly into a local var, then compare against the captured value in the if. VM-verified in =test-results/20260519-115318/=: =mkinitcpio-firmware (AUR)= and =tidaler (AUR)= now report =error code: 1= (yay's actual exit) instead of the misleading =error code: 0=. The same packages still appear in the summary because yay returns non-zero when sub-deps fail to build (e.g. =aic94xx-firmware=), but the codes are accurate now. If the underlying sub-dep failures stay noisy, that's a separate concern — open a new task. - -*** 2026-05-16 Sat @ 09:00:41 -0500 AI Response: Surfaced the expanded AUR-exit-0 pattern -2026-05-16 07:40 VM run passed (52/0/5) with the same warning profile as the 2026-05-11 18:36 run. Error count went 7 → 13: 5 fixed/unchanged, +5 new AUR-exit-0 entries (broadens the existing tidaler item into the dedicated =[#B]= subtask above), +1 genuinely new error in =setting up emacs configuration files= (=git pull= ran in =~/.emacs.d= which existed from stow but had no =.git=). Patched =archsetup:1932-1945= with a three-branch check: clone if missing/empty, pull if =.git= exists, =git init=/=fetch=/=checkout= in place if the dir came from stow. - -*** 2026-05-19 Tue @ 01:25:26 -0500 Verified the b9907c7 emacs-stow fix end-to-end -=make test= 21:44 → 22:29 (42 min), =test-results/20260518-214516/=. 52/0/5, =ArchSetup Exit Code: 0=. The third-branch path fired correctly — install log =archsetup-2026-05-18-21-45-46.log:14358-14365= shows =From https://git.cjennings.net/dotemacs= → =[new branch] main -> origin/main= → =Reset branch 'main'= → =branch 'main' set up to track 'origin/main'=. No exit-128, no =fatal: not a git repository=. Error Summary down to 7 (was 13 on 2026-05-16); the emacs entry is gone. AUR exit-0 logging triggered for 2 packages this run (mkinitcpio-firmware, tidaler) vs 6 on 2026-05-16 — same bug class, fewer triggers, still tracked under =[#B] AUR exit-0 logged as error=. Issue Attribution: 1 ARCHSETUP entry (Proton VPN Daemon failed — known VM-no-VPN-config artifact). Cleanup ran clean via the normal path. - -** TODO [#B] Osbot camera configuration :chore:quick: -SCHEDULED: <2026-08-08 Sat> -:PROPERTIES: -:LAST_REVIEWED: 2026-08-08 +:LAST_REVIEWED: 2026-09-17 :END: Re-graded [#C] → [#B] and scheduled at the 2026-08-08 review: Craig leaves on vacation in a week (~2026-08-15), so this needs to land before then. Blocked @@ -2269,12 +3748,17 @@ only on the camera being physically plugged in (no /dev/video node as of Craig's roam capture 2026-07-20, routed via .emacs.d sentry inbox-zero as archsetup-owned device setup: "configure osbot camera." Scope to define at pickup (device model, what "configure" covers — kernel module, v4l settings, default framing). *** 2026-07-21 Tue @ 08:10:00 -0500 Scoped (Craig): tiny — it works, just needs configuring via a panel Craig: the camera works (bought for being Linux-friendly), it just needs configuring. There IS a config panel — =cameractrls= 0.6.10 is installed (a GTK GUI for camera controls: exposure, white balance, PTZ, focus, framing) plus =v4l-utils= for the CLI path. Caveat found 2026-07-21: no =/dev/video*= device is present right now, so the camera isn't currently plugged in / its UVC node isn't enumerated. Task: with the camera connected, confirm it enumerates as a /dev/video node, then set defaults in cameractrls. Small, mostly a live-hardware step. +*** 2026-09-17 Thu @ 08:59:59 -0400 Re-graded [#B] → [#C] and unscheduled +The 08-08 raise was only for the vacation deadline, which has passed, and the +camera is at home. Pick it up the next time the camera is plugged in. ** TODO [#C] Re-check python-lyricsgenius --skipinteg workaround :chore:solo: :PROPERTIES: -:LAST_REVIEWED: 2026-07-09 +:LAST_REVIEWED: 2026-08-17 :END: archsetup installs =python-lyricsgenius= with =--mflags --skipinteg=, skipping makepkg integrity + PGP checks — a workaround originally for an expired-signature issue upstream (surfaced by the 2026-06-23 --noconfirm audit). Periodically test whether the cause has cleared: if a plain =aur_install python-lyricsgenius= builds without complaint, drop the =--skipinteg= workaround. Removal needs a real AUR build to confirm, so it isn't a blind change. +*** 2026-08-17 Mon @ 10:08:51 -0700 Rechecked: still needed, cause unchanged +Fresh AUR clone, =makepkg --verifysource= on 3.7.0-1: the PyPI tarball passes, =LICENSE.txt= still FAILS its b2sum. The PKGBUILD still pins the license at github master, so its checksum drifts whenever upstream touches the file. =--skipinteg= stays. *** 2026-07-23 Thu @ 02:20:00 -0500 Rechecked: still needed, unchanged Fresh AUR clone, =makepkg --verifysource= on 3.7.0-1 (PKGBUILD still unchanged since the last check): the PyPI tarball passes, =LICENSE.txt= still FAILS its b2sum. Same structural cause — the source pins the license at github master, so its checksum drifts whenever upstream touches the file. =--skipinteg= stays. Nothing to change in the installer. @@ -2305,15 +3789,6 @@ The goal is a single place to edit each config, not two. *** 2026-07-21 Tue @ 08:00:00 -0500 Audit reconcile: single-theme now, so the switch-revert pressure is lower The theme system went single-theme — Hudson was retired (dotfiles e03436e; documented archsetup 98142bb, 2026-07-18), leaving Dupre as the only theme. So "edits get overwritten on theme switch" now bites only rarely (switches are effectively nonexistent). The structural two-places-to-edit problem still stands and is worth fixing, but the urgency the original grading implied has dropped. -** CANCELLED [#C] Review current tool pain points annually -CLOSED: [2026-08-08 Sat] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-06 -:END: -Killed at the 2026-08-08 task review: an undated annual intention that never -fired — pain points get surfaced organically as they bite. -Once-yearly systematic inventory of known deficiencies and friction points in current toolset - ** TODO [#D] Installer + scripts refactor opportunities :refactor:solo: Grading: no behavior change; parking lot. 10 refactors remain from the sentry audit — duplicated GPU-modalias scan, triple hand-rolled retry loop, stow x4, display_server/window_manager dispatch dup, Maia ELO range x3, per-script log helpers, GRUB/snapper/fsck sed clusters, waybar-battery positional sed. Full list with line numbers in [[file:docs/design/2026-07-19-sentry-code-findings.org][sentry code findings]] (High/Medium/Low tagged). Pull individual ones out as their own tasks when tackled. The system-mutation sed clusters (snapper/fsck/GRUB/waybar) want characterization coverage before any rewrite. *** 2026-07-20 Mon @ 16:35:00 -0500 Extracted validate_yesno and the NVIDIA_MIN_DRIVER constant @@ -2366,11 +3841,28 @@ Parse yay errors and provide specific, actionable fixes instead of generic error ** TODO [#D] Improve progress indicators throughout install Enhance existing indicators to show what's happening in real-time -** TODO [#C] Telega coredump recurrence tell :bug:maint: +** TODO [#D] Telega coredump recurrence tell :bug:maint: :PROPERTIES: -:LAST_REVIEWED: 2026-07-21 +:LAST_REVIEWED: 2026-09-17 :END: +*** 2026-09-17 Thu @ 08:59:59 -0400 Re-graded C → D: the tell has been quiet on both machines since at least 08-28 +Grading: Minor severity (a crashed telega-server restarts; the cost is coredump +noise and a chat client that blinks) x rare edge case (no occurrence on either +machine in the ~3 weeks the records cover) = P4 = [#D]. + +This moves the letter by moving an input, not by overruling the 2026-07-21 raise +below. That raise was correct on its own frequency row: the tell had just fired, +repeatedly, across five days. It has since stopped, so the row moved from +most-users-frequently to rare-edge-case and the letter follows it. If the +assertion reappears the row moves back and so does the grade. + +Checked at the 2026-09-17 review. ratio: =coredumpctl list telega-server= finds +nothing, and its coredump records reach back to 2026-08-28 (the oldest is an +unrelated Hyprland SIGABRT); =~/.telega/telega-server.log= (3.0 MB, last written +09-13) holds zero =tdat_plist_value= assertions. velox: no telega-server +coredumps either. Back to a watch item; the tell and the durable escape are +unchanged. *** 2026-07-21 Tue @ 08:10:00 -0500 Re-graded D → C (Craig): treat as the actionable version-skew recurrence Craig's call: the fired tell counts as the version-skew recurrence this task predicts (zevlg =:latest= outran the installed elisp again), not just benign host noise — so it bumps [#D] → [#C]. Action: re-pin / upgrade the TDLib side (upgrade the elisp telega package on ratio+velox, or move to a host-native pinned TDLib build) so the server's plist parser and the installed elisp agree again. The durable escape remains the host-native pinned TDLib build. @@ -2384,1208 +3876,190 @@ Verified rather than assumed: =~/.telega/telega-server.log= carries zero =tdat_p Re-graded =[#C]= → =[#D]= per the bug matrix. There is no defect to fix here; it is a watch item with a named tell, and the severity × frequency read is cosmetic (host coredump noise on a metric we own) × rare edge case → P4 → =[#D]=. It stays on the list only so the tell isn't lost. The maintenance console's coredump metric flagged telega-server on ratio (8 coredumps) and velox (18). Root cause was a version skew: the Dockerized =zevlg/telega-server:latest= is frozen at the 2026-06-05 build while the installed elisp lagged at 20260513, so the newer server's plist parser choked on the older elisp's output. .emacs.d fixed it by upgrading telega to 20260706 on both machines (docker kept, =docker pull= is a no-op against the frozen image). Host-coredump pollution should stop. If zevlg later pushes a =:latest= that outruns the installed elisp, the skew and the coredumps recur — the tell is a fresh =tdat_plist_value:500= assertion in =~/.telega/telega-server.log=. The durable escape is a host-native pinned TDLib build, at the cost of an AUR source build. -* Archsetup Resolved - -** DONE [#C] Net panel: Enterprise error never dismisses :bug:dotfiles:network: -CLOSED: [2026-07-12 Sun] -Fixed in dotfiles =a157bed=. Root cause: error toasts are sticky by design (so background refreshes can't wipe an unread error), but the enterprise join hint's flow posts no follow-up status and row clicks post none either, so nothing ever replaced it. Fix: a window-wide capture-phase click gesture dismisses a sticky toast on the user's next interaction; policy in =viewmodel.toast_action_plan= (unit-tested), timed toasts and background clears unchanged. Panel smoke run confirms launch/doctor/close with the gesture installed. Pointer-level dismiss is a manual-testing child (AT-SPI can't drive pointer gestures). Repro screenshot: =~/pictures/screenshots/2026-07-10_195911.png=. -** DONE [#C] Net diagnostics leak connection names + SSIDs into copyable report and --json :bug:dotfiles:network:solo: -CLOSED: [2026-07-12 Sun] -Resolved in dotfiles =df1543a=: the =redact_ssid= toggle now scrubs saved profile names, active SSIDs, envelope-carried names, and =.nmconnection= keyfile basenames from the copyable report and the diag/doctor =--json= envelopes (one systemic pass in =redact.py=; MAC/IP scrub applies to those envelopes too). On-screen output and functional envelopes (status/list) unchanged. 15 new tests; live-verified on ratio (toggle on removes the active connection name from =diagnose --json=, default unchanged). -The net doctor's copyable report (=report.py=, =scrub_text=) scrubs only MAC/IP, and =net diag/doctor --json= (=cli.py=) dumps the raw dict with no redaction. SSID redaction lives only in the event log (=redact_event=, gated on =redact_ssid=, default off). So a connection name (usually the SSID) appears in the clear in the link-step evidence and in every =--json= consumer — the copyable report is exactly the text a user pastes into a bug report. Secrets (PSK/password/token/portal URL) are already stripped, so this is names, not credentials: Minor severity, graded on severity alone per the privacy carve-out. - -Split out of the 2026-07-11 net-doctor-expansion spec review: that spec's new rival-manager/keyfile-perms verdicts keep parity with this pre-existing behavior rather than half-solve it. Fix shape: extend redaction to cover the connection name + keyfile basename across the copyable report and =--json= (one systemic pass, not per-verdict), with a redaction test. Engine-wide, so it wants one coherent change rather than being bolted onto the expansion work. -** DONE [#B] Bt doctor expansion v1 — build the READY spec :feature:dotfiles:bluetooth: -CLOSED: [2026-07-12 Sun] -:PROPERTIES: -:SPEC_ID: 3d4d61c4-e5df-44e9-b8e0-40b31452c3f7 -:END: -Build the [[file:docs/specs/2026-07-11-bt-doctor-expansion-spec.org][bt doctor expansion]] (IMPLEMENTED). Adds a dmesg firmware-hint probe (names the missing blob on a no-adapter fault) and a boot-enablement probe (catches an adapter disabled at boot) to the shipped bt doctor (=~/.dotfiles/bluetooth/=). Archsetup owns the dotfiles work end to end. All phases shipped and fake-verified (d19fdca, f05a9b4, d7d859f); the live reboot-persistence half is on the manual-testing checklist. -*** 2026-07-11 Sat @ 03:06:32 -0500 Built the two read-only probes -New module =bluetooth/src/bt/probes.py= plus =doctor.py= wiring, on dotfiles main (=d19fdca=, pushed). Two reads the diagnose chain never did: =firmware_hint()= scans the current boot's kernel log for per-vendor firmware-load failures (Intel ibt-*.sfi, MediaTek BT_RAM_CODE, Realtek rtl_bt, Broadcom .hcd, Qualcomm QCA), returning the named blob via a bounded =cmd.run(journalctl -k)= that reuses the =doctor.py:84= precedent; =boot_enablement()= reads three boot-persistence signals (bluez AutoEnable from main.conf [Policy], =systemctl is-enabled bluetooth=, whether TLP lists bluetooth in =DEVICES_TO_DISABLE_ON_STARTUP=). =diagnose()= gates the firmware read to the no-adapter branch and the boot read to the soft-blocked/powered-off branch, so a healthy run reads neither; the raw signals ride a new =probes= key that =doctor()= carries into =--json=. Detection only: no verdict, formatter, or repair change (that's Phase 1). Every read degrades to None on an unreadable tool/file, so a probe that can't see never invents a fault. AutoEnable absent/unset reads None, not false, matching bluez's compiled default of true, so only an explicit =AutoEnable=false= is the fault. New env roots for tests (=BT_MAIN_CONF=, =BT_TLP_CONF=, defaulting to absent temp paths in the Sandbox base so no test reads real /etc); =fake-journalctl= branches on =-k=, =fake-systemctl= answers =is-enabled bluetooth=. 125 bt tests, full =make test= green; =/review-code= approved (no Critical/Important; one Minor noting the firmware read also covers the btctl-unavailable branch, harmless). Inbox note sent to dotfiles. -*** 2026-07-11 Sat @ 03:16:59 -0500 Built the firmware-hint Guide verdict -On dotfiles main (=f05a9b4=, pushed). The no-adapter step now names the blob: a new =_no_adapter_step= consults =probes.firmware_hint()= on a genuine no-adapter fault and, on a per-vendor signature match, sets =evidence= to "no Bluetooth adapter found — <Vendor> firmware <blob> failed to load" and =next_action= to "update linux-firmware and reboot", tagged with a new =code="no-adapter-firmware"= so a =--json= consumer can branch without string-matching. A clean log keeps the generic hardware/driver verdict. A Guide, not a repair: the step carries no =repair= action, so =--fix= never touches it, and it needs no privilege model (so Phase 1 lands independently of the shared cross-panel model). A missing =bluetoothctl= (=BtctlError=) short-circuits before the firmware read, so the verdict fires only on a real no-adapter fault, not a broken install — this also tightened Phase 0 (which read the log on both None branches) to the genuine no-adapter case. =_mk= gained a uniform =code= key (default None) added to every diagnose step, mirroring the existing =repair= key; no test asserts an exact step key-set, verified. =format_doctor_human= already renders =evidence=/=next_action=, so no formatter change. 133 bt tests (+8), full =make test= green; =/review-code= clean. Inbox note sent to dotfiles. -*** 2026-07-11 Sat @ 06:48:21 -0500 Built the persistent-power verdict + fix -On dotfiles main (=d7d859f=, pushed). =_powered_step= consumes the bt Phase 0 boot-enablement probe: AutoEnable explicitly false, service disabled at boot, or TLP listing bluetooth → =powered-off-persistent= (code + evidence naming the cause) carrying the =persist-power= repair; absent/unset config reads as auto-enable-on (bluez default) → the plain =power-on=, so healthy machines can't false-positive. The =persist-power= repair fixes only the causes set (AutoEnable=true, =systemctl enable bluetooth=, drop bluetooth from the TLP list), powers the adapter on now, and verifies each cause cleared. Config edits are pure idempotent text transforms in a new =bootconf= module (comments preserved), staged and installed via a fixed-destination =cp= verb so the write needs root but the mutation is unit-testable. =priv.py= gained three narrow verbs (=enable-bluetooth=, =write-main-conf=, =write-tlp-conf=). The repair is Privileged and resolves through =panelkit= before running: can't-elevate degrades to the guide. The bt shim gained =panelkit= on its path. 287 bt tests + 65 make-test suites green, review-code Approve, voice. Inbox note sent. Live half (real reboot persistence) is the VM/manual checklist. -*** 2026-07-12 Sun @ 09:14:00 -0500 Flipped the bt spec to IMPLEMENTED and logged the vNext items -Spec status heading now IMPLEMENTED (dated history line + Status mirror); all four phase headings DONE. vNext items (stale-bond signature, connection-parameter hints, bt-audio-profile expansion) logged as the "Bt doctor vNext" task. The live half (real reboot persistence) remains on the manual-testing checklist — findings there come back as bugs. -** DONE [#B] One copy + close control pair on every output wall :feature:dotfiles:solo: -CLOSED: [2026-07-12 Sun] -Resolved in dotfiles =dccd744=: every wall carries the o-copy/o-clear overlay pair. bluetooth gained the copy key (transcript via the new =viewmodel.step_copy_line=, CLI-shaped); maint traded its header COPY key for the overlay pair, kept HIDE, and its ✕ clears the session log via =PanelModel.wall_clear=; the four hand-rolled wl-copy calls collapsed into =panelkit.clipboard.copy_text= (PANELKIT_WLCOPY test seam), moving maint off the GTK clipboard so its copies survive the panel closing too. 15 new tests (8 clipboard, 4 step_copy_line, 3 wall_clear); full suite 66 green; all four panel smokes run live off-workspace — maint + audio fully pass, net + bt fail only the pre-existing state-word startup race. - -Converge all four instrument-console output walls on the net panel's well controls: a copy glyph and a ✕ close, as an overlay at the top right, hidden until content lands. Craig's call, 2026-07-10, while reviewing the audio doctor's wall: "network panels as standard across all others, make it consistent." - -Where they stand today, no two alike. net has copy + ✕ (=net/src/net/gui.py=, the =o-copy= / =o-clear= overlay). bluetooth has ✕ but no copy. maint has COPY + HIDE as keys in a header row, and no ✕. audio just gained copy + ✕ (dotfiles =bd33440=). - -Work: give bluetooth a copy key, give maint the overlay pair, and lift the four hand-rolled =_copy_output= implementations into one shared helper rather than a fifth copy. maint keeps HIDE alongside close, because its wall is a persistent session action log you collapse and keep, where net's, bt's, and audio's are per-run results you dismiss. - -Copy text is per panel but one rule: it pastes as that panel's CLI prints, so the paste lines up with the terminal a user is already looking at. audio's =viewmodel.wall_copy_text()= is the worked example. - -Consistent with the 2026-07-07 scope note on the sibling task above: net's compact glyph overlay is what standardizes, not maint's wide COPY-key header row. -** DONE [#C] Org-capture float popup grows too large :bug:hyprland:quick:solo: -CLOSED: [2026-07-14 Tue] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-13 -:END: -Craig answered the pre-flight (2026-07-14): cap at 120 wide, height proportional. Applied as 120 Emacs columns (11 px/col measured from the live daemon) = 1320 px wide, height 653 px from the old rule's aspect. Ratio's size rule shrank from the 1892x936 scratchpad match and both hosts gained a max_size growth cap (the field is max_size — bare "maxsize" is invalid and hyprctl reload won't say so; check hyprctl configerrors). Verified live: config clean, a probe window floats at exactly 1320x653. Dotfiles 9c4dc2f. -** DONE [#C] Panel smoke: faceplate state-word assertion fails on the live compositor :bug:dotfiles:test: -CLOSED: [2026-07-14 Tue] -Diagnosed and fixed within the session: not a race — test drift. Dotfiles b581d5d (2026-07-05) made the faceplate word the static subsystem identity (NETWORKING / BLUETOOTH / AUDIO) and updated the audio smoke, but the net and bt smokes kept asserting the retired live-state words and had failed on every run since. Both now assert the identity word like audio's (dotfiles 32cd99f); both smokes run RESULT: OK end to end, which also green-gates the doctor-streaming change. -** DONE [#C] Realtime lamp output for the net + bt doctors :feature:solo: -CLOSED: [2026-07-14 Tue] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-09 -:END: -Shipped in dotfiles 0318a91. Both doctors stream: diagnose() emits each step as it completes (bt streams the first diagnosis only — the fix loop's re-diagnoses would replay the chain), and a repair's row goes up amber at attempt start and settles green/red with narration + evidence at completion, so the lamp blinks for the repair's real duration. The 3.5-entry height cap turned out to already be in both wells (it landed with the doctor expansions), so only the streaming half needed building. 5 new tests across net + bt; both suites green; AT-SPI smokes at parity with HEAD (one pre-existing state-word failure, filed separately). -Retrofit the net doctor (=~/.dotfiles/net/src/net/doctor.py=) and bluetooth doctor (=~/.dotfiles/bluetooth/src/bt/doctor.py=) to stream results as a live output wall — one lamp per escalation step, amber while running, green on success, red on failure — instead of a final summary. Matches the maintenance-console doctor design (see [[file:docs/design/maintenance-console-design-ideas.org][maintenance-console-design-ideas.org]], "Doctor = live output wall"). Goal: every doctor in the system reads the same way. Both doctors already step through an escalation chain re-probing after each, so the steps are natural lamp boundaries. - -Scope note (Craig, 2026-07-07): realtime lamp *behavior* only. The maintenance console's wider results-wall layout (date+time stamp column, COPY, persistent history) does NOT backport — the net/bt panels are ~400px wide and lack the horizontal real estate. Their existing output wells keep their compact layout; this task just makes them stream live. - -Addendum (Craig, 2026-07-07): DO backport the 3.5-entry height convention — every panel's output well caps at 3.5 visible entries, the half-visible entry being the scroll cue, with the dark slate-on-black scrollbar. Layout stays compact per above; only the height cap + scroll affordance carries over. -** DONE [#B] Absorb the clock-panel project into the dotfiles :feature:waybar:dotfiles: -CLOSED: [2026-07-18 Sat] -Absorbed into =~/.dotfiles= (commit 3fab11d): package =clock/src/clock/= (renamed from clock_panel), the six PNG watchface layers packaged inside the module at =clock/src/clock/assets/=, a stowed =clock-panel= shell shim (LD_PRELOADs gtk4-layer-shell), waybar left-click now =clock-panel toggle= with the absolute path dropped, tests converted pytest→unittest into =tests/clock/= plus an asset-load guard. Kept the layer-shell overlay and the socket toggle. The standalone repo is archived (ARCHIVED.md), kept for its design history. Verified live: the bar click renders the polished watchface. -** DONE [#A] Velox boot recovery — no kernel in BE :bug:velox:zfs: -CLOSED: [2026-07-19 Sun] -Recovered. Velox boots linux-lts 6.18.38 and is back on the tailnet (up 1d+, /boot holds initramfs-linux-lts.img). The pre-pacman ZFS snapshot rollback restored the kernel from the ZBM recovery shell. -Velox won't boot: ZBM prompts for the passphrase, unlocks, then reports no bootable environment with a kernel. Cause: an interrupted kernel =-Syu= removed the old kernel and never installed the new one — /mnt/be/boot (from zroot/ROOT/default) holds ONLY intel-ucode.img; vmlinuz-linux + both initramfs are gone. /boot lives inside zroot/ROOT/default (no separate boot dataset), so root-dataset snapshots capture it. - -Status 2026-07-15: a first rollback attempt did NOT fix it (square zero after reboot) — suspected typo in the snapshot name, so the rollback likely errored and did nothing. NOT verified. Next session: verify state in the ZBM recovery shell BEFORE any reboot. - -Recovery lever: the pre-pacman ZFS snapshot hook (live on velox since 2026-06-29) snapshots zroot/ROOT/default@pre-pacman_<ts> before every pacman transaction. The newest =pre-pacman_<ts>= predating the failed upgrade holds the intact old kernel — roll back to it. - -Morning steps (Craig at velox ZBM → recovery shell, Ctrl+R): -#+begin_src sh -# 1. pool writable + key loaded -zpool get readonly zroot -zfs get -H -o value keystatus zroot/ROOT/default -# if readonly=on: zpool export zroot && zpool import -f -N zroot -# if keystatus=unavailable: zfs load-key zroot - -# 2. list snapshots — COPY THE EXACT NAME (the typo bit here last time) -zfs list -t snapshot -o name,creation zroot/ROOT/default | grep pre-pacman - -# 3. see current /boot state (read-only mount) -umount /mnt/be 2>/dev/null; mkdir -p /mnt/be -mount -t zfs -o zfsutil,ro zroot/ROOT/default /mnt/be -ls -la /mnt/be/boot - -# 4. if /boot still shows only intel-ucode.img: redo rollback with the exact name -umount /mnt/be 2>/dev/null -zfs rollback -r zroot/ROOT/default@pre-pacman_<EXACT-TS> # -r, NOT -R - -# 5. VERIFY before reboot — remount RO, confirm the kernel is back -mount -t zfs -o zfsutil,ro zroot/ROOT/default /mnt/be -ls -la /mnt/be/boot # MUST show vmlinuz-linux + initramfs-linux.img -umount /mnt/be - -# 6. only once /boot shows a kernel: -zpool export zroot && reboot -#+end_src -Scope: only zroot/ROOT/default reverts; /home, /var, /media are separate datasets, untouched. After boot: =pacman -Syu= attended, confirm /boot holds vmlinuz-linux + initramfs before any shutdown. Full diagnosis: =inbox/PROCESSED-2026-07-15-0002-from-.emacs.d-velox-boot-failure-handoff.org=; ZBM photo: =inbox/PROCESSED-2026-07-15-0002-from-.emacs.d-PXL_20260715_043758976.jpg= (local on ratio; inbox is gitignored). -** DONE [#C] Restore date-format scrolling on the waybar date module :feature:waybar:dotfiles:quick: -CLOSED: [2026-07-19 Sun] -Shipped dotfiles 9dfe082: date-only ring (ordinal/full/longdate), on-scroll rewired, layout guard flipped. UTC/time stay on the time module. -Date and time are separate fixed-position controls. The time display cycles its -own formats, including UTC; the date/calendar control cycles date-only formats -and never displays a second time. Implement the dedicated format rings, -tooltip behavior, and tests together in the dotfiles Waybar configuration. -Reference material for the compact clock/chronograph treatment is filed in -[[file:working/clock-display-references/][working/clock-display-references/]]. - -*** 2026-07-19 Sun @ 04:36:26 -0500 Folded clock-panel interaction direction -The clock-panel handoff settled the prior open question: UTC belongs only to -the time ring, while the date ring is date-only. The existing task is therefore -a focused follow-up, not a two-line restoration of the old combined ring. -** DONE [#C] Notification sound loudness :chore:audio:quick:solo: -CLOSED: [2026-07-19 Sun] -Shipped dotfiles 808ca23: NOTIFY_VOLUME default 65536->39322 (0.6 gain) in both notify copies. -Reduce notification-sound playback loudness by 40% (0.6 gain, approximately --4.4 dB). Change the =NOTIFY_VOLUME= playback control rather than re-encoding -the normalized sound files; verify each notification type still plays clearly. -** DONE [#C] Show the active wired interface in the Waybar network module :feature:waybar:network: -CLOSED: [2026-07-19 Sun] -Shipped dotfiles 22867f9: select_device prefers connected wifi -> connected ethernet -> wifi fallback, so a live cable shows the wired glyph+iface instead of Offline. -When Ethernet is active, replace the offline-WiFi presentation with the wired -interface glyph and interface name. -** DONE [#C] Let the clock panel dismiss itself on right click :feature:clock:waybar: -CLOSED: [2026-07-19 Sun] -Shipped dotfiles fc9a2b7: secondary-button gesture -> ClockApplication._dismiss hides the open panel. Live-verified with Craig 2026-07-19. -Make a right click inside the open clock panel toggle it closed. Preserve left -click for its established interaction; the Waybar time module remains the -explicit way to reopen the panel. -** DONE [#C] Make the WiFi toggle connect the best available profile :feature:network: -CLOSED: [2026-07-19 Sun] -Shipped dotfiles 9105361: manage.wifi_radio -> _connect_best_saved activates the strongest in-range saved profile on enable; nothing in range falls back to NM autoconnect. -When enabling WiFi, automatically connect to the highest-priority available -saved network instead of requiring a panel selection first. -** DONE [#A] Tracked WireGuard private keys in repo — public leak, resolved :bug:security:network: -CLOSED: [2026-07-20 Mon] -Confirmed a live public leak, not just at-risk: git.cjennings.net runs cgit (scan-path=/var/git), so archsetup.git was anonymously cloneable over https. An unauthenticated clone pulled the configs with intact PrivateKeys. Exposed 2026-07-05 (c7b7d16) to 2026-07-20. Regraded to P1/[#A] (public credential exposure, severity-alone carve-out) from the initial [#B]. -Scope was wider than first found: the current 3 configs (assets/wireguard-config/wg-*.conf) plus 7 older ones at the pre-reorg path assets/wireguard/ (switzerland x2, USCALA/USCASF/USDC/USGAAT/USNY) — 10 config files, all with real keys. -Resolution: Craig expired all the Proton WireGuard configs (keys dead). Purged all 10 from every commit with git filter-repo, force-pushed main + v0.5, and ran git gc --prune=now on the server bare repo. Verified via anonymous clone: zero real-key blobs reachable, all old exposed commits gone. Stopped tracking plaintext (gitignore + README, out-of-band configs only). -Follow-ups filed below: harden cgit exposure; installer no longer ships configs. -** DONE [#C] Installer chpasswd unguarded — unloggable primary user :bug:solo:quick: -CLOSED: [2026-07-20 Mon] -Fixed (fa3135a): extracted set_user_password, which guards the chpasswd with error_fatal so a failure aborts loudly instead of silently leaving no password. Fake-chpasswd test pins the guard fires on failure and stays quiet on success. -Grading: Major severity (fresh system's primary user can't log in) x rare edge case (chpasswd seldom fails) = P3 = [#C]. -archsetup:1168 runs =echo "$user:$pass" | chpasswd= with no guard, then unsets the password next line; set -e is off (line 21), so a silent failure leaves no password and no log entry. Fix: guard with error_fatal (report + "set it by hand: passwd $user") before unsetting. See findings doc (S2). -** DONE [#C] Installer nvme early module never built into initramfs :bug:solo: -CLOSED: [2026-07-20 Mon] -Fixed in e0d22bd: extracted ensure_nvme_early_module, which rebuilds the initramfs whenever it changed the conf (regardless of ZFS root) and scopes the presence check to the MODULES line. TDD via tests/installer-steps/test_ensure_nvme_early_module.py. -Grading: Minor severity (module autoload still boots the system) x most-machines (all Craig's ZFS-root boxes) = P3 = [#C]. -archsetup:2910 writes MODULES=(nvme) but the only mkinitcpio -P in boot_ux runs =if ! is_zfs_root=, so on ZFS-root non-Framework machines the early-load hardening is never compiled in. Also archsetup:2918 greps the whole file for "nvme" (not the MODULES line). Fix: rebuild initramfs after the MODULES edit regardless of ZFS; scope the presence grep to =^MODULES=(=. See findings doc (S3). -** DONE [#C] Installer disk-space pre-flight check is fragile :bug:solo:quick: -CLOSED: [2026-07-20 Mon] -Fixed in aef074f: extracted check_disk_space using df -P (wrap-safe) and a KB comparison (no truncation bias); non-numeric df output falls back to zero so a malformed read aborts loudly. TDD via tests/installer-steps/test_check_disk_space.py. -Grading: Major severity (aborts a valid install) x some (df wraps long device names on a live ISO / device-mapper root) = P3 = [#C]. -archsetup:487 parses =df / | awk 'NR==2'=, which reads the device-name line (empty $4 -> 0 GB) when df wraps; archsetup:488 also integer-truncates the GB compare against the 20 GB floor. Fix: =df -P /= (single-line) or =df --output=avail=; compare in KB to avoid the rounding bias. See findings doc (S1). -** DONE [#C] Installer run_step state + exit-code handling :bug:solo: -CLOSED: [2026-07-20 Mon] -Fixed in 6de55d2: run_step records the state marker whenever the step function returns (a return past error_fatal's exit means only a non-fatal warning is left), added local to run_step/show_status, and captured pacman's real exit in the refresh loop. TDD via tests/installer-steps/test_run_step.py. -Grading: Major severity (resume re-runs steps and can abort on a survivable warning) x some (a step whose last action is a non-fatal failure) = P3 = [#C]. -archsetup:298 marks a step complete only when its function returns 0, but error_warn/run_task return 1, so a non-fatal-failing step never writes its marker and re-runs on resume. Also archsetup:1034 reports =$?= of the =false= test, not pacman's real exit code; and run_step locals (290/318) leak to global scope. Fix: step functions =return 0= explicitly (or gate run_step on a per-step error flag); capture the real exit code; add =local=. See findings doc (S1). -** DONE [#C] cmail password decrypted world-readable before chmod :bug:security:solo:quick:cmail: -CLOSED: [2026-07-20 Mon] -Already fixed in dffecf5 (before this session): decrypt_to_secure wraps the gpg decrypt in a 0077-umask subshell so the file is 0600 from creation, with tests/cmail/ verifying the umask at write time. The task was stale; verified green and closed. -Grading: security carve-out — brief local plaintext exposure of the mail password, requires a concurrent local shell during install; narrow window = low severity = P3 = [#C]. -scripts/cmail-setup-finish.sh:52 gpg-decrypts to ~/.config/.cmailpass at the process umask (often 0644), then chmod 600 on the next line. Fix: =(umask 077; gpg ... --output ...)= or decrypt to a mktemp 0600 file and mv into place (mirror the import-wireguard mktemp -d 0700 pattern). See findings doc (S4). -** DONE [#C] Installer sudoers.pacnew blind copy risks lockout :bug:solo:quick: -CLOSED: [2026-07-20 Mon] -Fixed in c80e855: extracted replace_sudoers_pacnew, which runs visudo -cf on the pacnew and only copies a validated file (warns and keeps the working sudoers otherwise). TDD via tests/installer-steps/test_replace_sudoers_pacnew.py. -Grading: Major severity (a malformed sudoers locks out privilege escalation) x rare edge case = P3 = [#C]. -archsetup:1146 does =[ -f /etc/sudoers.pacnew ] && cp /etc/sudoers.pacnew /etc/sudoers= with no validation, right before the NOPASSWD rule at 1183. Fix: =visudo -cf /etc/sudoers.pacnew && cp ... || error_warn=. See findings doc (S2). -** DONE [#C] WireGuard import leaves full-tunnel VPN live on failure :bug:solo:network: -CLOSED: [2026-07-20 Mon] -Fixed in 36daf76: the down now runs before the rename modify (targets the stable UUID), so a failed modify under set -e can't leave a live full-tunnel VPN. Added a connection-down case to fake-nmcli and two ordering tests. -Grading: Major severity (all traffic silently routed through Proton until manual cleanup) x rare (nmcli modify failure) = P3 = [#C]. -scripts/import-wireguard-configs.sh:51-62 imports (which brings the 0.0.0.0/0 tunnel up), renames, then deactivates; under set -e a failed modify aborts before the down, leaving the tunnel live. Fix: bring the connection down right after parsing the UUID, before the rename. See findings doc (S4). -** DONE [#C] net-scenarios diagnose failure exits green :bug:test:solo: -CLOSED: [2026-07-20 Mon] -Fixed in cf211cd: a diagnose miss sets a per-scenario rc carried to the subshell exit, so the run fails honestly while still running fix + assert. New harness at tests/net-scenarios/ drives the real script with stubbed ssh/rsync/jq. -Grading: Major severity (a net-doctor diagnosis regression is reported as a passing run — false green on a diagnostic tool) x rare edge case (only when a diagnosis regresses and this first-draft harness is relied on) = P3 = [#C]. -scripts/testing/run-net-scenarios.sh:103 — the scenario_diagnose_expect else-branch prints fail "...diagnose did NOT name it" but never forces a non-zero subshell exit, so ( ... ) || fails=... leaves fails unincremented and the script prints "all scenarios passed" + exit 0. Fix: exit 1 in that branch like the other two checks. See findings doc (S5). -** DONE [#C] pacman-hook-order test is a tautology :test:solo:quick: -CLOSED: [2026-07-20 Mon] -Fixed in 1b7236b: the test now extracts the hook filenames the installer writes and compares them against the stock 60-mkinitcpio-remove name (pacman's filename ordering is the real invariant, not source position). Mutation-verified: a 05->70 rename fails the new compare where the old literal compare stayed true. -Grading: Major severity (guards boot-critical hook ordering — a reorder that removes the current initramfs without a rebuild is unbootable, and this test would ship it green) x rare (hook order rarely changes) = P3 = [#C]. -tests/installer-steps/test_pacman_hook_order.py:20 — the two assertLess calls compare string literals ("05..." < "60..."), a constant ASCII fact always true regardless of file content; the ordering the test exists to protect is never measured. Only the assertIn presence checks do real work. Fix: assert on positions — text.index("05-zfs-snapshot.hook") < text.index("60-mkinitcpio-remove.hook") (and the guard hook). See findings doc (S6). -** DONE [#C] Add inetutils to install base :feature:solo:quick:network: -CLOSED: [2026-07-20 Mon] -Already done in 1115543 (earlier today): inetutils sits in install_required_software, with tests/installer-steps/test_required_software.py pinning it (test_installs_inetutils_for_ftp, green). The task was stale; verified and closed. The next full VM run covers the install-path verification. -Original context: TRAMP's /ftp: method needs =/usr/bin/ftp= (GNU inetutils); dirvish has an FTP quick-access entry. Installed manually on ratio 2026-07-14. From .emacs.d handoff 2026-07-14-1751. -** DONE [#D] Installer resume-idempotency cluster :bug:solo: -CLOSED: [2026-07-20 Mon] -Fixed in 8917f2f: extracted crontab_append_once (dedup guard), zfs_scrub_timer_units (one timer per pool, warn on none instead of @.timer), and enable_user_service (wants-symlink; gamemode now uses it and syncthing folds into the shared helper). TDD via tests/installer-steps/test_idempotency_cluster.py. -Grading: Minor severity x rare edge case (re-run after a mid-step failure) = P4 = [#D]. Group of small non-idempotent / wrong-target spots. -crontab log-cleanup line duplicates on resume (archsetup:1713 — guard on absence); zfs scrub timer picks an arbitrary pool via =head -1= and yields =@.timer= when empty (archsetup:1857); gamemode enabled via =systemctl --user= which the script itself documents fails at install time (archsetup:2419 — use the manual wants-symlink like syncthing). See findings doc (S2, S3). -** DONE [#D] Installer unguarded chmod/cp after non-fatal ops :bug:solo:quick: -CLOSED: [2026-07-20 Mon] -Fixed in dd41036: extracted install_executable (guarded cp + chmod +x) for the two zfs scripts; guarded the two hypr-live-update-guard chmods inline with error_warn. TDD via tests/installer-steps/test_install_executable.py. -Grading: Minor severity x rare edge case (only when a preceding non-fatal cp/clone failed) = P4 = [#D]. -With set -e off, unguarded chmod/cp hit missing/partial files silently: hypr-live-update-guard chmods (archsetup:2108/2144), zfs-replicate cp (archsetup:1820) leaving a service with a dead ExecStart, zfs-pre-snapshot cp (archsetup:1943) leaving a broken pacman hook. Fix: wrap each in =(...) >> log 2>&1 || error_warn=. See findings doc (S2, S3). -** DONE [#D] normalize-notify-sounds temp/atomicity can corrupt tracked file :bug:solo:quick: -CLOSED: [2026-07-20 Mon] -Fixed in a29769e: resolves the real target via readlink -f, stages the temp beside it, guards on a non-empty encode, and atomically mv's into place (preserving the stow symlink); an EXIT trap cleans a leaked temp. TDD via tests/normalize-notify/ with fake ffmpeg. -Grading: Minor severity (corrupts a repo-tracked sound file, recoverable via git) x rare (ffmpeg failure/interrupt) = P4 = [#D]. -scripts/normalize-notify-sounds.sh:39-46 has no EXIT trap on the mktemp and does =cat "$tmp" > "$f"= (truncate-first) where $f is a stow symlink into the repo; a zero-byte/failed encode writes a corrupt file. Fix: EXIT trap; =[ -s "$tmp" ]= guard; write $f.tmp and overwrite on success. See findings doc (S4). -** DONE [#D] VM test-framework robustness cluster :bug:test:solo: -CLOSED: [2026-07-20 Mon] -Fixed in 866d327: profile-suffixed PID/monitor/serial paths, kill_qemu reaps-or-polls to death before the snapshot restore, debug-vm uses DISK_PATH, and both runners report an honest ARCHSETUP_COMPLETED marker instead of a fake exit code. TDD via tests/vm-framework/test_vm_utils.py (suffix red->green; kill_qemu as a contract pin). -Grading: Minor severity x rare edge case (each fires only in a narrow test-harness path) = P4 = [#D]. Group of four small framework bugs from the S5 audit. -scripts/testing/debug-vm.sh:49 hardcodes the btrfs base disk, ignoring the profile-correct DISK_PATH from init_vm_paths (FS_PROFILE=zfs boots the wrong base or fatals); lib/vm-utils.sh:284 kill_qemu -9's and deletes the PID file without waiting, so a force-kill restore races the dying qemu's qcow2 lock and silently leaves the base image dirty (fix: wait for the PID); lib/vm-utils.sh:69 leaves PID_FILE/MONITOR_SOCK/SERIAL_LOG un-suffixed so parallel btrfs+zfs runs collide (fix: suffix by FS_PROFILE like DISK_PATH); run-test.sh:287 (and run-test-baremetal.sh:234) reports a completion-marker grep as ARCHSETUP_EXIT_CODE, not the installer's real exit — misleading since the installer runs set -e off and can error then still write the marker (fix: rename + capture the true status). Testinfra remains the real pass/fail backstop. See findings doc (S5). -** DONE [#D] Gallery-widget prototype elisp bugs :bug:design:solo:quick: -CLOSED: [2026-07-20 Mon] -Fixed in 552736e: shared clamp feeds needle + readout (150 renders 100%), explicit cl-lib require, and gallery-widget--source-dir with a default-directory fallback. TDD: 3 new ERT tests (clamp red->green; the other two land as pins since svg.el transitively loads cl-lib). -Grading: Minor severity x rare edge case (out-of-range input / cold byte-compile / interactive re-eval) = P4 = [#D]. Prototype code, all three Minor. -docs/prototypes/gallery-widget.el:139 renders the readout from the unclamped value while the needle clamps 0-100, so at value 150 the needle pins at +60 degrees but the text reads "150%" (fix: clamp once, format both from it); :69 calls cl-loop without (require 'cl-lib) — works only via the autoload cookie, bites on a cold byte-compile (fix: add the require); :29 computes its dir from (or load-file-name buffer-file-name), both nil on interactive re-eval outside a load/file buffer (fix: fall back to default-directory). See findings doc (S7). -** DONE [#D] Audit test-quality cluster (Python + elisp) :test:solo: -CLOSED: [2026-07-20 Mon] -Fixed in 179fbd5 (plus 552736e for the gauge-level clamp test): socket check via find -type s, gen_tokens degenerate case pinned exactly as characterization, tick count as direct occurrences, and write-svg covered. All five items dispositioned. -Grading: no runtime behavior change; test-suite quality. Group of five weak/missing tests from the S6/S7 audit. -scripts/testing/tests/test_desktop.py:96 passes a shell glob to `test -S`, which breaks on zero or multiple sockets (masked today because the test always skips); tests/gallery-tokens/test_gen_tokens.py:181 asserts properties too weak to notice the marker output is garbled (impossible input, so low); tests/gallery-widgets/test-gallery-widget.el:77 counts ticks via split-string + cl-count-if :start 1 (a coincidence of split semantics, not a match count); :47 tests the needle-angle helper's clamp but never the rendered readout at an out-of-range value (exactly why the S7 readout/needle bug ships green — add a gauge-level boundary case); :159 leaves gallery-widget-write-svg uncovered (add a Normal write-to-temp case). See findings doc (S6, S7). -** DONE [#B] Installer GRUB_CMDLINE overwrite drops boot params :bug:solo: -CLOSED: [2026-07-21 Tue] -Fixed in f9da097: update_grub_cmdline merges the current value with archsetup's tokens (existing tokens survive, same-key conflicts resolve to archsetup's value) behind a refuse-to-write safety check, via awk + mv with a backup_system_file first. TDD via tests/installer-steps/test_grub_cmdline.py (8 cases incl. cryptdevice/resume/zfs survival and idempotence). -Grading: Critical severity (unbootable) x some-users-sometimes (machines whose base install set a cryptdevice=/resume=/zfs= cmdline param) = P2 = [#B]. -archsetup:3054 rewrites the whole GRUB_CMDLINE_LINUX_DEFAULT line with a fixed string; nothing re-adds a pre-existing cryptdevice/resume/zfs token, so grub-mkconfig (3059) can bake an unbootable config. Fix: read the current value and append only the missing tokens; assert any pre-existing boot-critical token survives before grub-mkconfig. See [[file:docs/design/2026-07-19-sentry-code-findings.org][sentry code findings]] (S3). -** DONE [#C] Maint status wall copy buttons :feature:maint:dotfiles: -CLOSED: [2026-07-21 Tue] -Shipped in dotfiles 8bc79ba per Craig's calls (one global button, rendered text): COPY on the doctor row serializes every category band via the same card_spec the GUI renders, through panelkit clipboard. TDD tests/maint/test_status_copy.py, full dotfiles make test green, inbox note sent. Live check pending: open the maint panel, press COPY, paste. -Craig's roam capture 2026-07-20, routed via .emacs.d sentry inbox-zero as archsetup-owned UI work. Dotfiles maint panel work; archsetup drives it end-to-end per the standing rule. -** DONE [#B] Build: desktop-settings panel :feature:hyprland:dotfiles: -CLOSED: [2026-07-22 Wed] -:PROPERTIES: -:SPEC_ID: d6bb1e73-ec90-4327-85ee-bfa762da5bce -:END: -The GTK build of the desktop-settings panel per the spec (docs/specs/2026-07-02-desktop-settings-panel-spec.org, DOING; normative reference: prototype 37). Work happens in dotfiles settings/ — archsetup drives the lifecycle. Two non-blocking build-time picks live in the spec's Review findings (wallpaper setter tool; store location/format) — decide in phase 1 and record there. -*** 2026-07-22 Wed @ 13:14:01 -0500 Built the backings engine (phase 1) — dotfiles 7a15237 -Landed as dotfiles settings/src/settings (10 modules) + tests/settings (118 tests against fake binaries, auto-discovered by make test — 81 suites green). Covers brightness/kbd (5% floor, x10 drum), toggles (dim, pointer cycle via toggle-touchpad, caffeine), DND class-split (dunst pause level 60, close-all before unpause, alarms punch through live), powerprofilesctl, nightlight (resident gammastep), hypridle.conf renderer + symlink-safe write + caffeine-respecting reload + hyprlock grace, suntimes (pure NOAA math), and the wallpaper engine (awww/mpvpaper/projector adapters, galleries, random draw, atomic JSON store). All three build-time picks recorded as DONE findings in the spec (setter=awww, store=state.json, nightlight=gammastep). Handoff note in ~/.dotfiles/inbox/. -*** 2026-07-22 Wed @ 15:26:44 -0500 Built the presenters (phase 2) — dotfiles 5172289 -Three GTK-free models per prototype 37, all at 100% line coverage (tests/settings/test_presenters.py, 100 tests; full repo suite green before and after). programs.py: the matrix — eight complete programs (Craig's four factory scenes drafted here per the pre-flight pick, slots 1-4 first-class), pin rows + power radio row, activate returns the full sets, member writes return apply/updated with active-is-live surviving. bench.py: drum mapping (screen never reads 0, floor 5%; kbd floors at 0), idle rail order clamping between enabled neighbors, park/unpark with re-clamp, caffeine bypass, view-state builder tolerant of no-backlight None. channels.py: the eight-channel bank, per-mode sources visibility, alpha/recency sort (unlabeled last), the shared mint/edit/delete grammar for pairs/sets/colors (press arm-cycle, two-picture set minimum, dup rejection, selection clamping), sources guardrails, interval wheel, previews. Handoff note in ~/.dotfiles/inbox/. -*** 2026-07-22 Wed @ 16:03:52 -0500 Ported prototype 37's instruments to GTK (phase 3) — dotfiles 33d82eb -The panel renders P37 end to end. New instruments.py carries the three Cairo instruments as clock-free humble objects: ProgramMatrix (glyph/numbered heads over jewel pins + CPU POWER paper letter wheels), DrumRoller (paper drums, drag-to-set, dimmed n/a on no-backlight machines), TripDial (sqrt 300° scale, colored stage tabs, OFF-notch parking, exact-minutes drag counter, BYPASSED · CAFFEINE stamp, bottom legend). gui.py rebuilt to P37's layout with the wallpaper sub-view: channel bank with drawn faces, minted pair/color/set trays (alpha/time sort, edit/delete chip feet), the three presses (pair arm-cycle, color picker, set press + interval wheel), sources with a folder picker. New GTK-free glue all unit-tested (test_panel_glue.py, 33 tests): dial geometry in bench, matrix/idle/wallpaper wiring in panel, presenter-vocabulary channels (pair/solid/random-from-set) in wallpaper.apply. AT-SPI smoke (make test-panel-settings) drives the real wiring against faked backings + a sandboxed store, pinned to its own child pid so it can never fire a live panel's backings. Visually verified on a headless output against P37 captures (main + pair/single/solid/random). Adaptations recorded in the handoff: five-stage dial (WATCH gets its own green — the engine runs watch separately, P37 merged the label), DESKTOP_SETTINGS_START_VIEW test seam. Full suite 84 suites green; window rule widened for the 540px panel. Handoff note in ~/.dotfiles/inbox/. -*** 2026-07-22 Wed @ 16:47:54 -0500 Integrated phase 4 — dotfiles 680b50d -Bar consolidation had landed early (74f723e); this pass shipped the rest. settings-project hosts the watch/clock/world channels as HTML faces (settings/faces/) on a gtk-layer-shell background window over WebKit2 — all three visually verified on a headless output, world reading the waybar worldclock roster via query param. settings-watch is the hypridle watch-stage host: throwaway-profile chrome kiosk that reveals only after its window maps behind the lock and relocks before teardown — a failed face degrades to the plain lock, never a bare desktop (unlocked lifecycle verified live; the locked swap goes to the manual checklist). Sun-pair location reads whereami live per transition with last-good cache in state.json (verified live: 9.5s first beat, New Orleans coords, Gogh day side applied); desktop-settings-tick.timer (2 min, enabled on ratio, added to the installer) drives flips and random draws — 23ms no-op beats. dunstrc history_length 100 protects held alarms (full DND cycle verified against live dunst; wtimer alarms already CRITICAL via the notify wrapper, no promotion rule needed). Live hypridle rewrite verified — five-stage regime rendered through the stow symlink, caffeine respected (found engaged, daemon correctly left stopped). Refresh signals needed no rewiring (touchpad signals itself via toggle-touchpad). 45 new tests; suite 84 suites green; smoke 13/13. Handoff note in ~/.dotfiles/inbox/. -Velox one-time steps (sync doesn't carry): mpvpaper (AUR), optionally power-profiles-daemon (service off), and systemctl --user enable --now desktop-settings-tick.timer. -*** 2026-07-22 Wed @ 17:05:58 -0500 Landed the 17-point end-to-end pass — dotfiles 9038eee -Prototype 37's 17-point suite re-derived against the real panel (the original Playwright script wasn't preserved; the functional surface in the spec's Final prototype section is the source). tests/settings/panel_e2e.py + run-panel-e2e.sh + =make test-panel-e2e=: points 1-14 drive the running panel over AT-SPI (program recall with per-backing verification across FOCUS/BATTERY/slot1, pointer console keys, all eight wallpaper channels including projected watch/world stop/start ordering, close); points 15-17 cover the Cairo instruments (drums, tripper dial clamp/park/render/reload, matrix pins + letter wheels with active-is-live) at the backing layer, since AT-SPI can't reach a DrawingArea's hit-tests. Same safety posture as the smoke: sandboxed store, faked backings, pid-pinned a11y node. 17/17 green on ratio's live compositor; full suite 85 green; smoke 13/13; ruff clean. The drag gestures go to the manual checklist below. Handoff note in ~/.dotfiles/inbox/. -*** 2026-07-22 Wed @ 17:05:58 -0500 Flipped the spec to IMPLEMENTED -docs/specs/2026-07-02-desktop-settings-panel-spec.org DOING → IMPLEMENTED with a dated history line naming the shipping commits (dotfiles 7a15237 / 74f723e / 5172289 / 33d82eb / 680b50d / 9038eee) and the verification evidence (85 suites, smoke 13/13, e2e 17/17). The four panel drag-gesture checks and the locked-path night-watch swap live under "Manual testing and validation" — human-eye checks, not implementation blockers. -** CANCELLED [#B] Hyprland layoutmsg crash — bad_variant_access (upstream) :bug:hyprland: -CLOSED: [2026-07-21 Tue] -Dropped 2026-07-21 (Craig's call) — not tracking the upstream report. The crash evidence (both reports + tmpfs log excerpts) and the voice-passed issue draft stay preserved in [[file:working/hyprland-layoutmsg-crash/][working/hyprland-layoutmsg-crash/]] if it recurs and is worth reviving. -Grading: Critical severity (SIGSEGV kills the whole desktop session; every GUI app's unsaved state lost) x rare edge case (twice in ~4.5 months: 2026-03-07 on v0.54.1, 2026-07-20 on v0.55.4) = P2 = [#B]. Upstream Hyprland bug, not this repo's code — the task tracks reporting it and picking up the fix. -A layoutmsg mfact dispatch (layout-resize, mod+H/L) throws std::bad_variant_access inside Layout::CAlgorithm::layoutMsg, uncaught, SIGSEGV. Both crashes fired from the layout-resize mfact path (keycode 104 shrink today, 108 grow in March). Layout at crash was master and the identical mfact had worked seconds earlier; the pre-crash window held monocle<->master toggles, two window closes dropping focus to "[Window nullptr]", and togglefloating x2. Monocle is a registered v0.55 layout (log shows graceful "Unknown monocle layoutmsg" rejects), so the config is not at fault; related edges are guarded ("mfact -> no window") while this path misses its variant guard. Repo has no newer build (0.55.4-1 installed and repo). -Evidence preserved in [[file:working/hyprland-layoutmsg-crash/][working/hyprland-layoutmsg-crash/]] (both crash reports + excerpts from the tmpfs session log, extracted before reboot loses it). -Next: Craig posts the issue himself (2026-07-20 decision) — the voice-passed draft is [[file:working/hyprland-layoutmsg-crash/issue-draft.md][issue-draft.md]], with both crash reports and the log excerpts beside it for attaching. Watch the repo for a fixed release and close on confirmation. The layout-resize script guard was declined (a script can't observe the internal desync). -** DONE [#C] WireGuard import is now config-less — decide feature fate :feature:network: -CLOSED: [2026-07-21 Tue] -Decided 2026-07-21 (Craig): KEEP the import feature. The out-of-band flow is already in place — =assets/wireguard-config/= carries a README documenting "drop plaintext =*.conf= locally at install time (gitignored); ship encrypted =*.conf.gpg= to track", and its =.gitignore= enforces it (=*.conf= blocked, =!*.conf.gpg= allowed). The script also already no-ops gracefully on an empty dir (=shopt -s nullglob= + a =found= flag), so nothing ships and nothing errors when no configs are present. Nothing to build; the fate decision was the whole task. -scripts/import-wireguard-configs.sh reads assets/wireguard-config/*.conf, but no configs ship in the repo anymore (removed as a public-leak fix; .gitignore blocks plaintext). -** DONE [#C] Dupre theme waybar.css drifted from live style.css :bug:dotfiles:waybar: -CLOSED: [2026-07-21 Tue] -Fixed in dotfiles 3e4e7ff (2026-07-20, "fix(theme): sync dupre waybar.css with the live weather rules") — dupre/waybar.css is byte-identical to live again, restoring the =#custom-weather= selectors/hover/gold divider, so =tests/theme-css= is green. -Grading: Minor severity (cosmetic, reverts only on a theme switch) × rare edge case (dupre is already the active theme) = P4 = [#D] on user impact, bumped to [#C] because the dotfiles =make test= stays RED until synced, poisoning the green baseline for every future commit. -The weather-kit work added =#custom-weather= selectors to =hyprland/.config/waybar/style.css= but never mirrored them into =hyprland/.config/themes/dupre/waybar.css=. =tests/theme-css= asserts the two files are identical (set-theme copies the theme file over the live one), so switching to dupre would silently revert the weather chip styling. Fix: sync the theme file to live. Pre-existing; found 2026-07-19 during an unrelated commit's green-baseline run. -** DONE [#A] Hyprlock lockout: AMD-iGPU DPMS invalidates the lock, session wedges :bug:hyprland:installer:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as archsetup =a9391c9= + dotfiles =3046c9c=, both pushed; applied live to ratio and velox. Reboot ratio to activate the root fix (=amdgpu.runpm=0=); the watchdog covers until then. - -WHAT HAPPENED. Ratio's screen idle-locked, then wedged: hyprlock gone, the compositor still holding the ext-session-lock, no password prompt, recoverable only from a console. Recovered live with =hyprctl dispatch exec hyprlock= (=allow_session_lock_restore=true= was already set, so a replacement client adopted the dead lock). - -ROOT CAUSE (evidence, not the first guess). My first read was "hyprlock crashed on its screenshot buffer" — WRONG. Coredumps are captured here (two telega SIGSEGVs the same afternoon) and there is NO hyprlock coredump, so it did not segfault; memory was fine, so not OOM. The hyprland log shows the real chain: =Modesetting DP-4= / =Restoring crtc 86= (a display modeset) → =color management protocol is enabled and outputs changed= → =SessionLock.cpp:50 SessionLockSurface object remains but surface is being destroyed=. A display power cycle tore down the lock surface. Online research confirms it's a documented AMD-integrated-Radeon issue (hyprlock#953, Hyprland#5822): the GPU resources the lock client holds become invalid when the display powers down and back up. Ratio is a Strix Halo Radeon 8060S — exactly that hardware, and its cmdline already carried =amdgpu.dcdebugmask=0x10= + =no_vpe_idle_pg=1= display workarounds, a history of the same fragility. - -THE FIX, four layers, research-validated: -1. Root cause: =amdgpu.runpm=0= on the kernel cmdline (AMD only, added in =update_grub_cmdline= behind =detect_gpu_vendors=). Keeps GPU runtime PM from invalidating the resources on a display cycle. Live in ratio's grub.cfg; effective next boot. -2. Separate crash cause: =configure_hyprlock_pam= writes a complete =/etc/pam.d/hyprlock= (auth/account/session). The package default is =auth include login= only, so pam_end() crashes on uninitialised handles. Applied live to both machines. -3. Recovery net: the =screen-lock= watchdog (dotfiles) relaunches hyprlock on a non-zero exit; hypridle's =lock_cmd= routes through it. Independently the same shape as the community's watchdog layer. -4. NOT done, deliberately: the =dpms off= listener stays in the committed hypridle — =runpm=0= makes it safe on AMD, and it's wanted on Intel/velox for idle display-off. Ratio's test rail already removed it as a local choice. - -REVERTED a wrong turn: I'd first built a screenshot-to-file change (grim the desktop, point hyprlock at the file) on the theory the live screencopy buffer crashed. The research showed the cause is GPU runtime PM, not the background source, so I dropped it and reverted hyprlock.conf to =path = screenshot=. - -PROCESS NOTE — I hit the pathspec-commit trap AGAIN (the one the =Two agent sessions sharing one repo= VERIFY documents). After surgically staging only the =lock_cmd= line via =git update-index=, I ran =git commit <path> -m ...=, which commits the WORKING TREE of that path, not the index — so it committed ratio's test rail (dpms-off removed, timeout 450) with a message claiming dpms-off stays. Caught it before push, =git reset --soft=, re-verified. The rule: after =update-index=, commit with =git commit= (no pathspec), never =git commit <path>=. - -Tests: archsetup 372 (test_grub_cmdline AMD-runpm cases + test_hyprlock_pam, both call sites in CALL_SITES); dotfiles 3687 incl. tests/screen-lock. Each guard proven by deletion. - -Grading: Critical severity (full session lockout, console-only recovery) x rare edge case (needs an idle lock plus a display modeset on the AMD iGPU) = P2 = [#B by the matrix]. Raised to [#A] here because it stranded a live machine and the root fix needs a reboot to arm — worth Craig seeing at the top until he reboots ratio. -** DONE [#B] Adversarial review of the sentry run — six fixes reworked :bug:test:tooling:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Craig asked for a skeptical review of every sentry change. Eight agents covered all 23 code commits, each told to disbelieve by default and to answer three questions per commit: does the problem exist and is it reachable, is the fix correct or is there a better one, would each test fail with the fix reverted. Every finding below was re-verified by hand before acting on it. - -SIX COMMITS NEEDED WORK, now fixed: archsetup =1207ca5= (wipedisk), =96e12b5= (firmware trim), =560e1dd= (autologin), =3c2155d= (initramfs tabs); dotfiles =ec7229b= (tunnel import), =a81aa0e= (thumbnail sweep), =56807e5= (three residual guards), =c90ee34= (event-log isolation). Both suites green: archsetup 341, dotfiles 3687 on both gates. - -THE ONE THAT MATTERED MOST. =wipedisk= ran =blkdiscard -f= BEFORE the busy check. =-f= disables the exclusive open util-linux has used since 2.36, so on the exact case the round-11 commit reasoned about — the user picked the wrong disk — it discarded a live filesystem and only then let sgdisk fail, printing "could not clear the partition table ... run this again". Data gone, user told nothing happened. The ordering predates the sentry commit, but round 11 wrote reasoning about the busy-disk case into the comment and error text while leaving the discard first, which made the misreport worse in the one direction that costs something. Dropping =-f= makes the kernel's own O_EXCL the gate. - -THREE PATTERNS WORTH MORE THAN THE INDIVIDUAL FIXES: - -1. CALL SITES WENT UNTESTED IN FIVE SUITES. Every helper had thorough tests; not one proved it was called. Deleting the call left everything green — including the guard on a =pacman -Rdd= of twelve firmware packages, whose removal would have run the trim on ratio. Closed with =CALL_SITES= in =test_orchestrators= (nine pairs, static) and a wiring assertion in the settings suite. Static on purpose: the behavioural harness runs un-stubbed bodies for real, which is fine for an orchestrator and not for a leaf that removes packages. - -2. A NEW OUTCOME VALUE NEEDS EVERY CONSUMER WALKED, EVERY TIME. Done for the portal enum in round 3, skipped for the tunnel-import one in round 4 — where =import_configs= folded a disarm failure into "none imported (N failed)", the opposite of what happened, in the multi-select flow the GUI actually uses. - -3. MY FIXTURES TWICE CLAIMED A FIDELITY THEY DID NOT HAVE. The wipedisk fixture used this machine's real disk names, so five of six tests passed with the seam removed. The mkplaylist fake does a full =cat > /dev/null= drain while its docstring says it "drains stdin exactly when the real one would" — which is what let the wrong failure mode survive. - -AND ONE FINDING WAS DISPROVED OUTRIGHT: round 1's =a57c443= claimed ffmpeg drains the read loop so only the first track is processed. Measured under strace and driven end to end with real ffmpeg (three runs of three, four 120s mp3s), the loop never truncates. The hazard is real and =-nostdin= is right; the symptom was reasoned from shellcheck SC2095 and never run. Corrected in =a30741a=, along with the OpenVPN autoconnect claim and the "four consumers" undercount. - -ALL THREE NOW CLOSED, in dotfiles =c7cb40d= (pushed). =_restore_dot='s =noop= split into =already-on= and =not-managed=, so the step stops claiming a restore that never happened. =_disable_dot= checks its restart as well as its move, since moving the drop-in aside does nothing until resolved reloads. - -The thumbnail one could not be built as described, and that is worth recording. The cache name is a SHA-1 of realpath plus mtime, so no filename says which source it came from; per-source sweeping would mean changing the key format and invalidating every cached thumbnail. Bounding the growth gets the same result for less: a deferred sweep now trims to a 500-file ceiling, oldest first, because eviction is safe exactly where sweeping is not (an evicted thumbnail is rebuilt on the next warm pass, costing one decode and never a file). WHEN A FIX CANNOT BE BUILT AS SPECIFIED, SAY SO AND SOLVE THE ACTUAL HAZARD — the hazard here was unbounded growth, not imprecise attribution. -** DONE [#D] Repair tiers call an unverifiable service restart a failed one :bug:network:bluetooth:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as dotfiles =041d6b9= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). 7 new tests across =tests/bt/test_bt.py= and =tests/net/test_net.py=; dotfiles suite 3665 -> 3672, =make test= exit 0 on both gates. Each of the three guards proven a real gate by deleting it and watching the suite go red. - -Found in the 2026-07-24 sentry bug-hunt, round 14, on the cross-package =repair.py= diff that rounds 5-13 had left unspent. - -=cmd.service_active= is tri-state in both the net and bt packages, and its docstring says so outright: True, False, or None when systemctl itself can't answer (absent binary, or a timeout). Six callers. Three rule on it correctly — =bt/doctor._service_step= branches on None with "systemctl unavailable — can't check the service", and =net/diag= compares =is False= at both its call sites. Three tested it with plain truthiness: - -- =bt/repair.py= =repair_service_restart= -- =net/repair.py= =_service_restart= (the nm-restart and resolved-restart tiers) -- =net/repair.py= =repair_unmask_nm= - -So an unanswerable systemctl was reported as "bluetooth.service is still not active" / "NetworkManager still isn't running after a restart" — a statement about the service made on no evidence at all. Each then pointed the user at =journalctl -u <unit>=, which is the same systemd client stack that had just failed to answer. That last part is round 10's read again: an error message advertising a remedy it cannot honour. - -All three now report =warn= on None, with evidence naming the verification rather than the service, and a next action of checking systemd is reachable and re-running the doctor. Control flow is unchanged: =warn= was already a status both packages emit, both CLIs already exit non-zero on anything but =pass=, and =net/doctor= only inspects a repair step's status for the =dns-test= tier — every consumer was checked before the change, not after. (An adversarial re-review counted twelve, not four; all twelve handle =warn= correctly, so the conclusion held while the claim understated the work.) - -THE SEAM FOR THE TESTS, worth reusing: both suites already carry an exec-failure harness that plants a non-executable file on an emptied PATH, which is exactly what makes =cmd.run= return None. So the None case is reachable through the real code path with no mocking at all. Each test class asserts that premise first (=service_active= really is None in the sandbox) rather than assuming it. - -Grading: Minor severity (the claim is wrong but errs pessimistic — it says a repair failed when it may have worked, rather than falsely reassuring; nothing is damaged) x rare edge case = P4 = [#D]. Fixed rather than filed because the change is three branches and it completes a class — leaving two of three sites collapsed is the failure mode the round-6 =c2eb3e1= commit exists to remember. - -NOT PART OF THIS CLASS, checked and left alone: =settings/toggles.dim_state= is the only other genuine True/False/None helper in the tree, and both its callers pass the value through to the viewmodel rather than collapsing it. Every other "or None" in the packages is two-state (a value or nothing), where falsy handling is correct. -** DONE [#B] Firmware trim gated on a DMI field that never carries the vendor :bug:tooling:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as archsetup =2e228f7= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). New =tests/installer-steps/test_framework_firmware_trim.py=, 12 tests carrying the real DMI strings off both daily drivers. Each of the three conditions proven load-bearing by deleting it and watching the suite go red, and the old gate proven wrong by restoring it (4 failures). - -Found in the 2026-07-24 sentry bug-hunt, round 13, reading archsetup's remaining state-mutating steps. =trim_firmware= gated on =grep -qi "framework" /sys/class/dmi/id/product_name= and no Framework machine has "framework" in =product_name= — it lives in =sys_vendor=. Read live: velox is =Framework= / ="Laptop (13th Gen Intel Core)"=, ratio is =Framework= / ="Desktop (AMD Ryzen AI Max 300 Series)"=. The gate returns false on both, so the step has been a silent no-op on the exact hardware it was written for. velox IS trimmed today (=linux-firmware-{atheros,intel,realtek,whence}= and nothing else) but not by this code path. - -THE REPAIR IS WHERE THE DANGER IS, which is why this is worth reading twice. Swapping =product_name= for =sys_vendor= is the obvious one-word fix and it is wrong: ratio is a Framework Desktop, and =trim_firmware= runs =pacman -Rdd linux-firmware-amdgpu=, which takes the firmware its Ryzen AI Max iGPU needs to bring up a display. Today only the =grep -qi intel /proc/cpuinfo= second gate stands between ratio and that. So =is_framework_intel_laptop= wants three DMI facts — vendor Framework, and a model naming both Laptop and Intel — and the cpuinfo read stays as an independent second gate rather than the only one. - -Verified live after the change: velox TRIM=yes, ratio TRIM=no, where the old gate said no to both. - -Grading: Minor severity (the trim never happens; nothing breaks, the machine just carries ~550MB it was meant to shed) x every user, every time (every Framework Intel install, which is the whole population the step targets) = P2 = [#B]. The AMD-firmware removal is not graded separately because it never shipped — it is the hazard the fix is shaped to avoid. -** DONE [#B] Fresh install leaves the dotfiles repo permanently dirty :bug:tooling:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as archsetup =c3b3617= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). New =tests/installer-steps/test_mark_volatile_configs.py=, 8 tests against a fixture git repo with =sudo= stubbed on PATH. Every guard proven a real gate by deletion. A note went to =~/.dotfiles/inbox/= because =skip-volatile= now has an outside caller. - -Found in the 2026-07-24 sentry bug-hunt, round 13, diffing archsetup's =stow_dotfiles= against the dotfiles Makefile's =stow= target — two implementations of one operation, which is round 5's read applied across repos rather than across packages. - -The Makefile's =stow= target ends with =$(MAKE) skip-volatile=, setting git's skip-worktree bit on the four configs their apps rewrite in place (=btop=, =qalculate=, =calibre=, =waypaper=; the list is =volatile-configs=). archsetup stows inline with raw =stow= calls and never ran that step. So a machine archsetup installed goes dirty the first time one of those apps writes its config, and every later =git pull --ff-only= trips over paths the user never edited. Confirmed by grep: archsetup contains no =skip-volatile=, no =volatile=, and no =make stow= — yet both daily drivers carry the bits, so they came from a hand-run =make stow=, not the installer. ratio in fact carries seven, three more than =volatile-configs= lists, which is evidence the churn is real and ongoing. - -The fix calls the dotfiles target rather than copying its logic, so the volatile list stays single-source. Two details that are load-bearing: it runs *after* =git restore .= so the bit lands on a pristine tree, and it runs as the user, because root writing =.git/index= leaves it root-owned and the user's next git command then cannot update the index at all. A checkout with no Makefile is a quiet no-op — nothing to delegate to is not an error. - -DELIBERATELY NOT DONE: replacing the whole inline stow with =make -C "$dotfiles_dir" stow "$desktop_env"=. The Makefile stows =--target=$(HOME)=, which during an install is root's home, and it carries interactive conflict handling; archsetup stows =--target=/home/$username --adopt= as root on purpose. =skip-volatile= is the one target with no such coupling — it works on the repo through =git -C= and never reads HOME. - -Grading: Minor severity (a repo that reads dirty forever and pulls that need a stash; the workaround is one command) x every user, every time (every fresh install that stows dotfiles) = P2 = [#B]. -** DONE [#C] Unattended install blocks on an interactive prompt :bug:tooling:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as archsetup =cbcb53f= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). New =tests/installer-steps/test_configure_autologin.py= (11) and =tests/installer-steps/test_select_locale.py= (11). Every guard proven a real gate by breaking it and watching the suite go red: dropping the autologin unattended branch fails 1 (on the leftover-stdin assertion, which is the real gate — the drop-in still gets written because the read swallows the sentinel and treats it as "yes"); dropping the locale unattended branch fails 1; breaking either precedence rule fails 2. - -Found in the 2026-07-24 sentry bug-hunt, round 12, continuing through archsetup's own installer. Two members of one class, which is the point: round 10 fixed the third member and left these. - -THE CLASS: an advisory prompt — one that carries its own default — still reading stdin under =--config-file=, the documented unattended mode. Round 10 ruled on it for =nvidia_preflight='s rc-10 prompt. Two sites never got the ruling. - -1. =configure_autologin=. When =enable_autologin= is unset (=AUTOLOGIN= is optional, and =archsetup.conf.example= line 31 ships it commented out) and the root is encrypted, it prompted =Enable automatic console login for $username? [Y/n]= on a bare =read=. It runs from =configure_encrypted_autologin=, inside =boot_ux=, the last entry in =STEPS= — so an unattended install of an encrypted machine works for 40-60 minutes and then sits at a prompt nobody is watching. Under =curl | bash= it is worse: stdin is the script itself, so the read eats a line of source. - -2. =select_locale= (extracted from =preflight_checks= by this commit). The =Choice [1]:= menu fired whenever =/etc/locale.conf= carried no =LANG== and =LOCALE= was unset — also commented out in the example config. archsetup does not require an archangel install, and =configure_build_environment='s own "no LANG=" branch is proof it expects that state. - -Both now take the prompt's own default under =--config-file= and print an =[OK] ... (unattended, --config-file)= line saying so. An explicit =AUTOLOGIN=yes/no= or =LOCALE== still wins; the default only answers a question nobody can. - -WHAT MADE THEM TESTABLE, which is round 10's read (d) applied again: =configure_autologin= hardcoded =/etc/systemd/system/getty@tty1.service.d= and =select_locale= hardcoded =/etc/locale.conf=, so neither could run against a fixture — while their siblings =replace_sudoers_pacnew= and =ensure_nvme_early_module= both take a defaulted path argument for exactly that reason. Both now do. Zero shellcheck delta against HEAD; =make test-unit= 276 -> 298, exit 0. - -Grading: Major severity (unattended installation, a documented feature, does not complete; recoverable by pressing a key, no data loss) x some users, sometimes (needs unattended mode plus an omitted key) = P3 = [#C]. - -THE PROMPTS DELIBERATELY LEFT ALONE, because the class is "prompts with a default", not "all prompts": username (line 636) and password (648/650) have no default to take — there is no sane fallback for either, and =archsetup.conf.example= documents both as "If not set, you will be prompted". They also fire in =preflight_checks=, in the first second of the run, where a blocked prompt is visible rather than silent. The "Enter locale" sub-prompt is reachable only from menu choice 9, which unattended never picks. -** DONE [#C] wipedisk says "Disk erased." when it erased nothing :bug:tooling:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as archsetup =HEAD= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). New =tests/wipedisk/test_wipedisk.py=, 6 tests running the real script against a fixture device directory with fake blkdiscard/sgdisk on PATH. All four guards proven real by deleting each and watching the suite go red. - -Found in the 2026-07-24 sentry bug-hunt, round 11, reading =scripts/= — 30 lines, no tests, and the most destructive script in the repo. Not installed by the installer; it is run by hand from the checkout, which is why the frequency axis stays low. - -Three defects, all of which make the script's final word untrue: - -1. =sgdisk --zap-all= had its result discarded, and "Disk erased." printed unconditionally. sgdisk refuses a busy device — a mounted filesystem or a live md/LVM/ZFS holder — which is exactly what a user hits after picking the wrong disk. So the tool announced an erase it had not performed and exited 0. - -2. "Disk erased." overstates what the tool does even on success. =sgdisk --zap-all= destroys partition tables, not data, and =blkdiscard -f ... || true= deliberately tolerates a device that cannot discard. On a disk without discard support the script cleared the partition table and left every byte readable, while telling the user the disk was erased. That is the one path where the wrong belief has a privacy consequence — someone trusting the message before disposing of a drive. - -3. The prompt says "Select the disk id to use" and then listed every entry in =/dev/disk/by-id=. On this machine that is 18 entries of which 12 are =-partN= partitions (verified by listing it). The menu promised disks and offered partitions. - -Fix: whole disks only (globbed rather than =ls | grep=, so a name with whitespace cannot split into two menu entries); the zap's result is checked and a failure exits 1 naming the busy-device cause; the closing message reports what actually happened, and when discard was unsupported it says the data is still recoverable and points at =nvme format= / =hdparm= for a disposal-grade wipe. - -Grading: Major severity (the tool reports an outcome it did not achieve; in the disposal case that is a data-exposure consequence) × rare edge case (a hand-run helper the installer does not install, and defect 1 additionally needs sgdisk to fail) = P3 = [#C]. - -Worth recording about the tests rather than the code: two of the six passed against the unmodified script for the wrong reason. Without the =WIPEDISK_BY_ID= override the script read the real =/dev/disk/by-id=, so the harness was driving a menu of this machine's actual disks (harmless — the fake blkdiscard/sgdisk shadowed the real ones on PATH — but it was not testing the fixture). And =test_empty_by_id_directory= was not a gate at first: with the guard deleted the empty select menu still falls through to the confirm prompt, reads EOF and declines, so exit code and call log alone pass either way. It now asserts the message. -** DONE [#B] zfs-replicate reports success when every backup failed :bug:backup:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as archsetup =HEAD= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). Diagnostics moved to stderr; the loop counts failures and exits 1 when any dataset failed. New =tests/zfs-replicate/test_zfs_replicate.py=, 9 tests driving the real script with a fake syncoid and a fake ping on PATH (the =tests/zfs-pre-snapshot/fake-zfs= pattern). Both fixes proven real gates by reverting them: dropping the counter fails 3, putting =error()= back on stdout fails 1. - -Found in the 2026-07-24 sentry bug-hunt, round 11, reading =scripts/= — 73 lines with no test file, installed by =configure_zfs_snapshots= as =/usr/local/bin/zfs-replicate= and run by =zfs-replicate.service=, a =Type=oneshot= on a nightly timer. Its exit code and its journal output are the only signals anyone ever sees. - -Two defects, both verified by running the script rather than argued: - -1. The full-replication loop caught each =syncoid= failure, warned, carried on, then printed "Replication complete." and exited 0 regardless. Driven with a fake syncoid failing all four datasets: four =[WARN] Failed= lines, then "Replication complete.", exit code 0. systemd records =Result=success=. A backup that has not run for months is indistinguishable from a working one — and the whole point of the tool is to have a copy when the primary is gone. - -2. =determine_host= runs inside a command substitution (=TRUENAS_HOST=$(determine_host)=) and its =error()= wrote to stdout. On an unreachable TrueNAS the message was captured into =TRUENAS_HOST= and discarded, and =set -e= then killed the script. Driven with both hosts unreachable: exit 1 and completely empty output. A nightly service failing with nothing in the journal to say why. - -Same class as three bugs already fixed this session — =_restore_dot= claiming "DNS-over-TLS restored" without checking, =portal_restore_watch= discarding its outcome, =import_config= returning ok on an unchecked modify. A mutating operation that reports a success it did not get. - -Grading: Critical severity (a backup system that reports success while backing nothing up; the failure surfaces only when the backup is needed — graded on the harm once in the failure state, not on how rarely it is entered) × rare edge case (needs a ZFS root, a reachable TrueNAS, and the user enabling the timer by hand — archsetup deliberately does not enable it, and =findmnt -n -o FSTYPE /= on this machine says btrfs, so it is latent here) = P2 = [#B]. - -Left alone: =BACKUP_PATH="backups" # TODO: Configure actual path= is still an unresolved TODO in the destination, and single-dataset mode relies on =set -e= to propagate a syncoid failure rather than reporting it. Neither is a defect in the sense above; the TODO is Craig's call. -** DONE [#D] Wireless regdom is silently unset for a three-letter-language locale :bug:installer:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as archsetup =249bb93=. =locale_country= matches the =_CC= group instead of counting characters, and =set_wireless_regdom= verifies the substitution landed rather than trusting sed's exit code. 16 tests. -=configure_networking= derives the wireless regulatory domain by fixed offset: =wireless_region="${current_lang:3:2}"=, with a comment reading "extract country code (positions 3-4)". That is correct only for a two-letter language code. - -=validate_config= accepts =^[a-z]{2,3}(_[A-Z]{2})?...=, so a three-letter language is a legal =LOCALE=, and glibc ships 75 of them (=agr_PE=, =ast_ES=, =ber_DZ=, =ayc_PE=, ...). Verified by running the expansion: =ber_DZ.UTF-8= yields =_D=, =ayc_PE.UTF-8= yields =_P=, =C= yields the empty string, =POSIX= yields =IX=. - -The sed that follows only uncomments an existing =#WIRELESS_REGDOM="XX"= line in =/etc/conf.d/wireless-regdom= (176 of them, owned by wireless-regdb). A garbage region matches nothing, sed exits 0, and the =|| error_warn= never fires — so the regdom is never set and nothing says so. The task line does print the garbage region ("configuring wireless regulatory domain (_D)"), so it is visible in the log rather than fully silent. - -Confirmed the mechanism itself works for the normal case: line 168 of this machine's =/etc/conf.d/wireless-regdom= reads =WIRELESS_REGDOM="US"= uncommented, which is archsetup's own edit. - -Grading: Minor severity (WiFi falls back to the conservative "00" regdomain — fewer channels and lower tx power, but WiFi works) × rare edge case (one of 75 three-letter-language locales, or a =LOCALE= with no country) = P4 = [#D]. - -Fix when it comes up: derive the country from the =_CC= group by pattern rather than by offset, and warn when it cannot be derived or when the sed changed nothing. Worth doing together with the sibling gap — nothing in the installer verifies that a =sed -i= uncomment actually matched, so a distro reshuffling one of these config files would fail the same silent way. All 22 =sed -i= sites share that stance, so it is a uniform design choice rather than an odd one out. -** DONE [#B] Initramfs hook swap can leave a LUKS machine unbootable :bug:installer:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as archsetup =HEAD= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). The swap moved into =switch_udev_hook_to_systemd=, which declines when =hooks_need_busybox_init= sees a standalone =encrypt= token, and the caller now rebuilds the initramfs only when the conf actually changed. New =tests/installer-steps/test_switch_udev_hook.py=, 10 tests; both guards proven real by breaking them (removing the refusal: 4 failures; loosening the token match to a bare =encrypt= substring: 1 failure). - -Found in the 2026-07-24 sentry bug-hunt, round 10. =configure_initramfs_hook= ran =sed -i '/^HOOKS=/ s/\budev\b/systemd/'= on any non-ZFS root, then =mkinitcpio -P=. Its only guard was =is_zfs_root=. - -Why that breaks a LUKS machine, verified against the installed mkinitcpio rather than argued: -- =/usr/lib/initcpio/install/systemd= line 70 is =add_symlink /init usr/lib/systemd/systemd=, so the systemd hook replaces the busybox init outright. -- =/usr/lib/initcpio/hooks/encrypt= is an =#!/usr/bin/ash= script whose entire body is a =run_hook()= function — the busybox init's mechanism. Under systemd init nothing calls it. -- =mkinitcpio= carries no conflict check for the pairing (grepped; nothing), so the rebuild succeeds and archsetup reports success. -- This machine's own =/etc/mkinitcpio.conf= documents the two valid pairings as separate examples: =udev= + =encrypt= (line 45) and =systemd= + =sd-encrypt= (line 51). The sed converted half of the first pairing and produced neither. - -Effect: on a LUKS root using the standard busybox =encrypt= hook, archsetup rewrites HOOKS to =systemd= while leaving =encrypt= behind, rebuilds the initramfs, and exits cleanly. At the next boot the root is never unlocked. The machine needs live media and manual mkinitcpio surgery to recover. - -The sibling asymmetry: =is_encrypted_root()= already exists in this script and =configure_autologin= uses it to branch on exactly this condition. The initramfs step consulted neither it nor HOOKS. =merge_grub_cmdline='s own comment names =cryptdevice== as a boot-critical parameter to preserve — and =cryptdevice== is read only by the =encrypt= hook, so archsetup explicitly anticipates the configuration that another of its steps then breaks. - -Grading: Critical severity (the machine will not boot and recovery needs external media — graded on the harm once in the failure state, not on how rarely it is entered) × some users, sometimes (LUKS-encrypted non-ZFS root using the busybox =encrypt= hook; deterministic for those machines, absent everywhere else) = P2 = [#B]. - -Deliberately not attempted: migrating =encrypt= to =sd-encrypt=. That means rewriting the kernel cmdline from =cryptdevice== to =rd.luks.name== against the volume's UUID, which is a real migration and not a mechanical edit. Refusing the cosmetic swap keeps a working machine working, which is the right trade against quieter fsck output. -** DONE [#D] keymap and consolefont hooks are inert under the systemd initramfs :bug:installer:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as archsetup =249bb93=. The swap rewrites both to =sd-vconsole=, collapsing them into one entry and never duplicating an existing one. The open question is answered: this machine is KEYMAP=us with no encrypt hook, but the function runs on LUKS machines where a non-US layout at the passphrase prompt is exactly what sd-vconsole restores. 7 tests. -Same class as the =encrypt= bug above, but cosmetic rather than boot-critical, so it was filed rather than bundled into that fix. - -Enumerating the busybox-only hooks on this machine (every hook under =/usr/lib/initcpio/hooks/= defining =run_hook=/=run_earlyhook=/=run_latehook=) gives: btrfs, consolefont, encrypt, grub-btrfs-overlayfs, keymap, memdisk, resume, sleep, udev, usr. All go inert once =/init= is systemd. Of those, =encrypt= is the only boot-critical one — =resume= is handled natively by systemd's hibernate-resume generator, and =btrfs= by udev rules (this machine runs =btrfs= alongside =systemd= and boots fine). - -=keymap= and =consolefont= are the live leftovers. Run =grep '^HOOKS=' /etc/mkinitcpio.conf= on this machine: the line carries =systemd= plus =keymap consolefont= and no =udev=, so archsetup's swap has already run here and both hooks are installed into the image and never executed. The systemd equivalent is the single =sd-vconsole= hook, which is what the distro's own systemd example on line 51 of =/etc/mkinitcpio.conf= uses. - -Effect: the early-boot console keeps the default font and keymap until =systemd-vconsole-setup= runs in the real root. =add_nvme_early_module= sets =FONT=ter-132n= in =/etc/vconsole.conf= expecting it to apply at that stage, so the configured font is briefly not what archsetup asked for. - -Grading: Cosmetic severity (a few seconds of default console font on a machine that boots normally) × some users, sometimes = P4 = [#D]. - -Fix when it comes up: have =switch_udev_hook_to_systemd= also rewrite =keymap consolefont= to =sd-vconsole= when it performs the swap, and add the fixture cases to =tests/installer-steps/test_switch_udev_hook.py=. Worth confirming first whether a non-US keymap is ever needed at the initramfs prompt on a machine that reaches this path. -** DONE [#B] NVIDIA Wayland preflight blocks dwm and headless installs :bug:installer:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as archsetup =HEAD= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). The NVIDIA block moved out of =preflight_checks= into a new =nvidia_preflight= function that returns early unless =desktop_env= is =hyprland= and archsetup is the one installing drivers. New =tests/nvidia-preflight/test_nvidia_preflight_gate.py=, 11 tests; each of the three guards was proven a real gate by deleting it and watching the suite go red (3, 1, and 1 failures respectively). - -Found in the 2026-07-24 sentry bug-hunt, round 10, reading archsetup's own installer. =preflight_checks= called =nvidia_preflight_report= unconditionally and exited 1 on rc 11 (repo driver below the 535 Wayland floor, or =pacman -Si nvidia-utils= unable to answer). The check is Wayland-specific — every line it prints names Wayland/Hyprland — but it ran before any =desktop_env= branch and consulted neither =desktop_env= nor =skip_gpu_drivers=. - -Effect, proven empirically rather than argued (three scenarios driven against the extracted block): =DESKTOP_ENV=dwm= plus =--no-gpu-drivers= on an NVIDIA machine with an old repo driver aborts the install; so does =DESKTOP_ENV=none=. Neither install ever runs a compositor, and =--no-gpu-drivers= means the user installs the driver themselves. Worse, the abort's own fix hint reads "install with DESKTOP_ENV=dwm (X11) instead" — the one remedy it prints is the one it refuses to honor, so the user has no working workaround short of editing the script. - -The sibling asymmetry that makes it an oversight rather than a decision: =install_gpu_drivers= returns early on =skip_gpu_drivers=, and =display_server= / =window_manager= both branch on =desktop_env= with a =none= arm that skips outright. The preflight gate applied neither ruling. - -Second defect at the same site, fixed in the same commit: the rc-10 path (card detected, driver fine) prompts with a bare =read=. =--config-file= is documented as "unattended installation", and =aur_install= already rules that a prompt not covered by =--noconfirm= "blocks forever waiting for input" on a headless install. The rc-10 prompt is advisory, so it now answers itself with its own =[Y/n]= default when a config file was supplied. rc 11 stays a hard stop either way. - -Grading: Critical severity (archsetup cannot be run at all on that machine, and the printed workaround does not work — graded on the harm once in the failure state, not on how rarely it is entered) × rare edge case (needs an NVIDIA card, a repo driver below the floor or an unsynced pacman db, and a non-hyprland =desktop_env=; hyprland is the default and Craig's own machines are AMD and Intel) = P2 = [#B]. - -Noted, not fixed: =display_server= and =window_manager= both point their unknown-value hint at a =--desktop-env= flag that the argument parser does not implement. Both arms are unreachable today (=validate_config= rejects a bad =DESKTOP_ENV=, and without a config file the value is always the default), so it is a stale string rather than a live defect. -** DONE [#B] mkplaylist retags only the first file :bug:music:quick:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as dotfiles =a57c443= (committed locally, deliberately NOT pushed — held for Craig's morning review of the sentry run). =ffmpeg -nostdin= on the conversion call. New =tests/mkplaylist= suite, 12 tests; removing the flag turns the suite red (verified by reverting: 5 failures, green on restore). NOTE: the fake ffmpeg does a full =cat > /dev/null= drain, which the real one does not do — so the suite gates the flag's presence, not the production failure mode. The docstring claiming the fake "drains stdin exactly when the real one would" is false and should be corrected. -Found in the 2026-07-24 sentry bug-hunt (shellcheck SC2095). =common/.local/bin/mkplaylist=: =generate_music_m3u= pipes the file list into =tag_music_file= (line 130), which consumes it with =while IFS= read -r file=. Inside that loop, =ffmpeg -i "$file" -vn -c:a flac "$outputfile"= (line 46) reads stdin by default for its interactive keyboard controls, so it consumes bytes the loop is relying on. - -CORRECTION (2026-07-24, from an adversarial re-review): the failure mode stated above — "the loop sees EOF and exits after the first file" — is WRONG, and this task originally asserted it. Measured under strace, ffmpeg polls fd 0 and reads roughly one byte per half-second of transcode wall time; flac encoding runs about 2000x realtime, so a ten-minute mp3 converts in ~0.28s and yields zero or one stolen byte, never a drain. Driven end to end with real ffmpeg against four 120s mp3s, three runs of three: all four were converted and retagged every time. The loop never truncated. - -What is real is the hazard, not the observed symptom: one stolen byte mangles a path, which makes mid3v2/metaflac fail and =set -e= abort the run loudly. =-nostdin= is still the right fix and the commit still stands. The original finding came from shellcheck SC2095 plus reasoning, and was never run — which is exactly what "verify before filing" exists to prevent. - -Effect: on a directory of non-flac audio, only the first file is converted and retagged. Files 2..N are silently skipped — no error, no output, and the playlist itself still generates (a separate =find=), so nothing signals that the retagging stopped. - -Grading: Major severity (the retagging feature is broken past the first file, and it fails silently) × most users frequently (the script exists to batch-process a directory, so more than one non-flac file is the normal case) = P2 = [#B]. - -Fix: =ffmpeg -nostdin= (or =< /dev/null= on the call). Verifiable with a fake =ffmpeg= on PATH asserting it is invoked once per input file. -** DONE [#C] timezone-change prints command-not-found instead of its help :bug:tooling:quick:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as dotfiles =15d2b63= (committed locally, deliberately NOT pushed — held for Craig's morning review), together with the Portugal-zone defect below. New =tests/timezone-change= suite, 12 tests. -Found in the 2026-07-24 sentry bug-hunt (shellcheck SC2288). =common/.local/bin/timezone-change=, default =*)= case (lines 63-67): =echo= sits alone on its own line, so the following quoted string runs as a *command* rather than as its argument. - -#+begin_src sh -*) - echo - "Invalid option chosen." - echo - "Some valid options are: eastern, central, pacific, rome, london, st_lucia, italy, france, spain ." - ;; -#+end_src - -The user gets two blank lines and two =command not found= errors; the list of valid options never prints. The timezone is correctly left unchanged, so this is an output defect only. - -Grading: Minor severity (wrong output on an error path, nothing corrupted) × some users sometimes (only on an unrecognized option) = P3 = [#C]. - -Fix: fold each string into its =echo=. Verifiable by running the script with a bogus argument and asserting the option list appears on stdout. -** DONE [#C] Thumbnail sweep wipes the whole cache when a wallpaper source is unreadable :bug:settings:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as dotfiles =0bd8c67= (committed locally, deliberately NOT pushed — held for Craig's morning review). 8 new tests. -Found in the 2026-07-24 sentry bug-hunt, reviewing the orphan sweep shipped the night before (dotfiles =e752a16=). =os.walk= stays silent about a directory it cannot enter, so =wallpaper.scan_sources= returns =[]= for a source that is missing, renamed, or permission-denied — the same answer it gives for a gallery the user emptied on purpose. =settings/cli.py= tick then hands that empty list to =thumbstore.sweep_orphans=, =live_names= comes back empty, and every cache-shaped file is classified an orphan. - -Proven empirically rather than reasoned: seeding three well-formed thumbnails plus a stray README, then sweeping against a nonexistent source directory, deleted all three (the README survived, so the cache-name regex guard works — it just doesn't help here). - -Effect once entered: the entire persistent thumbnail cache is deleted, so the next wallpaper-view open pays the cold-decode cost the cache was built to remove (measured at 3.7s for a viewport of Craig's largest 8, which is what tripped the compositor's kill prompt), and the tick needs roughly ten idle beats — about twenty minutes — to rewarm at =WARM_PER_BEAT= 8. - -Grading: Major severity (grading the being-in-it, per the don't-double-count-rarity rule: the cache is gone, the original freeze returns, and recovery is unattended and slow) × rare edge case (both configured sources — =~/videos/wallpaper= and =~/pictures/wallpaper= — are local directories, so this needs one deleted, renamed, or made unreadable while a beat fires; a removable or network source would hit it routinely) = P3 = [#C]. - -Fixed in this session: new =wallpaper.sources_available(sources)= tells "readable and empty" apart from "could not read", and =sweep_orphans= grew a =sources_ok= parameter that declines to sweep when it is False. Deferring a sweep costs only some stale files; sweeping wrongly costs the whole cache. -** DONE [#C] timezone-change sets a nonexistent zone for Portugal :bug:tooling:quick:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as dotfiles =15d2b63= (committed locally, deliberately NOT pushed — held for Craig's morning review). =Europe/Lisbon=. The suite also pins the general invariant: every zone the script can emit must exist in tzdata, so a future bad entry fails at test time rather than in Craig's hands. -Found in the 2026-07-24 sentry bug-hunt, validating every zone the script sets against =/usr/share/zoneinfo=. =common/.local/bin/timezone-change= line 39 maps =portugal= / =lisbon= to =Europe/Portugal=, which is not a tzdata identifier — the real one is =Europe/Lisbon= (a bare =Portugal= legacy alias also exists at the top level, but not under =Europe/=). =timedatectl set-timezone "Europe/Portugal"= fails, so the timezone is never changed. - -The other 17 zones the script sets all resolve correctly, so this is the single bad entry. - -Grading: Major severity (the option is wholly broken — the zone is not set and the command errors) × rare edge case (one option of eighteen, hit only when actually switching to Portugal) = P3 = [#C]. - -Fix: =Europe/Lisbon=. Verifiable by asserting the argument handed to a fake =timedatectl=, plus a suite-wide check that every zone the script names exists in the tzdata database. -** DONE [#C] settings-project stop() can SIGTERM an unrelated process :bug:settings:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Found in the 2026-07-24 sentry bug-hunt, round 2, reviewing =settings/src/settings/project.py=. =stop()= read a pid out of =$XDG_RUNTIME_DIR/settings-project.pid= and SIGTERMed it with no check that the pid still belonged to the projection. A projection that dies without running =stop()= (crash, OOM, a failed =execvpe= on the clock path — that last one was already noted as tolerated residue) leaves the file behind, so once the kernel wraps its pid counter that pid can name something else entirely, and the next =start= or =stop= kills it. - -This is a hazard the codebase had already ruled on elsewhere and simply hadn't applied here: =maint/src/maint/doctor.py= revalidates =/proc/<pid>/comm= against the expected name before its KILL remedy fires, explicitly to refuse recycled pids. - -Grading: Major severity (grading the being-in-it — an arbitrary user process takes a SIGTERM, and an editor with unsaved work is a plausible victim) × rare edge case (needs an unclean exit *and* pid reuse; =pid_max= here is 4194304, so wrap-around takes a very long time) = P3 = [#C]. - -Fixed as dotfiles =722994e= (committed locally, deliberately NOT pushed — held for Craig's morning review). The pidfile now records the process start time from =/proc/<pid>/stat= next to the pid, and =stop()= fires only when the recorded value still matches the live process. Start time is the right token rather than =comm=: it is mode-independent (the clock channel execs into =python3=, so comm changes while comm-matching would have needed per-mode knowledge) and it is exec-stable, verified directly — pid and start time were identical either side of an =execvpe=. A recycled pid cannot reproduce it. Legacy bare-pid pidfiles keep the old unconditional behavior so the upgrade never strands a live projection. -** DONE [#C] wtimer alarms fire an hour off on the eve of a DST change :bug:timer:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Found in the 2026-07-24 sentry bug-hunt, round 3, reading =timer/src/timer/engine.py=. =parse_alarm= resolves a bare wall-clock time ("07:00") to its next occurrence: it builds today's instant, and when that is already past it rolled forward with =epoch += 86400=. A DST day is 23 or 25 hours long, so a fixed 86400 lands on the wrong wall time whenever tomorrow crosses a transition. - -Reproduced against America/Chicago and the two 2026 US transitions. Asking for =07:00= at 08:00 on Sat 2026-03-07 (spring forward that Sunday) gave 08:00 Sunday — an hour late. Asking for =07:00= at 08:00 on Sat 2026-10-31 (fall back that Sunday) gave 06:00 Sunday — an hour early. - -The recurring path was never affected, which is what makes this an oversight rather than a design choice: =next_alarm= walks candidate days and rebuilds =datetime(y, m, d, hh, mm)= per day, so it is already DST-correct. Only the one-shot rollover took the shortcut. Both were pinned by the new tests. - -Grading: Major severity (grading the being-in-it — an alarm that fires an hour off has wholly failed at the one thing an alarm does, and the fall-back direction wakes you early while the spring-forward direction lets you oversleep) × rare edge case (two nights a year, and only when the requested wall time has already passed today) = P3 = [#C]. - -Fixed as dotfiles =9b6c2c9= (committed locally, deliberately NOT pushed — held for Craig's morning review). The rollover now rebuilds the local time on tomorrow's calendar date, the same construction =next_alarm= uses. Eight tests pin =TZ=America/Chicago= (saved and restored around each case), covering both transitions, the twelve-hour form, an ordinary-day control, a DST eve where the requested time is still ahead, and two characterization cases asserting the recurring path stays DST-safe. -** DONE [#B] net portal-restore claims encrypted DNS is back without checking :bug:net:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Found in the 2026-07-24 sentry bug-hunt, round 3, reading =net/src/net/repair.py=. A captive-portal login moves the DNS-over-TLS drop-in aside so plain DNS can reach the venue's login page, and =_restore_dot()= moves it back afterwards. It fired both privileged steps — the =mv= and the =systemctl restart systemd-resolved= — and returned ="restored"= without reading either result. =repair_portal_restore()= then rendered a pass step reading "DNS-over-TLS restored". - -So a declined or failed =sudo -n mv= left DNS-over-TLS off while the tool told the user it was back on. The same for a resolved restart that fails: the drop-in is on disk but the running resolver is still serving plain DNS. - -The asymmetry is what makes it an oversight rather than a decision. The sibling =_disable_dot()=, twenty lines up, checks its own move with =_ok()= and returns False rather than claiming a success it did not get. The restore half simply never got the same treatment, and it is the half where the failure is silent — the disable path's failure is visible immediately because the portal page won't load. - -Grading: graded on severity alone under the privacy carve-out. DNS queries continue in cleartext to the venue resolver on an untrusted network, and the affirmative "restored" message is what removes the user's reason to check. Bounded by =net diagnose='s =encrypted-dns= step, which exists precisely to catch a portal run that never restored, so the exposure ends at the next diagnose rather than persisting unseen forever. Major severity = P2 = [#B]. - -Fixed as dotfiles =018c0c5= (committed locally, deliberately NOT pushed — held for Craig's morning review). Both privileged steps are now checked, with two new outcomes: ="failed"= when the move back fails (encrypted DNS still off, rendered as a fail step) and ="unapplied"= when the drop-in is back but resolved would not restart (rendered as a warn step). Each names the command to run by hand. Four tests cover both failures at the =_restore_dot()= and step levels, mirroring the existing declined-move test on the disable side. -** DONE [#B] the portal restore watcher fails silently, so DNS stays in the clear :bug:net:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Found in the 2026-07-24 sentry bug-hunt, round 4, reading the rest of =net/src/net/repair.py= after the round-3 fix above. =portal_restore_watch()= polls until the link comes back online, calls =_restore_dot()=, and discards the outcome entirely. - -Three things compound into a silent failure. The watcher is spawned detached with =stdin=, =stdout=, and =stderr= all on =/dev/null=, so nothing it could print reaches anyone. It runs outside the =repair()= dispatch, so unlike every other mutating tier it never wrote an event-log line either. And =repair_portal_login= tells the user "encrypted DNS restores itself once you're online", which is precisely what removes their reason to check. A ="failed"=, ="unapplied"=, or ="ambiguous"= restore therefore left the machine on plain DNS on a venue network with no signal at any level. - -This is the round-3 finding one layer out, and the asymmetry is the tell: =018c0c5= taught =repair_portal_restore()= — the *manual fallback* — to stop claiming a success it did not get, while the *automatic* path, the one that actually runs in the normal flow, kept dropping the same result on the floor. Fixing the fallback and leaving the primary silent is a worse split than the original bug. - -Grading: graded on severity alone under the privacy carve-out, exactly as the round-3 sibling. Same exposure (cleartext DNS to an untrusted venue resolver), same bound (=net diagnose='s =encrypted-dns= step catches the stranded state), and the same affirmative promise removing the reason to look. Major severity = P2 = [#B]. - -Fixed as dotfiles =601c5b4= (committed locally, deliberately NOT pushed — held for Craig's morning review). The watcher now returns the outcome, appends a =portal-restore-watch= event with it, and fires a persistent =notify security= alert on each of the three failing outcomes, each naming the command to run by hand. A clean restore stays silent. Five tests: one per failing outcome, one pinning the silence on a clean restore, and one on the event-log line. The whole =TestPortalLogin= class now shadows =notify= with a logging fake, so no future watcher test can fire a real desktop notification mid-suite. -** DONE [#D] dns-override failure path says "reverted" without checking :bug:net:quick:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as dotfiles =2cf3fb3=. The revert is checked; a declined one now says 1.1.1.1 is still set and names =resolvectl revert <iface>=. -Found in the 2026-07-24 sentry bug-hunt, round 3, sweeping for siblings of the portal-restore finding above. =net/src/net/repair.py=, =repair_dns_override()= failure path: when the 1.1.1.1 override doesn't restore resolution, it calls =priv.run("dns-revert", iface)=, discards the result, and returns evidence reading "override didn't restore resolution — reverted". A failed revert leaves 1.1.1.1 set on the link while the step says it was removed. - -Same defect class as the portal-restore bug, three hundred lines up in the same file, and it survived the sweep only because the consequence is much smaller. Every other mutating repair in this file verifies by re-measuring afterwards rather than by reading an exit code, which is the stronger pattern and is why the sweep otherwise came back dry. - -Grading: Minor severity (a stale per-link override sends DNS to Cloudflare instead of the venue resolver, it dies on the next reconnect, and =net diagnose='s =dns-override-present= step exists specifically to catch it) × rare edge case (needs the override to fail *and* the revert to fail) = P4 = [#D]. - -Fix: the same idiom the portal-restore fix now uses. Wrap the revert in =_ok()= and drop the "— reverted" claim (or say the revert failed and name =resolvectl revert <iface>=) when it returns False. The existing =RepairHarness= makes the privileged call fail with =NET_SUDO="false"=, so the test is a near-copy of =test_restore_reports_failure_when_the_move_back_is_declined=. -** DONE [#B] a timezone-less Date header crashes the whole net diagnose run :bug:net:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Found in the 2026-07-24 sentry bug-hunt, round 4, reading =net/src/net/diag.py=. =_clock_skew_s()= fetches the probe server's =Date= header with =curl -sI=, parses it with =parsedate_to_datetime=, and subtracts it from a timezone-aware =datetime.now(timezone.utc)=. RFC 5322 allows a =Date= to carry =-0000=, which means UTC while explicitly claiming no local zone, and a =Date= with no zone at all parses leniently as well. Both come back *naive*, and subtracting a naive datetime from an aware one raises =TypeError=. - -The =try= wraps only the =parsedate_to_datetime= call, so the =TypeError= from the line below it is uncaught. It escapes =_clock_skew_s=, escapes =_steps_egress_edges=, and takes down the entire =diagnose()= run — no report, no steps, a Python traceback. =net doctor= runs diagnose first, so the panel's doctor button dies with it. - -Verified against Python 3.14.6 before writing the fix: =parsedate_to_datetime("Thu, 01 Jan 2020 00:00:00 -0000")= returns =tzinfo=None=, and the subtraction raises. The zoneless form behaves the same. Only the =GMT= form (which the well-behaved probe host sends) comes back aware, which is why this never showed up in normal use. - -What makes it more than a curiosity is *when* the code runs. =_steps_egress_edges= fires only after the http-probe has already failed, so the server answering that =HEAD= is frequently a captive portal's interception appliance rather than the real probe host — and a minimal embedded HTTP stack is exactly the kind that emits a non-GMT =Date=. The one path guaranteed to be talking to a non-standard server is the one that can't survive a non-standard header. - -Grading: Major severity (grading the being-in-it — the diagnostic tool produces no report at all, and =net doctor= goes with it, on precisely the broken network it exists to diagnose) × rare edge case (needs a failing probe *and* a portal appliance that omits a numeric offset) = P2 = [#B]. - -Fixed as dotfiles =8933500= (committed locally, deliberately NOT pushed — held for Craig's morning review). A naive parse is now read as UTC, which is what =-0000= means. Two tests, and the second is the one that matters: it drives a *current* =-0000= timestamp and asserts no clock row, so a lazy "catch =TypeError= and return None" fix would fail it while the correct reading passes. -** DONE [#C] a tunnel import that can't be disarmed still reports success :bug:net:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Found in the 2026-07-24 sentry bug-hunt, round 4, reading =net/src/net/manage.py=. =import_config()= imports a WireGuard or OpenVPN config as an NM profile, then fires =nmcli connection modify <uuid> connection.id <name> connection.autoconnect no= — and discarded the result, returning =ok=True= regardless. - -That modify is the whole safety of the feature, and the module's own docstring says so: =nmcli connection import= *auto-activates* the profile it creates, "which nobody asked for by picking a file", so "every import here ends with the profile deactivated and autoconnect off". A failed modify inverts that. For WireGuard — a device-type connection — autoconnect stays on, so the tunnel re-arms itself at the next boot and takes the default route with it, and the profile keeps the transient staged interface name (=wgpvpn=) while the envelope reports the config's real name, so the panel names a profile that isn't there. - -CORRECTION (2026-07-24, from an adversarial re-review): the blanket claim originally written here — that a failed disarm re-arms the tunnel at boot — is wrong for OpenVPN. =man 5 nm-settings-nmcli= states autoconnect is not implemented for VPN profiles, and an OpenVPN import is an NM VPN profile, so the modify is near-cosmetic on that half. The bug is real and security-relevant for WireGuard, which is the primary case; the severity as stated overreached to cover both. - -The asymmetry, again the tell: =_nmcli_import()=, twenty lines up in the same file, checks its own =returncode= and raises rather than return a UUID it did not get. The modify below it never got the same treatment. - -Grading: Major severity (grading the being-in-it — a full-tunnel VPN the user never asked to connect arms on every boot and carries all their egress, it persists across reboots rather than self-healing, and the affirmative "imported X" is what removes the reason to check) × rare edge case (needs the modify to fail after the import succeeded) = P3 = [#C]. - -Fixed as dotfiles =e0d4d8a= (committed locally, deliberately NOT pushed — held for Craig's morning review). New =_disarm()= returns whether the modify took. On failure the profile is still deactivated first — the import already brought it up, and the verdict shouldn't decide whether it keeps running — and then a =disarm-failed= envelope names the UUID and the exact command to finish the job. Three tests: the failing verdict, =import_configs= counting it as failed rather than imported, and a characterization test pinning that the deactivate still runs on the failure path. -** DONE [#C] a binary that can't be exec'd crashes the panels instead of degrading :bug:net:bluetooth:audio:maint:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Found in the 2026-07-24 sentry bug-hunt, round 5, comparing the four panel packages' subprocess wrappers against each other. - -Every wrapper in the panels states the same contract: an unusable tool becomes a degraded result, never an exception. =cmd.run= returns None; =nmcli.run=, =btctl.run= and =pactl.run= raise their own domain error, which every caller already guards on; =speedtest.run_speedtest= returns an error envelope. All of them caught only =FileNotFoundError=, so they kept the contract for a tool that is *absent* and broke it for a tool that is *present but unusable*. - -Verified against Python 3.14.6 rather than argued. =subprocess.run= raises =PermissionError= for a file without its execute bit, =OSError= (ENOEXEC, "Exec format error") for an executable file that is neither a binary nor a script with a shebang, =NotADirectoryError= when a path component is a plain file, and =OSError= when a fork is refused under memory or PID pressure. None of the four is =FileNotFoundError=, so each escapes the guard: waybar's net/bt/audio modules die rather than dimming, and a maint probe takes the whole envelope with it — in exactly the machine state maint exists to report on. - -The asymmetry, and this codebase had already ruled on it three separate times: =net/iw.py='s =signal_dbm= and =settings/spawn.py='s =detached= both catch =(OSError, subprocess.TimeoutExpired)=, and =audio/cmd.py='s doctor-tier =probe()= enumerates =FileNotFoundError=, =NotADirectoryError= and =PermissionError= as "absent" under a docstring promising it never raises. Its sibling =run()=, twenty lines up in the same file, kept the narrow catch — as did all five copies of =run()= and all three tool wrappers. =audio/status.py='s docstring records that this same class already bit once ("the bar's audio module died rather than dimming"); that fix widened the guard's *scope* and left its *exception set* alone. - -Grading: Major severity (grading the being-in-it — the status surface is dead while the condition holds, and for maint the tool that reports the fault is the one that dies of it; no data loss, and it clears when the tool or the pressure does) × rare edge case (needs a binary with wrong permissions, a lost shebang, or a fork refused under pressure) = P3 = [#C]. - -Fixed as dotfiles =44fdae1= (committed locally, deliberately NOT pushed — held for Craig's morning review). Widened to =OSError= across net, bt, audio, maint and panelkit — five =cmd.run= helpers, the three tool wrappers, =probe._curl= and =speedtest.run_speedtest=. The domain-error wrappers keep their "<tool> not found" message for a genuinely absent binary and add a second arm naming the errno for an unusable one, so the report can still tell the two apart. 28 tests, one class per package, driving all three exec failures against real files on a temp PATH; each was watched failing against unmodified production code first (27 red). Audio's class carries a characterization case pinning =cmd.probe='s existing behavior, so the sibling that got this right can't regress into the one that didn't. -** DONE [#C] a failed pty-backed spawn strands both ends of the pty :bug:net:bluetooth:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Found in the 2026-07-24 sentry bug-hunt, round 6, auditing the =subprocess.Popen= sites the round-5 fix didn't reach. - -Two spawns open a pty before launching and catch only =FileNotFoundError= around the =Popen=: =bt/pairing.py='s =pair_interactive= (bluetoothctl under a pty so the passkey agent is interactive) and =net/speedtest.py='s =run_speedtest_stream= (speedtest-go under a pty because it buffers everything to exit when piped). Both are the same exec-failure class as =44fdae1= — a binary present but not executable raises =PermissionError=, a lost shebang raises =OSError= — and neither is =FileNotFoundError=. - -What makes these worse than the =run= wrappers is where the cleanup lives. =os.close(master)= and =os.close(slave)= sit *inside* the =FileNotFoundError= arm, so an escaping =OSError= skips them: every failed attempt strands two descriptors. Both call sites are buttons in a long-lived panel process — the pairing flow and the console's SPEED key — and a user who gets no feedback presses again, so the leak accumulates under exactly the conditions that caused it. - -Grading: Major severity (grading the being-in-it — a descriptor leak in a process meant to run for days, on a path the user retries, plus the exception escaping a documented "(ok, detail)" / error-envelope contract) × rare edge case (needs an unusable bluetoothctl or speedtest-go) = P3 = [#C]. - -Fixed as dotfiles =c2eb3e1= (committed locally, deliberately NOT pushed — held for Craig's morning review). An =OSError= arm on each closes both ends and returns the module's own failure shape, naming the errno. Four tests: two pin the return contract, two count =/proc/self/fd= across three attempts — the fd count is what actually fails against unmodified code, and it was watched failing before the fix. - -The wider sweep this came from is recorded so it isn't repeated: every =except FileNotFoundError= in production was enumerated. The other exec sites were already correct (=maint/gui.py= x3, =net/kick.py=, =timer/engine.py= x2, =timer/gui.py=, =net/repair.py= x2, =audio/peak.py= all catch =OSError=), and the remaining hits are file-open catches, not exec. =clock/__main__.py='s =toggle()= has no guard at all but spawns =sys.executable=, which is by definition runnable; not filed. -** DONE [#C] one impatient client kills the clock panel's toggle listener for good :bug:clock:waybar:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Found in the 2026-07-24 sentry bug-hunt, round 6, sweeping every acquired resource (pty, socket, mkstemp, tempdir) for cleanup that isn't in a =finally=. - -=clock/src/clock/app.py='s =_listen()= guards =accept()= with =except OSError: return= and leaves the request body — =recv=, =runtime_log=, =sendall= — outside any guard. =send_toggle()= in =__main__.py= gives the panel 0.25s to acknowledge, then closes. An ack later than that hits a dead peer and raises =BrokenPipeError=, which escapes the =while= loop and ends the listener thread. - -Verified empirically, not argued: a client that connects, sends, and gives up after 250ms makes the server's =sendall= raise =BrokenPipeError= (errno 32) and the listener thread exits. - -What makes it Major rather than a nuisance is that it neither self-heals nor announces itself. The socket file stays bound, so every later =clock toggle= still *connects* — then stalls the full 250ms, gets no reply, and falls through to spawning =clock serve=. GTK's single-instance forwarding turns that into =do_activate= on the running service, and =do_activate= calls =show_clock()=, not =toggle()=. So from the first bad client onward, clicking the waybar time module opens the panel every time and never closes it; the only ways out are the right-click dismiss inside the panel or restarting the service. Nothing logs it. - -Grading: Major severity (grading the being-in-it — the toggle is one-way from then on, it persists for the life of the service, and there is no signal it happened) × rare edge case (needs a reply to miss the 250ms budget: a busy main loop mid-redraw, a slow runtime-log write, or an interrupted =clock toggle=) = P3 = [#C]. - -Fixed as dotfiles =7c02614= (committed locally, deliberately NOT pushed — held for Craig's morning review). An =OSError= arm around the request body scopes a dead peer to its own request, mirroring the guard =accept()= already had. =GLib.idle_add= runs before the ack, so the user's click still takes effect — only the acknowledgement is lost. New =tests/clock/test_socket.py=, 3 tests driving the real =_listen= against a stand-in owner (it touches only =self._socket= and =self.toggle=, so no Gtk.Application is needed). The gate is the second toggle after an impatient first: it times out on unmodified code because no listener is left. The other two pin what the fix must preserve — the toggle fires even when the ack can't be delivered, and an unknown command is still answered without toggling. - -Left alone deliberately: =do_activate= calling =show_clock()= rather than =toggle()=. Changing it would alter what a cold =clock toggle= does on first launch, which is a design call for Craig rather than part of this defect. Worth raising if he ever wants the spawn path to toggle too. -** DONE [#B] fuzzel breaks the pinentry protocol loop on every passphrase :bug:security:gpg:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Found in the 2026-07-24 sentry bug-hunt, round 7 — from the live journal rather than from reading. Grepping this boot for tracebacks turned up four instances of =pinentry-fuzzel: line 36: read: 0: read error: Resource temporarily unavailable=, and every one sits 4-7 seconds after a =GETPIN= (the time it takes to type a passphrase). The =BYE= handler's log line never appears once. - -=hyprland/.local/bin/pinentry-fuzzel= speaks the Assuan pinentry protocol on a pipe gpg-agent keeps open, reading one command per iteration of =while read cmd rest=. The =GETPIN= arm shells out to fuzzel, which *inherits that pipe as its stdin*. fuzzel runs an event loop over its own input, so it sets =O_NONBLOCK= on fd 0 — and =--dmenu= would read the pipe as menu items besides. The flag lands on the shared open file description and outlives fuzzel, so the shell's next =read= fails with =EAGAIN= and the loop ends mid-protocol. - -Grading: Minor severity (the passphrase is delivered *before* the break, so decrypts still succeed and nothing is corrupted — what's lost is everything after: =BYE= is never acknowledged, and gpg-agent's same-connection retry after a wrong passphrase, =SETERROR= then =GETPIN= again, can't be served; that retry is what the script's "reenter" label exists for, and it has never once been reachable) × every user, every time (four for four in the journal, and the test reproduces it deterministically) = P2 = [#B]. - -Fixed as dotfiles =e727dcd= (committed locally, deliberately NOT pushed — held for Craig's morning review). =< /dev/null= on the fuzzel call, so the non-blocking flag lands somewhere harmless; =--lines 0= was already there, so no menu input was ever wanted. =ENABLE_LOGGING= became env-overridable as a test seam — the script logs through an absolute =/usr/bin/logger= that PATH can't shadow, so without it every test run would write ten lines into the real journal. - -New =tests/pinentry-fuzzel/=, 8 tests driving the real script over a live pipe the way gpg-agent does. The fake fuzzel sets =O_NONBLOCK= on whatever fd 0 it is handed, exactly as the real one does, which is what makes them a gate rather than a restatement of the fix. Four fail against unmodified code — one reproducing the journal's message verbatim — and one records the fd fuzzel was given, pinning the cause rather than the symptom. - -THE CALIBRATION NOTE, and it is about my own earlier sweep. This is the same shape as round 1's =a57c443= (ffmpeg draining the pipe a =while read= loop was consuming). Round 1 swept both repos for siblings of that bug and came back empty — because it searched for the *mechanism* (a child that drains stdin) rather than the *shape* (a child that inherits stdin at all inside a read loop). Two different mechanisms, one shape, and the narrower search missed a live daily-use instance. Scope a class sweep by shape, not by the mechanism of the first instance found. -** DONE [#C] a truncated webcam record strands every camera off :bug:settings:privacy:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Found in the 2026-07-24 sentry bug-hunt, round 8, sweeping production for non-atomic file writes. - -=settings/src/settings/webcam.py='s =_record()= wrote =~/.local/state/settings/webcam.json= with a plain truncate-in-place =open(path, "w")=. That record is the only route back on, and the module docstring says so: deauthorizing a camera removes its video4linux nodes, so =usb_devices()= returns nothing afterward and =_recorded()= becomes the sole source of the paths to re-authorize. A write that truncated and then failed left an empty file; =_recorded()= caught the resulting =JSONDecodeError= and returned =[]=; =_known_devices()= then had nothing; and =set_power(True)= returned None without re-authorizing anything. Every camera stranded off, with no way back through the panel until a replug or a reboot. - -The asymmetry, seventh instance of this read: six other state writers in the tree already write through a temp file and a rename — =maint/cache=, =net/cache=, =audio/ptt=, =timer/engine=, =settings/store=, =maint/curation=. The one whose loss is most expensive was the one that didn't. - -Grading: Major severity (grading the being-in-it — the privacy switch becomes one-way, the panel offers no route back, and the user has to know to replug the camera or write sysfs by hand; bounded by the fact that a reboot re-enumerates USB and restores authorized=1) × rare edge case (needs a crash or ENOSPC inside a microsecond-wide write window) = P3 = [#C]. - -Fixed as dotfiles =8b40b79= (committed locally, deliberately NOT pushed — held for Craig's morning review). =_record= now mirrors =store.save=: =mkstemp= in the target directory, write, =os.replace=, unlink the temp on any failure. Four tests; the gate is a =_record= whose =json.dump= raises, after which the previous record must still be readable — it isn't on the old code. The other three pin what the fix must preserve: no temp-file residue, the =_recorded()= round trip, and the end-to-end power-off/power-on with the class symlinks removed, which is the scenario the record exists for. - -HOW IT WAS FOUND, and it confirms round 7's lesson twice over. Round 4 ran an atomic-write sweep and reported "nine sites, six unique-per-writer, three sharing a fixed =.tmp=" — it enumerated the writers that *were* atomic and compared their temp-file naming, and never asked which state writers aren't atomic at all. Same narrowing that made round 1's stdin sweep miss the pinentry bug: the sweep was scoped to a property of the instances already found rather than to the shape of the hazard. -** DONE [#D] a failed wallpaper apply reports "nothing to apply" :bug:settings:quick:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as dotfiles =2cf3fb3=. The decision moved to =gui.wallpaper_apply_toast=, a module-level pure helper, because the callback lives inside a GTK widget where no test can reach it. 5 tests. -Found in the 2026-07-24 sentry bug-hunt, round 7, sweeping the settings panel's worker callbacks. - -=settings/gui.py='s =_async= passes an exception through as the *result* rather than as a separate error argument, so every =done= callback has to test =isinstance(res, Exception)=. Five do — =_mx_pin=, =_mx_letter=, =_after_matrix=, =_set_pointer=, the drum/dial/gallery/refresh callbacks. =_wp_apply= is the one that doesn't: - -#+begin_src python -def _wp_apply(self, note="Wallpaper set"): - self._async(lambda: panel.wallpaper_apply(self.state), - lambda ok: self._toast( - note if ok is True else "nothing to apply", - good=ok is True)) -#+end_src - -=panel.wallpaper_apply= calls =store.save=, which can raise =OSError= (disk full, a permissions change on the config dir). The exception then arrives as =ok=, =ok is True= is False, and the toast reads "nothing to apply" — describing a no-op when the apply actually failed. The toast is at least marked =good=False= (red), so the user gets a negative signal; what's lost is the reason, which every sibling callback surfaces via =str(res)=. - -Grading: Minor severity (wrong text on an error path, correctly marked as a failure, nothing corrupted) × rare edge case (needs =store.save= or =wallpaper.apply= to raise rather than return False) = P4 = [#D]. - -Fix: give it the same =isinstance(res, Exception)= arm its five siblings have — toast =str(res)= on an exception, keep the current two-way message otherwise. One callback, three lines. -** DONE [#D] two manage.py nmcli reads sit outside their own error conversion :bug:net:quick:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as dotfiles =2cf3fb3=. =_key_mgmt= converts both nmcli exceptions to "", which both call sites already treat as neither wpa-eap nor sae. 2 tests, including one driving =_classify_up_failure= end to end. -Found in the 2026-07-24 sentry bug-hunt, round 4, reading =net/src/net/manage.py=. =nmcli.run()= raises =NmcliTimeout= on timeout and =NmcliError= on a missing binary, and every mutation in this module is written to convert both into a result envelope. Two calls escape that conversion because they run through =_key_mgmt()=, which wraps =nmcli.get_value= and catches nothing: - -- =edit()= line 243 calls =_key_mgmt(uuid)= for the enterprise-profile refusal *before* its own =try=, while the next four lines catch exactly those two exceptions around =nmcli.run=. -- =_classify_up_failure()= calls it on =up()='s failure path, so a slow =connection show= turns a classifiable activation failure into an exception. - -Consequence is a leaked exception where the caller expected an envelope. The panel absorbs it — =gui.bg()= catches =Exception= and renders =str(e)= — so there it degrades to a worse message rather than a crash. =net edit= from the CLI has no such catch and prints a traceback. - -Grading: Minor severity (the operation fails either way; what's lost is the classified message, and only the CLI path shows a traceback) × rare edge case (=connection show= has a 2s timeout and nmcli's presence is already established by the time either site runs) = P4 = [#D]. - -Fix: give =_key_mgmt= the same conversion its callers use — catch =(nmcli.NmcliError, nmcli.NmcliTimeout)= and return "", which both call sites already handle correctly (neither "wpa-eap" nor "sae"). One =try= in one helper covers both sites. -** DONE [#D] three atomic writers share one fixed .tmp name :bug:quick:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Fixed as dotfiles =2cf3fb3=. All three carry =.tmp.$(getpid)=, matching the six writers that already did. 6 tests across audio and maint. -Found in the 2026-07-24 sentry bug-hunt, round 4, sweeping both repos for the temp-file half of the atomic-write idiom. The tree writes state atomically in nine places, and six of them make the temp path unique per writer: =net/cache.py= and =timer/engine.py= both use =f"{path}.tmp.{os.getpid()}"=, and =settings/store.py=, =settings/idle.py=, =bt/repair.py=, =net/probe.py= all use =tempfile.mkstemp=/=NamedTemporaryFile=. Three use a bare =path + ".tmp"=: - -- =audio/src/audio/ptt.py= =write_state= (the lead carried over from round 3's Next Steps) -- =maint/src/maint/cache.py= =put= -- =maint/src/maint/curation.py= =_write_user= - -=os.replace= makes the *rename* atomic, but a shared temp name is not: two writers open the same path, the second truncates under the first, and the file that gets renamed into place is a blend of both. The loser's own =os.replace= then raises =FileNotFoundError=, because the winner already renamed the name out from under it. - -Real concurrent-writer pairs exist for two of the three. =maint/cache.py= =updates_repo= is written by =maint-net-scan.timer= hourly and again by =doctor._fresh_pending()= at UPDATE fire time. =audio/ptt.py= has three writers by design (the CLI toggle bound to a key, the waybar right-click, and the GTK panel) — its module docstring says so. =curation.py= is written by panel key presses and CLI verbs. - -Grading: Minor severity (every reader degrades rather than crashes — =cache.get= catches =ValueError= and reports no data, =read_state= reads a torn file as disarmed, and both recover on the next write; the sharpest edge is the loser's =FileNotFoundError= aborting the rest of =scan_net=, which the next hourly run repairs) × rare edge case (the write window is a millisecond or two, and the overlapping writers are an hourly timer against a human keypress) = P4 = [#D]. - -Fix: give all three the =f"{path}.tmp.{os.getpid()}"= form the two careful siblings already use. It is three one-line changes and needs no new abstraction. Note this closes the torn-file half only — the read-modify-write in =ptt.toggle_plan= and =curation.set_preference= can still lose an update between two writers, which wants a lock rather than a temp-name change and should stay a separate decision. -** DONE [#C] dmenuexitmenu word-splits its menu so no entry matches :bug:dwm:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -Found in the 2026-07-24 sentry bug-hunt (shellcheck SC2128). =dwm/.local/bin/dmenuexitmenu= line 4 expands the menu unquoted: =choice=$(echo -e $menuitems | dmenu ...)=. Word-splitting collapses the runs of spaces the labels carry, so dmenu shows =Lock= where the =case= arm expects =Lock = (two spaces) and =Logout = where the arm expects a trailing space. No arm matches, so choosing an entry does nothing at all. - -CORRECTION (2026-07-24): THE BUG AS FILED DOES NOT EXIST. Ran it. =echo= rejoins the words split off the unquoted expansion with single spaces, and no label carries two spaces, so quoted and unquoted produce byte-identical output — verified against the exact literals from git rather than a retyped copy. Every =case= arm matches and every menu action works. - -What is real is latent. An unquoted expansion collapses a double space and glob-expands a =*=; the second was demonstrated turning a label into a directory listing. No current label triggers either. - -Hardened anyway in dotfiles =2cf3fb3= as robustness, not as a bug fix: the expansion is quoted and the bogus one-element array is now a plain string. Output confirmed unchanged byte-for-byte. New =tests/dmenuexitmenu/= (10 tests) pins the working behaviour, and shellcheck on the file drops from three findings to one. - -SECOND SENTRY FILING DISPROVED BY RUNNING IT, after =a57c443= (mkplaylist). Both came from a shellcheck hit plus reasoning, neither was executed. A static-analysis finding says a construct is unsafe, not that it currently misbehaves, and both filings treated the first as the second. -** DONE [#C] Timer module hero hierarchy :feature:waybar:timer:quick:solo: -CLOSED: [2026-07-24 Fri] -From the roam inbox (Craig, claimed 2026-07-22). Which display ("hero") wins the waybar timer module when several timer modes run simultaneously: pomodoro wins over everything (the user is actively working; it's likely their main focus). The rest rank in chronological order of when they would ring. Worked example: with a just-started 15-min timer, a 1-hr timer at 10 minutes left, a pomodoro, and an alarm ringing in 12 minutes — show the pomodoro; when it completes, the 1-hr timer (rings first), then the alarm, then the 15-min timer. Feeds the timer-panel spec (docs/specs/2026-07-02-timer-panel-spec.org). - -Shipped as dotfiles =9eedb39=. Pomodoro wins the hero, then soonest-to-ring, in both selectors (=engine.select_primary= for the bar, =panel.primary_id= for the GTK hero). Craig's worked example is a test. FLAGGED FOR CRAIG: the two selectors diverge on a *ringing* alarm (the bar excludes it, the panel gives it the hero) and I left that as-is rather than reverse a deliberate choice. Whether to unify them is your call. -** DONE [#C] Timer module: drop RING message, persistent notifications :bug:waybar:timer:quick:solo: -CLOSED: [2026-07-24 Fri] -From the roam inbox (Craig, claimed 2026-07-22). Remove the RING message from the timer module display; verify all timer and alarm notifications are persistent; the icon returns to normal once the notification has fired. Rationale: keeps timers and pomodoros from interfering with one another's displays (pairs with the hero-hierarchy task above). - -Shipped as dotfiles =9eedb39=. The tooltip no longer prints RING or a (ringing) suffix; a fired alarm shows its clock time and its persistent notification carries the alert. Verified the timer and alarm completion notes already set persist=True. -** DONE [#C] PTT icon outline removal :bug:waybar:quick:solo: -CLOSED: [2026-07-24 Fri] -From the roam inbox (Craig, claimed 2026-07-22): the waybar PTT icon should not have an outline. Cosmetic × every-glance = P3 = [#C]. - -Shipped as dotfiles =e63c0cf= (live style.css + dupre theme source). Removed the amber/green text-shadow glow from the armed/talk states, the only outline-like effect on the icon. FLAGGED FOR CRAIG: this is my read of "outline" (the glow). If you meant the glyph shape itself, it's a one-line revert. Confirm live by pressing PTT. -** DONE [#B] Dotfiles tests leak state across files :bug:test:dotfiles:solo: -CLOSED: [2026-07-23 Thu] -Resolved 2026-07-23 as dotfiles =c333598=. The polluter was =tests/weather/test_weather.py=, and it accounted for all 38 failures on its own. - -The mechanism was not the env leak the body below guessed at — tests/weather never writes =os.environ=. Its whereami fake did =weather.subprocess.run = ...= on a freshly-loaded module object. The fresh module isolated the weather code, but =weather.subprocess= is the one shared stdlib module object every module in the process holds, so the assignment replaced =subprocess.run= process-wide and never restored it. Every later test file got weather's fake result back from =subprocess.run=; the tell was wtimer asserting on =r.returncode= and getting "'R' object has no attribute 'returncode'", where =R= is weather's fake result class. - -Triage: TEST HYGIENE, not production global state. The weather script reads env at import and never writes, so no long-lived-process caching defect sits behind it. A scan for the same pattern (patching a stdlib module attribute reached through another module's namespace) finds exactly one instance in the suite — the three other =setattr= sites all snapshot and restore. So the planned shared env helper across 28 files was aimed at the wrong target and wasn't needed. - -Fix: rebind the loaded module's own =subprocess= name to a stub namespace, so nothing outside that module changes and there is nothing to restore. - -Gate: =make test= now runs two gates per the add-don't-replace decision — =test-forked= (one process per file, catches order dependence) and the new =test-shared= (every suite in one process, catches leakage). Built on stdlib unittest rather than pytest, since pytest was only the diagnostic tool and isn't a project dependency. Verified as a real gate, not just green today: with the defect deliberately reintroduced it goes red, and green once restored. A focused test in tests/weather pins the invariant on the culprit as well, because the shared gate alone blames the three victim files. - -Verification: 3500 tests, both gates, exit 0. - -Original finding follows. - -Found 2026-07-23 during the speedrun. =make test= is green, but it runs each test file in its own =python3 -m unittest= process, which hides cross-file state leakage. A single-process whole-tree run (=python3 -m pytest tests/ -p no:randomly=) fails 38: 22 in =tests/wtimer/test_wtimer.py=, 10 in =tests/zoom-web/test_zoom_web.py=, 6 in =tests/wlogout-menu/test_wlogout_menu.py=. - -Not a regression — a worktree at the pre-speedrun commit produces the identical 22/10/6 profile, so this predates tonight's work. Those three files also pass cleanly when run together (170 passed), so the polluter is a fourth file somewhere in the tree that mutates global state (env var, cwd, or a module-level patch) without restoring it. 28 test files write =os.environ= directly. - -Why it matters: the green gate can't see this class of bug, so a real isolation defect — or a genuine failure that only appears under a different order — passes CI silently. Bisect by running the tree with subsets until the polluter is identified (pytest's =-p no:randomly= keeps the order stable while bisecting), fix its cleanup, then decide whether =make test= should gain a single-process pass so the gate covers it. -** DONE [#B] Wallpaper view freezes the panel — thumbnail decode :bug:dotfiles:solo: -CLOSED: [2026-07-23 Thu] -Craig reported 2026-07-23: selecting the wallpaper button freezes the module and the compositor asks whether to kill it. Root cause proven: =_Thumb._draw= decoded each source image with =new_from_file_at_scale= on the GTK main thread. Measured on Craig's 78 wallpapers — a viewport of the 8 largest takes 3.7s, the whole set 13s. That block trips Hyprland's "not responding" watchdog. - -Grading: Critical severity (panel unusable, watchdog kill) × every user every time the wallpaper view opens = P1 = [#A] by the matrix. Held at [#B] because step 1 already shipped and removes the user-visible freeze; the remainder is a latency enhancement, not a showstopper. - -*** 2026-07-23 Thu @ 15:40 Step 1 — async decode (dotfiles f45f321) -Moved the decode to a worker thread via a new =settings/thumbcache.py= (pure, injected decode/scheduler/thread; 6 tests). The thumb shows its dark ground until the pixbuf lands, then redraws. Verified live on a headless output: worst main-loop stall opening the pair view dropped from multi-second to 68ms; the cache filled with 81 decoded pixbufs (the one miss is a .webm, correctly falling back to the ▶ glyph). Full suite 3512, both gates, smoke OK. This alone fixes the reported freeze. - -*** 2026-07-23 Thu @ 16:30 Step 2 — persistent on-disk cache (dotfiles 463cc4f) -Built the persistent layer: =settings/thumbstore.py= decodes each source once to a 512px PNG under =~/.cache/settings/thumbs=, keyed by path + mtime so an edited wallpaper self-invalidates. The hot-path decode reads that PNG and scales in-memory. Warming rides the existing =settings tick= CLI verb (the 2-min timer already runs it), building up to =WARM_PER_BEAT=8= missing thumbnails per beat — best-effort, journals a line on failure, never blocks the wallpaper flip. thumbstore is pure (stat/decode/load/save injected); 10 tests. - -Went with incremental warming (8/beat, ~10 beats to full) as the safe default rather than full-warm-on-change — the per-beat cap is a one-line flip if Craig wants it faster. Measured: hot-path decode of a viewport dropped from 3.7s cold to 47ms warm. No installer change (the tick service already runs =settings tick=); cache lives outside the repo. Full suite 3522, both gates, smoke OK, live panel verified (81 pixbufs render, 48ms worst stall warm). -** DONE [#C] Panel scrollbars too short :bug:dotfiles:quick:solo: -CLOSED: [2026-07-23 Thu] -Shipped 2026-07-23 as dotfiles =0d64837= (22px scrollbar, 16px trough, 14px slider thickness with a 48px floor along the travel axis). Left open by oversight during the speedrun; closing now. - -Follow-on, and my own regression: enlarging the bar to 22px is what made it start covering the thumbnails, because nothing grew the tray to match. Craig reported it the same day ("scrollbars that obscure the images") and it's fixed in =c0ddf57= — the tray now reserves a 22px lane for the bar as a margin on the scrolled box, so the bar sits below the images instead of across them. Measured before: tray 68px, content 68px, a visible 14px bar inside the same 68px. After: tray 90, content 68, bar clear. The lane is a constant under the scrollbar CSS with a note to keep the two in step, since the coupling between bar thickness and tray height is exactly what broke. - -From the roam inbox (Craig, claimed 2026-07-23): all scrollbars need to be much taller than before. The always-visible scrollbars shipped in 7e8eb4a set =min-height: 10px; min-width: 10px= on the slider (=settings/src/settings/gui.py=, the =.dupre-panel scrollbar slider= rule) — that's the floor for a short slider, and the trough itself is thin. Raise both the slider floor and the trough thickness so the bar is comfortably grabbable. Cosmetic × every glance at the wallpaper trays = P3 = [#C]. -** DONE [#C] Video wallpapers don't fit the desktop :bug:dotfiles:solo: -CLOSED: [2026-07-24 Fri] -From the roam inbox (Craig, claimed 2026-07-23): videos don't fit the desktop in desktop-settings. The video channel drives mpvpaper (=settings/src/settings/wallpaper.py=); mpvpaper passes options through to mpv, so the fit is a =--panscan=/=--video-unscaled=/keepaspect question rather than a layout one. Reproduce with a video whose aspect differs from the output, pick the mode that fills without distorting (cover, matching how the image channels behave), and cover it in the wallpaper tests. Minor severity × whenever the video channel is selected = P3 = [#C]. - -Shipped as dotfiles =04d1489=. =set_video= now passes =panscan=1.0=, so mpvpaper fills the output and crops the overflow instead of letterboxing; keepaspect stays on so nothing stretches. Tested against the mpvpaper arg log. -** DONE [#C] World-clock wallpaper arrangement :feature:dotfiles: -CLOSED: [2026-07-24 Fri] -Shipped 2026-07-24 as dotfiles =6afbe09=, iterated live with Craig. The grid of boxed mini-clocks became a centered vertical clock line: cities down a spine, west (Honolulu) top to east (Wellington) bottom, labels alternating both sides, no boxes. Each shows city / time (12h) / day+date / timezone region name ("US Central"). Day/night dimming + amber home carried over, title dropped, cursor restored over the desktop. Prototypes archived in archsetup 40216e7. The face is parameterized (=?layout=vertical|horizontal=, =?hour12=1|0=) so the panel pickers below can drive it. -** DONE [#C] Floating layout — should we? :feature:hyprland: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -From the roam inbox (Craig, claimed 2026-07-23): consider whether Hyprland should offer a floating layout — how it would work, the benefits, and the complexity. A brainstorm/spike, not a build: the deliverable is an assessment Craig reads and decides on, not a shipped layout. Not :solo:. When picked up, run it as a brainstorm — how a floating mode coexists with the current tiling binds (toggle keybind, per-workspace vs global, window-rule interactions), what it buys over the existing =togglefloating=, and the config/muscle-memory cost — then bring Craig the recommendation. - -CONCRETE PROPOSAL from a second roam item (Craig, 2026-07-24 via work) — "floating mode as the easiest mode": -- Can't select floating until at least one window is displayed. -- Entering floating freezes each window's position and floats it exactly where it is. -- During floating, drag windows with mod+mouse-drag. -- Exiting floating switches to tiling or monocle and lets that layout take over. -Craig's note: "simple, could be useful for different reasons." This is the design the brainstorm should evaluate first — assess feasibility against Hyprland's actual float/tile transitions (does freezing current geometry survive the tiling↔floating switch, does re-tiling on exit reflow cleanly) before recommending. - -ASSESSED, dotfiles =8cf4728=: =docs/2026-07-24-floating-layout-assessment.org=. Verdict: buildable and worth building on a capture-then-restore of window geometry (=hyprctl clients -j= gives at/size), which is a real gesture plain =togglefloating= can't express. Craig's four-rule proposal is folded in and each rule assessed. One taste call flagged (exit to previous layout vs always monocle). Ready to file a build task on Craig's go. -** DONE [#C] World clock wallpaper: bold the city names :feature:dotfiles:quick:solo: -CLOSED: [2026-07-24 Fri] +** TODO [#B] Proton static WireGuard profiles pass no traffic :chore:network: :PROPERTIES: -:LAST_REVIEWED: 2026-07-24 +:LAST_REVIEWED: 2026-09-09 :END: -From the roam inbox (Craig, 2026-07-24 via work): bold the city names on the world-clock wallpaper face (=settings/faces/world.html=, shipped =6afbe09=). Cosmetic × every glance at the world face = P3 = [#C]. Solo — a CSS weight change, screenshot-verifiable — but it's a visual call, so build it and show the render rather than close off a green suite. Pairs with the open world-face picker task. +wg-US-CA-144, wg-US-TX-714 and wg-NL-781 (all on wgpvpn) complete a WireGuard handshake and answer ICMP at 10.2.0.1, then forward nothing: no DNS on any transport, no HTTPS payload, no IPv6. The same account over the Proton CLI works, so the static configs are what Proton stopped honoring (the shape of an expired certificate on the profile). Diagnosed 2026-09-01. -Shipped as dotfiles =e63c0cf=. =.lbl .city= is now =font-weight:700=. Rendered offscreen and confirmed the bold reads well over the time/zone lines; home city stays amber. Comparison render was on ws5 for Craig. -** DONE [#C] Floating clock toggles on control+mod+c :feature:dotfiles:hyprland:solo: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -From the roam inbox (Craig, 2026-07-24 via work): a control+mod+c keychord should toggle the floating clock, the same as clicking the time waybar module. +The net doctor now names these as a dead tunnel and brings them down (dotfiles f56fd1a), which gets the machine back online but doesn't restore the tunnels. Two ways out: re-download the WireGuard configs from the Proton dashboard and re-import them (nmcli connection import type wireguard file ...), or drop the static profiles and use the Proton CLI only. Needs the Proton account, so not solo. -This answers the design question the round-6 clock-toggle fix deliberately left open (see the =clock toggle listener= DONE task above): =do_activate= calls =show_clock()= rather than =toggle()=, and the note there flagged "worth raising if he ever wants the spawn path to toggle too." He does. Build: a hyprland keybind bound to =clock toggle=, and confirm the toggle path (not show-only) fires whether the service is cold or warm. Solo — buildable and locally verifiable. +*** 2026-09-12 Sat @ 23:26:50 -0500 The NM tunnel DoT drop-in already covers wgpvpn +=/etc/NetworkManager/conf.d/tunnel-dns-over-tls.conf= (live on both daily +drivers since 2026-09-10/12, and now written by the installer) matches +=interface-name:wgpvpn= as well as =proton0=, so a re-imported static profile +gets per-link DNS over TLS turned off automatically. Their dead-forwarding +problem is separate and unchanged. -Shipped as dotfiles =e73a70e=. =bind = $mod CONTROL, C, exec, clock-panel toggle= reuses the exact command the time module's click runs, so it toggles identically. Registered clean on reload. Live keypress is Craig's to confirm. -** DONE [#C] Calculator scratchpad won't toggle closed on mod+x :bug:hyprland:solo: -CLOSED: [2026-07-24 Fri] +** TODO [#C] Declined dot-link-restore branch untested :test:network:dotfiles:solo:quick: :PROPERTIES: -:LAST_REVIEWED: 2026-07-24 +:LAST_REVIEWED: 2026-09-09 :END: -From the roam inbox (Craig, 2026-07-24 via work): =mod+x= opens the calculator scratchpad but doesn't close it — Craig has to kill the window by hand. A second =mod+x= should toggle it shut. Almost certainly a =togglespecialworkspace= vs plain =exec= binding in the hyprland config, or a scratchpad window-rule mismatch. Minor severity (a workaround exists: kill the window) × every time the calc scratchpad is used = P3 = [#C]. Solo — a keybind/window-rule fix, locally verifiable. - -Shipped as dotfiles =e73a70e=. New =calc-toggle= script (mirrors fuzzel-toggle: pgrep -x, pkill or launch), and =mod+X= now points at it, so a second press closes the calculator. 3 tests in tests/calc-toggle. -** DONE [#C] Saving and recalling window configurations :feature:hyprland: -CLOSED: [2026-07-24 Fri] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-24 -:END: -From the roam inbox (Craig, 2026-07-24 via work), a research idea: Craig wants to save a specific window+app arrangement and have it reappear on demand. What has to be known and built to make that happen — is there prior art (another WM or OS that does session/layout save-restore), what information do those need (app identity, geometry, workspace, launch command), and what are their rules. Explore how far Hyprland can get (hyprctl clients + dispatch, exec rules, window rules by class/title), document thoroughly, and review with Craig next time. Not :solo: — the deliverable is an assessment he reads and decides on, and it may spawn a build task once the shape is clear. Offer to file the build separately if part of it turns out urgent. - -RESEARCHED, dotfiles =8cf4728=: =docs/2026-07-24-window-config-save-recall-assessment.org=. Prior art surveyed (i3/sway =append_layout= swallow, KDE window rules, macOS Moom). Three tiers from cheapest: (1) reposition open windows — buildable + testable now; (2) relaunch + place by class rule; (3) full swallow-by-title, which hits the same-class ambiguity every tool hands back to the user. Recommends shipping tier 1; tiers 2-3 need Craig's call on how much manual disambiguation he'll accept. -** DONE [#C] Velox refresh sweep :chore:maint: -CLOSED: [2026-07-23 Thu] -From the roam inbox (Craig, claimed 2026-07-23): velox needs bringing up to date, the mouse/touchpad module is still there, investigate what else didn't move over. - -Resolved 2026-07-23 by a full sweep over tailscale. The touchpad module was already gone — velox's running waybar (started 01:05, after the reboot) and its tracked config both carry zero =custom/touchpad= entries; what Craig saw was the pre-restow waybar process from before the reboot, and the reboot cleared it. Sweep results: both machines at dotfiles f9b6404 (all three hyprland lock/exit fixes live on velox, config errors clean, =allow_session_lock_restore= reads true); stow restow clean, only the expected skip-worktree files; rulesets pulled to 50fc7ca and =make install= run (agent-text verified working by invoking it — an earlier "MISSING" reading was a PATH artifact of the non-interactive ssh shell, not a real gap); desktop-settings tick timer active; mpvpaper, power-profiles-daemon, gtk4-layer-shell, webkit2gtk all present. - -Genuine remaining differences, all per-machine installs rather than sync failures: =cmail-action=, =gcalcli=, and =playwright= aren't installed on velox, and =obsbot-wb-guard.service= isn't enabled there (the OBSBOT lives on ratio). None block anything; file separately if velox should send mail or drive browser tests. -** DONE [#C] Weather tooltip sunrise and sunset :feature:waybar:weather:quick:solo: -CLOSED: [2026-07-23 Thu] -Shipped 2026-07-23 as dotfiles =de62e9d=. The two rows sit directly below Humidity in the current-conditions block, rendered in the footer's 12-hour format (=%-I:%M %p=) so the tooltip reads one way throughout. - -Confirmed the no-extra-round-trip premise held: =sunrise,sunset= joined the existing =&daily== block. Split =forecast_url= and =reading_from= out of =fetch= so both the request and the reading are testable without network — that's what let the new cases cover a payload missing the fields. Six tests (Normal/Boundary/Error): row placement and format, a pre-change cache with no sun fields, an unparseable stamp, today's pair picked out of the six-day arrays, and the API omitting them. Reused the existing =_at= helper rather than adding a near-duplicate =_first=. - -Live-verified against the real API: sunrise 6:14 AM, sunset 7:59 PM for today in New Orleans, rendering in the actual tooltip. Full suite 3506 tests, both gates, exit 0. - -Open, not blocking: every other header row carries a glyph (thermometer, droplet, wind arrow) and the sun rows are plain text. The file's glyphs are marked font-confirmed codepoints, and I haven't verified a sunrise/sunset glyph renders rather than showing tofu, so I left them bare. Craig's call. - -From the roam inbox (Craig, claimed 2026-07-23): in the weather module's hover text, the section immediately after the location ends with the current humidity. Add the sunrise and sunset times for the current location directly below it. - -Cheap to source: the module already calls Open-Meteo with a =&daily== block (=common/.local/bin/weather=, the forecast URL around line 336), so =sunrise,sunset= joins that same request with no extra round trip — normalise_daily already parses the daily arrays. Times arrive as local ISO strings; render in Craig's canonical clock format rather than re-deriving one. The settings package's =suntimes.py= (pure NOAA math, no network) stays the offline fallback path if the API field is ever absent — don't duplicate its math here. -** DONE [#C] Maint doctor-row copy button :refactor:maint:quick:solo: -CLOSED: [2026-07-23 Thu] -Shipped 2026-07-23 as dotfiles =761fa5c=, "fix(maint): drop the COPY key from the doctor row" — the key and its orphaned handler removed from =maint/src/maint/gui.py=. =viewmodel.status_copy_text= stays: it's a tested pure serializer and the obvious source if a copy surface returns somewhere better placed. - -Correction to the body below: it describes a per-row button and a separate global one. There is only one COPY key, and it IS the global one Craig added in 8bc79ba two days earlier. He tried it and wanted it gone, so the row now reads DOCTOR · CLEAN UP · REVIEW & FIX. - -From the roam inbox (Craig, claimed 2026-07-23): remove the per-doctor-row copy button (next to REVIEW and FIX) from the maint status wall. The global COPY key (dotfiles 8bc79ba, "one global button copying rendered text") stays the one copy surface — the per-row button turned out to be clutter next to it. -** DONE [#C] WiFi tooltip signal strength :feature:waybar:network: -CLOSED: [2026-07-22 Wed] -From the roam inbox (Craig, claimed 2026-07-22): add signal strength to the WiFi tooltip. - -Resolved 2026-07-22: the tooltip's signal line existed but never fired on ratio — the mt7925 driver leaves /proc/net/wireless empty (legacy WEXT procfs unimplemented), so the dBm read returned None and the bar glyph fell to the weakest tier. Fix in dotfiles net/: an iw-dev-link nl80211 fallback (only spawns when procfs is empty), a signal_percent mapping, and an enriched line — Signal: ▂▄▆█ 100% · -32 dBm (excellent) — bars by band, percent, raw dBm, band word. The bar icon tier fixed itself as a side effect. -** DONE [#B] Desktop-settings dropdown panel :feature:waybar: -CLOSED: [2026-07-22 Wed] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-22 -:END: -Resolved 2026-07-22: shipped end to end via the "Build: desktop-settings panel" task (dotfiles 7a15237 → 9038eee; spec IMPLEMENTED, 85 suites + smoke 13/13 + e2e 17/17). Every open question below got settled in the spec: bar consolidation landed (74f723e), the wallpaper manager became the in-panel sub-view, and the format pickers split into their own sibling spec ([[file:docs/specs/2026-07-19-display-format-single-source-of-truth-spec.org]], DRAFT stub). Remaining human-eye checks live under "Manual testing and validation". - -Original body follows as the record. - -Initial spec written 2026-07-02: [[file:docs/specs/2026-07-02-desktop-settings-panel-spec.org]] (DRAFT — four decisions await Craig's review before build; architecture updated to the net panel's Blueprint/GTK4 stack). - -One waybar dropdown gathering the desktop toggles and sliders into a single settings panel, opened from a gear/settings glyph on the bar. Incorporate: -- *Auto-dim* toggle (the =custom/dim= feature just shipped — fold in here, or keep the standalone indicator and mirror it). -- *Brightness* slider (backlight, via brightnessctl). -- *Keyboard-backlight* brightness slider (brightnessctl on the kbd_backlight class). -- *Mouse* enable/disable toggle — shown only when a mouse is connected. -- *Trackpad* enable/disable toggle — shown only when a trackpad is connected (mirror =toggle-touchpad= / =touchpad-auto=). -- *Idle inhibitor* (the =custom/idle= module that replaced the built-in =idle_inhibitor= 2026-06-24 — toggles the hypridle daemon, state-synced icon). -- *Airplane mode* (the existing =airplane-mode= toggle; laptop-only). - -The conditional rows (mouse, trackpad, airplane) appear only when their hardware/context applies — reuse the laptop/device detection the airplane and touchpad indicators already do. - -Design / open questions (propose before building): -- Panel tech: sliders need a real toolkit (waybar can't host a slider), so a GTK4 + gtk4-layer-shell app like pocketbook is the likely shape. -- Which existing standalone bar modules (dim, touchpad, airplane, idle_inhibitor) collapse INTO this panel vs. stay on the bar as quick-access indicators. Craig's call. - -Implementation notes: a small GTK layer-shell app (mirror pocketbook's structure: src-layout Python package, pytest, Makefile) talking to brightnessctl / hyprctl / the touchpad + airplane helpers. Lives in the dotfiles repo or in-tree like pocketbook. TDD the backing toggle/slider logic. Sizable — worth a design doc first. - -Home handoff 2026-07-19 (inbox, resolving the open "few other things" decision — fold into the spec, close the open decision, extend the controls table, then run spec-review, may flip DRAFT→READY). Ownership: home drives the build (dotfiles settings/), archsetup keeps the canonical spec. Full reconciliation in home docs/design/2026-07-19-desktop-settings-module-brainstorm.org. -- ADD controls: night-light / color temperature; Do Not Disturb / notifications (dunst); lock / suspend quick actions; power profile (performance/balanced/saver); scenes/profiles — one control flipping several toggles at once (Focus, Presentation, Battery-saver, Night). Scenes are the payoff of consolidating everything. -- OUT (record reasons): volume / master-mute stays with the audio panel (no mirror here); theme light/dark goes to the theme-studio task. -- FORMAT PICKERS pulled to their own future sibling spec — time/date/weather format is out of THIS panel. Rationale: format settings live in many programs, so the design problem is a single source of truth for the canonical format. Track a future sibling-spec stub in docs/specs (time/date/weather format single-source-of-truth); Craig thinking it through separately, not started. -- STILL OPEN (spec already flags): wallpaper manager confirmed in scope, but row-that-opens-a-sub-view vs its own sub-spec undecided — resolve at spec-review. -** DONE [#C] Gallery probe: the fader-drag check is flaky :bug:test:design:quick:solo: -CLOSED: [2026-07-23 Thu] -Fixed 2026-07-23. Root cause confirmed rather than suspected: =panel-widget-gallery.html= line 74 sets =html{scroll-behavior:smooth}=, so =scrollIntoView= animates and the fixed 200ms sleep sometimes read =getBoundingClientRect= mid-scroll. The drag then dispatched at stale coordinates, the press missed the fader, and the check reported a dead widget. - -Fix: scroll with =behavior:'instant'=. The probe never needed the animation, so this removes the race instead of waiting it out. Also added a =settledRect= guard (rect stable across two reads AND on-screen) for zoom/column relayout, and a =hits()= assertion that the press actually lands on the fader before the drag goes out. - -Applied to the toggle-click check too — it shares the same fixed-sleep shape, and it failed for this exact reason during the diagnosis, so fixing only the fader would have left half the defect. - -Worth recording: my FIRST fix was wrong and made it worse. Polling until the rect stopped changing returned pre-scroll coordinates every time, because two identical samples are also what you get before the animation starts — an intermittent failure became a consistent one. The new hit-test assertion is what caught it, printing the press point at y=1326 against a 1200px window. That's the argument for asserting the press landed rather than only asserting the readout moved. - -Verified against the measured 1-in-6 failure rate: 8 consecutive runs, all three checks passing, with the press point identical every run (429,480) — deterministic, not lucky. Full probe 96 PASS, 0 FAIL, exit 0. - -=probe.mjs= check 3 ("fader drag tracks at 3x") intermittently reports =level 68 -> level 68=, i.e. the synthetic drag never registers. It has presumably been doing this all along unnoticed, since the suite is normally run once per batch. - -Grading: *Minor* severity (a false FAIL costs a re-run and a few minutes, and never ships a defect) x *most users, frequently* = P3 = =[#C]=. - -Frequency measured 2026-07-16, not estimated: 1 failure in 6 consecutive runs, having already fired twice in about fifteen that afternoon. The first grading guessed "some users, sometimes" (~1 in 10); at ~1 in 6, both people who run this suite hit it most sessions, so the row is "most users, frequently". The letter lands on =[#C]= either way, but the input was wrong and the matrix is only worth anything if its inputs are measured. - -Suspected cause: the check clicks the 3x size chip, calls =scrollIntoView=, waits a fixed 200ms, then reads =getBoundingClientRect= and dispatches the drag against those coordinates. If the zoom relayout or the smooth scroll hasn't settled, the rect is stale and the press lands off the fader — so the drag is a no-op and the readout never moves. The other timing-sensitive checks share the same fixed-sleep shape. - -*Do not fix this by raising the sleep.* That hides the race rather than removing it and leaves the check failing again on a slower run. Wait on the actual condition instead: poll until the rect stops changing between frames, or assert the press landed on the fader before dispatching the drag (the probes' own README already warns that a =find()= miss dispatches into nothing and reports as a widget bug). - -Why it matters beyond the annoyance: a gate that cries wolf gets its real failures ignored, and this suite is the only thing standing between the gallery and a silent regression. - -Recurrences: 2026-07-18 batch-6 gate (first run, passed 3 reruns); 2026-07-18 batch-9 gate (first cold run, =level 68 -> level 68=, passed 2 reruns); 2026-07-21 double-speedrun run (flashed one RED mid-run, passed on rerun). All were a session's first/early probe run — consistent with the stale-rect theory (cold-start relayout settling slower than the fixed 200ms sleep). -** DONE [#C] Dotfiles stow conflicts: first-launch risk + restow directory handling :bug:dotfiles:quick:solo: -CLOSED: [2026-07-23 Thu] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-14 -:END: -Closed 2026-07-23. The last open item was the velox check, and velox is reachable again. Pulled it from =02df01a= to =de62e9d= (clean tree, fast-forward), then ran =make conflicts hyprland=, which dry-runs every tier: "No stow conflicts", exit 0. Its old conflict copy had already cleared against the updated repo, so there was nothing for =make reset= to do. - -Verified the pull is live through the symlinks rather than just present in the repo: =~/.local/bin/weather= resolves into the dotfiles tree and returns today's sun times on velox. - -Note for the record: =make conflicts common= is not a valid invocation — =check-de= rejects it, because common and the host tier are auto-included in the DE-scoped run. =make conflicts hyprland= is the whole check. -*** 2026-07-14 Tue @ 00:51:51 -0500 Ratio calibre check passed; waypaper canonical decided (dark-lion) -Ratio's ~/.config/calibre is a directory symlink into the dotfiles repo (stow folded the whole dir), so the first-launch gap never existed there — check closed. Craig decided dark-lion.jpg is the canonical waypaper wallpaper; the repo config.ini updated from the that-one-up-there.jpg placeholder (the file is skip-worktree volatile, unskipped for the commit and re-flagged). Remaining: when velox is back online, run make conflicts / make reset there so its old conflict copy clears against the updated repo. -*** 2026-07-02 Thu @ 17:30:00 -0400 Shipped the Makefile hardening + first-launch guard (dotfiles 42a82d2) -The solo-able subset landed in the speedrun. =make conflicts <de>= is the loud first-launch guard: dry-runs all tiers, parses all four stow error shapes (plain file conflict, foreign symlink, dir-over-file, and restow's unstow_contents non-directory ERROR), lists each blocker with a directory/foreign-symlink marker, exits 1 when any exist. =make reset= now pre-clears the directory and foreign-symlink blockers =--adopt= aborts atomically on (removals printed; repo version wins per the target's contract), then adopts + git-checkouts as before. =make restow='s overwrite path switched rm -f → rm -rf so directory conflicts clear. 8 sandbox tests drive the real Makefile against a throwaway HOME (44 suites green). Also verified on velox: the whereami and mpd-playlists conflicts noted in this task were already hand-converted 2026-06-29 — =make conflicts hyprland= reports clean live. REMAINING (deferred per Craig's speedrun pre-flight): the waypaper canonical decision (live velox dark-lion.jpg vs repo that-one-up-there.jpg) and the ratio calibre-symlink check (ratio paused). -From the velox calibre incident (2026-06-27, note in ~/.dotfiles/inbox/processed/): calibre was launched before =make stow= ran, wrote its own default config into =~/.config/calibre/=, and silently blocked its own stow — it ran on factory defaults while the rest of common/ stowed fine. General pattern: any GUI app that auto-creates config on first run, launched before stow, blocks its own stow the same way. Velox was repaired by hand (=ln -srf= symlinks byte-identical to =stow --no-folding= output). - -Remaining work (re-graded C 2026-07-02 — the first-launch risk and the Makefile handling shipped in the speedrun; what's left is a paused-machine check): -- Waypaper canonical decision (Craig): RESOLVED 2026-07-14 — dark-lion.jpg is canonical (dotfiles fea3e93), repo config.ini updated off the that-one-up-there.jpg placeholder. -- Ratio check: RESOLVED 2026-07-14 — ratio's =~/.config/calibre= is a directory symlink into the repo (stow folded the dir), so the first-launch gap never existed there. -- When velox is back online: run =make conflicts= / =make reset= there so its old conflict copy clears against the updated repo. (velox carries a separate boot-recovery task; check once it's reachable.) -** DONE [#B] Weather tooltip caching :feature:waybar:weather:solo: -CLOSED: [2026-07-25 Sat 10:53] -From the roam inbox (Craig, claimed 2026-07-22): retrieve the weather tooltip data once per hour and cache it. If the network is unavailable, display the cached tooltip with explanatory text saying so. Dotfiles-side work (archsetup owns the lifecycle); touches common/.local/bin/weather. -Verified complete in the 2026-07-25 batch: the weather CLI already had the hourly default TTL, fresh-cache no-fetch path, stale fallback, and explicit offline footer. Its 33-test suite and the full dotfiles suite pass. -** DONE [#B] Settings gear becomes four device toggles :feature:waybar:dotfiles:solo: -CLOSED: [2026-07-25 Sat 10:53] -From the roam inbox (Craig, claimed 2026-07-23): the waybar gear should become four icons — touchpad, mouse, webcam, and a notification bubble. Clicking each toggles that setting directly. The first three turn red when disabled; the bubble turns red when DND is enabled. - -Today =custom/settings= (=hyprland/.config/waybar/config=) is one gear glyph () whose only job is =on-click: settings-panel=. The toggles themselves already exist and are tested — the settings package owns touchpad, mouse, and webcam (=webcam.py= is the USB-authorized kill switch from 2026-07-22), so this is a bar-side surface over existing backends rather than new capability. +repair_tunnel_dot_off (dotfiles net/src/net/repair.py) puts a tunnel link back to its DoT mode when turning DoT off didn't bring names back, and the evidence says "put back to <mode>" only when that restore succeeded. The restore-declined branch has no direct test. Give fake-resolvectl a second failure switch (FAKE_RESOLVECTL_DOT_RESTORE_FAIL) so the wording can be asserted absent as well as present. Follow-up from the f56fd1a review. -Note the state-polarity split when wiring the colors: three read "red = off" and DND reads "red = on". That asymmetry is deliberate (red means "something is disabled that normally isn't, or suppressed that normally isn't"), so encode it per-icon rather than deriving one rule. - -Decided 2026-07-23 (Craig): the gear STAYS alongside the four toggles as the panel launcher. So the bar's right side grows from 12 modules to 16 — the four toggles are net-new, the gear keeps its =on-click: settings-panel=. Open sub-question for build time, not blocking: whether the four toggles are four separate waybar modules or one custom module rendering four glyphs (fewer layout entries, one exec). Pick at build; the four-module shape is simplest and matches how mic/net already sit as individual modules. -Shipped in the 2026-07-25 batch as four independent JSON modules over the existing verified settings backends. Touchpad, mouse, and webcam turn terracotta when disabled; DND uses the deliberate inverse polarity; unavailable hardware dims. The gear remains the panel launcher. The live and Dupre theme CSS copies stay byte-identical. -** DONE [#C] Wallpaper panel selection and scroll state :feature:dotfiles:solo: -CLOSED: [2026-07-25 Sat 10:53] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-25 -:END: -From the roam inbox (Craig, 2026-07-25). Screenshot: =~/pictures/screenshots/2026-07-25_013041.png=. Three related behaviors in the settings wallpaper panel (=settings/src/settings/wallpaper.py=): -1. Open at the wallpaper currently displayed, not the top of the list. -2. Highlight that wallpaper as selected in the scrollable pane while it shows in the preview. -3. Keep the scroll position when a picture is selected. Today selecting a picture snaps the scroll back to the top, which is the bug half of this. -Grade: minor scroll-reset defect x every panel selection = P3 = [#C]; the open-at-current and select-current behaviors are enhancements at the same level. One type tag, so filed =:feature:= with the scroll-reset called out as the bug. Solo: buildable in the settings GTK panel, agent-verifiable via headless capture plus the wallpaper.py tests, no design call — swww query gives the current wallpaper, and scroll-position preservation and row selection are standard GTK. -Shipped in the 2026-07-25 batch. The panel queries =awww query= off the UI thread, prefers the actually displayed image over stale stored state, highlights it, scrolls it into view on first open, and remembers the horizontal adjustment across selection-triggered rebuilds. -** DONE [#C] Net tooltip IPs and line order :feature:waybar:network:solo: -CLOSED: [2026-07-25 Sat 10:53] -From the roam inbox (Craig, claimed 2026-07-23): in the wifi hover, add the internal IP, external IP, and gateway IP just below the Interface line; move the Signal line to just above the keyboard-shortcuts line. Design constraint: the bar's hot path does no network I/O (status.py deliberately skips _address_facts on the 2s beat) — internal IP + gateway can ride cheap local reads, but the external IP must come from a cache the connectivity probe refreshes, never a live lookup in waybar-net. -Shipped in the 2026-07-25 batch. The slow connectivity probe caches local addressing and a validated external IP with the network identity; the Waybar hot path only reads that valid cache. Tooltip order is Interface, internal/external/gateway IPs, connectivity detail, throughput, Signal, shortcut. -** DONE [#B] Dupre Kit merge — casting additions :feature:tooling:solo: -CLOSED: [2026-07-25 Sat 10:53] -Fold docs/prototypes/dupre-kit-additions.js back into the kit proper: detentFader (NEW — multi-detent slide attenuator with speedbump drag physics: magnet + escape hysteresis, parked tick glow) and the drumRoller redefinition (UPGRADE — 1..N channels and min/max range; stock hardcodes two drums and throws on one, defaults reproduce stock exactly) and the guardedToggle redefinition (UPGRADE — lever throws with rotateX so it flips toward the viewer instead of the stock 180° planar spin that sweeps sideways mid-transition; contract unchanged). Merge means: builders into widgets.js, the additions CSS into DUPRE_CSS, additions-scoped gradients into the shared defs plate, gallery cards for both in panel-widget-gallery.html, and POLICY entries. Origin: the desktop-settings casting sitting 2026-07-21 — Craig's direction is that components get finished by being needed ("the ones needed most will have had the most attention"), so more additions may accrue here before the merge; batch them. -Shipped in the 2026-07-25 batch. =widgets.js= now owns all three builders, shared gradients/CSS, contracts, and policy records; additions no longer redefines them when older casting pages load it. The gallery has a three-detent fader card and a three-channel 0–12 drum demonstration (112 cards total). Static ownership tests, JS syntax checks, and the complete headless interaction probe pass. -** DONE [#C] Maint live-refresh hairline replacement :feature:maint:solo: -CLOSED: [2026-07-25 Sat 10:53] -:PROPERTIES: -:LAST_REVIEWED: 2026-07-14 -:END: -From the roam inbox (routed 2026-07-13): the memory-killer section seemed to update too often, and "it's a bit unclear what the line is doing; consider something else." Diagnosis (2026-07-14): the data cadence is already the requested 3s (gui live tier, _LIVE_SECONDS); the perceived churn is the live-refresh hairline — the 2px bar under the live sections that drains full-to-empty over each 3s window, redrawn at 150ms (gui._hair_tick, viewmodel.refresh_fraction). It exists to tell a stale board from a frozen one (2026-07-09), but it reads as constant unexplained motion. Design call for Craig: replace the draining line with something whose meaning is legible — candidates: a dot that blinks once per refresh, a "3s" age caption that only appears when refresh is overdue, slowing the drain redraw, or dropping the indicator on live tiers and keeping it only when data goes stale. Keep the stale-vs-frozen distinguishability that motivated the hairline. -*** 2026-07-21 Tue @ 08:35:00 -0500 Decided (Craig): silent-until-stale age caption -Replace the draining 2px hairline with an age caption that shows ONLY when refresh is overdue (e.g. "3s", "8s" once past the expected window) and shows nothing while the board is healthy. This keeps the stale-vs-frozen signal — a frozen board surfaces a growing age number, a live one stays clean — while removing the constant motion the hairline created. Implementation (dotfiles, archsetup-owned): drop =gui._hair_tick= / the hairline draw, add an overdue-age caption driven off =viewmodel.refresh_fraction= (or the last-refresh timestamp) rendered only past the live window. Now unblocked; needs a live visual check on the panel after. -Shipped in the 2026-07-25 batch. The animated draw area and 150ms timer are gone; the memory section header stays silent through the healthy three-second window, then shows a once-per-second growing age caption. Pure boundary tests and the full maintenance suite pass. -** DONE [#D] Test-framework + prototype refactor cluster :refactor:solo: -CLOSED: [2026-07-25 Sat 10:53] -Grading: no behavior change; parking lot. Refactors from the S5-S7 audit, distinct from the installer refactor rollup above. -scripts/testing/run-test.sh + run-test-baremetal.sh duplicate the run/poll/report skeleton and have drifted (VM uses setsid + copy helpers, baremetal uses nohup + hand-rolled sshpass scp) — extract the shared core so baremetal inherits the sturdier paths; run-maint-nspawn.sh:66 + run-maint-scenarios.sh:78 duplicate the transport-independent _scenario_var/_validate_scenario/run_scenario (a sourced lib/maint-scenario.sh); run-test.sh:251,265 uses two different mechanisms (pgrep vs ps|grep) for the same liveness check; docs/prototypes/gen_tokens.py:78 repeats the section-iteration skeleton across four emitters; gallery-widget.el:95,136 hardcodes SVG arc/hub path strings that duplicate the cx/cy/radius geometry (dial desyncs silently on a constant change); gallery-widget.el:72,84 leans on the private svg--append. See findings doc (S5, S6, S7). -Completed test-first in the 2026-07-25 batch. QEMU and bare-metal runners share liveness/report helpers; maintenance transports share scenario validation/execution; token emitters share ordered section traversal; and the Emacs SVG gauge shares semicircle geometry and uses the public DOM append API. Every fast Python/ERT suite passes. -** DONE [#B] Two agent sessions sharing one git repo :chore:tooling: -CLOSED: [2026-07-26 Sun] -Craig approved the shared-rules-layer solution on 2026-07-26. - -Use one repository-scoped publish lock for every session and worktree sharing a clone. Derive the lock name from the real Git common-directory path; hold it across reconcile, stage, staged review, and commit; track the owning session and reviewed staged-tree fingerprint; refresh it after conversational waits; and repeat the staged review if ownership or the fingerprint changed. Ordinary working-tree edits remain concurrent. - -An approval waiver never waives the staged review, because that review is the gate that reads the actual hunks entering the commit. Rulesets owns the implementation in =commits.md=, =agent-lock=, and its Bats coverage; archsetup sent the approved implementation package through the rulesets inbox. -** DONE [#A] Reboot ratio to activate amdgpu.runpm=0 :bug:hyprland:ratio: -CLOSED: [2026-07-28 Tue] DEADLINE: <2026-07-28 Tue> -:PROPERTIES: -:CREATED: [2026-07-28 Tue] -:LAST_REVIEWED: 2026-07-28 -:END: -Craig's plan: close everything down, run topgrade, then reboot. Alarm set for 08:00 (=at= job 56, persistent desktop notify). - -=amdgpu.runpm=0= sits in =/etc/default/grub= and in the generated =/boot/grub/grub.cfg= (5 occurrences, so the reboot will actually apply it) but is absent from =/proc/cmdline=. The box has been up since 2026-07-22 21:10 and the fix landed 2026-07-24, so the running kernel predates it. The GPU is AMD Strix Halo (Radeon 8060S, =1002:1586=), exactly what the parameter targets: runtime power management invalidates the GPU resources hyprlock holds across a display power-cycle, so hyprlock exits without unlocking. - -That is the root cause under the 2026-07-27 lockdead screen. The screen-lock flock fix (dotfiles =ec18fd7=) stops one dead client from becoming a lockdead screen, but it treats the symptom -- this reboot treats the cause. - -Not :solo: — Craig closes his own session and runs topgrade first. - -Rebooted 2026-07-28 08:59. =amdgpu.runpm=0= confirmed present in =/proc/cmdline= afterward, so the parameter is finally live. - -Correction, 2026-07-29: the claim above and in the body that this is "the root cause under the 2026-07-27 lockdead screen" is wrong, and superseded. hyprlock was never crashing. Every logged exit is =rc=143=, SIGTERM, from =settings-watch= killing it by design. See =[#B] Night watch and the lock watchdog fight each other=. The reboot was still worth doing (the parameter is a genuine mitigation for a real AMD defect) but it did not fix this, and the lockdead screens continued after it. -** DONE [#B] Caffeine state is unreadable on both surfaces :bug:dotfiles:design:solo: -CLOSED: [2026-07-28 Tue] -:PROPERTIES: -:CREATED: [2026-07-28 Tue] -:LAST_REVIEWED: 2026-07-28 -:END: -Neither surface that reports caffeine tells the truth reliably, so there is no way to know at a glance whether the screen will lock. Found while investigating the 2026-07-27 lockout, where Craig believed caffeine was on and the screen locked anyway. - -Defect 1 — the settings panel shows a frozen value. =gui.py= calls =_refresh_async()= once during window construction (line 454) and again only after the user's own actions (=_after_matrix=, line 663). The only two =GLib.timeout_add= calls are one-shots (the 2400ms toast hide and a 350ms fire), so nothing re-reads state on a timer. An open panel therefore displays the caffeine value from the moment it opened, forever. Any external flip -- the waybar click, Super+I, the =caffeine-toggle= script -- leaves it stale with no self-correction. The panel is the only surface in the repo carrying a caffeine control (=panel.py:25=); maint has none and does not embed these toggles, so this is the display Craig read. - -Defect 2 — there is no caffeine indicator on the bar at all. =custom/caffeine= appears in neither the stowed =hyprland/.config/waybar/config= nor the live generated =/run/user/1000/waybar/config=, and no =custom/caffeine= block is defined anywhere in the waybar config dir. The =waybar-caffeine= script exists, works, and has its own passing test suite, but nothing displays it. So the bar has never been a source of caffeine state, and the keybind and script have been signalling (=pkill -RTMIN+8 waybar=) a module that isn't there. - -(An earlier read of this task said the bar showed two near-identical glyphs. That was wrong: the module is absent, not merely unstyled. The script's class names are still backwards -- =active= when caffeine is OFF, =inhibited= when ON -- and neither class is styled, but both points are moot until the module is actually in the bar.) - -Grading: Major severity (the panel reports state wrongly while it is open, and the only other surface does not exist, so there is no reliable source for a setting Craig actively manages) x most-of-the-time (any external toggle while the panel is open; the bar never shows it) = P2 = [#B]. - -Fix all three. Wire =custom/caffeine= into the bar, rename its classes so they describe caffeine rather than idle, and style them from the existing palette. Give the panel's toggle row a re-read on a timer or on focus-in. Solo -- buildable and testable, and the direction is settled by the defects rather than a taste call, though the bar color is worth a glance from Craig once it renders. - -All three shipped as dotfiles =033076c=, pushed. =custom/caffeine= now sits in the bar between DND and settings on =interval: 2=; classes renamed =on=/=off= and both styled, caffeine-ON in the theme's gold =#dab53d=; the panel re-reads live state every 3s while visible. Verified live in the stowed config and the generated =/run/user/1000/waybar/config=. The full suite caught a theme-copy regression (=themes/dupre/waybar.css= out of sync with =waybar/style.css=) that the focused suites missed. -** DONE [#B] hyprlock still exits mid-lock; the watchdog relaunch is silent :bug:hyprland:dotfiles: -CLOSED: [2026-07-29 Wed] -:PROPERTIES: -:CREATED: [2026-07-28 Tue] -:LAST_REVIEWED: 2026-07-28 -:END: -Craig, 2026-07-28 ~15:00: saw the Hyprland lockdead/error text blurred *behind* a working lock screen; it vanished when he authenticated. - -That ordering is the diagnosis. hyprlock's blur samples what the compositor is currently rendering, so the compositor was already showing lockdead when the new hyprlock attached. Sequence: hyprlock exits non-zero (no coredump, so it exits rather than crashing), Hyprland renders lockdead because the client is gone while the session stays locked, =screen-lock='s watchdog relaunches within =LOCK_RELAUNCH_DELAY= (0.5s), and the new client draws over the lockdead frame and blurs it. - -*The recovery worked.* On 2026-07-27 this same hyprlock exit produced two contending clients and a session recoverable only from another console. It now self-heals in half a second, and the residue is cosmetic. Both the flock guard (dotfiles =ec18fd7=) and the watchdog did their jobs — verified in this session's compositor log, where all four lock events created exactly one =sessionLock= and one =sessionLockSurface= each, against two of each on 2026-07-27. - -Two things remain. - -*Why hyprlock exits.* The wrapper's header blames GPU-resource invalidation across a display power-cycle (hyprlock#953), which =amdgpu.runpm=0= targets — and that parameter is live as of the 2026-07-28 08:59 reboot, confirmed in =/proc/cmdline=. There is also no DPMS idle rule any more (=e900903=), so idling never power-cycles the display. Yet hyprlock still exited. Strongest untested candidate: a screen recording (=wf-recorder= into =~/sync/recordings/2026-07-28-12-53-57.mkv=, running 12:53 until Craig killed it) held screencopy sessions on DP-4 across the lock. The compositor log carries 2454 screenshare sessions and a =CScreencopyProtocol= bind in the window between the last two locks. A screencopy client churning dmabufs alongside hyprlock's own is a plausible way to invalidate them, and it was the one large new variable that day. - -*The relaunch is silent.* The watchdog loop re-runs hyprlock and logs nothing, so there is no record of how often this fires, when, or with what exit code — which is exactly why the frequency couldn't be established from the logs. Log the exit code and a timestamp on each relaunch. - -Grading: Major severity (the lock client dies mid-lock, and the pre-fix version of this wedged a session unrecoverably) x most users frequently (twice in three days, and this is a single-user machine, so every occurrence lands on the only user) = P2 = [#B]. Downgraded from the 2026-07-27 [#A] because the wedge is fixed and the failure now self-heals. - -An earlier draft of this grading said "some users sometimes", which the matrix maps to P3 = [#C], not the [#B] written beside it. The frequency row was the wrong input rather than the letter: on a one-user machine a fault hitting twice in three days is frequent, not occasional. Corrected the input per the rule that a disputed grade is fixed at its inputs. - -Solo for the instrumentation half only: adding the relaunch logging is buildable, testable against the existing =tests/screen-lock= suite, and needs no decision. Diagnosing the exit is not solo — it needs a reproduction, and the likely trigger is Craig recording his screen. - -Next step when picked up: land the relaunch logging first so the next occurrence produces evidence, then try to reproduce by locking with =wf-recorder= running. - -Superseded 2026-07-29 by =[#A] Night watch and the lock watchdog fight each other=. The logging landed (dotfiles =5bbe2c3=) and answered it within hours: three =rc=143= entries, SIGTERM, from =settings-watch= killing hyprlock by design. Nothing was crashing, so both the AMD-iGPU and the screen-recorder hypotheses in this task are wrong. Kept closed rather than deleted because the reasoning that led here is worth the record. -** DONE [#A] Idle commits silently drop the screen-lock wrapper :bug:hyprland:dotfiles:security: -CLOSED: [2026-08-04 Tue] DEADLINE: <2026-07-29 Wed> +* Archsetup Resolved +** DONE [#B] Function keys issue media actions instead of F-keys :bug:velox: +CLOSED: [2026-09-01 Tue] +:PROPERTIES: +:CREATED: [2026-08-19 Wed] +:LAST_REVIEWED: 2026-08-19 +:END: + +From the roam inbox, Craig's words: "function keys should issue F+number +functionality rather than their media functionality when the button is hit. +currently it's reversed and I have to hit function and the f button for F+number +functionality." + +Check first whether this belongs to archsetup at all. On a Framework the Fn-lock +is a firmware-level toggle held in the keyboard itself (Fn+Esc on most +revisions), not something the OS sets, in which case this is one keystroke +rather than a change here. If it is instead a hid/keyboard-module quirk, it is +ours. + +Grading: Minor severity (the keys work, they are on the wrong layer, and there +is a workaround) x every user every time (every F-key press) = P2 = [#B]. + +Resolved 2026-09-01: not ours, as the body suspected. The Fn layer is decided +in the EC (the keyboard reaches Linux as a plain AT keyboard on i8042), so no +OS-side knob exists. One keystroke: Fn+Esc toggles Fn Lock; Craig confirmed +F1-F12 now send F-keys by default. The EC holds the state across reboots; it +reverts only if the EC loses power (battery disconnect or mainboard reset), +which is likely why it flipped around the August reinstall. +** DONE [#A] Lock-screen clock stale after a real sleep :bug:hyprland:dotfiles:velox: +CLOSED: [2026-09-13 Sun] SCHEDULED: <2026-08-25 Tue> +:PROPERTIES: +:CREATED: [2026-08-25 Tue] +:LAST_REVIEWED: 2026-08-25 +:END: +After waking velox from a real sleep, the hyprlock clock shows a stale time +(Craig confirmed 2026-08-24: the wake-from-sleep case, not an idle-locked +screen). Three isolated tests on 2026-08-24 failed to reproduce it — hyprlock +0.9.6 repainted within a second of a display power-cycle, a three-minute +SIGSTOP, and both together with the screenshot background — so it needs a real +suspend on the real hardware. + +Grading: Minor severity (cosmetic-to-confusing, the screen still unlocks) × +most users frequently (every wake) = P3 = [#C] by the matrix; held at [#A] at +Craig's direction on 2026-08-25 so it gets run while velox is the daily driver +on the road. Revisit the letter once the manual check has an answer. + +Not :solo: — the distinguishing observation is Craig's. The check lives under +Manual testing and validation: "Lock screen after a real sleep: is the clock +frozen, or is all of hyprlock frozen?". Its three outcomes each name a +different fix: stale-then-corrects → repaint interval; frozen with live input +→ clock rendering; frozen with dead input → hyprlock hung, a crash/hang +recovery bug the =screen-lock= watchdog doesn't cover. + +*** 2026-09-13 Sun @ 07:57:04 -0500 Closed: fixed, per Craig +Craig confirmed on 2026-09-13 Sun that the stale clock after a real sleep is fixed on +velox. I could not identify the commit from here (nothing in the dotfiles +or archsetup log since 2026-08-24 names hyprlock, the lock clock, sleep or +resume), so this closes on his report rather than a cited change; the +manual-testing check for it is retired with it. +** DONE [#A] Velox reinstall — DR test of archangel + archsetup :velox:chore: +CLOSED: [2026-09-13 Sun] DEADLINE: <2026-08-15 Sat> :PROPERTIES: -:CREATED: [2026-07-29 Wed] -:LAST_REVIEWED: 2026-07-29 +:CREATED: [2026-08-13 Thu] +:LAST_REVIEWED: 2026-08-13 :END: -Caught live 2026-07-29 05:30, seconds after it happened, while verifying that Craig's watch-stage change had landed. - -=idle.py= renders the *whole* hypridle.conf, including a hardcoded =GENERAL= block. That block said =lock_cmd = pidof hyprlock || hyprlock=. The live config said =|| screen-lock=. So every idle-stage commit through the panel rewrote =lock_cmd= and dropped the wrapper out of the chain. - -The wrapper is not incidental. It carries the flock duplicate guard (the fix for the 2026-07-27 unrecoverable wedge), the crash-relaunch watchdog, and the relaunch log. Parking one stage removed all three in a single write, and nothing said so. - -=tests/settings/test_settings.py:562= asserted the bare =|| hyprlock= form, so the suite *enforced* the regression. That is why 3845 tests stayed green through a day of work on exactly this subsystem. A test can pin the bug as readily as the fix. - -The false-negative this sets up is worth naming: with the wrapper gone the relaunch log stops receiving entries, and an empty log reads as "the problem is fixed" when it means "the instrument was removed". The =screen-lock= header already warns that an empty file is not proof; this is the mechanism that would have produced one. - -Fixed by dotfiles =ab059fb= (2026-07-29 05:59). Verified 2026-08-04 against the tree rather than the commit message: =idle.py:47= renders =lock_cmd = pidof hyprlock || screen-lock || hyprlock=, the live =hypridle.conf= matches, and =test_settings.py= now asserts the wrapper is in the chain plus a second test for the bare-hyprlock fallback. The test that used to pin the bug now pins the fix. - -The evidence that matters is the one this task named: =~/.local/var/log/screen-lock.log= is *receiving entries*, so the instrument is present. An empty log was the false negative to fear, and it did not happen. - -FIXED here, TDD, in the working tree pending commit: -- =idle.py= =GENERAL= now names =screen-lock=, with a comment saying why the line is load-bearing. -- The test now pins the wrapper form. Red first against the old template. -- Live config rewritten through the panel's own path and hypridle restarted; =lock_cmd= confirmed back to =screen-lock=, one hypridle running. - -Grading: Critical severity (=write_conf= truncates, so any hypridle key the renderer does not model is silently deleted rather than preserved — that is configuration data loss, and the =lock_cmd= case proved it happens in the field) x some users sometimes (only when an idle stage is committed, which is rare) = P2 = [#B]. - -An earlier draft graded this [#A] on a "security carve-out". That was wrong: disarming the guard is an availability problem, not a leak, and the carve-out is for privacy, security, compliance and safety. The severity band is what carries the weight here, and silent deletion of configuration is the =Critical= band's data-loss case. - -Two further fixes came out of an independent review of the first one: - -- *Fail-open restored.* =pidof hyprlock || screen-lock= made the wrapper the end of the chain, and =screen-lock= is a stow symlink in =~/.local/bin=, not a system binary. An unstowed tree, or a hypridle started without =~/.local/bin= on PATH, resolves it to 127 — so the screen would never lock *at all*. That is worse than the duplicate client the wrapper prevents. The chain now ends =|| hyprlock=, matching the wrapper's own fail-open discipline. -- *The file now says it is generated.* Three comment lines at the top of the rendered output name the renderer and warn that edits are overwritten. The absence of that header is how the divergence survived unnoticed. - -Still open, and why this stays a task rather than closing with the fixes: the header warns, but nothing *prevents* the next divergence, and the exposure is wider than =lock_cmd= alone. The review enumerated it: - -- =before_sleep_cmd= and =after_sleep_cmd= sit in the same hardcoded block, at identical risk. -- Every stage command is hardcoded in =_stage_commands= (brightness level, lock, watch, dpms, suspend), same one-way overwrite. -- =write_conf= *truncates* rather than merges, so any hypridle key the renderer does not know about (=ignore_dbus_inhibit=, =ignore_systemd_inhibit=, =inhibit_sleep=, =on-lock=, =on-unlock=) is deleted rather than preserved. That is the largest hole: a key nobody has added yet would vanish the first time a stage is parked. +Mainboard swapped Intel→AMD (Ryzen AI 9 HX 370); new NVRAM has no boot entry. +Decision: full reinstall via archangel+archsetup, run deliberately as a +disaster-recovery drill before the Sunday flight. Runbook (live checklist): +[[file:docs/2026-08-13-velox-reinstall-runbook.org][docs/2026-08-13-velox-reinstall-runbook.org]] +Done 2026-08-13: ISO rebuilt (archangel-2026-08-13, archsetup baked with AMD +microcode detection, velox profiles at /root/, .ai/inbox excluded — build.sh +edits pending commit in archangel), contents verified, dotfiles swept clean of +Intel assumptions. +Finding folded in: velox's truenas backups silently stopped ~Jul 6 (newest is +DAILY.0 Jul 6; wolf.conf.gpg from Jul 29 is in NO backup). Salvage pass in the +runbook is therefore REQUIRED before partitioning, and the fresh install must +fix + verify the backup timer (runbook Phase 5). -Options: have the renderer preserve the existing general block and unknown keys instead of emitting its own, or accept the template as the single source and move every hypridle setting into the panel. A design call for Craig, and the truncation half is the part that will bite next. +*** 2026-09-13 Sun @ 07:16:27 -0500 Applied the two live convergence steps on velox +Ratio got both by hand on 2026-09-12 while velox was off the tailnet; velox +came back on 2026-09-13 and got them over ssh: =systemctl disable --now +wsdd.service= (no Samba host to advertise; 43acf51 stops the installer +enabling it) and =systemctl mask passim.service= (the unit is static, so the +09-12 disable was a no-op; 38b1758 masks it in the installer, but +supplemental_software is a completed step there and doesn't re-run). +Verified after: wsdd inactive/disabled, passim inactive/masked, zero +listeners on 5357 and 27500. Same pass fast-forwarded velox's dotfiles to +f56fd1a and archsetup to dc00a62, and confirmed the headless Proton Bridge +service is still disabled there. + +*** 2026-09-13 Sun @ 07:56:37 -0500 Closed the drill and filed its working-dir artifacts +The reinstall itself finished on 2026-08-14; every finding it surfaced is +its own task, and the live convergence steps landed on 2026-09-13, so nothing +was left in this task but the record. Filed per the working-files +convention: the runbook to docs/2026-08-13-velox-reinstall-runbook.org (with +an Outcome section, since the checklist was never ticked during the run), +the boot-entry reference to docs/2026-08-15-velox-uefi-boot-entry-reference.org, +the three gap reports to docs/design/2026-08-14-velox-reinstall-gaps-1 to 3, +and the wttrin bundle to working/emacs-wttrin-rescue/ under the task that +owns it. working/velox-reinstall/ is gone and every inbound link repointed. +** DONE [#B] Land the rescued emacs-wttrin commit :chore:velox: +CLOSED: [2026-09-13 Sun] +:PROPERTIES: +:CREATED: [2026-08-14 Fri] +:LAST_REVIEWED: 2026-08-17 +:END: +bf0457f "feat: add wttrin-hide-follow-line to hide the wttr.in follow line" +(2026-06-24) was the only genuinely unpushed commit anywhere on the old +velox — 3 files, 103 insertions, with a test file. Rescued as a verified +git bundle before the disk was wiped (wttrin-bf0457f.bundle, deleted on +2026-09-13 once the commit was on the remote): +To land it: clone emacs-wttrin, =git fetch <bundle> --branches=, review the +commit, then push to git@cjennings.net:emacs-wttrin.git. Delete the bundle +once it's on the remote. + +*** 2026-08-17 Mon @ 19:57:42 -0700 Re-checked: still unlanded, and the bundle is still the only copy +Cloned the remote bare and asked it for the object directly: =git cat-file -t +bf0457f= returns "Not a valid object name", so the commit has never reached +=git@cjennings.net:emacs-wttrin.git=. Remote =main= is =ee8fdeb=. + +That makes =working/emacs-wttrin-rescue/wttrin-bf0457f.bundle= (moved there 2026-09-13 when the reinstall task filed its artifacts) the sole surviving +copy of 103 insertions across three files, on one laptop that is travelling. +Worth doing sooner than its =[#B]= suggests for that reason alone. It also +pinned the reinstall working directory open until 2026-09-13, when the bundle +moved into its own working dir here and the reinstall task closed. + +*** 2026-09-13 Sun @ 09:47:20 -0500 Landed on emacs-wttrin release/0.4.0 and deleted the bundle +Cherry-picked unchanged onto release/0.4.0 (which already contains main; the +remote's default branch is still main, so main doesn't carry it yet) as +cb70193, patch-id identical to bf0457f. Review turned up two +defects the rescued commit carried, both fixed test-first and pushed with it: +f9449f6 makes the existing-F check case-sensitive (case-fold-search defaults +to t, so a lowercase f read as F), and 9c8d23e keys the cache on the effective +display options (toggling the setting kept serving a cached buffer with the +Follow line, against the 944a52f rule). Pushed 37e1c94..9c8d23e; full suite +green at 73 files. + +Two things checked before landing: an old emacs-wttrin note that F broke ANSI +colour is stale (wttr.in sends identical colour codes with and without F), and +the suite's smoke failure was only missing Emacs 31.1 eask deps. The bundle's +main and release/0.4.0 heads are on the remote as-is, and its only other commit +(bf0457f) is there as the patch-identical cb70193, so the bundle and its +working dir are deleted. +emacs-wttrin has a handoff note in its inbox covering all of it. +** DONE [#A] Pre-vacation fix list — morning review +CLOSED: [2026-09-17 Thu] +:PROPERTIES: +:LAST_REVIEWED: 2026-09-17 +:END: +Reviewed 2026-08-08; the decisions it drew are recorded in the tasks below. +Closed at the 2026-09-17 review because the ~08-15 departure it ranked work for +has passed. Where each of the seven items went: + +1. Velox reliability → the [#A] sleep/suspend task. +2. Velox machine health for travel → moot: velox was wiped and reinstalled on + 2026-08-13. +3. Remote access from outside the LAN → the wolf WireGuard rider under the + sleep/suspend task. Tailscale to ratio works off-LAN (checked 2026-09-17); + truenas and truenas-kvm weren't re-checked. +4. cgit secrets audit → the [#A] audit task and its rotation VERIFY. +5. Osbot camera → its own task; the podman socket and camera udev rule are + installed (sleep/suspend task, 08-17 entry). +6. Hotspot/metered WiFi and network ordering → the held design calls under + Next Session Focus. +7. The orchestrator sequence-pin gap → filed 2026-09-17 as [#C] Orchestrator + sequence pin misses an added step. diff --git a/working/clock-display-references/2026-07-30-maeda-cosmos-notes.org b/working/clock-display-references/2026-07-30-maeda-cosmos-notes.org new file mode 100644 index 0000000..04d8629 --- /dev/null +++ b/working/clock-display-references/2026-07-30-maeda-cosmos-notes.org @@ -0,0 +1,11 @@ +#+TITLE: Companion to the Line applet sent a few minutes ago: a copy +#+SOURCE: from website +#+DATE: 2026-07-30 18:48:00 -0500 + +Companion to the Line applet sent a few minutes ago: a copy of Cosmos (C1, 1995), the orbits piece — 'numbers in orbit'. Delivered as 2026-07-30-maeda-cosmos-standalone.html in this inbox. + +Same shape as the Line copy: the <maeda-cosmos> web component plus its model/step/render modules esbuild-bundled inline as an IIFE, so it opens straight from file:// with no server and no network. Verified in Chrome from file:// — canvas 400x300, the month's days riding the rotated ellipse with today red on the sweep hand, seconds odometer live at the right edge, zero console errors. + +Controls: hover reshapes every orbit; press-drag pulls a new loop out of the system; the mark at bottom-left clears them and replays the fly-in credit; click the canvas then ESC toggles strip-move mode. Note this one has no readout event, unlike Line. + +Source of truth stays ~/code/maeda-tribute (src/pieces/cosmos/); regenerate rather than edit. Reference only — no action needed. diff --git a/working/clock-display-references/2026-07-30-maeda-cosmos-standalone.html b/working/clock-display-references/2026-07-30-maeda-cosmos-standalone.html new file mode 100644 index 0000000..bfa9fbb --- /dev/null +++ b/working/clock-display-references/2026-07-30-maeda-cosmos-standalone.html @@ -0,0 +1,422 @@ +<!doctype html> +<html lang="en"> +<head> +<meta charset="utf-8"> +<meta name="viewport" content="width=device-width, initial-scale=1"> +<title>Cosmos (C1, 1995) — Maeda × Shiseido tribute</title> +<style> +:root{ + --ground:#151311; + --panel:#100f0f; + --well:#0a0c0d; + --raise:#1a1917; + --silver:#bfc4d0; + --cream:#f3e7c5; + --steel:#969385; + --dim:#7c838a; + --wash:#2c2f32; + --gold:#e2a038; + --mono:"BerkeleyMono Nerd Font","Berkeley Mono",monospace; +} +*{box-sizing:border-box;margin:0;padding:0} +html{background:var(--ground);font-size:112%} +body{font-family:var(--mono);color:var(--silver);padding:2.4rem 2rem 4rem;line-height:1.45; + background:radial-gradient(1200px 600px at 70% -10%,#1c1915 0%,transparent 60%),var(--ground)} +.wrap{max-width:820px;margin:0 auto} +.eyebrow{color:var(--steel);font-size:.72rem;letter-spacing:.28em;text-transform:uppercase} +h1{font-size:1.5rem;color:var(--cream);font-weight:600;margin:.35rem 0 .2rem} +.sub{color:var(--dim);font-size:.9rem;margin-bottom:1.6rem} +.card{background:linear-gradient(180deg,var(--raise),var(--panel));border:1px solid #262320; + border-radius:12px;padding:1.2rem;margin-bottom:1.4rem} +.stage{background:var(--well);border:1px solid #22201d;border-radius:8px;padding:1rem; + display:flex;justify-content:center} +.wnote{color:var(--dim);font-size:.85rem;margin-top:1rem} +.wnote b{color:var(--steel);font-weight:600} +.igrid{display:grid;grid-template-columns:11rem 1fr;gap:.35rem .9rem;margin-top:1rem; + font-size:.82rem} +.ik{color:var(--steel)} +.iv{color:var(--dim)} +</style> +</head> +<body> +<div class="wrap"> + <p class="eyebrow">Waste Time Beautifully · C1</p> + <h1>Cosmos — numbers in orbit</h1> + <p class="sub">John Maeda for Shiseido, 1995 (Cal0.class) — rebuilt as a web component. Standalone copy, no build step.</p> + + <div class="card"> + <div class="stage"><maeda-cosmos id="piece"></maeda-cosmos></div> + <p class="wnote"><b>Controls:</b> hover to reshape every orbit — the axes track 1.5× the + cursor's distance from each orbit's center; press and drag to pull a new loop out of the + system; the mark at bottom-left clears them and replays the fly-in credit. Click the + canvas first, then <b>ESC</b> toggles strip-move mode (the ground goes orange and the + clock strip comes free).</p> + <p class="wnote">The days of the month ride a rotated ellipse, today in red on the sweep + hand, with a seconds odometer scrolling at the right edge. The month as a gravitational + system, days as bodies in orbit, the user as a hand that perturbs the heavens. Time is + periodic, not linear.</p> + </div> + + <div class="card"> + <p class="eyebrow">Spec</p> + <div class="igrid"> + <span class="ik">original</span><span class="iv">Cal0.class + UniverseCal, orbit, eint, efloatsin (Java 1.1)</span> + <span class="ik">dimensions</span><span class="iv">400×300, black ground, Helvetica 10</span> + <span class="ik">collection</span><span class="iv">SFMOMA 99.550</span> + <span class="ik">revolution</span><span class="iv">7,500 ms per sweep, phase = wall clock mod cycle — every orbit in lockstep forever; orientation wobbles ±10° on a 15,000 ms sine</span> + <span class="ik">interaction</span><span class="iv">every mouse move (not just drag) restyles all orbits; press spawns a zero-axis orbit that inflates as the cursor pulls away</span> + <span class="ik">period quirks</span><span class="iv">a cursor-driven rotation angle is computed and stored but never drawn — a dead store in the 1995 bytecode; Mac Java Date bug detected and adjusted</span> + <span class="ik">credit</span><span class="iv">chars fly in at 800+80i ms from random scatter, hold 1.5 s, vanish; logo click replays</span> + <span class="ik">tests</span><span class="iv">38 (Vitest + fast-check)</span> + </div> + </div> +</div> + +<script> +(() => { + // src/pieces/cosmos/model.js + var TAU = Math.PI * 2; + var CREDIT = "designed by john maeda"; + var FAITHFUL_1995 = { + faceW: 400, + // applet tag WIDTH=400 HEIGHT=300 + faceH: 300, + cycleMs: 7500, + // one revolution, wall-clock modulo + wobbleAmp: 10 * Math.PI / 180, + // efloatsin(0, 10°, 15000) + wobblePeriodMs: 15e3, + stretch: 1.5, + // axes = 1.5× cursor distance from center + initialAxes: [300, 100], + // the startup orbit at panel center + fontSize: 10, + // applet param default; Helvetica plain + stripWidth: 20, + // stringWidth("59")+6 at 10 px, rounded + clockReach: 120, + // readout + weekday row extent left of the strip + logo: { dx: 8, dy: 8, w: 70, h: 16 } + // shiseido.gif box, bottom-left + }; + function sweepPhase(cfg, nowMs) { + return nowMs % cfg.cycleMs / cfg.cycleMs * TAU; + } + function wobbleAngle(cfg, elapsedMs) { + return cfg.wobbleAmp * Math.sin(elapsedMs * TAU / cfg.wobblePeriodMs); + } + function orbitPoint(o, theta, alpha) { + const ct = Math.cos(theta); + const st = Math.sin(theta); + const ca = Math.cos(alpha); + const sa = Math.sin(alpha); + return { + x: o.lmaj * ct * ca - o.lmin * st * sa + o.h, + y: o.lmaj * ct * sa + o.lmin * st * ca + o.k + }; + } + function dayTheta(day, today, ndays, phase) { + return phase + (day - today) * TAU / ndays; + } + function makeOrbit(x, y, lmaj = 0, lmin = 0) { + return { h: x + 0.5, k: y + 0.5, lmaj, lmin, wobbleT: 0 }; + } + function stretchOrbit(cfg, o, mx, my) { + return { ...o, lmaj: (mx - o.h) * cfg.stretch, lmin: (my - o.k) * cfg.stretch }; + } + function creditDurations(len) { + const chars = []; + for (let i = 0; i < len; i++) chars.push(800 + 80 * i); + return { chars, sentinel: 800 + 80 * len + 1500 }; + } + function easeLinear(t, a, b, dur) { + if (t <= 0) return a; + if (t >= dur) return b; + return a + (b - a) * t / dur; + } + function creditDone(len, t) { + return t >= creditDurations(len).sentinel; + } + function stripOffset(cfg, sec, msWithin) { + const ld = cfg.fontSize + 1; + const siddy = 3 + cfg.fontSize; + return -ld * sec - msWithin * ld / 1e3 + siddy - ld; + } + function logoHit(cfg, x, y) { + const { dx, dy, w, h } = cfg.logo; + return x < dx + w && y > cfg.faceH - h - dy; + } + function clockHit(cfg, strip, x, y) { + const ld = cfg.fontSize + 1; + const left = cfg.faceW - strip.sidex - cfg.clockReach; + const right = cfg.faceW - strip.sidex + cfg.stripWidth; + return x > left && x < right && y < strip.siddy && y > strip.siddy - ld; + } + + // src/pieces/cosmos/step.js + function scatterStarts(cfg, rng) { + const starts = []; + for (let i = 0; i < CREDIT.length; i++) { + starts.push({ + x: Math.floor(rng() * cfg.faceW) * (rng() < 0.5 ? 1 : -1) + cfg.faceW / 2, + y: Math.floor(rng() * cfg.faceH) * (rng() < 0.5 ? 1 : -1) + cfg.faceH / 2 + }); + } + return starts; + } + function initState(cfg, rng = Math.random) { + return { + orbits: [makeOrbit(cfg.faceW / 2, cfg.faceH / 2, cfg.initialAxes[0], cfg.initialAxes[1])], + credit: { t: 0, starts: scatterStarts(cfg, rng) }, + strip: { sidex: cfg.stripWidth, siddy: 3 + cfg.fontSize }, + killPending: false, + movable: false, + draggingStrip: false, + prevHeld: false, + prevCursor: null, + rng + }; + } + function toggleMovable(state) { + state.movable = !state.movable; + return state; + } + function step(cfg, state, dtMs, input) { + const s = { ...state, strip: { ...state.strip }, credit: { ...state.credit } }; + const { x, y, held } = input; + s.credit.t += dtMs; + s.orbits = s.orbits.map((o) => ({ ...o, wobbleT: o.wobbleT + dtMs })); + if (x == null) { + s.prevHeld = held; + return s; + } + if (held && !s.prevHeld) { + if (logoHit(cfg, x, y)) { + s.killPending = true; + s.credit = { t: 0, starts: scatterStarts(cfg, s.rng) }; + } else if (clockHit(cfg, s.strip, x, y) || s.movable) { + s.draggingStrip = true; + } else { + s.orbits = [...s.orbits, makeOrbit(x, y)]; + } + } + if (held && s.draggingStrip && s.prevCursor) { + s.strip.sidex -= x - s.prevCursor.x; + s.strip.siddy += y - s.prevCursor.y; + } + s.orbits = s.orbits.map((o) => stretchOrbit(cfg, o, x, y)); + if (!held && s.prevHeld) { + if (s.killPending) { + s.orbits = [s.orbits[0]]; + s.strip = { sidex: cfg.stripWidth, siddy: 3 + cfg.fontSize }; + } + s.killPending = false; + s.draggingStrip = false; + } + s.prevHeld = held; + s.prevCursor = { x, y }; + return s; + } + + // src/lib/calendar.js + function isLeapYear(y) { + return y % 4 === 0 && (y % 100 !== 0 || y % 400 === 0); + } + var MDAYS = [31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31]; + function daysInMonth(y, m) { + return m === 2 && isLeapYear(y) ? 29 : MDAYS[m - 1]; + } + + // src/pieces/cosmos/render.js + var FACE = { w: 400, h: 300 }; + var MONTHS = ["jan ", "feb ", "mar ", "apr ", "may ", "jun ", "jul ", "aug ", "sep ", "oct ", "nov ", "dec "]; + var WEEKDAYS = ["S ", "M ", "T ", "W ", "T ", "F ", "S "]; + function font(cfg) { + return `${cfg.fontSize}px Helvetica, Arial, sans-serif`; + } + function dateParts(nowMs) { + const d = new Date(nowMs); + return { + year: d.getFullYear(), + month: d.getMonth() + 1, + // 1-12 + day: d.getDate(), + weekday: d.getDay(), + hours: d.getHours(), + minutes: d.getMinutes(), + seconds: d.getSeconds(), + msWithin: nowMs % 1e3 + }; + } + function drawOrbit(ctx, cfg, o, dp, phase, dim) { + const ndays = daysInMonth(dp.year, dp.month); + const alpha = wobbleAngle(cfg, o.wobbleT); + const daystr = `${dp.day} ${dp.year}`; + const monthstr = MONTHS[dp.month - 1]; + for (let d = 1; d <= ndays; d++) { + const p = orbitPoint(o, dayTheta(d, dp.day, ndays, phase), alpha); + if (d === dp.day) { + ctx.fillStyle = "#f00"; + ctx.fillText(daystr, p.x, p.y); + ctx.fillText(monthstr, p.x - ctx.measureText(monthstr).width, p.y); + } else { + ctx.fillStyle = dim ? "#808080" : "#fff"; + ctx.fillText(String(d), p.x, p.y); + } + } + } + function drawStrip(ctx, cfg, state, dp) { + const ld = cfg.fontSize + 1; + const { sidex, siddy } = state.strip; + const x = FACE.w - sidex; + const tileH = ld * 60; + const off = stripOffset(cfg, dp.seconds, dp.msWithin) + (siddy - (3 + cfg.fontSize)); + ctx.save(); + ctx.beginPath(); + ctx.rect(x, 0, cfg.stripWidth, FACE.h); + ctx.clip(); + ctx.fillStyle = "#404040"; + ctx.fillRect(x, 0, cfg.stripWidth, FACE.h); + ctx.fillStyle = "#808080"; + for (const tile of [-1, 0, 1]) { + for (let i = 0; i < 60; i++) { + const y = off + tile * tileH + i * ld + ld; + if (y > -ld && y < FACE.h + ld) { + ctx.fillText(String(i).padStart(2, "0"), x + 3, y); + } + } + } + ctx.restore(); + ctx.fillStyle = "#f00"; + ctx.fillText(String(dp.seconds).padStart(2, "0"), x + 3, siddy); + const h12 = dp.hours % 12 === 0 ? 12 : dp.hours % 12; + const clock = `${h12}${dp.minutes < 10 ? ":0" : ":"}${dp.minutes}:`; + const cx = x + 3 - ctx.measureText(clock).width; + ctx.fillText(clock, cx, siddy); + let wx = cx - ctx.measureText("S M T W T F S ").width; + for (let i = 0; i < 7; i++) { + ctx.fillStyle = i === dp.weekday ? "#f00" : "#808080"; + ctx.fillText(WEEKDAYS[i], wx, siddy); + wx += ctx.measureText(WEEKDAYS[i]).width; + } + } + function drawCredit(ctx, cfg, state) { + const { t, starts } = state.credit; + if (creditDone(CREDIT.length, t)) return; + const durs = creditDurations(CREDIT.length); + let fx = (FACE.w - ctx.measureText(CREDIT).width) / 2 + 1; + const fy = FACE.h / 2 + cfg.fontSize / 2; + ctx.fillStyle = "#fff"; + for (let i = 0; i < CREDIT.length; i++) { + const ch = CREDIT[i]; + const x = easeLinear(t, starts[i].x, fx, durs.chars[i]); + const y = easeLinear(t, starts[i].y, fy, durs.chars[i]); + ctx.fillText(ch, x, y); + fx += ctx.measureText(ch).width; + } + } + function drawLogo(ctx, cfg) { + const { dx, dy, h } = cfg.logo; + ctx.fillStyle = "#808080"; + ctx.fillText("shiseido", dx, FACE.h - dy - h / 2 + cfg.fontSize / 2); + } + function draw(ctx, cfg, state, nowMs) { + const dp = dateParts(nowMs); + ctx.fillStyle = state.movable ? "#e08000" : "#000"; + ctx.fillRect(0, 0, FACE.w, FACE.h); + ctx.font = font(cfg); + ctx.textBaseline = "alphabetic"; + drawStrip(ctx, cfg, state, dp); + const phase = sweepPhase(cfg, nowMs); + state.orbits.forEach((o, i) => { + drawOrbit(ctx, cfg, o, dp, phase, state.killPending && i > 0); + }); + drawLogo(ctx, cfg); + drawCredit(ctx, cfg, state); + ctx.strokeStyle = "#fff"; + ctx.lineWidth = 1; + ctx.strokeRect(0.5, 0.5, FACE.w - 1, FACE.h - 1); + } + + // src/lib/loop.js + var STEP_MS = 30; + function startLoop({ stepFn, renderFn, stepMs = STEP_MS }) { + let acc = 0; + let last = null; + let raf = null; + let running = true; + function frame(ts) { + if (!running) return; + if (last === null) last = ts; + acc += Math.min(ts - last, 250); + last = ts; + while (acc >= stepMs) { + stepFn(stepMs); + acc -= stepMs; + } + renderFn(); + raf = requestAnimationFrame(frame); + } + raf = requestAnimationFrame(frame); + return () => { + running = false; + if (raf) cancelAnimationFrame(raf); + }; + } + + // src/pieces/cosmos/index.js + var MaedaCosmos = class extends HTMLElement { + connectedCallback() { + const canvas = document.createElement("canvas"); + canvas.width = FACE.w; + canvas.height = FACE.h; + canvas.style.cssText = "display:block;width:100%;max-width:400px;cursor:crosshair;touch-action:none"; + canvas.tabIndex = 0; + this.appendChild(canvas); + const ctx = canvas.getContext("2d"); + const frozen = this.getAttribute("frozen-now"); + const nowMs = () => frozen ? Number(frozen) * 1e3 : Date.now(); + const cfg = FAITHFUL_1995; + let state = initState(cfg); + const input = { x: null, y: null, held: false }; + const toFace = (ev) => { + const r = canvas.getBoundingClientRect(); + return { + x: (ev.clientX - r.left) * (FACE.w / r.width), + y: (ev.clientY - r.top) * (FACE.h / r.height) + }; + }; + canvas.addEventListener("pointermove", (ev) => { + Object.assign(input, toFace(ev)); + }); + canvas.addEventListener("pointerdown", (ev) => { + canvas.setPointerCapture(ev.pointerId); + canvas.focus(); + Object.assign(input, toFace(ev)); + input.held = true; + }); + const release = () => { + input.held = false; + }; + canvas.addEventListener("pointerup", release); + canvas.addEventListener("pointercancel", release); + canvas.addEventListener("keydown", (ev) => { + if (ev.key === "Escape") toggleMovable(state); + }); + this._stop = startLoop({ + stepFn: (dtMs) => { + state = step(cfg, state, dtMs, input); + }, + renderFn: () => draw(ctx, cfg, state, nowMs()) + }); + } + disconnectedCallback() { + if (this._stop) this._stop(); + } + }; + customElements.define("maeda-cosmos", MaedaCosmos); +})(); + +</script> +</body> +</html> diff --git a/working/clock-display-references/2026-07-30-maeda-line-notes.org b/working/clock-display-references/2026-07-30-maeda-line-notes.org new file mode 100644 index 0000000..ba50a80 --- /dev/null +++ b/working/clock-display-references/2026-07-30-maeda-line-notes.org @@ -0,0 +1,11 @@ +#+TITLE: Copy of the Line applet (C2, 1997) from the Maeda x Shiseido +#+SOURCE: from website +#+DATE: 2026-07-30 18:38:41 -0500 + +Copy of the Line applet (C2, 1997) from the Maeda x Shiseido tribute — the diagonal timeline piece, 'zoom through time'. Delivered as 2026-07-30-1838-from-website-2026-07-30-maeda-line-standalone.html in this inbox. + +It is a single self-contained file: the <maeda-line> web component and its model/step/render modules are esbuild-bundled inline as an IIFE, so it opens straight from file:// with no server, no build step, and no network. Verified rendering in Chrome from file:// — canvas 460x400, zero console errors, live readout wired. + +Controls: move the pointer along the diagonal to aim, press and hold to zoom continuously into that moment, release to float back out. Full zoom is exactly 60 seconds and the release is symmetric. + +Source of truth stays ~/code/maeda-tribute (src/pieces/line/); this copy is a snapshot, so regenerate rather than edit it if the piece changes. Sent for reference — no action needed unless you want it wired into the panel-widget gallery. diff --git a/working/clock-display-references/2026-07-30-maeda-line-standalone.html b/working/clock-display-references/2026-07-30-maeda-line-standalone.html new file mode 100644 index 0000000..783939c --- /dev/null +++ b/working/clock-display-references/2026-07-30-maeda-line-standalone.html @@ -0,0 +1,457 @@ +<!doctype html> +<html lang="en"> +<head> +<meta charset="utf-8"> +<meta name="viewport" content="width=device-width, initial-scale=1"> +<title>Line (C2, 1997) — Maeda × Shiseido tribute</title> +<style> +:root{ + --ground:#151311; + --panel:#100f0f; + --well:#0a0c0d; + --raise:#1a1917; + --silver:#bfc4d0; + --cream:#f3e7c5; + --steel:#969385; + --dim:#7c838a; + --wash:#2c2f32; + --gold:#e2a038; + --mono:"BerkeleyMono Nerd Font","Berkeley Mono",monospace; +} +*{box-sizing:border-box;margin:0;padding:0} +html{background:var(--ground);font-size:112%} +body{font-family:var(--mono);color:var(--silver);padding:2.4rem 2rem 4rem;line-height:1.45; + background:radial-gradient(1200px 600px at 70% -10%,#1c1915 0%,transparent 60%),var(--ground)} +.wrap{max-width:820px;margin:0 auto} +.eyebrow{color:var(--steel);font-size:.72rem;letter-spacing:.28em;text-transform:uppercase} +h1{font-size:1.5rem;color:var(--cream);font-weight:600;margin:.35rem 0 .2rem} +.sub{color:var(--dim);font-size:.9rem;margin-bottom:1.6rem} +.card{background:linear-gradient(180deg,var(--raise),var(--panel));border:1px solid #262320; + border-radius:12px;padding:1.2rem;margin-bottom:1.4rem} +.stage{background:var(--well);border:1px solid #22201d;border-radius:8px;padding:1rem; + display:flex;justify-content:center} +.readout{margin-top:.8rem;color:var(--gold);font-size:.82rem;letter-spacing:.06em} +.wnote{color:var(--dim);font-size:.85rem;margin-top:1rem} +.wnote b{color:var(--steel);font-weight:600} +.igrid{display:grid;grid-template-columns:11rem 1fr;gap:.35rem .9rem;margin-top:1rem; + font-size:.82rem} +.ik{color:var(--steel)} +.iv{color:var(--dim)} +</style> +</head> +<body> +<div class="wrap"> + <p class="eyebrow">Waste Time Beautifully · C2</p> + <h1>Line — zoom through time</h1> + <p class="sub">John Maeda for Shiseido, 1997 (cal1.class) — rebuilt as a web component. Standalone copy, no build step.</p> + + <div class="card"> + <div class="stage"><maeda-line id="piece"></maeda-line></div> + <div class="readout" id="readout">scale — · counter —</div> + <p class="wnote"><b>Controls:</b> move the pointer along the diagonal to aim; press and + hold to zoom continuously into that moment; release and you float back out. Full zoom + is exactly 60 seconds, and the release is symmetric.</p> + <p class="wnote">Time as a single line — <b>diagonal</b>, bottom-left to top-right, the + detail every written description of this applet gets wrong. Maeda's note: it reflects + the relativity and comparability of time intervals. The zoom is the content.</p> + </div> + + <div class="card"> + <p class="eyebrow">Spec</p> + <div class="igrid"> + <span class="ik">original</span><span class="iv">cal1.class, 12,339 bytes, single class, self-contained (Java 1.1)</span> + <span class="ik">dimensions</span><span class="iv">460×400, black ground, Helvetica throughout</span> + <span class="ik">collection</span><span class="iv">SFMOMA 99.554</span> + <span class="ik">zoom plateaus</span><span class="iv">72px per day / hour / minute / second — scn 1 → 72 → 1,728 → 103,680 → 6,220,800 px/day</span> + <span class="ik">easing</span><span class="iv">counter at 20 increments/sec over segments [150, 250, 350, 450]</span> + <span class="ik">labels</span><span class="iv">font = min(72, unit spacing); text under 5px degrades to a line; gray ramps 128→255 as a unit matures; the current unit is always red</span> + <span class="ik">period quirks</span><span class="iv">Mac JVM patch capped zoom at hour level; an unused HAPPY NEW YEAR string ships in the bytecode</span> + <span class="ik">tests</span><span class="iv">33 (Vitest + fast-check)</span> + </div> + </div> +</div> + +<script> +(() => { + // src/lib/calendar.js + function epochDayFromCivil(y, m, d) { + const yy = y - (m <= 2 ? 1 : 0); + const era = Math.floor(yy / 400); + const yoe = yy - era * 400; + const doy = Math.floor((153 * (m + (m > 2 ? -3 : 9)) + 2) / 5) + d - 1; + const doe = yoe * 365 + Math.floor(yoe / 4) - Math.floor(yoe / 100) + doy; + return era * 146097 + doe - 719468; + } + function civilFromEpochDay(ed) { + const z = ed + 719468; + const era = Math.floor(z / 146097); + const doe = z - era * 146097; + const yoe = Math.floor( + (doe - Math.floor(doe / 1460) + Math.floor(doe / 36524) - Math.floor(doe / 146096)) / 365 + ); + const y = yoe + era * 400; + const doy = doe - (365 * yoe + Math.floor(yoe / 4) - Math.floor(yoe / 100)); + const mp = Math.floor((5 * doy + 2) / 153); + const d = doy - Math.floor((153 * mp + 2) / 5) + 1; + const m = mp + (mp < 10 ? 3 : -9); + return { y: y + (m <= 2 ? 1 : 0), m, d }; + } + + // src/pieces/line/model.js + var DAY = 86400; + var FAITHFUL_1997 = { + stops: [1, 72, 1728, 103680, 6220800], + segments: [150, 250, 350, 450], + incsPerSec: 20 + }; + function totalIncrements(cfg) { + return cfg.segments.reduce((a, b) => a + b, 0); + } + function scaleForCounter(cfg, counter) { + const total = totalIncrements(cfg); + const c = Math.max(0, Math.min(total, counter)); + let acc = 0; + for (let i = 0; i < cfg.segments.length; i++) { + const seg = cfg.segments[i]; + if (c <= acc + seg) { + const f = (c - acc) / seg; + return cfg.stops[i] + f * (cfg.stops[i + 1] - cfg.stops[i]); + } + acc += seg; + } + return cfg.stops[cfg.stops.length - 1]; + } + function advanceCounter(cfg, counter, dtMs, held) { + const d = cfg.incsPerSec * dtMs / 1e3; + const next = held ? counter + d : counter - d; + return Math.max(0, Math.min(totalIncrements(cfg), next)); + } + function pxOfTime(view, t) { + return (t - view.focus) * view.scale / DAY + view.anchorPx; + } + function timeOfPx(view, px) { + return (px - view.anchorPx) * DAY / view.scale + view.focus; + } + var MAX_FONT = 72; + var MIN_TEXT_PX = 5; + function fontSizeFor(spacingPx) { + return Math.min(MAX_FONT, spacingPx); + } + function textVisible(fontSize) { + return fontSize >= MIN_TEXT_PX; + } + function grayFor(fontSize) { + return Math.floor(Math.min(MAX_FONT, fontSize) * 127 / MAX_FONT) + 128; + } + var FIXED_UNITS = { second: 1, minute: 60, hour: 3600, day: DAY }; + function* calendarTicks(unit, t0, t1) { + let { y, m } = civilFromEpochDay(Math.floor(t0 / DAY)); + const stride = unit === "decade" ? 10 : unit === "century" ? 100 : 1; + if (unit === "month") { + for (; ; ) { + const t = epochDayFromCivil(y, m, 1) * DAY; + if (t >= t0) break; + m++; + if (m > 12) m = 1, y++; + } + for (; ; ) { + const t = epochDayFromCivil(y, m, 1) * DAY; + if (t > t1) return; + yield t; + m++; + if (m > 12) m = 1, y++; + } + } else { + let yy = Math.ceil(y / stride) * stride; + if (epochDayFromCivil(yy, 1, 1) * DAY < t0) yy += stride; + while (epochDayFromCivil(yy - stride, 1, 1) * DAY >= t0) yy -= stride; + for (; ; ) { + const t = epochDayFromCivil(yy, 1, 1) * DAY; + if (t > t1) return; + if (t >= t0) yield t; + yy += stride; + } + } + } + function ticksInRange(unit, t0, t1) { + if (unit in FIXED_UNITS) { + const w = FIXED_UNITS[unit]; + const out = []; + for (let t = Math.ceil(t0 / w) * w; t <= t1; t += w) out.push(t); + return out; + } + return [...calendarTicks(unit, t0, t1)]; + } + var MONTHS = [ + "JANUARY", + "FEBRUARY", + "MARCH", + "APRIL", + "MAY", + "JUNE", + "JULY", + "AUGUST", + "SEPTEMBER", + "OCTOBER", + "NOVEMBER", + "DECEMBER" + ]; + function formatMonth(m) { + return MONTHS[m - 1]; + } + function formatDay(m, d) { + return `${m}/${d}`; + } + function formatHour(h) { + const twelve = h % 12 === 0 ? 12 : h % 12; + return { text: ` ${twelve}`, suffix: h >= 12 ? "pm" : "am" }; + } + function formatMinute(mm) { + return `:${String(mm).padStart(2, "0")}`; + } + + // src/pieces/line/step.js + function initState(cfg, range, lineLenPx) { + return { + counter: 0, + scale: scaleForCounter(cfg, 0), + focus: 0, + // epoch seconds under the anchor pixel + anchorPx: 0, + // pixel along the line where focus projects + range, + // {tMin, tMax} or null for the unbounded twist + lineLenPx + }; + } + function step(cfg, s, dtMs, input) { + const cursorPx = Math.max(0, Math.min(s.lineLenPx, input.cursorPx)); + let { focus, anchorPx } = s; + if (input.held || s.counter > 0) { + focus = timeOfPx({ focus: s.focus, scale: s.scale, anchorPx: s.anchorPx }, cursorPx); + anchorPx = cursorPx; + } + const counter = advanceCounter(cfg, s.counter, dtMs, input.held); + const scale = scaleForCounter(cfg, counter); + if (!input.held && counter < cfg.segments[0]) { + const target = cursorPx * DAY; + focus += (target - focus) / (counter + 1); + anchorPx = cursorPx; + } + if (s.range) { + focus = Math.max(s.range.tMin, Math.min(s.range.tMax, focus)); + } + return { ...s, counter, scale, focus, anchorPx }; + } + + // src/pieces/line/render.js + var LEVELS = [ + { unit: "month", secs: 30 * DAY }, + { unit: "day", secs: DAY }, + { unit: "hour", secs: 3600 }, + { unit: "minute", secs: 60 }, + { unit: "second", secs: 1 } + ]; + var FACE = { w: 460, h: 400 }; + function lineOrigin(state) { + const margin = (FACE.h - state.lineLenPx) / 2; + return { mx: FACE.w - FACE.h + margin, my: margin }; + } + function labelFor(unit, t) { + const dayIdx = Math.floor(t / DAY); + const civ = civilFromEpochDay(dayIdx); + const rem = t - dayIdx * DAY; + const hh = Math.floor(rem / 3600); + const mm = Math.floor(rem % 3600 / 60); + const ss = Math.floor(rem % 60); + switch (unit) { + case "month": + return formatMonth(civ.m); + case "day": + return formatDay(civ.m, civ.d); + case "hour": { + const h = formatHour(hh); + return h.text + h.suffix; + } + case "minute": + return formatMinute(mm); + case "second": + return formatMinute(ss); + default: + return ""; + } + } + function sameTick(unit, t, now) { + const w = { second: 1, minute: 60, hour: 3600, day: DAY }[unit]; + if (w) return Math.floor(t / w) === Math.floor(now / w); + const a = civilFromEpochDay(Math.floor(t / DAY)); + const b = civilFromEpochDay(Math.floor(now / DAY)); + return a.y === b.y && a.m === b.m; + } + function draw(ctx, state, nowSec) { + const { mx, my } = lineOrigin(state); + ctx.fillStyle = "#000"; + ctx.fillRect(0, 0, FACE.w, FACE.h); + const view = { focus: state.focus, scale: state.scale, anchorPx: state.anchorPx }; + const tA = Math.max( + state.range ? state.range.tMin : -Infinity, + state.focus - state.anchorPx * DAY / state.scale + ); + const tB = Math.min( + state.range ? state.range.tMax : Infinity, + state.focus + (state.lineLenPx - state.anchorPx) * DAY / state.scale + ); + ctx.strokeStyle = "#666"; + ctx.beginPath(); + ctx.moveTo(mx, FACE.h - my); + ctx.lineTo(mx + state.lineLenPx, FACE.h - my - state.lineLenPx); + ctx.stroke(); + const colWidth = {}; + for (let i = LEVELS.length - 1; i >= 0; i--) { + const { unit, secs } = LEVELS[i]; + const spacing = secs * state.scale / DAY; + if (spacing < 2 && unit !== "month") { + colWidth[unit] = 0; + continue; + } + const size = Math.floor(fontSizeFor(spacing)); + colWidth[unit] = textVisible(size) ? size * 2.2 : 0; + } + for (let i = LEVELS.length - 1; i >= 0; i--) { + const { unit, secs } = LEVELS[i]; + const spacing = secs * state.scale / DAY; + if (spacing < 2 && unit !== "month") continue; + const size = unit === "month" ? Math.floor(Math.min(72, 11 + state.scale)) : Math.floor(fontSizeFor(spacing)); + const gray = unit === "month" ? 255 : grayFor(size); + let offset = 0; + for (let j = LEVELS.length - 1; j > i; j--) offset += colWidth[LEVELS[j].unit]; + ctx.font = `${Math.max(size, 1)}px Helvetica, Arial, sans-serif`; + for (const t of ticksInRange(unit, tA, tB)) { + const n = pxOfTime(view, t); + if (n < 0 || n > state.lineLenPx) continue; + const x = mx + n; + const y = FACE.h - my - n; + const isNow = sameTick(unit, t, nowSec); + ctx.strokeStyle = ctx.fillStyle = isNow ? "#f00" : `rgb(${gray},${gray},${gray})`; + ctx.beginPath(); + ctx.moveTo(x, y); + ctx.lineTo(x + 3, y); + ctx.stroke(); + const text = labelFor(unit, t); + if (textVisible(size)) { + const w = ctx.measureText(text).width; + ctx.fillText(text, x - w - offset, y); + } else { + ctx.beginPath(); + ctx.moveTo(x - 4 - offset, y); + ctx.lineTo(x - offset, y); + ctx.stroke(); + } + } + } + ctx.fillStyle = "#ccc"; + ctx.font = "15px Helvetica, Arial, sans-serif"; + const label = `${Math.round(state.scale)}X${state.counter >= 1200 ? " (MAX)" : ""}`; + ctx.fillText(label, 3, 16); + } + + // src/lib/loop.js + var STEP_MS = 30; + function startLoop({ stepFn, renderFn, stepMs = STEP_MS }) { + let acc = 0; + let last = null; + let raf = null; + let running = true; + function frame(ts) { + if (!running) return; + if (last === null) last = ts; + acc += Math.min(ts - last, 250); + last = ts; + while (acc >= stepMs) { + stepFn(stepMs); + acc -= stepMs; + } + renderFn(); + raf = requestAnimationFrame(frame); + } + raf = requestAnimationFrame(frame); + return () => { + running = false; + if (raf) cancelAnimationFrame(raf); + }; + } + + // src/pieces/line/index.js + var MaedaLine = class extends HTMLElement { + connectedCallback() { + const canvas = document.createElement("canvas"); + canvas.width = FACE.w; + canvas.height = FACE.h; + canvas.style.cssText = "display:block;width:100%;max-width:460px;cursor:crosshair;touch-action:none"; + this.appendChild(canvas); + const ctx = canvas.getContext("2d"); + const frozen = this.getAttribute("frozen-now"); + const nowSec = () => frozen ? Number(frozen) : Date.now() / 1e3; + const { y } = (() => { + const d = new Date(nowSec() * 1e3); + return { y: d.getUTCFullYear() }; + })(); + const t0 = epochDayFromCivil(y, 1, 1) * DAY; + const t1 = epochDayFromCivil(y + 1, 1, 1) * DAY; + const lineLen = Math.round((t1 - t0) / DAY); + const cfg = FAITHFUL_1997; + let state = initState(cfg, { tMin: 0, tMax: t1 - t0 }, lineLen); + state.focus = nowSec() - t0; + state.anchorPx = state.focus / DAY; + const input = { cursorPx: state.anchorPx, held: false }; + const toLinePx = (ev) => { + const r = canvas.getBoundingClientRect(); + const sx = FACE.w / r.width; + const x = (ev.clientX - r.left) * sx; + const yy = (ev.clientY - r.top) * (FACE.h / r.height); + const margin = (FACE.h - lineLen) / 2; + const mx = FACE.w - FACE.h + margin; + return (x - mx + (FACE.h - margin - yy)) / 2; + }; + canvas.addEventListener("pointermove", (ev) => { + input.cursorPx = toLinePx(ev); + }); + canvas.addEventListener("pointerdown", (ev) => { + canvas.setPointerCapture(ev.pointerId); + input.cursorPx = toLinePx(ev); + input.held = true; + }); + const release = () => { + input.held = false; + }; + canvas.addEventListener("pointerup", release); + canvas.addEventListener("pointercancel", release); + this._stop = startLoop({ + stepFn: (dtMs) => { + state = step(cfg, state, dtMs, input); + }, + renderFn: () => { + draw(ctx, state, nowSec() - t0); + this.dispatchEvent(new CustomEvent("readout", { + detail: { scale: state.scale, counter: state.counter } + })); + } + }); + } + disconnectedCallback() { + if (this._stop) this._stop(); + } + }; + customElements.define("maeda-line", MaedaLine); +})(); + +</script> +<script> +document.getElementById("piece").addEventListener("readout", (e) => { + const { scale, counter } = e.detail; + document.getElementById("readout").textContent = + "scale " + Number(scale).toFixed(3) + " · counter " + Math.round(counter); +}); +</script> +</body> +</html> diff --git a/working/hyprland-lua-port/README.org b/working/hyprland-lua-port/README.org new file mode 100644 index 0000000..42779e5 --- /dev/null +++ b/working/hyprland-lua-port/README.org @@ -0,0 +1,131 @@ +#+TITLE: Hyprland .conf → Lua port — staged, not deployed +#+AUTHOR: Craig Jennings + +* Status + +Built and verified in a nested compositor. *Not deployed.* I deployed it to the +dotfiles tree on 2026-08-24 and then rolled it back the same afternoon, because +the switch had not been checked on real hardware and the machine needs to stay +usable. The dotfiles repo is untouched at =8f692f5=; the live config is the +original =hyprland.conf=. + +The port goes live only after the hardware check in =todo.org= under "Manual +testing and validation" passes. + +* What is here + +| File | What it is | +|----------------------------+---------------------------------------------------------------| +| =hyprland.lua= | The deliverable, 788 lines. Shared config. | +| =velox-local.lua= | velox host overrides, for =velox/.config/hypr/conf.d/=. | +| =ratio-local.lua= | ratio host overrides, for =ratio/.config/hypr/conf.d/=. | +| =reader-changes-for-lua.patch= | dotfiles-side: the four readers, ported and mutation-tested. | +| =test-desktop-for-lua.patch= | archsetup-side: the post-install desktop checks. | +| =hyprland.lua.generated= | Raw =hyprlang2lua= output, merging mode. Derivation evidence. | +| =nomerge.lua= | Same converter with =--no-merge=. Derivation evidence. | + +The two patches are the part that is easy to lose and expensive to redo. Both +were mutation-tested — every ported assertion was confirmed to go red when the +property it guards was removed — so replay them rather than rewriting the +assertions from scratch. + +* Redeploy, when the port is ready + +1. =cp hyprland.lua ~/.dotfiles/hyprland/.config/hypr/hyprland.lua= +2. =cp velox-local.lua ~/.dotfiles/velox/.config/hypr/conf.d/local.lua= +3. =cp ratio-local.lua ~/.dotfiles/ratio/.config/hypr/conf.d/local.lua= +4. Move the three =.conf= files out of their stow packages. Do NOT merely leave + them beside the =.lua=: with both present Hyprland 0.56.2 loads the =.lua= + (proven — see below), so leaving the =.conf= in place buys no rollback and + only creates ambiguity about which file is live. +5. =cd ~/.dotfiles && git apply <path>/reader-changes-for-lua.patch= +6. =cd ~/code/archsetup && git apply <path>/test-desktop-for-lua.patch= +7. Restow. Expect two traps, both hit on 2026-08-24 and both documented in + =todo.org=: =make restow hyprland= aborts on the pre-existing + =obsbot-wb-guard.service= conflict in =common= (restow =hyprland= and the host + package individually instead), and the running Hyprland rewrites a stub + =hyprland.conf= within a second of the symlink vanishing. Silence the stub with + =hyprctl keyword misc:disable_autoreload 1=, do the stow, then set it back to 0. +8. *Push the dotfiles change before committing archsetup.* The installer clones + the dotfiles *remote* (=archsetup:1481=), so until the push lands a fresh VM + stows a tree with only =hyprland.conf= and the post-install checks fail. + +* Two things already proven, so nobody re-derives them + +*With both files present, the =.lua= wins.* Tested in a nested Hyprland 0.56.2 +with a fixture whose =.conf= set =gaps_in=11= and whose =.lua= set =77=. Result +was 77, and the log read "[cfg] Using lua config found at ...hyprland.lua". This +is why step 4 moves the =.conf= out rather than leaving it as a fallback. + +*The converter is a draft, not an answer.* =hyprlang2lua= +(github.com/EIonTusk/hyprlang2lua) reported 100% coverage and still produced +three functional defects, two of which would have broken the desktop. The two +generated files are kept as evidence: both still carry the unfixed bind defect +(="CTRL" .. mod .. " + S"=, which collapses to an unparseable =CTRLSUPER + S=), +and the merging-mode file shows the source glob emitted mid-file where it silently +reverses every per-host override. =--no-merge= fixed the ordering structurally; +the rest were hand-fixed. + +* Review findings folded in (2026-08-24) + +An isolated review of the deployed diff, before the rollback. Four were fixed in +the files here; the rest are gates on redeploying, not on the port's correctness. + +** Fixed here + +- =hl_source_glob= now surfaces all three failure modes and survives them. A + host override that fails to parse was warned about and skipped, which on velox + means coming up with no =force_zero_scaling= and no monitor scale — looking + like the whole port failed rather than one file. A runtime error inside the + chunk was unprotected and would have taken the entire config down over a single + host file. Now: parse failure says SKIPPED, runtime failure says PARTIAL and is + caught with =pcall=, and a glob matching nothing says NO MATCH. hyprlang did + none of this. +- The =col.nogroup_border*= rationale sat *below* the =col= table, reading as a + preamble to =layout=. Moved above the two keys it explains — the same defect + the file header says was fixed for autostart. +- =Generated by hyprlang2lua. Review TODOs before reloading Hyprland= removed + from all three files. No TODOs exist, and in =ratio-local.lua= it had landed + mid-paragraph, splitting the DP-4 rationale from the =hl.monitor= call it + explains. +- =ratio-local.lua='s usage examples were still hyprlang syntax + (=monitor=DP-1,...=, =bind = $mod, L, ...=), which are syntax errors in a Lua + file, and the second named =$mod= — a variable a sourced chunk cannot see. + Rewritten in Lua, with a note on the scoping. Verified rather than assumed: + =loadfile= gives the chunk globals only, and both =mod= and =at_start= are + locals, so both read =nil= inside a sourced file. + +** Refuted by measurement + +- *Duplicate chords append; they do not replace.* The config binds Super+Z twice + on purpose (=exec pypr zoom= plus =submap zoom=), and the same for Escape + inside the submap — if the Lua API replaced rather than appended, Super+Z would + enter the submap without zooming and the pairing the config's own comment + relies on would be broken. Tested in a nested Hyprland 0.56.2 with exactly that + shape: both binds register. The live =.conf= session registers the same pair, + so behaviour matches. No action needed; recorded so nobody re-derives it. + +** Gates on redeploy — do these as part of the switch + +1. *Do not put the =.conf= files in a =retired/= directory inside the repo + without also excluding them from =dotfiles-validate=.* Its find uses + =-path '*/.config/hypr/*.conf'=, which globs across slashes and would match + the retired copies — 81 of 228 checked references came from the dead config + when this was tried. The validator would then fail pointing at a file kept + precisely because nothing loads it. Add =-not -path "$root/retired/*"= to both + finds, or park the =.conf= files outside the repo entirely. +2. *Add tests for the new =dotfiles-validate= Lua branch.* The 25 new lines ship + with none. Proven vacuous: replacing both new awk regexes with =NEVERMATCHES= + still leaves =tests/dotfiles-validate/= reporting 15 tests OK. An extractor + that matches nothing prints nothing and exits 0 — the same false-pass shape + the three ported test suites got vacuity guards for. This one has no guard. +3. *Guard the override ordering.* Nothing asserts =hl_source_glob= is the last + statement in =hyprland.lua=, and it is the single invariant the whole per-host + layer rests on. An edit that moves it above the =hl.config= blocks silently + reverses every host override — converter defect 1, the one that would have + reverted the Qt scaling fix. archsetup's VM suite structurally cannot catch + it, because the VM stows no host tier. The guard belongs in the dotfiles repo. +4. *Sweep the prose comments that still name =hyprland.conf=.* About fifteen + across live scripts. =hyprland/.local/bin/waybar-reserve:12= is the one that + matters: it documents "Wired as =exec = waybar-reserve= in hyprland.conf", + which is exactly the mechanism the port replaces with =at_reload=. diff --git a/working/hyprland-lua-port/hyprland.lua b/working/hyprland-lua-port/hyprland.lua new file mode 100644 index 0000000..c005193 --- /dev/null +++ b/working/hyprland-lua-port/hyprland.lua @@ -0,0 +1,812 @@ +-- Hyprland Configuration +-- Translated from DWM config.def.h and sxhkdrc +-- Craig Jennings <c@cjennings.net> + +-- ============================================================================ +-- Monitor Configuration +-- ============================================================================ + +-- hyprlang2lua polyfills — runtime helpers reproducing +-- hyprlang behaviour the typed Lua API doesn't expose directly. + +local function hl_source_glob(pattern) + -- 'source = path/*.conf' had hyprlang glob and inline-expand the + -- matches. require() can't glob, so we shell out to ls (matching + -- the user's brace-expansion behaviour) and dofile each result. + -- Paths with spaces or shell metacharacters in the directory + -- portion will misparse; typical ~/.config/hypr/ layouts don't + -- hit this. Swap to lfs.dir() or find -name if you need fancier. + -- + -- Both failure paths are surfaced loudly and neither is fatal. hyprlang did + -- neither, and the asymmetry matters in opposite directions: a syntax error + -- that only warns means velox comes up with no force_zero_scaling and no + -- monitor scale, which looks like the port failed rather than like one file + -- failed; and an uncaught runtime error inside the chunk would take the + -- entire config down over a single host override. So: report both, continue + -- past both. + local p = io.popen("ls " .. pattern .. " 2>/dev/null") + if not p then return end + local matched = 0 + for f in p:lines() do + matched = matched + 1 + local chunk, err = loadfile(f) + if not chunk then + io.stderr:write("hl_source_glob: SKIPPED " .. f .. + " -- it did not parse, so nothing in it applied: " .. + tostring(err) .. "\n") + else + local ok, rerr = pcall(chunk) + if not ok then + io.stderr:write("hl_source_glob: PARTIAL " .. f .. + " -- it errored partway, so some of it applied " .. + "and the rest did not: " .. tostring(rerr) .. "\n") + end + end + end + p:close() + if matched == 0 then + io.stderr:write("hl_source_glob: NO MATCH for " .. pattern .. + " -- every host override is missing. On a machine that " .. + "has a conf.d file this means the glob is wrong.\n") + end +end + +-- Autostart collectors. Each command stays under the comment that explains it, +-- in the order hyprlang ran them; the hl.on() handlers at the bottom replay the +-- lists. Written this way because the generated form hoisted every command into +-- one block at the end of the file and left the comments stranded where the +-- commands had been -- in this config that means paragraphs of rationale with +-- no code under them, and one cross-reference ("the two exec-once lines above") +-- that had become false. +local autostart, atshutdown, atreload = {}, {}, {} +local function at_start(cmd) autostart[#autostart + 1] = cmd end +local function at_shutdown(cmd) atshutdown[#atshutdown + 1] = cmd end +local function at_reload(cmd) atreload[#atreload + 1] = cmd end + +hl.monitor({ + output = "", + mode = "preferred", + position = "auto", + scale = "auto", +}) + +-- Waybar's strip (6px top margin + 54px bar) is reserved statically by +-- waybar-reserve, and waybar runs with "exclusive": false. The bar's own +-- exclusive zone would vanish and reappear on every SIGUSR2 reload (the +-- collapse mechanism) and on hide/crash/relaunch, snapping every tiled window +-- up and back down. The static reservation holds the clients in place; only +-- the bar itself changes. `exec` (not exec-once) reruns it on every config +-- reload, which is exactly when Hyprland resets dynamic reservations. The +-- script is idempotent, and a catch-all `monitor=,addreserved,...` rule can't +-- replace it (empty-name addreserved silently no-ops). +-- +-- Run three times over ~0.6s, not once: on reload Hyprland clears the +-- reservation AND re-fires this exec, and the two race. A single run that +-- fires before the clear no-ops (reserved still looks correct), the clear then +-- wins, and the non-exclusive bar drops off-screen. Re-applying past the clear +-- window makes the restore reliable; the script is idempotent so extra runs are +-- free. Applying a monitor rule (e.g. the DP-4 pin) also clears the reservation, +-- so this covers a reload that re-asserts monitors too. +at_reload("for i in 1 2 3; do sleep 0.2; waybar-reserve; done") + +-- ============================================================================ +-- Startup Applications +-- ============================================================================ +-- Portal and D-Bus setup FIRST, then waybar (needs portal for appearance query) +at_start("dbus-update-activation-environment --systemd WAYLAND_DISPLAY XDG_CURRENT_DESKTOP HYPRLAND_INSTANCE_SIGNATURE") +-- Start hyprland-session.target FIRST: it pulls up graphical-session.target, +-- which xdg-desktop-portal 1.22+ hard-requires (Requisite=). A bare-exec Hyprland +-- session has no session manager to raise that target, so without this the portal +-- fails its dependency at every login (screen-share + file pickers dead). +-- 'systemctl start' blocks until active, so the ';' sequence guarantees the target +-- is up before the portal restart runs. +-- +-- Portal restart (not start) reconnects stale portals on Hyprland restart. +-- Backend portals (GTK, Hyprland) restart BEFORE the main portal to avoid a 50s +-- GTK settings proxy timeout; the sequence keeps that ordering. Separated by ';' +-- not '&&' so a failing portal restart can't stop waybar from launching — waybar +-- degrades gracefully without the portal (only the appearance query is missed), +-- and gating the bar behind the portal left the desktop bar-less whenever +-- xdg-desktop-portal failed its dependency at login. Waybar stays gated on its +-- own config generation (waybar-active-config && waybar). +at_start("systemctl --user start hyprland-session.target; systemctl --user restart xdg-desktop-portal-hyprland xdg-desktop-portal-gtk; systemctl --user restart xdg-desktop-portal; waybar-active-config && waybar -c \"$XDG_RUNTIME_DIR/waybar/config\" -s ~/.config/waybar/style.css 2>&1 | grep -v \"LIBDBUSMENU-GLIB-WARNING\" > ~/.local/var/log/waybar-$(date +%Y-%m-%d-%H%M%S).log") + +-- Core services +at_start("/usr/lib/polkit-kde-authentication-agent-1") +at_start("/usr/bin/gnome-keyring-daemon --start --components=pkcs11,secrets,ssh") +at_start("dunst > ~/.local/var/log/dunst-$(date +%Y-%m-%d-%H%M%S).log 2>&1") + +-- Desktop appearance +-- `settings restore` replays the remembered toggles and reapplies the stored +-- wallpaper. It replaced `waypaper --restore` on 2026-08-14: waypaper keeps +-- its own config.ini and the settings store keeps another, neither knew about +-- the other, and the login replay always won — so a wallpaper chosen in the +-- panel came back as whatever the shell had last set. The store is the only +-- one of the two that can hold a sun pair, a video or a projected face, so it +-- owns the restore. set-wallpaper records into it for choices made outside +-- the panel. +-- +-- waypaper --restore stays as the fallback, not the owner. If the stored +-- wallpaper cannot be applied (an image deleted, a drive not mounted yet), +-- `settings restore` exits 3 and waypaper's independent copy still puts +-- something on the screen. Dropping it outright would trade this bug for a +-- bare desktop. +-- +-- The wallpaper half only. The toggle half runs from its own exec-once further +-- down, after hypridle and dunst — caffeine *is* "hypridle isn't running" and +-- DND *is* dunst's pause level, so replaying them here would spend the whole +-- re-assert budget correcting backings that have not launched yet, and would +-- replay them a second time besides. This slot exists for awww's timing, not +-- theirs. +at_start("awww-daemon & sleep 1 && { settings restore-wallpaper || waypaper --restore; }") + +-- Background services +at_start("touchpad-auto") +-- hypridle is reaped on both exit paths, because it outlives its compositor +-- otherwise. An orphaned daemon keeps firing idle actions at whatever session +-- is live next, and it holds its old logind session scope open (the scope can't +-- close while a process sits in it), so orphans accumulate one per abnormal +-- session death. On 2026-07-22 velox reached five concurrent hypridle daemons; +-- two of them racing to lock produced "Cannot re-lock" and a session wedged +-- locked with no client able to draw a password prompt — recoverable only from +-- another console. exec-shutdown covers a clean compositor exit; the pkill in +-- exec-once covers the paths where it never runs (crash, SIGKILL, TTY logout). +at_shutdown("pkill -x hypridle") +-- hypridle.conf is rendered here rather than tracked, because its contents +-- are this machine's stage times and hibernate setting. Tracking the render +-- meant every panel change dirtied the repo, and whichever machine +-- committed last imposed its policy on the others: a desktop ended up +-- carrying a laptop's suspend-then-hibernate line that it cannot run. +-- Rendering at session start makes the store the only source of truth and +-- the file a build artifact. +-- +-- hypridle-start owns the render, the fallback, and the ordering between +-- them, because that ordering is subtle enough to get wrong in a config +-- line nothing can test: a render can fail on purpose (a damaged store, to +-- avoid overwriting a real policy with defaults), and a fallback that +-- fired there would perform exactly the overwrite the render refused. +at_start("pkill -x hypridle; hypridle-start > ~/.local/var/log/hypridle-$(date +%Y-%m-%d-%H%M%S).log 2>&1") +at_start("/usr/lib/geoclue-2.0/demos/agent") +at_start("gammastep > ~/.local/var/log/gammastep-$(date +%Y-%m-%d-%H%M%S).log 2>&1") +at_start("mpd") +-- Replay the toggles that have no durable state of their own. Caffeine *is* +-- "hypridle isn't running" and DND *is* dunst's pause level, so the two +-- exec-once lines above (and dunst's) recreate both at a fixed default every +-- start — a deliberately-set caffeine was silently discarded on every login. +-- Ordered after those launches so it corrects a backing that exists; it also +-- re-asserts for a few seconds, which covers a backing that comes up late. +-- Logged like its neighbours: a silent exec-once failure here would look +-- exactly like the bug it fixes, and gammastep is the standing proof that a +-- launch dying quietly at session start can go unnoticed for a long time. +at_start("settings restore > ~/.local/var/log/settings-restore-$(date +%Y-%m-%d-%H%M%S).log 2>&1") + +-- Pyprland (scratchpads, magnify, etc.) +at_start("pypr > ~/.local/var/log/pypr-$(date +%Y-%m-%d-%H%M%S).log 2>&1") +at_start("hypr-refocus-scratchpad") + +-- Tray apps. wait-for-tray blocks until waybar's systray host is up (a fixed +-- sleep can't cover a slow cold-start waybar), so these register their icons +-- instead of opening as windows. Caps at ~30s, then launches anyway. +at_start("wait-for-tray && signal-desktop --start-in-tray --ozone-platform=wayland") +-- QT_FONT_DPI bumps the bridge's QML UI font (qt6ct General font is ignored by Qt Quick) +at_start("env QT_FONT_DPI=108 protonmail-bridge --no-window") + +-- ============================================================================ +-- Environment Variables +-- ============================================================================ +hl.env("XCURSOR_SIZE", "24") +hl.env("XCURSOR_THEME", "Bibata-Modern-Ice") +hl.env("XDG_CURRENT_DESKTOP", "Hyprland") +hl.env("XDG_SESSION_TYPE", "wayland") +hl.env("XDG_SESSION_DESKTOP", "Hyprland") +hl.env("_JAVA_AWT_WM_NONREPARENTING", "1") + +-- ============================================================================ +-- Appearance (matching DWM colors) +-- ============================================================================ +-- DWM colors: gray1=#222222, gray2=#444444, gray3=#bbbbbb, gray4=#eeeeee, cyan=#daa520 + +hl.config({ + general = { + gaps_in = 25, + gaps_out = 30, + border_size = 2, + col = { + active_border = "rgba(daa520ff)", + inactive_border = "rgba(444444ff)", + -- Pyprland 3.4+ applies `group deny` to scratchpads, which routes + -- their border through col.nogroup_border* instead of col.*_border. + -- Without these overrides Hyprland's defaults paint scratchpads + -- bright magenta. + nogroup_border_active = "rgba(daa520ff)", + nogroup_border = "rgba(444444ff)", + }, + layout = "master", + resize_on_border = true, + }, +}) + +hl.config({ + decoration = { + rounding = 10, + dim_inactive = true, + dim_strength = 0.4, + dim_special = 0.2, + blur = { + enabled = false, + }, + shadow = { + enabled = false, + }, + }, +}) + +hl.config({ + animations = { + enabled = true, + }, +}) + +hl.curve("myBezier", { type = "bezier", points = { { 0.05, 0.9 }, { 0.1, 1.05 } } }) +hl.animation({ + leaf = "windows", + enabled = true, + speed = 2, + bezier = "myBezier", +}) +hl.animation({ + leaf = "windowsOut", + enabled = true, + speed = 2, + bezier = "default", + style = "popin 80%", +}) +hl.animation({ + leaf = "fade", + enabled = true, + speed = 2, + bezier = "default", +}) +hl.animation({ + leaf = "workspaces", + enabled = true, + speed = 2, + bezier = "default", +}) +hl.animation({ + leaf = "specialWorkspace", + enabled = true, + speed = 2, + bezier = "default", + style = "slidevert", +}) + +-- ============================================================================ +-- Layout (master-stack like DWM tile) +-- ============================================================================ + +hl.config({ + master = { + new_status = "master", + new_on_top = true, + mfact = 0.55, + }, +}) + +hl.config({ + dwindle = { + preserve_split = true, + }, +}) + +-- ============================================================================ +-- Input +-- ============================================================================ + +hl.config({ + cursor = { + no_warps = true, + inactive_timeout = 2.0, + }, +}) + +hl.config({ + input = { + kb_layout = "us", + kb_options = "ctrl:nocaps", + numlock_by_default = true, + follow_mouse = 0, + -- 0, not the default 1: with follow_mouse off we never want focus to follow + -- the cursor. At 1, focus still jumps to the window under the pointer when it + -- crosses a floating<->tiled boundary, so launching a floating scratchpad (or + -- the org-capture popup) re-enabled focus-follows-mouse onto tiled windows. + float_switch_override_focus = 0, + mouse_refocus = false, + natural_scroll = true, + touchpad = { + natural_scroll = false, + }, + }, +}) + +-- ============================================================================ +-- Misc +-- ============================================================================ + +hl.config({ + misc = { + force_default_wallpaper = 0, + disable_hyprland_logo = true, + -- false so apps can't pull focus via activation requests. New windows still + -- focus on open (separate path); this stops e.g. a browser yanking focus + -- back off a freshly opened emacs frame. + focus_on_activate = false, + -- Let a fresh lock client adopt a session whose previous one died. The + -- default (off) is the strict reading of ext-session-lock: a dead lock + -- client leaves the session locked forever and refuses every replacement + -- ("Cannot re-lock"), so the screen stays up with nothing able to draw a + -- password prompt and the only way back in is another console. That is a + -- hard lockout, and it cost a session on velox 2026-07-22. On means a + -- replacement hyprlock re-attaches and prompts normally. The screen stays + -- locked either way — this decides whether the lock is recoverable, never + -- whether it holds. + allow_session_lock_restore = true, + }, +}) + +-- ============================================================================ +-- Debug (temporary - disable when stable) +-- ============================================================================ + +hl.config({ + debug = { + disable_logs = false, + }, +}) + +-- ============================================================================ +-- XWayland +-- ============================================================================ + +hl.config({ + xwayland = { + force_zero_scaling = true, + }, +}) + +-- ============================================================================ +-- Window Rules (Hyprland 0.53+ syntax: match:CONDITION, RULE) +-- ============================================================================ +-- Floating windows (from DWM rules) +hl.window_rule({ + match = { + class = "^(xdg-desktop-portal-gtk)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(Gimp)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(caffeine)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(qalculate-gtk)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + title = "^(Event Tester)$", + }, + float = true, +}) + +-- net / bluetooth instrument-console panels. Normal floating windows (formerly +-- gtk4-layer-shell overlays) so they drag to move and corner-drag to resize. +-- Opened top-right to match their old anchored spot: the panel is right-aligned +-- with a 44px gap, so x = 100% - (window width + 44). net is 420 wide, bt 380. +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.netpanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.netpanel)$", + }, + move = "100%-464 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.btpanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.btpanel)$", + }, + move = "100%-424 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.audiopanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.audiopanel)$", + }, + move = "100%-444 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.timerpanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.timerpanel)$", + }, + move = "100%-444 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.settingspanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.settingspanel)$", + }, + move = "100%-584 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.weatherpanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.weatherpanel)$", + }, + move = "100%-464 50", +}) + +-- maintenance console: the wide board (960), same right-aligned convention. +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.maintpanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.maintpanel)$", + }, + move = "100%-1004 50", +}) + +-- org-capture popup frame (quick-capture script names the frame) +-- Size is per-host in <host>/conf.d/local.lua: native window rules ignore +-- percentages (only pyprland honors them), so the popup is sized in absolute +-- pixels matching that host's terminal scratchpad. No size rule here means a +-- host without an override falls back to the script's char-cell geometry. +hl.window_rule({ + match = { + title = "^(org-capture)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + title = "^(org-capture)$", + }, + center = true, +}) + +-- dirvish popup frame (dirvish-popup script names the frame). No stay_focused — +-- it's a file manager that launches files into other apps, so focus must be free +-- to follow; q (cj/dirvish-popup-quit) closes the frame. +hl.window_rule({ + match = { + title = "^(dirvish)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + title = "^(dirvish)$", + }, + size = "1100 700", +}) + +hl.window_rule({ + match = { + title = "^(dirvish)$", + }, + center = true, +}) + +-- NOTE: center windowrules removed 2026-03-04 per pyprland maintainer suggestion +-- Testing whether pyprland handles scratchpad re-centering natively (issue #211) + +-- Gaming +hl.window_rule({ + match = { + class = "^(Civ5XP)$", + }, + fullscreen = true, +}) + +-- ============================================================================ +-- Key Bindings +-- ============================================================================ +local mod = "SUPER" + +-- Terminal and core apps (from DWM) +hl.bind(mod .. " + T", hl.dsp.exec_cmd("foot")) +hl.bind(mod .. " + E", hl.dsp.exec_cmd("emacsclient -c -a \"\" || emacs")) +-- Standalone emacs: its own process, not a frame on the daemon, so killing it +-- takes no other frame with it. init.el guards server-start on server-running-p, +-- so while the daemon holds the socket this process leaves it alone. With no +-- daemon up it finds no server and becomes one -- the guard lives in init.el, and +-- a keybind can't override it, since --eval runs after init. +hl.bind(mod .. " + SHIFT + E", hl.dsp.exec_cmd("emacs")) +hl.bind(mod .. " + N", hl.dsp.exec_cmd("quick-capture")) +hl.bind(mod .. " + W", hl.dsp.exec_cmd("$BROWSER")) +hl.bind(mod .. " + F", hl.dsp.exec_cmd("dirvish-popup")) +hl.bind(mod .. " + SHIFT + F", hl.dsp.exec_cmd("layout-cycle float-toggle")) + +-- From sxhkdrc +hl.bind(mod .. " + SPACE", hl.dsp.exec_cmd("fuzzel-toggle")) +hl.bind(mod .. " + SHIFT + W", hl.dsp.exec_cmd("$ALTBROWSER")) +hl.bind(mod .. " + P", hl.dsp.exec_cmd("media-toggle-all")) +hl.bind(mod .. " + SHIFT + L", hl.dsp.exec_cmd("calibre")) +hl.bind(mod .. " + SHIFT + P", hl.dsp.exec_cmd("toggle-touchpad")) + +-- Window management (from DWM) +-- Layout-aware navigation (works across master, scrolling) +hl.bind(mod .. " + J", hl.dsp.exec_cmd("layout-navigate next")) +hl.bind(mod .. " + K", hl.dsp.exec_cmd("layout-navigate prev")) +hl.bind(mod .. " + SHIFT + J", hl.dsp.exec_cmd("layout-navigate next move")) +hl.bind(mod .. " + SHIFT + K", hl.dsp.exec_cmd("layout-navigate prev move")) +hl.bind(mod .. " + H", hl.dsp.exec_cmd("layout-resize shrink")) +hl.bind(mod .. " + L", hl.dsp.exec_cmd("layout-resize grow")) +-- Swap focused window with master, then force focus onto the master slot. +-- swapwithmaster's own `master` focus param doesn't stick when invoked from +-- the master, so focusmaster master pins focus afterward. The 50ms sleep is +-- load-bearing: swapwithmaster fires an async focus event ~1-2ms after it +-- returns; without the delay that event lands AFTER focusmaster and flips +-- focus back to the detail. The sleep lets the swap's focus settle so +-- focusmaster runs last and wins. Proven via instrumented capture (19/19). +hl.bind(mod .. " + RETURN", hl.dsp.exec_cmd("hyprctl dispatch layoutmsg swapwithmaster && sleep 0.05 && hyprctl dispatch layoutmsg focusmaster master")) +hl.bind(mod .. " + G", hl.dsp.window.center()) +hl.bind(mod .. " + TAB", hl.dsp.focus({ workspace = "previous" })) +hl.bind(mod .. " + SHIFT + C", hl.dsp.window.close()) + +-- Layouts: master -> monocle +-- Cycle with Shift+arrows, or jump directly with Shift+T/M +-- (scrolling layout disabled until frame-fit + wrap-around work lands) +hl.bind(mod .. " + SHIFT + RIGHT", hl.dsp.exec_cmd("layout-cycle next")) +hl.bind(mod .. " + SHIFT + LEFT", hl.dsp.exec_cmd("layout-cycle prev")) +hl.bind(mod .. " + SHIFT + T", hl.dsp.exec_cmd("hyprctl keyword general:layout master && hyprctl keyword master:orientation left")) +hl.bind(mod .. " + SHIFT + M", hl.dsp.exec_cmd("hyprctl keyword general:layout monocle")) +hl.bind(mod .. " + SHIFT + SPACE", hl.dsp.window.float({ action = "toggle" })) + +-- Master layout adjustments +hl.bind(mod .. " + U", hl.dsp.layout("addmaster")) +hl.bind(mod .. " + D", hl.dsp.layout("removemaster")) + +-- Stash windows (hide to special workspace) +-- O = stash focused / Alt+O = stash others / Shift+O = restore all +hl.bind(mod .. " + O", hl.dsp.exec_cmd("stash-window")) +hl.bind(mod .. " + ALT + O", hl.dsp.exec_cmd("stash-others")) +hl.bind(mod .. " + SHIFT + O", hl.dsp.exec_cmd("stash-restore")) + +-- Gaps between windows only; window-gaps leaves the monitor-edge gap +-- (general:gaps_out) fixed, so widening/narrowing moves the space between +-- windows, not the screen-edge margin. +hl.bind(mod .. " + MINUS", hl.dsp.exec_cmd("window-gaps narrow")) +hl.bind(mod .. " + EQUAL", hl.dsp.exec_cmd("window-gaps widen")) +hl.bind(mod .. " + SHIFT + EQUAL", hl.dsp.exec_cmd("window-gaps reset")) +hl.bind(mod .. " + SHIFT + MINUS", hl.dsp.exec_cmd("window-gaps zero")) + +-- Auto-dim toggle (D = dim). Same action as clicking the waybar custom/dim icon. +hl.bind(mod .. " + SHIFT + D", hl.dsp.exec_cmd("dim-toggle")) +hl.bind(mod .. " + SHIFT + G", hl.dsp.exec_cmd("settings-panel")) + +-- Caffeine (keep-awake) toggle. Same action as clicking the waybar +-- custom/caffeine icon — flips the hypridle daemon so the screen will / won't +-- lock. Stays on $mod+I ($mod+C is taken by hyprpicker; no free caffeine key). +hl.bind(mod .. " + I", hl.dsp.exec_cmd("caffeine-toggle")) + +-- Airplane mode (low-power: wifi off + CPU/brightness/services). A deliberate +-- keybind, not a bar click — engaging it disconnects you, so it shouldn't be a +-- misclick away. The custom/net module shows the state; this toggles it. +-- On Super+Shift+X ("X" = everything off); Super+Shift+A toggles push-to-talk. +hl.bind(mod .. " + SHIFT + X", hl.dsp.exec_cmd("airplane-mode")) + +-- Toggle bar visibility, or relaunch waybar if it crashed (no exec-once respawn). +hl.bind(mod .. " + B", hl.dsp.exec_cmd("waybar-toggle")) + +-- Collapse / expand the left or right side of the bar to its base set +-- (same action as clicking the side's arrowhead). [ = left, ] = right. +hl.bind(mod .. " + bracketleft", hl.dsp.exec_cmd("waybar-collapse left")) +hl.bind(mod .. " + bracketright", hl.dsp.exec_cmd("waybar-collapse right")) + +-- Fullscreen +hl.bind(mod .. " + F11", hl.dsp.window.fullscreen({ mode = "fullscreen", action = "toggle" })) + +-- Workspaces 1-9 (from DWM TAGKEYS) +hl.bind(mod .. " + 1", hl.dsp.focus({ workspace = 1 })) +hl.bind(mod .. " + 2", hl.dsp.focus({ workspace = 2 })) +hl.bind(mod .. " + 3", hl.dsp.focus({ workspace = 3 })) +hl.bind(mod .. " + 4", hl.dsp.focus({ workspace = 4 })) +hl.bind(mod .. " + 5", hl.dsp.focus({ workspace = 5 })) +hl.bind(mod .. " + 6", hl.dsp.focus({ workspace = 6 })) +hl.bind(mod .. " + 7", hl.dsp.focus({ workspace = 7 })) +hl.bind(mod .. " + 8", hl.dsp.focus({ workspace = 8 })) +hl.bind(mod .. " + 9", hl.dsp.focus({ workspace = 9 })) + +-- Move window to workspace (from DWM tag) +hl.bind(mod .. " + SHIFT + 1", hl.dsp.window.move({ workspace = 1, follow = false })) +hl.bind(mod .. " + SHIFT + 2", hl.dsp.window.move({ workspace = 2, follow = false })) +hl.bind(mod .. " + SHIFT + 3", hl.dsp.window.move({ workspace = 3, follow = false })) +hl.bind(mod .. " + SHIFT + 4", hl.dsp.window.move({ workspace = 4, follow = false })) +hl.bind(mod .. " + SHIFT + 5", hl.dsp.window.move({ workspace = 5, follow = false })) +hl.bind(mod .. " + SHIFT + 6", hl.dsp.window.move({ workspace = 6, follow = false })) +hl.bind(mod .. " + SHIFT + 7", hl.dsp.window.move({ workspace = 7, follow = false })) +hl.bind(mod .. " + SHIFT + 8", hl.dsp.window.move({ workspace = 8, follow = false })) +hl.bind(mod .. " + SHIFT + 9", hl.dsp.window.move({ workspace = 9, follow = false })) + +-- Monitor focus (from DWM focusmon) +hl.bind(mod .. " + COMMA", hl.dsp.focus({ monitor = "-1" })) +hl.bind(mod .. " + PERIOD", hl.dsp.focus({ monitor = "+1" })) +hl.bind(mod .. " + SHIFT + COMMA", hl.dsp.window.move({ monitor = "-1" })) +hl.bind(mod .. " + SHIFT + PERIOD", hl.dsp.window.move({ monitor = "+1" })) + +-- ============================================================================ +-- Scratchpads (via pyprland) +-- ============================================================================ +-- Configured in ~/.config/hypr/pyprland.toml +-- Uses normal workspaces (not special), so new windows won't be captured +hl.bind(mod .. " + SHIFT + RETURN", hl.dsp.exec_cmd("pypr toggle term")) +hl.bind(mod .. " + A", hl.dsp.exec_cmd("audio-panel")) +hl.bind(mod .. " + R", hl.dsp.exec_cmd("pypr toggle monitor")) +hl.bind(mod .. " + SHIFT + N", hl.dsp.exec_cmd("net panel")) +hl.bind(mod .. " + SLASH", hl.dsp.exec_cmd("pypr toggle music")) + +-- Magnify (zoom) +-- mod+Z zooms and enters the "zoom" submap; inside it, Escape or mod+Z +-- unzooms and returns to the normal keymap. Exit forces `pypr zoom 1` +-- (factor 1) so submap state and zoom state can't desync. Note: while +-- zoomed, other Hyprland binds pause until you exit the submap. +hl.bind(mod .. " + Z", hl.dsp.exec_cmd("pypr zoom")) +hl.bind(mod .. " + Z", hl.dsp.submap("zoom")) + +hl.define_submap("zoom", function() + hl.bind("ESCAPE", hl.dsp.exec_cmd("pypr zoom 1")) + hl.bind("ESCAPE", hl.dsp.submap("reset")) + hl.bind(mod .. " + Z", hl.dsp.exec_cmd("pypr zoom 1")) + hl.bind(mod .. " + Z", hl.dsp.submap("reset")) +end) + +-- Calculator (not a scratchpad, just launches app) +hl.bind(mod .. " + X", hl.dsp.exec_cmd("calc-toggle")) +hl.bind(mod .. " + C", hl.dsp.exec_cmd("hyprpicker -a")) +hl.bind(mod .. " + CONTROL + C", hl.dsp.exec_cmd("clock-panel toggle")) + +-- Media/hardware keys +hl.bind("XF86AudioRaiseVolume", hl.dsp.exec_cmd("pactl set-sink-volume @DEFAULT_SINK@ +5%"), { locked = true, repeating = true }) +hl.bind("XF86AudioLowerVolume", hl.dsp.exec_cmd("pactl set-sink-volume @DEFAULT_SINK@ -5%"), { locked = true, repeating = true }) +hl.bind("XF86AudioMute", hl.dsp.exec_cmd("audio quick-mute"), { locked = true }) +hl.bind("XF86MonBrightnessUp", hl.dsp.exec_cmd("brightnessctl s +10%"), { locked = true, repeating = true }) +hl.bind("XF86MonBrightnessDown", hl.dsp.exec_cmd("brightnessctl s 10%-"), { locked = true, repeating = true }) + +-- Microphone mute toggle (waybar pulseaudio#mic indicator follows via PipeWire events). +-- On the hardware mic-mute key. Super+Shift+A used to duplicate this; it now +-- toggles push-to-talk mode instead (mic-toggle stays reachable on the hw key). +hl.bind("XF86AudioMicMute", hl.dsp.exec_cmd("mic-toggle"), { locked = true }) + +-- Push-to-talk toggle: enter PTT (mic muted, hold key armed) / exit (restore). +-- The hold key is configurable (audio config ptt_key = Control_R or mouse:NNN). +hl.bind(mod .. " + SHIFT + A", hl.dsp.exec_cmd("audio ptt-toggle")) + +-- Bluetooth panel (blueman retired in favor of the bt panel) +hl.bind(mod .. " + SHIFT + B", hl.dsp.exec_cmd("bt-panel")) + +-- Screenshots (grim + slurp + fuzzel menu) +-- Shift+S captures the whole desktop with no pointer interaction, so it +-- works on scratchpads and popups that region-select would dismiss +hl.bind(mod .. " + S", hl.dsp.exec_cmd("screenshot region")) +hl.bind(mod .. " + SHIFT + S", hl.dsp.exec_cmd("screenshot fullscreen")) +hl.bind("CTRL + " .. mod .. " + S", hl.dsp.exec_cmd("screenshot fullscreen")) + +-- Lock screen +hl.bind(mod .. " + ESCAPE", hl.dsp.exec_cmd("hyprlock")) + +-- Audio mute cycle (M for Mute): one key walks volume/mic through all four +-- on/off combinations. The waybar pulseaudio modules follow via PipeWire. +hl.bind(mod .. " + M", hl.dsp.exec_cmd("audio-cycle")) + +-- Exit/session +hl.bind(mod .. " + SHIFT + Q", hl.dsp.exec_cmd("pgrep -x wlogout || wlogout-menu")) +-- mod+Shift+Backspace no longer exits outright — it enters the "exitconfirm" +-- submap, where a second Backspace confirms and anything else backs out. +-- The bare bind killed two live sessions on 2026-07-22: it sits one key from +-- the mod+Shift cluster (Q wlogout, C killactive, Return terminal, G settings) +-- and took the whole session down with no prompt. wlogout owns the normal +-- exit path; this stays as the keyboard escape hatch for when wlogout won't +-- come up, so it's gated rather than removed. +hl.bind(mod .. " + SHIFT + BACKSPACE", hl.dsp.submap("exitconfirm")) + +hl.define_submap("exitconfirm", function() + hl.bind("BACKSPACE", hl.dsp.exit()) + hl.bind("ESCAPE", hl.dsp.submap("reset")) + hl.bind("catchall", hl.dsp.submap("reset")) +end) + +hl.bind(mod .. " + SHIFT + ESCAPE", hl.dsp.exec_cmd("hyprctl reload")) +hl.bind("CTRL + ALT + " .. mod .. " + K", hl.dsp.exec_cmd("hyprctl kill")) + +-- Mouse bindings (from DWM buttons) +hl.bind(mod .. " + mouse:272", hl.dsp.window.drag(), { mouse = true }) +hl.bind(mod .. " + mouse:273", hl.dsp.window.resize(), { mouse = true }) +hl.bind(mod .. " + SHIFT + mouse:272", hl.dsp.window.resize(), { mouse = true }) + +-- ============================================================================ +-- Machine-local overrides +-- ============================================================================ +-- Sourced last so machine-specific settings (monitor scale, gaps, keybinds) +-- override the defaults above. See conf.d/local.lua. + +-- Replay the collected autostart commands. +hl.on("hyprland.start", function() + for _, cmd in ipairs(autostart) do hl.exec_cmd(cmd) end +end) + +hl.on("hyprland.shutdown", function() + for _, cmd in ipairs(atshutdown) do hl.exec_cmd(cmd) end +end) + +-- hyprlang's `exec` ran at startup AND on every reload. config.reloaded fires on +-- the initial load too (verified in a nested Hyprland 0.56.2 on 2026-08-24: a +-- marker written from this handler appears at session start), so this single +-- handler covers both cases, which is what waybar-reserve needs. +hl.on("config.reloaded", function() + for _, cmd in ipairs(atreload) do hl.exec_cmd(cmd) end +end) + +hl_source_glob("$HOME/.config/hypr/conf.d/*.lua") diff --git a/working/hyprland-lua-port/hyprland.lua.generated b/working/hyprland-lua-port/hyprland.lua.generated new file mode 100644 index 0000000..04b30b9 --- /dev/null +++ b/working/hyprland-lua-port/hyprland.lua.generated @@ -0,0 +1,674 @@ +-- Hyprland Configuration +-- Translated from DWM config.def.h and sxhkdrc +-- Craig Jennings <c@cjennings.net> + +-- ============================================================================ +-- Monitor Configuration +-- ============================================================================ +-- Generated by hyprlang2lua. Review TODOs before reloading Hyprland. + +-- hyprlang2lua polyfills — runtime helpers reproducing +-- hyprlang behaviour the typed Lua API doesn't expose directly. + +local function hl_source_glob(pattern) + -- 'source = path/*.conf' had hyprlang glob and inline-expand the + -- matches. require() can't glob, so we shell out to ls (matching + -- the user's brace-expansion behaviour) and dofile each result. + -- Paths with spaces or shell metacharacters in the directory + -- portion will misparse; typical ~/.config/hypr/ layouts don't + -- hit this. Swap to lfs.dir() or find -name if you need fancier. + local p = io.popen("ls " .. pattern .. " 2>/dev/null") + if not p then return end + for f in p:lines() do + local chunk, err = loadfile(f) + if chunk then chunk() + else io.stderr:write("hl_source_glob: " .. tostring(err) .. "\n") end + end + p:close() +end + +hl.monitor({ + output = "", + mode = "preferred", + position = "auto", + scale = "auto", +}) + +-- Waybar's strip (6px top margin + 54px bar) is reserved statically by +-- waybar-reserve, and waybar runs with "exclusive": false. The bar's own +-- exclusive zone would vanish and reappear on every SIGUSR2 reload (the +-- collapse mechanism) and on hide/crash/relaunch, snapping every tiled window +-- up and back down. The static reservation holds the clients in place; only +-- the bar itself changes. `exec` (not exec-once) reruns it on every config +-- reload, which is exactly when Hyprland resets dynamic reservations. The +-- script is idempotent, and a catch-all `monitor=,addreserved,...` rule can't +-- replace it (empty-name addreserved silently no-ops). +-- +-- Run three times over ~0.6s, not once: on reload Hyprland clears the +-- reservation AND re-fires this exec, and the two race. A single run that +-- fires before the clear no-ops (reserved still looks correct), the clear then +-- wins, and the non-exclusive bar drops off-screen. Re-applying past the clear +-- window makes the restore reliable; the script is idempotent so extra runs are +-- free. Applying a monitor rule (e.g. the DP-4 pin) also clears the reservation, +-- so this covers a reload that re-asserts monitors too. + +-- ============================================================================ +-- Startup Applications +-- ============================================================================ +-- Portal and D-Bus setup FIRST, then waybar (needs portal for appearance query) +-- Start hyprland-session.target FIRST: it pulls up graphical-session.target, +-- which xdg-desktop-portal 1.22+ hard-requires (Requisite=). A bare-exec Hyprland +-- session has no session manager to raise that target, so without this the portal +-- fails its dependency at every login (screen-share + file pickers dead). +-- 'systemctl start' blocks until active, so the ';' sequence guarantees the target +-- is up before the portal restart runs. +-- +-- Portal restart (not start) reconnects stale portals on Hyprland restart. +-- Backend portals (GTK, Hyprland) restart BEFORE the main portal to avoid a 50s +-- GTK settings proxy timeout; the sequence keeps that ordering. Separated by ';' +-- not '&&' so a failing portal restart can't stop waybar from launching — waybar +-- degrades gracefully without the portal (only the appearance query is missed), +-- and gating the bar behind the portal left the desktop bar-less whenever +-- xdg-desktop-portal failed its dependency at login. Waybar stays gated on its +-- own config generation (waybar-active-config && waybar). + +-- Core services + +-- Desktop appearance +-- `settings restore` replays the remembered toggles and reapplies the stored +-- wallpaper. It replaced `waypaper --restore` on 2026-08-14: waypaper keeps +-- its own config.ini and the settings store keeps another, neither knew about +-- the other, and the login replay always won — so a wallpaper chosen in the +-- panel came back as whatever the shell had last set. The store is the only +-- one of the two that can hold a sun pair, a video or a projected face, so it +-- owns the restore. set-wallpaper records into it for choices made outside +-- the panel. +-- +-- waypaper --restore stays as the fallback, not the owner. If the stored +-- wallpaper cannot be applied (an image deleted, a drive not mounted yet), +-- `settings restore` exits 3 and waypaper's independent copy still puts +-- something on the screen. Dropping it outright would trade this bug for a +-- bare desktop. +-- +-- The wallpaper half only. The toggle half runs from its own exec-once further +-- down, after hypridle and dunst — caffeine *is* "hypridle isn't running" and +-- DND *is* dunst's pause level, so replaying them here would spend the whole +-- re-assert budget correcting backings that have not launched yet, and would +-- replay them a second time besides. This slot exists for awww's timing, not +-- theirs. + +-- Background services +-- hypridle is reaped on both exit paths, because it outlives its compositor +-- otherwise. An orphaned daemon keeps firing idle actions at whatever session +-- is live next, and it holds its old logind session scope open (the scope can't +-- close while a process sits in it), so orphans accumulate one per abnormal +-- session death. On 2026-07-22 velox reached five concurrent hypridle daemons; +-- two of them racing to lock produced "Cannot re-lock" and a session wedged +-- locked with no client able to draw a password prompt — recoverable only from +-- another console. exec-shutdown covers a clean compositor exit; the pkill in +-- exec-once covers the paths where it never runs (crash, SIGKILL, TTY logout). +hl.on("hyprland.shutdown", function() + hl.exec_cmd("pkill -x hypridle") +end) + +-- hypridle.conf is rendered here rather than tracked, because its contents +-- are this machine's stage times and hibernate setting. Tracking the render +-- meant every panel change dirtied the repo, and whichever machine +-- committed last imposed its policy on the others: a desktop ended up +-- carrying a laptop's suspend-then-hibernate line that it cannot run. +-- Rendering at session start makes the store the only source of truth and +-- the file a build artifact. +-- +-- hypridle-start owns the render, the fallback, and the ordering between +-- them, because that ordering is subtle enough to get wrong in a config +-- line nothing can test: a render can fail on purpose (a damaged store, to +-- avoid overwriting a real policy with defaults), and a fallback that +-- fired there would perform exactly the overwrite the render refused. +-- Replay the toggles that have no durable state of their own. Caffeine *is* +-- "hypridle isn't running" and DND *is* dunst's pause level, so the two +-- exec-once lines above (and dunst's) recreate both at a fixed default every +-- start — a deliberately-set caffeine was silently discarded on every login. +-- Ordered after those launches so it corrects a backing that exists; it also +-- re-asserts for a few seconds, which covers a backing that comes up late. +-- Logged like its neighbours: a silent exec-once failure here would look +-- exactly like the bug it fixes, and gammastep is the standing proof that a +-- launch dying quietly at session start can go unnoticed for a long time. + +-- Pyprland (scratchpads, magnify, etc.) + +-- Tray apps. wait-for-tray blocks until waybar's systray host is up (a fixed +-- sleep can't cover a slow cold-start waybar), so these register their icons +-- instead of opening as windows. Caps at ~30s, then launches anyway. +-- QT_FONT_DPI bumps the bridge's QML UI font (qt6ct General font is ignored by Qt Quick) + +-- ============================================================================ +-- Environment Variables +-- ============================================================================ +hl.env("XCURSOR_SIZE", "24") +hl.env("XCURSOR_THEME", "Bibata-Modern-Ice") +hl.env("XDG_CURRENT_DESKTOP", "Hyprland") +hl.env("XDG_SESSION_TYPE", "wayland") +hl.env("XDG_SESSION_DESKTOP", "Hyprland") +hl.env("_JAVA_AWT_WM_NONREPARENTING", "1") + +-- ============================================================================ +-- Appearance (matching DWM colors) +-- ============================================================================ +-- DWM colors: gray1=#222222, gray2=#444444, gray3=#bbbbbb, gray4=#eeeeee, cyan=#daa520 + +hl.curve("myBezier", { type = "bezier", points = { { 0.05, 0.9 }, { 0.1, 1.05 } } }) +hl.animation({ + leaf = "windows", + enabled = true, + speed = 2, + bezier = "myBezier", +}) +hl.animation({ + leaf = "windowsOut", + enabled = true, + speed = 2, + bezier = "default", + style = "popin 80%", +}) +hl.animation({ + leaf = "fade", + enabled = true, + speed = 2, + bezier = "default", +}) +hl.animation({ + leaf = "workspaces", + enabled = true, + speed = 2, + bezier = "default", +}) +hl.animation({ + leaf = "specialWorkspace", + enabled = true, + speed = 2, + bezier = "default", + style = "slidevert", +}) + +hl.window_rule({ + match = { + class = "^(xdg-desktop-portal-gtk)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(Gimp)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(caffeine)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(qalculate-gtk)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + title = "^(Event Tester)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.netpanel)$", + }, + float = true, + move = "100%-464 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.btpanel)$", + }, + float = true, + move = "100%-424 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.audiopanel)$", + }, + float = true, + move = "100%-444 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.timerpanel)$", + }, + float = true, + move = "100%-444 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.settingspanel)$", + }, + float = true, + move = "100%-584 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.weatherpanel)$", + }, + float = true, + move = "100%-464 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.maintpanel)$", + }, + float = true, + move = "100%-1004 50", +}) + +hl.window_rule({ + match = { + title = "^(org-capture)$", + }, + float = true, + center = true, +}) + +hl.window_rule({ + match = { + title = "^(dirvish)$", + }, + float = true, + size = "1100 700", + center = true, +}) + +hl.window_rule({ + match = { + class = "^(Civ5XP)$", + }, + fullscreen = true, +}) + +local mod = "SUPER" + +hl.bind(mod .. " + T", hl.dsp.exec_cmd("foot")) +hl.bind(mod .. " + E", hl.dsp.exec_cmd("emacsclient -c -a \"\" || emacs")) +hl.bind(mod .. " + SHIFT + E", hl.dsp.exec_cmd("emacs")) +hl.bind(mod .. " + N", hl.dsp.exec_cmd("quick-capture")) +hl.bind(mod .. " + W", hl.dsp.exec_cmd("$BROWSER")) +hl.bind(mod .. " + F", hl.dsp.exec_cmd("dirvish-popup")) +hl.bind(mod .. " + SHIFT + F", hl.dsp.exec_cmd("layout-cycle float-toggle")) + +hl.bind(mod .. " + SPACE", hl.dsp.exec_cmd("fuzzel-toggle")) +hl.bind(mod .. " + SHIFT + W", hl.dsp.exec_cmd("$ALTBROWSER")) +hl.bind(mod .. " + P", hl.dsp.exec_cmd("media-toggle-all")) +hl.bind(mod .. " + SHIFT + L", hl.dsp.exec_cmd("calibre")) +hl.bind(mod .. " + SHIFT + P", hl.dsp.exec_cmd("toggle-touchpad")) + +hl.bind(mod .. " + J", hl.dsp.exec_cmd("layout-navigate next")) +hl.bind(mod .. " + K", hl.dsp.exec_cmd("layout-navigate prev")) +hl.bind(mod .. " + SHIFT + J", hl.dsp.exec_cmd("layout-navigate next move")) +hl.bind(mod .. " + SHIFT + K", hl.dsp.exec_cmd("layout-navigate prev move")) +hl.bind(mod .. " + H", hl.dsp.exec_cmd("layout-resize shrink")) +hl.bind(mod .. " + L", hl.dsp.exec_cmd("layout-resize grow")) +hl.bind(mod .. " + RETURN", hl.dsp.layout("swapwithmaster && sleep 0.05 && hyprctl dispatch layoutmsg focusmaster master")) +hl.bind(mod .. " + G", hl.dsp.window.center()) +hl.bind(mod .. " + TAB", hl.dsp.focus({ workspace = "previous" })) +hl.bind(mod .. " + SHIFT + C", hl.dsp.window.close()) + +hl.bind(mod .. " + SHIFT + RIGHT", hl.dsp.exec_cmd("layout-cycle next")) +hl.bind(mod .. " + SHIFT + LEFT", hl.dsp.exec_cmd("layout-cycle prev")) +hl.bind(mod .. " + SHIFT + T", hl.dsp.exec_cmd("hyprctl keyword general:layout master && hyprctl keyword master:orientation left")) +hl.bind(mod .. " + SHIFT + M", hl.dsp.exec_cmd("hyprctl keyword general:layout monocle")) +hl.bind(mod .. " + SHIFT + SPACE", hl.dsp.window.float({ action = "toggle" })) + +hl.bind(mod .. " + U", hl.dsp.layout("addmaster")) +hl.bind(mod .. " + D", hl.dsp.layout("removemaster")) + +hl.bind(mod .. " + O", hl.dsp.exec_cmd("stash-window")) +hl.bind(mod .. " + ALT + O", hl.dsp.exec_cmd("stash-others")) +hl.bind(mod .. " + SHIFT + O", hl.dsp.exec_cmd("stash-restore")) + +hl.bind(mod .. " + MINUS", hl.dsp.exec_cmd("window-gaps narrow")) +hl.bind(mod .. " + EQUAL", hl.dsp.exec_cmd("window-gaps widen")) +hl.bind(mod .. " + SHIFT + EQUAL", hl.dsp.exec_cmd("window-gaps reset")) +hl.bind(mod .. " + SHIFT + MINUS", hl.dsp.exec_cmd("window-gaps zero")) + +hl.bind(mod .. " + SHIFT + D", hl.dsp.exec_cmd("dim-toggle")) +hl.bind(mod .. " + SHIFT + G", hl.dsp.exec_cmd("settings-panel")) + +hl.bind(mod .. " + I", hl.dsp.exec_cmd("caffeine-toggle")) + +hl.bind(mod .. " + SHIFT + X", hl.dsp.exec_cmd("airplane-mode")) + +hl.bind(mod .. " + B", hl.dsp.exec_cmd("waybar-toggle")) + +hl.bind(mod .. " + bracketleft", hl.dsp.exec_cmd("waybar-collapse left")) +hl.bind(mod .. " + bracketright", hl.dsp.exec_cmd("waybar-collapse right")) + +hl.bind(mod .. " + F11", hl.dsp.window.fullscreen({ mode = "fullscreen", action = "toggle" })) + +hl.bind(mod .. " + 1", hl.dsp.focus({ workspace = 1 })) +hl.bind(mod .. " + 2", hl.dsp.focus({ workspace = 2 })) +hl.bind(mod .. " + 3", hl.dsp.focus({ workspace = 3 })) +hl.bind(mod .. " + 4", hl.dsp.focus({ workspace = 4 })) +hl.bind(mod .. " + 5", hl.dsp.focus({ workspace = 5 })) +hl.bind(mod .. " + 6", hl.dsp.focus({ workspace = 6 })) +hl.bind(mod .. " + 7", hl.dsp.focus({ workspace = 7 })) +hl.bind(mod .. " + 8", hl.dsp.focus({ workspace = 8 })) +hl.bind(mod .. " + 9", hl.dsp.focus({ workspace = 9 })) + +hl.bind(mod .. " + SHIFT + 1", hl.dsp.window.move({ workspace = 1, follow = false })) +hl.bind(mod .. " + SHIFT + 2", hl.dsp.window.move({ workspace = 2, follow = false })) +hl.bind(mod .. " + SHIFT + 3", hl.dsp.window.move({ workspace = 3, follow = false })) +hl.bind(mod .. " + SHIFT + 4", hl.dsp.window.move({ workspace = 4, follow = false })) +hl.bind(mod .. " + SHIFT + 5", hl.dsp.window.move({ workspace = 5, follow = false })) +hl.bind(mod .. " + SHIFT + 6", hl.dsp.window.move({ workspace = 6, follow = false })) +hl.bind(mod .. " + SHIFT + 7", hl.dsp.window.move({ workspace = 7, follow = false })) +hl.bind(mod .. " + SHIFT + 8", hl.dsp.window.move({ workspace = 8, follow = false })) +hl.bind(mod .. " + SHIFT + 9", hl.dsp.window.move({ workspace = 9, follow = false })) + +hl.bind(mod .. " + COMMA", hl.dsp.focus({ monitor = -1 })) +hl.bind(mod .. " + PERIOD", hl.dsp.focus({ monitor = "+1" })) +hl.bind(mod .. " + SHIFT + COMMA", hl.dsp.window.move({ monitor = "-1" })) +hl.bind(mod .. " + SHIFT + PERIOD", hl.dsp.window.move({ monitor = "+1" })) + +hl.bind(mod .. " + SHIFT + RETURN", hl.dsp.exec_cmd("pypr toggle term")) +hl.bind(mod .. " + A", hl.dsp.exec_cmd("audio-panel")) +hl.bind(mod .. " + R", hl.dsp.exec_cmd("pypr toggle monitor")) +hl.bind(mod .. " + SHIFT + N", hl.dsp.exec_cmd("net panel")) +hl.bind(mod .. " + SLASH", hl.dsp.exec_cmd("pypr toggle music")) + +hl.bind(mod .. " + Z", hl.dsp.exec_cmd("pypr zoom")) +hl.bind(mod .. " + Z", hl.dsp.submap("zoom")) + +hl.define_submap("zoom", function() + hl.bind("ESCAPE", hl.dsp.exec_cmd("pypr zoom 1")) + hl.bind("ESCAPE", hl.dsp.submap("reset")) + hl.bind(mod .. " + Z", hl.dsp.exec_cmd("pypr zoom 1")) + hl.bind(mod .. " + Z", hl.dsp.submap("reset")) +end) + +hl.bind(mod .. " + X", hl.dsp.exec_cmd("calc-toggle")) +hl.bind(mod .. " + C", hl.dsp.exec_cmd("hyprpicker -a")) +hl.bind(mod .. " + CONTROL + C", hl.dsp.exec_cmd("clock-panel toggle")) + +hl.bind("XF86AudioRaiseVolume", hl.dsp.exec_cmd("pactl set-sink-volume @DEFAULT_SINK@ +5%"), { locked = true, repeating = true }) +hl.bind("XF86AudioLowerVolume", hl.dsp.exec_cmd("pactl set-sink-volume @DEFAULT_SINK@ -5%"), { locked = true, repeating = true }) +hl.bind("XF86AudioMute", hl.dsp.exec_cmd("audio quick-mute"), { locked = true }) +hl.bind("XF86MonBrightnessUp", hl.dsp.exec_cmd("brightnessctl s +10%"), { locked = true, repeating = true }) +hl.bind("XF86MonBrightnessDown", hl.dsp.exec_cmd("brightnessctl s 10%-"), { locked = true, repeating = true }) + +hl.bind("XF86AudioMicMute", hl.dsp.exec_cmd("mic-toggle"), { locked = true }) + +hl.bind(mod .. " + SHIFT + A", hl.dsp.exec_cmd("audio ptt-toggle")) + +hl.bind(mod .. " + SHIFT + B", hl.dsp.exec_cmd("bt-panel")) + +hl.bind(mod .. " + S", hl.dsp.exec_cmd("screenshot region")) +hl.bind(mod .. " + SHIFT + S", hl.dsp.exec_cmd("screenshot fullscreen")) +hl.bind("CTRL" .. mod .. " + S", hl.dsp.exec_cmd("screenshot fullscreen")) + +hl.bind(mod .. " + ESCAPE", hl.dsp.exec_cmd("hyprlock")) + +hl.bind(mod .. " + M", hl.dsp.exec_cmd("audio-cycle")) + +hl.bind(mod .. " + SHIFT + Q", hl.dsp.exec_cmd("pgrep -x wlogout || wlogout-menu")) +hl.bind(mod .. " + SHIFT + BACKSPACE", hl.dsp.submap("exitconfirm")) + +hl.define_submap("exitconfirm", function() + hl.bind("BACKSPACE", hl.dsp.exit()) + hl.bind("ESCAPE", hl.dsp.submap("reset")) + hl.bind("catchall", hl.dsp.submap("reset")) +end) + +hl.bind(mod .. " + SHIFT + ESCAPE", hl.dsp.exec_cmd("hyprctl reload")) +hl.bind("CTRL + ALT" .. mod .. " + K", hl.dsp.exec_cmd("hyprctl kill")) + +hl.bind(mod .. " + mouse:272", hl.dsp.window.drag()) +hl.bind(mod .. " + mouse:273", hl.dsp.window.resize()) +hl.bind(mod .. " + SHIFT + mouse:272", hl.dsp.window.resize()) + +-- Source: $HOME/.config/hypr/conf.d/*.conf (glob; resolved at runtime). Each matched .conf must be converted to .lua. +hl_source_glob("$HOME/.config/hypr/conf.d/*.lua") +hl.config({ + general = { + gaps_in = 25, + gaps_out = 30, + border_size = 2, + col = { + active_border = "rgba(daa520ff)", + inactive_border = "rgba(444444ff)", + nogroup_border_active = "rgba(daa520ff)", + nogroup_border = "rgba(444444ff)", + }, + -- Pyprland 3.4+ applies `group deny` to scratchpads, which routes their + -- border through col.nogroup_border* instead of col.*_border. Without + -- these overrides Hyprland's defaults paint scratchpads bright magenta. + layout = "master", + resize_on_border = true, + }, + decoration = { + rounding = 10, + dim_inactive = true, + dim_strength = 0.4, + dim_special = 0.2, + blur = { + enabled = false, + }, + shadow = { + enabled = false, + }, + }, + animations = { + enabled = true, + }, + -- ============================================================================ + -- Layout (master-stack like DWM tile) + -- ============================================================================ + master = { + new_status = "master", + new_on_top = true, + mfact = 0.55, + }, + dwindle = { + preserve_split = true, + }, + -- ============================================================================ + -- Input + -- ============================================================================ + cursor = { + no_warps = true, + inactive_timeout = 2.0, + }, + input = { + kb_layout = "us", + kb_options = "ctrl:nocaps", + numlock_by_default = true, + follow_mouse = 0, + -- 0, not the default 1: with follow_mouse off we never want focus to follow + -- the cursor. At 1, focus still jumps to the window under the pointer when it + -- crosses a floating<->tiled boundary, so launching a floating scratchpad (or + -- the org-capture popup) re-enabled focus-follows-mouse onto tiled windows. + float_switch_override_focus = 0, + mouse_refocus = false, + natural_scroll = true, + touchpad = { + natural_scroll = false, + }, + }, + -- ============================================================================ + -- Misc + -- ============================================================================ + misc = { + force_default_wallpaper = 0, + disable_hyprland_logo = true, + -- false so apps can't pull focus via activation requests. New windows still + -- focus on open (separate path); this stops e.g. a browser yanking focus + -- back off a freshly opened emacs frame. + focus_on_activate = false, + -- Let a fresh lock client adopt a session whose previous one died. The + -- default (off) is the strict reading of ext-session-lock: a dead lock + -- client leaves the session locked forever and refuses every replacement + -- ("Cannot re-lock"), so the screen stays up with nothing able to draw a + -- password prompt and the only way back in is another console. That is a + -- hard lockout, and it cost a session on velox 2026-07-22. On means a + -- replacement hyprlock re-attaches and prompts normally. The screen stays + -- locked either way — this decides whether the lock is recoverable, never + -- whether it holds. + allow_session_lock_restore = true, + }, + -- ============================================================================ + -- Debug (temporary - disable when stable) + -- ============================================================================ + debug = { + disable_logs = false, + }, + -- ============================================================================ + -- XWayland + -- ============================================================================ + xwayland = { + force_zero_scaling = true, + }, + -- ============================================================================ + -- Window Rules (Hyprland 0.53+ syntax: match:CONDITION, RULE) + -- ============================================================================ + -- Floating windows (from DWM rules) + -- net / bluetooth instrument-console panels. Normal floating windows (formerly + -- gtk4-layer-shell overlays) so they drag to move and corner-drag to resize. + -- Opened top-right to match their old anchored spot: the panel is right-aligned + -- with a 44px gap, so x = 100% - (window width + 44). net is 420 wide, bt 380. + -- maintenance console: the wide board (960), same right-aligned convention. + -- org-capture popup frame (quick-capture script names the frame) + -- Size is per-host in <host>/conf.d/local.conf: native window rules ignore + -- percentages (only pyprland honors them), so the popup is sized in absolute + -- pixels matching that host's terminal scratchpad. No size rule here means a + -- host without an override falls back to the script's char-cell geometry. + -- dirvish popup frame (dirvish-popup script names the frame). No stay_focused — + -- it's a file manager that launches files into other apps, so focus must be free + -- to follow; q (cj/dirvish-popup-quit) closes the frame. + -- NOTE: center windowrules removed 2026-03-04 per pyprland maintainer suggestion + -- Testing whether pyprland handles scratchpad re-centering natively (issue #211) + -- Gaming + -- ============================================================================ + -- Key Bindings + -- ============================================================================ + -- Terminal and core apps (from DWM) + -- Standalone emacs: its own process, not a frame on the daemon, so killing it + -- takes no other frame with it. init.el guards server-start on server-running-p, + -- so while the daemon holds the socket this process leaves it alone. With no + -- daemon up it finds no server and becomes one -- the guard lives in init.el, and + -- a keybind can't override it, since --eval runs after init. + -- From sxhkdrc + -- Window management (from DWM) + -- Layout-aware navigation (works across master, scrolling) + -- Swap focused window with master, then force focus onto the master slot. + -- swapwithmaster's own `master` focus param doesn't stick when invoked from + -- the master, so focusmaster master pins focus afterward. The 50ms sleep is + -- load-bearing: swapwithmaster fires an async focus event ~1-2ms after it + -- returns; without the delay that event lands AFTER focusmaster and flips + -- focus back to the detail. The sleep lets the swap's focus settle so + -- focusmaster runs last and wins. Proven via instrumented capture (19/19). + -- Layouts: master -> monocle + -- Cycle with Shift+arrows, or jump directly with Shift+T/M + -- (scrolling layout disabled until frame-fit + wrap-around work lands) + -- Master layout adjustments + -- Stash windows (hide to special workspace) + -- O = stash focused / Alt+O = stash others / Shift+O = restore all + -- Gaps between windows only; window-gaps leaves the monitor-edge gap + -- (general:gaps_out) fixed, so widening/narrowing moves the space between + -- windows, not the screen-edge margin. + -- Auto-dim toggle (D = dim). Same action as clicking the waybar custom/dim icon. + -- Caffeine (keep-awake) toggle. Same action as clicking the waybar + -- custom/caffeine icon — flips the hypridle daemon so the screen will / won't + -- lock. Stays on $mod+I ($mod+C is taken by hyprpicker; no free caffeine key). + -- Airplane mode (low-power: wifi off + CPU/brightness/services). A deliberate + -- keybind, not a bar click — engaging it disconnects you, so it shouldn't be a + -- misclick away. The custom/net module shows the state; this toggles it. + -- On Super+Shift+X ("X" = everything off); Super+Shift+A toggles push-to-talk. + -- Toggle bar visibility, or relaunch waybar if it crashed (no exec-once respawn). + -- Collapse / expand the left or right side of the bar to its base set + -- (same action as clicking the side's arrowhead). [ = left, ] = right. + -- Fullscreen + -- Workspaces 1-9 (from DWM TAGKEYS) + -- Move window to workspace (from DWM tag) + -- Monitor focus (from DWM focusmon) + -- ============================================================================ + -- Scratchpads (via pyprland) + -- ============================================================================ + -- Configured in ~/.config/hypr/pyprland.toml + -- Uses normal workspaces (not special), so new windows won't be captured + -- Magnify (zoom) + -- mod+Z zooms and enters the "zoom" submap; inside it, Escape or mod+Z + -- unzooms and returns to the normal keymap. Exit forces `pypr zoom 1` + -- (factor 1) so submap state and zoom state can't desync. Note: while + -- zoomed, other Hyprland binds pause until you exit the submap. + -- Calculator (not a scratchpad, just launches app) + -- Media/hardware keys + -- Microphone mute toggle (waybar pulseaudio#mic indicator follows via PipeWire events). + -- On the hardware mic-mute key. Super+Shift+A used to duplicate this; it now + -- toggles push-to-talk mode instead (mic-toggle stays reachable on the hw key). + -- Push-to-talk toggle: enter PTT (mic muted, hold key armed) / exit (restore). + -- The hold key is configurable (audio config ptt_key = Control_R or mouse:NNN). + -- Bluetooth panel (blueman retired in favor of the bt panel) + -- Screenshots (grim + slurp + fuzzel menu) + -- Shift+S captures the whole desktop with no pointer interaction, so it + -- works on scratchpads and popups that region-select would dismiss + -- Lock screen + -- Audio mute cycle (M for Mute): one key walks volume/mic through all four + -- on/off combinations. The waybar pulseaudio modules follow via PipeWire. + -- Exit/session + -- mod+Shift+Backspace no longer exits outright — it enters the "exitconfirm" + -- submap, where a second Backspace confirms and anything else backs out. + -- The bare bind killed two live sessions on 2026-07-22: it sits one key from + -- the mod+Shift cluster (Q wlogout, C killactive, Return terminal, G settings) + -- and took the whole session down with no prompt. wlogout owns the normal + -- exit path; this stays as the keyboard escape hatch for when wlogout won't + -- come up, so it's gated rather than removed. + -- Mouse bindings (from DWM buttons) + -- ============================================================================ + -- Machine-local overrides + -- ============================================================================ + -- Sourced last so machine-specific settings (monitor scale, gaps, keybinds) + -- override the defaults above. See conf.d/local.conf. +}) + +hl.on("hyprland.start", function() + hl.exec_cmd("dbus-update-activation-environment --systemd WAYLAND_DISPLAY XDG_CURRENT_DESKTOP HYPRLAND_INSTANCE_SIGNATURE") + hl.exec_cmd("systemctl --user start hyprland-session.target; systemctl --user restart xdg-desktop-portal-hyprland xdg-desktop-portal-gtk; systemctl --user restart xdg-desktop-portal; waybar-active-config && waybar -c \"$XDG_RUNTIME_DIR/waybar/config\" -s ~/.config/waybar/style.css 2>&1 | grep -v \"LIBDBUSMENU-GLIB-WARNING\" > ~/.local/var/log/waybar-$(date +%Y-%m-%d-%H%M%S).log") + hl.exec_cmd("/usr/lib/polkit-kde-authentication-agent-1") + hl.exec_cmd("/usr/bin/gnome-keyring-daemon --start --components=pkcs11,secrets,ssh") + hl.exec_cmd("dunst > ~/.local/var/log/dunst-$(date +%Y-%m-%d-%H%M%S).log 2>&1") + hl.exec_cmd("awww-daemon & sleep 1 && { settings restore-wallpaper || waypaper --restore; }") + hl.exec_cmd("touchpad-auto") + hl.exec_cmd("pkill -x hypridle; hypridle-start > ~/.local/var/log/hypridle-$(date +%Y-%m-%d-%H%M%S).log 2>&1") + hl.exec_cmd("/usr/lib/geoclue-2.0/demos/agent") + hl.exec_cmd("gammastep > ~/.local/var/log/gammastep-$(date +%Y-%m-%d-%H%M%S).log 2>&1") + hl.exec_cmd("mpd") + hl.exec_cmd("settings restore > ~/.local/var/log/settings-restore-$(date +%Y-%m-%d-%H%M%S).log 2>&1") + hl.exec_cmd("pypr > ~/.local/var/log/pypr-$(date +%Y-%m-%d-%H%M%S).log 2>&1") + hl.exec_cmd("hypr-refocus-scratchpad") + hl.exec_cmd("wait-for-tray && signal-desktop --start-in-tray --ozone-platform=wayland") + hl.exec_cmd("env QT_FONT_DPI=108 protonmail-bridge --no-window") +end) + +hl.on("config.reloaded", function() + hl.exec_cmd("for i in 1 2 3; do sleep 0.2; waybar-reserve; done") +end) + diff --git a/working/hyprland-lua-port/nomerge.lua b/working/hyprland-lua-port/nomerge.lua new file mode 100644 index 0000000..c85fd24 --- /dev/null +++ b/working/hyprland-lua-port/nomerge.lua @@ -0,0 +1,768 @@ +-- Hyprland Configuration +-- Translated from DWM config.def.h and sxhkdrc +-- Craig Jennings <c@cjennings.net> + +-- ============================================================================ +-- Monitor Configuration +-- ============================================================================ +-- Generated by hyprlang2lua. Review TODOs before reloading Hyprland. + +-- hyprlang2lua polyfills — runtime helpers reproducing +-- hyprlang behaviour the typed Lua API doesn't expose directly. + +local function hl_source_glob(pattern) + -- 'source = path/*.conf' had hyprlang glob and inline-expand the + -- matches. require() can't glob, so we shell out to ls (matching + -- the user's brace-expansion behaviour) and dofile each result. + -- Paths with spaces or shell metacharacters in the directory + -- portion will misparse; typical ~/.config/hypr/ layouts don't + -- hit this. Swap to lfs.dir() or find -name if you need fancier. + local p = io.popen("ls " .. pattern .. " 2>/dev/null") + if not p then return end + for f in p:lines() do + local chunk, err = loadfile(f) + if chunk then chunk() + else io.stderr:write("hl_source_glob: " .. tostring(err) .. "\n") end + end + p:close() +end + +hl.monitor({ + output = "", + mode = "preferred", + position = "auto", + scale = "auto", +}) + +-- Waybar's strip (6px top margin + 54px bar) is reserved statically by +-- waybar-reserve, and waybar runs with "exclusive": false. The bar's own +-- exclusive zone would vanish and reappear on every SIGUSR2 reload (the +-- collapse mechanism) and on hide/crash/relaunch, snapping every tiled window +-- up and back down. The static reservation holds the clients in place; only +-- the bar itself changes. `exec` (not exec-once) reruns it on every config +-- reload, which is exactly when Hyprland resets dynamic reservations. The +-- script is idempotent, and a catch-all `monitor=,addreserved,...` rule can't +-- replace it (empty-name addreserved silently no-ops). +-- +-- Run three times over ~0.6s, not once: on reload Hyprland clears the +-- reservation AND re-fires this exec, and the two race. A single run that +-- fires before the clear no-ops (reserved still looks correct), the clear then +-- wins, and the non-exclusive bar drops off-screen. Re-applying past the clear +-- window makes the restore reliable; the script is idempotent so extra runs are +-- free. Applying a monitor rule (e.g. the DP-4 pin) also clears the reservation, +-- so this covers a reload that re-asserts monitors too. + +-- ============================================================================ +-- Startup Applications +-- ============================================================================ +-- Portal and D-Bus setup FIRST, then waybar (needs portal for appearance query) +-- Start hyprland-session.target FIRST: it pulls up graphical-session.target, +-- which xdg-desktop-portal 1.22+ hard-requires (Requisite=). A bare-exec Hyprland +-- session has no session manager to raise that target, so without this the portal +-- fails its dependency at every login (screen-share + file pickers dead). +-- 'systemctl start' blocks until active, so the ';' sequence guarantees the target +-- is up before the portal restart runs. +-- +-- Portal restart (not start) reconnects stale portals on Hyprland restart. +-- Backend portals (GTK, Hyprland) restart BEFORE the main portal to avoid a 50s +-- GTK settings proxy timeout; the sequence keeps that ordering. Separated by ';' +-- not '&&' so a failing portal restart can't stop waybar from launching — waybar +-- degrades gracefully without the portal (only the appearance query is missed), +-- and gating the bar behind the portal left the desktop bar-less whenever +-- xdg-desktop-portal failed its dependency at login. Waybar stays gated on its +-- own config generation (waybar-active-config && waybar). + +-- Core services + +-- Desktop appearance +-- `settings restore` replays the remembered toggles and reapplies the stored +-- wallpaper. It replaced `waypaper --restore` on 2026-08-14: waypaper keeps +-- its own config.ini and the settings store keeps another, neither knew about +-- the other, and the login replay always won — so a wallpaper chosen in the +-- panel came back as whatever the shell had last set. The store is the only +-- one of the two that can hold a sun pair, a video or a projected face, so it +-- owns the restore. set-wallpaper records into it for choices made outside +-- the panel. +-- +-- waypaper --restore stays as the fallback, not the owner. If the stored +-- wallpaper cannot be applied (an image deleted, a drive not mounted yet), +-- `settings restore` exits 3 and waypaper's independent copy still puts +-- something on the screen. Dropping it outright would trade this bug for a +-- bare desktop. +-- +-- The wallpaper half only. The toggle half runs from its own exec-once further +-- down, after hypridle and dunst — caffeine *is* "hypridle isn't running" and +-- DND *is* dunst's pause level, so replaying them here would spend the whole +-- re-assert budget correcting backings that have not launched yet, and would +-- replay them a second time besides. This slot exists for awww's timing, not +-- theirs. + +-- Background services +-- hypridle is reaped on both exit paths, because it outlives its compositor +-- otherwise. An orphaned daemon keeps firing idle actions at whatever session +-- is live next, and it holds its old logind session scope open (the scope can't +-- close while a process sits in it), so orphans accumulate one per abnormal +-- session death. On 2026-07-22 velox reached five concurrent hypridle daemons; +-- two of them racing to lock produced "Cannot re-lock" and a session wedged +-- locked with no client able to draw a password prompt — recoverable only from +-- another console. exec-shutdown covers a clean compositor exit; the pkill in +-- exec-once covers the paths where it never runs (crash, SIGKILL, TTY logout). +hl.on("hyprland.shutdown", function() + hl.exec_cmd("pkill -x hypridle") +end) + +-- hypridle.conf is rendered here rather than tracked, because its contents +-- are this machine's stage times and hibernate setting. Tracking the render +-- meant every panel change dirtied the repo, and whichever machine +-- committed last imposed its policy on the others: a desktop ended up +-- carrying a laptop's suspend-then-hibernate line that it cannot run. +-- Rendering at session start makes the store the only source of truth and +-- the file a build artifact. +-- +-- hypridle-start owns the render, the fallback, and the ordering between +-- them, because that ordering is subtle enough to get wrong in a config +-- line nothing can test: a render can fail on purpose (a damaged store, to +-- avoid overwriting a real policy with defaults), and a fallback that +-- fired there would perform exactly the overwrite the render refused. +-- Replay the toggles that have no durable state of their own. Caffeine *is* +-- "hypridle isn't running" and DND *is* dunst's pause level, so the two +-- exec-once lines above (and dunst's) recreate both at a fixed default every +-- start — a deliberately-set caffeine was silently discarded on every login. +-- Ordered after those launches so it corrects a backing that exists; it also +-- re-asserts for a few seconds, which covers a backing that comes up late. +-- Logged like its neighbours: a silent exec-once failure here would look +-- exactly like the bug it fixes, and gammastep is the standing proof that a +-- launch dying quietly at session start can go unnoticed for a long time. + +-- Pyprland (scratchpads, magnify, etc.) + +-- Tray apps. wait-for-tray blocks until waybar's systray host is up (a fixed +-- sleep can't cover a slow cold-start waybar), so these register their icons +-- instead of opening as windows. Caps at ~30s, then launches anyway. +-- QT_FONT_DPI bumps the bridge's QML UI font (qt6ct General font is ignored by Qt Quick) + +-- ============================================================================ +-- Environment Variables +-- ============================================================================ +hl.env("XCURSOR_SIZE", "24") +hl.env("XCURSOR_THEME", "Bibata-Modern-Ice") +hl.env("XDG_CURRENT_DESKTOP", "Hyprland") +hl.env("XDG_SESSION_TYPE", "wayland") +hl.env("XDG_SESSION_DESKTOP", "Hyprland") +hl.env("_JAVA_AWT_WM_NONREPARENTING", "1") + +-- ============================================================================ +-- Appearance (matching DWM colors) +-- ============================================================================ +-- DWM colors: gray1=#222222, gray2=#444444, gray3=#bbbbbb, gray4=#eeeeee, cyan=#daa520 + +hl.config({ + general = { + gaps_in = 25, + gaps_out = 30, + border_size = 2, + col = { + active_border = "rgba(daa520ff)", + inactive_border = "rgba(444444ff)", + nogroup_border_active = "rgba(daa520ff)", + nogroup_border = "rgba(444444ff)", + }, + -- Pyprland 3.4+ applies `group deny` to scratchpads, which routes their + -- border through col.nogroup_border* instead of col.*_border. Without + -- these overrides Hyprland's defaults paint scratchpads bright magenta. + layout = "master", + resize_on_border = true, + }, +}) + +hl.config({ + decoration = { + rounding = 10, + dim_inactive = true, + dim_strength = 0.4, + dim_special = 0.2, + blur = { + enabled = false, + }, + shadow = { + enabled = false, + }, + }, +}) + +hl.config({ + animations = { + enabled = true, + }, +}) + +hl.curve("myBezier", { type = "bezier", points = { { 0.05, 0.9 }, { 0.1, 1.05 } } }) +hl.animation({ + leaf = "windows", + enabled = true, + speed = 2, + bezier = "myBezier", +}) +hl.animation({ + leaf = "windowsOut", + enabled = true, + speed = 2, + bezier = "default", + style = "popin 80%", +}) +hl.animation({ + leaf = "fade", + enabled = true, + speed = 2, + bezier = "default", +}) +hl.animation({ + leaf = "workspaces", + enabled = true, + speed = 2, + bezier = "default", +}) +hl.animation({ + leaf = "specialWorkspace", + enabled = true, + speed = 2, + bezier = "default", + style = "slidevert", +}) + +-- ============================================================================ +-- Layout (master-stack like DWM tile) +-- ============================================================================ + +hl.config({ + master = { + new_status = "master", + new_on_top = true, + mfact = 0.55, + }, +}) + +hl.config({ + dwindle = { + preserve_split = true, + }, +}) + +-- ============================================================================ +-- Input +-- ============================================================================ + +hl.config({ + cursor = { + no_warps = true, + inactive_timeout = 2.0, + }, +}) + +hl.config({ + input = { + kb_layout = "us", + kb_options = "ctrl:nocaps", + numlock_by_default = true, + follow_mouse = 0, + -- 0, not the default 1: with follow_mouse off we never want focus to follow + -- the cursor. At 1, focus still jumps to the window under the pointer when it + -- crosses a floating<->tiled boundary, so launching a floating scratchpad (or + -- the org-capture popup) re-enabled focus-follows-mouse onto tiled windows. + float_switch_override_focus = 0, + mouse_refocus = false, + natural_scroll = true, + touchpad = { + natural_scroll = false, + }, + }, +}) + +-- ============================================================================ +-- Misc +-- ============================================================================ + +hl.config({ + misc = { + force_default_wallpaper = 0, + disable_hyprland_logo = true, + -- false so apps can't pull focus via activation requests. New windows still + -- focus on open (separate path); this stops e.g. a browser yanking focus + -- back off a freshly opened emacs frame. + focus_on_activate = false, + -- Let a fresh lock client adopt a session whose previous one died. The + -- default (off) is the strict reading of ext-session-lock: a dead lock + -- client leaves the session locked forever and refuses every replacement + -- ("Cannot re-lock"), so the screen stays up with nothing able to draw a + -- password prompt and the only way back in is another console. That is a + -- hard lockout, and it cost a session on velox 2026-07-22. On means a + -- replacement hyprlock re-attaches and prompts normally. The screen stays + -- locked either way — this decides whether the lock is recoverable, never + -- whether it holds. + allow_session_lock_restore = true, + }, +}) + +-- ============================================================================ +-- Debug (temporary - disable when stable) +-- ============================================================================ + +hl.config({ + debug = { + disable_logs = false, + }, +}) + +-- ============================================================================ +-- XWayland +-- ============================================================================ + +hl.config({ + xwayland = { + force_zero_scaling = true, + }, +}) + +-- ============================================================================ +-- Window Rules (Hyprland 0.53+ syntax: match:CONDITION, RULE) +-- ============================================================================ +-- Floating windows (from DWM rules) +hl.window_rule({ + match = { + class = "^(xdg-desktop-portal-gtk)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(Gimp)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(caffeine)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(qalculate-gtk)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + title = "^(Event Tester)$", + }, + float = true, +}) + +-- net / bluetooth instrument-console panels. Normal floating windows (formerly +-- gtk4-layer-shell overlays) so they drag to move and corner-drag to resize. +-- Opened top-right to match their old anchored spot: the panel is right-aligned +-- with a 44px gap, so x = 100% - (window width + 44). net is 420 wide, bt 380. +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.netpanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.netpanel)$", + }, + move = "100%-464 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.btpanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.btpanel)$", + }, + move = "100%-424 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.audiopanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.audiopanel)$", + }, + move = "100%-444 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.timerpanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.timerpanel)$", + }, + move = "100%-444 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.settingspanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.settingspanel)$", + }, + move = "100%-584 50", +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.weatherpanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.weatherpanel)$", + }, + move = "100%-464 50", +}) + +-- maintenance console: the wide board (960), same right-aligned convention. +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.maintpanel)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + class = "^(net\\.cjennings\\.maintpanel)$", + }, + move = "100%-1004 50", +}) + +-- org-capture popup frame (quick-capture script names the frame) +-- Size is per-host in <host>/conf.d/local.conf: native window rules ignore +-- percentages (only pyprland honors them), so the popup is sized in absolute +-- pixels matching that host's terminal scratchpad. No size rule here means a +-- host without an override falls back to the script's char-cell geometry. +hl.window_rule({ + match = { + title = "^(org-capture)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + title = "^(org-capture)$", + }, + center = true, +}) + +-- dirvish popup frame (dirvish-popup script names the frame). No stay_focused — +-- it's a file manager that launches files into other apps, so focus must be free +-- to follow; q (cj/dirvish-popup-quit) closes the frame. +hl.window_rule({ + match = { + title = "^(dirvish)$", + }, + float = true, +}) + +hl.window_rule({ + match = { + title = "^(dirvish)$", + }, + size = "1100 700", +}) + +hl.window_rule({ + match = { + title = "^(dirvish)$", + }, + center = true, +}) + +-- NOTE: center windowrules removed 2026-03-04 per pyprland maintainer suggestion +-- Testing whether pyprland handles scratchpad re-centering natively (issue #211) + +-- Gaming +hl.window_rule({ + match = { + class = "^(Civ5XP)$", + }, + fullscreen = true, +}) + +-- ============================================================================ +-- Key Bindings +-- ============================================================================ +local mod = "SUPER" + +-- Terminal and core apps (from DWM) +hl.bind(mod .. " + T", hl.dsp.exec_cmd("foot")) +hl.bind(mod .. " + E", hl.dsp.exec_cmd("emacsclient -c -a \"\" || emacs")) +-- Standalone emacs: its own process, not a frame on the daemon, so killing it +-- takes no other frame with it. init.el guards server-start on server-running-p, +-- so while the daemon holds the socket this process leaves it alone. With no +-- daemon up it finds no server and becomes one -- the guard lives in init.el, and +-- a keybind can't override it, since --eval runs after init. +hl.bind(mod .. " + SHIFT + E", hl.dsp.exec_cmd("emacs")) +hl.bind(mod .. " + N", hl.dsp.exec_cmd("quick-capture")) +hl.bind(mod .. " + W", hl.dsp.exec_cmd("$BROWSER")) +hl.bind(mod .. " + F", hl.dsp.exec_cmd("dirvish-popup")) +hl.bind(mod .. " + SHIFT + F", hl.dsp.exec_cmd("layout-cycle float-toggle")) + +-- From sxhkdrc +hl.bind(mod .. " + SPACE", hl.dsp.exec_cmd("fuzzel-toggle")) +hl.bind(mod .. " + SHIFT + W", hl.dsp.exec_cmd("$ALTBROWSER")) +hl.bind(mod .. " + P", hl.dsp.exec_cmd("media-toggle-all")) +hl.bind(mod .. " + SHIFT + L", hl.dsp.exec_cmd("calibre")) +hl.bind(mod .. " + SHIFT + P", hl.dsp.exec_cmd("toggle-touchpad")) + +-- Window management (from DWM) +-- Layout-aware navigation (works across master, scrolling) +hl.bind(mod .. " + J", hl.dsp.exec_cmd("layout-navigate next")) +hl.bind(mod .. " + K", hl.dsp.exec_cmd("layout-navigate prev")) +hl.bind(mod .. " + SHIFT + J", hl.dsp.exec_cmd("layout-navigate next move")) +hl.bind(mod .. " + SHIFT + K", hl.dsp.exec_cmd("layout-navigate prev move")) +hl.bind(mod .. " + H", hl.dsp.exec_cmd("layout-resize shrink")) +hl.bind(mod .. " + L", hl.dsp.exec_cmd("layout-resize grow")) +-- Swap focused window with master, then force focus onto the master slot. +-- swapwithmaster's own `master` focus param doesn't stick when invoked from +-- the master, so focusmaster master pins focus afterward. The 50ms sleep is +-- load-bearing: swapwithmaster fires an async focus event ~1-2ms after it +-- returns; without the delay that event lands AFTER focusmaster and flips +-- focus back to the detail. The sleep lets the swap's focus settle so +-- focusmaster runs last and wins. Proven via instrumented capture (19/19). +hl.bind(mod .. " + RETURN", hl.dsp.layout("swapwithmaster && sleep 0.05 && hyprctl dispatch layoutmsg focusmaster master")) +hl.bind(mod .. " + G", hl.dsp.window.center()) +hl.bind(mod .. " + TAB", hl.dsp.focus({ workspace = "previous" })) +hl.bind(mod .. " + SHIFT + C", hl.dsp.window.close()) + +-- Layouts: master -> monocle +-- Cycle with Shift+arrows, or jump directly with Shift+T/M +-- (scrolling layout disabled until frame-fit + wrap-around work lands) +hl.bind(mod .. " + SHIFT + RIGHT", hl.dsp.exec_cmd("layout-cycle next")) +hl.bind(mod .. " + SHIFT + LEFT", hl.dsp.exec_cmd("layout-cycle prev")) +hl.bind(mod .. " + SHIFT + T", hl.dsp.exec_cmd("hyprctl keyword general:layout master && hyprctl keyword master:orientation left")) +hl.bind(mod .. " + SHIFT + M", hl.dsp.exec_cmd("hyprctl keyword general:layout monocle")) +hl.bind(mod .. " + SHIFT + SPACE", hl.dsp.window.float({ action = "toggle" })) + +-- Master layout adjustments +hl.bind(mod .. " + U", hl.dsp.layout("addmaster")) +hl.bind(mod .. " + D", hl.dsp.layout("removemaster")) + +-- Stash windows (hide to special workspace) +-- O = stash focused / Alt+O = stash others / Shift+O = restore all +hl.bind(mod .. " + O", hl.dsp.exec_cmd("stash-window")) +hl.bind(mod .. " + ALT + O", hl.dsp.exec_cmd("stash-others")) +hl.bind(mod .. " + SHIFT + O", hl.dsp.exec_cmd("stash-restore")) + +-- Gaps between windows only; window-gaps leaves the monitor-edge gap +-- (general:gaps_out) fixed, so widening/narrowing moves the space between +-- windows, not the screen-edge margin. +hl.bind(mod .. " + MINUS", hl.dsp.exec_cmd("window-gaps narrow")) +hl.bind(mod .. " + EQUAL", hl.dsp.exec_cmd("window-gaps widen")) +hl.bind(mod .. " + SHIFT + EQUAL", hl.dsp.exec_cmd("window-gaps reset")) +hl.bind(mod .. " + SHIFT + MINUS", hl.dsp.exec_cmd("window-gaps zero")) + +-- Auto-dim toggle (D = dim). Same action as clicking the waybar custom/dim icon. +hl.bind(mod .. " + SHIFT + D", hl.dsp.exec_cmd("dim-toggle")) +hl.bind(mod .. " + SHIFT + G", hl.dsp.exec_cmd("settings-panel")) + +-- Caffeine (keep-awake) toggle. Same action as clicking the waybar +-- custom/caffeine icon — flips the hypridle daemon so the screen will / won't +-- lock. Stays on $mod+I ($mod+C is taken by hyprpicker; no free caffeine key). +hl.bind(mod .. " + I", hl.dsp.exec_cmd("caffeine-toggle")) + +-- Airplane mode (low-power: wifi off + CPU/brightness/services). A deliberate +-- keybind, not a bar click — engaging it disconnects you, so it shouldn't be a +-- misclick away. The custom/net module shows the state; this toggles it. +-- On Super+Shift+X ("X" = everything off); Super+Shift+A toggles push-to-talk. +hl.bind(mod .. " + SHIFT + X", hl.dsp.exec_cmd("airplane-mode")) + +-- Toggle bar visibility, or relaunch waybar if it crashed (no exec-once respawn). +hl.bind(mod .. " + B", hl.dsp.exec_cmd("waybar-toggle")) + +-- Collapse / expand the left or right side of the bar to its base set +-- (same action as clicking the side's arrowhead). [ = left, ] = right. +hl.bind(mod .. " + bracketleft", hl.dsp.exec_cmd("waybar-collapse left")) +hl.bind(mod .. " + bracketright", hl.dsp.exec_cmd("waybar-collapse right")) + +-- Fullscreen +hl.bind(mod .. " + F11", hl.dsp.window.fullscreen({ mode = "fullscreen", action = "toggle" })) + +-- Workspaces 1-9 (from DWM TAGKEYS) +hl.bind(mod .. " + 1", hl.dsp.focus({ workspace = 1 })) +hl.bind(mod .. " + 2", hl.dsp.focus({ workspace = 2 })) +hl.bind(mod .. " + 3", hl.dsp.focus({ workspace = 3 })) +hl.bind(mod .. " + 4", hl.dsp.focus({ workspace = 4 })) +hl.bind(mod .. " + 5", hl.dsp.focus({ workspace = 5 })) +hl.bind(mod .. " + 6", hl.dsp.focus({ workspace = 6 })) +hl.bind(mod .. " + 7", hl.dsp.focus({ workspace = 7 })) +hl.bind(mod .. " + 8", hl.dsp.focus({ workspace = 8 })) +hl.bind(mod .. " + 9", hl.dsp.focus({ workspace = 9 })) + +-- Move window to workspace (from DWM tag) +hl.bind(mod .. " + SHIFT + 1", hl.dsp.window.move({ workspace = 1, follow = false })) +hl.bind(mod .. " + SHIFT + 2", hl.dsp.window.move({ workspace = 2, follow = false })) +hl.bind(mod .. " + SHIFT + 3", hl.dsp.window.move({ workspace = 3, follow = false })) +hl.bind(mod .. " + SHIFT + 4", hl.dsp.window.move({ workspace = 4, follow = false })) +hl.bind(mod .. " + SHIFT + 5", hl.dsp.window.move({ workspace = 5, follow = false })) +hl.bind(mod .. " + SHIFT + 6", hl.dsp.window.move({ workspace = 6, follow = false })) +hl.bind(mod .. " + SHIFT + 7", hl.dsp.window.move({ workspace = 7, follow = false })) +hl.bind(mod .. " + SHIFT + 8", hl.dsp.window.move({ workspace = 8, follow = false })) +hl.bind(mod .. " + SHIFT + 9", hl.dsp.window.move({ workspace = 9, follow = false })) + +-- Monitor focus (from DWM focusmon) +hl.bind(mod .. " + COMMA", hl.dsp.focus({ monitor = -1 })) +hl.bind(mod .. " + PERIOD", hl.dsp.focus({ monitor = "+1" })) +hl.bind(mod .. " + SHIFT + COMMA", hl.dsp.window.move({ monitor = "-1" })) +hl.bind(mod .. " + SHIFT + PERIOD", hl.dsp.window.move({ monitor = "+1" })) + +-- ============================================================================ +-- Scratchpads (via pyprland) +-- ============================================================================ +-- Configured in ~/.config/hypr/pyprland.toml +-- Uses normal workspaces (not special), so new windows won't be captured +hl.bind(mod .. " + SHIFT + RETURN", hl.dsp.exec_cmd("pypr toggle term")) +hl.bind(mod .. " + A", hl.dsp.exec_cmd("audio-panel")) +hl.bind(mod .. " + R", hl.dsp.exec_cmd("pypr toggle monitor")) +hl.bind(mod .. " + SHIFT + N", hl.dsp.exec_cmd("net panel")) +hl.bind(mod .. " + SLASH", hl.dsp.exec_cmd("pypr toggle music")) + +-- Magnify (zoom) +-- mod+Z zooms and enters the "zoom" submap; inside it, Escape or mod+Z +-- unzooms and returns to the normal keymap. Exit forces `pypr zoom 1` +-- (factor 1) so submap state and zoom state can't desync. Note: while +-- zoomed, other Hyprland binds pause until you exit the submap. +hl.bind(mod .. " + Z", hl.dsp.exec_cmd("pypr zoom")) +hl.bind(mod .. " + Z", hl.dsp.submap("zoom")) + +hl.define_submap("zoom", function() + hl.bind("ESCAPE", hl.dsp.exec_cmd("pypr zoom 1")) + hl.bind("ESCAPE", hl.dsp.submap("reset")) + hl.bind(mod .. " + Z", hl.dsp.exec_cmd("pypr zoom 1")) + hl.bind(mod .. " + Z", hl.dsp.submap("reset")) +end) + +-- Calculator (not a scratchpad, just launches app) +hl.bind(mod .. " + X", hl.dsp.exec_cmd("calc-toggle")) +hl.bind(mod .. " + C", hl.dsp.exec_cmd("hyprpicker -a")) +hl.bind(mod .. " + CONTROL + C", hl.dsp.exec_cmd("clock-panel toggle")) + +-- Media/hardware keys +hl.bind("XF86AudioRaiseVolume", hl.dsp.exec_cmd("pactl set-sink-volume @DEFAULT_SINK@ +5%"), { locked = true, repeating = true }) +hl.bind("XF86AudioLowerVolume", hl.dsp.exec_cmd("pactl set-sink-volume @DEFAULT_SINK@ -5%"), { locked = true, repeating = true }) +hl.bind("XF86AudioMute", hl.dsp.exec_cmd("audio quick-mute"), { locked = true }) +hl.bind("XF86MonBrightnessUp", hl.dsp.exec_cmd("brightnessctl s +10%"), { locked = true, repeating = true }) +hl.bind("XF86MonBrightnessDown", hl.dsp.exec_cmd("brightnessctl s 10%-"), { locked = true, repeating = true }) + +-- Microphone mute toggle (waybar pulseaudio#mic indicator follows via PipeWire events). +-- On the hardware mic-mute key. Super+Shift+A used to duplicate this; it now +-- toggles push-to-talk mode instead (mic-toggle stays reachable on the hw key). +hl.bind("XF86AudioMicMute", hl.dsp.exec_cmd("mic-toggle"), { locked = true }) + +-- Push-to-talk toggle: enter PTT (mic muted, hold key armed) / exit (restore). +-- The hold key is configurable (audio config ptt_key = Control_R or mouse:NNN). +hl.bind(mod .. " + SHIFT + A", hl.dsp.exec_cmd("audio ptt-toggle")) + +-- Bluetooth panel (blueman retired in favor of the bt panel) +hl.bind(mod .. " + SHIFT + B", hl.dsp.exec_cmd("bt-panel")) + +-- Screenshots (grim + slurp + fuzzel menu) +-- Shift+S captures the whole desktop with no pointer interaction, so it +-- works on scratchpads and popups that region-select would dismiss +hl.bind(mod .. " + S", hl.dsp.exec_cmd("screenshot region")) +hl.bind(mod .. " + SHIFT + S", hl.dsp.exec_cmd("screenshot fullscreen")) +hl.bind("CTRL" .. mod .. " + S", hl.dsp.exec_cmd("screenshot fullscreen")) + +-- Lock screen +hl.bind(mod .. " + ESCAPE", hl.dsp.exec_cmd("hyprlock")) + +-- Audio mute cycle (M for Mute): one key walks volume/mic through all four +-- on/off combinations. The waybar pulseaudio modules follow via PipeWire. +hl.bind(mod .. " + M", hl.dsp.exec_cmd("audio-cycle")) + +-- Exit/session +hl.bind(mod .. " + SHIFT + Q", hl.dsp.exec_cmd("pgrep -x wlogout || wlogout-menu")) +-- mod+Shift+Backspace no longer exits outright — it enters the "exitconfirm" +-- submap, where a second Backspace confirms and anything else backs out. +-- The bare bind killed two live sessions on 2026-07-22: it sits one key from +-- the mod+Shift cluster (Q wlogout, C killactive, Return terminal, G settings) +-- and took the whole session down with no prompt. wlogout owns the normal +-- exit path; this stays as the keyboard escape hatch for when wlogout won't +-- come up, so it's gated rather than removed. +hl.bind(mod .. " + SHIFT + BACKSPACE", hl.dsp.submap("exitconfirm")) + +hl.define_submap("exitconfirm", function() + hl.bind("BACKSPACE", hl.dsp.exit()) + hl.bind("ESCAPE", hl.dsp.submap("reset")) + hl.bind("catchall", hl.dsp.submap("reset")) +end) + +hl.bind(mod .. " + SHIFT + ESCAPE", hl.dsp.exec_cmd("hyprctl reload")) +hl.bind("CTRL + ALT" .. mod .. " + K", hl.dsp.exec_cmd("hyprctl kill")) + +-- Mouse bindings (from DWM buttons) +hl.bind(mod .. " + mouse:272", hl.dsp.window.drag()) +hl.bind(mod .. " + mouse:273", hl.dsp.window.resize()) +hl.bind(mod .. " + SHIFT + mouse:272", hl.dsp.window.resize()) + +-- ============================================================================ +-- Machine-local overrides +-- ============================================================================ +-- Sourced last so machine-specific settings (monitor scale, gaps, keybinds) +-- override the defaults above. See conf.d/local.conf. +-- Source: $HOME/.config/hypr/conf.d/*.conf (glob; resolved at runtime). Each matched .conf must be converted to .lua. +hl_source_glob("$HOME/.config/hypr/conf.d/*.lua") + +hl.on("hyprland.start", function() + hl.exec_cmd("dbus-update-activation-environment --systemd WAYLAND_DISPLAY XDG_CURRENT_DESKTOP HYPRLAND_INSTANCE_SIGNATURE") + hl.exec_cmd("systemctl --user start hyprland-session.target; systemctl --user restart xdg-desktop-portal-hyprland xdg-desktop-portal-gtk; systemctl --user restart xdg-desktop-portal; waybar-active-config && waybar -c \"$XDG_RUNTIME_DIR/waybar/config\" -s ~/.config/waybar/style.css 2>&1 | grep -v \"LIBDBUSMENU-GLIB-WARNING\" > ~/.local/var/log/waybar-$(date +%Y-%m-%d-%H%M%S).log") + hl.exec_cmd("/usr/lib/polkit-kde-authentication-agent-1") + hl.exec_cmd("/usr/bin/gnome-keyring-daemon --start --components=pkcs11,secrets,ssh") + hl.exec_cmd("dunst > ~/.local/var/log/dunst-$(date +%Y-%m-%d-%H%M%S).log 2>&1") + hl.exec_cmd("awww-daemon & sleep 1 && { settings restore-wallpaper || waypaper --restore; }") + hl.exec_cmd("touchpad-auto") + hl.exec_cmd("pkill -x hypridle; hypridle-start > ~/.local/var/log/hypridle-$(date +%Y-%m-%d-%H%M%S).log 2>&1") + hl.exec_cmd("/usr/lib/geoclue-2.0/demos/agent") + hl.exec_cmd("gammastep > ~/.local/var/log/gammastep-$(date +%Y-%m-%d-%H%M%S).log 2>&1") + hl.exec_cmd("mpd") + hl.exec_cmd("settings restore > ~/.local/var/log/settings-restore-$(date +%Y-%m-%d-%H%M%S).log 2>&1") + hl.exec_cmd("pypr > ~/.local/var/log/pypr-$(date +%Y-%m-%d-%H%M%S).log 2>&1") + hl.exec_cmd("hypr-refocus-scratchpad") + hl.exec_cmd("wait-for-tray && signal-desktop --start-in-tray --ozone-platform=wayland") + hl.exec_cmd("env QT_FONT_DPI=108 protonmail-bridge --no-window") +end) + +hl.on("config.reloaded", function() + hl.exec_cmd("for i in 1 2 3; do sleep 0.2; waybar-reserve; done") +end) + diff --git a/working/hyprland-lua-port/ratio-local.lua b/working/hyprland-lua-port/ratio-local.lua new file mode 100644 index 0000000..de190d5 --- /dev/null +++ b/working/hyprland-lua-port/ratio-local.lua @@ -0,0 +1,43 @@ +-- ratio — desktop, 1x scaling. Defaults in hyprland.lua are correct. +-- Sourced via conf.d/*.lua glob (last wins). +-- +-- Examples: +-- hl.monitor({ output = "DP-1", mode = "3440x1440@144", position = "auto", scale = 1 }) +-- hl.bind("SUPER + L", hl.dsp.exec_cmd("hyprlock")) +-- +-- Spell the modifier out. A sourced file is loaded by hl_source_glob via +-- loadfile, which gives the chunk globals only, so the shared config's `mod` +-- (and `at_start`, and the other locals) are NOT in scope here. + +-- DP-4 (Dell U3419W ultrawide) pinned to its native mode. Without an explicit +-- pin the shared catch-all monitor=,preferred,auto,auto lets an XWayland surface +-- (emacs runs X11-only on ratio) drive the mode down to 1280x720 at login. That +-- low mode also wipes DP-4's reserved area, which drops the non-exclusive waybar +-- off-screen (margin-top:-54 needs the 60px top reserve). Pinning holds both. + +hl.monitor({ + output = "DP-4", + mode = "3440x1440@60", + position = "0x0", + scale = "1", +}) + +-- org-capture popup: capped at 120 Emacs columns wide, height proportional +-- (Craig, 2026-07-14). 120 cols x 11 px/col = 1320 wide; height keeps the old +-- rule's aspect (1892:936) = 653 px (~27 lines at 24 px). The old 55% x 65% +-- scratchpad match (1892 x 936) was the "grows too large" complaint. max_size +-- holds the cap even if the frame tries to grow with capture content. +hl.window_rule({ + match = { + title = "^(org-capture)$", + }, + size = "1320 653", +}) + +hl.window_rule({ + match = { + title = "^(org-capture)$", + }, + max_size = "1320 653", +}) + diff --git a/working/hyprland-lua-port/reader-changes-for-lua.patch b/working/hyprland-lua-port/reader-changes-for-lua.patch new file mode 100644 index 0000000..747b047 --- /dev/null +++ b/working/hyprland-lua-port/reader-changes-for-lua.patch @@ -0,0 +1,241 @@ +diff --git a/common/.local/bin/dotfiles-validate b/common/.local/bin/dotfiles-validate +index 57a7505..d5e6695 100755 +--- a/common/.local/bin/dotfiles-validate ++++ b/common/.local/bin/dotfiles-validate +@@ -3,6 +3,7 @@ + # + # Walks the tree and extracts the commands that configs promise to launch: + # - hypr conf files: exec-once = CMD / exec = CMD / bind* = ..., exec, CMD ++# - hypr lua configs: at_start/at_shutdown/at_reload("CMD"), exec_cmd("CMD") + # - waybar config: "exec(-if)", "on-click*", "on-scroll-*", + # "on-double-click" values + # - systemd user units: Exec*= lines (leading -/@ modifiers stripped) +@@ -45,6 +46,30 @@ find "$root" -path '*/.config/hypr/*.conf' -type f 2>/dev/null | while read -r f + ' "$f" + done >> "$refs_file" + ++# --- hypr lua configs: autostart collectors and exec_cmd dispatchers --- ++# The Lua config manager (Hyprland 0.55+) spells the same two things as function ++# calls rather than assignments, so the .conf walk above sees none of them. Only ++# the first word is taken, as everywhere else here. The two matches are written ++# out rather than folded into a helper: awk cannot take a regex literal as a ++# function parameter -- it collapses to a boolean match against $0, which ++# silently turns every line into a bogus reference. ++find "$root" -path '*/.config/hypr/*.lua' -type f 2>/dev/null | while read -r f; do ++ awk -v file="$f" ' ++ match($0, /at_(start|shutdown|reload)\("/) { ++ rest = substr($0, RSTART + RLENGTH) ++ sub(/".*$/, "", rest) ++ n = split(rest, w, /[ \t]+/) ++ if (n > 0 && w[1] != "") print file ":" FNR ":" w[1] ++ } ++ match($0, /hl\.dsp\.exec_cmd\("/) { ++ rest = substr($0, RSTART + RLENGTH) ++ sub(/".*$/, "", rest) ++ n = split(rest, w, /[ \t]+/) ++ if (n > 0 && w[1] != "") print file ":" FNR ":" w[1] ++ } ++ ' "$f" ++done >> "$refs_file" ++ + # --- waybar configs: command-bearing JSON values --- + find "$root" -path '*/.config/waybar/*' -type f \( -name config -o -name '*.json' -o -name '*.jsonc' \) 2>/dev/null | while read -r f; do + awk -v file="$f" ' +diff --git a/tests/layout-cycle/test_layout_cycle.py b/tests/layout-cycle/test_layout_cycle.py +index 1454817..3740024 100644 +--- a/tests/layout-cycle/test_layout_cycle.py ++++ b/tests/layout-cycle/test_layout_cycle.py +@@ -18,6 +18,7 @@ Run from repo root: + + import json + import os ++import re + import subprocess + import tempfile + import unittest +@@ -25,7 +26,7 @@ import unittest + REPO_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")) + SCRIPT = os.path.join(REPO_ROOT, "hyprland/.local/bin/layout-cycle") + FAKE_HYPRCTL = os.path.join(os.path.dirname(__file__), "fake-hyprctl") +-HYPRLAND_CONF = os.path.join(REPO_ROOT, "hyprland/.config/hypr/hyprland.conf") ++HYPRLAND_CFG = os.path.join(REPO_ROOT, "hyprland/.config/hypr/hyprland.lua") + + FLASH = "rgba(ffd24aff)" + ACTIVE = "rgba(daa520ff)" +@@ -326,7 +327,7 @@ class TestFlashAllBorders(LayoutCycleHarness): + + + class TestScrollLayoutDisabledInConfig(unittest.TestCase): +- """Pin the hyprland.conf half of the scrolling disable. ++ """Pin the hyprland.lua half of the scrolling disable. + + The script tests above prove the ring skips scrolling; these prove no + keybinding reaches it either, and that the freed Super+Shift+S chord +@@ -335,10 +336,20 @@ class TestScrollLayoutDisabledInConfig(unittest.TestCase): + + @classmethod + def setUpClass(cls): +- with open(HYPRLAND_CONF) as f: ++ with open(HYPRLAND_CFG) as f: + cls.conf = f.read() ++ # Every bind in the Lua config is a top-level hl.bind() call; the ++ # locked/mouse/repeat variants that hyprlang spelled bindl/bindm/binde ++ # are the same call with an options table, so one prefix covers them. + cls.binds = [l for l in cls.conf.splitlines() +- if l.strip().startswith(("bind", "bindl", "bindm"))] ++ if l.strip().startswith("hl.bind")] ++ ++ def test_the_binds_were_actually_found(self): ++ """Guard the guard: a renamed call would empty the list and make both ++ assertions below pass against nothing.""" ++ self.assertGreater(len(self.binds), 50, ++ "found almost no hl.bind lines -- the two assertions " ++ "below would pass vacuously") + + def test_no_bind_selects_scrolling_layout(self): + offenders = [l for l in self.binds if "general:layout scrolling" in l] +@@ -346,7 +357,7 @@ class TestScrollLayoutDisabledInConfig(unittest.TestCase): + + def test_super_shift_s_is_fullscreen_screenshot(self): + shift_s = [l for l in self.binds +- if "$mod SHIFT, S," in l] ++ if re.search(r'\bmod\s*\.\.\s*"\s*\+\s*SHIFT\s*\+\s*S"', l)] + self.assertEqual(len(shift_s), 1, msg=f"binds found: {shift_s}") + self.assertIn("screenshot fullscreen", shift_s[0]) + +diff --git a/tests/settings/test_session_restore.py b/tests/settings/test_session_restore.py +index a989d52..b2a2721 100644 +--- a/tests/settings/test_session_restore.py ++++ b/tests/settings/test_session_restore.py +@@ -585,12 +585,32 @@ class TestCompositorWiring(unittest.TestCase): + script and passing tests while never being placed in the bar. + """ + +- CONF = os.path.join(REPO_ROOT, "hyprland/.config/hypr/hyprland.conf") ++ CONF = os.path.join(REPO_ROOT, "hyprland/.config/hypr/hyprland.lua") + + def _lines(self): ++ """The autostart commands, in launch order. ++ ++ The Lua config collects startup commands with at_start("CMD") and one ++ hl.on("hyprland.start") handler at the bottom replays the list in order, ++ so position in this file is still position at launch -- which is what ++ the ordering assertion below reads. Returning the command strings rather ++ than raw lines means every entry here is by construction an autostart ++ command, so the old startswith("exec-once") filter has nothing left to do. ++ """ ++ out = [] + with open(self.CONF) as f: +- return [ln.strip() for ln in f +- if ln.strip() and not ln.strip().startswith("#")] ++ for ln in f: ++ m = re.search(r'at_start\("(.*)"\)', ln.strip()) ++ if m: ++ out.append(m.group(1)) ++ return out ++ ++ def test_the_autostart_commands_were_actually_found(self): ++ """Guard the guard: a renamed collector would empty the list and make ++ every assertion below pass against nothing.""" ++ self.assertGreater(len(self._lines()), 10, ++ "found almost no at_start commands -- the assertions " ++ "below would pass vacuously") + + @staticmethod + def _is_toggle_restore(line): +@@ -607,22 +627,19 @@ class TestCompositorWiring(unittest.TestCase): + + def test_restore_runs_at_session_start(self): + self.assertTrue( +- any(ln.startswith("exec-once") and self._is_toggle_restore(ln) +- for ln in self._lines()), +- "hyprland.conf has no exec-once running `settings restore` — " ++ any(self._is_toggle_restore(ln) for ln in self._lines()), ++ "hyprland.lua has no at_start running `settings restore` — " + "remembered toggles would never be replayed") + + def test_wallpaper_restore_runs_at_session_start(self): + self.assertTrue( +- any(ln.startswith("exec-once") and "settings restore-wallpaper" in ln +- for ln in self._lines()), +- "hyprland.conf has no exec-once running `settings restore-wallpaper` " ++ any("settings restore-wallpaper" in ln for ln in self._lines()), ++ "hyprland.lua has no at_start running `settings restore-wallpaper` " + "— the stored wallpaper would never be put back") + + def test_the_toggle_restore_runs_only_once(self): + """Twice means the whole re-assert budget is spent twice per login.""" +- lines = [ln for ln in self._lines() +- if ln.startswith("exec-once") and self._is_toggle_restore(ln)] ++ lines = [ln for ln in self._lines() if self._is_toggle_restore(ln)] + self.assertEqual(len(lines), 1, lines) + + def test_restore_is_ordered_after_the_backings_it_corrects(self): +@@ -631,17 +648,15 @@ class TestCompositorWiring(unittest.TestCase): + # burn attempts on backings that aren't up yet. + lines = self._lines() + restore = next(i for i, ln in enumerate(lines) +- if ln.startswith("exec-once") +- and self._is_toggle_restore(ln)) ++ if self._is_toggle_restore(ln)) + for backing in ("hypridle", "dunst"): +- launch = next(i for i, ln in enumerate(lines) +- if ln.startswith("exec-once") and backing in ln) ++ launch = next(i for i, ln in enumerate(lines) if backing in ln) + self.assertLess(launch, restore, + f"`settings restore` is ordered before {backing}") + + + class DimEnv(TempEnv): +- """Adds the hyprctl fake, whose dim default is 1 -- as hyprland.conf's is.""" ++ """Adds the hyprctl fake, whose dim default is 1 -- as hyprland.lua's is.""" + + def setUp(self): + super().setUp() +diff --git a/tests/waybar-reserve/test_reserve_pairing.py b/tests/waybar-reserve/test_reserve_pairing.py +index 7421692..fbeb95d 100644 +--- a/tests/waybar-reserve/test_reserve_pairing.py ++++ b/tests/waybar-reserve/test_reserve_pairing.py +@@ -26,7 +26,7 @@ import unittest + + REPO_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")) + WAYBAR = os.path.join(REPO_ROOT, "hyprland/.config/waybar/config") +-HYPR = os.path.join(REPO_ROOT, "hyprland/.config/hypr/hyprland.conf") ++HYPR = os.path.join(REPO_ROOT, "hyprland/.config/hypr/hyprland.lua") + + + def waybar_config(): +@@ -35,13 +35,19 @@ def waybar_config(): + + + def reserve_exec_wired(): +- """True when hyprland.conf runs waybar-reserve via exec (not exec-once). ++ """True when the config re-runs waybar-reserve on every config reload. + +- Matches waybar-reserve invoked anywhere in an ``exec =`` line, so the +- reload-race-safe retry-loop form (``exec = for i in 1 2 3; do sleep 0.2; +- waybar-reserve; done``) counts the same as a bare ``exec = waybar-reserve``.""" ++ Matches waybar-reserve invoked anywhere in an ``at_reload(...)`` line, so the ++ reload-race-safe retry-loop form (``at_reload("for i in 1 2 3; do sleep 0.2; ++ waybar-reserve; done")``) counts the same as a bare ``at_reload("waybar-reserve")``. ++ ++ at_reload and not at_start, because hyprlang's ``exec`` ran at startup AND on ++ every reload, and the Lua port splits those two jobs. at_reload's handler is ++ ``hl.on("config.reloaded")``, which fires on the initial load as well, so it ++ alone covers both. Matching at_start instead would pass while reservations ++ died on the next reload -- which is exactly the defect this guards.""" + with open(HYPR) as f: +- return any(re.match(r"\s*exec\s*=.*\bwaybar-reserve\b", ln) for ln in f) ++ return any(re.match(r'\s*at_reload\(".*\bwaybar-reserve\b', ln) for ln in f) + + + def reserve_target(): +@@ -61,7 +67,7 @@ class ReservePairingHarness(unittest.TestCase): + + def test_hyprland_wires_the_reserve_script(self): + self.assertTrue(reserve_exec_wired(), +- "hyprland.conf lacks `exec = waybar-reserve`: reservations " ++ "hyprland.lua lacks `at_reload(... waybar-reserve ...)`: reservations " + "die on the next config reload") + + def test_reserve_target_covers_the_bar(self): diff --git a/working/hyprland-lua-port/test-desktop-for-lua.patch b/working/hyprland-lua-port/test-desktop-for-lua.patch new file mode 100644 index 0000000..5bf7612 --- /dev/null +++ b/working/hyprland-lua-port/test-desktop-for-lua.patch @@ -0,0 +1,50 @@ +diff --git a/scripts/testing/tests/test_desktop.py b/scripts/testing/tests/test_desktop.py +index 1538d6a..ca2f266 100644 +--- a/scripts/testing/tests/test_desktop.py ++++ b/scripts/testing/tests/test_desktop.py +@@ -11,6 +11,8 @@ installs `awww` (swww successor) and `pacman -Q swww` no longer matches — so + this checks awww. That divergence from the shell sweep is a correctness fix. + """ + ++import re ++ + import pytest + + +@@ -19,8 +21,12 @@ HYPRLAND_TOOLS = [ + "awww", "grim", "slurp", "gammastep", "foot", + ] + ++# Note: hyprland.lua sits beside two .conf files on purpose. Only Hyprland's own ++# config format is deprecated (removed in 0.57); hypridle and hyprlock are ++# separate projects still on hyprlang, so their .conf names are current. Don't ++# "fix" this list for consistency -- that reverts the Lua port. + HYPRLAND_CONFIGS = [ +- ".config/hypr/hyprland.conf", ++ ".config/hypr/hyprland.lua", + ".config/hypr/hypridle.conf", + ".config/hypr/hyprlock.conf", + ".config/waybar/config", +@@ -139,8 +145,19 @@ def test_bt_panel_wired(host, hyprland_installed, home): + waybar = host.file("%s/.config/waybar/config" % home) + assert "custom/bluetooth" in waybar.content_string, \ + "waybar config lacks the custom/bluetooth module" +- hyprconf = host.file("%s/.config/hypr/hyprland.conf" % home) +- assert "bt-panel" in hyprconf.content_string, \ +- "hyprland.conf lacks the bt-panel keybind" ++ hypr = host.file("%s/.config/hypr/hyprland.lua" % home).content_string ++ # Assert the dispatcher and the chord, not the bare binary name. "bt-panel" ++ # on its own still appears inside exec_cmd() when the chord that reaches it ++ # is mangled, so the weak form passes on a config where Super+Shift+B does ++ # nothing. That is not hypothetical: converting this config to Lua produced ++ # exactly that defect elsewhere ("CTRL" .. mod .. " + S" collapsing to an ++ # unparseable CTRLSUPER + S, which Hyprland rejects at parse and leaves dead). ++ # One pattern spanning the whole bind, not two independent checks: asserting ++ # the chord and the dispatcher separately passes when they sit on different ++ # lines, and a planned reshuffle of the panel keybinding family is exactly ++ # the change that would separate them. ++ assert re.search( ++ r'hl\.bind\(\s*mod\s*\.\.\s*" \+ SHIFT \+ B",\s*hl\.dsp\.exec_cmd\("bt-panel"\)', ++ hypr), "hyprland.lua lacks the Super+Shift+B bind reaching bt-panel" + assert host.file("%s/.config/themes/dupre/panel.css" % home).exists, \ + "shared panel css missing from the stowed theme" diff --git a/working/hyprland-lua-port/velox-local.lua b/working/hyprland-lua-port/velox-local.lua new file mode 100644 index 0000000..1da7359 --- /dev/null +++ b/working/hyprland-lua-port/velox-local.lua @@ -0,0 +1,64 @@ +-- velox — Framework 13, HiDPI 2256x1504. Sourced via conf.d/*.lua glob; +-- values here override hyprland.lua (last wins). + +hl.monitor({ + output = "eDP-1", + mode = "preferred", + position = "auto", + scale = "1.566667", +}) + +-- Scaling on this panel comes from the compositor alone. +-- +-- There used to be `env = GDK_SCALE,1.5` and `env = QT_SCALE_FACTOR,1.5` here, +-- to compensate for `xwayland:force_zero_scaling = true` in hyprland.lua, which +-- makes XWayland clients render unscaled and therefore tiny. That worked for +-- XWayland and broke everything else, because env vars reach every app, not just +-- the XWayland ones. Qt 6 and GTK on Wayland already take their scale from the +-- compositor, so the factor multiplied the monitor scale above: Qt saw a 960x640 +-- logical screen instead of 1440x960 and drew half again too large. Measured +-- with QScreen.geometry on 2026-08-19; 2256/1.566667 is exactly 1440, which is +-- what Qt reports once nothing overrides it. +-- +-- Turning force_zero_scaling off for this host is the other half. XWayland then +-- scales through the compositor, so those apps come out the right size and a +-- little soft rather than sharp and tiny. That cost lands only on XWayland, +-- which is the thing I avoid anyway, and correctly-sized beats crisp for the +-- occasional Zoom window the browser spawns. +-- +-- ratio needs neither: its monitor runs at scale 1, so nothing multiplies and +-- XWayland has nothing to compensate for. +hl.config({ + xwayland = { + force_zero_scaling = false, + }, +}) + +-- calibre renders oversized at the 1.57 compositor scale. Pin its own DPI so its +-- UI is comfortable without touching the desktop or other apps (validated at 96, +-- 2026-06-27). CALIBRE_OVERRIDE_DPI is calibre-only, so a session-wide env is safe. +hl.env("CALIBRE_OVERRIDE_DPI", "96") +-- No XCURSOR_SIZE here: Hyprland scales the native cursor by the monitor scale +-- already, so the shared default (24) renders correctly on HiDPI. Pre-scaling it +-- (the old 36 = 24 x 1.5) double-applied on top of the compositor's scale. + +-- org-capture popup: match the terminal scratchpad (75% x 70% of the 1437x958 +-- logical desktop = 1078 x 671 px). Native window rules ignore percentages, so +-- it's pinned in pixels here rather than in the shared hyprland.lua. +hl.window_rule({ + match = { + title = "^(org-capture)$", + }, + size = "1078 671", +}) + +-- Growth cap (Craig, 2026-07-14): the frame must never outgrow its pinned +-- size even when capture content pushes it. 1078 px is ~98 Emacs columns, +-- already under the 120-column cap ratio uses. +hl.window_rule({ + match = { + title = "^(org-capture)$", + }, + max_size = "1078 671", +}) + diff --git a/working/meeting-transcription-service/2026-09-19-work-handoff.org b/working/meeting-transcription-service/2026-09-19-work-handoff.org new file mode 100644 index 0000000..d4da1d9 --- /dev/null +++ b/working/meeting-transcription-service/2026-09-19-work-handoff.org @@ -0,0 +1,44 @@ +#+TITLE: Handoff from work: the meeting transcription service needs an install home +#+AUTHOR: Craig Jennings + +* What this is +A self-hosted meeting transcription service (whisper.cpp plus pyannote speaker diarization) that runs on +ratio, with velox as the offline fallback. It has been running on both machines since 2026-09-17. The code +arrives as meeting-transcription-service.tar.gz alongside this note. I picked archsetup as its home because +it is machine setup: a worker, two systemd user units, a Python venv and a model file. The client script and +its Emacs backend entry are handled separately. + +* How it works +- A job queue under ~/.local/state/meeting-transcribe/ with incoming/, work/, done/ and failed/. A systemd + user path unit (meeting-transcribe.path) starts a oneshot worker (meeting-transcribe.service) when a job + lands in incoming/. +- The worker (src/transcribe-worker) converts audio to 16 kHz mono, runs whisper-cli at word level, runs + pyannote (speaker-diarization-community-1) through src/diarize.py with the job's speaker count, and + merges the two by timestamp with src/merge_transcript.py. One job at a time, behind a lock. +- A finished job leaves done/<id>.txt. A failed job leaves failed/<id>.log. Nothing half-written reaches done/. +- No network listener. Tailscale ssh is the transport and the login, systemd is the daemon, the + filesystem is the queue. + +* What the install has to provide, per machine +- ~/.local/share/pyannote-diarize/.venv :: Python 3.12, torch (CPU build), pyannote.audio 4.0.7. About + 1.3 GB. Built with uv. +- ~/.local/share/whisper-models/ggml-large-v3-turbo-q5_0.bin :: the whisper model; ratio's and velox's + copies have the same checksum. whisper-cpp itself must be installed (it was already on velox). +- The src/ scripts placed where the units expect them, and the two user units enabled. +- Linger on, so the path unit runs without a login session. +- One-time, online, by hand: a Hugging Face token and acceptance of the pyannote model terms, so the + diarization model can be cached. After that neither machine needs Hugging Face. The token is a + credential: it must not be written into this repo, which is publicly cloneable. + +* State today +- Installed by hand on ratio and velox; both verified. No NVIDIA GPU and no ROCm on ratio, so it is CPU only. +- make check in the bundle runs pytest (129 tests; whisper, pyannote, ssh and scp are faked at the process + boundary), pyright and shellcheck. All green on 2026-09-19. +- Known rough edge: when two local runs overlap, the second finds the lock held, the worker returns + silently, and the client says "finished without producing a transcript". Rerunning fixes it, but the + message should say the lock was held. + +* What I'm asking archsetup for +Take ownership of the service side: the worker, diarize.py, merge_transcript.py, the two units, and an +install step that builds the venv, fetches the model and enables the units on both daily drivers. Keep +anything work-specific out of the repo; the code bundle has none (scanned before sending). diff --git a/working/meeting-transcription-service/Makefile b/working/meeting-transcription-service/Makefile new file mode 100644 index 0000000..2ed7dad --- /dev/null +++ b/working/meeting-transcription-service/Makefile @@ -0,0 +1,15 @@ +# Checks for the transcription service. transcribe-worker has no .py suffix, so +# pyright's directory scan skips it; it is named explicitly here. +.PHONY: check test types lint + +check: test types lint + +test: + python3 -m pytest tests -q + +types: + pyright + pyright src/transcribe-worker + +lint: + shellcheck src/ratio-transcribe diff --git a/working/meeting-transcription-service/pyrightconfig.json b/working/meeting-transcription-service/pyrightconfig.json new file mode 100644 index 0000000..8228120 --- /dev/null +++ b/working/meeting-transcription-service/pyrightconfig.json @@ -0,0 +1,4 @@ +{ + "extraPaths": ["src"], + "include": ["src", "tests"] +} diff --git a/working/meeting-transcription-service/src/diarize.py b/working/meeting-transcription-service/src/diarize.py new file mode 100644 index 0000000..a6ae85e --- /dev/null +++ b/working/meeting-transcription-service/src/diarize.py @@ -0,0 +1,114 @@ +#!/usr/bin/env python3 +"""Run pyannote speaker diarization on one audio file and write the turns as JSON. + +Usage: diarize.py AUDIO OUT_JSON [--speakers N | --min-speakers N --max-speakers N] + +Output is a list of {"start", "end", "speaker"} in seconds, which is what +merge_transcript.py reads. HF_TOKEN is only needed the first time, to download +the gated model; after that the cached copy loads offline. + +pyannote and torch are imported inside run(), so the pure helpers here can be +tested without the multi-gigabyte environment. +""" + +from __future__ import annotations + +import argparse +import json +import os +import sys +from collections.abc import Iterable +from pathlib import Path +from typing import Any + +MODEL = "pyannote/speaker-diarization-community-1" + + +def turns_from_tracks(tracks: Iterable[tuple[Any, Any, str]]) -> list[dict[str, Any]]: + """Convert pyannote (segment, track, label) triples into sorted turn dicts. + + Times are rounded to milliseconds. Segments with no length are dropped. + """ + turns = [ + {"start": round(float(seg.start), 3), "end": round(float(seg.end), 3), "speaker": str(label)} + for seg, _track, label in tracks + if float(seg.end) > float(seg.start) + ] + return sorted(turns, key=lambda t: (t["start"], t["end"])) + + +def _positive_int(value: str) -> int: + number = int(value) + if number < 1: + raise argparse.ArgumentTypeError("must be 1 or more") + return number + + +def parse_args(argv: list[str]) -> argparse.Namespace: + """Parse the command line. Exits with usage on bad or contradictory counts.""" + parser = argparse.ArgumentParser(description="Speaker diarization with pyannote.") + parser.add_argument("audio") + parser.add_argument("out") + parser.add_argument("--speakers", type=_positive_int, help="exact number of speakers") + parser.add_argument("--min-speakers", type=_positive_int) + parser.add_argument("--max-speakers", type=_positive_int) + args = parser.parse_args(argv) + if args.speakers is not None and (args.min_speakers or args.max_speakers): + parser.error("--speakers cannot be combined with --min-speakers/--max-speakers") + if args.min_speakers and args.max_speakers and args.min_speakers > args.max_speakers: + parser.error("--min-speakers cannot exceed --max-speakers") + return args + + +def pipeline_kwargs(args: argparse.Namespace) -> dict[str, int]: + """Only the speaker-count options that were actually given.""" + options = { + "num_speakers": args.speakers, + "min_speakers": args.min_speakers, + "max_speakers": args.max_speakers, + } + return {name: value for name, value in options.items() if value is not None} + + +def run(args: argparse.Namespace) -> list[dict[str, Any]]: + """Load the pipeline, diarize the audio, and return the turns.""" + # Heavy, and only installed in the service venv, so imported here on purpose. + from pyannote.audio import Pipeline # pyright: ignore[reportMissingImports] + + pipeline = Pipeline.from_pretrained(MODEL, token=os.environ.get("HF_TOKEN") or None) + if pipeline is None: + raise RuntimeError(f"could not load {MODEL}: check HF_TOKEN and that its terms are accepted") + output = pipeline(args.audio, **pipeline_kwargs(args)) + # The exclusive variant never overlaps two speakers, which is what a + # word-by-word merge wants. Older pipelines return the annotation itself. + annotation = getattr(output, "exclusive_speaker_diarization", None) + if annotation is None: + annotation = getattr(output, "speaker_diarization", output) + return turns_from_tracks(annotation.itertracks(yield_label=True)) + + +def main(argv: list[str]) -> int: + """CLI entry point. Writes OUT_JSON atomically; non-zero on any failure.""" + args = parse_args(argv) + if not Path(args.audio).is_file(): + print(f"Error: audio file not found: {args.audio}", file=sys.stderr) + return 1 + try: + turns = run(args) + except Exception as err: # noqa: BLE001 - report any model failure and exit non-zero + print(f"Error: diarization failed: {err}", file=sys.stderr) + return 1 + if not turns: + print("Error: diarization found no speech", file=sys.stderr) + return 1 + out = Path(args.out) + partial = out.with_name(out.name + ".partial") + partial.write_text(json.dumps(turns), encoding="utf-8") + partial.replace(out) + speakers = len({t["speaker"] for t in turns}) + print(f"{len(turns)} turns, {speakers} speakers -> {out}", file=sys.stderr) + return 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/working/meeting-transcription-service/src/merge_transcript.py b/working/meeting-transcription-service/src/merge_transcript.py new file mode 100644 index 0000000..13aec75 --- /dev/null +++ b/working/meeting-transcription-service/src/merge_transcript.py @@ -0,0 +1,329 @@ +#!/usr/bin/env python3 +"""Merge whisper word timings with speaker turns into transcript lines. + +Input is two JSON files: whisper-cli's ``-oj`` output (run at word level) and +the diarizer's list of speaker turns. Output is one line per stretch of speech, +``HH:MM:SS Speaker A: text``, the same shape the hosted services produced. + +Standard library only, so it runs under any Python 3.10+ without the venv. +""" + +from __future__ import annotations + +import json +import sys +from dataclasses import dataclass +from pathlib import Path + +DEFAULT_MAX_GAP_S = 3.0 +DEFAULT_SPEECH_MARGIN_S = 2.0 +DEFAULT_MAX_LOOP_S = 30.0 + + +@dataclass(frozen=True) +class Unit: + """A piece of transcribed text with its start and end in seconds.""" + + start: float + end: float + text: str + + +@dataclass(frozen=True) +class Turn: + """A stretch of audio the diarizer attributes to one speaker.""" + + start: float + end: float + speaker: str + + +def _timestamp(seconds: float) -> str: + """Render seconds as HH:MM:SS, floored.""" + whole = int(seconds) + return f"{whole // 3600:02d}:{whole % 3600 // 60:02d}:{whole % 60:02d}" + + +def _speaker_name(index: int) -> str: + """Name speakers A-Z in order of first speech, then by number.""" + return chr(ord("A") + index) if index < 26 else str(index) + + +def _speaker_for(unit: Unit, turns: list[Turn]) -> str: + """Pick the turn a unit belongs to. + + The turn overlapping most of the unit wins. A unit overlapping nothing (a + zero-length word, or one whisper timed into a silence) goes to the turn + nearest its midpoint, because whisper's word timings drift by a few hundred + milliseconds and dropping the word would be worse than a near guess. + """ + best = max(turns, key=lambda t: min(unit.end, t.end) - max(unit.start, t.start)) + if min(unit.end, best.end) - max(unit.start, best.start) > 0: + return best.speaker + mid = (unit.start + unit.end) / 2 + + def distance(turn: Turn) -> float: + if turn.start <= mid <= turn.end: + return 0.0 + return min(abs(mid - turn.start), abs(mid - turn.end)) + + return min(turns, key=distance).speaker + + +def merge(units: list[Unit], turns: list[Turn], max_gap_s: float = DEFAULT_MAX_GAP_S) -> list[str]: + """Return transcript lines for ``units`` labelled by ``turns``. + + Consecutive units from one speaker share a line. The line breaks when the + speaker changes, or when the speaker pauses longer than ``max_gap_s``, so a + long monologue still carries usable timestamps. + + Raises: + ValueError: if there are no turns, no spoken words, or ``max_gap_s`` is negative. + """ + if max_gap_s < 0: + raise ValueError("max_gap_s must not be negative") + spoken = sorted((u for u in units if u.text.strip()), key=lambda u: (u.start, u.end)) + if not spoken: + raise ValueError("no speech: the transcription holds no words") + if not turns: + raise ValueError("no speaker turns: the diarization is empty") + ordered_turns = sorted(turns, key=lambda t: (t.start, t.end)) + + names: dict[str, str] = {} + lines: list[tuple[float, str, list[str]]] = [] + previous_end = 0.0 + for unit in spoken: + speaker = _speaker_for(unit, ordered_turns) + name = names.setdefault(speaker, _speaker_name(len(names))) + if lines and lines[-1][1] == name and unit.start - previous_end <= max_gap_s: + lines[-1][2].append(unit.text) + else: + lines.append((unit.start, name, [unit.text])) + previous_end = max(previous_end, unit.end) + + return [ + f"{_timestamp(start)} Speaker {name}: {' '.join(''.join(parts).split())}" + for start, name, parts in lines + ] + + +def drop_outside_speech( + units: list[Unit], turns: list[Turn], margin_s: float = DEFAULT_SPEECH_MARGIN_S +) -> tuple[list[Unit], int]: + """Return the units that belong to speech, and how many were dropped. + + Whisper invents words ("Thank you.") when it is handed silence. The diarizer + marks where people actually spoke, so a unit that touches no turn and sits + more than ``margin_s`` from the nearest one is treated as invented. The margin + protects real words the diarizer clipped off the edge of a turn. + + Raises: + ValueError: if there are no turns, or ``margin_s`` is negative. + """ + if margin_s < 0: + raise ValueError("margin_s must not be negative") + if not turns: + raise ValueError("no speaker turns: the diarization is empty") + + def gap(unit: Unit) -> float: + # Seconds between the unit and its nearest turn; zero when they touch. + # Rounded to the millisecond, the resolution of whisper's offsets. + nearest = min(max(turn.start - unit.end, unit.start - turn.end, 0.0) for turn in turns) + return round(nearest, 3) + + kept = [unit for unit in units if gap(unit) <= margin_s] + return kept, len(units) - len(kept) + + +def _unit(item: dict) -> Unit: + """A Unit from a whisper segment or token dict (offsets in milliseconds). + + whisper-cli clamps a token's start to its segment's start without moving the + end, so some tokens arrive ending before they begin. Those become zero-length + at their start; left alone they corrupt both the overlap and the pause maths. + """ + start = item["offsets"]["from"] / 1000 + end = item["offsets"]["to"] / 1000 + return Unit(start, max(start, end), item["text"]) + + +def find_repetition( + text: str, min_words: int = 3, max_words: int = 12, min_repeats: int = 4 +) -> tuple[str, int] | None: + """Find a phrase repeated back to back, whisper's hallucination signature. + + Returns the phrase and its repeat count, or None. Four consecutive repeats of + a phrase of three or more words is the line: people say a thing two or three + times, and single-word runs ("yeah yeah yeah") are ordinary speech. + """ + words = text.split() + keys = [w.lower().strip(".,!?;:\"'") for w in words] + for size in range(min_words, max_words + 1): + for i in range(len(keys) - size * min_repeats + 1): + phrase = keys[i : i + size] + if len(set(phrase)) < 2: + continue + count = 1 + while keys[i + count * size : i + (count + 1) * size] == phrase: + count += 1 + if count >= min_repeats: + return " ".join(words[i : i + size]), count + return None + + +def collapse_repetitions( + units: list[Unit], + min_words: int = 3, + max_words: int = 12, + min_repeats: int = 4, + max_loop_s: float = DEFAULT_MAX_LOOP_S, +) -> tuple[list[Unit], list[tuple[str, int]]]: + """Collapse whisper's repetition loops, keeping one copy of the phrase. + + Whisper sometimes gets stuck and emits the same phrase over and over. A short + loop (up to ``max_loop_s`` of audio) costs a few seconds of speech, so it is + collapsed to a single occurrence and reported. A longer one means real speech + was lost for a stretch, and that is raised instead, so the job fails rather + than hand back a transcript with a hole in it. + + Returns the surviving units and a list of (phrase, repeat count) for every + loop collapsed. Units are matched word by word, so a phrase spread over word + units and a phrase sitting in one segment unit are both found. + + Raises: + ValueError: if a loop lasts longer than ``max_loop_s``, or the limit is negative. + """ + if max_loop_s < 0: + raise ValueError("max_loop_s must not be negative") + units = list(units) + collapsed: list[tuple[str, int]] = [] + + def find() -> tuple[int, int, int] | None: + # (first word index, phrase size, repeat count) of the earliest loop, or None + words = [(w, ui) for ui, u in enumerate(units) for w in u.text.split()] + keys = [w.lower().strip(".,!?;:\"'") for w, _ in words] + best: tuple[int, int, int] | None = None + for size in range(min_words, max_words + 1): + for i in range(len(keys) - size * min_repeats + 1): + if best is not None and i >= best[0]: + break + phrase = keys[i : i + size] + if len(set(phrase)) < 2: + continue + count = 1 + while keys[i + count * size : i + (count + 1) * size] == phrase: + count += 1 + if count >= min_repeats: + best = (i, size, count) + break + return best + + while True: + hit = find() + if hit is None: + return units, collapsed + i, size, count = hit + words = [(w, ui) for ui, u in enumerate(units) for w in u.text.split()] + phrase_text = " ".join(w for w, _ in words[i : i + size]) + doomed = set(range(i + size, i + count * size)) # word indexes of the repeats + touched = {ui for wi, (_, ui) in enumerate(words) if wi in doomed or i <= wi < i + size} + span_start = min(units[ui].start for ui in touched) + span_end = max(units[ui].end for ui in touched) + duration = round(span_end - span_start, 3) + if duration > max_loop_s: + raise ValueError( + f"whisper looped: {phrase_text!r} repeats {count} times over {duration:.0f} s; " + "rerun whisper with -mc 0" + ) + # Rebuild every touched unit from the words it keeps. A unit that held only + # repeats disappears; one that also held the first copy or later speech keeps + # those words. Every pass removes (count - 1) * size words, so this ends. + last_kept = max(ui for wi, (_, ui) in enumerate(words) if i <= wi < i + size) + rebuilt: list[Unit] = [] + for ui, unit in enumerate(units): + if ui not in touched: + rebuilt.append(unit) + continue + keep = [w for wi, (w, wui) in enumerate(words) if wui == ui and wi not in doomed] + if not keep: + continue + # The kept copy takes over the time the loop occupied, so the merge does + # not read the removed stretch as a pause and break the line there. + end = max(unit.end, span_end) if ui == last_kept else unit.end + rebuilt.append(Unit(unit.start, end, " " + " ".join(keep))) + units = rebuilt + collapsed.append((phrase_text, count)) + + +def load_whisper_json(path: str | Path) -> list[Unit]: + """Read whisper-cli JSON output. Offsets there are in milliseconds. + + With ``-ojf`` each segment carries its tokens and their offsets; those become + word-level units, which is what lets a speaker change land mid-segment. Plain + ``-oj`` output, or a segment with no tokens, falls back to the segment itself. + + Raises: + ValueError: if the file is not whisper's JSON shape. + """ + try: + data = json.loads(Path(path).read_text(encoding="utf-8")) + units: list[Unit] = [] + for item in data["transcription"]: + tokens = item.get("tokens") or [] + if not tokens: + units.append(_unit(item)) + continue + for token in tokens: + text = token["text"] + if not text or text.startswith("[_"): # [_BEG_], [_TT_123], [_EOT_] + continue + unit = _unit(token) + if units and not text.startswith(" "): + # A sub-word piece or punctuation: it belongs to the word before it. + previous = units[-1] + units[-1] = Unit(previous.start, max(previous.end, unit.end), previous.text + text) + else: + units.append(unit) + return units + except OSError as err: + raise ValueError(f"{path}: cannot read whisper output ({err.strerror})") from err + except (json.JSONDecodeError, KeyError, TypeError) as err: + raise ValueError(f"{path}: not whisper-cli JSON output ({err!r})") from err + + +def load_turns_json(path: str | Path) -> list[Turn]: + """Read the diarizer's turns: a list of {start, end, speaker}, in seconds. + + Raises: + ValueError: if the file is not that shape. + """ + try: + data = json.loads(Path(path).read_text(encoding="utf-8")) + if not isinstance(data, list): + raise TypeError("expected a list of turns") + return [Turn(float(t["start"]), float(t["end"]), str(t["speaker"])) for t in data] + except OSError as err: + raise ValueError(f"{path}: cannot read speaker turns ({err.strerror})") from err + except (json.JSONDecodeError, KeyError, TypeError, ValueError) as err: + raise ValueError(f"{path}: not a speaker-turns file ({err!r})") from err + + +def main(argv: list[str]) -> int: + """CLI: ``merge_transcript.py WHISPER_JSON TURNS_JSON`` prints the transcript.""" + if len(argv) != 2: + print("usage: merge_transcript.py WHISPER_JSON TURNS_JSON", file=sys.stderr) + return 2 + try: + turns = load_turns_json(argv[1]) + units, _dropped = drop_outside_speech(load_whisper_json(argv[0]), turns) + units, _collapsed = collapse_repetitions(units) + lines = merge(units, turns) + except ValueError as err: + print(f"Error: {err}", file=sys.stderr) + return 1 + print("\n".join(lines)) + return 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/working/meeting-transcription-service/src/ratio-transcribe b/working/meeting-transcription-service/src/ratio-transcribe new file mode 100755 index 0000000..8db59b6 --- /dev/null +++ b/working/meeting-transcription-service/src/ratio-transcribe @@ -0,0 +1,203 @@ +#!/usr/bin/env bash +# ratio-transcribe - Transcribe audio on my own transcription host, with speaker labels +# Usage: ratio-transcribe <audio-file> [language] +# +# Same contract as assemblyai-transcribe: the transcript goes to stdout, one line +# per speaker turn ("HH:MM:SS Speaker A: text"); progress and errors go to stderr; +# any failure exits non-zero with nothing on stdout. +# +# The work happens on a host that runs the meeting-transcribe queue (whisper-cpp +# plus pyannote). This script copies the audio over ssh, drops a job into the +# queue, waits, and prints the result. The job id is a hash of the audio and its +# options, so if the connection drops or the laptop sleeps, running the same +# command again just collects the finished transcript. If the host can't be +# reached at all, the same queue and worker run on this machine instead. +# +# Optional environment: +# SPEAKERS exact number of speakers, when you know it +# MIN_SPEAKERS, MAX_SPEAKERS a range instead +# TRANSCRIBE_HOST ssh name of the host (default: ratio) +# TRANSCRIBE_TIMEOUT seconds to wait for the job (default: 3600) +# TRANSCRIBE_POLL seconds between checks (default: 10) +# TRANSCRIBE_LOCAL=1 skip the host and run here +# TRANSCRIBE_WORKER path to the local worker + +set -euo pipefail + +AUDIO="${1:-}" +LANG_CODE="${2:-en}" +HOST="${TRANSCRIBE_HOST:-ratio}" +TIMEOUT="${TRANSCRIBE_TIMEOUT:-3600}" +POLL="${TRANSCRIBE_POLL:-10}" +WORKER="${TRANSCRIBE_WORKER:-$HOME/.local/share/pyannote-diarize/src/transcribe-worker}" +STATE=".local/state/meeting-transcribe" # relative to the home directory, on either machine + +if [[ -z "$AUDIO" ]]; then + echo "Usage: ratio-transcribe <audio-file> [language]" >&2 + echo "Example: SPEAKERS=3 ratio-transcribe meeting.m4a en" >&2 + exit 1 +fi + +if [[ ! -f "$AUDIO" ]]; then + echo "Error: Audio file not found: $AUDIO" >&2 + exit 1 +fi +# scp reads "name:with:colons" as host:path; an absolute path removes the ambiguity. +AUDIO="$(realpath -- "$AUDIO")" + +# Everything below ends up in a job file and on command lines, so check it first. +if [[ ! "$LANG_CODE" =~ ^[A-Za-z]{2,8}(-[A-Za-z0-9]{1,8})*$ ]]; then + echo "Error: Invalid language code: $LANG_CODE" >&2 + exit 1 +fi + +for name in SPEAKERS MIN_SPEAKERS MAX_SPEAKERS; do + value="${!name:-}" + if [[ -n "$value" && ! "$value" =~ ^[1-9][0-9]*$ ]]; then + echo "Error: $name must be a positive whole number of speakers, got: $value" >&2 + exit 1 + fi +done +if [[ -n "${SPEAKERS:-}" && ( -n "${MIN_SPEAKERS:-}" || -n "${MAX_SPEAKERS:-}" ) ]]; then + echo "Error: give an exact SPEAKERS count or a MIN/MAX speaker range, not both" >&2 + exit 1 +fi +if [[ -n "${MIN_SPEAKERS:-}" && -n "${MAX_SPEAKERS:-}" ]] && (( MIN_SPEAKERS > MAX_SPEAKERS )); then + echo "Error: MIN_SPEAKERS cannot exceed MAX_SPEAKERS (speaker range)" >&2 + exit 1 +fi + +for tool in jq sha256sum; do + if ! command -v "$tool" &> /dev/null; then + echo "Error: $tool command not found" >&2 + exit 1 + fi +done + +EXT="${AUDIO##*.}" +[[ "$EXT" =~ ^[A-Za-z0-9]{1,5}$ ]] || EXT="bin" +EXT="${EXT,,}" + +if [[ -n "${SPEAKERS:-}" ]]; then + COUNT_TAG="s${SPEAKERS}" +elif [[ -n "${MIN_SPEAKERS:-}${MAX_SPEAKERS:-}" ]]; then + COUNT_TAG="r${MIN_SPEAKERS:-x}-${MAX_SPEAKERS:-x}" +else + COUNT_TAG="auto" +fi +JOB_ID="$(sha256sum "$AUDIO" | cut -c1-16)-${LANG_CODE,,}-${COUNT_TAG}" + +JOB_JSON=$(jq -cn \ + --arg language "$LANG_CODE" \ + --arg name "$(basename "$AUDIO")" \ + --arg speakers "${SPEAKERS:-}" --arg min "${MIN_SPEAKERS:-}" --arg max "${MAX_SPEAKERS:-}" \ + '{language: $language} + + (if $speakers != "" then {speakers: ($speakers | tonumber)} else {} end) + + (if $min != "" then {min_speakers: ($min | tonumber)} else {} end) + + (if $max != "" then {max_speakers: ($max | tonumber)} else {} end) + + {original_name: $name}') + +# ssh reads stdin unless told not to, which would swallow the input of any loop +# this script is called from. Only the job-file upload needs stdin. +remote() { ssh -n -o BatchMode=yes -o ConnectTimeout=8 "$HOST" "$@"; } +remote_with_stdin() { ssh -o BatchMode=yes -o ConnectTimeout=8 "$HOST" "$@"; } + +# One word for where the job stands on the host: done, failed, queued or new. +remote_status() { + remote "cd $STATE 2>/dev/null || { echo new; exit 0; } + if [ -e done/$JOB_ID.txt ]; then echo done + elif [ -e failed/$JOB_ID.log ]; then echo failed + elif [ -d incoming/$JOB_ID ] || [ -d work/$JOB_ID ]; then echo queued + else echo new; fi" +} + +print_transcript() { # $1 = the transcript text + if [[ -z "${1//[[:space:]]/}" ]]; then + echo "Error: the transcript came back empty" >&2 + exit 1 + fi + echo "Transcription complete! (${SECONDS}s total)" >&2 + printf '%s\n' "$1" +} + +run_remote() { + local status + status=$(remote_status) + + if [[ "$status" == "failed" ]]; then + echo "An earlier attempt at this job failed; trying again..." >&2 + remote "rm -f $STATE/failed/$JOB_ID.log" + status="new" + fi + + if [[ "$status" == "new" ]]; then + echo "Uploading audio file to $HOST..." >&2 + # Copy into uploading/, then rename into incoming/. The queue only ever sees + # a complete job. + remote "mkdir -p $STATE/incoming $STATE/uploading/$JOB_ID" + scp -q -o BatchMode=yes "$AUDIO" "$HOST:$STATE/uploading/$JOB_ID/audio.$EXT" < /dev/null + printf '%s' "$JOB_JSON" | remote_with_stdin "cat > $STATE/uploading/$JOB_ID/job.json" + remote "mv $STATE/uploading/$JOB_ID $STATE/incoming/$JOB_ID" + echo "Job $JOB_ID queued. Waiting for completion..." >&2 + elif [[ "$status" == "queued" ]]; then + echo "Job $JOB_ID is already queued on $HOST. Waiting for completion..." >&2 + fi + + while true; do + # A dropped connection is not a failed job; keep asking until the timeout. + status=$(remote_status 2> /dev/null) || status="unreachable" + case "$status" in + done) + print_transcript "$(remote "cat $STATE/done/$JOB_ID.txt")" + return 0 + ;; + failed) + echo "Error: transcription failed on $HOST" >&2 + remote "cat $STATE/failed/$JOB_ID.log" >&2 || true + exit 1 + ;; + esac + if (( SECONDS >= TIMEOUT )); then + echo "Error: no result after ${TIMEOUT}s. The job is still with $HOST;" >&2 + echo "run the same command again to collect the transcript." >&2 + exit 1 + fi + sleep "$POLL" + [[ "$status" == "unreachable" ]] || echo "Processing... (${SECONDS}s elapsed)" >&2 + done +} + +run_local() { + if [[ ! -x "$WORKER" ]]; then + echo "Error: $HOST is unreachable and there is no local worker at $WORKER" >&2 + exit 1 + fi + local state="$HOME/$STATE" + if [[ ! -s "$state/done/$JOB_ID.txt" ]]; then + echo "Running the transcription locally (this machine is slower; expect a wait)..." >&2 + rm -f "$state/failed/$JOB_ID.log" + rm -rf "$state/uploading/$JOB_ID" + mkdir -p "$state/incoming" "$state/uploading/$JOB_ID" + cp "$AUDIO" "$state/uploading/$JOB_ID/audio.$EXT" + printf '%s' "$JOB_JSON" > "$state/uploading/$JOB_ID/job.json" + [[ -d "$state/incoming/$JOB_ID" ]] || mv "$state/uploading/$JOB_ID" "$state/incoming/$JOB_ID" + HF_HUB_OFFLINE=1 "$WORKER" >&2 < /dev/null + fi + if [[ -e "$state/failed/$JOB_ID.log" ]]; then + echo "Error: local transcription failed" >&2 + cat "$state/failed/$JOB_ID.log" >&2 + exit 1 + fi + if [[ ! -e "$state/done/$JOB_ID.txt" ]]; then + echo "Error: the local worker finished without producing a transcript" >&2 + exit 1 + fi + print_transcript "$(< "$state/done/$JOB_ID.txt")" +} + +if [[ -z "${TRANSCRIBE_LOCAL:-}" ]] && remote true 2> /dev/null; then + run_remote +else + [[ -n "${TRANSCRIBE_LOCAL:-}" ]] || echo "$HOST is unreachable." >&2 + run_local +fi diff --git a/working/meeting-transcription-service/src/transcribe-worker b/working/meeting-transcription-service/src/transcribe-worker new file mode 100755 index 0000000..12e64b0 --- /dev/null +++ b/working/meeting-transcription-service/src/transcribe-worker @@ -0,0 +1,255 @@ +#!/usr/bin/env python3 +"""Drain the meeting-transcription queue, one job at a time. + +Layout under the state directory (default ~/.local/state/meeting-transcribe): + + incoming/<id>/ audio.<ext> and an optional job.json, dropped by the client + work/<id>/ the job being processed + done/<id>.txt the transcript, plus done/<id>.json with run metadata + failed/<id>.log what went wrong, by stage + +The client uploads into uploading/<id>/ and renames the folder into incoming/ when +the copy is complete, so the queue only ever lists whole jobs. A folder whose name +starts with a dot is skipped as well, as a second line of defence. Nothing is +written into done/ except by rename, so a reader never sees half a transcript. + +Standard library only. whisper-cli and ffmpeg come from PATH; the diarizer runs in +its own virtualenv and is called as a subprocess. +""" + +from __future__ import annotations + +import argparse +import fcntl +import json +import os +import re +import shutil +import subprocess +import sys +import time +from dataclasses import dataclass, field +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent)) + +import merge_transcript # noqa: E402 + +MAX_ATTEMPTS = 2 +LANGUAGE_RE = re.compile(r"^[A-Za-z]{2,8}(-[A-Za-z0-9]{1,8})*$") +SHARE = Path.home() / ".local/share" + + +@dataclass +class Config: + """Where things live and how the tools are called.""" + + home: Path = Path.home() / ".local/state/meeting-transcribe" + whisper_model: Path = SHARE / "whisper-models/ggml-large-v3-turbo-q5_0.bin" + diarize_cmd: list[str] = field( + default_factory=lambda: [ + str(SHARE / "pyannote-diarize/.venv/bin/python"), + str(Path(__file__).resolve().parent / "diarize.py"), + ] + ) + threads: int = max(1, (os.cpu_count() or 4) // 2) + + +class JobError(Exception): + """A job failed at a named stage.""" + + def __init__(self, stage: str, detail: str) -> None: + super().__init__(f"{stage}: {detail}") + self.stage = stage + self.detail = detail + + +def _run(stage: str, argv: list[str]) -> None: + """Run one tool; raise JobError carrying the tail of its stderr on failure.""" + try: + result = subprocess.run(argv, capture_output=True, text=True, check=False, stdin=subprocess.DEVNULL) + except OSError as err: + raise JobError(stage, f"cannot run {argv[0]}: {err.strerror}") from err + if result.returncode != 0: + tail = "\n".join(result.stderr.strip().splitlines()[-15:]) + raise JobError(stage, f"exit {result.returncode}\n{tail}") + + +def _read_job(folder: Path) -> dict: + """Load and validate job.json. A missing file means all defaults.""" + path = folder / "job.json" + if not path.exists(): + return {} + try: + job = json.loads(path.read_text(encoding="utf-8")) + if not isinstance(job, dict): + raise ValueError("expected an object") + except (OSError, ValueError) as err: + raise JobError("job", f"job.json is unreadable: {err}") from err + language = job.get("language", "en") + if not isinstance(language, str) or not LANGUAGE_RE.match(language): + raise JobError("job", f"job.json: bad language {language!r}") + for key in ("speakers", "min_speakers", "max_speakers"): + value = job.get(key) + if value is not None and (isinstance(value, bool) or not isinstance(value, int) or value < 1): + raise JobError("job", f"job.json: {key} must be a positive whole number, got {value!r}") + return job + + +def _diarize_options(job: dict) -> list[str]: + options: list[str] = [] + if job.get("speakers") is not None: + return ["--speakers", str(job["speakers"])] + if job.get("min_speakers") is not None: + options += ["--min-speakers", str(job["min_speakers"])] + if job.get("max_speakers") is not None: + options += ["--max-speakers", str(job["max_speakers"])] + return options + + +def process(folder: Path, config: Config) -> tuple[list[str], dict]: + """Run one job folder through the pipeline. Returns transcript lines and metadata.""" + job = _read_job(folder) + audio = next((p for p in sorted(folder.iterdir()) if p.name.startswith("audio.")), None) + if audio is None: + raise JobError("job", "no audio file in the job folder") + + wav = folder / "speech.wav" + _run("ffmpeg", ["ffmpeg", "-v", "error", "-y", "-i", str(audio), "-ar", "16000", "-ac", "1", str(wav)]) + + # -ojf keeps normal segments (better text) and adds token offsets for the merge. + # -mc 0 stops whisper feeding its own output back in, which is what sends it + # into repetition loops on long meetings. + prefix = folder / "words" + _run("whisper", [ + "whisper-cli", "-m", str(config.whisper_model), "-f", str(wav), + "-l", job.get("language", "en"), "-ojf", "-mc", "0", + "-t", str(config.threads), "-of", str(prefix), + ]) + + turns = folder / "turns.json" + _run("diarize", [*config.diarize_cmd, str(wav), str(turns), *_diarize_options(job)]) + + try: + loaded_turns = merge_transcript.load_turns_json(turns) + # Whisper invents words in silence. Drop them before the loop guard, so a + # quiet meeting isn't mistaken for a repetition loop. + units, dropped = merge_transcript.drop_outside_speech( + merge_transcript.load_whisper_json(prefix.with_suffix(".json")), loaded_turns + ) + # A short stutter is collapsed and recorded; a long loop still fails the job. + units, collapsed = merge_transcript.collapse_repetitions(units) + lines = merge_transcript.merge(units, loaded_turns) + except ValueError as err: + raise JobError("merge", str(err)) from err + + meta = { + "original_name": job.get("original_name"), + "language": job.get("language", "en"), + "speakers_requested": job.get("speakers"), + "speakers_found": len({t.speaker for t in loaded_turns}), + "lines": len(lines), + "words": sum(len(line.split()) - 3 for line in lines), + "dropped_outside_speech": dropped, + "loops_collapsed": [[phrase, count] for phrase, count in collapsed], + } + return lines, meta + + +def _write_atomic(path: Path, text: str) -> None: + partial = path.with_name(path.name + ".partial") + partial.write_text(text, encoding="utf-8") + partial.replace(path) + + +def _fail(failed_dir: Path, job_id: str, stage: str, detail: str) -> None: + """Record a failed job. Falls back to stderr if even the log cannot be written.""" + try: + _write_atomic(failed_dir / f"{job_id}.log", f"stage: {stage}\n{detail}\n") + except OSError as err: + print(f"{job_id}: {stage}: {detail} (and the failure log could not be written: {err})", file=sys.stderr) + + +def _bump_attempts(folder: Path) -> int: + """Count this run against the job, tolerating a missing or broken job.json.""" + path = folder / "job.json" + try: + job = json.loads(path.read_text(encoding="utf-8")) if path.exists() else {} + if not isinstance(job, dict): + return 1 + except (OSError, ValueError): + return 1 # _read_job reports the real problem + try: + attempts = int(job.get("attempts") or 0) + 1 + except (TypeError, ValueError): + attempts = 1 # a malformed counter counts as a first try + job["attempts"] = attempts + path.write_text(json.dumps(job), encoding="utf-8") + return attempts + + +def drain(config: Config) -> dict[str, int]: + """Process every waiting job. Returns counts of done and failed jobs.""" + dirs = {name: config.home / name for name in ("incoming", "work", "done", "failed")} + for folder in dirs.values(): + folder.mkdir(parents=True, exist_ok=True) + + counts = {"done": 0, "failed": 0} + with open(config.home / "lock", "w", encoding="utf-8") as lock: + try: + fcntl.flock(lock, fcntl.LOCK_EX | fcntl.LOCK_NB) + except OSError: + return counts # another worker is draining; it will reach these jobs + + while True: + # Jobs stranded in work/ by a crash go first, then arrivals, oldest first. + waiting = sorted( + (p for d in (dirs["work"], dirs["incoming"]) for p in d.iterdir() + if p.is_dir() and not p.name.startswith(".")), + key=lambda p: (p.parent.name != "work", p.stat().st_mtime), + ) + if not waiting: + return counts + source = waiting[0] + job_id = source.name + done_txt = dirs["done"] / f"{job_id}.txt" + if done_txt.exists(): + shutil.rmtree(source) + continue + + folder = dirs["work"] / job_id + if source != folder: + source.rename(folder) + started = time.monotonic() + try: + if _bump_attempts(folder) > MAX_ATTEMPTS: + raise JobError("worker", f"gave up after {MAX_ATTEMPTS} attempts; the job kept crashing") + lines, meta = process(folder, config) + meta["seconds"] = round(time.monotonic() - started, 1) + _write_atomic(dirs["done"] / f"{job_id}.json", json.dumps(meta)) + _write_atomic(done_txt, "\n".join(lines) + "\n") + (dirs["failed"] / f"{job_id}.log").unlink(missing_ok=True) + counts["done"] += 1 + except JobError as err: + _fail(dirs["failed"], job_id, err.stage, err.detail) + counts["failed"] += 1 + except Exception as err: # noqa: BLE001 - one bad job must never stop the queue + _fail(dirs["failed"], job_id, "worker", f"{type(err).__name__}: {err}") + counts["failed"] += 1 + finally: + shutil.rmtree(folder, ignore_errors=True) + + +def main(argv: list[str]) -> int: + """CLI: drain the queue once and exit. systemd's path unit calls this.""" + parser = argparse.ArgumentParser(description="Drain the meeting-transcription queue.") + parser.add_argument("--home", type=Path, help="state directory (default ~/.local/state/meeting-transcribe)") + args = parser.parse_args(argv) + config = Config(home=args.home) if args.home else Config() + counts = drain(config) + print(f"done={counts['done']} failed={counts['failed']}", file=sys.stderr) + return 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/working/meeting-transcription-service/systemd/meeting-transcribe.path b/working/meeting-transcription-service/systemd/meeting-transcribe.path new file mode 100644 index 0000000..821b292 --- /dev/null +++ b/working/meeting-transcription-service/systemd/meeting-transcribe.path @@ -0,0 +1,12 @@ +[Unit] +Description=Watch the meeting-transcription queue for new jobs + +[Path] +# incoming/ only ever holds complete jobs: the client uploads into uploading/ and +# renames the finished folder across. So "not empty" always means real work, and the +# worker emptying the folder is what lets this unit go quiet again. +DirectoryNotEmpty=%h/.local/state/meeting-transcribe/incoming +MakeDirectory=yes + +[Install] +WantedBy=default.target diff --git a/working/meeting-transcription-service/systemd/meeting-transcribe.service b/working/meeting-transcription-service/systemd/meeting-transcribe.service new file mode 100644 index 0000000..92f5ef4 --- /dev/null +++ b/working/meeting-transcription-service/systemd/meeting-transcribe.service @@ -0,0 +1,9 @@ +[Unit] +Description=Transcribe queued meeting recordings (whisper + pyannote) + +[Service] +Type=oneshot +ExecStart=%h/.local/share/pyannote-diarize/src/transcribe-worker +# The pyannote model is cached after its first download; never reach for the network. +Environment=HF_HUB_OFFLINE=1 +Nice=5 diff --git a/working/meeting-transcription-service/tests/test_diarize.py b/working/meeting-transcription-service/tests/test_diarize.py new file mode 100644 index 0000000..ef009b8 --- /dev/null +++ b/working/meeting-transcription-service/tests/test_diarize.py @@ -0,0 +1,74 @@ +"""Tests for diarize's pure parts. The pyannote pipeline itself is not loaded here.""" + +import sys +from collections import namedtuple +from pathlib import Path + +import pytest + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "src")) + +import diarize # noqa: E402 + +Segment = namedtuple("Segment", "start end") + + +class TestTurnsFromTracks: + def test_diarize_turns_from_tracks_converts_and_sorts(self): + """Normal: (segment, track, label) triples become sorted turn dicts.""" + tracks = [(Segment(5.0, 9.25), "_", "SPEAKER_01"), (Segment(0.5, 4.0), "_", "SPEAKER_00")] + assert diarize.turns_from_tracks(tracks) == [ + {"start": 0.5, "end": 4.0, "speaker": "SPEAKER_00"}, + {"start": 5.0, "end": 9.25, "speaker": "SPEAKER_01"}, + ] + + def test_diarize_turns_from_tracks_rounds_to_milliseconds(self): + """Boundary: float noise from the model is rounded away.""" + tracks = [(Segment(0.03096875, 1.9980000000000002), "_", "SPEAKER_00")] + assert diarize.turns_from_tracks(tracks) == [ + {"start": 0.031, "end": 1.998, "speaker": "SPEAKER_00"} + ] + + def test_diarize_turns_from_tracks_drops_empty_segments(self): + """Boundary: zero or negative length segments carry no speech.""" + tracks = [(Segment(2.0, 2.0), "_", "SPEAKER_00"), (Segment(3.0, 2.5), "_", "SPEAKER_00")] + assert diarize.turns_from_tracks(tracks) == [] + + def test_diarize_turns_from_tracks_empty_input_is_empty_list(self): + """Boundary: no tracks.""" + assert diarize.turns_from_tracks([]) == [] + + +class TestParseArgs: + def test_diarize_parse_args_speakers_sets_exact_count(self): + """Normal: --speakers pins the count.""" + args = diarize.parse_args(["a.wav", "out.json", "--speakers", "3"]) + assert (args.audio, args.out, args.speakers) == ("a.wav", "out.json", 3) + + def test_diarize_parse_args_defaults_let_the_model_estimate(self): + """Normal: no count given.""" + args = diarize.parse_args(["a.wav", "out.json"]) + assert args.speakers is None and args.min_speakers is None and args.max_speakers is None + + @pytest.mark.parametrize("argv", [ + ["a.wav", "out.json", "--speakers", "0"], + ["a.wav", "out.json", "--speakers", "-2"], + ["a.wav", "out.json", "--speakers", "three"], + ["a.wav", "out.json", "--speakers", "3", "--max-speakers", "5"], + ["a.wav", "out.json", "--min-speakers", "4", "--max-speakers", "2"], + ["a.wav"], + ]) + def test_diarize_parse_args_rejects_bad_counts(self, argv): + """Error: non-positive, non-numeric, contradictory or missing arguments.""" + with pytest.raises(SystemExit): + diarize.parse_args(argv) + + +class TestPipelineKwargs: + def test_diarize_pipeline_kwargs_only_passes_what_was_given(self): + """Normal: unset options are not forwarded to the model.""" + args = diarize.parse_args(["a.wav", "o.json", "--min-speakers", "2", "--max-speakers", "4"]) + assert diarize.pipeline_kwargs(args) == {"min_speakers": 2, "max_speakers": 4} + args = diarize.parse_args(["a.wav", "o.json", "--speakers", "3"]) + assert diarize.pipeline_kwargs(args) == {"num_speakers": 3} + assert diarize.pipeline_kwargs(diarize.parse_args(["a.wav", "o.json"])) == {} diff --git a/working/meeting-transcription-service/tests/test_merge_transcript.py b/working/meeting-transcription-service/tests/test_merge_transcript.py new file mode 100644 index 0000000..a584043 --- /dev/null +++ b/working/meeting-transcription-service/tests/test_merge_transcript.py @@ -0,0 +1,510 @@ +"""Tests for merge_transcript: whisper words + speaker turns -> transcript lines.""" + +import json +import sys +from pathlib import Path + +import pytest + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "src")) + +import merge_transcript as mt # noqa: E402 +from merge_transcript import Turn, Unit # noqa: E402 + + +def words(start, *tokens, step=0.5): + """Evenly spaced word units beginning at ``start`` seconds.""" + return [ + Unit(start + i * step, start + (i + 1) * step, f" {token}") + for i, token in enumerate(tokens) + ] + + +class TestMergeNormal: + def test_merge_transcript_merge_two_speakers_yields_one_line_per_turn(self): + """Normal: words fall into the turn that contains them.""" + units = words(0.0, "Good", "morning.") + words(6.0, "Sounds", "good.") + turns = [Turn(0.0, 5.0, "SPEAKER_00"), Turn(5.5, 9.0, "SPEAKER_01")] + assert mt.merge(units, turns) == [ + "00:00:00 Speaker A: Good morning.", + "00:00:06 Speaker B: Sounds good.", + ] + + def test_merge_transcript_merge_letters_follow_order_of_first_speech(self): + """Normal: Speaker A is whoever talks first, whatever pyannote called them.""" + units = words(0.0, "First.") + words(4.0, "Second.") + words(8.0, "Again.") + turns = [ + Turn(0.0, 3.0, "SPEAKER_02"), + Turn(3.5, 7.0, "SPEAKER_00"), + Turn(7.5, 10.0, "SPEAKER_02"), + ] + assert mt.merge(units, turns) == [ + "00:00:00 Speaker A: First.", + "00:00:04 Speaker B: Second.", + "00:00:08 Speaker A: Again.", + ] + + def test_merge_transcript_merge_same_speaker_across_turns_stays_one_line(self): + """Normal: pyannote splits a speaker's run into turns; the line does not.""" + units = words(0.0, "One", "two") + words(1.2, "three.") + turns = [Turn(0.0, 1.0, "SPEAKER_00"), Turn(1.1, 2.0, "SPEAKER_00")] + assert mt.merge(units, turns) == ["00:00:00 Speaker A: One two three."] + + def test_merge_transcript_merge_long_pause_starts_a_new_line(self): + """Normal: a pause past max_gap_s breaks the line so timestamps stay useful.""" + units = words(0.0, "Before.") + words(10.0, "After.") + turns = [Turn(0.0, 12.0, "SPEAKER_00")] + assert mt.merge(units, turns, max_gap_s=3.0) == [ + "00:00:00 Speaker A: Before.", + "00:00:10 Speaker A: After.", + ] + + def test_merge_transcript_merge_word_in_a_gap_goes_to_nearest_turn(self): + """Normal: whisper's timing drifts; a word between turns joins the closer one.""" + units = [Unit(4.6, 4.8, " late"), Unit(5.2, 5.4, " early")] + turns = [Turn(0.0, 4.5, "SPEAKER_00"), Turn(5.5, 9.0, "SPEAKER_01")] + assert mt.merge(units, turns) == [ + "00:00:04 Speaker A: late", + "00:00:05 Speaker B: early", + ] + + def test_merge_transcript_merge_overlapping_turns_pick_the_larger_overlap(self): + """Normal: with overlapped speech, the word goes where most of it sits.""" + # The anchor word fixes SPEAKER_00 as A, so the contested word's label + # actually shows which turn won: 0.2 s of overlap with A, 0.9 s with B. + units = [Unit(1.0, 1.5, " anchor"), Unit(4.0, 5.0, " contested")] + turns = [Turn(0.0, 4.2, "SPEAKER_00"), Turn(4.1, 9.0, "SPEAKER_01")] + assert mt.merge(units, turns) == [ + "00:00:01 Speaker A: anchor", + "00:00:04 Speaker B: contested", + ] + + +class TestMergeBoundary: + def test_merge_transcript_merge_timestamp_past_an_hour_is_floored(self): + """Boundary: 3725.9 s renders as 01:02:05.""" + units = [Unit(3725.9, 3726.4, " Late.")] + turns = [Turn(3700.0, 3800.0, "SPEAKER_00")] + assert mt.merge(units, turns) == ["01:02:05 Speaker A: Late."] + + def test_merge_transcript_merge_zero_length_and_blank_units_are_handled(self): + """Boundary: whisper emits empty units and zero-length words.""" + units = [Unit(0.0, 0.0, ""), Unit(2.18, 2.18, " we"), Unit(2.18, 2.54, " have")] + turns = [Turn(0.0, 5.0, "SPEAKER_00")] + assert mt.merge(units, turns) == ["00:00:02 Speaker A: we have"] + + def test_merge_transcript_merge_twenty_seventh_speaker_gets_a_number(self): + """Boundary: letters run out after Z.""" + units, turns = [], [] + for i in range(27): + units += [Unit(i * 10.0, i * 10.0 + 1, f" s{i}")] + turns += [Turn(i * 10.0, i * 10.0 + 5, f"SPEAKER_{i:02d}")] + lines = mt.merge(units, turns) + assert lines[25] == "00:04:10 Speaker Z: s25" + assert lines[26] == "00:04:20 Speaker 26: s26" + + def test_merge_transcript_merge_unicode_and_inner_spacing_survive(self): + """Boundary: non-ASCII text; whitespace collapses to single spaces.""" + units = [Unit(0.0, 0.5, " Բարև,"), Unit(0.5, 1.0, " Երևան"), Unit(1.0, 1.5, " — café.")] + turns = [Turn(0.0, 2.0, "SPEAKER_00")] + assert mt.merge(units, turns) == ["00:00:00 Speaker A: Բարև, Երևան — café."] + + def test_merge_transcript_merge_unsorted_input_is_sorted_first(self): + """Boundary: neither list has to arrive in time order.""" + units = words(6.0, "Second.") + words(0.0, "First.") + turns = [Turn(5.0, 9.0, "SPEAKER_01"), Turn(0.0, 4.0, "SPEAKER_00")] + assert mt.merge(units, turns) == [ + "00:00:00 Speaker A: First.", + "00:00:06 Speaker B: Second.", + ] + + +class TestMergeError: + def test_merge_transcript_merge_no_turns_raises(self): + """Error: words with no diarization cannot be labelled.""" + with pytest.raises(ValueError, match="no speaker turns"): + mt.merge(words(0.0, "Hello."), []) + + def test_merge_transcript_merge_no_words_raises(self): + """Error: an empty transcription is a failure, not an empty transcript.""" + with pytest.raises(ValueError, match="no speech"): + mt.merge([Unit(0.0, 0.0, " ")], [Turn(0.0, 5.0, "SPEAKER_00")]) + + def test_merge_transcript_merge_negative_gap_setting_raises(self): + """Error: max_gap_s must not be negative.""" + with pytest.raises(ValueError, match="max_gap_s"): + mt.merge(words(0.0, "Hi."), [Turn(0.0, 1.0, "SPEAKER_00")], max_gap_s=-1) + + +class TestLoaders: + def test_merge_transcript_load_whisper_json_reads_millisecond_offsets(self, tmp_path): + """Normal: whisper-cli -oj stores offsets in milliseconds.""" + path = tmp_path / "w.json" + path.write_text(json.dumps({"transcription": [ + {"offsets": {"from": 0, "to": 190}, "text": " if"}, + {"offsets": {"from": 190, "to": 590}, "text": " there's"}, + ]})) + assert mt.load_whisper_json(path) == [Unit(0.0, 0.19, " if"), Unit(0.19, 0.59, " there's")] + + def test_merge_transcript_load_turns_json_reads_seconds(self, tmp_path): + """Normal: the diarizer writes start/end in seconds.""" + path = tmp_path / "t.json" + path.write_text(json.dumps([{"start": 0.5, "end": 4.25, "speaker": "SPEAKER_00"}])) + assert mt.load_turns_json(path) == [Turn(0.5, 4.25, "SPEAKER_00")] + + @pytest.mark.parametrize("content", ["not json", "{}", '{"transcription": [{"text": "x"}]}']) + def test_merge_transcript_load_whisper_json_malformed_raises(self, tmp_path, content): + """Error: a truncated or foreign file is rejected with its path named.""" + path = tmp_path / "w.json" + path.write_text(content) + with pytest.raises(ValueError, match="w.json"): + mt.load_whisper_json(path) + + @pytest.mark.parametrize("content", ["not json", "{}", '[{"start": 1}]']) + def test_merge_transcript_load_turns_json_malformed_raises(self, tmp_path, content): + """Error: same for the turns file.""" + path = tmp_path / "t.json" + path.write_text(content) + with pytest.raises(ValueError, match="t.json"): + mt.load_turns_json(path) + + +class TestLoadersMissingFile: + def test_merge_transcript_load_whisper_json_missing_file_raises_value_error(self, tmp_path): + """Error: a missing whisper file is a clean ValueError naming the path.""" + with pytest.raises(ValueError, match="gone.json"): + mt.load_whisper_json(tmp_path / "gone.json") + + def test_merge_transcript_load_turns_json_missing_file_raises_value_error(self, tmp_path): + """Error: a missing turns file (the diarizer failed upstream) is a clean ValueError.""" + with pytest.raises(ValueError, match="gone.json"): + mt.load_turns_json(tmp_path / "gone.json") + + def test_merge_transcript_main_missing_turns_file_exits_one_without_traceback(self, tmp_path, capsys): + """Error: the CLI reports it in one line and prints nothing on stdout.""" + w = tmp_path / "w.json" + w.write_text('{"transcription": [{"offsets": {"from": 0, "to": 500}, "text": " Hi."}]}') + assert mt.main([str(w), str(tmp_path / "gone.json")]) == 1 + captured = capsys.readouterr() + assert captured.out == "" + assert "gone.json" in captured.err and "Traceback" not in captured.err + + +class TestCli: + def _files(self, tmp_path, transcription, turns): + w = tmp_path / "w.json" + t = tmp_path / "t.json" + w.write_text(json.dumps({"transcription": transcription})) + t.write_text(json.dumps(turns)) + return str(w), str(t) + + def test_merge_transcript_main_prints_transcript_and_returns_zero(self, tmp_path, capsys): + """Normal: the CLI writes the lines to stdout.""" + w, t = self._files( + tmp_path, + [{"offsets": {"from": 0, "to": 500}, "text": " Hello."}], + [{"start": 0.0, "end": 2.0, "speaker": "SPEAKER_00"}], + ) + assert mt.main([w, t]) == 0 + assert capsys.readouterr().out == "00:00:00 Speaker A: Hello.\n" + + def test_merge_transcript_main_failure_prints_nothing_on_stdout(self, tmp_path, capsys): + """Error: a failed merge exits 1 with the reason on stderr only.""" + w, t = self._files(tmp_path, [], [{"start": 0.0, "end": 2.0, "speaker": "SPEAKER_00"}]) + assert mt.main([w, t]) == 1 + captured = capsys.readouterr() + assert captured.out == "" + assert "no speech" in captured.err + + def test_merge_transcript_main_wrong_argument_count_returns_two(self, capsys): + """Error: usage.""" + assert mt.main([]) == 2 + assert "usage" in capsys.readouterr().err.lower() + + +def tok(start_ms, end_ms, text): + return {"text": text, "offsets": {"from": start_ms, "to": end_ms}} + + +class TestLoadWhisperTokens: + """whisper-cli -ojf keeps normal segments and adds per-token offsets inside each.""" + + def _write(self, tmp_path, segments): + path = tmp_path / "full.json" + path.write_text(json.dumps({"transcription": segments})) + return path + + def test_merge_transcript_load_whisper_json_prefers_tokens_over_segments(self, tmp_path): + """Normal: with tokens present, units are words, not whole segments.""" + path = self._write(tmp_path, [{ + "offsets": {"from": 0, "to": 2000}, "text": " we basically", + "tokens": [tok(0, 0, "[_BEG_]"), tok(10, 400, " we"), tok(500, 1900, " basically"), tok(2000, 2000, "[_TT_100]")], + }]) + assert mt.load_whisper_json(path) == [Unit(0.01, 0.4, " we"), Unit(0.5, 1.9, " basically")] + + def test_merge_transcript_load_whisper_json_joins_subword_and_punctuation_tokens(self, tmp_path): + """Normal: a token with no leading space continues the previous word.""" + path = self._write(tmp_path, [{ + "offsets": {"from": 0, "to": 3000}, "text": " Saturday, yes.", + "tokens": [tok(0, 300, " Sat"), tok(300, 700, "urday"), tok(700, 750, ","), tok(900, 1300, " yes"), tok(1300, 1350, ".")], + }]) + assert mt.load_whisper_json(path) == [Unit(0.0, 0.75, " Saturday,"), Unit(0.9, 1.35, " yes.")] + + def test_merge_transcript_load_whisper_json_segment_without_tokens_falls_back(self, tmp_path): + """Boundary: plain -oj output, or a segment whose token list is empty.""" + path = self._write(tmp_path, [ + {"offsets": {"from": 0, "to": 1000}, "text": " First.", "tokens": []}, + {"offsets": {"from": 1000, "to": 2000}, "text": " Second."}, + ]) + assert mt.load_whisper_json(path) == [Unit(0.0, 1.0, " First."), Unit(1.0, 2.0, " Second.")] + + def test_merge_transcript_load_whisper_json_leading_continuation_token_stands_alone(self, tmp_path): + """Boundary: the very first token has nothing to attach to.""" + path = self._write(tmp_path, [{ + "offsets": {"from": 0, "to": 500}, "text": "ing on", + "tokens": [tok(0, 200, "ing"), tok(200, 500, " on")], + }]) + assert mt.load_whisper_json(path) == [Unit(0.0, 0.2, "ing"), Unit(0.2, 0.5, " on")] + + def test_merge_transcript_load_whisper_json_token_without_offsets_raises(self, tmp_path): + """Error: a malformed token is rejected, naming the file.""" + path = self._write(tmp_path, [{"offsets": {"from": 0, "to": 1}, "text": " x", "tokens": [{"text": " x"}]}]) + with pytest.raises(ValueError, match="full.json"): + mt.load_whisper_json(path) + + +class TestFindRepetition: + def test_merge_transcript_find_repetition_clean_text_returns_none(self): + """Normal: ordinary speech, even with stock phrases scattered about.""" + text = " ".join(f"point {i} and i don't know if that works for us" for i in range(8)) + assert mt.find_repetition(text) is None + + def test_merge_transcript_find_repetition_back_to_back_phrase_is_reported(self): + """Normal: whisper's hallucination loop, the same phrase again and again.""" + text = "okay so " + "We don't know where we're going to be. " * 6 + "anyway moving on" + hit = mt.find_repetition(text) + assert hit is not None + phrase, count = hit + assert count >= 6 + assert "where we're going to be" in phrase.lower() + + def test_merge_transcript_find_repetition_three_repeats_is_tolerated(self): + """Boundary: people do say a thing three times; four in a row is the line.""" + assert mt.find_repetition("go back to this area " * 3) is None + assert mt.find_repetition("go back to this area " * 4) is not None + + def test_merge_transcript_find_repetition_short_fillers_are_not_loops(self): + """Boundary: 'yeah yeah yeah yeah yeah' is speech, not a loop.""" + assert mt.find_repetition("yeah " * 9 + "no no no no no") is None + + def test_merge_transcript_find_repetition_empty_text_returns_none(self): + """Boundary: nothing to scan.""" + assert mt.find_repetition("") is None + + def test_merge_transcript_main_loop_in_transcript_fails(self, tmp_path, capsys): + """Error: a looping transcription exits 1 with nothing on stdout.""" + words_ = ("we don't know where we're going to be " * 5).split() + # 40 words at 3 s each: two minutes of loop, far past the 30 s a short stutter gets + segs = [{"offsets": {"from": i * 3_000, "to": i * 3_000 + 2_500}, "text": f" {w}"} for i, w in enumerate(words_)] + w = tmp_path / "w.json" + w.write_text(json.dumps({"transcription": segs})) + t = tmp_path / "t.json" + t.write_text(json.dumps([{"start": 0.0, "end": 130.0, "speaker": "SPEAKER_00"}])) + assert mt.main([str(w), str(t)]) == 1 + captured = capsys.readouterr() + assert captured.out == "" + assert "repeat" in captured.err.lower() + + +class TestInvertedTokenTimes: + """whisper-cli sometimes clamps a token's start to its segment, leaving end < start.""" + + def test_merge_transcript_load_whisper_json_inverted_token_becomes_zero_length(self, tmp_path): + """Boundary: an end before the start is treated as a zero-length word at the start.""" + path = tmp_path / "inv.json" + path.write_text(json.dumps({"transcription": [{ + "offsets": {"from": 13120, "to": 16160}, "text": " which you", + "tokens": [tok(13120, 9580, " which"), tok(13120, 10140, " you")], + }]})) + assert mt.load_whisper_json(path) == [Unit(13.12, 13.12, " which"), Unit(13.12, 13.12, " you")] + + def test_merge_transcript_merge_inverted_tokens_do_not_split_a_sentence(self, tmp_path): + """Normal: the real case, after a pause longer than max_gap_s.""" + path = tmp_path / "inv.json" + path.write_text(json.dumps({"transcription": [ + {"offsets": {"from": 5360, "to": 8640}, "text": " blue pixels", + "tokens": [tok(5360, 7000, " blue"), tok(7000, 8640, " pixels")]}, + {"offsets": {"from": 13120, "to": 16160}, "text": " which you know", + "tokens": [tok(13120, 9580, " which"), tok(13120, 10140, " you"), tok(13120, 10890, " know")]}, + ]})) + turns = [Turn(0.0, 8.8, "SPEAKER_00"), Turn(13.3, 16.0, "SPEAKER_00")] + assert mt.merge(mt.load_whisper_json(path), turns) == [ + "00:00:05 Speaker A: blue pixels", + "00:00:13 Speaker A: which you know", + ] + + +class TestDropOutsideSpeech: + """Whisper invents words in silence; the diarizer knows where the speech is.""" + + def test_merge_transcript_drop_outside_speech_removes_words_far_from_any_turn(self): + """Normal: a hallucinated run in a silent stretch goes; real words stay.""" + real = words(0.0, "Good", "morning.") + invented = words(60.0, "Thank", "you.", "Thank", "you.", step=5.0) + kept, dropped = mt.drop_outside_speech(real + invented, [Turn(0.0, 3.0, "SPEAKER_00")]) + assert kept == real + assert dropped == 4 + + def test_merge_transcript_drop_outside_speech_keeps_words_between_close_turns(self): + """Normal: a word in a short gap between two turns is speech the diarizer clipped.""" + units = words(0.0, "One.") + words(3.2, "and") + words(4.0, "two.") + turns = [Turn(0.0, 3.0, "SPEAKER_00"), Turn(4.0, 6.0, "SPEAKER_01")] + assert mt.drop_outside_speech(units, turns) == (units, 0) + + def test_merge_transcript_drop_outside_speech_word_exactly_at_the_margin_is_kept(self): + """Boundary: the margin is inclusive.""" + unit = Unit(5.0, 5.5, " edge") + assert mt.drop_outside_speech([unit], [Turn(0.0, 3.0, "S")], margin_s=2.0) == ([unit], 0) + + def test_merge_transcript_drop_outside_speech_word_just_past_the_margin_is_dropped(self): + """Boundary: one millisecond further and it goes.""" + unit = Unit(5.001, 5.5, " edge") + assert mt.drop_outside_speech([unit], [Turn(0.0, 3.0, "S")], margin_s=2.0) == ([], 1) + + def test_merge_transcript_drop_outside_speech_zero_margin_needs_contact_with_a_turn(self): + """Boundary: margin 0 keeps a word touching a turn and drops one that is not.""" + touching, apart = Unit(3.0, 3.4, " touch"), Unit(3.5, 3.9, " apart") + kept, dropped = mt.drop_outside_speech([touching, apart], [Turn(0.0, 3.0, "S")], margin_s=0.0) + assert (kept, dropped) == ([touching], 1) + + def test_merge_transcript_drop_outside_speech_before_the_first_turn_counts_too(self): + """Boundary: silence at the start of a recording.""" + early = Unit(1.0, 1.5, " Thanks.") + assert mt.drop_outside_speech([early], [Turn(30.0, 40.0, "S")]) == ([], 1) + + def test_merge_transcript_drop_outside_speech_no_units_is_a_no_op(self): + """Boundary: nothing in, nothing out.""" + assert mt.drop_outside_speech([], [Turn(0.0, 3.0, "S")]) == ([], 0) + + def test_merge_transcript_drop_outside_speech_no_turns_raises(self): + """Error: without turns there is no way to tell speech from silence.""" + with pytest.raises(ValueError, match="no speaker turns"): + mt.drop_outside_speech(words(0.0, "Hello."), []) + + def test_merge_transcript_drop_outside_speech_negative_margin_raises(self): + """Error: a negative margin is a caller bug.""" + with pytest.raises(ValueError, match="margin"): + mt.drop_outside_speech(words(0.0, "Hello."), [Turn(0.0, 3.0, "S")], margin_s=-1.0) + + def test_merge_transcript_main_silence_hallucinations_do_not_trip_the_loop_guard(self, tmp_path, capsys): + """Normal: a run of invented thank-yous in silence is dropped, not reported as a loop.""" + transcription = [{"offsets": {"from": 0, "to": 900}, "text": " Good morning."}] + [ + {"offsets": {"from": 60_000 + i * 5_000, "to": 64_000 + i * 5_000}, "text": " Thank you."} + for i in range(12) + ] + whisper = tmp_path / "w.json" + whisper.write_text(json.dumps({"transcription": transcription})) + turns = tmp_path / "t.json" + turns.write_text(json.dumps([{"start": 0.0, "end": 2.0, "speaker": "SPEAKER_00"}])) + assert mt.main([str(whisper), str(turns)]) == 0 + assert capsys.readouterr().out == "00:00:00 Speaker A: Good morning.\n" + + def test_merge_transcript_main_a_loop_inside_speech_still_fails(self, tmp_path, capsys): + """Error: a real repetition loop happens while someone is talking, and is still caught.""" + transcription = [ + {"offsets": {"from": i * 1_000, "to": i * 1_000 + 900}, "text": " where we're going to be"} + for i in range(60) # a full minute of the same phrase + ] + whisper = tmp_path / "w.json" + whisper.write_text(json.dumps({"transcription": transcription})) + turns = tmp_path / "t.json" + turns.write_text(json.dumps([{"start": 0.0, "end": 70.0, "speaker": "SPEAKER_00"}])) + assert mt.main([str(whisper), str(turns)]) == 1 + assert "looped" in capsys.readouterr().err + + +class TestCollapseRepetitions: + """A short stutter is collapsed to one occurrence; a long loop still fails the job.""" + + def test_merge_transcript_collapse_repetitions_short_loop_keeps_one_copy(self): + """Normal: a six-word phrase said four times in ten seconds becomes one phrase.""" + units = words(0.0, "So", "anyway,") + words(1.0, *("fair, it's not going to be".split() * 4), step=0.4) + words(12.0, "done.") + kept, collapsed = mt.collapse_repetitions(units) + assert " ".join(u.text.strip() for u in kept) == "So anyway, fair, it's not going to be done." + assert collapsed == [("fair, it's not going to be", 4)] + + def test_merge_transcript_collapse_repetitions_clean_units_are_untouched(self): + """Normal: ordinary speech passes through with nothing collapsed.""" + units = words(0.0, "We", "have", "detection", "today,", "and", "we", "have", "a", "plan.") + assert mt.collapse_repetitions(units) == (units, []) + + def test_merge_transcript_collapse_repetitions_two_loops_both_collapse(self): + """Normal: separate stutters are each collapsed and each reported.""" + units = (words(0.0, *("go back to this area".split() * 4), step=0.3) + + words(10.0, "then") + + words(11.0, *("where we're going to be".split() * 5), step=0.3)) + kept, collapsed = mt.collapse_repetitions(units) + assert " ".join(u.text.strip() for u in kept) == "go back to this area then where we're going to be" + assert collapsed == [("go back to this area", 4), ("where we're going to be", 5)] + + def test_merge_transcript_collapse_repetitions_three_repeats_are_speech(self): + """Boundary: three repeats is emphasis, not a loop, and stays.""" + units = words(0.0, *("this is the thing".split() * 3)) + assert mt.collapse_repetitions(units) == (units, []) + + def test_merge_transcript_collapse_repetitions_single_word_runs_stay(self): + """Boundary: "yeah yeah yeah yeah" is ordinary speech.""" + units = words(0.0, *(["yeah"] * 8)) + assert mt.collapse_repetitions(units) == (units, []) + + def test_merge_transcript_collapse_repetitions_loop_at_the_limit_is_collapsed(self): + """Boundary: a loop lasting exactly max_loop_s is still a short one.""" + units = words(0.0, *("we do not know where".split() * 4), step=1.5) # 20 words, 30.0 s + kept, collapsed = mt.collapse_repetitions(units, max_loop_s=30.0) + assert collapsed == [("we do not know where", 4)] and len(kept) == 5 + + def test_merge_transcript_collapse_repetitions_loop_past_the_limit_raises(self): + """Error: a loop longer than max_loop_s means real speech was lost, so the job fails.""" + units = words(0.0, *("we do not know where".split() * 4), step=1.6) # 32.0 s + with pytest.raises(ValueError, match="looped"): + mt.collapse_repetitions(units, max_loop_s=30.0) + + def test_merge_transcript_collapse_repetitions_segment_units_collapse_too(self): + """Boundary: whisper's segment fallback puts a whole phrase in one unit.""" + units = [Unit(i * 1.0, i * 1.0 + 0.9, " where we're going to be") for i in range(6)] + kept, collapsed = mt.collapse_repetitions(units) + assert len(kept) == 1 and collapsed == [("where we're going to be", 6)] + + def test_merge_transcript_collapse_repetitions_negative_limit_raises(self): + """Error: a negative limit is a caller bug.""" + with pytest.raises(ValueError, match="max_loop_s"): + mt.collapse_repetitions(words(0.0, "hi"), max_loop_s=-1.0) + + def test_merge_transcript_main_short_loop_is_collapsed_not_fatal(self, tmp_path, capsys): + """Normal: the CLI prints the collapsed transcript and exits 0.""" + words_ = "we have detection today " .split() + ("fair, it's not going to be " * 4).split() + "easy.".split() + segs = [{"offsets": {"from": i * 300, "to": i * 300 + 250}, "text": f" {w}"} for i, w in enumerate(words_)] + w = tmp_path / "w.json"; w.write_text(json.dumps({"transcription": segs})) + t = tmp_path / "t.json"; t.write_text(json.dumps([{"start": 0.0, "end": 60.0, "speaker": "SPEAKER_00"}])) + assert mt.main([str(w), str(t)]) == 0 + assert capsys.readouterr().out == "00:00:00 Speaker A: we have detection today fair, it's not going to be easy.\n" + + +class TestCollapseRepetitionsInsideUnits: + """The repeat can live inside one unit's text, which is what whisper's segment fallback emits.""" + + def test_merge_transcript_collapse_repetitions_loop_inside_one_unit_is_collapsed(self): + """Error case turned regression: a single unit holding the phrase four times must not hang.""" + unit = Unit(0.0, 8.0, " " + " ".join(["we do not know where"] * 4)) + kept, collapsed = mt.collapse_repetitions(unit and [unit]) + assert [u.text for u in kept] == [" we do not know where"] + assert collapsed == [("we do not know where", 4)] + + def test_merge_transcript_collapse_repetitions_partial_unit_keeps_its_other_words(self): + """Boundary: a unit holding the last repeat and real words after it keeps the real words.""" + units = words(0.0, *("go back to this area".split() * 3), step=0.5) + [ + Unit(7.5, 9.0, " go back to this area and then we stopped.") + ] + kept, collapsed = mt.collapse_repetitions(units) + assert " ".join(u.text.strip() for u in kept) == "go back to this area and then we stopped." + assert collapsed == [("go back to this area", 4)] diff --git a/working/meeting-transcription-service/tests/test_ratio_transcribe.py b/working/meeting-transcription-service/tests/test_ratio_transcribe.py new file mode 100644 index 0000000..1a10b3e --- /dev/null +++ b/working/meeting-transcription-service/tests/test_ratio_transcribe.py @@ -0,0 +1,263 @@ +"""Tests for ratio-transcribe, the client. + +ssh and scp are replaced by fakes that act on a temp directory standing in for the +remote home, so the client's real logic (job ids, upload-then-rename, polling, +collecting, the local fallback) runs against a filesystem it can't tell from the host. +""" + +import json +import os +import subprocess +from pathlib import Path + +import pytest + +SCRIPT = Path(__file__).resolve().parent.parent / "src" / "ratio-transcribe" +STATE = ".local/state/meeting-transcribe" + +FAKE_SSH = r"""#!/usr/bin/env bash +# ssh [-o k=v]... host command... -> run the command with HOME at the fake remote +[[ -n "${FAKE_SSH_DOWN:-}" ]] && exit 255 +no_stdin="" +while [[ "$1" == -* ]]; do [[ "$1" == "-n" ]] && no_stdin=1; [[ "$1" == "-o" ]] && shift; shift; done +shift # host +printf '%s\n' "$*" >> "$FAKE_SSH_LOG" +if [[ -n "$no_stdin" ]]; then + ( cd "$FAKE_REMOTE_HOME" && HOME="$FAKE_REMOTE_HOME" bash -c "$*" < /dev/null ) + rc=$? +else + ( cd "$FAKE_REMOTE_HOME" && HOME="$FAKE_REMOTE_HOME" bash -c "$*" ) + rc=$? + cat > /dev/null # like the real ssh, drain whatever stdin the command left behind +fi +if [[ "$*" == *"incoming/"* && "$*" == mv* && -n "${FAKE_ON_SUBMIT:-}" ]]; then + ( cd "$FAKE_REMOTE_HOME" && bash -c "$FAKE_ON_SUBMIT" ) +fi +exit $rc +""" + +FAKE_SCP = r"""#!/usr/bin/env bash +[[ -n "${FAKE_SSH_DOWN:-}" ]] && exit 255 +while [[ "$1" == -* ]]; do [[ "$1" == "-o" ]] && shift; shift; done +printf '%s -> %s\n' "$1" "$2" >> "$FAKE_SCP_LOG" +cp "$1" "$FAKE_REMOTE_HOME/${2#*:}" +""" + +# What the worker would do, compressed: finish or fail whatever sits in incoming/. +WORKER_OK = f"""cd {STATE}; mkdir -p done; for d in incoming/*/; do id=$(basename "$d"); + printf '00:00:00 Speaker A: Hello from the host.\\n' > done/$id.txt; rm -rf "$d"; done""" +WORKER_FAIL = f"""cd {STATE}; mkdir -p failed; for d in incoming/*/; do id=$(basename "$d"); + printf 'stage: whisper\\nexit 3\\n' > failed/$id.log; rm -rf "$d"; done""" +WORKER_EMPTY = f"""cd {STATE}; mkdir -p done; for d in incoming/*/; do id=$(basename "$d"); + : > done/$id.txt; rm -rf "$d"; done""" + + +class Rig: + def __init__(self, tmp_path): + self.tmp = tmp_path + self.remote = tmp_path / "remote-home" + self.remote.mkdir() + self.local_home = tmp_path / "local-home" + self.local_home.mkdir() + self.bin = tmp_path / "bin" + self.bin.mkdir() + for name, body in (("ssh", FAKE_SSH), ("scp", FAKE_SCP)): + path = self.bin / name + path.write_text(body) + path.chmod(0o755) + self.ssh_log = tmp_path / "ssh.log" + self.scp_log = tmp_path / "scp.log" + self.audio = tmp_path / "meeting.m4a" + self.audio.write_bytes(b"pretend audio") + + def run(self, args=None, on_submit=WORKER_OK, **env_extra): + env = { + "PATH": f"{self.bin}:{os.environ['PATH']}", + "HOME": str(self.local_home), + "FAKE_REMOTE_HOME": str(self.remote), + "FAKE_SSH_LOG": str(self.ssh_log), + "FAKE_SCP_LOG": str(self.scp_log), + "TRANSCRIBE_HOST": "testhost", + "TRANSCRIBE_POLL": "0", + "TRANSCRIBE_TIMEOUT": "5", + } + if on_submit: + env["FAKE_ON_SUBMIT"] = on_submit + env.update({k: str(v) for k, v in env_extra.items()}) + if args is None: + args = [str(self.audio)] + return subprocess.run([str(SCRIPT), *args], env=env, capture_output=True, text=True, timeout=60) + + def state(self, *parts): + return self.remote.joinpath(STATE, *parts) + + def uploads(self): + return self.scp_log.read_text().splitlines() if self.scp_log.exists() else [] + + +@pytest.fixture +def rig(tmp_path): + return Rig(tmp_path) + + +class TestClientNormal: + def test_ratio_transcribe_new_recording_is_uploaded_and_transcript_printed(self, rig): + """Normal: upload, wait, print. stdout is the transcript and nothing else.""" + result = rig.run() + assert result.returncode == 0, result.stderr + assert result.stdout == "00:00:00 Speaker A: Hello from the host.\n" + assert len(rig.uploads()) == 1 + + def test_ratio_transcribe_job_file_carries_language_speakers_and_name(self, rig): + """Normal: the options reach the host inside job.json.""" + rig.run(args=[str(rig.audio), "es"], on_submit=None, SPEAKERS=3, TRANSCRIBE_TIMEOUT=0) + jobs = list(rig.state("incoming").glob("*/job.json")) + assert len(jobs) == 1 + assert json.loads(jobs[0].read_text()) == { + "language": "es", "speakers": 3, "original_name": "meeting.m4a", + } + assert (jobs[0].parent / "audio.m4a").read_bytes() == b"pretend audio" + + def test_ratio_transcribe_finished_job_is_collected_without_uploading_again(self, rig): + """Normal: rerunning after a dropped connection costs one ssh round trip.""" + first = rig.run() + rig.scp_log.unlink() + second = rig.run(on_submit=None) + assert second.returncode == 0 + assert second.stdout == first.stdout + assert rig.uploads() == [] + + def test_ratio_transcribe_job_id_depends_on_audio_and_options(self, rig): + """Normal: same file with a different speaker count is a different job.""" + rig.run(on_submit=None, TRANSCRIBE_TIMEOUT=0) + rig.run(on_submit=None, TRANSCRIBE_TIMEOUT=0) + rig.run(on_submit=None, TRANSCRIBE_TIMEOUT=0, SPEAKERS=3) + assert len(list(rig.state("incoming").iterdir())) == 2 + + def test_ratio_transcribe_earlier_failure_is_retried(self, rig): + """Normal: a stale failure log does not block a fresh attempt.""" + assert rig.run(on_submit=WORKER_FAIL).returncode == 1 + result = rig.run(on_submit=WORKER_OK) + assert result.returncode == 0 + assert "Hello from the host" in result.stdout + assert list(rig.state("failed").glob("*.log")) == [] + + +class TestClientBoundary: + def test_ratio_transcribe_job_still_queued_is_not_uploaded_twice(self, rig): + """Boundary: a second run while the first job waits just joins the wait.""" + rig.run(on_submit=None, TRANSCRIBE_TIMEOUT=0) + rig.run(on_submit=None, TRANSCRIBE_TIMEOUT=0) + assert len(rig.uploads()) == 1 + + def test_ratio_transcribe_upload_lands_in_incoming_only_by_rename(self, rig): + """Boundary: the audio is copied into uploading/, never straight into incoming/.""" + rig.run(on_submit=None, TRANSCRIBE_TIMEOUT=0) + assert "/uploading/" in rig.uploads()[0] + assert list(rig.state("uploading").iterdir()) == [] + + def test_ratio_transcribe_awkward_filename_survives(self, rig): + """Boundary: spaces and quotes in the recording's name.""" + odd = rig.tmp / "someone's \"weekly\" sync.m4a" + odd.write_bytes(b"x") + result = rig.run(args=[str(odd)], on_submit=None, TRANSCRIBE_TIMEOUT=0) + job = json.loads(next(rig.state("incoming").glob("*/job.json")).read_text()) + assert job["original_name"] == "someone's \"weekly\" sync.m4a", result.stderr + + def test_ratio_transcribe_colon_in_a_relative_filename_still_uploads(self, rig): + """Boundary: scp reads "standup 14:30.m4a" as host:path; the client hands it an absolute path.""" + (rig.tmp / "standup 14:30.m4a").write_bytes(b"x") + env = { + "PATH": f"{rig.bin}:{os.environ['PATH']}", "HOME": str(rig.local_home), + "FAKE_REMOTE_HOME": str(rig.remote), "FAKE_SSH_LOG": str(rig.ssh_log), + "FAKE_SCP_LOG": str(rig.scp_log), "FAKE_ON_SUBMIT": WORKER_OK, + "TRANSCRIBE_HOST": "testhost", "TRANSCRIBE_POLL": "0", "TRANSCRIBE_TIMEOUT": "5", + } + result = subprocess.run( + [str(SCRIPT), "standup 14:30.m4a"], cwd=rig.tmp, env=env, capture_output=True, text=True, timeout=60, + ) + assert result.returncode == 0, result.stderr + assert rig.uploads()[0].startswith("/") + + def test_ratio_transcribe_host_unreachable_falls_back_to_local_worker(self, rig): + """Boundary: offline, the same queue and worker run on this machine.""" + worker = rig.tmp / "fake-worker" + worker.write_text( + "#!/usr/bin/env bash\n" + f"cd \"$HOME/{STATE}\" && mkdir -p done && for d in incoming/*/; do id=$(basename \"$d\");\n" + "printf '00:00:00 Speaker A: Local fallback.\\n' > done/$id.txt; rm -rf \"$d\"; done\n" + ) + worker.chmod(0o755) + result = rig.run(FAKE_SSH_DOWN=1, TRANSCRIBE_WORKER=str(worker)) + assert result.returncode == 0, result.stderr + assert result.stdout == "00:00:00 Speaker A: Local fallback.\n" + assert "local" in result.stderr.lower() + + +class TestClientStdin: + def test_ratio_transcribe_leaves_the_callers_stdin_alone(self, rig): + """Boundary: in a `while read` loop the client must not eat the loop's input.""" + env = { + "PATH": f"{rig.bin}:{os.environ['PATH']}", "HOME": str(rig.local_home), + "FAKE_REMOTE_HOME": str(rig.remote), "FAKE_SSH_LOG": str(rig.ssh_log), + "FAKE_SCP_LOG": str(rig.scp_log), "FAKE_ON_SUBMIT": WORKER_OK, + "TRANSCRIBE_HOST": "testhost", "TRANSCRIBE_POLL": "0", "TRANSCRIBE_TIMEOUT": "5", + } + result = subprocess.run( + ["bash", "-c", '"$0" "$1" > /dev/null 2>&1; cat', str(SCRIPT), str(rig.audio)], + env=env, input="next line of the caller's loop\n", capture_output=True, text=True, timeout=60, + ) + assert result.stdout == "next line of the caller's loop\n" + + +class TestClientError: + def test_ratio_transcribe_no_arguments_prints_usage(self, rig): + """Error: usage.""" + result = rig.run(args=[]) + assert result.returncode == 1 and "Usage: ratio-transcribe" in result.stderr and result.stdout == "" + + def test_ratio_transcribe_missing_file_fails_before_any_ssh(self, rig): + """Error: no such recording.""" + result = rig.run(args=[str(rig.tmp / "nope.m4a")]) + assert result.returncode == 1 and "not found" in result.stderr + assert not rig.ssh_log.exists() + + @pytest.mark.parametrize("env", [{"SPEAKERS": "0"}, {"SPEAKERS": "three"}, {"MIN_SPEAKERS": "-1"}, + {"MIN_SPEAKERS": "4", "MAX_SPEAKERS": "2"}, {"SPEAKERS": "3", "MAX_SPEAKERS": "5"}]) + def test_ratio_transcribe_bad_speaker_settings_are_rejected(self, rig, env): + """Error: non-positive, non-numeric or contradictory counts.""" + result = rig.run(**env) + assert result.returncode == 1 and "speaker" in result.stderr.lower() + assert not rig.ssh_log.exists() + + def test_ratio_transcribe_bad_language_is_rejected(self, rig): + """Error: the language ends up in a job file and a command line.""" + result = rig.run(args=[str(rig.audio), "en; rm -rf /"]) + assert result.returncode == 1 and "language" in result.stderr.lower() + assert not rig.ssh_log.exists() + + def test_ratio_transcribe_failed_job_reports_the_log_and_prints_nothing(self, rig): + """Error: the host's failure log comes back on stderr.""" + result = rig.run(on_submit=WORKER_FAIL) + assert result.returncode == 1 + assert result.stdout == "" + assert "stage: whisper" in result.stderr + + def test_ratio_transcribe_timeout_says_the_job_is_still_running(self, rig): + """Error: giving up waiting is not the job failing.""" + result = rig.run(on_submit=None, TRANSCRIBE_TIMEOUT=0) + assert result.returncode == 1 + assert result.stdout == "" + assert "again" in result.stderr.lower() + + def test_ratio_transcribe_empty_transcript_is_a_failure(self, rig): + """Error: a zero-byte result is never passed off as a transcript.""" + result = rig.run(on_submit=WORKER_EMPTY, TRANSCRIBE_TIMEOUT=1) + assert result.returncode == 1 + assert result.stdout == "" + + def test_ratio_transcribe_offline_without_local_worker_names_what_is_missing(self, rig): + """Error: no host and nothing installed locally.""" + result = rig.run(FAKE_SSH_DOWN=1, TRANSCRIBE_WORKER=str(rig.tmp / "absent")) + assert result.returncode == 1 + assert "absent" in result.stderr diff --git a/working/meeting-transcription-service/tests/test_transcribe_worker.py b/working/meeting-transcription-service/tests/test_transcribe_worker.py new file mode 100644 index 0000000..bfffd31 --- /dev/null +++ b/working/meeting-transcription-service/tests/test_transcribe_worker.py @@ -0,0 +1,359 @@ +"""Tests for transcribe-worker: the queue drain on the transcription host. + +ffmpeg, whisper-cli and the diarizer are replaced by small fake executables (the +process boundary). The queue handling, the merge and the loop guard run for real. +""" + +import importlib.machinery +import importlib.util +import json +import os +import stat +import sys +from pathlib import Path + +import pytest + +SRC = Path(__file__).resolve().parent.parent / "src" +sys.path.insert(0, str(SRC)) + + +def _load_worker(): + loader = importlib.machinery.SourceFileLoader("transcribe_worker", str(SRC / "transcribe-worker")) + spec = importlib.util.spec_from_loader("transcribe_worker", loader) + assert spec is not None + module = importlib.util.module_from_spec(spec) + sys.modules["transcribe_worker"] = module # dataclasses looks the module up by name + loader.exec_module(module) + return module + + +worker = _load_worker() + +FAKE_FFMPEG = """#!/usr/bin/env bash +# last argument is the output; the one after -i is the input +while [[ $# -gt 1 ]]; do [[ "$1" == "-i" ]] && in="$2"; shift; done +cp "$in" "$1" +exit "${FAKE_FFMPEG_EXIT:-0}" +""" + +FAKE_WHISPER = """#!/usr/bin/env bash +printf '%s\\n' "$*" >> "$FAKE_LOG" +while [[ $# -gt 0 ]]; do [[ "$1" == "-of" ]] && prefix="$2"; shift; done +[[ "${FAKE_WHISPER_EXIT:-0}" == "0" ]] && cp "$FAKE_WHISPER_JSON" "$prefix.json" +exit "${FAKE_WHISPER_EXIT:-0}" +""" + +FAKE_DIARIZE = """#!/usr/bin/env bash +printf 'diarize %s\\n' "$*" >> "$FAKE_LOG" +[[ "${FAKE_DIARIZE_EXIT:-0}" == "0" ]] && cp "$FAKE_TURNS_JSON" "$2" +echo "fake diarizer says hello" >&2 +exit "${FAKE_DIARIZE_EXIT:-0}" +""" + + +def whisper_json(*words, step_ms=400): + return {"transcription": [ + {"offsets": {"from": i * step_ms, "to": i * step_ms + 300}, "text": f" {w}"} + for i, w in enumerate(words) + ]} + + +class Rig: + def __init__(self, tmp_path, monkeypatch): + self.home = tmp_path / "state" + self.bin = tmp_path / "bin" + self.bin.mkdir() + for name, body in (("ffmpeg", FAKE_FFMPEG), ("whisper-cli", FAKE_WHISPER), ("fake-diarize", FAKE_DIARIZE)): + path = self.bin / name + path.write_text(body) + path.chmod(path.stat().st_mode | stat.S_IXUSR) + self.log = tmp_path / "calls.log" + self.whisper_json = tmp_path / "whisper.json" + self.turns_json = tmp_path / "turns.json" + self.set_whisper(whisper_json("Good", "morning.")) + self.turns_json.write_text(json.dumps([{"start": 0.0, "end": 60.0, "speaker": "SPEAKER_00"}])) + monkeypatch.setenv("PATH", f"{self.bin}:{os.environ['PATH']}") + monkeypatch.setenv("FAKE_LOG", str(self.log)) + monkeypatch.setenv("FAKE_WHISPER_JSON", str(self.whisper_json)) + monkeypatch.setenv("FAKE_TURNS_JSON", str(self.turns_json)) + self.monkeypatch = monkeypatch + self.config = worker.Config( + home=self.home, + whisper_model=tmp_path / "model.bin", + diarize_cmd=[str(self.bin / "fake-diarize")], + threads=2, + ) + + def set_whisper(self, data): + self.whisper_json.write_text(json.dumps(data)) + + def submit(self, job_id, job=None, audio=b"audio", where="incoming"): + folder = self.home / where / job_id + folder.mkdir(parents=True) + (folder / "audio.m4a").write_bytes(audio) + if job is not False: + (folder / "job.json").write_text(json.dumps(job if job is not None else {"language": "en"})) + return folder + + def calls(self): + return self.log.read_text().splitlines() if self.log.exists() else [] + + +@pytest.fixture +def rig(tmp_path, monkeypatch): + return Rig(tmp_path, monkeypatch) + + +class TestDrainNormal: + def test_transcribe_worker_drain_one_job_writes_transcript_to_done(self, rig): + """Normal: a job in incoming/ ends as done/<id>.txt and leaves nothing behind.""" + rig.submit("job1") + assert worker.drain(rig.config) == {"done": 1, "failed": 0} + assert (rig.home / "done" / "job1.txt").read_text() == "00:00:00 Speaker A: Good morning.\n" + assert not (rig.home / "incoming" / "job1").exists() + assert not (rig.home / "work" / "job1").exists() + assert not (rig.home / "failed" / "job1.log").exists() + + def test_transcribe_worker_drain_writes_metadata_beside_the_transcript(self, rig): + """Normal: done/<id>.json records what ran, for the client and for debugging.""" + rig.submit("job1", {"language": "en", "speakers": 3, "original_name": "standup.mkv"}) + worker.drain(rig.config) + meta = json.loads((rig.home / "done" / "job1.json").read_text()) + assert meta["original_name"] == "standup.mkv" + assert meta["speakers_found"] == 1 + assert meta["lines"] == 1 + assert meta["seconds"] >= 0 + + def test_transcribe_worker_drain_passes_language_and_speaker_count_through(self, rig): + """Normal: job options reach whisper and the diarizer.""" + rig.submit("job1", {"language": "es", "speakers": 3}) + worker.drain(rig.config) + whisper_call = next(c for c in rig.calls() if not c.startswith("diarize")) + diarize_call = next(c for c in rig.calls() if c.startswith("diarize")) + assert "-l es" in whisper_call + assert "-mc 0" in whisper_call and "-ojf" in whisper_call + assert diarize_call.endswith("--speakers 3") + + def test_transcribe_worker_drain_min_and_max_speakers_are_forwarded(self, rig): + """Normal: a range instead of an exact count.""" + rig.submit("job1", {"min_speakers": 2, "max_speakers": 5}) + worker.drain(rig.config) + diarize_call = next(c for c in rig.calls() if c.startswith("diarize")) + assert "--min-speakers 2" in diarize_call and "--max-speakers 5" in diarize_call + + def test_transcribe_worker_drain_processes_every_job_oldest_first(self, rig): + """Normal: the queue drains completely, in arrival order.""" + first = rig.submit("b-first") + rig.submit("a-second") + os.utime(first, (1, 1)) + assert worker.drain(rig.config) == {"done": 2, "failed": 0} + diarize_calls = [c for c in rig.calls() if c.startswith("diarize")] + assert "b-first" in diarize_calls[0] and "a-second" in diarize_calls[1] + + +class TestDrainBoundary: + def test_transcribe_worker_drain_empty_queue_is_a_no_op(self, rig): + """Boundary: nothing to do, and the folders get created.""" + assert worker.drain(rig.config) == {"done": 0, "failed": 0} + assert (rig.home / "incoming").is_dir() and (rig.home / "done").is_dir() + + def test_transcribe_worker_drain_ignores_uploads_still_in_flight(self, rig): + """Boundary: a dot-prefixed folder is left alone (second line of defence behind uploading/).""" + rig.submit(".tmp-job9") + assert worker.drain(rig.config) == {"done": 0, "failed": 0} + assert (rig.home / "incoming" / ".tmp-job9").exists() + + def test_transcribe_worker_drain_job_already_done_is_not_rerun(self, rig): + """Boundary: resubmitting finished work costs nothing.""" + (rig.home / "done").mkdir(parents=True) + (rig.home / "done" / "job1.txt").write_text("already here\n") + rig.submit("job1") + assert worker.drain(rig.config) == {"done": 0, "failed": 0} + assert (rig.home / "done" / "job1.txt").read_text() == "already here\n" + assert not (rig.home / "incoming" / "job1").exists() + assert rig.calls() == [] + + def test_transcribe_worker_drain_missing_job_file_uses_defaults(self, rig): + """Boundary: audio with no job.json still transcribes, in English, count estimated.""" + rig.submit("job1", job=False) + assert worker.drain(rig.config) == {"done": 1, "failed": 0} + + def test_transcribe_worker_drain_retries_a_job_left_in_work_by_a_crash(self, rig): + """Boundary: a job stranded in work/ is picked up again.""" + rig.submit("job1", where="work") + assert worker.drain(rig.config) == {"done": 1, "failed": 0} + + def test_transcribe_worker_drain_gives_up_on_a_job_that_keeps_crashing(self, rig): + """Boundary: the second stranding is a failure, not a loop.""" + rig.submit("job1", {"language": "en", "attempts": 2}, where="work") + assert worker.drain(rig.config) == {"done": 0, "failed": 1} + assert "attempt" in (rig.home / "failed" / "job1.log").read_text() + + +class TestDrainError: + def test_transcribe_worker_drain_whisper_failure_is_logged_with_its_stage(self, rig): + """Error: whisper exits non-zero.""" + rig.monkeypatch.setenv("FAKE_WHISPER_EXIT", "3") + rig.submit("job1") + assert worker.drain(rig.config) == {"done": 0, "failed": 1} + log = (rig.home / "failed" / "job1.log").read_text() + assert "whisper" in log + assert not (rig.home / "done" / "job1.txt").exists() + assert not (rig.home / "work" / "job1").exists() + + def test_transcribe_worker_drain_diarizer_failure_keeps_its_stderr(self, rig): + """Error: the diarizer fails; its own words land in the log.""" + rig.monkeypatch.setenv("FAKE_DIARIZE_EXIT", "1") + rig.submit("job1") + assert worker.drain(rig.config) == {"done": 0, "failed": 1} + log = (rig.home / "failed" / "job1.log").read_text() + assert "diarize" in log and "fake diarizer says hello" in log + + def test_transcribe_worker_drain_looping_transcription_fails(self, rig): + """Error: whisper's repetition loop is a failed job, never a transcript.""" + rig.set_whisper(whisper_json(*("we do not know where we are going".split() * 5), step_ms=1000)) # 40 s + rig.submit("job1") + assert worker.drain(rig.config) == {"done": 0, "failed": 1} + assert "looped" in (rig.home / "failed" / "job1.log").read_text() + assert not (rig.home / "done" / "job1.txt").exists() + + def test_transcribe_worker_drain_short_loop_is_collapsed_and_recorded(self, rig): + """Normal: a stutter of a few seconds is collapsed to one copy and noted in the metadata.""" + rig.set_whisper(whisper_json("Okay,", *("fair, it's not going to be".split() * 4), "easy.")) + rig.submit("job1") + assert worker.drain(rig.config) == {"done": 1, "failed": 0} + assert (rig.home / "done" / "job1.txt").read_text() == "00:00:00 Speaker A: Okay, fair, it's not going to be easy.\n" + meta = json.loads((rig.home / "done" / "job1.json").read_text()) + assert meta["loops_collapsed"] == [["fair, it's not going to be", 4]] + + def test_transcribe_worker_drain_malformed_job_file_fails_that_job_only(self, rig): + """Error: one bad job does not stop the queue.""" + bad = rig.submit("bad") + (bad / "job.json").write_text("{not json") + os.utime(bad, (1, 1)) + rig.submit("good") + assert worker.drain(rig.config) == {"done": 1, "failed": 1} + assert (rig.home / "done" / "good.txt").exists() + assert "job.json" in (rig.home / "failed" / "bad.log").read_text() + + def test_transcribe_worker_drain_job_without_audio_fails(self, rig): + """Error: a job folder holding no audio file.""" + folder = rig.submit("job1") + (folder / "audio.m4a").unlink() + assert worker.drain(rig.config) == {"done": 0, "failed": 1} + assert "audio" in (rig.home / "failed" / "job1.log").read_text() + + @pytest.mark.parametrize("bad", [{"speakers": 0}, {"speakers": "three"}, {"language": "en; rm -rf"}]) + def test_transcribe_worker_drain_rejects_bad_option_values(self, rig, bad): + """Error: options are validated before they reach a command line.""" + rig.submit("job1", bad) + assert worker.drain(rig.config) == {"done": 0, "failed": 1} + assert rig.calls() == [] + + +class TestDrainSilenceHallucinations: + def _silent_stretch(self, rig): + spoken = [{"offsets": {"from": 0, "to": 900}, "text": " Good morning."}] + invented = [ + {"offsets": {"from": 60_000 + i * 5_000, "to": 64_000 + i * 5_000}, "text": " Thank you."} + for i in range(12) + ] + rig.set_whisper({"transcription": spoken + invented}) + rig.turns_json.write_text(json.dumps([{"start": 0.0, "end": 2.0, "speaker": "SPEAKER_00"}])) + + def test_transcribe_worker_drain_words_invented_in_silence_are_dropped_not_failed(self, rig): + """Normal: a quiet meeting transcribes; the invented run never reaches the transcript.""" + self._silent_stretch(rig) + rig.submit("job1") + assert worker.drain(rig.config) == {"done": 1, "failed": 0} + assert (rig.home / "done" / "job1.txt").read_text() == "00:00:00 Speaker A: Good morning.\n" + + def test_transcribe_worker_drain_metadata_counts_the_dropped_words(self, rig): + """Normal: the count is on record, so a transcript that lost a lot is visible.""" + self._silent_stretch(rig) + rig.submit("job1") + worker.drain(rig.config) + assert json.loads((rig.home / "done" / "job1.json").read_text())["dropped_outside_speech"] == 12 + + def test_transcribe_worker_drain_recording_with_no_speech_at_all_fails_cleanly(self, rig): + """Error: everything whisper produced sits in silence.""" + rig.set_whisper({"transcription": [{"offsets": {"from": 60_000, "to": 61_000}, "text": " Thank you."}]}) + rig.turns_json.write_text(json.dumps([{"start": 0.0, "end": 2.0, "speaker": "SPEAKER_00"}])) + rig.submit("job1") + assert worker.drain(rig.config) == {"done": 0, "failed": 1} + assert "no speech" in (rig.home / "failed" / "job1.log").read_text() + + +class TestRunStdin: + def test_transcribe_worker_run_gives_tools_no_stdin(self, tmp_path): + """Boundary: ffmpeg reads stdin when it can, which would eat a calling loop's input.""" + out = tmp_path / "stdin-target" + read_end, write_end = os.pipe() + saved = os.dup(0) + os.dup2(read_end, 0) + try: + worker._run("probe", ["bash", "-c", f"readlink /proc/self/fd/0 > {out}"]) + finally: + os.dup2(saved, 0) + for fd in (saved, read_end, write_end): + os.close(fd) + assert out.read_text().strip() == "/dev/null" + + +class TestDrainResilience: + def test_transcribe_worker_drain_null_attempts_counts_as_a_first_try(self, rig): + """Error: a job.json with "attempts": null must not take the worker down; it is a first try.""" + odd = rig.submit("odd", {"language": "en", "attempts": None}) + os.utime(odd, (1, 1)) + rig.submit("good") + assert worker.drain(rig.config) == {"done": 2, "failed": 0} + assert (rig.home / "done" / "odd.txt").exists() + assert (rig.home / "done" / "good.txt").exists() + + def test_transcribe_worker_drain_unexpected_error_in_one_job_is_logged_and_the_queue_goes_on(self, rig): + """Error: an exception the pipeline never anticipated lands in failed/ with its type, not on the run.""" + first = rig.submit("first") + os.utime(first, (1, 1)) + rig.submit("second") + + def explode(job: dict) -> list[str]: + raise RuntimeError("boom") + + rig.monkeypatch.setattr(worker, "_diarize_options", explode) + assert worker.drain(rig.config) == {"done": 0, "failed": 2} # both reached, neither crashed the run + for job_id in ("first", "second"): + log = (rig.home / "failed" / f"{job_id}.log").read_text() + assert "RuntimeError" in log and "boom" in log + assert not (rig.home / "work" / job_id).exists() + + def test_transcribe_worker_drain_unwritable_done_dir_is_a_failed_job_not_a_crash(self, rig): + """Error: an OSError while writing the transcript lands in failed/, and the run survives.""" + if os.geteuid() == 0: + pytest.skip("root ignores directory permissions") + (rig.home / "done").mkdir(parents=True) + (rig.home / "done").chmod(0o500) + try: + rig.submit("job1") + assert worker.drain(rig.config) == {"done": 0, "failed": 1} + assert "job1" in "".join(p.name for p in (rig.home / "failed").iterdir()) + finally: + (rig.home / "done").chmod(0o700) + + def test_transcribe_worker_drain_ffmpeg_failure_is_logged_with_its_stage(self, rig): + """Error: ffmpeg exits non-zero.""" + rig.monkeypatch.setenv("FAKE_FFMPEG_EXIT", "2") + rig.submit("job1") + assert worker.drain(rig.config) == {"done": 0, "failed": 1} + assert "ffmpeg" in (rig.home / "failed" / "job1.log").read_text() + + def test_transcribe_worker_drain_returns_at_once_when_another_worker_holds_the_lock(self, rig): + """Boundary: a second drain does not touch the queue while the first holds the lock.""" + import fcntl + rig.home.mkdir(parents=True, exist_ok=True) + rig.submit("job1") + with open(rig.home / "lock", "w", encoding="utf-8") as held: + fcntl.flock(held, fcntl.LOCK_EX | fcntl.LOCK_NB) + assert worker.drain(rig.config) == {"done": 0, "failed": 0} + assert (rig.home / "incoming" / "job1").exists() + assert rig.calls() == [] diff --git a/working/velox-touchpad-interrupt/touchpad-module-underside-2026-08-15.jpg b/working/velox-touchpad-interrupt/touchpad-module-underside-2026-08-15.jpg Binary files differnew file mode 100644 index 0000000..40c71af --- /dev/null +++ b/working/velox-touchpad-interrupt/touchpad-module-underside-2026-08-15.jpg |
