aboutsummaryrefslogtreecommitdiff
diff options
context:
space:
mode:
authorCraig Jennings <c@cjennings.net>2026-08-26 22:36:27 -0600
committerCraig Jennings <c@cjennings.net>2026-08-26 22:36:27 -0600
commit6e5629df0f374a7cbe29f5f5aa05708f1af550e3 (patch)
tree150f259737901bc9e4ee89f8413811f484bf6597
parent3df22bf89d9d808ffd49eb38c7ef06f4f6e0e33b (diff)
downloadarchsetup-6e5629df0f374a7cbe29f5f5aa05708f1af550e3.tar.gz
archsetup-6e5629df0f374a7cbe29f5f5aa05708f1af550e3.zip
chore(tasks): close the ribbon reseat and review seven stale tasksHEADmain
I reseated the velox input-cover ribbon on 08-15 and never marked the task, so I closed it on the verification it already carried and left the HandlePowerKey shield in place. The review batch regraded the left-drag gesture assessment and the per-channel audio controls from B to C, pointed the boot-failure retrospective's /boot assertion at the topgrade spec's kernel gate, and kept the other four as graded.
-rw-r--r--todo.org280
1 files changed, 149 insertions, 131 deletions
diff --git a/todo.org b/todo.org
index f69127d..eade66e 100644
--- a/todo.org
+++ b/todo.org
@@ -404,127 +404,6 @@ Alternatives if it drags on: change Signal's tray setting so it keeps a
window (=~/.config/Signal/ephemeral.json= =system-tray-setting=), or run a
waybar carrying the #5240 fallback.
-** TODO [#A] Reseat velox input-cover ribbon — phantom power button :bug:velox:hardware:
-DEADLINE: <2026-08-26 Wed>
-:PROPERTIES:
-:CREATED: [2026-08-13 Thu]
-:LAST_REVIEWED: 2026-08-13
-:END:
-Machine off, lift the input cover (Framework QR-guided procedure, 5
-fasteners), reseat its ribbon connector to the mainboard — disturbed in the
-2026-08-13 board swap. Root cause of every "mystery reboot" that day:
-chassis flex (flash-drive touch, ethernet bump, lid partially lowered)
-fired phantom power-button presses — journalctl -b -1 showed "Power key
-pressed short." → orderly logind poweroff, then the glitching button
-powered it back on. While in there, reseat the USB expansion cards too —
-the flaky slot (two hard resets, one no-enumeration) is likely the same
-flex problem.
-THIRD SYMPTOM (2026-08-13 evening): touchpad delivers ZERO input events —
-15s synchronized libinput debug-events capture while swiping caught
-nothing, though i2c enumeration and a driver rebind handshake are clean.
-Signature of a dead interrupt line on the same ribbon. Keyboard + power
-LED lines work; BT mouse is the interim pointer.
-ESCALATED 2026-08-13 21:00: a fourth event killed the machine THROUGH the
-shield. Previous boot's journal ends mid-line (tailscaled chatter) with no
-shutdown sequence at all — a hard power cut, not logind acting. So the
-glitch now reaches the EC/hardware power path, which no software setting
-can intercept. The reseat is the only fix, and this is a
-lose-work-without-warning failure mode, not an inconvenience.
-Interim shield (already live): /etc/systemd/logind.conf.d/powerkey.conf
-sets HandlePowerKey=ignore — phantom presses log but do nothing; EC-level
-10s hold still force-cuts. Consider keeping it even after the repair.
-Verify after reseat: flex the chassis edges + partially lower the lid, then
-grep the journal for new "Power key pressed" lines — zero means fixed.
-Must be done before the Sunday flight — a phantom press mid-travel with the
-shield on is survivable, but the connector should not be trusted at 30,000
-feet on the loose setting.
-
-*** 2026-08-19 Wed @ 14:40:00 -0700 Retracted: the RTC reset is not this task's, and I should not have filed it here
-I attributed the 2026-08-19 network outage to this ribbon earlier today. Craig
-pushed back — he reseated it before the trip to get the touchpad working — and
-he is right. The evidence does not support the attribution and some of it points
-the other way.
-
-What actually holds. Boot -3 ended at 01:33:18 with no shutdown sequence: no
-power-off target, no unmounting. The next boot's kernel line reads =rtc_cmos
-00:01: setting system clock to 2025-01-01T00:00:16 UTC=, a firmware default, so
-the RTC was reset rather than drifted. No firmware update was applied
-(=fwupdmgr get-history= is empty) and the battery is fine.
-
-What refutes the ribbon. This boot logged *zero* =Power key pressed= events, and
-so did the four boots before it. The phantom-press symptom had genuinely stopped
-after 08-15, exactly as the 08-16 session recorded. The earlier events logged a
-power-key press and an orderly poweroff; this logged neither, which makes it a
-different signature, not a worse version of the same one.
-
-What I got wrong methodologically: I anchored on the most salient open hardware
-task and read association as evidence. I even wrote "I can't prove it is the
-same connector" and then filed it here anyway, which is the tell.
-
-Two things I checked and can rule out. There were no OOM kills — the 3,433
-matching lines are a systemd unit named "Periodically re-score Claude Code
-processes for the OOM-killer" firing on a timer, not memory pressure, and there
-is not a single "Killed process" line. Thermal is clean; the only mentions are
-boot-time zone registration at 34C and 45C.
-
-One real thing the same window did surface, tracked separately: a python3 crash
-loop, 251 core dumps in the final ten minutes, SIGABRT with =XFreeThreads= and
-=PyEval_RestoreThread= in the trace. It does not explain the RTC, because
-software cannot clear it, but it is its own problem.
-
-The open question that would settle the RTC is for Craig, not the journal: a
-long power-button hold on a Framework triggers an EC-level reset that clears the
-RTC, which fits a wedged machine being forced off. A 4-second hold would not.
-
-*** 2026-08-17 Mon @ 19:57:42 -0700 Not done, and the two symptoms now disagree
-The reseat did not happen before the flight, and velox is travelling. The
-deadline blew past on 08-14.
-
-The two symptoms have separated, which is worth recording because it changes
-what the evidence proves. The phantom presses have stopped: fifteen "Power key
-pressed" entries between 08-14 04:29 and 08-15 20:04, then nothing at all
-across five boots including today's. The touchpad has not — there is still no
-touchpad node under =/dev/input/by-path/=, which is the same dead interrupt
-line the body describes.
-
-So the quiet power button is not evidence the connector reseated itself. The
-interrupt line is the symptom that cannot be masked in software, and it is
-still dead, so the ribbon is still unseated. The most likely reason the
-presses stopped is that the machine has been sitting on hotel surfaces instead
-of being carried and flexed.
-
-The interim shield is still live (=HandlePowerKey=ignore=), and the escalation
-note stands: an EC-level glitch cuts power below systemd regardless of it.
-*** 2026-08-15 Sat @ 23:05:00 -0500 The reseat did happen, and the touchpad came back — this contradicts the 08-17 read
-Recording this because a parallel session concluded on 08-17 that the reseat had
-not happened and the touchpad was still dead. Both halves were done and verified
-that night, so the two accounts disagree and the disagreement should be visible
-rather than silently resolved by whichever session committed last.
-
-What was done: the input-cover ribbon was reseated first, which fixed the
-phantom power button — the 22:09 boot logged zero =Power key pressed= lines
-after Craig flexed the chassis, against nine on the boot before. The touchpad
-did not change, because the input-cover ribbon is not its connector. The 4-pin
-connector beside the printed =TOUCHPAD= label is silkscreened =PIN 1-2 GND /
-PIN 3-4 VCC= — pure power, so it cannot carry i2c or an interrupt. Reseating the
-ribbon that actually crosses to the mainboard fixed it.
-
-Measured, not assumed: the touchpad interrupt (=amd_gpio= pin 8) went from 0
-counts across all 24 CPUs to 1795, and =i2c_hid_acpi ... did not ack reset
-within 1000 ms= disappeared from the boot log. Craig confirmed the pointer moved.
-
-*Why the 08-17 probe likely misread it:* it checked for a node under
-=/dev/input/by-path/=. i2c-HID touchpads frequently get no =by-path= symlink
-even when fully working, so its absence is not evidence of a dead interrupt
-line. The falsifiable check is the interrupt count in =/proc/interrupts= while
-the pad is being touched, or the reset message in =dmesg=.
-
-*Left open rather than closed* — velox was refusing ssh at merge time on 08-20,
-so the current state could not be re-verified, and a later regression cannot be
-ruled out. One second of Craig's time settles it: move the pointer. If it works,
-close this; if it does not, the interrupt line went back down and that is new
-information.
-
** DOING [#A] Velox reinstall — DR test of archangel + archsetup :velox:chore:
DEADLINE: <2026-08-15 Sat>
:PROPERTIES:
@@ -1321,17 +1200,26 @@ measuring it. Verify placement at the time of the move rather than trusting a
recorded list.
** TODO [#B] Velox boot-failure retrospective — upgrade guard gaps :bug:zfs:maint:
:PROPERTIES:
-:LAST_REVIEWED: 2026-07-21
+:LAST_REVIEWED: 2026-08-26
:END:
Post-mortem for the 2026-07-15 velox no-kernel boot failure, from the archsetup/maint code review:
- maint's UPDATE remedy runs a plain =yay -Syu --noconfirm= (remedies.py:297). The live-update guard (guard.py) only matches mesa/hyprland (the 2026-06-07 live-swap class) — it never checks /boot, kernel, initramfs, or mkinitcpio exit. No post-upgrade /boot assertion exists. An interrupted kernel transaction slips straight through.
- Add a post-upgrade /boot assertion: after a transaction touching linux/linux-*, confirm vmlinuz-* + initramfs-*.img present and mkinitcpio exit 0; refuse to end the run (or page Craig) otherwise. Would have caught this.
- Sanoid-vs-actual dataset drift: configure_zfs_snapshots configures zroot/var/log + zroot/var/lib/pacman as separate datasets; velox's actual layout has neither separate (/var/log sits inside zroot/var). Reconcile.
- CONFIRMED (2026-07-21): the pre-pacman snapshot hook fired on velox — the 2026-07-15 no-kernel boot was recovered via the pre-pacman ZFS snapshot rollback, and velox is back on the tailnet running linux-lts 6.18.38 with initramfs present (2026-07-19 session). Root-cause hook-ordering fix shipped separately. Still open: the post-upgrade /boot assertion in guard.py and the sanoid-vs-actual dataset drift reconcile (the two bullets above).
-
-** TODO [#B] Assess a Hyprland left-drag window gesture :feature:hyprland:
+*** 2026-08-26 Wed @ 22:35:01 -0600 The /boot assertion now lives in the topgrade spec; the dataset drift is what remains here
+The post-upgrade /boot assertion is covered by the kernel-modules-check gate in
+[[file:docs/specs/2026-08-25-topgrade-guarded-upgrade-spec.org][the topgrade guarded-upgrade spec]]
+(dkms built for the new kernel, initramfs newer than vmlinuz, pre-pacman
+snapshot on a ZFS root), which ships with that spec's Phase 1 rather than here.
+What this task still owns is the sanoid-vs-actual dataset drift: whether to
+split zroot/var/log and zroot/var/lib/pacman out as configure_zfs_snapshots
+assumes, or change the config to match the layout velox actually has. That is
+a call I have not made, so the task stays [#B] and not solo.
+
+** TODO [#C] Assess a Hyprland left-drag window gesture :feature:hyprland:
:PROPERTIES:
-:LAST_REVIEWED: 2026-07-21
+:LAST_REVIEWED: 2026-08-26
:END:
Evaluate whether a global left-click drag can move ordinary windows without
breaking application selection, text interaction, or Wayland security
@@ -1380,15 +1268,15 @@ the audit plus wiring plus docs runs past thirty minutes on its own.
** TODO [#B] Add storage-capacity signals to the maintenance module :feature:maint:
:PROPERTIES:
-:LAST_REVIEWED: 2026-07-21
+:LAST_REVIEWED: 2026-08-26
:END:
Investigate capacity and growth diagnostics for full disks, identify the
appropriate remedies, and incorporate a clear storage signal into the
maintenance console.
-** TODO [#B] Add per-channel controls to the audio panel :feature:audio:
+** TODO [#C] Add per-channel controls to the audio panel :feature:audio:
:PROPERTIES:
-:LAST_REVIEWED: 2026-07-21
+:LAST_REVIEWED: 2026-08-26
:END:
Expose channel-level input and output volume controls without losing the
existing device-level workflow.
@@ -3203,7 +3091,7 @@ Grading (2026-08-25 review): Major severity — the doctor's verdict is silently
** TODO [#C] Weather chip color signals unclear + unenforced :bug:dotfiles:waybar:weather:
:PROPERTIES:
-:LAST_REVIEWED: 2026-07-21
+:LAST_REVIEWED: 2026-08-26
:END:
From the roam inbox (2026-07-20): the shipped Waybar weather chip's comfort coloring reads as noise — it shows amber for no clear reason, and some items are bolded, which isn't a legible signal. Craig's intended scheme (every item except the arrow key colored by whether the weather is comfortable; NO bold or italic anywhere):
- Normal — all text white: temp in 60-85; condition sunny/clear/etc.
@@ -3268,13 +3156,13 @@ Reproduced in ~1 minute of install: =dkms install zfs/2.3.3 -k 6.18.38-2-lts= ex
** TODO [#C] Waybar collapse control: replace the triangle glyph :feature:waybar:
:PROPERTIES:
-:LAST_REVIEWED: 2026-07-14
+:LAST_REVIEWED: 2026-08-26
:END:
From the 2026-07-04 roam capture. The waybar collapse mechanism (click the triangle, the bar sections redisplay shortened) works, but the triangle glyph doesn't match the instrument-console aesthetic the panels now use. Replace it with something in keeping with the console look. Aesthetic decision — bring Craig two or three concrete glyph/style options (a machined chevron, a console-key style expander, an engraved caret) before wiring. Dotfiles waybar config (handled per the archsetup-owns-dotfiles rule). Raised alongside the net-panel/audio speedrun; deferred from it because the glyph choice is a taste call.
** TODO [#C] Net panel: driver-health diagnostic tier :feature:network:
:PROPERTIES:
-:LAST_REVIEWED: 2026-07-14
+:LAST_REVIEWED: 2026-08-26
:END:
Follow-up from the 2026-07-04 net-panel hardening speedrun (Craig's cj question on the no-WiFi item). The shipped no-wifi-hardware verdict covers "no adapter at all." This tier covers "adapter present but the driver is wedged": read-only health signals — =ip link= (device present but no-carrier / down), =dmesg= / =journalctl -k= for firmware-load failures, =rfkill= for a hard block, =modinfo= / =lsmod= for the driver module — classified before a generic reset. Remedy actions: a privileged =modprobe -r <mod> && modprobe <mod>= reload of the wifi driver, and a firmware-package pointer when the failure is a missing/failed firmware load. Dotfiles net-package work (handled per the archsetup-owns-dotfiles rule). Design pass first to decide whether it's worth a repair tier vs a needs-user-action pointer.
@@ -3409,6 +3297,136 @@ Re-graded =[#C]= → =[#D]= per the bug matrix. There is no defect to fix here;
The maintenance console's coredump metric flagged telega-server on ratio (8 coredumps) and velox (18). Root cause was a version skew: the Dockerized =zevlg/telega-server:latest= is frozen at the 2026-06-05 build while the installed elisp lagged at 20260513, so the newer server's plist parser choked on the older elisp's output. .emacs.d fixed it by upgrading telega to 20260706 on both machines (docker kept, =docker pull= is a no-op against the frozen image). Host-coredump pollution should stop. If zevlg later pushes a =:latest= that outruns the installed elisp, the skew and the coredumps recur — the tell is a fresh =tdat_plist_value:500= assertion in =~/.telega/telega-server.log=. The durable escape is a host-native pinned TDLib build, at the cost of an AUR source build.
* Archsetup Resolved
+** DONE [#A] Reseat velox input-cover ribbon — phantom power button :bug:velox:hardware:
+CLOSED: [2026-08-26 Wed] DEADLINE: <2026-08-26 Wed>
+:PROPERTIES:
+:CREATED: [2026-08-13 Thu]
+:LAST_REVIEWED: 2026-08-13
+:END:
+Machine off, lift the input cover (Framework QR-guided procedure, 5
+fasteners), reseat its ribbon connector to the mainboard — disturbed in the
+2026-08-13 board swap. Root cause of every "mystery reboot" that day:
+chassis flex (flash-drive touch, ethernet bump, lid partially lowered)
+fired phantom power-button presses — journalctl -b -1 showed "Power key
+pressed short." → orderly logind poweroff, then the glitching button
+powered it back on. While in there, reseat the USB expansion cards too —
+the flaky slot (two hard resets, one no-enumeration) is likely the same
+flex problem.
+THIRD SYMPTOM (2026-08-13 evening): touchpad delivers ZERO input events —
+15s synchronized libinput debug-events capture while swiping caught
+nothing, though i2c enumeration and a driver rebind handshake are clean.
+Signature of a dead interrupt line on the same ribbon. Keyboard + power
+LED lines work; BT mouse is the interim pointer.
+ESCALATED 2026-08-13 21:00: a fourth event killed the machine THROUGH the
+shield. Previous boot's journal ends mid-line (tailscaled chatter) with no
+shutdown sequence at all — a hard power cut, not logind acting. So the
+glitch now reaches the EC/hardware power path, which no software setting
+can intercept. The reseat is the only fix, and this is a
+lose-work-without-warning failure mode, not an inconvenience.
+Interim shield (already live): /etc/systemd/logind.conf.d/powerkey.conf
+sets HandlePowerKey=ignore — phantom presses log but do nothing; EC-level
+10s hold still force-cuts. Consider keeping it even after the repair.
+Verify after reseat: flex the chassis edges + partially lower the lid, then
+grep the journal for new "Power key pressed" lines — zero means fixed.
+Must be done before the Sunday flight — a phantom press mid-travel with the
+shield on is survivable, but the connector should not be trusted at 30,000
+feet on the loose setting.
+
+*** 2026-08-19 Wed @ 14:40:00 -0700 Retracted: the RTC reset is not this task's, and I should not have filed it here
+I attributed the 2026-08-19 network outage to this ribbon earlier today. Craig
+pushed back — he reseated it before the trip to get the touchpad working — and
+he is right. The evidence does not support the attribution and some of it points
+the other way.
+
+What actually holds. Boot -3 ended at 01:33:18 with no shutdown sequence: no
+power-off target, no unmounting. The next boot's kernel line reads =rtc_cmos
+00:01: setting system clock to 2025-01-01T00:00:16 UTC=, a firmware default, so
+the RTC was reset rather than drifted. No firmware update was applied
+(=fwupdmgr get-history= is empty) and the battery is fine.
+
+What refutes the ribbon. This boot logged *zero* =Power key pressed= events, and
+so did the four boots before it. The phantom-press symptom had genuinely stopped
+after 08-15, exactly as the 08-16 session recorded. The earlier events logged a
+power-key press and an orderly poweroff; this logged neither, which makes it a
+different signature, not a worse version of the same one.
+
+What I got wrong methodologically: I anchored on the most salient open hardware
+task and read association as evidence. I even wrote "I can't prove it is the
+same connector" and then filed it here anyway, which is the tell.
+
+Two things I checked and can rule out. There were no OOM kills — the 3,433
+matching lines are a systemd unit named "Periodically re-score Claude Code
+processes for the OOM-killer" firing on a timer, not memory pressure, and there
+is not a single "Killed process" line. Thermal is clean; the only mentions are
+boot-time zone registration at 34C and 45C.
+
+One real thing the same window did surface, tracked separately: a python3 crash
+loop, 251 core dumps in the final ten minutes, SIGABRT with =XFreeThreads= and
+=PyEval_RestoreThread= in the trace. It does not explain the RTC, because
+software cannot clear it, but it is its own problem.
+
+The open question that would settle the RTC is for Craig, not the journal: a
+long power-button hold on a Framework triggers an EC-level reset that clears the
+RTC, which fits a wedged machine being forced off. A 4-second hold would not.
+
+*** 2026-08-17 Mon @ 19:57:42 -0700 Not done, and the two symptoms now disagree
+The reseat did not happen before the flight, and velox is travelling. The
+deadline blew past on 08-14.
+
+The two symptoms have separated, which is worth recording because it changes
+what the evidence proves. The phantom presses have stopped: fifteen "Power key
+pressed" entries between 08-14 04:29 and 08-15 20:04, then nothing at all
+across five boots including today's. The touchpad has not — there is still no
+touchpad node under =/dev/input/by-path/=, which is the same dead interrupt
+line the body describes.
+
+So the quiet power button is not evidence the connector reseated itself. The
+interrupt line is the symptom that cannot be masked in software, and it is
+still dead, so the ribbon is still unseated. The most likely reason the
+presses stopped is that the machine has been sitting on hotel surfaces instead
+of being carried and flexed.
+
+The interim shield is still live (=HandlePowerKey=ignore=), and the escalation
+note stands: an EC-level glitch cuts power below systemd regardless of it.
+*** 2026-08-15 Sat @ 23:05:00 -0500 The reseat did happen, and the touchpad came back — this contradicts the 08-17 read
+Recording this because a parallel session concluded on 08-17 that the reseat had
+not happened and the touchpad was still dead. Both halves were done and verified
+that night, so the two accounts disagree and the disagreement should be visible
+rather than silently resolved by whichever session committed last.
+
+What was done: the input-cover ribbon was reseated first, which fixed the
+phantom power button — the 22:09 boot logged zero =Power key pressed= lines
+after Craig flexed the chassis, against nine on the boot before. The touchpad
+did not change, because the input-cover ribbon is not its connector. The 4-pin
+connector beside the printed =TOUCHPAD= label is silkscreened =PIN 1-2 GND /
+PIN 3-4 VCC= — pure power, so it cannot carry i2c or an interrupt. Reseating the
+ribbon that actually crosses to the mainboard fixed it.
+
+Measured, not assumed: the touchpad interrupt (=amd_gpio= pin 8) went from 0
+counts across all 24 CPUs to 1795, and =i2c_hid_acpi ... did not ack reset
+within 1000 ms= disappeared from the boot log. Craig confirmed the pointer moved.
+
+*Why the 08-17 probe likely misread it:* it checked for a node under
+=/dev/input/by-path/=. i2c-HID touchpads frequently get no =by-path= symlink
+even when fully working, so its absence is not evidence of a dead interrupt
+line. The falsifiable check is the interrupt count in =/proc/interrupts= while
+the pad is being touched, or the reset message in =dmesg=.
+
+*Left open rather than closed* — velox was refusing ssh at merge time on 08-20,
+so the current state could not be re-verified, and a later regression cannot be
+ruled out. One second of Craig's time settles it: move the pointer. If it works,
+close this; if it does not, the interrupt line went back down and that is new
+information.
+
+*** 2026-08-26 Wed @ 22:30:46 -0600 Closed: the reseat was done on 08-15 and the task was never marked
+I reseated the ribbon on 2026-08-15 and never closed this. The 08-15 entry
+above already records the verification: zero =Power key pressed= lines on the
+22:09 boot after flexing the chassis, the touchpad interrupt count back up
+once the right connector was reseated. This boot shows zero presses as well.
+The interim shield (=HandlePowerKey=ignore= in
+=/etc/systemd/logind.conf.d/powerkey.conf=) is still live; I'm leaving it in
+place, since a phantom press with it on costs nothing and without it costs
+the session.
** DONE [#B] Velox touchpad interrupt line is dead — needs a part or a BIOS fix :bug:velox:hardware:
CLOSED: [2026-08-15 Sat]
:PROPERTIES: