aboutsummaryrefslogtreecommitdiff
diff options
context:
space:
mode:
-rw-r--r--archive/task-archive.org754
-rw-r--r--todo.org798
2 files changed, 776 insertions, 776 deletions
diff --git a/archive/task-archive.org b/archive/task-archive.org
index dfd273d..f2edf58 100644
--- a/archive/task-archive.org
+++ b/archive/task-archive.org
@@ -3104,3 +3104,757 @@ the stage defaults; render to a non-stowed path and have hypridle read
that; or keep it tracked but commit a machine-neutral render. The first
looks right — the store already holds the real source of truth, and the
rendered file is a build artifact.
+** DONE [#A] Reseat velox input-cover ribbon — phantom power button :bug:velox:hardware:
+CLOSED: [2026-08-26 Wed] DEADLINE: <2026-08-26 Wed>
+:PROPERTIES:
+:CREATED: [2026-08-13 Thu]
+:LAST_REVIEWED: 2026-08-13
+:END:
+Machine off, lift the input cover (Framework QR-guided procedure, 5
+fasteners), reseat its ribbon connector to the mainboard — disturbed in the
+2026-08-13 board swap. Root cause of every "mystery reboot" that day:
+chassis flex (flash-drive touch, ethernet bump, lid partially lowered)
+fired phantom power-button presses — journalctl -b -1 showed "Power key
+pressed short." → orderly logind poweroff, then the glitching button
+powered it back on. While in there, reseat the USB expansion cards too —
+the flaky slot (two hard resets, one no-enumeration) is likely the same
+flex problem.
+THIRD SYMPTOM (2026-08-13 evening): touchpad delivers ZERO input events —
+15s synchronized libinput debug-events capture while swiping caught
+nothing, though i2c enumeration and a driver rebind handshake are clean.
+Signature of a dead interrupt line on the same ribbon. Keyboard + power
+LED lines work; BT mouse is the interim pointer.
+ESCALATED 2026-08-13 21:00: a fourth event killed the machine THROUGH the
+shield. Previous boot's journal ends mid-line (tailscaled chatter) with no
+shutdown sequence at all — a hard power cut, not logind acting. So the
+glitch now reaches the EC/hardware power path, which no software setting
+can intercept. The reseat is the only fix, and this is a
+lose-work-without-warning failure mode, not an inconvenience.
+Interim shield (already live): /etc/systemd/logind.conf.d/powerkey.conf
+sets HandlePowerKey=ignore — phantom presses log but do nothing; EC-level
+10s hold still force-cuts. Consider keeping it even after the repair.
+Verify after reseat: flex the chassis edges + partially lower the lid, then
+grep the journal for new "Power key pressed" lines — zero means fixed.
+Must be done before the Sunday flight — a phantom press mid-travel with the
+shield on is survivable, but the connector should not be trusted at 30,000
+feet on the loose setting.
+
+*** 2026-08-19 Wed @ 14:40:00 -0700 Retracted: the RTC reset is not this task's, and I should not have filed it here
+I attributed the 2026-08-19 network outage to this ribbon earlier today. Craig
+pushed back — he reseated it before the trip to get the touchpad working — and
+he is right. The evidence does not support the attribution and some of it points
+the other way.
+
+What actually holds. Boot -3 ended at 01:33:18 with no shutdown sequence: no
+power-off target, no unmounting. The next boot's kernel line reads =rtc_cmos
+00:01: setting system clock to 2025-01-01T00:00:16 UTC=, a firmware default, so
+the RTC was reset rather than drifted. No firmware update was applied
+(=fwupdmgr get-history= is empty) and the battery is fine.
+
+What refutes the ribbon. This boot logged *zero* =Power key pressed= events, and
+so did the four boots before it. The phantom-press symptom had genuinely stopped
+after 08-15, exactly as the 08-16 session recorded. The earlier events logged a
+power-key press and an orderly poweroff; this logged neither, which makes it a
+different signature, not a worse version of the same one.
+
+What I got wrong methodologically: I anchored on the most salient open hardware
+task and read association as evidence. I even wrote "I can't prove it is the
+same connector" and then filed it here anyway, which is the tell.
+
+Two things I checked and can rule out. There were no OOM kills — the 3,433
+matching lines are a systemd unit named "Periodically re-score Claude Code
+processes for the OOM-killer" firing on a timer, not memory pressure, and there
+is not a single "Killed process" line. Thermal is clean; the only mentions are
+boot-time zone registration at 34C and 45C.
+
+One real thing the same window did surface, tracked separately: a python3 crash
+loop, 251 core dumps in the final ten minutes, SIGABRT with =XFreeThreads= and
+=PyEval_RestoreThread= in the trace. It does not explain the RTC, because
+software cannot clear it, but it is its own problem.
+
+The open question that would settle the RTC is for Craig, not the journal: a
+long power-button hold on a Framework triggers an EC-level reset that clears the
+RTC, which fits a wedged machine being forced off. A 4-second hold would not.
+
+*** 2026-08-17 Mon @ 19:57:42 -0700 Not done, and the two symptoms now disagree
+The reseat did not happen before the flight, and velox is travelling. The
+deadline blew past on 08-14.
+
+The two symptoms have separated, which is worth recording because it changes
+what the evidence proves. The phantom presses have stopped: fifteen "Power key
+pressed" entries between 08-14 04:29 and 08-15 20:04, then nothing at all
+across five boots including today's. The touchpad has not — there is still no
+touchpad node under =/dev/input/by-path/=, which is the same dead interrupt
+line the body describes.
+
+So the quiet power button is not evidence the connector reseated itself. The
+interrupt line is the symptom that cannot be masked in software, and it is
+still dead, so the ribbon is still unseated. The most likely reason the
+presses stopped is that the machine has been sitting on hotel surfaces instead
+of being carried and flexed.
+
+The interim shield is still live (=HandlePowerKey=ignore=), and the escalation
+note stands: an EC-level glitch cuts power below systemd regardless of it.
+*** 2026-08-15 Sat @ 23:05:00 -0500 The reseat did happen, and the touchpad came back — this contradicts the 08-17 read
+Recording this because a parallel session concluded on 08-17 that the reseat had
+not happened and the touchpad was still dead. Both halves were done and verified
+that night, so the two accounts disagree and the disagreement should be visible
+rather than silently resolved by whichever session committed last.
+
+What was done: the input-cover ribbon was reseated first, which fixed the
+phantom power button — the 22:09 boot logged zero =Power key pressed= lines
+after Craig flexed the chassis, against nine on the boot before. The touchpad
+did not change, because the input-cover ribbon is not its connector. The 4-pin
+connector beside the printed =TOUCHPAD= label is silkscreened =PIN 1-2 GND /
+PIN 3-4 VCC= — pure power, so it cannot carry i2c or an interrupt. Reseating the
+ribbon that actually crosses to the mainboard fixed it.
+
+Measured, not assumed: the touchpad interrupt (=amd_gpio= pin 8) went from 0
+counts across all 24 CPUs to 1795, and =i2c_hid_acpi ... did not ack reset
+within 1000 ms= disappeared from the boot log. Craig confirmed the pointer moved.
+
+*Why the 08-17 probe likely misread it:* it checked for a node under
+=/dev/input/by-path/=. i2c-HID touchpads frequently get no =by-path= symlink
+even when fully working, so its absence is not evidence of a dead interrupt
+line. The falsifiable check is the interrupt count in =/proc/interrupts= while
+the pad is being touched, or the reset message in =dmesg=.
+
+*Left open rather than closed* — velox was refusing ssh at merge time on 08-20,
+so the current state could not be re-verified, and a later regression cannot be
+ruled out. One second of Craig's time settles it: move the pointer. If it works,
+close this; if it does not, the interrupt line went back down and that is new
+information.
+
+*** 2026-08-26 Wed @ 22:30:46 -0600 Closed: the reseat was done on 08-15 and the task was never marked
+I reseated the ribbon on 2026-08-15 and never closed this. The 08-15 entry
+above already records the verification: zero =Power key pressed= lines on the
+22:09 boot after flexing the chassis, the touchpad interrupt count back up
+once the right connector was reseated. This boot shows zero presses as well.
+The interim shield (=HandlePowerKey=ignore= in
+=/etc/systemd/logind.conf.d/powerkey.conf=) is still live; I'm leaving it in
+place, since a phantom press with it on costs nothing and without it costs
+the session.
+** CANCELLED [#B] agent-text relay reports success for a message that went nowhere :bug:
+CLOSED: [2026-08-19 Wed]
+:PROPERTIES:
+:CREATED: [2026-08-19 Wed]
+:LAST_REVIEWED: 2026-08-19
+:END:
+
+Not a defect. rulesets refuted it with measurements and I reproduced theirs
+before accepting: on velox, whose account store is empty,
+=signal-cli -a +15550000000 send= exits 1 with "User +15550000000 is not
+registered", and =ssh 100.71.182.1 'exit 7'= returns 7, so a non-zero code
+propagates faithfully back through the relay. The loop's
+=[ "$rc" -eq 0 ] && break= therefore advances to the next host exactly as
+intended. signal-cli fails closed.
+
+I filed this off a conditional in their handoff — ".emacs.d raised a case
+neither of you tested ... *if* signal-cli send exits zero against an empty
+account store" — and turned the "if" into a graded [#B] with a =:blocked:= tag
+on another project, without running the one command that settles it. The
+machine that proves it was in front of me the whole time. Their ask is fair and
+I am recording it rather than the outcome alone: verify before filing a defect
+against someone else's work, especially one carrying a blocking tag.
+** DONE [#B] Clock/DNS bootstrap deadlock — recovery needs a second device :bug:velox:
+CLOSED: [2026-08-19 Wed]
+:PROPERTIES:
+:CREATED: [2026-08-19 Wed]
+:LAST_REVIEWED: 2026-08-19
+:END:
+
+The installer wrote both halves of a deadlock. =configure_dns= pins
+=DNSOverTLS=yes= with =DNSSEC=yes=, and both validate against the wall clock;
+the chrony step enables chronyd without writing a config, so the machine runs
+Arch's stock one whose only source is =pool 2.arch.pool.ntp.org= — a hostname.
+Boot with a wrong clock and DoT certificate validation fails, so nothing
+resolves; chrony then cannot resolve its pool, so the clock stays wrong.
+Neither side moves. It caught velox on the road 2026-08-19 and had to be
+diagnosed from a phone.
+
+Fixed at the root: the installer now writes
+=/etc/chrony.d/10-bootstrap-ip-ntp.conf= with two IP-addressed Cloudflare
+sources and points stock chrony.conf at the drop-in. An address needs no DNS
+and carries no certificate, so the escape hatch holds whatever broke the clock.
+velox has the same drop-in applied live, verified with =chronyc -n sources=
+(=162.159.200.1= selected) and =timedatectl= reporting synchronized.
+
+What is left here is the part I could not verify: the decisive test is a full
+power-down and cold boot, confirming the clock corrects itself untouched. See
+the manual-testing entry. Until that runs, the fix is sound by construction
+rather than demonstrated.
+
+Grading: Critical severity (total loss of network — no DNS means no egress, and
+recovery needs a second device) x some users sometimes (only machines that boot
+with a wrong clock, which is any RTC fault, BIOS reset, or drained cell) = P2 =
+[#B]. Graded on the being-in-it, not the getting-into-it: once the machine is in
+this state it is fully offline with no local path out.
+
+*** 2026-08-19 Wed @ 12:25:00 -0700 Reproduced it, and the mechanism was not what either of us said
+I wound velox's clock back 27 days with chronyd stopped and watched it fail.
+Resolution died outright, and plain UDP/53 to 1.1.1.1 kept answering throughout
+— the discriminator the doctor keys on, confirmed live rather than reasoned.
+
+The cause is DNSSEC, not DNS-over-TLS. resolved logged =signature-expired=
+against the root DNSKEY and every DS beneath it. The DoT handshake to
+=1.1.1.1:853= verified clean at that same clock, and the Cloudflare certificate
+runs Dec 2025 to Dec 2026, so it was never outside its window. An RRSIG window
+is days to weeks and a certificate is good for a year, so a skew that breaks
+DNSSEC normally leaves DoT untouched. The phone session blamed the certificate
+and I carried that forward into the first commit; both were wrong.
+
+=DNSSEC=allow-downgrade= does not rescue it either, which matters because it is
+the obvious reach and it is what ratio runs. resolved downgrades when a server
+lacks DNSSEC support, and a signature-window failure is a validation failure, so
+no downgrade fires. Six retries over eighteen seconds plus
+=resolvectl reset-server-features=, all dead. I briefly believed otherwise off a
+test whose success was a cache hit (=Data from: cache network=).
+
+So ratio was exposed after all, and I have given it the same drop-in. Its
+=162.159.200.1= is selected and its clock is synchronized.
+
+The fix itself is verified end to end: with the clock wound back and no DNS at
+all, chronyd reached the IP-addressed source and stepped the clock from
+2026-07-23 straight back to 2026-08-19. That is the whole claim, demonstrated
+rather than argued.
+
+Also settled: the clock landed on 2026-07-23 because that is systemd 261.2's
+build date to the minute (=/usr/lib/systemd/systemd=, 10:43:59), and systemd
+advances a garbage RTC to its own build epoch at boot. Not timesyncd's
+last-good-sync timestamp, which cannot be it — timesyncd is disabled here. That
+also confirms the RTC really was reading earlier than that, so the coin cell
+stays the prime suspect.
+
+*** 2026-08-19 Wed @ 10:12:00 -0700 Root fix, doctor verdict, and taxonomy entry landed
+The installer carries the drop-in; =post-rebuild-check= grew a sixth check that
+fails a machine whose every NTP source is a hostname; the net failure taxonomy
+gained the mode in its DNS layer plus a cluster 5 triage line, and its existing
+egress-layer clock entry now says outright that its remedy does not apply when
+DoT or DNSSEC is on.
+
+The doctor half is in dotfiles: =classify.py= reached "DNS not resolving → net
+repair dns-test" here, which cannot help, because every public resolver fails
+the same clock-sensitive validation — so the doctor sent you round a loop. It
+now emits a =clock-dns= row ahead of the generic DNS verdict. Detection is
+deliberately DNS-free: a local =timedatectl= read for sync state, and a bypass
+query addressed by IP over plain UDP/53 to tell "resolved is refusing to
+validate" apart from "DNS is genuinely dead".
+** DONE [#C] DNSSEC strictness on the travelling laptop :velox:
+CLOSED: [2026-08-19 Wed]
+:PROPERTIES:
+:CREATED: [2026-08-19 Wed]
+:LAST_REVIEWED: 2026-08-19
+:END:
+
+Craig chose =allow-downgrade= everywhere. Applied to velox, ratio, and the
+installer, and ratio's =DNSOverTLS= tightened from =opportunistic= to =yes= in
+the same pass, so all three now agree: encrypted DNS always, validation
+best-effort.
+
+The reasoning that settled it: the deadlock is fixed by the IP-addressed NTP
+source, and =allow-downgrade= was measured not to help with it at all. What
+=allow-downgrade= does buy is the venue-resolver case the taxonomy documents,
+where =yes= turns a resolver that mangles DNSSEC records into no answer at all.
+That is a hotel and airport problem, so it is velox's problem, and the
+encryption is the half worth being strict about.
+** CANCELLED [#C] Branch network policy on laptop vs desktop in the installer :feature:
+CLOSED: [2026-08-19 Wed]
+:PROPERTIES:
+:CREATED: [2026-08-19 Wed]
+:LAST_REVIEWED: 2026-08-19
+:END:
+
+Cancelled because the decision above emptied it. All three motivating cases now
+want the same value on every machine: =DNSSEC=allow-downgrade=, a stable
+per-network wifi MAC, and an IP-addressed NTP source. A branch with nothing to
+put on either side is machinery built for a divergence that does not exist, and
+it would be the kind of scaffolding that rots unread.
+
+Worth keeping the observation, which is the part with a shelf life: when a
+network default does need to differ by machine class, the test already exists.
+=ls /sys/class/power_supply/BAT*= is what =prune_waybar_battery=, the ppd mask,
+and the TLP config all key on. Reopen this then rather than building it now.
+** DONE [#C] Automate the clock/DNS deadlock repair in the net doctor :feature:
+CLOSED: [2026-08-19 Wed]
+:PROPERTIES:
+:CREATED: [2026-08-19 Wed]
+:LAST_REVIEWED: 2026-08-19
+:END:
+
+Shipped as the =clock-ip-ntp= repair, and the verdict is =fixable= rather than
+terminal. Both open questions got answered by driving a real deadlock instead of
+reasoning about it: =chronyc add server= returns =200 OK= against a running
+chronyd, and =makestep= needs a sample to land, so it took four calls and about
+eight seconds rather than working on the first. The repair retries accordingly.
+
+Verified end to end on velox against a genuine deadlock (wrong clock, chronyd
+running with only an unresolvable hostname source, DNS dead): the repair
+corrected the clock in 6.1 seconds and DNS came back.
+
+The live run also caught a defect no unit test would have. The doctor reported
+"Saved password for SpectrumSetup-3C was rejected" — because
+=_recent_auth_failure= greps =journalctl --since -5min=, and a clock weeks off
+windows onto a different incident's entries. It would have sent Craig to
+re-enter a password that was never wrong. Fixed twice over: the journal half is
+now skipped when the clock is untrustworthy, and the clock verdict is ordered
+above the auth verdict, since everything below it reasons over timestamps that
+only mean something once the clock is right. Airplane mode and hard rfkill stay
+above, being physical states the clock has no bearing on.
+** DONE [#D] net-scenarios harness times out under back-to-back suite runs :test:tooling:
+CLOSED: [2026-08-19 Wed]
+:PROPERTIES:
+:LAST_REVIEWED: 2026-07-24
+:END:
+=tests/net-scenarios/test_run_net_scenarios.py= errored on all 5 tests twice during round 12, each time =subprocess.TimeoutExpired= after its 20s budget on =scripts/testing/run-net-scenarios.sh --target root@fake-vm=. Both occurrences were in =make test-unit= runs launched immediately after a previous full run. It then passed 6 runs in a row (3 on a pristine tree, 3 with the round-12 change), and standalone it finishes in 0.09s, so this is not a regression from any code change.
+
+The harness stubs =ssh=, =rsync= and =jq= onto =PATH=, so nothing should touch the network at all — which is what makes a 20s timeout suspicious rather than merely slow. Worth reproducing under load before deciding whether the fix is a larger timeout or a real hang in the script. Evidence logs from the round: =/tmp/tu.log= and =/tmp/tu2.log= (tmpfs, gone after reboot).
+
+Not graded on the bug matrix: it is test infrastructure, not the shipped codebase.
+
+*** 2026-08-08 Sat @ 05:05:00 -0500 Recurred under concurrent load, same signature
+All 5 tests hit the 20s TimeoutExpired again during a =make test-unit= run
+that overlapped two review subagents running their own suites on the box.
+Standalone immediately after: 0.095s, all pass; the following quiet-machine
+full run was clean. Confirms the load-sensitivity read — reproduce under
+deliberate load before choosing between a bigger budget and a real hang.
+
+
+*** 2026-08-19 Wed @ 14:50:00 -0700 Root-caused and fixed: inherited stdin, not load
+Not load, and not the network. The harness stubs ssh as =cat >/dev/null=, which
+drains stdin to EOF. With no explicit stdin the stub inherits whatever the test
+runner had, so it returned instantly when stdin was redirected and blocked
+forever when it was a terminal or a live pipe. All five tests then burned their
+20-second budget.
+
+That is why it looked like a load effect: a run launched immediately after
+another inherited a different stdin than a standalone invocation. A/B measured
+today — =make test-unit </dev/null= exits 0, the same target with an open pipe
+on stdin hangs on all five. The note above guessed at "a larger timeout or a
+real hang in the script" and it was neither.
+
+Fixed by pinning =stdin=subprocess.DEVNULL= in =run_script=. Verified both ways:
+the previously-failing open-pipe case and the redirected case both pass in
+0.08s, and a full =make test-unit= under a live pipe is clean across 50 suites.
+** DONE [#A] powerprofilesctl crashes on a loop since ppd was masked :bug:velox:dotfiles:
+CLOSED: [2026-08-21 Fri]
+:PROPERTIES:
+:CREATED: [2026-08-17 Mon]
+:LAST_REVIEWED: 2026-08-17
+:END:
+Something polls power state every 10-30 seconds, and each poll runs
+=powerprofilesctl get=, which SIGABRTs. 47 coredumps on velox on 2026-08-17
+alone, the earliest at 08:34, four in one minute while I was watching.
+
+Cause is the 2026-08-16 fix that masked =power-profiles-daemon= so TLP
+survives on laptops. That fix is right and stays. What it did not account for
+is the settings module's power backing
+(=~/.dotfiles/settings/src/settings/power.py=), which shells out to
+=powerprofilesctl=. Against a masked unit the D-Bus activation fails with
+=NameHasNoOwner ... unit is masked=, and the caller aborts rather than
+degrading.
+
+Run by hand the same command exits 0 and prints the error, so the abort is
+context-dependent and the caller needs finding before the fix is written.
+Ratio does not mask ppd, which is why this is velox-only and why it appeared
+the day after the masking.
+
+Costs: journal spam, coredump disk churn, and repeated failed D-Bus
+activations on a travelling laptop's battery. It is also the leading suspect
+for the wedged user manager filed below.
+
+Fix shape: =power.py= should treat a masked or unavailable ppd as a
+first-class "no profile control here" state rather than an error path, and
+the poller should stop retrying a unit it has been told is masked. The
+machine-level half is already correct.
+
+Grading: Major severity (a crash loop burning battery and filling the
+journal, silently) x every user every time on any laptop with the TLP fix
+applied = P1 = [#A].
+
+*** 2026-08-17 Mon @ 19:57:42 -0700 The loop stopped at the reboot; the defect did not
+velox rebooted at 16:04 and there have been zero coredumps since, against 47
+in the twelve hours before it. So the loop is not currently burning anything.
+
+That is not a fix, and the distinction matters for whoever picks this up.
+=powerprofilesctl get= still fails exactly as recorded — =NameHasNoOwner ...
+unit is masked= — so every precondition for the loop is intact and it returns
+whenever the caller next polls. What the reboot cleared is the caller's state,
+not the bug.
+
+Narrowed the search the body asks for: =power.py= is the *only* file in
+dotfiles that shells out to =powerprofilesctl= (=SETTINGS_POWERPROFILESCTL=,
+line 14), so the caller is inside the settings module rather than waybar or a
+timer. Worth knowing that the coredumps are =powerprofilesctl= itself aborting
+— it is a python script, which is why they log as =/usr/bin/python3.14=
+SIGABRT rather than under its own name.
+
+Grade unchanged. The matrix inputs did not move: the severity is what happens
+while the machine is in that state, and the frequency row is every laptop
+carrying the TLP fix. A quiet interval since a reboot is not a frequency
+change.
+
+Fixed in dotfiles =e89d9db=. The caller was =waybar.py=, using =panel.read_state()= (the full snapshot of every control) to read one boolean, four bar modules deep on a 2-second interval. Two fixes, each needed alone: =panel.read_control()= reads a single control's backing, and =power.masked()= checks the mask symlink before shelling out. Verified with a logging stub: full snapshot unmasked calls powerprofilesctl once, masked calls it zero, and a waybar poll calls it zero even unmasked.
+** DONE [#A] The installer clones my two working repos shallow and read-only :bug:velox:
+CLOSED: [2026-08-21 Fri]
+:PROPERTIES:
+:CREATED: [2026-08-17 Mon]
+:LAST_REVIEWED: 2026-08-17
+:END:
+=archsetup:1432= clones the user's archsetup repo and =archsetup:1445= clones
+dotfiles, both with =--depth 1=. Those are not build directories. They are the
+two repos I actively develop in, and on velox they came back from the
+2026-08-13 rebuild with 7 commits of history each instead of 851.
+
+Found 2026-08-17, and found the worst way: I ran the credential-file history
+check that the GitHub-release task asks for, and it reported all five files
+absent from history with a clean exit. The real answer is that this clone
+cannot see the history those files live in. A shallow clone does not error on
+=git log -- <path>=, it answers "no commits" — so a security question came back
+falsely clean, and nothing about the output said otherwise.
+
+Everything else it breaks is quieter: =git log=, =blame=, =bisect=, and any
+archaeology past the boundary. The tree looks completely normal, which is why
+this survived four days on the machine.
+
+The right shape is already in the codebase. =scripts/post-install.sh:42-51=
+takes depth as a per-repo argument and defaults to a full clone, so wallpaper
+gets =--depth 1= and org does not. The AUR build clones (=archsetup:855=,
+=:1673=, =:1677=) are correctly shallow and stay that way. Only the two
+user-repo sites change.
+
+*Second defect, same two lines, found 2026-08-17 while pushing:* the dotfiles
+clone could not push at all. =archsetup:245= defaults =dotfiles_repo= to
+=https://git.cjennings.net/dotfiles.git=, the public read-only endpoint, so
+=git push= returned 403. Ratio uses =git@cjennings.net:dotfiles.git= and
+archsetup's own clone uses the matching ssh form, so velox was the odd one out
+purely because it was the machine rebuilt by the installer. Repointed velox's
+remote and pushed.
+
+That half needs a decision rather than a fix, which is why this task is no
+longer =:solo:=. The https default is *correct for a stranger* installing
+archsetup, who has no ssh key on the server, and this repo is being prepared
+for public release. It is wrong for my own machines, which need to push. The
+override already exists (=DOTFILES_REPO=, documented in
+=archsetup.conf.example=), so the question is only where my personal value
+lives: a config the personal ISO bakes in, a post-install step, or a detection
+that prefers ssh when a key is present. Craig's call.
+
+*Decided 2026-08-19: the ISO bakes the value, and a check nets the rest.*
+=archsetup:240= has the identical default for =archsetup_repo=, so this was
+always two repos rather than one. I ruled out detection — archsetup never
+restores =~/.ssh=, so key-presence at clone time depends on ordering it
+doesn't control, and "any key means ssh" would break a stranger who has an
+unrelated one. I ruled out a bare post-install step for the reason this whole
+class of bug exists: manual steps don't get run, which is why this sat four
+days. So the personal ISO carries =ARCHSETUP_REPO= / =DOTFILES_REPO= in the
+ssh form (noted on the secrets/ISO task), and =post-rebuild-check= check 8
+flags any working repo still on the read-only endpoint — covering curl|bash
+and stock-ISO installs, which the ISO value cannot reach.
+
+Repair on a machine already built: =git fetch --unshallow= in each repo, and
+=git remote set-url origin git@cjennings.net:<repo>.git= for dotfiles.
+
+Grading: Major severity (two working repos silently missing their history on
+the machine I develop on, and it returns confidently wrong answers to history
+questions rather than failing) x every user every time (every fresh install,
+both daily drivers) = P1 = [#A].
+
+Not :solo:. The depth half is (two lines plus tests in the existing
+=tests/installer-steps/= shape, verifiable by asserting the clone command
+carries no =--depth= for these two repos). The remote-URL half needs the
+decision above, so the task as a whole waits on it. Split it in two if the
+depth fix is wanted sooner.
+*** 2026-08-19 Wed @ 23:05:00 -0700 Dropped --depth from both user-repo clones
+=archsetup:1462= and =:1475= now clone full history;
+=tests/installer-steps/test_clone_user_repos.py= covers it with 8 cases, and
+one of them asserts the AUR build clones still carry =--depth 1= so the fix
+can't be over-applied by a careless repo-wide sed. Both my repos on velox were
+already unshallowed by hand last session, so this is prevention rather than
+repair.
+*** 2026-08-19 Wed @ 23:05:00 -0700 Settled the remote-URL half and netted it
+See the decision recorded above. The ISO half is a note on the secrets/ISO
+task; the net is =post-rebuild-check= check 8, which ships now.
+
+Both halves resolved. Depth: =a028aa5= drops =--depth 1= from both user-repo clones, with 8 tests including one asserting the AUR build clones stay shallow. Remote URL: decided 2026-08-19 (see above) — the personal ISO carries the ssh form, and =post-rebuild-check= check 8 (=87ff0b7=) flags any working repo still on the read-only endpoint, covering the install paths the ISO cannot reach.
+** DONE [#D] Worldclock tooltip blanks on one bad timezone row :bug:dotfiles:waybar:quick:solo:
+CLOSED: [2026-08-21 Fri]
+:PROPERTIES:
+:LAST_REVIEWED: 2026-07-25
+:END:
+Found by sentry (2026-07-25), verified by exercising. =hyprland/.local/bin/waybar-worldclock= builds each zone with =ZoneInfo(tz)= inside the loop (line ~99) with no guard, so a single malformed timezone row in =worldclock.conf= raises =ZoneInfoNotFoundError= and crashes the whole python pass. The tooltip then renders empty and *every* zone is lost, not just the bad row; the traceback only reaches stderr, where waybar never surfaces it.
+Repro: a conf with =America/Chicago|Home=, =Not/AZone|Bad=, =Europe/London|London= renders =tooltip: ""= (Home and London gone too).
+Grade: minor severity (one module's tooltip blanks, no data loss) x rare edge case (a malformed conf row) = P4 = [#D].
+Fix: wrap the per-row =ZoneInfo=/=datetime= in a try/except and =continue=, so a typo drops only that row and the valid zones still render. Solo + quick: the script already has an env-override test harness (=WAYBAR_TIME_EPOCH=, =WAYBAR_WORLDCLOCK_CONF=), so a red-first test is cheap.
+
+Fixed in dotfiles =8f692f5=. The per-row =ZoneInfo= is guarded, so a malformed row drops itself and the valid zones still render. Five cases, including a bad row first — the ordering that looks least like one typo and most like the module being broken. Caught the broad =except Exception= rather than =ZoneInfoNotFoundError=, because the row also parses floats and calls strftime and the contract wanted is "a bad row costs only itself".
+** DONE [#C] obsbot-wb-guard polls forever on machines with no OBSBOT :bug:dotfiles:quick:solo:
+CLOSED: [2026-08-21 Fri]
+:PROPERTIES:
+:LAST_REVIEWED: 2026-08-16
+:END:
+=obsbot-wb-guard.service= is =WantedBy=graphical-session.target= and lives in the shared =common/= stow tier, so it starts on every machine. Its main path is =while :; do check_once; sleep 2; done=, and =check_once= returns early when the camera node is absent. On a machine with no OBSBOT attached that is a process waking every two seconds forever to do nothing, which on a laptop is battery spend for zero benefit. No restart loop, though: the loop never exits, so =Restart=on-failure= never fires.
+
+Found 2026-08-16 on velox, after enabling it to match ratio and then having to disable it again by hand. A per-machine disable is the wrong shape, because it drifts velox from ratio permanently and a re-stow or a future audit will just put it back.
+
+Fix: give the unit =ConditionPathExists= on the camera node (=/dev/v4l/by-id/usb-Remo_Tech_Co.__Ltd._OBSBOT_PW106-video-index0=, the same default the script uses) so systemd skips it on any machine without the camera and starts it normally on ratio. Then re-enable it on velox, where it will simply be skipped. Note the limit: a camera plugged in later will not start it until the next login, which is the right trade against a permanent poll.
+
+Careful when disabling by hand in the meantime: =systemctl --user disable= on a *linked* unit deletes the unit symlink, and that symlink is stow-managed, so a bare disable silently removes a file from the dotfiles stow tree. Restore the link afterward or re-stow.
+
+Grade: minor severity (wasted wakeups and battery, no data loss, no failure) x every boot on any machine without the camera = P3 = [#C].
+
+Solo: buildable here (archsetup owns dotfiles end-to-end), verifiable by the agent (assert the unit is skipped on velox and still active on ratio), and no design call left open.
+
+Fixed in dotfiles =566dd14=. =ConditionPathExists= on the camera node, so systemd skips the unit where the camera is absent. velox is now =enabled= like ratio and reports =ConditionResult=no=; the stow symlink is untouched. Found while doing it: ratio has a Logitech BRIO and no OBSBOT on USB at all, so the 2-second poll was pointless on the desktop too, not merely costing laptop battery. A test asserts the unit's condition path and the script's =OBSBOT_WB_DEVICE= default stay equal, since drift there is invisible in both directions.
+** DONE [#C] Spine face tests decay against the wall clock :bug:test:dotfiles:solo:
+CLOSED: [2026-08-21 Fri]
+:PROPERTIES:
+:LAST_REVIEWED: 2026-08-02
+:END:
+=settings/faces/timeline-face-spine.test.mjs= has thirteen =SP.spineRows(g, h)= calls that omit the third argument, so =ref= falls back to its =new Date()= default while the file's events fixture is pinned to =JUL= (2026-07-31 18:30 UTC). Any assertion that depends on how much room the day needs is then measured against today's clock, and rots as the fixture recedes.
+
+One of them, "spacing is uniform everywhere except the gap home opens", had already rotted: green on 07-31 because that was the fixture's own date, red by 08-02. Fixed in place on 2026-08-02 by pinning =JUL=; the remaining thirteen pass today by luck. The measurement, for whoever picks this up — with =ref=now= the even step is 85.21 and home's gaps are 129.10 / 65.40 (the lower one collapses below a plain gap); with =ref=JUL= the step is 78.54 and the gaps are 129.10 / 145.46. Only the lower gap moves, because =up= does not depend on events and =down= does.
+
+Six other calls in the same file already pass =JUL= explicitly, so the convention exists and this is a miss, not a gap in the design. Fix: pass =JUL= at every call whose assertion reads geometry. Leave the call around line 747 alone — it sweeps =new Date(t0)= deliberately.
+
+Grade: minor severity (dev-facing only; no product behavior is wrong, the face itself is fine) x some users, sometimes (each call rots independently, whenever the fixture drifts far enough) = P3 = [#C]. Not merely cosmetic though: a suite that goes red for no real reason is how a genuine regression gets waved through.
+
+Solo — mechanical, an existing convention to copy, and verifiable by running the suite plus re-running it under a faked clock to prove the determinism actually holds.
+
+
+Fixed in dotfiles =c96a216=. All thirteen bare calls now pass =JUL=. The task's "line 747" was stale (the deliberate =t0= sweep is at 893 and already passed its own ref, so it was never at risk), and the continuation-form call closes its arguments on the next line, which is why a naive grep counts fourteen. Added a guard that reads the file and fails with the offending line numbers, and verified it bites by stripping =JUL= from one call and confirming it went red naming that line.
+** DONE [#A] Velox still carries the install placeholder passwords :bug:security:velox:
+CLOSED: [2026-08-23 Sun] SCHEDULED: <2026-08-20 Thu>
+:PROPERTIES:
+:CREATED: [2026-08-20 Thu]
+:LAST_REVIEWED: 2026-08-20
+:END:
+Closed 2026-08-23: I'd already rotated all three on the 08-14 bringup day, so
+this task was never live. Verified on velox before closing — =chage -l= puts the
+last password change for both =cjennings= and =root= at Aug 14 2026, and
+=/etc/zfs/zroot.key= was rewritten 2026-08-14 05:29 and no longer holds the
+placeholder (checked with a =grep -qx= that returns a yes/no without reading the
+key into a transcript).
+
+The premise below was wrong, and it's worth naming how. Nothing ever tested the
+credentials: the claim came from an unticked runbook item plus the archangel
+session handing back the values the *installer* had set, which reads as "these
+are current" only if you assume nobody changed them in between. An inference
+about a security exposure got recorded in the same voice as a measurement. The
+one command that settles it costs a second.
+
+Original body follows.
+
+The 2026-08-13 reinstall set placeholder credentials and the runbook's Phase 5
+item to replace them (=passwd=, =zfs change-key zroot=) was never ticked.
+Believed still live 2026-08-20 via the archangel handoff, which had to hand
+them back to Craig to get into the machine: =welcome1= for the pool, =welcome=
+for the accounts.
+
+So velox's full-disk encryption is currently protected by a dictionary word
+with a digit, on the machine that travels. Anyone who picks it up owns the pool
+and every account on it — the encryption is doing no work at all.
+
+Two commands, both on velox:
+- =passwd= for each account.
+- =zfs change-key zroot= for the pool passphrase. Note this is the ZBM unlock
+ passphrase, so get it right before rebooting.
+
+Grading: *severity-alone carve-out* — this is a security exposure, so the
+frequency row does not discount it (=todo-format.md=). Critical severity: total
+compromise of an encrypted-at-rest laptop from a guessable string, with the
+device leaving the house. = P1 = [#A].
+
+Distinct from the =VERIFY [#A] Rotate the credentials exposed by the 2026-08-09
+dotfiles leak= under the cgit audit — that one covers credentials a crawler
+already took from a public repo. This one is a local default never changed. Both
+are rotation work; neither substitutes for the other.
+** CANCELLED [#B] Consistent keybinding family for the panel console :feature:hyprland:
+CLOSED: [2026-08-21 Fri]
+:PROPERTIES:
+:LAST_REVIEWED: 2026-07-09
+:END:
+Merged into =[#B] Reconcile panel keybindings around Super+N=, which now carries
+this body's detail: the collision list, the velox plain-keyboard constraint, and
+maintenance-M as the priority chord. Cancelled rather than done — the work is
+still open, just tracked in one place instead of two.
+
+Consider putting every panel (net, bluetooth, audio, timer, and the coming maintenance console) on one consistent chord family — a shared modifier set (Super+Shift, Control+Alt, or similar) plus a mnemonic letter per panel (N/B/A/T/M). Today the panels open via waybar clicks only; a uniform chord family makes them keyboard-reachable and predictable. Watch for collisions with existing binds: Super+Shift+A is already PTT toggle, and the hold-to-talk grave bind is load-bearing. Decide the family, audit current hyprland binds for conflicts, wire via the dotfiles hyprland config, and document in the keybind reference. Both machines (velox can't QMK-remap, so chords must work on a plain laptop keyboard).
+*** 2026-07-14 Tue @ 00:31:36 -0500 Folded Craig's ask for a maintenance-panel keybinding; bumped [#C] → [#B]
+Craig asked (in session, 2026-07-14) for a maintenance keybinding specifically — the panel he's reaching for without one. Maintenance (M) is the priority chord when this task gets worked. The capture graduated the task from parking lot to active backlog.
+** DONE [#B] Waybar network module — custom/net :feature:waybar:network:
+CLOSED: [2026-08-21 Fri]
+:PROPERTIES:
+:LAST_REVIEWED: 2026-07-09
+:END:
+Closed 2026-08-21: the module shipped and is in daily use. Phases 1-4 all landed
+in dotfiles, and the tunnels track absorbed most of what Phase 5 originally
+covered. The one piece genuinely left — the =net vpn= CLI subcommand — is now its
+own task below, so the residual is tracked at its real size instead of holding a
+finished umbrella open.
+Unifies the old wifi-no-internet indicator (was =[#C]=) and the network-manager
+dropdown (was =[#B]=) into one =custom/net= module: a tested Python =net= engine
+(nmcli + diagnostics), a thin bar indicator, and a GTK4 layer-shell panel. Code
+lives in the dotfiles repo (hyprland tier + a =net/= package like pocketbook);
+archsetup only installs deps. Secrets stay in NetworkManager's own store (no
+separate credential store). The =captive= script becomes the diagnostics engine.
+Full design, acceptance criteria, and the failure-mode coverage table:
+[[file:docs/design/2026-06-29-waybar-network-module-spec.org][2026-06-29-waybar-network-module-spec.org]].
+
+Phases below, dependency order. Engine/unit work is agent-verifiable (=unittest=
++ fakes on PATH, coverage via venv); the live-network and visual states need real
+conditions, filed under "Manual testing and validation".
+
+*** 2026-06-29 Mon @ 20:19:11 -0400 Phase 1 shipped — indicator + console recovery
+Shipped to the dotfiles repo (10 commits, =5254bd8=..=c095a22=, pushed to main).
+The =net= engine is a src-layout Python package in-tree, imported by a bin shim
+that resolves the stow symlink back to the repo — so it runs from a bare TTY with
+no install, which the recovery path depends on.
+
+Landed: =net status= (fast path, one nmcli call + sysfs, degraded fallback in
+budget) + =net probe= (native captive probe, single-flight flock, atomic cache,
+fresh/stale/expired/unknown classes, iface/SSID/UUID invalidation); =waybar-net=
+replacing =custom/netspeed=, throughput → tooltip, CSS states in both themes +
+live; =net diagnose= (read-only steps) + =net repair= (rfkill/reset/bounce/
+dns-test, cleanup-verified) + =net doctor [--fix]= with the four terminal
+classifications; =net portal= + the =captive --probe-json= refactor; redacted
+JSONL event log; Makefile recovery targets (=make online= etc.); =~/.config/net/
+config=. Verified live: =make net-status= reads the real wlp170s0 / @Hyatt_WiFi.
+
+Airplane (Craig's call, option 1): =custom/net= absorbs only the *display* — net
+reads the airplane-mode state file and shows an airplane state/glyph. The
+airplane-mode toggle stays (it's a low-power mode — radios + CPU + brightness +
+services — not a radio switch), now on =custom/net='s right-click + signal 15.
+Deleted: =waybar-airplane=, =waybar-netspeed=, =custom/airplane=, their tests +
+css. =airplane-mode= kept.
+
+Tests: 160 in =tests/net/= (fake nmcli/curl/rfkill/resolvectl/ping/getent/
+systemctl on a temp PATH; doctor-classification fixtures; degraded-under-slow-
+nmcli benchmark) + the =captive= probe-mode tests; full dotfiles suite green (32
+suites). Coverage-gap pass via throwaway venv: pure modules ≥90% branch
+(classify 100%), IO-error branches excused in the test docstring.
+Deferred to Phase 2/3: archsetup deps (gtk4-layer-shell/python-gobject Phase 2,
+speedtest-go-bin Phase 3 — not added before the code that needs them).
+Verify (manual, live): see Manual testing and validation.
+
+*** 2026-06-29 Mon @ 22:19:25 -0400 Phase 2 shipped — panel shell + connection management
+Shipped to dotfiles (commits =4e7740f=..=24bcac5=, pushed). Engine: =net list= (saved
+MRU + in-range wifi scan, infrastructure types filtered), =net up/down= (UUID-keyed,
+mutation safety — keep prior link until target activates, classify wrong-password vs
+generic, report auto-reactivation), =net add/edit/remove/rescan= (open + WPA-PSK;
+enterprise activate-only; secret to NM's store, never our JSON/log — tested).
+
+Panel: a GTK-free PanelModel (selection, four state machines, the UX-flow enable
+rules, terminal states) + a GTK4 gtk4-layer-shell window (=net panel=) anchored
+top-right under the bar — Connections section with MRU list, active marked, signal
+glyph, row-click select, Connect/Add/Forget/Rescan, confirm-on-forget, worker-thread
+engine calls via GLib.idle_add. GTK imported lazily so the CLI/tests stay GTK-free.
+
+Bar interactions (settled with Craig over live iteration): left = =net-panel= toggle,
+middle = =net portal=, right = =net-fix= (notify the doctor result when one-way; open
+a terminal only when the outcome is fixable — the sudo/interactive case). Airplane on
+Super+Shift+A. archsetup adds =gtk4-layer-shell= + =python-gobject= (this commit);
+already on velox.
+
+Tests: 204 in tests/net (merge ordering/dedup, up/down mutation safety, no-secret-leak
+on add/edit, panel model + state machines, gui row-format helpers). Full dotfiles suite
+green (32 suites). Live-verified on velox: panel opens/toggles, list shows real 24
+profiles, right-click notification delivers (Craig confirmed). Phase 3 (diagnose/repair/
+speedtest IN the panel) is next; the engine for it already exists from Phase 1.
+
+*** 2026-06-29 Mon @ 22:43:40 -0400 Phase 3 shipped — diagnostics + speed test in the panel
+Shipped to dotfiles (=91277cf=..=691abcb=) + archsetup (=48052d6=, speedtest-go-bin),
+pushed. Engine: =net speedtest= (parses speedtest-go --json → ping from latency ns,
+down/up from per-server byte rates; missing-backend / offline / malformed → error
+envelope per the failure table). Panel grew a section switcher with four pages:
+- Connections (Phase 2).
+- Diagnose: =net diagnose= on a worker thread, each step a row (✓/✗/… glyph + title +
+ redacted evidence), read-only; Open-portal button when captive.
+- Repair: "Get me online" (=net doctor --fix=) + tiers (rfkill/reset/bounce/dns-test)
+ + force portal. Confirmations in-panel with the spec's exact wording; the privileged
+ tiers run via =net-popup= terminal (where the sudo prompt + step output, incl.
+ cleanup-verified, show) — a panel has no tty, and pkexec would mean a prompt per op.
+- Speed test: in-process =net speedtest= (no privilege → inline result: ↓/↑ Mbps + ping
+ + server), Run/Cancel (Cancel pkills the child), error envelope shown.
+
+213 net tests; pure helpers (step_indicator, format_speedtest) unit-tested. Full
+dotfiles suite green (32 suites). One unverified assumption: speedtest-go's dl/ul unit
+(taken as bytes/s; =BYTES_PER_SEC= flips it) — needs one real run vs a reference. The
+in-panel repair streaming (vs terminal) is a named future polish once the GUI-privilege
+story settles.
+
+The waybar network module ([#B] parent) is now COMPLETE through Phase 3. Phase 4
+(in-app help + user guide) and Phase 5 (VPN/WireGuard) remain as future work; the core
+feature (indicator + recovery + panel + diagnostics + speed test) is done.
+Verify (manual, live): see Manual testing and validation.
+
+*** 2026-07-09 Thu @ 16:32:54 -0500 Audit reconcile: Phase 4 is filed on the dotfiles side, waiting on them
+The dotfiles project accepted the Phase 4 handoff and filed it as a =[#C]= task in their own =todo.org= (their note, 2026-07-08 16:56): the help-text audit + panel help affordance, the user-guide/README, and the ratio rollout doc. Not started there. They ping when it lands, and this task's Phase 4 child closes then. Nothing to do here meanwhile.
+
+*** 2026-08-17 Mon @ 19:57:42 -0700 Landed on the dotfiles side; the block is cleared
+dotfiles shipped it as =138da7b= and closed its own task, so this one closes
+with it and the =:blocked:= tag comes off. Found by checking their =todo.org=
+rather than waiting for the ping — their close-out note says "archsetup pinged
+so its Phase 4 task can close", so the handoff worked and only this end was
+left open.
+
+All three acceptance criteria are met on their side: the help audit found and
+fixed a stale =net repair= action list (nine of nineteen actions were named;
+both the CLI help and =repair.py='s docstring now generate from the ACTIONS
+registry), =net/README.md= covers every command plus the recovery targets, and
+the ratio rollout is documented with both daily drivers verified current.
+
+They split the panel help affordance out rather than inventing it — no sibling
+panel has one, so its shape is a design call. It is tracked on their side, not
+here.
+
+Original deliverable, for the record: in-app help (=net --help= + per-command,
+panel help affordance); README/user-guide; archsetup Hyprland dep install
+(=gtk4-layer-shell=, =python-gobject=, =speedtest-go-bin=); ratio manual dep +
+stow step. Handed off 2026-07-04 with the archsetup deps already confirmed
+installed.
+
+*** 2026-08-21 Fri @ 14:18:03 -0700 Promoted the Phase 5 residual out to its own task
+Rescoped 2026-07-04 (audit): the tunnels track already shipped most of the original Phase 5. Panel tunnel bring-up/down and detection landed (dotfiles 2d9d060 probes tailscale/NM-wireguard/Proton; 21db05a brings overlays up/down from the panel's Tunnels sub-view; 31ba056 diagnose/doctor understand tunnel routes; archsetup 2e40781 wireguard config import; the net-panel-other-interfaces spec is IMPLEMENTED). What remains for Phase 5 is only the =net vpn ...= CLI subcommand — cli.py still has no vpn/tunnel parser. Fold the panel's existing tunnel operations into a CLI surface; spec separately when picked up.
+** DONE [#A] Ratio: pull .emacs.d before upgrading Emacs to 31.1 :chore:ratio:emacs:
+CLOSED: [2026-08-25 Tue]
+:PROPERTIES:
+:CREATED: [2026-08-25 Tue]
+:LAST_REVIEWED: 2026-08-25
+:END:
+Emacs 31.1's warnings.el defers daemon-startup warnings into a closure holding
+the =*Warnings*= buffer; the config's dashboard-only sweep killed that buffer,
+so the first client frame of every fresh 31.1 daemon failed on Wayland and
+emacsclient silently fell back to =$DISPLAY= (XWayland, pgtk warning dialog).
+Fixed in =.emacs.d= commit =63831060= (2026-08-25, velox verified live:
+=GdkWaylandDisplay=). Ratio is still on 30.2, which lacks the deferring code,
+so it is fine until it upgrades — then it hits the same trap once per daemon
+start unless the fix is pulled first.
+
+Order on ratio: =git -C ~/.emacs.d pull= (the push from velox is the telega
+session's; confirm =63831060= is on origin first), then the =pacman -Syu= that
+brings =emacs-wayland 31.1=, then restart the daemon. Check afterwards:
+=emacsclient -e '(pgtk-backend-display-class)'= → =GdkWaylandDisplay=.
+
+*** 2026-08-25 18:10 — pull already landed; the upgrade half remains
+Checked ratio over tailscale: =~/.emacs.d= is clean at =91fbac72= (= =origin/main=),
+and =63831060= is an ancestor of HEAD — =modules/undead-buffers.el= carries the
+=*Warnings*= entry. Ratio is on =emacs-wayland 30.2-3= with =31.1-1= pending among
+720 updates (last full upgrade 2026-08-01; kernel 7.1.5 → 7.1.9 also pending,
+btrfs root, uptime 3.5 weeks). The daemon is a plain =emacs --daemon= (not a
+user unit) holding 2 live frames, so the restart step will drop those frames.
+What remains: the =pacman -Syu= on ratio, the daemon restart, and the
+=(pgtk-backend-display-class)= check.
+
+*** 2026-08-25 Tue @ 18:35:00 -0600 Upgraded ratio to Emacs 31.1 and verified the Wayland backend
+Ran the upgrade over tailscale as a transient unit (=ratio-upgrade.service=,
+log at =/var/log/ratio-upgrade.log=): 714 packages, =--ignore= on the six
+packages the live-update guard would have blocked (aquamarine, hyprland,
+hyprutils, mesa, vulkan-radeon, wayland — still pending, apply from a TTY
+before the reboot). One orphan cleared first: =qemu-block-gluster= had been
+dropped from the repo and pinned =qemu-common=; the new =qemu-full= no
+longer needs it. Killed the plain =emacs --daemon= (no modified buffers, no
+graphical frames), started =emacs.service= instead so the daemon carries the
+systemd user environment, and probed from a throwaway frame:
+=(pgtk-backend-display-class)= → =GdkWaylandDisplay=, =*Warnings*= alive.
+Ratio still wants a reboot for =linux 7.1.9=. Pacnews to review there:
+=/etc/ssh/sshd_config.pacnew= and two =/etc/tpm2-tss/fapi-profiles/*.json=.
diff --git a/todo.org b/todo.org
index 40aad3c..0f769e9 100644
--- a/todo.org
+++ b/todo.org
@@ -1437,28 +1437,6 @@ Related: =[#B] Caffeine state is unreadable on both surfaces= covers display acc
Launched by hand afterward it runs fine and survives, so gammastep is not broken — it loses a race against compositor readiness at session start. Nothing relaunches it, so night light is simply off for the whole session, silently. (An earlier read of this said night light "has likely never worked from the config". That was wrong: the failure is a startup race, not a permanent break.)
Worth its own task — the fix is a readiness wait or a retry around that exec-once, not a persistence change. Filed here for now because it surfaced during this investigation.
-** DONE [#A] Pre-vacation fix list — morning review
-CLOSED: [2026-09-17 Thu]
-:PROPERTIES:
-:LAST_REVIEWED: 2026-09-17
-:END:
-Reviewed 2026-08-08; the decisions it drew are recorded in the tasks below.
-Closed at the 2026-09-17 review because the ~08-15 departure it ranked work for
-has passed. Where each of the seven items went:
-
-1. Velox reliability → the [#A] sleep/suspend task.
-2. Velox machine health for travel → moot: velox was wiped and reinstalled on
- 2026-08-13.
-3. Remote access from outside the LAN → the wolf WireGuard rider under the
- sleep/suspend task. Tailscale to ratio works off-LAN (checked 2026-09-17);
- truenas and truenas-kvm weren't re-checked.
-4. cgit secrets audit → the [#A] audit task and its rotation VERIFY.
-5. Osbot camera → its own task; the podman socket and camera udev rule are
- installed (sleep/suspend task, 08-17 entry).
-6. Hotspot/metered WiFi and network ordering → the held design calls under
- Next Session Focus.
-7. The orchestrator sequence-pin gap → filed 2026-09-17 as [#C] Orchestrator
- sequence pin misses an added step.
** TODO [#C] Re-apply the active program at session start :refactor:dotfiles:hyprland:
:PROPERTIES:
:CREATED: [2026-07-30 Thu]
@@ -3920,760 +3898,6 @@ problem is separate and unchanged.
repair_tunnel_dot_off (dotfiles net/src/net/repair.py) puts a tunnel link back to its DoT mode when turning DoT off didn't bring names back, and the evidence says "put back to <mode>" only when that restore succeeded. The restore-declined branch has no direct test. Give fake-resolvectl a second failure switch (FAKE_RESOLVECTL_DOT_RESTORE_FAIL) so the wording can be asserted absent as well as present. Follow-up from the f56fd1a review.
* Archsetup Resolved
-** DONE [#A] Reseat velox input-cover ribbon — phantom power button :bug:velox:hardware:
-CLOSED: [2026-08-26 Wed] DEADLINE: <2026-08-26 Wed>
-:PROPERTIES:
-:CREATED: [2026-08-13 Thu]
-:LAST_REVIEWED: 2026-08-13
-:END:
-Machine off, lift the input cover (Framework QR-guided procedure, 5
-fasteners), reseat its ribbon connector to the mainboard — disturbed in the
-2026-08-13 board swap. Root cause of every "mystery reboot" that day:
-chassis flex (flash-drive touch, ethernet bump, lid partially lowered)
-fired phantom power-button presses — journalctl -b -1 showed "Power key
-pressed short." → orderly logind poweroff, then the glitching button
-powered it back on. While in there, reseat the USB expansion cards too —
-the flaky slot (two hard resets, one no-enumeration) is likely the same
-flex problem.
-THIRD SYMPTOM (2026-08-13 evening): touchpad delivers ZERO input events —
-15s synchronized libinput debug-events capture while swiping caught
-nothing, though i2c enumeration and a driver rebind handshake are clean.
-Signature of a dead interrupt line on the same ribbon. Keyboard + power
-LED lines work; BT mouse is the interim pointer.
-ESCALATED 2026-08-13 21:00: a fourth event killed the machine THROUGH the
-shield. Previous boot's journal ends mid-line (tailscaled chatter) with no
-shutdown sequence at all — a hard power cut, not logind acting. So the
-glitch now reaches the EC/hardware power path, which no software setting
-can intercept. The reseat is the only fix, and this is a
-lose-work-without-warning failure mode, not an inconvenience.
-Interim shield (already live): /etc/systemd/logind.conf.d/powerkey.conf
-sets HandlePowerKey=ignore — phantom presses log but do nothing; EC-level
-10s hold still force-cuts. Consider keeping it even after the repair.
-Verify after reseat: flex the chassis edges + partially lower the lid, then
-grep the journal for new "Power key pressed" lines — zero means fixed.
-Must be done before the Sunday flight — a phantom press mid-travel with the
-shield on is survivable, but the connector should not be trusted at 30,000
-feet on the loose setting.
-
-*** 2026-08-19 Wed @ 14:40:00 -0700 Retracted: the RTC reset is not this task's, and I should not have filed it here
-I attributed the 2026-08-19 network outage to this ribbon earlier today. Craig
-pushed back — he reseated it before the trip to get the touchpad working — and
-he is right. The evidence does not support the attribution and some of it points
-the other way.
-
-What actually holds. Boot -3 ended at 01:33:18 with no shutdown sequence: no
-power-off target, no unmounting. The next boot's kernel line reads =rtc_cmos
-00:01: setting system clock to 2025-01-01T00:00:16 UTC=, a firmware default, so
-the RTC was reset rather than drifted. No firmware update was applied
-(=fwupdmgr get-history= is empty) and the battery is fine.
-
-What refutes the ribbon. This boot logged *zero* =Power key pressed= events, and
-so did the four boots before it. The phantom-press symptom had genuinely stopped
-after 08-15, exactly as the 08-16 session recorded. The earlier events logged a
-power-key press and an orderly poweroff; this logged neither, which makes it a
-different signature, not a worse version of the same one.
-
-What I got wrong methodologically: I anchored on the most salient open hardware
-task and read association as evidence. I even wrote "I can't prove it is the
-same connector" and then filed it here anyway, which is the tell.
-
-Two things I checked and can rule out. There were no OOM kills — the 3,433
-matching lines are a systemd unit named "Periodically re-score Claude Code
-processes for the OOM-killer" firing on a timer, not memory pressure, and there
-is not a single "Killed process" line. Thermal is clean; the only mentions are
-boot-time zone registration at 34C and 45C.
-
-One real thing the same window did surface, tracked separately: a python3 crash
-loop, 251 core dumps in the final ten minutes, SIGABRT with =XFreeThreads= and
-=PyEval_RestoreThread= in the trace. It does not explain the RTC, because
-software cannot clear it, but it is its own problem.
-
-The open question that would settle the RTC is for Craig, not the journal: a
-long power-button hold on a Framework triggers an EC-level reset that clears the
-RTC, which fits a wedged machine being forced off. A 4-second hold would not.
-
-*** 2026-08-17 Mon @ 19:57:42 -0700 Not done, and the two symptoms now disagree
-The reseat did not happen before the flight, and velox is travelling. The
-deadline blew past on 08-14.
-
-The two symptoms have separated, which is worth recording because it changes
-what the evidence proves. The phantom presses have stopped: fifteen "Power key
-pressed" entries between 08-14 04:29 and 08-15 20:04, then nothing at all
-across five boots including today's. The touchpad has not — there is still no
-touchpad node under =/dev/input/by-path/=, which is the same dead interrupt
-line the body describes.
-
-So the quiet power button is not evidence the connector reseated itself. The
-interrupt line is the symptom that cannot be masked in software, and it is
-still dead, so the ribbon is still unseated. The most likely reason the
-presses stopped is that the machine has been sitting on hotel surfaces instead
-of being carried and flexed.
-
-The interim shield is still live (=HandlePowerKey=ignore=), and the escalation
-note stands: an EC-level glitch cuts power below systemd regardless of it.
-*** 2026-08-15 Sat @ 23:05:00 -0500 The reseat did happen, and the touchpad came back — this contradicts the 08-17 read
-Recording this because a parallel session concluded on 08-17 that the reseat had
-not happened and the touchpad was still dead. Both halves were done and verified
-that night, so the two accounts disagree and the disagreement should be visible
-rather than silently resolved by whichever session committed last.
-
-What was done: the input-cover ribbon was reseated first, which fixed the
-phantom power button — the 22:09 boot logged zero =Power key pressed= lines
-after Craig flexed the chassis, against nine on the boot before. The touchpad
-did not change, because the input-cover ribbon is not its connector. The 4-pin
-connector beside the printed =TOUCHPAD= label is silkscreened =PIN 1-2 GND /
-PIN 3-4 VCC= — pure power, so it cannot carry i2c or an interrupt. Reseating the
-ribbon that actually crosses to the mainboard fixed it.
-
-Measured, not assumed: the touchpad interrupt (=amd_gpio= pin 8) went from 0
-counts across all 24 CPUs to 1795, and =i2c_hid_acpi ... did not ack reset
-within 1000 ms= disappeared from the boot log. Craig confirmed the pointer moved.
-
-*Why the 08-17 probe likely misread it:* it checked for a node under
-=/dev/input/by-path/=. i2c-HID touchpads frequently get no =by-path= symlink
-even when fully working, so its absence is not evidence of a dead interrupt
-line. The falsifiable check is the interrupt count in =/proc/interrupts= while
-the pad is being touched, or the reset message in =dmesg=.
-
-*Left open rather than closed* — velox was refusing ssh at merge time on 08-20,
-so the current state could not be re-verified, and a later regression cannot be
-ruled out. One second of Craig's time settles it: move the pointer. If it works,
-close this; if it does not, the interrupt line went back down and that is new
-information.
-
-*** 2026-08-26 Wed @ 22:30:46 -0600 Closed: the reseat was done on 08-15 and the task was never marked
-I reseated the ribbon on 2026-08-15 and never closed this. The 08-15 entry
-above already records the verification: zero =Power key pressed= lines on the
-22:09 boot after flexing the chassis, the touchpad interrupt count back up
-once the right connector was reseated. This boot shows zero presses as well.
-The interim shield (=HandlePowerKey=ignore= in
-=/etc/systemd/logind.conf.d/powerkey.conf=) is still live; I'm leaving it in
-place, since a phantom press with it on costs nothing and without it costs
-the session.
-** CANCELLED [#B] agent-text relay reports success for a message that went nowhere :bug:
-CLOSED: [2026-08-19 Wed]
-:PROPERTIES:
-:CREATED: [2026-08-19 Wed]
-:LAST_REVIEWED: 2026-08-19
-:END:
-
-Not a defect. rulesets refuted it with measurements and I reproduced theirs
-before accepting: on velox, whose account store is empty,
-=signal-cli -a +15550000000 send= exits 1 with "User +15550000000 is not
-registered", and =ssh 100.71.182.1 'exit 7'= returns 7, so a non-zero code
-propagates faithfully back through the relay. The loop's
-=[ "$rc" -eq 0 ] && break= therefore advances to the next host exactly as
-intended. signal-cli fails closed.
-
-I filed this off a conditional in their handoff — ".emacs.d raised a case
-neither of you tested ... *if* signal-cli send exits zero against an empty
-account store" — and turned the "if" into a graded [#B] with a =:blocked:= tag
-on another project, without running the one command that settles it. The
-machine that proves it was in front of me the whole time. Their ask is fair and
-I am recording it rather than the outcome alone: verify before filing a defect
-against someone else's work, especially one carrying a blocking tag.
-** DONE [#B] Clock/DNS bootstrap deadlock — recovery needs a second device :bug:velox:
-CLOSED: [2026-08-19 Wed]
-:PROPERTIES:
-:CREATED: [2026-08-19 Wed]
-:LAST_REVIEWED: 2026-08-19
-:END:
-
-The installer wrote both halves of a deadlock. =configure_dns= pins
-=DNSOverTLS=yes= with =DNSSEC=yes=, and both validate against the wall clock;
-the chrony step enables chronyd without writing a config, so the machine runs
-Arch's stock one whose only source is =pool 2.arch.pool.ntp.org= — a hostname.
-Boot with a wrong clock and DoT certificate validation fails, so nothing
-resolves; chrony then cannot resolve its pool, so the clock stays wrong.
-Neither side moves. It caught velox on the road 2026-08-19 and had to be
-diagnosed from a phone.
-
-Fixed at the root: the installer now writes
-=/etc/chrony.d/10-bootstrap-ip-ntp.conf= with two IP-addressed Cloudflare
-sources and points stock chrony.conf at the drop-in. An address needs no DNS
-and carries no certificate, so the escape hatch holds whatever broke the clock.
-velox has the same drop-in applied live, verified with =chronyc -n sources=
-(=162.159.200.1= selected) and =timedatectl= reporting synchronized.
-
-What is left here is the part I could not verify: the decisive test is a full
-power-down and cold boot, confirming the clock corrects itself untouched. See
-the manual-testing entry. Until that runs, the fix is sound by construction
-rather than demonstrated.
-
-Grading: Critical severity (total loss of network — no DNS means no egress, and
-recovery needs a second device) x some users sometimes (only machines that boot
-with a wrong clock, which is any RTC fault, BIOS reset, or drained cell) = P2 =
-[#B]. Graded on the being-in-it, not the getting-into-it: once the machine is in
-this state it is fully offline with no local path out.
-
-*** 2026-08-19 Wed @ 12:25:00 -0700 Reproduced it, and the mechanism was not what either of us said
-I wound velox's clock back 27 days with chronyd stopped and watched it fail.
-Resolution died outright, and plain UDP/53 to 1.1.1.1 kept answering throughout
-— the discriminator the doctor keys on, confirmed live rather than reasoned.
-
-The cause is DNSSEC, not DNS-over-TLS. resolved logged =signature-expired=
-against the root DNSKEY and every DS beneath it. The DoT handshake to
-=1.1.1.1:853= verified clean at that same clock, and the Cloudflare certificate
-runs Dec 2025 to Dec 2026, so it was never outside its window. An RRSIG window
-is days to weeks and a certificate is good for a year, so a skew that breaks
-DNSSEC normally leaves DoT untouched. The phone session blamed the certificate
-and I carried that forward into the first commit; both were wrong.
-
-=DNSSEC=allow-downgrade= does not rescue it either, which matters because it is
-the obvious reach and it is what ratio runs. resolved downgrades when a server
-lacks DNSSEC support, and a signature-window failure is a validation failure, so
-no downgrade fires. Six retries over eighteen seconds plus
-=resolvectl reset-server-features=, all dead. I briefly believed otherwise off a
-test whose success was a cache hit (=Data from: cache network=).
-
-So ratio was exposed after all, and I have given it the same drop-in. Its
-=162.159.200.1= is selected and its clock is synchronized.
-
-The fix itself is verified end to end: with the clock wound back and no DNS at
-all, chronyd reached the IP-addressed source and stepped the clock from
-2026-07-23 straight back to 2026-08-19. That is the whole claim, demonstrated
-rather than argued.
-
-Also settled: the clock landed on 2026-07-23 because that is systemd 261.2's
-build date to the minute (=/usr/lib/systemd/systemd=, 10:43:59), and systemd
-advances a garbage RTC to its own build epoch at boot. Not timesyncd's
-last-good-sync timestamp, which cannot be it — timesyncd is disabled here. That
-also confirms the RTC really was reading earlier than that, so the coin cell
-stays the prime suspect.
-
-*** 2026-08-19 Wed @ 10:12:00 -0700 Root fix, doctor verdict, and taxonomy entry landed
-The installer carries the drop-in; =post-rebuild-check= grew a sixth check that
-fails a machine whose every NTP source is a hostname; the net failure taxonomy
-gained the mode in its DNS layer plus a cluster 5 triage line, and its existing
-egress-layer clock entry now says outright that its remedy does not apply when
-DoT or DNSSEC is on.
-
-The doctor half is in dotfiles: =classify.py= reached "DNS not resolving → net
-repair dns-test" here, which cannot help, because every public resolver fails
-the same clock-sensitive validation — so the doctor sent you round a loop. It
-now emits a =clock-dns= row ahead of the generic DNS verdict. Detection is
-deliberately DNS-free: a local =timedatectl= read for sync state, and a bypass
-query addressed by IP over plain UDP/53 to tell "resolved is refusing to
-validate" apart from "DNS is genuinely dead".
-** DONE [#C] DNSSEC strictness on the travelling laptop :velox:
-CLOSED: [2026-08-19 Wed]
-:PROPERTIES:
-:CREATED: [2026-08-19 Wed]
-:LAST_REVIEWED: 2026-08-19
-:END:
-
-Craig chose =allow-downgrade= everywhere. Applied to velox, ratio, and the
-installer, and ratio's =DNSOverTLS= tightened from =opportunistic= to =yes= in
-the same pass, so all three now agree: encrypted DNS always, validation
-best-effort.
-
-The reasoning that settled it: the deadlock is fixed by the IP-addressed NTP
-source, and =allow-downgrade= was measured not to help with it at all. What
-=allow-downgrade= does buy is the venue-resolver case the taxonomy documents,
-where =yes= turns a resolver that mangles DNSSEC records into no answer at all.
-That is a hotel and airport problem, so it is velox's problem, and the
-encryption is the half worth being strict about.
-** CANCELLED [#C] Branch network policy on laptop vs desktop in the installer :feature:
-CLOSED: [2026-08-19 Wed]
-:PROPERTIES:
-:CREATED: [2026-08-19 Wed]
-:LAST_REVIEWED: 2026-08-19
-:END:
-
-Cancelled because the decision above emptied it. All three motivating cases now
-want the same value on every machine: =DNSSEC=allow-downgrade=, a stable
-per-network wifi MAC, and an IP-addressed NTP source. A branch with nothing to
-put on either side is machinery built for a divergence that does not exist, and
-it would be the kind of scaffolding that rots unread.
-
-Worth keeping the observation, which is the part with a shelf life: when a
-network default does need to differ by machine class, the test already exists.
-=ls /sys/class/power_supply/BAT*= is what =prune_waybar_battery=, the ppd mask,
-and the TLP config all key on. Reopen this then rather than building it now.
-** DONE [#C] Automate the clock/DNS deadlock repair in the net doctor :feature:
-CLOSED: [2026-08-19 Wed]
-:PROPERTIES:
-:CREATED: [2026-08-19 Wed]
-:LAST_REVIEWED: 2026-08-19
-:END:
-
-Shipped as the =clock-ip-ntp= repair, and the verdict is =fixable= rather than
-terminal. Both open questions got answered by driving a real deadlock instead of
-reasoning about it: =chronyc add server= returns =200 OK= against a running
-chronyd, and =makestep= needs a sample to land, so it took four calls and about
-eight seconds rather than working on the first. The repair retries accordingly.
-
-Verified end to end on velox against a genuine deadlock (wrong clock, chronyd
-running with only an unresolvable hostname source, DNS dead): the repair
-corrected the clock in 6.1 seconds and DNS came back.
-
-The live run also caught a defect no unit test would have. The doctor reported
-"Saved password for SpectrumSetup-3C was rejected" — because
-=_recent_auth_failure= greps =journalctl --since -5min=, and a clock weeks off
-windows onto a different incident's entries. It would have sent Craig to
-re-enter a password that was never wrong. Fixed twice over: the journal half is
-now skipped when the clock is untrustworthy, and the clock verdict is ordered
-above the auth verdict, since everything below it reasons over timestamps that
-only mean something once the clock is right. Airplane mode and hard rfkill stay
-above, being physical states the clock has no bearing on.
-** DONE [#D] net-scenarios harness times out under back-to-back suite runs :test:tooling:
-CLOSED: [2026-08-19 Wed]
-:PROPERTIES:
-:LAST_REVIEWED: 2026-07-24
-:END:
-=tests/net-scenarios/test_run_net_scenarios.py= errored on all 5 tests twice during round 12, each time =subprocess.TimeoutExpired= after its 20s budget on =scripts/testing/run-net-scenarios.sh --target root@fake-vm=. Both occurrences were in =make test-unit= runs launched immediately after a previous full run. It then passed 6 runs in a row (3 on a pristine tree, 3 with the round-12 change), and standalone it finishes in 0.09s, so this is not a regression from any code change.
-
-The harness stubs =ssh=, =rsync= and =jq= onto =PATH=, so nothing should touch the network at all — which is what makes a 20s timeout suspicious rather than merely slow. Worth reproducing under load before deciding whether the fix is a larger timeout or a real hang in the script. Evidence logs from the round: =/tmp/tu.log= and =/tmp/tu2.log= (tmpfs, gone after reboot).
-
-Not graded on the bug matrix: it is test infrastructure, not the shipped codebase.
-
-*** 2026-08-08 Sat @ 05:05:00 -0500 Recurred under concurrent load, same signature
-All 5 tests hit the 20s TimeoutExpired again during a =make test-unit= run
-that overlapped two review subagents running their own suites on the box.
-Standalone immediately after: 0.095s, all pass; the following quiet-machine
-full run was clean. Confirms the load-sensitivity read — reproduce under
-deliberate load before choosing between a bigger budget and a real hang.
-
-
-*** 2026-08-19 Wed @ 14:50:00 -0700 Root-caused and fixed: inherited stdin, not load
-Not load, and not the network. The harness stubs ssh as =cat >/dev/null=, which
-drains stdin to EOF. With no explicit stdin the stub inherits whatever the test
-runner had, so it returned instantly when stdin was redirected and blocked
-forever when it was a terminal or a live pipe. All five tests then burned their
-20-second budget.
-
-That is why it looked like a load effect: a run launched immediately after
-another inherited a different stdin than a standalone invocation. A/B measured
-today — =make test-unit </dev/null= exits 0, the same target with an open pipe
-on stdin hangs on all five. The note above guessed at "a larger timeout or a
-real hang in the script" and it was neither.
-
-Fixed by pinning =stdin=subprocess.DEVNULL= in =run_script=. Verified both ways:
-the previously-failing open-pipe case and the redirected case both pass in
-0.08s, and a full =make test-unit= under a live pipe is clean across 50 suites.
-** DONE [#A] powerprofilesctl crashes on a loop since ppd was masked :bug:velox:dotfiles:
-CLOSED: [2026-08-21 Fri]
-:PROPERTIES:
-:CREATED: [2026-08-17 Mon]
-:LAST_REVIEWED: 2026-08-17
-:END:
-Something polls power state every 10-30 seconds, and each poll runs
-=powerprofilesctl get=, which SIGABRTs. 47 coredumps on velox on 2026-08-17
-alone, the earliest at 08:34, four in one minute while I was watching.
-
-Cause is the 2026-08-16 fix that masked =power-profiles-daemon= so TLP
-survives on laptops. That fix is right and stays. What it did not account for
-is the settings module's power backing
-(=~/.dotfiles/settings/src/settings/power.py=), which shells out to
-=powerprofilesctl=. Against a masked unit the D-Bus activation fails with
-=NameHasNoOwner ... unit is masked=, and the caller aborts rather than
-degrading.
-
-Run by hand the same command exits 0 and prints the error, so the abort is
-context-dependent and the caller needs finding before the fix is written.
-Ratio does not mask ppd, which is why this is velox-only and why it appeared
-the day after the masking.
-
-Costs: journal spam, coredump disk churn, and repeated failed D-Bus
-activations on a travelling laptop's battery. It is also the leading suspect
-for the wedged user manager filed below.
-
-Fix shape: =power.py= should treat a masked or unavailable ppd as a
-first-class "no profile control here" state rather than an error path, and
-the poller should stop retrying a unit it has been told is masked. The
-machine-level half is already correct.
-
-Grading: Major severity (a crash loop burning battery and filling the
-journal, silently) x every user every time on any laptop with the TLP fix
-applied = P1 = [#A].
-
-*** 2026-08-17 Mon @ 19:57:42 -0700 The loop stopped at the reboot; the defect did not
-velox rebooted at 16:04 and there have been zero coredumps since, against 47
-in the twelve hours before it. So the loop is not currently burning anything.
-
-That is not a fix, and the distinction matters for whoever picks this up.
-=powerprofilesctl get= still fails exactly as recorded — =NameHasNoOwner ...
-unit is masked= — so every precondition for the loop is intact and it returns
-whenever the caller next polls. What the reboot cleared is the caller's state,
-not the bug.
-
-Narrowed the search the body asks for: =power.py= is the *only* file in
-dotfiles that shells out to =powerprofilesctl= (=SETTINGS_POWERPROFILESCTL=,
-line 14), so the caller is inside the settings module rather than waybar or a
-timer. Worth knowing that the coredumps are =powerprofilesctl= itself aborting
-— it is a python script, which is why they log as =/usr/bin/python3.14=
-SIGABRT rather than under its own name.
-
-Grade unchanged. The matrix inputs did not move: the severity is what happens
-while the machine is in that state, and the frequency row is every laptop
-carrying the TLP fix. A quiet interval since a reboot is not a frequency
-change.
-
-Fixed in dotfiles =e89d9db=. The caller was =waybar.py=, using =panel.read_state()= (the full snapshot of every control) to read one boolean, four bar modules deep on a 2-second interval. Two fixes, each needed alone: =panel.read_control()= reads a single control's backing, and =power.masked()= checks the mask symlink before shelling out. Verified with a logging stub: full snapshot unmasked calls powerprofilesctl once, masked calls it zero, and a waybar poll calls it zero even unmasked.
-** DONE [#A] The installer clones my two working repos shallow and read-only :bug:velox:
-CLOSED: [2026-08-21 Fri]
-:PROPERTIES:
-:CREATED: [2026-08-17 Mon]
-:LAST_REVIEWED: 2026-08-17
-:END:
-=archsetup:1432= clones the user's archsetup repo and =archsetup:1445= clones
-dotfiles, both with =--depth 1=. Those are not build directories. They are the
-two repos I actively develop in, and on velox they came back from the
-2026-08-13 rebuild with 7 commits of history each instead of 851.
-
-Found 2026-08-17, and found the worst way: I ran the credential-file history
-check that the GitHub-release task asks for, and it reported all five files
-absent from history with a clean exit. The real answer is that this clone
-cannot see the history those files live in. A shallow clone does not error on
-=git log -- <path>=, it answers "no commits" — so a security question came back
-falsely clean, and nothing about the output said otherwise.
-
-Everything else it breaks is quieter: =git log=, =blame=, =bisect=, and any
-archaeology past the boundary. The tree looks completely normal, which is why
-this survived four days on the machine.
-
-The right shape is already in the codebase. =scripts/post-install.sh:42-51=
-takes depth as a per-repo argument and defaults to a full clone, so wallpaper
-gets =--depth 1= and org does not. The AUR build clones (=archsetup:855=,
-=:1673=, =:1677=) are correctly shallow and stay that way. Only the two
-user-repo sites change.
-
-*Second defect, same two lines, found 2026-08-17 while pushing:* the dotfiles
-clone could not push at all. =archsetup:245= defaults =dotfiles_repo= to
-=https://git.cjennings.net/dotfiles.git=, the public read-only endpoint, so
-=git push= returned 403. Ratio uses =git@cjennings.net:dotfiles.git= and
-archsetup's own clone uses the matching ssh form, so velox was the odd one out
-purely because it was the machine rebuilt by the installer. Repointed velox's
-remote and pushed.
-
-That half needs a decision rather than a fix, which is why this task is no
-longer =:solo:=. The https default is *correct for a stranger* installing
-archsetup, who has no ssh key on the server, and this repo is being prepared
-for public release. It is wrong for my own machines, which need to push. The
-override already exists (=DOTFILES_REPO=, documented in
-=archsetup.conf.example=), so the question is only where my personal value
-lives: a config the personal ISO bakes in, a post-install step, or a detection
-that prefers ssh when a key is present. Craig's call.
-
-*Decided 2026-08-19: the ISO bakes the value, and a check nets the rest.*
-=archsetup:240= has the identical default for =archsetup_repo=, so this was
-always two repos rather than one. I ruled out detection — archsetup never
-restores =~/.ssh=, so key-presence at clone time depends on ordering it
-doesn't control, and "any key means ssh" would break a stranger who has an
-unrelated one. I ruled out a bare post-install step for the reason this whole
-class of bug exists: manual steps don't get run, which is why this sat four
-days. So the personal ISO carries =ARCHSETUP_REPO= / =DOTFILES_REPO= in the
-ssh form (noted on the secrets/ISO task), and =post-rebuild-check= check 8
-flags any working repo still on the read-only endpoint — covering curl|bash
-and stock-ISO installs, which the ISO value cannot reach.
-
-Repair on a machine already built: =git fetch --unshallow= in each repo, and
-=git remote set-url origin git@cjennings.net:<repo>.git= for dotfiles.
-
-Grading: Major severity (two working repos silently missing their history on
-the machine I develop on, and it returns confidently wrong answers to history
-questions rather than failing) x every user every time (every fresh install,
-both daily drivers) = P1 = [#A].
-
-Not :solo:. The depth half is (two lines plus tests in the existing
-=tests/installer-steps/= shape, verifiable by asserting the clone command
-carries no =--depth= for these two repos). The remote-URL half needs the
-decision above, so the task as a whole waits on it. Split it in two if the
-depth fix is wanted sooner.
-*** 2026-08-19 Wed @ 23:05:00 -0700 Dropped --depth from both user-repo clones
-=archsetup:1462= and =:1475= now clone full history;
-=tests/installer-steps/test_clone_user_repos.py= covers it with 8 cases, and
-one of them asserts the AUR build clones still carry =--depth 1= so the fix
-can't be over-applied by a careless repo-wide sed. Both my repos on velox were
-already unshallowed by hand last session, so this is prevention rather than
-repair.
-*** 2026-08-19 Wed @ 23:05:00 -0700 Settled the remote-URL half and netted it
-See the decision recorded above. The ISO half is a note on the secrets/ISO
-task; the net is =post-rebuild-check= check 8, which ships now.
-
-Both halves resolved. Depth: =a028aa5= drops =--depth 1= from both user-repo clones, with 8 tests including one asserting the AUR build clones stay shallow. Remote URL: decided 2026-08-19 (see above) — the personal ISO carries the ssh form, and =post-rebuild-check= check 8 (=87ff0b7=) flags any working repo still on the read-only endpoint, covering the install paths the ISO cannot reach.
-** DONE [#D] Worldclock tooltip blanks on one bad timezone row :bug:dotfiles:waybar:quick:solo:
-CLOSED: [2026-08-21 Fri]
-:PROPERTIES:
-:LAST_REVIEWED: 2026-07-25
-:END:
-Found by sentry (2026-07-25), verified by exercising. =hyprland/.local/bin/waybar-worldclock= builds each zone with =ZoneInfo(tz)= inside the loop (line ~99) with no guard, so a single malformed timezone row in =worldclock.conf= raises =ZoneInfoNotFoundError= and crashes the whole python pass. The tooltip then renders empty and *every* zone is lost, not just the bad row; the traceback only reaches stderr, where waybar never surfaces it.
-Repro: a conf with =America/Chicago|Home=, =Not/AZone|Bad=, =Europe/London|London= renders =tooltip: ""= (Home and London gone too).
-Grade: minor severity (one module's tooltip blanks, no data loss) x rare edge case (a malformed conf row) = P4 = [#D].
-Fix: wrap the per-row =ZoneInfo=/=datetime= in a try/except and =continue=, so a typo drops only that row and the valid zones still render. Solo + quick: the script already has an env-override test harness (=WAYBAR_TIME_EPOCH=, =WAYBAR_WORLDCLOCK_CONF=), so a red-first test is cheap.
-
-Fixed in dotfiles =8f692f5=. The per-row =ZoneInfo= is guarded, so a malformed row drops itself and the valid zones still render. Five cases, including a bad row first — the ordering that looks least like one typo and most like the module being broken. Caught the broad =except Exception= rather than =ZoneInfoNotFoundError=, because the row also parses floats and calls strftime and the contract wanted is "a bad row costs only itself".
-** DONE [#C] obsbot-wb-guard polls forever on machines with no OBSBOT :bug:dotfiles:quick:solo:
-CLOSED: [2026-08-21 Fri]
-:PROPERTIES:
-:LAST_REVIEWED: 2026-08-16
-:END:
-=obsbot-wb-guard.service= is =WantedBy=graphical-session.target= and lives in the shared =common/= stow tier, so it starts on every machine. Its main path is =while :; do check_once; sleep 2; done=, and =check_once= returns early when the camera node is absent. On a machine with no OBSBOT attached that is a process waking every two seconds forever to do nothing, which on a laptop is battery spend for zero benefit. No restart loop, though: the loop never exits, so =Restart=on-failure= never fires.
-
-Found 2026-08-16 on velox, after enabling it to match ratio and then having to disable it again by hand. A per-machine disable is the wrong shape, because it drifts velox from ratio permanently and a re-stow or a future audit will just put it back.
-
-Fix: give the unit =ConditionPathExists= on the camera node (=/dev/v4l/by-id/usb-Remo_Tech_Co.__Ltd._OBSBOT_PW106-video-index0=, the same default the script uses) so systemd skips it on any machine without the camera and starts it normally on ratio. Then re-enable it on velox, where it will simply be skipped. Note the limit: a camera plugged in later will not start it until the next login, which is the right trade against a permanent poll.
-
-Careful when disabling by hand in the meantime: =systemctl --user disable= on a *linked* unit deletes the unit symlink, and that symlink is stow-managed, so a bare disable silently removes a file from the dotfiles stow tree. Restore the link afterward or re-stow.
-
-Grade: minor severity (wasted wakeups and battery, no data loss, no failure) x every boot on any machine without the camera = P3 = [#C].
-
-Solo: buildable here (archsetup owns dotfiles end-to-end), verifiable by the agent (assert the unit is skipped on velox and still active on ratio), and no design call left open.
-
-Fixed in dotfiles =566dd14=. =ConditionPathExists= on the camera node, so systemd skips the unit where the camera is absent. velox is now =enabled= like ratio and reports =ConditionResult=no=; the stow symlink is untouched. Found while doing it: ratio has a Logitech BRIO and no OBSBOT on USB at all, so the 2-second poll was pointless on the desktop too, not merely costing laptop battery. A test asserts the unit's condition path and the script's =OBSBOT_WB_DEVICE= default stay equal, since drift there is invisible in both directions.
-** DONE [#C] Spine face tests decay against the wall clock :bug:test:dotfiles:solo:
-CLOSED: [2026-08-21 Fri]
-:PROPERTIES:
-:LAST_REVIEWED: 2026-08-02
-:END:
-=settings/faces/timeline-face-spine.test.mjs= has thirteen =SP.spineRows(g, h)= calls that omit the third argument, so =ref= falls back to its =new Date()= default while the file's events fixture is pinned to =JUL= (2026-07-31 18:30 UTC). Any assertion that depends on how much room the day needs is then measured against today's clock, and rots as the fixture recedes.
-
-One of them, "spacing is uniform everywhere except the gap home opens", had already rotted: green on 07-31 because that was the fixture's own date, red by 08-02. Fixed in place on 2026-08-02 by pinning =JUL=; the remaining thirteen pass today by luck. The measurement, for whoever picks this up — with =ref=now= the even step is 85.21 and home's gaps are 129.10 / 65.40 (the lower one collapses below a plain gap); with =ref=JUL= the step is 78.54 and the gaps are 129.10 / 145.46. Only the lower gap moves, because =up= does not depend on events and =down= does.
-
-Six other calls in the same file already pass =JUL= explicitly, so the convention exists and this is a miss, not a gap in the design. Fix: pass =JUL= at every call whose assertion reads geometry. Leave the call around line 747 alone — it sweeps =new Date(t0)= deliberately.
-
-Grade: minor severity (dev-facing only; no product behavior is wrong, the face itself is fine) x some users, sometimes (each call rots independently, whenever the fixture drifts far enough) = P3 = [#C]. Not merely cosmetic though: a suite that goes red for no real reason is how a genuine regression gets waved through.
-
-Solo — mechanical, an existing convention to copy, and verifiable by running the suite plus re-running it under a faked clock to prove the determinism actually holds.
-
-
-Fixed in dotfiles =c96a216=. All thirteen bare calls now pass =JUL=. The task's "line 747" was stale (the deliberate =t0= sweep is at 893 and already passed its own ref, so it was never at risk), and the continuation-form call closes its arguments on the next line, which is why a naive grep counts fourteen. Added a guard that reads the file and fails with the offending line numbers, and verified it bites by stripping =JUL= from one call and confirming it went red naming that line.
-** DONE [#A] Velox still carries the install placeholder passwords :bug:security:velox:
-CLOSED: [2026-08-23 Sun] SCHEDULED: <2026-08-20 Thu>
-:PROPERTIES:
-:CREATED: [2026-08-20 Thu]
-:LAST_REVIEWED: 2026-08-20
-:END:
-Closed 2026-08-23: I'd already rotated all three on the 08-14 bringup day, so
-this task was never live. Verified on velox before closing — =chage -l= puts the
-last password change for both =cjennings= and =root= at Aug 14 2026, and
-=/etc/zfs/zroot.key= was rewritten 2026-08-14 05:29 and no longer holds the
-placeholder (checked with a =grep -qx= that returns a yes/no without reading the
-key into a transcript).
-
-The premise below was wrong, and it's worth naming how. Nothing ever tested the
-credentials: the claim came from an unticked runbook item plus the archangel
-session handing back the values the *installer* had set, which reads as "these
-are current" only if you assume nobody changed them in between. An inference
-about a security exposure got recorded in the same voice as a measurement. The
-one command that settles it costs a second.
-
-Original body follows.
-
-The 2026-08-13 reinstall set placeholder credentials and the runbook's Phase 5
-item to replace them (=passwd=, =zfs change-key zroot=) was never ticked.
-Believed still live 2026-08-20 via the archangel handoff, which had to hand
-them back to Craig to get into the machine: =welcome1= for the pool, =welcome=
-for the accounts.
-
-So velox's full-disk encryption is currently protected by a dictionary word
-with a digit, on the machine that travels. Anyone who picks it up owns the pool
-and every account on it — the encryption is doing no work at all.
-
-Two commands, both on velox:
-- =passwd= for each account.
-- =zfs change-key zroot= for the pool passphrase. Note this is the ZBM unlock
- passphrase, so get it right before rebooting.
-
-Grading: *severity-alone carve-out* — this is a security exposure, so the
-frequency row does not discount it (=todo-format.md=). Critical severity: total
-compromise of an encrypted-at-rest laptop from a guessable string, with the
-device leaving the house. = P1 = [#A].
-
-Distinct from the =VERIFY [#A] Rotate the credentials exposed by the 2026-08-09
-dotfiles leak= under the cgit audit — that one covers credentials a crawler
-already took from a public repo. This one is a local default never changed. Both
-are rotation work; neither substitutes for the other.
-** CANCELLED [#B] Consistent keybinding family for the panel console :feature:hyprland:
-CLOSED: [2026-08-21 Fri]
-:PROPERTIES:
-:LAST_REVIEWED: 2026-07-09
-:END:
-Merged into =[#B] Reconcile panel keybindings around Super+N=, which now carries
-this body's detail: the collision list, the velox plain-keyboard constraint, and
-maintenance-M as the priority chord. Cancelled rather than done — the work is
-still open, just tracked in one place instead of two.
-
-Consider putting every panel (net, bluetooth, audio, timer, and the coming maintenance console) on one consistent chord family — a shared modifier set (Super+Shift, Control+Alt, or similar) plus a mnemonic letter per panel (N/B/A/T/M). Today the panels open via waybar clicks only; a uniform chord family makes them keyboard-reachable and predictable. Watch for collisions with existing binds: Super+Shift+A is already PTT toggle, and the hold-to-talk grave bind is load-bearing. Decide the family, audit current hyprland binds for conflicts, wire via the dotfiles hyprland config, and document in the keybind reference. Both machines (velox can't QMK-remap, so chords must work on a plain laptop keyboard).
-*** 2026-07-14 Tue @ 00:31:36 -0500 Folded Craig's ask for a maintenance-panel keybinding; bumped [#C] → [#B]
-Craig asked (in session, 2026-07-14) for a maintenance keybinding specifically — the panel he's reaching for without one. Maintenance (M) is the priority chord when this task gets worked. The capture graduated the task from parking lot to active backlog.
-** DONE [#B] Waybar network module — custom/net :feature:waybar:network:
-CLOSED: [2026-08-21 Fri]
-:PROPERTIES:
-:LAST_REVIEWED: 2026-07-09
-:END:
-Closed 2026-08-21: the module shipped and is in daily use. Phases 1-4 all landed
-in dotfiles, and the tunnels track absorbed most of what Phase 5 originally
-covered. The one piece genuinely left — the =net vpn= CLI subcommand — is now its
-own task below, so the residual is tracked at its real size instead of holding a
-finished umbrella open.
-Unifies the old wifi-no-internet indicator (was =[#C]=) and the network-manager
-dropdown (was =[#B]=) into one =custom/net= module: a tested Python =net= engine
-(nmcli + diagnostics), a thin bar indicator, and a GTK4 layer-shell panel. Code
-lives in the dotfiles repo (hyprland tier + a =net/= package like pocketbook);
-archsetup only installs deps. Secrets stay in NetworkManager's own store (no
-separate credential store). The =captive= script becomes the diagnostics engine.
-Full design, acceptance criteria, and the failure-mode coverage table:
-[[file:docs/design/2026-06-29-waybar-network-module-spec.org][2026-06-29-waybar-network-module-spec.org]].
-
-Phases below, dependency order. Engine/unit work is agent-verifiable (=unittest=
-+ fakes on PATH, coverage via venv); the live-network and visual states need real
-conditions, filed under "Manual testing and validation".
-
-*** 2026-06-29 Mon @ 20:19:11 -0400 Phase 1 shipped — indicator + console recovery
-Shipped to the dotfiles repo (10 commits, =5254bd8=..=c095a22=, pushed to main).
-The =net= engine is a src-layout Python package in-tree, imported by a bin shim
-that resolves the stow symlink back to the repo — so it runs from a bare TTY with
-no install, which the recovery path depends on.
-
-Landed: =net status= (fast path, one nmcli call + sysfs, degraded fallback in
-budget) + =net probe= (native captive probe, single-flight flock, atomic cache,
-fresh/stale/expired/unknown classes, iface/SSID/UUID invalidation); =waybar-net=
-replacing =custom/netspeed=, throughput → tooltip, CSS states in both themes +
-live; =net diagnose= (read-only steps) + =net repair= (rfkill/reset/bounce/
-dns-test, cleanup-verified) + =net doctor [--fix]= with the four terminal
-classifications; =net portal= + the =captive --probe-json= refactor; redacted
-JSONL event log; Makefile recovery targets (=make online= etc.); =~/.config/net/
-config=. Verified live: =make net-status= reads the real wlp170s0 / @Hyatt_WiFi.
-
-Airplane (Craig's call, option 1): =custom/net= absorbs only the *display* — net
-reads the airplane-mode state file and shows an airplane state/glyph. The
-airplane-mode toggle stays (it's a low-power mode — radios + CPU + brightness +
-services — not a radio switch), now on =custom/net='s right-click + signal 15.
-Deleted: =waybar-airplane=, =waybar-netspeed=, =custom/airplane=, their tests +
-css. =airplane-mode= kept.
-
-Tests: 160 in =tests/net/= (fake nmcli/curl/rfkill/resolvectl/ping/getent/
-systemctl on a temp PATH; doctor-classification fixtures; degraded-under-slow-
-nmcli benchmark) + the =captive= probe-mode tests; full dotfiles suite green (32
-suites). Coverage-gap pass via throwaway venv: pure modules ≥90% branch
-(classify 100%), IO-error branches excused in the test docstring.
-Deferred to Phase 2/3: archsetup deps (gtk4-layer-shell/python-gobject Phase 2,
-speedtest-go-bin Phase 3 — not added before the code that needs them).
-Verify (manual, live): see Manual testing and validation.
-
-*** 2026-06-29 Mon @ 22:19:25 -0400 Phase 2 shipped — panel shell + connection management
-Shipped to dotfiles (commits =4e7740f=..=24bcac5=, pushed). Engine: =net list= (saved
-MRU + in-range wifi scan, infrastructure types filtered), =net up/down= (UUID-keyed,
-mutation safety — keep prior link until target activates, classify wrong-password vs
-generic, report auto-reactivation), =net add/edit/remove/rescan= (open + WPA-PSK;
-enterprise activate-only; secret to NM's store, never our JSON/log — tested).
-
-Panel: a GTK-free PanelModel (selection, four state machines, the UX-flow enable
-rules, terminal states) + a GTK4 gtk4-layer-shell window (=net panel=) anchored
-top-right under the bar — Connections section with MRU list, active marked, signal
-glyph, row-click select, Connect/Add/Forget/Rescan, confirm-on-forget, worker-thread
-engine calls via GLib.idle_add. GTK imported lazily so the CLI/tests stay GTK-free.
-
-Bar interactions (settled with Craig over live iteration): left = =net-panel= toggle,
-middle = =net portal=, right = =net-fix= (notify the doctor result when one-way; open
-a terminal only when the outcome is fixable — the sudo/interactive case). Airplane on
-Super+Shift+A. archsetup adds =gtk4-layer-shell= + =python-gobject= (this commit);
-already on velox.
-
-Tests: 204 in tests/net (merge ordering/dedup, up/down mutation safety, no-secret-leak
-on add/edit, panel model + state machines, gui row-format helpers). Full dotfiles suite
-green (32 suites). Live-verified on velox: panel opens/toggles, list shows real 24
-profiles, right-click notification delivers (Craig confirmed). Phase 3 (diagnose/repair/
-speedtest IN the panel) is next; the engine for it already exists from Phase 1.
-
-*** 2026-06-29 Mon @ 22:43:40 -0400 Phase 3 shipped — diagnostics + speed test in the panel
-Shipped to dotfiles (=91277cf=..=691abcb=) + archsetup (=48052d6=, speedtest-go-bin),
-pushed. Engine: =net speedtest= (parses speedtest-go --json → ping from latency ns,
-down/up from per-server byte rates; missing-backend / offline / malformed → error
-envelope per the failure table). Panel grew a section switcher with four pages:
-- Connections (Phase 2).
-- Diagnose: =net diagnose= on a worker thread, each step a row (✓/✗/… glyph + title +
- redacted evidence), read-only; Open-portal button when captive.
-- Repair: "Get me online" (=net doctor --fix=) + tiers (rfkill/reset/bounce/dns-test)
- + force portal. Confirmations in-panel with the spec's exact wording; the privileged
- tiers run via =net-popup= terminal (where the sudo prompt + step output, incl.
- cleanup-verified, show) — a panel has no tty, and pkexec would mean a prompt per op.
-- Speed test: in-process =net speedtest= (no privilege → inline result: ↓/↑ Mbps + ping
- + server), Run/Cancel (Cancel pkills the child), error envelope shown.
-
-213 net tests; pure helpers (step_indicator, format_speedtest) unit-tested. Full
-dotfiles suite green (32 suites). One unverified assumption: speedtest-go's dl/ul unit
-(taken as bytes/s; =BYTES_PER_SEC= flips it) — needs one real run vs a reference. The
-in-panel repair streaming (vs terminal) is a named future polish once the GUI-privilege
-story settles.
-
-The waybar network module ([#B] parent) is now COMPLETE through Phase 3. Phase 4
-(in-app help + user guide) and Phase 5 (VPN/WireGuard) remain as future work; the core
-feature (indicator + recovery + panel + diagnostics + speed test) is done.
-Verify (manual, live): see Manual testing and validation.
-
-*** 2026-07-09 Thu @ 16:32:54 -0500 Audit reconcile: Phase 4 is filed on the dotfiles side, waiting on them
-The dotfiles project accepted the Phase 4 handoff and filed it as a =[#C]= task in their own =todo.org= (their note, 2026-07-08 16:56): the help-text audit + panel help affordance, the user-guide/README, and the ratio rollout doc. Not started there. They ping when it lands, and this task's Phase 4 child closes then. Nothing to do here meanwhile.
-
-*** 2026-08-17 Mon @ 19:57:42 -0700 Landed on the dotfiles side; the block is cleared
-dotfiles shipped it as =138da7b= and closed its own task, so this one closes
-with it and the =:blocked:= tag comes off. Found by checking their =todo.org=
-rather than waiting for the ping — their close-out note says "archsetup pinged
-so its Phase 4 task can close", so the handoff worked and only this end was
-left open.
-
-All three acceptance criteria are met on their side: the help audit found and
-fixed a stale =net repair= action list (nine of nineteen actions were named;
-both the CLI help and =repair.py='s docstring now generate from the ACTIONS
-registry), =net/README.md= covers every command plus the recovery targets, and
-the ratio rollout is documented with both daily drivers verified current.
-
-They split the panel help affordance out rather than inventing it — no sibling
-panel has one, so its shape is a design call. It is tracked on their side, not
-here.
-
-Original deliverable, for the record: in-app help (=net --help= + per-command,
-panel help affordance); README/user-guide; archsetup Hyprland dep install
-(=gtk4-layer-shell=, =python-gobject=, =speedtest-go-bin=); ratio manual dep +
-stow step. Handed off 2026-07-04 with the archsetup deps already confirmed
-installed.
-
-*** 2026-08-21 Fri @ 14:18:03 -0700 Promoted the Phase 5 residual out to its own task
-Rescoped 2026-07-04 (audit): the tunnels track already shipped most of the original Phase 5. Panel tunnel bring-up/down and detection landed (dotfiles 2d9d060 probes tailscale/NM-wireguard/Proton; 21db05a brings overlays up/down from the panel's Tunnels sub-view; 31ba056 diagnose/doctor understand tunnel routes; archsetup 2e40781 wireguard config import; the net-panel-other-interfaces spec is IMPLEMENTED). What remains for Phase 5 is only the =net vpn ...= CLI subcommand — cli.py still has no vpn/tunnel parser. Fold the panel's existing tunnel operations into a CLI surface; spec separately when picked up.
-** DONE [#A] Ratio: pull .emacs.d before upgrading Emacs to 31.1 :chore:ratio:emacs:
-CLOSED: [2026-08-25 Tue]
-:PROPERTIES:
-:CREATED: [2026-08-25 Tue]
-:LAST_REVIEWED: 2026-08-25
-:END:
-Emacs 31.1's warnings.el defers daemon-startup warnings into a closure holding
-the =*Warnings*= buffer; the config's dashboard-only sweep killed that buffer,
-so the first client frame of every fresh 31.1 daemon failed on Wayland and
-emacsclient silently fell back to =$DISPLAY= (XWayland, pgtk warning dialog).
-Fixed in =.emacs.d= commit =63831060= (2026-08-25, velox verified live:
-=GdkWaylandDisplay=). Ratio is still on 30.2, which lacks the deferring code,
-so it is fine until it upgrades — then it hits the same trap once per daemon
-start unless the fix is pulled first.
-
-Order on ratio: =git -C ~/.emacs.d pull= (the push from velox is the telega
-session's; confirm =63831060= is on origin first), then the =pacman -Syu= that
-brings =emacs-wayland 31.1=, then restart the daemon. Check afterwards:
-=emacsclient -e '(pgtk-backend-display-class)'= → =GdkWaylandDisplay=.
-
-*** 2026-08-25 18:10 — pull already landed; the upgrade half remains
-Checked ratio over tailscale: =~/.emacs.d= is clean at =91fbac72= (= =origin/main=),
-and =63831060= is an ancestor of HEAD — =modules/undead-buffers.el= carries the
-=*Warnings*= entry. Ratio is on =emacs-wayland 30.2-3= with =31.1-1= pending among
-720 updates (last full upgrade 2026-08-01; kernel 7.1.5 → 7.1.9 also pending,
-btrfs root, uptime 3.5 weeks). The daemon is a plain =emacs --daemon= (not a
-user unit) holding 2 live frames, so the restart step will drop those frames.
-What remains: the =pacman -Syu= on ratio, the daemon restart, and the
-=(pgtk-backend-display-class)= check.
-
-*** 2026-08-25 Tue @ 18:35:00 -0600 Upgraded ratio to Emacs 31.1 and verified the Wayland backend
-Ran the upgrade over tailscale as a transient unit (=ratio-upgrade.service=,
-log at =/var/log/ratio-upgrade.log=): 714 packages, =--ignore= on the six
-packages the live-update guard would have blocked (aquamarine, hyprland,
-hyprutils, mesa, vulkan-radeon, wayland — still pending, apply from a TTY
-before the reboot). One orphan cleared first: =qemu-block-gluster= had been
-dropped from the repo and pinned =qemu-common=; the new =qemu-full= no
-longer needs it. Killed the plain =emacs --daemon= (no modified buffers, no
-graphical frames), started =emacs.service= instead so the daemon carries the
-systemd user environment, and probed from a throwaway frame:
-=(pgtk-backend-display-class)= → =GdkWaylandDisplay=, =*Warnings*= alive.
-Ratio still wants a reboot for =linux 7.1.9=. Pacnews to review there:
-=/etc/ssh/sshd_config.pacnew= and two =/etc/tpm2-tss/fapi-profiles/*.json=.
** DONE [#B] Function keys issue media actions instead of F-keys :bug:velox:
CLOSED: [2026-09-01 Tue]
:PROPERTIES:
@@ -4817,3 +4041,25 @@ main and release/0.4.0 heads are on the remote as-is, and its only other commit
(bf0457f) is there as the patch-identical cb70193, so the bundle and its
working dir are deleted.
emacs-wttrin has a handoff note in its inbox covering all of it.
+** DONE [#A] Pre-vacation fix list — morning review
+CLOSED: [2026-09-17 Thu]
+:PROPERTIES:
+:LAST_REVIEWED: 2026-09-17
+:END:
+Reviewed 2026-08-08; the decisions it drew are recorded in the tasks below.
+Closed at the 2026-09-17 review because the ~08-15 departure it ranked work for
+has passed. Where each of the seven items went:
+
+1. Velox reliability → the [#A] sleep/suspend task.
+2. Velox machine health for travel → moot: velox was wiped and reinstalled on
+ 2026-08-13.
+3. Remote access from outside the LAN → the wolf WireGuard rider under the
+ sleep/suspend task. Tailscale to ratio works off-LAN (checked 2026-09-17);
+ truenas and truenas-kvm weren't re-checked.
+4. cgit secrets audit → the [#A] audit task and its rotation VERIFY.
+5. Osbot camera → its own task; the podman socket and camera udev rule are
+ installed (sleep/suspend task, 08-17 entry).
+6. Hotspot/metered WiFi and network ordering → the held design calls under
+ Next Session Focus.
+7. The orchestrator sequence-pin gap → filed 2026-09-17 as [#C] Orchestrator
+ sequence pin misses an added step.