aboutsummaryrefslogtreecommitdiff
path: root/todo.org
diff options
context:
space:
mode:
Diffstat (limited to 'todo.org')
-rw-r--r--todo.org99
1 files changed, 90 insertions, 9 deletions
diff --git a/todo.org b/todo.org
index 1bb810d..fa3252d 100644
--- a/todo.org
+++ b/todo.org
@@ -78,6 +78,41 @@ with a wrong clock, which is any RTC fault, BIOS reset, or drained cell) = P2 =
[#B]. Graded on the being-in-it, not the getting-into-it: once the machine is in
this state it is fully offline with no local path out.
+*** 2026-08-19 Wed @ 12:25:00 -0700 Reproduced it, and the mechanism was not what either of us said
+I wound velox's clock back 27 days with chronyd stopped and watched it fail.
+Resolution died outright, and plain UDP/53 to 1.1.1.1 kept answering throughout
+— the discriminator the doctor keys on, confirmed live rather than reasoned.
+
+The cause is DNSSEC, not DNS-over-TLS. resolved logged =signature-expired=
+against the root DNSKEY and every DS beneath it. The DoT handshake to
+=1.1.1.1:853= verified clean at that same clock, and the Cloudflare certificate
+runs Dec 2025 to Dec 2026, so it was never outside its window. An RRSIG window
+is days to weeks and a certificate is good for a year, so a skew that breaks
+DNSSEC normally leaves DoT untouched. The phone session blamed the certificate
+and I carried that forward into the first commit; both were wrong.
+
+=DNSSEC=allow-downgrade= does not rescue it either, which matters because it is
+the obvious reach and it is what ratio runs. resolved downgrades when a server
+lacks DNSSEC support, and a signature-window failure is a validation failure, so
+no downgrade fires. Six retries over eighteen seconds plus
+=resolvectl reset-server-features=, all dead. I briefly believed otherwise off a
+test whose success was a cache hit (=Data from: cache network=).
+
+So ratio was exposed after all, and I have given it the same drop-in. Its
+=162.159.200.1= is selected and its clock is synchronized.
+
+The fix itself is verified end to end: with the clock wound back and no DNS at
+all, chronyd reached the IP-addressed source and stepped the clock from
+2026-07-23 straight back to 2026-08-19. That is the whole claim, demonstrated
+rather than argued.
+
+Also settled: the clock landed on 2026-07-23 because that is systemd 261.2's
+build date to the minute (=/usr/lib/systemd/systemd=, 10:43:59), and systemd
+advances a garbage RTC to its own build epoch at boot. Not timesyncd's
+last-good-sync timestamp, which cannot be it — timesyncd is disabled here. That
+also confirms the RTC really was reading earlier than that, so the coin cell
+stays the prime suspect.
+
*** 2026-08-19 Wed @ 10:12:00 -0700 Root fix, doctor verdict, and taxonomy entry landed
The installer carries the drop-in; =post-rebuild-check= grew a sixth check that
fails a machine whose every NTP source is a hostname; the net failure taxonomy
@@ -93,6 +128,51 @@ deliberately DNS-free: a local =timedatectl= read for sync state, and a bypass
query addressed by IP over plain UDP/53 to tell "resolved is refusing to
validate" apart from "DNS is genuinely dead".
+** VERIFY [#C] DNSSEC strictness on the travelling laptop :velox:
+:PROPERTIES:
+:CREATED: [2026-08-19 Wed]
+:LAST_REVIEWED: 2026-08-19
+:END:
+
+I changed velox to =DNSSEC=allow-downgrade= today and then put it back to =yes=,
+because the reason I changed it turned out to be false. It does not prevent the
+clock deadlock. The IP-addressed NTP source does, and that is already in place
+on both machines.
+
+What remains is a different question the taxonomy already documents: =DNSSEC=yes=
+hard-fails against venue resolvers that mangle DNSSEC records, which is a hotel
+and airport problem and therefore velox's problem more than ratio's.
+=allow-downgrade= trades authenticated answers for staying online. ratio already
+runs =allow-downgrade=, so the fleet disagrees with itself and with the
+installer, and I do not know whether ratio's setting was a deliberate policy or
+a forgotten workaround for one bad network.
+
+I have not made this call. It is a security posture change and it should be made
+knowingly rather than as a side effect of a theory I disproved an hour later.
+
+** TODO [#C] Branch network policy on laptop vs desktop in the installer :feature:
+:PROPERTIES:
+:CREATED: [2026-08-19 Wed]
+:LAST_REVIEWED: 2026-08-19
+:END:
+
+Three defaults were chosen for a desktop and then applied to the machine that
+travels: =DNSSEC=yes= (hard-fails on venue resolvers), a fresh wifi MAC per
+connection (every hotel reconnect looks like a new device, so the portal login
+starts over), and hostname-only NTP (the deadlock). Two are now fixed for every
+machine, and the third is the VERIFY above.
+
+The installer already branches on battery presence in three places:
+=prune_waybar_battery=, the ppd mask in =configure_tlp_power=, and the TLP config
+itself, all keyed on =ls /sys/class/power_supply/BAT*=. The precedent exists and
+network policy simply does not use it. Wiring the same test around
+=configure_networking= would let laptop and desktop defaults diverge deliberately
+instead of by drift, and would stop the next instance of this from happening.
+
+Grading: Minor severity (nothing is broken today; this prevents a recurrence) x
+some users sometimes (bites when a new default suits one machine class and not
+the other) = P3 = [#C].
+
** TODO [#C] Automate the clock/DNS deadlock repair in the net doctor :feature:
:PROPERTIES:
:CREATED: [2026-08-19 Wed]
@@ -1817,15 +1897,16 @@ Craig's standing checklist of everything that isn't agent-verifiable. Each child
Priority and type tag added by that audit: the task carried neither, which kept the project's largest live container out of the agenda entirely.
*** Clock/DNS deadlock: does velox recover its clock from a cold boot, untouched?
-What we're verifying: that the IP-addressed NTP drop-in actually breaks the
-bootstrap deadlock on a real cold start. This is the one test no agent can run —
-it needs a full power-down, which is exactly the event that empties a failing
-RTC. Everything else about the fix is verified; this is the part that rests on
-construction (an address needs no DNS, NTP carries no certificate) rather than
-on having been seen work.
-
-Do this before relying on it away from home — the failure mode strands the
-machine with no network and no way to look anything up.
+What we're verifying: the coin cell, not the fix. The fix itself is already
+demonstrated — on 2026-08-19 I wound the clock back 27 days with no DNS at all,
+and chronyd reached the IP-addressed source and stepped it straight back. What a
+cold boot adds is the hardware question: whether the RTC actually loses time
+across a full power-down, which is the thing that started this.
+
+Read the outcome carefully, because only one branch is informative. An RTC time
+that comes up wrong and then self-corrects tells you the cell is dying and the
+fix is holding. An RTC that comes up correct tells you nothing about the
+deadlock at all, only that the cell survived this particular night.
- Confirm the drop-in is in place and chrony is using it (block below).
- Shut all the way down — =poweroff=, not suspend, not reboot. The RTC only
loses time when the machine is actually off.