aboutsummaryrefslogtreecommitdiff
path: root/todo.org
diff options
context:
space:
mode:
authorCraig Jennings <c@cjennings.net>2026-08-19 12:32:07 -0700
committerCraig Jennings <c@cjennings.net>2026-08-19 12:32:07 -0700
commit7b68f77a9b99e5400472986cb80bae5fb92e2320 (patch)
treec9beef06a73bd6b107753c1e8e3e6350ad290aad /todo.org
parentec3a63caca4f2d955e594318a9a690e4c28af19e (diff)
downloadarchsetup-7b68f77a9b99e5400472986cb80bae5fb92e2320.tar.gz
archsetup-7b68f77a9b99e5400472986cb80bae5fb92e2320.zip
docs: correct the clock/DNS deadlock mechanism to DNSSEC
I reproduced the failure by winding velox's clock back 27 days with chronyd stopped, and the cause is not what I recorded. Resolved logged signature-expired against the root DNSKEY and every DS beneath it. The DoT handshake to 1.1.1.1:853 verified clean at that same clock, and the Cloudflare certificate runs Dec 2025 to Dec 2026, so it was never outside its window. An RRSIG window is days to weeks while a certificate is good for a year, so a skew that breaks DNSSEC normally leaves DoT untouched. DNSSEC=allow-downgrade does not rescue it either. Resolved downgrades when a server lacks DNSSEC support, and a signature-window failure is a validation failure, so no downgrade fires. Six retries over eighteen seconds plus a reset-server-features, all dead. I briefly believed otherwise off a test whose success was a cache hit. The fix itself is verified end to end. With the clock wound back and no DNS at all, chronyd reached the IP-addressed source and stepped the clock straight back. Also settled: the clock landed on 2026-07-23 because that is systemd 261.2's build date to the minute, and systemd advances a garbage RTC to its own build epoch at boot.
Diffstat (limited to 'todo.org')
-rw-r--r--todo.org99
1 files changed, 90 insertions, 9 deletions
diff --git a/todo.org b/todo.org
index 1bb810d..fa3252d 100644
--- a/todo.org
+++ b/todo.org
@@ -78,6 +78,41 @@ with a wrong clock, which is any RTC fault, BIOS reset, or drained cell) = P2 =
[#B]. Graded on the being-in-it, not the getting-into-it: once the machine is in
this state it is fully offline with no local path out.
+*** 2026-08-19 Wed @ 12:25:00 -0700 Reproduced it, and the mechanism was not what either of us said
+I wound velox's clock back 27 days with chronyd stopped and watched it fail.
+Resolution died outright, and plain UDP/53 to 1.1.1.1 kept answering throughout
+— the discriminator the doctor keys on, confirmed live rather than reasoned.
+
+The cause is DNSSEC, not DNS-over-TLS. resolved logged =signature-expired=
+against the root DNSKEY and every DS beneath it. The DoT handshake to
+=1.1.1.1:853= verified clean at that same clock, and the Cloudflare certificate
+runs Dec 2025 to Dec 2026, so it was never outside its window. An RRSIG window
+is days to weeks and a certificate is good for a year, so a skew that breaks
+DNSSEC normally leaves DoT untouched. The phone session blamed the certificate
+and I carried that forward into the first commit; both were wrong.
+
+=DNSSEC=allow-downgrade= does not rescue it either, which matters because it is
+the obvious reach and it is what ratio runs. resolved downgrades when a server
+lacks DNSSEC support, and a signature-window failure is a validation failure, so
+no downgrade fires. Six retries over eighteen seconds plus
+=resolvectl reset-server-features=, all dead. I briefly believed otherwise off a
+test whose success was a cache hit (=Data from: cache network=).
+
+So ratio was exposed after all, and I have given it the same drop-in. Its
+=162.159.200.1= is selected and its clock is synchronized.
+
+The fix itself is verified end to end: with the clock wound back and no DNS at
+all, chronyd reached the IP-addressed source and stepped the clock from
+2026-07-23 straight back to 2026-08-19. That is the whole claim, demonstrated
+rather than argued.
+
+Also settled: the clock landed on 2026-07-23 because that is systemd 261.2's
+build date to the minute (=/usr/lib/systemd/systemd=, 10:43:59), and systemd
+advances a garbage RTC to its own build epoch at boot. Not timesyncd's
+last-good-sync timestamp, which cannot be it — timesyncd is disabled here. That
+also confirms the RTC really was reading earlier than that, so the coin cell
+stays the prime suspect.
+
*** 2026-08-19 Wed @ 10:12:00 -0700 Root fix, doctor verdict, and taxonomy entry landed
The installer carries the drop-in; =post-rebuild-check= grew a sixth check that
fails a machine whose every NTP source is a hostname; the net failure taxonomy
@@ -93,6 +128,51 @@ deliberately DNS-free: a local =timedatectl= read for sync state, and a bypass
query addressed by IP over plain UDP/53 to tell "resolved is refusing to
validate" apart from "DNS is genuinely dead".
+** VERIFY [#C] DNSSEC strictness on the travelling laptop :velox:
+:PROPERTIES:
+:CREATED: [2026-08-19 Wed]
+:LAST_REVIEWED: 2026-08-19
+:END:
+
+I changed velox to =DNSSEC=allow-downgrade= today and then put it back to =yes=,
+because the reason I changed it turned out to be false. It does not prevent the
+clock deadlock. The IP-addressed NTP source does, and that is already in place
+on both machines.
+
+What remains is a different question the taxonomy already documents: =DNSSEC=yes=
+hard-fails against venue resolvers that mangle DNSSEC records, which is a hotel
+and airport problem and therefore velox's problem more than ratio's.
+=allow-downgrade= trades authenticated answers for staying online. ratio already
+runs =allow-downgrade=, so the fleet disagrees with itself and with the
+installer, and I do not know whether ratio's setting was a deliberate policy or
+a forgotten workaround for one bad network.
+
+I have not made this call. It is a security posture change and it should be made
+knowingly rather than as a side effect of a theory I disproved an hour later.
+
+** TODO [#C] Branch network policy on laptop vs desktop in the installer :feature:
+:PROPERTIES:
+:CREATED: [2026-08-19 Wed]
+:LAST_REVIEWED: 2026-08-19
+:END:
+
+Three defaults were chosen for a desktop and then applied to the machine that
+travels: =DNSSEC=yes= (hard-fails on venue resolvers), a fresh wifi MAC per
+connection (every hotel reconnect looks like a new device, so the portal login
+starts over), and hostname-only NTP (the deadlock). Two are now fixed for every
+machine, and the third is the VERIFY above.
+
+The installer already branches on battery presence in three places:
+=prune_waybar_battery=, the ppd mask in =configure_tlp_power=, and the TLP config
+itself, all keyed on =ls /sys/class/power_supply/BAT*=. The precedent exists and
+network policy simply does not use it. Wiring the same test around
+=configure_networking= would let laptop and desktop defaults diverge deliberately
+instead of by drift, and would stop the next instance of this from happening.
+
+Grading: Minor severity (nothing is broken today; this prevents a recurrence) x
+some users sometimes (bites when a new default suits one machine class and not
+the other) = P3 = [#C].
+
** TODO [#C] Automate the clock/DNS deadlock repair in the net doctor :feature:
:PROPERTIES:
:CREATED: [2026-08-19 Wed]
@@ -1817,15 +1897,16 @@ Craig's standing checklist of everything that isn't agent-verifiable. Each child
Priority and type tag added by that audit: the task carried neither, which kept the project's largest live container out of the agenda entirely.
*** Clock/DNS deadlock: does velox recover its clock from a cold boot, untouched?
-What we're verifying: that the IP-addressed NTP drop-in actually breaks the
-bootstrap deadlock on a real cold start. This is the one test no agent can run —
-it needs a full power-down, which is exactly the event that empties a failing
-RTC. Everything else about the fix is verified; this is the part that rests on
-construction (an address needs no DNS, NTP carries no certificate) rather than
-on having been seen work.
-
-Do this before relying on it away from home — the failure mode strands the
-machine with no network and no way to look anything up.
+What we're verifying: the coin cell, not the fix. The fix itself is already
+demonstrated — on 2026-08-19 I wound the clock back 27 days with no DNS at all,
+and chronyd reached the IP-addressed source and stepped it straight back. What a
+cold boot adds is the hardware question: whether the RTC actually loses time
+across a full power-down, which is the thing that started this.
+
+Read the outcome carefully, because only one branch is informative. An RTC time
+that comes up wrong and then self-corrects tells you the cell is dying and the
+fix is holding. An RTC that comes up correct tells you nothing about the
+deadlock at all, only that the cell survived this particular night.
- Confirm the drop-in is in place and chrony is using it (block below).
- Shut all the way down — =poweroff=, not suspend, not reboot. The RTC only
loses time when the machine is actually off.