diff options
Diffstat (limited to 'todo.org')
| -rw-r--r-- | todo.org | 99 |
1 files changed, 90 insertions, 9 deletions
@@ -78,6 +78,41 @@ with a wrong clock, which is any RTC fault, BIOS reset, or drained cell) = P2 = [#B]. Graded on the being-in-it, not the getting-into-it: once the machine is in this state it is fully offline with no local path out. +*** 2026-08-19 Wed @ 12:25:00 -0700 Reproduced it, and the mechanism was not what either of us said +I wound velox's clock back 27 days with chronyd stopped and watched it fail. +Resolution died outright, and plain UDP/53 to 1.1.1.1 kept answering throughout +— the discriminator the doctor keys on, confirmed live rather than reasoned. + +The cause is DNSSEC, not DNS-over-TLS. resolved logged =signature-expired= +against the root DNSKEY and every DS beneath it. The DoT handshake to +=1.1.1.1:853= verified clean at that same clock, and the Cloudflare certificate +runs Dec 2025 to Dec 2026, so it was never outside its window. An RRSIG window +is days to weeks and a certificate is good for a year, so a skew that breaks +DNSSEC normally leaves DoT untouched. The phone session blamed the certificate +and I carried that forward into the first commit; both were wrong. + +=DNSSEC=allow-downgrade= does not rescue it either, which matters because it is +the obvious reach and it is what ratio runs. resolved downgrades when a server +lacks DNSSEC support, and a signature-window failure is a validation failure, so +no downgrade fires. Six retries over eighteen seconds plus +=resolvectl reset-server-features=, all dead. I briefly believed otherwise off a +test whose success was a cache hit (=Data from: cache network=). + +So ratio was exposed after all, and I have given it the same drop-in. Its +=162.159.200.1= is selected and its clock is synchronized. + +The fix itself is verified end to end: with the clock wound back and no DNS at +all, chronyd reached the IP-addressed source and stepped the clock from +2026-07-23 straight back to 2026-08-19. That is the whole claim, demonstrated +rather than argued. + +Also settled: the clock landed on 2026-07-23 because that is systemd 261.2's +build date to the minute (=/usr/lib/systemd/systemd=, 10:43:59), and systemd +advances a garbage RTC to its own build epoch at boot. Not timesyncd's +last-good-sync timestamp, which cannot be it — timesyncd is disabled here. That +also confirms the RTC really was reading earlier than that, so the coin cell +stays the prime suspect. + *** 2026-08-19 Wed @ 10:12:00 -0700 Root fix, doctor verdict, and taxonomy entry landed The installer carries the drop-in; =post-rebuild-check= grew a sixth check that fails a machine whose every NTP source is a hostname; the net failure taxonomy @@ -93,6 +128,51 @@ deliberately DNS-free: a local =timedatectl= read for sync state, and a bypass query addressed by IP over plain UDP/53 to tell "resolved is refusing to validate" apart from "DNS is genuinely dead". +** VERIFY [#C] DNSSEC strictness on the travelling laptop :velox: +:PROPERTIES: +:CREATED: [2026-08-19 Wed] +:LAST_REVIEWED: 2026-08-19 +:END: + +I changed velox to =DNSSEC=allow-downgrade= today and then put it back to =yes=, +because the reason I changed it turned out to be false. It does not prevent the +clock deadlock. The IP-addressed NTP source does, and that is already in place +on both machines. + +What remains is a different question the taxonomy already documents: =DNSSEC=yes= +hard-fails against venue resolvers that mangle DNSSEC records, which is a hotel +and airport problem and therefore velox's problem more than ratio's. +=allow-downgrade= trades authenticated answers for staying online. ratio already +runs =allow-downgrade=, so the fleet disagrees with itself and with the +installer, and I do not know whether ratio's setting was a deliberate policy or +a forgotten workaround for one bad network. + +I have not made this call. It is a security posture change and it should be made +knowingly rather than as a side effect of a theory I disproved an hour later. + +** TODO [#C] Branch network policy on laptop vs desktop in the installer :feature: +:PROPERTIES: +:CREATED: [2026-08-19 Wed] +:LAST_REVIEWED: 2026-08-19 +:END: + +Three defaults were chosen for a desktop and then applied to the machine that +travels: =DNSSEC=yes= (hard-fails on venue resolvers), a fresh wifi MAC per +connection (every hotel reconnect looks like a new device, so the portal login +starts over), and hostname-only NTP (the deadlock). Two are now fixed for every +machine, and the third is the VERIFY above. + +The installer already branches on battery presence in three places: +=prune_waybar_battery=, the ppd mask in =configure_tlp_power=, and the TLP config +itself, all keyed on =ls /sys/class/power_supply/BAT*=. The precedent exists and +network policy simply does not use it. Wiring the same test around +=configure_networking= would let laptop and desktop defaults diverge deliberately +instead of by drift, and would stop the next instance of this from happening. + +Grading: Minor severity (nothing is broken today; this prevents a recurrence) x +some users sometimes (bites when a new default suits one machine class and not +the other) = P3 = [#C]. + ** TODO [#C] Automate the clock/DNS deadlock repair in the net doctor :feature: :PROPERTIES: :CREATED: [2026-08-19 Wed] @@ -1817,15 +1897,16 @@ Craig's standing checklist of everything that isn't agent-verifiable. Each child Priority and type tag added by that audit: the task carried neither, which kept the project's largest live container out of the agenda entirely. *** Clock/DNS deadlock: does velox recover its clock from a cold boot, untouched? -What we're verifying: that the IP-addressed NTP drop-in actually breaks the -bootstrap deadlock on a real cold start. This is the one test no agent can run — -it needs a full power-down, which is exactly the event that empties a failing -RTC. Everything else about the fix is verified; this is the part that rests on -construction (an address needs no DNS, NTP carries no certificate) rather than -on having been seen work. - -Do this before relying on it away from home — the failure mode strands the -machine with no network and no way to look anything up. +What we're verifying: the coin cell, not the fix. The fix itself is already +demonstrated — on 2026-08-19 I wound the clock back 27 days with no DNS at all, +and chronyd reached the IP-addressed source and stepped it straight back. What a +cold boot adds is the hardware question: whether the RTC actually loses time +across a full power-down, which is the thing that started this. + +Read the outcome carefully, because only one branch is informative. An RTC time +that comes up wrong and then self-corrects tells you the cell is dying and the +fix is holding. An RTC that comes up correct tells you nothing about the +deadlock at all, only that the cell survived this particular night. - Confirm the drop-in is in place and chrony is using it (block below). - Shut all the way down — =poweroff=, not suspend, not reboot. The RTC only loses time when the machine is actually off. |
