aboutsummaryrefslogtreecommitdiff
path: root/docs/design/2026-08-14-velox-reinstall-gaps-1.org
diff options
context:
space:
mode:
Diffstat (limited to 'docs/design/2026-08-14-velox-reinstall-gaps-1.org')
-rw-r--r--docs/design/2026-08-14-velox-reinstall-gaps-1.org137
1 files changed, 137 insertions, 0 deletions
diff --git a/docs/design/2026-08-14-velox-reinstall-gaps-1.org b/docs/design/2026-08-14-velox-reinstall-gaps-1.org
new file mode 100644
index 0000000..cf0d723
--- /dev/null
+++ b/docs/design/2026-08-14-velox-reinstall-gaps-1.org
@@ -0,0 +1,137 @@
+#+TITLE: What the velox reinstall left behind — four gaps the install could close
+#+AUTHOR: Craig Jennings
+
+* Heads-up: this was found from a .emacs.d session
+
+I opened a .emacs.d session on velox this morning, two days after the fresh
+Arch install, and the first thing it did was fail: there was no =.ai/=
+directory to read. Chasing that turned up four separate things the reinstall
+did not restore. Three I repaired from the session; one needs me at my phone.
+
+None of this is a .emacs.d bug. They are all install-side gaps, which is why
+they are landing in your inbox. Machine is velox; ratio was the reference for
+every comparison below.
+
+* Gap 1 — the gitignored tooling layer does not survive a reinstall
+
+=~/.emacs.d= was re-cloned on 2026-08-13. Git brought back every tracked file
+and none of the agent tooling, because =.gitignore= deliberately excludes it:
+=.ai/=, =.claude/=, =CLAUDE.md=, =todo.org=, and =inbox/= were all simply
+absent. That is the correct ignore policy — this repo relays to a public
+mirror — but it means a reinstall silently drops the entire working state of
+every gitignore-mode project.
+
+The damage on velox was total rather than partial: 374 files, 4.5 MB,
+including =todo.org= (556 KB) and 184 archived session files. Nothing carries
+it. Not git, not stow, not the bootstrap.
+
+I recovered it by rsyncing the set from ratio over the tailnet. Ratio was
+authoritative and velox held nothing, so there was no merge to adjudicate —
+which is luck, not design. Had velox held a few days of divergent state, this
+would have been a hand reconciliation. It has been one before: 2026-07-31, when
+the two machines' =.ai/= trees had forked to zero files in common.
+
+Worth knowing: this is fleet-general. Every project on the box that gitignores
+its =.ai/= has the same hole, not just =.emacs.d=.
+
+What the install could do: after cloning a project, check whether a sibling
+daily driver holds a =.ai/= for it, and offer to pull it across. Or at minimum,
+list the projects whose tooling layer is missing so the gap is visible on day
+one instead of at the first session that trips over it.
+
+* Gap 2 — stowed user timers come back linked but not enabled
+
+The unit files all arrived correctly through the dotfiles stow, symlinked into
+=~/.config/systemd/user/= and resolving fine. But being present is not being
+enabled, and the reinstall enabled only some of them:
+
+| unit | velox after reinstall | ratio |
+|---------------------------+-----------------------+----------|
+| calendar-sync.timer | enabled, active | enabled |
+| agenda-render-cache.timer | enabled, active | enabled |
+| roam-sync.timer | *linked, inactive* | enabled |
+| signal-receive.timer | *linked, inactive* | enabled |
+| emacs.service | linked, inactive | linked |
+
+=emacs.service= reads the same on both machines, so I take that one as
+intentional and left it alone. The other two are real drift: =systemctl --user
+enable= writes a =timers.target.wants= symlink into =~/.config/systemd/user/=,
+and that symlink is not stow-managed, so nothing in the dotfiles repo carries
+it. A stowed unit file is inert until something enables it.
+
+I enabled both with =systemctl --user enable --now=. Both fired immediately and
+exited clean, and both now show a next elapse.
+
+What the install could do: enable the units it stows, explicitly, as a named
+step. The inconsistency is the tell — two of four came back enabled, which
+suggests something enables a subset and nothing enumerates the rest.
+
+* Gap 3 — the roam clone was stale, and held a diff that would have destroyed data
+
+This one has an ordering constraint, so it matters more than its size suggests.
+
+velox's =~/org/roam= was ten commits behind ratio, stuck at the 2026-08-04
+auto-sync while ratio was at 2026-08-14 — a direct consequence of gap 2, since
+=roam-sync.timer= was never enabled here.
+
+The dangerous part: velox's clone also carried an *uncommitted* =inbox.org=
+that had been emptied. Seventeen deletions, file down to zero bytes, holding a
+pre-2026-08-04 state whose captures were long since processed on ratio.
+
+So the naive repair — enable =roam-sync.timer= and let it catch up — would have
+committed that emptying and pushed it, deleting the four live inbox items on
+ratio. The timer is the repo's only committer and it commits whatever it finds.
+
+I checked ratio's =inbox.org= first and confirmed it was a strict superset of
+velox's HEAD version (same three items plus an 2026-08-09 capture), which made
+the local change provably worthless. Then discarded it, fast-forwarded to
+=a411b43=, and only then enabled the timer. Clone is clean and current, first
+sync ran green.
+
+What the install could do: if it ever enables =roam-sync= on a rebuilt machine,
+reconcile the clone *before* enabling, not after. An auto-committing timer
+pointed at a stale dirty clone is a data-loss path, and the failure is silent
+and remote — it lands on the *other* machine.
+
+* Gap 4 — signal-cli lost its registration, and that breaks the whole fleet
+
+=signal-receive.service= ran for the first time and reported:
+
+: signal-receive: +15045173983 not registered on this machine — nothing to do
+
+velox's signal-cli data dir holds a 39-byte empty =accounts.json=. Ratio still
+has both numbers. So the reinstall wiped the registration, and per the design
+notes velox was supposed to be the *primary* — ratio is the linked device.
+
+The effect is wider than velox, because of how =agent-text= dispatches: if the
+local signal-cli holds the account it sends directly, otherwise it ssh-relays to
+a hardcoded velox. Velox no longer holds it, so a send from here relays to
+itself and fails; a send from any third machine relays to velox and fails the
+same way. Only ratio still works, and only via the direct branch. The error text
+blames "velox down or unreachable", which is misleading — velox is up and on the
+tailnet, it just is not registered.
+
+This is the one I could not repair from the session: re-linking needs me at my
+phone (Signal → Settings → Linked Devices, scanning the QR from =signal-cli
+link -n velox=). Filed in .emacs.d's todo.org as [#B].
+
+What the install could do: verify =signal-cli listAccounts= is non-empty after a
+rebuild and say so loudly if it is not. Silent loss of the phone channel is
+exactly the kind of thing nobody notices until the page that mattered never
+arrives.
+
+* Summary of what I changed on velox
+
+- Restored =.ai/=, =.claude/=, =CLAUDE.md=, =todo.org=, =inbox/= to
+ =~/.emacs.d= by rsync from ratio.
+- Discarded the stale local =inbox.org= diff in =~/org/roam= and fast-forwarded
+ the clone to current.
+- Enabled and started =roam-sync.timer= and =signal-receive.timer=.
+
+Left alone, deliberately: =emacs.service= (matches ratio), and velox's Signal
+registration (needs the phone).
+
+One unrelated thing I noticed while comparing the machines: ratio's signal-cli
+warns its messages were last received twelve days ago, even though its
+=signal-receive.timer= is enabled and active. That may be nothing, but the
+receive cadence there is worth a look.