aboutsummaryrefslogtreecommitdiff
path: root/working/context-engineering-rightsizing/rollout.org
diff options
context:
space:
mode:
authorCraig Jennings <c@cjennings.net>2026-07-27 13:15:14 -0500
committerCraig Jennings <c@cjennings.net>2026-07-27 13:15:14 -0500
commit79ed3b09a9ee2a63fb55d2354aa0c77ea24c6efa (patch)
tree928273e02ea3d40a2aee2bdde4a8f922aec9762f /working/context-engineering-rightsizing/rollout.org
parent7ea1d7b1402eb68a13479f8073b84819c1d59ec8 (diff)
downloadrulesets-79ed3b09a9ee2a63fb55d2354aa0c77ea24c6efa.tar.gz
rulesets-79ed3b09a9ee2a63fb55d2354aa0c77ea24c6efa.zip
docs: add context-engineering rightsizing analysis and rollout
I wrote three working documents while reviewing the Claude 5 context-engineering post, the Opus 5 prompting guide, and the Fable field guide against what this repo ships downstream. In proposals.org I measured the always-loaded surface at 32,123 words and ranked six changes, plus two places where a post contradicts something we arrived at deliberately. I also reread the rules as your prompts rather than as agent context, which is where the sharper finding is: 41 execution and hygiene workflows against 6 discovery and design ones, on a system whose bottleneck has moved. In rollout.org I phased the work and named the seven decisions that gate it. My lead finding is that the harness system prompt already carries most of what the Opus 5 guide recommends adding, so the posts' value here is subtractive. Applying them additively would make the duplication worse. In metrics.org I split the posts into separable claims and marked which are testable here and which are judgment calls, rather than inventing a metric for the ones an eval harness would be needed to settle. I also stated the pilot's stop conditions in advance, including a zero-tolerance threshold on the load-bearing rules, so I can't renegotiate that threshold later under pressure to make a phase succeed.
Diffstat (limited to 'working/context-engineering-rightsizing/rollout.org')
-rw-r--r--working/context-engineering-rightsizing/rollout.org306
1 files changed, 306 insertions, 0 deletions
diff --git a/working/context-engineering-rightsizing/rollout.org b/working/context-engineering-rightsizing/rollout.org
new file mode 100644
index 0000000..4afc27f
--- /dev/null
+++ b/working/context-engineering-rightsizing/rollout.org
@@ -0,0 +1,306 @@
+#+TITLE: Context-Engineering Rightsizing — Rollout Schedule
+#+AUTHOR: Craig Jennings
+#+DATE: 2026-07-27
+
+* Why this is phased rather than done in one pass
+
+The change is a bet that on-demand loading fires as reliably as always-resident
+text. That bet is cheap to test and expensive to assume. Every phase below
+either produces evidence or spends evidence already earned. Nothing
+load-bearing moves before the mechanism has been watched working.
+
+The second reason is blast radius. Everything here rides the template sync into
+every project on its next startup, so a bad phase is not contained to this
+repo. The early phases are chosen so a failure is visible and harmless.
+
+* A finding that changes the plan
+
+The harness system prompt already carries most of what the Opus 5 guide
+recommends adding. Its task-scope block, its correction-narration block, and
+its subagent-delegation cap are present nearly verbatim. The
+context-engineering post's replacement comment guidance ("write code that reads
+like the surrounding code") is present as the post's own new wording.
+
+Two consequences:
+
+1. *Do not "apply the posts" by adding their suggested prompt blocks.* They are
+ already live. Adding them to =claude-rules/= would create exactly the
+ duplicate-and-conflict problem the first post opens with, while making the
+ token count worse.
+2. *There is a third deduplication axis.* The proposals named =protocols.org=
+ against =claude-rules/=. There is also =claude-rules/= against the harness
+ system prompt, and that one is invisible from inside the repo. Any rule that
+ restates harness guidance is pure cost.
+
+This is why the posts' value here is subtractive, not additive.
+
+* Phase 0 — Free wins (this session, no dependencies)
+
+*Scope.* Three items that interact with nothing.
+
+1. Fix the review-finding pre-filter (P2). =review-code/SKILL.md= lines 251 and
+ 434 tell the reviewer to drop low-confidence findings before reporting.
+ Replace with report-everything-labelled, filter in a named second pass.
+2. Run =/doctor=. The context-engineering post says Anthropic shipped these
+ practices as a command that rightsizes skills and =CLAUDE.md=. Its output is
+ free evidence, and it may disagree with this plan, which is worth knowing
+ before executing it.
+3. Record the harness-overlap finding above where it will be seen at the moment
+ it matters — a note in the rules index, not buried in this document.
+
+*Reasoning.* None of these depend on the pilot's outcome, and item 2 could
+change the plan.
+
+*Decision needed:* none.
+
+*Success criteria.* Review skill reports with confidence labels and a separate
+filter step. =/doctor= output read and reconciled against this schedule.
+
+*Rollback.* Single revert; nothing downstream depends on it.
+
+* Phase 1 — The pilot migration (next working session)
+
+*Scope.* Six rule files move from always-loaded to on-demand skills. Roughly
+3,000 words, about 12% of the rules surface.
+
+| File | Words | How a silent miss would be caught |
+|-----------------------+-------+-------------------------------------------|
+| =org-tables.md= | 464 | =lint-org= checker =org-table-standard= |
+|-----------------------+-------+-------------------------------------------|
+| =docs-lifecycle.md= | 582 | spec status-board grep; =lint-org= checkers |
+|-----------------------+-------+-------------------------------------------|
+| =ui-prototyping.md= | 696 | =spec-review= verifies the process ran |
+|-----------------------+-------+-------------------------------------------|
+| =keybinding-display.md= | 505 | you see the wrong format immediately |
+|-----------------------+-------+-------------------------------------------|
+| =desktop-capture.md= | 458 | a window lands on your active workspace |
+|-----------------------+-------+-------------------------------------------|
+| =patterns.md= | 291 | already only a pointer; nothing to miss |
+|-----------------------+-------+-------------------------------------------|
+
+*Reasoning — the selection rule matters more than the list.* These were not
+picked for being small or cheap. They were picked because *a failure to fire is
+detectable*. Four have a mechanical checker or workflow gate that catches the
+miss; two produce an error you see within seconds. That is what makes the pilot
+an experiment rather than a hope.
+
+=daily-drivers.md= and =emacs.md= were considered and held back. Both are
+low-risk in content, but a miss on either surfaces slowly — as drift on the
+other machine, or as a stale daemon — so neither would tell us anything within
+the trial window.
+
+*Decisions needed.*
+
+- *D1 — Confirm the pilot set.* Six files as listed, or trim further. My
+ recommendation is the six: fewer than that and the trial may not exercise the
+ mechanism enough to learn from.
+- *D2 — Does the always-loaded core carry a skill index?* A one-line-per-skill
+ list naming what exists and when it applies. It costs perhaps 200 words and
+ should materially improve trigger reliability, since the model can see that a
+ rule exists even when its content isn't loaded. My recommendation is yes, and
+ the pilot is the right place to test whether the index is what does the work.
+
+*Success criteria.* Always-loaded surface drops to about 29,000 words. All six
+skills exist with trigger descriptions. Suite green, sync clean, every project
+picks up the change on next startup without drift.
+
+*Rollback.* One revert restores the files to =claude-rules/=. The skills can
+stay in place harmlessly.
+
+* Phase 2 — Live trial (one week of real sessions, no work required)
+
+*Scope.* Use the system normally. Do not compensate for the pilot by mentioning
+the moved rules — that would invalidate the result.
+
+*Reasoning.* This is the phase that buys everything after it. The question is
+narrow and answerable: when work touches one of the six domains, does the skill
+fire without prompting?
+
+*What gets recorded.* Each session that touches a pilot domain notes one line
+in the session log: which domain, whether the skill fired, and whether the
+detector caught anything. At the end of the week that's a short table rather
+than an impression.
+
+*Decision needed:* none during the trial.
+
+*Success criteria.* Defined in advance so the verdict isn't argued after the
+fact:
+
+- *Pass* — no detector fires on a moved rule, or any miss is caught by its
+ detector and corrected in the same session.
+- *Fail* — a miss reaches a commit, or the same rule misses twice.
+- *Ambiguous* — no session touched the domain. That is not a pass; extend the
+ window or move a rule that gets exercised more.
+
+*Rollback.* Revert on a Fail, and the plan stops at Phase 0.
+
+* Phase 3 — Go/no-go and the gate separation (one session)
+
+*Scope.* Read the trial table, decide whether the mechanism is trusted, and
+separate the approval gates.
+
+*Reasoning.* The gate separation is the highest-leverage input in the whole
+plan, and it sits here rather than earlier for one reason: if Phase 2 fails,
+the question is moot, because nothing more moves either way.
+
+*Decision needed.*
+
+- *D3 — Which gates are preference and which are guardrail?* Every approval gate
+ in the system reads identically in the files. Some you would keep even if the
+ agent were perfectly reliable, because you want to see what goes out under
+ your name. Others exist because the worst case used to be worse. The list to
+ walk: the publish approval gate, the inbox shared-asset approval, the
+ spec-review flip, the wrap certification, and the no-approvals mode's carve
+ outs.
+
+ Preference gates are untouchable and stay always-loaded regardless of length.
+ Guardrail gates are candidates for relaxation on the posts' argument. I can
+ prepare the list with my read of each, but the answers are yours.
+
+*Success criteria.* Every gate labelled. The label determines what Phase 4 may
+move.
+
+* Phase 4 — The load-bearing files (two or three sessions)
+
+*Scope.* =commits.md= (5,561), =todo-format.md= (4,494), =testing.md= (2,824),
+=working-files.md= (950), =subagents.md= (1,041). About 15,000 words, the bulk
+of the remaining surface.
+
+The split within each file is by blast radius, not by length. =commits.md= is
+the worked example: the AI-attribution ban and the content-scope rule stay
+always-loaded and get shorter, while the publish flow, the message format, and
+the voice mechanics become the publish skill that loads when a commit is in
+play.
+
+*Reasoning.* This is where the token math actually pays. It runs last because
+it is where a silent miss is expensive: an unattributed commit, a leaked path,
+an ungraded task.
+
+*Decision needed.*
+
+- *D4 — Resolve the =verification.md= conflict (C1).* The Opus 5 guide says
+ explicit verification instructions cause over-verification and should be
+ removed. Your standing direction is never guess, always check. My read is
+ that these are compatible because they address different things: the honesty
+ core (don't claim a green suite you didn't run) stays, and the process
+ injection (green baseline before starting, suite as its own step) moves into
+ the publish skill. But it is your rule and your call, and this decision blocks
+ =commits.md= moving because the two files reference each other.
+
+*Success criteria.* Always-loaded surface under about 8,000 words. Two full
+weeks of sessions with no attribution, scope, or grading miss.
+
+*Rollback.* Per-file, since each moves independently.
+
+* Phase 5 — Deduplication (one session)
+
+*Scope.* Three axes, in increasing order of payoff:
+
+1. =protocols.org= against =claude-rules/= — the cross-project boundary,
+ working-files, AI-attribution, and inbox cadence are each stated twice.
+2. =claude-rules/= against the harness system prompt — the finding at the top
+ of this document. Invisible from inside the repo and therefore never audited.
+3. Within =claude-rules/= — rules that restate each other.
+
+*Decision needed.*
+
+- *D5 — Which surface owns each duplicated rule.* Generally the more specific
+ one should own the content and the more general should carry a pointer, but
+ there are cases where the reverse is right.
+
+*Success criteria.* Each rule stated once. A stated rule for where new rules go,
+so the duplication doesn't regrow.
+
+* Phase 6 — Terseness and positive framing (rides along with Phases 4 and 5)
+
+*Scope.* Rewrite prohibitions that aren't hard invariants as positive
+descriptions. Cut the files that don't practice what they demand: =commits.md=
+arguing terseness at 5,561 words, =testing.md= arguing an eight-row table
+against rationalizations the model no longer needs talked out of, 591 bold
+markers in files that ban bold in output.
+
+*Reasoning.* Not a separate campaign. Every file opened in Phases 4 and 5 gets
+this pass while it's open, because doing it separately means editing everything
+twice.
+
+*Decision needed:* none. This is style, and the voice skill already owns the
+standard.
+
+* Phase 7 — Discovery practices (after the surface is down)
+
+*Scope.* The field guide's missing practices: a blind-spot pass, an interview
+pattern, an implementation-notes convention for long builds, and possibly the
+quiz.
+
+*Reasoning.* Deliberately last, for two reasons. It adds surface, which fights
+every phase before it, so it should land only once there is room. And the
+seven-to-one execution-to-discovery ratio is the finding most likely to change
+how the system actually feels to use, which makes it worth doing carefully
+rather than early.
+
+*Decision needed.*
+
+- *D6 — Which practices you actually want.* I have low confidence on the quiz
+ fitting how you work, and medium-high on the rest.
+
+* Phase 8 — Effort calibration (independent, any time after Phase 2)
+
+*Scope.* Set effort levels for the unattended loops: sentry's hourly fires,
+work-the-backlog, the no-approvals speedrun.
+
+*Reasoning.* Independent of everything else and safe to run in parallel. The
+guide says to sweep effort on your own evals rather than trusting carried-over
+defaults; we have no evals, so this is a trial and a judgment.
+
+*Decision needed.*
+
+- *D7 — Accepted quality floor for unattended passes.* A sentry sweep that runs
+ cheaper but misses one finding per night may be a good trade or a bad one.
+ That's a preference, not a measurement.
+
+* Ongoing — The consistency sweep
+
+Runs alongside, not as a phase. Each file opened in Phases 4 through 6 gets read
+for contradictions and stale facts while it's open, and findings go to a running
+list rather than being fixed opportunistically. The expensive item — instructions
+that contradict each other across files — is what the first post opens with and
+what this whole exercise is downstream of.
+
+* Decisions, collected
+
+| ID | Decision | Needed by |
+|----+----------------------------------------------+-----------|
+| D1 | Confirm the six-file pilot set | Phase 1 |
+|----+----------------------------------------------+-----------|
+| D2 | Skill index in the always-loaded core? | Phase 1 |
+|----+----------------------------------------------+-----------|
+| D3 | Which gates are preference vs guardrail | Phase 3 |
+|----+----------------------------------------------+-----------|
+| D4 | Resolve the verification.md conflict | Phase 4 |
+|----+----------------------------------------------+-----------|
+| D5 | Which surface owns each duplicated rule | Phase 5 |
+|----+----------------------------------------------+-----------|
+| D6 | Which discovery practices you want | Phase 7 |
+|----+----------------------------------------------+-----------|
+| D7 | Quality floor for unattended passes | Phase 8 |
+|----+----------------------------------------------+-----------|
+
+Only D1 and D2 are needed to start.
+
+* The number this is aiming at
+
+| Stage | Always-loaded words |
+|-------------------+---------------------|
+| Today | 32,123 |
+|-------------------+---------------------|
+| After Phase 1 | 29,100 |
+|-------------------+---------------------|
+| After Phase 4 | under 8,000 |
+|-------------------+---------------------|
+| After Phase 5 | under 6,000 |
+|-------------------+---------------------|
+
+Roughly an 80% reduction, which lands near what Anthropic reported. That
+symmetry is a coincidence worth distrusting rather than aiming for: the target
+is whatever survives the blast-radius test, and if that turns out to be 12,000
+words then 12,000 words is the right answer.