diff options
| author | Craig Jennings <c@cjennings.net> | 2026-07-27 13:15:14 -0500 |
|---|---|---|
| committer | Craig Jennings <c@cjennings.net> | 2026-07-27 13:15:14 -0500 |
| commit | 79ed3b09a9ee2a63fb55d2354aa0c77ea24c6efa (patch) | |
| tree | 928273e02ea3d40a2aee2bdde4a8f922aec9762f /languages/go/tests | |
| parent | 7ea1d7b1402eb68a13479f8073b84819c1d59ec8 (diff) | |
| download | rulesets-79ed3b09a9ee2a63fb55d2354aa0c77ea24c6efa.tar.gz rulesets-79ed3b09a9ee2a63fb55d2354aa0c77ea24c6efa.zip | |
docs: add context-engineering rightsizing analysis and rollout
I wrote three working documents while reviewing the Claude 5 context-engineering post, the Opus 5 prompting guide, and the Fable field guide against what this repo ships downstream.
In proposals.org I measured the always-loaded surface at 32,123 words and ranked six changes, plus two places where a post contradicts something we arrived at deliberately. I also reread the rules as your prompts rather than as agent context, which is where the sharper finding is: 41 execution and hygiene workflows against 6 discovery and design ones, on a system whose bottleneck has moved.
In rollout.org I phased the work and named the seven decisions that gate it. My lead finding is that the harness system prompt already carries most of what the Opus 5 guide recommends adding, so the posts' value here is subtractive. Applying them additively would make the duplication worse.
In metrics.org I split the posts into separable claims and marked which are testable here and which are judgment calls, rather than inventing a metric for the ones an eval harness would be needed to settle. I also stated the pilot's stop conditions in advance, including a zero-tolerance threshold on the load-bearing rules, so I can't renegotiate that threshold later under pressure to make a phase succeed.
Diffstat (limited to 'languages/go/tests')
0 files changed, 0 insertions, 0 deletions
