The N=3 A/B (previous commits) found the precision-weighted shape
combination improved both example programmes (harbor-house -1.4%,
health-centre -13.9%), but N=3 is a thin sample by this project's own
standard (xyu/9yx use N=15). Two confirmations:
- N=15, plain search, budget=3000 (mirrors xyu/9yx's own protocol exactly):
both programmes trend NEGATIVE (harbor +6.1%, health-centre +6.6%,
Wilcoxon p=0.044)
- N=15, staged search, budget=20000 (true same-conditions replication --
identical to the original A/B except seed count): both programmes AGAIN
trend negative (harbor +6.6% p=0.15, health-centre +4.7% p=0.48)
The same-conditions replication disagrees with the original result's
direction on both programmes. Conclusion: the N=3 positive signal was
sampling noise, not a real effect -- health-centre's -13.9% was driven
substantially by one seed (71->43 fails) that didn't hold up.
multi_use stays default OFF and is not recommended even as a promising
lever -- this is a clean NULL, closing out both halves of §26's original
multi-use-leaves question (path a was NULL/NEGATIVE, path b is NULL after
replication). Mechanism itself is unchanged, complete, and fully tested.
DESIGN.md §33 rewritten with all three measurements and the honest verdict.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY