homemaker-layout/experiments/results
Claude 84907694d0
Record what the crinkliness examination found: 39.13, 39.14, 39.15
39.13 -- the tail rescale, and its verdict. NULL, and not for want of power:
12 of 12 pairs byte-identical, both programmes, same trajectories. The failing
tail is 0.034% of corpus value, so making it orderable cannot move a search.
Kept (free, and k54 needs the region orderable) but recorded as correct and
inert, not as a fix. It also corrects gvb's premise: zero exposure is NOT
beyond the inner loop's reach -- perturbing division ratios alone moves the
zero-exposure set on 6-12 of 12 trials at +-25%, and in the direction wanted
(harbor s1 8 -> 6 buried leaves).

39.14 -- what the factor actually rewards. 1/crink is the room's depth from
its daylit wall in storey-heights, so the variable is sound and its fail
boundary (1.62h = 4.86 m) agrees with 38.3's independently-derived frontage
bound. The two-sided gaussian on it is not: the near side penalises surplus
daylight that edge_cost and outside_edge_cost already bill at 100 and 133.3
per m2, it has never once produced a fail (it needs crink > 21.5; corpus max
is 3.95), and its peak sits at a 2.5 m deep room -- an ordinary 4 m room
scores 0.395 and the corpus's realised median depth is 2.95 m. The search
built what it was paid for. A/B at pilot budget is underpowered rather than
null: the arms reach different layouts but the same fail counts.

39.15 -- the magic numbers. A sigma is not a preference, it is an acceptance
interval target +- 2.1460*sigma, so it decides failures. The blanket
hypothesis does not survive -- programme-house reaches 1 fail, structural on
two of three seeds. The specific one does, and it shows 39.1's CLEAN verdict
answered a weaker question: sweeping a spec's tolerance box asks whether SOME
shape is feasible, and all 67 pass, but at the DECLARED target area and aspect
harbor needs 7 corner rooms and maple 6, while health-centre and
programme-house need none. Within-programme, those codes fail 62% and 78% of
their instances against 26% and 31% for all others. Three declared quantities
are jointly contradictory and nothing said so; the resolution is an author
decision, not a retuned constant.

Also recorded: 82% of size fails are rooms larger than target, which is the
same shape of double-charge but explicitly NOT the same case -- size's upper
bound is the main brake on growth and must not be removed on the analogy.

405 passed, 72 skipped.

Closes homemaker-py-9gj, homemaker-py-u5q.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 15:11:11 +00:00
..
0wr_qpk_harbor_n24.tsv Harbor A/Bs at n=3 could never resolve their own margins 2026-08-29 19:34:30 +00:00
ab_9gj_crinkliness.csv Record what the crinkliness examination found: 39.13, 39.14, 39.15 2026-09-05 15:11:11 +00:00
ab_9gj_ramp.csv Make the crinkliness factor one-sided: stop billing the daylit wall twice 2026-09-05 07:27:02 +00:00
ab_cpsat_assign_20k_harbor_maple.log homemaker-py-2g7.5: full acceptance-criteria A/B (harbor+maple, 3 seeds, 20k budget) 2026-08-05 00:01:00 +01:00
ab_ssz_power.csv ssz: 61% of reported crinkliness fails are not defects; two corrections 2026-08-26 18:07:01 +00:00
ab_ssz_search.csv ssz: record the A/B result honestly -- not a pass at n=3 2026-08-26 17:57:11 +00:00
coldstart_baseline.tsv coldstart maple-court seed 2 @ 500000: 55 fails (11h/44s) 2026-09-03 10:06:06 +01:00
d86_1ph_iiofix.tsv Re-verify 1ph: the iio bug could never have touched it 2026-08-29 13:21:29 +00:00
d86_1ph_preiio.tsv Re-verify 1ph: the iio bug could never have touched it 2026-08-29 13:21:29 +00:00
ioe_1ph_current_objective.tsv Re-validate collapse_insearch's default under the current objective 2026-08-29 13:57:45 +00:00