homemaker-layout/experiments/ab_9gj_crinkliness.py

272 lines
12 KiB
Python
Raw Permalink Normal View History

Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
"""Search A/B for the crinkliness reformulation (`homemaker-py-9gj`).
Two independent changes to `quality_uncrinkliness`, either or both:
crinkliness_tail="ramp" below FAIL_THRESHOLD. The factor evaluates a
gaussian at `x = 1/crink`, whose exponent
grows like 1/crink^2, so the failing tail
spans quality 1e-300..1e-1 -- all of it
numerically zero beside a passing leaf's ~1.
The ramp makes that tail a straight line in
crinkliness. (DESIGN.md §39.13. Measured a
complete null on its own: 12 of 12 pairs
byte-identical.)
crinkliness_shape="daylight" above it. `1/crink` is the room's mean depth
from its daylit wall in storey-heights, and
the stock gaussian is TWO-sided on it, so a
room with more daylight than target is
penalised for it -- while the cost model
already charges that wall via
`exterior_wall`/`boundary_wall`. "daylight"
clips that side to 1.0. (DESIGN.md §39.14.)
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
**Why stock scoring is valid here** (the §38.9 trap, and the one case the
`9gj` bead flags as exempt): the ramp is continuous at FAIL_THRESHOLD and
strictly below it, so no leaf changes which side of the threshold it is on.
The fail set is byte-identical on every corpus artefact -- asserted in
`tests/test_fitness_crinkliness_tail.py`, not assumed here. An arm therefore
cannot win by deleting a fail category, and both arms are scored under stock.
**Two experiment shapes**, because they answer different questions:
--start plateau (default) seed each run from that programme's
`coldstart-500000-s<k>.dom`. This is the ESCAPE
test the bead asks for: the ramp has signal only
where a search has already built partially-lit
rooms, and the corpus `init.dom` files show a
+0.000% score delta -- there is nothing for it to
grade at the start of a search.
--start init cold start, for comparison.
Pairing is on (starting layout, RNG seed), so `--seeds N` over `--starts M`
gives N*M paired samples per programme.
Usage::
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
python experiments/ab_9gj_crinkliness.py --budget 8000 --seeds 2 --starts 3
python experiments/ab_9gj_crinkliness.py --start init --budget 8000 --seeds 6
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
Sharding, because one run is minutes and the job list is 4x that. Each shard
keeps BOTH arms of a pair together, so the two halves of a comparison never
land on differently-loaded processes::
for i in 0 1 2 3; do
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
python experiments/ab_9gj_crinkliness.py --budget 8000 --seeds 2 --starts 3 \
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
--shard $i --nshards 4 &
done; wait
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
python experiments/ab_9gj_crinkliness.py --report
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
A POWERED run needs more than the in-session pilot could afford. §39.12 puts
harbor's minimum detectable difference at n=3 at 13.7 fails; the plateau-escape
deltas here are single-digit, so budget for n >= 8 pairs per programme
(`--seeds 3 --starts 3` gives 9) and a budget large enough for either arm to
move at all -- the pilot's 8000 evals is 1.6% of what produced the plateau::
for i in $(seq 0 3); do
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
python experiments/ab_9gj_crinkliness.py --budget 100000 --seeds 3 --starts 3 \
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
--shard $i --nshards 4 &
done; wait
"""
from __future__ import annotations
import argparse
import collections
import copy
import csv
import time
from pathlib import Path
from homemaker_layout import dom as dom_mod
from homemaker_layout import driver, fitness
CORPUS = ["examples/harbor-house", "examples/maple-court"]
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
# Each arm names a crinkliness configuration. "stock" must stay first: it is
# the baseline every other arm is paired against, and the yardstick all arms
# are SCORED under.
ARM_CONF = {
"stock": {},
"ramp": {"crinkliness_tail": "ramp"},
"daylight": {"crinkliness_shape": "daylight"},
"daylight+ramp": {"crinkliness_shape": "daylight",
"crinkliness_tail": "ramp"},
Quality aggregation: divide out how many questions a leaf was asked 39.17 left the search's storey choice unexplained and blamed value_rate. It is not the rate, or not only. Measured over the twelve baseline runs, value/cost by leaf kind: outside ground 7.40, roof terrace 2.69, room 0.34, circulation 0.02. A terrace returns 2.7x its cost where a room returns a third of it, so filling upper storeys with terrace is not the search leaving value on the table -- it is by a wide margin the most profitable thing the objective offers. 7% of the corpus area produces 32% of its value. Most of that gap is mean quality: 0.986 for a terrace against 0.223 for a room. Quality is a PRODUCT of factors and the kinds are not asked the same number of questions -- an outside leaf is exempt from size, crinkliness and access, so 3 of 7 factors can ever bite it against a room's 6. Each exemption is individually right (no programme size target; uncovered outside is lit by definition; ground-level outside needs no access). The consequence is not: a leaf exempt from the two harshest factors out-scores one judged on them and doing well, purely by not being asked, and quality multiplies the value rate. Stated generally, and this is not about outside space: under a product, adding any new quality criterion mechanically devalues every leaf it applies to, including leaves that score 1.0 on it. The objective's scale should not depend on how many things it measures. quality_aggregate="geometric_mean" (default OFF, "product" is stock) divides that out. Computed in log space so six small factors cannot underflow the product before the root is taken; a zero factor still gives zero, so a fully buried leaf is worth nothing either way. Telling "exempt" from "asked and scored 1.0" needs factor_is_asked, which restates conditions that live inside the quality_* methods. That duplication can drift, so tests/test_fitness_aggregate.py pins it against every leaf in the corpus: wherever the predicate says exempt, the factor really is 1.0. Fail set byte-identical everywhere, and for a stronger reason than 39.13/39.14 had: evaluate_leaf emits each fail from the factor itself before anything is combined, so no aggregation can move one. Score effect +37% to +169%, reaching all four programmes where the crinkliness changes reached two; room value/cost 0.34 -> 0.66, circulation 0.02 -> 0.07. Deliberately not fixed: a terrace still out-earns a room 4:1, which is the rates (value_supported = value_inside = 300 against costs of 110 and 200), not the aggregation. That is a design judgement for the programme author, and 39.16 is a standing reminder that "this inherited constant looks wrong" has been wrong twice already in this section. Left open on ecx with the numbers. A/B running; verdict to follow. Refs homemaker-py-ecx. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 18:57:00 +00:00
# homemaker-py-ecx (§39.18): orthogonal to the crinkliness arms -- it
# changes how the factors are COMBINED, not what any of them says.
"geomean": {"quality_aggregate": "geometric_mean"},
"geomean+daylight": {"quality_aggregate": "geometric_mean",
"crinkliness_shape": "daylight",
"crinkliness_tail": "ramp"},
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
}
ARMS = ["stock", "daylight", "daylight+ramp"]
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
def _with_arm(arm: str):
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
"""Patch `fitness.load_config` so every evaluator built during the run --
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
the driver's, the inner loop's, the seeder's -- sees the arm's overrides.
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
`driver.search` has no parameter for them and `driver._fitness_for` is
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
lru_cached, so the cache is cleared around the patch (see ab_ssz_search).
"""
orig = fitness.load_config
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
arm_ov = ARM_CONF[arm]
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
def patched(directory, overrides=None):
ov = dict(overrides or {})
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
ov.update(arm_ov)
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
return orig(directory, overrides=ov)
return orig, patched
def tiers(fails) -> tuple[int, int]:
c = collections.Counter(fitness.classify_fail_tier(f) for f in fails)
return c["hard"], c["soft"]
def run_arm(progdir: str, start: Path, seed: int, tail: str, budget: int,
child_budget: int) -> dict:
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
orig, patched = _with_arm(tail)
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
fitness.load_config = patched
driver._fitness_for.cache_clear()
t0 = time.perf_counter()
try:
res = driver.search(dom_mod.load(str(start)), progdir, budget=budget,
seed=seed, child_budget=child_budget, n_workers=1)
root = copy.deepcopy(res.best.root)
finally:
fitness.load_config = orig
driver._fitness_for.cache_clear()
conf, cost = orig(progdir, overrides={"leaf_sharing": True,
"collapse_insearch": True})
score, fails = fitness.Fitness(conf, cost).score_with_fails(copy.deepcopy(root))
h, s = tiers(fails)
return dict(programme=Path(progdir).name, start=start.name, seed=seed,
tail=tail, hard=h, soft=s, total=h + s, score=score,
elapsed_s=round(time.perf_counter() - t0, 1))
def starting_points(progdir: str, kind: str, n: int) -> list[Path]:
"""The distinct layouts to start from.
`--starts` applies to plateau mode only: there is exactly one `init.dom`,
and repeating it would give several jobs the same (start, seed) pairing key,
which `_report` would silently collapse to one. In init mode the RNG seed
is the only sampling dimension, so use `--seeds`.
"""
d = Path(progdir)
if kind == "init":
return [d / "init.dom"]
found = sorted(d.glob("coldstart-500000-s*.dom"))[:n]
if not found:
raise SystemExit(f"no coldstart-500000-s*.dom in {d} for --start plateau")
return found
def main() -> None:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--budget", type=int, default=8000)
ap.add_argument("--child-budget", type=int, default=80)
ap.add_argument("--seeds", type=int, default=2, help="RNG seeds per start")
ap.add_argument("--starts", type=int, default=3,
help="plateau layouts to start from (plateau mode only)")
ap.add_argument("--start", choices=("plateau", "init"), default="plateau")
ap.add_argument("--corpus", nargs="+", default=CORPUS)
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
ap.add_argument("--arms", nargs="+", default=ARMS,
choices=sorted(ARM_CONF), help="first arm is the baseline")
ap.add_argument("--out", default="experiments/results/ab_9gj_crinkliness.csv")
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
ap.add_argument("--shard", type=int, default=0,
help="run only jobs i where i %% nshards == shard")
ap.add_argument("--nshards", type=int, default=1,
help="split the job list across N processes; each writes "
"<out>.shard<i>. Use --report to merge and analyse.")
ap.add_argument("--report", action="store_true",
help="merge <out>.shard* (or <out>) and print the analysis "
"only -- runs nothing")
args = ap.parse_args()
out = Path(args.out)
out.parent.mkdir(parents=True, exist_ok=True)
if args.report:
rows = []
shards = sorted(out.parent.glob(out.name + ".shard*")) or (
[out] if out.exists() else [])
for sh in shards:
with sh.open() as fh:
for r in csv.DictReader(fh):
for k in ("total", "hard", "soft", "seed"):
r[k] = int(r[k])
r["score"] = float(r["score"])
rows.append(r)
if len(shards) > 1:
with out.open("w", newline="") as fh:
w = csv.DictWriter(fh, fieldnames=list(rows[0]))
w.writeheader()
w.writerows(rows)
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
_report(rows, args.corpus, args.arms)
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
print(f"\nmerged {len(rows)} runs from {len(shards)} file(s) into {out}")
return
# Build the whole job list first so sharding is deterministic and every
# (start, seed) pair keeps BOTH arms in the same shard -- the comparison is
# paired, and splitting a pair across processes would let machine load
# differ between the two halves of one pair.
jobs = []
for progdir in args.corpus:
for start in starting_points(progdir, args.start, args.starts):
for seed in range(args.seeds):
jobs.append((progdir, start, seed))
rows: list[dict] = []
dest = (out if args.nshards == 1
else out.with_name(out.name + f".shard{args.shard}"))
for i, (progdir, start, seed) in enumerate(jobs):
if i % args.nshards != args.shard:
continue
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
for tail in args.arms:
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
r = run_arm(progdir, start, seed, tail, args.budget,
args.child_budget)
rows.append(r)
print(f" {r['programme']:<14} {start.name:<26} seed={seed} "
f"{tail:<9} {r['hard']}h/{r['soft']}s = {r['total']:3d} "
f"score {r['score']:.4g} {r['elapsed_s']}s", flush=True)
with dest.open("w", newline="") as fh:
w = csv.DictWriter(fh, fieldnames=list(rows[0]))
w.writeheader()
w.writerows(rows)
if args.nshards == 1:
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
_report(rows, args.corpus, args.arms)
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
print(f"\nwrote {dest}")
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
def _report(rows, corpus, arms=None) -> None:
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
"""homemaker-py-tco: state what this N could resolve, beside the result."""
from ab_report import format_report, paired_report
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
seen = []
for r in rows:
if r["tail"] not in seen:
seen.append(r["tail"])
arms = [a for a in (arms or seen) if a in seen] or seen
base = arms[0]
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
for progdir in corpus:
name = Path(progdir).name
by: dict[tuple, dict] = {}
for r in rows:
if r["programme"] == name:
by.setdefault((r["start"], r["seed"]), {})[r["tail"]] = r["total"]
Make the crinkliness factor one-sided: stop billing the daylit wall twice The tail rescale shipped in cd392e7 is a measured NULL as a search intervention -- 12 of 12 pairs byte-identical on harbor and maple, 8000 evals from a plateau, not merely underpowered. Of course it is: the whole failing tail is 0.034% of corpus value. Looking at the rest of the factor, prompted by the owner, found something much larger above the threshold. crink = area_outside/area = (L*h)/A, so 1/crink = A/(L*h) is the room's mean depth from its daylit wall in storey-heights. That is the right variable for a daylight rule, and the fail boundary it implies (1/crink = 1.62, i.e. 4.86 m at h=3) is a sensible one that agrees with 38.3's frontage bound derived independently. What is wrong is hanging a TWO-sided gaussian on it: * The near side penalises a room for having MORE daylit wall than target -- while leaf_cost's siblings edge_cost and outside_edge_cost already charge that same wall at exterior_wall=100 and boundary_wall=133.3 per m2. The wall is billed once in cost and again as lost value. * It never earns its keep as a failure either: the over-exposed branch only reaches FAIL_THRESHOLD above crinkliness 21.5, and the corpus maximum is 3.95. It has never produced a single fail; it only removes value. * 133 of the 318 passing graded leaves in the 500k baseline (42%) sit on that side, mean quality 0.810. crinkliness_shape="daylight" (default OFF, "gaussian" is stock) clips it: a room shallower than the gaussian's peak scores 1.0, because daylight is a sufficiency requirement and surplus is the cost model's business, not this factor's. Clipping at the PEAK rather than at FAIL_THRESHOLD is deliberate -- it keeps the factor continuous and preserves the graded approach to the daylight limit, where clipping at the threshold would put a 10x cliff on the exact boundary the 0.5**n fail multiplier already steps on. Fail set byte-identical on all 21 corpus artefacts for all four shape/tail combinations, so stock stays a valid yardstick for every arm. Area-weighted crinkliness quality 0.480 -> 0.513, leaf quality product 0.2722 -> 0.2831; per-artefact score +0.2%..+19.6%, and unlike the ramp it reaches health-centre and programme-house, where the tail change was 0.000%. Note "daylight" clips the OPPOSITE side from 38.1's superseded compact_ok, which forgives being buried; composing either with those modes is refused. ab_9gj_ramp.py becomes ab_9gj_crinkliness.py and takes named arms, since it now covers both changes; its first arm is the baseline and the yardstick. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 07:27:02 +00:00
for arm in arms[1:]:
keys = sorted(k for k, v in by.items() if base in v and arm in v)
if len(keys) < 2:
continue
print(f"\n--- {name}: {arm} vs {base} (stock-scored total fails) ---")
print(format_report(paired_report(
[by[k][base] for k in keys], [by[k][arm] for k in keys],
base, arm)))
Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
if __name__ == "__main__":
main()