Commit graph

107 commits

Author SHA1 Message Date
85c1183d4c homemaker-py-2g7.4: shape-curve DP prototype (Otten/Stockmeyer) — PASS
Prototype + validation for an exact size/width/proportion feasibility DP
over a frozen slicing topology, replacing the ~80-200 eval Nelder-Mead
inner loop's approximate answer to the same question with one bottom-up
pass (experiments/shapecurve_spike.py). Leaf feasible regions are exact
FAIL_THRESHOLD-inversions of fitness.py's quality_size/width/proportion;
internal-node composition runs on a shared discretised grid.

Validated on harbor-house-l0 (experiments/validate_shapecurve.py, 200
random topologies vs NM minimising shape-fail-count directly): 99.0%
agreement (0 false negatives), 93.6x speedup at grid_n=150, plot-level
bbox approximation error quantified at +7.5% (root-causing both observed
false positives). All three acceptance criteria cleared -- see DESIGN.md
§37.2 for full results and the caveats/scope not covered (multi-storey,
leaf_sharing/co_type, true skew-quad regions). Kept as a reference spike,
same status as experiments/autodiff_spike.py (§34); production wiring
into driver.py filed as homemaker-py-6xh.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-02 23:43:30 +01:00
8efdc02fd9 homemaker-py-2g7.3: hard/soft fail tiering behind --use-tiers flag
Splits the flat outer-search comparator (-n_fails, fitness) into a tiered
(-n_hard, -n_soft, fitness) so search budget stops being spent polishing
SOFT shape fails (crinkliness/proportion/size/width/edge-too-long/
staircase-volume) while HARD structural fails (missing space, wrong/
required level, level/circulation/vertical connectivity, adjacency,
stairs, covered-outside, storey limits, public access) remain unfixed.

fitness.classify_fail_tier/tier_counts classify every fail string emitted
across fitness.py and graph.py, raising on anything unrecognised so new
fail sites must declare a tier. Validated against all real fail strings in
the checked-in corpus plus every fail-emission call site read from source.

driver.Individual gains n_hard/n_soft (populated from innerloop.Result.
fail_lines); search(use_tiers=...) swaps the comparator when set (default
off, so existing runs are unaffected — inner-loop 0.5^n cliff untouched).
evolve.py exposes --use-tiers / HOMEMAKER_USE_TIERS.

experiments/tier_ab_2g7_3.py runs the acceptance A/B (harbor+maple, 3
seeds, 20k evals) in the background; results pending.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-02 16:00:39 +01:00
b0bd1a896b homemaker-py-91f: residual diagnostic on the current full default stack
Re-ran the §13.1/§13.2-style per-leaf fail-breakdown diagnostic on real
driver.search_staged runs (budget 20000, seeds 0-2, harbor-house and
maple-court) under the current full default stack (leaf-sharing x3,
depth-balanced, interior-O, share-aware edge cap) -- never decomposed by
category since those defaults were flipped on.

Finding: crinkliness (48%) and size (20.6%) now dominate the residual on
both programmes (~69% combined); construction-completeness fails
(missing space, adjacency, level, connectivity) are down to a small
tail (<=6% each). This revises erc.1's old recommendation to deprioritise
compactness-cuts in favour of leaf-sharing -- leaf-sharing is now fully
deployed and crinkliness is proportionally more dominant than ever, so
DESIGN.md §13.11 recommends reopening a compactness/crinkliness-targeted
construction lever as the next concrete step.

Also files two bugs found while validating the methodology: dumping and
reloading a .dom under leaf_sharing+collapse_insearch does not reproduce
the search's own in-process fail count (homemaker-py-iio), and
run_staged_search.py's own sanity rescore omits the collapse_insearch
override (homemaker-py-7ua). experiments/run_and_capture_91f.py sidesteps
this by capturing the true in-process fails list instead of rescoring
from disk; experiments/diag_residual_91f.py tallies fail categories from
those sidecars.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY
2026-08-01 19:38:46 +01:00
0ec404c73d Spike: torch-autodiff inner-loop ratio optimisation (homemaker-py-2ax) — negative
Build a torch-differentiable local proxy for the ratio-to-fitness path (exact
port of geometry.py's coordinate recursion + the 5 continuous per-leaf quality
factors, with discrete/structural facts frozen from a real fitness.py
snapshot and the 0.5^n cliff relaxed to a sigmoid) and compare Adam ascent
against nm_search on frozen topologies from programme-house and harbor-house.

Result: ~30-35x slower per unit of search progress than nm_search at both
6 DOF and 36 DOF (per-op torch tensor dispatch overhead with no batching
opportunity, plus snapshot/resnapshot cost on par with a full oracle eval),
and no better quality at matched budget. A step-size sensitivity check
confirmed the flagged 0.5^n cliff risk is real, but autodiff doesn't make the
gradient direction any cheaper to obtain here. Not recommended; kept as
reference only, not wired into innerloop.py. Full writeup in DESIGN.md §34.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY
2026-08-01 10:26:43 +01:00
3d141a619d homemaker-py-1s3: multi-use leaves N=15 confirmation -- does not replicate
The N=3 A/B (previous commits) found the precision-weighted shape
combination improved both example programmes (harbor-house -1.4%,
health-centre -13.9%), but N=3 is a thin sample by this project's own
standard (xyu/9yx use N=15). Two confirmations:

- N=15, plain search, budget=3000 (mirrors xyu/9yx's own protocol exactly):
  both programmes trend NEGATIVE (harbor +6.1%, health-centre +6.6%,
  Wilcoxon p=0.044)
- N=15, staged search, budget=20000 (true same-conditions replication --
  identical to the original A/B except seed count): both programmes AGAIN
  trend negative (harbor +6.6% p=0.15, health-centre +4.7% p=0.48)

The same-conditions replication disagrees with the original result's
direction on both programmes. Conclusion: the N=3 positive signal was
sampling noise, not a real effect -- health-centre's -13.9% was driven
substantially by one seed (71->43 fails) that didn't hold up.

multi_use stays default OFF and is not recommended even as a promising
lever -- this is a clean NULL, closing out both halves of §26's original
multi-use-leaves question (path a was NULL/NEGATIVE, path b is NULL after
replication). Mechanism itself is unchanged, complete, and fully tested.
DESIGN.md §33 rewritten with all three measurements and the honest verdict.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY
2026-08-01 08:52:24 +01:00
27719975b0 homemaker-py-1s3: multi-use leaves as permanent design goal (§26 path b)
Builds path (b) from §26 -- a leaf permanently serving two DIFFERENT
compatible programme codes at once, extending leaf-sharing's same-code
multiplicity mechanism to different-but-compatible codes. Architect-declared
`co_locate` pairs (validated against interchangeable()'s S1-S4 bounds, no
transitive closure so the b3v chain problem can't recur), threaded through
graph.py's checks via a new leaf_codes() resolver and fitness.py's quality
terms (additive size, stricter-of-both width/proportion). Construction-time
only, gated behind `multi_use` (default OFF, bit-identical when off).

End-to-end A/B (20k evals x 3 seeds x 2 programmes) came back net negative:
harbor-house -4.0% but health-centre +24.5% worse (3/3 seeds), because
fusing different codes' shape targets via stricter-of-both can impose a
tighter joint constraint than either code needed alone, which the tightly-
packed health-centre programme can't absorb. Written up as DESIGN.md §33;
multi_use stays default OFF, no default-flip recommended.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY
2026-07-31 00:16:12 +01:00
ce3fed166f homemaker-py-9yx: health-centre — non-synthetic third example programme
A small primary-care health centre: 19 distinct, individually-sized room
codes at n=20 room instances (only one deliberate duplication: two public
WCs), filling the gap between programme-house's duplicated-count sweep
sizes and harbor-house's out-of-range 37 real-diversity instances. Widths
are deliberately tiered (>1.3x gaps at three boundaries) so the auto-derived
interchange relation resolves to three bounded utility/office/clinical
classes instead of one whole-building chain, which a first pass produced.

experiments/run_9yx_sweep.sh repeats xyu's ruin_recreate ON/OFF protocol
(budget 3000, 4 workers, N=15 seeds) against this programme.
2026-07-30 00:09:37 +01:00
e5b7bc810e xyu: n=18 ruin_recreate N=15 confirmation — weakened, not resolved
Extends y51's n=18 synthetic sweep (strongest of four sizes at N=10) to
N=15 seeds, matching f1d's own confirmation sample size. Effect shrank
(9.3%->6.4%, two-sided Wilcoxon p 0.098->0.059) but didn't evaporate or
reverse — an ambiguous middle case, not a clean confirm or null. Refiled
option (b) (non-synthetic third example programme) as homemaker-py-9yx
since extending N alone doesn't address the interchangeable-room-code
confound §24 already flagged. enable_ruin_recreate stays default OFF.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY
2026-07-29 10:10:23 +01:00
15abba9679 e01: harbor-house beam-width larger-N confirmation (null)
N=15 driver.search sweep of construction_beam_width 1 vs 4 (protocol
identical to c94's original 5-seed run, DESIGN.md §29): 6W/4L/5T, mean
fails 57.0->56.6, Wilcoxon p=0.84. Excluding seed 2's outlier the mean
flips slightly negative (56.3->56.9), confirming the §29 5-seed
"improvement" was that one outlier. construction_beam_width stays
default 1 on confirmed rather than precautionary grounds. DESIGN.md §30.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY
2026-07-29 00:22:55 +01:00
bc59849666 y51: synthetic room-count sweep finds no clean ruin_recreate size threshold
Locates the threshold f1d's programme-house/harbor-house split implied,
using four synthetic sizes (10/14/18/22 rooms) derived from programme-house
by scaling its bedroom+ensuite module count, since no natural third example
programme sits between the two. Results are noisy and non-monotonic (n=10
mild win, n=14 clean null, n=18 strongest trend at p=0.098, n=22 near-null)
rather than a clean decay with room count -- documented in DESIGN.md #24.
enable_ruin_recreate stays default OFF; filed homemaker-py-xyu as a
low-priority follow-up (larger-N at n=18, or a non-synthetic third example).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 15:35:03 +01:00
0d94e58119 f1d: ruin-and-recreate LNS operator, validated positive on programme-house
Adds operators.mutate_ruin_recreate: un-divides one wing of a storey and
rebuilds it with the same adjacency-aware constructor the seeders use
(_assign_adjacency_aware, generalised with a new `scope` param), seeded
from the surviving circulation bordering the wing. Gated behind
enable_ruin_recreate (default off) / --ruin-recreate, same pattern as
reassociate/bridge_circulation.

A/B (qpk protocol, DESIGN.md §23): initial uniform-weight run was
underpowered (fired ~1/32 children), null. A weight=3.0 follow-up
(_MUTATION_WEIGHTS["ruin_recreate"]) showed a statistically significant
win on programme-house across 15 seeds (8W/1L/6T, mean fails 7.07->6.00,
Wilcoxon p=0.041) but no consistent effect on harbor-house across 8 seeds
(3W/2L/3T). Kept default off pending a size-threshold follow-up.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 09:31:42 +01:00
1e2fa2adfc lj3/qjg: larger-N + weight sweep finds bridge_circulation effect is noise
Combined follow-up to 8sh (DESIGN.md §22): raised bridge_circulation's
_MUTATION_WEIGHTS entry to 2.0 (lj3) and re-ran the qi6/qpk-protocol A/B
at 4x the sample size (qjg) in one sweep, since the two variables were
confounded if tested separately. Result is null in the opposite direction
from 8sh's small-N signal -- no total-fail benefit (p=0.71 programme-house
N=20, p=0.69 harbor-house N=12) and a higher rate of trajectory-divergence
-induced new not-connected fails than at the original uniform weight.
Reverted the weight bump; enable_bridge_circulation stays default off.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-25 12:20:22 +01:00
a103e0114a 8sh: insert/relocate-circulation repair operator (qi6 mechanism (a))
Adds operators.mutate_bridge_circulation: retypes the cheapest path between
two disconnected circulation components to circulation, directly clearing a
'level N not connected' fail instead of relying on the qi6 graded comparator
key (measured negative, DESIGN.md §18). Gated off by default via
driver.search's enable_bridge_circulation flag and evolve.py
--bridge-circulation, mirroring enable_reassociate's clean-toggle pattern.

qi6/qpk-protocol A/B (DESIGN.md §21) is directionally positive but mixed at
N=3/N=5 (never worse on total fails; clears 2/5 baseline not-connected fails
vs qi6's 0/4; one seed's RNG-trajectory divergence adds 2 new not-connected
fails) — kept default off pending a larger-N confirmation sweep
(homemaker-py-qjg) and a mutation-weight bump experiment (homemaker-py-lj3).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-24 19:48:13 +01:00
8328ac1b69 qi6: full-budget A/B for graded circulation-connectivity signal, negative result
conn_grade ON vs OFF (qpk protocol, experiments/run_qi6_ab.sh): harbor-house
(budget 2500, seeds 1-3) byte-identical output in every seed — the secondary
comparator key never fired. programme-house (budget 3000, seeds 1-5) 3/5 seeds
tie exactly; seeds 1/2 diverge to a different topology but the fail delta is
adjacency/crinkliness/width/access/size, never connectivity. Zero of 4 cases
where a not-connected fail was present got cleared by the grade.

Mechanism (b)/(c) (graded proximity as tertiary comparator key) is falsified,
not just unconfirmed. Kept default OFF (already was). Closed qi6; filed
homemaker-py-8sh for the remaining candidate (mechanism (a): an explicit
insert/relocate-circulation operator that doesn't depend on the search
stumbling onto a fail-count tie).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-23 18:29:45 +01:00
e09221051c 6zy/§11.8: co-tune topology diversity × tournament pressure — null robust
Expose tournament_k (default 2) on search()/search_staged(), threaded into
both _tournament call sites and the staged path's internal search() calls;
HOMEMAKER_TOURNAMENT_K env knob in the scaled/staged harnesses; run_6zy_ab.sh
joint niche×k grid (RESUME-able).

Result (negative, acceptable): no (niche,k) cell beats the legacy (off,k=2)
baseline. Blank-slate programme-house (5 seeds) baseline mean 4.80 fails is the
best of the 6-cell grid; every k>2 and every niche=on cell is 6.0-7.0. Niching
bites (pop_distinct 16/16 vs 4-11) but sharper pressure does not convert it to
lower fails — §11.5 'diffuses effort' null is robust to selection pressure;
plateau stays reachability-bound (confirms §11.4/§11.5).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 22:56:23 +01:00
d627ee5fb2 psk/§14: island model — null (best-of-N at equal budget wins)
Prime a population from N independent converged elites + crossover-heavy
migration phase, vs best-of-N at equal total budget. Island does NOT win:
harbor 68 vs control 67 (within parallel noise), maple 124 vs control 116
(decisive). Default-off child_probe hook on driver.search instruments the
deciding mechanism: area-matched crossover across independently-converged
elites rarely synthesizes (1/65 harbor, 3/63 maple beat the better parent,
max fail-drop 2-5), confirming the alignment hypothesis (non-canonical 9gp
encoding -> disruptive splice). Search-machinery null #3; residual stays
geometry/shape-bound. 233 tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 06:20:29 +01:00
bb9b355f14 x3b/§13.10: productionise leaf-sharing — per-code share grain + CLI wiring
Make the §13.3 lever a first-class feature, not experiment-only.

- programme.py: SpaceReq.share (default 1) + has_share, parsed from
  patterns.config 'share: N'.
- operators._share_grain: resolve per-code grain from leaf_share_factor
  selector — 0 = per-code opt-in (share iff share:N>=2), >=2 = global with
  per-code override (share:1 opts OUT, share:N sets grain). _share_rooms
  groups per resolved grain.
- End-to-end conf injection without monkeypatch: load_config(overrides=)
  merges run-level keys last; driver.search / innerloop.optimise /
  NativeEvaluator / _fitness_for thread conf_overrides={leaf_sharing:True}
  through both inner-loop and off-tree scorers when sharing is on.
- homemaker-evolve: --leaf-sharing/--no-leaf-sharing + --leaf-share-factor
  (env HOMEMAKER_LEAF_SHARING / HOMEMAKER_LEAF_SHARE_FACTOR).
- Example programmes untouched (§13.3/§13.9 stay reproducible). Experiment
  load_config monkeypatches updated to accept overrides=.

Tests: grain modes, opt-out, default-OFF parity, load_config overrides,
programme parse, CLI parse. 233 pass. Smoke: harbor 37 vs 95 fails on/off.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 22:04:35 +01:00
f43de001fb rq2/§13.9: flip share_edge_cap default-ON for leaf-sharing runs
§13.8 verdict was positive and monotone-harmless, so default the share-aware
edge-too-long cap to leaf_sharing when share_edge_cap is unset — mirrors the
pll bal+share and §13.6 interior_outside default flips. Explicit
share_edge_cap=False still reproduces the pre-flip control arm.

- fitness.Fitness.__init__: cap defaults to self._leaf_sharing when the conf
  key is unset (None); explicit True/False honoured.
- run_staged_search.py: pin conf["share_edge_cap"] = share_edge in both A/B
  arms so SHAREEDGE=0 stays a clean control post-flip.
- tests: control arm now pins share_edge_cap=False; new
  test_edge_cap_defaults_on_under_leaf_sharing guards the flip.
- DESIGN.md §13.9: rebaseline §13.x floor (maple 80.3→74.0, harbor 34.7→31.0).

Non-sharing runs untouched: programme-house control re-score reproduces
bit-for-bit. 222 tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 21:38:53 +01:00
393183b356 hph/§13.8: share-aware edge-too-long cap — shared leaves no longer penalised for aggregate wall length
§13.7 flagged edge-too-long as harbor's top fail class. Dissection showed the
bulk are a leaf-sharing REPRESENTATION ARTIFACT: a share=k leaf aggregates k
same-code rooms, so its walls run ~k× the flat 8 m cap purely for being big —
the same §13.3 leak (size/missing relaxed for shared leaves) on the wall measure,
since edge_cost/outside_edge_cost ignored leaf.share.

Fix: Fitness._edge_cap(*leaves) scales the 8 m cap by the largest type-guarded
leaf_share among adjoining leaves, mirroring quality_size's k×target; non-shared
leaves keep the flat cap so genuine narrow/oversize pathologies stay flagged.
Gated behind a share_edge_cap config knob (SHAREEDGE env), default OFF so the
§13.x controls reproduce.

A/B (full Phase-8 stack, staged, 20k evals, seeds 0/1/2): control reproduces
§13.7 (maple 80.3 exact, harbor 34.7≈34.0); share-aware arm maple 80.3→74.0
(−7.9%), harbor 34.7→31.0 (−10.6%), zero regressions across 6 seeds. Positive
and monotone-harmless (only ever removes a false-positive fail). Verdict:
recommend default-ON; follow-up issue flips the default + rebaselines the floor.

Tests: 6 new unit tests for _edge_cap (221 pass).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JygRv4n2dcyDQqMiDRe7TN
2026-06-28 21:24:51 +01:00
dba85094ea hph: dissect §13.7 edge-too-long — leaf-sharing artifact + narrow sliver
experiments/diag_edge_too_long.py: the 6 harbor edge-too-long fails are 2
locations — a share=3 combined leaf (247 m², aspect 1.2; flat 8 m cap not
share-aware, unlike quality_size's k×target) accounting for ~4, and one
1.2×16.7 m narrow sliver (~2, also caught by width/proportion). No corridors.
Files homemaker-py-hph (share-aware edge-too-long fix).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JygRv4n2dcyDQqMiDRe7TN
2026-06-28 14:54:55 +01:00
27d7a0f771 erc.7d/§13.7: high-budget harbor floor probe — close 71d NO-GO, wrap Phase 8
500k serial full-stack harbor probe (probe_harbor_floor.py): 20 fails,
crinkliness 13→4, landlocked crinkliness ~13→2 of 20. Interior-O (default-ON,
erc.8) is 71d's named fix and dissolved its landlocked-crinkliness target;
residual now diffuse (top class edge-too-long). NO-GO on 71d.

Cumulative Phase-8 floor vs §12.2 baseline (leaf-share-relaxed): maple
136.0→80.3 (−41%), harbor 74.0→34.0 (−54%) — all from construction levers,
none from search machinery, per the epic thesis.

Closes erc epic: 71d/7u5/jrb/u8x superseded-by-construction; erc.5/erc.6
wont-fix (Diag A/B revisit conditions unmet). DESIGN §13.7.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JygRv4n2dcyDQqMiDRe7TN
2026-06-28 14:23:34 +01:00
2491a9be12 erc/ld2: interior-O light-well seeding — §13.6 positive on dense floors
Seed O as interior light wells (most-landlocked leaves first, count scaled
by room count via outside_divisor) instead of one peripheral O, attacking the
erc crinkliness residual: seed diagnostic confirms every crinkliness fail is
under-exposed (landlocked), none over-exposed.

A/B (20k evals, seeds 0/1/2, bal+share stack, §13.6): control reproduces §13.5;
interior odiv=3 gives harbor -16.4% (all seeds improve) and maple -2.8%
(net-neutral). Default-optimal divisor 3 found by seed sweep (6 was null).

Lever default OFF; default-ON flip tracked as erc.8.

- operators: interior_outside + outside_divisor through constructive_topology,
  lift_base_to_storeys, _assign_adjacency_aware (fix n_circ budget for >1 O)
- driver.search/search_staged threading; run_staged_search.py INTERIORO/ODIV env
- test_interior_outside_seeds_landlocked_wells_and_scales_count
- experiments/run_interioro_ab.sh; DESIGN.md §13.6

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 07:20:20 +01:00
34a7b2ecf9 erc.7: leaf-sharing × depth-balancing synergy CONFIRMED end-to-end (§13.5)
Synergy A/B: bal+share vs share-alone, factor 3, seeds 0/1/2, staged 20k.
maple 86.3->82.3 (-4.6%), harbor 50.7->40.0 (-21.1%, non-overlapping arms).
Control reproduces §13.3. Adds run_synergy_ab.sh + run_sharefactor_sweep.sh
(factor 2/4 sweep under bal+share, running).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 16:47:14 +01:00
4aaa295dc1 erc.4: depth-balanced construction mechanism + floor probe (§13.4)
_grow_leaves grew a random caterpillar, so equal-target rooms landed at
wildly different binary-tree depths — the depth-driven size maldistribution
Diagnostic B (§13.2) localized (same code at 0.05x and 14.7x target). The
depth_balanced flag always splits a shallowest leaf instead, growing a
near-complete tree so the proportion-aware sizing pass hits each target with
cut fractions near their proportional value.

Floor probe (diag_depth_balance.py): depth spread collapses 7->1, the giant
ratio falls (maxR 12->8 harbor / 16->6 maple), %undersize 54->25 / 42->22,
and the achievable floor drops -12% harbor / -11% maple at EQUAL leaf count.
Additive with leaf-sharing (bal+sh3 beats §13.3 share3-alone). Default OFF,
214 tests pass; threaded through driver.search/search_staged and exposed via
DEPTHBAL in run_staged_search.py. End-to-end 20k A/B running.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 22:36:24 +01:00
e983229857 erc.3: explicit per-leaf multiplicity closes the leak; driver A/B wired (§13.3)
Replace the area-derived share recovery with explicit, type-guarded per-leaf
multiplicity: construction stamps leaf.share=k and leaf.share_type=code; the
fitness (graph.leaf_share) honours k only while leaf.type==share_type, so any
retype/undivide auto-invalidates a stale share — no operator resets, and a
small leaf cannot retype its way into covering rooms it does not provide. Two
Node fields survive the whole search via deepcopy (genome.decode is unused in
the hot path); .dom emits `share` only on a live shared leaf.

This closes the §13.3 missing-fail leak: floor probe missing 17–44 → 0, and the
achievable floor drops −39% harbor (120.3→73.3) / −32% maple (194.7→133.0) with
no re-emergence as size fails.

Flag threaded through driver.search/search_staged → constructive_topology /
lift_base_to_storeys, exposed via LEAFSHARE/LEAFSHAREFAC in run_staged_search.py
(injects the objective into inner-loop + final-score fitness so both A/B arms
share one programme dir). run_leafshare_ab.sh runs the staged 20k A/B.
Smoke-tested end-to-end (harbor, factor 3, re-score OK). 214 tests pass;
default-OFF reproduces baseline.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 18:16:17 +01:00
bf3ff43837 erc.3: leaf-sharing mechanism + floor probe (§13.3)
Same-code rooms collapse into fewer, larger SHARED leaves so the ~1.8/leaf
shape tax (§13.1) is paid once per group. Multiplicity k is recovered from
area (k=clamp(round(area/target),1,max_share)) — no genome change — and used
in two default-OFF sites: graph.check_space_counts counts coverage (Σk vs
req.count) so one leaf covers several rooms without a missing fail, and
fitness.quality_size centres on k×target (σ scaled by k). Construction:
operators._share_rooms groups instances; _size_divisions_from_targets sizes
shared leaves to k×target via leaf_mult.

Floor probe (experiments/diag_leaf_sharing.py, harbor+maple, seeds 0/1/2,
+innerloop): total fails −27% harbor / −16% maple at share3, shape factors
fall ~linearly with leaf count (confirms §13.1). Cap: 17–44 missing fails
leak because depth maldistribution (§13.2) keeps shared leaves below k×target
so round() undercounts; inner loop can't close it. Net still positive.

Default-OFF reproduces baseline exactly (214 tests pass). Driver plumbing +
staged 20k A/B remain; §13.3 records the next design fork.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 08:30:26 +01:00
be6857414d erc.2: Diagnostic B — undersize-despite-slack localization (§13.2)
The "56% empty plot" is a misreading: sized rooms already hold 1.4-1.5x
their aggregate target area; ~46% of plot is circulation, not claimable
void. Size fails are depth-driven MALDISTRIBUTION — the same type/target
leaf lands 0.05x..14.7x by binary-tree position. The inner loop cannot
repair it (frozen topology, budget-80 size fails move only -1.6/-3.7).

=> Falsifies plot-fill-as-claim-void: re-scope erc.4 to depth-balanced /
giant-splitting construction; deprioritise erc.6 (inner-loop term, wrong
DOF). Reinforces erc.3 leaf-sharing for the starved tail.

Script: experiments/diag_slack_localization.py (self-contained evidence).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 22:47:34 +01:00
7bd4adf32a erc.1: Diagnostic A — per-leaf shape-fail vs density (§13.1)
Controlled synthetic sweep (maple-court, room set fixed, circ_divisor 2->9)
shows per-leaf shape-fail is FLAT vs slicing density (1.72-1.94, no trend)
while TOTAL shape fails track leaf count linearly (139->116). Crinkliness
dominates (~0.8/leaf) and is flat; cuts are already squarest yet still pay
~1.8 fails/leaf. Floor is INTRINSIC to per-leaf slicing, not cut quality.

Verdict: prioritise leaf-sharing (erc.3); deprioritise compactness-cuts
(erc.5 -> P4). Adds experiments/diag_leaf_shapefail.py.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 22:06:04 +01:00
e9684ea7ef c3g: circ-per-room granularity knob (circ_divisor) + A/B harness
Threads circ_divisor (default 3 = unchanged) through
operators.constructive_topology/lift_base_to_storeys and
driver.search/search_staged; env CIRCDIV in run_staged_search.py. Adds
experiments/run_c3g_ab.sh.

Motivation (DESIGN.md §12.3 diagnostic): the maple shape residual is
over-granular construction (73 small leaves -> crinkliness+size). Cheap raw-seed
probe: a coarser spine lowers the SHAPE floor (maple 135->110, harbor 83->66)
but raises access/adjacency, leaving the raw TOTAL floor flat-to-worse. Because
§12.3 showed shape is the HARD residual and access/adjacency are cheap to
repair, only an end-to-end A/B settles whether trading them pays — this is the
plumbing for that run. Tests green (default path byte-identical).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-21 21:10:18 +01:00
7e39bf5870 Phase 7 §12.3: 9gp A/B measured — NEGATIVE; close 9gp + epic leu
24-run sweep (maple-court + harbor, seeds 0/1/2, 20000 evals): M3 reassociate
and the shape-feasibility filter are both neutral-to-slightly-worse vs the
§12.2 baseline (maple 136.0 -> 139-140, harbor 74.0 -> 77-78). Baseline controls
reproduce §12.2 exactly, so the negative is real.

Verdict: the Phase-7 residual is the geometry/shape floor of the constructed
slicing layouts, not reachability/feasibility-bound — third independent negative
on search machinery (§11.4/§11.5/§12.3) vs four construction/seed wins
(§11.2/§11.6/§11.7/§12.2). A full canonical Polish rewrite is not justified: its
one testable promise (associativity reachability) was tested and did not pay.
Both operators kept default-OFF.

Closes 9gp.1, 9gp.2, 9gp; epic leu (Phase 7) auto-closed (3/3). Adds the
reproducible sweep harness experiments/run_9gp_ab.sh.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-21 07:21:51 +01:00
6ee5d4b4ae Phase 7 §12.3: re-scoped 9gp — shape-feasibility filter + M3 reassociate (9gp.1, 9gp.2)
Land the two evidence-supported parts of the re-scoped 9gp capstone as
operators on the existing decoded Node tree (no Polish-expression rewrite),
each default-OFF and measured against the §12.2 leu.2 baseline.

9gp.1 shape-feasibility pre-filter: operators.predicted_shape_fails lays a
topology out at its proportion-aware target geometry and counts shape fails
(size/width/proportion/crinkliness); driver._evaluate prunes clearly-infeasible
topologies before the inner loop (1 eval vs ~80), guarded so nothing that could
beat the incumbent is discarded. search/search_staged feasibility_filter,
feasibility_max_shape_fails (env FEAS/MAXSHAPE), default OFF.

9gp.2 M3 Wong-Liu reassociate: operators.mutate_reassociate adds associativity
(a|b)|c <-> a|(b|c) on same-orientation live cuts — the canonical-slicing move
missing from swap(M1)/rotate(M2), attacking the §11.4/§11.5 reachability
bottleneck. enable_reassociate (env REASSOC), default OFF (weight 0 -> baseline
byte-identical).

Unit tests (operators + driver) green, full suite 211 passed; maple-court smoke
run clean under native fitness. A/B sweep handed off per the plan; DESIGN.md
§12.3 documents the design and the pending measurement.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 18:54:48 +01:00
995342d0a4 Phase 7 §12.2: proportion-aware constructive seeding + storey_minimum fix (leu.2, cq1)
Size each constructive-seed cut from leaf TARGET areas (division=[f,f] gives
left area-fraction f) and pick each cut's rotation for child squareness — both
derived from target dims, topology/type assignment untouched. Area-only
regressed (slivers); rotation choice is what makes it pay.

End-to-end (20000 evals, 3 seeds, staged): harbor 85.3->74.0 (-13%, best 69),
maple-court 151.7->136.0 (-10%, best 126). PROP=0 reproduces the §11.7/§12.1
baselines exactly. programme-house regresses at fixed budget (deeper local
optimum walls off the undivide restructuring path) but a budget sweep shows
it's convergence speed, not a worse asymptote (PROP=1 reaches 1 fail at 150k).
Default-on (seed_proportion_aware=True, env PROP=1).

cq1: n_storeys now honours storey_minimum, not just level: keys — programme-house
(storey_minimum:2, all rooms level:0) was seeded one storey short and fell
through to plain search. New programme.storey_minimum()/n_storeys_for();
driver.search passes min_storeys to the seeder; search_staged routes on the max.
No-op for harbor/maple; programme-house single-stage 8.0->5.0.

New maple-court best (126) saved as generated.dom. 204 tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 14:04:42 +01:00
d004e4c937 Phase 6 §11.7: adjacency-aware lift + secondary adjacencies (ld5)
_assign_adjacency_aware gains fixed_circ (seed the connected-dominating-set from
given circulation leaves) and secondary-adjacency-aware room placement: codes
with the most non-c adjacency requirements are placed first, each onto the open
slot satisfying the most of its requirements against already-typed neighbours
(clustering k1<->da1, da1<->o). lift_base_to_storeys(reqs, adjacency_aware=True)
grows the upper-floor circulation spine off the inherited vertical core and
assigns rooms around it; threaded through driver.search_staged
(seed_adjacency_aware) and run_staged_search.py (ADJ env).

End-to-end staged harbor, 20000 evals, mean total fails over 3 seeds:
ADJ=0 99.0 (reproduces the §11.4 staged lex baseline exactly), ADJ=1 85.3
(-13.7, -14%; best 78). New best harbor configuration overall: staged baseline
99.0 -> single-stage adjacency-aware (§11.6) 90.7 -> staged + adjacency-aware
lift 85.3. Staging and adjacency-aware seeding compose.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 11:47:40 +01:00
c1586237ca Phase 6 §11.6: adjacency-aware constructive seeding (s44)
operators._assign_adjacency_aware spends ~one extra leaf per three rooms on a
greedy connected-dominating-set of circulation leaves (read from the geometric
leaf_graph, type-independent), so every room borders a connected circulation
spine and adjacency-to-c + access are satisfied by construction. Default-on via
constructive_topology(adjacency_aware=True), threaded through
driver.search(seed_adjacency_aware) and run_search_scaled.py (ADJ env).

End-to-end single-stage, 20000 evals, mean total fails over 3 seeds:
harbor 110.0 -> 90.7 (-17.5%; ADJ=0 reproduces the §11.2 105 baseline exactly),
programme-house 12.3 -> 9.3 (-24%). Adjacency-aware single-stage harbor (mean
90.7, best 85) beats the §11.3 staged best of 95 — the first Phase-6 fail-count
reduction from seeding. Follow-ups (lift_base_to_storeys, secondary adjacencies)
filed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 09:23:12 +01:00
059964ee05 Phase 6 §11.5: structural niching + restarts — negative result (c4c.5)
genome.signature: ratio-invariant structural topology hash (per-storey tree
shape + cut orientation + leaf types), the cheap stand-in for the 9gp canonical
encoding. driver gains niche_by_signature (one individual per topology, replaces
the fitness-scalar dedup) and restart_patience (soft restart: keep elites,
refill with fresh seeds); SearchResult gains n_distinct_signatures /
diversity_history / n_restarts.

Diversity criterion MET (final-pop distinct ~5/16 -> 16/16). Gate NOT met:
blank-slate programme-house mean fails 12.3(legacy)/12.7(niche)/13.0(restart)
over 3 seeds at 20000 evals; harbor staged 95/94/108. Niching is a tie within
seed noise, restarts strictly worse — falsifies the premise that the
fitness-scalar dedup causes premature convergence. Both flags default-off,
kept for reuse. Epic c4c complete.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 23:42:39 +01:00
ed2869074b Phase 6 §11.4: graded high-fail objective — negative result (c4c.4)
Implement a graded proximity comparator key (-n_fails, grade, fitness) behind
a default-off use_grade flag: fitness._leaf_grade / score_with_grade sum
f/FAIL_THRESHOLD over failing per-leaf quality factors; scalar fitness and fail
count stay untouched so the inner-loop 0.5^n cliff (§5.4) is unaffected (0/9
regression check: PASS). Read once per child in driver._evaluate off the
already-optimised tree; threaded through search_staged (Stage 2 only).

Harbor staged A/B (20000 evals, seeds 0/1/2): lex 95/96/106 (mean 99.0) vs
lex+grade 99/98/102 (mean 99.7) — grade wins 1/3, no plateau escape. Premise
falsified: within a fixed fail-tier 0.5^n is constant so fitness still spans
~6 orders of magnitude; grade above fitness displaces that working signal.
Verdict: reject; lexicographic (-n_fails, fitness) stands. Flag kept default-off
for reproducibility / possible reuse as a §11.5 diversity signal.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 22:33:29 +01:00
6ed9e0b4b1 Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).

New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.

Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
3c8f7aba07 Lexicographic outer-search comparison, preserve inner-loop cliff (homemaker-py-yg5)
Outer search now ranks individuals by (-n_fails, fitness) instead of raw
fitness scalar.  This prevents high-score 3-fail designs from displacing
2-fail designs in tournament selection and population replacement — the
root cause of the §4.8 pathology where flag count dominates geometry.

Inner loop is unchanged: it still optimises against the raw 0.5^n fitness
scalar, so the cliff that prevents trading into new failures remains intact
(0/9 regressions in experiments/penalty_reshape.py).

Also removes stale _CHILD_INNER_KW = {"sigmas": (0.05,)}: this was left
over from the CMA-ES era; the NM inner loop default (homemaker-py-d6d)
does not accept a sigmas parameter.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-14 09:20:03 +01:00
0e5e607c4f Swap inner loop default from CMA-ES to Nelder-Mead (homemaker-py-d6d)
Bakeoff with native fitness shows NM wins at all DOF sizes: +9% at
child_budget=80 for programme-house (6-7 DOF), and decisively at
harbor-house scale (35-40 DOF) where CMA-ES exhausts its convergence
detector after ~3 generations (46 evals) and adds failures on 12/15
runs.  NM uses the full budget, is parameter-free, and has zero new
failures across all test cases.

- Add nm_search() to innerloop.py; change optimise() default to "nm"
- Add nm_search to parametrised test cases
- Add bakeoff_native.py and bakeoff_harbor.py experiments with results

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-14 08:51:22 +01:00
646ee30ab6 Rename package: homemaker → homemaker-layout
- src/homemaker/ → src/homemaker_layout/; all imports updated
- pyproject.toml: name = homemaker-layout, entry point updated
- .beads/config.yaml: dolt sync.remote updated to homemaker-layout.git
- Delete temporary debug/perl scripts from project root
- README.md, DESIGN.md: package path references updated
- GitHub repo renamed; git remote updated

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-14 08:18:06 +01:00
7796e795a5 Phase 3 gate (homemaker-py-ccw): scaled search on native fitness
programme-house budget=20000: 1.04e-02 (2 fails), 1.36× over Phase-2
oracle run and 2.60× over urb-evolve p128. Winning topology found via
rotate at eval 10357, unreachable within Phase-2 budget. 71.8 evals/s
(~140× faster than batched oracle).

harbor-house (16 rooms): 3.73e-18 (49 fails) at budget 10000 in 633s.
This programme is beyond the oracle's capability; native fitness makes
it feasible. 638 topologies explored.

Adds experiments/run_search_scaled.py (native-only search runner, no
oracle dependency). DESIGN.md records Phase 3 gate result.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 22:10:38 +01:00
8e762b80d8 Phase-2 gate results: 2/3 seeds → REVIEW; fix patterns.config re-score bug
benchmark_vs_urbevolve.py results (2026-06-13, budget=2000, URB_NO_OCCLUSION=1):
- Seeded designs: memetic beats urb-evolve 1.91× (c964435) and 1.63× (2f45907)
- Blank slate init.dom: memetic at 18 fails vs urb-evolve at 6 fails (topology
  diversity gap from single-seed mutation chain vs random-population init)

Bug fixed: run_search.py was calling oracle.score on out.parent without
patterns.config present — causing the re-score to return near-zero instead of
the correct tracked fitness. Added shutil.copy to propagate patterns.config
alongside the output .dom before the standalone re-score.

Gate recorded in DESIGN.md §7. Closes homemaker-py-way.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 09:56:01 +01:00
bc61f8cb73 Bake-off: CMA-ES confirmed as inner-loop optimiser (homemaker-py-d0s)
4-way comparison (NM / CMA-ES / compass / compass-ms) over 3 corpus files ×
3 seeds at budget 200, cold-start, URB_NO_OCCLUSION=1. CMA-ES wins on
batch-efficiency (18 oracle calls vs 200 for NM, 12x speedup on Perl startup
amortisation per §4.6) with acceptable quality (x1.41 @200 vs NM's x1.56).
Compass stalls on narrow-valley landscapes and introduces fail regressions.
NM flagged as Phase 3+ candidate once native fitness removes oracle call overhead.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 09:47:15 +01:00
c01a8a0887 Native fitness: leaf quality terms + cost model (homemaker-py-gnw)
Port Urb's programme-driven fitness leaf quality factors (perpendicular,
proportion, size, width, crinkliness, daylight, access), value rates,
and cost model (per-leaf area costs, interior/exterior wall edge costs,
boundary costs) to Python.  Passes 0-mismatch parity against the Urb
oracle across all 35 corpus files (407 leaves, 2849 factors), using
URB_NO_OCCLUSION=1 simple crinkliness (illumination factor pinned to 1).

Key fixes: _dist must use math.sqrt not math.hypot (1-ULP difference
flips boundary overlap predicates); leaf-scope fail regex requires ^\d+/
prefix to exclude building-level failure messages.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 07:59:21 +01:00
3bf507a483 Fix benchmark cell arg order and mutate_swap on undivided trees
run_urbevolve took (seed, budget, pop, cell) but cells call
fn(seed, budget, cell, **kw) — every urb-evolve cell died on TypeError,
deferred silently by pool.map. mutate_swap lacked the empty-candidates
noop guard the other operators have, crashing on init.dom-style bare
plots. Regression test: every mutation survives an undivided tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-12 23:26:22 +01:00
e2b3e20070 Phase-2 gate benchmark: memetic loop vs urb-evolve at equal eval budgets
experiments/benchmark_vs_urbevolve.py (homemaker-py-way): 3 seed designs
(init from scratch, c964435 weak, 2f45907 strong) x 2000-eval runs, the
1000-eval tier read from each run's best-so-far log; urb-evolve gets two
population sizes (default 128 = ~16 generations at this budget, and 16 =
~130) and credit for its better one. Counts via the MAX_EVALS counter
patch in urb-evolve.pl (committed in the urb repo); both systems under
URB_NO_OCCLUSION=1; all finals re-scored through urb-fitness.pl as the
common deterministic yardstick.

run_search.py generalised to (budget, rng-seed, seed.dom, out.dom);
innerloop.optimise now handles 0-DOF topologies (an undivided plot like
init.dom scores once instead of crashing CMA).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-12 22:22:16 +01:00
f160c6dc9e Use Urb's canonical UPPERCASE generic types (C/O); case-insensitive class checks
Bruno's correction: 'C' was never a 'covered' type — Is_Covered is a
geometric predicate. Urb generics are canonically uppercase (get_space_types
qw/C O S/; corpus 100% uppercase). The driver/operator type pool emitted
lowercase 'c'/'o', creating mixed-case designs that fragmented Dom->Ratios
class buckets and fired the latent ratio_type first-match nondeterminism
(which the search promptly reward-hacked). Operators now emit uppercase
generics only and class checks match case-insensitively (t[0].lower() in
'cos', cf. Is_Circulation/Is_Outside). The Urb-side class-sum patch remains
as defensive hardening, zero-impact on canonical designs (35/35 parity).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-12 19:01:53 +01:00
0beb005a23 Memetic search driver: steady-state GA over topology, warm-started inner loop
driver.py (homemaker-py-b39): tournament selection, operators.mutate (storey
ops down-weighted) + area-matched crossover, every child's geometry
delegated to the warm-started inner loop (Lamarckian write-back; children
use a single local CMA phase - the exploratory ladder phase exists for cold
projections children never face). Budget stated and accounted in oracle
evaluations; near-duplicate fitness guard against population collapse
(neutral mutations are common, per 8cs).

free_with_keys/ratio_map/warm_x0 promoted from the 8cs experiment into
innerloop.py as the Lamarckian inheritance API; alignment with
solver.free_branches asserted across the corpus.

tests/test_driver.py fakes the inner loop: budget accounting, monotone
improvement history, warm-start + sigma plumbing, valid .dom output.
31 tests pass. experiments/run_search.py is the end-to-end acceptance run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-12 14:22:26 +01:00
92cc63348e Topology operators: 7 mutations + area-matched subtree crossover
operators.py (homemaker-py-nyb): divide/undivide/retype/swap/rotate/
level_add/level_delete + Urb-style area-matched base-storey crossover.
Operators edit the decoded Node tree; genome.encode absorbs all repair
(dangling deltas, storey misalignment) so every child is a valid genome
by construction. Geometry moves deliberately absent — the inner loop owns
continuous DOF, and 8cs made Lamarckian re-optimisation mandatory.

Fixes dom._link to CLEAR stale below-links when a path vanishes from the
storey below (undividing a base branch left upper nodes pointing at
orphaned quads; oracle scoring unaffected but in-process geometry crashed).

Acceptance (experiments/operator_locality.py, flag-on): 115/115 children
scored without error; geometry perturbation small for core ops (retype
0.07, divide/undivide 0.14, swap/crossover 0.16-0.17), fitness
perturbation large for all (0.68-0.99 rel) — the 0.5^n cliff flags most
raw moves, confirming warm-started re-optimisation + penalty reshaping
as the load-bearing design choices. 27 tests pass.

Closes homemaker-py-nyb.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-12 14:07:35 +01:00
13f73be771 Topology genome: base tree + per-storey deltas + type assignment
genome.py (homemaker-py-k2g): Genome = base-floor GNode tree + per-storey
StoreyDelta (undivides, divide subtrees, leaf retypes, height) + base
metadata. encode/decode round-trips dom.py Node trees.

Key empirical finding baked into the design: upper-storey nodes carry
heavily drifted DEAD fields (97 inherited-cut divisions, 187 rotations
differ from the owning node below across the corpus) — dead because
geometry delegates to below before reading them. decode canonicalises
them; encode stores only owned state, so genomes from drifted sources
compare equal (fixed-point test).

Acceptance: 35/35 corpus files fitness-identical after round-trip through
the oracle (experiments/genome_parity.py, URB_NO_OCCLUSION=1); owned-cut
projection + genome fixed-point + storey counts in tests/test_genome.py
(16 tests pass).

Closes homemaker-py-k2g.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-12 13:52:32 +01:00