3l6: leaf-sharing runs auto-finish (unfold+polish) before write

Leaf-sharing evolve runs optimised a sharing-credited objective (a shared
leaf of code X with share=k counts as k programme rooms, size re-centred on
k*target) but wrote the un-materialised genome, which the canonical
homemaker-fitness rates catastrophically worse (harbor-house: internal
1.03e-05 vs canonical 6.73e-29, 15 critical missing-room fails).

Fix: before write, driver.polish_finish() unfolds every live shared leaf
into k distinct rooms (operators.unfold_shared_leaves — pays down the count
deficit) then warm-starts a leaf_sharing=False polish search from the
unfolded genome so the materialised rooms get proportion/width/size cleanup.
Returned best.fitness is the canonical score (leaf_sharing off => internal
== canonical). This is yaa's proven unfold-then-polish path, made automatic.

evolve.py: new --polish-budget (env HOMEMAKER_POLISH_BUDGET; -1=auto=
budget//2, 0=unfold+rescore only). Interrupt forces polish_budget=0 for a
fast honest output. Default stays --leaf-sharing on (its topology-search
speed retained; output made honest by the finish). Schedule B in-run
annealing remains homemaker-py-kpu.

Verified (harbor-house 3000+1500): reported polish fitness 4.79788e-27
matches canonical homemaker-fitness exactly, 0 critical fails. Tests: 254
pass (+3 polish_finish).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
This commit is contained in:
Bruno Postle 2026-07-15 10:21:58 +01:00
parent 053ace2f53
commit 479d4a57d5
4 changed files with 197 additions and 14 deletions

View file

@ -22,7 +22,7 @@
{"id":"homemaker-py-1p0","title":"Geometry inner loop: full-objective equal-offset ratio optimiser","description":"DESIGN.md §5.1, §7 Phase 1. Productionise experiments/optimize_fullfitness.py into homemaker: optimise(topology, x0=None) -\u003e (geometry, fitness). DOF = equal-offset division ratios of free branches (solver.free_branches, lowest-storey cut ownership), clipped to [eps, 1-eps]. Objective = full oracle fitness (never a proxy — §4.2 falsified). Must support warm-start x0 (§5.6) and a population/batch evaluation mode so each iteration scores via one batched oracle call (§4.6).","acceptance_criteria":"Reproduces or exceeds §4.5 gains (x1.24x1.67, no new failures) on 2f45907, candidate-002, c964435; works as a library call on any corpus .dom","status":"closed","priority":1,"issue_type":"feature","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-06-11T23:36:58Z","created_by":"Bruno Postle","updated_at":"2026-06-12T08:46:31Z","started_at":"2026-06-12T00:14:19Z","closed_at":"2026-06-12T08:46:31Z","close_reason":"innerloop.optimise() lands: batched CMA-ES sigma ladder (0.05/0.15, IPOP popsize doubling, deterministic seeding) over equal-offset free-branch ratios vs full oracle fitness; warm-start x0 supported. Acceptance vs unprojected originals: x1.65/x1.66/x1.58 against bars x1.24/x1.67/x1.59, no new failures, 46 oracle calls vs NM's 200. Two near-bar results accepted as reproduced-within-noise (1% tol) — draw spread brackets the single-NM-draw bars; approved by Bruno 2026-06-12. Gotchas: equal-offset projection of legacy unequal cuts loses fitness/adds failures (midpoint projection used); pycma seed=0 means clock-seeded.","dependencies":[{"issue_id":"homemaker-py-1p0","depends_on_id":"homemaker-py-av5","type":"blocks","created_at":"2026-06-12T00:39:33Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":3,"comment_count":0}
{"id":"homemaker-py-8cs","title":"Experiment: warm-vs-cold start of inner loop (Lamarckian inheritance)","description":"DESIGN.md §5.6, §4.6. Warm-starting a child topology's inner loop from the parent's optimised ratios is the main lever for cutting per-topology cost (~3 min/topology cold). Apply single topology mutations to optimised corpus designs, re-optimise warm (surviving cuts keep values, new cuts get heuristic defaults) vs cold, compare oracle-call counts to convergence at equal final fitness.","acceptance_criteria":"Speedup factor measured across \u003e=10 mutated topologies; decision recorded (expect order-of-magnitude; if \u003c2x, revisit §4.6 Phase-2 scoping)","notes":"Experiment script committed (experiments/warm_vs_cold.py, 1cc86c8) and machinery validated oracle-free; one mutated child scored through the oracle OK. Waiting on homemaker-py-gp2 reference run to finish, then execute under URB_NO_OCCLUSION=1 (3 parents x 400 evals + 12 children x 2 x 200 evals, ~1.5-2 h oracle time). Default budgets: parent 400, child 200; target = evals to 95% of best final.","status":"closed","priority":1,"issue_type":"task","owner":"bruno@postle.net","created_at":"2026-06-11T23:36:58Z","created_by":"Bruno Postle","updated_at":"2026-06-12T11:44:45Z","closed_at":"2026-06-12T11:44:45Z","close_reason":"Measured (URB_NO_OCCLUSION=1, parent budget 400, child 200, 12 single mutations across 3 designs): cold start reached 95% of warm final in 0/12 cases within budget — speedup unbounded at practical budgets; warm finals beat cold finals x1.2-x4 in 12/12; 6/12 warm starts were within 95% at 1 eval (near-neutral mutations). Decision: Lamarckian warm-starting is MANDATORY in the memetic driver (homemaker-py-b39), not an optimisation; cold starts produce strictly worse geometry at equal budget. Note: 2 undivides were exactly fitness-neutral (same-type merge == Merge_Divided equivalence) — locality datum for homemaker-py-nyb.","dependencies":[{"issue_id":"homemaker-py-8cs","depends_on_id":"homemaker-py-1p0","type":"blocks","created_at":"2026-06-12T00:39:34Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
{"id":"homemaker-py-av5","title":"Batched oracle: score many .dom files per invocation","description":"oracle.py currently scores one .dom per urb-fitness.pl call (~1.65 s/dom). DESIGN.md §4.6: batching amortises Perl startup to ~0.99 s/dom and is required so population/batch optimisers can score a whole generation in one oracle call. Extend oracle.py with a batch API: write N .dom files, one perl invocation, parse N .score/.fails pairs. Keep the single-file path for compatibility.","acceptance_criteria":"Batch of 35 corpus files scores in one perl invocation; per-file results identical to single-file calls; measured s/dom reported","status":"closed","priority":1,"issue_type":"task","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-06-11T23:36:56Z","created_by":"Bruno Postle","updated_at":"2026-06-12T00:14:06Z","started_at":"2026-06-11T23:50:40Z","closed_at":"2026-06-12T00:14:06Z","close_reason":"score_batch() lands in oracle.py; 35-file corpus parity verified single-vs-batch (1e-12 rel fitness, exact fail sets); 0.98 s/dom batched vs 1.27 single, x1.30","dependency_count":0,"dependent_count":1,"comment_count":0}
{"id":"homemaker-py-3l6","title":"Leaf-sharing default makes internal fitness diverge from canonical homemaker-fitness score","description":"Default is --leaf-sharing (on). Leaf sharing is a fitness-evaluation knob (fitness.py:415 quality_size, plus edge cap and count check): a shared leaf of code X is credited as satisfying k programme entries, with its size Gaussian re-centred on k*target. The evolve internal objective therefore rewards genomes that under-materialise the programme. When the winning .dom is re-scored by the canonical homemaker-fitness (leaf_sharing off), those un-materialised copies become 'missing required space ... (critical)' fails.\n\nObserved on examples/harbor-house (init.dom, budget 3M, workers 2):\n - leaf sharing ON (default): internal best 1.03e-05, but canonical score 6.73e-29 with 90 fails (15 critical missing-room).\n - --no-leaf-sharing (warm-started to full budget): internal and canonical agree at 4.19e-06, 15 fails, 0 critical -- ~9500x better than the prior best 3m.dom (4.41e-10).\n\nSo the default silently optimises an objective the canonical scorer does not credit, and writes a catastrophically worse .dom than its reported internal fitness implies.\n\nOptions to consider:\n 1. Make --no-leaf-sharing the default (strict per-leaf baseline agrees with canonical scorer).\n 2. Before writing output, re-score best-so-far with leaf_sharing off and warn (or refuse) if it regresses vs internal fitness.\n 3. Materialise/unfold shared leaves into k distinct rooms when writing the .dom, so the output satisfies the per-room programme.\n 4. Keep sharing as an early-phase relaxation only and anneal leaf_share_factor down to 0 before finishing (see related annealing investigation).","status":"open","priority":2,"issue_type":"bug","owner":"bruno@postle.net","created_at":"2026-07-05T16:24:34Z","created_by":"Bruno Postle","updated_at":"2026-07-05T16:24:34Z","dependency_count":0,"dependent_count":1,"comment_count":0}
{"id":"homemaker-py-3l6","title":"Leaf-sharing default makes internal fitness diverge from canonical homemaker-fitness score","description":"Default is --leaf-sharing (on). Leaf sharing is a fitness-evaluation knob (fitness.py:415 quality_size, plus edge cap and count check): a shared leaf of code X is credited as satisfying k programme entries, with its size Gaussian re-centred on k*target. The evolve internal objective therefore rewards genomes that under-materialise the programme. When the winning .dom is re-scored by the canonical homemaker-fitness (leaf_sharing off), those un-materialised copies become 'missing required space ... (critical)' fails.\n\nObserved on examples/harbor-house (init.dom, budget 3M, workers 2):\n - leaf sharing ON (default): internal best 1.03e-05, but canonical score 6.73e-29 with 90 fails (15 critical missing-room).\n - --no-leaf-sharing (warm-started to full budget): internal and canonical agree at 4.19e-06, 15 fails, 0 critical -- ~9500x better than the prior best 3m.dom (4.41e-10).\n\nSo the default silently optimises an objective the canonical scorer does not credit, and writes a catastrophically worse .dom than its reported internal fitness implies.\n\nOptions to consider:\n 1. Make --no-leaf-sharing the default (strict per-leaf baseline agrees with canonical scorer).\n 2. Before writing output, re-score best-so-far with leaf_sharing off and warn (or refuse) if it regresses vs internal fitness.\n 3. Materialise/unfold shared leaves into k distinct rooms when writing the .dom, so the output satisfies the per-room programme.\n 4. Keep sharing as an early-phase relaxation only and anneal leaf_share_factor down to 0 before finishing (see related annealing investigation).","notes":"FIXED (option 3+2 combined, auto-finish): leaf-sharing runs now unfold+polish+rescore before write so output is honest under the canonical scorer.\n\nImplementation:\n- driver.polish_finish(result, programme_dir, polish_budget, ...): deep-copies best, operators.unfold_shared_leaves() to materialise the count deficit, then warm-starts a leaf_sharing=False search (bootstrap=False) from the unfolded genome. polish_budget\u003c=0 -\u003e single rescore eval only (used on interrupt). Stitches evals/topologies/sigs/restarts/history onto the sharing run; history tagged share:/polish: since the two objectives are not comparable. Returned best.fitness is canonical (leaf_sharing off =\u003e internal==canonical).\n- evolve.py: new --polish-budget flag (env HOMEMAKER_POLISH_BUDGET, default -1=auto=budget//2, 0=unfold+rescore only). main() calls polish_finish when --leaf-sharing on; interrupt forces polish_budget=0 for a fast honest output.\n\nVerified end-to-end (harbor-house, budget 3000 + polish 1500): reported polish fitness 4.79788e-27 EXACTLY matches canonical homemaker-fitness, 0 critical fails (was: internal 1.03e-05 vs canonical 6.73e-29 w/ 15 critical). Tiny budget so absolute quality low but honesty restored. Tests: driver.polish_finish x3 (test_driver.py), full suite 254 pass.\n\nDefault kept --leaf-sharing on per decision (sharing's topology-search speed retained; output made honest by the finish). Schedule B in-run annealing remains kpu.","status":"in_progress","priority":2,"issue_type":"bug","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-07-05T16:24:34Z","created_by":"Bruno Postle","updated_at":"2026-07-15T08:00:00Z","started_at":"2026-07-15T06:56:29Z","dependency_count":0,"dependent_count":1,"comment_count":0}
{"id":"homemaker-py-rq2","title":"Flip share_edge_cap default-ON + rebaseline §13.x floor (hph follow-up)","description":"hph/§13.8 A/B confirmed the share-aware edge-too-long cap is positive and monotone-harmless (maple 80.3→74.0, harbor 34.7→31.0, zero regressions across 6 seeds). The fix shipped behind the SHAREEDGE/share_edge_cap knob (default OFF) so controls reproduce. This issue flips the default ON for leaf-sharing runs — it completes the §13.3 leaf-share objective relaxation on the wall measure, mirroring the pll/interior_outside default flips. Rebaselines the §13.x full-stack floor numbers (harbor 34.7→31.0, maple 80.3→74.0 become the new baseline). Couple with INTERIORO/odiv3 if those are also being default-flipped. Verify the test suite + a control re-score still reproduce post-flip.","status":"closed","priority":2,"issue_type":"task","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-06-28T20:01:11Z","created_by":"Bruno Postle","updated_at":"2026-06-28T20:39:00Z","started_at":"2026-06-28T20:32:40Z","closed_at":"2026-06-28T20:39:00Z","close_reason":"Closed","dependency_count":0,"dependent_count":0,"comment_count":0}
{"id":"homemaker-py-hph","title":"edge-too-long not share-aware: shared leaves (share\u003e1) penalised for aggregate wall length (§13.7 follow-up)","description":"DESIGN §13.7 flagged edge-too-long as harbor's top fail class (6). Dissection (experiments/diag_edge_too_long.py on the 500k probe best) shows the 6 fails are only 2 distinct locations:\n\n(1) DOMINANT ~4/6: leaf 'lllr' on both levels is a share=3 leaf (one quad = 3 rooms, 247 m2, edges 15-17 m, aspect 1.2 NEARLY SQUARE). Its walls exceed the flat 8 m cap purely because it aggregates 3 rooms — a leaf-sharing REPRESENTATION ARTIFACT, not a design flaw. §13.3 relaxed size/missing for shared leaves (quality_size centres on k*target) but edge_cost (fitness.py:474) and outside_edge_cost (fitness.py:490) still use a flat 8.0 m regardless of leaf.share. So a shared leaf is penalised for being big — the same leak §13.3 closed, on a different measure.\n\n(2) ~2/6: leaf 'llll' is a 1.2 m x 16.7 m sliver (aspect 14) at correct area — a REAL narrow-room pathology, already caught by width/proportion. Its edge-too-long is the wall it shares with lllr.\n\nNo corridors involved.\n\nPROPOSED FIX: make edge-too-long share-aware — exempt or scale the 8 m cap by leaf.share (type-guarded, as graph.leaf_share does) in edge_cost/outside_edge_cost, mirroring quality_size's k*target. Clears the ~4 artifact fails without masking the narrow sliver. Optional separate lever: lift/parametrise the flat 8 m cap for non-domestic programmes (harbor-house) — blunter, lower priority. A/B under §13 protocol (controls reproduce harbor 34.0 / maple 80.3); record verdict. Repro: experiments/diag_edge_too_long.py.","status":"closed","priority":2,"issue_type":"bug","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-06-28T13:51:26Z","created_by":"Bruno Postle","updated_at":"2026-06-28T20:03:00Z","started_at":"2026-06-28T14:46:18Z","closed_at":"2026-06-28T20:03:00Z","close_reason":"Closed","dependency_count":0,"dependent_count":0,"comment_count":0}
{"id":"homemaker-py-71d.1","title":"Diagnostic: high-budget harbor floor on full default stack — does landlocked crinkliness still dominate after interior-O?","description":"71d go/no-go probe. 71d targets landlocked crinkliness (area_outside=0, ratio-invariant) which its named fix (interior O courtyards) addresses. interior_outside now ships default-ON (erc.8), so re-measure: run harbor full default stack at high budget (1M evals, n_workers=4, seed 0) and break down the at-convergence residual — fail-type histogram + landlocked-vs-under-exposed split of crinkliness fails. If landlocked still dominates -\u003e 71d worth it; if interior-O dissolved it -\u003e 71d redundant. Verdict to DESIGN.md.","notes":"VERDICT (DESIGN §13.7): NO-GO on 71d. 500k serial full-stack harbor probe (seed 0) -\u003e 20 fails. Crinkliness collapsed 13-\u003e4, landlocked crinkliness ~13-\u003e2 of 20. Interior-O (now default) IS 71d's named fix (interior O courtyards) and already dissolved the target block. Residual now diffuse (top class edge-too-long 6), no concentrated ratio-invariant block for a targeted operator. Recommend close 71d + 7u5/jrb/u8x as superseded-by-construction.","status":"closed","priority":2,"issue_type":"task","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-06-28T06:57:44Z","created_by":"Bruno Postle","updated_at":"2026-06-28T13:19:08Z","started_at":"2026-06-28T06:58:10Z","closed_at":"2026-06-28T13:19:08Z","close_reason":"Closed","dependencies":[{"issue_id":"homemaker-py-71d.1","depends_on_id":"homemaker-py-71d","type":"parent-child","created_at":"2026-06-28T07:57:44Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":0,"dependent_count":0,"comment_count":0}
@ -76,23 +76,23 @@
{"id":"homemaker-py-erc.6","title":"Experiment: inner-loop slack-expansion objective term","description":"Inner-loop counterpart to plot-fill construction. If Diagnostic B shows the inner loop has room to expand leaves into slack but no objective gradient to do so (the scalar rewards hitting target area but not exceeding it where slack exists), add a term/incentive so the ratio optimiser pushes leaf boundaries out to consume neighbouring slack and satisfy size, rather than parking at target.\n\nCONDITIONAL on Diagnostic B: build this only if B localizes the gap to the inner loop (room to expand, no gradient); if B shows construction targets too-small dims, prefer the plot-fill construction sibling. Must preserve the §5.4 inner-loop cliff / §4.9 lexicographic protection — the term sits where it cannot displace the fail-count ordering. A/B vs §12.2 baseline, seeds 0/1/2, 20000 evals, staged, default-OFF. Record DESIGN.md §13.6.","notes":"DEPRIORITISED by Diagnostic B (§13.2). B shows the inner loop CANNOT repair undersize: the slack is depth-driven maldistribution baked into the frozen topology, and the equal-offset ratio DOF cannot shrink a 14x leaf to feed a starved one without trading into shape fails (0.5^n cliff). Wrong DOF and wrong direction — the blocker is slicing POSITION, not a missing expansion reward. Fix belongs upstream in construction/topology (erc.4 re-scoped, erc.3). Keep as a low-priority follow-up only if a depth-balanced construction still leaves a residual size gradient the inner loop could pick up.","status":"closed","priority":4,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-06-22T23:16:24Z","created_by":"Bruno Postle","updated_at":"2026-06-28T13:22:22Z","closed_at":"2026-06-28T13:22:22Z","close_reason":"wont-fix (DESIGN §13.7): Diag B (§13.2) showed the inner loop cannot repair undersize (wrong DOF — slicing position, frozen-topology ratios). Superseded by depth-balanced construction (erc.4). Condition unmet.","dependencies":[{"issue_id":"homemaker-py-erc.6","depends_on_id":"homemaker-py-erc","type":"parent-child","created_at":"2026-06-23T00:16:23Z","created_by":"Bruno Postle","metadata":"{}"},{"issue_id":"homemaker-py-erc.6","depends_on_id":"homemaker-py-erc.2","type":"blocks","created_at":"2026-06-23T00:16:47Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
{"id":"homemaker-py-erc.5","title":"Experiment: compactness-aware cuts (minimize leaf perimeter/area)","description":"Attacks the #1 factor, crinkliness (346) — a per-leaf perimeter/area property DISTINCT from proportion (aspect ratio). Proportion-aware seeding (leu.2) sizes splits but does not bias toward balanced, square-ish subdivision. Add a KD-tree-style 'keep both children compact' cut rule (prefer the cut orientation/position that minimises summed child perimeter/area) in construction.\n\nCONDITIONAL on Diagnostic A: if A shows per-leaf shape-fail is FLAT across densities (floor intrinsic to slicing density), better cuts at the same leaf count will not pay → this should be closed wont-fix in favour of leaf-sharing. Only build if A shows shape-fail RISES with density. A/B vs §12.2 baseline, seeds 0/1/2, 20000 evals, staged, default-OFF. Record DESIGN.md §13.5.","notes":"DEPRIORITISED by erc.1 verdict (§13.1): per-leaf shape-fail flat vs slicing density and cuts already squarest (_size_divisions_from_targets picks squarest rotation) yet still ~1.8 fails/leaf =\u003e little compactness headroom at fixed leaf count. Floor is intrinsic to leaf COUNT, not cut quality. Revisit only if leaf-sharing (erc.3) underdelivers.","status":"closed","priority":4,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-06-22T23:16:21Z","created_by":"Bruno Postle","updated_at":"2026-06-28T13:22:17Z","closed_at":"2026-06-28T13:22:17Z","close_reason":"wont-fix (DESIGN §13.7): Diag A (§13.1) showed the floor is intrinsic to leaf COUNT not cut quality; revisit condition was 'only if leaf-sharing underdelivers' but leaf-sharing OVER-delivered (32…39%, §13.3). Condition unmet.","dependencies":[{"issue_id":"homemaker-py-erc.5","depends_on_id":"homemaker-py-erc","type":"parent-child","created_at":"2026-06-23T00:16:21Z","created_by":"Bruno Postle","metadata":"{}"},{"issue_id":"homemaker-py-erc.5","depends_on_id":"homemaker-py-erc.1","type":"blocks","created_at":"2026-06-23T00:16:43Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
{"id":"homemaker-py-2g5","title":"Rebuild occlusion/daylight/sun subsystem in Python (post-Phase-5, after optimisation fully native)","description":"DESIGN.md §6 port scope — a whole subsystem, not a term. quality_daylight (Leaf.pm:281-296) needs Urb::Misc::Sun + Urb::Field::Occlusion (+CIESky); quality_uncrinkliness also takes the occlusion object. Indoor spaces return 1 for daylight; cost is outdoor spaces + crinkliness. Port Sun_horizontal (262980-minute normalisation) and the occlusion wall set from Dom-\u003eWalls.","acceptance_criteria":"Daylight and crinkliness factors match Perl (float tolerance) across the corpus, including multi-storey cases","notes":"Re-scoped 2026-06-12: occlusion disabled in the Urb oracle instead of ported (see homemaker-py-gp2). Native fitness ships with simple crinkliness (illumination factor = 1, in homemaker-py-gnw). This issue is now the eventual Python occlusion rebuild, only after optimisation works entirely in Python. Restores outdoor-daylight and shaded-wall selection pressure.\nReframed 2026-06-17: orthogonal to epic homemaker-py-c4c. This is fitness FIDELITY (restoring daylight + shaded-wall selection pressure to match Perl), not search CAPABILITY — it changes what 'good' means, not the search's ability to find good. It will NOT improve final designs in the sense currently sought. Stays P4, deferred until the topology-search-quality epic lands and optimisation is fully native.","status":"open","priority":4,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-06-11T23:38:25Z","created_by":"Bruno Postle","updated_at":"2026-06-17T19:14:48Z","dependency_count":0,"dependent_count":0,"comment_count":0}
{"_type":"memory","key":"experiment-seeding-pitfall-run-search-scaled-py-s","value":"Experiment seeding pitfall: run_search_scaled.py's default PH_SEED (c964…dom) is a FINISHED programme-house design — passing it warm-starts and floors at ~3 fails, NOT a blank-slate topology search. For blank-slate runs comparable to §11.5/§11.6 baselines, seed from examples/programme-house/init.dom (a bare undivided plot; driver bootstrap auto-triggers only on bare plots). Bit the 6zy sweep — first pass used c964 and falsely showed 3-fail floor across the whole grid."}
{"_type":"memory","key":"multi-storey-staircase-consistency-when-dividing-or-retyping","value":"Multi-storey staircase consistency: when dividing or retyping a circulation (C) leaf at one level, the same structural change should be propagated to the matching leaf on ALL other storeys so the stair core path is maintained. The optimizer cannot fix staircase disruptions through trial-and-error geometry alone — it requires a synchronized multi-level operator that applies the same topology change to every storey simultaneously."}
{"_type":"memory","key":"run-to-run-reproducibility-in-homemaker-layout-serial","value":"Run-to-run reproducibility in homemaker-layout: serial search (workers=1) is byte-for-byte deterministic; parallel (workers\u003e1) is now deterministic too AFTER fixing driver._run_batch to admit futures in submission order (was as_completed/completion order, bug xcy). Reproducibility holds only for a FIXED worker count — serial vs parallel differ because children-per-iteration is 1 vs n_workers (different batch granularity), which is expected, not a bug. The constructive seeder was NEVER nondeterministic: _assign_adjacency_aware has unique idx tiebreaks; comparing topologies with Python builtin hash() of the signature STRING is invalid (PYTHONHASHSEED salts str hashing per process) — use a stable hash (sha1) or genome.signature equality."}
{"_type":"memory","key":"9o5-multi-use-leaves-is-path-a-superposition","value":"9o5 multi-use leaves is path (a) — superposition as SEARCH RELAXATION that COLLAPSES to specific usage at the end, NOT path (b) loose-fit/no-collapse. Bruno's intent: codes with SIMILAR leaf requirements form an interchangeable equivalence class; during evolution the solver doesn't commit which leaf serves which specific usage (smoother landscape, no fighting over exact leaf usage); at the end the layout is CONDENSED to specific usages by brute-forcing the in-class assignment (3 interchangeable usages over 3 leaves = 3! = 6 combinations to check, pick best). 'Derive automatically' compatibility = requirement-similarity grouping. This reverses the issue's stated 'path b preferred' note."}
{"_type":"memory","key":"cli-tool-style-prefer-python-m-homemaker-module","value":"CLI tool style: prefer python -m homemaker.module --parameters pattern, installable via pip install -e . with pyproject.toml entry_points. Not standalone bin/ scripts."}
{"_type":"memory","key":"island-model-psk-14-is-a-null-priming","value":"Island model (psk, §14) is a NULL: priming a population from N converged independent elites + crossover-heavy migration does not beat best-of-N at equal total budget (maple island 124 vs control 116). The child_probe instrument shows WHY: area-matched crossover across independently-converged elites almost never synthesizes (1-3 of ~64 children beat the better parent, max drop 2-5) because the slicing encoding is non-canonical (9gp), so splices are disruptive not combinatorial. Search-machinery null #3 after graded-objective and niching/restarts; residual stays geometry/shape-bound."}
{"_type":"memory","key":"programme-house-optimisation-result-2026-06-14-15","value":"Programme-house optimisation result (2026-06-14/15): best achievable is 1 fail (l1 wrong level, score ~0.005). 0 fails is geometrically impossible: l1 (min 27m²) must occupy ll (~23m²) at level 0, which eliminates the t3-adj-C provider; dividing ll into lll(l1)+llr(C) gives llr proportion ~6:1 (fails). Python memetic optimizer achieves 1 fail in 50k evals vs Perl optimiser's 2-3 fails. Winning topology: TWO C nodes at level 0 — ll(C) for t3-adj-C via geometric contact, rl(C) for staircase via tree-sibling adjacency to rrr(O). Best .dom: scratch/from-warmstart-fixed.dom and scratch/from-compound3-fixed.dom."}
{"_type":"memory","key":"proportion-aware-constructive-seeding-leu-2-12-2","value":"Proportion-aware constructive seeding (leu.2/§12.2): sizing seed cuts from target AREAS only regresses (thin slivers wreck aspect); you must ALSO pick each cut's rotation for child squareness. It is a convergence ACCELERATOR via a deeper local optimum around the constructed topology: wins where that topology is roughly right and budget is scarce (harbor -13%, maple -10% at 20k evals) but DELAYS small programmes where the seed must be restructured by undivide (programme-house regresses at fixed budget, yet reaches the floor given budget - speed, not asymptote). Default-on. Also: n_storeys must honour storey_minimum, not just level: keys (programme-house storey_minimum:2, all rooms level:0 - was seeded 1 storey short; cq1)."}
{"_type":"memory","key":"user-preference-bruno-this-is-a-fedora-system","value":"User preference (Bruno): this is a Fedora system — NEVER install Python packages via pip without asking first; always ask whether to install the rpm via dnf (e.g. python3-cma) before considering pip. Applies to any dependency additions."}
{"_type":"memory","key":"run-to-run-reproducibility-in-homemaker-layout-serial","value":"Run-to-run reproducibility in homemaker-layout: serial search (workers=1) is byte-for-byte deterministic; parallel (workers\u003e1) is now deterministic too AFTER fixing driver._run_batch to admit futures in submission order (was as_completed/completion order, bug xcy). Reproducibility holds only for a FIXED worker count — serial vs parallel differ because children-per-iteration is 1 vs n_workers (different batch granularity), which is expected, not a bug. The constructive seeder was NEVER nondeterministic: _assign_adjacency_aware has unique idx tiebreaks; comparing topologies with Python builtin hash() of the signature STRING is invalid (PYTHONHASHSEED salts str hashing per process) — use a stable hash (sha1) or genome.signature equality."}
{"_type":"memory","key":"cli-tool-style-prefer-python-m-homemaker-module","value":"CLI tool style: prefer python -m homemaker.module --parameters pattern, installable via pip install -e . with pyproject.toml entry_points. Not standalone bin/ scripts."}
{"_type":"memory","key":"experiment-harness-gotcha-the-leaf-sharing-relaxed-objective","value":"Experiment harness gotcha: the leaf-sharing RELAXED objective (§13.3) is injected ONLY by monkeypatching fitness.load_config in the parent process (run_staged_search.py / probe scripts). This is parent-process-only and does NOT propagate into ProcessPoolExecutor workers (n_workers\u003e1), which re-import fitness fresh and score under the STRICT on-disk patterns.config -\u003e r.n_fails MISMATCH (worker strict vs parent relaxed re-score). ALL §13.x floor runs were therefore SERIAL. Any future PARALLEL leaf-sharing experiment will silently mis-score until leaf_sharing lives on disk/CLI (tracked: homemaker-py-x3b). The parallel driver itself is correct; both paths score via load_config(programme_dir)."}
{"_type":"memory","key":"homemaker-py-pythonpath-set-pythonpath-home-bruno-src","value":"homemaker-layout PYTHONPATH: package installed as 'homemaker-layout' via pip install -e . so 'import homemaker_layout' works from anywhere without PYTHONPATH. For running tests use 'python -m pytest' from project root /home/bruno/src/homemaker-layout (pyproject.toml adds src/ automatically). Never try pip show homemaker — that's the old homemaker-addon conflict."}
{"_type":"memory","key":"ld2-13-6-interior-o-seed-diagnostic-all","value":"ld2/§13.6 interior-O seed diagnostic: ALL crinkliness fails in the constructed bal+share seed are UNDER-exposed (crink\u003c0.62, landlocked rooms with no facade + no uncovered-O neighbour) — zero over-exposed sliver fails. So the erc crinkliness residual is genuine under-daylighting, validating the interior light-well premise. Default outside_divisor=6 was too sparse (null: harbor 147-\u003e142, crinkliness even rose). odiv=3 is the seed-optimal joint setting: harbor seed fails 147-\u003e129 (-18), maple 219-\u003e206 (-14), landlocked fails drop, at cost of more leaves (harbor +4, maple +8). Because it ADDS leaves it carries the §13.4 wash-out risk; A/B to convergence pending."}
{"_type":"memory","key":"multi-storey-staircase-consistency-when-dividing-or-retyping","value":"Multi-storey staircase consistency: when dividing or retyping a circulation (C) leaf at one level, the same structural change should be propagated to the matching leaf on ALL other storeys so the stair core path is maintained. The optimizer cannot fix staircase disruptions through trial-and-error geometry alone — it requires a synchronized multi-level operator that applies the same topology change to every storey simultaneously."}
{"_type":"memory","key":"urb-oracle-nondeterminism-urb-fitness-pl-output-varies","value":"Urb oracle nondeterminism: urb-fitness.pl output varies run-to-run from Perl hash-order randomisation — .fails line ORDER shuffles (compare sorted, use oracle.Score.fail_lines) and the score float can flip by ~1 ULP (compare with math.isclose rel_tol=1e-12, never ==). Not a batching artifact; affects single runs too. Matters for the Phase 3 native-fitness parity gate (homemaker-py-uxz)."}
{"_type":"memory","key":"strategy-decision-2026-06-12-bruno-occlusion-daylight","value":"Strategy decision 2026-06-12 (Bruno): occlusion/daylight is ORTHOGONAL to building a scalable optimiser. Disable it in Urb (env flag, homemaker-py-gp2) rather than port it; native fitness uses simple crinkliness (illumination factor = 1); rebuild occlusion in Python only after optimisation is fully native (homemaker-py-2g5, now P4). Consequence: all scores change when the flag flips — re-baseline corpus/.score, DESIGN \\$4.5 gains, gate bars at one clean boundary AFTER homemaker-py-1p0 closes; Phase-2 urb-evolve benchmark must run with the same flag."}
{"_type":"memory","key":"user-preference-bruno-this-is-a-fedora-system","value":"User preference (Bruno): this is a Fedora system — NEVER install Python packages via pip without asking first; always ask whether to install the rpm via dnf (e.g. python3-cma) before considering pip. Applies to any dependency additions."}
{"_type":"memory","key":"warm-x0-initialization-bug-pattern-when-a-topology","value":"warm_x0 initialization bug pattern: when a topology operator explicitly sets division ratios on a newly-created node (e.g. compound_fix sets node.division=[0.25,0.25] for t3), parent.ratios has no entry for that node (it was a leaf). warm_x0 defaults it to 0.5, corrupting the inner loop's starting point and making the operator invisible to lex comparison. Fix: only propagate child ratios for nodes where the parent node was NOT already divided; stale hidden nodes revealed by structural mutations (swap flipping b.below) must NOT contribute their pre-writeback values. See driver.py lines 259-267 (fixed 2026-06-14)."}
{"_type":"memory","key":"deceptive-valleys-in-topology-search-when-every-single","value":"Deceptive valleys in topology search: when every single-step mutation from a target state passes through a high-fail intermediary (e.g. level_fix displaces a room into 5+ new fails), a compound operator that atomically applies two coordinated changes can escape. Design compound operators to land on the low-fail state directly, bypassing the deceptive gradient. Programme-house example: level_compound_fix atomically moves the level-constrained room AND re-inserts the displaced room adjacent to C in one step (operators.py, 2026-06-14)."}
{"_type":"memory","key":"island-model-psk-14-is-a-null-priming","value":"Island model (psk, §14) is a NULL: priming a population from N converged independent elites + crossover-heavy migration does not beat best-of-N at equal total budget (maple island 124 vs control 116). The child_probe instrument shows WHY: area-matched crossover across independently-converged elites almost never synthesizes (1-3 of ~64 children beat the better parent, max drop 2-5) because the slicing encoding is non-canonical (9gp), so splices are disruptive not combinatorial. Search-machinery null #3 after graded-objective and niching/restarts; residual stays geometry/shape-bound."}
{"_type":"memory","key":"homemaker-py-pythonpath-set-pythonpath-home-bruno-src","value":"homemaker-layout PYTHONPATH: package installed as 'homemaker-layout' via pip install -e . so 'import homemaker_layout' works from anywhere without PYTHONPATH. For running tests use 'python -m pytest' from project root /home/bruno/src/homemaker-layout (pyproject.toml adds src/ automatically). Never try pip show homemaker — that's the old homemaker-addon conflict."}
{"_type":"memory","key":"urb-fitness-bug-found-fixed-2026-06-12","value":"Urb fitness bug found+fixed 2026-06-12 (patch in /home/bruno/src/urb, uncommitted): ProgrammeDriven.pm ratio_o/ratio_type grepped case-insensitively over the ratios hash and took the FIRST key — nondeterministic (x4.5 score swings) for designs with mixed-case type classes (both 'c' circulation and 'C' covered). Fixed to SUM the class (matches Is_Circulation//Is_Outside semantics); 35/35 corpus scores unchanged. CRITICAL for homemaker-py-3y7/gnw: the native port must implement class-SUM ratios. Building.pm has the same unpatched pattern (site-driven path, not used by our oracle). Also: the memetic search reward-hacked this bug before the fix — search results predating it are noise artifacts."}
{"_type":"memory","key":"urb-oracle-nondeterminism-urb-fitness-pl-output-varies","value":"Urb oracle nondeterminism: urb-fitness.pl output varies run-to-run from Perl hash-order randomisation — .fails line ORDER shuffles (compare sorted, use oracle.Score.fail_lines) and the score float can flip by ~1 ULP (compare with math.isclose rel_tol=1e-12, never ==). Not a batching artifact; affects single runs too. Matters for the Phase 3 native-fitness parity gate (homemaker-py-uxz)."}
{"_type":"memory","key":"adjacency-in-binary-slicing-tree-is-structural-not","value":"Adjacency in binary slicing tree is structural, not geometric: the inner-loop NM cannot fix topological adjacency failures. Two paths exist: (1) tree-sibling adjacency — a node is adjacent to its sibling in the tree; (2) cross-zone geometric adjacency — leaves from different subtrees that happen to share a boundary. Staircase/adjacency fails require a topology mutation that changes which nodes are siblings or which zones touch. This was proved empirically on programme-house: staircase fail from rot=0 layout could not be fixed by NM but was fixed by level_retype creating a two-C topology (2026-06-14/15)."}
{"_type":"memory","key":"correction-to-urb-fitness-bug-memory-bruno-2026","value":"CORRECTION to urb-fitness-bug memory (Bruno, 2026-06-12): 'C' is NOT a 'covered' type — Is_Covered is a geometric predicate (indoor space above). Urb's generic types are canonically UPPERCASE: C=circulation, O=outside, S=sahn (get_space_types qw/C O S/; corpus is 100% uppercase, never 'c'/'o' leaves). The mixed-case designs that fired the latent ratio_type first-match bug were created by homemaker's own operator type pool emitting lowercase 'c'/'o' — fixed: driver/operators now emit uppercase generics only, and class checks use t[0].lower() in 'cos'. The Urb class-sum patch stays as defensive hardening (zero impact on canonical designs). Native port (3y7/gnw): treat type classes case-insensitively, generics canonically uppercase."}
{"_type":"memory","key":"urb-fitness-bug-found-fixed-2026-06-12","value":"Urb fitness bug found+fixed 2026-06-12 (patch in /home/bruno/src/urb, uncommitted): ProgrammeDriven.pm ratio_o/ratio_type grepped case-insensitively over the ratios hash and took the FIRST key — nondeterministic (x4.5 score swings) for designs with mixed-case type classes (both 'c' circulation and 'C' covered). Fixed to SUM the class (matches Is_Circulation//Is_Outside semantics); 35/35 corpus scores unchanged. CRITICAL for homemaker-py-3y7/gnw: the native port must implement class-SUM ratios. Building.pm has the same unpatched pattern (site-driven path, not used by our oracle). Also: the memetic search reward-hacked this bug before the fix — search results predating it are noise artifacts."}
{"_type":"memory","key":"9o5-multi-use-leaves-is-path-a-superposition","value":"9o5 multi-use leaves is path (a) — superposition as SEARCH RELAXATION that COLLAPSES to specific usage at the end, NOT path (b) loose-fit/no-collapse. Bruno's intent: codes with SIMILAR leaf requirements form an interchangeable equivalence class; during evolution the solver doesn't commit which leaf serves which specific usage (smoother landscape, no fighting over exact leaf usage); at the end the layout is CONDENSED to specific usages by brute-forcing the in-class assignment (3 interchangeable usages over 3 leaves = 3! = 6 combinations to check, pick best). 'Derive automatically' compatibility = requirement-similarity grouping. This reverses the issue's stated 'path b preferred' note."}
{"_type":"memory","key":"deceptive-valleys-in-topology-search-when-every-single","value":"Deceptive valleys in topology search: when every single-step mutation from a target state passes through a high-fail intermediary (e.g. level_fix displaces a room into 5+ new fails), a compound operator that atomically applies two coordinated changes can escape. Design compound operators to land on the low-fail state directly, bypassing the deceptive gradient. Programme-house example: level_compound_fix atomically moves the level-constrained room AND re-inserts the displaced room adjacent to C in one step (operators.py, 2026-06-14)."}
{"_type":"memory","key":"experiment-harness-gotcha-the-leaf-sharing-relaxed-objective","value":"Experiment harness gotcha: the leaf-sharing RELAXED objective (§13.3) is injected ONLY by monkeypatching fitness.load_config in the parent process (run_staged_search.py / probe scripts). This is parent-process-only and does NOT propagate into ProcessPoolExecutor workers (n_workers\u003e1), which re-import fitness fresh and score under the STRICT on-disk patterns.config -\u003e r.n_fails MISMATCH (worker strict vs parent relaxed re-score). ALL §13.x floor runs were therefore SERIAL. Any future PARALLEL leaf-sharing experiment will silently mis-score until leaf_sharing lives on disk/CLI (tracked: homemaker-py-x3b). The parallel driver itself is correct; both paths score via load_config(programme_dir)."}
{"_type":"memory","key":"never-use-corpus-filenames-candidate-001-dom-candidate","value":"Never use corpus filenames (candidate-001.dom, candidate-002.dom, generated.dom, init.dom, etc.) as --output targets when running experiments. These are test fixtures. Always write experimental outputs to scratch/ or a timestamped path. Lesson from 2026-06-14: warm-start runs overwrote candidate-001/002.dom and broke graph tests."}
{"_type":"memory","key":"strategy-decision-2026-06-12-bruno-occlusion-daylight","value":"Strategy decision 2026-06-12 (Bruno): occlusion/daylight is ORTHOGONAL to building a scalable optimiser. Disable it in Urb (env flag, homemaker-py-gp2) rather than port it; native fitness uses simple crinkliness (illumination factor = 1); rebuild occlusion in Python only after optimisation is fully native (homemaker-py-2g5, now P4). Consequence: all scores change when the flag flips — re-baseline corpus/.score, DESIGN \\$4.5 gains, gate bars at one clean boundary AFTER homemaker-py-1p0 closes; Phase-2 urb-evolve benchmark must run with the same flag."}
{"_type":"memory","key":"experiment-seeding-pitfall-run-search-scaled-py-s","value":"Experiment seeding pitfall: run_search_scaled.py's default PH_SEED (c964…dom) is a FINISHED programme-house design — passing it warm-starts and floors at ~3 fails, NOT a blank-slate topology search. For blank-slate runs comparable to §11.5/§11.6 baselines, seed from examples/programme-house/init.dom (a bare undivided plot; driver bootstrap auto-triggers only on bare plots). Bit the 6zy sweep — first pass used c964 and falsely showed 3-fail floor across the whole grid."}

View file

@ -537,6 +537,88 @@ def search(
return result
def polish_finish(
result: SearchResult,
programme_dir: str | Path,
*,
polish_budget: int,
pop_size: int = 8,
child_budget: int = 80,
p_crossover: float = 0.2,
seed: int = 0,
n_workers: int = 1,
superpose: bool = False,
rescore_budget: int = 200,
log=None,
) -> SearchResult:
"""homemaker-py-3l6: convert a leaf-sharing run's dishonest best into a
canonically-scored, materialised output.
A sharing run's internal objective credits a shared leaf (``share=k``) as k
programme rooms with its size target re-centred on ``k*target``, so
``result.best`` looks good internally but is k1 rooms short per shared leaf
under the canonical (sharing-off) scorer the divergence this bug is about.
This:
1. **Unfolds** every live shared leaf into k distinct sibling rooms
(:func:`operators.unfold_shared_leaves`), paying down the materialisation
deficit that otherwise leaves the de-shared genome deep in the missing-room
fail hole (yaa: naive warm-start without unfold stalls ~60× worse).
2. **Polishes** the unfolded genome with a warm-started ``leaf_sharing=False``
search (``polish_budget`` evals) so the freshly materialised children get
their proportion/width/size cleaned up. yaa proved this unfold-then-polish
path catches the direct no-sharing route (harbor-house 4.19e-06).
With ``polish_budget <= 0`` the polish is skipped: the unfolded genome is just
re-optimised once and canonically scored (honest output, no extra search
used on interrupt). Either way the returned result's ``best.fitness`` is the
canonical score (leaf_sharing off internal == canonical), and eval /
topology / history accounting is stitched onto the sharing run.
"""
def _log(msg: str) -> None:
if log:
log(msg)
if result.best is None:
return result
unfolded = copy.deepcopy(result.best.root)
n_created = operators.unfold_shared_leaves(unfolded)
_log(f"[finish] unfold: materialised {n_created} shared-leaf "
f"{'copy' if n_created == 1 else 'copies'}")
if polish_budget > 0:
r2 = search(
unfolded, programme_dir, budget=polish_budget, pop_size=pop_size,
child_budget=child_budget, p_crossover=p_crossover, seed=seed,
n_workers=n_workers, bootstrap=False, leaf_sharing=False,
superpose=superpose, log=log,
)
else:
# No polish: re-optimise the unfolded genome's ratios once and score it
# canonically so the written .dom and reported fitness are honest.
ind, used = _evaluate(
unfolded, programme_dir, None, x0=None, budget=rescore_budget,
inner_kw={}, lineage="unfold", leaf_sharing=False, superpose=superpose)
r2 = SearchResult(best=ind, population=[ind], n_evals=used, n_topologies=1)
r2.n_distinct_signatures = 1
r2.history = [(0, ind.fitness, ind.lineage)]
# Stitch the polish/rescore onto the sharing run so totals are cumulative and
# the history shows the phase change (sharing fitness is not comparable to the
# canonical polish fitness, so the two phases are tagged, not merged linearly).
r2.history = (
[(e, f, f"share:{lin}") for e, f, lin in result.history]
+ [(e + result.n_evals, f, f"polish:{lin}") for e, f, lin in r2.history]
)
r2.n_evals += result.n_evals
r2.n_topologies += result.n_topologies
r2.n_distinct_signatures += result.n_distinct_signatures
r2.n_restarts += result.n_restarts
r2.interrupted = r2.interrupted or result.interrupted
return r2
def search_staged(
seed_root: dom.Node,
programme_dir: str | Path,

View file

@ -17,6 +17,11 @@ Options:
--child-budget N per-child budget (default: $HOMEMAKER_CHILD_BUDGET or 80)
--workers N parallel workers (default: $HOMEMAKER_WORKERS or 1)
--seed N RNG seed (default: $HOMEMAKER_SEED or 0)
--polish-budget N after a leaf-sharing run, unfold the shared leaves and
run N extra no-sharing evals so the written .dom is honest
under the canonical scorer (homemaker-py-3l6). -1 = auto
(budget//2); 0 = unfold + rescore only. Ignored with
--no-leaf-sharing. (default: $HOMEMAKER_POLISH_BUDGET or -1)
--output PATH output .dom path (default: <seed_stem>_evolved.dom
next to seed; use - for stdout)
@ -90,6 +95,16 @@ def _parse_args(argv=None) -> argparse.Namespace:
"requirements) form equivalence classes and each candidate "
"collapses every superposed leaf to its best in-class usage "
"before scoring (default: off)")
p.add_argument("--polish-budget", type=int,
default=_env_int("HOMEMAKER_POLISH_BUDGET", -1),
metavar="N",
help="homemaker-py-3l6: after a leaf-sharing run, unfold the "
"shared leaves and run this many extra evals of "
"no-sharing local search to clean up the materialised "
"rooms before write, so the output is honest under the "
"canonical scorer. -1 = auto (budget//2); 0 = unfold + "
"rescore only, no polish. Ignored with --no-leaf-sharing "
"(default: -1)")
p.add_argument("--output", type=Path, default=None, metavar="PATH",
help="output .dom path (- for stdout)")
return p.parse_args(argv)
@ -151,6 +166,34 @@ def main(argv=None) -> int:
log=lambda m: print(m, file=sys.stderr, flush=True),
)
# homemaker-py-3l6: a leaf-sharing run's internal best is scored against a
# sharing-credited objective (a shared leaf counts as k programme rooms), so
# r.best is dishonest under the canonical scorer — it is k-1 rooms short per
# shared leaf. Unfold those leaves and warm-start a no-sharing polish so the
# written .dom is honest AND its materialised rooms are cleaned up (yaa: the
# unfold-then-polish path catches the direct no-sharing route). After this,
# r.best.fitness is the canonical score (leaf_sharing off ⇒ internal == canon).
if args.leaf_sharing and r.best is not None:
polish_budget = args.budget // 2 if args.polish_budget < 0 else args.polish_budget
# An interrupted sharing run still needs an honest output, but the user
# asked to stop — unfold and rescore only, skip the long polish phase.
if r.interrupted:
polish_budget = 0
print(file=sys.stderr)
print(f"--- finishing (homemaker-py-3l6): unfold + polish "
f"{polish_budget} evals ---", file=sys.stderr, flush=True)
r = driver.polish_finish(
r, programme_dir,
polish_budget=polish_budget,
pop_size=args.pop,
child_budget=args.child_budget,
p_crossover=0.2,
seed=args.seed,
n_workers=args.workers,
superpose=args.superpose,
log=lambda m: print(m, file=sys.stderr, flush=True),
)
elapsed = time.perf_counter() - t0
print(file=sys.stderr)

View file

@ -256,3 +256,61 @@ def test_search_parallel_is_reproducible():
a = run()
b = run()
assert a == b, "parallel search is not reproducible run-to-run"
def _shared_best_result() -> driver.SearchResult:
"""A SearchResult whose best carries a live 3-room shared leaf (share=3),
plus a distinct C leaf the harbor-house pathology in miniature."""
root = dom.Node(node=[[0, 0], [12, 0], [12, 8], [0, 8]],
height=2.7, wall_outer=0.25, wall_inner=0.08,
rotation=0, division=[0.5, 0.5])
root.left = dom.Node(type="n", share=3, share_type="n")
root.right = dom.Node(type="C")
dom._link(root)
best = driver.Individual(root=root, fitness=1e-5, n_fails=3, ratios={},
lineage="construct/0")
r = driver.SearchResult(best=best, population=[best], n_evals=1000,
n_topologies=5, n_distinct_signatures=4, n_restarts=1)
r.history = [(80, 1e-6, "construct/0"), (160, 1e-5, "core_divide noop")]
return r
def test_polish_finish_unfolds_and_stitches_rescore(fake_inner):
# homemaker-py-3l6, polish_budget<=0: unfold the shared leaf, rescore once
# under leaf_sharing off, and stitch accounting/history onto the sharing run.
r0 = _shared_best_result()
r = driver.polish_finish(r0, CORPUS, polish_budget=0, rescore_budget=150)
# the shared leaf is materialised into 3 distinct n rooms, stamps cleared
leaves = r.best.root.leaves()
assert sum(1 for lf in leaves if lf.type == "n") == 3
assert all(lf.share == 1 for lf in leaves)
# accounting is cumulative (1000 sharing evals + one 150-eval rescore)
assert r.n_evals == 1000 + 150
assert r.n_topologies == 5 + 1
assert r.n_distinct_signatures == 4 + 1
assert r.n_restarts == 1
# history keeps both phases, tagged so the objective change is visible
assert [lin for *_, lin in r.history[:2]] == [
"share:construct/0", "share:core_divide noop"]
assert r.history[-1][2].startswith("polish:")
# the rescore ran with leaf_sharing off (no sharing override reaches the inner)
assert not fake_inner[-1]["kw"].get("conf_overrides")
def test_polish_finish_runs_polish_search(fake_inner):
# polish_budget>0: a warm-started no-sharing search runs from the unfolded
# genome and its evals accrue on top of the sharing run.
r0 = _shared_best_result()
r = driver.polish_finish(r0, CORPUS, polish_budget=400, pop_size=3,
child_budget=80, seed=1)
assert r.n_evals > 1000 + 400 - 80 # sharing 1000 + ~400 polish evals
assert r.best.root.leaves() # a valid materialised genome survived
assert sum(1 for lf in r.best.root.leaves() if lf.share > 1) == 0
assert r.history[0][2].startswith("share:")
assert any(lin.startswith("polish:") for *_, lin in r.history)
def test_polish_finish_noop_without_best():
empty = driver.SearchResult(best=None, population=[], n_evals=0, n_topologies=0)
assert driver.polish_finish(empty, CORPUS, polish_budget=100) is empty