2026-06-12 14:22:26 +01:00
|
|
|
|
"""Memetic search driver, small-scale (DESIGN.md §5, §7 Phase 2).
|
|
|
|
|
|
|
|
|
|
|
|
Steady-state memetic GA over topology: the outer loop owns *topology only*
|
|
|
|
|
|
(operators.py moves on decoded Node trees); every child's geometry is
|
|
|
|
|
|
delegated to the warm-started inner loop (innerloop.optimise), and the
|
|
|
|
|
|
optimised ratios are written back into the individual (Lamarckian — measured
|
|
|
|
|
|
mandatory, homemaker-py-8cs: cold starts never catch up at equal budget).
|
|
|
|
|
|
|
|
|
|
|
|
Budgets are stated and accounted in **oracle evaluations** (scored .dom
|
|
|
|
|
|
files), never generations (§4.6 arithmetic). This driver is deliberately
|
|
|
|
|
|
small-scale for the Phase-2 proof on the batched Perl oracle; scaling up
|
|
|
|
|
|
waits for the native fitness (Phase 3).
|
2026-06-13 23:29:12 +01:00
|
|
|
|
|
|
|
|
|
|
Cold-start bootstrap (homemaker-py-0px): when the seed is an undivided bare
|
|
|
|
|
|
plot, the search auto-generates a diverse initial population by randomly
|
|
|
|
|
|
applying divide mutations until each topology has approximately the programme
|
|
|
|
|
|
room count, then evaluates all pop_size individuals before the memetic loop
|
|
|
|
|
|
begins. This crosses the zero-feasibility region that single-seed chaining
|
|
|
|
|
|
cannot escape.
|
2026-06-14 06:55:58 +01:00
|
|
|
|
|
|
|
|
|
|
Parallelism (homemaker-py-5l6): ``n_workers > 1`` evaluates a batch of
|
|
|
|
|
|
children per iteration using ``concurrent.futures.ProcessPoolExecutor``.
|
|
|
|
|
|
Each worker is independent (NativeEvaluator has no shared mutable state).
|
|
|
|
|
|
The geometry module-level cache is cleared in each worker after fork to
|
|
|
|
|
|
prevent stale id-keyed entries inherited from the parent process.
|
2026-06-12 14:22:26 +01:00
|
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
|
|
from __future__ import annotations
|
|
|
|
|
|
|
|
|
|
|
|
import copy
|
2026-06-18 22:33:29 +01:00
|
|
|
|
import functools
|
2026-06-12 14:22:26 +01:00
|
|
|
|
from dataclasses import dataclass, field
|
|
|
|
|
|
from pathlib import Path
|
|
|
|
|
|
|
|
|
|
|
|
import numpy as np
|
|
|
|
|
|
|
homemaker-py-6xh: wire shape-curve DP into driver.py as an NM warm-start
Promotes the validated shape-curve DP (experiments/shapecurve_spike.py,
2g7.4, DESIGN.md §37.2) from a reference-only spike into
src/homemaker_layout/shapecurve.py, and wires it into driver._evaluate as a
warm-start for innerloop.optimise: when eligible (single storey, no
leaf_sharing/superpose/max_share/multi_use) and no caller-supplied x0, the
DP's exact shape-feasible ratio point is written onto the tree before NM
runs, off by default (shapecurve_warmstart=/--shapecurve-warmstart).
Caught and fixed a latent bug promoting the spike: realise() could leave
numpy.float64 in `division`, which yaml.safe_dump can't serialise — the
original spike never round-tripped through dom.dumps so this was never hit.
A/B on harbor-house-l0 (experiments/ab_shapecurve_warmstart.py, budget=2000,
5 seeds): mean total fails 16.6 (on) vs 19.6 (off), ~3.5x mean fitness
improvement; mean hard-fail count alone was a noise-level wash at this
sample size. Full writeup in DESIGN.md §37.4.
Deliberately deferred to new tracked beads (children of 2g7): DP-exact hard
pre-filter (wkh), multi-storey below-link support (koo), leaf_sharing/
co_type modelling (tym), true skew-quad polygon algebra (ekc) — 6xh stays
in_progress pending those.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-03 18:43:28 +01:00
|
|
|
|
from . import dom, fitness, genome, innerloop, operators, programme, shapecurve
|
2026-06-12 14:22:26 +01:00
|
|
|
|
|
2026-06-14 09:20:03 +01:00
|
|
|
|
_CHILD_INNER_KW: dict = {}
|
2026-06-12 14:22:26 +01:00
|
|
|
|
|
2026-06-18 22:33:29 +01:00
|
|
|
|
|
2026-07-16 08:38:08 +01:00
|
|
|
|
def _overrides_for(leaf_sharing: bool, superpose: bool,
|
2026-07-18 18:44:24 +01:00
|
|
|
|
max_share: int | None = None,
|
2026-07-19 20:35:18 +01:00
|
|
|
|
conn_grade: bool = False,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
collapse_insearch: bool = True,
|
|
|
|
|
|
multi_use: bool = False) -> dict | None:
|
2026-07-16 08:38:08 +01:00
|
|
|
|
"""Run-level conf overrides for the native evaluator (None when all off).
|
|
|
|
|
|
|
|
|
|
|
|
``max_share`` (homemaker-py-kpu) overrides the evaluator's ``leaf_share_max``
|
|
|
|
|
|
grain cap for the in-run annealing ramp; ``None`` leaves the config default.
|
2026-07-18 18:44:24 +01:00
|
|
|
|
``conn_grade`` (homemaker-py-qi6) turns the graded proximity scalar into the
|
2026-07-19 20:35:18 +01:00
|
|
|
|
circulation-connectivity signal (§18). ``collapse_insearch`` (homemaker-py-
|
|
|
|
|
|
qpk) runs the 94g global cell<->room collapse inside every fitness eval
|
|
|
|
|
|
instead of once at finish time.
|
2026-07-16 08:38:08 +01:00
|
|
|
|
"""
|
9o5: type superposition + per-eval collapse (multi-use leaves)
Interchangeable codes (similar size/width/proportion, compatible level/stack,
no adjacency edge) form equivalence classes derived from the programme. With
--superpose (default off), each fitness eval COLLAPSES every superposed leaf to
its best in-class usage via an optimal supply->demand assignment (brute force
<=C! within cap C=4, scipy Hungarian beyond), then scores the condensed types.
Because collapse re-types on the unmerged tree before all checks, counts /
adjacency / quality are unchanged downstream -- no Node field, no graph/operator
changes -- and default OFF is bit-identical.
- programme.py: derive_interchange_classes + interchangeable (S1-S4, locked
thresholds R_SIZE=1.5/R_WIDTH=1.3/R_PROP=1.5, CLASS_CAP=4)
- fitness.py: collapse_superposition, _best_assignment, _usage_quality;
superpose/superpose_class_cap conf knobs; collapse hooked into _evaluate_full
- driver.py/evolve.py: superpose flag plumbed beside leaf_sharing; --superpose
- tests/test_superposition.py: 17 tests (derivation, assignment, end-to-end)
Closes homemaker-py-9o5 (build); validation A/B is homemaker-py-xi7.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:08:46 +01:00
|
|
|
|
ov: dict = {}
|
|
|
|
|
|
if leaf_sharing:
|
|
|
|
|
|
ov["leaf_sharing"] = True
|
|
|
|
|
|
if superpose:
|
|
|
|
|
|
ov["superpose"] = True
|
2026-07-16 08:38:08 +01:00
|
|
|
|
if max_share is not None:
|
|
|
|
|
|
ov["leaf_share_max"] = int(max_share)
|
2026-07-18 18:44:24 +01:00
|
|
|
|
if conn_grade:
|
|
|
|
|
|
ov["conn_grade"] = True
|
2026-07-19 20:35:18 +01:00
|
|
|
|
if collapse_insearch:
|
|
|
|
|
|
ov["collapse_insearch"] = True
|
2026-07-31 00:16:12 +01:00
|
|
|
|
if multi_use:
|
|
|
|
|
|
ov["multi_use"] = True
|
9o5: type superposition + per-eval collapse (multi-use leaves)
Interchangeable codes (similar size/width/proportion, compatible level/stack,
no adjacency edge) form equivalence classes derived from the programme. With
--superpose (default off), each fitness eval COLLAPSES every superposed leaf to
its best in-class usage via an optimal supply->demand assignment (brute force
<=C! within cap C=4, scipy Hungarian beyond), then scores the condensed types.
Because collapse re-types on the unmerged tree before all checks, counts /
adjacency / quality are unchanged downstream -- no Node field, no graph/operator
changes -- and default OFF is bit-identical.
- programme.py: derive_interchange_classes + interchangeable (S1-S4, locked
thresholds R_SIZE=1.5/R_WIDTH=1.3/R_PROP=1.5, CLASS_CAP=4)
- fitness.py: collapse_superposition, _best_assignment, _usage_quality;
superpose/superpose_class_cap conf knobs; collapse hooked into _evaluate_full
- driver.py/evolve.py: superpose flag plumbed beside leaf_sharing; --superpose
- tests/test_superposition.py: 17 tests (derivation, assignment, end-to-end)
Closes homemaker-py-9o5 (build); validation A/B is homemaker-py-xi7.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:08:46 +01:00
|
|
|
|
return ov or None
|
|
|
|
|
|
|
|
|
|
|
|
|
2026-06-18 22:33:29 +01:00
|
|
|
|
@functools.lru_cache(maxsize=None)
|
9o5: type superposition + per-eval collapse (multi-use leaves)
Interchangeable codes (similar size/width/proportion, compatible level/stack,
no adjacency edge) form equivalence classes derived from the programme. With
--superpose (default off), each fitness eval COLLAPSES every superposed leaf to
its best in-class usage via an optimal supply->demand assignment (brute force
<=C! within cap C=4, scipy Hungarian beyond), then scores the condensed types.
Because collapse re-types on the unmerged tree before all checks, counts /
adjacency / quality are unchanged downstream -- no Node field, no graph/operator
changes -- and default OFF is bit-identical.
- programme.py: derive_interchange_classes + interchangeable (S1-S4, locked
thresholds R_SIZE=1.5/R_WIDTH=1.3/R_PROP=1.5, CLASS_CAP=4)
- fitness.py: collapse_superposition, _best_assignment, _usage_quality;
superpose/superpose_class_cap conf knobs; collapse hooked into _evaluate_full
- driver.py/evolve.py: superpose flag plumbed beside leaf_sharing; --superpose
- tests/test_superposition.py: 17 tests (derivation, assignment, end-to-end)
Closes homemaker-py-9o5 (build); validation A/B is homemaker-py-xi7.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:08:46 +01:00
|
|
|
|
def _fitness_for(programme_dir: str, leaf_sharing: bool = False,
|
2026-07-16 08:38:08 +01:00
|
|
|
|
superpose: bool = False,
|
2026-07-18 18:44:24 +01:00
|
|
|
|
max_share: int | None = None,
|
2026-07-19 20:35:18 +01:00
|
|
|
|
conn_grade: bool = False,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
collapse_insearch: bool = True,
|
|
|
|
|
|
multi_use: bool = False) -> "fitness.Fitness":
|
2026-06-28 22:04:35 +01:00
|
|
|
|
"""Cached Fitness evaluator per (programme dir, leaf_sharing) (config load is
|
|
|
|
|
|
the cost).
|
|
|
|
|
|
|
|
|
|
|
|
Used only to read the graded proximity scalar (§11.4) and the shape-fail
|
|
|
|
|
|
feasibility proxy off an already-optimised tree in :func:`_evaluate`; the
|
|
|
|
|
|
inner loop's own NativeEvaluator is untouched. ``leaf_sharing`` (homemaker-py-
|
|
|
|
|
|
x3b) injects the run-level flag so this off-tree scorer agrees with the
|
|
|
|
|
|
inner loop instead of reading the on-disk (sharing-free) patterns.config.
|
|
|
|
|
|
Cached per process — workers fork their own copy.
|
2026-06-18 22:33:29 +01:00
|
|
|
|
"""
|
2026-07-19 20:35:18 +01:00
|
|
|
|
overrides = _overrides_for(leaf_sharing, superpose, max_share, conn_grade,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
collapse_insearch, multi_use)
|
2026-06-28 22:04:35 +01:00
|
|
|
|
conf, cost = fitness.load_config(programme_dir, overrides=overrides)
|
2026-06-18 22:33:29 +01:00
|
|
|
|
return fitness.Fitness(conf, cost)
|
|
|
|
|
|
|
2026-06-20 18:54:48 +01:00
|
|
|
|
|
|
|
|
|
|
@functools.lru_cache(maxsize=None)
|
|
|
|
|
|
def _reqs_for(programme_dir: str) -> dict:
|
|
|
|
|
|
"""Cached programme requirements per dir, for the §12.3 shape-feasibility
|
|
|
|
|
|
pre-filter (homemaker-py-9gp.1). Cached per process — workers fork a copy."""
|
|
|
|
|
|
return programme.load_programme_dir(programme_dir)
|
|
|
|
|
|
|
2026-06-12 14:22:26 +01:00
|
|
|
|
# storey add/delete are drastic (geometry perturbation 0.25-0.33 and a
|
2026-06-17 22:51:58 +01:00
|
|
|
|
# deleted storey stacks missing-space failures) — sample them rarely.
|
|
|
|
|
|
# place_missing is the high-leverage §11.2 repair: it noops cheaply once the
|
|
|
|
|
|
# required set is complete, so over-sampling it costs little and directly
|
2026-07-25 12:20:22 +01:00
|
|
|
|
# attacks the dominant missing-space failure mode. bridge_circulation was
|
|
|
|
|
|
# tried at the same 2.0 weight (homemaker-py-lj3) but a larger-N A/B
|
|
|
|
|
|
# (homemaker-py-qjg, DESIGN.md §22) found no total-fail benefit and MORE
|
|
|
|
|
|
# trajectory-divergence-induced new not-connected fails than at the
|
|
|
|
|
|
# uniform default weight -- reverted, left at implicit uniform weight.
|
2026-07-26 09:31:42 +01:00
|
|
|
|
_MUTATION_WEIGHTS = {"level_add": 0.2, "level_delete": 0.2, "place_missing": 2.0,
|
|
|
|
|
|
"ruin_recreate": 3.0}
|
2026-06-12 14:22:26 +01:00
|
|
|
|
|
|
|
|
|
|
|
2026-06-14 06:55:58 +01:00
|
|
|
|
def _worker_init() -> None:
|
|
|
|
|
|
"""Clear the geometry cache in each forked worker process.
|
|
|
|
|
|
|
|
|
|
|
|
geometry._cache is keyed by id(node) (Python memory address). After
|
|
|
|
|
|
fork the inherited cache holds parent-process ids that could collide
|
|
|
|
|
|
with freshly allocated nodes in the worker, producing wrong hits.
|
|
|
|
|
|
"""
|
|
|
|
|
|
from . import geometry
|
|
|
|
|
|
geometry.clear_cache()
|
|
|
|
|
|
|
|
|
|
|
|
|
2026-06-12 14:22:26 +01:00
|
|
|
|
@dataclass
|
|
|
|
|
|
class Individual:
|
|
|
|
|
|
root: dom.Node
|
|
|
|
|
|
fitness: float
|
|
|
|
|
|
n_fails: int
|
|
|
|
|
|
ratios: dict[tuple[int, str], float]
|
|
|
|
|
|
lineage: str = "seed"
|
2026-06-18 22:33:29 +01:00
|
|
|
|
grade: float = 0.0 # §11.4 graded proximity; secondary comparator key only
|
2026-06-18 23:42:39 +01:00
|
|
|
|
sig: str = "" # §11.5 structural topology signature; niching key
|
homemaker-py-2g7.3: hard/soft fail tiering behind --use-tiers flag
Splits the flat outer-search comparator (-n_fails, fitness) into a tiered
(-n_hard, -n_soft, fitness) so search budget stops being spent polishing
SOFT shape fails (crinkliness/proportion/size/width/edge-too-long/
staircase-volume) while HARD structural fails (missing space, wrong/
required level, level/circulation/vertical connectivity, adjacency,
stairs, covered-outside, storey limits, public access) remain unfixed.
fitness.classify_fail_tier/tier_counts classify every fail string emitted
across fitness.py and graph.py, raising on anything unrecognised so new
fail sites must declare a tier. Validated against all real fail strings in
the checked-in corpus plus every fail-emission call site read from source.
driver.Individual gains n_hard/n_soft (populated from innerloop.Result.
fail_lines); search(use_tiers=...) swaps the comparator when set (default
off, so existing runs are unaffected — inner-loop 0.5^n cliff untouched).
evolve.py exposes --use-tiers / HOMEMAKER_USE_TIERS.
experiments/tier_ab_2g7_3.py runs the acceptance A/B (harbor+maple, 3
seeds, 20k evals) in the background; results pending.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-02 16:00:39 +01:00
|
|
|
|
n_hard: int = 0 # homemaker-py-2g7.3: hard-fail count (structural, tiered comparator)
|
|
|
|
|
|
n_soft: int = 0 # homemaker-py-2g7.3: soft-fail count (shape/quality, tiered comparator)
|
2026-06-12 14:22:26 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@dataclass
|
|
|
|
|
|
class SearchResult:
|
|
|
|
|
|
best: Individual
|
|
|
|
|
|
population: list[Individual]
|
|
|
|
|
|
n_evals: int
|
|
|
|
|
|
n_topologies: int
|
|
|
|
|
|
history: list[tuple[int, float, str]] = field(default_factory=list)
|
|
|
|
|
|
# (oracle evals consumed, new best fitness, lineage) per improvement
|
2026-06-14 07:48:13 +01:00
|
|
|
|
interrupted: bool = False
|
2026-06-18 23:42:39 +01:00
|
|
|
|
n_distinct_signatures: int = 0 # §11.5 total distinct topologies ever admitted
|
|
|
|
|
|
diversity_history: list[tuple[int, int, int]] = field(default_factory=list)
|
|
|
|
|
|
# (evals, distinct sigs in population, cumulative distinct sigs seen)
|
|
|
|
|
|
n_restarts: int = 0 # §11.5 diversity restarts triggered
|
2026-06-12 14:22:26 +01:00
|
|
|
|
|
|
|
|
|
|
|
2026-06-13 23:29:12 +01:00
|
|
|
|
def random_topology(seed_root: dom.Node, n_leaves: int,
|
|
|
|
|
|
rng: np.random.Generator, types: list[str]) -> dom.Node:
|
|
|
|
|
|
"""Grow a random topology from ``seed_root`` by repeated divide mutations.
|
|
|
|
|
|
|
|
|
|
|
|
Applies ``mutate_divide`` until the total leaf count across all storeys
|
|
|
|
|
|
reaches ``n_leaves``. The result is a deep copy; ``seed_root`` is
|
|
|
|
|
|
unchanged.
|
|
|
|
|
|
"""
|
|
|
|
|
|
root = copy.deepcopy(seed_root)
|
|
|
|
|
|
while sum(len(lvl.leaves()) for lvl in dom.levels(root)) < n_leaves:
|
|
|
|
|
|
root, _ = operators.mutate_divide(root, rng, types)
|
|
|
|
|
|
return root
|
|
|
|
|
|
|
|
|
|
|
|
|
2026-06-12 14:22:26 +01:00
|
|
|
|
def _evaluate(root: dom.Node, programme_dir, urb_root, x0, budget, inner_kw,
|
2026-06-20 18:54:48 +01:00
|
|
|
|
lineage: str, want_grade: bool = False,
|
|
|
|
|
|
feasibility_max_shape_fails: int | None = None,
|
2026-06-28 22:04:35 +01:00
|
|
|
|
best_n_fails: int | None = None,
|
9o5: type superposition + per-eval collapse (multi-use leaves)
Interchangeable codes (similar size/width/proportion, compatible level/stack,
no adjacency edge) form equivalence classes derived from the programme. With
--superpose (default off), each fitness eval COLLAPSES every superposed leaf to
its best in-class usage via an optimal supply->demand assignment (brute force
<=C! within cap C=4, scipy Hungarian beyond), then scores the condensed types.
Because collapse re-types on the unmerged tree before all checks, counts /
adjacency / quality are unchanged downstream -- no Node field, no graph/operator
changes -- and default OFF is bit-identical.
- programme.py: derive_interchange_classes + interchangeable (S1-S4, locked
thresholds R_SIZE=1.5/R_WIDTH=1.3/R_PROP=1.5, CLASS_CAP=4)
- fitness.py: collapse_superposition, _best_assignment, _usage_quality;
superpose/superpose_class_cap conf knobs; collapse hooked into _evaluate_full
- driver.py/evolve.py: superpose flag plumbed beside leaf_sharing; --superpose
- tests/test_superposition.py: 17 tests (derivation, assignment, end-to-end)
Closes homemaker-py-9o5 (build); validation A/B is homemaker-py-xi7.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:08:46 +01:00
|
|
|
|
leaf_sharing: bool = False,
|
2026-07-16 08:38:08 +01:00
|
|
|
|
superpose: bool = False,
|
2026-07-18 18:44:24 +01:00
|
|
|
|
max_share: int | None = None,
|
2026-07-19 20:35:18 +01:00
|
|
|
|
conn_grade: bool = False,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
collapse_insearch: bool = True,
|
homemaker-py-6xh: wire shape-curve DP into driver.py as an NM warm-start
Promotes the validated shape-curve DP (experiments/shapecurve_spike.py,
2g7.4, DESIGN.md §37.2) from a reference-only spike into
src/homemaker_layout/shapecurve.py, and wires it into driver._evaluate as a
warm-start for innerloop.optimise: when eligible (single storey, no
leaf_sharing/superpose/max_share/multi_use) and no caller-supplied x0, the
DP's exact shape-feasible ratio point is written onto the tree before NM
runs, off by default (shapecurve_warmstart=/--shapecurve-warmstart).
Caught and fixed a latent bug promoting the spike: realise() could leave
numpy.float64 in `division`, which yaml.safe_dump can't serialise — the
original spike never round-tripped through dom.dumps so this was never hit.
A/B on harbor-house-l0 (experiments/ab_shapecurve_warmstart.py, budget=2000,
5 seeds): mean total fails 16.6 (on) vs 19.6 (off), ~3.5x mean fitness
improvement; mean hard-fail count alone was a noise-level wash at this
sample size. Full writeup in DESIGN.md §37.4.
Deliberately deferred to new tracked beads (children of 2g7): DP-exact hard
pre-filter (wkh), multi-storey below-link support (koo), leaf_sharing/
co_type modelling (tym), true skew-quad polygon algebra (ekc) — 6xh stays
in_progress pending those.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-03 18:43:28 +01:00
|
|
|
|
multi_use: bool = False,
|
2026-08-03 21:10:28 +01:00
|
|
|
|
shapecurve_warmstart: bool = False,
|
|
|
|
|
|
shapecurve_prune: bool = False) -> tuple[Individual, int]:
|
2026-06-20 18:54:48 +01:00
|
|
|
|
# §12.3 shape-feasibility pre-filter (homemaker-py-9gp.1): if even the best
|
|
|
|
|
|
# achievable (proportion-aware) geometry of this topology already has at least
|
|
|
|
|
|
# as many shape fails as the incumbent's TOTAL fails — and exceeds the tunable
|
|
|
|
|
|
# threshold — it cannot beat the incumbent, so prune it for one feasibility
|
|
|
|
|
|
# eval instead of spending the full inner-loop budget. The best_n_fails guard
|
|
|
|
|
|
# makes the proxy safe: a topology whose shape-fail floor is still below the
|
|
|
|
|
|
# incumbent is never discarded. Pruned individuals are tagged and never admitted.
|
2026-07-19 20:35:18 +01:00
|
|
|
|
overrides = _overrides_for(leaf_sharing, superpose, max_share, conn_grade,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
collapse_insearch, multi_use)
|
2026-08-03 23:30:18 +01:00
|
|
|
|
# §37.4/§37.6 shape-curve DP warm-start (homemaker-py-6xh/koo, DESIGN.md
|
|
|
|
|
|
# §37.2/§37.4/§37.6): when eligible (any storey count since homemaker-py-koo
|
|
|
|
|
|
# — none of leaf_sharing/superpose/max_share/multi_use, which the DP still
|
|
|
|
|
|
# doesn't model) and no caller-supplied x0 (never override an
|
homemaker-py-6xh: wire shape-curve DP into driver.py as an NM warm-start
Promotes the validated shape-curve DP (experiments/shapecurve_spike.py,
2g7.4, DESIGN.md §37.2) from a reference-only spike into
src/homemaker_layout/shapecurve.py, and wires it into driver._evaluate as a
warm-start for innerloop.optimise: when eligible (single storey, no
leaf_sharing/superpose/max_share/multi_use) and no caller-supplied x0, the
DP's exact shape-feasible ratio point is written onto the tree before NM
runs, off by default (shapecurve_warmstart=/--shapecurve-warmstart).
Caught and fixed a latent bug promoting the spike: realise() could leave
numpy.float64 in `division`, which yaml.safe_dump can't serialise — the
original spike never round-tripped through dom.dumps so this was never hit.
A/B on harbor-house-l0 (experiments/ab_shapecurve_warmstart.py, budget=2000,
5 seeds): mean total fails 16.6 (on) vs 19.6 (off), ~3.5x mean fitness
improvement; mean hard-fail count alone was a noise-level wash at this
sample size. Full writeup in DESIGN.md §37.4.
Deliberately deferred to new tracked beads (children of 2g7): DP-exact hard
pre-filter (wkh), multi-storey below-link support (koo), leaf_sharing/
co_type modelling (tym), true skew-quad polygon algebra (ekc) — 6xh stays
in_progress pending those.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-03 18:43:28 +01:00
|
|
|
|
# explicit Lamarckian warm-start), solve for an exact shape-feasible ratio
|
|
|
|
|
|
# point and write it onto the tree in place. `x0=None` below then picks it up
|
|
|
|
|
|
# as the inner loop's start point. On infeasible or ineligible, `root` is left
|
|
|
|
|
|
# untouched — falls through to today's cold/proportion-aware start exactly.
|
2026-08-03 21:10:28 +01:00
|
|
|
|
dp_eligible = ((shapecurve_warmstart or shapecurve_prune)
|
|
|
|
|
|
and shapecurve.eligible(root, leaf_sharing, superpose, max_share, multi_use))
|
|
|
|
|
|
dp_feasible = None
|
|
|
|
|
|
if dp_eligible and shapecurve_warmstart and x0 is None:
|
|
|
|
|
|
dp_feasible, _ = shapecurve.solve(root, _fitness_for(
|
homemaker-py-6xh: wire shape-curve DP into driver.py as an NM warm-start
Promotes the validated shape-curve DP (experiments/shapecurve_spike.py,
2g7.4, DESIGN.md §37.2) from a reference-only spike into
src/homemaker_layout/shapecurve.py, and wires it into driver._evaluate as a
warm-start for innerloop.optimise: when eligible (single storey, no
leaf_sharing/superpose/max_share/multi_use) and no caller-supplied x0, the
DP's exact shape-feasible ratio point is written onto the tree before NM
runs, off by default (shapecurve_warmstart=/--shapecurve-warmstart).
Caught and fixed a latent bug promoting the spike: realise() could leave
numpy.float64 in `division`, which yaml.safe_dump can't serialise — the
original spike never round-tripped through dom.dumps so this was never hit.
A/B on harbor-house-l0 (experiments/ab_shapecurve_warmstart.py, budget=2000,
5 seeds): mean total fails 16.6 (on) vs 19.6 (off), ~3.5x mean fitness
improvement; mean hard-fail count alone was a noise-level wash at this
sample size. Full writeup in DESIGN.md §37.4.
Deliberately deferred to new tracked beads (children of 2g7): DP-exact hard
pre-filter (wkh), multi-storey below-link support (koo), leaf_sharing/
co_type modelling (tym), true skew-quad polygon algebra (ekc) — 6xh stays
in_progress pending those.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-03 18:43:28 +01:00
|
|
|
|
str(programme_dir), leaf_sharing, superpose, max_share,
|
|
|
|
|
|
conn_grade, collapse_insearch, multi_use))
|
2026-06-20 18:54:48 +01:00
|
|
|
|
if (feasibility_max_shape_fails is not None and best_n_fails is not None):
|
2026-08-03 21:10:28 +01:00
|
|
|
|
# §37.5 DP-exact hard prune (homemaker-py-wkh, DESIGN.md §37.5): the
|
|
|
|
|
|
# shape-curve DP gives an EXACT feasible/infeasible verdict (0/200
|
|
|
|
|
|
# measured false negatives on harbor-house-l0, §37.2) for the same
|
|
|
|
|
|
# size/width/proportion family predicted_shape_fails only heuristically
|
|
|
|
|
|
# counts at one (proportion-aware) layout. Composed conservatively —
|
|
|
|
|
|
# DP feasible VETOES the heuristic prune outright (a real feasible
|
|
|
|
|
|
# point exists, so the heuristic's high count was a false signal from
|
|
|
|
|
|
# an unlucky single layout, never the true floor); DP infeasible only
|
|
|
|
|
|
# licenses an exact prune when the incumbent already has zero total
|
|
|
|
|
|
# fails (best_n_fails<=0) — infeasible proves the shape-fail floor is
|
|
|
|
|
|
# >=1, which alone beats a zero-fail incumbent, but does not by itself
|
|
|
|
|
|
# establish the floor reaches an arbitrary best_n_fails>0, so that case
|
|
|
|
|
|
# still defers to the heuristic count (unchanged behaviour).
|
|
|
|
|
|
if dp_eligible and shapecurve_prune and dp_feasible is None:
|
|
|
|
|
|
dp_feasible = shapecurve.is_feasible(root, _fitness_for(
|
|
|
|
|
|
str(programme_dir), leaf_sharing, superpose, max_share,
|
|
|
|
|
|
conn_grade, collapse_insearch, multi_use))
|
|
|
|
|
|
if shapecurve_prune and dp_feasible is True:
|
|
|
|
|
|
prune = False
|
|
|
|
|
|
pred = 0
|
|
|
|
|
|
elif shapecurve_prune and dp_feasible is False and best_n_fails <= 0:
|
|
|
|
|
|
prune = True
|
|
|
|
|
|
pred = max(1, best_n_fails)
|
|
|
|
|
|
else:
|
|
|
|
|
|
pred = operators.predicted_shape_fails(
|
|
|
|
|
|
root, _reqs_for(str(programme_dir)),
|
|
|
|
|
|
_fitness_for(str(programme_dir), leaf_sharing, superpose, max_share,
|
|
|
|
|
|
conn_grade, collapse_insearch, multi_use))
|
|
|
|
|
|
prune = pred > feasibility_max_shape_fails and pred >= best_n_fails
|
|
|
|
|
|
if prune:
|
homemaker-py-2g7.3: hard/soft fail tiering behind --use-tiers flag
Splits the flat outer-search comparator (-n_fails, fitness) into a tiered
(-n_hard, -n_soft, fitness) so search budget stops being spent polishing
SOFT shape fails (crinkliness/proportion/size/width/edge-too-long/
staircase-volume) while HARD structural fails (missing space, wrong/
required level, level/circulation/vertical connectivity, adjacency,
stairs, covered-outside, storey limits, public access) remain unfixed.
fitness.classify_fail_tier/tier_counts classify every fail string emitted
across fitness.py and graph.py, raising on anything unrecognised so new
fail sites must declare a tier. Validated against all real fail strings in
the checked-in corpus plus every fail-emission call site read from source.
driver.Individual gains n_hard/n_soft (populated from innerloop.Result.
fail_lines); search(use_tiers=...) swaps the comparator when set (default
off, so existing runs are unaffected — inner-loop 0.5^n cliff untouched).
evolve.py exposes --use-tiers / HOMEMAKER_USE_TIERS.
experiments/tier_ab_2g7_3.py runs the acceptance A/B (harbor+maple, 3
seeds, 20k evals) in the background; results pending.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-02 16:00:39 +01:00
|
|
|
|
# predicted_shape_fails only counts the size/width/proportion/
|
|
|
|
|
|
# crinkliness SOFT family (operators._SHAPE_FAIL_SUFFIXES), so the
|
|
|
|
|
|
# proxy carries no HARD information — tier it all soft.
|
2026-06-20 18:54:48 +01:00
|
|
|
|
ind = Individual(root=root, fitness=0.0, n_fails=pred, ratios={},
|
|
|
|
|
|
lineage=f"pruned/{lineage}", grade=0.0,
|
homemaker-py-2g7.3: hard/soft fail tiering behind --use-tiers flag
Splits the flat outer-search comparator (-n_fails, fitness) into a tiered
(-n_hard, -n_soft, fitness) so search budget stops being spent polishing
SOFT shape fails (crinkliness/proportion/size/width/edge-too-long/
staircase-volume) while HARD structural fails (missing space, wrong/
required level, level/circulation/vertical connectivity, adjacency,
stairs, covered-outside, storey limits, public access) remain unfixed.
fitness.classify_fail_tier/tier_counts classify every fail string emitted
across fitness.py and graph.py, raising on anything unrecognised so new
fail sites must declare a tier. Validated against all real fail strings in
the checked-in corpus plus every fail-emission call site read from source.
driver.Individual gains n_hard/n_soft (populated from innerloop.Result.
fail_lines); search(use_tiers=...) swaps the comparator when set (default
off, so existing runs are unaffected — inner-loop 0.5^n cliff untouched).
evolve.py exposes --use-tiers / HOMEMAKER_USE_TIERS.
experiments/tier_ab_2g7_3.py runs the acceptance A/B (harbor+maple, 3
seeds, 20k evals) in the background; results pending.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-02 16:00:39 +01:00
|
|
|
|
sig=genome.signature(root), n_hard=0, n_soft=pred)
|
2026-06-20 18:54:48 +01:00
|
|
|
|
return ind, 1
|
2026-06-12 14:22:26 +01:00
|
|
|
|
r = innerloop.optimise(root, programme_dir, x0=x0, budget=budget,
|
2026-06-28 22:04:35 +01:00
|
|
|
|
urb_root=urb_root, conf_overrides=overrides, **inner_kw)
|
2026-06-18 22:33:29 +01:00
|
|
|
|
# §11.4: read the graded proximity scalar off the optimised tree. The inner
|
|
|
|
|
|
# loop left ``root`` at the optimum (Lamarckian write-back), so re-scoring a
|
|
|
|
|
|
# copy reproduces r.fitness/r.n_fails exactly and adds the grade. One extra
|
|
|
|
|
|
# native eval per child (~1/child_budget overhead); skipped unless requested.
|
|
|
|
|
|
grade = 0.0
|
|
|
|
|
|
if want_grade:
|
9o5: type superposition + per-eval collapse (multi-use leaves)
Interchangeable codes (similar size/width/proportion, compatible level/stack,
no adjacency edge) form equivalence classes derived from the programme. With
--superpose (default off), each fitness eval COLLAPSES every superposed leaf to
its best in-class usage via an optimal supply->demand assignment (brute force
<=C! within cap C=4, scipy Hungarian beyond), then scores the condensed types.
Because collapse re-types on the unmerged tree before all checks, counts /
adjacency / quality are unchanged downstream -- no Node field, no graph/operator
changes -- and default OFF is bit-identical.
- programme.py: derive_interchange_classes + interchangeable (S1-S4, locked
thresholds R_SIZE=1.5/R_WIDTH=1.3/R_PROP=1.5, CLASS_CAP=4)
- fitness.py: collapse_superposition, _best_assignment, _usage_quality;
superpose/superpose_class_cap conf knobs; collapse hooked into _evaluate_full
- driver.py/evolve.py: superpose flag plumbed beside leaf_sharing; --superpose
- tests/test_superposition.py: 17 tests (derivation, assignment, end-to-end)
Closes homemaker-py-9o5 (build); validation A/B is homemaker-py-xi7.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:08:46 +01:00
|
|
|
|
_, _, grade = _fitness_for(
|
2026-07-18 18:44:24 +01:00
|
|
|
|
str(programme_dir), leaf_sharing, superpose, max_share,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
conn_grade, collapse_insearch, multi_use).score_with_grade(
|
2026-06-18 22:33:29 +01:00
|
|
|
|
copy.deepcopy(root))
|
homemaker-py-2g7.3: hard/soft fail tiering behind --use-tiers flag
Splits the flat outer-search comparator (-n_fails, fitness) into a tiered
(-n_hard, -n_soft, fitness) so search budget stops being spent polishing
SOFT shape fails (crinkliness/proportion/size/width/edge-too-long/
staircase-volume) while HARD structural fails (missing space, wrong/
required level, level/circulation/vertical connectivity, adjacency,
stairs, covered-outside, storey limits, public access) remain unfixed.
fitness.classify_fail_tier/tier_counts classify every fail string emitted
across fitness.py and graph.py, raising on anything unrecognised so new
fail sites must declare a tier. Validated against all real fail strings in
the checked-in corpus plus every fail-emission call site read from source.
driver.Individual gains n_hard/n_soft (populated from innerloop.Result.
fail_lines); search(use_tiers=...) swaps the comparator when set (default
off, so existing runs are unaffected — inner-loop 0.5^n cliff untouched).
evolve.py exposes --use-tiers / HOMEMAKER_USE_TIERS.
experiments/tier_ab_2g7_3.py runs the acceptance A/B (harbor+maple, 3
seeds, 20k evals) in the background; results pending.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-02 16:00:39 +01:00
|
|
|
|
n_hard, n_soft = fitness.tier_counts(r.fail_lines)
|
2026-06-12 14:22:26 +01:00
|
|
|
|
ind = Individual(root=root, fitness=r.fitness, n_fails=r.n_fails,
|
2026-06-18 22:33:29 +01:00
|
|
|
|
ratios=innerloop.ratio_map(root), lineage=lineage,
|
homemaker-py-2g7.3: hard/soft fail tiering behind --use-tiers flag
Splits the flat outer-search comparator (-n_fails, fitness) into a tiered
(-n_hard, -n_soft, fitness) so search budget stops being spent polishing
SOFT shape fails (crinkliness/proportion/size/width/edge-too-long/
staircase-volume) while HARD structural fails (missing space, wrong/
required level, level/circulation/vertical connectivity, adjacency,
stairs, covered-outside, storey limits, public access) remain unfixed.
fitness.classify_fail_tier/tier_counts classify every fail string emitted
across fitness.py and graph.py, raising on anything unrecognised so new
fail sites must declare a tier. Validated against all real fail strings in
the checked-in corpus plus every fail-emission call site read from source.
driver.Individual gains n_hard/n_soft (populated from innerloop.Result.
fail_lines); search(use_tiers=...) swaps the comparator when set (default
off, so existing runs are unaffected — inner-loop 0.5^n cliff untouched).
evolve.py exposes --use-tiers / HOMEMAKER_USE_TIERS.
experiments/tier_ab_2g7_3.py runs the acceptance A/B (harbor+maple, 3
seeds, 20k evals) in the background; results pending.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-02 16:00:39 +01:00
|
|
|
|
grade=grade, sig=genome.signature(root),
|
|
|
|
|
|
n_hard=n_hard, n_soft=n_soft)
|
2026-06-12 14:22:26 +01:00
|
|
|
|
return ind, r.n_evals
|
|
|
|
|
|
|
|
|
|
|
|
|
2026-06-14 09:20:03 +01:00
|
|
|
|
def _tournament(pop: list[Individual], rng: np.random.Generator, key_fn, k: int = 2) -> Individual:
|
2026-06-12 14:22:26 +01:00
|
|
|
|
picks = rng.integers(len(pop), size=k)
|
2026-06-14 09:20:03 +01:00
|
|
|
|
return max((pop[int(i)] for i in picks), key=key_fn)
|
2026-06-12 14:22:26 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def search(
|
|
|
|
|
|
seed_root: dom.Node,
|
|
|
|
|
|
programme_dir: str | Path,
|
|
|
|
|
|
budget: int = 2000,
|
|
|
|
|
|
pop_size: int = 8,
|
|
|
|
|
|
child_budget: int = 80,
|
|
|
|
|
|
seed_budget: int = 200,
|
2026-06-13 23:29:12 +01:00
|
|
|
|
bootstrap: bool | None = None,
|
|
|
|
|
|
bootstrap_n_leaves: int | None = None,
|
2026-06-12 14:22:26 +01:00
|
|
|
|
p_crossover: float = 0.2,
|
|
|
|
|
|
seed: int = 0,
|
|
|
|
|
|
types: list[str] | None = None,
|
|
|
|
|
|
inner_kw: dict | None = None,
|
|
|
|
|
|
urb_root=None,
|
|
|
|
|
|
log=None,
|
2026-06-14 06:55:58 +01:00
|
|
|
|
n_workers: int = 1,
|
2026-06-14 09:20:03 +01:00
|
|
|
|
use_lex: bool = True,
|
homemaker-py-2g7.3: hard/soft fail tiering behind --use-tiers flag
Splits the flat outer-search comparator (-n_fails, fitness) into a tiered
(-n_hard, -n_soft, fitness) so search budget stops being spent polishing
SOFT shape fails (crinkliness/proportion/size/width/edge-too-long/
staircase-volume) while HARD structural fails (missing space, wrong/
required level, level/circulation/vertical connectivity, adjacency,
stairs, covered-outside, storey limits, public access) remain unfixed.
fitness.classify_fail_tier/tier_counts classify every fail string emitted
across fitness.py and graph.py, raising on anything unrecognised so new
fail sites must declare a tier. Validated against all real fail strings in
the checked-in corpus plus every fail-emission call site read from source.
driver.Individual gains n_hard/n_soft (populated from innerloop.Result.
fail_lines); search(use_tiers=...) swaps the comparator when set (default
off, so existing runs are unaffected — inner-loop 0.5^n cliff untouched).
evolve.py exposes --use-tiers / HOMEMAKER_USE_TIERS.
experiments/tier_ab_2g7_3.py runs the acceptance A/B (harbor+maple, 3
seeds, 20k evals) in the background; results pending.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-02 16:00:39 +01:00
|
|
|
|
use_tiers: bool = False,
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
rank_bonus_fn=None,
|
|
|
|
|
|
rank_bonus_weight: float = 1.0,
|
|
|
|
|
|
seed_factory=None,
|
|
|
|
|
|
base_p: float = 1.0,
|
2026-06-29 06:20:29 +01:00
|
|
|
|
child_probe=None,
|
2026-06-18 22:33:29 +01:00
|
|
|
|
use_grade: bool = False,
|
2026-07-18 18:44:24 +01:00
|
|
|
|
conn_grade: bool = False,
|
2026-06-29 22:56:23 +01:00
|
|
|
|
tournament_k: int = 2,
|
2026-06-18 23:42:39 +01:00
|
|
|
|
niche_by_signature: bool = False,
|
|
|
|
|
|
restart_patience: int | None = None,
|
|
|
|
|
|
restart_elite: int = 1,
|
2026-06-19 09:23:12 +01:00
|
|
|
|
seed_adjacency_aware: bool = True,
|
Phase 7 §12.2: proportion-aware constructive seeding + storey_minimum fix (leu.2, cq1)
Size each constructive-seed cut from leaf TARGET areas (division=[f,f] gives
left area-fraction f) and pick each cut's rotation for child squareness — both
derived from target dims, topology/type assignment untouched. Area-only
regressed (slivers); rotation choice is what makes it pay.
End-to-end (20000 evals, 3 seeds, staged): harbor 85.3->74.0 (-13%, best 69),
maple-court 151.7->136.0 (-10%, best 126). PROP=0 reproduces the §11.7/§12.1
baselines exactly. programme-house regresses at fixed budget (deeper local
optimum walls off the undivide restructuring path) but a budget sweep shows
it's convergence speed, not a worse asymptote (PROP=1 reaches 1 fail at 150k).
Default-on (seed_proportion_aware=True, env PROP=1).
cq1: n_storeys now honours storey_minimum, not just level: keys — programme-house
(storey_minimum:2, all rooms level:0) was seeded one storey short and fell
through to plain search. New programme.storey_minimum()/n_storeys_for();
driver.search passes min_storeys to the seeder; search_staged routes on the max.
No-op for harbor/maple; programme-house single-stage 8.0->5.0.
New maple-court best (126) saved as generated.dom. 204 tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 14:04:42 +01:00
|
|
|
|
seed_proportion_aware: bool = True,
|
2026-06-20 18:54:48 +01:00
|
|
|
|
enable_reassociate: bool = False,
|
2026-07-22 17:43:36 +01:00
|
|
|
|
enable_shape_repair: bool = False,
|
2026-07-24 19:48:13 +01:00
|
|
|
|
enable_bridge_circulation: bool = False,
|
2026-07-26 09:31:42 +01:00
|
|
|
|
enable_ruin_recreate: bool = False,
|
2026-06-20 18:54:48 +01:00
|
|
|
|
feasibility_filter: bool = False,
|
|
|
|
|
|
feasibility_max_shape_fails: int | None = None,
|
2026-06-21 21:10:18 +01:00
|
|
|
|
circ_divisor: int = 3,
|
2026-06-27 21:15:50 +01:00
|
|
|
|
leaf_sharing: bool = True,
|
|
|
|
|
|
leaf_share_factor: int = 3,
|
9o5: type superposition + per-eval collapse (multi-use leaves)
Interchangeable codes (similar size/width/proportion, compatible level/stack,
no adjacency edge) form equivalence classes derived from the programme. With
--superpose (default off), each fitness eval COLLAPSES every superposed leaf to
its best in-class usage via an optimal supply->demand assignment (brute force
<=C! within cap C=4, scipy Hungarian beyond), then scores the condensed types.
Because collapse re-types on the unmerged tree before all checks, counts /
adjacency / quality are unchanged downstream -- no Node field, no graph/operator
changes -- and default OFF is bit-identical.
- programme.py: derive_interchange_classes + interchangeable (S1-S4, locked
thresholds R_SIZE=1.5/R_WIDTH=1.3/R_PROP=1.5, CLASS_CAP=4)
- fitness.py: collapse_superposition, _best_assignment, _usage_quality;
superpose/superpose_class_cap conf knobs; collapse hooked into _evaluate_full
- driver.py/evolve.py: superpose flag plumbed beside leaf_sharing; --superpose
- tests/test_superposition.py: 17 tests (derivation, assignment, end-to-end)
Closes homemaker-py-9o5 (build); validation A/B is homemaker-py-xi7.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:08:46 +01:00
|
|
|
|
superpose: bool = False,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
multi_use: bool = False,
|
2026-06-27 21:15:50 +01:00
|
|
|
|
depth_balanced: bool = True,
|
2026-06-28 07:29:42 +01:00
|
|
|
|
interior_outside: bool = True,
|
2026-06-28 07:20:20 +01:00
|
|
|
|
outside_divisor: int = 3,
|
2026-07-28 00:00:51 +01:00
|
|
|
|
construction_beam_width: int = 1,
|
2026-07-16 08:38:08 +01:00
|
|
|
|
max_share: int | None = None,
|
|
|
|
|
|
seed_pop: list[dom.Node] | None = None,
|
2026-07-24 09:55:34 +01:00
|
|
|
|
collapse_insearch: bool = True,
|
homemaker-py-6xh: wire shape-curve DP into driver.py as an NM warm-start
Promotes the validated shape-curve DP (experiments/shapecurve_spike.py,
2g7.4, DESIGN.md §37.2) from a reference-only spike into
src/homemaker_layout/shapecurve.py, and wires it into driver._evaluate as a
warm-start for innerloop.optimise: when eligible (single storey, no
leaf_sharing/superpose/max_share/multi_use) and no caller-supplied x0, the
DP's exact shape-feasible ratio point is written onto the tree before NM
runs, off by default (shapecurve_warmstart=/--shapecurve-warmstart).
Caught and fixed a latent bug promoting the spike: realise() could leave
numpy.float64 in `division`, which yaml.safe_dump can't serialise — the
original spike never round-tripped through dom.dumps so this was never hit.
A/B on harbor-house-l0 (experiments/ab_shapecurve_warmstart.py, budget=2000,
5 seeds): mean total fails 16.6 (on) vs 19.6 (off), ~3.5x mean fitness
improvement; mean hard-fail count alone was a noise-level wash at this
sample size. Full writeup in DESIGN.md §37.4.
Deliberately deferred to new tracked beads (children of 2g7): DP-exact hard
pre-filter (wkh), multi-storey below-link support (koo), leaf_sharing/
co_type modelling (tym), true skew-quad polygon algebra (ekc) — 6xh stays
in_progress pending those.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-03 18:43:28 +01:00
|
|
|
|
shapecurve_warmstart: bool = False,
|
2026-08-03 21:10:28 +01:00
|
|
|
|
shapecurve_prune: bool = False,
|
2026-08-04 09:19:36 +01:00
|
|
|
|
assign_solver: str = "greedy",
|
|
|
|
|
|
enable_reassign: bool = False,
|
§39.10: preserving constructed connectivity is NULL — and it reframes §39.9
§39.9 named the upstream fix: keep circulation connected DURING the resize
rather than rebuilding it after. Built and measured. It does not help, and the
reason matters more than the lever.
Both halves of the re-cut do damage, in different proportions per programme.
Freezing rotations and letting only ratios move (% levels connected, 12 seeds):
harbor 100 -> 71 -> 50, health-centre 100 -> 8 -> 8, maple 100 -> 92 -> 67. So
health-centre is destroyed entirely by the ratio and maple mostly by the
rotation; a fix must be able to give back either.
operators._size_divisions_preserving_circulation snapshots every cut, resizes,
then reverts the cuts on the tree path between each circulation pair the resize
broke -- programme fully intact, no retyping, only geometry given back. It works
on connectivity (harbor 50->92%, maple 67->97%, health-centre 8->17%) and costs
area accuracy: constructed-seed fails harbor 96.6->141.5, maple 141.8->175.8,
size fails roughly double. (A greedy single-cut revert barely moved -- it stalls
where no ONE revert helps though two would. Targeting the broken pairs is what
made connectivity work.)
The obvious defence -- raw constructed seeds understate it, the resize is only a
warm start, the inner loop should recover -- was TESTED AND FAILS. Full search,
harbor-house, 12000 evals, seed 1:
OFF 43 fails, 9 hard, 3 connectivity
ON 65 fails, 26 hard, 4 connectivity
Worse on every axis, including connectivity itself.
REFRAMING: §39.9's fact stands (the resize destroys 41 of 49 circulation edges)
but is NOT ACTIONABLE, because construction-time connectivity does not determine
final connectivity. The search discards and rebuilds the seeder's circulation
either way, and constraining the seed only spends area quality the search cannot
recover. Together with §39.8 (not an incentive problem) that retires the framing
this thread inherited from §38: connectivity is neither a construction problem
nor an incentive one.
Both flags (repair_circulation, preserve_circulation) stay default off with the
numbers recorded, plus byte-identical-default tests. Do not revisit either
without a new formulation -- the standing this document gives bubble.py.
356 passed (+1 new), same 7 pre-existing fixture failures, lint unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-08-26 15:11:46 +00:00
|
|
|
|
preserve_circulation: bool = False,
|
Checkpoint long searches; the cold-start runs were lost to a reclaimed box
All four 500k runs died about 10 minutes in when the container was
reclaimed. No SIGTERM fired, so no .dom was written and 0 of 12 runs
completed. My plan committed results per finished run, which protected
nothing because no run reached its commit point. The bad assumption was
reading "reclaimed after inactivity" as CPU inactivity; it is conversation
inactivity, and background compute does not hold the box open.
Progress reached before the loss (from the tracked logs): harbor 24,960
evals / 40 fails, maple 14,880 / 79, health-centre 25,920 / 33,
programme-house 138,800 / 2.
The underlying gap is not environmental: a search's only output lands at
the very end or on SIGTERM, so ANY abrupt loss -- reclaimed container, OOM,
power cut -- takes the whole run with it. On a 3M-eval search that is 2.4
days of compute with no recoverable artefact.
- driver.search gains checkpoint=/checkpoint_every=: the current best is
handed to a callback at most every N evals. Rate-limited by evals, not
improvements, which come in bursts early. A failing checkpoint is logged
and swallowed -- losing a checkpoint is bad, losing the search because a
checkpoint failed is worse.
- homemaker-evolve --checkpoint-every N writes <out>.dom.checkpoint via
mkstemp + os.replace, so a crash can never catch it half-written. It is
deliberately NOT the output path: a checkpoint is a leaf-sharing run's
internal best, dishonest under the canonical scorer until the finish
stage unfolds it (homemaker-py-3l6), and must not be mistaken for the
finished article.
- Verified the written checkpoint re-loads as a valid .dom.
Default off, so behaviour is unchanged without the flag.
Lint at parity (46); tests 372 passed (3 new), same 2 pre-existing failures.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-08-29 05:43:58 +00:00
|
|
|
|
checkpoint=None,
|
|
|
|
|
|
checkpoint_every: int = 0,
|
2026-06-12 14:22:26 +01:00
|
|
|
|
) -> SearchResult:
|
|
|
|
|
|
"""Run the memetic loop from ``seed_root`` until ``budget`` oracle
|
|
|
|
|
|
evaluations are consumed. Returns the best individual found; its ``root``
|
2026-06-13 23:29:12 +01:00
|
|
|
|
carries the optimised geometry and dumps to a valid ``.dom``.
|
|
|
|
|
|
|
|
|
|
|
|
``bootstrap=None`` (default) auto-detects: if ``seed_root`` is an
|
|
|
|
|
|
undivided bare plot, generates a diverse initial population of ``pop_size``
|
|
|
|
|
|
random topologies (each with approximately ``bootstrap_n_leaves`` leaves)
|
|
|
|
|
|
before the memetic loop starts. Pass ``bootstrap=False`` to force the
|
|
|
|
|
|
legacy single-seed path (appropriate for warm starts from existing designs).
|
2026-06-14 06:55:58 +01:00
|
|
|
|
|
|
|
|
|
|
``n_workers=1`` (default) runs serially; ``n_workers > 1`` evaluates
|
|
|
|
|
|
children in parallel using ``ProcessPoolExecutor``. The bootstrap batch
|
|
|
|
|
|
is fully parallel; the main loop generates ``n_workers`` children per
|
|
|
|
|
|
iteration from the current population snapshot and evaluates them in
|
|
|
|
|
|
parallel. Results are admitted in completion order (fastest first), so
|
|
|
|
|
|
later children in each batch see an already-updated population.
|
2026-06-18 23:42:39 +01:00
|
|
|
|
|
|
|
|
|
|
``niche_by_signature`` (DESIGN.md §11.5, default ``False`` — REJECTED, kept
|
|
|
|
|
|
for reuse) replaces the legacy fitness-scalar duplicate guard with structural
|
|
|
|
|
|
niching: the population holds at most one individual per
|
|
|
|
|
|
:func:`genome.signature` (topology), keeping the better of any collision, so
|
|
|
|
|
|
distinct topologies whose fitness scalars coincide (common in the high-fail
|
|
|
|
|
|
``0.5^n`` regime) are no longer discarded. ``restart_patience`` (default
|
|
|
|
|
|
``None`` = off) triggers a soft restart when the best has not improved for
|
|
|
|
|
|
that many evals: the top ``restart_elite`` incumbents are kept and the rest of
|
|
|
|
|
|
the population is refilled with fresh constructive/random seeds, the
|
|
|
|
|
|
soft-restart analog of urb-evolve's upfront random-population diversity.
|
|
|
|
|
|
|
|
|
|
|
|
Both default off: §11.5 measured that they raise structural diversity as
|
|
|
|
|
|
designed (final-population distinct topologies ~5/16 → 16/16) but do **not**
|
|
|
|
|
|
lower the fail count — a tie within seed noise on blank-slate programme-house
|
|
|
|
|
|
(mean 12.3 → 12.7) and harbor (95 → 94), with restarts strictly worse. The
|
|
|
|
|
|
high-fail plateau is therefore not a population-diversity deficit; the lever
|
|
|
|
|
|
is the canonical encoding (``homemaker-py-9gp``) and richer operators.
|
2026-07-16 08:38:08 +01:00
|
|
|
|
|
|
|
|
|
|
``max_share`` (homemaker-py-kpu) overrides the evaluator's ``leaf_share_max``
|
|
|
|
|
|
grain cap for this phase; ``None`` uses the config default. ``seed_pop`` (also
|
|
|
|
|
|
kpu) supplies an explicit initial population of decoded roots — evaluated
|
|
|
|
|
|
under this phase's evaluator instead of bootstrapping or single-seeding — so a
|
|
|
|
|
|
grain-anneal ramp can hand a whole population from one phase to the next.
|
2026-07-19 20:35:18 +01:00
|
|
|
|
|
2026-07-24 09:55:34 +01:00
|
|
|
|
``collapse_insearch`` (homemaker-py-qpk, default on) runs the 94g global
|
|
|
|
|
|
cell<->room collapse inside every fitness eval instead of once at finish
|
|
|
|
|
|
time, so search optimises the collapsed objective directly. A/B-validated
|
|
|
|
|
|
positive on harbor-house (3/3, mean fails 80.3->72.0) and, after the
|
|
|
|
|
|
homemaker-py-1ph larger-N seed sweep, on programme-house too (11/17
|
|
|
|
|
|
non-tied wins, mean fails 7.95->7.10) — DESIGN.md §17/§20.
|
2026-07-22 17:43:36 +01:00
|
|
|
|
|
|
|
|
|
|
``enable_shape_repair`` (homemaker-py-161, EXPERIMENTAL, default off) threads
|
|
|
|
|
|
a ``fitness.Fitness`` instance into ``operators.mutate`` so the ``shape_rotate``
|
|
|
|
|
|
and ``deslim`` repair operators (homemaker-py-7fm) can fire during the GA
|
|
|
|
|
|
instead of no-opping. 7fm's finish-time hill-climb found these operators never
|
|
|
|
|
|
improve an already-co-evolved layout (every candidate move traded one fail for
|
|
|
|
|
|
another); this flag tests whether in-search selection pressure lets a
|
|
|
|
|
|
locally-worse move survive to be completed by a later step or crossover — a
|
|
|
|
|
|
different regime. Mirrors ``enable_reassociate``'s clean-toggle A/B pattern
|
|
|
|
|
|
(§12.3, 9gp.2): default off reproduces prior runs byte-for-byte.
|
2026-07-24 19:48:13 +01:00
|
|
|
|
|
|
|
|
|
|
``enable_bridge_circulation`` (homemaker-py-8sh, EXPERIMENTAL, default off)
|
|
|
|
|
|
un-mutes the ``bridge_circulation`` repair operator (homemaker-py-qi6
|
|
|
|
|
|
mechanism (a)): it retypes leaves on the cheapest path between two
|
|
|
|
|
|
disconnected circulation components to clear a ``level N not connected``
|
|
|
|
|
|
fail directly, instead of relying on the outer search to discover
|
|
|
|
|
|
connectivity via a comparator-key gradient (qi6 mechanism (b)/(c), measured
|
|
|
|
|
|
NEGATIVE — DESIGN.md §18). Gated the same way as ``reassociate`` (zero
|
|
|
|
|
|
mutation weight unless enabled) rather than ``shape_repair``'s style,
|
|
|
|
|
|
because it needs no ``fitness.Fitness`` instance — only the tree's own
|
|
|
|
|
|
adjacency graph — so it is otherwise unconditionally live once landed in
|
|
|
|
|
|
``operators.MUTATIONS``.
|
2026-07-26 09:31:42 +01:00
|
|
|
|
|
|
|
|
|
|
``enable_ruin_recreate`` (homemaker-py-f1d, EXPERIMENTAL, default off) un-
|
|
|
|
|
|
mutes ``operators.mutate_ruin_recreate``: a large-neighbourhood-search move
|
|
|
|
|
|
that un-divides one wing of a storey and rebuilds it with the same
|
|
|
|
|
|
adjacency-aware constructor the seeders use (``operators.
|
|
|
|
|
|
_assign_adjacency_aware``, seeded from the surviving circulation bordering
|
|
|
|
|
|
the wing), instead of relying only on the small local mutation operators to
|
|
|
|
|
|
discover an improving rearrangement. Gated like ``reassociate`` (zero
|
|
|
|
|
|
mutation weight unless enabled) — it needs only ``reqs``, no
|
|
|
|
|
|
``fitness.Fitness`` instance.
|
2026-07-28 00:00:51 +01:00
|
|
|
|
|
|
|
|
|
|
``construction_beam_width`` (homemaker-py-c94, EXPERIMENTAL, default 1)
|
|
|
|
|
|
forwarded to ``operators.constructive_topology``/``lift_base_to_storeys``'s
|
|
|
|
|
|
same-named parameter, in turn ``_assign_adjacency_aware``'s ``beam_width``:
|
|
|
|
|
|
a width-K beam search over which leaf a room lands on during construction,
|
|
|
|
|
|
instead of one irrevocable greedy pass. ``1`` (default) reproduces the
|
|
|
|
|
|
prior greedy seeding exactly.
|
2026-08-04 09:19:36 +01:00
|
|
|
|
|
|
|
|
|
|
``assign_solver`` (homemaker-py-2g7.5, EXPERIMENTAL, default "greedy")
|
|
|
|
|
|
forwarded to ``operators.constructive_topology``/``lift_base_to_storeys``'s
|
|
|
|
|
|
same-named parameter: ``"cpsat"`` replaces the greedy/beam room-code
|
|
|
|
|
|
placement with an exact OR-Tools CP-SAT solve (DESIGN.md §37.7),
|
|
|
|
|
|
falling through to the greedy/beam path on any solver failure.
|
|
|
|
|
|
``"greedy"`` (default) reproduces prior seeding exactly.
|
|
|
|
|
|
|
|
|
|
|
|
``enable_reassign`` (homemaker-py-2g7.5, EXPERIMENTAL, default off)
|
|
|
|
|
|
un-mutes ``operators.mutate_reassign``: the CP-SAT analogue of
|
|
|
|
|
|
``enable_ruin_recreate`` — re-solves one wing's room-code labelling
|
|
|
|
|
|
exactly instead of un-dividing and regrowing it. Gated the same way as
|
|
|
|
|
|
``ruin_recreate`` (zero mutation weight unless enabled).
|
2026-06-13 23:29:12 +01:00
|
|
|
|
"""
|
2026-06-12 14:22:26 +01:00
|
|
|
|
from .oracle import DEFAULT_URB_ROOT
|
|
|
|
|
|
|
|
|
|
|
|
urb_root = urb_root or DEFAULT_URB_ROOT
|
|
|
|
|
|
rng = np.random.default_rng(seed)
|
|
|
|
|
|
inner_kw = dict(_CHILD_INNER_KW, **(inner_kw or {}))
|
2026-06-20 18:54:48 +01:00
|
|
|
|
# §12.3 M3 reassociate (homemaker-py-9gp.2) is default-OFF: force its weight to
|
|
|
|
|
|
# 0 unless enabled, so the leu.2 baseline reproduces byte-for-byte (the operator
|
|
|
|
|
|
# never fires) and the A/B is a clean single-variable toggle.
|
|
|
|
|
|
mutation_weights = dict(_MUTATION_WEIGHTS)
|
|
|
|
|
|
if not enable_reassociate:
|
|
|
|
|
|
mutation_weights["reassociate"] = 0.0
|
2026-07-24 19:48:13 +01:00
|
|
|
|
if not enable_bridge_circulation:
|
|
|
|
|
|
mutation_weights["bridge_circulation"] = 0.0
|
2026-07-26 09:31:42 +01:00
|
|
|
|
if not enable_ruin_recreate:
|
|
|
|
|
|
mutation_weights["ruin_recreate"] = 0.0
|
2026-08-04 09:19:36 +01:00
|
|
|
|
if not enable_reassign:
|
|
|
|
|
|
mutation_weights["reassign"] = 0.0
|
2026-07-22 17:43:36 +01:00
|
|
|
|
# homemaker-py-161: shape_rotate/deslim are gated by operators.mutate itself
|
|
|
|
|
|
# (fit_ops go to zero probability when fit=None) — only build the Fitness
|
|
|
|
|
|
# instance, and thus only let them fire, when explicitly enabled.
|
|
|
|
|
|
shape_repair_fit = (
|
|
|
|
|
|
_fitness_for(str(programme_dir), leaf_sharing, superpose, max_share,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
conn_grade, collapse_insearch, multi_use)
|
2026-07-22 17:43:36 +01:00
|
|
|
|
if enable_shape_repair else None)
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
# Optional ranking bonus (DESIGN.md §11.3 Stage 1): bias selection toward
|
|
|
|
|
|
# individuals with high substrate-readiness via a multiplicative factor
|
|
|
|
|
|
# (1 + W·bonus) on fitness. The reported fitness/history stay the TRUE
|
|
|
|
|
|
# fitness; only the comparison key changes. rank_bonus_fn=None (default) ⇒
|
|
|
|
|
|
# the key is unchanged, so normal/Stage-2/programme-house runs are unaffected.
|
|
|
|
|
|
def _rank_fitness(ind: Individual) -> float:
|
|
|
|
|
|
if rank_bonus_fn is None:
|
|
|
|
|
|
return ind.fitness
|
|
|
|
|
|
return ind.fitness * (1.0 + rank_bonus_weight * rank_bonus_fn(ind.root))
|
|
|
|
|
|
|
2026-06-18 22:33:29 +01:00
|
|
|
|
# §11.4 graded objective (EXPERIMENT, default off — REJECTED, see DESIGN.md
|
|
|
|
|
|
# §11.4): a continuous proximity bonus (ind.grade) inserted as a secondary key
|
|
|
|
|
|
# BENEATH fail-count and ABOVE fitness, ordering neighbours by how close their
|
|
|
|
|
|
# failing constraints are to satisfaction. Hypothesis was that fitness is
|
|
|
|
|
|
# ~flat (0.5^n) in the high-fail regime; this was FALSIFIED — within a fixed
|
|
|
|
|
|
# fail-tier 0.5^n is constant so fitness still spans ~6 orders of magnitude,
|
|
|
|
|
|
# and grade above it merely displaces that working signal (no plateau escape).
|
|
|
|
|
|
# Kept default-off for reproducibility. Strictly beneath -n_fails ⇒ the
|
|
|
|
|
|
# missing-space hierarchy (§6) is preserved and the inner-loop cliff (§5.4)
|
|
|
|
|
|
# is untouched.
|
2026-07-18 18:44:24 +01:00
|
|
|
|
# homemaker-py-qi6 §18: the connectivity signal rides the same grade channel,
|
|
|
|
|
|
# so enabling it enables the grade secondary key.
|
|
|
|
|
|
use_grade = use_grade or conn_grade
|
homemaker-py-2g7.3: hard/soft fail tiering behind --use-tiers flag
Splits the flat outer-search comparator (-n_fails, fitness) into a tiered
(-n_hard, -n_soft, fitness) so search budget stops being spent polishing
SOFT shape fails (crinkliness/proportion/size/width/edge-too-long/
staircase-volume) while HARD structural fails (missing space, wrong/
required level, level/circulation/vertical connectivity, adjacency,
stairs, covered-outside, storey limits, public access) remain unfixed.
fitness.classify_fail_tier/tier_counts classify every fail string emitted
across fitness.py and graph.py, raising on anything unrecognised so new
fail sites must declare a tier. Validated against all real fail strings in
the checked-in corpus plus every fail-emission call site read from source.
driver.Individual gains n_hard/n_soft (populated from innerloop.Result.
fail_lines); search(use_tiers=...) swaps the comparator when set (default
off, so existing runs are unaffected — inner-loop 0.5^n cliff untouched).
evolve.py exposes --use-tiers / HOMEMAKER_USE_TIERS.
experiments/tier_ab_2g7_3.py runs the acceptance A/B (harbor+maple, 3
seeds, 20k evals) in the background; results pending.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-02 16:00:39 +01:00
|
|
|
|
# homemaker-py-2g7.3 (DESIGN.md §37): tiered comparator, EXPERIMENT default off.
|
|
|
|
|
|
# Splits the flat -n_fails key into (-n_hard, -n_soft) so search budget stops
|
|
|
|
|
|
# being spent polishing SOFT shape fails (crinkliness/proportion/size/width/
|
|
|
|
|
|
# edge-too-long/staircase-volume) while HARD structural fails (missing space,
|
|
|
|
|
|
# wrong/required level, level/circulation/vertical connectivity, adjacency,
|
|
|
|
|
|
# stairs, covered-outside, storey limits, public access — fitness.py's
|
|
|
|
|
|
# classify_fail_tier) remain unfixed. Does not change the scalar fitness or
|
|
|
|
|
|
# total fail count, so the inner-loop 0.5^n cliff protection (§5.4) and the
|
|
|
|
|
|
# §4.9 outer A/B baseline are untouched when this flag is off.
|
|
|
|
|
|
if use_lex and use_tiers and use_grade:
|
|
|
|
|
|
_key = lambda ind: (-ind.n_hard, -ind.n_soft, ind.grade, _rank_fitness(ind))
|
|
|
|
|
|
elif use_lex and use_tiers:
|
|
|
|
|
|
_key = lambda ind: (-ind.n_hard, -ind.n_soft, _rank_fitness(ind))
|
|
|
|
|
|
elif use_lex and use_grade:
|
2026-06-18 22:33:29 +01:00
|
|
|
|
_key = lambda ind: (-ind.n_fails, ind.grade, _rank_fitness(ind))
|
|
|
|
|
|
elif use_lex:
|
|
|
|
|
|
_key = lambda ind: (-ind.n_fails, _rank_fitness(ind))
|
|
|
|
|
|
else:
|
|
|
|
|
|
_key = lambda ind: _rank_fitness(ind)
|
2026-06-13 23:29:12 +01:00
|
|
|
|
# Always load reqs so bootstrap_n_leaves can be auto-derived from programme.
|
2026-06-14 07:50:39 +01:00
|
|
|
|
reqs = programme.load_programme_dir(programme_dir)
|
Phase 7 §12.2: proportion-aware constructive seeding + storey_minimum fix (leu.2, cq1)
Size each constructive-seed cut from leaf TARGET areas (division=[f,f] gives
left area-fraction f) and pick each cut's rotation for child squareness — both
derived from target dims, topology/type assignment untouched. Area-only
regressed (slivers); rotation choice is what makes it pay.
End-to-end (20000 evals, 3 seeds, staged): harbor 85.3->74.0 (-13%, best 69),
maple-court 151.7->136.0 (-10%, best 126). PROP=0 reproduces the §11.7/§12.1
baselines exactly. programme-house regresses at fixed budget (deeper local
optimum walls off the undivide restructuring path) but a budget sweep shows
it's convergence speed, not a worse asymptote (PROP=1 reaches 1 fail at 150k).
Default-on (seed_proportion_aware=True, env PROP=1).
cq1: n_storeys now honours storey_minimum, not just level: keys — programme-house
(storey_minimum:2, all rooms level:0) was seeded one storey short and fell
through to plain search. New programme.storey_minimum()/n_storeys_for();
driver.search passes min_storeys to the seeder; search_staged routes on the max.
No-op for harbor/maple; programme-house single-stage 8.0->5.0.
New maple-court best (126) saved as generated.dom. 204 tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 14:04:42 +01:00
|
|
|
|
# Constructive seed must honour storey_minimum, not just level: keys (§12.2).
|
|
|
|
|
|
min_storeys = programme.storey_minimum(programme_dir)
|
2026-06-12 14:22:26 +01:00
|
|
|
|
if types is None:
|
2026-06-12 19:01:53 +01:00
|
|
|
|
# Urb's generic types are canonically UPPERCASE (get_space_types:
|
|
|
|
|
|
# qw/C O S/; the corpus is 100% uppercase). Predicates match
|
|
|
|
|
|
# case-insensitively but Dom->Ratios keys raw strings — mixing cases
|
|
|
|
|
|
# fragments the class buckets, so never emit lowercase generics.
|
|
|
|
|
|
types = sorted(reqs) + ["C", "O"]
|
2026-06-12 14:22:26 +01:00
|
|
|
|
|
2026-06-13 23:29:12 +01:00
|
|
|
|
do_bootstrap = (not seed_root.divided) if bootstrap is None else bootstrap
|
|
|
|
|
|
|
Checkpoint long searches; the cold-start runs were lost to a reclaimed box
All four 500k runs died about 10 minutes in when the container was
reclaimed. No SIGTERM fired, so no .dom was written and 0 of 12 runs
completed. My plan committed results per finished run, which protected
nothing because no run reached its commit point. The bad assumption was
reading "reclaimed after inactivity" as CPU inactivity; it is conversation
inactivity, and background compute does not hold the box open.
Progress reached before the loss (from the tracked logs): harbor 24,960
evals / 40 fails, maple 14,880 / 79, health-centre 25,920 / 33,
programme-house 138,800 / 2.
The underlying gap is not environmental: a search's only output lands at
the very end or on SIGTERM, so ANY abrupt loss -- reclaimed container, OOM,
power cut -- takes the whole run with it. On a 3M-eval search that is 2.4
days of compute with no recoverable artefact.
- driver.search gains checkpoint=/checkpoint_every=: the current best is
handed to a callback at most every N evals. Rate-limited by evals, not
improvements, which come in bursts early. A failing checkpoint is logged
and swallowed -- losing a checkpoint is bad, losing the search because a
checkpoint failed is worse.
- homemaker-evolve --checkpoint-every N writes <out>.dom.checkpoint via
mkstemp + os.replace, so a crash can never catch it half-written. It is
deliberately NOT the output path: a checkpoint is a leaf-sharing run's
internal best, dishonest under the canonical scorer until the finish
stage unfolds it (homemaker-py-3l6), and must not be mistaken for the
finished article.
- Verified the written checkpoint re-loads as a valid .dom.
Default off, so behaviour is unchanged without the flag.
Lint at parity (46); tests 372 passed (3 new), same 2 pre-existing failures.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-08-29 05:43:58 +00:00
|
|
|
|
last_checkpoint = [0] # list so the nested recorder can rebind it
|
|
|
|
|
|
|
2026-06-12 14:22:26 +01:00
|
|
|
|
def _log(msg: str) -> None:
|
|
|
|
|
|
if log:
|
|
|
|
|
|
log(msg)
|
|
|
|
|
|
|
|
|
|
|
|
n_evals = 0
|
|
|
|
|
|
n_topologies = 0
|
2026-06-18 23:42:39 +01:00
|
|
|
|
last_improve = 0 # n_evals at the last best-fitness improvement (restart clock)
|
|
|
|
|
|
seen_sigs: set[str] = set() # §11.5 cumulative distinct topologies ever admitted
|
2026-06-12 14:22:26 +01:00
|
|
|
|
result = SearchResult(best=None, population=[], n_evals=0, n_topologies=0)
|
|
|
|
|
|
|
|
|
|
|
|
def admit(ind: Individual, pop: list[Individual]) -> None:
|
2026-06-18 23:42:39 +01:00
|
|
|
|
nonlocal n_topologies, last_improve
|
2026-06-12 14:22:26 +01:00
|
|
|
|
n_topologies += 1
|
2026-06-18 23:42:39 +01:00
|
|
|
|
seen_sigs.add(ind.sig)
|
2026-06-20 18:54:48 +01:00
|
|
|
|
# §12.3 pruned by the shape-feasibility filter: counted as an explored
|
|
|
|
|
|
# topology (so the prune rate is visible) but never bred from or ranked.
|
|
|
|
|
|
if ind.lineage.startswith("pruned/"):
|
|
|
|
|
|
return
|
2026-06-14 09:20:03 +01:00
|
|
|
|
if result.best is None or _key(ind) > _key(result.best):
|
2026-06-12 14:22:26 +01:00
|
|
|
|
result.best = ind
|
2026-06-18 23:42:39 +01:00
|
|
|
|
last_improve = n_evals
|
2026-06-12 14:22:26 +01:00
|
|
|
|
result.history.append((n_evals, ind.fitness, ind.lineage))
|
2026-06-18 23:42:39 +01:00
|
|
|
|
result.diversity_history.append(
|
|
|
|
|
|
(n_evals, len({p.sig for p in pop} | {ind.sig}), len(seen_sigs)))
|
2026-06-12 14:22:26 +01:00
|
|
|
|
_log(f"[{n_evals:6d} evals] best {ind.fitness:.6g} "
|
|
|
|
|
|
f"(fails {ind.n_fails}) via {ind.lineage}")
|
Checkpoint long searches; the cold-start runs were lost to a reclaimed box
All four 500k runs died about 10 minutes in when the container was
reclaimed. No SIGTERM fired, so no .dom was written and 0 of 12 runs
completed. My plan committed results per finished run, which protected
nothing because no run reached its commit point. The bad assumption was
reading "reclaimed after inactivity" as CPU inactivity; it is conversation
inactivity, and background compute does not hold the box open.
Progress reached before the loss (from the tracked logs): harbor 24,960
evals / 40 fails, maple 14,880 / 79, health-centre 25,920 / 33,
programme-house 138,800 / 2.
The underlying gap is not environmental: a search's only output lands at
the very end or on SIGTERM, so ANY abrupt loss -- reclaimed container, OOM,
power cut -- takes the whole run with it. On a 3M-eval search that is 2.4
days of compute with no recoverable artefact.
- driver.search gains checkpoint=/checkpoint_every=: the current best is
handed to a callback at most every N evals. Rate-limited by evals, not
improvements, which come in bursts early. A failing checkpoint is logged
and swallowed -- losing a checkpoint is bad, losing the search because a
checkpoint failed is worse.
- homemaker-evolve --checkpoint-every N writes <out>.dom.checkpoint via
mkstemp + os.replace, so a crash can never catch it half-written. It is
deliberately NOT the output path: a checkpoint is a leaf-sharing run's
internal best, dishonest under the canonical scorer until the finish
stage unfolds it (homemaker-py-3l6), and must not be mistaken for the
finished article.
- Verified the written checkpoint re-loads as a valid .dom.
Default off, so behaviour is unchanged without the flag.
Lint at parity (46); tests 372 passed (3 new), same 2 pre-existing failures.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-08-29 05:43:58 +00:00
|
|
|
|
# Crash safety for long runs. A 3M-eval search is days of compute
|
|
|
|
|
|
# whose only output lands at the very end (or on SIGTERM), so an
|
|
|
|
|
|
# abrupt loss -- a reclaimed container, an OOM, a power cut --
|
|
|
|
|
|
# takes everything with it. When `checkpoint` is given it is
|
|
|
|
|
|
# handed the current best at most every `checkpoint_every` evals,
|
|
|
|
|
|
# so the run always has a recoverable artefact on disk. Rate-limited
|
|
|
|
|
|
# by evals, not by improvements, because improvements come in
|
|
|
|
|
|
# bursts early on. A failing checkpoint must never kill the search.
|
|
|
|
|
|
if checkpoint is not None and (
|
|
|
|
|
|
n_evals - last_checkpoint[0] >= checkpoint_every):
|
|
|
|
|
|
last_checkpoint[0] = n_evals
|
|
|
|
|
|
try:
|
|
|
|
|
|
checkpoint(result.best, n_evals)
|
|
|
|
|
|
except Exception as exc: # noqa: BLE001
|
|
|
|
|
|
_log(f"[{n_evals:6d} evals] checkpoint failed: {exc!r}")
|
2026-06-18 23:42:39 +01:00
|
|
|
|
if niche_by_signature:
|
|
|
|
|
|
# §11.5 structural niching: at most one individual per topology
|
|
|
|
|
|
# signature, keeping the better of any collision. This preserves
|
|
|
|
|
|
# STRUCTURAL diversity directly — distinct topologies whose fitness
|
|
|
|
|
|
# scalars happen to coincide (common in the high-fail 0.5^n regime)
|
|
|
|
|
|
# are no longer wrongly discarded, and neutral geometry variants of an
|
|
|
|
|
|
# incumbent topology can never crowd out a rival topology.
|
|
|
|
|
|
for i, p in enumerate(pop):
|
|
|
|
|
|
if p.sig == ind.sig:
|
|
|
|
|
|
if _key(ind) > _key(p):
|
|
|
|
|
|
pop[i] = ind
|
|
|
|
|
|
return
|
|
|
|
|
|
else:
|
|
|
|
|
|
# legacy fitness-scalar dedup (population collapse guard —
|
|
|
|
|
|
# neutral mutations are common, homemaker-py-8cs)
|
|
|
|
|
|
if any(abs(ind.fitness - p.fitness) <= 1e-9 * max(abs(p.fitness), 1e-300)
|
|
|
|
|
|
for p in pop):
|
|
|
|
|
|
return
|
2026-06-12 14:22:26 +01:00
|
|
|
|
if len(pop) < pop_size:
|
|
|
|
|
|
pop.append(ind)
|
|
|
|
|
|
return
|
2026-06-14 09:20:03 +01:00
|
|
|
|
worst = min(range(len(pop)), key=lambda i: _key(pop[i]))
|
|
|
|
|
|
if _key(ind) > _key(pop[worst]):
|
2026-06-12 14:22:26 +01:00
|
|
|
|
pop[worst] = ind
|
|
|
|
|
|
|
|
|
|
|
|
pop: list[Individual] = []
|
2026-06-14 06:55:58 +01:00
|
|
|
|
|
2026-06-29 06:20:29 +01:00
|
|
|
|
# homemaker-py-psk (island model §14): optional per-child instrumentation
|
|
|
|
|
|
# hook, default off (no behaviour change). ``child_probe(ind)`` is called
|
|
|
|
|
|
# once per evaluated child. Used by the island-migration A/B to measure
|
|
|
|
|
|
# whether area-matched crossover across independently-converged elites EVER
|
|
|
|
|
|
# yields a child that beats max(parent fails) — distinguishing a mechanistic
|
|
|
|
|
|
# (alignment) null from a budget null. The crossover parents' fail counts are
|
|
|
|
|
|
# appended to the child's lineage as ``|pf=a,b`` (only when the probe is set),
|
|
|
|
|
|
# so the signal survives the ProcessPoolExecutor pickle round-trip that an
|
|
|
|
|
|
# id(root) key cannot (the worker returns a deserialised, distinct object).
|
|
|
|
|
|
|
2026-06-14 06:55:58 +01:00
|
|
|
|
# Set up optional process pool for parallel child evaluation.
|
|
|
|
|
|
_pool = None
|
|
|
|
|
|
if n_workers > 1:
|
|
|
|
|
|
from concurrent.futures import ProcessPoolExecutor
|
|
|
|
|
|
_pool = ProcessPoolExecutor(max_workers=n_workers, initializer=_worker_init)
|
|
|
|
|
|
|
|
|
|
|
|
def _run_batch(
|
|
|
|
|
|
tasks: list[tuple], # (root, x0, budget_, inner_kw_, lineage)
|
2026-06-20 18:54:48 +01:00
|
|
|
|
filter_on: bool = False,
|
2026-06-14 06:55:58 +01:00
|
|
|
|
) -> None:
|
2026-06-20 18:54:48 +01:00
|
|
|
|
"""Evaluate a batch of tasks and admit results; parallel when _pool set.
|
|
|
|
|
|
|
|
|
|
|
|
``filter_on`` enables the §12.3 shape-feasibility pre-filter for this
|
|
|
|
|
|
batch — used for mutation children only, never for the seed/bootstrap or
|
|
|
|
|
|
restart batches (construction invariants must survive)."""
|
2026-06-14 06:55:58 +01:00
|
|
|
|
nonlocal n_evals
|
2026-06-20 18:54:48 +01:00
|
|
|
|
mx = feasibility_max_shape_fails if (filter_on and feasibility_filter) else None
|
|
|
|
|
|
best_nf = result.best.n_fails if result.best is not None else None
|
2026-06-14 06:55:58 +01:00
|
|
|
|
full = [
|
2026-06-28 22:04:35 +01:00
|
|
|
|
(root, programme_dir, urb_root, x0, budget_, kw_, lin, use_grade,
|
2026-07-19 20:35:18 +01:00
|
|
|
|
mx, best_nf, leaf_sharing, superpose, max_share, conn_grade,
|
2026-08-03 21:10:28 +01:00
|
|
|
|
collapse_insearch, multi_use, shapecurve_warmstart, shapecurve_prune)
|
2026-06-14 06:55:58 +01:00
|
|
|
|
for root, x0, budget_, kw_, lin in tasks
|
|
|
|
|
|
]
|
|
|
|
|
|
if _pool is not None:
|
2026-06-22 23:25:50 +01:00
|
|
|
|
# Submit the whole batch in parallel, but admit results in SUBMISSION
|
|
|
|
|
|
# order, not completion order (homemaker-py-xcy). ``admit`` is
|
|
|
|
|
|
# order-sensitive — it accrues ``n_evals`` per result and keeps the
|
|
|
|
|
|
# FIRST individual of any equal-key tie as ``best`` — so consuming
|
|
|
|
|
|
# futures as they complete made a parallel run non-reproducible
|
|
|
|
|
|
# (completion order varies run-to-run; measured 167 vs 161 fails for
|
|
|
|
|
|
# maple-court seed 0). Iterating ``futs`` in order blocks on each in
|
|
|
|
|
|
# turn while all still run concurrently, reproducing the serial
|
|
|
|
|
|
# admission sequence exactly (verified byte-identical .dom).
|
2026-06-14 06:55:58 +01:00
|
|
|
|
futs = [_pool.submit(_evaluate, *t) for t in full]
|
2026-06-22 23:25:50 +01:00
|
|
|
|
for f in futs:
|
2026-06-14 06:55:58 +01:00
|
|
|
|
ind, used = f.result()
|
|
|
|
|
|
n_evals += used
|
2026-06-29 06:20:29 +01:00
|
|
|
|
if child_probe is not None:
|
|
|
|
|
|
child_probe(ind)
|
2026-06-14 06:55:58 +01:00
|
|
|
|
admit(ind, pop)
|
|
|
|
|
|
else:
|
|
|
|
|
|
for t in full:
|
|
|
|
|
|
ind, used = _evaluate(*t)
|
|
|
|
|
|
n_evals += used
|
2026-06-29 06:20:29 +01:00
|
|
|
|
if child_probe is not None:
|
|
|
|
|
|
child_probe(ind)
|
2026-06-14 06:55:58 +01:00
|
|
|
|
admit(ind, pop)
|
|
|
|
|
|
|
2026-06-18 23:42:39 +01:00
|
|
|
|
# A fresh seed individual (used for the initial bootstrap and for §11.5
|
|
|
|
|
|
# restart injections). Mirrors the construction order: custom seed_factory >
|
|
|
|
|
|
# programme-aware construction > random divide-grown topology.
|
§39.4: tighten generic-type matching, reverting the harbor rename
Supersedes the previous commit's approach. Renaming harbor's four colliding
codes fixed one programme; tightening the matching rule fixes the rule, so a
room may be called anything. cr1/of/st1/st2 are restored and the examples are
byte-identical to their pre-§39 state -- which also means existing .dom
artefacts (evolved-3M*) stay valid, so migrate_ju3_rename.py is deleted.
The rule: Urb has exactly three GENERIC structural types (get_space_types:
qw/C O S/), the leaves the search creates. Measured across the corpus: 154 C,
110 O, 1 S, not one lowercase generic -- while every programme code is
lowercase, including single-character ones (r, t, m, n). Case is the
discriminator, not length. Every generic test was type[0].lower() in (...), a
case-insensitive PREFIX that swept up any programme code starting with those
letters; they now match the generic set exactly. 30 sites across dom, fitness,
graph, operators, programme, shapecurve and bubble.
NOT applied to the SEMANTIC prefixes: l/k/b/t classify programme codes by first
letter (graph.py builds bedroom<->toilet and kitchen<->living relations from
them) and stay prefix-based. Where the namespaces were mixed in one expression
they were split -- has_circulation's ("b","l","k","c") is three semantic
prefixes plus dom.is_circulation; access()'s ("l","c","s") is semantic l plus
the generic circulation set.
New: dom.GENERIC_{CIRCULATION,OUTSIDE,TYPES} + is_generic(); fitness.
_generic_class(), replacing the _t0 dispatch in quality_size/quality_width/
quality_proportion/value_rate -- the four terms that mattered most and that a
first sweep missed, since they dispatch through a t0 variable rather than an
inline test. graph._adjacency_target resolves a generic adjacency requirement
(programmes write "adjacency: [c, o]") to the generic set while every other
requirement keeps Perl's prefix semantics.
Two subtleties: S is in both generic sets but takes the OUTSIDE parameter
families -- a first translation tested circulation first and silently gave S
the circulation params, caught by test_get_space_params_sahn_proportion. And
validate_codes survives, narrowed to a code spelled exactly C/O/S, which is a
genuine ambiguity; merely starting with c/o/s is now fine.
Invariant asserted as a test: test_scoring_is_invariant_under_programme_code_
spelling relabels one tree and its config together and re-scores. Bit-identical
across 12 comparisons (6 seeds x collapse on/off).
Re-baseline (seed 1, 20k, original names): 58 fails (15 hard / 43 soft) against
the real 37-instance programme, with cr1 at 79.1 m2 vs declared 80 (was 32.9
and 17.1), of/st1/st2 all present and in band, and one fail naming any of them.
57 -> 58 on a 5-instance-harder programme is within noise: "did not regress".
Fallout (§39.5): 2g7.5's CP-SAT seeder win does not survive. Over 6 seeds --
harbor real 102/114 (cpsat loses), harbor old-effective 98/99 (tie, so the win
was already marginal), maple-court 156/144 (cpsat wins). maple is the control:
the solver did not regress, harbor's programme changed. Test xfail'd with that
reason plus a maple companion; both assign_solver flags stay default off.
Filed homemaker-py-w6x to re-check other narrow-margin harbor A/Bs.
345 passed, 1 xfailed, same 7 pre-existing fixture failures.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-08-26 09:45:28 +00:00
|
|
|
|
prog = {c: r for c, r in reqs.items() if not dom.is_generic(c)}
|
2026-06-18 23:42:39 +01:00
|
|
|
|
n_target = bootstrap_n_leaves or max(len(reqs), 3)
|
|
|
|
|
|
|
|
|
|
|
|
def _make_seed_task(tag: str) -> tuple:
|
|
|
|
|
|
if seed_factory is not None:
|
|
|
|
|
|
# Custom seed (DESIGN.md §11.3 Stage 2: lift the evolved base into a
|
|
|
|
|
|
# full multi-storey design with the upper room sets instantiated by
|
|
|
|
|
|
# construction).
|
|
|
|
|
|
return (seed_factory(rng), None, child_budget, {}, f"lift/{tag}")
|
|
|
|
|
|
if prog:
|
2026-06-19 09:23:12 +01:00
|
|
|
|
topo = operators.constructive_topology(
|
Phase 7 §12.2: proportion-aware constructive seeding + storey_minimum fix (leu.2, cq1)
Size each constructive-seed cut from leaf TARGET areas (division=[f,f] gives
left area-fraction f) and pick each cut's rotation for child squareness — both
derived from target dims, topology/type assignment untouched. Area-only
regressed (slivers); rotation choice is what makes it pay.
End-to-end (20000 evals, 3 seeds, staged): harbor 85.3->74.0 (-13%, best 69),
maple-court 151.7->136.0 (-10%, best 126). PROP=0 reproduces the §11.7/§12.1
baselines exactly. programme-house regresses at fixed budget (deeper local
optimum walls off the undivide restructuring path) but a budget sweep shows
it's convergence speed, not a worse asymptote (PROP=1 reaches 1 fail at 150k).
Default-on (seed_proportion_aware=True, env PROP=1).
cq1: n_storeys now honours storey_minimum, not just level: keys — programme-house
(storey_minimum:2, all rooms level:0) was seeded one storey short and fell
through to plain search. New programme.storey_minimum()/n_storeys_for();
driver.search passes min_storeys to the seeder; search_staged routes on the max.
No-op for harbor/maple; programme-house single-stage 8.0->5.0.
New maple-court best (126) saved as generated.dom. 204 tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 14:04:42 +01:00
|
|
|
|
seed_root, reqs, rng, types, min_storeys=min_storeys,
|
|
|
|
|
|
adjacency_aware=seed_adjacency_aware,
|
2026-06-21 21:10:18 +01:00
|
|
|
|
proportion_aware=seed_proportion_aware,
|
2026-06-24 18:16:17 +01:00
|
|
|
|
circ_divisor=circ_divisor,
|
2026-06-25 22:36:24 +01:00
|
|
|
|
leaf_sharing=leaf_sharing, leaf_share_factor=leaf_share_factor,
|
2026-06-28 07:20:20 +01:00
|
|
|
|
depth_balanced=depth_balanced,
|
2026-07-28 00:00:51 +01:00
|
|
|
|
interior_outside=interior_outside, outside_divisor=outside_divisor,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
construction_beam_width=construction_beam_width,
|
§39.10: preserving constructed connectivity is NULL — and it reframes §39.9
§39.9 named the upstream fix: keep circulation connected DURING the resize
rather than rebuilding it after. Built and measured. It does not help, and the
reason matters more than the lever.
Both halves of the re-cut do damage, in different proportions per programme.
Freezing rotations and letting only ratios move (% levels connected, 12 seeds):
harbor 100 -> 71 -> 50, health-centre 100 -> 8 -> 8, maple 100 -> 92 -> 67. So
health-centre is destroyed entirely by the ratio and maple mostly by the
rotation; a fix must be able to give back either.
operators._size_divisions_preserving_circulation snapshots every cut, resizes,
then reverts the cuts on the tree path between each circulation pair the resize
broke -- programme fully intact, no retyping, only geometry given back. It works
on connectivity (harbor 50->92%, maple 67->97%, health-centre 8->17%) and costs
area accuracy: constructed-seed fails harbor 96.6->141.5, maple 141.8->175.8,
size fails roughly double. (A greedy single-cut revert barely moved -- it stalls
where no ONE revert helps though two would. Targeting the broken pairs is what
made connectivity work.)
The obvious defence -- raw constructed seeds understate it, the resize is only a
warm start, the inner loop should recover -- was TESTED AND FAILS. Full search,
harbor-house, 12000 evals, seed 1:
OFF 43 fails, 9 hard, 3 connectivity
ON 65 fails, 26 hard, 4 connectivity
Worse on every axis, including connectivity itself.
REFRAMING: §39.9's fact stands (the resize destroys 41 of 49 circulation edges)
but is NOT ACTIONABLE, because construction-time connectivity does not determine
final connectivity. The search discards and rebuilds the seeder's circulation
either way, and constraining the seed only spends area quality the search cannot
recover. Together with §39.8 (not an incentive problem) that retires the framing
this thread inherited from §38: connectivity is neither a construction problem
nor an incentive one.
Both flags (repair_circulation, preserve_circulation) stay default off with the
numbers recorded, plus byte-identical-default tests. Do not revisit either
without a new formulation -- the standing this document gives bubble.py.
356 passed (+1 new), same 7 pre-existing fixture failures, lint unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-08-26 15:11:46 +00:00
|
|
|
|
multi_use=multi_use, assign_solver=assign_solver,
|
|
|
|
|
|
preserve_circulation=preserve_circulation)
|
2026-06-18 23:42:39 +01:00
|
|
|
|
return (topo, None, child_budget, {}, f"construct/{tag}")
|
|
|
|
|
|
n = int(rng.integers(max(1, n_target - 1), n_target + 2))
|
|
|
|
|
|
return (random_topology(seed_root, n, rng, types), None, child_budget,
|
|
|
|
|
|
{}, f"bootstrap/{tag}")
|
|
|
|
|
|
|
2026-06-14 07:48:13 +01:00
|
|
|
|
interrupted = False
|
2026-06-14 06:55:58 +01:00
|
|
|
|
try:
|
2026-07-16 08:38:08 +01:00
|
|
|
|
if seed_pop is not None:
|
|
|
|
|
|
# homemaker-py-kpu (Schedule B): carry a whole population across a
|
|
|
|
|
|
# grain-anneal phase change. Each root is re-optimised and re-scored
|
|
|
|
|
|
# under THIS phase's evaluator (leaf_sharing/max_share) as the initial
|
|
|
|
|
|
# population, so gross topology/adjacency continuity is preserved while
|
|
|
|
|
|
# the effective problem is refined — not restarted from a single best.
|
|
|
|
|
|
_run_batch([(copy.deepcopy(r), None, seed_budget, {},
|
|
|
|
|
|
f"anneal-seed/{i}") for i, r in enumerate(seed_pop)])
|
|
|
|
|
|
elif do_bootstrap:
|
2026-06-14 06:55:58 +01:00
|
|
|
|
# Bootstrap: diverse initial population from random topologies.
|
|
|
|
|
|
# Each individual is a cold start, so use the exploratory sigma
|
|
|
|
|
|
# schedule (inner_kw={} → cma_search defaults: sigmas=(0.05, 0.15)).
|
|
|
|
|
|
# Leaf count varied ±1 around the target to increase structural diversity.
|
2026-06-17 22:51:58 +01:00
|
|
|
|
# Programme-aware constructive seeding (§11.2): when the programme
|
|
|
|
|
|
# has required spaces, instantiate each by construction so the seed
|
|
|
|
|
|
# population starts with ~zero missing-space failures instead of a
|
|
|
|
|
|
# random divide+retype walk that leaves required rooms absent.
|
2026-06-18 23:42:39 +01:00
|
|
|
|
_run_batch([_make_seed_task(str(i)) for i in range(pop_size)])
|
2026-06-12 14:22:26 +01:00
|
|
|
|
else:
|
2026-06-14 06:55:58 +01:00
|
|
|
|
seed_ind, used = _evaluate(copy.deepcopy(seed_root), programme_dir, urb_root,
|
|
|
|
|
|
x0=None, budget=seed_budget,
|
2026-06-18 22:33:29 +01:00
|
|
|
|
inner_kw={}, lineage="seed",
|
2026-06-28 22:04:35 +01:00
|
|
|
|
want_grade=use_grade,
|
9o5: type superposition + per-eval collapse (multi-use leaves)
Interchangeable codes (similar size/width/proportion, compatible level/stack,
no adjacency edge) form equivalence classes derived from the programme. With
--superpose (default off), each fitness eval COLLAPSES every superposed leaf to
its best in-class usage via an optimal supply->demand assignment (brute force
<=C! within cap C=4, scipy Hungarian beyond), then scores the condensed types.
Because collapse re-types on the unmerged tree before all checks, counts /
adjacency / quality are unchanged downstream -- no Node field, no graph/operator
changes -- and default OFF is bit-identical.
- programme.py: derive_interchange_classes + interchangeable (S1-S4, locked
thresholds R_SIZE=1.5/R_WIDTH=1.3/R_PROP=1.5, CLASS_CAP=4)
- fitness.py: collapse_superposition, _best_assignment, _usage_quality;
superpose/superpose_class_cap conf knobs; collapse hooked into _evaluate_full
- driver.py/evolve.py: superpose flag plumbed beside leaf_sharing; --superpose
- tests/test_superposition.py: 17 tests (derivation, assignment, end-to-end)
Closes homemaker-py-9o5 (build); validation A/B is homemaker-py-xi7.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:08:46 +01:00
|
|
|
|
leaf_sharing=leaf_sharing,
|
2026-07-16 08:38:08 +01:00
|
|
|
|
superpose=superpose,
|
2026-07-18 18:44:24 +01:00
|
|
|
|
max_share=max_share,
|
2026-07-19 20:35:18 +01:00
|
|
|
|
conn_grade=conn_grade,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
collapse_insearch=collapse_insearch,
|
homemaker-py-6xh: wire shape-curve DP into driver.py as an NM warm-start
Promotes the validated shape-curve DP (experiments/shapecurve_spike.py,
2g7.4, DESIGN.md §37.2) from a reference-only spike into
src/homemaker_layout/shapecurve.py, and wires it into driver._evaluate as a
warm-start for innerloop.optimise: when eligible (single storey, no
leaf_sharing/superpose/max_share/multi_use) and no caller-supplied x0, the
DP's exact shape-feasible ratio point is written onto the tree before NM
runs, off by default (shapecurve_warmstart=/--shapecurve-warmstart).
Caught and fixed a latent bug promoting the spike: realise() could leave
numpy.float64 in `division`, which yaml.safe_dump can't serialise — the
original spike never round-tripped through dom.dumps so this was never hit.
A/B on harbor-house-l0 (experiments/ab_shapecurve_warmstart.py, budget=2000,
5 seeds): mean total fails 16.6 (on) vs 19.6 (off), ~3.5x mean fitness
improvement; mean hard-fail count alone was a noise-level wash at this
sample size. Full writeup in DESIGN.md §37.4.
Deliberately deferred to new tracked beads (children of 2g7): DP-exact hard
pre-filter (wkh), multi-storey below-link support (koo), leaf_sharing/
co_type modelling (tym), true skew-quad polygon algebra (ekc) — 6xh stays
in_progress pending those.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSwQwpEaHFBkeVSDDWd75S
2026-08-03 18:43:28 +01:00
|
|
|
|
multi_use=multi_use,
|
2026-08-03 21:10:28 +01:00
|
|
|
|
shapecurve_warmstart=shapecurve_warmstart,
|
|
|
|
|
|
shapecurve_prune=shapecurve_prune)
|
2026-06-14 06:55:58 +01:00
|
|
|
|
n_evals += used
|
|
|
|
|
|
admit(seed_ind, pop)
|
|
|
|
|
|
|
|
|
|
|
|
while n_evals < budget:
|
2026-06-18 23:42:39 +01:00
|
|
|
|
# §11.5 diversity restart: if the best has not improved for
|
|
|
|
|
|
# restart_patience evals, keep the top restart_elite incumbents and
|
|
|
|
|
|
# refill the population with fresh constructive/random seeds. This
|
|
|
|
|
|
# re-injects the upfront structural diversity a single mutation chain
|
|
|
|
|
|
# loses (the blank-slate gap, §7 Phase 2) — the soft-restart analog of
|
|
|
|
|
|
# urb-evolve's random initial population. Off by default
|
|
|
|
|
|
# (restart_patience=None) so existing experiments are unaffected.
|
|
|
|
|
|
if (restart_patience is not None and pop
|
|
|
|
|
|
and n_evals - last_improve >= restart_patience
|
|
|
|
|
|
and n_evals + child_budget <= budget):
|
|
|
|
|
|
keep = sorted(pop, key=_key, reverse=True)[:max(1, restart_elite)]
|
|
|
|
|
|
pop[:] = keep
|
|
|
|
|
|
result.n_restarts += 1
|
|
|
|
|
|
last_improve = n_evals # reset clock; avoid immediate re-trigger
|
|
|
|
|
|
n_fresh = min(pop_size - len(pop),
|
|
|
|
|
|
max(0, (budget - n_evals) // child_budget))
|
|
|
|
|
|
_log(f"[{n_evals:6d} evals] restart #{result.n_restarts}: "
|
|
|
|
|
|
f"keep {len(keep)}, inject {n_fresh} fresh seeds")
|
|
|
|
|
|
if n_fresh:
|
|
|
|
|
|
_run_batch([_make_seed_task(f"r{result.n_restarts}.{i}")
|
|
|
|
|
|
for i in range(n_fresh)])
|
|
|
|
|
|
continue
|
2026-06-14 06:55:58 +01:00
|
|
|
|
# How many children to generate this iteration: n_workers in parallel,
|
|
|
|
|
|
# but cap at what the remaining budget can afford (ceiling division).
|
|
|
|
|
|
batch_n = (
|
|
|
|
|
|
min(n_workers,
|
|
|
|
|
|
max(1, (budget - n_evals + child_budget - 1) // child_budget))
|
|
|
|
|
|
if _pool is not None else 1
|
|
|
|
|
|
)
|
|
|
|
|
|
tasks = []
|
|
|
|
|
|
for _ in range(batch_n):
|
|
|
|
|
|
if len(pop) >= 2 and rng.random() < p_crossover:
|
2026-06-29 22:56:23 +01:00
|
|
|
|
a, b = (_tournament(pop, rng, _key, k=tournament_k),
|
|
|
|
|
|
_tournament(pop, rng, _key, k=tournament_k))
|
2026-06-14 06:55:58 +01:00
|
|
|
|
child_root, _, desc = operators.crossover(a.root, b.root, rng)
|
2026-06-29 06:20:29 +01:00
|
|
|
|
if child_probe is not None:
|
|
|
|
|
|
desc = f"{desc}|pf={a.n_fails},{b.n_fails}"
|
2026-06-14 06:55:58 +01:00
|
|
|
|
ratios = {**b.ratios, **a.ratios} # primary parent wins
|
|
|
|
|
|
else:
|
2026-06-29 22:56:23 +01:00
|
|
|
|
parent = _tournament(pop, rng, _key, k=tournament_k)
|
2026-06-14 06:55:58 +01:00
|
|
|
|
child_root, desc = operators.mutate(parent.root, rng, types,
|
2026-06-20 18:54:48 +01:00
|
|
|
|
weights=mutation_weights,
|
2026-07-22 17:43:36 +01:00
|
|
|
|
reqs=reqs, base_p=base_p,
|
|
|
|
|
|
fit=shape_repair_fit)
|
2026-06-15 07:27:03 +01:00
|
|
|
|
# Carry operator-specified ratios for nodes that are genuinely
|
|
|
|
|
|
# newly divided (existed as leaves in the parent, are now
|
|
|
|
|
|
# divided in the child). Structural mutations (e.g. swap) can
|
|
|
|
|
|
# reveal previously-hidden nodes whose stale pre-writeback
|
|
|
|
|
|
# ratios must NOT be propagated — those default to 0.5.
|
|
|
|
|
|
parent_lvls = dom.levels(parent.root)
|
|
|
|
|
|
new_splits = {
|
|
|
|
|
|
(li, path): val
|
|
|
|
|
|
for (li, path), val in innerloop.ratio_map(child_root).items()
|
|
|
|
|
|
if li >= len(parent_lvls)
|
|
|
|
|
|
or not (pn := parent_lvls[li].by_id(path))
|
|
|
|
|
|
or not pn.divided
|
|
|
|
|
|
}
|
|
|
|
|
|
ratios = {**new_splits, **parent.ratios}
|
2026-06-14 06:55:58 +01:00
|
|
|
|
x0 = innerloop.warm_x0(child_root, ratios)
|
|
|
|
|
|
tasks.append((child_root, x0, child_budget, inner_kw, desc))
|
2026-06-20 18:54:48 +01:00
|
|
|
|
_run_batch(tasks, filter_on=True)
|
2026-06-14 07:48:13 +01:00
|
|
|
|
except KeyboardInterrupt:
|
|
|
|
|
|
interrupted = True
|
|
|
|
|
|
_log(f"[{n_evals:6d} evals] interrupted — returning best-so-far")
|
2026-06-14 06:55:58 +01:00
|
|
|
|
finally:
|
|
|
|
|
|
if _pool is not None:
|
|
|
|
|
|
_pool.shutdown(wait=True)
|
2026-06-12 14:22:26 +01:00
|
|
|
|
|
2026-06-14 09:20:03 +01:00
|
|
|
|
result.population = sorted(pop, key=_key, reverse=True)
|
2026-06-12 14:22:26 +01:00
|
|
|
|
result.n_evals = n_evals
|
|
|
|
|
|
result.n_topologies = n_topologies
|
2026-06-18 23:42:39 +01:00
|
|
|
|
result.n_distinct_signatures = len(seen_sigs)
|
2026-06-14 07:48:13 +01:00
|
|
|
|
result.interrupted = interrupted
|
2026-06-12 14:22:26 +01:00
|
|
|
|
return result
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
|
|
|
|
|
|
|
2026-07-15 10:21:58 +01:00
|
|
|
|
def polish_finish(
|
|
|
|
|
|
result: SearchResult,
|
|
|
|
|
|
programme_dir: str | Path,
|
|
|
|
|
|
*,
|
|
|
|
|
|
polish_budget: int,
|
|
|
|
|
|
pop_size: int = 8,
|
|
|
|
|
|
child_budget: int = 80,
|
|
|
|
|
|
p_crossover: float = 0.2,
|
|
|
|
|
|
seed: int = 0,
|
|
|
|
|
|
n_workers: int = 1,
|
|
|
|
|
|
superpose: bool = False,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
multi_use: bool = False,
|
2026-07-24 09:55:34 +01:00
|
|
|
|
collapse_insearch: bool = True,
|
2026-07-15 10:21:58 +01:00
|
|
|
|
rescore_budget: int = 200,
|
|
|
|
|
|
log=None,
|
|
|
|
|
|
) -> SearchResult:
|
|
|
|
|
|
"""homemaker-py-3l6: convert a leaf-sharing run's dishonest best into a
|
|
|
|
|
|
canonically-scored, materialised output.
|
|
|
|
|
|
|
|
|
|
|
|
A sharing run's internal objective credits a shared leaf (``share=k``) as k
|
|
|
|
|
|
programme rooms with its size target re-centred on ``k*target``, so
|
|
|
|
|
|
``result.best`` looks good internally but is k−1 rooms short per shared leaf
|
|
|
|
|
|
under the canonical (sharing-off) scorer — the divergence this bug is about.
|
|
|
|
|
|
This:
|
|
|
|
|
|
|
|
|
|
|
|
1. **Unfolds** every live shared leaf into k distinct sibling rooms
|
|
|
|
|
|
(:func:`operators.unfold_shared_leaves`), paying down the materialisation
|
|
|
|
|
|
deficit that otherwise leaves the de-shared genome deep in the missing-room
|
|
|
|
|
|
fail hole (yaa: naive warm-start without unfold stalls ~60× worse).
|
|
|
|
|
|
2. **Polishes** the unfolded genome with a warm-started ``leaf_sharing=False``
|
|
|
|
|
|
search (``polish_budget`` evals) so the freshly materialised children get
|
|
|
|
|
|
their proportion/width/size cleaned up. yaa proved this unfold-then-polish
|
|
|
|
|
|
path catches the direct no-sharing route (harbor-house 4.19e-06).
|
|
|
|
|
|
|
|
|
|
|
|
With ``polish_budget <= 0`` the polish is skipped: the unfolded genome is just
|
|
|
|
|
|
re-optimised once and canonically scored (honest output, no extra search —
|
|
|
|
|
|
used on interrupt). Either way the returned result's ``best.fitness`` is the
|
|
|
|
|
|
canonical score (leaf_sharing off ⇒ internal == canonical), and eval /
|
|
|
|
|
|
topology / history accounting is stitched onto the sharing run.
|
|
|
|
|
|
"""
|
|
|
|
|
|
def _log(msg: str) -> None:
|
|
|
|
|
|
if log:
|
|
|
|
|
|
log(msg)
|
|
|
|
|
|
|
|
|
|
|
|
if result.best is None:
|
|
|
|
|
|
return result
|
|
|
|
|
|
|
|
|
|
|
|
unfolded = copy.deepcopy(result.best.root)
|
|
|
|
|
|
n_created = operators.unfold_shared_leaves(unfolded)
|
|
|
|
|
|
_log(f"[finish] unfold: materialised {n_created} shared-leaf "
|
|
|
|
|
|
f"{'copy' if n_created == 1 else 'copies'}")
|
|
|
|
|
|
|
|
|
|
|
|
if polish_budget > 0:
|
|
|
|
|
|
r2 = search(
|
|
|
|
|
|
unfolded, programme_dir, budget=polish_budget, pop_size=pop_size,
|
|
|
|
|
|
child_budget=child_budget, p_crossover=p_crossover, seed=seed,
|
|
|
|
|
|
n_workers=n_workers, bootstrap=False, leaf_sharing=False,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
superpose=superpose, multi_use=multi_use,
|
|
|
|
|
|
collapse_insearch=collapse_insearch, log=log,
|
2026-07-15 10:21:58 +01:00
|
|
|
|
)
|
|
|
|
|
|
else:
|
|
|
|
|
|
# No polish: re-optimise the unfolded genome's ratios once and score it
|
|
|
|
|
|
# canonically so the written .dom and reported fitness are honest.
|
|
|
|
|
|
ind, used = _evaluate(
|
|
|
|
|
|
unfolded, programme_dir, None, x0=None, budget=rescore_budget,
|
2026-07-19 20:35:18 +01:00
|
|
|
|
inner_kw={}, lineage="unfold", leaf_sharing=False, superpose=superpose,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
multi_use=multi_use, collapse_insearch=collapse_insearch)
|
2026-07-15 10:21:58 +01:00
|
|
|
|
r2 = SearchResult(best=ind, population=[ind], n_evals=used, n_topologies=1)
|
|
|
|
|
|
r2.n_distinct_signatures = 1
|
|
|
|
|
|
r2.history = [(0, ind.fitness, ind.lineage)]
|
|
|
|
|
|
|
|
|
|
|
|
# Stitch the polish/rescore onto the sharing run so totals are cumulative and
|
|
|
|
|
|
# the history shows the phase change (sharing fitness is not comparable to the
|
|
|
|
|
|
# canonical polish fitness, so the two phases are tagged, not merged linearly).
|
|
|
|
|
|
r2.history = (
|
|
|
|
|
|
[(e, f, f"share:{lin}") for e, f, lin in result.history]
|
|
|
|
|
|
+ [(e + result.n_evals, f, f"polish:{lin}") for e, f, lin in r2.history]
|
|
|
|
|
|
)
|
|
|
|
|
|
r2.n_evals += result.n_evals
|
|
|
|
|
|
r2.n_topologies += result.n_topologies
|
|
|
|
|
|
r2.n_distinct_signatures += result.n_distinct_signatures
|
|
|
|
|
|
r2.n_restarts += result.n_restarts
|
|
|
|
|
|
r2.interrupted = r2.interrupted or result.interrupted
|
|
|
|
|
|
return r2
|
|
|
|
|
|
|
|
|
|
|
|
|
94g: public-access pin + keep-better wrapper + CLI/finish-hook wiring
Public-access term (preserve_public_access, default on): when the building's
only street access is an l/k ROOM neighbour of a public outside leaf (no
circulation fallback — an existential building-level check the per-leaf
objective can't see), that leaf is pinned (kept, its demand slot decremented)
so the collapse can't drop "no outside public access". Best layout 15→13
becomes 15→12 with zero new fails; sweep total 172→171, still monotone.
collapse_finish(root, **kw) -> (tree, base, coll, applied): keep-better wrapper,
scores on throwaway copies (score_with_fails merges in place), returns the
collapse only if fails don't increase.
Wiring: driver.collapse_best updates result.best (lineage +collapse, canonical
re-score); evolve.py runs it after the sharing polish behind --collapse/
--no-collapse (default on). New homemaker-collapse CLI (collapse_cmd.py) applies
it to an existing .dom, writing <stem>.collapsed.dom.
tests/test_collapse_global.py: demand-set relabel, level hard constraint, c/o/s
exclusion, no-op safety, keep-better/unmerged. 267 pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 10:29:44 +01:00
|
|
|
|
def collapse_best(
|
|
|
|
|
|
result: SearchResult,
|
|
|
|
|
|
programme_dir: str | Path,
|
|
|
|
|
|
*,
|
|
|
|
|
|
leaf_sharing: bool = False,
|
|
|
|
|
|
superpose: bool = False,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
multi_use: bool = False,
|
2026-08-05 07:52:44 +01:00
|
|
|
|
max_share: int | None = None,
|
|
|
|
|
|
conn_grade: bool = False,
|
94g: public-access pin + keep-better wrapper + CLI/finish-hook wiring
Public-access term (preserve_public_access, default on): when the building's
only street access is an l/k ROOM neighbour of a public outside leaf (no
circulation fallback — an existential building-level check the per-leaf
objective can't see), that leaf is pinned (kept, its demand slot decremented)
so the collapse can't drop "no outside public access". Best layout 15→13
becomes 15→12 with zero new fails; sweep total 172→171, still monotone.
collapse_finish(root, **kw) -> (tree, base, coll, applied): keep-better wrapper,
scores on throwaway copies (score_with_fails merges in place), returns the
collapse only if fails don't increase.
Wiring: driver.collapse_best updates result.best (lineage +collapse, canonical
re-score); evolve.py runs it after the sharing polish behind --collapse/
--no-collapse (default on). New homemaker-collapse CLI (collapse_cmd.py) applies
it to an existing .dom, writing <stem>.collapsed.dom.
tests/test_collapse_global.py: demand-set relabel, level hard constraint, c/o/s
exclusion, no-op safety, keep-better/unmerged. 267 pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 10:29:44 +01:00
|
|
|
|
log=None,
|
|
|
|
|
|
**collapse_kw,
|
|
|
|
|
|
) -> SearchResult:
|
|
|
|
|
|
"""homemaker-py-94g: finish-time global cell→room collapse on the best layout.
|
|
|
|
|
|
|
|
|
|
|
|
Relabels the best tree's room cells to the programme rooms they fit best via
|
|
|
|
|
|
one optimal assignment (hard level constraint, adjacency relaxation, and
|
|
|
|
|
|
public-access pinning — see :meth:`fitness.Fitness.collapse_global`), keeping
|
|
|
|
|
|
the result only if the fail count does not increase (:meth:`collapse_finish`).
|
|
|
|
|
|
A strictly monotone finish-time polish that searches only labels, not
|
|
|
|
|
|
geometry, so it cannot touch shape-intrinsic fails (long-thin cells, etc.).
|
|
|
|
|
|
|
|
|
|
|
|
Updates ``result.best`` in place with the canonically re-scored relabelling
|
2026-08-05 07:52:44 +01:00
|
|
|
|
when it helps; otherwise leaves the result untouched.
|
|
|
|
|
|
|
|
|
|
|
|
homemaker-py-sd3: the evaluator built here is deliberately CANONICAL
|
|
|
|
|
|
(``collapse_insearch=False``) regardless of whether the run being finished
|
|
|
|
|
|
used in-search collapse — both the keep-better guard and the reported
|
|
|
|
|
|
post-collapse fail count must match what ``homemaker-fitness`` reports for
|
|
|
|
|
|
the written ``.dom`` (no in-search override on disk), not this run's
|
|
|
|
|
|
in-search objective. ``max_share``/``conn_grade`` are still threaded through
|
|
|
|
|
|
so the evaluator's config otherwise matches the run (matters when
|
|
|
|
|
|
``leaf_sharing`` is on, e.g. a kpu/anneal grain that hasn't been unfolded
|
|
|
|
|
|
yet, or qi6's graded scalar in the reported grade).
|
|
|
|
|
|
"""
|
94g: public-access pin + keep-better wrapper + CLI/finish-hook wiring
Public-access term (preserve_public_access, default on): when the building's
only street access is an l/k ROOM neighbour of a public outside leaf (no
circulation fallback — an existential building-level check the per-leaf
objective can't see), that leaf is pinned (kept, its demand slot decremented)
so the collapse can't drop "no outside public access". Best layout 15→13
becomes 15→12 with zero new fails; sweep total 172→171, still monotone.
collapse_finish(root, **kw) -> (tree, base, coll, applied): keep-better wrapper,
scores on throwaway copies (score_with_fails merges in place), returns the
collapse only if fails don't increase.
Wiring: driver.collapse_best updates result.best (lineage +collapse, canonical
re-score); evolve.py runs it after the sharing polish behind --collapse/
--no-collapse (default on). New homemaker-collapse CLI (collapse_cmd.py) applies
it to an existing .dom, writing <stem>.collapsed.dom.
tests/test_collapse_global.py: demand-set relabel, level hard constraint, c/o/s
exclusion, no-op safety, keep-better/unmerged. 267 pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 10:29:44 +01:00
|
|
|
|
if result.best is None:
|
|
|
|
|
|
return result
|
|
|
|
|
|
|
2026-08-05 07:52:44 +01:00
|
|
|
|
fit = _fitness_for(str(programme_dir), leaf_sharing, superpose, max_share,
|
|
|
|
|
|
conn_grade, collapse_insearch=False, multi_use=multi_use)
|
94g: public-access pin + keep-better wrapper + CLI/finish-hook wiring
Public-access term (preserve_public_access, default on): when the building's
only street access is an l/k ROOM neighbour of a public outside leaf (no
circulation fallback — an existential building-level check the per-leaf
objective can't see), that leaf is pinned (kept, its demand slot decremented)
so the collapse can't drop "no outside public access". Best layout 15→13
becomes 15→12 with zero new fails; sweep total 172→171, still monotone.
collapse_finish(root, **kw) -> (tree, base, coll, applied): keep-better wrapper,
scores on throwaway copies (score_with_fails merges in place), returns the
collapse only if fails don't increase.
Wiring: driver.collapse_best updates result.best (lineage +collapse, canonical
re-score); evolve.py runs it after the sharing polish behind --collapse/
--no-collapse (default on). New homemaker-collapse CLI (collapse_cmd.py) applies
it to an existing .dom, writing <stem>.collapsed.dom.
tests/test_collapse_global.py: demand-set relabel, level hard constraint, c/o/s
exclusion, no-op safety, keep-better/unmerged. 267 pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 10:29:44 +01:00
|
|
|
|
tree, base_fails, coll_fails, applied = fit.collapse_finish(
|
|
|
|
|
|
result.best.root, **collapse_kw
|
|
|
|
|
|
)
|
|
|
|
|
|
if log:
|
|
|
|
|
|
verb = "applied" if applied else "reverted — no improvement"
|
|
|
|
|
|
log(f"[finish] collapse: {base_fails} → {coll_fails} fails ({verb})")
|
|
|
|
|
|
if applied:
|
|
|
|
|
|
score, fails, grade = fit.score_with_grade(copy.deepcopy(tree))
|
|
|
|
|
|
result.best = Individual(
|
|
|
|
|
|
root=tree,
|
|
|
|
|
|
fitness=score,
|
|
|
|
|
|
n_fails=len(fails),
|
|
|
|
|
|
ratios=result.best.ratios,
|
|
|
|
|
|
lineage=result.best.lineage + "+collapse",
|
|
|
|
|
|
grade=grade,
|
|
|
|
|
|
sig=result.best.sig,
|
|
|
|
|
|
)
|
|
|
|
|
|
return result
|
|
|
|
|
|
|
|
|
|
|
|
|
2026-07-16 08:38:08 +01:00
|
|
|
|
def search_annealed(
|
|
|
|
|
|
seed_root: dom.Node,
|
|
|
|
|
|
programme_dir: str | Path,
|
|
|
|
|
|
*,
|
|
|
|
|
|
budget: int,
|
|
|
|
|
|
polish_budget: int,
|
|
|
|
|
|
grain_ladder: tuple[int, ...] = (4, 3, 2),
|
|
|
|
|
|
pop_size: int = 8,
|
|
|
|
|
|
child_budget: int = 80,
|
|
|
|
|
|
seed_budget: int = 200,
|
|
|
|
|
|
p_crossover: float = 0.2,
|
|
|
|
|
|
seed: int = 0,
|
|
|
|
|
|
types: list[str] | None = None,
|
|
|
|
|
|
inner_kw: dict | None = None,
|
|
|
|
|
|
n_workers: int = 1,
|
|
|
|
|
|
superpose: bool = False,
|
|
|
|
|
|
log=None,
|
|
|
|
|
|
**search_kw,
|
|
|
|
|
|
) -> SearchResult:
|
|
|
|
|
|
"""homemaker-py-kpu (DESIGN.md §16): in-run leaf-share grain annealing.
|
|
|
|
|
|
|
|
|
|
|
|
Schedule B from ``homemaker-py-yaa``. Instead of a single hard sharing→off
|
|
|
|
|
|
transition (§15's unfold+polish finish), ramp the leaf-share grain **down**
|
|
|
|
|
|
across phases within one continuous run — e.g. ``grain_ladder=(4, 3, 2)`` then
|
|
|
|
|
|
off — carrying the whole population across each step. This is graduated
|
|
|
|
|
|
non-convexity: the coarse early grain fixes gross topology/adjacency on a
|
|
|
|
|
|
small effective problem; each step refines it, so no single fitness cliff has
|
|
|
|
|
|
to be crossed at once.
|
|
|
|
|
|
|
|
|
|
|
|
Each grain step lowers the evaluator's ``leaf_share_max`` cap and, *before*
|
|
|
|
|
|
resuming, unfolds every population leaf whose ``share`` exceeds the new cap
|
|
|
|
|
|
(:func:`operators.unfold_shared_leaves` with ``above=cap``) so the carried
|
|
|
|
|
|
population stays materialised — the leaves the lower cap would under-credit
|
|
|
|
|
|
become real rooms instead of fresh missing fails. The final phase de-shares
|
|
|
|
|
|
entirely (``leaf_sharing=False``, unfold ``above=1``) and polishes under the
|
|
|
|
|
|
canonical objective, so the returned ``best.fitness`` is the honest canonical
|
|
|
|
|
|
score exactly as §15's finish guarantees.
|
|
|
|
|
|
|
|
|
|
|
|
``grain_ladder`` is deduped and sorted descending; entries < 2 are dropped
|
|
|
|
|
|
(no-op grain). ``budget`` is split evenly across the sharing phases (remainder
|
|
|
|
|
|
to the first); ``polish_budget`` funds the final de-share phase (``<= 0`` or an
|
|
|
|
|
|
interrupt ⇒ unfold + single rescore only, no search — honest but unpolished).
|
|
|
|
|
|
Extra keyword args forward to :func:`search`.
|
|
|
|
|
|
"""
|
|
|
|
|
|
def _log(msg: str) -> None:
|
|
|
|
|
|
if log:
|
|
|
|
|
|
log(msg)
|
|
|
|
|
|
|
|
|
|
|
|
ladder = sorted({int(g) for g in grain_ladder if int(g) >= 2}, reverse=True)
|
|
|
|
|
|
if not ladder:
|
|
|
|
|
|
# Degenerate ladder (all grains < 2) ⇒ nothing to anneal: a plain
|
|
|
|
|
|
# no-sharing search over the full budget, honest by construction.
|
|
|
|
|
|
return search(
|
|
|
|
|
|
seed_root, programme_dir, budget=budget + max(0, polish_budget),
|
|
|
|
|
|
pop_size=pop_size, child_budget=child_budget, seed_budget=seed_budget,
|
|
|
|
|
|
p_crossover=p_crossover, seed=seed, types=types, inner_kw=inner_kw,
|
|
|
|
|
|
n_workers=n_workers, leaf_sharing=False, superpose=superpose, log=log,
|
|
|
|
|
|
**search_kw)
|
|
|
|
|
|
|
|
|
|
|
|
n_phases = len(ladder)
|
|
|
|
|
|
base = budget // n_phases
|
|
|
|
|
|
phase_budgets = [base] * n_phases
|
|
|
|
|
|
phase_budgets[0] += budget - base * n_phases # remainder to phase 0
|
|
|
|
|
|
|
|
|
|
|
|
def _stitch(acc: "SearchResult | None", r: SearchResult, tag: str) -> SearchResult:
|
|
|
|
|
|
"""Concatenate phase ``r`` onto ``acc`` with cumulative accounting and a
|
|
|
|
|
|
tagged history (objectives differ across grains, so histories are tagged
|
|
|
|
|
|
and concatenated, never merged linearly — as §15's finish does)."""
|
|
|
|
|
|
r.history = [(e, f, f"{tag}:{lin}") for e, f, lin in r.history]
|
|
|
|
|
|
r.diversity_history = list(r.diversity_history)
|
|
|
|
|
|
if acc is None:
|
|
|
|
|
|
return r
|
|
|
|
|
|
prev = acc.n_evals
|
|
|
|
|
|
r.n_evals += prev
|
|
|
|
|
|
r.n_topologies += acc.n_topologies
|
|
|
|
|
|
r.n_distinct_signatures += acc.n_distinct_signatures
|
|
|
|
|
|
r.n_restarts += acc.n_restarts
|
|
|
|
|
|
r.interrupted = r.interrupted or acc.interrupted
|
|
|
|
|
|
r.history = acc.history + [(e + prev, f, lin) for e, f, lin in r.history]
|
|
|
|
|
|
r.diversity_history = (
|
|
|
|
|
|
acc.diversity_history
|
|
|
|
|
|
+ [(e + prev, d, c) for e, d, c in r.diversity_history])
|
|
|
|
|
|
return r
|
|
|
|
|
|
|
|
|
|
|
|
combined: SearchResult | None = None
|
|
|
|
|
|
prev_pop: list[Individual] = []
|
|
|
|
|
|
|
|
|
|
|
|
for i, cap in enumerate(ladder):
|
|
|
|
|
|
if i == 0:
|
|
|
|
|
|
_log(f"[anneal] phase 1/{n_phases}: grain {cap}, budget "
|
|
|
|
|
|
f"{phase_budgets[0]} (construct population)")
|
|
|
|
|
|
r = search(
|
|
|
|
|
|
seed_root, programme_dir, budget=phase_budgets[0], pop_size=pop_size,
|
|
|
|
|
|
child_budget=child_budget, seed_budget=seed_budget,
|
|
|
|
|
|
p_crossover=p_crossover, seed=seed, types=types, inner_kw=inner_kw,
|
|
|
|
|
|
n_workers=n_workers, leaf_sharing=True, leaf_share_factor=cap,
|
|
|
|
|
|
max_share=cap, superpose=superpose, log=log, **search_kw)
|
|
|
|
|
|
else:
|
|
|
|
|
|
roots = [copy.deepcopy(ind.root) for ind in prev_pop]
|
|
|
|
|
|
created = sum(operators.unfold_shared_leaves(rt, above=cap) for rt in roots)
|
|
|
|
|
|
_log(f"[anneal] phase {i + 1}/{n_phases}: grain {cap}, budget "
|
|
|
|
|
|
f"{phase_budgets[i]} — unfolded {created} leaf-"
|
|
|
|
|
|
f"{'copy' if created == 1 else 'copies'} (share>{cap})")
|
|
|
|
|
|
r = search(
|
|
|
|
|
|
seed_root, programme_dir, budget=phase_budgets[i], pop_size=pop_size,
|
|
|
|
|
|
child_budget=child_budget, seed_budget=seed_budget,
|
|
|
|
|
|
p_crossover=p_crossover, seed=seed, types=types, inner_kw=inner_kw,
|
|
|
|
|
|
n_workers=n_workers, leaf_sharing=True, leaf_share_factor=cap,
|
|
|
|
|
|
max_share=cap, superpose=superpose, log=log, seed_pop=roots,
|
|
|
|
|
|
**search_kw)
|
|
|
|
|
|
combined = _stitch(combined, r, tag=f"g{cap}")
|
|
|
|
|
|
prev_pop = r.population
|
|
|
|
|
|
if r.interrupted:
|
|
|
|
|
|
break
|
|
|
|
|
|
|
|
|
|
|
|
if combined is None or combined.best is None:
|
|
|
|
|
|
return combined or SearchResult(
|
|
|
|
|
|
best=None, population=[], n_evals=0, n_topologies=0)
|
|
|
|
|
|
|
|
|
|
|
|
# Final honesty phase: de-share entirely. Unfold ALL remaining shared leaves
|
|
|
|
|
|
# and polish (or just rescore) under the canonical sharing-off objective, so
|
|
|
|
|
|
# the returned best is the honest canonical score (§15's guarantee).
|
|
|
|
|
|
if polish_budget > 0 and not combined.interrupted:
|
|
|
|
|
|
roots = [copy.deepcopy(ind.root) for ind in prev_pop]
|
|
|
|
|
|
created = sum(operators.unfold_shared_leaves(rt, above=1) for rt in roots)
|
|
|
|
|
|
_log(f"[anneal] finish: de-share (grain off), polish {polish_budget} "
|
|
|
|
|
|
f"evals — unfolded {created} leaf-"
|
|
|
|
|
|
f"{'copy' if created == 1 else 'copies'}")
|
|
|
|
|
|
r = search(
|
|
|
|
|
|
seed_root, programme_dir, budget=polish_budget, pop_size=pop_size,
|
|
|
|
|
|
child_budget=child_budget, seed_budget=seed_budget,
|
|
|
|
|
|
p_crossover=p_crossover, seed=seed, types=types, inner_kw=inner_kw,
|
|
|
|
|
|
n_workers=n_workers, leaf_sharing=False, superpose=superpose, log=log,
|
|
|
|
|
|
seed_pop=roots, **search_kw)
|
|
|
|
|
|
else:
|
|
|
|
|
|
best_root = copy.deepcopy(combined.best.root)
|
|
|
|
|
|
created = operators.unfold_shared_leaves(best_root, above=1)
|
|
|
|
|
|
_log(f"[anneal] finish: de-share (grain off), rescore only — unfolded "
|
|
|
|
|
|
f"{created} leaf-{'copy' if created == 1 else 'copies'}")
|
2026-08-05 07:52:44 +01:00
|
|
|
|
# homemaker-py-sd3: forward collapse_insearch/multi_use from the phase
|
|
|
|
|
|
# kwargs (same family as the collapse_best bug) — omitting them left
|
|
|
|
|
|
# this rescore silently defaulting to _evaluate's collapse_insearch=True
|
|
|
|
|
|
# even on a --no-collapse-insearch run, contradicting the run's own
|
|
|
|
|
|
# objective on interrupt/no-polish exits.
|
2026-07-16 08:38:08 +01:00
|
|
|
|
ind, used = _evaluate(
|
|
|
|
|
|
best_root, programme_dir, None, x0=None, budget=seed_budget,
|
2026-08-05 07:52:44 +01:00
|
|
|
|
inner_kw={}, lineage="unfold", leaf_sharing=False, superpose=superpose,
|
|
|
|
|
|
collapse_insearch=search_kw.get("collapse_insearch", True),
|
|
|
|
|
|
multi_use=search_kw.get("multi_use", False))
|
2026-07-16 08:38:08 +01:00
|
|
|
|
r = SearchResult(best=ind, population=[ind], n_evals=used, n_topologies=1)
|
|
|
|
|
|
r.n_distinct_signatures = 1
|
|
|
|
|
|
r.history = [(0, ind.fitness, ind.lineage)]
|
|
|
|
|
|
return _stitch(combined, r, tag="polish")
|
|
|
|
|
|
|
|
|
|
|
|
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
def search_staged(
|
|
|
|
|
|
seed_root: dom.Node,
|
|
|
|
|
|
programme_dir: str | Path,
|
|
|
|
|
|
budget: int = 20000,
|
|
|
|
|
|
pop_size: int = 16,
|
|
|
|
|
|
child_budget: int = 80,
|
|
|
|
|
|
seed_budget: int = 300,
|
|
|
|
|
|
stage1_frac: float = 0.4,
|
|
|
|
|
|
base_p: float = 0.15,
|
|
|
|
|
|
rank_bonus_weight: float = 1.0,
|
|
|
|
|
|
p_crossover: float = 0.2,
|
|
|
|
|
|
seed: int = 0,
|
|
|
|
|
|
types: list[str] | None = None,
|
|
|
|
|
|
inner_kw: dict | None = None,
|
|
|
|
|
|
log=None,
|
|
|
|
|
|
n_workers: int = 1,
|
2026-06-18 22:33:29 +01:00
|
|
|
|
use_grade: bool = False,
|
2026-06-29 22:56:23 +01:00
|
|
|
|
tournament_k: int = 2,
|
2026-06-18 23:42:39 +01:00
|
|
|
|
niche_by_signature: bool = False,
|
|
|
|
|
|
restart_patience: int | None = None,
|
|
|
|
|
|
restart_elite: int = 1,
|
2026-06-19 11:47:40 +01:00
|
|
|
|
seed_adjacency_aware: bool = True,
|
Phase 7 §12.2: proportion-aware constructive seeding + storey_minimum fix (leu.2, cq1)
Size each constructive-seed cut from leaf TARGET areas (division=[f,f] gives
left area-fraction f) and pick each cut's rotation for child squareness — both
derived from target dims, topology/type assignment untouched. Area-only
regressed (slivers); rotation choice is what makes it pay.
End-to-end (20000 evals, 3 seeds, staged): harbor 85.3->74.0 (-13%, best 69),
maple-court 151.7->136.0 (-10%, best 126). PROP=0 reproduces the §11.7/§12.1
baselines exactly. programme-house regresses at fixed budget (deeper local
optimum walls off the undivide restructuring path) but a budget sweep shows
it's convergence speed, not a worse asymptote (PROP=1 reaches 1 fail at 150k).
Default-on (seed_proportion_aware=True, env PROP=1).
cq1: n_storeys now honours storey_minimum, not just level: keys — programme-house
(storey_minimum:2, all rooms level:0) was seeded one storey short and fell
through to plain search. New programme.storey_minimum()/n_storeys_for();
driver.search passes min_storeys to the seeder; search_staged routes on the max.
No-op for harbor/maple; programme-house single-stage 8.0->5.0.
New maple-court best (126) saved as generated.dom. 204 tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 14:04:42 +01:00
|
|
|
|
seed_proportion_aware: bool = True,
|
2026-06-20 18:54:48 +01:00
|
|
|
|
enable_reassociate: bool = False,
|
2026-07-22 17:43:36 +01:00
|
|
|
|
enable_shape_repair: bool = False,
|
2026-07-24 19:48:13 +01:00
|
|
|
|
enable_bridge_circulation: bool = False,
|
2026-07-26 09:31:42 +01:00
|
|
|
|
enable_ruin_recreate: bool = False,
|
2026-06-20 18:54:48 +01:00
|
|
|
|
feasibility_filter: bool = False,
|
|
|
|
|
|
feasibility_max_shape_fails: int | None = None,
|
2026-06-21 21:10:18 +01:00
|
|
|
|
circ_divisor: int = 3,
|
2026-06-27 21:15:50 +01:00
|
|
|
|
leaf_sharing: bool = True,
|
|
|
|
|
|
leaf_share_factor: int = 3,
|
9o5: type superposition + per-eval collapse (multi-use leaves)
Interchangeable codes (similar size/width/proportion, compatible level/stack,
no adjacency edge) form equivalence classes derived from the programme. With
--superpose (default off), each fitness eval COLLAPSES every superposed leaf to
its best in-class usage via an optimal supply->demand assignment (brute force
<=C! within cap C=4, scipy Hungarian beyond), then scores the condensed types.
Because collapse re-types on the unmerged tree before all checks, counts /
adjacency / quality are unchanged downstream -- no Node field, no graph/operator
changes -- and default OFF is bit-identical.
- programme.py: derive_interchange_classes + interchangeable (S1-S4, locked
thresholds R_SIZE=1.5/R_WIDTH=1.3/R_PROP=1.5, CLASS_CAP=4)
- fitness.py: collapse_superposition, _best_assignment, _usage_quality;
superpose/superpose_class_cap conf knobs; collapse hooked into _evaluate_full
- driver.py/evolve.py: superpose flag plumbed beside leaf_sharing; --superpose
- tests/test_superposition.py: 17 tests (derivation, assignment, end-to-end)
Closes homemaker-py-9o5 (build); validation A/B is homemaker-py-xi7.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:08:46 +01:00
|
|
|
|
superpose: bool = False,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
multi_use: bool = False,
|
2026-06-27 21:15:50 +01:00
|
|
|
|
depth_balanced: bool = True,
|
2026-06-28 07:29:42 +01:00
|
|
|
|
interior_outside: bool = True,
|
2026-06-28 07:20:20 +01:00
|
|
|
|
outside_divisor: int = 3,
|
2026-07-28 00:00:51 +01:00
|
|
|
|
construction_beam_width: int = 1,
|
2026-08-04 09:19:36 +01:00
|
|
|
|
assign_solver: str = "greedy",
|
|
|
|
|
|
enable_reassign: bool = False,
|
Staged harness re-scored under a different objective than it searched
run_staged_search.py reported MISMATCH on its BASELINE arm -- the
LEAFSHARE=0/MULTIUSE=0 control every A/B compares against. Two facts
combined: driver.search_staged had no collapse_insearch parameter at all,
so every inner search() call inherited search()'s True default
unconditionally; and no example patterns.config sets the key, so the final
_native_score rescore got False from a bare load_config. Search optimised
one objective, the rescore graded another.
The 7ua fix pinned the key inside a fitness.load_config monkeypatch, but
that patch was installed only `if leaf_share or multi_use` -- so it fixed
every arm except the control.
Fixed in the right place: search_staged now HAS the parameter (default
True, byte-identical to the inherited default), threaded into all three
internal search() calls. The harness chooses the arm explicitly (COLLAPSE,
default 1), passes it to the search, and passes the SAME value to
_native_score, which overrides the key rather than hoping the config
carries it. The rescore mirrors the search by construction.
Verified on programme-house, budget 150:
baseline MISMATCH 1.56663e-08 vs 1.51708e-08 -> OK
COLLAPSE=0 (knob did not exist) -> OK 1.66216e-08
LEAFSHARE=1 / MULTIUSE=1 -> OK
COLLAPSE=0 scoring differently confirms the knob is not a no-op, and the
default arm's search result is unchanged, so no prior staged number moves.
Audited the other three search_staged callers: run_and_capture_91f.py
already pins collapse_insearch: True; run_island_ab.py never re-scores;
probe_harbor_floor.py did NOT pin it and had the same bug -- now fixed, and
that is the harness which produced every 13.x floor number.
The recorded mitigating factor -- only the continuous score moved, the fail
count matched, and the run_*_ab.sh greps read only the count -- is true and
is exactly what made it dangerous: a harness that reports MISMATCH on its
own control, invisibly to the metric of record, trains everyone to ignore
the warning.
Closes homemaker-py-4ok.
Lint at parity (46); tests 381 passed, 0 failed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-08-29 12:11:38 +00:00
|
|
|
|
collapse_insearch: bool = True,
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
) -> SearchResult:
|
|
|
|
|
|
"""Staged per-floor topology search (DESIGN.md §11.3, ``homemaker-py-c4c.3``).
|
|
|
|
|
|
|
|
|
|
|
|
Searches the genome in causal dependency order:
|
|
|
|
|
|
|
|
|
|
|
|
- **Stage 1** (``stage1_frac`` of the budget): a single-storey base over the
|
|
|
|
|
|
level-0 room set (a programme auto-derived to a tempdir), ranked with a
|
|
|
|
|
|
substrate-readiness bonus so the base is selected as a good *substrate* —
|
|
|
|
|
|
a reserved, vertically-alignable core and enough divisible footprint for the
|
|
|
|
|
|
upper floors — not merely a good ground floor (anti-bungalow, §4.2).
|
|
|
|
|
|
- **Stage 2** (remaining budget): the best base is lifted into a full
|
|
|
|
|
|
multi-storey design with each upper storey's required room set instantiated
|
|
|
|
|
|
by construction (``operators.lift_base_to_storeys``); the deltas are searched
|
|
|
|
|
|
with the base kept mutable at low probability (``base_p``).
|
|
|
|
|
|
|
|
|
|
|
|
Single-storey programmes (e.g. programme-house) have no upper floors to stage,
|
|
|
|
|
|
so this falls through to a plain :func:`search` — guaranteeing no regression.
|
|
|
|
|
|
"""
|
|
|
|
|
|
import shutil
|
|
|
|
|
|
import tempfile
|
|
|
|
|
|
|
|
|
|
|
|
from . import graph
|
|
|
|
|
|
|
|
|
|
|
|
reqs = programme.load_programme_dir(programme_dir)
|
Phase 7 §12.2: proportion-aware constructive seeding + storey_minimum fix (leu.2, cq1)
Size each constructive-seed cut from leaf TARGET areas (division=[f,f] gives
left area-fraction f) and pick each cut's rotation for child squareness — both
derived from target dims, topology/type assignment untouched. Area-only
regressed (slivers); rotation choice is what makes it pay.
End-to-end (20000 evals, 3 seeds, staged): harbor 85.3->74.0 (-13%, best 69),
maple-court 151.7->136.0 (-10%, best 126). PROP=0 reproduces the §11.7/§12.1
baselines exactly. programme-house regresses at fixed budget (deeper local
optimum walls off the undivide restructuring path) but a budget sweep shows
it's convergence speed, not a worse asymptote (PROP=1 reaches 1 fail at 150k).
Default-on (seed_proportion_aware=True, env PROP=1).
cq1: n_storeys now honours storey_minimum, not just level: keys — programme-house
(storey_minimum:2, all rooms level:0) was seeded one storey short and fell
through to plain search. New programme.storey_minimum()/n_storeys_for();
driver.search passes min_storeys to the seeder; search_staged routes on the max.
No-op for harbor/maple; programme-house single-stage 8.0->5.0.
New maple-court best (126) saved as generated.dom. 204 tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 14:04:42 +01:00
|
|
|
|
# Honour storey_minimum even when no room is pinned to an upper level (§12.2):
|
|
|
|
|
|
# e.g. programme-house is storey_minimum:2 with all rooms level:0, so its
|
|
|
|
|
|
# valid solutions are multi-storey and it must stage, not fall through.
|
|
|
|
|
|
n_storeys = max(programme.n_storeys_required(reqs),
|
|
|
|
|
|
programme.storey_minimum(programme_dir))
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
|
|
|
|
|
|
def _log(msg: str) -> None:
|
|
|
|
|
|
if log:
|
|
|
|
|
|
log(msg)
|
|
|
|
|
|
|
|
|
|
|
|
if n_storeys < 2:
|
|
|
|
|
|
_log("[staged] single-storey programme — falling back to plain search")
|
|
|
|
|
|
return search(seed_root, programme_dir, budget=budget, pop_size=pop_size,
|
|
|
|
|
|
child_budget=child_budget, seed_budget=seed_budget,
|
|
|
|
|
|
p_crossover=p_crossover, seed=seed, types=types,
|
2026-06-18 22:33:29 +01:00
|
|
|
|
inner_kw=inner_kw, log=log, n_workers=n_workers,
|
2026-06-29 22:56:23 +01:00
|
|
|
|
use_grade=use_grade, tournament_k=tournament_k,
|
|
|
|
|
|
niche_by_signature=niche_by_signature,
|
2026-06-19 11:47:40 +01:00
|
|
|
|
restart_patience=restart_patience, restart_elite=restart_elite,
|
Phase 7 §12.2: proportion-aware constructive seeding + storey_minimum fix (leu.2, cq1)
Size each constructive-seed cut from leaf TARGET areas (division=[f,f] gives
left area-fraction f) and pick each cut's rotation for child squareness — both
derived from target dims, topology/type assignment untouched. Area-only
regressed (slivers); rotation choice is what makes it pay.
End-to-end (20000 evals, 3 seeds, staged): harbor 85.3->74.0 (-13%, best 69),
maple-court 151.7->136.0 (-10%, best 126). PROP=0 reproduces the §11.7/§12.1
baselines exactly. programme-house regresses at fixed budget (deeper local
optimum walls off the undivide restructuring path) but a budget sweep shows
it's convergence speed, not a worse asymptote (PROP=1 reaches 1 fail at 150k).
Default-on (seed_proportion_aware=True, env PROP=1).
cq1: n_storeys now honours storey_minimum, not just level: keys — programme-house
(storey_minimum:2, all rooms level:0) was seeded one storey short and fell
through to plain search. New programme.storey_minimum()/n_storeys_for();
driver.search passes min_storeys to the seeder; search_staged routes on the max.
No-op for harbor/maple; programme-house single-stage 8.0->5.0.
New maple-court best (126) saved as generated.dom. 204 tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 14:04:42 +01:00
|
|
|
|
seed_adjacency_aware=seed_adjacency_aware,
|
2026-06-20 18:54:48 +01:00
|
|
|
|
seed_proportion_aware=seed_proportion_aware,
|
|
|
|
|
|
enable_reassociate=enable_reassociate,
|
2026-07-22 17:43:36 +01:00
|
|
|
|
enable_shape_repair=enable_shape_repair,
|
Staged harness re-scored under a different objective than it searched
run_staged_search.py reported MISMATCH on its BASELINE arm -- the
LEAFSHARE=0/MULTIUSE=0 control every A/B compares against. Two facts
combined: driver.search_staged had no collapse_insearch parameter at all,
so every inner search() call inherited search()'s True default
unconditionally; and no example patterns.config sets the key, so the final
_native_score rescore got False from a bare load_config. Search optimised
one objective, the rescore graded another.
The 7ua fix pinned the key inside a fitness.load_config monkeypatch, but
that patch was installed only `if leaf_share or multi_use` -- so it fixed
every arm except the control.
Fixed in the right place: search_staged now HAS the parameter (default
True, byte-identical to the inherited default), threaded into all three
internal search() calls. The harness chooses the arm explicitly (COLLAPSE,
default 1), passes it to the search, and passes the SAME value to
_native_score, which overrides the key rather than hoping the config
carries it. The rescore mirrors the search by construction.
Verified on programme-house, budget 150:
baseline MISMATCH 1.56663e-08 vs 1.51708e-08 -> OK
COLLAPSE=0 (knob did not exist) -> OK 1.66216e-08
LEAFSHARE=1 / MULTIUSE=1 -> OK
COLLAPSE=0 scoring differently confirms the knob is not a no-op, and the
default arm's search result is unchanged, so no prior staged number moves.
Audited the other three search_staged callers: run_and_capture_91f.py
already pins collapse_insearch: True; run_island_ab.py never re-scores;
probe_harbor_floor.py did NOT pin it and had the same bug -- now fixed, and
that is the harness which produced every 13.x floor number.
The recorded mitigating factor -- only the continuous score moved, the fail
count matched, and the run_*_ab.sh greps read only the count -- is true and
is exactly what made it dangerous: a harness that reports MISMATCH on its
own control, invisibly to the metric of record, trains everyone to ignore
the warning.
Closes homemaker-py-4ok.
Lint at parity (46); tests 381 passed, 0 failed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-08-29 12:11:38 +00:00
|
|
|
|
collapse_insearch=collapse_insearch,
|
2026-07-24 19:48:13 +01:00
|
|
|
|
enable_bridge_circulation=enable_bridge_circulation,
|
2026-07-26 09:31:42 +01:00
|
|
|
|
enable_ruin_recreate=enable_ruin_recreate,
|
2026-06-20 18:54:48 +01:00
|
|
|
|
feasibility_filter=feasibility_filter,
|
2026-06-21 21:10:18 +01:00
|
|
|
|
feasibility_max_shape_fails=feasibility_max_shape_fails,
|
2026-06-24 18:16:17 +01:00
|
|
|
|
circ_divisor=circ_divisor,
|
|
|
|
|
|
leaf_sharing=leaf_sharing,
|
2026-06-25 22:36:24 +01:00
|
|
|
|
leaf_share_factor=leaf_share_factor,
|
9o5: type superposition + per-eval collapse (multi-use leaves)
Interchangeable codes (similar size/width/proportion, compatible level/stack,
no adjacency edge) form equivalence classes derived from the programme. With
--superpose (default off), each fitness eval COLLAPSES every superposed leaf to
its best in-class usage via an optimal supply->demand assignment (brute force
<=C! within cap C=4, scipy Hungarian beyond), then scores the condensed types.
Because collapse re-types on the unmerged tree before all checks, counts /
adjacency / quality are unchanged downstream -- no Node field, no graph/operator
changes -- and default OFF is bit-identical.
- programme.py: derive_interchange_classes + interchangeable (S1-S4, locked
thresholds R_SIZE=1.5/R_WIDTH=1.3/R_PROP=1.5, CLASS_CAP=4)
- fitness.py: collapse_superposition, _best_assignment, _usage_quality;
superpose/superpose_class_cap conf knobs; collapse hooked into _evaluate_full
- driver.py/evolve.py: superpose flag plumbed beside leaf_sharing; --superpose
- tests/test_superposition.py: 17 tests (derivation, assignment, end-to-end)
Closes homemaker-py-9o5 (build); validation A/B is homemaker-py-xi7.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:08:46 +01:00
|
|
|
|
superpose=superpose,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
multi_use=multi_use,
|
2026-06-28 07:20:20 +01:00
|
|
|
|
depth_balanced=depth_balanced,
|
|
|
|
|
|
interior_outside=interior_outside,
|
2026-07-28 00:00:51 +01:00
|
|
|
|
outside_divisor=outside_divisor,
|
2026-08-04 09:19:36 +01:00
|
|
|
|
construction_beam_width=construction_beam_width,
|
|
|
|
|
|
assign_solver=assign_solver,
|
|
|
|
|
|
enable_reassign=enable_reassign)
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
|
|
|
|
|
|
if types is None:
|
|
|
|
|
|
types = sorted(reqs) + ["C", "O"]
|
|
|
|
|
|
rng = np.random.default_rng(seed)
|
|
|
|
|
|
buckets = programme.partition_rooms_by_storey(reqs, n_storeys, rng)
|
|
|
|
|
|
|
|
|
|
|
|
tmp = Path(tempfile.mkdtemp(prefix="homemaker_stage1_"))
|
|
|
|
|
|
try:
|
|
|
|
|
|
programme.write_stage1_programme(programme_dir, tmp, buckets[0])
|
|
|
|
|
|
|
|
|
|
|
|
# Stage 1 — single-storey base, readiness-biased ranking.
|
|
|
|
|
|
b1 = max(1, int(budget * stage1_frac))
|
|
|
|
|
|
_log(f"[staged] stage 1: base floor, budget {b1} "
|
|
|
|
|
|
f"(rooms {sum(buckets[0].values())}, +readiness bonus)")
|
|
|
|
|
|
r1 = search(
|
|
|
|
|
|
seed_root, tmp, budget=b1, pop_size=pop_size,
|
|
|
|
|
|
child_budget=child_budget, seed_budget=seed_budget,
|
|
|
|
|
|
p_crossover=p_crossover, seed=seed, types=None,
|
|
|
|
|
|
inner_kw=inner_kw, log=log, n_workers=n_workers,
|
Staged harness re-scored under a different objective than it searched
run_staged_search.py reported MISMATCH on its BASELINE arm -- the
LEAFSHARE=0/MULTIUSE=0 control every A/B compares against. Two facts
combined: driver.search_staged had no collapse_insearch parameter at all,
so every inner search() call inherited search()'s True default
unconditionally; and no example patterns.config sets the key, so the final
_native_score rescore got False from a bare load_config. Search optimised
one objective, the rescore graded another.
The 7ua fix pinned the key inside a fitness.load_config monkeypatch, but
that patch was installed only `if leaf_share or multi_use` -- so it fixed
every arm except the control.
Fixed in the right place: search_staged now HAS the parameter (default
True, byte-identical to the inherited default), threaded into all three
internal search() calls. The harness chooses the arm explicitly (COLLAPSE,
default 1), passes it to the search, and passes the SAME value to
_native_score, which overrides the key rather than hoping the config
carries it. The rescore mirrors the search by construction.
Verified on programme-house, budget 150:
baseline MISMATCH 1.56663e-08 vs 1.51708e-08 -> OK
COLLAPSE=0 (knob did not exist) -> OK 1.66216e-08
LEAFSHARE=1 / MULTIUSE=1 -> OK
COLLAPSE=0 scoring differently confirms the knob is not a no-op, and the
default arm's search result is unchanged, so no prior staged number moves.
Audited the other three search_staged callers: run_and_capture_91f.py
already pins collapse_insearch: True; run_island_ab.py never re-scores;
probe_harbor_floor.py did NOT pin it and had the same bug -- now fixed, and
that is the harness which produced every 13.x floor number.
The recorded mitigating factor -- only the continuous score moved, the fail
count matched, and the run_*_ab.sh greps read only the count -- is true and
is exactly what made it dangerous: a harness that reports MISMATCH on its
own control, invisibly to the metric of record, trains everyone to ignore
the warning.
Closes homemaker-py-4ok.
Lint at parity (46); tests 381 passed, 0 failed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-08-29 12:11:38 +00:00
|
|
|
|
collapse_insearch=collapse_insearch,
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
rank_bonus_fn=lambda root: graph.substrate_readiness(root, reqs, n_storeys),
|
|
|
|
|
|
rank_bonus_weight=rank_bonus_weight,
|
2026-06-29 22:56:23 +01:00
|
|
|
|
tournament_k=tournament_k,
|
2026-06-18 23:42:39 +01:00
|
|
|
|
niche_by_signature=niche_by_signature,
|
|
|
|
|
|
restart_patience=restart_patience, restart_elite=restart_elite,
|
2026-06-19 11:47:40 +01:00
|
|
|
|
seed_adjacency_aware=seed_adjacency_aware,
|
Phase 7 §12.2: proportion-aware constructive seeding + storey_minimum fix (leu.2, cq1)
Size each constructive-seed cut from leaf TARGET areas (division=[f,f] gives
left area-fraction f) and pick each cut's rotation for child squareness — both
derived from target dims, topology/type assignment untouched. Area-only
regressed (slivers); rotation choice is what makes it pay.
End-to-end (20000 evals, 3 seeds, staged): harbor 85.3->74.0 (-13%, best 69),
maple-court 151.7->136.0 (-10%, best 126). PROP=0 reproduces the §11.7/§12.1
baselines exactly. programme-house regresses at fixed budget (deeper local
optimum walls off the undivide restructuring path) but a budget sweep shows
it's convergence speed, not a worse asymptote (PROP=1 reaches 1 fail at 150k).
Default-on (seed_proportion_aware=True, env PROP=1).
cq1: n_storeys now honours storey_minimum, not just level: keys — programme-house
(storey_minimum:2, all rooms level:0) was seeded one storey short and fell
through to plain search. New programme.storey_minimum()/n_storeys_for();
driver.search passes min_storeys to the seeder; search_staged routes on the max.
No-op for harbor/maple; programme-house single-stage 8.0->5.0.
New maple-court best (126) saved as generated.dom. 204 tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 14:04:42 +01:00
|
|
|
|
seed_proportion_aware=seed_proportion_aware,
|
2026-06-20 18:54:48 +01:00
|
|
|
|
enable_reassociate=enable_reassociate,
|
2026-07-22 17:43:36 +01:00
|
|
|
|
enable_shape_repair=enable_shape_repair,
|
2026-07-24 19:48:13 +01:00
|
|
|
|
enable_bridge_circulation=enable_bridge_circulation,
|
2026-07-26 09:31:42 +01:00
|
|
|
|
enable_ruin_recreate=enable_ruin_recreate,
|
2026-06-20 18:54:48 +01:00
|
|
|
|
feasibility_filter=feasibility_filter,
|
|
|
|
|
|
feasibility_max_shape_fails=feasibility_max_shape_fails,
|
2026-06-21 21:10:18 +01:00
|
|
|
|
circ_divisor=circ_divisor,
|
2026-06-24 18:16:17 +01:00
|
|
|
|
leaf_sharing=leaf_sharing,
|
|
|
|
|
|
leaf_share_factor=leaf_share_factor,
|
9o5: type superposition + per-eval collapse (multi-use leaves)
Interchangeable codes (similar size/width/proportion, compatible level/stack,
no adjacency edge) form equivalence classes derived from the programme. With
--superpose (default off), each fitness eval COLLAPSES every superposed leaf to
its best in-class usage via an optimal supply->demand assignment (brute force
<=C! within cap C=4, scipy Hungarian beyond), then scores the condensed types.
Because collapse re-types on the unmerged tree before all checks, counts /
adjacency / quality are unchanged downstream -- no Node field, no graph/operator
changes -- and default OFF is bit-identical.
- programme.py: derive_interchange_classes + interchangeable (S1-S4, locked
thresholds R_SIZE=1.5/R_WIDTH=1.3/R_PROP=1.5, CLASS_CAP=4)
- fitness.py: collapse_superposition, _best_assignment, _usage_quality;
superpose/superpose_class_cap conf knobs; collapse hooked into _evaluate_full
- driver.py/evolve.py: superpose flag plumbed beside leaf_sharing; --superpose
- tests/test_superposition.py: 17 tests (derivation, assignment, end-to-end)
Closes homemaker-py-9o5 (build); validation A/B is homemaker-py-xi7.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:08:46 +01:00
|
|
|
|
superpose=superpose,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
multi_use=multi_use,
|
2026-06-25 22:36:24 +01:00
|
|
|
|
depth_balanced=depth_balanced,
|
2026-06-28 07:20:20 +01:00
|
|
|
|
interior_outside=interior_outside,
|
|
|
|
|
|
outside_divisor=outside_divisor,
|
2026-07-28 00:00:51 +01:00
|
|
|
|
construction_beam_width=construction_beam_width,
|
2026-08-04 09:19:36 +01:00
|
|
|
|
assign_solver=assign_solver,
|
|
|
|
|
|
enable_reassign=enable_reassign,
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
)
|
|
|
|
|
|
best_base = r1.best.root
|
|
|
|
|
|
_log(f"[staged] stage 1 done: base {r1.best.fitness:.6g} "
|
|
|
|
|
|
f"({r1.best.n_fails} fails), readiness "
|
|
|
|
|
|
f"{graph.substrate_readiness(best_base, reqs, n_storeys):.3f}")
|
|
|
|
|
|
finally:
|
|
|
|
|
|
shutil.rmtree(tmp, ignore_errors=True)
|
|
|
|
|
|
|
|
|
|
|
|
# Stage 2 — lift base into full multi-storey, search deltas, base low-prob.
|
|
|
|
|
|
b2 = max(1, budget - r1.n_evals)
|
|
|
|
|
|
upper = buckets[1:]
|
|
|
|
|
|
|
|
|
|
|
|
def _seed_factory(rng2):
|
2026-06-19 11:47:40 +01:00
|
|
|
|
return operators.lift_base_to_storeys(
|
|
|
|
|
|
best_base, upper, rng2, types, reqs=reqs,
|
Phase 7 §12.2: proportion-aware constructive seeding + storey_minimum fix (leu.2, cq1)
Size each constructive-seed cut from leaf TARGET areas (division=[f,f] gives
left area-fraction f) and pick each cut's rotation for child squareness — both
derived from target dims, topology/type assignment untouched. Area-only
regressed (slivers); rotation choice is what makes it pay.
End-to-end (20000 evals, 3 seeds, staged): harbor 85.3->74.0 (-13%, best 69),
maple-court 151.7->136.0 (-10%, best 126). PROP=0 reproduces the §11.7/§12.1
baselines exactly. programme-house regresses at fixed budget (deeper local
optimum walls off the undivide restructuring path) but a budget sweep shows
it's convergence speed, not a worse asymptote (PROP=1 reaches 1 fail at 150k).
Default-on (seed_proportion_aware=True, env PROP=1).
cq1: n_storeys now honours storey_minimum, not just level: keys — programme-house
(storey_minimum:2, all rooms level:0) was seeded one storey short and fell
through to plain search. New programme.storey_minimum()/n_storeys_for();
driver.search passes min_storeys to the seeder; search_staged routes on the max.
No-op for harbor/maple; programme-house single-stage 8.0->5.0.
New maple-court best (126) saved as generated.dom. 204 tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 14:04:42 +01:00
|
|
|
|
adjacency_aware=seed_adjacency_aware,
|
2026-06-21 21:10:18 +01:00
|
|
|
|
proportion_aware=seed_proportion_aware,
|
2026-06-24 18:16:17 +01:00
|
|
|
|
circ_divisor=circ_divisor,
|
2026-06-25 22:36:24 +01:00
|
|
|
|
leaf_sharing=leaf_sharing, leaf_share_factor=leaf_share_factor,
|
2026-06-28 07:20:20 +01:00
|
|
|
|
depth_balanced=depth_balanced,
|
2026-07-28 00:00:51 +01:00
|
|
|
|
interior_outside=interior_outside, outside_divisor=outside_divisor,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
construction_beam_width=construction_beam_width,
|
2026-08-04 09:19:36 +01:00
|
|
|
|
multi_use=multi_use, assign_solver=assign_solver)
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
|
|
|
|
|
|
_log(f"[staged] stage 2: upper floors as deltas, budget {b2}, base_p {base_p}")
|
|
|
|
|
|
r2 = search(
|
|
|
|
|
|
best_base, programme_dir, budget=b2, pop_size=pop_size,
|
|
|
|
|
|
child_budget=child_budget, seed_budget=seed_budget,
|
|
|
|
|
|
p_crossover=p_crossover, seed=seed, types=types,
|
|
|
|
|
|
inner_kw=inner_kw, log=log, n_workers=n_workers,
|
Staged harness re-scored under a different objective than it searched
run_staged_search.py reported MISMATCH on its BASELINE arm -- the
LEAFSHARE=0/MULTIUSE=0 control every A/B compares against. Two facts
combined: driver.search_staged had no collapse_insearch parameter at all,
so every inner search() call inherited search()'s True default
unconditionally; and no example patterns.config sets the key, so the final
_native_score rescore got False from a bare load_config. Search optimised
one objective, the rescore graded another.
The 7ua fix pinned the key inside a fitness.load_config monkeypatch, but
that patch was installed only `if leaf_share or multi_use` -- so it fixed
every arm except the control.
Fixed in the right place: search_staged now HAS the parameter (default
True, byte-identical to the inherited default), threaded into all three
internal search() calls. The harness chooses the arm explicitly (COLLAPSE,
default 1), passes it to the search, and passes the SAME value to
_native_score, which overrides the key rather than hoping the config
carries it. The rescore mirrors the search by construction.
Verified on programme-house, budget 150:
baseline MISMATCH 1.56663e-08 vs 1.51708e-08 -> OK
COLLAPSE=0 (knob did not exist) -> OK 1.66216e-08
LEAFSHARE=1 / MULTIUSE=1 -> OK
COLLAPSE=0 scoring differently confirms the knob is not a no-op, and the
default arm's search result is unchanged, so no prior staged number moves.
Audited the other three search_staged callers: run_and_capture_91f.py
already pins collapse_insearch: True; run_island_ab.py never re-scores;
probe_harbor_floor.py did NOT pin it and had the same bug -- now fixed, and
that is the harness which produced every 13.x floor number.
The recorded mitigating factor -- only the continuous score moved, the fail
count matched, and the run_*_ab.sh greps read only the count -- is true and
is exactly what made it dangerous: a harness that reports MISMATCH on its
own control, invisibly to the metric of record, trains everyone to ignore
the warning.
Closes homemaker-py-4ok.
Lint at parity (46); tests 381 passed, 0 failed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-08-29 12:11:38 +00:00
|
|
|
|
collapse_insearch=collapse_insearch,
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
bootstrap=True, seed_factory=_seed_factory, base_p=base_p,
|
2026-06-18 22:33:29 +01:00
|
|
|
|
# §11.4: the graded objective targets the dense two-floor quality-fail
|
|
|
|
|
|
# regime, which is Stage 2. Stage 1 keeps its readiness-biased key so the
|
|
|
|
|
|
# substrate-selection semantics (§11.3) are unchanged.
|
2026-06-29 22:56:23 +01:00
|
|
|
|
use_grade=use_grade, tournament_k=tournament_k,
|
|
|
|
|
|
niche_by_signature=niche_by_signature,
|
2026-06-18 23:42:39 +01:00
|
|
|
|
restart_patience=restart_patience, restart_elite=restart_elite,
|
2026-06-20 18:54:48 +01:00
|
|
|
|
enable_reassociate=enable_reassociate,
|
2026-07-22 17:43:36 +01:00
|
|
|
|
enable_shape_repair=enable_shape_repair,
|
2026-07-24 19:48:13 +01:00
|
|
|
|
enable_bridge_circulation=enable_bridge_circulation,
|
2026-07-26 09:31:42 +01:00
|
|
|
|
enable_ruin_recreate=enable_ruin_recreate,
|
2026-06-20 18:54:48 +01:00
|
|
|
|
feasibility_filter=feasibility_filter,
|
|
|
|
|
|
feasibility_max_shape_fails=feasibility_max_shape_fails,
|
2026-06-21 21:10:18 +01:00
|
|
|
|
circ_divisor=circ_divisor,
|
2026-06-24 18:16:17 +01:00
|
|
|
|
leaf_sharing=leaf_sharing,
|
|
|
|
|
|
leaf_share_factor=leaf_share_factor,
|
9o5: type superposition + per-eval collapse (multi-use leaves)
Interchangeable codes (similar size/width/proportion, compatible level/stack,
no adjacency edge) form equivalence classes derived from the programme. With
--superpose (default off), each fitness eval COLLAPSES every superposed leaf to
its best in-class usage via an optimal supply->demand assignment (brute force
<=C! within cap C=4, scipy Hungarian beyond), then scores the condensed types.
Because collapse re-types on the unmerged tree before all checks, counts /
adjacency / quality are unchanged downstream -- no Node field, no graph/operator
changes -- and default OFF is bit-identical.
- programme.py: derive_interchange_classes + interchangeable (S1-S4, locked
thresholds R_SIZE=1.5/R_WIDTH=1.3/R_PROP=1.5, CLASS_CAP=4)
- fitness.py: collapse_superposition, _best_assignment, _usage_quality;
superpose/superpose_class_cap conf knobs; collapse hooked into _evaluate_full
- driver.py/evolve.py: superpose flag plumbed beside leaf_sharing; --superpose
- tests/test_superposition.py: 17 tests (derivation, assignment, end-to-end)
Closes homemaker-py-9o5 (build); validation A/B is homemaker-py-xi7.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:08:46 +01:00
|
|
|
|
superpose=superpose,
|
2026-07-31 00:16:12 +01:00
|
|
|
|
multi_use=multi_use,
|
2026-06-25 22:36:24 +01:00
|
|
|
|
depth_balanced=depth_balanced,
|
2026-06-28 07:20:20 +01:00
|
|
|
|
interior_outside=interior_outside,
|
|
|
|
|
|
outside_divisor=outside_divisor,
|
2026-07-28 00:00:51 +01:00
|
|
|
|
construction_beam_width=construction_beam_width,
|
2026-08-04 09:19:36 +01:00
|
|
|
|
assign_solver=assign_solver,
|
|
|
|
|
|
enable_reassign=enable_reassign,
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
# Stitch the two stages into one accounting (total evals, tagged history).
|
|
|
|
|
|
r2.n_evals += r1.n_evals
|
|
|
|
|
|
r2.n_topologies += r1.n_topologies
|
2026-06-18 23:42:39 +01:00
|
|
|
|
r2.n_distinct_signatures += r1.n_distinct_signatures
|
|
|
|
|
|
r2.n_restarts += r1.n_restarts
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
r2.history = (
|
|
|
|
|
|
[(e, f, f"S1:{lin}") for e, f, lin in r1.history]
|
|
|
|
|
|
+ [(e + r1.n_evals, f, f"S2:{lin}") for e, f, lin in r2.history]
|
|
|
|
|
|
)
|
2026-06-18 23:42:39 +01:00
|
|
|
|
r2.diversity_history = (
|
|
|
|
|
|
[(e, d, c) for e, d, c in r1.diversity_history]
|
|
|
|
|
|
+ [(e + r1.n_evals, d, c) for e, d, c in r2.diversity_history]
|
|
|
|
|
|
)
|
Phase 6 §11.3: staged per-floor search (c4c.3)
Search the genome in causal dependency order. Stage 1 evolves a single-storey
base over the level-0 room set (programme auto-derived to a tempdir), ranked
with a substrate-readiness bonus (reserved core × divisible capacity) so the
base is selected as a good substrate, not just a good ground floor (anti-§4.2).
Stage 2 lifts the best base into a full multi-storey design — preserving the
inherited core, instantiating each upper storey's required set by construction —
and searches the deltas with the base mutable at low probability (base_p=0.15).
New: programme.{n_storeys_required,partition_rooms_by_storey,write_stage1_programme},
graph.substrate_readiness, operators.{lift_base_to_storeys,_pick_weighted_by_storey},
base_p threading, driver.search rank_bonus_fn/seed_factory/base_p hooks +
search_staged orchestrator, experiments/run_staged_search.py, tests/test_staging.py.
Result (harbor, 20000 evals, seed 0): staged 95 fails vs single-stage 105
(-10, -9.5%), gain in crinkliness 27->18 + edge 12->8. Anti-bungalow confirmed
(Stage-2 core moves all noop — core inherited, not carved). Programme-house
regression PASS (warmstart-2f4 still reaches whole-pop 1-fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 06:05:53 +01:00
|
|
|
|
return r2
|