Commit graph

64 commits

Author SHA1 Message Date
eec075c4a7 h10: confirm §12.3 A/B was already at fixed worker count, no re-run needed
run_9gp_ab.sh never threaded a worker count through run_staged_search.py, so
every §12.3 arm ran at n_workers=1 (serial) — the one mode §12.4 already
proved byte-for-byte reproducible even before the completion-order
determinism fix (that bug is ProcessPoolExecutor as_completed-only).
Spot-checked empirically: same config run twice gave identical fail counts
at every checkpoint. Closes homemaker-py-h10 as confirmed-null without
re-spending the ~8 core-hours a full sweep re-run would cost.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY
2026-07-30 09:51:27 +01:00
391f5108ef homemaker-py-9yx: health-centre sweep resolves to clean null (DESIGN.md §32)
N=15 seeds, xyu's own ruin_recreate ON/OFF protocol, against a real 20-room
diverse programme instead of programme-house's duplicated-code sweep: 2.2%
mean-fails delta, p=0.40 two-sided — much weaker than xyu's own inconclusive
6.4%/p=0.059 reading at the same room count, and converging with
harbor-house's null-to-negative result instead. Closes the diversity-axis
gap xyu's larger-N pass could not reach; enable_ruin_recreate stays OFF on a
now-broader evidence base. Issue closed.
2026-07-30 08:08:05 +01:00
e5b7bc810e xyu: n=18 ruin_recreate N=15 confirmation — weakened, not resolved
Extends y51's n=18 synthetic sweep (strongest of four sizes at N=10) to
N=15 seeds, matching f1d's own confirmation sample size. Effect shrank
(9.3%->6.4%, two-sided Wilcoxon p 0.098->0.059) but didn't evaporate or
reverse — an ambiguous middle case, not a clean confirm or null. Refiled
option (b) (non-synthetic third example programme) as homemaker-py-9yx
since extending N alone doesn't address the interchangeable-room-code
confound §24 already flagged. enable_ruin_recreate stays default OFF.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY
2026-07-29 10:10:23 +01:00
15abba9679 e01: harbor-house beam-width larger-N confirmation (null)
N=15 driver.search sweep of construction_beam_width 1 vs 4 (protocol
identical to c94's original 5-seed run, DESIGN.md §29): 6W/4L/5T, mean
fails 57.0->56.6, Wilcoxon p=0.84. Excluding seed 2's outlier the mean
flips slightly negative (56.3->56.9), confirming the §29 5-seed
"improvement" was that one outlier. construction_beam_width stays
default 1 on confirmed rather than precautionary grounds. DESIGN.md §30.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY
2026-07-29 00:22:55 +01:00
5b9545a314 docs: DESIGN.md §29 correction — c94 end-to-end run overturns raw-seed-only claim
The raw-seed diagnostic (single constructed tree, byte-identical across
beam widths) was wrongly taken as proof a full driver.search run would
also be byte-identical. Actually running it (5 seeds/programme, budget
1500, n_workers=1) shows harbor-house diverges: 2 wins/1 loss/2 ties vs
greedy — a full bootstrap population hits beam-vs-greedy tie-breaks a
lone raw seed sample missed. programme-house stayed tied 5/5. Verdict
corrected from a confident null to inconclusive/mixed on harbor-house;
practical disposition unchanged (construction_beam_width stays default 1).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY
2026-07-28 09:38:23 +01:00
6d719e03ab c94: beam/best-first search over adjacency-aware room placement (null)
operators._assign_adjacency_aware gains beam_width (default 1 = exact
prior greedy behaviour), threaded through constructive_topology/
lift_base_to_storeys/driver.search/search_staged as
construction_beam_width. Verified functioning on an adversarial
synthetic case, but byte-identical raw-seed output to greedy at every
width tested (1/4/8/20) on both example programmes -- no headroom for
the beam to find on this repo's programmes. DESIGN.md section 29.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY
2026-07-28 00:00:51 +01:00
a04fb2bc8e docs: DESIGN.md §28 — backfill cdl (finish-time local_search default)
§25 still said the evolve.py wiring and broader sweep were pending on
homemaker-py-cdl; that closed in the previous commit, so update §25's
stale forward-references and add §28 with the 46-file sweep results and
the collapse_insearch hot-path reasoning for leaving collapse_global's
own default off.
2026-07-27 21:26:01 +01:00
9c6b1552eb docs: DESIGN.md §26/§27 — backfill 9o5/xi7/b3v (type superposition) and mi7 (bubble-diagram signal)
Two closed, substantive experiments were missing their DESIGN.md write-up
despite being referenced as prior art by later sections:

- 9o5/xi7/b3v (closed 2026-06-30/07-17): multi-use-leaf type superposition,
  a full feature build + real A/B validation (negative — OFF beats ON on
  both programme-house and harbor-house) + a veto-hatch follow-up for the
  one genuine false-positive interchange class found. §17 and §20 both cite
  its verdict directly but it never got its own section.

- mi7 (closed 2026-07-25): 3D bubble-diagram / topological-hop-distance
  fitness signal prototype, tested against real evolved trajectories on two
  programmes, both formulations null. bubble.py was left in the tree
  uncommitted "as documented reference" by the closing session -- committing
  it now (with two trivial ruff fixes: unused import, ambiguous var name) so
  the reference this write-up makes to it is actually resolvable, plus a
  CLAUDE.md module-list entry.

Numbered §26/§27 (appended, not inserted chronologically) to avoid
renumbering every cross-reference in §14-§25.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 23:16:21 +01:00
5c8a5a5e09 docs: DESIGN.md §25 — 2-opt local search past collapse_global Jacobi plateau (9wi)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 21:15:21 +01:00
bc59849666 y51: synthetic room-count sweep finds no clean ruin_recreate size threshold
Locates the threshold f1d's programme-house/harbor-house split implied,
using four synthetic sizes (10/14/18/22 rooms) derived from programme-house
by scaling its bedroom+ensuite module count, since no natural third example
programme sits between the two. Results are noisy and non-monotonic (n=10
mild win, n=14 clean null, n=18 strongest trend at p=0.098, n=22 near-null)
rather than a clean decay with room count -- documented in DESIGN.md #24.
enable_ruin_recreate stays default OFF; filed homemaker-py-xyu as a
low-priority follow-up (larger-N at n=18, or a non-synthetic third example).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 15:35:03 +01:00
0d94e58119 f1d: ruin-and-recreate LNS operator, validated positive on programme-house
Adds operators.mutate_ruin_recreate: un-divides one wing of a storey and
rebuilds it with the same adjacency-aware constructor the seeders use
(_assign_adjacency_aware, generalised with a new `scope` param), seeded
from the surviving circulation bordering the wing. Gated behind
enable_ruin_recreate (default off) / --ruin-recreate, same pattern as
reassociate/bridge_circulation.

A/B (qpk protocol, DESIGN.md §23): initial uniform-weight run was
underpowered (fired ~1/32 children), null. A weight=3.0 follow-up
(_MUTATION_WEIGHTS["ruin_recreate"]) showed a statistically significant
win on programme-house across 15 seeds (8W/1L/6T, mean fails 7.07->6.00,
Wilcoxon p=0.041) but no consistent effect on harbor-house across 8 seeds
(3W/2L/3T). Kept default off pending a size-threshold follow-up.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 09:31:42 +01:00
1e2fa2adfc lj3/qjg: larger-N + weight sweep finds bridge_circulation effect is noise
Combined follow-up to 8sh (DESIGN.md §22): raised bridge_circulation's
_MUTATION_WEIGHTS entry to 2.0 (lj3) and re-ran the qi6/qpk-protocol A/B
at 4x the sample size (qjg) in one sweep, since the two variables were
confounded if tested separately. Result is null in the opposite direction
from 8sh's small-N signal -- no total-fail benefit (p=0.71 programme-house
N=20, p=0.69 harbor-house N=12) and a higher rate of trajectory-divergence
-induced new not-connected fails than at the original uniform weight.
Reverted the weight bump; enable_bridge_circulation stays default off.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-25 12:20:22 +01:00
a103e0114a 8sh: insert/relocate-circulation repair operator (qi6 mechanism (a))
Adds operators.mutate_bridge_circulation: retypes the cheapest path between
two disconnected circulation components to circulation, directly clearing a
'level N not connected' fail instead of relying on the qi6 graded comparator
key (measured negative, DESIGN.md §18). Gated off by default via
driver.search's enable_bridge_circulation flag and evolve.py
--bridge-circulation, mirroring enable_reassociate's clean-toggle pattern.

qi6/qpk-protocol A/B (DESIGN.md §21) is directionally positive but mixed at
N=3/N=5 (never worse on total fails; clears 2/5 baseline not-connected fails
vs qi6's 0/4; one seed's RNG-trajectory divergence adds 2 new not-connected
fails) — kept default off pending a larger-N confirmation sweep
(homemaker-py-qjg) and a mutation-weight bump experiment (homemaker-py-lj3).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-24 19:48:13 +01:00
a78409df90 1ph: larger-N seed sweep confirms collapse_insearch positive, flip default ON
20-seed programme-house sweep (vs the original 5) resolves the qpk A/B's
mixed 3/5 result as small-sample noise around a true small positive: mean
fails 7.95->7.10 (~10.7%), 11W/6L/3T, paired t-test p~0.028. Flips
collapse_insearch's default from OFF to ON in evolve.py and driver.py
(_overrides_for/_fitness_for/_evaluate/search/polish_finish); opt out with
--no-collapse-insearch. fitness.Fitness itself is unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-24 09:55:34 +01:00
372bebd0a3 DESIGN.md: backfill §19 with homemaker-py-161's in-search A/B verdict
161 (2026-07-22) answered the "remaining open question" §19 left dangling —
threading fit into driver.search so shape_rotate/deslim can fire mid-GA —
but the result only ever landed in bd notes, never here. Also negative:
full-budget harbor-house A/B (seeds 0-3) shows no improvement, confirming
the finish-time finding at in-search scale. Both halves of §19's mechanism
space are now closed negative in the doc, matching bd state.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-23 22:39:08 +01:00
8328ac1b69 qi6: full-budget A/B for graded circulation-connectivity signal, negative result
conn_grade ON vs OFF (qpk protocol, experiments/run_qi6_ab.sh): harbor-house
(budget 2500, seeds 1-3) byte-identical output in every seed — the secondary
comparator key never fired. programme-house (budget 3000, seeds 1-5) 3/5 seeds
tie exactly; seeds 1/2 diverge to a different topology but the fail delta is
adjacency/crinkliness/width/access/size, never connectivity. Zero of 4 cases
where a not-connected fail was present got cleared by the grade.

Mechanism (b)/(c) (graded proximity as tertiary comparator key) is falsified,
not just unconfirmed. Kept default OFF (already was). Closed qi6; filed
homemaker-py-8sh for the remaining candidate (mechanism (a): an explicit
insert/relocate-circulation operator that doesn't depend on the search
stumbling onto a fail-count tie).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-23 18:29:45 +01:00
2b7a7d2926 qpk: in-search global collapse — run collapse_global per-eval during search
Runs the 94g finish-time cell↔room collapse inside every fitness eval
(collapse_insearch conf flag, default off, bit-identical when off) instead
of once at the end, so search optimises the collapsed objective directly.
Plumbed through fitness.py/driver.py/evolve.py the same way superpose/
conn_grade are; --collapse-insearch CLI flag.

A/B validated against the xi7 protocol (equal budget, both arms finished
with standard finish-time --collapse): POSITIVE, opposite of the 9o5/xi7
prior. harbor-house ON wins 3/3 (mean fails 80.3->72.0); programme-house
mixed 3/5 (mean fails 8.4->7.8). Kept default off pending a larger
programme-house sample; documented as a working opt-in for harbor-house-
scale-or-larger programmes. Full writeup in DESIGN.md §20.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-19 20:35:18 +01:00
07a4739576 7fm: targeted shape-repair operators (shape_rotate/deslim), negative finish-time result
Diagnosed the geometry-intrinsic residual from 94g's collapse: ratio
re-optimisation isn't the bottleneck (1500-eval NM makes zero difference on
the 12-fail collapsed best layout); the causes are upstream area starvation
and cut-orientation mismatch. Added mutate_shape_rotate/mutate_deslim
targeting each, gated on a Fitness instance like the existing reqs-gated
repair ops.

Evaluated as a finish-time exhaustive hill-climb on the same 6-layout
harbor-house sweep 94g used: zero improving moves found anywhere — every
candidate move traded the shape fail for a new adjacency/access fail on the
co-evolved layout (§4.2's lesson, now confirmed for topology repair). Closes
homemaker-py-7fm; spun homemaker-py-161 for the open in-search-GA question.

See DESIGN.md §19 for the full writeup.
2026-07-19 11:06:56 +01:00
94d4223a55 qi6: graded circulation-connectivity signal (§18)
The dominant post-collapse fail is the binary "level N not connected",
which is flat across fragmentation (a 7-component storey scores the same
as a 2-component one), so the outer search has no gradient toward
connected circulation. A finish-time convert-to-circulation repair was
prototyped and measured NEGATIVE (195->560 fails: bridging needed rooms
costs more missing-room fails than the one binary fail it clears).

Instead add graph.circulation_connectivity(G) = largest-circ-component
fraction, summed over storeys onto the score_with_grade proximity channel
(conf flag conn_grade; replaces the §11.4 leaf-grade there). It is a
secondary comparator key only — scalar fitness and fail count stay
byte-identical — restoring the gradient the binary fail lacks. Threaded
through driver (_overrides_for/_fitness_for/_evaluate/search; enabling it
implies the grade key) and evolve --conn-grade (default off).

A/B on full-budget runs pending; short smoke run confirms plumbing.

Tests: tests/test_conn_grade.py x9 (fraction contract, non-circ ignored,
monotone under (dis)connection, score/fail invariance); 276 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 18:44:24 +01:00
1ae8faac7c docs: DESIGN.md §17 — finish-time global cell→room collapse (94g)
Document the collapse as built (default-on in evolve + homemaker-collapse CLI):
label-relative vs geometry-intrinsic fails, collapse_global mechanism (c/o/s
partition, hard level, adjacency relaxation, threshold objective, public-access
pin), the two measurement corrections, keep-better wrapper + wiring, and the
6-layout sweep verification. DESIGN.md is the system-of-record for users without
beads access.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 12:45:35 +01:00
c980100589 kpu: Schedule B A/B DONE — graduated grain ramp falsified (negative)
harbor-house 3M (500k/grain x3 + 1.5M polish, workers 4, ~22h): the in-run
grain anneal reached 1.26e-08 / 23 fails (canonical byte-for-byte), losing
decisively to both the direct --no-leaf-sharing baseline (5.14e-06 / 15) and
yaa's single-hard-transition warm chain (4.19e-06 / 15) — ~400x worse fitness,
+8 fails.

Each grain step spikes the fail count as its unfolded leaves acquire
independent shape fails (phase-end 19->21->27, final de-share 27->36); the
per-phase budget re-polishes a partially-materialised state the next step
materialises further, so coarse-grain gains do not carry forward. The polish
phase started from a deeper hole (36) than the warm chain's single clean
transition and 1.5M evals recovered only to 23. The sharing-phase topology
skeleton is best cashed in once, at full grain — not annealed.

Machinery retained (search_annealed, --anneal-grain, unfold above=, seed_pop,
max_share override): correct, tested, honest, reusable. Default finish stays
§15's single-transition unfold+polish. DESIGN §16 records the verdict; closes
homemaker-py-kpu.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-17 16:33:36 +01:00
3b3eef68a7 kpu: Schedule B in-run leaf-share grain annealing (search_annealed)
Ramp the leaf-share grain down within one continuous run (e.g. 4->3->2->off),
carrying the whole population across each step — graduated non-convexity over
the single hard sharing->off transition of the §15 finish.

- operators.unfold_shared_leaves(above=cap): unfold only leaves whose share
  exceeds the new grain cap, leaving smaller-share leaves collapsed for the
  next step. above=1 (default) keeps the full-unfold §15 behaviour.
- driver: max_share override threaded through _overrides_for/_fitness_for/
  _evaluate so a phase can rebuild the evaluator at a lower leaf_share_max cap;
  search(seed_pop=) evaluates an explicit initial population so a phase hands
  its whole population to the next instead of restarting from a single best.
- driver.search_annealed: one phase per descending grain then a de-share
  polish; unfold-above-cap between steps; cumulative accounting + grain-tagged
  history; honest canonical best (byte-for-byte verified vs homemaker-fitness).
- evolve: --anneal-grain LADDER CLI (self-finishing; §15 finish not applied).

8iv settled the primitive (grid unfold beat the circulation-aware slice), so
the ramp reuses the plain balanced-grid unfold at every step.

Tests: unfold above-cap selectivity, seed_pop seeding, search_annealed phase
stitching / honest finish / degenerate-ladder fallback. 258 pass. DESIGN §16.

Head-to-head A/B on harbor-house still to run; verdict pending (issue open).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-16 08:38:08 +01:00
48261142df 3l6: document unfold+polish auto-finish in DESIGN §15
Add DESIGN.md §15 recording the leaf-sharing output-honesty bug (internal
sharing objective diverges from canonical scorer), yaa's conclusive
unfold-then-polish investigation, and the driver.polish_finish auto-finish
fix + --polish-budget CLI knob. Matches §13.10's documentation of the
original leaf-sharing feature.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-15 10:35:14 +01:00
e09221051c 6zy/§11.8: co-tune topology diversity × tournament pressure — null robust
Expose tournament_k (default 2) on search()/search_staged(), threaded into
both _tournament call sites and the staged path's internal search() calls;
HOMEMAKER_TOURNAMENT_K env knob in the scaled/staged harnesses; run_6zy_ab.sh
joint niche×k grid (RESUME-able).

Result (negative, acceptable): no (niche,k) cell beats the legacy (off,k=2)
baseline. Blank-slate programme-house (5 seeds) baseline mean 4.80 fails is the
best of the 6-cell grid; every k>2 and every niche=on cell is 6.0-7.0. Niching
bites (pop_distinct 16/16 vs 4-11) but sharper pressure does not convert it to
lower fails — §11.5 'diffuses effort' null is robust to selection pressure;
plateau stays reachability-bound (confirms §11.4/§11.5).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 22:56:23 +01:00
d627ee5fb2 psk/§14: island model — null (best-of-N at equal budget wins)
Prime a population from N independent converged elites + crossover-heavy
migration phase, vs best-of-N at equal total budget. Island does NOT win:
harbor 68 vs control 67 (within parallel noise), maple 124 vs control 116
(decisive). Default-off child_probe hook on driver.search instruments the
deciding mechanism: area-matched crossover across independently-converged
elites rarely synthesizes (1/65 harbor, 3/63 maple beat the better parent,
max fail-drop 2-5), confirming the alignment hypothesis (non-canonical 9gp
encoding -> disruptive splice). Search-machinery null #3; residual stays
geometry/shape-bound. 233 tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 06:20:29 +01:00
bb9b355f14 x3b/§13.10: productionise leaf-sharing — per-code share grain + CLI wiring
Make the §13.3 lever a first-class feature, not experiment-only.

- programme.py: SpaceReq.share (default 1) + has_share, parsed from
  patterns.config 'share: N'.
- operators._share_grain: resolve per-code grain from leaf_share_factor
  selector — 0 = per-code opt-in (share iff share:N>=2), >=2 = global with
  per-code override (share:1 opts OUT, share:N sets grain). _share_rooms
  groups per resolved grain.
- End-to-end conf injection without monkeypatch: load_config(overrides=)
  merges run-level keys last; driver.search / innerloop.optimise /
  NativeEvaluator / _fitness_for thread conf_overrides={leaf_sharing:True}
  through both inner-loop and off-tree scorers when sharing is on.
- homemaker-evolve: --leaf-sharing/--no-leaf-sharing + --leaf-share-factor
  (env HOMEMAKER_LEAF_SHARING / HOMEMAKER_LEAF_SHARE_FACTOR).
- Example programmes untouched (§13.3/§13.9 stay reproducible). Experiment
  load_config monkeypatches updated to accept overrides=.

Tests: grain modes, opt-out, default-OFF parity, load_config overrides,
programme parse, CLI parse. 233 pass. Smoke: harbor 37 vs 95 fails on/off.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 22:04:35 +01:00
f43de001fb rq2/§13.9: flip share_edge_cap default-ON for leaf-sharing runs
§13.8 verdict was positive and monotone-harmless, so default the share-aware
edge-too-long cap to leaf_sharing when share_edge_cap is unset — mirrors the
pll bal+share and §13.6 interior_outside default flips. Explicit
share_edge_cap=False still reproduces the pre-flip control arm.

- fitness.Fitness.__init__: cap defaults to self._leaf_sharing when the conf
  key is unset (None); explicit True/False honoured.
- run_staged_search.py: pin conf["share_edge_cap"] = share_edge in both A/B
  arms so SHAREEDGE=0 stays a clean control post-flip.
- tests: control arm now pins share_edge_cap=False; new
  test_edge_cap_defaults_on_under_leaf_sharing guards the flip.
- DESIGN.md §13.9: rebaseline §13.x floor (maple 80.3→74.0, harbor 34.7→31.0).

Non-sharing runs untouched: programme-house control re-score reproduces
bit-for-bit. 222 tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 21:38:53 +01:00
393183b356 hph/§13.8: share-aware edge-too-long cap — shared leaves no longer penalised for aggregate wall length
§13.7 flagged edge-too-long as harbor's top fail class. Dissection showed the
bulk are a leaf-sharing REPRESENTATION ARTIFACT: a share=k leaf aggregates k
same-code rooms, so its walls run ~k× the flat 8 m cap purely for being big —
the same §13.3 leak (size/missing relaxed for shared leaves) on the wall measure,
since edge_cost/outside_edge_cost ignored leaf.share.

Fix: Fitness._edge_cap(*leaves) scales the 8 m cap by the largest type-guarded
leaf_share among adjoining leaves, mirroring quality_size's k×target; non-shared
leaves keep the flat cap so genuine narrow/oversize pathologies stay flagged.
Gated behind a share_edge_cap config knob (SHAREEDGE env), default OFF so the
§13.x controls reproduce.

A/B (full Phase-8 stack, staged, 20k evals, seeds 0/1/2): control reproduces
§13.7 (maple 80.3 exact, harbor 34.7≈34.0); share-aware arm maple 80.3→74.0
(−7.9%), harbor 34.7→31.0 (−10.6%), zero regressions across 6 seeds. Positive
and monotone-harmless (only ever removes a false-positive fail). Verdict:
recommend default-ON; follow-up issue flips the default + rebaselines the floor.

Tests: 6 new unit tests for _edge_cap (221 pass).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JygRv4n2dcyDQqMiDRe7TN
2026-06-28 21:24:51 +01:00
27d7a0f771 erc.7d/§13.7: high-budget harbor floor probe — close 71d NO-GO, wrap Phase 8
500k serial full-stack harbor probe (probe_harbor_floor.py): 20 fails,
crinkliness 13→4, landlocked crinkliness ~13→2 of 20. Interior-O (default-ON,
erc.8) is 71d's named fix and dissolved its landlocked-crinkliness target;
residual now diffuse (top class edge-too-long). NO-GO on 71d.

Cumulative Phase-8 floor vs §12.2 baseline (leaf-share-relaxed): maple
136.0→80.3 (−41%), harbor 74.0→34.0 (−54%) — all from construction levers,
none from search machinery, per the epic thesis.

Closes erc epic: 71d/7u5/jrb/u8x superseded-by-construction; erc.5/erc.6
wont-fix (Diag A/B revisit conditions unmet). DESIGN §13.7.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JygRv4n2dcyDQqMiDRe7TN
2026-06-28 14:23:34 +01:00
2491a9be12 erc/ld2: interior-O light-well seeding — §13.6 positive on dense floors
Seed O as interior light wells (most-landlocked leaves first, count scaled
by room count via outside_divisor) instead of one peripheral O, attacking the
erc crinkliness residual: seed diagnostic confirms every crinkliness fail is
under-exposed (landlocked), none over-exposed.

A/B (20k evals, seeds 0/1/2, bal+share stack, §13.6): control reproduces §13.5;
interior odiv=3 gives harbor -16.4% (all seeds improve) and maple -2.8%
(net-neutral). Default-optimal divisor 3 found by seed sweep (6 was null).

Lever default OFF; default-ON flip tracked as erc.8.

- operators: interior_outside + outside_divisor through constructive_topology,
  lift_base_to_storeys, _assign_adjacency_aware (fix n_circ budget for >1 O)
- driver.search/search_staged threading; run_staged_search.py INTERIORO/ODIV env
- test_interior_outside_seeds_landlocked_wells_and_scales_count
- experiments/run_interioro_ab.sh; DESIGN.md §13.6

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 07:20:20 +01:00
30a29387d5 erc.7: factor sweep — factor 3 confirmed default under bal+share (§13.5, close)
leaf_share_factor 2/4 under bal+share, seeds 0/1/2. factor 2 regresses both
(maple +10.4, harbor +13.0); factor 3 and 4 tied within seed noise. Keep
factor 3. leaf_share_max=4 covers factor<=4, no missing-fail leak. Closes erc.7.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 10:57:02 +01:00
34a7b2ecf9 erc.7: leaf-sharing × depth-balancing synergy CONFIRMED end-to-end (§13.5)
Synergy A/B: bal+share vs share-alone, factor 3, seeds 0/1/2, staged 20k.
maple 86.3->82.3 (-4.6%), harbor 50.7->40.0 (-21.1%, non-overlapping arms).
Control reproduces §13.3. Adds run_synergy_ab.sh + run_sharefactor_sweep.sh
(factor 2/4 sweep under bal+share, running).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 16:47:14 +01:00
a5dfc18f6c erc.4: depth-balanced construction A/B verdict — modest standalone (§13.4, close)
End-to-end 20k A/B (seeds 0/1/2): depth-balancing gives −5.8% maple / −3.2%
harbor with OVERLAPPING arms — far less than the −11/−12% seed-floor probe,
because the 20k search erodes most of the seed advantage via divide/undivide
mutations (unlike leaf-sharing's structural leaf-count cut, which the search
cannot undo). Baseline reproduces §12.2 (maple 137.0 vs 136.0, harbor 74.0).

Promise is the additive floor with leaf-sharing (probe: bal+sh3 << share3-alone);
the decisive test is erc.7 synergy. Keep depth_balanced default OFF; close erc.4,
advance erc.7.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 07:07:20 +01:00
4aaa295dc1 erc.4: depth-balanced construction mechanism + floor probe (§13.4)
_grow_leaves grew a random caterpillar, so equal-target rooms landed at
wildly different binary-tree depths — the depth-driven size maldistribution
Diagnostic B (§13.2) localized (same code at 0.05x and 14.7x target). The
depth_balanced flag always splits a shallowest leaf instead, growing a
near-complete tree so the proportion-aware sizing pass hits each target with
cut fractions near their proportional value.

Floor probe (diag_depth_balance.py): depth spread collapses 7->1, the giant
ratio falls (maxR 12->8 harbor / 16->6 maple), %undersize 54->25 / 42->22,
and the achievable floor drops -12% harbor / -11% maple at EQUAL leaf count.
Additive with leaf-sharing (bal+sh3 beats §13.3 share3-alone). Default OFF,
214 tests pass; threaded through driver.search/search_staged and exposed via
DEPTHBAL in run_staged_search.py. End-to-end 20k A/B running.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 22:36:24 +01:00
605385a1ba erc.3: leaf-sharing A/B verdict — −37% maple / −32% harbor (§13.3, close)
Staged 20k A/B (seeds 0/1/2, factor 3) vs default-OFF baseline:
  maple-court  137.0 → 86.3  (−37%)
  harbor-house  74.0 → 50.3  (−32%)
Baseline arm reproduces §12.2 exactly (maple 137 vs 136, harbor 74.0 vs 74.0);
total separation (every share run beats every baseline run same-programme);
~35% faster at equal budget. First Phase-8 floor-mover, 5th construction win.

Closes erc.3. Follow-ups: dyh (productionise on evolve CLI / patterns.config),
erc.7 (erc.4 depth-balancing synergy + factor/max_share sweep).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 21:52:19 +01:00
e983229857 erc.3: explicit per-leaf multiplicity closes the leak; driver A/B wired (§13.3)
Replace the area-derived share recovery with explicit, type-guarded per-leaf
multiplicity: construction stamps leaf.share=k and leaf.share_type=code; the
fitness (graph.leaf_share) honours k only while leaf.type==share_type, so any
retype/undivide auto-invalidates a stale share — no operator resets, and a
small leaf cannot retype its way into covering rooms it does not provide. Two
Node fields survive the whole search via deepcopy (genome.decode is unused in
the hot path); .dom emits `share` only on a live shared leaf.

This closes the §13.3 missing-fail leak: floor probe missing 17–44 → 0, and the
achievable floor drops −39% harbor (120.3→73.3) / −32% maple (194.7→133.0) with
no re-emergence as size fails.

Flag threaded through driver.search/search_staged → constructive_topology /
lift_base_to_storeys, exposed via LEAFSHARE/LEAFSHAREFAC in run_staged_search.py
(injects the objective into inner-loop + final-score fitness so both A/B arms
share one programme dir). run_leafshare_ab.sh runs the staged 20k A/B.
Smoke-tested end-to-end (harbor, factor 3, re-score OK). 214 tests pass;
default-OFF reproduces baseline.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 18:16:17 +01:00
bf3ff43837 erc.3: leaf-sharing mechanism + floor probe (§13.3)
Same-code rooms collapse into fewer, larger SHARED leaves so the ~1.8/leaf
shape tax (§13.1) is paid once per group. Multiplicity k is recovered from
area (k=clamp(round(area/target),1,max_share)) — no genome change — and used
in two default-OFF sites: graph.check_space_counts counts coverage (Σk vs
req.count) so one leaf covers several rooms without a missing fail, and
fitness.quality_size centres on k×target (σ scaled by k). Construction:
operators._share_rooms groups instances; _size_divisions_from_targets sizes
shared leaves to k×target via leaf_mult.

Floor probe (experiments/diag_leaf_sharing.py, harbor+maple, seeds 0/1/2,
+innerloop): total fails −27% harbor / −16% maple at share3, shape factors
fall ~linearly with leaf count (confirms §13.1). Cap: 17–44 missing fails
leak because depth maldistribution (§13.2) keeps shared leaves below k×target
so round() undercounts; inner loop can't close it. Net still positive.

Default-OFF reproduces baseline exactly (214 tests pass). Driver plumbing +
staged 20k A/B remain; §13.3 records the next design fork.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 08:30:26 +01:00
be6857414d erc.2: Diagnostic B — undersize-despite-slack localization (§13.2)
The "56% empty plot" is a misreading: sized rooms already hold 1.4-1.5x
their aggregate target area; ~46% of plot is circulation, not claimable
void. Size fails are depth-driven MALDISTRIBUTION — the same type/target
leaf lands 0.05x..14.7x by binary-tree position. The inner loop cannot
repair it (frozen topology, budget-80 size fails move only -1.6/-3.7).

=> Falsifies plot-fill-as-claim-void: re-scope erc.4 to depth-balanced /
giant-splitting construction; deprioritise erc.6 (inner-loop term, wrong
DOF). Reinforces erc.3 leaf-sharing for the starved tail.

Script: experiments/diag_slack_localization.py (self-contained evidence).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 22:47:34 +01:00
7bd4adf32a erc.1: Diagnostic A — per-leaf shape-fail vs density (§13.1)
Controlled synthetic sweep (maple-court, room set fixed, circ_divisor 2->9)
shows per-leaf shape-fail is FLAT vs slicing density (1.72-1.94, no trend)
while TOTAL shape fails track leaf count linearly (139->116). Crinkliness
dominates (~0.8/leaf) and is flat; cuts are already squarest yet still pay
~1.8 fails/leaf. Floor is INTRINSIC to per-leaf slicing, not cut quality.

Verdict: prioritise leaf-sharing (erc.3); deprioritise compactness-cuts
(erc.5 -> P4). Adds experiments/diag_leaf_shapefail.py.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 22:06:04 +01:00
e95a3477a8 Fix parallel search nondeterminism; re-diagnose homemaker-py-xcy
The constructive seeder was never nondeterministic: _assign_adjacency_aware
ends every max/min with a unique leaf-idx tiebreak and uses set unions only
for membership, so iteration order never leaks. constructive_topology(seed=0)
is byte-identical across processes for every example programme. The cited
"sig 4480 vs 16064" was a measurement artifact — Python's builtin hash() of a
str is salted per process (PYTHONHASHSEED), so an identical signature hashes to
different ints run-to-run.

The real run-to-run noise was parallel-only: driver._run_batch admitted futures
via as_completed (completion order), and admit() is order-sensitive (accrues
n_evals per result; keeps the first individual of an equal-key tie as best). A
long parallel run diverged 167 vs 161 fails (maple seed 0). Fix: admit futures
in submission order (block on each result in turn; all still run concurrently),
reproducing the serial admission sequence. Two workers=4 runs are now
byte-identical. Serial (workers=1) was already byte-for-byte reproducible.

Per-seed numbers are reproducible only at a fixed worker count; serial != parallel
is expected (children/iteration 1 vs n_workers changes batch granularity).

- driver: iterate futs in submission order, not as_completed
- test: test_search_parallel_is_reproducible (fails on pre-fix, passes on fix)
- DESIGN.md §12.4: corrected the reproducibility note

Closes homemaker-py-xcy

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 23:25:50 +01:00
cfb0518531 §12.4: construction-granularity A/B — NULL; close c3g; note determinism bug
End-to-end A/B (maple div6/div8, harbor div6, seeds 0/1/2, 20000 evals) vs the
§12.3 div=3 baseline: every arm within ±1.7 of baseline (maple 136.0 -> 137.0 /
134.3; harbor 74.0 -> 75.3), inside the measured ±3 noise floor with large
per-seed spread. Coarsening the circulation spine lowers the raw shape floor but
raises access/adjacency by as much; end-to-end they wash out. Verdict: keep
circ_divisor=3; the maple/harbor residual is the geometry floor of the slicing
representation at this room density — neither search machinery (§12.3) nor
construction granularity (§12.4) moves it beyond noise.

En route: the div=3 control (129 vs §12.3's 126) exposed a reproducibility bug —
_assign_adjacency_aware iterates id()-ordered sets of Node objects, so the
constructive seed is nondeterministic across processes (~±3 fail noise). Filed
homemaker-py-xcy (P2); per-seed ledger numbers are not reproducible, only
multi-seed means.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 00:49:56 +01:00
e700090c5c §12.3 residual diagnostic: over-granularity, not placement; file c3g
Per-leaf breakdown of maple-court constructive seeds (6 seeds) overturns the
earlier 'shape-aware placement' handoff guess: shape fails are UNIFORM
(~68/73 leaves fail) at only 0.44 plot utilisation, dominated by crinkliness
(perimeter/area) then size (undersize). So the residual is neither a room->leaf
placement mismatch (no well-shaped leaves to place into) nor density-bound — it
is over-granular construction (73 small leaves for 52 rooms). Corrected the
§12.3 verdict accordingly and filed homemaker-py-c3g (construction granularity /
leaf-shape lever) as an unproven, must-be-A/B'd hypothesis.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-21 20:56:49 +01:00
7e39bf5870 Phase 7 §12.3: 9gp A/B measured — NEGATIVE; close 9gp + epic leu
24-run sweep (maple-court + harbor, seeds 0/1/2, 20000 evals): M3 reassociate
and the shape-feasibility filter are both neutral-to-slightly-worse vs the
§12.2 baseline (maple 136.0 -> 139-140, harbor 74.0 -> 77-78). Baseline controls
reproduce §12.2 exactly, so the negative is real.

Verdict: the Phase-7 residual is the geometry/shape floor of the constructed
slicing layouts, not reachability/feasibility-bound — third independent negative
on search machinery (§11.4/§11.5/§12.3) vs four construction/seed wins
(§11.2/§11.6/§11.7/§12.2). A full canonical Polish rewrite is not justified: its
one testable promise (associativity reachability) was tested and did not pay.
Both operators kept default-OFF.

Closes 9gp.1, 9gp.2, 9gp; epic leu (Phase 7) auto-closed (3/3). Adds the
reproducible sweep harness experiments/run_9gp_ab.sh.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-21 07:21:51 +01:00
6ee5d4b4ae Phase 7 §12.3: re-scoped 9gp — shape-feasibility filter + M3 reassociate (9gp.1, 9gp.2)
Land the two evidence-supported parts of the re-scoped 9gp capstone as
operators on the existing decoded Node tree (no Polish-expression rewrite),
each default-OFF and measured against the §12.2 leu.2 baseline.

9gp.1 shape-feasibility pre-filter: operators.predicted_shape_fails lays a
topology out at its proportion-aware target geometry and counts shape fails
(size/width/proportion/crinkliness); driver._evaluate prunes clearly-infeasible
topologies before the inner loop (1 eval vs ~80), guarded so nothing that could
beat the incumbent is discarded. search/search_staged feasibility_filter,
feasibility_max_shape_fails (env FEAS/MAXSHAPE), default OFF.

9gp.2 M3 Wong-Liu reassociate: operators.mutate_reassociate adds associativity
(a|b)|c <-> a|(b|c) on same-orientation live cuts — the canonical-slicing move
missing from swap(M1)/rotate(M2), attacking the §11.4/§11.5 reachability
bottleneck. enable_reassociate (env REASSOC), default OFF (weight 0 -> baseline
byte-identical).

Unit tests (operators + driver) green, full suite 211 passed; maple-court smoke
run clean under native fitness. A/B sweep handed off per the plan; DESIGN.md
§12.3 documents the design and the pending measurement.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 18:54:48 +01:00
995342d0a4 Phase 7 §12.2: proportion-aware constructive seeding + storey_minimum fix (leu.2, cq1)
Size each constructive-seed cut from leaf TARGET areas (division=[f,f] gives
left area-fraction f) and pick each cut's rotation for child squareness — both
derived from target dims, topology/type assignment untouched. Area-only
regressed (slivers); rotation choice is what makes it pay.

End-to-end (20000 evals, 3 seeds, staged): harbor 85.3->74.0 (-13%, best 69),
maple-court 151.7->136.0 (-10%, best 126). PROP=0 reproduces the §11.7/§12.1
baselines exactly. programme-house regresses at fixed budget (deeper local
optimum walls off the undivide restructuring path) but a budget sweep shows
it's convergence speed, not a worse asymptote (PROP=1 reaches 1 fail at 150k).
Default-on (seed_proportion_aware=True, env PROP=1).

cq1: n_storeys now honours storey_minimum, not just level: keys — programme-house
(storey_minimum:2, all rooms level:0) was seeded one storey short and fell
through to plain search. New programme.storey_minimum()/n_storeys_for();
driver.search passes min_storeys to the seeder; search_staged routes on the max.
No-op for harbor/maple; programme-house single-stage 8.0->5.0.

New maple-court best (126) saved as generated.dom. 204 tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 14:04:42 +01:00
7d7994e7a3 Phase 7 §12.1: larger-than-house benchmark maple-court + baseline (leu.1)
26 programme entries / 52 rooms / 3 storeys (~1015 m2 internal). Mirrors
harbor's adjacency-to-c load + secondary adjacencies; room codes avoid the
generic c/o/s leading-letter trap. Staged adjacency-aware baseline (20000
evals, URB_NO_OCCLUSION=1): 145/158/152 fails, mean 151.7; all native
re-score OK. Best (145) saved as generated.dom. Recorded in DESIGN.md §12.1.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 14:00:21 +01:00
d004e4c937 Phase 6 §11.7: adjacency-aware lift + secondary adjacencies (ld5)
_assign_adjacency_aware gains fixed_circ (seed the connected-dominating-set from
given circulation leaves) and secondary-adjacency-aware room placement: codes
with the most non-c adjacency requirements are placed first, each onto the open
slot satisfying the most of its requirements against already-typed neighbours
(clustering k1<->da1, da1<->o). lift_base_to_storeys(reqs, adjacency_aware=True)
grows the upper-floor circulation spine off the inherited vertical core and
assigns rooms around it; threaded through driver.search_staged
(seed_adjacency_aware) and run_staged_search.py (ADJ env).

End-to-end staged harbor, 20000 evals, mean total fails over 3 seeds:
ADJ=0 99.0 (reproduces the §11.4 staged lex baseline exactly), ADJ=1 85.3
(-13.7, -14%; best 78). New best harbor configuration overall: staged baseline
99.0 -> single-stage adjacency-aware (§11.6) 90.7 -> staged + adjacency-aware
lift 85.3. Staging and adjacency-aware seeding compose.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 11:47:40 +01:00
c1586237ca Phase 6 §11.6: adjacency-aware constructive seeding (s44)
operators._assign_adjacency_aware spends ~one extra leaf per three rooms on a
greedy connected-dominating-set of circulation leaves (read from the geometric
leaf_graph, type-independent), so every room borders a connected circulation
spine and adjacency-to-c + access are satisfied by construction. Default-on via
constructive_topology(adjacency_aware=True), threaded through
driver.search(seed_adjacency_aware) and run_search_scaled.py (ADJ env).

End-to-end single-stage, 20000 evals, mean total fails over 3 seeds:
harbor 110.0 -> 90.7 (-17.5%; ADJ=0 reproduces the §11.2 105 baseline exactly),
programme-house 12.3 -> 9.3 (-24%). Adjacency-aware single-stage harbor (mean
90.7, best 85) beats the §11.3 staged best of 95 — the first Phase-6 fail-count
reduction from seeding. Follow-ups (lift_base_to_storeys, secondary adjacencies)
filed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 09:23:12 +01:00
059964ee05 Phase 6 §11.5: structural niching + restarts — negative result (c4c.5)
genome.signature: ratio-invariant structural topology hash (per-storey tree
shape + cut orientation + leaf types), the cheap stand-in for the 9gp canonical
encoding. driver gains niche_by_signature (one individual per topology, replaces
the fitness-scalar dedup) and restart_patience (soft restart: keep elites,
refill with fresh seeds); SearchResult gains n_distinct_signatures /
diversity_history / n_restarts.

Diversity criterion MET (final-pop distinct ~5/16 -> 16/16). Gate NOT met:
blank-slate programme-house mean fails 12.3(legacy)/12.7(niche)/13.0(restart)
over 3 seeds at 20000 evals; harbor staged 95/94/108. Niching is a tie within
seed noise, restarts strictly worse — falsifies the premise that the
fitness-scalar dedup causes premature convergence. Both flags default-off,
kept for reuse. Epic c4c complete.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 23:42:39 +01:00
ed2869074b Phase 6 §11.4: graded high-fail objective — negative result (c4c.4)
Implement a graded proximity comparator key (-n_fails, grade, fitness) behind
a default-off use_grade flag: fitness._leaf_grade / score_with_grade sum
f/FAIL_THRESHOLD over failing per-leaf quality factors; scalar fitness and fail
count stay untouched so the inner-loop 0.5^n cliff (§5.4) is unaffected (0/9
regression check: PASS). Read once per child in driver._evaluate off the
already-optimised tree; threaded through search_staged (Stage 2 only).

Harbor staged A/B (20000 evals, seeds 0/1/2): lex 95/96/106 (mean 99.0) vs
lex+grade 99/98/102 (mean 99.7) — grade wins 1/3, no plateau escape. Premise
falsified: within a fixed fail-tier 0.5^n is constant so fitness still spans
~6 orders of magnitude; grade above fitness displaces that working signal.
Verdict: reject; lexicographic (-n_fails, fitness) stands. Flag kept default-off
for reproducibility / possible reuse as a §11.5 diversity signal.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 22:33:29 +01:00