Commit graph

221 commits

Author SHA1 Message Date
6d719e03ab c94: beam/best-first search over adjacency-aware room placement (null)
operators._assign_adjacency_aware gains beam_width (default 1 = exact
prior greedy behaviour), threaded through constructive_topology/
lift_base_to_storeys/driver.search/search_staged as
construction_beam_width. Verified functioning on an adversarial
synthetic case, but byte-identical raw-seed output to greedy at every
width tested (1/4/8/20) on both example programmes -- no headroom for
the beam to find on this repo's programmes. DESIGN.md section 29.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R8agJBT2ZpmF3ErW7wi2wY
2026-07-28 00:00:51 +01:00
a04fb2bc8e docs: DESIGN.md §28 — backfill cdl (finish-time local_search default)
§25 still said the evolve.py wiring and broader sweep were pending on
homemaker-py-cdl; that closed in the previous commit, so update §25's
stale forward-references and add §28 with the 46-file sweep results and
the collapse_insearch hot-path reasoning for leaving collapse_global's
own default off.
2026-07-27 21:26:01 +01:00
2ede6ed514 cdl: default local_search on for finish-time collapse call sites
homemaker-py-9wi's 2-opt adjacency polish (collapse_global's local_search
kwarg) was validated on harbor-house alone; this closes homemaker-py-cdl
by extending the A/B sweep to programme-house (34 more files, 46 total):
0 regressions, 2 improvements, rest identical.

collapse_global's own default stays False since it also runs every
fitness eval via collapse_insearch (qpk) on the unmerged tree, where the
2-opt pass would add cost to a hot path the sweep never measured.
Instead default it on at the two one-shot finish-time call sites:
homemaker-collapse --local-search, and a new homemaker-evolve
--collapse-local-search wired through driver.collapse_best's
**collapse_kw.
2026-07-26 23:33:40 +01:00
9c6b1552eb docs: DESIGN.md §26/§27 — backfill 9o5/xi7/b3v (type superposition) and mi7 (bubble-diagram signal)
Two closed, substantive experiments were missing their DESIGN.md write-up
despite being referenced as prior art by later sections:

- 9o5/xi7/b3v (closed 2026-06-30/07-17): multi-use-leaf type superposition,
  a full feature build + real A/B validation (negative — OFF beats ON on
  both programme-house and harbor-house) + a veto-hatch follow-up for the
  one genuine false-positive interchange class found. §17 and §20 both cite
  its verdict directly but it never got its own section.

- mi7 (closed 2026-07-25): 3D bubble-diagram / topological-hop-distance
  fitness signal prototype, tested against real evolved trajectories on two
  programmes, both formulations null. bubble.py was left in the tree
  uncommitted "as documented reference" by the closing session -- committing
  it now (with two trivial ruff fixes: unused import, ambiguous var name) so
  the reference this write-up makes to it is actually resolvable, plus a
  CLAUDE.md module-list entry.

Numbered §26/§27 (appended, not inserted chronologically) to avoid
renumbering every cross-reference in §14-§25.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 23:16:21 +01:00
5c8a5a5e09 docs: DESIGN.md §25 — 2-opt local search past collapse_global Jacobi plateau (9wi)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 21:15:21 +01:00
699b049333 bd: sync issues.jsonl export after 9wi close / cdl file
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 21:09:20 +01:00
67f0c38f45 9wi: 2-opt local search past collapse_global's Jacobi adjacency plateau
The Jacobi adjacency relaxation in collapse_global (94g) re-solves a linear
assignment each round holding neighbours' labels fixed from the previous
round, which can 2-cycle between labellings that satisfy zero adjacency
requirements even when a fully-satisfying permutation exists (proved by
test_two_opt_polish_escapes_jacobi_plateau on a minimal 4-cell chain).

Fitness._two_opt_adjacency_polish runs after the Jacobi fixpoint and tries
swapping the labels of every same-level pair of supply leaves, keeping a
swap only on strict improvement -- monotone by construction. Gated behind
collapse_global(local_search=...) / homemaker-collapse --local-search,
default off pending a broader sweep (homemaker-py-cdl). Swept the 11
harbor-house evolved-*/3m/materialised .dom files: 0 regressions, 1 real
improvement (evolved-anneal-3M.dom 21->19 fails).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 20:59:07 +01:00
bf2654c037 bd: sync issues.jsonl export after y51 close / xyu file
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 15:39:07 +01:00
bc59849666 y51: synthetic room-count sweep finds no clean ruin_recreate size threshold
Locates the threshold f1d's programme-house/harbor-house split implied,
using four synthetic sizes (10/14/18/22 rooms) derived from programme-house
by scaling its bedroom+ensuite module count, since no natural third example
programme sits between the two. Results are noisy and non-monotonic (n=10
mild win, n=14 clean null, n=18 strongest trend at p=0.098, n=22 near-null)
rather than a clean decay with room count -- documented in DESIGN.md #24.
enable_ruin_recreate stays default OFF; filed homemaker-py-xyu as a
low-priority follow-up (larger-N at n=18, or a non-synthetic third example).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 15:35:03 +01:00
d7884e6638 bd: file y51, f1d size-threshold follow-up
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 09:38:47 +01:00
0d94e58119 f1d: ruin-and-recreate LNS operator, validated positive on programme-house
Adds operators.mutate_ruin_recreate: un-divides one wing of a storey and
rebuilds it with the same adjacency-aware constructor the seeders use
(_assign_adjacency_aware, generalised with a new `scope` param), seeded
from the surviving circulation bordering the wing. Gated behind
enable_ruin_recreate (default off) / --ruin-recreate, same pattern as
reassociate/bridge_circulation.

A/B (qpk protocol, DESIGN.md §23): initial uniform-weight run was
underpowered (fired ~1/32 children), null. A weight=3.0 follow-up
(_MUTATION_WEIGHTS["ruin_recreate"]) showed a statistically significant
win on programme-house across 15 seeds (8W/1L/6T, mean fails 7.07->6.00,
Wilcoxon p=0.041) but no consistent effect on harbor-house across 8 seeds
(3W/2L/3T). Kept default off pending a size-threshold follow-up.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 09:31:42 +01:00
bc11394499 bd: sync issues.jsonl export after dolt push
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-25 12:30:24 +01:00
1e2fa2adfc lj3/qjg: larger-N + weight sweep finds bridge_circulation effect is noise
Combined follow-up to 8sh (DESIGN.md §22): raised bridge_circulation's
_MUTATION_WEIGHTS entry to 2.0 (lj3) and re-ran the qi6/qpk-protocol A/B
at 4x the sample size (qjg) in one sweep, since the two variables were
confounded if tested separately. Result is null in the opposite direction
from 8sh's small-N signal -- no total-fail benefit (p=0.71 programme-house
N=20, p=0.69 harbor-house N=12) and a higher rate of trajectory-divergence
-induced new not-connected fails than at the original uniform weight.
Reverted the weight bump; enable_bridge_circulation stays default off.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-25 12:20:22 +01:00
a103e0114a 8sh: insert/relocate-circulation repair operator (qi6 mechanism (a))
Adds operators.mutate_bridge_circulation: retypes the cheapest path between
two disconnected circulation components to circulation, directly clearing a
'level N not connected' fail instead of relying on the qi6 graded comparator
key (measured negative, DESIGN.md §18). Gated off by default via
driver.search's enable_bridge_circulation flag and evolve.py
--bridge-circulation, mirroring enable_reassociate's clean-toggle pattern.

qi6/qpk-protocol A/B (DESIGN.md §21) is directionally positive but mixed at
N=3/N=5 (never worse on total fails; clears 2/5 baseline not-connected fails
vs qi6's 0/4; one seed's RNG-trajectory divergence adds 2 new not-connected
fails) — kept default off pending a larger-N confirmation sweep
(homemaker-py-qjg) and a mutation-weight bump experiment (homemaker-py-lj3).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-24 19:48:13 +01:00
a78409df90 1ph: larger-N seed sweep confirms collapse_insearch positive, flip default ON
20-seed programme-house sweep (vs the original 5) resolves the qpk A/B's
mixed 3/5 result as small-sample noise around a true small positive: mean
fails 7.95->7.10 (~10.7%), 11W/6L/3T, paired t-test p~0.028. Flips
collapse_insearch's default from OFF to ON in evolve.py and driver.py
(_overrides_for/_fitness_for/_evaluate/search/polish_finish); opt out with
--no-collapse-insearch. fitness.Fitness itself is unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-24 09:55:34 +01:00
bba71f1b69 bd: sync issues.jsonl export after dolt push
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-24 07:40:09 +01:00
26eb334450 bd: file 1ph (qpk seed-sweep follow-up)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-23 23:52:54 +01:00
372bebd0a3 DESIGN.md: backfill §19 with homemaker-py-161's in-search A/B verdict
161 (2026-07-22) answered the "remaining open question" §19 left dangling —
threading fit into driver.search so shape_rotate/deslim can fire mid-GA —
but the result only ever landed in bd notes, never here. Also negative:
full-budget harbor-house A/B (seeds 0-3) shows no improvement, confirming
the finish-time finding at in-search scale. Both halves of §19's mechanism
space are now closed negative in the doc, matching bd state.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-23 22:39:08 +01:00
d125d2f19c bd: sync issues.jsonl export after dolt push
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-23 18:32:39 +01:00
8328ac1b69 qi6: full-budget A/B for graded circulation-connectivity signal, negative result
conn_grade ON vs OFF (qpk protocol, experiments/run_qi6_ab.sh): harbor-house
(budget 2500, seeds 1-3) byte-identical output in every seed — the secondary
comparator key never fired. programme-house (budget 3000, seeds 1-5) 3/5 seeds
tie exactly; seeds 1/2 diverge to a different topology but the fail delta is
adjacency/crinkliness/width/access/size, never connectivity. Zero of 4 cases
where a not-connected fail was present got cleared by the grade.

Mechanism (b)/(c) (graded proximity as tertiary comparator key) is falsified,
not just unconfirmed. Kept default OFF (already was). Closed qi6; filed
homemaker-py-8sh for the remaining candidate (mechanism (a): an explicit
insert/relocate-circulation operator that doesn't depend on the search
stumbling onto a fail-count tie).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-23 18:29:45 +01:00
baa9109c67 bd: sync issues.jsonl export after dolt push
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-22 17:44:51 +01:00
f22691e93d 161: thread fit into GA for in-search shape_rotate/deslim, negative result
driver.search()/search_staged() gain enable_shape_repair (default off),
mirroring the enable_reassociate clean-toggle pattern: only builds a
fitness.Fitness instance and passes it to operators.mutate() when
enabled, so shape_rotate/deslim (7fm) can actually be selected mid-GA
instead of always no-opping on fit=None.

Full A/B sweep (harbor-house, budget=1M, 4 seeds) shows no improvement:
mean fails 14.50 (off) vs 14.75 (on), within seed noise. Confirms 7fm's
finish-time finding at in-search scale — these operators don't rescue
harbor-house's residual fails even with GA selection pressure.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDZjAATDWW1xFfc7xnJqSt
2026-07-22 17:43:36 +01:00
d95df4b59a bd: sync issues.jsonl export after dolt push
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-19 20:38:46 +01:00
2b7a7d2926 qpk: in-search global collapse — run collapse_global per-eval during search
Runs the 94g finish-time cell↔room collapse inside every fitness eval
(collapse_insearch conf flag, default off, bit-identical when off) instead
of once at the end, so search optimises the collapsed objective directly.
Plumbed through fitness.py/driver.py/evolve.py the same way superpose/
conn_grade are; --collapse-insearch CLI flag.

A/B validated against the xi7 protocol (equal budget, both arms finished
with standard finish-time --collapse): POSITIVE, opposite of the 9o5/xi7
prior. harbor-house ON wins 3/3 (mean fails 80.3->72.0); programme-house
mixed 3/5 (mean fails 8.4->7.8). Kept default off pending a larger
programme-house sample; documented as a working opt-in for harbor-house-
scale-or-larger programmes. Full writeup in DESIGN.md §20.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-19 20:35:18 +01:00
07a4739576 7fm: targeted shape-repair operators (shape_rotate/deslim), negative finish-time result
Diagnosed the geometry-intrinsic residual from 94g's collapse: ratio
re-optimisation isn't the bottleneck (1500-eval NM makes zero difference on
the 12-fail collapsed best layout); the causes are upstream area starvation
and cut-orientation mismatch. Added mutate_shape_rotate/mutate_deslim
targeting each, gated on a Fitness instance like the existing reqs-gated
repair ops.

Evaluated as a finish-time exhaustive hill-climb on the same 6-layout
harbor-house sweep 94g used: zero improving moves found anywhere — every
candidate move traded the shape fail for a new adjacency/access fail on the
co-evolved layout (§4.2's lesson, now confirmed for topology repair). Closes
homemaker-py-7fm; spun homemaker-py-161 for the open in-search-GA question.

See DESIGN.md §19 for the full writeup.
2026-07-19 11:06:56 +01:00
94d4223a55 qi6: graded circulation-connectivity signal (§18)
The dominant post-collapse fail is the binary "level N not connected",
which is flat across fragmentation (a 7-component storey scores the same
as a 2-component one), so the outer search has no gradient toward
connected circulation. A finish-time convert-to-circulation repair was
prototyped and measured NEGATIVE (195->560 fails: bridging needed rooms
costs more missing-room fails than the one binary fail it clears).

Instead add graph.circulation_connectivity(G) = largest-circ-component
fraction, summed over storeys onto the score_with_grade proximity channel
(conf flag conn_grade; replaces the §11.4 leaf-grade there). It is a
secondary comparator key only — scalar fitness and fail count stay
byte-identical — restoring the gradient the binary fail lacks. Threaded
through driver (_overrides_for/_fitness_for/_evaluate/search; enabling it
implies the grade key) and evolve --conn-grade (default off).

A/B on full-budget runs pending; short smoke run confirms plumbing.

Tests: tests/test_conn_grade.py x9 (fraction contract, non-circ ignored,
monotone under (dis)connection, score/fail invariance); 276 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 18:44:24 +01:00
1ae8faac7c docs: DESIGN.md §17 — finish-time global cell→room collapse (94g)
Document the collapse as built (default-on in evolve + homemaker-collapse CLI):
label-relative vs geometry-intrinsic fails, collapse_global mechanism (c/o/s
partition, hard level, adjacency relaxation, threshold objective, public-access
pin), the two measurement corrections, keep-better wrapper + wiring, and the
6-layout sweep verification. DESIGN.md is the system-of-record for users without
beads access.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 12:45:35 +01:00
7c2b9402de docs: add collapse_cmd/homemaker-collapse to module lists (94g)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 12:39:23 +01:00
ec7f646a3b bd: file 7fm/qi6/qpk (shape-fix, circulation, in-search WFC), close 94g
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 11:14:41 +01:00
880a214d96 94g: public-access pin + keep-better wrapper + CLI/finish-hook wiring
Public-access term (preserve_public_access, default on): when the building's
only street access is an l/k ROOM neighbour of a public outside leaf (no
circulation fallback — an existential building-level check the per-leaf
objective can't see), that leaf is pinned (kept, its demand slot decremented)
so the collapse can't drop "no outside public access". Best layout 15→13
becomes 15→12 with zero new fails; sweep total 172→171, still monotone.

collapse_finish(root, **kw) -> (tree, base, coll, applied): keep-better wrapper,
scores on throwaway copies (score_with_fails merges in place), returns the
collapse only if fails don't increase.

Wiring: driver.collapse_best updates result.best (lineage +collapse, canonical
re-score); evolve.py runs it after the sharing polish behind --collapse/
--no-collapse (default on). New homemaker-collapse CLI (collapse_cmd.py) applies
it to an existing .dom, writing <stem>.collapsed.dom.

tests/test_collapse_global.py: demand-set relabel, level hard constraint, c/o/s
exclusion, no-op safety, keep-better/unmerged. 267 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 10:29:44 +01:00
da18ef744e 94g: threshold objective for collapse_global (fail-count, not continuous)
Add objective="quality"|"threshold" to collapse_global and make threshold
the finish-time default. The continuous-quality objective maximises
sum(usage_quality*area), which can trade one leaf just over the 0.1 fail
threshold for another just under (a fail SHUFFLE). The threshold objective
maximises the COUNT of passing size/width/proportion factors directly, with
continuous fit only as a tiebreak. A satisfied adjacency and a passing factor
share one weight (_COLLAPSE_FAIL_W) so both fail classes are minimised jointly.

Sweep over 6 harbor-house evolved layouts (total fails, base 195):
  adj_off/quality 192  adj_on/quality 185  adj_off/thresh 181  adj_on/thresh 172
adj_on/threshold is monotone across all 6 (never worse than baseline), so it
is the new default. Residual on the best layout (15→13) is the building-level
"no outside public access" constraint, outside the per-leaf model.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 08:37:02 +01:00
d52cce6863 94g: finish-time global cell→room collapse (Fitness.collapse_global)
Global relabel of inside-room leaves via one optimal assignment over the
full leaf set — the 9o5 per-class collapse generalised to N leaves ↔ M
required rooms — as a one-shot finish-time polish on a committed layout.

Assignable room_codes exclude any starting c/o/s to match the scorer's
own partition (check_space_counts skips those; cr1/st1/st2 collide with
the circulation/structure convention). Hard level constraint via a -1e12
forbid penalty. Adjacency handled as an iterated relaxation: geometry is
fixed at finish time so each leaf's graph neighbours are fixed; warm-start
from evolved labels, each pass a linear assignment over quality + an
adjacency bonus (has_adjacency vs current labels), Jacobi to a fixpoint.

Measured (level+adjacency): best evolved layout 15→14 fails, rougher ones
32→28 and 90→83; adjacency-on beats adjacency-off everywhere (off regresses
the best layout +1). Substrate only — not wired into search or a CLI yet.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-18 01:18:36 +01:00
5ee6b62070 bd: file 94g (global cell↔room collapse / WFC), link to xi7+b3v
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-17 18:17:24 +01:00
f283406190 bd: close b3v (interchange veto hatch)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-17 16:54:31 +01:00
9cbdf880a4 b3v: interchange:false veto hatch for interchange-class derivation
9o5 §7.5 escape hatch: a per-space `interchange: false` opt-out in
patterns.config removes a code from auto-derived interchange classes,
letting the architect veto a harmful grouping (harbor-house's transitive
8-code chain) without disabling superposition globally.

SpaceReq gains an `interchange` bool (default True). Honoured as an S0
short-circuit in interchangeable() and by filtering derive_interchange_
classes() input. Superpose default stays OFF regardless (xi7 verdict), so
this only bites when superposition is enabled on a real config.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-17 16:53:20 +01:00
c980100589 kpu: Schedule B A/B DONE — graduated grain ramp falsified (negative)
harbor-house 3M (500k/grain x3 + 1.5M polish, workers 4, ~22h): the in-run
grain anneal reached 1.26e-08 / 23 fails (canonical byte-for-byte), losing
decisively to both the direct --no-leaf-sharing baseline (5.14e-06 / 15) and
yaa's single-hard-transition warm chain (4.19e-06 / 15) — ~400x worse fitness,
+8 fails.

Each grain step spikes the fail count as its unfolded leaves acquire
independent shape fails (phase-end 19->21->27, final de-share 27->36); the
per-phase budget re-polishes a partially-materialised state the next step
materialises further, so coarse-grain gains do not carry forward. The polish
phase started from a deeper hole (36) than the warm chain's single clean
transition and 1.5M evals recovered only to 23. The sharing-phase topology
skeleton is best cashed in once, at full grain — not annealed.

Machinery retained (search_annealed, --anneal-grain, unfold above=, seed_pop,
max_share override): correct, tested, honest, reusable. Default finish stays
§15's single-transition unfold+polish. DESIGN §16 records the verdict; closes
homemaker-py-kpu.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-17 16:33:36 +01:00
3b3eef68a7 kpu: Schedule B in-run leaf-share grain annealing (search_annealed)
Ramp the leaf-share grain down within one continuous run (e.g. 4->3->2->off),
carrying the whole population across each step — graduated non-convexity over
the single hard sharing->off transition of the §15 finish.

- operators.unfold_shared_leaves(above=cap): unfold only leaves whose share
  exceeds the new grain cap, leaving smaller-share leaves collapsed for the
  next step. above=1 (default) keeps the full-unfold §15 behaviour.
- driver: max_share override threaded through _overrides_for/_fitness_for/
  _evaluate so a phase can rebuild the evaluator at a lower leaf_share_max cap;
  search(seed_pop=) evaluates an explicit initial population so a phase hands
  its whole population to the next instead of restarting from a single best.
- driver.search_annealed: one phase per descending grain then a de-share
  polish; unfold-above-cap between steps; cumulative accounting + grain-tagged
  history; honest canonical best (byte-for-byte verified vs homemaker-fitness).
- evolve: --anneal-grain LADDER CLI (self-finishing; §15 finish not applied).

8iv settled the primitive (grid unfold beat the circulation-aware slice), so
the ramp reuses the plain balanced-grid unfold at every step.

Tests: unfold above-cap selectivity, seed_pop seeding, search_annealed phase
stitching / honest finish / degenerate-ladder fallback. 258 pass. DESIGN §16.

Head-to-head A/B on harbor-house still to run; verdict pending (issue open).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-16 08:38:08 +01:00
3d0250817c bd: 8iv circulation-aware unfold falsified (A/B loss), close; correct kpu
Investigated homemaker-py-8iv (route access to unfolded shared-leaf children).
Built + A/B-tested circulation-aware slicing vs the existing balanced grid; the
150k-eval warm-start polish shows slice loses decisively (41 fails/3.5e-14 vs
grid 25 fails/2.4e-09, grid ahead at every milestone). Reverted the code to the
grid unfold; closed 8iv negative and corrected kpu (Schedule B) to use the grid
unfold, not slicing. Also closed yaa (investigation complete).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-16 07:42:57 +01:00
48261142df 3l6: document unfold+polish auto-finish in DESIGN §15
Add DESIGN.md §15 recording the leaf-sharing output-honesty bug (internal
sharing objective diverges from canonical scorer), yaa's conclusive
unfold-then-polish investigation, and the driver.polish_finish auto-finish
fix + --polish-budget CLI knob. Matches §13.10's documentation of the
original leaf-sharing feature.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-15 10:35:14 +01:00
479d4a57d5 3l6: leaf-sharing runs auto-finish (unfold+polish) before write
Leaf-sharing evolve runs optimised a sharing-credited objective (a shared
leaf of code X with share=k counts as k programme rooms, size re-centred on
k*target) but wrote the un-materialised genome, which the canonical
homemaker-fitness rates catastrophically worse (harbor-house: internal
1.03e-05 vs canonical 6.73e-29, 15 critical missing-room fails).

Fix: before write, driver.polish_finish() unfolds every live shared leaf
into k distinct rooms (operators.unfold_shared_leaves — pays down the count
deficit) then warm-starts a leaf_sharing=False polish search from the
unfolded genome so the materialised rooms get proportion/width/size cleanup.
Returned best.fitness is the canonical score (leaf_sharing off => internal
== canonical). This is yaa's proven unfold-then-polish path, made automatic.

evolve.py: new --polish-budget (env HOMEMAKER_POLISH_BUDGET; -1=auto=
budget//2, 0=unfold+rescore only). Interrupt forces polish_budget=0 for a
fast honest output. Default stays --leaf-sharing on (its topology-search
speed retained; output made honest by the finish). Schedule B in-run
annealing remains homemaker-py-kpu.

Verified (harbor-house 3000+1500): reported polish fitness 4.79788e-27
matches canonical homemaker-fitness exactly, 0 critical fails. Tests: 254
pass (+3 polish_finish).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-15 10:21:58 +01:00
053ace2f53 bd: yaa conclusive (unfold catches direct route); file Schedule B (kpu)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-15 07:50:15 +01:00
4f06037c09 bd: sync issues.jsonl (file yaa unfold follow-up 8iv)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-12 16:54:07 +01:00
aaaaad2e53 yaa: unfold shared leaves at sharing->no-sharing transition
operators.unfold_shared_leaves(): materialise each live shared leaf
(share=k) into k distinct same-code sibling leaves, splitting its
footprint into k equal-target children with squarest per-cut rotation
(_size_subtree_equal), clearing the share stamp. Surgical — siblings'
evolved geometry is untouched.

Investigation result (harbor-house, evolved-3M sharing seed):
- unfold alone (zero search) closes all 15 critical missing-room fails
  and lifts the canonical score 6.73e-29 -> 1.46e-19 (90->59 fails).
- warm-starting a --no-leaf-sharing evolve from the unfolded seed runs
  ~7-8 orders of magnitude ahead of the naive (un-materialised) warm
  start at equal budget. The count deficit, not the sizing, was what
  stranded Schedule A deep in the fail hole.

Adds test_unfold_shared_leaves_materialises_deficit.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-12 16:34:33 +01:00
4c200389e3 bd: file leaf-sharing score-mismatch bug (3l6) + annealing/unfold investigation (yaa)
Ran harbor-house init.dom under 3M-eval memetic search. Default --leaf-sharing
scored 6.7e-29 (90 fails, 15 missing rooms) when re-scored by canonical
homemaker-fitness, vs 5.14e-06 (15 fails, 0 critical) for the honest
--no-leaf-sharing warm-start chain. Head-to-head confirmed a naive
sharing->no-sharing warm-start plateaus ~60x behind, motivating programmatic
unfold of shared leaves at the phase transition (yaa).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-07-12 15:45:30 +01:00
c9320cd7d3 xi7: 9o5 validation run complete — superpose NULL/negative; file b3v veto hatch
A/B at equal budget (collapsed score): --superpose vs --no-superpose.
- programme-house (budget 3000, seeds 1-5): OFF wins 4/5
- harbor-house  (budget 2500, seeds 1-3): OFF wins 2/3
Relaxation gap (§7.4) small (ratio 1.01-1.23); per-eval collapse removes it
by construction, so the failure mode is geometry-floor dominance, not the gap.
harbor-house 8-code chain misgroups and adds fails -> filed b3v (interchange:false).
Verdict: keep --superpose default OFF.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8566xAxTnwtJTkpXjYNZm
2026-06-30 08:24:43 +01:00
9c414834c7 bd: sync issues.jsonl (9o5 closed, xi7 follow-up)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:12:45 +01:00
c3635634e8 9o5: type superposition + per-eval collapse (multi-use leaves)
Interchangeable codes (similar size/width/proportion, compatible level/stack,
no adjacency edge) form equivalence classes derived from the programme. With
--superpose (default off), each fitness eval COLLAPSES every superposed leaf to
its best in-class usage via an optimal supply->demand assignment (brute force
<=C! within cap C=4, scipy Hungarian beyond), then scores the condensed types.
Because collapse re-types on the unmerged tree before all checks, counts /
adjacency / quality are unchanged downstream -- no Node field, no graph/operator
changes -- and default OFF is bit-identical.

- programme.py: derive_interchange_classes + interchangeable (S1-S4, locked
  thresholds R_SIZE=1.5/R_WIDTH=1.3/R_PROP=1.5, CLASS_CAP=4)
- fitness.py: collapse_superposition, _best_assignment, _usage_quality;
  superpose/superpose_class_cap conf knobs; collapse hooked into _evaluate_full
- driver.py/evolve.py: superpose flag plumbed beside leaf_sharing; --superpose
- tests/test_superposition.py: 17 tests (derivation, assignment, end-to-end)

Closes homemaker-py-9o5 (build); validation A/B is homemaker-py-xi7.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 07:08:46 +01:00
87d309771e 9o5: lock similarity thresholds, confirm C=4, defer veto hatch — spec build-ready
R_size=1.5 / R_width=1.3 / R_prop=1.5 for the interchangeable-class similarity
gate (S2); class-size cap C=4 confirmed; interchange:false veto hatch deferred
to a later fix only if auto-derivation misgroups on real configs. All open
questions resolved.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 23:20:59 +01:00
4ddb193fbf 9o5: lock collapse cadence — per-eval with class-size cap C=4
Resolve open Q1: collapse runs per fitness eval (search optimises the
condensed objective directly, removing the relaxation gap), bounded by a
derivation-time class-size cap C=4 (<=24 perms/eval). Note the collapse is a
separable linear-sum assignment, so Hungarian solves it exactly beyond the
cap if a real class ever exceeds it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 23:18:20 +01:00
bfe7a8e329 9o5: attach path-(a) superposition+collapse design spec
Spec the multi-use-leaves feature per Bruno's framing: superposition as a
search relaxation over auto-derived interchangeable equivalence classes
(requirement-similarity), condensed to specific usage at the end by
brute-forcing the in-class assignment (3 usages/3 leaves = 6 perms). Records
the reversal of the issue's 'path b preferred' note, the relaxation-gap /
0-3 search-easing prior, default-OFF baseline gate, and open Qs (collapse
cadence, similarity thresholds, veto hatch).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 23:15:22 +01:00