homemaker-py-r5a: canonicalise stale leaf-share stamps before any relabel
collapse_global's own commit could relabel a leaf back to the code its stale share_type names, making share_type == type true again and resurrecting a multiplicity credit for area never sized for it -- the commit-door companion to the iio valuation bug. dom.canonicalize_shares() drops share/share_type whenever share_type != type; called at the top of collapse_global (covers collapse_global's own commit, 2-opt, and standalone finish-time use) and _evaluate_full (covers collapse_superposition and ordinary retype mutations) so the guard is an actual invariant instead of a per-reader check. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Dq3WAXft8RszMG2CLH7VkU
This commit is contained in:
parent
522f9b1d68
commit
12247248d5
4 changed files with 105 additions and 24 deletions
|
|
@ -1,4 +1,4 @@
|
|||
{"id":"homemaker-py-2g7.4","title":"Exact shape-curve inner loop (Otten/Stockmeyer DP) replacing Nelder-Mead","description":"The classic slicing-floorplan result applied to our exact representation: each leaf's size/width/proportion constraints define a feasible-shape region; these compose bottom-up through the slicing tree as piecewise shape curves, yielding in ONE linear pass (no iteration): (a) whether ANY ratio assignment satisfies all per-leaf shape constraints, and (b) the ratios that realize a chosen point on the root curve. Today the same question costs an 80-eval NM run per child (~all of the 3M-eval budget) and answers it only approximately. Plan: (1) prototype on harbor-house-l0 with a rectangular plot approximation; (2) validate against innerloop.optimise — DP-feasible topologies must score \u003e= NM result when polished, DP-infeasible must never reach 0 shape fails under NM; (3) wire as a PRE-FILTER: prune shape-infeasible children before any native eval, and warm-start NM from DP ratios (or replace NM entirely where the plot is near-rectangular; keep NM as final polish for skew). CAVEATS to model honestly: crinkliness/access/adjacency are NOT in the DP (graph terms, not per-leaf shape) — the DP handles the size/width/proportion family only, which is fine for pruning; equal-offset skew-quad geometry means DP areas are approximate — measure the approximation error on real plots first (harbor plot is a near-rect quad). Expected payoff: 100-1000x cheaper feasibility, turning topology search into enumerate-and-prune and unlocking the racing/MAP-Elites/CP issues. Cf. §34: autodiff failed on wall-clock; this is a different attack — exactness via structure, not gradients.","acceptance_criteria":"on harbor-house-l0: DP verdict agrees with NM-polished shape-fail outcome on \u003e=95% of 200 random topologies; measured speedup \u003e=50x per feasibility decision; approximation error on the skew plot quantified","status":"open","priority":1,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-08-02T09:15:04Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:15:04Z","dependencies":[{"issue_id":"homemaker-py-2g7.4","depends_on_id":"homemaker-py-2g7","type":"parent-child","created_at":"2026-08-02T10:15:04Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-2g7.4","title":"Exact shape-curve inner loop (Otten/Stockmeyer DP) replacing Nelder-Mead","description":"The classic slicing-floorplan result applied to our exact representation: each leaf's size/width/proportion constraints define a feasible-shape region; these compose bottom-up through the slicing tree as piecewise shape curves, yielding in ONE linear pass (no iteration): (a) whether ANY ratio assignment satisfies all per-leaf shape constraints, and (b) the ratios that realize a chosen point on the root curve. Today the same question costs an 80-eval NM run per child (~all of the 3M-eval budget) and answers it only approximately. Plan: (1) prototype on harbor-house-l0 with a rectangular plot approximation; (2) validate against innerloop.optimise — DP-feasible topologies must score \u003e= NM result when polished, DP-infeasible must never reach 0 shape fails under NM; (3) wire as a PRE-FILTER: prune shape-infeasible children before any native eval, and warm-start NM from DP ratios (or replace NM entirely where the plot is near-rectangular; keep NM as final polish for skew). CAVEATS to model honestly: crinkliness/access/adjacency are NOT in the DP (graph terms, not per-leaf shape) — the DP handles the size/width/proportion family only, which is fine for pruning; equal-offset skew-quad geometry means DP areas are approximate — measure the approximation error on real plots first (harbor plot is a near-rect quad). Expected payoff: 100-1000x cheaper feasibility, turning topology search into enumerate-and-prune and unlocking the racing/MAP-Elites/CP issues. Cf. §34: autodiff failed on wall-clock; this is a different attack — exactness via structure, not gradients.","acceptance_criteria":"on harbor-house-l0: DP verdict agrees with NM-polished shape-fail outcome on \u003e=95% of 200 random topologies; measured speedup \u003e=50x per feasibility decision; approximation error on the skew plot quantified","status":"open","priority":1,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-08-02T09:15:04Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:15:04Z","dependencies":[{"issue_id":"homemaker-py-2g7.4","depends_on_id":"homemaker-py-2g7","type":"parent-child","created_at":"2026-08-02T10:15:04Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":0,"dependent_count":1,"comment_count":0}
|
||||
{"id":"homemaker-py-2g7.3","title":"Hard/soft fail tiering: 'solved' = zero hard fails","description":"Lex-by-total-count treats a crinkly wall the same as a missing room, so search polishes shape taxes instead of fixing structure — the 3M-run best still carries 'level 0/1 not connected' and wrong-level fails after 1.7M evals. Split fails into HARD (missing space, wrong/required level, level connectivity, circulation connectivity, stairs, covered-outside) and SOFT (crinkliness, proportion, size, width, edge-too-long) tiers. Outer comparator becomes (-hard, -soft, fitness); 'solved' is defined as zero hard fails. GUARDS: (1) the inner-loop 0.5^n cliff must keep protecting against trading into new fails (§4.5/§4.9 — rerun the 0/9 inner-loop-protection check); (2) rerun the §4.9 outer A/B: the scheme must not reintroduce the scalar pathology; (3) §11.4 warns comparator reshaping alone does not escape topology basins — the claim here is narrower: budget stops being spent on soft fails while hard fails remain, and reporting becomes meaningful. The tier map lives in fitness.py next to the fail emission sites so new fail strings must declare a tier. Can start before the calibration issue lands but final tier assignments should be reviewed against its findings.","acceptance_criteria":"tiered comparator behind a flag with A/B on harbor+maple (3 seeds, 20k evals): hard-fail count at budget strictly better or equal on mean, no §4.9 regression; report shows hard/soft split","status":"open","priority":1,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-08-02T09:14:14Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:14:14Z","dependencies":[{"issue_id":"homemaker-py-2g7.3","depends_on_id":"homemaker-py-2g7","type":"parent-child","created_at":"2026-08-02T10:14:14Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-2g7.2","title":"Calibrate the objective against human reference designs","description":"Score the traced human solutions (from the plan-\u003edom composer issue) and classify EVERY fail they raise as one of: (a) genuine spec violation (fix the trace or accept), (b) representation artifact (fix scoring, cf. §13.3/§13.8 share leaks), or (c) miscalibrated threshold (fix the constant/curve). Prime suspect: crinkliness — 48% of the evolved residual (§13.11), flat ~0.8/leaf tax even on squarest layouts (§13.1); if a real human plan pays it broadly, the gaussian on 1/crink is mis-tuned, not the designs. Outcome: either the human reference scores at/near 0 hard fails (objective validated, search is the gap) or a concrete list of scoring fixes. This finally makes 'the examples are solvable' a measured statement. Also record the human design's score as the per-programme target line on all future runs.","acceptance_criteria":"every fail on each human reference classified with evidence; miscalibrations filed/fixed; per-programme target scores recorded in DESIGN.md","status":"open","priority":1,"issue_type":"task","owner":"bruno@postle.net","created_at":"2026-08-02T09:14:11Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:14:11Z","dependencies":[{"issue_id":"homemaker-py-2g7.2","depends_on_id":"homemaker-py-2g7","type":"parent-child","created_at":"2026-08-02T10:14:11Z","created_by":"Bruno Postle","metadata":"{}"},{"issue_id":"homemaker-py-2g7.2","depends_on_id":"homemaker-py-2g7.1","type":"blocks","created_at":"2026-08-02T10:14:11Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-2g7.1","title":"Human reference corpus: plan-\u003edom composer + first traced human solutions","description":"There are NO human-generated plans in the corpus — every non-empty .dom is evolution output, so the system has no ground truth for what a good design scores. Build the missing pipeline: (1) a plan-\u003edom composer — input a traced rectangular partition (rooms as rects/quads with type codes, per storey), validate it, extract the binary slicing tree by recursive guillotine-cut detection, and emit a .dom (levels, heights, perimeter, divisions). Non-slicible partitions are REPORTED with the offending region rather than rejected silently — whether human plans even lie in the slicing class is itself a first-order representability finding. (2) Trace at least one human-drawn solution for harbor-house (the plateau benchmark) and one for programme-house. Practical input path: trace in Inkscape over the scan and parse SVG rects (examples/harbor-house/drawings/ already holds SVG assets), or a simple YAML room list; avoid automatic raster vectorization for now. Uses: (a) calibration ground truth for the objective, (b) search seeds, (c) representability test of the slicing-tree phenotype, (d) later, few-shot examples for the LLM repair operator.","acceptance_criteria":"composer round-trips a synthetic slicible partition to a scoring .dom; at least one human harbor-house solution traced, composed, and scored with homemaker-fitness; non-slicible input produces a diagnostic naming the unsliceable region","status":"open","priority":1,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-08-02T09:14:10Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:14:10Z","dependencies":[{"issue_id":"homemaker-py-2g7.1","depends_on_id":"homemaker-py-2g7","type":"parent-child","created_at":"2026-08-02T10:14:09Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":0,"dependent_count":1,"comment_count":0}
|
||||
|
|
@ -27,11 +27,12 @@
|
|||
{"id":"homemaker-py-1p0","title":"Geometry inner loop: full-objective equal-offset ratio optimiser","description":"DESIGN.md §5.1, §7 Phase 1. Productionise experiments/optimize_fullfitness.py into homemaker: optimise(topology, x0=None) -\u003e (geometry, fitness). DOF = equal-offset division ratios of free branches (solver.free_branches, lowest-storey cut ownership), clipped to [eps, 1-eps]. Objective = full oracle fitness (never a proxy — §4.2 falsified). Must support warm-start x0 (§5.6) and a population/batch evaluation mode so each iteration scores via one batched oracle call (§4.6).","acceptance_criteria":"Reproduces or exceeds §4.5 gains (x1.24–x1.67, no new failures) on 2f45907, candidate-002, c964435; works as a library call on any corpus .dom","status":"closed","priority":1,"issue_type":"feature","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-06-11T23:36:58Z","created_by":"Bruno Postle","updated_at":"2026-06-12T08:46:31Z","started_at":"2026-06-12T00:14:19Z","closed_at":"2026-06-12T08:46:31Z","close_reason":"innerloop.optimise() lands: batched CMA-ES sigma ladder (0.05/0.15, IPOP popsize doubling, deterministic seeding) over equal-offset free-branch ratios vs full oracle fitness; warm-start x0 supported. Acceptance vs unprojected originals: x1.65/x1.66/x1.58 against bars x1.24/x1.67/x1.59, no new failures, 46 oracle calls vs NM's 200. Two near-bar results accepted as reproduced-within-noise (1% tol) — draw spread brackets the single-NM-draw bars; approved by Bruno 2026-06-12. Gotchas: equal-offset projection of legacy unequal cuts loses fitness/adds failures (midpoint projection used); pycma seed=0 means clock-seeded.","dependencies":[{"issue_id":"homemaker-py-1p0","depends_on_id":"homemaker-py-av5","type":"blocks","created_at":"2026-06-12T00:39:33Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":3,"comment_count":0}
|
||||
{"id":"homemaker-py-8cs","title":"Experiment: warm-vs-cold start of inner loop (Lamarckian inheritance)","description":"DESIGN.md §5.6, §4.6. Warm-starting a child topology's inner loop from the parent's optimised ratios is the main lever for cutting per-topology cost (~3 min/topology cold). Apply single topology mutations to optimised corpus designs, re-optimise warm (surviving cuts keep values, new cuts get heuristic defaults) vs cold, compare oracle-call counts to convergence at equal final fitness.","acceptance_criteria":"Speedup factor measured across \u003e=10 mutated topologies; decision recorded (expect order-of-magnitude; if \u003c2x, revisit §4.6 Phase-2 scoping)","notes":"Experiment script committed (experiments/warm_vs_cold.py, 1cc86c8) and machinery validated oracle-free; one mutated child scored through the oracle OK. Waiting on homemaker-py-gp2 reference run to finish, then execute under URB_NO_OCCLUSION=1 (3 parents x 400 evals + 12 children x 2 x 200 evals, ~1.5-2 h oracle time). Default budgets: parent 400, child 200; target = evals to 95% of best final.","status":"closed","priority":1,"issue_type":"task","owner":"bruno@postle.net","created_at":"2026-06-11T23:36:58Z","created_by":"Bruno Postle","updated_at":"2026-06-12T11:44:45Z","closed_at":"2026-06-12T11:44:45Z","close_reason":"Measured (URB_NO_OCCLUSION=1, parent budget 400, child 200, 12 single mutations across 3 designs): cold start reached 95% of warm final in 0/12 cases within budget — speedup unbounded at practical budgets; warm finals beat cold finals x1.2-x4 in 12/12; 6/12 warm starts were within 95% at 1 eval (near-neutral mutations). Decision: Lamarckian warm-starting is MANDATORY in the memetic driver (homemaker-py-b39), not an optimisation; cold starts produce strictly worse geometry at equal budget. Note: 2 undivides were exactly fitness-neutral (same-type merge == Merge_Divided equivalence) — locality datum for homemaker-py-nyb.","dependencies":[{"issue_id":"homemaker-py-8cs","depends_on_id":"homemaker-py-1p0","type":"blocks","created_at":"2026-06-12T00:39:34Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-av5","title":"Batched oracle: score many .dom files per invocation","description":"oracle.py currently scores one .dom per urb-fitness.pl call (~1.65 s/dom). DESIGN.md §4.6: batching amortises Perl startup to ~0.99 s/dom and is required so population/batch optimisers can score a whole generation in one oracle call. Extend oracle.py with a batch API: write N .dom files, one perl invocation, parse N .score/.fails pairs. Keep the single-file path for compatibility.","acceptance_criteria":"Batch of 35 corpus files scores in one perl invocation; per-file results identical to single-file calls; measured s/dom reported","status":"closed","priority":1,"issue_type":"task","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-06-11T23:36:56Z","created_by":"Bruno Postle","updated_at":"2026-06-12T00:14:06Z","started_at":"2026-06-11T23:50:40Z","closed_at":"2026-06-12T00:14:06Z","close_reason":"score_batch() lands in oracle.py; 35-file corpus parity verified single-vs-batch (1e-12 rel fitness, exact fail sets); 0.98 s/dom batched vs 1.27 single, x1.30","dependency_count":0,"dependent_count":1,"comment_count":0}
|
||||
{"id":"homemaker-py-2g7.7","title":"LLM repair operator at stagnation (dom+fails -\u003e targeted compound edits)","description":"Generalize the §4.10 lesson: deceptive valleys are crossed by COMPOUND edits (move room + re-home displaced room + fix ratios atomically), which we currently hand-code one per valley (mutate_level_compound_fix). Our fail messages are semantically rich and localized ('me1 on wrong level', 'level 1 not connected', '0/rlrlr proportion') and the .dom is readable — ideal LLM input. Loop: on stagnation (no fail-tier improvement for N evals), serialize best individual + .fails + programme summary -\u003e LLM proposes 3-5 multi-step repairs as structured edit scripts (a small DSL over existing operator primitives: swap/divide/retype/rotate with explicit paths — NOT freeform dom text, so proposals are always well-formed) -\u003e apply, inner-loop, lex-accept as usual. Native fitness disposes; a bad proposal costs one child budget. Cost discipline: one LLM call ~ thousands of native evals, so plateau-only, cache by (signature, fails) key. Benchmark: the 3M-run best sat on 'level 0 not connected' + 'me1 on wrong level' for \u003e1M evals — moves a plan-reader fixes in one edit. Use claude via API (see claude-api skill); temperature\u003e0 for diverse proposals. Later extension (separate issue): AlphaEvolve-style operator-code synthesis using our existing A/B harness as the evaluator.","acceptance_criteria":"on the evolved-3M-nols-3 15-fail plateau seed: repair loop reduces hard-fail count where 1M+ blind evals did not, within \u003c=20 LLM calls; edit-DSL rejects malformed proposals; A/B at equal native-eval budget shows strictly better final fails on \u003e=2/3 seeds","status":"open","priority":2,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-08-02T09:15:54Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:15:54Z","dependencies":[{"issue_id":"homemaker-py-2g7.7","depends_on_id":"homemaker-py-2g7","type":"parent-child","created_at":"2026-08-02T10:15:53Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-2g7.9","title":"Parallel best-of-N + racing harness (use all cores, kill stragglers early)","description":"§14 measured islands \u003c= best-of-N, and the 3M runs used workers=1-2 on a 4-core box — independent seeds are the proven shape and we are not even using the local machine. Build a harness: launch N independent search_staged seeds across all cores (processes, not threads — mind the cvw id()-keyed cache bug), checkpoint fail-counts periodically, successively halve (hyperband-style: kill runs above median hard-fail count at each rung, reallocate budget to survivors). Fix/respect homemaker-py-b8g (parallel non-determinism) and homemaker-py-cvw first or work around with process isolation. This multiplies whatever eval cost the shape-curve DP issue achieves; on its own it is a free 4x locally and scales to any box. Report best + variance across seeds (the seed-variance in §12-§13 tables is huge — 78 vs 97 same config — so best-of-N is worth several levers combined).","acceptance_criteria":"harness runs N=16 seeds on 4 cores with racing; at equal total native-eval budget beats the single-seed mean on harbor by at least the observed seed spread; deterministic per-seed replay","status":"open","priority":2,"issue_type":"task","owner":"bruno@postle.net","created_at":"2026-08-02T09:15:58Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:15:58Z","dependencies":[{"issue_id":"homemaker-py-2g7.9","depends_on_id":"homemaker-py-2g7","type":"parent-child","created_at":"2026-08-02T10:15:58Z","created_by":"Bruno Postle","metadata":"{}"},{"issue_id":"homemaker-py-2g7.9","depends_on_id":"homemaker-py-b8g","type":"blocks","created_at":"2026-08-02T10:16:16Z","created_by":"Bruno Postle","metadata":"{}"},{"issue_id":"homemaker-py-2g7.9","depends_on_id":"homemaker-py-cvw","type":"blocks","created_at":"2026-08-02T10:16:15Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":2,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-2g7.7","title":"LLM repair operator at stagnation (dom+fails -\u003e targeted compound edits)","description":"Generalize the §4.10 lesson: deceptive valleys are crossed by COMPOUND edits (move room + re-home displaced room + fix ratios atomically), which we currently hand-code one per valley (mutate_level_compound_fix). Our fail messages are semantically rich and localized ('me1 on wrong level', 'level 1 not connected', '0/rlrlr proportion') and the .dom is readable — ideal LLM input. Loop: on stagnation (no fail-tier improvement for N evals), serialize best individual + .fails + programme summary -\u003e LLM proposes 3-5 multi-step repairs as structured edit scripts (a small DSL over existing operator primitives: swap/divide/retype/rotate with explicit paths — NOT freeform dom text, so proposals are always well-formed) -\u003e apply, inner-loop, lex-accept as usual. Native fitness disposes; a bad proposal costs one child budget. Cost discipline: one LLM call ~ thousands of native evals, so plateau-only, cache by (signature, fails) key. Benchmark: the 3M-run best sat on 'level 0 not connected' + 'me1 on wrong level' for \u003e1M evals — moves a plan-reader fixes in one edit. Use claude via API (see claude-api skill); temperature\u003e0 for diverse proposals. Later extension (separate issue): AlphaEvolve-style operator-code synthesis using our existing A/B harness as the evaluator.","acceptance_criteria":"on the evolved-3M-nols-3 15-fail plateau seed: repair loop reduces hard-fail count where 1M+ blind evals did not, within \u003c=20 LLM calls; edit-DSL rejects malformed proposals; A/B at equal native-eval budget shows strictly better final fails on \u003e=2/3 seeds","status":"open","priority":2,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-08-02T09:15:54Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:15:54Z","dependencies":[{"issue_id":"homemaker-py-2g7.7","depends_on_id":"homemaker-py-2g7","type":"parent-child","created_at":"2026-08-02T10:15:53Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":0,"dependent_count":1,"comment_count":0}
|
||||
{"id":"homemaker-py-2g7.6","title":"Spike: graph-first construction — adjacency-realizing slicing trees / rectangular dualization","description":"Research spike, timeboxed. Literature: rectangular dualization (planar triangulated graph -\u003e rectangular floorplan) and characterizations of slicible adjacency graphs. Our programme already IS an adjacency graph (every room wants c, plus secondary pairs); instead of mutating trees hoping adjacency emerges, construct trees that realize the required adjacency by construction — the direction §11.6/§11.7 crawled toward greedily. Deliverable is a WRITTEN assessment (DESIGN.md section): can harbor's programme graph (16 rooms + spine, 2 storeys with stacking constraint) be dualized into slicing trees, how many, and is enumeration of realizing trees tractable? Prototype only if the answer is clearly yes. Watch for: multi-storey Below-inheritance constrains both floors' trees jointly; circulation spine is a connected dominating set requirement, not a simple adjacency.","acceptance_criteria":"DESIGN.md section with go/no-go verdict, the relevant algorithms named, and complexity estimate for harbor-scale programmes","status":"open","priority":2,"issue_type":"task","owner":"bruno@postle.net","created_at":"2026-08-02T09:15:07Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:15:07Z","dependencies":[{"issue_id":"homemaker-py-2g7.6","depends_on_id":"homemaker-py-2g7","type":"parent-child","created_at":"2026-08-02T10:15:07Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-2g7.5","title":"CP-SAT type assignment for a fixed tree (replace swap/retype random walk)","description":"For a FIXED topology, assigning room codes to leaves subject to counts, required levels, adjacency-to-circulation-spine, secondary adjacencies (k1-da1, da1-o...), and share grouping is a small discrete problem (~30-70 leaves, ~16-26 codes) — well within OR-Tools CP-SAT range, solvable optimally in milliseconds. Today swap/retype/level_retype random-walk this space; §11.6/§11.7's greedy constructive assignment was the single biggest fail-count win of Phase 6, and CP-SAT is its exact big brother. Plan: model leaf-graph adjacency (geometry.leaf_graph) as fixed at seed geometry; objective = weighted satisfied adjacencies + level compliance; use as (a) seeder replacing the greedy _assign_adjacency_aware, (b) periodic 'reassign' operator inside search (the assignment analogue of ruin_recreate §23), (c) post-collapse repair. Note the §11.2 lesson: assignment quality at SEED geometry can shift after the inner loop moves ratios — re-run assignment after geometry settles (alternating minimization).","acceptance_criteria":"A/B vs greedy seeder (harbor+maple, 3 seeds, 20k evals): adjacency+access seed fails strictly lower; end-to-end mean fails no worse; reassign operator fires and is accepted at least once per run","status":"open","priority":2,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-08-02T09:15:06Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:15:06Z","dependencies":[{"issue_id":"homemaker-py-2g7.5","depends_on_id":"homemaker-py-2g7","type":"parent-child","created_at":"2026-08-02T10:15:05Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-cvw","title":"Parallel staged runs: substrate_readiness reads stale id()-keyed geometry cache in the parent process","description":"Found by the homemaker-py-zrx expert review. geometry._cache is keyed by (id(node), idx) and relies on every reader being preceded by clear_cache(). In driver.search_staged stage 1 with n_workers\u003e1 that contract breaks: the PARENT process never runs score_with_fails (children are scored in the pool workers), so its cache is never cleared, yet _rank_fitness -\u003e rank_bonus_fn -\u003e graph.substrate_readiness(ind.root) reads geometry.area()/coordinate() in the parent on every tournament/admit comparison. Evicted individuals are eventually gc'd (Node trees are parent\u003c-\u003echild reference cycles, freed by the cycle collector) while their cache entries linger; freshly unpickled worker results reuse those addresses, and substrate_readiness then serves another (dead) tree's coordinates.\n\nVerified with a probe simulating the parent's allocation pattern (unpickle jittered harbor-house trees, pop-16 eviction churn, periodic gc.collect): 24/300 readiness computations returned a corrupted value, worst absolute error 0.999 on the [0,1] readiness scale (i.e. completely wrong), and the parent cache grew without bound (38k entries after 300 children — it is never cleared for the whole run). Serial staged runs are safe (every in-process score_with_fails clears the cache between children, and live/dead id coexistence prevents collisions).\n\nImpact: silently biases stage-1 substrate selection in every parallel staged run (run_staged_search.py with WORKERS\u003e1 — the default experimental harness), and makes the bias address-dependent, i.e. NON-DETERMINISTIC across byte-identical re-runs. This is a concrete, static-read-visible candidate mechanism for part of homemaker-py-b8g's irreproducibility (it is not BLAS): it perturbs the stage-1 trajectory, not a single fixed-genome score. Reported fitness numbers are unaffected (the bonus only reorders the comparator).\n\nRecommended fix: geometry.clear_cache() at substrate_readiness entry (cheap: the readiness read is a handful of areas on the base level; serial-mode behaviour is unchanged because the cache there is already cold at that point). The durable fix for the whole bug class — also covering the (unobserved but real) gc-timing hazard in collapse_finish's cand deepcopy, probed 0/6 today only because cyclic trees outlive the deepcopy window — is to cache on the Node object itself (as Urb does, per geometry.py's own comment) or key by a per-tree epoch, so a recycled address can never alias. Also add a defensive geometry.clear_cache() at collapse_global entry (one line, zero practical cost: finish-time it is one-shot, in-search the cache was just cleared by _evaluate_full).","status":"open","priority":2,"issue_type":"bug","owner":"bruno@postle.net","created_at":"2026-08-02T08:19:18Z","created_by":"Bruno Postle","updated_at":"2026-08-02T08:19:18Z","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-r5a","title":"Stale leaf-share stamp resurrects when collapse_global commits a leaf back to its stamped code","description":"Found by the homemaker-py-zrx expert review. The homemaker-py-iio fix stops _collapse_value/_usage_quality PROBES from seeing a stale share (share\u003e1, share_type != type), but the COMMIT path still resurrects it: when collapse_global's assignment (or a 2-opt swap in _two_opt_adjacency_polish) relabels a leaf back to its stale share_type, leaf.type == share_type again, graph.leaf_share goes live, and the leaf immediately counts as k rooms with a k*target size centre — a credit the Hungarian matrix valued at 1x (the iio guard cleared the stamp for exactly that probe). The resurrected stamp then SERIALISES (dom._emit's guard passes once type == share_type), so it persists in the output.\n\nConsequences: (1) in-process eval vs dump/reload eval of the SAME tree diverge again — the exact 91f/iio divergence class, reopened through the commit door. Repro (verified today): 12x8 two-leaf tree, left leaf typed b1 carrying stale share=3/share_type=n, programme n(count 3, size 24+-5, w 3+-0.8, p 2+-0.6) + b1(count 1, size 24+-5, w 12+-0.5, p 2+-0.6), leaf_sharing+collapse_insearch on: live eval = 12 fails / score 7.43e-08; dump+reload twin = 19 fails / 7.29e-11 (twin gains '0/l size', 'missing required space n#1' + critical + 3 would-need lines; live instead has 'too many spaces: n (found 4, expected 3)'). (2) The Jacobi valuation (1x, post-iio) and the committed reality (kx) disagree, so assignments are made under one objective and scored under another; the 2-opt reward() sees the kx credit during trial swaps while the Jacobi matrix never did — the two phases of the same optimiser price the same relabel differently. (3) An ordinary retype mutation that happens to restore a leaf's old code resurrects the stamp the same way (no collapse needed), with the same live-vs-reloaded divergence.\n\nRecommended fix: canonicalise stale stamps instead of guarding readers one by one — at _evaluate_full entry (or minimally at collapse_global entry over the supply set), drop share/share_type whenever share_type is set and != type, exactly mirroring dom._emit's serialisation guard, so the in-memory tree can never disagree with its canonical dumped form. Add a dump/reload-agreement regression test in the style of test_collapse_global_dump_reload_agree_with_stale_share but driving the COMMIT (use the repro above: assignment must relabel the stamped leaf back to its stamped code). Note this slightly changes search dynamics (accidental resurrection credit disappears), so re-run a quick harbor-house sanity A/B when landing. Feeds homemaker-py-d86 (historical re-verification should use the post-fix semantics).","status":"open","priority":2,"issue_type":"bug","owner":"bruno@postle.net","created_at":"2026-08-02T08:18:36Z","created_by":"Bruno Postle","updated_at":"2026-08-02T08:18:36Z","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-cvw","title":"Parallel staged runs: substrate_readiness reads stale id()-keyed geometry cache in the parent process","description":"Found by the homemaker-py-zrx expert review. geometry._cache is keyed by (id(node), idx) and relies on every reader being preceded by clear_cache(). In driver.search_staged stage 1 with n_workers\u003e1 that contract breaks: the PARENT process never runs score_with_fails (children are scored in the pool workers), so its cache is never cleared, yet _rank_fitness -\u003e rank_bonus_fn -\u003e graph.substrate_readiness(ind.root) reads geometry.area()/coordinate() in the parent on every tournament/admit comparison. Evicted individuals are eventually gc'd (Node trees are parent\u003c-\u003echild reference cycles, freed by the cycle collector) while their cache entries linger; freshly unpickled worker results reuse those addresses, and substrate_readiness then serves another (dead) tree's coordinates.\n\nVerified with a probe simulating the parent's allocation pattern (unpickle jittered harbor-house trees, pop-16 eviction churn, periodic gc.collect): 24/300 readiness computations returned a corrupted value, worst absolute error 0.999 on the [0,1] readiness scale (i.e. completely wrong), and the parent cache grew without bound (38k entries after 300 children — it is never cleared for the whole run). Serial staged runs are safe (every in-process score_with_fails clears the cache between children, and live/dead id coexistence prevents collisions).\n\nImpact: silently biases stage-1 substrate selection in every parallel staged run (run_staged_search.py with WORKERS\u003e1 — the default experimental harness), and makes the bias address-dependent, i.e. NON-DETERMINISTIC across byte-identical re-runs. This is a concrete, static-read-visible candidate mechanism for part of homemaker-py-b8g's irreproducibility (it is not BLAS): it perturbs the stage-1 trajectory, not a single fixed-genome score. Reported fitness numbers are unaffected (the bonus only reorders the comparator).\n\nRecommended fix: geometry.clear_cache() at substrate_readiness entry (cheap: the readiness read is a handful of areas on the base level; serial-mode behaviour is unchanged because the cache there is already cold at that point). The durable fix for the whole bug class — also covering the (unobserved but real) gc-timing hazard in collapse_finish's cand deepcopy, probed 0/6 today only because cyclic trees outlive the deepcopy window — is to cache on the Node object itself (as Urb does, per geometry.py's own comment) or key by a per-tree epoch, so a recycled address can never alias. Also add a defensive geometry.clear_cache() at collapse_global entry (one line, zero practical cost: finish-time it is one-shot, in-search the cache was just cleared by _evaluate_full).","status":"open","priority":2,"issue_type":"bug","owner":"bruno@postle.net","created_at":"2026-08-02T08:19:18Z","created_by":"Bruno Postle","updated_at":"2026-08-02T08:19:18Z","dependency_count":0,"dependent_count":1,"comment_count":0}
|
||||
{"id":"homemaker-py-r5a","title":"Stale leaf-share stamp resurrects when collapse_global commits a leaf back to its stamped code","description":"Found by the homemaker-py-zrx expert review. The homemaker-py-iio fix stops _collapse_value/_usage_quality PROBES from seeing a stale share (share\u003e1, share_type != type), but the COMMIT path still resurrects it: when collapse_global's assignment (or a 2-opt swap in _two_opt_adjacency_polish) relabels a leaf back to its stale share_type, leaf.type == share_type again, graph.leaf_share goes live, and the leaf immediately counts as k rooms with a k*target size centre — a credit the Hungarian matrix valued at 1x (the iio guard cleared the stamp for exactly that probe). The resurrected stamp then SERIALISES (dom._emit's guard passes once type == share_type), so it persists in the output.\n\nConsequences: (1) in-process eval vs dump/reload eval of the SAME tree diverge again — the exact 91f/iio divergence class, reopened through the commit door. Repro (verified today): 12x8 two-leaf tree, left leaf typed b1 carrying stale share=3/share_type=n, programme n(count 3, size 24+-5, w 3+-0.8, p 2+-0.6) + b1(count 1, size 24+-5, w 12+-0.5, p 2+-0.6), leaf_sharing+collapse_insearch on: live eval = 12 fails / score 7.43e-08; dump+reload twin = 19 fails / 7.29e-11 (twin gains '0/l size', 'missing required space n#1' + critical + 3 would-need lines; live instead has 'too many spaces: n (found 4, expected 3)'). (2) The Jacobi valuation (1x, post-iio) and the committed reality (kx) disagree, so assignments are made under one objective and scored under another; the 2-opt reward() sees the kx credit during trial swaps while the Jacobi matrix never did — the two phases of the same optimiser price the same relabel differently. (3) An ordinary retype mutation that happens to restore a leaf's old code resurrects the stamp the same way (no collapse needed), with the same live-vs-reloaded divergence.\n\nRecommended fix: canonicalise stale stamps instead of guarding readers one by one — at _evaluate_full entry (or minimally at collapse_global entry over the supply set), drop share/share_type whenever share_type is set and != type, exactly mirroring dom._emit's serialisation guard, so the in-memory tree can never disagree with its canonical dumped form. Add a dump/reload-agreement regression test in the style of test_collapse_global_dump_reload_agree_with_stale_share but driving the COMMIT (use the repro above: assignment must relabel the stamped leaf back to its stamped code). Note this slightly changes search dynamics (accidental resurrection credit disappears), so re-run a quick harbor-house sanity A/B when landing. Feeds homemaker-py-d86 (historical re-verification should use the post-fix semantics).","status":"closed","priority":2,"issue_type":"bug","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-08-02T08:18:36Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:44:06Z","started_at":"2026-08-02T09:25:44Z","closed_at":"2026-08-02T09:44:06Z","close_reason":"Fixed: dom.canonicalize_shares() drops share/share_type whenever share_type != type, called at the top of collapse_global and _evaluate_full so a leaf relabelled back to its stale share_type (collapse commit, collapse_superposition, or a retype mutation) can no longer resurrect a multiplicity credit. Added regression test test_collapse_global_commit_does_not_resurrect_stale_share; confirmed via monkeypatch that it fails without the fix. Full suite (338 tests) passes; harbor-house A/B (evolved-3M/-nols/-anneal) shows identical scores pre/post-fix (no stale stamps on those files).","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-zrx","title":"Targeted expert review of core numeric/scoring path (fitness/solver/collapse) for silent correctness bugs","description":"Run a deep, expensive-model code review scoped to the core numeric/scoring\nlogic: fitness.py, solver.py, collapse_cmd.py, and the collapse_insearch\npath in innerloop.py/driver.py. Motivated by homemaker-py-iio: a stale\nleaf-share leak in collapse_global's probe valuation silently corrupted\nscores for an unknown period before being caught by manual diagnostic\nreview (see DESIGN.md §35 for the retroactive-impact writeup). That bug\nclass - subtle numeric/state bugs that don't crash, just quietly bias\nscores - is exactly what a careful full-context review with a stronger\nmodel is suited to catching, and exactly what a quick pass would miss.\n\nScope: read fitness.py, solver.py, collapse_cmd.py, and the\ncollapse_insearch code path end-to-end looking for:\n- other stale-state/leak bugs analogous to iio (shared mutable state\n reused across probes/leaves without proper reset)\n- valuation/accounting mismatches between search-time scoring and\n finish-time collapse scoring (the class of bug behind 7ua)\n- non-determinism sources under n_workers\u003e1 (b8g) if visible from a\n static read\n- anything else that would bias .score output without raising an\n exception or failing a test\n\nOut of scope: CLI wrappers, dom.py parsing, genome/operators (topology\nsearch), occlusion/daylight (2g5) - not on the numeric-correctness path.\n\nRelated: d86 (re-verify qpk/1ph historical numbers against the iio fix)\nand 7ua (false MISMATCH bug) are follow-ups from the same root cause\nclass this review is meant to catch earlier next time. This review\nshould probably run before/alongside d86 so any new findings feed into\nthe historical re-verification rather than requiring a second pass.","notes":"REVIEW COMPLETE (2026-08-02). Files read end-to-end: fitness.py, solver.py, collapse_cmd.py, innerloop.py, driver.py (collapse_insearch path), plus geometry.py/graph.py/dom.py support and evolve.py plumbing. Three confirmed bugs and one hygiene task filed:\n\n- homemaker-py-r5a (P2, CONFIRMED by minimal repro): stale leaf-share stamps resurrect when collapse_global's COMMIT relabels a leaf back to its stamped code — the iio fix guarded the probes but not the commit; live vs dump/reload evals of the same tree diverge again (12 vs 19 fails, score 7.4e-08 vs 7.3e-11 in the repro), and the resurrected stamp serialises and persists. Recommend canonicalising stale stamps at _evaluate_full (or collapse_global) entry, mirroring dom._emit.\n- homemaker-py-cvw (P2, CONFIRMED by probe): n_workers\u003e1 search_staged stage 1 — parent process never clears geometry._cache but substrate_readiness reads geometry there every ranking comparison; dead individuals' id()-keyed entries alias freshly unpickled children (24/300 readiness values corrupted, worst error ~1.0). Address-dependent selection bias; candidate mechanism for part of b8g. Serial runs safe.\n- homemaker-py-sd3 (P3, CONFIRMED on 5 evolved files): driver.collapse_best builds its evaluator with collapse_insearch=True baked in (no way to thread the run flag); the 94g keep-better guard is vacuous (base==coll 5/5, e.g. logs 12-\u003e12 where canonical shows 15-\u003e12) and a canonically fail-increasing collapse would be silently applied. Same gap in search_annealed's final rescore.\n- homemaker-py-pek (P3): fitness.py has two process_storey definitions; the first (~line 1146) is dead code silently shadowed by the second (~1452).\n\nReviewed clean (no defect found): solver.py (experiments-only, not on the scoring path; its residuals ignore share/co_type but nothing in search calls it); gaussian/truncated-e and clipped-gaussian ports; _gaussian_product combination; check_space_counts coverage arithmetic and missing-id suppression; collapse_global's pin/slot accounting, forbid handling, Jacobi synchronous update and the xcy submission-order determinism fix; merge_divided (merges only o/s leaves, so no share-stamp loss); NativeEvaluator deepcopy hygiene (per-eval clear_cache at _evaluate_full entry protects the whole in-eval path including collapse_insearch); collapse_finish's cand-deepcopy id-reuse hazard probed 0/6 (cyclic trees outlive the deepcopy window) — defensive clear recommended in cvw. Cross-links added to b8g and d86.","status":"closed","priority":2,"issue_type":"task","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-08-02T07:46:56Z","created_by":"Bruno Postle","updated_at":"2026-08-02T08:20:42Z","started_at":"2026-08-02T07:58:03Z","closed_at":"2026-08-02T08:20:42Z","close_reason":"Closed","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-iio","title":"Rescoring a dumped .dom under leaf_sharing+collapse_insearch does not reproduce the search's own reported n_fails","description":"Discovered during homemaker-py-91f. driver.search_staged's own r.best.n_fails (computed in-process via NativeEvaluator -\u003e Fitness.score_with_fails on copy.deepcopy(self.root) each eval) is NOT reproduced by dom.dump(r.best.root)+dom.load()+Fitness.score_with_fails on a fresh deepcopy, even with an IDENTICAL, fully-correct conf (leaf_sharing=True, share_edge_cap=True, collapse_insearch=True) and even within the SAME process (no cross-process/hash-seed effects -- verified PYTHONHASHSEED 0-4 all give the identical, stable, WRONG number). Concretely (harbor-house seed=0, budget=20000, full default stack): search reports 37 fails; copy.deepcopy(r.best.root) rescored immediately in-process also gives 37 (exact match, verified 5x); but dom.dump(r.best.root, f)+dom.load(f) then rescored gives a stable, reproducible 53 -- 15 extra fails, dominated by 'missing required space: m*' / 'missing m: would need adjacency/level' for a level-0 count=3 code ('m', Meeting Room) that must be getting satisfied via collapse_insearch's collapse_global relabelling in the live tree but is NOT literally present as a raw leaf.type in the tree (grep of the dumped .dom confirms no leaf typed exactly 'm' anywhere). Ruled out: hash-seed randomness (stable across PYTHONHASHSEED 0-4), naive YAML float-precision loss (dom.load+dom.dump round-trip is byte-stable once loaded), and the known separate collapse_insearch-conf-omission bug in run_staged_search.py's own rescore (homemaker-py-7ua, which produces a DIFFERENT wrong number, 55, via a different mechanism -- missing the collapse_insearch override entirely). This is a THIRD, distinct issue: even with the conf fully correct, dump/reload of the raw (pre-collapse) topology changes what collapse_global's Jacobi-relaxation/adjacency-relabelling converges to. Leading hypothesis (not yet confirmed): dom.py's _link()-reconstructed parent/below/position linkage after a fresh parse does not exactly match the linkage the live, incrementally-mutated search tree carries (module docstring notes 'multi-storey wall-stacking where an upper quad inherits its coordinates from the matching quad below' -- a below-pointer or traversal-order difference could change collapse_global's adjacency graph or leaf iteration order). Needs focused investigation with debug instrumentation inside collapse_global comparing the live vs reloaded tree's leaf order/adjacency graph on the SAME topology. Impact: any workflow that dumps a .dom under the leaf_sharing+collapse_insearch stack and later rescores it from disk (homemaker-fitness CLI, homemaker-collapse, ad-hoc diagnostics) gets a WRONG, but stable/reproducible-looking, fail count -- silently more pessimistic than what the search actually achieved. homemaker-py-91f's fail-category tally works around this by scoring driver.search_staged's r.best.root in-process, immediately, never via a dump/reload round trip (see experiments/run_and_capture_91f.py).","notes":"Correction: the rigorous historical re-verification follow-up is filed as\nhomemaker-py-d86 (not a placeholder ID as in the previous note).","status":"closed","priority":2,"issue_type":"bug","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-08-01T15:29:46Z","created_by":"Bruno Postle","updated_at":"2026-08-02T06:54:04Z","started_at":"2026-08-01T19:05:07Z","closed_at":"2026-08-01T20:07:32Z","close_reason":"Closed","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-91f","title":"Residual diagnostic on current full default construction stack","description":"Re-run the §13.1/§13.2-style per-leaf fail-breakdown diagnostic (experiments/diag_leaf_shapefail.py, diag_slack_localization.py) on the CURRENT full default stack (proportion-aware + adjacency-aware seeding, depth-balanced, leaf-sharing factor 3, interior-O odiv=3, share-aware edge cap — post homemaker-py-rq2/x3b), on harbor-house and maple-court. The last such diagnostic predates hph/rq2 (share-aware edge cap) and the erc.7 depth-balance+leaf-sharing synergy default flip, so the current floor (harbor 31.0, maple 74.0 per §13.9) has never been decomposed by failure category/leaf. This is the same read-only methodology that found leaf-sharing (erc.3), depth-balancing (erc.4), interior-O (ld2), and the edge-cap fix (hph) — DESIGN.md's own diagnostic-first discipline. Expected output: which fail category now dominates the residual, informing the next concrete construction lever (the way §13.7's edge-too-long finding directly produced hph). No code changes, no A/B — pure measurement.","design":"Reference: DESIGN.md §13.1 (erc.1), §13.2 (erc.2), §13.7 (71d.1), §13.9 (rq2). Scripts to reuse/extend: experiments/diag_leaf_shapefail.py, experiments/diag_slack_localization.py, experiments/diag_edge_too_long.py.","notes":"Progress note (2026-08-01T16:31:04+01:00): found a real reproducibility gap (filed as homemaker-py-iio, P2) -- rescoring a dumped .dom under the full leaf_sharing+collapse_insearch stack does NOT reproduce the search's own in-process reported n_fails (verified: 37 search-time vs 53 stable-but-wrong after dump+reload, same conf, same process, not hash-seed noise). Also filed homemaker-py-7ua (P3) for a narrower, separate bug: run_staged_search.py's own final sanity rescore omits the collapse_insearch override entirely. Both mean the first batch of 6 staged-search runs I did via the external run_staged_search.py + dump + reload-rescore pipeline (results in /tmp scratchpad .../91f/*.dom) are NOT trustworthy for a fail-category tally -- reloading them and rescoring inflates certain categories (missing-required-space cascade) artificially. Relaunched all 6 (harbor-house + maple-court, seeds 0/1/2, budget 20000, full default stack: leaf_sharing/leaf_share_factor=3/depth_balanced/interior_outside/outside_divisor=3/share_edge_cap) via a new script (experiments/run_and_capture_91f.py) that captures the TRUE in-process fails list immediately off driver.search_staged's r.best.root (verified this matches the search's own reported n_fails exactly, 5/5 repeats) instead of dumping+reloading. This is now running in the background; ETA ~2-2.5h total. Will tally fail categories from the *.fails.json outputs once complete.","status":"closed","priority":2,"issue_type":"task","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-08-01T10:06:37Z","created_by":"Bruno Postle","updated_at":"2026-08-01T18:37:15Z","started_at":"2026-08-01T11:19:09Z","closed_at":"2026-08-01T18:37:15Z","close_reason":"Diagnosed via real driver.search_staged runs (budget 20000, seeds 0/1/2, harbor-house + maple-court, full default stack). Fails/seed: harbor 37/33/30 (mean 33.3, vs cited 31.0); maple 82/84/78 (mean 81.3, vs cited 74.0) -- good sanity check on the stack wiring. Fail-category breakdown (combined): crinkliness 48.0%, size 20.6%, everything else (adjacency/proportion/access/edge-too-long/missing-space/connectivity/stairs) each \u003c=6%. Shape-intrinsic fails (crinkliness+size ~69%) now completely dominate the residual; construction-completeness fails are a small tail. This revises erc.1's (§13.1) old recommendation to deprioritise compactness-cuts (erc.5) in favour of leaf-sharing (erc.3) -- leaf-sharing/depth-balancing/interior-O/edge-cap are now all fully deployed as defaults, yet crinkliness is proportionally MORE dominant than ever. Recommendation written up in DESIGN.md §13.11: reopen erc.5-style compactness-aware cutting, or a crinkliness-targeted lever specifically (crinkliness:size ~2.3:1), as the next concrete construction lever. Two bugs found and filed along the way: homemaker-py-iio (P2, open) -- dumping+reloading a .dom under leaf_sharing+collapse_insearch does not reproduce the search's own in-process fail count (root cause not yet found); homemaker-py-7ua (P3, open) -- run_staged_search.py's own final sanity rescore omits the collapse_insearch override. New scripts: experiments/run_and_capture_91f.py (in-process fails capture, avoids the iio pitfall), experiments/diag_residual_91f.py (category tally from the *.fails.json sidecars).","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
|
|
@ -73,11 +74,13 @@
|
|||
{"id":"homemaker-py-nyb","title":"High-locality topology operators (mutation + subtree crossover)","description":"DESIGN.md §5, §7 Phase 2, §8.4. Mutation moves: divide/undivide leaf, swap children, rotate cut, retype leaf, per-floor delta edits, storey add/delete (cf. Urb Mutate.pm — but geometry sliding belongs to the inner loop, not the operator set). Crossover: area-matched subtree exchange (a subtree = a contiguous region, so crossover is meaningful — Crossover.pm). Operators must be high-locality: small genome change =\u003e small phenotype change, so warm-started inner loops stay cheap.","acceptance_criteria":"Each operator produces valid genomes (oracle scores them without error); locality measured (mean fitness/geometry perturbation per operator)","status":"closed","priority":2,"issue_type":"feature","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-06-11T23:37:27Z","created_by":"Bruno Postle","updated_at":"2026-06-12T13:07:37Z","started_at":"2026-06-12T12:54:23Z","closed_at":"2026-06-12T13:07:37Z","close_reason":"operators.py lands: 7 mutations + area-matched crossover, valid-by-construction via genome.encode repair. 115/115 oracle-valid children; locality measured: geom-pert 0.07-0.33 per op, fitness-pert 0.68-0.99 (0.5^n cliff flags raw moves — warm restart + penalty reshaping confirmed load-bearing). Also fixed dom._link stale below-links on structural mutation.","dependencies":[{"issue_id":"homemaker-py-nyb","depends_on_id":"homemaker-py-k2g","type":"blocks","created_at":"2026-06-12T00:39:36Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":1,"comment_count":0}
|
||||
{"id":"homemaker-py-k2g","title":"Topology genome: base-floor tree + per-floor deltas + type assignment","description":"DESIGN.md §5.2, §7 Phase 2. Genome = base-floor slicing topology (primary) + per-leaf type assignment + per-floor divide/undivide deltas (Below-inheritance as regulariser; cut owned by lowest storey where its path is divided — §10). Must round-trip to/from dom.py Node trees so the oracle and inner loop consume it directly. Includes storey count and per-floor type overrides.","acceptance_criteria":"Genome \u003c-\u003e .dom round-trip on all 35 corpus files preserves fitness; multi-storey wall stacking preserved","status":"closed","priority":2,"issue_type":"feature","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-06-11T23:37:26Z","created_by":"Bruno Postle","updated_at":"2026-06-12T12:52:34Z","started_at":"2026-06-12T10:55:21Z","closed_at":"2026-06-12T12:52:34Z","close_reason":"genome.py encode/decode lands. 35/35 oracle fitness parity after round-trip (flag-on); genome fixed-point + owned-projection tests. Dead-field discovery: corpus upper storeys carry drifted dead divisions (97) and rotations (187) — canonicalised by decode, validated fitness-neutral.","dependency_count":0,"dependent_count":1,"comment_count":0}
|
||||
{"id":"homemaker-py-d0s","title":"Experiment: inner-loop optimiser bake-off at equal oracle budgets","description":"DESIGN.md §7 Phase 1, §8.3. DOF is only ~rooms-1 (6–7 on corpus). Compare Nelder-Mead vs CMA-ES vs batched multi-start pattern search at equal oracle-call budgets, measuring fitness gained per oracle call and wall-clock (batch-friendliness matters — §4.6). Measure, don't commit blind.","acceptance_criteria":"Table of fitness-per-budget across \u003e=3 candidates; one optimiser chosen and recorded in DESIGN.md","status":"closed","priority":2,"issue_type":"task","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-06-11T23:36:59Z","created_by":"Bruno Postle","updated_at":"2026-06-13T08:48:13Z","started_at":"2026-06-12T21:22:15Z","closed_at":"2026-06-13T08:48:13Z","close_reason":"Bake-off complete: CMA-ES confirmed as Phase 1/2 optimiser. NM wins quality per eval but sequential architecture incompatible with batching (§4.6). Compass stalls on narrow valleys. Results in DESIGN.md §8.3 and experiments/bakeoff_innerloop.*","dependencies":[{"issue_id":"homemaker-py-d0s","depends_on_id":"homemaker-py-1p0","type":"blocks","created_at":"2026-06-12T00:39:35Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-2g7.10","title":"MAP-Elites archive over (hard-fail profile, leaf count, circulation fraction)","description":"Quality-diversity as the population-level answer to the §4.10 deceptive-valley problem: an archive keeps the elite per behavior niche, so 'transiently worse but structurally different' stepping stones survive — exactly what lex selection provably discards (§11.4's own analysis). DISTINCT from the failed §11.5/§11.8 niching: that kept diverse individuals under ONE selection pressure; MAP-Elites keeps the BEST individual per niche with no cross-niche competition. Descriptors to try: hard-fail category histogram (bucketed), total leaf count, circulation area fraction, storey balance. Emit from the existing genome.signature/score_with_grade machinery (kept default-off for exactly this reuse, §11.4 verdict). Blocked on the shape-curve DP: archive-filling needs cheap evals to be meaningful. Gate honestly per the ledger discipline: 3 seeds, control = current default stack.","acceptance_criteria":"A/B at equal budget (harbor+maple, 3 seeds): archive best hard-fails \u003c= default-stack best on mean; archive demonstrably contains the stepping stone for at least one accepted valley-crossing (traceable lineage)","status":"open","priority":3,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-08-02T09:16:00Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:16:00Z","dependencies":[{"issue_id":"homemaker-py-2g7.10","depends_on_id":"homemaker-py-2g7","type":"parent-child","created_at":"2026-08-02T10:15:59Z","created_by":"Bruno Postle","metadata":"{}"},{"issue_id":"homemaker-py-2g7.10","depends_on_id":"homemaker-py-2g7.4","type":"blocks","created_at":"2026-08-02T10:15:59Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-2g7.8","title":"LLM operator synthesis (AlphaEvolve-style): evolve mutation-operator code against the A/B harness","description":"Second LLM role, after the repair operator proves the plumbing: let the LLM propose new OPERATOR CODE (python functions with the mutate_* signature) and evaluate candidates with the exact experiment discipline DESIGN.md already enforces (control reproduces baseline, 3 seeds, 20k evals, verdict). The project's ledger of 20+ operator experiments with verdicts is unusually good few-shot material: feed it the §11-§13 history so it learns what already failed (niching, grading, annealing...) and why. Sandbox the generated code; acceptance purely empirical via the harness. This is compute-hungry — schedule after the shape-curve DP lands so each A/B is cheap.","acceptance_criteria":"one synthesized operator survives the standard 3-seed A/B gate on harbor or maple (mean fails strictly better, control reproduces baseline)","status":"open","priority":3,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-08-02T09:15:56Z","created_by":"Bruno Postle","updated_at":"2026-08-02T09:15:56Z","dependencies":[{"issue_id":"homemaker-py-2g7.8","depends_on_id":"homemaker-py-2g7","type":"parent-child","created_at":"2026-08-02T10:15:56Z","created_by":"Bruno Postle","metadata":"{}"},{"issue_id":"homemaker-py-2g7.8","depends_on_id":"homemaker-py-2g7.7","type":"blocks","created_at":"2026-08-02T10:15:56Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-pek","title":"fitness.py: delete the dead first process_storey definition (silently shadowed)","description":"Found by the homemaker-py-zrx expert review. class Fitness defines process_storey TWICE: the original gnw-scope version at fitness.py:1146 and the extended hgg version at fitness.py:1452. Python keeps only the second; the first ~45 lines are dead code that still reads as live. This is a silent-bug vector: an edit to the first definition (e.g. a fix to the covered-outside failure emission, which is duplicated verbatim in both) changes nothing at runtime and no test would notice. Delete the first definition (its docstring notes are preserved in the second). No behaviour change; run the suite to confirm 337 pass.","status":"open","priority":3,"issue_type":"task","owner":"bruno@postle.net","created_at":"2026-08-02T08:19:56Z","created_by":"Bruno Postle","updated_at":"2026-08-02T08:19:56Z","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-sd3","title":"driver.collapse_best bakes collapse_insearch=True into its finish evaluator, making the 94g keep-better guard vacuous","description":"Found by the homemaker-py-zrx expert review; same family as homemaker-py-7ua but in the PRODUCT (driver.py), not the experiment script. driver.collapse_best builds its evaluator as _fitness_for(str(programme_dir), leaf_sharing, superpose, multi_use=multi_use) — so collapse_insearch silently takes _fitness_for's default True. collapse_best has no collapse_insearch parameter, so evolve.py cannot thread the run's --collapse-insearch flag through even if it wanted to.\n\nConsequences, verified on 5 evolved harbor-house trees today: (1) the keep-better guard of collapse_finish is VACUOUS — base_fails is measured on a deepcopy that _evaluate_full re-collapses in-eval, so base == collapsed on 5/5 files (e.g. evolved-3M-nols-3.dom logs '12 -\u003e 12 (applied)' where the canonical evaluator shows the collapse actually did 15 -\u003e 12). The 94g safety property 'kept only if the fail count does not increase' is therefore not being checked against the true pre-collapse tree: a canonically fail-INCREASING collapse would be silently applied (collapse_global is 'monotone on harbor-house but not proven in general' per its own docstring — the guard exists precisely for that case). (2) The '[finish] collapse: N -\u003e M' log line under-reports the collapse's real effect (experiment logs quoting it understate 94g's contribution). (3) In a --no-collapse-insearch run the finish evaluator contradicts the run's objective outright — the deterministic 7ua mechanism, now in the default pipeline. (4) Minor: max_share and conn_grade are also not forwarded (matters for kpu/anneal and qi6 runs). Same pattern in search_annealed's final rescore branch: _evaluate(..., leaf_sharing=False, superpose=superpose) leaves _evaluate's collapse_insearch default True, and search_annealed has no way to pass the flag to it.\n\nOn the 5 probed files the returned tree's canonical fails happened to equal the reported number (the tree is a collapse fixpoint after iters=6 + 2-opt, so the extra in-eval collapse found nothing) — but that is not guaranteed, and the vacuous guard + misleading log line are unconditional.\n\nRecommended fix: add a collapse_insearch (and max_share/conn_grade) parameter to collapse_best, thread it from evolve.py, and make collapse_finish's keep-better measurement use a CANONICAL (collapse_insearch=False) evaluator regardless — the guard's job is to protect the canonical fail count of the written .dom, which homemaker-fitness scores with the on-disk config (no insearch override). Decide explicitly which objective the final 'best: N fails' report should quote (canonical is what the .dom.fails sidecar will say).","status":"open","priority":3,"issue_type":"bug","owner":"bruno@postle.net","created_at":"2026-08-02T08:19:41Z","created_by":"Bruno Postle","updated_at":"2026-08-02T08:19:41Z","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-d86","title":"Rigorously re-verify qpk/1ph historical numbers against the homemaker-py-iio fix","description":"homemaker-py-iio (fixed 2026-08-02) found a stale-leaf-share metadata leak\nin Fitness._collapse_value/_usage_quality that could corrupt one cell of\ncollapse_global's Hungarian assignment during any leaf_sharing+collapse\nrun -- i.e. essentially the entire \"full default stack\" used from\nhomemaker-py-x3b (leaf_sharing default-on) onward, including the very\nstudies that justified defaulting collapse_insearch on (94g, qpk/1ph, 8sh).\n\nA same-codebase fix-vs-no-fix re-run of the qpk protocol (harbor-house,\nbudget 2500, seeds 1-3) confirmed the bug demonstrably perturbs real\nper-seed outcomes under collapse_insearch=ON (2/3 seeds diverged by 5-8\nfails, non-directionally) -- see DESIGN.md §35 for full details. That\nre-run used TODAY's codebase, not the actual historical commit, and only 3\nharbor-house seeds, not the original seed sets -- so it establishes the bug\nwas real and non-trivial but does NOT establish whether 1ph's aggregate\nN=20 programme-house verdict (mean 7.95-\u003e7.10, paired t-test p~=0.028)\nwould have changed under the fix.\n\nThis issue is to do the rigorous version: check out the codebase near the\n1ph commit (~2026-07-24, \"post-qpk commits through 161\"), backport the iio\nfix there in an isolated worktree, and re-run the ACTUAL historical seed\nsets (programme-house N=20 seeds 1-20, harbor-house N=3 seeds 1-3) at the\n1ph protocol's exact parameters, comparing per-seed and aggregate results\nagainst the published numbers. Low priority: the qualitative direction of\nthe qpk/1ph conclusion is probably still right (noise is non-directional\nand the N=20 statistical margin is comfortably above the observed per-seed\nswing), this is about tightening confidence, not expecting a reversal.","notes":"homemaker-py-zrx review (2026-08-02): before re-running historical seeds, note homemaker-py-r5a — the iio fix's semantics are incomplete (stale share stamps RESURRECT when collapse_global commits a leaf back to its stamped code, diverging live vs dump/reload evals by e.g. 12 vs 19 fails in a minimal repro). Decide whether the backported 'fix' for the historical re-run includes the r5a canonicalisation or only the iio probe guard; comparing against published numbers with only the partial fix may still carry the same class of noise.","status":"open","priority":3,"issue_type":"task","owner":"bruno@postle.net","created_at":"2026-08-02T06:53:51Z","created_by":"Bruno Postle","updated_at":"2026-08-02T08:20:40Z","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-7ua","title":"run_staged_search.py final rescore omits collapse_insearch override, causing false MISMATCH under leaf-sharing","description":"experiments/run_staged_search.py's _native_score() (used for the final 're-scored (native): ... -\u003e OK/MISMATCH' sanity line) calls fitness.load_config(programme_dir) with NO overrides, but driver.search_staged's internal evaluator always runs with collapse_insearch=True (baked into driver.search's default, search_staged has no param to disable it). The script's monkeypatched fitness.load_config only injects leaf_sharing/share_edge_cap/multi_use, not collapse_insearch, so the final rescore conf silently diverges from the search-time conf whenever leaf_sharing is on (the current default stack). Observed during homemaker-py-91f: a WORKERS=4 budget=2000 harbor-house run reported best fails=38 during search but re-scored fails=34 -\u003e MISMATCH (partly parallel non-determinism per homemaker-py-b8g, but the missing collapse_insearch override is a separate, deterministic contributor). Fix: add collapse_insearch=True to the monkeypatched conf alongside leaf_sharing/share_edge_cap.","status":"open","priority":3,"issue_type":"bug","owner":"bruno@postle.net","created_at":"2026-08-01T11:32:58Z","created_by":"Bruno Postle","updated_at":"2026-08-01T11:32:58Z","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-b8g","title":"Investigate parallel/BLAS non-determinism noise source in n_workers\u003e1 runs","description":"DESIGN.md §14 (psk, island-model experiment) flagged a real, uninvestigated noise source: 'Phase A is unaffected by the probe, yet harbor seed 2 scored 71 then 73 on byte-identical re-runs -- parallel/BLAS non-determinism, the same +/-2-3 effect §12.4 flagged.' This is DISTINCT from the homemaker-py-xcy bug (ProcessPoolExecutor as_completed ordering), which was fixed and made same-worker-count parallel runs reproducible for the SEARCH TRAJECTORY. This remaining noise is at the SCORING level (a single fitness eval on a fixed genome apparently returning different fail counts across runs), plausibly numpy/scipy BLAS thread nondeterminism in the geometry/inner-loop math. It was never root-caused or fixed, and it widens the error bars on every A/B in this log run at n_workers\u003e1 (the great majority of them, since serial sweeps are expensive). Investigate: reproduce minimally (score the same frozen .dom N times under workers\u003e1), bisect whether it's BLAS threading (try OMP_NUM_THREADS=1/OPENBLAS_NUM_THREADS=1), floating-point summation order, or something else; fix or document a mitigation (e.g. pin thread count in worker processes).","design":"Reference: DESIGN.md §14 'Noise caveat (carry forward)', §12.4 (homemaker-py-xcy, the related-but-distinct trajectory-ordering bug already fixed). If the cause is BLAS thread count, the fix is likely a one-line env pin in the worker pool initializer (driver.py's ProcessPoolExecutor setup).","notes":"homemaker-py-zrx review (2026-08-02) found a concrete, non-BLAS candidate mechanism for part of this noise in PARALLEL STAGED runs: homemaker-py-cvw — substrate_readiness in the parent process reads stale id()-keyed geometry cache entries (24/300 corrupted in a churn probe, worst error ~1.0), perturbing stage-1 selection address-dependently across byte-identical re-runs. Does not explain fixed-genome single-eval divergence (if that was ever actually isolated); re-test after cvw lands before chasing BLAS.","status":"open","priority":3,"issue_type":"bug","owner":"bruno@postle.net","created_at":"2026-08-01T10:07:45Z","created_by":"Bruno Postle","updated_at":"2026-08-02T08:20:13Z","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-b8g","title":"Investigate parallel/BLAS non-determinism noise source in n_workers\u003e1 runs","description":"DESIGN.md §14 (psk, island-model experiment) flagged a real, uninvestigated noise source: 'Phase A is unaffected by the probe, yet harbor seed 2 scored 71 then 73 on byte-identical re-runs -- parallel/BLAS non-determinism, the same +/-2-3 effect §12.4 flagged.' This is DISTINCT from the homemaker-py-xcy bug (ProcessPoolExecutor as_completed ordering), which was fixed and made same-worker-count parallel runs reproducible for the SEARCH TRAJECTORY. This remaining noise is at the SCORING level (a single fitness eval on a fixed genome apparently returning different fail counts across runs), plausibly numpy/scipy BLAS thread nondeterminism in the geometry/inner-loop math. It was never root-caused or fixed, and it widens the error bars on every A/B in this log run at n_workers\u003e1 (the great majority of them, since serial sweeps are expensive). Investigate: reproduce minimally (score the same frozen .dom N times under workers\u003e1), bisect whether it's BLAS threading (try OMP_NUM_THREADS=1/OPENBLAS_NUM_THREADS=1), floating-point summation order, or something else; fix or document a mitigation (e.g. pin thread count in worker processes).","design":"Reference: DESIGN.md §14 'Noise caveat (carry forward)', §12.4 (homemaker-py-xcy, the related-but-distinct trajectory-ordering bug already fixed). If the cause is BLAS thread count, the fix is likely a one-line env pin in the worker pool initializer (driver.py's ProcessPoolExecutor setup).","notes":"homemaker-py-zrx review (2026-08-02) found a concrete, non-BLAS candidate mechanism for part of this noise in PARALLEL STAGED runs: homemaker-py-cvw — substrate_readiness in the parent process reads stale id()-keyed geometry cache entries (24/300 corrupted in a churn probe, worst error ~1.0), perturbing stage-1 selection address-dependently across byte-identical re-runs. Does not explain fixed-genome single-eval divergence (if that was ever actually isolated); re-test after cvw lands before chasing BLAS.","status":"open","priority":3,"issue_type":"bug","owner":"bruno@postle.net","created_at":"2026-08-01T10:07:45Z","created_by":"Bruno Postle","updated_at":"2026-08-02T08:20:13Z","dependency_count":0,"dependent_count":1,"comment_count":0}
|
||||
{"id":"homemaker-py-7xb","title":"Validate full winning construction stack generalises to health-centre","description":"The whole positive construction-quality stack (adjacency-aware + proportion-aware seeding, depth-balanced growth, leaf-sharing factor 3, interior-O odiv=3, share-aware edge cap) has only ever been measured end-to-end on harbor-house and maple-court (DESIGN.md §11-§13, cumulative -54%/-41% vs the leu.2 baseline per §13.7). examples/health-centre exists (built for homemaker-py-9yx, a non-synthetic ~20-room programme of a different building type -- primary care, not house/co-housing) but has only ever been used to NULL-test ruin_recreate; the positive stack itself has never been run there. Run the current default full stack (staged search, matching the §13.9/§13.10 default config) on health-centre at a comparable budget/seed count to harbor/maple's Phase-8 measurements, and report whether the fail-count reduction pattern (dominated by leaf-sharing, then depth-balance synergy, then interior-O) holds on a structurally different programme mix, or whether health-centre's room-type diversity (19 distinct codes, mostly single-instance, per §32) changes which lever dominates.","design":"Reference: DESIGN.md §13.3/§13.5/§13.6/§13.9 (the levers to validate), §32 (9yx, health-centre's construction and room-code tiering). No new code expected -- this is a measurement run with the existing default-on stack, comparable to the leu.1/§12.1 benchmark-establishment style.","status":"open","priority":3,"issue_type":"task","owner":"bruno@postle.net","created_at":"2026-08-01T10:07:29Z","created_by":"Bruno Postle","updated_at":"2026-08-01T10:07:29Z","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-fe2","title":"Experiment: 2-opt local-search polish inside collapse_insearch hot loop","description":"collapse_global's optional 2-opt adjacency polish (homemaker-py-9wi, §25) is proven positive and default-ON at finish-time (homemaker-py-cdl, §28: 46-file sweep, 0 regressions, 2 improvements incl. harbor evolved-anneal-3M 21-\u003e19). In-search collapse (collapse_insearch, homemaker-py-qpk/1ph, §20) is separately proven positive and default-ON (~11% mean fail reduction on both example programmes at N=15/20). But the two have never been combined: §28 explicitly left collapse_global's method-level local_search default OFF because 2-opt running inside the per-eval hot loop (thousands of calls per search) was 'untested and likely-costly, out of scope' for that issue -- only the one-shot finish-time cost (\u003c1s even on the largest file) was measured. This issue is the measurement: A/B collapse_insearch with local_search=True vs False (both already default-on baseline), on harbor-house and maple-court, staged search, matching the qpk/1ph protocol (equal budget, keep-better guard already monotone by construction). Report both the wall-clock cost multiplier and any fail-count effect; only recommend a default flip if positive and the cost is not prohibitive.","design":"Reference: DESIGN.md §20 (qpk), §25 (9wi), §28 (cdl) 'Where the default did NOT change' paragraph. Protocol: mirror experiments/run_qi6_ab.sh / run_lj3_qjg_ab.sh style equal-budget A/B, finish with standard --collapse, canonical homemaker-fitness re-score.","status":"open","priority":3,"issue_type":"task","owner":"bruno@postle.net","created_at":"2026-08-01T10:07:12Z","created_by":"Bruno Postle","updated_at":"2026-08-01T10:07:12Z","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-9yx","title":"Non-synthetic third example programme to isolate ruin_recreate room-count threshold","description":"y51/xyu follow-up (option b, not run by xyu). The synthetic n=10/14/18/22 sweep scales room count by duplicating already-interchangeable programme-house room codes (count: on b1/t1/b2/t2) -- the same mechanism harbor-house itself uses 'to reduce complexity'. xyu extended n=18 to N=15 seeds (DESIGN.md 31): trend weakened but did not evaporate (9.3%-\u003e6.4%, two-sided Wilcoxon p 0.098-\u003e0.059), still ambiguous. A genuinely distinct third example programme with real room-type diversity at an intermediate room count (not a duplicated-code scale-up) would avoid the interchangeable-room confound and better isolate room count as the driving variable behind the wing-rebuild-fraction hypothesis from f1d (DESIGN.md 23).","notes":"RESOLVED (2026-07-30, DESIGN.md §32): built examples/health-centre, a 19-code/\nn=20 real health-centre programme (not duplicated-count). Wilcoxon N=15 vs\nxyu's own protocol: 8W/5L/2T, mean fails 46.13-\u003e45.13, delta=2.2%, two-sided\np=0.40, one-sided p=0.20 -- a clean null, weaker even than xyu's own\ninconclusive 6.4%/p=0.059 reading at the same room count. Converges with\nharbor-house's null-to-negative result rather than y51's synthetic sweep.\nConclusion: the y51/xyu signal was substantially an artifact of the\nduplicated-interchangeable-code mechanism, not a real room-count effect.\nenable_ruin_recreate stays OFF. No further follow-up filed.\n\nNote en route: first draft of health-centre's room sizes auto-derived into\none 19-code interchange class (9o5's transitive chain) -- fixed by tiering\nroom widths with \u003e1.3x gaps at 3 boundaries into 3 bounded classes. Worth\nremembering for any future non-synthetic programme design.","status":"closed","priority":3,"issue_type":"task","assignee":"Bruno Postle","owner":"bruno@postle.net","created_at":"2026-07-29T09:07:48Z","created_by":"Bruno Postle","updated_at":"2026-07-30T07:07:09Z","started_at":"2026-07-29T14:05:00Z","closed_at":"2026-07-30T07:07:09Z","close_reason":"Closed","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
|
|
@ -118,27 +121,27 @@
|
|||
{"id":"homemaker-py-erc.6","title":"Experiment: inner-loop slack-expansion objective term","description":"Inner-loop counterpart to plot-fill construction. If Diagnostic B shows the inner loop has room to expand leaves into slack but no objective gradient to do so (the scalar rewards hitting target area but not exceeding it where slack exists), add a term/incentive so the ratio optimiser pushes leaf boundaries out to consume neighbouring slack and satisfy size, rather than parking at target.\n\nCONDITIONAL on Diagnostic B: build this only if B localizes the gap to the inner loop (room to expand, no gradient); if B shows construction targets too-small dims, prefer the plot-fill construction sibling. Must preserve the §5.4 inner-loop cliff / §4.9 lexicographic protection — the term sits where it cannot displace the fail-count ordering. A/B vs §12.2 baseline, seeds 0/1/2, 20000 evals, staged, default-OFF. Record DESIGN.md §13.6.","notes":"DEPRIORITISED by Diagnostic B (§13.2). B shows the inner loop CANNOT repair undersize: the slack is depth-driven maldistribution baked into the frozen topology, and the equal-offset ratio DOF cannot shrink a 14x leaf to feed a starved one without trading into shape fails (0.5^n cliff). Wrong DOF and wrong direction — the blocker is slicing POSITION, not a missing expansion reward. Fix belongs upstream in construction/topology (erc.4 re-scoped, erc.3). Keep as a low-priority follow-up only if a depth-balanced construction still leaves a residual size gradient the inner loop could pick up.","status":"closed","priority":4,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-06-22T23:16:24Z","created_by":"Bruno Postle","updated_at":"2026-06-28T13:22:22Z","closed_at":"2026-06-28T13:22:22Z","close_reason":"wont-fix (DESIGN §13.7): Diag B (§13.2) showed the inner loop cannot repair undersize (wrong DOF — slicing position, frozen-topology ratios). Superseded by depth-balanced construction (erc.4). Condition unmet.","dependencies":[{"issue_id":"homemaker-py-erc.6","depends_on_id":"homemaker-py-erc","type":"parent-child","created_at":"2026-06-23T00:16:23Z","created_by":"Bruno Postle","metadata":"{}"},{"issue_id":"homemaker-py-erc.6","depends_on_id":"homemaker-py-erc.2","type":"blocks","created_at":"2026-06-23T00:16:47Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-erc.5","title":"Experiment: compactness-aware cuts (minimize leaf perimeter/area)","description":"Attacks the #1 factor, crinkliness (346) — a per-leaf perimeter/area property DISTINCT from proportion (aspect ratio). Proportion-aware seeding (leu.2) sizes splits but does not bias toward balanced, square-ish subdivision. Add a KD-tree-style 'keep both children compact' cut rule (prefer the cut orientation/position that minimises summed child perimeter/area) in construction.\n\nCONDITIONAL on Diagnostic A: if A shows per-leaf shape-fail is FLAT across densities (floor intrinsic to slicing density), better cuts at the same leaf count will not pay → this should be closed wont-fix in favour of leaf-sharing. Only build if A shows shape-fail RISES with density. A/B vs §12.2 baseline, seeds 0/1/2, 20000 evals, staged, default-OFF. Record DESIGN.md §13.5.","notes":"DEPRIORITISED by erc.1 verdict (§13.1): per-leaf shape-fail flat vs slicing density and cuts already squarest (_size_divisions_from_targets picks squarest rotation) yet still ~1.8 fails/leaf =\u003e little compactness headroom at fixed leaf count. Floor is intrinsic to leaf COUNT, not cut quality. Revisit only if leaf-sharing (erc.3) underdelivers.","status":"closed","priority":4,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-06-22T23:16:21Z","created_by":"Bruno Postle","updated_at":"2026-06-28T13:22:17Z","closed_at":"2026-06-28T13:22:17Z","close_reason":"wont-fix (DESIGN §13.7): Diag A (§13.1) showed the floor is intrinsic to leaf COUNT not cut quality; revisit condition was 'only if leaf-sharing underdelivers' but leaf-sharing OVER-delivered (−32…−39%, §13.3). Condition unmet.","dependencies":[{"issue_id":"homemaker-py-erc.5","depends_on_id":"homemaker-py-erc","type":"parent-child","created_at":"2026-06-23T00:16:21Z","created_by":"Bruno Postle","metadata":"{}"},{"issue_id":"homemaker-py-erc.5","depends_on_id":"homemaker-py-erc.1","type":"blocks","created_at":"2026-06-23T00:16:43Z","created_by":"Bruno Postle","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
|
||||
{"id":"homemaker-py-2g5","title":"Rebuild occlusion/daylight/sun subsystem in Python (post-Phase-5, after optimisation fully native)","description":"DESIGN.md §6 port scope — a whole subsystem, not a term. quality_daylight (Leaf.pm:281-296) needs Urb::Misc::Sun + Urb::Field::Occlusion (+CIESky); quality_uncrinkliness also takes the occlusion object. Indoor spaces return 1 for daylight; cost is outdoor spaces + crinkliness. Port Sun_horizontal (262980-minute normalisation) and the occlusion wall set from Dom-\u003eWalls.","acceptance_criteria":"Daylight and crinkliness factors match Perl (float tolerance) across the corpus, including multi-storey cases","notes":"Re-scoped 2026-06-12: occlusion disabled in the Urb oracle instead of ported (see homemaker-py-gp2). Native fitness ships with simple crinkliness (illumination factor = 1, in homemaker-py-gnw). This issue is now the eventual Python occlusion rebuild, only after optimisation works entirely in Python. Restores outdoor-daylight and shaded-wall selection pressure.\nReframed 2026-06-17: orthogonal to epic homemaker-py-c4c. This is fitness FIDELITY (restoring daylight + shaded-wall selection pressure to match Perl), not search CAPABILITY — it changes what 'good' means, not the search's ability to find good. It will NOT improve final designs in the sense currently sought. Stays P4, deferred until the topology-search-quality epic lands and optimisation is fully native.","status":"open","priority":4,"issue_type":"feature","owner":"bruno@postle.net","created_at":"2026-06-11T23:38:25Z","created_by":"Bruno Postle","updated_at":"2026-06-17T19:14:48Z","dependency_count":0,"dependent_count":0,"comment_count":0}
|
||||
{"_type":"memory","key":"adjacency-in-binary-slicing-tree-is-structural-not","value":"Adjacency in binary slicing tree is structural, not geometric: the inner-loop NM cannot fix topological adjacency failures. Two paths exist: (1) tree-sibling adjacency — a node is adjacent to its sibling in the tree; (2) cross-zone geometric adjacency — leaves from different subtrees that happen to share a boundary. Staircase/adjacency fails require a topology mutation that changes which nodes are siblings or which zones touch. This was proved empirically on programme-house: staircase fail from rot=0 layout could not be fixed by NM but was fixed by level_retype creating a two-C topology (2026-06-14/15)."}
|
||||
{"_type":"memory","key":"homemaker-py-3l6-fix-leaf-sharing-evolve-runs","value":"homemaker-py-3l6 fix: leaf-sharing evolve runs now auto-finish before write via driver.polish_finish — unfold_shared_leaves() then a warm-started leaf_sharing=False polish search (--polish-budget, default budget//2). Makes the written .dom honest under canonical homemaker-fitness (internal==canonical when leaf_sharing off). Interrupt path forces polish_budget=0 (unfold+rescore only). This is yaa's unfold-then-polish, made automatic; Schedule B annealing is still kpu."}
|
||||
{"_type":"memory","key":"run-to-run-reproducibility-in-homemaker-layout-serial","value":"Run-to-run reproducibility in homemaker-layout: serial search (workers=1) is byte-for-byte deterministic; parallel (workers\u003e1) is now deterministic too AFTER fixing driver._run_batch to admit futures in submission order (was as_completed/completion order, bug xcy). Reproducibility holds only for a FIXED worker count — serial vs parallel differ because children-per-iteration is 1 vs n_workers (different batch granularity), which is expected, not a bug. The constructive seeder was NEVER nondeterministic: _assign_adjacency_aware has unique idx tiebreaks; comparing topologies with Python builtin hash() of the signature STRING is invalid (PYTHONHASHSEED salts str hashing per process) — use a stable hash (sha1) or genome.signature equality."}
|
||||
{"_type":"memory","key":"unfold-strategy-for-shared-leaves-homemaker-py-8iv","value":"Unfold strategy for shared leaves (homemaker-py-8iv, resolved 2026-07-16): use the BALANCED GRID (operators._grow_balanced/_size_subtree_equal), NOT circulation-aware slicing. Slicing a shared leaf perpendicular to its access edge so every child touches the corridor was implemented + A/B-tested and LOST decisively (150k-eval warm-start polish from evolved-3M: slice 41 fails/3.5e-14 vs grid 25 fails/2.4e-09, grid ahead at every milestone). Reason: k rooms all touching one wall are intrinsically thin slices; that geometric debt (proportion/long/width) is unfixable without topology change, while the grid's squarer children let local search re-route access cheaply via level_retype/place_missing/level_fix. Lesson: at the sharing-\u003eno-sharing transition, prioritise squarer children and leave access to local search; do not reintroduce slicing in Schedule B (kpu)."}
|
||||
{"_type":"memory","key":"multi-storey-staircase-consistency-when-dividing-or-retyping","value":"Multi-storey staircase consistency: when dividing or retyping a circulation (C) leaf at one level, the same structural change should be propagated to the matching leaf on ALL other storeys so the stair core path is maintained. The optimizer cannot fix staircase disruptions through trial-and-error geometry alone — it requires a synchronized multi-level operator that applies the same topology change to every storey simultaneously."}
|
||||
{"_type":"memory","key":"collapse-global-s-jacobi-adjacency-relaxation-homemaker-py","value":"collapse_global's Jacobi adjacency relaxation (homemaker-py-94g) is a synchronous per-round linear-assignment re-solve, which can 2-cycle indefinitely between two labellings that each satisfy ZERO adjacency requirements even though a permutation satisfying ALL of them exists -- proven on a minimal 4-cell chain (p1-q1-p2-q2, two disjoint adjacency pairs p1\u003c-\u003ep2/q1\u003c-\u003eq2) in test_two_opt_polish_escapes_jacobi_plateau. homemaker-py-9wi added Fitness._two_opt_adjacency_polish: a same-level pairwise-swap local search run after the Jacobi fixpoint, gated behind collapse_global(local_search=True) (default off, exposed as homemaker-collapse --local-search). Monotone by construction (a swap is kept only if it strictly increases total reward). Empirically on the 11 harbor-house evolved-*.dom/3m.dom/materialised-3M.dom layouts: 10 matched Jacobi-only exactly, 0 regressed, and evolved-anneal-3M.dom improved 21-\u003e19 fails (fixed a genuine mutual da1\u003c-\u003ek1 adjacency miss the Jacobi loop couldn't reach)."}
|
||||
{"_type":"memory","key":"ld2-13-6-interior-o-seed-diagnostic-all","value":"ld2/§13.6 interior-O seed diagnostic: ALL crinkliness fails in the constructed bal+share seed are UNDER-exposed (crink\u003c0.62, landlocked rooms with no facade + no uncovered-O neighbour) — zero over-exposed sliver fails. So the erc crinkliness residual is genuine under-daylighting, validating the interior light-well premise. Default outside_divisor=6 was too sparse (null: harbor 147-\u003e142, crinkliness even rose). odiv=3 is the seed-optimal joint setting: harbor seed fails 147-\u003e129 (-18), maple 219-\u003e206 (-14), landlocked fails drop, at cost of more leaves (harbor +4, maple +8). Because it ADDS leaves it carries the §13.4 wash-out risk; A/B to convergence pending."}
|
||||
{"_type":"memory","key":"correction-to-urb-fitness-bug-memory-bruno-2026","value":"CORRECTION to urb-fitness-bug memory (Bruno, 2026-06-12): 'C' is NOT a 'covered' type — Is_Covered is a geometric predicate (indoor space above). Urb's generic types are canonically UPPERCASE: C=circulation, O=outside, S=sahn (get_space_types qw/C O S/; corpus is 100% uppercase, never 'c'/'o' leaves). The mixed-case designs that fired the latent ratio_type first-match bug were created by homemaker's own operator type pool emitting lowercase 'c'/'o' — fixed: driver/operators now emit uppercase generics only, and class checks use t[0].lower() in 'cos'. The Urb class-sum patch stays as defensive hardening (zero impact on canonical designs). Native port (3y7/gnw): treat type classes case-insensitively, generics canonically uppercase."}
|
||||
{"_type":"memory","key":"proportion-aware-constructive-seeding-leu-2-12-2","value":"Proportion-aware constructive seeding (leu.2/§12.2): sizing seed cuts from target AREAS only regresses (thin slivers wreck aspect); you must ALSO pick each cut's rotation for child squareness. It is a convergence ACCELERATOR via a deeper local optimum around the constructed topology: wins where that topology is roughly right and budget is scarce (harbor -13%, maple -10% at 20k evals) but DELAYS small programmes where the seed must be restructured by undivide (programme-house regresses at fixed budget, yet reaches the floor given budget - speed, not asymptote). Default-on. Also: n_storeys must honour storey_minimum, not just level: keys (programme-house storey_minimum:2, all rooms level:0 - was seeded 1 storey short; cq1)."}
|
||||
{"_type":"memory","key":"urb-fitness-bug-found-fixed-2026-06-12","value":"Urb fitness bug found+fixed 2026-06-12 (patch in /home/bruno/src/urb, uncommitted): ProgrammeDriven.pm ratio_o/ratio_type grepped case-insensitively over the ratios hash and took the FIRST key — nondeterministic (x4.5 score swings) for designs with mixed-case type classes (both 'c' circulation and 'C' covered). Fixed to SUM the class (matches Is_Circulation//Is_Outside semantics); 35/35 corpus scores unchanged. CRITICAL for homemaker-py-3y7/gnw: the native port must implement class-SUM ratios. Building.pm has the same unpatched pattern (site-driven path, not used by our oracle). Also: the memetic search reward-hacked this bug before the fix — search results predating it are noise artifacts."}
|
||||
{"_type":"memory","key":"deceptive-valleys-in-topology-search-when-every-single","value":"Deceptive valleys in topology search: when every single-step mutation from a target state passes through a high-fail intermediary (e.g. level_fix displaces a room into 5+ new fails), a compound operator that atomically applies two coordinated changes can escape. Design compound operators to land on the low-fail state directly, bypassing the deceptive gradient. Programme-house example: level_compound_fix atomically moves the level-constrained room AND re-inserts the displaced room adjacent to C in one step (operators.py, 2026-06-14)."}
|
||||
{"_type":"memory","key":"experiment-harness-gotcha-the-leaf-sharing-relaxed-objective","value":"Experiment harness gotcha: the leaf-sharing RELAXED objective (§13.3) is injected ONLY by monkeypatching fitness.load_config in the parent process (run_staged_search.py / probe scripts). This is parent-process-only and does NOT propagate into ProcessPoolExecutor workers (n_workers\u003e1), which re-import fitness fresh and score under the STRICT on-disk patterns.config -\u003e r.n_fails MISMATCH (worker strict vs parent relaxed re-score). ALL §13.x floor runs were therefore SERIAL. Any future PARALLEL leaf-sharing experiment will silently mis-score until leaf_sharing lives on disk/CLI (tracked: homemaker-py-x3b). The parallel driver itself is correct; both paths score via load_config(programme_dir)."}
|
||||
{"_type":"memory","key":"homemaker-py-pythonpath-set-pythonpath-home-bruno-src","value":"homemaker-layout PYTHONPATH: package installed as 'homemaker-layout' via pip install -e . so 'import homemaker_layout' works from anywhere without PYTHONPATH. For running tests use 'python -m pytest' from project root /home/bruno/src/homemaker-layout (pyproject.toml adds src/ automatically). Never try pip show homemaker — that's the old homemaker-addon conflict."}
|
||||
{"_type":"memory","key":"island-model-psk-14-is-a-null-priming","value":"Island model (psk, §14) is a NULL: priming a population from N converged independent elites + crossover-heavy migration does not beat best-of-N at equal total budget (maple island 124 vs control 116). The child_probe instrument shows WHY: area-matched crossover across independently-converged elites almost never synthesizes (1-3 of ~64 children beat the better parent, max drop 2-5) because the slicing encoding is non-canonical (9gp), so splices are disruptive not combinatorial. Search-machinery null #3 after graded-objective and niching/restarts; residual stays geometry/shape-bound."}
|
||||
{"_type":"memory","key":"user-preference-bruno-this-is-a-fedora-system","value":"User preference (Bruno): this is a Fedora system — NEVER install Python packages via pip without asking first; always ask whether to install the rpm via dnf (e.g. python3-cma) before considering pip. Applies to any dependency additions."}
|
||||
{"_type":"memory","key":"experiment-seeding-pitfall-run-search-scaled-py-s","value":"Experiment seeding pitfall: run_search_scaled.py's default PH_SEED (c964…dom) is a FINISHED programme-house design — passing it warm-starts and floors at ~3 fails, NOT a blank-slate topology search. For blank-slate runs comparable to §11.5/§11.6 baselines, seed from examples/programme-house/init.dom (a bare undivided plot; driver bootstrap auto-triggers only on bare plots). Bit the 6zy sweep — first pass used c964 and falsely showed 3-fail floor across the whole grid."}
|
||||
{"_type":"memory","key":"programme-house-optimisation-result-2026-06-14-15","value":"Programme-house optimisation result (2026-06-14/15): best achievable is 1 fail (l1 wrong level, score ~0.005). 0 fails is geometrically impossible: l1 (min 27m²) must occupy ll (~23m²) at level 0, which eliminates the t3-adj-C provider; dividing ll into lll(l1)+llr(C) gives llr proportion ~6:1 (fails). Python memetic optimizer achieves 1 fail in 50k evals vs Perl optimiser's 2-3 fails. Winning topology: TWO C nodes at level 0 — ll(C) for t3-adj-C via geometric contact, rl(C) for staircase via tree-sibling adjacency to rrr(O). Best .dom: scratch/from-warmstart-fixed.dom and scratch/from-compound3-fixed.dom."}
|
||||
{"_type":"memory","key":"cli-tool-style-prefer-python-m-homemaker-module","value":"CLI tool style: prefer python -m homemaker.module --parameters pattern, installable via pip install -e . with pyproject.toml entry_points. Not standalone bin/ scripts."}
|
||||
{"_type":"memory","key":"strategy-decision-2026-06-12-bruno-occlusion-daylight","value":"Strategy decision 2026-06-12 (Bruno): occlusion/daylight is ORTHOGONAL to building a scalable optimiser. Disable it in Urb (env flag, homemaker-py-gp2) rather than port it; native fitness uses simple crinkliness (illumination factor = 1); rebuild occlusion in Python only after optimisation is fully native (homemaker-py-2g5, now P4). Consequence: all scores change when the flag flips — re-baseline corpus/.score, DESIGN \\$4.5 gains, gate bars at one clean boundary AFTER homemaker-py-1p0 closes; Phase-2 urb-evolve benchmark must run with the same flag."}
|
||||
{"_type":"memory","key":"urb-oracle-nondeterminism-urb-fitness-pl-output-varies","value":"Urb oracle nondeterminism: urb-fitness.pl output varies run-to-run from Perl hash-order randomisation — .fails line ORDER shuffles (compare sorted, use oracle.Score.fail_lines) and the score float can flip by ~1 ULP (compare with math.isclose rel_tol=1e-12, never ==). Not a batching artifact; affects single runs too. Matters for the Phase 3 native-fitness parity gate (homemaker-py-uxz)."}
|
||||
{"_type":"memory","key":"collapse-global-94g-and-any-label-usage-optimisation","value":"collapse_global (94g) and any label/usage optimisation CANNOT fix geometry-intrinsic fails. The harbor-house 15-fail best layout contains long-thin cells that are useless whatever room usage is assigned — their width/proportion/crinkliness fails are shape-bound, not label slack. Two consequences: (1) do not over-claim collapse gains — only ~2-3 of that layout's fails are reclaimable relabel slack, the rest are geometry- or building-level bound; (2) the threshold objective must not be tuned to 'pass' a degenerate cell via a permissive room type — a metric-pass on a physically useless space is gaming, not a fix. Real remedies for these are geometry/topology search (cell shape) and circulation placement, filed separately, not the collapse."}
|
||||
{"_type":"memory","key":"warm-x0-initialization-bug-pattern-when-a-topology","value":"warm_x0 initialization bug pattern: when a topology operator explicitly sets division ratios on a newly-created node (e.g. compound_fix sets node.division=[0.25,0.25] for t3), parent.ratios has no entry for that node (it was a leaf). warm_x0 defaults it to 0.5, corrupting the inner loop's starting point and making the operator invisible to lex comparison. Fix: only propagate child ratios for nodes where the parent node was NOT already divided; stale hidden nodes revealed by structural mutations (swap flipping b.below) must NOT contribute their pre-writeback values. See driver.py lines 259-267 (fixed 2026-06-14)."}
|
||||
{"_type":"memory","key":"9o5-multi-use-leaves-is-path-a-superposition","value":"9o5 multi-use leaves is path (a) — superposition as SEARCH RELAXATION that COLLAPSES to specific usage at the end, NOT path (b) loose-fit/no-collapse. Bruno's intent: codes with SIMILAR leaf requirements form an interchangeable equivalence class; during evolution the solver doesn't commit which leaf serves which specific usage (smoother landscape, no fighting over exact leaf usage); at the end the layout is CONDENSED to specific usages by brute-forcing the in-class assignment (3 interchangeable usages over 3 leaves = 3! = 6 combinations to check, pick best). 'Derive automatically' compatibility = requirement-similarity grouping. This reverses the issue's stated 'path b preferred' note."}
|
||||
{"_type":"memory","key":"homemaker-py-3l6-fix-leaf-sharing-evolve-runs","value":"homemaker-py-3l6 fix: leaf-sharing evolve runs now auto-finish before write via driver.polish_finish — unfold_shared_leaves() then a warm-started leaf_sharing=False polish search (--polish-budget, default budget//2). Makes the written .dom honest under canonical homemaker-fitness (internal==canonical when leaf_sharing off). Interrupt path forces polish_budget=0 (unfold+rescore only). This is yaa's unfold-then-polish, made automatic; Schedule B annealing is still kpu."}
|
||||
{"_type":"memory","key":"never-use-corpus-filenames-candidate-001-dom-candidate","value":"Never use corpus filenames (candidate-001.dom, candidate-002.dom, generated.dom, init.dom, etc.) as --output targets when running experiments. These are test fixtures. Always write experimental outputs to scratch/ or a timestamped path. Lesson from 2026-06-14: warm-start runs overwrote candidate-001/002.dom and broke graph tests."}
|
||||
{"_type":"memory","key":"9o5-multi-use-leaves-is-path-a-superposition","value":"9o5 multi-use leaves is path (a) — superposition as SEARCH RELAXATION that COLLAPSES to specific usage at the end, NOT path (b) loose-fit/no-collapse. Bruno's intent: codes with SIMILAR leaf requirements form an interchangeable equivalence class; during evolution the solver doesn't commit which leaf serves which specific usage (smoother landscape, no fighting over exact leaf usage); at the end the layout is CONDENSED to specific usages by brute-forcing the in-class assignment (3 interchangeable usages over 3 leaves = 3! = 6 combinations to check, pick best). 'Derive automatically' compatibility = requirement-similarity grouping. This reverses the issue's stated 'path b preferred' note."}
|
||||
{"_type":"memory","key":"deceptive-valleys-in-topology-search-when-every-single","value":"Deceptive valleys in topology search: when every single-step mutation from a target state passes through a high-fail intermediary (e.g. level_fix displaces a room into 5+ new fails), a compound operator that atomically applies two coordinated changes can escape. Design compound operators to land on the low-fail state directly, bypassing the deceptive gradient. Programme-house example: level_compound_fix atomically moves the level-constrained room AND re-inserts the displaced room adjacent to C in one step (operators.py, 2026-06-14)."}
|
||||
{"_type":"memory","key":"strategy-decision-2026-06-12-bruno-occlusion-daylight","value":"Strategy decision 2026-06-12 (Bruno): occlusion/daylight is ORTHOGONAL to building a scalable optimiser. Disable it in Urb (env flag, homemaker-py-gp2) rather than port it; native fitness uses simple crinkliness (illumination factor = 1); rebuild occlusion in Python only after optimisation is fully native (homemaker-py-2g5, now P4). Consequence: all scores change when the flag flips — re-baseline corpus/.score, DESIGN \\$4.5 gains, gate bars at one clean boundary AFTER homemaker-py-1p0 closes; Phase-2 urb-evolve benchmark must run with the same flag."}
|
||||
{"_type":"memory","key":"collapse-global-94g-and-any-label-usage-optimisation","value":"collapse_global (94g) and any label/usage optimisation CANNOT fix geometry-intrinsic fails. The harbor-house 15-fail best layout contains long-thin cells that are useless whatever room usage is assigned — their width/proportion/crinkliness fails are shape-bound, not label slack. Two consequences: (1) do not over-claim collapse gains — only ~2-3 of that layout's fails are reclaimable relabel slack, the rest are geometry- or building-level bound; (2) the threshold objective must not be tuned to 'pass' a degenerate cell via a permissive room type — a metric-pass on a physically useless space is gaming, not a fix. Real remedies for these are geometry/topology search (cell shape) and circulation placement, filed separately, not the collapse."}
|
||||
{"_type":"memory","key":"multi-storey-staircase-consistency-when-dividing-or-retyping","value":"Multi-storey staircase consistency: when dividing or retyping a circulation (C) leaf at one level, the same structural change should be propagated to the matching leaf on ALL other storeys so the stair core path is maintained. The optimizer cannot fix staircase disruptions through trial-and-error geometry alone — it requires a synchronized multi-level operator that applies the same topology change to every storey simultaneously."}
|
||||
{"_type":"memory","key":"proportion-aware-constructive-seeding-leu-2-12-2","value":"Proportion-aware constructive seeding (leu.2/§12.2): sizing seed cuts from target AREAS only regresses (thin slivers wreck aspect); you must ALSO pick each cut's rotation for child squareness. It is a convergence ACCELERATOR via a deeper local optimum around the constructed topology: wins where that topology is roughly right and budget is scarce (harbor -13%, maple -10% at 20k evals) but DELAYS small programmes where the seed must be restructured by undivide (programme-house regresses at fixed budget, yet reaches the floor given budget - speed, not asymptote). Default-on. Also: n_storeys must honour storey_minimum, not just level: keys (programme-house storey_minimum:2, all rooms level:0 - was seeded 1 storey short; cq1)."}
|
||||
{"_type":"memory","key":"adjacency-in-binary-slicing-tree-is-structural-not","value":"Adjacency in binary slicing tree is structural, not geometric: the inner-loop NM cannot fix topological adjacency failures. Two paths exist: (1) tree-sibling adjacency — a node is adjacent to its sibling in the tree; (2) cross-zone geometric adjacency — leaves from different subtrees that happen to share a boundary. Staircase/adjacency fails require a topology mutation that changes which nodes are siblings or which zones touch. This was proved empirically on programme-house: staircase fail from rot=0 layout could not be fixed by NM but was fixed by level_retype creating a two-C topology (2026-06-14/15)."}
|
||||
{"_type":"memory","key":"experiment-harness-gotcha-the-leaf-sharing-relaxed-objective","value":"Experiment harness gotcha: the leaf-sharing RELAXED objective (§13.3) is injected ONLY by monkeypatching fitness.load_config in the parent process (run_staged_search.py / probe scripts). This is parent-process-only and does NOT propagate into ProcessPoolExecutor workers (n_workers\u003e1), which re-import fitness fresh and score under the STRICT on-disk patterns.config -\u003e r.n_fails MISMATCH (worker strict vs parent relaxed re-score). ALL §13.x floor runs were therefore SERIAL. Any future PARALLEL leaf-sharing experiment will silently mis-score until leaf_sharing lives on disk/CLI (tracked: homemaker-py-x3b). The parallel driver itself is correct; both paths score via load_config(programme_dir)."}
|
||||
{"_type":"memory","key":"urb-fitness-bug-found-fixed-2026-06-12","value":"Urb fitness bug found+fixed 2026-06-12 (patch in /home/bruno/src/urb, uncommitted): ProgrammeDriven.pm ratio_o/ratio_type grepped case-insensitively over the ratios hash and took the FIRST key — nondeterministic (x4.5 score swings) for designs with mixed-case type classes (both 'c' circulation and 'C' covered). Fixed to SUM the class (matches Is_Circulation//Is_Outside semantics); 35/35 corpus scores unchanged. CRITICAL for homemaker-py-3y7/gnw: the native port must implement class-SUM ratios. Building.pm has the same unpatched pattern (site-driven path, not used by our oracle). Also: the memetic search reward-hacked this bug before the fix — search results predating it are noise artifacts."}
|
||||
{"_type":"memory","key":"user-preference-bruno-this-is-a-fedora-system","value":"User preference (Bruno): this is a Fedora system — NEVER install Python packages via pip without asking first; always ask whether to install the rpm via dnf (e.g. python3-cma) before considering pip. Applies to any dependency additions."}
|
||||
{"_type":"memory","key":"correction-to-urb-fitness-bug-memory-bruno-2026","value":"CORRECTION to urb-fitness-bug memory (Bruno, 2026-06-12): 'C' is NOT a 'covered' type — Is_Covered is a geometric predicate (indoor space above). Urb's generic types are canonically UPPERCASE: C=circulation, O=outside, S=sahn (get_space_types qw/C O S/; corpus is 100% uppercase, never 'c'/'o' leaves). The mixed-case designs that fired the latent ratio_type first-match bug were created by homemaker's own operator type pool emitting lowercase 'c'/'o' — fixed: driver/operators now emit uppercase generics only, and class checks use t[0].lower() in 'cos'. The Urb class-sum patch stays as defensive hardening (zero impact on canonical designs). Native port (3y7/gnw): treat type classes case-insensitively, generics canonically uppercase."}
|
||||
{"_type":"memory","key":"experiment-seeding-pitfall-run-search-scaled-py-s","value":"Experiment seeding pitfall: run_search_scaled.py's default PH_SEED (c964…dom) is a FINISHED programme-house design — passing it warm-starts and floors at ~3 fails, NOT a blank-slate topology search. For blank-slate runs comparable to §11.5/§11.6 baselines, seed from examples/programme-house/init.dom (a bare undivided plot; driver bootstrap auto-triggers only on bare plots). Bit the 6zy sweep — first pass used c964 and falsely showed 3-fail floor across the whole grid."}
|
||||
{"_type":"memory","key":"ld2-13-6-interior-o-seed-diagnostic-all","value":"ld2/§13.6 interior-O seed diagnostic: ALL crinkliness fails in the constructed bal+share seed are UNDER-exposed (crink\u003c0.62, landlocked rooms with no facade + no uncovered-O neighbour) — zero over-exposed sliver fails. So the erc crinkliness residual is genuine under-daylighting, validating the interior light-well premise. Default outside_divisor=6 was too sparse (null: harbor 147-\u003e142, crinkliness even rose). odiv=3 is the seed-optimal joint setting: harbor seed fails 147-\u003e129 (-18), maple 219-\u003e206 (-14), landlocked fails drop, at cost of more leaves (harbor +4, maple +8). Because it ADDS leaves it carries the §13.4 wash-out risk; A/B to convergence pending."}
|
||||
{"_type":"memory","key":"programme-house-optimisation-result-2026-06-14-15","value":"Programme-house optimisation result (2026-06-14/15): best achievable is 1 fail (l1 wrong level, score ~0.005). 0 fails is geometrically impossible: l1 (min 27m²) must occupy ll (~23m²) at level 0, which eliminates the t3-adj-C provider; dividing ll into lll(l1)+llr(C) gives llr proportion ~6:1 (fails). Python memetic optimizer achieves 1 fail in 50k evals vs Perl optimiser's 2-3 fails. Winning topology: TWO C nodes at level 0 — ll(C) for t3-adj-C via geometric contact, rl(C) for staircase via tree-sibling adjacency to rrr(O). Best .dom: scratch/from-warmstart-fixed.dom and scratch/from-compound3-fixed.dom."}
|
||||
{"_type":"memory","key":"run-to-run-reproducibility-in-homemaker-layout-serial","value":"Run-to-run reproducibility in homemaker-layout: serial search (workers=1) is byte-for-byte deterministic; parallel (workers\u003e1) is now deterministic too AFTER fixing driver._run_batch to admit futures in submission order (was as_completed/completion order, bug xcy). Reproducibility holds only for a FIXED worker count — serial vs parallel differ because children-per-iteration is 1 vs n_workers (different batch granularity), which is expected, not a bug. The constructive seeder was NEVER nondeterministic: _assign_adjacency_aware has unique idx tiebreaks; comparing topologies with Python builtin hash() of the signature STRING is invalid (PYTHONHASHSEED salts str hashing per process) — use a stable hash (sha1) or genome.signature equality."}
|
||||
{"_type":"memory","key":"unfold-strategy-for-shared-leaves-homemaker-py-8iv","value":"Unfold strategy for shared leaves (homemaker-py-8iv, resolved 2026-07-16): use the BALANCED GRID (operators._grow_balanced/_size_subtree_equal), NOT circulation-aware slicing. Slicing a shared leaf perpendicular to its access edge so every child touches the corridor was implemented + A/B-tested and LOST decisively (150k-eval warm-start polish from evolved-3M: slice 41 fails/3.5e-14 vs grid 25 fails/2.4e-09, grid ahead at every milestone). Reason: k rooms all touching one wall are intrinsically thin slices; that geometric debt (proportion/long/width) is unfixable without topology change, while the grid's squarer children let local search re-route access cheaply via level_retype/place_missing/level_fix. Lesson: at the sharing-\u003eno-sharing transition, prioritise squarer children and leave access to local search; do not reintroduce slicing in Schedule B (kpu)."}
|
||||
{"_type":"memory","key":"urb-oracle-nondeterminism-urb-fitness-pl-output-varies","value":"Urb oracle nondeterminism: urb-fitness.pl output varies run-to-run from Perl hash-order randomisation — .fails line ORDER shuffles (compare sorted, use oracle.Score.fail_lines) and the score float can flip by ~1 ULP (compare with math.isclose rel_tol=1e-12, never ==). Not a batching artifact; affects single runs too. Matters for the Phase 3 native-fitness parity gate (homemaker-py-uxz)."}
|
||||
{"_type":"memory","key":"warm-x0-initialization-bug-pattern-when-a-topology","value":"warm_x0 initialization bug pattern: when a topology operator explicitly sets division ratios on a newly-created node (e.g. compound_fix sets node.division=[0.25,0.25] for t3), parent.ratios has no entry for that node (it was a leaf). warm_x0 defaults it to 0.5, corrupting the inner loop's starting point and making the operator invisible to lex comparison. Fix: only propagate child ratios for nodes where the parent node was NOT already divided; stale hidden nodes revealed by structural mutations (swap flipping b.below) must NOT contribute their pre-writeback values. See driver.py lines 259-267 (fixed 2026-06-14)."}
|
||||
|
|
|
|||
|
|
@ -138,6 +138,27 @@ def levels(root: Node) -> list[Node]:
|
|||
return out
|
||||
|
||||
|
||||
def canonicalize_shares(root: Node) -> None:
|
||||
"""Drop any leaf-share stamp that no longer matches its leaf's current
|
||||
type (``share = 1``, ``share_type = None``).
|
||||
|
||||
homemaker-py-r5a: ``_emit`` already treats a leaf's share as live only
|
||||
while ``share_type == type`` (a READ-time guard), but a stale stamp left
|
||||
on the live tree can be COMMITTED back to life: if something later
|
||||
relabels the leaf back to the code its stale ``share_type`` names
|
||||
(``collapse_global``'s assignment, ``collapse_superposition``, or an
|
||||
ordinary retype mutation restoring an old code), ``share_type == type``
|
||||
becomes true again and the leaf resurrects a multiplicity credit for
|
||||
area that was never re-verified against it. Calling this before any
|
||||
relabelling pass makes the guard an actual invariant instead of a
|
||||
per-reader check, so a resurrection can never happen."""
|
||||
for lvl in levels(root):
|
||||
for leaf in lvl.leaves():
|
||||
if leaf.share_type is not None and leaf.share_type != leaf.type:
|
||||
leaf.share = 1
|
||||
leaf.share_type = None
|
||||
|
||||
|
||||
def _link(root: Node) -> None:
|
||||
lvls = levels(root)
|
||||
for lvl in lvls:
|
||||
|
|
|
|||
|
|
@ -642,6 +642,12 @@ class Fitness:
|
|||
from collections import Counter
|
||||
from . import graph as graph_mod
|
||||
|
||||
# homemaker-py-r5a: drop any stale share/share_type BEFORE this pass
|
||||
# relabels anything, so a leaf relabelled back to the code its stale
|
||||
# stamp names cannot resurrect a multiplicity credit (see
|
||||
# dom.canonicalize_shares).
|
||||
dom_mod.canonicalize_shares(root)
|
||||
|
||||
prog = self._programme or {}
|
||||
if not prog:
|
||||
return
|
||||
|
|
@ -1671,6 +1677,10 @@ class Fitness:
|
|||
from . import graph as graph_mod
|
||||
|
||||
geometry.clear_cache()
|
||||
# homemaker-py-r5a: canonicalise stale share stamps before any
|
||||
# relabelling pass (collapse_superposition/collapse_global) or read
|
||||
# can resurrect one -- see dom.canonicalize_shares.
|
||||
dom_mod.canonicalize_shares(root)
|
||||
|
||||
failures: list[str] = []
|
||||
tracking: dict = {
|
||||
|
|
|
|||
|
|
@ -249,6 +249,53 @@ def test_collapse_global_dump_reload_agree_with_stale_share(tmp_path):
|
|||
assert live_types == ["b2", "b1"]
|
||||
|
||||
|
||||
def test_collapse_global_commit_does_not_resurrect_stale_share(tmp_path):
|
||||
# homemaker-py-r5a: unlike the iio bug above (a stale stamp swaying which
|
||||
# CANDIDATE code wins), this is the COMMIT door -- collapse_global's own
|
||||
# assignment relabels the leaf back to the code its stale share_type
|
||||
# names, so share_type == type becomes true again "for real" and the
|
||||
# leaf would resurrect a share=3 credit for area that was never sized
|
||||
# for 3 rooms. Demand/sizes mirror test_relabels_to_demand_set (two
|
||||
# identically-typed leaves of different area spread across two demand
|
||||
# codes by size fit) so collapse is EXPECTED to move the smaller (left)
|
||||
# leaf onto "n" -- exactly the stale share_type stamped on it below.
|
||||
from homemaker_layout import dom
|
||||
|
||||
conf = _conf({
|
||||
"b1": {"size": [16.0, 4.0], "width": [4.0, 1.0], "proportion": [1.5, 0.5]},
|
||||
"n": {"size": [12.0, 3.0], "width": [3.5, 0.8], "proportion": [1.5, 0.5]},
|
||||
}, leaf_sharing=True)
|
||||
|
||||
def _make_root():
|
||||
root = _two_leaf_root("b1", "b1")
|
||||
left, _right = root.leaves()
|
||||
left.share = 3
|
||||
left.share_type = "n" # stale: left is currently typed "b1", not "n"
|
||||
return root
|
||||
|
||||
live = _make_root()
|
||||
Fitness(conf=conf).collapse_global(live)
|
||||
left, right = live.leaves()
|
||||
# Collapse did relabel left back onto "n" (the scenario the bug needs)...
|
||||
assert left.type == "n"
|
||||
# ...but the resurrected-looking match must not carry a share credit --
|
||||
# the stamp predates this assignment and was never re-verified.
|
||||
assert not (left.share > 1 and left.share_type == left.type)
|
||||
|
||||
path = tmp_path / "stale_share_commit.dom"
|
||||
dumped = _make_root()
|
||||
dom.dump(dumped, str(path))
|
||||
reloaded = dom.load(str(path))
|
||||
Fitness(conf=conf).collapse_global(reloaded)
|
||||
|
||||
assert [lf.type for lf in live.leaves()] == [lf.type for lf in reloaded.leaves()]
|
||||
assert [lf.share for lf in live.leaves()] == [lf.share for lf in reloaded.leaves()]
|
||||
assert (
|
||||
[lf.share_type for lf in live.leaves()]
|
||||
== [lf.share_type for lf in reloaded.leaves()]
|
||||
)
|
||||
|
||||
|
||||
def test_collapse_finish_is_keep_better_and_unmerged():
|
||||
# collapse_finish returns (tree, base, collapsed, applied); the tree it hands
|
||||
# back is unmerged (leaves still carry their divisions), and collapsed<=base.
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue