38.18 confirmed the 1ph default-flip was sound for its own era, but that
measurement predates three changes to the objective it was measured
against -- 39.4's namespace fix, 38.10/38.11's per-space crinkliness, and
38.12's missing-space cascade -- and collapse_insearch runs collapse_global
inside every eval, valued against exactly the factors those touched. A
default carried on a superseded measurement is an assumption, not a result.
Re-ran the 1ph protocol as published on the current codebase:
N OFF ON W/L/T diff t p
published 1ph 20 7.95 7.10 11/6/3 +0.85 2.38 0.028
current objective 20 7.85 7.15 10/7/3 +0.70 1.82 0.069
current objective 40 7.60 7.03 21/14/5 +0.57 2.01 0.045
current objective 60 7.58 7.02 29/19/12 +0.57 2.45 0.017
Verdict: the default STANDS. At N=60, mean diff +0.567 fails/seed, paired
t=2.454 (df=59), p=0.0171 exact, 95% CI [+0.105, +1.029] excluding zero;
Wilcoxon signed-rank cross-check agrees (p=0.0138), which matters because
fail counts are small integers and normality is not obvious.
Two caveats. The effect is about a third smaller than published (+0.57 vs
+0.85) -- partly regression from a lucky N=20 draw, partly plausible real
erosion, since several fails collapse_global used to clear have been
redefined out of existence or made harder.
More usefully: the published N=20 can no longer detect its own effect. At
exactly that sample size the current answer is p ~= 0.069, a null by the
conventional threshold. Had I stopped at N=20 the honest report would have
been "the 1ph verdict no longer reproduces" and the default would have
looked unjustified. It took N=60 to resolve. That is the 8sh/1ph/qi6/lj3
pattern this log warns about, now biting the flagship result itself: any
future re-validation of this default needs N >= 40.
20 annotated in place so a reader of the original claim sees the current
figure. Harness takes a seed range now (APPEND=1 to extend a sweep).
Closes homemaker-py-ioe.
Lint at parity (46); tests 384 passed, 0 failed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
|
||
|---|---|---|
| .beads | ||
| .claude | ||
| examples | ||
| experiments | ||
| src/homemaker_layout | ||
| tests | ||
| .gitignore | ||
| AGENTS.md | ||
| CLAUDE.md | ||
| DESIGN.md | ||
| pyproject.toml | ||
| README.md | ||
homemaker-layout
Programme-driven building-layout search over slicing trees. A clean-room Python successor to the Perl Urb project, intended to eventually be 100% Python.
Why a rewrite
Urb represents a building as a binary slicing tree where room sizes are derived top-down from division ratios. That makes room area an emergent property of every cut above it, which:
- gives the genome low locality (a cut near the root rescales every descendant),
- makes target room sizes nearly impossible to hit, so the gaussian size penalty dominates fitness, and
- defeats crossover (transplanted subtrees lose their proportions).
homemaker inverts this: leaves carry target dimensions from the programme and division ratios are solved bottom-up for a fixed topology. The evolutionary search then only explores topology + types + adjacency.
Phase plan
Solver experiment: port Urb's geometry, re-solve ratios from programme targets, score the result against the original via the Perl oracle.✓Native Python fitness (retire the Perl oracle).✓Memetic search: canonical slicing genome + high-locality operators + Nelder-Mead inner loop.✓Penalty reshaping: lexicographic✓(-n_fails, fitness)outer-search comparison.Representation upgrade: canonical slicing encoding + bottom-up shape feasibility, scaled to larger programmes.✓- Search-quality experiments (current): a long running series of
opt-in levers tried against the
harbor-house,health-centre, andprogramme-houseexample corpora — leaf-sharing, finish-time cell→room collapse, ruin-and-recreate LNS, 2-opt polish, multi-use/co-located leaves, adjacency-graph and bubble-diagram fitness signals, and more. Most of these are negative/null results kept as opt-in flags or reference code rather than defaults. SeeDESIGN.md§11 onward for the full, numbered experiment log with methodology and results for each.
Layout
src/homemaker_layout/dom.py— read/write Urb.domYAML into aNodetree.src/homemaker_layout/geometry.py— faithful port of Urb's top-down geometry.src/homemaker_layout/programme.py— parsepatterns.configspace requirements.src/homemaker_layout/solver.py— bottom-up ratio solve (scipy).src/homemaker_layout/fitness.py— native Python fitness evaluator.src/homemaker_layout/fitness_cmd.py—homemaker-fitnessCLI (drop-in forurb-fitness.pl).src/homemaker_layout/collapse_cmd.py—homemaker-collapseCLI: finish-time global cell→room relabel of a.dom.src/homemaker_layout/graph.py— leaf-adjacency graph for programme-driven checks.src/homemaker_layout/genome.py— topology genome: base-floor tree + per-storey deltas.src/homemaker_layout/operators.py— high-locality mutation and subtree crossover.src/homemaker_layout/innerloop.py— ratio optimisation inner loop (Nelder-Mead / CMA-ES).src/homemaker_layout/driver.py— memetic search outer loop.src/homemaker_layout/evolve.py—homemaker-evolveCLI entry point.src/homemaker_layout/oracle.py— legacy Perl shim, kept for cross-validation only.src/homemaker_layout/bubble.py— 3D bubble-diagram adjacency fitness-signal prototype (DESIGN.md §27); validated null, not wired intofitness.py— reference only.
Room codes and reserved names
Leaf types live in three namespaces that share a first character. Only the first is enforced; the other two are conventions the fitness function reads, so a room's spelling can change how it is scored.
1. Generic structural types — C, O, S (reserved). The leaves the
search itself creates: C circulation, O outside, S sahn (an outside court
that also serves as circulation). Always uppercase. A programme code spelled
exactly C, O or S is rejected at load.
2. Programme room codes — anything else, lowercase. k1, b1, cr1,
of, and single-character codes like r or t. These may start with any
letter: since DESIGN.md §39.4 the generic tests match C/O/S exactly, so
naming a room cr1 no longer makes it circulation. (Before that fix it did —
and silently dropped it from the required-space check entirely.)
3. Access requirements — the usage: attribute. Every space declares one
of living, kitchen, bedroom, toilet, utility, none. Mandatory, no
fallback, and a missing or unknown value is a load error. It replaced a
first-character convention (b/t/l/k) under which a room silently
inherited another room's connectivity rules from its spelling — la1 "Laundry
Room" was trimmed as a living room (DESIGN.md §39.7).
spaces:
la1:
usage: utility # controlled, drives engine behaviour
name: Laundry Room # free text, building-specific
A usage value exists only where the engine treats it differently, so the vocabulary is closed: a new access class means new code, not new config. Check a programme with:
python experiments/audit_programme_config.py
which reports reserved-name collisions, the usage class each code picks up, and whether each room's size/width/proportion/crinkliness targets are mutually satisfiable at all.