homemaker-layout/examples/maple-court/coldstart-500000-s0.log
Claude e5eb397b52
Checkpoint long searches; the cold-start runs were lost to a reclaimed box
All four 500k runs died about 10 minutes in when the container was
reclaimed. No SIGTERM fired, so no .dom was written and 0 of 12 runs
completed. My plan committed results per finished run, which protected
nothing because no run reached its commit point. The bad assumption was
reading "reclaimed after inactivity" as CPU inactivity; it is conversation
inactivity, and background compute does not hold the box open.

Progress reached before the loss (from the tracked logs): harbor 24,960
evals / 40 fails, maple 14,880 / 79, health-centre 25,920 / 33,
programme-house 138,800 / 2.

The underlying gap is not environmental: a search's only output lands at
the very end or on SIGTERM, so ANY abrupt loss -- reclaimed container, OOM,
power cut -- takes the whole run with it. On a 3M-eval search that is 2.4
days of compute with no recoverable artefact.

  - driver.search gains checkpoint=/checkpoint_every=: the current best is
    handed to a callback at most every N evals. Rate-limited by evals, not
    improvements, which come in bursts early. A failing checkpoint is logged
    and swallowed -- losing a checkpoint is bad, losing the search because a
    checkpoint failed is worse.
  - homemaker-evolve --checkpoint-every N writes <out>.dom.checkpoint via
    mkstemp + os.replace, so a crash can never catch it half-written. It is
    deliberately NOT the output path: a checkpoint is a leaf-sharing run's
    internal best, dishonest under the canonical scorer until the finish
    stage unfolds it (homemaker-py-3l6), and must not be mistaken for the
    finished article.
  - Verified the written checkpoint re-loads as a valid .dom.

Default off, so behaviour is unchanged without the flag.

Lint at parity (46); tests 372 passed (3 new), same 2 pre-existing failures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-08-29 05:43:58 +00:00

42 lines
2.2 KiB
Text

WARNING: 338 m2 per storey needs ~70 m of daylit wall; the plot's non-private perimeter gives 55 m. Roughly 22 m2 of courtyard closes the gap. (DESIGN.md §38.3)
seed : /home/user/homemaker-layout/examples/maple-court/init.dom
programme : maple-court
budget : 500000
pop : 16
child_budget : 80
workers : 1
rng seed : 0
leaf sharing : True (factor=3)
superpose : False
multi_use : False
conn grade : False
use tiers : False
bridge circulation : False
ruin recreate : False
collapse in-search : True
shapecurve warmstart : False
shapecurve prune : False
output : /home/user/homemaker-layout/examples/maple-court/coldstart-500000-s0.dom
[ 80 evals] best 1.5923e-44 (fails 136) via construct/0
[ 160 evals] best 4.80951e-39 (fails 119) via construct/1
[ 1360 evals] best 2.33063e-35 (fails 108) via level_compound_fix noop
[ 1600 evals] best 4.62166e-35 (fails 107) via undivide 0/lrrl
[ 2400 evals] best 2.12969e-34 (fails 105) via core_undivide noop
[ 2640 evals] best 8.34232e-34 (fails 103) via crossover lrlrr<->lrrll
[ 3120 evals] best 1.18679e-33 (fails 102) via retype 2/rrrr->ws1
[ 3520 evals] best 2.83542e-32 (fails 98) via core_undivide noop
[ 4240 evals] best 5.90625e-32 (fails 97) via swap 0/lrlr
[ 4800 evals] best 6.75777e-32 (fails 97) via core_divide lrlr (2 floors)
[ 5120 evals] best 9.68203e-30 (fails 90) via core_divide lrlr (2 floors)
[ 6720 evals] best 1.06032e-29 (fails 90) via crossover lrll<->lrll
[ 6800 evals] best 2.32821e-29 (fails 89) via swap 0/lrrr
[ 7440 evals] best 9.83875e-29 (fails 87) via divide 0/rlll
[ 7520 evals] best 1.98761e-28 (fails 86) via core_undivide noop
[ 8960 evals] best 3.45107e-27 (fails 82) via level_fix ef1: lvl2/lrlrr → lvl0/rllr
[ 9600 evals] best 3.48472e-27 (fails 82) via core_undivide noop
[ 12320 evals] best 3.50776e-27 (fails 82) via level_compound_fix noop
[ 12640 evals] best 3.62472e-27 (fails 82) via swap 2/rlll
[ 13440 evals] best 3.64565e-27 (fails 82) via core_undivide noop
[ 14320 evals] best 3.67293e-27 (fails 82) via retype 0/rllll->py
[ 14640 evals] best 1.36771e-26 (fails 80) via level_compound_fix py: lvl0 → lvl1/lrr
[ 14880 evals] best 2.61906e-26 (fails 79) via divide 1/llrr