homemaker-layout/experiments/ab_9gj_ramp.py

232 lines
9.6 KiB
Python
Raw Normal View History

Rescale the underflowing crinkliness tail so the failing region has an ordering quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 06:29:02 +00:00
"""Search A/B for the crinkliness tail rescale (`homemaker-py-9gj`).
`quality_uncrinkliness` evaluates a gaussian at `x = 1/crink`, so its exponent
grows like `1/crink^2` as exposure falls. Measured over the 500k cold-start
baseline (DESIGN.md §39.12), the FAILING compact tail spans crinkliness
0.12..0.59 and quality 1e-300..1e-1 -- every value of which is numerically
zero beside a passing leaf's ~1. `crinkliness_tail="ramp"` replaces that tail,
and only that tail, with a straight line in crinkliness.
**Why stock scoring is valid here** (the §38.9 trap, and the one case the
`9gj` bead flags as exempt): the ramp is continuous at FAIL_THRESHOLD and
strictly below it, so no leaf changes which side of the threshold it is on.
The fail set is byte-identical on every corpus artefact -- asserted in
`tests/test_fitness_crinkliness_tail.py`, not assumed here. An arm therefore
cannot win by deleting a fail category, and both arms are scored under stock.
**Two experiment shapes**, because they answer different questions:
--start plateau (default) seed each run from that programme's
`coldstart-500000-s<k>.dom`. This is the ESCAPE
test the bead asks for: the ramp has signal only
where a search has already built partially-lit
rooms, and the corpus `init.dom` files show a
+0.000% score delta -- there is nothing for it to
grade at the start of a search.
--start init cold start, for comparison.
Pairing is on (starting layout, RNG seed), so `--seeds N` over `--starts M`
gives N*M paired samples per programme.
Usage::
python experiments/ab_9gj_ramp.py --budget 8000 --seeds 2 --starts 3
python experiments/ab_9gj_ramp.py --start init --budget 8000 --seeds 6
Sharding, because one run is minutes and the job list is 4x that. Each shard
keeps BOTH arms of a pair together, so the two halves of a comparison never
land on differently-loaded processes::
for i in 0 1 2 3; do
python experiments/ab_9gj_ramp.py --budget 8000 --seeds 2 --starts 3 \
--shard $i --nshards 4 &
done; wait
python experiments/ab_9gj_ramp.py --report
A POWERED run needs more than the in-session pilot could afford. §39.12 puts
harbor's minimum detectable difference at n=3 at 13.7 fails; the plateau-escape
deltas here are single-digit, so budget for n >= 8 pairs per programme
(`--seeds 3 --starts 3` gives 9) and a budget large enough for either arm to
move at all -- the pilot's 8000 evals is 1.6% of what produced the plateau::
for i in $(seq 0 3); do
python experiments/ab_9gj_ramp.py --budget 100000 --seeds 3 --starts 3 \
--shard $i --nshards 4 &
done; wait
"""
from __future__ import annotations
import argparse
import collections
import copy
import csv
import time
from pathlib import Path
from homemaker_layout import dom as dom_mod
from homemaker_layout import driver, fitness
CORPUS = ["examples/harbor-house", "examples/maple-court"]
ARMS = ["gaussian", "ramp"]
def _with_tail(tail: str):
"""Patch `fitness.load_config` so every evaluator built during the run --
the driver's, the inner loop's, the seeder's -- sees `crinkliness_tail`.
`driver.search` has no parameter for it and `driver._fitness_for` is
lru_cached, so the cache is cleared around the patch (see ab_ssz_search).
"""
orig = fitness.load_config
def patched(directory, overrides=None):
ov = dict(overrides or {})
ov["crinkliness_tail"] = tail
return orig(directory, overrides=ov)
return orig, patched
def tiers(fails) -> tuple[int, int]:
c = collections.Counter(fitness.classify_fail_tier(f) for f in fails)
return c["hard"], c["soft"]
def run_arm(progdir: str, start: Path, seed: int, tail: str, budget: int,
child_budget: int) -> dict:
orig, patched = _with_tail(tail)
fitness.load_config = patched
driver._fitness_for.cache_clear()
t0 = time.perf_counter()
try:
res = driver.search(dom_mod.load(str(start)), progdir, budget=budget,
seed=seed, child_budget=child_budget, n_workers=1)
root = copy.deepcopy(res.best.root)
finally:
fitness.load_config = orig
driver._fitness_for.cache_clear()
conf, cost = orig(progdir, overrides={"leaf_sharing": True,
"collapse_insearch": True})
score, fails = fitness.Fitness(conf, cost).score_with_fails(copy.deepcopy(root))
h, s = tiers(fails)
return dict(programme=Path(progdir).name, start=start.name, seed=seed,
tail=tail, hard=h, soft=s, total=h + s, score=score,
elapsed_s=round(time.perf_counter() - t0, 1))
def starting_points(progdir: str, kind: str, n: int) -> list[Path]:
"""The distinct layouts to start from.
`--starts` applies to plateau mode only: there is exactly one `init.dom`,
and repeating it would give several jobs the same (start, seed) pairing key,
which `_report` would silently collapse to one. In init mode the RNG seed
is the only sampling dimension, so use `--seeds`.
"""
d = Path(progdir)
if kind == "init":
return [d / "init.dom"]
found = sorted(d.glob("coldstart-500000-s*.dom"))[:n]
if not found:
raise SystemExit(f"no coldstart-500000-s*.dom in {d} for --start plateau")
return found
def main() -> None:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--budget", type=int, default=8000)
ap.add_argument("--child-budget", type=int, default=80)
ap.add_argument("--seeds", type=int, default=2, help="RNG seeds per start")
ap.add_argument("--starts", type=int, default=3,
help="plateau layouts to start from (plateau mode only)")
ap.add_argument("--start", choices=("plateau", "init"), default="plateau")
ap.add_argument("--corpus", nargs="+", default=CORPUS)
ap.add_argument("--out", default="experiments/results/ab_9gj_ramp.csv")
ap.add_argument("--shard", type=int, default=0,
help="run only jobs i where i %% nshards == shard")
ap.add_argument("--nshards", type=int, default=1,
help="split the job list across N processes; each writes "
"<out>.shard<i>. Use --report to merge and analyse.")
ap.add_argument("--report", action="store_true",
help="merge <out>.shard* (or <out>) and print the analysis "
"only -- runs nothing")
args = ap.parse_args()
out = Path(args.out)
out.parent.mkdir(parents=True, exist_ok=True)
if args.report:
rows = []
shards = sorted(out.parent.glob(out.name + ".shard*")) or (
[out] if out.exists() else [])
for sh in shards:
with sh.open() as fh:
for r in csv.DictReader(fh):
for k in ("total", "hard", "soft", "seed"):
r[k] = int(r[k])
r["score"] = float(r["score"])
rows.append(r)
if len(shards) > 1:
with out.open("w", newline="") as fh:
w = csv.DictWriter(fh, fieldnames=list(rows[0]))
w.writeheader()
w.writerows(rows)
_report(rows, args.corpus)
print(f"\nmerged {len(rows)} runs from {len(shards)} file(s) into {out}")
return
# Build the whole job list first so sharding is deterministic and every
# (start, seed) pair keeps BOTH arms in the same shard -- the comparison is
# paired, and splitting a pair across processes would let machine load
# differ between the two halves of one pair.
jobs = []
for progdir in args.corpus:
for start in starting_points(progdir, args.start, args.starts):
for seed in range(args.seeds):
jobs.append((progdir, start, seed))
rows: list[dict] = []
dest = (out if args.nshards == 1
else out.with_name(out.name + f".shard{args.shard}"))
for i, (progdir, start, seed) in enumerate(jobs):
if i % args.nshards != args.shard:
continue
for tail in ARMS:
r = run_arm(progdir, start, seed, tail, args.budget,
args.child_budget)
rows.append(r)
print(f" {r['programme']:<14} {start.name:<26} seed={seed} "
f"{tail:<9} {r['hard']}h/{r['soft']}s = {r['total']:3d} "
f"score {r['score']:.4g} {r['elapsed_s']}s", flush=True)
with dest.open("w", newline="") as fh:
w = csv.DictWriter(fh, fieldnames=list(rows[0]))
w.writeheader()
w.writerows(rows)
if args.nshards == 1:
_report(rows, args.corpus)
print(f"\nwrote {dest}")
def _report(rows, corpus) -> None:
"""homemaker-py-tco: state what this N could resolve, beside the result."""
from ab_report import format_report, paired_report
for progdir in corpus:
name = Path(progdir).name
by: dict[tuple, dict] = {}
for r in rows:
if r["programme"] == name:
by.setdefault((r["start"], r["seed"]), {})[r["tail"]] = r["total"]
keys = sorted(k for k, v in by.items() if len(v) == 2)
if len(keys) < 2:
continue
print(f"\n--- {name}: ramp vs gaussian (stock-scored total fails) ---")
print(format_report(paired_report(
[by[k]["gaussian"] for k in keys], [by[k]["ramp"] for k in keys],
"gaussian", "ramp")))
if __name__ == "__main__":
main()