quality_uncrinkliness evaluates a gaussian at x = 1/crink, so its exponent grows like 1/crink^2 and underflows a double to exactly zero below crink ~ 1/15. Measured over the twelve 500k cold-start runs (39.12): 430 leaves carry a minimum-exposure requirement, 112 fail it, and those 112 span quality 1e-300..1e-1 while contributing 0.034% of total value on 23% of the floor area. Every value in that range is numerically zero beside a passing leaf's ~1, so the search cannot rank two layouts that differ only in how exposed their under-lit rooms are. This is wider than the bead's diagnosis (a flat 0.0 for zero-exposure leaves) and it explains why 38.1's `floor` mode measured as a no-op: max(q, 0.01) maps 110 of the 112 onto one constant, replacing a flat zero with a flat 0.01. crinkliness_tail="ramp" (default OFF, "gaussian" is stock) replaces the tail -- only the tail, only below FAIL_THRESHOLD, only on the compact side -- with a straight line in crinkliness meeting the gaussian exactly at the crossing. _crink_at_fail_threshold inverts the gaussian there using the same truncated _E the factor is evaluated with. Deliberately conservative: nothing at or above FAIL_THRESHOLD moves, so no calibration changes and no leaf crosses the threshold. The fail set is byte-identical on all 21 committed corpus artefacts, the four init.dom seeds included -- asserted in tests/test_fitness_crinkliness_tail.py, not assumed. That invariance is also what makes it legal to score both arms of the A/B under stock (the 38.9 trap's one exemption). A fully buried leaf still scores exactly 0; this restores an ordering within the failing region, it does not forgive it. Composing with 38.1's superseded modes is refused, since both rewrite the same tail. Score effect on the baseline artefacts: +0.3%..+2.8% on harbor and maple, exactly +0.000% on health-centre, programme-house, and every init.dom -- a programme with no partially-exposed failing rooms has nothing to grade, and neither does any starting layout. The ramp is a mid-search signal by construction, so experiments/ab_9gj_ramp.py defaults to seeding each run from a 500k plateau artefact rather than cold. The module-level math import replaces a now-redundant local one. DESIGN.md 39.13 and the A/B verdict follow in a separate commit. Refs homemaker-py-9gj. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
231 lines
9.6 KiB
Python
231 lines
9.6 KiB
Python
"""Search A/B for the crinkliness tail rescale (`homemaker-py-9gj`).
|
|
|
|
`quality_uncrinkliness` evaluates a gaussian at `x = 1/crink`, so its exponent
|
|
grows like `1/crink^2` as exposure falls. Measured over the 500k cold-start
|
|
baseline (DESIGN.md §39.12), the FAILING compact tail spans crinkliness
|
|
0.12..0.59 and quality 1e-300..1e-1 -- every value of which is numerically
|
|
zero beside a passing leaf's ~1. `crinkliness_tail="ramp"` replaces that tail,
|
|
and only that tail, with a straight line in crinkliness.
|
|
|
|
**Why stock scoring is valid here** (the §38.9 trap, and the one case the
|
|
`9gj` bead flags as exempt): the ramp is continuous at FAIL_THRESHOLD and
|
|
strictly below it, so no leaf changes which side of the threshold it is on.
|
|
The fail set is byte-identical on every corpus artefact -- asserted in
|
|
`tests/test_fitness_crinkliness_tail.py`, not assumed here. An arm therefore
|
|
cannot win by deleting a fail category, and both arms are scored under stock.
|
|
|
|
**Two experiment shapes**, because they answer different questions:
|
|
|
|
--start plateau (default) seed each run from that programme's
|
|
`coldstart-500000-s<k>.dom`. This is the ESCAPE
|
|
test the bead asks for: the ramp has signal only
|
|
where a search has already built partially-lit
|
|
rooms, and the corpus `init.dom` files show a
|
|
+0.000% score delta -- there is nothing for it to
|
|
grade at the start of a search.
|
|
--start init cold start, for comparison.
|
|
|
|
Pairing is on (starting layout, RNG seed), so `--seeds N` over `--starts M`
|
|
gives N*M paired samples per programme.
|
|
|
|
Usage::
|
|
|
|
python experiments/ab_9gj_ramp.py --budget 8000 --seeds 2 --starts 3
|
|
python experiments/ab_9gj_ramp.py --start init --budget 8000 --seeds 6
|
|
|
|
Sharding, because one run is minutes and the job list is 4x that. Each shard
|
|
keeps BOTH arms of a pair together, so the two halves of a comparison never
|
|
land on differently-loaded processes::
|
|
|
|
for i in 0 1 2 3; do
|
|
python experiments/ab_9gj_ramp.py --budget 8000 --seeds 2 --starts 3 \
|
|
--shard $i --nshards 4 &
|
|
done; wait
|
|
python experiments/ab_9gj_ramp.py --report
|
|
|
|
A POWERED run needs more than the in-session pilot could afford. §39.12 puts
|
|
harbor's minimum detectable difference at n=3 at 13.7 fails; the plateau-escape
|
|
deltas here are single-digit, so budget for n >= 8 pairs per programme
|
|
(`--seeds 3 --starts 3` gives 9) and a budget large enough for either arm to
|
|
move at all -- the pilot's 8000 evals is 1.6% of what produced the plateau::
|
|
|
|
for i in $(seq 0 3); do
|
|
python experiments/ab_9gj_ramp.py --budget 100000 --seeds 3 --starts 3 \
|
|
--shard $i --nshards 4 &
|
|
done; wait
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import argparse
|
|
import collections
|
|
import copy
|
|
import csv
|
|
import time
|
|
from pathlib import Path
|
|
|
|
from homemaker_layout import dom as dom_mod
|
|
from homemaker_layout import driver, fitness
|
|
|
|
CORPUS = ["examples/harbor-house", "examples/maple-court"]
|
|
ARMS = ["gaussian", "ramp"]
|
|
|
|
|
|
def _with_tail(tail: str):
|
|
"""Patch `fitness.load_config` so every evaluator built during the run --
|
|
the driver's, the inner loop's, the seeder's -- sees `crinkliness_tail`.
|
|
|
|
`driver.search` has no parameter for it and `driver._fitness_for` is
|
|
lru_cached, so the cache is cleared around the patch (see ab_ssz_search).
|
|
"""
|
|
orig = fitness.load_config
|
|
|
|
def patched(directory, overrides=None):
|
|
ov = dict(overrides or {})
|
|
ov["crinkliness_tail"] = tail
|
|
return orig(directory, overrides=ov)
|
|
|
|
return orig, patched
|
|
|
|
|
|
def tiers(fails) -> tuple[int, int]:
|
|
c = collections.Counter(fitness.classify_fail_tier(f) for f in fails)
|
|
return c["hard"], c["soft"]
|
|
|
|
|
|
def run_arm(progdir: str, start: Path, seed: int, tail: str, budget: int,
|
|
child_budget: int) -> dict:
|
|
orig, patched = _with_tail(tail)
|
|
fitness.load_config = patched
|
|
driver._fitness_for.cache_clear()
|
|
t0 = time.perf_counter()
|
|
try:
|
|
res = driver.search(dom_mod.load(str(start)), progdir, budget=budget,
|
|
seed=seed, child_budget=child_budget, n_workers=1)
|
|
root = copy.deepcopy(res.best.root)
|
|
finally:
|
|
fitness.load_config = orig
|
|
driver._fitness_for.cache_clear()
|
|
|
|
conf, cost = orig(progdir, overrides={"leaf_sharing": True,
|
|
"collapse_insearch": True})
|
|
score, fails = fitness.Fitness(conf, cost).score_with_fails(copy.deepcopy(root))
|
|
h, s = tiers(fails)
|
|
return dict(programme=Path(progdir).name, start=start.name, seed=seed,
|
|
tail=tail, hard=h, soft=s, total=h + s, score=score,
|
|
elapsed_s=round(time.perf_counter() - t0, 1))
|
|
|
|
|
|
def starting_points(progdir: str, kind: str, n: int) -> list[Path]:
|
|
"""The distinct layouts to start from.
|
|
|
|
`--starts` applies to plateau mode only: there is exactly one `init.dom`,
|
|
and repeating it would give several jobs the same (start, seed) pairing key,
|
|
which `_report` would silently collapse to one. In init mode the RNG seed
|
|
is the only sampling dimension, so use `--seeds`.
|
|
"""
|
|
d = Path(progdir)
|
|
if kind == "init":
|
|
return [d / "init.dom"]
|
|
found = sorted(d.glob("coldstart-500000-s*.dom"))[:n]
|
|
if not found:
|
|
raise SystemExit(f"no coldstart-500000-s*.dom in {d} for --start plateau")
|
|
return found
|
|
|
|
|
|
def main() -> None:
|
|
ap = argparse.ArgumentParser(description=__doc__,
|
|
formatter_class=argparse.RawDescriptionHelpFormatter)
|
|
ap.add_argument("--budget", type=int, default=8000)
|
|
ap.add_argument("--child-budget", type=int, default=80)
|
|
ap.add_argument("--seeds", type=int, default=2, help="RNG seeds per start")
|
|
ap.add_argument("--starts", type=int, default=3,
|
|
help="plateau layouts to start from (plateau mode only)")
|
|
ap.add_argument("--start", choices=("plateau", "init"), default="plateau")
|
|
ap.add_argument("--corpus", nargs="+", default=CORPUS)
|
|
ap.add_argument("--out", default="experiments/results/ab_9gj_ramp.csv")
|
|
ap.add_argument("--shard", type=int, default=0,
|
|
help="run only jobs i where i %% nshards == shard")
|
|
ap.add_argument("--nshards", type=int, default=1,
|
|
help="split the job list across N processes; each writes "
|
|
"<out>.shard<i>. Use --report to merge and analyse.")
|
|
ap.add_argument("--report", action="store_true",
|
|
help="merge <out>.shard* (or <out>) and print the analysis "
|
|
"only -- runs nothing")
|
|
args = ap.parse_args()
|
|
|
|
out = Path(args.out)
|
|
out.parent.mkdir(parents=True, exist_ok=True)
|
|
|
|
if args.report:
|
|
rows = []
|
|
shards = sorted(out.parent.glob(out.name + ".shard*")) or (
|
|
[out] if out.exists() else [])
|
|
for sh in shards:
|
|
with sh.open() as fh:
|
|
for r in csv.DictReader(fh):
|
|
for k in ("total", "hard", "soft", "seed"):
|
|
r[k] = int(r[k])
|
|
r["score"] = float(r["score"])
|
|
rows.append(r)
|
|
if len(shards) > 1:
|
|
with out.open("w", newline="") as fh:
|
|
w = csv.DictWriter(fh, fieldnames=list(rows[0]))
|
|
w.writeheader()
|
|
w.writerows(rows)
|
|
_report(rows, args.corpus)
|
|
print(f"\nmerged {len(rows)} runs from {len(shards)} file(s) into {out}")
|
|
return
|
|
|
|
# Build the whole job list first so sharding is deterministic and every
|
|
# (start, seed) pair keeps BOTH arms in the same shard -- the comparison is
|
|
# paired, and splitting a pair across processes would let machine load
|
|
# differ between the two halves of one pair.
|
|
jobs = []
|
|
for progdir in args.corpus:
|
|
for start in starting_points(progdir, args.start, args.starts):
|
|
for seed in range(args.seeds):
|
|
jobs.append((progdir, start, seed))
|
|
|
|
rows: list[dict] = []
|
|
dest = (out if args.nshards == 1
|
|
else out.with_name(out.name + f".shard{args.shard}"))
|
|
for i, (progdir, start, seed) in enumerate(jobs):
|
|
if i % args.nshards != args.shard:
|
|
continue
|
|
for tail in ARMS:
|
|
r = run_arm(progdir, start, seed, tail, args.budget,
|
|
args.child_budget)
|
|
rows.append(r)
|
|
print(f" {r['programme']:<14} {start.name:<26} seed={seed} "
|
|
f"{tail:<9} {r['hard']}h/{r['soft']}s = {r['total']:3d} "
|
|
f"score {r['score']:.4g} {r['elapsed_s']}s", flush=True)
|
|
with dest.open("w", newline="") as fh:
|
|
w = csv.DictWriter(fh, fieldnames=list(rows[0]))
|
|
w.writeheader()
|
|
w.writerows(rows)
|
|
if args.nshards == 1:
|
|
_report(rows, args.corpus)
|
|
print(f"\nwrote {dest}")
|
|
|
|
|
|
def _report(rows, corpus) -> None:
|
|
"""homemaker-py-tco: state what this N could resolve, beside the result."""
|
|
from ab_report import format_report, paired_report
|
|
for progdir in corpus:
|
|
name = Path(progdir).name
|
|
by: dict[tuple, dict] = {}
|
|
for r in rows:
|
|
if r["programme"] == name:
|
|
by.setdefault((r["start"], r["seed"]), {})[r["tail"]] = r["total"]
|
|
keys = sorted(k for k, v in by.items() if len(v) == 2)
|
|
if len(keys) < 2:
|
|
continue
|
|
print(f"\n--- {name}: ramp vs gaussian (stock-scored total fails) ---")
|
|
print(format_report(paired_report(
|
|
[by[k]["gaussian"] for k in keys], [by[k]["ramp"] for k in keys],
|
|
"gaussian", "ramp")))
|
|
|
|
|
|
if __name__ == "__main__":
|
|
main()
|