Quality aggregation: divide out how many questions a leaf was asked
39.17 left the search's storey choice unexplained and blamed value_rate. It is
not the rate, or not only.
Measured over the twelve baseline runs, value/cost by leaf kind: outside ground
7.40, roof terrace 2.69, room 0.34, circulation 0.02. A terrace returns 2.7x
its cost where a room returns a third of it, so filling upper storeys with
terrace is not the search leaving value on the table -- it is by a wide margin
the most profitable thing the objective offers. 7% of the corpus area produces
32% of its value.
Most of that gap is mean quality: 0.986 for a terrace against 0.223 for a
room. Quality is a PRODUCT of factors and the kinds are not asked the same
number of questions -- an outside leaf is exempt from size, crinkliness and
access, so 3 of 7 factors can ever bite it against a room's 6. Each exemption
is individually right (no programme size target; uncovered outside is lit by
definition; ground-level outside needs no access). The consequence is not: a
leaf exempt from the two harshest factors out-scores one judged on them and
doing well, purely by not being asked, and quality multiplies the value rate.
Stated generally, and this is not about outside space: under a product, adding
any new quality criterion mechanically devalues every leaf it applies to,
including leaves that score 1.0 on it. The objective's scale should not depend
on how many things it measures.
quality_aggregate="geometric_mean" (default OFF, "product" is stock) divides
that out. Computed in log space so six small factors cannot underflow the
product before the root is taken; a zero factor still gives zero, so a fully
buried leaf is worth nothing either way.
Telling "exempt" from "asked and scored 1.0" needs factor_is_asked, which
restates conditions that live inside the quality_* methods. That duplication
can drift, so tests/test_fitness_aggregate.py pins it against every leaf in the
corpus: wherever the predicate says exempt, the factor really is 1.0.
Fail set byte-identical everywhere, and for a stronger reason than 39.13/39.14
had: evaluate_leaf emits each fail from the factor itself before anything is
combined, so no aggregation can move one. Score effect +37% to +169%, reaching
all four programmes where the crinkliness changes reached two; room value/cost
0.34 -> 0.66, circulation 0.02 -> 0.07.
Deliberately not fixed: a terrace still out-earns a room 4:1, which is the
rates (value_supported = value_inside = 300 against costs of 110 and 200), not
the aggregation. That is a design judgement for the programme author, and
39.16 is a standing reminder that "this inherited constant looks wrong" has
been wrong twice already in this section. Left open on ecx with the numbers.
A/B running; verdict to follow.
Refs homemaker-py-ecx.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 18:57:00 +00:00
|
|
|
"""`quality_aggregate="geometric_mean"` (homemaker-py-ecx, DESIGN.md §39.18).
|
|
|
|
|
|
|
|
|
|
Quality is a product over the factors, and leaf kinds face different numbers of
|
|
|
|
|
them: a room is judged on size, crinkliness and access, an outside leaf is
|
|
|
|
|
exempt from all three. Exemption alone therefore buys a higher quality, and
|
|
|
|
|
quality multiplies the value rate. The geometric mean divides that out.
|
|
|
|
|
|
|
|
|
|
Two invariants matter and both are asserted here:
|
|
|
|
|
|
|
|
|
|
* the fail set cannot move, because `evaluate_leaf` emits each fail from the
|
|
|
|
|
factor itself before anything is combined;
|
|
|
|
|
* `factor_is_asked` must agree with the `quality_*` methods -- whenever it says
|
|
|
|
|
a factor is exempt, that factor really is exactly 1.0. It is a separate
|
|
|
|
|
statement of the same conditions, so it can drift; this pins it.
|
|
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
from __future__ import annotations
|
|
|
|
|
|
|
|
|
|
import copy
|
|
|
|
|
import math
|
|
|
|
|
from pathlib import Path
|
|
|
|
|
|
|
|
|
|
import pytest
|
|
|
|
|
|
|
|
|
|
from homemaker_layout import dom as dom_mod
|
|
|
|
|
from homemaker_layout.fitness import Fitness, load_config
|
|
|
|
|
|
|
|
|
|
EXAMPLES = Path(__file__).resolve().parent.parent / "examples"
|
|
|
|
|
PROGRAMMES = ["harbor-house", "maple-court", "health-centre", "programme-house"]
|
|
|
|
|
pytestmark = pytest.mark.skipif(not (EXAMPLES / "harbor-house").is_dir(),
|
|
|
|
|
reason="examples absent")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _artefacts():
|
|
|
|
|
for name in PROGRAMMES:
|
|
|
|
|
d = EXAMPLES / name
|
|
|
|
|
if not d.is_dir():
|
|
|
|
|
continue
|
|
|
|
|
for p in sorted(d.glob("coldstart-500000-s*.dom")) + [d / "init.dom"]:
|
|
|
|
|
if p.exists():
|
|
|
|
|
yield d, p
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_exempt_factors_really_are_one():
|
|
|
|
|
"""The invariant `factor_is_asked` rests on, checked against every leaf in
|
|
|
|
|
the corpus rather than assumed from reading the code."""
|
|
|
|
|
checked = 0
|
|
|
|
|
for d, p in _artefacts():
|
A terrace is no longer worth more per m2 than a real internal room
Owner's ruling. Measured over the twelve baseline layouts as realised value per
m2 (rate x quality, not the rate alone):
as shipped before this room 67.0 terrace 294.6 violates, 4.39x
value_supported=100 only room 67.0 terrace 98.2 violates, 1.46x
geometric mean only room 132.9 terrace 296.8 violates, 2.23x
both room 132.9 terrace 98.9 satisfies
The 4.4x is roughly 2.2x aggregation and 2.0x rate, so neither half alone is
enough. That is why 39.18's geometric-mean aggregation moves from default-OFF
to default-ON here rather than waiting on its own A/B: it is not an optional
improvement, it is half of a ruling.
value_supported 300 -> 100, and set to value_outside rather than to a number
that makes the inequality come out -- back-solving from the corpus's measured
mean room quality would rot the moment either changed. Outdoor space is worth
the same to an occupant whatever level it sits on; the real difference between
a ground garden and a roof terrace is what it takes to BUILD, and cost already
says that (outside 10.0 vs outside_supported 110.0). Value describes worth,
cost describes structure, and the level belongs in the second.
Changed in CONF_DEFAULTS and the four corpus patterns.config files, which all
declared 300.0 explicitly. NOT changed in harbor-house-l0 (a shape-curve test
fixture) or y51-sweep-* (historical fixtures that exist to reproduce past
measurements) -- repricing those would destroy what they are for.
Neither change can move a fail, structurally rather than luckily: value rates
never enter fail emission, and evaluate_leaf emits each fail from its factor
before anything is combined. Verified corpus-wide: identical fail sets, scores
+11% to +169% (and -5% once, on a layout that is mostly terrace).
tests/test_terrace_value_ruling.py pins the ruling as an invariant of the
objective, and asserts that reverting the aggregation breaks it again, so
neither half can be quietly dropped.
The 500k cold-start baseline (39.12) is superseded -- this changes what "good"
means. The layouts stay valid and their fail counts are unchanged, but a fresh
corpus run is needed before any new number is compared with them.
Still untouched: circulation returns 0.07 per unit cost against a room's 0.66,
by far the worst thing a building can contain. That is homemaker-py-hxi.
Refs homemaker-py-ecx.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 19:25:02 +00:00
|
|
|
conf, cost = load_config(d, overrides={"quality_aggregate": "product"})
|
Quality aggregation: divide out how many questions a leaf was asked
39.17 left the search's storey choice unexplained and blamed value_rate. It is
not the rate, or not only.
Measured over the twelve baseline runs, value/cost by leaf kind: outside ground
7.40, roof terrace 2.69, room 0.34, circulation 0.02. A terrace returns 2.7x
its cost where a room returns a third of it, so filling upper storeys with
terrace is not the search leaving value on the table -- it is by a wide margin
the most profitable thing the objective offers. 7% of the corpus area produces
32% of its value.
Most of that gap is mean quality: 0.986 for a terrace against 0.223 for a
room. Quality is a PRODUCT of factors and the kinds are not asked the same
number of questions -- an outside leaf is exempt from size, crinkliness and
access, so 3 of 7 factors can ever bite it against a room's 6. Each exemption
is individually right (no programme size target; uncovered outside is lit by
definition; ground-level outside needs no access). The consequence is not: a
leaf exempt from the two harshest factors out-scores one judged on them and
doing well, purely by not being asked, and quality multiplies the value rate.
Stated generally, and this is not about outside space: under a product, adding
any new quality criterion mechanically devalues every leaf it applies to,
including leaves that score 1.0 on it. The objective's scale should not depend
on how many things it measures.
quality_aggregate="geometric_mean" (default OFF, "product" is stock) divides
that out. Computed in log space so six small factors cannot underflow the
product before the root is taken; a zero factor still gives zero, so a fully
buried leaf is worth nothing either way.
Telling "exempt" from "asked and scored 1.0" needs factor_is_asked, which
restates conditions that live inside the quality_* methods. That duplication
can drift, so tests/test_fitness_aggregate.py pins it against every leaf in the
corpus: wherever the predicate says exempt, the factor really is 1.0.
Fail set byte-identical everywhere, and for a stronger reason than 39.13/39.14
had: evaluate_leaf emits each fail from the factor itself before anything is
combined, so no aggregation can move one. Score effect +37% to +169%, reaching
all four programmes where the crinkliness changes reached two; room value/cost
0.34 -> 0.66, circulation 0.02 -> 0.07.
Deliberately not fixed: a terrace still out-earns a room 4:1, which is the
rates (value_supported = value_inside = 300 against costs of 110 and 200), not
the aggregation. That is a design judgement for the programme author, and
39.16 is a standing reminder that "this inherited constant looks wrong" has
been wrong twice already in this section. Left open on ecx with the numbers.
A/B running; verdict to follow.
Refs homemaker-py-ecx.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 18:57:00 +00:00
|
|
|
fit = Fitness(conf, cost)
|
|
|
|
|
seen = []
|
|
|
|
|
orig = Fitness.evaluate_leaf
|
|
|
|
|
|
|
|
|
|
def ev(self, leaf, G, level_id, groups, fail, _o=orig, _s=seen):
|
|
|
|
|
q, f = _o(self, leaf, G, level_id, groups, fail)
|
|
|
|
|
_s.append((leaf, dict(f)))
|
|
|
|
|
return q, f
|
|
|
|
|
|
|
|
|
|
Fitness.evaluate_leaf = ev
|
|
|
|
|
try:
|
|
|
|
|
fit.score_with_fails(dom_mod.load(str(p)))
|
|
|
|
|
finally:
|
|
|
|
|
Fitness.evaluate_leaf = orig
|
|
|
|
|
|
|
|
|
|
for leaf, factors in seen:
|
|
|
|
|
for name, value in factors.items():
|
|
|
|
|
if not fit.factor_is_asked(name, leaf):
|
|
|
|
|
assert value == 1.0, (
|
|
|
|
|
f"{p.name}: {name} is marked exempt for leaf "
|
|
|
|
|
f"{leaf.id!r} ({leaf.type!r}) but scored {value}")
|
|
|
|
|
checked += 1
|
|
|
|
|
assert checked > 100, "expected plenty of exempt factors to check"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_fail_set_is_byte_identical():
|
|
|
|
|
for d, p in _artefacts():
|
|
|
|
|
root = dom_mod.load(str(p))
|
A terrace is no longer worth more per m2 than a real internal room
Owner's ruling. Measured over the twelve baseline layouts as realised value per
m2 (rate x quality, not the rate alone):
as shipped before this room 67.0 terrace 294.6 violates, 4.39x
value_supported=100 only room 67.0 terrace 98.2 violates, 1.46x
geometric mean only room 132.9 terrace 296.8 violates, 2.23x
both room 132.9 terrace 98.9 satisfies
The 4.4x is roughly 2.2x aggregation and 2.0x rate, so neither half alone is
enough. That is why 39.18's geometric-mean aggregation moves from default-OFF
to default-ON here rather than waiting on its own A/B: it is not an optional
improvement, it is half of a ruling.
value_supported 300 -> 100, and set to value_outside rather than to a number
that makes the inequality come out -- back-solving from the corpus's measured
mean room quality would rot the moment either changed. Outdoor space is worth
the same to an occupant whatever level it sits on; the real difference between
a ground garden and a roof terrace is what it takes to BUILD, and cost already
says that (outside 10.0 vs outside_supported 110.0). Value describes worth,
cost describes structure, and the level belongs in the second.
Changed in CONF_DEFAULTS and the four corpus patterns.config files, which all
declared 300.0 explicitly. NOT changed in harbor-house-l0 (a shape-curve test
fixture) or y51-sweep-* (historical fixtures that exist to reproduce past
measurements) -- repricing those would destroy what they are for.
Neither change can move a fail, structurally rather than luckily: value rates
never enter fail emission, and evaluate_leaf emits each fail from its factor
before anything is combined. Verified corpus-wide: identical fail sets, scores
+11% to +169% (and -5% once, on a layout that is mostly terrace).
tests/test_terrace_value_ruling.py pins the ruling as an invariant of the
objective, and asserts that reverting the aggregation breaks it again, so
neither half can be quietly dropped.
The 500k cold-start baseline (39.12) is superseded -- this changes what "good"
means. The layouts stay valid and their fail counts are unchanged, but a fresh
corpus run is needed before any new number is compared with them.
Still untouched: circulation returns 0.07 per unit cost against a room's 0.66,
by far the worst thing a building can contain. That is homemaker-py-hxi.
Refs homemaker-py-ecx.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 19:25:02 +00:00
|
|
|
c_prod, cost = load_config(d, overrides={"quality_aggregate": "product"})
|
Quality aggregation: divide out how many questions a leaf was asked
39.17 left the search's storey choice unexplained and blamed value_rate. It is
not the rate, or not only.
Measured over the twelve baseline runs, value/cost by leaf kind: outside ground
7.40, roof terrace 2.69, room 0.34, circulation 0.02. A terrace returns 2.7x
its cost where a room returns a third of it, so filling upper storeys with
terrace is not the search leaving value on the table -- it is by a wide margin
the most profitable thing the objective offers. 7% of the corpus area produces
32% of its value.
Most of that gap is mean quality: 0.986 for a terrace against 0.223 for a
room. Quality is a PRODUCT of factors and the kinds are not asked the same
number of questions -- an outside leaf is exempt from size, crinkliness and
access, so 3 of 7 factors can ever bite it against a room's 6. Each exemption
is individually right (no programme size target; uncovered outside is lit by
definition; ground-level outside needs no access). The consequence is not: a
leaf exempt from the two harshest factors out-scores one judged on them and
doing well, purely by not being asked, and quality multiplies the value rate.
Stated generally, and this is not about outside space: under a product, adding
any new quality criterion mechanically devalues every leaf it applies to,
including leaves that score 1.0 on it. The objective's scale should not depend
on how many things it measures.
quality_aggregate="geometric_mean" (default OFF, "product" is stock) divides
that out. Computed in log space so six small factors cannot underflow the
product before the root is taken; a zero factor still gives zero, so a fully
buried leaf is worth nothing either way.
Telling "exempt" from "asked and scored 1.0" needs factor_is_asked, which
restates conditions that live inside the quality_* methods. That duplication
can drift, so tests/test_fitness_aggregate.py pins it against every leaf in the
corpus: wherever the predicate says exempt, the factor really is 1.0.
Fail set byte-identical everywhere, and for a stronger reason than 39.13/39.14
had: evaluate_leaf emits each fail from the factor itself before anything is
combined, so no aggregation can move one. Score effect +37% to +169%, reaching
all four programmes where the crinkliness changes reached two; room value/cost
0.34 -> 0.66, circulation 0.02 -> 0.07.
Deliberately not fixed: a terrace still out-earns a room 4:1, which is the
rates (value_supported = value_inside = 300 against costs of 110 and 200), not
the aggregation. That is a design judgement for the programme author, and
39.16 is a standing reminder that "this inherited constant looks wrong" has
been wrong twice already in this section. Left open on ecx with the numbers.
A/B running; verdict to follow.
Refs homemaker-py-ecx.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJ84Feep79Hhm3E4zZJmnB
2026-09-05 18:57:00 +00:00
|
|
|
c_geo, _ = load_config(d, overrides={"quality_aggregate": "geometric_mean"})
|
|
|
|
|
_, f_prod = Fitness(c_prod, cost).score_with_fails(copy.deepcopy(root))
|
|
|
|
|
_, f_geo = Fitness(c_geo, cost).score_with_fails(copy.deepcopy(root))
|
|
|
|
|
assert f_prod == f_geo, f"{p} changed its fail set under the geometric mean"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_geometric_mean_is_the_product_when_every_factor_is_asked():
|
|
|
|
|
"""No free lunch: a leaf asked all six should agree with `prod ** (1/6)`."""
|
|
|
|
|
fit = Fitness(*load_config(EXAMPLES / "harbor-house"))
|
|
|
|
|
leaf = dom_mod.Node(type="r")
|
|
|
|
|
factors = {"perpendicular": 0.9, "proportion": 0.8, "size": 0.5,
|
|
|
|
|
"width": 0.95, "crinkliness": 0.4, "access": 1.0, "daylight": 1.0}
|
|
|
|
|
asked = [v for k, v in factors.items() if fit.factor_is_asked(k, leaf)]
|
|
|
|
|
expected = math.prod(asked) ** (1.0 / len(asked))
|
|
|
|
|
assert fit._aggregate_geometric(leaf, factors) == pytest.approx(expected)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_a_zero_factor_still_makes_the_leaf_worthless():
|
|
|
|
|
"""A fully buried leaf is worth nothing under either aggregation -- the
|
|
|
|
|
geometric mean must not launder a zero into 0.4-ish."""
|
|
|
|
|
fit = Fitness(*load_config(EXAMPLES / "harbor-house"))
|
|
|
|
|
leaf = dom_mod.Node(type="r")
|
|
|
|
|
factors = {"perpendicular": 1.0, "proportion": 1.0, "size": 1.0,
|
|
|
|
|
"width": 1.0, "crinkliness": 0.0, "access": 1.0, "daylight": 1.0}
|
|
|
|
|
assert fit._aggregate_geometric(leaf, factors) == 0.0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_it_does_not_underflow_where_the_product_would():
|
|
|
|
|
"""The point of computing in log space: six small factors multiply to a
|
|
|
|
|
denormal, but their geometric mean is an ordinary number."""
|
|
|
|
|
fit = Fitness(*load_config(EXAMPLES / "harbor-house"))
|
|
|
|
|
leaf = dom_mod.Node(type="r")
|
|
|
|
|
tiny = 1e-60
|
|
|
|
|
factors = {k: tiny for k in ("perpendicular", "proportion", "size",
|
|
|
|
|
"width", "crinkliness", "access")}
|
|
|
|
|
factors["daylight"] = 1.0
|
|
|
|
|
assert math.prod(factors[k] for k in factors) == 0.0 # product underflows
|
|
|
|
|
assert fit._aggregate_geometric(leaf, factors) == pytest.approx(tiny, rel=1e-6)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_unknown_aggregate_is_rejected():
|
|
|
|
|
conf, cost = load_config(EXAMPLES / "harbor-house",
|
|
|
|
|
overrides={"quality_aggregate": "mean"})
|
|
|
|
|
with pytest.raises(ValueError, match="unknown quality_aggregate"):
|
|
|
|
|
Fitness(conf, cost)
|