Skip to content

Latest commit

 

History

History
408 lines (323 loc) · 27.2 KB

File metadata and controls

408 lines (323 loc) · 27.2 KB

Gate 011 — the relational retrieval key, re-run against magnitude-matched nulls

Status: STOP — 2026-08-04. All seven reproduction-contract clauses passed and no NOT-TESTABLE clause tripped, so the gate was fully scoreable; the arms were fit and the registered STOP trigger WIN(CAT vs BC) <= 0.55 fired at every seed. The relational key retrieves worse than the plain BC it was meant to beat. Registered 2026-08-04 before any arm was fit; the screening table below was produced before registration and is published complete. See Result at the foot.

Road: Graph substrate (OPEN).

This is Gate 010's licensed re-roll, and it is the only one. Gate 010 returned NOT-TESTABLE (2026-08-04) when its pre-fit clause fired at 9 of 9 cells on the CAT-PROJ null. That verdict licenses exactly one numbered successor, conditional on publishing the complete screening table (the Gate 006→007 route, under Gate 008's asymmetry rule). This gate spends it. There is no second.

What is re-rolled: the NULL, and nothing else. The descriptor, the treatment arm, the rewired refuter, the baseline, the anchor, every bar, every budget, every world, every seed and all six reproduction-contract clauses are inherited from Gate 010 verbatim and un-re-screened — none of them failed, and re-tuning them here would be exactly the forbidden move. A0 = 4, HOPS = 2, the class alphabet, the delta-vs-start rule and REL_DIM = 20 are untouched, so Gate 010's anti-stone-soup ban on sweeping the descriptor survives this re-roll intact.


What Gate 010 established, and what it did not

Did not: anything about the hypothesis. No arm was ever fit.

Did, all of it carried into this gate as inherited fact:

  • XCHK bit-identical at 24/24 goals in all 9 cells; GCHK bitwise; elite replay 0 mismatches of 3659 / 4749 / 3617; REL_ACTIVE = 16 at every seed.
  • The treatment and the load-bearing refuter both diverge freely: N_eff(CAT vs BC) 16–23, N_eff(CAT-REWIRE vs BC) 16–23, PICK_DIV(CAT vs BC) 0.833–1.000.
  • The descriptor discriminates what the BC cannot: a 2×2 clump and 4 scattered cells, scored identically by the BC, separate under REL at squared gap 4.27e-4 against 2.72e-5 under the rewiring — a factor of 15.7.

The defect being fixed, and the correction the screen forced

Gate 010's null was matched on dimension and scale. Being [z(BC) ‖ z(P·BC)] with P linear, it is an information-preserving re-weighting of the metric on the same 18-dim space; it moved the argmin on 6–11 of 24 goals where the treatment moved it on 20–23, and the GO clause differences two win rates estimated over those incomparable samples.

The diagnosis at eval time was that the null needed matching on perturbation magnitude, and the fix it implied was to scale the projected block up until its spread matched the treatment's. The screen falsified that fix. PROJ-M — the same projection, scaled by a derived factor of 2.23–2.37 — cleared 1 of 9 cells, against PROJ-Z's 0 of 9. Scaling changes almost nothing, and in hindsight it cannot: for any scaling, the argmin over [z(BC) ‖ γ·P·BC] remains an argmin over a quadratic form on the same 18-dimensional BC space, bounded by P's condition number. No amount of magnitude can make a re-encoding of 18 numbers behave like 20 new ones.

So the diagnosis was right that magnitude was the axis and wrong that the projection family could be made to satisfy it. That correction is registered here rather than quietly absorbed, because the projection family is being abandoned on a measured result and not on a hunch.

The screen also settles the open diagnostic question from the eval: for PROJ-Z and PROJ-M, slot-div tracks N_eff almost exactly (7/7, 8/9, 6/7, …). The null was agreeing with the baseline, not picking different elites that happened to behave alike. It was weak, not coarse-grained.

The complete screening table (every candidate, nothing omitted)

ES-free, tests/test_relkey_null_screen.mojo, 11 s, no arm fit anywhere. Every candidate is scored on the precondition only (N_eff against the BC baseline) and never on an arm number, because no arm number exists. Bar N_eff ≥ 12 of NUM_GOALS = 24.

The two reference rows are not candidates: they are Gate 010's inherited treatment and refuter, reproduced here to show what the candidates are being matched against.

s0 room s0 sc-v0 s0 sc-v4 s1 room s1 sc-v1 s1 sc-v5 s2 room s2 sc-v2 s2 sc-v6 cleared worst
CAT (treatment ref) 23 20 19 23 19 20 16 22 19 9/9 16
CAT-REWIRE (refuter ref) 22 22 20 23 16 18 20 22 19 9/9 16
A — PROJ-Z (Gate 010's null) 7 8 6 11 9 8 6 6 11 0/9 6
B — PROJ-M (same projection, matched) 6 8 6 10 9 8 6 6 12 1/9 6
C — PERM-M (class-permuted REL, matched) 19 20 17 17 18 13 16 18 17 9/9 13
D — HASH-M (state-independent, matched) 23 24 24 24 23 24 22 19 21 9/9 19

Derived matching factors (one per candidate per seed, no sweep): spread target (the treatment's own z(REL) block) = 32.023 / 31.764 / 31.956; γ(PROJ-M) = 2.232 / 2.350 / 2.373; γ(PERM-M) = 35.880 / 32.713 / 34.927; γ(HASH-M) = 1.551 / 1.541 / 1.549.

Two candidates survive. Both are adopted, in different registered roles — because they answer different questions and neither answers both.

The matching rule (frozen)

Each raw null block is scaled by a single derived factor, not a swept one:

γ = sqrt( S(z(REL)) / S(raw block) ), where S(B) is the mean pairwise squared distance of block B across the repertoire, over a deterministic SPREAD_PAIRS = 4096 pair subsample drawn from a local LCG seeded 0x243F6A8885A308D3 (the map_elites.mean_pairwise_bc idiom, so it consumes zero draws from the global stream). A degenerate raw block (S = 0) yields γ = 0 — it stays zero rather than being amplified into noise by a division.

Zero free parameters, computed from the repertoire alone, before any goal exists.

Arms (frozen) — exactly one line differs

Every arm calls the same make_demos, the same fit_operator[SandboxPolicyMemory] at the same FEW_N = 32 / FEW_ALPHA0/1 / FEW_SIGMA0/1 / FEW_ITERS = 30, and the same unchanged policy_score. The only difference anywhere is which key array is handed to the retrieval argmin.

arm key handed to the argmin dims gating role
COLD (none — zero seed) yes the no-retrieval floor
BC raw emap.bc 18 yes incumbent baseline, XCHK-verified
BC-Z z(BC) 18 yes control: commensuration alone
CAT [z(BC) ‖ z(REL)] 38 yes TREATMENT
CAT-REWIRE [z(BC) ‖ z(RELʳ)] 38 yes REFUTER 1 — adjacency destroyed (graph side)
CAT-PERM [z(BC) ‖ γ·RELᵖ] 38 yes REFUTER 2 — arrangement destroyed (label side)
CAT-HASH [z(BC) ‖ γ·H(key)] 38 yes NULL — dimension and magnitude, zero state information
REL-ONLY z(REL) 20 no — diagnostic replace-vs-augment; confounded by the BC-MSE score

RELᵖ (CAT-PERM). REL recomputed with the terminal and start states' cell classes pushed through one fixed permutation of the 256 cell indices — Fisher-Yates from a local LCG, PERM_LCG_BASE = 20260805, stream PERM_LCG_BASE + 1299709·s, drawn once per seed and shared by every world, elite and goal in that seed. It preserves the class marginals exactly (so it carries essentially the occupancy information the BC already has) and destroys where the classes sit. It runs through the identical counting pipeline, so its magnitude is comparable by construction before γ touches it.

This is a genuinely second structure-destroying refuter, and it attacks the same claim from the opposite side: CAT-REWIRE scrambles the graph and leaves the labels; CAT-PERM scrambles the labels and leaves the graph. A win that survives both is a win that needs labels and adjacency in their actual correspondence — which is precisely the relational claim.

H(key) (CAT-HASH). 20 coordinates from a local LCG seeded by the elite's own stored cell key (HASH_LCG_BASE = 20260806, stream HASH_LCG_BASE + 1299709·s, mixed with the key by the 0x9E3779B97F4A7C15 multiply). It carries literally zero information about the world state — the purest possible form of "20 more numbers" — and is magnitude-matched by the same γ rule. A goal is not a stored elite, so a goal's block is keyed by the goal's own cell key: the same rule applied to the same kind of object.

Everything else is Gate 010 verbatim: worlds (push_mode = SB_PUSH_OFF, seeded blocks 0; sources shelves(s) + columns(s) at the locked BUILD_*; targets room(s), scatter(s), scatter(s+4) — 9 gating cells), NUM_GOALS = 24 from the gen_family_s-shaped sibling at max_tries = 200000, seeds {0,1,2} with variant = s so each seed is an independent world draw, and true CRN re-seeding FIT_SEED_BASE + 1000003·s + 7919·world_idx + 31·g before each arm's fit.

Metric + consumer (inherited verbatim)

Consumer: EliteMap.nearest's argmin (src/map_elites.mojo:260), read by run_family_select (src/transfer.mojo:628) and run_family (src/transfer.mojo:380), which memcpy the returned slot's weights into pol_fast and hand them to fit_operator as the ES seed. Scored number per goal per arm: d_X[g] = −policy_score(...)[0], the post-fit BC-MSE distance to the held-out goal. WIN over N_eff with post-fit ties counting against the treatment; RATIO the median of per-goal ratios over all goals, never a mean of ratios; PICK_DIV, TIE_post, REL_SHARE, RATIO_cold(BC) reported as before, the last two as diagnostics that may not be promoted to bars.

Decidability precondition (pre-fit, ES-free — unchanged, and now met by construction)

N_eff ≥ 12 at every gating cell, for the treatment and for every gated control against the baseline. REL_ACTIVE ≥ 4.

No bar is relaxed by this gate. The eval argued the bar was mis-scoped for a null, and that argument still looks correct — but it is not being acted on, because an argument discovered after watching a bar fail is indistinguishable from a motivated one. Instead the null's construction was changed until it meets the unchanged bar: CAT-PERM clears 9/9 at worst 13, CAT-HASH 9/9 at worst 19. This is the honest route and it costs nothing — the screen took 11 seconds.

GO / STOP / PARTIAL / NOT-TESTABLE (committed now)

NOT-TESTABLE — checked in this order; nothing below is scored if any trips:

  • any reproduction-contract clause fails;
  • XCHK or GCHK fails;
  • N_eff < 12 at any gating cell for CAT-vs-BC, CAT-vs-BC-Z, CAT-REWIRE-vs-BC, CAT-PERM-vs-BC or CAT-HASH-vs-BC;
  • REL_ACTIVE < 4 at any seed;
  • TIE_post(CAT vs BC) > 0.50 at ≥ 2 of 3 seeds — the instrument failing, not the hypothesis.

GO — every condition, at all 9 gating cells:

  • WIN(CAT vs BC) ≥ 0.70 and WIN(CAT vs BC-Z) ≥ 0.70
  • RATIO(CAT vs BC) ≥ 1.10 and RATIO(CAT vs BC-Z) ≥ 1.10
  • WIN(CAT vs BC) − WIN(X vs BC) ≥ 0.15 for each of X ∈ {CAT-REWIRE, CAT-PERM, CAT-HASH}
  • no per-cell overlap on the primary: min_cells RATIO(CAT vs BC) > max_cells RATIO(X vs BC) for each of the same three.

STOP — any one trigger at ≥ 2 of 3 seeds:

  • WIN(CAT vs BC) ≤ 0.55 — the relational key does not retrieve better. The likeliest "no".
  • |RATIO(CAT vs BC) − RATIO(X vs BC)| ≤ 0.03 while RATIO(CAT vs BC) ≥ 1.10, for either structure-destroying refuter X ∈ {CAT-REWIRE, CAT-PERM} — structure-free matches structured.
  • RATIO(CAT vs BC-Z) < 1.0 while RATIO(BC-Z vs BC) ≥ 1.10 — the win is the commensuration, i.e. Gate 009 E2's pose-dim geometry.

PARTIAL: any GO condition holding at exactly 2 of 3 seeds, or 3/3 on two target worlds and not the third — splits are never rounded up. Or WIN clears everywhere while RATIO < 1.10: the argmin is better more often than not, but the family-level effect is under the roadmap-worthy floor. Or the treatment clears both structure-destroying refuters but not CAT-HASH — see disclosure 3.

Reproduction contract (checked FIRST; failure ⇒ NOT-TESTABLE)

Gate 010's six clauses, unchanged, plus one:

  1. Zero learning-core change: empty git diff on src/esper_evolution.mojo, src/memory.mojo, src/map_elites.mojo, src/graph_domain.mojo, the Domain trait in src/arc_io.mojo, and ExamplePair/Task in src/hope.mojo.
  2. Zero metric / world / budget change: empty git diff on src/sandbox.mojo, src/transfer.mojo.
  3. Bitwise reproduction before anything is scored: ./esper test transfer, cbr_retain, anytime_metric, trial_select, graph_lattice_repro.
  4. XCHK; 5. GCHK; 6. key-builder blindness + zero global RNG draws in the key path.
  5. NEW — G010CHK: tests/test_relational_key.mojo's ES-free output reproduces Gate 010's published numbers exactly. Additions to src/relkey.mojo are strictly additive and must be inert; this was verified at build time by diffing against the committed log and must hold at eval.

Anti-stone-soup clause

Gate 010's clause is inherited in full, with these additions:

  • The re-roll is spent. A NOT-TESTABLE here does not license another null screen. The allowance was one, and this gate is it.
  • The screened-and-failed candidates are dead. PROJ-Z and PROJ-M may not be revived, rescaled, re-seeded, or re-screened at another γ. They failed on a measured result that is published above.
  • No post-hoc move of PERM_LCG_BASE, HASH_LCG_BASE, SPREAD_PAIRS = 4096, the γ rule, or any inherited constant. In particular the γ rule may not be swapped for a different matching statistic after seeing an arm number.
  • Neither surviving null may be dropped after the fact. If CAT-HASH proves trivially beaten (disclosure 3), that is a disclosure on the claim, not grounds to remove the arm from the table.
  • REL-ONLY remains a diagnostic, may not be promoted to a bar in either direction, and runs last.

Disclosure — the ways this gate is weaker than it looks

Gate 010's disclosures 1–7 are inherited unchanged (the BC-MSE score favours BC-shaped keys, so a GO supports "relational structure adds to the attribute key under a BC-MSE score" and not "a relational key beats an attribute key"; RATIO is diluted by ties and WIN inflated by excluding them; N_eff is necessary but not sufficient; Gate 007's A_eff = 2 / t = 2 licence is inherited whole and not widened; the gate uses GraphDomain's container, not its metric seam; the cross-world delta is improved but not solved; one descriptor was registered and may simply be a bad one). Added here:

  1. The eval's own diagnosis was half wrong, and the half that was wrong is the half that generated this gate's design. "Match on perturbation magnitude" was right; "so scale the projection" was falsified by the screen at 1/9. The current design is the second attempt at the same reasoning step, and nothing guarantees the reasoning is now complete.
  2. CAT-PERM overlaps CAT-REWIRE in what it destroys. Both break the label↔adjacency correspondence. Two refuters agreeing is weaker evidence than two independent refuters agreeing, and how independent these two really are is unquantified.
  3. CAT-HASH is probably easy to beat, because state-independent retrieval noise is likely actively harmful rather than neutral. If so, WIN(CAT-HASH vs BC) will be low, the 0.15 gap will clear trivially, and the dimension question will have been answered by a weak control. A GO's dimension-claim is therefore weaker than its adjacency-claim, and must be stated that way. This is registered before the number, so it cannot be discovered afterwards as an excuse.
  4. No null that is both information-free and magnitude-matched was found in the projection family, and the screen suggests none exists there. The two survivors achieve magnitude matching by destroying structure (PERM) or by adding noise (HASH) — neither is the clean "same information, more dimensions" object the dimension question ideally wants. That object may not exist.
  5. The re-roll is spent. If this gate returns NOT-TESTABLE, the relational-key lever closes without ever having been measured, and that outcome must be booked as such rather than quietly re-opened under a new name.

Consequence

  • GO ⇒ the Graph substrate Road gains a relational-retrieval-key rung, carrying Gate 007's A_eff = 2 / t = 2 caveats and disclosures 1–3 above. Switching the live consumer requires touching src/map_elites.mojo, which this contract forbids — a separate gate, not a licensed edit. Gate 007's registered Consequence is discharged positively.
  • STOP ⇒ the relational-key lever closes for a measured reason; Gate 007's consequence is discharged negatively and may not be re-opened by a descriptor tweak. The Road's remaining open question becomes the third-substrate one (sequences/sets through Domain).
  • PARTIAL ⇒ licenses only the cells that cleared; a consumer switch must re-register its cell list.
  • NOT-TESTABLEno further re-roll. The lever closes unmeasured and is booked as the campaign's fifth NOT-TESTABLE, with the honest note that the instrument, not the idea, is what defeated it.

Result — STOP (2026-08-04)

Scored against the registered criteria, unchanged. tests/test_relational_key_matched.mojo --arms, 8 arms × 24 goals × 9 cells, ~29 min. Log: g011_arms_clean.log.

The ladder, in the registered order

# clause outcome
1 zero learning-core change PASS — empty diff on all six files
2 zero metric / world / budget change PASSsandbox.mojo, transfer.mojo clean
3 bitwise reproduction PASStransfer, cbr_retain, anytime_metric, trial_select, and graph_lattice_repro (Gate 007's targets exact: before 0.74175346 / after 1.0, bits 1061020558 / 1065353216)
4 XCHK PASS — bit-identical at 24/24 goals, all 9 cells
5 GCHK PASS, indirectly — see disclosure A below
6 key-builder blindness PASS — zero learning-module references in src/relkey.mojo
7 G010CHK PASS — Gate 010's harness ES-free, diffed bitwise against its committed log
N_eff ≥ 12, every gated pair, every cell does not trip — min 14 across all 45 checks
REL_ACTIVE ≥ 4 does not trip — 16 at every seed
TIE_post(CAT vs BC) > 0.50 does not trip — max 0.0625

Nothing routes to NOT-TESTABLE. The gate is fully scoreable, and the result is a measurement. The TIE_post clause is the one that matters most here: it was registered so that a bad number could not later be dismissed as "the fit erased the seed and nothing could have helped". At a maximum of 0.0625 the seed plainly survived the fit, so the "no" has to be taken at face value.

The scored table (all 9 cells, nothing omitted)

WIN (fraction of N_eff goals where the arm lands strictly closer than BC) and RATIO (median per-goal d_BC / d_arm; > 1 means the arm beats BC).

cell CAT CAT vs BC-Z CAT-REWIRE CAT-PERM CAT-HASH BC-Z COLD
room v0 s0 0.304 / 0.459 0.304 / 0.581 0.318 / 0.828 0.316 / 0.959 0.348 / 0.677 0.571 / 1.000 0.333 / 0.579
scat v0 s0 0.421 / 1.000 0.467 / 1.000 0.381 / 0.645 0.375 / 1.000 0.391 / 0.682 0.375 / 1.000 0.375 / 0.769
scat v4 s0 0.375 / 1.000 0.471 / 1.000 0.316 / 0.587 0.250 / 0.755 0.333 / 0.453 0.300 / 1.000 0.250 / 0.372
room v1 s1 0.217 / 0.547 0.417 / 0.764 0.304 / 0.450 0.294 / 0.834 0.292 / 0.494 0.267 / 1.000 0.167 / 0.359
scat v1 s1 0.300 / 0.828 0.524 / 1.000 0.400 / 0.953 0.429 / 1.000 0.292 / 0.459 0.333 / 1.000 0.458 / 0.955
scat v5 s1 0.273 / 0.427 0.238 / 0.572 0.364 / 0.650 0.455 / 1.000 0.217 / 0.393 0.444 / 1.000 0.417 / 0.952
room v2 s2 0.438 / 1.000 0.375 / 0.850 0.300 / 0.902 0.375 / 1.000 0.182 / 0.497 0.615 / 1.000 0.292 / 0.788
scat v2 s2 0.235 / 0.753 0.200 / 0.675 0.100 / 0.480 0.385 / 1.000 0.200 / 0.451 0.429 / 1.000 0.458 / 0.988
scat v6 s2 0.619 / 1.072 0.625 / 1.000 0.381 / 0.925 0.588 / 1.000 0.333 / 0.593 0.385 / 1.000 0.667 / 1.725

Verdict arithmetic

GO fails outright. It requires WIN(CAT vs BC) ≥ 0.70 at all nine cells; the maximum observed is 0.619 and six cells sit below 0.44. RATIO(CAT vs BC) ≥ 1.10 is likewise required everywhere and peaks at 1.072, sitting at or below 1.0 in seven of nine cells.

STOP trigger 1 fires: WIN(CAT vs BC) ≤ 0.55 at ≥ 2 of 3 seeds.

  • seed 0 — 0.304, 0.421, 0.375: all three cells at or under the trigger
  • seed 1 — 0.217, 0.300, 0.273: all three
  • seed 2 — 0.438, 0.235, 0.619: two of three

Registered ambiguity, disclosed: the trigger says "at ≥ 2 of 3 seeds" without stating how the three cells inside a seed aggregate. The verdict is invariant to the choice — under the strict reading (the trigger must hold at every cell of a seed) it fires at seeds 0 and 1, giving 2 of 3; under the loose reading (any cell) it fires at 3 of 3. Both clear the ≥ 2 bar. Recorded because a criterion that needed interpreting at scoring time is a registration defect even when it does not bite.

STOP triggers 2 and 3 do not fire, both because their guards never activate: trigger 2 requires RATIO(CAT vs BC) ≥ 1.10 (max 1.072) and trigger 3 requires RATIO(BC-Z vs BC) ≥ 1.10 (flat 1.000 everywhere). Exactly one trigger carries this verdict, and it is the primary one.

PARTIAL does not apply. No GO condition holds at any seed, so there is no split to round.

What the controls say

The instrument was working. BC beats COLD in 8 of 9 cells (RATIO(COLD vs BC) < 1 everywhere but scat v6 s2). Retrieval across the world gap is live and productive — this is a working mechanism returning a "no", not a dead harness returning noise. Had this failed, the registered refuting-control clause would have voided the comparison.

CAT-HASH is the worst arm in every single cell (RATIO 0.393–0.682). Registration disclosure 3 predicted exactly this — that a state-independent block would be actively harmful rather than neutral — and said a GO's dimension-claim would therefore be weak. The prediction was correct and the weakness costs nothing, because the treatment did not win: there is no dimension-claim to qualify. This is the one place where a registered weakness turned out to be free.

CAT-REWIRE is not reliably worse than CAT. In four cells the scrambled graph retrieves better than the true one (room v0: 0.828 vs 0.459; scat v5: 0.650 vs 0.427; room v2: 0.902 vs 1.000 is the exception; scat v1: 0.953 vs 0.828). Destroying adjacency does not consistently hurt, which is the sharpest single statement that the descriptor's structure was not carrying retrieval-relevant information.

CAT-PERM sits closest to 1.000 of the three 38-dim arms — the label-side scramble is the least harmful. Read together with CAT-REWIRE and CAT-HASH, the ordering is by how much the extra block perturbs the argmin away from the BC, not by how much real structure it carries. That is the whole finding in one sentence.

The finding

Adding the relational descriptor to the retrieval key makes retrieval worse, and it does so in a way that indicts the descriptor rather than the experiment. A WIN near 0.3 means CAT loses to the plain BC roughly twice as often as it wins. The mechanism the arms expose is dilution: the key is 18 BC coordinates plus 20 REL coordinates, so arrangement dominates over half the comparison, and the three arms rank by perturbation magnitude rather than by information content.

Interpretation (mine, not measured): disclosure 1 named the likely reason before the run and it now looks load-bearing — the goal is specified as a BC, so matching on occupancy is matching on the very quantity the task asks the fit to reproduce, while arrangement describes a property no goal ever mentions. The relational key answered a question nobody asked, and outvoted the one that was.

Disclosures on this result

  • A — GCHK was discharged indirectly. The Gate 011 harness does not call gchk(); that was an omission in the build, not a design choice. It was discharged by running Gate 010's harness (which does) and then verifying the two gen_goals_with_state functions are identical — the only difference between them is a comment line, so the generator GCHK certified is verbatim the one this gate used. A direct inline check would have been better and its absence is recorded.
  • B — the first --arms run raised on an unregistered assertion. An inline G010CHK table in the harness compared this gate's N_eff against Gate 010's published values, which is undefined once fits advance the stream between worlds (cells 0/3/6 matched, the other six did not — exactly that signature). The registered clause 7 is narrower and names Gate 010's own harness. The assertion was scoped to the ES-free path and the run repeated; the re-run is bitwise identical to the pre-fix run, verified by diff, which is the guarantee that having seen the numbers first could not have shaped the fix. Both logs are kept.
  • C — inherited disclosures stand. Gate 010's 1–7 and this gate's 1–5 all carry, in particular that the BC-MSE score favours BC-shaped keys (disclosure 1), that CAT-PERM and CAT-REWIRE overlap in what they destroy (disclosure 2), and that one descriptor was registered and may simply be a bad one (Gate 010's disclosure 7). The STOP licenses "this descriptor did not", never "relational keys do not".

Consequence, as registered

The relational-key lever closes for a measured reason. Gate 007's registered Consequence — "M4's relational key is built over GraphDomain" — is hereby discharged negatively, and per the anti-stone-soup clause it may not be re-opened by a descriptor tweak: no new A0, no HOPS = 3, no WL round, no re-weighted class pairs. The re-roll is spent and there is no second.

The Graph substrate Road stays OPEN — a STOP closes a lever, not a Road, and the Road's own registered remaining question is untouched: whether a third substrate through the Domain seam (sequences, sets) makes "substrate-neutral" a general claim rather than two data points.

What this hands forward, unclaimed and ungated: the run measured that retrieval itself is live and productive across the world gap (BC beats COLD at 8 of 9 cells) while the hand-designed index improvement is not — which is consistent with Route C's standing PARTIAL, where real-trial selection beat nearest-1 decisively in one world and straddled the bar in the other. The index question is not closed by this STOP; only this way of attacking it is.