From 97307ebb00d825a4bf225dd0e69115d16e0ee78a Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 12 Aug 2026 20:11:59 +0000 Subject: [PATCH 1/3] =?UTF-8?q?plan:=20rebuild=20substrate-comfort-zones?= =?UTF-8?q?=20=C2=A72=20as=20a=20cross-swap,=20not=20a=20horse=20race?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Operator ruling, two messages: the method is cross-swap and hypothesis testing under the premise that the model CAPTURES the phenomenon but is NOT calibrated; and constancy is relative, so the design deliberately manufactures strong correlation differences on the assumption those differences are fit to evaluate the hypothesis. Under that premise v1 §2 was wrong in kind, not in detail. It raced four calibration arms on reconstruction RMSE — but a deliberately wrong calibration produces bad error BY DEFINITION, so CAL-ABS-FOREIGN is the measurement CONDITION, not a competitor that might lose. Racing a condition guarantees the answer and measures a tautology. Rebuilt: - §2 is a 4x4 donor x target transfer matrix per geometry arm. Primary metric Spearman rho (structure preservation); the hypothesis's quantity is the derived transfer loss L[D][T] = rho[T][T] - rho[D][T]. RMSE/bias in Pa stay per cell as evidence the swap hurt, never as the verdict. - CAL-ABS-OWN / CAL-ABS-FOREIGN collapse into ONE arm's diagonal and off-diagonal. Splitting them was the tell the design was still a race. - The dynamic arms' rows are flat by construction (no donor exists). C2 now requires BOTH that L == 0 exactly AND that the identical code path be proven able to produce a non-flat row -- the can-it-DIFFER gate W5 paid for once already. - C1b measures constancy (separation = between-box range / mean within-box sigma >= 3) instead of claiming it. C1c gives the suitability ASSUMPTION a falsifier: the regimes must differ in autocorrelation decay and rank-distribution shape, not merely in |grad p|, or the ladder is four copies of one condition and a null there VOIDS the reading. - The crossover scores against the DIAGONAL on rho (own calibration, the hardest opponent); the weak form against the swapped cell is reported separately and labelled evidence of wiring, not of merit. Unchanged: the §1 regime ladder and its three preflight corrections, equal budget, the geometry axis, the C0 control gate, the store-the-operands output contract, §5's non-claims. The scaffold was sound; the question was inverted. v1's framing is preserved in place as §2's correction note. Every D-CZ row was still Queued when this happened -- no measured result was reinterpreted. Old->new bar mapping recorded on the board. Board: EPIPHANIES E-A-HORSE-RACE-IS-NOT-A-CROSS-SWAP-1 (prepend, suffix- checked), INTEGRATION_PLANS revision entry (prepend, suffix-checked), STATUS_BOARD D-CZ rows re-cut with the renumbering note. Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi --- .claude/board/EPIPHANIES.md | 80 +++++++ .claude/board/INTEGRATION_PLANS.md | 50 ++++ .claude/board/STATUS_BOARD.md | 28 ++- .claude/plans/substrate-comfort-zones-v1.md | 239 ++++++++++++++++---- 4 files changed, 349 insertions(+), 48 deletions(-) diff --git a/.claude/board/EPIPHANIES.md b/.claude/board/EPIPHANIES.md index 8e466264..8f9b6262 100644 --- a/.claude/board/EPIPHANIES.md +++ b/.claude/board/EPIPHANIES.md @@ -1,3 +1,83 @@ +## 2026-08-12 — E-A-HORSE-RACE-IS-NOT-A-CROSS-SWAP-1 + +**Status:** FINDING `[H]` — methodological, operator-ruled. Caught on +`substrate-comfort-zones-v1` BEFORE any bar ran, so nothing measured had to +be reinterpreted. Grade is `[H]` because the reasoning is airtight given the +premise but the premise itself ("the model captures the phenomenon") is an +assumption this plan does not test. + +**The defect, in one line: when the premise is *the model captures the +phenomenon but is not calibrated*, an error metric measured under a +deliberately wrong calibration is bad BY DEFINITION, so scoring it answers +nothing.** The first draft of that plan compared four calibration arms on +reconstruction RMSE — a horse race. But `CAL-ABS-FOREIGN` is not a +competitor that might lose; it is the measurement CONDITION the whole +hypothesis is about. Racing it guarantees the answer and measures the +tautology. + +**What the design must be instead.** Miscalibration is a *transfer* +question, so the instrument is a transfer matrix, not a ranking: `M[D][T]` = +read regime `T`'s field through a codebook derived from regime `D`. The +diagonal is own-calibration; every off-diagonal cell is a swap. The primary +metric becomes **structure preservation** (Spearman `ρ` of reconstructed vs +true), and the quantity the hypothesis actually concerns is the derived +**transfer loss** `L[D][T] = ρ[T][T] − ρ[D][T]` — how much ordering survives +being read through the wrong calibration. Absolute error is retained as +*evidence that the swap genuinely hurt*, never as the verdict. + +**The sharp consequence that only appears in the matrix framing.** A +window-local encoding (rank-normalised, Fisher-z) has **no donor at all**, +so its matrix row is flat and `L ≡ 0` **by construction**. That is not an +artifact to hide — it IS the property under test: a substrate that carries +no absolute anchor cannot have that anchor be wrong. Which reframes the +whole comparison. The question stops being "which formula wins" and becomes +"in which regime does the anchor-free encoding's own-window `ρ` exceed the +absolute encoding's, measured against the DIAGONAL (its own calibration, +the hardest available opponent) rather than against the swapped cell (which +it beats almost by definition)." + +**And the degeneracy must be VERIFIED both ways** — this is the +`E-A-CONTROL-THAT-CANNOT-LOSE-IS-NO-CONTROL-1` lesson in its can-it-DIFFER +form, which W5 already paid for once. A flat row is only evidence if the +identical code path is proven able to produce a NON-flat one (the absolute +arm must show real off-diagonal degradation). Otherwise "flat" means the +harness cannot vary anything, and the encoding was never tested. + +**Second half of the ruling, on the regime ladder itself.** The operator's +framing: *in science you hold variables constant to test the others; +constancy is relative, so the design deliberately manufactures strong +correlation differences — on the assumption that those differences are fit +to evaluate the hypothesis.* Three things follow, and the plan had only the +first: + +1. **The 9.3× `|∇p|` spread across the four regimes is the design's PURPOSE, + not a by-product of box selection.** At small spread every effect is in + the noise; at large spread it is legible or it does not exist. +2. **"Constant" is relative, so it is MEASURED** — within-box spread of the + discriminator must be small relative to the between-box spread + (`separation ≥ 3`). Below that, "held constant" is a label on a box, not + a condition, and the caveat must ride every downstream number. +3. **"Fit to evaluate the hypothesis" is an ASSUMPTION and therefore needs a + falsifier.** The ladder is built on a pressure-gradient discriminator but + the hypothesis is about correlation STRUCTURE. If two regimes are + indistinguishable in autocorrelation decay and rank-distribution shape, + they are ONE condition for this question however far apart their + gradients sit — and the ladder would be measuring four copies of the same + thing. A null there VOIDS the reading rather than weakening it: it would + mean the spread was manufactured on the wrong axis. + +**The transferable rule.** Before scoring arms against each other, ask which +of them is a *competitor* and which is a *condition*. A condition put on the +starting line will always lose, and its loss carries no information. Score +conditions by what SURVIVES them, not by how much they cost. + +**Cross-ref:** `.claude/plans/substrate-comfort-zones-v1.md` §2 (the rebuilt +instrument, with v1's framing preserved in place as a correction note rather +than deleted) and §3 C1b/C1c/C2/C3/C4; `STATUS_BOARD.md` D-CZ-* (rows re-cut +while still Queued, with the old→new mapping recorded); +`E-A-CONTROL-THAT-CANNOT-LOSE-IS-NO-CONTROL-1`; +`E-R²-IS-NEAR-BLIND` (why the physical unit is retained alongside `ρ`). + ## 2026-08-12 — E-THE-DISPLACEMENT-FILTER-ATE-THE-STRANDED-STRATUM-1 **Status:** FINDING `[G]` — W6 RUN (`comet_tail_w6.py`/`.json`), audited diff --git a/.claude/board/INTEGRATION_PLANS.md b/.claude/board/INTEGRATION_PLANS.md index 5340b616..de6ac45c 100644 --- a/.claude/board/INTEGRATION_PLANS.md +++ b/.claude/board/INTEGRATION_PLANS.md @@ -1,3 +1,53 @@ +## 2026-08-12 — substrate-comfort-zones-v1 REVISED (§2 rebuilt: horse race → cross-swap) + +Same file, `.claude/plans/substrate-comfort-zones-v1.md`, still **ACTIVE**, +still exploratory tier. Not a new plan and not a v2 — a pre-registration +sharpened **before any bar ran** (only D-CZ-0, the regime preflight, was +DONE; everything else was Queued), which is precisely when a design should +be allowed to change. + +**Why.** Operator ruling, two messages: the method is *cross-swap and +hypothesis testing, under the premise that the model captures the phenomenon +but is not calibrated*; and *in science you hold variables constant to test +the others — constancy is relative, so the design deliberately manufactures +strong correlation differences, on the assumption those differences are fit +to evaluate the hypothesis*. Against that, v1 §2 was a **horse race**: four +calibration arms scored on reconstruction RMSE. Under the stated premise +that is vacuous — a deliberately wrong calibration produces bad error BY +DEFINITION, so `CAL-ABS-FOREIGN` is the measurement CONDITION, not a +competitor that might lose. + +**What changed.** +- §2 is now a **4×4 donor × target transfer matrix** per geometry arm. + Primary metric **Spearman `ρ`** (structure preservation); the quantity the + hypothesis concerns is the derived **transfer loss** + `L[D][T] = ρ[T][T] − ρ[D][T]`. RMSE/bias in Pa are retained per cell as + evidence the swap genuinely hurt — never as the verdict. +- `CAL-ABS-OWN` / `CAL-ABS-FOREIGN` stop being two arms: they are the + diagonal and the off-diagonal of ONE arm's matrix. Splitting them was the + tell that the design was still a race. +- The dynamic arms' rows are **flat by construction** (`L ≡ 0`, no donor + exists) — that is the property under test, and C2 now requires BOTH that + it hold exactly AND that the identical code path be proven able to produce + a non-flat row (the can-it-DIFFER gate W5 already paid for). +- Two NEW bars for the second operator message: **C1b** measures constancy + (`separation = between-box range / mean within-box σ ≥ 3`) instead of + claiming it; **C1c** gives the suitability ASSUMPTION a falsifier — the + regimes must differ in autocorrelation decay and rank-distribution shape, + not merely in `|∇p|`, or the ladder is four copies of one condition. +- The crossover bar now scores against the **diagonal** (own-calibration, + the hardest opponent) on `ρ`; the weak form against the swapped cell is + reported separately and labelled as evidence of wiring, not of merit. + +**What did NOT change:** the §1 regime ladder and its three preflight +corrections, the equal-budget discipline, the geometry axis, the C0 control +gate, the output contract's store-the-operands rule, and §5's +non-claims. The scaffold was sound; the question was inverted. + +**Preserved, not deleted:** v1's framing survives in place as §2's +correction note, so a future session can see what the design was and why it +moved. Finding recorded as `E-A-HORSE-RACE-IS-NOT-A-CROSS-SWAP-1`. + ## 2026-08-12 — substrate-comfort-zones-v1 (PLAN; where does each substrate formula feel at home?) Plan: `.claude/plans/substrate-comfort-zones-v1.md`. Status **ACTIVE**, diff --git a/.claude/board/STATUS_BOARD.md b/.claude/board/STATUS_BOARD.md index af13cd0c..5fa7e4e0 100644 --- a/.claude/board/STATUS_BOARD.md +++ b/.claude/board/STATUS_BOARD.md @@ -1,7 +1,8 @@ ## substrate-comfort-zones-v1 — the comfort-zone map (PRE-REGISTERED 2026-08-12) -Plan: `.claude/plans/substrate-comfort-zones-v1.md`. Regime × geometry × -calibration → where does each substrate formula feel at home. §1 preflight +Plan: `.claude/plans/substrate-comfort-zones-v1.md`. A **cross-swap** +(donor × target) transfer matrix per geometry arm → where does each +substrate formula feel at home, and how badly does it travel. §1 preflight already run and it corrected two regime definitions before any bar existed. | D-id | Deliverable | Status | Feeds | @@ -9,10 +10,25 @@ already run and it corrected two regime definitions before any bar existed. | D-CZ-0 | §1 regime preflight (`\|∇p\|` ladder, elevation-confound screen, speed-is-not-the-discriminator finding) | **DONE** — ladder R1 Amazon 10.2 → R2 ocean 14.9 → R3 W Siberia 43.8 → R4 storm 95.6 (9.3× range); 4 land candidates excluded on elev σ > 150 m | the regime axis all other rows score on | | D-CZ-1 | C0 controls (shuffled codebook + degenerate geometry), losability-smoke-tested BEFORE the full run | Queued | gates every cell — a control that can't lose voids its cell | | D-CZ-2 | C1 regime-ladder stability across ≥3 timesteps | Queued | anti-cherry-pick on the whole regime axis | -| D-CZ-3 | **C2 the crossover** — `RANK-DYN − ABS-OWN` sign flip calm↔storm | Queued | **the operator's hypothesis, two-sided** | -| D-CZ-4 | C3 miscalibration penalty vs turbulence (`ABS-FOREIGN`/`ABS-OWN` ratio per tier) | Queued | "storms forgive bad calibration" | -| D-CZ-5 | C4 geometry floor on a SAMPLING-fidelity metric (NULL expected-plausible per W5 B4) | Queued | does the index floor bite where smoothing didn't | -| D-CZ-6 | C5 the comfort matrix (the deliverable) | Queued | read-off answer to "where is each formula at home" | +| D-CZ-2b | **C1b constancy is measured** — `separation = (between-box range)/(mean within-box σ)` ≥ 3 | Queued | earns the phrase "held constant"; otherwise a caveat rides every cell | +| D-CZ-2c | **C1c the suitability ASSUMPTION** — regimes must differ in autocorrelation decay + rank-distribution shape, not only in `\|∇p\|` | Queued | a null here VOIDS the cross-swap reading (spread manufactured on the wrong axis) | +| D-CZ-3 | **C2 degenerate-row verification, both halves** — dynamic arms `L ≡ 0` exactly, AND `CAL-ABS` proven non-degenerate through the same code path | Queued | the can-it-DIFFER gate; without half (ii) a flat row proves nothing | +| D-CZ-4 | C3 **transfer loss** vs turbulence — `L̄[T] = mean_{D≠T} L[D][T]` strictly smaller in R4 than R1, reported with `occupancy`/`saturation` | Queued | "storms forgive bad calibration", with a mechanism attached | +| D-CZ-5 | **C4 the crossover** — `ρ(CAL-RANK) − ρ(CAL-ABS, D=T)` sign flip calm↔storm (vs the DIAGONAL, the hardest opponent) | Queued | **the operator's hypothesis, two-sided**; weak form vs `D≠T` reported separately | +| D-CZ-6 | C5 geometry floor on a SAMPLING-fidelity metric (NULL expected-plausible per W5 B4) | Queued | does the index floor bite where smoothing didn't | +| D-CZ-7 | C6 the transfer matrix (the deliverable) | Queued | comfort read off the diagonal, travel-cost off the off-diagonal | + +> **2026-08-12 — rows re-cut, not restated.** The operator ruled that the +> plan's §2 was built as a horse race where a cross-swap diagnostic belongs +> (premise: the model *captures* the phenomenon but is *not calibrated*, so +> miscalibration is the condition of measurement, not an arm). §2 was +> rebuilt around the 4×4 donor×target transfer matrix and the bars were +> renumbered; two new bars (C1b, C1c) exist because "held constant" and +> "the manufactured spread is on the right axis" are now measured rather +> than assumed. **Every row above was still Queued when this happened — no +> measured result was reinterpreted.** Old→new: C2→C4 (crossover, now +> against the diagonal on `ρ` rather than against `ABS-OWN` on RMSE), +> C3→C3 (now transfer loss, not an RMSE ratio), C4→C5, C5→C6. ## golden-vs-tempered-stride-v1 — head-vs-gut queue — RUN 2026-08-12 diff --git a/.claude/plans/substrate-comfort-zones-v1.md b/.claude/plans/substrate-comfort-zones-v1.md index 34117d0a..3b481010 100644 --- a/.claude/plans/substrate-comfort-zones-v1.md +++ b/.claude/plans/substrate-comfort-zones-v1.md @@ -12,11 +12,26 @@ > 2. *good geometry vs badly calibrated* > 3. *then we find out where the substrate formulas etc. feel at home* > +> **Operator correction (2026-08-12), two further messages that reshaped +> the design — v1's §2 was rebuilt around them, see §2's correction note:** +> 4. the method is **cross-swap and hypothesis testing, under the premise +> that the model captures the phenomenon but is not calibrated** +> 5. in science you hold variables constant to test the others; **constancy +> is relative**, so the design deliberately manufactures strong +> correlation differences — **on the assumption that those differences +> are fit to evaluate the hypothesis** (that assumption gets its own +> falsifier, C1c) +> > **The hypothesis to falsify:** a badly-calibrated substrate that maps -> DYNAMICALLY performs BETTER in strong storms than a well-calibrated -> absolute one — i.e. miscalibration is not uniformly a defect; in a -> high-variance regime an anchor-free adaptive encoding may win precisely -> because the fixed one saturates. +> DYNAMICALLY preserves MORE STRUCTURE in strong storms than a +> well-calibrated absolute one — i.e. miscalibration is not uniformly a +> defect; in a high-variance regime an anchor-free adaptive encoding may +> win precisely because the fixed one saturates. +> +> **Stated as the instrument** (§2): under cross-swap, the absolute +> encoding's transfer loss `L` should shrink as turbulence rises, while the +> dynamic encoding's is zero by construction — so the crossover is a +> statement about `ρ` on the diagonal, not about RMSE anywhere. --- @@ -100,11 +115,91 @@ one exists. --- -## §2 THE TWO AXES (the operator's "good geometry vs badly calibrated") +## §2 THE INSTRUMENT — CROSS-SWAP (Kreuztausch), not a horse race + +> **Correction of this plan's first draft (operator, 2026-08-12).** v1 §2 +> was built as a *comparison of formulas* — four calibration arms scored +> against each other on RMSE. That reads the premise backwards. The premise +> is: **assume the model DOES capture the phenomenon, but is NOT +> calibrated.** Under that premise miscalibration is the *condition of +> measurement*, not one arm in a race — and RMSE under a deliberately wrong +> calibration is bad **by definition**, so scoring it answers nothing. The +> informative quantity is how much of the captured **structure** survives +> the swap. The regime ladder, the budget discipline and the control gate +> from v1 are unchanged; only the question is inverted. + +### §2.1 What is held constant, what is swapped + +The operator's methodological frame: *you hold variables constant to test +the others; constancy is relative, so the design deliberately manufactures +strong correlation differences — on the assumption that those differences +are fit to evaluate the hypothesis.* Both halves are load-bearing here, and +the assumption in the second half is given its own falsifier (C1c). + +Each cell holds everything fixed except one thing: + +| held constant | varied | +|---|---| +| the box (regime), the timestep, the geometry arm, the sample count | **only the calibration's donor regime** | + +Notation: `M[D][T]` = read regime **T**'s field through a codebook derived +from regime **D**. The diagonal `D = T` is own-calibration; every +off-diagonal cell is a swap. One full matrix per geometry arm, so geometry +never blends into the calibration answer. + +**Constancy is operationalized, never assumed** (C1b): the within-box +spread of the discriminator must be small relative to the between-box +spread. A box that fails that is not a control condition — it is just +another data point wearing the label. + +### §2.2 The matrix and its metrics + +The 4×4 `M[D][T]` over `{R1 CALM, R2 OCEAN, R3 ACTIVE, R4 STORM}`. Per +cell, measured and stored raw: + +| quantity | what it answers | +|---|---| +| **`ρ` Spearman(reconstructed, true)** | **PRIMARY** — how much ordering survived | +| `occupancy` = fraction of the 256 levels actually used | the mechanism: a foreign codebook collapses the field onto few levels | +| `saturation` = fraction clipped at level 0 or 255 | the other half of the mechanism: the field runs off the donor's range | +| RMSE / bias in Pa | **secondary** — evidence the swap genuinely hurts in absolute terms, never the verdict | -The two axes are **orthogonal by construction** and are varied -independently, so a result can be attributed to one or the other rather -than to their blend. +**The derived quantity the whole plan turns on:** + +``` +transfer loss L[D][T] = ρ[T][T] − ρ[D][T] +``` + +— how much structure regime `T` loses when read through `D`'s calibration. +`L` is the cross-swap statement of "badly calibrated": *low* `L` means the +substrate carried the structure even though the numbers were wrong. + +### §2.3 The dynamic row is degenerate — and that IS the point + +A rank-normalised or Fisher-z encoding re-derived **inside the window** has +no donor at all. Its matrix row is therefore identical across `D` and +`L ≡ 0` **by construction**. That is not a defect to hide; it is precisely +the property under test — a dynamically-mapped substrate cannot be +mis-calibrated, because it carries no absolute anchor to get wrong. + +Two consequences, both mandatory: + +1. **The degeneracy must be VERIFIED, not asserted.** A nonzero + off-diagonal on a dynamic arm means a donor parameter leaked into the + window — a bug, and the run is void until it is found. +2. **…and the pipeline must be proven able to produce a NON-degenerate + row** through the identical code path (the absolute arms must show real + off-diagonal degradation). A row that is constant because the harness + cannot vary anything is the `E-A-CONTROL-THAT-CANNOT-LOSE` defect in its + can-it-DIFFER form, measured in W5. + +So the real comparison is not "which formula wins" but: + +> In which regime does `ρ(dynamic, own window)` exceed +> `ρ(absolute, foreign donor)` — and does the margin grow with turbulence? + +That is the operator's hypothesis stated as a cross-swap, and it is the +only form in which a mis-calibrated arm can be scored fairly. ### Axis A — GEOMETRY (where the samples sit) @@ -121,20 +216,26 @@ budget silently advantages the arm with more samples — measured there as ### Axis B — CALIBRATION (how the 256 palette levels are placed) -| arm | construction | absolute anchor? | dynamic? | -|---|---|---|---| -| `CAL-ABS-OWN` | 256 uniform levels over THIS box's own min/max | yes | no | -| `CAL-ABS-FOREIGN` | 256 uniform levels over a DIFFERENT regime's min/max | yes, **wrong one** | no | -| `CAL-RANK-DYN` | rank-normalised within the window, re-derived per box | **no** | **yes** | -| `CAL-FISHERZ-DYN` | Fisher-z on within-window ranks (the arc's analytic codebook) | **no** | **yes** | - -`CAL-ABS-FOREIGN` is the literal reading of "badly calibrated"; -`CAL-RANK-DYN` / `CAL-FISHERZ-DYN` are "badly calibrated in absolute terms -BUT dynamically mapping" — the operator's actual candidate. +Three encodings, each run through the **full 4×4 donor matrix**: -**Metric:** reconstruction RMSE in **Pa** (the physical unit, per the -`E-R²-IS-NEAR-BLIND` lesson — never R² alone), plus Spearman ρ of the -reconstructed vs true field, plus mean **bias** in Pa. +| arm | construction | absolute anchor? | matrix shape | +|---|---|---|---| +| `CAL-ABS` | 256 uniform levels over donor `D`'s min/max | yes | **full** — diagonal = own, off-diagonal = swap | +| `CAL-RANK` | rank-normalised within the target window | **no** | **degenerate** — `L ≡ 0` by construction | +| `CAL-FISHERZ` | Fisher-z on within-window ranks (the arc's analytic codebook) | **no** | **degenerate** — same | + +v1 listed `CAL-ABS-OWN` and `CAL-ABS-FOREIGN` as two separate arms. They +are not two arms — they are the **diagonal and the off-diagonal of one +arm's matrix**, and splitting them was the tell that the design was still a +race. `CAL-ABS` with `D = T` is own-calibration; `CAL-ABS` with `D ≠ T` is +the literal "badly calibrated"; the two dynamic arms are "badly calibrated +in absolute terms BUT dynamically mapping" — the operator's actual +candidate, and the ones whose rows are flat by construction. + +**Metrics:** primary `ρ` and the derived transfer loss `L` (§2.2); +secondary RMSE and mean **bias** in **Pa** (the physical unit, per the +`E-R²-IS-NEAR-BLIND` lesson — never R² alone), reported for every cell but +never used as the verdict. --- @@ -155,29 +256,67 @@ reconstructed vs true field, plus mean **bias** in Pa. R1 < R2 < R3 < R4 must hold on **≥3 independent timesteps**, not just the preflight's one. If the ladder inverts on any timestep, the regime axis is not stable and every downstream cell is reported with that caveat. -- **C2 THE CROSSOVER — the operator's hypothesis, two-sided:** - `Δ = RMSE(CAL-RANK-DYN) − RMSE(CAL-ABS-OWN)` must be **> 0 in R1/R2 - (calm: dynamic loses) AND < 0 in R4 (storm: dynamic wins)** — a genuine - sign flip. **Both failure directions are reportable results, not - disappointments:** no flip = the hypothesis is refuted on this data and - says so; flip in the *opposite* direction = dynamic encoding is a - calm-regime tool, which would be a real and surprising finding. -- **C3 THE MISCALIBRATION PENALTY SHRINKS WITH TURBULENCE:** the ratio - `RMSE(CAL-ABS-FOREIGN) / RMSE(CAL-ABS-OWN)` must be **strictly smaller in - R4 than in R1** — the direct statement of "storms are more forgiving of - bad calibration." Reported with the ratio at every tier, so a monotone - trend (or its absence) is visible rather than inferred from two endpoints. -- **C4 GEOMETRY FLOOR BITES HERE (or it does not):** `GEO-GOLDEN-LO` must +- **C1b CONSTANCY IS RELATIVE — so it is MEASURED, not claimed.** For the + discriminator (`|∇p|`), the **within-box** spread must be small relative + to the **between-box** spread: report + `separation = (between-box range) / (mean within-box σ)` and require + **≥ 3**. Below that, "holding the regime constant" is a label rather than + a condition, and every downstream cell inherits that caveat explicitly. + *(This is the operationalization of the operator's point that constancy + is relative — a box is a control condition only insofar as its internal + variation is dominated by the spread the design manufactured.)* +- **C1c THE REGIMES MUST DIFFER IN CORRELATION STRUCTURE, not merely in + `|∇p|` — the suitability ASSUMPTION, made falsifiable.** The ladder is + built on a pressure-gradient discriminator, but the hypothesis is about + **structure**. So before any swap runs: measure each box's own + **autocorrelation decay length** and **rank-distribution shape** (Gini / + tail ratio of `|∇p|`). If R1 and R4 are indistinguishable on those, they + are ONE regime for this question no matter how far apart their gradients + are, and the ladder measures four copies of the same condition. + **Pre-registered honest reading:** a null here VOIDS the cross-swap + interpretation rather than weakening it — it would mean the manufactured + spread was manufactured on the wrong axis. Reported first, not last. +- **C2 THE DEGENERATE ROW — verified, and proven capable of being + non-degenerate.** Two halves, both required: + (i) every dynamic arm (`CAL-RANK`, `CAL-FISHERZ`) must show `L[D][T] = 0` + for all `D` **exactly**; any nonzero off-diagonal means a donor parameter + leaked into the window and the run is VOID until it is found; + (ii) through the **identical code path**, `CAL-ABS` must show a + **non-zero** off-diagonal in at least one regime. Half (i) alone is the + can-it-DIFFER defect measured in W5 — a row that is flat because the + harness cannot vary anything proves nothing about the encoding. +- **C3 TRANSFER LOSS SHRINKS WITH TURBULENCE:** for `CAL-ABS`, the mean + off-diagonal transfer loss `L̄[T] = mean_{D ≠ T} L[D][T]` must be + **strictly smaller in R4 than in R1** — the cross-swap statement of + "storms are more forgiving of bad calibration." Reported at every tier so + a monotone trend (or its absence) is visible rather than inferred from + two endpoints, and reported **alongside `occupancy` and `saturation`**, so + a shrinking loss can be attributed to a mechanism rather than asserted. +- **C4 THE CROSSOVER — the operator's hypothesis, two-sided:** + `Δ[T] = ρ(CAL-RANK, T) − ρ(CAL-ABS, D=T, T)` must be **< 0 in R1/R2 + (calm: own-calibration absolute wins) AND > 0 in R4 (storm: dynamic + wins)** — a genuine sign flip against the *diagonal*, which is the + hardest available opponent. **Both failure directions are reportable + results, not disappointments:** no flip = the strong hypothesis is + refuted on this data and says so; flip the *other* way = dynamic encoding + is a calm-regime tool, which would be real and surprising. + **The weak form is reported separately and never conflated with it:** + `ρ(CAL-RANK, T) > ρ(CAL-ABS, D ≠ T, T)` — dynamic beats a *mis-calibrated* + absolute. That one is nearly guaranteed by C2(i) and is therefore + evidence of wiring, not of merit. +- **C5 GEOMETRY FLOOR BITES HERE (or it does not):** `GEO-GOLDEN-LO` must be worse than `GEO-GOLDEN-HI` at equal budget. **Pre-registered honest reading:** W5's B4 already found the floor to be a *safety margin, not a mechanism* on a smoothing metric — so a NULL here is expected-plausible and must be reported plainly, not buried. What would be genuinely informative is the floor biting on a *sampling-fidelity* metric where it did not bite on a *smoothing* one. -- **C5 THE COMFORT MATRIX (descriptive, the deliverable):** the full - `regime × (geometry × calibration)` RMSE table, plus each cell normalized - by its regime's best arm — so "where does this formula feel at home" is - read directly off the matrix rather than argued. +- **C6 THE TRANSFER MATRIX (descriptive, the deliverable):** the full + `geometry × (donor × target)` table of `ρ`, `L`, `occupancy`, + `saturation`, RMSE and bias — every cell raw, plus the derived `L̄[T]` + column. "Where does this formula feel at home" is then **read off the + diagonal**, and "how badly does it travel" **off the off-diagonal** — + neither argued. --- @@ -187,10 +326,15 @@ Per the repeated finding that a first artifact ships summaries and omits the operands its headline rests on (W6's per-storm predictors; W5's family-B histogram; the chat-only 99.38 %), the JSON **must** carry: -- every cell's **raw** RMSE / bias / Spearman ρ, in Pa where dimensional -- the per-regime **codebook edges actually used** (so a miscalibration - claim is auditable without a re-fetch) -- the **measured** `|∇p|`, spd σ, elev σ and lsm per box per timestep +- every cell's **raw** `ρ` / `occupancy` / `saturation` / RMSE / bias, + keyed by `(geometry, donor D, target T)` — the full matrix, not the + diagonal plus a summary. `L[D][T]` is DERIVED in the report from stored + `ρ`, never stored alone (per the W6 lesson: store the operands, so a + headline can be re-derived without a re-fetch) +- the per-regime **codebook edges actually used**, for every donor — a + miscalibration claim is auditable only if the wrong codebook is on disk +- the **measured** `|∇p|`, spd σ, elev σ and lsm per box per timestep, + plus C1b's `separation` ratio and C1c's decay length + tail ratio - the **sample count actually drawn** per arm (the equal-budget proof, not the intent) - **units on every dimensional field name**, per the `c_bow`-is-km⁻¹ lesson @@ -217,3 +361,14 @@ already fetched in preflight. silent substitution. - **Not a substitute for CT-F17.** Nothing here touches the directional claim; it is a substrate-fidelity map, a different question entirely. +- **Not a claim that the four boxes are the same condition minus one + knob.** Real regimes differ in more than the discriminator. C1b bounds + how far the "held constant" label is earned, C1c bounds whether the + manufactured spread lies on the axis the hypothesis is about — and + whatever those two report travels with every downstream number rather + than being dropped once the matrix is filled. +- **Not a claim that transfer loss isolates calibration alone.** `L` is + measured with `occupancy` and `saturation` beside it precisely because a + shrinking `L` could also mean the target's field happens to sit inside + the donor's range by luck of that timestep. Three timesteps bound that; + they do not eliminate it. From 522efab49ade72a2b3a8263740c675987d691362 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 12 Aug 2026 20:14:38 +0000 Subject: [PATCH 2/3] plan: pre-register the negative-transfer-loss case before it can be explained away Self-check while verifying the sign conventions (all three correct: low L = structure survived; C3's L-bar shrinking = storms forgive miscalibration; C4's inequality correctly inverted from v1's RMSE convention to rho's higher-is-better) surfaced a gap the conventions themselves hide. L = rho[T][T] - rho[D][T] is NOT guaranteed non-negative. Nothing forces the diagonal to be the best cell: a box whose own min/max is set by one outlier spreads its 256 levels badly, while a donor with a wider range may spread them better -- so a foreign codebook can legitimately beat the own one. Pre-registered as informative rather than as a bug, with the consequence stated: if it occurs, rho[T][T] is not a valid reference point for that box's L-bar, so C3's trend must be re-read against the best-available cell and the substitution said out loud. occupancy/saturation are reported beside it because that is where the mechanism would show. Naming it now is the whole point -- an unexpected sign discovered mid-analysis is precisely the result that gets rationalized after the fact. Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi --- .claude/plans/substrate-comfort-zones-v1.md | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/.claude/plans/substrate-comfort-zones-v1.md b/.claude/plans/substrate-comfort-zones-v1.md index 3b481010..9d84fc6a 100644 --- a/.claude/plans/substrate-comfort-zones-v1.md +++ b/.claude/plans/substrate-comfort-zones-v1.md @@ -174,6 +174,18 @@ transfer loss L[D][T] = ρ[T][T] − ρ[D][T] `L` is the cross-swap statement of "badly calibrated": *low* `L` means the substrate carried the structure even though the numbers were wrong. +**`L < 0` is possible and is pre-registered as informative, not as a bug.** +Nothing forces the diagonal to be the best cell: a box whose own min/max is +set by a single outlier spreads its 256 levels badly, while a donor with a +wider range may spread them better — so a *foreign* codebook can beat the +own one. If that occurs it is reported as measured, with `occupancy` and +`saturation` beside it (which is where the mechanism would show), and it +carries a specific consequence: **`ρ[T][T]` is then not a valid reference +point for that box's `L̄[T]`, so C3's trend must be re-read against the +best-available cell and the substitution stated.** Naming this in advance is +the point — an unexpected sign discovered mid-analysis is exactly the kind +of result that gets explained away. + ### §2.3 The dynamic row is degenerate — and that IS the point A rank-normalised or Fisher-z encoding re-derived **inside the window** has From 1756c518c16ae8b8dea2136aaf21d5b7e2ec45d4 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 12 Aug 2026 20:16:00 +0000 Subject: [PATCH 3/3] =?UTF-8?q?plan:=20fix=20the=20contradiction=20CodeRab?= =?UTF-8?q?bit=20caught=20--=20=C2=A72.3=20named=20the=20WEAK=20form=20as?= =?UTF-8?q?=20the=20hypothesis?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Blocking consistency issue on #944, and it is the exact defect this PR exists to remove. §2.3's summary passage defined the hypothesis as rho(dynamic, own window) > rho(absolute, FOREIGN donor). C4, STATUS_BOARD and the INTEGRATION_PLANS entry all score against the DIAGONAL and demote the foreign comparison to a wiring check. Mechanism: §2.3 was written before C4 was sharpened, and the sharpening was not propagated backwards. Left standing, the plan carried two incompatible verdict criteria with the WEAKER one holding the headline -- the swapped cell establishing merit, which is precisely the tautology §2 was rebuilt to delete. C2(i) makes the dynamic arm win the foreign comparison almost by construction: an encoding with no anchor cannot be out-transferred by one whose anchor is deliberately wrong. Fixed in place, correction note kept beside it rather than a silent patch. EPIPHANIES entry extended (unmerged, composes in place; suffix-verified): the defect is STICKY -- it survived the rewrite built to remove it, and a reviewer caught it, not the author. Fifth instance in this arc of a claim consistent with its own operands but inconsistent with a sibling claim. Generalization recorded: when a design is revised, the summary sentences are the last thing updated and the first thing a reader believes; a revision is complete only when every sentence that NAMES the conclusion has been re-derived from the new machinery. The other four requested checks came back consistent and are unchanged: the L sign convention, C4's inverted inequality, the C-bar-to-D-CZ 1:1 mapping, and C2's both-halves requirement. Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi --- .claude/board/EPIPHANIES.md | 21 +++++++++++++++++ .claude/plans/substrate-comfort-zones-v1.md | 25 ++++++++++++++++++++- 2 files changed, 45 insertions(+), 1 deletion(-) diff --git a/.claude/board/EPIPHANIES.md b/.claude/board/EPIPHANIES.md index 8f9b6262..5b95b4d1 100644 --- a/.claude/board/EPIPHANIES.md +++ b/.claude/board/EPIPHANIES.md @@ -66,6 +66,27 @@ first: thing. A null there VOIDS the reading rather than weakening it: it would mean the spread was manufactured on the wrong axis. +**The defect is STICKY — it survived the rewrite built to remove it, and a +REVIEWER caught it, not the author (CodeRabbit on #944, same day).** §2.3's +summary passage still named the *foreign-donor* comparison as "the +operator's hypothesis" while C4, `STATUS_BOARD` and the `INTEGRATION_PLANS` +entry all correctly scored against the diagonal. Mechanism: §2.3 was written +before C4 was sharpened, and the sharpening was not propagated backwards. +Left standing, the plan would have carried **two incompatible verdict +criteria with the weaker one holding the headline** — the swapped cell +establishing merit, i.e. the exact tautology the rewrite existed to delete. + +That makes this the **fifth** instance in this arc of a claim that is +internally consistent with its own operands but inconsistent with a SIBLING +claim (#930's fused relation, #927's decimal claim, #928's audit figures, +#941's dropped qualifier, now this). The generalization is sharper than the +individual fixes: **when a design is revised, the summary sentences are the +last thing to be updated and the first thing a reader believes.** A revision +is not complete when the new machinery is right; it is complete when every +sentence that *names the conclusion* has been re-derived from the new +machinery. Self-review kept missing it because the author re-reads the part +that changed, not the part that merely still sounds right. + **The transferable rule.** Before scoring arms against each other, ask which of them is a *competitor* and which is a *condition*. A condition put on the starting line will always lose, and its loss carries no information. Score diff --git a/.claude/plans/substrate-comfort-zones-v1.md b/.claude/plans/substrate-comfort-zones-v1.md index 9d84fc6a..ce5b6717 100644 --- a/.claude/plans/substrate-comfort-zones-v1.md +++ b/.claude/plans/substrate-comfort-zones-v1.md @@ -208,11 +208,34 @@ Two consequences, both mandatory: So the real comparison is not "which formula wins" but: > In which regime does `ρ(dynamic, own window)` exceed -> `ρ(absolute, foreign donor)` — and does the margin grow with turbulence? +> **`ρ(absolute, OWN donor — the diagonal)`** — and does the margin grow +> with turbulence? + +**Against the DIAGONAL, deliberately: it is the hardest available +opponent.** The comparison against a *foreign* donor is the weak form. It +is reported separately (C4) and labelled as evidence of **wiring, not of +merit**, because C2(i) makes the dynamic arm win it almost by construction +— an encoding with no anchor cannot be out-transferred by one whose anchor +is deliberately wrong. Scoring the hypothesis on the swapped cell would +re-introduce exactly the tautology this whole section removes. That is the operator's hypothesis stated as a cross-swap, and it is the only form in which a mis-calibrated arm can be scored fairly. +> **Caught by review, and it is the same defect one level up (2026-08-12, +> CodeRabbit on #944).** The first version of this passage named the +> foreign-donor comparison as "the operator's hypothesis" — contradicting +> C4, `STATUS_BOARD`, and the `INTEGRATION_PLANS` entry, all of which +> correctly score against the diagonal. The passage was written BEFORE C4 +> was sharpened, and the sharpening was not propagated back. Left standing, +> the plan would have carried two incompatible verdict criteria with the +> weaker one holding the headline — the swapped cell establishing merit, +> which is the precise thing §2 exists to prevent. Recorded rather than +> silently patched: it is the fifth instance of a claim that is consistent +> with its own operands but inconsistent with a sibling claim (cf. #930's +> fused relation, #927's decimal claim, #928's audit figures, #941's +> dropped qualifier), and it is a REVIEWER who caught it, not the author. + ### Axis A — GEOMETRY (where the samples sit) | arm | construction | a-priori quality |