diff --git a/.claude/board/EPIPHANIES.md b/.claude/board/EPIPHANIES.md index 8e466264..5b95b4d1 100644 --- a/.claude/board/EPIPHANIES.md +++ b/.claude/board/EPIPHANIES.md @@ -1,3 +1,104 @@ +## 2026-08-12 — E-A-HORSE-RACE-IS-NOT-A-CROSS-SWAP-1 + +**Status:** FINDING `[H]` — methodological, operator-ruled. Caught on +`substrate-comfort-zones-v1` BEFORE any bar ran, so nothing measured had to +be reinterpreted. Grade is `[H]` because the reasoning is airtight given the +premise but the premise itself ("the model captures the phenomenon") is an +assumption this plan does not test. + +**The defect, in one line: when the premise is *the model captures the +phenomenon but is not calibrated*, an error metric measured under a +deliberately wrong calibration is bad BY DEFINITION, so scoring it answers +nothing.** The first draft of that plan compared four calibration arms on +reconstruction RMSE — a horse race. But `CAL-ABS-FOREIGN` is not a +competitor that might lose; it is the measurement CONDITION the whole +hypothesis is about. Racing it guarantees the answer and measures the +tautology. + +**What the design must be instead.** Miscalibration is a *transfer* +question, so the instrument is a transfer matrix, not a ranking: `M[D][T]` = +read regime `T`'s field through a codebook derived from regime `D`. The +diagonal is own-calibration; every off-diagonal cell is a swap. The primary +metric becomes **structure preservation** (Spearman `ρ` of reconstructed vs +true), and the quantity the hypothesis actually concerns is the derived +**transfer loss** `L[D][T] = ρ[T][T] − ρ[D][T]` — how much ordering survives +being read through the wrong calibration. Absolute error is retained as +*evidence that the swap genuinely hurt*, never as the verdict. + +**The sharp consequence that only appears in the matrix framing.** A +window-local encoding (rank-normalised, Fisher-z) has **no donor at all**, +so its matrix row is flat and `L ≡ 0` **by construction**. That is not an +artifact to hide — it IS the property under test: a substrate that carries +no absolute anchor cannot have that anchor be wrong. Which reframes the +whole comparison. The question stops being "which formula wins" and becomes +"in which regime does the anchor-free encoding's own-window `ρ` exceed the +absolute encoding's, measured against the DIAGONAL (its own calibration, +the hardest available opponent) rather than against the swapped cell (which +it beats almost by definition)." + +**And the degeneracy must be VERIFIED both ways** — this is the +`E-A-CONTROL-THAT-CANNOT-LOSE-IS-NO-CONTROL-1` lesson in its can-it-DIFFER +form, which W5 already paid for once. A flat row is only evidence if the +identical code path is proven able to produce a NON-flat one (the absolute +arm must show real off-diagonal degradation). Otherwise "flat" means the +harness cannot vary anything, and the encoding was never tested. + +**Second half of the ruling, on the regime ladder itself.** The operator's +framing: *in science you hold variables constant to test the others; +constancy is relative, so the design deliberately manufactures strong +correlation differences — on the assumption that those differences are fit +to evaluate the hypothesis.* Three things follow, and the plan had only the +first: + +1. **The 9.3× `|∇p|` spread across the four regimes is the design's PURPOSE, + not a by-product of box selection.** At small spread every effect is in + the noise; at large spread it is legible or it does not exist. +2. **"Constant" is relative, so it is MEASURED** — within-box spread of the + discriminator must be small relative to the between-box spread + (`separation ≥ 3`). Below that, "held constant" is a label on a box, not + a condition, and the caveat must ride every downstream number. +3. **"Fit to evaluate the hypothesis" is an ASSUMPTION and therefore needs a + falsifier.** The ladder is built on a pressure-gradient discriminator but + the hypothesis is about correlation STRUCTURE. If two regimes are + indistinguishable in autocorrelation decay and rank-distribution shape, + they are ONE condition for this question however far apart their + gradients sit — and the ladder would be measuring four copies of the same + thing. A null there VOIDS the reading rather than weakening it: it would + mean the spread was manufactured on the wrong axis. + +**The defect is STICKY — it survived the rewrite built to remove it, and a +REVIEWER caught it, not the author (CodeRabbit on #944, same day).** §2.3's +summary passage still named the *foreign-donor* comparison as "the +operator's hypothesis" while C4, `STATUS_BOARD` and the `INTEGRATION_PLANS` +entry all correctly scored against the diagonal. Mechanism: §2.3 was written +before C4 was sharpened, and the sharpening was not propagated backwards. +Left standing, the plan would have carried **two incompatible verdict +criteria with the weaker one holding the headline** — the swapped cell +establishing merit, i.e. the exact tautology the rewrite existed to delete. + +That makes this the **fifth** instance in this arc of a claim that is +internally consistent with its own operands but inconsistent with a SIBLING +claim (#930's fused relation, #927's decimal claim, #928's audit figures, +#941's dropped qualifier, now this). The generalization is sharper than the +individual fixes: **when a design is revised, the summary sentences are the +last thing to be updated and the first thing a reader believes.** A revision +is not complete when the new machinery is right; it is complete when every +sentence that *names the conclusion* has been re-derived from the new +machinery. Self-review kept missing it because the author re-reads the part +that changed, not the part that merely still sounds right. + +**The transferable rule.** Before scoring arms against each other, ask which +of them is a *competitor* and which is a *condition*. A condition put on the +starting line will always lose, and its loss carries no information. Score +conditions by what SURVIVES them, not by how much they cost. + +**Cross-ref:** `.claude/plans/substrate-comfort-zones-v1.md` §2 (the rebuilt +instrument, with v1's framing preserved in place as a correction note rather +than deleted) and §3 C1b/C1c/C2/C3/C4; `STATUS_BOARD.md` D-CZ-* (rows re-cut +while still Queued, with the old→new mapping recorded); +`E-A-CONTROL-THAT-CANNOT-LOSE-IS-NO-CONTROL-1`; +`E-R²-IS-NEAR-BLIND` (why the physical unit is retained alongside `ρ`). + ## 2026-08-12 — E-THE-DISPLACEMENT-FILTER-ATE-THE-STRANDED-STRATUM-1 **Status:** FINDING `[G]` — W6 RUN (`comet_tail_w6.py`/`.json`), audited diff --git a/.claude/board/INTEGRATION_PLANS.md b/.claude/board/INTEGRATION_PLANS.md index 5340b616..de6ac45c 100644 --- a/.claude/board/INTEGRATION_PLANS.md +++ b/.claude/board/INTEGRATION_PLANS.md @@ -1,3 +1,53 @@ +## 2026-08-12 — substrate-comfort-zones-v1 REVISED (§2 rebuilt: horse race → cross-swap) + +Same file, `.claude/plans/substrate-comfort-zones-v1.md`, still **ACTIVE**, +still exploratory tier. Not a new plan and not a v2 — a pre-registration +sharpened **before any bar ran** (only D-CZ-0, the regime preflight, was +DONE; everything else was Queued), which is precisely when a design should +be allowed to change. + +**Why.** Operator ruling, two messages: the method is *cross-swap and +hypothesis testing, under the premise that the model captures the phenomenon +but is not calibrated*; and *in science you hold variables constant to test +the others — constancy is relative, so the design deliberately manufactures +strong correlation differences, on the assumption those differences are fit +to evaluate the hypothesis*. Against that, v1 §2 was a **horse race**: four +calibration arms scored on reconstruction RMSE. Under the stated premise +that is vacuous — a deliberately wrong calibration produces bad error BY +DEFINITION, so `CAL-ABS-FOREIGN` is the measurement CONDITION, not a +competitor that might lose. + +**What changed.** +- §2 is now a **4×4 donor × target transfer matrix** per geometry arm. + Primary metric **Spearman `ρ`** (structure preservation); the quantity the + hypothesis concerns is the derived **transfer loss** + `L[D][T] = ρ[T][T] − ρ[D][T]`. RMSE/bias in Pa are retained per cell as + evidence the swap genuinely hurt — never as the verdict. +- `CAL-ABS-OWN` / `CAL-ABS-FOREIGN` stop being two arms: they are the + diagonal and the off-diagonal of ONE arm's matrix. Splitting them was the + tell that the design was still a race. +- The dynamic arms' rows are **flat by construction** (`L ≡ 0`, no donor + exists) — that is the property under test, and C2 now requires BOTH that + it hold exactly AND that the identical code path be proven able to produce + a non-flat row (the can-it-DIFFER gate W5 already paid for). +- Two NEW bars for the second operator message: **C1b** measures constancy + (`separation = between-box range / mean within-box σ ≥ 3`) instead of + claiming it; **C1c** gives the suitability ASSUMPTION a falsifier — the + regimes must differ in autocorrelation decay and rank-distribution shape, + not merely in `|∇p|`, or the ladder is four copies of one condition. +- The crossover bar now scores against the **diagonal** (own-calibration, + the hardest opponent) on `ρ`; the weak form against the swapped cell is + reported separately and labelled as evidence of wiring, not of merit. + +**What did NOT change:** the §1 regime ladder and its three preflight +corrections, the equal-budget discipline, the geometry axis, the C0 control +gate, the output contract's store-the-operands rule, and §5's +non-claims. The scaffold was sound; the question was inverted. + +**Preserved, not deleted:** v1's framing survives in place as §2's +correction note, so a future session can see what the design was and why it +moved. Finding recorded as `E-A-HORSE-RACE-IS-NOT-A-CROSS-SWAP-1`. + ## 2026-08-12 — substrate-comfort-zones-v1 (PLAN; where does each substrate formula feel at home?) Plan: `.claude/plans/substrate-comfort-zones-v1.md`. Status **ACTIVE**, diff --git a/.claude/board/STATUS_BOARD.md b/.claude/board/STATUS_BOARD.md index af13cd0c..5fa7e4e0 100644 --- a/.claude/board/STATUS_BOARD.md +++ b/.claude/board/STATUS_BOARD.md @@ -1,7 +1,8 @@ ## substrate-comfort-zones-v1 — the comfort-zone map (PRE-REGISTERED 2026-08-12) -Plan: `.claude/plans/substrate-comfort-zones-v1.md`. Regime × geometry × -calibration → where does each substrate formula feel at home. §1 preflight +Plan: `.claude/plans/substrate-comfort-zones-v1.md`. A **cross-swap** +(donor × target) transfer matrix per geometry arm → where does each +substrate formula feel at home, and how badly does it travel. §1 preflight already run and it corrected two regime definitions before any bar existed. | D-id | Deliverable | Status | Feeds | @@ -9,10 +10,25 @@ already run and it corrected two regime definitions before any bar existed. | D-CZ-0 | §1 regime preflight (`\|∇p\|` ladder, elevation-confound screen, speed-is-not-the-discriminator finding) | **DONE** — ladder R1 Amazon 10.2 → R2 ocean 14.9 → R3 W Siberia 43.8 → R4 storm 95.6 (9.3× range); 4 land candidates excluded on elev σ > 150 m | the regime axis all other rows score on | | D-CZ-1 | C0 controls (shuffled codebook + degenerate geometry), losability-smoke-tested BEFORE the full run | Queued | gates every cell — a control that can't lose voids its cell | | D-CZ-2 | C1 regime-ladder stability across ≥3 timesteps | Queued | anti-cherry-pick on the whole regime axis | -| D-CZ-3 | **C2 the crossover** — `RANK-DYN − ABS-OWN` sign flip calm↔storm | Queued | **the operator's hypothesis, two-sided** | -| D-CZ-4 | C3 miscalibration penalty vs turbulence (`ABS-FOREIGN`/`ABS-OWN` ratio per tier) | Queued | "storms forgive bad calibration" | -| D-CZ-5 | C4 geometry floor on a SAMPLING-fidelity metric (NULL expected-plausible per W5 B4) | Queued | does the index floor bite where smoothing didn't | -| D-CZ-6 | C5 the comfort matrix (the deliverable) | Queued | read-off answer to "where is each formula at home" | +| D-CZ-2b | **C1b constancy is measured** — `separation = (between-box range)/(mean within-box σ)` ≥ 3 | Queued | earns the phrase "held constant"; otherwise a caveat rides every cell | +| D-CZ-2c | **C1c the suitability ASSUMPTION** — regimes must differ in autocorrelation decay + rank-distribution shape, not only in `\|∇p\|` | Queued | a null here VOIDS the cross-swap reading (spread manufactured on the wrong axis) | +| D-CZ-3 | **C2 degenerate-row verification, both halves** — dynamic arms `L ≡ 0` exactly, AND `CAL-ABS` proven non-degenerate through the same code path | Queued | the can-it-DIFFER gate; without half (ii) a flat row proves nothing | +| D-CZ-4 | C3 **transfer loss** vs turbulence — `L̄[T] = mean_{D≠T} L[D][T]` strictly smaller in R4 than R1, reported with `occupancy`/`saturation` | Queued | "storms forgive bad calibration", with a mechanism attached | +| D-CZ-5 | **C4 the crossover** — `ρ(CAL-RANK) − ρ(CAL-ABS, D=T)` sign flip calm↔storm (vs the DIAGONAL, the hardest opponent) | Queued | **the operator's hypothesis, two-sided**; weak form vs `D≠T` reported separately | +| D-CZ-6 | C5 geometry floor on a SAMPLING-fidelity metric (NULL expected-plausible per W5 B4) | Queued | does the index floor bite where smoothing didn't | +| D-CZ-7 | C6 the transfer matrix (the deliverable) | Queued | comfort read off the diagonal, travel-cost off the off-diagonal | + +> **2026-08-12 — rows re-cut, not restated.** The operator ruled that the +> plan's §2 was built as a horse race where a cross-swap diagnostic belongs +> (premise: the model *captures* the phenomenon but is *not calibrated*, so +> miscalibration is the condition of measurement, not an arm). §2 was +> rebuilt around the 4×4 donor×target transfer matrix and the bars were +> renumbered; two new bars (C1b, C1c) exist because "held constant" and +> "the manufactured spread is on the right axis" are now measured rather +> than assumed. **Every row above was still Queued when this happened — no +> measured result was reinterpreted.** Old→new: C2→C4 (crossover, now +> against the diagonal on `ρ` rather than against `ABS-OWN` on RMSE), +> C3→C3 (now transfer loss, not an RMSE ratio), C4→C5, C5→C6. ## golden-vs-tempered-stride-v1 — head-vs-gut queue — RUN 2026-08-12 diff --git a/.claude/plans/substrate-comfort-zones-v1.md b/.claude/plans/substrate-comfort-zones-v1.md index 34117d0a..ce5b6717 100644 --- a/.claude/plans/substrate-comfort-zones-v1.md +++ b/.claude/plans/substrate-comfort-zones-v1.md @@ -12,11 +12,26 @@ > 2. *good geometry vs badly calibrated* > 3. *then we find out where the substrate formulas etc. feel at home* > +> **Operator correction (2026-08-12), two further messages that reshaped +> the design — v1's §2 was rebuilt around them, see §2's correction note:** +> 4. the method is **cross-swap and hypothesis testing, under the premise +> that the model captures the phenomenon but is not calibrated** +> 5. in science you hold variables constant to test the others; **constancy +> is relative**, so the design deliberately manufactures strong +> correlation differences — **on the assumption that those differences +> are fit to evaluate the hypothesis** (that assumption gets its own +> falsifier, C1c) +> > **The hypothesis to falsify:** a badly-calibrated substrate that maps -> DYNAMICALLY performs BETTER in strong storms than a well-calibrated -> absolute one — i.e. miscalibration is not uniformly a defect; in a -> high-variance regime an anchor-free adaptive encoding may win precisely -> because the fixed one saturates. +> DYNAMICALLY preserves MORE STRUCTURE in strong storms than a +> well-calibrated absolute one — i.e. miscalibration is not uniformly a +> defect; in a high-variance regime an anchor-free adaptive encoding may +> win precisely because the fixed one saturates. +> +> **Stated as the instrument** (§2): under cross-swap, the absolute +> encoding's transfer loss `L` should shrink as turbulence rises, while the +> dynamic encoding's is zero by construction — so the crossover is a +> statement about `ρ` on the diagonal, not about RMSE anywhere. --- @@ -100,11 +115,126 @@ one exists. --- -## §2 THE TWO AXES (the operator's "good geometry vs badly calibrated") - -The two axes are **orthogonal by construction** and are varied -independently, so a result can be attributed to one or the other rather -than to their blend. +## §2 THE INSTRUMENT — CROSS-SWAP (Kreuztausch), not a horse race + +> **Correction of this plan's first draft (operator, 2026-08-12).** v1 §2 +> was built as a *comparison of formulas* — four calibration arms scored +> against each other on RMSE. That reads the premise backwards. The premise +> is: **assume the model DOES capture the phenomenon, but is NOT +> calibrated.** Under that premise miscalibration is the *condition of +> measurement*, not one arm in a race — and RMSE under a deliberately wrong +> calibration is bad **by definition**, so scoring it answers nothing. The +> informative quantity is how much of the captured **structure** survives +> the swap. The regime ladder, the budget discipline and the control gate +> from v1 are unchanged; only the question is inverted. + +### §2.1 What is held constant, what is swapped + +The operator's methodological frame: *you hold variables constant to test +the others; constancy is relative, so the design deliberately manufactures +strong correlation differences — on the assumption that those differences +are fit to evaluate the hypothesis.* Both halves are load-bearing here, and +the assumption in the second half is given its own falsifier (C1c). + +Each cell holds everything fixed except one thing: + +| held constant | varied | +|---|---| +| the box (regime), the timestep, the geometry arm, the sample count | **only the calibration's donor regime** | + +Notation: `M[D][T]` = read regime **T**'s field through a codebook derived +from regime **D**. The diagonal `D = T` is own-calibration; every +off-diagonal cell is a swap. One full matrix per geometry arm, so geometry +never blends into the calibration answer. + +**Constancy is operationalized, never assumed** (C1b): the within-box +spread of the discriminator must be small relative to the between-box +spread. A box that fails that is not a control condition — it is just +another data point wearing the label. + +### §2.2 The matrix and its metrics + +The 4×4 `M[D][T]` over `{R1 CALM, R2 OCEAN, R3 ACTIVE, R4 STORM}`. Per +cell, measured and stored raw: + +| quantity | what it answers | +|---|---| +| **`ρ` Spearman(reconstructed, true)** | **PRIMARY** — how much ordering survived | +| `occupancy` = fraction of the 256 levels actually used | the mechanism: a foreign codebook collapses the field onto few levels | +| `saturation` = fraction clipped at level 0 or 255 | the other half of the mechanism: the field runs off the donor's range | +| RMSE / bias in Pa | **secondary** — evidence the swap genuinely hurts in absolute terms, never the verdict | + +**The derived quantity the whole plan turns on:** + +``` +transfer loss L[D][T] = ρ[T][T] − ρ[D][T] +``` + +— how much structure regime `T` loses when read through `D`'s calibration. +`L` is the cross-swap statement of "badly calibrated": *low* `L` means the +substrate carried the structure even though the numbers were wrong. + +**`L < 0` is possible and is pre-registered as informative, not as a bug.** +Nothing forces the diagonal to be the best cell: a box whose own min/max is +set by a single outlier spreads its 256 levels badly, while a donor with a +wider range may spread them better — so a *foreign* codebook can beat the +own one. If that occurs it is reported as measured, with `occupancy` and +`saturation` beside it (which is where the mechanism would show), and it +carries a specific consequence: **`ρ[T][T]` is then not a valid reference +point for that box's `L̄[T]`, so C3's trend must be re-read against the +best-available cell and the substitution stated.** Naming this in advance is +the point — an unexpected sign discovered mid-analysis is exactly the kind +of result that gets explained away. + +### §2.3 The dynamic row is degenerate — and that IS the point + +A rank-normalised or Fisher-z encoding re-derived **inside the window** has +no donor at all. Its matrix row is therefore identical across `D` and +`L ≡ 0` **by construction**. That is not a defect to hide; it is precisely +the property under test — a dynamically-mapped substrate cannot be +mis-calibrated, because it carries no absolute anchor to get wrong. + +Two consequences, both mandatory: + +1. **The degeneracy must be VERIFIED, not asserted.** A nonzero + off-diagonal on a dynamic arm means a donor parameter leaked into the + window — a bug, and the run is void until it is found. +2. **…and the pipeline must be proven able to produce a NON-degenerate + row** through the identical code path (the absolute arms must show real + off-diagonal degradation). A row that is constant because the harness + cannot vary anything is the `E-A-CONTROL-THAT-CANNOT-LOSE` defect in its + can-it-DIFFER form, measured in W5. + +So the real comparison is not "which formula wins" but: + +> In which regime does `ρ(dynamic, own window)` exceed +> **`ρ(absolute, OWN donor — the diagonal)`** — and does the margin grow +> with turbulence? + +**Against the DIAGONAL, deliberately: it is the hardest available +opponent.** The comparison against a *foreign* donor is the weak form. It +is reported separately (C4) and labelled as evidence of **wiring, not of +merit**, because C2(i) makes the dynamic arm win it almost by construction +— an encoding with no anchor cannot be out-transferred by one whose anchor +is deliberately wrong. Scoring the hypothesis on the swapped cell would +re-introduce exactly the tautology this whole section removes. + +That is the operator's hypothesis stated as a cross-swap, and it is the +only form in which a mis-calibrated arm can be scored fairly. + +> **Caught by review, and it is the same defect one level up (2026-08-12, +> CodeRabbit on #944).** The first version of this passage named the +> foreign-donor comparison as "the operator's hypothesis" — contradicting +> C4, `STATUS_BOARD`, and the `INTEGRATION_PLANS` entry, all of which +> correctly score against the diagonal. The passage was written BEFORE C4 +> was sharpened, and the sharpening was not propagated back. Left standing, +> the plan would have carried two incompatible verdict criteria with the +> weaker one holding the headline — the swapped cell establishing merit, +> which is the precise thing §2 exists to prevent. Recorded rather than +> silently patched: it is the fifth instance of a claim that is consistent +> with its own operands but inconsistent with a sibling claim (cf. #930's +> fused relation, #927's decimal claim, #928's audit figures, #941's +> dropped qualifier), and it is a REVIEWER who caught it, not the author. ### Axis A — GEOMETRY (where the samples sit) @@ -121,20 +251,26 @@ budget silently advantages the arm with more samples — measured there as ### Axis B — CALIBRATION (how the 256 palette levels are placed) -| arm | construction | absolute anchor? | dynamic? | -|---|---|---|---| -| `CAL-ABS-OWN` | 256 uniform levels over THIS box's own min/max | yes | no | -| `CAL-ABS-FOREIGN` | 256 uniform levels over a DIFFERENT regime's min/max | yes, **wrong one** | no | -| `CAL-RANK-DYN` | rank-normalised within the window, re-derived per box | **no** | **yes** | -| `CAL-FISHERZ-DYN` | Fisher-z on within-window ranks (the arc's analytic codebook) | **no** | **yes** | - -`CAL-ABS-FOREIGN` is the literal reading of "badly calibrated"; -`CAL-RANK-DYN` / `CAL-FISHERZ-DYN` are "badly calibrated in absolute terms -BUT dynamically mapping" — the operator's actual candidate. +Three encodings, each run through the **full 4×4 donor matrix**: -**Metric:** reconstruction RMSE in **Pa** (the physical unit, per the -`E-R²-IS-NEAR-BLIND` lesson — never R² alone), plus Spearman ρ of the -reconstructed vs true field, plus mean **bias** in Pa. +| arm | construction | absolute anchor? | matrix shape | +|---|---|---|---| +| `CAL-ABS` | 256 uniform levels over donor `D`'s min/max | yes | **full** — diagonal = own, off-diagonal = swap | +| `CAL-RANK` | rank-normalised within the target window | **no** | **degenerate** — `L ≡ 0` by construction | +| `CAL-FISHERZ` | Fisher-z on within-window ranks (the arc's analytic codebook) | **no** | **degenerate** — same | + +v1 listed `CAL-ABS-OWN` and `CAL-ABS-FOREIGN` as two separate arms. They +are not two arms — they are the **diagonal and the off-diagonal of one +arm's matrix**, and splitting them was the tell that the design was still a +race. `CAL-ABS` with `D = T` is own-calibration; `CAL-ABS` with `D ≠ T` is +the literal "badly calibrated"; the two dynamic arms are "badly calibrated +in absolute terms BUT dynamically mapping" — the operator's actual +candidate, and the ones whose rows are flat by construction. + +**Metrics:** primary `ρ` and the derived transfer loss `L` (§2.2); +secondary RMSE and mean **bias** in **Pa** (the physical unit, per the +`E-R²-IS-NEAR-BLIND` lesson — never R² alone), reported for every cell but +never used as the verdict. --- @@ -155,29 +291,67 @@ reconstructed vs true field, plus mean **bias** in Pa. R1 < R2 < R3 < R4 must hold on **≥3 independent timesteps**, not just the preflight's one. If the ladder inverts on any timestep, the regime axis is not stable and every downstream cell is reported with that caveat. -- **C2 THE CROSSOVER — the operator's hypothesis, two-sided:** - `Δ = RMSE(CAL-RANK-DYN) − RMSE(CAL-ABS-OWN)` must be **> 0 in R1/R2 - (calm: dynamic loses) AND < 0 in R4 (storm: dynamic wins)** — a genuine - sign flip. **Both failure directions are reportable results, not - disappointments:** no flip = the hypothesis is refuted on this data and - says so; flip in the *opposite* direction = dynamic encoding is a - calm-regime tool, which would be a real and surprising finding. -- **C3 THE MISCALIBRATION PENALTY SHRINKS WITH TURBULENCE:** the ratio - `RMSE(CAL-ABS-FOREIGN) / RMSE(CAL-ABS-OWN)` must be **strictly smaller in - R4 than in R1** — the direct statement of "storms are more forgiving of - bad calibration." Reported with the ratio at every tier, so a monotone - trend (or its absence) is visible rather than inferred from two endpoints. -- **C4 GEOMETRY FLOOR BITES HERE (or it does not):** `GEO-GOLDEN-LO` must +- **C1b CONSTANCY IS RELATIVE — so it is MEASURED, not claimed.** For the + discriminator (`|∇p|`), the **within-box** spread must be small relative + to the **between-box** spread: report + `separation = (between-box range) / (mean within-box σ)` and require + **≥ 3**. Below that, "holding the regime constant" is a label rather than + a condition, and every downstream cell inherits that caveat explicitly. + *(This is the operationalization of the operator's point that constancy + is relative — a box is a control condition only insofar as its internal + variation is dominated by the spread the design manufactured.)* +- **C1c THE REGIMES MUST DIFFER IN CORRELATION STRUCTURE, not merely in + `|∇p|` — the suitability ASSUMPTION, made falsifiable.** The ladder is + built on a pressure-gradient discriminator, but the hypothesis is about + **structure**. So before any swap runs: measure each box's own + **autocorrelation decay length** and **rank-distribution shape** (Gini / + tail ratio of `|∇p|`). If R1 and R4 are indistinguishable on those, they + are ONE regime for this question no matter how far apart their gradients + are, and the ladder measures four copies of the same condition. + **Pre-registered honest reading:** a null here VOIDS the cross-swap + interpretation rather than weakening it — it would mean the manufactured + spread was manufactured on the wrong axis. Reported first, not last. +- **C2 THE DEGENERATE ROW — verified, and proven capable of being + non-degenerate.** Two halves, both required: + (i) every dynamic arm (`CAL-RANK`, `CAL-FISHERZ`) must show `L[D][T] = 0` + for all `D` **exactly**; any nonzero off-diagonal means a donor parameter + leaked into the window and the run is VOID until it is found; + (ii) through the **identical code path**, `CAL-ABS` must show a + **non-zero** off-diagonal in at least one regime. Half (i) alone is the + can-it-DIFFER defect measured in W5 — a row that is flat because the + harness cannot vary anything proves nothing about the encoding. +- **C3 TRANSFER LOSS SHRINKS WITH TURBULENCE:** for `CAL-ABS`, the mean + off-diagonal transfer loss `L̄[T] = mean_{D ≠ T} L[D][T]` must be + **strictly smaller in R4 than in R1** — the cross-swap statement of + "storms are more forgiving of bad calibration." Reported at every tier so + a monotone trend (or its absence) is visible rather than inferred from + two endpoints, and reported **alongside `occupancy` and `saturation`**, so + a shrinking loss can be attributed to a mechanism rather than asserted. +- **C4 THE CROSSOVER — the operator's hypothesis, two-sided:** + `Δ[T] = ρ(CAL-RANK, T) − ρ(CAL-ABS, D=T, T)` must be **< 0 in R1/R2 + (calm: own-calibration absolute wins) AND > 0 in R4 (storm: dynamic + wins)** — a genuine sign flip against the *diagonal*, which is the + hardest available opponent. **Both failure directions are reportable + results, not disappointments:** no flip = the strong hypothesis is + refuted on this data and says so; flip the *other* way = dynamic encoding + is a calm-regime tool, which would be real and surprising. + **The weak form is reported separately and never conflated with it:** + `ρ(CAL-RANK, T) > ρ(CAL-ABS, D ≠ T, T)` — dynamic beats a *mis-calibrated* + absolute. That one is nearly guaranteed by C2(i) and is therefore + evidence of wiring, not of merit. +- **C5 GEOMETRY FLOOR BITES HERE (or it does not):** `GEO-GOLDEN-LO` must be worse than `GEO-GOLDEN-HI` at equal budget. **Pre-registered honest reading:** W5's B4 already found the floor to be a *safety margin, not a mechanism* on a smoothing metric — so a NULL here is expected-plausible and must be reported plainly, not buried. What would be genuinely informative is the floor biting on a *sampling-fidelity* metric where it did not bite on a *smoothing* one. -- **C5 THE COMFORT MATRIX (descriptive, the deliverable):** the full - `regime × (geometry × calibration)` RMSE table, plus each cell normalized - by its regime's best arm — so "where does this formula feel at home" is - read directly off the matrix rather than argued. +- **C6 THE TRANSFER MATRIX (descriptive, the deliverable):** the full + `geometry × (donor × target)` table of `ρ`, `L`, `occupancy`, + `saturation`, RMSE and bias — every cell raw, plus the derived `L̄[T]` + column. "Where does this formula feel at home" is then **read off the + diagonal**, and "how badly does it travel" **off the off-diagonal** — + neither argued. --- @@ -187,10 +361,15 @@ Per the repeated finding that a first artifact ships summaries and omits the operands its headline rests on (W6's per-storm predictors; W5's family-B histogram; the chat-only 99.38 %), the JSON **must** carry: -- every cell's **raw** RMSE / bias / Spearman ρ, in Pa where dimensional -- the per-regime **codebook edges actually used** (so a miscalibration - claim is auditable without a re-fetch) -- the **measured** `|∇p|`, spd σ, elev σ and lsm per box per timestep +- every cell's **raw** `ρ` / `occupancy` / `saturation` / RMSE / bias, + keyed by `(geometry, donor D, target T)` — the full matrix, not the + diagonal plus a summary. `L[D][T]` is DERIVED in the report from stored + `ρ`, never stored alone (per the W6 lesson: store the operands, so a + headline can be re-derived without a re-fetch) +- the per-regime **codebook edges actually used**, for every donor — a + miscalibration claim is auditable only if the wrong codebook is on disk +- the **measured** `|∇p|`, spd σ, elev σ and lsm per box per timestep, + plus C1b's `separation` ratio and C1c's decay length + tail ratio - the **sample count actually drawn** per arm (the equal-budget proof, not the intent) - **units on every dimensional field name**, per the `c_bow`-is-km⁻¹ lesson @@ -217,3 +396,14 @@ already fetched in preflight. silent substitution. - **Not a substitute for CT-F17.** Nothing here touches the directional claim; it is a substrate-fidelity map, a different question entirely. +- **Not a claim that the four boxes are the same condition minus one + knob.** Real regimes differ in more than the discriminator. C1b bounds + how far the "held constant" label is earned, C1c bounds whether the + manufactured spread lies on the axis the hypothesis is about — and + whatever those two report travels with every downstream number rather + than being dropped once the matrix is filled. +- **Not a claim that transfer loss isolates calibration alone.** `L` is + measured with `occupancy` and `saturation` beside it precisely because a + shrinking `L` could also mean the target's field happens to sit inside + the donor's range by luck of that timestep. Three timesteps bound that; + they do not eliminate it.