Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
101 changes: 101 additions & 0 deletions .claude/board/EPIPHANIES.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,104 @@
## 2026-08-12 — E-A-HORSE-RACE-IS-NOT-A-CROSS-SWAP-1

**Status:** FINDING `[H]` — methodological, operator-ruled. Caught on
`substrate-comfort-zones-v1` BEFORE any bar ran, so nothing measured had to
be reinterpreted. Grade is `[H]` because the reasoning is airtight given the
premise but the premise itself ("the model captures the phenomenon") is an
assumption this plan does not test.

**The defect, in one line: when the premise is *the model captures the
phenomenon but is not calibrated*, an error metric measured under a
deliberately wrong calibration is bad BY DEFINITION, so scoring it answers
nothing.** The first draft of that plan compared four calibration arms on
reconstruction RMSE — a horse race. But `CAL-ABS-FOREIGN` is not a
competitor that might lose; it is the measurement CONDITION the whole
hypothesis is about. Racing it guarantees the answer and measures the
tautology.

**What the design must be instead.** Miscalibration is a *transfer*
question, so the instrument is a transfer matrix, not a ranking: `M[D][T]` =
read regime `T`'s field through a codebook derived from regime `D`. The
diagonal is own-calibration; every off-diagonal cell is a swap. The primary
metric becomes **structure preservation** (Spearman `ρ` of reconstructed vs
true), and the quantity the hypothesis actually concerns is the derived
**transfer loss** `L[D][T] = ρ[T][T] − ρ[D][T]` — how much ordering survives
being read through the wrong calibration. Absolute error is retained as
*evidence that the swap genuinely hurt*, never as the verdict.

**The sharp consequence that only appears in the matrix framing.** A
window-local encoding (rank-normalised, Fisher-z) has **no donor at all**,
so its matrix row is flat and `L ≡ 0` **by construction**. That is not an
artifact to hide — it IS the property under test: a substrate that carries
no absolute anchor cannot have that anchor be wrong. Which reframes the
whole comparison. The question stops being "which formula wins" and becomes
"in which regime does the anchor-free encoding's own-window `ρ` exceed the
absolute encoding's, measured against the DIAGONAL (its own calibration,
the hardest available opponent) rather than against the swapped cell (which
it beats almost by definition)."

**And the degeneracy must be VERIFIED both ways** — this is the
`E-A-CONTROL-THAT-CANNOT-LOSE-IS-NO-CONTROL-1` lesson in its can-it-DIFFER
form, which W5 already paid for once. A flat row is only evidence if the
identical code path is proven able to produce a NON-flat one (the absolute
arm must show real off-diagonal degradation). Otherwise "flat" means the
harness cannot vary anything, and the encoding was never tested.

**Second half of the ruling, on the regime ladder itself.** The operator's
framing: *in science you hold variables constant to test the others;
constancy is relative, so the design deliberately manufactures strong
correlation differences — on the assumption that those differences are fit
to evaluate the hypothesis.* Three things follow, and the plan had only the
first:

1. **The 9.3× `|∇p|` spread across the four regimes is the design's PURPOSE,
not a by-product of box selection.** At small spread every effect is in
the noise; at large spread it is legible or it does not exist.
2. **"Constant" is relative, so it is MEASURED** — within-box spread of the
discriminator must be small relative to the between-box spread
(`separation ≥ 3`). Below that, "held constant" is a label on a box, not
a condition, and the caveat must ride every downstream number.
3. **"Fit to evaluate the hypothesis" is an ASSUMPTION and therefore needs a
falsifier.** The ladder is built on a pressure-gradient discriminator but
the hypothesis is about correlation STRUCTURE. If two regimes are
indistinguishable in autocorrelation decay and rank-distribution shape,
they are ONE condition for this question however far apart their
gradients sit — and the ladder would be measuring four copies of the same
thing. A null there VOIDS the reading rather than weakening it: it would
mean the spread was manufactured on the wrong axis.

**The defect is STICKY — it survived the rewrite built to remove it, and a
REVIEWER caught it, not the author (CodeRabbit on #944, same day).** §2.3's
summary passage still named the *foreign-donor* comparison as "the
operator's hypothesis" while C4, `STATUS_BOARD` and the `INTEGRATION_PLANS`
entry all correctly scored against the diagonal. Mechanism: §2.3 was written
before C4 was sharpened, and the sharpening was not propagated backwards.
Left standing, the plan would have carried **two incompatible verdict
criteria with the weaker one holding the headline** — the swapped cell
establishing merit, i.e. the exact tautology the rewrite existed to delete.

That makes this the **fifth** instance in this arc of a claim that is
internally consistent with its own operands but inconsistent with a SIBLING
claim (#930's fused relation, #927's decimal claim, #928's audit figures,
#941's dropped qualifier, now this). The generalization is sharper than the
individual fixes: **when a design is revised, the summary sentences are the
last thing to be updated and the first thing a reader believes.** A revision
is not complete when the new machinery is right; it is complete when every
sentence that *names the conclusion* has been re-derived from the new
machinery. Self-review kept missing it because the author re-reads the part
that changed, not the part that merely still sounds right.

**The transferable rule.** Before scoring arms against each other, ask which
of them is a *competitor* and which is a *condition*. A condition put on the
starting line will always lose, and its loss carries no information. Score
conditions by what SURVIVES them, not by how much they cost.

**Cross-ref:** `.claude/plans/substrate-comfort-zones-v1.md` §2 (the rebuilt
instrument, with v1's framing preserved in place as a correction note rather
than deleted) and §3 C1b/C1c/C2/C3/C4; `STATUS_BOARD.md` D-CZ-* (rows re-cut
while still Queued, with the old→new mapping recorded);
`E-A-CONTROL-THAT-CANNOT-LOSE-IS-NO-CONTROL-1`;
`E-R²-IS-NEAR-BLIND` (why the physical unit is retained alongside `ρ`).

## 2026-08-12 — E-THE-DISPLACEMENT-FILTER-ATE-THE-STRANDED-STRATUM-1

**Status:** FINDING `[G]` — W6 RUN (`comet_tail_w6.py`/`.json`), audited
Expand Down
50 changes: 50 additions & 0 deletions .claude/board/INTEGRATION_PLANS.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,53 @@
## 2026-08-12 — substrate-comfort-zones-v1 REVISED (§2 rebuilt: horse race → cross-swap)

Same file, `.claude/plans/substrate-comfort-zones-v1.md`, still **ACTIVE**,
still exploratory tier. Not a new plan and not a v2 — a pre-registration
sharpened **before any bar ran** (only D-CZ-0, the regime preflight, was
DONE; everything else was Queued), which is precisely when a design should
be allowed to change.

**Why.** Operator ruling, two messages: the method is *cross-swap and
hypothesis testing, under the premise that the model captures the phenomenon
but is not calibrated*; and *in science you hold variables constant to test
the others — constancy is relative, so the design deliberately manufactures
strong correlation differences, on the assumption those differences are fit
to evaluate the hypothesis*. Against that, v1 §2 was a **horse race**: four
calibration arms scored on reconstruction RMSE. Under the stated premise
that is vacuous — a deliberately wrong calibration produces bad error BY
DEFINITION, so `CAL-ABS-FOREIGN` is the measurement CONDITION, not a
competitor that might lose.

**What changed.**
- §2 is now a **4×4 donor × target transfer matrix** per geometry arm.
Primary metric **Spearman `ρ`** (structure preservation); the quantity the
hypothesis concerns is the derived **transfer loss**
`L[D][T] = ρ[T][T] − ρ[D][T]`. RMSE/bias in Pa are retained per cell as
evidence the swap genuinely hurt — never as the verdict.
- `CAL-ABS-OWN` / `CAL-ABS-FOREIGN` stop being two arms: they are the
diagonal and the off-diagonal of ONE arm's matrix. Splitting them was the
tell that the design was still a race.
- The dynamic arms' rows are **flat by construction** (`L ≡ 0`, no donor
exists) — that is the property under test, and C2 now requires BOTH that
it hold exactly AND that the identical code path be proven able to produce
a non-flat row (the can-it-DIFFER gate W5 already paid for).
- Two NEW bars for the second operator message: **C1b** measures constancy
(`separation = between-box range / mean within-box σ ≥ 3`) instead of
claiming it; **C1c** gives the suitability ASSUMPTION a falsifier — the
regimes must differ in autocorrelation decay and rank-distribution shape,
not merely in `|∇p|`, or the ladder is four copies of one condition.
- The crossover bar now scores against the **diagonal** (own-calibration,
the hardest opponent) on `ρ`; the weak form against the swapped cell is
reported separately and labelled as evidence of wiring, not of merit.

**What did NOT change:** the §1 regime ladder and its three preflight
corrections, the equal-budget discipline, the geometry axis, the C0 control
gate, the output contract's store-the-operands rule, and §5's
non-claims. The scaffold was sound; the question was inverted.

**Preserved, not deleted:** v1's framing survives in place as §2's
correction note, so a future session can see what the design was and why it
moved. Finding recorded as `E-A-HORSE-RACE-IS-NOT-A-CROSS-SWAP-1`.

## 2026-08-12 — substrate-comfort-zones-v1 (PLAN; where does each substrate formula feel at home?)

Plan: `.claude/plans/substrate-comfort-zones-v1.md`. Status **ACTIVE**,
Expand Down
28 changes: 22 additions & 6 deletions .claude/board/STATUS_BOARD.md
Original file line number Diff line number Diff line change
@@ -1,18 +1,34 @@
## substrate-comfort-zones-v1 — the comfort-zone map (PRE-REGISTERED 2026-08-12)

Plan: `.claude/plans/substrate-comfort-zones-v1.md`. Regime × geometry ×
calibration → where does each substrate formula feel at home. §1 preflight
Plan: `.claude/plans/substrate-comfort-zones-v1.md`. A **cross-swap**
(donor × target) transfer matrix per geometry arm → where does each
substrate formula feel at home, and how badly does it travel. §1 preflight
already run and it corrected two regime definitions before any bar existed.

| D-id | Deliverable | Status | Feeds |
|---|---|---|---|
| D-CZ-0 | §1 regime preflight (`\|∇p\|` ladder, elevation-confound screen, speed-is-not-the-discriminator finding) | **DONE** — ladder R1 Amazon 10.2 → R2 ocean 14.9 → R3 W Siberia 43.8 → R4 storm 95.6 (9.3× range); 4 land candidates excluded on elev σ > 150 m | the regime axis all other rows score on |
| D-CZ-1 | C0 controls (shuffled codebook + degenerate geometry), losability-smoke-tested BEFORE the full run | Queued | gates every cell — a control that can't lose voids its cell |
| D-CZ-2 | C1 regime-ladder stability across ≥3 timesteps | Queued | anti-cherry-pick on the whole regime axis |
| D-CZ-3 | **C2 the crossover** — `RANK-DYN − ABS-OWN` sign flip calm↔storm | Queued | **the operator's hypothesis, two-sided** |
| D-CZ-4 | C3 miscalibration penalty vs turbulence (`ABS-FOREIGN`/`ABS-OWN` ratio per tier) | Queued | "storms forgive bad calibration" |
| D-CZ-5 | C4 geometry floor on a SAMPLING-fidelity metric (NULL expected-plausible per W5 B4) | Queued | does the index floor bite where smoothing didn't |
| D-CZ-6 | C5 the comfort matrix (the deliverable) | Queued | read-off answer to "where is each formula at home" |
| D-CZ-2b | **C1b constancy is measured** — `separation = (between-box range)/(mean within-box σ)` ≥ 3 | Queued | earns the phrase "held constant"; otherwise a caveat rides every cell |
| D-CZ-2c | **C1c the suitability ASSUMPTION** — regimes must differ in autocorrelation decay + rank-distribution shape, not only in `\|∇p\|` | Queued | a null here VOIDS the cross-swap reading (spread manufactured on the wrong axis) |
| D-CZ-3 | **C2 degenerate-row verification, both halves** — dynamic arms `L ≡ 0` exactly, AND `CAL-ABS` proven non-degenerate through the same code path | Queued | the can-it-DIFFER gate; without half (ii) a flat row proves nothing |
| D-CZ-4 | C3 **transfer loss** vs turbulence — `L̄[T] = mean_{D≠T} L[D][T]` strictly smaller in R4 than R1, reported with `occupancy`/`saturation` | Queued | "storms forgive bad calibration", with a mechanism attached |
| D-CZ-5 | **C4 the crossover** — `ρ(CAL-RANK) − ρ(CAL-ABS, D=T)` sign flip calm↔storm (vs the DIAGONAL, the hardest opponent) | Queued | **the operator's hypothesis, two-sided**; weak form vs `D≠T` reported separately |
| D-CZ-6 | C5 geometry floor on a SAMPLING-fidelity metric (NULL expected-plausible per W5 B4) | Queued | does the index floor bite where smoothing didn't |
| D-CZ-7 | C6 the transfer matrix (the deliverable) | Queued | comfort read off the diagonal, travel-cost off the off-diagonal |

> **2026-08-12 — rows re-cut, not restated.** The operator ruled that the
> plan's §2 was built as a horse race where a cross-swap diagnostic belongs
> (premise: the model *captures* the phenomenon but is *not calibrated*, so
> miscalibration is the condition of measurement, not an arm). §2 was
> rebuilt around the 4×4 donor×target transfer matrix and the bars were
> renumbered; two new bars (C1b, C1c) exist because "held constant" and
> "the manufactured spread is on the right axis" are now measured rather
> than assumed. **Every row above was still Queued when this happened — no
> measured result was reinterpreted.** Old→new: C2→C4 (crossover, now
> against the diagonal on `ρ` rather than against `ABS-OWN` on RMSE),
> C3→C3 (now transfer loss, not an RMSE ratio), C4→C5, C5→C6.

## golden-vs-tempered-stride-v1 — head-vs-gut queue — RUN 2026-08-12

Expand Down
Loading