diff --git a/.claude/board/LATEST_STATE.md b/.claude/board/LATEST_STATE.md index 73fdeaa2..c0247c8c 100644 --- a/.claude/board/LATEST_STATE.md +++ b/.claude/board/LATEST_STATE.md @@ -1,3 +1,16 @@ +## 2026-08-12 — lance-graph #926 (MERGED) — the storm spine, the 12-byte L4 carrier, and the propagation failure mode + +### Current Contract Inventory — no new types (probes + report + board; zero Rust, zero product code) + +- **The spine `[G]`:** a surface low = **center address + 14 logical fit values** (~12 ring means + a 2-value wn-1 dipole) = **90.9–94.3 %** of in-disk MSLP variance, replicated across three independent blind samples (1980–2021, 41+ storms, four seasons). Not 93–97 % — that figure was a **36-parameter per-ring fit**, not the 14-value model the storage claim describes. +- **It fits the REAL carrier `[H]`:** le-contract §3 **L4 is a PAIR** — `6 × (8:8)`, `palette256²`. The single byte is the *selector*; the pair is a centroid-tile cell. A **12-byte facet** recovers the f64 spine to **0.07 Pa RMSE** with a **+1.59 Pa bias**. **14 logical values ≠ 14 bytes** — keep model size and carrier budget apart. +- **Fisher-z is per-read, not universal `[H]`:** **8.3× tighter** than plain rank in the storm tail on the raw field; **4.7× worse** than uniform on ring means. **Wins a RANK/TAIL read, loses an INTERPOLATE/LEVEL read** — which is why an L4 ClassView **MAY** (not must) declare an analytic codebook. +- **The directional claim is `[S]`, measured-unsupported:** CT-F14 (the properly-powered pre-registered test) = **0.684, p=0.0835 → NO-VERDICT**. The pooled figure crosses p<0.05 but is **gated** — the pre-registration had no contingency for its largest component failing its own n≥20 floor. +- **⚠ Two methodological rules this PR banks, both worth carrying forward:** + - **R² is structurally near-blind to encoder bias.** `var(y−ŷ)` (used at 11 sites) drops the squared mean residual; in-disk variance ~1e5 Pa² means a **+92.76 Pa** offset moves R² in the 5th decimal. "Lossless" was inferred from the one statistic that could not detect the loss. **Report RMSE and mean bias in the physical unit beside every R².** + - **The dominant defect class is PROPAGATION, not judgment.** All five documentation defects were *a claim corrected in one home and left standing in another* — prose vs artifact, body vs heading, code vs JSON, report vs PR description. When correcting a claim, **grep for its twins first.** +- **Still open, operator calls:** CT-F16 (score the dipole against **steering-level** motion — the one measured, unwired moderator) · the moist budget (θe/precip are **proxies**, not a closed entropy budget) before CT-M1..M3 · C-register climatological calibration (single-timestep today) · `domino.rs` is a *fixed tridiagonal kernel*, so the moderator model is unbuilt, not unwired. + ## 2026-08-11 — lance-graph #924 (MERGED) — evaluation plan ACTIVE after a 0-of-11 spec audit ### Current Contract Inventory — no new types (plan + audit + corrections) diff --git a/.claude/board/PR_ARC_INVENTORY.md b/.claude/board/PR_ARC_INVENTORY.md index 81027cb1..632ec162 100644 --- a/.claude/board/PR_ARC_INVENTORY.md +++ b/.claude/board/PR_ARC_INVENTORY.md @@ -1,3 +1,20 @@ +## 2026-08-12 — lance-graph #926 (MERGED) — the storm spine (90.9–94.3 %), the 12-byte L4 carrier, and five corrections that were each a claim fixed in ONE of its homes + +- **Added.** `probes/weather-p1/COMET_TAIL_REPORT.md` (the canonical document: falsification ledger §4–§5.11, **§6.1** the L4 carrier, **§6.2** the three registers, §9 the spine/moderators reframe) + **10 probe scripts with committed JSON results**. EPIPHANIES: `E-THE-HEADLINE-NUMBER-MEASURED-A-MODEL-NOBODY-CLAIMED-1`, `E-SPINE-FOUND-MODERATORS-MISSING-1`, `E-THE-BYTE-WAS-ONLY-THE-SELECTOR-THE-PAIR-IS-THE-CARRIER-1`, `E-ZERO-FOR-ELEVEN-…` cross-refs. **9 919 insertions / 28 files: 47 % results-JSON, 36 % probe scripts, 17 % prose — zero Rust, zero library/product code, no test harness.** This PR changes no product surface; it is measurement apparatus and its output. +- **Locked — the spine `[G]`.** A surface low compresses to **a center address + 14 logical fit values** (~12 ring-profile means + a 2-value wn-1 dipole) **= 90.9–94.3 % of in-disk MSLP variance**, replicated across **three independent blind samples, 1980–2021, 41+ storms, four seasons, never shaken**. Physical basis is textbook and stated up front (§2): a translating vortex in a steering flow, geostrophy making the geometry *signed*, and a linear background gradient being **pure wavenumber-1 with `a₁ ∝ r`** — three independent falsifiable predictions, all three probed. +- **Locked — the carrier `[H]`.** le-contract §3 row **L4 is a PAIR**: `6 × (8:8)`, `palette256²`, "similarity = ONE table read". A single palette byte is only the **selector**; the pair is a cell in the centroid tile. The spine fits: a **12-byte facet** (10 ring bytes spread over the radius with 2 interpolated + a 2-byte dipole rail) recovers the f64 spine to **0.07 Pa RMSE (0.03 %)** with a **+1.59 Pa bias**. **14 logical values ≠ 14 bytes** — model size and carrier budget are distinct and were conflated in three places before this landed. +- **Locked — the Fisher-z demarcation `[H]`.** Fisher-z centroid axes are **8.3× TIGHTER** than plain rank in the storm tail on the raw field (24.74 vs 204.54 Pa) and **4.7× WORSE** than uniform on the ring means (18.07 vs 3.84 Pa). Not a contradiction — **Fisher-z wins a RANK/TAIL read and loses an INTERPOLATE/LEVEL read**, which is exactly why le-contract says a ClassView **MAY** declare an analytic codebook. Corrects an over-generalization made earlier in the same session that Fisher-z is *the* L4 axis. A post-hoc rescue (rank against the field rather than the encoded values) was measured and **also failed** — recorded rather than dropped. +- **The directional claim is NOT established `[S]`.** CT-F14 — the single properly-powered, pre-registered, displacement-filtered test — returned **13/19 = 0.684, p = 0.0835 → NO-VERDICT** (one short of its own n ≥ 20 floor, and it would have failed the 0.70 bar regardless). The pooled 3-sample figure crosses p < 0.05 but is **gated, not promoted**: the pre-registration had no contingency for its largest component failing its own floor, and that gap is named rather than exploited. Trajectory across the arc: dead (F3) → alive ±3–7° (F4) → general → reversed → pooled-but-unsupported (F14). +- **⊘ Five corrections, and they are ONE failure mode.** Every genuine defect this PR's review surfaced was **a claim corrected in one of its homes and left standing in another**: (1) the **93–97 %** headline — a 36-parameter per-ring fit, not the 14-value model claimed — corrected in the block and left in the prose beneath it, in EPIPHANIES ×2, and in the **PR title/body**; (2) *"a null does not produce a ladder"* corrected in EPIPHANIES, left in the report §9 **and** in the PR body; (3) *"the statistically correct reading"* surviving in a **heading one line above its own correction**; (4) `CT_F12 pass:true` at n=3 and CT-F14's `ESTABLISHED` verdict — **code fixed, JSON artifact never regenerated**; (5) the values/bytes conflation in three documents. In each case the claim was **already known to be wrong** — what failed was propagation, not judgment. +- **⊘ R² used `var()` instead of the uncentered MSE at 11 sites (8 files).** `var(y−ŷ)` discards the squared *mean* residual, so any BIASED reconstruction is flattered. **Zero effect** wherever a ring-mean profile is present (`mean(resid)` = 1e-12 by construction) — so every f64 headline is unchanged and no other JSON needed regenerating — but it concealed a **+92.76 Pa systematic bias** in carve A (0.9212 → 0.9129). **The 12-byte facet had been called "lossless" on the strength of an R² agreeing to four decimals: in-disk variance is ~1e5 Pa², so R² is structurally near-blind to exactly the defect that matters for an ENCODER.** The probes now emit RMSE and mean bias in Pa beside every R². +- **⊘ Nine vacuous falsifiers across the arc**, the last three found by review: L4x passed **comparing a codebook array against ITSELF** (a uniform codebook is fixed by its population's min/max, and storm 1's range contains storm 2's) — and what MADE it vacuous was switching to the codebook the *previous bar had just named best*; E6's decay test accepted a **flat** tail; R5 quantified over 256 bytes while sampling 5, understating the worst spread **9×**. +- **Deferred / not claimed.** `domino.rs` does **not** implement the moderator model — its `W` is a *fixed* tridiagonal smoothing kernel (`domino.rs:113`), no learned weights, gates, hidden or cell state; the tile-GEMM **shape and substrate** exist, the model does not. θe / TCWV / precipitation / vertical velocity are **proxies**, not a closed entropy budget — that budget must be written before the CT-M1..M3 diabatic gate is used as a moderator. CT-F16 (score the dipole against **steering-level** motion, the one moderator this chain measured) NOT run. C-register calibration is single-timestep, not climatological — an operator call. +- **Docs.** `COMET_TAIL_REPORT.md` §1/§4/§5.1–5.11/§6.1/§6.2/§8b/§9; EPIPHANIES ×3 new + 2 regraded in place; PR title/body corrected post-merge-review. + +**Correction (2026-08-12, same-day, caught by verifying this entry's own figures against the committed JSONs before landing it):** the Fisher-z ring-mean ratio is **4.7×**, not the *5×* first written here and in `COMET_TAIL_REPORT.md` §6.1 / `EPIPHANIES.md` — 18.07 / 3.84 = 4.71. A 6 % rounding in the FAVOURABLE direction, about to be frozen into an append-only record. Fixed in the two NEW board entries and in `COMET_TAIL_REPORT.md` §6.1 (not a board file). **`EPIPHANIES.md`'s merged entry is deliberately left reading *5×*** — it merged with #926, and the append-only rule freezes a merged entry except its Status/Confidence lines. **My first attempt edited that number in place, unmarked, which is history rewriting, not correction** — reverted the same commit. THIS entry is where the correction lives; a newest-first ledger is read from the top, so the frozen line is not a trap. That makes this PR mixed (hygiene + one correction) rather than hygiene-only, and the correction is recorded here rather than as a separate entry because it corrects #926's own content. **Nine of ten figures verified exactly; this was the tenth** — and it is the arc's signature defect one more time, committed by the person who spent the arc naming it. + +**Confidence (2026-08-12):** merged. The **structural** claim is `[G]` — wn-1 dominance and the R² lift replicated across three independent samples and survived every apparatus gate, and the f64 headlines are unaffected by the `var()`→MSE fix. The **carrier fit** is `[H]`: 2 storms, 1 timestep, 1 variable — structural fit to the 12-byte facet, saying nothing about forecast skill. The **directional predictor is `[S]` and measured-unsupported at this power**, not merely pending. External review (CodeRabbit) found 8 real defects across 23 threads; the two most valuable were the `var()`/MSE numerator and the CT-F14 artifact contradicting its own report. Docstring-coverage reported 69.23 % for three runs while measuring 100 % five ways locally; it self-corrected to 100 % — **declining to manufacture docstrings for an unreproducible metric was the right call.** + ## 2026-08-11 — lance-graph #924 (MERGED) — the audit that failed 11 of 11 test specs, and the plan it made ACTIVE - **Added.** `weather-substrate-evaluation-v1.md` **§8 audit record** + a full **§3 rewrite to v2 specs**; the plan header flips **DRAFT → ACTIVE**. Knowledge doc **§12.18**. EPIPHANIES `E-ZERO-FOR-ELEVEN-THE-AUTHOR-CANNOT-AUDIT-HIS-OWN-FALSIFIERS-1`. `AGENT_LOG` run entry. Board hygiene for #923. diff --git a/probes/weather-p1/COMET_TAIL_REPORT.md b/probes/weather-p1/COMET_TAIL_REPORT.md index d9feff01..a8a181d9 100644 --- a/probes/weather-p1/COMET_TAIL_REPORT.md +++ b/probes/weather-p1/COMET_TAIL_REPORT.md @@ -871,7 +871,7 @@ Three results worth more than the headline: is far starker than R² suggests: carve A's held outer rings cost **+92.76 Pa of bias**, carve D's **+1.59 Pa**. - **L3 FAILED, and so did my proposed rescue.** Fisher-z centroid axes are - **5× worse** than uniform on the ring means (18.07 vs 3.84 Pa). I + **4.7× worse** than uniform on the ring means (18.07 vs 3.84 Pa). I hypothesised the population was wrong — ranks taken against the 24 encoded values instead of the field — and measured that too (L3b): **19.00 Pa, no rescue.** Mechanism: ring means are a smooth *narrow-band* quantity sitting