From 4b80db503ea94d9bbf1e8e81ebd43189475f686d Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 12 Aug 2026 10:49:50 +0000 Subject: [PATCH 1/4] board: record merged PR #929 (arc entry + LATEST_STATE) Post-merge board hygiene per the Mandatory Board-Hygiene Rule. #929 was a mixed PR (2 probes + report sections + epiphany), so its entry records the non-hygiene half: * CT-F16: the steering rescue failed both pre-registered bars; the level sweep is monotone toward the SURFACE (best 850 hPa). The ladder survives as a measurement; its operational reading is dead. * F16c's rotated control scored 0.684 = CT-F14's headline -- the sign test could not distinguish the real referent from a wrong one at n=19. Rule banked: a control's score is the floor any rate-headline must clear. * The circular resultant resolves the same 19 rows at p=0.0050, estimates the offset (-30.2 deg) instead of penalizing it, and separates the rotated control by 100.3 deg. Post-hoc, instrument demo, not a promotion -- but sign-consistency numbers are retired as verdict-grade instruments across the arc. All 13 figures in the entry verified against the committed JSONs BEFORE landing (13/13 exact -- the #927 self-verification, now standard practice). Append-only audit: zero removed lines across both board files (pure prepend). THIS PR IS PURE HYGIENE -- no type, plan, deliverable, epiphany or code. Per the termination clause it generates no further obligations; the chain stops here. Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi --- .claude/board/LATEST_STATE.md | 10 ++++++++++ .claude/board/PR_ARC_INVENTORY.md | 12 ++++++++++++ 2 files changed, 22 insertions(+) diff --git a/.claude/board/LATEST_STATE.md b/.claude/board/LATEST_STATE.md index 13a42abe3..c72b696b8 100644 --- a/.claude/board/LATEST_STATE.md +++ b/.claude/board/LATEST_STATE.md @@ -1,3 +1,13 @@ +## 2026-08-12 — lance-graph #929 (MERGED) — steering rescue dead; sign test retired as primary instrument + +### Current Contract Inventory — no new types (2 probes + report §5.12/§5.13 + 1 epiphany) + +- **CT-F16 `[G]`:** the steering-level moderator FAILED both pre-registered bars on CT-F14's own 19 storms (0.579 vs 0.70; residual +28.4 % wider; level sweep monotone toward the surface, best 850 hPa). §9.2's "single most promising fix" is superseded in place. The height ladder survives as a measurement; its operational reading died. +- **The instrument calibration `[G]`:** a 90°-rotated control scored **0.684 = CT-F14's headline**. At n=19 the sign test separates from a wrong answer only at 14/19. **Rule: attach a deliberately-wrong referent to any rate-headline; its score is the floor.** +- **The replacement instrument:** circular resultant (R̄, μ, Rayleigh) resolves the same rows at **p=0.0050** (sign test: 0.0835), estimates the offset (**−30.2°±36.5°**) instead of penalizing it, and makes the rotated control visible (μ shifted 100.3° at identical R̄). Post-hoc — instrument demo, NOT a promotion. **Sign-consistency numbers are no longer verdict-grade anywhere in this arc.** +- **Faltung frame:** resultant = first circular Fourier coefficient; CT-W6 = deconvolution (components ⊛ apparatus noise); `Z_256` circular Faltung is `DistanceLut::circular()`-native. +- **Queued, operator go-ahead pending:** **CT-W6** (dipole = neighbor far-field + bow-wave vector sum, global fit, 38 obs — the mechanistic candidate for the −30° offset and the 0.684 plateau) · CT-W2s/W5/W7 (sunflower collision facets, spiral-ADI via Fibonacci-stride pairs, Gegendruck) · CT-F17 (fresh-sample verdict) · C-register climatological calibration. + ## 2026-08-12 — lance-graph #927 (MERGED) — board hygiene for #926 + the living-vs-ledger correction rule ### Current Contract Inventory — no new types (board entries + one report correction) diff --git a/.claude/board/PR_ARC_INVENTORY.md b/.claude/board/PR_ARC_INVENTORY.md index 9b33b5b60..5a07e1e3e 100644 --- a/.claude/board/PR_ARC_INVENTORY.md +++ b/.claude/board/PR_ARC_INVENTORY.md @@ -1,3 +1,15 @@ +## 2026-08-12 — lance-graph #929 (MERGED) — CT-F16 kills the steering rescue; the control scores the headline; the resultant resolves what the sign test collapsed + +- **Added.** `comet_tail_f16.py`/`.json` (bars committed BEFORE the run, `05f09005`), `comet_tail_resultant_instrument.py`/`.json` (deterministic bootstrap, seed pinned), report **§5.12 + §5.13**, §9.2's steering row superseded **in place**, EPIPHANIES `E-THE-CONTROL-SCORED-THE-HEADLINE-1`. All 13 headline figures in THIS entry verified against the committed JSONs before landing (13/13 exact — the #927 lesson, now standard). +- **Locked — the leading moderator is DEAD `[G]`.** §9.2 had named the steering level *"the single most promising fix"*. Paired re-scoring of CT-F14's own 19 storms with ONLY the motion reference swapped (surface 6h displacement → 500/600/700 hPa disk-mean flow; stored centres, byte-identical `err_deg`): **F16a 0.579** vs a 0.70 bar (worse than surface's 0.684), **F16b residual 28.4 % WIDER** (68.29° → 87.71°) where ≥10 % tightening was predicted, **F16d level sweep monotone toward the SURFACE** — best 850 hPa, outside the ladder's 400–650 hPa band, no mid-tropospheric optimum. The height ladder itself (field-per-level, different quantity) is NOT falsified; its operational reading is. +- **Locked — the control calibrated the instrument `[G]`.** F16c's deliberately **90°-rotated** reference scored **13/19 = 0.684, p=0.0835 — numerically identical to CT-F14's headline**. At n=19 the sign-test ladder is 11→0.324, 13→0.0835, 14→0.0318: **CT-F14 was never one storm short of significance; it was one storm short of distinguishability from an answer built to be wrong.** Rule banked: *an anti-vacuity control measures the RESOLVING POWER of the instrument, not only the test it guards* — attach a deliberately-wrong referent to any rate-headline and read the control's score as the floor. +- **Locked — the 0.684 plateau was the STATISTIC, not the data `[G]`.** Operator: *"die irrationale Aufsummierung hilft, dass der Dipol nicht auf 0.68 kollabiert."* Measured on the SAME 19 rows, no fetch: circular resultant **R̄=0.516, μ=−30.2°±36.5°, Rayleigh p=0.0050** where the sign test gave NO-VERDICT (p=0.0835). The rotated control — conflated at 0.684=0.684 by the sign test — shows identical R̄ with **μ shifted 100.3°**: the wrong referent is VISIBLE. Permuted control collapses to R̄=0.142, under the uniform floor 0.203. Mechanism: binarization discards magnitude AND mean direction; the vector sum keeps both, so the arc's systematic offset is now ESTIMATED (−30.2°) instead of penalized. A one-sided sign test on a distribution not centred at zero reports which SIDE the bias falls on — this arc used it as the primary instrument since §4; every prior sign number (2/2, 6/10, 8/10, 13/19) is bounded by §5.13 in BOTH directions. +- **The Faltung frame (operator, same exchange).** The resultant IS the first circular Fourier coefficient (Faltung with `e^{iθ}`); the queued CT-W6 fit is a DEconvolution (component mix ⊛ apparatus noise, ±3–7° from CT-F4); on `Z_256` circular Faltung is substrate-native (`DistanceLut::circular()`'s domain). +- **⊘ Scope, held.** §5.13's p=0.0050 is **post-hoc on the same sample — an instrument demonstration, NOT a promotion** of the directional claim; the report says so in three places. Sign-test conflation second finding: weak flow (<10 m/s, n=6) sign 0.833 / median |err| 103°; strong (≥10, n=13) sign 0.462 / |err| 55°; corr(speed,|err|)=−0.407 — magnitude improves with steering as physics expects while sign moves opposite. +- **Deferred / queued.** **CT-W6** — two-component vector deconvolution of the dipole (neighbor far-field + bow-wave `½ρv_rel²`, GLOBAL coefficients over 38 observations, circular metrics, F16c-standard controls) — the candidate mechanism for the −30° offset AND the plateau (a vector sum cannot be scored against one bearing). CT-W1 folds in as the `c_geo` component. CT-W2s/W5/W7 (sunflower collision nodes with the 2-rail-address + 8-DOF facet carve, spiral-ADI aligned/orthogonal Fibonacci strides, Gegendruck), **CT-F17** (fresh-sample verdict). None run. + +**Confidence (2026-08-12):** merged. F16's failure verdicts are `[G]` (pre-registered, paired). The resultant numbers are `[G]` as measurements, `[S]` as evidence for the directional claim (post-hoc). The steering ladder survives as an n=2 observation. Structural claim (§9.1) untouched throughout. + ## 2026-08-12 — lance-graph #927 (MERGED) — #926's board hygiene, one rounding caught by self-verification, and the rule I broke while citing it - **Added.** `PR_ARC_INVENTORY` + `LATEST_STATE` entries for the merged #926. One correction to `COMET_TAIL_REPORT.md` §6.1. **This entry exists because #927 was MIXED, not hygiene-only** — it landed hygiene *and* a correction, and the rule's termination clause exempts only the pure case. #927's own PR description called itself "hygiene-only"; that was written before the correction landed and was wrong by the time it merged. From 25ab52b6340d0c2411822b9c1e9bb80adb1d9ac8 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 12 Aug 2026 10:57:16 +0000 Subject: [PATCH 2/4] board+report: a FALSE relation between two true numbers -- codex catch, fixed in all three homes Codex P2 on #930: "identical R-bar with mu shifted 100.3 deg" fused two different comparisons. Identical R-bar belongs to steering<->rotated (0.343 = 0.343, mu shifted exactly 90 -- rotation preserves concentration by construction); the 100.3 deg separation belongs to surface<->rotated, where R-bar is NOT identical (0.516 vs 0.343). The composite was false in all three homes: LATEST_STATE, the #929 arc entry, report SS5.13. Stated correctly the finding is STRONGER -- the control separates from surface in BOTH channels. Unmerged board entries composed in place (append-only rule's unmerged-PR allowance); the merged report gets a dated correction note. Lesson banked: the 13/13 figure verification checked every NUMBER and still missed this -- a relation between two individually correct numbers can be false. Verify comparative claims AS claims: every "identical/same/larger" must name both operands, and the check must evaluate the relation. Also codex P2 #2: the LATEST_STATE shipped-PR table had stalled at #780. Added #926-#929 rows plus an explicit gap-note row for #781-#925 (carried by PR_ARC_INVENTORY) -- honest gap, not silent reconstruction of ~150 rows. Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi --- .claude/board/LATEST_STATE.md | 7 ++++++- .claude/board/PR_ARC_INVENTORY.md | 4 +++- probes/weather-p1/COMET_TAIL_REPORT.md | 18 ++++++++++++++---- 3 files changed, 23 insertions(+), 6 deletions(-) diff --git a/.claude/board/LATEST_STATE.md b/.claude/board/LATEST_STATE.md index c72b696b8..43c7055b3 100644 --- a/.claude/board/LATEST_STATE.md +++ b/.claude/board/LATEST_STATE.md @@ -4,7 +4,7 @@ - **CT-F16 `[G]`:** the steering-level moderator FAILED both pre-registered bars on CT-F14's own 19 storms (0.579 vs 0.70; residual +28.4 % wider; level sweep monotone toward the surface, best 850 hPa). §9.2's "single most promising fix" is superseded in place. The height ladder survives as a measurement; its operational reading died. - **The instrument calibration `[G]`:** a 90°-rotated control scored **0.684 = CT-F14's headline**. At n=19 the sign test separates from a wrong answer only at 14/19. **Rule: attach a deliberately-wrong referent to any rate-headline; its score is the floor.** -- **The replacement instrument:** circular resultant (R̄, μ, Rayleigh) resolves the same rows at **p=0.0050** (sign test: 0.0835), estimates the offset (**−30.2°±36.5°**) instead of penalizing it, and makes the rotated control visible (μ shifted 100.3° at identical R̄). Post-hoc — instrument demo, NOT a promotion. **Sign-consistency numbers are no longer verdict-grade anywhere in this arc.** +- **The replacement instrument:** circular resultant (R̄, μ, Rayleigh) resolves the same rows at **p=0.0050** (sign test: 0.0835), estimates the offset (**−30.2°±36.5°**) instead of penalizing it, and separates the rotated control in **BOTH channels** — vs surface (the pair the sign test conflated at 0.684 = 0.684): R̄ 0.343 vs 0.516 AND μ −130.5° vs −30.2° (100.3° apart); the steering↔rotated pair confirms rotation-invariance by construction (identical R̄ 0.343, μ shifted exactly 90°). Post-hoc — instrument demo, NOT a promotion. **Sign-consistency numbers are no longer verdict-grade anywhere in this arc.** - **Faltung frame:** resultant = first circular Fourier coefficient; CT-W6 = deconvolution (components ⊛ apparatus noise); `Z_256` circular Faltung is `DistanceLut::circular()`-native. - **Queued, operator go-ahead pending:** **CT-W6** (dipole = neighbor far-field + bow-wave vector sum, global fit, 38 obs — the mechanistic candidate for the −30° offset and the 0.684 plateau) · CT-W2s/W5/W7 (sunflower collision facets, spiral-ADI via Fibonacci-stride pairs, Gegendruck) · CT-F17 (fresh-sample verdict) · C-register climatological calibration. @@ -841,6 +841,11 @@ Membrane consumers can now pull BOTH halves of a render `classid` BBB-safely fro | PR | Merged | Title | What it added | |---|---|---|---| +| *gap note* | — | **#781–#925 are NOT in this table** — carried by the dated sections above + `PR_ARC_INVENTORY.md`. Recorded 2026-08-12 (codex P2 on #930) rather than silently reconstructed; the table had stalled at #780. | — | +| **#929** | 2026-08-12 | CT-F16 kills the steering rescue; the control scores the headline; circular resultant replaces the sign test | 2 probes (`comet_tail_f16`, `comet_tail_resultant_instrument`) + report §5.12/§5.13 + `E-THE-CONTROL-SCORED-THE-HEADLINE-1`. Steering FAILED both bars (0.579 vs 0.70; residual +28.4 % wider; sweep monotone to surface, best 850 hPa); rotated control scored 0.684 = CT-F14's headline; resultant resolves the same rows at p=0.0050, offset estimated −30.2°. | +| **#928** | 2026-08-12 | Board hygiene for #927 — chain terminator | Pure prepend (+13/−0, +10/−0). Living-documents vs append-only-ledgers correction discipline banked; zero-removed-lines diff = structural append-only audit. | +| **#927** | 2026-08-12 | Board hygiene for #926 + 4.7×-not-5× correction + append-only revert | Mixed. 9/10 figure self-verification caught the 6 % favourable rounding; an in-place edit of a MERGED EPIPHANIES entry reverted same-PR (the rule being cited was the rule broken). | +| **#926** | 2026-08-12 | weather-p1: the storm spine (90.9–94.3 %), the 12-byte L4 carrier, moderators identified but unwired | 9 919 insertions, zero Rust/product code. Spine `[G]` (3 blind samples, 41+ storms); 12-byte `6×(8:8)` facet at 0.07 Pa RMSE / +1.59 Pa bias; Fisher-z rank/tail-vs-level demarcation; CT-F14 NO-VERDICT held against a pooling rule that said ESTABLISHED; `var()`→MSE R² fix at 11 sites (hid +92.76 Pa bias). | | **#780** | 2026-07-21 | Recipes audited → 3 real tenants wired (A9 24-loci CausalWitnessFacet + SPO + qualia) → rung-ordered NaN-gated causal ladder | 4 commits, merge `8a00988`. `recipe_claim_audit` proves the 34 kernels sound-on-proxy (28/34; 2 INERT CAS/ETD + 4 CONSTANT ARE/ZCF/ICR/HKF) but 0/34 read a real organ → wiring gap, not composition. New contract modules: `causal_witness` (A9 = 24 signed-i4 loci, THIRD ClassView reading of the #729 12-byte register, no stored bytes), `recipe_substrate` (SubstrateView projects SPO+witness+qualia; **qualia additive + stakes-only**, causality = KAUSAL witness edge; missing tenant → NaN), `recipe_dispatch` (rung-ordered NaN-gated ladder keyed by NARS inference; ICR #31 unshadowed; the systematic-sweep mode COMPLEMENTING `select_tactic`'s 8/34 saccade — E-LADDER-UNSHADOWS-SELECTOR-1). +18 tests → 970 green, clippy clean; 3 probes all-gates-green. Theory anchors: PCRLLM 2511.08392 (proof-carrying single-step, Pei Wang), Causal-ML survey 2206.15475 (SCM ladder), MME-Reasoning 2505.21327 + NARS-Swift flagged; cc-thinking-skills mapped as rung-4 macro catalogue (E-THINKING-SKILLS-ARE-RUNG-4-MACROS-1). D-REC-WIRE-1. | | **#676** | 2026-07-10 | Post-#674 doc arc: E-NOBODY-WAITS-1 + VISION.md (graded AGI canon) + ancestry census + D-MTS-1 design inputs | Doc-only, 4 commits. E-NOBODY-WAITS-1 banked (no messages, no actors; ractor = compile-time ownership only; `&mut` IS the serialization; prime invariant "nobody waits for anything or any scheduling"; the ack-gated advance is canonical for the **ticket tier only** — `kanbanstep` stays canonical for stream reasoning, and the ack (the **SLA gate**) is reserved for the OGAR ticket tier, repurposed consumer-side as the **actionhandler queue** (see E-ACK-HARD-GATE-VS-KANBANSTEP-STREAM-1 + E-ACK-SLA-GATE-ACTIONHANDLER-QUEUE-1); supervisor KanbanMsg drivers = TD-MESSAGE-RESIDUE, leave-as-is per operator). `.claude/v3/VISION.md` = the AGI-aspiring canon, every claim graded [G]/[G-on-proxies]/[RULING]/[ASPIRATION], 2-Sonnet-preflight + 2-Opus-filigree provenance, all 9 review fixes applied. MODULE-TABLE ancestry census: thinking-engine (51) + p64-bridge (1) + cognitive-shader-driver (22 src + build.rs) with gem-status column. Plan addendum: D-MTS-1 frozen comparison — AriGraph context V3-TENANT-SHAPED; arm-discovery + DeepNSM ingest legs; CAM-PQ 6×8=48-bit path codes (address side) vs 6×palette256² = 12 B = one V3 tenant (value side). coderabbit 5/5 findings fixed in `093489c`. Merge `001839e`. | | **#674** | 2026-07-10 | V3 W2–W6 continuation + comma quorum/awareness measured + 5+3 council + StyleFamily dedup (D-TSC-1) | 16 commits. `contract::style_family::StyleFamily` (12 orchestration families, frozen ordinals Deliberate=0..Metacognitive=11) + `ThinkingStyle::family()` (36→12 total) + `default_runbook()` — five mutually divergent hand-rolled 12-style tables replaced (E-FIVE-STYLE-TABLES-1); first live 5+3 council run (spec v1→v2→v3). D-MTS-5 comma quorum MEASURED GREEN (N_eff 11.00/12; boundary = spectral participation) + D-MTS-6 comma awareness MEASURED GREEN (k*=1: one stored truth bit/comma level ≈ aligned k=4; lattice buys ≈log₂12 effective bits; D-MTS-6b gates real CE64 shrink). `BatchWriter::ack_and_propose` first-ack-wins dedup (codex P2). 1549 tests green. Merge `cd5178e`. | diff --git a/.claude/board/PR_ARC_INVENTORY.md b/.claude/board/PR_ARC_INVENTORY.md index 5a07e1e3e..0e418b491 100644 --- a/.claude/board/PR_ARC_INVENTORY.md +++ b/.claude/board/PR_ARC_INVENTORY.md @@ -3,11 +3,13 @@ - **Added.** `comet_tail_f16.py`/`.json` (bars committed BEFORE the run, `05f09005`), `comet_tail_resultant_instrument.py`/`.json` (deterministic bootstrap, seed pinned), report **§5.12 + §5.13**, §9.2's steering row superseded **in place**, EPIPHANIES `E-THE-CONTROL-SCORED-THE-HEADLINE-1`. All 13 headline figures in THIS entry verified against the committed JSONs before landing (13/13 exact — the #927 lesson, now standard). - **Locked — the leading moderator is DEAD `[G]`.** §9.2 had named the steering level *"the single most promising fix"*. Paired re-scoring of CT-F14's own 19 storms with ONLY the motion reference swapped (surface 6h displacement → 500/600/700 hPa disk-mean flow; stored centres, byte-identical `err_deg`): **F16a 0.579** vs a 0.70 bar (worse than surface's 0.684), **F16b residual 28.4 % WIDER** (68.29° → 87.71°) where ≥10 % tightening was predicted, **F16d level sweep monotone toward the SURFACE** — best 850 hPa, outside the ladder's 400–650 hPa band, no mid-tropospheric optimum. The height ladder itself (field-per-level, different quantity) is NOT falsified; its operational reading is. - **Locked — the control calibrated the instrument `[G]`.** F16c's deliberately **90°-rotated** reference scored **13/19 = 0.684, p=0.0835 — numerically identical to CT-F14's headline**. At n=19 the sign-test ladder is 11→0.324, 13→0.0835, 14→0.0318: **CT-F14 was never one storm short of significance; it was one storm short of distinguishability from an answer built to be wrong.** Rule banked: *an anti-vacuity control measures the RESOLVING POWER of the instrument, not only the test it guards* — attach a deliberately-wrong referent to any rate-headline and read the control's score as the floor. -- **Locked — the 0.684 plateau was the STATISTIC, not the data `[G]`.** Operator: *"die irrationale Aufsummierung hilft, dass der Dipol nicht auf 0.68 kollabiert."* Measured on the SAME 19 rows, no fetch: circular resultant **R̄=0.516, μ=−30.2°±36.5°, Rayleigh p=0.0050** where the sign test gave NO-VERDICT (p=0.0835). The rotated control — conflated at 0.684=0.684 by the sign test — shows identical R̄ with **μ shifted 100.3°**: the wrong referent is VISIBLE. Permuted control collapses to R̄=0.142, under the uniform floor 0.203. Mechanism: binarization discards magnitude AND mean direction; the vector sum keeps both, so the arc's systematic offset is now ESTIMATED (−30.2°) instead of penalized. A one-sided sign test on a distribution not centred at zero reports which SIDE the bias falls on — this arc used it as the primary instrument since §4; every prior sign number (2/2, 6/10, 8/10, 13/19) is bounded by §5.13 in BOTH directions. +- **Locked — the 0.684 plateau was the STATISTIC, not the data `[G]`.** Operator: *"die irrationale Aufsummierung hilft, dass der Dipol nicht auf 0.68 kollabiert."* Measured on the SAME 19 rows, no fetch: circular resultant **R̄=0.516, μ=−30.2°±36.5°, Rayleigh p=0.0050** where the sign test gave NO-VERDICT (p=0.0835). The rotated control — conflated with SURFACE at 0.684=0.684 by the sign test — is separated from surface in **BOTH channels**: R̄ 0.343 vs 0.516 AND μ −130.5° vs −30.2° (100.3° apart). The identical-R̄ statement belongs to the steering↔rotated pair (0.343 = 0.343, μ shifted exactly 90° — rotation preserves concentration by construction): the wrong referent is VISIBLE. *(This sentence originally fused the two comparisons into "identical R̄ with μ shifted 100.3°" — a FALSE composite of two individually-true numbers, caught by codex on #930 while unmerged, composed in place per the append-only rule's unmerged-PR allowance.)* Permuted control collapses to R̄=0.142, under the uniform floor 0.203. Mechanism: binarization discards magnitude AND mean direction; the vector sum keeps both, so the arc's systematic offset is now ESTIMATED (−30.2°) instead of penalized. A one-sided sign test on a distribution not centred at zero reports which SIDE the bias falls on — this arc used it as the primary instrument since §4; every prior sign number (2/2, 6/10, 8/10, 13/19) is bounded by §5.13 in BOTH directions. - **The Faltung frame (operator, same exchange).** The resultant IS the first circular Fourier coefficient (Faltung with `e^{iθ}`); the queued CT-W6 fit is a DEconvolution (component mix ⊛ apparatus noise, ±3–7° from CT-F4); on `Z_256` circular Faltung is substrate-native (`DistanceLut::circular()`'s domain). - **⊘ Scope, held.** §5.13's p=0.0050 is **post-hoc on the same sample — an instrument demonstration, NOT a promotion** of the directional claim; the report says so in three places. Sign-test conflation second finding: weak flow (<10 m/s, n=6) sign 0.833 / median |err| 103°; strong (≥10, n=13) sign 0.462 / |err| 55°; corr(speed,|err|)=−0.407 — magnitude improves with steering as physics expects while sign moves opposite. - **Deferred / queued.** **CT-W6** — two-component vector deconvolution of the dipole (neighbor far-field + bow-wave `½ρv_rel²`, GLOBAL coefficients over 38 observations, circular metrics, F16c-standard controls) — the candidate mechanism for the −30° offset AND the plateau (a vector sum cannot be scored against one bearing). CT-W1 folds in as the `c_geo` component. CT-W2s/W5/W7 (sunflower collision nodes with the 2-rail-address + 8-DOF facet carve, spiral-ADI aligned/orthogonal Fibonacci strides, Gegendruck), **CT-F17** (fresh-sample verdict). None run. +**Correction lesson (2026-08-12, codex P2 on #930):** the 13/13 figure verification checked every NUMBER against the JSONs and still missed this — because **a relation between two individually-correct numbers can be false**. "Identical R̄" and "100.3° apart" were each true of SOME pair, just not the same pair. Rule: *verify comparative claims as claims — every "identical/same/larger" must name both operands, and the check must evaluate the RELATION, not only its operands.* + **Confidence (2026-08-12):** merged. F16's failure verdicts are `[G]` (pre-registered, paired). The resultant numbers are `[G]` as measurements, `[S]` as evidence for the directional claim (post-hoc). The steering ladder survives as an n=2 observation. Structural claim (§9.1) untouched throughout. ## 2026-08-12 — lance-graph #927 (MERGED) — #926's board hygiene, one rounding caught by self-verification, and the rule I broke while citing it diff --git a/probes/weather-p1/COMET_TAIL_REPORT.md b/probes/weather-p1/COMET_TAIL_REPORT.md index cf5e3acc2..23e044e40 100644 --- a/probes/weather-p1/COMET_TAIL_REPORT.md +++ b/probes/weather-p1/COMET_TAIL_REPORT.md @@ -875,10 +875,20 @@ Three things, in order of importance: and offset (μ) come out as two numbers instead of eating each other. The systematic ≈−30° offset the arc has chased since §5.1 is now *estimated* (−30.2° ± 36.5°) instead of *penalizing the score*. -2. **The wrong referent is now visible.** F16c's rotated control scored - 0.684 = indistinguishable under the sign test. Under the resultant it has - the same R̄ (rotation preserves concentration, by construction) but μ - shifted **100.3°** — separated by well over both CIs. The instrument +2. **The wrong referent is now visible — in BOTH channels.** F16c's rotated + control scored 0.684 = indistinguishable from SURFACE under the sign test. + Under the resultant the control differs from surface in concentration + (R̄ 0.343 vs 0.516) AND direction (μ −130.5° vs −30.2°, **100.3°** apart, + over both CIs). Separately, the steering↔rotated pair confirms + rotation-invariance by construction: identical R̄ (0.343 = 0.343), μ + shifted exactly 90°. + + > **⚠ Corrected 2026-08-12 (codex on #930).** This item originally read + > "it has the same R̄ … but μ shifted 100.3°" — fusing the + > steering↔rotated comparison (where R̄ IS identical) with the + > surface↔rotated one (where the 100.3° lives) into a false composite of + > two true numbers. The 13/13 figure verification missed it because it + > checked operands, not relations. The instrument hierarchy is clean: real referent (0.516) > structured-but-wrong (0.343, wrong μ) > permuted (0.142, below the uniform floor). 3. **NOT a promotion.** Same sample, post-hoc — the p=0.0050 demonstrates the From 755a3c9674d4c5df253b9abab62945a623c66229 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 12 Aug 2026 11:03:27 +0000 Subject: [PATCH 3/4] =?UTF-8?q?report=20=C2=A710=20+=20weather-w-probes-v1?= =?UTF-8?q?:=20product-lead=20program=20+=20Sonnet=20worker=20briefs=20wit?= =?UTF-8?q?h=20stranded=20rescue?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Operator ask: "auf product lead level dokumentation erweitern und die followup testreihen für sonnet agents runterbrechen zum token sparen … und stranded rescue." - COMET_TAIL_REPORT.md §10 — the product-lead program: 10.1 measurement standard (circular resultant primary, control scores are floors, comparative RELATIONS verified as claims), 10.2 dipole vector-sum model incl. the STRANDED regime (v_rel = v_storm − v_env covers cut-off lows natively), 10.3 corridor α physics, 10.4 queue-and-bow anchors, 10.5 sunflower/facet/spiral-ADI substrate with [H] flags, 10.6 roadmap, 10.7 worker model + stranded rescue. - .claude/plans/weather-w-probes-v1.md — §0 verbatim-paste worker preamble (iron rules, fetch()/t-guard/wrap_deg/circular() code, controls-FIRST with VOID rule, checkpoint-per-storm + resume-skip + exec-runs heartbeat = the stranded-rescue protocol); full briefs W5 / W2s-a / W6; outlines W2s-b / W7; audit-gated CT-F17 spec; execution order + token budget. - Board hygiene per the new-plan rule: INTEGRATION_PLANS.md prepend + STATUS_BOARD.md rows D-W5 / D-W2sA / D-W6 / D-W2sB / D-W7 / D-CT-F17 (all Queued; F17 gated on W6 + mandatory independent adversarial audit). Exploratory tier ACTIVE; verdict tier gated. No probe was run in this commit. Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi --- .claude/board/INTEGRATION_PLANS.md | 26 +++ .claude/board/STATUS_BOARD.md | 16 ++ .claude/plans/weather-w-probes-v1.md | 302 +++++++++++++++++++++++++ probes/weather-p1/COMET_TAIL_REPORT.md | 136 +++++++++++ 4 files changed, 480 insertions(+) create mode 100644 .claude/plans/weather-w-probes-v1.md diff --git a/.claude/board/INTEGRATION_PLANS.md b/.claude/board/INTEGRATION_PLANS.md index 67ea92024..92a84be3d 100644 --- a/.claude/board/INTEGRATION_PLANS.md +++ b/.claude/board/INTEGRATION_PLANS.md @@ -1,3 +1,29 @@ +## 2026-08-12 — weather-w-probes-v1 (WORKER-BRIEF PLAN; the W-probe wave after CT-F16's failed bars + the resultant-instrument demo) + +Plan: `.claude/plans/weather-w-probes-v1.md`. Status **ACTIVE for the +exploratory tier**; the verdict tier (CT-F17) is GATED on CT-W6's result AND a +mandatory independent adversarial spec audit before any bar is committed. +Companion product-lead doc: `probes/weather-p1/COMET_TAIL_REPORT.md` §10 +(measurement standard · dipole vector-sum model incl. the stranded-storm regime +via v_rel · corridor α physics · queue-and-bow anchors · sunflower/facet/ +spiral-ADI substrate · roadmap · worker model). Operator ask: *"auf product +lead level dokumentation erweitern und die followup testreihen für sonnet +agents runterbrechen zum token sparen … und stranded rescue."* + +Shape: **§0 is a verbatim-paste worker preamble** (iron rules incl. +commit-bars-before-run; no board writes — one-writer rule; the fetch() code, +t-index guard, wrap_deg/err_deg, circular(); controls-FIRST with the VOID rule; +comparative sentences must name both operands — the #930 relation lesson baked +into the brief; **stranded rescue** = checkpoint every storm to +`.partial.jsonl` with resume-skip + tag-file heartbeat in `exec-runs/`, +so a killed worker resumes instead of re-fetching ~GBs). Wave 1 (parallel, +synthetic/stored, cheap): **W5** spiral-ADI anisotropy on Vogel N=4096 · +**W2s-a** golden-vs-grid pairing on the real cos-lat metric · **W6** the +two-component (geo + bow) deconvolution on the stored 19 storms — the +identification test for the vector-sum model, with the stranded stratification +(|v_storm| < 8 m/s vs ≥ 8 m/s) as bar B3. Gated: W2s-b (α window), W7, CT-F17 +(fresh 1959–1979 sample, V-test p<0.05 ∧ R̄≥0.35, audit mandatory). Doc-only. + ## 2026-08-11 — weather-substrate-evaluation-v1 (EVALUATION PLAN; the known-vs-test ledger for the whole #915→#922 arc) Plan: `.claude/plans/weather-substrate-evaluation-v1.md`. Operator ask: *"create diff --git a/.claude/board/STATUS_BOARD.md b/.claude/board/STATUS_BOARD.md index b31247df6..1756d2a44 100644 --- a/.claude/board/STATUS_BOARD.md +++ b/.claude/board/STATUS_BOARD.md @@ -1,3 +1,19 @@ +## weather-w-probes-v1 — W-probe queue (PRE-REGISTERED 2026-08-12) + +Plan: `.claude/plans/weather-w-probes-v1.md` (worker briefs; §0 preamble is +verbatim-paste for every Sonnet worker, incl. the stranded-rescue checkpoint +protocol). Product-lead frame: `probes/weather-p1/COMET_TAIL_REPORT.md` §10. +Wave 1 = parallel, no operator gate beyond go-ahead; gated rows named. + +| D-id | Deliverable | Wave | Status | Feeds | +|---|---|---|---|---| +| D-W5 | Spiral-ADI anisotropy: Vogel N=4096, iso ≤0.15 & aniso ≤1.25, non-Fibonacci stride control ≥1.5× | 1 | Queued | domino.rs gather design; [H] flags §10.5 | +| D-W2sA | Golden-vs-grid pairing on real cos-lat metric (zero-ties G1, CV G2) | 1 | Queued | facet-node addressing [G] extension | +| D-W6 | Two-component deconvolution (geo + bow, global lstsq, 38 eqs / 2 params; B3 = stranded stratification via v_rel) | 1 | Queued | dipole vector-sum identification; F17 gate | +| D-W2sB | α-window sweep β∈[0.85,1.15] | gated (W2s-a) | Queued | corridor α discriminator | +| D-W7 | Corridor two-regime α field probe | gated (W6) | Queued | §10.3 physics | +| D-CT-F17 | FRESH-sample verdict (1959–1979, N=70 candidates, V-test p<0.05 ∧ R̄≥0.35; independent adversarial spec audit MANDATORY before bars commit) | gated (W6 + audit) | Queued | the directional claim's verdict path | + ## 2026-08-11 — EV queue re-graded to v2 specs (post-audit; rows below unchanged in STATUS) All eleven EV rows in the block below still read **Queued** — correctly, none has diff --git a/.claude/plans/weather-w-probes-v1.md b/.claude/plans/weather-w-probes-v1.md new file mode 100644 index 000000000..1224ba741 --- /dev/null +++ b/.claude/plans/weather-w-probes-v1.md @@ -0,0 +1,302 @@ +# weather-w-probes-v1 — the W series as self-contained Sonnet worker briefs + +> **Status:** ACTIVE for exploratory probes (W5, W2s-a, W6, W2s-b, W7). +> **CT-F17 is verdict-tier and MUST NOT run** until its spec passes an +> independent adversarial audit (the 0-of-11 lesson, +> `E-ZERO-FOR-ELEVEN-THE-AUTHOR-CANNOT-AUDIT-HIS-OWN-FALSIFIERS-1`: the +> author is structurally the wrong person to find his spec's vacuous pass +> routes). All other bars here are author-written and unaudited — fine for +> exploratory tier, said out loud. +> +> **Why this file exists:** token economy + stranded rescue (operator, +> 2026-08-12). A worker receives §0 + ONE brief pasted verbatim — never this +> repo's history, never `COMET_TAIL_REPORT.md`. Product-lead context lives in +> the report §10; this file is executable instructions only. +> +> **Provenance of the physics:** report §10.2–§10.5 (dipole vector-sum model, +> Windkanal two-regime α, sunflower/collision facets, spiral-ADI). + +--- + +## §0 SHARED WORKER PREAMBLE — paste this block verbatim into every worker prompt + +You are a Sonnet grindworker executing ONE pre-specified measurement probe. +Bounded input, known output shape, no synthesis, no scope growth. + +**SCOPE — iron rules:** +1. You create/modify ONLY: your probe file `probes/weather-p1/.py`, its + outputs `.json` + `.partial.jsonl`, and your tag-file + `probes/weather-p1/exec-runs/.txt`. NOTHING else. Never touch + `.claude/board/*`, `COMET_TAIL_REPORT.md`, or another probe. +2. No `cargo` anything. No Rust. Python 3 + numpy + numcodecs only. +3. Commit the probe file with its bars BEFORE executing it (commit message: + `probes/weather-p1: pre-registered BEFORE the run`). Then run. Then + commit results. NEVER adjust a bar after seeing output; a failed bar is + reported as FAIL with the same prominence as a pass. +4. Do not claim anything you did not run. `NO-VERDICT` ≠ `FAIL`. If a fetch + or step dies, record it in the tag-file and stop — do not improvise. +5. Every module and every `def` gets a docstring (coverage gate). Docstrings + state WHAT is measured and WHY the bar can fail (no vacuous assertions). + +**DATA ACCESS (copy verbatim):** +```python +import json, pathlib, urllib.request +import numcodecs, numpy as np + +B = ("https://storage.googleapis.com/weatherbench2/datasets/era5/" + "1959-2022-6h-1440x721.zarr") +op = urllib.request.build_opener(urllib.request.ProxyHandler({})) +meta = json.loads(op.open(B + "/.zmetadata", timeout=90).read())["metadata"] + +def fetch(var, key): + """Fetch and decode one zarr chunk from the WB2 store.""" + za = meta[f"{var}/.zarray"] + raw = op.open(f"{B}/{var}/{key}", timeout=900).read() + dec = numcodecs.get_codec(za["compressor"]).decode(raw) + return np.frombuffer(dec, dtype=np.dtype(za["dtype"])).reshape(za["chunks"]) +``` +Chunk keys: surface vars `f"{t}.0.0"`; 13-level vars (`u_component_of_wind`, +`v_component_of_wind`, `geopotential`, …) `f"{t}.0.0.0"` — ALL 13 levels +arrive in one ~42 MB chunk. Levels: `[50,100,150,200,250,300,400,500,600, +700,850,925,1000]` hPa. Grid: 721×1440, 0.25°, lat from `fetch("latitude", +"0")`. Anchor guard (include verbatim, run before any fetch): +```python +import datetime +EPOCH = datetime.datetime(1959, 1, 1) +def t_index(dt): + """WB2 time index: 6-hourly steps since 1959-01-01.""" + return int(round((dt - EPOCH).total_seconds() / 3600 / 6)) +assert t_index(datetime.datetime(2021, 6, 15, 12)) == 91246 +``` +Store coverage ENDS 2021-12-31 18Z despite the filename; guard any t against +`meta["mean_sea_level_pressure/.zarray"]["shape"][0] - 1`. + +**GEOMETRY + SPINE (copy verbatim where needed):** `wrap_deg(d) = (d+180.0) % +360.0 - 180.0` (range `[-180,180)`); `err_deg(lp, ref) = +wrap_deg(rad2deg(lp - (ref + π/2)))`; the disk/ring decomposition is copied +verbatim from `probes/weather-p1/comet_tail_f16.py` (`geom_ll`/`geom`, +`low_pole_bearing`/`spine`) — R_E=6371.0, R_DISK=1200.0, RING=100.0. + +**STATISTICS STANDARD (report §10.1 — binding):** +```python +def circular(errs_deg, n): + """Resultant length R_bar, mean direction mu (deg), Rayleigh p (Zar/Mardia).""" + th = np.deg2rad(np.asarray(errs_deg)) + c, s = np.cos(th).mean(), np.sin(th).mean() + r = float(np.hypot(c, s)); mu = float(np.rad2deg(np.arctan2(s, c))) + z = n * r * r + p = float(np.exp(-z) * (1 + (2*z - z*z)/(4*n) + - (24*z - 132*z**2 + 76*z**3 - 9*z**4)/(288*n*n))) + return r, mu, max(min(p, 1.0), 0.0) +``` +Sign fractions may be PRINTED as descriptive context, NEVER used as a +pass/fail bar. Every bearing-related bar ships with TWO controls through the +identical pipeline: a +90°-rotated referent and a deterministically permuted +one (`(i+7) % n`). The controls' scores are reported FIRST; if a control +matches the real referent's score, the probe is VOID regardless of its bars. +Comparative sentences in your output must name BOTH operands ("R̄ identical +to X", never bare "identical") — a relation between two correct numbers can +be false. + +**OUTPUT + STRANDED RESCUE:** +- Final JSON beside the script: + `with open(pathlib.Path(__file__).with_name(".json"), "w") as fh: json.dump(out, fh, indent=2)` +- **Checkpointing (mandatory for any probe with per-unit fetches):** after + each unit of work (storm / node batch), append ONE line to + `.partial.jsonl` (`json.dumps(row)` + newline, flush). On startup, + read the partial file if it exists and SKIP completed keys (`t0` or unit + id). All RNG via `np.random.default_rng()`. + A stranded run is then rescued by re-invoking the same script — it resumes, + never re-fetches finished units. +- Tag-file `exec-runs/.txt`: append start line, one line per ~10 units + (heartbeat), and a final line with every bar verdict. This is YOUR only + log; the orchestrator consolidates boards. + +--- + +## §1 BRIEF W5 — spiral-ADI: two Fibonacci-stride sweeps ≈ one 2D diffusion? (Sonnet, zero fetch) + +**File:** `spiral_adi_probe.py`. **Seed:** 20260812. **No network.** + +**Objective.** On a Vogel lattice (`r = c·√k`, `θ = k·2π(1−1/φ)`, N=4096, +c chosen so max radius = 1.0), test whether alternating tridiagonal smoothing +sweeps along the two parastichy stride families approximate an isotropic 2D +diffusion — and that the result *depends on the strides being Fibonacci*. + +**Steps.** +1. Build the lattice. For each radius band (8 equal-area annuli), find the + dominant nearest-neighbor index-difference pair by measuring, for each + point, `argmin_j |x_{k+j} − x_k|` over `j ∈ {1..60}`; record the two most + frequent j per band (expect adjacent Fibonacci numbers; REPORT the bands + where they transition). +2. Per band, measure the crossing angle between the two stride directions at + each point (angle between `x_{k+j1}−x_k` and `x_{k+j2}−x_k`); report the + distribution (median, IQR, per band). +3. Sweep operator: for stride j, order points into chains `(start, j)` + within a band; one sweep = `y_i ← 0.25·y_{prev} + 0.5·y_i + 0.25·y_{next}` + along each chain (open ends: hold). One ADI iteration = sweep family A + then family B, band-appropriate strides. +4. Test field: Gaussian bump `exp(−|x−x0|²/2σ²)`, σ=0.08, x0 at radius 0.45. + Run 8 ADI iterations. Reference: sample the analytic heat-kernel-blurred + Gaussian (σ_ref² = σ² + 8·s² where s = the measured mean neighbor spacing + × 0.5 — CALIBRATE σ_ref by least-squares over σ_ref, then judge SHAPE) at + the lattice points. +5. Anisotropy metric: fit the blurred bump's second-moment tensor; ratio of + eigenvalues λ_max/λ_min. + +**Bars (pre-registered, commit before run):** +- **B1** (descriptive, no pass/fail): stride pairs per band + transition map + + crossing-angle table. +- **B2 ISO:** after best-σ_ref calibration, relative L2 error between ADI + result and isotropic reference ≤ **0.15**, AND second-moment anisotropy + λ_max/λ_min ≤ **1.25**. +- **B3 CONTROL (can-it-fail):** identical run with strides forced to + **12 and 18** (non-Fibonacci, non-coprime-ish) must give anisotropy ≥ + **1.5×** the Fibonacci run's. If the wrong strides smooth just as + isotropically, the Fibonacci claim measures nothing — say VOID. + +**Output JSON:** `{bands: [...], crossing_angles: {...}, iso_error, aniso_fib, +aniso_control, verdicts: {B2, B3}}`. + +--- + +## §2 BRIEF W2s-a — golden two-lattice pairing on REAL lat/lon geometry (Sonnet, zero fetch) + +**File:** `sunflower_pairing_probe.py`. **Seed:** 20260812. **No network.** + +**Objective.** The collision-node construction assumes the two sunflower +lattices pair generically (no ties, even pair distances). Prove it survives +the real cos-lat metric — the #921 lesson says disk properties do NOT +automatically transfer. + +**Steps.** +1. Two Vogel lattices, N=2048 each, disk radius 1500 km, centers at + (55.0N, 340.0E) and (55.0N, ~361.9E) → 1400 km apart at that latitude. + Project to km via the metric `dx = R_E·cos(lat_c)·Δlon_rad`, + `dy = R_E·Δlat_rad` (R_E=6371.0) — the same `geom_ll` convention. +2. Overlap band: points of lattice H within 900 km of center T and vice + versa. Nearest-pair map: for each H-point in band, its nearest T-point. +3. Control: TWO axis-aligned square grids of identical point density over + the same two disks, same metric, same pairing procedure. + +**Bars:** +- **G1 TIES:** count of exact-duplicate nearest-pair distances (float64 + equality after rounding to 1e-9 km): golden = **0**; grid control **> 0** + (if the grid also has zero ties, the tie test is vacuous on this geometry — + report VOID for G1 and rely on G2). +- **G2 EVENNESS:** coefficient of variation of nearest-pair distances: + golden CV **< grid CV** (strict). +- **G3** (descriptive): χ² of pair-midpoint density against uniform across + the corridor band, both constructions, reported not judged. + +**Output JSON:** `{n_pairs, ties_golden, ties_grid, cv_golden, cv_grid, +chi2_golden, chi2_grid, verdicts: {G1, G2}}`. + +--- + +## §3 BRIEF W6 — the dipole deconvolution: neighbor + bow, global fit (Sonnet, ~40 chunks) + +**File:** `comet_tail_w6.py`. **Seed:** 20260812. +**Inputs:** `comet_tail_f16.json` rows (19 storms; fields per row: `date`, +`t0`, `center_lat`, `center_lon`, `displacement_km`, `err_surface_deg`, +`low_pole_rad`, `steer_bearing_rad`, `steer_speed_ms`). + +**Objective.** Test the report-§10.2 model: the dipole vector is a GLOBAL +linear combination of a neighbor-far-field predictor and a bow-wave +predictor. This is a mechanistic test on stored storms — NOT a verdict +(that is CT-F17). + +**Per storm (checkpoint one JSONL row each):** +1. Fetch MSLP `f"{t0}.0.0"`. Recompute the CONSTRAINED spine dipole vector + `D_i = (a1, b1)` via the verbatim `spine()` code from + `comet_tail_f16.py` (ring means + lstsq on `[r·cosθ, r·sinθ]`) about the + stored center. +2. **Motion bearing recovered by ALGEBRA, no tracking:** + `mth_deg = wrap_deg(rad2deg(low_pole_rad) − 90 − err_surface_deg)`. + (Exact inversion of `err = wrap(lp − (mth+90))`.) Motion speed = + `displacement_km` per 6 h → m/s. +3. Fetch u/v `f"{t0}.0.0.0"`; disk-mean 850 hPa wind vector `v_env850` + (1200 km disk). **`v_rel = v_storm − v_env850`** (vector, m/s). Bow + predictor: `P_bow,i = (0.5·1.2·|v_rel|²) · û(bearing(v_rel) + 180°)` + [Pa · unit-vector — low pole BEHIND relative motion]. +4. Neighbor: zonal-anomaly field `p − mean(p, axis=1)`; within the annulus + **600–2500 km** of the center and lat 20–80N, find the strongest POSITIVE + anomaly cell → `A_H` [Pa], distance `d_H` [km], bearing `θ_H`. Neighbor + predictor: `P_geo,i = (A_H/d_H) · û(θ_H + 180°)` [Pa/km · unit-vector — + background low pole points AWAY from the H]. +5. Store row: `t0, D, P_geo, P_bow, |v_storm|, |v_rel|, A_H, d_H`. + +**Global fit:** stack 19×2 = 38 scalar equations +`D = c_geo·P_geo + c_bow·P_bow`; solve `(c_geo, c_bow)` by lstsq (2 free +parameters — units absorb into the c's; state this). Define +`R²_vec(model) = 1 − Σ|D_i − D̂_i|² / Σ|D_i − mean(D)|²`. Fit also the two +single-predictor models. + +**Bars (pre-registered):** +- **B0 CONTROLS FIRST:** joint fit with per-storm PERMUTED `P_bow` + (`(i+7) % 19`) and, separately, with `P_bow` rotated +90°. Either control's + joint `R²_vec` must stay ≤ single-geo `R²_vec` + **0.03**; otherwise the + joint gain is degrees-of-freedom, probe VOID. +- **B1 IDENTIFIABILITY:** joint `R²_vec` ≥ best single `R²_vec` + **0.10**. +- **B2 SIGN:** `c_bow > 0` AND `c_geo > 0` (both components in their + physically-predicted directions). +- **B3** (descriptive): resultant `(R̄, μ, p)` of residual bearings, overall + AND stratified by `|v_storm|` < / ≥ 8 m/s — the **stranded stratum**: if + §10.2's stranded-rescue reading is right, the weak-`|v_storm|` residuals + should NOT be worse than the strong ones once `v_rel` carries the bow. +- **B4** (descriptive): per-storm table `bearing(D)` vs `bearing(D̂)`. + +**Cost:** 19 MSLP (~2 MB ea) + 19 u/v chunks (~42 MB ea) ≈ 850 MB, ~3 min. +Checkpoint after every storm; resume skips completed `t0`. + +--- + +## §4 BRIEF W2s-b — the α-field on a real H–T pair (GATED on W2s-a G2 pass) + +**File:** `corridor_alpha_probe.py`. Outline — finalize bars at spawn time +using W2s-a's measured pairing stats: +- One timestep (T0=91246), fields: MSLP + 10m u/v. Deepest NH low + + strongest H within 3000 km (zonal anomaly), per W6 §3.4's finder. +- Sunflower pair per W2s-a geometry; nodes in the corridor band. +- Per node: `|∇p|` (centred differences on the 0.25° grid, cos-lat metric, + longitude WRAPS via np.roll — the #926 lesson), measured `|v|`, + `f = 2Ω·sin(lat)`. +- **Core bar (fixed now):** regression of measured `|v|` on geostrophic + `|∇p|/(ρf)` over open-corridor nodes (|lat| ∈ 30–75, ∇p above its median): + slope β ∈ **[0.85, 1.15]** — the geostrophic regime must recover itself, + else the α machinery measures nothing here. Permuted-node control R² < + **0.1**. Descriptive: per-segment α map; overlay vs the temperature- + gradient contested band (`go_territory` §-style gradT). + +## §5 BRIEF W7 — Gegendruck (GATED on W2s-b) + +Outline only: two consecutive timesteps; corridor lanes = aligned-family +spiral bands; per lane, along-lane mass-flux convergence at t0 vs Δp +(t1−t0) in lanes i±1. Bars to be written AFTER W2s-b fixes the lane +geometry; must include rotated-lane and time-shuffled controls. + +## §6 CT-F17 — the fresh-sample VERDICT (GATED: W6 result + independent adversarial spec audit) + +The only verdict-tier probe. Shape (parameters frozen NOW, bars audited +before run): mechanically generated dates in **1959–1979** (the only unused +window; samples 1–3 used 2015–21 / 1980–95 / 1996–2010), fixed start +1959-02-01 12Z + fixed stride 53 days, N_CANDIDATES=70 targeting ≥20 +qualifying at the observed 0.28 rate; displacement ≥ 250 km/6h a priori; +score each storm's dipole bearing against the **W6-fitted model's predicted +bearing** (fallback branch if W6 fails B1: against surface motion, stated in +advance). Statistics per §10.1: V-test toward the training-fixed μ, p < +0.05, R̄ ≥ 0.35, rotated + permuted controls with floors. **The spec text +goes to an independent adversarial audit (5+3 or codex-style) BEFORE +execution — non-negotiable.** + +## §7 Execution order + token budget + +| wave | probes | parallel? | worker tokens (est.) | +|---|---|---|---| +| 1 | W5, W2s-a, W6 | yes — disjoint files, no shared state | ~15–25k each (brief + code + run logs) | +| 2 | W2s-b | after W2s-a | ~20k | +| 3 | W7 | after W2s-b | ~20k | +| 4 | F17 spec audit → run | after W6 | audit ~30k, run ~25k | + +Orchestrator (Opus) does: spawn with §0+brief, consolidate tag-files, write +boards, interpret results, decide F17's branch. Workers never interpret. diff --git a/probes/weather-p1/COMET_TAIL_REPORT.md b/probes/weather-p1/COMET_TAIL_REPORT.md index 23e044e40..1f7f05d8a 100644 --- a/probes/weather-p1/COMET_TAIL_REPORT.md +++ b/probes/weather-p1/COMET_TAIL_REPORT.md @@ -1355,3 +1355,139 @@ The shapes already exist and are proven: Gate, unchanged: train/test on disjoint decades, pre-registered bars, adversarial audit (plan §8) before any of it is called more than a probe. +--- + +## 10. The follow-up program — instrument standard, working model, corridor physics, compute substrate `[mixed grades, per item]` + +Product-lead summary of where the arc stands after §5.12/§5.13 and what the +W-probe series tests. Worker-executable briefs live in +`.claude/plans/weather-w-probes-v1.md`; this section is the WHY, that plan is +the HOW. + +### 10.1 Measurement standard (binding for every probe from here) + +| rule | reason | +|---|---| +| **Primary statistic: circular resultant** (R̄ concentration, μ mean direction, Rayleigh/V-test p; bootstrap CI on μ) | the sign test collapsed vectors to bits and saturated at 0.684 — the same rows resolve at p=0.0050 under the resultant (§5.13) | +| **Control floors mandatory**: a +90°-rotated AND a permuted referent through the *identical* pipeline; the control's score is the floor any headline must clear | a rotated control scored CT-F14's exact headline (§5.12) — a control measures the instrument's RESOLVING POWER, not just the test it guards | +| Magnitude metrics alongside (median \|error\|, sd); sign fractions may be *reported*, never verdict-grade | a sign test on an offset distribution reports which SIDE the bias falls on, not correctness (§5.12, weak/strong stratification) | +| Pre-registration committed **before** execution, commit hash cited | the arc's standing discipline; bars never move after output exists | +| **Verify comparative claims as claims** — every "identical/same/larger" names both operands and the check evaluates the relation | two individually-true numbers made one false sentence (§5.13 correction); a figure check cannot catch a relation error | + +### 10.2 The working model: the dipole is a vector sum `[H]` + +Every single-referent scoring (motion 0.684, steering 0.579, everything +between) plateaued because a **vector sum cannot be scored against one +bearing**. The candidate decomposition: + +``` +D = c_geo · (A_H/d_H) · û_awayFromH the neighbor far-field (Hochdruck) + + c_bow · (½ρ|v_rel|²) · (−v̂_rel) the bow wave (Bugwelle: high ahead, low behind) + + r the Nahkampf residual (cold pools, collision turbulence) +``` + +- **The neighbor term** is §2's own derivation read backwards: the linear + background gradient IS the first-order far-field of the adjacent high — the + H–T collision was inside the spine all along, filed under "background". +- **The bow term** is stagnation pressure `½ρ·v_rel²` ahead of relative + motion — meteorology's own *bow echo*; the superposition rotates the summed + low pole beyond 90°-left toward the rear, which is the **shape of the + −30° offset** (μ = −30.2° ± 36.5°, §5.13) the arc has chased since §5.1. +- **The stranded regime is covered natively** (operator: *"stranded + rescue"*): the bow variable is `v_rel = v_storm − v_env`, so a cut-off / + quasi-stationary low in ambient flow has a rock-in-river bow + (`½ρ|v_env|²` upstream) even at zero displacement. §5.12's weak-flow + paradox — sign 0.833 with median \|error\| 103° at <10 m/s steering — is + the signature of scoring *stranded* systems against a motion referent that + barely exists. CT-W6 stratifies by \|v_storm\| to test exactly this. +- **Identifiability**: fitted with GLOBAL coefficients across storms (38 + observations, 2 parameters), conditioned by the natural phase diversity + between neighbor- and motion-bearings. Per-storm free amplitudes would be + an exact fit and meaningless — the global constraint is the test. + +### 10.3 Corridor physics (Windkanal) — two regimes, one measured exponent `[G] physics, [S] wiring` + +| | geostrophic corridor (synoptic) | gap flow / Düseneffekt | +|---|---|---| +| balance | Coriolis ⊥ ∇p | ageostrophic, Bernoulli | +| law | `v = (1/ρf)·\|∇p\|` | `v ≈ √(2Δp/ρ)` | +| direction | along isobars, low on the left (NH) | through the gap, H → T | +| scaling | **linear** in the gradient | **square-root** in the drop | + +In the corridor the two gradients ADD (`\|∇p\| ≈ \|dp_H/dr\| + \|dp_T/dr\|` at +the midpoint — both computable from the two ring profiles), which is why the +wind maximum lies *between* the centers. Instead of choosing a regime, fit +the exponent `α` per corridor segment: **α ≈ 1 geostrophic, α ≈ ½ Bernoulli, +and the α-DEVIATION field marks where volumetric math stops** — the measured +boundary of the Nahkampf sector. Note `√(2Δp/ρ)` is also the cold-pool +density-current speed: gap flow and the Abregnen gust front are the same +ageostrophic law at two scales. + +### 10.4 The queue-and-bow mechanics — all four metaphors have [G] anchors + +| mechanism | physics | anchor | +|---|---|---| +| Domino in der Schlange | kinematic waves in a 1D chain (density waves run backwards) | Lighthill–Whitham 1955 — one paper, rivers AND traffic; jamitons ≡ roll waves | +| Ausweichen → Gegendruck Nebenspur | blocked along-lane flux deflects laterally, pressurizing the neighbor streamtube | mass continuity; MOBIL lane-change back-pressure | +| Flugzeug-Formel | Bernoulli conversion at the evasion node (speed ↔ pressure) | `p + ½ρv² = const`; NOT lift-on-a-free-vortex (§ demarcation: a free vortex advects — bound-vortex lift needs a surface) | +| Bugwelle → Windentstehung | stagnation ridge `½ρv_rel²` ahead of a moving system drives new wind | bow echo; gust front; isallobaric wind | + +Discretization: lanes = aligned spiral-arm family, Gegendruck = the +orthogonal family, Bernoulli = the node nonlinearity, Bugwelle = the +moving-boundary source term — i.e. the machinery of §10.5, no new substrate. + +### 10.5 Sampling & compute substrate — sunflower, collision facets, spiral-ADI + +**Why the golden lattice** (operator: *Verteilung und Kollision im +irrationalen Raum*), three `[G]` properties: (1) three-distance evenness — +at any N the azimuthal gaps take ≤3 values; discrepancy `O(log N/N)` vs +Monte-Carlo's `O(1/√N)`; (2) two lattices over different centers are +**incommensurate** — no moiré, no ties, generic unique nearest-pairings +(grids produce tie families exactly where the corridor bands are); (3) +**prefix extensibility** — any prefix is well-distributed, so per-node +accuracy grows monotonically by extending the sequence, no re-meshing. + +**The collision node is ONE V3 facet** (12 B): rails 0–1 = `(k_H:k_T)` — the +pair address, from which position, radii and azimuths are *implied* (place +deterministic, residue stored); rails 2–5 = **8 state bytes** (the 8 +Freiheitsgrade). Two sanctioned carvings — spine-pair +(`p_H, g_H, p_T, g_T, u, v, α, resid`) vs kinematic/frontogenesis +(`u, v, div, ζ, stretch, shear, p, ∇T` — the Petterssen machinery) — and the +ClassView picks the reading per class, which is exactly what content-blind +facets are for. + +**Spiral-ADI on domino.rs**: stride-1 is the *scatter* ordering (sampling, +not physics); the spatial axes are the **Fibonacci stride families** +`F_j`/`F_{j+1}` (the two parastichy arm families, winding oppositely, +crossing quasi-orthogonally). Two tridiagonal sweeps — aligned then +orthogonal — are ADI (Peaceman–Rachford): a 2D diffusion/elliptic operator +from two 1D passes of domino's **existing fixed tridiagonal kernel**, which +is *correct for this use* (demarcated from §9.4's moderator non-claim: the +learned model is still absent; the physics operator is not). Missing piece: +only the facet→lane gather. Honest `[H]` flags, each gated by a W-probe +bar: parastichy strides change with radius (piecewise segments), crossing +angles vary, cos-lat distorts equal-area on the real grid (the #921 lesson). + +### 10.6 Roadmap + +| probe | question | core pre-registered bar | inputs | est. cost | tier | gated on | +|---|---|---|---|---|---|---| +| **W5** spiral-ADI | do two Fibonacci-stride tridiagonal sweeps ≈ isotropic 2D diffusion? | iso-error ≤ 0.15; non-Fibonacci stride control ≥ 1.5× more anisotropic | none (synthetic) | minutes, 0 fetch | Sonnet | — | +| **W2s-a** pairing geometry | does the golden two-lattice pairing survive real lat/lon metric? | zero ties + CV(pair distances) < grid CV on identical geometry | none (lattice math) | minutes, 0 fetch | Sonnet | — | +| **W6** dipole deconvolution | is D = c_geo·neighbor + c_bow·bow identifiable? | joint R²_vec ≥ best single + 0.10; c_bow > 0; permuted-v_rel control collapses | 19 stored rows + ~40 chunks | ~5 min fetch | Sonnet (spec-bound) | — | +| **W2s-b** α-field | does the geostrophic regime recover itself on a real H–T pair? | α ∈ [0.85, 1.15] in the open corridor; permuted control R² < 0.1 | 1 timestep, 3 vars | ~3 min | Sonnet | W2s-a | +| **W7** Gegendruck | does lane-i convergence pressurize lane i±1? | outlined in plan; bars TBD at spec time | 2 timesteps | ~5 min | Sonnet | W2s-b | +| **CT-F17** fresh-sample verdict | does the W6-fitted model predict on UNSEEN storms? | V-test toward the pre-registered μ, p < 0.05, R̄ ≥ 0.35, controls | fresh 1959–1979 sample | ~30 min | Sonnet run, **Opus spec** | W6 + **independent adversarial spec audit** (the 0-of-11 lesson: the author cannot audit his own falsifiers) | + +### 10.7 Worker execution model — token economy + stranded rescue + +The briefs in `weather-w-probes-v1.md` are **self-contained**: an orchestrator +pastes §0 (shared preamble, ~70 lines) + one brief (~60–90 lines) into a +Sonnet worker — the worker never loads this report or the session history. +Every brief carries the **stranded-rescue protocol**: probes checkpoint one +JSONL row per unit of work (`.partial.jsonl`), fetches are resumable +(completed `t0`s skipped on restart), all randomness is seeded, and a +heartbeat tag-file lets the orchestrator detect a stranded run and hand the +*partial* state to a fresh worker instead of re-paying the fetches. Workers +write ONLY their probe files + their own tag-file — never board files +(one-writer rule). From 15f2e32f5e30a1ea715c10695adb4cb58bd37d95 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 12 Aug 2026 11:13:36 +0000 Subject: [PATCH 4/4] board+report: open-review sweep #920-#930 -- three frozen figures wrong, and the append-only audit could not fail Operator: "930 has comments / check also previous 5 / if you want go back another 5." All 68 review comments on #920-#930 enumerated and checked against the TREE, not the merge. Depth is non-uniform and said so: #922-#930 finding-by-finding; #920 (27) and #921 (5) spot-checked on P1/governance only (both clean), ~30 older findings there left explicitly UNVERIFIED. Clean: #922/#924/#925/#929 zero comments; all 27 of #926's findings fixed in the tree (equal-budget grid_pts + regenerated E2 JSON, seam-wrapping subgrid_min, find_center->None, F7d 35->40, CT_F12 NO-VERDICT, persisted storm metadata, np.roll longitude, __file__-relative write, net-decay E6, the 93-97%->90.9-94.3% headline); #923's plan-status P2 resolved. Three open, all frozen in append-only ledgers, all corrected in NEW entries: 1. "+92.76 Pa moves R2 in the 5th decimal" is refuted by the report's own carve table 15 lines above it: carve A's +92.76 Pa moved R2 0.9212->0.9129 (0.0083, THIRD decimal); carve D's +1.59 Pa moved it 2.4e-06 (SIXTH). The var() BUG was blind at every magnitude; the STATISTIC is near-blind only in the single-digit-Pa regime -- which is exactly where "lossless" was claimed. 2. "10 probe scripts with committed JSON" undercounts: 15 .py added, 11 with a committed .json. 3. "+13/-0, +10/-0, +0/-0": measured +13/-0, +17/-0, and EPIPHANIES absent from the net diff -- it was +1/-1 at 0f9e6bcd (the in-place edit), zeroed by the revert. The +0/-0 was a revert artefact, not evidence of purity. And the audit METHOD is replaced: zero-deletions proves ADDITIVE, not PREPEND. The suffix check (new.endswith(old)) proves it. Both halves measured before banking: a constructed end-append scores zero-del=True/suffix=False; a true prepend True/True. It also fires on THIS PR (LATEST_STATE suffix=False at +15/-0) -- correctly, since #930 inserts table rows mid-file and composes an unmerged entry in place. A False is a demand for justification, not a verdict. Same defect shape as #930's relation error, twice in three days: two individually-true numbers asserted of one pair when each belongs to a different one. A figure-by-figure check cannot catch it; the relation must be evaluated as a claim. Already standing in report SS10.1 and the worker preamble of weather-w-probes-v1 (both verified present, not assumed). Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi --- .claude/board/LATEST_STATE.md | 13 ++++++++++++ .claude/board/PR_ARC_INVENTORY.md | 14 +++++++++++++ probes/weather-p1/COMET_TAIL_REPORT.md | 28 ++++++++++++++++++++------ 3 files changed, 49 insertions(+), 6 deletions(-) diff --git a/.claude/board/LATEST_STATE.md b/.claude/board/LATEST_STATE.md index 43c7055b3..7a35551fc 100644 --- a/.claude/board/LATEST_STATE.md +++ b/.claude/board/LATEST_STATE.md @@ -1,3 +1,16 @@ +## 2026-08-12 — open-review sweep of #920–#929 — three ledger figures were wrong, and one audit method cannot do what it claims + +### Current Contract Inventory — no new types (corrections to merged #927/#928 entries; append-only, corrected here not in place) + +Operator: *"930 has comments … check also previous 5 … if you want go back another 5."* All 68 review comments on #920–#930 enumerated and checked against the tree rather than against the merge — **with non-uniform depth, stated rather than glossed:** #922–#930 finding-by-finding; **#920 (27) and #921 (5) spot-checked on their P1/governance items only** (both clean — `p1_noise_floor.py` is explicitly `SUPERSEDED by p1_ci_vs_floor.py`; the merged #917 entry has exactly one commit in its history, so it was never mutated), **leaving ~30 older findings there UNVERIFIED**. Within the verified range: **#922, #924, #925, #929: zero review comments. #926's 27 findings: all verified fixed in the tree** (equal-budget `grid_pts`, seam-wrapping `subgrid_min`, `find_center → None`, F7d 35→40, CT-F12 `NO-VERDICT-INSUFFICIENT-N`, persisted storm metadata, `np.roll` longitude, `__file__`-relative JSON write, net-decay E6, the 93–97 %→90.9–94.3 % headline). **#923's plan P2 is resolved** (`Status: ACTIVE (audited 2026-08-11)`). Three findings on **#927/#928 were still open**, all of them numbers frozen in append-only ledgers: + +- **⊘ CORRECTION to #926's entry (line 29 below), codex P2 on #927 — "a +92.76 Pa offset moves R² in the 5th decimal" is FALSE, and the report's own carve table refutes it.** Measured: **+92.76 Pa (carve A) moved R² 0.9212 → 0.9129 = 0.0083, the THIRD decimal**; **+1.59 Pa (carve D) moved it 0.943406 → 0.943403 = 2.4e-06, the SIXTH.** The sentence fused carve A's magnitude with carve D's insensitivity. **The true and sharper statement: the `var()` BUG was blind at EVERY magnitude — that is what hid +92.76 Pa — while the STATISTIC is near-blind only in the single-digit-Pa regime, which is exactly where "lossless" was claimed on four-decimal agreement.** The banked doctrine (RMSE + mean bias in Pa beside every R²) is unchanged; only its illustration was mispaired. **Third instance this week of the same shape as #930's relation error** — two individually-true numbers asserted of one pair when each belongs to a different one. +- **⊘ CORRECTION to #927's `Added` line, codex P2 — "10 probe scripts with committed JSON results" undercounts.** Measured on the #926 first-parent diff: **15 `.py` added, 11 of them with a committed `.json`** (the four `ev*` scripts have none). The immutable scope summary omitted delivered work. +- **⊘ CORRECTION to #928's audit figures (line 17 below), codex P2 — "+13/−0, +10/−0, +0/−0" is wrong in two places.** Measured `a4e264c5..3dae97be`: LATEST_STATE **+13/−0** ✓, PR_ARC_INVENTORY **+17/−0** (never +10 at any commit in the range), and EPIPHANIES **does not appear in the net diff at all** — it was **+1/−1 at `0f9e6bcd`**, the in-place edit of a merged entry, which the revert zeroed. So the "+0/−0" was the *product of a revert*, not evidence of purity — and read correctly it is the stronger result: **the audit WOULD have caught the violation mid-PR.** +- **⚠ AND THE AUDIT METHOD ITSELF IS WEAKER THAN CLAIMED (codex P2 on #928).** *Zero removed lines proves ADDITIVE, not PREPEND* — an insert in the middle of a ledger, or an append at the end, also deletes nothing. Since #928 banked it as "worth reusing", a future session would inherit a test that cannot detect a mid-file insertion. **The correct one-command pure-prepend test is the SUFFIX check** (CodeRabbit's own analysis chain used it): `git show origin/main:F` must be a suffix of `git show HEAD:F` — i.e. `new.endswith(old)`. Zero-deletions stays useful as the cheap first screen; the suffix check is the one that proves the property. + - **Both halves measured before banking this, per the falsifiability rule.** Constructed falsifier: a true prepend gives `zero-del=True, suffix=True`; an **append at the END** gives `zero-del=True, suffix=False` — caught only by the suffix check, which is the discrimination the old method lacks. + - **And it fires on THIS PR, correctly.** `origin/main..HEAD` for #930: `PR_ARC_INVENTORY` / `STATUS_BOARD` / `INTEGRATION_PLANS` all `suffix=True`; **`LATEST_STATE` `suffix=False` at `+15/−0`** — the old screen passed it, the new one flags it. The flag is right and the change is still legitimate: #930 inserts shipped-PR table rows **mid-file** (§Recently Shipped PRs) and composes an **unmerged** entry in place (the PR-878 allowance). **A `False` is not automatically a violation — it is a demand for the justification to be stated.** That is the property worth having; the old test could not even ask. + ## 2026-08-12 — lance-graph #929 (MERGED) — steering rescue dead; sign test retired as primary instrument ### Current Contract Inventory — no new types (2 probes + report §5.12/§5.13 + 1 epiphany) diff --git a/.claude/board/PR_ARC_INVENTORY.md b/.claude/board/PR_ARC_INVENTORY.md index 0e418b491..78b59a4b9 100644 --- a/.claude/board/PR_ARC_INVENTORY.md +++ b/.claude/board/PR_ARC_INVENTORY.md @@ -1,3 +1,17 @@ +## 2026-08-12 — open-review sweep of #920–#930 — three frozen figures corrected, and the append-only audit replaced with one that can actually fail + +- **Scope, stated at the coverage actually achieved.** Operator asked for the open comments on #930, then the previous five, then five more. All 68 review comments on **#920–#930** were enumerated; the checking depth is **not uniform and is not claimed to be**: #922–#930 were verified finding-by-finding against the tree, while **#920 (27 findings) and #921 (5) were spot-checked only** on their P1/governance items — both clean (`p1_noise_floor.py` carries an explicit `SUPERSEDED by p1_ci_vs_floor.py` header that fixes the decoded-error framing *and* the 255-vs-256 bucket-edge bug; the merged **#917** entry has exactly one commit in its whole history, `4be809c8`, so it was never mutated). **The remaining ~30 older coderabbit findings on #920/#921 are UNVERIFIED here** — recorded as such rather than absorbed into a clean-sweep claim. Result within the verified range: **#922 / #924 / #925 / #929 carry zero review comments**; **all 27 findings on #926 are fixed in the tree** (verified individually: `grid_pts` returns exactly `n` and the E2 JSON is regenerated at 64/64/64 · 256/256/256 · 1024/1024/1024; `subgrid_min` wraps longitude via `np.take(..., mode="wrap")`; `find_center` returns `None` on an empty mask; F7d tests `>= 40.0` matching its key; `CT_F12` emits `NO-VERDICT-INSUFFICIENT-N` instead of `pass: true`; storm metadata persisted; `d_dx` uses `np.roll`; `go_territory` writes `__file__`-relative; E6 requires a net decline; the 93–97 % headline superseded by 90.9–94.3 %); **#923's plan-status P2 is resolved**. Three findings on **#927/#928** were genuinely open, and all three were numbers already frozen in append-only ledgers. + +- **Corrected — the R² decimal claim `[G]` (codex P2 on #927; #926's entry + report §6.1).** *"A +92.76 Pa offset moves R² in the 5th decimal"* is refuted by the carve table 15 lines above it in the report: **+92.76 Pa (carve A) moved R² 0.9212 → 0.9129 — 0.0083, the THIRD decimal**; **+1.59 Pa (carve D) moved it 0.943406 → 0.943403 — 2.4e-06, the SIXTH.** Carve A's magnitude was fused with carve D's insensitivity. **True statement: the `var()` BUG was blind at every magnitude — that is what concealed +92.76 Pa — while the STATISTIC is near-blind only in the single-digit-Pa regime, which is precisely where "lossless" was inferred from four-decimal agreement.** The doctrine it supports (RMSE + mean bias in Pa beside every R²) is untouched; only the illustration was mispaired. **This is the same failure shape as #930's relation error, found two days apart** — see the correction lesson below. + +- **Corrected — the delivered-scope count `[G]` (codex P2 on #927).** #927's `Added` line records *"10 probe scripts with committed JSON results"*. Measured on #926's first-parent diff: **15 `.py` added, 11 with a committed `.json`** (the four `ev*` scripts ship none). An immutable scope summary that undercounts delivered work permanently hides it. + +- **Corrected — the audit figures `[G]` (codex P2 on #928).** *"+13/−0, +10/−0, +0/−0"*: measured `a4e264c5..3dae97be` gives LATEST_STATE **+13/−0** ✓, PR_ARC **+17/−0** (never +10 at any commit in the range), EPIPHANIES **absent from the net diff** — it stood at **+1/−1 at `0f9e6bcd`**, the in-place edit of a merged entry, which the revert zeroed. The "+0/−0" was therefore *the product of a revert*, not evidence of purity — and read correctly it is the **stronger** result: the audit would have caught the violation while the PR was live. + +- **⚠ REPLACED — the append-only audit itself `[G]`.** #928 banked *"zero removed lines proves a pure prepend"* as reusable. It does not: **zero deletions proves ADDITIVE, not PREPEND** — a mid-file insert or an end-append deletes nothing either. The property is proved by the **SUFFIX check**: `git show origin/main:F` must be a suffix of `git show HEAD:F` (`new.endswith(old)`). **Both halves measured before banking, per the falsifiability rule:** a constructed end-append scores `zero-del=True, suffix=False` — caught only by the new test; a true prepend scores `True/True`. **It also fires on #930 itself**: `LATEST_STATE` is `suffix=False` at `+15/−0` where the old screen passed, because #930 inserts shipped-PR rows mid-file and composes an unmerged entry in place (the PR-878 allowance). **A `False` is a demand for justification, not an automatic violation** — and the old test could not even raise the question. Zero-deletions survives as the cheap first screen. + +**Correction lesson (2026-08-12, second statement in three days).** #930's *"identical R̄, μ shifted 100.3°"* and #927's *"+92.76 Pa moves R² in the 5th decimal"* are the **same defect**: two individually-correct numbers asserted of one pair when each belongs to a different pair. Verifying every figure against its JSON — done, 13/13 and 9/10 — cannot catch it, because **every operand is true**. The check that catches it: *for each comparative claim, name both operands explicitly and evaluate the RELATION*. Now standing in the report's §10.1 measurement standard and in the §0 worker preamble of `weather-w-probes-v1`, so the fleet inherits it rather than rediscovering it. + ## 2026-08-12 — lance-graph #929 (MERGED) — CT-F16 kills the steering rescue; the control scores the headline; the resultant resolves what the sign test collapsed - **Added.** `comet_tail_f16.py`/`.json` (bars committed BEFORE the run, `05f09005`), `comet_tail_resultant_instrument.py`/`.json` (deterministic bootstrap, seed pinned), report **§5.12 + §5.13**, §9.2's steering row superseded **in place**, EPIPHANIES `E-THE-CONTROL-SCORED-THE-HEADLINE-1`. All 13 headline figures in THIS entry verified against the committed JSONs before landing (13/13 exact — the #927 lesson, now standard). diff --git a/probes/weather-p1/COMET_TAIL_REPORT.md b/probes/weather-p1/COMET_TAIL_REPORT.md index 1f7f05d8a..04a07ecd7 100644 --- a/probes/weather-p1/COMET_TAIL_REPORT.md +++ b/probes/weather-p1/COMET_TAIL_REPORT.md @@ -998,12 +998,28 @@ insensitivity. > - Carve D moved 0.943406 → 0.943403 (2.4e-06). > > **The deeper lesson, which is why the wording changed and not just the -> digits:** in-disk variance here is ~1e5 Pa², so a systematic offset of tens -> of Pa perturbs R² in the 5th decimal. **R² is structurally near-blind to -> exactly the defect that matters for an encoder**, and "lossless" was inferred -> from the one statistic that could not detect the loss. The probe now reports -> **RMSE and mean bias in Pa alongside every R²**, because those are what -> distinguish the carves. +> digits:** in-disk variance here is ~1e5 Pa², so a systematic offset of a few +> Pa perturbs R² in the 6th decimal. **R² is near-blind, in the single-digit-Pa +> regime, to exactly the defect that matters for an encoder** — and that is the +> regime the "lossless" claim was made in, inferred from the one statistic that +> could not detect the loss at that magnitude. The probe now reports **RMSE and +> mean bias in Pa alongside every R²**, because those are what distinguish the +> carves. +> +> > **⚠ CORRECTION 2026-08-12 (codex P2 on PR #927) — the first version of this +> > paragraph said "an offset of TENS of Pa perturbs R² in the 5th decimal", +> > which the table 15 lines above refutes.** The two halves belong to different +> > carves: **+92.76 Pa (carve A) moved R² 0.9212 → 0.9129 — 0.0083, the THIRD +> > decimal**, plainly visible; **+1.59 Pa (carve D) moved it 0.943406 → +> > 0.943403 — 2.4e-06, the SIXTH decimal.** Fusing carve A's magnitude with +> > carve D's insensitivity materially understated how well a *correct* R² +> > detects a large bias. The sharper and true statement: **the `var()` BUG was +> > blind at every magnitude — that is what hid +92.76 Pa — while the STATISTIC +> > is blind only in the single-digit-Pa regime, which is exactly where +> > "lossless" was claimed on four-decimal agreement.** Same failure mode as the +> > #930 relation error: two individually-correct numbers asserted of one pair +> > when each belongs to a different one. The doctrine (report RMSE + bias in Pa +> > beside every R²) is unchanged and is what the correction rests on. Three results worth more than the headline: