From 33e2b6f114c0f25b0ccdc74b99e594a92233682a Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 12 Aug 2026 19:45:59 +0000 Subject: [PATCH 1/2] board: record merged #942 (arc entry + shipped row) #942 was mixed -- the CodeRabbit "mean"-qualifier fix AND the new substrate-comfort-zones-v1 plan -- so it gets its own arc entry (the termination clause covers hygiene-only PRs, not this). The entry records what a future session needs the "why" for: the plan states the operator's hypothesis as an INTERACTION term, not a main effect (worse-everywhere-but-least-worse-in-storms and better-only-in- storms are different findings), and D-CZ-0's three preflight corrections are this session's own lessons applied rather than re-learned -- sample-composition arithmetic before the first fetch, both controls smoke-tested for losability, units annotated at coefficient definition. Also records what is NOT done: the report SS10 reconciliation sweep. SS10.2's model is disconfirmed, SS10.5's "gather-design unblocked" is retracted, SS10.6's roadmap still shows pre-wave gates -- the corrections live in plan RUN sections and epiphanies, and the report itself is stale. Named in Deferred so it does not get lost. Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi --- .claude/board/LATEST_STATE.md | 1 + .claude/board/PR_ARC_INVENTORY.md | 43 +++++++++++++++++++++++++++++++ 2 files changed, 44 insertions(+) diff --git a/.claude/board/LATEST_STATE.md b/.claude/board/LATEST_STATE.md index 005ef2a75..6d9401bab 100644 --- a/.claude/board/LATEST_STATE.md +++ b/.claude/board/LATEST_STATE.md @@ -855,6 +855,7 @@ Membrane consumers can now pull BOTH halves of a render `classid` BBB-safely fro | PR | Merged | Title | What it added | |---|---|---|---| | *gap note* | — | **#781–#925 are NOT in this table** — carried by the dated sections above + `PR_ARC_INVENTORY.md`. Recorded 2026-08-12 (codex P2 on #930) rather than silently reconstructed; the table had stalled at #780. | — | +| **#942** | 2026-08-12 | "mean" qualifier fix + `substrate-comfort-zones-v1` (calibration × regime, made falsifiable) | MIXED. Plan states the operator's hypothesis as an INTERACTION (a poorly-calibrated adaptive substrate does relatively better in strong/turbulent storms than in calm regimes) across 3 held-constant regimes × formula × calibration quality — "worse everywhere but least-worse in storms" ≠ "better only in storms", and the plan distinguishes them. D-CZ-0 preflight DONE with 3 corrections applied from this session's own lessons (sample-composition arithmetic before first fetch; both controls smoke-tested for losability; units annotated at coefficient definition). D-CZ-1..6 Queued. Also fixed CodeRabbit's Minor on #941 — #940's 25 %/9 % figures are MEANS over 19 storms, qualifier restored (4th instance this week of scope-lost-in-summary). | | **#940** | 2026-08-12 | W6 lands — vector-sum dipole model VOID by its own anti-vacuity control; stranded stratum empty by CT-F14's filter arithmetic; same-PR sign/units correction round | B0 VOID: single-geo R²=−0.104 (worse than the mean); both controls (permuted P_bow, P_bow rotated 90°) clear the `≤single-geo+0.03` ceiling. B3 stranded (`\|v_storm\|<8 m/s`): n=0 — `displacement_km≥250`/6h implies `\|v_storm\|≥11.57 m/s` for every admitted storm, `E-THE-DISPLACEMENT-FILTER-ATE-THE-STRANDED-STRATUM-1`. Same-PR: codex+CodeRabbit caught a sign-convention bug (`D=-spine(...)`, matching `low_pole_bearing()`'s own flip; verified offline that this leaves R²/B0/B1/B3 unchanged, confirmed on the actual re-run) and a units error (`c_bow` is km⁻¹ not dimensionless; replaced with the dimensionally valid `\|c_bow·P_bow\|` vs `\|D\|` metric — mean-over-19-storms: geo≈25%, bow≈9% of mean `\|D\|`). CT-F17's gate now moot for this model form. | | **#938** | 2026-08-12 | W5 v2 RUN lands — B2 REVERSES to genuine FAIL, B3's VOID CONFIRMED at full headline scale, both link families now verified | B2: real diffusion resolved (raw rel-L2 vs unsmoothed input = 0.190), operator's own anisotropy = **1.5251 vs the 1.25 bar** (baseline through the clean 3.35σ mask = 1.0046 — the operator alone contributes ~0.52). `domino.rs` gather-design claim REFUTED at this test point. B3: 99.68 % (family A) / 99.56 % (family B) of the QUALIFYING population's control links land on a pure Fibonacci offset (n_qualifying=4 782 017, out of the headline lattice N=7 651 227 — not a 62k sub-sample) — dominated by the two discovered strides 2584=F(18) / 4181=F(19) respectively. Same-PR fixes for 3 more codex findings: family-B histogram was previously uncomputed (now measured + JSON patched); "four orders of magnitude" corrected to ~1.9 (76.9×); B4 downgraded from verdict to explicit descriptive reading (n=19 was dropped without pre-authorization — B4 stays open). | | **#936** | 2026-08-12 | W5 v1 RUN (B2 PASS / B3 VOID / B4 smooth) — ⊘ SUPERSEDED same day, all three verdicts compromised; v2 CODE fix merged in this PR, v2 RUN/results in flight | Codex found 4 real defects: B3 control subsampled at 250k/band (~74 % self-linked at headline, ratio 0.9996 uninformative); B2's fit floor = input σ (8 iters ≈ inert, "resolved" indistinguishable from "untouched"); bump only 1.72σ from the mask edge (analytic truncation ratio 1.2082 ≈ the measured "1.213 asymptote" — likely the MASK, not the operator); 99.38 % Fibonacci-membership figure was chat-only. Fixed same day, code merged HERE (`106ca605`, verified an ancestor of this PR's merge commit — corrects two overclaims caught by codex P2 on PR #937: the fix was described as "landing in a follow-up PR" when it had already landed, and the histogram artifact as "committed" when only the SCRIPT that produces it had landed — the tracked JSON is still v1's until the v2 run completes): full-band control, V-matched iteration scaling, bump moved to 3.35σ clearance + baseline-through-mask computed, offset histogram now computed by the script. The epiphany's MECHANISM survives (sub-cap 99.38 % measurement was never subsampled) but its "does not exist" phrasing was also softened (codex P2 on #937: 99.38 % ≠ 100 %, one N/one construction ≠ universal proof) — only the headline evidentiary number is retracted; the claim is restated at its actual evidentiary scope, not deleted. | diff --git a/.claude/board/PR_ARC_INVENTORY.md b/.claude/board/PR_ARC_INVENTORY.md index 9feaf873d..f4eb672f7 100644 --- a/.claude/board/PR_ARC_INVENTORY.md +++ b/.claude/board/PR_ARC_INVENTORY.md @@ -1,3 +1,46 @@ +## 2026-08-12 — lance-graph #942 (MERGED) — the "mean" qualifier fix + `substrate-comfort-zones-v1` (the operator's calibration-vs-regime hypothesis, made falsifiable) + +- **Added.** `.claude/plans/substrate-comfort-zones-v1.md` — the operator's + hypothesis stated as a testable interaction: *a poorly-calibrated + substrate that adapts dynamically should do RELATIVELY better in strong, + turbulent storms than in calm regimes.* Three held-constant regimes + (open water / flatland / storm-with-high-shear) × substrate formula × + calibration quality. The claim is an **interaction term**, not a main + effect — "worse everywhere but least-worse in storms" and "better only + in storms" are different findings and the plan distinguishes them. + D-CZ-0 (preflight) ran and is **DONE**; D-CZ-1..6 Queued. +- **Locked — D-CZ-0's three preflight corrections, found before any run.** + Applying this session's own hard-won rules to the new plan rather than + re-learning them: (1) the sample-composition check (the + `E-THE-DISPLACEMENT-FILTER-ATE-THE-STRANDED-STRATUM-1` lesson) run + BEFORE the first fetch — a regime split must be arithmetically reachable + in the source data, not discovered empty after the run; (2) both + controls smoke-tested for LOSABILITY at negligible cost before any + expensive arm (the W2s-a degenerate-grid and W5 locality-is-Fibonacci + lessons: a control that cannot lose or cannot differ carries zero + information when it "passes"); (3) every coefficient annotated with its + units at definition, and magnitude claims restricted to dimensionally + consistent contributions (the #940 `c_bow`-is-km⁻¹ lesson). +- **Fixed — CodeRabbit Minor on #941**: #940's summary rows dropped the + **mean** qualifier the source RUN entry and the committed JSON's + `fitted_contribution_Pa_per_km` keys carry (means over 19 storms). + PR_ARC's merged entry got an appended dated correction (append-only); + LATEST_STATE's living row was fixed in place (precedent: #939's + N-vs-n_qualifying fix). **Fourth instance this week of one defect + class** — a qualifier or operand pairing true in the source, lost in + the summary; the figure is never wrong, its scope goes unstated. +- **Deferred.** Every D-CZ probe (operator go-ahead pending). The report + §10 reconciliation sweep (§10.2's model is disconfirmed, §10.5's + "gather-design unblocked" is retracted, §10.6's roadmap still shows the + pre-wave gates — the corrections live in plan RUN sections and + epiphanies, and the report itself has not yet been brought current). +- **Confidence.** High on the qualifier fix (verified against the + committed JSON's own keys). The plan is a SPECIFICATION — nothing in + it has run beyond D-CZ-0's preflight checks. + +**Status:** MERGED (`61a60618`). Branch `claude/jirak-math-theorems-harvest-rfii13` +→ `main`. Plan + board — zero product code, zero probe runs. + ## 2026-08-12 — lance-graph #940 (MERGED) — W6 lands: VOID by its own control, sample-composition lesson, and a same-PR sign/units correction round - **Added.** `comet_tail_w6.py`/`.json` — full detail already recorded in From 5313fddffa9ff7ec35a681d7036907be4c07df05 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 12 Aug 2026 19:49:11 +0000 Subject: [PATCH 2/2] board: fix 2 codex P2 on #943 -- regime count (4 not 3) and a prescribed-vs-done control claim Both real, both verified against the sources before fixing: 1. REGIME COUNT. The adopted ladder in substrate-comfort-zones-v1.md has FOUR regimes, not three: R1 CALM Amazon (|grad p| 10.2), R2 OCEAN S-Pacific gyre (14.9), R3 ACTIVE W-Siberian lowland (43.8), R4 STORM (95.6) -- the preflight DELIBERATELY split the operator's "flatland" into a calm and an active tier, and the plan's output contract budgets for four. Both summary rows said three, so a session executing from the board would have dropped either the calm-land or the active-flatland tier. Fixed in both, with the actual gradient ladder named so the count is checkable rather than asserted. 2. PRESCRIBED vs DONE. Both rows read as if the control-losability smoke test had already been performed, while STATUS_BOARD correctly marks D-CZ-1 as Queued and the same entry elsewhere said nothing beyond D-CZ-0 ran. A future session could have skipped the gate. Fixed by splitting D-CZ-0's bullet into what it ACTUALLY EXECUTED (sample-composition arithmetic -- which is what forced the 4-way split; units-at-definition discipline) versus what it PRESCRIBED for runs that have not happened (D-CZ-1's control smoke test, Queued, gating every later cell). Stated explicitly: the requirement existing in the plan is not the requirement being satisfied. Both are the same defect class this session keeps hitting -- a summary losing a distinction the source carries. Fifth instance; the mechanical counter-check (re-read the source, count/verify, don't paraphrase from memory) remains the lever. Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi --- .claude/board/LATEST_STATE.md | 2 +- .claude/board/PR_ARC_INVENTORY.md | 45 ++++++++++++++++++------------- 2 files changed, 28 insertions(+), 19 deletions(-) diff --git a/.claude/board/LATEST_STATE.md b/.claude/board/LATEST_STATE.md index 6d9401bab..75c5510e4 100644 --- a/.claude/board/LATEST_STATE.md +++ b/.claude/board/LATEST_STATE.md @@ -855,7 +855,7 @@ Membrane consumers can now pull BOTH halves of a render `classid` BBB-safely fro | PR | Merged | Title | What it added | |---|---|---|---| | *gap note* | — | **#781–#925 are NOT in this table** — carried by the dated sections above + `PR_ARC_INVENTORY.md`. Recorded 2026-08-12 (codex P2 on #930) rather than silently reconstructed; the table had stalled at #780. | — | -| **#942** | 2026-08-12 | "mean" qualifier fix + `substrate-comfort-zones-v1` (calibration × regime, made falsifiable) | MIXED. Plan states the operator's hypothesis as an INTERACTION (a poorly-calibrated adaptive substrate does relatively better in strong/turbulent storms than in calm regimes) across 3 held-constant regimes × formula × calibration quality — "worse everywhere but least-worse in storms" ≠ "better only in storms", and the plan distinguishes them. D-CZ-0 preflight DONE with 3 corrections applied from this session's own lessons (sample-composition arithmetic before first fetch; both controls smoke-tested for losability; units annotated at coefficient definition). D-CZ-1..6 Queued. Also fixed CodeRabbit's Minor on #941 — #940's 25 %/9 % figures are MEANS over 19 storms, qualifier restored (4th instance this week of scope-lost-in-summary). | +| **#942** | 2026-08-12 | "mean" qualifier fix + `substrate-comfort-zones-v1` (calibration × regime, made falsifiable) | MIXED. Plan states the operator's hypothesis as an INTERACTION (a poorly-calibrated adaptive substrate does relatively better in strong/turbulent storms than in calm regimes) across **4** held-constant regimes × formula × calibration quality — the preflight SPLIT "flatland" into calm and active tiers: R1 CALM Amazon (`\|∇p\|` 10.2) → R2 OCEAN S-Pacific (14.9) → R3 ACTIVE W-Siberia (43.8) → R4 STORM (95.6), a 9.3× range. "Worse everywhere but least-worse in storms" ≠ "better only in storms", and the plan distinguishes them. D-CZ-0 preflight DONE (sample-composition arithmetic run before the first fetch — it is what forced the 4-way split; units annotated at coefficient definition). **D-CZ-1, the control-losability smoke test, is PRESCRIBED BY the plan but Queued — not yet run**, and gates every later cell. D-CZ-2..6 Queued. Also fixed CodeRabbit's Minor on #941 — #940's 25 %/9 % figures are MEANS over 19 storms, qualifier restored (4th instance this week of scope-lost-in-summary). | | **#940** | 2026-08-12 | W6 lands — vector-sum dipole model VOID by its own anti-vacuity control; stranded stratum empty by CT-F14's filter arithmetic; same-PR sign/units correction round | B0 VOID: single-geo R²=−0.104 (worse than the mean); both controls (permuted P_bow, P_bow rotated 90°) clear the `≤single-geo+0.03` ceiling. B3 stranded (`\|v_storm\|<8 m/s`): n=0 — `displacement_km≥250`/6h implies `\|v_storm\|≥11.57 m/s` for every admitted storm, `E-THE-DISPLACEMENT-FILTER-ATE-THE-STRANDED-STRATUM-1`. Same-PR: codex+CodeRabbit caught a sign-convention bug (`D=-spine(...)`, matching `low_pole_bearing()`'s own flip; verified offline that this leaves R²/B0/B1/B3 unchanged, confirmed on the actual re-run) and a units error (`c_bow` is km⁻¹ not dimensionless; replaced with the dimensionally valid `\|c_bow·P_bow\|` vs `\|D\|` metric — mean-over-19-storms: geo≈25%, bow≈9% of mean `\|D\|`). CT-F17's gate now moot for this model form. | | **#938** | 2026-08-12 | W5 v2 RUN lands — B2 REVERSES to genuine FAIL, B3's VOID CONFIRMED at full headline scale, both link families now verified | B2: real diffusion resolved (raw rel-L2 vs unsmoothed input = 0.190), operator's own anisotropy = **1.5251 vs the 1.25 bar** (baseline through the clean 3.35σ mask = 1.0046 — the operator alone contributes ~0.52). `domino.rs` gather-design claim REFUTED at this test point. B3: 99.68 % (family A) / 99.56 % (family B) of the QUALIFYING population's control links land on a pure Fibonacci offset (n_qualifying=4 782 017, out of the headline lattice N=7 651 227 — not a 62k sub-sample) — dominated by the two discovered strides 2584=F(18) / 4181=F(19) respectively. Same-PR fixes for 3 more codex findings: family-B histogram was previously uncomputed (now measured + JSON patched); "four orders of magnitude" corrected to ~1.9 (76.9×); B4 downgraded from verdict to explicit descriptive reading (n=19 was dropped without pre-authorization — B4 stays open). | | **#936** | 2026-08-12 | W5 v1 RUN (B2 PASS / B3 VOID / B4 smooth) — ⊘ SUPERSEDED same day, all three verdicts compromised; v2 CODE fix merged in this PR, v2 RUN/results in flight | Codex found 4 real defects: B3 control subsampled at 250k/band (~74 % self-linked at headline, ratio 0.9996 uninformative); B2's fit floor = input σ (8 iters ≈ inert, "resolved" indistinguishable from "untouched"); bump only 1.72σ from the mask edge (analytic truncation ratio 1.2082 ≈ the measured "1.213 asymptote" — likely the MASK, not the operator); 99.38 % Fibonacci-membership figure was chat-only. Fixed same day, code merged HERE (`106ca605`, verified an ancestor of this PR's merge commit — corrects two overclaims caught by codex P2 on PR #937: the fix was described as "landing in a follow-up PR" when it had already landed, and the histogram artifact as "committed" when only the SCRIPT that produces it had landed — the tracked JSON is still v1's until the v2 run completes): full-band control, V-matched iteration scaling, bump moved to 3.35σ clearance + baseline-through-mask computed, offset histogram now computed by the script. The epiphany's MECHANISM survives (sub-cap 99.38 % measurement was never subsampled) but its "does not exist" phrasing was also softened (codex P2 on #937: 99.38 % ≠ 100 %, one N/one construction ≠ universal proof) — only the headline evidentiary number is retracted; the claim is restated at its actual evidentiary scope, not deleted. | diff --git a/.claude/board/PR_ARC_INVENTORY.md b/.claude/board/PR_ARC_INVENTORY.md index f4eb672f7..19d265f0a 100644 --- a/.claude/board/PR_ARC_INVENTORY.md +++ b/.claude/board/PR_ARC_INVENTORY.md @@ -3,24 +3,33 @@ - **Added.** `.claude/plans/substrate-comfort-zones-v1.md` — the operator's hypothesis stated as a testable interaction: *a poorly-calibrated substrate that adapts dynamically should do RELATIVELY better in strong, - turbulent storms than in calm regimes.* Three held-constant regimes - (open water / flatland / storm-with-high-shear) × substrate formula × - calibration quality. The claim is an **interaction term**, not a main - effect — "worse everywhere but least-worse in storms" and "better only - in storms" are different findings and the plan distinguishes them. - D-CZ-0 (preflight) ran and is **DONE**; D-CZ-1..6 Queued. -- **Locked — D-CZ-0's three preflight corrections, found before any run.** - Applying this session's own hard-won rules to the new plan rather than - re-learning them: (1) the sample-composition check (the - `E-THE-DISPLACEMENT-FILTER-ATE-THE-STRANDED-STRATUM-1` lesson) run - BEFORE the first fetch — a regime split must be arithmetically reachable - in the source data, not discovered empty after the run; (2) both - controls smoke-tested for LOSABILITY at negligible cost before any - expensive arm (the W2s-a degenerate-grid and W5 locality-is-Fibonacci - lessons: a control that cannot lose or cannot differ carries zero - information when it "passes"); (3) every coefficient annotated with its - units at definition, and magnitude claims restricted to dimensionally - consistent contributions (the #940 `c_bow`-is-km⁻¹ lesson). + turbulent storms than in calm regimes.* **FOUR** held-constant regimes + × substrate formula × calibration quality — the preflight deliberately + SPLIT the operator's "flatland" into a calm and an active tier, so the + adopted ladder is **R1 CALM (Amazon basin, `|∇p|` 10.2) → R2 OCEAN + (S Pacific gyre, 14.9) → R3 ACTIVE (W Siberian lowland, 43.8) → R4 + STORM (CT-F14 centres, 95.6)**, a 9.3× dynamic range. The claim is an + **interaction term**, not a main effect — "worse everywhere but + least-worse in storms" and "better only in storms" are different + findings and the plan distinguishes them. D-CZ-0 (preflight) ran and is + **DONE**; D-CZ-1..6 Queued. +- **Locked — D-CZ-0's preflight corrections, found before any run; and + what it PRESCRIBED for the runs that have not happened yet.** Applying + this session's own hard-won rules to the new plan rather than + re-learning them. **Already executed in D-CZ-0:** the + sample-composition check (the + `E-THE-DISPLACEMENT-FILTER-ATE-THE-STRANDED-STRATUM-1` lesson) — the + regime ladder was verified arithmetically reachable in the source data + BEFORE the first fetch, and it is what forced the calm/active split; + plus the units discipline (every coefficient annotated at definition, + magnitude claims restricted to dimensionally consistent contributions — + the #940 `c_bow`-is-km⁻¹ lesson). **Prescribed but NOT yet run:** + **D-CZ-1 is the control-losability smoke test and is Queued** — the + W2s-a degenerate-grid and W5 locality-is-Fibonacci lessons are written + into the plan as a mandatory prerequisite gating every later cell, not + as work already done. A future session must run D-CZ-1 before any + expensive arm; the requirement existing in the plan is not the + requirement being satisfied. - **Fixed — CodeRabbit Minor on #941**: #940's summary rows dropped the **mean** qualifier the source RUN entry and the committed JSON's `fitted_contribution_Pa_per_km` keys carry (means over 19 storms).