Skip to content

board: record merged #944 (arc entry + shipped row) - #945

Merged
AdaWorldAPI merged 1 commit into
mainfrom
claude/jirak-math-theorems-harvest-rfii13
Aug 12, 2026
Merged

board: record merged #944 (arc entry + shipped row)#945
AdaWorldAPI merged 1 commit into
mainfrom
claude/jirak-math-theorems-harvest-rfii13

Conversation

@AdaWorldAPI

Copy link
Copy Markdown
Owner

Board hygiene for #944.

#944 was MIXED, not pure hygiene — it landed a plan revision, an epiphany, and a review-caught correction — so the termination clause does not apply and it owes both obligations of the merged-PR row.

obligation file done
arc entry PR_ARC_INVENTORY.md PREPEND, suffix-verified (+5150 chars)
shipped row LATEST_STATE.md inserted above #942

What the arc entry carries (the "why" a future session would otherwise re-derive):

  • the cross-swap inversion and the premise that forces it — the model captures the phenomenon but is not calibrated, so miscalibration is the measurement condition and RMSE under a deliberately wrong calibration is bad by definition;
  • the collapse of CAL-ABS-OWN / CAL-ABS-FOREIGN into one arm's diagonal and off-diagonal (splitting them was the tell the design was still a race);
  • C2's both-halves requirement — the dynamic arms' L ≡ 0 must hold exactly and CAL-ABS must be proven non-degenerate through the same code path, the can-it-DIFFER gate W5 already paid for;
  • the two new constancy bars (C1b measures constancy instead of claiming it; C1c gives the suitability assumption a falsifier, and a null there VOIDS rather than weakens);
  • the pre-registered L < 0 case — a donor's wider range can legitimately beat an outlier-set diagonal, so the consequence is stated before it can be explained away;
  • the CodeRabbit catch: §2.3 still named the weak form as the hypothesis after C4 was sharpened — the fifth sibling-claim inconsistency in this arc, and one a reviewer found rather than the author.

Confidence: [H] on the methodological finding (airtight given the premise; the premise itself is untested by this plan), [G] on nothing — #944 measured nothing.

Recorded in both homes: D-CZ-1, the control-losability smoke test, is still Queued and still gates every later cell. #944 did not move it.

🤖 Generated with Claude Code

https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi


Generated by Claude Code

#944 was MIXED, not pure hygiene -- it landed a plan revision, an epiphany,
and a review-caught correction -- so it owes both obligations of the
merged-PR row. The termination clause does not apply.

Arc entry carries the why a future session would need: the cross-swap
inversion and the premise that forces it, the CAL-ABS-OWN/FOREIGN collapse
into one arm's diagonal and off-diagonal, C2's both-halves requirement, the
two new constancy bars, the pre-registered negative-L case, and the
CodeRabbit catch (the summary passage still naming the weak form as the
hypothesis after C4 was sharpened -- fifth sibling-claim inconsistency this
arc). Confidence [H] on the methodological finding, [G] on nothing: this PR
measured nothing.

D-CZ-1 recorded as still Queued in both homes -- it gates every later cell
and this PR did not move it.

PR_ARC_INVENTORY prepend suffix-verified.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi
@coderabbitai

coderabbitai Bot commented Aug 12, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AdaWorldAPI, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 53 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 06eb42f9-1676-4560-af33-c11a115caf23

📥 Commits

Reviewing files that changed from the base of the PR and between ac5d95f and c57c076.

📒 Files selected for processing (2)
  • .claude/board/LATEST_STATE.md
  • .claude/board/PR_ARC_INVENTORY.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cursor

cursor Bot commented Aug 12, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_3a12d592-63c7-4e8f-a81d-e3259b0b622d)

Copy link
Copy Markdown
Owner Author

Bot findings on this PR: none actionable — Cursor Bugbot hit a usage limit, CodeRabbit skips drafts (and was rate-limited earlier today). Not re-triggering it for two board files.

Instead, self-audited the arc entry against the merged #944 range, because this arc's recurring defect is precisely a summary claim that does not survive checking against its source — five instances so far, the most recent one caught by a reviewer rather than by me.

Every checkable claim in the entry verified against 4304590c..ac5d95fd:

claim result
zero code, zero probe runs, zero fetches ✅ 4 files, all .claude/*.md
the fix landed in 1756c518 ✅ present in range (3 commits)
4×4 donor × target matrix
primary metric Spearman ρ
L[D][T] = ρ[T][T] − ρ[D][T] ✅ verbatim
C1b separation ≥ 3
regime ladder 10.2 / 14.9 / 43.8 / 95.6, 9.3× ✅ all five figures
C2 both halves mandatory
L < 0 pre-registered
§2.3 scores against the diagonal
foreign donor = weak form, "wiring, not merit"
D-CZ-0 the only DONE row DONE: [D-CZ-0]
D-CZ-1 still Queued Queued: [D-CZ-1, 2, 2b, 2c, 3, 4, 5, 6, 7]

No discrepancies. The DONE/Queued split was read off the board programmatically rather than from memory — that specific pair (prescribed-vs-done) was the codex P2 on #943, so it is the one most worth not asserting.


Generated by Claude Code

@AdaWorldAPI
AdaWorldAPI marked this pull request as ready for review August 12, 2026 20:28
@AdaWorldAPI
AdaWorldAPI merged commit 65c52cc into main Aug 12, 2026
2 checks passed
AdaWorldAPI pushed a commit that referenced this pull request Aug 12, 2026
Two findings, one run (probes/weather-p1/substrate_comfort_d_cz_0_1.py).

1. D-CZ-0 was marked DONE with NO committed script or JSON. Its figures were
quoted in the plan, the arc entry, the LATEST_STATE row and three PR bodies.
Worse: the #945 self-audit -- run specifically because this arc keeps
shipping unbacked summary claims -- reported those figures as "verified"
while only comparing the arc entry to the PLAN. Both are prose. A figure
cited in two documents is cited twice, not confirmed once.

Reproduced now: 4 of 9 rows (ratios 1.004/1.022/0.994/0.931). The five
EXCLUDED land candidates are unreproducible -- their box centres were never
recorded anywhere -- and no coordinates were invented to fake them. The
|grad p| definition was never committed either, so four candidates were
computed and the winner decided from data: Pa per grid cell with NO cos(lat)
metric (max dev 0.069; next-best 0.398). That understates high latitude by
1/cos(lat), so R3 is ~40% low; metric-corrected the ladder is
10.3/15.5/61.2/100.9 -- ORDER survives, range widens 9.3x -> 9.8x.

Two defects in the reproduction caught by its own guards: measuring all 19
storms at the preflight timestep instead of each storm's own t0 (inverted
R3/R4 and would have read as C1 failing), and a seam assertion firing on a
real storm at 353.4 E where a plain slice gives a one-column box.

2. D-CZ-1 PASSES. Both controls lose to both real arms on both metrics in
all four regimes; GEO-DEGENERATE saturates 92-97%, so the mechanism is
visible rather than assumed. C1b separation = 6.28 vs the >= 3 bar.

3. ...and the same run AMENDS C4. rho is SATURATED on the diagonal --
real-arm spread 3e-6..4.7e-5 -- so C4 as pre-registered could not have
fired. rho keeps four orders of magnitude of range OFF the diagonal, so L
keeps it; C4 moves to RMSE in Pa with rho as a floor check. Legitimate
because D-CZ-1's whole purpose is to test the apparatus before the expensive
cells and no C4 cell has been scored; illegitimate the instant one has.
Recorded as an amendment carrying its trigger, and propagated to the header
and the C4 bar rather than left contradicting them.

A HINT, explicitly not a result: CAL-ABS beats CAL-RANK on RMSE everywhere
and its margin shrinks as the field gets more active (3.96 -> 1.85 -> 1.14),
R4 breaking the monotone at 1.49. One timestep, no control on that
comparison, and diagonal-vs-diagonal is not what the hypothesis is about.
Recorded so a later run cannot present it as a confirmation.

EPIPHANIES: E-A-FIGURE-CITED-TWICE-IS-NOT-CONFIRMED-ONCE-1 and
E-THE-METRIC-THAT-SEPARATES-ONE-COMPARISON-IS-BLIND-TO-ANOTHER-1 (prepend,
suffix-verified). STATUS_BOARD D-CZ-0/D-CZ-1 updated.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi
AdaWorldAPI pushed a commit that referenced this pull request Aug 12, 2026
#946 was MIXED -- probe code, real measurements, two epiphanies and a
synthesis document -- so it owes both obligations of the merged-PR row.

The arc entry carries what a future session would otherwise re-derive: that
D-CZ-0 was marked DONE with no artifact behind it and that #945's own audit
compared prose to prose; the gradient definition identified FROM DATA as
Pa-per-cell-without-cos(lat) and its bounded consequence (order survives,
magnitudes do not); D-CZ-1's gate passing 19/19 after the
sample-composition fix; the C4 amendment WITH its legitimacy scope (valid
only while no C4 cell has been scored); the disable-verified selftest
including the tie-averaging break that flips a sign silently; and the
exploratory correlation recorded with its not-a-result status string so a
later run cannot restate it as a confirmation.

Also records the formula matrix: 46 rated primitives, 20 D and 3 V, i.e.
half the inventory is a negative result -- with the C tier explained (a
comfort zone is a map, not a grade) and the D/V distinction stated (D lost
its test; V distinguished nothing).

Confidence split recorded honestly: [G] on the gate, the reproduction, the
definition identification and the rho-saturation measurement; [H] on the C4
amendment's scope; and the two exploratory items explicitly NOT results.

PR_ARC_INVENTORY prepend suffix-verified.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CcpLeEC3XK8Eye53GKBVvi
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants