Skip to content

R2IL wave: the chain carrier wins, the vocabulary survives optimization, and the boundary that qualifies both - #1016

Merged
AdaWorldAPI merged 1 commit into
mainfrom
claude/wave-r2il-generalization
Aug 23, 2026
Merged

R2IL wave: the chain carrier wins, the vocabulary survives optimization, and the boundary that qualifies both#1016
AdaWorldAPI merged 1 commit into
mainfrom
claude/wave-r2il-generalization

Conversation

@AdaWorldAPI

Copy link
Copy Markdown
Owner

Autoattended 6-slice wave: 4 Sonnet probe workers, 1 Sonnet scribe, 1 Opus synthesis worker, 1 Haiku guarded executor, 1 read-only Opus meta-reviewer. Orchestrator compiled centrally, adjudicated every gate, applied every fix, and is the sole writer of all board files.

The named next measurement was blocked. Wider corpora need r2sleigh to lift new binaries and that sibling is absent from this checkout. Re-planned around what the existing corpus can actually answer: it holds two builds of one source at different optimization levels (71 fns / 3040 ops vs 72 / 2300, disjoint address keys) — a real held-out split.

Findings

F-10 — the def-use chain carrier confirms #1014's prescription as code. 1,872 length-3 chains collapse to 27 signatures against the window's 97 (3.6× tighter); top-10 occupancy 0.887 vs 0.505; and 95.9% of the top chain's occurrences skip at least one intervening op (median span 6, max 148) — invisible to a window matcher at any width the data would justify.

F-9 — the idiom vocabulary survives optimization completely. 10/10 top-K transfer; 61 of 64 TRAIN trigram types survive; TEST carries 94, of which 33 are TEST-ONLY. The optimizer added to the vocabulary while cutting ops/function 42.82 → 31.94. The partial-survival pre-registration was refuted and recorded in place.

F-11 — the boundary, and it qualifies the others. The convention's residual (36,747 count units) is of comparable-or-larger magnitude than its classified output (17,557 rows), 88.1% one named reason, conservation exact. Consequence: F-1, F-2 and F-10 are claims about a seven-opcode projection of two binaries, not about x86-64.

F-12 — the Stamp loss curve. 0 at N≤64, 33.3% at 96, 55.2% at 143, 87.5% at 512. Modelled wider registers recover 15/0 dropped at N=143 — modelled only, not a proposal, costs unmeasured. Any width change is the operator's ruling.

The meta-review changed the findings, not just the code

Verdict FIX-THEN-LAND, with 4 P0s — three gates that no input could fail (one of which labelled a true-by-construction identity "the falsifier for this whole section"), plus this arc's own knowledge-doc header stating F-1 as an unqualified claim about "real machine code" that the same document later contradicts. Also caught: an unchecked prov_op_site uniqueness assumption underpinning every span figure, and a gate awarding PASS with no assertion at all (an all-comment TSV would have passed it).

All P0s and the load-bearing P1s fixed; every gate re-run green with strictly harder assertions (top1*2 < total, distinct*10 < total, frac > 0.5, n > 64, and a shared reason_for join now exercised by both the loop and its can-fire).

Two findings were weakened by the review and are recorded weakened: F-10 drops "prescription confirmed" — the chain carrier's over-admission is 0 by construction, so what remains is an uncontrolled concentration comparison with no occurrence-matched control; F-11 gains an explicit units caveat (ore rows vs residual count units, never established equal).

Process notes

  • A worker caught an error in its own brief — the slag by_address section is 4 columns, not the 5 I specified — and parsed what the file declares rather than what it was told.
  • The Haiku executor stopped at command 1 on a formatting failure, wrote its receipt, attempted no fix, and escalated. Contract behaviour. Supervisor fixed and re-carded: 8/8 GREEN, including the deliberate corpus-absent guard (exit 2, the never-fabricate path proven live).
  • Anti-fabrication rule in every brief: workers cannot run anything, so they know no number — every gate had to be a relation with the value printed. All four complied.

Also in this PR


Generated by Claude Code

…on, and the boundary that qualifies both

Autoattended 6-slice wave (4 Sonnet probe workers, 1 Sonnet scribe, 1 Opus
synthesis worker, 1 Haiku guarded executor). The named next measurement —
wider corpora — was BLOCKED: the r2sleigh sibling is absent, so no new
binaries can be lifted. Re-planned around what the existing corpus can
answer: it holds two builds of ONE source at different optimization levels,
a real held-out split.

F-10 — the def-use chain carrier CONFIRMS #1014's prescription as code:
1,872 chains collapse to 27 signatures against the window's 97 (3.6x
tighter), top-10 occupancy 0.887 vs 0.505, and 95.9% of the top chain's
occurrences skip past adjacency (median span 6, max 148) — invisible to a
window matcher at any width the data justifies.

F-9 — the idiom vocabulary survives optimization COMPLETELY. 10/10 top-K
transfer; 61 of 64 TRAIN trigram types survive; TEST carries 94, of which
33 are TEST-ONLY. The optimizer ADDED to the vocabulary while cutting
ops/function 42.82 -> 31.94. The partial-survival pre-registration was
refuted and is recorded in place.

F-11 — THE BOUNDARY. The convention's residual (36,747 count units) is of
comparable or larger magnitude than its classified output (17,557 rows),
88.1% of it one named reason, conservation exact. Consequence: F-1, F-2 and
F-10 are claims about a SEVEN-OPCODE PROJECTION of two binaries, not about
x86-64.

F-12 — the Stamp loss curve: 0 at N<=64, 33.3% at 96, 55.2% at 143, 87.5%
at 512. Modelled wider registers recover 15/0 dropped at N=143 — modelled
only, not a proposal, costs unmeasured. The width ruling is the operator's.

Meta-review returned FIX-THEN-LAND with 4 P0s — three gates no input could
fail (one labelling a true-by-construction identity "the falsifier for this
whole section") plus a knowledge-doc header overclaiming F-1 beyond its
projection. All fixed; every gate re-run green with strictly harder
assertions (top1*2<total, distinct*10<total, frac>0.5, n>64, a shared
reason_for join exercised by both the loop and its can-fire). Two findings
were WEAKENED by the review and are recorded weakened: F-10 drops
"prescription confirmed" (over-admission is 0 by construction; no
occurrence-matched control run), F-11 gains an explicit units caveat.

Process notes: a worker caught an error in its own brief (the slag
by_address section is 4 columns, not the 5 I specified) and parsed what the
file declares. The Haiku executor STOPPED at command 1 on a formatting
failure, wrote its receipt, attempted no fix, escalated — contract
behaviour; supervisor fixed and re-carded to 8/8 GREEN including the
deliberate corpus-absent guard.

Board hygiene same-commit (orchestrator sole writer): EPIPHANIES
E-THE-SEVEN-OPCODE-PROJECTION-IS-NOT-X86-AND-THE-CHAIN-CARRIER-WINS-1,
INTEGRATION_PLANS, AGENT_LOG consolidating the executor receipt, the new
knowledge doc .claude/knowledge/r2il-behavioral-carrier.md (F-1..F-12), and
the PR_ARC gap-fill for #977..#1005 from the scribe's primary-source
reconstruction (#976 absent from history; #978/#979 landed as direct
commits; one figure discrepancy with #975 flagged unadjudicated).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KCGhDYoQBXs3poaR7sFuqp
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, you can upgrade your account or add credits to your account and enable them for code reviews in your settings.

@coderabbitai

coderabbitai Bot commented Aug 23, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: f0e35f20-187c-4217-b499-288fdcab3bbc


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cursor

cursor Bot commented Aug 23, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_d85861b9-97fc-4db2-9306-5605746094fb)

@AdaWorldAPI
AdaWorldAPI merged commit 6fc8507 into main Aug 23, 2026
8 checks passed
AdaWorldAPI pushed a commit that referenced this pull request Aug 23, 2026
#1016 was taken by another session's R2IL generalization wave, created and
merged while this one was in flight. The board entries named a number that
belongs to someone else's finding; fixed in place rather than appended,
because the entries are this PR's own and are not yet merged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ArVbbq3DsToBM7r79zGeEf
AdaWorldAPI pushed a commit that referenced this pull request Aug 23, 2026
…doc-ir

Self-correction inside this arc's own (unmerged) entries, in place rather than
appended, because they are this PR's and not yet merged.

1. The seam minted `source_id`/`span_id`. `ogar_doc_ir::DocIr` already answers
   all three questions a tokenization receipt asks — `content_sha256`,
   `(DocPage::number, Region::reading_order)`, `Region::text`. Re-cut to read
   them; the receipt now mints nothing. 41 gates, 18 disable-runs.

2. "The OCR boundary supplies no byte offsets" is RETIRED as a gap. It supplies
   no PAGE-wide offset and does not need to: a region owns its text, so an
   offset is region-local, and ogar-from-docv1::region_text is where the
   leading_space-aware join already lives. What remains is a sub-region span
   needing a non-zero byte_from, which the receipt already carries.

Also corrected: "247 of 255" was the count of ids APPEARING in the lane, not
the vocabulary size. Measured, the table is FULL at 255/255 on Alice and
180/255 on the KJV fixture, where the corpus rather than the cap set it — which
is what #1016's own record of that fixture says. The saturation conclusion
holds and is stronger; the number was the wrong quantity.

The banked patch is refreshed to both commits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ArVbbq3DsToBM7r79zGeEf
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants