R2IL wave: the chain carrier wins, the vocabulary survives optimization, and the boundary that qualifies both - #1016
Conversation
…on, and the boundary that qualifies both Autoattended 6-slice wave (4 Sonnet probe workers, 1 Sonnet scribe, 1 Opus synthesis worker, 1 Haiku guarded executor). The named next measurement — wider corpora — was BLOCKED: the r2sleigh sibling is absent, so no new binaries can be lifted. Re-planned around what the existing corpus can answer: it holds two builds of ONE source at different optimization levels, a real held-out split. F-10 — the def-use chain carrier CONFIRMS #1014's prescription as code: 1,872 chains collapse to 27 signatures against the window's 97 (3.6x tighter), top-10 occupancy 0.887 vs 0.505, and 95.9% of the top chain's occurrences skip past adjacency (median span 6, max 148) — invisible to a window matcher at any width the data justifies. F-9 — the idiom vocabulary survives optimization COMPLETELY. 10/10 top-K transfer; 61 of 64 TRAIN trigram types survive; TEST carries 94, of which 33 are TEST-ONLY. The optimizer ADDED to the vocabulary while cutting ops/function 42.82 -> 31.94. The partial-survival pre-registration was refuted and is recorded in place. F-11 — THE BOUNDARY. The convention's residual (36,747 count units) is of comparable or larger magnitude than its classified output (17,557 rows), 88.1% of it one named reason, conservation exact. Consequence: F-1, F-2 and F-10 are claims about a SEVEN-OPCODE PROJECTION of two binaries, not about x86-64. F-12 — the Stamp loss curve: 0 at N<=64, 33.3% at 96, 55.2% at 143, 87.5% at 512. Modelled wider registers recover 15/0 dropped at N=143 — modelled only, not a proposal, costs unmeasured. The width ruling is the operator's. Meta-review returned FIX-THEN-LAND with 4 P0s — three gates no input could fail (one labelling a true-by-construction identity "the falsifier for this whole section") plus a knowledge-doc header overclaiming F-1 beyond its projection. All fixed; every gate re-run green with strictly harder assertions (top1*2<total, distinct*10<total, frac>0.5, n>64, a shared reason_for join exercised by both the loop and its can-fire). Two findings were WEAKENED by the review and are recorded weakened: F-10 drops "prescription confirmed" (over-admission is 0 by construction; no occurrence-matched control run), F-11 gains an explicit units caveat. Process notes: a worker caught an error in its own brief (the slag by_address section is 4 columns, not the 5 I specified) and parsed what the file declares. The Haiku executor STOPPED at command 1 on a formatting failure, wrote its receipt, attempted no fix, escalated — contract behaviour; supervisor fixed and re-carded to 8/8 GREEN including the deliberate corpus-absent guard. Board hygiene same-commit (orchestrator sole writer): EPIPHANIES E-THE-SEVEN-OPCODE-PROJECTION-IS-NOT-X86-AND-THE-CHAIN-CARRIER-WINS-1, INTEGRATION_PLANS, AGENT_LOG consolidating the executor receipt, the new knowledge doc .claude/knowledge/r2il-behavioral-carrier.md (F-1..F-12), and the PR_ARC gap-fill for #977..#1005 from the scribe's primary-source reconstruction (#976 absent from history; #978/#979 landed as direct commits; one figure discrepancy with #975 flagged unadjudicated). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KCGhDYoQBXs3poaR7sFuqp
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_d85861b9-97fc-4db2-9306-5605746094fb) |
#1016 was taken by another session's R2IL generalization wave, created and merged while this one was in flight. The board entries named a number that belongs to someone else's finding; fixed in place rather than appended, because the entries are this PR's own and are not yet merged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ArVbbq3DsToBM7r79zGeEf
…doc-ir Self-correction inside this arc's own (unmerged) entries, in place rather than appended, because they are this PR's and not yet merged. 1. The seam minted `source_id`/`span_id`. `ogar_doc_ir::DocIr` already answers all three questions a tokenization receipt asks — `content_sha256`, `(DocPage::number, Region::reading_order)`, `Region::text`. Re-cut to read them; the receipt now mints nothing. 41 gates, 18 disable-runs. 2. "The OCR boundary supplies no byte offsets" is RETIRED as a gap. It supplies no PAGE-wide offset and does not need to: a region owns its text, so an offset is region-local, and ogar-from-docv1::region_text is where the leading_space-aware join already lives. What remains is a sub-region span needing a non-zero byte_from, which the receipt already carries. Also corrected: "247 of 255" was the count of ids APPEARING in the lane, not the vocabulary size. Measured, the table is FULL at 255/255 on Alice and 180/255 on the KJV fixture, where the corpus rather than the cap set it — which is what #1016's own record of that fixture says. The saturation conclusion holds and is stronger; the number was the wrong quantity. The banked patch is refreshed to both commits. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ArVbbq3DsToBM7r79zGeEf
Autoattended 6-slice wave: 4 Sonnet probe workers, 1 Sonnet scribe, 1 Opus synthesis worker, 1 Haiku guarded executor, 1 read-only Opus meta-reviewer. Orchestrator compiled centrally, adjudicated every gate, applied every fix, and is the sole writer of all board files.
The named next measurement was blocked. Wider corpora need
r2sleighto lift new binaries and that sibling is absent from this checkout. Re-planned around what the existing corpus can actually answer: it holds two builds of one source at different optimization levels (71 fns / 3040 ops vs 72 / 2300, disjoint address keys) — a real held-out split.Findings
F-10 — the def-use chain carrier confirms #1014's prescription as code. 1,872 length-3 chains collapse to 27 signatures against the window's 97 (3.6× tighter); top-10 occupancy 0.887 vs 0.505; and 95.9% of the top chain's occurrences skip at least one intervening op (median span 6, max 148) — invisible to a window matcher at any width the data would justify.
F-9 — the idiom vocabulary survives optimization completely. 10/10 top-K transfer; 61 of 64 TRAIN trigram types survive; TEST carries 94, of which 33 are TEST-ONLY. The optimizer added to the vocabulary while cutting ops/function 42.82 → 31.94. The partial-survival pre-registration was refuted and recorded in place.
F-11 — the boundary, and it qualifies the others. The convention's residual (36,747 count units) is of comparable-or-larger magnitude than its classified output (17,557 rows), 88.1% one named reason, conservation exact. Consequence: F-1, F-2 and F-10 are claims about a seven-opcode projection of two binaries, not about x86-64.
F-12 — the Stamp loss curve. 0 at N≤64, 33.3% at 96, 55.2% at 143, 87.5% at 512. Modelled wider registers recover 15/0 dropped at N=143 — modelled only, not a proposal, costs unmeasured. Any width change is the operator's ruling.
The meta-review changed the findings, not just the code
Verdict FIX-THEN-LAND, with 4 P0s — three gates that no input could fail (one of which labelled a true-by-construction identity "the falsifier for this whole section"), plus this arc's own knowledge-doc header stating F-1 as an unqualified claim about "real machine code" that the same document later contradicts. Also caught: an unchecked
prov_op_siteuniqueness assumption underpinning every span figure, and a gate awarding PASS with no assertion at all (an all-comment TSV would have passed it).All P0s and the load-bearing P1s fixed; every gate re-run green with strictly harder assertions (
top1*2 < total,distinct*10 < total,frac > 0.5,n > 64, and a sharedreason_forjoin now exercised by both the loop and its can-fire).Two findings were weakened by the review and are recorded weakened: F-10 drops "prescription confirmed" — the chain carrier's over-admission is 0 by construction, so what remains is an uncontrolled concentration comparison with no occurrence-matched control; F-11 gains an explicit units caveat (ore rows vs residual count units, never established equal).
Process notes
by_addresssection is 4 columns, not the 5 I specified — and parsed what the file declares rather than what it was told.Also in this PR
.claude/knowledge/r2il-behavioral-carrier.md— the consolidated reference (F-1..F-12) so the next session loads it instead of rediscovering the arc.Generated by Claude Code