the integration half of #1012: one receipt, three borrowed consumers - #1017
Conversation
`PROBE-TOKEN-SEAM-1` (37 gates, 13 disable-runs verified red-then-green) asks what `E-TOKEN-BPE-CAN-FIT-NOT-YET-BUY-1` explicitly did not: can ONE versioned BPE tokenization of a span serve several consumers at once? Measured yes — 313 source tokenizations for 308 spans plus 5 fixtures, and Tantivy, DeepNSM-v2 and a forward-prediction surface each added ZERO. The probe itself lives in AdaWorldAPI/paperless-rs (crates/paperless-token, docs/TOKEN-SEAM-ARCHITECTURE.md) because it needs a Tantivy dependency this workspace does not carry. This commit is the lance-graph-side record. Two facts make it adoptable and neither was designed for it: DeepNSM-v2's library is already tokenizer-free (`parse_to_spo(&[Tagged])` takes `(WordId, Pos)` and no string — the split_whitespace lives only in its examples), and Tantivy's indexer never reads offset_from/offset_to at all, so an index cannot become the ABI even by accident. The number that bounds prior work: the 8-bit id lane SATURATES — 247 of 255 ids on 75 KB of ordinary English, compression 3.18x -> 2.03x. #1012's 3.35x is a 1 KB figure and must not be read as corpus-independent. The hi-byte PAGE lane is untested and is the next probe. Four gaps named rather than worked around: no shipped token continuation mechanism (RailCarving::AxisSlab caps at 24 levels, under the measured p50 of 4 particles); ValueTenant has no token variant; there is no callable PoS surface because deepnsm_v2::lexicon was deliberately deleted and the insight_coca_read grounding cited for that is itself an example binary outside a lean consumer's dependency barrier; and the cam96 codebook/codes are absent, so the semantic half never ran. Polars refuted: zero occurrences across nine checkouts. There was nothing to remove. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ArVbbq3DsToBM7r79zGeEf
The probe that E-ONE-RECEIPT-MANY-BORROWED-CONSUMERS-1 cites lives in AdaWorldAPI/paperless-rs, whose push is denied at the org/App level (verified through the session proxy AND proxy-bypassed; the in-environment GH_TOKEN is a 14-character placeholder, so the usual "a 403 here is the proxy" escape does not apply). The commit exists only in an ephemeral container. Banked as the documented plateau pattern: the full patch minus the two corpus fixtures, which are 350 KB and exactly reproducible by two recorded commands. SHA256s for all three files are in the README so a reconstruction that differs by one byte is caught rather than silently changing every measured number. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ArVbbq3DsToBM7r79zGeEf
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_063382cf-9de5-49ff-a8e2-3e60b017c3b2) |
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
#1016 was taken by another session's R2IL generalization wave, created and merged while this one was in flight. The board entries named a number that belongs to someone else's finding; fixed in place rather than appended, because the entries are this PR's own and are not yet merged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ArVbbq3DsToBM7r79zGeEf
…on-architecture-3xd4eh # Conflicts: # .claude/board/PR_ARC_INVENTORY.md
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ArVbbq3DsToBM7r79zGeEf
…doc-ir Self-correction inside this arc's own (unmerged) entries, in place rather than appended, because they are this PR's and not yet merged. 1. The seam minted `source_id`/`span_id`. `ogar_doc_ir::DocIr` already answers all three questions a tokenization receipt asks — `content_sha256`, `(DocPage::number, Region::reading_order)`, `Region::text`. Re-cut to read them; the receipt now mints nothing. 41 gates, 18 disable-runs. 2. "The OCR boundary supplies no byte offsets" is RETIRED as a gap. It supplies no PAGE-wide offset and does not need to: a region owns its text, so an offset is region-local, and ogar-from-docv1::region_text is where the leading_space-aware join already lives. What remains is a sub-region span needing a non-zero byte_from, which the receipt already carries. Also corrected: "247 of 255" was the count of ids APPEARING in the lane, not the vocabulary size. Measured, the table is FULL at 255/255 on Alice and 180/255 on the KJV fixture, where the corpus rather than the cap set it — which is what #1016's own record of that fixture says. The saturation conclusion holds and is stronger; the number was the wrong quantity. The banked patch is refreshed to both commits. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ArVbbq3DsToBM7r79zGeEf
The banked patch carried two of the four commits on the paperless-rs branch. It now carries all four — the seam probe, the ogar-doc-ir re-cut, the two measurement probes, and the intake wiring that put the S-2 gate in front of both retinas. The insurance is only insurance if it works, so it was checked rather than assumed: `git am` of all four onto the base commit applies clean, and `cargo test -p paperless-intake -p paperless-kv` on the reconstructed tree is green. The README now says so, and says what the check does NOT cover — the token probe needs the two corpus fixtures the patch deliberately omits. Also recorded there: the doc-IR pair floats on `branch = "main"` rather than a rev because cargo does not unify a `branch` and a `rev` source even at the same commit, and two SourceIds mean two incompatible `DocIr` types; and `--features dom` does not compile because `spider_doc_ir` predates `TableCell::confidence`. A reconstruction that hits either should know it is upstream, not its own mistake. Push to AdaWorldAPI/paperless-rs is still 403 at the GitHub App level — re-tested this session, through the proxy. Nothing has changed there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ArVbbq3DsToBM7r79zGeEf
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_1733c764-990a-4e53-a545-838a662c899e) |
…ed verified The paperless-token code now lives in AdaWorldAPI/tesseract-rs as `crates/tesseract-paperless`, feature-gated, with CI on all three tiers. tesseract-rs accepts pushes; paperless-rs does not. So this directory stops being the way to obtain the code and becomes what it should have been called from the start: forensics. The README now says that up front, and says the other half too — do not sync anything back to paperless-rs, which is a dead copy. Two things the move fixed rather than carried: the `[patch]` section is gone entirely (tesseract-rs path-deps its siblings, so the escaping-relative-path trap does not arise), and `ingest_html` became a producer-agnostic `ingest_doc_ir(bytes, index, build)` — the caller brings its own producer as a closure, so the broken `spider_doc_ir` build is nobody's dependency now. Also corrected: the fixture table gave `coca_academic_20k.tsv` as 180 361 bytes. It is 226 651. The sha256 was right, which is why the reconstruction check passed and why nothing downstream was wrong — but the check I ran covered patch-applies and tests-pass, not the byte counts, and "verified" without that scope was a broader claim than the evidence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ArVbbq3DsToBM7r79zGeEf
PROBE-TOKEN-SEAM-1— 41 gates, 18 disable-runs each verified red-then-green — asks whatE-TOKEN-BPE-CAN-FIT-NOT-YET-BUY-1(#1012) explicitly did not: can ONE versioned BPE tokenization of a source span serve several consumers at once?Measured yes. One source tokenization per span — Tantivy, DeepNSM-v2 and a forward-prediction input surface each added zero. Query analysis is counted on a separate counter, because a query is different bytes and folding it into one number would make the claim a lie.
This PR is the board record for that work: the epiphany, the arc entry, the state table, and the archived patch.
The identity comes from
ogar-doc-ir— the receipt mints nothingThe first cut invented
source_id: u32/span_id: u32. That was a second population wearing the document layer's job;ogar_doc_ir::DocIralready answers all three questions a receipt asks, and the seam was re-cut to read them:ogar-doc-irsuppliesDocIr::content_sha256, interned once per documentRegion, at(DocPage::number, Region::reading_order)Region::textTwo consequences worth the PR:
content_sha256is a PER-ACQUISITION dedup key, not a cross-retina identity — the crate's own docs correct its plan's first sketch on exactly this point. For a tokenization receipt that is the right reading: you tokenize bytes, so different bytes are a different tokenization. Cross-retina convergence is a facts question (converges_on_facts) and is not this seam's business.ogar-from-docv1::region_textis where theleading_space-aware join already lives. What remains is far smaller — a sub-region span needs a non-zerobyte_from, which the receipt already carries and no producer emits.The seam is source-agnostic for free:
docir.rshas no line that knows which retina produced the IR.The two facts that make it adoptable, and neither was designed for it
parse_to_spo(&[Tagged])consumes(WordId, Pos)pairs and touches no string; thesplit_whitespace/normaliselogic lives ONLY in its two examples. The crate needed no change — and in the new home that shows up structurally:deepnsm-v2is a dependency of the PROBES and not ofsrc/.Token::textandToken::position, usesposition_lengthtransiently, and readsoffset_from/offset_tonowhere outside its own tests. An index cannot become the ABI even by accident.The number that bounds prior work
The 8-bit id table is FULL at 75 KB of ordinary English (and still full on the whole 170 KB file), while on the KJV fixture the corpus rather than the cap set it at 180 — which is what #1016's own record of that fixture says. The two corpora sit on opposite sides of that line. #1012's 3.35× is a 1 KB figure and must not be read as corpus-independent.
(An earlier revision of this body quoted "247 of 255" as the vocabulary. That was the count of ids appearing in the lane, not the table size. Corrected here and in the board entry.)
The follow-up probe named there has since RUN, and it inverts the headline: at the 256:256 rail (one token per
(u8:u8), two separate bytes, never a widenedu16)alice-fullcompresses 6.13× against the flat-u8 1.95×, uses 7 675 of 65 536 ids — corpus-bound with 88.3 % of the tile free, where the u8 table was CAP-bound — and its resident particles are 35.2 % smaller. So "saturates at 75 KB" was the cap's property, not the corpus's.Gaps, named rather than worked around
rail_geometry::RailCarving::AxisSlab { reg, cont }caps atRAIL_MAX_DEPTH = 24— below the measured p50 of 4 particles. Two framings are lawful and the trade is exact:particle_countalone bounds the run and a PAD scan inside it is already exact because PAD is reserved (cost: one vocabulary slot, which at a cap-bound table is not free); ortoken_countcosts 4 bytes and frees the slot.ValueTenanthas no token variant (16 discriminants, none for text).deepnsm_v2::lexiconwas deliberately deleted, and theinsight_coca_readgrounding cited for that deletion is itself an example binary in a crate carryingserde/serde_yml/tokio/ndarray, outside a lean consumer's dependency barrier. Recorded rather than re-litigated.cam96_codebook.bin/cam96_codes.binare absent, so palette256² distance is unexercised.DocIrfrom text, not from a retina. Half-closed — see below.Update 2026-08-24 — the move, and what it changed on purpose
Four commits landed before the code moved (the seam probe; the
ogar-doc-irre-cut; the two measurement probes; the intake wiring). The archived patch carries all four. In the new home three things are deliberately different, and each is a fix rather than a port:The whole
[patch]section is gone. tesseract-rs path-deps its siblings, so the escaping-relative-path trap does not arise — and with it goes the branch-vs-rev problem: cargo does not unify abranchand arevsource even at the identical commit (measured: twoogar-doc-irentries in the lock, same#719471dbsuffix), and two SourceIds mean two incompatibleDocIrtypes in one binary, which defeats a source-agnostic IR entirely.ingest_htmlbecameingest_doc_ir(bytes, index, build). The old form calledspider_doc_irdirectly, which bound intake to ONE crawler and inherited its build — and that build is broken (E0063: missing field 'confidence', a fieldogar-doc-iradded after spider's only commit; spider floats onbranch = "main", so it broke the moment the field landed and nothing builds the pair to notice). The new form takes the producer as a closure: any producer that can build aDocIris admitted and the crate depends on none of them. That is what a source-agnostic IR is FOR, and hard-coding one crawler had quietly spent the property. The closure — rather than a ready-madeDocIr— is what keeps S-2 true for a producer the crate has never heard of, because on a duplicate it is never called.One assertion became a runtime guard. The old code merely tested that the producer's hash matched the gate's.
IntakeError::IdentityMismatchnow refuses the mismatch, because a producer keying its IR by a normalised or re-encoded form of its input would leave the gate and the subtree silently addressing two different documents.Both new guards are disable-verified, 12 of 13 passing under each, so each is pinned by the test that names it rather than by collateral damage:
content_sha256 != hashcheckbuild()ahead of the short-circuitThe evidence reproduced in the new home, checked rather than trusted:
probe_token_seamreports ALL 41 GATES GREEN with numbers identical to the original run, and both measurement probes reproduce. 13 lib tests; all three tiers clippy-clean at-D warnings; fmt clean;tesseract-core27/27 unaffected.Gap 5 is half-closed. Intake now runs the gate in front of real producers and hands back one
DocIreither way. What is still open is the comparison that would actually test source-agnosticism end to end — the SAME content through two different retinas — and it is blocked on the upstream one-liner inspider, which this session cannot push.Worth carrying past this PR: a floating branch dep with no CI that exercises it is drift with no detector.
ogar-doc-ir's closedRegionKindvocabulary and version marker guard the DATA shape at load time; neither can see a struct field added at compile time. The new crate is a workspace MEMBER with optional deps for exactly this reason — an excluded crate is never compiled by CI, and an uncompiled crate rots invisibly.Polars
Refuted, and the honest form is weaker than the question invited: zero occurrences across nine checkouts; every
DataFramemention is prose; neither the intake tree nortesseract-rsdeclares arrow/datafusion/lance/lancedb at all. There was nothing to remove.Method notes
An independent vacuity audit of the finished probe found five holes, all real — a gate asserting byte counts but never span counts (the CRLF bug that collapsed 300 spans into 1 would have re-passed it), a threshold true by construction, an assertion about a type signature rather than behaviour, an unconditional prefix check, and an unexercised ASCII-vs-Unicode whitespace divergence. All fixed.
Three disable-runs were themselves wrong first. The sharpest is from the re-cut:
T-DOCIR-KEYcompared each receipt against the samespans()call it was validating — an implementation checked against itself — and stayed green whenspans()was changed to renumber by position. It now walks theDocIrindependently, and the fixture'sreading_orderis2i+1rather than the positional index so the two are distinguishable at all. A fourth disable (flattening table cells) found nothing because neither text corpus contains a table or a figure — a fixture-coverage gap, closed with a purpose-built three-region page.A knob that does not bind is not a disable; a fixture's SHAPE is part of a test's coverage; and an implementation compared against itself is not a check.
Board hygiene
EPIPHANIES(E-ONE-RECEIPT-MANY-BORROWED-CONSUMERS-1, with its in-entry self-correction),PR_ARC_INVENTORY,LATEST_STATE,AGENT_LOG, and the archived patch + its superseded notice.🤖 Generated with Claude Code
https://claude.ai/code/session_01ArVbbq3DsToBM7r79zGeEf