diff --git a/.claude/board/AGENT_LOG.md b/.claude/board/AGENT_LOG.md index 5f115814..85e43ba1 100644 --- a/.claude/board/AGENT_LOG.md +++ b/.claude/board/AGENT_LOG.md @@ -1,3 +1,19 @@ +## 2026-08-19 — RP-SEAL first pass: 15 researchers (5 domains × builder/adversary/scout; orchestrator-consolidated) + +15/15 completed, 0 errors, 0 empty results (~5.59M subagent tokens, 842 tool +uses, run `wf_ca974718-1b4`). Adversary cells Opus, builder/scout Sonnet, per +the model-policy split. First pass fully independent: no cross-talk, blind to +`docs/lotus/**` and all board files; every brief carried the no-cargo + +tag-file-only rules — 0 cargo invocations measured across all transcripts; +this entry is the orchestrator consolidating, per the one-writer rule. +Mid-run operator scope pivot (rustynum struck / symbiont deprecated) applied +as a consolidation hard filter: audit found zero findings sourced from +either; nothing re-run. One agent death (A3, permission denial on the +sandboxed workflow driver) self-healed on resume with full cache reuse. +Outcome: `docs/lotus/RP-SEAL-CONSOLIDATION-PASS1.md` (matrix M1–M15, +cross-domain graph, DELETED list, Tier 0–4 work order) + Appendix H +(`docs/lotus/rp-seal-v1/`, 16 reports incl. the orchestrator pre-pass). + ## 2026-08-13 — weather-poc W0/W1 fleet (4 Sonnet workers, disjoint files; orchestrator-consolidated) 4/4 completed, 0 errors. Plan `.claude/plans/weather-soa-bake-v1.md`; the first diff --git a/.claude/board/EPIPHANIES.md b/.claude/board/EPIPHANIES.md index 5d5e6ecb..5e0135ec 100644 --- a/.claude/board/EPIPHANIES.md +++ b/.claude/board/EPIPHANIES.md @@ -1,3 +1,85 @@ +## 2026-08-19 — E-REPLAY-IS-CANONICAL-COMPACTION-IS-ECONOMICS-1 + +**Status:** RULING `[operator]` (STORNO of pass-1 consolidation framing). + +**STORNO'd:** the RP-SEAL consolidation's framing that physical-layout +work gates the cognitive/meta arc. Compaction is NOT required for the +target architecture. The D2 defect, restated narrower and more serious: +**PHYSICAL ITERATION ORDER MUST NOT DEFINE SEMANTIC REPLAY ORDER.** Fix +by canonical replay coordinates/order — never by physical reordering or +compaction. Separately: permanent temporal replay requires a +retention/tombstone policy under which historical knowledge states +remain reconstructible indefinitely. + +**The desired model (verbatim):** physical placement may remain +arbitrary · physical fragmentation may remain arbitrary · tombstoned +state remains historically addressable · semantic replay order is +canonical · current visibility is a projection · historical visibility +is a QueryReference projection. Compaction, if ever enabled, is only +optional storage economics — neither semantic repair nor a prerequisite +for cognition. + +**Applied:** consolidation doc §0 (`docs/lotus/RP-SEAL-CONSOLIDATION- +PASS1.md`) with ⊘-marked in-place corrections; Tier 1 reordered +(canonical coordinates №1, retention/tombstone №2, QueryReference +projection fixes №3); all layout/compaction items reclassified optional +economics; F-NOCOMPACT promoted from soak test to the default operating +description. Corrects the gating framing in +`E-RP-SEAL-PASS1-THE-MAXIM-WORKED-1` (its findings stand; their +consequence inverts — M2/M12 are reasons compaction stays off). + +## 2026-08-19 — E-RP-SEAL-PASS1-THE-MAXIM-WORKED-1 + +**Status:** FINDING (consolidated), first independent pass complete. + +**The event.** The RP-SEAL 15-researcher independent pass (run +`wf_ca974718-1b4`, 15/15 cells, 0 errors, no cross-talk, blind to boards) +completed and was consolidated per charter §12–§13: +`docs/lotus/RP-SEAL-CONSOLIDATION-PASS1.md`, with all 15 reports + the +orchestrator pre-pass committed as Appendix H under `docs/lotus/rp-seal-v1/`. + +**What died (charter §17 executed):** the phase/comma coefficient schedule +(C2 adversary: 0/6 conditions, syndrome separation impossible under +permutation; C1+C3 independently bound cascades ≥33% overhead / priced in +distance — three disjoint routes, one kill); unconditional Morton/SFC +default (B2 measured + B3 literature: identical page partition to Hilbert +at 4^j pages, 20–45× quantizer hot-spots, three proven negative results); +"compaction = one manifest publication" (reserve_fragment_ids is itself a +commit); the naive query+repair one-grouping hope (Morton locality +MANUFACTURES the correlated failure that kills locality-aligned parity). + +**What was found:** the coalescing path writes MORE bytes than no +coalescing — physical amplification = b+1 (65× at b=64) from 512 B null +padding per landing row, and 52% of seal CPU is a non-amortizing FNV hash +(E2); 11 arrival/physical-order leaks into durable coordinates, worst +L2: durable semantic replay order IS Lance physical scan order, so +compaction here MUTATES semantics (D2) — this gates all layout work; +five shipped seal paths with defective/false-accept behavior wearing +strong names (C2 §B1–B5, firefly independently by C3); Lance 9.0.0 +already ships the whole compaction/remap/stable-row-id/per-row-version +apparatus unused by our append-only path (5 cells concordant). + +**What survived, sharpened:** boring hash + row/column P+Q over the native +64×64 (6.25%, 63× repair amplification vs flat-RS 4095×) as the seal +baseline; the caller-supplied compaction-rewrite ordering key as the ONE +open precedented upstream seam (in-repo R-tree Hilbert leaves + Delta +Z-order as precedent; IndexRemapper usable today); two genuine +NOVEL-CANDIDATES with no prior art found by independent searches — the +joint query-scatter × repair-scatter address-order question (refined by +C2's anti-synergy into "query-aligned order, ANTI-aligned parity groups") +and the STRICT/AWARE/RETRO reader-rung admission tier (D3: no literature +counterpart; D1: currently inert in code, T_now missing as a type, +hlc_tick not actually HLC). + +**Scope-pivot filter:** PASSED — zero findings sourced from rustynum or +symbiont; the two mentions cite the exclusion ruling, and D1 converts it +into a finding (the version→kanban OUT-bridge lost its only caller). + +**Tier 0 before anything else:** pin D2's leaks (F-PHYS-ORDER first), +re-verify E2, X-C2-1 injection harness over the five seal paths, E3's +locality-debt trigger metric, perf-event or strike the cache-miss metric. +Full matrix M1–M15 + tiering in the consolidation doc. + ## 2026-08-18 — E-PIN-LANCE9-LANCEDB033-DF541-ARROW58-NO-DF53-1 **Status:** RULING `[operator]` + same-day discharge, measured. diff --git a/.claude/board/STATUS_BOARD.md b/.claude/board/STATUS_BOARD.md index edcc1ab2..3c9d941c 100644 --- a/.claude/board/STATUS_BOARD.md +++ b/.claude/board/STATUS_BOARD.md @@ -1,4 +1,4 @@ -## erasure-seals-compaction-research-v1 — RP-SEAL (operator charter, DISPATCHED 2026-08-18) +## erasure-seals-compaction-research-v1 — RP-SEAL (operator charter; PASS 1 CONSOLIDATED 2026-08-19 + §0 STORNO: canonical replay coordinates core, compaction = optional economics → docs/lotus/RP-SEAL-CONSOLIDATION-PASS1.md; Tier 0 probes next) Plan: `.claude/plans/erasure-seals-compaction-research-v1.md` (the operator's 15-researcher program; supersedes lotus deliverables 3-9's sequencing). diff --git a/.claude/plans/erasure-seals-compaction-research-v1.md b/.claude/plans/erasure-seals-compaction-research-v1.md index b7142c2d..4e64279e 100644 --- a/.claude/plans/erasure-seals-compaction-research-v1.md +++ b/.claude/plans/erasure-seals-compaction-research-v1.md @@ -3,6 +3,16 @@ > **Status: DISPATCHED 2026-08-18** — independent pass running as a 15-agent > background workflow (run `wf_ca974718-1b4`; reports land under the session > scratchpad `rp-seal-v1/`, committed at consolidation as Appendix H). +> +> **PASS 1 COMPLETE + CONSOLIDATED (2026-08-19).** 15/15 cells, 0 errors. +> Consolidation: `docs/lotus/RP-SEAL-CONSOLIDATION-PASS1.md` (M1–M15, +> DELETED list, Tier 0–4). Appendix H committed: `docs/lotus/rp-seal-v1/`. +> The comma schedule is DELETED per §17; Tier 0 probes are the next phase. +> **§0 STORNO (operator, same day):** canonical replay coordinates are +> the core fix (physical iteration order must not define semantic replay +> order); retention/tombstone for indefinite historical reconstructibility +> is core; compaction = optional storage economics, never semantic repair, +> never a cognition prerequisite. See the consolidation doc §0. > Tier allocation per the operator's "Opus for filigrane, Sonnet for > grindwork": the five ADVERSARY cells (A2/B2/C2/D2/E2) = strongest tier; > the five BUILDER and five SCOUT cells = grindwork tier. Consolidation diff --git a/docs/lotus/RP-SEAL-CONSOLIDATION-PASS1.md b/docs/lotus/RP-SEAL-CONSOLIDATION-PASS1.md new file mode 100644 index 00000000..9b72c5c8 --- /dev/null +++ b/docs/lotus/RP-SEAL-CONSOLIDATION-PASS1.md @@ -0,0 +1,493 @@ +# RP-SEAL Consolidation — First Independent Pass (2026-08-19) + +> Consolidation of the 15-researcher independent pass mandated by +> `.claude/plans/erasure-seals-compaction-research-v1.md` (§12–§13). +> Run `wf_ca974718-1b4`: 15/15 cells completed, 0 errors, 0 empty results, +> ~5.59M subagent tokens, 842 tool uses. Full per-cell reports committed as +> **Appendix H** under `docs/lotus/rp-seal-v1/` (A1–E3 plus +> `A0-orchestrator-prepass.md`, the orchestrator's private pre-dispatch +> cross-check, unseen by any researcher). +> +> **Independence discipline held.** First pass ran with no cross-talk, blind +> to `docs/lotus/**` and all board files. Builder/adversary/scout roles per +> domain. **Scope-pivot filter applied and PASSED**: the operator's mid-run +> rulings (rustynum struck — everything ndarray; `crates/symbiont` +> deprecated) bind this consolidation as a hard filter; audit of all 15 +> reports found zero findings *sourced from* either — the only two mentions +> (D1, C3) cite the exclusion ruling itself, and D1 converts it into a +> finding (see M9). Nothing required re-anchoring or re-running. +> **Model-tier discipline held**: 0 cargo invocations by any researcher. +> +> Charter maxim, §17: *do not make the beautiful idea true; make it survive +> fifteen people trying to make it false.* This document records what +> survived, what died, and what was refined into a sharper question. + +--- + +## 0. Operator ruling (STORNO, 2026-08-19) — the target architecture + +> Issued immediately after pass-1 consolidation landed; this section +> CORRECTS the consolidation's framing and outranks any conflicting +> sentence below. Storno discipline: the corrected sentences are marked +> ⊘ in place, not deleted. + +**STORNO'd:** the framing that physical-layout work (reordering, +compaction) gates the cognitive/meta arc. **Compaction is NOT required +for the target architecture.** The D2 defect is narrower and more +serious than the consolidation's framing: + +> **PHYSICAL ITERATION ORDER MUST NOT DEFINE SEMANTIC REPLAY ORDER.** + +The fix is **canonical replay coordinates/order** — never physical +reordering, never compaction. Separately, **permanent temporal replay +requires a retention/tombstone policy under which historical knowledge +states remain reconstructible indefinitely.** + +The desired model (operator, verbatim): + +- physical placement may remain arbitrary +- physical fragmentation may remain arbitrary +- tombstoned state remains historically addressable +- semantic replay order is canonical +- current visibility is a projection +- historical visibility is a QueryReference projection + +Compaction, if ever enabled, is only optional storage economics. It is +neither semantic repair nor a prerequisite for cognition. + +**Consequences applied below:** + +1. Tier 1's durable-order-key item is THE core fix and moves to №1, + restated as *canonical replay coordinates*. F-PHYS-ORDER becomes the + probe of the ruling's invariant itself. +2. A NEW core Tier 1 item: the retention/tombstone policy for indefinite + historical reconstructibility — absorbing M14 (cleanup currently + deletes exactly the manifests the temporal horizons resolve to; under + this ruling, cleanup must never delete state needed for historical + addressability) and elevating the QueryReference defects (L6, L8) from + hygiene to core: QueryReference IS the historical-visibility mechanism. +3. M3's layout-key seam and every Tier 2/3 layout/compaction item are + reclassified **optional storage economics** — real, precedented, and + non-gating. E3's locality-debt metric moves to that track. F-NOCOMPACT + is promoted from soak test to the DEFAULT operating description. +4. M2 and M12 remain true as findings; their consequence is inverted: + they are reasons compaction stays OFF, not obstacles to clear so it + can turn on. + +--- + +## 1. Evidence matrix + +Columns per charter §12. "Independent?" marks genuine independent +rediscovery (high-value) vs single-source claims. Cells cite their own +file:line evidence in the Appendix H reports; only the consolidated claim +and its strongest anchors are repeated here. + +### M1 — Lance 9.0.0 already ships the apparatus; the consumer uses none of it + +**Claim.** Deferred remap via Fragment-Reuse-Index, rank-based +O(#fragments) row-address remap (`RowAddrRemap`), stable row IDs, +size-bounded incremental compaction, `IndexRemapper` (caller-receivable +old→new remap), and per-row `created_at_version` / `last_updated_at_version` +provenance (RLE, compaction-correct) are all public and un-feature-gated in +the pinned 9.0.0. lance-graph's production write path +(`LanceCycleWriter`/`cycle_sink.rs`) is bootstrap-create-then-forever-append +and invokes none of it; grep for the machinery across `crates/` returns +zero non-vendored hits. `excluded_fragment_ids` / `max_source_rows` / +`max_source_bytes` are upstream-main-only (11.0.0-beta.14). +**Support:** A1 (two independent runs, concordant), A2, A3, E1, E3 — five +cells, independent. **Dissent:** none. **Confidence:** HIGH (settled +archaeology). **Falsifier:** none needed; source-anchored. +**Next:** consume, don't rebuild — see Tier table. + +### M2 — Locality-repair-by-compaction is structurally constrained, twice over + +**Claim.** (a) Compaction at both versions cannot reorder rows; under +`IndexRemapMode::Compact` order preservation is a *correctness contract* +(ascending-id guard, `row_addr_remap.rs:35-45,139-160`), so +locality-repairing compaction is prohibited in the cheap mode — the escapes +are Direct's O(rows) coordinator RAM or stable row IDs. (b) Stable row IDs +hard-conflict with `defer_index_remap` (thrown `Err` by design) and delete +the FragReuseGroup permutation witness. The two obvious remap-cost levers +cannot be composed, and the witness is anti-correlated with the cheap +configuration. +**Support:** A2 (primary), A1 (the defer/stable conflict, independent). +**Dissent:** none. **Confidence:** HIGH. +**Falsifier:** attempt a reordering compaction under Compact remap → guard +rejection. **Next:** any layout work routes through fragment *rewrite with +an ordering key* (M3), never through remap. + +### M3 — The one open, precedented upstream seam: a compaction-time layout key + +**Claim.** A caller-supplied physical-layout ordering key applied at +compaction-rewrite time is the genuinely open seam: Lance already ships the +mechanism scoped to one index (R-tree Hilbert-sorted leaf pages), Delta +ships the base-table analogue (Z-order), and `IndexRemapper` is a shipped +inversion-of-control trait that lets an external consumer receive the exact +old→new remap on every compaction commit today — a free first "permutation +witness" experiment with zero upstream changes. The secondary-index remap +cost that motivates all of this traces to one documented upstream gap: +indices key on physical row addresses instead of the stable row IDs Lance +already has. +**Support:** A3 (primary), E3 (LSM mapping, independent), B1 (zero-cost +sort-by-own-key arm, independent). **Dissent:** none. **Confidence:** HIGH +for the seam's existence; the *benefit* is unmeasured (see M10). +**Falsifier:** F-AMPLIFY. **Next:** Tier 2 IndexRemapper experiment before +any Tier 3 RFC. + +### M4 — The SFC default died; the question moved + +**Claim.** (a) Measured: at every 4^j page size this architecture uses, +Morton and Hilbert induce the IDENTICAL page partition (they differ only in +page order), so pages-touched is identical for every query — the +Morton-vs-Hilbert debate is moot at power-of-4 granularity. (b) Measured: +the linear (non-rank) quantizer produces 20–45× petal hot-spots and 70–84% +empty fields on clustered identities; `morton_slot` over a hash is a bit +permutation, not a locality code. (c) Literature: an unbounded clustering +gap for near-full-extent queries, the Arrwwid d≥3 disadvantage for the +whole cube-recursive family, and a proven impossibility of one curve +serving mixed query shapes. (d) Measured against expectation: Hilbert's +mean code-span is 18–22% WORSE than Morton on square queries even where +its median is better; Z-order beats Hilbert 2.0× on dyadic-parity +selections. +**Support:** B2 (measurement) + B3 (literature) — independent routes to the +same kill; B1 concurs (the project's own non-interleaved tiered key beat a +real Morton variant on the range-query metric in the one prior in-repo +experiment). **Dissent:** none on the kill; B2 retains Morton as *minimax +default* only. **Confidence:** HIGH. +**Falsifier:** F-SFC is now partially discharged — the surviving form is +query-mix- and dimensionality-conditioned selection, with the quantizer +(rank vs linear), not the curve, as the first-order variable. +**Next:** B1's pre-registered arms (own-key-bytes sort first; true +bit-level Morton, Hilbert, spectral as open arms) under E1's harness. + +### M5 — The comma/phase coefficient schedule is DELETED + +**Claim.** The project-specific phase/comma coefficient schedule fails its +own charter gate: 0 of 6 necessary conditions non-trivially met; +cross-level syndrome separation is impossible under permutation (an +XOR-fold seal is permutation-invariant, so completion-order scrambling — +which the architecture explicitly declares non-semantic — erases exactly +the structure the schedule would need). Locus+version-bound per-chunk hash ++ flat MDS RS strictly dominates cascade/product codes on every detection +axis in the C2 truth table. +**Support:** C2 (adversary, primary). C1 independently bounds cascade +parity below at ≥33% overhead (Gopalan–Huang–Simitci–Yekhanin locality +bound) vs 6.25% for row/column P+Q; C3 independently shows hierarchy is +purchased with distance (generalized Singleton, hierarchical VIII.5, +availability bounds), never free. Three cells, three routes, one verdict. +**Dissent:** none. **Confidence:** HIGH. +**Falsifier:** X-C2-2 (the DELETE gate) remains pre-registered if anyone +wants the corpse re-examined; the burden is now on the schedule. +**Next:** nothing — per charter §17 the schedule is removed from the +program. The seal baseline is C1's recommendation (M6). + +### M6 — The seal baseline: shipped hash + row/column P+Q over the native 64×64 + +**Claim.** The low-risk seal is the already-shipped content hash (~0.07% +overhead) plus row/column P+Q parity over the native 64×64 field grid: +6.25% overhead, 63× repair amplification vs flat-RS's 4095×. Repair- +bandwidth optimizations layered on an UNMODIFIED MDS seal +(Guruswami–Wootters trace repair, piggybacking) are proven and free; +hierarchical/cascade structure is a priced escalation, not a default. The +GF(2^8) SIMD idiom needed for RS constant-multiply (nibble-split + PSHUFB ++ XOR) already ships in `ndarray/src/simd_avx2.rs` (Harley–Seal popcount), +per the everything-ndarray ruling. One program-level question gates real +deployment: whether the target object store already erasure-codes one +layer down (double-coding wastes the overhead). +**Support:** C1 (primary), C3 (bounds, independent), C2 (flat-RS dominance, +independent). **Dissent:** none. **Confidence:** HIGH for the +recommendation shape; MEDIUM for the numbers until X-C2-1/X-C2-3 run. +**Falsifier:** F-RS + X-C2-3. **Next:** X-C2-1 injection harness (the +prerequisite), then X-C2-3. + +### M7 — Query-locality × repair-locality: anti-synergy found, novel question refined + +**Claim.** The naive cross-domain hope — one grouping serving both query +locality and repair locality — is an ANTI-synergy: Morton-style locality +manufactures exactly the correlated failure that destroys locality-aligned +parity groups (C2). Independently, C3 and E1 both searched and found NO +prior art posing the joint "one address order minimizing both query +scatter and repair scatter" question. The refined open question is +therefore: an address order serving query locality while parity groups +deliberately ANTI-align. This is the program's strongest Tier 4 candidate. +**Support:** C2 (anti-synergy), C3 + E1 (novelty, independent searches). +**Dissent:** none — the three compose rather than conflict. +**Confidence:** HIGH that the naive version is dead; the refined question +is open by construction. **Falsifier:** the E1 harness's repair-scatter +axis vs query-scatter axis, jointly. **Next:** only after Tier 0/1. + +### M8 — The coalescing/seal implementation inverts its own headline (E2) + +**Claim.** Amortization is real at the durable boundary — one fsync/cycle +pays above C≈100 casts — and false everywhere else in the current +implementation: (a) the coalescing path writes MORE payload-column bytes +than no coalescing at all, (1+C+D)·512 > C·512 always, with physical +amplification equal to b+1 — 65× at the repo's own b=64 — caused by 512 B +of null padding per landing row; (b) 52% of seal CPU is a byte-at-a-time +FNV-1a hash that is invariant under coalescing (47.5 ms at b=1, 4, 64 +alike) and structurally forbids an incremental variant; (c) the parity +GRANULE, not the seal granule, sets the incremental-vs-batch crossover +(≈1.46% dirty-row density for 8 KiB units at the stated geometry). +**Support:** E2 (adversary, measured on the real code). **Dissent:** none; +no other cell touched these paths. **Confidence:** HIGH (measured), single +source — re-verification is cheap and required before the fix lands. +**Falsifier/next:** Tier 0 re-run of E2's measurements, then Tier 1 fixes: +sparse landing rows (kill the b+1 padding term) and a blockwise, +incrementally-composable seal hash (keeping fail-closed semantics). + +### M9 — The temporal ladder: Layer 1 real, Layer 2 inert, one genuine novelty + +**Claim.** `temporal.rs` Layer 1 (causal deinterlacing / crash recovery) is +production-wired and tested. Layer 2 (STRICT/AWARE/RETRO epistemic +projection) has exactly one non-test exerciser, whose own source admits the +interesting axes are inert; `T_now` (the hypothesis's third coordinate) +exists nowhere as a type; `hlc_tick` is HLC-named but not HLC-implemented +(no physical component, no merge rule). Lance itself already stamps the +W_write/A_last primitives per row (M1), so the missing coordinates need no +external stamping layer for the single-writer case. The version→kanban +OUT-bridge (`LanceVersionScheduler`) is tested library code whose only +non-test caller was in operator-deprecated `symbiont` — it now has zero +live callers (D1, converting the scope ruling into a wiring fact). +Literature: the reader-rung-selectable STRICT/AWARE/RETRO admission tier +has no counterpart in the surveyed systems (XTDB's current/history split +is binary; classical isolation is per-transaction) — the program's second +Tier 4 candidate. The retroactive-structures bound Θ(min(√m, n·log m)) is +a forward falsifier only (no retroactive-write mechanism exists). +**Support:** D1 (code), D3 (literature) — independent; A1 (upstream +primitives). **Dissent:** none. **Confidence:** HIGH. +**Next:** Tier 1 — introduce `T_now` as a type and a real HLC merge rule, +or explicitly rule both out; re-home or delete the caller-less OUT-bridge. + +### M10 — Random placement is cheaper than assumed; irreversibility is the real cost + +**Claim.** E2's own adversarial cache-tiling thesis was self-DISPROVED: +random identity placement costs ~nothing at the design's 32 MiB geometry +(0.97× vs sequential) and only 2.30× at 512 MiB. B1: today's arrival order +is architecturally closer to a random control than to any spatial order. +E3: vendor cost warnings (Snowflake/Iceberg/Delta) triangulate that +unconditional locality-key ordering is known-costly maintenance. A2: the +real cost of random placement is IRREVERSIBILITY (M2), not read latency. +Consequence: physical-layout work needs a measured read-pattern +justification at real geometry before any engineering; E3's +entropy/scatter "locality debt" compaction trigger — the one genuine +literature gap in the LSM mapping — is the cheap first experiment. +**Support:** E2, B1, E3, A2 — four cells, independent, convergent. +**Dissent:** none. **Confidence:** HIGH. **Falsifier:** E3's pre-registered +locality-debt experiment; F-NOCOMPACT. **Next:** Tier 0. + +### M11 — The integrity surface has defective seals wearing strong names + +**Claim.** Five shipped seal/integrity paths, three defective in +false-accept terms (C2 §B1–B5): `firefly_frame.rs` advertises +"Reed-Solomon ECC" while `compute_ecc` is a 4-way XOR fold with zero +detection capability (independently found by C3); `merkle_tree.rs` (the +in-tree multi-scale syndrome) carries four defects including a +100%-false-accept path wearing a BCH decoder's name; `unified_audit.rs` +chains FNV-1a as tamper evidence (trivially forgeable — though its +callcenter threat model is drift, not adversaries); `seal.rs` is 48-bit +truncated BLAKE3 with an alpha-mask caveat; `persist_sink.rs`'s cycle +seal is FNV-1a-based (see M8's CPU finding). Hash-only seals false-accept +on exactly the faults this system generates (wrong-slot/wrong-version +writes of valid bytes) — locus+version binding is the fix, not wider +hashes. +**Support:** C2 (primary), C3 (firefly, independent). **Dissent:** none. +**Confidence:** HIGH (source-anchored). **Next:** Tier 0 — fix or +truthfully re-document each path; X-C2-1 harness pins false-accept rates. + +### M12 — Eleven order leaks into durable coordinates; L2 escalates first + +**Claim.** D2's leak table (report §"The leak table", L1–L11): 11 +arrival/physical-order leaks into durable coordinates, 10 unpinned. The +escalation is **L2**: durable semantic replay order IS Lance physical scan +order (`cycle_sink.rs:971-980` `scan_in_order(true)` with no `order_by`; +`persist_sink.rs:665-666` "no sort is done here") — so compaction here is +a semantic-order-MUTATING operation, not locality repair. Close behind: +L3 (same-row coalescing resolved by arrival under a non-injective +`row_of`), L6 (`QueryReference::default()` = `u64::MAX` admits future rows +even under Strict), L8 (deinterlace sort key mixes HLC ticks and Lance +versions as commensurable). The one pinned leak (batch_hash, F-ORD-REAL) +is the safest member of its family — it fails closed. +**Support:** D2 (primary); A2 independently establishes the +compaction-order contract that makes L2 unfixable-by-accident (M2). +**Dissent:** none. **Confidence:** HIGH (file:line-anchored). +**Falsifiers (pre-registered in D2):** F-PHYS-ORDER (highest value), +F-ORD-2, F-RESTART, F-HINDSIGHT family. **Next:** ⊘ corrected by §0 — +the fix is canonical replay coordinates; compaction/layout is optional +economics that simply stays off. The original "gates all physical-layout +work" framing survives only in the trivial sense that the optional track +also benefits. + +### M13 — Measurement infrastructure: 80% exists; two hazards recorded + +**Claim.** The harness family is specifiable today from shipped primitives +(`measure_wal_curve.rs`, 3,558 lines, plus Lance's own compaction/remap +machinery); the one real gap is hardware cache-miss counting (perf-event / +iai-callgrind, neither wired). Hazards: (a) the 512-byte witness payload +sits between the cloud (1000 B) and local (10 B) early-materialization +thresholds, so file:// benchmarks INVERT production S3 behavior — every +probe must run both storage schemes; (b) charter §14's "do not optimize a +metric that was not measured" currently strikes cache misses until the +counter lands. +**Support:** E1 (primary), A2 (threshold hazard, independent). +**Dissent:** none. **Confidence:** HIGH. +**Next:** Tier 0 — wire perf-event or strike the metric; add the +dual-scheme rule to every probe design. + +### M14 — Cleanup and temporal horizons are one resource (design coupling) + +**Claim.** `cleanup.rs:498-514` deletes exactly the manifests that +checkout_version-based STRICT/AWARE/RETRO horizons resolve to: +stale-version retention bytes are pinned by the cognitive layer's temporal +window, not by an operational retention policy. Retention must be derived +from the temporal window. Related seed: `FragReuseIndexDetails.Group. +changed_row_addrs` already records which rows changed per version, so a +retained per-version parity delta can triage stale-vs-corrupt (C2). +**Support:** A2, C2 — independent halves of one design coupling. +**Dissent:** none. **Confidence:** HIGH for the coupling; the parity-delta +triage is a design seed, unmeasured. **Next:** Tier 1 design note in the +temporal work (M9). + +### M15 — Orchestrator pre-pass reconciliation (disagreement record) + +The pre-pass (`A0-orchestrator-prepass.md`) claimed "prepared fragments → +one manifest publication" via `execute_uncommitted` + `CommitBuilder:: +execute_batch`. A2 found compaction's own path commits twice +(`reserve_fragment_ids` is itself a `commit_transaction`; an abort between +the two permanently advances `max_fragment_id`). Both are true: the +two-phase INSERT path exists exactly as the pre-pass reports; the +COMPACTION path does not get the same shape. The prepared-artifact +publication doc (D-LOTUS-6) stands for inserts; any compaction design must +budget two versions per pass. No other pre-pass claim conflicted with the +fleet; the fleet exceeded the pre-pass everywhere else. + +--- + +## 2. Cross-domain relationship graph + +Edges marked ⊕ are independent rediscoveries (high-value per §12); +⊗ marks composition of findings that never met in one cell. + +- **M2 ⊗ M12 (A2 × D2):** two independent reasons the current substrate + blocks layout work — compaction cannot reorder (correctness contract), + and if it could, it would mutate semantic replay order (L2). ⊘ §0 reframes + the consequence: these are reasons compaction stays OFF (its default), + and the L2 fix is canonical replay coordinates regardless of whether + any layout work ever happens. +- **M5 ⊕ (C1 × C2 × C3):** three independent routes (truth table; locality + bound; distance bounds) to the same kill of hierarchical/cascade coding + as a default. Genuine independent confirmation. +- **M7 ⊗ (C2 × C3 × E1):** the novelty (no joint prior art) survives, but + C2's anti-synergy inverts its naive form — the composition is sharper + than any single cell's claim. +- **M4 ⊕ (B2 × B3 × B1):** measurement, literature, and the one prior + in-repo experiment all against an unconditional SFC default. +- **M1 ⊕ (A1 × A1' × A2 × A3 × E1 × E3):** the double-run of A1 (resume + artifact) plus four other cells — concordant archaeology; treated as one + claim with unusually deep support, not six discoveries. +- **M9 ⊗ (D1 × D3 × A1):** code inertness + literature novelty + upstream + per-row provenance compose into the temporal work order: the novel tier + is worth building precisely because the primitives are already stored. +- **M8 ⊗ M6 (E2 × C1):** any parity overhead percentage is meaningless + until the b+1 padding amplification is fixed — 6.25% parity on top of + 65× padding is noise. E2's fix precedes C1's seal. +- **B1 ⊕ E3 ⊕ B3 (convergent shape):** the recursive-bisection → + HEEL/HIP/TWIG address shape was independently built four times in this + workspace (B1), matches the packed-memory-array amortized bound lineage + (E3), and is the same recursive-block decomposition as cache-oblivious + layouts (B3). The workspace keeps rebuilding one structure; naming it + once is a documentation deliverable, not new engineering. + +Shared-ancestry caution: all five E/A cells read the same Lance source +tree, so their agreement on M1 is expected (same primary source), unlike +the M5 triple, whose three routes are methodologically disjoint. + +--- + +## 3. Tiering (charter §13) + +**DELETED (survived nothing):** +- The phase/comma coefficient schedule (M5). C2's X-C2-2 remains + pre-registered should anyone want to re-litigate; the default is gone. +- Unconditional Morton (or any SFC) as a layout default (M4). Survives + only as a query-mix-conditioned choice with a rank quantizer. +- "Compaction = one manifest publication" (M15). Two commits, budgeted. +- The naive "one grouping serves query AND repair locality" (M7). + +**TIER 0 — measurement/falsifiers only (gates everything):** +1. Pin D2's unpinned leaks: F-PHYS-ORDER (L2, highest value), F-ORD-2 + (L3), F-RESTART (L1/L5/L9), F-HINDSIGHT family (L6-L11 subset). (M12) +2. Re-verify E2's three measurements (b+1 amplification, FNV CPU share, + parity-granule crossover) on the committed tree. (M8) +3. X-C2-1 injection harness; pin false-accept rates for the five seal + paths of M11 (fix or re-document each). +4. E3's locality-debt (remap-entropy) trigger metric — the cheap + falsifier before ANY layout/movement engineering. (M10) +5. Wire perf-event cache-miss counting or strike the metric; adopt the + dual-storage-scheme rule for every probe. (M13) + +**TIER 1 — lance-graph internal experiments/fixes (order per §0):** +1. **Canonical replay coordinates/order** — the durable semantic order + key that fixes L2/L9 by construction; physical iteration order must + not define semantic replay order. NOT solved by reordering or + compaction. (M12, §0) +2. **Retention/tombstone policy for indefinite historical + reconstructibility** — tombstoned state stays historically + addressable; cleanup derives from this requirement, never from + operational defaults (absorbs M14; constrained by A2's + cleanup/horizon coupling). (§0, M14) +3. **Historical visibility as a QueryReference projection** — fix L6 + (default `u64::MAX` admits future rows under Strict) and L8 + (incommensurable sort key); `T_now` as a type + a real HLC merge + rule, or their explicit rejection; re-home or delete the caller-less + OUT-bridge. (M9, M12, §0) +4. Sparse landing rows — kill the b+1 null-padding amplification. (M8) +5. Blockwise, incrementally-composable seal hash with locus+version + binding (replaces FNV; keeps fail-closed semantics). (M8, M11) + +**TIER 2 — generic Lance prototypes (no upstream changes needed; +items 2–3 are the OPTIONAL storage-economics track per §0):** +1. `IndexRemapper`-consumer permutation witness (A3's free experiment). +2. Row/column P+Q seal over the 64×64 grid as an external layer, after + Tier 0 №2/№3 (C1's 6.25%/63× numbers, verified by X-C2-3). +3. B1's layout arms under the E1 harness: own-key-bytes sort (zero-cost + arm) vs arrival baseline vs bit-level Morton/Hilbert/spectral — only + if Tier 0 №4 shows measurable locality debt. + +**TIER 3 — credible upstream RFC/PR candidates (OPTIONAL +storage-economics track per §0; each must answer the +§13 questionnaire before filing):** +1. Caller-supplied compaction-rewrite ordering key (M3; precedents: + in-repo R-tree Hilbert leaves, Delta Z-order). +2. Stable-row-id keying for secondary indices (closes the documented + remap-cost gap A3 identified upstream). +3. Overlay-staleness as a compaction trigger (A3's novel-candidate seed; + the docs already name it, no code consumes it). + +**TIER 4 — paper-worthy questions (after Tiers 0–2 produce data):** +1. The refined joint question: address order for query locality with + deliberately ANTI-aligned parity groups (M7). +2. Reader-rung epistemic admission tiers (STRICT/AWARE/RETRO) as a + temporal-database primitive (M9; D3 found no counterpart). +3. The b+1 coalescing amplification analysis as a negative result (M8). + +--- + +## 4. Work order + +Tier 0 items 1–3 are independent of each other and of Tier 1 №4; they can +run as separate probe PRs. Nothing in Tier 2+ starts before its named +Tier 0 gate is green. The charter's phase structure (probes → design → +paper) resumes from here; §16's paper skeleton acquires its §10 (negative +results) and §11 (cross-domain discoveries) content from this document. + +## 5. What the maxim did + +Fifteen researchers were pointed at the beautiful ideas. The comma schedule +died in one afternoon (M5) — with three independent proofs, which is a +cheaper funeral than one shipped defect. The SFC default died (M4). The +amortization headline inverted (M8). What survived is better than what was +proposed: a seal that is boring and correct (M6), a layout seam with real +upstream precedent (M3), two genuinely novel questions with no prior art +(M7, M9), and a Tier 0 list that fixes real, file:line-anchored defects +before any new structure is built on them. diff --git a/docs/lotus/rp-seal-v1/A0-orchestrator-prepass.md b/docs/lotus/rp-seal-v1/A0-orchestrator-prepass.md new file mode 100644 index 00000000..eec720f7 --- /dev/null +++ b/docs/lotus/rp-seal-v1/A0-orchestrator-prepass.md @@ -0,0 +1,75 @@ +# Orchestrator pre-pass — lance 9.0.0 capability audit (PRIVATE until consolidation) + +> Purpose: cross-check against RP-SEAL Domain-A independent reports at +> consolidation. NOT shown to researchers (independence rule). All findings +> from /tmp/sources/lance-9 (upstream tag v9.0.0; workspace Cargo.lock +> checksum 23d04bed056e254bc6e31264b031c8492507ca57939586f016924081dcf221a9). + +## (a) Fragment-level write-without-commit: EXISTS in 9.0.0 + +- `InsertBuilder::execute_uncommitted(data: Vec) -> Result` + — rust/lance/src/dataset/write/insert.rs:132; doc example shows the + two-phase pattern explicitly (execute_uncommitted → CommitBuilder::execute). +- `InsertBuilder::execute_uncommitted_stream(source) -> Result` + — insert.rs:181; doc: "Write data files, but don't commit the transaction + yet. Use CommitBuilder to commit." +- `FragmentCreateBuilder` — rust/lance/src/dataset/fragment/write.rs:71; + fragment-level `write_fragments` at :120. +- Top-level `write_fragments` (dataset/write.rs:586) is DEPRECATED since + 0.20.0 in favor of execute_uncommitted_stream; its doc: fragments "have + not yet been assigned an ID... so this function can be called in + parallel, and the IDs can be assigned after writing is complete." +- `do_write_fragments` (write.rs:597+) is the internal parallel writer. + +## (b) Two-phase prepared commit: EXISTS + +- `Operation` enum — dataset/transaction.rs:320 (Append among variants). +- `CommitBuilder::execute(Transaction)` — dataset/write/commit.rs. +- **`CommitBuilder::execute_batch(Vec) -> BatchCommitResult`** + — commit.rs:475 — MANY prepared transactions in ONE commit; the + many-petals→one-root shape natively. (BatchCommitResult at :512; a + commented-out `rejected: Vec` field at :518 suggests + partial-acceptance semantics are not final — verify behavior.) +- `.with_skip_auto_cleanup(...)` exists on CommitBuilder (seen in + insert.rs do_commit). + +## (c) Orphan/uncommitted cleanup: EXISTS, with an in-flight guard + +- `cleanup_old_versions` — dataset/cleanup.rs:1350: removes "files that are + not referenced by any valid manifest" → abandoned prepared fragments are + collectable garbage BY DEFINITION. +- Unverified-file retention: cleanup.rs:32 "Otherwise we will leave the file + unless delete_unverified is set to true"; RemovalStats tracks `unverified` + (:122-124) — files younger than a retention threshold are NOT deleted + unless explicitly forced. This is the guard that protects IN-FLIGHT + prepared fragments from GC mid-preparation → F-GC's answer candidate. +- `auto_cleanup_hook` (:1356+) runs per lance.auto_cleanup config on commit. + +## (d) Blob storage: EXISTS + +- `Dataset::take_blobs` / `take_blobs_by_addresses` / `take_blobs_by_indices` + — dataset.rs:1737/1752/1761; blob.rs has ReadBlob/ReadBlobRange builders; + write-side ExternalBlobMode + `with_blob_pack_file_size_threshold` + (write.rs:560-575, blob v2 .blob pack sidecar files). + +## Reader-invisibility consequence (the epistemic point) + +Uncommitted data files exist on storage but are invisible to every reader +because readers resolve exclusively through manifests. So "durable but +unpublished" exists NATIVELY in Lance without breaking snapshot atomicity — +the F-VISIBILITY guarantee moves from "structural absence" to "guarded by +manifest-reference", which is the git-object pattern precisely. + +## Open questions the researchers should hit independently + +- Transaction recoverability across restart (is a prepared Transaction + re-derivable from its files, or in-memory only?). +- execute_batch conflict/rebase semantics under concurrent commits. +- Whether any of this surface changed between 9.0.0 and current upstream + (two-column discipline). +- Row-address stability under compaction (stable row IDs? remapping?). + +## Also banked: deltalake ceiling measurement + +deltalake-core newest = 0.32.4, datafusion req ^53.1.0 (registry sparse +index, 2026-08-18) — no DF-54 deltalake exists. diff --git a/docs/lotus/rp-seal-v1/A1.md b/docs/lotus/rp-seal-v1/A1.md new file mode 100644 index 00000000..169d5705 --- /dev/null +++ b/docs/lotus/rp-seal-v1/A1.md @@ -0,0 +1,750 @@ +# A1 — Lance 9.0.0 Storage/Compaction Archaeology (BUILDER) + +Two-column discipline observed throughout: **Column A** = exact pinned +`lance = "=9.0.0"` at `/tmp/sources/lance-9` (verified `version = "9.0.0"` in +root `Cargo.toml`, matching the workspace's `Cargo.lock` pin, checksum +`23d04bed056e254bc6e31264b031c8492507ca57939586f016924081dcf221a9`). **Column +B** = current upstream at `/tmp/sources/lance-main` (shallow clone, +`version = "11.0.0-beta.14"`, HEAD `a8e2a78` "fix(encoding): reject zero +fixed-size-list dimension in repdef decimation (#8618)", dated +2026-08-18). Every claim below is tagged `[9.0.0]`, `[main-only]`, or +`[both]`. + +--- + +## SOURCE ARCHAEOLOGY + +### 1. Fragments and DataFiles — the physical layout unit + +`rust/lance-table/src/format/fragment.rs` [both, structurally identical shape] + +- `DataFile` (fragment.rs:29-58): one physical file backing a subset of a + fragment's columns. Fields: `path: String`, `fields: Arc<[i32]>` (field ids + present in this file), `column_indices: Arc<[i32]>` (file-local column + offsets; `-1` marks a field with no top-level column, e.g. blob payload + columns), `file_major_version`/`file_minor_version` (the Lance **file** + format version, decoupled from the **table** format version), `file_size_bytes: + CachedFileSize`, `base_id: Option` (which registered base path this + file lives under, for multi-base/shallow-clone datasets, feature-flagged + via `FLAG_BASE_PATHS`). +- `DataFileFieldInterner` (fragment.rs:237-330): a linear-scan-then-HashMap + cache that Arc-shares identical `fields`/`column_indices`/version-metadata + byte payloads across fragments on manifest deserialization. Doc comment + states the motivation explicitly: *"At 20M fragments the deduplication + typically saves multiple GB of heap because every fragment in a homogeneous + table carries the same field list."* This is a manifest-memory optimization, + not a data-layout change — directly relevant to the research program's + "logical & physical bytes" and "peak RSS" metrics at high fragment counts. +- `Fragment` (fragment.rs:489-520-ish): `id: u64`, `files: Vec`, + `overlays: Vec` (see §6), `deletion_file: Option`, + `row_id_meta: Option` (see §2), `physical_rows: Option` + (original row count, including deleted rows — **public API**, used + everywhere for O(1) row-count arithmetic instead of re-scanning files), + `last_updated_at_version_meta` / `created_at_version_meta: Option` + (see §7 — **per-row write-time provenance already lives in the fragment + metadata**). +- `DeletionFile` (fragment.rs, near 400): `read_version: u64`, `id: u64`, + `file_type: {Array, Bitmap}`, `num_deleted_rows: Option`, + `base_id: Option`. Soft-deletes are files, not in-place mutation of the + data file — public API, feature-flagged `FLAG_DELETION_FILES`. + +`num_rows()` on `Fragment` (fragment.rs, ~510) is `physical_rows - +deletion_file.num_deleted_rows` when both are known — a fully O(1), +metadata-only computation. This is the number the compaction planner uses to +decide fragment candidacy without touching file bytes (§3). + +### 2. Row addressing — two independent addressing schemes, feature-switched + +`rust/lance-core/src/utils/address.rs:23-66` [both]: `RowAddress(u64)` packs +`(fragment_id: u32) << 32 | (row_offset: u32)`. `FRAGMENT_SIZE = 1<<32`, +`TOMBSTONE_FRAG = 0xffffffff`. This is **physical addressing**: it changes +whenever a row is rewritten into a different fragment/offset (compaction, +update). This is what indices store when `FLAG_STABLE_ROW_IDS` is unset +(default in older tables). + +`rust/lance-table/src/rowids.rs` + `rowids/index.rs` [both]: **stable row +IDs** (`FLAG_STABLE_ROW_IDS = 2`, `feature_flags.rs:13-14`, doc: *"Row ids are +stable for both moves and updates. Fragments contain an index mapping row ids +to row addresses."*). `RowIdSequence(Vec)` (rowids.rs:49) is a +per-fragment sequence of u64 row ids, run-length/range-optimized via the +private `U64Segment` enum (contiguous `Range`, encoded array, etc.). +`RowIdIndex` (rowids/index.rs:27) is `RangeInclusiveMap` — disjoint row-id ranges mapping to (row-id-segment, +address-segment) pairs; `get(row_id) -> Option` is a range-map +lookup + segment `.position()`/`.get()`, i.e. O(log F) where F = number of +disjoint chunks, not O(rows). `get_many()` (rowids/index.rs:73+) sorts the +input batch once and walks the range map linearly, amortizing to +`O(F + N log N)`. `FragmentRowIdIndex { fragment_id, row_id_sequence, +deletion_vector }` is the per-fragment unit fed into `RowIdIndex::new`. + +Under stable row ids, compaction moves rows to new physical addresses but the +**logical row id never changes**, so `commit_compaction` (optimize.rs:2010, +see §3) sets `needs_remapping = !dataset.manifest.uses_stable_row_ids() && +!options.defer_index_remap` — **stable-row-id tables skip index remapping +entirely on compaction**; indices keyed by row id resolve addresses through +the always-current `RowIdIndex` at query time instead. This is the single +biggest lever in the whole file-set for the research program's "index remap +bytes/CPU" metric: it can be driven to **zero** by choosing stable row ids, +independent of any erasure-coding or SFC-locality work. + +### 3. The compaction planner and pipeline + +`rust/lance/src/dataset/optimize.rs` (8,509 lines incl. tests) [both, `pub +mod`, **not** feature-gated in `dataset.rs:82`]. + +Module doc (optimize.rs:4-73) states the two triggers for compaction +candidacy (fewer rows than `target_rows_per_fragment` with mergeable +neighbors; deletion percentage over a threshold) and the distributed +execution shape: + +``` +plan_compaction() -> CompactionPlan -> {CompactionTask.execute() -> RewriteResult}* -> commit_compaction() +``` + +- `CompactionMode` (optimize.rs:134-156) `[both]`: `Reencode` (default, + decode+re-encode), `TryBinaryCopy` (binary copy with Reencode fallback), + `ForceBinaryCopy` (binary copy or hard fail). **This directly answers the + "reencode / binary-copy modes" rumor: it exists verbatim in 9.0.0.** +- `IndexRemapMode` (optimize.rs:159-186) `[both]`: `Compact` (default is + actually `Direct`; `Compact` trades lookup CPU for O(#fragments) memory via + `RowAddrRemap::Compact`, see §5) vs `Direct` (materializes a full + `HashMap>`, O(#rows) memory, O(1) lookup). +- `CompactionOptions` (optimize.rs:194-320ish) `[9.0.0 fields]`: + `target_rows_per_fragment` (default `1024*1024`), `max_rows_per_group` + (default 1024), `max_bytes_per_file`, `materialize_deletions` (default + true), `materialize_deletions_threshold` (default 0.1), `num_threads`, + `batch_size`, `io_buffer_size`, **`defer_index_remap: bool`** (answers the + "deferred index remapping" rumor — present in 9.0.0), `index_remap_mode`, + `compaction_mode: Option`, deprecated + `enable_binary_copy`/`enable_binary_copy_force`, + `binary_copy_read_batch_bytes` (default 16 MiB), + **`max_source_fragments: Option`** (answers "incremental compaction + limits" — present in 9.0.0: *"allows for incremental compaction (e.g., + compact 20 fragments at a time)"*), **`max_overlays_per_fragment: + Option`** (default `Some(10)` — a fragment with too many attached + overlay files (§6) is force-compacted regardless of size/deletions), + `transaction_properties`. All fields are also settable via dataset config + keys `lance.compaction.*` (`CompactionOptions::from_dataset_config`, + `apply_dataset_config`, optimize.rs:~350-420). +- `DefaultCompactionPlanner::plan()` (optimize.rs:685-830ish) `[9.0.0]`: + streams `collect_metrics` over every fragment (buffered by + `io_parallelism()`), computes `CompactionCandidacy::{CompactItself, + CompactWithNeighbors}` per fragment from three signals — overlay-count + overflow, deletion-percentage overflow, undersize — bins adjacent + same-index-coverage fragments into `CandidateBin`s (fragments covered by + different sets of indices are **never** merged into one compaction task — + "we cannot combine indexed fragments with unindexed fragments", module doc + line 21-23), splits each bin by `target_rows_per_fragment` + (`CandidateBin::split_for_size`), then truncates the task list by + `max_source_fragments` if set. +- `compact_files()` / `compact_files_with_planner()` (optimize.rs:855-900) + `[9.0.0]`: `futures::stream::iter(tasks).map(rewrite_files).buffer_unordered(num_threads)`, + i.e. genuinely parallel task execution within one process, `try_collect`, + then a single `commit_compaction` call. +- `CompactionPlan { tasks: Vec, read_version: u64, options: + CompactionOptions }` (optimize.rs:938) — serializable (`Serialize + + Deserialize`), so it is the wire-format for the distributed-execution story + (module doc: split the plan, ship `CompactionTask`s to other machines, + ship `RewriteResult`s back, commit any subset — *"as long as the tasks + don't rewrite any of the same fragments, they can be committed in any + order"*, doc lines 78-82). +- `commit_compaction()` (optimize.rs:2002-2231) `[9.0.0]` is the single most + information-dense function for this research program: + - Anchors `read_version` to **`min(task.read_version)`** across all + completed tasks rather than `dataset.manifest.version`, with a load-bearing + comment explaining why (distributed callers open a fresh `Dataset` handle + for `commit_compaction` that may be far ahead of the handle used to + `plan_compaction`; anchoring to the max would let the OCC conflict checker + silently skip a concurrent DELETE and resurrect deleted rows). + - Branches three ways on `(needs_remapping, defer_index_remap)`: + 1. **stable row ids** (`!needs_remapping`, no defer): no remap at all, + just `reserve_fragment_ids` for the new fragments. + 2. **immediate remap** (`needs_remapping`): builds either a `Direct` + `HashMap` or a `Compact` `GroupInput` list (see §5) from each task's + `row_addrs: Option>` (a serialized `RoaringTreemap` of the old + row addresses that were actually rewritten — deleted rows are simply + absent), calls `remap_options.create_remapper(dataset)`, and produces + `RewrittenIndex` entries that go straight into the `Transaction`. + 3. **deferred remap** (`defer_index_remap`): builds `FragReuseGroup`s + (§4) instead, and — the "all-or-nothing" invariant, explicit in a code + comment — only actually **writes a Fragment Reuse Index at all** if at + least one rewritten group is covered by an existing data index or by + the FRI's own existing chain; otherwise it logs `"skipping + fragment-reuse index: no rewritten fragments were covered by an + index"` and commits with `frag_reuse_index: None`. The comment is + explicit about *why* partial-FRI is unsound: *"a concurrent reindex + can make a skipped fragment indexed and the conflict resolver's + FRI-present path won't re-check it."* + - Builds one `Transaction` with `Operation::Rewrite { groups, + rewritten_indices, frag_reuse_index }` and calls `dataset.apply_commit`; + on failure, calls `cleanup_data_fragments` on the newly-written (now + orphaned) fragment files before propagating the error — i.e. compaction + writes data files speculatively and only reconciles success/failure at + the very end (see §8, "prepared writes"). + +### 4. Deferred index remapping via the Fragment Reuse Index (FRI) + +`rust/lance-table/src/system_index/frag_reuse.rs` (480 lines) [9.0.0, +structurally present in main too] is the **format-layer** definition; consumed +from `rust/lance/src/dataset/index/frag_reuse.rs` (599 lines) and built in +`rust/lance/src/index/frag_reuse.rs` (not read in full, referenced from +optimize.rs `build_new_frag_reuse_index`). + +- `FRAG_REUSE_INDEX_NAME = "__lance_frag_reuse"` (frag_reuse.rs:18) — a + **system index** (`lance_index::is_system_index`), not a user-visible + secondary index. +- `FragDigest { id: u64, physical_rows: usize, num_deleted_rows: usize }` + (frag_reuse.rs:22-26) — a lightweight fragment fingerprint, not the full + `Fragment`. +- `FragReuseGroup { changed_row_addrs: Vec, old_frags: Vec, + new_frags: Vec }` (frag_reuse.rs:63-66) — `changed_row_addrs` + is a serialized `RoaringTreemap` of the old physical addresses that moved + (i.e. **exactly the answer to the rumor's "old-row-address -> new-row- + address remapping" — this is the literal, on-disk payload of that + mapping**, one group per compaction task). +- `FragReuseVersion { dataset_version: u64, groups: Vec }` + (frag_reuse.rs:100-104) — one FRI "version" per compaction transaction; the + FRI itself is a **chain** of these versions, accumulated across multiple + deferred compactions until every data index has caught up. +- `cleanup_frag_reuse_index()` (`dataset/index/frag_reuse.rs:44`) — doc: + *"If all the indices currently available are already caught up to as a + specific reuse version, all older reuse versions (inclusive) can be cleaned + up."* Caught-up test: an index is caught up against a FRI version if its + fragment-bitmap coverage is disjoint from that version's touched fragments, + **or** the index is at/past the version's dataset version and holds none of + the version's old fragment ids. This is a live, lazy "repair debt" GC over + a chain of pending remaps — a close structural analog to LSM-tree + tombstone/GC-generation bookkeeping (see PRIOR ART). + +The mechanism (confirmed by external Lance design-discussion search, see +PRIOR ART): compaction with `defer_index_remap = true` skips remapping the +indices synchronously; the FRI records old->new address deltas; a stale +index is corrected **lazily at index-load time** by consulting the FRI chain +in memory. This decouples compaction cadence from index-build cadence at the +cost of index-load-time translation (paid once per process, not per query +once cached) — directly the "repair CPU/wall" vs. "compaction CPU/wall" +tradeoff the research program's metric list separates. + +### 5. Row-address remap encoding — rank-based, not per-row hashmap + +`rust/lance-core/src/utils/row_addr_remap.rs` [both] is the general-purpose +remap consumer both the immediate-remap and (indirectly, via FRI replay) +deferred-remap paths feed into. + +- `RowAddrRemap::Direct(HashMap>)` — O(#rewritten rows) + memory, O(1) lookup. +- `RowAddrRemap::Compact(CompactRowAddrRemap)` — O(#fragments) memory. Module + doc (row_addr_remap.rs:1-44) spells the algorithm out precisely: per old + fragment, store `(old_offsets: bitmap-like segment, old_rows_before: u64)`; + per new fragment (in write order), store `(fragment_id, new_rows_before, + physical_rows)`. A lookup **ranks** the old offset within its fragment's + offset set (`old_offsets.rank(offset) - 1`), adds `old_rows_before` to get + a global rewritten-row index `k`, then finds which new-fragment range + contains `k`. The doc is explicit that this depends on order preservation: + *"old_frag_ids must match the order old fragments are read... new_frags + must match the order new rows are written... Current compaction satisfies + this because it scans selected fragments in order and writes the resulting + stream without reordering rows."* + +This is a genuine rank-based succinct-style remap (see PRIOR ART: rank/select +bit-vector literature) already living inside Lance 9.0.0, independent of and +predating any erasure-coding proposal. It is a strong candidate substrate for +the research program's "index remap bytes/CPU" experiments, since switching +`IndexRemapMode` from `Direct` to `Compact` is a one-line config change with +measurable memory/CPU tradeoffs already wired end-to-end. + +### 6. Data overlay files — unreleased, cell-level, rank-addressed updates + +`rust/lance-table/src/format/overlay.rs` [9.0.0, present but **explicitly +unreleased**] — `DataOverlayFile` supplies new values for a subset of +`(physical offset, field)` cells **without rewriting the fragment's base data +files**. Coverage is Roaring-bitmap-encoded over **physical offsets** (stable +across deletions, like deletion vectors), either `Shared` (one bitmap for all +fields, dense case) or `PerField` (sparse case). Values are stored +**rank-addressed** within the overlay's own data file: *"a covered offset's +value sits at its rank — the 0-based count of set bits below it in that +field's coverage bitmap"* — the exact same rank-in-bitmap addressing pattern +as §5's `RowAddrRemap::Compact`, independently reinvented for a different +purpose inside the same codebase (flagged as a SURPRISE below). Overlays are +ordered newest-last (`sort_overlays_newest_last`) and a later `committed_version` +wins on conflict; a field tombstone sentinel (`TOMBSTONE_FIELD_ID = -2`, +overlay.rs:51) invalidates a stale overlay column when the base data for that +field is later rewritten. + +Feature-gating (`feature_flags.rs:22-33`): `FLAG_UNSTABLE_DATA_OVERLAY_FILES += 64`. Debug builds always understand it; **release builds treat it as an +unknown flag and refuse the dataset** unless the environment variable +`LANCE_ENABLE_UNSTABLE_DATA_OVERLAY_FILES` is set. This is a genuinely +unreleased, benchmark-only feature as of 9.0.0/main — a real prior-art match +for the research program's "can a layout key derived from identity make +physical canonicalization constructive rather than a repair sort" question, +in that overlays let a writer avoid rewriting a fragment for a small cell +update, at the cost of read-path fan-in across overlay layers, which +`CompactionOptions.max_overlays_per_fragment` (§3) exists specifically to +bound. + +### 7. Per-row temporal provenance already inside Lance + +`rust/lance-table/src/rowids/version.rs` [both], module doc: *"Row version +tracking for cross-version diff functionality... track the latest update +version for each row... enabling efficient cross-version diff operations."* + +- `RowDatasetVersionRun { span: U64Segment, version: u64 }` (version.rs:30) + — a run of identical dataset-version stamps over a **contiguous span of row + offsets** (positions within `RowIdSequence` order, not row ids themselves). +- `RowDatasetVersionSequence { runs: Vec }` + (version.rs:56) — RLE-encoded per-fragment. +- Each `Fragment` (§1) carries **two independent instances** of this: + `created_at_version_meta` (the version at which the row was first written — + the research program's `W_write` for row birth) and + `last_updated_at_version_meta` (the version at which the row's value was + most recently touched, e.g. an update or an overlay landing on it). +- Both are rechunked during compaction (`optimize.rs`, ~1950-1992, in the + block right before `commit_compaction`): `rechunk_version_sequences` maps + the old per-old-fragment sequences onto the new fragment boundaries (using + `chunk_sizes` derived from the new fragments' `physical_rows`), with a + `debug_assert_eq!(old_total, new_total)` invariant that row counts are + conserved across the rewrite. + +**Cross-domain relevance (OGAR grounding).** OGAR's +`docs/TEMPORAL-TIME-TRAVEL.md` (pinned commit `386a6fd8`, read literally per +task instructions) frames the temporal-epistemology layer as strictly a +**planner-level query annotation over the raw Lance version counter** — *"No +new storage, no new contract, no new container"* — with `KnowledgeItem.created_at` +mapped to *"Lance version V at which the row landed"* and cross-server +ordering deferred to a hybrid logical clock stamped on an external +`CognitiveEventRow`, explicitly because Lance itself was assumed to expose +only the single monotonic dataset-version counter, not per-row write-time +metadata. **That assumption is only half true**: Lance 9.0.0 already +persists exactly a per-row `W_write` value (`created_at_version_meta`) and a +per-row `A_last`-adjacent value (`last_updated_at_version_meta`), both +RLE-encoded and rechunked correctly through compaction, entirely inside the +table format — no planner-side bookkeeping, no HLC stamping, needed for the +**single-writer, single-server** case the "interlacing window" hypothesis +describes. See SURPRISES. + +### 8. "Prepared but uncommitted" — three distinct real mechanisms, no unified name + +The rumored "prepared/uncommitted writes" concept maps onto **three +structurally different real mechanisms** in 9.0.0, not one: + +1. **Ordinary write-then-commit.** `Dataset::write` / fragment writers + (`rust/lance/src/dataset/write.rs`, `fragment/write.rs`) write data files + to object storage *before* any manifest references them; a `Transaction` + is only durable once its manifest is committed (`io/commit.rs`, doc: + *"a transaction is committed by writing the next manifest file... care + should be taken to ensure that the manifest file is written only once"*). + `commit_compaction`'s failure path (`cleanup_data_fragments`, §3) proves + this: compacted fragment files are written speculatively and deleted on + commit failure. `cleanup.rs`'s module doc explicitly names the ambiguity + this creates: *"it is difficult to distinguish between a data/tx/idx file + which was leftover from an abandoned transaction and a data file which is + part of an ongoing operation... If the file is at least 7 days old then we + assume it is probably not part of any ongoing operation."* — a real, + *heuristic*, age-based orphan-file GC policy, not a hard guarantee. +2. **Detached manifests** (`rust/lance-table/src/format/manifest.rs:106-111`, + `DETACHED_VERSION_MASK = 0x8000_0000_0000_0000`; `Dataset::commit_detached` + at `rust/lance/src/dataset.rs:1587`; `list_detached_manifests` at + `dataset.rs:2577-2591`, doc: *"used for staging changes"*). A detached + commit writes a fully-formed manifest whose version has the top bit set — + it is a real, durable, queryable snapshot but **never becomes part of the + dataset's linear version chain** (`is_detached_version`, manifest.rs:109). + This is the closest thing in Lance to "prepared fragments, immutable, one + publication event later" — except detached manifests are a terminal + staging artifact, not (in 9.0.0) something later "promoted" into the main + chain by a second explicit publish call. +3. **The MemWAL subsystem** (`rust/lance/src/dataset/mem_wal/`, `pub mod + mem_wal` unconditionally registered, `dataset.rs:80`) — a genuinely + separate durability layer ahead of the Lance table: `wal.rs` writes Arrow + IPC batches to object storage with `writer_epoch` fencing metadata and + **bit-reversed file naming** (explicit doc: *"distribute files evenly + across S3 keyspace"* — i.e. it already solves an SFC-adjacent hot-partition + problem, just for WAL segment naming, not row layout). `is_terminal_failure` + checks `error.fence_reason().is_some()` for the classic + split-brain-writer fencing pattern. This is out of scope for compaction + proper but is the actual "data durable, not yet a queryable Lance + fragment" staging tier in 9.0.0 — worth flagging for whichever cell + investigates write-ahead durability specifically. + +None of these three composes into anything resembling "immutable code +symbols before one manifest publication, RS-encoded" — that specific +combination (prepared fragments as erasure-coding source symbols, published +atomically) has **no counterpart anywhere in Lance's source**, in either +column (see PRIOR ART, §"erasure coding: absent"). + +### 9. Commit protocol / version publication + +`rust/lance/src/io/commit.rs` [both]: trait `CommitHandler`; default +`ConditionalPutCommitHandler` (conditional-put-to-temp-then-rename, doc +lines 11-15) for object stores that support it; `CommitLock` as a simpler +alternative for stores that only offer locking. `conflict_resolver.rs` +(`TransactionRebase`) implements optimistic concurrency control: a +transaction records its `read_version`, and on commit the resolver scans +every manifest between `read_version` and the current head for conflicting +operations before allowing the write to proceed — this is the mechanism +`commit_compaction`'s `tasks_read_version` anchoring (§3) is protecting. + +### 10. Cleanup / GC + +`rust/lance/src/dataset/cleanup.rs` (4,334 lines incl. tests) [both]. Deletes, +in order of conservatism: old (non-latest) manifests past an age threshold; +data files unreferenced by any valid manifest; delete files unreferenced by +any valid manifest; index files unreferenced by any valid manifest. The +7-day-old heuristic (§8.1) is the load-bearing fallback for files that are +unreferenced but too young to trust as abandoned. This is the closest match +to the research program's "stale-version retention bytes" and "restart +recovery work" metrics — it is entirely reference-counting-by-manifest-scan, +no reachability index, no generation counters beyond raw age. + +### 11. Manifest structure (version publication unit) + +`rust/lance-table/src/format/manifest.rs:35-104` [both]: `schema`, `version: +u64`, `branch: Option`, `writer_version`, `fragments: Arc>` +(sorted by fragment id, gaps allowed), `version_aux_data`, `index_section`, +`timestamp_nanos: u128`, `tag`, `reader_feature_flags`/`writer_feature_flags` +(u64 bitmasks, §1/§2/§6/§8.2 flags all live here), `max_fragment_id`, +`transaction_file`/`transaction_section` (inline-vs-external transaction +provenance, `FLAG_DISABLE_TRANSACTION_FILE` controls which), private +`fragment_offsets: Vec` (precomputed prefix-sum for O(log F) fragment +lookup by global row offset — direct evidence Lance already treats the +fragment sequence as address space requiring locality-aware lookup), +`next_row_id: u64` (the stable-row-id allocator high-water mark), +`data_storage_format`, `config`/`table_metadata` (two separate key-value +maps — client config vs. arbitrary table metadata), `base_paths: HashMap`. + +### 12. What's public API vs internal + +Public, stable-surface (re-exported from `lance::dataset::optimize` per the +module's own doctest, which is a **compiled, CI-run example** — +`rust/lance/src/dataset/optimize.rs:26-56`): `compact_files`, +`CompactionOptions`, `CompactionMetrics`, `CompactionMode`, `IndexRemapMode`, +`IgnoreRemap`, `IndexRemapper`/`IndexRemapperOptions`/`RemappedIndex` +(re-exported from the `remapping` submodule, `pub mod remapping` at +optimize.rs:126), `plan_compaction`-family types (`CompactionPlan`, +`CompactionTask`, `RewriteResult`), `commit_compaction`, +`CompactionPlanner` trait + `DefaultCompactionPlanner`. + +Internal/crate-private: `binary_copy` (`mod binary_copy;` — no `pub`, +optimize.rs:125), `CandidateBin`/`CompactionCandidacy`/`FragmentMetrics` +(planning-internal), `reserve_fragment_ids`, `rewrite_files`. The Fragment +Reuse Index internals (`lance-table::system_index::frag_reuse`) are `pub` at +the crate level (needed by `lance-index`'s consumer side) but explicitly a +**system index**, invisible through the normal user-facing index API +(`lance_index::is_system_index` filters it out of listings). + +`optimize` and `cleanup` are both plain `pub mod` in `dataset.rs` +(lines 74, 82) — **neither is behind a Cargo feature flag**; they compile +into every build of the `lance` crate regardless of `default-features`. + +--- + +## MECHANISM + +Putting §§3-7 together, Lance 9.0.0's compaction/remap mechanism is: + +``` +plan_compaction (metadata-only scan, candidacy binning by size+deletion%+overlay-count, + index-coverage-aware grouping) + | + v +CompactionTask (serializable, distributable) + | + v +rewrite_files: read old fragments -> {reencode | binary-copy} -> write new fragments, + capture RoaringTreemap of rewritten OLD row addresses + | + v +RewriteResult { new_fragments, row_addrs: Option>, metrics, read_version } + | + v +commit_compaction: branch on (stable_row_ids?, defer_index_remap?) + -- stable row ids: no remap needed at all + -- immediate + Direct: HashMap> (O(rows) mem, O(1) lookup) + -- immediate + Compact: rank-based per-fragment remap (O(frags) mem, O(log) lookup) + -- deferred: FragReuseGroup{changed_row_addrs} appended to a + chain (__lance_frag_reuse system index); + consumed lazily at each index's next load + | + v +Transaction{Operation::Rewrite} -> apply_commit -> OCC conflict check against + [tasks_read_version, head] -> new Manifest version +``` + +The binary-copy path (§3, `binary_copy.rs`) additionally page-copies raw +encoded bytes with 64-byte alignment padding for V2.1+ files, rejecting input +files carrying extra global buffers (footer is always regenerated, never +copied) — i.e. compaction already distinguishes "logical row rewrite" from +"physical byte copy" as two genuinely different code paths with different +cost profiles, which is exactly the logical/physical byte distinction the +research program's metric list separates. + +## EXPECTED BENEFIT + +For the research program specifically: Lance 9.0.0 already offers, with zero +new code, a large fraction of the "amortized compaction + deferred remap" +half of the hypothesis: +- `defer_index_remap` decouples compaction cadence from index-remap cost + (§3, §4) — directly testable today. +- `IndexRemapMode::Compact` vs `Direct` is a real, already-implemented + memory/CPU tradeoff for the "index remap bytes/CPU" metric (§5). +- Stable row ids eliminate remap cost entirely for a large class of + workloads (§2) — the natural control arm/baseline for any remap-cost + experiment. +- `max_source_fragments` (9.0.0) and upstream-only `excluded_fragment_ids`/ + `max_source_rows`/`max_source_bytes` (§ COLUMN DIFF) give bounded, + incremental compaction "for free" as a planning-boundary primitive that a + Morton-cascade amortization scheme could sit directly on top of, rather + than reimplementing fragment-selection logic. +- Per-row `created_at_version`/`last_updated_at_version` (§7) removes the + need for a bespoke external write-time stamp for the single-writer, + single-dataset temporal-deinterlacing case. + +## EXPECTED FAILURE + +- **No erasure coding of any kind exists in Lance** (§8, verified by + exhaustive grep in both columns — zero hits for `reed_solomon`, `erasure`, + `parity_group`, `ec_shard`). Any RS-seal proposal is 100% new surface + grafted onto this substrate, not an extension of an existing partial + implementation — so its integration cost (new file format, new manifest + fields, new feature flag, new commit-path hooks) should be estimated at + "greenfield," not "extend existing machinery." +- **`defer_index_remap` is explicitly incompatible with stable row ids** + (`optimize.rs`, `DefaultCompactionPlanner::plan`, hard error: *"stable row + IDs do not require index remapping during compaction, so there is nothing + to defer"*) — any experiment design combining "deferred remap" and "stable + row ids" as independent knobs will hit a hard `Err`, not silently degrade. +- **All-or-nothing FRI writing** (§3) means a compaction that touches only + not-yet-indexed fragments writes *no* FRI at all — an experiment measuring + "seal metadata bytes" against deferred compaction must condition on + whether any rewritten fragment was actually index-covered, or it will + measure zero cost for a real deferred-remap run and wrongly conclude the + mechanism is free. +- **The row-address-remap "order preservation" invariant is load-bearing and + undocumented at the type level** (§5): `RowAddrRemap::Compact` is silently + wrong (not panicking, just wrong) if a future rewrite path reorders rows + during the read-then-write pipeline. Any research code building on this + primitive needs its own reordering test, since nothing in the type system + enforces it. +- Data overlay files (§6) are release-gated behind an env var specifically + because they are pre-production; any comparison baseline using them must + disclose that and should expect upstream churn. + +## EXPERIMENT + +Pre-registered designs, scoped to what column-A Lance 9.0.0 can measure +directly without new Rust code (glass-box harness over `lance::dataset::optimize`): + +**E1 — Index-remap cost vs. mode, at fixed compaction shape.** +- *Metric*: index remap CPU (wall time of the `Direct`/`Compact` branch in + `commit_compaction`), index remap bytes (`RewrittenIndex` serialized size), + peak RSS during `commit_compaction`. +- *Baseline*: `IndexRemapMode::Direct` (the crate default) on a table with an + IVF-PQ vector index and 20% of fragments undersized. +- *Control*: `IndexRemapMode::Compact` on the identical fragment set/index. +- *Kill condition*: if `Compact` mode's wall time exceeds `Direct` mode's by + more than the RSS savings would justify at the target fragment count (i.e. + crossover point never reached below realistic table sizes), the + research program's implicit assumption that "compact remap is a free + memory win" is falsified for that workload shape. + +**E2 — Deferred remap: repair debt accumulation and drain cost.** +- *Metric*: FRI chain length (versions), FRI metadata bytes, index-load-time + CPU with N pending FRI versions (0, 1, 5, 20), point-lookup latency through + an index with a live FRI chain vs. a freshly-remapped index. +- *Baseline*: `defer_index_remap = false` (immediate remap every compaction). +- *Control*: `defer_index_remap = true`, run 20 compactions before any + `remap_column_index`/reindex, then measure. +- *Kill condition*: if index-load-time FRI-replay CPU grows worse than + linearly in chain length (the doc's O(1)-per-query claim after caching + would predict linear-in-chain-length load cost, flat per-query cost after), + the "compaction and index building no longer conflict" claim from the + design discussion (see PRIOR ART) does not hold at scale, which bounds how + aggressively any Morton-cascade "amortize compaction" scheme can defer. + +**E3 — Binary-copy vs. reencode: physical vs. logical byte divergence.** +- *Metric*: write amplification (bytes written / logical bytes compacted), + compaction CPU/wall, files touched. +- *Baseline*: `CompactionMode::Reencode`. +- *Control*: `CompactionMode::ForceBinaryCopy` on a homogeneous-version + fragment set (the only case it accepts, per `binary_copy.rs`'s version- + and global-buffer-compatibility checks). +- *Kill condition*: if binary copy's wall-time win over reencode is smaller + than the extra complexity of maintaining format-compatibility checks + (measured as: fraction of real-world fragment sets that pass the + `is_binary_copy_compatible` check at all), binary copy is not a generalizable + lever for any repair/rewrite scheme layered on top. + +**E4 — Stable row ids as the zero-remap control arm.** +- *Metric*: index remap bytes/CPU (should measure exactly 0 for stable row + ids), point/random-take latency (row id -> address indirection cost vs. + direct physical addressing). +- *Baseline*: legacy physical row addressing (no `FLAG_STABLE_ROW_IDS`). +- *Control*: stable row ids enabled. +- *Kill condition*: if point/random-take latency under stable row ids is + worse by more than the remap-cost savings at realistic query rates, "stable + row ids make remap free" is true but not a net win — reframes any + seal/compaction proposal that assumes stable-row-id tables as the default + substrate. + +**E5 — Rank-based remap vs. hashmap remap, at scale, isolated microbenchmark.** +- *Metric*: peak RSS, remap-construction CPU, per-lookup CPU, for + `RowAddrRemap::Compact` vs `RowAddrRemap::Direct` at 10^5 / 10^7 / 10^9 + rewritten rows, independent of the rest of the compaction pipeline (direct + unit-level microbench against `lance-core::utils::row_addr_remap`). +- *Baseline*: `Direct`. +- *Control*: `Compact`. +- *Kill condition*: crossover row count where `Compact`'s extra per-lookup + cost outweighs its memory savings — this number directly bounds any + claim that a Morton/SFC-ordered rank scheme would out-perform Lance's + existing generic rank-based remap at the row counts the research program's + baseline geometry (65,536 rows/cycle, 32 MiB) actually produces per unit of + work. + +## PRIOR ART + +Searches actually run (with queries, verbatim): + +1. `WebSearch: "rank select bitmap row address remapping compaction storage + engine succinct data structure"` — returned academic rank/select + literature (Engineering Compact Data Structures for Rank and Select + Queries on Bit Vectors, arXiv 2206.01149; A Fast x86 Implementation of + Select, arXiv 1706.00990; several Springer chapters on succinct rank/ + select). **Classification for §5/§6's rank-in-bitmap addressing: KNOWN** + — rank/select over bit vectors is 40+ years of established succinct-data- + structure theory (Jacobson-style constant-time rank/select in n + o(n) + bits); Lance's application of it to row-address remapping and to overlay- + cell addressing is a straightforward, uncredited-but-standard + **TRANSFER** of that primitive into a storage-engine remap/update context. +2. `WebSearch: "Lance format stable row IDs fragment reuse index deferred + remap design"` — returned the Lance/lancedb GitHub Discussions and + Issues that originated this design: "Re-evaluating stable row ID + implementation" (Discussion #6933), "Stable Row ID" (Discussion #3694), + "Remap row ids during compaction" (Issue #1378), "Epic: Move-Stable Row + Ids" (Issue #2307), plus `lance.org/format/index/` and + `lance.org/format/table/`. These confirm the FRI mechanism's public + design rationale verbatim matches what the source shows: *"allowing + compaction to skip the index remap step, instead recording a mapping from + old fragment row addresses to new ones, and when indices are loaded into + the cache, the FRI is applied to translate the old row addresses to the + current ones... adds a small cost to index load time but does not affect + query performance once the index is cached... decoupling means compaction + and index building no longer conflict."* **Classification: KNOWN** — this + is documented, intentional Lance design, not a hidden or accidental + mechanism; the research program should cite Lance's own FRI as prior art + for "deferred repair," not propose it as novel. +3. `WebSearch: "Iceberg Delta Lake Hudi compaction manifest rewrite + comparison Lance table format versioned storage"` — comparative material + on the three dominant lakehouse table formats. **Classification: KNOWN / + TRANSFER** — Iceberg's manifest-list -> manifest-file -> data-file 3-tier + metadata tree is structurally richer than Lance's 2-tier Manifest -> + Fragment -> DataFile (§11), but serves the same "avoid rewriting a stale + pointer to every data file on every commit" goal; Hudi's record-level + index -> file-group direct lookup is a near-exact **TRANSFER** analog of + Lance's `RowIdIndex` -> `RowAddress` lookup (§2), independently arrived + at in a different codebase for the same reason (mutable-row workloads + need row-stable addressing under a copy-on-write/merge-on-read storage + layer); Delta Lake's flat sequential `_delta_log` is structurally closer + to Lance's `_transactions/` directory + numbered manifests (§9) than to + Iceberg's tree. +4. Read literally (per task instruction, not a search but recorded as + grounding): `OGAR docs/TEMPORAL-TIME-TRAVEL.md` at pinned commit + `386a6fd8`. Confirms the exact `T_now`/`A_last`/`W_write` framing + originates from a documented planner-layer design (`CONTEMPORARY`/ + `ANACHRONISTIC`/`SPOILER` `TemporalStatus`, `QueryReference{ref_version, + mode, rung}`) that was, at time of writing, assumed to require an + *external* per-row write-time stamp (HLC) because Lance was believed to + expose only the single dataset-version counter. §7 above is the + falsification-relevant finding against that assumption for the + single-writer case. + +No search was run specifically for "erasure coding + compaction locality" +literature (e.g. Reed-Solomon repair-bandwidth papers, regenerating codes) — +that is squarely domain B/C territory (bidirectional erasure-coded seals is +not a Lance mechanism, confirmed absent in §8/EXPECTED FAILURE) and is +flagged here as an explicit gap for whichever researcher owns that axis, not +silently skipped. + +## SURPRISES + +1. **Finding**: Lance already tracks per-row `created_at_version` and + `last_updated_at_version` (RLE-encoded, rechunked correctly through + compaction) directly in fragment metadata — the exact `W_write`/`A_last` + primitives the research program's temporal-deinterlacing hypothesis (and + OGAR's own documented design) assumed would need external HLC + bookkeeping. + **Classification: NOVEL-CANDIDATE** (as a *connection*, not as a Lance + feature — the feature itself is KNOWN/documented Lance functionality; the + novel part is that no design doc I read, in this repo's board files + excluded per the independence rule or in OGAR's temporal doc, appears to + have identified this specific Lance primitive as a substitute for + external HLC stamping in the single-writer case. Flagged as + NOVEL-CANDIDATE pending confirmation that the consolidation phase doesn't + find this already known elsewhere in the 15-cell corpus.) +2. **Finding**: The exact same rank-in-bitmap addressing trick + (`rank(bit_position) -> dense value-array index`) is independently + implemented twice in Lance's own source for two unrelated purposes — + compaction's `RowAddrRemap::Compact` (§5) and the unreleased overlay + files' per-cell value storage (§6) — with no shared abstraction between + them. + **Classification: KNOWN** (rank/select itself is textbook; the specific + *reuse pattern within one codebase* is a TRANSFER-of-a-transfer, not + independently novel, but worth flagging to whoever designs a unified + locality/remap primitive for the research architecture: Lance's own + internal precedent argues for factoring this into one shared + rank-addressed-value-store abstraction rather than a third bespoke + implementation). +3. **Finding**: `defer_index_remap` is hard-incompatible with stable row + ids by design (an explicit `Err`, not a silent no-op) — the two levers + the research program might reach for to reduce remap cost are mutually + exclusive in the current API, not composable. + **Classification: KNOWN** (directly documented in the source error + message, not inferred) — but worth flagging because a naive experiment + grid crossing both as independent factors will silently lose cells to + errors unless designed around this. +4. **Finding**: lance-graph's actual production write path + (`crates/lance-graph/src/graph/cycle_sink.rs`) uses **only** + `Dataset::write(WriteMode::Create)` and `dataset.append()` — grepping the + entire `crates/` tree for `CompactionOptions`, `compact_files`, + `RowIdSequence`, `RowAddrRemap`, `FragReuseGroup`, `IndexRemapMode`, or + `CompactionMode::` returns **zero hits outside `lance`/`lance-table` + themselves**. None of §§2-7's machinery is invoked anywhere in the + consumer codebase today. + **Classification: KNOWN** (this is a direct, mechanical grep result, not + an inference) — but it reframes "EXPECTED BENEFIT" materially: the + research program's proposed architecture would be the **first** consumer + of Lance's compaction/remap/FRI machinery in this workspace, not an + optimization of an existing hot path. Any benefit projection should be + against "currently unused capability, cost zero today" rather than + "currently-paid cost we can reduce." +5. **Finding**: `excluded_fragment_ids` (compaction planning boundaries) and + `max_source_rows`/`max_source_bytes` (finer incremental-compaction caps + beyond `max_source_fragments`) exist **only** in upstream main + (11.0.0-beta.14), added after the 9.0.0 tag; `max_source_fragments`, + `defer_index_remap`, `CompactionMode`, the FRI, and `RowAddrRemap` are all + already present in 9.0.0. + **Classification: DISPROVED** (as a "9.0.0 has this" rumor) for + `excluded_fragment_ids`/`max_source_rows`/`max_source_bytes` specifically + — anyone designing against the pinned 9.0.0 dependency must not assume + these three fields exist; they are a genuine main-only capability gap + that would require either an upstream version bump (against this + workspace's stated exact-lockstep-pin discipline) or a local + reimplementation of fragment-exclusion as a pre-filter on + `dataset.get_fragments()` before calling the 9.0.0 planner. + +## VERDICT + +Lance 9.0.0's compaction/repair/remap subsystem is substantially more mature +and closer to the research program's hypothesis space than a from-scratch +design would assume — deferred index remap via a versioned Fragment Reuse +Index, rank-based O(#fragments) address remapping, stable row ids that +eliminate remap cost entirely, incremental/bounded compaction via +`max_source_fragments`, and per-row birth/update version provenance are all +real, shipped, `pub`, un-feature-gated code in the exact pinned version — +but every one of these levers is presently **unused** by lance-graph's +production write path, erasure coding has zero footprint anywhere in Lance +(either column), `excluded_fragment_ids`/`max_source_rows`/`max_source_bytes` +are upstream-only and not available at the pinned 9.0.0, and the FRI's +all-or-nothing write rule plus the stable-row-id/defer-remap incompatibility +are two concrete traps any experiment or architecture design must account +for rather than discover empirically. diff --git a/docs/lotus/rp-seal-v1/A2.md b/docs/lotus/rp-seal-v1/A2.md new file mode 100644 index 00000000..5acc932f --- /dev/null +++ b/docs/lotus/rp-seal-v1/A2.md @@ -0,0 +1,725 @@ +# Cell A2 — Domain A (Lance storage/compaction), Role: ADVERSARY + +**Thesis under attack:** *"random first-come chunk placement is cheap in Lance."* + +**Source discipline.** Column A = exact pinned `lance` **9.0.0** at `/tmp/sources/lance-9` +(`git describe` → `v9.0.0`, HEAD `7653c20`, workspace `Cargo.toml` `version = "9.0.0"`). +Column B = current upstream at `/tmp/sources/lance-main` (`a8e2a78`, `version = +"11.0.0-beta.14"`, cloned 2026-08-18). Every claim below is column A unless the row says +otherwise; column B is consulted **only** to ask "is this already fixed upstream?" — the +answer, for every defect in this report, is **no**. + +Consumer source: `/home/user/lance-graph` working tree (read-only). + +**Verdict up front:** the thesis is **DISPROVED as stated** and survives only in a much +narrower form. Random first-come placement is cheap *to write* and cheap *for +within-one-file random take*; it is **not** cheap for anything else, and — the sharpest +finding — Lance 9.0.0 **cannot repair it later**, because compaction has no reordering +capability at all and the compact index-remap makes order preservation a *correctness +invariant*, not an implementation detail. + +--- + +## SOURCE ARCHAEOLOGY + +### A. The consumer's actual write path (what is being defended) + +| Fact | Evidence | +|---|---| +| One cycle = 64k fleet | `crates/lance-graph-supervisor/src/cycle_driver.rs:1549` `const FLEET: u32 = 65_536;` | +| Cycle image is keyed by logical row | `crates/lance-graph-planner/src/persist_sink.rs:377-392` — `freeze` sorts casts by `stream_position`, then folds into `image: BTreeMap>` | +| Batch = 1 frame row + L landing rows + I image rows | `crates/lance-graph/src/graph/cycle_sink.rs:459` `let n = 1 + batch.landings.len() + batch.image.len();` | +| Payload column is `FixedSizeBinary(512)` | `cycle_sink.rs:112` `EPISODIC_WITNESS_BYTES = 512`; schema at `cycle_sink.rs:124-145` | +| The store mutation is `Dataset::append` with **default** `WriteParams` | `cycle_sink.rs:726` `ds.append(reader, None)`; create path `cycle_sink.rs:710-725` uses `..Default::default()` | +| Reads are **unindexed full scans with a filter** | `cycle_sink.rs:984` `scan.filter("kind = 1 AND cycle > N")`; `cycle_sink.rs:1116` `scan.filter("kind = 2 AND cycle = N")`; `find_frame` `cycle_sink.rs:560` | +| **No `compact_files`, no `create_index`, no `cleanup_old_versions` anywhere in the write arc** | `grep -rn "compact_files\|create_index\|cleanup_old_versions" crates/lance-graph{,-planner,-supervisor}/src` → zero hits (the only `compaction` hits are `surreal_container`'s own WAL fold and a comment in `arigraph/sensorium.rs:372`) | + +Two immediate consequences that frame everything below: + +1. **Nothing in the shipped arc ever compacts, indexes, or GCs.** The "cheap" claim has + never met the machinery that would price it. Every cost model in this report is + therefore a *prediction about the first time that machinery is switched on*, and the + pre-registered experiments are designed to be run before it is. +2. `WriteParams::default()` sets `enable_stable_row_ids: false` + (`rust/lance/src/dataset/write.rs:430`) and `max_rows_per_file: 1024*1024` + (`write.rs:418`). At ~131,073 rows/cycle this is **exactly one fragment and exactly one + version per cycle**, with **address-style (unstable) row ids** — the mode that *requires* + index remapping on compaction. + +### B. Lance 9.0.0 compaction — the load-bearing reads + +**B1. Bins are formed by adjacency in the fragment-id-sorted list, nothing else.** +`rust/lance/src/dataset/optimize.rs:707-812`. `get_fragments()` is asserted sorted by id +(`optimize.rs:713-716`); a running `current_bin` accumulates consecutive candidates. Two +things terminate a bin: + +- `optimize.rs:794-797` — `(None, Some(_)) => { candidate_bins.push(...) }`: **a single + non-candidate fragment ends the run.** +- `optimize.rs:777-790` — `if bin.indices == indices { …push… } else { …close bin, start + new… }`: **a change in the set of indices covering the fragment ends the run.** + +**B2. There is no sort, no Z-order, no clustering option — in either column.** +`CompactionOptions` (`optimize.rs:191-280`) has 16 fields; none of them is an ordering key. +`grep -rn "z_order\|zorder\|hilbert\|cluster_by\|sort_by:" rust/lance/src/dataset/optimize.rs` +returns nothing in column A **and nothing in column B**. Compaction merges; it never +re-sorts. + +**B3. Order preservation is a *correctness invariant* of the compact remap, not an +accident.** `rust/lance-core/src/utils/row_addr_remap.rs:35-45`: + +> *"Compact remap does not store each old-to-new row mapping. It computes `k` from the +> old-row layout, then maps it to the k-th row written to the new fragments. **This +> requires the reader-to-writer pipeline to preserve row order.** … Current compaction +> satisfies this because it scans selected fragments in order and writes the resulting +> stream without reordering rows."* + +`GroupRemap::new` (`row_addr_remap.rs:139-160`) hard-*errors* if new fragment ids are not +ascending. So a hypothetical locality-repairing compaction would silently mis-map every +index entry under `IndexRemapMode::Compact` unless the remap were redesigned. + +**B4. The default remap mode materializes one hash entry per rewritten row, at the +coordinator, for the whole compaction.** `IndexRemapMode::default()` is `Direct` +(`optimize.rs:171-176`). `commit_compaction` accumulates `direct_row_id_map: +HashMap>` across **all** tasks before remapping +(`optimize.rs:2049`, `2092-2098`), fed by `transpose_row_addrs` +(`optimize/remapping.rs:155-194`), which pre-sizes to +`rewritten_rows + deleted_rows` (`remapping.rs:180-188`). + +**B5. `reserve_fragment_ids` is itself a full commit.** `optimize.rs:1541-1573` builds +`Operation::ReserveFragments` and calls `commit_transaction` — a manifest write and a +version bump — **before** the `Rewrite` commit that `commit_compaction` later performs +(`optimize.rs:2036-2045`). Identical in column B (`optimize.rs:2120-2132`). Fragment ids +are only ever advanced; nothing recycles them. + +**B6. `Rewrite` conflict surface.** `rust/lance/src/io/commit/conflict_resolver.rs:729-880`. +Compatible-with-anything: `Append`, `ReserveFragments`, `Project`, `Clone`, `UpdateConfig`, +`UpdateMemWalState`, `UpdateBases`. Conflicting: `Merge` unconditionally (`:821-823`); +concurrent `Rewrite` touching the same old fragments (`:790-795`) **or** any concurrent +`Rewrite` that also carries a fragment-reuse index (`:796-801`); `DataOverlay` on a touched +fragment (`:771-784`); `Delete`/`Update` on a touched fragment (`:750-770`); +`DataReplacement` on a touched fragment (`:806-819`); `CreateIndex` under several +straddling conditions (`:823-880`). Default `num_retries` is **20** +(`rust/lance-table/src/io/commit.rs:1550`), and *every* attempt re-lists and re-reads the +new transactions (`rust/lance/src/io/commit.rs:958-975`, comment: *"We are pessimistic +here"*). + +**B7. Stable row ids and the fragment-reuse witness are mutually exclusive.** +`needs_remapping = !uses_stable_row_ids && !defer_index_remap` (`optimize.rs:2013`), and +`plan()` hard-rejects `defer_index_remap` on a stable-row-id dataset +(`optimize.rs:697-706`). So turning on stable row ids to *avoid* the remap simultaneously +removes the durable `FragReuseGroup` old→new map (`optimize.rs:2141-2146`) — the only +persisted permutation witness the format produces. + +### C. Manifest, cleanup, retention + +| Fact | Evidence | +|---|---| +| Manifest carries the **complete** fragment list | `rust/lance-table/src/format/manifest.rs:52` `pub fragments: Arc>` (identical column B, `manifest.rs:55`) | +| Every commit re-serializes the whole manifest | `rust/lance-table/src/io/manifest.rs:185-233` → `do_write_manifest` | +| Per-fragment protobuf ≈ path string + two int32 vectors + versions + sizes | `protos/table.proto:313-343` (`DataFragment`), `:365-400` (`DataFile`) — ~100–140 B/fragment for this 13-column schema (**to be measured, not asserted**) | +| Cleanup **reads every retained manifest** | `rust/lance/src/dataset/cleanup.rs:498-514` `process_manifests` → `list_manifest_locations` → `process_manifest_file` → `read_manifest` per version | +| Cleanup also full-LISTs five subtrees | `cleanup.rs:660-668` (`versions/`, `transactions/`, `data/`, `indices/`, `deletions/`) | +| Unverified-file protection is a 7-day wall-clock window | `cleanup.rs:319` `const UNVERIFIED_THRESHOLD_DAYS: i64 = 7;` used at `cleanup.rs:625-626, 648` | +| `delete_unverified(true)` removes that protection entirely | `cleanup.rs:1300-1306`, doc: *"should only be done if the caller can guarantee there are no updates happening at the same time"* | +| Auto-cleanup default cadence would be **every 20 versions** | `rust/lance/src/dataset/write.rs:258-265` `AutoCleanupParams { interval: 20, older_than: 14 days }` — but `WriteParams::default()` leaves `auto_cleanup: None` (`write.rs:433`), so the consumer has it **off** | +| `Dataset::commit` does **not** verify the fragments exist | `rust/lance/src/dataset.rs:1530-1531`: *"This method will not verify that the provided fragments exist and correct, that is the caller's responsibility."* | + +### D. Read path physics + +| Fact | Evidence | +|---|---| +| `RowAddress` = `frag_id:u32 \|\| offset:u32` | `rust/lance-core/src/utils/address.rs:24-37`, `FRAGMENT_SIZE = 1 << 32` | +| `take` has three branches: contiguous / sorted / scatter | `rust/lance/src/dataset/take.rs:180`, `:210`, `:232`, `:282`; `check_row_addrs` at `:451-470` | +| Sorted/scatter branches issue **one `do_take` future per fragment**, buffered at `io_parallelism` | `take.rs:255-281` | +| Opening a data file costs a `size()` + a `block_size` tail read | `rust/lance-file/src/reader.rs:749-757` `read_tail` | +| `block_size` = 4 KiB local, **64 KiB cloud** | `rust/lance-io/src/object_store.rs:66`, `:77`, `:1166-1167` | +| `io_parallelism` = 8 local, **64 cloud** | `object_store.rs:62-64` | +| Range coalescing window is exactly `block_size` | `rust/lance-io/src/scheduler.rs:1139-1142` `is_close_together`; merge loop `:1173-1184`; split at `max_iop_size` = 16 MiB (`object_store.rs:79-83`) | +| Metadata cache default 1 GiB, index cache 6 GiB | `rust/lance/src/dataset.rs:157`, `:161` | +| Statistics-based filter pushdown is **v1-only** | `rust/lance/src/io/exec/pushdown_scan.rs:664-665` `.expect("Operation does not yet support v2 fragments")`; selection at `scanner.rs:2818-2825` | +| For v2, an unindexed filter is a `FilteredReadExec` refine over a scan | `scanner.rs:2903-2966` `new_filtered_read`; dispatch `scanner.rs:2971-2999` | +| Zone maps exist **only** as an explicitly built scalar index | `rust/lance-index/src/scalar/zonemap.rs:52,107` (`zonemap.lance`, `ZoneMapIndex`) | +| **Cloud** late-materialization threshold is `byte_width < 1000` ⇒ *early* | `scanner.rs:2350-2375` `is_early_field`; applied at `scanner.rs:2403` `subtract_predicate(\|f\| !self.is_early_field(f))` | +| `FixedSizeBinary(512).byte_width_opt() == Some(512)` | `rust/lance-arrow/src/lib.rs:170` | +| **Local** threshold is `byte_width < 10` ⇒ 512 B is *late* | `scanner.rs:2373` | +| 512 B ≥ 256 B ⇒ the payload gets **fullzip** structural encoding (per-value repetition index) | `rust/lance-encoding/src/encodings/logical/primitive.rs:3860-3872` `is_narrow`, `MINIBLOCK_MAX_BYTE_LENGTH_PER_VALUE = 256`; `prefers_miniblock` `:3875-3885` | + +The last row is the one genuinely **favourable** finding and it is stated honestly: +Lance's adaptive structural encoding does give the 512-byte witness column a repetition +index, so a `take` of scattered rows *inside one file* is cheap by construction. Every +failure below is across files, across versions, or in metadata — never inside a page. + +--- + +## MECHANISM + +The adversarial mechanism is a single sentence with three legs: + +> **In Lance 9.0.0 the physical order of rows inside a fragment is fixed at write time and +> is never changed again by any operation the format offers; compaction merges runs of +> *arrival-adjacent* fragments without reordering; and the only mechanism that would let +> an index survive a reordering (`IndexRemapMode::Compact`) is *defined* by the +> assumption that no reordering happened.** + +Leg 1 — **write-time freeze.** `DetachedCycleBatch::freeze` (`persist_sink.rs:377-392`) +canonicalizes by `stream_position` and then emits image rows in `BTreeMap` (row-ascending) +order, so *within one cycle* the image is grid-ordered. But each cycle is its own fragment +(`Dataset::append`, default `max_rows_per_file`), so the **global** physical order is +`(cycle, row)` — i.e. arrival-major, grid-minor. A neighborhood query over a grid +region spanning K cycles touches K fragments no matter how small the region is. + +Leg 2 — **compaction merges arrival order.** `DefaultCompactionPlanner::plan` +(`optimize.rs:707-812`) bins consecutive fragment ids; `rewrite_files` scans the bin in +order and writes the stream through. The output is arrival-major, grid-minor at a larger +granularity. Compaction reduces *fragment count*; it does not reduce *scatter*. Calling +this "repair of physical locality" is precisely the analogy this cell was asked to test, +and it fails: Lance compaction is a **fixative**, not a repair. + +Leg 3 — **the remap forbids the repair.** `row_addr_remap.rs:35-45` and the ascending-id +guard at `:139-160` make order preservation a contract. Any design that later wants +"compaction = locality repair" must either (a) run only in `Direct` mode and pay +§F3's RAM, (b) use stable row ids and thereby lose the `FragReuseGroup` witness +(`optimize.rs:697-706`, `:2013`), or (c) rewrite the whole dataset outside the compaction +API (`Overwrite`), which is a full read-modify-write and conflicts with everything +(`conflict_resolver.rs:734-750` — `Overwrite` is *not* in `Rewrite`'s compatible set). + +Around that spine sit six independent amplifiers, each with its own trigger. + +--- + +## EXPECTED BENEFIT + +Stated so the attack is not a strawman. Random first-come placement genuinely wins on: + +- **Write path.** One append per cycle, one manifest publication, no read-modify-write, no + sort, no shuffle. `freeze` is `sort_by_key` over ≤64k `u64`s plus a `BTreeMap` fold — + microseconds against a network round trip. +- **Within-file random take.** The 512-byte payload lands on **fullzip** structural + encoding with a per-value repetition index (`primitive.rs:3860-3872`), which is exactly + the mechanism the Lance paper (arXiv 2504.15247) measures as beating Parquet on random + access for values ≥ 16 B. A `take` of scattered offsets **inside one fragment** is close + to optimal. +- **Commit conflict surface.** A pure `Append` is compatible with every other operation in + the conflict matrix that matters (`conflict_resolver.rs:743-749` lists `Append` as + unconditionally compatible with `Rewrite`); the single-writer topology + (`cycle_sink.rs:200-262`) means zero contention in practice. +- **Crash story on the append path.** `Dataset::append` is one MVCC commit; a crash before + it leaves an unreferenced data file that ages out. The `batch_hash` idempotency key + (`persist_sink.rs:401-430`) plus reconciliation-first `commit_cycle` gives a genuinely + clean retry. This is real and should be preserved by whatever replaces the placement + policy. + +Everything in the next section is a cost the above does **not** pay for. + +--- + +## EXPECTED FAILURE + +Ten scenarios. Each gives the mechanism at `file:line`, the workload that triggers it, the +metric that measures it, and the **kill condition** — the measurement that, if it comes +back the other way, proves *this cell* wrong rather than the thesis. + +### F1 — Compaction cements arrival order; it cannot repair locality +**Mechanism.** `optimize.rs:707-812` (adjacency binning), no ordering option in +`CompactionOptions` (`optimize.rs:191-280`, columns A **and** B), +`row_addr_remap.rs:35-45` (order preservation is a contract). +**Workload.** Write N cycles where each cycle touches a uniform random subset of the 65,536 +grid rows. Compact to a single fragment. Then run a neighborhood query over a contiguous +Morton range of the grid. +**Metric.** *pages touched per query* and *files touched per query*, before vs after +compaction, vs a control dataset written grid-sorted. +**Kill condition.** If post-compaction pages-touched for a grid-local query drops to within +1.5× of the grid-sorted control, this cell is wrong and compaction *does* repair locality. +**Prediction.** It does not move at all except by the fragment-count factor: the rows are +in the same relative order, just in a bigger file. + +### F2 — One large fragment splits the compaction run forever +**Mechanism.** `optimize.rs:794-797` `(None, Some(_))` closes the bin; +`optimize.rs:777-790` closes it again on any index-membership change. +**Workload.** Alternate cycles: 99 small cycles (few dirty rows) then one full-sweep cycle +whose fragment exceeds `target_rows_per_fragment` (default 1 Mi rows, +`optimize.rs:288`). Repeat. +**Metric.** *fragment count* after `compact_files`, and `CompactionPlan.tasks.len()`. +**Kill condition.** If fragment count after compaction is O(total_rows / +target_rows_per_fragment), the binning is not adjacency-fragile. **Prediction:** it is +O(number of large cycles), because every large fragment is a barrier. +**Second-order.** The same barrier is created by *partial* indexing: index one prefix of +the store and every bin boundary lands on the indexed/unindexed frontier. + +### F3 — Default index remap is an O(rows) coordinator-resident hash map +**Mechanism.** `IndexRemapMode::default() = Direct` (`optimize.rs:171-176`); +`direct_row_id_map` accumulates across **all** tasks (`optimize.rs:2049`, `:2092-2098`); +`transpose_row_addrs` sizes to `rewritten + deleted` (`remapping.rs:180-188`); entry is +`u64 → Option` = 8 + 16 B plus hashbrown control byte, at ≤ 87.5 % load and +power-of-two capacity ⇒ **~29–57 B/row resident**. +**Workload.** Any store with ≥ 1 scalar/vector index and unstable row ids (the default, +`write.rs:430`) compacted in one pass. 131,073 rows/cycle × 1,000 cycles = 1.31×10⁸ rows. +**Metric.** *peak RSS* of the compaction coordinator; *index remap bytes*; *index remap CPU*. +**Kill condition.** Peak RSS stays sub-linear in rewritten rows — i.e. `Direct` is not +actually materializing one entry per row. **Prediction:** 3.7–7.5 GB at 1,000 cycles, +OOM well before the store is interesting. +**Aggravator.** `IndexRemapMode::Compact` exists and is O(#fragments) +(`row_addr_remap.rs:53-64`) — but it is **not** the default, and it is precisely the mode +that forbids reordering (F1). The design is offered a choice between "cannot afford the +remap" and "can never reorder". + +### F4 — On cloud storage the 512-byte payload is read **eagerly**, and a local benchmark inverts the result +**Mechanism.** `scanner.rs:2350-2375`: `Heuristic` materialization treats a field as +*early* when `byte_width < 1000` on cloud, `< 10` on local. +`FixedSizeBinary(512).byte_width_opt() == Some(512)` (`lance-arrow/src/lib.rs:170`). +`scanner.rs:2403` subtracts only the *non*-early fields from the eager projection. +So on cloud the witness payload is **not** late-materialized; on local it **is**. +**Workload.** `scan_image(cycle = N)` (`cycle_sink.rs:1108-1140`) against a store of T +image rows, once on `file://` and once on an S3-compatible endpoint. +**Metric.** *logical bytes read* and *physical bytes read* per query, plus *files & pages +touched*, as a function of T with the matched-row count held constant at 65,536. +**Kill condition.** Cloud bytes-read stays proportional to *matched* rows rather than +*total* rows — i.e. `FilteredReadExec` demotes the column despite `is_early_field`. +**Prediction:** cloud bytes-read grows as 512·T while local stays near 512·65,536. This is +the single most dangerous item in the report, because the local number *looks fine* and is +the one a developer will measure. +**Column B:** identical heuristic and identical constants (`lance-main` +`scanner.rs:2549-2553`). Not fixed upstream. + +### F5 — There is no index, so every read is a scan; adding one re-arms F2 and F3 +**Mechanism.** No `create_index` call anywhere in the consumer write arc (grep, §A). +Statistics pushdown is v1-only (`pushdown_scan.rs:664-665`), so v2 unindexed filters are a +`FilteredReadExec` refine over a full scan (`scanner.rs:2903-2966`). Zone maps are an +opt-in scalar index (`lance-index/src/scalar/zonemap.rs:52,107`). +**Workload.** `find_frame(cycle)` and `scan_image(cycle)` — both already on the recovery +and read paths — as T grows. +**Metric.** *point-read latency*, *files & pages touched per query*, *cache misses*. +**Kill condition.** Latency flat in T (would mean some pruning exists that this reading +missed). +**The trap:** the obvious fix — build a `BTree`/`ZoneMap` index on `cycle` — makes +compaction bins split on index membership (F2), makes remapping mandatory (F3), and puts +`Rewrite` into conflict with `CreateIndex` (`conflict_resolver.rs:823-880`). There is no +configuration in which all three are cheap. + +### F6 — Manifest bytes are quadratic in cycle count; cold open is linear +**Mechanism.** `manifest.rs:52` (full fragment vector) + `io/manifest.rs:185-233` (full +re-serialization every commit). One fragment and one version per cycle (§A). +**Workload.** N cycles of steady append. Let b = bytes/fragment in the manifest protobuf +(`protos/table.proto:313-400`; expected 100–140 B for this 13-column schema — **measure, +do not assume**). +**Metric.** *logical bytes written* (cumulative manifest bytes ≈ b·N²/2), *version count*, +*restart recovery work* (bytes read by `Dataset::open` ≈ b·N), *stale-version retention +bytes*. +**Kill condition.** Cumulative manifest bytes sub-quadratic in N — i.e. some delta/ +checkpoint manifest exists that this reading missed. **Checked column B: it does not** +(`lance-main/rust/lance-table/src/format/manifest.rs:55`). +**Calibration.** At b = 130 B and N = 10⁶ cycles: latest manifest ≈ 130 MB read on *every* +cold open; cumulative manifest bytes ≈ 65 TB written. At a modest one cycle/minute that N +is ~2 years. + +### F7 — Cleanup is O(retained versions × fragments) and reads every manifest; and it is in direct conflict with the temporal thesis +**Mechanism.** `cleanup.rs:498-514` reads **every** retained manifest; `cleanup.rs:660-668` +full-LISTs five subtrees. Default auto-cleanup cadence, if enabled, is every **20 versions** +(`write.rs:258-265`). +**Workload.** Enable `lance.auto_cleanup.interval = 20` on a one-version-per-cycle store. +**Metric.** *compaction/GC CPU and wall*, *stale-version retention bytes*, *version count*. +**Kill condition.** Per-run cleanup cost sub-linear in retained manifests. +**The deeper conflict.** OGAR `docs/TEMPORAL-TIME-TRAVEL.md` @ `386a6fd8` maps +`KnowledgeHorizon(T)` → `dataset.checkout_version(V_ref)` and `EpistemicMode::RETRO/AWARE` +→ reading rows whose `lance_version > ref_version`. **Every manifest cleanup deletes an +epistemic horizon.** A design that treats temporal deinterlacing as first-class cannot run +version cleanup on the range it deinterlaces over — so `stale-version retention bytes` is +not an operational afterthought, it is a *designed* liability, and it must appear in the +architecture's own cost model. This is not a Lance defect; it is a genuine, unresolved +tension between two of the program's own hypotheses. + +### F8 — "Prepared fragments, then ONE manifest publication" is structurally false for compaction and unsafe in general +**Mechanism (a) — it is not one publication.** `reserve_fragment_ids` is itself a +`commit_transaction` (`optimize.rs:1541-1573`), executed inside `commit_compaction` +(`optimize.rs:2036-2045`), *before* the `Rewrite` commit. Compaction therefore costs +**≥ 2 versions**, and a failure between them leaves a version that changed nothing except +a permanently advanced `max_fragment_id`. +**Mechanism (b) — commit does not verify.** `dataset.rs:1530-1531`: *"This method will not +verify that the provided fragments exist and correct."* +**Mechanism (c) — cleanup can delete a prepared file.** Files newer than the 7-day +`UNVERIFIED_THRESHOLD_DAYS` (`cleanup.rs:319`) are spared **only** when +`delete_unverified == false` (`cleanup.rs:648`). The API offers `delete_unverified(true)` +(`cleanup.rs:1300-1306`). Combined with (b): a cleanup run with `delete_unverified` racing +a preparation window produces a **committed manifest referencing deleted files** — a +version that is permanently unreadable and does not self-heal. +**Mechanism (d) — concurrent commit.** 20 retries by default +(`lance-table/src/io/commit.rs:1550`); each attempt writes a transaction file +(`rust/lance/src/io/commit.rs:954-956`), and the *pessimistic* re-list happens every +attempt (`:958-975`). +**Workload.** (i) Kill the coordinator between `ReserveFragments` and `Rewrite`. (ii) Run +`cleanup_old_versions(older_than = 1h, delete_unverified = true)` concurrently with a +multi-hour prepare. (iii) Run two compactions with overlapping fragment sets. +**Metric.** *version count* (expect +1 per aborted compaction), *fragment count*, orphan +data/transaction-file bytes, and a binary **corruption flag** (does `Dataset::open` at the +new version succeed?). +**Kill condition.** Compaction completes in exactly one manifest publication, **or** a +committed manifest referencing a deleted data file is rejected/repaired at commit time. +**Prediction:** neither holds. + +### F9 — Object-store request amplification is set by fragment count, not by row count +**Mechanism.** Per data file per cache miss: a `size()` (HEAD) plus a `block_size` tail GET +— 64 KiB on cloud (`reader.rs:749-757`, `object_store.rs:77`). `take` issues one +`do_take` per fragment (`take.rs:255-281`). Coalescing merges only within `block_size` +(`scheduler.rs:1139-1142`). +**Workload.** A K-row point-take whose rows are spread over F fragments, sweeping F from 1 +to 10⁴ with K fixed. +**Metric.** *point-take latency*, *files & pages touched per query*, *physical bytes read*, +*cache misses*, plus request count from the object-store IO tracker +(`object_store.rs:1029`). +**Kill condition.** Requests and bytes flat in F (would mean metadata caching fully absorbs +it; the 1 GiB default, `dataset.rs:161`, makes that plausible only while F is small). +**The pincer.** `is_close_together` merges any two ranges within `block_size`, so scattered +rows either (a) fail to coalesce → request amplification, or (b) coalesce across a gap → +byte amplification. Random placement guarantees you are always paying one of the two. A +grid-ordered layout is the only regime where both are small at once. + +### F10 — Fragment ids are u32 and monotonically consumed +**Mechanism.** `RowAddress::FRAGMENT_SIZE = 1 << 32` (`address.rs:27`) ⇒ fragment id is 32 +bits. `reserve_fragment_ids` only advances `max_fragment_id` (`optimize.rs:1566-1572`); +nothing recycles, and every *aborted* compaction still consumes its reservation. +**Workload.** Long-horizon append at F fragments/cycle. +**Metric.** *fragment count* over time; extrapolate to 2³². +**Kill condition.** Fragment ids are reused or the space is wider than u32. +**Prediction.** At 1 fragment/cycle this is ~4.3×10⁹ cycles (not a near-term risk); at one +fragment per *work field* (4,096/cycle, the geometry in the program brief) it is ~10⁶ +cycles — inside the design horizon. + +--- + +## EXPERIMENT (pre-registered) + +Seven designs. Each names metric, baseline, control, and kill condition **before** any +number is produced. All are read-only against the pinned column-A source plus purpose-built +fixture datasets; none touches the consumer repo. No experiment is run in this cell. + +Common fixture generator (call it `GEN`): write N cycles into a fresh Lance dataset with +the `cycle_store_schema()` shape (13 columns, `payload FixedSizeBinary(512)`), one +`Dataset::append` per cycle, default `WriteParams`. Two placement arms: + +- **arm R (random / first-come):** each cycle's image rows are a uniform random subset of + the 65,536 grid rows. +- **arm G (grid-ordered control):** the *same* logical row→payload content, but written + so that physical order equals grid (Morton) order — realized by buffering C cycles and + emitting one sorted fragment per C-cycle super-batch. Arm G is the **control**, and it + is deliberately *not* an implementation proposal; it exists to bound how much locality + is on the table. + +### X1 — Does compaction repair locality? (kills or confirms F1) +- **Metric:** *pages touched per query* + *files touched per query* + *neighborhood-take + latency* for a Morton-contiguous range of 4,096 grid rows. +- **Baseline:** arm R before `compact_files`. +- **Control:** arm G before compaction (the locality ceiling). +- **Procedure:** measure; run `compact_files(target_rows_per_fragment = 1<<20)`; measure + again; report the ratio R_after/G. +- **Kill condition:** `R_after` within **1.5×** of `G` on pages-touched ⇒ compaction *is* + a locality repair and F1 is wrong. +- **Pre-registered prediction:** `R_after/G` stays within 10 % of `R_before/G` on + pages-touched (only *files* touched improves). + +### X2 — Adjacency fragility (F2) +- **Metric:** *fragment count* after compaction; `plan.tasks.len()`; *compaction CPU/wall*. +- **Baseline:** 1,000 uniform small cycles. +- **Control:** the same 1,000 cycles with one over-target fragment inserted every 100. +- **Kill condition:** final fragment counts differ by < 2× ⇒ binning is not + adjacency-fragile. +- **Second arm:** repeat with a scalar index covering only fragments `[0, 500)`, to + isolate the `bin.indices == indices` split (`optimize.rs:777-790`). + +### X3 — Remap cost curve (F3) +- **Metric:** *peak RSS*, *index remap bytes*, *index remap CPU*. +- **Baseline:** `IndexRemapMode::Direct` (the default). +- **Control:** `IndexRemapMode::Compact`, and a third arm with + `enable_stable_row_ids = true` (no remap at all). +- **Sweep:** rewritten rows from 10⁶ to 10⁸. +- **Kill condition:** `Direct` peak RSS grows sub-linearly, or stays under 1 GB at 10⁸ + rewritten rows. +- **Reported alongside:** whether the `Compact` arm can be used *at all* by a design that + wants reordering (it cannot — `row_addr_remap.rs:35-45`), and whether the stable-row-id + arm still produces a `FragReuseGroup` witness (it does not — `optimize.rs:697-706`). + +### X4 — The cloud/local materialization inversion (F4) — highest priority +- **Metric:** *logical bytes read* + *physical bytes read* + *files & pages touched per + query* for `scan_image(cycle = N/2)`. +- **Baseline:** local `file://` store. +- **Control:** the **same bytes** served through an S3-compatible endpoint (MinIO), so + `ObjectStore::is_cloud()` is true and `block_size` is 64 KiB. +- **Sweep:** T (total image rows) from 10⁵ to 10⁸ with matched rows fixed at 65,536. +- **Kill condition:** cloud bytes-read grows no faster than local bytes-read ⇒ F4 is wrong. +- **Third arm (the fix probe):** re-run cloud with + `MaterializationStyle::AllLate` (`scanner.rs:273`) and with the payload widened past + 1,000 B, to check that the threshold is the actual cause rather than a coincidence. +- **Why first:** this failure is *invisible* on a developer laptop and would otherwise be + discovered in production. + +### X5 — Metadata quadratic (F6, F7) +- **Metric:** *logical bytes written* (manifest only), *version count*, *restart recovery + work* (bytes + wall for `Dataset::open` cold), *stale-version retention bytes*, + *compaction/GC CPU and wall* for one `cleanup_old_versions` run. +- **Baseline:** N ∈ {10³, 10⁴, 10⁵} cycles, no cleanup. +- **Control:** same N with `lance.auto_cleanup.interval = 20`, `older_than = 14 days` + (the documented defaults, `write.rs:258-265`) and a clock that makes cleanup actually + fire. +- **First deliverable:** measure **b**, the real bytes/fragment in the manifest protobuf + for this schema. Every quadratic estimate in this report is `b·N²/2` and **b is currently + an estimate, not a measurement.** +- **Kill condition:** cumulative manifest bytes sub-quadratic in N, or cold-open bytes + sub-linear in N. + +### X6 — Prepared-fragments crash and race matrix (F8) +Six cells, each a binary outcome plus the metrics named: + +| # | Injected event | Outcome to record | +|---|---|---| +| 1 | kill between `ReserveFragments` commit and `Rewrite` commit | version count Δ, `max_fragment_id` Δ, orphan data bytes | +| 2 | kill after data files written, before any commit | orphan data bytes; do they age out at 7 d? | +| 3 | `cleanup(older_than=1h, delete_unverified=true)` during a 2 h prepare, then commit | **corruption flag**: does `Dataset::open(new_version)` succeed? | +| 4 | same as 3 with `delete_unverified=false` | expect: safe (the 7-day window holds) | +| 5 | two `Rewrite`s over overlapping fragments | retry count, transaction-file orphan bytes, final version count | +| 6 | `Rewrite` + concurrent `Merge` | expect unconditional retryable conflict (`conflict_resolver.rs:821-823`) | + +- **Kill condition for the "one publication" claim:** cell 1 shows **zero** extra versions. +- **Kill condition for the corruption claim:** cell 3 shows the commit rejected, or the + version readable. + +### X7 — Request amplification vs fragment count (F9) +- **Metric:** object-store *request count* (via `ObjectStore::io_tracker`, + `object_store.rs:1029`), *physical bytes read*, *point-take latency p50/p99*, + *cache misses*. +- **Baseline:** K = 1,024 rows taken from F = 1 fragment. +- **Control:** identical K, F swept 1 → 10⁴ (arm R naturally produces high F; arm G with + super-batching produces low F for the same logical content). +- **Second sweep:** metadata cache size from 1 GiB (default) down to 16 MiB, to find the F + at which the cache stops absorbing the per-file tail reads. +- **Kill condition:** requests flat in F across the whole sweep. + +**Metrics deliberately NOT claimed by this cell** (they belong to other domains and are +listed so the consolidation can see the gap): *seal CPU & metadata bytes*, *repair read/ +write bytes & helpers touched*, *write amplification of the erasure-coding layer*. This +cell measures only the storage substrate those would sit on. + +--- + +## PRIOR ART + +Searches actually run, with what they returned and how it changed the classification. + +**S1 — `Delta Lake checkpoint quadratic metadata growth full transaction log replay small +files` (WebSearch).** Returned Databricks' transaction-log article, Denny Lee on +checkpoints, Capital One, Conduktor. Established: Delta writes *deltas* and periodically +*checkpoints* (default every 10 commits) precisely because replaying the log is expensive; +a checkpoint at version 10 stores 38 live files rather than 62 historical operations. +**Effect on my classification:** F6 is **KNOWN** as a problem class. Lance's design is the +*dual* — always a full manifest, so read cost is bounded and write cost is not. That dual +framing is worth stating but is not novel. + +**S2 — `Apache Iceberg orphan files two-phase commit prepared data files manifest +publication cleanup race` (WebSearch).** Returned Dremio's commit/ACID walkthrough, IOMETE's +maintenance runbook and their "why we rebuilt orphan cleanup" post, iceberglakehouse.com. +Established: orphan files from write-then-fail are a standard failure mode; a **3-day safety +window** is the standard mitigation; running orphan cleanup with a retention shorter than +the longest write "can corrupt the table by deleting in-progress files"; and there is a +canonical safe ordering (expire snapshots → remove orphans → rewrite manifests). +**Effect:** F8(c) is **KNOWN** in the general lakehouse sense — Lance's 7-day +`UNVERIFIED_THRESHOLD_DAYS` is the same idea with a different constant. What is *not* +covered by this prior art is F8(a), the `ReserveFragments`-is-itself-a-commit finding, +which is Lance-specific. + +**S3 — `LSM-tree compaction key-range grouping versus insertion-order grouping locality +columnar lakehouse Z-order clustering` (WebSearch).** Returned the 2025 LSM survey +(arXiv 2507.09642), the VLDB compaction design-space paper (Sarkar et al., PVLDB 14), ArceKV, +and an LSM spatial-index paper. Established: LSM compaction merges **sorted runs by key**, +which is exactly the property Lance compaction lacks — LSM compaction *is* a locality +repair because the merge is key-ordered by construction. +**Effect:** this is the sharpest contrast in the report and it makes F1 non-obvious: a +reader who reasons by analogy from LSM will *assume* compaction restores order. In Lance it +does the opposite. + +**S4 — `"compaction" data lake does not reorder rows preserves write order clustering must +be explicit ZORDER OPTIMIZE separate from compaction` (WebSearch).** Returned Delta's +optimization docs, Microsoft Fabric Z-order, IOMETE's Z-order-during-compaction post, +delta.io's Z-order article. Established plainly: *"OPTIMIZE does not automatically apply +ZORDER"*; *"Never run Z-Ordering without OPTIMIZE"*; they are combined into one command +only because both are expensive shuffle-and-rewrite operations. +**Effect:** "compaction ≠ clustering" is firmly **KNOWN**. Therefore F1's *first half* +(Lance compaction does not cluster) is **KNOWN**. Its *second half* — that Lance's +`IndexRemapMode::Compact` structure makes reordering a *correctness* violation rather than +merely an unimplemented feature — did **not** appear in any result, and is classified +NOVEL-CANDIDATE below. + +**S5 — `secondary index remap row identity compaction "order preserving" constraint prevents +reordering rewrite rowid mapping compressed` (WebSearch).** Returned lancedb/lance issue +#1378 ("Remap row ids during compaction"), order-preserving dictionary-compression patents, +GeoWave secondary indexing, the HyPer transactional-compaction paper (arXiv 1208.0224). +Nothing describes a remap *data structure whose compactness is purchased with an +order-preservation assumption that then forbids reordering compaction*. This is the +strongest evidence for the NOVEL-CANDIDATE call on S-1 below, and I record the search +explicitly rather than asserting novelty. + +**S6 — `late materialization threshold heuristic column width filter selectivity read +amplification wrong side of threshold columnar` (WebSearch).** Returned the arrow-rs +late-materialization deep dive (Dec 2025), Abadi's *Materialization Strategies in a +Column-Oriented DBMS* (ICDE 2007), *Selective Late Materialization in Modern Analytical +Databases* (VLDB 2025, Tsinghua), Cloudera's lazy-materialization docs, and flash-optimized +columnar patents with a 1.5 % selectivity threshold. Established: threshold-driven +materialization heuristics are standard and their failure mode ("wrong side of the +threshold") is documented in general. +**Effect:** the *class* of failure in F4 is **KNOWN**. What I could not find is the specific +pattern of a **cloud-vs-local threshold pair that inverts the answer between a developer's +laptop and production** (1000 B vs 10 B, with a 512 B payload sitting between them). That +specific benchmark-validity trap is classified NOVEL-CANDIDATE. + +**S7 — `Lance format compaction fragment adjacency small fragments point lookup latency +object storage request amplification` (WebSearch).** Returned lance.org docs, LanceDB +performance guide, Daft/Ray integration docs, and a LanceDB operations post naming +*"fragment sprawl"*: *"Continuous writes leave behind hundreds of small files (fragments) +that slowly hurt search performance… Cleaning this up means running heavy compaction jobs +that read and rewrite large amounts of data."* +**Effect:** fragment sprawl is **KNOWN** and publicly documented by the vendor. F9's +existence is not a discovery; its *magnitude for this workload* (one fragment per cycle, +forever, with no compaction wired) is the contribution. + +**S8 — alphaXiv `discover_papers` on lakehouse / compaction / Z-order / data skipping / +columnar storage formats.** Returned, most relevantly: +`2504.15247` **Lance: Efficient Random Access in Columnar Storage through Adaptive +Structural Encodings** (LanceDB et al.) — the format's own paper; I read it in full. It +confirms the miniblock/fullzip split at ~16 B measured, matching the source's +`MINIBLOCK_MAX_BYTE_LENGTH_PER_VALUE = 256` gate, and it frames random-access cost as +{read amplification, IOPS, decode}. It is the direct source for the *favourable* finding in +EXPECTED BENEFIT, and it says nothing about cross-fragment or metadata cost — the paper's +scope is deliberately *within a file*, which is exactly the boundary this cell attacks. +`2504.04186` **AutoComp: Automated Data Compaction for Log-Structured Tables in Data Lakes** +(Microsoft / UMD / LinkedIn) — small-file proliferation as a first-class production problem; +prior art for "compaction policy is a decision problem", i.e. F2's binning is a *policy* +choice with known alternatives. +`2602.23289` **Workload-Aware Incremental Reclustering in Cloud Data Warehouses** — the +closest published treatment of "clustering decays and must be incrementally repaired, +separately from file consolidation". This is prior art for the *remedy* implied by F1, and +it reinforces that reclustering is a distinct operation from compaction. +`2608.08639` **Smart Compaction: Predicting Compaction Utility from Lakehouse Table +Metadata** — confirms "threshold-driven compaction decisions" is the current state of the +art, i.e. Lance's `target_rows_per_fragment` heuristic is typical, not anomalous. +`2504.11540` **Pruning in Snowflake** and `2304.05028` **An Empirical Evaluation of Columnar +Storage Formats** — background for the pruning/zone-map framing in F5. + +**Not searched, declared:** I did not search erasure-coding repair-locality literature +(LRC/Azure/Clay codes) — that is another cell's domain, and any claim I made there would be +unanchored. + +--- + +## SURPRISES + +**S-1 — The compact index-remap makes order preservation a *correctness invariant*, so a +locality-repairing compaction is not merely unimplemented in Lance — it is structurally +prohibited in the default-adjacent mode.** +`row_addr_remap.rs:35-45` states it outright: the compact remap "does not store each +old-to-new row mapping… **This requires the reader-to-writer pipeline to preserve row +order**", and `GroupRemap::new` (`:139-160`) errors on non-ascending new fragment ids. The +system's *memory saving on the remap* is purchased with *the ability to ever reorder*. The +escape hatches are worse: `Direct` mode costs ~29–57 B per rewritten row at the coordinator +(F3), and stable row ids remove the remap but also remove the `FragReuseGroup` witness +(`optimize.rs:697-706`, `:2013`). Search S5 found no prior description of this specific +coupling. +→ **NOVEL-CANDIDATE.** (The weaker statement — "compaction ≠ clustering" — is KNOWN via S4 +and must not be conflated with this.) + +**S-2 — The 512-byte witness payload sits *below* the cloud early-materialization threshold +(1000 B) and *above* the local one (10 B), so a local benchmark of the read path measures +the opposite regime from production.** +`scanner.rs:2350-2375` + `lance-arrow/src/lib.rs:170` + `scanner.rs:2403`. On `file://` the +payload is late-materialized (cheap unindexed filter); on S3 it is read eagerly for the +whole scanned range. The payload width, 512 B, was chosen for the node ABI +(`cycle_sink.rs:112`, `key(16)|edges(16)|value(480)`) with no knowledge of this threshold — +it is coincidence that it lands in the inverting band. Search S6 found the general +"threshold heuristic" class but not this cloud/local inversion as a benchmark-validity trap. +→ **NOVEL-CANDIDATE.** + +**S-3 — "Prepare fragments, then ONE manifest publication" is false for Lance compaction: +`reserve_fragment_ids` is itself a `commit_transaction`.** +`optimize.rs:1541-1573`, called from `commit_compaction` at `optimize.rs:2036-2045`, +identical in column B (`lance-main` `optimize.rs:2120-2132`). Compaction costs at least two +versions, and an abort between them advances `max_fragment_id` permanently with no data +change. The pattern the thesis names is real for the *append* path (`Dataset::append` = one +commit) and false for the *compaction* path — and it is the compaction path where the +"prepare many, publish once" story was doing the argumentative work. +→ **DISPROVED** (as a general claim about Lance; survives only for plain appends). + +**S-4 — Version cleanup and temporal deinterlacing are the same resource, spent twice.** +OGAR `docs/TEMPORAL-TIME-TRAVEL.md` @ `386a6fd8` §2 maps `KnowledgeHorizon(T)` → +`checkout_version(V_ref)` and defines `ANACHRONISTIC` as `row.lance_version > V_ref`. Lance +cleanup (`cleanup.rs:498-514`) deletes exactly the manifests those horizons resolve to. +Every byte of "stale-version retention" the storage domain wants to reclaim is a byte the +temporal domain has already spent. Cross-domain, this means the two hypotheses cannot be +costed independently: the program's metric list already contains both *stale-version +retention bytes* and *version count*, and this cell asserts they are **not** free variables +— they are pinned by whatever temporal window the AWARE/RETRO modes are specified over. +→ **TRANSFER** (retention-vs-time-travel tension is standard in versioned stores; what is +specific here is that a *cognitive* subsystem, not a compliance policy, is what pins the +window). + +**S-5 — The fragment-reuse index is already a durable permutation witness, and turning on +the feature that would make it cheap destroys it.** +The program's discovery mandate asks *"can fragment-reuse maps serve as permutation +witnesses?"* — mechanically, yes: `FragReuseGroup { changed_row_addrs, old_frags, new_frags }` +(`optimize.rs:2141-2146`) is a persisted, versioned, system-indexed old→new address map +(`FRAG_REUSE_INDEX_NAME`). But it exists **only** when `defer_index_remap` is on, which is +**rejected outright** on stable-row-id datasets (`optimize.rs:697-706`), and it is written +all-or-nothing per compaction with an explicit unsoundness note about partial FRIs +(`optimize.rs:2055-2061`). Two concurrent compactions that both produce an FRI conflict +unconditionally (`conflict_resolver.rs:796-801`). So the witness is real but is available +only in the configuration that also forces the expensive remap regime. +Relatedly, the answer to *"can parity-group boundaries double as compaction boundaries?"* is +**no, not freely**: compaction bins are already pinned by three unrelated constraints — +fragment-id adjacency (`optimize.rs:794-797`), index-membership equality +(`optimize.rs:777-790`), and `target_rows_per_fragment` splitting (`split_for_size`, +`optimize.rs:1452-1484`) — none of which aligns with any semantic or parity grouping. A +parity group would be a fourth constraint layered on three that already fight each other. +→ **TRANSFER** (the map exists as a known mechanism; using it as a coding-layer permutation +witness is a different layer — and this cell's contribution is the *negative* result that +its availability is anti-correlated with the cheap configuration). + +--- + +## VERDICT + +**The thesis "random first-come chunk placement is cheap in Lance" is DISPROVED as a +general claim, and survives only in the narrow form: *cheap to write, and cheap for random +take within a single fragment*.** + +What kills it, in order of confidence: + +1. **It is not reversible.** Lance 9.0.0 — and 11.0.0-beta.14 — offer no reordering + compaction, and `IndexRemapMode::Compact` makes order preservation a correctness + contract (`row_addr_remap.rs:35-45`). Placement chosen at write time is *permanent* + short of a full `Overwrite`. "We can compact it later" is the load-bearing assumption + behind "cheap", and it is false. (S-1) +2. **The read path is currently unindexed full scans of a column that cloud storage reads + eagerly.** `cycle_sink.rs:984/1116` filter without any index; `pushdown_scan` is v1-only; + `is_early_field` puts a 512 B column on the eager side on cloud + (`scanner.rs:2350-2375`). The cost is O(total store) per lookup, and a local benchmark + will not show it. (S-2, F5) +3. **Metadata is quadratic and nothing in the shipped arc bounds it.** Full manifest per + commit × one fragment per cycle, with no compaction, no index, and no cleanup wired + anywhere in `lance-graph`. (F6, F7) +4. **The "prepare many, publish once" pattern is false where it matters.** (S-3, F8) + +**What survives and should be protected.** The fullzip structural encoding of the 512-byte +witness (`primitive.rs:3860-3872`) genuinely gives good within-file random access; the +single-writer append with a content-hash idempotency key (`persist_sink.rs:401-430`, +`cycle_sink.rs:200-262`) is a clean, low-conflict commit story; and per-cycle coalescing +(one image row per dirty row, `cycle_sink.rs:105-108`) is a real and correct amortization. +The write-side design is sound. It is the *placement* claim that fails, and it fails on the +read and maintenance sides. + +**The falsifiable core of this cell.** X4 (cloud/local materialization inversion) and X1 +(does compaction repair locality) are the two experiments that decide it. If X1 comes back +showing post-compaction locality within 1.5× of the grid-ordered control, F1 is wrong and +most of this report collapses. If X4 comes back showing cloud bytes-read proportional to +matched rows, the read-path argument collapses. Both are cheap to run and neither has been +run. Until they are, the honest statement is: **the "cheap" claim has never been priced +against the machinery that would price it, because none of that machinery is wired.** + +**One thing I could not settle and am flagging rather than guessing:** the real +bytes-per-fragment `b` in the manifest protobuf for this 13-column schema. Every quadratic +number in F6 is `b·N²/2` with `b ≈ 100–140 B` inferred from `protos/table.proto:313-400`. +That is an **estimate, not a measurement**, and X5's first deliverable is to replace it. diff --git a/docs/lotus/rp-seal-v1/A3.md b/docs/lotus/rp-seal-v1/A3.md new file mode 100644 index 00000000..f8907d0f --- /dev/null +++ b/docs/lotus/rp-seal-v1/A3.md @@ -0,0 +1,511 @@ +# A3 — Lance upstream evolution beyond 9.0.0: extension seams for locality-aware compaction + +Role: PRIOR-ART / UPSTREAM SCOUT. Cell A3, Domain A (Lance storage/compaction). + +## SOURCE ARCHAEOLOGY + +**Trees compared** (both shallow single-commit clones, so `git log`/`git diff` across +them is unavailable; all comparisons below are plain-file `diff` against matched +paths): + +- Column A (pinned baseline): `/tmp/sources/lance-9`, commit + `7653c2065d6951e0ef236ba9587c6832f7456947`, tag `v9.0.0`, `Cargo.toml` version + `9.0.0`, committed 2026-07-24. +- Column B (current upstream): `/tmp/sources/lance-main`, commit + `a8e2a7857d3a95ac71446024fbff9c61a7fd51a9`, `main`/`origin/HEAD`, `Cargo.toml` + version `11.0.0-beta.14`, cloned 2026-08-18. So the window covered is roughly + three and a half weeks and two major-version bumps' worth of upstream churn + (9 → 10 → 11-beta), which is consistent with Lance's fast release cadence. + +**Method.** Top-level crate layout is identical between the two trees (`ls +rust/` diff is empty — no crate added or removed). I therefore diffed matched +files/dirs directly rather than trying to reconstruct history: + +- `rust/lance/src/dataset/optimize.rs` (compaction core): 8,509 → 9,786 lines + (+1,277). Full diff captured to `optimize_diff.txt` (2,592 diff lines). +- `rust/lance/src/dataset/optimize/binary_copy.rs`: 602 → 441 lines (net + shrink; refactor, not feature removal — see MECHANISM). +- `rust/lance/src/dataset/optimize/remapping.rs`: 568 → 597 lines. +- `rust/lance/src/dataset/transaction.rs`: 6,733 → 9,607 lines (+2,874). Diffed + to `transaction_diff.txt`. +- `rust/lance/src/dataset/cleanup.rs`: 4,334 → 4,832 lines (+498). Diffed to + `cleanup_diff.txt` — turned out to be almost entirely new tests plus overlay + file retention (`lineage_retention_covers_inherited_overlay_files`, + `keep_set_covers_referenced_overlay_files`) and empty-index-directory + cleanup; no new *public* surface (`grep '^+.*pub fn\|pub struct\|pub enum'` + on the diff returns nothing). +- `rust/lance/src/dataset/fragment.rs`: 6,238 → 6,825 (+587). +- `rust/lance/src/dataset/rowids.rs`: 760 → 1,090 (+330); `dataset/rowids/` + gained a `validate.rs` submodule not present in 9.0.0; `dataset/versions/` + (a new `mod.rs`) also appeared. +- `rust/lance-table/src/io/commit/external_manifest.rs`: 728 → 1,704 lines + (+976, more than doubled). +- `rust/lance-core/src/utils/row_addr_remap.rs`: **byte-identical**, 413 → 413 + lines, `diff -q` empty. +- `rust/lance-index/src/frag_reuse.rs`, `rust/lance/src/index/frag_reuse.rs`, + `rust/lance/src/dataset/index/frag_reuse.rs`, + `rust/lance-table/src/system_index/frag_reuse.rs`: **all four + byte-identical**, unchanged. +- `rust/lance-index/src/scalar/rtree.rs`: 1,325 → 1,637 (+312, one new + function `merge_rtree_indices`); `rtree/sort/hilbert_sort.rs`: + byte-identical. +- Protobuf: `protos/table.proto` and `protos/transaction.proto` differ (new + feature flag 128 `FLAG_MEM_WAL_INDEX_CATCHUP`, `MergedGeneration` → + `CompactedSsTable` / `FlushedGeneration` → `SsTable` rename + semantics + hardening for the MemWAL subsystem). `protos/file2.proto`, + `protos/encodings_v2_1.proto`, `protos/index_old.proto` also differ but were + not investigated in depth (out of scope for this cell — no compaction/GC + surface referenced them). +- `docs/`: **file listing is identical minus one PNG rename** + (`mem_wal_regional.png` → `mem_wal_shard.png`); every markdown page that + exists in 9.0.0 exists verbatim-named in main, and the two format-versioning + docs that were spot-checked byte-diffed cleanly (`row_id_lineage.md` + identical; `data_overlay_file.md` identical; `versioning.md` differs by + exactly one new feature-flag table row). **This means every new + compaction-relevant capability found below (`excluded_fragment_ids`, + `max_source_rows`, `max_source_bytes`) shipped in code and in the Python + binding docstrings but has no dedicated prose page yet** — i.e. it is real, + tested, public API, just undocumented at the format-spec level, which is + itself useful signal: these are recent, still-settling additions rather + than long-stable format concepts. +- Confirmed by direct `grep` across the whole `lance-main` tree (excluding + `target/`): **zero occurrences** of `reed-solomon`, `erasure`, `galois`, or + `parity group`/`parity_group` in any `.rs` file. Lance's own "repair" + vocabulary (`grep -rn '\brepair\b'`) is entirely about **metadata** + self-healing — the external-manifest staging/finalize protocol and the + MemWAL index-catchup protocol — never about reconstructing data bytes from + redundancy. This is a load-bearing negative result for the whole 15-person + program: **Lance today delegates all data-durability/erasure concerns to + the underlying object store; there is no internal redundancy-encoding layer + to extend, patch, or conflict with.** Any RS-seal proposal is necessarily a + new layer, not a modification of something that exists. +- `python/python/lance/optimize.py`, `python/python/tests/test_optimize.py`, + and `python/src/dataset/optimize.rs` all reference `excluded_fragment_ids` + and `max_source_bytes` — confirms these are public, bound-and-tested API, + not internal-only scaffolding. +- Two `WebSearch` calls were run to ground the cross-domain questions in + established literature (see PRIOR ART). + +## MECHANISM + +The concrete, file-and-line-anchored deltas found (all in `lance-main`, +`rust/lance/src/dataset/optimize.rs` unless noted): + +**Genuinely new since 9.0.0** (absent by `grep -c` in `lance-9`'s +`optimize.rs`, confirmed zero hits for `max_source_rows`, `max_source_bytes`, +`excluded_fragment_ids`; `max_source_fragments` was already present in +9.0.0): + +1. `CompactionOptions::excluded_fragment_ids: Vec` (field ~294, doc + comment ~289–295): *"Excluded fragments act as boundaries between adjacent + compaction candidates, so fragments on opposite sides of an exclusion are + never combined into the same task."* Implemented in + `DefaultCompactionPlanner::plan` (~745–765): the fragment stream is mapped + to `Option<(Fragment, Metrics)>`, `None` for an excluded id, and the + `while let Some(res) = ...` loop explicitly flushes (`candidate_bins.push`) + the current adjacency bin whenever it sees `None`, with the comment + *"Exclusions preserve adjacency semantics: they terminate the current bin + instead of allowing candidates on either side to be planned together."* + This is — almost verbatim — the "optional group boundaries" seam named in + my brief, and it is **already shipped, tested** + (`test_excluded_fragments_are_planning_boundaries`), and **already public** + through the Python binding. +2. `CompactionOptions::max_source_rows: Option` and + `max_source_bytes: Option` (fields ~266–283), joining the + pre-existing `max_source_fragments`. Enforced by a new function + `limit_tasks_to_source_budget` (~993–1050): walks the already-planned task + list in order, accumulates `total_fragments`/`total_rows`/`total_bytes`, + and truncates the plan the moment any configured budget would be exceeded + — explicitly for **incremental compaction** ("compact 20 fragments at a + time" per the doc comment), not for locality. `task_source_bytes` (~1063) + sums `DataFile::file_size_bytes` (a value **already cached in the manifest + since ≤9.0.0** — confirmed unchanged in `lance-table/src/format/fragment.rs`) + over a task's base files *and* overlay files, restricted to fields still + present in the current schema (dropped-column files are skipped). Row/byte + sizes are read from manifest metadata, never from an extra object-store + HEAD — the function's own doc comment states this is deliberate: an + HTTP round trip per file "would turn planning into one round trip per + file." Tested by `test_max_source_rows`, `test_max_source_bytes`. +3. Nested/struct-aware Blob v2 rewriting during compaction + (`BlobV2FieldRewritePlan`, `transform_blob_v2_list_array`, + `BlobV2BatchRewritePlan`) — a real capability expansion (Blob v2 fields + nested inside structs are now correctly rewritten during compaction, + whereas 9.0.0's `BlobV2Descriptor` machinery only handled top-level blob + columns), but orthogonal to locality/erasure and not pursued further here. +4. `rust/lance-index/src/scalar/rtree.rs::merge_rtree_indices` — a new + function to merge multiple R-tree scalar indices (relevant when + spatially-indexed fragments are compacted together); tangential to my + scope but consistent with the general direction of "index maintenance + keeps pace with compaction," see PRIOR ART item P6 below for the sort + machinery it merges. +5. MemWAL protobuf hardening: `FLAG_MEM_WAL_INDEX_CATCHUP` (bit 128, new), + `FlushedGeneration`→`SsTable`, `MergedGeneration`→`CompactedSsTable`, + with an explicit safety-semantics rewrite: **absence in + `IndexCatchupProgress.caught_up_generations` now means "unknown, must be + treated as not-caught-up and repaired" on a table with the new feature + bit, rather than "assumed fully caught up"** (the old, silently-unsafe + default) on legacy tables. `transaction.rs` grew ~2,874 lines almost + entirely to support this (the `LogicalIndexSegments`, + `apply_mem_wal_index_coverage`, `withdraw_coverage_invalidated_after_build`, + `require_index_catchup` family, plus ~80 new unit tests with names like + `an_index_short_of_the_read_version_is_not_credited`, + `credit_never_reaches_past_the_read_version`, + `carried_positions_do_not_leak_between_shards`). +6. `external_manifest.rs` more than doubled (728→1,704 lines) rewriting and + hardening the staging→atomic-select→materialize→best-effort-repair commit + protocol for object stores without conditional PUT, with an explicit + "correctness model" doc block (4-step protocol, ETag-is-not-identity + discipline, idempotent helper completion by any reader). + +**Already present since ≤9.0.0** (unchanged, or present-but-renamed; +important as *prior art the program can build on* even though it is not new +churn): + +7. `IndexRemapMode::{Compact, Direct}` + `lance_core::utils::row_addr_remap` + (file byte-identical across both trees). `Direct` materializes a full + `HashMap>` (O(rows) memory, O(1) lookup); `Compact` + builds a `RoaringBitmap`/rank-based structure keyed per rewrite-group + (`GroupInput { rewritten_old_row_addrs, old_frag_ids, new_frags }`), + O(#fragments) memory, each lookup paying a bitmap-rank plus one binary + search over new-fragment ranges. This is a **caller-selectable, already + generalized answer to "index remap bytes vs. CPU"** — it directly answers + one of the program's core measurement axes, in a form any GridLake-style + scheme could reuse verbatim for its own auxiliary row-address-keyed + structures (see seam G4 below). +8. `defer_index_remap: bool` + the fragment-reuse system index + (`FRAG_REUSE_INDEX_NAME`, byte-identical `frag_reuse.rs` across four + files/crates). When set, compaction proceeds without touching secondary + indices at all; instead it records enough information in a dedicated + system index (`FragReuseGroup`) for indices to be remapped **later**, out + of the hot commit path — rejected outright for datasets with stable row + IDs, since those never need remapping (`"stable row IDs do not require + index remapping during compaction, so there is nothing to defer"`). This + is a working, shipped instance of "deferred incremental locality repair" — + deferring costly reconciliation work into a queue a background process + drains — for the index-freshness problem specifically. +9. `CompactionMode::{Reencode, TryBinaryCopy, ForceBinaryCopy}` + + `can_use_binary_copy_current`/`rewrite_files_binary_copy` + (`optimize/binary_copy.rs`). When source fragments share an exact file + version, identical field/column-index layout, no deletion files, and no + extra global buffers, compaction **copies page and column-buffer bytes + directly** rather than decode-then-reencode, re-generating only the + footer. The main-branch refactor (602→441 lines) pushed v2.0-vs-v2.1+ + structural-header handling behind a new `ConcreteFileVersion` / + `file_versions::{create_lazy_writer, finalize_external_metadata_column, + data_file_columns, physical_column_count}` abstraction — i.e. Lance is + actively generalizing "which concrete file-format dialect governs this + copy" into a pluggable dispatch point, which is exactly the kind of + abstraction layer a "permute while copying, don't reencode" extension + would need to hook. Today, however, binary copy is strictly **file + concatenation, not permutation**: rows are laid down in read order, never + reordered. +10. **Stable row IDs** (`FLAG_STABLE_ROW_IDS`, present ≤9.0.0, docs + byte-identical). `row_id_lineage.md` is explicit that this already + decouples logical row identity from physical row address for base-table + reads (`_rowid` stays constant across compaction/updates), **but**: + *"Row address is currently the primary form of identifier used for + indexing purposes… Work to support stable row IDs in indices is in + progress."* This is the crux fact for the whole compaction/erasure-repair + coupling question: **the entire reason compaction needs `IndexRemapper`/ + `RowAddrRemap`/`defer_index_remap` machinery at all is that secondary + indices are keyed by physical row *address*, not by the stable logical + row *id* that already exists.** No evidence found that this gap closed + between 9.0.0 and main (no `StableRowId`/`uses_stable_row_ids` hits + inside `lance-index`'s own source). +11. `IndexRemapper`/`IndexRemapperOptions` trait pair + (`optimize/remapping.rs`, trait position unchanged since ≤9.0.0; only the + `create_remapper` factory method gained a `&Dataset` parameter). Doc + comment: *"When compaction runs the row ids will change… The details of + how this happens are not a part of the compaction process and so a trait + is defined here to allow for **inversion of control**."* Concretely: + `async fn remap_indices(&self, index_map: RowAddrRemap, affected_fragment_ids: + &[u64]) -> Result>`. **This is a today-usable public + extension point**: any external consumer (a hypothetical RS + parity-locality index, or GridLake's own auxiliary structures) can + implement `IndexRemapper` and receive the exact old→new row-address + remap and the affected fragment set on every compaction commit, with zero + upstream code change required. `IgnoreRemap` is the shipped no-op + implementation. +12. `rust/lance-index/src/scalar/rtree/sort/hilbert_sort.rs` (byte-identical + across both trees, so present ≤9.0.0): the R-tree scalar index sorts 2-D + geometry rows by a **Hilbert space-filling curve** value + (`hilbert_curve(x, y)`, 16-bit-per-axis quantization) before paging them + into fixed-size (`DEFAULT_RTREE_PAGE_SIZE`) leaves. Docs + (`docs/src/format/index/scalar/rtree.md`, unchanged): *"Hilbert sorting + imposes a linear order on 2D items using a space-filling Hilbert curve to + maximize locality in both axes. This improves leaf clustering, which + benefits query pruning."* **This is a fully worked, shipped example of + SFC-based physical clustering inside Lance** — it is scoped to one + index's leaf layout over 2-D geometry values, not to base-fragment + physical row order, but it is direct proof the pattern ("sort by an SFC + key, then cut fixed-size pages") is already idiomatic Lance engineering, + not an alien concept a compaction proposal would be introducing cold. +13. Data overlay files (`FLAG_UNSTABLE_DATA_OVERLAY_FILES`, bit 64, + experimental, docs byte-identical across both trees — present ≤9.0.0). + Cell-level sparse updates keyed by `committed_version` (the dataset + version at which the overlay became effective, not the version it was + read from — chosen specifically so overlay-vs-index staleness comparisons + are correct under commit retries). The doc's own "Versioning and + ordering" section states: *"the gap between an overlay's + `committed_version` and an index's `dataset_version`... is a staleness + measure **the compaction scheduler can use**."* — Lance's own design docs + already name, but do not yet wire up, exactly a locality/staleness-debt + trigger concept. +14. `CompactionOptions::max_overlays_per_fragment: Option` (default + `Some(10)`, present ≤9.0.0, unchanged): when a fragment's overlay *count* + exceeds the threshold it is force-compacted regardless of size/deletion + ratio. This is a real, shipped, **coarse debt trigger** — just counting + overlays, not measuring the staleness/locality signal item 13 names. +15. `CompactionMetrics { fragments_removed, fragments_added, files_removed, + files_added }` (unchanged struct, both trees) — deliberately thin. No + byte counts, no row counts, no timing, no locality/entropy statistic of + any kind is reported today. + +## EXPECTED BENEFIT + +For the program's actual research question (RS-seal + locality-aware +compaction over versioned Lance datasets), the benefit of this cell's finding +is **de-risking and re-scoping**, not a new mechanism of its own: + +- Several of the seams the program hypothesized as "would need to be + invented" are **already shipped, generic, tested, and Python-bound** + (`excluded_fragment_ids`, size-bounded incremental compaction). A + GridLake-flavored proposal riding on these gets them for free today and + should say so rather than re-inventing them. +- The single largest *missing* piece — a caller-supplied ordering/grouping + key that changes physical row layout during a rewrite, the way Delta + Lake's `OPTIMIZE ... ZORDER BY` or its own R-tree Hilbert-sort already do + — is a well-scoped, precedented, backward-compatible, purely additive + extension (see seam G1 below), not a format-breaking change. +- The `IndexRemapper` inversion-of-control point means a first experimental + implementation of "keep an external locality/parity structure in sync with + physical row movement" **needs no upstream PR at all** to be tried; that + should be the first experiment run, before any upstream proposal is + drafted, because it is free and immediately falsifiable. + +## EXPECTED FAILURE + +- `excluded_fragment_ids` only prevents fragments from being *combined*; it + says nothing about *ordering* which surviving bins get scheduled next, so + it cannot by itself express "compact this locality shard now because its + debt crossed a threshold" — that needs seam G3. +- `max_source_bytes`/`max_source_rows` bound a **single planning run's** + input size for throughput/incrementality reasons; they are not a locality + concept and would not, by themselves, make compaction cluster related rows + — conflating "bounded batch size" with "locality-aware batching" in a + proposal would be a real category error worth flagging to the other + domains. +- Binary copy's compatibility check (`can_use_binary_copy_current`) rejects + *any* fragment with a deletion file. An RS-seal design that wants + page-level copy-without-reencode during a locality repack will collide + with ordinary delete traffic far more often than a naive read suggests — + worth an explicit experiment (does real workload delete-rate make binary + copy inapplicable often enough to matter?). +- The MemWAL "absence means unknown, must repair" hardening (item 5) is a + **metadata**-catchup safety pattern, not a data-repair pattern — reusing + its vocabulary for an RS-repair trigger is a structural analogy, not a + code reuse opportunity; conflating the two in a write-up would overstate + how much is actually shared. + +## EXPERIMENT + +Four seams, pre-registered per the program's metric list. All are additive +(no format-breaking field removal/reinterpretation), so backward +compatibility is "old readers/writers ignore an unset field," identical to +how `excluded_fragment_ids`/`max_source_*` themselves shipped (default +`None`/empty ⇒ old behavior byte-for-byte). + +**G1 — caller-provided sortable layout key for compaction rewrite.** +Add `CompactionOptions::layout_expr: Option` (DataFusion is +already a `lance` dependency via the datafusion-backed scanner used inside +`compact_files`), consumed by the rewrite stream that currently just +concatenates rows in read order (`write_fragments_internal` inside +`optimize.rs`'s task-execution path) as an optional sort stage before +writing new fragments. Generalizes exactly the pattern +`hilbert_sort.rs` already implements for R-tree leaves, one layer up, to +ordinary compaction output — the same relationship Delta Lake's +`OPTIMIZE...ZORDER BY` has to plain `OPTIMIZE`. + - *Metric*: fragment/neighborhood-take latency before vs. after a + compaction run with `layout_expr` set to a synthetic 2-D Morton/Hilbert + key over two correlated columns, on a dataset seeded with random + insertion order. + - *Baseline*: `layout_expr = None` (today's behavior, insertion-order + preserving compaction — `optimize.rs`'s own doc comment: *"This method + tries to preserve the insertion order of rows in the dataset"*). + - *Control*: compact with a **wrong/adversarial** key (e.g. reverse + insertion order) — must not improve neighborhood-take latency, else the + benchmark is measuring "any reshuffle helps" rather than "this key + helps." + - *Kill condition*: if sorted-write throughput (rows/sec written during + compaction) regresses by more than binary-copy's own foregone speedup + (i.e., losing binary-copy eligibility plus paying for a sort costs more + than the read-side win), the seam is not worth landing as-is and needs a + "binary-copy-compatible permutation" fallback instead (see G4). + - *Convincing-a-maintainer benchmark*: reproduce Delta Lake's own published + Z-order data-skipping win (min/max stats narrow, files-touched-per-query + drops) on an equivalent Lance scalar-index scan, citing `rtree.rs`'s + existing Hilbert-sort as in-repo precedent for the mechanism being + sound, not novel machinery. + +**G2 — locality/I-O statistics in `CompactionMetrics`.** +Add `bytes_read`, `bytes_written`, `rows_rewritten` (all cheap: already +computed internally per task, just never surfaced) as additive fields on +`CompactionMetrics`/`RewriteResult`; defaults preserve existing serialized +shape (`#[serde(default)]`, the same idiom `excluded_fragment_ids` used). + - *Metric*: reviewer-facing observability — no runtime cost, pure plumbing. + - *Baseline*: current 4-field struct. + - *Kill condition*: none functional; only risk is API churn cost to + existing consumers of the struct's `Debug`/serialized shape — mitigate + with `#[non_exhaustive]` or a builder, matching how `CompactionOptions` + itself already grew nine fields since inception without a breaking bump. + +**G3 — generalize the overlay-staleness signal (item 13) into an actual +compaction trigger.** +Expose per-fragment "overlay staleness debt" (max gap between any overlay's +`committed_version` and the fragment's base/index version) as a queryable +statistic, and let `CompactionOptions` accept a threshold on it, exactly +parallel to how `max_overlays_per_fragment` already thresholds *count*. + - *Metric*: stale-version retention bytes, read amplification (cells + resolved through N overlay hops before hitting base data) before/after + a debt-triggered compaction pass, vs. a fixed-schedule (cron) compaction + baseline. + - *Baseline*: `max_overlays_per_fragment` (count-only) as today's proxy. + - *Control*: threshold set so high it never fires — must reproduce + baseline exactly (byte-identical `CompactionMetrics`), else the trigger + has a side effect beyond "when to compact." + - *Kill condition*: if staleness-debt triggering does not reduce read + amplification more than the existing count-based trigger for a realistic + overlay-write workload, the added complexity (a second threshold + dimension) is not justified over the shipped one. + +**G4 — stabilize/surface a lightweight "move receipt" independent of +implementing `IndexRemapper`.** +Today the only way to observe *which physical rows moved where* during a +compaction commit is to implement the full `IndexRemapper` trait (async, +index-shaped). Propose a cheap, optional, coarse receipt — `Vec<(old_frag_id, +new_frag_id, RoaringBitmap_of_moved_offsets)>` or simply the `GroupInput` +list already constructed internally — attached to `RewriteResult`/emitted via +a lighter synchronous callback, for consumers (an out-of-process +locality/repair coordinator) who want "what moved" without paying for a full +index-remapper implementation or blocking the commit on it. + - *Metric*: index remap bytes/CPU an external consumer pays to reconstruct + "what moved" via the heavy path (implement `IndexRemapper`, materialize + `RowAddrRemap::Direct`) vs. the light receipt. + - *Baseline*: today's only option — implement `IndexRemapper`. + - *Kill condition*: if the coarse receipt's granularity (fragment-pair + + bitmap) is insufficient for a real consumer's needs (e.g. it needs exact + new offsets, not just "which old offsets moved somewhere in this new + fragment"), the seam must widen to exposing `RowAddrRemap` itself as + a public, serializable type — which is a bigger stability commitment and + should be flagged as such rather than silently scope-crept. + +## PRIOR ART + +**Searches run:** +- Full-tree `grep` (both trees, `.rs`/`.md`/`.py`, excluding `target/`) for: + `z-order|zorder|hilbert|morton|space-filling|space_filling|cluster_by| + liquid.cluster`, `reed.solomon|erasure.cod|erasure_cod|galois|xor.*parity| + parity.*group`, `\brepair\b`, `frag_reuse`, `StableRowId|uses_stable_row_ids` + (scoped to `lance-index`), `file_size_bytes`. +- `WebSearch`: "Delta Lake Z-order OPTIMIZE compaction data skipping liquid + clustering locality" and "HDFS erasure coding placement locality group + repair Reed-Solomon." +- Diff-based structural comparison (no keyword search) of `CompactionOptions`, + `Fragment`/`DataFile`, `RowAddrRemap`, `IndexRemapper`, `CompactionMetrics`, + `protos/table.proto` between the two pinned trees. + +**Findings from those searches:** +- Delta Lake `OPTIMIZE ... ZORDER BY` is the direct, shipped, widely-used + analog of seam G1: a caller-named set of columns is mapped through a Z-order + (Morton) space-filling curve and the compacted output files are physically + sorted by that key, which narrows per-file min/max stats and improves + data-skipping (sources: docs.delta.io/optimizations-oss, + delta.io/blog/2023-06-03-delta-lake-z-order, Databricks' `OPTIMIZE` doc). + Databricks' own newer "liquid clustering" replaces Z-order with an + incremental, tree-based layout that specifically targets uniform file size + and avoids needing a full-table rewrite each time — directly relevant to + the "incremental locality repair" question the program poses, and a strong + hint that a naive "always fully re-sort on every compaction" version of G1 + would be following the *older*, already-superseded design in that + ecosystem; an incremental variant is the more defensible target. +- Locally Repairable Codes (LRC), e.g. HDFS-Xorbas, are the established + answer to "can erasure repair locality and query locality share a + grouping": local parity groups are formed over small subsets of data + chunks specifically so a single-chunk repair reads only its local group + (measured ~2× reduction in repair I/O/network vs. plain Reed-Solomon in the + HDFS-Xorbas literature), with global parity layered on top for larger + failure tolerance, and stripe placement is explicitly constrained to + respect failure-domain boundaries. This is squarely on point for the + program's "can parity-group boundaries double as compaction boundaries?" + question — but it is a technique from **outside** Lance (HDFS/Azure + storage systems), with zero code or vocabulary overlap inside Lance itself + (confirmed by the erasure/parity grep above returning nothing). + +## SURPRISES + +1. **`excluded_fragment_ids` already implements "optional group boundaries" + almost verbatim, and it shipped between 9.0.0 and now, tested and + Python-bound.** — KNOWN (it is exactly the generic seam my brief asked me + to evaluate the plausibility of; here it is plausible because it already + exists). +2. **The entire reason compaction needs index-remap machinery + (`IndexRemapper`, `RowAddrRemap`, `defer_index_remap`) traces to one + documented, still-open gap: secondary indices are keyed by physical row + address, not by the already-existing stable logical row id.** — KNOWN + (stated plainly in Lance's own `row_id_lineage.md`, unchanged since + ≤9.0.0), but worth elevating: it reframes "index remap cost" from an + intrinsic compaction tax into a closable gap, which changes the long-run + cost model any RS-locality proposal should assume. +3. **Lance already ships a working Hilbert-curve SFC clustering mechanism — + just scoped to one scalar index's leaf layout (2-D geometry via + `rtree.rs`/`hilbert_sort.rs`), not to base-fragment physical row order.** + — TRANSFER (the technique — SFC sort then fixed-size paging — is well + known; its presence *inside Lance already*, one layer away from where the + program wants it, was not obviously predictable and de-risks G1 + substantially: the file-format and DataFusion-expression plumbing needed + already exists somewhere in this codebase). +4. **`IndexRemapper` is a today-usable, already-shipped inversion-of-control + extension point that lets an external consumer track exactly which + physical rows moved on every compaction commit, with zero upstream code + change.** — TRANSFER/KNOWN-as-a-pattern (trait-based inversion of control + is unsurprising as a technique), but its exact fit to the program's + "permutation witness" need is a genuine, actionable, low-cost first + experiment nobody needs to wait on an upstream PR to run. +5. **Lance's own design docs (`data_overlay_file.md`, unchanged since + ≤9.0.0) already name a locality/staleness-debt trigger concept — "the + gap between an overlay's committed_version and an index's dataset_version + ... is a staleness measure the compaction scheduler can use" — but nothing + in the code consumes it as a trigger; only a coarser overlay-*count* + threshold (`max_overlays_per_fragment`) is wired up today.** — + NOVEL-CANDIDATE as an actual upstream contribution (closing the gap + between "the docs name this metric" and "the scheduler uses it" appears + to be genuinely unbuilt, based on the byte-identical doc and the absence + of any staleness-threshold field alongside `max_overlays_per_fragment` in + `CompactionOptions`), though the *idea* of version-gap-as-staleness is a + restatement of ordinary MVCC/versioned-storage practice elsewhere, so the + novelty is narrowly in "Lance hasn't wired its own already-stated idea + into a trigger yet," not in the idea itself. + +## VERDICT + +Between `v9.0.0` and current `main` (11.0.0-beta.14), upstream Lance +independently shipped a size/exclusion-bounded incremental compaction +planner (`excluded_fragment_ids`, `max_source_rows`, `max_source_bytes`) that +already realizes two of the six candidate seams named in my brief almost +exactly as specified, with no GridLake-specific semantics required to get +there — strong evidence the *direction* (generic, additive, `Option`-gated +compaction-planning knobs) is one upstream maintainers accept. The genuinely +open extension seam is a caller-supplied physical-layout ordering key at +compaction-rewrite time (my G1): Lance already has this exact mechanism, +scoped to one 2-D scalar index's leaf pages (Hilbert sort in `rtree.rs`), and +Delta Lake ships the base-fragment-level version of it as `OPTIMIZE ... +ZORDER BY`, so proposing it for Lance's own compaction rewrite path is +neither novel nor architecturally alien — it is a precedented generalization +with a clear, additive, backward-compatible shape and an obvious +convince-a-maintainer benchmark (reduced files-touched-per-query via +narrower per-file min/max stats, exactly Delta's own published claim). The +single most important qualifier for the rest of the 15-person program: Lance +has zero internal erasure-coding/redundancy machinery to build on — the +"seal" half of the program's title is entirely new territory relative to +this codebase, while the "compaction/locality" half has more prior art +already shipped, generically, than the program's hypothesis list assumed. diff --git a/docs/lotus/rp-seal-v1/B1.md b/docs/lotus/rp-seal-v1/B1.md new file mode 100644 index 00000000..e9a05b2c --- /dev/null +++ b/docs/lotus/rp-seal-v1/B1.md @@ -0,0 +1,662 @@ +# B1 — Locality-Key Generators for Seal/Compaction Layout (spatial / Morton-cascade domain) + +**Role:** BUILDER (experimental design only — nothing below is implemented or run) +**Scope:** design experimental locality-key generators over the baseline 512B-row / +8KiB-field / 32MiB-cycle geometry; locate and characterize every project-native +Morton/HHTL/facet-tier/CurveRuler/bgz17 mechanism with file:line; specify a +harness (key cost, clustering metrics, query shapes, plug-in points into seal +and compaction); make no claim about the phase-progressive schedule's benefit. + +--- + +## SOURCE ARCHAEOLOGY + +### 0. The baseline geometry maps exactly onto an existing bit-width in this codebase + +512B row → 8KiB field (4×4=16 rows) → 32MiB cycle (4096 fields = 65,536 rows). +The stated 2-D hierarchy `4/16/64/256/1024/4096` is six levels of a fan-out-4 +(quadtree) reduction over the 4096 *fields*; folding in the field's own 4×4 +intra-field sub-grid (2 more fan-out-4 levels) gives **8 quadtree levels total += 4⁸ = 65,536**, i.e. the cycle's row-grid is naturally a **256×256** 2-D +address space (each axis 8 bits). This is not a coincidence I get to exploit +loosely — it is *exactly* the bit width the one general-purpose Morton +primitive already in this workspace produces (see §1). + +### 1. The one real, reusable, domain-agnostic Morton primitive: `FacetTier::morton()` + +`/home/user/lance-graph/crates/lance-graph-contract/src/facet.rs:54-74` + +```rust +pub const fn morton(self) -> u16 { + Self::spread8(self.lo) | (Self::spread8(self.hi) << 1) +} +const fn spread8(x: u8) -> u16 { /* magic-number bit-spread, 4 shift/or/and pairs */ } +``` + +`FacetTier { hi: u8, lo: u8 }` is one 8:8 tile of the 16-byte, `#[repr(C, +align(16))]` `FacetCascade` (`facet.rs:1-106`) — a **content-blind** register: +"the substrate is ALWAYS 8:8; only the CONSUMER projects meaning onto the +bytes" (`facet.rs:4-8`). `morton()` interleaves `lo` on even bits, `hi` on odd +— a textbook Z-order code, `u8 × u8 → u16`. This is the *exact* arithmetic a +256×256 grid's Morton key needs, already written, already unit-tested for +round-trip identity (`facet.rs:640-644`), with **zero caller today that uses +it for physical row placement** — its only consumer is the round-trip test. +The module doc also claims a GFNI AVX-512 acceleration path +(`vgf2p8affineqb`, `facet.rs:25`); grepped `ndarray/src` and the whole +`lance-graph` tree for `pdep|pext|gf2p8affine` — **zero hits outside this one +doc-comment**. The only implementation is the scalar magic-number trick; the +SIMD claim is aspirational, not shipped. + +### 2. `NiblePath` — a 16-ary radix-trie locality key, production-live, not spatial + +`/home/user/lance-graph/crates/lance-graph-contract/src/hhtl.rs` (1041 lines). +A `u64`-packed nibble sequence (`path: u64, depth: u8`, +`hhtl.rs:56-59`), fan-out 16, max depth 16 (`hhtl.rs:39-45`). Rich locality +operators, all `const fn`, O(depth): `is_ancestor_of` (`:176-183`), +`common_prefix_depth` (`:251-268`, "the operational form of `panCAKES ≡ +radix trie ≡ HHTL`"), `family_hop_count` (`:414-417`, explicitly "the +CLAM tree distance… the operator's HHTL CLAM via family-nodes hop count as +adjacency"), `common_ancestor` (`:459-503`). Real production consumer: +`WikidataClass::nibble_path()` → +`/home/user/lance-graph/crates/lance-graph-ontology/src/wikidata_hhtl.rs:66`. +`from_guid_prefix`/`from_guid_prefix_v3` (`hhtl.rs:302-402`) lower a +`NodeGuid`'s `classid‖HEEL‖HIP‖TWIG` prefix into this path directly — so +**every canonical node already carries a 16-ary hierarchical address for +free**, no extra bytes, no extra computation beyond a shift. Its "space" is +an *ontological/classification* hierarchy (or, generically, whatever the +`ClassView` says HEEL/HIP/TWIG mean), not literal x/y — but the operational +contract (le-contract §3, quoted in CLAUDE.md) says the SAME rail mechanism +is bound per-domain to "OSM: literal x/y" for spatial data and to a +"palette256:palette256" centroid pair for semantic data. §3 below shows the +literal-x/y instantiation already exists and was already measured. + +### 3. `weather_poc::key` — the literal-x/y instantiation, ALREADY BUILT and ALREADY MEASURED + +`/home/user/lance-graph/crates/weather-poc/src/key.rs` (348 lines, 4 tests, +all reading against the exhaustive real grid — not sampled). + +``` +key[4] = (lat_idx >> 6) as u8 // HEEL, lat axis ("tile" coordinate) +key[5] = (lon_idx >> 6) as u8 // HEEL, lon axis +key[6] = (lat_idx & 63) as u8 // HIP, lat axis ("within-tile" offset) +key[7] = (lon_idx & 63) as u8 // HIP, lon axis +``` +(`key.rs:100-108`). This is **not interleaved** — four separate bytes, +concatenated axis-major (all of HEEL, then all of HIP), each axis +independently tiled at 64. Sorting on these bytes lexicographically gives +**tile-major order, row-major within each tile** — a real, already-shipped, +already-tested "blocked raster" locality key, over a real 721×1440 (= +1,038,240-cell) ERA5 grid with ragged last tiles on both axes (`key.rs:33-37, +50-60`) and a cyclic-longitude seam handled by `box_ranges` returning **two** +non-wrapping `CellBox`es when a query crosses 0°/360° (`key.rs:154-197, +303-328`) rather than one approximate range. + +**This key costs nothing to compute beyond what a `NodeGuid` already stores.** +`encode_key` does not derive anything new — HEEL/HIP are literally the same +shift/mask split the canonical node key layout performs anyway +(CLAUDE.md's own CANON section: `classid(4) | HEEL(2) | HIP(2) | TWIG(2) | +family(3) | identity(3)`). Sorting by "the row's own key bytes, unmodified" +is therefore a **zero-cost locality key** — it is latent in every row's +identity already. + +### 4. `layout_probe.rs` — the experiment this task is asked to design was already run once, on a coarser variant, and the verdict is a documented KILL + +`/home/user/lance-graph/crates/weather-poc/examples/layout_probe.rs` (1172 +lines) is `D-WXS-2a` half A: three byte-order **arms** over the same four +axis bytes (`layout_probe.rs:1-140`): + +- **SHIPPED** — `[lat_tile, lon_tile, lat_hip, lon_hip]`, `encode_key` + verbatim (`:28-34`). +- **MORTON** — nibble-interleaved *per tier* (HEEL and HIP each + independently), **not** bit-level: "the alternative is BIT-level Morton… + true Z-order… Bit-level Morton was NOT implemented" (`:35-56`, explicit + in the module doc). This is a coarser, 4-bit-granularity Z-order than + `FacetTier::morton()` in §1, which *is* bit-level. +- **CONTROL-BAD** — `[lon_hip, lat_hip, lon_tile, lat_tile]`, SHIPPED + reversed, "deliberately locality-destroying" (`:57-59`). + +Two metrics, both computed exactly (not sampled/approximated) over the whole +real grid: + +1. **`range_count`** (`layout_probe.rs:578-599`) — sort a query box's own + cells by their rank under the arm's key order and count breaks in the + consecutive-integer run. This is *precisely* the literature's + **clustering number** `c(q,π)` (§ PRIOR ART) — same definition, derived + independently in-project. +2. **`neighbour_locality`** (`:601-674`) — median `|rank(a) − rank(b)|` + over the 4-neighbourhood (lat±1 clipped at poles, lon±1 always wraps), + pooled over a deterministic prime-stride sample (`SAMPLE_STRIDE = 677`, + chosen coprime to 64/1440/721 so the sample phase never aliases, + `:603-615`) of 1,534 cells. + +Four query-box kinds: **tile-aligned, non-tile-aligned, seam-crossing, +pole-adjacent** (`layout_probe.rs:295-337, 359-556`). + +**Recorded result** (`/home/user/lance-graph/probes/weather-p1/SUBSTRATE_FORMULA_MATRIX.md` +rows G16–G18, §1c): SHIPPED wins `range_count` on the realistic +non-tile-aligned box (median **140.00** vs Morton's **212.50**); Morton wins +`neighbour_locality` by 2× (**16 vs 32**); CONTROL-BAD loses both by +~20× (a control that can lose, and does). The pre-registered bar required +Morton to win *both* metrics to justify a migration — it split, so **KILL +fired, no migration** (row G17: "Verdict: C — no unambiguous win"). Half B +(the same three arms under a real computational stencil access pattern — a +ζ/vorticity finite-difference stencil, i.e. exactly this program's +"neighborhood-take" shape) is **explicitly not run**, gated on an unrelated +external classid-mint blocker (`D-WXS-9`→`D-WXS-0`), *not* on any limitation +of the key-space math itself (SUBSTRATE_FORMULA_MATRIX.md §7 N5, §5). + +### 5. Two more independent readings of "Morton"/"HHTL cascade," one of them dead code + +- `bgz_tensor::morton_cascade::v3::morton2` + (`/home/user/lance-graph/crates/bgz-tensor/src/morton_cascade/v3.rs:18-25`) + — a 2-bit×2-bit → 4-bit Morton tile, applied to `(i, i>>2) & 0b11` inside + the SPO-tenant scoring loop. Its output is **computed and discarded**: + `let _cell = morton2(...)` (`v3.rs:36`) — the underscore prefix is the + compiler-silencer for "unused"; nothing downstream reads `_cell`. This is + a third, semantic-centroid-pair reading of "Morton cascade" (not spatial), + vestigial as a locality mechanism today. +- `perturbation_sim::hhtl::HhtlKey` + (`/home/user/lance-graph/crates/perturbation-sim/src/hhtl.rs`, 193 lines) + — a **fourth** distinct generator producing the *same address shape* + (`heel: u16, hip: u16, twig: u16`) by **recursive Cheeger/Fiedler spectral + bisection of a graph Laplacian** (electrical-grid bus topology: `bus`, + `susceptance`, `limit` — this crate is a power-grid resilience simulator, + unrelated to storage). Its own doc is explicit about the relationship to + the canon: "the OGAR production form widens each tier to a 16-ary/ + 256-centroid tile… this is the spectral-bisection INSTANCE of the same + address, not that full encoding" (`hhtl.rs:1-13`). Mechanistically this is + the odd one out: not a bit-interleave of coordinates and not an + ontology-routed trie, but a **graph-topology-native** hierarchical + address that minimizes edge cut by construction (§ SURPRISES, § cross- + domain question on RS repair). + +**Duplicate-implementation note (explicitly requested):** the same +"HEEL/HIP/TWIG hierarchical address" *shape* is independently produced by +**four** unrelated mechanisms in this workspace — ontology-routing +(`NiblePath`, §2), literal coordinate tiling (`weather_poc::key`, §3), +semantic-centroid Morton (`bgz_tensor::morton_cascade`, §5, partly dead), +and graph spectral bisection (`perturbation_sim::hhtl`, §5) — with **no +shared trait or interface** unifying them; each was built independently for +its own domain and none cross-references the others as instances of one +pattern. + +### 6. The deterministic phase-progressive schedule — located and characterized, no benefit claim + +`/home/user/lance-graph/crates/helix/src/curve_ruler.rs` (117 lines). +`CurveRuler::index(k) = (start_offset + STRIDE·k) mod MODULUS`, `MODULUS=17`, +`STRIDE=4` (`curve_ruler.rs:16-28,50-56`; constants at +`/home/user/lance-graph/crates/helix/src/constants.rs:47,52`). +`gcd(4,17)=1` so one period visits all 17 residues (a full permutation, +tested at `curve_ruler.rs:90-100`). `from_hhtl(path, depth) = +from_place(path.wrapping_add(depth))`, `from_place(place) = place % 17` +(`:31-43`). **In-project measurement already exists and is directly +relevant to any temptation to feed an HHTL/Morton address into this ruler +as a locality key:** +`/home/user/lance-graph/crates/helix/examples/probe_hhtl_intake_blindness.rs:19-30` +(H1, pre-registered, confirmed) shows the ruler is **structure-blind**: it +collapses `(path, depth)` to one scalar *before* any nibble/level structure +can act, so two different HHTL carvings of the same address produce only a +constant rotation with zero per-cell variance — "the address enters the +ruler as a number, never as a hierarchy." I record this because the task +explicitly asks me to locate it; I make **no claim** that this stride +mechanism helps a compaction/seal locality key — that question is scoped to +Domain C, and the one measurement this workspace already ran on it points +the other way for the *use as an address carrier* (it says nothing about +its use as, e.g., a within-parity-group phase dither, which is a different +question again). + +### 7. `bgz17` — no independent Morton/spatial machinery found + +Checked every file under `/home/user/lance-graph/crates/bgz17/src/` +(`base17.rs`, `palette.rs`, `palette_matrix.rs`, `palette_csr.rs`, …) for +`morton|spread|interleave|hilbert` — only two incidental hits, both +unrelated prose ("spread information" in an RNG comment). `bgz17`'s +locality-relevant contribution is its 256×256 distance/compose table +architecture (cited throughout CLAUDE.md), which is a *value*-space lookup +structure, not a *key*-space (row-placement) locality generator, and is out +of this cell's scope. + +### 8. The write path this key must plug into — three files, one hard constraint + +- `/home/user/lance-graph/crates/lance-graph-planner/src/batch_writer.rs:89-138` + — `BatchWriter::cast()` records intent in strict `CastId` (monotonic) + insertion order, `BTreeMap`-backed. No placement decision anywhere in this + file; it is arrival-order bookkeeping only, and the module doc states + there is **zero production call site** for `cast()` today (`:12-20`) — + the mechanism this key must eventually plug alongside is itself declared + status, not yet a running path. +- `/home/user/lance-graph/crates/lance-graph-planner/src/persist_sink.rs:1-183` + — `order_cycle_stably(rows, key)` (`:121-138`) sorts a cycle's + casts by a **caller-supplied canonical key**, currently always + `stream_position: u64` (`:168-183, 351-378`), which the module doc + states explicitly is a *logical/epistemic* order key, not a physical + one: "64k thoughts run concurrently and finish in ARBITRARY CPU order… + the durable image is already canonical [after this sort]… `WalSink:: + scan_sealed` returns rows in that stored order and **never sorts**: + completion order is fixed to canonical order at WRITE time, not repaired + on every read" (`:25-35`). This is a **hard constraint** discovered by + archaeology, not a hypothesis to test: a locality key **cannot** replace + or perturb this write-side key without breaking the documented + physical-race-safety contract of the WAL append. Any locality-driven + physical reordering must happen **after** the seal, as a separate + compaction pass over already-durable rows. +- `/home/user/lance-graph/crates/lance-graph-planner/src/temporal.rs:1-135` + — `deinterlace()` operates on a *different* axis entirely (`server_id, + lance_version, hlc_tick`, `:17`) over already-materialized in-memory + rows (`:346-385`). This confirms temporal ordering and physical/spatial + placement are orthogonal in this codebase today — good, because it means + a locality key does not have to reconcile with temporal semantics, only + with the write-order constraint above. + +### 9. Lance 9.0.0 (upstream, pinned) — no native SFC, but the remap machinery a locality resort would ride + +`/tmp/sources/lance-9` (exact pinned tag v9.0.0). Grepped +`rust/lance/src/dataset/optimize/` and the whole `rust/lance/src/` tree for +`sort_by|SortColumn|sort_order|zorder|hilbert` — **zero hits** on anything +resembling a native space-filling-curve write/compaction ordering. +`rust/lance/src/dataset/optimize/remapping.rs:1-80` **does** carry the real +mechanism any locality-driven physical resort would use: `IndexRemapper` +trait + `RowAddrRemap` — "when compaction runs the row ids will change… +indices will need to be remapped" (`remapping.rs:47-50`). This is the +concrete, measurable cost center for "index remap bytes/CPU" in the metric +list below: it scales with how many rows a locality resort physically +relocates, not with the key's own compute cost. + +### 10. No erasure-coding implementation exists in-repo + +Grepped `reed.solomon|ReedSolomon|erasure.cod|parity_group|galois` (case- +insensitive) across `/home/user/lance-graph`. The only hits are the excluded +research plan file and one doc-comment ("ECC (Reed-Solomon for error +correction)") in `lance-graph-cognitive/src/fabric/udp_transport.rs:29` — +network FEC framing, not a storage codec, and not implemented there either +(prose only). **The RS/seal side of this program is a clean-slate design +space in this codebase**; nothing here already commits to a parity-group +shape I would be contradicting. + +--- + +## MECHANISM + +### The two coordinate spaces this geometry actually supports + +The 65,536-row cycle supports **two independent "spaces"** a locality key +can be defined over, and the workspace's own architecture (le-contract's +rail-binding doctrine, quoted in CLAUDE.md: "OSM: literal x/y" vs "a +semantic domain binds the same rail to a PQ subspace pair") already says +these are meant to be interchangeable readings of the *same* 8:8 tile +mechanism: + +1. **Coordinate space** — the natural 256×256 quadtree grid implied by the + baseline geometry (§ SOURCE ARCHAEOLOGY 0). Query shapes: axis-aligned + box, point, k-radius neighbourhood, uniform-random take. +2. **Graph space** — each row is a node in the SPO/edge-block adjacency + graph already carried by the canonical node's 16-byte edge block. Query + shapes: k-hop neighbourhood expansion, subject-lookup, random-take. + +A locality key generator is scored **per space**; nothing requires one +generator to serve both (that is itself an open question, § SURPRISES / +cross-domain). + +### Generator catalog (7 arms; #1 and #7 are zero-cost, #2–#6 require design/measurement) + +| id | name | space | formula | key-gen cost | already measured? | +|---|---|---|---|---|---| +| **A1 — IDENTITY-PREFIX ("SHIPPED")** | zero-cost | coord | sort by the row's own existing `classid‖HEEL‖HIP‖TWIG` bytes, verbatim, no transform | **0** (bytes already exist) | YES — `weather_poc::key` + `layout_probe.rs` SHIPPED arm, § SOURCE ARCHAEOLOGY 3–4 | +| **A2 — MORTON-BIT (Z-order)** | coord | `FacetTier::morton()` on `(y_byte, x_byte)` verbatim, `facet.rs:54-74` | ~10 scalar ALU ops/row, O(1), branch-free | NO (only the coarser nibble variant, A3, was ever run) | +| **A3 — MORTON-NIBBLE** | coord | 4-bit-granularity interleave per tier, `layout_probe.rs` MORTON arm | similar to A2, slightly more shuffling | YES — split verdict, KILL'd for migration (§ SOURCE ARCHAEOLOGY 4) | +| **A4 — HILBERT** | coord | standard `xy2d` bit-loop (Skilling/Wikipedia form): 8 iterations over a 256-side grid, each a compare + conditional rotate/reflect (XOR-swap) of the running `(x,y)` state | O(log₂ side) = 8 iterations, **data-dependent branches per bit** (no closed-form bit-spread trick exists for Hilbert) | NO — not implemented anywhere in this workspace | +| **A5 — NiblePath-ANCESTRY** | graph | sort by the packed `u64` `NiblePath.path` (`hhtl.rs:56-59`), treating SPO subject/predicate/object identity (or classid tiers) as the hierarchy | O(1) — value already exists per row via `from_guid_prefix*` | NO | +| **A6 — CHEEGER-BISECTION** | graph | recursive Fiedler-vector bisection of the row-adjacency Laplacian, generalized from `perturbation_sim::hhtl::HhtlKey` (`hhtl.rs` in that crate) to the SPO edge block | **O(N·k)** to build (k = iterative eigensolver steps per split, over ~`log₂(65536)=16` levels) — orders of magnitude more expensive than A1–A5, and re-derivable only when the graph structure changes | NO | +| **A7 — RANDOM CONTROL** | either | `splitmix64(NodeGuid bytes)`, low 16 bits as rank | O(1), one avalanche hash | NO, but see the finding below: it is arguably what the write path already produces | + +**A finding that reframes the whole "control" question:** `persist_sink.rs` +documents that a cycle's rows land in `stream_position`-order only *after* +"finishing in ARBITRARY CPU order" from 65,536 concurrently completing +producers (`persist_sink.rs:27-30`). `stream_position` is a logical +completion-order counter, not a function of any row's spatial or graph +coordinates. So **today's actual write-time order is, with respect to +either coordinate space above, indistinguishable in kind from A7 (random +control)** — not a designed baseline, but the *status quo*. This means A1 +(cost **zero**, reusing bytes the row already carries) is not merely "the +cheap option" in this catalog — it is a strict, free improvement over +what the write path does today, before any Morton/Hilbert/spectral +machinery is even considered. Framed as a kill condition below. + +### How a locality key plugs into (a) seal/parity grouping and (b) a compaction sort key + +**(b) Compaction sort key — the safe, well-precedented seam.** Given the +hard write-side constraint in § SOURCE ARCHAEOLOGY 8 (`WalSink::scan_sealed` +"never sorts" — the seal's row order is fixed at WAL-append time and is an +epistemic invariant, not a layout choice), a locality key **cannot** be +`order_cycle_stably`'s key. It plugs in exactly where Delta Lake's/Iceberg's +Z-ORDER plugs in (§ PRIOR ART): as the row-ordering **policy** fed into a +periodic, explicit background compaction pass — mechanically, into Lance's +own `IndexRemapper`/`RowAddrRemap` machinery (`remapping.rs:47-80`), which +already exists for exactly this purpose ("when compaction runs the row ids +will change"). No new Lance-side primitive is needed; only a new sort-key +function supplied to an existing remap call. This is the same seam +`layout_probe.rs`'s own gated "half B" (the ζ-stencil access pattern) was +scoped for, generalized from the 721×1440 ERA5 grid to this program's +256×256/65,536-row cycle geometry — and *unblocked* here, because (unlike +the weather bake) this experiment needs no classid mint: `encode_key`-style +functions already take `classid` as a caller-supplied parameter +(`key.rs:90`), so a synthetic/placeholder classid suffices for a pure +key-space probe exactly as it did in `layout_probe.rs`. + +**(a) Seal/parity grouping — a genuinely open design choice.** RS +protection is stated (program premise) to be defined over canonical LOGICAL +bytes, and no RS implementation exists in-repo to constrain this (§ SOURCE +ARCHAEOLOGY 10). Two distinct group-assignment policies are worth +distinguishing and measuring separately, because they answer different +questions: +- **Physical-contiguity groups** — a parity stripe = N *physically + consecutive* rows in the sealed (`stream_position`-then-possibly- + compacted) image. Cheap, but only as good as whatever order the image is + already in. +- **Logical/declustered groups (key-value groups)** — a parity stripe = all + rows whose locality key falls in one *quadtree cell* (e.g. one Morton + quadrant, or one `NiblePath` prefix at a fixed depth), **independent of + physical position**. This is the "logical group, physical scatter" shape + well known in declustered-parity RAID literature, generalized here to a + key-derived group rather than a hash-derived one. + +If a *background compaction pass* later physically resorts rows by the +**same** key used to define logical RS groups, the two collapse into one: +after compaction, an RS stripe is automatically physically contiguous too, +and "repair" and "range read" scatter become the same quantity. This is +this program's own cross-domain conjecture ("can a layout key derived from +identity make physical canonicalization constructive rather than a repair +sort?") — MECHANISM makes it *coherent*, not *true*; § EXPERIMENT below is +what would test it, and I found no prior art directly proposing this +specific coupling (§ PRIOR ART, § SURPRISES). + +--- + +## EXPECTED BENEFIT + +- **A1 (identity-prefix)**: strictly free improvement over the documented + arrival-order status quo for *any* subsequent range or neighbourhood + read, because it is not "no locality key" — it is "the cheapest + non-degenerate locality key," and the write path currently ships + something closer to A7. +- **A2 (bit-level Morton, reusing `FacetTier::morton()`)**: expected to + *dominate* A3 (nibble-level Morton) on both metrics — finer interleave + granularity should never *increase* worst-case run-fragmentation for a + box not aligned to the coarse nibble boundary, and A3's clustering + number on non-tile-aligned boxes (212.50) already lost to A1's 140.00, so + the open question is whether A2 closes that gap or only improves + `neighbour_locality` further while still losing `range_count`. +- **A4 (Hilbert)**: per the clustering-number literature (§ PRIOR ART), + expected to be the only coordinate-space candidate capable of beating A1 + on **both** metrics simultaneously — Z-order variants (A2, A3) are + well-documented to have a strictly worse worst-case clustering number + than Hilbert in 2-D. +- **A5/A6 (graph-space)**: expected to dominate on k-hop/neighbourhood-take + latency and on RS-repair scatter for graph-shaped access, where A1–A4 + (coordinate-only) have no reason to correlate with graph adjacency at + all. + +## EXPECTED FAILURE + +- **A2 vs A1 on range_count**: A2 could still lose to A1's `range_count` + the way A3 did — Morton's Z-shaped discontinuities (a query box straddling + a big power-of-two boundary can still fragment badly even at bit-level + granularity) are a *known* pathology, not fixed by finer granularity + alone. If measured, this would **replicate**, not merely extend, the + already-KILL'd A3 finding. +- **A4 (Hilbert)**: real-world implementations frequently do not realize + the asymptotic clustering-number advantage at the small `N`/coarse-tile + scale this program's per-field (16-row) or per-cycle (65,536-row) + granularity actually operates at — the clustering-number bounds in the + literature are asymptotic in the SFC's recursion depth, and 8 levels + (256×256) is modest. +- **A6 (Cheeger-bisection)**: the eigensolve cost may dominate every other + metric in the program's list (compaction CPU/wall) badly enough that any + query-locality or repair-locality win is not worth it relative to A1's + zero-cost baseline — this is a real risk given perturbation-sim's own + framing of Cheeger bisection as an offline, one-shot topology analysis, + not a per-cycle relayout operation. +- **The joint seal+compaction key (§ MECHANISM (a))**: may simply not + transfer — `layout_probe.rs`'s own finding (a key that wins one locality + metric loses another, on the SAME underlying grid) is a direct precedent + for "no single key serves two distinct locality demands (query-range vs + neighbour-adjacency) equally well," and RS-repair locality is a *third* + demand with no guarantee it aligns with either. + +--- + +## EXPERIMENT + +All designs below are pre-registered: metric, baseline, control, kill +condition, stated before any run. Costs and clustering numbers are +UNMEASURED unless flagged "already measured" (only A1/A3, § SOURCE +ARCHAEOLOGY 4). Nothing here is implemented. + +### Harness shape + +- **Corpus**: the synthetic 256×256 row-grid (coordinate space) and a + synthetic SPO adjacency graph over the same 65,536 rows (graph space, + degree distribution drawn from the existing `perturbation-sim` grid + generator's own edge model as a reusable fixture generator, `Grid::new` + pattern already present at `perturbation-sim/src/hhtl.rs:29-43`, reused + for graph-shape only — not its Cheeger machinery). +- **Query shapes**, generalizing `layout_probe.rs`'s own four box kinds + (§ SOURCE ARCHAEOLOGY 4) plus two the program's metric list names + explicitly: + 1. tile-aligned box (aligned to a quadtree cell boundary at some level) + 2. non-tile-aligned (arbitrary axis-aligned box) + 3. edge/corner-adjacent (this cycle's analogue of "pole-adjacent" — + boundary effects at the 256×256 grid's own edge, relevant because + seal boundaries are real edges here, unlike the ERA5 pole) + 4. point query (single-row lookup — sanity control, expect all arms + equal) + 5. k-radius neighbourhood (coordinate Chebyshev/Manhattan ball, OR + graph-space `family_hop_count`-ball for A5/A6 — this is the metric + program's "neighborhood-take latency" row, and the direct + generalization of `layout_probe.rs`'s still-gated "half B") + 6. uniform-random take (adversarial; expect near-identical clustering + number across ALL arms including A1 — this is the metric program's + "random-take latency" row, and doubles as a sanity check that no arm + is silently degenerate on this shape) + +### Metrics, defined precisely (reusing the in-project, literature-matching definitions from § SOURCE ARCHAEOLOGY 4) + +- **Clustering number `c(q,π)`** — for query `q` and key order `π`: sort + `q`'s cells by rank under `π`, count breaks in the consecutive-integer + run (`layout_probe.rs:581-599`, identical to Faloutsos & Roseman's + metric, § PRIOR ART). Reported as the median over each query-shape class + (per `layout_probe.rs`'s own reporting convention), never a single mean + that a fat tail can hide. +- **Neighbour locality** — median `|rank(a)−rank(b)|` over a deterministic, + coprime-stride sample of adjacency pairs (coordinate 4-neighbourhood for + A1–A4; graph 1-hop edges for A5/A6), `layout_probe.rs:601-674`. +- **Key-generation cost** — wall-time and instruction-count proxy (ALU op + count from the formula column above) per row, measured with the same + arm's cost isolated from sort cost. +- **Sort/compaction cost** — `O(N log N)` comparison sort for A1–A5; + eigensolve-dominated for A6; measured wall-time for `N=65,536` (one + cycle) and for a multi-cycle compaction batch. +- **Index remap bytes/CPU** — proxy: number of rows whose physical rank + changes between "as-sealed" order and "as-compacted" order under each + arm, fed conceptually through Lance's `RowAddrRemap` + (`remapping.rs:1-80`); A1 applied to an already-`stream_position`-ordered + image is expected to have the SMALLEST remap footprint of any non-trivial + arm specifically because it needs no transform — measure, don't assume. +- **Repair read/write bytes & helpers touched** — for a synthetic RS + stripe defined either as (i) N physically-consecutive rows post-compaction + or (ii) N rows sharing a locality-key prefix (§ MECHANISM (a)), measure + how many additional physical pages a repair must touch under each + grouping policy × each key arm. +- **Version count / stale-version retention bytes** — count how many rows + a resort under each arm displaces relative to the PREVIOUS cycle's + compacted layout (a proxy for how much churn a periodic re-sort would + cost in Lance version/GC terms); A1 applied on top of an + already-A1-ordered prior cycle is expected to be near-zero churn + (append-friendly); A4/A6 full resorts are expected to be high-churn + unless the new cycle's rows already fall near their predecessors' key + order. +- **Restart recovery work** — orthogonality control: this must be + UNCHANGED across all 7 arms, because `WalSink::scan_sealed` reads + sealed-cycle order, not a compaction-time key (§ SOURCE ARCHAEOLOGY 8). + If measured recovery cost DOES vary by arm, that is itself a kill + condition (§ below) — it would mean the locality key leaked into the + epistemic write path, which the source says must not happen. + +### Pre-registered kill conditions + +1. **KILL for "any locality key beats zero-cost A1"**: if A2's median + `range_count` on non-tile-aligned boxes is not ≥10% better than A1's, + *while also* not worse on `neighbour_locality` — i.e. if A2 merely + repeats A3's split verdict — reject bit-level Morton for migration too, + on the same bar `layout_probe.rs` already applied to A3. +2. **KILL for "Hilbert is worth its branch cost"**: if A4's key-generation + cost exceeds A2's by more than the ratio of their clustering-number + improvement (i.e. Hilbert's per-row ALU cost buys less than + proportional locality gain), prefer A2. +3. **KILL for "graph-space key serves coordinate-space queries too" (the + joint-key conjecture)**: if A5/A6's `range_count` on coordinate-space + boxes is not better than A7 (random control) — i.e. no better than + having no coordinate structure at all — the "one key for both spaces" + hypothesis is falsified for this geometry: spatial and graph locality + need separate keys. +4. **KILL for "locality-driven compaction is compatible with the one-WAL- + write-per-cycle discipline"**: this is a **structural finding, not a + hypothesis to run** — persist_sink.rs's own documented behaviour + (§ SOURCE ARCHAEOLOGY 8) already establishes that computing a locality + key inside `freeze()`/`order_cycle_stably()` itself would violate the + write-side physical-race-safety invariant. Any future design that + proposes fusing the locality key into the seal, rather than into a + post-seal compaction pass, should be rejected on this evidence without + re-deriving it. +5. **KILL for the RS-group coupling (§ MECHANISM (a))**: if a repair under + the "logical/declustered group" policy touches *more* physical pages + than under "physical-contiguity groups" post-A1-compaction (i.e. the + declustering buys nothing once even the cheap zero-cost key is applied), + the coupling conjecture is not worth its bookkeeping complexity — plain + physical-contiguity groups on an A1-ordered image suffice. + +--- + +## PRIOR ART + +**Searches actually run** (all via `mcp__alphaXiv__discover_papers` and +`WebSearch`, not memory): + +1. `discover_papers(["Hilbert curve","Z-order","space-filling curve", + "clustering property","range query"], prioritize=historical)` → + returned, among others, **Onion Curve: A Space Filling Curve with + Near-Optimal Clustering** (arXiv 1801.07399). Read its content directly + (fetched via `mcp__alphaXiv__get_paper_content`, grepped for the metric + definition rather than reading all 3,463 lines). It defines the + **clustering number** `c(q,π)` — "the minimum number of clusters of π + for query q" — states it as the standard measure for SFC range-query + performance, cites this as originating with Faloutsos & Roseman's 1989 + analysis of the Z-order (Hilbert vs Z-order clustering comparison) and + with Alber & Niedermeier's lower bounds on the Hilbert curve's + clustering number (both cited *within* the Onion Curve paper's related + work — I have not independently verified the 1989 paper's exact title + myself, only the citation as it appears in this secondary source, so I + flag this provenance explicitly rather than asserting first-hand + verification). Figure 1/2 of that paper show the canonical example: a + query region where Hilbert's clustering number is 2 and Z-order's is + higher, and an average-case sweep over all 7×7 query squares where + Hilbert's average clustering number is "much higher" for Z-order than + for their proposed Onion curve. This is the exact metric `range_count` + in `layout_probe.rs` (§ SOURCE ARCHAEOLOGY 4) independently reinvents — + **KNOWN, matched to established literature.** +2. Also surfaced in the same search: **Z-ordered Range Refinement for + Multi-dimensional Range Queries** (2305.12732) and **Efficient Cost + Modeling of Space-filling Curves** (2312.16355) — both post-date and + refine the clustering-number framework; not read in full, listed for + completeness of the search record. +3. `WebSearch("Delta Lake Z-ordering compaction file skipping locality + data layout 2024 2025")` → confirms the industry precedent for the + exact plug-in point this design proposes in § MECHANISM (b): Z-ORDER + is explicitly a **periodic, explicit `OPTIMIZE`-time** operation in + Delta Lake, distinct from ordinary small-file compaction, and its + effectiveness is documented to degrade with each additional + clustering column and to depend on collected per-file min/max stats — + directly analogous to this design's "compaction-time-only, not + write-time" constraint (§ SOURCE ARCHAEOLOGY 8) and to the concern that + a locality key not backed by usable data-skipping stats is wasted work. + Also surfaced "Liquid Clustering" as the 2025/2026-era evolution away + from manual Z-ORDER re-runs toward continuously-maintained clustering — + relevant context for whether this program's "background compaction + pass" framing (a discrete, periodic operation) or a continuous variant + is the better target, left open here. +4. `discover_papers(["erasure coding","data locality","repair locality", + "space-filling curve","columnar storage placement"])` → returned + several Locally Repairable Code (LRC) papers (2401.15308, 2505.06819, + 2512.10425, 2312.12803) — all about the **coding-theoretic** locality + of the *code itself* (rack-aware, wide-stripe LRC constructions), not + about coupling a data-placement/query-locality key to the parity-group + boundary. **No paper returned proposes using the SAME key for both + query-range locality and erasure-repair-group locality** — this + directly supports classifying the joint-key conjecture in § MECHANISM + (a) as NOVEL-CANDIDATE rather than KNOWN (§ SURPRISES). + +--- + +## SURPRISES + +1. **"The project's own already-shipped, non-interleaved tiered key beats + a real (if coarse) Morton variant on the exact range-query metric the + literature says Z-order should help with, on a real 1M-cell grid."** + — **KNOWN** (this is a project-internal, already-measured, already-on- + the-record finding — `SUBSTRATE_FORMULA_MATRIX.md` rows G16–G18 — not + something I am newly discovering; it existed before this search, and I + located and am reporting it, per the task's own instruction to + characterize project-native mechanisms with file:line). + +2. **"Today's actual write-path row order (`stream_position`, i.e. arrival + order of 65,536 concurrently-completing producers) is architecturally + closer in kind to a random-control baseline than to any spatial order, + by the write path's own documented behaviour."** — **TRANSFER**. The + "random control as the locality floor" concept is standard in the + clustering-number literature (and is exactly what `layout_probe.rs`'s + own CONTROL-BAD arm operationalizes, though CONTROL-BAD is a + *deliberately reversed* key, not an unordered/arrival-order one). Here + I am transferring that same floor-baseline framing onto a *pre-existing, + unrelated* project mechanism (`persist_sink.rs`'s documented + "ARBITRARY CPU order") that was never built with locality in mind — + a known technique recognized as already instantiated one layer away + from where it was designed. + +3. **"The same 'recursive-bisection → HEEL/HIP/TWIG address' shape was + independently built four times across this workspace, for four + unrelated domains (ontology routing, ERA5 grid tiling, semantic- + centroid scoring, and electrical-grid spectral resilience), with no + shared trait and no cross-reference between them."** — **NOVEL- + CANDIDATE** as a workspace-architecture observation (the search run was + exhaustive grep across the repo for `HHTL|Morton|CurveRuler`, § + SOURCE ARCHAEOLOGY throughout, and the individual instances are each + independently documented but never named as instances of one pattern + anywhere I found). + +4. **"Cheeger/Fiedler spectral bisection (already implemented in this + workspace for an unrelated power-grid simulator) is mechanistically the + right tool to jointly minimize graph-query scatter AND erasure-repair + scatter, because both objectives reduce to 'minimize edge cut across + the chosen partition' — yet no prior art proposes coupling ONE spectral + partition to BOTH a data-placement key and an RS parity-group boundary + in a general storage system."** — **NOVEL-CANDIDATE**, with the caveat + that its two halves are separately well-trodden: graph partitioning for + *query* locality in distributed graph stores (METIS/Fennel/PowerGraph- + style partitioning) is well-established general knowledge, and LRC-style + *coding* locality is its own established literature area (§ PRIOR ART + item 4) — the search specifically found no paper proposing the + combination, which is the actual novel-candidate claim, not either half + alone. + +5. **"A2 (true bit-level Morton, reusing `FacetTier::morton()` verbatim) + has never been run, even though it is a single already-written function + call away and the project already ran and KILL'd its coarser nibble- + level sibling."** — **DISPROVED-shaped caution, not a finding**: I flag + this as a gap rather than a surprise proper, because it would be easy + to over-read "the coarser variant lost, therefore Morton is bad" when + the actual tested arm was explicitly documented (`layout_probe.rs:46-56`) + as *not* the finer bit-level spelling. I classify the *inference* "Morton + is settled/killed in this workspace" as a risk to guard against, not + as something this search itself disproved — it is exactly the kind of + premature generalization § EXPERIMENT's A2 kill condition (§4) exists + to test honestly rather than assume either way. diff --git a/docs/lotus/rp-seal-v1/B2.md b/docs/lotus/rp-seal-v1/B2.md new file mode 100644 index 00000000..a5ff427e --- /dev/null +++ b/docs/lotus/rp-seal-v1/B2.md @@ -0,0 +1,891 @@ +# Cell B2 — Domain B (spatial layout / Morton cascade), role: ADVERSARY + +**Charge:** construct query distributions and identity distributions designed to make the +Morton/Z-order cascade LOSE; attack the 4-ary hierarchy (4/16/64/256/1024/4096) for +resonance and petal hot-spotting; deliver per attack the layout defeated, the metric, the +expected magnitude, and — two-sided — what result would REHABILITATE Morton. + +**Posture note.** Nothing here was built or run against the repositories. All source claims +are read-only citations with file:line. All numbers below come from a standalone +combinatorial harness I wrote in the scratchpad +(`b2work/{sfc,sfc2,sfc3,sfc4,petal}.py`), calibrated against the published literature +before any attack was scored (see §MECHANISM, "Harness calibration"). + +--- + +## SOURCE ARCHAEOLOGY + +### S1. There are (at least) four different "Morton" constructions in the workspace, and they are not the same object + +| # | Construction | Where | Axes | What it interleaves | +|---|---|---|---|---| +| M1 | `FacetTier::morton()` = `spread8(lo) \| (spread8(hi) << 1)` | `lance-graph/crates/lance-graph-contract/src/facet.rs:62-73` | 8 bit × 8 bit → `u16` | the two opaque bytes of ONE 8:8 tile | +| M2 | `CascadeKey::{family,leaf,identity}` + `morton48()` | `lance-graph/crates/perturbation-sim/src/cascade_key.rs:52-112` | 24 bit × 24 bit → 48 bit | a spectral-embedding coordinate pair, byte-blocked into three `u16` tiers | +| M3 | `morton_slot(h)` = `((x&0xF0)<<8)\|((y&0xF0)<<4)\|((x&0x0F)<<4)\|(y&0x0F)` | `lance-graph/crates/onebrc-probe/src/lane_f.rs:83-87` | 8 bit × 8 bit → `u16` | two bytes **of an FNV-1a hash** — and it is a *nibble*-block interleave, not a bit interleave | +| M4 | `morton2(x,y)` (true bit interleave, 16×16→32) | `lance-graph/crates/perturbation-sim/src/splat.rs:81-91`; `helix/examples/morton_shift_motion_probe.rs:74-90` | 16 bit × 16 bit | geometric / sprite coordinates | + +M2 is a genuine Z-order code: byte-blocked concatenation of per-tier bit interleaves is +bit-identical to a full 48-bit interleave of the two 24-bit axes, because block +concatenation preserves the interleave order. Verified by construction in the harness. + +M3 is **not a locality construction at all** — see attack A9. + +### S2. The one design decision that the source itself flags, and its cost is unmeasured + +`cascade_key.rs:69-75`: + +> Map a normalized coordinate `t ∈ [0,1]` to a 24-bit axis index `[0, 2^24)`. +> **Linear (not rank)** so a nibble prefix = a quad-tree quadrant — the +> `256 = 4⁴` hierarchical-ancestry condition. + +This is the load-bearing trade of the whole cascade: *quadtree ancestry requires linear +(equal-width) quantization; equal-width quantization has no defence against skew.* Attack +A7 measures what that costs. Industry practice went the other way (see PRIOR ART P6). + +### S3. A measured Morton-interleave REJECTION already exists in the repo + +`lance-graph/crates/lance-graph-contract/src/rail_geometry.rs:24-30`: + +> The slab variant is not decoration: a consumer bake **measured the pair reading and +> rejected it** for its hierarchy (**44.25 % of paths fit, vs 99.62 %** in twelve +> per-axis levels). A pair byte is two SEPARATE bytes — never widened to u16; a widened +> word has no axis. + +`RailCarving::{InterleavedPairs, AxisSlab}` (`rail_geometry.rs:66-75`) exists *because* +the interleaved (Morton-shaped, 6 levels × 2 axes) reading lost to the de-interleaved +(12 levels × 1 axis, i.e. **raster-per-axis**) reading by 2.25× on path-fit for a real +bake. This is an in-repo precedent that the raster/axis-slab layout beats the interleaved +layout when the hierarchy is deeper on one axis than the other — exactly attack A1's +shape, one layer up. + +### S4. The shipped persistence path does not order by Morton + +- `lance-graph/crates/lance-graph-planner/src/persist_sink.rs:133-135` — + `order_cycle_stably(rows, key)` is `rows.sort_by_key(key)`, generic. +- `:377-378` — `freeze()` calls `order_cycle_stably(&mut casts, |s| s.stream_position)`. +- `:168-183` — `SweepSlot::stream_position` is documented as *"the caller's EXISTING + canonical order key"*, *"the existing canonical (**textual/stream**) order key … NOT a + new coordinate"*, with the hard contract *"monotonically increasing per owner ACROSS + cycles"* because `recover_and_apply` uses it as the durable restart watermark + (`:653-654`, `:678-682`). +- `:357-358, :379-382` — the durable image is `BTreeMap>` keyed by + `SweepSlot::row`, so the emitted physical order **is ascending `row`**. + +Consequence, and it is the most important structural fact in this report: **`stream_position` +cannot be a Morton code** (a Morton code over identity is not monotone in time, so it +would break the recovery watermark), and it is not one today. The only place a +space-filling order can enter the shipped write path is `SweepSlot::row`, i.e. at +identity/locus minting — *upstream of persistence entirely*. Today no code in the persist +path consumes M1/M2/M3/M4. Every Morton layout benefit in this program is therefore +**prospective**, and the experiments below must be read as measuring a layout that does +not yet exist in the write path. + +`persist_sink.rs:128-131` already names the constructive hook: + +> when the key is already a dense slot index into the cycle image, a stable scatter into +> predetermined positions is cheaper than a general `O(n log n)` sort on the 64k path. + +### S5. Lance, two-column discipline + +| Fact | Column A: pinned 9.0.0 (`/tmp/sources/lance-9`, tag v9.0.0) | Column B: current upstream (`/tmp/sources/lance-main`, `a8e2a78`, cloned 2026-08-18) | +|---|---|---| +| Space-filling curve shipped | **Hilbert only.** `rust/lance-index/src/scalar/rtree/sort/hilbert_sort.rs` — `HilbertSorter` (:28-39), `hilbert_curve(x,y)` (:198-201), 16-bit axes (`hilbert_max = (1<<16)-1`, :174) | same file, same code | +| Z-order / Morton anywhere in Lance | **absent** (grep `z_order\|zorder\|morton` over `rust/` → zero hits) | absent | +| `Sorter` trait implementations | exactly one: `HilbertSorter` (`rust/lance-index/src/scalar/rtree/sort.rs:8-13`) | same | +| Public vs internal | `#[cfg(feature = "geo")] pub mod rtree;` — `rust/lance-index/src/scalar.rs:44-45`. Public API, gated behind the `geo` feature; used for R-tree bulk-load ordering, **not** as a dataset-layout facility | same | +| Compaction reordering | **none.** `rust/lance/src/dataset/optimize.rs:842` *"Compacts the files in the dataset **without reordering them**"*; `:849` *"tries to preserve the insertion order of rows"*; `:1296-1303` rewrite scan is `scan_in_order(true)` over the old fragments | same | + +Two consequences. (a) The storage engine this architecture sits on, faced with the exactly +analogous problem — linearize 2-D extents for an index — chose **Hilbert, not Z-order**, +in both columns. (b) Lance 9.0.0 offers **no reorder-during-compaction**: there is no +`OPTIMIZE ZORDER BY` equivalent. Whatever order the write path emits is the order that +survives; compaction can only concatenate and drop deletions. A cycle-partitioned writer +that emits Morton order *within* each cycle therefore never converges to a globally +Morton-ordered dataset — see attack A10. + +### S6. Existing in-repo defences I must not attack as if they were absent + +- **Anti-moiré dither.** `helix/src/constants.rs:46-51` and `curve_ruler.rs:1-27`: the + stride-4-over-17 coprime walk (`gcd(4,17)=1` ⇒ full permutation of all 17 residues; + the doc explicitly contrasts a non-coprime stride that "misses {6,7,10,11}"). This is a + real, shipped defence against the resonance class in attack A4 — but note it operates + on a **prime** modulus 17, while the cascade's own moduli are 4, 16, 256, 4096, i.e. + powers of two, where *no* stride is coprime to the modulus unless it is odd. + `perturbation-sim/examples/comma_quorum.rs:62-63` records exactly this nudge + ("nudged to 2395 to be coprime with M = 2^12"). +- **O(1) rigid translation.** `helix/examples/morton_shift_motion_probe.rs:20-31`: + dilated-lane add moves every pixel's Morton code with one delta. This is a genuine + Morton-only capability and attack A6 must not pretend otherwise — but it moves the + *code*, not the *storage*, which is the distinction A6 exploits. +- **EWA seam suppression.** `perturbation-sim/src/splat.rs:104-129` + + `ewa_suppresses_a_morton_seam_outlier` (:169-186): the repo already knows Morton has + "Z-jump seams" that a box average aliases across, and down-weights them. That is an + acknowledgement of the A2 defect at the coarsening layer, not at the query layer. +- **The measured control already exists.** `lane_f.rs:89-95, 245-252`: `lane_r_radix` is + described as *"the plain-radix control: identical pipeline, slot = low 16 hash bits, no + interleave. F−R isolates the addressing tax."* The repo built the right control; A9 + argues about what it can possibly measure. + +--- + +## MECHANISM + +### Harness + +A 256 × 256 grid (65,536 cells — exactly one cycle at the program's baseline geometry, +and exactly one cascade tier's 256×256 centroid tile). Five linearizations: + +- `morton` — true bit interleave (M2/M4 form), `lo`/x on even bits, `hi`/y on odd, matching + `facet.rs:62`; +- `hilbert` — **the exact `hilbert_curve` from Lance 9.0.0** + (`hilbert_sort.rs:198-245`), ported verbatim so the comparison is against the code the + storage engine actually ships; +- `raster` (row-major), `column` (column-major) — the two axis-aligned baselines; +- `nibble4` — the `morton_slot` nibble-block interleave from `lane_f.rs:83-87`. + +Metrics, all reported: **runs** (maximal contiguous intervals in the linear order — the +classic clustering number, and the count of sequential-IO ranges), **pages@16** (distinct +16-row pages touched = one 4×4 work field = 8 KiB at 512 B/row), **pages@1024** (distinct +1024-row pages), and **span** (`max−min` of the linear position — the min/max zone-map +proxy). + +### Harness calibration (done before scoring any attack) + +Mean cluster count for a q×q window over all offsets: + +| q | morton | hilbert | raster | literature | +|---|---|---|---|---| +| 2 | **2.6094** | **1.9994** | 2.0000 | z-curve 2.625, Gray 2.5, Hilbert 2.0 (Moon et al.; see P1) | +| 3 | 4.4961 | 3.0221 | 3.0000 | — | +| 4 | 6.2882 | 3.9670 | 4.0000 | — | +| 8 | 14.1345 | 7.9510 | 8.0000 | — | + +Morton 2.6094 vs the published 2.625 and Hilbert 1.9994 vs the published 2.0 confirm the +harness reproduces the standard result. Everything below is scored on the same harness. + +### The eleven attacks + +Each attack states: **layout it defeats · metric · magnitude · rehabilitation condition.** + +--- + +#### A1 — Full-length axis stripe ("scan one row of the grid") + +Defeats: **Morton and Hilbert both**, in favour of raster (horizontal) or column-major +(vertical). + +| query | morton runs | hilbert runs | raster runs | morton p@16 | raster p@16 | column p@16 | +|---|---|---|---|---|---|---| +| 1×256 horizontal | 128 | 86 | **1** | 64 | **16** | 256 | +| 256×1 vertical | 256 | 86 | 256 | 64 | 256 | **16** | + +Averaged over all offsets (pages@16): horizontal — morton 64.00, hilbert 64.00, raster +**16.00**, column 256.00; vertical — the mirror image. + +Magnitude: **4× more pages, up to 128× more IO ranges** than the matching axis-aligned +layout. + +The honest counter, and it is strong: Morton is the **minimax** choice — 64 both ways +against raster's {16, 256}. Geometric mean is identical (√(16·256) = 64). So Z-order buys +**variance reduction, not expected-value reduction** on axis-aligned thin queries. The +crossover is exact and computable: with fraction *p* of full-length stripes on the +favoured axis, raster costs 16p + 256(1−p) and Morton costs 64; they cross at + +> **p\* = 0.80.** + +Kill / rehabilitate: if the measured workload has **≥ 80 % of thin-query mass on one +axis**, ship the axis-slab layout — and note `rail_geometry.rs:24-30` records a bake that +already measured itself into that regime (44.25 % vs 99.62 %). Below 80 %, Morton wins +and A1 is rebutted. + +--- + +#### A2 — The off-by-one power-of-two straddle (the Z-jump seam, at the query layer) + +Defeats: **Morton**, and to a lesser degree Hilbert; favours raster on span. + +A 16×16 window, quadrant-aligned vs shifted by one cell across the grid centre: + +| window | morton runs | morton p@16 | morton span | hilbert runs | raster span | +|---|---|---|---|---|---| +| aligned @ (0,0) | **1** | 16 | 256 | 1 | 3856 | +| centred @ (120,120) | 4 | 16 | 32,896 | 3 | 3856 | +| off-centre @ (121,121) | **46** | **25** | 33,022 | 24 | 3856 | + +Magnitude: a **one-cell translation multiplies run count by 46×**, pages@16 by 1.56×, and +span by **129×** (256 → 33,022 = 50.4 % of the whole dataset). The same shift at 32×32: +runs 1 → 94, pages@16 64 → 81. + +This is the Databricks/Hudi "large min/max difference" complaint made numeric — but see +S8/A11 below, because it does **not** land where the folklore says it does. + +Rehabilitate: if the query generator is quadrant-aligned by construction — which it is +when the query is *"give me this cascade prefix"* rather than *"give me this rectangle"* — +A2 is void. The rehabilitation condition is therefore falsifiable and sharp: **show that +≥ 90 % of production range predicates are expressible as a nibble-prefix mask on the +cascade key rather than as an arbitrary rectangle.** If they are, Morton's aligned column +(runs = 1, span = |Q|) is the operative one and A2 never fires. + +--- + +#### A3 — The growing square window (the Moon et al. gap, and it widens) + +Defeats: **Morton**, in favour of Hilbert *and* raster, on run count. + +Morton/Hilbert cluster ratio: q=2 → 1.31×, q=3 → 1.49×, q=4 → 1.59×, q=8 → **1.78×**. +Morton/raster at q=8 → 1.77×. The deficit grows with query side length and, per the +literature, with dimension. + +Magnitude: **1.3–1.8× more IO ranges**, growing. + +Rehabilitate: A11 (below) is the rehabilitation, and it is decisive at page granularity. + +--- + +#### A4 — Strided lattice / cascade resonance (the 4-ary hierarchy attacked directly) + +Defeats: **Morton and Hilbert equally**, in favour of raster/column. + +Sample every s-th cell on both axes (a stride-s lattice — telemetry decimation, LOD +preview, every-Nth-tenant sweep, a checkerboarded parity group): + +| stride s | morton p@16 | hilbert p@16 | raster p@16 | morton/raster | +|---|---|---|---|---| +| 2 | 4096 | 4096 | 2048 | **2×** | +| 4 | 4096 | 4096 | 1024 | **4×** | +| 8 | 1024 | 1024 | 512 | 2× | +| 16 | 256 | 256 | 256 | 1× | + +Derived law (regime `s ≤ 2^j`, page = 4^j rows = a 2^j × 2^j aligned tile): a stride-s +lattice puts (2^j/s)² points in each Morton page and 2^j/s·2^j/… — concretely +`morton_pages = |Q|·s²/4^j` and `raster_pages = |Q|·s/2^j`, giving a ratio of exactly **s**, +maximised at **s = 2^j = √P**, where Morton touches one page per point. Above that the +ratio decays as 2^j/s. Predicted worst case ratio **√P**; measured √16 = 4 at s = 4. ✓ + +**This is the resonance the charge asked about, and it is real: the cascade's own 4-ary +period is the period a strided workload aliases against, and the penalty grows as the +square root of the page size.** Bigger pages make it strictly worse — which cuts directly +against "raise the field size to amortize the seal." + +Rehabilitate, two ways, both falsifiable: +1. **The coprime dither already in the repo.** `helix/src/constants.rs:46-51`'s + stride-4-over-17 walk breaks alias by construction *over a prime modulus*. Show that a + comma-phase permutation applied to the low tier converts the stride-s lattice into a + pseudo-random subset (predicted: Morton pages → min(|Q|, 4^j·(1−e^{−|Q|/4^j}))). If the + measured Morton page count under dithered addressing drops to raster's, A4 is + rebutted. Note the obstruction: **no stride is coprime to a power-of-two modulus unless + it is odd**, so the dither must be an odd-stride walk or a bit-mixing permutation, not + the literal 4-over-17. +2. Show that no production query is a lattice. Strided sampling is a *plausible* workload + (decimation, sharded sweeps, parity striping) but if the measured workload has none, A4 + is inert. + +--- + +#### A5 — Checkerboard: THE ATTACK BACKFIRES (recorded as a rehabilitation) + +I built the checkerboard expecting a wash. It is a Morton **win**, and a clean one. + +Full-grid checkerboard at block size b (32,768 cells selected), run counts: + +| b | morton | hilbert | raster | column | nibble4 | +|---|---|---|---|---|---| +| 1 | **16,385** | 32,768 | 32,641 | 32,641 | 30,721 | +| 2 | **4,097** | 8,192 | 16,321 | 16,321 | 15,361 | +| 4 | **1,025** | 2,048 | 8,161 | 8,161 | 7,681 | +| 8 | **257** | 512 | 4,081 | 4,081 | 3,841 | + +Morton beats Hilbert by **exactly 2.0×** at every b ≥ 1 and raster by **4–16×**. +Mechanism: a dyadic checkerboard selects cells with `(x_k ⊕ y_k) = 0` at bit level k; in +Morton those two bits are *adjacent* in the code (positions 2k, 2k+1), so the selected set +is a union of dyadic intervals; in Hilbert the parity is scrambled by the state machine and +every selected cell is isolated. Pages@16 also favours Morton at b ≥ 4 (2048 vs raster's +4096). + +Rehabilitation condition (i.e. this is what would *save* Morton): **any workload whose +selection predicate is a parity, XOR, or dyadic-residue condition on the coordinates — +parity groups, RS striping across a 2-D grid, alternating-tenant sweeps, bit-plane +extraction — is a regime where Z-order is the correct choice and Hilbert is strictly +worse.** Since erasure-coding stripes over a 2-D grid *are* exactly such a predicate, this +matters for the seal domain and I flag it for cross-domain follow-up (Domain C/D). + +--- + +#### A6 — Moving window (the sliding viewport) + +Defeats: **Morton**, in favour of Hilbert and — surprisingly — the nibble-block variant. + +A 16×16 window slid by one cell in x: + +| x₀ | morton runs | hilbert runs | nibble4 runs | morton p@16 | +|---|---|---|---|---| +| 112 (aligned) | 1 | 1 | 1 | 16 | +| 113 | **32** | 13 | **2** | 20 | +| 114 | 16 | 7 | 2 | 20 | +| 115 | **32** | 11 | 2 | 20 | + +Magnitude: Morton's cost **oscillates with window phase by 32×** (1 ↔ 32) and is 2.5–2.9× +Hilbert's off-phase. The nibble-block interleave `morton_slot` — which groups 16 consecutive +x-values before interleaving — is **16× better than Morton here**, because its granularity +matches the window. + +Two important qualifications against my own attack. (i) `morton_shift_motion_probe.rs` +correctly establishes that a rigid translation is a single dilated-lane add on the *code*, +O(1) in sprite size. That is true and is a real Morton capability. But it moves addresses, +not bytes: the *storage ranges* the moved window covers are still 32 runs. **Address-delta +O(1) and IO-range O(1) are different claims and only the first is proven.** (ii) At +pages@16 the oscillation is only 16 → 20 (1.25×), far milder than the run count. + +Rehabilitate: measure whether the consumer is range-request-bound (runs matter — Morton +loses 2.5×) or bytes-bound (pages matter — Morton loses 1.25× and Hilbert is identical, per +A11). If bytes-bound, A6 is a rounding error. Also: if the window is always tile-aligned by +construction, A6 is void — same escape as A2. + +--- + +#### A7 — Petal hot-spot under skewed identity (the strongest attack in this report) + +Defeats: **the cascade's uniformity premise**, independent of curve choice. + +Cycle = 65,536 rows over 4,096 fields of 16 ("the 16-slot petal"); field id = top 12 bits +of the 48-bit cascade code. Identity coordinates drawn from five distributions, quantized +two ways — `linear` (the shipped `axis24`, `cascade_key.rs:69-75`) and `rank` (per-axis +quantile transform, which is what the industry does — see P6): + +| distribution | quant | max occupancy | max/mean | **empty fields** | Gini | top-16 quadrant max/mean | +|---|---|---|---|---|---|---| +| uniform | linear | 32 | 2.00× | 0.0 % | 0.140 | 1.03 | +| uniform | rank | 31 | 1.94× | 0.0 % | 0.137 | 1.03 | +| single Gaussian cluster | **linear** | 712 | **44.50×** | **83.8 %** | **0.955** | 4.03 | +| single Gaussian cluster | rank | 33 | 2.06× | 0.0 % | 0.137 | 1.01 | +| 8 Gaussian clusters | **linear** | 339 | **21.19×** | **70.7 %** | 0.908 | 2.40 | +| 8 Gaussian clusters | **rank** | 151 | **9.44×** | **54.4 %** | 0.804 | 2.00 | +| Pareto/Zipf | linear | 55 | 3.44× | 0.0 % | 0.288 | 2.20 | +| Pareto/Zipf | rank | 31 | 1.94× | 0.0 % | 0.139 | 1.02 | +| lognormal | **linear** | 341 | **21.31×** | 9.7 % | 0.670 | 4.69 | +| lognormal | rank | 86 | 5.38× | 0.2 % | 0.146 | 1.01 | + +Magnitude: **20–45× occupancy skew and 70–84 % empty fields** under clustered identity +distributions. Mean occupancy is 16 rows/field; the hot field holds 712. + +This is not a hypothetical distribution. `cascade_key.rs:6-15` says the axes come from a +**spectral embedding** (two lowest non-trivial Laplacian eigenvectors) — and spectral +embeddings of graphs with community structure are *definitionally* multi-modal and +cluster-heavy. The "8 Gaussian clusters" row is the closest model of that and it is the +worst-behaved: **rank quantization only halves it (21.19× → 9.44×, 70.7 % → 54.4 % empty), +because a per-axis marginal quantile transform cannot decorrelate 2-D clusters.** + +Consequences that reach the rest of the program: +- **Storage**: if a parity group or a compaction unit is a field, 70–84 % of it is padding. + Physical bytes for a cycle inflate by up to 6× at the same logical bytes. +- **Seal metadata**: 4,096 field descriptors of which ~700 carry data. +- **Compaction**: the "compaction boundary = Morton prefix = parity-group boundary" + equivalence the discovery mandate asks about *requires* roughly uniform occupancy to be + useful; at Gini 0.955 the boundaries carve wildly unequal work. +- **The trilemma**: quadtree ancestry ⟹ linear quantization ⟹ hot-spots. You may have any + two. This is the sharpest statement I can make and it is the one I would put to the + proponents. + +Rehabilitate, in decreasing order of my confidence that it works: +1. **Measure the real identity distribution.** If the production identity coordinates are + near-uniform (e.g. because they are hash-derived, or because the codebook is trained to + equal-occupancy — which is exactly what the OGAR "hierarchical 4⁴ centroid codebook" + would achieve if it is trained rather than assumed), max/mean ≤ 2 and A7 evaporates. + **This is the single highest-value measurement in Domain B.** A hierarchically-trained + 4-ary centroid codebook *is* a per-level rank transform that preserves ancestry — if it + exists and is trained, the trilemma dissolves and A7 is fully rebutted. I found the + codebook *described* (`lance-graph-contract/src/codebook.rs:5`, OGAR CLAUDE.md "Tier + interpretation") but did not find a trainer. +2. Show that empty fields cost nothing because the cycle image is `BTreeMap`-sparse + (`persist_sink.rs:357`) and Lance stores only present rows — in which case the 84 % + empties are logical, not physical, and only the *addressing* is skewed. This is + plausible and would demote A7 from "storage catastrophe" to "load-imbalance". +3. Adopt Delta's `ADD_NOISE`-equivalent (P6) and accept the min/max-skipping loss. + +--- + +#### A8 — Morton's exact 2× axis anisotropy, mapped onto the workspace's own semantics + +Defeats: **Morton**, in favour of Hilbert, on the *odd-bit* axis specifically. + +Runs for a length-L line along each axis: + +| L | morton, vary the even-bit axis | morton, vary the odd-bit axis | hilbert (either) | +|---|---|---|---| +| 4 | 2 | 4 | 2 | +| 16 | 8 | 16 | 6 | +| 64 | 32 | 64 | 22 | +| 256 | 128 | 256 | 86 | + +Exactly **2.000× at every scale**; Hilbert is symmetric to within one run. + +Now bind it to the source. `facet.rs:54-64`: `morton() = spread8(lo) | (spread8(hi) << 1)` — +**`lo` on the even bits, `hi` on the odd bits**. And `facet.rs:33-36` / +`rail_geometry.rs:80-88` name the semantics: `lo` = *is_a / member / row / centroid-lo*, +`hi` = *part_of / group / column / centroid-hi*, with the zero-fallback putting +`RailAxis::Taxonomy` on byte 0 and `RailAxis::Mereology` on byte 1. + +Therefore: **"walk the mereology (part_of) chain at fixed taxonomy" costs exactly 2× the +IO ranges of "walk the taxonomy (is_a) chain at fixed mereology."** Nothing in the docs +records this. If the dominant traversal is part_of-major, the axis assignment is backwards +and a one-line swap (`spread8(hi) | (spread8(lo) << 1)`) halves it. + +Rehabilitate: measure the traversal mix. If taxonomy-major dominates, the current +assignment is already correct and A8 becomes a *confirmation* that someone chose well — +but it is currently unrecorded either way, which makes it luck rather than design. + +--- + +#### A9 — `morton_slot` over a hash is not a locality construction + +Defeats: **the premise**, not a layout. + +`lane_f.rs:83-87` computes a nibble-block interleave of hash bytes 0 and 1 of an FNV-1a +digest; `lane_f.rs:93-95` computes the control as the low 16 bits of the same digest. For a +uniform 64-bit hash, **any fixed bit/nibble permutation of 16 uniform bits is distributed +identically to the identity permutation.** The two slot functions are therefore equal in +distribution: identical expected collision count, identical probe-length distribution, +identical page-touch distribution, for every query, on every input. + +There is nothing for the interleave to preserve, because the hash has already destroyed +the locality. The repo states this correctly as far as it goes — +*"the one-line difference that isolates the addressing tax"* — but the framing invites the +reading that F is a *locality* lane. It is not; it is the same lane with four extra ALU +ops. `lane_g.rs:259` and `lane_h.rs:82` then use `morton_slot(h) * shards >> 16` as a shard +selector, which is a *top-bits-of-a-permuted-hash* selector — fine as a hash, but it is +selecting on the interleaved nibbles, i.e. on `x_hi`, which is hash byte 0's high nibble. + +Magnitude: **zero locality benefit; a small, measurable ALU cost.** Whatever F−R measures, +it cannot be a clustering effect, so any positive F−R must be attributed to code layout, +branch prediction, or noise. + +Rehabilitate: none available at the hash. Morton over a hash becomes meaningful only if the +"hash" is replaced by a **locality-sensitive** code (an LSH, a trained centroid id, a +CAM-PQ code) — at which point the interleave is doing real work. That is a concrete, +cheap, high-value change: interleave the *palette centroid pair*, never the digest. + +--- + +#### A10 — The temporal band: Morton locality is never realized across cycles + +Defeats: **the whole layout program**, on Lance 9.0.0 specifically. + +Writes are cycle-partitioned (`persist_sink.rs:136-143`, one cycle = one WAL append = one +`DatasetVersion`). Each cycle's image is `BTreeMap` keyed by `row` +(`:357-358, :379-382`), so each *fragment* is internally sorted by `row`. If `row` is a +Morton code, cycle *k* writes a Morton-scattered **subset** of the 48-bit space. Cycles +are appended, not merged. And Lance 9.0.0 compaction is explicitly +*"without reordering them"* / *"preserve the insertion order"* +(`rust/lance/src/dataset/optimize.rs:842, 849`, both columns). + +So after *n* cycles the dataset is *n* interleaved Morton-sorted runs. A range query on a +Morton prefix touches **up to n fragments**, one per cycle, regardless of how good the +curve is. Merging them into one globally-ordered run is precisely a merge sort — which no +Lance API in either column performs. + +Adversarial query: *"give me everything in Morton prefix P"* after 100 cycles. Files +touched ≈ min(100, fragments overlapping P). Compare a single-writer bulk load: 1 file. +Magnitude: **up to n× file-open and page-index cost, where n = version count.** + +Rehabilitate, and this is genuinely constructive: +- Show that the write path can assign `row` as a **dense slot index into the cycle image** + — which `persist_sink.rs:128-131` already anticipates — such that a cycle covers a + *contiguous* Morton range rather than a scattered subset. That converts each fragment + into a Morton *interval*, and then the interleaving disappears and prefix queries touch + O(1) fragments. **The measurement that decides it: what fraction of a cycle's 65,536 rows + fall in one 4^j Morton block?** If cycles are locus-coherent (a cycle sweeps a region), + A10 is void. If cycles are time-coherent (a cycle sweeps whatever arrived), A10 is fatal + and no curve choice can fix it — only a reordering compaction can, which Lance 9.0.0 does + not have and which would have to be built as an out-of-engine rewrite. +- This is the concrete answer to the discovery mandate's *"is compaction 'repair' of + physical locality?"* — **on Lance 9.0.0 it provably is not, because compaction is + order-preserving by contract.** Any locality repair must be a bespoke rewrite that + produces a new insertion order, and then it inherits the full index-remap cost. + +--- + +#### A11 — The attack on the premise: at page = 4^j, Morton and Hilbert are the SAME layout + +This is the result I did not expect and it reframes Domain B. + +**Measured page-partition identity.** For page size P, do Morton and Hilbert induce the +same *set* of pages (as sets of grid cells, ignoring page order)? + +| P | 4 | 16 | 64 | 256 | 1024 | 32 | 128 | +|---|---|---|---|---|---|---|---| +| identical partition | **yes** | **yes** | **yes** | **yes** | **yes** | no | no | + +Both curves are quadrant-refining: every aligned 2^j × 2^j block is contiguous under both. +So for P = 4^j they enumerate the *same blocks in a different order*. Consequently +**pages-touched is identical for every query**, which the sweep confirms exactly: + +| q | morton p@16 | hilbert p@16 | raster p@16 | morton p@1024 | hilbert p@1024 | raster p@1024 | +|---|---|---|---|---|---|---| +| 4 | 3.03 | **3.03** | 4.71 | 1.17 | **1.17** | 1.74 | +| 8 | 7.55 | **7.55** | 11.37 | 1.42 | **1.42** | 2.75 | +| 16 | 22.47 | **22.47** | 30.81 | 2.05 | **2.05** | 4.74 | +| 32 | 76.50 | **76.50** | 93.87 | 3.84 | **3.84** | 8.75 | + +Identical to the last digit, for every query family I tested (squares, stripes, hotspots, +checkerboards, strides). + +**The architecture's own baseline page is 4^j.** 8 KiB field ÷ 512 B row = 16 rows = 4². +The cascade's hierarchy is 4/16/64/256/1024/4096 — every one a power of 4 except 1024 +(=4^5) and 4096 (=4^6), which are also powers of 4. So the program sits *entirely* inside +the regime where the Morton-vs-Hilbert question is provably vacuous for bytes read. + +Two-sided consequences: + +- **Rehabilitates Morton against A3 (and the entire Moon et al. literature) for the + bytes-read metric.** The 1.78× cluster deficit is a *run-count* deficit — more IO ranges, + same pages. If the reader is a columnar page reader (which Lance is), it is invisible. +- **De-rehabilitates the choice of Morton as a differentiator.** It gains nothing over + Hilbert on pages either. The justification must therefore rest entirely on: dilated + arithmetic is a handful of ALU ops vs Hilbert's ~30-op state machine + (`hilbert_sort.rs:198-245` is 45 lines of bit-twiddling for one point), a nibble prefix + *is* a quadrant with no decode, and the A5 dyadic-parity family. Those are real. "Better + locality" is not among them. +- **And it identifies where the two curves DO differ: page ORDER.** Morton's page sequence + jumps at every quadrant boundary; Hilbert's is a Hamiltonian path on the page grid. That + matters for read-ahead, for the page-index min/max, and for whether *consecutive* pages + in a file are spatial neighbours. That, not cell-level locality, is the real question, + and no metric in this program currently measures it. + +**Caveat on the theorem's own boundary:** it holds only for P = 4^j *and* an aligned grid +*and* full occupancy. Under A7's 84 %-empty regime, page boundaries are defined by +*occupied-row counts*, not by geometry, so the partitions diverge and the theorem does not +apply. A7 and A11 interact: uniform occupancy is what makes the SFC choice moot; skewed +occupancy is what makes it matter again — and in the skewed regime nobody has measured +which curve wins. + +--- + +## EXPECTED BENEFIT + +Stated adversarially — the narrowest set of claims I could not break: + +1. **Minimax robustness on axis-aligned queries.** 64 pages either way vs raster's + {16, 256}. Real, bounded, and only a win below the p\* = 0.80 workload skew. +2. **Aligned-block queries are one run.** A prefix predicate on the cascade key is a single + contiguous interval, span = |Q|, zero decode. Nothing else in the comparison set does + this — raster needs 2^j runs for the same block. This is Morton's genuine advantage and + it is an advantage over *raster*, not over Hilbert. +3. **Dyadic-parity workloads (A5).** Morton beats Hilbert exactly 2× and raster 4–16×. + Directly relevant if parity groups are laid out as a 2-D residue class. +4. **O(1) rigid address translation** (`morton_shift_motion_probe.rs`). Real and + Morton-exclusive; applies to codes, not to storage ranges. +5. **Encode/decode cost.** ~5 ALU ops vs Hilbert's ~45-line state machine, with a GFNI + single-instruction path claimed at `facet.rs:25`. At 65,536 rows/cycle this is real CPU. +6. **Page-partition parity with Hilbert (A11).** Not a benefit of Morton — an absence of a + penalty. Worth stating because it removes the strongest published objection. + +## EXPECTED FAILURE + +1. **Skewed identity ⇒ 20–45× petal hot-spot, 70–84 % empty fields** (A7). The most likely + failure and the one the source's own `linear (not rank)` comment guarantees. Spectral + embeddings are cluster-heavy by construction. +2. **The cross-cycle interleave** (A10). Morton order within a fragment does not compose to + Morton order across fragments, and Lance 9.0.0 compaction is order-preserving by + contract. Locality is never repaired. +3. **Strided-lattice resonance with penalty √P** (A4). Grows with the page/field size, i.e. + worsens exactly when you enlarge the field to amortize the seal. +4. **The 2× axis anisotropy is currently unrecorded** (A8), so the taxonomy/mereology + assignment is a coin flip that happens to have landed. +5. **`morton_slot` over a hash measures nothing about locality** (A9), so any "Morton lane + wins" result from the 1BRC probes cannot support a layout conclusion. +6. **Off-alignment phase sensitivity: 32–46× run-count swings from a one-cell shift** + (A2, A6). Bounded to 1.25–1.56× at page granularity, so this fails loudly only for a + range-request-bound reader. +7. **Morton is not the engine's choice.** Lance ships Hilbert and only Hilbert, in both + columns. Any Morton layout is an out-of-engine convention the engine will neither + maintain nor exploit. + +## EXPERIMENT + +Pre-registered. Each: metric · baseline · control · kill condition. None requires building +the proposed architecture; all are measurable against existing artefacts or a synthetic +harness. + +**E-B2-1 — Identity occupancy census (highest value; gates A7).** +*Metric:* per-field occupancy distribution over 4,096 fields — max/mean, Gini, % empty, +p99 — for the REAL identity coordinate source (`spectral_embedding`, `cascade_key.rs`), +under both `axis24` linear and a per-axis rank transform. +*Baseline:* uniform random coordinates (predicted max/mean 2.00, 0 % empty — measured). +*Control:* the same points under a hierarchically-trained 4⁴ centroid codebook, if one can +be produced; if not, record its absence as a finding. +*Kill:* if measured max/mean ≤ 3 and empty ≤ 10 % on real data, A7 is dead and linear +quantization is vindicated. If max/mean ≥ 10 or empty ≥ 50 %, the cascade's uniformity +premise is falsified and the trilemma must be resolved before any seal/parity design that +assumes equal field occupancy. + +**E-B2-2 — Cycle locus coherence (gates A10).** +*Metric:* for each cycle, the fraction of its 65,536 rows falling inside a single 4^j +Morton block, for j = 2..6; and, over n cycles, the number of fragments overlapping a random +Morton-prefix range. +*Baseline:* a single bulk load of the same rows (1 fragment). +*Control:* random assignment of rows to cycles (predicted: every fragment overlaps every +prefix). +*Kill:* if fragments-touched grows linearly in version count, the layout program cannot +deliver query locality on Lance 9.0.0 without a reordering rewrite; report that as a +blocking dependency rather than a tuning parameter. + +**E-B2-3 — Runs-vs-pages decisiveness (gates A3, A6, A11).** +*Metric:* end-to-end point/neighbourhood/random-take latency and files-and-pages-touched +for the SAME query set under Morton and under the Lance 9.0.0 `hilbert_curve`, at page = +16 rows and page = 1024 rows. +*Baseline:* raster order. +*Control:* page = 32 rows (NOT a power of 4) — where the harness shows the partitions +*do* diverge, so any Morton/Hilbert difference must appear here and must not appear at 16. +*Kill:* if latency differs by < 5 % between Morton and Hilbert at 4^j pages, A11 is +confirmed operationally and the SFC choice must be re-justified on CPU and prefix +semantics alone. If it differs by > 20 %, A11 is wrong and I want to know why. + +**E-B2-4 — Workload axis-skew census (gates A1; decides against `RailCarving::AxisSlab`).** +*Metric:* the empirical distribution of query shapes — fraction that are thin on one axis, +their axis, their length; then compute pages@16 under Morton, raster, column, axis-slab. +*Baseline:* the p\* = 0.80 crossover computed here. +*Control:* the bake recorded at `rail_geometry.rs:24-30` (44.25 % vs 99.62 %) — re-derive +its numbers to confirm the harness agrees with a real measurement. +*Kill:* measured single-axis mass ≥ 80 % ⇒ ship axis-slab, not the interleaved cascade. + +**E-B2-5 — Strided-resonance sweep and the dither rebuttal (gates A4).** +*Metric:* pages-touched for stride-s lattices, s ∈ {2,4,8,16,32,64}, at page ∈ {16, 256, +1024}, undithered vs with an odd-stride comma permutation on the low tier. +*Baseline:* raster (analytic `|Q|·s/2^j`). +*Control:* random subsets of the same cardinality (the no-structure floor). +*Kill:* if the dither brings Morton within 1.2× of raster at the worst stride s = √P, A4 is +rebutted and the coprime walk should be made part of the addressing contract. If it cannot +— and note the power-of-two-modulus obstruction — record the √P penalty as a standing cost +of enlarging the field. + +**E-B2-6 — Axis-assignment A/B (gates A8).** +*Metric:* IO ranges and pages for taxonomy-major vs mereology-major traversals, under +`morton() = spread8(lo) | spread8(hi)<<1` and its swap. +*Baseline:* Hilbert (symmetric — predicted no change under swap). +*Control:* a synthetic traversal mix with a known 50/50 split (predicted: no difference, so +any measured difference is harness bias). +*Kill:* if the dominant traversal is mereology-major, the current bit assignment costs 2× +and should be swapped; either way, the result must be written into `facet.rs`'s doc so the +choice stops being implicit. + +**E-B2-7 — The F−R attribution test (gates A9).** +*Metric:* wall time and cache misses for `lane_f_morton` vs `lane_r_radix`, plus the +*distributional* check — collision counts and probe-length histograms for both slot +functions over the same input. +*Baseline:* the analytic prediction: identical distributions, F slower by the interleave's +ALU cost. +*Control:* replace the FNV digest with a locality-sensitive code (a CAM-PQ centroid pair) +and re-run; predicted: F now beats R by a real margin. +*Kill:* if the distributions differ at all, my analysis is wrong. If they are identical and +F is faster, the difference is code layout/branch effects and must not be reported as +clustering. + +**E-B2-8 — Page ORDER (the gap A11 opens; nothing currently measures it).** +*Metric:* read-ahead effectiveness and page-index min/max pruning as a function of *page +sequence*: mean spatial distance between consecutive pages in file order, and pruning rate +of a zone map built over page-level (x,y) extents. +*Baseline:* Morton page order. +*Control:* Hilbert page order over the *identical* page partition (A11 guarantees they +exist), plus a random page permutation as the floor. +*Kill:* if Hilbert page order prunes ≥ 15 % better or halves read-ahead misses, the correct +design is **Morton within a page, Hilbert across pages** — a hybrid neither the repo nor +Lance currently expresses, and the most interesting constructive outcome available in +Domain B. + +--- + +## PRIOR ART + +Searches actually run (WebSearch unless noted; alphaXiv `get_paper_content` for P5): + +1. `"Hilbert curve versus Z-order clustering properties Moon Jagadish Faloutsos analysis average number of clusters range query"` +2. `"Databricks liquid clustering Hilbert curve replaces Z-order ZORDER limitations"` +3. `"learned space filling curve LMSFC BMTree bit merging tree beats Z-order Hilbert query-aware layout VLDB SIGMOD"` +4. `"space filling curve locality query shape dependence no single curve optimal elongated stripe query Z-order better than Hilbert"` +5. `"Xu Tirthapura On the optimality of clustering properties of space filling curves PODS 2012 lower bound no curve optimal all query sizes"` +6. `"QUILTS space filling curve query-aware Nishimura Yokota design SFC for specific query workload skewed"` +7. `"Morton order recursive array layout matrix cache performance Chatterjee Wise Frens versus blocked layout dilated arithmetic overhead"` +8. `"Z-order curve skewed data distribution load imbalance quadtree hot spot uniform grid quantization versus rank space quantile partitioning"` +9. `"Delta Lake ZORDER implementation interleave range partition id rank not raw value rangePartitionId z-order skew"` +10. `"Z-order Hilbert identical page partition quadrant aligned block contiguous both curves same set of blocks difference only ordering of blocks"` +11. `"Z-order Morton curve outperforms Hilbert for checkerboard dyadic parity selection queries block-alternating pattern clustering"` +12. alphaXiv full text: *Efficient Cost Modeling of Space-filling Curves* (arXiv 2312.16355) + +**P1 — Moon, Jagadish, Faloutsos, Saltz, "Analysis of the Clustering Properties of the +Hilbert Space-Filling Curve," IEEE TKDE 13(1), 2001.** Closed-form cluster counts for +arbitrary rectilinear query shapes; the canonical 2×2 numbers z-curve **2.625**, Gray +**2.5**, Hilbert **2.0**. My harness reproduces 2.6094 / 1.9994. *Every attack in A3 is +this result restated at larger q; A3 is KNOWN, not novel.* + +**P2 — Xu & Tirthapura, "On the Optimality of Clustering Properties of Space Filling +Curves," PODS 2012.** Corollary 2: for a rectangular query of fixed size, averaged over all +rotations, **any continuous SFC is optimal**. Critical asymmetry for this program: Z-order +is **not** continuous (it has jumps), so it is excluded from that optimality class — which +is precisely the A2 seam. The follow-on *Onion Curve* paper (arXiv 1801.07399) records that +even Hilbert "can be far from optimal, even for cube-shaped queries." + +**P3 — Databricks Liquid Clustering / Delta ZORDER.** Liquid Clustering replaced Z-order +with **Hilbert**, citing exactly the A2 mechanism ("Z-Order jumps generate files with a +large min/max difference, which will make it useless"), plus incremental ZCube clustering +and a reported 2.5× clustering speedup. Apache Hudi likewise offers both Z-order and +Hilbert with Hilbert recommended. *A2 is KNOWN; my contribution is only the magnitude on +this specific geometry (129× span, 46× runs).* + +**P4 — QUILTS (Nishimura & Yokota, SIGMOD 2017); LMSFC; BMTree (VLDB 16:2158, arXiv +2505.01697).** Query-aware and learned bit-merging curves. LMSFC reports **2–4× lower query +time than fixed Z-order or Hilbert**; BMTree learns per-subspace curves via MCTS. *This is +the "a learned order beats Z-order" leg of my charge, and the literature already settles +it in the affirmative — with an important caveat below.* + +**P5 — *Efficient Cost Modeling of Space-filling Curves* (arXiv 2312.16355 / PVLDB).** Read +in full. Two findings that cut both ways: (i) LBMC beats BMTree by **28×** block accesses +on a SKEW dataset (111 vs 3,084) and 6× QUILTS — so on skewed data the learned curves' +advantage is enormous, which amplifies A7; (ii) but *"The BMTree and QUILTS outperform LC, +ZC, and HC on real data such as NYC … **However, there are no consistent results across the +different datasets**"* — i.e. **fixed Z-order sometimes beats the learned curves**, and the +learned methods need per-dataset tuning. Also: LC (lexicographic = raster) is uniformly the +worst on 2-D range queries, which is the honest counterweight to A1. + +**P6 — Delta Lake's ZORDER interleaves `range_partition_id`, not raw values.** The +optimizer samples the RDD to find range boundaries and interleaves the *rank bucket ids*; +`OPTIMIZE_ZORDERBY_NUM_RANGE_IDS` controls their granularity, and +`OPTIMIZE_ZORDERBY_ADD_NOISE` appends a random byte specifically "to deal with skew … but +may have a negative impact on overall min/max skipping effectiveness." **This is direct +prior art against `cascade_key.rs`'s linear choice**: the largest production Z-order +implementation uses the rank transform the workspace's comment explicitly rejects, and +ships an explicit knob for the residual skew. The trade the workspace made silently is the +trade Delta made loudly and in the opposite direction. + +**P7 — Morton layouts for arrays (Wise & Frens; Chatterjee et al.; "Improving the +Performance of Morton Layout by Array Alignment and Loop Unrolling," LCPC 2003; arXiv +2309.07002 on evolved generalized Morton layouts).** Recursive/Morton layouts beat +canonical layouts by 1.2–2.5× for standard matmul, and the known cost is that +"recursive data layouts require non-traditional address computation … bit-level +manipulations that are not supported in current processors," handled by dilated arithmetic +or table lookup. Confirms benefit #5's premise is real but paid for. *Morton-hybrid* layouts +(Morton down to a tile, then raster inside) are the standard fix and are exactly the +"Morton across pages, raster within" shape that E-B2-8's control probes from the other side. + +**P8 — Skew/partitioning literature.** Fixed-grid partitioning assumes uniformity and +degrades under skew; quadtree/R-tree splits are more balanced; ε-hot-spot tracking is a +multidimensional generalization of approximate quantile summaries (SpatialHadoop skewness +work; adaptive spatial partitioning). *A7's mechanism is KNOWN; its magnitude on this +geometry is new only in the sense that nobody had measured this geometry.* + +**P9 — In-repo prior art, which I weight highest.** `rail_geometry.rs:24-30` records a real +bake choosing the de-interleaved per-axis layout over the interleaved pairs, 99.62 % vs +44.25 %. `splat.rs:169-186` records the Morton seam as a known aliasing hazard. +`lane_f.rs:89-95` records the correct radix control. The workspace has been adversarial +about this before; my job here was mostly to quantify. + +**Search that returned nothing:** no prior art found for the A5 result (Z-order beating +Hilbert by exactly 2× on dyadic-parity/checkerboard selections) — search 11 explicitly +returned "does not appear in the search results." Classified accordingly below. + +--- + +## SURPRISES + +**S-B2-1 — At page = 4^j, Morton and Hilbert induce the *identical* page partition, so +pages-touched is identical for every query; they differ only in page ORDER. The entire +Hilbert-over-Z clustering advantage is a run-count effect, invisible to a paged columnar +reader — and the program's 8 KiB / 16-row field is exactly 4².** +*Classification:* **KNOWN** for the partition fact (both curves are quadrant-refining; a +patent in the search-10 results states plainly that both "partition pages into the same sets +of aligned blocks at quadrant boundaries … they traverse the same block structure but in a +different order"). **NOVEL-CANDIDATE** for the *consequence* — I found no source stating +that this makes the Moon et al. advantage vanish on the bytes-read metric at 4^j pages, and +searches 1, 4, and 10 were aimed squarely at it. Measured, verified two-sided (partitions +diverge at P = 32 and P = 128). + +**S-B2-2 — The measured mean code-span is *worse* for Hilbert than Morton on square +queries: 2120 vs 1737 (q=8), 4794 vs 4034 (q=16), 10345 vs 8777 (q=32) — 18–22 % worse — +while Hilbert's *median* span is better (344 vs 424). The industry's stated reason for +preferring Hilbert (better min/max ⇒ better file skipping) is a tail-behaviour claim, and +on the mean it points the other way at this scale.** +*Classification:* **DISPROVED** (the popular formulation, in this regime). The literature's +defensible claim is about *cluster count*, which Hilbert genuinely wins; the min/max +framing repeated in the lakehouse blogs does not survive measurement here. Caveat stated +plainly: my span metric models a min/max over the *clustering-key column*; a zone map over +the underlying (x, y) columns behaves differently, and raster's span (1800 / 3856 / 7968, +constant across offsets) beats both curves on this metric at every q. + +**S-B2-3 — Z-order beats Hilbert by exactly 2.0× and raster by 4–16× on dyadic-parity +(checkerboard) selections, at every block size. The attack I designed to hurt Morton +rehabilitates it — and the mechanism (the two parity bits are adjacent in a Morton code, +scrambled in a Hilbert code) predicts the same win for any XOR/residue predicate, including +2-D erasure-stripe membership.** +*Classification:* **NOVEL-CANDIDATE.** Search 11 documented no prior art; the result is +elementary once seen, so I expect it exists somewhere, but I could not find it. Flagged for +Domain C/D: if parity groups are a 2-D residue class, this is a genuine cross-domain +argument for Z-order over Hilbert that has nothing to do with range-query locality. + +**S-B2-4 — `cascade_key.rs`'s documented "linear (not rank)" requirement is a trilemma: +quadtree ancestry ⟹ equal-width quantization ⟹ 20–45× petal hot-spots with 70–84 % empty +fields on clustered identities. Delta Lake resolved the same trilemma in the opposite +direction (interleaving rank buckets, plus an explicit noise knob), and a per-axis rank +transform only halves the multi-modal case (21.19× → 9.44×) because marginal quantiles +cannot decorrelate 2-D clusters.** +*Classification:* **TRANSFER.** The skew-vs-uniform-grid mechanism is KNOWN (P6, P8); what +transfers is the observation that the workspace's ancestry constraint *forbids* the standard +fix, and that the constraint is stated in a doc comment as a benefit with no measurement of +its cost. The escape — a hierarchically trained 4⁴ centroid codebook, which is a +rank transform that *preserves* ancestry — is described in the OGAR canon but I found no +trainer; E-B2-1 is designed to settle it. + +**S-B2-5 — Morton's axis anisotropy is exactly 2.000× at every scale, and `facet.rs:62` +puts `lo`/is_a/Taxonomy on the cheap even bits and `hi`/part_of/Mereology on the expensive +odd bits. A mereology-major traversal therefore costs exactly twice a taxonomy-major one, +and nothing records the choice.** +*Classification:* **TRANSFER.** Z-order's dimension-order sensitivity is KNOWN (it is why +"which column comes first in ZORDER BY" is a documented tuning question, and it is the +entire premise of QUILTS' bit-merging patterns, P4). The transfer is binding it to the +workspace's own semantic axis assignment and to the `RailAxis::{Taxonomy, Mereology}` +zero-fallback at `rail_geometry.rs:80-88`, where it is currently accidental. + +--- + +## VERDICT + +**Morton survives as a defensible minimax default — and three of the four things it is +credited with do not survive.** + +What survives: aligned-prefix queries as a single run (a real win over raster, worth 2^j +runs); dyadic-parity selections, where it beats Hilbert 2× and raster 4–16× (A5, and it may +matter to the erasure domain); ~5 ALU ops of encode against Hilbert's 45-line state machine; +O(1) rigid address translation; and minimax behaviour on axis-aligned thin queries at 64 +pages against raster's {16, 256}. + +What does not survive. **(1) "Morton gives better locality" is either false or vacuous +depending on the metric**: 1.31–1.78× *worse* than Hilbert on run count, growing with query +size, and provably *identical* to Hilbert on pages touched at every 4^j page size — which is +every page size this architecture uses (A3, A11). **(2) The uniformity premise is the +weakest link, not the curve.** The source's own `linear (not rank)` requirement, applied to +spectral-embedding coordinates that are cluster-heavy by construction, produces 20–45× +petal hot-spots and 70–84 % empty fields, against Delta's contrary and battle-tested choice +(A7, P6). **(3) `morton_slot` over an FNV digest is a bit permutation of a uniform variate, +not a locality code** — the 1BRC "Morton lane" cannot support a layout conclusion (A9). +**(4) The layout question is currently moot in the shipped write path**: the persist seam +orders by `stream_position`, contractually a monotone stream key that a Morton code cannot +be, and the only hook is `SweepSlot::row` at identity-minting time (S4). + +And the environment is hostile in two specific ways worth stating plainly: Lance ships +**Hilbert and only Hilbert** in both columns, gated behind `geo` and scoped to R-tree +bulk-load, with no Z-order anywhere; and Lance 9.0.0 compaction is contractually +**order-preserving**, so cross-cycle Morton interleaving is never repaired by the engine +(A10). *"Is compaction repair of physical locality?"* — on this engine, provably not. + +Two attacks failed and I record them as such: the checkerboard **backfired** into Morton's +best result (A5), and the strided-lattice attack has a plausible in-repo rebuttal already +built (the coprime comma walk, S6/A4) that I could not evaluate because the cascade's moduli +are powers of two where the literal stride-4-over-17 does not apply. + +**The decisive next measurement is not a curve comparison.** It is E-B2-1: the real +identity-occupancy census. If occupancy is near-uniform, A7 dies, A11 makes the +Morton/Hilbert choice a CPU question, and the cascade is sound. If occupancy is as skewed as +a spectral embedding suggests, then no space-filling curve is the problem or the solution — +the quantizer is — and the seal, parity-group, and compaction-boundary designs that assume +equal field occupancy all need re-founding before layout is worth another hour. Second +priority is E-B2-8, the one genuinely open question A11 exposes: **page order**, where a +"Morton within a page, Hilbert across pages" hybrid is expressible, unmeasured, and +unclaimed by anyone. diff --git a/docs/lotus/rp-seal-v1/B3.md b/docs/lotus/rp-seal-v1/B3.md new file mode 100644 index 00000000..94589581 --- /dev/null +++ b/docs/lotus/rp-seal-v1/B3.md @@ -0,0 +1,692 @@ +# B3 — Space-Filling Curve Literature Scout (Domain B: Spatial Layout / Morton Cascade) + +Role: Literature Scout. This report does not test or implement the proposed +Morton-cascade / bidirectional-seal architecture. It retrieves and extracts the +three assigned papers plus adjacent literature, and asks a narrow question: +**does the published record support treating a single Hilbert- or Morton-style +curve as a safe default for a 2-D→hierarchical grid store's compaction / +locality layer, or does it document specific, provable conditions under which +that default fails?** + +--- + +## SOURCE ARCHAEOLOGY + +Retrieved via the alphaXiv MCP tool (full text unless noted) and WebSearch/WebFetch. +Depth of engagement is stated per source so downstream researchers can judge +confidence. + +**Target packet (full text obtained and read in full):** + +1. Pan Xu, Cuong Nguyen, Srikanta Tirthapura, "Onion Curve: A Space Filling + Curve with Near-Optimal Clustering," *ICDE 2018*, pp. 1236–1239. + arXiv:1801.07399v2 [cs.CG]. — full text, all sections read (definitions, + both 2-D and 3-D theorems, lower bounds, experiments, conclusion). +2. Christian Böhm, "Space-filling Curves for High-performance Data Mining," + arXiv:2008.01684v1 [cs.LG]. — full text, all sections. +3. Liang Zhou, Chris R. Johnson, Daniel Weiskopf, "Data-Driven Space-Filling + Curves," *IEEE TVCG* (Vis 2020 track), arXiv:2009.06309v1 [cs.GR]. — full + text, all sections including algorithm listings and evaluation tables. + +**Extended literature (full text or substantial-fraction read):** + +4. Herman Haverkort, "Recursive tilings and space-filling curves with little + fragmentation," *Journal of Computational Geometry* 2(1):92–127, 2011. + arXiv:1002.1843v1 [cs.CG]. — abstract, introduction, related work, and full + results statement (§1.4) read in detail; deep proof sections (§2–3, tilings + with Arrwwid number 3) sampled but not fully verified line-by-line. +5. Igor V. Netay, "Cyclic space-filling curves and their clustering property," + arXiv:2006.10286v1 [cs.DS], 2020. — full text, including its independent + reproduction of the Moon-Jagadish-Faloutsos-Saltz simulation table. +6. Jian Gao, Xin Cao, Xin Yao, Gong Zhang, Wei Wang, "LMSFC: A Novel + Multidimensional Index based on Learned Monotonic Space Filling Curves," + *PVLDB* 16(10):2605–2617, 2023. arXiv:2304.12635v4 [cs.DB]. — abstract + + introduction read in full (motivation, monotonic-vs-non-monotonic SFC + framing, three stated limitations of prior learned-SFC work); body sections + not read. +7. Guanli Liu, Lars Kulik, Christian S. Jensen, Tianyi Li, Jianzhong Qi, + "Efficient Cost Modeling of Space-filling Curves," arXiv:2312.16355v1 + [cs.DB], 2023. — abstract + introduction read in full, including the + explicit q1/q2/q3 worked counterexample figure description. + +**Cited-through / secondary-source engagement (not independently fetched in full):** + +8. B. Moon, H. V. Jagadish, C. Faloutsos, J. H. Saltz, "Analysis of the + Clustering Properties of the Hilbert Space-Filling Curve," *IEEE TKDE* + 13(1):124–141, 2001. — the canonical clustering-formula paper. A direct + WebFetch of the PDF failed (binary/compressed stream, no extractable text + layer available to the fetch tool). Its results are triangulated from two + independent citing sources that quote or reproduce its content: (a) the + Onion Curve paper's Related Work (§I-B), which states Moon et al.'s + asymptotic result in words; (b) Netay (2020) §3.2, which explicitly + re-runs Moon et al.'s experimental protocol and reproduces a full + d=2,3,4 simulation table attributed to "[3]" = Moon et al. This is + corroboration, not a substitute for reading the original formulas — + flagged as **secondary** throughout. +9. Chan, Har-Peled, Jones, "On Locality-Sensitive Orderings and their + Applications," arXiv:1809.11147, ITCS 2019 / SICOMP. — abstract-level via + WebFetch (the fetch tool summarized the abstract and definitions; full + proof content not retrieved). +10. Bender, Demaine, Farach-Colton, "Cache-Oblivious B-Trees," FOCS 2000 / + SICOMP; Frigo, Leiserson, Prokop, Ramachandran, "Cache-Oblivious + Algorithms," FOCS 1999 — engaged only via WebSearch result summaries + (van Emde Boas layout mechanism, ideal-cache model definition). Not + independently fetched; standard, extremely well-established results, + treated as KNOWN background rather than closely audited primary sources. + +**Titles located but not analyzed** (found via WebSearch listings only, cited +for completeness of the survey, not used as evidentiary basis for any claim +below): "The Case for Learned Spatial Indexes" (arXiv:2008.10349); "Evaluating +Learned Spatial Indexes" (arXiv:2606.19034); "A Survey of Learned Indexes for +the Multi-dimensional Space" (arXiv:2403.06456); WAZI, ZMINDEX, Flood, +Tsunami, Qd-tree (named and one-line-characterized in WebSearch synthesis +only). + +--- + +## MECHANISM + +### A. The three target papers, field by field + +**1801.07399 "Onion Curve" — SFC clustering metric, query-shape dependence, counterexamples** + +- **Clustering metric.** For an SFC π (bijection U → {0,…,n−1}) and query + q ⊆ U, the *clustering number* c(q,π) is the minimum number of maximal + runs ("clusters") of π-consecutive cells that partition q — i.e., disk + seeks needed to retrieve q if data is stored in π order. For a query + *set* Q, c(Q,π) is the average over q ∈ Q, and the *approximation ratio* + η(Q,π) = c(Q,π)/OPT(Q) is the metric compared across curves. +- **Query-shape dependence, proved not assumed.** The paper's central + finding is that η is not curve-shape-invariant: it depends jointly on + (a) the query's aspect ratio (cube vs. general rectangle) and (b) the + query's **size relative to the universe**, a variable prior clustering + analyses (including Moon et al.) had held constant. Table II in the + paper tabulates η(Q,Onion) and (partially) η(Q,Hilbert) across five + regimes of ℓ (query side length) parameterized as ℓ = φ·n^{1/d·μ} + ψ, + μ ∈ [0,1] — i.e., the query can grow at any power-law rate relative to + the universe, from constant-size (μ=0) to full-size (μ=1, φ→1). +- **Counterexample to Hilbert optimality (Lemma 5, §IV).** For cube + queries whose side length ℓ = ᵈ√n − O(1) (i.e. nearly spanning the whole + grid), the Hilbert curve's clustering number is **Ω(n^{(d−1)/d})** — i.e. + **Ω(√n) in 2-D, Ω(n^{2/3}) in 3-D** — while the onion curve stays + **Θ(1)** on the *same query class*. This is an unbounded gap: as n grows, + Hilbert's approximation ratio for these queries diverges. The paper's own + framing: "the approximation ratio of the Hilbert curve can be unbounded, + even for the case of cube queries in two dimensions" (§I-A, item B). +- **Second counterexample — impossibility, not just Hilbert's failure + (§V-B, Lemmas 10–11).** No single SFC can be near-optimal for all + rectangular query shapes: there exist rectangular query sets such that + optimizing an SFC for one shape provably degrades its performance on a + different shape. This is stronger than "Hilbert isn't perfect" — it is + a proof that *no* curve, onion or otherwise, can dodge the trade-off. +- **Onion curve construction (§III-A).** Cells are ordered by *layer*: + distance-to-boundary ∇(α) = min(x+1, √n−x, y+1, √n−y); layer S(t) = {α : + ∇(α)=t}; the curve visits all of S(1), then S(2), … S(m). This is a + fundamentally different partition axis from bit-interleaving (Z-order) + or recursive-quadrant subdivision (Hilbert/Peano): it is *radial/onion* + rather than *recursive/quadrant*. +- **Quantified bounds.** 2-D onion curve: η ≤ 2.32 for cube queries of any + size. 3-D: η ≤ 3.4. These are proven upper bounds, not simulation + results, though the paper also reports matching simulation confirmation + (Figs. 5–7): for ℓ > n^{1/d}/2, onion beats Hilbert by "more than 200 + times" in 3-D at the largest tested cube sizes, and is never worse than + Hilbert at any tested size. + +**2008.01684 "Space-filling Curves for High-performance Data Mining" (Böhm) — locality preservation, cache-oblivious applications, transformation cost** + +- **Locality preservation mechanism.** The paper frames space-filling + curves as an alternative to explicit cache-blocking: a Hilbert loop over + (i,j) acts "like the nesting of a high number (2·log₂ n) of loops going + forward and backward with different step-sizes," so a single traversal + order is simultaneously good at *every* cache/block granularity — the + cache-oblivious property. Fig. 1 in the paper shows measured cache-miss + curves: nested loops show a sharp knee once the working set exceeds + cache size; the Hilbert loop degrades gracefully across a wide range of + cache sizes (5–20% of main memory cited as the realistic regime where + the improvement is largest). +- **Transformation cost — the paper's specific technical contribution.** + Naive Hilbert coordinate↔index conversion via a Mealy automaton costs + **O(log n)** per pair (bit-serial state machine, Fig. 3). The paper's + contribution is collapsing this to **O(1) amortized** two ways: (a) a + context-free-grammar / Lindenmayer-system loop generator (§4) that + emits the traversal as a constant-time-per-step recursive/iterative + procedure rather than re-deriving coordinates from the index every + iteration; (b) a fully non-recursive, constant-space, constant-time + formulation (§5, Fig. 5) using `tzcnt` (trailing-zero-count hardware + instruction) to directly determine the current recursion level and + movement direction from the incrementing Hilbert value — explicitly + citing PDEP/PEXT (parallel bit deposit/extract, BMI2) as the relevant + hardware primitive family for bit-interleave curves generally. +- **Non-square-grid cost.** Standard Hilbert curves require n = 2^L per + axis. The paper's "overlay-grid" (FUR-Hilbert-Loop) and "jump-over" + (FGF-Hilbert-Loop) techniques handle arbitrary/non-power-of-two and even + triangular iteration domains at constant amortized overhead per element, + avoiding the naive "round up to next power of two, discard the rest" + approach's up-to-3× (near-square) or unbounded (highly rectangular) + waste. +- **Applications demonstrated cache-oblivious via this technique:** dense + matrix multiplication, Cholesky decomposition, Floyd-Warshall transitive + closure, k-means clustering, and — via the jump-over variant — similarity + join over index-structure-restricted candidate pairs (cites Perdacher, + Plant, Böhm, SIGMOD 2019, for the join application specifically). + +**2009.06309 "Data-Driven Space-Filling Curves" (Zhou, Johnson, Weiskopf) — adaptive/data-conditioned ordering, multiscale/quadtree implications** + +- **Adaptive/data-conditioned mechanism.** Rather than a fixed geometric + recursion, the curve is the (approximate) minimum-weight Hamiltonian + path over the data's adjacency graph, with per-edge weight + W = (1−α)·N(value-difference) + α·R(positional coherence), α ∈ [0,1] + user-tunable. This directly instantiates the top-level prompt's + "data-driven/adaptive ordering" question: the curve is literally + computed from the data instance, not from grid geometry alone. + Regular-grid construction (§4): grid → 2×2 (2-D) / 2×2×2 (3-D) "circuit" + graph → dual graph → minimum spanning tree (Prim's) → Hamiltonian cycle + via cover-and-merge → cut to path. Complexity/cost is not asymptotically + characterized by the paper in closed form; empirically Table 1 reports + wall-clock times from 12 s (64×64) to 4h31m (128³ voxels) on a 2019 + laptop, i.e. the adaptive curve is **not** a constant-per-element + computation like Hilbert/Z-order — it is a one-time O(graph-size) + preprocessing cost, amortized over subsequent queries, explicitly not + suitable for online/incremental re-derivation. +- **Multiscale/quadtree implications (§5).** For non-uniform-depth data + (quadtree/octree), a single global minimum-Hamiltonian-path formulation + is infeasible (uneven vertex counts per dimension defeat the regular-grid + circuit-graph trick, and a Hamiltonian cycle may not even exist). The + paper's fallback is a **top-down, block-by-block greedy approximation**: + solve the coarsest level's SFC first, then recursively refine each block, + matching entry points across scale boundaries via nearest-value + heuristic (Algorithm 2/3). Crucially, the paper's own evaluation + (§6, Fig. 13) finds the **regular-grid technique has strictly better + data-value coherency than its own multiscale/quadtree technique** — the + top-down greedy refinement demonstrably breaks coherent spatial features + at block boundaries (Fig. 12) — but the multiscale variant is faster and + is the only option when data genuinely has non-uniform density/scale + (e.g., particle simulations, where octree-per-particle resolution is + natural but a uniform regular grid is not). +- **Locality vs. feature-coherency trade-off, quantified.** Autocorrelation + of *data value* (feature coherency) favors the data-driven curve; the + autocorrelation of *radial Euclidean distance* (pure spatial locality) + favors the fixed Peano-Hilbert curve (Fig. 13, both 2-D and 3-D). The + data-driven curve is a deliberate trade of pure locality for + data-feature coherency, tunable via α — it does **not** dominate Hilbert + on the locality axis, only on the value-coherency axis. This is an + important nuance: "data-driven beats fixed curves" is true only for the + metric the paper optimizes for (value coherency under brushing/boxplot + visualization), not universally. + +### B. Beyond the packet + +**Moon, Jagadish, Faloutsos, Saltz (2001) — the canonical clustering-formula result (secondary-sourced, see caveat in Source Archaeology).** +The Onion Curve paper's Related Work states the headline result in words: +*"if the size of the query remains constant, the clustering number of the +Hilbert curve is asymptotically equal to the surface area of the query shape +divided by twice the number of dimensions."* Netay (2020) independently +reproduces the underlying simulation protocol for cube queries and gets, e.g. +at d=2, ℓ=15 on a 1024×1024 grid: Z-curve 28.17 clusters, Hilbert 15.04, +H-curve (Netay's own curve) 15.02 — i.e., for *constant-size* cube queries +Hilbert tracks its own asymptotic formula almost exactly and clearly beats +raw Z-order (roughly a 1.87× gap at ℓ=15, widening at larger ℓ within the +constant-size regime — e.g. ℓ=2: Z=2.62 vs Hilbert=2.00, ratio 1.31×; ℓ=15: +ratio 1.87×). **The critical scope condition, made explicit by the Onion +Curve paper's own Related Work comparison, is that Moon et al.'s asymptotic +result assumes query size is held *constant relative to the universe* as +n→∞** — precisely the regime the Onion Curve paper shows Hilbert is optimal +in, and precisely the regime that breaks down (Ω(√n) gap) once query size is +allowed to scale with the universe. + +**Haverkort (2011) — the Arrwwid number: a formal, dimension-general counterexample to Hilbert/Morton optimality.** +Defines the *Arrwwid number* a(tiling or curve) as the smallest a such that +any ball Q is covered by a curve-fragments (contiguous ranges) with total +volume O(vol(Q)) — a scale-free reformulation of clustering number that, +unlike Moon et al.'s framework, permits comparing curves built on +*different* underlying tilings (not just different orderings of the same +4-square-per-square recursion Hilbert and Z-order share). +- **2-D:** no recursive tiling/curve can beat Arrwwid number 3 (proven lower + bound, Theorem 1). The standard uniform-square tiling underlying both + Hilbert and Z-order has Arrwwid number **4**; Asano et al.'s prior AR²W² + ordering achieves 3 on that same tiling (matching the lower bound); + Haverkort constructs additional tilings (Gosper-island fractal tiles, + the "Daun" rectangular tiling) also achieving the optimal 3, with + *simpler, non-diagonal* curve structure (diagonal connections in AR²W² + are separately shown, citing Bugnion et al., to be harmful to disk-head + random-walk cost models). +- **d ≥ 3, the strongest counterexample found in this sweep (§1.4, + quoted directly):** *"In d dimensions, no recursive tiling has Arrwwid + number better than d+1 … Regular (Hyper)cube-based tilings have Arrwwid + number 2^d, and any space-filling curve based on such a tiling has + Arrwwid number at least 2^d − 1 … For d ≥ 3 this is more than what can be + achieved with rectangular blocks"* (which achieve (3/4)·2^d). Concretely + at d=3: cube-tiling-based curves (which is the family Hilbert and Morton + both belong to) cannot beat Arrwwid 7; a rectangular-block tiling can + reach 6. At d=4: 15 vs. 12. **This is a formally proven, dimension-scaling + disadvantage of the entire cube-recursive curve family (Hilbert + included, Morton/Z-order included) relative to non-cube tilings, once + the indexed space is 3-dimensional or higher** — directly relevant if a + "Morton cascade" is applied jointly across more than 2 real axes (e.g. + two spatial axes plus a temporal/version axis, or the stated 2-D + hierarchy composed with a third orthogonal dimension). + +**Netay (2020) — a practical near-tie alternative, and an evaluation-methodology lesson.** +The H-curve (cyclic construction, no cell-symmetry/reflection bookkeeping) +matches Hilbert's clustering number to within measurement noise across d=2,3,4 +and ℓ up to 15 (see table above/full data in the paper), while its own +profiling shows encode/decode 4–8× faster in wall-clock than the author's +Hilbert baseline implementation (0.06–0.08 ms/1000-calls vs 0.31–0.47 ms/1000-calls, +d=n=7). The methodological point transferable to any layout redesign: a +curve construction can be *computationally* much cheaper than the classical +Hilbert automaton at **statistically indistinguishable** clustering quality — +i.e., "we need Hilbert specifically" and "we need Hilbert's clustering +number" are two different, separable claims, and the literature contains at +least one curve that gets the clustering benefit without paying Hilbert's +per-element automaton cost. + +**Cache-oblivious layout, beyond loops: van Emde Boas layout and cache-oblivious B-trees.** +Frigo, Leiserson, Prokop, Ramachandran's ideal-cache model (FOCS 1999) +formalizes what Böhm's paper exploits informally for loop nests: an algorithm +analyzed against an idealized two-level (cache size M, block size B, +optimal offline replacement) hierarchy, without the algorithm itself knowing +M or B, automatically performs well on *every* real multi-level hierarchy. +Bender, Demaine, Farach-Colton extend the same idea from *array traversal +order* (what Böhm's Hilbert loops do) to *pointer-based tree memory layout*: +the van Emde Boas layout recursively splits a search tree at its middle +level of edges (top recursive subtree + several bottom recursive subtrees of +half height), producing an array embedding that achieves O(1 + log_{B+1} N) +memory transfers per search for *any* B, M simultaneously — the same +"good-at-every-scale-without-knowing-the-scale" guarantee Böhm's paper claims +for Hilbert loops, but applied to trees rather than dense arrays. This is +the same underlying idea (recursive block decomposition, obliviousness to +the specific hierarchy parameters) appearing independently in two different +sub-literatures (dense-array SFC traversal vs. pointer-tree layout). + +**Multidimensional indexing context.** +Standard multidimensional index families (R-tree, kd-tree, quadtree) are +overlap-based spatial partitioning structures, distinct from — but often +composed with — SFC-based linearization (an SFC turns multidimensional keys +into a single sortable key consumable by an ordinary B+-tree). The +LMSFC paper's introduction crisply states the operational trade-off this +sweep kept surfacing: **monotonic SFCs (Z-order/Morton) permit O(1)/cheap +range decomposition of a query into contiguous key ranges via direct bit +manipulation, whereas non-monotonic SFCs (Hilbert) require walking/enumerating +the query boundary to determine which ranges to scan — better clustering, +worse (and shape-dependent) decomposition cost.** This is precisely why +production systems the cost-modeling paper cites (Apache Hudi, Redshift, +SparkSQL) default to Z-order despite its documented inferior clustering +number: the *engineering* cost of query decomposition is a separate axis +from the *clustering-number* metric the theory papers optimize, and Z-order +wins on the former even while provably losing on the latter. + +**Learned / workload-adaptive SFCs.** +The cost-modeling paper (Liu et al. 2023) states the generalizing principle +explicitly and gives a minimal worked counterexample (their Fig. 1): three +different SFCs over the *same* data space are each optimal for a different +one of three queries q1/q2/q3 (single contiguous segment vs. two vs. four +segments touched) — *"no single SFC is optimal for all queries."* This is +the same qualitative claim the Onion Curve paper proves formally for +rectangle shapes (Lemmas 10–11), arrived at independently via a different +(cost-estimation/RL) framing five years later. LMSFC (Gao et al. 2023) +responds by learning a *parameterized monotonic* SFC family per +dataset+workload instance (not selecting among a small fixed menu), explicitly +motivated by exactly this impossibility: "a fixed SFC does not necessarily +work the best on a given dataset instance." The broader survey landscape +(named, not analyzed in depth — see Source Archaeology) includes WAZI +(workload-aware Z-indexing), Flood/Tsunami (learn a preferred dimension + +per-cell 1-D model rather than a full SFC), and Qd-tree (reinforcement +learning over partitioning strategy directly, bypassing SFCs altogether) — +i.e., the field's recent trajectory has moved *away* from "pick one fixed +curve" toward either (a) learned/parameterized curve families or (b) +abandoning the single-linearization framing for direct learned partitioning. + +**Locality-Sensitive Orderings (Chan, Har-Peled, Jones) — the generalization that resolves the impossibility, at a cost.** +Rather than accept "no single curve works for all queries," this line +proves that **O(1/ε^d) simultaneous Z-order-like orderings** (explicitly +described as "extensions of the Z-order") jointly guarantee: for any two +points p, q within Euclidean distance r of each other, at least one of the +O(ε^{-d}) orderings has every point "between" p and q (in that ordering) +within εr of the p–q segment. This directly answers, from a different +literature, the top-level prompt's cross-domain question "can one SFC order +minimize both query locality and RS repair locality" with a documented +**"no, but a small constant number of orderings jointly can"** — a +qualitatively different mitigation strategy than either (a) picking one +best-average curve or (b) an expensive fully learned/instance-optimal curve. + +### C. Synthesis table — query shape × curve → published winner + +| Query shape / regime | Published winner | Basis | +|---|---|---| +| Cube, **constant size** relative to universe (n→∞, ℓ fixed) | Hilbert curve — asymptotically optimal among continuous SFCs | Moon et al. 2001 (via Onion Curve §I-B + Netay 2020 reproduction); Xu & Tirthapura 2014 (cited in Onion Curve §I-B) | +| Cube, **size scaling with universe** (ℓ = Θ(n^{1/d}), esp. ℓ → full side length) | Onion curve (η ≤ 2.32 in 2D / ≤ 3.4 in 3D); Hilbert **unbounded** (Ω(√n) / Ω(n^{2/3})) | Xu, Nguyen, Tirthapura 2018, Lemma 5 + Theorems 4–6 | +| **General rectangular** query set (elongated, arbitrary aspect ratio) | **No single curve** can be near-optimal across all such sets — proven impossibility | Xu, Nguyen, Tirthapura 2018, §V-B Lemmas 10–11 | +| Isotropic ball/region coverage, **d ≥ 3** (Arrwwid sense) | Rectangular-block tilings ((3/4)·2^d) strictly beat any cube-recursive tiling/curve (≥ 2^d − 1), incl. Hilbert & Morton | Haverkort 2011, §1.4 | +| Isotropic region coverage, **d = 2** | AR²W² (Asano et al.) / Haverkort's Gosper- or Daun-tiling curves reach the proven optimum (Arrwwid 3); standard Hilbert/Z-order tiling is Arrwwid 4 | Haverkort 2011 | +| Single contiguous-segment queries (one dominant access shape known in advance) | Z-order (bit-interleave) — by construction the query maps to few, cheaply-decomposed ranges | Liu et al. 2023, Fig. 1 worked example | +| **Mixed** workload of differently-shaped queries, unknown a priori | No fixed curve; either O(1/ε^d) simultaneous locality-sensitive orderings, or a learned/workload-fit monotonic SFC family | Chan/Har-Peled/Jones 2019; Gao et al. 2023 (LMSFC); Liu et al. 2023 | +| Non-uniform / skewed data density, objective = feature-value coherency (not pure locality) | Data-driven Hamiltonian-path curve beats fixed Hilbert/Peano/scanline on value autocorrelation; **loses** to Hilbert on pure spatial-distance autocorrelation | Zhou, Johnson, Weiskopf 2020, §6 Fig. 13 | +| Multiscale/hierarchical data (quadtree/octree, non-uniform resolution) | Top-down adaptive refinement necessary but **measurably worse locality** than a regular-grid curve on data that *can* be regular-gridded; use only when data genuinely can't be | Zhou, Johnson, Weiskopf 2020, §6, §8 | +| Cache/memory-hierarchy-oblivious dense loop traversal (any B, M) | Hilbert loop (or any continuous SFC) — genuinely near-curve-agnostic result, not shape-dependent | Böhm 2020, Fig. 1; Frigo/Leiserson/Prokop/Ramachandran ideal-cache model | +| Pointer-based tree memory layout, cache-oblivious | van Emde Boas layout (not an SFC per se, but the same recursive-oblivious idea applied to trees) | Bender, Demaine, Farach-Colton, FOCS 2000 | + +--- + +## EXPECTED BENEFIT + +What the literature would predict *if* a system commits to a Hilbert- or +Morton-family curve over its logical row/grid space, under the conditions +where the theory says these curves are strong: + +- **Cache/memory-hierarchy obliviousness for dense traversal** (Böhm): + compaction or bulk-scan loops laid out via a continuous SFC should show + graceful cache-miss degradation across a wide range of working-set sizes, + with no single "cliff," versus a naive row-major/nested-loop traversal. + This is a genuinely curve-family-general result (any continuous SFC gets + it), not specific to Hilbert. +- **Near-optimal clustering for small/constant-size range queries** (Moon + et al., reproduced by Netay): if the dominant query shape is a small, + roughly-cube neighborhood (e.g., a local HHTL-tier lookup or a bounded + cascade radius that does *not* scale with total grid size), a continuous + curve (Hilbert or a comparably-cheap alternative such as the H-curve) + should deliver clustering close to its own asymptotic formula, and + markedly (≈1.3×–1.9× in the measured 2-D regime) fewer fragments than + raw Z-order/Morton bit-interleave. +- **O(1)-amortized coordinate↔index conversion** (Böhm §5): with the + non-recursive `tzcnt`-based formulation, the per-element overhead of a + Hilbert-family traversal is asymptotically no worse than a raw bit-interleave + (Z-order/Morton) walk — the classical "Hilbert is slower to compute" + objection is specifically rebutted for the *loop-traversal* use case (not + for arbitrary point-to-index lookups). +- **Bounded fragmentation, formally, in exactly 2 dimensions** (Haverkort): + if the indexed structure is genuinely 2-D and a good ordering is chosen + on the standard tiling, Arrwwid number 3 (the proven 2-D optimum) is + achievable — a hard guarantee, not an asymptotic hand-wave, on the *worst + case* number of fragments any ball-shaped query touches. + +## EXPECTED FAILURE + +Grounded specifically in the proven negative/scope-limiting results found +in this sweep: + +1. **Unbounded clustering degradation as query size approaches grid size** + (Onion Curve Lemma 5): any compaction/query workload whose access + pattern scans a large fraction of one axis (e.g., "read a whole + classid's worth of rows," "scan from the current tick back to a + deep-history awareness horizon," or any range whose extent grows with + the store rather than staying bounded) should show Hilbert- or + Morton-family clustering number growing as Ω(√n) (2-D) / Ω(n^{2/3}) + (3-D+) relative to the achievable optimum, even though the *shape* + (cube-like) is exactly the shape these curves are supposed to be good + at. The failure is triggered by *size relative to universe*, not shape + — an easy variable to overlook if benchmarking only at small/fixed + query sizes. +2. **d ≥ 3 sub-optimality is a *proof*, not a heuristic concern** + (Haverkort §1.4): any layout that jointly indexes more than two + genuinely orthogonal axes (two spatial + one temporal/version axis is + the generic case in a versioned store) via a cube-recursive + Hilbert/Morton-style curve cannot achieve the best possible worst-case + fragmentation bound — a rectangular-block tiling provably does better, + and the gap widens geometrically with dimension (2^d − 1 vs. (3/4)·2^d). + If the architecture's cascade is described as "2-D" but is in practice + composed with a third orthogonal indexing axis (e.g. Lance + DatasetVersion, or a separate compaction/parity-group axis), this + counterexample applies directly and should be checked, not assumed away. +3. **No single curve for heterogeneous query shapes — a proof of + impossibility, twice independently arrived at** (Onion Curve Lemmas + 10–11; Liu et al. 2023 Fig. 1): if the system serves both narrow + single-column/single-classid scans *and* roughly-square + neighborhood/cascade-radius lookups from the same physical layout, no + fixed curve (Hilbert, Morton, onion, or any other single choice) can be + near-optimal for both simultaneously. A benchmark that only exercises + one query shape class will not detect this; the mixed-workload + regime is where the failure appears. +4. **Raw Z-order/Morton (bit-interleave, discontinuous) is measurably + worse than any continuous curve on clustering, for every tested cube + size** (Netay's reproduction of Moon et al.'s protocol: ~1.3×–1.9× + more clusters at d=2, growing gaps at d=3,4). If "Morton cascade" in + the hypothesis means literal interleaved-bit ordering rather than a + continuous curve built on the same recursive tiling, this is a + documented, not merely theoretical, clustering deficit — traded, in + production systems, for cheaper range decomposition (a genuinely + different, non-clustering cost axis per the LMSFC/cost-modeling + papers), which is a legitimate but *separate* justification that + should be argued on its own terms rather than conflated with "good + locality." +5. **Adaptive/data-driven curves that improve on Hilbert do so at real, + non-incremental cost** (Zhou et al.: 12 s to 4h31m one-time + preprocessing depending on grid size) and **trade away pure spatial + locality for value coherency** — they are not a drop-in win, and their + own top-down multiscale variant (needed for non-uniform hierarchical + data) is measurably *worse* at locality than their own regular-grid + variant. Any claim that "a data-driven/adaptive ordering strictly + dominates a fixed geometric one" should be checked against which axis + (locality vs. feature coherency) is actually being optimized. + +## EXPERIMENT + +Three pre-registered designs, drawn from the program's metric list. + +**EXP-B3-1: Clustering-number / physical-pages-touched sweep across query-size regimes.** +- *Metric*: fragment (cluster) count per query, as a direct proxy for + files-and-pages-touched-per-query and point/neighborhood/random-take + latency. +- *Baseline*: current row-major / lexicographic key order (or whatever the + system's actual un-cascaded default is). +- *Design*: replay three query-shape classes — (a) small cube/near-cube + (side length constant, e.g. one HHTL tier or one 4×4 field), (b) cube + scaling with grid extent up to ℓ ≈ 0.9·(grid side), (c) thin/elongated + rectangular ranges — against three candidate layouts: raw Morton + (bit-interleave), Hilbert (or the cheaper H-curve variant), and a + rectangular-block or onion-style layout, at 3+ grid sizes to see scaling + behavior, not just one point. +- *Control*: identical query set replayed against every layout; + fragment-count ratios computed pairwise, not just absolute numbers. +- *Kill condition*: if the Morton:Onion (or Morton:Hilbert) fragment-count + ratio for regime (b) at the largest tested ℓ is **< 1.5×** — i.e., far + short of reproducing even a fraction of the theory's Ω(√n) asymptotic gap + — the claim that "near-full-extent scans are a real locality liability + for this system's Morton cascade" is not empirically supported at the + scales this system actually operates at, and the amortization argument + for the cascade should be re-derived from measured, not asymptotic, + numbers. + +**EXP-B3-2: Arrwwid-style fragmentation at effective d ≥ 3.** +- *Metric*: index remap bytes/CPU (directly downstream of fragment count), + compaction CPU/wall. +- *Baseline*: the system's actual cube/Morton-style tiling across however + many jointly-indexed orthogonal axes it really has (spatial × spatial × + temporal/version, if that composition exists). +- *Control*: an explicit rectangular-block (non-cube) tiling variant sized + per Haverkort's construction, at the same effective d. +- *Design*: measure fragment count for isotropic (ball-shaped, in the + joint index space) queries at d = 2 vs. d = 3 (with the third axis + actually populated with realistic version/temporal spread, not held at + a single value) — the theory predicts a *qualitative* regime change at + d = 3, not just a quantitative one. +- *Kill condition*: if the rectangular-block control shows **no measurable + reduction** in fragment count or compaction CPU relative to the cube + baseline at realistic (finite, not asymptotic) query and grid sizes, the + Arrwwid counterexample — while formally proven — does not transfer to + this system's actual operating regime, and finite-size effects (small + constant multipliers, cache effects unrelated to the asymptotic bound) + dominate instead. + +**EXP-B3-3: Mixed-workload curve selection.** +- *Metric*: logical & physical bytes read, write/read amplification, + point/neighborhood/random-take latency, aggregated over a *realistic + mixed* query trace (not a single shape class). +- *Baseline*: single fixed Morton curve; single fixed Hilbert curve + (two baselines, to separate "any continuous curve" effects from + "Hilbert specifically" effects). +- *Control*: (c) a small ensemble of O(1/ε^d) locality-sensitive orderings + per Chan/Har-Peled/Jones, each a Z-order variant, selected per-query by + the cheapest applicable one; (d) a workload-fit monotonic curve in the + LMSFC style, refit periodically as the query trace evolves. +- *Design*: the trace must genuinely mix query shapes/sizes (point lookups, + bounded-radius neighborhood, near-full-column scans) in proportions + measured from real production access patterns, not synthesized to favor + any one candidate. +- *Kill condition*: if the multi-ordering ensemble (c) does **not** beat + the single-curve baselines on the mixed trace, first check whether the + trace is actually mixed (a degenerate, single-shape-dominated trace + would falsify the *relevance* of the LSO result to this workload, not + the LSO theorem itself, which is proven for worst-case adversarial + mixes) before concluding the technique doesn't help here. + +--- + +## PRIOR ART + +Searches actually run this session (verbatim intent, not exhaustive strings): + +- alphaXiv `get_paper_content` on arXiv:1801.07399 (Onion Curve) — full text, + read across two tool calls (initial 400-line window + targeted offsets at + 700, 1376, 1550, 2900 to reach the 2-D/3-D theorem statements, Hilbert + lower bound, and experimental/conclusion sections of a 3,463-line + extraction). +- alphaXiv `get_paper_content` on arXiv:2008.01684 (Böhm) — returned in full + directly (fit under the token cap). +- alphaXiv `get_paper_content` on arXiv:2009.06309 (Zhou/Johnson/Weiskopf) — + large-output path; read from the persisted JSON in full. +- alphaXiv `get_paper_content` on arXiv:1002.1843 (Haverkort) — large-output + path; grepped for `Arrwwid|Theorem|optimal|abstract` to locate the results + statement, then read lines 1–260 in full for abstract/intro/related + work/§1.4 results summary. +- alphaXiv `get_paper_content` on arXiv:2006.10286 (Netay, H-curve) — full + text, fit directly. +- alphaXiv `get_paper_content` on arXiv:2304.12635 (LMSFC) and arXiv:2312.16355 + (cost-modeling) — both large-output; read the first ~150 lines + (abstract+introduction) of each from the persisted extraction. +- WebSearch: `Moon Jagadish Faloutsos Saltz "Analysis of the clustering + properties of the Hilbert space-filling curve" 2001 TKDE` — located the + paper's bibliographic identity and abstract summary; direct WebFetch of + the PDF (aiichironakano.github.io mirror) subsequently failed to yield + extractable text (binary/compressed stream) — documented as a negative + result, triangulated via citing sources instead (see Source Archaeology + item 8). +- WebSearch: `Haverkort "Recursive tilings and space-filling curves with + little fragmentation" Journal of Computational Geometry` — confirmed + venue/date and got the Arrwwid-number one-line definition independently + of the full-text read, as a cross-check. +- WebSearch: `learned space-filling curve database index arxiv Z-order + Hilbert query optimality` — surfaced LMSFC, ZMINDEX, WAZI, the cost- + modeling paper, and the learned-index survey. +- WebSearch: `van Emde Boas layout cache-oblivious search trees Bender + Demaine Farach-Colton` — cache-oblivious-B-tree family, beyond the packet's + array/loop-centric framing. +- WebSearch: `"cache-oblivious" B-tree survey Frigo Leiserson Prokop + Ramachandran 1999 FOCS ideal cache` — confirmed the ideal-cache model + definition underlying Böhm's informal cache-oblivious-loop claim. +- WebSearch: `learned space-filling curve Qd-tree Flood LISA multidimensional + index arxiv workload-aware` — located the broader learned-multidimensional- + index survey landscape (not deep-read; titles only, see Source Archaeology). +- WebSearch: `"Z-order" versus "Hilbert curve" thin rectangle query shape when + Z-order wins database benchmark` — **documented negative result**: could + not find an academic source specifically isolating "thin rectangle ⇒ + Z-order wins" as a benchmarked claim; industry blog consensus (Databricks/ + Hudi) is "Hilbert generally clusters better, at higher decomposition cost." + This search is the basis for *not* asserting a "thin query ⇒ Z-order wins + on clustering" row in the synthesis table — the closest substantiated + claim found is Z-order's *decomposition-cost* advantage (cost-modeling + paper), not a clustering-number win for any particular shape. +- WebSearch: `"onion curve" space filling curve compaction OR versioned OR + temporal storage application` — **documented negative result**: no + application of the onion curve (or radial/layer-ordering generally) to + compaction, LSM/versioned storage, or temporal-horizon indexing was found. + Used as the basis for classifying the onion-curve↔version-horizon + connection noted in Surprises as NOVEL-CANDIDATE rather than KNOWN. +- WebSearch: `space-filling curve locality query scatter erasure coding + repair locality shared ordering` — **documented negative result**: no + paper connecting SFC clustering-number theory directly to erasure-coding + repair locality was found in this pass; surfaced the adjacent, more + general Locality-Sensitive-Orderings line instead. +- WebFetch on arXiv:1809.11147 (Locality-Sensitive Orderings) — abstract + + definitions extracted; full proof content not retrieved (fetch tool + summarized rather than returning raw text). +- WebFetch on the Moon et al. PDF mirror — failed (binary stream, no text + layer available to the tool); recorded as a negative/failed retrieval, + not silently dropped. + +--- + +## SURPRISES + +1. **Finding:** Regular cube-based tilings — the family both Hilbert *and* + raw Morton/Z-order belong to — provably cannot achieve the best possible + Arrwwid (worst-case fragmentation) number once the indexed space is 3 + or more dimensions; a rectangular-block tiling strictly beats them, + (3/4)·2^d vs. ≥ 2^d − 1, with the gap widening geometrically in d. + **Classification: KNOWN** — a rigorously proven result in Haverkort + (2011), itself extending Asano, Ranjan, Roos, Welzl, Widmayer (1997). + Not new to the literature; likely to be a surprise only relative to an + *implicit assumption* that "Hilbert/Morton-style curves are the right + default regardless of dimension." + +2. **Finding:** A single curve's approximation ratio relative to optimal + clustering is not determined by query *shape* alone (as most classical + analyses, including Moon et al., implicitly assume by holding query + size constant) — it depends jointly on shape *and* size relative to the + universe, and can be literally unbounded (Ω(√n)) for the same shape + (cube) that the curve is asymptotically optimal for at constant size. + **Classification: KNOWN** — proven in Xu, Nguyen, Tirthapura (2018); + flagged here because it is the kind of result that is easy to miss if a + benchmark only ever tests small, fixed-size queries and extrapolates. + +3. **Finding:** The "obliviousness across scale" idea underlying + cache-oblivious space-filling-curve loop traversal (Böhm, building on + Frigo/Leiserson/Prokop/Ramachandran's ideal-cache model) and the idea + underlying cache-oblivious pointer-tree layout (van Emde Boas layout, + Bender/Demaine/Farach-Colton) are the *same* recursive-block-decomposition + principle, arrived at independently for dense-array traversal order + versus pointer-tree memory placement. **Classification: TRANSFER** — a + known technique (block-recursive layout that performs well at every + scale without parameterizing on the scale) appearing at a different + layer (trees vs. arrays) of the same problem family; worth checking + whether any tree-shaped structures in the architecture under study + (e.g. an index remap structure, a version-history tree) could borrow + the van Emde Boas layout directly rather than only the SFC-traversal + idea. + +4. **Finding:** The Onion Curve's core idea — order cells by + *distance-to-boundary* (radial/layer decomposition) rather than by + recursive quadrant subdivision (Hilbert/Peano) or bit-interleave + (Morton) — is structurally a nested/onion-peel ordering. No published + application of this radial-layer ordering principle to versioned/ + temporal storage, compaction ordering, or "distance from current + awareness horizon" indexing was found in this literature pass (two + independent negative searches, see Prior Art). Given that "distance + from a moving reference point" (current tick, current write horizon, + compaction frontier) is a natural analogue of "distance from the grid + boundary," and the onion curve is specifically near-optimal for + queries that scale with the whole universe — precisely the failure + regime of Hilbert/Morton — this is worth a dedicated probe rather than + assumed transfer. **Classification: NOVEL-CANDIDATE** (documented + negative search; the structural analogy is this session's own + observation, not sourced from any paper read, and is offered as a + testable hypothesis, not a finding). + +5. **Finding:** No single space-filling curve is optimal across a + heterogeneous query workload — proven twice, independently, five years + apart, via completely different methods: a formal impossibility proof + over rectangle shapes (Xu, Nguyen, Tirthapura 2018, Lemmas 10–11) and + an empirical/cost-modeling worked counterexample (Liu et al. 2023, + Fig. 1). The mitigation the literature converges on is *not* "pick the + single best-average curve" but either (a) a small constant number of + simultaneous orderings with a joint worst-case guarantee + (Locality-Sensitive Orderings, O(1/ε^d) orderings) or (b) a + learned/instance-fit parameterized curve family (LMSFC). **Classification: + KNOWN** (each half is separately established) but the *convergence* of + two independent lines of work on the same impossibility, five years + apart, with different tools, is itself a mild corroborating surprise — + this is not a fragile or contested result. + +--- + +## VERDICT + +The literature does not support treating any single Hilbert- or +Morton-family curve as a universally safe default: it contains a proven +unbounded clustering gap for near-full-extent queries (Onion Curve), a +proven dimension-scaling fragmentation disadvantage for the entire +cube-recursive curve family at d≥3 (Haverkort), and a proven impossibility +of any single curve serving heterogeneous query shapes near-optimally +(Onion Curve + independently, cost-modeling literature) — all of which are +falsifiable, pre-registered, and directly testable against the actual query +mix, dimensionality, and size-scaling regime this architecture will +operate under before a Morton-cascade default is locked in. diff --git a/docs/lotus/rp-seal-v1/C1.md b/docs/lotus/rp-seal-v1/C1.md new file mode 100644 index 00000000..c9a305f3 --- /dev/null +++ b/docs/lotus/rp-seal-v1/C1.md @@ -0,0 +1,809 @@ +# C1 — Erasure Coding / Seals (Domain C, Builder) + +**Role**: design (not implement) reference erasure-coding constructions over the +baseline geometry, evaluate the seven named candidates, and inventory existing +project mechanisms with file:line precision. + +**Independence note**: this report was produced without reading any other +researcher's output, any `docs/lotus/**`, any `.claude/board/**`, or +`lance-graph/.claude/plans/**`. All code claims below are grounded in direct +reads of committed source in `/home/user/ndarray`, `/home/user/lance-graph`, +`/home/user/OGAR`, and the two pinned Lance columns at `/tmp/sources/lance-9` +(tag v9.0.0) and `/tmp/sources/lance-main` (shallow clone, 2026-08-18). All +literature claims are grounded in the fetched paper text or web search results +listed in PRIOR ART, not recalled-and-assumed. + +--- + +## Baseline geometry — restated and cross-checked + +- row = 512 B +- field = 4×4 rows = 16 rows × 512 B = **8192 B = 8 KiB** +- cycle = 4096 fields × 8 KiB = **32 MiB = 65,536 rows** (4096 × 16 = 65,536 ✓) +- 2-D hierarchy: 4 / 16 / 64 / 256 / 1024 / 4096 fields = 4¹ … 4⁶ (a 6-level + 4-ary/quad-tree over the cycle's fields) +- inverse reduction 4096 → 1024 → 256 → 64 → 16 → 4 → 1 + +The one geometric fact every candidate below leans on: **4096 = 64 × 64**. The +cycle's 4096 fields form a natural **64×64 grid**, which is exactly the shape +row/column and product codes want. I treat the **field (8 KiB)** as the coding +symbol/chunk unit throughout (RS/parity operates byte-for-byte *across* +corresponding bytes of same-sized chunks — a "symbol" in the GF(2^w) sense is +one byte within one chunk; the *chunk* is what gets counted as k, m, n). + +--- + +## SOURCE ARCHAEOLOGY + +### Existing integrity/redundancy mechanisms found, with file:line + +| # | Mechanism | Location | What it actually does | +|---|---|---|---| +| 1 | `Seal` / `MerkleRoot` (blake3, 48-bit truncated) | `ndarray/src/hpc/seal.rs:1-59` | `Plane::merkle()` masks bits by `alpha` and blake3-hashes → 6-byte root. `verify()` returns `Wisdom`/`Staunen`. **Pure detection.** No repair, no localization below whole-`Plane` granularity. | +| 2 | `MerkleTree` (8 branches × 64 leaves, blake3) | `ndarray/src/hpc/merkle_tree.rs:1-100+` | Hierarchical hash over named regions (identity/nars/edges/rl/bloom/qualia/adjacency/content) of the 256-word (2 KB) container. `StaunenType` classifies *which branch* changed. **Detection + coarse localization** (branch/leaf granularity), still zero repair. This is the closest existing analogue to a "hierarchical cascade" structure — but it is a **hash tree, not a parity tree**: it can tell you a branch is wrong, never reconstruct it. | +| 3 | `MerkleRoot` (XOR-fold, non-cryptographic) | `lance-graph/crates/lance-graph/src/graph/spo/merkle.rs:24-39` | `h = h.rotate_left(7) ^ w; h = h.wrapping_mul(K)` folded over a `Fingerprint`'s words → one `u64`. Doc-comment explicitly records a **known gap**: `verify_lineage()` does NOT re-hash (test `test_verify_lineage_gap` pins this as *expected* behaviour, lines 228-245). Detection-only, and even detection is optional/bypassable by construction. | +| 4 | Wide checksum, XOR-fold of a 256-word container | `lance-graph/crates/bgz17/src/container.rs:436-452` (`compute_wide_checksum`, `verify_wide_checksum`) | `container[254] = XOR-fold(words[128..253])`. **This is literally a RAID-4/5 "P" parity computation** wearing a checksum's name: an XOR-fold of many words into one word is exactly what a single-parity XOR scheme computes over the same word set — the difference is purely in how it's *used* (compared against itself for detection, vs. XOR'd against survivors for reconstruction). It is currently used only for detection. | +| 5 | `xor_parity` / `recover` — **explicit RAID-5-style XOR parity** | `lance-graph/crates/lance-graph-cognitive/src/container_bs/delta.rs:8-27` | `parity = c[0] ⊕ c[1] ⊕ … ⊕ c[n]`; `recover(survivors, parity) = parity ⊕ (all survivors)`. Module doc: *"The parity container enables single-container recovery (RAID-5 style)."* This is **Candidate F/E-adjacent territory already implemented**: flat, single-parity (m=1), pure GF(2) XOR, no GF(2^8) multiply, protects exactly one erasure per group, and — critically — **has no localization mechanism of its own**: `recover()` must be told externally which index is missing (it takes `survivors: &[&Container]`, not "all containers, some possibly corrupt"). It is an *erasure* code consumer (needs known-location loss), not an *error*-correcting one. | +| 6 | GF(2) XOR-bind semirings (VSA, blasgraph) | `lance-graph/crates/holograph/src/graphblas/semiring.rs:409`; `lance-graph/crates/lance-graph/src/graph/blasgraph/{types.rs:402,semiring.rs:52}`; `lance-graph-contract/src/grammar/role_keys.rs:83,126` | All "GF(2) field arithmetic" hits in the codebase are **GF(2)**, i.e. plain XOR/bind algebra for VSA superposition and graph semirings — **not** GF(2^8)/GF(2^16) erasure-coding arithmetic. Zero hits for `galois`, `gf256`, `reed_solomon`, `erasure` as erasure-coding terms anywhere in `ndarray`, `lance-graph`, or `OGAR` (`grep -rli` across both trees, confirmed empty except the plan doc I am barred from reading and this program's own prompt). **Finding: GF(2^8)/GF(2^16) Reed-Solomon-class arithmetic does not exist anywhere in this codebase today.** Every existing "field arithmetic" or "parity" hit is either GF(2) XOR or a hash. | +| 7 | `"seal"` as a *durability* term, unrelated to erasure coding | `lance-graph/crates/lance-graph-supervisor/src/cycle_driver.rs:114-235` (`SealedTransition`, `SealedCycle`, `SealFailure`, `SealRecovery`, `seal_cycle`); `lance-graph/crates/lance-graph-planner/src/persist_sink.rs` module doc: *"one semantic generation = one globally-aligned 64k-row cycle that seals into exactly one `DatasetVersion`"* | **Naming collision worth flagging explicitly.** In this codebase "seal" already means *"the single WAL write that durably commits a cycle"* — a transactional/commit concept, with its own `SealFailure`/`SealRecovery` taxonomy (Regenerate / ResubmitFrozen / Permanent / Escalate). The research program's own vocabulary ("seal CPU & metadata bytes" in the metric list) is a **different, unrelated sense** of "seal" — an integrity/redundancy stamp attached to a finalized physical unit. Any implementation proposal MUST NOT name a new type `Seal`/`SealedX` inside `lance-graph-supervisor`/`-planner` without a disambiguating prefix (e.g. `RedundancySeal`, `EcSeal`), or it will silently collide with `SealedCycle`'s existing, load-bearing meaning. | + +### GF(2^8) SIMD building blocks that already exist (and don't), by file:line + +The task explicitly asks what `ndarray::simd` would need for GF(2^8)/GF(2^16) +arithmetic "in SIMD terms." I checked this concretely rather than assuming. + +The textbook fast technique for **GF(2^8) multiply-by-a-fixed-constant** is +Plank/Greenan/Miller's split-table method (see PRIOR ART): split each byte +into low/high nibble, do two independent 16-entry table lookups (one per +nibble) via a byte-shuffle instruction, XOR the two results. Concretely: +`b·c = T_lo[b & 0x0F] ^ T_hi[(b >> 4) & 0x0F]` for two tables precomputed once +per constant `c`. This requires exactly four primitive kinds: (1) a per-byte +low-nibble mask, (2) a per-16-bit-lane right-shift + mask for the high +nibble, (3) a 16-entry-table byte shuffle, (4) XOR-combine. + +I checked whether **all four already exist as public `ndarray::simd` API**: + +- `U8x32::shuffle_bytes` (AVX2, `_mm256_shuffle_epi8`) — `ndarray/src/simd_avx2.rs:2649-2651`. **Exists, public.** +- `U8x64::shuffle_bytes` (AVX-512, `_mm512_shuffle_epi8`) — `ndarray/src/simd_avx512.rs:738-740`. **Exists, public.** +- `U8x32::shr_epi16` — `ndarray/src/simd_avx2.rs` (listed in the `impl U8x32` method inventory); `U8x64::shr_epi16` — `ndarray/src/simd_avx512.rs:593`. **Exists, public** — this is exactly the "16-bit-lane shift + AND cleanup" half of nibble extraction (the AND-mask-0x0F cleanup step is separate, see next). +- `BitAnd`/`BitXor` on `U8x32` — `ndarray/src/simd_avx2.rs:2726` (`impl core::ops::BitAnd for U8x32`), `:2746` (`BitXor`). **Exists, public**, via operator overload (`a & mask`, `a ^ b`). +- `BitAnd`/`BitXor` on `U8x64` — `ndarray/src/simd_avx512.rs:827-828` (`impl_bin_op!(U8x64, BitAnd, bitand, _mm512_and_si512)`, same line for `BitXor`). **Exists, public**, macro-generated. + +**Finding: all four primitives needed for a nibble-split GF(2^8) +constant-multiply kernel already exist as PUBLIC `ndarray::simd` API on both +AVX2 (`U8x32`) and AVX-512 (`U8x64`) — composing them costs zero new +intrinsics.** This is not hypothetical: `ndarray/src/simd_avx2.rs:353-393` +(`pub fn popcount`) already implements the *identical* nibble-split → +dual-shuffle → combine idiom for a *different* table (Harley-Seal popcount +LUT) using raw intrinsics internally; `U8x32::nibble_popcount_lut()` exposes +the table-construction half as a named public helper. A GF(2^8) +constant-multiply kernel is the *same shape of code* with a different pair of +16-entry tables and XOR instead of ADD to combine — realistically ~30-50 new +lines per backend, not a new primitive class. + +What is **missing**: + +- **No GF(2^8) table/generator/log-exp infrastructure anywhere** — no + primitive-element table, no `gf256_mul(a, b)` (variable×variable, needed to + *derive* the two nibble tables for an arbitrary Vandermonde/Cauchy + coefficient at encode-configuration time), no inverse table (needed for + decode when solving a small linear system over erased symbols). +- **No fused `gf256_muladd_region(dst: &mut [u8], src: &[u8], coeff: u8)`** — + the actual hot-path primitive an encoder needs (ISA-L calls this shape + `gf_vect_mul`/`ec_encode_data`), i.e. the W1a-consumer-contract-shaped ask: + struct method on a typed wrapper, three backends, parity test, documented + saturation semantics (none needed here — GF(2^8) has no overflow, unlike + the workspace's `saturating_abs` VPABSB lesson, which is a genuinely + different hazard class worth NOT conflating). +- **NEON gap.** `ndarray/src/simd_neon.rs` has **no** `shuffle_bytes` / + `vqtbl1q_u8`-equivalent (checked via grep for `vqtbl|shuffle_bytes|vtbl` — + zero hits). ARMv8 NEON's `vqtbl1q_u8` is the direct analogue of + `_mm256_shuffle_epi8`/`_mm512_shuffle_epi8` and is baseline-NEON (no extra + feature flag), so this is a real, closeable gap, not a hardware limitation + — but it means **today the GF(2^8) SIMD path is x86-only**; a NEON + `shuffle_bytes` primitive is a genuine prerequisite for the "all three + backends" contract, not yet satisfied. +- **`U8x64::Mul` exists (`ndarray/src/simd_avx512.rs:789`) and is a TRAP, not + a shortcut.** It is plain wrapping `u8` multiply (`a*b mod 256`), which is + the WRONG multiplication for GF(2^8) — a naive implementer reaching for + "there's already a `Mul` impl on `U8x64`" would silently build a broken, + non-MDS erasure code. Worth flagging explicitly as a naming hazard + symmetric to the workspace's own VPABSB lesson. + +**GF(2^16) cost, in SIMD terms.** The nibble-split trick does not simply +"double" for 16-bit symbols: a 16-bit value has 4 nibbles, and multiply-by-constant +either needs ~4 nibble-table lookups combined via more XORs (roughly 2-4× +the GF(2^8) SIMD op count per byte of data for the classical +Vandermonde/Cauchy multiply-accumulate construction — Plank et al. report +GF(2^16) SIMD tables that are correspondingly larger and slower per byte, +see PRIOR ART), **or** a structurally different approach: additive-FFT-based +Reed-Solomon over GF(2^16) with a Cantor basis (Lin-Chung-Han 2014; shipped +as Leopard-RS/`reed-solomon-16`), which changes the *complexity class* from +O(n·k) constant-multiplies to O(n log n) butterfly operations. At this +codebase's k=4096 scale (see Candidate A below), the FFT-based construction +is the one real production systems actually use — "GF(2^16) via nibble +tables, scaled up" is the wrong mental model for k this large; "GF(2^16) via +additive FFT" is the right one, and it is architecturally a *different* +kernel shape (needs a Cantor-basis multiplication table + a +butterfly-network traversal, not a per-byte muladd loop), not currently +present in any form in `ndarray::simd`. + +**Self-derived cost estimate** (not from a paper — a direct consequence of +the construction above, stated as such): computing a GF(2^8) **P** (pure XOR, +1 op-family: XOR-accumulate) costs roughly **5-6× fewer** SIMD ops per source +chunk than computing a GF(2^8) **Q** (nibble-split multiply: AND + shift+AND ++ 2×shuffle + XOR = 5 ops, *then* XOR-accumulate into the running sum) for +the same output size. This is the same asymmetry RAID-6 implementations +universally exploit (P is "free" relative to Q) — confirmed structurally, not +just by reputation. + +### Lance 9.0.0 / current-upstream check — no page-level integrity exists to build on + +Checked both pinned columns for any existing checksum/CRC/erasure machinery +in the file format itself: + +- `grep -rn "checksum" rust/lance-file/src` → **zero hits** in either + `/tmp/sources/lance-9` or `/tmp/sources/lance-main`. +- `grep -rln "crc\|Crc"` under `lance-file/src` → **zero hits**, both columns. +- `grep -rli "erasure\|reed.solomon\|reed_solomon"` across all of + `rust/` in both columns → **zero hits**. +- `lance-file/src/format.rs` (33 lines total) defines only `MAGIC = b"LANC"`, + `MAJOR_VERSION`/`MINOR_VERSION`, and re-exports two protobuf modules + (`pb`, `pbfile`) — the actual page/footer layout is protobuf-schema-defined, + not hand-rolled, and carries **no checksum field visible at this level**. +- The only integrity-adjacent hit anywhere in `lance-file`/`lance-core`/ + `lance-encoding` is `lance-core/src/utils/bloomfilter/sbbf.rs` (a + split-block Bloom filter — a *membership* structure for zone maps, not an + integrity check) and two `statistics.rs`/`primitive.rs` references that are + column statistics, not checksums. + +**Finding: Lance 9.0.0 (and current upstream) provides no page-level +checksum, no CRC, and no erasure protection of its own.** Any seal/parity +scheme sits entirely ABOVE Lance, cannot piggy-back on an existing Lance +mechanism, and cannot assume Lance will refuse to hand back silently-corrupted +bytes — Lance's durability story is `DatasetVersion`-level (a version either +committed or it didn't; see `persist_sink.rs`), not byte-level. This also +means the *object store underneath* Lance is the only layer today doing +anything erasure-coding-shaped for this stack (see EXPECTED FAILURE). + +--- + +## MECHANISM + +I evaluate the seven named candidates against the 64×64 field grid (4096 +fields/cycle). All GF(2^8) feasibility claims use Anvin's documented ceiling: +a systematic P+Q construction over GF(2^8) supports **at most 255 data +elements** per group (PRIOR ART). All redundancy percentages are `m/k` on +**physical field bytes** unless stated otherwise — I do not call any of these +"free"; each is stated with its exact `(k, m)`. + +### Existing-mechanism baseline (not a new candidate — what's already shipped) + +- **Plain checksum/hash** (blake3-truncated `Seal`/`MerkleRoot`, + `ndarray/src/hpc/seal.rs`; XOR-fold `MerkleRoot`, + `lance-graph/.../spo/merkle.rs`; XOR-fold wide checksum, + `bgz17/src/container.rs`): **redundancy overhead ≈ 0*** (a 6-48-bit hash + per chunk is negligible relative to an 8 KiB field — 6 bytes / 8192 bytes + = **0.073%**). Detection only. Localization = whatever granularity you hash + at (per-field hashing gives field-precise localization, i.e. 8 KiB). Zero + repair capability — needs a second, independent redundancy mechanism to act + on a detected/localized failure. + +### A — Flat RS(k+2, k) + +A single group of k data fields protected by 2 parity fields (P = XOR, Q = +GF(2^8) Vandermonde/Cauchy syndrome), evaluated at four k values to show the +overhead/locality/feasibility curve rather than picking one number and +calling it representative: + +| k | overhead (2/k) | GF(2^8) feasible? | repair helpers (k−1) | bytes read to fix 1 field | amplification | +|---|---|---|---|---|---| +| 16 | 12.5% | yes (16≤255) | 15 | 120 KiB | 15× | +| 64 | 3.125% | yes | 63 | 504 KiB | 63× | +| 254 | 0.787% | yes (at the ceiling) | 253 | 2,024 KiB (≈1.98 MiB) | 253× | +| 4096 (whole cycle) | 0.0488% | **no** — needs GF(2^16) or FFT | 4095 | 32,760 KiB (≈32 MiB) | **4095×** | + +**The k=4096 row is the important one to not skip past.** Storage overhead is +essentially free (0.05%), and that is *exactly* the trap: to repair a single +lost 8 KiB field you must read almost the entire 32 MiB cycle — a 4095× +read amplification for an 8 KiB payload. This is the textbook RS +repair-bandwidth problem that directly motivated regenerating codes +(Dimakis et al. 2010) and locally repairable codes (Huang et al. 2012, +Gopalan et al. 2012) in the first place — I am not asserting this from +memory, it is the explicit stated motivation in both papers (see PRIOR ART). +"Flat RS(k+2,k)" at cycle scale is a real, well-documented anti-pattern, not +a strawman I am inventing to make hierarchical schemes look good. + +GF cost: P is one XOR-accumulate pass over k chunks; Q is one GF(2^8) +nibble-split-multiply-accumulate pass (≈5-6× the SIMD-op cost of P per +source chunk, see SOURCE ARCHAEOLOGY). At k=4096 the group **cannot** use +classical GF(2^8) Vandermonde/Cauchy Q at all (>255 data elements) — it must +either (a) split into ≥17 sub-groups of ≤254 (which is then not "flat" any +more — it's B/C/D wearing a trenchcoat), or (b) move to GF(2^16) +(2-4× the per-byte SIMD cost, classical construction) or an FFT-based +GF(2^16) code (O(n log n), different kernel shape entirely — Leopard-RS +reports >1.2 GB/s/core on AVX2 for exactly this regime). + +### B — Row P/Q + Column P/Q + +Treat the 4096 fields as a 64×64 grid. Every row of 64 fields gets its own +P+Q (2 parity fields); every column of 64 fields gets its own P+Q. Each of +the 128 lines (64 rows + 64 columns) is an independent RS(66,64) group — +GF(2^8) feasible with huge margin (64 ≪ 255). + +- Parity fields added: 64×2 (row) + 64×2 (col) = **256**. +- Overhead: 256/4096 = **6.25%**. +- Repair locality for one lost field: **63 helpers via either axis** (pick + whichever line is healthier) — 504 KiB read, **63× amplification**, a + 64× improvement over flat-4096's 4095× for a storage cost that is still + two orders of magnitude below "protect everything with its own dedicated + group" would otherwise imply. +- Fault tolerance: any single line's own P+Q corrects up to 2 erasures + *within that line* algebraically (no iterative/propagation decode needed); + because every field sits on two independent lines (its row and its + column), the system additionally survives many multi-failure patterns + flat-64 partitioning could not (two losses in *different* flat-64 groups + are each individually fine under flat partitioning too, but two losses in + the *same* flat-64 group that also happen to share neither row nor column + under B's grid can be resolved via whichever axis is unaffected — the + redundant covering is the actual benefit over flat partitioning at the + same k=64, not raw overhead, which is identical: flat-64's 2/64=3.125% + *per axis* matches B's row-only or column-only cost exactly; B pays 2× + that (6.25%) for the redundant orthogonal covering). + +### C — Product code (classic 2-D XOR, single parity per axis) + +Same 64×64 grid, but only ONE parity per axis (pure XOR row-parity + pure +XOR column-parity), plus one corner cell (parity-of-parities): + +- Parity fields added: 64 (row XOR) + 64 (col XOR) + 1 (corner) = **129**. +- Overhead: 129/4096 = **3.15%** — essentially half of B, because each axis + carries 1 syndrome (P only) instead of 2 (P+Q). +- Repair locality for the *common* single-loss case: identical to B — 63 + helpers via either axis, 504 KiB, 63× amplification. +- Fault tolerance: **structurally weaker than B in a specific, well-known, + quantifiable way** — see EXPECTED FAILURE for the "four-corner rectangle" + counterexample. Two erasures within one row can be recovered easily only + if at most one of them lacks column-side help; the classic degenerate + pattern (2 rows × 2 columns, exactly one erasure at each of the 4 + row/column intersections) leaves every affected row AND every affected + column with exactly 2 unknowns and only 1 equation (its single XOR + parity) — algebraically under-determined regardless of how many + additional reads are attempted. This is not a bandwidth problem to solve + by reading more; it is a rank-deficiency the code's minimum distance does + not cover. +- GF cost: **cheapest of B/C/D at the field-count level** — pure XOR, zero GF + multiply anywhere, so the Q-vs-P 5-6× SIMD cost differential (SOURCE + ARCHAEOLOGY) never applies. This is a genuine, real, computable reason C + might be preferred over B for pure single-erasure protection despite its + documented worst-case gap. + +### D — Local P/Q + higher-level cascade parity (following the native 4-ary reduction) + +This is where the baseline geometry's own 6-level 4-ary hierarchy +(4096→1024→256→64→16→4→1) becomes directly load-bearing, and where the +Gopalan-Huang-Simitci-Yekhanin locality bound (PRIOR ART) gives an exact, +checkable *lower* bound on overhead as a function of which cascade tier you +attach local parity to. The bound: for a systematic code with information +locality r (repair reads r other symbols) over k data symbols at minimum +distance d, `n − k ≥ ⌈k/r⌉ + d − 2` (Gopalan et al. 2012, Theorem 5), attained +with equality by "canonical"/Pyramid-style constructions. + +Applying it at every cascade tier (r = the number of *sibling leaf fields* +a repair at that tier would read; d=2 = single-erasure-only floor, d=3 = the +correction for "2 erasures within one local group" using local P+Q instead +of local P-only): + +| attach local parity at | sibling group size (r) | Gopalan floor at d=2 | Gopalan floor at d=3 | overhead (d=2) | overhead (d=3) | repair bytes (r×8 KiB) | amplification | +|---|---|---|---|---|---|---|---| +| level 1 (groups of 4) | r=3 | ⌈4096/3⌉=1366 | 1367 | **33.4%** | 33.4% | 24 KiB | **3×** | +| level 2 (groups of 16) | r=15 | ⌈4096/15⌉=274 | 275 | **6.7%** | 6.7% | 120 KiB | 15× | +| level 3 (groups of 64) | r=63 | ⌈4096/63⌉=66 | 67 | **1.61%** | 1.64% | 504 KiB | 63× | + +**The single most important number for this candidate: attaching local +parity at the native level-1 tier (group-of-4, matching the codebase's own +existing 4-ary spatial/semantic cascade exactly) costs a MINIMUM of 33.4% +storage overhead by the Gopalan bound alone — before any global/cascade +tier is added on top.** This is not a design flaw of a particular +construction; it is the information-theoretic floor for buying repair +locality of 3 at this k. It is 5-10× the overhead of candidates B/C for a +repair amplification that is 20-21× cheaper (3× vs 63×). This is the +canonical locality/overhead trade the LRC literature exists to quantify +(Gopalan et al.; corroborated empirically by Azure LRC(12,2,2)'s real +16-data-fragment deployment landing at almost exactly the same ~33% +overhead figure at a *different* r=6 — see PRIOR ART — which is a useful +real-world sanity check on the bound's practical tightness, not a +coincidence to lean on as if it were the same number). + +**Convergence with C at the coarse end.** At r=63 (level-3 grouping), D's +Gopalan-optimal single-parity-per-group construction is essentially +**identical in shape to one axis of candidate C** (a flat XOR-per-group-of-64 +covering, no redundant orthogonal covering) — 66 parity fields vs C's 64+1 +row-XOR fields, both ≈1.6% overhead, both 63-helper repair. D at its coarsest +practical tier and "half of C" (row-parity only, no column-parity) are the +same construction wearing different names. This is a genuine, checkable +cross-domain finding — see SURPRISES. + +**Cascading to the root.** Rolling local-tier parity chunks up through the +remaining levels (1024→256→64→16→4→1) adds a *bounded, rapidly-shrinking* +additional cost: each higher tier operates on the PRIOR tier's parity count, +shrunk by ≈4× per level, so the total across all higher tiers is bounded by +the geometric series `tier1 × (1 + 1/4 + 1/16 + …) ≤ tier1 × 4/3`. At the +r=3 (level-1) setting this puts a full-cascade-to-root estimate at roughly +33.4% × 4/3 ≈ **44.5% overhead upper bound** — a real, non-trivial +additional ~11 points beyond the local-tier-only floor, but not a doubling. +(Stated explicitly as an order-of-magnitude geometric-series estimate, not a +precise derivation of a specific decode algorithm's real redundancy — the +EXPERIMENT section below turns this into a measured number.) + +GF cost: identical per-group asymmetry to A/B (P cheap, Q ≈5-6× P) at +whichever tier(s) use P+Q rather than pure-P; small groups (r=3) make the +per-group GF-table setup overhead (building the nibble tables once per +constant) a larger fraction of total work than at large-k flat RS, since +there are many more, smaller groups — worth measuring, not assuming, in the +EXPERIMENT (metric: "seal CPU & metadata bytes" should show D's per-cycle +constant-count scaling as O(k) small groups vs A's O(1) large group). + +### E — Hash-only syndromes + +Exactly what mechanisms #1-4 in SOURCE ARCHAEOLOGY already are: a +strong/cryptographic hash (blake3) or a cheap XOR-fold, per chunk, with zero +redundant data. Overhead ≈0.07% (48-bit truncated blake3 / 8 KiB field). +**Detection + localization, zero repair.** Not a standalone candidate for a +"seal that survives loss" — it is the detection half every other candidate +needs paired with it (see G), and it is already shipped in three different +forms in this codebase (see SOURCE ARCHAEOLOGY #1, #3, #4). + +### F — Payload parity + +Reading this candidate as a **scope** decision, not a coding-algebra +decision: protect only the variable/bulk *value payload* bytes of a chunk, +leaving small fixed-width keys/metadata (the 16-byte GUID key, the 16-byte +edge block per the V3 canon) either unprotected (already durable via +Lance's own version-level commit) or under plain replication (16 bytes is +cheap to replicate 3× outright — 48 bytes total — which is not worth +erasure-coding at all; erasure coding's win only materializes at chunk sizes +where `m/k` beats `(replicas−1)×100%`). **F composes with A-D**; it is not a +fifth algebra. `container_bs/delta.rs`'s existing `xor_parity`/`recover` +(SOURCE ARCHAEOLOGY #5) is *already* an F-shaped choice in practice: it +operates on whole `Container` objects (which mix key-shaped and +payload-shaped words in the same 2 KB blob per `bgz17/src/container.rs`'s +own word-layout comments), so the existing code has not actually drawn the +F distinction — it protects the WHOLE container, keys included. A genuine F +implementation would need to first define which byte ranges of a field are +"payload" vs "key/metadata," which for a V3-canon row means resolving +`ClassView::edge_codec_flavor` per class before deciding what to protect — +non-trivial, class-dependent, and outside this cell's scope to design further. + +### G — Hybrids: strong per-chunk hash + erasure parity + +This is the practically-recommended combination, and it follows directly +from a precise, textbook coding-theory fact worth stating exactly rather +than hand-waved: **to correct t erasures at KNOWN locations costs t parity +symbols; to correct t errors at UNKNOWN locations costs 2t parity symbols** +(classical result underlying Berlekamp-Massey/Peterson-Gorenstein-Zierler +decoding — every RS/BCH text states this; I am not citing a single paper for +it because it is foundational coding theory, not a specific contribution). + +Consequence: relying on an erasure code's OWN syndrome to both *detect* and +*localize* silent corruption (candidates A-D used as self-correcting codes, +with no separate hash) needs **double** the parity of the same code used +purely as an *erasure* repair mechanism fed known-bad locations from an +external source. A cheap per-chunk hash (mechanism #1/#3/#4, ~0.07% +overhead) turns every candidate above from "error-correcting" mode into +"erasure" mode for free, halving its effective parity requirement for the +same fault coverage — or equivalently, doubling its fault coverage for the +same parity budget. + +**G is not a new algebra — it is "add mechanism #1 (already shipped, ~0.07% +overhead) in front of whichever of A/B/C/D you pick."** Concretely: hash +every field at seal time (reusing `Plane::merkle`'s pattern or +`compute_wide_checksum`'s XOR-fold, whichever CPU/collision-resistance +trade-off the write path can afford); at read/repair time, per-field hash +mismatch tells you EXACTLY which field(s) are bad (localization, essentially +free); feed those as known erasures into whichever redundancy scheme (B/C/D) +is configured. This is also the industry-standard pattern (ZFS checksums + +separate RAID-Z parity; Ceph's per-object CRC + separate erasure-coded +pool) — not something invented here. + +--- + +## EXPECTED BENEFIT + +- **B and C both buy a ~65× reduction in repair read-amplification (4095×→63×) + relative to flat whole-cycle RS, for single-digit-percent storage + overhead (6.25% / 3.15%).** This is the highest-leverage, lowest-risk move + available and requires no departure from classical, decades-proven RAID-6 + construction — genuinely low technical risk. +- **D buys a further ~20× reduction in repair amplification over B/C + (63×→3×) but at 5-10× the storage cost (33-45% vs 3-6%)** — a real, + quantifiable trade that the Gopalan bound makes falsifiable: any proposed + D-shaped construction that claims BOTH r=3 locality AND <20% overhead at + k=4096 is provably impossible and should be rejected on inspection, not + benchmarked. +- **All four algebra candidates (A-D) are buildable from EXISTING + `ndarray::simd` public API on x86 with zero new intrinsics** — the + `shuffle_bytes`/`shr_epi16`/`BitAnd`/`BitXor` primitives needed for GF(2^8) + are already public on both `U8x32` and `U8x64`, and the exact nibble-split + idiom is already implemented (for a different table) in + `ndarray/src/simd_avx2.rs:353-393`. This substantially lowers the + implementation-risk estimate for any of A-D relative to "we'd need to add + raw intrinsics," which was the a priori worst case. +- **G (hash+parity) is a strict improvement over any bare A-D at equal + parity budget** — it is provably not worse (adds a fixed, tiny ~0.07% + overhead) and is provably better whenever silent (unknown-location) + corruption is a real threat, per the erasures-cost-t/errors-cost-2t rule. + There is no honest argument for shipping A-D bare when three + hash-mechanism implementations already exist in-tree to compose with. +- **The 4-ary cascade structure this codebase already uses for spatial/HHTL + purposes (per the OGAR/lance-graph CLAUDE.md context and + `merkle_tree.rs`'s branch/leaf hierarchy) is genuinely reusable as the + ORGANIZATIONAL skeleton for D** — the group boundaries, not the algebra, + transfer directly (4096→1024→256→…). Whether the *algebra itself* should + reuse anything from the hash-tree code is a separate, smaller question + (see SURPRISES: it should NOT reuse hash-tree code for the parity + computation itself, only the grouping topology). + +--- + +## EXPECTED FAILURE + +1. **The premise itself may be partially redundant with the object store.** + Lance has zero page-level integrity of its own (confirmed in SOURCE + ARCHAEOLOGY), which means Lance's *files* are opaque blobs to whatever + object store holds them — and modern object stores (S3, GCS, Azure Blob) + already erasure-code data at rest for exactly the "protect against + media/hardware failure" threat model, at "11 nines" durability SLAs. + Azure Blob Storage's erasure coding IS Huang et al.'s LRC (PRIOR ART) — + the exact algebra candidate D is reaching for is *already running one + layer down* on that deployment target. **If this system's actual + deployment target is an object store with its own erasure coding, an + application-level scheme protects only against a narrower threat class: + in-process bit-flips before the write, application-logic corruption that + writes wrong-but-well-formed bytes, or sub-object repair-locality wins + the object store's own API doesn't expose (partial-object self-heal + without a full-object GET).** This is a real, falsifiable pre-condition + the EXPERIMENT below must check FIRST, not an argument against building + anything — but a design that doesn't name which of these three narrower + threats it addresses is solving a problem the deployment target may + already have solved, and claiming credit for durability the object store + was already providing. +2. **Candidate C's four-corner rectangle failure is not a corner case to + shrug off** — for a 64×64 grid it requires exactly 2 rows and 2 columns + to each contribute exactly one of the 4 losses, which under an + independent-per-field failure model (not correlated to physical + proximity) has non-trivial probability at realistic multi-failure rates + and grows with grid size (more row/column pairs to choose from at larger + k). This must be measured against a real failure-rate model, not assumed + negligible. +3. **A at k=4096 is presented in the candidate list as if it were a + reasonable baseline to benchmark against; it is not, at production + scale, without first splitting into ≤255-sized sub-groups — at which + point it stops being "flat" and the comparison against B/C/D becomes + apples-to-oranges unless every candidate is measured at the SAME + effective sub-group size.** The EXPERIMENT design below pins this by + testing flat RS at k∈{16,64,254} — matching B/C/D's own group + sizes — rather than only at k=4096. +4. **D's cost estimate assumes an "ideal" (Gopalan-optimal / Pyramid-code + canonical) construction attains the bound with equality.** A concrete + implementation (e.g. a naive recursive-XOR-of-XORs following the + 4-ary tree literally) may NOT attain the bound — the bound is a + *lower* bound on ANY valid (r,d)-code, not a recipe. A real + implementation could easily land at 40-50% overhead for r=3 rather than + the theoretical 33.4% floor if it does not use the Pyramid-code + "canonical" support-graph structure (disjoint hyperedges of size + exactly r+1, per Gopalan et al. Theorem 7/9). This gap between "the + bound says X is possible" and "our actual construction achieves X" is + exactly what the EXPERIMENT must measure, not assume closed. +5. **GF(2^16) at k=4096 is a genuinely different kernel (FFT/Cantor-basis), + not a bigger version of the GF(2^8) nibble-table kernel** — if candidate + A is pursued at whole-cycle scale, budgeting engineering time as "GF(2^8) + work but wider" will be wrong; it is closer to "port an FFT library," + a materially larger scope than the P+Q constructions in B/C/D. +6. **The Xorbas/Azure numbers (14% more storage, halved repair I/O; ~33% + overhead at r=6) are NOT this system's numbers** — they were measured at + different (k,m,r) on different hardware for different workloads (HDFS + block sizes in the 10s-100s of MB, not 8 KiB fields). Citing them as if + they transfer directly to this system's 8 KiB/32 MiB geometry would be + exactly the kind of unearned transfer this research program exists to + catch. I have deliberately re-derived this system's own numbers from + the Gopalan bound at its OWN geometry rather than reusing Azure's/ + Xorbas's percentages, and flagged where the two happen to land close + (level-1 D ≈ Azure's ~33%) as a coincidence worth noting, not a proof. + +--- + +## EXPERIMENT + +Every design below is pre-registered: metric, baseline, control, kill +condition, stated before any measurement. None of these have been run — this +is a design document, per the BUILDER posture. + +### EXP-C1-1: Storage overhead vs. repair amplification curve (A/B/C/D) + +- **Metric**: logical bytes vs. physical bytes (storage overhead, `m/k`); + repair read bytes for a single simulated field loss (helper chunks × + 8 KiB); write/seal CPU (wall-clock, per-cycle, single core, both P-only + and P+Q paths measured separately per the 5-6× asymmetry claim above). +- **Baseline**: candidate E alone (hash-only, ~0.07% overhead, infinite + repair amplification — cannot repair — used as the "detection floor"). +- **Control**: flat RS at k∈{16,64,254} (NOT just k=4096) run alongside + B/C/D so all four candidates are compared at matched group sizes, per + EXPECTED FAILURE #3. +- **Kill condition**: if candidate D's measured overhead at r=3 comes in + BELOW the Gopalan floor (33.4%) for a construction claiming distance + d≥2, the implementation is provably violating the bound and has a bug + (undercounting parity, wrong r, or the code is not actually + systematic/(r,d) in the paper's technical sense) — this is a hard, + mechanical, falsifiable check, not a judgment call. If D's measured + overhead exceeds the geometric-series estimate (~44.5%) by a wide margin + (say >60%), the cascade-to-root implementation is carrying dead weight + worth investigating before shipping. + +### EXP-C1-2: Object-store redundancy overlap check (gates the whole program) + +- **Metric**: does the deployment target (name it explicitly — S3, GCS, + Azure Blob, local disk, other) already erasure-code or replicate at the + object/file granularity Lance writes at? If yes, what durability SLA does + it publish, and does it expose partial-object repair (byte-range GET + cheaper than full-object GET) that an app-level scheme could exploit + instead of building its own repair-locality machinery from scratch? +- **Baseline**: "no app-level erasure coding, rely entirely on the object + store" — measured/estimated durability and repair-locality of THIS + baseline first. +- **Control**: the same measurement with app-level G (hash + B, the + cheapest credible option) layered on top. +- **Kill condition**: if the object store already provides equal-or-better + durability AND equal-or-better repair locality than the app-level scheme + under test, for the SAME threat model (media/hardware failure only), the + program's justification narrows to the three specific narrower threats + named in EXPECTED FAILURE #1 (pre-object-store bit-flips, app-logic + corruption, sub-object partial-repair) — and the experiment plan must be + re-scoped to demonstrate value against THOSE threats specifically, not + against "durability" generically. This is the single highest-value + negative result this cell could produce if it comes back positive + (object store already sufficient) — it would redirect the entire + program's justification, not just this cell's candidate ranking. + +### EXP-C1-3: Rectangle-pattern vulnerability of C, measured not assumed + +- **Metric**: for the 64×64 grid, enumerate (or Monte-Carlo sample at + realistic scale) the fraction of random k-field failure sets (k=2,3,4,…) + that hit an unresolvable pattern under C's iterative XOR-peeling decode + vs. the fraction B resolves via its per-line P+Q. +- **Baseline**: B (same repair locality, no known unresolvable pattern at + d=3 per line). +- **Control**: C with an added corner-diagonal parity (a cheap + EVENODD/RDP-style third syndrome, still pure XOR, no GF multiply) to see + whether it closes the rectangle gap at lower cost than moving fully to B's + P+Q-per-axis. +- **Kill condition**: if C's unresolvable-pattern rate at realistic + multi-failure rates (whatever the deployment's actual measured field-loss + rate turns out to be, from EXP-C1-2's object-store data) is below, say, + 10⁻⁶ per cycle, the "half the overhead of B" argument for C wins outright + and the rectangle problem is academically real but operationally + irrelevant at this system's scale — a legitimate, falsifiable way for C + to win despite EXPECTED FAILURE #2. + +### EXP-C1-4: GF(2^8) SIMD kernel — build vs. measure against Plank's reported speedup + +- **Metric**: wall-clock throughput (bytes/sec) of a `gf256_muladd_region` + built purely from the four existing public `ndarray::simd` primitives + identified in SOURCE ARCHAEOLOGY, on AVX2 and AVX-512, vs. a + naive scalar table-lookup baseline. +- **Baseline**: scalar `for byte in src { dst[i] ^= GF_MUL_TABLE[byte as usize][coeff as usize] }`. +- **Control**: same kernel measured on the Q (GF multiply) path vs. the P + (pure XOR) path, to confirm or refute the self-derived 5-6× ratio. +- **Kill condition**: if measured SIMD speedup over scalar falls + substantially short of Plank et al.'s reported 2.7×-12× range (say, + <2×), something about ndarray's existing `shuffle_bytes`/`shr_epi16` + composition is not achieving the intended instruction selection (e.g. + LLVM not folding the nibble-mask-shift sequence into the efficient form) + and needs investigation before any candidate ships — this would be a + concrete, actionable NEGATIVE result about the *feasibility* claim in + MECHANISM/EXPECTED BENEFIT, not just a performance nit. + +### EXP-C1-5: NEON parity check (closes or confirms the cross-backend gap) + +- **Metric**: does a `U8x16`/`U8x32`-equivalent `shuffle_bytes` exist or + get added to `simd_neon.rs` via `vqtbl1q_u8`, and does the SAME + nibble-split GF(2^8) kernel reproduce bit-identical output to the AVX2/ + AVX-512 path on the same input (byte-for-byte parity test, matching this + workspace's own established `E-OCR-*`-style parity-proof discipline seen + in the tesseract-rs precedent). +- **Baseline**: scalar GF(2^8) multiply (portable, always correct, + reference oracle). +- **Kill condition**: any backend disagreement on ANY input byte pair + is an immediate correctness bug, not a performance question — GF(2^8) + arithmetic has no rounding/saturation ambiguity, so parity must be exact, + not "close enough." + +--- + +## PRIOR ART + +Searches actually run (tool + query), not just papers named from memory: + +1. `WebSearch`: "Plank Screaming Fast Galois Field Arithmetic SIMD + Instructions FAST 2013" → confirmed Plank, Greenan, Miller, *Screaming + Fast Galois Field Arithmetic Using Intel SIMD Instructions*, USENIX FAST + 2013 — GF(2^w) for w∈{4,8,16,32} via SIMD, 2.7×-12× speedup over + table/log-based scalar. This is the canonical source for the nibble-split + PSHUFB technique I mapped onto `ndarray::simd::shuffle_bytes` above. + [Screaming Fast Galois Field Arithmetic (USENIX)](https://www.usenix.org/conference/fast13/technical-sessions/presentation/plank_james_simd) +2. `WebSearch`: "Azure Storage Local Reconstruction Codes Huang Simitci 2012 + erasure coding" → Huang, Simitci, Xu, Ogus, Calder, Gopalan, Li, Yekhanin, + *Erasure Coding in Windows Azure Storage*, USENIX ATC 2012 (Best Paper). + LRC(12,2,2): 12 data + 2 local + 2 global, 1.33× overhead, single-fragment + reconstruction from 6 fragments (its own local group) instead of a full + RS(12,4) stripe. [Erasure Coding in Windows Azure Storage (USENIX)](https://www.usenix.org/conference/atc12/technical-sessions/presentation/huang) +3. `WebSearch` + `mcp__alphaXiv__get_paper_content` on arXiv:1106.3625 → + Gopalan, Huang, Simitci, Yekhanin, *On the Locality of Codeword Symbols* + (2011/2012), full text retrieved. Theorem 5: `n − k ≥ ⌈k/r⌉ + d − 2` for + any [n,k,d]-code with information locality r; Theorem 9's "canonical + code" structure theorem (disjoint (r+1)-regular hyperedges) is what a + real D implementation needs to attain the bound with equality, per + EXPECTED FAILURE #4. This is the single most load-bearing citation in + this report — it is what turns Candidate D's cost from "a design choice" + into "a provable floor," used directly in the MECHANISM table above. + [arXiv:1106.3625](https://arxiv.org/abs/1106.3625) +4. `WebSearch`: "Xorbas HDFS-RAID Sathiamoorthy 2013 locally repairable + codes Facebook VLDB repair bandwidth reduction" → Sathiamoorthy, + Asteris, Papailiopoulos, Dimakis et al., *XORing Elephants: Novel + Erasure Codes for Big Data*, VLDB 2013 — LRC construction, 14% more + storage than RS at their (k,m), disk I/O and network traffic halved, + deployed on a real 3000-node/45 PB Facebook cluster. Used above as an + explicitly-flagged non-transferable real-world data point (EXPECTED + FAILURE #6), not as this system's own number. [XORing Elephants (VLDB PDF)](https://www.vldb.org/pvldb/vol6/p325-sathiamoorthy.pdf) +5. `WebSearch`: "RDP row diagonal parity Corbett NetApp EVENODD Blaum two + disk failure product code storage" → Corbett et al., *Row-Diagonal + Parity for Double Disk Failure Correction*, FAST 2004 (Best Paper); + confirmed as EVENODD's closest relative, pure-XOR (no GF multiply at + all) double-failure construction — cited in MECHANISM/Candidate C as a + cheaper-than-P+Q alternative worth considering if C's rectangle gap + needs closing without moving fully to GF(2^8) Q. [Row-Diagonal Parity (USENIX PDF)](https://www.usenix.org/legacy/events/fast04/tech/corbett/corbett.pdf) +6. `WebSearch`: "Anvin mathematics of RAID-6 GF(2^8) 255 disks limit + Reed-Solomon P Q parity" → H. Peter Anvin, *The Mathematics of RAID-6* — + confirmed the exact GF(2^8) capacity ceiling (255 data disks max for a + systematic P+Q construction) used throughout MECHANISM/Candidate A to + determine GF(2^8) feasibility at each tested k. [The mathematics of RAID-6](https://alamos.math.arizona.edu/RTG16/ECC/raid6.pdf) +7. `WebSearch`: "Leopard-RS FFT Reed-Solomon GF(2^16) Cantor basis fast + erasure coding large k" → catid/leopard (Leopard-RS), O(N log N) MDS + RS via additive FFT with Cantor basis {2}, >1.2 GB/s/core AVX2, up to + 65536 total shards in GF(2^16); rooted in Lin-Chung-Han, *Novel + Polynomial Basis With Fast Fourier Transform and Its Application to + Reed-Solomon Erasure Codes*. Used in MECHANISM/Candidate A and the + GF(2^16) cost subsection to establish that whole-cycle-scale RS in + practice means FFT-based codes, not scaled-up nibble tables. + [github.com/catid/leopard](https://github.com/catid/leopard) +8. `WebSearch`: "Hellerstein 1994 Coding Techniques for Handling Failures + in Large Disk Arrays two-dimensional parity rectangle uncorrectable + pattern" → confirmed Hellerstein, Gibson, Karp, Katz, Patterson, + *Coding Techniques for Handling Failures in Large Disk Arrays*, + Algorithmica 12(2/3):182-208, 1994 — the foundational paper analyzing + exactly this class of parity-check-matrix structures for disk arrays. + The specific "four-corner rectangle" counterexample presented in + EXPECTED FAILURE #2 is derived here directly from the linear-algebra + structure of the 2-D XOR construction (an under-determined system: 2 + unknowns, 1 equation, per affected row/column) rather than lifted + verbatim from the paper's text, which the search results did not + surface in enough detail to quote precisely — flagged as + self-derived-but-standard, not mis-attributed. + [Coding Techniques for Handling Failures (Algorithmica)](https://link.springer.com/article/10.1007/BF01185210) +9. In-repo `grep`/`WebSearch`-free source reads: exhaustive `grep -rli` + sweeps across `/home/user/ndarray`, `/home/user/lance-graph`, + `/home/user/OGAR` for `galois|gf256|gf(2|reed.solomon|reed_solomon| + erasure|checksum|crc32|xxhash|parity` (see SOURCE ARCHAEOLOGY for full + hit list and exact file:line), plus targeted `grep`/`wc -l`/file reads + of `/tmp/sources/lance-9` and `/tmp/sources/lance-main` for any existing + page-level integrity mechanism (`checksum`, `crc`, `erasure`). + +**What I explicitly did NOT find prior art for**, after the above searches +(i.e. the honest "documented search came up empty" half of the +NOVEL-CANDIDATE bar): a **recursive, multi-level** (beyond the flat +local+global two-tier shape of Pyramid codes / Azure LRC) locally-repairable +code where the tier boundaries are inherited from an *already-existing, +unrelated* spatial/semantic quad-tree in the same system (as opposed to a +tier structure designed purpose-built for the erasure code itself). Pyramid +codes (Huang, Chen, Li, NCA 2007, cited as reference [5] inside the Gopalan +paper) and Azure LRC are both explicitly two-tier (local groups + one flat +global layer); I found no paper describing an N-tier recursive LRC that +*reuses* an existing application-level 4-ary cascade for reasons other than +the erasure code's own overhead/locality optimization. See SURPRISES. + +--- + +## SURPRISES + +1. **[TRANSFER]** The exact SIMD idiom needed for GF(2^8) constant-multiply + (nibble-split + dual PSHUFB-family shuffle + XOR) is not hypothetical — + it is the SAME idiom already shipped in `ndarray/src/simd_avx2.rs:353-393` + for Harley-Seal popcount, just with a different pair of 16-entry tables + and ADD instead of XOR to combine. All four building-block primitives + (`shuffle_bytes`, `shr_epi16`, `BitAnd`, `BitXor`) are already public on + both `U8x32` (AVX2) and `U8x64` (AVX-512). This is a genuine, checkable, + low-risk transfer from a known technique (Plank et al. 2013) into a + pre-existing code shape in this exact codebase, at a different layer + (popcount → erasure-coding arithmetic) than either was originally built + for — textbook TRANSFER, not novel, but concretely load-bearing for the + EXPECTED BENEFIT claim about implementation risk. +2. **[TRANSFER]** `lance-graph-cognitive/src/container_bs/delta.rs`'s + existing `xor_parity`/`recover` functions are, structurally, an + already-implemented instance of RAID-5 (candidate F/E-adjacent) — the + module doc-comment even says "RAID-5 style" outright. This means part of + candidate F/the-XOR-half-of-B/C is not a design proposal at all; it is + an audit of existing code against the taxonomy this research program + uses, revealing the existing implementation has NOT drawn the + payload-vs-metadata scope distinction candidate F implies (it protects + whole `Container` blobs, keys included) and has NO localization + mechanism of its own (`recover()` must be told which index is missing + externally) — i.e. it is Candidate F glued to nothing (no candidate G + hash pairing exists for it today). +3. **[NOVEL-CANDIDATE, documented search came up empty]** A recursive + N-tier locally-repairable code whose tier boundaries are inherited from + an ALREADY-EXISTING, purpose-unrelated spatial/semantic quad-tree + cascade (this codebase's own 4-ary HHTL/Base17/palette-centroid + structure, evidenced structurally by `merkle_tree.rs`'s 8-branch/64-leaf + hash hierarchy and referenced repeatedly in the OGAR/lance-graph + CLAUDE.md context as a first-class organizing principle) is, as far as + the searches in PRIOR ART could establish, NOT a documented technique — + Pyramid codes and Azure LRC are both purpose-built two-tier + constructions, not N-tier reuses of an unrelated tree. Whether the + REUSE (sharing tier boundaries with an existing structure, for the sake + of one shared mental model / one shared set of group-membership + lookups, rather than for any coding-theoretic reason) has any actual + coding-theoretic benefit over a purpose-built tier count is an open, + testable question this cell does not resolve — flagged NOVEL-CANDIDATE + specifically for the *reuse-of-an-unrelated-existing-tree* framing, not + for "hierarchical/nested LRC" in general (which is well-trodden — see + Pyramid codes). +4. **[KNOWN, but worth stating precisely since the task's own candidate list + conflates two axes]** Candidates A-D are coding-ALGEBRA choices; + candidates E and F are SCOPE/composition choices (what to protect, and + whether to pair detection with repair); candidate G names the + composition of E with any of A-D explicitly. This is not a + surprising fact in coding theory, but stating it precisely resolves what + would otherwise read as "seven independent candidates" into "four real + algebra choices, one detection primitive, one scope decision, and one + correct recommendation to combine them" — worth flagging so the + consolidation pass doesn't score E/F/G as competitors to A-D on the + same axis. +5. **[DISPROVED]** My working assumption entering this cell was that + "flat RS(k+2,k)" at k=4096 would be a reasonable strawman baseline to + benchmark against real numbers — i.e. that its main defect would show + up as merely "worse repair locality" in a benchmark. Checking Anvin's + documented GF(2^8) ceiling (255 data elements) disproved the premise + that it is even a well-formed classical construction at that k: it is + not just slower to repair, it **cannot be built at all** in GF(2^8) + without first silently becoming a multi-group scheme (i.e. secretly + becoming B/C/D) or moving to a structurally different GF(2^16)/FFT + kernel. The EXPERIMENT design was revised (EXP-C1-1) specifically to + test flat RS at matched sub-group sizes rather than only at cycle + scale, precisely because the original "just benchmark A at k=4096" + framing was itself disproved as a coherent baseline. + +--- + +## VERDICT + +No single candidate wins outright; the geometry makes the trade-offs +unusually clean to state as a decision rule rather than a single pick. +**G (hash + B) is the lowest-risk, highest-confidence recommendation** for a +first implementation: pair the already-shipped, near-zero-cost detection +mechanisms (blake3/XOR-fold hashing, three variants already in-tree) with +row/column P+Q over the native 64×64 field grid — 6.25% storage overhead, +63× repair amplification (down from flat-RS's 4095×), buildable today from +public `ndarray::simd` primitives on x86 with a real, closeable, well-scoped +NEON gap, and zero exposure to candidate C's rectangle-pattern hole. **D is +the correct escalation path if measured repair-locality requirements +demand better than 63×** (e.g. if EXP-C1-2 shows the object store's own +partial-repair story is worse than assumed), but its cost is not a design +nicety to negotiate down — 33%+ overhead at r=3 is an information-theoretic +floor (Gopalan et al.), not an implementation target to optimize past. +**The single highest-value open question for the consolidation pass is +EXP-C1-2**: whether the deployment's object store already provides +equivalent-or-better durability for the threat model this whole domain is +implicitly assuming, which would redirect the entire erasure-coding +program's justification toward the three narrower threats (pre-store +bit-flips, app-logic corruption, sub-object partial-repair) rather than +generic "durability." diff --git a/docs/lotus/rp-seal-v1/C2.md b/docs/lotus/rp-seal-v1/C2.md new file mode 100644 index 00000000..04251235 --- /dev/null +++ b/docs/lotus/rp-seal-v1/C2.md @@ -0,0 +1,801 @@ +# C2 — Domain C (erasure coding / seals), role: ADVERSARY + +**Cell:** C2 +**Charge:** build the injection matrix every candidate seal scheme must survive; produce +the scheme × injection truth table with coding-theory justification; adjudicate the +phase/comma hypothesis against boring controls; report DELETE where warranted. + +**Posture note.** I am the adversary. I concede what survives explicitly, because a report +that finds everything broken is as useless as one that finds nothing broken. Three things +in the existing code survive my attack and I say so in place. Everything else below is an +attempt to make the idea false. + +--- + +## SOURCE ARCHAEOLOGY + +Every claim in this section is a read of committed code at the paths given. No cargo was +run; nothing was written to any repo. + +### A. There is no integrity layer in Lance 9.0.0 to build on, or to be redundant with + +Two-column discipline applied. + +| Column | Path | Finding | +|---|---|---| +| **A — pinned 9.0.0** | `/tmp/sources/lance-9` | `grep -riE "checksum\|crc"` over `rust/lance-file/`, `rust/lance-table/`, `rust/lance-io/`, `rust/lance-encoding/src/decoder.rs` → **zero hits**. `protos/table.proto` → zero `checksum`/`crc`/`digest` hits except `FragmentReuseIndexDetails.FragmentDigest`. | +| **B — current upstream** | `/tmp/sources/lance-main` (shallow clone 2026-08-18) | Same greps → one hit, a *comment* at `rust/lance-table/src/io/commit.rs:244-247` explaining that **"An ETag is not necessarily a content checksum"**. Still no data checksum. | + +`FragmentDigest` (`protos/table.proto:565-571` in column A; `:572` in column B) is +`{ id: uint64, physical_rows: uint64, num_deleted_rows: uint64 }` — **rewrite bookkeeping, +not integrity**. It sits inside `FragmentReuseIndexDetails.Group` alongside +`changed_row_addrs: bytes` ("a roaring treemap of the changed row addresses"). That group +*is* an old-row-address → new-row-address permutation witness, but it carries no digest of +any byte. + +Only hashing anywhere near the data path is functional, not protective: +`rust/lance-core/src/utils/bloomfilter/sbbf.rs:36` (XxHash64, Parquet SBBF spec), +`rust/lance-encoding/src/statistics.rs:244` (xxh3 for HyperLogLog cardinality). + +**Consequence for the whole program.** A seal layer here has *no baseline to beat and no +existing syndrome to compose with*. It is also not redundant with anything Lance does. The +residual threat model is therefore narrow and must be stated before any scheme is chosen: +object stores already give per-object checksums on PUT/GET, TLS gives transport integrity, +and local filesystems may or may not (ZFS/btrfs yes, ext4/xfs no). What is genuinely +uncovered is (i) bit rot on non-checksumming local storage, (ii) *software* faults — +wrong-slot writes, lost writes, stale reads, coalescing bugs, and (iii) the "did my commit +land?" ambiguity the writers already fight. **(ii) and (iii) dominate, and they are +identity/version problems, not channel-noise problems.** This distinction decides the +whole truth table below. + +### B. The shipped seals are hash-only, and three of them are defective + +#### B1. `ndarray/src/hpc/seal.rs` — 48-bit truncated BLAKE3, masked + +100 lines. `MerkleRoot([u8; 6])`, doc comment: *"48 bits = collision-safe for <10M nodes."* +`Plane::merkle` hashes `bits[i] & alpha[i]` (alpha-masked), `Plane::verify` returns +`Seal::Wisdom | Seal::Staunen`. + +Two arithmetic checks (computed, `python3`, not asserted): + +* **As a verifier** — probability a corrupted plane hashes to its own stored root is + `2^-48 = 3.55e-15` per check. Fine. The doc's number is defensible *in this reading*. +* **As a content address** — `MerkleRoot` derives `Hash + Eq + PartialOrd`, so it is + usable as a map key. Birthday over `N = 10^7` distinct roots: + `P(any collision) = 1 - exp(-N(N-1)/2^49) = 0.163`. **A 16.3% chance of at least one + collision at exactly the population the doc calls "collision-safe."** The doc does not + say which reading it means; both readings are reachable from the public API. + +`alpha` masking has an asymmetric blind spot: flipping `alpha` 1→0 where `bits`=1 changes +the masked byte (detected); flipping `alpha` 0→1 where `bits`=0 does not (undetected). +Alpha corruption is therefore half-invisible by construction. + +#### B2. `ndarray/src/hpc/merkle_tree.rs` — the multi-scale syndrome that already exists, and its four defects + +This is the closest in-tree thing to the program's "multi-scale syndrome" idea: root (48b) +→ 8 branches (8×48b) → 64 leaves (64×48b), packed into 8 Kbit for SIMD Hamming, with +`typed_staunen()` claiming semantic localization (`ContentChanged` / `NarsChanged` / +`EdgesChanged` / `QualiaChanged` / `MultipleChanges`). + +**Defect 1 — 71.88% of the sealed object is not sealed.** `BRANCH_REGIONS` +(`merkle_tree.rs:22-31`) covers metadata words +`[0,16) ∪ [4,8) ∪ [16,32) ∪ [32,40) ∪ [40,48) ∪ [56,64) ∪ [96,112)` += `[0,48) ∪ [56,64) ∪ [96,112)` = **72 of 256 words**. Uncovered runs: `[48,56)`, +`[64,96)`, `[112,256)` = **184 words = 1472 of 2048 bytes**. And +`root = BLAKE3(b0‖b1‖…‖b7)` (`:173-177`) — the root hashes *the branch hashes*, never the +raw container. Therefore **corruption anywhere in those 1472 bytes changes no leaf, no +branch, and no root. False-accept probability 1.0, not `2^-48`.** Leaves are subdivisions +of the same regions, so the leaf level does not rescue it. + +**Defect 2 — nested check groups destroy localization.** Branch 0 = `(0,16)`, branch 1 = +`(4,8)`; branch 1's support is a strict subset of branch 0's. Any real change to the NARS +words changes *both*, so `typed_staunen` can only ever return `MultipleChanges([0,1])`. +**The `NarsChanged` arm is dead code on real input.** The in-tree test admits this in a +14-line comment (`:443-471`) and then *fabricates* the condition by assigning +`tree_b.branches[1] = MerkleRoot([0xFF; 6])` directly. By this workspace's own +falsifiability rule that test is vacuous: no input reaches the branch it claims to cover. +In coding terms the check-matrix rows have nested supports, so the error-locus → syndrome +map is not injective onto realizable syndromes. **Overlapping parity groups do not add +localization; they subtract it.** + +**Defect 3 — four of branch 1's eight leaves are structurally identical.** The leaf split +is `w_start = start + (leaf·len)/8`, `w_end = start + ((leaf+1)·len)/8` with a degenerate +fallback. For branch 1 (`len = 4`) the sub-ranges compute to +`(4,4) (4,5) (5,5) (5,6) (6,6) (6,7) (7,7) (7,8)`; the degenerate `(4,4)` falls to +`meta[4..5]` — the same input as `(4,5)`. So leaves 8≡9, 10≡11, 12≡13, 14≡15. Eight leaf +slots carry four leaves' worth of information. Integer-division subdivision silently +degenerates whenever `region_len < leaves_per_branch`. + +**Defect 4 — `xor_diff` fabricates a tree that is not one.** `:287-332` XORs two trees' +packed bits and then *re-reads* the XORed bytes back into `root`, `branches`, `leaves`. +The XOR of two truncated BLAKE3 values is not the BLAKE3 of anything. The returned struct +satisfies `root ≠ H(branches)` and every field is a type-correct lie. Nothing downstream +checks the invariant, so any consumer that treats a diff tree as a tree is accepting +garbage with full type safety. `test_xor_diff_self_zero` only asserts `a ⊕ a = 0`, which is +true of any XOR and proves nothing about the representation. + +#### B3. `lance-graph-cognitive/src/fabric/firefly_frame.rs` — an ECC with literally zero detection capability + +Header (`udp_transport.rs:29`) advertises *"0x80-0x9F: ECC (Reed-Solomon for error +correction)"*; `compute_ecc` (`:478-494`) is commented *"Real implementation would use +BCH(1247, 1024)."* What is actually there: + +```rust +fn verify_ecc(words: &[u64; Self::WORDS]) -> Option<[u64; Self::WORDS]> { + let corrected = *words; // <- never modified + ... + if syndrome == 0 { Some(corrected) } + else if syndrome.count_ones() <= 1 { + // Single bit error - correctable + // Find and flip the bit <- never implemented + Some(corrected) } + else { Some(corrected) } // "may be uncorrectable" +} +``` + +**All three arms return the input unchanged, and `decode()` treats `Some(_)` as success.** +`verify_ecc` cannot fail. It is a 100% false-accept path wearing the name of a BCH decoder. + +The fold underneath is also blind in the way that matters most here. +`ecc[i % 4] ^= word` is GF(2)-linear and **commutative**: it is invariant under *any +permutation of the 16 words*. The one property this whole system needs a seal to +protect — that worker completion order must not become semantic order — is exactly the +property an XOR fold cannot see. On top of that, `syndrome ^= expected[i] ^ received[i]` +folds four u64 words down to one, so any error affecting `ecc[a]` and `ecc[b]` identically +cancels; and `ecc[i] = (ecc[i] & …FE) | (parity & 1)` overwrites bit 0 of the fold with a +self-referential popcount parity, discarding one bit. + +#### B4. `lance-graph-callcenter/src/unified_audit.rs` — FNV-1a as a tamper-evidence chain + +`AuditMerkleRoot(pub u64)`, `chain(prev, salt, entry) = FNV1a64(prev ‖ salt ‖ entry)` +(`:97-127`), doc: *"cross-domain audit logs are unlinkable"*, *"safe to persist + +cross-binary verify."* + +FNV-1a is `h ← (h ⊕ b)·PRIME mod 2^64` with **no finalization and no state truncation**: +the output *is* the internal state. Therefore `chain` is **length-extendable by +construction** — anyone holding the published root can compute the root of +`(prev ‖ salt ‖ entry ‖ suffix)` without knowing `salt`. The salt is a prefix of a +Merkle-Damgård-shaped iteration, i.e. the textbook broken-MAC construction. The +"unlinkability" claim rests on salt secrecy that the construction does not provide. It is +also 64 bits and not collision-resistant (RFC 9923 / draft-eastlake-fnv states FNV is +non-cryptographic outright). + +*Determinism is a real requirement here and FNV does satisfy it.* The fix is not "drop +FNV"; it is "use a keyed, non-extendable, deterministic PRF" — BLAKE3 keyed mode is +already vendored in this workspace (`ndarray/src/hpc/blake3.rs`) and is deterministic +across platforms and Rust versions. + +#### B5. `lance-graph-planner/src/persist_sink.rs:392-428` — the shipped cycle seal + +`DetachedCycleBatch::content_hash` — FNV-1a 64 over `(cycle, base_version)` then, per +canonical landing, `(stream_position, owner, row, move-tag, [mailbox, wcp], +payload.len(), payload)`. This is the durable idempotency key: on an ambiguous commit, +`find_frame(cycle)` returns the stored hash and equality ⇒ `Reconciled` (the batch is +declared already-landed), inequality ⇒ `HashConflict` (fail closed). + +**Three concessions, stated plainly because they are real.** + +1. **The encoding is injective.** Every field is fixed-width or length-prefixed and the + `Option` is tagged, so each record is self-delimiting; a concatenation of self-delimiting + records is uniquely decodable. There is no framing-ambiguity collision here. I looked + for one and it is not there. +2. **The comparison is scoped by `cycle`,** so an accidental false accept requires two + *different* batches for the *same* cycle id to collide: `2^-64 = 5.4e-20` per honest + retry. Negligible. The 64-bit width is not the practical risk it would be for a global + content-address. +3. **`base_version` is deliberately inside the hash** (`:392-400`), which correctly makes a + re-derived frame fail closed rather than launder a divergence. That is good design and + I could not break it. + +**Two attacks that do land.** + +* **The seal covers the recipe, not the dish.** `freeze` (`:377-389`) computes the hash over + `casts` (the landings), then separately folds `image.insert(s.row, s.payload.clone())` + — "later stream position wins". `build_batch` (`cycle_sink.rs:457-520`) writes + `1 + landings.len() + image.len()` rows: the frame row, the landing rows, **and the image + rows**. `batch_hash` is never computed over `image`. A defect in the coalescing fold — + a changed tie-break rule, a map-ordering change, a partial fold — produces a *different + durable image* under an *identical* `batch_hash`, and the reconciliation path will report + `Reconciled` on it. The idempotency key seals the inputs and not the artifact. +* **The order-independence claim is conditional and unstated.** `:359-361` claims + *"Identical completed sets yield identical hashes regardless of worker completion order + (the freeze canonicalizes first)."* The canonicalizer is + `order_cycle_stably(rows, key) = rows.sort_by_key(key)` (`:133-135`) — a *stable* sort, + and its own test asserts *"equal keys keep arrival order"* (`:1567-1572`). So the claim + holds **iff `stream_position` is unique within the cycle**. The doc on + `stream_position` (`:175-183`) guarantees only *"monotonically increasing per owner + across cycles"* and explicitly says *"this layer does not re-key by + `(CycleId, stream_position)`."* Per-owner monotonicity does not imply + cross-owner uniqueness. If two owners ever share a `stream_position`, the canonical + order — and hence `batch_hash` — becomes arrival-order dependent, and an honest retry + after a lost acknowledgement produces a *different* hash for the *same* work → + permanent `HashConflict` → the cycle can never be reconciled. **Fails closed (safety + preserved) but wedges liveness.** This is a one-line falsifier, pre-registered below. + +`batch_hash` is also written into **every one of the `1 + |landings| + |image|` rows** +(`push_common`, `cycle_sink.rs:481-489`). For a 64 Ki-row cycle that is ≥ 512 KiB of a +repeated 8-byte constant before Lance encoding. Whether Lance's encoders collapse it is a +measurement, not an assumption (metric: seal metadata bytes). + +### C. What the substrate can and cannot execute + +`ndarray/src/hpc/blake3.rs` exists (vendored BLAKE3, with `hash_many` SIMD transposes). + +GF(2^8) SIMD multiply — the Plank/Greenan/Miller split-table technique (FAST'13), which is +how every fast software RS is built — needs a 16-entry in-lane byte table lookup +(`_mm256_shuffle_epi8` / `vqtbl1q_u8`). Measured availability in `ndarray::simd`: + +| tier | primitive | status | +|---|---|---| +| AVX2 | `U8x32::permute_bytes` (`simd_avx2.rs:2645`, *"Matches `_mm256_shuffle_epi8`"*, high-bit zeroing) | **native — usable** | +| AVX-512 | `U8x64::permute_bytes` (`simd_avx512.rs:695-711`) | VBMI → `_mm512_permutexvar_epi8`; **AVX-512F without VBMI (Skylake-X, Cascade Lake, Ice Lake-SP) → scalar loop over a stack array** | +| 64-byte composed type | `simd_avx2.rs:1435` | scalar loop | +| **NEON** | — | `grep -c 'vqtbl' src/simd_neon.rs` → **0**; `grep -c 'permute_bytes'` → **0** (2519 lines). **No byte-table primitive exists.** | +| wasm / scalar | `simd_scalar.rs:1528` | scalar loop | + +`.cargo/config.toml` baseline is **x86-64-v3 (AVX2)**, not v4 — the root `CLAUDE.md` +claim of "`target-cpu=x86-64-v4` (AVX-512 mandatory)" is stale relative to the committed +config, which explains the v3 choice at length. + +**Consequence.** Any RS scheme is fast on AVX2 and on VBMI-class AVX-512, and collapses to +a scalar byte loop on plain AVX-512F, NEON, and wasm. The workspace's "all SIMD from +`ndarray::simd`" invariant forbids reaching for `vqtbl1q_u8` in a consumer crate, and the +missing-capability STOP rule says the consumer must not grow the substrate. **A NEON +`permute_bytes` is a substrate-first prerequisite of every RS candidate.** This is a +concrete, cheap, non-negotiable gate that no amount of scheme design substitutes for. + +### D. The temporal machinery, read literally + +`git show 386a6fd8…:docs/TEMPORAL-TIME-TRAVEL.md` (OGAR) and +`lance-graph-planner/src/temporal.rs`. What is actually there, for the stale-chunk +injection: + +* `QueryReference { ref_version, mode: STRICT|AWARE|RETRO, rung }` is a **query-time + annotation**, explicitly *"No new storage, no new contract, no new container."* +* `CONTEMPORARY / ANACHRONISTIC / SPOILER` are decided by comparing + `row.lance_version` to `V_ref` — i.e. **by reading a stored version column**, never by + any algebraic property of the data. +* HLC `(server_id, lance_version, hlc_tick)` is the cross-server deinterlace key, again a + stored coordinate. +* `temporal.rs` names two layers: **Layer 1 causal deinterlacing** (`local_trajectories`: + split the interleaved global log back into per-owner chains) and **Layer 2 epistemic + projection** (`classify` / `deinterlace`). + +**This is the decisive archaeological fact for my cell.** Every temporal distinction in the +design is made by *comparing a stored version coordinate*. Nothing in the temporal layer is +a syndrome, and nothing in it is derivable from data bytes. Any proposal that a parity +syndrome will distinguish stale-version contamination from corruption is proposing to +re-derive, algebraically, a fact the system already stores explicitly — and, as §MECHANISM +shows, it cannot. + +--- + +## MECHANISM + +### The six candidate schemes + +* **S1 hash-only** — a `b`-bit digest per chunk. (Shipped shape: `seal.rs` b=48; + `persist_sink.rs` b=64 per cycle.) +* **S2 flat RS(n,k)**, `m = n−k` parity symbols over the whole cycle, **no per-chunk hash**. +* **S3 row/col P+Q** — a 2-D array with 2 parities per row and 2 per column (RAID-6-per-line). +* **S4 product code** RS(n₁,k₁) ⊗ RS(n₂,k₂). +* **S5 cascade parity** — local parity per Morton level (4/16/64/256/1024/4096) plus a + global parity. This is an LRC / hierarchical code. +* **S6 hybrid** — per-chunk hash **bound to (locus, version)** + flat RS. + +### The nine injections + +`I1` single erasure (position known) · `I2` double erasure · `I3` silent corruption +(1 chunk, position unknown) · `I4` wrong-slot substitution (a valid chunk at the wrong +locus) · `I5` stale chunk (right locus, right shape, old version) · `I6` duplicate chunk · +`I7` two correlated failures (one failure domain) · `I8` boundary-correlated failures +(faults at parity-group boundaries) · `I9` **phase-aligned failures** (the fault pattern is +congruent to the cascade/parity-group structure itself). + +Outcome codes: **FA** false-accept (verifier says OK, data is wrong) · **D** detect · +**L** localize (name the faulty symbol) · **R** repair · **—** not applicable. + +### The truth table + +| | S1 hash-only | S2 flat RS (no hash) | S3 row/col P+Q | S4 product | S5 cascade / LRC | S6 hash+RS (locus+ver bound) | +|---|---|---|---|---|---|---| +| **I1** single erasure | D L, no R | **D L R** | **D L R** (local) | **D L R** (local) | **D L R**, *r* reads | **D L R** | +| **I2** double erasure | D L, no R | **R** if m≥2 | **R** | **R** if pattern ∉ rectangle | R if ≤ local capacity, else escalate global | **R** if m≥2 | +| **I3** silent corruption | **D L**, FA `2^-b` | m=1: **D only, no L, parity pollution**; m≥2: D L R via key equation, corrects ⌊m/2⌋ | D L R (row∩col) | D L R (1 error) | needs hashes; local single parity ⇒ D only | **D L R** — hash turns the error into an erasure, so **m** faults not ⌊m/2⌋ | +| **I4** wrong-slot | **FA** if hash not locus-bound (certain for identical/default content); else D L | D L of destination only; a *swap* is 2 errors ⇒ needs m≥4 | D L destination | D L destination | D L destination | **D L + positively diagnoses wrong-slot** (embedded locus ≠ found locus) | +| **I5** stale chunk | D if content differs; **cannot classify as stale** | indistinguishable from an error; **beyond-radius miscorrection ⇒ FA with silent alteration** | same | same | same | **D L + positively diagnoses stale** (embedded version < expected) — *the only scheme that can* | +| **I6** duplicate chunk | FA if digest is over an unordered set or an **XOR fold**; D if framing is length/count-bound | as I4 | as I4 | as I4 | as I4 | **D** | +| **I7** two correlated failures | D L per chunk | **R** for any m erasures, correlation-agnostic (MDS) | good when the correlation unit is a line | as S3 | **degrades to the local group's capacity** | **R** up to m, any pattern | +| **I8** boundary-correlated | D L per chunk | **R** (MDS is pattern-agnostic) | pattern-dependent | pattern-dependent | **worst case** | **R** up to m | +| **I9** phase-aligned | D L per chunk (a genuine strength: hashes have no preferred pattern) | **R** — MDS has no adversarial pattern | rectangle pattern uncorrectable at weight `d₁d₂` | **rectangle uncorrectable** | **catastrophic** — see below | **R** up to m | + +### The coding theory behind each column + +**MDS and the error/erasure gap.** RS(n,k) has minimum distance `d = n−k+1 = m+1` +(Singleton, met with equality). It corrects **any m erasures** (positions known) but only +`⌊m/2⌋ errors` (positions unknown). Attaching a per-chunk hash converts every error into an +erasure and therefore **doubles the effective correction capacity for silent corruption**. +This is not a subtle result and it is the single largest lever in the whole table. The +patent literature states it plainly for product codes too: *"Data and parity symbols can +have internal checksums to detect corruption, so upon decoding, bad symbols can be +identified"* (US 9,600,365, local erasure codes for data storage). **S6's dominance over +S2, S3, S4, S5 on the detection axis is textbook, not a discovery.** + +**Product codes are not MDS, and the gap is not academic.** For RS(n₁,k₁)⊗RS(n₂,k₂), +`d = d₁·d₂`, but the Singleton bound for the same parameters is `n₁n₂ − k₁k₂ + 1`, which is +strictly larger. Worked instance at comparable overhead: flat **RS(16,12)** stores 12 data +symbols and corrects **any 4 erasures**. Product **(4,3)⊗(4,3)** stores 9 data symbols in +16 (a *worse* rate), has `d = 4`, and **fails on a 2×2 erasure rectangle** — one of the +weight-4 patterns RS handles. The product code stores less and protects worse. It buys +exactly one thing: repair locality. + +**The direct answer to the charge's product-code question.** *Do a 2-D product code's +row/column syndromes genuinely localize better than per-chunk hashes + flat RS?* **No — +they localize worse beyond one fault.** One silent error at `(i,j)`: row-syndrome `{i}` ∩ +col-syndrome `{j}` gives a unique cell — as good as a hash, using `r+c` parity symbols +instead of `r·c` hashes. Two errors at `(i,j)` and `(i′,j′)` with `i≠i′, j≠j′`: the +syndromes are `{i,i′}` and `{j,j′}`, and the candidate set is the **four corners** +`{(i,j),(i,j′),(i′,j),(i′,j′)}`. Over a large field the RS row-syndrome carries the error +*magnitude*, which usually breaks the tie — but a rectangle pattern with matched magnitudes +is exactly the weight-`d₁d₂` uncorrectable word, and it is not a contrived case: it is what +a two-device failure looks like when the array is laid out device-per-column. Per-chunk +hash + flat RS localizes **every** faulty chunk independently, with `2^-b` false-accept +each, no ambiguity at any multiplicity. Product syndromes win only on parity *volume* +(`r+c` vs `r·c` hashes), which is a metadata-bytes trade, not a localization win. + +**The locality bound is the ceiling on cascade parity.** For a code with locality `r`, +`d ≤ n − k − ⌈k/r⌉ + 2` (Gopalan et al.). Every level of locality costs distance. Nested / +hierarchical constructions (Pyramid codes, Azure-LRC, Xorbas, Duminuco–Biersack +hierarchical codes) are the named prior art, and none of them is MDS. **A cascade parity +over Morton levels 4/16/…/4096 cannot be MDS, by theorem, and its distance shortfall is +computable in advance from `⌈k/r⌉`.** This is not an empirical question. + +**I9 is where the whole architecture's central tension lives.** Morton order exists in this +design to make spatially/semantically adjacent rows *physically adjacent*. Physical +adjacency is what a fragment, an object, and a device are made of. So a device-scale +failure **is** a phase-aligned failure with respect to any parity group that was drawn on +Morton boundaries. **The very locality that Morton order buys for queries is exactly the +correlation that destroys a locality-aligned parity group.** Flat RS is immune (MDS: any +`m` erasures, no preferred pattern). Cascade parity is maximally exposed. Stated as a +constraint the program cannot escape: + +> Query locality and repair locality want the **same** grouping. Failure independence wants +> the **opposite**. One order cannot deliver all three. + +The classical resolution is **declustering / interleaving**: group by Morton for queries, +then assign parity-group membership by a stride coprime to the group size so that one +failure domain contributes at most one member to any group. CIRC (CD-ROM) does exactly +this; declustered-parity RAID does exactly this; interleaved RS does exactly this. And this +workspace already ships the generator: `helix/src/curve_ruler.rs` — `(start + 4k) mod 17`, +`gcd(4,17)=1`, proven by test to visit all 17 residues. **This is the honest kernel of the +"incommensurate schedule" intuition, and its correct job is interleaving, not coefficients.** + +### The phase/comma hypothesis, adjudicated + +The admissible framing is: does a deterministic incommensurate schedule for **RS coefficient +orientation across cascade levels** (1) preserve matrix rank, (2) preserve MDS/correction, +(3) reduce cross-level syndrome collisions, (4) improve localization, (5) cost ≤ SIMD work, +(6) reconstruct from locus/level rather than storage? Controls it must beat: standard RS +coefficients, sequential finite-field powers, random-but-seeded valid coefficients, no +modulation. + +| # | Condition | Verdict | +|---|---|---| +| 1 | rank preserved | **Yes — but vacuously.** A Vandermonde/Cauchy matrix over *distinct* evaluation points is full-rank; permuting the points permutes the generator's columns, yielding a **permutation-equivalent** code. **Every control also passes.** Not evidence. | +| 2 | MDS / correction preserved | **Yes — vacuously.** Permutation-equivalent codes have identical weight enumerators, hence identical `d` and identical MDS status. **Every control also passes.** Not evidence. | +| 3 | fewer cross-level syndrome collisions | **IMPOSSIBLE under permutation.** For RS, a single-symbol error at position `i` produces syndrome `(e·α_i^0, …, e·α_i^{2t−1})`; the locator is recovered as the ratio `S₁/S₀ = α_i`. A permutation of the evaluation points **relabels which index maps to which locator but leaves the achievable syndrome SET invariant.** Two levels sharing one syndrome register therefore collide exactly as much after the permutation as before. Distinguishability requires **disjoint** syndrome sets, i.e. GRS *column multipliers* `β_ℓ` in distinct multiplicative cosets — a different construction, not an orientation schedule. That construction is (a) known (coset-based GRS evaluation sets are standard), (b) hard-capacity-bounded at `⌊(q−1)/n⌋` disjoint levels (for `GF(2^8)`, `n=64` ⇒ **3 levels**), and (c) strictly dominated by a **1-byte level tag** — or by nothing at all, since the parity's own storage address already names its level. **Fails, and the problem it solves is self-inflicted.** | +| 4 | better localization | **No.** Localization comes from converting errors to erasures (per-symbol hashes) or from 2-D syndrome intersection. Coefficient orientation contributes to neither. | +| 5 | cost ≤ SIMD work | **Fails, and predictably.** Software RS multiplies via 4-bit split tables held in shuffle registers. A level-dependent coefficient schedule means a **different table set per level**: table regeneration or reload at every level boundary, against a fixed-coefficient encoder whose tables stay L1/register-resident. The schedule *adds* cache and table pressure proportional to the level count. And on this substrate the tables only run at speed on AVX2 / AVX-512-VBMI at all (§SOURCE ARCHAEOLOGY C). | +| 6 | reconstruct from locus/level | **Yes — vacuously.** Boring RS already derives coefficient `α^i` from index `i`; nothing is stored. **Every control also passes.** Not evidence. | + +**Score: zero of six conditions are both non-trivial and satisfied.** Three are satisfied by +every control including "no modulation"; one is impossible as framed; one is dominated by a +one-byte tag; one fails on cost. Against the required control groups the schedule wins on +nothing. + +> **VERDICT ON THE PHASE/COMMA SCHEDULE: DELETE.** +> The salvage — and it is a real one — is that the coprime-stride generator already in +> `curve_ruler.rs` is a sound full-permutation interleaver, and interleaving is precisely +> the mechanism that defuses I9. Keep the stride; move it from the coefficients to the +> parity-group membership map. That relocation is KNOWN prior art (CIRC, interleaved RS, +> declustered parity), so it enters as engineering, not as a claim. + +### The stale-chunk question, answered from coding theory + +*Can any syndrome distinguish stale-version contamination from corruption?* + +**In general: no, and the reason is structural.** A syndrome is `S = H·r = H·(c + e) = H·e` +— a function of the *error* alone. A stale chunk contributes `e = c_old|_i − c_new|_i`; a +corrupt chunk contributes an arbitrary `e`. There is no algebraic predicate on `H·e` that +separates "this difference happens to be a previous codeword's symbol" from "this +difference is arbitrary", because the stale symbol is drawn from the same alphabet and its +difference is an ordinary field element. Worse, for RS **without** per-symbol hashes, a +number of stale symbols exceeding `⌊m/2⌋` can decode to a *wrong nearby codeword* — a +beyond-radius miscorrection that **silently alters data while reporting success** (`FA` in +the I5/S2 cell). Staleness is an identity fact; identity is not in the syndrome. + +**The known correct mechanism is self-describing storage**, not coding: put the version *in +the chunk* and compare. ZFS is the canonical instance — a 256-bit checksum stored in the +*parent* block pointer (so the location is part of the verification context) plus the birth +transaction group, which is exactly what lets ZFS detect lost writes and misdirected writes +that a same-block checksum cannot. "Parity Lost and Parity Regained" (Krioukov et al., +FAST'08) is the definitive taxonomy: it model-checks real RAID protection schemes, *finds +flaws in every one*, and names **parity pollution** — recomputing parity from already-corrupt +data spreads a single fault across the stripe. That failure mode applies verbatim to any +cascade-parity scrub in this design that recomputes a level's parity without first +validating its members. **S6 is the scheme the literature converged on, and my table +reproduces that convergence independently.** + +**One partial exception worth pre-registering rather than dismissing.** If the system +retains the *previous version's parity* (or a per-version parity delta), then a stale +symbol at locus `i` produces syndrome `H·(c_old|_i − c_new|_i)` — and `c_old|_i − c_new|_i` +is **exactly the per-row delta the system already knows**, because Lance's +`FragmentReuseIndexDetails.Group.changed_row_addrs` is a roaring treemap of precisely which +row addresses changed. So the check "is the implied error magnitude at the implied locus +equal to the known version delta?" is computable, and a match is *evidence of staleness +rather than corruption*. Honest bounds: it is **diagnostic, not authenticating** (random +corruption matches by chance with probability `2^-symbol_bits` per locus, and an adversary +who writes the old version produces a match by construction); its value is operational +triage — re-read at the correct version vs. invoke repair — which is a real cost difference +in read amplification. Classified in SURPRISES; the prior-art search is documented there. + +--- + +## EXPECTED BENEFIT + +Stated as what would survive if the program executes the boring path, since that is what my +analysis says wins. + +1. **S6 closes a real gap that Lance does not close.** With no checksum anywhere in the + 9.0.0 format (§A), a locus+version-bound per-chunk digest is the *only* thing in the + stack that can distinguish "the bytes I got are not the bytes I wrote" from "the bytes I + wrote were wrong". This is the highest-value deliverable in Domain C and it does not + need erasure coding at all. +2. **Errors → erasures doubles correction capacity** for the same parity overhead + (`m` vs `⌊m/2⌋`). Free, once (1) exists. +3. **Positive diagnosis of wrong-slot and stale** — uniquely available to S6, and these are + the two failure modes this architecture actually generates (a versioned, relocating, + compacting store with a coalescing writer), as opposed to channel noise, which the object + store already covers. +4. **The coprime-stride interleaver defuses I9 at ~zero cost** — a permutation of group + membership, one modular multiply per assignment, no change to the code's algebra, and it + converts a device-aligned failure from "one group loses everything" to "every group loses + one". +5. **Four in-tree defects are fixable cheaply and independently of any scheme choice:** the + 1472-byte coverage hole (B2/defect 1), the nested branch regions (defect 2), the + degenerate leaves (defect 3), and `verify_ecc`'s unconditional `Some` (B3). None requires + a research programme; all are currently silent. + +--- + +## EXPECTED FAILURE + +Ordered by my confidence that they will actually happen. + +1. **Cascade parity loses to flat RS on reliability at equal overhead, by theorem + (Gopalan bound) and by phase-alignment (I9), and it is the phase-alignment that will + bite** because Morton locality *manufactures* the correlated failure domain. +2. **The phase/comma coefficient schedule loses to every control.** Adjudicated above; the + three conditions it "passes" are passed identically by "no modulation". +3. **Product / row-col codes lose to hash+RS on localization at multiplicity ≥ 2.** The + four-corner ambiguity is structural. +4. **Hash-only seals false-accept on the exact faults this system generates.** Wrong-slot + substitution of identical content (all-zero rows, default payloads) is a *certainty*, not + a probability, for any digest that does not bind the locus. `seal.rs::merkle` and + `merkle_tree.rs::hash_words` both take no position input. +5. **XOR-fold seals cannot see reordering.** Commutativity is the whole problem: the + property the architecture most needs sealed (completion order ≠ semantic order) is in + the fold's kernel. `firefly_frame::compute_ecc` is the in-tree instance. +6. **RS on this substrate is 5–20× slower than the papers on NEON, wasm, and plain + AVX-512F**, because the byte-table primitive is missing or scalar there (§C). If the + program benchmarks on an AVX2 dev box and ships to ARM, the seal CPU number is wrong by + an order of magnitude. +7. **The seal will be computed over the wrong object.** Already true in-tree: `batch_hash` + covers the landings, the writer persists the image. Any new seal must name, explicitly, + *which bytes* it is a function of, and a test must prove that changing the persisted + artifact without changing the inputs changes the seal. +8. **`stream_position` collisions wedge reconciliation.** Fails closed, so not a corruption + risk — a liveness risk, and an invisible one until it fires in production. +9. **Scrub-induced parity pollution.** Any cascade scrub that recomputes a level's parity + before validating its members converts one silent corruption into a whole group's worth. + FAST'08 named this; it is the default outcome of a naive implementation. + +--- + +## EXPERIMENT + +Pre-registered. Metrics drawn from the program list. Each has a stated kill condition; a +kill means the branch is deleted, not re-scoped. + +### X-C2-1 — The injection harness (prerequisite for everything else) + +*Design.* A fault-injection layer over a synthetic cycle (65,536 × 512 B = 32 MiB) that +applies each of I1–I9 at a controlled multiplicity and, for I9, at strides congruent to +each cascade level (4/16/64/256/1024/4096). Each candidate scheme S1–S6 is asked to emit +one of {OK, DETECT, LOCALIZE(set), REPAIR(bytes)}. **Ground truth is the injection record.** +*Metric.* False-accept rate per (scheme, injection, multiplicity); localization precision +and recall. +*Baseline.* S1 with a 64-bit locus-bound digest. +*Control.* The **null injection** (no fault): every scheme must report OK on 10⁶ clean +trials. A scheme that flags a clean cycle is as broken as one that accepts a dirty one. +*Kill.* Any scheme with a **non-zero false-accept rate on I1–I3 at multiplicity 1** is +deleted immediately — that is the floor, and no downstream benefit compensates. +*Anti-vacuity.* The harness must first reproduce the four **known** in-tree false accepts +as positive controls: (a) corruption in metadata words `[112,256)` invisible to +`MerkleTree`; (b) `verify_ecc` accepting a 32-bit-flipped frame; (c) an XOR-fold accepting a +word permutation; (d) an unbound digest accepting a wrong-slot substitution of identical +content. **If the harness cannot detect these four, it is measuring nothing** and no other +result from it may be reported. + +### X-C2-2 — Does the phase schedule beat boring RS? (the DELETE gate) + +*Design.* Four arms at identical `(n,k)`, identical field, identical data: **A** standard RS +coefficients; **B** sequential finite-field powers; **C** random-but-seeded valid +coefficients; **D** the project's incommensurate orientation schedule. +*Metrics.* (i) rank / MDS verified exhaustively on small parameters; (ii) **cross-level +syndrome collision rate** — the fraction of single-fault syndromes at level ℓ that are also +realizable at level ℓ′; (iii) localization precision under I3 at multiplicity 1–4; +(iv) **encode/decode ns per MiB and L1d cache misses** (metric: seal CPU, cache misses). +*Baseline.* Arm A. *Control.* Arm C — a *seeded random valid* schedule is the honest null: +if D beats A but not C, the benefit was "any bijection", not this one. +*Kill (pre-committed, all four are DELETE conditions).* + - D's collision rate (ii) is **not strictly lower than C's** → DELETE. + - D's localization (iii) is **within noise of A** → DELETE. + - D's throughput (iv) is **worse than A by > 5%** or its L1d misses are higher → DELETE. + - D's advantage on (ii) is **matched by A plus a 1-byte level tag** → DELETE. +*Prediction (registered before running).* D loses on (ii) by construction — a permutation +preserves the syndrome set — and loses on (iv) from table pressure. I expect DELETE on two +independent grounds. + +### X-C2-3 — Product/row-col syndrome localization vs per-chunk hash + flat RS + +*Design.* Equal storage overhead across S3, S4, S6. Inject I3 at multiplicities 1, 2, 3, 4, +including the **rectangle pattern** explicitly. +*Metric.* Localization precision/recall; **repair read bytes and helper chunks touched** +(the axis product codes are supposed to win); decode CPU. +*Baseline.* S6. *Control.* S2 (flat RS, no hashes) — isolates how much of S6's benefit is +the hash rather than the code. +*Kill.* If S3/S4 localization at multiplicity 2 is **not better** than S6's, the "2-D +syndromes localize better" hypothesis is DISPROVED and product codes are retained, if at +all, **solely** on repair bandwidth. *Prediction:* S6 wins localization at every +multiplicity ≥ 2; S3/S4 win repair read bytes at multiplicity 1 by roughly `c/k`. + +### X-C2-4 — Phase-aligned failure sweep (the cascade gate) + +*Design.* Fix storage overhead. Sweep the failure-domain size across the Morton level sizes +(4/16/64/…/4096) and across a **coprime-stride-declustered** membership map. Arms: flat RS, +cascade parity Morton-aligned, cascade parity stride-declustered. +*Metric.* Fraction of injected patterns that are unrepairable; **repair read/write bytes and +helpers touched** per repair; **fragment count** touched. +*Baseline.* Flat RS. *Control.* A **random** membership permutation — if the coprime stride +does not beat a random permutation, the stride is decoration and only declustering mattered. +*Kill.* If Morton-aligned cascade parity's unrepairable fraction exceeds flat RS's at any +tested failure-domain size, **Morton-aligned parity groups are DELETED** (declustering +becomes mandatory, not optional). If stride-declustered cascade parity does not beat the +random control, **the stride is deleted** and plain declustering is used. + +### X-C2-5 — Stale-vs-corrupt discrimination (the temporal gate) + +*Design.* Inject I5 (stale chunk) and I3 (random corruption) at matched multiplicities into +a store carrying (a) no version stamp, (b) an in-chunk version stamp, (c) a retained +previous-version parity plus the FRI `changed_row_addrs` delta set. +*Metric.* Confusion matrix stale ↔ corrupt; **read amplification of the recovery action** +taken (re-read at version vs. full repair); **stale-version retention bytes** required by +arm (c). +*Baseline.* Arm (b) — the ZFS-shaped answer. +*Control.* Arm (a) must show **chance-level** discrimination. If it does not, the harness is +leaking ground truth and the whole experiment is void. +*Kill.* If arm (c)'s discrimination is not **strictly better than chance** *and* its +retention-byte cost is below the read-amplification it saves, the delta-syndrome +discriminator is DELETED and version stamping stands alone. *Prediction:* (b) ≈ perfect; +(c) better than chance but with a false-"stale" rate on random corruption of `2^-symbol_bits` +per locus, and a retention cost that only pays on workloads with a high stale-read rate. + +### X-C2-6 — Substrate gate (blocking; cheap; run first) + +*Design.* Implement the GF(2^8) split-table multiply **exclusively** through +`ndarray::simd` and measure encode throughput on AVX2, AVX-512-with-VBMI, +AVX-512-without-VBMI, NEON, and scalar. +*Metric.* GiB/s encode; seal CPU per 32 MiB cycle. +*Baseline.* A published software-RS number (ISA-L / jerasure class). +*Kill.* If NEON or plain AVX-512F is **> 5× below** the AVX2 arm, **no RS scheme may be +selected until `permute_bytes`/`vqtbl` lands in `ndarray::simd`** — substrate-first, per the +missing-capability STOP rule. Measured today: NEON has **no** byte-table primitive +(`grep -c vqtbl src/simd_neon.rs` → 0) and AVX-512F-without-VBMI falls to a scalar stack +loop (`simd_avx512.rs:703-710`), so I expect this gate to fire. + +### X-C2-7 — Falsifiers for the four in-tree defects (independent of scheme choice) + +Each is a single test with a stated red-then-green disable run. + +| # | Assertion | Disable that must turn it red | +|---|---|---| +| a | corruption in metadata word 200 changes the `MerkleTree` root | extend `BRANCH_REGIONS` to cover `[112,256)` → green | +| b | a real change to the NARS words yields `NarsChanged`, not `MultipleChanges` | make branch 0 and branch 1 supports disjoint → green | +| c | `MerkleTree::from_cogrecord` yields 64 **distinct** leaves on distinct-content input | fix the degenerate sub-range → green | +| d | `FireflyFrame::decode` returns `None` on a 32-bit-flipped frame | implement a real syndrome → green | +| e | `batch_hash` changes when the persisted **image** changes with identical inputs | hash the image too → green | +| f | `content_hash` is invariant to arrival order when two owners share a `stream_position` | key by `(cycle, stream_position, owner)` → green | + +*Kill condition for (f):* if the test **passes as written** — i.e. duplicate +`stream_position` is impossible by an invariant I did not find — then the liveness concern in +§B5 is withdrawn and the invariant should be documented at the `stream_position` field. +I could not find that invariant in the source; the field's own doc explicitly declines to +key by `(CycleId, stream_position)`. + +--- + +## PRIOR ART + +Searches actually run, verbatim queries, and what each returned. Where a search returned +nothing on point, I say so — a documented empty search is what licenses a NOVEL-CANDIDATE +label, and an undocumented one licenses nothing. + +1. **`"Parity Lost and Parity Regained" disk corruption taxonomy lost write misdirected write RAID`** + → Krioukov, Bairavasundaram et al., FAST'08 (research.cs.wisc.edu/adsl/Publications/parity-fast08.pdf). + Model-checks real RAID protection schemes under single-error conditions, **finds flaws in + every scheme**, names **parity pollution**. This is the canonical taxonomy for I3/I4/I5 + and it is the source of my scrub-pollution warning. **Domain C's failure taxonomy is + KNOWN and eighteen years old.** +2. **`locally repairable codes LRC Azure Xorbas pyramid codes repair locality minimum distance bound`** + → Kolosov et al., USENIX ATC'18 *On Fault Tolerance, Locality, and Optimality in LRCs*; + Kadekodi et al., FAST'23 *Practical Design Considerations for Wide LRCs*; + Papailiopoulos–Dimakis arXiv:1206.3804; errorcorrectionzoo.org LRC entry. + Establishes Pyramid/Azure-LRC/Xorbas as data-LRCs, the locality-vs-distance bound, and + average-locality metrics. **Cascade parity = LRC. Fully KNOWN; the bound is a theorem.** +3. **`erasure coding stale version detection versioned storage syndrome distinguish stale replica from corruption`** + → arXiv:1411.4762 (sparsity-exploiting erasure coding for delta-based versioning), + arXiv:1503.05434 (compressed differential erasure codes), Springer DEC chapter, and work + on Generalized Pyramid Codes for multi-version storage. **Every hit uses version deltas + for STORAGE EFFICIENCY. The search returned nothing on using a syndrome to discriminate + staleness from corruption** — the summary said so explicitly. This documented gap is what + supports the classification in SURPRISES #3. +4. **`FNV-1a 64 collision attack non-cryptographic hash invertible length extension weakness`** + → RFC 9923 / draft-eastlake-fnv-35 (the FNV spec itself), plus multiple sources stating + FNV-1a *"is not suitable for cryptographic purposes… vulnerable to hash collisions and + pre-image attacks."* Length extension follows directly from the construction (output = + internal state, no finalization) rather than from a citation. **KNOWN; the spec says so.** +5. **`self-validating writes physical identity block checksum lost write detection ZFS birth txg version mirror`** + → ZFS documentation (FreeBSD Handbook ch. 22/23, Klara Systems, Rossmann): 256-bit block + checksum **stored in the parent block pointer, not with the block**, forming a + self-validating Merkle tree; enables identifying *which* mirror copy is correct; + TXG-based consistency. Plus US 7,454,668 *Techniques for data signature and protection + against lost writes* and US 9,892,153 *Detecting lost writes*. + **Locus-bound + version-bound checksums are the KNOWN answer to I4 and I5, with patents.** +6. **`product code row column parity error localization versus per-chunk checksum plus Reed-Solomon storage`** + → US 9,600,365 (local erasure codes for data storage) states directly: *"Data and parity + symbols can have internal checksums to detect corruption, so upon decoding, bad symbols + can be identified."* Also US 9,214,964 / 9,490,849 (product codes in HDD), CD-ROM RS + Product Code (P/Q dimensions), robust product codes. + **"Hashes turn errors into erasures" is KNOWN and patented. The product-code localization + hypothesis is DISPROVED by the same literature that defines product codes.** +7. **`space filling curve Hilbert Morton order data placement erasure coding repair locality query locality joint optimization`** + → Extensive SFC locality literature (Hilbert vs Morton locality, bit-interleaving, + partitioning, point-cloud neighbourhood search). The result summary states explicitly + that the hits *"don't specifically address the integration with erasure coding repair + locality, query locality, or joint optimization."* **Documented empty search on the joint + question** — basis for SURPRISES #1. +8. **`generalized Reed-Solomon code multiplier coset distinct evaluation points syndrome disjoint identify which code stripe`** + → GRS definition with evaluation points `α_i` and **column multipliers `β_i`** (Hall, MSU + ch. 5; SageMath coding-theory reference); self-dual GRS constructions built from + **subgroups of finite fields and their cosets**; US 5,872,799 *Global parity symbol for + interleaved Reed-Solomon coded data*. **Coset-multiplier constructions and global-parity- + over-interleaved-RS are both KNOWN. The phase schedule has no room left to be novel in.** +9. **`algebraic signatures Schwarz Miller "Store, Forget, and Check" remote storage integrity Galois field signature parity commutes`** + → Schwarz & Miller, ICDCS'06 (ssrc.ucsc.edu/Papers/schwarz-icdcs06.pdf). *"the signature + of parity gives the same result as the parity of the signatures"*, over GF(2^16), combined + with m/n erasure coding, blinded against collusion. Plus a decade of follow-on + remote-data-possession work built on it. + **This is the "multi-scale composable syndrome" idea, published in 2006. Any hierarchical + syndrome proposal in this program must cite it and beat it, or be labelled a + reimplementation.** + +**In-repo prior art** (this workspace's own, which a session is more likely to re-derive +than to find): `ndarray/src/hpc/merkle_tree.rs` already implements a 3-level multi-scale +digest with semantic localization; `ndarray/src/hpc/seal.rs` already implements the +verification predicate; `helix/src/curve_ruler.rs` already implements the coprime-stride +permutation; `lance-graph-planner/src/persist_sink.rs:392` already implements a content +seal with locus binding. **Four of the program's "new" ideas exist in the tree, three of +them defective.** Fixing them is cheaper than designing successors. + +--- + +## SURPRISES + +1. **Morton locality manufactures exactly the correlated failure that kills locality-aligned + parity.** Query locality and repair locality want the same grouping; failure independence + wants the opposite; no single order gives all three, and the resolution (decluster the + parity-group membership by a coprime stride while keeping Morton for queries) is available + in-tree via `curve_ruler.rs`. — **TRANSFER.** Declustered parity and interleaved RS + (CIRC) are decades old; the SFC-locality literature is vast; search #7 returned no work + *joining* them, but the join is an application of two known mechanisms, not a new one. + +2. **An XOR-fold "ECC" is blind to permutation, and permutation-blindness is precisely this + architecture's cardinal sin.** `firefly_frame::compute_ecc` folds with `^=`, which is + commutative, so scrambled worker-completion order — the one thing the design insists must + never become semantic order — lies exactly in the fold's kernel. And + `verify_ecc`'s three arms all return `Some(*words)` unmodified, so it cannot fail at all. + — **KNOWN** (linear checksums have known kernels; RFC-level knowledge), **but the + in-tree instance is a live 100% false-accept path** that no test covers. + +3. **A retained per-version parity delta can discriminate stale from corrupt, because the + delta the syndrome implies is already stored** — Lance's + `FragmentReuseIndexDetails.Group.changed_row_addrs` is a roaring treemap of exactly which + rows changed, so "does the implied error magnitude at the implied locus equal the known + version delta?" is computable, giving triage (re-read vs. repair) rather than blind + repair. Bounded honestly: diagnostic, not authenticating; random corruption false-matches + at `2^-symbol_bits`. — **TRANSFER.** Differential/compressed erasure codes for versioned + data are published (search #3), but every hit uses deltas for storage efficiency; the + search returned nothing on using the delta *as a staleness discriminator*. The mechanism + is entirely known; only the role is new, so this is TRANSFER, not NOVEL-CANDIDATE. + +4. **A hierarchical syndrome whose check groups are NESTED localizes worse than one flat + digest.** `BRANCH_REGIONS[0] = (0,16) ⊃ BRANCH_REGIONS[1] = (4,8)` makes the + `NarsChanged` verdict unreachable on real input, and the in-tree test proves it by + fabricating the state it cannot produce. Multi-scale syndromes require **disjoint or + properly laminar** supports; overlap is not redundancy, it is destroyed injectivity of the + locus→syndrome map. — **KNOWN** in coding theory (independence of check equations), but + it is an easy and repeated design error, and it is live here. + +5. **"Multi-scale composable syndrome" is Schwarz & Miller 2006 — algebraic signatures, + where the signature of the parity equals the parity of the signatures.** Any hierarchical + seal in this program that does not cite it is reinventing a twenty-year-old primitive. + — **KNOWN, DISPROVED as novel.** + +*(Adjacent findings that did not make the top five: 71.88% of the `MerkleTree`-sealed +container is under no hash at all — false-accept probability 1.0, not `2^-48`; the shipped +`batch_hash` seals the landings while the writer persists the coalesced image, so the seal +covers the recipe and not the dish; `MerkleRoot`'s "collision-safe for <10M nodes" is +`P=0.163` under the birthday bound if used as a content address; `NEON` has no byte-table +primitive in `ndarray::simd`, blocking any RS scheme substrate-first.)* + +--- + +## VERDICT + +**On the phase/comma coefficient schedule: DELETE.** Zero of six admissible conditions are +both non-trivial and satisfied. Rank preservation, MDS preservation, and +reconstruct-from-locus are passed identically by *every* control including "no modulation" — +they are properties of permutation-equivalence and of RS itself, not of the schedule. The +one non-trivial condition, reducing cross-level syndrome collisions, is **impossible** under +a permutation of evaluation points, because permutation leaves the achievable syndrome set +invariant; achieving it needs disjoint GRS multiplier cosets, which is known, is +capacity-bounded at `⌊(q−1)/n⌋` levels, and is dominated by a one-byte level tag or by the +parity's own address. Cost fails too: level-dependent coefficients mean level-dependent +split-tables and table churn where fixed coefficients stay resident. **Salvage the coprime +stride as a declustering interleaver** (`curve_ruler.rs` already provides it) — that is where +an incommensurate schedule does real work, it defuses the phase-aligned failure mode, and it +is known engineering rather than a claim. + +**On scheme selection: the boring answer wins and the program should adopt it before +researching anything else.** Per-chunk digest **bound to (locus, version)** + flat MDS RS +(S6) strictly dominates every alternative on detection, localization, wrong-slot diagnosis, +and stale diagnosis; it is the only scheme in the table that can *positively diagnose* the +two failure modes this architecture actually generates; and it converts errors to erasures, +doubling correction capacity for free. Cascade parity and product codes buy repair +**bandwidth**, at a distance cost that is a theorem (Gopalan) and an I9 exposure that Morton +ordering actively creates. **Adopt them only after X-C2-4 shows a measured repair-bytes win +that survives declustering.** + +**On the product-code localization hypothesis: DISPROVED.** Row/column syndromes match a +per-chunk hash at multiplicity 1 and lose at multiplicity ≥ 2 to four-corner ambiguity; the +rectangle pattern at weight `d₁d₂` is uncorrectable where equal-overhead MDS RS is not. The +patent literature that defines product codes already prescribes per-symbol checksums for +exactly this reason. + +**On the stale-vs-corrupt question: no syndrome can do it, and the system already stores the +answer.** The syndrome is a function of the error alone and cannot separate "a previous +codeword's symbol" from "an arbitrary symbol". Every temporal distinction in this design +(`QueryReference`, `CONTEMPORARY`/`ANACHRONISTIC`/`SPOILER`, the HLC tick) is made by +comparing a *stored version coordinate*. Put the version in the chunk and compare — the ZFS +answer, with patents. The one interesting residue is the delta-parity discriminator (SURPRISES #3), +which is triage-grade, not authentication-grade, and is gated on X-C2-5. + +**Blocking, and cheaper than any of the above: fix the four live false-accept paths and run +X-C2-6 first.** `verify_ecc` cannot fail. 1472 of 2048 sealed bytes are unsealed. Four leaf +slots are duplicates. `xor_diff` returns a type-correct lie. `batch_hash` seals the inputs +and not the artifact. `AuditMerkleRoot` is a length-extendable FNV chain sold as +tamper-evidence and unlinkability. And no RS scheme can be selected at all until +`ndarray::simd` grows a NEON byte-table primitive, because the workspace's own SIMD +invariant forbids the consumer-side workaround. **A research programme that designs a new +seal while `verify_ecc` returns `Some` on every input has mis-ordered its work.** diff --git a/docs/lotus/rp-seal-v1/C3.md b/docs/lotus/rp-seal-v1/C3.md new file mode 100644 index 00000000..ab49cdb3 --- /dev/null +++ b/docs/lotus/rp-seal-v1/C3.md @@ -0,0 +1,649 @@ +# C3 — Coding-Theory Scout: Constraints Cheat-Sheet for Bidirectional Erasure-Coded Seals + +**Cell:** C3 (Domain C — erasure coding / seals; role: coding-theory scout) +**Posture:** research only. No implementation attempted. This report extracts the +theorems and impossibility results that any candidate erasure/locality scheme in +the wider program (built by other cells) must obey, and grounds them against the +actual source of the three sibling repos. + +--- + +## SOURCE ARCHAEOLOGY + +**Headline finding: no erasure-coding, Reed-Solomon, or Galois-field(>2) machinery +exists anywhere in `lance-graph`, `ndarray`, or `OGAR` today.** "Bidirectional +erasure-coded seals" is a hypothesis to be tested against an architecture that +currently has no coding-theoretic layer at all. This matters for scope: every +constraint below is a constraint on a *design that does not yet exist*, not a +audit of a live implementation. + +**1. "Seal" in the actual codebase means a WAL/version commit, not an erasure +code.** `crates/lance-graph-planner/src/persist_sink.rs` defines the term +precisely: + +> "cycle — the 64k-row frame. Its seal is ONE WAL write → one version." +> "Every thought in an open cycle reads only the sealed predecessor +> (`frame.base_version = Vn`); the seal publishes `Vn+1` once." + +This is the real temporal substrate a "bidirectional erasure-coded seal" would +sit on top of: `WalSink::scan_sealed` excludes unsealed cycles, so the +`A_last < W_write <= T_now` shorthand in the task prompt is a reasonable +*informal* gloss on `read Vn / write Vn+1` — but two things it get wrong if taken +literally as implemented: (a) there is **no single global `A_last`** — I traced +this to OGAR's temporal doc below, where the awareness cutoff (`knowable_from`) +is stamped **per class**, not globally; (b) `batch_writer.rs`'s own header +comment (2026-08-04 self-correction, `⊘ CORRECTED`) documents that its +`cast()`/deinterlace machinery has **zero production call sites** today — the +temporal-window plumbing this program wants to hang an erasure code off of is +itself only partially wired, per the repo's own admission. + +**2. OGAR's `docs/TEMPORAL-TIME-TRAVEL.md` at the pinned commit +(`386a6fd8`) confirms and refines the temporal model, and it is NOT a single +`T_now × A_last × W_write` tuple as shorthand'd in the task.** The real +primitives: `QueryReference { ref_version: u64, mode: STRICT|AWARE|RETRO, rung: +u8 }` (a planner-layer query annotation, not storage); `checkout_version(V_ref)` +as the free primitive (time travel costs nothing extra — it is just a pinned +Lance version read); `knowable_from` populated **at class-registration time**, +one `u64` column per class, not one global watermark; and a `(server_id, +local_lance_version, hlc_tick)` HLC tuple proposed (not yet built) for +cross-server causal ordering. **Consequence for any erasure scheme:** if +repair/seal windows are scoped by "how far behind is a reader allowed to be," +that scoping is naturally **per-class**, not global — a parity group spanning +rows of different classes would need to explicitly reconcile potentially +different `knowable_from` cutoffs, which is a real design constraint, not a +detail. + +**3. A genuine "docstring claims Reed-Solomon, code implements XOR parity" gap +already exists in this workspace — worth citing as a concrete instance of the +exact failure mode this program should guard against.** +`crates/lance-graph-cognitive/src/fabric/udp_transport.rs` and +`firefly_frame.rs` document a frame layout with: + +``` +0x80-0x9F: ECC (Reed-Solomon for error correction) +... +/// ECC: Hamming error correction (words 16-19, 224 bits used) +``` + +but the actual implementation (`firefly_frame.rs::compute_ecc`) is: + +```rust +// Simplified ECC: XOR-based parity +let mut ecc = [0u64; 4]; +for (i, &word) in data.iter().enumerate() { + ecc[i % 4] ^= word; +} +// ... a single bit-parity checksum per lane +``` + +This is **not** Reed-Solomon (no polynomial evaluation, no Vandermonde/generator +matrix, no distance-3+ guarantee) and **not** a real Hamming code (no syndrome +table, no single-error-correction decode shown) — it is a 4-way XOR fold plus a +parity bit, which is distance-2 at best and cannot correct anything, only +detect. This is exactly the workspace's own documented anti-pattern (its +`CLAUDE.md` files elsewhere catalogue "a doc-comment claim is not a behaviour" +as a house-wide, repeatedly-rediscovered failure mode). **Recommendation for +the consolidator:** any cell's design doc that names "Reed-Solomon" or "MDS" for +a seal/repair mechanism must be checked against an actual generator/parity-check +matrix with the Singleton property verified, not asserted by name. + +**4. `GF(2)` appears repeatedly across the workspace, but it is always the +trivial XOR field for VSA binding — never a coding-theoretic GF(2^m).** +`crates/holograph/src/graphblas/semiring.rs`, `crates/lance-graph/src/graph/ +blasgraph/{types,mod,semiring}.rs`, and `crates/lance-graph-contract/src/ +grammar/role_keys.rs` all use "GF(2) algebra" to mean bind = XOR, bundle = +majority-vote — a *cognitive* VSA operation, not an erasure code. This is +relevant because RAID5-style single-XOR parity (bind-with-XOR) tops out at +distance 2 (tolerates exactly one erasure); a scheme claiming distance-3+ +"erasure-coded" protection cannot be built from this substrate's existing +XOR primitives without moving to a genuine extension field or an array-code +construction (see EVENODD/RDP below) — a plain second XOR fold over the same +data does **not** buy a second erasure's worth of protection (the classic RAID6 +lesson: two independent XOR parities computed the same way are linearly +dependent and protect nothing beyond one erasure; RDP/EVENODD's second parity +is deliberately **diagonal**, not a repeated row-XOR). + +**5. Lance 9.0.0 (both the exact-pinned column and current upstream) already +owns a row-id remap / physical-relocation mechanism that any erasure scheme +must compose with, not duplicate.** `rust/lance/src/dataset/rowids.rs` + +`lance_table::rowids::{RowIdSequence, FragmentRowIdIndex}` + +`lance_core::utils::deletion::DeletionVector` give Lance's own stable +logical-row-id ↔ physical-fragment-offset indirection, used by `optimize.rs` +(compaction) and `cleanup.rs`. This is precisely the "index-remapped rewrite" +layer the research program's metrics list names (`index remap bytes/CPU`) — +it is native to Lance, pre-existing, and not itself erasure-coded. A candidate +scheme's "physical canonicalization" work is additive on top of this, and its +cost should be measured *net of* what Lance's own remap already costs, not +conflated with it. + +**6. `crates/helix/src/curve_ruler.rs`'s `CurveRuler` (stride-4-over-17, +`gcd(4,17)=1`) is a genuine, already-implemented **coprime-integer permutation +generator** over a small residue field.** Its doc comment: "gcd(4,17)=1 → full +permutation of all 17 residues... unlike the banned Fibonacci-mod-17 walk." +This is directly relevant to the cross-domain question "can one SFC order +minimize both query scatter and RS repair scatter?" and to Tamo-Barg-style LRC +construction (§ MECHANISM below), which explicitly requires the length-n +support to be **partitioned into disjoint groups of size (r+1)** on which a +"good polynomial" is constant — a coprime stride walk over a small modulus is +exactly the right *shape* of generator for producing such balanced partitions +deterministically from an address, though 17 is not of the form m(r+1) for +typical small r and would need to be chosen to match a partition size, not +reused as-is. + +**Net read of the source map:** the architecture as it exists today gives (a) a +version-sealed WAL-cycle temporal substrate with per-class awareness cutoffs, +(b) a native Lance row-remap/compaction layer, (c) a coprime cyclic-permutation +primitive, and (d) zero erasure coding. The coding-theory constraints below are +therefore constraints on a **clean-slate design**, evaluated against literature +that already solved (and bounded) very similar problems for RAID/cloud storage. + +--- + +## MECHANISM + +The three requested papers plus the wider search converge on a small number of +composable primitives. Stated compactly (all variable names as in the source +papers, translated to a common notation at the end): + +**(A) The classical MDS Singleton bound.** An `[n,k]` code has `d ≤ n-k+1`; +equality defines MDS. Reed-Solomon (evaluation of a degree-`(k-1)` polynomial at +`n` distinct field points) achieves it for any `n ≤ q` (field size `q`). + +**(B) Locally Repairable Codes (LRC), scalar/vector-general Singleton-type bound +(Papailiopoulos & Dimakis 2012, arXiv:1206.3804, Theorem 1).** For an +`(n, r, d, M, α)`-LRC (file size `M` bits, `n` coded symbols of `α` bits each, +each symbol locally repairable from `≤ r` others): + +``` +d ≤ n − ⌈M/α⌉ − ⌈M/(rα)⌉ + 2 +``` + +This is **universal**: holds for linear and nonlinear codes, scalar and vector +alphabets. Setting `M=k, α=1` recovers the earlier Gopalan et al. scalar bound +`d ≤ n − k − ⌈k/r⌉ + 2`. **Achievability** (their Theorem 2): when `(r+1) | n` +and `r ≤ n−d`, LRCs meeting this bound exist over a sufficiently large field, via +a locality-aware flow-graph/multicast-capacity argument (a specialization of the +Dimakis et al. cut-set network-coding technique, see (D)). **Explicit +construction** (their §5): partition the file into `r` sub-blocks, RS-encode each +sub-block independently to length `n`, place `r` blocks + one XOR-parity block +per repair group of size `r+1` — repair of any single symbol is a plain XOR of +`r` others. Rate loss vs. an `(n,k)`-MDS code is exactly `1/r · k/n`. + +**(C) Generalized `(r,δ)` Singleton bound (survey 1806.04437, Theorem VII.1, +citing Prakash-Kamath-Lalitha-Kumar and the original Gopalan et al./Papailiopoulos- +Dimakis δ=2 case).** For an `[n,k]` code with `(r,δ)` information-symbol +locality (each information symbol sits in a local code of dimension `≤ r` and +minimum distance `≥ δ`, so the local code alone tolerates `δ-1` erasures): + +``` +d_min ≤ (n − k + 1) − (⌈k/r⌉ − 1)(δ − 1) +``` + +`δ=2` recovers (B)'s scalar case. **Optimal, explicit constructions exist** +(Tamo-Barg, survey §VII.B.2, Theorem VII.3): partition the `n` evaluation +points into `m = n/(r+1)` disjoint groups `A_1..A_m` of size `r+1`; pick a +"**good polynomial**" `g(x)` — one that is *constant on each `A_i`* and has +degree exactly `r+1` — and build the message polynomial as +`f(x) = Σ a_ij · [g(x)]^j · x^i` (a "layered" polynomial in the good-polynomial +variable). Every local group's restriction is then itself effectively a +low-degree (hence MDS-checkable) sub-code, and the construction meets Theorem +VII.1 with equality. **This is the formal object the task's "coefficient +re-orientation across levels" question is really about** — see the dedicated +subsection below. + +**(D) Cut-set bound for regenerating codes (Dimakis, Godfrey, Wu, Wainwright, +Ramchandran 2010 — reproduced in the survey, §II.A, eq. (1)).** For a +functional-repair regenerating code with parameters `((n,k,d),(α,β),B)` — `n` +total nodes, any `k` reconstruct the file, node repair contacts `d` helpers and +downloads `β` per helper, each node stores `α` — the achievable file size is +bounded by a min-cut over an information-flow graph: + +``` +B ≤ Σ_{i=0}^{k-1} min{α, (d−i)β} +``` + +Achievable (for finite repair horizons, and even infinite via Wu's later +strengthening) with linear network coding over a large field. Two extremal +points: **MSR** (`α = B/k`, the minimum possible storage, i.e. MDS-tight +storage) and **MBR** (`β` minimized, storage slightly above MDS). This bound is +the *ceiling*; whether an EXACT-repair (not merely functional-repair) code can +achieve it is a separate, harder question — which is exactly what paper (E) +resolves for MDS/RS codes specifically. + +**(E) Repair bandwidth of Reed-Solomon codes under exact linear repair +(Guruswami & Wootters 2015, arXiv:1509.04764).** +- **Achievability (their Theorem 1 / Corollary 2):** for `RS(F,k)` with rate + `1-1/|B|` using the *entire field* `F` as evaluation points (`n=|F|`), there is + a linear exact-repair scheme over a subfield `B` with bandwidth exactly `n-1` + sub-symbols of `B`, i.e. `(n-1)·log2(1/ε)` bits where `ε = |B|^{-1}` is the + redundancy fraction. The trick is a **trace-polynomial characterization**: a + linear repair scheme for a coordinate `α*` corresponds to finding `t` dual + codewords whose evaluations span a *low*-dimensional subspace everywhere + except at `α*`, where they span a *high*-dimensional one (their Theorem 4 — a + full "if and only if" characterization of linear repair schemes for any MDS + code via its dual code). +- **Matching lower bound for ANY MDS code with a linear repair scheme (their + Theorem 3, the load-bearing result for a cheat-sheet):** + +``` +b ≥ (n−1)·log_|B|( (n−1)/(n−k) ) [subsymbols of B] +b·log2|B| ≥ (n−1)·log2( (1/ε)·(1−1/n) ) [bits] +``` + + Their Remark 6 gives a one-line derivation identical in spirit to the + Dimakis cut-set bound (D): for any repair scheme, some `n-k`-size subset of + helpers must jointly supply enough information to distinguish all possible + values, so the *average* download over a random helper set is at least + `t·(n-1)/(n-k)`. +- **Naive RS repair, for contrast:** `k·log2(n)` bits (download `k` full + symbols) — the Guruswami-Wootters scheme is a genuine, often order-of- + magnitude, improvement, **without changing `n,k,d` at all** — a strict + algorithmic win layered on the *same* MDS code. +- Subsequent work cited within: Ye & Barg extend the same trace-polynomial + framework to the high-sub-packetization ("large-t") regime and *match the + cut-set bound (D) exactly* there — so RS codes are competitive with + purpose-built MSR array codes in **both** the low- and high-sub-packetization + regimes, once repaired correctly. + +**(F) Hierarchical/multi-tier LRC bound (survey §VIII.E, Theorem VIII.5, +Sasidharan-Agarwal-Kumar).** For a 2-tier hierarchy — every symbol sits in a +tight "local" code (dimension ≤ `r1`, distance ≥ `δ1`) that is itself nested +inside a wider "middle" code (dimension ≤ `r2`, distance ≥ `δ2`, `r1 ≤ r2`): + +``` +d ≤ n − k + 1 − (⌈k/r2⌉ − 1)(δ2 − 1) − (⌈k/r1⌉ − 1)(δ1 − δ2) +``` + +This is the **direct, load-bearing formal tool for the Morton-cascade / +2-D-hierarchy question** — see below. + +**(G) Availability / product codes (survey §VIII.B.2).** The `[(r+1)^t, r^t]` +product code in `t` dimensions is an `(n=(r+1)^t, k=r^t, r, t)` **availability** +code: every symbol has `t` *disjoint* size-`(r+1)` recovery sets (one per axis), +rate `(r/(r+1))^t`. Wang et al.'s alternative construction achieves rate +`r/(r+t)` with block length `C(r+t,t)` (strictly smaller for the same `(r,t)`) +via a combinatorial parity-check design on `t`- and `(t-1)`-subsets of an +`(r+t)`-element ground set, at the cost of losing the clean per-axis +decomposability of the plain product code. + +**(H) Classical MDS array codes — EVENODD, RDP, RAID-6 P/Q (web search, not in +the three seed papers; see PRIOR ART).** `p+2`-column (EVENODD, `p` prime) or +`p+1`-column (RDP, `p` prime) constructions achieve MDS distance-3 over `GF(2)` +using **horizontal row-parity plus diagonal parity** — critically, the two +parities are **not symmetric under axis transpose**: EVENODD's diagonal parity +needs an "adjusting factor S" computed along one specific diagonal direction; +naively computing "row XOR" then "column XOR" of the same data does **not** +give an MDS distance-3 code (two linearly-dependent single-parity computations +protect against only one erasure, the classic RAID-4-stacked-twice failure). +RDP improves on EVENODD by avoiding the adjusting-factor computation but "fails +to attain the optimal update complexity" per source. Both operate over `GF(2)` +(XOR-only), unlike RS which needs `GF(2^8)` or larger to get an MDS distance +beyond 2 for `n > q+1` symbol alphabets. + +**(I) Piggybacking (survey §V, [Rashmi-Shah-Ramchandran], cited via the +Hitchhiker practical system).** Start from `α` independent codewords of the +*same* `[n,k]` MDS code; add small linear combinations ("piggybacks") of one +codeword's symbols onto *another's coded (parity) symbols*, chosen so +decodability of each codeword individually is unaffected. **This changes +nothing about `(n,k,d)` or sub-packetization — it is a strictly additive +repair-bandwidth optimization on top of an already-MDS code**, reported to save +25-50% repair bandwidth/disk-read. This is the cleanest existing precedent for +"can a scheme be layered on top of an unmodified MDS/seal mechanism without +touching its distance guarantee." + +**(J) Practical GF SIMD kernels (ISA-L, Jerasure — web search).** Intel ISA-L +implements general `[n,k]` RS-type erasure coding in `GF(2^8)` with primitive +polynomial `x^8+x^4+x^3+x^2+1` (0x1D), SIMD-accelerated Galois-field multiply; +Jerasure is the portable-C precedent with the same field convention; both are +cited directly in the survey's practice section (§XII) as the acceleration +substrate under Clay code / Butterfly / PM-RBT / Beehive production systems. +**Constraint for a candidate design:** if the payload alphabet ever exceeds 256 +distinct symbols in a single stripe (`n > 255` for a `GF(2^8)`-based RS code), +either an extension field (`GF(2^16)`, etc.) or a Cauchy/array-code alternative +is required — this workspace's native V3 rows are byte-addressed (12-byte +content-blind payload, `6×(u8:u8)` rails), which is naturally compatible with +`GF(2^8)` symbol size, but any parity group spanning more than 255 "columns" +(e.g. striping across an entire 4096-field cycle rather than one row) would +need to be re-derived over a larger field or restructured as an array code. + +--- + +## Constraints Cheat-Sheet (for builders/adversary) + +| # | Bound / theorem | Formula | Binds | +|---|---|---|---| +| 1 | Singleton (MDS) | `d ≤ n−k+1` | any linear code | +| 2 | LRC, universal vector form (Papailiopoulos-Dimakis Thm 1) | `d ≤ n − ⌈M/α⌉ − ⌈M/(rα)⌉ + 2` | any code, linear/nonlinear, scalar/vector, `δ=2` implicit | +| 3 | LRC, `(r,δ)` generalized Singleton (survey Thm VII.1) | `d ≤ (n−k+1) − (⌈k/r⌉−1)(δ−1)` | any `[n,k]` linear code with `(r,δ)` locality | +| 4 | Hierarchical (2-tier) LRC (survey Thm VIII.5) | `d ≤ n−k+1 − (⌈k/r2⌉−1)(δ2−1) − (⌈k/r1⌉−1)(δ1−δ2)`, `r1≤r2`, `δ1≥δ2` | nested local-inside-middle code | +| 5 | Cooperative-recovery LRC (survey, [153]) | `d_min(n,k,r,t) ≤ n−k+1 − t·⌊(k−t)/r⌋` | `t`-simultaneous-erasure local repair | +| 6 | Cut-set bound, regenerating codes (Dimakis et al.) | `B ≤ Σ_{i=0}^{k−1} min{α,(d−i)β}` | any `(n,k,d),(α,β)` regenerating code, functional repair | +| 7 | MDS linear exact-repair lower bound (Guruswami-Wootters Thm 3) | `b ≥ (n−1)·log_{\|B\|}((n−1)/(n−k))` subsymbols | any linear exact-repair scheme for any MDS code | +| 8 | MDS linear exact-repair achievability (GW Thm 1 / Cor 2) | `b = n−1` subsymbols over `B` (rate `1−1/\|B\|`, `n=\|F\|`), TIGHT vs #7 | RS with evaluation points = whole field | +| 9 | Availability/product code | `n=(r+1)^t, k=r^t`, rate `(r/(r+1))^t`, `t` disjoint recovery sets | codes needing `t`-fold parallel local repair | +| 10 | Availability, min-distance (survey [146]) | `d_min(n,k,r,t) ≤ n−k+2−⌈(k+1)/r⌉·` (⌈·⌉-family bound, tightened in [111]/[144] for large `t`) | `(n,k,r,t)` availability codes | + +**Rate/locality/distance is a three-way trade — no free axis.** Every row above +subtracts *more* from `n-k+1` as `r` (or `r1,r2`) shrinks or as more tiers/ +disjoint-recovery-sets are demanded. A design that claims low locality (small +`r`), high rate (`k/n → 1`), *and* MDS-level distance simultaneously is +claiming to beat one of rows 2-6 and should be treated as almost certainly +wrong until the exact `(n,r,δ,k)` are plugged in and checked. + +--- + +## The precise conditions under which "coefficient re-orientation across levels" preserves or breaks MDS + +This is the sharpest question the task asked, and the literature gives a +mechanically checkable answer, not a vibe. + +**A code's Singleton-type bound is a function of `(n,k,r,δ)` (or the +hierarchical `(n,k,r1,δ1,r2,δ2)` tuple) alone — it says nothing about *which* +axis of a multi-dimensional layout carries the parity.** So "re-orientation" +(e.g. reading the same 12-byte V3 payload once as `6×(u8:u8)` rails and once as +`4×(u8:u8:u8)` SPO triples — exactly the pattern this workspace's own +`EdgeCodecFlavor::{CoarseOnly, CoarseResidue, Pq32x4}` already does for +*content interpretation*, per `lance-graph`'s CANON section) is **MDS-neutral +if and only if every one of the following holds simultaneously**: + +1. **Fixed support partition.** Tamo-Barg's construction (mechanism C above) + requires the `n` evaluation points to be split into disjoint groups `A_i` of + size `r+1` on which a *single* good polynomial `g(x)` is constant. If + "re-orientation" changes which physical symbols belong to which repair + group (e.g. row-groups become column-groups), that is a **different** + partition and needs a **new** good polynomial satisfying the same degree + constraint on the **new** groups — this is not automatic and must be + re-derived, not merely relabeled. + +2. **Every size-`k` (or size-`r`) submatrix stays full rank.** The MDS/LRC + property reduces to a linear-algebra fact about the generator (or parity- + check) matrix: no `k` columns are linearly dependent (equivalently, every + `k×k` submatrix of the generator matrix is invertible — a Vandermonde-style + condition, guaranteed by distinct evaluation points over a large-enough + field). A re-orientation that reuses the *same physical bytes* under a + *different* `(n,k)` factorization (say, once as `n=6,k=4` rails and once as + `n=4,k=3` SPO triples) is asking the same byte matrix to be simultaneously + MDS-optimal for **two different rank conditions**. In general these are + **not** the same matrix up to coordinate permutation, so a single physical + layout cannot be independently MDS-optimal under both readings unless one + is explicitly derived as a **sub-code or super-code** of the other (e.g. one + reading's parity symbols are a strict linear image of the other's) — this + must be checked, not assumed, per candidate design. + +3. **Isotropy is NOT free for array codes.** EVENODD/RDP (mechanism H) are the + concrete counter-example that a re-orientation "obviously" preserving + symmetry can silently break MDS-ness: their diagonal parity is direction- + dependent (EVENODD's "adjusting factor S" is computed along one specific + diagonal), so transposing rows↔columns in an array code is **not** + automatically distance-preserving — it must be re-verified as its own + construction, exactly like re-deriving RDP from EVENODD was a nontrivial, + separate contribution in the literature, not a relabeling. + +4. **Nesting (hierarchy) always costs distance, monotonically, and the cost is + formula (row 4 above).** If "levels" means literally nesting a tighter local + code inside a wider middle code (the Morton-cascade / HEEL-HIP-TWIG framing), + the correct bound is **not** the flat single-tier bound reused at each + level — it is the strictly tighter Theorem VIII.5 expression, which + subtracts an *additional* term `(⌈k/r1⌉−1)(δ1−δ2)` beyond the single-tier + middle-code cost. **A hierarchical design that claims the *same* worst-case + `d` as a flat `(r2,δ2)` code at every level is wrong by construction** — + the extra tier necessarily costs distance in exchange for cheaper + single-erasure (tier-1) repair. This is quantifiable and should be the + first number any hierarchical/cascade proposal reports. + +5. **Additive (piggyback-style) re-orientation is the one safe pattern.** + Mechanism (I) is the literature's existence proof that a repair-bandwidth + optimization *can* be layered on an unmodified `(n,k,d)`-MDS code with zero + distance cost — but only because it is explicitly **additive on the parity + symbols of a fixed code**, never a re-partition of which symbols are + systematic vs. parity, and never a change to `n` or `k`. If the candidate + architecture's "bidirectional" seal wants a free lunch, this is the shape + it must take: extra small linear combinations added to existing parity, + not a re-carving of the support set. + +**Summary answer:** coefficient re-orientation preserves MDS exactly when it +is (a) a coordinate relabeling of the *same* fixed-rank generator matrix +(trivially safe), or (b) an additive overlay on a fixed code's parity symbols +in the piggybacking sense (safe, proven, 25-50% bandwidth wins available for +free); it breaks MDS (or silently degrades the achieved distance below what +was claimed) whenever it changes which symbols belong to which repair group +without re-deriving the good-polynomial/generator-matrix for the new +partition, or whenever it treats a genuinely anisotropic array-code +construction (row/diagonal parity) as if it were direction-symmetric. + +--- + +## EXPECTED BENEFIT + +If a design in this program adopts (E)'s trace-repair scheme or (I)'s +piggybacking **without changing the underlying MDS `(n,k,d)`**, the literature +supports large, real wins with **zero** distance cost: +- Repair bandwidth for high-rate RS: naive `k·log2(n)` bits → `(n-1)·log2(1/ε)` + bits (order-of-magnitude for high rate, small `ε`), *provably optimal* + among linear schemes (row 7/8 match). +- Piggybacking: 25-50% repair-bandwidth/disk-read reduction on unmodified MDS, + in production (Hitchhiker/HDFS). +- If locality (LRC) is adopted instead of pure MDS repair-bandwidth + optimization, Azure's own production numbers (survey §XII) are the + reference: storage overhead 1.29 vs. 1.5 for comparable repair degree — + a real, measured, non-hypothetical number this program's baseline can be + checked against. + +## EXPECTED FAILURE + +- Any claim of "hierarchical erasure protection across the 2-D/Morton cascade + with no distance cost" is very likely to be **row 4's bound violated in + practice**, i.e. an implicit, unverified assumption that nesting is free. +- Any "we re-oriented the same 12-byte payload as both an LRC and a + product-code axis" claim needs an explicit rank check (constraint #2 above); + absent that check it is presumptively **broken**, not merely unproven. +- A "GF(2) XOR-based Reed-Solomon" claim (per the workspace's own + `firefly_frame.rs` precedent) is near-certain to be a **distance-2, not + distance-`k`, mechanism** mislabeled — this specific failure mode has + *already happened once* in this exact codebase. +- Reusing a single stride/coprime-walk generator (e.g. `CurveRuler`, + stride-4-mod-17) as *both* the SFC ordering for query locality *and* the + partition generator for LRC repair groups is plausible in principle (both + want balanced, deterministic, address-derived groupings) but the literature + gives no existing proof that one ordering is simultaneously optimal for + both objectives — treat as NOVEL-CANDIDATE requiring its own falsifier (see + SURPRISES). + +--- + +## EXPERIMENT + +Five pre-registered, falsifiable probes a builder/adversary cell should be able +to run once a candidate scheme exists. Each ties to the program's metric list. + +**EXP-C3-1 — Rank/MDS-preservation check under re-orientation.** +*Metric:* boolean pass/fail (every size-`k` submatrix of the generator matrix, +under BOTH readings of the layout, is full rank) + minimum distance actually +achieved (measured by exhaustive erasure-pattern search up to the design's `d`, +on a small `n` instance). +*Baseline:* the flat, single-tier `(n,k)` MDS code with no re-orientation. +*Control:* a deliberately-broken re-orientation (row/column-transposed EVENODD +diagonal parity, per constraint #3) as the known-negative case. +*Kill condition:* if the "preserved" re-orientation's measured minimum +distance is `< d` claimed by the design doc, the re-orientation claim is +falsified — full stop, before any performance number is trusted. + +**EXP-C3-2 — Hierarchical-cost accounting.** +*Metric:* measured minimum distance of the actual 2-tier (or `t`-tier) +construction vs. (a) the flat single-tier bound at the middle-tier `(r2,δ2)` +and (b) Theorem VIII.5's tighter bound. +*Baseline:* single-tier `(r2,δ2)` LRC at the same rate. +*Control:* single-tier `(r1,δ1)` LRC at the same rate (the other extreme). +*Kill condition:* if the hierarchical scheme's distance does *not* sit exactly +at Theorem VIII.5's bound (either strictly better = bug in the check, or +strictly worse = suboptimal construction), record which and stop claiming +optimality until reconciled. + +**EXP-C3-3 — Repair bandwidth vs. the Guruswami-Wootters lower bound.** +*Metric:* measured repair bandwidth (bytes moved during a single-symbol +repair) vs. row 7's closed-form lower bound for the same `(n,k,B)`. +*Baseline:* naive full-row-download repair (`k` symbols). +*Control:* the GW trace-repair scheme itself (row 8), implemented as a +reference, if the design's evaluation points are the whole field. +*Kill condition:* any claimed scheme beating row 7's bound for a genuinely +*linear* repair scheme is either non-linear (must be disclosed) or wrong. + +**EXP-C3-4 — Product-code vs. hand-rolled 2-D parity, availability.** +*Metric:* number of disjoint recovery sets actually realized per symbol +(`t`, per row 9), and rate achieved, vs. the `[(r+1)^t, r^t]` product-code +closed form and the Wang et al. improved-rate construction. +*Baseline:* plain row+column XOR parity over the 2-D grid (the "obvious" +un-optimized design). +*Control:* the exact Wang et al. parity-check-matrix construction (smaller +`n` for the same `(r,t)`). +*Kill condition:* if the candidate's rate is *worse* than the plain product +code's `(r/(r+1))^t` for the same `(n,k,r,t)`, the "smarter" 2-D design has +failed to beat the naive baseline and should be dropped. + +**EXP-C3-5 — SFC-order dual-use falsifier (query locality vs. repair +locality).** *Metric:* two independent scatter metrics computed over the +*same* candidate ordering (e.g. `CurveRuler`-style stride walk, or a Morton +order): (a) query-locality scatter — number of distinct physical +pages/fragments touched by a representative neighborhood query; (b) +repair-locality scatter — number of distinct physical pages/fragments touched +by a representative single-symbol repair under the erasure scheme built on top +of that ordering. +*Baseline:* two SEPARATE orderings, one hand-tuned for (a) alone and one for +(b) alone. +*Control:* a random (non-structured) ordering, as the negative control for +both metrics. +*Kill condition:* if the single shared ordering's combined (a)+(b) scatter is +not simultaneously within, say, 20% of each metric's own best single-purpose +ordering, the "one SFC order serves both" hypothesis (a NOVEL-CANDIDATE +cross-domain question per the task) is disproved for that ordering family, +and the two concerns should be treated as requiring independently chosen +orderings. + +--- + +## PRIOR ART + +Searches actually run (query, tool, and what came back): + +1. **alphaXiv `get_paper_content`** on `arxiv.org/abs/1206.3804` (Locally + Repairable Codes, Papailiopoulos & Dimakis) — full text retrieved (3,018 + lines), read in full for abstract, Theorem 1 (universal Singleton-type + bound), Theorem 2 (achievability via locality-aware flow-graph), and the + explicit MDS-pre-coding + XOR construction (§5). + +2. **alphaXiv `get_paper_content`** on `arxiv.org/abs/1509.04764` (Repairing + Reed-Solomon Codes, Guruswami & Wootters) — full text retrieved (2,722 + lines), read in full for abstract, Theorem 1/Corollary 2 (achievability), + Theorem 3 (matching lower bound for linear repair of any MDS code, with the + cut-set-bound-style proof in Remark 6), and Theorem 4 (the dual-code + characterization of all linear repair schemes). + +3. **alphaXiv `get_paper_content`** on `arxiv.org/abs/1806.04437` (Erasure + Coding for Distributed Storage: An Overview, Balaji, Krishnan, Vajha, + Ramkumar, Sasidharan, Kumar) — full text retrieved (5,964 lines); targeted + reads of: §II.A cut-set bound derivation (eq. 1), §VII.A generalized + `(r,δ)` Singleton bound (Theorem VII.1) and §VII.B.2 Tamo-Barg construction + (Theorem VII.3, "good polynomials"), §VIII.B.2 product/availability codes + (Wang et al. construction), §VIII.E hierarchical-locality bound (Theorem + VIII.5), §V piggybacking framework definition, and §XII "Codes in Practice" + (Azure LRC production numbers, Xorbas/Hitchhiker/Clay/Butterfly/PM-RBT, + ISA-L usage). + +4. **WebSearch**: `"EVENODD RDP RAID-6 array code P Q parity construction + MDS"` — confirmed EVENODD (prime `p`, `p+2` columns, diagonal-parity + adjusting factor `S`) and RDP (prime `p`, `p+1` columns, NetApp 2DFT array, + avoids the adjusting factor but loses optimal update complexity) as the + two canonical horizontal MDS array codes for RAID-6; also surfaced P-Code, + MDR codes, C-Codes (cyclic lowest-density) as further array-code variants + not pursued further given scope. + +5. **WebSearch**: `"ISA-L Jerasure Reed-Solomon GF(2^8) SIMD erasure coding + kernel Intel"` — confirmed ISA-L's field convention (`GF(2^8)`, primitive + polynomial `0x1D`), SIMD-accelerated Galois multiply, and its role as the + Jerasure-comparable practical kernel library cited in the survey's practice + section. + +6. **Source grep** (not literature, but disciplined "is this novel or already + known internally" search per the task's mandate) across + `/home/user/lance-graph`, `/home/user/ndarray`, `/home/user/OGAR` for + `reed.solomon|erasure|galois|gf256|gf\(2` — surfaced the `firefly_frame.rs` + doc/code mismatch and confirmed no genuine coding-theory implementation + exists in either repo (§ SOURCE ARCHAEOLOGY above). + +**Not retrieved / explicitly out of scope for this cell** (per task +instructions): rustynum (superseded-donor ruling), `crates/symbiont` +(deprecated), and any file under `docs/lotus/**`, other cells' `.md` outputs, +or `.claude/board/**` / `.claude/plans/**` (independence rule). + +--- + +## SURPRISES + +1. **KNOWN.** The workspace's own `firefly_frame.rs` docstring claims + "Reed-Solomon" / "Hamming" ECC while the implementation is 4-way XOR-fold + parity. This is not a coding-theory surprise (XOR-parity-mislabeled-as-RS + is a well-known real-world anti-pattern) but it *is* a surprising, + concrete, in-repo instance of exactly the risk this whole research program + should be defending against, found by simple archaeology rather than by + analyzing any candidate design. + +2. **TRANSFER.** The Tamo-Barg "good polynomial" construction (constant on + disjoint evaluation-point groups, used to build optimal `(r,δ)`-LRCs) is + structurally the same shape of object as this workspace's own `CurveRuler` + coprime-stride generator (constant-modulus, deterministic partition of a + residue class into balanced groups) — a known LRC-construction technique + *transferred* onto an unrelated primitive already present in the codebase + for a different purpose (φ-spiral curve addressing). Whether `CurveRuler`'s + specific modulus (17) and stride (4) can be repurposed as a real good- + polynomial partition generator for some `(r+1) | 17`-shaped code is a real, + checkable question, not yet checked. + +3. **NOVEL-CANDIDATE.** "Can one SFC/address order simultaneously minimize + query-neighborhood scatter *and* erasure-repair scatter?" — searched the + three seed papers and the survey; found no paper that poses or answers this + as a joint-optimization question (the LRC/regenerating-code literature + optimizes repair locality against *storage rate and distance*, never + against a *separate* query-access-pattern objective simultaneously). This + is a genuine candidate for new ground, not covered by classical coding + theory, and EXP-C3-5 above is the falsifier that should be run before it + is trusted. + +4. **DISPROVED (attractive but wrong).** The intuitive analogy "row parity + + column parity over the 2-D grid = a legitimate 2-erasure-tolerant MDS code, + the same way RAID-6 works" is **false** as stated — plain independent + row-XOR and column-XOR parities are *not* automatically MDS; RAID-6's real + second parity (EVENODD/RDP) is deliberately diagonal, computed with an + axis-dependent adjusting term, specifically *because* naive orthogonal-axis + XOR parity fails to be MDS beyond a single erasure for more than a handful + of disks. Any 2-D-hierarchy design in this program that reaches for "row + parity, column parity, done" for its base erasure layer is reaching for a + disproved shortcut. + +5. **KNOWN, but easy to get backwards.** Nesting local codes inside a wider + middle code (hierarchical locality) is *not* free — Theorem VIII.5 shows it + strictly *costs* distance relative to a flat middle-tier code, in exchange + for cheaper single-erasure repair. It would be an easy, attractive mistake + for a design doc to claim "the hierarchy gives us cheap local repair at no + distance cost" — the bound says the opposite, and by a term that is + exactly computable from `(k, r1, r2, δ1, δ2)`, so this is the one + over-claim in this space with a one-line, mechanical check. + +--- + +## VERDICT + +The erasure-coding side of the program has strong, directly-applicable +theoretical support for the "layer repair-bandwidth optimizations or locality +onto an unmodified MDS seal" direction (Guruswami-Wootters trace repair, +piggybacking — both proven, both zero-distance-cost, both with production +precedent), but **no existing theorem licenses a "hierarchical, re-oriented, +multi-axis erasure structure with no distance cost" claim** — every relevant +bound (generalized Singleton, hierarchical-locality Singleton, availability +min-distance) shows locality/hierarchy/availability is purchased with +distance, monotonically and exactly quantifiably, and the workspace's own +codebase already contains one real, checkable instance (`firefly_frame.rs`) +of an "ECC" claim that does not survive inspection. Any Domain-C design that +wants to nest repair tiers across the program's 2-D Morton hierarchy must +report its distance cost against Theorem VIII.5 explicitly, and any +"coefficient re-orientation" must pass the rank/partition checks in the +dedicated subsection above before its MDS claim is trusted. diff --git a/docs/lotus/rp-seal-v1/D1.md b/docs/lotus/rp-seal-v1/D1.md new file mode 100644 index 00000000..4b73c315 --- /dev/null +++ b/docs/lotus/rp-seal-v1/D1.md @@ -0,0 +1,208 @@ +# D1 — Temporal Deinterlacing: `temporal.rs` × OGAR Time-Travel × Lance `DatasetVersion` + +**Cell:** D1 (Domain D — temporal deinterlacing) · **Role:** Builder +**Scope discipline honored:** no code run/changed; only `git show`/`log`/`diff` and read-only file exploration; independent of other researchers' outputs; rustynum and `crates/symbiont` excluded per operator scope ruling. + +--- + +## SOURCE ARCHAEOLOGY + +### A. `lance-graph-planner/src/temporal.rs` (871 lines, fully read) + +The module is explicit that it implements **two distinct temporal layers** and states this itself in its module doc (lines 1–66): + +- **Layer 2 — "EPISTEMIC projection"** (`classify`, `deinterlace`, `QueryReference`, `EpistemicMode`, `TemporalStatus`) — answers *"what may this reader know?"*. This is the layer the task brief's `T_now × A_last × W_write` hypothesis targets. +- **Layer 1 — "CAUSAL deinterlacing"** (`local_trajectories`, `local_trajectory_of`, `LocalCausalRow`) — answers *"what did one owner actually do, in its own order?"* by splitting a globally-interleaved durable log `A@s0, C@s0, B@s0, A@s1` back into per-owner chains. **This layer is a completely different mechanism** (owner+cast_seq grouping, not version/horizon filtering) and is production-wired (see §C). + +Key types, with line anchors: + +| Type / fn | Lines | Role | +|---|---|---| +| `LanceVersion = u64` | 72 | The storage-frame clock tick. | +| `EpistemicMode { Strict, Aware, Retro }` | 76–113 | `for_rung(rung)`: 0–4→Strict, 5–8→Aware, 9+→Retro (91–97). `admits(status)`: Strict admits only Contemporary; Aware additionally admits Anachronistic; Retro additionally admits Spoiler; **Unknowable is never admitted by any mode** (99–112). | +| `TemporalStatus { Contemporary, Anachronistic, Spoiler, Unknowable }` | 116–126 | The per-row verdict. | +| `QueryReference { server_id: u16, ref_version: LanceVersion, hlc_tick: Option, mode: EpistemicMode, rung: u8 }` | 133–176 | The reader's stamped view. `default()` = `server_id:0, ref_version:u64::MAX, hlc_tick:None, mode:Strict, rung:0` (150–161) — i.e. "latest, strict, single-server." `at(ref_version, rung)` derives `mode` from `rung` (163–176). | +| `classify(row_version, knowable_from, v_ref)` | 184–199 | The core decision function (reproduced verbatim below). | +| `DependsClosure` trait + `NoDeps` | 213–277 | The DATA-causal axis seam; `NoDeps` is the only shipped impl and is `satisfied: true` unconditionally. | +| `Classification::dispatchable(mode)` | 289–297 | `mode.admits(temporal) && data_ready` — conjunction of the TIME and DATA axes. | +| `DeinterlaceRow` trait | 318–330 | `subject()`, `lance_version()`, `knowable_from()`, `hlc_tick() -> Option { None }` (default). | +| `deinterlace(rows, v_ref, deps)` | 345–376 | Filters by `dispatchable`, then `sort_by_key((hlc_tick.unwrap_or(lance_version), lance_version))` (369–374) — HLC-first ordering with an explicit, tested fallback to `lance_version` for rows carrying no tick (Codex P2 on #468, tested at 728–752). | +| `LocalCausalRow` trait, `local_trajectories`, `local_trajectory_of` | 397–448 | Layer 1 — owner (`MailboxId`) + `cast_seq` (monotonic **per owner across restarts**, explicitly NOT a per-process counter, 404–411). | + +`classify` itself (185–199), the exact decision the whole layer reduces to: + +```rust +pub fn classify(row_version, knowable_from, v_ref) -> TemporalStatus { + if knowable_from > v_ref.ref_version { Unknowable } + else if row_version <= v_ref.ref_version { Contemporary } + else if v_ref.mode == Retro { Spoiler } + else { Anachronistic } +} +``` + +Note precisely what is and is not compared: **only `ref_version` (=`A_last` below) ever appears on the right-hand side.** There is no separate "now" bound anywhere in this function. + +The test suite (450–870) is genuinely a falsifier suite, not decoration: `layer1_deinterlaces_interleaved_global_log_into_owner_local_chain` (524–578) is the operator's own worked example; `no_hindsight_streamed_known_game` (754–869) replays a 10-ply chess-opening stream and asserts a `Strict` reader at `v` sees exactly plies `0..=v` and that the *same* future ply reads `Anachronistic`+refused under `Strict` but `Spoiler`+admitted under `Retro` at the *identical* `v_ref` — this is the sharpest falsifier of the STRICT/RETRO distinction in the tree. + +### B. `lance-graph-contract/src/temporal_pov.rs` (315 lines, fully read) — zero-dep mirror + +`VersionRange { from, to }` (half-open) + `TemporalPov { range, rung }` deliberately re-derive **only** the version-range half of `QueryReference`/`classify` (module doc 18–27: *"The richer per-row classification … is deliberately left to `classify`/`deinterlace` on the planner side"*). `TemporalPov::at(ref_version, rung)` is `[0, ref_version+1)` — i.e. it hard-codes `A_last`'s lower admission bound at `0`; there is still no independent "now" concept here either. + +### C. The real write path — where `DatasetVersion` (transaction time) actually comes from + +- `lance_graph_contract::scheduler::DatasetVersion(pub u64)` (`crates/lance-graph-contract/src/scheduler.rs:33–34`) — doc: *"the surreal Timeline tick, i.e. one entry of `Dataset::versions()`."* +- `lance_graph_planner::persist_sink` (1752 lines; read to line 1342) defines `CycleFrame { cycle: CycleId, base_version: DatasetVersion }` — the sealed predecessor `Vn` every thought in an open cycle reads (144–163) — and `WalSink::commit_cycle` publishes exactly `Vn+1` atomically per cycle (519–578). The module doc is explicit that **write-side ordering is not `temporal.rs`'s job** (25–35): completion-order races are resolved by `order_cycle_stably` (stable sort on `stream_position`) *before* the WAL append, so `temporal.rs`'s read-side machinery never has to repair order. +- `lance_graph::graph::cycle_sink::LanceCycleWriter` (2208 lines) is the concrete `WalSink` — it wraps a REAL `lance::Dataset` (`use lance::dataset::{Dataset, WriteMode, WriteParams}`, line 89) and reports `DatasetVersion(self.ds.as_ref().map_or(0, |d| d.version().version))` (line 437) — i.e. `DatasetVersion` really is Lance's own `.version().version`, not a shadow counter, in this one concrete writer. +- `lance_graph_planner::batch_writer::BatchWriter

` (233 lines, fully read) is the **pre-write intent stager** ("ahead-firing" — module doc 1–77): `cast()` records intent moves before storage completes; the module doc contains an explicit self-correction (44–71) about a wrong claimed mechanism from an earlier session, and states plainly that `cast()` "has **zero production call sites**" as of the last verified audit (line 18). +- `lance_graph_callcenter::version_watcher::LanceVersionWatcher` (313 lines, fully read) is a **separate, unrelated notification bus**: `std::sync::{RwLock, Mutex, Condvar}`-based always-latest fan-out over `CognitiveEventRow` (never `tokio::sync`, enforced by a compile-time `Send+Sync` test, 302–312). It is bumped by `LanceMembrane::project()` on every commit and has its own subscriber/wake API (`subscribe`, `wait_changed`, `try_changed`). **It is never called from, or by, `temporal.rs`.** It is the closest thing in the tree to a live "T_now changed" signal, but it lives on a completely different type (`CognitiveEventRow`) and a completely different crate boundary. + +### D. `lance-graph/src/graph/scheduler.rs` — the mirror-image, also unwired in production + +`LanceVersionScheduler` (the "OUT-direction… IN-direction core impl" per its own doc, 1–41) is a **real** Lance-backed `VersionScheduler` impl (reads `Dataset` via `VersionedGraph::versions()`). Every call site of `.drive_once(`/`.drive_at_latest(`/`LanceVersionScheduler::new(` in the repo is inside `scheduler.rs`'s own test module (lines 301–420, verified by grep). The only non-test reference elsewhere is `crates/symbiont/src/kanban_loop.rs:42` — and `crates/symbiont` is operator-deprecated as of 2026-08-18 (excluded from this study's scope per the task brief). So **both directions of the version↔kanban bridge that the OGAR doc names ("Sprint 7 … `LanceVersionWatcher`" and its dual) are shipped as tested library code with no live non-deprecated caller.** + +### E. `blw_fusion.rs` — the ONE real (harness-scale) exercise of `classify`/`deinterlace` + +`crates/lance-graph-planner/examples/blw_fusion.rs` (1534 lines; read in full via targeted sections) states outright, in its own header (17–19): *"This is the FIRST `DeinterlaceRow` implementor and the FIRST `deinterlace` caller anywhere in the tree."* It: + +- Seeds a synthetic KJV-verse tenant (`MailboxSoA`) over `S_CYCLES=8` sealed cycles (`SLICE=250` verses/cycle), computing two independent rank-based binary "verdict" projections (`A` = SIGHT basin, `B` = BODY basin, plus an inertness control `Z`) per verse per horizon (110–370). +- Reads the resulting `VerdictRow` stream through **both** `QueryReference::at(V4, RUNG_STRICT=0)` and `QueryReference::at(V4, RUNG_AWARE=5)` — the *same pin*, only the mode differs — and measures Hamming distance / Cohen's-κ-style binary association (`jc::stats::binary_association`, a local crate — `Cargo.toml:77`) between the Strict and Aware readings (735–1071). +- Runs a **drop test across all 8 horizons** (C7, 1410–1507) with a pre-registered `DROP_THRESHOLD=0.01`. +- Closes with an eleven-point **"§6 not-claimed" block** (1509–1530) that is the single most important honesty artifact for this cell: + 1. κ/φ measure *reliability* (overlap), never *validity*. + 2. No p-value (domain-correlated data; cites the workspace's own `I-NOISE-FLOOR-JIRAK` rule). + 3. "…**No zero-copy claim** — `deinterlace` `.cloned()`s the admitted rows (`temporal.rs:363`)." + 4. **"No durability claim — `MemWal` is in-process `Mutex`/`Vec`; 'versions' are sequence numbers, not Lance versions."** + 5. **"No HLC/multi-writer claim… `knowable_from`… constant here… the Unknowable branch is never reached."** + 6. And its own methodological correction (C3, 1525–1530): *"under this corpus `deinterlace` reduces to `filter(v <= ref) + stable sort`; `Unknowable`/`DependsClosure`/HLC axes are inert, not exercised."* + +This is decisive, first-party evidence, independently corroborating the `reasoning_loop.rs` and `batch_writer.rs` self-audits below. + +### F. The repository's own honesty audits (found, not asserted by me) + +- `lance-graph/examples/reasoning_loop.rs:51–66` — **"STATUS: TESTED-ONLY (verified 2026-07-27, §12 substrate trace)."** Quoted verbatim: *"There is no production moment-read: every `deinterlace` call site lives in `temporal.rs`'s own `#[cfg(test)]` module, the only `DeinterlaceRow` implementor is a test struct, and no production code sets a non-zero `server_id` or a `Some` `QueryReference::hlc_tick`… `temporal.rs:504` asserts it."* (Note: `blw_fusion.rs` post-dates this note and is the first exception, itself heavily caveated — see §E.) +- `lance-graph-planner/src/batch_writer.rs:12–20` — **"STATUS: DECLARED — the read side is UNWIRED (verified 2026-07-27)."** + +### G. OGAR `docs/TEMPORAL-TIME-TRAVEL.md` @ anchor `386a6fd8` (full doc read) + +Status line: *"CARVED v0 (2026-06-04). No code — boundary + alignment record."* Contents, condensed: + +- §1: `LanceVersionWatcher` + `std::sync::Condvar` (I-2 invariant: tokio reserved for Layer-3 outbound sinks only) shipped and corrected OGAR's earlier tokio-based Sprint-7 sketch. Sole writer = `LanceMembrane` (callcenter); OGAR (`ogar-runtime`) is a subscriber/reactor only, never a writer. +- §2 — the exact mapping table this cell was asked to verify (reproduced, with **[STATUS]** annotations added from source verification above): + + | Python framework concept | Lance / standing-wave equivalent | **[Verified status, 2026-08]** | + |---|---|---| + | `KnowledgeItem.created_at` | Lance version `V` the row landed at | **REAL** — `DeinterlaceRow::lance_version()`, backed by real `Dataset::version().version` in `LanceCycleWriter`. | + | `KnowledgeItem.knowable_from` | Lance version `V` the row's *class* was registered | **TYPE-VISIBLE, NEVER POPULATED** — every shipped `DeinterlaceRow` impl (the test `Row`, `blw_fusion`'s `VerdictRow`) hardcodes it to a compile-time constant (`0`). No producer reads a real class-registration event. | + | `KnowledgeHorizon` (at time T) | `dataset.checkout_version(V_ref)` | **PARTLY REAL** — `checkout_version` is real Lance API (verified in both source columns, §H below), but `temporal.rs` never calls it; `QueryReference::ref_version` plays the analogous role purely as an in-memory filter over rows already fetched by some other path. | + | `TemporalStatus.CONTEMPORARY/ANACHRONISTIC/SPOILER` | `row.lance_version {≤,>,>(Retro)} V_ref` | **REAL**, exactly as coded (`classify`). | + | `EpistemicMode.STRICT/AWARE/RETRO` | planner-level query annotation | **REAL as a type + tested library**, **UNDEPLOYED** as a system behavior (§E, §F). | + | `EpistemicPolicy.for_rung(N)` | which `ThinkingStyle` opts into which mode | **REAL** — `EpistemicMode::for_rung`, table fixed 0-4/5-8/9+. | + | `CausalChain (depends_on/enables)` | standing-wave tier transitions | **PARTLY REAL** — see §I. | + +- §2 "Cross-server hindsight": names the exact HLC tuple `(server_id, local_lance_version, hlc_tick)` and states the planner should ask *"as of HLC tick T_ref"* rather than *"as of local version V_ref."* **This is precisely `QueryReference`'s shape** (`server_id`, `hlc_tick: Option`) — confirming the code's field names were taken directly from this doc, six weeks before `temporal.rs` shipped (the doc predates the module by inspection of the repo's own dating conventions). +- §3 — Decision #4, **surfaced, not actioned**: should `ActionInvocation.emitted_at_millis` (a plain wall-clock `i64`) become an HLC tuple? Doc's own answer: *"wall-clock `i64` is fine [for now]… single-server causal order is the Lance version sequence itself."* **This decision has still not been revisited** — `lance-graph-contract/src/action.rs:207` still carries `emitted_at_millis: Option` as a plain wall-clock stamp (verified by grep; not HLC-typed). +- §6: *"Both sessions: FYI absorbed; not building yet… OGAR does not build the planner/callcenter epistemology layers."* + +### H. Lance `DatasetVersion` — what the storage engine actually gives you (Column A = pinned 9.0.0, Column B = current upstream) + +Both columns checked (`/tmp/sources/lance-9`, tag v9.0.0; `/tmp/sources/lance-main`, commit `a8e2a7857d3a95ac71446024fbff9c61a7fd51a9`, `Cargo.toml` reports `11.0.0-beta.14`) — **identical on every point below**, so this is a stable Lance-API finding, not a 9.0.0 quirk: + +- `rust/lance/src/dataset.rs:213–224` — `pub struct Version { pub version: u64, pub timestamp: DateTime, pub metadata: BTreeMap }`. `version` is the pure transaction-time sequence number (monotonic, one MVCC manifest per commit). `timestamp` is wall-clock commit time, per-writer, **not causally comparable across writers** without an external synchronization assumption (exactly OGAR §2's own warning). `metadata` is an opaque app-writable key/value bag attached to the manifest at commit time. +- `rust/lance/src/dataset.rs:456` — `checkout_version(version: impl Into)`. +- `rust/lance/src/dataset/refs.rs:30–39` (identical in both columns) — `pub enum Ref { VersionNumber(u64), Version(Option, Option), Tag(String) }`. **There is no time-based variant.** A `grep -i "checkout.*timestamp|as_of"` over `rust/lance/src/` in both columns returns nothing relevant to a timestamp-indexed checkout. +- `rust/lance-table/src/format/transaction.rs:17–22` — `Transaction { inner: pb::Transaction }`, an opaque protobuf wrapper. No HLC field anywhere in the transaction record. + +**What this means for the transaction-time axis:** Lance gives you exactly one clean coordinate — a totally-ordered, monotonic, per-dataset sequence number, checked out by number or tag, with an auxiliary (non-authoritative, non-queryable-by-range) wall-clock stamp. This is what `LanceVersion`/`DatasetVersion` faithfully wraps throughout the `lance-graph` planner/contract/callcenter stack. + +**What would be required for a genuine valid-time axis (Snodgrass's second temporal dimension — "when a fact was true in the modeled reality," independent of when it was stored):** Lance itself supplies **nothing**. `checkout_version`/`Ref` only ever addresses *storage* history. A valid-time query ("what did we believe was true as of 2026-01-01, using everything we know *today*") needs an ordinary **row-level Arrow column pair** (`valid_from`, `valid_to`) queried by an ordinary DataFusion predicate — entirely outside the version/checkout machinery, and it would compose *orthogonally* with `checkout_version` (transaction time) to form a true bitemporal table, exactly the "Bitemporal Property Graphs" pattern found in prior art (§ PRIOR ART). **The codebase does not currently do this.** `knowable_from` looks like a first step toward valid-time semantics but is measured in `LanceVersion` units (§G table) — it is a *coarser transaction-time* coordinate (when the *class* was registered), not an independent valid-time domain. So today the system has **two transaction-time-flavored coordinates** (row version, class-registration version) and **zero true valid-time coordinates**. + +### I. The DATA-causal axis has a real candidate backing store, wired to nothing + +`crates/lance-graph/src/graph/spo/link_chain.rs` parses genuine `depends_on` SPO triples out of Odoo field-dependency declarations (`odoo_ontology.spo.ndjson` sits alongside it) and already ships `split_all_depends_on` / per-subject dependency-chain decomposition, tested (`link_chain.rs:319+`). This is exactly the shape `temporal.rs`'s `DependsClosure` trait wants (`(subject, "depends_on", object)` edges, module doc 212–223). **A `grep` across the whole tree for `impl DependsClosure for` outside `temporal.rs` itself returns nothing** — no adapter connects `link_chain.rs`'s real graph to the trait. `NoDeps` (universally `satisfied: true`) is the only impl anywhere. + +### J. A second, unrelated "temporal window" mechanism exists in the same repo + +`lance-graph-contract/src/causal_witness.rs` (A9 `CausalWitnessFacet`, EXPERIMENTAL status per its own header) reads a **different** 12-byte register as 24 signed 4-bit loci naming an offset into a **fixed ±8-slot Markov window**, and is explicitly cross-referenced from `causal_witness.rs`'s own doc as "the `±8` `temporal.rs` Markov window" even though `temporal.rs` (the file studied here) has **no `±8` window concept at all** — that phrase is a leftover from the pre-2026-07-10 VSA-braid design (superseded per this repo's own `lance-graph/CLAUDE.md` "2026-07-10 supersession" note quoted in this session's system context, which states the Markov trajectory "moves off the VSA braid onto the `temporal.rs` sorted stream… generalizes the ±5 window"). This is a live naming-collision hazard: two genuinely different bounded/unbounded temporal mechanisms share a "window" vocabulary in cross-references that have not been kept in sync. + +--- + +## MECHANISM — precise `T_now × A_last × W_write` mapping + +| Hypothesis coordinate | Code reality | +|---|---| +| **`W_write`** (a row's write tick) | `DeinterlaceRow::lance_version()` → real, `LanceVersion = u64`. In the one concrete writer (`LanceCycleWriter`) this is a genuine Lance MVCC version number. **Exists, is real, is the least controversial coordinate.** | +| **`A_last`** (reader's last-awareness / reference horizon) | `QueryReference::ref_version`. Real, tested, is exactly what `classify`'s `row_version <= v_ref.ref_version` compares against. **Exists, is real.** | +| **`T_now`** (current tick, i.e. the live store head) | **Does not exist as a type anywhere.** `QueryReference` carries no field for "the actual current version." The closest thing is `ref_version: u64::MAX` — a *sentinel meaning "unbounded," not a value tracking the real head*. The one place an actual current head is computed is `LanceCycleWriter::head_version()`-shaped code (`self.ds.as_ref().map_or(0, |d| d.version().version)`, `cycle_sink.rs:437`) — but that value **never flows into a `QueryReference`.** The live-notification analogue, `LanceVersionWatcher`, tracks change on a *different type* (`CognitiveEventRow`) on a *different crate boundary*, and is never consulted by `temporal.rs`. | + +**Consequence for the proposed window `A_last < W_write ≤ T_now`:** the code implements only the **left** boundary of this window, explicitly (`classify`'s `row_version <= ref_version` test defines Contemporary; `EpistemicMode::Aware.admits(Anachronistic)` admits *everything* with `row_version > ref_version`, unconditionally). There is **no code-level upper bound** at `T_now`. In a real deployment the upper bound would be enforced only *physically* — a store cannot contain a row whose version exceeds its own current head — never *logically*, by any field or comparison in `QueryReference`/`classify`/`deinterlace`. So the hypothesis's three-coordinate window is **not modeled as three coordinates; it is modeled as two** (`A_last`, `W_write`), with the third (`T_now`) present only as an ambient physical invariant of "whatever the caller happened to hand to `deinterlace`." **This is a genuine finding, not a nitpick:** nothing stops a caller from constructing `QueryReference::at(ref_version, 9)` (Retro) with `ref_version` set arbitrarily far beyond any version that will ever exist, and nothing in the type system distinguishes "Spoiler because I peeked at a real future row" from "Spoiler because I pinned a horizon nothing has reached yet" — `Anachronistic`/`Spoiler` classification is oblivious to whether `T_now` has caught up to `row_version` at all. The distinction is invisible to the code and only meaningful if a caller *happens* to only ever pass rows that genuinely exist in the live store — which is exactly the caveat `blw_fusion.rs` states outright ("versions are sequence numbers, not Lance versions" — i.e., even the one real exerciser sidesteps a live store). + +**HLC reality check** (Kulkarni/Demirbas 2014, "Logical Physical Clocks" — searched and confirmed, §PRIOR ART): a real HLC is a `(physical_ms, logical_counter)` pair with a defined update rule (`max(local_physical, received_physical, prev_hlc) + increment-on-tie`) that provably stays within bounded skew of NTP time while preserving happens-before. `QueryReference::hlc_tick: Option` is a **bare, externally-minted, opaque integer** with only a comparison-and-fallback rule (`deinterlace`'s sort key). There is no update rule, no physical-clock component, no monotonicity proof, and — per grep — **no `HybridLogicalClock` struct anywhere in the tree.** The field is *named* after HLC (taken verbatim from the OGAR doc's §2 vocabulary) but the mechanism that would make the name accurate is absent. This is squarely a case of a documented-but-unimplemented axis; see SURPRISES. + +**`knowable_from` reality check:** conceptually this is the schema/valid-time axis OGAR's mapping intends, but it is typed `LanceVersion` (i.e. still transaction time — see §H) and, per every production `DeinterlaceRow` impl found, is a hardcoded constant. The `Unknowable` branch of `classify` — the one branch that would make `knowable_from` load-bearing — is stated by `blw_fusion.rs` itself to be "never reached" on any real corpus in the tree. + +--- + +## EXPECTED BENEFIT + +1. **Reproducible epistemic pinning for AI-agent forensics.** `nars/insight.rs`'s `VersionedSnapshot`/`steppable_to` (comparing `arena_id`, `mode`, `rung`, `server_id`) is a real, if narrow, consumer that stamps a reasoning-arena reading with the epistemic view it was taken under — enabling "was this insight computed under the same rules as that one" audits, cheaply (field equality, no `classify` call needed). +2. **A working, falsifiable no-hindsight discipline for replay/backtesting.** `reasoning_loop.rs`'s and `blw_fusion.rs`'s streamed-known-game / seated-corpus tests demonstrate the STRICT/RETRO distinction actually holds under adversarial construction (a known-future fact IS excluded under Strict at the identical pin where Retro admits it) — this is a real, tested guarantee, not aspiration, *at the library level*. +3. **Layer 1 (crash recovery) already delivers value independent of Layer 2.** `persist_sink::recover_and_apply` + `local_trajectory_of` are genuinely production-shaped (idempotent watermark, cyclic-lap-safe, tested against scrambled completion order) and do not depend on any of the unproven Layer-2 machinery. +4. **A DataFusion-native valid-time extension is cheap to add later precisely because Lance keeps transaction-time and row data orthogonal** (§H) — nothing in the current design would need to change to add `valid_from`/`valid_to` Arrow columns; they would compose as an ordinary predicate alongside `checkout_version`. + +## EXPECTED FAILURE + +1. **Silent collapse to a version filter.** On any corpus resembling what exists in the tree today (one verdict-row class, `knowable_from` constant, no cross-server writer), the entire epistemic-projection machinery — as `blw_fusion.rs` admits about *itself* — "reduces to `filter(v <= ref) + stable sort`." A team building atop the `EpistemicMode`/`Unknowable`/`DependsClosure` abstraction on the belief that these axes are load-bearing will be surprised when they discover, in production, that they never fire. +2. **HLC-shaped field, non-HLC guarantee.** If a second writer ever mints `hlc_tick` values independently (the stated cross-server use case), nothing in the code enforces the skew-bounded, monotonic-merge property a real HLC provides. `deinterlace`'s sort by raw `hlc_tick` would then silently produce a causally-inconsistent ordering under clock skew, and — because no cross-server integration test exists anywhere in the grep results — nothing would catch it before deployment. +3. **Not zero-copy at the scale the surrounding system cares about.** `deinterlace` `.cloned()`s every admitted row (`temporal.rs:363`, flagged by `blw_fusion.rs` itself). `persist_sink.rs`'s own module doc describes 64k-thought cycles as the target scale; a per-query `O(admitted rows)` clone was never measured against that scale (§EXPERIMENT Design 1). +4. **Two independently-evolving "temporal window" vocabularies in one repo** (`temporal.rs`'s open version-range vs. `causal_witness.rs`'s fixed ±8-slot Markov window) risk a future session conflating them, especially since a stale cross-reference already does (§J). +5. **`OUT`-direction wiring loss.** `LanceVersionScheduler`'s only non-test, non-deprecated consumer status is now zero (its sole external caller was in the operator-deprecated `symbiont` crate) — a genuinely-Lance-backed, tested bridge from version ticks to kanban moves currently has no live caller at all. + +## EXPERIMENT — three pre-registered designs + +All three build directly on `blw_fusion.rs`'s existing pre-registration discipline (fixed constants before any number exists; degeneracy/collapse guards; an inertness control) rather than re-inventing it — the metric names below are drawn from the program's metric list. + +### Design D1-A — "Is deinterlace worth its clone, at cycle scale?" + +- **Setup:** reuse `blw_fusion.rs`'s `VerdictRow`/`Proj{A,B,Z}` corpus generator, but seat it into a real `LanceCycleWriter` (§C) instead of the synthetic `MemWal`, so `DeinterlaceRow::lance_version()` values are genuine `Dataset::version().version` numbers (closing the harness's own "not a durability claim" gap). +- **Metric:** point/neighborhood/random-take latency and CPU for (a) `classify`+`deinterlace` over an in-process `Vec` fetched once per horizon, vs (b) an equivalent DataFusion predicate (`WHERE lance_version <= :ref_version`) pushed to the physical scan, at population sizes 250 / 2,500 / 25,000 / 65,536 rows (the repo's own canonical cycle size, `.claude` context: 65,536 rows = 32 MiB). +- **Baseline:** (b) above (predicate pushdown, no Rust-level post-filter). +- **Control:** an inertness arm — `deinterlace` called with `EpistemicMode::Strict` and `ref_version = u64::MAX` (admits everything) should match (a) with *no* filter at all, isolating the classify/sort overhead from the filter's selectivity. +- **Kill condition:** if (a) is slower than (b) by more than 2× at 65,536 rows, the in-process epistemic layer is not the right place to enforce this filter at cycle scale — the correct fix is a physical-plan-level `EpistemicMode` predicate, not a post-fetch Rust filter, and the finding should be written back onto `temporal.rs`'s own module doc (which currently makes no performance claim either way). + +### Design D1-B — "Does `hlc_tick` actually behave like an HLC under skew?" + +- **Setup:** two synthetic writers, A and B, with independently-advancing "physical" clocks (B's clock rate perturbed ±20%/±50ms relative to A's, matching realistic cross-datacenter NTP skew). Writer A emits a row R1 that is causally acknowledged by writer B (an explicit happens-before edge recorded out-of-band, e.g. "B's row R2 cites R1's content") before B emits R2. +- **Metric:** rate of happens-before violations in the `deinterlace`-produced order — i.e. how often the sort-by-`(hlc_tick.unwrap_or(lance_version), lance_version)` (`temporal.rs:369-374`) places R2 before R1 despite the explicit causal edge. +- **Baseline:** wall-clock-only ordering (sort by `Version.timestamp`, §H) — the alternative decision #4 in the OGAR doc explicitly considered and deferred. +- **Control:** a genuine HLC implementation (Kulkarni/Demirbas update rule) fed the same skew trace, as the theoretical zero-violation reference. +- **Kill condition:** if the current opaque-`u64` `hlc_tick` scheme shows a violation rate statistically indistinguishable from the wall-clock-only baseline (i.e., it buys nothing over naive timestamps once tick values are externally minted rather than merged by an HLC update rule), the field's name is actively misleading and either (i) a real HLC merge function needs to be built and made the sole minting path, or (ii) the field should be renamed/documented as "opaque cross-server tie-break, causality NOT guaranteed" rather than HLC. + +### Design D1-C — "Does the DATA-causal axis do anything on real dependency data?" + +- **Setup:** implement a `DependsClosure` adapter over `link_chain.rs`'s real `split_all_depends_on` output from `odoo_ontology.spo.ndjson` (§I) — the one already-shipped, non-trivial dependency graph in the tree. +- **Metric:** fraction of `TIME`-Contemporary rows that `classify_ready`'s conjunction (`Classification::dispatchable`) actually withholds due to an unsatisfied `depends_on` edge, versus the `NoDeps` baseline (which withholds 0% by construction). +- **Baseline:** `NoDeps` (current universal impl). +- **Control:** a randomly-shuffled dependency graph of identical edge count/degree distribution (Erdős–Rényi-style null), to separate "real Odoo dependency structure withholds rows" from "any graph of this density withholds rows." +- **Kill condition:** if the real-graph withhold rate is statistically indistinguishable from the shuffled-null withhold rate, the axis's current per-subject `satisfied: bool` shape (no cycle handling, no fixed-point iteration — `DepClosure` is computed once per `classify_ready` call, `temporal.rs:303-314`) is not capturing anything the graph structure actually contains, and the trait likely needs a recursive/fixed-point redesign before it is worth wiring to a real dependency source. + +--- + +## PRIOR ART (searches actually run) + +1. **WebSearch: "hybrid logical clocks HLC Kulkarni causality distributed systems paper"** → confirmed the canonical reference (Kulkarni, Demirbas et al., "Logical Physical Clocks and Consistent Snapshots in Globally Distributed Databases," 2014) and its structure: 64-bit `(physical_ms, logical_counter)`, monotonic, NTP-bounded, used by CockroachDB/MongoDB/YugabyteDB. Used to establish that the repo's `hlc_tick: Option` is **not** an implementation of this — it has no physical component and no merge rule. +2. **WebSearch: "bitemporal database valid time transaction time Snodgrass Datomic as-of query"** → confirmed the standard valid-time/transaction-time vocabulary (Snodgrass, TSQL2; XTDB's bitemporality docs; a 2026 "Bitemporal Property Graphs" paper on dealing with both valid and transaction time in graph stores) and that valid-time and transaction-time are queried by structurally different mechanisms (`AS OF ts` vs `AS OF tx-ts`). Used directly to characterize what Lance's `checkout_version` does/does not give (§H) and to show `knowable_from` is transaction-time-flavored, not true valid-time. +3. **WebSearch: "deinterlacing" metaphor distributed systems multi-writer log merge causal order** → returned only generic causal-ordering/vector-clock material; **no direct prior art found** for "deinterlacing" as a named pattern for multi-writer log merge outside video-engineering usage. Classified as a NOVEL-CANDIDATE naming/metaphor choice (not a mechanism claim — the underlying merge-sort-by-clock mechanism itself is KNOWN, per vector-clock/HLC literature). +4. **WebSearch: epistemic access control "spoiler" hindsight bias distributed database read isolation as-of query security** → returned conventional snapshot-isolation and access-control-policy material; **no direct prior art found** for a three-tier (Strict/Aware/Retro) reader-permission ladder gating access to future-dated rows as a first-class database access-control primitive. The nearest KNOWN neighbors are ordinary snapshot isolation (`AS OF` reads) and vector-clock causal consistency, neither of which frames "seeing the future" as a graded, intentional, auditable permission. +5. **In-repo search** (not external, but a documented "prior art" pass in the sense the task requires): exhaustive grep for `HybridLogicalClock`, `hlc`, `server_id`, `impl DependsClosure`, `impl DeinterlaceRow`, `.drive_once(`/`.drive_at_latest(`, `checkout.*timestamp`/`as_of` across both `lance-graph` and both pinned/upstream Lance columns — the source of every "unwired"/"real"/"absent" verdict in this report. + +## SURPRISES + +1. **The OGAR doc's `(server_id, local_lance_version, hlc_tick)` triple was carried into `QueryReference`'s field names essentially verbatim, but only the middle field (`ref_version`) ever got a real semantics — the other two shipped as inert placeholders that are still zero/`None` in every production code path six-plus weeks later.** — **TRANSFER** (a documented cross-server design was faithfully transferred into a type signature; the *mechanism* behind two of its three fields was not transferred, and no code path currently exercises them.) +2. **`deinterlace`'s hypothesis window `A_last < W_write ≤ T_now` is realized as only a one-sided inequality in code (`W_write ≤ A_last` for admission), with the upper bound `T_now` enforced nowhere — not even as a type.** — **NOVEL-CANDIDATE** (searched: no prior-art hit for a named "epistemic window" primitive with an explicit, enforced upper temporal bound in mainstream distributed-database literature; this repo's own implementation likewise does not enforce the upper bound, so the "window" is a real gap, not merely an unresearched pattern). +3. **A real, non-trivial `DependsClosure` backing store already exists in the tree (`link_chain.rs`'s Odoo `depends_on` graph) and has zero lines of code connecting it to `temporal.rs`'s DATA-causal axis.** — **KNOWN gap class** (missing-adapter pattern is common; noted here because it is a *ready-made, real* dataset sitting one file away from the trait it would satisfy, which is an unusually cheap unwired seam). +4. **The repository's own `blw_fusion.rs` admits that, on the one corpus it was ever run against, `deinterlace` "reduces to `filter(v <= ref) + stable sort`" — i.e., the harness built specifically to exercise the epistemic machinery independently discovered, and stated in its own source, that the machinery's distinguishing branches (`Unknowable`, `DependsClosure`, HLC-ordering) are inert on real-shaped data.** — **DISPROVED** (this is the mechanism's own authors falsifying the claim that the richer axes matter, on the best evidence currently in the tree — a strong, first-party negative result, exactly the kind the research program values). +5. **Lance's `checkout_version`/`Ref` API is identical across the pinned 9.0.0 and the current (11.0.0-beta.14) upstream — no timestamp-indexed checkout has been added in the intervening major-version jump — meaning any future valid-time support genuinely has to be built entirely above Lance, never inside it.** — **KNOWN** (consistent with Lance's documented design as a transaction-time-only MVCC store; confirmed here by direct two-column source comparison rather than assumed). + +## VERDICT + +`temporal.rs` is a carefully engineered, honestly self-documented, and *genuinely tested* implementation of a two-layer temporal model — but as of this snapshot it is a **shadow system**: Layer 1 (causal deinterlacing / crash recovery) is production-wired through `persist_sink::recover_and_apply`; Layer 2 (the STRICT/AWARE/RETRO epistemic projection this cell was asked to study) has exactly one non-test exerciser in the whole tree (`blw_fusion.rs`), and that exerciser's own source code states, in eleven numbered disclaimers, that none of the interesting axes (`Unknowable`, `DependsClosure`, HLC) are exercised by it. Of the hypothesis's three coordinates, two (`A_last`=`ref_version`, `W_write`=`lance_version`) are real and precisely implemented; the third (`T_now`) does not exist as a type or field anywhere and the "window" it would bound is enforced only as an ambient physical property never checked in code. The `hlc_tick` field borrows its name from a real academic mechanism (Kulkarni/Demirbas HLC) it does not implement. Lance itself supplies a clean, verified transaction-time axis and nothing resembling valid-time — any bitemporal extension is entirely future work, orthogonal to everything currently built. Recommend: before any further Layer-2 feature work, run Design D1-A (is the in-process filter worth its clone at cycle scale) and D1-B (does `hlc_tick` need a real HLC merge rule before a second writer ever exists) — both are cheap, both would directly settle open questions the repository's own comments already flag as unresolved. diff --git a/docs/lotus/rp-seal-v1/D2.md b/docs/lotus/rp-seal-v1/D2.md new file mode 100644 index 00000000..b05f9d8f --- /dev/null +++ b/docs/lotus/rp-seal-v1/D2.md @@ -0,0 +1,676 @@ +# D2 — Domain D (temporal deinterlacing), ROLE: ADVERSARY + +**Thesis under attack:** *scheduler chronology must not become semantic chronology.* + +**Short answer:** the thesis is **CONFIRMED as a real and currently-violated +requirement**, not as an achieved property. I found **eleven** distinct places where +scheduler/arrival/physical order leaks into a durable coordinate, an identity, a +recovery decision, or a reader-visible order. **One** of them is already pinned +in-tree (`F-ORD-REAL`, `batch_hash`). The other ten are not, and two of those are, +in my judgement, more dangerous than the pinned one — because the pinned one +**fails closed** (hash mismatch → refuse) while they **fail silent** (drop a +transition, reorder a replay, return a stale projection). + +The second, larger finding is scoping: the epistemic layer +(`STRICT/AWARE/RETRO`) is **declared but unwired**. There is no production +`deinterlace` caller and no production `DeinterlaceRow`. So a hindsight claim about +this system today is a claim about a test double, and the F-HINDSIGHT family below +must be designed to say so. + +--- + +## SOURCE ARCHAEOLOGY + +Read directly, this session, no summaries. Repo `lance-graph` at commit `300afa7`. + +| Artifact | Path | What it actually is | +|---|---|---| +| Temporal epistemology + 2-layer deinterlace | `crates/lance-graph-planner/src/temporal.rs` (870 L) | `EpistemicMode`, `TemporalStatus`, `QueryReference`, `classify`, `deinterlace` (layer 2); `LocalCausalRow`, `local_trajectories` (layer 1) | +| Ahead-firing intent recorder | `crates/lance-graph-planner/src/batch_writer.rs` (233 L) | `CastId(u64)` from a per-process `next_id`; `BTreeMap` board; **no durable state** | +| Cycle seal / durable contract | `crates/lance-graph-planner/src/persist_sink.rs` (1751 L) | `SweepSlot`, `DetachedCycleBatch::freeze`, `content_hash`, `recover_and_apply`, `WalSink` | +| Cycle driver | `crates/lance-graph-supervisor/src/cycle_driver.rs` (2618 L) | `collect_casts`, `seal_cycle`, `apply_sealed_transitions`, `run_cycle`; **the F-ORD-REAL pin lives at 2491-2618** | +| The real sink | `crates/lance-graph/src/graph/cycle_sink.rs` (2208 L) | `LanceCycleWriter`, the only non-fake `WalSink`; `scan_sealed` at 971 | +| Zero-dep mirror | `crates/lance-graph-contract/src/temporal_pov.rs` | `TemporalPov` / `VersionRange` — the surface external consumers (stockfish-rs) use | +| OGAR temporal model | `OGAR docs/TEMPORAL-TIME-TRAVEL.md` @ `386a6fd8` | **CARVED v0, "No code — boundary + alignment record"**; explicitly carves the epistemology OUT of OGAR into the planner | + +### Archaeology finding A0 — the OGAR anchor is a boundary record, not a design + +`TEMPORAL-TIME-TRAVEL.md` §2 maps the Python framework onto Lance versions and +concludes *"query-level annotation, not storage… No new storage, no new contract, +no new container."* §6 is titled **"Holding pattern"**: *"FYI absorbed; not building +yet."* Its one live technical warning is decision #4: `emitted_at_millis` is +**wall-clock `i64`**, which *"is not causally ordered across servers"*. + +The brief's shorthand — `T_now × A_last × W_write` with an interlacing window +`A_last < W_write ≤ T_now` — **does not exist in the implementation**. The real +`classify` is a **two**-input comparison plus a mode: + +```rust +// temporal.rs:185-199 +if knowable_from > v_ref.ref_version { Unknowable } +else if row_version <= v_ref.ref_version { Contemporary } +else if Retro { Spoiler } else { Anachronistic } +``` + +There is no `A_last`. `knowable_from` is a *schema* clock (when a class was +registered), not an awareness clock. **A researcher who relies on the three-clock +shorthand will design experiments against a machine that isn't there.** I flag this +as the single most likely source of program-wide error. + +### Archaeology finding A1 — the wiring status, in the source's own words + +`batch_writer.rs:12-20`: +> **STATUS: DECLARED — the read side is UNWIRED (verified 2026-07-27)… `deinterlace` +> has no production caller (all call sites are in `temporal.rs`'s own `#[cfg(test)]` +> module), there is no production `DeinterlaceRow` implementor, and `cast()` itself +> has **zero production call sites**.** + +`lance-graph/examples/reasoning_loop.rs:55-66` repeats it and adds: *"no production +code sets a non-zero `server_id` or a `Some` `QueryReference::hlc_tick`."* + +I verified both independently: `grep 'impl.*DeinterlaceRow'` returns three hits — +`temporal.rs` (test mod), `examples/blw_fusion.rs`, `examples/measure_wal_curve.rs`. +`grep 'impl.*LocalCausalRow'` returns three, all tests/examples. `grep -rn hlc` +outside `temporal.rs` returns only doc comments saying it is unset. + +**Consequence for the program:** the write path (`cast → collect_casts → freeze → +commit`) is real and Lance-backed. The read/epistemic path is not. Any D-domain +benefit claim must be split accordingly. + +### Archaeology finding A2 — a cited type that does not exist + +`temporal.rs:404-412` states the `cast_seq` precondition and prescribes the fix: + +> A per-process counter (e.g. `BatchWriter`'s `CastId`, which resets to 0 on +> restart) is therefore **NOT a valid key**… Use a durable-log position instead +> (`persist_sink::LandedWitness` keys on `DurableCoordinate::log_order`, the WAL +> entry position, which never resets). + +`grep -rn "LandedWitness\|DurableCoordinate" --include=*.rs crates/` returns **exactly +one hit: that doc comment itself.** Neither type exists anywhere in the workspace. +The module correctly diagnoses the hazard and then points the fixer at vapor. The +only real key is `SweepSlot::stream_position`, whose monotonicity is a *caller +obligation* enforced by nothing. + +--- + +## MECHANISM + +### The actual chain, end to end (this is the object under test) + +``` +thought → BatchWriter::cast(on_behalf, moves, payload) + └─ CastId(next_id); next_id += 1 ← ARRIVAL ORDER, per-process, resets to 0 + batch_writer.rs:132-138 + → collect_casts(writer, cycle, position_base, row_of) + └─ stream_position = position_base + cast.0 ← THE LEAK SITE + cycle_driver.rs:385 + └─ first move per owner (cast order) → paired_move; rest → held + → DetachedCycleBatch::freeze(frame, artifact_casts) + └─ order_cycle_stably(by stream_position) persist_sink.rs:378 + └─ image.insert(row, payload) (later stream position wins) :382-384 + └─ batch_hash = FNV1a(frame ++ ∀slot: stream_position ‖ owner ‖ row ‖ …) + persist_sink.rs:406-434 + → LanceCycleWriter::commit_cycle → ONE DatasetVersion + → apply_sealed_transitions(fleet, sealed, watermarks) + └─ per applied transition: watermark = stream_position cycle_driver.rs:563-568 +--- crash --- + → scan_sealed() → PHYSICAL Lance scan order, never sorted cycle_sink.rs:971-1055 + → recover_and_apply(owner, sealed, applied_through) + └─ "already in canonical stream order… so no sort is done here" + └─ if stream_position <= watermark { continue } ← SILENT SKIP + persist_sink.rs:665-682 +``` + +`stream_position` is therefore **not one thing**. It is simultaneously: +(1) the write-side canonical sort key, (2) a *field folded into the publication +identity hash*, (3) the same-row coalescing tiebreak, (4) the in-memory apply order, +and (5) the durable recovery watermark. **It is minted from arrival.** Every one of +those five is a channel through which scheduler chronology becomes semantic +chronology; the in-tree pin covers exactly one of them. + +### The leak table + +Severity: **HIGH** = silent semantic loss or divergence reachable in the current +design; **MED** = requires an unimplemented caller obligation to be got wrong, or is +latent behind an unwired path; **LOW** = hygiene / API-shape. + +| # | Leak | Site | Sev | Pinned? | +|---|---|---|---|---| +| **L1** | `stream_position = position_base + CastId` — arrival mints a durable coordinate | `cycle_driver.rs:385`; `batch_writer.rs:132-138` | HIGH | **partly** — only the `batch_hash` consequence | +| **L2** | Recovery replay order **is Lance physical scan order** | `cycle_sink.rs:971-980` (`scan_in_order(true)`, no `order_by`) + `persist_sink.rs:665-666` ("no sort is done here") | HIGH | **no** | +| **L3** | Same-row coalescing resolved by arrival when `row_of` is not injective | `persist_sink.rs:382-384`; `collect_casts` `row_of` param `cycle_driver.rs:361` | HIGH | **no** — the pin asserts `image` equality *only under an injective `row_of`* | +| **L4** | Durable watermark advanced by an **intent-only** transition that never became durable | `cycle_driver.rs:442-452` (transitions over **all** casts) + `563-568`; filtered out at `persist_sink.rs:620-623` | MED | **no** | +| **L5** | `next_position_base` returns **0** for an empty cycle; the `max(prev, next)` carry is a caller obligation | `cycle_driver.rs:453-457`, contract at `:168-172` | MED | one test at `:2313` | +| **L6** | `QueryReference::default()` = `ref_version: u64::MAX` ⇒ **every** row is `Contemporary`, incl. future rows, even under `Strict` | `temporal.rs:150-161` + `190-199` | MED-HIGH | **no** (the test asserts the *value*, not the consequence) | +| **L7** | `server_id` is carried but **never read** by `classify`/`deinterlace`; cross-server `lance_version`s are incomparable | `temporal.rs:134-137, 185-199, 346-376` | MED | **no** | +| **L8** | `deinterlace` sort key mixes an HLC tick and a Lance version as if commensurable (`hlc.unwrap_or(lance_version)`) | `temporal.rs:369-374` | MED | **no** — the test picks HLC ≈ version, concealing it | +| **L9** | `local_trajectories` stable-sort ⇒ on a `cast_seq` tie, **scan order becomes replay order**; the prescribed durable key does not exist | `temporal.rs:424-433`, `404-412` | MED | **no** | +| **L10** | Divergent mirror: `TemporalPov::at(u64::MAX,0)` **excludes** `u64::MAX`; `classify` **includes** it | `temporal_pov.rs:177-186` vs `temporal.rs:192` | LOW-MED | documented as "one divergence", not tested as a contradiction | +| **L11** | `mode` and `rung` are independent public fields; `for_rung` policy is bypassable by struct literal | `temporal.rs:133-148` | LOW | **no** | + +### L1 — independently confirmed, and its scope is wider than the pin + +The in-tree pin `f_ord_real_defect_pin_arrival_order_leaks_into_batch_hash` +(`cycle_driver.rs:2557-2600`) is correct and well-built: it perturbs *the order of +`cast()` calls* (upstream of the mint, per the operator's own sharpening — "do not +test permutation after the order key already exists"), carries an anti-vacuity check +that the perturbation reached the mint, and pins that `batch_hash` differs. I +reproduce its reasoning from source and agree. + +**But its two "legs that hold" are conditional, and it says so only implicitly.** + +- `assert_eq!(b1.image, b2.image, "the coalesced image is arrival-independent")` + holds **because the fixture passes `row_of = u64::from`, which is injective.** + The comment credits "`row_of(owner)` discipline" — but the discipline that makes + it true is *injectivity*, not identity-derivation. A banked/sharded layout + (`|o| o % NBANKS`), a delegated write-on-behalf where two mailboxes resolve to one + owner row, or any many-to-one `row_of` breaks it immediately, and then + `image.insert(row, payload)` (last-in-stream-order wins) makes the **durable + content** arrival-dependent. That is **L3**, and it is strictly worse than the + pinned defect: a wrong hash fails closed, a wrong image is durable and silent. +- `semantic_set` deliberately strips `stream_position` before comparing. So the pin + proves the *content* is stable while *the coordinate* is not — which is exactly + the setup for L2/L4/L5, where the coordinate is what recovery consumes. + +**Judgement:** the pinned defect is the least harmful member of its own family. + +### L2 — the one I would escalate first + +`recover_and_apply` documents: *"The landings arrive already ordered (write-side +deinterlace at seal), so no sort is done here."* The only real supplier is +`LanceCycleWriter::scan_sealed`, which issues a filtered projection with +`scan_in_order(true)` and **no `order_by`**. + +Pinned-column check (`/tmp/sources/lance-9`, tag v9.0.0), `rust/lance/src/dataset/scanner.rs:1407-1423`: + +> *"Set whether to read data in order (default: true). A scan will always read from +> the disk concurrently. If this property is true then a ready batch … will only be +> returned if it is the next batch in the sequence."* + +`ordered` means **"do not let parallelism scramble the batch sequence"** — it is a +statement about *stream assembly*, not about any column. The sequence it preserves +is the manifest's fragment order. So: + +> **The durable semantic replay order of the kanban lifecycle is the physical +> fragment order of a Lance dataset.** + +Today that happens to coincide with `(cycle, stream_position)` because cycles are +appended monotonically and `freeze` sorted within a cycle. It is coincidence, not +contract. And Lance itself contains the operation that breaks it — +`rust/lance/src/dataset/transaction.rs:3075-3084`: + +```rust +if let Some(replace_range) = replace_range { + final_fragments.splice(replace_range, new_fragments); // contiguous: order preserved +} else { + // Slower path for non-contiguous ranges + for fragment in group.old_fragments.iter() { + final_fragments.retain(|f| f.id != fragment.id); + } + final_fragments.extend(new_fragments); // ← APPENDED TO THE END +} +``` + +**Both columns identical.** Current upstream (`/tmp/sources/lance-main`, +`version = 11.0.0-beta.14`) carries the same branch verbatim at +`transaction.rs:3804-3813`. This is not a 9.0.0 wart that a bump removes. + +*Honest reachability leg (this is where I constrain my own claim):* the **default** +planner cannot reach it. `DefaultCompactionPlanner` (`optimize.rs:734-800`) builds +bins from **contiguous runs** of candidate fragments — `(None, Some(_)) => bin is +complete` closes the bin the moment a non-candidate is met. So default +`compact_files` produces contiguous groups → the `splice` path → order preserved. +The non-contiguous branch is reachable via `compact_files_with_planner` +(`optimize.rs:862`, public) with a custom `CompactionPlanner`, and by any future +planner change upstream. + +So L2's severity is **HIGH-conditional**: the dependency is real, undocumented, +unpinned, and defended only by an implementation detail of a *pluggable* component +in a dependency the workspace treats as upstream-authoritative and bumps in +lockstep. The failure is not a crash. `recover_and_apply` walks an out-of-order +stream, meets a **high** `stream_position` first, sets the watermark there, and then +`continue`s past every lower position — **silently discarding durable kanban +transitions on recovery.** No error, no counter, no log. + +### L4 — a non-durable event advances a durable cursor + +`seal_cycle` computes `transitions` over **all** collected casts, artifact *and* +intent-only, and says so (`cycle_driver.rs:160-167`): *"an intent-only +(empty-payload) cast's move still seals here even when the cycle as a whole is +`CommitOutcome::NoChange`."* `persist_cycle` then filters exactly those casts out of +the durable batch (`persist_sink.rs:620-623`). `apply_sealed_transitions` advances +`watermarks[owner] = t.stream_position` for every successfully applied transition — +including the intent-only ones. + +Net: the watermark can point at a `stream_position` that has **no durable landing**. +The design intends phase and watermark to "move together" as one owner state; for an +intent-only move the phase moves **in memory only**. If the watermark is persisted +(the doc requires the caller to persist it "WITH the owner's sealed cycle state") and +the phase is rebuilt from durable data, the two diverge across a crash, and the next +artifact landing's `mv.from` will not match the rebuilt phase → `StalePhase` → the +doc's own words, *"a permanent `StalePhase` stall"*. + +I grade this **MED, not HIGH**, deliberately: watermark persistence has **no +implementation in-tree** (`grep` finds only the in-memory `HashMap` in +`run_cycle`). It is a contract-level trap waiting for the first real implementor, +not a live bug. That distinction should survive into the program's consolidation. + +### L6 — the default reference is not a horizon + +`QueryReference::default()` sets `ref_version: u64::MAX` as a "latest" sentinel, with +`mode: Strict`. Feed that to `classify`: `knowable_from > u64::MAX` is unsatisfiable, +and `row_version <= u64::MAX` is universally true. **Every row classifies +`Contemporary`.** A `Strict` reader at the default reference has *zero* epistemic +protection — including against rows written after it began reading. The no-hindsight +guarantee is a property of `QueryReference::at(v, rung)` **only**, and the type's +`Default` is the one shape that silently opts out of it. For a stale-awareness or +long-running-reader scenario this is precisely the leak the module exists to prevent. + +A fix must resolve "latest" to a **concrete observed head** at reference-construction +time, not carry a sentinel into the comparison. + +### L8 — the HLC fallback compares incommensurable magnitudes + +```rust +// temporal.rs:369-374 +out.sort_by_key(|r| (r.hlc_tick().unwrap_or_else(|| r.lance_version()), r.lance_version())); +``` + +The existing test `deinterlace_mixed_hlc_falls_back_to_lance_version` uses +HLC `{500, 900}` against version `750` — values chosen in the same numeric range, so +the interleave looks correct. A real HLC is a physical-time-seeded 64-bit quantity +(µs or ns since epoch, ~10^15); a Lance version is a small commit counter (~10^3). +Under an actual partial rollout **every** legacy row sorts *before* every HLC row — +reintroducing exactly the failure the `unwrap_or(0)` fix (Codex P2 on #468) removed, +just displaced by 15 orders of magnitude. The prior fix repaired the symptom at the +tested magnitudes and left the type error. + +--- + +## EXPECTED BENEFIT + +Stated adversarially — what actually survives if the leaks are fixed. + +1. **A real separation of five order-like quantities.** The codebase already names + them distinctly (cast / chunk / cycle; completion order vs stream position vs + `DatasetVersion` vs physical row) and the two-layer split (causal vs epistemic) + is genuinely useful and, as far as I can find, not a standard packaging. That + conceptual scaffolding is the durable contribution and it survives every finding + above. +2. **Write-side canonicalization is the right place.** `persist_sink`'s argument + — canonicalize once at seal, never repair on read — is correct and is what makes + the 64k→1-WAL amortization possible. My L2 finding does not attack it; it attacks + the *unstated assumption that storage preserves the canonicalization*. +3. **Fail-closed publication identity.** `(cycle, batch_hash)` with reconcile-first, + `HashConflict` never promoted, and the deliberate refusal to invent a version for + `Reconciled` is unusually disciplined. The arrival-dependence of `batch_hash` is a + defect *in the key's derivation*, not in the protocol around it. +4. **If (and only if) `batch_hash` becomes content-addressed**, a genuinely new + capability appears: two replicas that computed the same semantic cycle would + publish the same identity, making cross-replica divergence detectable by hash + comparison rather than by data diff. That is a real prize and the strongest reason + to fix L1 beyond hygiene. +5. **Epistemic modes as a query annotation with zero storage cost** is a sound design + and the OGAR doc's placement argument is right. + +**Benefits I do NOT expect and would push back on:** +- Any claim of hindsight-freedom *today*. Unwired (A1) + L6 means the guarantee is + currently demonstrated on a test struct at a non-default reference. +- Any claim that temporal machinery gives cross-server causal order. `server_id` is + inert (L7), HLC is unproduced, and the OGAR anchor explicitly defers it. + +--- + +## EXPECTED FAILURE + +The scenarios I was asked to construct. Each is written to be executable. + +**S1 — Reordered workers (the pinned one, extended).** 64 owners, identical semantic +set, arrival order forward vs reversed. `batch_hash` differs (pinned). *Extension:* +per-owner watermark **values** differ, so any consumer treating a watermark as a +comparable global cursor is arrival-dependent. + +**S2 — Same-row contention under a non-injective `row_of` (unpinned).** Two owners +resolve to one row (bank/shard layout, or write-on-behalf delegation). Both cast +artifacts in one cycle. `image.insert(row, …)` keeps whichever arrived later. **The +durable payload is chosen by the scheduler.** No error. This is the sharpest new +falsifier I can construct and it is a two-line change to the existing fixture. + +**S3 — Restart with a resettable key.** Process restarts; `BatchWriter::next_id` +returns to 0. If the caller takes `position_base` from `SealedCycle::next_position_base` +without the documented `max(previous_base, …)`, and the last cycle was empty +(`unwrap_or(0)`, L5), positions **repeat below an existing watermark** and +`recover_and_apply` silently skips real landings. The type returns `0` and the +correctness carry is prose. + +**S4 — Compaction as semantic corruption (L2).** Seal N cycles; run a compaction +whose group is non-contiguous; `scan_sealed` now yields rows out of `stream_position` +order; `recover_and_apply` sets the watermark from the first (high) row and drops the +rest. **A physical-locality operation destroyed semantic replay order.** This +directly answers the program's cross-domain prompt "is compaction *repair* of +physical locality?" — under this recovery contract it is also a *mutation of +semantic order*, which is the opposite of repair. + +**S5 — Crash between an intent-only advance and the next artifact (L4).** Phase moved +in memory, watermark moved, nothing durable. Rebuild → phase reverts, watermark does +not → next sealed landing trips `StalePhase` → permanent stall. + +**S6 — Stale awareness / long-running reader (L6).** A reader constructed with +`QueryReference::default()` (or `..Default::default()`) sees rows written *after* it +began, under `Strict`. The no-hindsight guarantee silently does not apply. + +**S7 — Cross-server ambiguity (L7).** Rows from `server_id` 0 and 1 with independent +version streams. `classify` compares `row_version <= ref_version` across frames; +`deinterlace` merges them. The result is arithmetically defined and semantically +meaningless. The type surface advertises coverage the body does not implement. + +**S8 — Clock skew, precisely scoped.** There is **no wall clock anywhere in this +path** — `stream_position`, `CycleId`, `DatasetVersion` are all counters. So classic +clock skew cannot corrupt the current write path. It enters only via the *deferred* +HLC, and there L8 (magnitude mixing) bites before skew does. I record this as a +**negative result**: the naive "clock skew breaks it" attack fails against this +design, and that is a point in the design's favour worth stating plainly. + +**S9 — Two "canonical" epistemologies disagreeing (L10).** `TemporalPov::at(u64::MAX, 0)` +computes `checked_add(1)` → `None` → half-open `[0, u64::MAX)` and **rejects** +`u64::MAX`; `classify` with `ref_version = u64::MAX` **accepts** it. An external +consumer on the contract mirror and an internal reader on the planner return +different answers for the same row at the sentinel. + +--- + +## EXPERIMENT (pre-registered) + +All designs are read-only/offline; none requires the architecture to be built. Each +names metric, baseline, control, kill condition. Anti-vacuity guard stated where the +trap is real (this workspace's own falsifiability rule demands the can-fire *and* +can-stay-silent halves). + +### F-ORD-2 — coalescing under a non-injective `row_of` (attacks L3) +- **Metric:** `DetachedCycleBatch::image` byte-equality across two arrival orders. +- **Baseline:** the existing `freeze_with_arrival_order` fixture with `row_of = u64::from` + (injective) — must stay equal (this is the control that proves the perturbation is + not simply breaking everything). +- **Control/perturbation:** same fixture, `row_of = |o| (o % 8) as u64`, distinct + per-owner payloads, forward vs reversed arrival. +- **Anti-vacuity:** assert ≥2 owners share a row, and that the two orders assign the + same owner different `stream_position`s (perturbation reached the mint). +- **Kill condition:** if `image` is byte-equal under the non-injective map, L3 is + **DISPROVED** and `row_of` injectivity is not load-bearing. If it differs, the + in-tree pin's "image leg holds" comment must be re-scoped to "holds iff `row_of` is + injective", and injectivity must become a checked precondition. + +### F-PHYS-ORDER — physical order vs replay order (attacks L2) — *highest value* +- **Metric:** `scan_sealed(None).is_sorted_by_key(|ls| (ls.cycle, ls.slot.stream_position))` + before and after compaction; plus `recover_and_apply` applied-move count and the + program metrics **restart recovery work**, **compaction CPU/wall**, **files & pages + touched per query**, **index remap bytes/CPU**. +- **Baseline:** N sealed cycles, no compaction → sorted, recovery applies all M moves. +- **Control:** default `compact_files` (contiguous bins) → expected still sorted; this + control is what distinguishes "compaction breaks it" from "the non-contiguous + branch breaks it", and I expect it to pass. +- **Perturbation:** `compact_files_with_planner` with a planner emitting one + non-contiguous group (or a middle fragment made non-candidate), forcing + `transaction.rs:3081-3083`. +- **Kill condition:** if `scan_sealed` remains sorted after a *non-contiguous* rewrite, + L2 is **DISPROVED** and Lance provides a stronger order guarantee than its source + reads. If it is unsorted, the finding is confirmed and the fix must preserve: (a) + the write-side-canonicalization principle (do **not** move the sort into + `recover_and_apply` as a blanket read-time repair — that is the amortization + `persist_sink` exists to protect), but (b) either bound `scan_sealed` with an + explicit `order_by(stream_position)` **or** make `recover_and_apply` order-robust by + keying on a per-owner max rather than a first-seen scan. +- **Secondary metric worth capturing:** `order_by` on a scan costs a sort; measure + **compaction CPU/wall** and **pages touched** for the `order_by` variant vs an + index on `stream_position` vs recovery-side robustness. This is the + locality-vs-correctness tradeoff the RP-SEAL program is actually about. + +### F-RESTART — resettable key + empty-cycle reset (attacks L1/L5/L9) +- **Metric:** count of landings `recover_and_apply` **skips** (positions ≤ watermark) + that were never previously applied — i.e. **silently lost transitions**; plus + **restart recovery work** and **version count**. +- **Baseline:** `position_base = max(prev_base, next_position_base)` (correct carry) → + zero lost. +- **Perturbation A:** naive `position_base = sealed.next_position_base` with an + intervening empty cycle (`unwrap_or(0)`). + **Perturbation B:** simulated process restart (`BatchWriter::new()`) with a stale + cached base. +- **Kill condition:** if losses are zero under both perturbations, L1/L5 are + contained by something I did not find. If non-zero, the fix must make the carry a + **type obligation** (e.g. an opaque `StreamCursor` that can only be advanced, never + constructed from a bare `u64`), because prose has already failed once here + (`temporal.rs:410` prescribes a type that does not exist). + +### F-HINDSIGHT family (the epistemic side) + +**F-HINDSIGHT-0 — honesty gate (run first).** Assert, mechanically, that +`deinterlace` has no non-test caller and no non-test `DeinterlaceRow`. **Kill: if a +production caller appears, every later F-HINDSIGHT result must be re-run against it.** +This exists so no F-HINDSIGHT result is ever reported as a claim about production +without the reader knowing it is a claim about a double. + +**F-HINDSIGHT-1 — projection exactness across all three modes.** +- **Metric:** for a version stream `0..N` with a mid-stream `knowable_from = K`, + the exact set returned by `deinterlace` at `v_ref = v` for each of + `Strict / Aware / Retro`. +- **Pre-registered expectation:** `Strict` = `{r : r.version ≤ v ∧ r.knowable_from ≤ v}`; + `Aware` = `Strict ∪ {r.version > v}`; `Retro` = same membership as `Aware` but with + those rows classified `Spoiler`; **`Unknowable` excluded in all three**. +- **Anti-vacuity (both halves):** the three sets must be **pairwise distinct** + (`|Strict| < |Aware|`) — otherwise the test passes on a fixture with no future rows; + **and** at least one row must be `Unknowable` and excluded in all three — otherwise + `knowable_from` is never exercised. The existing + `no_hindsight_streamed_known_game` (`temporal.rs:790`) sets `knowable_from = 0` on + every row and never constructs `Aware`, so it is *not* this test. +- **Kill:** any set differs from the pre-registered one. + +**F-HINDSIGHT-2 — frontier invariance (THE falsifier; nothing in-tree does this).** +- **Claim under test:** *a hindsight-rung read is bit-identical whether or not any + in-flight/frontier activity existed.* +- **Metric:** byte-equality of the serialized projection at fixed `(v_ref, mode)`. +- **Arms:** (a) sealed store, empty writer; (b) same sealed store, an **open** cycle + staging K casts (never sealed); (c) same, staging a **different** K' casts; + (d) same, an open cycle that was staged and then **failed** to seal + (`CommitError::Io`). +- **Baseline:** arm (a). **Kill:** any of (b)/(c)/(d) differs from (a) by one byte. +- **Anti-vacuity:** prove the frontier is non-empty and *would* change the projection + if admitted — i.e. seal arm (b)'s casts in a follow-up cycle and show the projection + at a *later* `v_ref` then does change. Without this, "identical" could hold because + the frontier was empty. +- **Why it matters:** `persist_sink`'s epistemic-race argument is made on the *sink* + side (`scan_sealed` excludes unsealed cycles). This tests it from the *reader* side, + which is where the guarantee is actually consumed, and it is the only design here + that could falsify the sealed-read-horizon claim end to end. + +**F-HINDSIGHT-3 — the default is not a horizon (pins L6).** +- **Metric:** `classify(row_version = head + 1, knowable_from = 0, &QueryReference::default())`. +- **Pre-registered current behaviour:** `Contemporary` (the defect). +- **Kill/flip:** when a fix lands, this must become `Anachronistic` and the test is + re-pinned deliberately. Pair with a can-stay-silent half: `QueryReference::at(head, 0)` + must **still** classify the same row `Anachronistic` both before and after, proving + the fix changed only the sentinel path. + +**F-HINDSIGHT-4 — mode/rung coherence (pins L11).** Construct +`QueryReference { rung: 0, mode: Retro, ..default() }` and show a rung-0 reader taking +a `Spoiler`. Kill: if `classify` consults `rung`, L11 is DISPROVED. + +**F-HINDSIGHT-5 — cross-server refusal (pins L7).** Two rows, `server_id` 0 and 1, +independent version streams. Assert current behaviour (silent comparison), then +require that a fixed `classify` either refuses cross-frame rows or routes them +through an HLC. **Kill:** if any current code path already refuses, L7 is DISPROVED. + +**F-HINDSIGHT-6 — HLC magnitude (pins L8).** Rerun +`deinterlace_mixed_hlc_falls_back_to_lance_version` with HLC ticks at realistic +epoch-µs magnitude (~1.7e15) and versions ~1e3. **Kill:** if the legacy row still +interleaves correctly, L8 is DISPROVED. Expected: every legacy row sorts first. + +**F-HINDSIGHT-7 — mirror agreement (pins L10).** Property test: for `v` over a +boundary-weighted sample including `0, 1, u64::MAX-1, u64::MAX`, +`TemporalPov::at(v, rung).admits(r)` ≡ `classify(r, 0, &QueryReference::at(v, rung)) == Contemporary`. +**Kill:** any disagreement. Expected: exactly one, at `r = v = u64::MAX`. + +### Program-metric mapping (which of the listed metrics each design moves) +- F-PHYS-ORDER → **restart recovery work**, **compaction CPU/wall**, **files & pages + touched per query**, **index remap bytes/CPU**, **fragment count**. +- F-RESTART → **restart recovery work**, **version count**, **stale-version retention bytes**. +- F-ORD-2 → **logical bytes** (durable payload identity), **write amplification** (a + re-published cycle after a hash mismatch). +- F-HINDSIGHT-2 → **point/neighborhood latency** and **pages touched** are the *cost* + side of any fix that adds an `order_by` or a horizon resolution. + +--- + +## PRIOR ART + +Searches I actually ran this session (WebSearch), with what each returned and how it +changed a classification. + +1. `"external consistency" vs "internal consistency" … STRICT AWARE RETRO temporal query` + → Spanner external consistency / TrueTime; nothing on epistemic read modes. The + Spanner framing is directly relevant to L7: external consistency exists precisely + because *"commit time order can be reversed from real-time order"* across machines + — the problem `server_id` is carried for and does not solve. +2. `bitemporal "valid time" "transaction time" "decision time" as-of hindsight lookahead` + → XTDB / bitemporality literature; multitemporal DBs with a **decision time** axis. + **This is the closest prior art to the whole temporal module** and it is decisive + for classification: as-of queries against valid time, retroactive writes, and the + rule *"you cannot write a new transaction with a transaction-time in the past"* are + the standard lookahead-prevention mechanism. `STRICT/AWARE/RETRO` is a **naming and + policy layer over as-of reads**, not a new temporal model. +3. `deterministic database recovery replay order … compaction reorders log replay` + → deterministic-DB survey (CACM), main-memory recovery survey (ACM), TiDB log + compaction. Key line: **"LSN monotonicity ensures a globally ordered sequence that + enables deterministic replay"** — i.e. the standard solution is an explicit + monotonic log coordinate, *not* physical order. Also: compaction copying data + "cannot guarantee that indexes on the duplicate data will be maintained in proper + order." This makes L2 a **known hazard class**, and makes lance-graph's reliance on + physical order a departure from the established answer. +4. `content-addressed idempotency key derived from arrival order … exactly-once anti-pattern` + → Morling on idempotency keys; several practitioner treatments. The consensus + prescription is *"a deterministic hash of the event content"* as the key, and the + named anti-pattern is *"do not use timestamps as the key."* Arrival-ordinal is the + same category error as a timestamp. So L1 is a **known anti-pattern**, instantiated + here. +5. `hybrid logical clock partial deployment … mixing HLC with non-HLC rows` + → HLC fundamentals (Kulkarni et al. lineage, `uhlc-rs`), Raft+HLC session + guarantees. **No result addressed mixed HLC/non-HLC sort keys.** The search engine + itself concluded it is "a practical systems engineering challenge rather than a + widely-discussed theoretical topic." That absence is what supports L8's + NOVEL-CANDIDATE grading below — with the honest caveat that absence of a paper is + weak evidence and this is far more likely folklore than novelty. +6. `"relying on physical row order" / "SELECT without ORDER BY" implicit ordering dependency` + → jOOQ "Don't do this", IBM Db2 support note, Postgres list threads, Michael Swart. + Canonical statements: *"If your query depends on row order and you didn't write + ORDER BY, your query is already broken"*; IBM's note that behaviour changed after an + OS release and *"applications that relied on unspecified ordering might fail after + upgrading."* L2 is textbook — which **raises** its severity (it is a well-known way + to lose data) while **lowering** its novelty to zero. + +**Not searched, and I flag the gap:** I did not search for prior art on the *two-layer* +split itself (per-owner causal deinterlacing distinct from epistemic projection). I +suspect stream-processing literature (per-key substreams, Kafka per-partition +ordering, Flink watermarks) covers it, and I decline to grade it NOVEL without that +search. A follow-up should run it before anyone claims the layering is new. + +--- + +## SURPRISES + +**1. Compaction is a semantic-order-mutating operation under the current recovery +contract.** `scan_sealed` returns physical order (`cycle_sink.rs:971-980`), +`recover_and_apply` explicitly never sorts (`persist_sink.rs:665-666`), and Lance's +`handle_rewrite_fragments` appends non-contiguously-rewritten fragments to the end +(`transaction.rs:3075-3084`, **identical in 9.0.0 and 11.0.0-beta.14**). The program +asked "is compaction *repair* of physical locality?" — here it is also *damage to +semantic order*. Directionally, this argues the two must be **decoupled by an explicit +order key**, not unified. +→ **KNOWN** (implicit-ordering-dependency; search 6, plus the deterministic-replay +LSN answer from search 3). Known as a hazard class; **its instantiation here is +unpinned and, I judge, the most dangerous finding in this cell.** + +**2. The pinned defect is the safest member of its own family.** `batch_hash` +arrival-dependence fails **closed** (`HashConflict`, never promoted). The *unpinned* +siblings — coalesced-image choice under a non-injective `row_of`, the recovery +watermark, the replay order — fail **silent**. A team reading the in-tree pin could +reasonably conclude the arrival leak is characterised; it is characterised on the one +channel where the system already defends itself. +→ **TRANSFER** — "the tested failure mode is the one that was already safe" is a known +review pathology; transferring it to *ordering-key* defects (where the same coordinate +feeds an identity, a content choice, and a cursor) is the specific move here. + +**3. A doc comment prescribes a fix using a type that does not exist.** +`temporal.rs:404-412` correctly diagnoses that `CastId` is not a valid `cast_seq` and +directs the reader to `persist_sink::LandedWitness` / `DurableCoordinate::log_order`. +Workspace-wide grep: **that string appears once, in the doc comment itself.** Meanwhile +`cycle_driver.rs:385` derives the durable coordinate from exactly the `CastId` the doc +forbids, adding a base whose carry is prose. The hazard was seen, named, and then +mitigated by a citation. +→ **NOVEL-CANDIDATE** (as a *review artifact*, not as computer science). Searches 3–4 +turned up no discussion of "documentation citing a non-existent mitigation" as a named +defect class. I grade it NOVEL-CANDIDATE with low confidence and note it is more +plausibly unnamed folklore than genuinely new. + +**4. Clock skew is not an attack surface here, and that is a positive result.** I was +asked to construct a skew scenario and could not: every coordinate on the live write +path (`CastId`, `stream_position`, `CycleId`, `DatasetVersion`) is a counter. Skew +enters only through the deferred HLC — and there the *magnitude* bug (L8) fires before +skew does. The design's avoidance of wall-clock in durable coordinates is exactly what +Spanner needs TrueTime for and what the idempotency literature warns about +("do not use timestamps as the key"). +→ **DISPROVED** (the attack, not the design). The naive skew attack fails; recording +this protects the program from an experiment that cannot produce a result. + +**5. Two "canonical" epistemologies disagree at the sentinel.** +`TemporalPov::at(u64::MAX, 0)` computes `checked_add(1) → None → [0, u64::MAX)` and +**rejects** `u64::MAX` (`temporal_pov.rs:177-186`, asserted at `:299-301`); `classify` +with the same reference **accepts** it (`temporal.rs:192`, `row_version <= ref_version`). +The mirror is the zero-dep surface external consumers use. A sentinel chosen to mean +"no horizon" acquires an off-by-one the moment a second implementation makes it a +range. +→ **KNOWN** (sentinel-as-in-band-value; half-open vs closed interval mismatch). Classic; +worth pinning precisely because it is classic and is currently only prose-documented as +"one documented divergence" rather than tested as a contradiction. + +--- + +## VERDICT + +**The central thesis survives, sharpened and inverted: "scheduler chronology must not +become semantic chronology" is not a property this system has — it is the property it +is currently missing in at least eleven places, ten of them unpinned.** The single +durable coordinate `stream_position = position_base + CastId` is minted from arrival +and then reused as a publication-identity input, a content-coalescing tiebreak, an +apply order, and a recovery watermark; the in-tree `F-ORD-REAL` pin covers only the +first, which is also the only one that fails closed. + +The highest-value finding is **not** the arrival leak, which is known prior art and +already half-characterised. It is **L2**: the durable semantic replay order of the +kanban lifecycle is the *physical fragment order of a Lance dataset*, defended only by +an implementation detail of a pluggable compaction planner, in both the pinned 9.0.0 +column and current upstream. That makes physical-locality maintenance — the very +subject of this research program — a semantic-correctness hazard rather than a pure +performance concern, and it argues **against** the program's attractive hypothesis +that compaction boundaries, parity groups, and query locality can share one grouping +without an explicit, storage-independent order key. The deterministic-replay +literature already gives the answer (a monotonic log coordinate); this design +substituted physical order for it. + +On the epistemic side the verdict is **not proven, and not currently provable**: +`deinterlace` has no production caller and no production `DeinterlaceRow`, so +STRICT/AWARE/RETRO behaviour today is a statement about a test struct — and the one +shape most likely to be reached by accident, `QueryReference::default()`, disables the +no-hindsight guarantee entirely by making `ref_version = u64::MAX` admit every row as +`Contemporary`. **F-HINDSIGHT-2 (frontier invariance) is the experiment I would run +first on that side**, because it is the only design here that tests the sealed-read +horizon from the reader's side, where the guarantee is actually consumed; and +**F-PHYS-ORDER** is the one I would run first overall. + +Two things I will not let the program overclaim. First, the brief's +`T_now × A_last × W_write` shorthand **is not the implementation** — `classify` is a +two-clock comparison plus a mode, with no awareness clock at all; experiments built on +the shorthand will test a machine that does not exist. Second, the temporal model is +**bitemporal as-of querying with a policy layer on top**, well covered by the XTDB / +bitemporality literature; the plausibly-new part is the two-layer split (per-owner +causal deinterlacing separate from epistemic projection), and I explicitly decline to +grade that NOVEL because I did not run the stream-processing prior-art search it needs. diff --git a/docs/lotus/rp-seal-v1/D3.md b/docs/lotus/rp-seal-v1/D3.md new file mode 100644 index 00000000..0098147f --- /dev/null +++ b/docs/lotus/rp-seal-v1/D3.md @@ -0,0 +1,514 @@ +# D3 — Temporal-Literature Scout: HLC, Bitemporal Databases, Retroactivity, and Modern Time-Travel vs. the Project's Actual `temporal.rs` + +**Cell:** D3 (Domain D — temporal deinterlacing) · **Role:** TEMPORAL-LITERATURE SCOUT +**Posture:** research only — no code was written or run; `cargo` was never invoked. + +--- + +## SOURCE ARCHAEOLOGY + +### 1. The project's own temporal doctrine, read literally + +`/home/user/OGAR/docs/TEMPORAL-TIME-TRAVEL.md` at the pinned commit +`386a6fd848334b1d880c8408b3810f045d135cfe` (fetched via `git show`, not the +working tree, per instructions) is a **"CARVED v0" boundary record**, not an +implementation doc. Its content, verbatim in substance: + +- A Python "temporal-epistemology framework" + (`epistemology.py`/`detector.py`/`awareness.py`/`hydration.py`) is mapped + onto Lance versions via a table: `KnowledgeItem.created_at` → the Lance + version a row landed at; `KnowledgeItem.knowable_from` → the Lance version + at which the row's **class** was registered; `KnowledgeHorizon` (at time T) + → `dataset.checkout_version(V_ref)`; `TemporalStatus.{CONTEMPORARY, + ANACHRONISTIC, SPOILER}`; `EpistemicMode.{STRICT, AWARE, RETRO}` as a + **planner-level query annotation**, not storage. +- It explicitly proposes, but does **not build**, a + `QueryReference { ref_version: u64, mode: STRICT|AWARE|RETRO, rung: u8 }` + annotation, and a **cross-server hindsight** extension: "a hybrid logical + clock (HLC) per writer stamps `(server_id, local_lance_version, hlc_tick)` + on the `CognitiveEventRow`; multi-server reads sort by `hlc_tick` for a + deterministic causal-time ordering." It flags, as **Decision #4, surfaced + not blocking**, that `ActionInvocation.emitted_at_millis` is plain + wall-clock `i64` and is *not* causally ordered across servers — an HLC + variant is deferred until cross-server hindsight is a real workload. +- It is explicit that OGAR does **not** own any of this — "OGAR does NOT + build `QueryReference` / `EpistemicMode` / the planner filter... nor the + HLC stamp." Both are named as another session's (`lance-graph-planner` / + `lance-graph-callcenter`) responsibility. + +**This means the task prompt's shorthand — "T_now × A_last × W_write with an +interlacing window resembling A_last < W_write ≤ T_now" — is not literal +project vocabulary.** No field named `A_last`, `W_write`, or a +current-vs-write-vs-awareness triple exists in the OGAR doc or (see below) +in the code that implements it. The doc's real vocabulary is +`(ref_version, hlc_tick, mode, rung)` read against `(row_version, +knowable_from)`. I investigated the actual implementation rather than +relying on the shorthand, as instructed. + +### 2. The implementation that carried the doctrine forward (read as source, not as prior-session narrative) + +`/home/user/lance-graph/crates/lance-graph-planner/src/temporal.rs` (870 +lines) **is** the promised planner-layer module — its own module doc opens +"Temporal epistemology + deinterlacing — query-time, planner-layer policy," +directly continuing the OGAR doc's boundary. Key types, read from source: + +```rust +pub type LanceVersion = u64; + +pub enum EpistemicMode { Strict, Aware, Retro } +impl EpistemicMode { + pub fn for_rung(rung: u8) -> Self { + match rung { 0..=4 => Strict, 5..=8 => Aware, _ => Retro } + } + pub fn admits(self, status: TemporalStatus) -> bool { /* Contemporary always; + Anachronistic only Aware|Retro; Spoiler only Retro; Unknowable never */ } +} + +pub enum TemporalStatus { Contemporary, Anachronistic, Spoiler, Unknowable } + +pub struct QueryReference { + pub server_id: u16, + pub ref_version: LanceVersion, // "KnowledgeHorizon" — u64::MAX = latest + pub hlc_tick: Option, // None single-server; cross-server deferred + pub mode: EpistemicMode, + pub rung: u8, +} + +pub fn classify(row_version: LanceVersion, knowable_from: LanceVersion, + v_ref: &QueryReference) -> TemporalStatus { + if knowable_from > v_ref.ref_version { Unknowable } + else if row_version <= v_ref.ref_version { Contemporary } + else if v_ref.mode == Retro { Spoiler } + else { Anachronistic } +} +``` + +`deinterlace(rows, v_ref, deps)` +filters rows to those that are **both** time-admitted (`classify` + +`EpistemicMode::admits`) and data-ready (a `DependsClosure` trait, currently +only implemented trivially as `NoDeps`), then **sorts** the survivors by +`hlc_tick.unwrap_or(lance_version)` — an explicit, code-commented decision +("Codex P2 on #468") to fall back to the row's own version rather than `0`, +so a missing-HLC row does not get forced ahead of every HLC-bearing row. + +A second axis — `local_trajectories`/`LocalCausalRow` — reconstructs each +**owner's own** causal chain out of a globally-interleaved durable log (crash +recovery), which the module doc labels "Layer 1 — CAUSAL deinterlacing," kept +explicitly distinct from "Layer 2 — EPISTEMIC projection" (`classify`/ +`deinterlace`). This two-layer split (who-wrote-it vs. what-may-a-reader-see) +is architecturally significant and is addressed in SURPRISES. + +A dedicated regression test, `no_hindsight_streamed_known_game` (lines +754–843), models a 10-ply "known game" as a Lance version stream and proves, +for three representative present-readers (`v_ref ∈ {2,5,8}`), that +`EpistemicMode::Strict` structurally excludes every future ply +(`TemporalStatus::Anachronistic`, refused by `admits`), while a `Retro` +reader (rung ≥ 9) at the identical `v_ref` sees the same future row +classified `Spoiler` and **admitted** — i.e., the same physical row carries a +different label depending on whether the peek was structural leakage or an +opted-in horizon break. The test's own doc comment cross-references +`AdaWorldAPI/stockfish-rs examples/hindsight_stream.rs` (`D-SF-HINDSIGHT-1`) +as consuming "the zero-dep `TemporalPov` mirror of this exact machinery" +against real lichess games — evidence the abstraction is meant to travel +beyond this one repo. + +### 3. What the code does **not** yet contain + +- No two-component `(l, c)` HLC struct, no bounded-counter update rule + (Kulkarni et al.'s Algorithm 1), no physical-clock read. `hlc_tick` is a + bare `Option` placeholder; the doc-comment is candid that it exists + only so "the cross-server policy... wakes them up with no breaking + signature change." The HLC *algorithm* itself is not implemented here. +- No retroactive-write path. Every row arrives at the write frontier; + `classify`/`deinterlace` are pure *read-side* filters over an + already-committed version stream. There is no operation that inserts an + update **into the past** of an already-sealed prior version and asks the + structure to recompute — the formal subject of retroactive data + structures (§ below). +- `deinterlace` is a **linear scan + sort** over the candidate row set per + query — no interval index, no monotone "stable frontier" variable + analogous to GentleRain's GST or CausalSpartanX's DSV, no reverse-log + early-termination analogous to XTDB's resolution algorithm (§ below). + +--- + +## MECHANISM + +Stated precisely, using the vocabulary actually in the code (not the task +prompt's shorthand): a reader is a `(ref_version, hlc_tick, mode, rung)` +tuple. A row is a `(row_version, knowable_from, [hlc_tick])` tuple. `classify` +computes a 4-way visibility predicate that composes **two independent +temporal filters**: + +1. A **schema-registration gate**: `knowable_from > ref_version ⇒ Unknowable` + — the row's *class* did not exist yet at the reader's horizon, regardless + of the row's own version. +2. A **version-order gate**: `row_version ≤ ref_version ⇒ Contemporary`, + else `Spoiler` (if the reader's mode is `Retro`, i.e. an intentional + horizon break) or `Anachronistic` (a structural refusal for `Strict`/ + `Aware` readers with `Anachronistic` further split into "admitted" + (`Aware`) vs "refused" (`Strict`)). + +`deinterlace` then merge-sorts the admitted set by `hlc_tick` (falling back +to `lance_version`), producing the "standing wave" — the causally-coherent +per-reader projection the rest of the planner dispatches on. + +--- + +## EXPECTED BENEFIT + +If HLC is actually wired in (currently deferred): bounded-divergence-from- +physical-time ordering across independently-clocked writers (§ CausalSpartanX +below) without the write-delay pathology single-scalar physical clocks incur +under clock skew — directly relevant to the "4 independently-clocked +producers" framing in `temporal.rs`'s own module doc (lance / surrealql / +ractor / thinking). The `Strict`/`Aware`/`Retro` reader-selectable admission +control gives a single storage substrate multiple simultaneous "isolation +levels" per *reader competence* rather than per *transaction*, which none of +the classical literature surveyed offers as a first-class primitive (see +SURPRISES). + +## EXPECTED FAILURE + +1. **Linear-scan cost.** `deinterlace`'s filter+sort has no shortcut + structure; as row count and version count grow, every query pays O(n log + n) over the full candidate set, unlike GentleRain/CausalSpartanX's O(1) + GST/DSV check or XTDB's early-terminating reverse-log walk. +2. **No retroactive-write model.** If the wider seal/compaction research + programme (Domains A–C of this same study) ever wants to inject a + corrected/repaired symbol into an *already-sealed* prior version and have + downstream reads reflect it, `temporal.rs` currently has no vocabulary for + that operation at all — everything is append-only at the write frontier. + `Retro`/`Spoiler` is about a reader looking *forward* past its own + horizon, not about a writer reaching *backward* into a sealed past. +3. **Undefined behavior at the compaction/GC horizon.** Nothing in the read + code establishes what happens when a `Strict` reader's `ref_version` is + older than the oldest Lance version retained after compaction — the + classical MVCC "vacuum horizon" problem (§ PostgreSQL/InnoDB analogy, + not independently confirmed here — flagged as an open question, not a + verified defect, since `persist_sink.rs`/`cycle_driver.rs` were not + fully read in this session). + +--- + +## EXPERIMENT + +All four are **pre-registered designs**, not results. None require the +proposed architecture to be built; each is stated so a later session running +the actual code (or a toy harness reproducing its types) can execute it +without ambiguity. + +### E1 — HLC-absence cross-writer ordering fault +- **Metric:** correctness of causal order (does `deinterlace`'s merge-sort + ever place a row before a row it causally depends on?), plus write/read + amplification if a fix requires re-sorting. +- **Baseline:** current code — `hlc_tick: None` always, sort key falls back + to `lance_version`. +- **Control:** a toy harness feeding two independent monotonic + `lance_version` counters (simulating two Lance datasets/writers, as the + module doc's "4 independently-clocked producers" implies) into + `DeinterlaceRow` with `hlc_tick = None`, then checking whether any known + cross-writer causal pair (`B` was created causally after reading `A`) can + be assigned `lance_version(A) > lance_version(B)` such that `deinterlace` + orders `B` before `A`. +- **Kill condition:** if such an inversion is constructible (it should be, + by CausalSpartanX's own argument — raw per-writer version counters carry + no cross-writer causal information), that **confirms** the doc's own + stated rationale for eventually wiring HLC in; it does not falsify + anything currently claimed, since the doc already labels this deferred. + This experiment's value is to produce a concrete, reproducible inversion + case as a regression fixture for whichever session wires HLC. + +### E2 — Retroactivity overhead, if/when a "correct a sealed version" + operation is added +- **Metric:** index-remap CPU/bytes and compaction CPU/wall, as a function of + `m` (operations since the corrected point) and `n` (structure size) — the + program's own metric list. +- **Baseline:** append-only (current `temporal.rs` behavior — no retroactive + writes exist). +- **Control:** a hypothetical **partially retroactive** implementation + (corrections land in the past, only the *present* may be queried) vs. a + hypothetical **fully retroactive** one (any past state, post-correction, + is queryable) — the exact Demaine/Iacono/Langerman dichotomy. +- **Kill condition:** per Chen/Demaine/Gu/Vassilevska Williams/Xu/Yu + (arXiv:1804.06932, read in full), the *tight* achievable overhead for + general-purpose full retroactivity over partial is + `Θ(min(√m, n·log m))` up to `m^{o(1)}` factors, and — conditioned on any of + three standard fine-grained-complexity conjectures (a weak Circuit-SAT + hypothesis, online (min,+)-product, or 3-SUM) — this bound is **provably + not improvable** for a general data-structure problem. If a proposed + "cheap retroactive correction into a sealed erasure-coded version" design + claims sub-`√m` (or sub-`n·log m`) amortized cost for genuinely + *arbitrary* corrections, that is either (a) exploiting problem-specific + structure the general lower bound does not cover (a real, checkable + possible win — the design should say what property it exploits) or (b) an + error. This is the report's sharpest, most directly falsifiable claim to + carry into the consolidation pass. + +### E3 — Scan-vs-index cost of `deinterlace` under load +- **Metric:** point/neighborhood/random-take latency, cache misses, CPU, as + row count `n` and version count `m` scale. +- **Baseline:** current linear filter+sort. +- **Control:** (a) a GST/DSV-style single monotone "stable frontier" + variable per producer (GentleRain/CausalSpartanX-style — cheap but only + correct when visibility reduces to a scalar comparison, which it does + **not** here because of the `knowable_from` schema gate and the + 4-producer merge); (b) an interval-tree or reverse-log-with-early- + termination index (XTDB-style). +- **Kill condition:** if (a) cannot be made correct without re-deriving + something isomorphic to the `knowable_from` check per row (very likely, + since GST/DSV assume a single scalar dependency frontier and this system + has a second, class-registration-keyed frontier), that specifically + **disproves** "just add a GST" as an easy fix and redirects toward (b). + +### E4 — Session-guarantee / compaction-horizon fault at the read boundary +- **Metric:** stale-version retention bytes, restart recovery work, and a + binary correctness outcome (loud failure vs. silent wrong answer). +- **Baseline:** query at `ref_version` strictly within the currently + retained Lance version range. +- **Control:** query at a `ref_version` older than the oldest retained + version after a simulated compaction/GC pass. +- **Kill condition:** per Terry et al.'s per-key monotonic-read guarantee + (1994, formalized below) and the classical MVCC "vacuum horizon" problem, + a correct system must either retain enough history to serve the pinned + `ref_version` or **refuse the read explicitly**. Any code path that + silently substitutes a *different* (later) version's data for a + `ref_version` that has been vacuumed away would violate the very + no-hindsight guarantee `no_hindsight_streamed_known_game` proves in the + other direction (a `Strict` reader must never see the future) — its + temporal mirror-image (a pinned reader must never silently see an + arbitrary substitute past). Not evaluated against the actual + `persist_sink.rs`/compaction code in this session — flagged for a + mechanism-focused cell to check directly against source. + +--- + +## PRIOR ART + +Searches actually run this session (tool + query, verbatim or paraphrased): + +- `alphaXiv.get_paper_content` on `arxiv.org/abs/1812.07123` (CausalSpartanX) + — full text retrieved and read across ~700 lines (intro, background, + causal-consistency definitions, clock-anomaly analysis, full HLC recap + incl. Algorithm 1, ROTX protocol, impossibility results, related work, + conclusion). +- `alphaXiv.discover_papers` — keywords `["hybrid logical clocks", "causal + consistency", "snapshot isolation"]`, question "Hybrid logical clocks for + consistent snapshots in globally distributed databases; the Kulkarni et + al. HLC paper" — surfaced `1808.05698` (Session Guarantees with Raft and + HLC, same author group), which was then fetched and read for its formal + per-key session-guarantee definitions (monotonic-read, read-your-write, + monotonic-write, write-follows-read; reference [30] confirmed by direct + grep of the fetched text to be Terry, Demers, Petersen, Spreitzer, + Theimer, Welch, "Session Guarantees for Weakly Consistent Replicated + Data," PDIS 1994, pp. 140–149). +- `WebSearch` — `Kulkarni Demirbas "Logical Physical Clocks" "Consistent + Snapshots in Globally Distributed Databases" arxiv` — confirmed the + original HLC paper (Kulkarni, Demirbas, Madappa, Avva, Leone, OPODIS 2014, + pp. 17–32; SUNY Buffalo Tech. Rep. 2014-04) has **no arXiv mirror**; its + PDF (`cse.buffalo.edu/tech-reports/2014-04.pdf`) was fetched but both + `WebFetch` and the `Read` tool (no `pdftoppm`/poppler-utils in this + sandbox) failed to extract text from it. The paper's content was instead + recovered faithfully through CausalSpartanX's Section 4, which explicitly + states it "recall[s] HLCs from [21]" and reproduces the formal `⟨l.e, c.e⟩` + tuple definition and the full update algorithm (Algorithm 1) verbatim — + treated here as a faithful secondary transcription of the primary source, + not read from the primary PDF directly. This limitation is disclosed + rather than papered over. +- `WebSearch` — `"Logical Physical Clocks" Kulkarni Demirbas site:arxiv.org` + — confirmed again: no arXiv-hosted copy exists; surfaced adjacent HLC- + lineage papers (`2104.15099` "Achieving Causality with Physical Clocks," + `1808.05698` again). +- `WebSearch` — `Snodgrass Jensen bitemporal database transaction time valid + time consensus glossary temporal database concepts` — surfaced "A + Consensus Glossary of Temporal Database Concepts" (Jensen, Clifford, + Gadia, Segev, Snodgrass; ACM SIGMOD Record 23(1):52–64, 1994) and the + TSQL2/BCDM data model chapter; the glossary PDF itself + (`sigmodrecord.org/.../140979.140996.pdf`) failed to extract via + `WebFetch` (binary/compressed stream) — the valid-time/transaction-time/ + bitemporal definitions used in this report are instead cross-confirmed via + `WebFetch` on the Wikipedia "Temporal database" article, which quotes the + same canonical definitions and additionally names the SQL:2011 + standardization (application-time period tables, system-versioned tables, + system-versioned application-time period tables). +- `WebSearch` — `Demaine retroactive data structures partially retroactive + fully retroactive nonoblivious` — located the original Demaine/Iacono/ + Langerman SODA 2004 / ACM TALG 2007 paper (definitions of partial vs. full + retroactivity, and Tangwongsan's "non-oblivious" variant) and the 2018 + follow-up. +- `alphaXiv.get_paper_content` on `arxiv.org/abs/1804.06932` ("Nearly + Optimal Separation Between Partially And Fully Retroactive Data + Structures," Chen/Demaine/Gu/Vassilevska Williams/Xu/Yu) — full text + retrieved and read in full: the O(√m) partial→full transformation from the + original 2004 paper, the new `O(n·log m)` transformation, and the + conditional `Ω(min(√m, n)^{1-o(1)})` lower bounds under three fine-grained- + complexity conjectures (a weak Circuit-SAT hypothesis, online (min,+)- + product, 3-SUM), closing the gap to `Θ(min(√m, n·log m))` up to `m^{o(1)}`. +- `WebFetch` on `en.wikipedia.org/wiki/Temporal_database` — valid time / + transaction time / bitemporal definitions, the SQL:2011 standard names, + confirmation that Datomic/XTDB/Delta/Iceberg are *not* mentioned there + (so those had to be sourced separately). +- `WebSearch` — `Datomic bitemporal "as-of" "history" database time travel + valid time vs transaction time architecture` — established Datomic is + **unitemporal** (system/transaction time only, one immutable timeline per + entity) and that XTDB adds the second (valid/business/application-time) + axis Datomic lacks. +- `WebFetch` on `xtdb.com/blog/building-a-bitemp-index-2-resolution` — XTDB's + bitemporal index resolution algorithm: an append-only event log keyed on + `(system_from, valid_time range, doc)`, reverse-system-time playback that + terminates early for current-state queries, retroactive corrections + modeled as new immutable interval entries (never rewriting prior events), + avoiding the "2× storage amplification" of naive four-column bitemporal + schemas. +- `WebSearch` — `Apache Iceberg Delta Lake time travel snapshot metadata + "VERSION AS OF" "TIMESTAMP AS OF" implementation manifest` — Iceberg's + `metadata.json` → snapshot list → per-snapshot manifest-list → manifest + file architecture; confirmed time-travel is "read an older snapshot + instead of the current one: no special storage or infrastructure + required" (i.e., the storage-side mechanism is identical in shape to + Lance's own version-ladder `checkout_version`, which is precisely the + mapping the OGAR doc already draws). +- `WebSearch` — `Berenson "A Critique of ANSI SQL Isolation Levels" snapshot + isolation definition write skew` — confirmed the classical formal + definition of Snapshot Isolation (a transaction reads a consistent + snapshot as of its start, permits Write Skew where Repeatable Read does + not) — the paper Berenson/Bernstein/Gray/Melton/O'Neil/O'Neil, SIGMOD 1995 + (also mirrored as `arxiv.org/abs/cs/0701157`, not separately fetched in + full this session; used at citation-confirmation depth only). +- `WebSearch` — `Cheney Chiticariu Tan "Provenance in Databases" survey + why-provenance where-provenance how-provenance temporal` — confirmed + the survey (*Foundations and Trends in Databases* 1(4), 2009) and its + why/how/where taxonomy, including "how-provenance" via **provenance + semirings** — noted in SURPRISES as a direct structural echo of this + same workspace's own semiring inventory (`docs/SEMIRING_ALGEBRA_SURFACE.md` + lists 15 semirings across the stack, including a `TruthPropagating` + semiring with NARS deduction/revision — read as file content only, not as + a "prior session conclusion" under the independence rule, since it is the + project's own currently-committed architecture documentation, functionally + equivalent to reading the source). +- `Read`/`Grep` on `/home/user/lance-graph/crates/lance-graph-planner/src/ + temporal.rs` (full read of the module doc, all type/enum/struct + definitions, `classify`, `classify_ready`, `deinterlace`, and the + `no_hindsight_streamed_known_game` test with its own extensive doc + comments) — the primary grounding source for everything in MECHANISM and + the mapping table below. +- `Bash`/`Grep` across `lance-graph-planner/src/` and + `lance-graph-supervisor/src/` for the task prompt's literal shorthand + terms (`awareness_tick`, `write_tick`, `current_tick`, `last_awareness`, + `reference_horizon`, `hindsight`) — confirmed `hindsight` appears + (doc-comments and one test name) but the tick-triple vocabulary does not; + the real field names are `ref_version`/`hlc_tick`/`mode`/`rung` (§ SOURCE + ARCHAEOLOGY item 1). + +**Not done, and stated as such rather than guessed at:** a full read of +`persist_sink.rs` (1751 lines), `cycle_driver.rs` (2618 lines), or +`batch_writer.rs` — these are more properly a MECHANISM cell's territory, +and this report does not claim any finding about compaction/GC interacting +with the temporal read-horizon beyond flagging it as an open question in +EXPECTED FAILURE / E4. + +--- + +## SURPRISES + +1. **The "deinterlacing" framing is a video-editing metaphor for a + mechanism that is, underneath, MVCC snapshot-visibility filtering + composed with a bitemporal schema-registration gate — not a new + primitive.** `classify`'s `row_version ≤ ref_version ⇒ Contemporary` + test is structurally identical to Snapshot Isolation's definition + (Berenson et al. 1995: a transaction reads only versions committed + before its own start) and to the classical rollback-database "as of + transaction time T" read (Snodgrass/Jensen). The genuinely different + part — merge-sorting across `hlc_tick` from **multiple independently- + clocked producers**, and gating additionally on `knowable_from` (a + second, schema-level transaction-time axis, not the domain valid-time + axis bitemporal literature usually pairs with transaction-time) — is a + **composition** of two known techniques (HLC-ordered causal + consistency + bitemporal schema-evolution gating) applied to a novel + 4-producer setting, not a new mechanism in either half. + **Classification: KNOWN (each half) / TRANSFER (the specific + composition across 4 producers onto one Lance version ladder).** + +2. **`EpistemicMode::{Strict, Aware, Retro}` as a *reader-selectable, + rung-keyed* admission-control tier has no direct classical-DB analog + found in this search.** Classical isolation levels (Read Uncommitted / + Read Committed / Repeatable Read / Snapshot / Serializable) are + properties of a *transaction*, chosen once per transaction, uniform + for all data the transaction touches. XTDB's `history` API and + Datomic's `as-of`/`history` distinguish "current-state read" from + "explicit hindsight query" but do not gate the *same* query differently + by a numeric competence tier of the reader, nor do they have a + middle tier ("Aware" — may see hindsight but not intentional spoilers) + between the two. **Classification: NOVEL-CANDIDATE**, with the caveat + that the search performed (general web search + alphaXiv discovery for + "session guarantees," "bitemporal," "epistemic access control") did not + turn up a close match, but was not exhaustive of, e.g., the + differential-privacy or capability-security literatures, which were + out of this cell's scope. + +3. **The `Anachronistic` vs. `Spoiler` distinction is a genuine, testable + correctness property that the project's own test suite already proves + two-sided (both "can it fire" and "can it stay silent," per this + workspace's own falsifiability doctrine) — and it maps cleanly onto + the classical MVCC "read-only anomaly" vocabulary once you see it.** + `Spoiler` only ever appears when the reader's own `mode` is already + `Retro` (an *opted-in* horizon break); the identical future row under + `Strict` is `Anachronistic` and refused. This is the same distinction + XTDB draws between a normal `as-of` query (never sees the future by + construction, since `system_from` bounds the log walk) and an explicit + `history`/audit query (deliberately walks past that bound) — but XTDB + does not expose this as a *per-row label* the way `temporal.rs` does. + **Classification: TRANSFER** (the underlying distinction — accidental + vs. intentional horizon break — exists in XTDB's API surface; its + reification as a per-row enum value is this project's own choice, not + found elsewhere in the literature searched). + +4. **The retroactive-data-structures literature's tight `Θ(min(√m, + n·log m))` bound is a ready-made, directly citable kill condition for + any future claim that this project's seal/compaction design offers + "cheap" corrections into an already-sealed prior version** — and + the project's own `temporal.rs`, as it stands today, has **no** + retroactive-write path at all, so this bound is not yet being tested + by anything in the current code; it is a forward-looking gate for + whichever cell/wave adds one. **Classification: KNOWN** (the bound + itself, Chen/Demaine/Gu/Vassilevska Williams/Xu/Yu 2018), flagged here + as a **TRANSFER** opportunity into this project's design space — + not yet exercised by anything that exists. + +5. **A provenance-semiring echo:** Cheney/Chiticariu/Tan's how-provenance + is formalized as evaluating relational-algebra expressions over a + commutative semiring of provenance tokens (Green/Karvounarakis/Tannen's + provenance semirings) — and this workspace's own `docs/ + SEMIRING_ALGEBRA_SURFACE.md` independently lists 15 semirings across + the stack (7 HDR semirings, an SPO truth semiring, a palette semiring, + a `TruthPropagating` planner semiring with NARS deduction/revision, an + attention semiring, 5 contract semiring choices) — none of them, + on the file's own description, framed as *provenance* semirings, but + structurally in the same algebraic family (idempotent/commutative + semiring-valued annotation propagation). This is a plausible, not yet + verified, bridge: could the project's existing `TruthPropagating` + semiring be *read* as a provenance semiring for temporal + contemporary/anachronistic/spoiler/unknowable classification, giving + `classify`'s 4-way output a principled algebraic composition rule + across joins instead of the current per-row `if`/`else if` chain? + **Classification: NOVEL-CANDIDATE** (no evidence in the code that this + connection has been made; flagged as a genuine research question for + Domain D's consolidation, not asserted as already true). + +--- + +## VERDICT + +The project's actual `temporal.rs` implements a coherent, testable, +already-partially-self-falsifying (`no_hindsight_streamed_known_game`) +**read-side epistemic visibility filter** that composes two known +techniques — HLC-style causal ordering (currently only reserved as an +`Option` slot, not yet algorithmically implemented) and bitemporal +schema-registration gating — across an unusually wide fan-in of four +independently-clocked producers; it has **no retroactive-write mechanism at +all**, so the retroactive-data-structures literature's tight cost bounds are +a forward-looking falsifier for later waves rather than a live constraint +today, and the `Strict`/`Aware`/`Retro` per-reader-rung admission tiering +appears to be a genuine departure from anything found in the classical or +modern-systems literature surveyed. + diff --git a/docs/lotus/rp-seal-v1/E1.md b/docs/lotus/rp-seal-v1/E1.md new file mode 100644 index 00000000..fa1ba362 --- /dev/null +++ b/docs/lotus/rp-seal-v1/E1.md @@ -0,0 +1,574 @@ +# E1 — Benchmark Harness Family for Cascade-Threshold SoA Compaction (BUILDER) + +## SOURCE ARCHAEOLOGY + +**Baseline geometry is already load-bearing production code, not a proposal.** +`lance-graph-supervisor/tests/probe_ignition_64k.rs:66-69` and +`lance-graph-supervisor/examples/measure_wal_curve.rs:115-119` both pin +`FLEET_OWNERS = 65_536`, `CANONICAL_ROW_BYTES = 512`, +`CANONICAL_FRAME_BYTES = 65_536 × 512 = 32 MiB` — exactly the cell brief's +"cycle = 4096 fields × 8 KiB = 32 MiB = 65,536 rows" (4096 fields × 16 +rows/field × 512 B/row = 32 MiB, consistent to the byte). + +**A five-axis measurement harness already exists and is the single richest +piece of prior art for this cell.** `measure_wal_curve.rs` (3,558 lines, +`cargo run --release -p lance-graph-supervisor --features cycle-driver +--example measure_wal_curve`) is a real, runnable binary that: + +- Reads `/proc/self/status` for `VmHWM`/`VmRSS` and `/proc/self/stat` fields + 10/12 for minor/major page faults (`measure_wal_curve.rs:245-295`, + function `read_proc_status` / `read_proc_stat_faults`). **No hardware + performance counters are read anywhere in this repo** — confirmed by + `grep -rn "perf_event\|dhat\|jemalloc\|cachegrind" crates --include="*.rs"` + returning zero hits outside doc-comment mentions of an "allocator" concept. + Page faults are the current cache/memory-pressure proxy; they are **not** + L1/L2/LLC cache misses. +- Already documents and fixed a real methodological trap I would otherwise + repeat: `VmHWM` is process-monotonic, so differencing two `VmHWM` readings + inside one process yields the historical maximum twice, not two footprints + — "An earlier revision differenced VmHWM and printed a NEGATIVE + 'overhead' … That figure is retracted, not reported" (`:250-252`, + `:3450-3454`). Every per-arm memory number in that file is a **VmRSS + delta**, never a VmHWM difference. +- Already runs exactly the loop-order / physical-layout experiment this + cell is asked to design, at a coarser grain: the **M-arm** + (`:2671-2900`, "Morton reorder inserted before the seal") compares a + natural (arrival-order) write pipeline against one with a Morton reorder + step inserted before seal, times `cast/collect/reorder/seal/write/sync` + phases separately, asserts **semantic digest identity** between the two + physical orders (`:2681-2691`, `assert_eq!(digest_natural, digest_morton, + …)` — the exact correctness pin any reordering experiment needs), and + reports a **pre-registered SUM verdict** — reorder cost vs. downstream + savings, never the downstream gain alone (`:2836-2854`). The **O-arm** + (`:2944-3300`) asks a directly analogous question to this cell's + loop-interchange question: does ordering sourced *after* seal (today's + pipeline) differ from ordering sourced *before* seal from a + `local_trajectories` (temporal.rs) replay — with a compile-time + "firewall" self-scan (`FIREWALL-START`/`FIREWALL-END` sentinels, + `:2965-2995`) proving the alternate arm's ordering source is not secretly + borrowed from the arm it's being compared against. +- Ships a WAL-curve segment-size sweep (`SEGMENT_TABLE`, + `:1066-1090`: 2,048 / 4,096 / 8,192 / 16,384 / 65,536 rows/segment) with + an **honesty discipline** worth copying verbatim: the plateau is reported + as a "descriptive knee, never PASS/KILL" (`:3474-3475` in the run() + docstring; `:3491-3520` in the closing summary), and when p95/median + spread exceeds a ceiling across repeats the file prints "NOT MEASURABLE + ON THIS HOST" rather than fabricating a knee (`:3499-3506`). +- Establishes the CSV `Row` schema (`:320-357`) already carrying most of + the metric list this cell needs: `segment_rows`/`segment_bytes`, + `wal_syscalls`, `fsync_calls`, `dataset_versions`, `peak_rss_bytes`, + `minor_faults`/`major_faults`, `context_switches`, `result_digest` + (FNV1a-64 semantic digest), `morton_reorder_ns`. + +**Compaction, fragment reuse, and index-remap are real, shipped Lance 9.0.0 +mechanisms — not something I need to invent an analogue for.** Verified in +the exact-pinned tree at `/tmp/sources/lance-9` (tag v9.0.0, matching +`Cargo.lock`'s checksum), never in current upstream without checking the +column: + +- `lance/src/dataset/optimize.rs` (8,509 lines). Module doc (`:5-23`) + states the two compaction triggers Lance actually uses: a fragment with + fewer rows than `target_rows_per_fragment`, or a fragment whose deleted + fraction exceeds `materialize_deletions_threshold` — and explicitly names + the failure mode this cell's H2 targets: *"if a series of small streaming + appends are performed, eventually there will be a large number of small + files"* (`:9-11`). +- `IndexRemapMode` (`:162-180`) is a real two-way cost tradeoff already + shipped: `Compact` — "store a compact remap and compute row-address + mappings during lookup… uses less memory, but each lookup does extra + bitmap/range computation", vs `Direct` (the default) — "store the full + row-address remap in memory for fast direct lookups… uses more peak + memory". This is directly the "index remap bytes/CPU" metric pair the + brief asks for, already named as a memory-vs-CPU knob by Lance itself. +- `lance-core/src/utils/row_addr_remap.rs` (413 lines) implements + `RowAddrRemap::{compact, direct}` precisely: compact mode is + `O(#fragments)` via `old_fragment_id -> (old_offsets: RoaringBitmap, + old_rows_before)` plus ordered new-fragment ranges, with an explicit + ordering precondition documented in the module doc (`:1-45`): *"Compact + remap does not store each old-to-new row mapping… requires the + reader-to-writer pipeline to preserve row order."* Direct mode is a plain + `HashMap>` (`:73`). Both are real, benchmarkable code + paths, not something to mock. +- `CompactionMetrics` (`:637-646`) is minimal by design — only + `fragments_removed/added`, `files_removed/added` — **no byte-level or + CPU-time field at all**. Any byte/CPU/cache metric for compaction has to + come from instrumenting the call site, not from Lance's own return type. +- `manifest.rs:492-495`, `uses_stable_row_ids()` reads a reader-feature + flag bit. `optimize.rs`'s planner explicitly refuses + `defer_index_remap=true` on a stable-row-id dataset (`:701-704` in the + planner impl) — stable row IDs and index remap are two genuinely + different cost regimes for the same compaction event, and both are real, + flaggable dataset configurations. +- **Reed-Solomon / erasure coding does not exist anywhere in Lance.** + `grep -rl "reed.solomon\|erasure\|ReedSolomon" /tmp/sources/lance-9/rust + /tmp/sources/lance-main/rust` returns zero hits in **both** the pinned + 9.0.0 tree and the current upstream clone. Lance's durability model is + object-storage redundancy (S3/GCS-class), not application-level erasure + coding. Any "repair locality" experiment against a real erasure code is + therefore necessarily a **standalone probe outside Lance**, not a Lance + benchmark (see H5 below — this is stated as a finding, not assumed). + +**The temporal/deinterlace layer this cell's brief paraphrases as "T_now × +A_last × W_write" is real code, but the paraphrase needs a correction.** +`lance-graph-planner/src/temporal.rs:185-206` (`classify`) takes exactly +two version numbers — `row_version` (write tick) and `v_ref.ref_version` +(the reader's pinned horizon, i.e. "last awareness tick") — plus a mode +(`EpistemicMode::{Strict, Aware, Retro}`, `:79-86`, matching STRICT/AWARE/ +RETRO in the brief verbatim) and a `knowable_from` bound. **There is no +third "current version" argument** — `QueryReference::default()` +(`:150-161`) sets `ref_version = u64::MAX` ("latest") when a caller wants +"now"; the function itself never distinguishes "the dataset's live head" +from "the reader's pinned horizon" as two separate inputs. OGAR's +`docs/TEMPORAL-TIME-TRAVEL.md` at the pinned commit `386a6fd8` (§2, table) +independently derived the same two-input model (`ref_version` / +`knowable_from` / mode) from a Python temporal-epistemology framework +before the Rust code existed, and both sources agree the HLC coordinate +(`hlc_tick: Option`, `temporal.rs:143-145`) is type-visible but +**unwired** — `server_id` and `hlc_tick` are carried for a future +cross-server policy that does not exist yet. This is important for +E1 specifically: the interlacing-window shorthand `A_last < W_write <= +T_now` in the brief is a reasonable gloss for intuition but is **not** +what `classify()` implements, and a harness that hard-codes three distinct +version pointers would be testing a mechanism the code doesn't have. + +**Write-side canonical ordering is a real, already-shipped discipline** +(`persist_sink.rs:133`, `order_cycle_stably` — stable sort of a cycle's +casts by a caller-supplied key, here `stream_position`, before the single +WAL append; module doc `:27-33` explains this is what makes the durable +image canonical regardless of arrival order). This is the natural +"canonical sort" arm for H3 below and is not hypothetical — it runs on +every `run_cycle` call today. + +**ndarray already reasons in 64-byte cache-line tiles for HPC kernels** +(`ndarray/src/hpc/amx_matmul.rs:11,25,40-41,62,111,351-390`, AMX tile +config; `ndarray/src/hpc/fingerprint.rs:205-207,645`, 64-byte SIMD chunk +iteration) — this is local precedent that the codebase already thinks in +cache-line/tile-sized units for compute kernels, but **no existing +`ndarray` bench (`benches/bench1.rs`, `crates/p64/benches/p64_bench.rs`, +`crates/fractal/benches/manifold_bench.rs`) measures cache misses** — they +are all Criterion wall-clock benches (`black_box`/`criterion_group` +idiom, confirmed by grep). `simd_soa.rs` (`ndarray/src/simd_soa.rs`, +539 lines) is a layout-only SoA carrier (`MultiLaneColumn`, `Arc<[u8]>`, +64-byte chunk iteration) with **no bench file at all** — a real gap: the +one module in the workspace literally named for this cell's domain has +zero measurement coverage today. + +**No global-allocator instrumentation exists anywhere in the workspace.** +`grep -rln "impl.*GlobalAlloc" /home/user/ndarray /home/user/lance-graph` +(excluding `target/`) returns nothing. An allocation-count/byte harness +component is a genuine gap to design, not something to wire up from an +existing crate. + +## MECHANISM (design only — nothing below is implemented) + +### Shared substrate (built once, reused by H1–H5) + +- **Fixture builder**: one 32 MiB cycle = 65,536 rows × 512 B, addressed as + 4,096 fields (16 rows × 512 B = 8 KiB each). Fields are given a 2-D + coordinate in a 64×64 grid (`4096 = 64×64`), matching the brief's + "2-D hierarchy 4/16/64/256/1024/4096 (2×2..64×64)". A `FieldOrder` enum + parameterizes physical placement: `RowMajor`, `Morton` (reusing the + M-arm's proven reorder primitive rather than re-deriving bit-interleave), + `Scatter` (a fixed pseudo-random permutation, seeded, reproducible), + `IndexRemapped` (built by applying a synthetic `RowAddrRemap` after an + initial `RowMajor` write, simulating one compaction event). +- **Cascade grouping**: for threshold `T ∈ {4,16,64,256,1024,4096}`, fields + are partitioned into contiguous (in `FieldOrder`) blocks of `T`; the + inverse-reduction ladder `4096→1024→256→64→16→4→1` is realized as + successive 4-way grouping passes over these blocks — i.e. the harness + builds the pyramid bottom-up once per fixture and every H1/H4 cell reads + a level of it, rather than rebuilding per cell. +- **Metrics struct** (a superset of `measure_wal_curve.rs`'s `Row`, + additive — no existing field renamed): adds + `alloc_count`, `alloc_bytes`, `dealloc_bytes` (from a `#[global_allocator]` + wrapper around `System` with per-arm `AtomicU64` counters, reset + immediately before and read immediately after the timed region — **never** + inside the same pass as digest verification, see EXPECTED FAILURE), + `cache_references`, `cache_misses`, `llc_load_misses`, `instructions`, + `branch_misses` (see measurement method below), `fragment_count`, + `files_touched`, `pages_touched_4k` (distinct 4 KiB-aligned byte ranges + read, computed from the fixture's own address log — no OS support + needed), `index_remap_build_ns`, `index_remap_bytes`, + `index_remap_lookup_ns_per_row`, `compaction_cpu_ns`, `compaction_wall_ns`, + `read_amplification` (physical bytes read ÷ logical bytes needed for the + query under test), `write_amplification` (physical bytes written ÷ + logical bytes changed), `result_digest` (FNV1a-64, reusing + `measure_wal_curve.rs`'s own `fnv1a64` + `semantic_digest` idiom + verbatim). +- **Repeat / stability protocol**: median-of-7 for wall-clock timings + (Criterion-style, matching `ndarray/benches/bench1.rs`'s existing + `black_box` idiom for anything JIT/inlining-sensitive); median-of-5 with + an explicit p95/median spread ceiling for anything reading `/proc` or + hardware counters, printing "NOT MEASURABLE ON THIS HOST" rather than a + fabricated number when the ceiling is exceeded — this is a direct reuse + of `measure_wal_curve.rs:3499-3520`'s own gate, not a new convention. +- **Correctness pin**: every physical-layout arm (H1, H3, H4) asserts + `result_digest` equality against a `RowMajor`, un-reordered reference + build of the *same logical content*, exactly as the M-arm's + `assert_eq!(digest_natural, digest_morton, "M-arm can-fire: …")` + (`measure_wal_curve.rs:2681-2691`). A layout change that is not + digest-identical is a correctness bug, not a performance data point, and + the harness must fail loudly rather than report a number for it. + +### H1 — Row/field/cascade geometry sweep + +Sweeps working-set size by varying the cascade threshold `T ∈ +{4,16,64,256,1024,4096}` (bytes-per-unit `= T × 8 KiB`: 32 KiB, 128 KiB, +512 KiB, 2 MiB, 8 MiB, 32 MiB), crossed with `FieldOrder ∈ {RowMajor, +Morton}`. For each `(T, order)` cell, treats one `T`-field block as one +flush/seal/index unit and measures bytes touched, alloc_count, and cache +misses per logical row processed. + +- **Metric set**: `cache_misses`/`instructions` per row, `alloc_count`, + `pages_touched_4k`, wall_ns/row. +- **Measurement method**: primary = `perf-event` crate + (`Hardware::CACHE_MISSES`/`CACHE_REFERENCES`/`INSTRUCTIONS`) wrapping the + timed region per repeat, gated by the spread-ceiling protocol above. + **Constraint stated honestly**: `perf_event_open` requires either + `CAP_PERFMON`/root or a permissive + `/proc/sys/kernel/perf_event_paranoid`; this could not be verified as + available inside the current research sandbox (no benchmark was run — + BUILDER role does not execute), so the design's fallback path is not + optional decoration. Fallback = `iai-callgrind` (Cachegrind-simulated + `Dr`/`D1mr`/`DLmr`), which is deterministic across hosts and CI-safe — + the T-sweep specifically should be run under **both** where possible, + since simulated and real counters can diverge (simulated caches use an + idealized model; real caches have prefetchers, non-inclusive LLCs, etc.) + and a divergence between them is itself a finding. +- **Baseline**: `RowMajor`, `T = 4096` — today's shipped behavior (one + seal for the whole 65,536-row cycle, `persist_sink.rs`'s "cycle" level). +- **Control**: a page-cache-cold run (fresh process per repeat) alongside + the page-cache-warm run (repeats within one process), since + `measure_wal_curve.rs`'s own WAL-curve section already found the write + phase host-noise-dominated by "page-cache / dirty-writeback state, not + segment size" (`:3499-3502`) — the same confound plausibly applies to + read-side cache-miss measurement and must be controlled for, not + assumed away. +- **Kill condition**: if `cache_misses`/row is flat (within the measured + noise ceiling) across all six `T` values for both orders, the cascade + threshold geometry carries no cache-locality signal at this row/field + size on the tested host, and any downstream hierarchy built on the + 4/16/64/256/1024/4096 ladder for *performance* reasons (as opposed to + semantic/addressing reasons) is falsified. + +### H2 — First-come chunk vs whole-64K batch processing + +Directly operationalizes `optimize.rs`'s own documented failure mode +(small streaming appends → many small fragments) and +`persist_sink.rs`'s "cast ⊂ chunk ⊂ cycle" architecture. Two arms over `N` += 200 cycles: **FIRST-COME** flushes/seals as soon as a chunk of `C` rows +has arrived (`C` swept over 512 / 8,192 / 65,536 = 1 row / 1 field / 1 +full cycle); **WHOLE-BATCH** accumulates the full 65,536-row cycle before +any seal (today's shipped `run_cycle`). + +- **Metric set**: `fragment_count` after N cycles (from Lance's real + `Dataset::get_fragments().len()` and `CompactionMetrics`, zero + instrumentation needed — this is the same quantity the `optimize.rs` + module doctest itself asserts, `:47-56`), `wal_syscalls`, `fsync_calls` + (already-shipped fields, `measure_wal_curve.rs:344-345`), point-read + p50/p99 latency for a single owner's local trajectory (the exact query + `local_trajectories`/`deinterlace` in `temporal.rs` perform) measured + **with compaction OFF for the entire soak window**. +- **Baselines**: `C = 65,536` (whole-batch, today's shipped path) is the + reference row in every table, matching the WAL-curve section's own + "W0-current numbers are beside it in the CSV, never substituted" + discipline (`:3512-3520`). +- **Control (compaction disabled for a soak window)**: run the full N=200 + cycles with `compact_files` never called; measure fragment-count growth + and point-read latency degradation as a curve over cycles, not a single + endpoint number. Then flip compaction on once and measure the one-shot + repair cost (`compaction_cpu_ns`, `compaction_wall_ns`, + `index_remap_build_ns`) to restore fragment count to baseline — this is + the direct measurement for the cross-domain question "is compaction + repair of physical locality," continued in H5. +- **Kill condition**: if small-`C` first-come flushing shows no measurable + fragment-count blowup relative to whole-batch (i.e. Lance's own + `Scanner::batch_size`/`io_buffer_size` internal buffering, or + `target_rows_per_fragment`'s default of 1,048,576 rows, already absorbs + the effect at this scale), the chunk-vs-batch axis is not a locality + lever at 65,536-row cycles and any hierarchy justified by it needs a + different justification (semantic grouping, not I/O locality). + +### H3 — Canonical sort vs scatter vs index-remapped random placement + +Three arms, reusing the M-arm's proven natural/reordered comparison +machinery rather than re-deriving it: + +- **Arm A — CANONICAL**: `order_cycle_stably` by `stream_position` + (`persist_sink.rs:133`, today's shipped path). +- **Arm B — SCATTER**: no reorder; rows land in arrival (CPU-completion) + order. This is exactly the M-arm's existing "natural" pipeline + (`measure_wal_curve.rs`'s `run_m_arm_pipeline(false, …)`) — H3-B + extends that harness with the neighborhood-scatter metric below rather + than re-implementing the pipeline. +- **Arm C — INDEX-REMAPPED**: simulate one compaction event that moves + rows to a different physical fragment layout, forcing a real + `RowAddrRemap` build, under **both** `IndexRemapMode::Compact` and + `IndexRemapMode::Direct` (`optimize.rs:162-180`), swept over what + fraction of the 65,536-row cycle is rewritten (10% / 50% / 100%). +- **Metric set**: A/B — read latency and `pages_touched_4k` for the local- + trajectory query (the "neighborhood scatter" proxy: fraction of a + query's rows landing in physically-adjacent 4 KiB ranges). C — + `index_remap_build_ns`, `index_remap_bytes` (Direct mode's + `HashMap>` is measured via the allocator counters, never + computed from a per-entry byte-size assumption, per this workspace's own + falsifiability rule against unverified claims), `index_remap_lookup_ns` + per row, at both remap modes and all three rewrite fractions — a 2×3 + grid directly against Lance's own documented memory-vs-CPU tradeoff. + Also run once more with `uses_stable_row_ids()` = true (the manifest + flag, `manifest.rs:492-495`), where `optimize.rs` itself refuses + `defer_index_remap` (`:701-704`) — i.e. measure the SAME compaction + event under the regime where row identity is decoupled from physical + address, as the natural fourth arm this flag makes available for free. +- **Baseline**: Arm A. +- **Kill condition**: if Arm B shows no measurable read-latency or + page-touch degradation vs. Arm A at this row/page size, the write-side + canonical-ordering discipline `persist_sink.rs`'s own module doc leads + with as an architectural rationale ("physical race safety") has no + measured **performance** payoff at this geometry — its justification + would then rest entirely on the *epistemic* correctness argument the + same doc makes (never on a locality argument), and that should be + stated plainly rather than left implied. + +### H4 — Loop-interchange: cycle-major vs row-tile-major at pipeline depth k + +The pre-registered decision procedure for "does the L1-resident 8 KiB +petal-tile amortization claim survive measurement." + +- **CYCLE-MAJOR** (today's only shipped order — there is no multi-cycle + batching anywhere in `persist_sink.rs`'s three-level cast⊂chunk⊂cycle + model): `for c in 0..k { for f in 0..4096 { process(cycle[c], field[f]) + } }`. Working set per inner-loop pass = 32 MiB (the whole cycle). +- **ROW-TILE-MAJOR**: `for tile in tiles(4096, T) { for c in 0..k { + process(cycle[c], tile) } }`. Working set per inner-loop pass = `T × + 8 KiB × k` bytes — the same field-tile is revisited across `k` cycles + before moving to the next tile. +- **Grid**: `T ∈ {4,16,64,256,1024,4096}` (H1's cascade ladder, reused, not + reinvented) `× k ∈ {1,4,16}` = 18 cells. `k=1` collapses cycle-major and + tile-major to the same computation for every `T` — a built-in sanity + check (both orders MUST produce byte-identical digests and near-identical + metrics at `k=1`; a divergence there is a harness bug, not a finding). + This crosses realistic L1 (32–64 KiB)/L2 (256 KiB–2 MiB)/LLC (8–32 MiB + typical single-socket) boundaries systematically: `T=4,k=4→128 KiB` + (L2), `T=4,k=16→512 KiB` (L2 boundary), `T=64,k=1→512 KiB`, `T=1024, + k=1→8 MiB` (LLC boundary), `T=4096,k=1→32 MiB` (exceeds typical LLC). +- **Metric set**: `cache_misses`/row (primary), `instructions`/row (to + separate "genuinely fewer misses" from "just fewer instructions" — + the classic loop-tiling confound), wall_ns/row. +- **Measurement method**: same primary/fallback split as H1 + (`perf-event` / `iai-callgrind`), because this is the noisiest + measurement in the family and needs the strongest stability gate. +- **Baseline**: cycle-major, `k=1` (today's shipped per-cycle processing — + literally the only mode that exists in the codebase today; every other + cell in this grid is hypothetical). +- **Control**: CPU-affinity pinning and, where the environment permits, + frequency-scaling/turbo disabled, to reduce host noise — stated as a + control to apply, not guaranteed available in every run environment + (this could not be verified as available in the current research + sandbox). +- **Kill condition**: if, for every `(T,k)` cell, tile-major's total + cache-miss count is ≥ cycle-major's at the same `k`, the loop-interchange + direction the hypothesis claims beneficial is wrong outright on this + geometry/host. A weaker, partial kill: amortization (tile-major beats + cycle-major) appears **only** at `k=1` — meaning the effect is ordinary + same-buffer reuse dressed up as "petal tiles," not a genuine multi-cycle + amortization effect, since `k=1` by construction cannot distinguish the + two orderings. + +### H5 — Compaction-as-repair-of-locality / RS-repair-locality probe + +Two parts, deliberately split because only one is measurable against real +shipped code: + +- **Part 1 (measurable, pre-registered and kill-gated)**: run Lance's + real `compact_files`/`plan_compaction`/`CompactionTask::execute`/ + `commit_compaction` (`optimize.rs`) against fixtures whose locality was + degraded by H2's small-`C` first-come arm or H3's Scatter arm. Define + `locality_debt` = `fragment_count − 1` measured immediately before + compaction (a direct, already-available Lance quantity, no new + instrumentation). Regress `compaction_cpu_ns`/`compaction_wall_ns` + against `locality_debt` across the swept degradation levels (Spearman + ρ, since the relationship need not be linear). This operationalizes + the brief's "can a temporal closure event trigger compaction" / + "locality debt estimated from index-remap entropy" questions with a + concrete, computable proxy instead of leaving them rhetorical. + **Kill condition**: `|ρ| < 0.3` across the swept range falsifies + "compaction cost is a locality-debt function" as a COST MODEL at this + geometry — the alternative (supported by `optimize.rs`'s own design, + which triggers on discrete thresholds like `max_overlays_per_fragment` + and `materialize_deletions_threshold` rather than a continuous debt + score) is that compaction cost is dominated by fixed per-file overhead + and threshold crossings, not a smooth function of scatter. +- **Part 2 (design-only, explicitly NOT pre-registered or kill-gated)**: + Lance has **no** erasure coding (confirmed absent, SOURCE ARCHAEOLOGY + above, both pinned 9.0.0 and current upstream), so an RS-repair-locality + experiment cannot be run against Lance itself. The honest design is a + **standalone** micro-benchmark, outside Lance, implementing a toy + `(k,l,g)`-style local-parity scheme (per the Xorbas/Azure-LRC pattern — + PRIOR ART) over the *same* 4,096-field / cascade-threshold geometry, + measuring whether an SFC order (RowMajor vs Morton) that minimizes query + scatter also minimizes *repair* scatter (distinct physical pages/ + fragments touched to reconstruct one lost field or local-parity group). + This is flagged as future work needing new code that doesn't exist + anywhere in the source map — not something this cell's kill-gated + experiment set can respons­ibly claim to test today. + +## EXPECTED BENEFIT + +Building this harness family would (a) turn the pipeline-depth-`k` +"petal-tile amortization" claim from an unverified assertion into a +measured, literature-grounded finding — exactly the falsifiability +standard this workspace's own governance already requires ("a doc-comment +claim is not a behaviour… a test must exercise the claim") but has not yet +applied to this specific claim; (b) reuse the majority of already-proven +measurement machinery (VmRSS-delta method, digest-identity pinning, +median-of-N with an explicit spread-ceiling gate, the CSV `Row` schema, +the M-arm's natural-vs-reordered comparison shape) rather than reinventing +it, keeping marginal engineering cost low and inheriting hard-won lessons +(the VmHWM retraction) for free; (c) is the first benchmark in this +workspace to exercise `CompactionOptions`/`IndexRemapMode`/`RowAddrRemap` +at all — confirmed zero existing criterion bench touches compaction or +index remap anywhere in `lance-graph`, `lance-graph-benches`, or +`lance-graph-contract`; (d) gives the T×k sweep a principled framing (the +ideal-cache model, Frigo/Leiserson/Prokop/Ramachandran) instead of the +single ad hoc Morton-vs-natural point comparison that is all that exists +today. + +## EXPECTED FAILURE + +(a) Hardware performance counters may be unavailable in the actual run +environment (`perf_event_paranoid`, missing `CAP_PERFMON`) — this was not +verified as available in the current research sandbox, so the +`iai-callgrind`/Cachegrind fallback is load-bearing, not decorative, and +the harness must degrade to an honest "NOT MEASURABLE" report rather than +silently substituting a different (page-fault) proxy for cache misses, as +the current substrate does today. (b) At `T=4096` the working set is +32 MiB, already at or past typical single-socket LLC size (8–32 MiB) — +the T-sweep plausibly shows a knee only in the low range (`T≤64`) and goes +flat/bandwidth-bound beyond that; this is a real, useful, but less +dramatic result than a hierarchy-wide amortization curve, and the harness +design must not be read as promising the dramatic result. (c) Lance's +compaction path performs its own internal batched scanning +(`Scanner::batch_size`/`io_buffer_size`), which may already absorb H2's +small-fragment inefficiency once compaction runs — meaning the real signal +could live entirely in the pre-compaction soak-window point-read curve, +not in compaction cost itself, so the harness must report both curves +rather than treating compaction cost as the single dependent variable. +(d) The digest-identity correctness pin, while mandatory, must run in a +pass **separate** from the timed region — hashing every row during a +cache-miss-sensitive measurement would itself perturb the metric it exists +to validate; the M-arm's own separation of "digest identity (an assert, +not a print)" from the timed T1 window is the model to copy, and any +harness implementation that interleaves the two invalidates its own +numbers. + +## EXPERIMENT + +See H1–H5 above; each already states metric set, baseline, control, and +kill condition inline per the program's pre-registration requirement. Full +metric-to-mechanism map: + +| Metric (program list) | Mechanism / source | +|---|---| +| Logical & physical bytes | `CANONICAL_FRAME_BYTES` (exact arithmetic) vs. `cycle_sink.rs`'s documented Arrow-builder copy (`:80-86`) and `RewriteResult::new_fragments` byte sizes | +| Write/read amplification | Derived: physical bytes written/read ÷ logical bytes changed/needed per query | +| Point/neighborhood/random-take latency | `local_trajectories`/`deinterlace` query timing (temporal.rs), median-of-N | +| Fragment count | `Dataset::get_fragments().len()` / `CompactionMetrics` (native, zero instrumentation) | +| Files & pages touched per query | `CompactionMetrics::files_removed/added` (native) + `pages_touched_4k` (computed from the fixture's address log) | +| Index remap bytes/CPU | `RowAddrRemap::{compact,direct}` build + lookup, wrapped in `Instant` + allocator counters | +| Compaction CPU/wall | `compact_files` call wrapped in `Instant`, process vs. thread CPU time | +| Repair read/write bytes & helpers touched | H5 Part 1 (Lance compaction only) / H5 Part 2 (design-only, standalone LRC toy) | +| Seal CPU & metadata bytes | `persist_sink::run_cycle`'s seal phase (already timed as `seal_ns`/`freeze_ns` in the existing `Row` schema) | +| Cache misses | `perf-event` crate primary, `iai-callgrind` Cachegrind fallback (neither exists in the workspace today) | +| Peak RSS | `/proc/self/status` VmRSS **delta**, never VmHWM (per the workspace's own retracted-VmHWM lesson) | +| Restart recovery work | `persist_sink::recover_and_apply`, reusing the existing "recovery replays an owners chain in stream order" test idiom as the correctness pin | +| Version count / stale-version retention bytes | `WalSink::timeline()` → `FrameMeta` list (native); on-disk manifest history size as a real, measurable quantity | + +## PRIOR ART + +Searches actually run (WebSearch, this session): + +1. `"locally repairable codes LRC Xorbas repair locality erasure coding + survey 2023 2024"` → confirmed Xorbas (implied local parity) and + Azure-LRC `(k,l,g)` as the standard local-parity constructions; a + Dec-2024 survey (Cheng et al., ACM TOS) covers the design space. Used + to ground H5 Part 2's toy scheme rather than inventing an ad hoc code. +2. `"cache-oblivious algorithms loop tiling working set amortization + benchmark methodology"` → confirmed the ideal-cache model + (Frigo/Leiserson/Prokop/Ramachandran) as the standard framing for + exactly H4's question: block-size sweep against an unknown cache + size/associativity, hybrid tiled+recursive designs are known-good. + Used to justify H1/H4's T-sweep-across-cache-levels design rather than + a single Morton-vs-natural point comparison. +3. `"Rust perf_event iai-callgrind cache miss microbenchmark hardware + counters criterion 2024 2025"` → confirmed `perf-event` (raw + `perf_event_open` wrapper, cache-references/misses/instructions/ + branch-misses/page-faults/context-switches) and `iai-callgrind` + (Cachegrind-simulated, deterministic, CI-safe) as the two live Rust + options; confirmed neither exists anywhere in this workspace today + (grep, zero hits). +4. `"Z-order Morton curve layout query locality repair locality erasure + coded storage tradeoff"` → **no direct prior art found** connecting + space-filling-curve query locality to erasure-code repair locality — + see SURPRISES. Confirms H5 Part 2 is genuinely open, not a rediscovery. +5. `"LSM-tree compaction write amplification RUM conjecture cost model + locality columnar"` → confirmed "compaction trades write amplification + for reduced read/space amplification" is standard, well-established + LSM-tree literature (RUM conjecture; Sarkar et al., VLDB'21, + "Constructing and Analyzing the LSM Compaction Design Space"). Used to + classify H5 Part 1's hypothesis as TRANSFER, not novel — see below. + +## SURPRISES + +1. **Finding**: no hardware cache-miss counters (`perf_event`, + Cachegrind, or otherwise) exist anywhere in this workspace's + benchmarking substrate today — every existing measurement (WAL-curve + harness, `ndarray` Criterion benches) uses either wall-clock time or + `/proc`-derived page-fault/RSS proxies. **Classification: KNOWN** (this + is simply an observed gap in existing tooling, not a research claim). +2. **Finding**: Lance 9.0.0's `IndexRemapMode::{Compact,Direct}` is + already, explicitly, a memory-vs-lookup-CPU tradeoff shipped in + production Lance — the exact "index remap bytes/CPU" axis this + program's metric list asks for already exists as a named, documented + knob, not something to synthesize. **Classification: KNOWN** (found + directly in source, `optimize.rs:162-180`). +3. **Finding**: "compaction reduces read/space amplification at the cost + of write amplification" (this cell's H5 Part 1 premise) is standard + LSM-tree cost-model literature (RUM conjecture), being applied here to + a fragment/manifest-versioned columnar format (Lance) that is + architecturally distinct from a leveled LSM tree. **Classification: + TRANSFER** (known technique/framing, applied at a different layer — + Lance is not an LSM tree; its compaction trigger is threshold-based on + fragment size/deletion fraction, not level-merge based). +4. **Finding**: no prior art was found connecting space-filling-curve + query-locality optimization (Z-order/Morton) to erasure-code + repair-locality optimization (LRC/Xorbas-style local parity) — the two + literatures (spatial indexing and distributed-storage coding theory) + appear not to have been cross-pollinated in searchable form. + **Classification: NOVEL-CANDIDATE** (documented search run, see PRIOR + ART item 4; H5 Part 2 is scoped explicitly as future/unmeasured work, + not claimed as tested). +5. **Finding**: the brief's own temporal shorthand + ("T_now × A_last × W_write", "interlacing window A_last < W_write <= + T_now") does not match what `temporal.rs::classify` actually implements + — the real function has two version inputs (write tick, reader-pinned + horizon) plus a mode enum, not three independent version pointers. + **Classification: DISPROVED** (an attractive shorthand from the task + brief that the source code does not support as literally stated; the + underlying STRICT/AWARE/RETRO + `knowable_from` mechanism is real and + correctly named, only the three-pointer framing is wrong). + +## VERDICT + +The benchmark harness family is fully specifiable today without new +research: H1–H4 compose almost entirely from already-shipped, already- +measured primitives (the M-arm/O-arm comparison shape, the VmRSS-delta +protocol, Lance's real `CompactionOptions`/`IndexRemapMode`/ +`RowAddrRemap`), and the one genuine gap in the existing substrate +(hardware cache-miss counting) has two concrete, off-the-shelf Rust +answers (`perf-event`, `iai-callgrind`) neither of which is wired up yet. +H5 is honestly two-tiered: Part 1 (compaction-as-locality-repair) is +measurable against real Lance code today; Part 2 (RS-repair-locality vs. +query-locality under one SFC order) requires new standalone code Lance +does not provide and is correctly scoped as unmeasured future work rather +than folded into the pre-registered, kill-gated experiment set. diff --git a/docs/lotus/rp-seal-v1/E2.md b/docs/lotus/rp-seal-v1/E2.md new file mode 100644 index 00000000..b592e73f --- /dev/null +++ b/docs/lotus/rp-seal-v1/E2.md @@ -0,0 +1,683 @@ +# E2 — Domain E (SoA / cache / amortization), role: ADVERSARY + +**Charge.** Attack the amortization story quantitatively, from the real code, before +anyone builds it. Derive first-principles cost models for (a) whole-batch +freeze-and-seal, (b) first-come chunk/petal processing with incremental seal, +(c) index-remapped random placement. Name the regimes where each LOSES, with +crossover parameters and the measurement that decides. + +**Measurement platform for every number below** (stated so nothing is portable by +accident): Intel Xeon @ 2.10 GHz, 4 vCPU, L1d 192 KiB total, L2 8 MiB total, +**L3 reported 260 MiB (shared/virtualised)**, ext4 on `/dev/vda`, `rustc 1.94.1 +-O -C target-cpu=native`. Standalone probes, not the repo's test harness (no +cargo was run). Sources: the session scratchpad's `e2work/bench{,2,3,4}.rs`. +**A 260 MiB LLC is unusually large; several conclusions below are LLC-sensitive +and I flag each one.** + +--- + +## SOURCE ARCHAEOLOGY + +### A. The write path, pass-by-pass (lance-graph, working tree) + +I counted the passes and copies myself rather than trusting the module docs. + +| # | Site | What it does per cycle | +|---|---|---| +| 1 | `lance-graph-planner/src/batch_writer.rs:132-138` `cast()` | per cast: one `BTreeMap` insert (`board`), one `Vec` push (`pending_payloads`). `P` is **owned** here in practice. | +| 2 | `lance-graph-supervisor/src/cycle_driver.rs:364-391` `collect_casts` | drains payloads, then **two `BTreeMap` lookups per cast** (`on_behalf_of`, `intent_moves`), builds `Vec` (payload moved). | +| 3 | `cycle_driver.rs:459` `let frozen = casts.clone();` | **full deep clone of every payload** — C × 512 B memcpy + C allocations. Held until commit success as a "retry cache". | +| 4 | `persist_sink.rs:619-622` artifact filter | `into_iter().filter().collect()` — a second `Vec` (72 B/elem moves; payload pointers only). | +| 5 | `persist_sink.rs:133-135, 378` `order_cycle_stably` | `sort_by_key` on 72-byte structs, O(C log C) moves + closure calls. | +| 6 | `persist_sink.rs:379-382` image fold | `image.insert(s.row, s.payload.clone())` **per cast** → a **second full deep clone of all C payloads**, of which C−D are allocated, copied, then immediately dropped. Coalescing is done by overwrite, so the CPU cost is O(C), not O(D). | +| 7 | `persist_sink.rs:401-429` `content_hash` | **FNV-1a, one byte at a time**, over frame + every canonical landing **including its full 512-byte payload**. Single multiply-dependency chain; not combinable, not parallelisable. | +| 8 | `lance-graph/src/graph/cycle_sink.rs:458-556` `build_batch` | 13 Arrow builders sized `n = 1 + C + D`; `FixedSizeBinaryBuilder::with_capacity(n, 512)` (`:472`); `append_null()` for the frame row (`:491`) and **every landing row** (`:521`); `append_value` for each image row (`:532`). | +| 9 | `cycle_sink.rs:706-728` `raw_append` | one `Dataset::write(Create)` or `ds.append` → **one file, one manifest, one fsync**. | + +**The load-bearing discovery is at step 8.** In `arrow-array-58.3.0`, +`src/builder/fixed_size_binary_builder.rs:91-95`: + +```rust +pub fn append_null(&mut self) { + self.values_builder.extend(std::iter::repeat_n(0u8, self.value_length as usize)); + self.null_buffer_builder.append_null(); +} +``` + +and `:59-69` `with_capacity(capacity, byte_width)` → `Vec::with_capacity(capacity * byte_width)`. + +**Every null payload slot costs a full 512 zero bytes.** The batch therefore +materialises `(1 + C + D) × 512` payload-column bytes for `D × 512` bytes of real +data. + +### B. Does Lance squeeze the zeros back out? At the pinned column A: **no.** + +Column A = `/tmp/sources/lance-9`, upstream tag **v9.0.0**, HEAD `7653c20`, +`Cargo.toml version = "9.0.0"` (matches the workspace lock pin). + +| Fact | File:line (column A) | +|---|---| +| `LanceFileVersion::default()` is **`V2_1`** | `rust/lance-encoding/src/version.rs:19-40` (`#[default] V2_1`) | +| `WriteParams::default().data_storage_version = None` → `.unwrap_or_default()` = V2_1 | `rust/lance/src/dataset/write.rs:429`, `:457-459` | +| Automatic general block compression is gated **`version >= V2_2`** | `rust/lance-encoding/src/compression.rs:536-542` (`MIN_BLOCK_SIZE_FOR_GENERAL_COMPRESSION = 32 KiB`, `:86`) | +| Constant-page layout is gated **`>= V2_2`** | `rust/lance-encoding/src/encodings/logical/primitive.rs:5425-5437` | +| RLE requires `bits_per_value ∈ {8,16,32,64}`; FSB(512) = **4096 bits** → excluded | `rust/lance-encoding/src/compression.rs:279-282` | +| `FixedSizeBinary` → `encode_flat_data` → `FixedWidthDataBlock { bits_per_value = byte_width*8 }`; **nulls are not compacted out** | `rust/lance-encoding/src/data.rs:1507-1529` | + +`LanceCycleWriter` passes `WriteParams { mode: Create, ..Default::default() }` +(`cycle_sink.rs:713-716`) and `ds.append(reader, None)` (`:726`) — no storage-version +override anywhere in the file. **So the zero padding reaches disk verbatim under +lance 9.0.0 defaults.** (Column B / current upstream was NOT consulted for this +claim; it is a 9.0.0 statement only.) + +Two further 9.0.0 facts that bound design (c): + +* `Manifest.fragments: Arc>` — the **full** fragment list is in + **every** manifest (`rust/lance-table/src/format/manifest.rs:35-52`). Cumulative + manifest bytes are O(commits × fragments). +* Prepared-fragment staging **does exist**: `InsertBuilder::execute_uncommitted_stream` + → `Transaction` (`rust/lance/src/dataset/write/insert.rs:181`), and + `CommitBuilder::execute_batch` merges N `Operation::Append` transactions into + ONE manifest (`rust/lance/src/dataset/write/commit.rs:475-509`, Append-only). +* `RowAddrRemap::{Compact (O(#fragments)), Direct (HashMap>, O(#rows))}` + (`rust/lance-core/src/utils/row_addr_remap.rs:59-65`); `CompactionOptions.index_remap_mode` + **defaults to `Direct`** (`rust/lance/src/dataset/optimize.rs:~243`), and + `WriteParams::enable_stable_row_ids` **defaults to `false`** (`write.rs:430`). +* Page statistics / zone maps exist **only in the legacy v1 reader** + (`rust/lance-file/src/previous/reader.rs:375,393`). No zone maps on the v2 path, + and `LanceCycleWriter` creates **no scalar index** (grep for `create_index` in + `cycle_sink.rs`: zero hits). `scan_image(cycle)` / `scan_sealed(after_cycle)` + are therefore full-column scans over all history. + +### C. Substrate capability gaps that price the erasure half + +`ndarray` `src/simd_ops.rs` + `src/simd_int_ops.rs` exhaustive inventory: f32/f64 +arithmetic, i8/i16 ops, GEMM, `eq_u32_to_mask`, `gt_i32_to_mask`, +`mask_and/or/andnot(_assign)`, `masked_sum_i32`, `array_chunks/windows`, bf16 tile +GEMM. **There is no `mask_xor`, no `xor_bytes`, no GF(2^8) multiply, no CLMUL, no +CRC, no SIMD hash.** An RS path under the "all SIMD from `ndarray::simd`" +invariant needs new primitives before a single parity byte is computed. + +Conversely `ndarray/src/hpc/blake3.rs` (864 lines) is a **real tree hash** — +`parent_cv` at `:330`, `cv_stack: [[u32;8];54]` at `:421`. A combinable +incremental seal digest already exists in-tree; the seal does not use it. + +### D. Temporal is a read-side cost, not a seal-side one + +OGAR `docs/TEMPORAL-TIME-TRAVEL.md` @ `386a6fd8` is explicit: the epistemology +layer is a **planner query annotation** — `QueryReference { ref_version, mode, rung }`, +STRICT/AWARE/RETRO, HLC `(server_id, local_lance_version, hlc_tick)` — and +"**No new storage, no new contract, no new container.**" `deinterlace` +(`lance-graph-planner/src/temporal.rs:346-376`) is a filter + `.cloned()` + full +comparison sort: at 64 k rows that is 32 MiB of clones plus O(n log n) — **entirely +on the read path**. Nothing in the seal cost model should be charged to temporal, +and nothing in temporal amortises with the seal. + +--- + +## MECHANISM + +Parameters used throughout: + +* `N` = 65,536 rows/cycle (32 MiB at 512 B/row), `P` = 4096 petals/cycle + (16 rows = 8 KiB each) — the program's baseline geometry. +* `C` = artifact casts (breaths) sealed in the cycle. +* `D` = distinct dirty rows, `D ≤ C`. +* `b = C/D` = **breaths per dirty row** (the coalescing factor; the repo's own + falsifier uses b = 64). +* `δ = D/N` = **dirty-row density**. The supervisor's own ruling is + `E-COMPLETE-CYCLE-IS-PHYSICALLY-SPARSE-NOT-A-FULL-REWRITE-1`, i.e. δ ≪ 1 is the + claimed normal case. +* `g` = parity-unit granularity (512 B row, or 8 KiB petal). +* `G` = number of parity groups per cycle. + +### (a) Whole-batch freeze-and-seal — measured cost model + +Per cycle, from the pass list above: + +``` +CPU(a) = C·(2 lg C tree lookups) [collect_casts] + + C·512 B memcpy + C allocs [frozen = casts.clone()] + + O(C log C) 72-B moves [order_cycle_stably] + + C·512 B memcpy + C allocs [image insert clones] + + (C·568 B) serial FNV [content_hash] + + (1 + C + D)·512 B buffer writes [arrow payload column] +Bytes(a, payload column) = (1 + C + D) · 512 +Durable ops(a) = 1 file + 1 manifest + 1 fsync +``` + +Measured, C = 65,536 (`bench.rs`, ms): + +| term | D=C (b=1) | D=C/4 (b=4) | D=C/64 (b=64) | +|---|---|---|---| +| `frozen = casts.clone()` | 21.4 | 17.2 | 16.8 | +| freeze image clones | 29.0 | 10.9 | 5.4 | +| `sort_by_key` | 8.3 | 7.9 | 8.5 | +| **`content_hash` (FNV-1a)** | **47.6** | **47.5** | **47.4** | +| arrow payload column | 36.4 | 18.5 | 13.6 | +| **CPU subtotal** | **142.7** | **102.0** | **91.7** | +| FNV share of subtotal | 33 % | 47 % | **52 %** | +| payload-column bytes | 64.0 MiB | 40.0 MiB | 32.5 MiB | +| **amplification vs. D·512** | **2.00×** | **5.00×** | **65.0×** | + +Anchors: FNV-1a serial = **0.70 GB/s** (≈ 3 cycles/byte, the `imul` dependency +chain); an 8-independent-lane ILP variant of the same arithmetic reaches +**2.57 GB/s** (3.6× — an ILP ceiling probe, *not* a valid substitute digest); +`memcpy`-shaped payload cloning ≈ 1.6–2.0 GB/s including allocator traffic. + +Durable boundary, measured on ext4 (`bench4.rs`): small append + `fsync` = +**0.126–0.159 ms**; one 32 MiB append + `fsync` = **202 ms** (≈165 MB/s). + +### (b) First-come chunk/petal with incremental seal — cost model + +``` +Durable ops(b) = F fragments/cycle (+ 1..F manifests) +Bytes(b, payload) = C·512 (no cross-petal coalescing if a petal closes + between two breaths on the same row) +Manifest(b) after K cycles = K·F fragments listed in EVERY manifest +Seal digest(b) requires a COMBINABLE hash — FNV cannot combine +RSS(b) = K_open · 8 KiB +``` + +Measured file/metadata tax (`bench4.rs`, create + 8 KiB + `fsync`): +**0.197–0.249 ms/file**, essentially flat in count. + +* 4096 files/cycle = **910 ms**, vs **202 ms** for one 32 MiB append+fsync → + **4.5× worse wall-clock for the same 32 MiB**, before any manifest bytes. +* Manifest arithmetic at ~130 B/`DataFragment` (`protos/table.proto:313-365`: + `id` + `DataFile{path ≈ 41 B uuid.lance, fields, column_indices, versions, + size}`): with F = 4096 and one merged manifest per cycle, cycle K writes + `K·4096·130 B` — **53 MB at K=100, 533 MB at K=1000, per commit**. With one + manifest per petal it is `O(K·F²)` — 21.8 GB/cycle at K=10. Both are dead + without per-cycle compaction. +* Break-even from the measured constants: F fragments cost `0.22·F` ms; that + equals the 202 ms whole-batch fsync at **F ≈ 920**. To hold the metadata tax + under 15 % of the seal, **F ≲ 140 fragments/cycle** — i.e. petals must be + grouped **≥ 30:1** before becoming fragments. **4096 is 30–60× too many.** + +### (c) Index-remapped random placement — cost model + +``` +Cache: penalty(random 512-B row scatter) = f(working set / LLC) +Remap: RowAddrRemap::Direct ≈ 24 B payload/entry, hashbrown ≤7/8 load, + pow-2 table ⇒ 32–64 B resident per row +Identity: enable_stable_row_ids = false (9.0.0 default) ⇒ row ADDRESSES move +``` + +Measured scatter curve (`bench3.rs`, full random permutation, 512-B rows): + +| image | seq GB/s | rand GB/s | **rand/seq** | +|---|---|---|---| +| 8 MiB | 10.28 | 9.79 | 1.05× | +| **32 MiB = one cycle** | 9.12 | 9.35 | **0.97× (no penalty)** | +| 128 MiB | 9.27 | 6.56 | 1.41× | +| 512 MiB | 7.39 | 3.22 | **2.30×** | +| 2 GiB | 7.19 | 3.19 | 2.25× | + +`BTreeMap` image build with random vs sequential row keys: 28.1 vs 23.8 ms = **1.18×**. + +--- + +## EXPECTED BENEFIT + +Stated fairly, because an adversary that concedes nothing has measured nothing. + +1. **Whole-batch's durable-boundary amortisation is real and large.** One fsync + per cycle at 0.126–0.202 ms fixed vs. a per-cast marginal of ~2.2 µs + (142.7 ms / 65,536) means the fixed cost is amortised out at **C ≈ 100 casts**. + Above that the durable boundary is genuinely free per cast. +2. **Coalescing's *logical* claim holds.** `DetachedCycleBatch::freeze` really does + fold b breaths into one image row; `scan_image` really does return one 512-byte + value. The reader-visible image is minimal. +3. **Write-side canonical ordering really does eliminate read-time repair.** + `scan_sealed` never sorts (`cycle_sink.rs:971-1004`), and the sort it avoids is + the 8.3 ms measured above — per reader, per read, forever. +4. **Random identity placement costs nothing at the program's own geometry.** + 0.97× at 32 MiB. My prior — "cache tiling defeated by random arrival" — is + **falsified at N = 1 cycle on this machine**. Reported honestly as a negative. +5. **Prepared-fragment staging with one publication is available at 9.0.0** + (`execute_uncommitted_stream` + `execute_batch`), so the "immutable code symbols + before one manifest publication" shape is buildable, not hypothetical. +6. **A combinable seal digest already exists in-tree** (`ndarray::hpc::blake3`, + `parent_cv`), so the incremental-seal prerequisite is a wiring problem, not a + research problem. + +--- + +## EXPECTED FAILURE + +### F1 — The coalescing win is exactly cancelled, and then some, by null padding + +**This is the strongest single finding and it is on the code as written.** + +Payload-column bytes: coalescing saves `(C − D)` **image** rows but pays `C` +**landing** rows of 512 zero bytes each (`build_batch` `:521`), plus the frame row. + +``` +bytes_current = (1 + C + D)·512 +bytes_naive = C·512 (one payload row per cast, no coalescing at all) +bytes_ideal = D·512 +``` + +`bytes_current > bytes_naive` **for every C, D ≥ 1**. Amplification vs ideal is +`(1+C+D)/D = b + 1 + 1/D` — **the amplification factor IS the coalescing factor**. +At the repo's own b = 64: 65×. At b = 1: 2×. + +The existing falsifier does not catch it. `cycle_sink.rs:1247-1285` +`sixty_four_breaths_on_one_row_cost_one_image_row` calls itself "the measured +bytes-written falsifier" and then asserts +`image.values().map(Vec::len).sum() == 512` — where `image` is a +`BTreeMap>` keyed by row, so **for one row it can only ever be 512, +by construction**. It measures the logical image, never a file. Predicted true +file content for that test: 1 frame + 64 landings + 1 image = **66 × 512 = +33,792 payload-column bytes on disk**, 512 of them real. By this repo's own +falsifiability rule, that assertion is implied by the code it tests. + +**Kill/repair (cheap):** separate landings and image into two Arrow batches (or +two datasets), or make `payload` a `Binary` column where a null costs an offset, +not 512 bytes, or write at `LanceFileVersion::V2_2+` so the >32 KiB block +auto-compressor sees the zero runs. + +### F2 — The seal's dominant CPU term does not amortise at all + +`content_hash` is byte-at-a-time FNV-1a over **every canonical landing including +its payload** — 47.5 ms, **invariant across b = 1, 4, 64** (measured). It is 33 % +of the seal at b=1 and **52 % at b=64**: the better coalescing gets, the more the +seal is just hashing. FNV's serial multiply chain also makes it the one term that +**cannot** be split per petal, so it is simultaneously the biggest cost and the +hard blocker on design (b). Repairs: `xxh3` (~10 GB/s, ≈14× faster) or +`ndarray::hpc::blake3` tree-hashed per petal and combined with `parent_cv`. + +### F3 — Two full payload deep-copies before Lance ever sees a byte + +`cycle_driver.rs:459` (`frozen`) and `persist_sink.rs:381` (`image.insert(...clone())`) +each copy all C payloads; `build_batch` copies D more. Total ≈ `(2C + D)·512` +plus the `(1+C)·512` zero fill. At C=D=65,536 that is **~160 MiB of memory +traffic and ~131 k allocations to make one 32 MiB durable image — a ~5× RSS +multiplier and a ~5× memory-bandwidth multiplier.** `frozen` is explicitly an +*optional* retry cache the module doc says is always safe to drop +(`cycle_driver.rs:420-425`), so 21 ms and 32 MiB are spent per cycle on a +convenience. + +### F4 — Parity update amplification: the crossover is at ~1.5 % dirty density + +Under the classical RAID-6 partial-stripe result (3 reads + 3 writes per +partial-stripe update), one touched parity group costs `6·g` bytes of I/O. + +Distinct groups touched, D dirty rows uniform over G groups (balls-in-bins): +`E[G_touched] = G·(1 − e^(−D/G))`. + +With `G = P = 4096` petal groups, `g = 8 KiB`, RS(8,2) so a full-stripe rewrite +of the whole 32 MiB cycle costs `1.25 × 32 MiB = 40 MiB`: + +``` +incremental RMW cheaper ⟺ G_touched · 6 · 8192 < 41,943,040 + ⟺ G_touched < 853 + ⟺ D < ~956 rows ⟺ δ < 1.46 % +``` + +**Above ~1.5 % dirty-row density, recomputing the whole cycle's parity is cheaper +than updating it incrementally.** And below it, per-row amplification is brutal: +at D = 1024, `G_touched ≈ 906` (0.885 groups per dirty row — essentially no +sharing), so 906 × 48 KiB = **42.5 MiB of parity I/O for 512 KiB of logical +change = 85×**. + +Moving the parity granule to 512 B changes the crossover by 14× (`6·512 = 3072 B` +per group; break-even at 13,653 rows = **20.8 % density**). **The parity granule, +not the seal granule, sets the crossover.** No parity granule is currently chosen +anywhere, and no adaptive policy exists — and an adaptive one would itself need an +inertness test (raising the threshold must silence something). + +Prerequisite cost, separately: `ndarray::simd` has **no** XOR-over-bytes, GF(2^8), +CLMUL or CRC primitive. Scalar `xor_bytes` measures 14–20 GB/s (`bench.rs`), so +P-parity CPU is not the problem; Q-parity GF multiply without PSHUFB tables is +~1 B/cycle and would be. + +### F5 — Per-petal fragments at 4096/cycle are not viable in Lance's manifest model + +Measured 910 ms/cycle of `create+fsync` for 4096 files (4.5× the whole-batch +fsync), and `Manifest.fragments` carries the **full** list forever, so cumulative +manifest bytes are `O(K·F)` per commit and `O(K²·F)` in total. Even with +`execute_batch` collapsing the commits to one manifest, the **fragment count** is +the O(N²) driver, not the commit count. Consequence: **per-petal fragmentation +makes compaction mandatory every cycle**, and a compaction of a 32 MiB cycle is a +full read + full write + index remap — the incremental saving is repaid as +compaction debt inside the same cycle. Viable F is **≲ 140/cycle**, i.e. petals +grouped ≥30:1. + +### F6 — First-come petal processing and coalescing are in direct conflict + +If a petal closes between two breaths on the same row, the second breath is a new +durable image row **and** a new parity RMW. Expected durable image rows per +logical row under an open-window `T`: + +``` +E[rows] = 1 + (b − 1)·P(inter-breath gap > T) +``` + +So the b-fold coalescing is recovered only if `T` exceeds the ~p99 inter-breath +gap for that row. Under arrival-driven ("first-come") closure that gap is +unknowable without a barrier — and a barrier is whole-batch. **"First-come" and +"coalescing" cannot both be maximised; b is the exchange rate.** At the repo's +own b = 64 a fully first-come design writes 64× the payload bytes and performs +64× the parity RMWs. + +### F7 — RSS advantage of open petals evaporates under adversarial identity arrival + +`K_open · 8 KiB` (measured: 0.5 / 2.0 / 8.0 / 32.0 MiB at K = 64 / 256 / 1024 / +4096). The adversary is trivial: make each arriving cast target a distinct petal. +`K_open` is bounded below by the coupon-collector count `G(1 − e^(−D_window/G))`, +so a uniformly-scattered identity stream drives `K_open → P` and the resident set +to **32 MiB — identical to whole-batch's image, plus all the per-petal metadata.** +The RSS advantage is a property of identity *clustering*, not of the incremental +design; it is exactly the LOTUS locality claim, and if that claim fails the +incremental design has no RSS story left. + +### F8 — Random placement is free at 32 MiB and expensive above ~½ LLC + +Measured 0.97× at one cycle, 1.41× at 128 MiB, **2.30× at 512 MiB**. The geometry +is already sized to the cache — which is a benefit, and also a **fragility**: the +design loses the moment (i) more than ~4 cycles are open concurrently, (ii) the +row grows past 512 B, or (iii) it runs on a normal 32–64 MiB-L3 server part, where +**one** cycle already sits at the LLC boundary. This machine's 260 MiB L3 flatters +the design by roughly 8× in working-set headroom. **Do not generalise the 0.97×.** + +### F9 — Physical relocation *does* change logical identity, by default + +`enable_stable_row_ids: false` (9.0.0 default, `write.rs:430`) means row +**addresses** move on compaction. Anything keyed by row address — a mask, a +Morton→address table, a parity-group membership map — is silently invalidated by +every compaction. Combined with `index_remap_mode` defaulting to `Direct` +(`HashMap>`, ~32–64 B resident per row): one 65,536-row cycle +costs 2–4 MiB of transient hash table, a 100-cycle compaction **210–420 MiB**, a +1000-cycle compaction **2.1–4.2 GB**. `Compact` mode exists and is O(#fragments); +it is not the default. + +### F10 — Point reads are O(total history), not O(1) + +`scan_image(cycle)` filters `kind = 2 AND cycle = X` with **no scalar index and no +v2 zone maps**. Reading cycle #10,000's image scans the projected columns of all +10,000 cycles. Since the payload column is 2–65× amplified (F1), that scan is also +2–65× larger than the logical data. This is a read-amplification term that grows +without bound in cycle count and is invisible to every existing test. + +### Two-sided: where WHOLE-BATCH clearly loses + +| # | Regime | Quantity | Crossover | +|---|---|---|---| +| W1 | Low fan-out cycles | fixed 0.126–0.22 ms + `(1+C+D)·512` per cycle regardless of C | amortisation only pays at **C ≳ 100**; below that whole-batch is pure overhead | +| W2 | Write-visible p99 latency | first cast waits the entire cycle: ≥ 92–143 ms CPU + up to 202 ms fsync ⇒ **~300 ms** | any SLA < ~300 ms. There is **no release rule** (timer/size) — the release *is* the cycle boundary | +| W3 | Bytes-at-risk | a crash loses the whole open cycle (up to C casts) | the "deterministic regeneration" defence (`cycle_driver.rs:28-44`) is void the moment a thought reads a clock, an LLM, or an external source | +| W4 | Peak RSS | `(3C + 2D + 1)·512` ≈ **160 MiB** for 32 MiB of data | vs `K_open·8 KiB` = 0.5 MiB at K=64 — **~300× worse** when identity *is* clustered | +| W5 | Restart recovery work | 100 % of the cycle re-run | incremental re-runs ~½ a petal group | +| W6 | Parity, sparse regime | full-stripe recompute = 40 MiB vs RMW = `G_touched·48 KiB` | **below δ ≈ 1.5 %, incremental parity wins** — precisely the regime `E-COMPLETE-CYCLE-IS-PHYSICALLY-SPARSE-…` claims is normal | + +W6 against F4 is the design's real bind: **no single parity policy is right on both +sides of δ ≈ 1.5 %**, and δ is a workload property nobody has measured. + +--- + +## EXPERIMENT (pre-registered) + +Metrics are drawn from the program's list. Each has a baseline, a control, and a +kill condition stated before the run. + +### X1 — Physical bytes of the null-padded payload column *(kills or confirms F1)* +* **Metric:** physical bytes; write amplification. +* **Procedure:** write one cycle at C=65,536 for b ∈ {1, 4, 64}; `du -sb` the + dataset directory, and read the per-column encoded size from the Lance file + footer. Compare to `D·512`. +* **Baseline:** `bytes_ideal = D·512`. **Control:** an identical run with the + landing rows moved to a second table (no `payload` column at all). +* **Prediction:** `(1+C+D)·512` ± Lance framing; amplification 2.00 / 5.00 / 65.0×. +* **KILL:** if the b=64 dataset is within 1.5× of `D·512`, F1 is disproved — + Lance is squeezing the zeros somewhere I did not find, and the padding argument + collapses. +* **Second arm:** repeat at `data_storage_version = V2_2`. Prediction: the >32 KiB + auto-compressor removes most of the padding on disk but **not** the in-memory + `(1+C+D)·512` builder buffer or the RSS. + +### X2 — Seal CPU decomposition and the hash share *(kills or confirms F2)* +* **Metric:** seal CPU (ms) per stage; cache misses (`perf stat`); peak RSS. +* **Procedure:** instrument the five stages of `persist_cycle` at + C=65,536 × b ∈ {1,4,64}; then swap `content_hash` for (i) `xxh3`, (ii) + `ndarray::hpc::blake3` per-petal + `parent_cv` combine. +* **Baseline:** current FNV. **Control:** a null-hash build (returns 0) to bound + the maximum achievable saving. +* **Prediction:** FNV is 33 %/47 %/52 % of seal CPU; xxh3 removes ≥ 85 % of it; + BLAKE3-tree removes ≥ 60 % single-threaded and ≥ 85 % on 4 cores. +* **KILL:** if FNV is < 15 % of measured seal CPU on the real path (i.e. Lance + encode + I/O dominates by 10×), F2 is a micro-optimisation, not a finding. + +### X3 — Fragment-granularity sweep *(kills or confirms F5)* +* **Metric:** fragment count; files & pages touched per query; compaction CPU/wall; + version count; stale-version retention bytes; manifest bytes per commit. +* **Procedure:** F ∈ {1, 16, 64, 256, 1024, 4096} fragments/cycle via + `execute_uncommitted_stream` + `execute_batch`; run K = 200 cycles; measure + per-commit manifest size, total wall time, and then the cost of + `compact_files` back to F=1. +* **Baseline:** F=1 (today). **Control:** F=4096 with compaction disabled, to show + the manifest curve unbounded. +* **Prediction:** total wall ≈ `202 + 0.22·F` ms/cycle; manifest bytes/commit + ≈ `130·K·F`; the knee is F ≈ 140. +* **KILL:** if wall time is flat in F up to 4096 (object-store batching hides the + metadata op), F5 is a local-filesystem artefact and must be re-run on the real + object store before it is quoted. + +### X4 — Parity granule × dirty density *(kills or confirms F4 / W6)* +* **Metric:** repair read/write bytes & helpers touched; write amplification; seal CPU. +* **Procedure:** simulate RS(8,2) over the cycle at `g ∈ {512 B, 8 KiB}`, sweeping + δ ∈ {0.1 %, 0.5 %, 1.5 %, 5 %, 25 %, 100 %} under two identity distributions: + uniform, and Morton-clustered (the LOTUS claim). Count parity bytes read+written + for incremental RMW vs full-stripe recompute. +* **Baseline:** full-stripe recompute = `1.25 × 32 MiB`. **Control:** no parity. +* **Prediction:** uniform crossover at δ ≈ 1.46 % (g=8 KiB) and δ ≈ 20.8 % + (g=512 B); Morton clustering shifts the crossover **up** in proportion to the + reduction in `G_touched`. +* **KILL:** if Morton clustering does **not** reduce `G_touched` by ≥ 4× at + δ = 1 %, the locality premise that makes RS affordable is disproved, and RS over + the logical grid should be abandoned in favour of RS over the physical append + batch (which is always full-stripe, 1.25×, at the cost of tying repair locality + to write time rather than identity). + +### X5 — Coalescing loss vs petal-open window *(kills or confirms F6)* +* **Metric:** logical & physical bytes; write amplification; version count. +* **Procedure:** replay a real cast trace; sweep petal-open window + T ∈ {0, 1, 10, 100 ms, ∞(=whole-batch)}; measure durable image rows per logical + row and the empirical inter-breath gap distribution per row. +* **Baseline:** T=∞. **Control:** T=0 (pure first-come). +* **Prediction:** `E[rows] = 1 + (b−1)·P(gap > T)`; T must exceed the p99 gap to + recover ≥ 95 % of the coalescing. +* **KILL:** if the measured p99 inter-breath gap is < 1 ms, first-come petals cost + almost nothing and F6 is moot — the barrier can be a 1 ms timer, not a cycle. + +### X6 — Scatter penalty vs LLC *(bounds F8, already partly run)* +* **Metric:** cache misses; point/neighborhood/random-take latency; GB/s. +* **Procedure:** the `bench3` sweep on ≥3 machines with distinct L3 (32 / 64 / + ≥256 MiB), plus a `K_open ∈ {1,2,4,8}` concurrent-cycle arm. +* **Baseline:** sequential placement. **Control:** `cpupower`/`resctrl`-restricted + LLC to emulate a small part. +* **Prediction:** penalty ≈ 1.0× below ½ LLC, rising to ~2.3× above 2× LLC. +* **KILL:** if a 32 MiB-L3 part shows ≥1.5× at K_open = 1, the geometry is + **not** cache-sized in general and the 32 MiB cycle constant must be re-derived + per deployment, not fixed. + +### X7 — Restart recovery + stale-version retention *(W5, F9, F10)* +* **Metric:** restart recovery work; version count; stale-version retention bytes; + index remap bytes/CPU; point-read latency vs history length. +* **Procedure:** K = 1000 cycles; measure `scan_image(cycle_1)` latency at + K ∈ {10, 100, 1000}; run `compact_files` with `index_remap_mode ∈ {Direct, Compact}` + and `enable_stable_row_ids ∈ {false, true}`, recording peak RSS and whether + externally-held row addresses survive. +* **Baseline:** K=10. **Control:** the same dataset with a scalar index on `cycle`. +* **Prediction:** point-read latency linear in K; `Direct` remap RSS ≈ 32–64 B × + total rows; with `enable_stable_row_ids = false`, externally-held addresses break. +* **KILL:** if point-read latency is flat in K, Lance is pruning by some mechanism + I did not find in 9.0.0, and F10 is withdrawn. + +--- + +## PRIOR ART + +Searches actually run (web + alphaXiv), with what each returned. + +1. **"erasure coding partial stripe update parity read-modify-write amplification + small write penalty RAID-6"** → the Small Write Penalty is textbook: 2 reads + + 2 writes for RAID-5, **3 reads + 3 writes (6 physical I/Os per logical write) + for RAID-6**; Reconstruct-Write is preferred once most of a stripe is updated. + H-Code (partial-stripe-write-optimised MDS array code) and **PARIX: Speculative + Partial Writes in Erasure-Coded Systems** (USENIX ATC'17) are the direct prior + art. ⇒ **F4's mechanism is KNOWN.** What is not in the literature is the + specific δ ≈ 1.5 % crossover for *this* geometry — that is arithmetic, not + novelty. +2. **"Local Reconstruction Codes Azure repair locality Huang 2012"** → Huang et + al., *Erasure Coding in Windows Azure Storage*, ATC'12 (Best Paper): (12,2,2) + LRC, local groups of six, "a 10 % increase in storage overhead can halve the + cost of all degraded reads and most block repairs". Kolosov et al. ATC'18 on + optimality of LRCs. ⇒ **repair locality via grouping is thoroughly KNOWN**; any + petal↔parity-group identification is an instance, not an invention. +3. **"columnar file format null padding fixed-size-binary write amplification + sparse column Arrow validity"** → Arrow docs on validity bitmaps + padding, and + *Mainlining Databases* (arXiv 2004.14471): "**There are two sources of write + amplification in Arrow: (1) it disallows gaps in a column**, and (2) it stores + variable-length values consecutively." ⇒ **F1's mechanism is KNOWN at format + level.** Its instantiation — a coalescing optimisation whose saving is exactly + cancelled by the null padding of its own metadata rows — I did not find named + anywhere. +4. **"Iceberg / Delta Lake manifest rewrite O(n²) small commits metadata + amplification streaming"** → extensively known: per-commit metadata files, + "1,440 manifest files per partition over 24 hours", "a table with 500+ manifests + forces the planner to open and parse each one", and the Iceberg v4 roadmap + explicitly targeting single-file commits / adaptive metadata trees. ⇒ **F5 is + KNOWN** and Lance 9.0.0's `Arc>`-in-every-manifest is the same + shape. +5. **"group commit amortization crossover batch size latency throughput fsync"** → + "a flush has large fixed cost, group commit batches many committing transactions + into one flush"; the throughput/latency/bytes-at-risk triangle; release rules + (timer T, count K, adaptive). *Group Commit Self-Clocks* (arXiv 2606.18187) + argues tuning is unnecessary above a device-set load threshold. ⇒ **the + whole-batch seal is textbook group commit**, and its missing release rule (W2) + is the textbook omission. +6. **"write-combining buffer partial cache line store … 64-byte granularity + penalty"** → "if the buffer has only been partially filled before being flushed, + the memory controller will internally perform a **read/modify/write cycle** to + merge the new data" — the cache-level twin of partial-stripe parity RMW. +7. **"FNV-1a byte at a time … versus xxHash BLAKE3 tree hash"** → "FNV-1a's serial + recurrence prevents parallelizing a single hash across input chunks, which is + its fundamental limitation versus tree-structured hashes like BLAKE3 or + block-parallel hashes like xxh3"; BLAKE3's Merkle structure permits incremental + and parallel hashing. ⇒ **F2 is KNOWN**; the finding is that this seal picked + the one hash that forecloses design (b). +8. **alphaXiv discovery** (keywords: space-filling curve, erasure coding, repair + locality, compaction, write amplification, Morton order; difficulty 8) → 11 + results, all LRC/locally-repairable-code theory (2101.10028, 2312.12803, + 2401.07835, 2401.15308, 2505.06819, 2603.09605 *Nemo*, 2603.05162 *RESYSTANCE*, + 2603.05439 *O³-LSM*, 2602.10020 *METTLE*, 2605.31214, 2608.04403) plus LSM + compaction-offload work. **Nothing returned on: one SFC order chosen to minimise + both query-scan and RS-repair scatter; compaction framed as repair of physical + locality; parity-group boundaries doubling as compaction boundaries.** Recorded + as a documented negative search, not as proof of novelty — I ran one discovery + call plus six web searches, which is not an exhaustive survey of the + coding-theory literature. +9. **"merge-on-read delta log record amplification cancels compaction benefit"** → + Hudi MoR background; nothing matching the specific "metadata row cancels the + coalescing saving" pattern. + +--- + +## SURPRISES + +1. **The coalescing optimisation writes MORE payload-column bytes than no + coalescing at all, and its amplification factor equals its coalescing factor + (b+1).** `(1+C+D)·512 > C·512` always, because `FixedSizeBinaryBuilder::append_null` + writes 512 zero bytes per landing row and lance 9.0.0's default V2_1 has neither + auto-compression (≥V2_2 gated) nor an RLE path for 4096-bit values. + → **KNOWN mechanism, but this instance appears unnoticed** — the repo's own + "measured bytes-written falsifier" asserts a quantity that is 512 by + construction. Classified **KNOWN** (Arrow's no-gaps write amplification, + arXiv 2004.14471) — the *instance* is a defect, not a discovery. + +2. **The seal's dominant CPU term is invariant under the very optimisation the + seal exists to perform.** FNV-1a over all C landings measures 47.5 ms at b=1, + b=4 and b=64 alike, rising from 33 % to 52 % of seal CPU as coalescing improves. + The one hash that cannot be combined per chunk is also the hard blocker on + incremental sealing — while a tree hash (`ndarray::hpc::blake3::parent_cv`) is + already in the workspace, unused. → **KNOWN** (FNV serial recurrence vs + BLAKE3/xxh3 is standard hashing folklore); the *coupling* to the seal-granularity + choice is the useful part. + +3. **Random identity placement costs nothing at the design's own 32 MiB geometry + (0.97×) and 2.30× at 512 MiB.** My adversarial prior — "cache tiling defeated by + random arrival" — is falsified at N = 1 cycle on a 260 MiB-L3 machine. The + geometry is accidentally or deliberately LLC-sized; on a normal 32 MiB-L3 part + one cycle already sits at the boundary. → **DISPROVED** as an attack at baseline + geometry; survives only as a fragility (K_open > ~4, or a smaller LLC). + +4. **The parity granule, not the seal granule, sets the incremental-vs-batch + crossover — and it lands at 1.5 % dirty density.** With RS(8,2) and 8 KiB parity + units, full-stripe recompute beats per-touch RMW above δ ≈ 1.46 % (from + `G(1−e^{−D/G})·6g < 1.25·32 MiB`); with 512 B units the crossover moves 14× to + δ ≈ 20.8 %. Since the design's own ruling is that a cycle is *physically sparse*, + the two policies straddle exactly the claimed operating point, and no adaptive + policy exists. → **TRANSFER** (RAID-6 SWP arithmetic + balls-in-bins, applied to + this geometry); the crossover value is new to this design, the method is not. + +5. **Partial-write read-modify-write is the same penalty at three granules — + 64 B write-combining buffer, 512 B row, 8 KiB petal / RS stripe — and a design + that mismatches its write granule to any one of them pays it three times.** The + memory controller's partial-WC-flush RMW, the columnar format's + no-gaps padding, and the parity partial-stripe RMW are one law read at three + scales. → **TRANSFER**. Each granule's penalty is textbook; naming the granule + *ladder* as a single design constraint (choose one write granule that divides + all three) is the transferable part, and it makes an immediate prediction: + the only safe petal size is a multiple of the parity unit which is a multiple of + the row which is a multiple of 64 B — 8 KiB / 512 B / 64 B already satisfies it, + so the ladder is *currently* clean and would be broken by any non-power-of-two + parity unit or by a row size change. + +--- + +## VERDICT + +**The amortisation story is real at the durable boundary and false everywhere +else, and the current implementation inverts its own headline claim.** + +Concretely: + +* **What survives.** One fsync per cycle is genuine and pays off above C ≈ 100 + casts (measured fixed 0.126–0.202 ms against ~2.2 µs/cast marginal). Write-side + canonical ordering genuinely retires an 8.3 ms per-read repair sort. Prepared + fragments with one manifest publication are buildable at lance 9.0.0. +* **What does not.** The coalescing win — the design's signature claim — is + cancelled and reversed in physical bytes by 512 B of null padding per landing + row, with amplification `b+1` (2× at b=1, **65× at b=64**), and the test that + claims to measure it cannot fail. The seal's largest CPU term (52 % at b=64) + does not amortise at all and is the one hash that forbids incremental sealing. + Two full payload deep-copies and a `(1+C+D)·512` builder buffer make peak RSS + ~5× the data. +* **What the LOTUS-shaped alternative loses.** Per-petal fragments at 4096/cycle + cost 910 ms/cycle of filesystem metadata (4.5× the whole-batch fsync) and drive + Lance's every-manifest-lists-every-fragment model into O(K²·F) — making per-cycle + compaction mandatory and repaying the incremental saving immediately. First-come + closure and coalescing are mutually exclusive with exchange rate b. The open-petal + RSS advantage is a restatement of the identity-clustering claim and vanishes + (32 MiB, i.e. parity with whole-batch) under uniform scatter. +* **The decision that actually matters is one nobody has made:** the *parity + granule* and whether parity is defined over the **logical grid** (identity-stable, + repair-local, but partial-stripe RMW at 6× and 85× amplification below + δ ≈ 1.5 %) or over the **physical append batch** (always full-stripe at 1.25×, no + RMW ever, but repair locality tied to write time and destroyed by compaction). + X4 decides it; nothing else should be built until it runs. + +**Recommended order of work, cheapest-first, each with a kill condition already +written above:** X1 (physical bytes — one afternoon, and it either kills F1 or kills +the coalescing claim), X2 (hash swap — ~50 % of seal CPU for a one-line change and +it unblocks design (b)), X4 (parity granule × density — the architectural fork), +then X3/X5/X6/X7. **Do not build incremental petal sealing before X2 and X3.** + +**Confidence.** Source facts (file:line, both repos and lance 9.0.0) — high. +Microbenchmark ratios — medium-high, single machine, single run each, no +statistical treatment; every absolute is platform-bound and the 260 MiB LLC is +atypical. Derived crossovers (C≈100, F≈140, δ≈1.46 %, LLC/2) — medium: the +arithmetic is checkable but the inputs are this machine's constants, and none has +been run against real Lance I/O. diff --git a/docs/lotus/rp-seal-v1/E3.md b/docs/lotus/rp-seal-v1/E3.md new file mode 100644 index 00000000..724bbc99 --- /dev/null +++ b/docs/lotus/rp-seal-v1/E3.md @@ -0,0 +1,186 @@ +# E3 — Systems-Literature Scout: LSM Compaction Design Space vs. Lance/lotus Locality-Debt Compaction + +**Domain E (SoA / cache / amortization), Role: SYSTEMS-LITERATURE SCOUT** +**Program:** *An Experimental Study of Bidirectional Erasure-Coded Seals and Locality-Aware Compaction in Versioned Structure-of-Arrays Grid Stores* + +--- + +## SOURCE ARCHAEOLOGY + +### The mandatory paper's four primitives (Sarkar, Staratzis, Zhu, Athanassoulis, *Constructing and Analyzing the LSM Compaction Design Space*, PVLDB 14(11), 2021, arXiv:2202.04522) + +The paper decomposes **every** LSM compaction strategy into an ensemble of four orthogonal design primitives (§3.1, lines 405–410 of the extracted text): + +1. **Compaction trigger** — *when* to re-organize: level saturation, #sorted-runs threshold, file staleness, space amplification, tombstone-TTL. +2. **Data layout** — *how* data is organized after compaction: leveling (one sorted run/level), tiering (multiple sorted runs/level), 1-leveling, L-leveling, hybrid (per-level choice). +3. **Compaction granularity** — *how much* moves per job: level, sorted-run, single file, multiple files. +4. **Data movement policy** — *which* block moves: round-robin, least-overlap-with-parent, least-overlap-with-grandparent, coldest, oldest, tombstone-density, tombstone-TTL. + +Cardinality of the induced design space with typical primitive value counts: **>10⁴** distinct strategies (§3.2), of which the paper catalogs 24 real systems occupying only a thin slice (Table 1) and benchmarks 10 representative points (Table 2) inside RocksDB v6.11.4. + +Twelve numbered observations (O1–O12) and seven takeaways (TA I–VII) are the paper's empirical payload; the ones most load-bearing for this program: + +- **O1**: `Full` (whole-level, leveled) compaction moves **63×** the ingested data (32× reads + 31× writes); `Tier` (universal/tiered) moves **23×**. Compaction, not ingestion, dominates device-bandwidth consumption. +- **O2 / TA I**: Partial compaction (smaller granularity) cuts data movement 34–56% vs. `Full` but **multiplies compaction-job COUNT** by ≈L (tree depth) because every buffer flush cascades through all levels instead of triggering only every T flushes. *Granularity and trigger frequency trade against each other, not for free.* +- **O3 / TA II**: `Tier` has the lowest mean compaction latency but the **highest tail write latency** (~2.5× `Full`, 5–12× partial-leveling) — eager-write designs push worst-case cost into occasional catastrophic merges. +- **TA III/IV**: Point-lookup latency is *largely independent of the data layout itself* once caches are warm; compactions help lookups mainly by **warming the block cache**, not by the compaction policy per se. +- **O10**: Tiering **scales poorly** vs. leveled/hybrid as data volume grows — the tiering→leveling asymmetry sharpens with scale, matching the Dostoevsky/Monkey asymptotic story below. + +### Lance 9.0.0 (pinned, column A, `/tmp/sources/lance-9`, tag `v9.0.0`, commit `7653c2065d...`) mapped onto the four primitives + +Grepped `rust/lance/src/dataset/optimize.rs` (8,509 lines) directly — this is the actual compaction implementation the "lotus"/Lance store this program targets would sit on top of. + +| Primitive | Lance 9.0.0 mechanism | File:line | +|---|---|---| +| **Trigger** | Threshold ensemble, evaluated **only when `compact_files`/`plan_compaction` is called externally** — no background daemon: (a) `fragment.overlays.len() > max_overlays_per_fragment` (default `Some(10)`) — a literal instance of the paper's **"#sorted runs" trigger type (ii)**; (b) `deletion_percentage() > materialize_deletions_threshold` (default 0.1) — a literal instance of **space-amplification trigger (iv)**; (c) `physical_rows < target_rows_per_fragment` (default 2²⁰) — a small-fragment trigger with no direct analog in the paper's list (closest: an inverse "level saturation," triggering on *under*-fullness rather than over-fullness). | `optimize.rs:741–761` | +| **Data layout** | **Flat, single-tier fragment list** — Lance has no level hierarchy at all. Structurally closest to the paper's `Tier`/"universal compaction" data layout (one flat pool, no leveled cascade) rather than leveling. | `optimize.rs:194–306` (`CompactionOptions`), no level concept anywhere in the crate | +| **Granularity** | **Adjacent-fragment bins**: `DefaultCompactionPlanner::plan` walks the fragment list in id order and accumulates a `CandidateBin` of *physically adjacent* candidate fragments, split to `target_rows_per_fragment` via `bin.split_for_size(...)`. Closest paper analog: "multiple files" granularity — but constrained to **contiguous** fragments only, never a global top-k selection. | `optimize.rs:707–820` | +| **Data movement policy** | **None in the paper's sense.** There is no least-overlap / coldest / oldest / tombstone-density selection — Lance's only "policy" is adjacency-in-list-order plus the three trigger predicates above; index compatibility (`bin.indices == indices`) is the sole non-trivial grouping constraint. This is a real, measurable **absence** relative to the design space the paper documents; 6 of 7 named data-movement policies (round-robin, least-overlap×2, coldest, oldest, tombstone-density) have **no Lance equivalent**. | `optimize.rs:763–806` | + +Two Lance-native mechanisms fall *outside* the paper's four primitives entirely and are directly relevant to the program's own metric list: + +- **`DataOverlayFile`** (`rust/lance-table/src/format/overlay.rs`) — a fragment can carry **multiple overlay files**, each supplying new values for a `(physical offset, field)` subset without rewriting the base data, stored **newest-last**, `committed_version`-ordered, higher-version-wins on conflict (lines 6–34). This is, structurally, **a per-fragment LSM run set living one abstraction level below the paper's unit of analysis** — the paper analyzes whole-dataset LSM trees; Lance nests essentially the same tiered-run/append-then-merge idea *inside a single fragment*, gated by exactly the "#sorted runs" trigger (`max_overlays_per_fragment`). +- **Fragment Reuse Index (FRI)** (`docs/src/format/index/system/frag_reuse.md`, `rust/lance-index/src/frag_reuse.rs`) — when `defer_index_remap=true`, compaction defers the old→new row-address remap into an accumulating index instead of eagerly rewriting every scalar/vector index; the FRI "grows with each compaction" and needs periodic trimming once dependent indices catch up (doc lines 49–57). This is a direct, already-shipped instance of the program's **"index remap bytes/CPU"** and **"stale-version retention bytes"** metrics, with a documented growth/cleanup tension. +- `lance-graph`'s own write path (`crates/lance-graph-planner/src/persist_sink.rs`, `batch_writer.rs`) contains **zero compaction logic of its own** — it is a pure single-writer WAL-append sink (`CommitOutcome`/`CommitError`, one commit per 64K-row cycle) that delegates all physical reorganization to Lance's `optimize.rs`. There is currently **no locality-aware or debt-triggered compaction anywhere in this stack** — the proposed mechanism would be wholly new engineering, not a variant of something already partially built. + +--- + +## MECHANISM + +The proposed idea under test — *"locality-debt triggered, locality-key ordered, incremental compaction"* — decomposes cleanly into the paper's own four slots, which is itself useful (see SURPRISES): + +1. **Trigger** = a *locality-debt* metric (some function of index-remap entropy / scatter of physically co-accessed rows across fragments/files) crossing a threshold, replacing size/count/staleness/space-amp/TTL. +2. **Layout** = fragments physically ordered by a **locality key** (e.g., a space-filling-curve coordinate over the grid's 2-D hierarchy) rather than append/insertion order. +3. **Granularity** = **incremental** — small, continuously-amortized rewrites rather than periodic full re-sorts. +4. **Data movement policy** = pick the fragment(s)/rows whose locality-key rank is most "out of place" relative to their neighbors (a debt-repair policy, structurally a new named policy not in the paper's list of 7). + +The broader literature converges on this same four-part shape from several independent directions: + +### Write-amplification theory (Monkey, Dostoevsky) + +- **Monkey** (Dayan, Athanassoulis, Idreos, SIGMOD'17) shows worst-case lookup cost ∝ Σ(Bloom-filter false-positive rates across levels), and that the *optimal* allocation of a fixed memory budget across per-level filters is **non-uniform** — larger levels get lower bits/key. Empirically 50–80% lookup-latency reduction over LevelDB's uniform allocation as data volume grows. +- **Dostoevsky** (Dayan, Idreos, SIGMOD'18) generalizes leveling/tiering into the **Fluid LSM-tree**: a continuum parameterized independently at each level, and introduces **Lazy Leveling** — leveled merging ONLY at the largest level, tiered (lazy) merging everywhere else — which *improves the asymptotic update cost* while holding point-lookup, long-range-lookup, and space-amplification bounds fixed. The mechanism directly generalizes both the paper's "1-leveling"/"L-leveling"/"hybrid" data-layout options and is the closed-form answer to "how eager should merging be, level by level." + +Both papers make the write/read/space three-way trade-off **quantitative** rather than qualitative — this is the missing piece the design-space paper itself only measures empirically (O1–O12) without deriving. + +### Adaptive, workload-driven reorganization (database cracking, adaptive merging) + +- **Database cracking** (Idreos, Kersten, Manegold) reorganizes a column *as a side effect of query execution* — each range-selection predicate physically partitions the column further ("incremental quicksort"), with **zero explicit administration** and continuous convergence toward a sorted index driven purely by which queries actually ran. +- **Adaptive merging** (Graefe & Kuno) is the merge-sort analog: "incremental external merge sort," partition/merge logic layered on top of cracking. The synthesis paper (Idreos, Manegold, Kuno, Graefe, *"Merging What's Cracked, Cracking What's Merged"*) unifies both into a spectrum from very-active to very-lazy adaptive indexing, with an explicit acceptance criterion from the companion benchmarking paper (Graefe, Idreos, Kuno, Manegold): **(a) lightweight initialization** (first-query cost bounded against a full scan) and **(b) fast convergence** (bounded against a full index build) — *both* must hold, not just one (extracted, arXiv:1203.0055 lines 337–344). + +This is the direct literature precedent for a **debt-triggered, incremental** (never full-rebuild) compaction: the acceptance criterion above is exactly the falsifiable form the EXPERIMENT section below borrows for kill conditions. + +### Cache-oblivious structures (Frigo et al.) and the packed memory array + +- The **ideal-cache model** (Frigo, Leiserson, Prokop, Ramachandran, 1999) analyzes algorithms against an unknown cache of size M, block size B, optimal replacement — an algorithm ignorant of M, B that still achieves the *cache-aware-optimal* transfer count is "cache-oblivious." +- The **van Emde Boas (vEB) memory layout** (Prokop) achieves O(log_B N) search transfers for a *static* tree by recursively laying subtrees in contiguous memory. +- The **packed memory array (PMA)** (Bender et al., building on Itai et al.) is the dynamic counterpart: N ordered elements stored in a contiguous array **interleaved with gaps**, sized so density stays in a target band; insert/delete costs **O(1 + (log²N)/B) amortized** memory transfers while preserving the same asymptotic locality as a static sorted array (extracted, arXiv:2209.09166 lines 74–87). This is precisely the abstract problem *"maintain a physically sorted-by-key structure under continuous insertion, cheaply, amortized"* — i.e., the theoretical tool for exactly what "incremental locality-key ordered compaction" is trying to build ad hoc. +- The **Cache-Oblivious B-Tree** (Bender et al.) stores the *whole search tree* (not just leaves) in a PMA ordered by vEB layout, giving O(log_B N) search + O((log²N)/B) amortized update — I/O-optimal on both axes simultaneously, without knowing B. + +### Engineering practice: clustering keys / sort orders / Z-order (industrial, not academic — engineering blogs cited as sources per instructions) + +- **Snowflake Automatic Clustering** — micro-partitions (50–500 MB uncompressed each) are re-clustered by a background *service*, triggered continuously, that "only performs automated maintenance if the table will benefit"; Snowflake's own docs explicitly warn reclustering "can incur substantial costs" and to evaluate trade-offs before enabling. [Snowflake docs — clustering keys, automatic clustering, micro-partitions] +- **Apache Iceberg** — `sort order` is table-metadata that engines/compaction honor; `rewrite_data_files` optionally applies **Z-order (Morton-curve) clustering** across multiple columns simultaneously — explicitly documented as **more expensive** than plain bin-packing compaction because it requires an actual multi-dimensional sort, not a linear one. [Dremio, IOMETE, Alex Merced/iceberglakehouse engineering posts] +- **Delta Lake `OPTIMIZE ... ZORDER BY`** — colocates rows so column min/max statistics (data-skipping) become tight and non-overlapping per file; explicitly noted that Z-order **effectiveness degrades per additional column** and is wasted on columns without collected stats. [Delta Lake docs, Databricks docs, Salesforce Engineering blog] +- **Peloton / NoisePage self-driving DBMS** (Pavlo et al., CIDR'17; CMU DB Group) — the direct precedent for **workload-forecast-triggered** physical-layout adaptation (hybrid row/column layout, index/materialized-view selection) rather than size/age-triggered; NoisePage's successor architecture explicitly frames this as cost/benefit-modeled tree search (MCTS/RL) over candidate physical configurations, not a fixed heuristic trigger. + +All four industrial systems independently confirm the same empirical fact the design-space paper measures analytically (O2/TA I): **maintaining a sort/clustering order under continuous ingestion costs materially more than plain size-based compaction, and every vendor treats it as an explicit, separately-priced, opt-in operation** — never the unconditional default. + +--- + +## EXPECTED BENEFIT + +If "locality-debt triggered, locality-key ordered, incremental compaction" is built as literature suggests it should be (Lazy-Leveling-style asymmetric eagerness + PMA-style amortized incremental reinsertion + cracking-style workload-driven triggering rather than a periodic daemon): + +- **Bounded, provable write amplification** in the Dostoevsky/Fluid-LSM sense — the "locality-key" dimension becomes analogous to the tiering/leveling axis parameterized *per grid tier* (the program's own 4/16/64/256/1024/4096 hierarchy is a natural per-level Fluid-LSM parameterization target), rather than an unbounded global re-sort. +- **Reduced query-time file/page-touch counts** for spatially/temporally local queries — directly evidenced by Iceberg/Delta Z-order case studies and by the design-space paper's O1 observation that compaction's main payoff is bounding read amplification, not write throughput. +- **Amortized incremental cost** in the strict PMA sense (O(1 + log²N/B) per insertion) IF the locality-key ordering is built on a packed-array-style substrate rather than ad hoc file rewriting — this is a concrete, citable asymptotic target the program can hold the design to. +- **A trigger with an actual falsifiable acceptance test** — cracking/adaptive-merging's two-part criterion (bounded initialization cost vs. full scan; bounded convergence cost vs. full rebuild) is a ready-made pass/fail bar, not something the program needs to invent from scratch. + +## EXPECTED FAILURE + +- **The trigger has no daemon to live in.** Lance 9.0.0's compaction is externally invoked (`compact_files`/`plan_compaction`), and `lance-graph`'s `persist_sink.rs` is a pure append-only WAL sink with zero background reorganization logic. A "locality-debt triggered" scheme needs either (a) a new always-on watcher process (contradicting the workspace's own I-2 std::sync-not-tokio hot-path discipline noted in OGAR's `TEMPORAL-TIME-TRAVEL.md`), or (b) piggybacking the check onto every write cycle, which the design-space paper's O2/TA I already shows **multiplies compaction-job count** relative to a periodic/threshold trigger. +- **"Locality-key ordered" pulls toward Leveling's worst case.** Maintaining a *single* global physical order under continuous insertion is structurally the Dostoevsky/Fluid-LSM **leveling extreme** (merge at every level to preserve one sorted run) — the design-space paper's own numbers (O1: `Full` moves 63× ingested bytes) and every industrial source (Snowflake/Iceberg/Delta) independently confirm this is the *most expensive* point in the space, not a free win. An "incremental" claim does not remove this — TA I is explicit that partial/incremental granularity reduces peak burst cost but does **not** reduce the *total* data movement, which remains dominated by how eager the ordering constraint is, not by batch size. +- **A locality-debt metric with no established closed form.** Nothing found in the searched literature computes compaction eligibility from an *entropy/scatter-of-index-remap* signal; every trigger in the paper's taxonomy (and every industrial system) is size-, count-, age-, or space-amp-based. Absence of precedent means the metric's cost (to compute AND to keep fresh) is itself unmeasured — it could dominate the very I/O it's meant to save. +- **Fragment Reuse Index growth is already a documented liability at the trigger's neighbor.** Lance's own docs (`frag_reuse.md` lines 49–57) state the FRI "grows with each compaction" and needs a **separate periodic trim**, i.e., deferred-remap compaction already trades one unbounded-growth structure for another. A locality-debt trigger firing *more often* than size-based compaction (because locality decays faster than fragment size grows) would accelerate this exact liability. + +--- + +## EXPERIMENT + +Pre-registered designs, using the program's metric list. All experiments assume the actual pinned Lance 9.0.0 `optimize.rs` mechanism as baseline (grounded above), not an idealized LSM. + +### EXP-E3-1 — Replicate O1/O2/TA-I on Lance's real trigger knobs + +- **Metric:** write amplification (physical bytes written / logical bytes ingested), fragment count, compaction CPU/wall, index-remap bytes/CPU. +- **Baseline:** `CompactionOptions::default()` (`target_rows_per_fragment=2²⁰`, `max_overlays_per_fragment=Some(10)`, `materialize_deletions_threshold=0.1`). +- **Control:** sweep `target_rows_per_fragment` and `max_overlays_per_fragment` independently across ≥5 values each (the design-space paper's own "Full vs. partial" and "#sorted-runs trigger" axes), holding the other fixed. +- **Kill condition:** if write amplification does **not** decrease monotonically as `max_overlays_per_fragment` decreases (i.e., tighter "#sorted-runs" trigger fails to reduce amplification the way O2 predicts for the paper's own systems), the direct transfer of O1/O2 to Lance's overlay mechanism is **falsified** and the "DataOverlayFile ≈ per-fragment LSM run" analogy (see SURPRISES) needs revision before any locality-debt work builds on it. + +### EXP-E3-2 — Locality-key ordering cost vs. Fluid-LSM eagerness prediction + +- **Metric:** write amplification, point/neighborhood-lookup latency, files & pages touched per query, peak RSS. +- **Baseline:** current adjacency-only bin-packing (no locality key at all). +- **Treatment arms:** (a) full Z-order-style rewrite on every compaction (the "Leveling extreme" per EXPECTED FAILURE); (b) Dostoevsky-style Lazy-Leveling analog — locality-key ordering enforced only at the *coarsest* grid tier (e.g., the 4096→1024 reduction level), left tiered/unordered at finer tiers; (c) PMA-style incremental reinsertion maintaining local order without ever doing a full rewrite. +- **Kill condition:** if (c)'s measured write amplification does not asymptotically undercut (a) as row count N grows (i.e., does not trend toward the PMA's O(log²N/B) bound relative to (a)'s O(N) full-rewrite cost), the cache-oblivious transfer (Frigo/Bender) does not hold at this system's actual block size/fragment granularity, and the "incremental" framing is cosmetic rather than substantive — falsifies the EXPECTED BENEFIT claim. +- **Kill condition 2 (cross-check against O10):** if arm (a)/(b)'s neighborhood-query latency does NOT improve materially over the untuned baseline as N grows, the design-space paper's own O10 finding ("Tier scales poorly vs leveled/hybrid") transfers unfavorably — locality-key ordering may need to be leveled-at-every-tier to pay off at all, which reopens the write-amplification risk from EXPECTED FAILURE rather than resolving it. + +### EXP-E3-3 — Locality-debt trigger vs. cracking's two-part acceptance test + +- **Metric:** compaction CPU/wall (initialization cost) vs. a full-scan baseline; convergence steps/wall-time to reach steady-state query latency vs. a full-index baseline (the exact Graefe/Idreos/Kuno/Manegold criterion, arXiv:1203.0055 lines 337–344). +- **Baseline:** Lance's current threshold triggers (size/overlay-count/deletion%). +- **Treatment:** a candidate locality-debt metric (e.g., normalized entropy of the physical-offset↔locality-key rank correlation per fragment window). +- **Kill condition:** if the locality-debt metric's own computation cost (CPU per evaluation × evaluation frequency) exceeds a fixed fraction (pre-register: 5%) of the compaction CPU it is meant to save, OR if first-query "lightweight initialization" cost exceeds a full scan of the affected fragments, the metric fails cracking's own acceptance bar and the trigger primitive is not viable as specified — this is the sharpest, cheapest kill test to run first. + +--- + +## PRIOR ART + +Searches actually run (alphaXiv `discover_papers` + `WebSearch`, this session): + +1. `discover_papers(["write amplification","Dostoevsky","Monkey","LSM tree","tiered leveled compaction"], prioritize=historical)` — surfaced 2202.04522 (mandatory paper) plus 9 follow-on LSM papers (Bloom-filter adaptivity, learned Bloom filters, key-value separation, ML-driven compaction characterization) — none pre-date or supersede Monkey/Dostoevsky as the canonical space-time-tradeoff references. +2. `discover_papers(["database cracking","adaptive merging","adaptive indexing"], prioritize=historical)` — surfaced 1203.0055 (Stochastic Database Cracking, Idreos/Kersten/Manegold lineage) as the strongest historical-adjacent source with a full related-work section citing the original cracking, adaptive-merging, and unifying papers directly (retrieved and read). +3. `discover_papers(["cache-oblivious algorithms","cache-oblivious B-tree","Frigo"], prioritize=historical)` — surfaced 2209.09166 (Cache-Oblivious Representation of B-Tree Structures) with a clean restatement of the Frigo et al. ideal-cache model, vEB layout, and PMA bounds (retrieved and read). +4. `WebSearch("Dostoevsky LSM tree Dayan Idreos SIGMOD 2018 ...")` and `WebSearch("Monkey optimal navigable key-value store ...")` — Dostoevsky/Monkey are **not** on arXiv under those titles (checked: alphaXiv discovery did not surface them directly by title search either); primary sources are the authors' own hosted PDFs (nivdayan.github.io) and ACM DL, cross-confirmed via Harvard's Stratos Idreos publication pages — used as the authoritative summaries since the arXiv-first sourcing instruction has no arXiv ID to target for these two specific papers. +5. `WebSearch("Snowflake automatic clustering micro-partitions ...")`, `WebSearch("Apache Iceberg sort order compaction rewriteDataFiles z-order ...")`, `WebSearch("Delta Lake OPTIMIZE ZORDER BY ...")` — engineering-blog/vendor-doc sources, per the task's explicit allowance ("engineering blogs count as sources here"). +6. `WebSearch("Peloton self-driving database NoisePage adaptive physical design ...")` and a follow-up `WebSearch` for the CIDR'17 paper directly — confirmed Peloton (CIDR'17) → NoisePage (2018 successor) lineage and the workload-forecast-triggered (not threshold-triggered) physical-design-adaptation framing. +7. Direct source-code archaeology (not a literature search but load-bearing for PRIOR ART's "what does this system actually do today" half): grepped and read `rust/lance/src/dataset/optimize.rs` (pinned v9.0.0, `/tmp/sources/lance-9`), `rust/lance-table/src/format/overlay.rs`, `docs/src/format/index/system/frag_reuse.md`, and `crates/lance-graph-planner/src/{persist_sink,batch_writer}.rs` in `/home/user/lance-graph`. Confirmed `overlay.rs` is present byte-identical-in-spirit in both the pinned v9.0.0 tag and current `/tmp/sources/lance-main` (both checked directly), so the overlay/FRI mechanism is not a post-9.0.0 addition being mistakenly attributed to the pin. + +### Per-source extracted-fields table + +| Source | Trigger | Layout | Granularity | Movement/reorg policy | Key quantitative claim | +|---|---|---|---|---|---| +| Design-space paper (2202.04522) | 5 named types (§3.1.1) | leveling/tiering/hybrid (§3.1.2) | level/sorted-run/file/multi-file (§3.1.3) | 7 named policies (§3.1.4) | Full: 63× data movement; partial: −34–56%; O10: tiering scales poorly | +| Monkey (SIGMOD'17) | space-amp / lookup-cost driven filter re-tuning | leveling (paper's baseline) | N/A (memory allocation, not data movement) | Bloom-filter bits/level optimized, non-uniform | 50–80% point-lookup latency reduction vs. uniform allocation, at scale | +| Dostoevsky (SIGMOD'18) | merge-eagerness parameter per level | **Fluid LSM-tree**: continuum, Lazy Leveling = tiered-except-largest-level | per-level tunable | implicit in eagerness parameter | improves worst-case update complexity while holding lookup/space bounds fixed | +| Database cracking (1203.0055 + cited lineage) | query-predicate-driven (implicit, continuous) | column-store, in-place partition | per-query, incremental | "crack" at predicate boundary, no external admin | acceptance test: init cost ≤ full scan; convergence ≤ full index build | +| Cache-oblivious B-tree/PMA (2209.09166, Frigo et al.) | N/A (structural invariant, not a "trigger") | PMA-ordered, vEB-recursive | per-insert/delete | density-band rebalancing | O(log_B N) search, O(1+log²N/B) amortized update | +| Snowflake clustering (vendor docs) | continuous background service, cost-gated | micro-partition, key-defined | per-micro-partition (50–500MB) | re-cluster by key value proximity | explicitly costed, opt-in, "substantial cost" warning | +| Iceberg sort-order/Z-order (vendor/blog) | manual `rewrite_data_files` invocation | table-metadata sort order, optional Z-curve | file-level rewrite, target-sized (512MB–1GB) | multi-dim Morton-order clustering | Z-order "more expensive" than plain bin-packing (explicit trade-off) | +| Delta `OPTIMIZE ZORDER` (vendor/blog) | manual invocation | Z-order within partition | per-partition file merge | min/max stat tightening | effectiveness "drops with each extra column" | +| Peloton/NoisePage (CIDR'17 lineage) | workload-forecast (ML), not threshold | hybrid row/column, adaptive | index/materialized-view/layout selection | cost/benefit search (MCTS/RL) | first *learned-trigger* precedent, not a fixed heuristic | +| Lance 9.0.0 `optimize.rs` (this repo, pinned) | overlay-count / deletion% / under-size (§ table above) | flat fragment pool (≈"tiered", no levels) | adjacent-fragment bins | **none** (adjacency + index-compat only) | `max_overlays_per_fragment` default 10 = literal "#sorted runs" trigger instance | + +--- + +## SURPRISES + +1. **Lance's `DataOverlayFile` is a per-fragment LSM run set, one abstraction level below what the design-space paper analyzes.** — **TRANSFER.** The paper's taxonomy operates on whole-dataset LSM trees (levels of files); Lance nests functionally the same "append new run, defer merge, higher-version-wins" idea *inside a single fragment*, and the trigger that flushes it (`max_overlays_per_fragment=10`) is a byte-for-byte instance of the paper's trigger-type (ii) "#sorted runs." This was not found stated anywhere in the searched literature or in Lance's own docs (`frag_reuse.md` discusses the FRI's cost implications but never frames overlays as "LSM runs"); it is a structural observation from reading the source directly, not a claim from prior art — flagging as TRANSFER rather than NOVEL-CANDIDATE because the underlying pattern (deferred-merge append log) is thoroughly KNOWN, just not previously named at this granularity for this system. + +2. **The design-space paper's own O10 ("tiering scales poorly vs. leveled/hybrid") plus every industrial Z-order source's cost warnings triangulate to the same prediction: an unconditionally "locality-key ordered" store degrades toward the *worst* point in the published space, not a novel improvement on it.** — **KNOWN.** Three independent lineages (academic LSM measurement, LSM theory, industrial engineering blogs) agree; this is not new information, it's convergent confirmation, and it directly sharpens the EXPECTED FAILURE section's kill conditions. + +3. **Nothing in the searched literature computes compaction eligibility from an index-remap-entropy / locality-debt signal — every trigger found is size/count/age/space-amp-based, including the workload-*forecast*-driven Peloton/NoisePage line (which triggers on predicted future access patterns, not on a measured scatter/entropy of current physical layout).** — **NOVEL-CANDIDATE**, conditionally. Search performed: `discover_papers` calls #1–3 above plus targeted re-reads of the design-space paper's full trigger list (§3.1.1) and Monkey/Dostoevsky's trigger framing (merge-eagerness/space-amp, not entropy) found no closed-form "locality debt" metric anywhere. This is the one place in the program's hypothesis space this scout's search genuinely could not find prior art for — but the search was narrow (LSM + cracking + cache-oblivious + 4 vendor blogs); a broader search specifically in the spatial-index / GIS-clustering-quality literature (Hilbert-curve clustering metrics, fragmentation indices in spatial DBs) was **not** run this session and should be treated as an open gap before this is promoted past NOVEL-CANDIDATE. + +4. **The packed memory array (PMA) is, almost too literally, the pre-existing theoretical answer to "maintain a locality-key sorted order under continuous insertion, incrementally, with a proven amortized bound" — yet nothing in the LSM-compaction or Lakehouse-clustering literature searched this session cites or builds on it.** — **TRANSFER**, and a strong one: the PMA's O(1+log²N/B) amortized-update bound is a directly citable, falsifiable target for EXP-E3-2's kill condition, imported wholesale from a 25-year-old cache-oblivious-algorithms result into a domain (Lakehouse/LSM compaction) that currently solves the same problem with heuristic, unbounded-cost full rewrites (Z-order) instead. + +5. **Database cracking's two-part acceptance criterion (bounded init cost vs. full scan; bounded convergence vs. full rebuild) is directly reusable, unmodified, as the falsifiability test for a "locality-debt trigger" — the cracking literature already solved "how do we know an adaptive/incremental reorganization scheme is actually worth it" for a structurally analogous problem (query-adaptive column reorganization) three technique-generations before this program's proposal.** — **TRANSFER.** Confirmed by direct extraction from the Graefe/Idreos/Kuno/Manegold benchmarking paper as cited in 1203.0055 (lines 337–344); no evidence found that any LSM-compaction paper (including the mandatory one) states an equivalent falsifiable acceptance bar for its own triggers — the design-space paper *measures* triggers empirically (O1–O12) but never proposes a pass/fail test for a *new* trigger, which cracking's literature does. + +--- + +## VERDICT + +The four-primitive taxonomy (trigger/layout/granularity/movement-policy) from the mandatory paper maps onto Lance 9.0.0's actual, pinned `optimize.rs` almost cleanly — with one structurally real absence (no data-movement policy beyond adjacency) and one structurally real extra (per-fragment `DataOverlayFile` runs, a nested LSM the paper's own taxonomy does not anticipate). Every literature-adjacent piece of the proposed "locality-debt triggered, locality-key ordered, incremental compaction" idea is KNOWN or TRANSFER — the write-amplification tradeoff (Dostoevsky/Monkey), the incremental-reorganization acceptance test (cracking/adaptive merging), and the amortized-bound target for incremental ordering (PMA/cache-oblivious) are all directly reusable, well-established results, not gaps. The one place a documented search came back empty — a locality-debt/entropy-of-scatter trigger metric, as opposed to size/count/age/space-amp/forecast triggers — is a genuine, narrowly-scoped NOVEL-CANDIDATE, but it is also exactly the piece EXP-E3-3's kill condition is built to falsify cheaply and first, before any engineering investment in the layout/granularity/movement machinery (which the literature already predicts, convergently across three independent lineages, will be expensive if built as an unconditional global reordering). + +--- + +*No model identifiers appear in this document. Report authored per the E3 (SYSTEMS-LITERATURE SCOUT) research-cell brief.*