Skip to content

feat(index): configure IVF shuffle offset preload budget - #9434

Merged
jackye1995 merged 2 commits into
lance-format:mainfrom
jackye1995:codex/configure-ivf-shuffle-offset-preload
Sep 19, 2026
Merged

jackye1995 merged 2 commits into
lance-format:mainfrom
jackye1995:codex/configure-ivf-shuffle-offset-preload

Conversation

@jackye1995

@jackye1995 jackye1995 commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Allow IVF shuffle offset preloading to be tuned with LANCE_SHUFFLE_MAX_PRELOADED_OFFSETS_BYTES. The default remains 256 MiB; 536870912 allows 512 MiB and 0 disables preloading. Reject malformed values, apply the setting to writer handoff and reopened readers, and document its per-shuffle memory cost.

Replayed the same complete 1B-row IVF_RQ shuffle: 31,623 partitions and 381 MiB decoded offsets, with 72 CPU threads and unchanged merge budgets.

Merge measurement #9431 + this PR, 256 MiB #9431 + this PR, 512 MiB This PR alone, 512 MiB
Wall time 61.01 min 13.38 min 13.50 min
CPU time 15,726 s 2,038 s 2,058 s
Average CPU cores 4.30 2.54 2.54
Observed RSS high-water mark 14.21 GiB 9.64 GiB 9.43 GiB

The final column excludes both #9431 changes: it uses the original sorted-vector scheduler and sequential structural-page loading. All three full outputs passed exact row-ID/partition-count checks and code/factor fingerprint checks. The standalone run also captured a successful process exit and complete resource accounting.

At 512 MiB, this input's offset table fits in memory and enables coalesced partition reads. This PR alone completed within 1% of the elapsed and CPU time measured with both PRs. This establishes that #9431 is unnecessary for this particular preloaded replay; it does not fix the original fallback path when offsets exceed the configured cap. The 61-minute baseline already includes #9431 and is not an unfixed baseline.

These are single measurements with warm caches: the first two ran sequentially on one worker, and the standalone run used a matching instance in us-west-2. Background input writeback continued during the 256 MiB baseline; the two 512 MiB runs began after it finished. All merges recorded zero physical input reads. Startup and output auditing are excluded; memory uses one-second RSS high-water observations. Small differences between the 512 MiB runs should not be treated as established performance effects.

Formatting, all 39 shuffle tests, full-workspace Rust clippy, and all 37 CI checks passed.

@github-actions github-actions Bot added A-index Vector index, linalg, tokenizer A-docs Documentation enhancement New feature or request labels Sep 19, 2026

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Gate recommendation: approve.

The hard-coded 256 MiB threshold can push large shuffles back to per-partition reads. This keeps the default and file behavior unchanged while giving operators control over the per-shuffle offset-table allowance for both newly written and reopened shuffles. Keeping this opt-in is preferable to raising memory for every workload, and coverage exercises fallback, coalescing, invalid input, reopening, and row preservation.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 19, 2026
@jackye1995
jackye1995 merged commit 1f7b847 into lance-format:main Sep 19, 2026
44 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

A-docs Documentation A-index Vector index, linalg, tokenizer enhancement New feature or request K-approved Latest Gatekeeper recommendation permits acceptance.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants