feat: reduce open file handles during IVF training - #6169
Conversation
PR ReviewNice improvement — reducing open file descriptors from O(partitions) to O(1) is a meaningful usability win for large-scale IVF training. The two-file design with sorted data + cumulative offsets is clean and well-tested. A few items to consider: P1: Potential u32 overflow in
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
| let offsets = offsets_stream.try_collect::<Vec<_>>().await?; | ||
| let offsets = if offsets.len() == 0 { | ||
| // We should not hit this path if there is no batches | ||
| unreachable!() |
There was a problem hiding this comment.
should it panic with a message?
There was a problem hiding this comment.
unreachable panics with the message "I didn't expect to get here" and the line number which is pretty much what I'd use if I did an explicit panic.
|
|
||
| let mut ranges = Vec::with_capacity(self.num_batches as usize); | ||
| for batch_idx in 0..self.num_batches { | ||
| if batch_idx == 0 && is_uneven { |
There was a problem hiding this comment.
is the is_uneven stuff equivalent to checking if partition_id = 0 and batch_idx = 0?
There was a problem hiding this comment.
Yes, that's maybe simpler. Updated.
This is a legitimate concern. Unfortunately, we don't actually support take with u64 offsets. This is because we're reusing some old paths from the v0.1 days where we used u32 offsets. It's all fixable but more follow-up territory. By my math I think, even with 1536-dimension vectors, we are ok until we hit trillions of rows. 1T rows => sqrt(1T) * 96 * 1T bytes of PQ data => 96EB of data which, split into 128MB chunks, would give us ~1B chunks. Either way, I turned it into a |
Xuanwo
left a comment
There was a problem hiding this comment.
Thank you for this PR! Only two small questions.
| /// | ||
| /// This default is likely to be fine for most use cases. | ||
| fn shuffle_batch_bytes() -> usize { | ||
| std::env::var("LANCE_SHUFFLE_BATCH_BYTES") |
There was a problem hiding this comment.
I think it's a good idea to add some protection here. This would prevent users from setting LANCE_SHUFFLE_BATCH_BYTES to 0, which could lead to using batch_size_bytes = 0 and producing incorrect results.
There was a problem hiding this comment.
Good idea, now it will log a warning and use the default.
| let mut partition_counts = vec![0u64; np]; | ||
| for i in 0..part_ids.len() { | ||
| let pid = part_ids.value(i) as usize; | ||
| if pid < np { |
There was a problem hiding this comment.
Is it possible for pid >= np? It looks like we will just write those data without offsets. global_row_count will always include them.
There was a problem hiding this comment.
It shouldn't be possible (np is num_partitions) but I now log a warning instead of silently ignoring.
a692013 to
203b465
Compare
…num_partitions file handles Actually write the offsets to a file and don't accumulate Change progress reporting to report number of rows shuffled and not number of batches processed Address review suggestions More PR suggestions Remove dead code Address PR review
203b465 to
08715e5
Compare
|
CI failure seems unrelated. Will merge if remaining CI job passes. |


The previous shuffler used one open file per partition. At large scales this meant tedious re-adjusting of OS limits. The new shuffler uses two open files. We potentially introduce a bit more random access in the later read phase but the overall performance hasn't changed significantly.