Bug: scans of COPY-loaded, not-yet-checkpointed CSR rel groups fail with read past EOF
Ladybug version
main @ 08915c1 (observed on a PR branch based on it; the failing code paths are unmodified main code)
What operating system are you using?
linux (x86_64); observed in CI minimal-test (RelWithDebInfo)
What happened?
After batch-COPY-loading rels with no checkpoint in between, a plain forward scan fails:
Assertion failed in file "src/common/file_system/local_file_system.cpp" on line 423:
localFileInfo->getFileSize() >= position + numBytes
(In release builds the DASSERT is compiled out, but the short read still throws IOException, so this fails loudly everywhere.)
Are there known steps to reproduce?
CALL force_checkpoint_on_close=false;
CALL auto_checkpoint=false;
CREATE NODE TABLE N(id INT64 PRIMARY KEY);
CREATE REL TABLE R(FROM N TO N, w INT64);
UNWIND range(0, 999) AS i CREATE (:N {id: i});
CHECKPOINT;
-- rels.csv: 1000 rows `i,(i+1)%1000,i` (endpoints match N PKs)
COPY R FROM 'rels.csv';
MATCH (a:N), (b:N) WHERE a.id = 0 AND b.id = 500 CREATE (a)-[:R {w: 1000}]->(b);
MATCH (a:N)-[e:R]->(b:N) RETURN COUNT(e), CAST(SUM(e.w) AS INT64);
-- ^ fails with the read-past-EOF assertion above; expected 1001 / 500500
A failed-then-retried checkpoint in the same state also misbehaves: the checkpoint's own
scans of the flushed state hit the same broken pages and error out before any page
allocation, so failure injection keyed on allocation never fires.
Why this looks pre-existing (not caused by #1051 / #1058)
No checkpoint runs between the COPY and the failing read (auto_checkpoint=false, no
explicit CHECKPOINT), so none of the CSR checkpoint-rollback code executes on this path.
The flushed state is produced entirely by unmodified code (RelBatchInsert flush paths).
Contrasting shapes that work:
- 3-rel PK-matched
COPY + immediate scan (CopyTest.RelCopyWithoutDefaultHashIndex) passes.
- The same 1000 rels loaded via single-row
CREATEs + CHECKPOINT + scan passes.
- Reads after a successful checkpoint of the
COPY-loaded data pass (the checkpoint
rewrite appears to normalize whatever is off).
So the gap is specifically: multi-page flushed (not yet checkpointed) CSR state vs the
scan/checkpoint-scan readers. Hypothesis (unconfirmed): the flushed chunks' page-range
metadata disagrees with what is actually on disk (single-page flushes stay within file
slack/preallocation, which is why tiny cases pass). Candidate code:
Column::flushData / ColumnChunkData::flushBuffer page math vs
initScanForCommittedPersistent / scanCSRHeader page reads.
Expected
Reads (and checkpoint scans) of flushed-but-uncheckpointed CSR groups work, or the flush
path produces state identical in observable behavior to checkpointed state.
Bug: scans of COPY-loaded, not-yet-checkpointed CSR rel groups fail with read past EOF
Ladybug version
main @ 08915c1 (observed on a PR branch based on it; the failing code paths are unmodified main code)
What operating system are you using?
linux (x86_64); observed in CI minimal-test (RelWithDebInfo)
What happened?
After batch-
COPY-loading rels with no checkpoint in between, a plain forward scan fails:(In release builds the
DASSERTis compiled out, but the short read still throwsIOException, so this fails loudly everywhere.)Are there known steps to reproduce?
A failed-then-retried checkpoint in the same state also misbehaves: the checkpoint's own
scans of the flushed state hit the same broken pages and error out before any page
allocation, so failure injection keyed on allocation never fires.
Why this looks pre-existing (not caused by #1051 / #1058)
No checkpoint runs between the
COPYand the failing read (auto_checkpoint=false, noexplicit
CHECKPOINT), so none of the CSR checkpoint-rollback code executes on this path.The flushed state is produced entirely by unmodified code (
RelBatchInsertflush paths).Contrasting shapes that work:
COPY+ immediate scan (CopyTest.RelCopyWithoutDefaultHashIndex) passes.CREATEs +CHECKPOINT+ scan passes.COPY-loaded data pass (the checkpointrewrite appears to normalize whatever is off).
So the gap is specifically: multi-page flushed (not yet checkpointed) CSR state vs the
scan/checkpoint-scan readers. Hypothesis (unconfirmed): the flushed chunks' page-range
metadata disagrees with what is actually on disk (single-page flushes stay within file
slack/preallocation, which is why tiny cases pass). Candidate code:
Column::flushData/ColumnChunkData::flushBufferpage math vsinitScanForCommittedPersistent/scanCSRHeaderpage reads.Expected
Reads (and checkpoint scans) of flushed-but-uncheckpointed CSR groups work, or the flush
path produces state identical in observable behavior to checkpointed state.