Skip to content

Bug: scans of COPY-loaded, not-yet-checkpointed CSR rel groups fail with read past EOF #1059

Description

@adsharma

Bug: scans of COPY-loaded, not-yet-checkpointed CSR rel groups fail with read past EOF

Ladybug version

main @ 08915c1 (observed on a PR branch based on it; the failing code paths are unmodified main code)

What operating system are you using?

linux (x86_64); observed in CI minimal-test (RelWithDebInfo)

What happened?

After batch-COPY-loading rels with no checkpoint in between, a plain forward scan fails:

Assertion failed in file "src/common/file_system/local_file_system.cpp" on line 423:
localFileInfo->getFileSize() >= position + numBytes

(In release builds the DASSERT is compiled out, but the short read still throws IOException, so this fails loudly everywhere.)

Are there known steps to reproduce?

CALL force_checkpoint_on_close=false;
CALL auto_checkpoint=false;
CREATE NODE TABLE N(id INT64 PRIMARY KEY);
CREATE REL TABLE R(FROM N TO N, w INT64);
UNWIND range(0, 999) AS i CREATE (:N {id: i});
CHECKPOINT;
-- rels.csv: 1000 rows `i,(i+1)%1000,i` (endpoints match N PKs)
COPY R FROM 'rels.csv';
MATCH (a:N), (b:N) WHERE a.id = 0 AND b.id = 500 CREATE (a)-[:R {w: 1000}]->(b);
MATCH (a:N)-[e:R]->(b:N) RETURN COUNT(e), CAST(SUM(e.w) AS INT64);
-- ^ fails with the read-past-EOF assertion above; expected 1001 / 500500

A failed-then-retried checkpoint in the same state also misbehaves: the checkpoint's own
scans of the flushed state hit the same broken pages and error out before any page
allocation, so failure injection keyed on allocation never fires.

Why this looks pre-existing (not caused by #1051 / #1058)

No checkpoint runs between the COPY and the failing read (auto_checkpoint=false, no
explicit CHECKPOINT), so none of the CSR checkpoint-rollback code executes on this path.
The flushed state is produced entirely by unmodified code (RelBatchInsert flush paths).

Contrasting shapes that work:

  • 3-rel PK-matched COPY + immediate scan (CopyTest.RelCopyWithoutDefaultHashIndex) passes.
  • The same 1000 rels loaded via single-row CREATEs + CHECKPOINT + scan passes.
  • Reads after a successful checkpoint of the COPY-loaded data pass (the checkpoint
    rewrite appears to normalize whatever is off).

So the gap is specifically: multi-page flushed (not yet checkpointed) CSR state vs the
scan/checkpoint-scan readers. Hypothesis (unconfirmed): the flushed chunks' page-range
metadata disagrees with what is actually on disk (single-page flushes stay within file
slack/preallocation, which is why tiny cases pass). Candidate code:
Column::flushData / ColumnChunkData::flushBuffer page math vs
initScanForCommittedPersistent / scanCSRHeader page reads.

Expected

Reads (and checkpoint scans) of flushed-but-uncheckpointed CSR groups work, or the flush
path produces state identical in observable behavior to checkpointed state.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions