Skip to content

Bug: SIGSEGV opening a preserved interrupted-checkpoint journal on Linux ARM64 (0.19.0 and 0.21.2) #1098

Description

@noamsiegel

Ladybug version

0.19.0 and 0.21.2, official PyPI Linux ARM64 wheels. The preserved database/journal originated under 0.19.0; the same fixture was tested with both readers.

What operating system are you using?

Linux aarch64, kernel 7.0.0-34-generic, Python 3.12.14. The original application is Cognee 1.6.1 in a Linux ARM64 container.

What happened?

An existing graph has a preserved interrupted-checkpoint journal. Opening a copy of the completed main database succeeds and MATCH (n) RETURN count(n) returns 45,846 nodes. Adding copies of its .wal.checkpoint and .shadow files causes a native segmentation fault during ladybug.Database(...), before the constructor returns. The subprocess exits with signal 11 (returncode == -11).

This occurs on both 0.19.0 and 0.21.2 using buffer_pool_size=134217728 and max_num_threads=2, in separate subprocesses and on separate copies. No Cognee imports, live writers, or MCP connections are involved in these isolated opens.

Fixture sizes:

File Bytes
graph.pkl 1,322,979,328
graph.pkl.wal.checkpoint 1,743,877
graph.pkl.shadow 0

The application had previously experienced an OOM kill and graph-worker replay timeouts. We have not established that the OOM caused this interrupted checkpoint, or identified the original write sequence.

Expected behavior: recover the valid journal records, or return an actionable recovery/validation exception rather than terminating the process with SIGSEGV. The completed main graph is at least openable and count-queryable without these sidecars; this is not a claim that every stored property has been validated.

Known steps to reproduce

The attached-in-body Python script takes a fixture directory containing the three files above. It runs a main-file-only control and a checkpoint-recovery case, using fresh copies and isolated subprocesses. It never modifies the supplied fixture. Run it once with each installed Ladybug version:

python ladybug-checkpoint-repro.py /path/to/fixture

Observed output for both versions:

base-only exit 0
Ladybug <0.19.0 or 0.21.2>
OPEN_OK
NODE_COUNT [45846]

with-checkpoint exit -11
Ladybug <0.19.0 or 0.21.2>
Fatal Python error: Segmentation fault
Current thread ... (most recent call first):
  <no Python frame>

The fixture contains personal memory, so no database/journal files or core dumps are attached. This is a fixture-dependent reproduction, not yet a self-contained synthetic write/crash reproducer. In particular, we have not reproduced a new writer failure with 0.21.2.

Related issues and diagnostic question

#714 reports journals that become unreplayable after hard termination. #1022 fixes loss of recovered frozen-WAL records. Both are closed. Please advise whether this preserved-fixture recovery SIGSEGV belongs to those issues or represents a separate validation/recovery defect, and which sanitized diagnostics would help identify the failing WAL record.

Reproduction script

#!/usr/bin/env python3
"""Test copied graph/checkpoint fixtures without modifying the supplied files."""
import argparse
from pathlib import Path
import resource
import shutil
import subprocess
import sys
import tempfile

parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("fixture", type=Path, help="directory containing graph.pkl and its checkpoint files")
args = parser.parse_args()
source = args.fixture / "graph.pkl"
if not source.is_file() or not Path(str(source) + ".wal.checkpoint").is_file():
    parser.error("fixture must contain graph.pkl and graph.pkl.wal.checkpoint")

resource.setrlimit(resource.RLIMIT_CORE, (0, 0))
child = """
import ladybug, sys
print("Ladybug", ladybug.__version__, flush=True)
db = ladybug.Database(sys.argv[1], buffer_pool_size=134217728, max_num_threads=2)
print("OPEN_OK", flush=True)
conn = ladybug.Connection(db)
result = conn.execute("MATCH (n) RETURN count(n)")
print("NODE_COUNT", result.get_next(), flush=True)
conn.close()
db.close()
"""
failed = False
for case in ("base-only", "with-checkpoint"):
    # Retain private scratch copies for inspection. Never mutate the source fixture.
    directory = Path(tempfile.mkdtemp(prefix="ladybug-" + case + "-"))
    copied = directory / "graph.pkl"
    shutil.copy2(source, copied)
    if case == "with-checkpoint":
        for suffix in (".wal.checkpoint", ".shadow"):
            sidecar = Path(str(source) + suffix)
            if sidecar.is_file():
                shutil.copy2(sidecar, Path(str(copied) + suffix))
    try:
        result = subprocess.run(
            [sys.executable, "-X", "faulthandler", "-c", child, str(copied)],
            capture_output=True, text=True, timeout=45,
        )
        print(case, "exit", result.returncode, flush=True)
        print(result.stdout.strip(), flush=True)
        print(result.stderr.strip(), flush=True)
        failed |= result.returncode != 0
    except subprocess.TimeoutExpired:
        print(case, "TIMEOUT after 45s", flush=True)
        failed = True
sys.exit(1 if failed else 0)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions