Skip to content

feat(db): advance an empty partition catch-all under bounded locks - #7993

Open
Areson wants to merge 5 commits into
block:mainfrom
Areson:Areson/partition-catchall-advancement
Open

Areson wants to merge 5 commits into
block:mainfrom
Areson:Areson/partition-catchall-advancement

Conversation

@Areson

@Areson Areson commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds bounded, serialized advancement of an empty right-edge partition catch-all for events and delivery_log. The repaired catalogs currently end in an empty catch-all from 2027-01-01, and the merged manager (#6515) only audits them. Without this change, January writes land in the catch-all, after which only an operator can move them.

When BUZZ_PARTITION_MANAGER_ADVANCE_ENABLED=true, each relay keeps six dedicated monthlies after the current month. It does this at startup and on every audit interval by replacing the empty catch-all with canonical monthlies and a new canonical <table>_p_future after the horizon. From today's catalogs the first run creates January–March 2027 and moves the catch-all to April 2027.

Safety design

  • Refuse before locking. The following return operator_required without taking any table lock:

    • a populated catch-all, or one whose occupancy is unknown;
    • a catch-all with child-only triggers, which the replacements would not inherit;
    • a misaligned bound;
    • a DEFAULT, anomalous or pending-detach child;
    • a name collision.
  • One winner. A schema-scoped pg_try_advisory_xact_lock is taken before any DDL. Losers report skipped_locked and never issue DDL. The winner re-plans from a fresh audit before taking any table lock and checks for collisions. A runner that lost a race therefore sees the committed layout as a no-op, not as colliding names.

  • Writer lock order. First ONLY the parent is locked ACCESS EXCLUSIVE under a 2 s lock_timeout. Then each foreign-key counterpart (today only communities) is locked in OID order SHARE ROW EXCLUSIVE NOWAIT, then the catch-all NOWAIT.

    • The counterparts are read again once the parent is locked, and that set is the one locked. Any FK DDL on or to the parent conflicts with the parent lock, so the set can't change after that. An FK committed while maintenance waited is therefore locked NOWAIT too, so attaching never waits on it under the parent lock. The read before the lock remains as a cheap early refusal.
    • This is the order an insert takes locks (parent, then the FK check on communities), so maintenance cannot deadlock with ingest. The internal manual runbook used to lock communities first; it now uses this order too.
    • SHARE ROW EXCLUSIVE is the strongest lock the DDL takes on communities; measured on PG 16 and 17, the catch-all DROP takes none, as long as the catch-all's only foreign keys are those inherited from the parent. A catch-all with its own foreign key, or one that another table's foreign key references, is refused: once before any lock, and again after the catch-all lock, where the check is authoritative. Readers and other writers' FOR KEY SHARE checks proceed. A concurrent communities writer makes the run fail fast with lock_timeout.
    • A 40P01 deadlock victim is also reported as lock_timeout.
    • Parents that are referenced by a foreign key, or that reference an inherited table, are refused, including when such a key commits while maintenance waits on the parent.
  • Re-plan under lock. The catalog is re-audited under those locks. The transaction proceeds only if the plan is unchanged.

  • Bounded. statement_timeout 5 s, idle_in_transaction_session_timeout 5 s, and a 10 s client deadline per table. A deadline-abandoned connection is closed, never returned to the pool, and the server rolls back anything uncommitted. If the deadline fires after COMMIT was sent, the change may have landed; the next audit reports the truth. The orphaned backend keeps its locks until its in-flight statement ends, at most statement_timeout.

  • Verify before commit. The changed catalog is re-audited inside the transaction for bounds, canonical kinds, trigger parity and catch-all emptiness, and serving safety must not regress. It also checks that every new child has valid, ready clones of each parent index and validated clones of each PK/unique/FK constraint.

  • Audits tolerate a concurrent advance. Catalog rows come from the statement snapshot, but pg_get_expr renders from current catalog state. So an audit that overlaps another relay's DROP of a catch-all can see a partition that no longer exists. The read-only audit detects this (a NULL bound, or 42P01 from the emptiness probe) and retries up to twice with a fresh snapshot.

  • Fail fast on audit probes. If another relay's audit is probing the catch-all's emptiness, the catch-all NOWAIT lock fails and the run reports lock_timeout. The parent lock is then released immediately, and the next runner or interval completes the advance.

Uncovered-month creation (BUZZ_PARTITION_MANAGER_CREATE_ENABLED, default true) now runs through the same locked transaction, replacing the unlocked per-month CREATE. That flag remains the kill switch for all automatic partition DDL: with it off, advancement is disabled too. With advancement disabled (the default) the periodic loop stays read-only and startup behaves as before, apart from the locking.

A runner that planned DDL always re-audits before returning, even when a peer did the work, so it never reports the pre-maintenance layout. A serialized re-plan that finds only disabled work reports skipped_disabled.

Behavior changes

  • The relay audit horizon and the buzz-admin partition-audit default are now 6 months (PARTITION_MANAGER_MONTHS_AHEAD), up from 3. Until advancement is enabled, months already covered by the catch-all make degraded true. That correctly reports a runway under six months.
  • New metric buzz_partition_maintenance_runs_total{table,outcome} with outcomes noop | skipped_disabled | created | advanced | skipped_locked | operator_required | lock_timeout | deadline | error. buzz_partition_create_attempts_total keeps its per-month created | skipped_covered | error labels. error there still means a DDL or collision failure; lock_timeout and deadline runs add nothing to it, and only buzz_partition_maintenance_runs_total records them.
  • ensure_future_partitions (free function and Db method) is removed; maintain_partitions is the one entry point.
  • New advisory-lock label lock_type="partition_maintenance".

Rollout

Enable BUZZ_PARTITION_MANAGER_ADVANCE_ENABLED per environment in BPCI: bb-block-staging first, then production. The staging canary proves a real advance because the six-month horizon already extends past January. Fallback: rerun the internal manual runbook, which uses the same parent-first lock order as of 2026-10-05. It keeps ACCESS EXCLUSIVE on the counterparts, because its detach/attach can drop FK triggers on them.

Testing

  • cargo test -p buzz-db --lib partition: 17 unit tests. The 10 new ones cover:
    • plans: repaired layout, policy gating, no-op, refusal states including child-only triggers, December→January;
    • the create kill switch also stopping advancement;
    • the month cap at exactly 120;
    • per-month metric outcomes for transient failures versus DDL failures;
    • the mid-audit-drop retry classifier.
  • cargo test -p buzz-db --lib partition -- --ignored (PostgreSQL 17): 48/48 on 9 consecutive runs, including 17 new tests:
    • advance to the canonical layout, with routing and idempotent rerun;
    • replacing a canonically named catch-all in one transaction;
    • populated catch-all refused in under 1 s while the parent is locked;
    • a communities writer making the run fail fast and roll back;
    • a communities reader not blocking the advance;
    • an in-flight insert's FK check, and its write to communities, behind queued maintenance completing without a deadlock;
    • a foreign key committed while maintenance waits on the parent, with a writer on its new counterpart, failing fast as lock_timeout on that table instead of waiting under the parent lock;
    • a foreign key referencing events, committed while maintenance waits on the parent, being refused under the parent lock;
    • a catch-all's own foreign key, or a foreign key referencing it, being refused before the parent lock with the key intact, and a catch-all key committed while maintenance waits on the parent being refused after the catch-all lock;
    • a 40P01 deadlock victim being classified as lock_timeout;
    • parent contention timing out without blocking the other table;
    • held advisory lock skipping with no DDL;
    • concurrent runners converging, with at most a transient lock_timeout;
    • a stale audit after a concurrent advance being a no-op, and the losing runner returning the committed layout;
    • a re-plan that finds only disabled work reporting skipped_disabled;
    • an audit retrying when a listed partition is dropped mid-audit;
    • a catch-all with a child-only trigger being refused;
    • client deadline abandoning the attempt and closing the session (scoped to the test's own blocked backend);
    • catch-all name collision;
    • incoming foreign key refused.
  • Mutation checks:
    • restoring the old counterpart-first ACCESS EXCLUSIVE order fails the reader test (lock_timeout) and the writer test (a ~1 s deadlock wait);
    • dropping NOWAIT from the counterpart lock, or allowing a populated catch-all, fails the matching tests;
    • removing the parent-locked counterpart re-read fails the committed-FK tests (the parent is held 2.06 s waiting on the new counterpart), and ignoring its refusal fails the referencing-FK test;
    • locking the counterparts SHARE ROW EXCLUSIVE before the parent fails the writer test;
    • removing either catch-all foreign-key check fails its test.
  • Full PostgreSQL lane: 716/718. Both failures are timing-sensitive tests outside partition code, and each passes 3/3 when run alone:
    • hanging_redis_peer_times_out_before_the_health_listener_binds;
    • readiness_check_cancellation_balances_waiter_and_inflight_connection.
  • cargo clippy -p buzz-db -p buzz-relay -p buzz-admin --all-targets -D warnings and cargo fmt --check are clean.

Related

Refs #2396.

This PR conflicts textually with #7994 (partition audit observability follow-ups) in crates/buzz-db/src/store/partition.rs and crates/buzz-relay/src/main.rs. Whichever PR merges second will be rebased onto the other.

🤖 Generated with Claude Code

Ian Oberst and others added 2 commits September 30, 2026 07:06
The repaired events and delivery_log catalogs end in an empty catch-all
from 2027-01-01 that the merged manager only audits. Add opt-in
maintenance (BUZZ_PARTITION_MANAGER_ADVANCE_ENABLED, default false) that
keeps six dedicated monthlies after the current month by replacing the
empty catch-all with canonical monthlies and a later <table>_p_future,
at startup and on every audit interval.

Each table's change runs in one transaction that takes a schema-scoped
advisory lock (losers report skipped_locked), locks foreign-key
counterparts ACCESS EXCLUSIVE NOWAIT in OID order before ONLY the parent
under a 2s lock_timeout and the catch-all NOWAIT, re-plans from a
locked re-audit, and verifies bounds, trigger parity, and index and
constraint inheritance before commit. Statement, idle-in-transaction,
and a 10s client deadline bound every attempt; a deadline-abandoned
connection is closed rather than pooled. A populated catch-all,
misaligned bound, unsafe catalog, or name collision is refused as
operator_required before any lock.

Uncovered-month creation now uses the same locked path. The relay and
buzz-admin audit horizon becomes six months, and maintenance outcomes
are exported as buzz_partition_maintenance_runs_total{table,outcome}.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Ian Oberst <ioberst@block.xyz>
- Refuse to replace a catch-all that carries child-only triggers; the
  replacement partitions inherit only parent triggers.
- Re-plan from a fresh audit once the maintenance advisory lock is held,
  before any table lock, so a runner that lost a race sees the winner's
  layout as a no-op instead of reporting its monthlies as collisions.
  Name collisions on the refusal path are still reported before any lock.
- Retry a read-only table audit (up to twice) when a partition it listed
  was dropped concurrently: pg_get_expr renders from current catalog
  state and yields NULL for a relation missing since the snapshot, and
  the emptiness probe then fails with 42P01.
- Scope the deadline test to the backend blocked by its own holder, and
  let concurrent-runner convergence tolerate a fail-fast lock_timeout
  from the other runner's audit probing the catch-all.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Ian Oberst <ioberst@block.xyz>

@TheSentinel454 TheSentinel454 left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review of this PR at ee1fb7e4ad8d3d75dda12ce4d9da949e38075ea9.

The direction is good. Partition DDL now goes through one advisory-locked transaction that re-plans under its locks and verifies before commit. Refusing a populated catch-all keeps row moves an operator job. The re-plan and verify logic and the close_on_drop deadline path held up under review.

One P1 (inline at maintenance.rs L545): the lock order and mode in apply_change can deadlock with foreground events → communities traffic. This was reproduced against a real relay on PostgreSQL 17.11, where maintenance was the 40P01 victim. From the source, the reverse interleaving can instead abort a foreground ingest write. The details and suggested fix are inline. There are also several P2s inline.

Evidence at this SHA:

  • cargo test -p buzz-db --lib partition -- --ignored: 44/44 on PG17.11.
  • A live relay advanced the repaired layout to Jan–Apr 2027 monthlies with *_p_future from May 2027. A rerun was a no-op.

Separately, Run Codex Security Review failed because the workflow refuses to check out fork code under pull_request_target. No security review output exists for this PR.

Comment thread crates/buzz-db/src/store/partition/maintenance.rs Outdated
Comment thread crates/buzz-db/src/store/partition/maintenance.rs Outdated
Comment thread crates/buzz-db/src/store/partition/maintenance.rs
Comment thread crates/buzz-db/src/store/partition/maintenance.rs Outdated
Comment thread crates/buzz-db/src/store/partition/maintenance.rs
Comment thread crates/buzz-db/src/store/partition/maintenance.rs Outdated
Comment thread crates/buzz-db/src/store/partition/maintenance.rs
Comment thread crates/buzz-db/src/store/partition.rs
Comment thread crates/buzz-db/src/store/partition.rs Outdated
Comment thread crates/buzz-relay/src/main.rs
Address review on catch-all advancement:

- Lock the parent first, then counterparts in SHARE ROW EXCLUSIVE NOWAIT,
  then the catch-all. This matches the order writers take locks, so
  maintenance cannot deadlock with ingest, and readers and FK key checks on
  communities are no longer blocked. Classify 40P01 as lock_timeout.
- Count only DDL and collision failures as per-month create errors.
- Re-audit whenever a table planned DDL, and report skipped_disabled when
  the serialized re-plan finds only disabled work.
- Make the create kill switch stop advancement too.
- Fix the off-by-one in the advancement month cap.
- Retry audits only on a dedicated PartitionDroppedMidAudit error.
- Do not claim a deadline-abandoned attempt rolled back.
- Remove the ensure_future_partitions wrapper.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Ian Oberst <ioberst@block.xyz>

@TheSentinel454 TheSentinel454 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Thanks for the replies and the quick fixes. Two small follow-ups at 2b53073f are inline, both P2.

Rollout note: the PR's fallback, "rerun the manual runbook", still locks communities first. That's the order that deadlocks with ingest. Whoever owns the runbook should switch it to the order used here: parent ACCESS EXCLUSIVE, then counterparts SHARE ROW EXCLUSIVE NOWAIT, then the catch-all NOWAIT.

return Ok(LockedAttempt::Refused(collision_message(&name)));
}

let parent = maintenance_parent(&mut transaction, table).await?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 P2: maintenance_parent reads the FK counterparts before the parent lock, and the locked re-audit at L615 only re-checks the partition layout. If an FK to or from events is committed between this read and the parent lock, the new table isn't in counterparts. CREATE … PARTITION OF then takes SHARE ROW EXCLUSIVE on it with a wait (no NOWAIT) while holding the parent's ACCESS EXCLUSIVE, which is the stall this ordering is meant to rule out. The window is small, since FK DDL on events itself needs a lock that conflicts with the parent lock. Re-running maintenance_parent after the parent lock and returning Refused if the result changed would close it. Keep this early call as the cheap pre-lock refusal.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Confirmed and fixed in 7b6b199. I went a little further than you suggested. Once the parent is locked, maintenance reads the counterparts again and uses that set:

  • if the new read now refuses (for example, a new FK that references events), the run returns Refused;
  • if the set has grown, the new counterparts are locked SHARE ROW EXCLUSIVE NOWAIT like the rest.

Refusing whenever the set changed would report a harmless race as operator_required. The parent-locked read is enough on its own: any FK DDL on or to events needs a lock that conflicts with the parent's ACCESS EXCLUSIVE, so the set can't change after that point. The transaction is READ COMMITTED, so the second read sees what committed while we waited. The early read stays as the cheap pre-lock refusal.

New test: foreign_key_committed_while_waiting_on_the_parent_is_locked_nowait.

  • A migration holds an uncommitted ALTER TABLE events ADD FOREIGN KEY … REFERENCES reviewers.
  • A reviewers writer queues behind the migration.
  • Maintenance queues on the parent, then the migration commits.

Maintenance now fails with events lock_timeout within 1 s, and delivery_log still advances. Without the re-read, it held the parent for 2.06 s waiting on reviewers, and the test fails.

),
)
.await?;
// Attaching a partition adds its foreign key under SHARE ROW EXCLUSIVE on

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 P2 (doc): this explains SHARE ROW EXCLUSIVE only by what attaching needs. It would help to also say why dropping the catch-all can't need more. Suggested addition: "Dropping the catch-all removes only its inherited constraint and check triggers, which lock the catch-all itself, never the referenced table (PG14–17)." Without that line, it's natural to suspect the DROP upgrades the lock on communities. We did, and the probes showed it doesn't.

@Areson Areson Oct 5, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Added in 7b6b199: dropping the catch-all removes only its inherited foreign key and check triggers, which lock the catch-all itself, never the referenced table. I cited the versions we probed ("lock probes on PostgreSQL 16 and 17") rather than 14–17, since 14 and 15 weren't measured.

The counterpart set was read before the parent lock, and the locked
re-audit checks only the partition layout. A foreign key on or to the
parent that committed while maintenance waited for the parent lock was
missing from the set, so attaching a partition waited on its table without
NOWAIT while holding the parent ACCESS EXCLUSIVE, the stall the lock order
is meant to rule out.

Read the counterparts again once the parent is locked and use that set.
Foreign-key DDL on or to the parent conflicts with the parent lock, so the
set is stable from there. The pre-lock read remains as the cheap refusal.
A new test commits a foreign key while maintenance waits on the parent,
with a writer on the new counterpart: maintenance now fails fast with
lock_timeout instead of waiting out the 2 s lock timeout under the parent.

Also document why dropping the catch-all needs no stronger counterpart lock.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Ian Oberst <ioberst@block.xyz>
@Areson

Areson commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

🤖 Re: the rollout note in the latest review (link) that the manual runbook still locks communities first: agreed, and done. The runbook is internal, not in this repo. Both procedures, the events catch-all repair and the delivery_log empty-catch-all advance, now lock the parent first with the 2 s budget, then the counterparts NOWAIT, then the children NOWAIT. They also re-check the FK inventory once the parent is locked.

The runbook keeps ACCESS EXCLUSIVE on the counterparts rather than SHARE ROW EXCLUSIVE. Its detach/attach of tables with pre-created FKs can drop FK action triggers on communities, which needs ACCESS EXCLUSIVE; our probes covered only this PR's drop-and-create.

The per-environment execution records now carry a "do not reuse this SQL" banner. Their rendered SQL is unchanged so the recorded hashes still match. The PR's fallback line now points to the revised order.

@TheSentinel454 TheSentinel454 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two optional nits: in the new test, asserting message.contains("reviewers") would be a stronger check than the 1 s bound. That bound also covers delivery_log's whole advance, and it failed 3 of 20 runs under 4× load. And the post-lock refusal branch at :597-600 has no test.

lock(
&mut transaction,
&format!(
"LOCK TABLE ONLY {} IN ACCESS EXCLUSIVE MODE NOWAIT",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

An FK added with ALTER TABLE ONLY <catch-all> ADD FOREIGN KEY … REFERENCES x is missed in two places. maintenance_parent misses it because it's not on the parent. The child-only-trigger refusal misses it because RI triggers are tgisinternal. The parent lock doesn't block adding it. DROP TABLE at :657 then takes AccessExclusiveLock on x without NOWAIT while holding the parent (2.06 s against a plain x reader on PG17.11). With no contention, it silently drops the FK. Suggest refusing here, after the catch-all lock, if the catch-all has any FK with conparentid = 0 or is referenced by one. Use the existing "lock set is not proven" refusal, and qualify the comment at :605-607 and PR body line 19.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Confirmed and fixed in d2178b0. Maintenance now refuses with the existing "lock set is not proven" reason when the catch-all has a foreign key with conparentid = 0, or is referenced by one. The refusal message names the constraints.

The check runs twice:

  • Before any lock. A standing case is then refused without taking the parent outage on every run.
  • After the catch-all lock, as you suggested. That's the authoritative point, because FK DDL on the catch-all conflicts with that lock but not with the parent's.

The DROP comment and the PR description now say the no-stronger-lock claim covers inherited keys only, and that other keys are refused.

Tests:

  • catch_all_foreign_keys_outside_the_parent_are_refused covers both an ALTER TABLE ONLY <catch-all> ADD FOREIGN KEY and a table that references the catch-all. A reader holds the parent throughout, so the refusal has to come before the parent lock. The key survives, and once it's removed the advance goes through.
  • catch_all_foreign_key_committed_while_waiting_on_the_parent_is_refused commits the catch-all key while maintenance waits on the parent, so only the post-lock check can see it.

Removing either check fails its test.

A foreign key added to the catch-all alone, or one referencing it, is
invisible to the parent's counterpart set and to the child-only trigger
refusal, since foreign-key triggers are internal, and adding it conflicts
with no parent lock. Dropping the catch-all then removed it silently and,
under contention, locked its other table ACCESS EXCLUSIVE without NOWAIT
while holding the parent.

Refuse such a catch-all with the existing "lock set is not proven"
reason: once before any lock, so a standing case never takes the parent
outage, and again after the catch-all lock, where the check is
authoritative because foreign-key DDL on the catch-all conflicts with it.

Tests:
- Both kinds of key are refused without reaching the parent lock, and the
  key survives.
- A catch-all key committed while maintenance waits on the parent is
  refused by the post-lock check.
- A foreign key referencing the parent, committed while maintenance waits,
  hits the post-lock counterpart refusal, which had no test.
- The FK-check writer test also writes to communities, so locking the
  counterparts SHARE ROW EXCLUSIVE before the parent now deadlocks it.
- The committed-counterpart test asserts that NOWAIT names the new table
  instead of a 1 s bound that also covered delivery_log's advance.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Ian Oberst <ioberst@block.xyz>
@Areson

Areson commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

🤖 Re: the two optional nits in the latest review (link). Took both in d2178b0:

  • The new test asserts on the lock that failed. foreign_key_committed_while_waiting_on_the_parent_is_locked_nowait now requires could not obtain lock on relation "….reviewers" instead of the 1 s bound. Without the re-read, the failure is a lock-timeout cancel with no relation name, so the check doesn't depend on load.
  • The post-lock refusal has a test. foreign_key_referencing_the_parent_committed_while_waiting_is_refused commits an FK that references events while maintenance waits on the parent, and expects events is referenced by a foreign key. If the post-lock refusal is ignored, the advance goes through and the test fails.

The partition Postgres suite (52 tests) passed five times in a row. The four concurrent tests and the writer test passed 20 of 20 runs. That was without your 4× load, but the assertion no longer depends on timing.

This branch is waiting to be deployed

1 waiting deployment
codex-review — d2178b00 Waiting Oct 6, 2026 by Areson via Run Codex Security Review #7058
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants