Repository navigation
test(finding): accept either deadlock victim in the bulk-delete vs tag writer race - #16100
Merged
Merged
Conversation
…g writer race test_tag_writer_holding_a_tag_row_does_not_fail_the_chunk assumed Postgres always aborts the bulk-delete chunk when it deadlocks with bulk_add_tags_to_instances. Postgres does not promise that. A waiting backend runs the deadlock check once, deadlock_timeout after it started waiting. The chunk starts waiting first, so it is the victim only when the writer reaches its COMMIT within deadlock_timeout. On a slow or loaded machine the chunk's one check runs before the cycle exists, and the writer is aborted instead. That failed the test about 1 run in 17 locally (3 of 50). Both orders satisfy what #16091 guarantees: the chunk deletes, no tag through row survives, and the tag count stays consistent. The test now asserts exactly that, plus that exactly one side was the deadlock victim (the chunk's attempts are counted through lock_findings_for_delete). A new test forces the writer-victim order deterministically by holding the writer's COMMIT past the chunk's deadlock check, so both orders stay covered. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
blakeaowens
approved these changes
Sep 26, 2026
devGregA
approved these changes
Sep 26, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
[sc-15987]
Description
TestBulkDeleteRacesConcurrentM2MWriter.test_tag_writer_holding_a_tag_row_does_not_fail_the_chunk(added in #16091) is flaky. Its failure isOperationalError: deadlock detected, raised in the concurrent tag writer thread.The test forces a real deadlock.
bulk_add_tags_to_instancesholds a tag's count row, and the bulk-delete chunk holds the finding rows and waits on that tag row. The writer then commits, and its deferred foreign-key check waits on the chunk's finding lock. The test assumed Postgres always aborts the chunk, whose conflict retry then finishes the delete.Postgres does not guarantee which side it aborts. A waiting backend runs the deadlock check once,
deadlock_timeoutafter it started waiting, and aborts itself if it finds a cycle. The chunk starts waiting first, so it is the victim only when the writer reaches its COMMIT withindeadlock_timeout. On a slow or loaded machine the writer gets there later: the chunk's one check has already run and found no cycle, so the writer's check finds the cycle and the writer is aborted.Both orders meet the guarantee #16091 describes: the chunk deletes, no tag through row survives, and the tag count stays consistent. The product code needs no change. When the writer is the victim its tag add rolls back and the chunk commits on its first attempt, which is also a consistent result.
Changes, all in
unittests/test_dedupe_delete_commit_race.py:lock_findings_for_deleteand checks the writer for a transient conflict error. Any other writer error still fails the test.test_tag_writer_aborted_as_the_deadlock_victim_leaves_the_chunk_to_finish, forces the writer-victim order. The writer holds its COMMIT for twice the server'sdeadlock_timeout, which is read at run time. The test then asserts the writer was aborted and the chunk did not retry. So the order that used to be the flake is covered on purpose.The follow-up that #16091 mentions still applies:
bulk_add_tags_to_instancescould take FOR KEY SHARE on its findings before the tag UPDATE and remove this cycle entirely.Test results
These ran in the OSS unit-test compose (
ossfix-unittestsimage) with--keepdb:OperationalError: deadlock detectedin the writer.unittests.test_dedupe_delete_commit_race: 9 tests OK, withDD_V3_FEATURE_LOCATIONSboth off and on.ruff check --config ruff.tomlwith ruff 0.16.5 (the pinned version) passes.Documentation
Test only, no user-facing change.
🤖 Generated with Claude Code