Dedupe: row-lock the duplicate_finding self-FK so import dedup and duplicate deletion stop failing at COMMIT - #16064
Merged
Merged
Conversation
… bulk delete The duplicate_finding self-FK is ON DELETE DO_NOTHING and DEFERRABLE INITIALLY DEFERRED, and both code paths that guard it read without a lock and wrote afterwards: - _drop_links_to_deleted_originals checked the matched originals with a plain SELECT. A delete still uncommitted at that moment was invisible to it, so the flush wrote its links and then failed at COMMIT with "Key (duplicate_finding_id)=(N) is not present", rolling back the whole batch's deduplication. - resolve_inbound_duplicate_references read the findings pointing into a chunk, then the chunk was deleted. A dedup flush committing a new link into the chunk in between failed the chunk at COMMIT with "Key (id)=(N) is still referenced" (the excess-duplicate delete task). The flush now takes FOR KEY SHARE (the lock Postgres itself takes for an FK check) on the originals and the rows it writes, in one ascending-id statement inside its transaction. Each bulk-delete chunk takes FOR UPDATE on its findings, also in one ascending-id statement, right before resolving inbound references; the resolver moved after the child-row cascade so finding rows remain the last rows a chunk locks. A flush that locks first commits first and the resolver sees its links; a flush that comes later waits, finds the original gone, and drops the link. Tests reproduce both interleavings with two real connections and real commits, plus a lock-order case that deadlocks if the flush locked only the originals. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
blakeaowens
approved these changes
Sep 23, 2026
devGregA
approved these changes
Sep 23, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
[sc-15672]
Description
Two background jobs fail at COMMIT on the
dojo_finding.duplicate_finding_idself-FK, one from each side of the same race:Key (duplicate_finding_id)=(N) is not present in table "dojo_finding".async_dupe_delete):Key (id)=(N) is still referenced from table "dojo_finding".The self-FK is
ON DELETE DO_NOTHINGandDEFERRABLE INITIALLY DEFERRED, so nothing is checked until COMMIT. Both sides already guard it with a read-then-write step, but neither step held a lock:_drop_links_to_deleted_originals(dojo/finding/deduplication.py) checked that each matched original still exists with a plain SELECT, then bulk-updated the links. A delete that was still uncommitted when the SELECT ran was invisible to it, so the flush wrote its links and its COMMIT failed once the delete committed.resolve_inbound_duplicate_references(dojo/finding/helper.py) read which findings still point into a chunk, re-pointed them, then deleted the chunk. A dedup flush that committed a new link into the chunk after that read made the chunk's COMMIT fail.When either COMMIT fails the whole transaction rolls back: the entire batch's deduplication on the import side, the whole chunk on the delete side.
Fix
Both sides now lock the rows the FK joins, in one statement each, ordered by id.
SELECT id ... WHERE id = ANY(originals + rows being written) ORDER BY id FOR KEY SHARE, inside the flush's existing transaction. FOR KEY SHARE is the lock Postgres itself takes on a referenced row when it checks a foreign key. It conflicts only with DELETE and key updates, so ordinary updates to an original and other flushes linking to it are not blocked. If a concurrent transaction is deleting an original, the check waits for it. Once the delete commits, the row is not returned and its links are dropped the same way the guard already dropped links to originals deleted earlier.lock_findings_for_delete(chunk_ids)takesSELECT ... FOR UPDATE ORDER BY idon the chunk right beforeresolve_inbound_duplicate_references. Any flush that wants to link into the chunk then waits until the chunk commits, and finds the original gone. A flush that locked first commits first, and the resolver's read sees its links.duplicate_findingvalue.Options considered and not taken:
on_delete=SET_NULL: Django emulates it in the Collector, and the raw-SQL chunked delete bypasses the Collector. A database-levelON DELETE SET NULLwould not fix the import side either, because the flush's uncommitted link is not visible to the delete. It would also leaveduplicate=Truerows with no original, and it needs a migration.Deadlock analysis:
set_duplicate). Locking only the originals produced a realdeadlock detectedin the new tests; the combined lock passes.FOR UPDATEon their own chunk, and chunks commit separately, so there is no cycle, including withorder_desc=True.Finding.delete/finding_delete) is unchanged. The Collector's DELETE now waits for a flush holding FOR KEY SHARE on the finding. If that flush commits a link to it, the delete fails at COMMIT with the same FK error it can hit today, whichdelete_finding_with_conflict_retryalready retries.post_process_findings_batchand_bulk_delete_findings_internal.No migration and no schema change.
Test results
New
unittests/test_dedupe_delete_commit_race.pyreproduces both interleavings with two real connections and real commits. ATestCasenever checks a deferred constraint andTransactionTestCasecannot run in this suite, so the class commits its own rows and removes them afterwards, including the lazily created notification settings row that would otherwise shift later query counts. The interleaving is deterministic: the side that has to go second waits until the other has finished its step or is blocked on a lock (pg_blocking_pids).Key (duplicate_finding_id)=(11) is not present in table "dojo_finding". After, the link is dropped, the rest of the batch is written, and no links dangle.Key (id)=(6) is still referenced from table "dojo_finding". After, the chunk deletes and the late flush drops its link.Local results:
test_dedupe_flush_missing_original,test_bulk_delete_*,test_deduplication_logic,test_duplication_loops,test_finding_helper,test_prepare_duplicates_for_delete,test_importers_deduplication, reimport flush/drain and others): 275 tests OK.test_async_deleteandtest_cascade_delete: 31 OK.docker/entrypoint-unit-tests.shphases: parallel 7836 OK (504 skipped), non-parallel 45 OK, transactional 7 OK (all skipped), performance 11 OK with query counts unchanged.test_importers_performancepasses both before and after the new module runs. The only rows left behind are pghistory audit events, andtest_auditlogandtest_flush_auditlogpass with them present.bugfix(which includes Stream uid/hash dedupe candidates to bound reimport memory #16042, touching the same dedup module): the new module plustest_dedupe_candidate_streaming,test_dedupe_flush_missing_original,test_deduplication_logic,test_bulk_delete_duplicate_references,test_duplication_loops,test_prepare_duplicates_for_deleteandtest_importers_deduplicationran 218 tests, OK (2 skipped). Ruff 0.16.5 passes on the changed files.Documentation
No documentation change. Deduplication and duplicate-deletion behavior is unchanged apart from these jobs no longer failing.
Checklist
dev. (Bug fix: rebased on the latestbugfix.)dev.bugfixbranch.🤖 Generated with Claude Code