Repository navigation
perf(importers): look up only the endpoints a report can match - #16225
Conversation
EndpointManager.get_or_create_endpoints read every endpoint of the product into Python on each import and reimport flush to match the handful the report names. Only opening the cursor counts as SQL time, so on a product with millions of endpoints a flush spent minutes iterating rows while the SQL time stayed small. A large product's reimport took hundreds of seconds even when nothing had to be created. The lookup now selects only the product's endpoints whose lower(host) matches a queued key's host (plus null/empty hosts when a queued key has none), which the existing (product, lower(host)) index serves, and keeps only rows whose full key was queued. Matching is unchanged: case-insensitive protocol and host, the scheme's default port equal to no port, first by id wins, and other products' endpoints never match. Measured on a product with 5M endpoints: a 1000-finding reimport with nothing changed went from 574 s to 22 s wall, an import of 1000 findings from 754 s to 112 s, and a reimport with 20% changed from 431 s to 50 s. Query counts are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Adversarial review (independent session, per review-pr) Recommendation: request-changes. There is one blocker. On real data the scoped lookup matches the old full-product scan, and the speedup holds up. The blocker is that the SQL prefilter lowercases with Postgres while the key check lowercases with Python. For some hosts the two disagree, and the importer then creates a new duplicate endpoint on every flush. Blocker1. Python Checked on Postgres 16,
These hosts reach the table with their case intact. Reproduced with the old loop and the new code side by side on the same rows, in a rolled-back transaction. One existing endpoint
The duplicates also reach Pro's endpoint-manager subclass, which reprioritizes the product whenever Suggested fix: make the SQL candidate set a superset of what the Python key check accepts. The smallest change that stops the growth is to also match Postgres's own lowercase of the raw queued hosts, so a host spelled the same way always finds its row: raw_hosts = {kwargs["host"] for kwargs in self._endpoints_to_create.values() if kwargs["host"]}
host_filter = Q(lower_host__in=lowered) | Q(lower_host__in=[Lower(Value(h)) for h in raw_hosts])I tried that filter, together with the null-host change from item 2, as a patch to the method on the same rows (three flushes of Nice to have2. A queued key with no host makes the lookup scan every endpoint in the database.
So 3. The branch no longer contains the current Adversarial attempts
Suites
Verified hands-on vs static
🤖 Generated with Claude Code |
An adversarial review found that the scoped endpoint lookup filtered on LOWER(host) with Python-lowered key hosts. The two disagree for some characters (a dotted capital I, a final sigma, any non-ASCII letter under a C collation), so the existing endpoint was missed and every flush created another duplicate. The filter now also accepts the database's LOWER() of the raw queued hosts, computed inside the same query, so no round trip is added and the importer query counts are unchanged. The null/empty-host branch is written against lower_host so the (product, lower(host)) index serves it instead of a scan of the whole endpoint table. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Both points are fixed in e89121f. Blocker (Python vs database lowercase). The filter now accepts both the Python-lowered key hosts and New test
Should-fix (null/empty-host branch). This branch is now Re-run with 🤖 Generated with Claude Code |
Description
EndpointManager.get_or_create_endpointsruns on every import and reimport flush. To match the endpoints the report names, it read every endpoint of the product into Python. Only opening the cursor counts as SQL time, so on a product with millions of endpoints a flush spent minutes iterating rows while SQL time stayed small. Production diagnostics show this shape: reimports with wall time far above SQL time, at worst about 33 s of SQL in a request that ran for about 25 minutes.The lookup now selects only the product's endpoints whose
lower(host)matches the host of a queued key. A queued key with no host also matches null or empty hosts. The existing(product, lower(host))index serves the lookup. Only rows whose full key was queued are kept. Matching does not change:Test results
New
TestGetOrCreateEndpointsScopedLookupinunittests/test_endpoint_manager.py:dojo_endpointdo not grow when unrelated endpoints are added to the product. Ondevit fails with1 != 61(60 unrelated endpoints added);On
devthe module gives 13 tests and 2 failures; with this change all 13 pass, run with the CI tag filters (--exclude-tag=non-parallel --exclude-tag=performance).Wider runs pass in both modes:
test_endpoint_manager,test_importers_performance,test_import_reimport,test_tag_inheritance_perfandtest_endpoint_meta_import: 212 tests withDD_V3_FEATURE_LOCATIONS=False, 180 with it on.test_importers_performancequery counts are unchanged.Scale check through
POST /api/v2/import-scan/and/reimport-scan/(Generic Findings JSON, one endpoint per finding) into a product holding 5M endpoints:devwall timeQuery counts are the same before and after; only the rows read change.
Ruff (
ruff==0.16.9) check passes on the changed files.Documentation
No user-facing change.
Checklist
dev.devbranch.🤖 Generated with Claude Code