fix(jobs): isolate reconcilers on a dedicated queue so bulk fan-outs cannot starve them - #2
Merged
Merged
Conversation
…cannot starve them
Friction-tested 2026-05-11: deploy_status_reconcile was scheduled every
30s on the default queue (MaxWorkers=5). A weekly_digest fan-out (one
row per team, ~tens of thousands of rows) pinned all 5 worker slots
indefinitely. Result: 232,130 weekly_digest rows in 'available' state,
96 deploy_status_reconcile rows starved behind them — and customer
deployments stuck at status='building' even after pods were Ready.
This commit moves the two fast reconcilers (deploy-status every 30s,
custom-domain every 5min) to a dedicated 'reconcile' queue with
MaxWorkers=2. Bulk jobs (email digests, storage scans, expiry sweeps)
keep their 5 default-queue slots. The two pools are independent — a
backlog on one cannot drain the other.
Also fixes the misleading startup log that hard-coded "queues":"[default]"
even when more queues are configured.
Tests:
- TestQueueReconcileConst guards the queue name so a typo in either
the InsertOpts closures or the QueueConfig map doesn't silently route
reconcilers back to the default queue, reintroducing the bug.
Live verification before this PR (v2.1.0-reconcile-queue deployed):
1. Triggered POST /deploy/323a85a5/redeploy
2. Within 30s: log entry
"jobs.deploy_status_reconcile.transition","from":"building","to":"healthy"
3. GET /deploy/323a85a5 returned status="healthy" in under a minute
— previously it stayed "building" indefinitely.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805
added a commit
that referenced
this pull request
May 17, 2026
P2 worker-side fixes from BUGHUNT-REPORT-2026-05-17.md:
1. GeoLite2 refresh corrupted the .mmdb after every run — MaxMind serves a
gzipped tarball (suffix=tar.gz) but the job io.Copy'd the raw bytes
straight to the .mmdb path. Now gunzip + untar and extract only the
*.mmdb member (geodb.go: extractGeoLite2MMDB).
2. razorpay_webhook_events dedup table was never pruned (migration 033
envisioned it; no job shipped). Added RazorpayWebhookPruneWorker — a
daily DELETE of rows > 30 days, registered alongside uptime_retention.
3. Billing reconciler free-downgrade stranded paid resources: a terminal
Razorpay status with PaidCount==0 set plan_tier='free' (the 24h-TTL
ephemeral tier) while resources kept a paid tier + expires_at=NULL.
Now every terminal status downgrades to 'hobby' (lowest paid tier),
matching the non-zero-paid branch — keeps team-tier and resource-tier
coherent.
4. ExpireStacksWorker hard-deleted the stacks row even when namespace
teardown was skipped (k8sClient==nil) — orphaning a live namespace
with no DB pointer. Now skips the DELETE too, leaving the row for a
later in-cluster run.
5. Storage suspend/unsuspend flapped at the boundary (suspend and
unsuspend both at limitBytes). Added a hysteresis band: unsuspend only
below 90% of the limit; the unsuspend loop now also skips resources the
suspend loop flipped in the same Work() tick.
6. Expiry worker ignored paused/suspended anonymous resources past TTL —
query was status='active' only, orphaning them. Now expires
status IN ('active','paused','suspended') past TTL (TTL wins over
lifecycle state).
7. Confirmed worker storage limits are MiB throughout (*1024*1024); no
decimal-MB sites — no change needed.
Tests: extraction unit tests for #1, hysteresis dead-band test for #5,
paused/suspended expiry test for #6, prune-job tests for #2. go build /
go vet / go test ./... -short all green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
`deploy_status_reconcile` was scheduled every 30s on the default queue (MaxWorkers=5). A `weekly_digest` fan-out (one row per team) pinned all 5 worker slots indefinitely — 232K `available` rows piled up, and the reconciler never ran. Customer deployments stuck at `status="building"` even after pods were Ready.
Fix: move fast reconcilers (deploy-status every 30s, custom-domain every 5min) to a dedicated `reconcile` queue with MaxWorkers=2. Bulk jobs keep their 5 default-queue slots. Backlog on one cannot drain the other.
Also fixes a misleading startup log that hard-coded `"queues":"[default]"` even when more queues are present.
Live verification (before this PR)
Worker v2.1.0-reconcile-queue deployed to prod:
Test plan
🤖 Generated with Claude Code