Skip to content

fix(jobs): isolate reconcilers on a dedicated queue so bulk fan-outs cannot starve them - #2

Merged
mastermanas805 merged 1 commit into
masterfrom
fix/reconcile-queue
May 11, 2026
Merged

mastermanas805 merged 1 commit into
masterfrom
fix/reconcile-queue

Conversation

@mastermanas805

Copy link
Copy Markdown
Member

Summary

`deploy_status_reconcile` was scheduled every 30s on the default queue (MaxWorkers=5). A `weekly_digest` fan-out (one row per team) pinned all 5 worker slots indefinitely — 232K `available` rows piled up, and the reconciler never ran. Customer deployments stuck at `status="building"` even after pods were Ready.

Fix: move fast reconcilers (deploy-status every 30s, custom-domain every 5min) to a dedicated `reconcile` queue with MaxWorkers=2. Bulk jobs keep their 5 default-queue slots. Backlog on one cannot drain the other.

Also fixes a misleading startup log that hard-coded `"queues":"[default]"` even when more queues are present.

Live verification (before this PR)

Worker v2.1.0-reconcile-queue deployed to prod:

  1. Triggered `POST /deploy/323a85a5/redeploy`.
  2. Within 30s: `{"msg":"jobs.deploy_status_reconcile.transition","from":"building","to":"healthy"}`.
  3. `GET /deploy/323a85a5` returned `status: "healthy"` — previously stuck `building` for hours.

Test plan

  • `TestQueueReconcileConst` — guards the queue name string so a typo can't silently route reconcilers back to the starving queue
  • `go build ./...` passes
  • Live: status flip verified end-to-end on v2.1.0

🤖 Generated with Claude Code

…cannot starve them

Friction-tested 2026-05-11: deploy_status_reconcile was scheduled every
30s on the default queue (MaxWorkers=5). A weekly_digest fan-out (one
row per team, ~tens of thousands of rows) pinned all 5 worker slots
indefinitely. Result: 232,130 weekly_digest rows in 'available' state,
96 deploy_status_reconcile rows starved behind them — and customer
deployments stuck at status='building' even after pods were Ready.

This commit moves the two fast reconcilers (deploy-status every 30s,
custom-domain every 5min) to a dedicated 'reconcile' queue with
MaxWorkers=2. Bulk jobs (email digests, storage scans, expiry sweeps)
keep their 5 default-queue slots. The two pools are independent — a
backlog on one cannot drain the other.

Also fixes the misleading startup log that hard-coded "queues":"[default]"
even when more queues are configured.

Tests:
- TestQueueReconcileConst guards the queue name so a typo in either
  the InsertOpts closures or the QueueConfig map doesn't silently route
  reconcilers back to the default queue, reintroducing the bug.

Live verification before this PR (v2.1.0-reconcile-queue deployed):
  1. Triggered POST /deploy/323a85a5/redeploy
  2. Within 30s: log entry
     "jobs.deploy_status_reconcile.transition","from":"building","to":"healthy"
  3. GET /deploy/323a85a5 returned status="healthy" in under a minute
     — previously it stayed "building" indefinitely.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@mastermanas805
mastermanas805 merged commit 729e3c8 into master May 11, 2026
@mastermanas805
mastermanas805 deleted the fix/reconcile-queue branch May 11, 2026 08:54
mastermanas805 added a commit that referenced this pull request May 17, 2026
P2 worker-side fixes from BUGHUNT-REPORT-2026-05-17.md:

1. GeoLite2 refresh corrupted the .mmdb after every run — MaxMind serves a
   gzipped tarball (suffix=tar.gz) but the job io.Copy'd the raw bytes
   straight to the .mmdb path. Now gunzip + untar and extract only the
   *.mmdb member (geodb.go: extractGeoLite2MMDB).

2. razorpay_webhook_events dedup table was never pruned (migration 033
   envisioned it; no job shipped). Added RazorpayWebhookPruneWorker — a
   daily DELETE of rows > 30 days, registered alongside uptime_retention.

3. Billing reconciler free-downgrade stranded paid resources: a terminal
   Razorpay status with PaidCount==0 set plan_tier='free' (the 24h-TTL
   ephemeral tier) while resources kept a paid tier + expires_at=NULL.
   Now every terminal status downgrades to 'hobby' (lowest paid tier),
   matching the non-zero-paid branch — keeps team-tier and resource-tier
   coherent.

4. ExpireStacksWorker hard-deleted the stacks row even when namespace
   teardown was skipped (k8sClient==nil) — orphaning a live namespace
   with no DB pointer. Now skips the DELETE too, leaving the row for a
   later in-cluster run.

5. Storage suspend/unsuspend flapped at the boundary (suspend and
   unsuspend both at limitBytes). Added a hysteresis band: unsuspend only
   below 90% of the limit; the unsuspend loop now also skips resources the
   suspend loop flipped in the same Work() tick.

6. Expiry worker ignored paused/suspended anonymous resources past TTL —
   query was status='active' only, orphaning them. Now expires
   status IN ('active','paused','suspended') past TTL (TTL wins over
   lifecycle state).

7. Confirmed worker storage limits are MiB throughout (*1024*1024); no
   decimal-MB sites — no change needed.

Tests: extraction unit tests for #1, hysteresis dead-band test for #5,
paused/suspended expiry test for #6, prune-job tests for #2. go build /
go vet / go test ./... -short all green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant