Skip to content

feat(worker): backend-agnostic storage scanner (any S3-compat service) - #5

Merged
mastermanas805 merged 1 commit into
masterfrom
feat/storage-backend-agnostic
May 11, 2026
Merged

mastermanas805 merged 1 commit into
masterfrom
feat/storage-backend-agnostic

Conversation

@mastermanas805

Copy link
Copy Markdown
Member

Pairs with api PR #18. Storage scanner reads OBJECT_STORE_* env vars (with MINIO_* fallback) and auto-detects TLS for managed S3 vendors. Build + tests pass.

…/GCS)

The storage_bytes scanner now works against any S3-compatible service
without code changes — only env-var swap. Pairs with the api repo's
storage-provider refactor on the same branch name.

Changes:

  config.go     adds OBJECT_STORE_* env vars (ENDPOINT / ACCESS_KEY /
                SECRET_KEY / BUCKET / REGION / SECURE), with the legacy
                MINIO_* names kept as fallback for backward compat.

  workers.go    StorageScanner constructor now reads OBJECT_STORE_*
                instead of MINIO_*. The MinIO admin client used by the
                ExpireAnonymousWorker stays gated on legacy MINIO_*
                because per-IAM-user cleanup is only relevant for
                self-hosted MinIO (the shared-key backend doesn't
                create per-customer IAM in the first place).

  storage_minio.go  NewMinIOStorageScanner auto-detects TLS:
                explicit https:// / http:// prefixes win; otherwise
                a heuristic enables TLS for known managed S3 vendor
                hostnames (digitaloceanspaces.com, amazonaws.com,
                cloudflarestorage.com, googleapis.com, wasabisys.com,
                backblazeb2.com). Self-hosted in-cluster MinIO
                continues to use plain HTTP. The minio-go SDK speaks
                plain S3 against all of these.

Verified: go build + go test ./internal/jobs/... pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@mastermanas805
mastermanas805 merged commit 38b831b into master May 11, 2026
@mastermanas805
mastermanas805 deleted the feat/storage-backend-agnostic branch May 11, 2026 16:05
mastermanas805 added a commit that referenced this pull request May 17, 2026
P2 worker-side fixes from BUGHUNT-REPORT-2026-05-17.md:

1. GeoLite2 refresh corrupted the .mmdb after every run — MaxMind serves a
   gzipped tarball (suffix=tar.gz) but the job io.Copy'd the raw bytes
   straight to the .mmdb path. Now gunzip + untar and extract only the
   *.mmdb member (geodb.go: extractGeoLite2MMDB).

2. razorpay_webhook_events dedup table was never pruned (migration 033
   envisioned it; no job shipped). Added RazorpayWebhookPruneWorker — a
   daily DELETE of rows > 30 days, registered alongside uptime_retention.

3. Billing reconciler free-downgrade stranded paid resources: a terminal
   Razorpay status with PaidCount==0 set plan_tier='free' (the 24h-TTL
   ephemeral tier) while resources kept a paid tier + expires_at=NULL.
   Now every terminal status downgrades to 'hobby' (lowest paid tier),
   matching the non-zero-paid branch — keeps team-tier and resource-tier
   coherent.

4. ExpireStacksWorker hard-deleted the stacks row even when namespace
   teardown was skipped (k8sClient==nil) — orphaning a live namespace
   with no DB pointer. Now skips the DELETE too, leaving the row for a
   later in-cluster run.

5. Storage suspend/unsuspend flapped at the boundary (suspend and
   unsuspend both at limitBytes). Added a hysteresis band: unsuspend only
   below 90% of the limit; the unsuspend loop now also skips resources the
   suspend loop flipped in the same Work() tick.

6. Expiry worker ignored paused/suspended anonymous resources past TTL —
   query was status='active' only, orphaning them. Now expires
   status IN ('active','paused','suspended') past TTL (TTL wins over
   lifecycle state).

7. Confirmed worker storage limits are MiB throughout (*1024*1024); no
   decimal-MB sites — no change needed.

Tests: extraction unit tests for #1, hysteresis dead-band test for #5,
paused/suspended expiry test for #6, prune-job tests for #2. go build /
go vet / go test ./... -short all green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 2, 2026
- orphan_sweep PASS 4: young-namespace (within grace) + age-lookup-error are
  NOT reaped (#8 grace branches).
- emitDeployFailedAudit: dedup-hit skips the INSERT (#15 idempotency branch).
- dbGracePeriodOpener.TerminateActiveGracePeriod: UPDATE success + error-wrap.
- billing terminal downgrade still succeeds when the grace-close errors (#5
  fail-open warn branch).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 2, 2026
…dup, cursor (#78)

* fix(worker): bug-bash batch 2 — grace close, namespace reaper grace, autopsy dedup, cursor

Four confirmed bugs from the 2026-06-02 platform bug bash:

- #5 (P1) billing_reconciler: a terminal Razorpay status downgraded the team
  but left the active payment_grace_periods row open, so payment_grace_reminder
  emitted dunning emails forever and the terminator later re-acted on an
  already-cancelled subscription. Add TerminateActiveGracePeriod to the
  gracePeriodOpener interface (status→'terminated', terminated_at=now()) and
  call it in the terminal-downgrade branch (fail-open).

- #8 (P1) orphan_sweep PASS 4: the customer-namespace reaper excluded 'pending'
  resources from the live-token set AND had no creation-grace, so a sweep
  during two-phase provisioning could DELETE a live, mid-provision namespace.
  Add 'pending' to fetchLiveResourceTokens and a namespace-age grace check
  (skip if younger than orphanNoDBRowGrace) mirroring PASS 3.

- #15 (P2) deploy_failure_autopsy: emitDeployFailedAudit inserted a new
  deploy.failed audit row (new id) on every reconciler retry, and the forwarder
  dedups by audit_id (not deployment) → duplicate failure emails. Make it
  idempotent: skip the INSERT when a deploy.failed row already exists for the
  deployment (metadata->>'deploy_id'). Fail-open on probe error.

- #18 (P2) billing_reconciler scanChargeUndeliverable: jumping the cursor to
  now() on an empty window skipped rows that became visible a moment later
  (clock skew / late commit). Leave the cursor unchanged on count==0 — the 1h
  look-back re-applies and re-scanning the small indexed window is cheap.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(worker): cover bug-bash batch-2 changed lines (100% patch gate)

- orphan_sweep PASS 4: young-namespace (within grace) + age-lookup-error are
  NOT reaped (#8 grace branches).
- emitDeployFailedAudit: dedup-hit skips the INSERT (#15 idempotency branch).
- dbGracePeriodOpener.TerminateActiveGracePeriod: UPDATE success + error-wrap.
- billing terminal downgrade still succeeds when the grace-close errors (#5
  fail-open warn branch).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 4, 2026
) (#86)

A deployments row stuck at status='building' with an EMPTY provider_id is
never reaped: the deploy_status_reconcile sweep skips empty-provider_id rows
unconditionally (correct for a build in flight), so a deploy whose api
goroutine DIED before UpdateDeploymentProviderID (crash mid-runDeploy) keeps
the row 'building' forever and permanently consumes the team's
deployments_apps tier cap.

Fix: in the empty-provider_id branch, if status='building' AND the row is
older than a 15m grace window, reap it to 'failed' via a guarded UPDATE
(double-guarded on status='building' AND empty provider_id, so a row that
raced and acquired a provider_id is a no-op). This frees the tier cap; the
idempotent deploy_failure_autopsy job then emits the deploy.failed audit ->
failure email. Fresh (<15m) builds are still left alone — 15m is well beyond
the normal ~30-90s build.

- add createdAt to activeDeployment + created_at to listActiveDeployments SELECT
- add stuckBuildingGrace (15m) + stuckBuildingReapMessage consts
- add guarded reapStuckBuilding helper (does NOT reuse updateStatus, which
  doesn't gate on provider_id)
- log jobs.deploy_status_reconcile.stuck_building_reaped + reaped counter

Coverage block:
  Symptom:       building deploy w/ empty provider_id never reaped; tier cap leak
  Enumeration:   rg 'if d.providerID == ""' / grep listActiveDeployments SELECT callers
  Sites found:   1 (single sweep loop + its SELECT/scan)
  Sites touched: 1
  Coverage test: TestDeployStatusReconciler_Work_StuckBuildingReaped (+ Fresh/
                 ReapFailed/StuckDeploying/RowWithProviderIDUnaffected)
  Live verified: awaiting post-merge auto-deploy + worker /healthz SHA gate (rule 14)

100% patch coverage (Work/listActiveDeployments/reapStuckBuilding all 100.0%);
jobs package 97.2% (>=95%). make gate green.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant