feat(worker): backend-agnostic storage scanner (any S3-compat service) - #5
Merged
Merged
Conversation
…/GCS)
The storage_bytes scanner now works against any S3-compatible service
without code changes — only env-var swap. Pairs with the api repo's
storage-provider refactor on the same branch name.
Changes:
config.go adds OBJECT_STORE_* env vars (ENDPOINT / ACCESS_KEY /
SECRET_KEY / BUCKET / REGION / SECURE), with the legacy
MINIO_* names kept as fallback for backward compat.
workers.go StorageScanner constructor now reads OBJECT_STORE_*
instead of MINIO_*. The MinIO admin client used by the
ExpireAnonymousWorker stays gated on legacy MINIO_*
because per-IAM-user cleanup is only relevant for
self-hosted MinIO (the shared-key backend doesn't
create per-customer IAM in the first place).
storage_minio.go NewMinIOStorageScanner auto-detects TLS:
explicit https:// / http:// prefixes win; otherwise
a heuristic enables TLS for known managed S3 vendor
hostnames (digitaloceanspaces.com, amazonaws.com,
cloudflarestorage.com, googleapis.com, wasabisys.com,
backblazeb2.com). Self-hosted in-cluster MinIO
continues to use plain HTTP. The minio-go SDK speaks
plain S3 against all of these.
Verified: go build + go test ./internal/jobs/... pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805
added a commit
that referenced
this pull request
May 17, 2026
P2 worker-side fixes from BUGHUNT-REPORT-2026-05-17.md:
1. GeoLite2 refresh corrupted the .mmdb after every run — MaxMind serves a
gzipped tarball (suffix=tar.gz) but the job io.Copy'd the raw bytes
straight to the .mmdb path. Now gunzip + untar and extract only the
*.mmdb member (geodb.go: extractGeoLite2MMDB).
2. razorpay_webhook_events dedup table was never pruned (migration 033
envisioned it; no job shipped). Added RazorpayWebhookPruneWorker — a
daily DELETE of rows > 30 days, registered alongside uptime_retention.
3. Billing reconciler free-downgrade stranded paid resources: a terminal
Razorpay status with PaidCount==0 set plan_tier='free' (the 24h-TTL
ephemeral tier) while resources kept a paid tier + expires_at=NULL.
Now every terminal status downgrades to 'hobby' (lowest paid tier),
matching the non-zero-paid branch — keeps team-tier and resource-tier
coherent.
4. ExpireStacksWorker hard-deleted the stacks row even when namespace
teardown was skipped (k8sClient==nil) — orphaning a live namespace
with no DB pointer. Now skips the DELETE too, leaving the row for a
later in-cluster run.
5. Storage suspend/unsuspend flapped at the boundary (suspend and
unsuspend both at limitBytes). Added a hysteresis band: unsuspend only
below 90% of the limit; the unsuspend loop now also skips resources the
suspend loop flipped in the same Work() tick.
6. Expiry worker ignored paused/suspended anonymous resources past TTL —
query was status='active' only, orphaning them. Now expires
status IN ('active','paused','suspended') past TTL (TTL wins over
lifecycle state).
7. Confirmed worker storage limits are MiB throughout (*1024*1024); no
decimal-MB sites — no change needed.
Tests: extraction unit tests for #1, hysteresis dead-band test for #5,
paused/suspended expiry test for #6, prune-job tests for #2. go build /
go vet / go test ./... -short all green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805
added a commit
that referenced
this pull request
Jun 2, 2026
- orphan_sweep PASS 4: young-namespace (within grace) + age-lookup-error are NOT reaped (#8 grace branches). - emitDeployFailedAudit: dedup-hit skips the INSERT (#15 idempotency branch). - dbGracePeriodOpener.TerminateActiveGracePeriod: UPDATE success + error-wrap. - billing terminal downgrade still succeeds when the grace-close errors (#5 fail-open warn branch). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805
added a commit
that referenced
this pull request
Jun 2, 2026
…dup, cursor (#78) * fix(worker): bug-bash batch 2 — grace close, namespace reaper grace, autopsy dedup, cursor Four confirmed bugs from the 2026-06-02 platform bug bash: - #5 (P1) billing_reconciler: a terminal Razorpay status downgraded the team but left the active payment_grace_periods row open, so payment_grace_reminder emitted dunning emails forever and the terminator later re-acted on an already-cancelled subscription. Add TerminateActiveGracePeriod to the gracePeriodOpener interface (status→'terminated', terminated_at=now()) and call it in the terminal-downgrade branch (fail-open). - #8 (P1) orphan_sweep PASS 4: the customer-namespace reaper excluded 'pending' resources from the live-token set AND had no creation-grace, so a sweep during two-phase provisioning could DELETE a live, mid-provision namespace. Add 'pending' to fetchLiveResourceTokens and a namespace-age grace check (skip if younger than orphanNoDBRowGrace) mirroring PASS 3. - #15 (P2) deploy_failure_autopsy: emitDeployFailedAudit inserted a new deploy.failed audit row (new id) on every reconciler retry, and the forwarder dedups by audit_id (not deployment) → duplicate failure emails. Make it idempotent: skip the INSERT when a deploy.failed row already exists for the deployment (metadata->>'deploy_id'). Fail-open on probe error. - #18 (P2) billing_reconciler scanChargeUndeliverable: jumping the cursor to now() on an empty window skipped rows that became visible a moment later (clock skew / late commit). Leave the cursor unchanged on count==0 — the 1h look-back re-applies and re-scanning the small indexed window is cheap. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(worker): cover bug-bash batch-2 changed lines (100% patch gate) - orphan_sweep PASS 4: young-namespace (within grace) + age-lookup-error are NOT reaped (#8 grace branches). - emitDeployFailedAudit: dedup-hit skips the INSERT (#15 idempotency branch). - dbGracePeriodOpener.TerminateActiveGracePeriod: UPDATE success + error-wrap. - billing terminal downgrade still succeeds when the grace-close errors (#5 fail-open warn branch). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805
added a commit
that referenced
this pull request
Jun 4, 2026
) (#86) A deployments row stuck at status='building' with an EMPTY provider_id is never reaped: the deploy_status_reconcile sweep skips empty-provider_id rows unconditionally (correct for a build in flight), so a deploy whose api goroutine DIED before UpdateDeploymentProviderID (crash mid-runDeploy) keeps the row 'building' forever and permanently consumes the team's deployments_apps tier cap. Fix: in the empty-provider_id branch, if status='building' AND the row is older than a 15m grace window, reap it to 'failed' via a guarded UPDATE (double-guarded on status='building' AND empty provider_id, so a row that raced and acquired a provider_id is a no-op). This frees the tier cap; the idempotent deploy_failure_autopsy job then emits the deploy.failed audit -> failure email. Fresh (<15m) builds are still left alone — 15m is well beyond the normal ~30-90s build. - add createdAt to activeDeployment + created_at to listActiveDeployments SELECT - add stuckBuildingGrace (15m) + stuckBuildingReapMessage consts - add guarded reapStuckBuilding helper (does NOT reuse updateStatus, which doesn't gate on provider_id) - log jobs.deploy_status_reconcile.stuck_building_reaped + reaped counter Coverage block: Symptom: building deploy w/ empty provider_id never reaped; tier cap leak Enumeration: rg 'if d.providerID == ""' / grep listActiveDeployments SELECT callers Sites found: 1 (single sweep loop + its SELECT/scan) Sites touched: 1 Coverage test: TestDeployStatusReconciler_Work_StuckBuildingReaped (+ Fresh/ ReapFailed/StuckDeploying/RowWithProviderIDUnaffected) Live verified: awaiting post-merge auto-deploy + worker /healthz SHA gate (rule 14) 100% patch coverage (Work/listActiveDeployments/reapStuckBuilding all 100.0%); jobs package 97.2% (>=95%). make gate green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pairs with api PR #18. Storage scanner reads OBJECT_STORE_* env vars (with MINIO_* fallback) and auto-detects TLS for managed S3 vendors. Build + tests pass.