Skip to content

resources: pause/resume endpoints (Pro+ tier) — suspend without deletion - #60

Merged
mastermanas805 merged 1 commit into
masterfrom
feat/resource-pause-resume-fresh
May 13, 2026
Merged

mastermanas805 merged 1 commit into
masterfrom
feat/resource-pause-resume-fresh

Conversation

@mastermanas805

Copy link
Copy Markdown
Member

Summary

Adds POST /api/v1/resources/:id/pause and /resume so customers can suspend a resource (stop billing the slot, keep the data) instead of deleting it. Tier-gated to Pro+.

  • Paused rows stop counting against per-type resource quotas; storage_bytes still counts toward the storage cap (iron rule — no pause-and-bloat).
  • Connection URL is preserved on resume: same host, same password, same database name. No re-issuance, agent's existing config keeps working.
  • Atomic semantics: provider-side revoke runs before the DB status flip; a provider failure leaves the row 'active' and surfaces a 503. A DB-flip failure after a successful revoke triggers a best-effort rollback so wire and DB state never disagree.

Provider-side semantics

Resource type Pause action Resume action
postgres REVOKE CONNECT ON DATABASE + pg_terminate_backend GRANT CONNECT
redis ACL SETUSER <user> off (keeps password + key-pattern rules intact) ACL SETUSER <user> on
mongodb revokeRolesFromUser (drops readWrite on db_; user stays) grantRolesToUser
queue / storage / webhook status flip only (no provider call) status flip only

Pushback — Redis ACL DELUSER vs SETUSER off: The task brief suggested ACL DELUSER + recreate-on-resume with a stored password hash. We deliberately did not do that. DELUSER is a one-way trip: if the password hash were ever lost (encrypted blob corruption, key rotation gap), the resume path would be unrecoverable. ACL SETUSER <user> off is the lossless reversible primitive — keeps the password and key-pattern rules in place, just disables the user. Symmetric on for resume.

Iron rules + idempotency

  • Pause-on-paused → 409 already_paused with agent_action.
  • Resume-on-active → 409 not_paused with agent_action.
  • Cross-team → 403 forbidden.
  • Hobby / free → 402 upgrade_required (same shape as /provision-twin, carries upgrade_url + agent_action).

Migration 024

ALTER TABLE resources to permit 'paused' in the status check constraint plus a paused_at TIMESTAMPTZ column and a partial index idx_resources_paused. Idempotent (IF EXISTS / IF NOT EXISTS); applied cleanly to a re-run baseline.

Test plan

  • make test-unit green for every package except a pre-existing flake (TestAdminList_AdminUserSees200 — fails on baseline master too).
  • 10 new tests cover: Pro success on pause + resume (with DB-state assertions on status + paused_at), hobby 402, already-paused 409, resume-not-paused 409, cross-team 403, unauthenticated 401, invalid UUID 400, not-found 404, and the storage-still-counts iron rule asserted at the model layer (SumStorageBytesByTeamAndType includes paused rows).
  • TestAgentActionContract green — three new constants (AgentActionPauseRequiresPro, AgentActionResourceAlreadyPaused, AgentActionResourceNotPaused) all satisfy the U3 contract (open with "Tell the user", contain https://instanode.dev/, < 280 chars, contain an action verb).
  • Live verification against the deployed cluster pending /instant-ship.

🤖 Generated with Claude Code

POST /api/v1/resources/:id/pause and /resume let customers suspend a resource
without deleting the data. Paused resources stop counting against per-type
quota (so the slot is "free"), but storage_bytes still counts toward the per-
team storage cap so pause-and-bloat isn't an escape hatch. Connection URL is
preserved on resume — same password, no re-issuance.

Provider-side semantics:
  - postgres: REVOKE CONNECT ON DATABASE (and pg_terminate_backend), reversed
    by GRANT CONNECT on resume.
  - redis:    ACL SETUSER <user> off (and back to on for resume). Chose this
    over ACL DELUSER+recreate because DELUSER drops the password — the
    re-create path requires a stored hash and breaks if the encrypted blob
    is ever lost. ACL off is the lossless reversible primitive.
  - mongodb:  revokeRolesFromUser / grantRolesToUser. The user itself stays;
    only the readWrite role on db_<token> is flipped.
  - queue / storage / webhook: pure status flips. No provider action needed
    (NATS subject, DO Spaces bucket, stored token).

Tier-gated to Pro+ via the same multiEnvTierAllowed helper used by /promote
and /provision-twin so the 402 response shape is symmetric across the Pro+
surface. Carries `agent_action` + `upgrade_url`.

Iron-rule atomicity: provider revoke runs BEFORE the status flip; failure
keeps the row 'active' and returns 503. A DB-flip failure after a successful
revoke triggers a best-effort rollback (resume the provider) so the wire
state and DB state never disagree silently.

Idempotency-as-error: pause-on-paused → 409 already_paused; resume-on-active
→ 409 not_paused. Three new agent_action constants (Pause requires Pro,
Already paused, Not paused) — all pass TestAgentActionContract.

Migration 024 adds 'paused' to the resources.status check constraint and a
new paused_at TIMESTAMPTZ column with a partial index. SumStorageBytesByTeam
AndType now sums status IN ('active','paused') so paused storage stays
visible to quota enforcement.

Tests (10 new) cover: Pro success on pause + resume + DB-state assertions,
hobby 402 (with agent_action shape), already-paused 409, resume-not-paused
409, cross-team 403, unauth 401, invalid UUID 400, not-found 404, and the
storage-still-counts iron rule asserted at the model layer. Plus the new
constants are wired into agent_action_contract_test.go.

`make test-unit` green except for pre-existing flake TestAdminList_AdminUser
Sees200 (fails on baseline master too — independent of this change).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@mastermanas805
mastermanas805 merged commit ec001cd into master May 13, 2026
mastermanas805 added a commit that referenced this pull request Jun 5, 2026
…ent integration (docs/ci/01-CI-INTEGRATION-DESIGN.md) (#265)

* test(ci): Wave 3a+4 — multi-tier e2e factory with_resources + Razorpay test-card payment integration

Wave 3a (#60) + Wave 4 (#61) api-side, per docs/ci/01-CI-INTEGRATION-DESIGN.md.
Both are test-infrastructure, token/secret-gated, inert in prod.

Part A — multi-tier test-user factory (extends /internal/e2e/account)
- The tier param + team/growth-400 gate + paid-tier elevation already shipped
  in the prior factory PR; this adds the remaining brief item: with_resources.
- with_resources=true pre-seeds a small set of FAST, row-only resources
  (webhook + cache) on the minted team — synchronous, no backend RPC, tier-
  snapshotted at the team's tier via the same CreateResource→MarkResourceActive
  two-phase lifecycle a real provision uses — so a journey can start populated.
  Seeded tokens surface on the response (seeded_tokens/seeded_count, always a
  non-null array). Rows are reaped with the team by ReapAccount.
- Default tier stays "free" (NOT pro): the brief assumed current default was
  pro, but it has always been free — changing it would silently hand CI a paid
  tier and break the existing empty-body test. Kept free; documented.
- Tests (registry-iterating, rule 18): every allowed tier round-trips +
  reflects in the minted team's plan_tier (anonymous→free team plan);
  every blocked tier (team/growth) → 400 tier_not_allowed; with_resources
  seeds exactly the handler's seed set as active/owned/tier-snapshotted rows;
  omitting it seeds nothing. team-tier-rejected proof retained.

Part B — Razorpay test-card payment integration (Approach B, webhook-injection)
- New billing_testcard_payment_test.go: real-backend, runs in the api test gate
  against test Postgres (NOT the live-k8s e2e suite). Razorpay TEST MODE needs
  no recurring approval, so this is unblocked despite the prod live-checkout
  operator gate.
- Constructs the EXACT subscription.charged body (raw bytes, signed in place,
  created_at inside the ±5-min replay window), signs hex(HMAC-SHA256(rawBody,
  secret)) via the shared signRazorpayPayload (parity with verifyRazorpaySignature),
  POSTs to /razorpay/webhook with X-Razorpay-Signature + x-razorpay-event-id.
- Handler maps subscription→team via resolveTeamFromNotes: notes.team_id (UUID,
  primary) → fallback GetTeamByRazorpaySubscriptionID(sub.id). Test stamps
  notes.team_id.
- Asserts: tier→pro AND an active permanent resource elevated to pro
  (ElevateResourceTiersByTeam, inside UpgradeTeamAllTiersWithSubscription) AND
  the upgrade contract surfaces (plans.Registry resolves pro limits > free for
  the elevated resource — the /api/v1/capabilities source of truth).
- Negatives: tampered body → 400 no upgrade; wrong secret → 400 no upgrade.
- Idempotency: same x-razorpay-event-id twice → deduped:true + exactly one
  razorpay_webhook_events row + tier upgraded once.
- Failure path: payment.failed → 200, NO upgrade (state asserted, not delivery;
  failure email is webhook-gated per project_payment_failure_email_coverage).

Verify: go build/vet ./... green; new factory + payment tests pass on a fresh
test DB. (Local full -p 1 surfaces pre-existing pollution flakes in untouched
packages — models TestLinkGitHubID, storage anon-quota cap — that pass on a
clean DB; CI's ephemeral DB is authoritative.)

Awaiting web wave to drive UI journeys against the multi-tier factory.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(ci): 100% patch coverage for with_resources seed error arms

CI diff-cover flagged 80% patch coverage: the seed-failure branches were
uncovered (internal_e2e_account.go:297-300 seed_failed 503 arm; 374-379
CreateResource/MarkResourceActive error returns).

- Add e2eSeedFastResources package-var seam (mirrors the e2eSignSessionJWT
  seam) so a test can force CreateAccount's seed_failed 503 arm without making
  the real resources table reject an insert mid-request.
- TestE2EAccount_Create_WithResources_SeedFailure_Returns503: seam → 503
  seed_failed.
- internal_e2e_account_seed_whitebox_test.go (sqlmock): the two
  seedFastResources error arms — CreateResource error + MarkResourceActive
  error after a successful insert.

seedFastResources + CreateAccount now report 100.0% func coverage.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(ci): satisfy CreateResource-callsite + error-code registry guards for the seed path

build-and-test surfaced two registry-iterating guard failures from the new
with_resources seed path:

- TestEveryCreateResourceCallSiteIsFollowedByFinalizeProvision: seedFastResources
  calls models.CreateResource without finalizeProvision. That guard targets the
  orphan-generator shape (insert row → backend provision → 201 without atomic
  credential persistence). The seed is CI-only, row-only (webhook/cache), has NO
  backend RPC and NO credential to persist, so the shape can't occur. Added
  internal_e2e_account.go to the allowList with a justifying comment naming the
  alternate path (CreateResource→MarkResourceActive two-phase lifecycle).

- TestErrorCode_HasAgentAction: the new seed_failed code had no codeToAgentAction
  entry. Added it to coverageAllowlist alongside the other CI-only
  /internal/e2e/account codes (machine-to-machine, not customer-facing).

Both guards green locally.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant