Skip to content

promote: email-link approval workflow for non-dev promotions (dev-env unchanged) - #65

Merged
mastermanas805 merged 2 commits into
masterfrom
feat/promote-approval-workflow-fresh
May 13, 2026
Merged

mastermanas805 merged 2 commits into
masterfrom
feat/promote-approval-workflow-fresh

Conversation

@mastermanas805

Copy link
Copy Markdown
Member

Summary

  • Adds an email-link approval gate to POST /api/v1/stacks/:slug/promote and POST /api/v1/resources/:id/provision-twin. Target envs other than development now persist a pending row in a new promote_approvals table, emit a promote.approval_requested audit row (Brevo forwarder consumes it), and return 202 with approval_id + expires_at + agent_action. Dev-env promotes execute immediately, unchanged.
  • New public route GET /approve/:token (no auth, token IS the credential, 10 req/sec per-IP rate limit via Redis INCR/EXPIRE) atomically flips status to approved and 302-redirects to https://instanode.dev/app/promotions/<id>?approved=1. Expired and used links render HTML pages.
  • Admin routes GET /api/v1/<admin-prefix>/promotions?status=&limit= and POST /api/v1/<admin-prefix>/promotions/:id/reject live behind the existing ADMIN_PATH_PREFIX + RequireAdmin gates.
  • Migration 026_promote_approvals.sql: id, token (UNIQUE), team_id, requested_by_email, promote_kind (stack | resource_twin), promote_payload JSONB, from_env, to_env, status, created_at, expires_at (24h), approved_at, executed_at, rejected_at. Two partial indexes back the token lookup and the worker's pending-execution poll.
  • Worker-side auto-execute polling is out of scope; flagged as follow-up. Manual trigger today: re-call the original endpoint with approval_id in the body — consumeApprovedPromote / consumeApprovedTwin verify (kind,from,to) match before flipping the row to executed.

Agent actions

  • newAgentActionPromoteApprovalSent(toEnv, email) → "Tell the user the promote to <to_env> requires email approval. Check <email> for a link expiring in 24h. Dev-env promotes skip this step. Track at https://instanode.dev/app/promotions."
  • AgentActionPromoteTokenExpired → "Tell the user the approval link expired. Re-request the promote at https://instanode.dev/app — links are valid for 24h."

Both registered in TestAgentActionContract and TestAgentActionContract_RegistryCoverage.

Iron rules

  • Tests in same PR. make test-unit all-green (test-unit: all packages green).
  • Cryptographic tokens via crypto/rand only; TestPromoteApproval_TokenGeneration_Unique asserts non-repetition across 32 calls.
  • Single-use approval enforced atomically: UPDATE ... WHERE status='pending' AND expires_at > now(); TestPromoteApproval_ApproveIsAtomic proves exactly one of two concurrent approve calls wins.
  • Per-IP rate-limit on /approve/:token defends the 32-byte token space.

Test plan

  • make test-unit — all packages green
  • TestAgentActionContract includes both new strings
  • 12 new tests in promote_approval_test.go cover dev-bypass, non-dev pending, audit emit, approve flow, expired flow, used-token flow, reject, no-auth on /approve/:token, token uniqueness, atomic single-use, list-newest-first, contract compliance
  • Existing TestStackPromote_* + TestProvisionTwin_* retargeted at to="development" to keep exercising the underlying contract (the non-dev approval flow has its own coverage)
  • (follow-up) Worker-side polling job consumes status='approved' AND executed_at IS NULL and runs the cached promote_payload

🤖 Generated with Claude Code

mastermanas805 and others added 2 commits May 13, 2026 09:42
Adds an email-link approval gate to POST /api/v1/stacks/:slug/promote and
POST /api/v1/resources/:id/provision-twin. Any target env other than
"development" now persists a pending row in promote_approvals, emits a
promote.approval_requested audit row (the Brevo forwarder picks this up
and sends the email), and returns 202 with an approval_id + expires_at +
agent_action telling the user to check their inbox. Dev-env promotes
execute immediately and are entirely unchanged.

Migration 026 backs a single table with status (pending|approved|rejected|
expired|executed), a 32-byte crypto/rand token, 24h expiry, and JSONB
payload so a worker (separate PR) can replay the original POST. Single-
use approval is enforced via atomic UPDATE ... WHERE status='pending' AND
expires_at > now() — concurrent clicks resolve to exactly one approval.

Public route GET /approve/:token requires no auth (the token IS the
credential) and is rate-limited to 10 req/sec per IP via a Redis INCR
with 2s TTL — defends the 32-byte token space against brute force while
failing open on Redis errors. Admin routes GET /promotions and POST
/promotions/:id/reject live behind the existing ADMIN_PATH_PREFIX +
ADMIN_EMAILS gates.

Existing stack_promote / provision_twin tests retargeted at to="development"
so they exercise the pre-approval contract that's still under test;
non-dev coverage lives in the new promote_approval_test.go suite (12
tests, all 9 prompt cases plus token-uniqueness, atomic-single-use, and
the admin-list shape).

Worker-side polling that auto-executes approved rows is out of scope for
this PR — flag as follow-up. Manual trigger today: re-call the original
endpoint with approval_id in the body; consumeApprovedPromote /
consumeApprovedTwin verify (kind,from,to) match before flipping to
executed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…al-workflow-fresh

# Conflicts:
#	internal/handlers/agent_action.go
#	internal/handlers/agent_action_contract_test.go
#	internal/models/audit_kinds.go
@mastermanas805
mastermanas805 merged commit e484984 into master May 13, 2026
mastermanas805 added a commit that referenced this pull request May 14, 2026
…obby_plus copy (#107)

Wave FIX-H wraps up the BugBash B36 backup-correctness items so customer-
facing restore stops being one accidental retry away from a destroyed
database.

#57/#Q45 — restore replay guard
  Adds models.HasInflightRestore; a second POST /restore for the same
  resource while a prior row is pending/running returns 409
  restore_in_progress + AgentActionRestoreInflight. Fail-CLOSED on DB
  error (concurrent pg_restore --clean races itself).

#58/#A2 — restore to a new DB
  Accepts optional target_resource_id. Worker restores into the target
  (same team only); audit row carries both source + target ids. Skips
  the destructive-ack ceremony because the agent already opted into a
  fresh database.

#59 — backup integrity
  Migration 043 adds a nullable sha256 TEXT column to resource_backups.
  Worker stamps the digest at finalize; restore handler verifies before
  pg_restore. NULL on legacy rows is logged + accepted (fail-open on
  pre-043 data).

#64/#Q46 — cross-tenant 404
  GetBackupByIDForTeam joins resources to scope by team; cross-tenant
  backup_id guess now returns 404 backup_not_found instead of
  400 backup_resource_mismatch. Matches FIX-B tenant-isolation posture.

#65/#Q47 — refund quota on failure
  New internal endpoint POST /internal/teams/:id/backup-quota/refund
  (WORKER_INTERNAL_JWT_SECRET HS256, fail-closed when unset). The
  worker calls this when a MANUAL backup fails terminally so the
  team's daily manual-backups counter is credited back.

#66/#Q48 — hobby agent_action points to Hobby Plus
  AgentActionRestoreRequiresHobbyPlus added. The Hobby-tier restore 402
  now nudges Hobby Plus ($19/mo, restore enabled) instead of skipping
  the customer past the cheapest restore-enabled plan onto Pro ($49).

#67/#Q49 — destructive ack required for in-place restore
  In-place restore (no target_resource_id) now requires
  destructive_acknowledgment: true in the body. pg_restore --clean drops
  every table — refusing without an explicit ack prevents an agent
  testing a backup from wiping a live customer DB.

#Q50 — RPO/RTO on /capabilities
  Adds rpo_minutes + rto_minutes per tier in plans.yaml (anonymous/free
  = 0/0, hobby/hobby_plus = 1440/30, pro/team = 60/15) wired into the
  /api/v1/capabilities matrix via plans.Registry.RPOMinutes/RTOMinutes
  (added in instant.dev/common/plans#13).

Tests
  - +5 backup_test.go cases: ReplayBlocked, TargetNewDB,
    RequiresDestructiveAck, HobbyAgentActionPointsToHobbyPlus,
    CrossTenantBackupID_404
  - All prior backup tests updated to include destructive_acknowledgment
  - Contract test enforces the four new agent_action constants

DO NOT TOUCH list respected — no edits to email/, dpop.go, circuit/.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request May 21, 2026
…obby_plus copy

Wave FIX-H wraps up the BugBash B36 backup-correctness items so customer-
facing restore stops being one accidental retry away from a destroyed
database.

#57/#Q45 — restore replay guard
  Adds models.HasInflightRestore; a second POST /restore for the same
  resource while a prior row is pending/running returns 409
  restore_in_progress + AgentActionRestoreInflight. Fail-CLOSED on DB
  error (concurrent pg_restore --clean races itself).

#58/#A2 — restore to a new DB
  Accepts optional target_resource_id. Worker restores into the target
  (same team only); audit row carries both source + target ids. Skips
  the destructive-ack ceremony because the agent already opted into a
  fresh database.

#59 — backup integrity
  Migration 043 adds a nullable sha256 TEXT column to resource_backups.
  Worker stamps the digest at finalize; restore handler verifies before
  pg_restore. NULL on legacy rows is logged + accepted (fail-open on
  pre-043 data).

#64/#Q46 — cross-tenant 404
  GetBackupByIDForTeam joins resources to scope by team; cross-tenant
  backup_id guess now returns 404 backup_not_found instead of
  400 backup_resource_mismatch. Matches FIX-B tenant-isolation posture.

#65/#Q47 — refund quota on failure
  New internal endpoint POST /internal/teams/:id/backup-quota/refund
  (WORKER_INTERNAL_JWT_SECRET HS256, fail-closed when unset). The
  worker calls this when a MANUAL backup fails terminally so the
  team's daily manual-backups counter is credited back.

#66/#Q48 — hobby agent_action points to Hobby Plus
  AgentActionRestoreRequiresHobbyPlus added. The Hobby-tier restore 402
  now nudges Hobby Plus ($19/mo, restore enabled) instead of skipping
  the customer past the cheapest restore-enabled plan onto Pro ($49).

#67/#Q49 — destructive ack required for in-place restore
  In-place restore (no target_resource_id) now requires
  destructive_acknowledgment: true in the body. pg_restore --clean drops
  every table — refusing without an explicit ack prevents an agent
  testing a backup from wiping a live customer DB.

#Q50 — RPO/RTO on /capabilities
  Adds rpo_minutes + rto_minutes per tier in plans.yaml (anonymous/free
  = 0/0, hobby/hobby_plus = 1440/30, pro/team = 60/15) wired into the
  /api/v1/capabilities matrix via plans.Registry.RPOMinutes/RTOMinutes
  (added in instant.dev/common/plans#13).

Tests
  - +5 backup_test.go cases: ReplayBlocked, TargetNewDB,
    RequiresDestructiveAck, HobbyAgentActionPointsToHobbyPlus,
    CrossTenantBackupID_404
  - All prior backup tests updated to include destructive_acknowledgment
  - Contract test enforces the four new agent_action constants

DO NOT TOUCH list respected — no edits to email/, dpop.go, circuit/.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant