Skip to content

hotfix: allow 'reaped' status in resources_status_check (migration 024) - #67

Merged
mastermanas805 merged 1 commit into
masterfrom
hotfix/migration-024-allow-reaped
May 13, 2026
Merged

mastermanas805 merged 1 commit into
masterfrom
hotfix/migration-024-allow-reaped

Conversation

@mastermanas805

Copy link
Copy Markdown
Member

Summary

  • v6.0.0 rollout crash-looped on 024_resources_paused_status.sql because the new CHECK constraint omits the legacy 'reaped' status that ~220 prod rows carry.
  • Add 'reaped' to the allowed set to preserve historical signal (worker-reaped vs user-deleted distinction).

Verification

Repro on cluster:

db.migrations.applying file=024_resources_paused_status.sql
main.migrations_failed error="pq: check constraint resources_status_check of relation resources is violated by some row"

Prod status distribution:

 deleted | 232
 reaped  | 220
 active  |  11
 expired |   5

No code references 'reaped' — purely legacy state. Constraint patch is non-destructive.

Test plan

  • grep for 'reaped' in api+worker — no code emits or reads it
  • post-merge: rebuild v6.0.1, redeploy, verify migration 024 applies cleanly
  • verify GET /healthz reports v6.0.1

Migration 024 added a CHECK constraint with only ('active','paused','expired','deleted'),
but ~220 prod rows carry status='reaped' from a prior worker cleanup job. The
migration failed in v6.0.0 rollout with:
  check constraint "resources_status_check" of relation "resources" is violated by some row

Keep 'reaped' in the allowed set rather than collapsing it into 'deleted' —
losing the worker-reaped vs user-deleted distinction would erase historical signal.

Verified: ENVIRONMENT=production, v6.0.0-2026-05-13 image, instance crash-looped on this exact constraint after 25 migrations succeeded.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@mastermanas805
mastermanas805 merged commit cb634f1 into master May 13, 2026
mastermanas805 added a commit that referenced this pull request May 14, 2026
Migration 024 added a CHECK constraint with only ('active','paused','expired','deleted'),
but ~220 prod rows carry status='reaped' from a prior worker cleanup job. The
migration failed in v6.0.0 rollout with:
  check constraint "resources_status_check" of relation "resources" is violated by some row

Keep 'reaped' in the allowed set rather than collapsing it into 'deleted' —
losing the worker-reaped vs user-deleted distinction would erase historical signal.

Verified: ENVIRONMENT=production, v6.0.0-2026-05-13 image, instance crash-looped on this exact constraint after 25 migrations succeeded.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request May 14, 2026
…#79)

Migration 024 added a CHECK constraint with only ('active','paused','expired','deleted'),
but ~220 prod rows carry status='reaped' from a prior worker cleanup job. The
migration failed in v6.0.0 rollout with:
  check constraint "resources_status_check" of relation "resources" is violated by some row

Keep 'reaped' in the allowed set rather than collapsing it into 'deleted' —
losing the worker-reaped vs user-deleted distinction would erase historical signal.

Verified: ENVIRONMENT=production, v6.0.0-2026-05-13 image, instance crash-looped on this exact constraint after 25 migrations succeeded.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request May 14, 2026
…obby_plus copy (#107)

Wave FIX-H wraps up the BugBash B36 backup-correctness items so customer-
facing restore stops being one accidental retry away from a destroyed
database.

#57/#Q45 — restore replay guard
  Adds models.HasInflightRestore; a second POST /restore for the same
  resource while a prior row is pending/running returns 409
  restore_in_progress + AgentActionRestoreInflight. Fail-CLOSED on DB
  error (concurrent pg_restore --clean races itself).

#58/#A2 — restore to a new DB
  Accepts optional target_resource_id. Worker restores into the target
  (same team only); audit row carries both source + target ids. Skips
  the destructive-ack ceremony because the agent already opted into a
  fresh database.

#59 — backup integrity
  Migration 043 adds a nullable sha256 TEXT column to resource_backups.
  Worker stamps the digest at finalize; restore handler verifies before
  pg_restore. NULL on legacy rows is logged + accepted (fail-open on
  pre-043 data).

#64/#Q46 — cross-tenant 404
  GetBackupByIDForTeam joins resources to scope by team; cross-tenant
  backup_id guess now returns 404 backup_not_found instead of
  400 backup_resource_mismatch. Matches FIX-B tenant-isolation posture.

#65/#Q47 — refund quota on failure
  New internal endpoint POST /internal/teams/:id/backup-quota/refund
  (WORKER_INTERNAL_JWT_SECRET HS256, fail-closed when unset). The
  worker calls this when a MANUAL backup fails terminally so the
  team's daily manual-backups counter is credited back.

#66/#Q48 — hobby agent_action points to Hobby Plus
  AgentActionRestoreRequiresHobbyPlus added. The Hobby-tier restore 402
  now nudges Hobby Plus ($19/mo, restore enabled) instead of skipping
  the customer past the cheapest restore-enabled plan onto Pro ($49).

#67/#Q49 — destructive ack required for in-place restore
  In-place restore (no target_resource_id) now requires
  destructive_acknowledgment: true in the body. pg_restore --clean drops
  every table — refusing without an explicit ack prevents an agent
  testing a backup from wiping a live customer DB.

#Q50 — RPO/RTO on /capabilities
  Adds rpo_minutes + rto_minutes per tier in plans.yaml (anonymous/free
  = 0/0, hobby/hobby_plus = 1440/30, pro/team = 60/15) wired into the
  /api/v1/capabilities matrix via plans.Registry.RPOMinutes/RTOMinutes
  (added in instant.dev/common/plans#13).

Tests
  - +5 backup_test.go cases: ReplayBlocked, TargetNewDB,
    RequiresDestructiveAck, HobbyAgentActionPointsToHobbyPlus,
    CrossTenantBackupID_404
  - All prior backup tests updated to include destructive_acknowledgment
  - Contract test enforces the four new agent_action constants

DO NOT TOUCH list respected — no edits to email/, dpop.go, circuit/.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request May 21, 2026
…obby_plus copy

Wave FIX-H wraps up the BugBash B36 backup-correctness items so customer-
facing restore stops being one accidental retry away from a destroyed
database.

#57/#Q45 — restore replay guard
  Adds models.HasInflightRestore; a second POST /restore for the same
  resource while a prior row is pending/running returns 409
  restore_in_progress + AgentActionRestoreInflight. Fail-CLOSED on DB
  error (concurrent pg_restore --clean races itself).

#58/#A2 — restore to a new DB
  Accepts optional target_resource_id. Worker restores into the target
  (same team only); audit row carries both source + target ids. Skips
  the destructive-ack ceremony because the agent already opted into a
  fresh database.

#59 — backup integrity
  Migration 043 adds a nullable sha256 TEXT column to resource_backups.
  Worker stamps the digest at finalize; restore handler verifies before
  pg_restore. NULL on legacy rows is logged + accepted (fail-open on
  pre-043 data).

#64/#Q46 — cross-tenant 404
  GetBackupByIDForTeam joins resources to scope by team; cross-tenant
  backup_id guess now returns 404 backup_not_found instead of
  400 backup_resource_mismatch. Matches FIX-B tenant-isolation posture.

#65/#Q47 — refund quota on failure
  New internal endpoint POST /internal/teams/:id/backup-quota/refund
  (WORKER_INTERNAL_JWT_SECRET HS256, fail-closed when unset). The
  worker calls this when a MANUAL backup fails terminally so the
  team's daily manual-backups counter is credited back.

#66/#Q48 — hobby agent_action points to Hobby Plus
  AgentActionRestoreRequiresHobbyPlus added. The Hobby-tier restore 402
  now nudges Hobby Plus ($19/mo, restore enabled) instead of skipping
  the customer past the cheapest restore-enabled plan onto Pro ($49).

#67/#Q49 — destructive ack required for in-place restore
  In-place restore (no target_resource_id) now requires
  destructive_acknowledgment: true in the body. pg_restore --clean drops
  every table — refusing without an explicit ack prevents an agent
  testing a backup from wiping a live customer DB.

#Q50 — RPO/RTO on /capabilities
  Adds rpo_minutes + rto_minutes per tier in plans.yaml (anonymous/free
  = 0/0, hobby/hobby_plus = 1440/30, pro/team = 60/15) wired into the
  /api/v1/capabilities matrix via plans.Registry.RPOMinutes/RTOMinutes
  (added in instant.dev/common/plans#13).

Tests
  - +5 backup_test.go cases: ReplayBlocked, TargetNewDB,
    RequiresDestructiveAck, HobbyAgentActionPointsToHobbyPlus,
    CrossTenantBackupID_404
  - All prior backup tests updated to include destructive_acknowledgment
  - Contract test enforces the four new agent_action constants

DO NOT TOUCH list respected — no edits to email/, dpop.go, circuit/.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant