Skip to content

billing: payment-failure dunning + 7d grace + auto-cancel (with NR dashboard + alerts) - #66

Merged
mastermanas805 merged 3 commits into
masterfrom
feat/payment-failure-dunning-fresh
May 13, 2026
Merged

mastermanas805 merged 3 commits into
masterfrom
feat/payment-failure-dunning-fresh

Conversation

@mastermanas805

Copy link
Copy Markdown
Member

Summary

  • What: wires the dunning state machine on the api side. Razorpay subscription.charged_failed opens a 7-day grace row; subscription.charged during grace flips it to recovered. Four new audit kinds (payment.grace_started, payment.grace_reminder, payment.grace_recovered, payment.grace_terminated) feed the Brevo email forwarder.
  • Why: today's billing flow assumes the happy path. There's no in-between for a card that declines while the customer is otherwise in good standing — they discover their account is gone only when they next visit the dashboard. This PR is the trigger half of the standard SaaS dunning pattern (NOT customer-initiated cancel, which remains support-only per project rules).
  • Migration 027 adds payment_grace_periods with a partial-unique index on team_id WHERE status='active'. Razorpay webhook redeliveries hit the constraint and silently no-op so customers don't get duplicate reminder emails.

What ships in this PR (api side)

  • internal/db/migrations/027_payment_dunning.sql — table + two indexes (status/expires_at for worker sweeps, partial-unique team_id WHERE active for idempotency).
  • internal/models/payment_grace_periods.go — Create (returns ErrPaymentGraceAlreadyActive on duplicate-active), GetActive, MarkRecovered, MarkTerminated. State transitions enforced by the application; the partial-unique index is the concurrency guard.
  • internal/handlers/billing.go — extended Razorpay webhook switch: new case for subscription.charged_failed, recovery path tacked onto subscription.charged. All audit emits fail-open.
  • internal/models/audit_kinds.go — 4 new constants with full docstrings.
  • internal/handlers/openapi.go — webhook description updated.
  • internal/testhelpers/testhelpers.go — table mirror so handler tests work without depending on the SQL-migration step.

What ships in a separate follow-up PR (worker side)

  • payment_grace_reminder River job — every-30-min sweep that finds active grace rows with last_reminder_at < now() - interval '6 hours', emits payment.grace_reminder audit row (Brevo forwarder picks it up), bumps reminders_sent + last_reminder_at. Up to 28 reminders over the 7-day window.
  • payment_grace_terminator River job — hourly sweep over status='active' AND expires_at < now(). Calls Razorpay subscription cancel (best-effort, log+audit on failure), soft-deletes all team resources (transactionally), emits payment.grace_terminated. This is the destructive half — kept out of api per brief.

Pushback / confirmation: the worker-side jobs are flagged as follow-ups as the brief required. Iron rule honoured — this PR only ships the trigger + grace state + recovery + reminder-emit-via-audit-kind, no destructive work.

Audit kinds + Brevo templates

Audit kind Brevo template Fires when
payment.grace_started instanode-payment-grace-started-v1 api handles subscription.charged_failed
payment.grace_reminder instanode-payment-grace-reminder-v1 worker reminder job (follow-up PR) every 6h
payment.grace_recovered instanode-payment-grace-recovered-v1 api handles subscription.charged during active grace
payment.grace_terminated instanode-payment-grace-terminated-v1 worker terminator job (follow-up PR) at expires_at

NR observability (same PR per memory rule)

infra/newrelic/ lives in this repo, so dashboards + alerts ship here:

  • dashboards/billing-dunning.json — customers in active grace, reminders/day, recovery-rate billboard, auto-terminations/week, webhook p95 latency.
  • alerts/payment-failure-spike.json — critical if >10 grace_started in 1h (Razorpay outage or webhook double-fire).
  • alerts/dunning-recovery-rate-low.json — critical if 7d rolling recovery rate < 30% (unhealthy dunning UX).

bake_test.sh updated to include the new files — 42/42 schema-shape assertions green.

Test plan

  • make test-unit green across all 20 packages
  • 12 model tests pass (happy path, duplicate-active rejection, after-recovered-allows-new, cross-team isolation, transition immutability, redelivery idempotency)
  • 6 webhook tests pass (charge_failed opens grace + emits audit, redelivery is no-op, charged-during-grace flips to recovered, charged-without-grace emits no recovery audit, cross-team isolation, fail-open audit-miss does not roll back grace row)
  • TestAgentActionContract green
  • bake_test.sh 42/42 NR schema assertions pass
  • Worker-side reminder + terminator jobs ship in a separate follow-up PR (flagged above)

🤖 Generated with Claude Code

mastermanas805 and others added 3 commits May 13, 2026 09:45
Wires the dunning state machine on the api side: Razorpay subscription.charged_failed
opens a 7-day grace row (payment_grace_periods), subscription.charged during the
grace window flips it to recovered. Emits four new audit kinds (grace_started,
grace_reminder, grace_recovered, grace_terminated) consumed by the Brevo forwarder
for the dunning email sequence.

Idempotency comes from a partial-unique index on team_id WHERE status='active' —
Razorpay webhook redeliveries hit the constraint and silently no-op so customers
don't get duplicate reminder emails. Cross-team isolation + fail-open audit emit
are covered.

The worker-side reminder cadence (every 6h, up to 28 over 7d) and the destructive
terminator job (Razorpay cancel + resource soft-delete) ship in a separate worker
PR — this PR delivers the trigger + state machine + recovery path.

NR dashboard + two alerts ship in the same PR (billing-dunning.json,
payment-failure-spike.json for Razorpay-outage detection, dunning-recovery-rate-low.json
for unhealthy dunning UX). bake_test.sh updated; 42/42 schema-shape assertions
green.

Test counts: 12 model tests + 6 handler tests = 18 new tests. make test-unit
green across all 20 packages. TestAgentActionContract green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@mastermanas805
mastermanas805 merged commit b8ace35 into master May 13, 2026
mastermanas805 added a commit that referenced this pull request May 14, 2026
…obby_plus copy (#107)

Wave FIX-H wraps up the BugBash B36 backup-correctness items so customer-
facing restore stops being one accidental retry away from a destroyed
database.

#57/#Q45 — restore replay guard
  Adds models.HasInflightRestore; a second POST /restore for the same
  resource while a prior row is pending/running returns 409
  restore_in_progress + AgentActionRestoreInflight. Fail-CLOSED on DB
  error (concurrent pg_restore --clean races itself).

#58/#A2 — restore to a new DB
  Accepts optional target_resource_id. Worker restores into the target
  (same team only); audit row carries both source + target ids. Skips
  the destructive-ack ceremony because the agent already opted into a
  fresh database.

#59 — backup integrity
  Migration 043 adds a nullable sha256 TEXT column to resource_backups.
  Worker stamps the digest at finalize; restore handler verifies before
  pg_restore. NULL on legacy rows is logged + accepted (fail-open on
  pre-043 data).

#64/#Q46 — cross-tenant 404
  GetBackupByIDForTeam joins resources to scope by team; cross-tenant
  backup_id guess now returns 404 backup_not_found instead of
  400 backup_resource_mismatch. Matches FIX-B tenant-isolation posture.

#65/#Q47 — refund quota on failure
  New internal endpoint POST /internal/teams/:id/backup-quota/refund
  (WORKER_INTERNAL_JWT_SECRET HS256, fail-closed when unset). The
  worker calls this when a MANUAL backup fails terminally so the
  team's daily manual-backups counter is credited back.

#66/#Q48 — hobby agent_action points to Hobby Plus
  AgentActionRestoreRequiresHobbyPlus added. The Hobby-tier restore 402
  now nudges Hobby Plus ($19/mo, restore enabled) instead of skipping
  the customer past the cheapest restore-enabled plan onto Pro ($49).

#67/#Q49 — destructive ack required for in-place restore
  In-place restore (no target_resource_id) now requires
  destructive_acknowledgment: true in the body. pg_restore --clean drops
  every table — refusing without an explicit ack prevents an agent
  testing a backup from wiping a live customer DB.

#Q50 — RPO/RTO on /capabilities
  Adds rpo_minutes + rto_minutes per tier in plans.yaml (anonymous/free
  = 0/0, hobby/hobby_plus = 1440/30, pro/team = 60/15) wired into the
  /api/v1/capabilities matrix via plans.Registry.RPOMinutes/RTOMinutes
  (added in instant.dev/common/plans#13).

Tests
  - +5 backup_test.go cases: ReplayBlocked, TargetNewDB,
    RequiresDestructiveAck, HobbyAgentActionPointsToHobbyPlus,
    CrossTenantBackupID_404
  - All prior backup tests updated to include destructive_acknowledgment
  - Contract test enforces the four new agent_action constants

DO NOT TOUCH list respected — no edits to email/, dpop.go, circuit/.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request May 21, 2026
…obby_plus copy

Wave FIX-H wraps up the BugBash B36 backup-correctness items so customer-
facing restore stops being one accidental retry away from a destroyed
database.

#57/#Q45 — restore replay guard
  Adds models.HasInflightRestore; a second POST /restore for the same
  resource while a prior row is pending/running returns 409
  restore_in_progress + AgentActionRestoreInflight. Fail-CLOSED on DB
  error (concurrent pg_restore --clean races itself).

#58/#A2 — restore to a new DB
  Accepts optional target_resource_id. Worker restores into the target
  (same team only); audit row carries both source + target ids. Skips
  the destructive-ack ceremony because the agent already opted into a
  fresh database.

#59 — backup integrity
  Migration 043 adds a nullable sha256 TEXT column to resource_backups.
  Worker stamps the digest at finalize; restore handler verifies before
  pg_restore. NULL on legacy rows is logged + accepted (fail-open on
  pre-043 data).

#64/#Q46 — cross-tenant 404
  GetBackupByIDForTeam joins resources to scope by team; cross-tenant
  backup_id guess now returns 404 backup_not_found instead of
  400 backup_resource_mismatch. Matches FIX-B tenant-isolation posture.

#65/#Q47 — refund quota on failure
  New internal endpoint POST /internal/teams/:id/backup-quota/refund
  (WORKER_INTERNAL_JWT_SECRET HS256, fail-closed when unset). The
  worker calls this when a MANUAL backup fails terminally so the
  team's daily manual-backups counter is credited back.

#66/#Q48 — hobby agent_action points to Hobby Plus
  AgentActionRestoreRequiresHobbyPlus added. The Hobby-tier restore 402
  now nudges Hobby Plus ($19/mo, restore enabled) instead of skipping
  the customer past the cheapest restore-enabled plan onto Pro ($49).

#67/#Q49 — destructive ack required for in-place restore
  In-place restore (no target_resource_id) now requires
  destructive_acknowledgment: true in the body. pg_restore --clean drops
  every table — refusing without an explicit ack prevents an agent
  testing a backup from wiping a live customer DB.

#Q50 — RPO/RTO on /capabilities
  Adds rpo_minutes + rto_minutes per tier in plans.yaml (anonymous/free
  = 0/0, hobby/hobby_plus = 1440/30, pro/team = 60/15) wired into the
  /api/v1/capabilities matrix via plans.Registry.RPOMinutes/RTOMinutes
  (added in instant.dev/common/plans#13).

Tests
  - +5 backup_test.go cases: ReplayBlocked, TargetNewDB,
    RequiresDestructiveAck, HobbyAgentActionPointsToHobbyPlus,
    CrossTenantBackupID_404
  - All prior backup tests updated to include destructive_acknowledgment
  - Contract test enforces the four new agent_action constants

DO NOT TOUCH list respected — no edits to email/, dpop.go, circuit/.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant