Skip to content

worker: churn predictor daily job flags 7d-inactive teams via audit_log - #15

Merged
mastermanas805 merged 1 commit into
masterfrom
feat/churn-prediction-job-fresh
May 13, 2026
Merged

mastermanas805 merged 1 commit into
masterfrom
feat/churn-prediction-job-fresh

Conversation

@mastermanas805

Copy link
Copy Markdown
Member

Summary

  • New ChurnPredictorWorker (River, daily 03:00 UTC) scans every non-Team team for churn-risk signal and writes one churn.risk_flagged audit_log row per qualifier
  • Single-query candidate scan (no N+1): plan_tier != 'team' AND MAX(activity) < now()-7d AND COUNT(active resources) > 0 AND NOT EXISTS churn.risk_flagged in last 30d
  • Existing event-email forwarder picks up churn.risk_flagged and triggers the we_miss_you Brevo template — event_email_mapping.go gains auditKindChurnRiskFlagged + buildChurnRiskFlagged so the forwarder invariants stay consistent
  • Activity-kind matching uses LIKE patterns (deploy%, vault.%, experiment.%) so the query is forward-compatible the day those audit producers wire up (see "Production gap" below)

Rationale

7-day inactivity threshold: dev workflows touch the platform at least weekly; shorter windows would noise-flag busy/vacationing customers, longer windows would miss drifting users before they fully detach.

30-day dedupe window: at most one "we miss you" email per team per month. Faster cadence feels spammy and re-fires the same copy to customers who've decided not to return; slower cadence lets a still-silent team go uncontacted month after month. 30d also aligns with monthly billing cycles, which is the standard B2B reactivation rate-limit.

Production gap (pushback)

The brief enumerated activity kinds as provision, deploy.*, vault.*, experiment.*, login. Grep of api/internal/handlers/ confirms only provision and experiment.conversion are written today (plus env_policy.updated). deploy.*, vault.*, and login are listed as intended kinds in migration 012_audit_log.sql but no producer emits them yet.

This means today's churn signal effectively reduces to "team hasn't provisioned anything or converted an experiment in 7+ days." Operators should be aware that "user hasn't logged in" cannot be detected until the auth path emits an audit row. The LIKE patterns in the SELECT make this a zero-line-change upgrade once those producers land.

The brief's example query also referenced users.is_primary = true. The users table has no is_primary column (migration 001); the worker uses the "oldest user on the team" convention from expire_imminent.go for consistency.

Test plan

  • go test ./... green — 89 tests, 0 failures
  • 9 hermetic sqlmock scenarios pass: flagged 8d team, 2d active team (no-op), no resources (no-op), flagged 15d ago in-dedupe (no-op), flagged 35d ago re-flagged, team-tier excluded (no-op), no-owner-email skip, top-level query error retries, per-row INSERT failure fail-open
  • TestEventEmail_AllSupportedKindsHaveBuilder confirms the new kind has a builder and vice-versa
  • go vet ./... clean
  • Operator wires BREVO_TEMPLATE_IDS={"churn.risk_flagged": <id>} before the Brevo forwarder will actually send (forwarder returns SkippedNoTemplate until then)

🤖 Generated with Claude Code

…t_log

Adds ChurnPredictorWorker — runs daily at 03:00 UTC, writes one
churn.risk_flagged audit_log row per qualifying team. The existing
event-email forwarder drains those rows into the "we_miss_you"
Brevo template (operator wires the template id).

Churn-risk SELECT in one query (no N+1): plan_tier != 'team',
MAX(audit_log.created_at for activity kinds) < now() - 7d,
COUNT(active resources) > 0, and NOT EXISTS a churn.risk_flagged
row in the last 30 days. Per-row INSERTs are fail-open so one bad
row never blocks the rest of the batch.

Activity-kind matching uses LIKE patterns (deploy%, vault.%,
experiment.%) so the query is forward-compatible the day those
producers start writing audit rows — only `provision` and
`experiment.conversion` actually fire today.

Also extends event_email_mapping.go with auditKindChurnRiskFlagged
and buildChurnRiskFlagged so the forwarder picks up the new kind.

Tests: 9 hermetic sqlmock cases cover the brief's scenario matrix
(flagged, dedupe in-window, dedupe expired, no resources, team-tier
excluded, no-email skip, fail-open insert, top-level error retry,
metadata shape).
@mastermanas805
mastermanas805 merged commit 772abbd into master May 13, 2026
mastermanas805 added a commit that referenced this pull request Jun 2, 2026
- orphan_sweep PASS 4: young-namespace (within grace) + age-lookup-error are
  NOT reaped (#8 grace branches).
- emitDeployFailedAudit: dedup-hit skips the INSERT (#15 idempotency branch).
- dbGracePeriodOpener.TerminateActiveGracePeriod: UPDATE success + error-wrap.
- billing terminal downgrade still succeeds when the grace-close errors (#5
  fail-open warn branch).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 2, 2026
…dup, cursor (#78)

* fix(worker): bug-bash batch 2 — grace close, namespace reaper grace, autopsy dedup, cursor

Four confirmed bugs from the 2026-06-02 platform bug bash:

- #5 (P1) billing_reconciler: a terminal Razorpay status downgraded the team
  but left the active payment_grace_periods row open, so payment_grace_reminder
  emitted dunning emails forever and the terminator later re-acted on an
  already-cancelled subscription. Add TerminateActiveGracePeriod to the
  gracePeriodOpener interface (status→'terminated', terminated_at=now()) and
  call it in the terminal-downgrade branch (fail-open).

- #8 (P1) orphan_sweep PASS 4: the customer-namespace reaper excluded 'pending'
  resources from the live-token set AND had no creation-grace, so a sweep
  during two-phase provisioning could DELETE a live, mid-provision namespace.
  Add 'pending' to fetchLiveResourceTokens and a namespace-age grace check
  (skip if younger than orphanNoDBRowGrace) mirroring PASS 3.

- #15 (P2) deploy_failure_autopsy: emitDeployFailedAudit inserted a new
  deploy.failed audit row (new id) on every reconciler retry, and the forwarder
  dedups by audit_id (not deployment) → duplicate failure emails. Make it
  idempotent: skip the INSERT when a deploy.failed row already exists for the
  deployment (metadata->>'deploy_id'). Fail-open on probe error.

- #18 (P2) billing_reconciler scanChargeUndeliverable: jumping the cursor to
  now() on an empty window skipped rows that became visible a moment later
  (clock skew / late commit). Leave the cursor unchanged on count==0 — the 1h
  look-back re-applies and re-scanning the small indexed window is cheap.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(worker): cover bug-bash batch-2 changed lines (100% patch gate)

- orphan_sweep PASS 4: young-namespace (within grace) + age-lookup-error are
  NOT reaped (#8 grace branches).
- emitDeployFailedAudit: dedup-hit skips the INSERT (#15 idempotency branch).
- dbGracePeriodOpener.TerminateActiveGracePeriod: UPDATE success + error-wrap.
- billing terminal downgrade still succeeds when the grace-close errors (#5
  fail-open warn branch).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 4, 2026
…icks (#84)

Finding #6 (SWEEP-BACKLOG-2026-06-04). The idempotency guard in
emitDeployFailedAudit (deploy_failure_autopsy.go, bug-bash #15) already
short-circuits the audit_log INSERT when a deploy.failed row exists for the
deployment, so the email forwarder can't fan out a duplicate failure email
per reconciler tick. This adds the missing positive unit coverage:

- TestEmitDeployFailedAudit_FirstTickInserts — no existing row → EXISTS probe
  false → exactly one INSERT.
- TestEmitDeployFailedAudit_SecondTickIsNoOp — existing row → EXISTS probe
  true → NO INSERT (a second INSERT would fail the ordered sqlmock).

No production code change — the guard was shipped 2026-06-02; this closes the
backlog item's "two ticks → exactly ONE deploy.failed row" test requirement.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant