worker: churn predictor daily job flags 7d-inactive teams via audit_log - #15
Merged
Merged
Conversation
…t_log Adds ChurnPredictorWorker — runs daily at 03:00 UTC, writes one churn.risk_flagged audit_log row per qualifying team. The existing event-email forwarder drains those rows into the "we_miss_you" Brevo template (operator wires the template id). Churn-risk SELECT in one query (no N+1): plan_tier != 'team', MAX(audit_log.created_at for activity kinds) < now() - 7d, COUNT(active resources) > 0, and NOT EXISTS a churn.risk_flagged row in the last 30 days. Per-row INSERTs are fail-open so one bad row never blocks the rest of the batch. Activity-kind matching uses LIKE patterns (deploy%, vault.%, experiment.%) so the query is forward-compatible the day those producers start writing audit rows — only `provision` and `experiment.conversion` actually fire today. Also extends event_email_mapping.go with auditKindChurnRiskFlagged and buildChurnRiskFlagged so the forwarder picks up the new kind. Tests: 9 hermetic sqlmock cases cover the brief's scenario matrix (flagged, dedupe in-window, dedupe expired, no resources, team-tier excluded, no-email skip, fail-open insert, top-level error retry, metadata shape).
mastermanas805
added a commit
that referenced
this pull request
Jun 2, 2026
- orphan_sweep PASS 4: young-namespace (within grace) + age-lookup-error are NOT reaped (#8 grace branches). - emitDeployFailedAudit: dedup-hit skips the INSERT (#15 idempotency branch). - dbGracePeriodOpener.TerminateActiveGracePeriod: UPDATE success + error-wrap. - billing terminal downgrade still succeeds when the grace-close errors (#5 fail-open warn branch). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805
added a commit
that referenced
this pull request
Jun 2, 2026
…dup, cursor (#78) * fix(worker): bug-bash batch 2 — grace close, namespace reaper grace, autopsy dedup, cursor Four confirmed bugs from the 2026-06-02 platform bug bash: - #5 (P1) billing_reconciler: a terminal Razorpay status downgraded the team but left the active payment_grace_periods row open, so payment_grace_reminder emitted dunning emails forever and the terminator later re-acted on an already-cancelled subscription. Add TerminateActiveGracePeriod to the gracePeriodOpener interface (status→'terminated', terminated_at=now()) and call it in the terminal-downgrade branch (fail-open). - #8 (P1) orphan_sweep PASS 4: the customer-namespace reaper excluded 'pending' resources from the live-token set AND had no creation-grace, so a sweep during two-phase provisioning could DELETE a live, mid-provision namespace. Add 'pending' to fetchLiveResourceTokens and a namespace-age grace check (skip if younger than orphanNoDBRowGrace) mirroring PASS 3. - #15 (P2) deploy_failure_autopsy: emitDeployFailedAudit inserted a new deploy.failed audit row (new id) on every reconciler retry, and the forwarder dedups by audit_id (not deployment) → duplicate failure emails. Make it idempotent: skip the INSERT when a deploy.failed row already exists for the deployment (metadata->>'deploy_id'). Fail-open on probe error. - #18 (P2) billing_reconciler scanChargeUndeliverable: jumping the cursor to now() on an empty window skipped rows that became visible a moment later (clock skew / late commit). Leave the cursor unchanged on count==0 — the 1h look-back re-applies and re-scanning the small indexed window is cheap. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(worker): cover bug-bash batch-2 changed lines (100% patch gate) - orphan_sweep PASS 4: young-namespace (within grace) + age-lookup-error are NOT reaped (#8 grace branches). - emitDeployFailedAudit: dedup-hit skips the INSERT (#15 idempotency branch). - dbGracePeriodOpener.TerminateActiveGracePeriod: UPDATE success + error-wrap. - billing terminal downgrade still succeeds when the grace-close errors (#5 fail-open warn branch). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805
added a commit
that referenced
this pull request
Jun 4, 2026
…icks (#84) Finding #6 (SWEEP-BACKLOG-2026-06-04). The idempotency guard in emitDeployFailedAudit (deploy_failure_autopsy.go, bug-bash #15) already short-circuits the audit_log INSERT when a deploy.failed row exists for the deployment, so the email forwarder can't fan out a duplicate failure email per reconciler tick. This adds the missing positive unit coverage: - TestEmitDeployFailedAudit_FirstTickInserts — no existing row → EXISTS probe false → exactly one INSERT. - TestEmitDeployFailedAudit_SecondTickIsNoOp — existing row → EXISTS probe true → NO INSERT (a second INSERT would fail the ordered sqlmock). No production code change — the guard was shipped 2026-06-02; this closes the backlog item's "two ticks → exactly ONE deploy.failed row" test requirement. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
ChurnPredictorWorker(River, daily 03:00 UTC) scans every non-Team team for churn-risk signal and writes onechurn.risk_flaggedaudit_log row per qualifierplan_tier != 'team'ANDMAX(activity) < now()-7dANDCOUNT(active resources) > 0ANDNOT EXISTS churn.risk_flagged in last 30dchurn.risk_flaggedand triggers thewe_miss_youBrevo template —event_email_mapping.gogainsauditKindChurnRiskFlagged+buildChurnRiskFlaggedso the forwarder invariants stay consistentdeploy%,vault.%,experiment.%) so the query is forward-compatible the day those audit producers wire up (see "Production gap" below)Rationale
7-day inactivity threshold: dev workflows touch the platform at least weekly; shorter windows would noise-flag busy/vacationing customers, longer windows would miss drifting users before they fully detach.
30-day dedupe window: at most one "we miss you" email per team per month. Faster cadence feels spammy and re-fires the same copy to customers who've decided not to return; slower cadence lets a still-silent team go uncontacted month after month. 30d also aligns with monthly billing cycles, which is the standard B2B reactivation rate-limit.
Production gap (pushback)
The brief enumerated activity kinds as
provision, deploy.*, vault.*, experiment.*, login. Grep ofapi/internal/handlers/confirms onlyprovisionandexperiment.conversionare written today (plusenv_policy.updated).deploy.*,vault.*, andloginare listed as intended kinds in migration012_audit_log.sqlbut no producer emits them yet.This means today's churn signal effectively reduces to "team hasn't provisioned anything or converted an experiment in 7+ days." Operators should be aware that "user hasn't logged in" cannot be detected until the auth path emits an audit row. The LIKE patterns in the SELECT make this a zero-line-change upgrade once those producers land.
The brief's example query also referenced
users.is_primary = true. Theuserstable has nois_primarycolumn (migration 001); the worker uses the "oldest user on the team" convention fromexpire_imminent.gofor consistency.Test plan
go test ./...green — 89 tests, 0 failuresTestEventEmail_AllSupportedKindsHaveBuilderconfirms the new kind has a builder and vice-versago vet ./...cleanBREVO_TEMPLATE_IDS={"churn.risk_flagged": <id>}before the Brevo forwarder will actually send (forwarder returnsSkippedNoTemplateuntil then)🤖 Generated with Claude Code