Skip to content

api: truth pass — revenue mechanism + agent-action errors + env promotion - #21

Merged
mastermanas805 merged 2 commits into
masterfrom
truth-pass/api-2026-05-12
May 12, 2026
Merged

mastermanas805 merged 2 commits into
masterfrom
truth-pass/api-2026-05-12

Conversation

@mastermanas805

Copy link
Copy Markdown
Member

Summary

Rollup of multi-agent retro work on the api side, landing alongside instanode-web#30 (dashboard truth pass), worker, and mcp v0.8.0.

Revenue mechanism (pay-from-day-one)

  • Trial removed from onboarding.go::Claim — was giving away the product
  • Claim transfers team_id only; tier stays free until Razorpay webhook fires elevation
  • ElevateResourceTiersByTeam elevates both anonymous-TTL and free-tier resources, always clears expires_at, guards reaper race
  • Cancellation downgrades to free (not anonymous) for claimed teams
  • 5 new tests in resource_elevate_test.go

Tier model

  • free tier added as distinct real plan (same limits as anonymous; for claimed-but-unpaid users)
  • 001_initial.sql default plan_tier hobby → free
  • openapi.go tier enum includes free

Agent-friendly error responses (§10.15)

  • respondError now optionally emits agent_action + upgrade_url so the LLM literally knows what to tell the user
  • 15 error codes mapped (quota_exceeded, storage_limit_reached, vault_*, member_limit, upgrade_required, etc.)
  • Transient infra codes deliberately NOT mapped (agent retries)
  • Validation codes NOT mapped (message already names the field)

Env promotion as Pro-tier feature (§10.17, plumbed but waiting)

  • Migration 016_stack_env.sql adds stacks.env + parent_stack_id
  • POST /api/v1/stacks/:slug/promote — Pro/Team/Growth only, 402+agent_action on hobby/free/anonymous
  • POST /api/v1/vault/copy — same tier gate, supports dry_run + overwrite + key allowlist
  • 15 new tests
  • ⚠️ Not wired: redeploy hook depends on Phase-1 POST /deploy/new. Contract works, rows land, tier-gate enforced; production stack won't run new code until compute hook lands. "Plumbed but waiting."

Migrator removal

  • migratorclient/ + migrator_e2e_test.go deleted
  • billing.go::triggerMigrationsForTeam removed
  • Phase-2 dedicated-infra-per-tier story replaced by Phase-1 growth-tier-on-demand

Test infra fixes

  • testhelpers.go schema backfill (was pre-existing test pollution)
  • billing_test.go uses unique subscription IDs

Migrations

  • ✅ 015_resource_expiry_reminded.sql and 016_stack_env.sql already applied to the live managed PG before this lands. Safe to roll.

Test plan

  • make test-unit — all packages green
  • go build + go vet clean
  • After merge: docker build, push to ghcr, kubectl rollout instant-api in instant ns, verify https://api.instanode.dev/healthz

🤖 Generated with Claude Code

mastermanas805 and others added 2 commits May 11, 2026 22:36
Captures the full anonymous → claim → upgrade → cancel funnel in a single
runnable test. Codifies the contract for /claim's session_token, /whoami's
tier+email enrichment, /api/v1/billing's trial/active/cancelled status,
and ElevateResourceTiersByTeam's promote-on-upgrade behaviour.

Adds three tests:

  - TestE2E_FullCustomerFlow_AnonymousToProToCancelled — 11-step happy
    path. Provisions /db/new anonymously, claims with email, uses the
    returned session_token to hit /whoami, /billing, /resources, fires
    a subscription.charged webhook, asserts tier flips to pro on all
    surfaces and existing resources are elevated, fires
    subscription.cancelled, asserts downgrade to hobby and that
    existing resources keep tier=pro (documented snapshot behaviour).

  - TestE2E_FullCustomerFlow_WhoamiBeforeClaim — regression guard that
    the anonymous upgrade_jwt cannot auth against /api/v1/whoami; must
    return 401.

  - TestE2E_FullCustomerFlow_StoragePathReturnsSpacesCreds — post-Spaces
    switch sanity check that /storage/new returns S3-compatible creds
    and the public endpoint does not leak the in-cluster MinIO hostname.

All three skip cleanly when the required secrets are absent
(E2E_JWT_SECRET, E2E_RAZORPAY_WEBHOOK_SECRET) and pass against the live
local k8s cluster (verified: 26s wall-clock, 3/3 PASS).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…tion

Rollup of multi-agent retro work landing across the api on 2026-05-12.
Shipped together because the dashboard truth-pass (instanode-web#30) +
worker expiry-reminder + MCP v0.8.0 all interlock with these changes.

Revenue mechanism (pay-from-day-one):
  - onboarding.go::Claim() no longer calls StartTrial. The 14-day trial
    was giving away the product per the no-trial memory rule.
  - Claim only transfers team_id ownership; tier stays 'free' (new tier
    keeping limits identical to anonymous) until subscription.charged
    webhook fires ElevateResourceTiersByTeam.
  - models.ElevateResourceTiersByTeam now elevates BOTH anonymous-TTL
    and free-tier resources (not just already-permanent ones), always
    clears expires_at, guards a reaper-race via expires_at > now().
  - billing.handleSubscriptionCancelled downgrades to 'free' (not
    'anonymous') for zero-paid teams. Claimed teams cannot drop back.
  - 5 new resource_elevate_test.go cases lock the elevation contract.

Tier model:
  - plans.yaml adds 'free' as a real plan with limits mirroring
    anonymous (pre-claim users get anonymous; claimed-but-unpaid get
    free; both expire at 24h).
  - 001_initial.sql default plan_tier changed from 'hobby' to 'free'.
  - openapi.go tier enum extended with 'free'.

Agent-friendly error responses (§10.15):
  - respondError extended to optionally emit `agent_action` (copy the
    agent can show the user) + `upgrade_url`.
  - 15 error codes mapped: quota_exceeded, storage_limit_reached,
    vault_quota_exceeded, vault_not_available, vault_env_not_allowed,
    member_limit, upgrade_required, tier_unavailable, rate_limit_exceeded,
    unauthorized, auth_required, invalid_token, missing_token,
    vault_requires_auth, invitation_invalid, already_accepted,
    already_claimed, webhook_inactive, resource_not_found, forbidden,
    last_owner.
  - Transient infra codes deliberately NOT mapped (agent should retry).
  - Validation codes NOT mapped (agent-side request-construction bugs;
    `message` already names the field).
  - middleware/quota.go::PaymentRequired emits agent_action+upgrade_url.
  - 7 new agent_action_test.go cases + quota+openapi guards.

Env promotion as Pro-tier feature (§10.17, plumbed but waiting):
  - Migration 016_stack_env.sql adds stacks.env + parent_stack_id.
  - POST /api/v1/stacks/:slug/promote — Pro/Team/Growth only, returns
    402 with agent_action on hobby/free/anonymous; creates a sibling
    stack with parent_stack_id linking back to the family.
  - POST /api/v1/vault/copy — same tier gate, supports dry_run +
    overwrite + key allowlist.
  - 15 new tests (stack_promote_test.go, vault_copy_test.go,
    openapi multi-env documentation guard).
  - NOT WIRED: actual redeploy hook depends on Phase-1 POST /deploy/new
    which is on the roadmap. Contract works, rows land correctly,
    tier-gate enforced; production stack won't run new code until the
    compute hook lands. "Plumbed but waiting."

Migrator removal (audit recommended drop):
  - migratorclient/ + migrator_e2e_test.go deleted.
  - billing.go::triggerMigrationsForTeam removed.
  - dev.go migClient param removed.
  - router.go init removed.
  - config.go MigratorAddr + MigratorSecret fields removed.
  - The Phase-2 dedicated-infra-per-tier story is replaced by Phase-1
    growth-tier-on-demand provisioning; migrator was phantom architecture.

Secret hygiene:
  - testhelpers.go schema backfill (app_id, env, key_prefix,
    vault_secrets, vault_audit_log, etc.) fixes pre-existing test
    pollution that blocked the test gate.
  - billing_test.go uses unique subscription IDs to avoid
    teams_stripe_customer_id_key unique-constraint races.
  - storage.go + vault.go + stack.go use respondErrorWithAgentAction
    for tier-aware overrides.

Test gate:
  - make test-unit: all packages green
  - go build + go vet: clean
  - 15+ new test cases across handlers/middleware/models

Migrations 015 + 016 already applied to the live managed PG before
this commit lands. The api code is safe to roll.

Reference: /Users/manassrivastava/Documents/InstaNode/RETRO-2026-05-12.md

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@mastermanas805
mastermanas805 merged commit 477e3c2 into master May 12, 2026
@mastermanas805
mastermanas805 deleted the truth-pass/api-2026-05-12 branch May 12, 2026 02:14
mastermanas805 added a commit that referenced this pull request Jun 3, 2026
…te race (bug-bash #21)

The middleware did GET(miss) → run handler → SET with no atomic reservation
between the GET and the SET. Two requests carrying the same Idempotency-Key
(or the same body-fingerprint) racing in that window both saw redis.Nil and
both ran the handler — double-creating real backend resources on the
authenticated /db|/cache|/nosql provision paths, which have no other
per-burst dedup gate (the per-fingerprint INCR cap lives only on the anon
branch; CreateDeployment already has its own FOR UPDATE cap).

Fix: on a cache miss, write an in-flight reservation marker via SETNX
(idemEntry.InFlight, 60s TTL) the instant the handler starts. A concurrent
same-key request reads the marker in the top GET hit-block and returns 409
idempotency_key_in_progress (retryable — same envelope as the existing
conflict 409) instead of re-running the handler. The real response overwrites
the marker on completion (plain SET); non-cacheable paths (5xx, handler
error, mutable 4xx) DELETE the marker so an immediate legitimate retry isn't
blocked. Applied symmetrically to both the explicit-key and fingerprint
branches. SETNX is best-effort/fail-open — a Redis hiccup never blocks
provisioning (documented dedup posture, never a correctness gate).

Tests: TestIdempotency_{ExplicitKey,Fingerprint}_ConcurrentInFlight_Returns409
hold request A in the handler (blocked on a channel) so B observes A's live
reservation → 409, handler runs exactly once. Pass under -race; new code 100%
covered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 3, 2026
…te race (bug-bash #21) (#232)

The middleware did GET(miss) → run handler → SET with no atomic reservation
between the GET and the SET. Two requests carrying the same Idempotency-Key
(or the same body-fingerprint) racing in that window both saw redis.Nil and
both ran the handler — double-creating real backend resources on the
authenticated /db|/cache|/nosql provision paths, which have no other
per-burst dedup gate (the per-fingerprint INCR cap lives only on the anon
branch; CreateDeployment already has its own FOR UPDATE cap).

Fix: on a cache miss, write an in-flight reservation marker via SETNX
(idemEntry.InFlight, 60s TTL) the instant the handler starts. A concurrent
same-key request reads the marker in the top GET hit-block and returns 409
idempotency_key_in_progress (retryable — same envelope as the existing
conflict 409) instead of re-running the handler. The real response overwrites
the marker on completion (plain SET); non-cacheable paths (5xx, handler
error, mutable 4xx) DELETE the marker so an immediate legitimate retry isn't
blocked. Applied symmetrically to both the explicit-key and fingerprint
branches. SETNX is best-effort/fail-open — a Redis hiccup never blocks
provisioning (documented dedup posture, never a correctness gate).

Tests: TestIdempotency_{ExplicitKey,Fingerprint}_ConcurrentInFlight_Returns409
hold request A in the handler (blocked on a channel) so B observes A's live
reservation → 409, handler runs exactly once. Pass under -race; new code 100%
covered.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant