Skip to content

fix(compute): explicit kaniko resources + imagePullPolicy Always - #2

Merged
mastermanas805 merged 1 commit into
masterfrom
fix/deploy-compute-correctness
May 11, 2026
Merged

mastermanas805 merged 1 commit into
masterfrom
fix/deploy-compute-correctness

Conversation

@mastermanas805

Copy link
Copy Markdown
Member

Summary

Two friction-point fixes that an unaided AI agent hits when redeploying via `/deploy/new`:

1. Kaniko build pod was throttled to 50m CPU / 256Mi RAM

The per-namespace `LimitRange` (hobby default) was the only thing sizing the kaniko container — non-trivial `npm install` hung 5+ minutes. Sets explicit `Requests=250m/256Mi` and `Limits=1/512Mi` on the kaniko container. Build runs as a Job, so the bigger allocation doesn't inflate the app's permanent footprint.

2. `imagePullPolicy: IfNotPresent` + `:latest` tag → silent stale redeploys

When kaniko pushed a new image to GHCR under the same `:latest` tag, k8s served the cached old layer because the tag string was unchanged. Redeploys appeared to succeed while pods kept the prior code. Switched to `PullAlways` everywhere (single-app + stack).

3. `K8sProvider.clientset` is now `kubernetes.Interface`

Drop-in change — `*Clientset` already implements `Interface`. Enables tests to inject `fake.NewSimpleClientset`.

Test plan

  • `TestKanikoJobHasExplicitResources`: asserts the kaniko container has Requests and Limits set, with the CPU request at the 250m floor
  • `TestAppDeploymentUsesPullAlways`: asserts `ImagePullPolicy == PullAlways` on the app container
  • `go build ./...` passes
  • Follow-up: sha-pin image tags so we can switch back to `PullIfNotPresent` and save bandwidth

🤖 Generated with Claude Code

Two related fixes that an agent hits when redeploying via /deploy/new:

1. Kaniko build pod was sized by the per-namespace LimitRange (hobby
   tier default: 50m CPU / 256Mi RAM). Non-trivial npm install hung
   5+ minutes. Now sets explicit Requests=250m/256Mi and
   Limits=1/512Mi on the kaniko container. Build runs as a Job, so
   the higher allocation does not inflate the app's permanent quota.

2. App container had imagePullPolicy=IfNotPresent on the :latest tag.
   After a redeploy pushed a new image to GHCR, k8s served the cached
   old layer because the tag string was unchanged. Redeploys silently
   succeeded while pods kept the prior code. Switched to PullAlways.

Also makes the K8sProvider.clientset field a kubernetes.Interface so
tests can inject fake.NewSimpleClientset.

Tests cover the resource floor + PullAlways guarantee.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@mastermanas805
mastermanas805 merged commit 082f9fe into master May 11, 2026
@mastermanas805
mastermanas805 deleted the fix/deploy-compute-correctness branch May 11, 2026 08:07
mastermanas805 added a commit that referenced this pull request May 11, 2026
…s package (#12)

User-facing pain that drove this: every domain rename or proxy hostname
tweak required scraping the codebase with grep + sed. The instant.dev →
instanode.dev rename earlier today touched 28 files and still missed
some (we're still finding stragglers). The "Use named constants, not
inline strings" memo applies — extracting was overdue.

New package internal/urls/urls.go centralises:

  Public hostnames:
    PublicAPIBase        = "https://api.instanode.dev"
    PublicMarketingBase  = "https://instanode.dev"
    StartURLPrefix       = PublicMarketingBase + "/start"
    DeploymentWildcard   = "deployment.instanode.dev"
    StoragePublicHost    = "s3.instanode.dev"

  Cluster-internal proxy FQDNs (for in-cluster workloads — friction PR #2):
    InternalPGProxy      = "instant-pg-proxy.instant.svc.cluster.local:5432"
    InternalRedisProxy   = "instant-redis-proxy.instant.svc.cluster.local:6379"
    InternalMongoProxy   = "instant-mongo-proxy.instant.svc.cluster.local:27017"
    InternalNATSProxy    = "instant-nats-proxy.instant.svc.cluster.local:4222"
    InternalMinIO        = "minio.instant-data.svc.cluster.local:9000"

  Helper:
    UpgradeStartURL(jwt) — canonical builder for /start?t=<jwt>, replaces
                          12 sites that had identical fmt.Sprintf calls.

Call sites refactored:
  - 12 fmt.Sprintf upgrade-URL calls across 6 provisioning handlers
    → urls.UpgradeStartURL(jwtToken)
  - proxiedInternalURL in handlers/internal_url.go → uses 4 urls.Internal*
  - upgradeNote / limitExceededNote bare links → urls.StartURLPrefix
  - canonicalAPIBase (handlers/auth.go) → alias of urls.PublicAPIBase
  - defaultCanonicalResourceURL (middleware/auth.go) → alias of
    urls.PublicAPIBase  (eliminates the long-standing duplication of the
    same string in two different files)
  - 15 user-facing error strings ("Sign up at https://instanode.dev/start")
    → string concat with urls.StartURLPrefix

Audit results:
  Before:  28+ occurrences of "instanode.dev/start" in non-test handler code
  After:   1 occurrence (an inline example URL inside the OpenAPI JSON
           description string — appropriate to leave as a literal since it's
           sample documentation, not a runtime-constructed URL)

  Before:  4 proxy FQDN string literals in handlers/internal_url.go
  After:   0 (all 4 use urls.* constants)

Tests:
  internal/urls/urls_test.go (12 sub-tests):
    TestPublicHostnames_MatchExpectedShape         — guards 5 public hosts
    TestInternalProxyHostnames_CorrectPortsAndService — guards 5 internal
    TestUpgradeStartURL_Composition                — builder semantics

  Existing tests still pass:
    TestProxiedInternalURL          — 7 sub-cases unchanged after refactor
    TestUpgradeNote_DoesNotMentionTrial / TestLimitExceededNote_… — 6 cases
    TestWhoami_*, TestDeployNew_EnvVars_*, all TestOpenAPI_* — pass

Live verification on v2.1.0-urls-package in prod:
  POST /db/new (fresh) → upgrade host = "instanode.dev"
                         internal_url host = "instant-pg-proxy.instant.svc.cluster.local:5432"
                         note starts with "Works for 24h free"
  GET /api/v1/whoami → 200 with team_id + plan_tier + team_name
                       (proves the middleware audience lookup still works
                        via the urls.PublicAPIBase alias)

Out of scope (separate follow-ups):
  - Email templates in internal/email/email.go (45 occurrences) — these
    are marketing copy strings with their own care; need different
    extraction approach (probably a templated config).
  - SDK references (sdk-go, sdk-node, mcp) — each is its own repo;
    consolidating across modules needs a shared common/urls package and
    cross-repo coordination.
  - Worker / provisioner repos — they have far fewer URL string
    instances and can do their own thing.
  - The OpenAPI bearerAuth description's inline example URL — sample
    documentation, not a runtime URL.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request May 16, 2026
…tract bugs

The api Deploy workflow had been red on every run. Root causes, all fixed:

Schema-mirror drift (the bulk of the ~25 failures)
  testhelpers.runMigrations hand-mirrors the prod schema; CI ran against a
  bare Postgres so any migration that added a table/column without a mirror
  edit silently broke the gate. Completed the mirror (email_events,
  pending_deletions, deployment_events, api_keys, magic_links, custom_domains,
  service_components, uptime_samples + 9 drifted columns) and added
  TestRunMigrationsMirrorsEveryMigrationTable — a guard that enumerates every
  CREATE TABLE in the migration files and fails if one is unmirrored.
  deploy.yml now also applies the real migration files before testing, so CI
  runs the same schema developers do (kills column drift too).

Genuine product/test bugs the lax mirror had been masking:
  - rate_limit.go / idempotency.go: a nil *redis.Client SIGSEGV'd the whole
    API on the first request. Both now fail open (CLAUDE.md convention #1).
  - cache/redis.go provisionLocal/StorageBytes: nil Redis client now returns
    a clean error → handler 503, never a panic (convention #2).
  - models.AcceptInvitation: did not clear is_primary when moving an invitee
    into a team, violating uq_users_one_primary_per_team → 500. Now cleared.
  - resource_elevate_test.go inserted stack/deployment rows missing the real
    NOT NULL columns (namespace, app_id); now valid.
  - Fixed 5 handler tests with stale tier/seat/mock assumptions and dropped
    the stale -skip list (TestOpenAPI/TestCrossTeam/TestCustomDomainCreate
    all pass now).

go test ./... -short: 25 packages, 0 FAIL, 0 panic. build + vet clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request May 20, 2026
…t no-op

Two contract-drift P0s flagged in the bug-hunt brief, both shipped in the
same PR because they share the test gate.

B13-P0-F1: POST /auth/cli auth_url leaked dead-brand
  * Old: frontendURL() returned "https://instant.dev" in production —
    instant.dev is the dead-brand marketing host (404s). An agent
    following the auth_url landed on a parking page and gave up.
  * Fix: read cfg.DashboardBaseURL (DASHBOARD_BASE_URL env var); fall
    back to "https://instanode.dev" when DASHBOARD_BASE_URL is unset
    in production. internal/handlers/cli_auth.go.
  * Coverage: TestAuth_CLI_ReturnsInstanodeDomain (4 cases —
    explicit base, empty-fallback, dev default, trailing-slash strip)
    + TestAuth_CLI_NoLegacyInstantDevInResponses (rule-17 source-scan
    that fails if a new handler emits instant.dev/<x> anywhere).
  * Also cleaned: internal/handlers/logs.go (error message URL),
    internal/handlers/openapi.go (storage endpoint example).

B7-P0-1: PATCH /stacks/:slug/env was a silent no-op
  * Old: handler logged stack.env.noted, returned 200, persisted
    nothing — the next /stacks/:slug/redeploy rebuilt with stale env.
    Silent-data-loss failure mode.
  * Fix (Option A from the brief): migration 062 adds
    stacks.env_vars JSONB DEFAULT '{}'::jsonb. Handler now loads
    existing env, merges (empty-string value deletes), validates
    every key against isValidEnvKey (POSIX [A-Z_][A-Z0-9_]*),
    persists, emits stack.env.updated audit_log row, and returns
    the full merged map.
  * 64KiB cap on serialized payload (ErrStackEnvVarsTooLarge → 413).
  * OpenAPI updated to describe PATCH semantics + new 409/413
    responses + the merged-env response shape.
  * Coverage: TestStack_PatchEnv_PersistsAndReturns (7 subtests —
    auth, first patch, merge, delete-via-empty, invalid_env_key,
    missing_env, 404 unknown slug). All assertions verify the DB
    round-trip — pre-fix the handler returned 200 but persisted
    nothing, which is exactly what subtest #2 now fails on.

Coverage block per CLAUDE.md rule 17:
  Symptom:        instant.dev/<x> in handler-emitted strings
  Enumeration:    rg -nF "instant.dev/" internal/handlers/ --type go
  Sites found:    3 (cli_auth.go:339, logs.go:136, openapi.go:2784)
  Sites touched:  3 — all three migrated to instanode.dev
  Coverage test:  TestAuth_CLI_NoLegacyInstantDevInResponses scans
                  every .go file under internal/handlers/ and fails
                  if any non-test, non-comment, non-label,
                  non-import-path mention of instant.dev/<x> appears.
  Live verified:  pending — see follow-up commit after deploy

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request May 23, 2026
…ward 95%

Add 12 _final2 test files covering the biggest reachable uncovered chunks in
the api internal/handlers package, lifting the package from 93.12% → 93.9%
under CI conditions (pgvector/redis/mongo/nats, migrations applied,
go test ./internal/handlers/... -short -count=1 -p 1).

Covered arms (all reachable; no waivers):
- auth.go: GitHub/Google find-or-create LinkID errors, email-lookup DB errors
  (fault DB), Google empty-name team-name fallback. (OAuth flows driven via the
  existing SetOAuthURLsForTest + startFakeOAuth seams.)
- sns_verify.go: full snsVerifier.verify matrix (guards, cert-fetch error,
  sig decode, SignatureVersion 1/unknown/2, happy RSA verify) + getCert
  cache hit/miss/error via a generated throwaway RSA cert + fetchCert seam.
- internal_terminate.go: dunning/downgrade DB-error arms (staged openFaultDB),
  razorpay canceler-not-configured / cancel-error / happy paths.
- internal_resend_magic_link.go: lookup/update-hash DB errors + mark-failed /
  attempts-lookup / mark-abandoned best-effort error arms (fault DB + fake mailer).
- deploy_ttl.go: MakePermanent/SetTTL update_failed + refresh_failed +
  lookupDeployment fetch_failed (seeded deploy + staged fault DB).
- deploy_teardown_reconciler.go: begin_tx_failed, list_failed, empty-tx commit,
  mark_failed (fault DB + existing fakeTeardownProvider).
- env_policy.go: Get fetch_failed, Put role_lookup_failed / owner_required /
  invalid_body / invalid_env_policy / persist_failed / happy.
- billing.go (CreateCheckoutAPI): CreateSubscription circuit-open / razorpay-error
  / incomplete-response / persistence-failure (fake CreateSubscription seam,
  HMAC-signed payloads only, never a real key).
- stack.go: UpdateEnv fetch/persist DB errors + Family fetch_failed (seeded
  stack + staged fault DB).
- audit.go: bad-team-id unauthorized + List/CSV happy paths with metadata +
  actor-email lookup.
- provision_helper.go denyProvisionOverCap (was 0%) + per-handler over-cap
  (no-existing → 429; with-existing → dedup-return) + backend-failure
  soft-delete + finalizeProvision-failure persist arms across
  db/cache/nosql/vector/queue/storage/webhook (anon + authenticated).

No new exported test symbols were required (all arms reachable via existing
export_*_test.go seams or white-box package handlers tests), so no
export_final2_test.go was added.

Provably-unreachable arms left uncovered (documented): defaultFetchCert
(real network), generateOAuthState/issueSessionJWT failures (crypto/rand +
HS256 []byte key cannot fail), the k8s ComputeProvider constructor branch
(needs a live cluster; noop fallback is covered).

Known pre-existing shared-dev-DB flakes unrelated to this change: TestAdminList_*
(limit=100 paging on accumulated data) and TestQueue_* (NATS :8222 monitor not
reachable locally) — verify on a fresh DB per CLAUDE.md.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request May 23, 2026
…copy list-failed

- family_bulk_twin happy path: seed active postgres+redis+mongodb parents in
  one env and POST /families/bulk-twin to another with working local backends,
  exercising twinOneParent + ProvisionForTwinCore success for all three
  backends (db/cache/nosql 81.8%→89.3%, twinOneParent →82.6%).
- vault.go CopySecrets list_failed arm: rename vault_secrets away after the
  team/tier checks pass (isolated DB) so ListVaultSecretKeys errors → 5xx.

Package 93.9% → 93.93% under CI conditions; only the documented pre-existing
shared-dev-DB flakes remain (TestAdminList_* paging, TestQueue_* NATS :8222).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request May 23, 2026
…mp (#163)

* test(coverage): api handlers provisioning-arms slice → ≥95% (#161)

Covers db/cache/nosql/queue/queue_provider/storage/storage_presign/
provision_helper/family_bulk_twin handler files via bufconn fake
provisioner, sqlmock, miniredis, and package-var seams. New exports
live in export_provarms_test.go (not the shared export_test.go).

Co-authored-by: Claude <noreply@anthropic.com>

* chore(deps): bump x/crypto 0.50->0.52, x/net 0.53->0.55 (OSV)

Clears OSV-Scanner/govulncheck (GO-2026-5005..5033 crypto, 5025..5030 net).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): api handlers residual files + middleware → ≥95% (#162)

Raise the handler files the prior slice left below 95% (resource, webhook,
onboarding, admin_customers, admin_impersonate, admin_promos_audit,
admin_customer_notes, billing) plus internal/middleware to ≥95%.

Seams added (production behavior unchanged — vars default to the real fn):
- admin_impersonate.go: signImpersonationToken var (drives sign_failed 503).
- webhook.go: cryptoEncrypt var (drives storeEncryptedURL encrypt-fail).

Techniques: brokenDB closed-pool for db_failed arms; DATA-DOG/go-sqlmock for
mid-sequence failures (mark-converted / team-create / user-create, soft-delete,
List scan/rows-err, Detail per-query failures); bufconn fakeProvisioner for the
Delete gRPC-deprovision arm; a dead MinIO endpoint for the storage-deprovision
warn arm; dead-Redis clients for fail-open arms; real onboarding JWTs minted
in-process for claim/preview/start branches; the cov2 webhook harness for the
payment-failed + charged-receipt arms.

All new test files carry a _residual suffix; new re-exports live in
export_residual_test.go (no edits to the shared export_test.go /
export_rbw_test.go / export_billing_test.go).

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): wire 0%-under-CI handlers hermetically (api_keys/logs/internal_*/custom_domain/storage)

Recovers handler coverage that was 0% under CI's postgres-only matrix because
the routes were k8s/storage-gated and never wired into a test app:
- api_keys, internal_resend_magic_link, internal_backup_refund, magic_link_circuit
  ctors: pure DB + fake-mailer, now driven end-to-end.
- logs.go: typed clientset field against kubernetes.Interface + SetClientset
  seam so a k8s fake clientset exercises the tier/status/pod-list/stream arms.
- custom_domain.go: stubbed CustomDomainProvider + real DB drives
  Create/List/Verify(ingress+cert+fail)/Delete.
- storage.go + storage_presign.go: real do-spaces-backed provider (zero network
  at provision/presign time) drives broker-mode + authenticated + presign arms.

internal/handlers under CI conditions: 80.9% -> 84.6%.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): recover queue + vector + audit + backup arms under CI

- coverage.yml: add a NATS service (8222 monitor) so /queue/new's NATS health
  check passes — without it every queue provision 503s and the handler's
  provision + per-tenant-credential arms never run under CI. Switch the
  postgres service to pgvector/pgvector:pg16 (drop-in superset) so /vector/new's
  CREATE EXTENSION vector succeeds instead of skipping (~44% -> covered).
- queue.go: add SetCredProvider seam; new test injects a fake
  QueueCredentialProvider to cover the AuthMode=isolated success + creds-error
  fallback arms (legacy_open provider can't reach them).
- audit.go: filter/tier/CSV/masked-email arms via query-param matrix + all tiers.
- backup.go: parseListCursor/parseIntStrict arms via bad/huge ?limit + ?before.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): deletion_confirm pure helpers + redirect handler

White-box coverage for confirmationLinkBase / buildConfirmationLink / teamIsPaid
/ deletionAuditKind* / deletionAuditResourceType / EmailConfirmDeletionRedirect-
Handler — pure string/URL composition + a no-side-effect 302 that the
deploy/stack delete-confirm flow only reaches when provisioning is available.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): github_deploy Get/Disconnect + vault GetSecret/List/Delete arms

- github_deploy.go: Get (connected/not-connected/not-found/cross-team) and
  Disconnect (happy/idempotent/not-found/cross-team) — existing tests only
  covered Connect + Receive.
- vault.go: GetSecret version-fetch + validation arms, ListKeys, DeleteSecret
  arms via the existing vaultTestApp + makeTeamUser helpers.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): twin approval gate + promote-approval rate-limit/HTML + family-binding helpers

- twin.go: beginTwinApproval (202 pending-approval, hermetic — short-circuits
  before provisioning) + consumeApprovedTwin invalid/not-found arms (were 0%).
- promote_approval.go: checkApproveRateLimit (real redis, under + over budget) +
  the 5 approval HTML helpers (white-box).
- family_bindings.go: BindingError.Error, mapBindingError every kind + default +
  empty-key label, nameOrType, nameOrEmpty (white-box pure helpers).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): webhook authed provision + vector authed provision + brevo webhook arms

- webhook.go: newWebhookAuthenticated (team-owned provision + receive verbs +
  list-requests) — anonymous-path tests didn't reach it.
- vector.go: newVectorAuthenticated (pro-team pgvector provision) — now reachable
  with the pgvector CI image; skips cleanly if backend unavailable.
- email_webhooks.go: Brevo invalid-payload / missing-email-skip / insert-fail
  fail-open arms via sqlmock.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): stack delete-confirm/cancel/UpdateEnv arms + validateHostname + SNS pure helpers

- stack.go: Delete email-confirm 202 + Cancel + immediate-skip-header + not-found
  + cross-team; UpdateEnv not-found/invalid-body/missing-env/invalid-key/merge+
  delete/deleting-409 (custom app wires the noop email client + confirm routes).
- custom_domain.go: validateHostname every rejection arm (white-box).
- sns_verify.go: parseSNSCertPEM / buildSNSSigningString / snsFieldValue
  white-box (the network defaultFetchCert path stays out by design).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): stack New validation/cap arms + vector anonymous-dedup path

- stack.go New: missing-manifest / invalid-manifest / missing-tarball / per-tier
  deployment-cap 402 (noop provider).
- vector.go: anonymous dedup path (vectorAnonymousLimits + decryptConnectionURL
  on the limit-exceeded branch) after exceeding the per-fingerprint cap.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): internal_terminate auth arms + vault copy validation + audit pure helpers

- internal_terminate.go: empty-token / bad-purpose / missing-iat / future-iat-skew
  verify arms (existing test covered wrong-secret/expired/mismatch/missing-bearer).
- vault.go CopySecrets: missing/invalid from-to env, same-env, invalid-key-in-
  allowlist, invalid-body validation arms.
- audit.go: maskEmail all 4 arms + tierLookbackDays blocked/unlimited/bounded
  (white-box).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): github webhook Receive branch/event filter arms

Covers Receive's non-push-event ignore, branch-mismatch ignore, zero-SHA
branch-delete ignore, and invalid-push-payload 400 — the idempotency/ping/
signature tests didn't reach these branch-filter arms.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): make new handler tests re-run-safe on the shared test DB

custom_domains.hostname is GLOBALLY unique and SetupTestDB reuses one shared
instant_dev_test DB, so hardcoded hostnames collided across consecutive runs
(and the rate-limit IP key survived its 2s TTL). Switch custom-domain hostnames
+ the promote-approval rate-limit IPs to per-run uuid-derived values. Verified
3x consecutive runs green without a DB reset.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): raise handler coverage for billing/vault/twin/custom_domain/deletion_confirm/webhook/storage

Build fake-client scaffolding (no waivers) to drive previously-unreachable
error and network arms under CI conditions:

- billing.go: add BillingPortal interface + billingPortalFactory seam
  (SetBillingPortalForTestPortal) so ListInvoicesAPI / UpdatePaymentMethodAPI /
  ChangePlanAPI post-subscription arms (success / circuit-open / razorpay-error)
  run against a FAKE Razorpay portal — never a real network call, never rzp_live.
  93.4% -> 94.9%.
- vault.go: bad-AES-key app + closed-DB app drive encrypt/decrypt/persist/fetch/
  list/delete/copy error arms; invalid-team JWT + validate-arms + quota-blocked
  copy. 79.9% -> 93.7%.
- deletion_confirm.go: BV* export seams drive requestEmailConfirmedDeletion /
  resolveEmailConfirmedDeletion / cancelEmailConfirmedDeletion through a real
  fiber.Ctx + fake Mailer + crafted pending_deletions rows. 74.8% -> 89.3%.
- twin.go: consumeApprovedTwin approval-gate arms (not-approved/mismatch/expired/
  wrong-team/success) + family-validate twin_exists + redis dispatch. 65.1% -> 80.2%.
- custom_domain.go: stubbed CustomDomainProvider cert-poll/still-issuing/cap/
  ingress-teardown arms + validateHostname rejections + closed-DB arms. 76.8% -> 82.6%.
- webhook.go: auth invalid-team + garbage-ring-item decode-skip. 90.6% -> 91.0%.
- storage.go: authenticated quota-exceeded / team-lookup-fail / invalid-team arms.

All new tests carry the _bvwave suffix; re-exports live in export_bvwave_test.go
(no collisions with existing export_*_test.go). go build ./... + go vet clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): _vecwave handler arm coverage for vector/backup/github_deploy/resource

Adds black-box + targeted-arm tests pushing internal/handlers coverage on the
four target files via fake-provider scaffolding (no production code touched):

- vector.go (67.9%→84.2%): bufconn fakeProvisioner gRPC arm of provisionVectorDB
  + createPgvectorExtension stub (now 100%); decryptConnectionURL three arms via
  exported VectorDecryptConnectionURLForTest; anon/auth validation + 402 + gRPC-
  error + persist-failure arms.
- backup.go (78.9%→85.3%): CreateRestore target_resource_id arms (invalid/
  not-found/cross-team/type-mismatch/happy/missing-sha256); list/map all-optional-
  fields; requireOwnedResource fetch_failed (brokenDB); parseIntStrict too-large;
  parseUserIDFromCtx malformed.
- github_deploy.go (73.3%→81.8%): Connect invalid_body/invalid_branch/
  already_connected(409)/installation_id; Receive deploy-triggered(202)+rate-
  limit(429); not-found arms (bad uuid / unknown connection / cross-team).
- resource.go (92.8%→93.4%): pause/resume provider HAPPY paths against real
  pg/redis/mongo (revoke/grant CONNECT, ACL on/off, role grant/revoke) via
  exported Call{Pause,Resume}ProviderForTest; RotateCredentials postgres+redis
  happy paths.

New test-only export file export_vecwave_test.go (no duplicate re-exports).

Residual <95% arms are error-injection (DB-mid-call failure, races, recycle/
dedup birthday-collision) and the unreachable crypto/rand.Read arm
(RotateCredentials, Go 1.26) — documented in the brief report.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): billing CreateCheckoutAPI invalid-body + dedup-guard arm

Adds a CreateCheckoutAPI test that wires Redis (exercising the SETNX dedup
guard success path) and posts malformed JSON to hit the invalid_body 400 arm.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): drive deploy/stack async-pipeline handlers toward 95%

Adds the deploy/stack async-build-pipeline coverage slice (suffix
_deployasync). Builds a fault-injecting database/sql driver
(faultdb_deployasync_test.go) that proxies lib/pq and forces query/exec
failures after the Nth call, so the mid-handler 'first query succeeds →
a later query errors → 503' arms — unreachable with a plain closed-DB
handle — are exercised across Get/Logs/List/UpdateEnv/Delete/Redeploy/
Promote/Family/ConfirmDelete/CancelDelete.

Per-file coverage (CI conditions, pgvector pg16 + redis7 + mongo6 + nats):
  deploy.go            86.3% → 95.2%
  promote_approval.go  72.5% → 93.5%
  stack.go             80.9% → 92.5%

Covers runStackDeploy/runStackRedeploy/captureAutopsy/runDeploy +
per-service onUpdate/onImageBuilt callbacks via the existing noop
StackProvider/compute.Provider doubles + SetStackProvider/
SetComputeProvider seams, plus ghost-team requireTeam error arms,
TTL/env/tarball validation arms, promote create/in-place/vault paths,
and email-confirmed deletion deprovision flows.

Residual <95% arms are test-resistant: k8s-provider constructors
(need a live cluster), corrupted-multipart tarball open/read arms,
crypto/rand + marshal-of-controlled-data branches (cannot fail), and
CAS-lost-race !ok branches (concurrency-only). Documented in the
agent report.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): final handlers pass — twin/github/custom_domain/vector/backup/storage/stack/resource/webhook/api_keys + DNS+minio+twin seams

Drives the api internal/handlers package coverage up via:
- twin.go: faultdb DB-error arms (source/team/approval lookups, execute,
  validate-family) + bad-team-id + named-source carry-forward + derefUUID unit.
- github_deploy.go: Connect/Get/Disconnect/Receive fetch/create/delete/enqueue/
  decrypt DB-error arms via faultdb + httptest-signed push; requireTeam arms.
- custom_domain.go: NEW lookupTXT package-var seam → TXT-match success +
  cert->ready transition; faultdb DB-error arms across requireTeam/Stack/Domain,
  Create/List/Delete; serialize optional fields.
- vector.go / storage.go: bufconn fakeProvisioner + NEW hermetic prefix-scoped
  storage impl (storage.NewWithImpl seam) → credential mode + per-tenant IAM
  audit arm; soft-delete + persist-fail arms.
- backup.go: count-fail-while-list-succeeds + insert/team/list DB-error arms.
- stack.go: k8s-provider-fallback constructor branch, checkStackDeployLimit
  redis-error, stackOwnerCheck anon-mismatch, consumeApprovedPromote DB-error.
- resource.go: Pause/Resume/Rotate faultdb DB-error + rollback arms.
- webhook.go / api_keys.go / family_bulk_twin.go: auth + DB-error arms.
- agent_action / maskSourceIP / buildContextConfig pure-function default branches.

New seams: custom_domain.lookupTXT package var; storage.NewWithImpl (hermetic
prefix-scoped impl); reuse of existing faultdb driver + bufconn fakeProvisioner.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): more handlers DB-error arms — status/api_keys/internal_terminate + pure-func default branches

- status.go: compute per-component fail-open + listComponents-error 500.
- api_keys.go: Create/List/Revoke db_failed via faultdb + bad-id + not-found.
- internal_terminate.go: invalid-team-id + team-lookup/idempotency/pause DB-error arms.
- agent_action/maskSourceIP/buildContextConfig pure-function default branches.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): deletion_confirm resolveEmailConfirmedDeletion DB-error arms

ConfirmDelete → resolveEmailConfirmedDeletion missing-token / lookup-failed /
mark-failed arms via faultdb + seeded pending_deletions row.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): vector anonymous over-cap dedup decrypt-fail fallthrough arm

Drives vector.go's over-cap dedup branch where the existing resource's stored
connection_url fails to decrypt (corrupted post-provision), exercising the
fail-closed fallthrough to fresh provision (vector.go:294-298). DB-exposed
bufconn fixture lets the test corrupt the row + hammer the cap.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): vector over-cap dedup + cross-service fallback; backup CreateRestore/ListRestores DB-error arms

- vector.go: over-cap dedup decrypt-fail fallthrough (corrupt all active rows
  then over-cap) + cross-service-fallback 429 (retype rows so type-lookup
  misses but any-lookup hits); auth gRPC-error soft-delete via DB-exposed
  failing bufconn fixture.
- backup.go (now 95.8%): CreateRestore team/backup/inflight/insert DB-error
  arms, no-user 401, missing-ack 400; ListRestores count-fail; bad-team/bad-id
  across all four routes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): storage decideStorageMode bucket-per-tenant + quota arms; storage/vector anon over-cap dedup

- storage.go: decideStorageMode BucketPerTenant+paid branch; auth quota-check
  fail-open + storage-limit-reached 402; anonymous over-cap dedup happy +
  decrypt-fail fallthrough.
- vector.go: anonymous over-cap dedup happy path.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): stack New deployment-limit/count-fail + ConfirmDelete email-disabled; webhook anon over-cap dedup

- stack.go New: deployment_limit_reached 402 (over-cap team), quota_check_failed
  503 (faultdb count error); ConfirmDelete deletion_email_disabled (no email
  client wired).
- webhook.go NewWebhook: anonymous over-cap dedup branch.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): resource RotateCredentials update_failed arm

RotateCredentials persist-error arm (UpdateConnectionURL fails after a
decryptable URL rotate) via faultdb + rfEncryptURL seed helper.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): audit List/CSV DB-error; vault PutSecret bad-AES; resource_family Family/ListFamilies DB-error arms

- audit.go: List + ListCSV db_failed via faultdb.
- vault.go: PutSecret encryptPlaintext aes-key-invalid 500.
- resource.go (Family/ListFamilies): fetch_failed + invalid_id + list DB-error.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): whoami enrichment/nil-db; usage_wall + deploys_audit DB-error arms

- whoami.Get: nil-db early return + tier/email enrichment success.
- usage_wall.GetWall: db_failed via faultdb.
- deploys_audit.List: db_failed via faultdb.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): final serial pass #2 — handlers DB-error/edge arms toward 95%

Add 12 _final2 test files covering the biggest reachable uncovered chunks in
the api internal/handlers package, lifting the package from 93.12% → 93.9%
under CI conditions (pgvector/redis/mongo/nats, migrations applied,
go test ./internal/handlers/... -short -count=1 -p 1).

Covered arms (all reachable; no waivers):
- auth.go: GitHub/Google find-or-create LinkID errors, email-lookup DB errors
  (fault DB), Google empty-name team-name fallback. (OAuth flows driven via the
  existing SetOAuthURLsForTest + startFakeOAuth seams.)
- sns_verify.go: full snsVerifier.verify matrix (guards, cert-fetch error,
  sig decode, SignatureVersion 1/unknown/2, happy RSA verify) + getCert
  cache hit/miss/error via a generated throwaway RSA cert + fetchCert seam.
- internal_terminate.go: dunning/downgrade DB-error arms (staged openFaultDB),
  razorpay canceler-not-configured / cancel-error / happy paths.
- internal_resend_magic_link.go: lookup/update-hash DB errors + mark-failed /
  attempts-lookup / mark-abandoned best-effort error arms (fault DB + fake mailer).
- deploy_ttl.go: MakePermanent/SetTTL update_failed + refresh_failed +
  lookupDeployment fetch_failed (seeded deploy + staged fault DB).
- deploy_teardown_reconciler.go: begin_tx_failed, list_failed, empty-tx commit,
  mark_failed (fault DB + existing fakeTeardownProvider).
- env_policy.go: Get fetch_failed, Put role_lookup_failed / owner_required /
  invalid_body / invalid_env_policy / persist_failed / happy.
- billing.go (CreateCheckoutAPI): CreateSubscription circuit-open / razorpay-error
  / incomplete-response / persistence-failure (fake CreateSubscription seam,
  HMAC-signed payloads only, never a real key).
- stack.go: UpdateEnv fetch/persist DB errors + Family fetch_failed (seeded
  stack + staged fault DB).
- audit.go: bad-team-id unauthorized + List/CSV happy paths with metadata +
  actor-email lookup.
- provision_helper.go denyProvisionOverCap (was 0%) + per-handler over-cap
  (no-existing → 429; with-existing → dedup-return) + backend-failure
  soft-delete + finalizeProvision-failure persist arms across
  db/cache/nosql/vector/queue/storage/webhook (anon + authenticated).

No new exported test symbols were required (all arms reachable via existing
export_*_test.go seams or white-box package handlers tests), so no
export_final2_test.go was added.

Provably-unreachable arms left uncovered (documented): defaultFetchCert
(real network), generateOAuthState/issueSessionJWT failures (crypto/rand +
HS256 []byte key cannot fail), the k8s ComputeProvider constructor branch
(needs a live cluster; noop fallback is covered).

Known pre-existing shared-dev-DB flakes unrelated to this change: TestAdminList_*
(limit=100 paging on accumulated data) and TestQueue_* (NATS :8222 monitor not
reachable locally) — verify on a fresh DB per CLAUDE.md.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): final pass #2 batch 2 — bulk-twin happy path + vault copy list-failed

- family_bulk_twin happy path: seed active postgres+redis+mongodb parents in
  one env and POST /families/bulk-twin to another with working local backends,
  exercising twinOneParent + ProvisionForTwinCore success for all three
  backends (db/cache/nosql 81.8%→89.3%, twinOneParent →82.6%).
- vault.go CopySecrets list_failed arm: rename vault_secrets away after the
  team/tier checks pass (isolated DB) so ListVaultSecretKeys errors → 5xx.

Package 93.9% → 93.93% under CI conditions; only the documented pre-existing
shared-dev-DB flakes remain (TestAdminList_* paging, TestQueue_* NATS :8222).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): final serial pass #3 — handlers 93.9%→94.5% via behavior-preserving seams

Production seams (package-var indirection, prod path unchanged, each new line covered):
- helpers.go: `randRead = rand.Read` — routes generateAppID/generateOAuthState/
  generateSessionID through it so the rand.Read error arm is testable.
- stack.go: `openMultipartFile` (tarball open) + `newK8sStackProvider` factory.
- deploy.go: routes both fh.Open() sites through openMultipartFile; adds
  `newK8sComputeProvider` factory; routes generateAppID through randRead.

Both k8s factory seams cover their default closure body AND the noop-fallback
error arm + the success arm via injected fakes; all seam lines are 100% covered.

New _final3 tests close confirmed-reachable arms: stack/deploy multipart New
error arms, stack Redeploy invalid-form + tarball-open, internal-JWT verifier
wrong-method/bad-uuid arms (terminate/backup-refund/resend), OAuth-start +
CLI-session rand.Read errors, /claim missing-email/invalid-token/already-claimed,
resolveResourceBindings rejection matrix, resource Family not-found/cross-team,
familyMember map name/parent branches, agent-action deployment-limit tiers,
requireName/sanitizeName invalid-UTF8, respondProvisionFailed circuit-open,
twin ProvisionForTwinCore create/validate fault arms, webhook anon over-cap dedup.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): final pass #3 cont. — k8s seam closures+error arms, team_members + promote audit

- Cover the newK8sStackProvider / newK8sComputeProvider default closure bodies
  (InvokeDefault*ForTest) AND the noop-fallback error arms (seam forced to
  error) so every new seam production line is 100% covered (patch-coverage).
- team_members cacheInviteResponse (nil-rdb/dead-rdb store-error) + emitInviteAudit
  insert-error arms via exporters + fault DB.
- emitPromoteAuditEvent InsertAuditEvent-error arm via exporter + fault DB.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): final pass #3 — experiments.Converted validation arms

missing-fields / unknown-experiment / invalid-variant / action-truncation arms
of POST /api/v1/experiments/converted.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): final pass #3 — vault.audit AppendVaultAudit-error arm

Drives (*VaultHandler).audit best-effort warn arm via exporter + fault DB.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): final pass #3 — injectMessageID empty-id/unmarshal-error arms

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): final pass #3 — deploy Patch invalid_body arm

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): final pass #3 — ResourceEnvByTokenForMiddleware all 3 arms

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): final pass #3 — env_policy.Get nil-policy + bad-team-id arms

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): final pass #3 — CancelForTeam no-subscription + DB-error arms

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(coverage): seam2 — quota.CheckStorageQuota + k8s logs clientset seams

Add two behavior-preserving package-var seams to lift internal/handlers coverage:
- var checkStorageQuota = quota.CheckStorageQuota (seams.go); route db/cache/nosql
  call sites through it so StorageExceeded warning arms are reachable in tests.
- var buildLogsClientset over buildLogsK8sClientset; NewLogsHandler + the
  in-cluster/kubeconfig fallback now coverable via a client-go fake clientset.

NewLogsHandler + buildLogsK8sClientset now 100%. Local handlers coverage
93.93% -> 94.8% (capped: TestQueue_* requires NATS, unavailable locally;
authoritative number needs CI with NATS + customer-DB).

Recovered from a stalled background agent (stalled on the full ./... measure).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci: run NATS (monitoring :8222) in build-and-test so TestQueue_* pass

The TestQueue_IsolatedCreds_EmbeddedInResponse / _CredIssueError_FallsBackToLegacyOpen
coverage tests provision via queueprovider.Provision, which health-checks
http://localhost:8222/healthz (NATSHost defaults to localhost). CI had no NATS
service, so they 503'd on every run — the required build-and-test check has been
red on this branch since they landed. Add a nats-server -m 8222 step (service
containers can't pass -m). Repoint the two NATS-DOWN tests (finalize/softdelete
final2) off 127.0.0.1 to the reserved non-resolvable host nats.test (matching the
7 other queue-fault tests) so a live localhost NATS doesn't satisfy their 503
expectation.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 3, 2026
…ug-bash #2)

buildLogCache snapshots kaniko build logs of FAILED builds for the
failure autopsy. It was TTL-bounded (30m) and swept stale entries on each
new failure — but had NO size cap. A burst of failing builds inside one
TTL window (a broken base image, a wedged registry, a fork-bomb of bad
deploys) accumulates one ≤200-line snapshot per failure with no ceiling:
unbounded memory growth on a long-lived api pod.

Add buildLogCacheMaxEntries=256 and capBuildLogCacheSize(), called
evict-after-store from snapshotBuildLogs. When a store pushes the live
count past the cap, evict oldest-first (a recent failure's autopsy is far
likelier to be read than one from hundreds of failures ago). sync.Map has
no length, so we snapshot keys+times in one Range pass, sort by
capturedAt, and delete the excess oldest.

Tests: TestCapBuildLogCacheSize_EvictsOldestOverCap (over-cap → newest
survive, oldest evicted), TestCapBuildLogCacheSize_NoOpUnderCap. Both new
funcs 100% covered; verified locally.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 3, 2026
…ug-bash #2)

buildLogCache snapshots kaniko build logs of FAILED builds for the
failure autopsy. It was TTL-bounded (30m) and swept stale entries on each
new failure — but had NO size cap. A burst of failing builds inside one
TTL window (a broken base image, a wedged registry, a fork-bomb of bad
deploys) accumulates one ≤200-line snapshot per failure with no ceiling:
unbounded memory growth on a long-lived api pod.

Add buildLogCacheMaxEntries=256 and capBuildLogCacheSize(), called
evict-after-store from snapshotBuildLogs. When a store pushes the live
count past the cap, evict oldest-first (a recent failure's autopsy is far
likelier to be read than one from hundreds of failures ago). sync.Map has
no length, so we snapshot keys+times in one Range pass, sort by
capturedAt, and delete the excess oldest.

Tests: TestCapBuildLogCacheSize_EvictsOldestOverCap (over-cap → newest
survive, oldest evicted), TestCapBuildLogCacheSize_NoOpUnderCap. Both new
funcs 100% covered; verified locally.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 3, 2026
…ug-bash #2) (#230)

buildLogCache snapshots kaniko build logs of FAILED builds for the
failure autopsy. It was TTL-bounded (30m) and swept stale entries on each
new failure — but had NO size cap. A burst of failing builds inside one
TTL window (a broken base image, a wedged registry, a fork-bomb of bad
deploys) accumulates one ≤200-line snapshot per failure with no ceiling:
unbounded memory growth on a long-lived api pod.

Add buildLogCacheMaxEntries=256 and capBuildLogCacheSize(), called
evict-after-store from snapshotBuildLogs. When a store pushes the live
count past the cap, evict oldest-first (a recent failure's autopsy is far
likelier to be read than one from hundreds of failures ago). sync.Map has
no length, so we snapshot keys+times in one Range pass, sort by
capturedAt, and delete the excess oldest.

Tests: TestCapBuildLogCacheSize_EvictsOldestOverCap (over-cap → newest
survive, oldest evicted), TestCapBuildLogCacheSize_NoOpUnderCap. Both new
funcs 100% covered; verified locally.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 5, 2026
common #46 retired every unlimited (-1) resource limit to finite caps
(Team/Growth now finite). Four api handler tests still assumed the old
"team/growth = unlimited" reality and reded build-and-test, which broke
api master (every PR checks out common@master with the new numbers).

Fixes (TESTS only — the -1-handling production code is preserved as
correct defensive utilities, still reachable via provisions_per_day: -1
and future tiers):

- billing_usage_coverage: TestBillingUsage_UnlimitedTier_MbToBytesNegative
  routed the team tier (was -1) through the HTTP usage handler to hit
  mbToBytes(-1). No real tier carries -1 now, so the -1 path can't be
  reached via a tier. Moved to a direct internal test
  (mb_to_bytes_internal_test.go) that feeds synthetic -1/finite inputs —
  keeps the defensive "-1 -> ∞" branch covered.
- misc_routes_block_integration: the "team tier short-circuits to
  near_wall=false" subtest assumed the unlimited early-return. That
  early-return was removed (Team is finite -> walls apply); subtest now
  asserts near_wall=true with a seeded wall row.
- small_handlers_final: TestUsageWallFinal_DBError_503 used failAfter=1
  (team-tier pre-query #1 then audit query #2). With the early-return
  gone the audit query is the FIRST DB call -> failAfter=0.
- team_coverage_mock: TestTeamMembers_InviteMember_LegacyMemberSuccess
  assumed team_members=-1 so withinMemberLimit skipped the count query.
  Team is now 25 -> added the teamSeatTotal (members + pending) mock
  expectations; 1 seat < 25 -> within limit -> 201.

New rule-18 guard (strict_margin_finite_limits_test.go): iterates the
LIVE plans.yaml registry and fails if any tier ships a -1 on any costed
resource limit (only provisions_per_day may be -1). Re-introducing an
unlimited resource cap — or a new tier with one — now reds here instead
of silently shipping unbounded COGS.

Team stays GATED (no checkout change). Unblocks api master.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 5, 2026
common #46 retired every unlimited (-1) resource limit to finite caps
(Team/Growth now finite). Four api handler tests still assumed the old
"team/growth = unlimited" reality and reded build-and-test, which broke
api master (every PR checks out common@master with the new numbers).

Fixes (TESTS only — the -1-handling production code is preserved as
correct defensive utilities, still reachable via provisions_per_day: -1
and future tiers):

- billing_usage_coverage: TestBillingUsage_UnlimitedTier_MbToBytesNegative
  routed the team tier (was -1) through the HTTP usage handler to hit
  mbToBytes(-1). No real tier carries -1 now, so the -1 path can't be
  reached via a tier. Moved to a direct internal test
  (mb_to_bytes_internal_test.go) that feeds synthetic -1/finite inputs —
  keeps the defensive "-1 -> ∞" branch covered.
- misc_routes_block_integration: the "team tier short-circuits to
  near_wall=false" subtest assumed the unlimited early-return. That
  early-return was removed (Team is finite -> walls apply); subtest now
  asserts near_wall=true with a seeded wall row.
- small_handlers_final: TestUsageWallFinal_DBError_503 used failAfter=1
  (team-tier pre-query #1 then audit query #2). With the early-return
  gone the audit query is the FIRST DB call -> failAfter=0.
- team_coverage_mock: TestTeamMembers_InviteMember_LegacyMemberSuccess
  assumed team_members=-1 so withinMemberLimit skipped the count query.
  Team is now 25 -> added the teamSeatTotal (members + pending) mock
  expectations; 1 seat < 25 -> within limit -> 201.

New rule-18 guard (strict_margin_finite_limits_test.go): iterates the
LIVE plans.yaml registry and fails if any tier ships a -1 on any costed
resource limit (only provisions_per_day may be -1). Re-introducing an
unlimited resource cap — or a new tier with one — now reds here instead
of silently shipping unbounded COGS.

Team stays GATED (no checkout change). Unblocks api master.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 5, 2026
…d (-1) (#262)

* feat(plans): strict-≥80%-margin tier redesign — retire every "unlimited" (-1)

CEO directive (2026-06-05, LOCKED): retire every `-1` ("unlimited") limit in
api/plans.yaml (the source of truth) except provisions_per_day (a
per-fingerprint rate limit, no per-GB COGS) and apply the locked number table.
Hard-caps only (no metered overage). Team stays GATED — these are the eventual
contract numbers, NOT an un-gate. Above Team = Enterprise (contact sales).

plans.yaml changes (monthly + *_yearly mirror):
- anonymous/free: queue 1024→64 MB; queue_count -1→1 (0=unlimited via
  QueueCountLimit, so 1).
- hobby: queue 5120→2048 MB. pro: queue 10240→5120 MB.
- growth: vector 20480→10240; redis_commands -1→5M; mongo 20480/50/50k;
  queue 20480/50; storage 153600; webhooks 100000 (all -1 retired).
- team: every -1 → finite (pg 51200/100, vector 30720/100, redis 1536/10M,
  mongo 40960/50/50k, queue 40960/100, storage 307200, webhooks 100000,
  members 25, vault 1000, deploys 100).

Other surfaces (rule 22):
- openapi.go: drop "$199 unlimited" + "team is unlimited" wording.
- usage_wall.go: remove the team-tier short-circuit (Team is finite now, so it
  has quota walls like every other tier) — falls through to the audit query.
- Registry-iterating regression: TestPlansYAML_NoUnlimitedExceptProvisionsPerDay
  reflects over every Limits int field on every tier and fails the build if any
  -1 reappears (or a new int field is added with -1) except provisions_per_day.
- Synced pinning tests: QueueCountLimit, vector limits, deploys/queue caps,
  webhook (team 100k; -1→10000 clamp arm now covered via a synthetic registry).

Rule-22 synchronized pricing redesign (strict-80%, retire unlimited) — see
docs/sessions/2026-06-04/PRICING-MARGIN-MODEL-AND-TEAM-REDESIGN.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(openapi): regenerate snapshot for strict-margin description edits

snapshot-drift CI gate caught the openapi.go changes (Team "$199 unlimited" →
"$199 finite, not unlimited" + usage-wall "team is unlimited" → "Team has
finite limits"). Regenerated via `make openapi-snapshot` (rule 22 — the
dashboard + instanode-web typed clients depend on this file).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(plans): fix tier-assumption tests for strict-80%-margin redesign

common #46 retired every unlimited (-1) resource limit to finite caps
(Team/Growth now finite). Four api handler tests still assumed the old
"team/growth = unlimited" reality and reded build-and-test, which broke
api master (every PR checks out common@master with the new numbers).

Fixes (TESTS only — the -1-handling production code is preserved as
correct defensive utilities, still reachable via provisions_per_day: -1
and future tiers):

- billing_usage_coverage: TestBillingUsage_UnlimitedTier_MbToBytesNegative
  routed the team tier (was -1) through the HTTP usage handler to hit
  mbToBytes(-1). No real tier carries -1 now, so the -1 path can't be
  reached via a tier. Moved to a direct internal test
  (mb_to_bytes_internal_test.go) that feeds synthetic -1/finite inputs —
  keeps the defensive "-1 -> ∞" branch covered.
- misc_routes_block_integration: the "team tier short-circuits to
  near_wall=false" subtest assumed the unlimited early-return. That
  early-return was removed (Team is finite -> walls apply); subtest now
  asserts near_wall=true with a seeded wall row.
- small_handlers_final: TestUsageWallFinal_DBError_503 used failAfter=1
  (team-tier pre-query #1 then audit query #2). With the early-return
  gone the audit query is the FIRST DB call -> failAfter=0.
- team_coverage_mock: TestTeamMembers_InviteMember_LegacyMemberSuccess
  assumed team_members=-1 so withinMemberLimit skipped the count query.
  Team is now 25 -> added the teamSeatTotal (members + pending) mock
  expectations; 1 seat < 25 -> within limit -> 201.

New rule-18 guard (strict_margin_finite_limits_test.go): iterates the
LIVE plans.yaml registry and fails if any tier ships a -1 on any costed
resource limit (only provisions_per_day may be -1). Re-introducing an
unlimited resource cap — or a new tier with one — now reds here instead
of silently shipping unbounded COGS.

Team stays GATED (no checkout change). Unblocks api master.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant