Skip to content

api: §10.20 cached aggregation endpoints + cache helper - #22

Merged
mastermanas805 merged 1 commit into
masterfrom
caching-layer-2026-05-12
May 12, 2026
Merged

mastermanas805 merged 1 commit into
masterfrom
caching-layer-2026-05-12

Conversation

@mastermanas805

Copy link
Copy Markdown
Member

Summary

Adds the server-side aggregation layer the dashboard needs to retire its client-side resource-list scans for the Usage panel + sidebar counts. Honors the §13 caching-and-consistency memory rule: gating paths stay real-time, read aggregates become eventual-consistent with visible as_of.

  • internal/cache.GetOrSet[T] — typed Redis cache wrapper with singleflight collapse + fail-open on Redis errors.
  • GET /api/v1/billing/usage — 30s Redis cache, Cache-Control: private, max-age=30, stale-while-revalidate=60. Storage bytes per service + deployment/webhook/vault/member counts. Returns as_of + freshness_seconds.
  • GET /api/v1/team/summary — 5m Redis cache, Cache-Control: private, max-age=300. Resource breakdown by type, deployment / member / vault counts. Drives the dashboard sidebar.

§14 caching-review checklist

GET /api/v1/billing/usage

  1. Fresh-required or eventual-OK? — Eventual. Not on a gating path.
  2. Staleness window? — 30s server-side + up to 60s stale-while-revalidate on the browser.
  3. Cache layer? — Redis billing:usage:<team_id> + HTTP Cache-Control.
  4. Invalidation trigger? — TTL expiry. Aggregates self-heal on next miss.
  5. Cache-down behavior? — Fail-open: GetOrSet falls through to fn (DB) so the endpoint stays up.
  6. N concurrent requests? — singleflight.Group collapses to 1 fn invocation per key.
  7. Cache key scope? — Team-scoped (billing:usage:<team_id>). No cross-team leak.

GET /api/v1/team/summary — same answers, longer TTL (300s server, 300s browser), no stale-while-revalidate (window already wide).

Test plan

  • make test-unit green — all packages, including new internal/cache.
  • cache_test.go: singleflight collapses 20 concurrent callers to 1 fn invocation; Redis-down falls through; corrupt cache entry heals; nil client passes through; negative caching works; Invalidate clears key.
  • billing_usage_test.go: same team hit twice in <30s = ONE set of DB queries (sqlmock strict mode); Redis-down → two DB roundtrips, still 200 OK; different teams → separate cache entries.
  • team_summary_test.go: same.
  • Live-cluster sanity: hit the endpoint, observe Cache-Control headers, verify the dashboard renders the as of Ns ago footnote. (Manas to verify post-merge.)

Quota check paths (POST /db/new, /cache/new, etc.) are intentionally not modified — they read fresh per §13 because they gate provisioning.

🤖 Generated with Claude Code

Adds the server-side aggregation layer the dashboard needs to retire its
client-side resource-list scans for the Usage panel and sidebar counts.
The §13 freshness matrix calls for eventual-consistent caching on these
read paths (gating paths stay real-time per their own bucket).

cache.GetOrSet[T](ctx, rdb, key, ttl, fn) collapses concurrent identical
requests through golang.org/x/sync/singleflight and fails open when Redis
is unreachable (logs + falls through to fn). Negative caching is allowed;
corrupt cache entries are treated as misses and healed on the next SET.

GET /api/v1/billing/usage  — Redis-cached 30s, Cache-Control: private,
max-age=30, stale-while-revalidate=60. Aggregates storage bytes per
service (postgres/redis/mongodb), counts of deployments/webhooks/vault/
members. Carries `as_of` ISO timestamp + `freshness_seconds` so the
dashboard can render an "as of Ns ago" footnote.

GET /api/v1/team/summary  — Redis-cached 5m, Cache-Control: private,
max-age=300. Resource breakdown by type (one GROUP BY query), deployment
count, member count, vault key count. Consumed by the sidebar
SidebarUpgradeCard for resource/member counts that previously didn't
render because the dashboard had no live source.

Cache keys are team-scoped (`billing:usage:<team_id>`,
`team:summary:<team_id>`) so two teams never share an entry.

Tests:
  - internal/cache: singleflight collapses concurrent callers to one fn
    invocation; Redis-down falls through to fn; corrupt entry heals;
    zero-value (negative) caching works; Invalidate clears the key.
  - internal/handlers: two requests for the same team in <30s/<5min run
    exactly one set of DB queries (asserted via sqlmock strict mode);
    different teams trigger separate aggregations.

Quota check paths (POST /db/new, /cache/new, etc.) are intentionally
NOT touched — they read fresh per §13 because they gate provisioning.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@mastermanas805
mastermanas805 merged commit 2d959b1 into master May 12, 2026
@mastermanas805
mastermanas805 deleted the caching-layer-2026-05-12 branch May 12, 2026 03:03
mastermanas805 added a commit that referenced this pull request May 21, 2026
…lit cleanup) (#134)

These directories are leftover from the original monorepo before
worker and provisioner were split into their own repos. The stale
go.mod files inside them were never actually built — but OSV-Scanner
indexes them and flags ~15 false-positive CVEs (stdlib 1.25.0 issues,
x/net@0.49.0, fiber CVEs that don't apply to the live api binary).

Verified safe to delete:
- api/provisioner/go.mod last modified 2026-05-12 (obs scaffolding)
- api/worker/go.mod same
- live provisioner / worker code lives in InstaNode-dev/provisioner
  and InstaNode-dev/worker repos
- no internal package imports from api/provisioner/* or api/worker/*

Closes pending task #22 (C1 cleanup api/worker + api/provisioner subdirs).

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 2, 2026
CI (full real-DB suite) caught three existing tests that encoded the
pre-fix behavior batch-1 deliberately changed:
- TestPlansRegistry_IsDedicatedTier: team is now dedicated (#12).
- TestDeployCancelDelete_CrossTeam: cross-tenant now 404 not 403 (#22).
- TestBrevo_Receive_UnknownMessageID (billing_coverage): delivered UPDATE now
  carries the terminal-class guard (8 args) + an existence-probe SELECT (#6).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mastermanas805 added a commit that referenced this pull request Jun 2, 2026
… 404 leak, Team dedicated (#218)

* fix(api): bug-bash batch — free TTL, brevo non-clobber, dedup expiry, 404 leak, Team dedicated

Five confirmed bugs from the 2026-06-02 platform bug bash:

- #4 (P1) free-tier resources never expired: authenticated provisions
  hardcoded ExpiresAt=nil even for free/anonymous tiers. Add
  resourceExpiryForTier (24h for ephemeral tiers, nil for paid) and apply it
  at all 10 authenticated CreateResource sites. Per product decision: enforce
  plans.yaml's documented 24h TTL for claimed-unpaid resources.
- #6 (P1) Brevo 'delivered' webhook clobbered a terminal bounce/complaint on
  out-of-order delivery, corrupting the email truth surface (rule 12). Guard
  the UPDATE against terminal classes; distinguish terminal-kept from unknown.
- #17/#20 (P2) fingerprint dedup-return handed back credentials for
  active-but-expired anonymous resources: add the expires_at filter to both
  GetActiveResourceByFingerprint[Type], matching GetAllActiveResourcesByFingerprint.
- #22 (P3) deploy CancelDelete returned 403 cross-tenant (leaking existence);
  now 404 like the other deploy endpoints.
- #12 (P2) Team tier gets dedicated infra: add dedicated:true to team +
  team_yearly in plans.yaml (pairs with common defaultYAML). Per product
  decision 2026-06-02.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(brevo): cover #6 terminal-kept non-clobber path + fix delivered mocks

The delivered UPDATE now carries the terminal-class guard (8 args) and a 0-row
result triggers an existence probe. Update expectDeliveredUpdate to the new arg
list, add the SELECT mock to the unknown-message test, and add a terminal-kept
regression (delivered-after-bounce → matched:true, class preserved). Closes the
batch-1 patch-coverage gap on brevo_webhook.go.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(api): update existing tests for batch-1 behavior changes

CI (full real-DB suite) caught three existing tests that encoded the
pre-fix behavior batch-1 deliberately changed:
- TestPlansRegistry_IsDedicatedTier: team is now dedicated (#12).
- TestDeployCancelDelete_CrossTeam: cross-tenant now 404 not 403 (#22).
- TestBrevo_Receive_UnknownMessageID (billing_coverage): delivered UPDATE now
  carries the terminal-class guard (8 args) + an existence-probe SELECT (#6).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(brevo): cover delivered-handler SELECT-probe error branch (#218 coverage)

diff-cover flagged brevo_webhook.go:530-531 — the `if qErr != nil` arm of the
delivered handler's existence probe (bug bash #6). When the terminal-class-
guarded UPDATE affects 0 rows, a follow-up SELECT distinguishes terminal-kept
from genuinely-unknown; a non-ErrNoRows fault on that probe must surface as an
error (→ 500) so Brevo retries rather than the message being mislabeled.

Adds TestBrevo_Receive_Delivered_ProbeError (sqlmock, hermetic): UPDATE → 0
rows, SELECT probe → generic error, asserts 500.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant