Skip to content

LTRAC-2154: Cache routing KV reads at the Cloudflare location for 5 minutes - #3261

Merged
jorgemoya merged 3 commits into
canaryfrom
LTRAC-2154/perf/kv-read-cache-ttl
Oct 2, 2026
Merged

jorgemoya merged 3 commits into
canaryfrom
LTRAC-2154/perf/kv-read-cache-ttl

Conversation

@jorgemoya

@jorgemoya jorgemoya commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Linear: LTRAC-2154

What/Why?

CloudflareKvAdapter read CATALYST_ROUTES_KV without a cacheTtl, so each Cloudflare location cached a read for Workers KV's 60s default. That's the same window as the in-process L1 in front of it (SHARED_STORE_RECHECK_MS). The L1 re-reads just as the location's copy expires, so in load testing nearly every routing-cache read went to central KV storage, about 88ms, with cold and warm reads looking the same.

Reads now pass cacheTtl: 300. It's longer than the L1 window, so L1 re-reads hit the location cache, and it's capped at the 5-minute storefront-status window, the shortest freshness window with-routes stores. Two tests pin both bounds.

Trade-off worth a look (also raised by Codex review): a value read just before its expiryTime stays in the location cache for the full 5 minutes. So in the worst case, a storefront status change (maintenance mode, launch) can take up to about 11 minutes to reach a location instead of about 7 (status window 5 min + location cache + 1 min L1). Route changes get the same 4 extra minutes on top of their 30-minute window. It's usually sooner: the request that sees the expired value triggers a background refresh, and Cloudflare says a write is usually, though not guaranteed to be, visible at once in the location that made it.

We chose to accept and document this rather than shorten the status window, because that window is set in with-routes and shared by every adapter. Shortening it would add Storefront API calls on Vercel and other hosts that don't have this location cache. Keys that don't exist in KV aren't helped; those reads still go to central storage.

Testing

Load-tested with k6 from Dallas (DFW) against a Native Hosting sandbox store with temporary Server-Timing instrumentation on with-routes. The instrumentation isn't part of this PR; it's on jorgemoya/kv-lookup-timing. Guest pages only, all 1,385 requests per run returned 200.

Cloudflare KV read (ms) p50 p90 p95 mean
steady 4 req/s, before 81 121 135 87.8
steady 4 req/s, after 8 106 143 51.2
each page every 75s, before 84 136 151 89.8
each page every 75s, after 6 14 63 12.8
key not in KV, before → after 87 → 74 125 → 106 136 → 118 86.6 → 78.5

The steady run's tail is the central refetch each key makes once per 5 minutes; that run also started right after a redeploy, with an empty location cache. In steady traffic about 97% of lookups are answered by the L1, so averaged over every request the KV cost went from about 2.9ms to 1.8ms. The larger gain is on low-traffic pages and fresh isolates.

pnpm vitest run   # in core
Tests  211 passed (211)

Migration

See the changeset. In core/lib/kv/adapters/cloudflare-kv.ts, RoutesKvNamespace.get now takes { type: 'json', cacheTtl } instead of 'json', and mget passes cacheTtl: 300.

🤖 Generated with Claude Code

… minutes

Workers KV caches reads at each location for 60s by default, the same
window as the in-process L1. The L1 therefore re-reads exactly as the
location's copy expires, so load testing showed nearly every read going
to central storage (~88ms mean, cold and warm alike). Read with a 300s
cacheTtl so L1 re-reads hit the location cache instead.

Refs LTRAC-2154
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jorgemoya
jorgemoya requested a review from a team as a code owner October 1, 2026 23:04
@vercel

vercel Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
catalyst Ready Ready Preview Oct 2, 2026 3:17pm UTC

Request Review

@changeset-bot

changeset-bot Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: e2cc228

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 1 package
Name Type
@bigcommerce/catalyst-core Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Unlighthouse Performance Comparison — Vercel

Comparing PR preview deployment Unlighthouse scores vs production Unlighthouse scores.

Summary Score

Aggregate score across all categories as reported by Unlighthouse.

Prod Desktop Prod Mobile Preview Desktop Preview Mobile
Score 98 99 99 99

Category Scores

Category Prod Desktop Prod Mobile Preview Desktop Preview Mobile
Performance 100 100 100 100
Accessibility 95 98 95 95
Best Practices 100 100 100 100
SEO 88 100 100 100

Core Web Vitals

Metric Prod Desktop Prod Mobile Preview Desktop Preview Mobile
LCP 0.3 s 1.7 s 0.4 s 0.4 s
CLS 0 0 0 0
FCP 0.3 s 0.4 s 0.4 s 0.4 s
TBT 0 ms 0 ms 0 ms 30 ms
Max Potential FID 40 ms 40 ms 50 ms 80 ms
Time to Interactive 0.3 s 0.4 s 0.4 s 0.6 s

Full Unlighthouse report →

…ache

A value read just before its expiryTime stays in the location cache for
the full cacheTtl, so a storefront status change can take up to ~11
minutes to reach a location instead of ~7. Note it in the adapter and the
changeset rather than shortening the status window, which every KV
adapter shares.

Refs LTRAC-2154
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread core/lib/kv/adapters/cloudflare-kv.ts Outdated
Comment on lines +32 to +41
// How long each Cloudflare location caches a read. The default (60s) matches
// the in-process L1 window, so the L1 always re-reads just as the location's
// copy expires and nearly every read goes to central storage. Outlasting the L1
// lets those re-reads hit the location cache. Kept at the 5-minute storefront
// status window, the shortest freshness window `with-routes` stores.
//
// Trade-off: a value read just before its `expiryTime` stays cached here for
// the full TTL, so a storefront status change (maintenance, launch) can take up
// to ~11 minutes to reach a location instead of ~7. A refresh write is usually,
// but not guaranteed to be, visible at once in the location that made it.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we shorten this to make sure the comment is providing good value?

Refs LTRAC-2154
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jorgemoya
jorgemoya enabled auto-merge October 2, 2026 15:19
@jorgemoya
jorgemoya added this pull request to the merge queue Oct 2, 2026
Merged via the queue into canary with commit 24dacae Oct 2, 2026
18 checks passed
@jorgemoya
jorgemoya deleted the LTRAC-2154/perf/kv-read-cache-ttl branch October 2, 2026 15:31

This branch was successfully deployed

1 active deployment
Preview — e2cc228f Deployed Oct 2, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants