Skip to content

Keep optimistic usage updates in track when the replicator is not ready - #202

Merged
bpapillon merged 1 commit into
mainfrom
ryan/replicator-track-offline
Sep 30, 2026
Merged

bpapillon merged 1 commit into
mainfrom
ryan/replicator-track-offline

Conversation

@ryanechternacht

Copy link
Copy Markdown
Member

Problem

track enqueues the event and then optimistically bumps the cached company metric via updateCompanyMetrics. That second step was gated on datastreamClient.isConnected(), which in replicator mode is the health endpoint's ready field. When a customer's Schematic account is closed or Schematic is unreachable, the replicator stays up with its Redis cache intact but reports ready: false. Flag checks keep evaluating from that cache, but usage stopped counting locally, so numeric limits never moved for as long as the replicator was not ready. The event was still enqueued to the API either way.

What I found

History of the gate:

Correctness checks:

  • Missing cache: updateCompanyMetrics reads the company from the cache and returns without writing if it is not there, so a not ready replicator with an empty or partial cache is a no-op. The test covers this.
  • Double counting: the bump is a local prediction. Full company messages overwrite the cached company, and partial messages upsert metrics by (event subtype, period, month reset) with the server's absolute value, so whatever the server sends next replaces the local figure instead of adding to it. While the replicator is not ready nothing is sent, so there is no race with a server update. When it recovers, its resync rewrites companies with server values. Duplicate events are already handled separately: trackWithReservation passes updateMetrics = false when the settle did not land locally (Close the lease-client gaps the port reviews found #194).

So I found no correctness reason for the gate.

One related issue this does not change: updateCompanyMetrics rewrites the company and its lookup keys with the SDK's cacheTTL (24 hours by default), while the replicator writes them with no TTL by default. That already happens today whenever track runs with the replicator ready, including for events that match no metric (the company is rewritten regardless). With this change it can also happen while the replicator is not ready, and during an outage longer than cacheTTL, a company that was tracked and then went idle for cacheTTL would drop out of Redis until the replicator resyncs. I'd suggest a follow-up so the metric write in replicator mode keeps the existing expiry and skips the write when no metric matched.

Change

The isConnected() condition is removed from the metric update in emitTrack, so the local update runs whenever a DataStream client exists, in both modes. In WebSocket mode a disconnected client also keeps evaluating cached companies in checkFlag, so skipping the bump there would under count in the same way. The cache is the SDK's own in that mode, so the TTL point above does not apply.

Test

New tests/unit/replicator/track-not-ready.test.ts (and a .fernignore entry for the directory). It uses a real DataStream client in replicator mode, the real WASM rules engine, and the in-memory fake Redis seeded the way the replicator writes it (snake_case JSON under versioned keys). fetch is stubbed so the health endpoint returns { ready: false, cache_version: "v-test" }.

  • track with quantity 2 moves the cached api_call metric from 3 to 5.
  • A flag gated on api_call >= 5 evaluates false, stays false after one track, and flips to true after the second, with the API mocked to fail.
  • Tracking a company that is not cached leaves Redis unchanged.

The first two fail on main and pass with this change. yarn build and yarn test pass (101 suites; the Redis integration suite is skipped locally without TEST_REDIS_URL).

This branch and #201 each add the same tests/unit/replicator/ line to .fernignore at the same spot, so they merge cleanly in either order.

🤖 Generated with Claude Code

track only bumped the cached company metric when isConnected() was true.
In replicator mode that is the health endpoint's ready field, which goes
false when Schematic is unreachable or the account is closed while the
replicator's Redis cache stays intact. Flag checks keep evaluating from
that cache, so usage stopped counting locally and numeric limits stopped
moving. The event itself was still enqueued either way.

The gate has no recorded rationale. It came from the Go SDK's optimistic
metric updates, written before replicator mode existed, when "connected"
was the only signal that the DataStream was in use at all. Replicator mode
later redefined isConnected() as replicator readiness without revisiting
it. The bump only touches a company already in the cache, and the next
company the server sends replaces the metric value instead of adding to
it, so running it while not ready cannot double count.

Drop the isConnected() gate so the local update runs whenever a DataStream
client exists, in both modes. In WebSocket mode a disconnected client also
keeps evaluating cached companies, so the same reasoning applies.

Add a replicator mode test with a real DataStream client, fake Redis
seeded in the replicator's key layout, the real WASM engine, and a health
endpoint reporting ready: false. track bumps the cached metric, a
metric-gated flag flips once the threshold is reached, and a company that
is not cached is left alone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bpapillon
bpapillon merged commit 2436d1c into main Sep 30, 2026
7 checks passed
@bpapillon
bpapillon deleted the ryan/replicator-track-offline branch September 30, 2026 15:29
ryanechternacht added a commit that referenced this pull request Sep 30, 2026
The note from #203 said bulk flag checks bypass the shared cache and track
skips local metric updates until the replicator is ready. Track no longer
skips them (#202), and single and bulk checks now gate the same way.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ryanechternacht added a commit that referenced this pull request Sep 30, 2026
The test from #202 expected a metric-gated flag check to flip from the
cache while the replicator is not ready. Flag checks now skip the cache
until it is ready, so the test tracks while not ready and checks the flag
after the replicator reports ready.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants