Skip to content

Serve replicator-mode flag checks only once the cache is ready - #116

Open
ryanechternacht wants to merge 2 commits into
mainfrom
ryan/replicator-cache-ready-gate
Open

ryanechternacht wants to merge 2 commits into
mainfrom
ryan/replicator-cache-ready-gate

Conversation

@ryanechternacht

@ryanechternacht ryanechternacht commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Depends on SchematicHQ/schematic-replicator#143. Do not merge before it. That PR changes /ready so ready: true (HTTP 200) means the cache is complete for its cache_version. Before the load completes it returns 503 with ready: false and a cache_version in the body.

Go counterpart: SchematicHQ/schematic-go#240.

Problem

In replicator mode, single (checkFlag, checkFlagWithEntitlement) and bulk (checkFlags with keys) flag checks already gated on isConnected(), which in replicator mode is replicator readiness. Both went to the API while not ready and to the cache once ready, so their outcomes did not differ. Two things were wrong around that gate:

  • checkReplicatorHealth only read the response body on a 2xx. A 503 from /ready carries ready: false and cache_version, so an SDK that started while the replicator was loading never recorded the cache version. When the replicator then turned ready, the SDK could read cache keys under the wrong version until the next poll.
  • Each check had its own copy of the gate (dataStreamClient != null && dataStreamClient.isConnected()), so nothing kept them in step, and there was no clearly named way to ask whether the replicator cache is ready to serve.

The health check also only caught IOException. Any other exception escaping the scheduled task would have stopped the health poll for good and left readiness stuck.

Change

  • DataStreamClient.checkReplicatorHealth parses the JSON body whatever the HTTP status. Readiness comes from the ready field. cache_version is recorded from any response with a non-empty one. A failed poll (connection error, timeout, unparseable body, or any other exception) sets not ready and keeps the last known cache version. The fail-closed behavior on an unreachable replicator is unchanged; it is an open question in #143.
  • New DataStreamClient.isCacheReady(). In replicator mode it returns replicator readiness from the health poll. Outside replicator mode it returns true, as in Go, since the SDK fills its own cache over the WebSocket and there is nothing to wait for. Schematic.isCacheReady() exposes it next to isDatastreamConnected().
  • New private Schematic.useDataStreamCache(), the one gate that single and bulk flag checks both call (tryDatastreamCheckFlag and the bulk datastream branch). It is the existing datastream check (configured and connected) plus isCacheReady(). Keeping the isConnected() term means WebSocket mode is unchanged: Go's datastream check never required a connection, while Java's always has. Not ready: the cache is not read and the existing API path runs, falling back to the flag default if the API fails. Ready: flags evaluate from the cache as before, including the existing API fallback when a flag is missing from the cache or evaluation errors.
  • isConnected() and isDatastreamConnected() behave as before. Their doc comments now say that in replicator mode they report replicator readiness, and point to isCacheReady().
  • README: new "Cache Readiness" section under Replicator Mode.

Java has no equivalent of Go's credit-lease prewarm. The only other cache read in Schematic is track's local metric update, which this PR does not touch. #115 changes that gate in the same file, and the two branches merge cleanly. The public DataStreamClient.getCachedFlag/Company/User accessors are not flag checks and stay ungated.

One existing difference between single and bulk is left as is, since Go does the same: a single check served from the cache enqueues a flag_check event, and a bulk check does not.

Tests

New ReplicatorCacheReadyTest. It seeds the cache under the replicator's key layout (flags:{cache_version}:{key}, and the company by ID plus a key lookup) and serves a health endpoint from a local HTTP server:

  • Not ready (503): single and bulk both skip the cache and return the mocked API's values, which are set opposite to the cached ones.
  • Not ready, API failing: single and bulk both return the SDK flag defaults.
  • Ready: single and bulk both evaluate from the cache with no API call and agree on value, reason, flag ID and company ID.
  • Ready, flag not in the cache: single and bulk both fall back to the API.
  • A 503 {"ready": false, "cache_version": "vX"} sets not ready and records vX, both at startup and after having been ready.
  • An unreachable health URL sets not ready and keeps the previous cache version. The same holds for an unparseable body. An empty cache_version keeps the previous one.
  • isCacheReady() is true outside replicator mode, and a WebSocket-mode client that is not connected still sends single and bulk checks to the API even with the flag cached.

I checked that the not-ready tests fail when the gate is removed, that the 503 test fails when the body is only read on 2xx, and that the WebSocket test fails if the gate drops the isConnected() term.

./gradlew compileJava spotlessCheck test passes on JDK 11 (257 tests, 0 failures).

🤖 Generated with Claude Code

Read the replicator health body whatever the HTTP status, so a 503
from /ready still records cache_version, and set readiness from the
ready field. A failed poll sets not ready and keeps the last known
cache version.

Add isCacheReady() on DataStreamClient and Schematic. Single and bulk
flag checks both gate on it: before the replicator reports ready they
skip the cache and use the API; once ready they evaluate from the
cache with the existing API fallback. isConnected() and
isDatastreamConnected() are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@ryanechternacht
ryanechternacht requested a review from a team as a code owner September 30, 2026 15:07
…e outside replicator mode

Align with SchematicHQ/schematic-go#240. isCacheReady() now returns true
outside replicator mode. Single and bulk flag checks share one helper,
useDataStreamCache(), which is the existing datastream check (configured
and connected, so WebSocket mode is unchanged) plus isCacheReady().

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant