Serve replicator-mode flag checks only once the cache is ready - #116
Open
ryanechternacht wants to merge 2 commits into
Open
ryanechternacht wants to merge 2 commits into
ryanechternacht wants to merge 2 commits into
Conversation
Read the replicator health body whatever the HTTP status, so a 503 from /ready still records cache_version, and set readiness from the ready field. A failed poll sets not ready and keeps the last known cache version. Add isCacheReady() on DataStreamClient and Schematic. Single and bulk flag checks both gate on it: before the replicator reports ready they skip the cache and use the API; once ready they evaluate from the cache with the existing API fallback. isConnected() and isDatastreamConnected() are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e outside replicator mode Align with SchematicHQ/schematic-go#240. isCacheReady() now returns true outside replicator mode. Single and bulk flag checks share one helper, useDataStreamCache(), which is the existing datastream check (configured and connected, so WebSocket mode is unchanged) plus isCacheReady(). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2 of 3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Depends on SchematicHQ/schematic-replicator#143. Do not merge before it. That PR changes
/readysoready: true(HTTP 200) means the cache is complete for itscache_version. Before the load completes it returns 503 withready: falseand acache_versionin the body.Go counterpart: SchematicHQ/schematic-go#240.
Problem
In replicator mode, single (
checkFlag,checkFlagWithEntitlement) and bulk (checkFlagswith keys) flag checks already gated onisConnected(), which in replicator mode is replicator readiness. Both went to the API while not ready and to the cache once ready, so their outcomes did not differ. Two things were wrong around that gate:checkReplicatorHealthonly read the response body on a 2xx. A 503 from/readycarriesready: falseandcache_version, so an SDK that started while the replicator was loading never recorded the cache version. When the replicator then turned ready, the SDK could read cache keys under the wrong version until the next poll.dataStreamClient != null && dataStreamClient.isConnected()), so nothing kept them in step, and there was no clearly named way to ask whether the replicator cache is ready to serve.The health check also only caught
IOException. Any other exception escaping the scheduled task would have stopped the health poll for good and left readiness stuck.Change
DataStreamClient.checkReplicatorHealthparses the JSON body whatever the HTTP status. Readiness comes from thereadyfield.cache_versionis recorded from any response with a non-empty one. A failed poll (connection error, timeout, unparseable body, or any other exception) sets not ready and keeps the last known cache version. The fail-closed behavior on an unreachable replicator is unchanged; it is an open question in #143.DataStreamClient.isCacheReady(). In replicator mode it returns replicator readiness from the health poll. Outside replicator mode it returns true, as in Go, since the SDK fills its own cache over the WebSocket and there is nothing to wait for.Schematic.isCacheReady()exposes it next toisDatastreamConnected().Schematic.useDataStreamCache(), the one gate that single and bulk flag checks both call (tryDatastreamCheckFlagand the bulk datastream branch). It is the existing datastream check (configured and connected) plusisCacheReady(). Keeping theisConnected()term means WebSocket mode is unchanged: Go's datastream check never required a connection, while Java's always has. Not ready: the cache is not read and the existing API path runs, falling back to the flag default if the API fails. Ready: flags evaluate from the cache as before, including the existing API fallback when a flag is missing from the cache or evaluation errors.isConnected()andisDatastreamConnected()behave as before. Their doc comments now say that in replicator mode they report replicator readiness, and point toisCacheReady().Java has no equivalent of Go's credit-lease prewarm. The only other cache read in
Schematicistrack's local metric update, which this PR does not touch. #115 changes that gate in the same file, and the two branches merge cleanly. The publicDataStreamClient.getCachedFlag/Company/Useraccessors are not flag checks and stay ungated.One existing difference between single and bulk is left as is, since Go does the same: a single check served from the cache enqueues a
flag_checkevent, and a bulk check does not.Tests
New
ReplicatorCacheReadyTest. It seeds the cache under the replicator's key layout (flags:{cache_version}:{key}, and the company by ID plus a key lookup) and serves a health endpoint from a local HTTP server:{"ready": false, "cache_version": "vX"}sets not ready and recordsvX, both at startup and after having been ready.cache_versionkeeps the previous one.isCacheReady()is true outside replicator mode, and a WebSocket-mode client that is not connected still sends single and bulk checks to the API even with the flag cached.I checked that the not-ready tests fail when the gate is removed, that the 503 test fails when the body is only read on 2xx, and that the WebSocket test fails if the gate drops the
isConnected()term../gradlew compileJava spotlessCheck testpasses on JDK 11 (257 tests, 0 failures).🤖 Generated with Claude Code