Skip to content

Evaluate flags from the replicator cache when the replicator is not ready - #114

Closed
ryanechternacht wants to merge 1 commit into
mainfrom
ryan/replicator-checkflags-not-ready
Closed

ryanechternacht wants to merge 1 commit into
mainfrom
ryan/replicator-checkflags-not-ready

Conversation

@ryanechternacht

Copy link
Copy Markdown
Member

Problem

In replicator mode the SDK polls the replicator's health endpoint and treats ready as "connected". When a customer's Schematic account is closed or Schematic is unreachable, the replicator stays up with its Redis cache intact but reports ready: false.

Today, when that happens:

  • checkFlag / checkFlagWithEntitlement skip the datastream (tryDatastreamCheckFlag returns null when isConnected() is false) and go to the API path: the local flag result cache, then features().checkFlag. With Schematic unreachable or the account closed, that call fails and the flag default is returned.
  • checkFlags has the same isConnected() gate, so it goes to the result cache plus the bulk features().checkFlags call, and returns defaults for every key when that fails.

The Go SDK (reference) does not gate flag checks on readiness. useDataStream() only checks that a datastream client exists, and in replicator mode entity resolution evaluates with whatever the cache holds.

A second issue made this worse. The replicator answers 503 with its JSON body while not ready, and checkReplicatorHealth only read the body on 2xx. An SDK that started while the replicator was not ready never learned cache_version, so it built cache keys under the local rules engine version instead of the replicator's. Go reads the body regardless of status.

Change

  • Schematic: new canEvaluateViaDatastream() used by both tryDatastreamCheckFlag and the checkFlags datastream path. In replicator mode it returns true whenever a datastream client exists. In direct WebSocket mode it still requires isConnected(), so WebSocket behavior is unchanged. A flag missing from the cache, or any evaluation error, still falls back to the API as before.
  • DataStreamClient.checkReplicatorHealth: reads and parses the JSON body for any status, updates the cache version from it (keeping the last known version when none is reported), and still treats non-2xx or ready: false as not ready.

Test

New ReplicatorNotReadyTest:

  • Replicator mode with an unreachable health URL (never ready), flag and company cached: checkFlag, checkFlagWithEntitlement, and checkFlags return the cached evaluation and never touch the API client.
  • Health endpoint returning 503 with {"ready": false, "cache_version": "v42"}: the client stays not connected but adopts v42 for cache keys.

All three fail on main and pass with this change. ./gradlew compileJava spotlessCheck test passes on JDK 11 (250 tests).

🤖 Generated with Claude Code

…eady

In replicator mode, checkFlag, checkFlagWithEntitlement, and checkFlags
skipped the datastream whenever the replicator's health endpoint reported
ready: false, and went to the API instead. The replicator keeps its Redis
cache when it loses its connection to Schematic, so those checks now
evaluate from the cache regardless of readiness, matching the Go SDK. A
flag missing from the cache still falls back to the API. Direct WebSocket
mode is unchanged and still requires an active connection.

The replicator answers 503 with a JSON body while not ready, and the
health check only read the body on 2xx. An SDK that started while the
replicator was not ready never learned cache_version and built cache keys
under the wrong version. The health check now reads the body regardless
of status.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@ryanechternacht

Copy link
Copy Markdown
Member Author

Closing in favor of fixing this in the Replicator. The SDK is right to gate flag checks on Replicator readiness; the Replicator should keep reporting ready in /health once its cache is fully loaded, including after it loses its connection to Schematic. A Replicator PR is coming for that. The separate track optimistic-update PR stays open.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants