Skip to content

Serve replicator-mode flag checks only once the cache is ready - #135

Open
ryanechternacht wants to merge 3 commits into
mainfrom
ryan/replicator-cache-ready-gate
Open

ryanechternacht wants to merge 3 commits into
mainfrom
ryan/replicator-cache-ready-gate

Conversation

@ryanechternacht

@ryanechternacht ryanechternacht commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Depends on SchematicHQ/schematic-replicator#143. Do not merge before it. #143 changes /ready so ready: true (200) means the cache is complete for its cache_version, and a replicator still loading answers 503 with ready: false and the cache_version it is loading.

Problem

In replicator mode (async client only; the sync client has no DataStream), single and bulk flag checks were gated differently:

  • check_flag / check_flag_with_entitlement evaluated from the replicator cache with no readiness gate, so they could serve from a partially loaded cache.
  • check_flags with keys only used the cache when ds.is_connected() was true, which in replicator mode is replicator readiness.

The health poll also called raise_for_status() before reading the body, so a 503 from /ready was treated as a failed poll and its cache_version was never recorded.

Change

  • Health poll (DataStreamClient._check_replicator_health) reads the JSON body whatever the HTTP status, sets readiness from the ready field, and records cache_version from any response that has a non-empty one. A failed poll (connection error, timeout, unparseable body) sets not ready and keeps the last known cache_version. get_replicator_cache_version_async also reads the body on a 503 now.
  • DataStreamClient.is_cache_ready() is new. In replicator mode it returns the replicator readiness above; outside replicator mode it returns true, since websocket mode fills and fetches its own cache.
  • is_connected() is unchanged, with a docstring saying that in replicator mode it reports replicator readiness and pointing to is_cache_ready().
  • One shared gate, AsyncSchematic._get_flag_check_datastream(), returns the DataStream client only when is_cache_ready() is true. check_flag_with_entitlement (and so check_flag), check_flags with keys, and the client-mode credit lease path in check() all use it (see the divergence note below). While the cache is not ready they take the existing API path, which falls back to the flag default if the API fails. Once it is ready they evaluate from the cache exactly as before, including the existing API fallback when evaluation errors (for example, flag not in cache).
  • Websocket mode is unchanged. check_flags keeps its existing extra is_connected() check there; in replicator mode that value equals is_cache_ready().
  • README's Replicator Mode section gains a Cache Readiness subsection.

Parity with Go

The Go counterpart is SchematicHQ/schematic-go#240, and this matches it. is_cache_ready() returns replicator readiness in replicator mode and true otherwise, so websocket mode is unchanged. One shared helper (_get_flag_check_datastream(), which returns the datastream only if one is configured and is_cache_ready() is true) replaces the datastream check at both the single and bulk flag-check branches, like Go's useDataStreamCache().

One divergence: check() with usage in client credit-lease mode also goes through the gate. It is a credit-aware flag check that evaluates the flag and company from the DataStream cache, so while the replicator cache is not ready it runs as a plain check through the API and takes no lease. Go's checkWithClientLease in client/check.go still reads the cache directly. This is documented in the README.

Like Go, prewarm is left ungated. It is not a flag check: it resolves a company ID through the DataStream (_resolve_company_id_with_wait) to acquire leases early, and still reads the cache without waiting for readiness.

Overlap with #134

#134 changes the track gate (_update_company_metrics). This PR does not touch track, and the two branches merge cleanly.

Tests

New tests/custom/test_replicator_cache_ready.py runs AsyncSchematic in replicator mode against RedisCache over fakeredis, seeded with JSON under the versioned keys the replicator writes, and a scripted /ready endpoint via httpx.MockTransport:

  • Not ready: single and bulk skip the cache (a DataStream check_flag spy is never called), return the API's values, and return flag defaults when the API fails. The same holds before any health response and after readiness is lost.
  • Ready: single and bulk evaluate from the cache through the real rules engine with no API call, and return identical results. A flag missing from the cache still falls back to the API.
  • A 503 with {"ready": false, "cache_version": "vX"} sets not ready and records vX. An unreachable URL and an unparseable body set not ready and keep the previous cache_version. An empty cache_version keeps the last one. The health-changed callback fires on transitions only.

A lease-client test in tests/custom/test_client.py checks that check() with usage runs as a plain API check and acquires no lease while the cache is not ready. The existing 503 health test was updated to the new semantics. poetry run mypy . passes, and poetry run pytest -n auto . passes (774 passed, 3 skipped).

🤖 Generated with Claude Code

ryanechternacht and others added 3 commits September 30, 2026 13:54
In replicator mode, check_flag and check_flag_with_entitlement read the
shared cache with no readiness gate, while check_flags with keys gated on
is_connected(). The health poll also called raise_for_status(), so a 503
from /ready never had its body read and its cache_version was dropped.

- Health poll reads the JSON body on any status, sets readiness from the
  ready field, records any non-empty cache_version, and on a failed poll
  sets not ready and keeps the last cache_version.
- Add DataStreamClient.is_cache_ready(). is_connected() is unchanged and
  documented as reporting replicator readiness in replicator mode.
- Single, bulk and the client-mode credit lease check share one gate,
  AsyncSchematic._get_flag_check_datastream(): while the cache is not
  ready they take the API path.
- README describes the readiness behavior.

Depends on SchematicHQ/schematic-replicator#143.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The test from #134 expected check_flag to read the cache while the
replicator is not ready. Flag checks now skip the cache until it is ready,
so the test checks the API path while not ready and the tracked usage from
the cache once the replicator reports ready.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant