Skip to content

refactor: consume Planner-owned physical DAG operators - #770

Merged
zzylol merged 0 commit into
test/promql-exact-function-coveragefrom
feat/shared-operator-foundation
Sep 26, 2026
Merged

zzylol merged 0 commit into
test/promql-exact-function-coveragefrom
feat/shared-operator-foundation

Conversation

@zzylol

@zzylol zzylol commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Dependency stack: main → #768 → #737 → #749 → #771 → #770 → #763 → #765 → #761 → #728 → #742 → #759

Independent follow-ups to #765: #756 (diagnostics), #766 (runtime controls and overhead inspection).

Before this PR

Shared execution and mathematical kernels lived in the backend, while Planner's IR could change independently. Candidate pruning and grouped ranking still had dedicated backend representations.

After this PR

Consume asap-physical-operators and asap_sketch_codec from Planner #462, pinned with Planner IR to revision 645cacb04a451e25147f0d23aaddd05da16b4746. Remove the backend copies. The library depends on Planner types and has no backend dependency.

The API follows Planner #461: phase belongs to node data state, candidate filtering uses a general semi-join, and grouped ranking composes Sort → grouped Limit. Old operator dispatch and wire compatibility paths are removed. Exact accumulator state is explicitly finalized before value consumers. Temporal exact TopK candidates also lower to grouped Sort → Limit, including integer Count scores.

The independent runtime owns shared producers, bounded delivery, cancellation and retained-output accounting. Native computation includes expressions, relations, window reductions and supported summary operations. The same implementation can run at ingestion time or query time. Custom summary values need not fit Arrow RecordBatch.

Backend #763 integrates ingestion and durable publication; #765 integrates query execution and completes relation/temporal/stored-state computation migration. Sources, storage and protocol conversion remain deployment responsibilities. Local raw Scan is deferred; blocking operators have no spill support.

Architecture and DataFusion comparison. Independent library tests run in Planner; backend tests exercise deployment bindings.

Validation

Planner: 246 shared-library unit tests, 18 library integration tests, one doc test, 432 mapping tests and the integration package passed. Both repositories pass strict workspace/all-target Clippy.

Backend stack: 118 type and 444 control-plane library tests passed. The data-plane run passed 925 tests; its remaining persistence synchronization test was fixed and passed separately. All 19 compatibility process tests and 59 other integration tests passed. Tests cover shared SQL sources, cancellation, stored-state coverage, temporal Sort/Limit, HLL/KLL confidence, and ingestion publication/recovery.

The mandatory #754 gate still fails on grouped-temporal-Sum producer shape and quantile-ratio local execution. Those assertions remain intact. Full #759 performance acceptance is not established. Local raw Scan remains explicitly deferred.

The publication process test verifies that a successor is cold before its own data arrives, rejects old-generation frames and returns the successor's new values; it does not assume cross-version payload reuse.

Planner #462 integration update

Pinned all Planner crates to b58f24268270a4dd7b7caaf0076ea3db842f3e9b from ASAPPlanner #462 and sketchlib to 5f03ccbd. Consumption uses physical_planner, shared summary kernels and the current wire-state shapes. Unsupported counter-weighted heap frontiers are reported to Planner before selection; raw counter values are not silently substituted for rate values. Validation: 429 control-plane library tests and data-plane all-target compilation passed.

Physical planning and deployment contract.

Physical planning integration status

Depends on ASAPPlanner #462 (991ef5c3032507e8913bc6f3644a4008c61435b7). Planner owns physical candidate compilation and workload-cost winner selection; backend owns deployment binding and operation. The reordered stack puts operator and accuracy support before the Level 1 test PR so each test branch can be validated independently.

Validation remains in progress. Level 1 passed on its own branch before the final counter-window dependency update. Level 2 currently passes six of seven live suites; the remaining sparse counter-window case is being repaired and rerun. Level 3 peak-memory acceptance cannot be established on the local Linux 5.15 host because memory.peak is unavailable; CI evidence is required. General stored scalar/result frontiers remain integration work, and these results do not establish that support.

@zzylol zzylol changed the title feat: independent shared physical operator DAG library refactor: consume Planner-owned physical DAG operators Sep 23, 2026
@zzylol
zzylol merged commit e9dc0af into test/promql-exact-function-coverage Sep 26, 2026
@zzylol
zzylol force-pushed the test/promql-exact-function-coverage branch from 3159bdf to ef58ad2 Compare September 26, 2026 04:42
@zzylol
zzylol force-pushed the feat/shared-operator-foundation branch from fef20a3 to b8ef5d4 Compare September 26, 2026 04:42
zzylol added a commit that referenced this pull request Sep 28, 2026
…774)

* docs: clarify Planner physical plan and SDS architecture

* docs: specify executable subplan materialization boundaries

* docs: scope migration to backend precompute and query plans

* docs: clarify window terminology migration

* Revert "docs: clarify window terminology migration"

This reverts commit daa5281.

* docs: focus physical plan and SDS designs

* docs: add physical compiler input example

* docs: add physical compiler output example

* docs: include query expression in compiler example

* docs: reorganize physical plan integration design

* docs: reorganize SDS and migration designs

* docs: define maintenance inputs before plan split example

* docs: use plan version consistently in backend design

* docs: remove standalone catalog materialization abstraction

* docs: add concise planner backend glossary

* refactor: split installed maintenance DAGs from query execution

* refactor: remove backend Collector dependency and normalize legacy DAGs

* test: restore whole-backend process coverage without Collector

* fix: migrate Planner main and reject legacy runtime artifacts

* test: send full sketch envelope in whole-backend E2E

* refactor: bind backend summary state through versioned SDS slots

* test: send complete sketch envelopes in process fixtures

* docs: clarify selected deployment guarantee terminology

* docs: explain missing planner maintenance guarantee

* docs: motivate selected producer maintenance decision

* docs: label catalog reads and SDS metadata ownership

* docs: align SDS ownership and lifecycle terminology

* docs: use current planner and plan-version names consistently

* docs: align integration diagram with SDS ownership

* docs: clarify instance identity and shared producer wording

* refactor: align plan bindings with current Planner and SDS contract

* docs: distinguish summary definitions from runtime stores

* docs: model one runtime summary store for DAG bindings

* docs: scope SDS lifecycle to read eligibility

* docs: tie stored summary examples directly to DAG outputs

* docs: name summary tables, stored records, and output references by role

* docs: illustrate summary definitions, stored records, and output references

* docs: limit v1 summary storage to definitions and stored summaries

* refactor: align plan bindings with summary store v1

* fix: validate selected DAG provenance

* fix: retain neutral sketch codec dependencies when syncing main

* fix: align derived DAG validation with current schema versions

* fix: keep maintenance document version distinct from complete DAG version

* refactor: adopt costed Planner selection without legacy API adapters

* test: verify exact process routing for uncertified Planner candidates

* docs: remove redundant Planner selection contract

* deps: pin Planner bounded HLL confidence model

* test: size transmitted KLL state from a certified accuracy contract

* test: retain KLL collector capability when using theoretical confidence

* refactor: implement typed summary semantics and physical plan lowering

* docs: separate Planner physical computation from backend deployment

* refactor: name the backend orchestration entry DeploymentPlanCompiler

* Restack PR #770 with implementation before standalone acceptance

* chore: consume Planner physical precompute candidate interfaces

* chore: consume Planner materialization frontier enumeration

* chore: consume exact temporal ranking physical candidates

* chore: use shared Planner candidate winner selection

* chore: consume sparse counter shared readout contracts

* test: declare collector ranking fixture capabilities explicitly

* Use Planner counter-window candidate execution

* Use shared keyed-counter omission contract

* Use Planner exact-counter population omission

* Construct fixture evidence for the standalone operator foundation

* build: pin Planner exact-state scratch merge implementation

* build: pin shared finalized-pane reconstruction fix

* docs: bind Planner physical DAGs without backend re-lowering

* docs: describe summary inputs with groups and pane duration

* docs: define summary semantic completeness beyond input scope

* docs: define SDS identity through canonical Planner computation

* docs: decouple SDS semantic identity from executable Planner IR

* docs: define Planner-owned SDS discovery for future ad hoc queries

* docs: streamline SDS design around definitions and stored results

* docs: track bound-query SDS migration across implementation PRs

* docs: identify active shared-library PR in bound-query migration

* build: align shared Planner dependencies with remote PR 462

* docs: separate bound SDS range lookup from state validation

* feat(sds): separate bound output routing from persisted semantic identity

* docs: state SDS migration responsibilities without stale implementation claims

* refactor(sds): keep semantic variants compact without changing wire format

* test: align evidence fixtures with selected exact count state

* docs: align precompute SDS description with semantic catalog

* test: bind imported-state fixtures to their actual semantic definition

* test: distinguish imported CMS transport from total-count planning

* docs: clarify Planner physical plan and SDS architecture

* docs: specify executable subplan materialization boundaries

* docs: scope migration to backend precompute and query plans

* docs: clarify window terminology migration

* Revert "docs: clarify window terminology migration"

This reverts commit daa5281.

* docs: focus physical plan and SDS designs

* docs: add physical compiler input example

* docs: add physical compiler output example

* docs: include query expression in compiler example

* docs: reorganize physical plan integration design

* docs: reorganize SDS and migration designs

* docs: define maintenance inputs before plan split example

* docs: use plan version consistently in backend design

* docs: remove standalone catalog materialization abstraction

* docs: add concise planner backend glossary

* docs: clarify selected deployment guarantee terminology

* docs: explain missing planner maintenance guarantee

* docs: motivate selected producer maintenance decision

* docs: label catalog reads and SDS metadata ownership

* docs: align SDS ownership and lifecycle terminology

* docs: use current planner and plan-version names consistently

* docs: align integration diagram with SDS ownership

* docs: clarify instance identity and shared producer wording

* docs: distinguish summary definitions from runtime stores

* docs: model one runtime summary store for DAG bindings

* docs: scope SDS lifecycle to read eligibility

* docs: tie stored summary examples directly to DAG outputs

* docs: name summary tables, stored records, and output references by role

* docs: illustrate summary definitions, stored records, and output references

* docs: limit v1 summary storage to definitions and stored summaries

* docs: separate Planner physical computation from backend deployment

* docs: bind Planner physical DAGs without backend re-lowering

* docs: describe summary inputs with groups and pane duration

* docs: define summary semantic completeness beyond input scope

* docs: define SDS identity through canonical Planner computation

* docs: decouple SDS semantic identity from executable Planner IR

* docs: define Planner-owned SDS discovery for future ad hoc queries

* docs: streamline SDS design around definitions and stored results

* docs: track bound-query SDS migration across implementation PRs

* docs: identify active shared-library PR in bound-query migration

* docs: separate bound SDS range lookup from state validation

* docs: state SDS migration responsibilities without stale implementation claims

* docs: preserve design index after rebasing onto main

* docs: clarify SDS source identity and version-scoped recovery

* docs: illustrate SDS identity and recovery decisions

* feat: bind Planner dataset semantics to deployment input identity

* fix: recognize shared native batch encoding at dependency boundary

* test: reject pre-dataset planning snapshot versions

* refactor: version dataset-bound catalog and update empty-plan fixtures

* docs: describe dataset-bound planning and installation inputs

* docs: clarify candidate selection and deployment ownership

* refactor: split backend deployment plans and bind stored outputs

* fix: restrict foundation SDS recovery to the installed generation

* fix: allocate fresh physical series when a plan version changes

* test: record foundation rebase and recovery regression evidence

* fix: complete shared state encoding adoption at the dependency boundary

* style: satisfy workspace formatting after the compiler rename

* feat(sds): separate bound output routing from persisted semantic identity

* feat(sds): require dataset-bound Planner definitions and versioned catalogs

* refactor: implement typed summary semantics and physical plan lowering

* fix: validate final SDS semantics across Planner adapters and recovery

* test: verify final SDS identity installation, serving and restart

* test: align shared execution fixtures with final #749 SDS foundation

* docs: record shared execution directly after the SDS foundation

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Remove obsolete SDS wire aliases and persisted metadata migration

* fix(data-plane): initialize dataset_identity in the bootstrap ingest contract

IngestContract gained an optional `dataset_identity`, but the bootstrap
envelope in data_plane's entry point was not updated, so the binary failed to
compile with E0063 and took the whole workspace build with it.

The bootstrap envelope describes the state before any plan is installed, where
every other field is a placeholder, so there is no dataset to bind to yet;
`None` is the accurate value. A published plan supplies the identity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Reject obsolete state-column syntax in compiler roundtrip coverage

* Remove unused materialization wire adapter and propagate recovery errors

* Keep projection fixtures on the canonical typed encoding

* Remove superseded deployment API field aliases

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit e627dbb)

* Document obsolete-format removals and cleanup validation

* Record passing restacked plan, storage and serving checks

* Remove generated evaluation archives and keep validation summaries

* Normalize validation summary formatting

* Remove PR process reports and defer execution design to shared-runtime PR

* Keep discovery and calibration on dataset-bound snapshot version 3

* Use typed deployment configuration in process E2E fixtures

* Align process assertions with dataset-bound SDS and cold successor activation

* fix: execute typed Planner fragments and preserve terminal resource errors

* fix: bind exact integer samples without losing input types

* test: verify exact integer protocol input binding

* refactor: name per-boundary physical fragments explicitly

* fix: consume finalized Planner query candidate outputs

* test: require Planner filters for bound protocol vectors

* test: use Planner schema lifting module

* fix: bind Planner filters and finalized query results

* docs: describe bound Planner filter execution

* fix: retain state binding and sharing beneath explicit query readouts

* fix: pin Planner query finalization for every candidate entry point

* docs: distinguish query values from stored accumulator boundaries

* test: assert exact Count state beneath its query readout

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Sep 29, 2026
…eads (#763)

* docs: remove redundant Planner selection contract

* deps: pin Planner bounded HLL confidence model

* test: size transmitted KLL state from a certified accuracy contract

* test: retain KLL collector capability when using theoretical confidence

* refactor: implement typed summary semantics and physical plan lowering

* docs: separate Planner physical computation from backend deployment

* refactor: name the backend orchestration entry DeploymentPlanCompiler

* Restack PR #770 with implementation before standalone acceptance

* Restack PR #763 with implementation before standalone acceptance

* chore: consume Planner physical precompute candidate interfaces

* chore: consume Planner materialization frontier enumeration

* chore: consume exact temporal ranking physical candidates

* chore: use shared Planner candidate winner selection

* chore: consume sparse counter shared readout contracts

* test: declare collector ranking fixture capabilities explicitly

* Use Planner counter-window candidate execution

* Use shared keyed-counter omission contract

* Use Planner exact-counter population omission

* Construct fixture evidence for the standalone operator foundation

* build: pin Planner exact-state scratch merge implementation

* build: pin shared finalized-pane reconstruction fix

* docs: bind Planner physical DAGs without backend re-lowering

* docs: describe summary inputs with groups and pane duration

* docs: define summary semantic completeness beyond input scope

* docs: define SDS identity through canonical Planner computation

* docs: decouple SDS semantic identity from executable Planner IR

* docs: define Planner-owned SDS discovery for future ad hoc queries

* docs: streamline SDS design around definitions and stored results

* docs: track bound-query SDS migration across implementation PRs

* docs: identify active shared-library PR in bound-query migration

* build: align shared Planner dependencies with remote PR 462

* docs: separate bound SDS range lookup from state validation

* feat(sds): separate bound output routing from persisted semantic identity

* docs: state SDS migration responsibilities without stale implementation claims

* refactor(sds): keep semantic variants compact without changing wire format

* test: align evidence fixtures with selected exact count state

* docs: align precompute SDS description with semantic catalog

* test: bind imported-state fixtures to their actual semantic definition

* test: bind imported-state fixtures to their actual semantic definition

* test: distinguish imported CMS transport from total-count planning

* docs: clarify Planner physical plan and SDS architecture

* docs: specify executable subplan materialization boundaries

* docs: scope migration to backend precompute and query plans

* docs: clarify window terminology migration

* Revert "docs: clarify window terminology migration"

This reverts commit daa5281.

* docs: focus physical plan and SDS designs

* docs: add physical compiler input example

* docs: add physical compiler output example

* docs: include query expression in compiler example

* docs: reorganize physical plan integration design

* docs: reorganize SDS and migration designs

* docs: define maintenance inputs before plan split example

* docs: use plan version consistently in backend design

* docs: remove standalone catalog materialization abstraction

* docs: add concise planner backend glossary

* docs: clarify selected deployment guarantee terminology

* docs: explain missing planner maintenance guarantee

* docs: motivate selected producer maintenance decision

* docs: label catalog reads and SDS metadata ownership

* docs: align SDS ownership and lifecycle terminology

* docs: use current planner and plan-version names consistently

* docs: align integration diagram with SDS ownership

* docs: clarify instance identity and shared producer wording

* docs: distinguish summary definitions from runtime stores

* docs: model one runtime summary store for DAG bindings

* docs: scope SDS lifecycle to read eligibility

* docs: tie stored summary examples directly to DAG outputs

* docs: name summary tables, stored records, and output references by role

* docs: illustrate summary definitions, stored records, and output references

* docs: limit v1 summary storage to definitions and stored summaries

* docs: separate Planner physical computation from backend deployment

* docs: bind Planner physical DAGs without backend re-lowering

* docs: describe summary inputs with groups and pane duration

* docs: define summary semantic completeness beyond input scope

* docs: define SDS identity through canonical Planner computation

* docs: decouple SDS semantic identity from executable Planner IR

* docs: define Planner-owned SDS discovery for future ad hoc queries

* docs: streamline SDS design around definitions and stored results

* docs: track bound-query SDS migration across implementation PRs

* docs: identify active shared-library PR in bound-query migration

* docs: separate bound SDS range lookup from state validation

* docs: state SDS migration responsibilities without stale implementation claims

* docs: preserve design index after rebasing onto main

* docs: clarify SDS source identity and version-scoped recovery

* docs: illustrate SDS identity and recovery decisions

* feat: bind Planner dataset semantics to deployment input identity

* fix: reject cross-version stored-state adoption

* fix: recognize shared native batch encoding at dependency boundary

* test: reject pre-dataset planning snapshot versions

* refactor: version dataset-bound catalog and update empty-plan fixtures

* docs: describe dataset-bound planning and installation inputs

* docs: clarify candidate selection and deployment ownership

* refactor: split backend deployment plans and bind stored outputs

* fix: restrict foundation SDS recovery to the installed generation

* fix: allocate fresh physical series when a plan version changes

* fix: reserve the persisted storage handle during recovery

* test: record foundation rebase and recovery regression evidence

* fix: complete shared state encoding adoption at the dependency boundary

* style: satisfy workspace formatting after the compiler rename

* feat(sds): separate bound output routing from persisted semantic identity

* feat(sds): require dataset-bound Planner definitions and versioned catalogs

* refactor: implement typed summary semantics and physical plan lowering

* fix: validate final SDS semantics across Planner adapters and recovery

* test: verify final SDS identity installation, serving and restart

* test: align shared execution fixtures with final #749 SDS foundation

* docs: record shared execution directly after the SDS foundation

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Remove obsolete SDS wire aliases and persisted metadata migration

* fix(data-plane): initialize dataset_identity in the bootstrap ingest contract

IngestContract gained an optional `dataset_identity`, but the bootstrap
envelope in data_plane's entry point was not updated, so the binary failed to
compile with E0063 and took the whole workspace build with it.

The bootstrap envelope describes the state before any plan is installed, where
every other field is a placeholder, so there is no dataset to bind to yet;
`None` is the accurate value. A published plan supplies the identity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Reject obsolete state-column syntax in compiler roundtrip coverage

* Remove unused materialization wire adapter and propagate recovery errors

* Keep projection fixtures on the canonical typed encoding

* Remove superseded deployment API field aliases

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit e627dbb)

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit e627dbb)

* Document obsolete-format removals and cleanup validation

* Record passing restacked plan, storage and serving checks

* Remove generated evaluation archives and keep validation summaries

* Normalize validation summary formatting

* Remove PR process reports and defer execution design to shared-runtime PR

* Keep discovery and calibration on dataset-bound snapshot version 3

* Use typed deployment configuration in process E2E fixtures

* Remove unused streaming-config fixture after physical-only startup

* Align process assertions with dataset-bound SDS and cold successor activation

* fix: execute typed Planner fragments and preserve terminal resource errors

* fix: bind exact integer samples without losing input types

* test: verify exact integer protocol input binding

* refactor: name per-boundary physical fragments explicitly

* fix: consume finalized Planner query candidate outputs

* test: require Planner filters for bound protocol vectors

* test: use Planner schema lifting module

* fix: bind Planner filters and finalized query results

* docs: describe bound Planner filter execution

* fix: retain state binding and sharing beneath explicit query readouts

* fix: pin Planner query finalization for every candidate entry point

* docs: distinguish query values from stored accumulator boundaries

* test: assert exact Count state beneath its query readout

* fix: preserve maintenance failures and stop failed worker admission

* docs: specify bounded revisions and consistent query snapshots

* feat: execute bounded precompute revisions with consistent SDS reads

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant