Conversation
…ieldDataType
Column -> Field, Schema.columns -> fields, SummaryFamilyType -> FieldDataType
(same variants); SummarySchema/SummaryField deleted. Post-ASAP node schemas
are built through Schema::lifted (no unique keys, closed). Adds the unified
operator IR module (ir::{OperatorNode, Operator, NonASAPOp, ASAPOp,
ScalarExpr}) and the lifecycle timing pass alongside the existing types.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…the unified IR; frontend-common resolver crate - ir::canonicalize / ir::cse port the pre-ASAP passes to Rc<OperatorNode>; CSE serves NonASAP and ASAP nodes alike and follows scalar references. - ir::export emits one post-ASAP node per operator (wire version 6): relational operators are Relational payloads, scalar references are ScalarRef edges, no embedded subtrees. - ir::timing writes execution timings from a LifecycleAssignment and validates every edge; nodes carry no timing before that pass. - asap-frontend-common owns the name-based tree front ends build (UnresolvedOp / UnresolvedScalar) and resolve_root, which binds names, derives schemas and canonicalizes into Rc<OperatorNode>. - ParsedWorkload roots are Rc<OperatorNode>. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One builder serves export / export_summary / export_post_asap; nodes are deduplicated by pointer, scalar operator references render as scalar_ref ids, every node carries structural_hash. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The SQL front end now emits ScalarExpr::{Exists, InSubquery, ScalarSubquery};
canonicalize rewrites them to the semi/anti/cross joins the planner saw before.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
SQL, PromQL and MetricsQL front ends build UnresolvedOp/UnresolvedScalar and return Rc<OperatorNode> from resolve_root. New: SQL Values (SELECT without FROM, VALUES), unary minus as Negative, ExprSemantics on every scalar comparison/arithmetic, uncorrelated EXISTS / IN / scalar subqueries as scalar expressions (lowered to joins by canonicalize); PromQL TimeRange.kind (instant vs range selector), the bool modifier (return_bool), scalar negation as Negative. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e unified IR (library)
Replacement::{Summary, Rewrite} collapse into Subtree(Rc<OperatorNode>);
assemble_residual becomes one rule (keep the operator, assemble its
children); nodes carry no timing, legality checks run through
ir::timing::validate_default and export through ir::export after
apply_lifecycle_timings. Test modules are migrated separately.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… kept subtree one node keep_pre_asap memoizes its exact-guarantee copy per input node, so a Scan kept by two candidates (AVG's SUM and COUNT branches) stays one Rc. The viewer contract test moves to crates/devtools/tests/viewer_contract.rs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Developer and architecture docs, the user guide and the viewer README now describe one OperatorNode DAG, lifecycle-assigned timing and wire version 6; the viewer renderer reads schema fields. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gressions Regressions caught by the migrated suites: - accuracy-evidence and runtime-support checks now cover summary decisions rooted in a relational operator above their readouts; - exact read boundaries keep the placement binding chose (query-time snapshot vs. maintenance feed); the timing pass honors it; - generic assembly keeps a value-computing operator (aggregate, binary op, window) exact over an approximate input instead of passing a guarantee through it; - table populations name their unified-IR input, so filtered SQL scans are recognized again. Adds the #468 acceptance tests (integration-tests/tests/operator_sharing.rs). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
QueryExpr, SummaryNode/SummaryExpr, ValueOperation, the wire-5 post-ASAP DAG, the old resolver, canonicalize and both CSE passes are gone; the unified ir module is the only operator language. pre_asap keeps the shared vocabulary (query_expr.rs renamed vocabulary.rs); post_asap keeps summary state and timing vocabulary. CSE compares a node's own fields with the signed-zero-aware check. cargo fmt --all. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Merge main into refactor/operator-flattening while retaining the unified OperatorNode graph proposed in #511. Port native physical compilation, aggregate filters, shared summaries and lifecycle placement to that graph. Remove ScalarBridge, preserve scalar workload roots as owned expressions, and lower mixed PromQL arithmetic and comparisons to Project/Filter with complete series identity and explicit vector/scalar conversions. Preserve native execution coverage and correct cross-selector maintenance eligibility and empty PromQL aggregation behavior. Validation: 1,569 workspace tests pass; workspace all-target/all-feature Clippy with warnings denied, rustfmt, and MetricsQL external consumer pass. Vendored MetricsQL baseline retains its main-branch sources and checker: lib baseline passes, but one existing doctest diagnostic hash differs under rustc 1.98.1 (same expected failing test set).
| use asap_types::pre_asap::query_expr::{GroupKeys, Source}; | ||
| use asap_types::pre_asap::schema::{Column, ColumnId, DataType, Schema}; | ||
| use asap_types::pre_asap::schema::{ColumnId, DataType, Field, Schema}; | ||
| use asap_types::pre_asap::vocabulary::{GroupKeys, Source}; |
There was a problem hiding this comment.
Why this is called vocabulary
There was a problem hiding this comment.
Does it reflect the actual content?
There was a problem hiding this comment.
“Vocabulary” was too vague and the module mixed different responsibilities. It contained operator parameter types (joins, windows, predicates, reductions, etc.), aggregate schema derivation, and IR errors. Split it into ir::operator_properties, ir::aggregate_schema, and ir::error; migrated imports and removed pre_asap::vocabulary. The modules now describe their actual contents and belong to the common IR. Change.
There was a problem hiding this comment.
ir::operator_properties, ir::aggregate_schema, and ir::error what does these three mean then?
There was a problem hiding this comment.
For your follow-up, these modules separate three concrete responsibilities:
ir::operator_propertiesdefines settings stored inside operator payloads: e.g.GroupKeyssays which columns an aggregate groups by,JoinKindselects inner/left/etc., andWindowFramedescribes a SQL window. It does not hold derived node metadata such as schema, guarantee, or timing.ir::aggregate_schemacomputes an aggregate's output schema from its input schema and reduction. For example, grouping byhostpreserves that column and adds the aggregate result column with its derived type. It does not compute aggregate values.ir::errordefinesQueryExprErrorfor schema/type derivation failures: invalid grouping-column indices, scalar signatures, sample columns, or an empty concatenation. My previous module description was too broad: structural and execution-timing validation have separate errors.
fc328b12 adds these explanations and examples to the module docs, the IR module index, and the acceptance document. This makes the boundaries explicit without introducing another abstraction or changing behavior. Workspace tests and clippy pass.
There was a problem hiding this comment.
Given that QueryExpr is removed, why QueryExprError still here?
There was a problem hiding this comment.
You are right: QueryExprError was a leftover name from the removed QueryExpr, not a separate concept that needed to survive the migration. Fixed in 150735f2: it is now ir::SchemaDerivationError, describing its actual role in deriving operator output schemas and scalar types. Updated signatures, imports, re-exports, error construction, and documentation throughout; no QueryExprError alias remains. Error variants and behavior are unchanged. Types/frontend-common tests, formatting, and workspace/all-target/all-feature clippy pass.
There was a problem hiding this comment.
How much stuffs is still pending on this thread?
|
The eight-PR split is implemented and pushed. Each PR targets the preceding branch, so its Files changed tab shows the incremental review scope.
The chain starts at Each layer was checked locally. The first six retain the existing production path while adding the new representation, lowerers, and compiler. #542 switches the compiled callers together and proves real batch summary sharing and execution of both result roots. #543 deletes the now-unused sources. Review #542 as the main integration change; some of its frontend/runtime diff is promotion of code already reviewed in earlier layers. Final validation: 1,588 workspace tests/doctests passed, formatting and workspace/all-target/all-feature Clippy passed, and viewer tests passed (24 passed; 6 Node-dependent tests skipped because Node is unavailable). GitHub CI is green for #535–#540; the latest layers are still running. Please review/merge in the table order. After each merge, retarget the next PR to |
|
The complete stack (#535, #536, #537, #539, #540, #541, #542, #543) is now rebased and pushed onto latest #535 also updates the two #511 design documents to distinguish Validation: all eight layers pass workspace/all-target/all-feature Clippy; final workspace tests/doctests: 1,590 passed; formatting passes; viewer tests: 25 passed, 6 Node-dependent skips. Remote CI has been retriggered. All eight branches were pushed atomically with explicit leases. No PRs were merged or closed. This supersedes the earlier base/tree-equality statements in the initial stack index: the stack now includes newer main changes and review fixes. |
|
Stack update: #535 has merged. The terminology-only Current review order: #549 → #536 → #537 → #539 → #540 → #541 → #542 → #543. All dependent branches have been restacked and pushed, preserving the merged schema / Also addressed the #536 constructor review: both operator categories now use |
Revised priority (2026-10-03)
This PR stays open as the #511 reference implementation until Phase A merges, then closes.
Status (2026-10-03, evening): Phase A is rebuilt and pushed as drafts: #567 → #560 → #537 → #539 → #540 → #541 → #542 → #543. Each tip passes clippy and the workspace tests, independently re-run. The design doc split out of #567 is #573, against main. Phase C is up as drafts on #543: #561 (Pass 1) → #575 (Stage 1 workload candidates +
stage_pipeline) → #576 (Stage 2/3) → #577 (Example 1 acceptance tests), plus the viewer #574. Phase B is complete as drafts: #578 (B1 types modules) → #579 (top-k readout consistency) → #581 (D1: stage-pipeline facade with tree-DP selection, MajorPass retired) → #582 (D2: Stage 1 without the cost model) → #583 (B2 logical-optimizer) → #584 (B3 physical-optimizer) → #585 (B4 plan-selection) → #586 (B4b facade in planner, asap-aware-mapping removed) → #587 (B5 executor). Follow-up: #580 (Pass 2 sharing and coverage parity).Phase A — finish #511 operator sharing
Phase B — #572 crate/module reorganization: five pure-move PRs (types modules → logical-optimizer → physical-optimizer → plan-selection → executor).
Phase C — #509 end to end: Example 1 through LogicalDAG → LogicalASAPDAG → PhysicalASAPDAG, with cost estimation, shown in a three-lane DAG viewer and checked by an e2e test. Example 4 (window materialization) comes next. Stage 1 starts from #561.
Parked: #552–#569 and #561 are now drafts, because they sit on the old
stack/528-legacy-physical-basechain and Phase B moves their files. Their content is re-scoped in Phase C. #551 is closed because it conflicts with the #572 accuracy decision.Implements the unified operator/scalar representation in #511 and resolves the main-branch merge.
Before this PR: ordinary and summary plans used separate node types and
KeepPreAsapwrappers; scalar computation could become operator nodes. For example,scalar(sum(up)) + 1required bridge nodes, and SQL scalar-subquery rewrites could lose empty-input/cardinality semantics.After this PR: both logical stages share
OperatorNodewithOperator::{NonASAP, ASAP}payloads, common typed schemas and explicit scalar producer edges.ScalarExprowns literals, arithmetic, evaluation context, conversions and subquery reads; scalar roots need no fabricated relation.KeepPreAsapis removed, and unchanged work is retained directly. Native compilation and flat wire export consume this representation.scalar(sum(up)) + 1.The acceptance matrix maps all four requested criteria and the proposal's language comparison tables to implementation/tests. It records parser/runtime gaps and compatibility changes (notably rejected ClickHouse stub signatures, native histograms and older persisted SUM payloads). This implements the proposal's representation scope; it does not claim full current DataFusion/Prometheus execution support. Generalizing analytical evidence to at-rest workloads was split into and merged separately as #532; this branch incorporates the updated main. Physical placement policy remains #520/#530.
Review follow-up: one multi-root workload DAG is explicit in
PlanOutput; replacements/CSE use sub-DAG terminology; shared operator properties, schema derivation and errors have separate modules.SketchStatisticand summary “evaluation” replace ambiguous query/readout names. Wire version 7 requires regenerated exports/native programs. A real SQL batch acceptance test selects a shared summary and executes both results; exact-count fixtures now use valid finalizers. Bulk retained evidence cannot hide any ASAP descendant (regression reproduced and fixed).Validation: after integrating #532, all 1,588 workspace tests pass with two existing ignores. Workspace/all-target/all-feature clippy passes with warnings denied.
cargo test --workspace --locked, fmt, clippy with warnings denied, external MetricsQL consumer, and DAG viewer Python tests. The viewer's six Node-dependent tests were skipped because Node.js is unavailable. The vendored MetricsQL baseline passes on Rust 1.99 (CI), verifying its existing 21 library and 3 doctest failures without changing fingerprints. Formatting and clippy also pass locally on Rust 1.99.Revised logical foundation (#509/#511)
The first five scopes are now ordered as:
Timing follows physical materialization; there is no intermediate timed-DAG stage. #541 and its existing descendants remain on the frozen stack/528-legacy-physical-base snapshot until the physical contract/materialization/selection scopes are reorganized. The first-five tip passes 1,961 workspace tests/doctests, formatting and full workspace Clippy.
Summary coverage prerequisite
#567 adds separate summary observation coverage metadata identified as missing in #535. Updated logical review order: #567 → #560 → #537 → #539 → #540 → #561. #560 now requires known disjoint coverage and derives its union; #537 preserves coverage in CSE and logical transport.