Skip to content

Design a universal ASAP primitive table: schema, physical data, and summary state #545

Description

@zzylol

We need a single, complete design for a universal ASAP primitive table that connects schema metadata with the physical data and summary state it describes.

This follows the design discussion on #535 and the common schema model introduced in #511/#535. Existing documents describe parts of the contract, but do not provide one complete table/container design covering ordinary data and ASAP primitive state.

Metadata versus physical data

Use the following distinction as the starting point:

Concept What it is Where it lives
Schema A collection of Fields. Pure metadata defining names, types, nullability, and other structural properties; holds no data values. Schema([Field, Field, ...])
Field The structural description of one column, e.g. name: user_id, type: Int64, nullable: false. Holds no data values. Inside a Schema
Table / RecordBatch A concrete container of columns holding actual values or summary states that conform to a schema. Table([Column, Column, ...])
Column / Array The physical vector of values, e.g. [101, 102, 103], or a defined representation of summary-state entries. Inside a Table / RecordBatch

The design must distinguish a physical Column/Array from a logical column reference. ColumnRef and resolved ColumnId identify inputs to expressions; they are not data containers. SchemaRef = Arc<Schema> changes ownership only, not the schema model.

Requested design

Produce one authoritative design document with concrete type/API sketches, invariants, and examples covering:

  1. Common schema contract. Define Schema, Field, and field types for ordinary values, exact accumulators, sketches, samples, and other supported ASAP primitive families. Explain where family parameters, nullability, grouping keys, time coordinates, and identity metadata belong.
  2. Physical table/container contract. Define Table versus RecordBatch, Column/Array, row counts, field-to-column correspondence, ownership, and validation. Specify how mixed ordinary-value and summary-state columns are represented. Explain whether columnar storage is required or whether the contract permits other layouts; map the existing row-oriented native Batch to the chosen model.
  3. Summary payload contract. Explain how a physical entry holds an exact accumulator, sketch, or sample; how the schema describes it; and how implementation identity, parameters, compatibility, and empty/NULL state are represented. Distinguish summary state from a finalized scalar/vector result.
  4. Operations and compatibility. Define which table/column operations apply to values and state, including projection, filtering, grouping, build/update, merge, supported deletion/subtraction, and summary evaluation. State when ordinary scalar operations must reject summary-state inputs.
  5. Physical layout and persistence boundaries. Separate the logical field type, in-memory representation, and serialized format. Specify what is universal versus executor/deployment-specific, including codecs/versioning, materialized state, and population/window coordinates. Explain how the existing $population, $window_end, value adapter fits without assuming it is the universal layout.
  6. Ownership and migration. Identify which contracts belong to ASAPPlanner/shared runtime versus deployment/storage integrations. Show how current Schema, SchemaRef, Batch, and summary payload APIs map to the proposed design, with compatibility and migration requirements.

Required worked examples

  • An ordinary data table, e.g. user_id: Int64 with values [101, 102, 103].
  • A grouped exact-summary table, showing grouping keys and SUM/COUNT accumulator state.
  • A grouped/windowed sketch table, showing population identity, window coordinates, and KLL or another concrete sketch state.
  • Evaluation of that summary table into an ordinary result table, including the input/output schemas and actual data/state containers.

For each example, show metadata and physical contents separately, explain how they correspond, and include empty/NULL behavior where applicable.

Acceptance criteria

  • One linked design document defines the complete metadata → container → summary payload → evaluation contract.
  • The four concepts in the table above have unambiguous names and responsibilities; metadata structs do not double as physical data containers.
  • Worked examples demonstrate ordinary, exact-summary, and approximate-summary data with explicit schema/layout correspondence.
  • Unsupported operations and compatibility checks are explicit rather than inferred from matching field names or numeric output types.
  • Existing design documents link to this contract and any conflicting ownership/layout statements are reconciled.
  • The document identifies implementation gaps and follow-up work. This issue requests the design; it does not assume the complete container abstraction is already implemented.

Existing material to consolidate

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions