Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion docs/design_docs/concepts/post-asap-ir.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,7 +49,8 @@ summary family supports incremental maintenance.
and selects the joined rows. Completeness evidence belongs to pruning, not ranking.

A `SummaryNode` carries its expression, schema and optional result guarantee.
State and query values have different contracts. Exact operations over
State and query values have different contracts; see
[Schema and physical data for ASAP primitives](../proposals/asap-primitive-schema.md). Exact operations over
approximate readouts still require composed accuracy guarantees. See the
[accuracy implementation companion](../../develop_docs/end-to-end-accuracy-guarantees.md)
and [physical-plan integration](../architecture/physical-plan-integration.md)
Expand Down
28 changes: 8 additions & 20 deletions docs/design_docs/physical-planning-and-deployment.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,26 +117,14 @@ operators from logical candidates has not completed this integration.

### Input semantics and summary semantics

`source`, `filter`, `grouping` and `window` describe input-data semantics:
where records originate, which records qualify, how they are grouped and which
time interval applies. They are not a complete description of arbitrary summary
computation. In particular, the same four fields can summarize different value
expressions or produce different states.

| Concern | Required semantic information |
| --- | --- |
| Input computation | Source identities and schemas, filters, joins/transforms and their order, or a reference to the canonical input sub-DAG |
| Values and grouping | Value expressions, item identities and weights where applicable, group keys and types, and operation-defined null/duplicate handling |
| Time | Time column and interpretation, interval bounds, evaluation alignment, and distinction between query range and maintained panes |
| Summary computation | Exact operation or sketch family, algorithm and parameters, and supported build/merge behavior |
| Output | State versus finalized value, output schema/type, and readout parameters when part of the output computation |

For example, KLL over `latency_seconds` and KLL over `log(latency_seconds)` differ
even with identical source, filter, grouping and window. Likewise, weighted
frequency state needs both item and weight expressions. More complex inputs
must retain their computation DAG; four descriptive fields cannot replace it.

The canonical selected computation is authoritative. These categories describe
`source`, `filter`, `grouping` and `window` describe input-data semantics but not
a complete summary computation: the same four fields can summarize different
value expressions or produce different states. The semantic information a summary
depends on, and where the IR records each part (field type, producing operator,
or coverage), is specified in
[Schema and physical data for ASAP primitives](proposals/asap-primitive-schema.md#4-proposed-node-field-design).

The canonical selected computation is authoritative. Those categories describe
what must be preserved, not a new flat IR or a second expression language.
Operator-defined behavior should be referenced through its canonical contract,
not independently configured in deployment metadata. Unsupported or unresolved
Expand Down
1 change: 1 addition & 0 deletions docs/design_docs/proposals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,3 +11,4 @@ extensions. A design document is not a promise of downstream runtime support.
- [Operator sharing](operator-sharing.md)
- [Decoupling operators from scalar expressions](decoupling_op_and_expr.md)
- [ASAPPlanner layering](planner-layering.md)
- [Schema and physical data for ASAP primitives](asap-primitive-schema.md)
585 changes: 585 additions & 0 deletions docs/design_docs/proposals/asap-primitive-schema.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion docs/design_docs/proposals/decoupling_op_and_expr.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,7 @@ operator inputs and scalar query-result references use `Rc<OperatorNode>`.
read by expressions. Keep `ScalarExpr::Column(ColumnId)`: the ID selects a field
for type checking and the corresponding input value for evaluation, independently
of the executor's row/column storage layout. See the
[fields versus column references contract](operator-sharing.md#21-one-schema-model-for-values-and-state).
[fields versus column references contract](asap-primitive-schema.md#3-proposed-schema-design).

Names are resolved to `ColumnId` before constructing these nodes. Parsing and
unresolved `ColumnRef` handling remain frontend concerns; no alternative generic
Expand Down
156 changes: 7 additions & 149 deletions docs/design_docs/proposals/operator-sharing.md
Original file line number Diff line number Diff line change
Expand Up @@ -365,155 +365,13 @@ caching or mutation mechanism.

### 2.1 One schema model for values and state

Use one `Schema` for operator outputs before and after optimization. Rename today's
`SummaryFamilyType` to `FieldDataType`: it types every field, and `Plain` is not a summary
family. Rename `Column` to `Field` and `Schema.columns` to `Schema.fields`: the struct
describes a column and holds none of its data. Retain the current `Schema` metadata.
The following is the proposed resolved interface; it is not the current Rust definition.

```rust
struct Field {
name: String,
dtype: FieldDataType,
nullable: bool,
table: Option<String>,
}

struct Schema {
fields: Vec<Field>,
time_index: Option<ColumnId>,
unique_keys: Vec<Vec<ColumnId>>,
closed: bool,
}

// Today's `SummaryFamilyType`, renamed; variants and payloads unchanged.
enum FieldDataType {
Plain(DataType),
ExactAggregate(ExactKind, ExactParams),
Sketch(SketchKind, GroupingStrategy),
Sample(SamplingKind, SamplingParams),
Wavelet(WaveletKind, WaveletParams),
StatModel(StatModelKind, StatModelParams),
}

// Proposed derived output classification, separate from column types.
enum OperatorResultKind {
Relation,
InstantVector,
RangeVector,
State,
}

impl Operator {
fn output_schema(&self) -> Result<Schema, QueryExprError>;
fn output_kind(&self) -> Result<OperatorResultKind, QueryExprError>;
fn validate_inputs(&self) -> Result<(), QueryExprError>;
}

impl OperatorNode {
fn validate_structure(&self) -> Result<(), QueryExprError>;
fn validate_execution_timing(&self) -> Result<(), QueryExprError>;
}

impl ScalarExpr {
fn scalar_type(&self, input: &Schema) -> Result<(DataType, bool), QueryExprError>;
}
```

**Fields versus column references.** These names describe different roles, not
competing representations of the same object:

| Name | Role | Holds runtime values? |
|---|---|---|
| `Schema` | Ordered `Field` metadata, plus key/time/closedness information | No |
| `Field` | Name, type, nullability and optional qualifier for one output column | No |
| `ColumnRef` | Unresolved logical reference: `Named`, `Qualified`, `SampleValue`, or `Wildcard` | No |
| `ColumnId = usize` | Resolved column position in a particular input/output schema | No |
| Runtime batch | Values conforming to a schema; storage layout is executor-specific | Yes |

Keep `ColumnRef`, `ColumnId`, and `ScalarExpr::Column(ColumnId)`. Renaming the
metadata struct `Column` to `Field` does not rename column references to field
references. The same position identifies metadata during planning and values
during execution; it is not a stable field identity across projections or joins.
Schema `unique_keys` and `time_index` also use these column positions.

For example, resolving `t.bytes` to position `1` produces `ColumnId = 1`.
`schema.fields[1]` supplies its type and nullability; evaluating
`ScalarExpr::Column(1)` reads the corresponding value. The native executor
currently reads `row[1]` from `Batch { schema, rows: Vec<Vec<Value>> }`. A columnar
executor would select array `1` instead. No physical `Column` container is
introduced by the metadata rename, and the old metadata `Column` struct is not
retained as a second type.

**Relationship to current types.** `Field` is today's pre-ASAP `Column` with `dtype`
widened from `DataType` to `FieldDataType`. `FieldDataType` is today's `SummaryFamilyType`
under a name that also fits its `Plain` case. The proposed common `Schema` replaces
the separate operator-edge roles of pre-ASAP `Schema` and post-ASAP `SummarySchema` /
`SummaryField`; it does not rename `DataType`. A pre-ASAP value column becomes
`Plain(dtype)`.
Frontend validation permits only ordinary value columns, preserving the current
pre-ASAP restriction even though the common schema can also express state.

| Field | Meaning and requirement |
|---|---|
| `fields` | Ordered named fields. `Plain(DataType)` is a readable value; other variants retain the identity and parameters of summary or exact-accumulator state. |
| `Field.nullable`, `Field.table` | Preserve SQL nullability and qualified column resolution. |
| `time_index` | Identifies the time column when present; it does not by itself distinguish an instant vector from a range vector. |
| `unique_keys` | Proven column combinations identifying rows; an empty list asserts no known key. Recompute these proofs when a rewrite changes identity. |
| `closed` | Whether `fields` completely describes the output. An open PromQL schema must retain unlisted labels through the existing complete-series-identity contract. |

`OperatorResultKind` is derived from the operation and its inputs and retained as
`OperatorNode.result_kind`. `State` describes an output carrying unfinalized state; its
schema may also contain ordinary grouping keys. `SummaryEstimate`,
`FinalizeExactAccumulator` and other readouts derive the appropriate relation or
vector kind from their operation and input context. Matching numeric columns do
not make those kinds interchangeable.

**Interface contracts.** `Operator::output_schema` and `output_kind` derive output
metadata from the payload and validated inputs. `validate_inputs` checks local
producer/consumer compatibility, such as vector inputs for `BinaryOp` or the
required state family for a summary readout. Scalar typing checks the input-kind
contract of `PromqlScalarFromVector` and other scalar plan reads.

| Validation entry | Scope and stage |
|---|---|
| `OperatorNode::validate_structure()` | Walks the reachable operator DAG, including scalar plan references; checks input contracts, scalar typing and agreement between retained and derived output metadata. Valid for logical and physical plans; permits `timing = None`. |
| `OperatorNode::validate_execution_timing()` | Includes structural validation, then requires assigned timing on every executable operator and checks phase dependencies. Used for executable physical candidates. |
| Existing planner assessment and selection (#509) | Establishes guarantees using the existing accuracy models and checks them against request requirements and deployment capabilities. Neither node method re-proves a guarantee or decides request feasibility. |

The two node methods need only the DAG and its annotations. Request requirements
and deployment models remain inputs to the existing planning/selection workflow,
not implicit globals of `validate_structure`. Passing the timing check alone does
not establish that a physical candidate satisfies the query's accuracy requirement.

`Scan.schema` declares the source columns; `Values.schema` declares the constructed
row shape. `OperatorNode.schema` is the derived output for any operation. A scan's
predicates cannot change its declared output columns; a Values row must match the
declared arity, types and nullability. These leaf outputs retain the declaration's
column layout and time/identity information, with only justified metadata changes.
The declaration and derived output therefore have distinct roles, and structural
validation rejects disagreement rather than trusting two independent schemas.

`scalar_type` keeps the existing method name and `(DataType, nullable)` result.
Its `input` is the applicable column scope: the child schema for a projection,
both input schemas for a join predicate, or aggregate outputs for `HAVING`.
Explicit subquery/conversion expressions validate their referenced producer using
the contracts above. Numeric expressions cannot consume state columns as numbers.
A standalone scalar expression is checked with an empty column scope and needs no fabricated
relation output schema. `QueryExprError` retains the existing error-type name;
result-kind, state-family, schema and execution-phase mismatches require
corresponding validation errors.

For example, a KLL build outputs `State` with a
`Sketch(SketchKind, GroupingStrategy)` column identifying KLL and its parameters.
Its p99 readout outputs an ordinary `Plain(Float64)` column in the appropriate
relation/vector schema. A numeric predicate can use that readout, but not the KLL
state. Exact accumulator state similarly requires `FinalizeExactAccumulator`.
An ordinary operator may pass state through only where its input/output contract
permits it. A bare-column projection can preserve the field's `FieldDataType`
directly during `output_schema` derivation; `scalar_type` applies when that column
is used as a scalar value and rejects state. Copying a state column does not turn
it into a readable scalar.
Every operator output, before and after optimization, uses one `Schema` whose
`Field`s are typed by `FieldDataType`: `Plain(DataType)` for a readable value, or
the family, algorithm and parameters of summary or exact-accumulator state.
`OperatorResultKind` marks state outputs, and state becomes a value only through
an explicit readout. The schema model, `ColumnRef` versus `ColumnId`, the
validation entry points and the readout boundary are specified in
[Schema and physical data for ASAP primitives](asap-primitive-schema.md).

### 2.2 Preserve existing accuracy semantics

Expand Down
3 changes: 2 additions & 1 deletion docs/design_docs/proposals/univmon-frequency-summary.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,8 @@ cardinality alternatives, and exact count remains the cheaper first count
candidate.

All four readouts have the same unit-weight update, input sub-DAG, grouping,
window, parameter identity and state schema. Existing post-ASAP structural
window, parameter identity and state schema
([ASAP primitive schema](asap-primitive-schema.md)). Existing post-ASAP structural
sharing can therefore intern their state producer while preserving distinct
readout nodes. Sharing is only legal within the same execution/data scope.
Precompute placement, SummaryCatalog installation, retention, and runtime
Expand Down
24 changes: 6 additions & 18 deletions docs/develop_docs/asap-aware-mapping-contracts.md
Original file line number Diff line number Diff line change
Expand Up @@ -325,24 +325,12 @@ backend inspection. Automatic selection skips those unproven ratios. Use

### Family, category, algorithm, and parameters

Sketches separate their query category from the concrete algorithm and its parameters:
A summary's identity has four levels: family (`FieldDataType` variant), sketch
category (`SketchCategory`), algorithm (`SketchAlgorithm`), and the validated
committed choice (`SketchKind`). The levels and their validation are specified in
[Schema and physical data for ASAP primitives](../design_docs/proposals/asap-primitive-schema.md#3-proposed-schema-design).

| Level | Type | Example |
| --- | --- | --- |
| **family** | `SummaryFamilyType` | `Sketch`, `Sample`, `Wavelet`, `StatModel`, `ExactAggregate` |
| **category** | `SketchCategory` | `Quantile`, `Cardinality`, `Frequency`, `TopK` |
| **algorithm** | `SketchAlgorithm` | `Kll` / `DDSketch` (both quantile); `Hll` (HyperLogLog) / `Theta` / `Kmv` (K-Minimum Values), all cardinality |
| **committed choice** | `SketchKind` | one validated category + algorithm + parameter combination |

A `SketchKind` is a validated committed choice. Its public constructor,
`SketchKind::new(algorithm, params)`, verifies that the parameter variant belongs
to the selected algorithm and classifies the pair into its category. The public
`.category()`, `.algorithm()`, and `.params()` accessors expose the committed
values without permitting an invalid combination.

Where this matters in practice: `CostModel::rank_candidates`, `CostModel::size_params`, and `SketchAlgorithmStrategy::replacements` operate at the **algorithm** level. `summary_candidates(intent)` returns a list of `SketchAlgorithm`s (`[Kll, DDSketch]` for a `Quantile` intent), never a bare `SketchKind` with nothing chosen underneath it. `SketchKind` appears after an algorithm has been selected and sized—on `Realization::Sketch(SketchKind)` and `SummaryFamilyType::Sketch(SketchKind)`.

`Sample`, `Wavelet`, and `StatModel` each use a flat `(Kind, Params)` pair. `Sketch` needs the additional algorithm level because multiple algorithms can serve the same purpose—for example, KLL and DDSketch both answer quantile queries.
Where this matters in practice: `CostModel::rank_candidates`, `CostModel::size_params`, and `SketchAlgorithmStrategy::replacements` operate at the **algorithm** level. `summary_candidates(intent)` returns a list of `SketchAlgorithm`s (`[Kll, DDSketch]` for a `Quantile` intent), never a bare `SketchKind` with nothing chosen underneath it. `SketchKind` appears after an algorithm has been selected and sized—on `Realization::Sketch(SketchKind)` and `FieldDataType::Sketch(SketchKind, GroupingStrategy)`.

---

Expand Down Expand Up @@ -378,7 +366,7 @@ The crate provides no default `Matcher` implementation because the answer depend

Concretely, `explanation.rs` reports three candidate kinds from each `TargetSubDAGCandidates`:

- `ExplanationKind::SketchApproximation` — the set contains a `Replacement::Summary` that realizes `SummaryFamilyType::Sketch(..)`, not just an exact/pass-through candidate.
- `ExplanationKind::SketchApproximation` — the set contains a `Replacement::Summary` that realizes `FieldDataType::Sketch(..)`, not just an exact/pass-through candidate.
- `ExplanationKind::CommonSubexpressionReuse` — `consumer_count >= 2` and the set contains `SharedSubDAGStrategy`'s "build once and share" candidate (the `Replacement::Rewrite` whose `Rc` is the set's `target`).

- `ExplanationKind::ExactComposition` — the candidate set contains an exact operation
Expand Down
24 changes: 4 additions & 20 deletions docs/develop_docs/pre-asap-ir.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,26 +17,10 @@ The pre-ASAP IR is defined using the `QueryExpr` enum. We discuss some of import

## Fields and column references

`Schema` owns `Field` metadata: name, type, nullability, and an optional table
qualifier. A `Field` contains no runtime values. The former schema `Column`
struct served this same metadata role; it was renamed to `Field`, not retained
as a second data container.

`ColumnRef` is an unresolved logical reference (`Named`, `Qualified`,
`SampleValue`, or `Wildcard`). Resolution binds a reference to `ColumnId`, a
`usize` position within a particular schema. `QueryExpr::Column(ColumnId)`
reads that column; the same position indexes `Schema::fields` for type checking
and a runtime row for its value. Group keys, unique keys, and `time_index` also
use these column positions. They are not stable identities across projections
or joins, so the positional reference remains `ColumnId`, not `FieldId`.

The native runtime names shared ownership `SchemaRef = Arc<Schema>` and stores
`Batch { schema: SchemaRef, rows: Vec<Vec<Value>> }`. `Schema` is the same metadata
model during planning and execution; the `Ref` suffix only distinguishes ownership.
It has no physical `Column`/array container. A column reference expresses what
to read independently of whether an executor stores its data as rows or arrays.
For example, resolving `t.bytes` to `ColumnId = 1` obtains its type from
`schema.fields[1]`; native execution reads `row[1]`.
`Schema` holds `Field` metadata (name, type, nullability, qualifier) and no
values; an unresolved `ColumnRef` resolves to a positional `ColumnId` within one
schema. The design, including how the same position selects a runtime value, is
in [Schema and physical data for ASAP primitives](../design_docs/proposals/asap-primitive-schema.md#3-proposed-schema-design).

## Node index

Expand Down
Loading